
Let me be honest with you. When I started in QA two decades ago, the job was already misunderstood. Most people thought we were just there to click buttons and write bug reports. Some still do. I’ve sat through enough sprint planning meetings where QA was an afterthought, where the assumption was that if developers write enough unit tests, we’d be redundant by next quarter.
It never happened. And now, with AI-generated code flooding every codebase on the planet, I’d argue we’re entering the era where QA finally gets its due reckoning, not as a safety net, but as a critical discipline that stands between functional software and catastrophic trust failures.
Let me tell you why I believe that, and more importantly, what I think we need to do about it.
The problem with AI-generated code isn’t the bugs. It’s the confidence.
I’ve spent the better part of my career testing enterprise applications sprawling, interconnected systems where a single misbehaving API endpoint can cascade into a financial discrepancy, a compliance violation, or worse, a patient getting the wrong information. These aren’t abstract risks. I’ve seen them up close.
When I started testing APIs manually before contract testing was a thing, before Postman had a UI, half of us recognized that you had to understand the business domain deeply. You had to ask: what does this data actually mean to the person on the other end? That question hasn’t changed. But the source of the code you’re testing has changed dramatically.
AI can produce a working implementation in seconds. It can write an endpoint that passes every unit test the developer asks for. What it struggles with, and I mean genuinely struggles, is knowing what it doesn’t know. It doesn’t know your legacy business rules buried in a 12-year-old stored procedure. It doesn’t know that there is one edge case in your tax calculation module that only surfaces during a specific date range in a leap year. It doesn’t know what “correct” means to your domain.
“Confidence without understanding is the most dangerous thing in software. AI can produce the first without the second. That’s exactly where we come in.”
We are the ones who know what “correct” means.
Here’s a question I want you to sit with for a moment: when was the last time you caught something that no one else was even looking for?
Not a regression. Not a broken test. Something subtler, a result that was technically valid, but somehow wrong. A number that passed validation but didn’t match reality. A flow that completed without error but left the user in a state that made no business sense. You felt it before you could explain it. And then you dug in, and you were right.
That instinct has a name. It’s domain knowledge. And it’s one of the things no AI model has, at least not yours.
I’ll give you some examples. Not hypothetical ones, the kind of things that happen in real projects, the kind that probably sound familiar.
Scenario 1 — The financial edge case
A refund API returns an HTTP 200 status code and a success payload. Every test passes. But you know, because you’ve been around long enough, that this particular transaction type shouldn’t be refundable after 30 days without a manual override flag being set. The API didn’t error. The model that generated it didn’t know the rule existed. You did. You checked. The flag was missing. A real customer would have been affected.
Scenario 2 — The date nobody told the system about
A calculation looks right in every test environment. As a reminder from two years ago, this module behaves differently on the last day of a financial quarter. You run it against that specific date. It breaks. The developer had no idea this rule even existed. It was buried in a comment in a stored procedure that hasn’t been reviewed in a while. You read it. Once. Two years ago. That’s what saved the release.
Scenario 3 — The integration that “worked” but lied
Two systems are talking to each other. The handshake completes, the data transfers, no errors in the logs. But you know the downstream system interprets a null value differently from an empty string, and the upstream is sending one when it should be sending the other. To any automated check, this is a pass. To you, it’s a ticking clock. Because you’ve seen what happens three steps down the pipeline when that distinction is ignored.
These aren’t examples of QA heroics. They’re examples of what happens when someone with genuine context pays attention. In that context, the mental map of how your systems actually behave, the institutional memory of past incidents, and the understanding of what the data means to a real human being accumulate over years. It doesn’t live in a repository. It doesn’t live in the documentation because most of the time it hasn’t been documented. It lives with people who’ve been paying attention.
Does any of this sound familiar?
You’ve caught a bug that wasn’t in any spec because you remembered an incident from a previous project.
You’ve questioned a “passing” result because something felt off, even before you could articulate why.
You’ve been the only person in the room who knew why a certain edge case mattered because you’d seen it fail before.
You’ve tested a feature built exactly to spec, yet still knew it was wrong for the user.
If you recognized yourself in any of those, you already understand the argument I’m making. That knowledge, the kind that doesn’t fit in a Jira ticket or a test case, is exactly what an AI model cannot replicate. It was trained on general internet data, not your company’s incident history. It doesn’t have visibility into what your users have complained about. It doesn’t remember the meeting where someone said, “We’ll handle that later,” and later never came.
“The most valuable thing a QA engineer carries isn’t a test suite. It’s a calibrated sense of what ‘right’ looks like built from years of watching software fail in ways nobody predicted.”
Automation is not enough. It never was.
I’ve built automation suites. I know what they can do and, more importantly, what they can’t. A regression suite tells you whether the things you already knew about still work. It cannot tell you whether the new thing does what it was meant to do in the first place. That distinction is enormous.
With AI-generated code, the risk of automation-only testing is that you’re running automated checks against software that was written by a model that also doesn’t know what “correct” means. You can end up in a loop where the code is wrong, the tests the model wrote are wrong in the same direction, and everything passes. Perfectly. Until a real user hits a real scenario.
This is where exploratory testing, something I’ve written about here before, becomes even more critical. We need humans who can think laterally, who can probe AI-generated features with curiosity rather than a checklist. Who can say: “This output looks plausible, but does it actually make sense?”
So — what do we need to do?
I’m not going to pretend I have all the answers. But I’ve been in this industry long enough to recognize a turning point when I see one. Here’s what I genuinely believe we need to start doing now:
Get closer to the AI tools your teams are using. Don’t be the person in the room who hasn’t touched the tools the developers are using. You don’t have to love them. Please make sure you understand their failure modes. How does your LLM handle ambiguous requirements? What does it do when the prompt is underspecified? These are testing questions. We know how to ask them.
Double down on domain knowledge. The more AI accelerates code production, the more valuable it becomes to be the person who deeply understands the business domain. Be that person. Read the specs that no one else reads. Talk to the people who actually use the software. Know the data flows end-to-end. This is not glamorous work. It’s foundational work.
Advocate for QA earlier in the AI loop. If your team is using AI to generate code, QA should have a voice in how requirements are structured before that code is ever generated. Garbage in, garbage out. We understand this principle in data, but it applies equally to the prompts feeding an LLM. A well-specified requirement with explicit edge cases produces better-generated code. That’s a QA contribution, even before a line is written.
Prepare to test AI outputs as a new class of artifact. This is a frontier, and I’ll be honest, most of us are still figuring it out. But testing a model’s output isn’t just functional testing. It involves evaluating consistency, reasoning quality, hallucination risk, and behavior under edge inputs. Some of these maps onto skills we already have. Some of it is genuinely new. Stay curious.
“The industry is producing code faster than it can be reasoned about. Someone has to do the reasoning. That’s us. That’s always been us.”
A personal note to close.
Twenty years ago, I chose QA not because it was the obvious path, but because I was genuinely fascinated by the question: how do you know that something works? That question has driven me across every project, every domain, every stack I’ve worked with. It turns out it’s a harder question than most people appreciate.
Artificial intelligence doesn’t make that question easier. If anything, it makes it more urgent. Because now, the things we’re building are more complex, built faster, and in many cases built by systems that cannot ask themselves whether what they’ve produced is actually right.
We can. We should. And I think, if we’re ready for it, this is the moment QA stops being misunderstood — and starts being recognized as one of the most important professions in the technology industry.
The AI era doesn’t shrink us. It needs us. More than ever.