AI

Your coding agent should call a QA system, not become one

Kirk Nathanson
October 5, 2026

Give Claude Code or Codex a browser and a feature to test and it'll do a respectable job. It'll click through the flow, read the DOM, write a Playwright test, run it, and fix the selector when it guesses wrong. Given how good the frontier models have gotten, a lot of people conclude you don’t need anything else for QA.

For one PR on one laptop, they're close to right.

The trouble starts when you multiply the demand across 10s or 100s of developers and agents shipping a few hundred PRs a week. At that scale, the team can't review every test those agents propose, can't run the suite on anyone's laptop, and can't have someone watching each agent while it works.

At that volume, a general-purpose coding agent runs into two limits. The first is what it knows about QA. QA Wolf's agents run on the same frontier models yours do, but they work with skills, rules, and evals built specifically for QA, which teach them what's worth testing, how to investigate a failure, and when coverage is actually complete. A model can't give itself any of that, and it's why a QA Wolf agent makes better QA decisions than the same model working alone.

The second is everything that isn't AI at all: a fleet of machines to run every test on every change, deterministic Playwright code so a failed run means something, and the plumbing that lets dozens of agents work at once without anyone steering. Making your coding agent your QA system means rebuilding both.

Your coding agent has to learn what's worth testing

A coding agent that just built a feature can write a decent test for it. The problem at volume is that every agent on every PR is doing the same thing, each in its own session with no view of the others, and all of them want to commit their tests to the same shared suite. Two agents touching checkout in the same week will each add a checkout test without knowing the other one did.

Somebody has to decide which of those tests belong, and nobody reads 300 proposed tests a week closely enough to catch the ones that check the wrong thing, duplicate a test that already exists, or protect something nobody cares about. So the suite fills up with tests that pass and prove very little, and every one of them still costs runtime and maintenance.

We saw how easily that happens while building Mapping AI, the QA Wolf agent that explores an application and outlines which workflows need tests. Pointed at an app with no constraints, a frontier model wrote down whatever it saw: "Display Search Dropdown," "View Hero Banner," every route to the checkout page and nothing checking that the cart made it there. It called products fully mapped with whole sections unvisited, stalled on a Google sign-in screen it didn't need to clear, and once it could type URLs, it started inventing addresses and outlining tests for 404 pages.

Those aren't reasoning failures, exactly. The model had no theory of what's worth testing, so we had to give it one: rules for what doesn't count as a test (arriving at a page, seeing that content exists, etc.), a two-pass exploration modeled on how a person tours an unfamiliar app, gap analysis that steers it toward the least-covered areas, and an eval gym that scores every change against the same fixed scenarios. Together those raised the flows found in a 15-minute window by about 50% on our standard scenario (95 to 143) and roughly doubled them on one benchmark site. We wrote up how we built it.

The map also has to agree with the suite you already have. When Mapping AI remaps an app, plain code, not a model, decides whether a new flow matches existing coverage, and flows the team already rejected stay rejected along with the reason. A coding agent working from one diff doesn't know the suite already has two versions of this test, or that someone threw out the same idea last month. You could keep that in a memory file, but then you're maintaining the coverage map by hand.

There's a simpler reason to keep the test-writer separate, too. Teams don't let authors approve their own PRs, and it isn't because authors are bad engineers. When the agent that built the feature also decides what the test checks, a misread ticket ends up in both.

Your coding agent needs a fleet of machines to run every test on every PR

A test that passed once on a developer's machine hasn't proven much. A suite earns its keep by running in full on every PR, against an environment that resembles production, fast enough that nobody's tempted to skip it.

The median passing QA Wolf test takes 99 seconds, and a typical customer has around 250 user flows (many have 4–5x that), so running them sequentially or in shards would take hours. QA Wolf runs them in parallel across isolated containers, and the median suite finishes in about 10 minutes.

Multiply that by every PR a team merges in a week, then add mobile (physical devices and emulators), multi-user flows that need two sessions in the same run, email and SMS verification, seeded test data, etc. That's a fleet of machines somebody has to build, pay for, and keep running, and it doesn't fit on a laptop.

What runs on that fleet matters as much as the fleet. A coding agent can check a feature by clicking through it, and for a quick look that's fine. It doesn't work as a release gate, because a model deciding what to click on each run can take a different path through the same build, and then a changed result could mean the product broke or the agent wandered off. QA Wolf's agents use the browser to write the test, then hand off to deterministic Playwright code that runs the same steps every time and that anyone on the team can read.

It would have to work with nobody watching

Most failures don't need a person to fix them, but something has to look at every one. When a PR renames "Place order" to "Complete purchase" and the checkout test breaks, the fix is a few minutes of work: read the PR, grep the tests for the old label, look at the new button in the app, update the selector, rerun, commit. Your coding agent can do that, and so can ours.

Selector changes cause fewer than 28% of failures, though. The rest are timing problems, runtime errors, bad data, flaky environments, and the occasional real bug, and diagnosing them takes the run's logs, trace, video, and network activity. Suites also decay as the product changes, and half of an automated suite goes obsolete within six months on a regular release schedule. At hundreds of PRs a week, that's a steady stream of investigations and repairs that has to get done whether or not anybody opens a terminal.

That's the job QA Wolf's agents do. A request like "fix every failing test from last night's run" gets split into independent pieces and handed to as many as 50 agents at once, each with its own machine, browser, and branch so they can't step on each other.

Fifty agents at once means nobody is watching any of them, and a lot of what keeps a coding agent safe on a laptop is the person sitting in front of it. You notice when it crashes, decide whether to rerun it, and stop the command that looks wrong. Take the person away and each of those has to be handled some other way.

QA Wolf's system prevents three problems your local coding agent would run into without someone at the keyboard: duplicate bug reports, commands that reach things they shouldn't, and work lost to a crash. An agent that files a bug and then crashes would file it again on a retry, so each unattended job gets one attempt and fails where the team can see it. The agents read PR descriptions anyone can write while holding test credentials, so every command runs with hard limits on what it can reach, and we don't rely on a model to spot the hostile ones. Work in progress is saved off the machine every few seconds, so a crash doesn't take it with it.

What calling QA Wolf looks like

This is what the QA Wolf MCP server is for. Connect QA Wolf to Claude Code, Codex, or any other MCP-compatible agent, and your coding agent can hand over what it knows about the change (the diff, the ticket, what the feature is supposed to do) and QA Wolf can take the hard work of actually testing.

A QA Wolf agent uses that context and specialized skills to find the coverage that already exists, add or update tests where it's missing, run them on QA Wolf's infrastructure, investigate failures, and file bugs with repro steps, a video replay, the Playwright trace, and network logs attached. The results go back to your coding agent, which still has the codebase and the task in context, so it can fix the code and send the new version through again. When a task needs local files or local execution, your agent can switch to the QA Wolf CLI, which has the same capabilities as the MCP server.

Your coding agent keeps doing what it's good at, which is understanding and changing your code. QA Wolf decides what's worth testing, runs it on every change, works out what failed, and keeps the suite current while your team ships.

It works with a free QA Wolf account. Connect your agent, point it at your staging environment, and ask it to test your next PR.

Connect QA Wolf to your coding agent

Try the AI testing platform that makes QA 12x faster.