Research & technology

QA Wolf vs Claude Code vs Codex

QA Wolf
October 6, 2026

We ran QA Wolf, Claude Code (Opus 5.5), and Codex (gpt-6.1-sol) through the same 1,072 evals covering test planning, automation, and failure investigation to benchmark the QA Wolf agent against the frontier models using the Playwright MCP. All three ran at medium reasoning effort.

QA Wolf vs coding agents on 1,072 evals covering test planning, automation, and failure investigation
  • QA Wolfgpt-6.1-sol95.5%
  • Claude CodeOpus 5.584.9%−10.6 points behind QA Wolf
  • Codexgpt-6.1-sol83.2%−12.3 points behind QA Wolf

How QA Wolf evals are developed

Evals are developed from real-world test cases and UI patterns encountered by the QA Wolf agent as it plans, automates, and reviews end-to-end Playwright and Appium tests.

When QA Wolf gets something wrong in the course of its work, we rebuild the problem as an eval and then update the agent so that it can successfully navigate that same situation in the future.

How new evals are identified

  • Detected automatically. Performance monitoring agents detect when a testing agent is hitting a wall, and when users are getting frustrated by the performance of the platform.
  • Reported by people. In-house engineers and customers report situations the testing agent doesn't handle well.
  • From scratch. When releasing new features we build starter evals that allow us to begin measurement from the start.

Building the eval

If the eval was inspired by a failure of the agent to accomplish a task, the first step is to trace the agent's activity back to where it departed from what the user wanted, not the last error it hit.

We then recreate that situation in a test app we built for gym evals, with checks that describe the correct outcome. The agent must fail 3 out of 5 gym runs for us to accept that the eval accurately reproduces the agent's limitation (a fix to the QA Wolf testing agent is accepted only when it passes 5 gym runs).

Examples

Examples of evals built from real work
What happened in real workThe eval it became
Asked to check a different version of a page, the QA Wolf agent changed a shared helper so it required a new argument. Every existing test that called the helper broke.A suite where one test needs a variation of an existing helper. It passes only if the original helper and all its existing callers still work.
A customer found tests that confirmed an editor had opened but never checked that the edit took effect.Evals that pass only if the test asserts the result of the action, not just that a control appeared.
Some suites have tests where three or four users work in the same app at once.Multi-user evals that check every user signs in at the start, in parallel, and that the test switches between their sessions correctly.
At a sign-in page, the QA Wolf agent asked users to configure saved settings instead of asking for the login it needed.An eval that reaches a sign-in page and passes only if the QA Wolf agent asks the user for credentials.
Some customer apps are harder to automate than our test apps, with fewer stable selectors and state that persists between visits.Test apps that keep state between visits, with more complex apps replacing simpler ones.

Ranking and running evals

Each eval is weighted by how often the problem comes up in real work and how much it matters. Evals that no longer measure an aspect of QA Wolf's approach are removed.

How QA Wolf runs its evals

Each eval is a request a user might make of a testing agent, plus the workspace and app it starts from. The agent works on the request, and its output is graded against a set of checks.

  1. The agent gets the request. It works the way it would for a real user: exploring the app, reading the existing code, writing or changing tests.
  2. The output is graded check by check. Each check is a single pass/fail condition, such as "the test asserts the confirmation message" or "the password is read from an environment variable." Some checks are code that inspects the files and test runs; others are judged by a model reading the result.
  3. Results are recorded per eval. An eval's result is the set of checks it passed and failed.

Methodology

The comparison covers 1,072 evals. All three agents ran at medium reasoning effort.

The three agents compared
AgentModelBrowser access
QA Wolfgpt-6.1-sol, with gpt-5.6-luna subagentsQA Wolf's own browser tooling
Claude CodeOpus 5.5Playwright MCP
Codexgpt-6.1-solPlaywright MCP

‍What the checks measure: Each check asks whether the work would hold up as a working test delivered to a customer:

  • Does the test do what was asked, and does it assert the right things?
  • Does it run?
  • Are credentials kept out of the code, and is test data cleaned up?
  • Does it work with the infrastructure a real test depends on, such as receiving email?
Try the AI testing platform that makes QA 12x faster.