Specific Labs has released Real-SWE, a coding benchmark built on private production code licensed from real companies rather than public repositories or purpose-built puzzles. Across 640 runs covering 8 model-and-harness combinations, the best result came from Anthropic's Fable 5.1 at a 38.8 percent resolution rate.
Measuring Work That Public Repositories Cannot Represent
Most coding benchmarks draw on public repositories or tasks written specifically for evaluation. The first kind may already sit in a model's training data; the second tends to strip away the constraints that make real engineering hard. Real-SWE takes a different route by licensing production code directly from companies. Code written inside a company does not exist on the public internet, which puts it natively outside the distribution any model was trained on.
The selected codebases include an event platform with more than 200,000 users that reached the top 100 on the App Store, a consumer fintech system that processes over 100,000 bank statements, and enterprise AI sales platforms. All of them carry logic tied directly to revenue and billing. One task, for example, asks the agent to fix invoice taxation so that each business on the platform is charged under its own tax arrangement, with queries to an external tax authority service, exempt customers handled correctly, and settled sales filed back so the returns reconcile.
A single task touches a median of 11 files, and instructions run around 1,742 characters. The prompt states what needs to change and leaves the implementation details to be discovered in the code and the surrounding services. Models are not run in isolation: each is paired with its native harness, such as Claude Code, Codex CLI, or Gemini CLI. Every run happens in an isolated sandbox exposing only the services that task needs, drawn from an AWS emulator, Docker, Kubernetes, PostgreSQL, Redis, Slack, Intercom, and others.
Even the Leader Clears Fewer Than 4 of 10 Tasks
The benchmark covers 10 tasks run 8 times each across 8 model-and-harness pairings, for 640 scored rollouts. Results are as follows.
| Rank | Model | Harness | Resolution Rate | Cost per Rollout |
|---|---|---|---|---|
| 1 | Fable 5.1 | Claude Code | 38.8 percent | 6.96 USD (about 1,070 yen) |
| 2 | GPT-6 Astra | Codex CLI | 33.8 percent | 4.67 USD (about 710 yen) |
| 3 | Gemini 3.8 Flash | Gemini CLI | 31.2 percent | 2.50 USD (about 380 yen) |
| 4 | GLM 5.3 | Claude Code | 28.8 percent | 5.12 USD (about 780 yen) |
| 5 | Grok 4.6 | Grok Build | 23.8 percent | 3.44 USD (about 530 yen) |
| 5 | Muse Spark 1.3 | Muse Code | 23.8 percent | 2.74 USD (about 420 yen) |
| 7 | Kimi K3 | Kimi Code | 18.8 percent | 3.90 USD (about 600 yen) |
| 8 | GPT-5.6 Sol | Codex CLI | 16.2 percent | 2.65 USD (about 410 yen) |
※1 USD = 153 JPY (as of September 10, 2026)
At 38.8 percent, the leading configuration still fails 6 of the 10 tasks. Six tasks sat below 15 percent for every model, and one task involving an analytics stream rewrite was never solved by any of them. Against the near-90 percent figures that public benchmarks now routinely produce, the gap is striking.
Models Fail in Different Ways
Classifying the failed runs, the single most common pattern across the board is a missed requirement — submitting work that does not satisfy a condition stated in the instructions or discoverable in the code.
The breakdown varies sharply by model. For Grok 4.6, missed requirements account for 67.2 percent of failures, and for Kimi K3, 53.8 percent. Gemini 3.8 Flash instead fails mostly on integration errors, at 49.1 percent, while GPT-5.6 Sol most often acts on an unverified assumption, at 43.3 percent. Fable 5.1 is more evenly spread, with 36.7 percent missed requirements and 34.7 percent integration errors.
That distinction matters in practice. Missed requirements are relatively easy for a human reviewer to catch, whereas a change that breaks integration with existing behavior can slip past tests and surface in production. Two models with the same resolution rate can demand very different review discipline.
Paying More Does Not Guarantee Better Results
Per-rollout costs land between 2.50 and 6.96 USD. Fable 5.1 leads the table at the highest price, but Gemini 3.8 Flash reaches 31.2 percent at roughly a third of that cost, which reverses the ranking on a cost-effectiveness basis. When a budget is fixed, working down the leaderboard from the top is not automatically the right call.
Time behaves the same way. Rollouts finishing in under 10 minutes failed 71.4 percent of the time, barely different from the 73.4 percent failure rate of longer ones. These are not tasks that come good with more thinking time; the failures cluster around not grasping what needed to be done in the first place.
How to Read the Numbers
Treating these results as a definitive ranking of the labs would be premature. Ten tasks and 640 rollouts is not a large sample, and the gap between adjacent models could easily fall within noise.
The pairing of each model with its own harness complicates interpretation further. Fable 5.1 and GLM 5.3 both ran on Claude Code, so model quality and tooling quality are not separated. The codebases are private, which means no third party can reproduce the run under the same conditions. These are numbers published by the party that built the benchmark, and that context belongs in any reading of them.
Even so, as a way to think about what happens when an agent is pointed at your own codebase, this is closer to reality than a leaderboard built on public repositories. For anyone evaluating adoption, the failure breakdown is more informative than the ranking itself.
Summary
Real-SWE evaluates coding agents on private production code licensed from real companies. Across 10 tasks and 640 rollouts, Fable 5.1 led at 38.8 percent, followed by GPT-6 Astra at 33.8 percent and Gemini 3.8 Flash at 31.2 percent. Six tasks stayed under 15 percent for every model, and missed requirements dominate the failure profile. More expensive models were not reliably stronger, and longer runs did not improve resolution rates. The small sample and the absence of independent replication are real limits, but the results help explain why high public-benchmark scores so often fail to match what teams see in their own repositories.
