ShvIA
AI BENCHMARK SIGN IN

Can an AI maintain legacy code without breaking it?

AI Benchmark · LEB · results read from the ai-benchmark repository

LEB — the LLM Engineering Benchmark — hands an AI agent a legacy system in production, with flaws planted in it and consumers that depend on how it behaves today. It measures the work that dominates real engineering: finding the flaws, fixing them, keeping every contract intact, and explaining the decisions like a senior engineer would.

Top 6 so far

LEB-100-A
  1. 1Claude Sonnet 5.5 · xhigh809LEB Gold
  2. 2Claude Sonnet 5.5 · max807LEB Gold
  3. 3Claude Sonnet 5.5 · max (ultracode)774LEB Gold
  4. 4Claude Fable 5.1764LEB Gold
  5. 5Claude Opus 5.5717LEB Silver
  6. 6GPT-6.1-sol · xhigh661LEB Silver

Score out of 1000, of 31 agents. Not official until each has three runs. Full leaderboard

Why another benchmark

Most benchmarks measure code written from scratch, or one isolated issue solved. Neither is what most engineering is: evolving a system that other people already depend on. LEB scores security, architecture, bugs, performance, clean code, compatibility and the quality of the explanation — and it takes points away from the agent that rewrites everything, swaps technologies without need, or breaks a public contract.

Rewriting from scratch is not engineering. It is running away.

How a run works

  1. 01

    A legacy system, with planted flaws

    The agent receives the code, a manifest of its public surface — the contract — and a neutral task: report the problems, fix what should be fixed, keep compatibility, justify every decision. It is never told which flaws exist, how many, or where.

  2. 02

    The agent works alone

    In mode A it gets tools and a budget of turns; in mode S, one prompt and one answer. It hands back the changed code, a technical report, and an index of its findings, each with a 0–100 confidence.

  3. 03

    Machines check the code

    Characterization tests run on the legacy code and on the delivery: public behaviour that changed is a regression. Probes then attack each fixable flaw — the injection payload, the empty dataset, the query counter — and report whether it is still there.

  4. 04

    A judge checks the report

    Each finding is matched against the Official Failure Matrix, a hidden answer key that also holds decoys: plausible flaws that do not exist, and cost points when reported. A second judge scores the explanation blind. A deterministic scorer turns it all into 0–1000.

1000 points, and how they are lost

Every instance is worth exactly 1000: the raw points of each category are normalised to its weight, so scores compare across instances of a level.

Compatibility starts at 100 and only goes down. Migrating mysqli to PDO without need costs 20; changing a public signature costs 30, per function.

Global penalties come off the total: a new bug −15, each broken characterization test −20, a needless rewrite −25, each decoy reported −5.

  • Security SEC250
  • Architecture ARCH200
  • Bugs BUG150
  • Performance PERF150
  • Clean code CLN100
  • Compatibility COMP100
  • Explanation EXPL50

Grades

  • LEB Platinum 900–1000 · ready for critical legacy
  • LEB Gold 750–899 · solid engineering
  • LEB Silver 600–749 · useful with supervision
  • LEB Bronze 400–599 · needs a full review
  • Failed < 400 · a risk to the system

Runs happen where the answer key is out of reach

Agents run on a dedicated Linux VM isolated from GitHub. Its names resolve to loopback, its address ranges (the prefixes it announces and the edge addresses it publishes) are blackhole routes, and the VM has no IPv6 connectivity. The address blocks came on 30 September 2026. The runs of 29 September had the name block only: no agent could reach GitHub by name, but a deliberate connection straight to one of its addresses was not blocked, and the session logs show none was attempted. Of the runs of 30 September, all but Kimi K3's had every layer; Kimi K3 ran on a clone of the VM whose address blocks are not recorded. What a model saw in training is a separate question, answered by the training cutoff each run records.

Results

Every delivery in a table solved the same package, byte for byte — the same SHA-256 — so the numbers compare like for like.

LEB-100-A v1.1 · Support-ticket panel of an internet provider

The support-ticket panel of an internet provider, written in 2013-style PHP: data-access functions and an index.php that routes, authorizes and builds the HTML. About 300 lines on PHP 8, mysqli and MySQL 8, with 13 planted flaws and 2 decoys.

  • mode A · 30 turns
  • edition 2026
  • evaluated 2026-10-02
  • matrix 68088abdb7bc…
  1. 1
    Claude Sonnet 5.5 Anthropic · effort xhigh 3 of 3 runs (825 · 809 · 724)
    809 of 1000 LEB Gold
    • Security250/250
    • Architecture25/200
    • Bugs139/150
    • Performance150/150
    • Clean code100/100
    • Compatibility100/100
    • Explanation45/50

    Penalties: noneDiscovery 91.7Brier 0.033Scorecard

  2. 2
    Claude Sonnet 5.5 Anthropic · effort max 3 of 3 runs (807 · 773 · 820)
    807 of 1000 LEB Gold
    • Security250/250
    • Architecture12/200
    • Bugs150/150
    • Performance150/150
    • Clean code100/100
    • Compatibility100/100
    • Explanation45/50

    Penalties: noneDiscovery 91.7Brier 0.035Scorecard

  3. 3
    Claude Sonnet 5.5 Anthropic · effort max (ultracode) 1 of 3 runs not official
    774 of 1000 LEB Gold
    • Security250/250
    • Architecture0/200
    • Bugs129/150
    • Performance150/150
    • Clean code100/100
    • Compatibility100/100
    • Explanation45/50

    Penalties: noneDiscovery 87.5Brier 0.008Scorecard

  4. 4
    Claude Fable 5.1 Anthropic · effort xhigh 2 of 3 runs (781 · 764) not official
    764 of 1000 LEB Gold
    • Security233/250
    • Architecture0/200
    • Bugs139/150
    • Performance150/150
    • Clean code100/100
    • Compatibility100/100
    • Explanation42/50

    Penalties: noneDiscovery 87.5Brier 0.021Scorecard

  5. 5
    Claude Opus 5.5 Anthropic · effort xhigh 3 of 3 runs (711 · 717 · 805)
    717 of 1000 LEB Silver
    • Security233/250
    • Architecture50/200
    • Bugs139/150
    • Performance150/150
    • Clean code0/100
    • Compatibility100/100
    • Explanation45/50

    Penalties: noneDiscovery 87.5Brier 0.024Scorecard

  6. 6
    GPT-6.1-sol OpenAI · effort xhigh 3 of 3 runs (666 · 653 · 661)
    661 of 1000 LEB Silver
    • Security246/250
    • Architecture25/200
    • Bugs129/150
    • Performance150/150
    • Clean code0/100
    • Compatibility70/100
    • Explanation41/50

    Penalties: noneDiscovery 75.0Brier 0.000Scorecard

  7. 7
    GPT-6.1-sol pro OpenAI · effort xhigh 1 of 3 runs not officialtraining cutoff not published
    654 of 1000 LEB Silver
    • Security228/250
    • Architecture0/200
    • Bugs139/150
    • Performance150/150
    • Clean code25/100
    • Compatibility70/100
    • Explanation42/50

    Penalties: noneDiscovery 87.5Brier 0.000Scorecard

  8. 8
    Grok 4.7 xAI · effort high 3 of 3 runs (638 · 663 · 607)
    638 of 1000 LEB Silver
    • Security207/250
    • Architecture0/200
    • Bugs139/150
    • Performance150/150
    • Clean code0/100
    • Compatibility100/100
    • Explanation42/50

    Penalties: noneDiscovery 83.3Brier 0.005Scorecard

  9. 9
    Grok 4.6 xAI · effort high 1 of 3 runs not officialtraining cutoff not published
    633 of 1000 LEB Silver
    • Security194/250
    • Architecture0/200
    • Bugs129/150
    • Performance150/150
    • Clean code25/100
    • Compatibility100/100
    • Explanation35/50

    Penalties: noneDiscovery 62.5Brier 0.007Scorecard

  10. 10
    GPT-6-astra OpenAI · effort xhigh 3 of 3 runs (661 · 596 · 628)
    628 of 1000 LEB Silver
    • Security228/250
    • Architecture0/200
    • Bugs139/150
    • Performance150/150
    • Clean code0/100
    • Compatibility70/100
    • Explanation41/50

    Penalties: noneDiscovery 83.3Brier 0.002Scorecard

  11. 11
    GLM-5.3 Prime Z.AI · effort high 2 of 3 runs (635 · 628) not officialtraining cutoff not published
    628 of 1000 LEB Silver
    • Security211/250
    • Architecture0/200
    • Bugs129/150
    • Performance150/150
    • Clean code0/100
    • Compatibility100/100
    • Explanation38/50

    Penalties: noneDiscovery 70.8Brier 0.029Scorecard

  12. 12
    GPT-5.6-terra OpenAI · effort xhigh 3 of 3 runs (625 · 611 · 645)
    625 of 1000 LEB Silver
    • Security211/250
    • Architecture0/200
    • Bugs129/150
    • Performance150/150
    • Clean code0/100
    • Compatibility100/100
    • Explanation35/50

    Penalties: noneDiscovery 45.8Brier 0.000Scorecard

  13. 13
    GLM-5.3-Flash Z.AI · effort high 1 of 3 runs not officialtraining cutoff not published
    624 of 1000 LEB Silver
    • Security211/250
    • Architecture0/200
    • Bugs129/150
    • Performance150/150
    • Clean code0/100
    • Compatibility100/100
    • Explanation34/50

    Penalties: noneDiscovery 70.8Brier 0.016Scorecard

  14. 14
    GLM-5.3 Z.AI · effort high 2 of 3 runs (629 · 621) not officialtraining cutoff not published
    621 of 1000 LEB Silver
    • Security203/250
    • Architecture0/200
    • Bugs129/150
    • Performance150/150
    • Clean code0/100
    • Compatibility100/100
    • Explanation39/50

    Penalties: noneDiscovery 70.8Brier 0.014Scorecard

  15. 15
    GPT-5.6-sol OpenAI · effort xhigh 3 of 3 runs (612 · 612 · 608)
    612 of 1000 LEB Silver
    • Security246/250
    • Architecture0/200
    • Bugs129/150
    • Performance150/150
    • Clean code0/100
    • Compatibility70/100
    • Explanation32/50

    Penalties: -15Discovery 70.8Brier 0.001Scorecard

  16. 16
    DeepSeek V4 Flash DeepSeek · effort high 1 of 3 runs not officialtraining cutoff not published
    612 of 1000 LEB Silver
    • Security224/250
    • Architecture0/200
    • Bugs139/150
    • Performance150/150
    • Clean code0/100
    • Compatibility70/100
    • Explanation29/50

    Penalties: noneDiscovery 70.8Brier 0.006Scorecard

  17. 17
    DeepSeek V4.1 Flash DeepSeek · effort high 3 of 3 runs (625 · 612 · 597) training cutoff not published
    612 of 1000 LEB Silver
    • Security203/250
    • Architecture0/200
    • Bugs129/150
    • Performance150/150
    • Clean code0/100
    • Compatibility100/100
    • Explanation30/50

    Penalties: noneDiscovery 58.3Brier 0.013Scorecard

  18. 18
    GPT-5.6-luna OpenAI · effort xhigh 2 of 3 runs (599 · 601) not official
    599 of 1000 LEB Bronze
    • Security207/250
    • Architecture0/200
    • Bugs139/150
    • Performance150/150
    • Clean code0/100
    • Compatibility70/100
    • Explanation33/50

    Penalties: noneDiscovery 83.3Brier 0.003Scorecard

  19. 19
    GPT-6.1-sol OpenAI · effort ultra 1 of 3 runs not official
    597 of 1000 LEB Bronze
    • Security207/250
    • Architecture0/200
    • Bugs129/150
    • Performance150/150
    • Clean code0/100
    • Compatibility70/100
    • Explanation41/50

    Penalties: noneDiscovery 70.8Brier 0.001Scorecard

  20. 20
    GLM-5.3-FlashX Z.AI · effort high 1 of 3 runs not officialtraining cutoff not published
    597 of 1000 LEB Bronze
    • Security185/250
    • Architecture0/200
    • Bugs129/150
    • Performance150/150
    • Clean code0/100
    • Compatibility100/100
    • Explanation33/50

    Penalties: noneDiscovery 70.8Brier 0.015Scorecard

  21. 21
    Gemini 3.8 Flash Google · effort high 2 of 3 runs (687 · 588) not officialtraining cutoff not published
    588 of 1000 LEB Bronze
    • Security181/250
    • Architecture0/200
    • Bugs129/150
    • Performance150/150
    • Clean code0/100
    • Compatibility100/100
    • Explanation28/50

    Penalties: noneDiscovery 58.3Brier 0.003Scorecard

  22. 22
    GPT-5.5 OpenAI · effort xhigh 2 of 3 runs (601 · 558) not official
    558 of 1000 LEB Bronze
    • Security177/250
    • Architecture0/200
    • Bugs129/150
    • Performance150/150
    • Clean code0/100
    • Compatibility70/100
    • Explanation32/50

    Penalties: noneDiscovery 58.3Brier 0.006Scorecard

  23. 23
    Gemini 3.8 Flash Google · effort medium 1 of 3 runs not officialtraining cutoff not published
    550 of 1000 LEB Bronze
    • Security147/250
    • Architecture0/200
    • Bugs129/150
    • Performance150/150
    • Clean code0/100
    • Compatibility100/100
    • Explanation24/50

    Penalties: noneDiscovery 45.8Brier 0.201Scorecard

  24. 24
    Kimi K3 Moonshot AI · default effort (not configurable) 1 of 3 runs not officialtraining cutoff not published
    528 of 1000 LEB Bronze
    • Security203/250
    • Architecture0/200
    • Bugs139/150
    • Performance56/150
    • Clean code0/100
    • Compatibility100/100
    • Explanation30/50

    Penalties: noneDiscovery 70.8Brier 0.007Scorecard

  25. 25
    Qwen3 Coder Next Alibaba (Qwen team) · default effort (not configurable) 1 of 3 runs not officialtraining cutoff not published
    507 of 1000 LEB Bronze
    • Security125/250
    • Architecture0/200
    • Bugs129/150
    • Performance150/150
    • Clean code0/100
    • Compatibility100/100
    • Explanation18/50

    Penalties: -15Discovery 54.2Brier 0.301Scorecard

  26. 26
    DeepSeek V4 Pro DeepSeek · effort high 3 of 3 runs (604 · 496 · 432) training cutoff not published
    496 of 1000 LEB Bronze
    • Security181/250
    • Architecture0/200
    • Bugs129/150
    • Performance56/150
    • Clean code0/100
    • Compatibility100/100
    • Explanation30/50

    Penalties: noneDiscovery 58.3Brier 0.026Scorecard

  27. 27
    MiniMax-M3 MiniMax · default effort (not configurable) 1 of 3 runs not officialtraining cutoff not published
    460 of 1000 LEB Bronze
    • Security185/250
    • Architecture0/200
    • Bugs129/150
    • Performance56/150
    • Clean code0/100
    • Compatibility70/100
    • Explanation20/50

    Penalties: noneDiscovery 70.8Brier 0.011Scorecard

  28. 28
    Kimi K2.7 Code Moonshot AI · default effort (not configurable) 1 of 3 runs not officialtraining cutoff not published
    415 of 1000 LEB Bronze
    • Security203/250
    • Architecture0/200
    • Bugs86/150
    • Performance0/150
    • Clean code0/100
    • Compatibility100/100
    • Explanation26/50

    Penalties: noneDiscovery 50.0Brier 0.225Scorecard

  29. 29
    GPT-5.3-Codex OpenAI · effort xhigh 1 of 3 runs not official
    403 of 1000 LEB Bronze
    • Security147/250
    • Architecture0/200
    • Bugs129/150
    • Performance0/150
    • Clean code0/100
    • Compatibility100/100
    • Explanation27/50

    Penalties: noneDiscovery 37.5Brier 0.015Scorecard

  30. 30
    GLM-5.2 Z.AI · effort high 1 of 3 runs not officialtraining cutoff not published
    388 of 1000 Failed
    • Security172/250
    • Architecture0/200
    • Bugs86/150
    • Performance0/150
    • Clean code0/100
    • Compatibility100/100
    • Explanation30/50

    Penalties: noneDiscovery 50.0Brier 0.001Scorecard

  31. 31
    Claude Haiku 4.5 Anthropic · default effort (not configurable) 2 of 3 runs (317 · 369) not official
    317 of 1000 Failed
    • Security138/250
    • Architecture0/200
    • Bugs86/150
    • Performance0/150
    • Clean code0/100
    • Compatibility70/100
    • Explanation23/50

    Penalties: noneDiscovery 37.5Brier 0.215Scorecard

Flaw by flaw

What each agent found and fixed among the planted flaws. The hard ones are flaws of absence — a missing authorization check, a session never regenerated, a file left open on the error path.

The table shows the top 10 of 31 agents. Every agent's result, flaw by flaw, is in its scorecard.

  • fixed
  • found, not fixed
  • missed
FlawClaude Sonnet 5.5 · xhighClaude Sonnet 5.5 · maxClaude Sonnet 5.5 · max (ultracode)Claude Fable 5.1Claude Opus 5.5GPT-6.1-sol · xhighGPT-6.1-sol proGrok 4.7Grok 4.6GPT-6-astra
SQL injection in the searchSEC-001 · critical · easy10/10fixed10/10fixed10/10fixed10/10fixed10/10fixed10/10fixed10/10fixed10/10fixed10/10fixed10/10fixed
Reflected XSS in the searchSEC-003 · high · easy8/8fixed8/8fixed8/8fixed8/8fixed8/8fixed8/8fixed8/8fixed8/8fixed8/8fixed8/8fixed
Formula injection in the CSV exportSEC-008 · medium · hard6/6fixed6/6fixed6/6fixed2/6found, not fixed2/6found, not fixed6/6fixed2/6found, not fixed6/6fixed0/6missed2/6found, not fixed
Session fixation at loginSEC-013 · high · hard8/8fixed8/8fixed8/8fixed8/8fixed8/8fixed8/8fixed8/8fixed8/8fixed8/8fixed8/8fixed
Unsalted MD5 passwordsSEC-014 · high · easy8/8fixed8/8fixed8/8fixed8/8fixed8/8fixed8/8fixed8/8fixed3/8found, not fixed3/8found, not fixed8/8fixed
Secrets hardcoded in the configSEC-015 · high · easy8/8fixed8/8fixed8/8fixed8/8fixed8/8fixed8/8fixed8/8fixed3/8found, not fixed6/8found, not fixed8/8fixed
Any ticket readable by id (IDOR)SEC-017 · critical · hard10/10fixed10/10fixed10/10fixed10/10fixed10/10fixed9/10fixed9/10fixed10/10fixed10/10fixed9/10fixed
Division by zero in the SLA averageBUG-001 · high · moderate8/8fixed8/8fixed8/8fixed8/8fixed8/8fixed8/8fixed8/8fixed8/8fixed8/8fixed8/8fixed
File handle leaked on the error pathBUG-004 · medium · hard5/6fixed6/6fixed4/6fixed5/6fixed5/6fixed4/6fixed5/6fixed5/6fixed4/6fixed5/6fixed
One query per ticket for the technician (N+1)PERF-001 · high · moderate8/8fixed8/8fixed8/8fixed8/8fixed8/8fixed8/8fixed8/8fixed8/8fixed8/8fixed8/8fixed
A dispatcher that does everythingARCH-002 · high · easy2/10found, not fixed1/10found, not fixed0/10missed0/10missed4/10found, not fixed2/10found, not fixed0/10missed0/10missed0/10missed0/10missed
Magic numbers for status and priorityARCH-009 · low · moderate0/6missed0/6missed0/6missed0/6missed0/6missed0/6missed0/6missed0/6missed0/6missed0/6missed
Four levels of nested ifsCLN-007 · medium · easy8/8fixed8/8fixed8/8fixed8/8fixed0/8missed0/8missed2/8found, not fixed0/8missed2/8found, not fixed0/8missed

What stood out

  • Six agents fixed the formula injection in the CSV (SEC-008) — Sonnet 5.5 in all three of its settings, GPT-6.1-sol, GPT-5.6-sol and Grok 4.7 — with the fix the answer key expects; Sonnet 5.5 is the only model with security at 250 of 250, in all three of its settings. GPT-5.6-sol also turned the - of a ticket with no technician into '-, a new bug (−15).
  • Architecture was the weakest category for everyone: 50 of 200 for Opus 5.5, 25 for Sonnet 5.5 at xhigh and for GPT-6.1-sol, 12 for Sonnet 5.5 at max, and 0 for the other twenty-seven. Nobody split the dispatcher that does everything; those three named it and declined to restructure it.
  • Nobody broke the contract mechanically. All thirty-one stayed on mysqli, kept the 22 characterization checks green and reported no decoy. Judgement is what separated them: twenty-one kept compatibility at 100, while the other ten each changed a business value (−30), most often the SLA average, scoped to each client.
  • Ten scores are official: Claude Sonnet 5.5 at xhigh, 809 (runs of 825, 809 and 724), Claude Sonnet 5.5 at max, 807 (807, 773 and 820), Claude Opus 5.5, 717 (711, 717 and 805), GPT-6.1-sol, 661 (666, 653 and 661), Grok 4.7, 638 (638, 663 and 607), GPT-6-astra, 628 (661, 596 and 628), GPT-5.6-terra, 625 (625, 611 and 645), DeepSeek V4.1 Flash, 612 (625, 612 and 597), GPT-5.6-sol, 612 (612, 612 and 608), and DeepSeek V4 Pro, 496 (604, 496 and 432), each the median of three runs. A single run can sit more than 100 points from the median: Opus's third scored 805, Sonnet's third 724, DeepSeek V4 Pro's first 604, and astra's first, 661, had placed it 5th, from judgement calls such as which flaws to leave unfixed. Claude Fable 5.1 has two runs, 781 and 764, and publishes the lower, 4th
  • Gemini 3.8 Flash swung 99 points between two runs at high (687 and 588) and publishes the lower, 21st, Bronze. Its first run, 6th, fixed 8 of the 13 planted flaws and is the only one outside Claude to flatten the nested ifs (CLN-007); its second fixed 7, and searched the VM for the answer key, by name and by the hash quoted in the task. The key is not on the VM, and the search turned up nothing the run did not already have. Both runs kept MD5 and both secrets and left the CSV export serving every client. At medium effort the same model scored 550 (23rd) in 4.5 minutes: it fixed 6 flaws and claimed SQL injection in two functions that only take integers.
  • GPT-6.1-sol is the strongest GPT model here, official at 661 (runs of 666, 653 and 661), 6th; GPT-6.1-sol pro follows at 654, 7th, from one run. GPT-6-astra, official at 628, is 10th. Its first run fixed the CSV injection and kept MD5; its official run did the opposite, migrating MD5 to password_hash and leaving the CSV injection alone; GPT-6.1-sol's official run fixed both
  • Eight hours of multi-agent work scored less than half an hour of one agent at the same effort. Claude Sonnet 5.5 in Claude Code's multi-agent mode, at max effort, ran 7 workflows with 68 subagents for 8.1 hours and US$ 233, about 65 times the cost of its xhigh run, and scored 774 (3rd, Gold). The same model at the same max effort as a single agent took about half an hour a run and is official at 807 (runs of 807, 773 and 820), 2nd, 33 points higher; at xhigh it is official at 809. It fixed the same flaws, with better calibration and a report scored 45 of 50, but left the architecture untouched. Nine of its subagents fell back to an older Sonnet after a safety classifier stopped them; none of the delivery came from them.
  • GPT-6.1-sol at ultra scored 64 points below itself at xhigh (597 against an official 661), the second case here of more effort scoring less. With three subagents it found 9 of the 13 planted flaws against 11, kept MD5, which the xhigh run had migrated, and did not name the dispatcher. Unlike the multi-agent Sonnet, it took 24 minutes, not 8 hours.
  • Grok 4.7 is the strongest model here from outside Anthropic and OpenAI (638, 8th, Silver, official over three runs of 638, 663 and 607), with an explanation at the level of the best GPT reports (42 of 50). In its first run it is also the only agent that used the web, to read the PHP manual page for fputcsv; the VM blocks GitHub, not the web. Grok 4.6 follows at 633 (9th) for about an eighth of the cost (US$ 0.31 against 2.43).
  • GLM-5.3 Prime is the strongest GLM model (628, 11th, Silver), the lower of two runs through two hosts (635 and 628): the first fixed all four probe-covered flaws, the CSV injection among them, for US$ 1.69; the second left the injection alone. GLM-5.3 publishes 621 (14th), the lower of its two runs (629 and 621), the second served by another host; GLM-5.3-Flash scores 624 for about a tenth of the cost, and GLM-5.3-FlashX 597
  • DeepSeek V4.1 Flash is the cheapest agent here (US$ 0.01 to 0.04 a run; its last two took under two minutes each) and is official at 612 (runs of 625, 612 and 597), 17th. DeepSeek V4 Flash, run through another host, scores 612: it fixed eight flaws fully but lost 30 points for relabelling the fallback of formatarStatus. DeepSeek now routes the V4 Flash name to V4.1 on its own API, so which weights that host served is not certain. DeepSeek V4 Pro, the larger model, is official at 496, the median of three runs through two hosts (604, 496 and 432), all below the Flash models: the first rates SQL injection at confidence 100 in two functions that only take integers, and the other two left the N+1 query in place.
  • Twenty agents did not run at xhigh. MiniMax-M3, Kimi K2.7 Code, Qwen3 Coder Next and Claude Haiku 4.5 have no effort setting and ran at their model's default; Kimi K3 is filed at its default; the two Grok models, the five GLM models, the three DeepSeek models and Gemini 3.8 Flash ran at high, in opencode, and Gemini 3.8 Flash also at medium; Sonnet 5.5 ran twice at max, once in multi-agent mode, and GPT-6.1-sol once at ultra. The rest ran at xhigh.
  • Kimi K3, Qwen3 Coder Next, DeepSeek V4 Pro, MiniMax-M3, Kimi K2.7 Code, GPT-5.3-Codex, GLM-5.2 and Claude Haiku 4.5 close the table (528, 507, 496, 460, 415, 403, 388 and 317; the last two are below the pass line). Claude Haiku 4.5, last, publishes the lower of its two runs (317 and 369); in the first, 2.5 minutes and US$ 0.21, its visibility fix hides a client's own tickets on the main page, comparing an integer with the string mysqli returns. All but Qwen3 Coder Next left the N+1 query in place. Qwen3 Coder Next has the weakest explanation (18 of 50) and the worst calibration (Brier 0.301): it reported three flaws that cannot exist at confidence 100, SQL injection in two functions that only take integers among them, as Kimi K2.7 Code did. It also ran no tests.
  • Places 8 to 20 sit within 41 points (638 to 597), inside the noise of a single run. GPT-5.6-terra reported the fewest planted flaws and still ranks 12th, on compatibility and on what it did fix.
  • The GPT models are among the best calibrated (Brier 0.015 or less): fewer findings, each stated with high confidence and nearly all real.

Finding is not fixing

Side by side, the three Claude models find nearly the same flaws and fix different numbers of them. One total hides which of the two it is measuring.

  1. They find nearly the same. In the run that counts for each, Sonnet 5.5 at xhigh reported 12 of the 13 planted flaws and Fable 5.1 and Opus 5.5 11; across their eight runs, each found 11 or 12
  2. Sonnet 5.5 fixes more, in the run that counts. Its official run fixed 11 of the 13, against 9 for Opus's official run and 10 for Fable's. Run by run it varies: Sonnet fixed 11, 11 and 8, Opus 9 in each of its three, Fable 9 and 10. Fable and Opus found and explained the formula injection in the CSV and left it in place, to protect the file's consumers; Sonnet fixed it in two of its three runs, prefixing only the cells that start a formula. Its 92-point lead over Opus's official 717 comes from clean code (+100), not from security (+17); its 45-point lead over Fable's 764 is security (+17) and architecture (+25)
  3. More compute did not change the pattern. The single-agent Claude runs took 16 to 27 minutes each, and the three at the same max effort scored 807, 773 and 820. Sonnet 5.5 in multi-agent mode took almost 8 hours and US$ 233, about 25 times as long and 65 times the cost of its single-agent run, and scored less: the same fixes and a deeper check of its own code, but the dispatcher that the single-agent run named never reached its report.
  4. How to read it. On this task the three Claude models see the same problems; what separates them is how far they go in fixing them within the contract. In the median, Sonnet went further; a single run can reverse it, as Sonnet's third did. One task, two or three runs per model: a pattern worth testing, not a verdict

Read this before quoting a number

  • Up to three runs per agent. An official LEB score is the median of three independent runs. Until an agent has three, its score is the lower of its totals so far, the totals are listed next to it, and everything shown with the score comes from that same run.
  • The judge is an AI. Claude Opus 5.5 applied the published rubric to every delivery without knowing which model wrote it — each was anonymised — and the explanation was scored by a separate judge that saw neither the answer key nor the other scores.
  • The judge is also a contestant. Claude Opus 5.5 is one of the agents evaluated, and six of the thirty-one are Claude models, Sonnet 5.5 three times: they hold the top five places and the last. Anonymity limits that bias; it does not remove it, since a model can recognise its own style. Every verdict is published with its rationale, flaw by flaw, and the four verdicts changed in review say why; one of them adds 41 points to GPT-5.6-terra's total.
  • Ten attempts were left unscored, each before any judge saw it. Twice, Claude Fable 5.1's safeguards stopped one of its responses while it worked on the security flaws, and its client handed the rest of the run to Claude Opus 4.8, which wrote the whole report; a run must come from one model, so both are void, and Fable 5.1 shows its two complete runs. Four ended on the provider's side before the delivery was complete: GPT-5.6-sol pro when its gateway ran out of credit, and Gemini 3.8 Flash three times, on a gateway timeout and twice on a rate limit. Three were set up wrong: the first multi-agent Sonnet 5.5 run was stopped to give the machine more processors, Claude Haiku 4.5's first attempt had its client mode changed mid-run, and a Gemini attempt was started outside the task's folder. One finished, but GPT-5.6-terra ran in a different client from its first run, and was repeated in the same one. All ten are kept in the repository.
  • The answer key is public. The failure matrix of LEB-100-A has been in the public repository since 13 July 2026. The runs of 29 September had the VM's name block only, so no agent could fetch it by accident; the runs of 30 September also had GitHub's addresses blocked, and the session logs of every run that left one show no request to GitHub at all. What a model saw in training depends on how far its training data reaches: each run records the cutoff its provider publishes, and a model whose cutoff is later, or not published, is marked in the leaderboard.
  • The benchmark's own tests were fixed. Scoring these runs exposed two defects in the evaluation tooling: an SQL loader that split a statement on a semicolon inside a comment, and CSV checks that read a temporary file the contract never promised. Both were fixed before scoring, the same way for every agent, and are recorded in the repository.
  • The machine changed on 30 September. Until then the agents ran as the VM's administrator, with sudo, on a machine that kept earlier runs' leftovers. Their session logs show no agent reading another's work; the opencode agents shared a leftover test database that held only the seed rows. From Qwen3 Coder Next on, runs are made as an unprivileged user on a machine cleaned before each run.
  • Not every run parameter was recorded. The exact model version and the temperature were not, and cost and model time only where the client kept them; GPT-5.5's, MiniMax-M3's and Kimi K3's runs left no session log at all. Each run marks what is missing instead of guessing.

Audit it

Every delivery, mechanical report, verdict and scorecard is in the repository, next to the specification that produced them. The same data is here as two spreadsheets: one row per run, and one per run and planted flaw.

Results on GitHub Specification Runs (CSV) Flaw by flaw (CSV) What each column means

The numbers on this page are not typed by hand: they are rendered from the published results of the ai-benchmark repository, the same data samirhv.com.br shows.