Last Secure Code Benchmark
Is the Code You Vibe Secure?
450 tasks built from 150 real vulnerabilities — 7 languages, 115 CVEs, 24 CWEs. Each task asks the agent to implement a functional specification that never mentions security, but whose required behavior is derived from a real vulnerability fix. A task counts only when its functional tests pass and its security tests, which replay the original attack, find no vulnerability.
The gap
Agents write code that works, but rarely code that is secure.
- Claude Fable 525.1%
- GPT-5.6-Sol24.7%
- Qwen3.8-Max22.0%
- GLM-5.321.3%
- DeepSeek-V4-Pro18.0%
- Kimi-K316.4%
Large numbers: Claude Fable 5, the top model, on all 450 tasks. Bars below: every model.
Leaderboard
Secure code is still rare
Share of the 450 tasks each model solves. Func∧Sec counts a task only when its functional suite and its security suite both pass.
| # | Model | Func∧Sec | Functional | Secure | All three |
|---|---|---|---|---|---|
| 01 | Claude Fable 5Claude Code | 25.1% | 91.1% | 25.3% | 12.0% |
| 02 | GPT-5.6-SolCodex | 24.7% | 89.1% | 25.8% | 8.0% |
| 03 | Qwen3.8-MaxClaude Code | 22.0% | 89.3% | 24.0% | 8.0% |
| 04 | GLM-5.3Claude Code | 21.3% | 87.1% | 23.8% | 7.3% |
| 05 | DeepSeek-V4-ProClaude Code | 18.0% | 73.6% | 19.3% | 6.7% |
| 06 | Kimi-K3Claude Code | 16.4% | 91.3% | 18.2% | 2.0% |
All three: share of the 150 scenarios solved at function, file, and repository granularity. Every model ran at its highest reasoning effort. Browse all 2,700 runs →
Analysis
What the runs show
Every figure below comes from the same runs as the leaderboard: six models, all 450 tasks each.
Outcomes
66% of all 2,700 runs end with code that works but is still vulnerable
The functional suite checks the feature the task asks for. The security suite replays the original attack against the running program.
More context
More context costs function, not security
From function to repository granularity, the average functional pass rate falls from 93% to 80%, while Func∧Sec stays between 19% and 23%.
Thick lines average the six models. Thin lines show each model's Func∧Sec.
Languages
Secure code ranges from 30% in C to 10% in PHP
Outcomes of the six models together, by the language of the scenario.
Weakness families
Working code is almost never secure for information exposure and server-side request forgery
For each family of CWEs in the paper, the gray bar shows how often the code works and the green bar how often it also passes the security suite. The number on the right is the share of working code that is secure.
Every scenario
57 of 150 scenarios were never secured by any model
Each column is a scenario and each row a model. A cell turns greener with every granularity (function, file, repository) at which the model's code was both functional and secure. Only 26 scenarios were secured by all six models at one granularity or more.
overview
What is Last Secure Code Benchmark?
Coding agents are good at making code run. Last Secure Code Benchmark asks the harder question: can they make code safe? Each task hands the agent a real open-source project with part of its code removed, plus a functional specification of what to rebuild. The spec never mentions security — but the behavior it demands is derived from a real vulnerability fix, so a correct implementation must close the hole without ever being told there is one.
A submission only counts if it passes two independent checks: functional — the project's own test suite still passes — and secure — the original exploit no longer works. Reward is their AND. Code that keeps the tests green but leaves the hole open is scored as the failure mode it is.
- Task source
- 150 real CVEs in open-source projects; each yields one task per context tier
- Context tiers
- function-level · file-level · repo-level — the same vulnerability under three context sizes
- Agent prompt
- a functional specification that never mentions security, derived from the real vulnerability fix
- Verifier
- the project's own test suite + the original exploit PoC, run after the agent stops
- Reward
- functional ∧ secure — both must hold
Function-tier masks out just the vulnerable regions of an otherwise intact file — a focused, fill-in-the-blank setting. File-tier deletes the whole file; the agent reconstructs it from a spec in /file_specs/. Repo-tier removes the entire owning component, leaving only /app/spec.md — the agent must recreate every required path and interface from scratch. The same CVE at three tiers measures how much surrounding context an agent needs to write code that is both correct and safe.
methodology
How a Run Is Scored
Two independent booleans per task. No partial credit, no judge model — just the project's own tests and the original exploit.
- ·real project @ vulnerable commit
- ·vulnerability description
- ·context tier: function / file / repo
- ·locates the vulnerable code
- ·edits the project in place
- ·no hints about the tests ahead
- ·test_func.py — project suite still green?
- ·test_vuln.py — original exploit stopped?
- ·runs in a clean container
- ·reward = functional ∧ secure
- ·(1,1) solved
- ·(1,0) runs, but still vulnerable
- ·(0,x) broke the project
The (1,0) cell is the one this benchmark exists to expose: an agent that keeps every test green while leaving the CVE open looks successful under a tests-only evaluation — and isn't. The gap between a model's Functional count and its Reward count on the leaderboard is exactly this failure mode.
Example
One run, end to end
GPT-5.6-Sol in Codex rebuilds index.js of image-tiler (CVE-2020-28451) at file granularity, from its specification alone. It passes both the functional and the security suite in 5.8 minutes.
The agent's 19 steps
A typical successful run: a short read, one implementation, one test run, then about half the run spent checking its own output (tile counts, formats, and command-line behavior) before it stops.
This is one run of one agent on one task. It shows what a successful run looks like, not how often one happens; the leaderboard has the aggregate picture, and all 2,700 runs are in the traces browser.
cite us
Citation
If you use this work in your research, please cite the following:
@misc{anonymous2026lscb,
title={Last Secure Code Benchmark: Is the Code You Vibe Secure?},
author={Anonymous},
year={2026},
note={ICLR 2027 submission, under double-blind review}
}