Last Secure Code BenchmarkLSCBench

Last Secure Code Benchmark

Is the Code You Vibe Secure?

450 tasks built from 150 real vulnerabilities — 7 languages, 115 CVEs, 24 CWEs. Each task asks the agent to implement a functional specification that never mentions security, but whose required behavior is derived from a real vulnerability fix. A task counts only when its functional tests pass and its security tests, which replay the original attack, find no vulnerability.

The gap

Agents write code that works, but rarely code that is secure.

91.1%of tasks end with working code
25.1%end with code that works and is secure
Works and secureWorks, still vulnerableDoes not work
  • Claude Fable 525.1%
  • GPT-5.6-Sol24.7%
  • Qwen3.8-Max22.0%
  • GLM-5.321.3%
  • DeepSeek-V4-Pro18.0%
  • Kimi-K316.4%

Large numbers: Claude Fable 5, the top model, on all 450 tasks. Bars below: every model.

Leaderboard

Secure code is still rare

Share of the 450 tasks each model solves. Func∧Sec counts a task only when its functional suite and its security suite both pass.

#ModelFunc∧SecFunctionalSecureAll three
01Claude Fable 5Claude Code
25.1%
91.1%
25.3%
12.0%
02GPT-5.6-SolCodex
24.7%
89.1%
25.8%
8.0%
03Qwen3.8-MaxClaude Code
22.0%
89.3%
24.0%
8.0%
04GLM-5.3Claude Code
21.3%
87.1%
23.8%
7.3%
05DeepSeek-V4-ProClaude Code
18.0%
73.6%
19.3%
6.7%
06Kimi-K3Claude Code
16.4%
91.3%
18.2%
2.0%

All three: share of the 150 scenarios solved at function, file, and repository granularity. Every model ran at its highest reasoning effort. Browse all 2,700 runs →

Analysis

What the runs show

Every figure below comes from the same runs as the leaderboard: six models, all 450 tasks each.

Outcomes

66% of all 2,700 runs end with code that works but is still vulnerable

The functional suite checks the feature the task asks for. The security suite replays the original attack against the running program.

Works and secureWorks, still vulnerableDoes not work
Claude Fable 5Claude Code
25%66%
GPT-5.6-SolCodex
25%64%
Qwen3.8-MaxClaude Code
22%67%
GLM-5.3Claude Code
21%66%
DeepSeek-V4-ProClaude Code
18%56%
Kimi-K3Claude Code
16%75%

More context

More context costs function, not security

From function to repository granularity, the average functional pass rate falls from 93% to 80%, while Func∧Sec stays between 19% and 23%.

0%25%50%75%100%FunctionFileRepositoryClaude Fable 5GPT-5.6-SolQwen3.8-MaxGLM-5.3DeepSeek-V4-ProKimi-K3Functional 80%Func∧Sec 23%

Thick lines average the six models. Thin lines show each model's Func∧Sec.

Languages

Secure code ranges from 30% in C to 10% in PHP

Outcomes of the six models together, by the language of the scenario.

Works and secureWorks, still vulnerableDoes not work
C23 scenarios
30%44%
Python22 scenarios
30%60%
JavaScript22 scenarios
27%60%
C++6 scenarios
19%69%
Java35 scenarios
18%79%
Go20 scenarios
14%70%
PHP22 scenarios
10%73%

Weakness families

Working code is almost never secure for information exposure and server-side request forgery

For each family of CWEs in the paper, the gray bar shows how often the code works and the green bar how often it also passes the security suite. The number on the right is the share of working code that is secure.

Memory safety24 scenarios · CWE-119 · CWE-125 · CWE-190 · CWE-416 · CWE-476 · CWE-787
Path and file handling24 scenarios · CWE-22 · CWE-434
Injection40 scenarios · CWE-77 · CWE-78 · CWE-79 · CWE-89 · CWE-94
Deserialization7 scenarios · CWE-502
Access control24 scenarios · CWE-284 · CWE-306 · CWE-352 · CWE-639 · CWE-862 · CWE-863
Input validation and limits16 scenarios · CWE-20 · CWE-400
Server-side request forgery7 scenarios · CWE-918
Information exposure8 scenarios · CWE-200

Every scenario

57 of 150 scenarios were never secured by any model

Each column is a scenario and each row a model. A cell turns greener with every granularity (function, file, repository) at which the model's code was both functional and secure. Only 26 scenarios were secured by all six models at one granularity or more.

NoneOne granularityTwoAll three
Claude Fable 5
GPT-5.6-Sol
Qwen3.8-Max
GLM-5.3
DeepSeek-V4-Pro
Kimi-K3
never secured · 57

overview

What is Last Secure Code Benchmark?

Coding agents are good at making code run. Last Secure Code Benchmark asks the harder question: can they make code safe? Each task hands the agent a real open-source project with part of its code removed, plus a functional specification of what to rebuild. The spec never mentions security — but the behavior it demands is derived from a real vulnerability fix, so a correct implementation must close the hole without ever being told there is one.

A submission only counts if it passes two independent checks: functional — the project's own test suite still passes — and secure — the original exploit no longer works. Reward is their AND. Code that keeps the tests green but leaves the hole open is scored as the failure mode it is.

450
total tasks
150
real-world CVEs
3
context tiers
2
checks per run
Task source
150 real CVEs in open-source projects; each yields one task per context tier
Context tiers
function-level · file-level · repo-level — the same vulnerability under three context sizes
Agent prompt
a functional specification that never mentions security, derived from the real vulnerability fix
Verifier
the project's own test suite + the original exploit PoC, run after the agent stops
Reward
functional ∧ secure — both must hold

Function-tier masks out just the vulnerable regions of an otherwise intact file — a focused, fill-in-the-blank setting. File-tier deletes the whole file; the agent reconstructs it from a spec in /file_specs/. Repo-tier removes the entire owning component, leaving only /app/spec.md — the agent must recreate every required path and interface from scratch. The same CVE at three tiers measures how much surrounding context an agent needs to write code that is both correct and safe.

methodology

How a Run Is Scored

Two independent booleans per task. No partial credit, no judge model — just the project's own tests and the original exploit.

CVE task
  • ·real project @ vulnerable commit
  • ·vulnerability description
  • ·context tier: function / file / repo
→
Agent patches
  • ·locates the vulnerable code
  • ·edits the project in place
  • ·no hints about the tests ahead
→
Verifier
  • ·test_func.py — project suite still green?
  • ·test_vuln.py — original exploit stopped?
  • ·runs in a clean container
→
Reward
  • ·reward = functional ∧ secure
  • ·(1,1) solved
  • ·(1,0) runs, but still vulnerable
  • ·(0,x) broke the project

The (1,0) cell is the one this benchmark exists to expose: an agent that keeps every test green while leaving the CVE open looks successful under a tests-only evaluation — and isn't. The gap between a model's Functional count and its Reward count on the leaderboard is exactly this failure mode.

Example

One run, end to end

GPT-5.6-Sol in Codex rebuilds index.js of image-tiler (CVE-2020-28451) at file granularity, from its specification alone. It passes both the functional and the security suite in 5.8 minutes.

The agent's 19 steps

A typical successful run: a short read, one implementation, one test run, then about half the run spent checking its own output (tile counts, formats, and command-line behavior) before it stops.

cve-2020-28451_file · GPT-5.6-Sol / CodexOpen in the traces browser →

Loading the trace…

This is one run of one agent on one task. It shows what a successful run looks like, not how often one happens; the leaderboard has the aggregate picture, and all 2,700 runs are in the traces browser.

cite us

Citation

If you use this work in your research, please cite the following:

@misc{anonymous2026lscb,
  title={Last Secure Code Benchmark: Is the Code You Vibe Secure?},
  author={Anonymous},
  year={2026},
  note={ICLR 2027 submission, under double-blind review}
}