In brief

LLMs trained on code have progressed at an unusual pace: over four years, the best systems went from 28% to over 90% on standard reference benchmarks. This trajectory conceals a more nuanced reality. Benchmarks are saturating, data contamination distorts scores, and controlled studies show that developers using AI assistants sometimes produce more vulnerable code, not less. The arrival of autonomous coding agents in 2024–2025 opened a new front, with results that are both impressive and difficult to interpret.


What models actually do

A code model is an LLM trained on corpora that include large quantities of source code. Its initial task was completion: continuing an existing code fragment. The scope then expanded to generation from natural language descriptions, bug fixing, cross-language translation, and code explanation.

The foundational model in this domain is Codex (Chen et al., OpenAI, 2021): a GPT-3 fine-tuned on 54 million public GitHub files. It reached 28.8% on the HumanEval benchmark — the reference of its era — and became the technical basis of GitHub Copilot at its commercial launch in 2022.

Since then, progress has been rapid. CodeLlama (Meta, 2023) raised that score to 53.7% on HumanEval with its 34-billion-parameter Python variant. DeepSeek-Coder (2024) reaches 90.2% with its 236-billion-parameter version — on par with GPT-4o. These models are available as open-weight releases, which significantly disrupted the market for proprietary solutions.


HumanEval: a test that has become too simple

HumanEval is a set of 164 Python problems with unit tests, designed by OpenAI in 2021. Its primary metric, pass@k, measures the proportion of problems solved when the model generates k different solutions. It is simple, reproducible, and now insufficient.

The problem is twofold. First, saturation: above 90% success for the best models, the measure no longer discriminates. Second, contamination: since HumanEval has been public since 2021, its solutions have very likely ended up in the training corpora of recent models. Scores therefore partly reflect memorization, not generalization.

To test this hypothesis, HumanEval+ (Liu et al., NeurIPS 2023) and EvoEval (Xia et al., 2024) automatically generate mutated variants of the original problems. The result is telling: GPT-4 loses approximately 20% performance on these variants, suggesting that a non-negligible fraction of its good scores stems from memorization.

In short: when a benchmark becomes public, models trained afterwards may “memorize” its answers without necessarily understanding them. A 90% score on HumanEval does not mean 90% understanding — it means at best that a good portion of the solutions were seen during training.


SWE-bench: the real difficulty

To move beyond toy problems, SWE-bench (Jimenez et al., Princeton/Chicago, 2023) proposes 2,294 real GitHub issues drawn from popular repositories like Django, scikit-learn, or sympy. The task: fix a real bug in a real codebase. This is considerably more complex than an isolated Python problem.

In 2024, the majority of systems resolved fewer than 5% of issues. Progress since then has been notable: the best agents approach 50% on SWE-bench Verified (Chowdhury et al., 2024), a manually verified subset of 500 issues guaranteeing ground truth quality. This still falls well short of 100%, and the hardest issues — those requiring deep architectural understanding of the codebase — remain out of reach.


Coding agents in 2024–2025

The qualitative leap of this period lies not in the models themselves but in how they are used. Coding agents combine an LLM with tools: shell access, file navigation, test execution. The agent can iterate on its own output until tests pass.

Devin (Cognition AI, March 2024) was presented as the first “AI software engineer” capable of working autonomously in a sandboxed environment. Its initial score of 13.86% on SWE-bench was contested and revised. SWE-agent (Yang et al., Princeton, 2024) is the open-source academic alternative, reaching 12.5% on SWE-bench lite.

On the tooling side, Cursor (Anysphere), Windsurf (Codeium), and Claude Code (Anthropic, 2025) renewed the market for augmented development environments. The SWE-agent + Claude 3.5 Sonnet combination is among the best-documented configurations on SWE-bench Verified in agentic mode.


The security problem

This is the most critical issue for professional use. Pearce et al. (2022, IEEE S&P) tested Codex on 89 scenarios involving known vulnerabilities — SQL injection, buffer overflow, path traversal. Result: 40% of code generated in sensitive contexts contained vulnerabilities.

The most counter-intuitive study comes from Perry et al. (2023, ACM CCS): in a controlled experiment, participants with access to GitHub Copilot produced significantly more vulnerable code than the group without an assistant, and expressed higher confidence in their incorrect code. The assistant made developers faster and more confident — not necessarily more secure.

Sandoval et al. (2023, USENIX Security) nuance this result: in their protocol, Copilot users did not produce more vulnerabilities than the control group. But the control group consisted of security experts — a selection bias that limits the conclusion’s scope.

The question remains unsettled. What is certain: code assistants do not actively detect vulnerabilities in the code they generate, and the increased confidence they induce can be a risk factor.

In short: an assistant that makes developers faster does not make them safer — the documented effect is the opposite. Generated code contains on average 40% vulnerabilities in sensitive contexts, and the confidence users place in it exceeds the actual quality. For production use, generated code must go through the same controls as a junior developer’s first commit.


The licensing question

Codex and GitHub Copilot sometimes reproduce GPL-licensed code verbatim. The Doe v. GitHub lawsuit (filed in 2022) alleges a violation of open-source license terms through reproduction without attribution. The legal question remains open: is a model trained on GPL code itself subject to copyleft?

The BigCode Project (Hugging Face + ServiceNow) responded by building The Stack v2 with an opt-out mechanism allowing developers to withdraw their code from the training corpus. This does not resolve the problem for models already trained on earlier versions.


Which benchmark to judge a code model

ContextRecommendationWhy
Fast internal evaluationHumanEval+ or EvoEvalMutations of the original problems — resistant to contamination
Realistic production evaluationSWE-bench VerifiedReal GitHub issues, complete codebase, audited ground truth
Reasoning test (vs memorization)CRUXEvalMeasures output prediction, not generation
Comparing complete agents (LLM + tools)SWE-agent or TerminalBench-2Includes shell, navigation, iteration over tests

Understanding or memorization?

The underlying question behind all these debates: do code models genuinely understand programs, or do they recognize memorized syntactic patterns?

CRUXEval (Gu et al., 2024) tests a different capability: predicting the output of a given program, which requires reasoning about its execution rather than generating a similar one. Scores are markedly lower than corresponding HumanEval scores, suggesting that code generation exploits syntactic memorization more than semantic understanding of programs.

This is an important limitation for complex use cases: a model can complete a function with correct syntax without having modeled what the program actually does.


Key takeaways

  • The best models exceed 90% on HumanEval, but this score is partly biased by training data contamination; mutated variants reveal significant performance degradation.
  • SWE-bench, which measures resolution of real GitHub issues, remains far harder: the best agents approach 50% on the verified subset, but complex tasks remain out of reach.
  • Controlled studies show that code assistants can increase the rate of vulnerabilities produced and developer confidence in incorrect code — a documented risk, not a hypothetical one.
  • The legal question of open-source licenses remains unresolved: no established case law on GPL code reproduction by models.
  • The difference between syntactic generation and semantic understanding is measurable (CRUXEval): models reason less well about program execution than they generate code.