In brief

The “harness” is the code surrounding an LLM — prompts, retrieval, memory, orchestration logic. Lee et al. (Stanford/MIT/KRAFTON, 2026) show that this code can produce a 6× performance gap on the same benchmark, with the same model. Meta-Harness is an automated system for optimizing these harnesses: a coding agent explores the complete history of past attempts and proposes successive iterations. Results across three domains: +7.7 pts in text classification with 4× fewer tokens, +4.7 pts in mathematical reasoning transferable to 5 unseen models, and first place among Claude Haiku 4.5 agents on TerminalBench-2.


The harness matters as much as the model

A harness refers to all the code that determines what a model receives at each step of a task: few-shot example selection, retrieval strategy, prompt format, memory management, branching logic. This code is rarely optimized systematically — it is designed manually, through ad hoc iterations, often modifying multiple dimensions simultaneously without isolating effects.

The authors document a 6× gap between the lowest-performing and best harness on the same benchmark, with the same base model. This gap exceeds in magnitude what most model choices produce. The problem of manual design is structural: the harness space is too large for exhaustive human exploration, and existing text optimization methods (OPRO, TextGrad, AlphaEvolve) do not directly apply — they optimize strings or discrete parameters, not arbitrary programs.

In short: changing the harness can multiply performance by 6 — a larger effect than changing the model. But this harness is arbitrary code (Python programs of 100-1000 lines), not a prompt or simple parameter. No existing method knew how to optimize it systematically.


How Meta-Harness works

The propose-evaluate-log loop

Meta-Harness is an outer-loop system. At each iteration, a coding agent (the proposer) generates one or more candidate harnesses as single-file Python files. Each harness is evaluated on a set of examples, and its results — source code, scores, and complete execution traces — are archived in a dedicated directory on the filesystem. The loop restarts, with the proposer having access to the full history.

A typical run evaluates between 40 and 109 candidate harnesses over 20 to 40 iterations, with 3 candidates proposed per iteration in text classification. The evaluated harness is a program between 100 and 1,000 lines of code, which can modify task-specific prompting, retrieval, memory, and orchestration. The base model remains frozen — only the code surrounding it varies.

The filesystem as feedback channel

The central design choice is to store feedback on the filesystem rather than in a prompt. Each evaluated harness contributes a directory containing its source code, scores, and execution traces (prompts, tool calls, model outputs, state updates). The proposer queries this filesystem via standard terminal tools — grep, cat — rather than ingesting everything into context.

This enables feedback at a qualitatively different scale from existing methods. According to the authors, a single evaluation token can produce up to 10 million tokens of diagnostic information — three orders of magnitude beyond the feedback budgets of TextGrad (0.015 MTok/iter) or AlphaEvolve (0.022 MTok/iter), versus 10.0 MTok/iter for Meta-Harness.

The coding agent proposer

The proposer is Claude Code with Opus 4.6, guided by a minimal domain-specific skill describing where to write new harnesses, how to inspect previous harnesses and their traces, and which files it can or cannot modify. It has no parent selection rule: it is free to inspect any previous harness. The population maintains a Pareto frontier over evaluated harnesses — when multiple objectives are active (accuracy and context cost), the search freely explores this frontier without hard-coded constraints.


Results

Text classification

Across three datasets (USPTO-50k, Symptom2Disease, LawBench), Meta-Harness achieves 48.6% average accuracy versus 40.9% for ACE, i.e., +7.7 points. Context efficiency is simultaneously better: 11.4K tokens on average versus 50.8K for ACE (+4.5×) and 28.5K for MCE (+2.5×). Meta-Harness occupies the dominant position on the accuracy–context Pareto frontier.

In terms of convergence speed, Meta-Harness reaches the final performance of OpenEvolve and TTT-Discover as early as the 4th evaluation (out of 40 total), and surpasses these methods by more than 10 points by end of search. On 9 out-of-distribution datasets not seen during search, the gain holds: 73.1% versus 70.2% for ACE (+2.9 points), with 7.3K tokens versus 11.7K.

Mathematical reasoning

A single harness discovered by Meta-Harness improves pass@1 accuracy by +4.7 points on average across 200 IMO-level problems, on 5 base models not seen during the search — from 34.1% (without retriever) to 38.8%. Individual gains range from +1.6 to +8.7 points depending on the model. Meta-Harness outperforms BM25 retrieval by +1.3 points on average (38.8% vs 37.5%), without an additional dense encoder.

Agentic coding (TerminalBench-2)

With Claude Haiku 4.5 as the base model, Meta-Harness achieves 37.6% on TerminalBench-2, ranking first among all reported Haiku 4.5 agents — ahead of Goose (35.5%), Terminus-KIRA (33.7%), Mini-SWE-Agent (29.8%), and Claude Code (27.5%). With Opus 4.6, Meta-Harness reaches 76.4%, second behind ForgeCode (81.8%) — the authors note they could not reproduce this score from ForgeCode’s public code, suggesting unpublished components.


Why it works: complete traces

The decisive ablation

The determining element is not the optimization loop itself, but access to raw execution traces. The ablation (Table 3) is clear: scores only → median 34.6 / best 41.3; scores + LLM summaries → median 34.9 / best 38.7; full Meta-Harness with traces → median 50.0 / best 56.7. The number of runs exceeding the zero-shot baseline goes from 26 (scores) to 23 (scores+summaries) to 39 (complete traces).

Summaries do not recover the missing signal — they can even be harmful by compressing diagnostically useful details. The raw richness of traces is irreplaceable.

Non-Markovian behavior

The proposer does not optimize via a Markov chain. During the TerminalBench-2 run, it reads a median of 82 files per iteration (range 69–99), split across 41% harness source code, 40% execution traces, 6% score/summary files, and 13% other. It references more than 20 previous candidates per iteration — not just the immediate parent. This global exploration of history is emergent behavior, not prescribed by the skill.

Causal reasoning

Qualitative analysis of the TerminalBench-2 run illustrates structured reasoning. At iteration 3, the proposer identifies a confound between prompt modifications and structural bugfixes, isolates the causal changes, and pivots toward an additive modification (environment bootstrapping) after 6 consecutive regressions. At iteration 10, it references results from a prior separate search run, performing explicit cross-run transfer. The math harness discovered autonomously merges two successful research lineages identified in the history.


When to use Meta-Harness versus the alternatives

ContextRecommendationWhy
Stable task, limited API budget (< $1k)Manual optimisation, or OPRO/TextGradMeta-Harness costs 10 MTok per iteration — poor ROI on a low-value task
High-volume production task (millions of calls a month)Meta-HarnessA 7.7-point gain, or 4× fewer tokens, amortises quickly at scale
New domain, no documented human harnessExploratory Meta-HarnessThe broad history compensates for the missing human baseline
Recent or little-known base modelMeta-Harness with an Opus proposerThe system discovers model-specific patterns that transfer to 5 unseen models

Limitations and open questions

A single proposer tested. All experiments use Claude Code with Opus 4.6 as the proposer. The effect of proposer variation (Sonnet, Haiku, other model) remains unmeasured.

High computational cost. Meta-Harness generates ~10 MTok/iter, versus 0.002 to 0.026 MTok/iter for the compared methods. The API cost of a complete search run is not explicitly reported in the paper — only the cost of evaluated harnesses is available, not that of the proposer itself.

Evaluation methodology on TerminalBench-2. The same benchmark serves for both search and final evaluation, due to the absence of a separate split (benchmark too small and costly). The authors verify the absence of overfitting through manual inspection and regex audits, but this methodology remains exposed to a critique of non-independent evaluation.

Co-evolution not addressed. The base model remains frozen throughout the search. The authors identify co-evolution of the harness and model weights as a natural direction, not covered by this work.

Skill quality as the primary lever. The authors note that the quality of the skill text (the description given to the proposer) is the dominant factor for search quality — more important than the number of iterations or population size. This sensitivity to the proposer’s prompt introduces a dependency on a parameter that is difficult to specify optimally.


Key takeaways

  • The harness (code surrounding the LLM) can produce a 6× gap on the same benchmark with the same model; Meta-Harness automates exploration of this space.
  • Access to raw execution traces is the key ingredient: LLM-generated summaries achieve performance comparable to scores alone, and well below complete traces (median 34.9 vs 50.0).
  • The system achieves +7.7 pts in text classification with 4× fewer tokens than the ACE baseline, and produces a harness transferable to 5 unseen models in mathematical reasoning.
  • The proposer exhibits emergent non-Markovian behavior: it consults the entire history (median 82 files/iteration) rather than following only the last iteration.
  • Main limitations: a single proposer tested, computational cost of 10 MTok/iter, and absence of independent evaluation on TerminalBench-2.