TSR Desk · science · 5 October 2026, 07:00 UTC
Rethinking the Evaluation of Harness Evolution for Agents
- What
- Rethinking the Evaluation of Harness Evolution for Agents
- Who
- arxiv.org
- When
- 5 October 2026, 04:00 UTC
- Category
- Science
- Primary source
- https://arxiv.org/abs/2607.12227
- What is not known
- This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.
Following prior work, we experiment on Terminal-Bench 2.1 and find that automatic harness evolution fails to outperform simple test-time scaling methods both with and without test cases, and exhibits limited generalization. It comes from a paper posted to arXiv on 5 October 2026. Harness evolution is an iterative search procedure that repeatedly evaluates and revises candidate harnesses used for LLM agents using task feedback. We revisit the evaluation of such automatic harness evolution procedures and identify two fundamental issues in the protocol. First, prior work does not compare these approaches with simple task-level search baselines under matched feedback and inference budgets. Second, prior work searches for harness configurations using verification signals (e.g., unit test cases) drawn from the same benchmarks on which it reports the final performance of the evolved harnesses, violating the standard separation between training and test data. To address this, we compare automatic harness evolution with simple test-time scaling and discovery baselines under comparable feedback and inference budgets, and evaluate evolved harnesses on held-out tasks to assess generalization. However, we find that long-horizon games are a promising setting for automatic harness evolution, as they are difficult enough to leave headroom, rely heavily on adaptation to out-of-distribution dynamics, and provide granular feedback by design. In these settings, task-specific harness evolution improves over the search baseline by 80.0% on ARC-AGI-3 and by 10.9% on EdgeBench under matched budgets. Together, these findings highlight the need for matched-budget baselines and held-out evaluation to distinguish genuine harness improvements from benchmark-specific search and overfitting, and point to a more careful characterization of when automatic harness evolution is actually useful. Our code is available at https://github.com/rethinking-harness-evolution.
Why it counts
Following prior work, we experiment on Terminal-Bench 2.1 and find that automatic harness evolution fails to outperform simple test-time scaling methods both with and without test cases, and exhibits limited generalization. In these settings, task-specific harness evolution improves over the search baseline by 80.0% on ARC-AGI-3 and by 10.9% on EdgeBench under matched budgets.
Sources
Primary source: primary source
What is not known
This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.
No clip. The article still stands.