Paper 1 of 5
Turbo Harness: Instance-Adaptive Harness Optimization
Problem
Existing harness optimization produces a single global harness applied uniformly, which may be suboptimal for individual task instances, and naïvely searching per instance is infeasible.
Approach
Turbo Harness reuses the exhaust (traces, reflections, evaluations) from a completed global harness search to build a structured playbook. A lightweight LLM harness editor is trained with reinforcement learning to condition on both the task instance and the playbook, generating a patch that adapts the global harness into an instance‑specific harness. At inference the editor produces the patch, which is applied (or falls back to the global harness) before the frozen execution model runs. The editor is trained with GRPO using task performance as reward, while the execution model remains frozen.
Result
On SWE‑smith‑MR with a frozen Claude Haiku 4.5 executor, the globally optimized Meta‑Harness achieves 50.7% pass rate. An untrained 9B editor without the playbook reaches 51.3%, with the playbook 50.0%, and RL‑trained without the playbook 49.3%. Combining RL training with playbook conditioning raises the pass rate to 64.0%, a 14.7‑point improvement over the RL‑trained editor without the playbook. Stronger frozen editors achieve 55.3% (Sonnet‑4.5) and 61.3% (Opus‑4.6), but the RL‑trained Qwen3.5‑9B matches the best.
Why it matters
Researchers building self‑improving agents and harness optimization pipelines should care because Turbo Harness shows that lightweight, RL‑trained editors can extract and apply global search knowledge to boost per‑instance performance without retraining the main model.
Method details
- Harness editor: Qwen3.5-9B full‑parameter fine‑tuned with FSDP and GRPO.
- Execution models: frozen LLMs per benchmark (Qwen3.5‑9B for agentic tasks, Claude Haiku 4.5 and Gemini 3.7 Flash for coding, Claude Sonnet 4.5 for TB2.1).
- Benchmarks: seven tasks (ALFWorld, ScienceWorld, DBBench, WebShop, SWE‑smith‑MR, SWE‑bench Verified, Terminal‑Bench‑2.1).
- Baselines: Meta‑Harness, Default (mini‑swe‑agent or basic tool‑calling), Terminus‑Kira, Terminus‑2, ReAct, Self‑Refine, Reflection, Harness‑R1.
- Ablations: compare untrained editor, playbook only, RL only, and combined RL+playbook; also vary editor model (Qwen3.5‑9B, Sonnet‑4.5, Opus‑4.6).
Numbers
- Pass rate, 64.0%, compared to Meta‑Harness 50.7%
- Pass rate, 61.3%, compared to Meta‑Harness 50.7%
- Pass rate, 51.3%, compared to Meta‑Harness 50.7%
Limitations
The paper does not claim to evaluate generalization to unseen repositories or to budgets beyond those tested.
Combining RL training with playbook conditioning raises the pass rate to 64.0%, a 14.7-point improvement over the same RL-trained editor without the playbook.Found in the source text, word for word.
Picked because: Introduces Turbo Harness, a concrete framework for instance-adaptive harness optimization that can be directly applied to improve LLM agent tooling.