Paper 1 of 5
SWE-Serve: Benchmarking Agentic Engineering For Production Inference Serving
Problem
Existing benchmarks either ignore inference engineering or focus only on isolated kernel generation and performance optimization, so they do not evaluate the full production inference serving stack.
Approach
SWE-Serve constructs 53 repository‑grounded tasks from recent SGLang production changes, each paired with hidden functional, regression, and where applicable end‑to‑end serving tests and calibrated performance gates. Agents receive a task instruction and the codebase at a base commit, produce a patch, and the verifier scores the patch. The benchmark runs in Harbor 0.13.1 sandboxes on CPU or a single H100 GPU, enforcing a closed‑book setting. Evaluation uses the mini‑SWE‑agent v2.4.3 harness, with no‑op and oracle controls plus adversarial verifier review to ensure task validity.
Result
The best‑performing configuration reaches a 75% mean pass@1, while the lowest model achieves 35% mean pass@1, showing a 40‑point spread. End‑to‑end serving tests reject roughly one‑third of patches that pass other tests (45.9% under the verifier versus 69.4% when E2E tests are excluded). Mean task wall‑clock time is 43.3 minutes, compared with 34.3 minutes on DeepSWE, and only 2.4% of attempts fail due to execution limits.
Why it matters
Researchers and engineers developing agentic code‑generation systems for inference serving should care, as SWE‑Serve quantifies the gap between local task completion and production‑correctness.
Method details
- Evaluated 11 models including Claude Opus 5, Claude Sonnet 5, GPT‑5.6 Sol, Luna, Terra, Kimi K3, DeepSeek V4 Flash 0731, GLM‑5.2, Gemini 3.6 Flash, Laguna S 2.1, and Inkling S.
- Ran 31 model‑effort configurations (low, medium, high, xhigh, max) across the 53 tasks, each repeated three times.
- Agents were limited to 350 steps or 210 minutes per task, with individual commands timing out after 120 seconds.
- Benchmark uses Harbor 0.13.1 to provide CPU or single H100 GPU resources for each sandbox.
- Baseline comparisons include DeepSWE benchmark metrics for wall‑clock time and task distribution.
Numbers
- mean pass@1, 75%, best‑performing configuration (Claude Opus 5 and GPT‑5.6 Sol)
- pass@1 range, 40 percentage points, across 11 models (75% to 35%)
- execution‑limit failures, 2.4% (42/1,749), across top‑per‑model configurations
- mean task wall‑clock time, 43.3 minutes, versus DeepSWE 34.3 minutes
- E2E test rejection rate, 45.9%, under verifier versus 69.4% with E2E tests excluded
- cost per task, $0.33, $17.40, across models
Limitations
The paper does not establish how agents would perform in open‑book settings or in real‑world deployment beyond the benchmark.
Across 11 models and 31 model‑effort configurations, the best‑performing configuration achieves 75% mean pass@1.Found in the source text, word for word.
Picked because: Introduces SWE-Serve, a released benchmark suite for evaluating agents on real‑world production inference engineering tasks, giving engineers concrete metrics and artifacts.