Paper 1 of 5
LLM Post-Training as Brownfield Maintenance: An Industrial Perspective on Dataware Engineering
Problem
Existing post‑training pipelines suffered low usable supervision (only 3,412 valid solutions from 4,064 synthetic problems) and a fixed mixture budget, so simply adding more distilled data failed because new examples must displace existing ones, leading to diminishing returns and data debt.
Approach
The authors apply yield engineering to create a coverage‑first synthetic patch (FDS‑3K‑Cov) that selects high‑yield examples via testcase rectification and constraint injection, then replaces 3.5% of the rehearsal mixture with these examples. The patch is evaluated by running a single fixed checkpoint 16 times per benchmark and reporting bootstrap confidence intervals. This approach directly addresses zero‑sum mixture design and improves the fraction of usable supervision while staying within the fixed compute budget.
Result
Yield engineering raised accepted supervision from 3,412 to 9,697 solutions (2.84× increase) and improved downstream metrics: CodeForces pass@1 +2.59 points, pass@3 +3.11 points; LiveCodeBench v6 pass@1 +6.11 points, pass@3 +8.05 points, all statistically significant across 16 stochastic runs.
Why it matters
Industrial teams that maintain large language models under fixed compute and data budgets should care, as the yield‑engineered patch shows a practical way to improve code‑generation performance without expanding resources.
Method details
- CodeForces benchmark with 65 tasks and LiveCodeBench v6 with 175 tasks are used for evaluation
- Baseline (Base) uses a checkpoint trained on a 269,198‑example rehearsal mixture alone
- Distill‑3K produces 3,412 Tree‑sitter‑valid synthetic solutions without yield engineering
- FDS‑3K‑Cov is a 3,412‑example coverage‑first subset generated with yield engineering
- Patch size is 3.5% of the continued‑training mixture (9,697 of 278,895 examples)
- Each condition is evaluated 16 stochastic generations per task with 95% bootstrap CIs
Numbers
- CodeForces pass@1 +2.59 points vs Base
- CodeForces pass@3 +3.11 points vs Base
- LiveCodeBench v6 pass@1 +6.11 points vs Base
- LiveCodeBench v6 pass@3 +8.05 points vs Base
- Accepted supervision increased 2.84× (3,412 to 9,697)
- Patch replaces 3.5% of mixture (9,697 of 278,895 examples)
Limitations
The study is limited to code‑generation benchmarks (CodeForces and LiveCodeBench v6) and does not demonstrate applicability to other tasks or long‑term maintenance scenarios.
the yield-engineered patch improved CodeForces pass@1 by +2.59 points (+3.11 pass@3) and held-out LiveCodeBench v6 pass@1 by +6.11 (+8.05 pass@3)Found in the source text, word for word.
Picked because: Provides concrete brownfield post‑training workflows, mixture‑patch tooling and released artifacts for maintaining LLMs in production environments.