Paper 1 of 5
Design Docs Are All You Need: An AI-native Machine-Learning Performance Tool
Problem
Existing ML performance‑modeling frameworks require constant refactoring because assumptions about models and hardware become invalid, and simply patching code incrementally accumulates technical debt and suffers from context‑window limited AI agents producing sub‑optimal code.
Approach
The authors make natural‑language design documents the single source of truth and store them as a DAG. An orchestrator agent walks the DAG in topological order, assigning a coding sub‑agent to regenerate each module from its doc. Design docs are written with step‑by‑step worked examples that serve as in‑context demonstrations for the generators. The system uses a minimal, recursively defined operator IR with symbolic SymPy cost expressions, supporting a fast analytical roll‑up mode for large sweeps and a slow modulo‑scheduling mode for fine‑grained studies. Leaves of the IR are priced for TPU resources, and a thin Python‑embedded tracing DSL builds graphs without manual Op construction.
Result
Regenerated implementations match hand‑audited reference models to round‑off precision, and a full library rebuild can be completed in 1.5 to 3 hours at a cost of roughly 100 USD, making continuous regeneration practical and economical.
Why it matters
Developers of ML performance‑modeling tools and teams using AI coding agents should care because the approach eliminates incremental technical debt and enables fast, cost‑effective full‑library regeneration.
Method details
- Reference model: DeepSeek‑V3 serving on a TPU pod slice.
- Library size: 50 design docs comprising 9,000 lines of specification prose.
- Regeneration time: between 1.5 and 3 hours for a full clean‑slate rebuild.
- API cost: approximately 100 USD per full rebuild using Claude Code.
- Cost share: about 20% of a weekly usage budget under a standard high‑tier plan.
- Operator IR leaves include MXU matmul tile, VMEM tile load, and ICI collective primitives.
Numbers
- regeneration time, 1.5 to 3 hours, full clean‑slate rebuild
- API cost, 100 USD, per complete rebuild using Claude Code
- budget share, 20%, of weekly usage budget under high‑tier plan
- design docs, 50, spanning 9,000 lines of specification prose
- library rebuild cost, 100 USD, compared to weekly budget
- regeneration time, 1.5 to 3 hours, compared to prior incremental patching
Limitations
The paper does not provide broader empirical performance evaluations beyond reproducing a single reference model, nor does it assess scalability to other hardware or model families.
Regenerated implementations reproduce hand-audited reference models-including DeepSeek-V3 serving on a TPU pod slice-to round-off precisionFound in the source text, word for word.
Picked because: Introduces an AI‑native performance‑modeling tool that engineers can use to profile and optimize ML workloads in production pipelines.