arXiv digest

Tuesday

September 8, 2026

Up to five new AI papers a day, read in full where arXiv renders them and checked against the source.

Paper 1 of 5

Design Docs Are All You Need: An AI-native Machine-Learning Performance Tool

Samuel Kushnir, Kimia Noorbakhsh, Kavya Sreedhar and 6 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.PL

Problem

Existing ML performance‑modeling frameworks require constant refactoring because assumptions about models and hardware become invalid, and simply patching code incrementally accumulates technical debt and suffers from context‑window limited AI agents producing sub‑optimal code.

Approach

The authors make natural‑language design documents the single source of truth and store them as a DAG. An orchestrator agent walks the DAG in topological order, assigning a coding sub‑agent to regenerate each module from its doc. Design docs are written with step‑by‑step worked examples that serve as in‑context demonstrations for the generators. The system uses a minimal, recursively defined operator IR with symbolic SymPy cost expressions, supporting a fast analytical roll‑up mode for large sweeps and a slow modulo‑scheduling mode for fine‑grained studies. Leaves of the IR are priced for TPU resources, and a thin Python‑embedded tracing DSL builds graphs without manual Op construction.

Result

Regenerated implementations match hand‑audited reference models to round‑off precision, and a full library rebuild can be completed in 1.5 to 3 hours at a cost of roughly 100 USD, making continuous regeneration practical and economical.

Why it matters

Developers of ML performance‑modeling tools and teams using AI coding agents should care because the approach eliminates incremental technical debt and enables fast, cost‑effective full‑library regeneration.

Method details
  • Reference model: DeepSeek‑V3 serving on a TPU pod slice.
  • Library size: 50 design docs comprising 9,000 lines of specification prose.
  • Regeneration time: between 1.5 and 3 hours for a full clean‑slate rebuild.
  • API cost: approximately 100 USD per full rebuild using Claude Code.
  • Cost share: about 20% of a weekly usage budget under a standard high‑tier plan.
  • Operator IR leaves include MXU matmul tile, VMEM tile load, and ICI collective primitives.
Numbers
  • regeneration time, 1.5 to 3 hours, full clean‑slate rebuild
  • API cost, 100 USD, per complete rebuild using Claude Code
  • budget share, 20%, of weekly usage budget under high‑tier plan
  • design docs, 50, spanning 9,000 lines of specification prose
  • library rebuild cost, 100 USD, compared to weekly budget
  • regeneration time, 1.5 to 3 hours, compared to prior incremental patching
Limitations

The paper does not provide broader empirical performance evaluations beyond reproducing a single reference model, nor does it assess scalability to other hardware or model families.

Regenerated implementations reproduce hand-audited reference models-including DeepSeek-V3 serving on a TPU pod slice-to round-off precisionFound in the source text, word for word.

Picked because: Introduces an AI‑native performance‑modeling tool that engineers can use to profile and optimize ML workloads in production pipelines.

Paper 2 of 5

Does Your Agent's Memory Survive a Model Upgrade? A Controlled Study of Memory Portability

Ankit Goyal, Jaideep Ray · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Memory migrations break when a model upgrade changes how old notes are interpreted or when mixed embedding versions are used, causing retrieval failures and repair to fail without the original evidence.

Approach

The study evaluates four memory formats-raw long‑context (LC‑RAW), retrieval‑augmented generation (RAG), compressed natural‑language notes (NOTES), and fixed‑schema knowledge graphs (KG‑fixed), across model upgrades using synthetic histories. Two open‑weight models (Llama‑3.1‑8B‑Instruct and Qwen2.5‑7B‑Instruct‑1M) serve as readers, and a writer‑swap experiment isolates model coupling. An embedding model upgrade (bge‑large‑en v1.0 to v1.5) is tested with full re‑embedding versus a 50/50 mixed index. Diagnostic decomposition attributes performance loss to construction information loss or retrieval failures. Repair experiments compare store‑only NOTES repair to raw‑history repair under budget constraints.

Result

Fixed‑schema KG‑fixed accuracy changed by only +0.0004 ± 0.0020 after a writer swap, while NOTES accuracy shifted asymmetrically by +9.91 or -13.28 percentage points. Full re‑embedding of the index yielded an 11.90‑point accuracy gain versus a 4.96‑point gain for the mixed index. Retrieval failures accounted for 81% of the RAG deficit, and raw‑history repair succeeded in 34 of 48 cases whereas store‑only NOTES repair never reached the 90% target.

Why it matters

Developers of agent memory systems and RAG pipelines should care because model upgrades can silently degrade memory portability, and the paper quantifies how different storage formats and migration strategies affect performance.

Method details
  • Two open‑weight models: Llama‑3.1‑8B‑Instruct and Qwen2.5‑7B‑Instruct‑1M (both <10B parameters).
  • 48 synthetic histories with randomized answer codes and exact scoring.
  • Memory formats: LC‑RAW, RAG, NOTES, KG‑fixed.
  • Embedding migration: BAAI/bge‑large‑en v1.0 → v1.5, both 1024‑dimensional vectors.
  • Evaluation contrasts H1, H2, H4a, H7a with one‑sample one‑sided t‑tests and Holm correction.
  • Bootstrap validation with 10,000 resamples.
Numbers
  • KG‑fixed accuracy change, +0.0004 ± 0.0020, compared against writer swap baseline
  • NOTES accuracy shift, +9.91 pp, compared against writer swap direction
  • NOTES accuracy shift, -13.28 pp, compared against opposite writer swap direction
  • Full re‑embed gain, 11.90 pp, compared against old embedding index
  • Mixed index gain, 4.96 pp, compared against old embedding index
  • Raw‑history repair success, 34/48 cases, compared against store‑only NOTES repair
Limitations

The experiments cover only two open‑weight models, a single cross‑family migration, and synthetic histories, so findings may not generalize to other model families or real‑world data.

fixed-schema structures transfer reliably, with KG-fixed accuracy changing by only $+0.0004 \pm 0.0020$ following a writer swap.Found in the source text, word for word.

Picked because: Provides a concrete study and methodology for preserving and validating agent memory across model upgrades, aiding LLM‑agent reliability.

Paper 3 of 5

The History Is the Detector: Executing CVE Patch History, End-to-End

Qiushi Wu, Kevin Eykholt, Youngja Park and 4 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.CR

Problem

Public vulnerability records are documented mainly for human inspection, so automated scanners cannot reuse the fix knowledge, and naive static signatures either over‑match or miss variants of the same flaw.

Approach

Bugstone‑E2E first mines fixing commits to synthesize reusable detection rules (CVE‑to‑Skill pipeline). Phase A indexes target code with Tree‑sitter and enumerates call sites matching rule anchors, then discards benign candidates with deterministic heuristics. Phase B invokes LLM‑based agents that use the rule’s verification contract to inspect taint flows and return structured verdicts. Phase C re‑triages the surviving candidates in isolated execution environments to obtain runtime evidence. Phase D generates scope‑checked patches and validates them with two‑sided differential tests.

Result

Bugstone‑E2E identified 2,710 fixing commits, built 1,033 rules, and when run on 14 programs produced runtime evidence for 644 findings, demonstrating that CVE history can be turned into an executable detection and repair workflow.

Why it matters

Security tool developers and organizations can leverage Bugstone‑E2E to automate detection of recurring vulnerabilities and obtain verified patches, reducing manual effort and improving remediation confidence.

Method details
  • Uses the cvelistV5 dataset of 19,325 high‑severity CVEs from 2022‑2026.
  • Phase B default model is gpt‑5‑mini; experiments also include GPT‑5.5, GPT‑5.6‑sol, Claude Sonnet 4.6, and Kimi‑K2.5.
  • Constructed 1,033 detection rules across 56 CWE families, packaged into 172 skills.
  • Compared against three agentic security harnesses: OpenAI/Codex security‑scan harness, Anthropic defending code reference harness, and Visa Vulnerability Agentic Harness.
  • Ablation: lightweight heuristics in Phase A remove benign sites without any LLM calls.
Numbers
  • high‑severity CVEs, 19,325, from 2022 to 2026
  • fixing commits, 2,710, identified by Bugstone‑E2E
  • detection rules, 1,033, across 56 CWE families
  • skills, 172, packaged from the rules
  • programs scanned, 14, across which runtime evidence was collected
  • findings with runtime evidence, 644, produced by the system
Limitations

The system only covers recurring patterns anchored to security‑relevant APIs, cannot detect one‑off design flaws or configuration errors, and Phase C runtime validation requires a reproducible build; Phase D is a prototype that does not guarantee regression‑free patches.

When applied across 14 programs, it produced runtime evidence for 644 findings.Found in the source text, word for word.

Picked because: Presents an end‑to‑end system that automatically extracts and reuses CVE patch history, offering a practical security‑automation artifact for DevOps.

Paper 4 of 5

CONTINUITY: Security-Context Contracts for Composable LLM Agent Controls

Chris Zheng, Geng Yang · abstract · pdf

quote verifiedfigures checkedread: full textcs.CR

Problem

Prior LLM agent systems assumed that individually correct security mechanisms would yield an end‑to‑end secure system, but security‑context discontinuities caused context to be dropped, widened, or reinterpreted across component boundaries, breaking security guarantees. The obvious fix of adding more isolated checks fails because the composition of those checks does not preserve guarantees across the full instruction‑to‑effect path.

Approach

The paper introduces CONTINUITY, a framework that models each component with an assume‑guarantee contract and carries authenticated security context across transitions. It uses signed root grants, provenance commitments, role‑bound transition receipts, bounded typed releases, transformation witnesses, and effect‑bound execution permits. A deterministic verifier checks root signatures, provenance, releases, transform relations, monotonicity, and finality before allowing an effect. The architecture places the planner outside the trusted path and enforces security at gateway, adapter, and finality layers. The reference implementation provides a Python 3.11+ verifier and a fault‑injection suite covering 32 fault classes across four domains.

Result

The full CONTINUITY configuration commits no harmful external effect in all 2,560 attack instances, contains every one of the 128 fault‑domain classes, completes all 700 benign tasks, and escalates all 200 ambiguous tasks, achieving 100% lifecycle correctness.

Why it matters

Security engineers and designers of LLM‑driven agents should care because CONTINUITY shows that explicit contracts are required to preserve security guarantees across composable components, preventing harmful effects that persist despite individual controls.

Method details
  • Python 3.11+ reference implementation; core module 1,523 lines and experiment script 1,697 lines
  • 30 regression tests and publication scripts are included
  • Fault suite instantiates 32 fault classes in four domains with 20 parameterized instances per fault‑domain pair
  • Evaluation runs 2,560 attack instances, 700 benign tasks, and 200 ambiguous tasks across seven configurations
  • Seven system configurations produce 24,220 system‑scenario runs
Numbers
  • Effect ASR, 0.0%, CONTINUITY vs 100.0% for Pass‑through
  • Classes, 128, CONTINUITY vs 0 for Pass‑through
  • Benign completion, 100%, CONTINUITY vs 100% for all configs
  • Ambiguous escalation accuracy, 100%, CONTINUITY vs 0% for Pass‑through
Limitations

The prototype does not provide a mechanically verified implementation and its integrity guarantees rely on trusted validators rather than semantic correctness of the underlying facts.

Continuity commits no harmful effect in 2,560 attack instances and contains every one of the 128 fault-domain classes.Found in the source text, word for word.

Picked because: Defines composable security‑context contracts for LLM agents, delivering a framework engineers can adopt to verify safe agent interactions.

Paper 5 of 5

Trace2Tower: Transition-Aware EigenTrace Induction of Multi-Level Skills for LLM Agents

Jiazheng Sun, Boyu Yang, Binhao Yuan and 2 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Current LLM agents rely on shallow trajectory retrieval and flat skill summarization, which ignore temporal dependencies and outcome‑conditioned topology of behavior. Simply adding more data or flattening trajectories does not capture the necessary execution order and success signals.

Approach

Trace2Tower first abstracts raw interaction steps into canonical events and replaces task‑specific entities with typed arguments. It then builds a directed graph that integrates semantic similarity, transition counts, and outcome tendencies without manual mixing weights. A parameter‑free contrastive affinity transformation combines success and failure evidence, after which a normalized Laplacian is formed and spectrally decomposed per connected component. The resulting eigenvectors define procedure and strategy groups that populate a hierarchical skill tower. Verifier‑guided feedback continuously refines the tower by editing graph edges and selecting relevant skills.

Result

On ALFWorld Trace2Tower reaches 87.31% success with an average of 10.35 steps and only 0.26 invalid actions, while on WebShop it attains 50.67% exact success. Removing transition evidence reduces success to 70.15%, and removing outcome or contrastive evidence reduces it to 73.88%. Cross‑model transfer shows a GPT‑5.4 authored tower raises Flash user success from 65.67% to 88.06% and Pro user success from 79.10% to 85.82%. Verifier‑guided updates increase ALFWorld success from 80.42% to 84.58% with TF‑IDF and to 83.75% with embeddings.

Why it matters

Researchers building LLM agents that need reusable hierarchical skills and temporal reasoning should consider Trace2Tower for more robust task mastery and efficient experience reuse.

Method details
  • GPT‑5.4 is used as the Skill Author and plan rewriter; DeepSeek‑V4‑Flash is the default Skill User
  • Training data: 1,240 No‑Skill trajectories from 310 ALFWorld training tasks and 400 trajectories from 100 WebShop training tasks
  • Baselines compared: No‑Skill, Expert‑Crafted, ExpeL (Zhao et al. 2024), SkillX (Wang et al. 2026a), Trace2Skill +Combined and +Error
  • Ablations: removing transition drops success to 70.15% and reduces hierarchy to 19 procedures/76 strategies; removing outcome or contrastive drops success to 73.88%
  • Interaction horizon is 20 steps; each task uses 4 rollouts, temperature 0, max 512 output tokens, 4,096‑dim embeddings processed in batches of 16
Numbers
  • ALFWorld success 87.31% vs baselines
  • ALFWorld steps 10.35 average
  • ALFWorld invalid actions 0.26
  • WebShop exact success 50.67%
  • Transition ablation success 70.15%
  • Outcome/contrastive ablation success 73.88%
Limitations

The paper does not claim to handle tasks beyond the evaluated benchmarks or to guarantee performance with substantially larger horizons.

Trace2Tower abstracts step-level interactions into canonical events, constructing a unified graph governed by semantic compatibility, transition dynamics, and outcome evidence.Found in the source text, word for word.

Picked because: Offers the Trace2Tower framework for extracting hierarchical skills from LLM agent traces, enabling actionable skill‑based tooling.