Pengfei Li, Naufal Suryanto, Sicheng Zhang and 1 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.CL
Problem
Existing evaluations only test knowledge or end-to-end agentic tasks and do not directly measure LLMs' ability to generate precise, executable CLI commands, so minor syntax errors invalidate execution.
Approach
The authors build KaliBench, a fine-grained NL‑to‑CLI benchmark with deterministic canonicalization and alias‑aware evaluation, and they add a multi‑stage verification pipeline that combines LLM‑based validation, sandboxed terminal execution, and human‑in‑the‑loop refinement; the benchmark also provides runtime‑free verifiable rewards for training models.
Result
No open‑weight model exceeds 42% exact‑command accuracy in the unrestricted setting, while supervised fine‑tuning and reinforcement learning with verifiable rewards bring an 8B model to performance comparable to a 685B MoE model; proprietary GPT‑5.6‑Sol reaches 61.68% EC on the full test set.
Why it matters
Researchers and practitioners building LLM‑driven cybersecurity agents should use KaliBench to benchmark and improve precise CLI generation, and to train models with verifiable, runtime‑free rewards.
Training: LoRA fine‑tuning with Unsloth on the RedSage‑Ins 8B backbone, three variants (SFT, RLVR with GRPO, SFT+GRPO), batch size 32, two epochs on a single H200 GPU (141 GB VRAM).
Inference: Evaluations run on an NVIDIA A100 multi‑GPU cluster using vLLM, temperature 0.2, max context 16,384 tokens (8,192 input, 8,192 output).
Baselines: Proprietary models GPT‑5.6‑Sol, Claude Opus 5, and Codex (GPT‑5.5) evaluated on 5,000 test queries in the unrestricted mode.
Numbers
Exact Correct (EC) ≤42%, compared to unrestricted open‑weight models
GPT‑5.6‑Sol EC (all 5K) 61.68%, compared to other proprietary models
Codex EC (all 5K) 51.68%, compared to GPT‑5.6‑Sol
Tool Accuracy (TA) average U→R 72.0% → 95.2%, showing improvement with tool hints
Hinted mode Opt‑F1 45.1% → 87.8%, showing large gain with tool hints
Hinted mode EC 22.3% → 73.1%, showing large gain with tool hints
Limitations
The paper does not demonstrate that open‑weight models can achieve high exact‑command accuracy without tool hints and does not evaluate long‑term robustness in live cybersecurity deployments.
no open-weight model exceeds 42% exact-command accuracy in the unrestricted settingFound in the source text, word for word.
Picked because: Provides a released benchmark (KaliBench) with runtime‑free verifiable rewards for evaluating LLM‑generated cybersecurity commands, giving engineers a concrete tool for agent verification.
Jichao Jiang, Cristian McGee, El Houcine Bergou and 2 others · abstract · pdf
quote unverifiedfigures checkedread: full textcs.LG
Problem
Full-parameter fine-tuning of LLMs incurs huge optimizer state memory, and existing fixes either compress state, drop first-order gradients, or change update geometry, which either degrade accuracy or reintroduce dense state.
Approach
TACO selects the sign of the largest-magnitude entry in each column of weight matrices to compute a column-wise one-sparse steepest-descent direction. It maintains a sparse gradient history stored in FP8 E4M3 values with int32 row indices, updating this history with an EMA and dynamic heavy hitters. The sparse history provides reliable winner selection without dense reconstruction. Updates are applied in-place and the dense gradient is released immediately, keeping persistent state near negligible. Scalar and vector parameters are handled by an auxiliary AdamW optimizer.
Result
TACO achieves comparable downstream accuracy while using far less memory; on SST-2 it reaches 94.2% accuracy with 276 tokens/second, versus 95.5% at 350 tokens/second for Adafactor and similar for other baselines. It enables full-parameter fine-tuning of 30‑32B models on a single 80 GB H100 GPU.
Why it matters
Practitioners who need to fine‑tune very large LLMs on limited GPU memory will benefit from TACO's drastic memory savings with minimal accuracy loss.
Method details
Evaluated on model sizes from 1.3B to 30B (OPT) and 8B to 32B (Qwen3) and up to 32B parameters across families OPT, Qwen3, Pythia, Llama, Mistral.
Datasets include SST-2, RTE, BoolQ, MultiRC, SQuAD, DROP, COPA, and CB.
Ablations cover optimizer step budget, EMA decay factor, heavy hitter threshold, and update sparsity.
Persistent optimizer state reduced by 174× (27.7 GB to 0.16 GB) relative to AdamW8bit.
Peak training memory reduced by 2.9× (80.6 GB to 27.5 GB) on OPT-13B.
Numbers
persistent optimizer state reduced 174× (27.7 GB to 0.16 GB) compared to AdamW8bit
peak training memory reduced 2.9× (80.6 GB to 27.5 GB) on OPT-13B
SST-2 accuracy 94.2% for TACO vs 95.5% for Adafactor, 95.7% for GaLore, 95.5% for FlashAdamW
throughput 276 tokens/second for TACO vs 350, 31, 547 tokens/second for Adafactor, GaLore, FlashAdamW respectively
full-parameter fine-tuning of 30‑32B‑parameter models on a single 80 GB H100 GPU
memory stays within 80 GB for OPT‑30B and Qwen3‑32B
Limitations
The paper does not establish performance on tasks beyond the eight downstream tasks evaluated.
The model could not quote the source for its claim, so nothing above has been checked against the paper. Read the abstract before trusting it.Citation check failed.
Picked because: Introduces TACO, a memory‑efficient optimizer for full‑parameter LLM fine‑tuning, enabling larger models on existing GPU infrastructure.
Xuan Zhang, Longtao Zheng, Cunxiao Du and 2 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.CL
Problem
Prior coding agents did not proactively compact context; the base model rarely invoked compact() before hitting the context limit, and simply adding length‑triggered compaction gave only limited gains.
Approach
AutoCompact adds a model‑invocable compact() action and trains the policy in two stages: first it collects on‑policy trajectories corrected by an online judge and fine‑tunes the model (SFT), then it applies outcome‑based reinforcement learning that rewards final task success while jointly optimizing coding and compaction decisions.
Result
AutoCompact reaches 39.6% pass rate on SWE‑bench Verified and 24.5% on SWE‑PolyBench Verified, which are absolute improvements of 9.2% and 5.0% over the base model respectively; outcome‑based RL adds another 7.4% and 2.8% gain over the SFT‑only checkpoint.
Why it matters
Developers of long‑horizon coding agents should adopt AutoCompact to achieve higher task success and better cost efficiency by learning when and how to compact context.
Method details
Base model: Qwen3-Coder-30B-A3B-Instruct
Training data: 1,052 judge‑corrected trajectories collected from 379 SWE‑rebench tasks
SFT: two epochs, batch size 8
RL: batch size 64, rollouts capped at 50 steps and 32K tokens
Context windows evaluated: 256K and 16K with forced‑compaction fallback
SWE‑bench Verified pass rate, 39.6%, AutoCompact vs Base 30.4%
SWE‑PolyBench Verified pass rate, 24.5%, AutoCompact vs Base 19.5%
AutoCompact‑SFT SWE‑bench Verified pass rate, 32.2%, vs SWE‑Compressor 31.0%
AutoCompact‑SFT SWE‑PolyBench Verified pass rate, 21.7%, vs SWE‑Compressor 20.1%
RL improvement SWE‑bench Verified, +7.4%, AutoCompact 39.6% vs AutoCompact‑SFT 32.2%
RL improvement SWE‑PolyBench Verified, +2.8%, AutoCompact 24.5% vs AutoCompact‑SFT 21.7%
Limitations
The paper does not establish performance beyond the evaluated benchmarks or under different model sizes, and it does not address scenarios where context management tools other than compact() are needed.
AutoCompact achieves the highest full‑run pass rates on both benchmarks, improving over Base by 9.2% on SWE‑bench Verified and 5.0% on SWE‑PolyBench Verified.Found in the source text, word for word.
Picked because: Presents AutoCompact, a method that learns when to compact context in long‑horizon coding agents, offering practical guidance for building self‑hosted LLM coding assistants.
Keyword‑matching benchmarks gave a false positive, crediting the 1B model with tool use it never performed, and the obvious fix of adding more tool‑SFT data failed because the web‑heavy phase erased the prior for the <|tool_call|> token.
Approach
The authors introduce a cheap, multi‑level diagnostic ladder (lenient keyword harness, verbatim‑reproduction check, first‑token probability probe, and embedding‑drift check) to expose the failure. They then apply a targeted SFT recipe that uses a diverse corpus, a learning rate five times higher than the original, and 2,202 training steps (~3.3 GPU‑hours) to restore the missing prior. The repair changes only the surrounding network while leaving the tied trigger‑token embedding 97.7% identical. This restores well‑formed tool calls without requiring massive token budgets.
Result
After repair, valid tool‑call emission on the full 269‑row corpus rises from 0.100 to 0.959 (the 600M achieves 0.926). On 238 unseen prompts the repaired 1B attains 0.536 versus 0.428 for the 600M (p = 0.004). The embedding‑drift check shows 97.7% of the bf16 table remains bit‑identical, confirming changes are localized.
Why it matters
Researchers and practitioners working with small language models should employ the cheap diagnostic ladder to verify tool‑use claims and can restore tool‑calling ability with minimal SFT effort.
Method details
661.6M‑parameter VectraYX‑600M decoder‑only model (65% code/technical text, no dedicated SFT)
1,109M‑parameter VectraYX‑1B decoder‑only model (web‑heavy curriculum, 6B‑token tool‑SFT phase)
Both share RMSNorm pre‑norm, GQA, SwiGLU, tied embeddings, RoPE/NoPE positional scheme, 2,048 sequence length, 32K BPE tokenizer, and special‑token layout
Targeted SFT uses a diverse corpus, learning rate 5× higher than baseline, 2,202 steps, ~3.3 GPU‑hours, three orders of magnitude fewer tokens than the failed tool‑SFT phase
Baseline lenient benchmark B4 scores: 0.660 (600M) vs 0.650 (1B)
Verbatim‑reproduction diagnostic: 6/6 successful calls for 600M, 0/4‑6 for 1B across checkpoints
Numbers
B4 tool‑use metric: 0.660 vs 0.650
Verbatim‑reproduction: 6/6 (600M) vs 0/4‑6 (1B)
First‑token probability on <|tool_call|>: 10^{-4}--10^{-5}
Valid emission on corpus: 0.100→0.959 (600M 0.926)
Unseen prompt pass rate: 0.536 vs 0.428 (p=0.004)
Embedding drift: 97.7% of bf16 table bit‑identical
Limitations
The benefit of a diverse corpus for suppressing over‑triggering remains a hypothesis due to seed sensitivity, and the repair is demonstrated on a single checkpoint with single‑seed results.
Both models over-trigger, rarely answering negative prompts without a call (0.09 for 600M, 0.17 for repaired 1B).Found in the source text, word for word.
Picked because: Offers a cheap diagnostic ladder to detect false‑positive tool‑use claims in small language models, directly supporting verification of LLM‑based automation.
Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing and 2 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.CL
Problem
Established text-to-SQL benchmarks only evaluate query generation and use single-table public datasets, which do not reflect the multi-table reasoning required in real enterprise data warehouses. The obvious fix of using larger public datasets still lacks the complex schema, actions, and ground‑truth consistency needed for realistic evaluation.
Approach
Argo-Bench simulates a large‑scale food‑delivery business and exports the simulated world into an Oracle E‑Business Suite warehouse with 235 tables and 7.5 billion rows. Tasks require agents to navigate this warehouse, reconstruct facts, and file concrete actions through a Python mission‑control library. A grader then scores each filing by its simulated consequences using nine grading modes. Reference solutions demonstrate solvability using only the warehouse. The framework isolates the skill of understanding and exploring enterprise data organization.
Result
Claude Opus 5.5 achieved the highest overall performance, solving 34.8 % of tasks and attaining an average score of 59.5 points. The next best models lag behind, with GPT‑6 Astra leading only in forecasting and Claude Sonnet 5.5 leading in compliance. More reasoning effort improves scores but with diminishing returns.
Why it matters
Researchers building data agents for enterprise environments should use Argo‑Bench to evaluate multi‑table reasoning, action filing, and end‑to‑end workflow capabilities.
Metrics: mean task score out of 100 and solved rate (tasks scoring ≥95).
Numbers
solved %, 34.8, Claude Opus 5.5 vs all models
score, 59.5, mean task score for Claude Opus 5.5
tables, 235, warehouse schema size
rows, 7.5 billion, warehouse data volume
tasks, 210, total number of Argo‑Bench tasks
steps, 82, average model calls per task for Claude Opus 5.5
Limitations
The benchmark covers only a single city, lacks foreign‑currency transactions, simulates only one year, and supports only the Oracle EBS format.
Argo-Bench goes beyond text-to-SQL: the agent files actions such as banning fraudulent accounts, allocating courier incentive budgets, or issuing back payFound in the source text, word for word.
Picked because: Introduces Argo‑Bench, an evaluation suite for data agents on enterprise‑scale workflows, with released artifacts that map to real‑world DevOps and data pipeline scenarios.