arXiv digest

Saturday

October 3, 2026

Up to five new AI papers a day, read in full where arXiv renders them and checked against the source.

Paper 1 of 5

KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards

Pengfei Li, Naufal Suryanto, Sicheng Zhang and 1 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.CL

Problem

Existing evaluations only test knowledge or end-to-end agentic tasks and do not directly measure LLMs' ability to generate precise, executable CLI commands, so minor syntax errors invalidate execution.

Approach

The authors build KaliBench, a fine-grained NL‑to‑CLI benchmark with deterministic canonicalization and alias‑aware evaluation, and they add a multi‑stage verification pipeline that combines LLM‑based validation, sandboxed terminal execution, and human‑in‑the‑loop refinement; the benchmark also provides runtime‑free verifiable rewards for training models.

Result

No open‑weight model exceeds 42% exact‑command accuracy in the unrestricted setting, while supervised fine‑tuning and reinforcement learning with verifiable rewards bring an 8B model to performance comparable to a 685B MoE model; proprietary GPT‑5.6‑Sol reaches 61.68% EC on the full test set.

Why it matters

Researchers and practitioners building LLM‑driven cybersecurity agents should use KaliBench to benchmark and improve precise CLI generation, and to train models with verifiable, runtime‑free rewards.

Method details
  • Dataset: 8,504 query‑command pairs covering 1,642 tools, 23 capability dimensions, and 5 security phases.
  • Training: LoRA fine‑tuning with Unsloth on the RedSage‑Ins 8B backbone, three variants (SFT, RLVR with GRPO, SFT+GRPO), batch size 32, two epochs on a single H200 GPU (141 GB VRAM).
  • Inference: Evaluations run on an NVIDIA A100 multi‑GPU cluster using vLLM, temperature 0.2, max context 16,384 tokens (8,192 input, 8,192 output).
  • Models: 14 general‑purpose models (e.g., Llama3.1‑Instruct 8B, Qwen‑3 8B/32B, Mistral‑3.2 24B, Gemma‑3 27B, DeepSeek‑V3.2 685B) and 7 cybersecurity‑oriented models (e.g., Foundation‑Sec‑Instruct 8B, Llama‑Primus‑Base 8B, RedSage‑Ins 8B).
  • Baselines: Proprietary models GPT‑5.6‑Sol, Claude Opus 5, and Codex (GPT‑5.5) evaluated on 5,000 test queries in the unrestricted mode.
Numbers
  • Exact Correct (EC) ≤42%, compared to unrestricted open‑weight models
  • GPT‑5.6‑Sol EC (all 5K) 61.68%, compared to other proprietary models
  • Codex EC (all 5K) 51.68%, compared to GPT‑5.6‑Sol
  • Tool Accuracy (TA) average U→R 72.0% → 95.2%, showing improvement with tool hints
  • Hinted mode Opt‑F1 45.1% → 87.8%, showing large gain with tool hints
  • Hinted mode EC 22.3% → 73.1%, showing large gain with tool hints
Limitations

The paper does not demonstrate that open‑weight models can achieve high exact‑command accuracy without tool hints and does not evaluate long‑term robustness in live cybersecurity deployments.

no open-weight model exceeds 42% exact-command accuracy in the unrestricted settingFound in the source text, word for word.

Picked because: Provides a released benchmark (KaliBench) with runtime‑free verifiable rewards for evaluating LLM‑generated cybersecurity commands, giving engineers a concrete tool for agent verification.

Paper 2 of 5

TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning

Jichao Jiang, Cristian McGee, El Houcine Bergou and 2 others · abstract · pdf

quote unverifiedfigures checkedread: full textcs.LG

Problem

Full-parameter fine-tuning of LLMs incurs huge optimizer state memory, and existing fixes either compress state, drop first-order gradients, or change update geometry, which either degrade accuracy or reintroduce dense state.

Approach

TACO selects the sign of the largest-magnitude entry in each column of weight matrices to compute a column-wise one-sparse steepest-descent direction. It maintains a sparse gradient history stored in FP8 E4M3 values with int32 row indices, updating this history with an EMA and dynamic heavy hitters. The sparse history provides reliable winner selection without dense reconstruction. Updates are applied in-place and the dense gradient is released immediately, keeping persistent state near negligible. Scalar and vector parameters are handled by an auxiliary AdamW optimizer.

Result

TACO achieves comparable downstream accuracy while using far less memory; on SST-2 it reaches 94.2% accuracy with 276 tokens/second, versus 95.5% at 350 tokens/second for Adafactor and similar for other baselines. It enables full-parameter fine-tuning of 30‑32B models on a single 80 GB H100 GPU.

Why it matters

Practitioners who need to fine‑tune very large LLMs on limited GPU memory will benefit from TACO's drastic memory savings with minimal accuracy loss.

Method details
  • Evaluated on model sizes from 1.3B to 30B (OPT) and 8B to 32B (Qwen3) and up to 32B parameters across families OPT, Qwen3, Pythia, Llama, Mistral.
  • Datasets include SST-2, RTE, BoolQ, MultiRC, SQuAD, DROP, COPA, and CB.
  • Baselines compared: AdamW8bit, Adafactor, GaLore, FlashAdamW.
  • Ablations cover optimizer step budget, EMA decay factor, heavy hitter threshold, and update sparsity.
  • Persistent optimizer state reduced by 174× (27.7 GB to 0.16 GB) relative to AdamW8bit.
  • Peak training memory reduced by 2.9× (80.6 GB to 27.5 GB) on OPT-13B.
Numbers
  • persistent optimizer state reduced 174× (27.7 GB to 0.16 GB) compared to AdamW8bit
  • peak training memory reduced 2.9× (80.6 GB to 27.5 GB) on OPT-13B
  • SST-2 accuracy 94.2% for TACO vs 95.5% for Adafactor, 95.7% for GaLore, 95.5% for FlashAdamW
  • throughput 276 tokens/second for TACO vs 350, 31, 547 tokens/second for Adafactor, GaLore, FlashAdamW respectively
  • full-parameter fine-tuning of 30‑32B‑parameter models on a single 80 GB H100 GPU
  • memory stays within 80 GB for OPT‑30B and Qwen3‑32B
Limitations

The paper does not establish performance on tasks beyond the eight downstream tasks evaluated.

The model could not quote the source for its claim, so nothing above has been checked against the paper. Read the abstract before trusting it.Citation check failed.

Picked because: Introduces TACO, a memory‑efficient optimizer for full‑parameter LLM fine‑tuning, enabling larger models on existing GPU infrastructure.

Paper 3 of 5

AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents

Xuan Zhang, Longtao Zheng, Cunxiao Du and 2 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.CL

Problem

Prior coding agents did not proactively compact context; the base model rarely invoked compact() before hitting the context limit, and simply adding length‑triggered compaction gave only limited gains.

Approach

AutoCompact adds a model‑invocable compact() action and trains the policy in two stages: first it collects on‑policy trajectories corrected by an online judge and fine‑tunes the model (SFT), then it applies outcome‑based reinforcement learning that rewards final task success while jointly optimizing coding and compaction decisions.

Result

AutoCompact reaches 39.6% pass rate on SWE‑bench Verified and 24.5% on SWE‑PolyBench Verified, which are absolute improvements of 9.2% and 5.0% over the base model respectively; outcome‑based RL adds another 7.4% and 2.8% gain over the SFT‑only checkpoint.

Why it matters

Developers of long‑horizon coding agents should adopt AutoCompact to achieve higher task success and better cost efficiency by learning when and how to compact context.

Method details
  • Base model: Qwen3-Coder-30B-A3B-Instruct
  • Training data: 1,052 judge‑corrected trajectories collected from 379 SWE‑rebench tasks
  • SFT: two epochs, batch size 8
  • RL: batch size 64, rollouts capped at 50 steps and 32K tokens
  • Context windows evaluated: 256K and 16K with forced‑compaction fallback
  • Baselines compared: Base, Fixed Compaction, SelfCompact, CompactionRL, SWE‑Compressor
Numbers
  • SWE‑bench Verified pass rate, 39.6%, AutoCompact vs Base 30.4%
  • SWE‑PolyBench Verified pass rate, 24.5%, AutoCompact vs Base 19.5%
  • AutoCompact‑SFT SWE‑bench Verified pass rate, 32.2%, vs SWE‑Compressor 31.0%
  • AutoCompact‑SFT SWE‑PolyBench Verified pass rate, 21.7%, vs SWE‑Compressor 20.1%
  • RL improvement SWE‑bench Verified, +7.4%, AutoCompact 39.6% vs AutoCompact‑SFT 32.2%
  • RL improvement SWE‑PolyBench Verified, +2.8%, AutoCompact 24.5% vs AutoCompact‑SFT 21.7%
Limitations

The paper does not establish performance beyond the evaluated benchmarks or under different model sizes, and it does not address scenarios where context management tools other than compact() are needed.

AutoCompact achieves the highest full‑run pass rates on both benchmarks, improving over Base by 9.2% on SWE‑bench Verified and 5.0% on SWE‑PolyBench Verified.Found in the source text, word for word.

Picked because: Presents AutoCompact, a method that learns when to compact context in long‑horizon coding agents, offering practical guidance for building self‑hosted LLM coding assistants.

Paper 4 of 5

Keyword Harnesses Fail Open: A Cheap Diagnostic Ladder for Tool-Use Claims in Small Language Models

Juan S. Santillana · abstract · pdf

quote verifiedfigures checkedread: full textcs.CL

Problem

Keyword‑matching benchmarks gave a false positive, crediting the 1B model with tool use it never performed, and the obvious fix of adding more tool‑SFT data failed because the web‑heavy phase erased the prior for the <|tool_call|> token.

Approach

The authors introduce a cheap, multi‑level diagnostic ladder (lenient keyword harness, verbatim‑reproduction check, first‑token probability probe, and embedding‑drift check) to expose the failure. They then apply a targeted SFT recipe that uses a diverse corpus, a learning rate five times higher than the original, and 2,202 training steps (~3.3 GPU‑hours) to restore the missing prior. The repair changes only the surrounding network while leaving the tied trigger‑token embedding 97.7% identical. This restores well‑formed tool calls without requiring massive token budgets.

Result

After repair, valid tool‑call emission on the full 269‑row corpus rises from 0.100 to 0.959 (the 600M achieves 0.926). On 238 unseen prompts the repaired 1B attains 0.536 versus 0.428 for the 600M (p = 0.004). The embedding‑drift check shows 97.7% of the bf16 table remains bit‑identical, confirming changes are localized.

Why it matters

Researchers and practitioners working with small language models should employ the cheap diagnostic ladder to verify tool‑use claims and can restore tool‑calling ability with minimal SFT effort.

Method details
  • 661.6M‑parameter VectraYX‑600M decoder‑only model (65% code/technical text, no dedicated SFT)
  • 1,109M‑parameter VectraYX‑1B decoder‑only model (web‑heavy curriculum, 6B‑token tool‑SFT phase)
  • Both share RMSNorm pre‑norm, GQA, SwiGLU, tied embeddings, RoPE/NoPE positional scheme, 2,048 sequence length, 32K BPE tokenizer, and special‑token layout
  • Targeted SFT uses a diverse corpus, learning rate 5× higher than baseline, 2,202 steps, ~3.3 GPU‑hours, three orders of magnitude fewer tokens than the failed tool‑SFT phase
  • Baseline lenient benchmark B4 scores: 0.660 (600M) vs 0.650 (1B)
  • Verb­atim‑reproduction diagnostic: 6/6 successful calls for 600M, 0/4‑6 for 1B across checkpoints
Numbers
  • B4 tool‑use metric: 0.660 vs 0.650
  • Verbatim‑reproduction: 6/6 (600M) vs 0/4‑6 (1B)
  • First‑token probability on <|tool_call|>: 10^{-4}--10^{-5}
  • Valid emission on corpus: 0.100→0.959 (600M 0.926)
  • Unseen prompt pass rate: 0.536 vs 0.428 (p=0.004)
  • Embedding drift: 97.7% of bf16 table bit‑identical
Limitations

The benefit of a diverse corpus for suppressing over‑triggering remains a hypothesis due to seed sensitivity, and the repair is demonstrated on a single checkpoint with single‑seed results.

Both models over-trigger, rarely answering negative prompts without a call (0.09 for 600M, 0.17 for repaired 1B).Found in the source text, word for word.

Picked because: Offers a cheap diagnostic ladder to detect false‑positive tool‑use claims in small language models, directly supporting verification of LLM‑based automation.

Paper 5 of 5

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing and 2 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.CL

Problem

Established text-to-SQL benchmarks only evaluate query generation and use single-table public datasets, which do not reflect the multi-table reasoning required in real enterprise data warehouses. The obvious fix of using larger public datasets still lacks the complex schema, actions, and ground‑truth consistency needed for realistic evaluation.

Approach

Argo-Bench simulates a large‑scale food‑delivery business and exports the simulated world into an Oracle E‑Business Suite warehouse with 235 tables and 7.5 billion rows. Tasks require agents to navigate this warehouse, reconstruct facts, and file concrete actions through a Python mission‑control library. A grader then scores each filing by its simulated consequences using nine grading modes. Reference solutions demonstrate solvability using only the warehouse. The framework isolates the skill of understanding and exploring enterprise data organization.

Result

Claude Opus 5.5 achieved the highest overall performance, solving 34.8 % of tasks and attaining an average score of 59.5 points. The next best models lag behind, with GPT‑6 Astra leading only in forecasting and Claude Sonnet 5.5 leading in compliance. More reasoning effort improves scores but with diminishing returns.

Why it matters

Researchers building data agents for enterprise environments should use Argo‑Bench to evaluate multi‑table reasoning, action filing, and end‑to‑end workflow capabilities.

Method details
  • Warehouse: 235 tables, 7.5 billion rows modeled on Oracle EBS 12.2.
  • Tasks: 210 data‑science and analytics tasks across five business areas.
  • Models evaluated: 14 frontier and open‑weight models, including GPT‑6 Astra and Claude Opus 5.5.
  • Evaluation setup: Inspect AI 0.3.263 on GKE, each agent in a gVisor sandbox with 25 preinstalled Python libraries, limited to 500 model turns.
  • Baseline comparison: proprietary models (GPT‑6 series, Claude series) versus open‑weight models (Kimi K3, GLM 5.3 Flash, DeepSeek V4.1 Flash, Qwen 3.8 Max).
  • Metrics: mean task score out of 100 and solved rate (tasks scoring ≥95).
Numbers
  • solved %, 34.8, Claude Opus 5.5 vs all models
  • score, 59.5, mean task score for Claude Opus 5.5
  • tables, 235, warehouse schema size
  • rows, 7.5 billion, warehouse data volume
  • tasks, 210, total number of Argo‑Bench tasks
  • steps, 82, average model calls per task for Claude Opus 5.5
Limitations

The benchmark covers only a single city, lacks foreign‑currency transactions, simulates only one year, and supports only the Oracle EBS format.

Argo-Bench goes beyond text-to-SQL: the agent files actions such as banning fraudulent accounts, allocating courier incentive budgets, or issuing back payFound in the source text, word for word.

Picked because: Introduces Argo‑Bench, an evaluation suite for data agents on enterprise‑scale workflows, with released artifacts that map to real‑world DevOps and data pipeline scenarios.