Narges Ahmadi, Yubo Jiao, Jônatas Augusto Manzolli and 2 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.AI
Problem
Travel behavior research often separates digital data collection from predictive modeling, evaluating each stage in isolation. Simply merging the datasets does not address the lack of coordinated, auditable workflows across collection, processing, and prediction.
Approach
The paper implements a three‑agent workflow: a Data Collection Agent runs a chatbot‑administered, image‑augmented stated‑preference survey; a Data Processing Agent transforms raw responses into analysis‑ready features; a Data Modeling Agent fits both conventional models (MNL, logistic regression, random forest) and LLM‑based predictors. The agents exchange standardized data objects and maintain versioned records, enabling a forward execution pass from survey design to demand prediction. LLMs are prompted under multiple framings (Expert vs Role‑Play, Base vs Richer context) and extended with persona, few‑shot, and vision inputs. Predictions are evaluated on a held‑out test set at the respondent level to avoid leakage.
Result
Random forest reached 69.6% accuracy, the top zero‑shot LLM slightly outperformed it at 69.9%, and adding visual context raised performance to 71.5%. Habitual travel information consistently improved results, Expert framing generally beat Role‑Play, and few‑shot prompting gave gains that stabilized after a small number of examples.
Why it matters
Travel behavior researchers and engineers building multimodal LLM pipelines should note that coordinated multi‑agent workflows can match or exceed conventional machine‑learning baselines without task‑specific fine‑tuning.
Method details
Nine locally deployed LLMs ranging from 2 to 35 billion parameters were evaluated.
Random forest baseline achieved 69.6% five‑class accuracy.
Best text‑only zero‑shot LLM reached 69.9% five‑class accuracy without task‑specific fitting.
Vision‑based configuration using weather images attained 71.5% five‑class accuracy.
Traditional models included an MNL estimated in Biogeme, logistic regression, and random forest.
Numbers
five‑class accuracy, 69.6%, random forest baseline
five‑class accuracy, 69.9%, best text‑only zero‑shot LLM
five‑class accuracy, 71.5%, best vision‑based configuration
LLMs, 9, locally run models ranging from 2 to 35 billion parameters
weather scenarios, 5, predefined conditions used in the survey
Limitations
The study implements only a single forward execution pass and does not evaluate iterative refinements; it also does not establish causal effects, focusing solely on predictive performance.
Random forest achieved 69.6% five-class accuracy, while the best text-only zero-shot LLM reached 69.9% without task-specific fitting.Found in the source text, word for word.
Picked because: Introduces a three‑agent workflow (chatbot data collection, structured processing, prediction) with released code and dataset, giving engineers a concrete pipeline for building self‑hosted LLM‑driven data collection systems.
Yucheng Jiang, Zora Zhiruo Wang, Ruishi Chen and 1 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.CL
Problem
Existing methods assume a given task or a single workflow and only produce step-level summaries, which cannot handle multi‑threaded, interleaved goals in naturalistic computer‑use traces. The obvious fix of applying a single workflow induction fails because it cannot disentangle concurrent activities or produce a structured hierarchical model.
Approach
Task Model Induction (TMI) first grounds low‑level mouse and keyboard events in their visual context and groups them into semantic actions and activities. It then performs latent task induction, assigning each activity to an existing task or creating a new one to keep tasks consistent across application shifts. For each discovered task, TMI independently induces an objective model via recursive goal decomposition and a procedure model using sequencing, for‑each, and while constructs. Finally, a model reconciliation step merges the two trees, aligning objectives with control‑flow operators at each node. The result is a single task model where every node carries both an objective and a control‑flow operator.
Result
Intrinsic evaluation on 38 HumanWork sessions shows TMI recovers interleaved tasks with 0.974 agreement against ground‑truth groupings and reconstructs 74.9% of observed execution steps, far exceeding the strongest workflow induction baseline. On procedural fidelity, TMI achieves 74.9% step description accuracy and 88.5% operator correctness, again well above baselines. Extrinsically, skills derived from TMI improve held‑out task accuracy by 30.0% over the strongest baseline.
Why it matters
Researchers building computer‑use agents and organizations needing auditable workflow models should care because TMI provides reusable hierarchical task models from raw interaction traces.
LLM judges for intrinsic evaluation: gpt‑5.5 and claude‑sonnet‑5
Numbers
Task grouping agreement, 0.974, against ground‑truth
Step description accuracy, 74.9%, versus baseline 30.3%
Operator correctness, 88.5%, versus baseline 52.7%
Held‑out task accuracy improvement, 30.0%, over strongest baseline
Objective node boundary correction, 64.5%, by procedure model
Procedure node boundary correction, 21.9%, by objective model
Limitations
The paper does not establish performance on completely unseen domains or with non‑English interfaces.
Our method substantially outperforms both baselines on procedural fidelity, achieving 74.9% step description accuracy versus 30.3% baseline and 88.5% operator correctness versus 52.7% under the gpt-5.5 judgeFound in the source text, word for word.
Picked because: Presents methods to derive symbolic task models from passive computer‑use traces, releasing tools that let engineers automate workflow capture and audit existing software processes.
Yizhe Chi, Wenyi Li, Deyao Hong and 7 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.AI
Problem
Existing benchmarks do not isolate the ability to design training algorithms; they are won by data collection or hyperparameter tuning, which does not address algorithmic redesign.
Approach
AI4AI-Bench provides ten frozen research repositories, each with a starting model and a fixed evaluation metric. An LLM agent gets four hours on a B300 GPU to rewrite the repository's training algorithm code. The patched code is then executed from scratch for up to twelve hours, keeping at most three checkpoints. A hidden evaluator scores the final model on the task-specific metric, mapping scores to a 0 to 1 scale where 0 is an uninformative model, 0.1 is the original repository algorithm, and 1.0 is the task optimum.
Result
Across 29 configurations of six systems on all ten tasks the mean score is 0.166, and the best system reaches 0.250. Submissions that modify the learning side average 0.226 versus 0.126 for those that only change run‑side aspects. Increasing reasoning effort raises the minority that reach the learning side from 8% to 64% and lifts the mean score from 0.094 to 0.196.
Why it matters
Benchmark designers and researchers building LLM agents for algorithmic improvement should care, as the suite quantifies current limits of automated training‑algorithm design.
Method details
OpenR1 task uses supervised fine‑tuning on Qwen2.5‑Coder‑1.5B‑Instruct evaluated with LiveCodeBench.
DDPO task uses diffusion RL on Stable Diffusion v1.5 evaluated with aesthetic score.
Model Soup task uses weight averaging over 72 CLIP checkpoints evaluated with ImageNet‑V2 top‑1.
Agents run on a single B300 GPU with a 4‑hour rewrite budget and up to 12‑hour rerun time.
Numbers
mean score, 0.166, baseline 0
best system score, 0.250, baseline 0.1 (repository algorithm)
learning‑side average, 0.226, compared to run‑side average 0.126
minority reaching learning side, 8%, increased to 64% with more effort
mean score with low effort, 0.094, increased to 0.196 with high effort
Limitations
The study does not establish that agents can close a large fraction of the gap to the task optimum or that performance scales with larger compute budgets.
the mean score is $0.166$, and the best system reaches $0.250$Found in the source text, word for word.
Picked because: Provides AI4AI‑Bench, a benchmark suite for evaluating LLM agents in algorithmic design and recursive self‑improvement, with artifacts that help verify and compare agent tooling.
Cheng Xu, Nan Yan, Liming Chen and 1 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.AI
Problem
Prior audits relied on a single greedy decode ledger, which fabricated capability changes due to inference batching. Simply decoding more carefully does not fix it because server‑side batching or nondeterminism still causes verdict flips.
Approach
The paper replaces the single‑decode ledger with a per‑problem exact test against a pooled baseline while controlling the false‑discovery rate. Baseline replicates from the frozen control are used to construct a separate measured null for each statistic. This null requires no additional experiments and is derived from existing multi‑arm study data. The audit then compares self‑training and distillation arms against these calibrated nulls. Statistical significance is assessed with regression and FDR‑controlled tests.
Result
External distillation improves problems the base model rarely reaches, while three self‑training variants do not. A regression rejects the asymmetry as a by‑product of distillation's larger overall gain (p < 10^{-8}). Self‑training corrupts baseline‑solved problems at rates well above the measured floor. The single‑decode ledger falsely reports a capability loss on an untrained model, with an expansion rate of 0.280.
Why it matters
Researchers auditing language‑model self‑improvement should adopt per‑problem nulls and FDR control to avoid spurious capability claims.
Method details
Model: Qwen3‑8B with rank‑32 LoRA self‑training.
Datasets: MATH‑500, AIME 2025/2026, and a difficulty band of problems.
Training: three rounds of rank‑32 LoRA self‑training, plus majority‑vote, distillation, STaR, SFT, and policy‑gradient arms.
Baseline: frozen control model run through identical pipeline.
Ablation: serializing inference requests versus batching to isolate flip cause.
Numbers
expansion rate, 0.280, compared to frozen control
regression p‑value, < 10^{-8}, compared to null hypothesis of no asymmetry
Limitations
The paper does not establish any positive gain from self‑training and provides no universal null that works without baseline replicates.
a regression rejects this asymmetry as a by-product of distillation's larger overall gain ($p < 10^{-8}$).Found in the source text, word for word.
Picked because: Describes a systematic audit of LLM self‑improvement (LoRA fine‑tuning) against a frozen control, releasing scripts that engineers can use to detect spurious gains in their own models.
Yiyang Feng, Biddut Sarker Bijoy, Niranjan Balasubramanian and 1 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.AI
Problem
LLM agents solved each task from scratch and did not improve with experience; existing skill induction methods often transferred unreliably and could even harm the agent, and the obvious fix of using whole-task (task-level) summaries does not work because task-level skills mostly reduce performance below the no‑memory baseline.
Approach
The paper systematically varies two axes of skill induction: level (task‑level vs subtask‑level) and format (text vs code). Agents store induced skills in a skill memory and retrieve them for new tasks. Six conditions (Task, Task+Text, Task+Code, Subtask, Subtask+Text, Subtask+Code) are compared against a no‑memory baseline. A skill utility score is computed from specificity and abstractness to predict transfer success without executing tasks. Experiments are run on three long‑horizon benchmarks with multiple LLM architectures.
Result
Subtask‑level skills consistently improve agent performance above the no‑memory baseline, whereas task‑level skills usually hurt performance; text‑format skills transfer better than code‑format skills. The proposed skill utility score correlates consistently with observed task success across all conditions. For example, on the Gemini-3.1-Pro model subtask‑level code skills achieve the highest score (77.5) compared to subtask‑level without skills (68.3).
Why it matters
Researchers and engineers building LLM agents should prefer subtask‑level, text‑based skill induction and can use the skill utility score as a cheap diagnostic before deployment.
Models: three MoE models (Qwen3-235B-A22B, GPT-OSS-120B, Nemotron-Super-120B), dense Qwen3 4B/8B/14B/32B, Gemma-3 4B/12B/27B, and Gemini-3.1-Pro.
Induction prompts: full prompt (L3) and two reduced versions (L1 minimal instruction, L2 instruction plus one demonstration) used for ablation.
Comparison conditions: Task, Task+Text, Task+Code, Subtask, Subtask+Text, Subtask+Code; Task and Subtask alone carry no memory.
Evaluation metrics: task success (average benchmark score), latency (wall‑clock time per task), and dependency (compute spent on growing context).
Baselines: agents without any skill memory (no‑memory baseline) and ablations with reduced induction prompts.
Numbers
AppWorld average Subtask+Text 16.5 vs Task+Text 9.3
OfficeBench average Subtask+Text 29.7 vs Task+Text 25.3
KramaBench average Subtask+Text 55.3 vs Task+Text 49.3
Gemini-3.1-Pro Subtask+Code 77.5 vs Subtask none 68.3
Task-level skills mostly reduce the agent’s performance below its no-memory baseline while subtask-level skills raise it above on average
Skill utility score correlates consistently with task success when skills are transferred
Limitations
The study is limited to three long‑horizon benchmarks and fixed skill‑memory mechanics; it does not cover other agentic settings such as computer use, evolving memories, or step‑level evaluation.
Task-level skills mostly reduce the agent’s performance below its no-memory baseline while subtask-level skills raise it above on averageFound in the source text, word for word.
Picked because: Offers a controlled study of cross‑task skill transfer in LLM agents, including released benchmarks and analysis code useful for building reliable agent skill libraries.