arXiv digest

Saturday

August 15, 2026

Up to five new AI papers a day, read in full where arXiv renders them and checked against the source.

Paper 1 of 3

MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and Coordination

Saisha Shetty, Satvik Tripathi, Austin Lin and 6 others · abstract · pdf

quote unverifiedfigures checkedread: abstract onlycs.AI

Problem

Monolithic LLM prompting was not working as expected, and the obvious fix of manual prompt engineering is not feasible.

Approach

MARC replaces monolithic LLM prompting with deterministic multi-agent orchestration for clinical reasoning, coordinating role-specialized agents for extraction, reasoning, answer generation, and evaluation. The framework includes a Decomposer module that generates task-specific agent prompts from a plain-language description. MARC supports both API-based and local CPU-compatible deployments and is entirely configurable via YAML. The framework is designed to be model-agnostic, interpretable, and accessible to clinical domain experts without programming expertise. MARC enables stage-wise failure attribution through explicit context passing and traceable intermediate outputs.

Result

The full framework is available at https://github.com/Penn-RAIL/MARC-v1, indicating the completion of the MARC framework.

Why it matters

Clinical domain experts without programming expertise should care about this framework as it is designed to be accessible to them.

Method details
  • The framework is model-agnostic
  • MARC is entirely configurable via YAML
  • The framework supports both API-based and local CPU-compatible deployments
Limitations

The paper does not establish any limitations.

The model could not quote the source for its claim, so nothing above has been checked against the paper. Read the abstract before trusting it.Citation check failed.

Picked because: This paper earns a slot because it presents MARC, a multi-agent framework that coordinates role-specialized agents for clinical reasoning, which is a concrete example of LLM agents and tool use.

Paper 2 of 3

TraVEL: Trajectory-Guided Video Embedding Learning for Driving-Video Retrieval

Yi-Chung Chen, Philip Jacobson, Tom Lampo and 6 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.CV

Problem

General-purpose multimodal embedding models often rely on shortcuts from static scene context and struggle to distinguish motion-centric events. Model scale alone does not resolve the driving-domain gap. The obvious fix of scaling up models does not work as shown by the results of off-the-shelf models.

Approach

The method works by first fine-tuning Qwen3-VL-Embedding on paired clips and reasoning traces from nuReasoning using an InfoNCE objective. Then it introduces TraVEL, a motion-aware fine-tuning framework that uses ego-trajectory similarity as a reward within Group Relative Policy Optimization. TraVEL combines physically grounded supervision with efficient embedding-based search. The approach involves supervised fine-tuning and trajectory-aware optimization. The trajectory reward is computed only for the selected video pairs after group construction.

Result

The results show that TraVEL improves motion-centric retrieval across model scales. At 2B, TraVEL raises longitudinal and lateral mAP by 9.8 and 4.7 points over SFT. The approach preserves the instance-level gains of SFT at both model scales.

Why it matters

This work is important for autonomous-driving data engines as it enables efficient retrieval of relevant clips from large-scale driving logs, which can support the development and evaluation of downstream driving systems.

Method details
  • The dataset used is nuReasoning which contains 20,000 long-tail driving clips.
  • The model architectures used are Qwen3-VL-Embedding, CLIP4Clip, InternVideo2-Stage2, and Cosmos-Embed1.
  • The model sizes used are 2B and 8B parameters for Qwen3-VL-Embedding.
  • The evaluation metrics used are recall at rank, median rank, and mean rank for instance-level retrieval, and average precision and mean average precision for motion-centric retrieval.
Numbers
  • R@1 of 8.7 for Qwen3-VL-Embed + TraVEL at 2B
  • Longitudinal mAP of 55.7 for Qwen3-VL-Embed + TraVEL at 2B, compared to 45.9 for SFT
  • Lateral mAP of 35.7 for Qwen3-VL-Embed + TraVEL at 2B, compared to 31.0 for SFT
  • R@1 of 10.1 for Qwen3-VL-Embed + TraVEL at 8B
Limitations

The current evaluation uses the available subset of nuReasoning and a retrieval pool of 1,715 videos, which may not establish scalability and generalization.

TraVEL thus combines physically grounded supervision with efficient embedding-based search.Found in the source text, word for word.

Picked because: This paper earns a slot because it proposes TraVEL, a trajectory-guided video embedding learning method for driving-video retrieval, which addresses the topic of retrieval and provides a released artifact.

Paper 3 of 3

DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data

Peter Schneider-Kamp, Jacob Nielsen, Gianluca Barmina and 2 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.CL

Problem

Current large language model development relies on massive, often non-permissible datasets, creating a high barrier for researchers committed to open-source and ethically sourced data. The obvious fix of using only permissible post-training data does not work because it would likely result in lower performance. This is due to the limited availability of such data.

Approach

The approach used in this paper is based on the Hierarchical Reasoning Model (HRM) architecture, which is trained from scratch using a mixture of 161 datasets. The model is trained using Fully Sharded Data Parallelism (FSDP) with a global batch size of 262,144 tokens. The training data is curated to include a mix of English and Danish instruction, knowledge, mathematics, and agentic-style post-training data. The model uses a hidden size of 1,536 and 12 attention heads per layer. The training process involves a 2,000-step linear warm-up and a constant schedule thereafter.

Result

The Mimir model outperforms all considered competitors on BoolQ, Winogrande, and DROP. On Math & Code, Mimir leads across its weight-class for GSM8K and HumanEval. On the Danish benchmarks, Mimir outperforms all competitors on grammatical tasks and question-answering tasks.

Why it matters

Researchers committed to open-source and ethically sourced data should care about this paper because it presents a model that achieves competitive performance using only permissible post-training data.

Method details
  • The model uses the Hierarchical Reasoning Model Text (HRM-Text) architecture with a hidden size of 1,536.
  • The training data consists of 161 datasets with a total of 70,479,308,606 tokens per epoch.
  • The model is trained using the AdamW optimizer with a peak learning rate of 2,000 and a weight decay of 0.1.
  • The model is trained for M steps on 8 NVIDIA B200 GPUs with 180 GB HBMe3.
  • The evaluation setup uses temperature 0 (greedy decoding) with shuffle seed 4242 on full datasets.
Numbers
  • Mimir yields a 36.7% improvement compared to HRM-Text (64.1 Mimir vs. 46.9 HRM-Text) on Math & Code.
  • Mimir is only 0.3 points behind Qwen 3.5 4B on English tasks.
  • Mimir is only 3.8% behind SmolLM3 3B on Math & Code.
  • Mimir achieves an accuracy of 87.8 on BoolQ.
  • Mimir achieves an accuracy of 73.5 on Winogrande.
Limitations

The paper does not establish the performance of the model on tasks that require reasoning beyond the capabilities of the Hierarchical Reasoning Model.

Mimir uses the Hierarchical Reasoning Model Text (HRM-Text) architecture with a hidden size of 1,536.Found in the source text, word for word.

Picked because: This paper earns a slot because it introduces DFM Mimir v1, a 1-billion-parameter language model that delivers competitive performance while being trained from scratch on permissible post-training data, which is a result that an engineer could apply to efficient inference on small or local models.