arXiv digest

Sunday

August 16, 2026

Up to five new AI papers a day, read in full where arXiv renders them and checked against the source.

Paper 1 of 6

LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure

Fanfei Li, Jana Zeller, Manuel Prada-Corral and 4 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.CL

Problem

Studying knowledge and skill acquisition in modern language models is difficult due to heterogeneous web-scale text corpora, making it hard to characterize prior exposure to related content. The obvious fix of using a smaller model or less data does not work because it does not provide a controlled environment for studying knowledge acquisition. This lack of control makes it difficult to understand how models acquire and use knowledge.

Approach

The method works by introducing LittleCurriculum, a curated 88B-token pretraining corpus tailored to U.S. elementary school material, and training a 5B-parameter language model called LittleLearner from scratch on this corpus. The LittleCurriculum corpus is constructed using a pipeline that includes age-of-acquisition pre-filtering, LLMJ annotation and classifier training, curriculum-based classification, symbolic filtering, and frequency sampling. The Qwen3 architecture is used for the model.

Result

The results show that LittleLearner has clear knowledge and capability boundaries mapped to interpretable curriculum guidelines. The model's performance on MathCAMPS shows that scaling model parameters helps performance within K, 5, but does not recover performance well in Beyond-K, 5. Post-training on unfiltered content increases K, 5 capabilities but does not overcome the Beyond-K, 5 gap.

Why it matters

This work is important for researchers who want to study knowledge acquisition in language models in a controlled environment, as it provides a developmentally restricted sandbox for this purpose. The results have implications for the development of more transparent and interpretable language models.

Method details
  • The model is trained for 100 hours on 8 NVIDIA B200 GPUs.
  • The LittleCurriculum corpus is an 88B-token pretraining corpus filtered from FineWeb-Edu to only contain elementary school (K, 5) data.
  • The Qwen3 architecture is used for the LittleLearner model.
  • The model is evaluated on MathCAMPS, a dataset for mathematical reasoning.
  • The model is compared against an unfiltered baseline.
Numbers
  • 88B-token pretraining corpus
  • 5B-parameter language model
  • 100 hours of training time
  • 8 NVIDIA B200 GPUs
  • 95% CIs for performance on MathCAMPS
Limitations

The paper does not establish how to bridge the performance gap between LittleLearner and the unfiltered baseline on out-of-scope content, and notes that certain emergent behaviors like in-context learning may be less pronounced than at frontier scales.

Consequently, studying knowledge and skill acquisition is difficult, as prior exposure to related content is hard to characterize.Found in the source text, word for word.

Picked because: This paper introduces LITTLECURRICULUM, a curated pretraining corpus that allows for controlled study of knowledge and skill acquisition in LLM agents.

Paper 2 of 6

Reduced Matrix Multiplication: Input-Adaptive Matrix-Product Reduction for LLM Inference

Zixuan Lan, Yanhong Li, Jiawei Zhou · abstract · pdf

quote verifiedfigures checkedread: full textcs.LG

Problem

Transformer-based language models incur substantial inference cost due to repeated high-dimensional matrix multiplications. The obvious fix of static pruning does not work as it leads to significant performance degradation. Dynamic selection is necessary to preserve performance.

Approach

The proposed method, Reduced Matrix Multiplication (RMM), reduces Transformer matrix products by selecting informative slices along their contraction dimensions. RMM works by dynamically selecting the retained feature dimensions based on activation statistics. It consists of attention-side and MLP-side components. The attention-side computations are more robust to pruning than MLP-side components. RMM provides a smooth and predictable accuracy-efficiency trade-off under a simple retention-ratio control.

Result

The results show that RMM remains robust across the evaluated discriminative, autoregressive generation, and long-context settings. At a retention ratio of 0.9, LLaMA 3.1 70B remains close to the full model on most benchmarks. The results also show that attention-side computations are substantially more reducible than MLP components.

Why it matters

This research is important for developers of large language models who need to optimize their models for inference efficiency. RMM provides a scalable direction for input-adaptive inference-time optimization.

Method details
  • Models evaluated include LLaMA 3.1 70B, LLaMA 3.1 8B, LLaMA 3.2 3B, and Qwen3 32B.
  • Datasets used include Copa, PiQA, CommonsenseQA, ARC-Easy, and WikiText.
  • Baselines compared against include SparseGPT, Wanda, SliceGPT, and magnitude pruning.
  • Ablations run include component-wise pruning analyses on attention and MLP blocks.
Numbers
  • RMM at retention ratio 0.9 achieves 84.4 accuracy on Copa, compared to the baseline.
  • RMM at retention ratio 0.5 achieves 49.4 accuracy on Copa, compared to the baseline.
  • The speedup of RMM over the dense model is 1.36 at sequence length 1024.
  • The perplexity of LLaMA 3.1 70B at retention ratio 0.5 is 167.78, compared to the baseline.
Limitations

The paper does not establish the applicability of RMM to other types of neural networks beyond Transformers.

Together, these results position RMM as a scalable direction for input-adaptive inference-time optimization.Found in the source text, word for word.

Picked because: This paper proposes Reduced Matrix Multiplication, a training-free method for reducing the inference cost of Transformer-based language models, which is a key challenge for efficient deployment of LLM agents.

Paper 3 of 6

MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification

Daniel Perkins, John Squires, Janou Milligan and 2 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.CV

Problem

Modern image classification models excel when trained on single task-specific datasets but often struggle to generalize across domains and difficulty levels. Standard ensembles are almost exclusively designed for narrow domains. Limited work has addressed the challenge of multi-domain image classification. Standard fusion strategies require passing every input through all models.

Approach

The proposed method ARMDIL uses a multimodal large language model (MLLM) agent to dynamically route each image to the most suitable vision backbone. The MLLM router predicts the visual domain of a given image and routes it to the optimal expert for classification. The framework provides an interpretable and adaptable alternative to standard ensembles. ARMDIL leverages an MLLM to predict the visual domain of a given image and then routes the image to the optimal expert for classification. The MLLM is used in conjunction with convolutional neural networks (ResNets), self-supervised representation learners (SSL), and vision-language models (VLMs).

Result

The routing accuracy for each ablation configuration is detailed in Table 6. The results show that self-consistency slightly reduces accuracy on CIFAR10, FER2013, and OrganAMNIST. The downstream classification accuracy and F1 scores for each ablation are detailed in Table 7. The results show that the self-consistency models achieve higher overall routing accuracies, but their classification accuracies and F1 scores lag behind the ARMDIL baseline.

Why it matters

This work is important for researchers and developers working on image classification tasks, especially those that require cross-domain robustness. The proposed method ARMDIL provides a novel approach to addressing the challenge of multi-domain image classification.

Method details
  • The MLLM used is Gemma-4-12B
  • The datasets used are CIFAR10, FER2013, EuroSAT, and OrganAMNIST
  • The experts used are ResNet, DINO, and CLIP
  • The ablation studies investigate the impact of self-consistency, chain-of-thought reasoning, and image-quality statistics
Numbers
  • 97.80, the domain classification accuracy of ARMDIL on CIFAR10
  • 98.07, the domain classification accuracy of SC w/o CoT & IQ on EuroSAT
  • 99.02, the classification accuracy of ARMDIL on CIFAR10
  • 90.78, the overall classification accuracy of ARMDIL
Limitations

The paper does not establish the efficacy of ARMDIL on other datasets or domains.

Crucially, we show that ARMDIL effectively navigates these trade-offs, performing competitively with specialized training-based routersFound in the source text, word for word.

Picked because: This paper presents ARMDIL, a heterogeneous ensemble approach that uses a multimodal large language model to dynamically route images to the most suitable vision backbone, demonstrating a concrete application of LLM agents in retrieval and evaluation.

Paper 4 of 6

OmniScientist: An Omni-Modal Omni-Discipline AI Scientist

Bobo Li, Hao Fei, Tianjie Ju and 2 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Existing systems typically reason over text, code, labels, or precomputed summaries, leaving scientifically decisive spatial, temporal, cross-channel, and procedural relations unavailable to the agent. Workflow coverage alone does not provide access to the full evidence on which scientific discovery depends. This limitation hinders the ability of AI scientists to automate complete research workflows.

Approach

The OmniScientist framework consists of a perception layer and 3 autonomous agents for ideation, experiment, and writeup, operating within a deterministic pipeline. The perception layer provides spatial, temporal, cross-channel, statistical, and dynamic analysis capabilities. The ideation stage formulates falsifiable hypotheses, the experiment stage designs tests and executes code, and the writeup stage compiles the final paper. The system enforces novelty screening, statistical validity, execution provenance, and numerical traceability through code-enforced checks.

Result

The system completes the full path from raw data to a compiled manuscript in all 36 cases and achieves a mean overall paper score of 6.3 with the reference reasoning backbone. The strongest alternate backbones, such as GLM and Kimi, fall within a remarkably similar performance range on their respective evaluated subsets.

Why it matters

Researchers and scientists should care about this work because it demonstrates the ability of an AI system to automate the entire research pipeline, from raw data to a compiled manuscript, and achieve high-quality results.

Method details
  • The perception model is pinned to Claude Sonnet 5 (Anthropic 2026) in every run.
  • The system uses a controlled run_python environment that manages subprocess execution and figure capture.
  • The system operates under explicit computational limits to ensure bounded exploration and convergence.
  • Each complete run outputs a structured JSON record, a Markdown summary, a replayable execution trace, and a compiled PDF report.
Numbers
  • mean overall paper score of 6.3
  • 36 cases
  • 85% of head-to-head judgments
  • 6.5 mean composite score for Sonnet 5 backbone
Limitations

The paper does not establish the ability of the system to handle cases outside the 36 datasets used in the evaluation.

By running idea, rigour, and claim checks in code, the system enforces novelty screening, statistical validity, execution provenance, and numerical traceability.Found in the source text, word for word.

Picked because: OmniScientist presents a concrete AI scientist system that automates research workflows, including hypothesis generation and manuscript preparation, making it a valuable read for those interested in LLM agents and tool use.

Paper 5 of 6

SAEVerbalizer: Generating Explanations for Sparse Autoencoder Features via Representation Verbalization

Weihan Meng, Hongzhu Guo, Yi Jing and 5 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.CL

Problem

Before this paper, explaining sparse autoencoder features relied primarily on external observation, leading to superficial explanations and computational inefficiency. The obvious fix of using the decoder directions directly does not work due to the lack of a framework to generate natural-language explanations from these directions. This limitation leads to a need for a new approach.

Approach

The SAEVerbalizer framework works by injecting sparse autoencoder decoder directions into a large language model's representations and fine-tuning the downstream layers to generate natural-language explanations. The framework consists of a verbalizer and an adapter, where the verbalizer receives a fixed prompt and injects the decoder direction, and the adapter maps decoder directions from another language model into the verbalizer's representation space. The verbalizer is fine-tuned on feature-explanation pairs, and the adapter is trained on aligned representation pairs. The framework enables the generation of explanations directly from decoder directions, addressing the limitations of previous approaches.

Result

The experiments show that the SAEVerbalizer framework achieves high Reference Agreement (RA) scores, indicating that the generated explanations are of high quality. The framework also exhibits transferability across sparse autoencoder dictionaries and language models. The results demonstrate the effectiveness of the SAEVerbalizer framework in generating natural-language explanations for sparse autoencoder features.

Why it matters

The SAEVerbalizer framework is important for researchers and practitioners working with sparse autoencoders and large language models, as it provides a way to generate high-quality natural-language explanations for the features learned by these models.

Method details
  • The verbalizer uses a Transformer architecture with a specific layer for injection, such as layer 16 in the 27B model.
  • The adapter is a single affine layer that maps source-LLM representations to the verbalizer's injection-layer representation space.
  • The training data for the verbalizer includes 12k, 24k, and 48k qualified feature-explanation pairs for each sparse autoencoder.
  • The evaluation metric used is Reference Agreement (RA), which measures the proportion of test features for which the generated explanation agrees with the reference explanation.
  • The test sets include the Global Train-Standard (GTS), Low-Index Gold (LIG), and Global Gold (GG) sets, each with a specific qualification standard.
Numbers
  • RA score of 52.3% on the GTS set for the 27B-L16 configuration with 48k feature-explanation pairs
  • RA score of 80.5% on the LIG set for the 27B-L16 configuration with 48k feature-explanation pairs
  • RA score of 56.1% on the GG set for the 27B-L16 configuration with 48k feature-explanation pairs
  • 12k, 24k, and 48k feature-explanation pairs used for training the verbalizer
Limitations

The paper does not establish the absolute correctness of generated explanations, as the references are inferred from activation examples rather than established ground truth.

Because each configuration uses a distinct SAE, both trends may also partly reflect variation across SAE feature distributions.Found in the source text, word for word.

Picked because: SAEVerbalizer introduces a framework for generating explanations for sparse autoencoder features via representation verbalization, providing a practical approach to understanding LLM representations.

Paper 6 of 6

Intern-S2-Preview: Scientific Agentic Foundation Model

Lei Bai, Jiaqi Cao, Chiyu Chen and 122 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.LG

Problem

Scientific discovery increasingly requires AI systems that can reason over scientific evidence of heterogeneous modalities, interact with scientific tools and environments, and sustain progress across long task horizons. The obvious fix of relying solely on textual descriptions or visualized signals does not work for time series understanding. General-purpose models lack the necessary capabilities for scientific intelligence.

Approach

The Intern-S2-Preview method works by first performing scientific multimodal pre-training over rendered scientific documents, interleaved image-text data, and diverse scientific corpora. Then, a unified post-training pipeline is applied, consisting of supervised fine-tuning, scalable multi-task reinforcement learning, black- and white-box agentic RL, and on-policy distillation. The pipeline is supported by practical techniques such as partial rollout with off-policy correction and adaptive length regularization. The architecture level extends time series modelling from efficient long-sequence understanding to numerical forecasting through upgraded time series modules. A separate memory-augmented path, Memory Decoder, is studied for rapid scientific specialization.

Result

The Intern-S2-Preview-397B model achieves competitive or leading results in multiple settings, including time series understanding and forecasting on the SciTS benchmark. The model outperforms general-purpose Text LLMs and Vision-Language LLMs, and achieves comparable or better performance than the trillion-parameter-scale Intern-S1-Pro. The separate Intern-MemDec-4B extension improves the Biology-Instructions average score from 56.92 to 60.32.

Why it matters

This research is important for scientists and researchers who require AI systems that can reason over scientific evidence and interact with scientific tools and environments. The model's capabilities can be used to support multimodal scientific understanding, reasoning, generation, and long-horizon tasks.

Method details
  • The model Intern-S2-Preview-397B has 397B parameters.
  • The training pipeline includes supervised fine-tuning, scalable multi-task reinforcement learning, and on-policy distillation.
  • The model is evaluated on benchmarks such as Biology-Instructions, SciTS, and ProteinBinder-9.
  • The separate Intern-MemDec-4B extension improves the Biology-Instructions average score.
  • The model achieves state-of-the-art results on multiple scientific and general-purpose benchmarks.
Numbers
  • 56.92, Biology-Instructions average score without Memory Decoder
  • 60.32, Biology-Instructions average score with Intern-MemDec-4B
  • 97.1, F1 score of Intern-S2-Preview-397B on ASU01 task
  • 66.9, F1 score of Intern-S2-Preview-397B on PHU01 task
  • 36.8, F1 score of Intern-S1-Pro on PHU01 task
Limitations

The paper does not establish the generalizability of the model to all scientific domains, but rather focuses on biology as a representative domain.

Scientific discovery increasingly requires AI systems that can reason over scientific evidence of heterogeneous modalities, interact with scientific tools and environments, and sustain progress across long task horizons.Found in the source text, word for word.

Picked because: Intern-S2-Preview presents a series of scientific agentic foundation models designed to support multimodal scientific understanding, reasoning, generation, and long-horizon tasks, making it a relevant read for those interested in LLM agents and tool use.