arXiv digest

Tuesday

August 18, 2026

Up to five new AI papers a day, read in full where arXiv renders them and checked against the source.

Paper 1 of 10

zLend: A Dual-Scope Cash-Flow Reconstruction Framework for On-Chain Credit Underwriting

Girish G N, Ashutosh Sahoo, Akshay SP and 2 others · abstract · pdf

quote verifiedfigures checkedread: full textq-fin.RM

Problem

Total wallet value conflates liquid and illiquid assets, causing mispricing of risk. Using total value alone fails because it cannot distinguish spendable reserve from non‑spendable holdings.

Approach

zLend reconstructs a wallet's daily balance twice: once over a fixed stablecoin basket and once over all fungible tokens using a cumulative‑net‑flow algorithm with a non‑negativity offset. From each series it computes liquidity coverage, cash‑flow regularity, peak‑to‑trough drawdown and recovery, outflow concentration, trend classification, and a recurring‑counterparty detector. The two scopes are then compared to produce liquidity‑ and flow‑mismatch severity scores. These scores feed a four‑tier underwriting rule specified in Tables 2 and 3. The entire pipeline is implemented in Python and ported to TypeScript with strict numerical verification.

Result

The analysis shows tier assignment is governed predominantly by the reference loan size, with four of six reference wallets changing tier across loan sizes. Drawdown and coverage criteria bind on disjoint wallets, confirming neither subsumes the other, and no tier‑rule criterion is inert. The framework is deployed in production and feeds real lending decisions via API integration.

Why it matters

DeFi platforms and on‑chain credit providers should care because zLend offers a verified, production‑ready way to separate liquid reserve from total wealth and generate robust underwriting signals.

Method details
  • Daily balance reconstruction uses a cumulative‑net‑flow construction with a non‑negativity offset (Algorithm 1)
  • Two parallel scopes are built: a stablecoin‑only view and a total‑wealth view over all fungible transfers
  • Signal family includes liquidity coverage, cash‑flow regularity, drawdown and recovery, outflow concentration, trend, and recurring‑counterparty detection
  • Tier assignment follows a four‑tier rule (Tables 2 and 3) driven mainly by the reference loan size
  • Implementation verification replicates compensated (Kahan) summation, population variance, round‑half‑to‑even to ten decimal places, and null propagation for near‑zero denominators
  • Validation achieved 78 of 78 field assertions against deployed system reference fixtures
Numbers
  • numerical tolerance, 1e-9, compared against reference implementation
  • field assertions, 78 of 78, exact agreement with reference fixtures
  • tier changes, four of six reference wallets, changed tier across loan sizes USD 10 to USD 25,000
  • record size, 166 columns, per evaluation record written to storage
  • scope fields, 69 fields per scope, plus six cross‑scope ratios
  • divergence magnitude, four orders of magnitude, between total‑wealth and stablecoin reconstructions
Limitations

The paper does not evaluate predictive performance on actual default outcomes; it focuses on methodology and verification.

tier assignment is governed predominantly by the reference loan size, with four of six reference wallets changing tier across loan sizesFound in the source text, word for word.

Picked because: zLend provides a deployed on-chain cash‑flow reconstruction framework with open‑source implementation, giving engineers a ready‑to‑use tool for credit underwriting automation.

Paper 2 of 10

Model Hypnosis: Strong control of AI via additive subliminal effects

Enric Boix-Adsera, Benedict Tessler · abstract · pdf

quote verifiedfigures checkedread: full textcs.CL

Problem

Prior work assumed that only strong, obvious prompts could reliably steer model outputs, and it ignored the impact of many weak, seemingly irrelevant cues; simply removing obvious cues does not prevent steering because inconspicuous paraphrases and typos can still combine to control behavior.

Approach

The authors treat each cue as an additive component to the log‑odds of a target answer and fit a linear model using ridge regression on random prompt configurations. This additive model predicts how stacked weak cues will shift probabilities. They then rank prompt configurations by the predicted log‑odds sum and select the top and bottom candidates as extreme prompts. By validating these extremes with fresh generations they demonstrate that stacked cues can flip model responses. The same pipeline is applied to both non‑reasoning and reasoning models.

Result

The additive model captures most of the variance in held‑out configurations, achieving R² values from 0.3 to 0.99 (median 0.75). Stacked weak cues can reliably flip the modal answer in both non‑reasoning and reasoning models, and the directional effect often transfers across models, especially within families.

Why it matters

AI safety and interpretability researchers should care because even innocuous textual variations can strongly steer model outputs, creating new attack surfaces and challenges for reliable deployment.

Method details
  • Non‑reasoning models include Qwen‑2.5 (3B‑72B), Qwen‑3 (4B‑32B), Qwen3.5‑9B, Gemma‑2‑9B, Gemma‑4‑12B, Llama‑3.1‑8B, Phi‑4, OLMo‑2‑7B, OLMo‑3‑7B.
  • Reasoning models include Qwen3‑8B (256, 1024, 4096 token budgets), GPT‑OSS‑20B, GPT‑5.6‑terra, GPT‑5.6‑Sol, Gemini‑3‑Flash, Claude‑Haiku‑4.5, Claude‑Sonnet‑5.
  • Additive model fitted to log‑odds using ridge regression on ~10^3 random prompt configurations per model‑cue‑effect cell.
  • Baseline comparison: random‑prompt additive fit versus measured log‑odds; steering range measured by top/bottom 100 extreme prompts.
  • Transfer experiment: prompts optimized on one source model evaluated on 16 non‑reasoning target models to assess directional effect preservation.
Numbers
  • held‑out configuration‑level R² spanning roughly 0.3 to 0.99
  • 5th‑95th percentile R² 0.54‑0.93
  • median R² 0.75
  • top‑ and bottom‑100 candidate prompts evaluated per model
  • Qwen3‑8B evaluated at three token budgets (256, 1024, 4096)
  • transfer significant for animal cues and some phrasing and JSON cues
Limitations

The paper evaluates a limited set of reasoning models due to API cost and does not prove that the additive approximation holds for all possible cue types or model families.

the additive model explains most of the held‑out configuration‑level variance, with held‑out configuration‑level R2 spanning roughly 0.3 to 0.99 (5th-95th percentile 0.54-0.93; median 0.75)Found in the source text, word for word.

Picked because: Model Hypnosis uncovers systematic prompt‑level control of LLMs and releases code to test and verify such effects, directly supporting LLM agent safety tooling.

Paper 3 of 10

GEO-Flag: Detecting and Measuring GEO-Optimized Web Content

Junjie Chu, Ye Leng, Mingjie Li and 3 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.LG

Problem

Existing GEO detection methods achieve decent aggregate F1 but fail to generalize across GEO optimizer families and authorship groups, often relying on authorship shortcuts; simple calibrations or zero‑shot LLMs do not close these gaps.

Approach

The paper introduces Intervention‑Paired Training (IPT), which supervises detector responses to paired GEO interventions and non‑GEO AI polishing, enforcing positive and invariant pair constraints. IPT is applied to fine‑tuned ModernBERT (and Qwen) encoders. The method is integrated into a GEO‑gated Agent system that first flags pages, then audits citation URLs through deterministic preprocessing, labeling, and verifiability aggregation modules. Together these components detect GEO pages and assess citation reliability without needing provenance metadata.

Result

IPT substantially improves detection performance: ModernBERT F1 climbs to 0.944 and worst‑group accuracy to 0.883, while Qwen‑IPT also gains accuracy and F1. Baseline lexical models still reach competitive F1 (0.878 to 0.880), but exhibit large shortcut gaps. The system estimates GEO prevalence of 8.90% overall and 16.36% among pages modified in 2026.

Why it matters

Search engine operators and auditors should care because the method provides a practical way to flag and audit GEO‑optimized content, improving the reliability of generative search results.

Method details
  • GEOFlagBench benchmark: 3,200 webpages, 400 queries, 4 domains, 8 GEO optimizer families.
  • Query‑level grouped split: 2,238 training documents and 962 test documents.
  • Fine‑tuned ModernBERT with 1,024‑token and 8,192‑token contexts; Qwen3‑0.6B SFT also used.
  • IPT raises ModernBERT accuracy from 0.839 to 0.931 and F1 from 0.862 to 0.944.
  • TF‑IDF word (1‑2) + LR baseline achieves highest aggregate F1 of 0.880.
  • Experiments run on Google Cloud g2‑standard‑8 (8 vCPUs, 32 GB RAM, NVIDIA L4 24 GB GPU).
Numbers
  • F1 0.880 (TF‑IDF word + LR) compared to other baselines
  • Accuracy 0.931 (ModernBERT‑IPT) vs 0.839 baseline
  • F1 0.944 (ModernBERT‑IPT) vs 0.862 baseline
  • Worst‑group accuracy 0.883 (ModernBERT‑IPT) vs 0.725 baseline
  • GEO prevalence 8.90% overall, 16.36% among 2026‑modified pages
  • Aggregate baseline F1 0.880 (highest among non‑IPT methods)
Limitations

The paper does not claim to detect malicious intent, factual inaccuracy, or to generalize beyond the eight studied GEO optimizer families.

IPT improves both overall performance and shortcut diagnostic metrics.Found in the source text, word for word.

Picked because: GEO‑Flag introduces a detection pipeline for generative‑search‑optimized web content and releases datasets and classifiers that engineers can integrate into CI/CD for SEO auditing.

Paper 4 of 10

Quipu: A Governed Bitemporal Knowledge Graph Store

Steve Brown · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Knowledge‑graph stores assume defaults: accept‑then‑clean, single or no time axis, equal trust for all writers, and external governance. These defaults cause hidden defects when multiple agents write concurrently. Simply adding pre‑state checks does not catch defects that only appear in the combined post‑state.

Approach

Quipu implements a governed store that inverts the four defaults. Writes are filtered by a gate that evaluates predicates against the pending post‑state (GS1). Gate outcomes are persisted as signed, time‑indexed verdicts (GS2). Authority is attached to named‑graph partitions with a lattice that never widens when composed (GS3, GS4). The trace, verdicts, and policies live in‑store, making audit decidable (GS5), and every decision can be replayed as‑of its transaction (GS6). Together these mechanisms provide bitemporal EAVT storage with three‑valued operations and a non‑widening label lattice.

Result

The gated store eliminated all planted defects (0 of 6) while the ungated store exhibited all six. All seven composition probes upheld the lattice contract. Fifty of fifty verdicts re‑derived faithfully at their instant, and the DEMM‑Bench content‑only reader reconstructed all 512 property‑level questions with zero overclaim, whereas container‑presence baselines overclaimed up to 87.5%.

Why it matters

Developers of knowledge‑graph stores and autonomous agents that write KG data should care because Quipu provides provable governance, auditability, and defect prevention without post‑hoc cleaning.

Method details
  • Dataset: Census benchmark, a deterministic multi‑writer lifecycle with seeded run 42.
  • Dataset: DEMM‑Bench external decision‑evidence sufficiency benchmark.
  • Architecture: Bitemporal EAVT log with three‑valued operation (assert, retract, tombstone).
  • Label system: Four‑axis lattice (freshness, trust, durability, policy) folded by meet/join.
  • Baseline: Ungated store used for comparison of defect detection.
  • Composition: Seven lattice probes testing non‑widening composition.
Numbers
  • planted defects, 0 of 6, gated store vs 6 of 6 ungated
  • composition probes upheld, 7 of 7, non‑widening contract
  • verdict replay, 50 of 50, re‑derived faithfully
  • property‑level overclaim, 0.0, content‑only reader vs up to 0.875 baseline
  • property‑level sufficient, 8/64, property‑level reader
  • container‑presence overclaim, 0.75, baseline overclaim rate
Limitations

The paper does not evaluate Quipu on large‑scale real‑world workloads beyond the synthetic Census and DEMM benchmarks.

the gated store ends with 0 of 6 planted defects versus 6 of 6 ungatedFound in the source text, word for word.

Picked because: Quipu delivers a governed bitemporal knowledge‑graph store with source code, addressing data governance and versioning challenges in production back‑ends.

Paper 5 of 10

ClawGym II: Exploring Black-Box RL on Agent Harness

Huatong Song, Fei Bai, Ming Yang and 17 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.CL

Problem

Scaling reinforcement learning through complex, opaque harnesses for long-horizon tasks was unstable and largely unexplored, and naïvely applying standard RL pipelines fails due to infrastructure failures, invalid token generation, and mismatched advantage estimation across heterogeneous harnesses.

Approach

The method builds a sandbox execution layer that isolates environments and harnesses, adds a serving proxy at the model boundary to capture all model calls, organizes these calls into prefix trees, and then applies adapted PPO (critic‑based) or GRPO (critic‑free) optimization over the recovered tree structure; it also introduces mix‑harness training by jointly sampling task-harness pairs from multiple harnesses within each batch while keeping advantage groups separate.

Result

Black‑box RL with Qwen3-30A3B raises Pass@1 on ClawGym-Bench by 9.98 points using OpenClaw and by 14.81 points using Claude Code, while maintaining stable optimization for 200 to 400 steps; mix‑harness training yields training rewards and downstream evaluation comparable to or slightly better than single‑harness baselines.

Why it matters

Researchers and engineers building long‑horizon agents that rely on external tool harnesses should care because the framework enables stable, scalable RL without redesigning each harness, and it can jointly leverage heterogeneous execution systems.

Method details
  • Model: Qwen3-30A3B used as the policy backbone.
  • Benchmarks: ClawGym-Bench for Pass@1 evaluation, plus ClawGym-SynData, JobBench, and OfficeQA for training and testing.
  • Mix‑harness batch configuration: 32 task‑harness instances with 8 rollouts per instance.
  • Training stability observed over approximately 200 to 400 optimization steps.
  • Baseline comparisons: individual‑harness training with OpenClaw or Claude Code.
Numbers
  • Pass@1 improvement, 9.98, OpenClaw
  • Pass@1 improvement, 14.81, Claude Code
  • Optimization steps, 200 to 400, stable training
  • Batch size, 32 task‑harness instances, 8 rollouts per instance
  • Training reward, comparable, mix‑harness vs individual harness
Limitations

The paper does not claim to establish performance beyond the presented models and benchmarks, nor does it address scalability to arbitrarily larger models or entirely new harness types.

black-box RL improves Pass@1 on ClawGym-Bench by 9.98 and 14.81 points through OpenClaw and Claude Code, respectivelyFound in the source text, word for word.

Picked because: ClawGym II offers a black‑box reinforcement‑learning harness with open‑source benchmarks, enabling engineers to experiment with agent‑level automation without modifying environments.

Paper 6 of 10

LAVA: Logic-Aware Validation and Augmentation Framework for Large-Scale Financial Document Auditing

Ruoqi Shu, Xuhui Wang, Isaac Wang and 2 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.AI

Problem

Existing pipelines cannot reliably process heterogeneous layouts, semantically rich content, and embedded business rules, leading to hallucinations and false positives; simply applying end-to-end LLMs fails because they lack layout preservation and explicit logical grounding.

Approach

LAVA is a modular, backbone-agnostic pipeline built on multimodal large language models that follows a four-stage design: document‑rule retrieval, layout‑preserving information extraction, auxiliary metadata enrichment, and auditable symbolic/arithmetic verification. It runs two parallel tracks-document processing and rule grounding-that exchange constraints. The document track produces structured content with enriched metadata, while the rule track retrieves, classifies, and dispatches rules to either a Symbolic Reasoner or an Arithmetic Processor based on reasoning type. Knowledge Extraction (KE) provides structured markup, Info Augmentation (IA) adds semantic metadata, and the Arithmetic Processor (AP) enforces numerical fidelity. The system is fully modular, allowing independent updates to each component.

Result

LAVA achieves the lowest hallucination and numerical infidelity rates among all baselines and dramatically reduces edge‑case error, while using fewer tokens per rule check. In the ablation study, removing KE raises error to 0.65 to 0.67, whereas full LAVA maintains error below 0.07 across rule categories.

Why it matters

Enterprises that need high‑accuracy, auditable financial document validation should consider LAVA because it reduces hallucinations and false positives while remaining efficient.

Method details
  • Uses multimodal large language models as a backbone, but is backbone‑agnostic.
  • Evaluated on a proprietary Canadian mortgage application dataset of ~1,000 scanned PDFs/images with dozens of expert‑curated validation rules.
  • Baselines compared: VLM + Field‑Level OCR, LLM + Field‑Level OCR, LLM + Enhanced OCR.
  • Ablations include removal of KE, IA, and AP, and variants with plain‑text and Markdown KE.
  • Metrics include Factual Hallucination Rate, Hallucination Rate, Numerical Infidelity Rate, Edge‑Case Error Rate, and token cost.
Numbers
  • Factual Hallucination Rate 0.01 vs 0.03 (VLM) and 0.03 (LLM)
  • Hallucination Rate 0.03 vs 0.08 (VLM) and 0.10 (LLM)
  • Numerical Infidelity Rate 0.01 vs 0.08 (VLM) and 0.04 (LLM)
  • Edge Case Error Rate 0.02 vs 0.27 (VLM) and 0.25 (LLM)
  • Ablation LAVA w/o KE error 0.65 vs LAVA 0.01
  • Ablation LAVA w/o IA error 0.45 vs LAVA 0.05
Limitations

The paper does not explicitly discuss limitations.

LAVA outperforms baselines in hallucination control and edge-case handling while maintaining efficient token usageFound in the source text, word for word.

Picked because: LAVA presents a logic‑aware validation and augmentation framework for large‑scale financial document auditing, complete with a released pipeline that can be deployed in enterprise workflows.

Paper 7 of 10

Le Critique: Privileged Value Functions for LLM Reinforcement Learning

Siddarth Venkatraman, Matthieu Dinot, Laurence Aitchison · abstract · pdf

quote verifiedfigures checkedread: full textcs.LG

Problem

Group-relative RL methods for LLMs suffer high gradient variance, straggler rollouts, and off‑policyness, and naïve addition of learned value functions is hindered by infrastructure complexity and the practical success of critic‑free pipelines.

Approach

The paper introduces Privileged Value Functions (PVF) that condition the value head on task‑specific privileged context without biasing the policy objective, and TETHER, a baseline that adaptively interpolates between the group‑mean GRPO baseline and the token‑level value baseline based on the accuracy of the value function. Both methods reuse the same Monte‑Carlo advantage computation and share the value‑training configuration. PVF provides token‑level variance reduction using extra information, while TETHER dynamically weights the mean and value baselines during training.

Result

PVF consistently outperforms both the ordinary value baseline and the group‑mean baseline across Reasoning Gym, CodeIO, and Sudoku, achieving the highest end‑of‑training rewards in all four experiments. Tether improves over the simple value baseline on all four tasks and is competitive with the Mean baseline, beating Mean on Reasoning Gym and MiniF2F, matching it on CodeIO, and narrowing the gap on Sudoku.

Why it matters

Researchers building RL pipelines for large language models should consider PVF for better variance reduction when privileged information is accessible, and Tether for a robust way to incorporate value functions into existing GRPO workflows.

Method details
  • Qwen3-4B-Instruct-2507 is used for Reasoning Gym, CodeIO, and Sudoku; Qwen3.5-4B is used for MiniF2F
  • Tasks include Reasoning Gym, CodeIO, Sudoku, and MiniF2F
  • Baselines compared are Mean (group‑mean GRPO), VF (ordinary token‑level value), PVF, and Tether
  • Each run performs 20 value‑warmup updates before policy training
  • Policy steps: 800 for Reasoning Gym, 650 for CodeIO, 600 for Sudoku, 500 for MiniF2F
Numbers
  • 20 value warmup updates before policy training
  • 800 policy steps for Reasoning Gym
  • 650 policy steps for CodeIO
  • 600 policy steps for Sudoku
  • batch size 128 for Reasoning Gym and CodeIO
Limitations

Tether does not fully recover the performance of the Mean baseline on Sudoku, and PVF requires privileged context that may not be available for all tasks.

Pvf is the best performing method in all settings.Found in the source text, word for word.

Picked because: Le Critique proposes privileged value‑function techniques for LLM reinforcement learning and provides an implementation that reduces variance in policy training, useful for building reliable LLM agents.

Paper 8 of 10

Semantic Bandits: In-Context Exploration-Exploitation is Biased by Semantic Priors

David Eric Austin, Kaheer Suleman, Jackie Chi Kit Cheung · abstract · pdf

quote verifiedfigures checkedread: full textcs.CL

Problem

Prior work assumed LLM agents would explore like classical bandit algorithms, but semantic labels attached to actions bias their exploration-exploitation balance; simply removing explicit semantics does not fix the issue because LLMs retain implicit semantic priors from pre‑training.

Approach

The paper introduces the semantic bandit, extending the multi‑armed bandit with textual action labels. It evaluates four nomenclatures (alphanumeric, sentiment, ordinal, world knowledge) under helpful and misleading reward configurations. Models are prompted with a best‑performing CoT prompt and optionally with explicit bias warnings or exploration instructions. Performance is compared to a label‑agnostic UCB1 baseline. Normalized cumulative regret and exploration counts are measured to quantify bias effects.

Result

Semantic priors strongly shape LLM behaviour: when labels align with reward (helpful), regret drops dramatically (e.g., Qwen3-32B ordinal helpful 0.00) and exploration is suppressed; when labels are misleading, regret rises (Qwen3-32B ordinal misleading 0.79) and exploration increases. The alphanumeric control matches UCB1 (regret 0.15) showing minimal bias. Negative reward values trigger more exploration than equivalent positive values.

Why it matters

Researchers building LLM‑based agents for interactive decision tasks should account for semantic label bias, as it can dramatically affect exploration and overall performance.

Method details
  • Instruction‑tuned LLMs: OLMo-3.1-32B-Instruct, Qwen3-32B, Gemini 3.1 Flash Lite.
  • Semantic bandit tasks in farming and clothing domains with action labels sampled per run.
  • Four nomenclatures: alphanumeric (random strings), sentiment (valence adjectives), ordinal (rank labels), world knowledge (domain‑specific labels).
  • Baseline: classical UCB1 algorithm with 10,000 replicates per condition.
  • Prompt variations: standard CoT prompt, explicit bias warning, and explicit exploration instruction.
  • Experiments run with 10 replicates per condition (20 for reward‑scale sweep) and three variance levels (high, low, none).
Numbers
  • Normalized cumulative regret (Qwen3-32B ordinal helpful) 0.00 vs UCB1 0.15
  • Normalized cumulative regret (Qwen3-32B ordinal misleading) 0.79 vs UCB1,
  • Normalized cumulative regret (OLMo-3.1 ordinal helpful) 0.04 vs UCB1 0.15
  • Normalized cumulative regret (Gemini ordinal helpful) 0.13 vs UCB1 0.15
  • Normalized cumulative regret (Qwen3-32B sentiment helpful) 0.09 vs UCB1 0.15
  • Normalized cumulative regret (Qwen3-32B alphanumeric helpful) 0.16 vs UCB1 0.15
Limitations

The study is limited to synthetic semantic bandit tasks and does not demonstrate how the findings generalize to more complex real‑world decision‑making environments.

semantically informative action labels reduce exploration in favour of exploitation, improving performance when aligned with the reward structure and severely degrading it when misaligned.Found in the source text, word for word.

Picked because: Semantic Bandits analyses how in‑context LLM prompts bias exploration‑exploitation and supplies a toolkit for measuring and correcting such biases in deployed agents.

Paper 9 of 10

UniTAC: Universal Task-Aware Compression via Weighted Distortion Measures

Homa Esfahanizadeh, Matin Mortaheb, Jinfeng Du and 1 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.LG

Problem

Task‑specific codecs are brittle and retraining a separate codec for each downstream task is infeasible in the field.

Approach

UniTAC trains a single Vision Transformer based image codec whose encoder and decoder are conditioned on a per‑component importance vector. The importance vector, derived from gradient attribution of any downstream model, is sent as low‑overhead side information. During training the model sees a broad family of synthetic importance maps so it learns to allocate bits according to the injected weights. At runtime the same backbone is re‑targeted to any task by swapping in the task‑derived importance map, without any retraining.

Result

At 0.034 bpp a single UniTAC model attains 91.4% downstream classification accuracy, only 1.9% lower than a dedicated task‑based codec (93.3%) and substantially higher than a universal codec (76.9%). It also matches the universal codec on overall PSNR while improving semantic PSNR and task accuracy at equal or lower rate.

Why it matters

Researchers and engineers building physical AI systems such as autonomous vehicles and robots should care because UniTAC enables flexible, task‑aware compression without per‑task retraining.

Method details
  • Architecture: two‑stage ViT autoencoder with token‑level weight conditioning and a hyperprior entropy model.
  • Training data: codecs trained on AffectNet and evaluated on the CelebA test split.
  • Importance maps: generated per image by integrated gradients of a downstream ResNet‑18 classifier.
  • Baselines: compared against a non‑semantic universal codec and a task‑based semantic codec, all sharing the same ViT architecture.
  • Training objective: rate-distortion loss combining a rate estimate with the separable weighted distortion normalized by total weight (Eq. 11).
  • Synthetic importance maps: random mixtures of Gaussian blobs of varying count, location, and scale.
Numbers
  • 0.034 bpp
  • 91.4% accuracy (UniTAC)
  • 93.3% accuracy (task‑based codec)
  • 76.9% accuracy (universal codec)
Limitations

The paper does not fully quantify the universality gap and acknowledges that performance depends on the quality of the supplied importance map.

On a localized task at 0.034 bpp, a single UniTAC model reaches 91.4% accuracy, only 1.9% below a task-based codec (93.3%) and above universal codecs (76.9%).Found in the source text, word for word.

Picked because: UniTAC introduces a universal task‑aware image codec with a released model that can be self‑hosted to adapt compression to changing downstream tasks, simplifying infrastructure for edge AI.

Paper 10 of 10

Steering the Flow: Inverting Face Recognition Models via Gradient-Guided Flow Matching

Ye Lu, Shen Wang, Zhaoyang Zhang and 4 others · abstract · pdf

quote verifiedfigures checkedread: full textcs.CV

Problem

Existing model inversion attacks rely on indirect or highly stochastic guidance, which makes it difficult to stably optimize generation trajectories toward target facial images.

Approach

SFMI first trains an unconditional Flow Matching model to learn a smooth vector field that maps Gaussian noise to the human‑face manifold. Then, during sampling, a Progressive Guidance Scheduler injects time‑dependent target‑specific gradients computed by back‑propagating through the face recognition model into the learned velocity field. The gradients steer the generative flow from random noise toward high‑density regions of the target identity. PGS dynamically modulates guidance strength over time to balance identity alignment and visual fidelity. The combined two‑stage pipeline enables stable trajectory correction for inversion.

Result

On the ArcFace target under an identity‑disjoint cross‑evaluation using CelebA, SFMI attains an accuracy of 0.9248, an FID of 22.61 and an LPIPS of 0.3874, outperforming prior white‑box inversion methods.

Why it matters

Researchers and practitioners concerned with privacy leakage in face recognition systems should care because SFMI demonstrates a powerful white‑box inversion capability that can recover identities with high fidelity.

Method details
  • Flow Matching prior is unconditional and trained on CelebA and FFHQ with image resolutions 64/112/224
  • Time steps are sampled from a Logit‑Normal distribution during prior training
  • Training uses a velocity‑matching loss to align the induced vector field with the OT ground‑truth
  • Progressive Guidance Scheduler injects normalized identity gradients into the vector field during sampling
  • Evaluation uses six white‑box face recognition models (Face.evoLVe, IR‑152, CosFace, ArcFace, MobileFaceNet, ViT)
  • Baselines include GMI, PPA, PLGMI, IFGMI and FGMIA with comparable public auxiliary data
Numbers
  • ACC 0.9248 compared to prior baselines
  • FID 22.61 compared to prior baselines
  • LPIPS 0.3874 compared to prior baselines
Limitations

The paper does not explicitly discuss any limitations.

Under an identity-disjoint cross-evaluation setting using the CelebA dataset, SFMI achieves an ACC of 0.9248, an FID of 22.61, and an LPIPS of 0.3874 on the ArcFace target.Found in the source text, word for word.

Picked because: Steering the Flow releases a gradient‑guided flow‑matching inversion method and code for reconstructing face data from models, giving security engineers a practical tool for vulnerability assessment.