Girish G N, Ashutosh Sahoo, Akshay SP and 2 others · abstract · pdf
quote verifiedfigures checkedread: full textq-fin.RM
Problem
Total wallet value conflates liquid and illiquid assets, causing mispricing of risk. Using total value alone fails because it cannot distinguish spendable reserve from non‑spendable holdings.
Approach
zLend reconstructs a wallet's daily balance twice: once over a fixed stablecoin basket and once over all fungible tokens using a cumulative‑net‑flow algorithm with a non‑negativity offset. From each series it computes liquidity coverage, cash‑flow regularity, peak‑to‑trough drawdown and recovery, outflow concentration, trend classification, and a recurring‑counterparty detector. The two scopes are then compared to produce liquidity‑ and flow‑mismatch severity scores. These scores feed a four‑tier underwriting rule specified in Tables 2 and 3. The entire pipeline is implemented in Python and ported to TypeScript with strict numerical verification.
Result
The analysis shows tier assignment is governed predominantly by the reference loan size, with four of six reference wallets changing tier across loan sizes. Drawdown and coverage criteria bind on disjoint wallets, confirming neither subsumes the other, and no tier‑rule criterion is inert. The framework is deployed in production and feeds real lending decisions via API integration.
Why it matters
DeFi platforms and on‑chain credit providers should care because zLend offers a verified, production‑ready way to separate liquid reserve from total wealth and generate robust underwriting signals.
Method details
Daily balance reconstruction uses a cumulative‑net‑flow construction with a non‑negativity offset (Algorithm 1)
Two parallel scopes are built: a stablecoin‑only view and a total‑wealth view over all fungible transfers
Signal family includes liquidity coverage, cash‑flow regularity, drawdown and recovery, outflow concentration, trend, and recurring‑counterparty detection
Tier assignment follows a four‑tier rule (Tables 2 and 3) driven mainly by the reference loan size
Implementation verification replicates compensated (Kahan) summation, population variance, round‑half‑to‑even to ten decimal places, and null propagation for near‑zero denominators
Validation achieved 78 of 78 field assertions against deployed system reference fixtures
Numbers
numerical tolerance, 1e-9, compared against reference implementation
field assertions, 78 of 78, exact agreement with reference fixtures
tier changes, four of six reference wallets, changed tier across loan sizes USD 10 to USD 25,000
record size, 166 columns, per evaluation record written to storage
scope fields, 69 fields per scope, plus six cross‑scope ratios
divergence magnitude, four orders of magnitude, between total‑wealth and stablecoin reconstructions
Limitations
The paper does not evaluate predictive performance on actual default outcomes; it focuses on methodology and verification.
tier assignment is governed predominantly by the reference loan size, with four of six reference wallets changing tier across loan sizesFound in the source text, word for word.
Picked because: zLend provides a deployed on-chain cash‑flow reconstruction framework with open‑source implementation, giving engineers a ready‑to‑use tool for credit underwriting automation.
Enric Boix-Adsera, Benedict Tessler · abstract · pdf
quote verifiedfigures checkedread: full textcs.CL
Problem
Prior work assumed that only strong, obvious prompts could reliably steer model outputs, and it ignored the impact of many weak, seemingly irrelevant cues; simply removing obvious cues does not prevent steering because inconspicuous paraphrases and typos can still combine to control behavior.
Approach
The authors treat each cue as an additive component to the log‑odds of a target answer and fit a linear model using ridge regression on random prompt configurations. This additive model predicts how stacked weak cues will shift probabilities. They then rank prompt configurations by the predicted log‑odds sum and select the top and bottom candidates as extreme prompts. By validating these extremes with fresh generations they demonstrate that stacked cues can flip model responses. The same pipeline is applied to both non‑reasoning and reasoning models.
Result
The additive model captures most of the variance in held‑out configurations, achieving R² values from 0.3 to 0.99 (median 0.75). Stacked weak cues can reliably flip the modal answer in both non‑reasoning and reasoning models, and the directional effect often transfers across models, especially within families.
Why it matters
AI safety and interpretability researchers should care because even innocuous textual variations can strongly steer model outputs, creating new attack surfaces and challenges for reliable deployment.
Additive model fitted to log‑odds using ridge regression on ~10^3 random prompt configurations per model‑cue‑effect cell.
Baseline comparison: random‑prompt additive fit versus measured log‑odds; steering range measured by top/bottom 100 extreme prompts.
Transfer experiment: prompts optimized on one source model evaluated on 16 non‑reasoning target models to assess directional effect preservation.
Numbers
held‑out configuration‑level R² spanning roughly 0.3 to 0.99
5th‑95th percentile R² 0.54‑0.93
median R² 0.75
top‑ and bottom‑100 candidate prompts evaluated per model
Qwen3‑8B evaluated at three token budgets (256, 1024, 4096)
transfer significant for animal cues and some phrasing and JSON cues
Limitations
The paper evaluates a limited set of reasoning models due to API cost and does not prove that the additive approximation holds for all possible cue types or model families.
the additive model explains most of the held‑out configuration‑level variance, with held‑out configuration‑level R2 spanning roughly 0.3 to 0.99 (5th-95th percentile 0.54-0.93; median 0.75)Found in the source text, word for word.
Picked because: Model Hypnosis uncovers systematic prompt‑level control of LLMs and releases code to test and verify such effects, directly supporting LLM agent safety tooling.
Junjie Chu, Ye Leng, Mingjie Li and 3 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.LG
Problem
Existing GEO detection methods achieve decent aggregate F1 but fail to generalize across GEO optimizer families and authorship groups, often relying on authorship shortcuts; simple calibrations or zero‑shot LLMs do not close these gaps.
Approach
The paper introduces Intervention‑Paired Training (IPT), which supervises detector responses to paired GEO interventions and non‑GEO AI polishing, enforcing positive and invariant pair constraints. IPT is applied to fine‑tuned ModernBERT (and Qwen) encoders. The method is integrated into a GEO‑gated Agent system that first flags pages, then audits citation URLs through deterministic preprocessing, labeling, and verifiability aggregation modules. Together these components detect GEO pages and assess citation reliability without needing provenance metadata.
Result
IPT substantially improves detection performance: ModernBERT F1 climbs to 0.944 and worst‑group accuracy to 0.883, while Qwen‑IPT also gains accuracy and F1. Baseline lexical models still reach competitive F1 (0.878 to 0.880), but exhibit large shortcut gaps. The system estimates GEO prevalence of 8.90% overall and 16.36% among pages modified in 2026.
Why it matters
Search engine operators and auditors should care because the method provides a practical way to flag and audit GEO‑optimized content, improving the reliability of generative search results.
Query‑level grouped split: 2,238 training documents and 962 test documents.
Fine‑tuned ModernBERT with 1,024‑token and 8,192‑token contexts; Qwen3‑0.6B SFT also used.
IPT raises ModernBERT accuracy from 0.839 to 0.931 and F1 from 0.862 to 0.944.
TF‑IDF word (1‑2) + LR baseline achieves highest aggregate F1 of 0.880.
Experiments run on Google Cloud g2‑standard‑8 (8 vCPUs, 32 GB RAM, NVIDIA L4 24 GB GPU).
Numbers
F1 0.880 (TF‑IDF word + LR) compared to other baselines
Accuracy 0.931 (ModernBERT‑IPT) vs 0.839 baseline
F1 0.944 (ModernBERT‑IPT) vs 0.862 baseline
Worst‑group accuracy 0.883 (ModernBERT‑IPT) vs 0.725 baseline
GEO prevalence 8.90% overall, 16.36% among 2026‑modified pages
Aggregate baseline F1 0.880 (highest among non‑IPT methods)
Limitations
The paper does not claim to detect malicious intent, factual inaccuracy, or to generalize beyond the eight studied GEO optimizer families.
IPT improves both overall performance and shortcut diagnostic metrics.Found in the source text, word for word.
Picked because: GEO‑Flag introduces a detection pipeline for generative‑search‑optimized web content and releases datasets and classifiers that engineers can integrate into CI/CD for SEO auditing.
Knowledge‑graph stores assume defaults: accept‑then‑clean, single or no time axis, equal trust for all writers, and external governance. These defaults cause hidden defects when multiple agents write concurrently. Simply adding pre‑state checks does not catch defects that only appear in the combined post‑state.
Approach
Quipu implements a governed store that inverts the four defaults. Writes are filtered by a gate that evaluates predicates against the pending post‑state (GS1). Gate outcomes are persisted as signed, time‑indexed verdicts (GS2). Authority is attached to named‑graph partitions with a lattice that never widens when composed (GS3, GS4). The trace, verdicts, and policies live in‑store, making audit decidable (GS5), and every decision can be replayed as‑of its transaction (GS6). Together these mechanisms provide bitemporal EAVT storage with three‑valued operations and a non‑widening label lattice.
Result
The gated store eliminated all planted defects (0 of 6) while the ungated store exhibited all six. All seven composition probes upheld the lattice contract. Fifty of fifty verdicts re‑derived faithfully at their instant, and the DEMM‑Bench content‑only reader reconstructed all 512 property‑level questions with zero overclaim, whereas container‑presence baselines overclaimed up to 87.5%.
Why it matters
Developers of knowledge‑graph stores and autonomous agents that write KG data should care because Quipu provides provable governance, auditability, and defect prevention without post‑hoc cleaning.
Method details
Dataset: Census benchmark, a deterministic multi‑writer lifecycle with seeded run 42.
The paper does not evaluate Quipu on large‑scale real‑world workloads beyond the synthetic Census and DEMM benchmarks.
the gated store ends with 0 of 6 planted defects versus 6 of 6 ungatedFound in the source text, word for word.
Picked because: Quipu delivers a governed bitemporal knowledge‑graph store with source code, addressing data governance and versioning challenges in production back‑ends.
Huatong Song, Fei Bai, Ming Yang and 17 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.CL
Problem
Scaling reinforcement learning through complex, opaque harnesses for long-horizon tasks was unstable and largely unexplored, and naïvely applying standard RL pipelines fails due to infrastructure failures, invalid token generation, and mismatched advantage estimation across heterogeneous harnesses.
Approach
The method builds a sandbox execution layer that isolates environments and harnesses, adds a serving proxy at the model boundary to capture all model calls, organizes these calls into prefix trees, and then applies adapted PPO (critic‑based) or GRPO (critic‑free) optimization over the recovered tree structure; it also introduces mix‑harness training by jointly sampling task-harness pairs from multiple harnesses within each batch while keeping advantage groups separate.
Result
Black‑box RL with Qwen3-30A3B raises Pass@1 on ClawGym-Bench by 9.98 points using OpenClaw and by 14.81 points using Claude Code, while maintaining stable optimization for 200 to 400 steps; mix‑harness training yields training rewards and downstream evaluation comparable to or slightly better than single‑harness baselines.
Why it matters
Researchers and engineers building long‑horizon agents that rely on external tool harnesses should care because the framework enables stable, scalable RL without redesigning each harness, and it can jointly leverage heterogeneous execution systems.
Method details
Model: Qwen3-30A3B used as the policy backbone.
Benchmarks: ClawGym-Bench for Pass@1 evaluation, plus ClawGym-SynData, JobBench, and OfficeQA for training and testing.
Mix‑harness batch configuration: 32 task‑harness instances with 8 rollouts per instance.
Training stability observed over approximately 200 to 400 optimization steps.
Baseline comparisons: individual‑harness training with OpenClaw or Claude Code.
Numbers
Pass@1 improvement, 9.98, OpenClaw
Pass@1 improvement, 14.81, Claude Code
Optimization steps, 200 to 400, stable training
Batch size, 32 task‑harness instances, 8 rollouts per instance
Training reward, comparable, mix‑harness vs individual harness
Limitations
The paper does not claim to establish performance beyond the presented models and benchmarks, nor does it address scalability to arbitrarily larger models or entirely new harness types.
black-box RL improves Pass@1 on ClawGym-Bench by 9.98 and 14.81 points through OpenClaw and Claude Code, respectivelyFound in the source text, word for word.
Picked because: ClawGym II offers a black‑box reinforcement‑learning harness with open‑source benchmarks, enabling engineers to experiment with agent‑level automation without modifying environments.
Ruoqi Shu, Xuhui Wang, Isaac Wang and 2 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.AI
Problem
Existing pipelines cannot reliably process heterogeneous layouts, semantically rich content, and embedded business rules, leading to hallucinations and false positives; simply applying end-to-end LLMs fails because they lack layout preservation and explicit logical grounding.
Approach
LAVA is a modular, backbone-agnostic pipeline built on multimodal large language models that follows a four-stage design: document‑rule retrieval, layout‑preserving information extraction, auxiliary metadata enrichment, and auditable symbolic/arithmetic verification. It runs two parallel tracks-document processing and rule grounding-that exchange constraints. The document track produces structured content with enriched metadata, while the rule track retrieves, classifies, and dispatches rules to either a Symbolic Reasoner or an Arithmetic Processor based on reasoning type. Knowledge Extraction (KE) provides structured markup, Info Augmentation (IA) adds semantic metadata, and the Arithmetic Processor (AP) enforces numerical fidelity. The system is fully modular, allowing independent updates to each component.
Result
LAVA achieves the lowest hallucination and numerical infidelity rates among all baselines and dramatically reduces edge‑case error, while using fewer tokens per rule check. In the ablation study, removing KE raises error to 0.65 to 0.67, whereas full LAVA maintains error below 0.07 across rule categories.
Why it matters
Enterprises that need high‑accuracy, auditable financial document validation should consider LAVA because it reduces hallucinations and false positives while remaining efficient.
Method details
Uses multimodal large language models as a backbone, but is backbone‑agnostic.
Evaluated on a proprietary Canadian mortgage application dataset of ~1,000 scanned PDFs/images with dozens of expert‑curated validation rules.
Ablations include removal of KE, IA, and AP, and variants with plain‑text and Markdown KE.
Metrics include Factual Hallucination Rate, Hallucination Rate, Numerical Infidelity Rate, Edge‑Case Error Rate, and token cost.
Numbers
Factual Hallucination Rate 0.01 vs 0.03 (VLM) and 0.03 (LLM)
Hallucination Rate 0.03 vs 0.08 (VLM) and 0.10 (LLM)
Numerical Infidelity Rate 0.01 vs 0.08 (VLM) and 0.04 (LLM)
Edge Case Error Rate 0.02 vs 0.27 (VLM) and 0.25 (LLM)
Ablation LAVA w/o KE error 0.65 vs LAVA 0.01
Ablation LAVA w/o IA error 0.45 vs LAVA 0.05
Limitations
The paper does not explicitly discuss limitations.
LAVA outperforms baselines in hallucination control and edge-case handling while maintaining efficient token usageFound in the source text, word for word.
Picked because: LAVA presents a logic‑aware validation and augmentation framework for large‑scale financial document auditing, complete with a released pipeline that can be deployed in enterprise workflows.
Siddarth Venkatraman, Matthieu Dinot, Laurence Aitchison · abstract · pdf
quote verifiedfigures checkedread: full textcs.LG
Problem
Group-relative RL methods for LLMs suffer high gradient variance, straggler rollouts, and off‑policyness, and naïve addition of learned value functions is hindered by infrastructure complexity and the practical success of critic‑free pipelines.
Approach
The paper introduces Privileged Value Functions (PVF) that condition the value head on task‑specific privileged context without biasing the policy objective, and TETHER, a baseline that adaptively interpolates between the group‑mean GRPO baseline and the token‑level value baseline based on the accuracy of the value function. Both methods reuse the same Monte‑Carlo advantage computation and share the value‑training configuration. PVF provides token‑level variance reduction using extra information, while TETHER dynamically weights the mean and value baselines during training.
Result
PVF consistently outperforms both the ordinary value baseline and the group‑mean baseline across Reasoning Gym, CodeIO, and Sudoku, achieving the highest end‑of‑training rewards in all four experiments. Tether improves over the simple value baseline on all four tasks and is competitive with the Mean baseline, beating Mean on Reasoning Gym and MiniF2F, matching it on CodeIO, and narrowing the gap on Sudoku.
Why it matters
Researchers building RL pipelines for large language models should consider PVF for better variance reduction when privileged information is accessible, and Tether for a robust way to incorporate value functions into existing GRPO workflows.
Method details
Qwen3-4B-Instruct-2507 is used for Reasoning Gym, CodeIO, and Sudoku; Qwen3.5-4B is used for MiniF2F
Tasks include Reasoning Gym, CodeIO, Sudoku, and MiniF2F
Baselines compared are Mean (group‑mean GRPO), VF (ordinary token‑level value), PVF, and Tether
Each run performs 20 value‑warmup updates before policy training
Policy steps: 800 for Reasoning Gym, 650 for CodeIO, 600 for Sudoku, 500 for MiniF2F
Numbers
20 value warmup updates before policy training
800 policy steps for Reasoning Gym
650 policy steps for CodeIO
600 policy steps for Sudoku
batch size 128 for Reasoning Gym and CodeIO
Limitations
Tether does not fully recover the performance of the Mean baseline on Sudoku, and PVF requires privileged context that may not be available for all tasks.
Pvf is the best performing method in all settings.Found in the source text, word for word.
Picked because: Le Critique proposes privileged value‑function techniques for LLM reinforcement learning and provides an implementation that reduces variance in policy training, useful for building reliable LLM agents.
David Eric Austin, Kaheer Suleman, Jackie Chi Kit Cheung · abstract · pdf
quote verifiedfigures checkedread: full textcs.CL
Problem
Prior work assumed LLM agents would explore like classical bandit algorithms, but semantic labels attached to actions bias their exploration-exploitation balance; simply removing explicit semantics does not fix the issue because LLMs retain implicit semantic priors from pre‑training.
Approach
The paper introduces the semantic bandit, extending the multi‑armed bandit with textual action labels. It evaluates four nomenclatures (alphanumeric, sentiment, ordinal, world knowledge) under helpful and misleading reward configurations. Models are prompted with a best‑performing CoT prompt and optionally with explicit bias warnings or exploration instructions. Performance is compared to a label‑agnostic UCB1 baseline. Normalized cumulative regret and exploration counts are measured to quantify bias effects.
Result
Semantic priors strongly shape LLM behaviour: when labels align with reward (helpful), regret drops dramatically (e.g., Qwen3-32B ordinal helpful 0.00) and exploration is suppressed; when labels are misleading, regret rises (Qwen3-32B ordinal misleading 0.79) and exploration increases. The alphanumeric control matches UCB1 (regret 0.15) showing minimal bias. Negative reward values trigger more exploration than equivalent positive values.
Why it matters
Researchers building LLM‑based agents for interactive decision tasks should account for semantic label bias, as it can dramatically affect exploration and overall performance.
Semantic bandit tasks in farming and clothing domains with action labels sampled per run.
Four nomenclatures: alphanumeric (random strings), sentiment (valence adjectives), ordinal (rank labels), world knowledge (domain‑specific labels).
Baseline: classical UCB1 algorithm with 10,000 replicates per condition.
Prompt variations: standard CoT prompt, explicit bias warning, and explicit exploration instruction.
Experiments run with 10 replicates per condition (20 for reward‑scale sweep) and three variance levels (high, low, none).
Numbers
Normalized cumulative regret (Qwen3-32B ordinal helpful) 0.00 vs UCB1 0.15
Normalized cumulative regret (Qwen3-32B ordinal misleading) 0.79 vs UCB1,
Normalized cumulative regret (OLMo-3.1 ordinal helpful) 0.04 vs UCB1 0.15
Normalized cumulative regret (Gemini ordinal helpful) 0.13 vs UCB1 0.15
Normalized cumulative regret (Qwen3-32B sentiment helpful) 0.09 vs UCB1 0.15
Normalized cumulative regret (Qwen3-32B alphanumeric helpful) 0.16 vs UCB1 0.15
Limitations
The study is limited to synthetic semantic bandit tasks and does not demonstrate how the findings generalize to more complex real‑world decision‑making environments.
semantically informative action labels reduce exploration in favour of exploitation, improving performance when aligned with the reward structure and severely degrading it when misaligned.Found in the source text, word for word.
Picked because: Semantic Bandits analyses how in‑context LLM prompts bias exploration‑exploitation and supplies a toolkit for measuring and correcting such biases in deployed agents.
Homa Esfahanizadeh, Matin Mortaheb, Jinfeng Du and 1 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.LG
Problem
Task‑specific codecs are brittle and retraining a separate codec for each downstream task is infeasible in the field.
Approach
UniTAC trains a single Vision Transformer based image codec whose encoder and decoder are conditioned on a per‑component importance vector. The importance vector, derived from gradient attribution of any downstream model, is sent as low‑overhead side information. During training the model sees a broad family of synthetic importance maps so it learns to allocate bits according to the injected weights. At runtime the same backbone is re‑targeted to any task by swapping in the task‑derived importance map, without any retraining.
Result
At 0.034 bpp a single UniTAC model attains 91.4% downstream classification accuracy, only 1.9% lower than a dedicated task‑based codec (93.3%) and substantially higher than a universal codec (76.9%). It also matches the universal codec on overall PSNR while improving semantic PSNR and task accuracy at equal or lower rate.
Why it matters
Researchers and engineers building physical AI systems such as autonomous vehicles and robots should care because UniTAC enables flexible, task‑aware compression without per‑task retraining.
Method details
Architecture: two‑stage ViT autoencoder with token‑level weight conditioning and a hyperprior entropy model.
Training data: codecs trained on AffectNet and evaluated on the CelebA test split.
Importance maps: generated per image by integrated gradients of a downstream ResNet‑18 classifier.
Baselines: compared against a non‑semantic universal codec and a task‑based semantic codec, all sharing the same ViT architecture.
Training objective: rate-distortion loss combining a rate estimate with the separable weighted distortion normalized by total weight (Eq. 11).
Synthetic importance maps: random mixtures of Gaussian blobs of varying count, location, and scale.
Numbers
0.034 bpp
91.4% accuracy (UniTAC)
93.3% accuracy (task‑based codec)
76.9% accuracy (universal codec)
Limitations
The paper does not fully quantify the universality gap and acknowledges that performance depends on the quality of the supplied importance map.
On a localized task at 0.034 bpp, a single UniTAC model reaches 91.4% accuracy, only 1.9% below a task-based codec (93.3%) and above universal codecs (76.9%).Found in the source text, word for word.
Picked because: UniTAC introduces a universal task‑aware image codec with a released model that can be self‑hosted to adapt compression to changing downstream tasks, simplifying infrastructure for edge AI.
Ye Lu, Shen Wang, Zhaoyang Zhang and 4 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.CV
Problem
Existing model inversion attacks rely on indirect or highly stochastic guidance, which makes it difficult to stably optimize generation trajectories toward target facial images.
Approach
SFMI first trains an unconditional Flow Matching model to learn a smooth vector field that maps Gaussian noise to the human‑face manifold. Then, during sampling, a Progressive Guidance Scheduler injects time‑dependent target‑specific gradients computed by back‑propagating through the face recognition model into the learned velocity field. The gradients steer the generative flow from random noise toward high‑density regions of the target identity. PGS dynamically modulates guidance strength over time to balance identity alignment and visual fidelity. The combined two‑stage pipeline enables stable trajectory correction for inversion.
Result
On the ArcFace target under an identity‑disjoint cross‑evaluation using CelebA, SFMI attains an accuracy of 0.9248, an FID of 22.61 and an LPIPS of 0.3874, outperforming prior white‑box inversion methods.
Why it matters
Researchers and practitioners concerned with privacy leakage in face recognition systems should care because SFMI demonstrates a powerful white‑box inversion capability that can recover identities with high fidelity.
Method details
Flow Matching prior is unconditional and trained on CelebA and FFHQ with image resolutions 64/112/224
Time steps are sampled from a Logit‑Normal distribution during prior training
Training uses a velocity‑matching loss to align the induced vector field with the OT ground‑truth
Progressive Guidance Scheduler injects normalized identity gradients into the vector field during sampling
Evaluation uses six white‑box face recognition models (Face.evoLVe, IR‑152, CosFace, ArcFace, MobileFaceNet, ViT)
Baselines include GMI, PPA, PLGMI, IFGMI and FGMIA with comparable public auxiliary data
Numbers
ACC 0.9248 compared to prior baselines
FID 22.61 compared to prior baselines
LPIPS 0.3874 compared to prior baselines
Limitations
The paper does not explicitly discuss any limitations.
Under an identity-disjoint cross-evaluation setting using the CelebA dataset, SFMI achieves an ACC of 0.9248, an FID of 22.61, and an LPIPS of 0.3874 on the ArcFace target.Found in the source text, word for word.
Picked because: Steering the Flow releases a gradient‑guided flow‑matching inversion method and code for reconstructing face data from models, giving security engineers a practical tool for vulnerability assessment.