Nikhil Sharma, Yunlin Gong, Xinyang Cheng and 1 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.HC
Problem
Prior work suffered from fragmented reporting of conversational AI systems and study configurations, which prevented reliable replication and extension. Simply improving documentation does not fix the issue because essential details are often missing, hindering faithful replication.
Approach
Gricea provides an open‑science platform that separates a visual authoring client, an execution runtime, and data/publication services. Researchers compose studies using a Study Flow (procedure, assignment, measurement) and a Task Flow (interactive conversational behavior). New capabilities are added via extensible node types with configuration definitions and execution handlers. Participant‑facing components expose configurable properties and emit interaction events that integrate with the existing flows. The platform stores a shared, executable representation that can be inspected, run, and reused.
Result
In a replication study Gricea could represent and execute 93% of eligible CUI 2026 papers, while it identified missing information in 96% of those papers. In a usability study with ten participants, all were able to author runnable studies within a 90‑minute session, indicating the platform’s expressivity and learnability.
Why it matters
Researchers building or replicating conversational AI studies should care because Gricea lowers technical barriers and provides a reusable, inspectable study representation, enabling more reliable and cumulative research.
Method details
Visual authoring client for designing Study Flow and Task Flow.
Execution runtime that runs the configured study and records data.
Data and publication services that version study definitions and execution records.
Extensible node types with configuration schemas and execution handlers.
Separation of Study Flow (assignment, ordering, measurement) from Task Flow (conversational task behavior).
Numbers
replication coverage, 93%, of eligible CUI 2026 papers
missing information flagged, 96%, of papers
participants in usability study, 10, researchers and practitioners
session duration, 90 minutes, total study length
participant compensation, $30, amazon gift cards
papers analyzed in formative analysis, 57, eligible participant‑facing CAI studies
Limitations
The evaluation involved only ten participants and a subspace of methodologies, so it does not capture the full range of prospective researchers and research practices.
We replicated configurations 93% of eligible CUI 2026 papers; while also flagging missing information in 96% of papers that hinder faithful replicationFound in the source text, word for word.
Picked because: Gricea offers an open‑science platform with configurable, deployable research artifacts for conversational AI, giving engineers a ready‑to‑self‑host infrastructure for experiments and reproducibility.
Shuai Bai, Jiayong Deng, Yikun Fu and 29 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.CL
Problem
Prior computer-use agents were split into either graphical interaction or code generation pipelines, and simply stacking those pipelines fails because real digital work requires interleaved exploration, implementation, and verification.
Approach
The authors build RecreationWorld, a five‑platform framework that provides a running reference as an oracle and a unified harness with native GUI control and coding tools. Agents autonomously decide when to explore the interface, write code, launch the artifact, and verify outputs using hidden behavioral tests. Trajectories are generated from high‑quality open‑source applications and used to train models. Evaluation uses two channels: Prog (structured automation assertions) and VLM (visual language model judgments). Scores are macro‑averaged across platforms.
Result
GPT‑6 Astra achieves the highest overall benchmark score of 58.06% but passes all programmatic tests on only 2.8% of tasks. Agents more reliably reproduce static interface structure than dynamic interactions, and generated applications are generally smaller than their references.
Why it matters
Researchers building hybrid computer‑use agents and benchmark designers should care because the work provides a scalable, verifiable environment and shows current models still struggle with full programmatic correctness.
Method details
RecreationWorld spans Ubuntu, macOS, Windows, Android, and Web platforms.
Training data consists of trajectories collected from open‑source applications across the five platforms.
Evaluation uses Prog assertions via AT‑SPI, AXUIElement, UI Automation, or UiAutomator, and VLM judgments from Qwen3.7‑Plus at temperature zero.
Ten frontier models are evaluated, including GPT‑6 Astra, Claude Opus 5, and GPT‑5.6 Sol.
Baseline scores: Claude Opus 5 44.16% overall, GPT‑5.6 Sol 42.06% overall.
Numbers
overall score, 58.06%, GPT‑6 Astra vs 44.16% Claude Opus 5
programmatic test pass rate, 2.8%, GPT‑6 Astra vs lower for other models
recreation size ratio, 16.9%, median recreation‑to‑reference LOC
percentage of recreations smaller than reference, 89.4%, across desktop and Android
window coverage, 5.23 distinct windows per trajectory, GPT‑6 Astra vs 2.21 Qwen3.8‑Max‑0902
Limitations
The paper does not establish that completing the strict final‑loop closure sequence correlates with higher benchmark performance.
GPT-6 Astra obtains the highest overall score at 58.06%Found in the source text, word for word.
Picked because: RecreationWorld provides a released hybrid computer‑use environment that lets developers test and verify agents that combine GUI interaction and code execution, directly supporting LLM agent tooling.
Jagadeesh Balam, Travis Bartley, Edresson Casanova and 46 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.CL
Problem
Prior spoken dialogue systems could not handle full-duplex interaction with integrated tool calling, and simply chaining separate speech, transcription, language, and TTS modules broke the temporal behavior needed for natural conversation.
Approach
NemotronLabs VoiceChat uses a streaming speech encoder feeding a decoder-only language model that produces two parallel output streams: one for agent text and one for structured function calls. An auxiliary RNN‑T branch predicts incremental user transcription from the shared encoder representation, while a streaming TTS decoder renders the agent text into speech. The full-duplex backbone is pretrained on pseudo‑dialogues (CPT) and then fine‑tuned on a weighted mix of turn‑taking, backchannel, and tool‑calling data (SFT). The audio‑codec prediction is disabled during backbone training and later combined at inference with the frozen TTS model. This unified streaming pipeline preserves real‑time turn management while enabling tool use.
Result
On Full‑Duplex‑Bench 1.0 V‑Model achieves the lowest pause‑handling takeover rate (15.3% synthetic, 25.5% CANDOR), 100% user‑interruption takeover, and the highest post‑interruption response quality (4.33/5). Its smooth‑turn takeover is 81.5% with latencies of 0.448 s (smooth) and 0.480 s (interruption). On Full‑Duplex‑Bench 1.5 it resumes after backchannels in 93% of cases. VoiceBench normalized average is 55.1 and Full‑Duplex‑Bench 3.0 tool‑selection F1 is 82.5%. ASR average WER improves from 9.02% (80 ms) to 8.28% (160 ms).
Why it matters
Researchers and engineers building real‑time conversational agents will benefit from a single open‑weight model that integrates speech recognition, generation, and tool calling without sacrificing latency.
Method details
Streaming speech encoder + decoder‑only LLM with parallel agent‑text and function heads.
Auxiliary RNN‑T branch for incremental user transcription.
Streaming TTS decoder trained separately and frozen during backbone training.
Training on 64 GPUs (8 nodes with 8 GPUs each) using bf16 precision.
CPT uses pseudo‑dialogues synthesized from LLM corpora; SFT uses weighted round‑robin sampling of multiple conversational datasets.
ASR evaluated with 80 ms and 160 ms chunk sizes on eight OpenASR benchmark datasets.
Numbers
pause‑handling TOR, 15.3%, lowest among open‑weight systems (synthetic)
pause‑handling TOR, 25.5%, lowest among open‑weight systems (CANDOR)
user‑interruption TOR, 100%, highest among open‑weight systems
response‑quality, 4.33/5, highest among open‑weight systems
resume rate on FDB 1.5, 93%, highest among open‑weight baselines
tool‑selection F1 on FDB 3.0, 82.5%, reported result
average WER (80 ms chunk), 9.02%, ASR benchmark
average WER (160 ms chunk), 8.28%, ASR benchmark
Limitations
Argument accuracy and end‑to‑end tool execution remain areas for improvement.
NemotronLabs VoiceChat combines a streaming speech encoder and decoder-only language model with parallel specialized output streams for agent text and structured function calls.Found in the source text, word for word.
Picked because: NemotronLabs VoiceChat is an open full‑duplex speech‑to‑speech model with native tool‑calling, delivering a self‑hostable stack engineers can integrate into voice‑based AI services.
Gabrielle O'Brien, Reed Milewicz, Nasir Eisty · abstract · pdf
quote verifiedfigures checkedread: full textcs.SE
Problem
There was little evidence about which programming tasks researchers delegate to generative AI and how they verify the generated code; existing validation relied on informal, individual judgment without shared testing infrastructure.
Approach
The authors surveyed 527 researchers, collected free‑text accounts of a single AI‑assisted task, and performed a two‑stage qualitative coding to create use‑case and evaluation codebooks. They then applied these codebooks to all responses, resolved ambiguities through consensus, and conducted quantitative analyses of code frequencies, experience effects, and confidence ratings using Welch's t‑tests and BH‑adjusted p‑values. The method captures both the primary AI use case and the reported validation strategy for each account.
Result
The five most common use cases were data handling (23.9%), visualization (19.8%), debugging (17.4%), mathematical/scientific computing (11.7%) and statistical analysis (10.9%). The most frequent evaluation strategy was running the code (52.8% of accounts), followed by inspecting outputs (22.5%), reading the code (19.4%), inspecting visualizations (16.2%) and using domain knowledge or intuition (11.1%). Confidence in the AI tool was unrelated to programming experience, while solo and evaluation confidence increased with experience.
Why it matters
Researchers and AI tool designers should care because the findings reveal that validation of AI‑generated scientific code is currently informal and individual, highlighting opportunities for better evaluation support in programming assistants.
Method details
527 free‑text survey responses were collected in 2025.
Two parallel codebooks (use‑case and evaluation) were developed through open coding of 106 accounts and iterative refinement.
Coding was performed by three authors with consensus meetings; median pairwise Krippendorff’s alpha with MASI distance was reported.
Quantitative analysis scripts were written with assistance from Claude Code and executed in R.
Statistical tests included Welch's t‑tests and chi‑square tests with Benjamini-Hochberg correction.
Numbers
Data handling, 23.9%, of coded accounts
Visualization, 19.8%, of coded accounts
Debugging, 17.4%, of coded accounts
Mathematical/scientific computing, 11.7%, of coded accounts
Statistical analysis, 10.9%, of coded accounts
Run code, 52.8%, of evaluation accounts
Limitations
The study relies on self‑reported, single‑task accounts and cannot observe the actual validation steps taken by respondents.
The modal strategy by a wide margin was running the code (, 52.8% of accounts)Found in the source text, word for word.
Picked because: The survey of how researchers use and validate AI coding assistants yields concrete verification practices and actionable guidelines for building reliable coding‑assistant pipelines.
Yiming Zhang, Jinghong Zhang, Haoran Zhao and 2 others · abstract · pdf
quote verifiedfigures checkedread: full textcs.CL
Problem
Memory-augmented LLMs retrieve conflicting positions but standard RAG blindly injects all retrieved memories, which markedly increases hallucinations compared to a memory‑free baseline.
Approach
The paper introduces the Memory Decision Layer (MDL), a zero‑parameter controller placed between retrieval and generation. MDL uses a three‑signal complementary encoder that fuses relevance, reliability, and risk via QR‑based orthogonal subspace projection and a meta‑working‑memory signal. It explicitly decouples confidence from consistency, applies risk inversion, and can abstain when risk is high. The resulting decision representation quantifies trustworthiness of each memory before generation.
Result
MDL cuts the hallucination rate under conflicting memories by about 56.04% in general scenarios and nearly eliminates hallucinations in high‑risk scenarios; decision accuracy reaches 61.2% for the full three‑signal combination (62.7% for a subset) versus 0.581 for any single signal.
Why it matters
Developers of retrieval‑augmented LLM pipelines should adopt MDL to obtain low‑latency, interpretable trust decisions that markedly suppress hallucinations.
Hyperparameters: subspace dimension 16, sharpness of value encoding 0.20, principal‑gain coefficient 0.10, risk subspace extra dimension (unspecified numeric).
Latency: MDL adds ~0.14 ms per decision, ~50× faster than the preceding embedding‑retrieval step and 4 to 5 orders of magnitude faster than an LLM self‑evaluation call.
Numbers
hallucination reduction, 56.04%, compared to standard RAG under conflicting memories
decision accuracy, 61.2%, full three‑signal MDL vs. 0.581% for single‑signal baseline
decision accuracy, 62.7%, best two‑signal combination vs. 0.581% single‑signal
latency addition, 0.14 ms per decision, 50× faster than embedding‑retrieval
latency advantage, four to five orders of magnitude faster than LLM self‑evaluation
ablation cost, 1.7 pp, when removing bounded second‑order risk modulation
Limitations
The evaluation is limited to a few open‑source LLMs and two benchmark datasets, so generalization to other models or domains is not demonstrated.
The controller is fully white-box: it relies purely on geometric operations, requires no trained parameters, and adds only about 0.14 ms per decision.Found in the source text, word for word.
Picked because: The Interpretable Memory Decision Controller introduces a practical memory‑trust mechanism for retrieval‑augmented LLM agents, with released code that engineers can adopt to reduce hallucinations.