Skip to content

arxiv-daily

Automated deployment @ 2026-10-05 00:22:35 Asia/Taipei

Welcome to contribute! Add your topics and keywords in topic.yml. You can also view historical data through the storage.

AI

Knowledge Graphs

Publish Date Title Authors Homepage Code
2026-10-01 KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards Pengfei Li et.al. 2610.02206v1 null
2026-10-01 Higher-Order Molecular Grammars for Generative and Foundation Models in Chemistry Yiming Huang et.al. 2610.02186v1 null
2026-10-01 From Knowledge Access to Source Learning: Developing Source-Specific Competence Lucheng Fu et.al. 2610.02150v1 null
2026-10-01 Mem++: Non-Destructive Memory for Long-Term Organizational LLM Agents Ahmad Yehia et.al. 2610.02002v1 null
2026-10-01 Can AI Oversight Be Zero Knowledge? Alessandro Chiesa et.al. 2610.01995v1 null
2026-10-01 Latent JEPA: Abstract Future Prediction for Latent Reasoning in Chemistry Xinjian Zhao et.al. 2610.01947v1 null
2026-10-01 Mapping the RAG Landscape: A Four Axis Taxonomy of Efficiency, Defense, Interactivity, and Reasoning Meghana Sunil et.al. 2610.01936v1 null
2026-10-01 From Network Intrusion Detection to Blockchain-Backed Endpoint Detection and Response: Mapping the Landscape of Decentralized Detection-and-Response Architectures Yahya Shahsavari et.al. 2610.01872v1 null
2026-10-01 Walking the Embedding Space: Datastore Extraction from Multimodal RAG Maria Carmen Jica et.al. 2610.01871v1 null
2026-10-01 Detecting Inconsistencies in Model Specifications with LLM-as-Verifier Reasoning Zichen Xie et.al. 2610.01847v1 null
2026-10-01 Code Owns the Simulation, Jev Owns the Evaluation Yaodong Yang et.al. 2610.01834v1 null
2026-10-01 The Asymptotics of Language Model Alignment with Memory Haricharan Balasundaram et.al. 2610.01828v1 null
2026-10-01 A Matryoshka Hierarchical RAG for Efficient Multi-Hop Question Answering Gianluca Bonifazi et.al. 2610.01767v1 null
2026-10-01 Task-Oriented Rank Adaptation for Continual Learning in Text Classification Rey Sanchez Lopez et.al. 2610.01702v1 null
2026-10-01 Iterative Policy Refinement through Semantic Rollout Analysis Feiyu Gavin Zhu et.al. 2610.01652v1 null
2026-10-01 Managing Context and Communication in Distributed Agentic UAV Swarms Andrea Iannoli et.al. 2610.01569v1 null
2026-10-01 From Rules to Neural Graphs: Scalable Structured Prediction for Patent Prior Art Search Nikolai Zenovkin et.al. 2610.01553v1 null
2026-10-01 Auto-Formalizing Neuro-Symbolic Predictors Samuele Bortolotti et.al. 2610.01519v1 null
2026-10-01 Decision Titan: Test-Time Training for Long-Term Memory in Offline Reinforcement Learning Jude Waide et.al. 2610.01513v1 null
2026-10-01 OverAct: Measuring and Mitigating Proactive Over-Authorization in LLM Tool-Calling Agents Taolin Zhang et.al. 2610.01508v1 null
2026-10-01 A Multi-Agent LLM Framework for Personalized Health Checkup Interpretation and Guidance HyungJun Kim et.al. 2610.01451v1 null
2026-10-01 LLM-Assisted Discovery of Typed Semantic Links for Ontology Network Construction Nouha Hayouni et.al. 2610.01393v1 null
2026-10-01 Generation Provenance Before Behavior Attribution: Auditing Synthetic Speech Research Objects Sidi Chang et.al. 2610.01378v1 null
2026-10-01 ARCCS: An Automated Regulatory Compliance Checking System Giorgos Filandrianos et.al. 2610.01345v1 null
2026-10-01 An ontology for cross-sectoral crisis management: core and public health modules Aldo Gangemi et.al. 2610.01326v1 null
2026-10-01 Dependency-Aware Reward Shaping for Agentic Reinforcement Learning Ziyi Chen et.al. 2610.01207v1 null
2026-10-01 Federated Agent Optimization Qiang Yang et.al. 2610.01195v1 null
2026-10-01 Temporally-Resolved Token Attribution Reveals the Generation Dynamics of Diffusion Language Models Darpan Aswal et.al. 2610.01177v1 null
2026-10-01 OrbitTAMP: Grounding Language Models for Task and Motion Planning in Spacecraft Rendezvous Yuji Takubo et.al. 2610.01093v1 null
2026-10-01 MOMAT: Mixture of Multiple Atlases for Low-Power Jailbreak Defense of Quantized LLMs Boyang Li et.al. 2610.01058v1 null
2026-10-01 Capturing In-Context Learning Dynamics with Task Operators Guangzhi Xiong et.al. 2610.01054v1 null
2026-10-01 Beyond Answer Confidence: A Controlled Audit of Self-Knowledge in a Black-Box Decision Model Sharath M Shankaranarayana et.al. 2610.01006v1 null
2026-10-01 Distilling Directional Verification Jungseob Lee et.al. 2610.00997v1 null
2026-10-01 Structure-agnostic Causal Representation Learning Arman Behnam et.al. 2610.00968v1 null
2026-10-01 ABDA-NL: A Natural-Language Scenario Explorer for Argument-Based Reasoning Shawn Bowers et.al. 2610.00947v1 null
2026-10-01 Screw Attention: Rigid-Body Algebra Inside a Transformer Aly Magassouba et.al. 2610.00904v1 null
2026-10-01 Match the Distribution, Not the Compute: Post-Training Multi-Token Prediction Heads Prachi Badarayani et.al. 2610.00888v1 null
2026-09-30 Backdoor Purification for LoRA-Tuned LLMs via Null-Space Projection Jianwei Li et.al. 2610.00685v1 null
2026-09-30 Ontology-Grounded, Reasoner-Verified Benchmarks for Evaluating LLM Reasoning in Scientific AI Nishtha N. Vaidya et.al. 2610.00682v1 null
2026-09-30 PhysicsMate: A Curriculum-Grounded Bengali Benchmark for Secondary Physics QA with Small-Model Adaptation Rashid Azraf Jahin et.al. 2610.00664v1 null
2026-09-30 Where's Waldo? Query-language Preference under Cross-lingual Knowledge Disparities Dayeon Ki et.al. 2610.00606v1 null
2026-09-30 Emergent Unfaithfulness: How Alignment Training Causes Language Models to Silently Override Task Faithfulness Pardis Sadat Zahraei et.al. 2610.00568v1 null
2026-09-30 Assessing the Impact of Language Disparity on Multilingual Linguistic Ability in Large Language Models Zhanyu Chen et.al. 2610.00540v1 null
2026-09-30 EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery Young-Jun Lee et.al. 2609.40340v1 null
2026-09-30 Linguistic Loopholes in LLM Unlearning: From a 174-Language Benchmark to Coverage-Aware Unlearning Tyler Skow et.al. 2609.40286v1 null
2026-09-30 EviRover: Reinforcing Agentic Perception Beyond a Glance Kaixuan Fan et.al. 2609.40230v1 null
2026-09-30 Learning from Research: Toward Lifelong Agent Harness Evolution Jingbo Yang et.al. 2609.40169v1 null
2026-09-30 On the (In)effectiveness of AMR Augmentation for Large Language Models Hoa Quynh Nhung Nguyen et.al. 2609.40121v1 null
2026-09-30 Persistent Context Graphs for Efficient Memory Compaction in LLM Agents Jingbo Yang et.al. 2609.40118v1 null
2026-09-30 JuryFlow: Disagreement-Guided Human-in-the-Loop Multi-Agent Evaluation Mufeng Yang et.al. 2609.40103v1 null
2026-09-30 AutoDataBench: A Data-centric Testbed for Accelerating Auto Research Ruifeng Yuan et.al. 2609.40097v1 null
2026-09-30 LongEmo: Towards Emotion Understanding and Reasoning in Long Videos Shuo Zhang et.al. 2609.40079v1 null
2026-09-30 TACTIC: Temporal and Context-Aware LLM Tactical Planning for Roadside LiDAR Attacks Yiming Gao et.al. 2609.39969v1 null
2026-09-30 AIMS: An Agentic AI Framework for Sim-to-Real Multi-Modal ISAC Yijie Bian et.al. 2609.39964v2 null
2026-09-30 Learning to Cover Locally: Graph Neural Combinatorial Optimization under a Hard Information Horizon Johannes F. Loevenich et.al. 2610.00422v1 null
2026-09-30 MCD: Causal Distillation of Multimodal In-Context Learning in Large Vision-Language Models Yanshu Li et.al. 2609.39920v1 null
2026-09-30 The Concrete-Arbitrary Gap: Kinship Reasoning in LLMs Is Not Indifferent to Presentation Thomas Pashby et.al. 2609.39913v1 null
2026-09-30 DoGBench: Can Agents Meet Expert Standards for User-Facing Documentation? Frances Liu et.al. 2609.39909v1 null
2026-09-30 Cognitive Enhancement: Rethinking the Necessity of Role-Playing for Large Language Models Xingjie Zhuang et.al. 2609.39853v1 null
2026-09-30 Explore-on-Graph: Hybrid Embedding-LLM Reasoning for Knowledge Graph Question Answering under Incompleteness Ola El Khatib et.al. 2609.39786v1 null
2026-09-30 MemCodex: Self-Programming Hierarchical Memory for Language Agents Xiaoqiang Wang et.al. 2609.39765v1 null
2026-09-30 OverForge: Reasoning Through Strategies and Tactics Helps Cooperative Lifelong Adaptation Oana Madalina Fron et.al. 2609.39727v1 null
2026-09-30 ArchitectureIQ: On the Measure of Training Intuition Zirui Ren et.al. 2609.39714v1 null
2026-09-30 ChronoGraph: Functional 4D Scene Graphs with Vision-Language Models for Interaction Understanding and Grounded Planning Chenyangguang Zhang et.al. 2609.39665v1 null
2026-09-30 Zero-Compute Cross-Lingual Transferability Estimation Using Typological Feature Proxies Dalton Raphael Harmsen et.al. 2609.39640v1 null
2026-09-30 RankEvolve: A Reliable Multi-Agent Auto-Research Harness for Evolving Ranking Models Zheng Chen et.al. 2609.39551v1 null
2026-09-30 Learning Normal Diffusion Dynamics for Backdoor Defense in Text-to-Image Models Junjian Li et.al. 2609.39548v1 null
2026-09-30 A Reusable Semantic Web Framework for Evidence-Grounded Fundamental Rights Impact Assessments under the EU AI Act Faith Olopade et.al. 2609.39537v1 null
2026-09-30 Beyond the Shadows of Plato's Cave: Evaluating False Memory in Autonomous Agents via Counterfactual Reasoning Quan M. Tran et.al. 2609.39473v1 null
2026-09-30 CAST: Causal Advantage-Structured Training with Spatially Grounded Compositional Rewards for Diffusion Models Shu Yu et.al. 2609.39441v1 null
2026-09-30 Inferring Causal Relations between Two Sequences of Events with Language Models Nishchal Prasad et.al. 2609.39406v1 null
2026-09-30 Exploring Heterogeneous Model Merging Approach for Complex Knowledge Transfer Jiahe Fan et.al. 2609.39369v1 null
2026-09-30 Offline Guidance, Online Reasoning: Reusing LLM Feedback for Small Language Models Bohan Zhang et.al. 2609.39346v1 null
2026-09-30 Understanding as No-Arbitrage: Bounded Dutch Books as a Definition and Training Objective for Language Models Daniel Dragonevskiy et.al. 2609.39341v1 null
2026-09-30 WorkGenesis: Building the Worlds That Teach Agents to Work Xinyu Zhu et.al. 2609.39325v1 null
2026-09-30 Faithful Dual-constrained Erasure for Robust LLM Safety Alignment Jiaqing Li et.al. 2609.39279v1 null
2026-09-30 Fyan: A Human--AI Harness with Semantic Auditing for Document-Level Formalization Wei Zhao et.al. 2609.39228v1 null
2026-09-30 DAGent: Evaluate-then-Grow Planning for Deep Research Agents Hanwen Liu et.al. 2609.39154v1 null
2026-09-30 Rep2Skill: Representation-Guided Skill Self-Evolution for LLM Agents Kaixing Zhang et.al. 2609.39149v1 null
2026-09-30 MASCRDM: Multi-Agent System for Compliance Risk Detection and Mitigation in Training Process of Large Language Models Yan Zhang et.al. 2609.39107v1 null
2026-09-30 Multi-LLM Collaborative Alignment via Stackelberg Games Christina Hahn et.al. 2609.39076v1 null
2026-09-30 CORE: Conflict-Oriented Reasoning Elimination for Verifiable Language-Model Search Siyu Song et.al. 2609.39069v1 null
2026-09-30 Structure-aware Reinforcement Learning for Protein Directed Evolution Zikun Nie et.al. 2609.39048v1 null
2026-09-30 SimEX: Simulation-Integrated Robotics AutoResearch Jiaheng Hu et.al. 2609.38982v1 null
2026-09-30 DrivingBench: Can Vision-Language Models Drive a Toyota Corolla? Aditya Ramabadran et.al. 2609.38948v1 null
2026-09-30 Prototype-guided Bilateral Alignment Multimodal Federated Learning Tianchi Liao Tianchi_Liao et.al. 2609.38925v1 null
2026-09-30 GraphForge: Training Working Agents with Graph-Anchored Workspace Synthesis Qisheng Su et.al. 2609.38923v1 null
2026-09-30 Unlearning Deceptive Behaviors in LLMs with Contrastive Forget Sets Haoran Tang et.al. 2609.38909v1 null
2026-09-30 K2P: Label-Free Knowledge to Prompt Distillation Yingchuan Zhang et.al. 2609.38898v1 null
2026-09-30 Unmerge: Efficient Machine Unlearning via Task Arithmetic Haoran Tang et.al. 2609.38895v1 null
2026-09-30 Right Answers, Costly Models: The Efficiency Gap in LLM-based Optimization Modeling Zhong Li et.al. 2609.38884v1 null
2026-09-30 Does Learning Protein Folding Generalize to Broader Reasoning? Yong Liu et.al. 2609.38879v1 null
2026-09-30 StateTree: Enhancing Long-Term Dialogue Reasoning via Reinforcement Learning Naen Xu et.al. 2609.38809v1 null
2026-09-30 GraphCert: Bootstrap Agentic Graph Reasoning with Certified Evidence Rubrics Weiqi Jiang et.al. 2609.38798v1 null
2026-09-30 Evaluating Persistent Calibration under Evolving Model Knowledge Victor Wang et.al. 2609.38797v1 null
2026-09-30 Self-Evolving Algorithm-Design Agents: Escaping In-Context Evolutionary Stagnation via Population-Curated Policy Optimization Chen Lu et.al. 2609.38757v1 null
2026-09-30 Learning to Route in Visual Space via Multi-Step Embedding Retrieval Tianyu Chen et.al. 2609.38743v1 null
2026-09-30 Concept-Grounded Attention: A Controlled Evaluation of Graph-Injected Attention, Temporal Versioning, and Epistemic Status Sachin Dev Duggal et.al. 2609.38684v1 null
2026-09-29 Bandits with Multiple Optimal Arms: Minimax Regret and Non-Adaptivity Kaixuan Ji et.al. 2609.38659v2 null
2026-09-29 Prompt2Skill: Unsupervised Skill Optimization From Natural Language Instructions Bo Ni et.al. 2609.38593v1 null

Abstracts

KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards

2610.02206v1 by Pengfei Li, Naufal Suryanto, Sicheng Zhang, Muzammal Naseer

LLMs are increasingly applied to cybersecurity workflows, where they are expected to translate analysts' intent into tool invocations. However, existing evaluations focus on knowledge-based assessments or end-to-end agentic tasks, and do not directly measure LLMs' ability to generate executable commands for real-world cybersecurity tools. This gap is critical because cybersecurity operations rely on strict command-line interfaces (CLIs), where minor syntax errors, incorrect flag--value bindings, or argument misordering can invalidate execution. We introduce KaliBench, a fine-grained benchmark and dataset for natural-language--to--CLI translation on Kali Linux, comprising 8,504 query--command pairs spanning 1,642 tools across 23 capability dimensions and 5 security phases. KaliBench is constructed via a manuscript-grounded pipeline with deterministic canonicalization and alias-aware evaluation, enabling precise and reproducible assessment of tool selection and argument construction. To ensure both semantic correctness and practical executability, we develop a multi-stage verification pipeline that combines LLM-based validation, sandboxed terminal execution, and human-in-the-loop refinement. Building on these fine-grained, deterministic signals, KaliBench further enables runtime-free verifiable rewards for training. Across three evaluation modes and 24 configurations of general-purpose and security-focused open-weight models, no open-weight model exceeds 42% exact-command accuracy in the unrestricted setting, highlighting the difficulty of accurate CLI-based cybersecurity tool use without explicit tool hints. We further show that supervised fine-tuning and reinforcement learning with verifiable rewards derived from KaliBench significantly improve an 8B model and achieve performance comparable to a 685B MoE model.

摘要:LLMs 正在越來越多地應用於網絡安全工作流程中,它們被期望將分析師的意圖轉化為工具調用。然而,現有的評估主要集中在基於知識的評估或端到端的代理任務上,並未直接測量 LLMs 生成可執行命令以供現實世界網絡安全工具使用的能力。這一差距至關重要,因為網絡安全操作依賴於嚴格的命令行界面 (CLIs),其中微小的語法錯誤、不正確的標誌--值綁定或參數錯序都可能使執行無效。我們介紹了 KaliBench,這是一個針對 Kali Linux 的自然語言到 CLI 翻譯的細粒度基準和數據集,包含 8,504 個查詢--命令對,涵蓋 1,642 種工具,跨越 23 個能力維度和 5 個安全階段。KaliBench 是通過一個基於手稿的管道構建的,具有確定性的標準化和別名感知評估,能夠精確且可重複地評估工具選擇和參數構建。為了確保語義正確性和實際可執行性,我們開發了一個多階段驗證管道,結合了基於 LLM 的驗證、沙盒終端執行和人類參與的精煉。基於這些細粒度的確定性信號,KaliBench 進一步使得無運行時的可驗證獎勵成為訓練的可能。在三種評估模式和 24 種通用及安全專注的開放權重模型配置中,沒有任何開放權重模型在不受限制的設置中超過 42% 的精確命令準確率,突顯了在沒有明確工具提示的情況下準確使用基於 CLI 的網絡安全工具的困難。我們進一步顯示,從 KaliBench 派生的可驗證獎勵的監督微調和強化學習顯著改善了一個 8B 模型,並達到了與 685B MoE 模型相當的性能。

Higher-Order Molecular Grammars for Generative and Foundation Models in Chemistry

2610.02186v1 by Yiming Huang, Yujie Zeng, Vijay Prakash Dwivedi, Simone Foti, Jianmin Wang, Jure Leskovec, Tolga Birdal

Molecular learning models are strongly shaped by their underlying representations. Yet standard sequential and graph formalisms struggle to explicitly encode higher-order topology, such as ring systems and recurring motifs. Existing higher-order representations can capture these structures directly, but they are often computationally demanding and difficult to decode into valid molecules. Here, we introduce Higher-order Grammar Representation (HGR), a principled, topology-aware framework that lifts molecules to combinatorial complexes and parses each complex into a compact sequence of production rules under a context-free higher-order grammar. By serialising higher-order topology into rule sequences, HGR makes these structures directly compatible with standard sequence models, avoiding the computational overhead of explicit higher-order encodings while preserving topological expressiveness. To reduce benchmark bias towards simple ring systems, we construct RingDiv, a ring-enriched benchmark containing 1.18 million molecules, including the curated RingDiv300k subset, and introduce the ring diversity index (RDI) to quantify ring-system coverage. In molecular generation, HGR-based models uniquely combine 100% validity by construction with leading distributional alignment, ranking first in FCD on all five generation benchmarks. In representation learning, HGR-FM achieves the highest mean AUC across seven MoleculeNet benchmarks under both transfer protocols, improving on the strongest baseline by 8.3 and 3.3 AUC points under probing and full fine-tuning, respectively. Collectively, these results establish HGR as an efficient higher-order representation for molecular generation and transferable representation learning.

摘要:分子學習模型受到其基礎表示的強烈影響。然而,標準的序列和圖形形式在明確編碼高階拓撲方面(如環系統和重複圖案)面臨挑戰。現有的高階表示可以直接捕捉這些結構,但它們通常計算需求高且難以解碼為有效的分子。在此,我們介紹高階語法表示(HGR),這是一個原則性、關注拓撲的框架,將分子提升為組合複合體,並將每個複合體解析為在上下文無關的高階語法下的緊湊生成規則序列。通過將高階拓撲序列化為規則序列,HGR使這些結構與標準序列模型直接兼容,避免了明確高階編碼的計算開銷,同時保留了拓撲表達能力。為了減少對簡單環系統的基準偏見,我們構建了RingDiv,這是一個包含118萬個分子的環增強基準,包括精心策劃的RingDiv300k子集,並引入環多樣性指數(RDI)來量化環系統的覆蓋範圍。在分子生成方面,基於HGR的模型獨特地結合了100%的有效性(由構造決定)與領先的分佈對齊,在所有五個生成基準中FCD排名第一。在表示學習方面,HGR-FM在七個MoleculeNet基準中,在兩種轉移協議下實現了最高的平均AUC,相較於最強基線分別提高了8.3和3.3 AUC點(在探測和完全微調下)。綜合這些結果,HGR確立了作為分子生成和可轉移表示學習的高效高階表示。

From Knowledge Access to Source Learning: Developing Source-Specific Competence

2610.02150v1 by Lucheng Fu, Kejing Xia, Yiyang Wang, Yiqiao Jin, Jinjin He, Xiyuan Yang, Haoxin Liu, Ye Yu, Haibo Jin, Yijia Xiao, Wenke Lee, B. Aditya Prakash, Haohan Wang

Large language model (LLM) agents increasingly rely on persistent external sources to solve sequences of knowledge-intensive tasks. Existing methods improve how source content is accessed and organized, while agent-memory systems preserve reusable knowledge from prior interactions, but repeated use of the same source is still largely treated as repeated access rather than an opportunity to progressively improve understanding of that source. We study source learning: developing reusable source-specific competence over a persistent authoritative source. We represent this competence with a persistent source model that captures reusable understanding of the source, including how its knowledge is structured, interpreted, and applied. To construct and progressively refine such models, we propose SourceLearn, which combines two complementary learning mechanisms. Self-Directed Source Learning identifies what remains incompletely understood and adaptively revisits the source, while Task-Guided Source Learning uses downstream experience to reveal local representational gaps and recurring needs in how source knowledge should be organized. In both cases, learning signals determine what should be reconsidered, while persistent updates are reconstructed from the authoritative source. Across five benchmarks and three LLM backends, SourceLearn achieves the best performance in 13 of 15 settings, with gains of up to 22.6 points over Hybrid RAG and substantial overall improvements over static source representations and experience-based memory baselines.

摘要:大型語言模型(LLM)代理越來越依賴持久的外部來源來解決一系列知識密集型任務。現有方法改善了如何訪問和組織來源內容,而代理記憶系統則保留了來自先前互動的可重用知識,但對同一來源的重複使用仍然主要被視為重複訪問,而不是逐步改善對該來源理解的機會。我們研究來源學習:在持久的權威來源上發展可重用的來源特定能力。我們用一個持久的來源模型來表示這種能力,該模型捕捉了對來源的可重用理解,包括其知識的結構、解釋和應用方式。為了構建和逐步完善這樣的模型,我們提出了SourceLearn,該模型結合了兩種互補的學習機制。自我導向來源學習識別尚未完全理解的內容並適應性地重新訪問來源,而任務引導來源學習則利用下游經驗揭示如何組織來源知識的局部表徵差距和重複需求。在這兩種情況下,學習信號決定了應該重新考慮的內容,而持久更新則是從權威來源重建的。在五個基準和三個LLM後端中,SourceLearn在15個設置中的13個中實現了最佳性能,與Hybrid RAG相比,增益高達22.6點,並且在靜態來源表示和基於經驗的記憶基準上有顯著的整體改進。

Mem++: Non-Destructive Memory for Long-Term Organizational LLM Agents

2610.02002v1 by Ahmad Yehia, Aly O. Abdelkareem, Islam Ahmed, Hesham Omran, Khaled Alashmouny, Christian Claudel, Abduallah Mohamed

Large Language Model (LLM) agents now take part in organizational work, where many authors record decisions across documents over months. Because a revised decision arrives as a new document rather than an edit, answering a question requires knowing which version held at a given time. However, most memory systems compress the record at write time. By distilling each document into facts, notes or graph edges, these methods fix what can be answered before any question is asked. To address this, we propose Mem++, a non-destructive memory framework shifting from write-time distillation to read-time selection. Mem++ stores every document whole with its date and author, and it calls no generative model at write time. At read time, it retrieves only documents dated up to the time a question asks about and fuses lexical and semantic rankings. Unlike systems that overwrite older versions, Mem++ keeps them and leaves the choice to the answering model. Evaluations on the organizational benchmark OrgMemBench demonstrate that Mem++ surpasses the strongest memory system baseline by 8.0 to 13.1 points across two answering models. With gpt-4.1-mini, it also achieves the best overall score, 2.6 points above RAG. In addition, Mem++ achieves the best average LLM-judge score on LoCoMo and ranks second on LongMemEval-S, behind only its entity-graph variant. Code for benchmark evaluation is available at https://github.com/AIDAChip-Inc/mem-plus-plus.

摘要:大型語言模型(LLM)代理現在參與組織工作,許多作者在數月內記錄決策於文件中。因為修訂的決策以新文件的形式出現,而不是編輯,因此回答問題需要知道在特定時間持有的是哪個版本。然而,大多數記憶系統在寫入時會壓縮記錄。通過將每個文件提煉成事實、筆記或圖邊,這些方法在任何問題被提出之前固定了可以回答的內容。為了解決這個問題,我們提出了Mem++,這是一個非破壞性的記憶框架,從寫入時的提煉轉向讀取時的選擇。Mem++ 將每個文件完整地存儲,並附上日期和作者,並且在寫入時不調用任何生成模型。在讀取時,它僅檢索在問題詢問時的日期之前的文件,並融合詞彙和語義排名。與覆蓋舊版本的系統不同,Mem++ 保留它們,並將選擇權留給回答模型。在組織基準測試 OrgMemBench 上的評估顯示,Mem++ 在兩個回答模型中超越了最強記憶系統基線 8.0 到 13.1 分。使用 gpt-4.1-mini 時,它還獲得了最佳整體分數,比 RAG 高出 2.6 分。此外,Mem++ 在 LoCoMo 上獲得了最佳平均 LLM-judge 分數,並在 LongMemEval-S 中排名第二,僅次於其實體圖變體。基準評估的代碼可在 https://github.com/AIDAChip-Inc/mem-plus-plus 獲得。

Can AI Oversight Be Zero Knowledge?

2610.01995v1 by Alessandro Chiesa, Ziyi Guan, Burcu Yildiz

AI systems increasingly produce outputs from confidential data, such as a fitness-for-duty assessment from medical records or the predicted properties of a drug candidate from its secret structure. It is important to verify that such outputs are correct without revealing the underlying data. A recent line of work studies verification of AI outputs via interactive proofs and debate for oracle-aided computation, where correctness may depend on an oracle such as human judgment, a physical experiment, or the web. These works focus on verification by a verifier that runs much faster than the computation. However, such efficient verification is impossible for general oracle-aided computation, and these works therefore rely on additional assumptions. We focus instead on privacy: allowing the verifier to run in time polynomial in the computation, we ask whether interactive arguments for oracle-aided computation can be zero knowledge, so that the verifier learns nothing about the confidential data beyond the correctness of the output. We prove that, in general, they cannot. In the random oracle model, there are no zero-knowledge proofs for all oracle-aided computations, even if both the prover and the verifier are allowed to run much longer than the computation itself. The impossibility extends to debate, a canonical model for scalable oversight. On the positive side, we show that if the oracle attaches a cryptographic signature to each of its answers, then every oracle-aided computation can be verified in zero knowledge with an efficient prover and verifier, assuming only collision-resistant hash functions. Beyond privacy, this also gives an alternative approach to scalable oversight that relies neither on an honest opponent, as in debate, nor on the robustness of the computation, as in prior single-prover protocols.

摘要:AI 系統越來越多地從機密數據中產生輸出,例如從醫療記錄中進行的適任性評估或從其秘密結構中預測的藥物候選物的性質。驗證這些輸出是否正確而不透露底層數據是很重要的。最近的一系列研究通過互動證明和辯論研究 AI 輸出的驗證,用於有 oracle 輔助的計算,其中正確性可能依賴於 oracle,例如人類判斷、物理實驗或網絡。這些研究專注於由運行速度遠快於計算的驗證者進行的驗證。然而,對於一般的有 oracle 輔助計算,這樣的高效驗證是不可能的,因此這些研究依賴於額外的假設。我們則專注於隱私:允許驗證者在計算的多項式時間內運行,我們詢問有 oracle 輔助計算的互動論證是否可以是零知識的,以便驗證者不會學到關於機密數據的任何信息,除了輸出的正確性。我們證明,通常情況下,它們是不可能的。在隨機 oracle 模型中,對於所有有 oracle 輔助的計算,沒有零知識證明,即使證明者和驗證者都被允許運行的時間遠超過計算本身。這種不可能性擴展到辯論,這是一個可擴展監督的典型模型。從積極的一面來看,我們展示了如果 oracle 為其每個答案附加加密簽名,那麼每個有 oracle 輔助的計算都可以在零知識中進行驗證,並且有高效的證明者和驗證者,只假設碰撞抗性哈希函數。除了隱私之外,這還提供了一種可擴展監督的替代方法,既不依賴於誠實的對手(如辯論中),也不依賴於計算的穩健性(如以前的單證明者協議中)。

Latent JEPA: Abstract Future Prediction for Latent Reasoning in Chemistry

2610.01947v1 by Xinjian Zhao, Yaoyao Xu, Xuemin Chen, Xiaozhuang Song, Tianshu Yu

Large language models offer a promising foundation for chemical reasoning, bringing together chemical knowledge and multistep problem solving. Chemical intuition can provide an initial sense of plausible outcomes before the details of a solution are fully worked out. Inspired by how such expectations complement explicit analysis, we study how continuous latent thoughts can be trained to anticipate informative aspects of future solutions without verbalizing every intermediate step. We introduce Latent JEPA, a framework that combines autoregressive learning with joint-embedding prediction of one or more future views. For chemical reasoning, we develop textual and molecular prediction objectives that connect latent thoughts to both subsequent reasoning and molecular outcomes. Experiments on ChemCoTBench show gains in molecular optimization and on several editing and reaction metrics. Representation analyses show that future prediction makes latent thoughts more informative about molecular outcomes and strengthens their correspondence with chemical structure. These findings support abstract future prediction as a learning principle for connecting continuous latent reasoning with scientific outcomes.

摘要:大型語言模型為化學推理提供了一個有前景的基礎,將化學知識和多步驟問題解決結合在一起。化學直覺能在解決方案的細節完全展開之前,提供一種合理結果的初步感知。受到這種期望如何補充明確分析的啟發,我們研究如何訓練連續潛在思維,以預測未來解決方案的資訊性方面,而不需要逐步口頭表達每一個中間步驟。我們介紹了潛在JEPA,一個將自回歸學習與一個或多個未來視圖的聯合嵌入預測相結合的框架。針對化學推理,我們開發了文本和分子預測目標,將潛在思維與後續推理和分子結果連接起來。在ChemCoTBench上的實驗顯示,在分子優化以及幾個編輯和反應指標上都有提升。表徵分析顯示,未來預測使潛在思維對分子結果的資訊性更強,並加強了它們與化學結構的對應性。這些發現支持將抽象的未來預測作為一種學習原則,以連接連續的潛在推理與科學結果。

Mapping the RAG Landscape: A Four Axis Taxonomy of Efficiency, Defense, Interactivity, and Reasoning

2610.01936v1 by Meghana Sunil, Shravya V, Shravan Venkatraman, Joe Dhanith PR

Large Language Models (LLMs) have demonstrated remarkable fluency across many tasks but remain limited by their static, parameter bound knowledge and their susceptibility to hallucinating information. Retrieval Augmented Generation (RAG) addresses these issues by incorporating external retrieval into the generation process, grounding model outputs in verifiable and up to date sources. While prior surveys primarily focus on core RAG architectures and standard pipelines, recent research explores broader challenges and capabilities that extend beyond these foundational designs. This survey provides a consolidated and structured examination of contemporary RAG developments, organizing the field into a four axis taxonomy: improving retrieval efficiency, strengthening robustness and security, supporting user driven and interactive workflows, and enabling multi step or complex reasoning. We formalize key components of the RAG framework and review methods spanning dense and sparse retrieval, fusion strategies, embedding optimizations, and reinforcement learning based retrieval policies, highlighting how these advances influence practical deployment and system design. We also synthesize evaluation practices, domain specific applications, and architectural variants such as Naive, Advanced, and Modular RAG. Finally, we outline persistent challenges related to retrieval quality, reliability, domain adaptation, scalability, and explainability, and identify opportunities for building RAG systems that are more reliable, adaptable, and transparent.

摘要:大型語言模型(LLMs)在許多任務中展現了卓越的流暢性,但仍然受到靜態的、參數限制的知識以及對虛假信息的易感性的限制。檢索增強生成(RAG)通過將外部檢索納入生成過程來解決這些問題,使模型輸出基於可驗證且最新的來源。雖然之前的調查主要集中在核心RAG架構和標準流程上,但最近的研究探討了超越這些基礎設計的更廣泛挑戰和能力。本調查提供了一個當代RAG發展的綜合和結構化檢視,將該領域組織為四個軸向的分類法:提高檢索效率、加強穩健性和安全性、支持用戶驅動和互動工作流程,以及實現多步驟或複雜推理。我們正式化了RAG框架的關鍵組件,並回顧了涵蓋密集和稀疏檢索、融合策略、嵌入優化和強化學習基於檢索政策的方法,強調這些進展如何影響實際部署和系統設計。我們還綜合了評估實踐、特定領域的應用以及如Naive、Advanced和Modular RAG等架構變體。最後,我們概述了與檢索質量、可靠性、領域適應性、可擴展性和可解釋性相關的持續挑戰,並確定了構建更可靠、可適應和透明的RAG系統的機會。

From Network Intrusion Detection to Blockchain-Backed Endpoint Detection and Response: Mapping the Landscape of Decentralized Detection-and-Response Architectures

2610.01872v1 by Yahya Shahsavari, Sara Rouhani, Kaiwen Zhang

While the literature on blockchain-assisted intrusion detection and prevention systems (IDS/IPS) for Internet of Things (IoT) and Industrial Internet of Things (IIoT) networks is mature, existing systematic reviews suffer from two critical limitations: they overlook the structural shift toward modern Endpoint Detection and Response (EDR) and Extended Detection and Response (XDR) architectures, and they conflate blockchain's distinct functional roles into a single monolithic category. This Systematization of Knowledge (SoK) addresses these gaps by proposing a three-axis taxonomy that classifies proposals by detection-system class (NIDS, HIDS, EDR/XDR), blockchain functional role, and response-automation maturity. Synthesizing research published in high-impact venues between 2019 and 2026, we provide a rigorous gap analysis exposing why a genuine per-endpoint blockchain-anchored response loop remains nearly nonexistent due to latency, deployment, and community mismatches. Furthermore, we evaluate structural, cross-cutting challenges persisting across the literature, including consensus latency on constrained devices, post-quantum cryptographic vulnerability, smart-contract attack surfaces, and the adversarial vulnerability of evolving LLM-based detection engines. Finally, we outline a comprehensive research agenda centered on hybrid on-chain/off-chain orchestration to bridge the gap between decentralized trust and rapid response automation.

摘要:雖然有關區塊鏈輔助的入侵檢測和預防系統(IDS/IPS)在物聯網(IoT)和工業物聯網(IIoT)網絡中的文獻已相當成熟,但現有的系統性評估存在兩個關鍵限制:它們忽視了向現代端點檢測與響應(EDR)和擴展檢測與響應(XDR)架構的結構性轉變,並且將區塊鏈的不同功能角色混淆為一個單一的整體類別。這項知識系統化(SoK)通過提出一個三軸分類法來解決這些空白,該分類法根據檢測系統類別(NIDS、HIDS、EDR/XDR)、區塊鏈功能角色和響應自動化成熟度對提案進行分類。綜合2019年至2026年間在高影響力期刊上發表的研究,我們提供了一個嚴謹的差距分析,揭示了為何真正的每個端點區塊鏈錨定響應循環幾乎不存在,原因在於延遲、部署和社群不匹配。此外,我們評估了文獻中持續存在的結構性、跨領域挑戰,包括在受限設備上的共識延遲、後量子密碼學脆弱性、智能合約攻擊面以及不斷演變的基於LLM的檢測引擎的對抗性脆弱性。最後,我們概述了一個以混合鏈上/鏈下協同為中心的全面研究議程,以彌合去中心化信任與快速響應自動化之間的鴻溝。

Walking the Embedding Space: Datastore Extraction from Multimodal RAG

2610.01871v1 by Maria Carmen Jica, Ali Satvaty, Suzan Verberne, Fatih Turkmen

Multimodal Retrieval-Augmented Generation (MRAG) has emerged as a reliable and cost-effective technique of grounding the generative capabilities of Multimodal Large Language Models (MLLMs) into relevant, up-to-date, external knowledge. Despite presenting several benefits, such as reducing hallucinatory behavior, they also introduce new attack surfaces, including leakage of private information and vulnerabilities against data extraction attacks. In this paper, we introduce $\immrag$, an adaptive and automatic data extraction attack procedure operating in a black box setting against \emph{image-returning} MRAG, a configuration in which the retrieved visual artifact is itself the response. Each query blends an attacker-held shadow image with an image already recovered from the system, and relevance-weighted resampling steers subsequent queries towards regions of the embedding space that still yield novel retrievals. Unlike current extraction attacks that aim to persuade the model towards data leakage by placing a malicious query as a textual prompt, $\immrag$ embeds the malicious instructions inside a user-given input image. We evaluate $\immrag$ on three plausible and distinct real-world scenarios: medical assistant, document-focused helper and general purpose tool. The experiments involve the study of the effectiveness of the attack on multiple CLIP-family retrievers, as well as the impact of various generators. A single 2500-query run reconstructs up to 611 distinct radiology images, 566 document scans and 416 general-purpose images under local-feature correspondence, and reaches up to $5.6\times$ as many distinct datastore items as a non-adaptive baseline. Our results show the urgent need for safeguards specifically designed for multimodal data.

摘要:多模態檢索增強生成(MRAG)已成為將多模態大型語言模型(MLLMs)的生成能力與相關的、最新的外部知識相結合的一種可靠且具成本效益的技術。儘管它提供了幾個好處,例如減少幻覺行為,但它們也引入了新的攻擊面,包括私密信息洩漏和對數據提取攻擊的脆弱性。 在本文中,我們介紹了 $\immrag$,這是一種適應性和自動化的數據提取攻擊程序,針對 \emph{圖像返回} MRAG 在黑箱環境中運作,這是一種檢索的視覺工件本身就是回應的配置。每個查詢將攻擊者持有的影像與系統中已恢復的影像混合,並且相關性加權重採樣引導後續查詢朝向仍能產生新穎檢索的嵌入空間區域。與目前旨在通過將惡意查詢作為文本提示來說服模型進行數據洩漏的提取攻擊不同,$\immrag$ 將惡意指令嵌入用戶提供的輸入影像中。我們在三個合理且不同的現實場景中評估了 $\immrag$:醫療助手、文件專注助手和通用工具。實驗涉及對多個 CLIP 家族檢索器的攻擊有效性以及各種生成器的影響進行研究。一次 2500 次查詢的運行重建了多達 611 幅不同的放射學影像、566 幅文件掃描和 416 幅通用影像,根據局部特徵對應,並達到高達 $5.6\times$ 的不同數據庫項目數量,相較於非適應性基準。我們的結果顯示出對專門為多模態數據設計的安全措施的迫切需求。

Detecting Inconsistencies in Model Specifications with LLM-as-Verifier Reasoning

2610.01847v1 by Zichen Xie, Mrigank Pawagi, Lize Shao, Yang Hu, Wenxi Wang

Model specifications define how large language models (LLMs) should behave, guiding alignment training, inference-time behavior, and evaluation. Yet these specifications may themselves contain defects: two individually reasonable principles may prescribe incompatible behavior when applied to the same situation, leaving no response that satisfies both. Detecting such inconsistencies is challenging. Formalizing natural-language specifications risks losing subtle distinctions, while behavior-based testing cannot reliably distinguish specification defects from differences in model behavior. We introduce VeriSpec, the first approach to directly detect inconsistencies in model specifications by auditing the specification text itself. Our key insight is to preserve the specification in natural language while using an LLM as a verifier. VeriSpec extracts structured, context-aware rules, constructs a topic-guided graph to cluster behaviorally related rules at the same authority level, and applies LLM-as-verifier reasoning to detect inconsistencies. Applying VeriSpec to the OpenAI Model Spec, we extract 405 rules and manually validate five inconsistencies, all reported to its developers, who responded positively and have initiated internal discussions. Compared with five baselines, VeriSpec identifies the most validated inconsistencies, achieves the highest precision (38.5%), and incurs the lowest cost per validated inconsistency ($11.12). These results establish direct specification auditing as a practical complement to behavioral alignment evaluation, catching defects at the source before they shape any model. The code is available at https://github.com/HIPREL-Group/VeriSpec.

摘要:模型規範定義了大型語言模型(LLMs)應該如何運作,指導對齊訓練、推理時的行為和評估。然而,這些規範本身可能包含缺陷:兩個各自合理的原則在應用於相同情境時可能會規定不相容的行為,導致沒有任何回應能同時滿足兩者。檢測這種不一致性是具有挑戰性的。將自然語言規範形式化可能會失去微妙的區別,而基於行為的測試則無法可靠地區分規範缺陷與模型行為的差異。我們引入了VeriSpec,這是第一種通過審核規範文本本身直接檢測模型規範中不一致性的方法。我們的關鍵見解是保留自然語言中的規範,同時使用LLM作為驗證者。VeriSpec提取結構化的、上下文感知的規則,構建主題引導的圖以聚類同一權威級別下行為相關的規則,並應用LLM作為驗證者的推理來檢測不一致性。將VeriSpec應用於OpenAI模型規範,我們提取了405條規則並手動驗證了五個不一致性,所有這些都已報告給其開發者,開發者對此做出了積極回應並已啟動內部討論。與五個基準相比,VeriSpec識別了最多的經過驗證的不一致性,達到了最高的精確度(38.5%),並且每個經過驗證的不一致性的成本最低($11.12)。這些結果確立了直接規範審核作為行為對齊評估的實用補充,在缺陷影響任何模型之前,及時捕捉到缺陷。代碼可在 https://github.com/HIPREL-Group/VeriSpec 獲得。

Code Owns the Simulation, Jev Owns the Evaluation

2610.01834v1 by Yaodong Yang, Hongyao Tang, Yi Ma, Xingyu Fan, Weixun Wang, Jinpeng Li, Tianpei Yang

Judgment models such as \jev{} return, in a single call and without reasoning text, a probability for each described option. This makes them attractive as an agent's action-selection layer, but it is unclear which decisions they can be trusted with. We test \jev{} on reflection tests, one-shot matrix games, the text game ALFWorld and robot control, and find a sharp boundary. \jev{} succeeds when the right option can be judged from what the input describes, which we call \emph{evaluation}. Specifically, it solves 99\% of the counterintuitive Cognitive Reflection Test questions. However, it fails when the right option depends on \emph{simulation} (i.e., predicting something not in the input), such as the opponent's action or the subgoal that must come first. In games, \jev{} plays suboptimally as if its rational opponent acted at random, because the opponent's action is not given. In ALFWorld, \jev{} favors commands that mention an object or place named in the task description. For example, given the task ``put a clean knife in the drawer'', \jev{} carries an unwashed knife straight to the drawer instead of first washing it at the sink. Surprisingly, many of these failures are not due to a lack of knowledge. Asked separately what the opponent will do, \jev{} usually answers correctly, and it responds well given the opponent's action. It fails when one call must both perform the simulation and evaluate based on it. This suggests letting code make the prediction or simulation. When code supplies it, such as a lookahead in ALFWorld and physics simulation in robot control, \jev{} becomes an expert controller through its general evaluation ability.

摘要:判斷模型如 \jev{} 在單次呼叫中返回每個描述選項的概率,且不需要推理文本。這使得它們作為代理的行動選擇層變得具有吸引力,但尚不清楚它們可以信任哪些決策。我們在反思測試、一回合矩陣遊戲、文本遊戲 ALFWorld 和機器人控制上測試 \jev{},並發現了一個明確的邊界。當正確選項可以從輸入描述中判斷時,我們稱之為 \emph{評估},\jev{} 成功地解決了 99\% 的反直覺認知反思測試問題。然而,當正確選項依賴於 \emph{模擬}(即預測輸入中不存在的事物)時,它則失敗,例如對手的行動或必須先完成的子目標。在遊戲中,\jev{} 表現得次優,彷彿其理性的對手隨機行動,因為對手的行動並未給出。在 ALFWorld 中,\jev{} 偏好提到任務描述中物體或地點的命令。例如,給定任務「將乾淨的刀放入抽屜」,\jev{} 直接將未洗的刀帶到抽屜,而不是先在水槽中清洗它。令人驚訝的是,這些失敗中的許多並非因為缺乏知識。單獨詢問對手會做什麼時,\jev{} 通常能正確回答,並且在給定對手的行動時反應良好。當一次呼叫必須同時執行模擬並基於此進行評估時,它則失敗。這表明應讓代碼進行預測或模擬。當代碼提供這些信息時,例如在 ALFWorld 中的前瞻和在機器人控制中的物理模擬,\jev{} 通過其一般評估能力成為專家控制器。

The Asymptotics of Language Model Alignment with Memory

2610.01828v1 by Haricharan Balasundaram, V. Arvind Rameshwar

Language model (LM) alignment broadly aims to perturb a given LM $Q$ into an aligned LM $q$ such that i) the outputs produced by $q$ and $Q$ are 'close' in probability, ii) $q$ has a higher expected reward than $Q$. Two common techniques for LM alignment are: KL-constrained RL, which requires knowledge of the LM distribution and is computationally expensive, and the best-of-$n$ algorithm, which requires only sampling from the LM. The work of Yang et al. established asymptotic closeness between the distributions produced by the two alignment methods for an $m$--length i.i.d. token sequence output by the LM, in the limit as $m$ increases to infinity. However, the i.i.d. assumption is not representative of practical LMs, whose output sequences often have memory. In this paper, we extend the asymptotic closeness result to the case when the $m$--length token sequence outputted by the LM is Markovian. Further, for finite-length output sequences -- particularly, when $m=1$ -- we provide a complete characterization of LM distributions and reward functions for which the KL-divergence between the distributions produced by the two alignment methods is zero -- a question first posed in Yang et al.

摘要:語言模型(LM)對齊的廣泛目標是將給定的 LM $Q$ 轉變為一個對齊的 LM $q$,使得 i) $q$ 和 $Q$ 所產生的輸出在概率上是「接近」的,ii) $q$ 的期望獎勵高於 $Q$。兩種常見的 LM 對齊技術是:KL 約束強化學習,這需要對 LM 分佈的了解並且計算上昂貴,以及最佳的 $n$ 算法,這僅需要從 LM 中進行取樣。Yang 等人的研究確立了在 $m$ 長度的獨立同分佈(i.i.d.)標記序列的情況下,兩種對齊方法所產生的分佈之間的漸近接近性,當 $m$ 增加到無限大時。然而,i.i.d. 假設並不代表實際的 LM,因為它們的輸出序列通常具有記憶性。在本文中,我們將漸近接近性結果擴展到 LM 輸出的 $m$ 長度標記序列為馬爾可夫過程的情況。此外,對於有限長度的輸出序列——特別是當 $m=1$ 時——我們提供了 LM 分佈和獎勵函數的完整特徵描述,對於這些情況,兩種對齊方法所產生的分佈之間的 KL 散度為零——這是一個最初由 Yang 等人提出的問題。

A Matryoshka Hierarchical RAG for Efficient Multi-Hop Question Answering

2610.01767v1 by Gianluca Bonifazi, Christopher Buratti, Michele Marchetti, Federica Parlapiano, Giulia Quaglieri, Davide Traini, Domenico Ursino, Luca Virgili

Retrieval-Augmented Generation (RAG) systems for multi-hop Question Answering (QA) must balance retrieval quality with computational cost. This cost is incurred during indexing time, through the use of expensive Knowledge Graphs (KGs) or Large Language Models (LLMs) to generate summaries, or during querying, through iterative LLM-driven retrieval. To reduce it while maintaining retrieval quality, we present MatRAG, a hierarchical framework that combines RAG systems with Matryoshka Representation Learning (MRL). MatRAG addresses both kinds of cost by aligning the semantic hierarchy of a clustering structure with the nested structure of MRL. Specifically, it organizes the corpus of documents into a Directed Acyclic Graph (DAG) of clusters with progressively coarser granularity. Each level is indexed by a lower Matryoshka dimension. MatRAG pairs an iterative, top-down traversal of the DAG with an entity-driven mechanism that controls the hop budget and re-ranks candidates. We evaluated MatRAG on three standard multi-hop QA benchmarks against seven representative baselines. MatRAG outperforms its strongest competitors in terms of retrieval quality; furthermore, it reduces indexing costs by avoiding KG construction and LLM-based summarization, and lowers query-time costs through dimension-aware similarity.

摘要:檢索增強生成(RAG)系統在多跳問題回答(QA)中必須平衡檢索質量與計算成本。這個成本在索引時產生,通過使用昂貴的知識圖譜(KG)或大型語言模型(LLM)來生成摘要,或在查詢時,通過迭代的LLM驅動檢索。為了在保持檢索質量的同時降低成本,我們提出了MatRAG,一個將RAG系統與馬特里奧什卡表示學習(MRL)相結合的分層框架。MatRAG通過將聚類結構的語義層次與MRL的嵌套結構對齊,解決了這兩種成本。具體而言,它將文檔語料庫組織成一個具有逐漸粗糙粒度的有向無環圖(DAG)聚類。每個層級由較低的馬特里奧什卡維度進行索引。MatRAG將DAG的迭代自上而下遍歷與一種驅動實體的機制相結合,該機制控制跳躍預算並重新排名候選者。我們在三個標準的多跳QA基準上評估了MatRAG,並與七個代表性的基準進行比較。MatRAG在檢索質量方面超越了其最強的競爭對手;此外,它通過避免KG構建和基於LLM的摘要來降低索引成本,並通過維度感知相似性來降低查詢時間成本。

Task-Oriented Rank Adaptation for Continual Learning in Text Classification

2610.01702v1 by Rey Sanchez Lopez, Eduardo Morales Manzanares, Hugo Jair Escalante

Continual learning (CL) in text classification faces two critical challenges: catastrophic forgetting and negative transfer across sequential tasks. Parameter-Efficient Fine-Tuning (PEFT) methods such as LoRA enable efficient adaptation by learning low-rank updates of the model parameters. However, these compact representations are normally trained in isolation, limiting their reuse across related tasks. We introduce Task-Oriented Rank Adaptation (TORA), a geometric routing framework that leverages the low-rank structure of LoRA adapters to decide whether to transfer knowledge from the most compatible expert (Boosting) or isolate the new task (Shielding) based on structural similarity. Evaluated across 15 diverse text classification benchmarks, TORA consistently avoids harmful routing decisions: compatible tasks exceed their isolated performance while reducing training time, and structurally distant tasks are protected from interference with no loss in accuracy. With a single geometric threshold and no reliance on task identities or predefined sequences, TORA provides a simple and effective approach for dynamic adapter routing in sequential text classification systems.

摘要:持續學習(CL)在文本分類中面臨兩個關鍵挑戰:災難性遺忘和在序列任務中的負轉移。參數高效微調(PEFT)方法如 LoRA 通過學習模型參數的低秩更新來實現高效適應。然而,這些緊湊的表示通常是在孤立的情況下訓練的,限制了它們在相關任務中的重用。我們引入了任務導向秩適應(TORA),這是一個幾何路由框架,利用 LoRA 適配器的低秩結構來決定是從最兼容的專家(提升)轉移知識,還是根據結構相似性隔離新任務(保護)。在 15 個不同的文本分類基準上進行評估,TORA 始終避免有害的路由決策:兼容任務的表現超過其孤立的性能,同時減少訓練時間,而結構上相距較遠的任務則受到保護,沒有準確度損失。TORA 以單一的幾何閾值運作,且不依賴於任務身份或預定序列,為序列文本分類系統中的動態適配器路由提供了一種簡單而有效的方法。

Iterative Policy Refinement through Semantic Rollout Analysis

2610.01652v1 by Feiyu Gavin Zhu, Qi Xu, Zhifei Deng, Zhigang Hua, Luke Simon, Jean Oh, Reid Simmons

Structured policies improve efficiency, robustness, and interpretability in imitation learning by introducing task-specific inductive bias, but existing structure generation methods rely either on extensive human input or on static domain knowledge encoded in LLMs, which may be inconsistent with the expert demonstrations. We propose a closed-loop framework that iteratively refines structured policies using LLM-guided analysis of policy rollouts. By logging rollouts as semantically meaningful tabular data and prompting the LLM to generate diagnostic analysis code, our method identifies suboptimalities in the policy structure and iteratively corrects them without requiring human instruction. Experiments on car racing and door opening tasks show that our approach improves imitation learning performance by up to 15% over zero-shot LLM-generated structures and requires 75% less compute to achieve the same reinforcement learning performance. These results demonstrate that tabular rollout analysis provides an effective feedback signal to align LLM-generated policy structures with expert demonstrations, and we can utilize it to generate good policy structures automatically.

摘要:結構化政策透過引入特定任務的歸納偏見來提升模仿學習的效率、穩健性和可解釋性,但現有的結構生成方法要麼依賴大量的人類輸入,要麼依賴於編碼在大型語言模型(LLMs)中的靜態領域知識,這可能與專家的示範不一致。我們提出了一個閉環框架,通過使用LLM引導的政策展開分析來迭代地改進結構化政策。通過將展開記錄為語義上有意義的表格數據,並提示LLM生成診斷分析代碼,我們的方法識別出政策結構中的次優性,並在不需要人類指導的情況下進行迭代修正。在賽車和開門任務上的實驗顯示,我們的方法在模仿學習性能上比零樣本LLM生成的結構提高了多達15%,並且需要75%更少的計算來達到相同的強化學習性能。這些結果表明,表格展開分析提供了一個有效的反饋信號,以使LLM生成的政策結構與專家示範對齊,我們可以利用它自動生成良好的政策結構。

Managing Context and Communication in Distributed Agentic UAV Swarms

2610.01569v1 by Andrea Iannoli, Ivan Zyrianoff, Angelo Trotta, Lorenzo Gigli, Marco Di Felice

Unmanned aerial vehicle (UAV) swarms increasingly rely on language-model agents to provide adaptive mission-level reasoning in uncertain environments. Fully distributed control, in which each UAV hosts an independent Small Language Model (SLM), removes reliance on a centralized coordinator but introduces an information-management problem: long-running interaction histories can degrade the reasoning context, while indiscriminate information dissemination increases communication and inference overhead. We address these challenges with a distributed UAV-agent architecture that enables continuous local SLM control through an event-driven reason-act-observe lifecycle. Runtime knowledge is represented as structured atomic notes and organized into core, local, and peer-specific memory. A deterministic interest-aware gossip engine selectively disseminates these notes according to recipient-specific semantic novelty and recency. We evaluate the architecture using ten UAVs in a simulated search-and-rescue mission. Our approach completes all experimental runs, whereas unrestricted flooding messages completes only 70-85\%, and delegating forwarding decisions to the SLM prevents mission completion in every run. Compared with unrestricted flooding, our approach approximately halves inference-token consumption, reduces transmitted data, and achieves lower survivor-count error.

摘要:無人機(UAV)群體越來越依賴語言模型代理在不確定環境中提供自適應的任務級推理。完全分散的控制中,每個UAV都擁有一個獨立的小型語言模型(SLM),這消除了對集中協調者的依賴,但引入了一個信息管理問題:長期的互動歷史可能會削弱推理上下文,而不加區別的信息傳播則增加了通信和推理的開銷。我們通過一種分散的UAV代理架構來解決這些挑戰,該架構通過事件驅動的推理-行動-觀察生命週期實現持續的本地SLM控制。運行時知識以結構化的原子筆記形式表示,並組織成核心、本地和對等特定的記憶。一個確定性的興趣感知八卦引擎根據接收者特定的語義新穎性和時效性選擇性地傳播這些筆記。我們使用十架UAV在模擬的搜索和救援任務中評估該架構。我們的方法完成了所有實驗運行,而不受限制的洪水消息僅完成了70-85\%,將轉發決策委託給SLM則在每次運行中都阻止了任務的完成。與不受限制的洪水相比,我們的方法大約減半了推理令牌的消耗,減少了傳輸數據,並實現了較低的生還者計數誤差。

2610.01553v1 by Nikolai Zenovkin, Sebastian Björkqvist

Patent search requires processing documents routinely exceeding tens of thousands of tokens. Most neural retrieval approaches operate on truncated inputs, limiting their effectiveness. Graph-based retrieval addresses this by representing each patent as a structured invention graph, but constructing these graphs relies on brittle rule-based parsers. We present the neural parser, which adapts biaffine attention from dependency parsing to predict invention graphs directly from patent text. Our local biaffine attention restricts pairwise scoring to a sliding window, reducing complexity from $O(n^2)$ to $O(n \cdot w)$. Since local and global scoring share the same weights, the model trains on short sequences and deploys on documents exceeding 40,000 tokens without retraining. Distilled from 1 million rule-parsed documents, it surpasses its teacher at 3$\times$ lower inference cost: neural graphs improve citation recall by 0.5% on short queries and 1.1% on full documents in a downstream Graph Transformer retrieval system.

摘要:專利搜尋需要處理的文件通常超過數萬個標記。大多數神經檢索方法在截斷的輸入上運作,限制了它們的有效性。基於圖的檢索通過將每個專利表示為結構化的發明圖來解決這個問題,但構建這些圖依賴於脆弱的基於規則的解析器。我們提出了神經解析器,它將雙仿射注意力從依賴解析中適應,以直接從專利文本預測發明圖。我們的局部雙仿射注意力將成對評分限制在滑動窗口內,將複雜度從 $O(n^2)$ 降低到 $O(n \cdot w)$。由於局部和全局評分共享相同的權重,該模型在短序列上進行訓練,並在不重新訓練的情況下部署在超過 40,000 個標記的文件上。從 100 萬個規則解析的文件中提煉出來,它在推理成本上超越了其教師,降低了 3$\times$:神經圖在下游圖Transformer檢索系統中,對於短查詢提高了 0.5% 的引用召回率,對於完整文件提高了 1.1%。

Auto-Formalizing Neuro-Symbolic Predictors

2610.01519v1 by Samuele Bortolotti, Weixin Chen, Han Zhao, Andrea Passerini, Stefano Teso, Antonio Vergari

Neuro-Symbolic (NeSy) predictors incorporate prior knowledge into the prediction process of neural networks, ensuring that outputs satisfy specified constraints, making them particularly suitable for high-stakes applications where compliance with domain knowledge is essential. A key bottleneck in this paradigm is the acquisition of symbolic constraints: encoding domain knowledge into logical formulas remains a manual and expert-intensive process. In this work, we investigate the extent to which auto-formalization via LLMs can systematically translate textual knowledge into symbolic knowledge that can be plugged into NeSy predictors. To this end, we introduce auto-nesy-bench, a new benchmark for evaluating constraint formalization and its impact on downstream accuracy of NeSy predictors. Through an extensive evaluation across several domains, we find that LLMs can formalize constraints to a meaningful extent, generating formulas that are often similar to those provided by human experts. Moreover, when the generated formulas are syntactically valid, they can lead to high-quality downstream predictions. The code and benchmark are available at https://unitn-sml.github.io/auto-nesy-bench/.

摘要:神經符號(NeSy)預測器將先前的知識納入神經網絡的預測過程中,確保輸出滿足特定的約束,使其特別適合於對領域知識遵循至關重要的高風險應用。這一範式中的一個關鍵瓶頸是符號約束的獲取:將領域知識編碼為邏輯公式仍然是一個手動且需要專家的過程。在本研究中,我們探討自動形式化通過大型語言模型(LLMs)在多大程度上可以系統性地將文本知識轉換為可以插入NeSy預測器的符號知識。為此,我們引入了auto-nesy-bench,一個新的基準,用於評估約束形式化及其對NeSy預測器下游準確性的影響。通過在幾個領域的廣泛評估,我們發現LLMs可以在有意義的程度上形式化約束,生成的公式通常與人類專家提供的公式相似。此外,當生成的公式在語法上有效時,它們可以導致高質量的下游預測。代碼和基準可在 https://unitn-sml.github.io/auto-nesy-bench/ 獲得。

Decision Titan: Test-Time Training for Long-Term Memory in Offline Reinforcement Learning

2610.01513v1 by Jude Waide, Robert Lieck

Long-term dependencies remain a major challenge for sequential decision-making in the field of AI: RNNs suffer from vanishing gradients and the limited expressivity of vector-based hidden states, whilst Transformer-based models are limited by the quadratic scaling of attention. Recent work has proposed tackling this problem with the Test-Time Training (TTT) framework, which stores episodic memories in the parameters of a neural network through gradient descent at both train and test-time. This approach has seen success in the domain of Natural Language Processing, however, to the best of our knowledge it has not yet been applied to the domain of Reinforcement Learning (RL), nor has there been a study analysing how this memory practically functions. In this paper, we study the potential of the TTT framework for offline RL by augmenting a Decision Transformer with TTT layers, dubbed the Decision Titan. We analyse performance and properties of the model in the X-Maze environment, an extension of T-Maze designed to test sequential memory, and investigate how the memory mechanism learns by visualising gate values over time. Our key findings are that Decision Titan can learn long-term dependencies with ranges 20x longer than the context window, generalises to lengths 1.7x the training data, but crucially temporal generalisation depends on the time embeddings used, and the ability to learn long-term dependencies depends on how the relevant information is encoded.

摘要:長期依賴性仍然是人工智慧領域中序列決策的一個主要挑戰:RNN 遭受消失梯度和基於向量的隱藏狀態表達能力有限的問題,而基於 Transformer 的模型則受到注意力的二次擴展限制。最近的研究提出了使用測試時訓練(Test-Time Training, TTT)框架來解決這個問題,該框架通過在訓練和測試期間的梯度下降將情節記憶儲存在神經網絡的參數中。這種方法在自然語言處理領域取得了成功,然而,據我們所知,它尚未應用於強化學習(Reinforcement Learning, RL)領域,也沒有研究分析這種記憶的實際運作方式。在本文中,我們通過增強決策 Transformer,並加入 TTT 層,稱之為 Decision Titan,來研究 TTT 框架在離線 RL 中的潛力。我們在 X-Maze 環境中分析模型的性能和特性,這是一個設計用來測試序列記憶的 T-Maze 擴展,並通過可視化門值隨時間的變化來調查記憶機制的學習方式。我們的主要發現是 Decision Titan 能夠學習長期依賴性,其範圍比上下文窗口長 20 倍,對訓練數據的長度進行 1.7 倍的泛化,但關鍵是時間泛化依賴於使用的時間嵌入,而學習長期依賴性的能力則取決於相關信息的編碼方式。

OverAct: Measuring and Mitigating Proactive Over-Authorization in LLM Tool-Calling Agents

2610.01508v1 by Taolin Zhang, Jiuheng Wan, Hanyu Wang, Tingyuan Hu, Chengyu Wang

LLM agents with tool-calling capabilities can access external services and private user data, but they may retrieve more information than a user's request explicitly requires. We study this behavior in structured tool-calling agents and term it proactive over-authorization. This setting differs from filesystem-level coding agents because the main risk is unnecessary access to private data. We introduce OverAct, a controlled benchmark spanning eight privacy-sensitive domains with deterministic, judge-free scoring, together with an interpretive decision-theoretic framework that yields three testable predictions. Across seven models from four families, all models significantly exceed authorized scope. Request specificity is the strongest predictor of severity, over-authorization grows sublinearly with tool-pool size, and decoding temperature has little effect. These patterns are consistent with a cost-asymmetry account, suggesting that over-authorization arises more from structural decision tendencies than from decoding randomness. We also propose SelfAudit, a zero-shot inference-time method that generates request-grounded justifications and filters unjustified calls before execution. Ablation shows that explicit filtering is the main driver of scope reduction. SelfAudit reduces privacy-oriented excess by 43% without oracle knowledge.

摘要:LLM 代理具有工具調用能力,可以訪問外部服務和私人用戶數據,但它們可能檢索比用戶請求明確要求的更多信息。我們研究這種行為在結構化工具調用代理中,並將其稱為主動過度授權。這種設置不同於文件系統級編碼代理,因為主要風險是對私人數據的不必要訪問。我們引入了 OverAct,一個涵蓋八個隱私敏感領域的受控基準,具有確定性、無評判的評分,並結合了一個解釋性決策理論框架,產生三個可測試的預測。在來自四個家族的七個模型中,所有模型的表現均顯著超出授權範圍。請求的具體性是嚴重程度的最強預測因子,過度授權隨著工具池大小的增長而次線性增長,而解碼溫度的影響很小。這些模式與成本不對稱的解釋一致,表明過度授權更多地源於結構性決策傾向,而非解碼隨機性。我們還提出了 SelfAudit,一種零樣本推斷時的方法,生成基於請求的理由並在執行前過濾不合理的調用。消融實驗顯示,明確過濾是範圍減少的主要驅動因素。SelfAudit 在沒有神諭知識的情況下將隱私導向的過剩減少了 43%。

A Multi-Agent LLM Framework for Personalized Health Checkup Interpretation and Guidance

2610.01451v1 by HyungJun Kim, Taehan Lee, Soojin Cheon

Personalized interpretation of health checkup results requires reasoning across longitudinal records, medical knowledge, lifestyle guidance, and healthcare navigation. We present a multi-agent large language model (LLM) system that identifies multiple intents, maps each to a task-specific agent, executes them in parallel, and synthesizes their outputs. We compared answers generated in Single Agent and Multi Agent settings on 120 Korean compound queries combining two to four requirements, using synthetic health checkup records. The Multi Agent improved the weighted LLM-judge score from 1.695 to 1.797 (p = 0.027), and three additional LLM judges showed consistent improvements ($Δ$ = +0.111 to +0.186, all p < 0.05). The gains came from usefulness, consistency, and the handling of every requirement in compound queries, whereas numerical accuracy and grounding improved significantly under only one of the four judges and medical safety did not differ, and critical failures occurred at similar rates (Single Agent 15.0% vs. Multi Agent 13.3%). Two human evaluators preferred Multi Agent in 66.7% and 68.3% of pairwise comparisons. Multi Agent execution increased latency and cost by 1.31$\times$ and 2.02$\times$, respectively. In exploratory subgroup analyses, the improvement was concentrated in queries involving personal-record lookup.

摘要:個性化的健康檢查結果解釋需要跨越長期記錄、醫學知識、生活方式指導和醫療導航的推理。我們提出了一個多代理大型語言模型(LLM)系統,該系統識別多個意圖,將每個意圖映射到特定任務的代理,並平行執行它們,最後綜合其輸出。我們比較了在單代理和多代理設置下,使用合成健康檢查記錄對120個韓國複合查詢生成的答案,這些查詢結合了兩到四個需求。多代理將加權LLM評審分數從1.695提高到1.797(p = 0.027),另外三位LLM評審顯示出一致的改善($Δ$ = +0.111到+0.186,所有p < 0.05)。這些增益來自於有用性、一致性以及對複合查詢中每個需求的處理,而數值準確性和基礎資料僅在四位評審中的一位顯著改善,醫療安全則沒有差異,且重大失誤的發生率相似(單代理15.0%對多代理13.3%)。兩位人類評估者在66.7%和68.3%的成對比較中偏好多代理。多代理執行使延遲和成本分別增加了1.31$\times$和2.02$\times$。在探索性子群分析中,改善集中在涉及個人記錄查詢的問題上。

2610.01393v1 by Nouha Hayouni, Sheeba Samuel, Alsayed Algergawy

Constructing typed, justified semantic links between ontologies is essential for enabling interoperability across heterogeneous and interdisciplinary knowledge domains. However, manually curating such links is difficult to scale. To address this challenge, we propose an end-to-end framework for ontology network construction that automates the discovery and generation of both intra-domain and inter-domain relationships. Our approach combines domain-adapted DistilBERT embeddings for dense contextual representation, clustering-based pre-filtering to reduce the candidate search space, and GPT-4o-driven relationship generation via iterative prompt engineering to produce semantically rich, interpretable links. Applied to ReproduceMeON - a network of 33 ontologies spanning machine learning, microscopy, computational science, and experimental workflow - the pipeline reduces approximately 800k raw concept pairs to 95k high-quality candidates. Human expert validation of 429 generated relationships by two independent annotators yields an overall precision of 80.19% (91.49% on high-certainty annotations) and an F1 of 0.890, with substantial inter-annotator agreement. Comparative experiments against five similarity-based baselines, including Sentence-BERT, show a substantial performance gap (best baseline F1 = 0.581), while an ablation study demonstrates that similarity-based methods alone fail to discriminate valid from invalid relationships (AUC approx 0.5) on the filtered candidate set. These findings highlight the necessity of LLM-based reasoning over concept roles and domain semantics for accurate relationship construction.

摘要:建構有類型、對齊的語義連結在本體之間對於實現異質和跨學科知識領域的互操作性至關重要。然而,手動策劃這些連結難以擴展。為了解決這一挑戰,我們提出了一個端到端的本體網絡建構框架,該框架自動發現和生成內域和跨域關係。我們的方法結合了針對特定領域調整的DistilBERT嵌入以獲得密集的上下文表示、基於聚類的預過濾以減少候選搜索空間,以及通過迭代提示工程驅動的GPT-4o關係生成,以產生語義豐富、可解釋的連結。應用於ReproduceMeON——一個涵蓋機器學習、顯微鏡學、計算科學和實驗工作流程的33個本體的網絡——該流程將大約80萬個原始概念對減少到9.5萬個高質量候選。由兩位獨立註釋者對429個生成關係進行的人類專家驗證產生了整體精確度80.19%(高確定性註釋為91.49%)和F1值0.890,並且註釋者之間的協議顯著。與五個基於相似性的基準進行的比較實驗,包括Sentence-BERT,顯示出顯著的性能差距(最佳基準F1 = 0.581),而消融研究表明,僅依賴相似性的方法無法區分有效和無效的關係(AUC約0.5)在過濾的候選集上。這些發現突顯了基於LLM的推理在概念角色和領域語義上的必要性,以實現準確的關係建構。

Generation Provenance Before Behavior Attribution: Auditing Synthetic Speech Research Objects

2610.01378v1 by Sidi Chang, Peiying Zhu

Attributing model behavior to synthetic training data requires knowing what produced each training item before estimating what that item caused. A waveform-label pair does not preserve this knowledge. We propose a generation-provenance substrate in which a synthetic research object binds source specification, generated content, waveform, target, fact requirements, quality signals, review lineage, and immutable manifest identity. Producer and selection mechanism determine evidentiary meaning; storage location and variable name do not. We audit this substrate in a private Japanese care-handoff pipeline. A 113-asset review population contains 1.552 hours of synthetic speech across six scenario families; all items have linked audio, transcripts, candidate notes, and fact checklists, but human evidence is selective and source-specific. Two faithful-only manifests are scenario-seed-disjoint and immutably versioned, while exact upstream attribution remains blocked by floating generator aliases, missing per-clip TTS and code stamps, and an unversioned checking prompt. We argue that generation provenance is necessary but not sufficient for behavior attribution: it defines the candidate causal graph and audit units, whereas contributive attribution still requires frozen training runs and intervention or influence evidence. The paper contributes a compact provenance contract, an audit protocol, and a bounded case study for synthetic-data attribution; controlled research access may be offered, but we do not claim causal training-data attribution, clinical validity, or unrestricted public release.

摘要:將模型行為歸因於合成訓練數據需要了解每個訓練項目是如何產生的,然後才能估計該項目造成了什麼。波形-標籤對並不保留這種知識。我們提出了一種生成來源基底,其中合成研究對象綁定了源規範、生成內容、波形、目標、事實要求、質量信號、審查血統和不可變的清單身份。生產者和選擇機制決定了證據意義;存儲位置和變量名稱則不然。我們在一個私有的日本護理交接管道中審計這一基底。一個包含113個資產的審查群體包含了六個情境系列中的1.552小時合成語音;所有項目都鏈接了音頻、轉錄、候選筆記和事實檢查清單,但人類證據是選擇性的且特定於來源。兩個僅限忠實的清單在情境種子上是不相交且不可變版本的,而精確的上游歸因仍然受到浮動生成器別名、缺失的每段TTS和代碼印記以及未版本化的檢查提示的阻礙。我們主張生成來源對於行為歸因是必要但不充分的:它定義了候選因果圖和審計單位,而貢獻性歸因仍然需要凍結的訓練運行和干預或影響證據。本文貢獻了一個簡潔的來源合約、一個審計協議和一個有界的合成數據歸因案例研究;可能會提供受控的研究訪問,但我們不聲稱因果訓練數據歸因、臨床有效性或不受限制的公開發布。

ARCCS: An Automated Regulatory Compliance Checking System

2610.01345v1 by Giorgos Filandrianos, José Menezes, Chrysoula Zerva, Alessandro Gianola

Regulatory compliance checking - deciding whether a target document satisfies the obligations of a regulation - requires interpreting dense legal text, identifying which provisions apply, and grounding each decision in explicit evidence. We present ARCCS, an end-to-end, automated, agentic, and regulation-agnostic Legal NLP system for compliance checking. ARCCS decomposes raw regulatory text into atomic, traceable requirements and evaluates a target document against them using retrieved evidence, confidence scores, and human-interpretable justifications. This design decouples compliance assessment from any fixed regulatory template or predefined rule set, enabling the pipeline to operate over regulations of varying size and structure. We evaluate ARCCS in two complementary settings. First, in a GDPR policy-document evaluation, LLM-based judges find its decisions and justifications legally and evidentially consistent in up to 96.67% of the assessed cases. Second, on an EU public-procurement benchmark comprising more than 1,200 individual rule checks, the system attains 98.8% accuracy in violation detection. ARCCS is, to our knowledge, the first fully open-source system for end-to-end regulatory compliance checking and auditable report generation.

摘要:監管合規檢查 - 決定目標文件是否滿足法規的義務 - 需要解釋密集的法律文本,識別適用的條款,並將每個決策基於明確的證據。我們提出了 ARCCS,一個端到端、自動化、主動且與法規無關的法律自然語言處理系統,用於合規檢查。ARCCS 將原始法規文本分解為原子、可追溯的要求,並使用檢索的證據、信心分數和人類可解釋的理由來評估目標文件。這一設計將合規評估與任何固定的法規模板或預定的規則集解耦,使得該流程能夠在不同大小和結構的法規上運行。我們在兩個互補的環境中評估 ARCCS。首先,在 GDPR 政策文件評估中,基於 LLM 的評審在多達 96.67% 的評估案例中發現其決策和理由在法律和證據上是一致的。其次,在一個包含超過 1,200 個個別規則檢查的歐盟公共採購基準上,該系統在違規檢測中的準確率達到 98.8%。據我們所知,ARCCS 是第一個完全開源的端到端監管合規檢查和可審計報告生成系統。

An ontology for cross-sectoral crisis management: core and public health modules

2610.01326v1 by Aldo Gangemi, Rita T. Sousa, Luigi Asprino, Giorgia Lodi, Andrea G. Nuzzolese, Valentina Presutti, Johannes Gysen, Diana F. Sousa, Luigi Spagnolo

This paper presents the European Crisis Management Ontology (ECMO), a modular OWL-based ontology intended as a cross-sectoral reference for disaster risk reduction and response. ECMO is designed to be organised as a network of ontological modules. Among the modules, ECMO-CORE captures fundamental crisis management concepts such as hazard, event, exposure, impact, and response measure and uses ontology design patterns and the OWL2 punning technique to resolve ambiguities between hazard types and event manifestations. In addition, domain-specific modules are defined as in the case of the public health module aligned with SNOMED CT and ICD-11. To demonstrate the resource's utility, we used ECMO to represent the data of the Epidemic Intelligence from Open Sources system of the Joint Research Centre to generate an end-to-end pipeline that populates an ECMO-compliant knowledge graph from unstructured epidemiological news. Initial results demonstrate that ECMO provides the formal guardrails necessary for consistent and unified knowledge representation and integration. The ontology is publicly available at https://doi.org/10.5281/zenodo.20070268 and is released under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.

摘要:這篇論文介紹了歐洲危機管理本體(ECMO),這是一個基於OWL的模組化本體,旨在作為災害風險減少和應對的跨領域參考。ECMO的設計是作為本體模組的網絡組織。 在這些模組中,ECMO-CORE捕捉了基本的危機管理概念,如危險、事件、暴露、影響和應對措施,並使用本體設計模式和OWL2的雙義技術來解決危險類型和事件表現之間的歧義。此外,還定義了特定領域的模組,例如與SNOMED CT和ICD-11對齊的公共衛生模組。為了展示該資源的實用性,我們使用ECMO來表示聯合研究中心的開放來源流行病情報系統的數據,以生成一個從非結構化流行病學新聞填充ECMO合規知識圖譜的端到端管道。初步結果顯示,ECMO提供了必要的正式框架,以實現一致和統一的知識表示和整合。該本體可在https://doi.org/10.5281/zenodo.20070268上公開獲得,並根據創用CC 4.0國際版(CC BY 4.0)授權發布。

Dependency-Aware Reward Shaping for Agentic Reinforcement Learning

2610.01207v1 by Ziyi Chen, Yan Zhang, Jianhui Wei, Daoan Zhang, Zuozhu Liu

When training large language models with reinforcement learning, terminal rewards provide little guidance about which steps matter. Common methods for assigning step credit overlook that work built on uncorrected mistakes is wasted while independent work remains valid. With only a final success/failure reward, every step in a failed episode has zero total future reward, even when it made progress. We propose Dependency-Aware Reward Shaping (DARS), which represents task progress as predicates linked by prerequisite relations and assigns step-level credit over the dependency graph. An annotator marks which predicates each step verifies, invalidates, or repairs. Verified predicates are discounted according to graph distance from the nearest broken prerequisite, while independent predicates are unaffected. Repairs update these weights based on any errors that remain; invalidated predicates need re-verification to regain credit. A fixed potential converts these annotations into signed per-step rewards. A common reward and annotation interface allows DARS to integrate with a range of reasoning and agentic training methods, such as GiGPO and ARPO/AEPO, without changing their rollout strategies or optimizers. Across five task families and models from 1.5B to 8B, DARS improves success by up to 10 points over GiGPO trained with the same budget and harness (ALFWorld), raises the WebShop task score and Search-R1 QA accuracy, complements AEPO's entropy-based training on AIME24/25 with a Python interpreter, and exceeds OmniOPD in controlled tool-free reasoning comparisons at 1.7B and 4B. Ablations show that step-level credit, dependency attenuation, and graph topology each contribute. On ALFWorld, a distilled 8B annotator matches the API annotator, enabling DARS to run efficiently without a frontier judge. Code is available at https://github.com/JianhuiWei7/DARS.

摘要:在使用強化學習訓練大型語言模型時,最終獎勵對於哪些步驟重要的指導作用有限。常見的步驟信用分配方法忽視了基於未修正錯誤的工作是浪費的,而獨立工作仍然有效。僅有的最終成功/失敗獎勵使得在失敗的情況下,每一步的總未來獎勵為零,即使它有所進展。我們提出了依賴感知獎勵塑造(DARS),它將任務進展表示為由前置關係連結的謂詞,並在依賴圖上分配步驟級別的信用。一名標註者標記每一步驗證、無效或修復了哪些謂詞。經過驗證的謂詞根據與最近的破損前置條件的圖距離進行折扣,而獨立謂詞則不受影響。修復根據仍然存在的任何錯誤更新這些權重;無效的謂詞需要重新驗證以恢復信用。一個固定的潛力將這些標註轉換為簽名的每步獎勵。共同的獎勵和標註介面使DARS能夠與一系列推理和代理訓練方法(如GiGPO和ARPO/AEPO)集成,而無需改變它們的展開策略或優化器。在五個任務系列和從1.5B到8B的模型中,DARS在與相同預算和設備(ALFWorld)訓練的GiGPO相比,成功率提高了最多10個點,提升了WebShop任務得分和Search-R1 QA準確性,並補充了AEPO在AIME24/25上基於熵的訓練,使用Python解釋器,並在1.7B和4B的受控無工具推理比較中超過了OmniOPD。消融實驗顯示,步驟級信用、依賴衰減和圖拓撲各自都有貢獻。在ALFWorld上,一個精煉的8B標註者與API標註者相匹配,使DARS能夠高效運行而無需前沿評判者。代碼可在 https://github.com/JianhuiWei7/DARS 獲得。

Federated Agent Optimization

2610.01195v1 by Qiang Yang, Zhiqiang Kou, Xueyi Zhang, Dong-Dong Wu, Hanlin Gu, Jing Guo, Yang Liu, Di Jiang, Qian Xu

Large language model (LLM) agents increasingly operate in private environments and accumulate valuable experience from task execution, tool use, feedback, and local knowledge. Yet such experience is distributed across organizations and cannot be directly shared because of privacy and proprietary constraints. Conventional federated learning is insufficient for this setting, as agent capabilities extend beyond model parameters to memory, tools, rewards, skills, and structured knowledge. In this paper, we formulate \textbf{Federated Agent Optimization (FAO)}, which studies how distributed agents can collaboratively improve through controlled information exchange while keeping raw data, complete trajectories, and private knowledge local. We define FAO as a multi-objective problem balancing agent utility, privacy leakage, and communication cost, and organize its optimization space across policy, memory, tool use, reward, and structured knowledge and skills. We further characterize how private experience can be abstracted, protected, aggregated, and adapted into transferable capabilities, providing a unified view of how agents can benefit from one another without direct experience sharing. Finally, we identify the key challenges of FAO and outline several promising directions for future research toward trustworthy federated agent systems.

摘要:大型語言模型 (LLM) 代理人越來越多地在私密環境中運作,並從任務執行、工具使用、反饋和本地知識中積累寶貴的經驗。然而,這些經驗分散在各個組織中,由於隱私和專有限制,無法直接共享。傳統的聯邦學習在這種情況下是不夠的,因為代理人的能力超出了模型參數,還包括記憶、工具、獎勵、技能和結構化知識。在本文中,我們提出了\textbf{聯邦代理優化 (FAO)},研究分散的代理人如何通過受控的信息交換協作改進,同時保持原始數據、完整的軌跡和私有知識的本地性。我們將FAO定義為一個多目標問題,平衡代理效用、隱私洩漏和通信成本,並在政策、記憶、工具使用、獎勵以及結構化知識和技能之間組織其優化空間。我們進一步描述了如何將私有經驗抽象化、保護、聚合和適應為可轉移的能力,提供了一個統一的視角,說明代理人如何在不直接共享經驗的情況下相互受益。最後,我們確定了FAO的主要挑戰,並概述了幾個有前景的未來研究方向,以促進可信的聯邦代理系統。

Temporally-Resolved Token Attribution Reveals the Generation Dynamics of Diffusion Language Models

2610.01177v1 by Darpan Aswal, Céline Hudelot

This work presents Diffusion Layer Integrated Gradients (DLIG), a token attribution method for diffusion language models (DLMs) that extends Integrated Gradients (IG~\cite{sundararajan2017axiomatic}) to arbitrary layers and denoising steps. DLIG attributes a DLM's progressive commitment to a self-generated or fixed completion for an input prompt. We establish direct correspondences between DLIG and the IG axioms of completeness, implementation invariance, linearity, and symmetry preservation. As a lightweight complement to interventional analysis, DLIG provides an inexpensive first check of mechanistic hypotheses across the denoising trajectory. We demonstrate this on word-sense disambiguation, multi-hop graph reasoning, and sentence infilling, revealing how DLMs draw on inputs across positions, layers, and denoising steps.

摘要:這項工作提出了擴展了整合梯度(Integrated Gradients, IG~\cite{sundararajan2017axiomatic})至任意層和去噪步驟的擴散層整合梯度(Diffusion Layer Integrated Gradients, DLIG),這是一種用於擴散語言模型(Diffusion Language Models, DLMs)的標記歸因方法。DLIG 將 DLM 對於自生成或固定完成的輸入提示的逐步承諾進行歸因。我們建立了 DLIG 與 IG 完整性、實現不變性、線性和對稱性保持的公理之間的直接對應關係。作為對介入分析的輕量補充,DLIG 提供了一種便宜的初步檢查,用於在去噪過程中檢驗機理假設。我們在詞義消歧、多跳圖推理和句子填充上展示了這一點,揭示了 DLM 如何在不同位置、層和去噪步驟中利用輸入。

OrbitTAMP: Grounding Language Models for Task and Motion Planning in Spacecraft Rendezvous

2610.01093v1 by Yuji Takubo, Daniele Gammelli, Marco Pavone, Simone D'Amico

Spacecraft rendezvous and proximity operations (RPO) are currently planned through an expertise-intensive process in which engineers translate high-level operational intent into safe, dynamically feasible trajectories, creating a bottleneck to scalable operations. Large language model (LLM)-based agents could offer an intuitive interface for this process, although their outputs are not inherently grounded in orbital dynamics, operational constraints, or the structure of admissible spacecraft maneuvers. To exploit their semantic reasoning while ensuring the generated plan's physical validity, this paper presents a hierarchical framework for spacecraft task-and-motion planning (TAMP) that grounds LLM reasoning in a graph of reusable behaviors and domain-specific planning modules. Within this framework, a pretrained LLM maps a natural-language command to a partial mission specification. The associated planners then resolve unspecified decisions within the admissible operational space. Finally, trajectory optimization converts the completed mission specification into a dynamically feasible trajectory. Numerical experiments demonstrate that this architecture substantially improves intent recovery over direct LLM generation, achieving 98% exact recovery of partial mission specifications across all evaluated splits when backed by frontier LLMs. Additional test-time-compute experiments show that, for a compact 9B model, verifier-guided revision increases exact recovery from 75% to 88%, while broader behavior-plan search independently improves selection among admissible trajectory realizations. Overall, these results establish a scalable and auditable foundation for language-driven agentic planning of spacecraft RPO.

摘要:太空船會合與接近操作(RPO)目前是通過一個需要專業知識的過程進行規劃,在這個過程中,工程師將高層次的操作意圖轉化為安全且動態可行的軌跡,這造成了可擴展操作的瓶頸。基於大型語言模型(LLM)的代理可以為這個過程提供直觀的界面,儘管它們的輸出並不固有地基於軌道動力學、操作限制或可接受的太空船機動結構。為了利用它們的語義推理,同時確保生成計劃的物理有效性,本文提出了一個太空船任務與運動規劃(TAMP)的分層框架,該框架將LLM推理基於可重用行為和特定領域規劃模塊的圖進行基礎化。在這個框架內,預訓練的LLM將自然語言命令映射到部分任務規範。相關的規劃者然後在可接受的操作空間內解決未指定的決策。最後,軌跡優化將完成的任務規範轉換為動態可行的軌跡。數值實驗表明,這個架構顯著改善了意圖恢復,相較於直接的LLM生成,當得到前沿LLM的支持時,在所有評估的分割中實現了98%的部分任務規範的精確恢復。額外的測試時間計算實驗顯示,對於一個緊湊的9B模型,驗證者引導的修訂將精確恢復從75%提高到88%,而更廣泛的行為計劃搜索則獨立改善了在可接受的軌跡實現中的選擇。總體而言,這些結果為基於語言驅動的太空船RPO代理規劃建立了一個可擴展和可審計的基礎。

MOMAT: Mixture of Multiple Atlases for Low-Power Jailbreak Defense of Quantized LLMs

2610.01058v1 by Boyang Li, Bingyu Shen, Weihao Hong, Zhiyuan Jiang, Xinlei Guan, Yan Ma, Miles Q. Li, Yi Sheng, Ruiyang Qin

Quantized large language models are increasingly deployed on edge devices for their low latency and energy efficiency. However, model quantization weakens alignment safeguards, leaving qLLMs (quantized large language models) highly vulnerable to jailbreak attacks. To address this challenge, we present MOMAT (Mixture of Multiple Atlases), a hardware-enhanced safety framework that combines structured knowledge retrieval with low-power defense acceleration. Each atlas represents a semantic cluster of harmful or benign sample sets and policy templates, enabling domain-localized Retrieval-Augmented Generation guarding that mitigates the curse of dimensionality and the resulting semantic sparsity problem in large, heterogeneous safety databases. MOMAT retrieves top-$k$ similarity features from all atlases for each prompt and evaluates them using a lightweight MoE (Mixture of Experts) detector, while a CiM (Compute-in-Memory)-accelerated similarity engine performs fast, low-power atlas-local retrieval. MOMAT's CiM-based retrieval accelerates a 100-query batch from 15,052.44 ms to 3,207.21 ns (a $4.69 \times 10^6\times$ speedup) and reduces energy from $8.1 \times 10^7$ $μ$J to 3.32 $μ$J, yielding an approximately $2.5 \times 10^5\times$ energy reduction over DRAM-based (Raspberry Pi) baselines. Red-team evaluations across standard benchmarks show that MOMAT matches the defense performance of state-of-the-art methods while avoiding benign overkill and providing substantial efficiency gains, demonstrating that CiM-based modular defenses can make edge-deployed qLLMs both safer and more energy-efficient. We will release the full 223.2k-sample dataset to foster future research.

摘要:量化的大型語言模型越來越多地部署在邊緣設備上,因為它們具有低延遲和能量效率。然而,模型量化削弱了對齊保護,讓qLLMs(量化大型語言模型)對越獄攻擊高度脆弱。為了解決這一挑戰,我們提出了MOMAT(多重圖譜混合),這是一個硬體增強的安全框架,結合了結構化知識檢索和低功耗防禦加速。每個圖譜代表一個有害或良性樣本集和政策模板的語義集群,實現了域本地化的檢索增強生成保護,減輕了維度詛咒和大型異構安全數據庫中產生的語義稀疏問題。MOMAT為每個提示從所有圖譜中檢索前$k$相似特徵,並使用輕量級的MoE(專家混合)檢測器對其進行評估,而CiM(內存計算)加速的相似性引擎則執行快速、低功耗的圖譜本地檢索。MOMAT的基於CiM的檢索將100查詢批次的時間從15,052.44毫秒加速到3,207.21納秒($4.69 \times 10^6\times$加速),並將能量從$8.1 \times 10^7$ $μ$J減少到3.32 $μ$J,實現了約$2.5 \times 10^5\times$的能量減少,相較於基於DRAM(樹莓派)的基準。針對標準基準的紅隊評估顯示,MOMAT的防禦性能與最先進的方法相當,同時避免了良性過度防護,並提供了顯著的效率增益,證明基於CiM的模組化防禦可以使邊緣部署的qLLMs更加安全和能量高效。我們將釋出完整的223.2k樣本數據集,以促進未來的研究。

Capturing In-Context Learning Dynamics with Task Operators

2610.01054v1 by Guangzhi Xiong, Zhenghao He, Bohan Liu, Sanchit Sinha, Wenqian Ye, Aidong Zhang

In-context learning (ICL) enables language models to perform new tasks from demonstrations without weight updates. However, every ICL inference requires processing the full set of examples, resulting in inefficient deployments, and how ICL works mechanistically is not fully understood. Prior work compresses ICL into fixed activation vectors extracted from specific layers or positions, but these input-independent interventions fail on complex tasks where the output depends on fine-grained interactions with the input. By analyzing the ICL forward pass, we show that each attention head's output is an affine transformation of its context-masked counterpart, and that the parameters of this transformation are empirically stable across samples for a given task. Building on this, we introduce Task Operator (TO), which replays this transformation as an analytically derived update to the attention output projection. Across lexical, algorithmic, and reasoning tasks, TO achieves the best overall performance among prior methods and substantially narrows the gap between zero-shot inference and ICL. We further show that the extracted knowledge concentrates in a task-specific sparse circuit across layers and positions, and that averaging operators from disjoint demonstration batches enables effective many-shot scaling without expanding the context window. Our code is available at https://github.com/gzxiong/task_operator.

摘要:在上下文學習(ICL)中,語言模型能夠從示範中執行新任務,而無需更新權重。然而,每次 ICL 推理都需要處理完整的示例集,導致部署效率低下,並且 ICL 的機制運作方式尚未完全理解。先前的研究將 ICL 壓縮為從特定層或位置提取的固定激活向量,但這些與輸入無關的干預在輸出依賴於與輸入的細緻互動的複雜任務上失敗。通過分析 ICL 的前向傳遞,我們顯示每個注意力頭的輸出是其上下文遮罩對應物的仿射變換,並且這種變換的參數在給定任務的樣本中經驗上是穩定的。基於此,我們引入了任務運算符(TO),它將這種變換重播為對注意力輸出投影的解析衍生更新。在詞彙、算法和推理任務中,TO 在先前方法中實現了最佳的整體性能,並大幅縮小了零-shot 推理和 ICL 之間的差距。我們進一步顯示,提取的知識集中在跨層和位置的任務特定稀疏電路中,並且從不相交的示範批次中平均運算符能夠有效地進行多次擴展,而無需擴大上下文窗口。我們的代碼可在 https://github.com/gzxiong/task_operator 獲得。

Beyond Answer Confidence: A Controlled Audit of Self-Knowledge in a Black-Box Decision Model

2610.01006v1 by Sharath M Shankaranarayana, Davor Runje, Jan Jannink

Decision models return probabilities intended for routing, abstention and automated action. Calibration makes those probabilities useful on average, but does not establish whether low confidence reflects chance or missing knowledge, nor whether confidence falls when a model moves beyond what it knows. We audit this distinction in Jev, a decision model, with over 15 public datasets and 6 generated task families, with paired interventions that vary the information supplied for a fixed item. Jev's confidence is calibrated on familiar closed-choice tasks but fails as an indicator of missing knowledge: with no answer-relevant information it assigns up to 0.80 to a salient option, and on news beyond an observed knowledge boundary it exceeds accuracy by 0.21--0.33, a gap that recalibration on earlier months does not close. Targeted yes/no questions give sharper readouts of the case: whether an outcome is settled (AUROC 1.00) and whether the evidence suffices (0.95, against 0.85 for confidence on the same items). Asking whether Jev knows the answer appears to flag fabricated entities and post-boundary news (0.91), but with realistic names or with dates removed it shows no advantage over answer uncertainty. Black-box knowledge audits therefore need explicit controls for surface cues. Code: https://github.com/Syntheme/beyond-answer-confidence.

摘要:決策模型返回旨在路由、放棄和自動行動的概率。校準使這些概率在平均情況下變得有用,但並未確定低信心是否反映了隨機性或缺失的知識,也未確定當模型超出其已知範疇時信心是否會下降。我們在Jev這個決策模型中審核這一區別,使用超過15個公共數據集和6個生成的任務系列,並配對干預,變化固定項目的信息供應。Jev的信心在熟悉的封閉選擇任務上經過校準,但作為缺失知識的指標卻失效:在沒有與答案相關的信息時,它對一個顯著選項賦予高達0.80的概率,並且在超過觀察知識邊界的新聞中,其準確性超過0.21--0.33,這一差距在對早期月份的重新校準中並未縮小。針對性的是/否問題提供了更清晰的案例讀數:結果是否已確定(AUROC 1.00)以及證據是否足夠(0.95,相較於對同一項目的信心為0.85)。詢問Jev是否知道答案似乎標記了虛構實體和邊界後的新聞(0.91),但在使用現實名稱或去除日期的情況下,並未顯示出相對於答案不確定性的優勢。因此,黑箱知識審核需要對表面線索進行明確控制。代碼:https://github.com/Syntheme/beyond-answer-confidence。

Distilling Directional Verification

2610.00997v1 by Jungseob Lee, Sugyeong Eo, Seongtae Hong, Seungyoon Lee, Chanjun Park, Jaehyung Seo, Heuiseok Lim

Knowledge distillation aims to transfer the factual knowledge of large language models to smaller models for efficient deployment. Yet a teacher may recall a relation in one direction while failing to generate the answer in the reverse direction. Distillation from its generated answers can therefore propagate this directional limitation to the student. The same teacher can nevertheless recognize such an answer by scoring the relation in the direction it knows. We introduce directional label distillation, in which frozen teachers score candidate answers in that known direction and the best-scoring candidate becomes the student's training target. On facts about parents and their children, known-direction scoring yields more accurate labels than scoring the requested direction, even after tuned corrections for name priors. With prior-corrected scores, the better direction depends on the facts rather than the template, and reverses on mined facts whose notable entity is the parent rather than the child. With the evaluated children's forward facts withheld, students trained on known-direction labels improve open-ended accuracy on their trained queries by 13 to 15 points over students trained on prior-corrected reverse labels. After generated answers are matched to a fixed name list by lexical similarity, students reproduce nearly all selected labels. Their accuracy largely follows label quality. The label advantage holds on unscreened queries and when candidates are retrieved without inserting correct answers. Our findings show that directional verification mitigates the transfer of errors from teacher-generated answers to students by providing more accurate training targets. Code is available at https://github.com/js-lee-AI/directional-verification.

摘要:知識蒸餾旨在將大型語言模型的事實知識轉移到較小的模型,以便高效部署。然後,教師可能在一個方向上回憶起一個關係,但未能在相反方向生成答案。因此,從其生成的答案進行蒸餾可能會將這一方向性限制傳播給學生。然而,同一位教師仍然可以通過在其已知方向上對關係進行評分來識別這樣的答案。我們引入了方向性標籤蒸餾,其中凍結的教師在已知方向上對候選答案進行評分,得分最高的候選者成為學生的訓練目標。在有關父母及其子女的事實中,已知方向的評分比要求方向的評分產生更準確的標籤,即使在對名稱先驗進行調整後也是如此。使用經過先驗修正的分數,更好的方向取決於事實而不是模板,並在挖掘的事實中反轉,當其顯著實體是父母而不是子女時。在評估的子女的前向事實被保留的情況下,基於已知方向標籤訓練的學生在其訓練查詢上的開放式準確率提高了13到15個點,超過了基於先驗修正的反向標籤訓練的學生。在通過詞彙相似性將生成的答案與固定名稱列表匹配後,學生幾乎重現了所有選定的標籤。他們的準確性在很大程度上取決於標籤質量。標籤優勢在未篩選的查詢上以及在檢索候選者時未插入正確答案的情況下依然存在。我們的研究結果顯示,方向性驗證通過提供更準確的訓練目標來減輕教師生成的答案向學生轉移錯誤的影響。代碼可在 https://github.com/js-lee-AI/directional-verification 獲得。

Structure-agnostic Causal Representation Learning

2610.00968v1 by Arman Behnam, Binghui Wang

Causal representation learning aims to discover robust features by exploiting the causal structure underlying data generation. Existing methods require specifying the causal structure a priori, yet different structures demand fundamentally incompatible invariance constraints, and misspecification leads to representations that discard predictive information. We introduce SaCRL, a framework that jointly identifies the causal structure and learns the corresponding invariant representation without prior structural knowledge. Our approach formulates structure selection as a soft optimization over candidate invariances using HSIC-based violation metrics, with adaptive weights that automatically concentrate on the achievable structure. We provide theoretical guarantees for structure identification, including under random-feature approximation, invariance satisfaction, and out-of-distribution generalization. Empirically, SaCRL recovers the true structure on synthetic and semi-synthetic Bayesian-network benchmarks, outperforms fixed-invariance baselines on Colored MNIST, achieves state-of-the-art accuracy on three DomainBed benchmarks (PACS, VLCS, OfficeHome), and degrades gracefully under structural misspecification and limited environment diversity. Code is available at: https://github.com/ArmanBehnam/sacrl.

摘要:因果表示學習旨在通過利用數據生成背後的因果結構來發現穩健的特徵。現有的方法需要事先指定因果結構,但不同的結構要求根本上不相容的不變性約束,且錯誤指定會導致丟失預測信息的表示。我們引入了SaCRL,一個框架,能夠在沒有先驗結構知識的情況下共同識別因果結構並學習相應的不變表示。我們的方法將結構選擇表述為對候選不變性的軟優化,使用基於HSIC的違規度量,並具有自適應權重,自動集中於可實現的結構。我們提供了結構識別的理論保證,包括隨機特徵近似、不變性滿足和分佈外泛化。在實證上,SaCRL在合成和半合成的貝葉斯網絡基準上恢復了真實結構,在Colored MNIST上超越了固定不變性基準,在三個DomainBed基準(PACS、VLCS、OfficeHome)上達到了最先進的準確率,並在結構錯誤指定和環境多樣性有限的情況下優雅降級。代碼可在以下網址獲得:https://github.com/ArmanBehnam/sacrl。

ABDA-NL: A Natural-Language Scenario Explorer for Argument-Based Reasoning

2610.00947v1 by Shawn Bowers, Martin Caminada, Haoyang Liu, Bertram Ludäscher

ABDA-NL adds a natural-language interface to ABDA, a system for argument-based discussion using ASPIC- knowledge bases under grounded semantics. Users see which conclusions are accepted, rejected, or undecided, open an interactive rendering of the grounded discussion game to learn why, explore what-if alternatives by suspending assumptions and rules or changing preferences, ask questions that are answered from a scenario's reference documents, and author new facts, assumptions, and rules in plain English. A large language model provides the bridge between language and formalism: it answers questions from the documents and the current state of the scenario, and it translates plain-English edits into candidate formal statements. The deterministic ABDA engine remains the sole source of arguments, attacks, and acceptance labels, and every proposal of the model is validated and confirmed by the user before it takes effect.

摘要:ABDA-NL 為 ABDA 添加了一個自然語言介面,這是一個基於論點的討論系統,使用 ASPIC 知識庫並基於基礎語義。用戶可以看到哪些結論被接受、拒絕或未決,打開一個互動式的基礎討論遊戲來了解原因,通過暫停假設和規則或改變偏好來探索假設的替代方案,提出問題,這些問題會從情境的參考文件中得到回答,並用簡單的英語創建新的事實、假設和規則。一個大型語言模型提供了語言與形式之間的橋樑:它回答來自文件和當前情境狀態的問題,並將簡單英語的編輯翻譯成候選的正式陳述。確定性的 ABDA 引擎仍然是論點、攻擊和接受標籤的唯一來源,模型的每一個提案在生效之前都需經用戶驗證和確認。

Screw Attention: Rigid-Body Algebra Inside a Transformer

2610.00904v1 by Aly Magassouba

Learned manipulation policies rediscover from data the spatial relations that rigid-body mechanics supplies in closed form. This costs data, and it leaves the policies fragile to geometric changes in the scene. We present Screw Attention, a transformer layer in which the relation between two bodies is a spatial transform rather than a graph edge. Every token is a body with a pose. Each pair of tokens carries the relative pose and, for robot joints, the joint screw. Messages are transported along this relation into the receiver's frame, while the attention scores see only frame-invariant quantities. By construction, the messages are equivariant to an independent change of frame at every token, and a single layer can express the velocity recursion of rigid-body mechanics. On simulated manipulation tasks, Screw Attention matches or outperforms controls of the same size, including graph, transformer and flat networks on LIBERO-Spatial. With 16,162 parameters it reaches 97.3% on LIBERO-Spatial from object poses (without images or language), above a flat network with 27x more parameters. Under a change of per-link frame convention its success is unchanged, while every other learned network falls below 3%. Placed on an analytic controller as a gated residual, it raises insertion success by 17.3 points. It is unaffected by pose noise up to 10,mm and by joint offsets within the factory calibration of a Franka arm. These results suggest a criterion: geometry is decisive when the task requires relations between frames that no other part of the system supplies. Code and trained policies will be released.

摘要:學習的操控策略從數據中重新發現剛體力學以封閉形式提供的空間關係。這需要數據,並且使得策略對場景中的幾何變化變得脆弱。我們提出了螺旋注意力(Screw Attention),這是一個Transformer層,其中兩個物體之間的關係是一種空間變換,而不是圖邊。每個標記都是一個具有姿態的物體。每對標記攜帶相對姿態,對於機器人關節,則是關節螺旋。消息沿著這種關係傳輸到接收者的框架中,而注意力分數僅查看框架不變的量。根據構造,這些消息對每個標記的獨立框架變化是等變的,且單層可以表達剛體力學的速度遞歸。在模擬操控任務中,螺旋注意力的表現與相同大小的控制器相匹配或超過,包括在LIBERO-Spatial上的圖形、Transformer和扁平網絡。擁有16,162個參數的它在LIBERO-Spatial上從物體姿態達到97.3%(不使用圖像或語言),超過了一個擁有27倍參數的扁平網絡。在每個連接的框架約定變化下,它的成功率保持不變,而其他學習的網絡則降至3%以下。作為一個門控殘差放置在分析控制器上,它將插入成功率提高了17.3個點。它對高達10mm的姿態噪聲和在Franka臂的工廠校準範圍內的關節偏移不受影響。這些結果暗示了一個標準:當任務需要框架之間的關係,而系統的其他部分無法提供時,幾何是決定性的。代碼和訓練好的策略將會發布。

Match the Distribution, Not the Compute: Post-Training Multi-Token Prediction Heads

2610.00888v1 by Prachi Badarayani, Aidan Jay, Chenghui Zhou, Dayquan Julienne, Yuan Gao, Tianwei Chen, George Zerveas, Ishmam Zabir, Xiren Zhou, Chris Quirk, Xia Song

Multi-token prediction (MTP) improves the throughput of autoregressive generation by enabling the language model to draft multiple next tokens per forward pass, while a verification step over draft tokens ensures that token distribution of the backbone is preserved. Every open MTP-family release (MiMo-7B, DeepSeek-V3, Qwen3) trains its heads jointly with the backbone over the full pretraining run of tens of trillions of tokens, thus setting the drafter quality at pretraining time. We ask whether a lightweight post-training pass on target-generated chain-of-thought is enough to reach the same expected throughput speedup on a frozen reasoning model, and study how a serving-time system built on such a checkpoint can be optimized. We present three findings. 1) On a frozen Qwen3-8B with $K{=}3$ chained MTP heads, we show that a post-training recipe with plain cross-entropy on $\approx!2.5$B tokens reaches or exceeds the expected speedup of jointly trained MiMo-7B on math, coding and knowledge benchmarks. Our post-training recipe utilizes $10^3$-$10^4\times$ less MTP-training tokens as compared with joint pre-training of MiMO-7B MTP baseline. 2) We propose a chain-aware relaxation of draft token verification rule that allows a bounded drift from backbone language model token distribution. We show that this relaxation lifts expected speedups by $+12$ to $+16\%$ per benchmark while preserving task accuracy. 3) We propose an adaptive controller that dynamically chooses the number of MTP heads to be engaged at inference time and demonstrate recovery of upto $11$--$14\%$ loss in speedup using fixed maximum MTP draft length.

摘要:多標記預測(MTP)透過使語言模型在每次前向傳遞中草擬多個下一個標記來提高自回歸生成的吞吐量,同時對草擬標記進行驗證步驟以確保骨幹的標記分佈得以保留。每個開放的MTP家族版本(MiMo-7B、DeepSeek-V3、Qwen3)在數十萬億標記的完整預訓練過程中,與骨幹共同訓練其頭部,從而在預訓練時設置草擬者的質量。我們詢問在目標生成的思維鏈上進行輕量級的後訓練過程是否足以在凍結的推理模型上達到相同的預期吞吐量加速,並研究基於此檢查點構建的服務時間系統如何進行優化。我們提出三個發現。1)在一個凍結的Qwen3-8B上,使用$K{=}3$鏈式MTP頭,我們顯示一個在$\approx!2.5$B標記上使用普通交叉熵的後訓練配方達到或超過了在數學、編碼和知識基準上共同訓練的MiMo-7B的預期加速。我們的後訓練配方使用的MTP訓練標記比MiMo-7B MTP基準的共同預訓練少了$10^3$-$10^4\times$。2)我們提出了一個鏈式感知的草擬標記驗證規則的放鬆,允許與骨幹語言模型標記分佈的有界漂移。我們顯示這種放鬆使每個基準的預期加速提高了$+12$到$+16\%$,同時保持任務準確性。3)我們提出了一個自適應控制器,動態選擇在推理時參與的MTP頭的數量,並展示使用固定的最大MTP草擬長度恢復高達$11$--$14\%$的加速損失。

Backdoor Purification for LoRA-Tuned LLMs via Null-Space Projection

2610.00685v1 by Jianwei Li, Jung-Eun Kim

With the rapid adoption of large language models (LLMs) and parameter-efficient fine-tuning (PEFT) methods, the risk of backdoor attacks has become more severe. Existing backdoor purification methods typically rely on at least one of the strong assumptions, such as prior knowledge of triggers, access to clean references, or aggressive retraining, and they often lack comprehensive evaluations. These constraints substantially limit their practical applicability. To overcome these challenges, our work proposes purifying LoRA-tuned LLMs without these assumptions and even without post-hoc retraining of the suspect parameters. Our objective is to significantly reduce the attack success rates (ASR) while preserving both (i) the base model's general capabilities and (ii) the new downstream skills learned through the adapter. Through a series of ablation studies, we progressively scale our approach from a single layer in a text classification setting to a full-parameter LLM in the generative task. Through careful data curation and feature approximation, we extract high-fidelity backdoor directions and, for each layer or head, construct orthogonal null spaces in both the input and output channels, onto which the LoRA updates are projected. Empirically, our null-space projection method reduces the ASR from nearly 100% to less than 10%, while preserving the base model's benign performance and the adapter's learned abilities during downstream task adaptation.

摘要:隨著大型語言模型(LLMs)和參數高效微調(PEFT)方法的快速採用,後門攻擊的風險變得更加嚴重。現有的後門淨化方法通常依賴於至少一個強假設,例如對觸發器的先驗知識、訪問乾淨參考資料或激進的再訓練,並且它們往往缺乏全面的評估。這些限制大大限制了它們的實際應用性。為了克服這些挑戰,我們的工作提出了在沒有這些假設的情況下淨化LoRA調整的LLMs,甚至不需要對可疑參數進行事後再訓練。我們的目標是顯著降低攻擊成功率(ASR),同時保留(i)基礎模型的一般能力和(ii)通過適配器學到的新下游技能。通過一系列的消融研究,我們逐步將我們的方法從文本分類設定中的單層擴展到生成任務中的全參數LLM。通過仔細的數據策劃和特徵近似,我們提取高保真度的後門方向,並為每一層或頭構建正交的零空間,這些零空間位於輸入和輸出通道上,LoRA更新將被投影到這些空間中。經驗上,我們的零空間投影方法將ASR從近乎100%降低到不到10%,同時在下游任務適應過程中保留了基礎模型的良性性能和適配器學到的能力。

Ontology-Grounded, Reasoner-Verified Benchmarks for Evaluating LLM Reasoning in Scientific AI

2610.00682v1 by Nishtha N. Vaidya, Stephan Grimm, Thomas Hubauer, Thomas A. Runkler

Large language models (LLMs) increasingly underpin scientific AI applications that reason over structured knowledge, from biomedical question answering to materials informatics. However, their logical reasoning often falls short, producing factual inaccuracies unacceptable in these settings. Reliable evaluation remains challenging: manual dataset construction scales poorly, and LLM-based generation risks embedding the very flaws it aims to measure. High-quality benchmarks must ground both correct and incorrect labelled examples in explicit background knowledge, formally verifiable by a standard reasoner. We propose a pipeline that automatically generates ontology-grounded multiple-choice question (MCQ) benchmarks from any sufficiently axiomatised OWL 2 ontology, with correct answers grounded in the ontology by design. Distractors are generated by perturbing the right-hand-side class expressions of class definition axioms, and their incorrectness is formally verified by an OWL reasoner via entailment checks. We evaluate the pipeline on three ontologies: Pizza (small, academic), PMDco (complex, materials science), and DOID (large, biomedical), generating 112, 2,491, and 15,216 MCQs respectively. Distractors span four semantic categories from class unsatisfiability to weakened subsumptions, enabling diagnostic evaluation of specific reasoning failures. Items meet natural language quality standards: mean LLM judge scores of 4.02, 4.36, and 3.36 out of 5 confirm fluency, and correct-answer-to-distractor similarity above 0.8 shows that wrong options cannot be dismissed on surface form alone. Six LLMs evaluated zero-shot achieve 41.1-76.8% accuracy, well above the 25% random-guessing baseline, indicating the benchmarks are challenging and discriminative. This work is a step towards more reliable benchmarks for assessing logical reasoning in scientific AI.

摘要:大型語言模型(LLMs)越來越多地支撐著科學人工智慧應用,這些應用在結構化知識上進行推理,從生物醫學問答到材料資訊學。然而,它們的邏輯推理經常不夠準確,在這些情境中產生的事實不準確是不可接受的。可靠的評估仍然具有挑戰性:手動數據集構建的擴展性差,而基於LLM的生成則有可能嵌入它所旨在測量的缺陷。高品質的基準必須將正確和不正確的標記示例基於明確的背景知識,並由標準推理器形式驗證。我們提出了一個管道,該管道自動從任何足夠公理化的OWL 2本體生成基於本體的多選題(MCQ)基準,正確答案在設計上基於本體。干擾項是通過擾動類定義公理的右側類表達式生成的,其不正確性通過OWL推理器通過推理檢查形式驗證。我們在三個本體上評估了該管道:Pizza(小型,學術)、PMDco(複雜,材料科學)和DOID(大型,生物醫學),分別生成112、2,491和15,216個MCQ。干擾項涵蓋了四個語義類別,從類不滿足到弱化的子類關係,使得特定推理失敗的診斷評估成為可能。項目符合自然語言質量標準:LLM評審的平均分數為4.02、4.36和3.36(滿分5分),確認了流暢性,且正確答案與干擾項的相似度超過0.8,顯示錯誤選項不能僅僅因表面形式而被忽視。六個評估的零樣本LLM達到41.1-76.8%的準確率,遠高於25%的隨機猜測基準,表明這些基準具有挑戰性和區分性。這項工作是朝著更可靠的基準邁出的一步,以評估科學人工智慧中的邏輯推理。

PhysicsMate: A Curriculum-Grounded Bengali Benchmark for Secondary Physics QA with Small-Model Adaptation

2610.00664v1 by Rashid Azraf Jahin, Saadman Sajid, Khan Raiyan Ibne Reza, Sumaiya Tabassum Nimi

Bengali secondary education lacks curriculum-grounded benchmarks for STEM question-solving, and general-purpose language models struggle with the precise terminology, unit conventions, and derivations that physics problems demand. We introduce PhysicsMate, a benchmark of 1834 question-answer pairs built from the National Curriculum and Textbook Board (NCTB) Grade 9-10 physics syllabus and grounded in a multi-relational knowledge graph of 1760 nodes and 2600 edges across ten ontological types. We Low-Rank Adapt at 0.6B, 1.7B, and 4B parameters, with a unified recipe and demonstrate a significant increase in closed-book accuracy in all scales (+5.5, +15.0, and +23.3 percentage points). A node-type analysis shows that the most benefited by adaptation is the structured curricular knowledge, which consists of physical quantities and named laws, while the least benefited is the loosely specified entity-level knowledge. The 4B model has been adapted and quantized to a small offline binary that can be used for local inference in resource constrained environments and offers a viable path to curriculum aligned physics support in environments with limited connectivity and hardware.

摘要:孟加拉的中學教育缺乏基於課程的STEM問題解決基準,而通用語言模型在物理問題所需的精確術語、單位慣例和推導方面表現不佳。我們介紹了PhysicsMate,這是一個由1834個問題-答案對組成的基準,基於國家課程和教科書委員會(NCTB)9-10年級物理課程大綱,並建立在一個包含1760個節點和2600條邊的多關係知識圖譜上,涵蓋十種本體類型。我們在0.6B、1.7B和4B參數下進行低秩適應,使用統一的配方,並在所有規模上顯示出閉卷準確率的顯著提高(+5.5、+15.0和+23.3個百分點)。節點類型分析顯示,適應中受益最多的是結構化課程知識,這包括物理量和命名法則,而受益最少的是鬆散指定的實體級知識。4B模型已被適應並量化為一個小型離線二進制文件,可用於資源受限環境中的本地推理,並為在連接性和硬體有限的環境中提供課程對齊的物理支持提供了一條可行的途徑。

Where's Waldo? Query-language Preference under Cross-lingual Knowledge Disparities

2610.00606v1 by Dayeon Ki, Ruochen Zhang, Silviu Cucerzan, Ryen W. White, Ning Gao

Large Language Models increasingly serve as interfaces for knowledge-intensive information seeking tasks across languages by synthesizing multilingual evidence. Prior work has shown that they often exhibit query-language preference -- the tendency to favor sources written in the language of the query -- but has largely examined this behavior in settings where equivalent knowledge is available across languages. However, this bias becomes consequential when sources in different languages provide incomplete or inconsistent accounts of the same fact, since the information users receive then depends on the sources a model selects to use. To characterize query-language preference under such cross-lingual knowledge disparities, we introduce Waldo, a multilingual Question-Answering (QA) benchmark constructed from Wikipedia. Waldo contains 12K QA pairs targeting knowledge gaps, where a fact is available in one language but absent in another, and knowledge conflicts, where language editions provide conflicting versions of the same fact. Evaluating eight models across five languages, we find that when one language edition merely lacks the relevant fact, models generally use evidence from the other language regardless of the query language. Under conflicting accounts, however, model responses strongly align with the document in the query language, causing semantically equivalent queries to elicit different accounts depending on the user's language. Finally, we explore two different approaches that could mitigate this preference under knowledge conflicts: a mechanistic intervention that ablates attention heads associated with query-language preference, and LoRA-based training, which reduces the preference gap by up to 61.5%.

摘要:大型語言模型越來越多地作為跨語言知識密集型信息搜尋任務的介面,通過綜合多語言證據來實現。先前的研究顯示,它們通常表現出查詢語言偏好——即偏向於使用以查詢語言撰寫的來源——但這種行為主要是在不同語言之間有等效知識的情況下進行的檢查。然而,當不同語言的來源提供相同事實的不完整或不一致的描述時,這種偏見變得至關重要,因為用戶所接收到的信息取決於模型選擇使用的來源。為了在這種跨語言知識差異下描述查詢語言偏好,我們引入了Waldo,一個基於維基百科構建的多語言問答(QA)基準。Waldo包含12K個針對知識空白的QA對,其中一種語言中有事實而另一種語言中缺失,以及知識衝突,其中語言版本提供相同事實的衝突版本。我們在五種語言中評估了八個模型,發現當一種語言版本僅缺少相關事實時,模型通常會使用來自另一種語言的證據,而不管查詢語言是什麼。然而,在衝突的描述下,模型的回應強烈對應於查詢語言中的文檔,導致語義上等價的查詢根據用戶的語言引出不同的描述。最後,我們探索了兩種不同的方法,可以減輕知識衝突下的這種偏好:一種機械干預,消除與查詢語言偏好相關的注意力頭,另一種基於LoRA的訓練,將偏好差距縮小至61.5%。

Emergent Unfaithfulness: How Alignment Training Causes Language Models to Silently Override Task Faithfulness

2610.00568v1 by Pardis Sadat Zahraei, Janvijay Singh, Gokhan Tur, Dilek Hakkani-Tur

Large language models are characterized by three key properties: capability, alignment, and faithfulness. Prior work studies the tradeoffs between capability and alignment, and between capability and faithfulness, but a third tension remains underexplored: the alignment-faithfulness conflict. We show that aligned models systematically deviate from their inputs on unsafe or sensitive content without disclosing the modification, a failure mode we call alignment-induced unfaithfulness (AIU). Unlike capability-driven unfaithfulness, which comes from errors in knowledge or reasoning, this is induced by post-training mechanisms that override adherence to the input. We introduce FaithConflict, a controlled dataset isolating both conflicts, and two complementary taxonomies: behavioral (B1-B8) and chain-of-thought reasoning (C0-C6). Across models, AIU increases with scale and more sharply than capability-driven unfaithfulness, a reverse scaling law; intermediate checkpoints show it is amplified during post-training, with DPO the stage at which the gap both grows most and becomes least visible. Prompting-based mitigation does not resolve it, revealing a capability-alignment-faithfulness trilemma in the design and evaluation of LLMs.

摘要:大型語言模型的特徵有三個關鍵屬性:能力、對齊和忠實性。先前的研究探討了能力與對齊之間的權衡,以及能力與忠實性之間的權衡,但第三種緊張關係仍然未被充分探討:對齊-忠實性衝突。我們展示了對齊模型在處理不安全或敏感內容時,系統性地偏離其輸入而不披露修改,這種失敗模式我們稱之為對齊引起的不忠實性(AIU)。與能力驅動的不忠實性不同,後者源於知識或推理的錯誤,這種不忠實性是由後訓練機制引起的,這些機制覆蓋了對輸入的遵循。我們引入了FaithConflict,一個控制數據集以隔離這兩種衝突,以及兩個互補的分類法:行為(B1-B8)和思維鏈推理(C0-C6)。在各模型中,AIU隨著規模的增長而增加,並且增長的幅度比能力驅動的不忠實性更為明顯,這是一種反向縮放法則;中間檢查點顯示它在後訓練期間被放大,DPO是這一差距增長最多且變得最不明顯的階段。基於提示的緩解措施無法解決此問題,揭示了大型語言模型設計和評估中的能力-對齊-忠實性三難問題。

Assessing the Impact of Language Disparity on Multilingual Linguistic Ability in Large Language Models

2610.00540v1 by Zhanyu Chen, Jaap Jumelet

Claims about the grammatical competence of multilingual language models vary sharply with how competence is measured, yet the interaction between evaluation paradigm, post-training, and language resource availability has not been systematically examined. We evaluate base and post-trained models from six families on MultiBLiMP, a syntactic minimal-pair benchmark covering 101 languages, using four evaluation methods. We report three principal findings. First, post-training degrades grammatical competence, but the magnitude of this effect is reduced unevenly by model scale, while low-resource languages bear the highest cost. Second, post-trained models retain grammatical knowledge they cannot articulate through explicit prompting, yet this is measurable only in high-resource languages, because near-chance baselines in low-resource settings leave little knowledge to hide. Third, native-language prompting recovers otherwise hidden competence on low-resource languages, demonstrating that only high-resource languages can be probed directly from unprompted probabilities. We conclude that multilingual grammatical evaluation must adopt language-informed, multi-paradigm protocols to avoid systematically underestimating low-resource abilities.

摘要:關於多語言模型的語法能力的主張,隨著能力測量方式的不同而有明顯差異,然而評估範式、後訓練和語言資源可用性之間的互動尚未被系統性地檢視。我們使用四種評估方法,對六個家族的基本模型和後訓練模型在 MultiBLiMP 上進行評估,這是一個涵蓋 101 種語言的句法最小對比基準。我們報告了三個主要發現。首先,後訓練會降低語法能力,但這一影響的程度因模型規模而不均勻地減少,而低資源語言承擔了最高的成本。其次,後訓練模型保留了它們無法通過明確提示表達的語法知識,但這僅在高資源語言中可測量,因為在低資源環境中接近隨機的基線幾乎沒有知識可隱藏。第三,母語提示恢復了在低資源語言上隱藏的能力,顯示只有高資源語言可以直接從未提示的概率中探測。我們總結認為,多語言語法評估必須採用語言知情的多範式協議,以避免系統性低估低資源能力。

EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery

2609.40340v1 by Young-Jun Lee, Jinheon Baek, Soyeong Jeong, Minki Kang, Seungyeon Jwa, Jonghyun Choi, Seungho Han, Dongyeop Kang

Evolutionary search with large language models (LLMs) can stall when progress requires external knowledge the model lacks. Supplying relevant documents helps, but simply adding web search tool can keep returning the same pages as solutions change. We introduce EvoDuet, a bi-level optimization method that co-evolves solutions and search queries with fixed model parameters. At each iteration, a retrieval gate lets the LLM assess its knowledge gap and choose to retrieve new documents, reuse stored ones, or proceed without them. An inner loop refines queries and ranks documents by the solution scores they are predicted to yield; an outer loop generates candidates in parallel from these documents and records the evaluated outcomes for later searches. Across 21 optimization tasks with one candidate per iteration, EvoDuet raises OpenEvolve's normalized discovery gain from 74.1% to 78.0% with GPT-5.6-Luna and from 61.3% to 82.3% with Gemini-3.8-Flash, whereas Qwen3.5-9B does not benefit. Our best runs surpass the previously reported best scores on eight tasks, including Swap Reduction on Q20 and Rosetta, and match them on three more. EvoDuet also improves with other scaffolds (e.g., Top-K, EvoX) on Sums/Diffs and Denoising, demonstrating its applicability across evolutionary search scaffolds.

摘要:進化搜尋與大型語言模型(LLMs)結合時,當進展需要模型缺乏的外部知識時,可能會停滯不前。提供相關文件有助於改善情況,但僅僅添加網路搜尋工具可能會在解決方案變化時持續返回相同的頁面。我們介紹EvoDuet,一種雙層優化方法,通過固定的模型參數共同進化解決方案和搜尋查詢。在每次迭代中,檢索閘讓LLM評估其知識差距,並選擇檢索新文件、重用存儲的文件或在沒有它們的情況下繼續。內部循環精煉查詢並根據預測產生的解決方案分數對文件進行排名;外部循環則從這些文件中平行生成候選項,並記錄評估結果以便後續搜尋。在21個優化任務中,每次迭代一個候選項,EvoDuet使OpenEvolve的標準化發現增益從74.1%提高到78.0%(使用GPT-5.6-Luna),並從61.3%提高到82.3%(使用Gemini-3.8-Flash),而Qwen3.5-9B則沒有受益。我們的最佳運行超過了先前報告的八個任務的最佳分數,包括Q20和Rosetta上的Swap Reduction,並在另外三個任務上達到相同的分數。EvoDuet在其他支架(例如Top-K、EvoX)上也在Sums/Diffs和去噪中有所改善,顯示其在進化搜尋支架中的適用性。

Linguistic Loopholes in LLM Unlearning: From a 174-Language Benchmark to Coverage-Aware Unlearning

2609.40286v1 by Tyler Skow, Shravan Chaudhari, Rama Chellappa, Abhay Yadav

Unlearning a fact in one language does not guarantee its removal in others as changing the query or even the requested answer language can reopen seemingly forgotten knowledge -- a cross-lingual loophole. The most straightforward solution to this challenge -- unlearning in all languages -- is neither scalable nor desirable as it amplifies damage to unrelated model capabilities. We introduce the task of language budgeted multilingual unlearning where the goal is to select a subset of languages that maximizes cross-lingual erasure. To study this task we introduce the Cross-Lingual Unlearning Tensor, an unlearning benchmark that spans 174 language--script pairs and 25 atomic paraphrase types to examine when forgetting generalizes across linguistic expressions of the same knowledge. We further propose COVER, which selects source languages to maximize predicted COVERage of languages receiving no forget supervision, enabling unlearning on a language budget. Surprisingly, we find naively selecting strong individual sources does not reliably compose into strong source sets motivating our development of COVER. At deployment COVER only requires benign calibration data and access to the frozen model. Across three model families and two disjoint forget sets, COVER reduces mean held-out residual access by 7.8--27.3% relative to uniform source selection. We find these gains extend beyond synthetic benchmarks to real news documents in low-resource language settings using human translated data from the Low Resource Languages for Emergent Incidents (LORELEI) corpus.

摘要:在一種語言中忘記一個事實並不保證它在其他語言中也會被移除,因為改變查詢或甚至請求的答案語言可能會重新打開看似被遺忘的知識——這是一種跨語言的漏洞。解決這一挑戰的最直接方案——在所有語言中忘記——既不可擴展也不可取,因為這會加劇對無關模型能力的損害。我們引入了語言預算多語言忘記的任務,其目標是選擇一組語言,以最大化跨語言的抹去。為了研究這一任務,我們引入了跨語言忘記張量,這是一個涵蓋174種語言-書寫對和25種原子改述類型的忘記基準,以檢查何時忘記在相同知識的語言表達中會泛化。我們進一步提出了COVER,它選擇源語言以最大化未接受忘記監督的語言的預測COVERage,從而實現語言預算下的忘記。令人驚訝的是,我們發現天真地選擇強大的個別源並不可靠地組成強大的源集合,這促使我們開發COVER。在部署時,COVER僅需要良性的校準數據和對凍結模型的訪問。在三個模型系列和兩個不相交的忘記集上,COVER相對於均勻源選擇減少了7.8-27.3%的平均保留殘餘訪問。我們發現這些增益超越了合成基準,擴展到使用來自低資源語言緊急事件(LORELEI)語料庫的人類翻譯數據的真實新聞文件。

EviRover: Reinforcing Agentic Perception Beyond a Glance

2609.40230v1 by Kaixuan Fan, Kaituo Feng, Tianshuo Peng, Yilei Jiang, Manyuan Zhang, Junke Wang, Xiangyu Yue

Visual perception is conventionally formulated as a one-shot prediction from a single glance at the image, under the assumption that the image content and the model's parametric knowledge suffice to resolve the query. This assumption often fails in real-world scenarios that hinge on fine-grained visual details or require knowledge-intensive and up-to-date information. We term such cases \textit{perception under insufficient evidence} and formulate perception as an agentic process that can obtain information beyond a single glance. To address the absence of data for this setting, we design two dedicated data generation pipelines, yielding EviRover-SFT-5K and EviRover-RL-12K for training. We further construct EviLens, a human-verified benchmark comprising 688 instances across five perception categories. Building on these data, we present EviRover, to our knowledge the first perception agent explicitly trained to resolve perceptual queries through interaction, using supervised fine-tuning followed by agentic reinforcement learning. Experiments show that the 4B EviRover outperforms its backbone by 30 points on average on EviLens, reaching performance comparable to advanced proprietary models. The gains transfer beyond EviLens to WebEyes, conventional perception benchmarks, and general multimodal benchmarks, including a 15-point improvement on BrowseComp-VL. All code, models, and data are released.

摘要:視覺感知通常被表述為從對圖像的一瞥中進行的一次性預測,假設圖像內容和模型的參數知識足以解決該查詢。這一假設在依賴細微視覺細節或需要知識密集和最新信息的現實場景中經常失效。我們將這類情況稱為\textit{在不足證據下的感知},並將感知表述為一種能夠獲取超越單一瞥見的信息的主動過程。為了解決這一情境下數據的缺乏,我們設計了兩個專門的數據生成管道,產生了用於訓練的EviRover-SFT-5K和EviRover-RL-12K。我們進一步構建了EviLens,一個經人類驗證的基準,包含五個感知類別的688個實例。在這些數據的基礎上,我們提出了EviRover,據我們所知,這是第一個明確訓練以通過互動解決感知查詢的感知代理,使用監督微調隨後進行主動強化學習。實驗顯示,4B的EviRover在EviLens上的表現平均比其基礎模型高出30分,達到與先進專有模型相當的性能。這些增益超越EviLens,轉移到WebEyes、傳統感知基準和一般多模態基準,包括在BrowseComp-VL上提高15分。所有代碼、模型和數據均已發布。

Learning from Research: Toward Lifelong Agent Harness Evolution

2609.40169v1 by Jingbo Yang, Kwei-Herng Lai, Xiaowen Wang, Yaar Harari, Evgeniy Gabrilovich, Shiyu Chang

Language agents are expected to solve increasingly complex tasks, creating a growing need for continual improvement. One promising approach is to evolve the agent harness, the software that governs tool use, memory management, and task execution, while keeping the underlying language model fixed. Recent methods automate this process by using a meta coding agent to modify the harness based on execution feedback. However, relying on that agent's existing knowledge and observed failures can restrict exploration and make adaptation reactive. Inspired by how human experts learn from the research literature for new solutions, we introduce ScholarEvolve, a framework that automatically draws on state-of-the-art research to guide harness evolution. ScholarEvolve organizes the harness evolution directions into functional modules and uses topic modeling to identify distinct improvement strategies for each module. It implements these strategies and evaluates their combinations to improve task performance. Moreover, the framework is designed to incorporate new publications over time, allowing research advances to drive proactive lifelong evolution. Experiments demonstrate improvements on AppWorld and Tau2-Bench. ScholarEvolve raises Qwen3.5-27B task goal completion from 49.6% to 63.6% on AppWorld Challenge, and raises GPT-5.4-mini pass@1 from 72.7% to 81.9% on Tau2-Bench Telecom.

摘要:語言代理預期能解決日益複雜的任務,這創造了持續改進的需求。一種有前景的方法是進化代理工具,即管理工具使用、記憶管理和任務執行的軟體,同時保持基礎語言模型不變。最近的方法通過使用元編碼代理自動化這一過程,根據執行反饋來修改工具。然而,依賴該代理的現有知識和觀察到的失敗可能會限制探索,並使適應變得被動。受到人類專家如何從研究文獻中學習新解決方案的啟發,我們引入了ScholarEvolve,一個自動利用最先進研究來指導工具進化的框架。ScholarEvolve將工具進化方向組織成功能模組,並使用主題建模來識別每個模組的不同改進策略。它實施這些策略並評估它們的組合以改善任務性能。此外,該框架設計為隨著時間的推移納入新出版物,允許研究進展推動主動的終身進化。實驗顯示在AppWorld和Tau2-Bench上有所改善。ScholarEvolve將Qwen3.5-27B在AppWorld Challenge上的任務目標完成率從49.6%提高到63.6%,並將GPT-5.4-mini在Tau2-Bench Telecom上的pass@1從72.7%提高到81.9%。

On the (In)effectiveness of AMR Augmentation for Large Language Models

2609.40121v1 by Hoa Quynh Nhung Nguyen, Jacopo Staiano, Michael Sullivan

While Abstract Meaning Representation (AMR) has historically improved performance on a range of NLP tasks, the benefit---or lack thereof---of AMR augmentation for modern LLMs is thus far unclear. In this paper, we attempt to reproduce recent work that reported substantial downstream gains from AMR augmentation, finding that these are likely due to specific choices in the experimental settings used: using a consistent and unified protocol for hyperparameter selection, we observe that text-only baselines consistently match or exceed the performance of AMR-augmented models. To investigate this null result, we introduce a perplexity-based probe measuring the degree to which AMR provides an LLM with supplemental relational knowledge not already available to the model. We find that AMR augmentation does not help LLMs improve their understanding of relational content in the sentence, indicating that augmenting these models with AMR offers no clear benefit on downstream tasks.

摘要:雖然抽象意義表示法(AMR)歷來在多種自然語言處理任務中提高了性能,但AMR增強對現代大型語言模型的好處——或缺乏好處——至今仍不明確。在本文中,我們試圖重現最近的研究,該研究報告了AMR增強帶來的顯著下游收益,發現這些收益可能是由於實驗設置中的特定選擇:使用一致且統一的超參數選擇協議,我們觀察到僅使用文本的基準模型在性能上始終與AMR增強模型相匹配或超過。為了調查這一無效結果,我們引入了一種基於困惑度的探測器,測量AMR為大型語言模型提供額外關聯知識的程度,而這些知識在模型中並不存在。我們發現AMR增強並未幫助大型語言模型改善對句子中關聯內容的理解,這表明用AMR增強這些模型在下游任務中並未提供明顯的好處。

Persistent Context Graphs for Efficient Memory Compaction in LLM Agents

2609.40118v1 by Jingbo Yang, Kwei-Herng Lai, Xiaowen Wang, Zhaoxuan Tan, Pei Zhou, Mengting Wan, Yaar Harari, Evgeniy Gabrilovich, Shiyu Chang

As LLM capabilities advance, agents are tackling increasingly complex tasks over longer horizons. Their growing interaction histories make memory compaction essential for staying within context windows and reducing prefill cost. Existing methods summarize the history or compress its KV cache, often adding model computation to preserve information for future requests. A new user request can change which history matters, but reassessing that history with the model requires re-encoding it if the KV cache has expired. Past attention provides signals of historical importance and dependencies between messages, while relevance to the current task must be assessed using the new user request. We introduce ReCAP, a memory compaction method that stores attention-derived importance scores and dependency links in a lightweight, persistent context graph. For each new request, ReCAP combines stored importance with relevance cues from the request and follows dependency links to select messages and their supporting context, without additional model calls for selection. Compared with Codex's default summarization-based compaction, ReCAP reduces estimated latency for compaction and cold restoration by approximately 95% on both Qwen3-Coder and gpt-oss. It also roughly halves the historical context per call on SWE-Together at comparable task quality and improves accuracy on the code tasks of Lost-in-Conversation over full history by 19.8 and 41.2 points.

摘要:隨著大型語言模型(LLM)能力的提升,代理正在處理越來越複雜的任務,並且時間範圍也越來越長。它們日益增長的互動歷史使得記憶壓縮對於保持在上下文窗口內和降低預填成本變得至關重要。現有的方法總結歷史或壓縮其KV快取,通常會增加模型計算以保留未來請求的信息。新的用戶請求可能會改變重要的歷史,但如果KV快取已過期,則需要重新編碼該歷史以重新評估它。過去的注意力提供了歷史重要性和消息之間依賴性的信號,而與當前任務的相關性必須使用新的用戶請求來評估。我們介紹了ReCAP,一種記憶壓縮方法,將基於注意力的衍生重要性分數和依賴鏈存儲在輕量級的持久上下文圖中。對於每個新的請求,ReCAP將存儲的重點與請求中的相關提示結合,並沿著依賴鏈選擇消息及其支持上下文,而無需額外的模型調用來進行選擇。與Codex的默認基於總結的壓縮相比,ReCAP在Qwen3-Coder和gpt-oss上將壓縮和冷恢復的估計延遲減少了約95%。它還在SWE-Together上將每次調用的歷史上下文大約減半,並在任務質量相當的情況下,將Lost-in-Conversation的代碼任務準確性提高了19.8和41.2個點。

JuryFlow: Disagreement-Guided Human-in-the-Loop Multi-Agent Evaluation

2609.40103v1 by Mufeng Yang, Junwei Yu, Yepeng Ding

Large language models (LLMs) are increasingly deployed as automated judges for AI-generated content, yet a single judge is unreliable and even a panel of judges leaves a hard residue: when judges disagree, majority voting discards the conflict instead of resolving it. We present JuryFlow, a disagreement-guided, human-in-the-loop multi-agent evaluation framework that treats inter-judge disagreement not as noise to be averaged away, but as a precise, claim-level signal indicating where an evaluation is uncertain. JuryFlow decomposes each candidate response into atomic claims, has a panel of heterogeneous judges assign per-claim verdicts, and builds a disagreement graph whose nodes are scored by verdict entropy and whose edges encode structural similarity between claims. A human acts as a structural guide, selecting which disagreement to resolve through a single, minimal intervention rather than re-labeling the response, after which the focal claim is re-evaluated, the correction propagates along graph edges and to historically similar cases, and is crystallized into reusable rubric entries that all judges inherit, making the evaluator progressively self-refining. To enable large-scale, reproducible benchmarking without human studies, we evaluate JuryFlow in an automatic configuration in which focal selection is made by entropy ranking. On MT-Bench and LLMBar, JuryFlow improves agreement with gold labels over single-judge and majority-vote panel baselines, and ablations isolate the contributions of disagreement-targeted re-evaluation, propagation, and rubric induction. We contribute (1) a human-in-the-loop paradigm that recasts the human from labeler to structural guide, (2) the JuryFlow framework operationalizing it through a disagreement graph, focal re-evaluation, and closed-loop rubric induction, and (3) an evaluation protocol with ablations that isolate where the gains originate.

摘要:大型語言模型(LLMs)越來越多地被用作自動評判AI生成內容的法官,但單一法官不可靠,即使是一組法官也會留下難以解決的問題:當法官意見不合時,多數投票會忽略衝突,而不是解決它。我們提出了JuryFlow,一個以不一致為指導的、人機協作的多代理評估框架,它將法官之間的不一致視為一種精確的、聲明級別的信號,指示評估的不確定性,而不是要被平均掉的噪音。JuryFlow將每個候選回應分解為原子聲明,讓一組異質法官對每個聲明給予裁決,並構建一個不一致圖,其節點由裁決熵進行評分,邊則編碼聲明之間的結構相似性。一名人類作為結構指導,選擇通過單一的最小干預來解決哪個不一致,而不是重新標記回應,之後重新評估焦點聲明,修正沿著圖邊和歷史上相似的案例進行傳播,並被凝結成所有法官繼承的可重用評分標準條目,使評估者逐漸自我精煉。為了實現大規模、可重複的基準測試而不需要人類研究,我們在自動配置中評估JuryFlow,其中焦點選擇由熵排名決定。在MT-Bench和LLMBar上,JuryFlow提高了與金標籤的協議,相較於單一法官和多數投票小組基準,並且消融實驗隔離了針對不一致的重新評估、傳播和評分標準引入的貢獻。我們貢獻了(1)一個人機協作的範式,將人類從標記者重新塑造為結構指導,(2)通過不一致圖、焦點重新評估和閉環評分標準引入來實現這一範式的JuryFlow框架,以及(3)一個評估協議,通過消融實驗隔離增益的來源。

AutoDataBench: A Data-centric Testbed for Accelerating Auto Research

2609.40097v1 by Ruifeng Yuan, Yizhi Li, Yaxin Du, Fengyu Cai, Yiqi Liu, Hou Pong Chan, Chenghua Lin, Yun Chen, Jian Yang, Bryan Dai, Pinyan Lu, Chenghao Xiao

Existing auto-research benchmarks often entangle multiple sources of improvement, including training frameworks, hyperparameters, compute budgets, and data, making it difficult to attribute why one frontier agent outperforms another to specific research capabilities. In this work, we isolate and systematically evaluate Data Intelligence: an agent's ability to understand, manipulate, and improve the data that shapes model capabilities. We introduce AutoDataBench, a controlled testbed built on a conceptual framework of data intelligence spanning data diagnosis, data organization, and data construction, instantiated through three highly curated optimization tasks while holding non-data factors fixed. Across tool use, retrieval, and knowledge injection, we evaluate frontier LLMs' ability to improve training data through iterative experimentation under task-specific resource budgets. Beyond optimization performance, we ask: do LLMs understand what their data interventions do? We compare predictions made before training with observed outcomes to seek evidence of data-effect reasoning beyond trial and error, and explore whether iterative feedback helps LLMs better understand how changes to training data affect model performance. Finally, we show that reusing AutoDataBench trajectories for mid-training improves downstream coding performance, highlighting its value in both evaluating data intelligence and generating high-quality training data. Code and resources are available at https://github.com/AutoDataBench/AutoDataBench.

摘要:現有的自動研究基準經常將多個改進來源糾纏在一起,包括訓練框架、超參數、計算預算和數據,這使得難以將一個前沿代理的優越表現歸因於特定的研究能力。在這項工作中,我們孤立並系統地評估數據智能:一個代理理解、操作和改善塑造模型能力的數據的能力。我們介紹了AutoDataBench,一個基於數據智能概念框架的受控測試平台,涵蓋數據診斷、數據組織和數據構建,通過三個高度策劃的優化任務實現,同時保持非數據因素不變。在工具使用、檢索和知識注入方面,我們評估前沿LLM在特定任務資源預算下通過迭代實驗改善訓練數據的能力。除了優化性能,我們還問:LLM是否理解它們的數據干預所做的事情?我們比較訓練前的預測與觀察到的結果,以尋求超越試錯的數據效果推理證據,並探索迭代反饋是否有助於LLM更好地理解訓練數據變化如何影響模型性能。最後,我們展示了在中期訓練中重用AutoDataBench軌跡改善下游編碼性能,突顯了其在評估數據智能和生成高質量訓練數據方面的價值。代碼和資源可在 https://github.com/AutoDataBench/AutoDataBench 獲得。

LongEmo: Towards Emotion Understanding and Reasoning in Long Videos

2609.40079v1 by Shuo Zhang, Yifan Zhou, Han Wang, Jinsong Zhang, Jingyu Li, Hongbing Li, Zhejun Zhang, Chengyi Zhao, Yuquan Hao, Yitong Liu, Jiyin Li, Ruiqi Tang, Zixuan Lin, Yi Luo, Xurui Zhang, Ronghao Chen, Huacan Wang, Lei Li

While recent Multimodal Large Language Models (MLLMs) have shown promise in affective computing, their reasoning capabilities are largely confined to short video clips with limited interactions. However, real-world emotions are not merely isolated instantaneous reactions but dynamic and cumulative processes deeply shaped by past experiences and ongoing events. To bridge this gap, we introduce LongEmoBench, a benchmark dedicated to emotion understanding and reasoning in long videos. It assesses progressive capabilities scaling from continuous scene interactions to complex episodic developments. Furthermore, we propose LongEmo, a novel memory-augmented agentic framework designed to tackle the immense challenges of long-range affective reasoning. LongEmo processes continuous video streams to construct an Event Memory Graph, explicitly modeling long-range dependencies and capturing emotional dynamics across discrete events. Given a question, the agent retrieves a query-relevant event stream from the graph, iteratively integrating multimodal memories and relational dependencies to deduce the final answer. Extensive evaluations of 17 representative methods reveal that they struggle significantly with emotion understanding and reasoning in long videos. In contrast, LongEmo achieves state-of-the-art performance, demonstrating the efficacy of its event-centric memory architecture.

摘要:最近的多模態大型語言模型(MLLMs)在情感計算方面顯示出潛力,但它們的推理能力主要限於短視頻片段,互動性有限。然而,現實世界中的情感並不僅僅是孤立的瞬時反應,而是深受過去經驗和當前事件影響的動態和累積過程。為了縮小這一差距,我們引入了LongEmoBench,一個專注於長視頻中的情感理解和推理的基準。它評估從連續場景互動到複雜情節發展的逐步能力。此外,我們提出了LongEmo,一個新穎的增強記憶的代理框架,旨在應對長期情感推理的巨大挑戰。LongEmo處理連續視頻流以構建事件記憶圖,明確建模長期依賴關係並捕捉離散事件中的情感動態。在給定問題的情況下,代理從圖中檢索與查詢相關的事件流,迭代整合多模態記憶和關係依賴,以推導最終答案。對17種代表性方法的廣泛評估顯示,它們在長視頻中的情感理解和推理方面面臨重大挑戰。相比之下,LongEmo實現了最先進的性能,展示了其以事件為中心的記憶架構的有效性。

TACTIC: Temporal and Context-Aware LLM Tactical Planning for Roadside LiDAR Attacks

2609.39969v1 by Yiming Gao, Shaocheng Luo

Physical LiDAR attacks are often evaluated using fixed primitives and manually selected parameters, despite their strong dependence on surrounding traffic. We present TACTIC, a scene-aware framework that uses a multimodal large language model (MLLM) to coordinate state-adaptive roadside LiDAR attacks. Under a gray-box threat model, TACTIC relies only on an attacker-operated roadside perception stack, without accessing the victim LiDAR's native point clouds or internal processing. Local perception provides metric vehicle states, while the MLLM combines these measurements with roadside imagery to infer relational traffic context and construct a semantic scene graph. Based on this representation, TACTIC selects and configures two complementary primitives: \emph{push-away}, which shifts the perceived range of a lead vehicle, and \emph{phantom-obstacle braking}, which triggers emergency braking through obstacle injection. Measured traffic states and empirically calibrated constraints ground the generated tactics in physically feasible operating regions. To accommodate MLLM latency, TACTIC overlaps reasoning and execution asynchronously while high-rate local perception detects scene changes and triggers replanning. Across 280 randomized CARLA trials, the full policy achieves a 100% collision rate, compared with 35% for a fixed rule, 60% for random selection, and 75% for a restricted LLM using mode selection with default parameters. Joint physical-and-image input achieves 100% success, versus 65% with physical measurements alone and 75% with imagery alone, while asynchronous $Δ$ refresh reduces scene-mutation response from 7.4 s to 2.0 s. These results show that scene-dependent tactical planning can expose context-sensitive LiDAR failure modes that fixed attack policies may miss.

摘要:物理LiDAR攻擊通常使用固定的原始元素和手動選擇的參數進行評估,儘管它們強烈依賴於周圍的交通。我們提出了TACTIC,一個場景感知框架,利用多模態大型語言模型(MLLM)來協調狀態自適應的路邊LiDAR攻擊。在灰盒威脅模型下,TACTIC僅依賴攻擊者操作的路邊感知堆疊,而不訪問受害者LiDAR的原始點雲或內部處理。當地感知提供度量車輛狀態,而MLLM將這些測量與路邊影像結合,以推斷關聯交通上下文並構建語義場景圖。基於這一表示,TACTIC選擇並配置兩個互補的原始元素:\emph{推開},它改變前方車輛的感知範圍,以及\emph{幻影障礙物制動},它通過障礙物注入觸發緊急制動。測量的交通狀態和經驗校準的約束將生成的戰術基於物理可行的操作區域。為了適應MLLM的延遲,TACTIC在高頻率的當地感知檢測場景變化並觸發重新規劃的同時,異步重疊推理和執行。在280次隨機化的CARLA試驗中,完整策略實現了100%的碰撞率,而固定規則為35%,隨機選擇為60%,使用默認參數的受限LLM的模式選擇為75%。聯合物理和影像輸入達到100%的成功率,而僅使用物理測量的成功率為65%,僅使用影像的成功率為75%,同時異步$Δ$刷新將場景變異響應從7.4秒減少到2.0秒。這些結果顯示,依賴場景的戰術規劃可以揭示固定攻擊政策可能忽略的上下文敏感LiDAR失效模式。

AIMS: An Agentic AI Framework for Sim-to-Real Multi-Modal ISAC

2609.39964v2 by Yijie Bian, Kai Zhang, Wei Guo, Zixin Wang, Shenghui Song, Jun Zhang, Khaled B. Letaief

Multi-modal integrated sensing and communication (ISAC) enables environmental perception and reliable connectivity for intelligent wireless networks. Data-driven multi-modal ISAC models depend heavily on annotated real-world data to learn relationships across sensing and wireless observations, thereby constraining scalable deployment. Although synthetic data generation reduces the burden, adapting existing simulation pipelines to a target deployment requires consistent scene, sensing, wireless, and learning configurations, while mismatches among these coupled components impair sim-to-real transferability. To address the challenge, we propose an agentic artificial intelligence (AI) framework for sim-to-real multi-modal ISAC, named AIMS. Given a natural-language deployment request specifying the target task, deployment conditions, and real-data budget, AIMS derives a deployment-specific sim-to-real configuration and coordinates its execution to produce a deployment-specific task model. A two-agent architecture coordinates scene construction with task learning. A scene construction agent generates geographically grounded, synchronized sensing and wireless records from shared physical states, while a scene understanding agent configures task-relevant modalities and mixture-of-experts (MoE) learning for zero-shot inference or few-shot adaptation. Structured domain knowledge guides dependency-aware planning, while validation evidence supports feedback-driven revision of affected decisions. Experiments on the real-world DeepSense 6G dataset demonstrate improved vehicle detection and beam prediction over the considered simulation and fusion baselines. A separate orchestration benchmark evaluates task interpretation, dependency reasoning, and feedback-driven replanning across diverse deployment requests, showing improved plan correctness with structured domain knowledge and validation feedback.

摘要:多模態整合感知與通信(ISAC)使智能無線網絡能夠進行環境感知和可靠連接。基於數據驅動的多模態ISAC模型在學習感知和無線觀測之間的關係時,重度依賴標註的真實世界數據,從而限制了可擴展的部署。儘管合成數據生成減輕了負擔,但將現有的模擬管道適應於目標部署需要一致的場景、感知、無線和學習配置,而這些耦合組件之間的錯配會損害模擬到現實的可轉移性。為了解決這一挑戰,我們提出了一個名為AIMS的代理人工智能(AI)框架,用於模擬到現實的多模態ISAC。給定一個自然語言的部署請求,具體說明目標任務、部署條件和真實數據預算,AIMS推導出一個特定於部署的模擬到現實配置,並協調其執行以生成特定於部署的任務模型。兩個代理架構協調場景構建與任務學習。一個場景構建代理從共享的物理狀態生成地理基礎的、同步的感知和無線記錄,而場景理解代理則配置與任務相關的模態和專家混合(MoE)學習,以進行零樣本推理或少樣本適應。結構化的領域知識指導依賴意識的規劃,而驗證證據支持受影響決策的反饋驅動修訂。在現實世界的DeepSense 6G數據集上的實驗顯示,與考慮的模擬和融合基準相比,車輛檢測和波束預測有所改善。一個單獨的協調基準評估任務解釋、依賴推理和跨多樣化部署請求的反饋驅動重新規劃,顯示出結構化的領域知識和驗證反饋提高了計劃的正確性。

Learning to Cover Locally: Graph Neural Combinatorial Optimization under a Hard Information Horizon

2610.00422v1 by Johannes F. Loevenich, Thies Moehlenhof, Laurin Holz, Maxime Schwarzer, Tobias Huerten, Roberto Rigolin F. Lopes

Neural combinatorial optimization typically assumes a centralized solver that reads the whole instance. We study the opposite: combinatorial optimization under a hard information horizon, where every node commits to its share of a global solution seeing only its $k$-hop neighborhood, and those commitments must compose into a globally feasible solution. We formalize this as local set cover and instantiate it on weighted multipoint relay (MPR) selection, the NP-hard 2-hop covering problem of the Optimized Link State Routing Protocol version 2 (OLSRv2) routing protocol (RFC~7181), whose horizon is imposed by the protocol, not chosen by the modeler. We prove two results. Any deterministic selector whose horizon is one hop short must either fail coverage or land a factor $Δ$ from optimal, and an $L$-layer graph neural network (GNN) read out at the deciding node is exactly an $L$-hop selector, so capacity cannot buy back radius. Conversely, at the horizon a \ac{GNN} of depth $O(Δ)$ reproduces the RFC~7181 covering greedy, and at width $O(c_{\max}Δ)$ its metric-aware weighted analogue, inheriting the $(1+\lnΔ_2)$-approximation in both cases. Empirically, a 3-layer \ac{GATv2} with a coverage-completing decoder, behavior-cloned from the CP-SAT optimum, reaches $\text{cost}/\text{opt}=1.030\pm0.001$ against greedy's $1.138$, closing $79.1\%$ of the gap at $100\%$ coverage. Restricting the same learner to one hop, on identical instances with the same decoder and demonstrations, collapses it to $1.344$, far worse than greedy. Two transfer checks target real-world networks. OLSRv2's unmodified selection code matches our cardinality greedy on $200/200$ unit-cost instances, and on $40{,}308$ instances of real battalion mobility the frozen model closes $48\%$ of the gap at full coverage. The information horizon, not the model capacity, is the most significant variable.

摘要:神經組合優化通常假設有一個集中式解算器,能夠讀取整個實例。我們研究相反的情況:在硬信息視野下的組合優化,其中每個節點僅能看到其 $k$ 跳鄰域,並承諾其在全局解中的份額,而這些承諾必須組合成一個全局可行的解。我們將其形式化為局部集合覆蓋,並在加權多點中繼(MPR)選擇上實例化,這是優化鏈路狀態路由協議版本 2(OLSRv2)路由協議(RFC~7181)的 NP 困難 2 跳覆蓋問題,其視野由協議強加,而非由建模者選擇。我們證明了兩個結果。任何其視野短一跳的確定性選擇器必須要麼失敗於覆蓋,要麼與最佳解相差一個因子 $Δ$,而在決策節點讀出的 $L$ 層圖神經網絡(GNN)恰好是一個 $L$ 跳選擇器,因此容量無法贖回半徑。相反,在視野內,深度為 $O(Δ)$ 的 \ac{GNN} 重現了 RFC~7181 的貪婪覆蓋,而在寬度為 $O(c_{\max}Δ)$ 時,其度量感知的加權類比也繼承了這兩種情況下的 $(1+\lnΔ_2)$ 近似。實證上,一個 3 層的 \ac{GATv2} 配備了一個完成覆蓋的解碼器,從 CP-SAT 最優解行為克隆而來,達到了 $\text{cost}/\text{opt}=1.030\pm0.001$,而貪婪的為 $1.138$,在 $100\%$ 覆蓋時縮小了 $79.1\%$ 的差距。將相同的學習者限制為一跳,在相同的實例中使用相同的解碼器和示範,則崩潰至 $1.344$,遠比貪婪差。兩個轉移檢查針對現實世界的網絡。OLSRv2 的未修改選擇代碼在 $200/200$ 單位成本實例中與我們的基數貪婪相匹配,而在 $40{,}308$ 個真實營移動實例中,凍結模型在完全覆蓋時縮小了 $48\%$ 的差距。信息視野,而非模型容量,是最重要的變量。

MCD: Causal Distillation of Multimodal In-Context Learning in Large Vision-Language Models

2609.39920v1 by Yanshu Li, Jiaqian Li, Canran Xiao, Xi Xiao, Tianyang Wang, Yongtai Liu

Large vision-language models (LVLMs) exhibit strong multimodal in-context learning (ICL) capabilities, yet this ability degrades substantially as model size decreases. Knowledge distillation offers a natural way to bridge this gap, but existing methods primarily align output distributions or hidden representations directly. Such alignment teaches the student what the teacher predicts without revealing which evidence in the complex context causally supports that prediction. Consequently, a student can imitate the teacher's answer while continuing to rely on language priors, prompt structure, or other spurious cues. To address this limitation, we introduce Multimodal Causal Distillation (MCD), a distillation framework that transfers how a strong teacher uses multimodal evidence during ICL. MCD uses structure-preserving token interventions to identify and verify causal evidence, then transfers how the teacher responds when that evidence is retained or removed. This design connects distillation to the causal patterns by which the model uses contextual evidence during multimodal ICL. Experiments across three LVLM families and seven benchmarks show that MCD improves student performance by 7.23 points on average and outperforms vanilla distillation by 4.68 points, while further analyses confirm the generalizability of these gains.

摘要:大型視覺-語言模型(LVLMs)展現出強大的多模態上下文學習(ICL)能力,但隨著模型大小的減少,這種能力會顯著下降。知識蒸餾提供了一種自然的方式來彌補這一差距,但現有的方法主要是直接對齊輸出分佈或隱藏表示。這種對齊教會學生教師的預測,而不揭示在複雜上下文中哪些證據因果地支持該預測。因此,學生可以模仿教師的答案,同時繼續依賴語言先驗、提示結構或其他虛假線索。為了解決這一限制,我們引入了多模態因果蒸餾(MCD),這是一種蒸餾框架,轉移強教師在ICL過程中如何使用多模態證據。MCD使用結構保持的標記干預來識別和驗證因果證據,然後轉移教師在保留或移除該證據時的反應。這一設計將蒸餾與模型在多模態ICL過程中使用上下文證據的因果模式聯繫起來。對三個LVLM家族和七個基準的實驗顯示,MCD平均提高了學生的表現7.23分,並且比普通蒸餾提高了4.68分,同時進一步分析確認了這些增益的可泛化性。

The Concrete-Arbitrary Gap: Kinship Reasoning in LLMs Is Not Indifferent to Presentation

2609.39913v1 by Thomas Pashby

We test whether large language models solve formally matched kinship problems equally well when relations are expressed in familiar vocabulary or by explicitly defined nonce predicates. Across 500 paired graphs, concrete accuracy exceeds arbitrary accuracy by 35.6 percentage points in local Qwen3.8-27B, 26.6 in Gemma 4 26B-A4B, 12.0 in Gemma 4 31B, and 5.4 in Qwen3.8-Max. All four paired gaps are statistically resolved. Reasoning budgets and prompt-language interventions can substantially reduce the difference, showing that it is modifiable rather than a fixed incapacity. The minimal conclusion is behavioral: on these tasks, the models' manifested relational competence is not indifferent to presentation. Explicit definitions provide the formal relations but do not make nonce predicates as usable as familiar vocabulary embedded in learned linguistic associations.

摘要:我們測試大型語言模型在關係以熟悉詞彙或明確定義的臨時謂詞表達時,是否能同樣有效地解決正式匹配的親屬關係問題。在500對圖形中,具體準確度在本地 Qwen3.8-27B 中超過任意準確度35.6個百分點,在 Gemma 4 26B-A4B 中為26.6,在 Gemma 4 31B 中為12.0,在 Qwen3.8-Max 中為5.4。這四個配對差距在統計上都是顯著的。推理預算和提示語言干預可以大幅減少這一差異,顯示出這是一種可調整的能力,而不是固定的無能。最基本的結論是行為性的:在這些任務中,模型所表現出的關係能力對於呈現方式並不無所謂。明確的定義提供了正式的關係,但並未使臨時謂詞如同嵌入在學習語言聯想中的熟悉詞彙那樣可用。

DoGBench: Can Agents Meet Expert Standards for User-Facing Documentation?

2609.39909v1 by Frances Liu, Manny Silva, Paige Calvert, Ayu Adiati, Sarah Sanders

We introduce DoGBENCH (Documentation Generation Benchmark), to our knowledge, the first benchmark for generating and maintaining real user-facing software documentation. It asks whether an agent can produce documentation that experienced technical writers would accept in review. The benchmark contains 292 items from open source projects, including Helm, PostHog, and Mautic. Each item gives the agent a pre-change repository and a trigger, such as a code pull request or a reported documentation gap. The agent must first decide whether the documentation needs an update. For items that need one, the agent must produce an acceptable patch in one attempt. For items that do not need updates, the agent must abstain. Task-specific rubrics, validated with project maintainers, score each patch on accuracy, completeness, reader guidance, placement, and repository conventions. The composite score combines patch quality with correct abstention, and a score of 100 means an agent meets every requirement for the task. Scores should not be interpreted as a percentage of an expert's capability. We evaluated seven agents. The highest-scoring agent reached 47.3 out of 100 on the 117-item held-out split. In a separate audit of 1,267 patches, the most common failure modes were task-completion gaps (45.5%), technical inaccuracies (36.6%), and incomplete conceptual or reference coverage (32.5%). Analysis of the corresponding trajectories identified three key patterns associated with these failures: (1) describing interfaces without examining how readers use them (36.0%), (2) missing decisive evidence and filling the gaps with plausible assumptions (33.1%), and (3) stopping after finding the first plausible documentation surface and leaving other affected pages stale (30.1%).

摘要:我們介紹 DoGBENCH(文檔生成基準),據我們所知,這是第一個用於生成和維護面向真實用戶的軟件文檔的基準。它詢問一個代理是否能夠生成經驗豐富的技術寫手在審查中會接受的文檔。該基準包含來自開源項目的 292 項內容,包括 Helm、PostHog 和 Mautic。每個項目給代理一個變更前的代碼庫和一個觸發器,例如代碼拉取請求或報告的文檔缺口。代理必須首先決定文檔是否需要更新。對於需要更新的項目,代理必須在一次嘗試中生成可接受的補丁。對於不需要更新的項目,代理必須選擇不作為。特定任務的評分標準經過項目維護者的驗證,根據準確性、完整性、讀者指導、位置和代碼庫慣例對每個補丁進行評分。綜合得分將補丁質量與正確的選擇不作為相結合,得分為 100 意味著代理滿足該任務的所有要求。得分不應被解釋為專家能力的百分比。我們評估了七個代理。得分最高的代理在 117 項保留分割中達到了 47.3 分(滿分 100)。在對 1,267 個補丁的單獨審核中,最常見的失敗模式是任務完成缺口(45.5%)、技術不準確(36.6%)和概念或參考覆蓋不完整(32.5%)。對應的軌跡分析確定了與這些失敗相關的三個關鍵模式:(1)描述接口而不檢查讀者如何使用它們(36.0%),(2)缺少決定性證據並用合理的假設填補空白(33.1%),以及(3)在找到第一個合理的文檔表面後停止,並讓其他受影響的頁面保持過時(30.1%)。

Cognitive Enhancement: Rethinking the Necessity of Role-Playing for Large Language Models

2609.39853v1 by Xingjie Zhuang, Jialong Tang, Chulun Zhou, Buchao Zhan, Zhirui Li, Junhui Li, Yazheng Yang, Jinsong Su

Role-playing prompting has become a popular yet simple technique for improving LLM reasoning and output quality. However, whether it consistently boosts performance across diverse domains remains unclear, as systematic validation is lacking. To fill this gap, we run multi-model, cross-domain, and multilingual experiments on MMLU and MMLU-Redux. We find that gains from role-play prompting depend heavily on model capacity, knowledge domain, and prompt language. Drawing on metacognition theory, we propose the persona-related cognitive alignment hypothesis: role-play works only when the LLM correctly grasps the designated persona and its associated knowledge domain. We test this hypothesis through persona information richness ablation, layer-wise entropy divergence analysis, and latent thought-space deflection observation. To reduce persona cognitive bias and stabilize role-play performance, we propose \textbf{M}ixed-\textbf{L}anguage \textbf{C}oncatenate \textbf{P}rediction \textbf{(MLCP}), a simple, training-free, and efficient multilingual prompt concatenation strategy. It aggregates semantically equivalent role prompts to enrich complementary representational cues. Extensive experiments show that MLCP consistently outperforms vanilla role-play prompting across all tested LLMs.

摘要:角色扮演提示已成為一種流行但簡單的技術,用於改善大型語言模型(LLM)的推理和輸出質量。然而,這種方法是否在不同領域中始終能提升性能仍不明確,因為缺乏系統性的驗證。為了填補這一空白,我們在MMLU和MMLU-Redux上進行了多模型、跨領域和多語言的實驗。我們發現,角色扮演提示的增益在很大程度上取決於模型的能力、知識領域和提示語言。根據元認知理論,我們提出了與角色相關的認知對齊假設:角色扮演僅在LLM正確理解指定角色及其相關知識領域時有效。我們通過角色信息豐富度消融、層級熵差異分析和潛在思維空間偏轉觀察來測試這一假設。為了減少角色認知偏見並穩定角色扮演表現,我們提出了\textbf{M}ixed-\textbf{L}anguage \textbf{C}oncatenate \textbf{P}rediction \textbf{(MLCP)},這是一種簡單、無需訓練且高效的多語言提示串聯策略。它聚合語義等價的角色提示,以豐富互補的表徵線索。大量實驗表明,MLCP在所有測試的LLM中始終優於傳統的角色扮演提示。

Explore-on-Graph: Hybrid Embedding-LLM Reasoning for Knowledge Graph Question Answering under Incompleteness

2609.39786v1 by Ola El Khatib, Djellel Difallah

Large language models (LLMs) are increasingly combined with knowledge graphs (KGs) to ground reasoning in structured evidence. However, most LLM-based KGQA methods rely on traversing existing graph edges and become unreliable when reasoning paths are broken by missing facts. Alternatives that ask LLMs to generate missing knowledge risk introducing hallucinated evidence. We introduce XoG (eXplore-on-Graph), a framework for multi-hop question answering over incomplete KGs that recovers missing reasoning paths from learned graph structure rather than LLM parametric knowledge. XoG combines type-level entity-relation statistics to identify candidate relations with KG embeddings to retrieve plausible missing entities, using the LLM as a semantic selector and reasoner. These mechanisms are integrated into an iterative planning-exploration-reasoning process. Experiments on WebQSP, CWQ, and the Wikidata-based BRINK benchmark show that XoG remains competitive on complete KGs and consistently outperforms comparable methods without task-specific KGQA training under KG incompleteness. These gains persist across multiple LLM backbones, indicating that stronger LLMs alone do not resolve missing graph evidence. XoG also reduces LLM token consumption by up to 33% compared with a closely related planning-based approach.

摘要:大型語言模型(LLMs)越來越多地與知識圖譜(KGs)結合,以在結構化證據中進行推理。然而,大多數基於LLM的KGQA方法依賴於遍歷現有的圖邊,當推理路徑因缺失事實而中斷時,這些方法便變得不可靠。要求LLM生成缺失知識的替代方案則有引入幻覺證據的風險。我們介紹了XoG(eXplore-on-Graph),這是一個針對不完整KG的多跳問題回答框架,它從學習到的圖結構中恢復缺失的推理路徑,而不是依賴於LLM的參數知識。XoG結合了類型級實體-關係統計來識別候選關係,並利用KG嵌入來檢索合理的缺失實體,使用LLM作為語義選擇器和推理者。這些機制被整合到一個迭代的規劃-探索-推理過程中。在WebQSP、CWQ和基於Wikidata的BRINK基準上的實驗表明,XoG在完整KG上保持競爭力,並在KG不完整的情況下,始終優於沒有特定任務KGQA訓練的可比方法。這些增益在多個LLM骨幹中持續存在,表明僅僅強大的LLM並不能解決缺失的圖證據。與一種密切相關的基於規劃的方法相比,XoG還將LLM的標記消耗減少了多達33%。

MemCodex: Self-Programming Hierarchical Memory for Language Agents

2609.39765v1 by Xiaoqiang Wang, Bang Liu

Agent memory faces heterogeneous access needs: a single-hop question may require one piece of evidence, whereas a multi-hop question must combine evidence from multiple sources. Predefined memory workflows cannot adapt to these varying needs. Recent adaptive methods search or learn over memory components and their compositions, but the design space itself remains predefined. We introduce MemCodex, a self-evolving hierarchical memory system that organizes experience into executable memory programs for summaries, relational knowledge, reusable skills, and latent memory. Open-ended program evolution searches the open design space of layer programs by rewriting how each layer is constructed, indexed, retrieved, and routed, thereby adapting both within-layer implementations and cross-layer composition. At query time, reads traverse the hierarchy from coarse to fine and stop once sufficient evidence is found, descending to the original history when needed. We further develop MemArena, a unified runtime that places heterogeneous data and memory systems behind a common interface. MemCodex improves average task success by 10.1% relative to the strongest adaptive-memory baseline, while using 3.4x fewer context tokens and achieving 2.1x faster inference.

摘要:代理記憶面臨異質的訪問需求:單跳問題可能需要一個證據,而多跳問題必須結合來自多個來源的證據。預定義的記憶工作流程無法適應這些不同的需求。最近的自適應方法在記憶組件及其組合上進行搜索或學習,但設計空間本身仍然是預定義的。我們介紹了MemCodex,一個自我演化的分層記憶系統,將經驗組織成可執行的記憶程序,用於摘要、關聯知識、可重用技能和潛在記憶。開放式程序演化通過重寫每一層的構建、索引、檢索和路由方式,搜索層程序的開放設計空間,從而適應層內實現和跨層組合。在查詢時,讀取從粗到細遍歷層級,並在找到足夠的證據後停止,必要時降回原始歷史。我們進一步開發了MemArena,一個統一的運行時,將異質數據和記憶系統置於共同接口後面。相較於最強的自適應記憶基準,MemCodex提高了平均任務成功率10.1%,同時使用了3.4倍更少的上下文標記,並實現了2.1倍更快的推理。

OverForge: Reasoning Through Strategies and Tactics Helps Cooperative Lifelong Adaptation

2609.39727v1 by Oana Madalina Fron, Ojas Shirekar, Chirag Raman

Cooperative language-model agents must coordinate over long horizons and adapt to changing environments and to partners with unfamiliar conventions, yet existing agents map observations to actions without separating persistent coordination strategies from their tactical execution. We introduce OverForge, a training-free hierarchical architecture that separates strategic reasoning over roles and divisions of labour from tactical reasoning over actions within each agent's private, partner-conditioned world model. A metacognitive Prefrontal Cortex Module couples the two levels by forming strategy-action branches, imagining their consequences with a forward model, and committing when confident. In OvercookedV2, OverForge delivers 7 soups in a connected kitchen versus 3 for each flat LLM baseline, retains agreed roles, and adopts roles proposed by unfamiliar partners. Ablations and a fixed-strategy probe show that persistent strategies guide tactical adaptation while each reasoning level contributes to coordination. Memory restarts show that cross-episode partner knowledge supports task performance and partner prediction, linking the hierarchy to continual adaptation.

摘要:合作語言模型代理必須在長期內協調並適應不斷變化的環境以及具有不熟悉慣例的夥伴,但現有的代理將觀察映射到行動,而不將持久的協調策略與其戰術執行分開。我們介紹了OverForge,一種無需訓練的分層架構,將角色和勞動分工的戰略推理與每個代理的私有、基於夥伴的世界模型中的行動戰術推理分開。元認知前額葉皮層模塊通過形成策略-行動分支來聯結這兩個層次,利用前向模型想像其後果,並在有信心時作出承諾。在OvercookedV2中,OverForge在連接的廚房中交付7碗湯,而每個平面LLM基準僅交付3碗,保持商定的角色,並採納不熟悉夥伴提出的角色。消融實驗和固定策略探針顯示,持久策略指導戰術適應,而每個推理層次都對協調有所貢獻。記憶重啟顯示,跨劇集的夥伴知識支持任務表現和夥伴預測,將層次結構與持續適應聯繫起來。

ArchitectureIQ: On the Measure of Training Intuition

2609.39714v1 by Zirui Ren, Shaoyang Guo, Chencheng Tang, Jinxin Wang, Chengyu Xiong, Shanbin Yu, Peihang Li, Yidi Wu, Bangzhe Huang, Qingyu Qu, Leqian Yang, Ziming Liu

Top researchers have good intuition, but do language models have as good intuition about model training as top AI researchers? To measure model intuition of LLMs and humans, we introduce the ArchitectureIQ benchmark. Each question presents a synthetic dataset and several training recipes, and the test-taker is asked to predict the recipe yielding the best test metric. Overall, we find that LLMs' model intuition is good but has four limitations: (1) The intuition is imperfect, or even sub-human in some cases. Frontier models achieve around 76% accuracy (random choice 33%) vs best human researcher (66.0%), yet remain far from perfect. For architecture-only questions, best human achieves 65% while GPT-6 Astra only has 38%. (2) The intuition is empirical, not structured, supported by the fact that more CoT compute does not lead to substantial improvement. Unlike math, we still lack a "Science of AI" language that enables structured reasoning on AI. (3) The intuition is not maximally condensed, and can be further compressed into a knoledge base. Our constructed knowledge base with only 20 items yields large gains for weak models: GPT-4o equipped with the accumulated knowledge almost matches the performance of Claude Opus 5. (4) The intuition is insensitive to dataset properties, but the best model should in general depend on data properties. This suggests that data is the real "dark matter" in AI -- LLMs (so do human researchers) understand too little about data, even less than model architectures.

摘要:頂尖研究者擁有良好的直覺,但語言模型對於模型訓練的直覺是否與頂尖AI研究者一樣出色?為了衡量LLM和人類的模型直覺,我們引入了ArchitectureIQ基準。每個問題都呈現一個合成數據集和幾個訓練配方,測試者被要求預測產生最佳測試指標的配方。總體而言,我們發現LLM的模型直覺良好,但有四個限制:(1) 直覺並不完美,甚至在某些情況下低於人類。前沿模型的準確率約為76%(隨機選擇33%),而最佳人類研究者為66.0%,但仍然遠未完美。對於僅限架構的問題,最佳人類達到65%,而GPT-6 Astra僅有38%。(2) 直覺是經驗性的,而不是結構化的,這一點得到了更多CoT計算並未帶來實質性改善的事實支持。與數學不同,我們仍然缺乏一種能夠對AI進行結構化推理的“AI科學”語言。(3) 直覺並未最大程度地濃縮,可以進一步壓縮成知識庫。我們構建的僅有20個項目的知識庫為弱模型帶來了巨大的增益:配備了累積知識的GPT-4o幾乎達到Claude Opus 5的性能。(4) 直覺對數據集特性不敏感,但最佳模型通常應依賴於數據特性。這表明數據是真正的AI“暗物質”——LLM(人類研究者也是)對數據的理解太少,甚至少於對模型架構的理解。

ChronoGraph: Functional 4D Scene Graphs with Vision-Language Models for Interaction Understanding and Grounded Planning

2609.39665v1 by Chenyangguang Zhang, Malgorzata Gwiazda, Guanlong Jiao, Yuanchen Ju, Federico Tombari, Koushil Sreenath, Marc Pollefeys, Sunghwan Hong

Embodied agents must determine where to act, anticipate the resulting scene changes, and interpret observed outcomes to guide subsequent actions. This requires connecting 4D interaction understanding, which explains how past actions changed the scene, with spatially grounded planning, which determines how and where to act toward a goal and anticipates the resulting scene changes. We introduce ChronoGraph, a functional 4D scene graph that links actions on affordance parts to semantic and geometric state changes. By representing observed and anticipated transitions in the same form, it provides a shared basis for understanding and planning. We construct ChronoGraphBench through an automatic data engine that converts human-interaction videos and simulated robot trajectories into graph-annotated questions for training and evaluating Vision-Language Models (VLMs) on both tasks. Using these annotations, we train ChronoGraphVLM by adapting pretrained VLMs in two stages. Graph-as-Chain-of-Thought supervised fine-tuning teaches the models to reconstruct observed transitions and predict future ones as graph traces before answering. Subsequent joint 4D graph reinforcement learning directly rewards graph properties and answer correctness. Experiments across model scales show improvements over the corresponding pretrained baselines and zero-shot transfer to VLM4D. Real-world demonstrations further show that graph-based planning and affordance grounding support mobile manipulation through existing robot skills without additional fine-tuning.

摘要:具身代理必須確定行動的地點,預測隨之而來的場景變化,並解釋觀察到的結果以指導後續行動。這需要將 4D 互動理解(解釋過去的行動如何改變場景)與空間基礎規劃(確定如何以及在哪裡朝著目標行動並預測隨之而來的場景變化)連接起來。我們介紹 ChronoGraph,一個功能性 4D 場景圖,將對可供性部分的行動與語義和幾何狀態變化聯繫起來。通過以相同的形式表示觀察到的和預期的轉變,它為理解和規劃提供了一個共同的基礎。我們通過一個自動數據引擎構建 ChronoGraphBench,該引擎將人類互動視頻和模擬機器人軌跡轉換為帶有圖形標註的問題,以便在兩個任務上訓練和評估視覺-語言模型(VLMs)。利用這些標註,我們通過在兩個階段適應預訓練的 VLMs 來訓練 ChronoGraphVLM。作為思維鏈的圖形監督微調教導模型重建觀察到的轉變並在回答之前預測未來的轉變作為圖形痕跡。隨後的聯合 4D 圖形強化學習直接獎勵圖形屬性和答案的正確性。跨模型規模的實驗顯示出相對於相應的預訓練基線的改進,以及對 VLM4D 的零樣本轉移。現實世界的演示進一步表明,基於圖形的規劃和可供性基礎支持通過現有的機器人技能進行移動操作,而無需額外的微調。

Zero-Compute Cross-Lingual Transferability Estimation Using Typological Feature Proxies

2609.39640v1 by Dalton Raphael Harmsen, Swier Garst, Thomas van Osch, Zarè Palanciyan, Joaquin Vanschoren

Cross-lingual transfer describes how knowledge in a source language benefits a target language. Measuring it quantitatively requires broad multilingual pre-training, as prior work has done with cross-lingual transfer matrices. We ask whether transfer is predictable from freely available typological features, and whether the prominence of high-resource source languages reflects typology or data quality and quantity. We show that typological databases contain cheap and dense signals about cross-lingual transfer. Our typology-only random forest on a 24-language prior-work transfer matrix scores leave-one-language-out $ρ{=}0.705$ and $R^2{=}0.49$, beating a non-typological control at $ρ{=}0.62$, which verifies the ability of typology-only predictions to reconstruct costly measured cross-lingual transfer. The signal survives leave-one-script-out and leave-one-family-out protocols, so script and family confounding do not explain the effect. By decomposing the transfer into a typology term and a resource-and-script bias term, we find the best-source ranking sensitive to this bias. In contrast, typology is not affected by this bias, which makes it a zero-compute screening tool that replaces hundreds of training runs with a model fit. Our code is available \href{https://github.com/dharmsen/typo-x-ling-transfer}{here}.

摘要:跨語言轉移描述了來源語言的知識如何惠及目標語言。定量測量需要廣泛的多語言預訓練,正如先前的工作所做的跨語言轉移矩陣。我們詢問是否可以從自由可用的類型特徵預測轉移,以及高資源來源語言的顯著性是否反映了類型學或數據質量和數量。我們展示了類型學數據庫包含有關跨語言轉移的廉價且密集的信號。我們的僅基於類型學的隨機森林在24語言的先前工作轉移矩陣上的得分為留一語言外 $ρ{=}0.705$ 和 $R^2{=}0.49$,超過了 $ρ{=}0.62$ 的非類型學控制,這證實了僅基於類型學的預測能夠重建昂貴的測量跨語言轉移的能力。該信號在留一腳本外和留一語系外的協議中仍然存在,因此腳本和語系的混淆並不能解釋這一效果。通過將轉移分解為類型學項和資源與腳本偏差項,我們發現最佳來源排名對此偏差敏感。相比之下,類型學不受此偏差影響,這使其成為一種零計算篩選工具,能夠用模型擬合取代數百次訓練運行。我們的代碼可在 \href{https://github.com/dharmsen/typo-x-ling-transfer}{這裡} 獲得。

RankEvolve: A Reliable Multi-Agent Auto-Research Harness for Evolving Ranking Models

2609.39551v1 by Zheng Chen, Linfeng Liu, Hong Li, Hong Yan

Auto-research agents, LLM systems that propose, implement, train, and evaluate model changes across iterations, promise to automate applied ML's experimental loop. Over long horizons, execution accuracy is a binding constraint: a change can silently leak held-out data, omit normalization, disconnect a gradient, or leave a train/eval flag unwired, invalidating expensive runs and compounding error across iterations. We present RankEvolve, an auto-research framework for evolving generative ranking models. An Executable Operating Protocol (EOP) declares phases, gates, branches, and loops, and the runtime enforces the compiled state machine. A meta-meta-harness composes complete black-box coding-agent products, including Claude Code and Codex, as execution-graph nodes that review and repair one another's work. In a budget-matched evaluation, heterogeneous composition raises all-oracle execution accuracy from the best single-product baseline of 45.8 percent to 62.5 percent (paired +16.7 points, 95 percent CI [6.6, 26.7]) while achieving a 10.4 percent silent critical-defect rate. An implemented knowledge layer carries findings, including negative results, across iterations. In a twelve-iteration deployment on the open-source HSTU recommender, RankEvolve reported NDCG@10 of 0.2192 on MovieLens-20M LARGE (+4.48 percent over the published anchor) and 0.1948 on BASE (+2.80 percent). ExecML-HSTU, seeded by incidents from that deployment, provides the oracle benchmark for the execution-accuracy evaluation. A pre-specified LitGPT transfer split replicates the heterogeneous-composition effect beyond recommendation (+12.5 points, 95 percent CI [3.0, 22.0]), and a paired ablation isolates per-step from full-protocol instruction injection. These results characterize when runtime-controlled composition of coding-agent products improves execution accuracy.

摘要:自動研究代理,LLM 系統在多次迭代中提議、實施、訓練和評估模型變更,承諾自動化應用機器學習的實驗循環。在長期的執行中,準確性是一個約束條件:一個變更可能會悄悄洩漏保留數據、忽略正規化、斷開梯度或使訓練/評估標誌未連接,從而使昂貴的運行失效並在迭代中累積錯誤。我們提出了 RankEvolve,一個用於演變生成排名模型的自動研究框架。一個可執行操作協議 (EOP) 聲明了階段、閘、分支和循環,並且運行時強制執行編譯的狀態機。一個元元框架組合了完整的黑箱編碼代理產品,包括 Claude Code 和 Codex,作為執行圖節點,彼此審查和修復工作。在一個預算匹配的評估中,異質組合將所有預言者的執行準確性從最佳單一產品基線的 45.8% 提高至 62.5%(配對 +16.7 點,95% 置信區間 [6.6, 26.7]),同時實現了 10.4% 的靜默關鍵缺陷率。一個實施的知識層在迭代中攜帶發現,包括負面結果。在對開源 HSTU 推薦系統的十二次迭代部署中,RankEvolve 在 MovieLens-20M LARGE 上報告的 NDCG@10 為 0.2192(比已發表的基準高出 +4.48%),在 BASE 上為 0.1948(高出 +2.80%)。ExecML-HSTU,基於該部署中的事件提供了執行準確性評估的預言者基準。一個預先指定的 LitGPT 轉移拆分在推薦之外複製了異質組合效應(+12.5 點,95% 置信區間 [3.0, 22.0]),而配對消融則將每步與完整協議指令注入隔離。這些結果表徵了何時運行時控制的編碼代理產品組合改善了執行準確性。

Learning Normal Diffusion Dynamics for Backdoor Defense in Text-to-Image Models

2609.39548v1 by Junjian Li, Xiaolong Liu, Peng Sun, Liantao Wu, Linghan Chen, Yudong Gao, Honglong Chen

Backdoor attacks pose a serious threat to the secure deployment of text-to-image (T2I) diffusion models. Existing defenses typically detect backdoors from specific abnormal patterns in internal representations, which may limit their generalizability with the emergence of increasingly diverse attack mechanisms. In this paper, we study backdoor defense of T2I diffusion models from a transition-dynamics perspective. We observe that benign diffusion trajectories exhibit structured and timestep-dependent transition patterns from cross-attention, latent and noise spaces, whereas backdoor attacks tend to induce deviations from such normal evolution. Motivated by these observations, we propose Normal Diffusion Dynamics Learning (NDDL), a novel backdoor defense framework that learns the normal transition dynamics of diffusion trajectories utilizing only benign samples. NDDL constructs compact multi-space trajectory representations and trains a timestep-conditioned dynamics model to predict the diffusion evolution. In the inference phase, deviations between the observed and predicted transitions are exploited to quantify dynamics inconsistency for backdoor detection. NDDL further enables trigger localization without any prior knowledge of the embedded backdoor by performing substitution with low-semantic words. Extensive experiments for diverse backdoor attacks demonstrate the effectiveness and generalizability of our proposed NDDL.

摘要:後門攻擊對文本到圖像(T2I)擴散模型的安全部署構成了嚴重威脅。現有的防禦通常通過檢測內部表示中的特定異常模式來識別後門,這可能會限制它們在不斷出現的多樣化攻擊機制中的通用性。本文從轉移動力學的角度研究T2I擴散模型的後門防禦。我們觀察到良性擴散軌跡在交叉注意、潛在和噪聲空間中顯示出結構化和時間步依賴的轉移模式,而後門攻擊則傾向於導致這種正常演變的偏差。受到這些觀察的啟發,我們提出了正常擴散動力學學習(NDDL),這是一種新穎的後門防禦框架,僅利用良性樣本學習擴散軌跡的正常轉移動力學。NDDL構建了緊湊的多空間軌跡表示,並訓練了一個時間步條件的動力學模型來預測擴散演變。在推斷階段,觀察到的轉移與預測轉移之間的偏差被用來量化動力學不一致性以進行後門檢測。NDDL進一步通過使用低語義詞進行替換,實現了無需任何嵌入後門的先驗知識的觸發器定位。針對多樣化後門攻擊的大量實驗證明了我們提出的NDDL的有效性和通用性。

A Reusable Semantic Web Framework for Evidence-Grounded Fundamental Rights Impact Assessments under the EU AI Act

2609.39537v1 by Faith Olopade, Delaram Golpayegani, David Lewis

The EU AI Act (Art. 27) requires deployers of high-risk AI systems to conduct Fundamental Rights Impact Assessments (FRIAs) before deployment, yet the evidence needed for credible assessments is fragmented across incompatible incident repositories, risk vocabularies, and legal texts. We present a reusable Semantic Web-based framework that consolidates this evidence for two high-risk public sector categories: employment and worker management (Annex III(4)) and access to essential public services (Annex III(5)(a)). A curated 150-record corpus is annotated along four axes using keyword, LLM, and hybrid methods and serialised as a SPARQL-queryable knowledge graph of 1,351 RDF triples. Five FRIA demonstration scenarios surface 103 records (68.7% coverage). Evaluation against a 69-record gold standard reveals that LLM-assisted classification of the employment domain achieves only $κ= 0.045$, a cautionary result for automated fairness-related evidence retrieval in this domain. All artefacts are released openly to support adoption by regulators, national authorities, and SMEs.

摘要:歐盟人工智慧法案(第27條)要求高風險人工智慧系統的部署者在部署前進行基本權利影響評估(FRIAs),然而,進行可信評估所需的證據在不相容的事件資料庫、風險詞彙和法律文本中是分散的。我們提出了一個可重用的基於語義網的框架,整合了兩個高風險公共部門類別的證據:就業和工人管理(附件III(4))以及獲取基本公共服務(附件III(5)(a))。一個經過策劃的150條記錄語料庫沿著四個軸進行了註釋,使用關鍵字、LLM和混合方法,並序列化為1,351個RDF三元組的SPARQL可查詢知識圖譜。五個FRIA示範場景顯示出103條記錄(68.7%的覆蓋率)。與69條記錄的金標準進行評估顯示,就業領域的LLM輔助分類僅達到$κ= 0.045$,這對於該領域自動化公平相關證據檢索是一個警示結果。所有產物均公開發布,以支持監管機構、國家當局和中小企業的採用。

Beyond the Shadows of Plato's Cave: Evaluating False Memory in Autonomous Agents via Counterfactual Reasoning

2609.39473v1 by Quan M. Tran, Zhuo Huang, Zhen Fang, Jing Zhang, Mingming Gong, Tongliang Liu

Autonomous agents increasingly rely on memory to generalize beyond their training environments. However, agents are bounded by what they have seen and believed, and leveraging such memories in unseen environments can introduce biases into their internal beliefs. We formalize this phenomenon as \textit{false memory}, which can arise from spurious correlations, environment shifts, and knowledge conflicts. Despite its importance, false memory is difficult to evaluate because it stems from agent internal beliefs and is easily confounded with ordinary generalization failures. Therefore, we propose FAME, a training-free framework that evaluates false memory through the evolution of agent beliefs under counterfactual reasoning. Specifically, counterfactual scenarios reveal how beliefs change as the latent concept of memory shifts under hypothetical interventions; thus, measuring the resulting concept drift provides a signal for distinguishing faithful versus false memory. Such concepts can be estimated from agent hidden states before answer generation, avoiding the need for reward design or answer sampling. Empirical experiments reveal that simply monitoring answers often fails to detect false memory, while FAME achieves AUROCs of 76.2% - 96.7% across false-memory settings, and outperforms the best baseline by 3.4% - 23.3% across realistic benchmarks, spanning math reasoning (GSM-Symbolic), code generation (GitChameleon), and complex reasoning (BigBench-Hard). We further release corresponding counterfactual templates and facilitate future research on false memory.

摘要:自主代理越來越依賴記憶來超越其訓練環境。然而,代理受到他們所見和所信的限制,並且在未見環境中利用這些記憶可能會將偏見引入他們的內部信念。我們將這一現象正式化為 \textit{錯誤記憶},這可能源於虛假的相關性、環境變化和知識衝突。儘管其重要性,錯誤記憶難以評估,因為它源於代理的內部信念,並且容易與普通的概化失敗混淆。因此,我們提出了 FAME,一個無需訓練的框架,通過在反事實推理下代理信念的演變來評估錯誤記憶。具體而言,反事實場景揭示了隨著記憶潛在概念在假設干預下的變化,信念如何改變;因此,測量由此產生的概念漂移提供了一個區分真實記憶和錯誤記憶的信號。這些概念可以從代理的隱藏狀態中估算,在答案生成之前,避免了獎勵設計或答案抽樣的需求。實證實驗顯示,僅僅監控答案往往無法檢測錯誤記憶,而 FAME 在錯誤記憶設置中達到了 76.2% - 96.7% 的 AUROC,並且在現實基準中比最佳基線高出 3.4% - 23.3%,涵蓋數學推理 (GSM-Symbolic)、代碼生成 (GitChameleon) 和複雜推理 (BigBench-Hard)。我們還發布了相應的反事實模板,並促進未來對錯誤記憶的研究。

CAST: Causal Advantage-Structured Training with Spatially Grounded Compositional Rewards for Diffusion Models

2609.39441v1 by Shu Yu, Chaochao Lu

Online reinforcement learning has been extended to flow matching for diffusion model (DM) image generation. However, this paradigm faces three limitations: (1) Window selection. Existing methods manually set the stochastic differential equation (SDE) sampling window, i.e., the denoising steps where exploration noise is injected. We instead determine it from each model's denoising trajectory. (2) Reward saturation. Current methods rely on scoring models trained on human annotations; we find that such scores are extremely high and nearly indistinguishable on the latest SOTA open-source DMs, making advantage estimation largely ineffective. (3) Sample inefficiency. A single scalar reward collapses different failure modes into almost identical scores, leaving minimal gradient guidance for targeted improvement. To address these issues, we propose CAST (Causal Advantage-Structured Training), an RL fine-tuning method for pretrained DMs, which (1) identifies the denoising step at which each model fixes the objects and their spatial arrangement in the image and uses that timing to set the SDE window, (2) decomposes each prompt via Causal Scene Graphs (CSG) into verifiable-atoms, i.e., minimal semantic units such as an object, count, attribute, or spatial relation that can each be checked independently, and rewards each atom separately, and (3) projects the signed atom-level advantages into pixel space through teacher-forced attention and uses them to spatially weight the SDE policy objective. We fine-tune two of the strongest open-source DMs, FLUX.2-dev and Qwen-Image-2512, with CAST, and evaluate them on GenEval 2, a compositional benchmark, and on Qwen-Image-Bench for overall quality. Within almost the same training budget, CAST's improvement over the base model on the most challenging GenEval 2 prompts is up to 3.07x that of Flow-GRPO, while overall generation quality also improves.

摘要:在線強化學習已擴展到擴散模型(DM)圖像生成的流匹配。然而,這一範式面臨三個限制:(1)窗口選擇。現有方法手動設置隨機微分方程(SDE)抽樣窗口,即注入探索噪聲的去噪步驟。我們則從每個模型的去噪軌跡中確定它。(2)獎勵飽和。目前的方法依賴於基於人類標註訓練的評分模型;我們發現這些分數極高,並且在最新的SOTA開源DM中幾乎無法區分,使得優勢估計在很大程度上無效。(3)樣本低效。一個單一的標量獎勵將不同的失敗模式壓縮為幾乎相同的分數,為有針對性的改進留下了最小的梯度指導。為了解決這些問題,我們提出了CAST(因果優勢結構訓練),這是一種用於預訓練DM的強化學習微調方法,該方法(1)確定每個模型在圖像中固定物體及其空間排列的去噪步驟,並利用該時機設置SDE窗口,(2)通過因果場景圖(CSG)將每個提示分解為可驗證的原子,即最小語義單元,如物體、數量、屬性或空間關係,每個都可以獨立檢查,並單獨獎勵每個原子,以及(3)通過教師強制注意力將簽名的原子級優勢投影到像素空間,並利用它們對SDE政策目標進行空間加權。我們使用CAST微調了兩個最強的開源DM,FLUX.2-dev和Qwen-Image-2512,並在組合基準GenEval 2和Qwen-Image-Bench上評估它們的整體質量。在幾乎相同的訓練預算下,CAST在最具挑戰性的GenEval 2提示上對基礎模型的改進高達Flow-GRPO的3.07倍,同時整體生成質量也有所提升。

Inferring Causal Relations between Two Sequences of Events with Language Models

2609.39406v1 by Nishchal Prasad, Eric Gaussier, Emilie Devijver, Alexander Obeid Guzman, Armen Aghasaryan, Gregor Gössler

Causal AI is a branch of Artificial Intelligence which helps understand and reason about cause and effect relationships, not just patterns or correlations. Causal discovery aims to infer elements of the underlying causal structure--often represented as a directed graph--from observational and, when available, interventional data. While causal discovery is the fundamental step for moving beyond mere associations toward genuine understanding, and thus the basic building block of causal AI, it becomes intrinsically difficult when causal relations must be inferred from single observations. In such situations, standard causal discovery methods cannot be used and one has to identify causal relations from limited amount of information. This is typically the case for, e.g., sequences of events produced by different alarms which need to be analyzed on the fly to detect abnormal phenomena, which are usually rare. We show in this study that it is possible to leverage the predictive power of Large Language Models (LLMs) to infer causal relations between only two sequences of events. This approach, which is validated on both synthetic and real data, provides better results than standard causal discovery algorithms on several time series data, even though these data were converted into smaller, single observed sequences.

摘要:因果 AI 是人工智慧的一個分支,幫助理解和推理因果關係,而不僅僅是模式或相關性。因果發現旨在從觀察數據以及在可用的情況下的干預數據中推斷潛在因果結構的元素——通常表示為有向圖。雖然因果發現是超越單純關聯走向真正理解的基本步驟,因此也是因果 AI 的基本構建塊,但當因果關係必須從單一觀察中推斷時,這變得本質上困難。在這種情況下,標準的因果發現方法無法使用,必須從有限的信息中識別因果關係。這通常適用於例如由不同警報產生的事件序列,這些序列需要即時分析以檢測異常現象,而這些現象通常是稀有的。我們在這項研究中顯示,利用大型語言模型 (LLMs) 的預測能力來推斷僅有兩個事件序列之間的因果關係是可能的。這種方法在合成數據和真實數據上都得到了驗證,並在幾個時間序列數據上提供了比標準因果發現算法更好的結果,即使這些數據被轉換為較小的單一觀察序列。

Exploring Heterogeneous Model Merging Approach for Complex Knowledge Transfer

2609.39369v1 by Jiahe Fan, Si Chen, Yinghao Hou, Wenbo Xia, Ke Xu, Hong Xie, Enhong Chen

Specialized models encode task-oriented behavior, but transferring that behavior to a general language model usually requires training, distillation, or representation alignment. We study whether such ability can instead be transferred directly at the parameter level. We apply two existing training-free heterogeneous merging methods, previously shown to transfer knowledge between general language models, to specialist-to-general transfer, projecting a specialist donor into the recipient's shape and interpolating backbone parameters without gradient updates or semantic alignment. Intersection-Merge (IM) injects a prefix-aligned donor slice matching the recipient shape, while Activate-Prune-Merge (APM) uses forward-pass activation statistics to select which donor dimensions to retain before injection. Across embedding, reranking, reward modeling, and MoE code-specialist transfer, both methods improve the general recipient, showing that simple heterogeneous merging can move capabilities across diverse specialist roles.

摘要:專門模型編碼任務導向行為,但將該行為轉移到一般語言模型通常需要訓練、蒸餾或表示對齊。我們研究這種能力是否可以直接在參數層面上轉移。我們應用兩種現有的無訓練異質合併方法,這些方法之前已顯示能在一般語言模型之間轉移知識,來進行專家到一般的轉移,將專家捐贈者投影到接收者的形狀,並在不進行梯度更新或語義對齊的情況下插值主幹參數。交集合併(IM)注入與接收者形狀匹配的前綴對齊捐贈者切片,而激活修剪合併(APM)則使用前向傳遞激活統計來選擇在注入之前保留哪些捐贈者維度。在嵌入、重新排序、獎勵建模和MoE代碼專家轉移中,這兩種方法都改善了一般接收者,顯示簡單的異質合併可以在多樣的專家角色之間轉移能力。

Offline Guidance, Online Reasoning: Reusing LLM Feedback for Small Language Models

2609.39346v1 by Bohan Zhang, Linan Yue, Weibo Gao, Pengyu Chen, Hong Guo, Yanqi Hao

Large language models (LLMs) offer strong reasoning capabilities but are often costly to access through commercial APIs, while small language models (SLMs) are easier to deploy locally yet remain weaker in reasoning. This capability-deployment gap has motivated LLM-SLM collaboration, which aims to improve SLM reasoning using LLM capabilities while preserving the deployment advantages of SLMs. Existing approaches mainly follow two paradigms. Knowledge distillation uses LLM-generated answers and reasoning trajectories to train SLMs offline, but requires parameter updates and additional training. Alternatively, online collaboration routes difficult problems to an LLM or leverages LLM-generated guidance and corrections when an SLM encounters difficulties. Although effective, online collaboration requires repeated LLM access. Moreover, the guidance produced for a particular problem is discarded after inference and cannot benefit subsequent problems involving similar reasoning states. In the paper, we focus on a more constrained setting in which the LLM is accessed only offline, the SLM parameters remain fixed, and online inference is performed solely by the SLM. To this end, we propose Reusable Latent Correction (RLC), which converts one-off natural-language guidance from a black-box LLM into persistent corrective experiences in the hidden space of an SLM. RLC stores these experiences in an external bank and retrieves them according to the SLM's current reasoning state, enabling the SLM to reuse LLM-derived corrections during inference without any online LLM calls. Experiments across multiple reasoning benchmarks and SLM scales show that RLC consistently improves SLM reasoning without parameter updates or online LLM calls. Code is available at https://github.com/ZBH031/reusable-latent-correction.

摘要:大型語言模型(LLMs)提供強大的推理能力,但通過商業API訪問的成本通常很高,而小型語言模型(SLMs)更易於本地部署,但在推理方面仍然較弱。這種能力與部署之間的差距促進了LLM-SLM的合作,旨在利用LLM的能力改善SLM的推理,同時保留SLM的部署優勢。現有的方法主要遵循兩種範式。知識蒸餾使用LLM生成的答案和推理軌跡來離線訓練SLM,但需要參數更新和額外的訓練。或者,線上合作將困難問題路由到LLM,或者在SLM遇到困難時利用LLM生成的指導和修正。雖然有效,但線上合作需要重複訪問LLM。此外,針對特定問題產生的指導在推理後會被丟棄,無法惠及後續涉及類似推理狀態的問題。在本文中,我們專注於一個更受限的設置,其中LLM僅在離線時訪問,SLM參數保持固定,並且線上推理僅由SLM執行。為此,我們提出了可重用潛在修正(RLC),它將來自黑箱LLM的一次性自然語言指導轉換為SLM隱藏空間中的持久修正經驗。RLC將這些經驗存儲在外部庫中,並根據SLM當前的推理狀態檢索它們,使SLM在推理過程中能夠重用LLM衍生的修正,而無需任何線上LLM調用。跨多個推理基準和SLM規模的實驗表明,RLC始終在不進行參數更新或線上LLM調用的情況下改善SLM的推理。代碼可在 https://github.com/ZBH031/reusable-latent-correction 獲得。

Understanding as No-Arbitrage: Bounded Dutch Books as a Definition and Training Objective for Language Models

2609.39341v1 by Daniel Dragonevskiy

Does a language model merely predict tokens, or does it understand what it says? We make this question measurable by defining "understanding" through the lens of no-arbitrage. A model understands a vocabulary to a certain degree if a computationally bounded trader cannot extract guaranteed profit by betting against the model's probabilities on logically related claims (a "Dutch book"). We establish three theoretical results: first, because full logical coherence is computationally intractable, understanding is inherently graded, not absolute. Second, we prove that the exact optimum of standard next-token prediction is inherently incoherent across different question formats; the flaw lies in the training objective, not the architecture. Third, we show that uncertainty accumulates predictably along reasoning chains, making unjustified overconfidence an arbitrage opportunity in itself. To address this, we introduce Arbitr, a training framework where an adversarial trader penalizes the model for logical inconsistencies, paired with a calibration anchor to prevent uninformative collapse. Across five pre-registered experiments on Qwen2.5 and Phi-3.5 models, we demonstrate that standard models are highly exploitable across different phrasings. Arbitr reduces this exploitability by orders of magnitude without sacrificing task accuracy, and the effect successfully transfers to unseen logical patterns and new model families. Crucially, we uncover a scaling illusion: at 7B parameters, near-zero measured incoherence often coincides with extreme, unjustified confidence. We conclude that while Arbitr enforces rigorous logical consistency, coherence is a necessary condition for knowledge, but not a sufficient one

摘要:一個語言模型僅僅是預測標記,還是它理解自己所說的內容?我們通過無套利的視角來定義“理解”,使這個問題可衡量。如果一個計算受限的交易者無法通過對模型在邏輯相關主張上的概率進行對賭來提取保證利潤(即“荷蘭書”),則模型在某種程度上理解詞彙。我們建立了三個理論結果:首先,由於完全的邏輯一致性在計算上是不可處理的,因此理解本質上是分級的,而不是絕對的。其次,我們證明標準的下個標記預測的精確最佳值在不同問題格式之間本質上是不一致的;這一缺陷在於訓練目標,而不是架構。第三,我們顯示不確定性在推理鏈中可預測地累積,使得不合理的過度自信本身成為一種套利機會。為了解決這個問題,我們引入了Arbitr,一個訓練框架,其中對手交易者因邏輯不一致而懲罰模型,並配備校準錨點以防止無信息的崩潰。在對Qwen2.5和Phi-3.5模型進行的五個預註冊實驗中,我們證明標準模型在不同措辭中高度可利用。Arbitr在不犧牲任務準確性的情況下,將這種可利用性降低了數個量級,並且這一效果成功轉移到未見的邏輯模式和新模型家族中。關鍵是,我們揭示了一種擴展錯覺:在70億參數下,接近零的測量不一致性常常與極端、不合理的自信相吻合。我們得出結論,雖然Arbitr強制執行嚴格的邏輯一致性,但一致性是知識的必要條件,而不是充分條件。

WorkGenesis: Building the Worlds That Teach Agents to Work

2609.39325v1 by Xinyu Zhu, Fenyi Liu, Yuzhu Cai, Shuo Tang, Rui Ye, Linfeng Zhang, Siheng Chen

The ability of Large Language Model (LLM) agents to complete daily and professional work is receiving increasing attention. Training such agents requires realistic work scenarios. Expert-authored occupational work is costly and slow to produce, while unconstrained synthesis often yields tasks with weak factual grounding or internally inconsistent requirements. To bridge this gap, we introduce WorkGenesis, a framework that constructs executable occupational work from real-world artifacts through two core technical innovations: (1) Evidence-Based Work Construction, which grounds each unit of work in real-world evidence by retrieving public files guided by O*NET occupational knowledge and synthesizing the surrounding context, companion materials, work request, and itemwise rubric around them; and (2) Execution-Guided Consistency Verification, which renders a reference deliverable inside the constructed work, attributes every unsatisfied rubric item to the agent, the task, or the rubric, and uses task and rubric defects as feedback to iteratively repair the work until it passes the audit. Experimental results demonstrate that Fx-Work-35B, trained with simple supervised fine-tuning (SFT) on only 20K units of work synthesized by WorkGenesis, achieves the highest scores among all comparable-scale baselines on the five reported metrics across GDPvalAA-v2, APEX-Agents-AA, and JobBench (31.00 versus 24.79 average score), and even surpasses frontier models such as the 1.6T DeepSeek-V4-Pro-Preview. These results show that WorkGenesis provides scalable training data for working agents.

摘要:大型語言模型(LLM)代理完成日常和專業工作的能力正受到越來越多的關注。訓練這樣的代理需要現實的工作場景。專家撰寫的職業工作成本高且生產緩慢,而不受限制的合成往往會產生事實基礎薄弱或內部不一致的要求的任務。為了填補這一空白,我們介紹了WorkGenesis,一個通過兩個核心技術創新來構建可執行職業工作的框架:(1) 基於證據的工作建構,通過根據O*NET職業知識檢索公共文件並合成周圍的上下文、伴隨材料、工作請求和逐項標準,將每個工作單元基於現實世界的證據進行基礎;(2) 執行引導的一致性驗證,這在構建的工作內部呈現參考交付物,將每個未滿足的標準項歸因於代理、任務或標準,並使用任務和標準缺陷作為反饋,迭代修復工作直到通過審核。實驗結果表明,Fx-Work-35B在僅用WorkGenesis合成的20K工作單元上進行簡單的監督微調(SFT)訓練,達到了在GDPvalAA-v2、APEX-Agents-AA和JobBench上報告的五個指標中所有可比規模基準中最高的分數(31.00對24.79的平均分),甚至超越了1.6T DeepSeek-V4-Pro-Preview等前沿模型。這些結果顯示,WorkGenesis為工作代理提供了可擴展的訓練數據。

Faithful Dual-constrained Erasure for Robust LLM Safety Alignment

2609.39279v1 by Jiaqing Li, Shide Zhou, Zhibo Zhang, Yuxi Li, Tianlong Yu, Kailong Wang

Machine unlearning has emerged as a crucial mechanism for removing hazardous knowledge and enforcing safety alignment in Large Language Models (LLMs). However, recent studies reveal a persistent security risk: unlearned models remain highly vulnerable to retraining attacks, where suppressed malicious behaviors rapidly resurface after benign fine-tuning. In this work, we investigate the optimization dynamics of unlearning and identify that this vulnerability stems from shallow alignment. Rather than effectively erasing target knowledge, models often exploit a shortcut by activating previously dormant parameters to act as spurious suppressors, forming a fragile inhibitory shell over intact malicious representations. To address this issue and enforce authentic memory deletion, we propose FDCU, a novel dual-constrained subspace projection framework. FDCU restricts parameter updates through a highly scalable, element-wise dual-masking rule: it preserves general knowledge manifolds via Fisher Information and strictly prohibits the abnormal activation of spurious suppressors via the Principle of Minimal Functional Intervention (PMFI). By reliably blocking the model's ability to superficially hide knowledge, FDCU promotes the authentic dismantling of target representations. Extensive experiments across specific knowledge erasure and safe output control tasks demonstrate that FDCU achieves state-of-the-art robustness against retraining attacks while maintaining near-lossless general utility, ensuring durable safety for LLMs.

摘要:機器去學習已成為移除危險知識和強化大型語言模型(LLMs)安全對齊的重要機制。然而,最近的研究揭示了一個持續的安全風險:去學習模型仍然對再訓練攻擊高度脆弱,抑制的惡意行為在良性微調後迅速重新出現。在這項工作中,我們研究了去學習的優化動態,並確定這一脆弱性源於淺層對齊。模型往往並未有效地抹去目標知識,而是通過激活先前靜止的參數來利用捷徑,作為虛假抑制器,形成一個脆弱的抑制外殼,覆蓋完整的惡意表徵。為了解決這一問題並強制實現真實的記憶刪除,我們提出了FDCU,一個新穎的雙約束子空間投影框架。FDCU通過一個高度可擴展的元素級雙遮罩規則來限制參數更新:它通過Fisher信息保留一般知識流形,並嚴格禁止虛假抑制器的異常激活,這是基於最小功能干預原則(PMFI)。通過可靠地阻止模型表面上隱藏知識的能力,FDCU促進了目標表徵的真實拆解。在特定知識刪除和安全輸出控制任務中的廣泛實驗表明,FDCU在抵抗再訓練攻擊方面達到了最先進的穩健性,同時保持近乎無損的一般效用,確保了LLMs的持久安全。

Fyan: A Human--AI Harness with Semantic Auditing for Document-Level Formalization

2609.39228v1 by Wei Zhao, Yangshuo Zou, Chengxiang Ding, Yifan Wu, Xuchuan Wang, Zimu Mao, Lei Zhang, Tao Luo

We present FYAN, a human--AI harness for document-level mathematical formalization. Rather than treating theorems in isolation, FYAN coordinates an end-to-end workflow spanning specification, proof planning, logical review, Lean proof construction, knowledge curation, and validation, with support for independent supervision and human guidance. A central component is evidence-grounded semantic auditing, which assesses whether formal statements faithfully preserve their informal specifications. A language model constructs structured evidence over local correspondences, omissions, scope, and logical relations, while a deterministic validator checks this evidence and produces reproducible judgments. When a substantive but admissible deviation is accepted, FYAN requires an explicit proof-transfer obligation connecting the formal statement back to a source-facing interpretation. With the same model (DeepSeek-V4.1-Flash) in every stage, FYAN proves 86 of 143 FormalTCS theorems under a strict Lean check, against 69 for a general agent harness, and raises the natural-language proof score from 0.501 to 0.851. On ConsistencyCheck, its semantic audit catches more inconsistent statements than a direct LLM judge, both on labels verified against the source (recall 0.777 vs. 0.636) and on the original labels (0.873 vs. 0.820), and localizes each mismatch it reports to a specific hypothesis, conclusion, or scope. FYAN also built ODENumLib, a 9,355-line Lean library for the numerical analysis of ordinary differential equation.

摘要:我們提出了 FYAN,一個用於文件級數學形式化的人類與 AI 結合工具。FYAN 不僅僅是將定理孤立地處理,而是協調了一個涵蓋規範、證明規劃、邏輯審查、Lean 證明構建、知識策展和驗證的端到端工作流程,並支持獨立監督和人類指導。一個核心組件是基於證據的語義審計,它評估正式陳述是否忠實地保留了其非正式規範。一個語言模型在局部對應、遺漏、範圍和邏輯關係上構建結構化證據,而一個確定性驗證器則檢查這些證據並產生可重複的判斷。當接受一個實質但可接受的偏差時,FYAN 要求一個明確的證明轉移義務,將正式陳述連接回源面向的解釋。在每個階段使用相同的模型(DeepSeek-V4.1-Flash),FYAN 在嚴格的 Lean 檢查下證明了 143 個 FormalTCS 定理中的 86 個,而一般代理工具僅證明了 69 個,並將自然語言證明分數從 0.501 提升至 0.851。在 ConsistencyCheck 上,其語義審計捕捉到的矛盾陳述比直接的 LLM 評判者更多,無論是在對照源驗證的標籤上(召回率 0.777 對 0.636)還是在原始標籤上(0.873 對 0.820),並將每個報告的不匹配定位到特定的假設、結論或範圍。FYAN 還構建了 ODENumLib,一個包含 9,355 行代碼的 Lean 庫,用於普通微分方程的數值分析。

DAGent: Evaluate-then-Grow Planning for Deep Research Agents

2609.39154v1 by Hanwen Liu, Yuanfu Sun, Qiaoyu Tan

Deep research tasks require agents to navigate large knowledge spaces, synthesize evidence across many sources, and adapt their plans as findings emerge. Directed acyclic graph (DAG)-based multi-agent systems suit this setting because they support parallel execution and isolate each sub-task within a focused dependency context. Yet existing DAG-based agents instantiate a task-level plan before execution and repair the graph only after failures or missing evidence are observed. This Plan-then-Patch strategy is brittle for deep research: the system commits most strongly when its evidence is weakest, and later revisions waste computation on branches that should not have been planned. We propose DAGent, a DAG-based multi-agent framework with Evaluate-then-Grow incremental planning: an Orchestrator grows the task graph one batch at a time, conditioning each expansion on confidence and uncertainty signals from completed nodes. A hierarchical context layer propagates compact QueryDocs by default while preserving full execution traces for on-demand recall. The recorded DAG topology admits structural RL signals that outcome-only recipes cannot define; DAGRPO, a GRPO adaptation, injects topology-conditioned credit on Executor rollouts and a structural compliance regularization on Orchestrator plans. Across BrowseComp-Plus, GAIA, and xbench-DeepSearch, DAGent surpasses the strongest open-source baseline by 5.3 / 5.8 / 2.0 points at the Qwen3-235B-A22B scale, and the lead replicates across four open-source backbones and extends to GPT-5 at 327K context. At the Qwen3-8B scale, DAGRPO improves over a same-budget outcome-only GRPO baseline by 3.0 average Pass@1 points. A same-architecture comparison shows that evidence-conditioned planning reaches higher accuracy at lower per-task token, tool-call, and step footprints than its Plan-then-Patch counterpart. Code: https://github.com/hanwenliu6825/DAGent

摘要:深度研究任務要求代理在大型知識空間中導航,綜合來自多個來源的證據,並隨著發現的出現調整計劃。基於有向無環圖(DAG)的多代理系統適合這種環境,因為它們支持並行執行並將每個子任務隔離在一個集中的依賴上下文中。然而,現有的基於DAG的代理在執行前會實例化一個任務級計劃,並僅在觀察到失敗或缺失證據後修復圖形。這種先計劃再修補的策略對於深度研究來說是脆弱的:系統在證據最薄弱時最強烈地承諾,而後續的修訂則在不應該計劃的分支上浪費計算。我們提出了DAGent,一個基於DAG的多代理框架,具有評估後增長的增量計劃:一個協調者一次增長一批任務圖,並根據已完成節點的信心和不確定性信號調整每次擴展。層次上下文層默認傳播緊湊的QueryDocs,同時保留完整的執行痕跡以便按需回憶。記錄的DAG拓撲允許結構性強化學習信號,這是僅基於結果的配方無法定義的;DAGRPO,GRPO的適應,將基於拓撲的信用注入到執行者的回滾中,並對協調者的計劃施加結構合規性正則化。在BrowseComp-Plus、GAIA和xbench-DeepSearch中,DAGent在Qwen3-235B-A22B規模上超越了最強的開源基線5.3 / 5.8 / 2.0點,並且這一優勢在四個開源骨幹中得到了重複,並擴展到327K上下文的GPT-5。在Qwen3-8B規模上,DAGRPO在同樣預算的僅基於結果的GRPO基線上提高了3.0的平均Pass@1點數。同一架構的比較顯示,基於證據的計劃在每個任務的標記、工具調用和步驟足跡上達到了更高的準確性。代碼:https://github.com/hanwenliu6825/DAGent

Rep2Skill: Representation-Guided Skill Self-Evolution for LLM Agents

2609.39149v1 by Kaixing Zhang, Changming Li, Yingdong Shi, Zheng Zhang, Kaitao Song, Wenjie Shi, Jingang Wang, Kan Ren

Textual skills enable large language model (LLM) based agents to accumulate reusable procedural knowledge without updating model parameters. Yet existing skill evolution remains largely confined to the text space: an optimizer must diagnose success and failure patterns, and revise skills solely from long execution trajectories and sparse task outcomes. This text-only paradigm leaves the agent's internal representations, which contain rich records of its evolving execution state, outside the skill optimization loop. We ask whether an agent can improve its external textual skills by reflecting on its own internal representations. We introduce Rep2Skill, a representation-guided framework for self-evolution on agent skills. Specifically, upon the collected agent rollouts, Rep2Skill models their internal model representation trajectories to localize turns that deviate from successful execution dynamics, and it further interprets these signals alongside the execution contexts as actionable textual feedback for targeted skill revision. Experiments on two agent environments with two open-source LLMs show that Rep2Skill consistently outperforms text-only approaches in the self-evolution setting, where the same LLM serves as both executor and optimizer without a stronger external model. This establishes a promising direction moving agent self-improvement beyond text-only reflection.

摘要:文本技能使基於大型語言模型(LLM)的代理能夠在不更新模型參數的情況下累積可重用的程序知識。然而,現有的技能演變在很大程度上仍然局限於文本空間:優化器必須診斷成功和失敗模式,並僅根據長期執行軌跡和稀疏的任務結果來修訂技能。這種僅限於文本的範式使代理的內部表徵(包含其不斷演變的執行狀態的豐富記錄)置於技能優化循環之外。我們詢問代理是否可以通過反思自身的內部表徵來改善其外部文本技能。我們介紹了Rep2Skill,一個基於表徵的自我演變框架,用於代理技能的自我提升。具體而言,在收集到的代理回合中,Rep2Skill建模其內部模型表徵軌跡,以定位偏離成功執行動態的轉折,並進一步將這些信號與執行上下文一起解釋為可行的文本反饋,以進行有針對性的技能修訂。在兩個代理環境和兩個開源LLM上的實驗顯示,Rep2Skill在自我演變設置中始終優於僅限文本的方法,其中同一LLM同時擔任執行者和優化器,而沒有更強的外部模型。這為推動代理自我改進超越僅限文本反思建立了一個有前景的方向。

MASCRDM: Multi-Agent System for Compliance Risk Detection and Mitigation in Training Process of Large Language Models

2609.39107v1 by Yan Zhang, Chuming Wei, Ruien Li, Yaoyao Peng, Wusheng Zhang, Guangwen Yang

Large Language Models (LLMs) have been applied in various fields. However, ensuring compliance and safety of LLMs, such as avoiding discrimination and bias, still remains a challenge. Current efforts mainly focus on detecting and filtering inputs and outputs of the trained models, rather than studying the intrinsic architecture of the models in real-time. To tackle this challenge, we analyze the LLMs training process and discover two critical issues: 1) Most of the existing methods are predominantly static in their approach to detection and filtering, achieving only localized optimizations without systematically enhancing the compliance of LLMs. 2) Another issue with existing approaches is the lack of real-time risk detection and mitigation across the full training process, which leads to limited flexibility. Motivated by these, we propose MASCRDM (Multi-Agent System for Compliance Risk Detection and Mitigation) during the LLM training process. Firstly, we develop a set of compliance rules based on existing Artificial Intelligence (AI) laws and a compliance-specific LLM with the instruction of compliance law experts. Then, we deconstruct LLMs into several components and identify key nodes based on the compliance knowledge graph. During LLMs training, we implement our multiple agents in the whole process, giving compliance risk alerts and suggestions for LLM developers. Experiments on discrimination and bias benchmark demonstrate that our multi-agent system can effectively improve the compliance while maintaining reasonable semantic performance. The results indicate that our method provides an executable path for mitigating compliance risk from within the LLMs systematically.

摘要:大型語言模型(LLMs)已被應用於各個領域。然而,確保LLMs的合規性和安全性,例如避免歧視和偏見,仍然是一個挑戰。目前的努力主要集中在檢測和過濾訓練模型的輸入和輸出,而不是實時研究模型的內在架構。為了解決這個挑戰,我們分析了LLMs的訓練過程,並發現了兩個關鍵問題:1)現有的大多數方法在檢測和過濾的方式上主要是靜態的,僅實現了局部優化,而未系統性地增強LLMs的合規性。2)現有方法的另一個問題是缺乏在整個訓練過程中的實時風險檢測和緩解,這導致了靈活性有限。受到這些問題的啟發,我們在LLM訓練過程中提出了MASCRDM(合規風險檢測和緩解的多代理系統)。首先,我們根據現有的人工智慧(AI)法律制定了一套合規規則,並與合規法律專家的指導下開發了一個合規專用的LLM。然後,我們將LLMs拆解為幾個組件,並根據合規知識圖譜識別關鍵節點。在LLMs的訓練過程中,我們在整個過程中實施了多個代理,為LLM開發者提供合規風險警報和建議。在歧視和偏見基準上的實驗表明,我們的多代理系統可以有效改善合規性,同時保持合理的語義性能。結果顯示,我們的方法為系統性地從LLMs內部緩解合規風險提供了一條可執行的途徑。

Multi-LLM Collaborative Alignment via Stackelberg Games

2609.39076v1 by Christina Hahn, Shangbin Feng, Dean Light, Swastik Roy, Hila Gonen, Yulia Tsvetkov

A pool of language models can collaborate and improve collectively by learning from one another's responses. These interactions depend on the instructions used during training. Existing methods typically sample instructions uniformly, even though their usefulness may change as the models improve: an instruction on which models' responses once differed in quality may later be answered equally well, while a previously difficult instruction may begin to provide a useful learning signal. We propose Stackelberg Alignment, a game-theory-inspired leader-follower framework that turns instruction selection into an adaptive curriculum. An EXP3 bandit acts as the leader, allocating a fixed sampling budget across instructions and updating its sampling distribution using a reward that combines instruction difficulty and response discriminability. The language models act as followers: they respond to the selected instructions, evaluate one another's responses, and learn from the resulting preference signals through DPO or GRPO. The framework uses Elo-style reputation-weighted peer judgment and reputation-based opponent matching to support reliable and competitive model interactions. Experiments across three heterogeneous model pools and 12 benchmarks spanning scientific discovery, reasoning, code, instruction following, and knowledge show that Stackelberg Alignment achieves the highest macro-average across three diverse model pools, outperforming the strongest training-time baseline by up to 7.4% and the best static inference baseline by 12-25%. Analysis confirms that the adaptive leader concentrates duels on the most informative instructions, and ablations show that both reputation-weighted judgment and reputation-based matching improve the effectiveness of multi-LLM evolution.

摘要:一組語言模型可以通過相互學習彼此的回應來協作並共同改進。這些互動依賴於訓練期間使用的指令。現有的方法通常均勻地抽樣指令,即使隨著模型的改進,其有用性可能會改變:曾經模型回應質量不同的指令,後來可能會得到同樣好的回答,而先前困難的指令可能開始提供有用的學習信號。我們提出了Stackelberg Alignment,一種受博弈論啟發的領導者-跟隨者框架,將指令選擇轉變為自適應課程。一個EXP3賭徒作為領導者,在指令之間分配固定的抽樣預算,並使用結合指令難度和回應可區分性的獎勵來更新其抽樣分佈。語言模型作為跟隨者:它們對所選指令做出回應,評估彼此的回應,並通過DPO或GRPO從結果偏好信號中學習。該框架使用Elo風格的聲譽加權同行評價和基於聲譽的對手匹配來支持可靠且具有競爭性的模型互動。在三個異質模型池和12個基準測試的實驗中,涵蓋科學發現、推理、代碼、指令遵循和知識,顯示Stackelberg Alignment在三個不同的模型池中達到了最高的宏觀平均,超越了最強的訓練時間基線高達7.4%,以及最佳靜態推理基線的12-25%。分析確認,自適應的領導者將決鬥集中在最具信息性的指令上,而消融實驗顯示,聲譽加權評價和基於聲譽的匹配都提高了多LLM演化的有效性。

2609.39069v1 by Siyu Song, Rui Xu, Jia Lin, Kai Liu, Weifang Wang

Test-time reasoning systems often respond to failure by restarting or revising the latest step, even when an earlier decision caused the error. We introduce CORE, a search controller that requests a certified conflict core from a verifier, backjumps to the latest decision in that core, and caches the conflict to avoid repeating it. Under sound verification, finite branching and depth, and exhaustive proposals, the uncapped search is complete and never prunes a valid solution. On 2,000 planted graph-coloring instances with matched proposals and an exact verifier, CORE reduces median verifier calls by 39.8% at 30 variables and 35.0% at 36 variables relative to chronological repair; caching further improves on backjumping alone. Across five reasoning tasks, CORE achieves 75.9% mean success with Qwen2.5-7B-Instruct and 84.2% with Qwen3-8B, compared with 72.5% and 81.8% for Tree of Thoughts. It also uses fewer verifier calls and generated tokens on both backbones. These results show the value of using certified failure explanations to direct language-model search.

摘要:測試時的推理系統經常通過重新啟動或修正最新步驟來應對失敗,即使早期的決策導致了錯誤。我們介紹了CORE,一個搜索控制器,它從驗證者那裡請求一個經過認證的衝突核心,回溯到該核心中的最新決策,並緩存衝突以避免重複。根據健全的驗證、有限的分支和深度以及徹底的提案,無上限的搜索是完整的,並且從不修剪有效解。對於2,000個植入的圖著色實例,使用匹配的提案和精確的驗證者,CORE在30個變數時將中位數驗證者調用減少了39.8%,在36個變數時減少了35.0%,相較於時間修復;緩存進一步改善了僅回跳的效果。在五個推理任務中,CORE在Qwen2.5-7B-Instruct上達到75.9%的平均成功率,在Qwen3-8B上達到84.2%,相比之下,Tree of Thoughts的成功率為72.5%和81.8%。它在兩個基礎架構上也使用了更少的驗證者調用和生成的標記。這些結果顯示了使用經過認證的失敗解釋來指導語言模型搜索的價值。

Structure-aware Reinforcement Learning for Protein Directed Evolution

2609.39048v1 by Zikun Nie, Suyuan Zhao, Yizhen Luo, Siqi Fan, Zaiqing Nie

Protein optimization remains a longstanding goal in life sciences. Existing machine learning-assisted directed evolution (MLDE) methods primarily rely on sequence-only features, overlooking the critical spatial constraints and co-evolutionary interactions encoded in protein structures. However, directly integrating structural information remains challenging due to the scarcity of reliable mutant structures. To address these issues, we propose StructEvo, a novel structure-aware reinforcement learning framework for protein directed evolution. StructEvo employs a delta-structure fusion encoder to approximate mutant structure features via feature differences, enabling dynamic incorporation of spatial knowledge. The vast mutation space is then decomposed into manageable subspaces through a structure-aligned hierarchical action network, while a geometric constraint further stabilizes delta feature learning. Our approach outperforms prior state-of-the-art methods by 9.2% and 16.3% on two challenging optimization benchmarks, and further identifies an experimentally validated epistasis pattern in GFP, highlighting the importance of structural guidance for effective protein directed evolution.

摘要:蛋白質優化仍然是生命科學中的一個長期目標。現有的機器學習輔助定向進化(MLDE)方法主要依賴於僅有序列的特徵,忽略了蛋白質結構中編碼的關鍵空間約束和共進化互動。然而,由於可靠突變體結構的稀缺,直接整合結構信息仍然具有挑戰性。為了解決這些問題,我們提出了StructEvo,一個新穎的結構感知強化學習框架,用於蛋白質定向進化。StructEvo採用一種增量結構融合編碼器,通過特徵差異來近似突變體結構特徵,從而實現空間知識的動態整合。然後,通過結構對齊的分層行動網絡,將廣泛的突變空間分解為可管理的子空間,而幾何約束進一步穩定了增量特徵學習。我們的方法在兩個具有挑戰性的優化基準上比之前的最先進方法提高了9.2%和16.3%,並進一步識別出GFP中的一個實驗驗證的表觀基因互作模式,突顯了結構指導對於有效蛋白質定向進化的重要性。

SimEX: Simulation-Integrated Robotics AutoResearch

2609.38982v1 by Jiaheng Hu, Roberto Martin-Martin, Peter Stone, Rocky Duan, Zhenyu Jiang, Guanya Shi

Coding agents powered by large language models (LLMs) have shown remarkable abilities to autonomously reason about and achieve goals in the digital world. However, bringing this success to the physical world remains challenging. On the one hand, direct generation methods (e.g., Code as Policies) often suffer from the LLMs' insufficient understanding of robots and physical environments. On the other hand, iterative trial-and-error tuning in the physical world (e.g., physical autoresearch) induces significant experimental cost and safety concerns. We introduce SimEX: Simulation-Integrated Robotics AutoResearch, an autoresearch framework that tightly integrates simulated experimentation, enabling coding agents to efficiently acquire physical capabilities for controlling real robots. SimEX operates in two stages. First, the agent conducts open-ended probe-and-optimize iterations in simulation, developing a robot toolbox with robust and generalizable capabilities. Second, the agent adapts the toolbox and the simulator together through only a few physical trials: each trial corrects the simulator, and the corrected simulator is used to diagnose failures and screen candidate repairs. We evaluate SimEX extensively in sim-to-sim settings and on physical robots. On challenging real-world manipulation tasks including towel folding, barcode scanning, and plate manipulation, SimEX enables coding agents to efficiently acquire robot skills without any demonstration and with only 10 minutes of real-robot interaction. These results suggest that simulation can be a critical component in achieving physical intelligence, not only as a source of training data that must closely replicate the real world, but also as a roughly correct laboratory where a coding agent develops the knowledge and procedures needed to act on the robot. More details and robot videos at https://robo-simex.github.io/

摘要:編碼代理由大型語言模型(LLMs)驅動,已顯示出在數位世界中自主推理和實現目標的卓越能力。然而,將這一成功帶入物理世界仍然面臨挑戰。一方面,直接生成方法(例如,將代碼視為政策)往往受到LLMs對機器人和物理環境理解不足的影響。另一方面,在物理世界中進行的迭代試錯調整(例如,物理自動研究)會產生顯著的實驗成本和安全問題。我們介紹SimEX:模擬整合機器人自動研究,這是一個自動研究框架,緊密整合模擬實驗,使編碼代理能夠高效獲取控制真實機器人的物理能力。SimEX分為兩個階段運作。首先,代理在模擬中進行開放式的探測和優化迭代,開發出具有穩健且可泛化能力的機器人工具箱。其次,代理通過僅進行幾次物理試驗來共同調整工具箱和模擬器:每次試驗都會修正模擬器,修正後的模擬器用於診斷故障並篩選候選修復方案。我們在模擬到模擬的設置和物理機器人上對SimEX進行了廣泛評估。在包括毛巾折疊、條碼掃描和盤子操作等具有挑戰性的現實世界操作任務中,SimEX使編碼代理能夠在沒有任何示範的情況下,僅用10分鐘的真實機器人互動高效獲取機器人技能。這些結果表明,模擬可以是實現物理智能的關鍵組成部分,不僅作為必須與現實世界緊密重複的訓練數據來源,還作為一個大致正確的實驗室,讓編碼代理發展出在機器人上行動所需的知識和程序。更多細節和機器人視頻請訪問 https://robo-simex.github.io/

DrivingBench: Can Vision-Language Models Drive a Toyota Corolla?

2609.38948v1 by Aditya Ramabadran, Simon Mahns, Tobias Gessler

Frontier models excel at many digital benchmarks, yet their ability to drive a real car, an everyday human skill, remains largely untested. We present DrivingBench, to our knowledge the first benchmark where general-purpose vision-language models must drive a real car. Through three tools, the models see camera frames from a Toyota Corolla and directly command its steering and velocity around a parking lot cone course at low speeds. The car may continue moving while the model thinks and new commands replace the currently running one, so inference latency is part of the task, testing the models' abilities to observe, act, monitor, recover, and complete a long-horizon objective under such constraints. We benchmark GPT-6 Astra, Claude Fable 5.1, GPT-5.6 Sol, and Grok 4.6 in vendor-native harnesses (Codex, Claude Code, Cursor) with up to three attempts each in one conversation; Astra is the only model to finish the course, on its second attempt, with no other attempt passing 50% of the course. Two of the four models improved materially across attempts with retained context. We also detail the design principles behind our action interface, and show how the tool output format and the framing of the task combined to determine whether models would drive at all or refuse. We release our harness, prompts, course map, and traces with video and telemetry for reproducibility.

摘要:前沿模型在許多數位基準測試中表現出色,但它們駕駛真實汽車的能力,這是一項日常人類技能,仍然在很大程度上未經測試。我們提出了DrivingBench,據我們所知,這是第一個基準測試,要求通用視覺-語言模型駕駛真實汽車。通過三個工具,模型可以看到來自豐田卡羅拉的攝影機畫面,並直接指揮其在停車場圓錐賽道上的轉向和速度,速度較低。當模型思考時,汽車可以繼續移動,新指令會取代當前正在運行的指令,因此推理延遲是任務的一部分,測試模型在此類限制下觀察、行動、監控、恢復和完成長期目標的能力。我們在廠商原生的鞍具(Codex、Claude Code、Cursor)中對GPT-6 Astra、Claude Fable 5.1、GPT-5.6 Sol和Grok 4.6進行基準測試,每個模型在一次對話中最多可嘗試三次;Astra是唯一在第二次嘗試中完成賽道的模型,其他模型的嘗試都未達到50%的賽道通過率。四個模型中有兩個在保留上下文的情況下在嘗試中有實質性改善。我們還詳細說明了我們行動介面的設計原則,並展示了工具輸出格式和任務框架如何結合來決定模型是否會駕駛或拒絕。我們釋放了我們的鞍具、提示、賽道地圖和視頻及遙測的追蹤數據,以便於重現。

Prototype-guided Bilateral Alignment Multimodal Federated Learning

2609.38925v1 by Tianchi Liao Tianchi_Liao, Lele Fu, Sheng Huang, Qing Hu, Hong-Ning Dai, Chuan Chen

Multimodal federated learning (MFL) has emerged as a pivotal paradigm for leveraging distributed data to enhance model performance. However, existing methods predominantly rely on idealized assumptions of model homogeneity and balanced modality distributions, rendering them ill-suited for practical scenarios characterized by heterogeneous client architectures and severe modality imbalance. To address these challenges, we propose a \textbf{M}ultimodal \textbf{Fed}erated learning Prototype-guided Bilateral Alignment (MFedPBA) framework. MFedPBA facilitates robust knowledge synergy through a dual alignment mechanism: (i) at the feature level, it aligns heterogeneous feature spaces via a projection encoder optimized by contrastive learning and the Gromov-Wasserstein distance; (ii) at the decision level, it employs an entropy-weighted aggregation of naturally aligned logit prototypes. This novel design achieves robust MFL by jointly tackling heterogeneous feature spaces and collectively aggregating decisions. Extensive experiments demonstrate that our method significantly outperforms state-of-the-art baselines under conditions of model heterogeneity and modality imbalance.

摘要:多模態聯邦學習(MFL)已成為利用分散數據以提升模型性能的關鍵範式。然而,現有的方法主要依賴於模型同質性和均衡模態分佈的理想假設,使其不適合於特徵異質客戶架構和嚴重模態不平衡的實際場景。為了解決這些挑戰,我們提出了一個\textbf{M}ultimodal \textbf{Fed}erated learning Prototype-guided Bilateral Alignment(MFedPBA)框架。MFedPBA通過雙重對齊機制促進穩健的知識協同:(i) 在特徵層面,它通過對比學習和Gromov-Wasserstein距離優化的投影編碼器對異質特徵空間進行對齊;(ii) 在決策層面,它使用自然對齊的logit原型的熵加權聚合。這一新穎的設計通過共同解決異質特徵空間和集體聚合決策來實現穩健的MFL。大量實驗表明,我們的方法在模型異質性和模態不平衡的條件下顯著超越了最先進的基準。

GraphForge: Training Working Agents with Graph-Anchored Workspace Synthesis

2609.38923v1 by Qisheng Su, Hanchen Wang, Guanru Zhu, Huicheng Jiang, Qiuyinzhe Zhang, Kou Shi, Zhen Fang, Ziao Zhang, Qingnan Ren, Zehui Chen, Tao Gui, Feng Zhao

Working agents need to read diverse files, coordinate tools, and produce deliverables. Training such agents requires tasks built on many real files with verifiable results, but few pipelines exist to synthesize this kind of data. Existing pipelines either generate files with models, which lack realism and diversity, or build tasks on real files without task-specific verifiers, leaving result quality unchecked. We introduce GraphForge, an evidence-graph based framework that grounds both the task and its verification in real files. Starting from occupation-grounded seeds for controlled diversity, GraphForge assembles a workspace of real files for each seed and builds an evidence graph over their relations. Since the task statement and rubrics are both derived from this graph, task requirements are backed by the workspace files and each criterion is anchored to the files needed to verify it. An initial rollout further tests executability, and a revision agent repairs the task and rubrics against the original files before trajectories are collected. Fine-tuning Qwen3.6-27B on 2,169 GraphForge trajectories brings GDPVal to 1445.7 (+65.7) under OpenHands, and Workspace-Bench-Lite and SpreadsheetBench II to 63.7 (+7.7) and 24.0 (+13.7) under Claude Code. Rejection fine-tuning on the SFT model's own rollouts, with candidates selected by the evidence-anchored rubrics, yields further improvements on all three benchmarks, suggesting that the rubrics provide a useful selection signal. The data and models are available.

摘要:工作代理需要閱讀多樣的文件、協調工具並產出交付物。訓練這些代理需要基於許多真實文件且具有可驗證結果的任務,但現有的管道很少能合成這類數據。現有的管道要麼生成缺乏現實感和多樣性的模型文件,要麼在沒有特定任務驗證者的情況下基於真實文件構建任務,導致結果質量無法檢查。我們介紹了 GraphForge,一個基於證據圖的框架,它將任務及其驗證根植於真實文件中。從以職業為基礎的種子開始以控制多樣性,GraphForge 為每個種子組裝一個真實文件的工作空間,並在它們的關係上構建一個證據圖。由於任務聲明和標準都是從這個圖中衍生的,任務要求得到了工作空間文件的支持,每個標準都與驗證所需的文件相連接。初步推出進一步測試可執行性,並且修訂代理在收集軌跡之前會根據原始文件修復任務和標準。在 2,169 個 GraphForge 軌跡上微調 Qwen3.6-27B,使 GDPVal 在 OpenHands 下達到 1445.7 (+65.7),而 Workspace-Bench-Lite 和 SpreadsheetBench II 在 Claude Code 下達到 63.7 (+7.7) 和 24.0 (+13.7)。對 SFT 模型自身推出的拒絕微調,通過證據錨定的標準選擇候選者,對所有三個基準進一步改善,這表明標準提供了一個有用的選擇信號。數據和模型均可用。

Unlearning Deceptive Behaviors in LLMs with Contrastive Forget Sets

2609.38909v1 by Haoran Tang, Rajiv Khanna

Large language models often know the truth and say otherwise: a model that answers correctly when asked neutrally will affirm a user's mistaken belief, or misstate a fact its system prompt wants hidden, once the context rewards it. Such deception is a behavior conditioned on context, not knowledge, yet machine unlearning, the natural tool for removing a behavior from the weights, is built to forget facts that a deceptive model still needs. We propose to unlearn when a model deceives rather than what it knows, with a contrastive forget unit built from the model's own realized deceptions: the same question under a deception-triggering and a neutral context, admitted only where belief holds and behavior flips. Standard objectives on this unit face a dilemma. Suppression objectives such as NPO leave much of the deception in place. Target-based objectives, which distill the model's neutral behavior into the pressured context, remove it but induce context blindness: a target generated without the context teaches the model to stop reading it, eroding benign system-prompt instructions, secret-keeping and the reasoning a monitor inspects, a failure invisible to deception rates and capability benchmarks. We introduce PACT, which trains toward pressure-aware counterfactual targets (the model's own honest response, with a trace that registers the pressure and resists it) while retaining the benign uses of the triggering context. On two 32B reasoning models, PACT reduces held-out deception from over 50% to under 3% while system-prompt adherence, secret-keeping and the reasoning trace stay at the base model's level. On a tug-of-war score of removal against retention, PACT reaches 0.94 and 0.86, against at most 0.77 and 0.60 for any baseline. Like removed knowledge, removed deception is shallow under relearning, and terms that simulate the attacker hold it only at a cost in context use.

摘要:大型語言模型經常知道真相卻說出相反的話:當中立地詢問時能正確回答的模型,會確認用戶的錯誤信念,或者在上下文獎勵它時,錯誤陳述其系統提示想要隱藏的事實。這種欺騙是基於上下文的行為,而非知識,然而,機器的遺忘,這一自然工具用於從權重中去除行為,卻是為了忘記一個欺騙模型仍然需要的事實。我們提議在模型欺騙時進行遺忘,而不是在它所知道的事情上,通過一個由模型自身實現的欺騙構建的對比遺忘單元:在一個觸發欺騙的上下文和一個中立上下文下的同一問題,僅在信念存在且行為翻轉的地方被承認。這個單元的標準目標面臨著困境。像NPO這樣的抑制目標會讓大部分欺騙保持不變。基於目標的目標,將模型的中立行為提煉到受壓上下文中,雖然去除了欺騙,但卻引發了上下文盲目性:在沒有上下文的情況下生成的目標教會模型停止閱讀它,侵蝕了良性的系統提示指令、保密和監控者檢查的推理,這是一種對欺騙率和能力基準來說是不可見的失敗。我們引入了PACT,它朝著壓力感知的反事實目標進行訓練(模型自身的誠實反應,帶有記錄壓力並抵抗它的痕跡),同時保留觸發上下文的良性用途。在兩個32B推理模型上,PACT將保留的欺騙從超過50%降低到低於3%,而系統提示的遵循、保密和推理痕跡仍保持在基礎模型的水平。在去除與保留的拔河得分中,PACT達到了0.94和0.86,而任何基線最多僅為0.77和0.60。像去除的知識一樣,去除的欺騙在重新學習下是淺薄的,模擬攻擊者的術語僅在上下文使用上付出代價。

K2P: Label-Free Knowledge to Prompt Distillation

2609.38898v1 by Yingchuan Zhang, Haoran Lu, Wenxuan Zhong, Ping Ma

Knowledge distillation can transfer reasoning from stronger teachers to frozen students through reusable prompts, but avoiding weight updates does not eliminate supervision. Without ground-truth answers, teacher solutions are unverified, and agreement with the teacher can reward shared mistakes. We introduce Knowledge-to-Prompt (K2P) for label-free knowledge distillation to prompts. K2P synthesizes reusable instructions from teacher solutions, refines them using paired teacher and student responses, and guides search and selection with answer agreement. It retains candidates that adaptive search may undervalue and selects on reserved questions. Deployment uses only the frozen student and selected prompt. Our theory separates generation and selection gaps and gives conditions under which agreement-guided construction yields accuracy guarantees despite imperfect teacher references. Across reasoning tasks and students, K2P outperforms label-free alternatives overall and remains competitive with supervised prompt optimization. Ablations and archive diagnostics assess the contributions of teacher solutions and refinement, while revealing the limits of agreement-guided selection.

摘要:知識蒸餾可以通過可重用的提示將推理從更強的教師轉移到凍結的學生,但避免權重更新並不消除監督。沒有真實答案的情況下,教師的解決方案是未經驗證的,與教師的一致性可能會獎勵共享錯誤。我們引入了無標籤知識蒸餾到提示的知識轉換(K2P)。K2P 從教師解決方案中合成可重用的指令,通過配對的教師和學生反應對其進行精煉,並用答案一致性指導搜索和選擇。它保留了自適應搜索可能低估的候選者,並在保留的問題上進行選擇。部署僅使用凍結的學生和選定的提示。我們的理論將生成和選擇差距分開,並給出一致性指導的構建在不完美教師參考下仍能產生準確性保證的條件。在推理任務和學生中,K2P 整體上表現優於無標籤的替代方案,並在監督提示優化中保持競爭力。消融和存檔診斷評估教師解決方案和精煉的貢獻,同時揭示一致性指導選擇的限制。

Unmerge: Efficient Machine Unlearning via Task Arithmetic

2609.38895v1 by Haoran Tang, Andrew Tan, Rajiv Khanna

Approximate machine unlearning seeks to remove the influence of a forget set from a trained model without full retraining. Existing gradient-based methods require data-dependent hyperparameter search, struggle when forget and retain knowledge are entangled, and offer little insight into where unlearning actually happens inside the network. We recast unlearning through the lens of task arithmetic: if finetuning produces a merged task vector $τ_m$ that combines learning on forget and retain sets, unlearning is the inverse operation that subtracts a learned forget component $τ_F$ to recover the retain task vector $τ_R$. The forget signal is concentrated: at every layer, forget activations lie in a subspace spanned by a handful of dominant directions, so we factorize $τ_F$ in a low-rank forget basis, which is faithful up to a small tail-eigenvalue residual and limits how far the correction can perturb retain. We then optimize three intuitive goals (match the merged vector inside the forget span, suppress leakage into the retain span, and bound the correction size) that provably bound forget leakage and retain damage in activation space. The resulting algorithm, Unmerge, is fast and powerful: on class-level unlearning with ResNet-50 on CIFAR-100 and Tiny ImageNet, it improves Tug-of-War by up to ~24% over a baseline of comparable runtime and by up to ~18% over stronger baselines that run ~5x slower, keeps membership-inference exposure at the level of retraining, and shrinks the feature-distribution gap to the retrained model, where relabeling methods leave forget features cleanly separable. Further studies show that Unmerge also applies to ViT-S/16 and scales to Llama-3.2-3B. The per-layer basis geometry that drives the algorithm also serves as a layerwise diagnostic for when and where unlearning becomes structurally hard.

摘要:近似機器遺忘旨在從訓練模型中去除忘記集的影響,而無需完全重新訓練。現有的基於梯度的方法需要依賴數據的超參數搜索,當忘記和保留知識交織在一起時會遇到困難,並且對於遺忘實際發生的位置提供的見解有限。我們通過任務算術的視角重新詮釋遺忘:如果微調產生了一個合併的任務向量 $τ_m$,該向量結合了對忘記和保留集的學習,則遺忘是減去學習到的忘記組件 $τ_F$ 的逆操作,以恢復保留任務向量 $τ_R$。忘記信號是集中在一起的:在每一層,忘記激活位於由少數主導方向所生成的子空間中,因此我們在低秩的忘記基底中對 $τ_F$ 進行因式分解,這對小尾特徵值殘差是忠實的,並限制了修正可以擾動保留的程度。我們然後優化三個直觀的目標(在忘記範圍內匹配合併向量,抑制對保留範圍的洩漏,以及限制修正大小),這些目標可以明確界定忘記洩漏和激活空間中的保留損害。最終的算法 Unmerge 既快速又強大:在 CIFAR-100 和 Tiny ImageNet 上使用 ResNet-50 進行類別級別的遺忘時,它在可比運行時間的基線上提高了 Tug-of-War 約 24%,在運行速度約慢 5 倍的更強基線上提高了約 18%,並將成員推斷暴露保持在重新訓練的水平,同時縮小了特徵分佈與重新訓練模型之間的差距,這使得重新標註方法能夠將忘記特徵清晰地分開。進一步的研究顯示,Unmerge 也適用於 ViT-S/16 並擴展到 Llama-3.2-3B。驅動算法的每層基底幾何形狀也作為一種層級診斷,幫助判斷何時以及在哪裡遺忘變得結構上困難。

Right Answers, Costly Models: The Efficiency Gap in LLM-based Optimization Modeling

2609.38884v1 by Zhong Li, Xin Huang, Jinhui Wan, Xiangyi Wang, Shenkai Zhang, Ruiqi Chen, Wenyu Liu, Zaiwen Wen, Ziyan Luo

Optimization modeling formulates real-world decision problems as mathematical programs that solvers can use to find optimal decisions. Large language models (LLMs) can automate this process, but the resulting correct formulations can require substantial time and memory to construct and solve, limiting practical scalability. Therefore, we systematically investigate whether LLMs can identify problem structure from natural-language descriptions and apply suitable optimization modeling techniques to generate mathematical models and solver code that solve the problems correctly and efficiently. To this end, we first curate OptTips, a knowledge base of 50 expert modeling techniques in eight families. Using this knowledge, we develop OptDachshund, a multi-agent framework that transforms problems from existing optimization benchmarks into new tasks for evaluating LLMs' use of modeling techniques. It constructs conventional and expert mathematical models with solver code for the same task and data, providing baselines for correctness and computational cost. The resulting EfficientOpt benchmark contains 561 expert-reviewed tasks with paired reference implementations. Evaluation of 11 representative LLMs reveals an efficiency gap on correctly solved tasks with comparable measurements: for every LLM, most generated programs take longer to solve than their expert counterparts. Within the comparable reference-size subset, 57\% of programs with correct objective values and fewer variables and linear constraints have longer recorded solver times. Case studies show that different modeling techniques can achieve the same optimal value at similar recorded cost. Faster solving may not reduce execution time if the code takes longer to prepare data and build the model. LLM optimization modeling should therefore be evaluated for both correctness and computational efficiency.

摘要:優化建模將現實世界的決策問題形式化為數學程序,解決者可以利用這些程序來尋找最佳決策。大型語言模型(LLMs)可以自動化這一過程,但生成的正確公式可能需要大量的時間和內存來構建和解決,從而限制了實際的可擴展性。因此,我們系統地研究LLMs是否能從自然語言描述中識別問題結構,並應用合適的優化建模技術來生成數學模型和解決器代碼,以正確且高效地解決問題。為此,我們首先整理了OptTips,這是一個包含八個類別中50種專家建模技術的知識庫。利用這些知識,我們開發了OptDachshund,一個多代理框架,將現有優化基準中的問題轉化為評估LLMs使用建模技術的新任務。它為相同的任務和數據構建常規和專家數學模型及解決器代碼,提供正確性和計算成本的基準。最終生成的EfficientOpt基準包含561個經專家審核的任務,並附有配對的參考實現。對11個具有代表性的LLMs的評估顯示,在正確解決的任務中存在效率差距,測量結果相似:對於每個LLM,大多數生成的程序解決所需時間比其專家對應物更長。在可比的參考大小子集中,57\%的程序具有正確的目標值且變量和線性約束較少,但記錄的解決時間更長。案例研究顯示,不同的建模技術可以在類似的記錄成本下達到相同的最佳值。如果代碼準備數據和構建模型所需的時間更長,則更快的解決可能不會減少執行時間。因此,LLM優化建模應該同時評估正確性和計算效率。

Does Learning Protein Folding Generalize to Broader Reasoning?

2609.38879v1 by Yong Liu, Zhanpeng Shi, Yizhou Dang, Zhongyue Zhang, Xiaoliang Shi, Zhijian Wei, Shuangjia Zheng

Large language models rely heavily on human text, which often conveys surface answers rather than the spatial and structural logic behind them. Protein folding is a natural testbed, because one solved structure yields thousands of exactly checkable spatial and topological statements. We ask: can learning to fold proteins teach general models reusable reasoning capabilities? To answer this, we build FoldingCorpus, a protein-derived question-answer dataset, and Fold2Reason, a recipe that post-trains on it through two complementary signals: discrete structural answers predicted via the model's native language head, and continuous 3D geometry decoded from the same shared representations. On FoldBench, Fold2Reason achieves structure prediction scores 2.7 to 3.5 times those of Qwen3.5-9B. Beyond protein structure prediction, it improves performance on all 10 benchmarks spanning spatial, graph, scientific, and general reasoning, raising macro-average accuracy from 45.09% to 48.33% (+3.23 pp), with positive gains on all 10 benchmarks, while matched controls built from random, synthetic, and shuffled structure yield substantially smaller or negative gains. Our work shows that non-linguistic, structure-dense scientific data can systematically improve broad reasoning in language models, making a solved scientific problem a practical source of post-training supervision.

摘要:大型語言模型在很大程度上依賴於人類文本,而這些文本往往傳達的是表面的答案,而不是其背後的空間和結構邏輯。蛋白質摺疊是一個自然的測試平台,因為一個已解決的結構可以產生數千個精確可檢查的空間和拓撲陳述。我們問:學習摺疊蛋白質能否教會通用模型可重用的推理能力?為了回答這個問題,我們構建了FoldingCorpus,一個基於蛋白質的問答數據集,以及Fold2Reason,一個通過兩種互補信號進行後訓練的配方:通過模型的原生語言頭預測的離散結構答案,以及從相同共享表示解碼的連續3D幾何。在FoldBench上,Fold2Reason的結構預測分數達到Qwen3.5-9B的2.7到3.5倍。除了蛋白質結構預測外,它還提高了在所有10個基準測試中的表現,這些基準涵蓋了空間、圖形、科學和一般推理,將宏觀平均準確率從45.09%提高到48.33% (+3.23 pp),在所有10個基準上均有正增益,而從隨機、合成和打亂結構建立的匹配對照則產生了顯著較小或負的增益。我們的工作表明,非語言的、結構密集的科學數據可以系統性地改善語言模型中的廣泛推理,使得已解決的科學問題成為實用的後訓練監督來源。

StateTree: Enhancing Long-Term Dialogue Reasoning via Reinforcement Learning

2609.38809v1 by Naen Xu, Wanqing Cui, Yibo Hu, Shixin Hong, Hengyu An, Meiguang Jin, Junfeng Ma, Tianyu Du

Large language models deployed as personalized assistants must reason over long, evolving interaction histories. However, in long-term dialogue reasoning, relevant evidence is scattered across sessions, preferences may be revised over time, and standard long-context training fails to address these challenges under data scarcity and prohibitive computational costs. We propose StateTree, a data-driven RL method that constructs a challenging auxiliary task from scarce dialogues with verifiable ground truth. StateTree augments multi-session dialogues with a tree-structured path-tracing task: key-value records are embedded across sessions to form a binary tree. Solving the task requires the model to traverse from root to leaf by retrieving records across sessions and comparing timestamps to resolve branches, then recover the hidden target question among distractor leaves. We apply curriculum RL training progressively increasing tree depth and introduce a compositional variant whose edges carry step-level reasoning fragments, training the model to compose partial cues into coherent queries. Trained on 10K-token contexts, StateTree generalizes to 128K tokens without full-length RL costs and exhibits capabilities including cross-session retrieval, temporal reasoning, knowledge update, and compositional multi-hop reasoning. StateTree outperforms both SFT and RL-based baselines while preserving short-context general reasoning. StateTree-7B achieves gains up to +23.60% on LongMemEval (128k), and StateTree-14B reaches 59.00% accuracy on LongMemEval, surpassing QwenLong-L1-32B (45.20%).

摘要:大型語言模型作為個人助理部署時,必須對長期且不斷演變的互動歷史進行推理。然而,在長期對話推理中,相關證據分散在不同的會話中,偏好可能隨時間而修訂,而標準的長上下文訓練未能在數據稀缺和高昂的計算成本下解決這些挑戰。我們提出了StateTree,一種數據驅動的強化學習方法,從稀缺的對話中構建出具有可驗證真實性的挑戰性輔助任務。StateTree通過一個樹狀結構的路徑追蹤任務來增強多會話對話:關鍵值記錄在會話之間嵌入,形成一個二叉樹。解決這個任務需要模型從根部遍歷到葉子,通過檢索會話中的記錄並比較時間戳來解析分支,然後在干擾葉子中恢復隱藏的目標問題。我們應用課程強化學習訓練,逐步增加樹的深度,並引入一種組合變體,其邊緣攜帶步驟級推理片段,訓練模型將部分提示組合成連貫的查詢。在10K標記上下文上訓練後,StateTree在不需全長強化學習成本的情況下,能夠推廣到128K標記,並展現出跨會話檢索、時間推理、知識更新和組合多跳推理等能力。StateTree在保持短上下文一般推理的同時,超越了SFT和基於強化學習的基準。StateTree-7B在LongMemEval(128k)上獲得了最高+23.60%的增益,而StateTree-14B在LongMemEval上達到了59.00%的準確率,超越了QwenLong-L1-32B(45.20%)。

GraphCert: Bootstrap Agentic Graph Reasoning with Certified Evidence Rubrics

2609.38798v1 by Weiqi Jiang, Yuchen Ying, Rui Wang, Kaixuan Chen, Bingde Hu, Shunyu Liu, Yu Wang, Tongya Zheng

Graph agents extend large language models (LLMs) with the ability to actively explore and reason over knowledge graphs through multi-step interactions with graph tools. However, training capable graph agents typically requires large collections of question-answer pairs and reasoning trajectories, whose manual construction is costly and difficult to scale. Moreover, employing proprietary LLMs to generate such supervision further risks exposing sensitive graph data to external services. Therefore, we propose GraphCert to bootstrap agentic graph reasoning with certified evidence rubrics during post-training. Specifically, the Bootstrapped Graph Quizzer guided by generation controls produces graph-grounded QA pairs and marks supporting evidence, which undergo execution certification and semantic curation. The accepted evidence is then canonicalized into certified evidence rubrics that later reward Graph Solver evidence alignment alongside answer correctness during GRPO training. Experiments on five graph reasoning domains in GRBENCH demonstrate that GraphCert consistently outperforms substantially larger LLM agents and post-training method. Furthermore, our analysis demonstrates that the learned policy transfers robustly across heterogeneous graph domains, suggesting that GraphCert acquires reusable graph-reasoning capabilities rather than domain-specific patterns. These results establish executable self-certification as an effective approach to self-training compact graph reasoning agents. Our code will be made publicly available.

摘要:圖形代理擴展了大型語言模型(LLMs),使其能夠通過與圖形工具的多步互動主動探索和推理知識圖譜。然而,訓練能夠的圖形代理通常需要大量的問答對和推理軌跡,其手動構建成本高且難以擴展。此外,使用專有的LLMs來生成這種監督進一步增加了將敏感圖形數據暴露給外部服務的風險。因此,我們提出了GraphCert,以在後訓練期間用經過認證的證據標準啟動代理圖形推理。具體而言,受生成控制指導的Bootstrapped Graph Quizzer生成基於圖形的問答對並標記支持證據,這些證據經過執行認證和語義策劃。被接受的證據隨後被標準化為經過認證的證據標準,這些標準在GRPO訓練期間獎勵圖形解決者的證據對齊以及答案的正確性。在GRBENCH的五個圖形推理領域的實驗表明,GraphCert始終超越了大得多的LLM代理和後訓練方法。此外,我們的分析表明,學習到的策略在異質圖形領域中穩健地轉移,這表明GraphCert獲得了可重用的圖形推理能力,而不是特定於領域的模式。這些結果確立了可執行的自我認證作為自我訓練緊湊圖形推理代理的有效方法。我們的代碼將公開提供。

Evaluating Persistent Calibration under Evolving Model Knowledge

2609.38797v1 by Victor Wang, Thomas Hofweber, Mohit Bansal, Elias Stengel-Eskin

As AI systems move from static repositories to agents that are capable of continual adaptation and learning, maintaining their trustworthiness means equipping the models backing them with the ability to produce confidence estimates that dynamically reflect their changing skills and knowledge. We introduce the problem of persistent calibration, which requires a confidence estimator to faithfully reflect the knowledge contained in a model as that knowledge changes, without recurring supervision. We operationalize this by examining persistent calibration across checkpoints of open models, asking whether confidence estimators trained on earlier checkpoints can generalize to later ones. Specifically, we aim to shed light on whether confidence is dependent on knowledge, a question with implications for the reliability of confidence estimates. To measure this relationship, we define and evaluate calibration on knowledge contrast sets: subsets containing questions that one checkpoint answers correctly and another checkpoint answers incorrectly, reflecting a change in knowledge. We show that both inference-time and fine-tuning methods fall short on contrast-set calibration compared to oracle methods trained on future checkpoints, even for methods that are well-calibrated on the full dataset. We provide evidence for the hypothesis that persistent calibration is challenging because there is a vast space of possible confidence functions that are well-calibrated on a given checkpoint, out of which only some rely on meta-knowledge features that would generalize to other checkpoints. Towards improving contrast-set calibration, we show that multi-checkpoint training helps, suggesting an avenue for identifying confidence features that remain robust across changing knowledge.

摘要:隨著人工智慧系統從靜態資料庫轉變為能夠持續適應和學習的代理,維持其可信度意味著需要為其背後的模型提供能夠產生動態反映其變化技能和知識的信心估計能力。我們引入了持續校準的問題,這要求信心估計器能夠忠實地反映模型中所包含的知識,隨著知識的變化而變化,而無需重複監督。我們通過檢查開放模型的檢查點來操作化這一點,詢問在早期檢查點上訓練的信心估計器是否能夠對後期檢查點進行泛化。具體而言,我們旨在闡明信心是否依賴於知識,這是一個對信心估計的可靠性有影響的問題。為了測量這種關係,我們在知識對比集上定義並評估校準:這些子集包含一個檢查點正確回答而另一個檢查點錯誤回答的問題,反映知識的變化。我們顯示,無論是推斷時的還是微調的方法,在對比集校準方面都不如在未來檢查點上訓練的神諭方法,即使對於在完整數據集上校準良好的方法也是如此。我們提供了證據支持持續校準是具有挑戰性的假設,因為在給定檢查點上有大量可能的信心函數是良好校準的,而其中只有一些依賴於能夠對其他檢查點進行泛化的元知識特徵。為了改善對比集校準,我們顯示多檢查點訓練是有幫助的,這暗示了一條識別在變化知識中保持穩健的信心特徵的途徑。

Self-Evolving Algorithm-Design Agents: Escaping In-Context Evolutionary Stagnation via Population-Curated Policy Optimization

2609.38757v1 by Chen Lu, Ke Xue, Siyuan Xu, Mingxuan Yuan, Chao Qian

Large language models are increasingly participating in complex real-world tasks in the form of algorithm-design agents, designing and refining algorithms. Many successful algorithm-design agents adopt pure in-context evolutionary frameworks, but they may quickly plateau in domains that require specialized knowledge. Parametric adaptation offers a way to internalize specialized knowledge, but conventional training requires abundant domain-specific corpora while high-quality algorithms are scarce in complex algorithm-design scenarios. In this paper, we propose sample-efficient parametric self-evolution where agents can explore and learn from self-generated algorithms. First, we characterize in-context evolutionary stagnation and analytically propose the Improvement Chain proposition, showing how learning successive self-generated algorithms can locally increase the likelihood of neighboring algorithms. Motivated by this local-transfer perspective, we further propose Population-Curated Policy Optimization (PCPO) to utilize a global population and a hybrid policy update scheme for retaining and reusing high-quality, diverse self-generated algorithms, shifting the policy towards stronger algorithms. In the task of learning rate schedule design for global placement in electronic design automation, trained only on 4 chip cases, PCPO outperforms the state-of-the-art in-context evolutionary methods (e.g., OpenEvolve and ShinkaEvolve) on average across 16 chip cases. With an 8B-size base model, PCPO achieves competitive performance compared to frontier closed-source models such as GPT-5.5. PCPO also reduces inference-time token cost by internalizing grounded domain knowledge and prompt distillation. Moreover, PCPO achieves significant speedups on four GPU kernel designs, with an average of 8.27$\times$ speedup against the PyTorch Eager baseline.

摘要:大型語言模型越來越多地參與複雜的現實世界任務,作為算法設計代理,設計和完善算法。許多成功的算法設計代理採用純粹的上下文進化框架,但在需要專業知識的領域中,它們可能很快達到瓶頸。參數適應提供了一種內化專業知識的方法,但傳統訓練需要大量特定領域的語料庫,而在複雜的算法設計場景中,高品質的算法卻稀缺。在本文中,我們提出了樣本高效的參數自我進化,讓代理可以探索並從自生成的算法中學習。首先,我們描述了上下文進化停滯的特徵,並分析性地提出了改進鏈命題,顯示學習連續自生成算法如何在局部上增加鄰近算法的可能性。受到這一局部轉移視角的啟發,我們進一步提出了人口策劃政策優化(PCPO),利用全球人口和混合政策更新方案來保留和重用高品質、多樣化的自生成算法,將政策轉向更強的算法。在電子設計自動化中進行全球佈局的學習率調度設計任務中,僅在4個晶片案例上訓練的PCPO在16個晶片案例中平均超越了最先進的上下文進化方法(例如,OpenEvolve和ShinkaEvolve)。使用8B大小的基礎模型,PCPO在性能上與前沿的封閉源模型(如GPT-5.5)相比具有競爭力。PCPO還通過內化基礎領域知識和提示蒸餾來降低推理時間的標記成本。此外,PCPO在四個GPU內核設計上實現了顯著的加速,與PyTorch Eager基準相比,平均加速達到8.27$\times$。

Learning to Route in Visual Space via Multi-Step Embedding Retrieval

2609.38743v1 by Tianyu Chen, Mingyuan Zhou, Jiaxing Wu

LLM agents rely on retrieval tools to access external knowledge, yet visual agentic search remains severely bottlenecked by standard single-step retrievers. In current pipelines, the agent must issue text queries for every intermediate step, struggling when visual clues are difficult to describe or when the retriever fails to surface necessary intermediate evidence within its top results. We hypothesize that offloading multi-step navigation across the entire embedding space directly to the retrieval tool resolves this performance bottleneck. To study this systematically, we introduce VHOP, a flexible data generation framework and benchmark with five core difficulty levels testing both visual matching and search planning. Using this framework, we develop VHOP-Router, an end-to-end training pipeline---combining supervised fine-tuning, online imitation learning, and reinforcement learning---that transforms a standard embedding model into an autoregressive multi-step retriever. Operating directly in the visual latent space, VHOP-Router retrieves linked image chains in a single tool call without requiring the agent to formulate intermediate text queries. Experiments show VHOP-Router boosts retrieval performance from under 5\% to 76.3\%. In agentic search, it improves task success rates by 52.7\% and reduces the average token length by 61\% from 1886 to 728, whereas upgrading the agent yields only a 3.7\% gain. Compared to a strong baseline where the agent retrieves the top 50 results per step, VHOP-Router maintains superior performance while reducing in-context images by $23\times$ and cutting the cumulative API payload by $35\times$. The models also generalize robustly to unseen difficulty levels and realistic test sets. Ultimately, VHOP and VHOP-Router provide an efficient and effective solution for visual agentic search that leaves native LLM capabilities entirely intact.

摘要:LLM 代理依賴檢索工具來訪問外部知識,但視覺代理搜索仍然受到標準單步檢索器的嚴重瓶頸。在當前的流程中,代理必須為每個中間步驟發出文本查詢,當視覺線索難以描述或檢索器未能在其頂部結果中顯示必要的中間證據時,代理會面臨困難。我們假設將整個嵌入空間的多步導航直接卸載到檢索工具上可以解決這一性能瓶頸。為了系統地研究這一點,我們介紹了 VHOP,一個靈活的數據生成框架和基準,具有五個核心難度級別,測試視覺匹配和搜索規劃。利用這一框架,我們開發了 VHOP-Router,一個端到端的訓練流程——結合了監督微調、在線模仿學習和增強學習——將標準嵌入模型轉變為自回歸多步檢索器。VHOP-Router 直接在視覺潛在空間中操作,能在一次工具調用中檢索鏈接的圖像鏈,而無需代理制定中間文本查詢。實驗顯示,VHOP-Router 將檢索性能從不足 5\% 提升至 76.3\%。在代理搜索中,它將任務成功率提高了 52.7\%,並將平均標記長度從 1886 減少到 728,減少了 61\%,而升級代理僅帶來 3.7\% 的增益。與一個強大的基準相比,該基準中代理每步檢索前 50 個結果,VHOP-Router 在保持優越性能的同時,將上下文中的圖像減少了 $23\times$,並將累積 API 負載減少了 $35\times$。這些模型在未見過的難度級別和現實測試集上也能穩健地泛化。最終,VHOP 和 VHOP-Router 為視覺代理搜索提供了一種高效且有效的解決方案,完全保留了原生 LLM 的能力。

Concept-Grounded Attention: A Controlled Evaluation of Graph-Injected Attention, Temporal Versioning, and Epistemic Status

2609.38684v1 by Sachin Dev Duggal, Pradyumna Swarnalatha Ramanna, Alexandros Vassiliades

Knowledge-intensive language-model systems typically represent external knowledge as text chunks or static graphs, with limited support for concept evolution, point-in-time reasoning, and distinctions between validated and inferred knowledge. We introduce the Concept Lifecycle Model (CLM), which represents concepts as persistent, graph-grounded, temporally versioned entities with explicit provenance and epistemic status, and Concept-Grounded Attention (CGA), which injects concept-graph structure into transformer computation through graph-biased self-attention (Form A) and gated cross-attention over concept nodes (Form B). We evaluate the framework in controlled settings using disabled-mechanism baselines. On 200 MuSiQue and HotpotQA questions with retrieval fixed, concept-graph retrieval recovers explicit multi-hop paths but does not improve evidence recall. Form A appears to steer attention, with 2.76 times more attention on gold than distractor concepts, but the same ratio occurs when Form A is disabled; the learned bias is negligible and no answers change. An identity-preserving Form B improves F1 from 0.188 to 0.221, but control concepts yield 0.213, indicating that most of the gain reflects added capacity. On LongMemEval, explicit temporal representation improves answer accuracy by 13 to 25 points across all tested generators, up to 122B parameters, while simplified CLM version resolution performs similarly to dated serialization because concept identity is not established reliably. On a synthetic source-independence task, protocol-derived epistemic status reduces unsupported assertions from 28% to 0.1% in a fine-tuned small model and from 19-68% to 0-5% in 72-122B models. Overall, the results support making temporal validity and epistemic status explicit, while showing that graph-attention diagnostics are not informative without disabled-mechanism controls.

摘要:知識密集型語言模型系統通常將外部知識表示為文本塊或靜態圖形,對於概念演變、特定時間推理以及驗證知識與推斷知識之間的區別支持有限。我們介紹了概念生命周期模型(CLM),它將概念表示為持久的、基於圖形的、具有時間版本的實體,並具有明確的來源和認識狀態,以及概念基礎注意力(CGA),它通過圖形偏見自注意力(形式A)和對概念節點的門控交叉注意力(形式B)將概念圖結構注入Transformer計算中。我們在受控環境中使用禁用機制基準評估該框架。在200個MuSiQue和HotpotQA問題中,當檢索固定時,概念圖檢索恢復了明確的多跳路徑,但並未改善證據召回。形式A似乎引導注意力,對金標概念的注意力比對干擾概念高出2.76倍,但當禁用形式A時,比例相同;學習的偏見微不足道,且沒有答案改變。保持身份的形式B將F1從0.188提高到0.221,但控制概念的F1為0.213,表明大多數增益反映了增加的容量。在LongMemEval上,明確的時間表示在所有測試的生成器中將答案準確性提高了13到25分,參數最多可達122B,而簡化的CLM版本解析的表現與過時的序列化相似,因為概念身份未能可靠地建立。在一個合成的源獨立任務中,協議衍生的認識狀態將不支持的斷言從28%減少到0.1%(在微調的小模型中),在72-122B的模型中則從19-68%減少到0-5%。總體而言,結果支持將時間有效性和認識狀態明確化,同時顯示圖形注意力診斷在沒有禁用機制控制的情況下並不具信息性。

Bandits with Multiple Optimal Arms: Minimax Regret and Non-Adaptivity

2609.38659v2 by Kaixuan Ji, Qiwei Di, Qingyue Zhao, Heyang Zhao, Quanquan Gu

We study multi-armed bandits (MAB) with multiple optimal arms, motivated by the fact that many practical decision making problems admit multiple correct answers. For $K$-armed bandits with $A$ optimal arms, we first provide a sharper analysis of previous sub-sampling algorithms (De Heide et al., 2021; Zhu and Nowak, 2020), establishing a $\tilde{O}\Big(\frac{K-A}{\sqrt{KA}}\sqrt{T} \Big)$ minimax regret, where $T$ is the total number of interactions and $\tilde O(\cdot)$ drops all constant and logarithmic factors, improving the previous $\tilde{O}(\sqrt{KT/A})$ regret. We then provide a matching lower bound up to logarithmic factors, indicating that our established rate is nearly minimax-optimal. We further show that the knowledge of $A$ up to $\tilde{O}(1)$ factors is necessary to achieve near-optimal regret, as near-optimal algorithms for one number of optimal arms must incur substantially larger regret than optimal regret for a smaller number. Overall, our results provide a comprehensive minimax characterization of $K$-armed bandits with $A$ over the entire range of $1 \leq A \leq K-1$.

摘要:我們研究具有多個最佳臂的多臂賭徒(MAB),這是因為許多實際的決策問題允許多個正確答案。對於具有 $A$ 個最佳臂的 $K$ 臂賭徒,我們首先對之前的子抽樣算法(De Heide et al., 2021; Zhu and Nowak, 2020)提供了更精確的分析,建立了 $\tilde{O}\Big(\frac{K-A}{\sqrt{KA}}\sqrt{T} \Big)$ 的最小最大後悔,其中 $T$ 是總互動次數,$\tilde O(\cdot)$ 刪除了所有常數和對數因子,改善了之前的 $\tilde{O}(\sqrt{KT/A})$ 後悔。然後,我們提供了一個與對數因子相匹配的下界,表明我們建立的速率幾乎是最小最大最佳的。我們進一步顯示,對 $A$ 的知識最多需要 $\tilde{O}(1)$ 的因子,才能實現接近最佳的後悔,因為對於一個最佳臂的數量,接近最佳的算法必須承擔比較小數量的最佳後悔大得多的後悔。總體而言,我們的結果提供了對於 $1 \leq A \leq K-1$ 整個範圍的 $K$ 臂賭徒與 $A$ 的全面最小最大特徵描述。

Prompt2Skill: Unsupervised Skill Optimization From Natural Language Instructions

2609.38593v1 by Bo Ni, Li Li, Ryan A. Rossi, Franck Dernoncourt, Tyler Derr

Skills are external artifacts that Large Language Models (LLMs) consume at inference time to improve their performance on specialized domains by incorporating relevant procedural and domain knowledge. Expert-authored skills are expensive to produce, and the resulting artifacts are not optimized for the specific model that consumes them, whose failure modes can vary with version, scale and training. In addition, emerging tasks may fall outside the scope of existing skill libraries, creating a need to develop new skills before curated training data become available. Recent works have explored automated skill optimization through reflection, but they require a curated, in-distribution training set, which users might not always have. To address these limitations, we present Prompt2Skill, a framework that builds skills from natural-language task description alone. From the prompt, the system derives a task specification, discovers or synthesizes datasets, and refines the skill in a closed loop of reflective editing. Across four domains spanning question answering, reading comprehension, spreadsheet manipulation, and mathematical reasoning, Prompt2Skill consistently outperforms the direct prompting baseline, achieving an average improvement of 10.8 across open-source and frontier models.

摘要:技能是大型語言模型(LLMs)在推理時消耗的外部產物,通過整合相關的程序和領域知識來提高其在專業領域的表現。專家撰寫的技能製作成本高昂,且所產生的產物並未針對消耗它們的特定模型進行優化,而這些模型的失效模式可能因版本、規模和訓練而異。此外,新興任務可能超出現有技能庫的範疇,因此在經過策劃的訓練數據可用之前,需要開發新的技能。最近的研究探討了通過反思進行自動化技能優化,但這需要一個策劃的、在分佈內的訓練集,而用戶可能並不總是擁有。為了解決這些限制,我們提出了Prompt2Skill,一個僅從自然語言任務描述構建技能的框架。系統從提示中推導出任務規範,發現或合成數據集,並在反思編輯的閉環中精煉技能。在涵蓋問題回答、閱讀理解、電子表格操作和數學推理的四個領域中,Prompt2Skill始終超越直接提示基線,平均提升達到10.8,適用於開源和前沿模型。

Medical

Publish Date Title Authors Homepage Code
2026-10-01 A Comparative Explainability Framework for DeBERTa-v3 in Zero-Shot Medical Abstract Classification Javier Diaz Esteban-Herreros et.al. 2610.02116v1 null
2026-10-01 Can AI Oversight Be Zero Knowledge? Alessandro Chiesa et.al. 2610.01995v1 null
2026-10-01 Counterfactual Auditing of Bias in Open-Source Large Language Models for Clinical Triage Manar Aljohani et.al. 2610.01963v1 null
2026-10-01 A rubric landscape for evaluating clinical reasoning in large language models: what exists, what is missing, and what needs to be combined Zhangshu Joshua Jiang et.al. 2610.01938v1 null
2026-10-01 Walking the Embedding Space: Datastore Extraction from Multimodal RAG Maria Carmen Jica et.al. 2610.01871v1 null
2026-10-01 On the Divergence of Accuracy and Mechanism Consistency in Time Series World Models Haochen Zhang et.al. 2610.01842v1 null
2026-10-01 iADD: Improving Alignment and Diversity in Diffusion Policy Optimization Ashok Prasad Neupane et.al. 2610.01789v1 null
2026-10-01 OpenMTB-Audit: Exposing Over-Refusal and Clinical Expert Perspectives in LLM-Based Molecular Tumor Board Safety Evaluation Negin Ashrafi et.al. 2610.01497v1 null
2026-10-01 A Multi-Agent LLM Framework for Personalized Health Checkup Interpretation and Guidance HyungJun Kim et.al. 2610.01451v1 null
2026-10-01 Generation Provenance Before Behavior Attribution: Auditing Synthetic Speech Research Objects Sidi Chang et.al. 2610.01378v1 null
2026-10-01 An ontology for cross-sectoral crisis management: core and public health modules Aldo Gangemi et.al. 2610.01326v1 null
2026-10-01 Model validation in machine learning: A scenario-based guide from hold-out splits to nested group cross-validation in biomedical and applied research Mehmet Baygin et.al. 2610.01284v1 null
2026-10-01 When Does Exercise-Specific Joint Selection Help? An Audit of Evaluation and Control Design Haotian Chen et.al. 2610.01188v1 null
2026-10-01 CineMR: Tool-Integrated Vision-Language Reasoning for Quantitative Cardiac MRI Assessment Kunyang Li et.al. 2610.01166v1 null
2026-10-01 A Matched-Budget Audit Framework for Recaptioned Image-Text Supervision Distributions Giyeong Oh et.al. 2610.00952v1 null
2026-09-30 Backdoor Purification for LoRA-Tuned LLMs via Null-Space Projection Jianwei Li et.al. 2610.00685v1 null
2026-09-30 Worse Together: How Performance Breaks Down in Multi-User Multi-Agent Teams Sahan Paliskara et.al. 2610.00583v1 null
2026-09-30 Can LLMs Reason Over Long Horizons? An Empirical Evaluation of Context Strategies for Longitudinal Clinical Reasoning Taye Akinrele et.al. 2610.00562v1 null
2026-09-30 Unlearnable, or Unmeasured? On the Reliability of Difficulty Labels in RLVR Chandak Chakma et.al. 2609.40115v1 null
2026-09-30 GateSPINE: Gated Cross-View Fusion for Lumbar Spine MRI Report Generation Hoang Nguyen Van et.al. 2609.40091v1 null
2026-09-30 Overview of BioASQ 2026: The fourteenth BioASQ Challenge on Large-Scale Biomedical Semantic Indexing and Question Answering Anastasios Nentidis et.al. 2609.39975v1 null
2026-09-30 Scores That Hold, Benchmarks That Leak: Measuring Dataset Contamination in Public Brain-Tumor MRI Classification Bhanu Prakash Vangala et.al. 2610.00421v1 null
2026-09-30 Trust Is Not a Score: Runtime Assurance Contracts for High-Risk AI Agents Serhii Zabolotnii et.al. 2609.39717v1 null
2026-09-30 Interpretable Synthetic Medical Tabular Data Generation for Clinical Decision Support Using Fuzzy Cognitive Maps Michael Vasilakakis et.al. 2610.00391v1 null
2026-09-30 CAMOS: Coupled Oscillatory State-Space Model for Multimodal Clinical Time-Series Maxx Richard Rahman et.al. 2609.39484v1 null
2026-09-30 Towards Trustworthy AI for Glioma Diagnosis: A Task-Aware Evaluation of Uncertainty Quantification Gonzalo Esteban Mosquera Rojas et.al. 2609.39429v1 null
2026-09-30 EHR-RobustGym: Benchmarking and Training Agents for Robust Clinical Reasoning Yitong Qiao et.al. 2609.39371v1 null
2026-09-30 Structure vs. Chain-of-Thought: Evaluating LLM Criteria Extraction for Depression Severity Xinkai Chen et.al. 2609.39049v1 null
2026-09-30 An Uncertainty-Guided Digital Twin Framework for Online Adaptive Proton Therapy in Head and Neck Cancer: A Feasibility Study Yizhou Wu et.al. 2609.39010v1 null
2026-09-30 Scoring Higher, Answering Worse: Mitigating Reward Hacking in Rubric-Based RL via Protocol-Level Rubrics Maoqi Liu et.al. 2609.38847v1 null
2026-09-30 CRAFT: Causal Responsibility and Failure Tracing in Medical Vision Language Models Chunzheng Zhu et.al. 2609.38810v1 null
2026-09-29 Sense and Sensitivity: Benchmarking LLM Clinical Triage Recommendations with Physician Experts Abinitha Gourabathina et.al. 2609.38600v1 null
2026-09-29 Defining and Categorising Human-AI Interactions in Clinical Trials: A Multidimensional Human-AI Classification Approach Sandra Woolley et.al. 2609.38559v1 null
2026-09-29 Personalized State-Transition-Aware Memory for Clinical Agents Maryam Haghifam et.al. 2609.38490v1 null
2026-09-29 KlinikeBench: Evaluating Language Models Beyond Diagnostic Accuracy Xueting Fang et.al. 2609.38480v1 null
2026-09-29 PrivMeSA: Privacy-Aware Self-Evolving Multi-Agent System for Medicine via Local-Remote LLM Collaboration Dannong Wang et.al. 2609.38458v1 null
2026-09-29 Colorectal Cancer Segmentation with Adaptive Augmentation and Multi-Resolution Ensemble Models Ümit Mert Çağlar et.al. 2609.38419v1 null
2026-09-29 Aegis: Generative Gradient Masking for Privacy-Preserving Medical Federated Learning Chaoyu Zhang et.al. 2609.38339v1 null
2026-09-29 A Proposed Rubric for Evaluating Expressed Clinical Reasoning in Large Language Model Responses Zhangshu Joshua Jiang et.al. 2609.37788v3 null
2026-09-29 Multi-Site Real-World Performance of Commercial AI for Pulmonary and Incidental Pulmonary Embolism Detection Aawez Mansuri et.al. 2609.37750v1 null
2026-09-29 Spatiotemporal Hyperedges for EEG Seizure Detection and Prediction Hyunju Kim et.al. 2609.37730v1 null
2026-09-29 Correct, Don't Delete: Mitigating Emergent Misalignment with Corrective Supervision Jacob Epifano et.al. 2609.37624v1 null
2026-09-29 ReLMem: Learning Recurrent Memory for Longitudinal EHR Modeling Zijie Meng et.al. 2609.37587v1 null
2026-09-29 Raw Imagery Impacting Your AI: Should You Care? Adrien Dorise et.al. 2609.38265v1 null
2026-09-29 Do Agent Benchmarks Do What They Say? An Executable-Contract Audit of Tool-Using Agent Environments Rohith Reddy Bellibatlu et.al. 2609.37315v1 null
2026-09-29 Information Bottleneck-Guided Adaptive Hypergraph Transformer for Brain Disease Diagnosis Jingxi Feng et.al. 2609.37220v1 null
2026-09-29 Physics-Informed Multi-Agent Coordination for Hospital Patient Flow Optimization Guoqing Zhang et.al. 2609.37022v1 null
2026-09-29 STAR-GRPO: Canonical Anchoring and Reliability-First Advantages against Representation-Dependent Reward Hacking Wan Tian et.al. 2609.36900v1 null
2026-09-29 Automated Screw Planning for Reduced Pelvic Fractures Based on Statistical Shape Models and Deep Learning Yang Gao et.al. 2609.36847v1 null
2026-09-29 How Medical VLMs Underutilize Their Vision Encoders: A Dermatology Perspective Janet Wang et.al. 2609.36557v1 null
2026-09-29 Reliability Testing of Medical Model Performance under Distributed Deployment Yifei Wang et.al. 2609.36525v1 null
2026-09-29 BRIDGE: Bilevel Retrieval-Credit-Aware Agentic Reinforcement Learning Quan Xiao et.al. 2609.36505v1 null
2026-09-28 ARCagent: An Adaptive Retrieval Calibration Agent for Clinical Question Answering Yuyan Chen et.al. 2609.36392v1 null
2026-09-28 Quantization Enables Private Dense Retrieval against Malicious Service Providers Louis Tremblay Thibault et.al. 2609.36376v1 null
2026-09-28 ThuRunel: Dynamic Decoupling for Structured Advisory Dialogue Yuyan Chen et.al. 2609.36340v1 null
2026-09-28 SCOUT: Synergizing Reasoning and Tool-Use for Computer-Use Safety Jianxing Chen et.al. 2609.36201v1 null
2026-09-28 PHASE: A Physiology-Guided Hierarchical Foundation Model for Intracranial EEG Yipeng Zhang et.al. 2609.36087v2 null
2026-09-28 IMC-CLINIC: Coupled Loss-Informed Newton Iterations for Clipping in Analog In-Memory Computing Yung-Chin Chen et.al. 2609.35586v1 null
2026-09-28 RareDx: Controlled Knowledge Integration and Graph-Grounded Policy Optimization for Rare-Disease Diagnosis Bo Zhang et.al. 2609.35549v2 null
2026-09-28 CLIMB: A Clinical Multimorbidity Benchmark for Diagnosing Co-occurring Conditions through Multiturn Conversations Yusuf Kesmen et.al. 2609.35462v1 null
2026-09-28 A decision-support system applied to Law: Reasoning and explainability of the decision Jeremy Bouche-Pillon et.al. 2609.35370v1 null
2026-09-28 Training-Free Clinical Reasoning through Medical Ontologies and Cognitive Mapping: A Symbolic-Probabilistic Knowledge Graph Framework Surajit Das et.al. 2609.35298v1 null
2026-09-28 CarveMix-RC: Addressing Rare-Class Imbalance Through Lesion-Aware Synthetic Augmentation for Brain Metastasis Segmentation Md Shibly Sadique et.al. 2609.35195v1 null
2026-09-28 DoAtlas-2: A Foundation for Self-Evolving Causal Biomedical Discovery Yulong Li et.al. 2609.35107v1 null
2026-09-28 VD-DeepStack: Bridging Visual Comparison and Language Reasoning for Few-Shot Anomaly Detection Mengyang Zhao et.al. 2609.34949v1 null
2026-09-28 Almost Human, Except When It Matters: VoxParity and the Decisions a Voice Should Change Bhavik Mangla et.al. 2609.35922v1 null
2026-09-28 Nociception as a Control Primitive: Afferent Channels and Nociceptive Memory for Agents Deployed in One Body Wolfgang Maass et.al. 2609.34840v1 null
2026-09-28 Applying Language Models in Clinical Medicine: Recent Trends and Perspectives Erik Aerts et.al. 2609.34780v2 null
2026-09-28 ResonAct: Streaming Metrics for Runtime Diagnosis and Self-Healing in Multi-Agent Systems Tarun Chintada et.al. 2609.34701v1 null
2026-09-28 SentZero: An Enhanced Sentence-Centric Vision-Language Pretraining for Multi-Task Zero-Shot Chest X-Ray Analysis Hangyul Yoon et.al. 2609.34479v1 null
2026-09-28 VL-AcneSeg: A Vision-Language Framework for Region-Aware Acne Lesion Segmentation Sukju Oh et.al. 2609.34472v1 null
2026-09-28 Evolving Support Priorities in Empathetic Reinforcement Learning Pengyu Huang et.al. 2609.34249v1 null
2026-09-28 Unknown is not normal: separating language-model extraction from rule-based decision logic for clinical risk scores Nicolás Vera Zúñiga et.al. 2609.34112v1 null
2026-09-28 The Devil is in the Spectrum Bias: Spectrum-Balanced Feature Matching for Robust Representation Distillation Kuniaki Saito et.al. 2609.34106v1 null
2026-09-28 TRACE: Expert-Aligned ECG Representation Learning with Rigorous Benchmarking and Real-World Validation in Acute Cardiac Care Lovely Yeswanth Panchumarthi et.al. 2609.34088v1 null
2026-09-28 Who Gets a Token, and What Does It Carry? Unequal Name Support and Concept Access in Large Language Models Mir Tafseer Nayeem et.al. 2609.34065v1 null
2026-09-28 Large Language Models for Structured Clinical Data Analysis: Dual-Agent Grounding and Validation Erfan D. Dehkalani et.al. 2609.34039v1 null
2026-09-27 Jev in Medicine: A Benchmark Evaluation Alfredo Madrid-García et.al. 2609.34024v2 null
2026-09-27 EHRAdapt: Adapting Pretrained Language Models to Electronic Health Records with Semantic Priors for Rare Clinical Events Andre R Goncalves et.al. 2609.34007v1 null
2026-09-27 Is your uncertainty map wrong, or is its target? Exact diagnostics for the Tweedie diagonal, and a gradient-free alternative Vicent Ribas et.al. 2609.33786v1 null
2026-09-27 BIRD: Distilling Decision Boundaries into Rationales for MLLM Adaptation Anglin Liu et.al. 2609.33713v1 null
2026-09-27 PPG-LM: A Photoplethysmography-Language Model with Multi-Level Clinical Alignment Xiaoda Wang et.al. 2609.33516v1 null
2026-09-27 Federated Multi-Modal Human Activity Recognition using Multi-Agent Reinforcement Learning Debasmita Dey et.al. 2609.33492v1 null
2026-09-27 A Cheap Verifier is Good Enough: LLM Post-training is Robust to Erroneous Rewards Andreas Plesner et.al. 2609.33467v1 null
2026-09-27 MAC-Net: A Multi-Task Deep Learning Framework for Modeling Cognitive Function From Task-Based fMRI Md. Tanvir Rahman et.al. 2609.33440v1 null
2026-09-27 Temporal Graph Learning of Wearable Actigraphy and Sleep Traces for Modelling Adolescent Crystallized Intelligence Md. Tanvir Rahman et.al. 2609.33428v1 null
2026-09-27 Explainable Deep Learning of Resting-State Functional Connectomes Reveals Network Biomarkers of Adolescent Intelligence Md. Tanvir Rahman et.al. 2609.33422v1 null
2026-09-27 QuPID: Quantum Parameter-Efficient Input-Dependent Retrieval Adaptation for Medical RAG Hyojun Ahn et.al. 2609.33351v1 null
2026-09-27 CHI: A Composite Hallucination Index Unifying Entity, Relation, and Quantity Dimensions for Summarization Evaluation Praveenkumar Katwe et.al. 2609.33343v1 null
2026-09-27 The Error You See Is Not the Error You Made: Progression-aware Reasoning Origin for Reasoning Error Localization Yiguo Wang et.al. 2609.33297v1 null
2026-09-27 CORTEX: A Verified Experience Layer for Generalist Agents Garapati Keerthana et.al. 2609.33260v1 null
2026-09-27 FOCUS: Benchmarking Retinal Model Generalization from Foundation Vision Encoders to Multimodal LLMs David Restrepo et.al. 2609.33158v1 null
2026-09-27 MedRouter: Demystifying Knowledge Differences Across Medical LLMs for Routing-Based Reasoning Lang Cao et.al. 2609.33119v1 null
2026-09-27 ECG-Scroll: A Long-Horizon, Streaming Benchmark and Agent Environment for Interpretation of Ambulatory Electrocardiograms Haitao Li et.al. 2609.33117v1 null
2026-09-27 SemReward-VL: Semantic Reward-Guided Video-Language Adaptation for Developmental Behavior Assessment De Jiang et.al. 2609.33082v1 null
2026-09-27 NutriVision: Ingredient-Conditioned Fusion and Prediction for Single-Image Food Nutrition Estimation Aman Kumar et.al. 2609.33076v1 null
2026-09-26 TCMQA: A 38K-Question Traditional Chinese Medicine Benchmark with a Licensed-Practitioner Reference Tzu-Heng Huang et.al. 2609.33014v1 null
2026-09-26 Multimodal LLMs Outperform Pathology Foundation Models in Cross-Domain Histological Similarity Yishu Zhang et.al. 2609.32876v2 null
2026-09-26 Counterfactual Self-Evolving Agents for Evidence-Grounded Reasoning Xing Han et.al. 2609.32870v1 null
2026-09-26 FinancialAuditBench: Benchmark Construction under Differential Privacy Using Real-World Priors Jerry Huang et.al. 2609.32835v1 null

Abstracts

A Comparative Explainability Framework for DeBERTa-v3 in Zero-Shot Medical Abstract Classification

2610.02116v1 by Javier Diaz Esteban-Herreros, David Muñoz-Valero, Raquel Martínez-España, Jose M. Juarez, Juan Moreno-Garcia

A comparative explainability framework is presented to audit DeBERTa-v3 under zero-shot classification of medical abstracts. The work addresses the disagreement problem in Explainable Artificial Intelligence, where different attribution methods produce divergent explanations for the same input and prediction. A natural language inference engine is implemented over the Medical Abstracts corpus with five enriched hypotheses per diagnostic category and a balanced sample of one thousand texts per class. Five explanation methods are compared: SHAP and LIME as model-agnostic approaches, occlusion and Input x Gradient as deep-learning-specific approaches, and Attention x Gradient as a transformer-specific approach. Explanations are standardized through top-token attribution, and pairwise agreement is quantified using the Jaccard index. High predictive accuracy is achieved across well-defined clinical domains, whereas performance degrades under high semantic ambiguity. Explanatory stability directly mirrors predictive certainty, exhibiting strong convergence in univalent categories and a marked drop under diagnostic uncertainty. Furthermore, qualitative error auditing uncovers three systemic failure mechanisms: lexical hypersensitivity, semantic overlap, and loss of attribution coherence. The results support the combined use of several explanation methods and quantitative agreement metrics when auditing transformer-based models in medical text classification, and suggest prioritizing specific clinical ontologies over broad diagnostic labels.

摘要:比較可解釋性框架被提出以審計 DeBERTa-v3 在醫學摘要的零樣本分類下。這項工作解決了可解釋人工智慧中的不一致問題,即不同的歸因方法對相同的輸入和預測產生不同的解釋。自然語言推理引擎在醫學摘要語料庫上實施,每個診斷類別有五個增強的假設,並且每個類別有一千篇文本的平衡樣本。比較了五種解釋方法:SHAP 和 LIME 作為模型無關的方法,遮蔽和輸入 x 梯度作為深度學習特定的方法,以及注意力 x 梯度作為Transformer特定的方法。通過頂部標記歸因標準化解釋,並使用 Jaccard 指數量化成對一致性。在明確定義的臨床領域中實現了高預測準確性,而在高語義模糊性下性能下降。解釋穩定性直接反映預測確定性,在單值類別中顯示出強烈的收斂,並在診斷不確定性下顯著下降。此外,定性錯誤審計揭示了三種系統性失效機制:詞彙過敏、語義重疊和歸因一致性的喪失。結果支持在醫學文本分類中審計基於Transformer的模型時,結合使用幾種解釋方法和定量一致性指標,並建議優先考慮特定的臨床本體論而非廣泛的診斷標籤。

Can AI Oversight Be Zero Knowledge?

2610.01995v1 by Alessandro Chiesa, Ziyi Guan, Burcu Yildiz

AI systems increasingly produce outputs from confidential data, such as a fitness-for-duty assessment from medical records or the predicted properties of a drug candidate from its secret structure. It is important to verify that such outputs are correct without revealing the underlying data. A recent line of work studies verification of AI outputs via interactive proofs and debate for oracle-aided computation, where correctness may depend on an oracle such as human judgment, a physical experiment, or the web. These works focus on verification by a verifier that runs much faster than the computation. However, such efficient verification is impossible for general oracle-aided computation, and these works therefore rely on additional assumptions. We focus instead on privacy: allowing the verifier to run in time polynomial in the computation, we ask whether interactive arguments for oracle-aided computation can be zero knowledge, so that the verifier learns nothing about the confidential data beyond the correctness of the output. We prove that, in general, they cannot. In the random oracle model, there are no zero-knowledge proofs for all oracle-aided computations, even if both the prover and the verifier are allowed to run much longer than the computation itself. The impossibility extends to debate, a canonical model for scalable oversight. On the positive side, we show that if the oracle attaches a cryptographic signature to each of its answers, then every oracle-aided computation can be verified in zero knowledge with an efficient prover and verifier, assuming only collision-resistant hash functions. Beyond privacy, this also gives an alternative approach to scalable oversight that relies neither on an honest opponent, as in debate, nor on the robustness of the computation, as in prior single-prover protocols.

摘要:AI 系統越來越多地從機密數據中產生輸出,例如從醫療記錄中進行的適任性評估或從其秘密結構中預測的藥物候選物的性質。驗證這些輸出是否正確而不透露底層數據是很重要的。最近的一系列研究通過互動證明和辯論研究 AI 輸出的驗證,用於有 oracle 輔助的計算,其中正確性可能依賴於 oracle,例如人類判斷、物理實驗或網絡。這些研究專注於由運行速度遠快於計算的驗證者進行的驗證。然而,對於一般的有 oracle 輔助計算,這樣的高效驗證是不可能的,因此這些研究依賴於額外的假設。我們則專注於隱私:允許驗證者在計算的多項式時間內運行,我們詢問有 oracle 輔助計算的互動論證是否可以是零知識的,以便驗證者不會學到關於機密數據的任何信息,除了輸出的正確性。我們證明,通常情況下,它們是不可能的。在隨機 oracle 模型中,對於所有有 oracle 輔助的計算,沒有零知識證明,即使證明者和驗證者都被允許運行的時間遠超過計算本身。這種不可能性擴展到辯論,這是一個可擴展監督的典型模型。從積極的一面來看,我們展示了如果 oracle 為其每個答案附加加密簽名,那麼每個有 oracle 輔助的計算都可以在零知識中進行驗證,並且有高效的證明者和驗證者,只假設碰撞抗性哈希函數。除了隱私之外,這還提供了一種可擴展監督的替代方法,既不依賴於誠實的對手(如辯論中),也不依賴於計算的穩健性(如以前的單證明者協議中)。

Counterfactual Auditing of Bias in Open-Source Large Language Models for Clinical Triage

2610.01963v1 by Manar Aljohani, Brandon Ho, Kenneth McKinley, Dennis Ren, Xuan Wang

Emergency department (ED) triage is a high-stakes prioritization task in which demographic, socioeconomic, and system-context information may improperly influence acuity assignment. Although open-source large language models (LLMs) are increasingly considered for local and privacy-preserving clinical decision support, it remains unclear how counterfactual bias varies across model families, sizes, medical-domain models, and domain-adapted models. We present a comparative counterfactual audit of ten open-source LLMs for pediatric Emergency Severity Index (ESI) prediction. Starting from real and handbook-style clinical vignettes, we construct paired counterfactual variants that change only one injected demographic, socioeconomic, healthcare-access, behavioral, social, or system-context variable while holding the clinical presentation fixed. Models include Qwen2.5-7B, Qwen2.5-14B-Instruct, a QLoRA fine-tuned Qwen2.5-7B, MedGemma variants, MedLLaMA2-7B, GPT-OSS-20B, and GPT-OSS-120B. We measure any counterfactual shift, undertriage, overtriage, shifts greater than one ESI level, mean shift, and mean absolute shift. Counterfactual sensitivity varied substantially and did not consistently decrease with larger model size or medical-domain pretraining. The fine-tuned Qwen2.5-7B showed the lowest overall sensitivity, with a 5.27% any-shift rate and mean absolute shift of 0.0534, versus 16.02% and 0.1706 for the base model. Several larger or medical-domain models showed more significant shifts. Stratified and correlation analyses further revealed clinically important directionality and shared failure patterns hidden by aggregate rates. These findings support counterfactual auditing as a lightweight, clinically interpretable framework for comparing fairness risks in open-source LLMs before clinical deployment.

摘要:急診部(ED)分診是一項高風險的優先排序任務,其中人口統計、社會經濟和系統背景信息可能不當影響急性程度的分配。儘管開源大型語言模型(LLMs)越來越被考慮用於本地和隱私保護的臨床決策支持,但目前尚不清楚反事實偏見在不同模型家族、大小、醫療領域模型和領域適應模型之間的變化情況。我們對十個開源LLM進行了針對兒科緊急嚴重性指數(ESI)預測的比較反事實審計。從真實和手冊風格的臨床小插曲開始,我們構建了配對的反事實變體,僅改變一個注入的人口統計、社會經濟、醫療訪問、行為、社會或系統背景變量,同時保持臨床表現不變。模型包括Qwen2.5-7B、Qwen2.5-14B-Instruct、經過QLoRA微調的Qwen2.5-7B、MedGemma變體、MedLLaMA2-7B、GPT-OSS-20B和GPT-OSS-120B。我們測量任何反事實變化、低估分診、過度分診、超過一個ESI級別的變化、平均變化和平均絕對變化。反事實敏感性變化顯著,且不一致地隨著模型大小或醫療領域預訓練的增大而減少。經過微調的Qwen2.5-7B顯示出最低的整體敏感性,任何變化率為5.27%,平均絕對變化為0.0534,而基礎模型則為16.02%和0.1706。幾個較大或醫療領域模型顯示出更顯著的變化。分層和相關分析進一步揭示了臨床上重要的方向性和由聚合率隱藏的共同失敗模式。這些發現支持反事實審計作為一種輕量級、臨床可解釋的框架,用於在臨床部署前比較開源LLM中的公平風險。

A rubric landscape for evaluating clinical reasoning in large language models: what exists, what is missing, and what needs to be combined

2610.01938v1 by Zhangshu Joshua Jiang, Zina Ibrahim, James T. Teo

Exam-style accuracy does not establish whether large language models (LLMs) reason well over clinical records. We define clinical reasoning as integrating and updating evidence across time and sources to form, revise and justify a patient's problem representation and a defensible plan. This structured narrative review maps three literatures: medical education assessment instruments, clinical LLM benchmarks published from 2023 onwards, and general-domain methods for evaluating long-form generation. We examine six dimensions: problem representation, temporal synthesis, differential and management reasoning, counterfactual reasoning, calibrated uncertainty, and reasoning faithfulness. Preprints are included and flagged. No single instrument covers all six dimensions. Problem representation and differential or management reasoning are reasonably covered, although reliability varies by instrument and setting. TIMER-Eval targets temporal synthesis, and ER-Reason assesses sequential diagnostic belief updating. Dedicated uncertainty and counterfactual evaluations are emerging, but their applicability to longitudinal free-text reasoning remains limited. Factual completeness is well theorised in general-domain evaluation, with early clinical evidence of important omissions. Faithfulness remains the weakest dimension, with one identified clinical causal-ablation study on multiple-choice questions. Existing tools should be combined through binary rubric items, separate completeness and correctness scores, case-specific importance weighting with non-compensable safety caps, temporal order-consistency checks, and chance-corrected reliability reporting. Further design work is needed for calibrated uncertainty, counterfactual reasoning and faithfulness over longitudinal free-text records. This review provides a design rationale, not a validated instrument.

摘要:考試風格的準確性並不能確定大型語言模型(LLMs)在臨床記錄上是否能夠進行良好的推理。我們將臨床推理定義為整合和更新跨時間和來源的證據,以形成、修訂和辯護病人的問題表述及可辯護的計劃。
這篇結構化的敘述性回顧映射了三個文獻領域:醫學教育評估工具、2023年以來發表的臨床LLM基準,以及評估長篇生成的一般領域方法。我們檢視了六個維度:問題表述、時間綜合、差異和管理推理、反事實推理、校準的不確定性,以及推理的忠實性。預印本已被納入並標記。
沒有單一的工具涵蓋所有六個維度。問題表述以及差異或管理推理的覆蓋相對合理,儘管可靠性因工具和環境而異。TIMER-Eval 針對時間綜合,而 ER-Reason 評估連續的診斷信念更新。專門的不確定性和反事實評估正在出現,但它們對於縱向自由文本推理的適用性仍然有限。事實的完整性在一般領域評估中有良好的理論基礎,並且早期臨床證據顯示出重要的遺漏。忠實性仍然是最薄弱的維度,其中有一項針對多選題的臨床因果消融研究被識別。
現有工具應通過二元評分項目、分開的完整性和正確性分數、特定案例的重要性加權(帶有不可補償的安全上限)、時間順序一致性檢查以及機會修正的可靠性報告進行結合。對於校準的不確定性、反事實推理和縱向自由文本記錄的忠實性,還需要進一步的設計工作。這篇回顧提供了一個設計的理由,而不是一個經過驗證的工具。

Walking the Embedding Space: Datastore Extraction from Multimodal RAG

2610.01871v1 by Maria Carmen Jica, Ali Satvaty, Suzan Verberne, Fatih Turkmen

Multimodal Retrieval-Augmented Generation (MRAG) has emerged as a reliable and cost-effective technique of grounding the generative capabilities of Multimodal Large Language Models (MLLMs) into relevant, up-to-date, external knowledge. Despite presenting several benefits, such as reducing hallucinatory behavior, they also introduce new attack surfaces, including leakage of private information and vulnerabilities against data extraction attacks. In this paper, we introduce $\immrag$, an adaptive and automatic data extraction attack procedure operating in a black box setting against \emph{image-returning} MRAG, a configuration in which the retrieved visual artifact is itself the response. Each query blends an attacker-held shadow image with an image already recovered from the system, and relevance-weighted resampling steers subsequent queries towards regions of the embedding space that still yield novel retrievals. Unlike current extraction attacks that aim to persuade the model towards data leakage by placing a malicious query as a textual prompt, $\immrag$ embeds the malicious instructions inside a user-given input image. We evaluate $\immrag$ on three plausible and distinct real-world scenarios: medical assistant, document-focused helper and general purpose tool. The experiments involve the study of the effectiveness of the attack on multiple CLIP-family retrievers, as well as the impact of various generators. A single 2500-query run reconstructs up to 611 distinct radiology images, 566 document scans and 416 general-purpose images under local-feature correspondence, and reaches up to $5.6\times$ as many distinct datastore items as a non-adaptive baseline. Our results show the urgent need for safeguards specifically designed for multimodal data.

摘要:多模態檢索增強生成(MRAG)已成為將多模態大型語言模型(MLLMs)的生成能力與相關的、最新的外部知識相結合的一種可靠且具成本效益的技術。儘管它提供了幾個好處,例如減少幻覺行為,但它們也引入了新的攻擊面,包括私密信息洩漏和對數據提取攻擊的脆弱性。 在本文中,我們介紹了 $\immrag$,這是一種適應性和自動化的數據提取攻擊程序,針對 \emph{圖像返回} MRAG 在黑箱環境中運作,這是一種檢索的視覺工件本身就是回應的配置。每個查詢將攻擊者持有的影像與系統中已恢復的影像混合,並且相關性加權重採樣引導後續查詢朝向仍能產生新穎檢索的嵌入空間區域。與目前旨在通過將惡意查詢作為文本提示來說服模型進行數據洩漏的提取攻擊不同,$\immrag$ 將惡意指令嵌入用戶提供的輸入影像中。我們在三個合理且不同的現實場景中評估了 $\immrag$:醫療助手、文件專注助手和通用工具。實驗涉及對多個 CLIP 家族檢索器的攻擊有效性以及各種生成器的影響進行研究。一次 2500 次查詢的運行重建了多達 611 幅不同的放射學影像、566 幅文件掃描和 416 幅通用影像,根據局部特徵對應,並達到高達 $5.6\times$ 的不同數據庫項目數量,相較於非適應性基準。我們的結果顯示出對專門為多模態數據設計的安全措施的迫切需求。

On the Divergence of Accuracy and Mechanism Consistency in Time Series World Models

2610.01842v1 by Haochen Zhang, Jiaheng Guo, Zhen Xu, Zachary Plotkin, Nicholas Konz, Zhen Tan, Tianlong Chen

A time series world model (TSWM) predicts a controlled system's state from its observed history and planned actions and exogenous inputs. Current approaches build forecasters with actions as covariates, trained and evaluated on prediction error under the executed plan. Yet world models compare unexecuted plans, but their responses to changed plans remain untested. We ask which design choices matter and whether accurate forecasters respond to changed plans as real systems do. We address both with a formalization and benchmark. The formalization separates state, actions and exogenous inputs, distinguishes continuous, mode and event actions, and introduces mechanism consistency, a metric built on declared action-state relations with known directions, such as a vasopressor raising blood pressure: it checks whether shifting an action moves the forecast in the declared direction. The benchmark consolidates eight public datasets with real actions from engineered infrastructure and clinical care, varying prediction space, plan fusion and plan encoding across seven backbones and five seeds. First, a frozen latent prediction space lowers MAE by 9.9% over observation space and gated output fusion lowers it by 12.7% over input concatenation on average, with both improving all eight datasets; temporal plan encoding changes average MAE by at most 2.2%. Second, prediction error and mechanism consistency diverge: the lowest-error configuration is at or below chance in consistency on four of five datasets with declared mechanisms, and no design choice avoids this. Finally, directional supervision, a loss penalizing the wrong-signed part of the response to a shifted action, significantly raises consistency on penalized mechanisms with no change in MAE. Together they give TSWMs a recipe: a frozen latent space and output-side fusion for accuracy, and a training objective for mechanism consistency.

摘要:時間序列世界模型 (TSWM) 從其觀察歷史、計畫行動和外部輸入預測受控系統的狀態。當前的方法建立了以行動為協變數的預測器,這些預測器在執行計畫下的預測誤差上進行訓練和評估。然而,世界模型比較未執行的計畫,但對於變更計畫的反應仍未經測試。我們詢問哪些設計選擇是重要的,以及準確的預測器是否像真實系統一樣對變更計畫做出反應。我們通過形式化和基準來解決這兩個問題。形式化將狀態、行動和外部輸入分開,區分連續、模式和事件行動,並引入機制一致性,這是一種基於已知方向的聲明行動-狀態關係構建的指標,例如一種升壓藥提高血壓:它檢查改變行動是否將預測移動到聲明的方向。基準整合了八個來自工程基礎設施和臨床護理的公共數據集,這些數據集中有真實行動,並在七個骨幹和五個種子中變化預測空間、計畫融合和計畫編碼。首先,凍結的潛在預測空間使 MAE 在觀察空間上降低了 9.9%,而門控輸出融合使其在輸入串接上平均降低了 12.7%,兩者都改善了所有八個數據集;時間計畫編碼的變化使平均 MAE 變化最多為 2.2%。其次,預測誤差和機制一致性出現分歧:在五個具有聲明機制的數據集中,最低誤差配置在一致性上與隨機相同或更低,且沒有任何設計選擇能避免這一點。最後,方向性監督,一種對於對移動行動的反應中錯誤符號部分進行懲罰的損失,顯著提高了懲罰機制的一致性,且 MAE 沒有變化。這些共同為 TSWM 提供了一個配方:凍結的潛在空間和輸出側融合以提高準確性,以及一個針對機制一致性的訓練目標。

iADD: Improving Alignment and Diversity in Diffusion Policy Optimization

2610.01789v1 by Ashok Prasad Neupane, Saugat Adhikari, Pramish Paudel, Ajad Chhatkuli, Danda Pani Paudel

Reinforcement learning based post training of diffusion models, such as Denoising Diffusion Policy Optimization (DDPO), optimizes a reverse diffusion process under a reward function. However, current approaches to reward optimizations do so at the cost of diversity and quality. In this paper, we provide better tradeoffs through careful theoretical considerations and method design. We analyze the theoretical framework and mathematically demonstrate that \emph{only-latter timestep} updates of diffusion model may be harmful for diversity contrary to the conclusions presented in a previous work. Additionally, we propose an incremental Feynman-Kac training based on strong theoretical foundations in order to achieve the best-yet alignment-diversity tradeoffs. We perform extensive experiments and compare our method against related diffusion policy optimization approaches in three different tasks and also provide strong ablations for each component, thus validating strong performance gains in both alignment and diversity.

摘要:基於強化學習的擴散模型後訓練,例如去噪擴散策略優化(DDPO),在獎勵函數下優化反向擴散過程。然而,當前的獎勵優化方法以多樣性和質量為代價。在本文中,我們通過仔細的理論考量和方法設計提供了更好的權衡。我們分析了理論框架,並數學上證明擴散模型的\emph{僅後時間步}更新可能對多樣性有害,這與之前工作的結論相反。此外,我們提出了一種基於強大理論基礎的增量費曼-卡克訓練,以實現最佳的對齊-多樣性權衡。我們進行了廣泛的實驗,並在三個不同的任務中將我們的方法與相關的擴散政策優化方法進行比較,還為每個組件提供了強有力的消融實驗,從而驗證了在對齊和多樣性方面的顯著性能提升。

OpenMTB-Audit: Exposing Over-Refusal and Clinical Expert Perspectives in LLM-Based Molecular Tumor Board Safety Evaluation

2610.01497v1 by Negin Ashrafi, Jia Luo, Stacey M. Frumm, Roxana Daneshjou

Molecular tumor boards integrate genomic findings, clinical context, and therapeutic evidence to support precision oncology. As AI enters this workflow, a key safety challenge is distinguishing truly unsupported recommendations from evidence-supported options that still require oncologist review because of incomplete information, poor ECOG performance status, or other clinical caveats. We introduce OpenMTB-Audit, an open-source benchmark of 500 synthetic non-small cell lung cancer cases spanning five adversarial error categories and four safety labels: Supported, Partially Supported, Unsupported, and Insufficient Information. Across eight large language model configurations, we identify pervasive over-refusal: all LLM configurations failed to retain the Partially Supported label in 83.3-100% of true Partially Supported cases, achieving high aggregate safety scores through label collapse rather than clinically calibrated reasoning. To address this limitation, we developed MTB-AuditAgent, a deterministic seven-module framework separating evidence verification, missing-information detection, safety classification, and abstention. It reduces over-refusal to 6.7% and achieves 91.2% accuracy (95% CI: 88.6-93.6%). A two-oncologist annotation study found disagreement concentrated at the boundary between information sufficiency and treatment optimization, underscoring the need to preserve clinically meaningful distinctions.

摘要:分子腫瘤委員會整合基因組發現、臨床背景和治療證據,以支持精準腫瘤學。隨著人工智慧進入這一工作流程,一個主要的安全挑戰是區分真正不被支持的建議與仍需腫瘤醫生審查的證據支持選項,因為信息不完整、ECOG表現狀態不佳或其他臨床警告。我們介紹了OpenMTB-Audit,一個開源基準,包含500個合成的非小細胞肺癌案例,涵蓋五個對抗性錯誤類別和四個安全標籤:支持、部分支持、不支持和信息不足。在八種大型語言模型配置中,我們發現普遍的過度拒絕:所有LLM配置在83.3-100%的真實部分支持案例中未能保留部分支持標籤,通過標籤崩潰而非臨床校準推理獲得高整體安全分數。為了解決這一限制,我們開發了MTB-AuditAgent,一個確定性的七模塊框架,將證據驗證、缺失信息檢測、安全分類和放棄分開。它將過度拒絕降低到6.7%,並達到91.2%的準確率(95% CI:88.6-93.6%)。一項由兩位腫瘤醫生進行的標註研究發現,分歧集中在信息充分性和治療優化之間的邊界,強調了保留臨床上有意義的區別的必要性。

A Multi-Agent LLM Framework for Personalized Health Checkup Interpretation and Guidance

2610.01451v1 by HyungJun Kim, Taehan Lee, Soojin Cheon

Personalized interpretation of health checkup results requires reasoning across longitudinal records, medical knowledge, lifestyle guidance, and healthcare navigation. We present a multi-agent large language model (LLM) system that identifies multiple intents, maps each to a task-specific agent, executes them in parallel, and synthesizes their outputs. We compared answers generated in Single Agent and Multi Agent settings on 120 Korean compound queries combining two to four requirements, using synthetic health checkup records. The Multi Agent improved the weighted LLM-judge score from 1.695 to 1.797 (p = 0.027), and three additional LLM judges showed consistent improvements ($Δ$ = +0.111 to +0.186, all p < 0.05). The gains came from usefulness, consistency, and the handling of every requirement in compound queries, whereas numerical accuracy and grounding improved significantly under only one of the four judges and medical safety did not differ, and critical failures occurred at similar rates (Single Agent 15.0% vs. Multi Agent 13.3%). Two human evaluators preferred Multi Agent in 66.7% and 68.3% of pairwise comparisons. Multi Agent execution increased latency and cost by 1.31$\times$ and 2.02$\times$, respectively. In exploratory subgroup analyses, the improvement was concentrated in queries involving personal-record lookup.

摘要:個性化的健康檢查結果解釋需要跨越長期記錄、醫學知識、生活方式指導和醫療導航的推理。我們提出了一個多代理大型語言模型(LLM)系統,該系統識別多個意圖,將每個意圖映射到特定任務的代理,並平行執行它們,最後綜合其輸出。我們比較了在單代理和多代理設置下,使用合成健康檢查記錄對120個韓國複合查詢生成的答案,這些查詢結合了兩到四個需求。多代理將加權LLM評審分數從1.695提高到1.797(p = 0.027),另外三位LLM評審顯示出一致的改善($Δ$ = +0.111到+0.186,所有p < 0.05)。這些增益來自於有用性、一致性以及對複合查詢中每個需求的處理,而數值準確性和基礎資料僅在四位評審中的一位顯著改善,醫療安全則沒有差異,且重大失誤的發生率相似(單代理15.0%對多代理13.3%)。兩位人類評估者在66.7%和68.3%的成對比較中偏好多代理。多代理執行使延遲和成本分別增加了1.31$\times$和2.02$\times$。在探索性子群分析中,改善集中在涉及個人記錄查詢的問題上。

Generation Provenance Before Behavior Attribution: Auditing Synthetic Speech Research Objects

2610.01378v1 by Sidi Chang, Peiying Zhu

Attributing model behavior to synthetic training data requires knowing what produced each training item before estimating what that item caused. A waveform-label pair does not preserve this knowledge. We propose a generation-provenance substrate in which a synthetic research object binds source specification, generated content, waveform, target, fact requirements, quality signals, review lineage, and immutable manifest identity. Producer and selection mechanism determine evidentiary meaning; storage location and variable name do not. We audit this substrate in a private Japanese care-handoff pipeline. A 113-asset review population contains 1.552 hours of synthetic speech across six scenario families; all items have linked audio, transcripts, candidate notes, and fact checklists, but human evidence is selective and source-specific. Two faithful-only manifests are scenario-seed-disjoint and immutably versioned, while exact upstream attribution remains blocked by floating generator aliases, missing per-clip TTS and code stamps, and an unversioned checking prompt. We argue that generation provenance is necessary but not sufficient for behavior attribution: it defines the candidate causal graph and audit units, whereas contributive attribution still requires frozen training runs and intervention or influence evidence. The paper contributes a compact provenance contract, an audit protocol, and a bounded case study for synthetic-data attribution; controlled research access may be offered, but we do not claim causal training-data attribution, clinical validity, or unrestricted public release.

摘要:將模型行為歸因於合成訓練數據需要了解每個訓練項目是如何產生的,然後才能估計該項目造成了什麼。波形-標籤對並不保留這種知識。我們提出了一種生成來源基底,其中合成研究對象綁定了源規範、生成內容、波形、目標、事實要求、質量信號、審查血統和不可變的清單身份。生產者和選擇機制決定了證據意義;存儲位置和變量名稱則不然。我們在一個私有的日本護理交接管道中審計這一基底。一個包含113個資產的審查群體包含了六個情境系列中的1.552小時合成語音;所有項目都鏈接了音頻、轉錄、候選筆記和事實檢查清單,但人類證據是選擇性的且特定於來源。兩個僅限忠實的清單在情境種子上是不相交且不可變版本的,而精確的上游歸因仍然受到浮動生成器別名、缺失的每段TTS和代碼印記以及未版本化的檢查提示的阻礙。我們主張生成來源對於行為歸因是必要但不充分的:它定義了候選因果圖和審計單位,而貢獻性歸因仍然需要凍結的訓練運行和干預或影響證據。本文貢獻了一個簡潔的來源合約、一個審計協議和一個有界的合成數據歸因案例研究;可能會提供受控的研究訪問,但我們不聲稱因果訓練數據歸因、臨床有效性或不受限制的公開發布。

An ontology for cross-sectoral crisis management: core and public health modules

2610.01326v1 by Aldo Gangemi, Rita T. Sousa, Luigi Asprino, Giorgia Lodi, Andrea G. Nuzzolese, Valentina Presutti, Johannes Gysen, Diana F. Sousa, Luigi Spagnolo

This paper presents the European Crisis Management Ontology (ECMO), a modular OWL-based ontology intended as a cross-sectoral reference for disaster risk reduction and response. ECMO is designed to be organised as a network of ontological modules. Among the modules, ECMO-CORE captures fundamental crisis management concepts such as hazard, event, exposure, impact, and response measure and uses ontology design patterns and the OWL2 punning technique to resolve ambiguities between hazard types and event manifestations. In addition, domain-specific modules are defined as in the case of the public health module aligned with SNOMED CT and ICD-11. To demonstrate the resource's utility, we used ECMO to represent the data of the Epidemic Intelligence from Open Sources system of the Joint Research Centre to generate an end-to-end pipeline that populates an ECMO-compliant knowledge graph from unstructured epidemiological news. Initial results demonstrate that ECMO provides the formal guardrails necessary for consistent and unified knowledge representation and integration. The ontology is publicly available at https://doi.org/10.5281/zenodo.20070268 and is released under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.

摘要:這篇論文介紹了歐洲危機管理本體(ECMO),這是一個基於OWL的模組化本體,旨在作為災害風險減少和應對的跨領域參考。ECMO的設計是作為本體模組的網絡組織。 在這些模組中,ECMO-CORE捕捉了基本的危機管理概念,如危險、事件、暴露、影響和應對措施,並使用本體設計模式和OWL2的雙義技術來解決危險類型和事件表現之間的歧義。此外,還定義了特定領域的模組,例如與SNOMED CT和ICD-11對齊的公共衛生模組。為了展示該資源的實用性,我們使用ECMO來表示聯合研究中心的開放來源流行病情報系統的數據,以生成一個從非結構化流行病學新聞填充ECMO合規知識圖譜的端到端管道。初步結果顯示,ECMO提供了必要的正式框架,以實現一致和統一的知識表示和整合。該本體可在https://doi.org/10.5281/zenodo.20070268上公開獲得,並根據創用CC 4.0國際版(CC BY 4.0)授權發布。

Model validation in machine learning: A scenario-based guide from hold-out splits to nested group cross-validation in biomedical and applied research

2610.01284v1 by Mehmet Baygin, Sengul Dogan, Turker Tuncer

Model validation estimates the performance of a complete learning procedure on new data. However, an invalid split can produce an optimistic and stable result. This tutorial reviews hold-out validation, train/validation/test designs, repeated random subsampling, k-fold and repeated stratified cross-validation, leave-one-out and leave-p-out schemes, group-aware validation, and nested group cross-validation. General machine-learning principles are linked to EEG epochs, paired-eye OCT images, repeated clinical measurements, and multicenter data. Eight controlled scenarios compare flawed and leakage-safe designs: seven use locked confusion matrices with auditable metrics, and one uses a reproducible repeated-study simulation. The scenarios cover global feature selection, normalization leakage, dependent records, center mixing, repeated test-set use, and estimator instability. Bias, variance, metric aggregation, uncertainty, and computational cost are also examined. A data-size matrix, a decision tree, and reporting checklists are provided. Reproducible MATLAB templates and scikit-learn counterparts are included. The results show that no validation method is universally best. The independent unit must match the intended deployment target. Every data-dependent operation must also exclude the observations used for performance estimation.

摘要:模型驗證評估完整學習程序在新數據上的表現。然後,無效的分割可能會產生樂觀且穩定的結果。本教程回顧了保留驗證、訓練/驗證/測試設計、重複隨機子抽樣、k-折和重複分層交叉驗證、留一法和留p法方案、群體感知驗證以及嵌套群體交叉驗證。一般的機器學習原則與EEG時期、配對眼睛OCT影像、重複臨床測量和多中心數據相關聯。八個控制場景比較了有缺陷和防洩漏的設計:七個使用鎖定的混淆矩陣和可審計的指標,一個使用可重複的重複研究模擬。這些場景涵蓋了全局特徵選擇、正規化洩漏、依賴記錄、中心混合、重複測試集使用和估計器不穩定性。還檢查了偏差、方差、指標聚合、不確定性和計算成本。提供了數據大小矩陣、決策樹和報告檢查清單。包括可重複的MATLAB模板和scikit-learn對應物。結果顯示,沒有一種驗證方法是普遍最佳的。獨立單元必須與預期的部署目標相匹配。每個依賴數據的操作也必須排除用於性能估計的觀察值。

When Does Exercise-Specific Joint Selection Help? An Audit of Evaluation and Control Design

2610.01188v1 by Haotian Chen, Jingkun Yu, Yuning Zhang, Bowen Ye

Exercise-specific joint selection can improve skeleton-based correctness classification, but what does that gain establish? We audit 1,057 repetitions from ten REHAB24-6 subjects, separating evaluation aggregation, subset structure, and temporal representation. The manual-subset kNN gain changes from 0.055 for pooled out-of-fold AUROC to 0.020 for equal-weight within-person AUROC; both paired intervals include zero. Among 1,000 dimension-matched random maps, 14 match or exceed the manual pooled result, versus 145 when bilateral structure and trunk inclusion are also matched. RBF-SVM retains a positive within-person gain, whereas logistic regression and a random-convolution comparator have negative point gains under that estimand. Sequence-order and paired-seed controls further qualify the interpretation. This exploratory audit shows why joint-selection claims require explicit estimands and structurally appropriate controls; it does not establish a new algorithm or clinical benefit.

摘要:運動特定的關節選擇可以改善基於骨架的正確性分類,但這樣的增益究竟建立了什麼?我們審核了來自十名 REHAB24-6 受試者的 1,057 次重複,分離評估聚合、子集結構和時間表示。手動子集 kNN 的增益從 0.055 變化到 0.020,對應於合併的折外 AUROC 和等權重的個體內 AUROC;這兩個配對區間都包含零。在 1,000 個維度匹配的隨機映射中,有 14 個匹配或超過手動合併結果,而當雙邊結構和軀幹包含也匹配時,則有 145 個。RBF-SVM 保持了正的個體內增益,而邏輯回歸和隨機卷積比較器在該估計下則有負的點增益。序列順序和配對種子控制進一步限定了解釋。這項探索性審核顯示為什麼關節選擇的主張需要明確的估計量和結構上適當的控制;它並未建立新的算法或臨床利益。

CineMR: Tool-Integrated Vision-Language Reasoning for Quantitative Cardiac MRI Assessment

2610.01166v1 by Kunyang Li, Hai Nguyen, Joshua Lowe, Chenguang Zhao, Peace C. Madueme, Mehdi Hedjazi Moghari, Mubarak Shah, Pegah Khosravi, Yuzhang Zhang

Cardiovascular magnetic resonance (CMR), including cine imaging, is a reference standard for the noninvasive assessment of cardiac morphology and ventricular function. Cine CMR interpretation integrates qualitative visual assessment with quantitative measurements of ventricular volumes, ejection fraction, myocardial mass, wall thickness, and regional wall motion. Current medical vision-language models (VLMs) cannot reliably derive quantitative measurements from multidimensional cine images without analysis tools. We present CineMR, a tool-augmented VLM that invokes cardiac image-analysis tools and integrates their outputs into interleaved reasoning for quantitative CMR assessment. We also construct a multi-cohort visual question answering benchmark covering quantitative metric extraction, multiclass diagnosis, and differential diagnosis, together with tools for segmentation, phase selection, volumetry, morphometry, and regional wall motion analysis. CineMR is trained with supervised fine-tuning (SFT) on tool-interaction traces followed by Group Relative Policy Optimization (GRPO) with conditional tool-use rewards. On the multi-cohort cine CMR benchmark, CineMR achieves 35.9% pass@1 and 58.9% pass@4, compared with 1.5% pass@1 for the Qwen3-VL-8B backbone and 0.0% and 7.0% pass@1 for LLaVA-Med v1.5 and MedGemma-4B, respectively. Correct tool invocation reaches 99.8% after GRPO, up from 78.9% after SFT. Live tool outputs improve ventricular measurement accuracy by 20.4--23.7% over direct model predictions, and removing all tools reduces pass@1 from 35.9% to 27.9%. These results highlight the importance of reliable tool use for quantitative cine CMR reasoning and support CineMR as a promising approach for assistive cardiac image assessment. Code, benchmark resources, and model weights are available at https://github.com/AI-MIND-Lab/CineMR.

摘要:心血管磁共振(CMR),包括動態影像,是非侵入性評估心臟形態和心室功能的參考標準。動態 CMR 解釋將定性視覺評估與心室體積、射血分數、心肌質量、壁厚和區域壁運動的定量測量相結合。目前的醫療視覺-語言模型(VLMs)在沒有分析工具的情況下,無法可靠地從多維動態影像中推導出定量測量。我們提出了 CineMR,一種增強工具的 VLM,調用心臟影像分析工具並將其輸出整合到交錯推理中,以進行定量 CMR 評估。我們還構建了一個涵蓋定量指標提取、多類別診斷和鑑別診斷的多隊列視覺問答基準,並提供分割、相位選擇、體積測量、形態測量和區域壁運動分析的工具。CineMR 在工具互動痕跡上進行了監督微調(SFT),隨後使用條件工具使用獎勵進行了群體相對策略優化(GRPO)。在多隊列動態 CMR 基準上,CineMR 的 pass@1 為 35.9%,pass@4 為 58.9%,而 Qwen3-VL-8B 的 pass@1 僅為 1.5%,LLaVA-Med v1.5 和 MedGemma-4B 的 pass@1 分別為 0.0% 和 7.0%。經過 GRPO 正確調用工具的比例達到 99.8%,而 SFT 後為 78.9%。實時工具輸出提高了心室測量的準確性,較直接模型預測提高了 20.4% 至 23.7%,而去除所有工具則使 pass@1 從 35.9% 降至 27.9%。這些結果突顯了可靠工具使用在定量動態 CMR 推理中的重要性,並支持 CineMR 作為輔助心臟影像評估的有前景方法。代碼、基準資源和模型權重可在 https://github.com/AI-MIND-Lab/CineMR 獲得。

A Matched-Budget Audit Framework for Recaptioned Image-Text Supervision Distributions

2610.00952v1 by Giyeong Oh, Junghun Park, Yuhan Bae, Youngjae Yu

Recaptioned image-text corpora are now standard for text-to-image (T2I) training, with vision--language model (VLM) captioners replacing sparse alt-text by dense descriptions. A recaptioned corpus is a supervision distribution induced by a documented captioning policy ($π$), captioner ($V_c$), and source corpus ($C$). Length-correlated proxies miss caption-register artifacts and downstream T2I benchmarks entangle the corpus with training choices, so this distribution is hard to audit at corpus scale. We introduce a reusable matched-budget audit framework for recaptioned supervision distributions $D_{π,V_c,C}$: at a fixed text budget of $B = 64$ it reports a five-axis profile spanning prompt-side coverage, image-conditioned faithfulness, and caption-surface health, with claimed controllable basic units (CBU) as the common claim unit. We instantiate the framework on seven paired comparisons over five public source corpora. Across the four cross-corpus pairs, the released surface raises supported CBU per caption by $+3.39$ to $+6.36$ under both Qwen and Gemma Judges, and on CC12M the same framework exposes a long-vs-dense frontier that is consistent across both judges and four budgets. We release the audited multi-source recap corpus ($\approx$ 490M) together with the audit-artifact bundle.

摘要:重新標題的影像-文本語料庫現在已成為文本到影像(T2I)訓練的標準,視覺-語言模型(VLM)標題生成器用密集的描述取代了稀疏的替代文本。重新標題的語料庫是由文件化的標題政策($π$)、標題生成器($V_c$)和來源語料庫($C$)所引發的監督分佈。與長度相關的代理錯過了標題註冊的工件,而下游的 T2I 基準則將語料庫與訓練選擇糾纏在一起,因此這種分佈在語料庫規模上很難進行審計。我們引入了一個可重用的匹配預算審計框架,用於重新標題的監督分佈 $D_{π,V_c,C}$:在固定的文本預算 $B = 64$ 下,它報告了一個涵蓋提示側覆蓋率、影像條件忠實度和標題表面健康的五軸概況,並以聲稱可控的基本單位(CBU)作為共同的聲明單位。我們在五個公共來源語料庫上進行了七個配對比較來實現該框架。在四對跨語料庫的比較中,釋放的表面在 Qwen 和 Gemma 評審下每個標題支持的 CBU 提高了 $+3.39$ 到 $+6.36$,而在 CC12M 上,同樣的框架揭示了一個長對密集的邊界,這在兩位評審和四個預算中都是一致的。我們釋放了經過審計的多來源重新標題語料庫($\approx$ 490M),以及審計工件包。

Backdoor Purification for LoRA-Tuned LLMs via Null-Space Projection

2610.00685v1 by Jianwei Li, Jung-Eun Kim

With the rapid adoption of large language models (LLMs) and parameter-efficient fine-tuning (PEFT) methods, the risk of backdoor attacks has become more severe. Existing backdoor purification methods typically rely on at least one of the strong assumptions, such as prior knowledge of triggers, access to clean references, or aggressive retraining, and they often lack comprehensive evaluations. These constraints substantially limit their practical applicability. To overcome these challenges, our work proposes purifying LoRA-tuned LLMs without these assumptions and even without post-hoc retraining of the suspect parameters. Our objective is to significantly reduce the attack success rates (ASR) while preserving both (i) the base model's general capabilities and (ii) the new downstream skills learned through the adapter. Through a series of ablation studies, we progressively scale our approach from a single layer in a text classification setting to a full-parameter LLM in the generative task. Through careful data curation and feature approximation, we extract high-fidelity backdoor directions and, for each layer or head, construct orthogonal null spaces in both the input and output channels, onto which the LoRA updates are projected. Empirically, our null-space projection method reduces the ASR from nearly 100% to less than 10%, while preserving the base model's benign performance and the adapter's learned abilities during downstream task adaptation.

摘要:隨著大型語言模型(LLMs)和參數高效微調(PEFT)方法的快速採用,後門攻擊的風險變得更加嚴重。現有的後門淨化方法通常依賴於至少一個強假設,例如對觸發器的先驗知識、訪問乾淨參考資料或激進的再訓練,並且它們往往缺乏全面的評估。這些限制大大限制了它們的實際應用性。為了克服這些挑戰,我們的工作提出了在沒有這些假設的情況下淨化LoRA調整的LLMs,甚至不需要對可疑參數進行事後再訓練。我們的目標是顯著降低攻擊成功率(ASR),同時保留(i)基礎模型的一般能力和(ii)通過適配器學到的新下游技能。通過一系列的消融研究,我們逐步將我們的方法從文本分類設定中的單層擴展到生成任務中的全參數LLM。通過仔細的數據策劃和特徵近似,我們提取高保真度的後門方向,並為每一層或頭構建正交的零空間,這些零空間位於輸入和輸出通道上,LoRA更新將被投影到這些空間中。經驗上,我們的零空間投影方法將ASR從近乎100%降低到不到10%,同時在下游任務適應過程中保留了基礎模型的良性性能和適配器學到的能力。

Worse Together: How Performance Breaks Down in Multi-User Multi-Agent Teams

2610.00583v1 by Sahan Paliskara, Nattaput Namchittai, Andrew Lampinen

People are increasingly delegating tasks to AI agents, and those agents are increasingly encountering other people's agents over shared resources such as a codebase, a calendar, or a budget. When each agent acts for a different user with different goals, coordination often fails, and the group ends up worse off than if a single agent had acted for everyone. We study this multi-user, multi-agent setting across five frontier models and 77 scenarios in four environments: an API key environment in which agents share a compute budget, a clinic in which they share a calendar, a personal assistant environment in which they share a group order or booking, and a merge queue in which they share a release cutoff. In each scenario, we compare a single agent that serves every user (a coordinator) to a team in which each agent serves one user, with and without a communication channel between the agents. Teams deliver worse group outcomes than the coordinator in every environment: without a channel, they completely collapse in two environments, and even with one, coordination overhead creates substantial gaps. For example, in the personal assistant environment, the coordinator fulfills a targeted user request about twice as often as teams. We identify distinct behaviors associated with this poor group-level performance, including stalling as teams grow, overriding each other's actions, and fabricating claims. We find effective but environment-specific mitigations, such as a team lead, explicit procedural instructions, and a platform check that makes an agent read its peers' messages before committing. We will release the API key, clinic, and personal assistant environments as MAMUBench, comprising 74 scenarios for evaluating multi-user, multi-agent coordination.

摘要:人們越來越多地將任務委派給 AI 代理,而這些代理也越來越多地與其他人的代理在共享資源上相遇,例如代碼庫、日曆或預算。當每個代理代表不同的用戶且目標不同時,協調往往失敗,結果小組的情況比由單一代理為所有人行動時更糟。我們研究了這種多用戶、多代理的設定,涵蓋五個前沿模型和四個環境中的 77 種情境:一個 API 金鑰環境,在這裡代理共享計算預算;一個診所,在這裡他們共享日曆;一個個人助理環境,在這裡他們共享團體訂單或預訂;以及一個合併隊列,在這裡他們共享發佈截止時間。在每個情境中,我們將為每個用戶服務的單一代理(協調者)與每個代理服務一位用戶的團隊進行比較,並考慮代理之間是否有通信渠道。團隊在每個環境中提供的群體結果都比協調者差:在沒有渠道的情況下,他們在兩個環境中完全崩潰,即使有一個,協調開銷也會造成相當大的差距。例如,在個人助理環境中,協調者滿足目標用戶請求的頻率約為團隊的兩倍。我們識別出與這種低群體表現相關的不同行為,包括隨著團隊增長而停滯、覆蓋彼此的行動以及捏造聲明。我們發現有效但特定於環境的緩解措施,例如團隊負責人、明確的程序指示,以及一個平台檢查,使代理在提交之前閱讀其同伴的消息。我們將發布 API 金鑰、診所和個人助理環境作為 MAMUBench,包含 74 種情境以評估多用戶、多代理的協調。

Can LLMs Reason Over Long Horizons? An Empirical Evaluation of Context Strategies for Longitudinal Clinical Reasoning

2610.00562v1 by Taye Akinrele, Noorbakhsh Amiri Golilarz, Subash Neupane, Sudip Mittal, Shahram Rahimi

Longitudinal clinical reasoning requires large language models (LLMs) to identify and integrate relevant evidence distributed across extended patient histories. Although long-context models can process increasingly large amounts of information, providing more history does not necessarily make relevant evidence more accessible or improve reasoning. We compare five context strategies (Full, Recent, Episodic, Semantic, and Hybrid) on MedLoCoMo across four open-weight LLMs, examining answer correctness, robustness to query-evidence distance, and abstention on questions with unsupported premises. Episodic and Hybrid generally achieve the strongest overall accuracy, while Recent Context degrades most as supporting evidence becomes more distant; Episodic and Hybrid maintain the highest accuracy at long distances. Analysis of adversarial questions further shows that strong performance on answerable questions does not necessarily translate to successful abstention when the available history does not support the requested conclusion. These findings show that reliable longitudinal reasoning depends not only on how much history an LLM can access, but critically on how relevant evidence is selected and presented for reasoning.

摘要:長期臨床推理需要大型語言模型(LLMs)識別並整合分佈在廣泛病歷中的相關證據。雖然長上下文模型可以處理越來越多的信息,但提供更多的歷史並不一定使相關證據更易於獲取或改善推理。我們在四個開放權重的LLM上比較五種上下文策略(完整、最近、情節、語義和混合)在MedLoCoMo上的表現,檢查答案的正確性、對查詢-證據距離的穩健性,以及對不支持前提的問題的棄權。情節和混合策略通常實現了最強的整體準確性,而最近上下文在支持證據變得更遙遠時下降最嚴重;情節和混合策略在長距離下保持最高的準確性。對對抗性問題的分析進一步顯示,對可回答問題的強勁表現並不一定轉化為在可用歷史不支持所要求結論時的成功棄權。這些發現表明,可靠的長期推理不僅依賴於LLM可以訪問多少歷史,還關鍵於如何選擇和呈現相關證據以進行推理。

Unlearnable, or Unmeasured? On the Reliability of Difficulty Labels in RLVR

2609.40115v1 by Chandak Chakma, Syed Nazmus Sakib, Nafiul Haque, Shifat E. Arman

Reinforcement learning with verifiable rewards (RLVR) has become an important approach for improving reasoning during post-training. Recent work suggests that some difficult prompts remain resistant to learning even when they occasionally produce correct solutions. We revisit this unlearnability phenomenon and find that the affected prompts do improve, at roughly one third of the learnable rate, while the difficulty-defined set used to study them is much less reproducible than expected. These difficulty labels are estimated from a limited number of sampled responses. Combining them across seeds can further change which prompts are selected instead of simply reducing measurement noise. We develop a sampling-based framework for quantifying this instability and determining how much evaluation is required for difficulty assignments to reproduce reliably. We also revisit the gradient-similarity evidence proposed to explain unlearnability and show that part of the observed separation arises because difficult prompts provide fewer correct rollouts from which their gradients can be estimated. Matching this sample count weakens the gradient difference but does not remove it. Overall, the slow-learning phenomenon survives our reanalysis, while both the prompts used to define it and the evidence used to explain it require more careful measurement.

摘要:強化學習與可驗證獎勵(RLVR)已成為改善後訓練推理的重要方法。最近的研究表明,即使某些困難的提示偶爾產生正確的解決方案,它們仍然對學習具有抵抗力。我們重新檢視這一不可學習現象,發現受影響的提示確實有所改善,改善速度約為可學習速率的三分之一,而用來研究它們的困難定義集的可重現性遠低於預期。這些困難標籤是從有限數量的樣本反應中估算得出的。跨種子結合它們可能進一步改變所選擇的提示,而不僅僅是減少測量噪音。我們開發了一個基於抽樣的框架來量化這種不穩定性,並確定為了使困難分配可靠地重現需要多少評估。我們還重新檢視了用於解釋不可學習的梯度相似性證據,並顯示觀察到的分離部分源於困難提示提供的正確回饋較少,從中無法估算其梯度。匹配這一樣本數量削弱了梯度差異,但並未消除它。總體而言,緩慢學習現象在我們的重新分析中仍然存在,而用來定義它的提示和用來解釋它的證據都需要更仔細的測量。

GateSPINE: Gated Cross-View Fusion for Lumbar Spine MRI Report Generation

2609.40091v1 by Hoang Nguyen Van, Cuong Vuong Tuan, Trang Mai Xuan, Bien Tran Van, Nam Tran Van, Thien Van Luong

Automated report generation can ease the burden radiolo gists face when interpreting multi-sequence MRI studies. Unlike CT, MRI examinations comprise multiple sequences and imaging planes, each con tributing complementary diagnostic information. Existing methods en code a study as a single volume and combine multiple acquisitions by fixed rules. Findings visible in only one plane are thus diluted and of ten missed, lowering recall on clinical efficacy metrics, where a missed abnormality is most costly. We propose GateSPINE, a vision-language framework that fuses sagittal T1 and T2 volumes with a training-free operator, encodes the fused sagittal and axial volumes with two parallel 3D encoders, and decodes their combined representation into a report. Its core mechanism is a gated cross view fusion module that predicts, per feature channel and token, how much of each view to admit, so the more informative view dominates at each spatial location. We evaluate GateSPINE on three lumbar MRI datasets, comprising two public bench marks and a private cohort collected from Phenikaa University Hospital, using both natural language generation (NLG) and clinical efficacy (CE) metrics. GateSPINE achieves the highest CE F1 through improved re call on all three datasets; on SPIDER, which lacks an axial sequence, this reflects the sagittal fusion component rather than the gated cross-view mechanism, which is validated on the two cohorts with both imaging planes. GateSPINE also remains competitive on standard NLG metrics.

摘要:自動報告生成可以減輕放射科醫生在解讀多序列MRI研究時所面臨的負擔。與CT不同,MRI檢查由多個序列和成像平面組成,每個平面提供互補的診斷信息。現有的方法將研究編碼為單一體積,並通過固定規則組合多個獲取結果。僅在一個平面上可見的發現因此被稀釋,並且經常被忽略,這降低了臨床效能指標的召回率,而漏掉的異常是最昂貴的。我們提出了GateSPINE,一個視覺-語言框架,融合了矢狀面T1和T2體積,使用無需訓練的運算符,並用兩個平行的3D編碼器編碼融合的矢狀面和軸向體積,然後將它們的組合表示解碼成報告。其核心機制是一個門控交叉視圖融合模塊,根據特徵通道和標記預測每個視圖應該接受多少,以便在每個空間位置上更具信息性的視圖占主導地位。我們在三個腰椎MRI數據集上評估GateSPINE,包括兩個公共基準和一個來自Phenikaa大學醫院的私有隊列,使用自然語言生成(NLG)和臨床效能(CE)指標。GateSPINE通過提高所有三個數據集的召回率,實現了最高的CE F1;在缺少軸向序列的SPIDER上,這反映了矢狀面融合組件,而不是門控交叉視圖機制,這在兩個具有成像平面的隊列中得到了驗證。GateSPINE在標準NLG指標上也保持競爭力。

Overview of BioASQ 2026: The fourteenth BioASQ Challenge on Large-Scale Biomedical Semantic Indexing and Question Answering

2609.39975v1 by Anastasios Nentidis, Georgios Katsimpras, Anastasia Krithara, Martin Krallinger, Miguel Rodríguez-Ortega, Eduard Rodriguez-López, Natalia Loukachevitch, Igor Rozhkov, Elena Tutubalina, Dimitris Dimitriadis, Vasiliki Patsiou, Grigorios Tsoumakas, George Giannakoulas, Alexandra Bekiaridou, Athanasios Samaras, Giorgio Maria Di Nunzio, Nicola Ferro, Stefano Marchesin, Marco Martinelli, Gianmaria Silvello, Georgios Paliouras

This paper presents an overview of the fourteenth edition of the BioASQ challenge, organized in the context of the Conference and Labs of the Evaluation Forum (CLEF) 2026. BioASQ is an international challenge series that supports progress in biomedical language processing tasks ranging from semantic indexing and information extraction to question answering and summarization. In 2026, BioASQ included six shared tasks: a) Task 14b on biomedical semantic question answering. b) Task Synergy14 on question answering for developing biomedical top- ics. c) Task MultiClinSum-2 on multilingual clinical summarization. d) Task BioNNE-R on extracting relations between nested named entities in Russian and English. e) Task ELCardioCC on clinical coding in cardiology. f) Task GutBrainIE on gut-brain interplay information extrac- tion. Across these six tasks, 87 distinct teams participated, submitting more than 1000 runs overall. As in previous editions, several submissions reached competitive performance, reflecting the continued progress of state-of-the-art methods across biomedical language processing tasks.

摘要:這篇論文概述了第十四屆BioASQ挑戰賽,該賽事在2026年評估論壇會議及實驗室(CLEF)的背景下舉辦。BioASQ是一系列國際挑戰,旨在支持生物醫學語言處理任務的進展,這些任務包括語義索引、信息提取、問題回答和摘要。在2026年,BioASQ包括六個共享任務:a) 任務14b,針對生物醫學語義問題回答。b) 任務Synergy14,針對發展生物醫學主題的問題回答。c) 任務MultiClinSum-2,針對多語言臨床摘要。d) 任務BioNNE-R,提取俄語和英語中嵌套命名實體之間的關係。e) 任務ELCardioCC,針對心臟病學的臨床編碼。f) 任務GutBrainIE,針對腸道與大腦相互作用的信息提取。在這六個任務中,共有87支不同的團隊參加,總共提交了超過1000次運行。與之前的版本一樣,幾個提交達到了競爭性的表現,反映了在生物醫學語言處理任務中最先進方法的持續進步。

Scores That Hold, Benchmarks That Leak: Measuring Dataset Contamination in Public Brain-Tumor MRI Classification

2610.00421v1 by Bhanu Prakash Vangala, Sowmya Guda, Latha Peddi, Navya Vangala

Automated classification of brain tumors from MRI is a heavily published application of deep learning in medical imaging, with reported accuracies on public benchmarks routinely exceeding 98%. However, accuracy does not capture a critical dimension of benchmark quality: dataset integrity, defined as the independence of test from training data at the image, patient, and acquisition-source levels. We introduce a three-layer contamination framework comprising duplicate, patient, and source-label leakage to assess the public corpora on which this literature rests. We audit the three most widely used corpora against a chest-radiograph negative control and quantify each layer's effect on measured performance across nine architectures and three evaluation conditions. Contamination is severe at every layer: 28.8% of the dominant corpus's official test split has a near-twin in its own training split, a second corpus leaks 22.3% of its test images byte-identically, 95.5% of traceable test images share a patient with training, and file-header features containing no anatomy separate tumor from no-tumor at 0.959 balanced accuracy, at parity with fine-tuned ResNet backbones. The unexpected result is that removing every identified leaked test image leaves balanced accuracy essentially unchanged: stable performance after deduplication does not establish benchmark integrity. Our findings establish dataset integrity as a distinct, measurable axis of benchmark quality that a stable leaderboard cannot certify. For biomedical research, reported accuracy on these corpora alone does not establish that a model has learned to recognize tumors rather than exploit dataset-specific cues. We release the contaminated-file lists, recovered patient identifiers, and deduplicated splits.

摘要:自動化分類腦腫瘤的MRI是深度學習在醫學影像中的一個廣泛發表的應用,報告的準確率在公共基準上通常超過98%。然而,準確率並未捕捉到基準質量的一個關鍵維度:數據集完整性,定義為在影像、病人和獲取來源層面上測試數據與訓練數據的獨立性。我們引入了一個三層污染框架,包括重複、病人和來源標籤洩漏,以評估這些文獻所依賴的公共語料庫。我們對三個最廣泛使用的語料庫進行審核,與胸部X光的陰性對照進行比較,並量化每一層對九種架構和三種評估條件下測量性能的影響。每一層的污染都很嚴重:主導語料庫的官方測試分割中有28.8%的部分在其自己的訓練分割中有近似的雙胞胎,第二個語料庫有22.3%的測試影像以字節相同的方式洩漏,95.5%的可追溯測試影像與訓練共享病人,且不包含任何解剖結構的檔案標頭特徵在0.959的平衡準確率下將腫瘤與非腫瘤區分開來,與微調的ResNet骨幹相當。意外的結果是,移除每一個識別出的洩漏測試影像後,平衡準確率基本保持不變:去重後的穩定性能並未建立基準完整性。我們的研究結果確立了數據集完整性作為基準質量的一個獨特可測量軸,穩定的領先榜無法證明。對於生物醫學研究,僅僅依賴這些語料庫報告的準確率並不能證明模型已經學會識別腫瘤,而不是利用數據集特定的線索。我們發布了污染檔案列表、恢復的病人標識符和去重的分割。

Trust Is Not a Score: Runtime Assurance Contracts for High-Risk AI Agents

2609.39717v1 by Serhii Zabolotnii

Benchmarks, audits, and agent protocols describe performance, permissions, and repair, but not how observed evidence should change an agent's authority during a consequential task. We call this the assurance-transition gap. We propose a Runtime Assurance Contract (RAC), a policy-level formal schema binding autonomy boundaries, component eligibility, evidence state, transition policy, human-review capacity, and non-compensatory gates. Under RAC, soft metrics may inform routing, whereas a failed or unknown mandatory gate forces retry, switch, escalation, deferral, or stop; aggregate performance cannot authorize action. We define the contract, an evidence record, a permission rule, and five invariants, and illustrate them in clinical, industrial, and judicial failure probes. We then report a deterministic failure-injection study in agentic coding: 280 constructed cases evaluated by a gate conjunction, a score-only rule, and a restricted protocol baseline. At the published example weights and threshold, the score rule admits 80 of 100 block-required injections and all 40 review-required injections. Tuned in hindsight, it matches the conjunction on this corpus. For positive weights, a positive threshold, binary risk signals, zero-signal controls, and an injected case firing each signal alone, we show that exact agreement holds if and only if the threshold does not exceed the smallest weight. A separate set of 18 hand-authored traces checks version-pinned evidence and review transitions against simpler policy variants. In a further prospective synthetic holdout of 24 episodes, two blinded LLM judges assign identical labels to all 72 action attempts; RAC and a separately implemented full stateful baseline both match these labels. These studies test mechanisms on synthetic cases; they establish neither deployed safety nor cross-domain effectiveness.

摘要:基準、審計和代理協議描述了性能、權限和修復,但並未說明在關鍵任務中,觀察到的證據應如何改變代理的權限。我們稱之為保證過渡差距。我們提出了一個運行時保證合約(RAC),這是一個政策層級的正式架構,約束自主邊界、組件資格、證據狀態、過渡政策、人類審查能力和非補償性閘門。在RAC下,軟指標可以用來指導路由,而失敗或未知的強制閘門則強迫重試、切換、升級、延遲或停止;總體性能無法授權行動。我們定義了合約、一個證據記錄、一條許可規則和五個不變量,並在臨床、工業和司法失敗探測中進行了說明。然後,我們報告了一項在代理編碼中的確定性失敗注入研究:280個構建的案例通過閘門聯合、一個僅計分的規則和一個受限的協議基線進行評估。在已發表的示例權重和閾值下,計分規則允許100個區塊所需注入中的80個和所有40個審查所需的注入。事後調整後,它在這個語料庫上與聯合匹配。對於正權重、正閾值、二元風險信號、零信號控制和每個信號單獨觸發的注入案例,我們顯示出精確一致性僅在閾值不超過最小權重時成立。一組18個手工編寫的痕跡檢查版本固定的證據和審查過渡,與更簡單的政策變體進行比較。在進一步的前瞻性合成保留中,24個集數中,兩位盲法LLM評審對所有72次行動嘗試分配了相同的標籤;RAC和一個單獨實施的完整狀態基線都與這些標籤相匹配。這些研究在合成案例上測試機制;它們既未建立已部署的安全性,也未建立跨領域的有效性。

Interpretable Synthetic Medical Tabular Data Generation for Clinical Decision Support Using Fuzzy Cognitive Maps

2610.00391v1 by Michael Vasilakakis, Dimitris K. Iakovidis

Synthetic medical tabular data generation has become essential for developing and validating computer-based medical systems (CBMSs) when real clinical data is restricted due to privacy, ethical, or data availability limitations. Existing probabilistic and deep generative models often lack interpretability and fail to preserve clinically meaningful dependencies, limiting their suitability for safety-critical applications. This paper proposes a novel application of Fuzzy Cognitive Maps (FCMs) in a framework for synthetic medical tabular data generation with explicit causality and privacy preservation. Clinical features are described using linguistically interpretable fuzzy sets, and inter-feature dependencies are encoded as FCM edge weights computed from fuzzy set intersections. Synthetic patient records are generated by propagating randomly initialized linguistic activation vectors through the FCM until convergence, followed by defuzzification to produce clinically coherent numerical values. The approach natively handles mixed data types, and domain constraints common in health records. Experimental evaluation on UCI medical benchmark datasets demonstrates competitive performance under a Train-on-Synthetic-Test-on-Real (TSTR) protocol. The proposed method achieves accuracy of up to 0.81 and AUROC of up to 0.90 on the Heart Disease dataset, matching or exceeding TVAE and Gaussian Copula baselines while running exclusively on CPU. Fidelity metrics including KS Complement (up to 0.91) and Correlation Similarity (up to 0.95) confirm strong statistical coherence, and DCR Baseline Protection scores consistently exceed those of TVAE, confirming adequate privacy guarantees. These results demonstrate that causally grounded, interpretable fuzzy modeling offers a computationally efficient and transparent alternative to deep generative models for trustworthy synthetic data generation in CBMSs.

摘要:合成醫療表格數據生成已成為開發和驗證基於計算機的醫療系統(CBMSs)的必要條件,尤其是在由於隱私、倫理或數據可用性限制而無法使用真實臨床數據的情況下。現有的概率和深度生成模型往往缺乏可解釋性,並且未能保持臨床上有意義的依賴性,這限制了它們在安全關鍵應用中的適用性。本文提出了一種在合成醫療表格數據生成框架中應用模糊認知圖(FCMs)的新方法,該方法具有明確的因果關係和隱私保護。臨床特徵使用語言可解釋的模糊集合來描述,特徵間的依賴性則作為從模糊集合交集計算出的FCM邊權重進行編碼。合成病歷通過將隨機初始化的語言激活向量在FCM中傳播直到收斂,然後進行去模糊化以生成臨床上連貫的數值。該方法本土處理混合數據類型以及健康記錄中常見的領域約束。在UCI醫療基準數據集上的實驗評估顯示,在合成訓練-真實測試(TSTR)協議下表現競爭力。所提出的方法在心臟病數據集上達到高達0.81的準確率和高達0.90的AUROC,與TVAE和高斯聯合基準相匹配或超過,且僅在CPU上運行。包括KS補充(高達0.91)和相關性相似度(高達0.95)在內的保真度指標確認了強大的統計一致性,而DCR基準保護分數始終超過TVAE的分數,確認了足夠的隱私保障。這些結果表明,基於因果關係的可解釋模糊建模為CBMSs中的可信合成數據生成提供了一種計算效率高且透明的替代方案。

CAMOS: Coupled Oscillatory State-Space Model for Multimodal Clinical Time-Series

2609.39484v1 by Maxx Richard Rahman, Mostafa Hammouda, Wolfgang Maass

Longitudinal clinical cohorts are multimodal, irregularly sampled and pervasively incomplete: in ADNI, positron emission tomography and cerebrospinal fluid assays are absent from roughly half of all visits. Linear state-space models handle irregular sampling gracefully but treat a missing modality by masking the input, leaving the transition operator untouched. We prove that this is a representational limitation: the latent state of any linear state-space layer whose transition operator does not depend on the availability pattern is an additive function of the availability indicators, so no such layer can represent an interaction between two modalities being jointly present or jointly absent. We propose CAMOS, which gives each modality a bank of second-order oscillators coupled through a matrix that sits inside the differential equation and is gated by availability, so the transition operator itself becomes a function of which measurements were taken. Coupling invalidates the analysis of uncoupled oscillatory models, and we restore it: a per-channel Gershgorin budget makes the effective stiffness positive definite uniformly over all $2^M$ availability patterns and all gaps, an energy argument charges amplification to availability transitions rather than sequence length, and a channel factorization preserves exact associative parallel scans. On ADNI, CAMOS outperforms uncoupled oscillatory state-space models and clinical fusion models on same-visit staging, landmark prediction and longitudinal forecasting, and under zero-shot transfer to OASIS-3 it is the only model that avoids collapse to the majority class.

摘要:長期臨床隊列是多模態的、不規則取樣的,且普遍不完整:在ADNI中,正電子發射斷層掃描和腦脊液檢測大約在一半的訪問中缺失。線性狀態空間模型優雅地處理不規則取樣,但通過遮蔽輸入來處理缺失的模態,保持轉換運算子不變。我們證明這是一個表徵限制:任何線性狀態空間層的潛在狀態,其轉換運算子不依賴於可用性模式,是可用性指標的加法函數,因此沒有這樣的層能夠表示兩個模態共同存在或共同缺失之間的交互。我們提出CAMOS,為每個模態提供一組通過一個位於微分方程內的矩陣耦合的二階振盪器,並由可用性進行開關,因此轉換運算子本身成為一個函數,取決於採取了哪些測量。耦合使得無耦合振盪模型的分析失效,而我們恢復了它:每通道的Gershgorin預算使得有效剛度在所有$2^M$可用性模式和所有間隙上均為正定,能量論證將放大歸因於可用性轉換而非序列長度,通道因式分解保留了精確的關聯並行掃描。在ADNI上,CAMOS在同訪問分期、地標預測和長期預測方面優於無耦合振盪狀態空間模型和臨床融合模型,並在零樣本轉移到OASIS-3時,它是唯一一個避免崩潰到多數類別的模型。

Towards Trustworthy AI for Glioma Diagnosis: A Task-Aware Evaluation of Uncertainty Quantification

2609.39429v1 by Gonzalo Esteban Mosquera Rojas, Sebastian R. van der Voort, Carolin M. Pirkl, Sandeep Kaushik, Marion Smits, Stefan Klein

Uncertainty Quantification (UQ) is a key requirement for trustworthy AI in high-stakes medical image analysis. In this work, we evaluate UQ in a multi-task Deep Learning framework for MRI-based glioma diagnosis that performs tumor segmentation and predicts IDH mutation status, 1p/19q co-deletion status, and tumor grade. Monte Carlo Dropout (MCD) is used for a detailed task-aware analysis of predictive, aleatoric, and epistemic uncertainty. We assess MC sample convergence, calibration, error detection, selective prediction, associations with segmentation performance, and the effect of voxel-wise uncertainty aggregation on case-level reliability. We also compare MCD with Deep Ensembles (DE) and Monte Carlo Deep Ensembles (MCDE), examine interactions between segmentation quality and classification, and evaluate a composite trust score integrating segmentation and classification uncertainty. Across tasks, uncertainty estimates supported meaningful error detection, while calibration depended on the dropout rate, with moderate rates yielding the most reliable probabilities. Uncertainty decomposition provided task-dependent interpretability but did not consistently improve error detection over predictive uncertainty alone. DE and MCDE showed comparable operational utility, with no method consistently dominating across tasks and metrics. The composite trust score did not consistently outperform classification uncertainty for selective prediction. Overall, our results provide a task-aware evaluation strategy and practical guidance for the development of trustworthy AI for glioma diagnosis.

摘要:不確定性量化(UQ)是高風險醫療影像分析中可信賴人工智慧的關鍵要求。在這項工作中,我們評估了一個多任務深度學習框架中的UQ,該框架用於基於MRI的膠質瘤診斷,執行腫瘤分割並預測IDH突變狀態、1p/19q共同缺失狀態和腫瘤等級。使用蒙特卡羅隨機失活(MCD)對預測性、隨機性和認知性不確定性進行詳細的任務感知分析。我們評估了蒙特卡羅樣本收斂性、校準、錯誤檢測、選擇性預測、與分割性能的關聯,以及體素級不確定性聚合對案例級可靠性的影響。我們還將MCD與深度集成(DE)和蒙特卡羅深度集成(MCDE)進行比較,檢查分割質量與分類之間的相互作用,並評估整合分割和分類不確定性的綜合信任分數。在各項任務中,不確定性估計支持有意義的錯誤檢測,而校準則依賴於隨機失活率,中等率產生最可靠的概率。不確定性分解提供了任務依賴的可解釋性,但並未始終改善僅依賴預測不確定性的錯誤檢測。DE和MCDE顯示出可比的操作效用,沒有一種方法在各任務和指標中始終佔優。綜合信任分數在選擇性預測中並未始終優於分類不確定性。總體而言,我們的結果提供了一種任務感知的評估策略和實用指導,旨在為膠質瘤診斷的可信賴人工智慧發展提供支持。

EHR-RobustGym: Benchmarking and Training Agents for Robust Clinical Reasoning

2609.39371v1 by Yitong Qiao, Yancheng Jin, Lei Liu, Yue Shen, Jian Wang, Jinjie Gu, Zhixuan Chu

In hospital workflows, electronic health records (EHRs) are often noisy, and may not contain the evidence needed to confirm events or measurements referenced in a clinical query. Even when database retrieval succeeds, clinical agents can overlook such discrepancies and return plausible but unsupported answers. We introduce EHR-RobustGym, a scalable and interactive environment for evaluating and training robust clinical agents grounded in noisy EHRs. Built on MIMIC-IV hospital records (365K patients, 31 tables, and over 500M records), EHR-RobustGym comprises 5,486 Clean-Noise pairs spanning six clinical intents and both patient-level and population-level queries. The pairs test robustness to Record-level, Value-level, and Query-level noise, while interactive SQL/Python execution and outcome verification support trajectory collection and training. Evaluating multiple LLMs reveals substantial robustness gaps: average task success across proprietary and large-scale open-weight models drops from 62.2% on Clean questions to 37.9% on Noise questions. At k=4, pass^k consistency falls below 50% for most evaluated models, exposing instability in clinical task completion. Supervised fine-tuning and reinforcement learning in EHR-RobustGym improve performance, with gains generalizing to five external EHR benchmarks. Together, these results position EHR-RobustGym as a testbed for evaluating and improving the evidence-grounded robustness of clinical agents.

摘要:在醫院工作流程中,電子健康紀錄(EHR)通常是嘈雜的,並且可能不包含確認臨床查詢中提到的事件或測量所需的證據。即使數據庫檢索成功,臨床代理也可能忽略這些差異,並返回看似合理但未得到支持的答案。我們介紹了EHR-RobustGym,一個可擴展的互動環境,用於評估和訓練基於嘈雜EHR的穩健臨床代理。EHR-RobustGym建立在MIMIC-IV醫院紀錄上(365K名患者,31個表格,超過5億條紀錄),包括5486對乾淨-嘈雜配對,涵蓋六種臨床意圖以及患者級和人口級查詢。這些配對測試對紀錄級、值級和查詢級噪聲的穩健性,同時互動的SQL/Python執行和結果驗證支持軌跡收集和訓練。對多個LLM的評估顯示出顯著的穩健性差距:在專有和大規模開放權重模型中,乾淨問題的平均任務成功率從62.2%下降到嘈雜問題的37.9%。在k=4時,大多數評估模型的pass^k一致性降至50%以下,顯示臨床任務完成的不穩定性。在EHR-RobustGym中進行的監督微調和強化學習改善了性能,並且增益在五個外部EHR基準上具有普遍性。總體而言,這些結果使EHR-RobustGym成為評估和改善臨床代理的證據基礎穩健性的測試平台。

Structure vs. Chain-of-Thought: Evaluating LLM Criteria Extraction for Depression Severity

2609.39049v1 by Xinkai Chen

A large language model (LLM) can rate depression severity directly from a social media post or mark which clinical criteria the post shows and let code turn the count into a label. The latter is easier to audit because a clinician can check each marked criterion. We compare these approaches on two Reddit corpora using three LLMs (from 9B to frontier scale) and two questionnaires (PHQ-9, BDI-II), and measure agreement with quadratic weighted kappa. For the two frontier models, criteria extraction scores above chain-of-thought on one corpus only when its decision thresholds are fitted on labeled data. Neither model's gain is significant, with or without recalibrating chain-of-thought on the same labels. With thresholds fixed a priori from PHQ-9's criteria, extraction shows no gain on either corpus, even where models mark over two criteria per post. The 9B model behaves differently on a corpus from depression communities. It labels most posts severe, whether prompted directly or with chain-of-thought, while the a priori rule beats both without labels. After chain-of-thought is recalibrated on the same labels, no significant gap remains, consistent with a calibration effect. Yet higher ordinal agreement does not ensure better detection of severe cases. PHQ-9 criteria extraction misses most severe posts, and moving from direct prompting to chain-of-thought and then to extraction increases misses in nearly all comparisons. On the primary corpus, a relabeled stress dataset, a model using that dataset's own features, including word counts from the text, is not significantly different from frontier criteria extraction under the a priori rule.

摘要:一個大型語言模型(LLM)可以直接從社交媒體帖子中評估抑鬱症的嚴重程度,或標記該帖子顯示的臨床標準,並讓代碼將計數轉換為標籤。後者更容易進行審核,因為臨床醫生可以檢查每個標記的標準。我們在兩個Reddit語料庫上比較這些方法,使用三個LLM(從9B到前沿規模)和兩個問卷(PHQ-9,BDI-II),並測量與二次加權kappa的一致性。對於這兩個前沿模型,標準提取分數在一個語料庫上僅在其決策閾值適配於標記數據時超過思考鏈。無論是否重新校準思考鏈,兩個模型的增益都不顯著。當閾值根據PHQ-9的標準事先固定時,提取在任何語料庫上都沒有增益,即使模型每個帖子標記超過兩個標準。9B模型在抑鬱社區的語料庫上表現不同。無論是直接提示還是使用思考鏈,它都將大多數帖子標記為嚴重,而事先規則在沒有標籤的情況下超越了兩者。在思考鏈在相同標籤上重新校準後,沒有顯著的差距,這與校準效應一致。然而,更高的序數一致性並不保證更好地檢測嚴重病例。PHQ-9標準提取錯過了大多數嚴重帖子,從直接提示轉向思考鏈再到提取幾乎在所有比較中都增加了漏掉的情況。在主要語料庫中,一個重新標記的壓力數據集,使用該數據集自身的特徵,包括來自文本的字數,與根據事先規則的前沿標準提取沒有顯著差異。

An Uncertainty-Guided Digital Twin Framework for Online Adaptive Proton Therapy in Head and Neck Cancer: A Feasibility Study

2609.39010v1 by Yizhou Wu, Ryan J. Sanford, Huiqiao Xie, Jie Ding, Shupeng Chen, Tung-Ho Wu, Ping-Hsiu Wu, Justin Roper, Jun Zhou, Minglei Kang, Bill Stokes, Sibo Tian, David S. Yu, Xiaofeng Yang, Chih-Wei Chang

Objective: Head and neck (HN) proton therapy spans six to seven weeks of anatomical change, while offline replanning takes about a week. We present an uncertainty-guided digital twin (UGDT) framework that forecasts treatment-day anatomy before treatment and evaluate whether it generates online adaptive proton therapy (APT) plans of clinical quality. Approach: A library of 302 longitudinal deformations from 88 previously treated HN patients was transported onto each new patient's treatment planning CT (TPCT) using two-step multi-atlas deformable image registration (DIR) built on a pretrained CT foundation model, generating about 284 predicted CTs (pdCTs) with contours per patient. Dispersion of propagated clinical target volume (CTV) contours defined a patient-specific robust margin. In ten patients, the quality assurance CT (QACT) triggering a replan represented treatment-day anatomy, and the physician-approved replan was the baseline. The pdCT most similar to the QACT (pdCT-H) and one from the lowest quartile (pdCT-L) were planned to within about 5% of baseline plan quality, forward-calculated on the QACT, and reoptimized to generate online APT plans. Main results: pdCT plans scored within -0.7% (pdCT-H) and -1.0% (pdCT-L) of baseline. Forward calculation on QACT reduced high-dose CTV D98% to 88.3% and 85.5%. After online reoptimization, D98% recovered to 98.3 +/- 0.3% and 98.2 +/- 0.3%, versus 98.5 +/- 0.4% at baseline. Spinal cord and brainstem doses remained below tolerance, and plan quality scores were within -1.1% (p = 0.19) and -1.7% (p = 0.01) of baseline. Significance: UGDT generated online APT plans comparable in quality to physician-approved offline replans using anatomy forecast before treatment, enabling a transition from reactive offline replanning toward anticipatory online adaptation.

摘要:目標:頭頸部(HN)質子治療的解剖變化持續六到七週,而離線重新規劃約需一週時間。我們提出了一個不確定性引導的數位雙胞胎(UGDT)框架,該框架預測治療當天的解剖結構,並評估其是否能生成臨床質量的在線自適應質子治療(APT)計劃。方法:從88名先前接受治療的HN患者中提取的302個縱向變形被轉移到每位新患者的治療計劃CT(TPCT),使用基於預訓練CT基礎模型的兩步多圖譜可變形影像註冊(DIR),為每位患者生成約284個預測CT(pdCT)及其輪廓。傳播的臨床靶區體積(CTV)輪廓的分散定義了患者特定的穩健邊界。在十名患者中,觸發重新規劃的質量保證CT(QACT)代表治療當天的解剖結構,並且醫生批准的重新規劃為基準。與QACT最相似的pdCT(pdCT-H)和來自最低四分位數的pdCT(pdCT-L)計劃的質量約在基準計劃質量的5%以內,並在QACT上進行前向計算,然後重新優化以生成在線APT計劃。主要結果:pdCT計劃的得分在基準的-0.7%(pdCT-H)和-1.0%(pdCT-L)之內。在QACT上的前向計算將高劑量CTV D98%降低至88.3%和85.5%。在線重新優化後,D98%恢復至98.3 +/- 0.3%和98.2 +/- 0.3%,而基準為98.5 +/- 0.4%。脊髓和腦幹的劑量保持在耐受範圍以下,計劃質量得分在基準的-1.1%(p = 0.19)和-1.7%(p = 0.01)之內。意義:UGDT生成的在線APT計劃在質量上可與醫生批准的離線重新規劃相媲美,並使用治療前的解剖預測,使得從反應式離線重新規劃向預測式在線適應的轉變成為可能。

Scoring Higher, Answering Worse: Mitigating Reward Hacking in Rubric-Based RL via Protocol-Level Rubrics

2609.38847v1 by Maoqi Liu, Junwei He, Bowen Zhang, Feiran Li, Wentao Ma, Rongyi Lin, Shuhan Zhong, Quan Fang

Rubric-based reinforcement learning (Rubric-RL) trains language models where no verifier exists. A judge checks each criterion of a rubric, and the verdicts are aggregated into a reward, most often by a weighted sum. We show that this additive aggregation is the weak point. Under a sum, criteria compensate for one another: a policy that misses the one decision that matters can buy the points back with advice nobody asked for. On clinical consultation, such a policy scores higher and answers worse. Rubric coverage rises while appropriateness on held-out physician criteria falls below the untrained model. The medical criteria are not to blame. Grouped so that they must hold together, the same criteria, unchanged to the word, recover a third of the loss; shorter answers recover almost none. We therefore propose Protocol-level Rubrics (ProRubric), which keeps what the criteria ask for and changes how they are aggregated. It groups a checklist into a few protocol-level dimensions. A dimension counts only when all of its criteria hold and its failure clause does not fire. The grouping is done once, offline, and leaves the optimizer unchanged. ProRubric raises appropriateness by 10.8 points without losing coverage and has the best seven-benchmark average at both scales. Reward validity is set not only by what a rubric verifies, but by how it aggregates. Code is available at https://github.com/Estrellajer/ProRubric

摘要:基於評分標準的強化學習(Rubric-RL)訓練語言模型,當中不存在驗證者。評審檢查評分標準的每一項準則,並將判決匯總為獎勵,通常是通過加權總和。我們顯示這種加法匯總是其弱點。在總和下,準則之間相互補償:一個錯過了關鍵決策的策略可以用沒有人要求的建議來彌補分數。在臨床諮詢中,這樣的策略得分較高,但回答卻較差。評分標準的覆蓋率上升,而在保留的醫師準則上的適當性則低於未經訓練的模型。醫療準則並不是問題所在。這些準則被分組在一起,必須保持一致,未改變的相同準則恢復了三分之一的損失;較短的回答幾乎沒有恢復。因此,我們提出了協議級別評分標準(ProRubric),它保留了準則所要求的內容並改變了它們的匯總方式。它將檢查清單分組為幾個協議級別的維度。只有當所有準則都成立且其失敗條款未觸發時,該維度才計算。分組一次性完成,離線進行,並不改變優化器。ProRubric在不失去覆蓋率的情況下提高了10.8分的適當性,並在兩個尺度上擁有最佳的七項基準平均值。獎勵的有效性不僅由評分標準所驗證的內容決定,還由其匯總方式決定。代碼可在 https://github.com/Estrellajer/ProRubric 獲得。

CRAFT: Causal Responsibility and Failure Tracing in Medical Vision Language Models

2609.38810v1 by Chunzheng Zhu, Jiaqi Zeng, Hongbo Zhao, Yihang Chen, Yijun Wang, Jianxin Lin

As vision language models are increasingly deployed in clinical diagnosis, understanding how they internally resolve competing visual and textual signals becomes a safety imperative. Existing mechanistic analyses remain confined to unimodal text and offer no explanation for why a single misleading sentence can override a correct image based diagnosis, or why a model commits to a confident answer despite insufficient visual evidence. We find that these two safety risks, arbitration failure where textual context overrides visual grounding and brake failure where the model commits without adequate evidence, are mediated by spatially disjoint attention head populations: arbitration heads form a mid-to-deep wideband reflecting cross-layer evidence competition, while brake heads concentrate in a narrow middle-to-late layer band that regulates evidence sufficiency and abstention behavior. To ground these observations in causal circuitry, we introduce CRAFT, which localizes each failure mode to a minimal causal head set via dual criteria and verifies necessity and sufficiency through temporal probes and Tuned Lens trajectory analysis. Excising arbitration heads sharply reduces conflict following with negligible degradation on clean inputs, while excising brake heads restores appropriate abstention under degraded visual evidence. The two interventions target spatially disjoint head sets and produce distinct corrective effects, underscoring the mechanistic separability of the failure modes. Experiments across multiple medical VQA benchmarks and VLM architectures validate both the localization and interventions, demonstrating that the identified heads causally drive each failure mode and that targeted modulation generalises without retraining. The code is available at GitHub repository.

摘要:隨著視覺語言模型在臨床診斷中的應用日益增多,了解它們如何內部解決競爭的視覺和文本信號變得至關重要。現有的機制分析仍然局限於單一模式的文本,並未解釋為何一個誤導性的句子可以覆蓋正確的影像基礎診斷,或為何模型在視覺證據不足的情況下仍然堅持給出自信的答案。我們發現這兩種安全風險,即文本上下文覆蓋視覺基礎的仲裁失敗,以及模型在缺乏充分證據的情況下做出承諾的剎車失敗,是由空間上不相交的注意力頭群體所介導:仲裁頭形成中到深的寬頻,反映跨層證據競爭,而剎車頭則集中在狹窄的中到後層帶,調節證據的充分性和放棄行為。為了將這些觀察與因果電路相結合,我們引入了CRAFT,該方法通過雙重標準將每個失敗模式定位到最小的因果頭集,並通過時間探測和調整透鏡軌跡分析驗證必要性和充分性。切除仲裁頭顯著減少了衝突,對於乾淨輸入的降解幾乎可以忽略不計,而切除剎車頭則在視覺證據降級的情況下恢復了適當的放棄行為。這兩種干預針對空間上不相交的頭集,產生了不同的修正效果,強調了失敗模式的機制可分性。在多個醫學VQA基準和VLM架構上的實驗驗證了定位和干預,證明所識別的頭因果驅動每個失敗模式,並且針對性的調節在不重新訓練的情況下具有普遍性。代碼可在GitHub倉庫中獲得。

Sense and Sensitivity: Benchmarking LLM Clinical Triage Recommendations with Physician Experts

2609.38600v1 by Abinitha Gourabathina, Haoran Zhang, Yuexing Hao, Walter Gerych, Marzyeh Ghassemi

As large language models (LLMs) are increasingly used in clinical settings, it is critical to evaluate their reliability under realistic variation in clinical text. We study this question in clinical triage, comparing LLMs to practicing physicians under text perturbations that preserve the underlying clinical setting. We introduce a benchmark of over 6,000 clinical scenarios, 7,000 physician annotations, and 225,000 model responses. Using this benchmark, we make two key observations. First, LLMs are more likely than physicians to recommend unnecessary care at baseline, and this tendency increases under perturbed inputs. Further, we find that LLM recommendations are more sensitive to gender and tone perturbations than human recommendations. Together, these results demonstrate that LLMs can vary under clinically irrelevant textual changes, highlighting the need for deployment-oriented evaluations grounded in expert physician behavior.

摘要:隨著大型語言模型(LLMs)在臨床環境中的使用日益增加,評估它們在臨床文本的現實變異下的可靠性變得至關重要。我們在臨床分診中研究這個問題,將LLMs與在保持基礎臨床環境的文本擾動下的執業醫生進行比較。我們引入了一個基準,包含超過6,000個臨床場景、7,000個醫生註釋和225,000個模型回應。利用這個基準,我們做出兩個關鍵觀察。首先,LLMs在基線下比醫生更可能推薦不必要的護理,這種傾向在擾動輸入下會增加。此外,我們發現LLM的建議對性別和語氣擾動的敏感度高於人類的建議。這些結果共同表明,LLMs在臨床無關的文本變化下可能會有所不同,突顯了基於專家醫生行為的部署導向評估的必要性。

Defining and Categorising Human-AI Interactions in Clinical Trials: A Multidimensional Human-AI Classification Approach

2609.38559v1 by Sandra Woolley, Tim Collins, Khalid Khattak, Illia Chernomorets, Ariane Arevalo, Chris Richardson

This paper examines human-AI interactions (HAIIs) in clinical trials and presents a multidimensional categorisation framework that classifies interactions according to AI tasks, human-AI relationships, interaction configurations and interacting human groups. We define HAII, examine existing taxonomies and extend existing categorisation approaches through this novel multidimensional framework. We purposively sampled 15 clinical trials from a previously reported dataset. Each trial was independently categorised by two human reviewers and six large language model (LLM) classifiers. The proposed categorisation provides a structured method for the consistent identification, comparison and synthesis of human-AI interactions across clinical-trial records. The framework is intended to support more consistent comparison and synthesis of AI-related clinical trials and to make explicit the different forms of human involvement associated with AI interventions. The results demonstrate the potential for LLM-assisted categorisation while indicating the continuing importance of human judgement where trial records are incomplete or ambiguous. The principal contribution is a proposed multidimensional framework that brings together AI tasks, human-AI relationships, interaction configurations and interacting human groups within a single approach designed for clinical-trial records. Its significance lies in its potential to support more systematic identification, comparison and synthesis of how humans and AI interact in clinical trials.

摘要:這篇論文探討了臨床試驗中的人類-人工智慧互動(HAIIs),並提出了一個多維度的分類框架,根據人工智慧任務、人類-人工智慧關係、互動配置和互動人群對互動進行分類。我們定義了HAII,檢視現有的分類法,並通過這個新穎的多維框架擴展現有的分類方法。我們有目的地從先前報告的數據集中抽取了15個臨床試驗。每個試驗由兩位人類評審和六個大型語言模型(LLM)分類器獨立進行分類。所提出的分類提供了一種結構化的方法,用於在臨床試驗記錄中一致地識別、比較和綜合人類-人工智慧互動。該框架旨在支持對人工智慧相關臨床試驗的更一致的比較和綜合,並明確不同形式的人類參與與人工智慧干預相關聯。結果顯示了LLM輔助分類的潛力,同時指出在人類判斷仍然重要的情況下,當試驗記錄不完整或模糊時,仍需依賴人類判斷。主要貢獻是一個提出的多維框架,將人工智慧任務、人類-人工智慧關係、互動配置和互動人群整合在一個針對臨床試驗記錄的單一方法中。其重要性在於它能支持對人類和人工智慧在臨床試驗中互動的更系統的識別、比較和綜合。

Personalized State-Transition-Aware Memory for Clinical Agents

2609.38490v1 by Maryam Haghifam, Zahra Rajabi, Yizhou Sun, Carlos Morato

Large language model (LLM) agents that reason over clinical records must track changes in a patient's state while preserving the history needed to understand them. Simply accumulating memories leaves it unclear which information still applies, whereas overwriting earlier memories can erase evidence needed to reconstruct treatment history and clinical trajectories. We introduce STAM, a state-transition-aware memory framework that records state changes as new clinical entries arrive. STAM combines semantic retrieval with typed clinical relations to identify affected memories, maintaining current information in Active and superseded or resolved information in History. At read time, a query-dependent gate selectively serves historical memory. Across four longitudinal clinical benchmarks, we evaluate STAM with downstream question answering, direct state-maintenance diagnostics, and comparisons at approximately matched context lengths.

摘要:大型語言模型 (LLM) 代理人必須在保留理解病歷所需的歷史的同時,追蹤病人狀態的變化。單純地累積記憶會使得不清楚哪些信息仍然適用,而覆蓋早期記憶則可能抹去重建治療歷史和臨床軌跡所需的證據。我們介紹了 STAM,一個狀態轉換感知的記憶框架,該框架在新的臨床條目到達時記錄狀態變化。STAM 結合語義檢索和類型化的臨床關係來識別受影響的記憶,並在主動信息中保持當前信息,在歷史中保持被取代或解決的信息。在讀取時,查詢依賴的閘門選擇性地提供歷史記憶。在四個縱向臨床基準中,我們通過下游問題回答、直接狀態維護診斷以及在大約匹配的上下文長度下進行比較來評估 STAM。

KlinikeBench: Evaluating Language Models Beyond Diagnostic Accuracy

2609.38480v1 by Xueting Fang, Zehui Li, Yang Yang, Camilla Giovino, Shubh K. Patel, Shailly Prajapati, Vallijah Subasri, Caihua Shan

Most clinical benchmarks evaluate language models (LMs) on diagnosis using complete case descriptions. In clinical practice, however, patients present information in different ways, and clinicians must obtain relevant history and determine which examinations are needed before reaching a diagnosis. Diagnostic accuracy alone therefore cannot establish whether an agent gathered essential information or conducted an appropriate clinical assessment. Furthermore, existing benchmarks lack professional clinicians' verification. To address this gap, we introduce KlinikeBench, a benchmark of 333 clinician-authored tasks, each providing an isolated sandbox environment with a virtual patient, clinical tools, and task-specific success criteria. More than 35 clinicians contributed to case authoring and benchmark evaluation. In an empirical study, clinicians gave simulated dialogues higher mean quality ratings than reference conversations, which is adapted from real conversation. In each task, an LM has a fixed budget of turns to communicate with the patient, ask about relevant history, request examinations, follow action constraints, and record a final diagnosis. We score these steps separately as well as together. Across 31 models and seven model families, the best-performing models (e.g., GPT-6-astra and Claude Opus 5) succeed on less than 30% of tasks, even though their diagnosis accuracy reaches 90.7%. Some models benefit from talking with the patient; others diagnose well from a complete chart but perform much worse in conversation. Overall, KlinikeBench provides a testbed for evaluating the full clinical encounter and reveals a substantial gap between diagnostic accuracy and performance in interactive clinical assessment.

摘要:大多數臨床基準在診斷上評估語言模型(LMs)時使用完整的病例描述。然而,在臨床實踐中,患者以不同的方式呈現信息,臨床醫生必須獲取相關病史並確定需要哪些檢查,才能達成診斷。因此,僅僅依賴診斷準確性無法確定一個代理是否收集了必要的信息或進行了適當的臨床評估。此外,現有的基準缺乏專業臨床醫生的驗證。為了解決這一問題,我們推出了KlinikeBench,這是一個包含333個臨床醫生撰寫的任務的基準,每個任務都提供一個獨立的沙盒環境,內有虛擬患者、臨床工具和特定任務的成功標準。超過35位臨床醫生參與了案例撰寫和基準評估。在一項實證研究中,臨床醫生給模擬對話的平均質量評分高於參考對話,而參考對話是從真實對話中改編而來。在每個任務中,LM有固定的回合預算來與患者溝通,詢問相關病史,請求檢查,遵循行動約束,並記錄最終診斷。我們分別對這些步驟進行評分,也會綜合評分。在31個模型和七個模型系列中,表現最佳的模型(例如,GPT-6-astra和Claude Opus 5)在不到30%的任務中成功,儘管它們的診斷準確率達到90.7%。一些模型在與患者交談時受益;其他模型從完整的病歷中診斷良好,但在對話中表現卻差得多。總體而言,KlinikeBench提供了一個評估完整臨床互動的測試平台,並揭示了診斷準確性與互動臨床評估表現之間的重大差距。

PrivMeSA: Privacy-Aware Self-Evolving Multi-Agent System for Medicine via Local-Remote LLM Collaboration

2609.38458v1 by Dannong Wang, Yuran Zhang, Bian Sun, Alex Stinard, Yuzhang Shang, Song Wang, Yu Tian

Clinical large language model (LLM) agents deployed locally can consult more capable remote models, but doing so risks exposing patient information. Privacy-conscious delegation places disclosure decisions with a local agent, yet removing explicit identifiers is insufficient: quasi-identifiers can accumulate across multi-turn consultations and repeated patient visits to enable re-identification. We introduce PrivMeSA, a privacy-aware self-evolving multi-agent system that learns to control disclosure and retains remote expertise for local reuse. A local agent manages each encounter and consults remote specialists that may request additional information. Reinforcement learning balances task accuracy against direct disclosure and registry-based re-identification risk, with privacy evaluated over the complete outbound transcript of each encounter. A local lesson memory distills completed consultations into generalized clinical guidance and retrieves relevant lessons before transmission, allowing subsequent cases to reuse expertise without another remote exchange. Memory grows without additional outcome labels or parameter updates. On an emergency-department benchmark built from MIMIC-IV-ED records, PrivMeSA improves mean task accuracy over delegation by up to 15.8 percentage points. In the same setting, PrivMeSA reduces the disclosure of personal details from 98.0% to 0.2% of cases and the share of cases in which the patient can be narrowed to ten or fewer registry patients from 74% to 0%.

摘要:臨床大型語言模型(LLM)代理在本地部署可以諮詢更強大的遠程模型,但這樣做有可能暴露患者信息。注重隱私的委派將披露決策交給本地代理,但去除明確的識別符號並不足夠:準識別符號可以在多輪諮詢和重複的患者訪問中累積,從而實現重新識別。我們介紹了PrivMeSA,一個隱私意識的自我演變多代理系統,能夠學習控制披露並保留遠程專業知識以供本地重用。本地代理管理每次接觸並諮詢可能要求額外信息的遠程專家。強化學習在任務準確性和直接披露及基於登記的重新識別風險之間取得平衡,並在每次接觸的完整外發記錄上評估隱私。本地教訓記憶將已完成的諮詢提煉成通用的臨床指導,並在傳輸前檢索相關的教訓,允許後續案例在不進行另一個遠程交換的情況下重用專業知識。記憶在沒有額外結果標籤或參數更新的情況下增長。在基於MIMIC-IV-ED記錄建立的急診部門基準上,PrivMeSA將任務準確性的平均值提高了最多15.8個百分點。在同一環境中,PrivMeSA將個人詳細信息的披露從98.0%降低到0.2%,並將患者可以縮小到十名或更少登記患者的案例比例從74%降低到0%。

Colorectal Cancer Segmentation with Adaptive Augmentation and Multi-Resolution Ensemble Models

2609.38419v1 by Ümit Mert Çağlar, Alptekin Temizel

Colorectal cancer (CRC) is the second most deadly and third most common cancer, and the leading cause of death among gastrointestinal cancers. Early diagnosis is crucial for the treatment of this cancer and increasing the survival rates. Although CRC is more common in developed regions, its occurrence is also increasing in developing regions as well. CRC diagnosis relies on histopathology assessment post-biopsy. Automated deep learning algorithms can significantly reduce diagnosis time, enhancing efficiency and supporting timely clinical decisions. We present an automated segmentation pipeline for whole-slide histopathology images that labels tumor grades 1-3 and normal mucosa. It utilizes dense prediction transformers with various encoder backbones, overlapping patches, and test-time augmentation. An adaptive augmentation policy, guided by large language models, further improves training. Top models were ensembled via soft voting, and mask refining post-processing steps, Gaussian blurring, morphological closing, and connected components analysis. On a colorectal cancer grade dataset, our method improved the F1 score from 62.92 to 69.84. Code is available here: github.com/caglarmert/ICIP2025

摘要:大腸癌(CRC)是第二大致死率和第三大常見癌症,也是腸胃道癌症中最主要的死亡原因。早期診斷對於這種癌症的治療和提高存活率至關重要。儘管CRC在發達地區更為常見,但在發展中地區的發生率也在增加。CRC的診斷依賴於活檢後的組織病理學評估。自動化深度學習算法可以顯著縮短診斷時間,提高效率並支持及時的臨床決策。
我們提出了一個自動化的整張幻燈片組織病理圖像分割管道,該管道標記腫瘤等級1-3和正常黏膜。它利用具有各種編碼器骨幹的密集預測Transformer、重疊補丁和測試時增強。由大型語言模型引導的自適應增強策略進一步改善了訓練。頂尖模型通過軟投票進行集成,並進行了掩膜精煉後處理步驟、高斯模糊、形態學關閉和連通組件分析。在一個大腸癌等級數據集上,我們的方法將F1分數從62.92提高到69.84。代碼可在此獲得:github.com/caglarmert/ICIP2025

Aegis: Generative Gradient Masking for Privacy-Preserving Medical Federated Learning

2609.38339v1 by Chaoyu Zhang, Shanghao Shi, Heng Jin, Ning Wang, Y. Thomas Hou, Wenjing Lou

Federated learning (FL) has become a foundational paradigm for multi-institutional medical AI, allowing hospitals and research centers to jointly train diagnostic models without exchanging patient records. This privacy promise, however, is increasingly contested: a malicious or honest-but-curious server can launch model inversion attacks (MIAs) that reconstruct private patient images directly from shared model updates, and recent scalable, closed-form attacks penetrate even secure aggregation at clinically realistic batch sizes. Existing defenses face an unsatisfactory dilemma. Gradient-perturbation methods such as differential privacy and pruning trade away the diagnostic accuracy on which clinical reliability depends, while cryptographic protocols add system complexity yet still leave updates exposed to these scalable attacks. We propose Aegis, a principled client-side defense that breaks this dilemma without perturbing patient data or modifying the FL protocol. Our key insight is that the success of every known MIA is fundamentally bounded by the local batch size relative to the model's leakage capacity; once this limit is exceeded, distinct samples collide and reconstructions collapse into indistinguishable mixtures. Aegis turns this universal bottleneck into a defense: each client superimposes onto its real update a masking gradient computed on locally synthesized, task-relevant data, deliberately pushing the effective batch beyond the attack's recovery capacity. We complement the design with theoretical convergence guarantees under standard convex assumptions and evaluate Aegis on MNIST, CIFAR-10, and three MedMNIST modalities (chest X-ray, abdominal CT, colon pathology). Aegis neutralizes three state-of-the-art MIAs while preserving model utility and incurring only modest overhead, offering a practical privacy primitive for medical FL.

摘要:聯邦學習(FL)已成為多機構醫療人工智慧的基礎範式,使醫院和研究中心能夠共同訓練診斷模型,而無需交換病歷記錄。 然而,這一隱私承諾正受到越來越多的質疑:一個惡意或誠實但好奇的伺服器可以發動模型反演攻擊(MIA),直接從共享的模型更新中重建私人病人影像,而最近可擴展的封閉形式攻擊甚至能夠穿透臨床現實批次大小的安全聚合。 現有的防禦面臨著不令人滿意的困境。 像差分隱私和修剪這樣的梯度擾動方法犧牲了臨床可靠性所依賴的診斷準確性,而加密協議則增加了系統的複雜性,卻仍然讓更新暴露於這些可擴展的攻擊之下。 我們提出了Aegis,一種原則性的客戶端防禦,打破了這一困境,而不擾動病人數據或修改FL協議。 我們的關鍵見解是,每個已知的MIA的成功在根本上受到相對於模型泄漏能力的本地批次大小的限制;一旦超過這一限制,不同的樣本將發生碰撞,重建將崩潰為無法區分的混合物。 Aegis將這一普遍瓶頸轉化為防禦:每個客戶端在其真實更新上疊加一個基於本地合成的、與任務相關的數據計算出的掩蔽梯度,故意將有效批次推向超過攻擊的恢復能力。 我們在標準凸假設下補充了理論收斂保證,並在MNIST、CIFAR-10和三種MedMNIST模態(胸部X光、腹部CT、結腸病理)上評估了Aegis。 Aegis中和了三種最先進的MIA,同時保留了模型的效用,並僅產生適度的開銷,為醫療FL提供了一種實用的隱私原語。

A Proposed Rubric for Evaluating Expressed Clinical Reasoning in Large Language Model Responses

2609.37788v3 by Zhangshu Joshua Jiang, Zina Ibrahim, James T. Teo

We propose a rubric for assessing expressed clinical reasoning in model responses, drawing on three bodies of work: medical education assessment frameworks (ART, SCT, Key Feature Problems and OSCE); clinical LLM benchmarks (MedR-Bench, HealthBench, TIMER-Bench, DR. BENCH, PrIME-LLM and PatientSafeBench); and general LLM reasoning evaluation research, including the Factuality-Validity-Coherence-Utility taxonomy, FaithCoT-Bench and C2-Faith. We use groundedness as a clinically oriented adaptation of the taxonomy's factuality category. The rubric brings these concepts together in a multidimensional framework for scoring free-text responses to gold-standard clinical vignettes. It includes provisional behavioural anchors, applicability rules and a separate flag for case-specific safety-critical errors. General-domain frameworks inform its design but are not treated as validated clinical instruments. The rubric does not replace case-specific reference criteria or the task-specific metrics of existing benchmarks. It has not yet been tested for inter-rater reliability, construct validity or clinical utility. Its immediate purpose is to make evaluation decisions explicit and open to scrutiny before empirical testing.

摘要:我們提出了一個用於評估模型回應中表達的臨床推理的評分標準,這基於三個研究領域:醫學教育評估框架(ART、SCT、關鍵特徵問題和OSCE);臨床LLM基準(MedR-Bench、HealthBench、TIMER-Bench、DR. BENCH、PrIME-LLM和PatientSafeBench);以及一般LLM推理評估研究,包括事實性-有效性-一致性-實用性分類法、FaithCoT-Bench和C2-Faith。我們使用基於實證的概念作為該分類法事實性類別的臨床導向調整。這個評分標準將這些概念整合到一個多維框架中,用於對金標準臨床小插曲的自由文本回應進行打分。它包括臨時行為錨點、適用性規則以及針對特定案例的安全關鍵錯誤的單獨標記。一般領域框架為其設計提供了指導,但不被視為經過驗證的臨床工具。該評分標準並未取代特定案例的參考標準或現有基準的任務特定指標。它尚未針對評分者間可靠性、構念效度或臨床實用性進行測試。其直接目的在於使評估決策明確並可接受檢視,然後再進行實證測試。

Multi-Site Real-World Performance of Commercial AI for Pulmonary and Incidental Pulmonary Embolism Detection

2609.37750v1 by Aawez Mansuri, Mohammadreza Chavoshi, Theodorus Dapamede, Wasif Bala, Beatrice Brown-Mulry, Rohan Isaac, Bardia Khosravi, Hanzhou Li, Frank Li, John T. Moon, Chad Robichaux, Dan I. G. Cohen-Addad, Ninad V. Salastekar, Janice Newsome, Judy W. Gichoya, Hari Trivedi

Pulmonary embolism (PE) is a leading cause of cardiovascular mortality, yet the real-world performance of FDA-cleared AI detection models remains incompletely characterized. We retrospectively evaluated two FDA-cleared AI algorithms from a single commercial platform (Aidoc Medical BriefCase), one for PE triage on dedicated CT pulmonary angiography (CTPA; n = 30,678) and one for incidental PE (iPE) detection on routine contrast-enhanced CTs (n = 37,191), across a 17-facility academic health system. Reference-standard labels were extracted from radiology reports using a validated LLM pipeline (97% accuracy, kappa = 0.94). The PE model achieved 86.8% sensitivity and 99.1% specificity, with sensitivity declining from 99.3% for saddle emboli to 72.9% for subsegmental PE, and from 89.7% for acute to 65.3% for non-acute PE. The iPE model achieved 73.5% sensitivity and 99.8% specificity. Both models demonstrated lower sensitivity than FDA-clearance benchmarks while exceeding cleared specificity, with diminishing performance for peripheral and non-acute emboli mirroring known human reader limitations and underscoring the need for standardized post-market surveillance of AI-enabled medical devices.

摘要:肺栓塞(PE)是心血管死亡的主要原因,但FDA批准的AI檢測模型在實際應用中的表現仍然未完全明確。我們回顧性地評估了來自單一商業平台(Aidoc Medical BriefCase)的兩個FDA批准的AI算法,一個用於專用CT肺動脈造影(CTPA;n = 30,678)的PE分流,另一個用於常規對比增強CT(n = 37,191)的偶然PE(iPE)檢測,涵蓋了17家學術醫療系統。參考標準標籤是通過經驗證的LLM管道(97%準確率,kappa = 0.94)從放射學報告中提取的。PE模型的敏感性達到86.8%,特異性為99.1%,其中敏感性從鞍狀栓塞的99.3%下降到亞段PE的72.9%,從急性PE的89.7%下降到非急性PE的65.3%。iPE模型的敏感性為73.5%,特異性為99.8%。兩個模型的敏感性均低於FDA批准的基準,但特異性超過批准標準,對於周邊和非急性栓塞的表現下降反映了已知的人類讀者限制,並強調了對AI驅動醫療設備標準化市場後監測的需求。

Spatiotemporal Hyperedges for EEG Seizure Detection and Prediction

2609.37730v1 by Hyunju Kim, Sheo Yon Jhin, Noseong Park, Nabil Imam

Seizure detection and prediction from EEG are clinically important but challenging because seizures are rare, temporally localized, and propagate as coordinated events across multiple channels. Recent dynamic graph neural networks model this by running a temporal model over a sequence of per-time-step pairwise channel edges. However, this pairwise construction misses the spatiotemporal coupling that constitutes a seizure, at substantial training cost. We propose HyBrain, which summarizes spatiotemporal EEG evidence through a small set of soft hyperedges rather than pairwise edges. A per-channel Mamba backbone produces one token per (channel, second), and a spatiotemporal hyperedge block pools these tokens into E_h shared group embeddings through soft memberships and broadcasts them back. The same encoder serves three downstream tasks: window-based detection, one-second point-wise detection, and preictal seizure prediction. On TUSZ and CHB-MIT, HyBrain achieves the best AUROC on every reported setting against ten baselines, with the largest gap on long-clip preictal prediction. It also matches the most efficient baselines in training time and peak GPU memory. A qualitative analysis shows that even a single learned hyperedge cleanly captures the preictal -> ictal -> postictal trajectory on a real seizure clip.

摘要:癲癇發作的檢測和預測從腦電圖(EEG)中提取是臨床上重要但具有挑戰性的,因為癲癇發作是罕見的、時間上局部的,並且作為協調事件在多個通道中傳播。最近的動態圖神經網絡通過在每個時間步的成對通道邊緣上運行一個時間模型來模擬這一點。然而,這種成對的構建錯過了構成癲癇發作的時空耦合,並且訓練成本相當高。我們提出了HyBrain,它通過一小組軟超邊而不是成對邊來總結時空EEG證據。每個通道的Mamba主幹為每個(通道,秒)生成一個標記,時空超邊塊通過軟成員資格將這些標記聚合為E_h共享的群組嵌入並將其廣播回去。相同的編碼器服務於三個下游任務:基於窗口的檢測、一秒點檢測和癲癇發作前預測。在TUSZ和CHB-MIT上,HyBrain在所有報告的設置中對比十個基準達到最佳的AUROC,在長片段癲癇發作前預測上差距最大。它在訓練時間和峰值GPU內存方面也與最有效的基準相匹配。質性分析顯示,即使是一個學習到的超邊也能清晰地捕捉到真實癲癇片段中的癲癇發作前 -> 發作 -> 發作後的軌跡。

Correct, Don't Delete: Mitigating Emergent Misalignment with Corrective Supervision

2609.37624v1 by Jacob Epifano

Fine-tuning a language model on a narrow set of harmful demonstrations, such as bad medical advice, can make it broadly misaligned on unrelated questions, a phenomenon known as emergent misalignment (EM). The usual defense is to find the offending rows and delete them, but a row locator failed our held-out test and deleting rows helps less than expected. We ask a different question: given a fixed set of poisoned rows, is it better to correct them than to remove them? We fine-tune Qwen2.5-14B-Instruct on a mixture of bad medical advice and benign chat data, select a quarter of the poison rows in advance, and either delete them or replace each with a corrected answer to the same prompt, keeping everything else the same. Replacing the rows cuts the EM rate by about a third and improves answers on held-out medical questions, while deleting the same rows has little measurable effect. The advantage is larger when half the poison rows are corrected, and it holds on a second base model and a second misaligned model organism. The content of the replacement appears to matter: paraphrasing the rows while keeping their bad advice shows no clear benefit, and the correct answers distributed with the dataset appear to do about as well as our rewriter's. Realigning an already-poisoned model with further fine-tuning is known to work, but which data does the work has not been compared directly. We find that a short round of training on corrections beats the same amount of training on generic chat data, that corrections on other medical prompts do roughly as well as corrections of the poisoned prompts themselves, and that instructing the correction writer to model a careful, harm-avoiding assistant adds no measurable benefit over plain corrections. In the settings we tested, correcting harmful training data reduces EM more than deleting it.

摘要:微調一個語言模型於一組狹窄的有害示範,例如不良的醫療建議,可能會使其在無關問題上廣泛地失調,這一現象被稱為新興失調(EM)。通常的防禦方法是找到有問題的行並刪除它們,但行定位器在我們的保留測試中失敗,刪除行的效果也不如預期。我們提出一個不同的問題:在給定一組固定的有毒行的情況下,修正它們是否比刪除它們更好?我們對Qwen2.5-14B-Instruct進行微調,使用不良醫療建議和良性聊天數據的混合,提前選擇四分之一的有毒行,並選擇刪除它們或用相同提示的修正答案替換每一行,保持其他一切不變。替換這些行將EM率降低了約三分之一,並改善了對保留醫療問題的回答,而刪除相同的行幾乎沒有可測量的效果。當修正一半的有毒行時,這一優勢更大,並且在第二個基準模型和第二個失調模型上也成立。替換內容似乎很重要:在保持其不良建議的同時對行進行意譯並未顯示出明顯的好處,與數據集一起分發的正確答案似乎與我們的重寫者的表現相當。在已經被污染的模型上進行進一步微調以重新對齊是已知有效的,但哪些數據發揮作用尚未直接比較。我們發現,對修正進行短期訓練的效果超過了對通用聊天數據進行相同量的訓練,對其他醫療提示的修正效果大致與對有毒提示本身的修正相當,並且指導修正作者模擬一個小心、避免傷害的助手並未帶來比普通修正更可測量的好處。在我們測試的設置中,修正有害的訓練數據比刪除它更能減少EM。

ReLMem: Learning Recurrent Memory for Longitudinal EHR Modeling

2609.37587v1 by Zijie Meng, Xiwei Dai, Yingying Zhang, Jian Wu, Xian Wu, Zuozhu Liu

Longitudinal electronic health record (EHR) modeling requires integrating new visits with an expanding patient history. Yet the continual accumulation of clinical information imposes increasing computational and memory costs on large language models (LLMs) when they process and retain complete patient histories. A practical alternative is visit-wise recurrent compression, which incorporates each incoming visit into a compact, continually updated patient memory. However, under a fixed memory budget, successive updates must integrate new information without progressively losing critical historical evidence needed to subsequent tasks. To address this challenge, we introduce Recurrent Longitudinal Memory (ReLMem), a framework that learns to maintain fixed-capacity patient memory for efficient downstream prediction with a frozen LLM. ReLMem equips this LLM with lightweight compression adapters to recurrently update the memory from its previous state and each incoming visit, without rereading earlier records. Specifically, we develop a multi-granularity optimization strategy to preserve task-relevant information throughout recurrent updates and support downstream prediction from the final memory. The intermediate supervision aligns attention outputs from compressed memory and the full history under identical queries, while prediction supervision minimizes cross-entropy with ground truth answers conditioned on the final memory. On EHR-based medication prediction, ReLMem approaches the F1 scores of full-history baseline while reducing average retained historical storage by 97.1%. Under the same memory budget, it improves macro- and micro-F1 over the strongest compressed-memory baseline by 4.66 and 4.75 percentage points, respectively. These results highlight the value of learning recurrent patient memory for efficient longitudinal EHR modeling.

摘要:長期電子健康紀錄 (EHR) 建模需要將新的就診整合到不斷擴展的病歷中。然而,不斷累積的臨床資訊對大型語言模型 (LLMs) 在處理和保留完整病歷時帶來了日益增加的計算和記憶成本。一個實用的替代方案是逐次就診的重複壓縮,這將每次進來的就診納入一個緊湊且持續更新的病人記憶中。然而,在固定的記憶預算下,連續的更新必須整合新資訊,而不會逐漸失去後續任務所需的關鍵歷史證據。為了解決這一挑戰,我們提出了重複長期記憶 (ReLMem) 框架,該框架學習維持固定容量的病人記憶,以便使用凍結的 LLM 進行高效的下游預測。ReLMem 為這個 LLM 配備了輕量級的壓縮適配器,以便從其先前狀態和每次進來的就診中重複更新記憶,而無需重新閱讀早期記錄。具體而言,我們開發了一種多粒度優化策略,以在重複更新過程中保留與任務相關的資訊,並支持從最終記憶中進行下游預測。中間監督將來自壓縮記憶和完整歷史的注意力輸出對齊在相同查詢下,而預測監督則最小化與最終記憶條件下的真實答案之間的交叉熵。在基於 EHR 的藥物預測中,ReLMem 的 F1 分數接近完整歷史基準,同時將平均保留的歷史存儲減少了 97.1%。在相同的記憶預算下,它分別提高了宏觀和微觀 F1 分數,超過最強壓縮記憶基準 4.66 和 4.75 個百分點。這些結果突顯了學習重複病人記憶在高效長期 EHR 建模中的價值。

Raw Imagery Impacting Your AI: Should You Care?

2609.38265v1 by Adrien Dorise, Marjorie Bellizzi, Stéphane May

Onboard AI is gaining interest for space applications such as vessel, wildfire, and cloud detection, where real-time processing can improve mission reactivity and reduce downlink needs. However, onboard models may operate on raw or minimally processed imagery rather than on restored ground products. This study evaluates how image degradation affects object detection by varying Signal-to-Noise Ratio (SNR), Modulation Transfer Function (MTF) at Nyquist, and Ground Sampling Distance (GSD). Controlled degradations are applied to Very High Resolution Maxar imagery, and three lightweight detectors, YOLOv5s, YOLOX-S, and NanoDet, are evaluated on the resulting operating points. The results show that the impact of image quality depends on the degradation mechanism, and that increasing degradation does not necessarily lead to a proportional decrease in vessel detection performance. GSD produces the most consistent performance shift, while MTF and SNR effects depend more on the model and resolution. Severe combinations of blur and noise produce the largest losses. These results provide task-level information that can support sensor, processing, and AI trade-offs for future onboard systems.

摘要:在太空應用中,機載人工智慧正受到關注,例如船隻、野火和雲層檢測,其中即時處理可以提高任務反應能力並減少下行鏈路需求。然而,機載模型可能在原始或經過最小處理的影像上運作,而不是在恢復的地面產品上。本研究評估影像退化如何影響物體檢測,通過改變信噪比(SNR)、奈奎斯特的調變傳遞函數(MTF)和地面取樣距離(GSD)。對非常高解析度的Maxar影像施加控制退化,並在結果操作點上評估三個輕量級檢測器,YOLOv5s、YOLOX-S和NanoDet。結果顯示,影像質量的影響取決於退化機制,且增加退化不一定會導致船隻檢測性能成比例下降。GSD產生最一致的性能變化,而MTF和SNR的影響則更多地依賴於模型和解析度。模糊和噪聲的嚴重組合會產生最大的損失。這些結果提供了任務級別的信息,可以支持未來機載系統的傳感器、處理和人工智慧的權衡。

Do Agent Benchmarks Do What They Say? An Executable-Contract Audit of Tool-Using Agent Environments

2609.37315v1 by Rohith Reddy Bellibatlu, Zichong Wang, Wenbin Zhang

Tool-using agents are entering settings where a wrong action carries real cost, and the benchmarks certifying them grade what each simulated tool call reports having done, assuming the tool did what its interface advertises. The audit taxonomies we survey publish no category for that assumption, and a defect beneath a score is present on every rerun. We treat a tool's advertised surfaces as an executable contract, check the implementation against it, and trace each score's provenance through the task files and evaluator code to the verdicts that derive from state a defective tool should have written. Across 34 audited mutating tools in four benchmarks we confirm seven tool defects and one evaluator property at pinned commits. On injected defects the checker raised no false positive in 25 flags, flagged 2 of 5 negative controls, and missed most: in 29 of 33 scored misses a clause covered the defect but no probe revealed it. The checker's own static half, run alone, flags 14 of 17 confirmed sites, so on these findings the dynamic half confirms and traces rather than discovers. Twelve further AgentDojo tools, with six held-out tools and the seven audited first, complete its 25-tool mutating surface, on which at least 5 tools diverge from their advertised surface as our contracts read it, a rate for AgentDojo alone. No gold trajectory reaches either tau2-bench defect; on 1,120 paths built to isolate the telecom defect, a number fixed by construction, the evaluator rewards a refuel of a suspended line and fails the repaired tool. The clearest case is a clinical benchmark whose tool tells the agent each write executed under a documented no-write design its interface does not disclose; its grader takes that message as evidence, so its action success rate records whether a request carried the expected payload, not whether any record changed.

摘要:使用工具的代理正在進入錯誤行動會帶來實際成本的環境,而認證它們的基準會評分每個模擬工具調用所報告的操作,假設該工具執行了其介面所宣傳的功能。我們調查的審核分類中沒有針對該假設的類別,並且每次重新運行時都會出現低於分數的缺陷。我們將工具宣傳的表面視為可執行的合約,檢查其實現是否符合,並追蹤每個分數的來源,通過任務檔案和評估代碼到應該由有缺陷工具寫出的判決。在四個基準中對34個經審核的變異工具進行的確認中,我們確認了七個工具缺陷和一個評估者屬性,這些都是在固定的提交上進行的。在注入的缺陷中,檢查器在25個標誌中沒有產生假陽性,標記了5個負控制中的2個,並且大多數都錯過了:在33次得分的錯過中,有29次條款涵蓋了缺陷,但沒有探測揭示它。檢查器自己的靜態部分,單獨運行時,標記了17個確認位置中的14個,因此根據這些發現,動態部分確認並追蹤,而不是發現。另外12個AgentDojo工具,與六個保留的工具和最初的七個經審核工具一起,完成了其25個工具的變異表面,其中至少有5個工具與我們合約所讀取的宣傳表面不同,這是AgentDojo單獨的比率。沒有金色軌跡達到任何tau2-bench缺陷;在1,120條為隔離電信缺陷而構建的路徑中,這個數字是由建設固定的,評估者對一條暫停線的重新加油給予獎勵,卻未能通過修復的工具。最明顯的案例是一個臨床基準,其工具告訴代理每次寫入在其介面未披露的文檔中執行的無寫入設計下進行;其評分者將該消息視為證據,因此其行動成功率記錄請求是否攜帶預期的有效載荷,而不是任何記錄是否發生變更。

Information Bottleneck-Guided Adaptive Hypergraph Transformer for Brain Disease Diagnosis

2609.37220v1 by Jingxi Feng, Xudong Chen, Yifan Zhang, Heming Xu, Hongcheng Han, Xijing Wang, Dong Zhang, Shaoyi Du

Exploring high-order correlations and long-range dependencies in brain networks holds significant value for both neuroscience research and clinical diagnosis. However, previous studies have lacked a unified integration of high-order and long-range dependency information in brain networks, and there is substantial redundancy behind various types of information. These issues limit their effectiveness in the diagnosis of brain diseases. To address this, we propose an Information Bottleneck-Guided Adaptive HyperGraph Transformer (IBAHGT). By incorporating the information bottleneck (IB) principle, this approach enables adaptive learning of high-order correlations and both short- and long-range dependencies within a unified framework for brain network analysis, achieving high-precision brain disease diagnosis. IBAHGT consists of three key components: an information bottleneck-guided adaptive hypergraph convolution, which introduces a novel hypergraph information bottleneck (HIB) principle to adaptively learn hypergraph message-passing weights between nodes and hyperedges, optimizes information flow and captures high-order information in brain networks that is maximally informative and minimally redundant (MIMR). The Transformer encoder captures global information within brain networks through the attention mechanism, specifically modeling short- and long-range dependencies. An information bottleneck-guided node-level adaptive fusion employs the IB principle to learn independent weights for each node, facilitating the fine-grained integration of high-order information and global information to obtain an efficient representation for downstream tasks. Extensive experiments demonstrate that the proposed method outperforms current state-of-the-art methods and can identify biomarkers for clinical applications.

摘要:探索大腦網絡中的高階相關性和長程依賴性對於神經科學研究和臨床診斷具有重要價值。然而,先前的研究缺乏對大腦網絡中高階和長程依賴信息的統一整合,並且各類信息之間存在大量冗餘。這些問題限制了它們在大腦疾病診斷中的有效性。為了解決這個問題,我們提出了一種信息瓶頸引導的自適應超圖Transformer(IBAHGT)。通過納入信息瓶頸(IB)原則,這種方法能夠在統一框架內自適應學習高階相關性以及短程和長程依賴性,實現高精度的大腦疾病診斷。IBAHGT由三個關鍵組件組成:一個信息瓶頸引導的自適應超圖卷積,該組件引入了一種新穎的超圖信息瓶頸(HIB)原則,以自適應地學習節點和超邊之間的超圖信息傳遞權重,優化信息流並捕捉大腦網絡中最大信息量和最小冗餘的高階信息(MIMR)。Transformer編碼器通過注意機制捕捉大腦網絡中的全局信息,特別是建模短程和長程依賴性。信息瓶頸引導的節點級自適應融合利用IB原則為每個節點學習獨立權重,促進高階信息和全局信息的細緻整合,以獲得下游任務的高效表示。廣泛的實驗表明,所提出的方法優於當前的最先進方法,並能夠識別臨床應用的生物標記。

Physics-Informed Multi-Agent Coordination for Hospital Patient Flow Optimization

2609.37022v1 by Guoqing Zhang, Rafik Hadfi, Takayuki Ito

Efficient patient flow coordination across autonomous hospital departments is critical for mitigating overcrowding and balancing resource utilization. While classical queueing theory, specifically open Baskett--Chandy--Muntz--Palacios (BCMP) networks, provides an interpretable mathematical topology for healthcare operations, analytical models rely on stationary assumptions and fixed routing matrices that degrade under state-dependent real-world dynamics. Conversely, centralized reinforcement learning approaches struggle to accommodate the decentralized structure of hospital governance, where individual clinical departments function with localized observations, heterogeneous resources, and divergent operational objectives. In this paper, we present a Multi-Agent Systems (MAS) framework titled \emph{Physics-Informed Multi-Agent Coordination}, which embeds empirically calibrated BCMP queueing topologies as physical priors within a decentralized multi-agent reinforcement learning architecture. Formulated as a Decentralized Partially Observable Markov Decision Process (Dec-POMDP) under coupled resource constraints, our method enables autonomous departmental agents to cooperatively negotiate patient routing and dynamic service scaling. To mitigate environmental non-stationarity without inducing excessive communication overhead, agents exchange localized action fingerprints along network edges and optimize a spatially decomposed reward structure. Empirical evaluations driven by real-world MIMIC-IV patient trajectories indicate that this cooperative multi-agent approach substantially reduces cumulative system delay compared to static Markovian approximations, heuristic dispatching, and independent multi-agent baselines, while maintaining clinical safety constraints.

摘要:有效的病人流動協調在自主醫院部門之間對於減輕擁擠和平衡資源利用至關重要。雖然經典的排隊理論,特別是開放的Baskett--Chandy--Muntz--Palacios (BCMP) 網絡,為醫療運營提供了可解釋的數學拓撲,但分析模型依賴於靜態假設和固定的路由矩陣,這在狀態依賴的現實世界動態下會退化。相反,集中式強化學習方法難以適應醫院治理的去中心化結構,其中各個臨床部門以本地觀察、異質資源和不同的運營目標運作。在本文中,我們提出了一個名為\emph{Physics-Informed Multi-Agent Coordination}的多智能體系統(MAS)框架,該框架將經驗校準的BCMP排隊拓撲作為物理先驗嵌入到去中心化的多智能體強化學習架構中。該方法在耦合資源約束下被表述為去中心化的部分可觀察馬爾可夫決策過程(Dec-POMDP),使自主部門智能體能夠合作協商病人路由和動態服務擴展。為了減輕環境非平穩性而不引入過多的通信開銷,智能體沿著網絡邊緣交換本地化的行動指紋,並優化一個空間分解的獎勵結構。基於現實世界MIMIC-IV病人軌跡的實證評估表明,這種合作的多智能體方法在保持臨床安全約束的同時,顯著減少了累積系統延遲,相較於靜態馬爾可夫近似、啟發式調度和獨立多智能體基準。

STAR-GRPO: Canonical Anchoring and Reliability-First Advantages against Representation-Dependent Reward Hacking

2609.36900v1 by Wan Tian, Zhongyi Li, Xiang Xu, Minhao Zou, Yijie Peng, Fuzhen Zhuang

Reward hacking occurs when policy optimization exploits a brittle reward interface or an overly permissive proxy objective, improving the training score without improving the underlying response quality. This phenomenon is amplified in group-relative policy optimization: an unsupported reward can shift the group baseline and alter the updates of other rollouts, while post-hoc or purely relative weighting cannot represent group-wide uncertainty. We propose \emph{Self-Tuned Anchored Reliability Group-Relative Policy Optimization} (STAR-GRPO), a reliability-first advantage estimator based on paired assessments of the same rollout. STAR separates the quality signal from its learning influence: score disagreement determines rollout reliability, relative reliability enters a self-tuned robust location--scale fit before group normalization, and absolute group reliability attenuates the resulting bounded advantage. The analysis establishes coordinate and second-moment bounds, characterizes exact centering through the weighted location equation, and gives reliability-dependent attenuation guarantees for outlying rewards. We evaluate STAR-GRPO in two complementary reward-hacking regimes. In token-interface exploitation, STAR prevents runaway optimization of the deployed-interface score while improving the canonical quality signal. In rubric-proxy overoptimization for medical reasoning, STAR improves independent semantic evaluation, narrows the proxy--judge discrepancy, and reduces overclaim while optimizing the same task proxy. Together, these results show that reliability-first normalization offers a principled way to limit unsupported reward influence on both group baselines and policy updates, while retaining the task reward as the optimization target.

摘要:獎勵駭客行為發生在政策優化利用脆弱的獎勵介面或過於寬鬆的代理目標時,這樣可以提高訓練分數,但並未改善基礎的反應質量。這一現象在群體相對政策優化中被放大:不受支持的獎勵可以改變群體基準並改變其他回合的更新,而事後或純相對的加權無法代表整個群體的不確定性。我們提出了\emph{自調整錨定可靠性群體相對政策優化}(STAR-GRPO),這是一種基於相同回合配對評估的可靠性優先優勢估計器。STAR將質量信號與其學習影響分開:分數不一致性決定回合的可靠性,相對可靠性進入自調整的穩健位置--尺度擬合,然後進行群體正規化,絕對群體可靠性減弱了結果的有界優勢。分析建立了坐標和二階矩界限,通過加權位置方程描述精確的中心化,並為異常獎勵提供了依賴於可靠性的減弱保證。我們在兩種互補的獎勵駭客制度中評估了STAR-GRPO。在令牌介面利用中,STAR防止了已部署介面分數的失控優化,同時改善了經典質量信號。在醫療推理的評分代理過度優化中,STAR改善了獨立語義評估,縮小了代理--評審之間的差距,並在優化相同任務代理的同時減少了過度聲明。總體而言,這些結果表明,可靠性優先的正規化提供了一種原則性的方法,以限制不受支持的獎勵對群體基準和政策更新的影響,同時保留任務獎勵作為優化目標。

Automated Screw Planning for Reduced Pelvic Fractures Based on Statistical Shape Models and Deep Learning

2609.36847v1 by Yang Gao, Sutuke Yibulayimu, Yanzhen Liu, Zian Zhao, Yudi Sang

Percutaneous iliosacral screw fixation is an important minimally invasive treatment for unstable pelvic fractures. Because the sacroiliac region has complex anatomy and narrow screw corridors, the accuracy and safety of screw placement directly affect surgical outcomes. Accurate and reliable preoperative screw planning is therefore essential to improve surgical success and reduce intraoperative risks. Conventional preoperative planning typically requires surgeons to determine screw trajectories through manual measurements, a labor-intensive process that depends on subjective clinical experience. To address these challenges, we propose a fully automated pipeline for preoperative iliosacral screw planning in patients with pelvic fractures. Using patient-specific three-dimensional anatomy, the pipeline automatically identifies safe screw corridors and generates individualized insertion trajectories to support clinical preoperative planning. We evaluated the proposed pipeline on 200 clinical cases of pelvic fractures. Compared with conventional manual measurements, the safety margin of the safe insertion corridors increased by 2% across the four screw types, the mean planning time decreased by more than 90%, and the clinical acceptance rate reached 95%.

摘要:經皮髖骶螺釘固定是一種重要的微創治療不穩定骨盆骨折的方法。由於骶髂區域具有複雜的解剖結構和狹窄的螺釘通道,螺釘放置的準確性和安全性直接影響手術結果。因此,準確可靠的術前螺釘規劃對於提高手術成功率和降低術中風險至關重要。傳統的術前規劃通常要求外科醫生通過手動測量來確定螺釘的軌跡,這是一個依賴主觀臨床經驗的勞動密集型過程。為了解決這些挑戰,我們提出了一個完全自動化的術前髖骶螺釘規劃流程,專為骨盆骨折患者設計。該流程利用患者特定的三維解剖結構,自動識別安全的螺釘通道並生成個性化的插入軌跡,以支持臨床術前規劃。我們在200例骨盆骨折的臨床案例中評估了所提出的流程。與傳統手動測量相比,安全插入通道的安全邊際在四種螺釘類型中增加了2%,平均規劃時間減少了90%以上,臨床接受率達到95%。

How Medical VLMs Underutilize Their Vision Encoders: A Dermatology Perspective

2609.36557v1 by Janet Wang, Yunbei Zhang, Xiao Wang, Jihun Hamm

Medical Vision-Language Models (VLMs) show significant promise for clinical image understanding, offering accurate diagnosis with interpretable reasoning. However, a critical performance gap exists between their strong vision encoders and the full multimodal model: in dermatology, the MedSigLIP encoder outperforms MedGemma by an average of 10.26 percentage points even when both use zero target-task labels; few-shot linear probing provides further evidence of strong visual representations. This gap motivates an investigation of how visual information is used in end-to-end diagnosis and why plausible-sounding predictions can lack grounding in image evidence. Using dermatology as our primary testbed, we systematically investigate three hypotheses for this phenomenon. We further provide a mechanistic analysis of the model's internal attention patterns, showing that a simple describe-then-decide prompting strategy increases vision attention by 30-40% during generation. Task-specific fine-tuning improves dermatology classification but reduces cross-domain medical question-answering performance in our evaluation. To address these challenges, we combine label-free prompting with low-label encoder-assisted reranking while keeping the VLM frozen. We validate the interventions across five VLM backbones in dermatology and provide supporting representation and attention analyses across additional medical modalities.

摘要:醫療視覺-語言模型(VLMs)在臨床影像理解方面顯示出顯著的潛力,提供可解釋的推理以進行準確診斷。然而,它們強大的視覺編碼器與完整的多模態模型之間存在一個關鍵的性能差距:在皮膚科,MedSigLIP 編碼器的表現平均超過 MedGemma 10.26 個百分點,即使兩者都使用零目標任務標籤;少量樣本線性探測進一步證明了強大的視覺表徵。這一差距促使我們調查視覺信息在端到端診斷中的使用方式,以及為什麼聽起來合理的預測可能缺乏影像證據的支持。以皮膚科作為我們的主要測試平台,我們系統地調查了這一現象的三個假設。我們進一步提供了模型內部注意力模式的機制分析,顯示簡單的描述-再決策提示策略在生成過程中將視覺注意力提高了 30-40%。特定任務的微調改善了皮膚科分類,但在我們的評估中降低了跨領域醫療問答的表現。為了解決這些挑戰,我們結合無標籤提示與低標籤編碼器輔助的重新排序,同時保持 VLM 凍結。我們在皮膚科的五個 VLM 骨幹上驗證了這些干預,並提供了支持的表徵和注意力分析,涵蓋其他醫療模態。

Reliability Testing of Medical Model Performance under Distributed Deployment

2609.36525v1 by Yifei Wang, Xiaohan Zhang, Youtao Ding, Tianlin Li, Xiaoyu Zhang, Yida Yang, Li Pan

Distributed inference has become an indispensable part of deploying medical models under practical latency, memory, and throughput constraints. Although modern frameworks improve serving efficiency through tensor parallelism, mixed precision, kernel fusion, and multi-device communication, they are generally assumed to preserve the behavior observed during centralized HuggingFace evaluation. This assumption creates an evaluation-deployment mismatch: a model may pass offline evaluation but produce a different output after the execution stack changes. To address this mismatch, we propose a testing framework and an improved, distributed-execution-sensitive medical-model benchmark that evaluates the same checkpoint and input under a centralized HuggingFace reference and matched distributed deployments. Extensive experiments across language, vision, and multimodal medical models show that execution changes can produce measurable output disagreements. Across supported visual settings, the test success rate ranges from 0.21 to 0.43 for single-modality models and from 0.32 to 0.98 for multimodal models. The benchmark is aimed at extending medical-model evaluation from capability and security to evaluation-deployment consistency.

摘要:分散推理已成為在實際延遲、記憶體和吞吐量限制下部署醫療模型不可或缺的一部分。儘管現代框架通過張量並行、混合精度、內核融合和多設備通信來提高服務效率,但通常假設它們能保持在集中式 HuggingFace 評估中觀察到的行為。這一假設造成了評估與部署之間的不匹配:一個模型可能在離線評估中通過,但在執行堆棧變更後產生不同的輸出。為了解決這一不匹配,我們提出了一個測試框架和一個改進的、對分散執行敏感的醫療模型基準,該基準在集中式 HuggingFace 參考和匹配的分散部署下評估相同的檢查點和輸入。針對語言、視覺和多模態醫療模型的廣泛實驗顯示,執行變更可能會產生可測量的輸出不一致。在支持的視覺設置中,單模態模型的測試成功率範圍為 0.21 到 0.43,而多模態模型的範圍為 0.32 到 0.98。該基準旨在將醫療模型評估從能力和安全性擴展到評估與部署的一致性。

BRIDGE: Bilevel Retrieval-Credit-Aware Agentic Reinforcement Learning

2609.36505v1 by Quan Xiao, Mingda Liu, Gaowen Liu, Katsuki Fujisawa, Tianyi Chen

Agentic reinforcement learning (ARL) with verifiable rewards improves the ability of large language models (LLMs) to tackle knowledge-intensive tasks by learning to interleave search and reasoning. However, most existing ARL methods optimize only LLM-generated tokens and treat retrieved evidence as environment observations. This creates an information-credit gap: failures caused by missing or misleading evidence are attributed to the LLM policy rather than to the retriever, which motivates training the LLM and the retriever jointly. In this paper, we show that retrieval and LLM policy learning are order-sensitive: adapting the retriever before optimizing the policy yields a larger reward gain than the reverse order. To preserve this hierarchy while allowing both components to co-adapt, we formulate retrieval-augmented agentic RL as a bilevel optimization problem. To solve it efficiently, we introduce BRIDGE, a memory-efficient first-order bilevel method motivated by a loss-landscape analysis of the RL and retrieval objectives. Across seven open-domain QA benchmarks, BRIDGE achieves the highest average accuracy with both 3B and 7B backbones, improving the multi-hop average over the strongest baseline by 9.6 and 3.4 EM points, respectively. It also achieves the best averaged answer accuracy and reasoning quality across medical QA benchmarks.

摘要:代理強化學習(ARL)與可驗證獎勵相結合,通過學習交替進行搜索和推理,提高了大型語言模型(LLMs)處理知識密集型任務的能力。然而,大多數現有的ARL方法僅優化LLM生成的標記,並將檢索到的證據視為環境觀察。這造成了一個信息信用差距:由於缺失或誤導性證據而導致的失敗被歸因於LLM策略,而不是檢索器,這促使了LLM和檢索器的聯合訓練。在本文中,我們展示了檢索和LLM策略學習對順序的敏感性:在優化策略之前調整檢索器,會比反向順序產生更大的獎勵增益。為了在允許兩個組件共同適應的同時保留這一層次結構,我們將檢索增強的代理RL公式化為一個雙層優化問題。為了高效解決它,我們引入了BRIDGE,一種基於RL和檢索目標的損失景觀分析的記憶高效的一階雙層方法。在七個開放域問答基準上,BRIDGE在3B和7B骨幹上都達到了最高的平均準確率,分別提高了對最強基線的多跳平均9.6和3.4 EM點。它還在醫學問答基準上實現了最佳的平均答案準確率和推理質量。

ARCagent: An Adaptive Retrieval Calibration Agent for Clinical Question Answering

2609.36392v1 by Yuyan Chen

In diseases where clinical guidelines are incomplete, contested, or mutually contradictory, knowledge completeness and dynamic conflict-aware synthesis are two safety-critical properties that standard Retrieval-Augmented Generation systems do not provide. Therefore, we present \sysname, an adaptive retrieval calibration clinical question-answering agent for ME/CFS, a disease where diagnostic frameworks coexist and major guidelines actively contradict each other on treatment. ARCagent contributes three components. First, a 1,706-chunk, 10-source knowledge base with a structured inter-guideline conflict registry spanning all active ME/CFS diagnostic frameworks. Second, a conflict-aware retrieval calibration pipeline that re-ranks retrieved evidence using query-specific focus and conflict signals. Third, a benchmark scored by LLM-as-Judge, avoiding systematic underestimation averaging 10.1 percentage points caused by keyword matching. ARCagent achieves 95.3%, outperforming all base LLMs. Code is available at https://github.com/Yukyin/ARCagent.

摘要:在臨床指導方針不完整、存在爭議或相互矛盾的疾病中,知識的完整性和動態衝突意識合成是標準增強檢索生成系統所不具備的兩個安全關鍵特性。因此,我們提出了\sysname,一個針對ME/CFS的自適應檢索校準臨床問答代理,這是一種診斷框架共存且主要指導方針在治療上積極矛盾的疾病。ARCagent貢獻了三個組件。首先,一個包含1,706個片段、10個來源的知識庫,擁有一個結構化的指導方針間衝突登記,涵蓋所有活躍的ME/CFS診斷框架。其次,一個衝突意識的檢索校準管道,使用查詢特定的焦點和衝突信號重新排名檢索到的證據。第三,一個由LLM-as-Judge評分的基準,避免了由關鍵字匹配造成的系統性低估,平均低10.1個百分點。ARCagent達到了95.3%的準確率,超越了所有基礎LLM。代碼可在https://github.com/Yukyin/ARCagent獲得。

Quantization Enables Private Dense Retrieval against Malicious Service Providers

2609.36376v1 by Louis Tremblay Thibault, Sofiane Azogagh, Marc-Olivier Killijian, Ulrich Aïvodji

Dense retrieval, the key component of Retrieval Augmented Generation (RAG), retrieves the most relevant documents by comparing dense vector representations of queries and passages from a large corpus. In privacy-sensitive applications, the server observes the query and controls which evidence is returned, creating both confidentiality and integrity risks. We formulate private dense retrieval as providing query privacy and retrieval integrity against a malicious server, and develop a two-round cryptographic protocol that provides both guarantees. Our protocol reduces private and verifiable retrieval to multiplication of a committed matrix by an encrypted vector and uses low-bit quantization to make this computation practical. We evaluate the resulting trade-off between cryptographic cost, retrieval quality, and downstream RAG accuracy across six embedding models, four language models, and corpora of up to 2.68 million passages. Our results show that, with a clipped quantizer, three-bit quantization largely preserves retrieval quality and downstream accuracy, while a private query over a corpus the size of a clinical reference requires one to three minutes of server time. These results suggest that private dense retrieval is already practical for moderately sized, privacy-sensitive corpora when minute-scale latency is acceptable.

摘要:密集檢索,檢索增強生成(RAG)的關鍵組件,通過比較查詢和來自大型語料庫的段落的密集向量表示來檢索最相關的文檔。在隱私敏感的應用中,伺服器觀察查詢並控制返回哪些證據,從而產生保密性和完整性風險。我們將私密密集檢索定義為在惡意伺服器面前提供查詢隱私和檢索完整性,並開發了一種雙輪加密協議,提供這兩項保證。我們的協議將私密且可驗證的檢索簡化為將已承諾的矩陣與加密向量相乘,並使用低位量化使這一計算變得實用。我們評估了六種嵌入模型、四種語言模型和多達268萬段落的語料庫中,加密成本、檢索質量和下游RAG準確性之間的權衡。我們的結果顯示,使用剪裁量化器的三位量化在很大程度上保持了檢索質量和下游準確性,而對於大小相當於臨床參考的語料庫,私密查詢需要一到三分鐘的伺服器時間。這些結果表明,當分鐘級延遲是可以接受的時候,私密密集檢索對於中等大小的隱私敏感語料庫已經是實用的。

ThuRunel: Dynamic Decoupling for Structured Advisory Dialogue

2609.36340v1 by Yuyan Chen

High-stakes advisory domains such as medical aesthetics, legal consultation, and educational planning exhibit a two-phase structure. The early phase requires empathetic elicitation and emotional support, and the late phase requires authoritative specialist judgment. Neither fully automated agents nor human junior consultants adequately address this structure at scale. We formalize the core design challenge as dynamic decoupling, asking how an AI advisory agent should decide what to ask, when to stop, what to resolve autonomously, and what to forward to the specialist. We present ThuRunel, an advisory agent combining a finite-state belief management framework, a chain-of-thought teacher synthesis protocol, and learned generation adapters. Against eleven baselines, ThuRunel achieves consistent improvements in elicitation completeness and specialist brief quality. ThuRunel is publicly deployed as a bilingual web application in which the same decoupling decisions operate from the client's side, grounded in a curated knowledge base that cites its sources in every answer.

摘要:高風險的諮詢領域如醫療美學、法律諮詢和教育規劃展示出雙階段結構。早期階段需要同理心的引導和情感支持,而後期階段則需要權威專家的判斷。無論是完全自動化的代理還是人類初級顧問,都無法在規模上充分應對這一結構。我們將核心設計挑戰形式化為動態解耦,詢問AI諮詢代理應如何決定詢問什麼、何時停止、什麼可以自主解決以及什麼需要轉交給專家。我們提出了ThuRunel,一個結合有限狀態信念管理框架、思維鏈教師綜合協議和學習生成適配器的諮詢代理。在十一個基準測試中,ThuRunel在引導完整性和專家簡報質量上實現了一致的改進。ThuRunel作為一個雙語網絡應用程序公開部署,其中相同的解耦決策從客戶端運作,基於一個策劃的知識庫,並在每個答案中引用其來源。

SCOUT: Synergizing Reasoning and Tool-Use for Computer-Use Safety

2609.36201v1 by Jianxing Chen, Xiao Yu, Shipra Agrawal, Zhou Yu

Computer-use agents (CUAs), while capable of completing computer tasks in everyday and professional workflows, can cause unintended harm even under benign instructions and environments. However, detecting such harm remains challenging. First, it requires careful, task-specific reasoning: verifiers guided only by general safety criteria often overlook many important but subtle harmful behaviors. Second, it requires active investigation: past trajectory screenshots show what the agent did but not always what actually changed in the environment, so LLM-as-a-judge verifiers that rely on screenshots alone may be unable to determine the actual consequences of actions. To address these challenges, we introduce SCOUT, a two-stage agentic safety verifier that synergizes reasoning-intensive rubric generation with tool-intensive evidence gathering. First, our SCOUT rubric generator extensively reasons over the task and the agent's trajectory to determine what successful and safe execution should entail, generating task-specific completion and safety rubrics. Then, our SCOUT probing agent follows these rubrics to interact with the post-execution environment and collect grounded evidence for final safety and completion judgments. We evaluate our framework on two computer-use safety benchmarks. On AutoElicit-Bench, SCOUT achieves 75.4 unsafe F1 and 74.5 completion F1, outperforming LLM-as-a-judge verifiers and naive tool-use verifiers. SCOUT leads on OS-Blind with 76.4% unsafe detection accuracy. Test-time reflection reduces final unsafe execution rates from 30.2% to 17.2% on AutoElicit-Bench. Ablations and analysis show that tool-free rubric generation in SCOUT elicits substantially more reasoning and is crucial for safety detection across verifier backbones, especially non-frontier ones. A preliminary extension to coding tasks shows that SCOUT can support safety verification beyond computer-use.

摘要:電腦使用代理(CUAs)雖然能夠在日常和專業工作流程中完成電腦任務,但即使在良性指令和環境下,也可能造成意想不到的傷害。然而,檢測這種傷害仍然具有挑戰性。首先,它需要仔細的、特定任務的推理:僅依賴一般安全標準的驗證者往往忽視許多重要但微妙的有害行為。其次,它需要主動調查:過去的軌跡截圖顯示代理所做的事情,但不總是顯示環境中實際發生的變化,因此僅依賴截圖的LLM作為判斷者的驗證者可能無法確定行動的實際後果。為了解決這些挑戰,我們引入了SCOUT,一種兩階段的代理安全驗證器,將推理密集的標準生成與工具密集的證據收集相結合。首先,我們的SCOUT標準生成器對任務和代理的軌跡進行廣泛推理,以確定成功和安全執行應該包含什麼,生成特定任務的完成和安全標準。然後,我們的SCOUT探測代理根據這些標準與執行後的環境互動,並收集基於證據的最終安全和完成判斷。我們在兩個電腦使用安全基準上評估了我們的框架。在AutoElicit-Bench上,SCOUT達到了75.4的危險F1和74.5的完成F1,超越了LLM作為判斷者的驗證者和天真的工具使用驗證者。SCOUT在OS-Blind上以76.4%的危險檢測準確率領先。測試時反思將AutoElicit-Bench上的最終危險執行率從30.2%降低到17.2%。消融和分析顯示,SCOUT中的無工具標準生成引發了顯著更多的推理,並且對於各種驗證者骨幹,尤其是非前沿的驗證者,對於安全檢測至關重要。對編碼任務的初步擴展顯示,SCOUT可以支持超越電腦使用的安全驗證。

PHASE: A Physiology-Guided Hierarchical Foundation Model for Intracranial EEG

2609.36087v2 by Yipeng Zhang, Chenda Duan, Yuanyi Ding, Tianyi Wang, Atsuro Daida, Masaki Izumi, Yuta Tanoue, Naoto Kuroda, Shaun A. Hussain, Nishant Sinha, Eishi Asano, Hiroki Nariai, Vwani Roychowdhury

Clinicians and neuroscientists have long analyzed intracranial electroencephalography (iEEG) through directly measurable physiological characteristics, which carry much of the information that downstream tasks depend on. Recent iEEG foundation models learn by reconstructing or predicting their inputs, which leaves the retention of these characteristics implicit. They are also evaluated mainly on cognitive decoding and a narrow clinical task, i.e., seizure detection. On a broad, clinically relevant benchmark such as Omni-iEEG, they remain below task-specific models when used frozen. We introduce PHASE, a physiology-guided foundation model that makes these characteristics explicit learning targets, pairing them with masked latent prediction in a temporal stage (PHASE-T) within each channel and a spatiotemporal stage (PHASE-ST) across synchronized channels. PHASE is pretrained on heterogeneous recordings from 222 participants at nine clinical sites. On all five Omni-iEEG clinical tasks, frozen PHASE-T outperforms every evaluated foundation model by up to 31\%, and fine-tuned PHASE-T surpasses the task-specific models, setting a new state of the art. PHASE-T benefits from physiological supervision, outperforming variants trained with latent prediction alone or auxiliary waveform reconstruction on every task in matched ablations. PHASE-T generalizes to unseen institutions, outperforming the compared models with few or no local labels. PHASE-ST further improves seizure-onset-zone identification over PHASE-T and, when frozen, decodes sound volume and pitch on BrainTreebank better than published models. Beyond task performance, PHASE learns to encapsulate the physiological characteristics clinicians recognize, from seizure onset and its propagation to anatomical region identity, even though its pretraining contains no ictal recordings or anatomical labels.

摘要:臨床醫生和神經科學家長期以來一直通過直接可測量的生理特徵分析顱內腦電圖 (iEEG),這些特徵攜帶了許多下游任務所依賴的信息。最近的 iEEG 基礎模型通過重建或預測其輸入來學習,這使得這些特徵的保留變得隱含。它們的評估主要集中在認知解碼和一個狹窄的臨床任務,即癲癇發作檢測。在像 Omni-iEEG 這樣的廣泛臨床相關基準上,當以凍結狀態使用時,它們的表現仍低於特定任務模型。我們介紹了 PHASE,一種生理引導的基礎模型,將這些特徵作為明確的學習目標,並在每個通道的時間階段 (PHASE-T) 和同步通道之間的空間時間階段 (PHASE-ST) 中將其與掩蔽潛在預測配對。PHASE 在九個臨床站點的 222 名參與者的異質錄音上進行了預訓練。在所有五個 Omni-iEEG 臨床任務中,凍結的 PHASE-T 的表現超過了每個評估的基礎模型,最高可達 31\%,而微調的 PHASE-T 超越了特定任務模型,創造了新的最先進水平。PHASE-T 受益於生理監督,在每個匹配的消融實驗中,超越了僅用潛在預測或輔助波形重建訓練的變體。PHASE-T 在未見過的機構中具有良好的泛化能力,超越了比較模型,並且幾乎沒有本地標籤。PHASE-ST 進一步改善了癲癇發作區域的識別,超過了 PHASE-T,並且在凍結狀態下,對 BrainTreebank 的聲音音量和音高的解碼表現優於已發表的模型。除了任務性能外,PHASE 學會了封裝臨床醫生識別的生理特徵,從癲癇發作及其傳播到解剖區域身份,即使其預訓練中不包含任何癲癇發作錄音或解剖標籤。

IMC-CLINIC: Coupled Loss-Informed Newton Iterations for Clipping in Analog In-Memory Computing

2609.35586v1 by Yung-Chin Chen, Chia-Yu Chen, Naveen Verma

Analog in-memory computing (IMC) offers a promising path toward energy-efficient large language model (LLM) inference by executing matrix multiplications (MatMul) directly within memory arrays in the analog domain. Its efficiency, however, comes with an additional source of error: limited-precision analog-to-digital converters (ADCs) quantize accumulated analog partial sums, introducing output-side error distinct from conventional activation and weight quantization at the MatMul inputs. Clipping can mitigate both operand and ADC quantization errors, but the optimal clipping factors must jointly balance activation rounding and clipping, weight rounding and clipping, and ADC quantization. Existing clipping methods, designed for digital quantization, do not explicitly optimize these coupled sources of IMC error and often rely on costly search-based calibration. We introduce IMC-CLINIC (Coupled Loss-Informed Newton Iterations for Clipping), a clipping calibration framework based on an analytical surrogate for IMC MatMul output error. The surrogate jointly models operand quantization, accumulated clipping-induced bias, and ADC quantization, enabling efficient evaluation of its gradient and approximate curvature from a small calibration set. IMC-CLINIC jointly optimizes activation and weight clipping factors using a safeguarded Newton-type method. Across multiple models and datasets, it improves average zero-shot accuracy by 6.5-11.5 percentage points over the grid search baseline while reducing calibration time by factors of 10.0-12.1. Its analytical surrogate closely tracks empirical IMC output error, and its optimizer is certified within 1% of the global optimum under the loss objective across all projections on two representative models.

摘要:類比記憶體計算(IMC)提供了一條有前景的途徑,以實現能效高的大型語言模型(LLM)推斷,通過在類比領域內直接在記憶體陣列中執行矩陣乘法(MatMul)。然而,它的效率伴隨著一個額外的誤差來源:有限精度的類比轉數字轉換器(ADC)對累積的類比部分和進行量化,這引入了與傳統的激活和權重量化在MatMul輸入時不同的輸出端誤差。剪裁可以減輕操作數和ADC量化誤差,但最佳剪裁因子必須共同平衡激活四捨五入和剪裁、權重四捨五入和剪裁,以及ADC量化。現有的剪裁方法旨在數位量化,並未明確優化這些耦合的IMC誤差來源,並且通常依賴於昂貴的基於搜索的校準。我們介紹IMC-CLINIC(耦合損失信息牛頓迭代剪裁),這是一個基於IMC MatMul輸出誤差的分析替代品的剪裁校準框架。該替代品共同建模操作數量化、累積的剪裁引起的偏差和ADC量化,從而能夠高效地評估其梯度和從小型校準集獲得的近似曲率。IMC-CLINIC使用受保護的牛頓類型方法共同優化激活和權重剪裁因子。在多個模型和數據集上,它將平均零-shot準確率提高了6.5-11.5個百分點,相較於網格搜索基準,並將校準時間減少了10.0-12.1倍。其分析替代品與實證IMC輸出誤差密切相關,並且其優化器在所有兩個代表性模型的損失目標下,經過所有投影後,證明在全球最優解的1%內。

RareDx: Controlled Knowledge Integration and Graph-Grounded Policy Optimization for Rare-Disease Diagnosis

2609.35549v2 by Bo Zhang, Yuchen Wang, Dongbai Li, Matthew Yu Heng Wong, Qingkai Zeng, Lijun Wang, Tien-Yin Wong, Peng Cui, Tianyu Liu

Rare-disease diagnosis is a long-tail reasoning problem: phenotypes are incomplete, individual disorders are sparsely documented, and relevant evidence is distributed across ontologies, gene annotations, and biomedical text. Language models consequently favor common conditions, miss rare candidates, or produce plausible but invalid names. We introduce RareDx, which couples controlled evidence use with knowledge-graph-grounded policy optimization. RareDx-Harness normalizes heterogeneous records into one ranked-diagnosis task and compares direct inference, static retrieval, adaptive tools, and structured phenotype-gene-disease reasoning over a shared knowledge layer. The training pipeline combines Top-10 post-training with RareDx-KGPO, our knowledge-graph-grounded policy optimization method. Its reward projects predictions into a canonical disease graph and integrates curated graded relevance, ontology proximity, biomedical similarity, and phenotype consistency. Vocabulary and output-budget constraints prevent dense partial credit from rewarding fabricated or overlong differentials. Across eight benchmarks, the complete RareDx system centered on Qwen3.5-9B reaches 38.34 macro Hit@10, 1.60 points above GPT-5.5 under the archived protocol; a disjoint validation-selection audit retains a 6.80-point routing gain over Direct on held-out cases. The 27B system reaches 23.53/36.56/40.76 at Hit@1/5/10. Controlled ablations show that retrieval is not uniformly helpful and that controlled routing is central to the gain. These results indicate that structured medical knowledge can turn a compact model into a competitive diagnostic ranker across heterogeneous long-tail settings in clinical practice.

摘要:罕見疾病的診斷是一個長尾推理問題:表型不完整,個別疾病的文獻記錄稀少,相關證據分散在本體論、基因註釋和生物醫學文本中。因此,語言模型偏向於常見病症,錯過罕見候選者,或產生看似合理但無效的名稱。我們介紹了RareDx,它將受控證據使用與基於知識圖的政策優化結合起來。RareDx-Harness將異質記錄標準化為一個排名診斷任務,並比較直接推理、靜態檢索、自適應工具和結構化表型-基因-疾病推理,這些都基於共享的知識層。訓練流程結合了Top-10後訓練與RareDx-KGPO,我們的基於知識圖的政策優化方法。其獎勵將預測投射到一個典範疾病圖中,並整合了策劃的分級相關性、本體接近性、生物醫學相似性和表型一致性。詞彙和輸出預算限制防止密集部分信用獎勵虛構或過長的差異。在八個基準測試中,完整的RareDx系統以Qwen3.5-9B為中心,達到38.34的宏觀Hit@10,比GPT-5.5在存檔協議下高出1.60分;一個不重疊的驗證選擇審計在保留案例中保持了比Direct高出6.80分的路由增益。27B系統在Hit@1/5/10上分別達到23.53/36.56/40.76。受控消融實驗顯示檢索並不總是有幫助,受控路由對增益至關重要。這些結果表明,結構化的醫學知識可以將一個緊湊的模型轉變為在臨床實踐中跨異質長尾環境的競爭性診斷排名器。

CLIMB: A Clinical Multimorbidity Benchmark for Diagnosing Co-occurring Conditions through Multiturn Conversations

2609.35462v1 by Yusuf Kesmen, Aniruddha Mukherjee, Yena Chang, David Sasu, Trevor Brokowski, Alexandra V. Kulinkina, Kristina Keitel, Akhil Arora, Lars Henning Klein, Mary-Anne Hartley

Patients often have several co-occurring clinical conditions, and the findings needed to identify and disambiguate them emerge over the course of a consultation. Evaluating clinical reasoning in this setting requires both multi-turn interaction and multi-label diagnosis. We introduce CLIMB, a benchmark in which a doctor model interviews a simulated patient to recover a ground truth set of co-occurring clinical conditions. Cases are synthesized from clinical decision algorithms and diagnostic datasets, grounding multimorbid presentations in structured clinical knowledge. Across six frontier and open models, none recovers the exact set of conditions in more than 10% of interactive cases. Diagnostic performance declines when conditions co-occur, even when models receive the full clinical record and the true number of conditions. Interaction reduces performance further. In controlled experiments, models behave like single-hypothesis trackers: they anchor on the diagnosis suggested by the opening findings, keep questioning around it, and recover a second condition mainly when a finding in view points to it. Questioning them further does not complete the set but adds mostly wrong diagnoses. We formalise this pattern with a theoretical reference model of single-hypothesis tracking. The benchmark, generator, and evaluation code are available at https://anonymous.4open.science/r/CLIMB-8340.

摘要:患者經常有多種共病臨床狀況,而識別和釐清這些狀況所需的發現會在諮詢過程中出現。在這種情境下評估臨床推理需要多輪互動和多標籤診斷。我們介紹了CLIMB,一個基準,其中醫生模型對模擬患者進行訪談,以恢復一組共病臨床狀況的真實基準。案例是從臨床決策算法和診斷數據集中合成的,將多重共病表現根植於結構化的臨床知識中。在六個前沿和開放模型中,沒有一個能在超過10%的互動案例中恢復出確切的狀況集。當狀況共存時,診斷表現會下降,即使模型獲得了完整的臨床記錄和真實的狀況數量。互動進一步降低了表現。在受控實驗中,模型的行為類似於單假設追蹤器:它們依賴於開頭發現所建議的診斷,圍繞此進行持續提問,並主要在某個發現指向它時恢復第二個狀況。進一步詢問並未完成該集合,而是主要增加了錯誤的診斷。我們用單假設追蹤的理論參考模型來形式化這一模式。基準、生成器和評估代碼可在 https://anonymous.4open.science/r/CLIMB-8340 獲得。

A decision-support system applied to Law: Reasoning and explainability of the decision

2609.35370v1 by Jeremy Bouche-Pillon, Pascale Zarat{é}, Yannick Chevalier, Nathalie Aussenac-Gilles

The emergence of the digital transition brought an increasing need to control the processing of digital information, including in Law Enforcement Agencies (LEAs). At the EU level, in recent years, many regulations have emerged to control data processing and exchange. Texts other than the GDPR, such as the ''Law Enforcement Directive (LED)'', appeared to regulate specifically how Law Enforcement Agencies (LEAs) could process data. A formal representation of these regulations can be part of decision systems that support LEAs in processing data in compliance with the regulations. Although many new formalisms have emerged to represent legal norms and rules, few are provided with a reasoning mechanism. Furthermore, systems used in decision-making processes in critical contexts such as medical diagnoses or legal decisions cannot be fully automated, and the explainability of their results is essential to ensure user confidence in decisions. This explainability aspect, while crucial, is lacking in most modern approaches that rely on machine learning. This paper describes a framework to operate formal rules from regulations, by focusing on explainability of the decision. After describing the general architecture of the proposed decision support framework, the paper showcases how symbolic AI and the SPARQL query language can support legal reasoning. It then describes an algorithm to generate a justification for the reasoning results, and outlines the procedure to be followed when the reasoning does not lead to a satisfactory conclusion. We notably focus on a method based on decision trees to determine what additional information to request from the user.

摘要:數位轉型的出現帶來了對數位資訊處理的日益需求,包括在執法機構(LEAs)中。在歐盟層面上,近年來出現了許多規範來控制數據處理和交換。除了GDPR之外,還出現了如“執法指令(LED)”等文本,專門規範執法機構(LEAs)如何處理數據。這些規範的正式表述可以成為支持執法機構在遵守規範的情況下處理數據的決策系統的一部分。儘管許多新的形式主義已經出現以表達法律規範和規則,但很少有配備推理機制的形式。此外,用於醫療診斷或法律決策等關鍵情境的決策過程中使用的系統不能完全自動化,其結果的可解釋性對於確保用戶對決策的信心至關重要。這一可解釋性方面雖然至關重要,但在大多數依賴機器學習的現代方法中卻缺乏。本文描述了一個運作規範形式規則的框架,重點在於決策的可解釋性。在描述所提議的決策支持框架的一般架構後,本文展示了符號人工智慧和SPARQL查詢語言如何支持法律推理。接著描述了一種生成推理結果的理由的算法,並概述了當推理未能導致令人滿意的結論時應遵循的程序。我們特別關注一種基於決策樹的方法,以確定需要向用戶請求的額外信息。

Training-Free Clinical Reasoning through Medical Ontologies and Cognitive Mapping: A Symbolic-Probabilistic Knowledge Graph Framework

2609.35298v1 by Surajit Das

Most clinical prediction systems learn patient-variable-outcome associations; we investigate a training-free diagnostic paradigm mapping patient observations to explicit medical knowledge. CKG Reasoner integrates candidate-specific Evidence Feature Nodes, patient-reference matching, a bounded Information Gate, knowledge-weighted evidence accumulation, disease similarity, and decisive clinical rules. Missing-aware normalization and coverage auditing distinguish absent from unavailable evidence. Candidate ranking is separate from outcome-label-independent K-means clustering, which uses four derived evidence coordinates (evidence strength, relative magnitude, directional similarity, and evidence completeness), not raw predictors or targets, to derive cohort-level assignments. Across six retrospective cohorts - four dengue (N = 1000, 1523, 989, 1018), malaria (N = 2190), and influenza (N = 4569) - a uniform, label-free, cohort-fitted K = 2 protocol yielded positive-class F1 scores of 0.996, 0.634, 0.936, 0.917, 0.695, and 0.842, and all-record accuracies of 0.996, 0.558, 0.914, 0.893, 0.707, and 0.906, respectively, with full partition-decision coverage using the frozen package and disease-specific knowledge representations. Neither scoring nor clustering uses outcome labels. Logistic regression provides a supervised baseline. Influenza incorporates confirmatory molecular PCR and is not independent pre-test prediction. Results characterize knowledge-grounded evidence separation, auditability, and sensitivity, not prospective clinical validity or comparative superiority. FOL/LLM-based clinical explanation remains unevaluated.

摘要:大多數臨床預測系統學習患者變數與結果之間的關聯;我們探討一種無需訓練的診斷範式,將患者觀察映射到明確的醫學知識。CKG Reasoner整合了候選特定的證據特徵節點、患者參考匹配、一個有界的信息閘、知識加權的證據累積、疾病相似性和決策臨床規則。缺失感知正規化和覆蓋審核將缺失證據與不可用證據區分開來。候選排名與結果標籤獨立的K均值聚類分開進行,該聚類使用四個衍生的證據坐標(證據強度、相對大小、方向相似性和證據完整性),而不是原始預測因子或目標,來導出隊列級別的分配。在六個回顧性隊列中——四個登革熱(N = 1000, 1523, 989, 1018)、瘧疾(N = 2190)和流感(N = 4569)——一個統一的無標籤、適合隊列的K = 2協議產生了正類F1分數分別為0.996、0.634、0.936、0.917、0.695和0.842,以及所有記錄的準確率分別為0.996、0.558、0.914、0.893、0.707和0.906,並且使用冷凍包和特定疾病知識表示達成了完全的分區決策覆蓋。無論是評分還是聚類都不使用結果標籤。邏輯回歸提供了一個監督的基準。流感結合了確認性分子PCR,並且不是獨立的預測測試。結果特徵化了以知識為基礎的證據分離、可審核性和敏感性,而不是前瞻性臨床有效性或比較優越性。基於FOL/LLM的臨床解釋仍未被評估。

CarveMix-RC: Addressing Rare-Class Imbalance Through Lesion-Aware Synthetic Augmentation for Brain Metastasis Segmentation

2609.35195v1 by Md Shibly Sadique, Md Fayaz Bin Hossen, Michael L. Evans, Walia Farzana, Asfaqur Rahman, Ahmed Temtam, Khan M. Iftekharuddin

Accurate segmentation of post-treatment brain metastases is essential for treatment planning, longitudinal disease monitoring, and quantitative assessment of therapeutic response. The BraTS-MET 2026 Task 1 challenge introduces a clinically relevant segmentation problem involving four anatomically distinct tumor subregions: non-enhancing tumor core (NETC), surrounding non-enhancing FLAIR hyperintensity (SNFH), enhancing tumor (ET), and the resection cavity (RC). Among these, RC segmentation is particularly challenging because of its low prevalence, heterogeneous postoperative appearance, and lesion-wise evaluation protocol, leading conventional segmentation networks to prioritize dominant tumor classes during optimization. The proposed nnU-Net-based framework explicitly addresses RC segmentation through four complementary components: (i) RC-weighted Dice and Cross-Entropy optimization to alleviate class imbalance, (ii) anatomically consistent cavity augmentation to increase the diversity of postoperative cavity appearances, (iii) a residual encoder architecture for enhanced multi-scale feature learning, and (iv) lesion-aware morphological post-processing to suppress false-positive cavity predictions while preserving anatomically plausible structures. The framework is evaluated on the BraTS-MET 2026 Task 1 online validation benchmark. Among the evaluated configurations, the ensemble model (Residual Encoder nnU-Net + nnU-Net + RC-aware CarveMix) achieves the best performance, with lesion-wise Dice scores of 0.732, 0.752, 0.708, and 0.575 and corresponding NSD scores of 0.794, 0.798, 0.727, and 0.474 for ET, TC, WT, and RC, respectively. These experimental results show that integrating RC-aware optimization, anatomically consistent augmentation, and lesion-aware post-processing provides an effective strategy for improving rare resection cavity segmentation in post-treatment brain metastases.

摘要:準確的術後腦轉移瘤分割對於治療計劃、長期疾病監測和治療反應的定量評估至關重要。BraTS-MET 2026 任務 1 挑戰引入了一個臨床相關的分割問題,涉及四個解剖上不同的腫瘤子區域:非增強腫瘤核心 (NETC)、周圍非增強 FLAIR 高信號 (SNFH)、增強腫瘤 (ET) 和切除腔 (RC)。在這些區域中,RC 分割特別具有挑戰性,因為其低發生率、異質的術後外觀以及病灶評估協議,導致傳統的分割網絡在優化過程中優先考慮主要腫瘤類別。所提出的基於 nnU-Net 的框架通過四個互補組件明確解決 RC 分割問題:(i) RC 加權的 Dice 和交叉熵優化以緩解類別不平衡,(ii) 解剖一致的腔體增強以增加術後腔體外觀的多樣性,(iii) 用於增強多尺度特徵學習的殘差編碼器架構,以及 (iv) 針對病灶的形態學後處理以抑制假陽性腔體預測,同時保留解剖上合理的結構。該框架在 BraTS-MET 2026 任務 1 在線驗證基準上進行評估。在評估的配置中,集成模型(殘差編碼器 nnU-Net + nnU-Net + RC-aware CarveMix)達到最佳性能,病灶-wise Dice 分數分別為 0.732、0.752、0.708 和 0.575,對應的 NSD 分數為 0.794、0.798、0.727 和 0.474,分別針對 ET、TC、WT 和 RC。這些實驗結果顯示,整合 RC-aware 優化、解剖一致的增強和病灶-aware 後處理提供了一種有效的策略,以改善術後腦轉移瘤中稀有切除腔的分割。

DoAtlas-2: A Foundation for Self-Evolving Causal Biomedical Discovery

2609.35107v1 by Yulong Li, Rong Xia, Yuxuan Zhang, Jianxu Chen, Xiwei Liu, Haochen Xue, Maosheng Li, Yuhang Liu, Yibo Yuan, Yutong Xie, Chong Li, Jionglong Su, Hagai Rossman, Eran Segal, Imran Razzak

We introduce DoAtlas-2, a foundation for self-evolving causal biomedical discovery that organizes knowledge around causal mechanisms and advances through external evidence from human populations. DoAtlas-2 integrates 771 research resources covering more than 720,000 participants in 48 countries, from longitudinal clinical phenotypes, medical imaging, and continuous physiological signals to eight molecular layers, together with an evidence network of approximately 4.7 million literature-derived records over 93,566 concepts and 149,383 candidate causal relations. DoAtlas-2 autonomously formulates research questions from evidence gaps and unresolved mechanisms, prespecifies their causal designs, and generates validated analyses. Supporting, challenging, and unresolved results continuously revise mechanistic interpretations, the causal evidence state, and the discovery frontier, so that DoAtlas-2 self-evolves within a closed loop of hypothesis generation, empirical testing, and renewed discovery. DoAtlas-2 has systematically evaluated 2,031 research questions. In the Human Phenotype Project (HPP), it formulated 4,014 candidate pathway questions across vascular, early-glycemic, and hepatic-metabolic systems, and screening of the first 1,079 yielded statistical support for 756. Representative studies identify blood pressure as a convergence node linking adiposity, hepatic, and lipid phenotypes to vascular outcomes, and show that an adiposity-inflammation-blood-pressure pathway is largely attenuated by joint adjustment for body mass index (BMI) and smoking. The discovered vascular network constitutes a completely interpretable predictive foundation, admitting exact attribution of every prediction and closed-form mediation effects. DoAtlas-2 thereby unifies causal mechanism discovery, population-evidence testing, and interpretable prediction within one continuously evolving foundation.

摘要:我們介紹 DoAtlas-2,這是一個自我演化的因果生物醫學發現基礎,圍繞因果機制組織知識,並通過來自人類群體的外部證據推進。DoAtlas-2 整合了 771 個研究資源,涵蓋來自 48 個國家的超過 720,000 名參與者,從縱向臨床表型、醫學影像和連續生理信號到八個分子層面,以及約 4.7 百萬個文獻衍生記錄的證據網絡,涉及 93,566 個概念和 149,383 個候選因果關係。DoAtlas-2 自主地從證據空白和未解決的機制中制定研究問題,預先指定其因果設計,並生成經過驗證的分析。支持、挑戰和未解決的結果不斷修訂機制解釋、因果證據狀態和發現前沿,使 DoAtlas-2 在假設生成、實證測試和新發現的封閉循環中自我演化。DoAtlas-2 系統性地評估了 2,031 個研究問題。在人類表型計畫 (HPP) 中,它制定了 4,014 個候選途徑問題,涵蓋血管、早期糖尿病和肝臟代謝系統,對首批 1,079 個問題的篩選產生了 756 個的統計支持。代表性研究確定血壓為一個匯聚節點,將肥胖、肝臟和脂質表型與血管結果聯繫起來,並顯示肥胖-炎症-血壓途徑在對體重指數 (BMI) 和吸煙進行聯合調整後大幅減弱。所發現的血管網絡構成了一個完全可解釋的預測基礎,允許對每個預測的精確歸因和封閉形式的中介效應。因此,DoAtlas-2 將因果機制發現、群體證據測試和可解釋預測統一於一個不斷演變的基礎之中。

VD-DeepStack: Bridging Visual Comparison and Language Reasoning for Few-Shot Anomaly Detection

2609.34949v1 by Mengyang Zhao, Zhuolin He, Haiyang Yu, Yuxuan Liang, Yifang Xu, Yuchuan Wu, Xiaolei Chen, Zhengtao Yao, Fan Shi, Yang Liu, Bin Li, Xiangyang Xue

Few-shot visual anomaly detection is fundamentally a visual comparison task, requiring fine-grained inspection of a query against normal references. Many recent methods based on large vision-language models (LVLMs) emphasize comparative reasoning through language chain-of-thought. Yet discrete, abstract descriptions may underrepresent dense, fine-grained visual differences, leaving a gap between visual comparison and its expression in language. To address this gap, we propose Visual Difference DeepStack (VD-DeepStack), which explicitly conditions language reasoning on query-reference visual differences. Specifically, we fuse DINO features with the LVLM visual hierarchy to strengthen fine-grained representations, then construct dense difference evidence from residuals between query features and softly matched reference features. The difference-evidence path injects spatially weighted difference vectors into query-image states at multiple decoder depths, while an auxiliary visual-context path provides fine-grained appearance information to support their interpretation. Experiments on 4 industrial and 2 medical anomaly benchmarks demonstrate substantial improvements in few-shot anomaly detection over baselines relying on textual comparative reasoning. These results support mitigating the visual comparison-reasoning gap through the joint design of comparison representations and their integration into the decoder. Code will be released upon acceptance.

摘要:少量樣本的視覺異常檢測基本上是一項視覺比較任務,需要對查詢與正常參考進行細緻的檢查。許多基於大型視覺-語言模型(LVLMs)的最新方法強調通過語言思維鏈進行比較推理。然而,離散的抽象描述可能無法充分表達密集的、細緻的視覺差異,從而在視覺比較與其語言表達之間留下了鴻溝。為了解決這一鴻溝,我們提出了視覺差異深層堆疊(VD-DeepStack),該方法明確地將語言推理條件化於查詢-參考的視覺差異。具體而言,我們將DINO特徵與LVLM視覺層級融合,以加強細緻的表示,然後從查詢特徵與柔性匹配的參考特徵之間的殘差構建密集的差異證據。差異證據路徑在多個解碼器深度將空間加權的差異向量注入查詢圖像狀態,而輔助視覺上下文路徑則提供細緻的外觀信息以支持其解釋。在4個工業和2個醫療異常基準上的實驗顯示,與依賴文本比較推理的基線相比,少量樣本異常檢測有了顯著的改進。這些結果支持通過比較表示的聯合設計及其在解碼器中的整合來減少視覺比較-推理之間的鴻溝。代碼將在接受後發布。

Almost Human, Except When It Matters: VoxParity and the Decisions a Voice Should Change

2609.35922v1 by Bhavik Mangla

A voice agent can handle almost every call on the words alone and still fail the few its sector's rules were written for. Emergency-call standards, fraud guidance, radio phraseology and vulnerability rules recognise that how a caller sounds, or what else is audible, can change the right action. VoxParity tests whether agents act on it. In 183 scenarios from 14 sectors, one transcript stays fixed while the audio changes (a coaching voice, a medical monitor beeping, a mayday under a radio check, noise over a drug name, a child's voice placing a bet, a frightened whisper), and with it the correct typed tool call. A words-only null test credits a system only if hearing the call moves its actions more than it moves a pipeline that only reads the words. Only 11 of the 23 systems that can also be run on the transcript pass. Descriptively, errors run toward the words: when the audio calls for protection, all 28 systems carry out the routine request more often than they over-react on clean calls (41% against 12% pooled; the words-only pipeline, 58% against 15%). Exploratory analyses place most of the leading systems' misses on cues they heard; systems beat the null almost entirely on items that state the rule; the leading systems overrule heard resignation or confusion far more often than acute alarm; and, in the models tested, describing the voice and stating the rule each recover part of the shortfall, leaving a gap on emotion.

摘要:一個語音代理可以僅依賴語言處理幾乎所有的通話,但仍然會在為其行業規則所寫的少數情況下失敗。緊急呼叫標準、詐騙指導、無線電術語和脆弱性規則認識到,來電者的聲音或其他可聽到的內容可能會改變正確的行動。VoxParity 測試代理是否會根據這些因素採取行動。在來自 14 個行業的 183 種情境中,一個文字記錄保持不變,而音頻卻在變化(教練的聲音、醫療監測器的嗶嗶聲、無線電檢查中的求救信號、藥品名稱上的噪音、一個孩子下注的聲音、一個驚恐的低語),隨之而來的是正確的鍵入工具呼叫。僅依賴文字的空白測試僅在聽到通話使其行動比僅閱讀文字的管道更有影響時,才會給系統加分。在可以基於文字記錄運行的 23 個系統中,只有 11 個通過了測試。描述性地說,錯誤傾向於文字:當音頻要求保護時,所有 28 個系統執行例行請求的頻率高於在清晰通話中過度反應的頻率(41% 對 12% 的總和;僅依賴文字的管道為 58% 對 15%)。探索性分析將大多數領先系統的失誤歸因於它們聽到的提示;系統幾乎完全在陳述規則的項目上超越了空白;領先系統在聽到的放棄或困惑上遠比在急性警報上更常推翻;而在測試的模型中,描述聲音和陳述規則各自彌補了部分短缺,但在情感上仍然存在差距。

Nociception as a Control Primitive: Afferent Channels and Nociceptive Memory for Agents Deployed in One Body

2609.34840v1 by Wolfgang Maass

An agent deployed in a single body cannot learn how fast that body wears, because every trial that would reveal its wear resistance wears the body it would protect. We study this \emph{epoch-one} setting, in which the parameters of a fixed-weight policy are set before the body is drawn and never updated in life. The agent carries a load-gated nociceptive channel and a memory that retains what was felt. We prove that felt cost moves the allocation to the best-\emph{paid} work not yet felt rather than the gentlest, that an agent without retention never sees the felt-cost constraint bind, and that the channel pays only where the threat is individually unpredictable, cheap to avoid and expensive to ignore. We measure per body, setting the agent with channel and memory against the same individual without them, where neither carries a schedule learned across lives. On $2{,}000$ simulated floor-layer knees, with wear anchored to published loss rates, feeling, retaining and substituting extends the working life from age $55.2$ to $59.6$ and raises career output from $33.7$ to $36.1$. $69.3\%$ of bodies gain and \textbf{none lose}. A body that feels but retains nothing past the day gains one of the $+4.4$ years, and retention carries the rest. A population-trained agent gains $+0.65$ years from the same channel at $-0.54$ output. The difference is what a species prior already supplies, and a single body has none. The two are related by an identity, the ablation mean reporting $(1-χ)$ of the per-body value with $χ$ the share a blind schedule already captures, so we report both. Where the regime map predicts value, a care robot sextuples its certified service life and a field-anchored fleet writes off $0.15$ of its machines instead of $0.55$. Where it predicts none, a rover gains little over blind caution, so the map holds in both directions.

摘要:一個部署在單一身體中的代理無法學習該身體的磨損速度,因為每一次試驗都會揭示其耐磨性,卻同時磨損了它所保護的身體。我們研究這種\emph{epoch-one}設置,在這種設置中,固定權重策略的參數在身體被繪製之前就已設定,並且在生命中從未更新。代理攜帶一個負載閘控的痛覺通道和一個保留所感知的記憶。我們證明了感知的成本將資源分配到尚未感知的最佳\emph{報酬}工作,而不是最溫和的工作;一個沒有記憶的代理從未見過感知成本約束的束縛;而且該通道僅在威脅是個別不可預測、避免成本低且忽視成本高的情況下才會支付。我們對每個身體進行測量,將具有通道和記憶的代理與同一個體進行比較,而兩者都沒有攜帶跨生命學習的時間表。在$2{,}000$個模擬的地板層膝蓋上,磨損基於已發表的損失率,感知、保留和替代將工作壽命從$55.2$歲延長至$59.6$歲,並將職業產出從$33.7$提高到$36.1$。$69.3\%$的身體獲得收益,\textbf{沒有一個損失}。一個感知但在當天之後不保留任何東西的身體獲得了$+4.4$年的增益,而保留則承擔了其餘的部分。一個經過人群訓練的代理從相同的通道中獲得$+0.65$年的增益,產出為$-0.54$。這一差異是物種先前已經提供的,而單一身體則沒有。這兩者通過一個身份相關聯,消融均值報告每個身體價值的$(1-χ)$,其中$χ$是盲目時間表已經捕獲的份額,因此我們報告兩者。在制度地圖預測價值的地方,一個護理機器人將其認證服務壽命增長六倍,而一個基於現場的艦隊則將$0.15$的機器報廢,而不是$0.55$。在預測為零的地方,一個探測器的增益僅略高於盲目謹慎,因此該地圖在兩個方向上都成立。

2609.34780v2 by Erik Aerts

The use and applicability of artificial intelligence (AI) in medical research and clinical practice has received increasing attention in the literature over recent years. The emergence of large language models (LLMs) has expanded discussions in regards to applications of AI within healthcare. While traditional deep learning based AI applications in medicine have often focused on specific and defined tasks, LLMs offer broader capabilities and flexibility in working with available data,. At the same time of writing, the integration of LLMs into medical settings raises important questions regarding their reliability, accuracy, transparency, safety, and appropriate role in a medical setting. This text presents and discusses recent talks and articles concerning the application of LLMs in medicine, with particular emphasis on their potential utility in research and clinical practice. It considers both the opportunities offered by these technologies and the challenges associated with their implementation, aiming to provide a perspective on the current and emerging role of LLMs within the medical field.

摘要:人工智慧(AI)在醫學研究和臨床實踐中的使用和適用性在近年來的文獻中受到越來越多的關注。大型語言模型(LLMs)的出現擴大了關於AI在醫療保健中應用的討論。雖然傳統基於深度學習的AI應用在醫學中往往專注於特定和明確的任務,但LLMs在處理可用數據方面提供了更廣泛的能力和靈活性。撰寫本文的同時,將LLMs整合進醫療環境中引發了有關其可靠性、準確性、透明度、安全性和在醫療環境中適當角色的重要問題。本文呈現並討論了有關LLMs在醫學中應用的最近演講和文章,特別強調它們在研究和臨床實踐中的潛在效用。它考慮了這些技術所提供的機會以及與其實施相關的挑戰,旨在提供對LLMs在醫療領域中當前和新興角色的看法。

ResonAct: Streaming Metrics for Runtime Diagnosis and Self-Healing in Multi-Agent Systems

2609.34701v1 by Tarun Chintada, Neelamadhav Gantayat, Ishaan Romil, Renuka Sindhgatta, Soujanya Soni, Sameep Mehta

Multi-agent systems (MAS) are increasingly used to automate enterprise workflows involving multiple specialized agents, external tools, and long-running task execution. Failures may arise from tool degradation, context propagation errors, coordination breakdowns, or repeated agent interactions that prevent task completion. While existing observability frameworks provide traces and logs, diagnosis and remediation are largely performed after execution completes, limiting opportunities for recovery during runtime. We present ResonAct, a runtime self-healing framework that enables continuous monitoring, diagnosis, and remediation of multi-agent systems through streaming operational metrics. ResonAct ingests execution traces, agent interactions, and tool invocations into a streaming analytics layer that continuously derives task progress, context health, and tool reliability metrics. These metrics serve as runtime control signals for detecting anomalous execution patterns and localizing root causes using a structured failure model. Based on the diagnosed failure, ResonAct dynamically selects remediation policies and performs actions. The framework operates as an external control plane, enabling intervention without modifying application agents or orchestration logic. We evaluate ResonAct across enterprise workflow scenarios and AppWorld benchmarks. The results show that the streaming metric-based analysis identifies execution degradations and localizes faults. Furthermore, policy-driven remediation improves task completion rates by up to 10.00 percentage points, with detection precision ranging from 70.59% to 82.91%, recall from 63.09% to 100%, recovery rates from 10.48% to 46.67%, and runtime overhead ranging from $-0.25%$ to 14.12% across the evaluated configurations.

摘要:多代理系統(MAS)越來越多地用於自動化涉及多個專門代理、外部工具和長時間運行任務執行的企業工作流程。失敗可能源於工具退化、上下文傳播錯誤、協調崩潰或重複的代理互動,這些都會阻礙任務的完成。雖然現有的可觀察性框架提供了追蹤和日誌,但診斷和修復主要是在執行完成後進行,這限制了在運行時進行恢復的機會。我們提出了ResonAct,一個運行時自我修復框架,通過流式操作指標實現多代理系統的持續監控、診斷和修復。ResonAct 將執行追蹤、代理互動和工具調用輸入到一個流式分析層,該層不斷推導任務進度、上下文健康狀況和工具可靠性指標。這些指標作為運行時控制信號,用於檢測異常執行模式並使用結構化故障模型定位根本原因。根據診斷出的故障,ResonAct 動態選擇修復政策並執行行動。該框架作為外部控制平面運行,允許在不修改應用代理或編排邏輯的情況下進行干預。我們在企業工作流程場景和AppWorld基準測試中評估了ResonAct。結果顯示,基於流式指標的分析能夠識別執行退化並定位故障。此外,基於政策的修復將任務完成率提高了最多10.00個百分點,檢測精度範圍為70.59%到82.91%,召回率範圍為63.09%到100%,恢復率範圍為10.48%到46.67%,運行時開銷範圍為$-0.25%$到14.12%,涵蓋了評估的配置。

SentZero: An Enhanced Sentence-Centric Vision-Language Pretraining for Multi-Task Zero-Shot Chest X-Ray Analysis

2609.34479v1 by Hangyul Yoon, Hyungyung Lee, Edward Choi, Eunho Yang

Vision-language (VL) pretraining using paired chest X-ray (CXR) images and radiology reports has shown strong potential for medical image understanding. However, existing methods often remain dependent on task-specific finetuning because radiology reports are lengthy, clinically dense, and difficult to align with simple zero-shot prompts. Recent sentence-level approaches partially address this limitation using clinical phrases extracted by large language models (LLMs), but they largely overlook the intrinsic characteristics of radiology discourse. In particular, limited positive-pair diversity constrains further gains, while clinically equivalent sentences frequently recur across patients, creating false negatives in contrastive learning. To address these issues, we propose SentZero, an enhanced sentence-centric VL pretraining framework for zero-shot, multi-task CXR analysis. SentZero introduces LLM-based abstract-level sentence structuring and mapping to expand positive-pair diversity, together with an additional loss term to mitigate false negatives. We further introduce sentence-conditioned residual modulation of visual embeddings, enabling visual features to adapt to the semantic characteristics of each input sentence. Across diverse downstream tasks and datasets, SentZero improves zero-shot generalization and outperforms prior multi-task zero-shot methods.

摘要:視覺-語言(VL)預訓練利用配對的胸部X光(CXR)影像和放射學報告顯示出對醫學影像理解的強大潛力。然而,現有的方法往往仍依賴於特定任務的微調,因為放射學報告冗長、臨床密集,並且難以與簡單的零樣本提示對齊。最近的句子級方法部分解決了這一限制,使用大型語言模型(LLMs)提取的臨床短語,但它們在很大程度上忽略了放射學話語的內在特徵。特別是,有限的正配對多樣性限制了進一步的增益,而臨床等效句子在不同患者之間經常重複,造成對比學習中的假陰性。為了解決這些問題,我們提出了SentZero,一個增強的以句子為中心的VL預訓練框架,用於零樣本的多任務CXR分析。SentZero引入基於LLM的抽象級句子結構和映射,以擴大正配對的多樣性,並增加一個額外的損失項以減輕假陰性。我們進一步引入句子條件的視覺嵌入殘差調制,使視覺特徵能夠適應每個輸入句子的語義特徵。在多樣的下游任務和數據集上,SentZero改善了零樣本泛化並超越了之前的多任務零樣本方法。

VL-AcneSeg: A Vision-Language Framework for Region-Aware Acne Lesion Segmentation

2609.34472v1 by Sukju Oh, Soo Ick Cho, Dae Hun Suh, Sukkyu Sun

Acne assessment is crucial for clinical decision-making, yet traditional grading and counting are subjective and fail to account for lesion size. While area-based assessment has emerged as a promising alternative, acne segmentation has continued to rely on general-purpose architectures. To address this gap, we propose VL-AcneSeg, a multimodal framework for acne lesion segmentation that leverages CLIP and region-level text prompts to incorporate spatial priors, enabling lesions to be localized across the whole face. Because region-level prompts indicate which facial areas contain lesions, we report a single global prompt, which requires no such information, as our primary setting. On our internal clinical dataset, VL-AcneSeg achieves a Dice score of 0.5082 and an IoU of 0.3407 under this protocol, the highest among all compared methods, including recent vision-language segmentation methods that are themselves given region-level prompts; region-level prompting raises these to 0.5296 and 0.3602. Moreover, lesion area measurements derived from our segmentation correlate with IGA scores at a level comparable to expert annotations (Pearson r = 0.719 versus 0.658). Notably, our framework maintains consistent performance across external validation datasets, performing reliably even on uncontrolled smartphone images without requiring additional training or fine-tuning. By pairing a protocol that requires no lesion-location information with area-based severity estimation, this work provides a foundation for objective acne assessment outside the clinic. Our implementation is publicly available at: https://github.com/sukjuoh/VL-AcneSeg

摘要:痤瘡評估對於臨床決策至關重要,但傳統的分級和計數方法主觀性強,且未能考慮病變大小。雖然基於面積的評估已成為一種有前景的替代方案,但痤瘡分割仍然依賴於通用架構。為了填補這一空白,我們提出了 VL-AcneSeg,一個多模態框架,用於痤瘡病變分割,利用 CLIP 和區域級文本提示來融入空間先驗,使病變能夠在整個面部進行定位。由於區域級提示指示哪些面部區域包含病變,我們報告了一個單一的全局提示,作為我們的主要設置,這不需要任何此類信息。在我們的內部臨床數據集中,VL-AcneSeg 在這一協議下達到了 0.5082 的 Dice 分數和 0.3407 的 IoU,這是所有比較方法中最高的,包括最近的視覺-語言分割方法,這些方法本身也提供了區域級提示;區域級提示將這些指標提高到 0.5296 和 0.3602。此外,從我們的分割中得出的病變面積測量與 IGA 分數的相關性達到與專家註釋相當的水平(Pearson r = 0.719 對比 0.658)。值得注意的是,我們的框架在外部驗證數據集上保持一致的性能,即使在無法控制的智能手機圖像上也能可靠地執行,而無需額外的訓練或微調。通過將不需要病變定位信息的協議與基於面積的嚴重程度評估相結合,這項工作為臨床外的客觀痤瘡評估提供了基礎。我們的實現已公開可用於:https://github.com/sukjuoh/VL-AcneSeg

Evolving Support Priorities in Empathetic Reinforcement Learning

2609.34249v1 by Pengyu Huang, Zhiyuan Han, Wenwen Tong, Hewei Guo, Jiangnan Chen, Sirui Chen, Lewei Lu, Beier Zhu, Xun Yang

We identify a fundamental mismatch in empathetic reinforcement learning: support priorities evolve with the dialogue state, yet existing methods typically optimize predefined reward specifications that remain fixed across turns. To model these evolving support priorities, we organize empathetic support along cognitive, affective, and proactive empathy, and propose Context-Adaptive Rubric Evolution (CARE). At each turn, CARE generates a context-adaptive rubric by adjusting both the weights of these three empathy dimensions and their fine-grained evaluation criteria. The rubric generator is trained with turn-level rubric supervision and human preference data through supervised fine-tuning followed by preference-based reinforcement learning, and then serves as an adaptive reward interface for online empathetic RL. Integrated with both RLVER and MICA, CARE achieves state-of-the-art performance across SentientBench, EQBench3, and EMPA under three independent LLM judges. Notably, on EMPA, CARE improves EPM-Idx over the strongest baseline by at least 13 points under all three judges, including an increase from 28.11 to 83.54 under Gemini-2.5-Pro. Further analyses show that learned rubric priorities systematically vary across dialogue stages and user emotions, demonstrating that CARE adapts what is rewarded as support needs evolve.

摘要:我們發現同理心強化學習中存在一個根本的不匹配:支持優先級隨著對話狀態而演變,但現有方法通常優化預定的獎勵規範,這些規範在各回合中保持固定。為了建模這些不斷演變的支持優先級,我們沿著認知、情感和主動同理心組織同理心支持,並提出了上下文自適應評分標準演變(CARE)。在每個回合中,CARE 通過調整這三個同理心維度的權重及其細緻的評估標準來生成一個上下文自適應的評分標準。評分標準生成器通過回合級評分標準監督和人類偏好數據進行監督微調,然後通過基於偏好的強化學習進行訓練,並作為在線同理心強化學習的自適應獎勵介面。與 RLVER 和 MICA 結合,CARE 在 SentientBench、EQBench3 和 EMPA 上實現了最先進的性能,並在三位獨立的 LLM 評審中表現出色。值得注意的是,在 EMPA 上,CARE 在所有三位評審下將 EPM-Idx 提高了至少 13 分,包括在 Gemini-2.5-Pro 下從 28.11 增加到 83.54。進一步的分析顯示,學習到的評分標準優先級在對話階段和用戶情緒之間系統性地變化,顯示出 CARE 會隨著支持需求的演變而調整獎勵內容。

Unknown is not normal: separating language-model extraction from rule-based decision logic for clinical risk scores

2609.34112v1 by Nicolás Vera Zúñiga

Large language models (LLMs) are increasingly used to compute clinical risk scores from free-text notes. Notes are often incomplete, and treating undocumented findings as normal can silently misclassify patients. We test whether separating three-state extraction (present, absent or unknown, by an LLM) from decision logic (deterministic code computing score bounds over unknown inputs) lets a system ask only questions that can change the decision. On 1,200 synthetic emergency cases across six calculators (HEART, CURB-65, qSOFA, PERC, Wells, Cockcroft-Gault), with a simulated clinician answering questions, we compared this bounds policy with asking for every missing input, a missing-equals-normal schema, and an end-to-end LLM agent (Claude Opus 5.5). With Claude Haiku 4.5 as extractor, the bounds policy matched ask-all accuracy (99.4% vs 99.4%) with half the questions (0.92 vs 1.78 per case) and no irrelevant ones. Treating missing as normal dropped accuracy to 91.2% and under-triaged 8.5% of patients (95% CI 7.1-10.2), and under-triage persisted under messy notes and a noisy clinician. The agent was equally accurate under ideal conditions (99.6%) but 9.5% of its questions were irrelevant; with a noisy clinician it was less accurate than the bounds policy (83.5% vs 87.0%, p<0.001) and committed prematurely in 2.7% of cases (bounds: 0%). A 9B local model as extractor reached oracle-level accuracy (99.8%). In 584 real case reports from MedCalc-Bench, only 52% contained enough information to determine the category (HEART 13%). Routing decisions through code that reasons explicitly about unknowns avoids premature commitment and irrelevant questions, halves the questions asked, and works with small local models.

摘要:大型語言模型(LLMs)越來越多地用於從自由文本筆記中計算臨床風險分數。筆記通常不完整,將未記錄的發現視為正常可能會悄悄地錯誤分類患者。我們測試了將三狀態提取(由LLM判斷的存在、缺失或未知)與決策邏輯(在未知輸入上計算分數邊界的確定性代碼)分開,是否能讓系統僅提出能改變決策的問題。在1200個合成緊急案例中,涵蓋六個計算器(HEART、CURB-65、qSOFA、PERC、Wells、Cockcroft-Gault),並由模擬臨床醫生回答問題,我們將這個邊界政策與要求每個缺失輸入的方式、缺失等於正常的模式以及一個端到端的LLM代理(Claude Opus 5.5)進行比較。使用Claude Haiku 4.5作為提取器,邊界政策在問題數量減半的情況下(每個案例0.92對1.78)達到了與要求所有問題相同的準確率(99.4%對99.4%),且沒有不相關的問題。將缺失視為正常使準確率降至91.2%,並使8.5%的患者被低估分級(95% CI 7.1-10.2),而且在雜亂的筆記和噪音臨床醫生的情況下,低估分級仍然存在。在理想條件下,代理的準確率同樣為99.6%,但其9.5%的問題是無關的;在噪音臨床醫生的情況下,其準確率低於邊界政策(83.5%對87.0%,p<0.001),並在2.7%的案例中過早做出承諾(邊界:0%)。一個9B本地模型作為提取器達到了神諭級的準確率(99.8%)。在來自MedCalc-Bench的584個真實案例報告中,只有52%包含足夠的信息來確定類別(HEART 13%)。通過對未知進行明確推理的代碼進行路由決策,可以避免過早承諾和不相關的問題,將提問數量減半,並且能與小型本地模型一起工作。

The Devil is in the Spectrum Bias: Spectrum-Balanced Feature Matching for Robust Representation Distillation

2609.34106v1 by Kuniaki Saito, Yoshitaka Ushiku

Large visual foundation models have demonstrated remarkable transferability across a wide range of downstream tasks. To deploy such models efficiently, feature matching has become a popular knowledge distillation approach that transfers teacher representations to smaller student models without requiring labeled data. However, we show that the conventional feature matching objective with L2-distance is inherently biased toward reconstructing dominant spectral directions of the teacher representation, while under-optimizing low-variance directions that often contain task-relevant information. To address this, we propose Spectrum-Balanced Feature Matching, SpecMatch, a simple objective that adaptively emphasizes under-optimized spectral directions while preserving the relative importance of dominant directions. SpecMatch is easy to implement and introduces negligible computational overhead. Extensive experiments on image recognition demonstrate that SpecMatch consistently improves downstream adaptation across diverse tasks, including image classification, anomaly detection, medical image analysis, and domain generalization. In particular, SpecMatch outperforms conventional feature matching in 40 of 42 teacher--student and training-setting combinations, while consistently improving over the original student model in all settings. We further demonstrate that the proposed objective generalizes beyond vision, improving downstream performance across six protein understanding tasks.

摘要:大型視覺基礎模型在各種下游任務中顯示出卓越的可轉移性。為了有效部署這些模型,特徵匹配已成為一種流行的知識蒸餾方法,該方法將教師表示轉移到較小的學生模型,而無需標記數據。然而,我們顯示,傳統的L2距離特徵匹配目標本質上偏向於重建教師表示的主導光譜方向,同時對通常包含任務相關信息的低方差方向進行優化不足。為了解決這個問題,我們提出了光譜平衡特徵匹配(Spectrum-Balanced Feature Matching,SpecMatch),這是一個簡單的目標,能夠自適應地強調未優化的光譜方向,同時保留主導方向的相對重要性。SpecMatch易於實現,並引入了微不足道的計算開銷。在圖像識別方面的廣泛實驗表明,SpecMatch在包括圖像分類、異常檢測、醫學影像分析和領域泛化等多種任務中,持續改善下游適應性。特別是,SpecMatch在42個教師-學生和訓練設置組合中的40個中超越了傳統的特徵匹配,同時在所有設置中持續改善了原始學生模型。我們進一步證明,所提出的目標在視覺之外也具有泛化能力,改善了六個蛋白質理解任務的下游性能。

TRACE: Expert-Aligned ECG Representation Learning with Rigorous Benchmarking and Real-World Validation in Acute Cardiac Care

2609.34088v1 by Lovely Yeswanth Panchumarthi, Andrew Lu, Saurabh Kataria, Delgersuren Bold, Minxiao Wang, Runze Yan, Patricia Dykes, Brian J. Gow, Tom J. Pollard, Jessica K. Zègre-Hemsey, Dillon J. Dzikowicz, Lekshmi Kumar, Xiao Hu, Ran Xiao

TRACE (Text-Reinforced Analysis of Cardio ECGs) is a multimodal electrocardiogram (ECG) representation model that learns clinically grounded signal embeddings for downstream cardiac classification. It is designed to address the limitations of existing CLIP-style training, which often struggles with noisy clinical text and fails to leverage the complementary strengths of unimodal (from ECG) and cross-modal (between ECG and matched cardiologist reports) learning. To bridge this gap, we propose a hybrid architecture that jointly learns unimodal and cross-modal representations via uncertainty-weighted multi-task learning while utilizing an LLM-based pipeline to extract high-fidelity findings from cardiologist reports. We evaluate TRACE across a spectrum of clinical urgency, establishing robust performance on public benchmarks for arrhythmia classification and structural abnormalities relative to existing unimodal and multimodal ECG models. To demonstrate real-world utility, we further validate the model on acute coronary occlusion (ACO), where the prevailing ST-elevation criteria miss 25-34% of true occlusions. Utilizing a large private ACO dataset with expert-annotated ground truth, TRACE significantly outperforms real-world clinical practice, yielding a 19.0% increase in sensitivity or a 62.6% reduction in false positive rates at the clinical baseline. This extensive evaluation confirms that TRACE delivers both strong performance on benchmark tasks and tangible clinical impact in the most acute, high-risk cardiac scenarios.

摘要:TRACE(文本強化心電圖分析)是一種多模態心電圖(ECG)表示模型,旨在學習臨床基礎的信號嵌入,以便用於下游心臟分類。它旨在解決現有CLIP風格訓練的局限性,這種訓練通常在嘈雜的臨床文本中掙扎,並未能利用單模態(來自ECG)和跨模態(在ECG與匹配的心臟病醫生報告之間)學習的互補優勢。為了填補這一空白,我們提出了一種混合架構,通過不確定性加權的多任務學習共同學習單模態和跨模態表示,同時利用基於LLM的管道從心臟病醫生報告中提取高保真結果。我們在臨床緊急程度的範疇內評估TRACE,並在公開基準上建立了對心律失常分類和結構異常的穩健性能,相較於現有的單模態和多模態ECG模型。為了展示其在現實世界中的實用性,我們進一步在急性冠狀動脈阻塞(ACO)上驗證該模型,當前的ST抬高標準錯過了25-34%的真實阻塞。利用一個大型私有ACO數據集,並由專家標註的真實情況,TRACE顯著超越了現實世界的臨床實踐,在臨床基線下實現了19.0%的靈敏度提升或62.6%的假陽性率降低。這一廣泛的評估確認TRACE在基準任務上提供了強大的性能,並在最急迫、高風險的心臟情況下帶來了實質性的臨床影響。

Who Gets a Token, and What Does It Carry? Unequal Name Support and Concept Access in Large Language Models

2609.34065v1 by Mir Tafseer Nayeem, Davood Rafiei

Names are personal identifiers, but they also carry social meaning and are widely used to evaluate how language models treat different people. Such evaluations typically assume that matched names are comparable model inputs. We show that this assumption often fails at the lexical interface: matched names are not necessarily matched inputs. Some names receive direct single-token access, while others are assembled from multiple subwords, creating unequal name-surface support. Across nearly half a million first names and 12 LLM-associated tokenizers, direct lexical access is highly selective, model dependent, and uneven across race- and gender-associated name metadata. We introduce NameTrace, a model-native, fine-grained, pre-behavioral framework for measuring whether unequal name-surface support remains a vocabulary property or becomes visible in task-relevant internal representations. NameTrace measures concept accessibility from the model's own probabilities over task-specific adjective axes with continuous task-aligned weights. On matched atomic and short-fragmented names within the same race/ethnicity--gender-associated strata, support predicts systematic differences in concept accessibility across fellowship, hiring, clinical assessment, and lending. These differences persist across all eight matched strata, extend across model families, and transfer to unseen names. Hidden-state interventions further show that the measured task directions have downstream leverage, shifting later constrained choices. Unequal lexical support is therefore demographically structured at the input and remains visible in task-relevant model computation. NameTrace makes lexical comparability measurable, supporting a broader principle: behavioral comparability begins with lexical comparability.

摘要:名字是個人識別符號,但它們也承載著社會意義,並廣泛用於評估語言模型如何對待不同的人。這些評估通常假設匹配的名字是可比較的模型輸入。我們顯示這一假設在詞彙介面上經常失效:匹配的名字不一定是匹配的輸入。有些名字獲得直接的單詞訪問,而另一些則由多個子詞組成,造成不平等的名字表面支持。在近五十萬個名字和12個與LLM相關的分詞器中,直接詞彙訪問高度選擇性,依賴於模型,並且在與種族和性別相關的名字元數據中不均勻。我們引入了NameTrace,這是一個模型原生的、細粒度的、前行為框架,用於測量不平等的名字表面支持是否仍然是一種詞彙特性,或在任務相關的內部表示中變得可見。NameTrace測量從模型自身的概率中對任務特定形容詞軸的概念可及性,並使用連續的任務對齊權重。在同一種族/民族—性別相關的層級內,對匹配的原子和短片段名字的支持預測在獎學金、招聘、臨床評估和貸款方面的概念可及性系統性差異。這些差異在所有八個匹配層級中持續存在,跨模型家族擴展,並轉移到未見過的名字。隱藏狀態的干預進一步顯示,測量的任務方向具有下游影響,改變後續的受限選擇。因此,不平等的詞彙支持在輸入層面上是人口結構化的,並在任務相關的模型計算中保持可見。NameTrace使詞彙可比性可測量,支持一個更廣泛的原則:行為可比性始於詞彙可比性。

Large Language Models for Structured Clinical Data Analysis: Dual-Agent Grounding and Validation

2609.34039v1 by Erfan D. Dehkalani, Seetha Shankaran, Abbot R. Laptook, C. Michael Cotten, P. Ellen Grant, Yangming Ou

Objective: To develop and characterize CLEAR-Med, a dual-agent framework for natural-language analysis of structured clinical data that separates SQL-based invocation from independent validation. Methods: CLEAR-Med uses one agent to translate a question into executable Structured Query Language (SQL), retain the executed query and database result, and produce a draft. Deterministic checks and a separately invoked cross-provider Validation Agent then accept the draft, request one bounded repair, or abstain. We formalized the system as a bounded selective pipeline and evaluated CLEAR-Med's configuration and scalability, and the Invocation Agent's accuracy and consistency on a 25-query development benchmark, using a harmonized 21-site neonatal hypoxic-ischemic encephalopathy table containing 532 de-identified infant records and approximately 1,300 variables. Results: CLEAR-Med completed all six nominal scalability configurations, including 500x1300. Across 25 development-benchmark queries repeated five times, the Invocation Agent answered 83 of 125 responses correctly (66.4%; query-cluster bootstrap 95% CI, 48.0-83.2%), compared with 15 of 125 (12.0%; 95% CI, 3.2-22.4%) for the ungrounded ChatGPT baseline, a paired improvement of 54.4 percentage points (95% CI, 36.8-72.0%). Conclusion: CLEAR-Med provides a general architecture for traceable analysis of structured clinical data: numerical claims remain linked to executed SQL, and unresolved cases can fail closed. The reported experiments characterize CLEAR-Med's configuration and scalability and the Invocation Agent's accuracy, while the formal analysis establishes the encoded-property guarantee of the complete control flow; a prospective full-pipeline evaluation of the validation and abstention stages is the next stage of this work.

摘要:目標:開發並表徵CLEAR-Med,這是一個雙代理框架,用於自然語言分析結構化臨床數據,將基於SQL的調用與獨立驗證分開。方法:CLEAR-Med使用一個代理將問題翻譯為可執行的結構化查詢語言(SQL),保留執行的查詢和數據庫結果,並生成草稿。確定性檢查和單獨調用的跨提供者驗證代理然後接受草稿,請求一次有限的修正,或選擇不進行修正。我們將系統形式化為一個有限的選擇性管道,並評估CLEAR-Med的配置和可擴展性,以及調用代理在25個查詢開發基準上的準確性和一致性,使用一個和諧的21個站點的新生兒缺氧缺血性腦病表,該表包含532個去標識的嬰兒記錄和約1300個變量。結果:CLEAR-Med完成了所有六個名義上的可擴展性配置,包括500x1300。在25個開發基準查詢中重複五次,調用代理正確回答了125個回應中的83個(66.4%;查詢集群自助法95%置信區間,48.0-83.2%),而未經驗證的ChatGPT基準僅正確回答了125個中的15個(12.0%;95%置信區間,3.2-22.4%),這是一個配對改善54.4個百分點(95%置信區間,36.8-72.0%)。結論:CLEAR-Med提供了一個可追溯的結構化臨床數據分析的通用架構:數值聲明仍然與執行的SQL相連,未解決的情況可以失敗關閉。報告的實驗表徵了CLEAR-Med的配置和可擴展性以及調用代理的準確性,而正式分析確立了完整控制流的編碼屬性保證;對驗證和放棄階段的前瞻性全管道評估是這項工作的下一階段。

Jev in Medicine: A Benchmark Evaluation

2609.34024v2 by Alfredo Madrid-García, Beatriz Merino-Barbancho

Jev is a non-generative "System One" model that assigns probabilities to predefined answer options and cannot answer outside them. Its accuracy and calibration on medical question-answering and case-based diagnostic-reasoning tasks are unknown. We evaluated Jev 1.13 on four medical benchmarks: MetaMedQA, PubMedQA, DiagnosisArena-MCQ and the NEJM Case Challenges. GPT-6 Sol, with (medium) and without reasoning, was the reference. The primary outcome was top-1 accuracy; key secondary outcomes were calibration, selective prediction and recognition of unanswerable questions. All 8,469 requests returned a valid answer. Jev's accuracy was similar to that of GPT-6 Sol with medium reasoning on PubMedQA (78.4% vs 78.2%;), lower on MetaMedQA (74.8% vs 82.7%) and much lower on DiagnosisArena-MCQ (59.8% vs 82.4%;) and the NEJM cases (61.8% vs 82.4%). On MetaMedQA, Jev's probabilities were the best calibrated (expected calibration error 0.063 vs 0.146), and its answers with a probability of at least 0.9 (52.9% of questions) were 93.4% accurate, but GPT-6 Sol was as accurate when it accepted a similar proportion of questions. On DiagnosisArena-MCQ, Jev's probabilities discriminated poorly (AUROC 0.645 vs 0.768). Of the 162 questions whose correct answer was "I don't know or cannot answer", Jev chose that option for 10.5% (GPT-6 Sol, 8.6%). Median latency was 0.27-0.31 s; all 2,823 items cost USD 0.08. Jev was fast and inexpensive, and its accuracy was similar to that of a frontier LLM on research abstracts but lower on examination questions and much lower on complex diagnostic cases. Task-specific validation is required before clinical use.

摘要:Jev是一個非生成的「系統一」模型,為預定的答案選項分配概率,並且無法在這些選項之外回答。它在醫學問答和基於案例的診斷推理任務上的準確性和校準情況尚不清楚。我們在四個醫學基準上評估了Jev 1.13:MetaMedQA、PubMedQA、DiagnosisArena-MCQ和NEJM案例挑戰。GPT-6 Sol,無論是(中等)還是沒有推理,都是參考標準。主要結果是前1準確率;關鍵的次要結果包括校準、選擇性預測和對無法回答問題的識別。所有8,469個請求都返回了有效的答案。Jev的準確性與GPT-6 Sol在PubMedQA上的中等推理相似(78.4%對78.2%;),在MetaMedQA上較低(74.8%對82.7%),在DiagnosisArena-MCQ上則低得多(59.8%對82.4%;)以及在NEJM案例上(61.8%對82.4%)。在MetaMedQA上,Jev的概率校準最佳(預期校準誤差0.063對0.146),其概率至少為0.9的答案(52.9%的問題)準確率為93.4%,但當GPT-6 Sol接受相似比例的問題時,其準確率也相當。在DiagnosisArena-MCQ上,Jev的概率區分能力較差(AUROC 0.645對0.768)。在162個正確答案為「我不知道或無法回答」的問題中,Jev選擇該選項的比例為10.5%(GPT-6 Sol為8.6%)。中位延遲為0.27-0.31秒;所有2,823個項目的成本為0.08美元。Jev速度快且成本低,其準確性與前沿LLM在研究摘要上的表現相似,但在考試問題上的準確性較低,對於複雜的診斷案例則低得多。在臨床使用之前,需要進行特定任務的驗證。

EHRAdapt: Adapting Pretrained Language Models to Electronic Health Records with Semantic Priors for Rare Clinical Events

2609.34007v1 by Andre R Goncalves, Vincent Liu, Priyadip Ray

Electronic health records (EHRs) encode clinical histories as (time, modality, code) tuples, whereas pretrained language models expect text tokens. Serializing them as text inflates sequence length and redundantly encodes structure. We introduce EHRAdapt, an adapter that maps tuples directly into a frozen language model's embedding space. Modality receives a learned embedding, time gaps enter through learned attention biases, and event codes receive dedicated vectors. Learning event vectors is the central challenge: clinical vocabularies are long-tailed, leaving rare events too few observations for reliable estimates. EHRAdapt therefore represents each event vector as the sum of a semantic prior and an evidence residual. The prior is a frozen embedding of the event's clinical description from a biomedical language model trained on clinical ontologies, mapped into the model's input space by a shared learned projection, so it supplies clinical meaning even when observations are scarce. The residual, a learned low-rank event-specific correction, refines it as evidence accumulates. We run continued pretraining on about 4 million patients' records with three frozen LLM backbones (OLMo2 1B, Llama3.2 1B, and OLMo2 7B), training only the adapter (0.1--0.6% of all parameters). The full adapter outperforms all ablations in held-out next-event prediction on every backbone. Removing the semantic pathway hurts rare events over ten times more than the most frequent ones, whereas removing the residual hurts overall prediction but improves it for the rarest events. On reportable infectious-disease and syndromic downstream classification tasks, EHRAdapt outperforms text-based LLM and count-based baselines, and both pathways improve rare-disease discrimination. The two pathways therefore play complementary roles, visible only when results are broken down by event frequency rather than averaged.

摘要:電子健康紀錄(EHRs)將臨床歷史編碼為(時間、模態、代碼)元組,而預訓練的語言模型則期望文本標記。將它們序列化為文本會增加序列長度並冗餘地編碼結構。我們引入了EHRAdapt,一個將元組直接映射到凍結語言模型嵌入空間的適配器。模態獲得一個學習的嵌入,時間間隙通過學習的注意力偏差進入,而事件代碼則獲得專用向量。學習事件向量是中心挑戰:臨床詞彙是長尾的,稀有事件的觀察次數太少,無法進行可靠的估計。因此,EHRAdapt將每個事件向量表示為語義先驗和證據殘差的總和。先驗是來自於在臨床本體上訓練的生物醫學語言模型的事件臨床描述的凍結嵌入,通過共享的學習投影映射到模型的輸入空間,因此即使觀察稀少也能提供臨床意義。殘差是一個學習的低秩事件特定修正,隨著證據的累積進行精細化。我們對約400萬名患者的記錄進行持續的預訓練,使用三個凍結的LLM骨幹(OLMo2 1B、Llama3.2 1B和OLMo2 7B),僅訓練適配器(佔所有參數的0.1--0.6%)。完整的適配器在每個骨幹的保留下一個事件預測中超越了所有的消融實驗。移除語義通路對稀有事件的影響超過最常見事件的十倍,而移除殘差則對整體預測造成損害,但對最稀有事件的預測有所改善。在可報告的傳染病和綜合徵下游分類任務中,EHRAdapt的表現超過基於文本的LLM和基於計數的基準,並且兩個通路都改善了稀有疾病的區分。因此,這兩個通路扮演互補的角色,只有當結果按事件頻率細分而非平均時才會顯現出來。

Is your uncertainty map wrong, or is its target? Exact diagnostics for the Tweedie diagonal, and a gradient-free alternative

2609.33786v1 by Vicent Ribas, Anna Oliveras Tous

A diffusion model can predict a follow-up medical scan from a baseline, but a clinician needs a per-voxel map of where that prediction can be trusted. Many such maps approximate the diagonal of the Tweedie posterior covariance, and are evaluated against another approximation of it, so whether the estimator or the target limits them is unclear. We compute the exact diagonal on six checkpoints across fourteen model-corpus conditions. Hutchinson at M=200 tracks it at rank agreement of at least 0.92 everywhere, yet in four of the fourteen the exact diagonal is anti-correlated with the denoising error, reaching -0.13, so a faithful estimator reproduces that reversal. All four are real-image conditions; on the models' own samples the reversal does not appear, so evaluating on generated samples flatters this family. What limits these maps is the target, not the estimator. We then introduce Tweedie Probe-Tangent (T-PT), a gradient-free residual probe that corrupts one model-supported prediction repeatedly and measures the voxel-wise variance of the denoiser's response. T-PT reads a different functional of the same Jacobian, and its exact second-order form ranks with the diagonal wherever the diagonal reverses; at thirty probes it returns a map too unstable to reproduce that ranking, while Hutchinson at M=5 already reproduces it, so T-PT there is not evidence against the reversal. We offer it as an instrument, not a better approximation. On brain MRI at full resolution, where every Jacobian-based estimator we test runs out of memory, T-PT leads a twenty-chain Monte-Carlo ensemble on five of eight endpoints inside tissue and trails it on none, at 16x fewer network evaluations; over the whole volume the ensemble leads, and fifty chains close the tissue gap. On lung CT the ensemble is ahead throughout. Both lose most of their discrimination where the change is, which remains open.

摘要:擴散模型可以從基線預測後續的醫學掃描,但臨床醫生需要一個每體素的地圖來確定該預測的可信度。許多這樣的地圖近似於Tweedie後驗協方差的對角線,並且是根據另一個近似進行評估,因此估計器或目標是否限制了它們尚不清楚。我們在十四個模型-語料條件下的六個檢查點計算了精確的對角線。Hutchinson在M=200的情況下,無論何處的排名一致性至少為0.92,但在十四個條件中的四個中,精確的對角線與去噪誤差呈反相關,達到-0.13,因此一個忠實的估計器會重現這一反轉。這四個都是實際影像條件;在模型自身的樣本中,這一反轉並不存在,因此在生成樣本上的評估使這一系列模型看起來更好。限制這些地圖的是目標,而不是估計器。
然後我們引入Tweedie Probe-Tangent (T-PT),這是一種無梯度的殘差探測器,反覆破壞一個模型支持的預測並測量去噪器響應的體素級變異性。T-PT讀取相同雅可比矩陣的不同函數,其精確的二階形式在對角線反轉的地方排名;在三十個探測點下,它返回的地圖不穩定到無法重現該排名,而Hutchinson在M=5的情況下已經重現了它,因此在這裡T-PT並不是反轉的證據。我們將其作為工具,而不是更好的近似。在全分辨率的腦部MRI中,我們測試的每個基於雅可比的估計器都耗盡了內存,T-PT在八個端點中的五個內部組織上引領了一個二十鏈的蒙特卡羅集成,並且在任何端點上都未落後,網絡評估數量少了16倍;在整個體積上,集成領先,而五十條鏈縮小了組織的差距。在肺部CT中,集成始終領先。兩者在變化發生的地方失去了大部分的區分能力,這一點仍然是開放的。

BIRD: Distilling Decision Boundaries into Rationales for MLLM Adaptation

2609.33713v1 by Anglin Liu, Yanlin Wu, Ruichao Chen, Yuting Zhang, Qingyuan Zeng, Pengxiang Cai, Ziqi Gong, Muchen Li, Jintai Chen

Adapting general-purpose multimodal large language models (MLLMs) to specialized domains requires learning domain-specific decision criteria, which often hinge on subtle visual distinctions between otherwise plausible answers. Rationale augmentation aims to expose such evidence through additional observations or inter-sample comparisons, yet a visually valid cue is not necessarily decision-relevant: it may describe how samples differ without changing the model's relative preference between competing answers. We therefore introduce BIRD, a self-improving Boundary-Informed Rationale Distillation framework that uses model-specific confusions to locate unresolved local decision boundaries and distills the evidence that resolves these confusions into rationales. For each sample, BIRD retrieves candidate neighbors from the target MLLM's own representation space and selects the most confusable one according to its answer preferences. It then generates answer-blind candidate evidence from their visual differences and functionally verifies which evidence most effectively strengthens the model's preference for the correct answer while avoiding inappropriate transfer across the pair. The verified evidence is then distilled into a single-sample rationale for standard supervised fine-tuning. Experiments on medical and chart VQA show that BIRD outperforms competing rationale-augmentation methods across two target MLLMs, while further analyses demonstrate clearer separation of confusable answers and stronger gains from model-matched supervision.

摘要:適應通用多模態大型語言模型(MLLMs)到專門領域需要學習特定於領域的決策標準,這通常依賴於在其他可行答案之間的微妙視覺區別。推理增強旨在通過額外的觀察或樣本間比較來揭示這種證據,然而,視覺上有效的線索不一定與決策相關:它可能描述樣本之間的差異,而不改變模型對競爭答案的相對偏好。因此,我們引入了BIRD,一個自我改善的邊界知情推理蒸餾框架,利用模型特定的混淆來定位未解決的局部決策邊界,並將解決這些混淆的證據提煉成推理。對於每個樣本,BIRD從目標MLLM自身的表示空間中檢索候選鄰居,並根據其答案偏好選擇最具混淆性的那一個。然後,它根據它們的視覺差異生成不依賴答案的候選證據,並功能性地驗證哪種證據最有效地增強模型對正確答案的偏好,同時避免在這對之間的不當轉移。經過驗證的證據然後被提煉成單樣本推理,用於標準的監督微調。在醫療和圖表VQA的實驗中,BIRD在兩個目標MLLM上超越了競爭的推理增強方法,而進一步的分析顯示出混淆答案之間更清晰的區分和來自模型匹配監督的更強增益。

PPG-LM: A Photoplethysmography-Language Model with Multi-Level Clinical Alignment

2609.33516v1 by Xiaoda Wang, Minxiao Wang, Maxwell A Xu, Patrick Langer, Kaiqiao Han, Defu Cao, Xiao Luo, Yuzhe Yang, Yan Liu, Xiao Hu, Yizhou Sun, Wei Wang, Carl Yang

Photoplethysmography (PPG) is widely recorded by clinical monitors and consumer wearables, providing a scalable source of continuous physiological information. These recordings offer an opportunity for physiological assessment at scale, but realizing this potential requires models to learn from both signal-derived physiological supervision and broader clinical context captured in electronic health records (EHRs). This involves aligning information spanning local observations, care events, and entire visits with PPG representations at corresponding temporal scales. However, existing PPG foundation models primarily rely on task-specific prediction heads, while the medical knowledge of large language models does not necessarily translate into waveform understanding. To bridge this gap, we introduce PPG-LM, the first PPG-language model family to learn physiological representations from both signal-derived supervision and broader clinical context captured in EHRs. To construct clinically grounded captions, we develop an automatic captioning pipeline that generates segment-, event-, and visit-level descriptions from signal measurements and structured EHR records. We then learn from these pairs through a two-stage framework that first establishes segment-language correspondence through contrastive learning and waveform-conditioned captioning, then extends alignment to events and visits through time-aware aggregation and temporal statement matching. Pretrained on approximately 73k hours of PPG, PPG-LM supports language-based recognition, cross-modal retrieval, and segment captioning. Experiments on MC-MED, MIMIC-III, and VitalDB show improved retrieval and caption factuality over language-model baselines and gains over PPG and time-series foundation models on multiple clinical prediction tasks.

摘要:光電容積描記法(PPG)被臨床監測器和消費者可穿戴設備廣泛記錄,提供了一個可擴展的持續生理信息來源。這些記錄為大規模的生理評估提供了機會,但實現這一潛力需要模型從信號衍生的生理監督和電子健康記錄(EHRs)中捕獲的更廣泛臨床背景中學習。這涉及到將涵蓋本地觀察、護理事件和整個就診的資訊與相應時間尺度上的PPG表示對齊。然而,現有的PPG基礎模型主要依賴於特定任務的預測頭,而大型語言模型的醫學知識並不一定能轉化為波形理解。為了彌補這一差距,我們介紹了PPG-LM,這是第一個從信號衍生的監督和EHRs中捕獲的更廣泛臨床背景中學習生理表示的PPG語言模型系列。為了構建臨床基礎的標題,我們開發了一個自動標題生成管道,從信號測量和結構化的EHR記錄中生成段落、事件和就診級別的描述。然後,我們通過一個兩階段框架從這些對中學習,首先通過對比學習和波形條件的標題建立段落-語言對應,然後通過時間感知聚合和時間語句匹配將對齊擴展到事件和就診。PPG-LM在約73,000小時的PPG上進行了預訓練,支持基於語言的識別、跨模態檢索和段落標題生成。在MC-MED、MIMIC-III和VitalDB上的實驗顯示,在語言模型基準上提高了檢索和標題的事實性,並在多個臨床預測任務上相對於PPG和時間序列基礎模型取得了增益。

Federated Multi-Modal Human Activity Recognition using Multi-Agent Reinforcement Learning

2609.33492v1 by Debasmita Dey, Tanmay Sen, Himel Mallick

Human Activity Recognition (HAR) from heterogeneous wearable sensors is fundamental to the Internet of Health Things (IoHT), supporting rehabilitation, elderly care, and smart healthcare. Existing multimodal fusion methods often assign fixed equal weights to sensor streams, overlooking differences in modality importance, acquisition cost, and sensor quality, which can vary due to movement, incorrect placement, or temporary blockage. We propose an adaptive and cost-aware multimodal HAR framework based on multi-agent reinforcement learning for centralized HAR and extend it to federated learning as FedMHAR. In the centralized setting, multimodal fusion is formulated as a cooperative Multi-Agent Reinforcement Learning (MARL) problem, where each sensing modality is assigned a PPO-based agent that learns per-sample fusion weights, enabling the model to emphasize informative modalities while down-weighting costly sensors when cheaper alternatives provide sufficient information. In the federated setting, we introduce BiFL-PPO, a bidirectional federated optimization strategy in which a server-side PPO policy learns client-specific trust weights and feeds them back to adapt local learning rates and proximal regularization. Unlike round-level optimization, BiFL-PPO uses dense batch-level rewards for more frequent feedback and stable training under heterogeneous client data. Evaluation on the MEx Rehabilitation and UTD Multimodal Human Action datasets shows that the centralized framework achieves 87.30% and 94.98% accuracy, respectively, outperforming conventional fusion methods and state-of-the-art HAR models. FedMHAR achieves 79.74% and 77.49% in the federated setting, consistently surpassing FedAvg, FedProx, FedBN, FedNova, and AdaFedProx, while providing more stable performance and reducing sensor acquisition cost.

摘要:人類活動識別(HAR)來自異質可穿戴傳感器,對健康物聯網(IoHT)至關重要,支持康復、老年護理和智慧醫療。現有的多模態融合方法通常對傳感器流分配固定的相等權重,忽略了模態重要性、獲取成本和傳感器質量的差異,這些差異可能因運動、不正確的放置或暫時的阻塞而變化。我們提出了一種基於多智能體強化學習的自適應和成本感知多模態HAR框架,用於集中式HAR,並將其擴展到聯邦學習,稱為FedMHAR。在集中式設置中,多模態融合被表述為一個合作的多智能體強化學習(MARL)問題,其中每個感知模態被分配一個基於PPO的智能體,該智能體學習每個樣本的融合權重,使模型能夠強調信息豐富的模態,同時在更便宜的替代方案提供足夠信息時降低成本傳感器的權重。在聯邦設置中,我們引入了BiFL-PPO,一種雙向聯邦優化策略,其中伺服器端的PPO策略學習客戶特定的信任權重並反饋以調整本地學習率和近端正則化。與回合級優化不同,BiFL-PPO使用密集的批次級獎勵以獲得更頻繁的反饋,並在異質客戶數據下實現穩定的訓練。在MEx康復和UTD多模態人類行動數據集上的評估顯示,集中式框架分別達到87.30%和94.98%的準確率,超越了傳統的融合方法和最先進的HAR模型。FedMHAR在聯邦設置中達到79.74%和77.49%的準確率,始終超越FedAvg、FedProx、FedBN、FedNova和AdaFedProx,同時提供更穩定的性能並降低傳感器獲取成本。

A Cheap Verifier is Good Enough: LLM Post-training is Robust to Erroneous Rewards

2609.33467v1 by Andreas Plesner, Curtis Northcutt, Francisco Guzmán, Anish Athalye

When post-training large language models on tasks with semi-verifiable rewards, there are many factors (training steps, base model size, training order, data quality, verifier accuracy, etc.) that practitioners must contend with to maximize model performance. Yet, it remains unclear how well verifier agreement predicts post-training performance on such tasks. In this paper, we explore this question with over 11k H100 GPU-hours, across HealthBench and PRBench tasks in medical, legal, and finance domains. Across the tested domains, Qwen3 trainees (1.7B-8B on HealthBench; 8B on PRBench), evaluation splits, and frontier LLM reference judges (which we call golden verifiers), higher verifier agreement does not consistently identify the best training verifier. Expensive verifiers need not outperform inexpensive ones, and open-weight Gemma verifiers produce strong training outcomes. We compare two low-cost choices retrospectively -- a cost-reducing choice and a balanced choice -- with estimated grading cost reductions of 98.8%-99.7% relative to the golden grading protocols and average post-training score gaps of 1-3 points from the best evaluated training verifier. These averages include larger losses in individual settings; they do not establish that verifier choices are interchangeable.

摘要:在對具有半可驗證獎勵的任務進行後訓練大型語言模型時,從業者必須面對許多因素(訓練步驟、基礎模型大小、訓練順序、數據質量、驗證者準確性等),以最大化模型性能。然而,目前尚不清楚驗證者的一致性在多大程度上預測此類任務的後訓練性能。在本文中,我們通過超過11,000小時的H100 GPU,探索這個問題,涵蓋醫療、法律和金融領域的HealthBench和PRBench任務。在測試的領域中,Qwen3訓練者(在HealthBench上為1.7B-8B;在PRBench上為8B)、評估拆分和前沿LLM參考評審(我們稱之為黃金驗證者)中,更高的驗證者一致性並不總是能夠準確識別最佳訓練驗證者。昂貴的驗證者不一定要優於便宜的驗證者,而開放權重的Gemma驗證者則產生了強大的訓練結果。我們回顧性比較了兩個低成本選擇——一個是降低成本的選擇,另一個是平衡的選擇——相對於黃金評分協議,估計的評分成本降低為98.8%-99.7%,而從最佳評估的訓練驗證者的平均後訓練分數差距為1-3分。這些平均數包括在個別設置中的更大損失;它們並未確立驗證者選擇是可互換的。

MAC-Net: A Multi-Task Deep Learning Framework for Modeling Cognitive Function From Task-Based fMRI

2609.33440v1 by Md. Tanvir Rahman, Nabil Anan Orka, Asaduzzaman Khan, Mohammad Ali Moni

Objective cognitive assessment from neural signals supports neurorehabilitation, but individual-level prediction from task-based fMRI (tfMRI) remains difficult because neural features coexist with substantial demographic and scanner-related variation. We present the Multi-task Activation and Contrast Network (MAC-Net), a covariate-aware deep learning framework for modeling individual cognitive function from regional tfMRI. By isolating tfMRI features into a dedicated neural pathway and restricting participant variables to a terminal late-fusion pathway, MAC-Net prevents dominant covariates from suppressing high-dimensional clinical representations during feature learning. Evaluating baseline data from 6,500 Adolescent Brain Cognitive Development Study participants under family-aware cross-validation, MAC-Net was benchmarked against linear models, random forests, and alternative deep architectures. The N-back plus Monetary Incentive Delay configuration achieved $R^{2}$ values of 0.174, 0.238, and 0.277 for fluid, crystallized, and total cognition, outperforming covariate-only baselines (0.178) and alternative deep models (0.217). N-back was the most informative paradigm, whereas incorporating the Stop Signal Task marginally degraded performance. Feature attributions via Integrated Gradients, DeepLIFT, and Input Gradient were highly concordant, localizing working-memory-related frontal, parietal, and cingulate regions. These findings demonstrate that covariate-aware multi-task modeling yields reproducible cognitive-function estimations, establishing a robust neural engineering framework for clinical translation.

摘要:目標認知評估來自神經信號,支持神經康復,但基於任務的功能性磁共振成像(tfMRI)在個體層面的預測仍然困難,因為神經特徵與顯著的人口統計和掃描儀相關變異共存。我們提出了多任務激活和對比網絡(MAC-Net),這是一個考慮協變量的深度學習框架,用於從區域tfMRI建模個體認知功能。通過將tfMRI特徵隔離到專門的神經通路並將參與者變量限制在終端的晚融合通路,MAC-Net防止主導協變量在特徵學習過程中壓制高維臨床表示。對6,500名青少年大腦認知發展研究參與者的基線數據進行家庭意識交叉驗證,MAC-Net與線性模型、隨機森林和其他深度架構進行了基準測試。N-back加上金錢激勵延遲配置在流動性、結晶性和總認知方面達到了$R^{2}$值分別為0.174、0.238和0.277,超越了僅考慮協變量的基準(0.178)和其他深度模型(0.217)。N-back是最具信息性的範式,而納入停止信號任務則輕微降低了性能。通過整合梯度、DeepLIFT和輸入梯度的特徵歸因高度一致,定位到與工作記憶相關的額葉、頂葉和扣帶區域。這些發現表明,考慮協變量的多任務建模產生可重複的認知功能估計,建立了一個穩健的神經工程框架以便於臨床轉化。

Temporal Graph Learning of Wearable Actigraphy and Sleep Traces for Modelling Adolescent Crystallized Intelligence

2609.33428v1 by Md. Tanvir Rahman, Nabil Anan Orka, Asaduzzaman Khan, Mohammad Ali Moni

Wearable actigraphy offers a scalable, ecologically valid alternative to episodic clinical assessment. However, predicting continuous adolescent crystallized intelligence ($G_c$) from such traces remains challenging due to irregular device adherence and complex behavioral-environmental interactions. We address this using daily summary data derived from 21-day Fitbit records of 6,091 adolescents in the Adolescent Brain Cognitive Development Study (Release 5.1). We propose SATURN, a Sleep-Activity Temporal Unified Regression Network. It represents participants as 21-node temporal graphs encoding daily behaviors and temporal adjacency. To prevent imputation artifacts, invalid-day edges are dynamically pruned during forward passes. Node embeddings are refined via residual GATv2 layers, aggregated through masked attention pooling, and fused with sociodemographic covariates. Under family-controlled, age-sex-BMI-stratified cross-validation, SATURN achieves $R^2 = 0.2783 \pm 0.0127$, consistently improving upon flattened machine learning (Gradient Boosting, $R^2 = 0.2372$) and sequential deep learning (BiLSTM, $R^2 = 0.2688$) baselines. Explainability analyses identify light activity, metabolic equivalents, and sleep duration as dominant predictors, while Monte Carlo dropout and subgroup analyses confirm equitable performance across sociodemographic strata. Ultimately, SATURN establishes a rigorous computational framework for digital cognitive phenotyping, offering a scalable pathway to complement traditional assessments by highlighting macro-level behavioral anomalies.

摘要:可穿戴行為測量提供了一種可擴展的、生態有效的替代方案,以取代臨床評估的偶發性。然而,從這些數據中預測持續的青少年結晶智力 ($G_c$) 仍然具有挑戰性,因為設備遵從性不規則且行為與環境之間的互動複雜。我們使用來自 6,091 名青少年在青少年大腦認知發展研究(版本 5.1)中,為期 21 天的 Fitbit 記錄所衍生的每日摘要數據來解決這個問題。我們提出了 SATURN,一個睡眠-活動時間統一回歸網絡。它將參與者表示為 21 節點的時間圖,編碼每日行為和時間相鄰性。為了防止插補伪影,在前向傳播過程中動態修剪無效日邊緣。節點嵌入通過殘差 GATv2 層進行精煉,通過遮罩注意力池化進行聚合,並與社會人口學協變量融合。在家庭控制、年齡-性別-BMI 分層的交叉驗證下,SATURN 的 $R^2 = 0.2783 \pm 0.0127$,持續優於扁平化的機器學習(梯度提升,$R^2 = 0.2372$)和序列深度學習(BiLSTM,$R^2 = 0.2688$)基準。可解釋性分析確定輕度活動、代謝當量和睡眠持續時間為主要預測因子,而蒙特卡羅隨機失活和子群分析則確認了在社會人口學層次上表現公平。最終,SATURN 建立了一個嚴謹的計算框架,用於數位認知表型,提供了一條可擴展的途徑,以通過突顯宏觀層面的行為異常來補充傳統評估。

Explainable Deep Learning of Resting-State Functional Connectomes Reveals Network Biomarkers of Adolescent Intelligence

2609.33422v1 by Md. Tanvir Rahman, Nabil Anan Orka, Asaduzzaman Khan, Mohammad Ali Moni

Mapping resting-state brain organization to individual differences in cognitive ability remains a major challenge in population neuroinformatics. Although deep learning enables flexible modeling of brain connectivity, limited interpretability restricts its scientific and clinical utility. To address this objective, we developed an explainable deep learning framework based on sparse projected residual networks to predict fluid, crystallized, and total intelligence from resting-state functional magnetic resonance imaging in 5,285 participants from the Adolescent Brain Cognitive Development study. We incorporated three complementary explainability methods (Integrated Gradients, Gradient Shapley Additive Explanations, and Occlusion) to interpret model behavior. The framework outperformed existing approaches, achieving Pearson correlations of 0.44, 0.58, and 0.56 for fluid, crystallized, and total intelligence, respectively, corresponding to predictive improvements of 6 to 9 percent. All three explainability methods produced near-identical feature rankings (pairwise rank correlations greater than 0.99). Consensus maps revealed a dual-layered functional architecture where primary predictive hubs localized within canonical systems, while the strongest global predictive pathways frequently bypassed these hubs through distributed, long-range relay connections. These findings suggest that intelligence emerges from the interaction between localized computational hubs and distributed communication pathways. Ultimately, these normative network architectures provide clinical reference maps to detect individual deviations, supporting earlier diagnosis, cognitive subtype stratification, and treatment monitoring in atypical neurodevelopment.

摘要:將靜息狀態下的大腦組織映射到個體在認知能力上的差異,仍然是人口神經資訊學中的一大挑戰。雖然深度學習使得大腦連接的靈活建模成為可能,但有限的可解釋性限制了其科學和臨床的實用性。為了達成這一目標,我們開發了一個基於稀疏投影殘差網絡的可解釋深度學習框架,從5,285名來自青少年大腦認知發展研究的參與者的靜息狀態功能性磁共振成像中預測流體智力、結晶智力和總智力。我們結合了三種互補的可解釋性方法(整合梯度、梯度沙普利加法解釋和遮蔽)來解釋模型行為。該框架的表現超過了現有的方法,對流體智力、結晶智力和總智力的皮爾森相關係數分別達到0.44、0.58和0.56,對應的預測改進為6%到9%。所有三種可解釋性方法產生了幾乎相同的特徵排名(成對排名相關係數大於0.99)。共識圖揭示了一種雙層功能架構,其中主要的預測樞紐位於典型系統內,而最強的全球預測通路則經常通過分散的長距離中繼連接繞過這些樞紐。這些發現表明,智力是由局部計算樞紐和分散通信通路之間的互動所產生的。最終,這些規範性網絡架構提供了臨床參考圖,以檢測個體偏差,支持早期診斷、認知亞型分層和在非典型神經發展中的治療監測。

QuPID: Quantum Parameter-Efficient Input-Dependent Retrieval Adaptation for Medical RAG

2609.33351v1 by Hyojun Ahn, Emily Jimin Roh, Soohyun Park, Walid Saad, Hyung-Chul Lee, Joongheon Kim

Fidelity-based quantum retrieval ranks candidates by the fidelity between query and archive states. Applying a shared input-independent unitary after fixed state encoding leaves that fidelity unchanged, so training the circuit cannot alter the ranking. Quantum parameter-efficient input-dependent retrieval adaptation (QuPID) repairs this by making the circuit input-dependent through data re-uploading and by comparing measurement readouts, vectors of local Pauli expectations, rather than states. The result is a small readout for adapting frozen image features to a local archive with limited data: training simulates the circuit classically, and inference runs on a GPU with fixed learned parameters. We characterize the class as a structured factorization of input-modulated quadratic feature maps, bound the frequency support of its re-uploading channel, and give a parameter-count generalization bound that motivates its small budget. Under a shared frozen backbone and a label-free protocol, QuPID's 60 parameters give higher precision-at-5 (P@5) on ChestX-ray14 and MURA than frozen medical encoders, and than adapters and low-rank adaptation (LoRA) with up to 5.25 million trainable parameters. On ChestX-ray14, the P@5 gain over the frozen encoder is +0.116, the lead over retuned adapters is widest at 512 adaptation examples (+0.040), and the full-budget margin over an equally compact classical rotation-plane head is +0.023 with a 95% interval excluding zero. Medical imaging is the primary testbed; the pattern recurs on two non-medical benchmarks, in report generation, and under simulated gate noise and finite-shot readout.

摘要:基於保真度的量子檢索通過查詢和檔案狀態之間的保真度對候選者進行排名。應用共享的輸入無關單位後,固定狀態編碼的保真度保持不變,因此訓練電路無法改變排名。量子參數高效的輸入依賴檢索適應(QuPID)通過數據重新上傳使電路依賴於輸入,並通過比較測量讀數、局部保利期望的向量,而不是狀態來修復這一點。結果是針對有限數據的本地檔案適應凍結圖像特徵的小型讀出:訓練在經典上模擬電路,而推理在具有固定學習參數的GPU上運行。我們將該類別表徵為輸入調製的二次特徵映射的結構因式分解,限制其重新上傳通道的頻率支持,並給出一個參數計數的泛化界限,這激勵了其小預算。在共享的凍結骨幹和無標籤協議下,QuPID的60個參數在ChestX-ray14和MURA上的精確度@5(P@5)高於凍結醫療編碼器,以及高達525萬可訓練參數的適配器和低秩適應(LoRA)。在ChestX-ray14上,與凍結編碼器相比,P@5的增益為+0.116,與重新調整的適配器相比,在512個適應示例下的領先幅度最大(+0.040),而與一個同樣緊湊的經典旋轉平面頭的全預算邊際為+0.023,95%區間不包括零。醫學影像是主要的測試平台;該模式在兩個非醫學基準、報告生成以及在模擬閘噪聲和有限次讀出下重複出現。

CHI: A Composite Hallucination Index Unifying Entity, Relation, and Quantity Dimensions for Summarization Evaluation

2609.33343v1 by Praveenkumar Katwe, Rakesh Chandra Balabantaray, Kali Prasad Vittala

Faithfulness evaluation of abstractive summaries remains an open challenge, with existing metrics addressing only isolated hallucination types: factual entity errors, relational inconsistencies, or numerical fabrications, without capturing their co-occurrence or interaction. We introduce CHI (Composite Hallucination Index), the first unified hallucination metric that decomposes faithfulness errors into three orthogonal dimensions: entity hallucination (EHI), relation hallucination (RHI*), and quantity hallucination (QHI). Each dimension employs a shared softmax-normalized architecture over Venn diagram-derived factors representing extractiveness, positive hallucination, over-focus, negative hallucination, and lost focus. The novel QHI component introduces tolerance-aware numerical matching with exact, epsilon, derived, and temporal comparison modes. We fuse the three dimensions via harmonic mean to produce a single composite score that penalizes weakness in any dimension. We validate CHI on 800 source articles spanning four domains (news, medical, legal, financial) with summaries from five generation systems. Empirical results demonstrate that: (i) the three dimensions are statistically orthogonal (mean rho = 0.148), confirming they capture distinct error types; (ii) CHI achieves the highest system-level correlation with human judgments (rho = 0.66, p = 0.006) on SummEval, outperforming ROUGE (rho = 0.53), EHI (rho = 0.58), and all individual components; and (iii) ablation studies confirm that all three dimensions contribute unique variance, with the full composite outperforming any individual component while providing decomposable error diagnostics unavailable from single-score baselines. CHI provides practitioners with a decomposable, interpretable, and efficient faithfulness metric suitable for both offline evaluation and online monitoring of summarization systems.

摘要:忠實度評估抽象摘要仍然是一個未解的挑戰,現有的指標僅針對孤立的幻覺類型:事實實體錯誤、關係不一致或數字虛構,而未捕捉它們的共現或互動。我們介紹了CHI(綜合幻覺指數),這是第一個統一的幻覺指標,將忠實度錯誤分解為三個正交維度:實體幻覺(EHI)、關係幻覺(RHI*)和數量幻覺(QHI)。每個維度都利用一個共享的softmax正規化架構,基於源自維恩圖的因素,這些因素代表了提取性、正幻覺、過度關注、負幻覺和失焦。新穎的QHI組件引入了容忍度感知的數字匹配,具有精確、epsilon、衍生和時間比較模式。我們通過調和平均將這三個維度融合,產生一個單一的綜合分數,對任何維度的弱點進行懲罰。我們在涵蓋四個領域(新聞、醫療、法律、金融)的800篇來源文章上驗證了CHI,這些文章的摘要來自五個生成系統。實證結果顯示:(i)這三個維度在統計上是正交的(平均rho = 0.148),確認它們捕捉到不同的錯誤類型;(ii)CHI在SummEval上達到與人類評價的最高系統級相關性(rho = 0.66,p = 0.006),超越了ROUGE(rho = 0.53)、EHI(rho = 0.58)和所有個別組件;以及(iii)消融研究確認這三個維度都貢獻了獨特的變異性,完整的綜合指標優於任何個別組件,同時提供了從單一分數基準無法獲得的可分解錯誤診斷。CHI為實踐者提供了一個可分解、可解釋且高效的忠實度指標,適用於離線評估和在線監控摘要系統。

The Error You See Is Not the Error You Made: Progression-aware Reasoning Origin for Reasoning Error Localization

2609.33297v1 by Yiguo Wang, Ziyuan Yang, Yi Zou, Dan Lin, Rongsheng Li, Yi Zhang

Verifying multi-step LLM reasoning requires more than determining whether a trace is correct: a useful verifier should identify where the reasoning first goes wrong. However, existing holistic methods provide little positional evidence, while forward sequential verification often treats the first rejected step as the error source. Under error propagation, this assumption can fail, since an earlier mistake may remain locally plausible and become observable only through its downstream consequences. We therefore rethink reasoning verification as a progression-aware error-source localization problem: rather than asking only where a reasoning trace first appears inconsistent, we ask which earlier step best explains how that inconsistency emerges along the trajectory. Based on this view, we propose Progression-aware Reasoning Origin (PRO), a training-free framework for first-error localization. PRO jointly models incoming support from the preceding context and outgoing compatibility with subsequent reasoning, selectively refines regions where these signals disagree, and finally performs detector-conditioned source attribution with intervention-based evidence to distinguish the true error origin from its propagated manifestations. We further formalize the gap between forward rejection and structural exposure, showing why incoming-side evidence alone is insufficient for reliable localization under error propagation. Experiments across open-form, medical, and structured reasoning tasks demonstrate consistent improvements over strong verification baselines, supporting progression-aware source attribution as a more faithful formulation of reasoning verification.

摘要:驗證多步驟 LLM 推理不僅需要確定一個痕跡是否正確:一個有用的驗證器應該能夠識別推理首次出錯的地方。然而,現有的整體方法提供的位置信息有限,而前向序列驗證通常將第一個被拒絕的步驟視為錯誤來源。在錯誤傳播的情況下,這一假設可能會失效,因為早期的錯誤可能在局部上仍然是合理的,並且只有通過其下游後果才能被觀察到。因此,我們重新思考推理驗證,將其視為一個進程感知的錯誤來源定位問題:我們不僅詢問推理痕跡首次出現不一致的地方,而是詢問哪一個早期步驟最能解釋沿著軌跡出現的不一致。基於這一觀點,我們提出了進程感知推理來源(PRO),這是一個無需訓練的首錯定位框架。PRO 共同建模來自前一上下文的支持和與後續推理的兼容性,選擇性地細化這些信號不一致的區域,並最終通過基於干預的證據進行檢測器條件的來源歸因,以區分真實的錯誤來源和其傳播的表現。我們進一步形式化了前向拒絕和結構曝光之間的差距,顯示為什麼僅依賴來自進入側的證據對於在錯誤傳播下的可靠定位是不足夠的。在開放式、醫療和結構化推理任務中的實驗顯示出對強驗證基準的一致改進,支持進程感知來源歸因作為推理驗證的更真實表述。

CORTEX: A Verified Experience Layer for Generalist Agents

2609.33260v1 by Garapati Keerthana, Manik Gupta

An agent can solve a task today and face the same task under new facts, tools, or governing knowledge tomorrow. Most agent systems can retrieve relevant text or recall prior conversations, but they lack a principled way to decide when a previous solution is still valid, when it must be adapted, and when it should be discarded. We introduce CORTEX (Contextual Orchestration and Reuse of Task EXperience), a general AI systems framework that connects specialized agents through an external layer of verified experience. Each episode records its task conditions, source and tool state, decisive predicates, proof trace, verifier, and outcome. A meta-controller chooses exact replay, checked adaptation, fresh synthesis, or escalation. Accepted episodes can become task patterns and procedural strategies through a challenge-driven development loop. This gives the system an implicit competence layer that can grow without changing model weights. We formalize system contracts for exact replay and source-version separation, and derive when reuse saves computation. A controlled two-domain implementation tests the exact-replay core on 1,000 synthetic cases. Complete-family holdouts test procedural transfer on 1,000 new-family cases across eight clinical and policy splits, with complete fresh-evidence grounding and perfect invariance to irrelevant-field and insertion-order perturbations. The transfer trace exposes the work required for verified strategy execution. These results establish an initial path toward general intelligence through reusable procedures, typed experience, and developmental transfer.

摘要:一個代理可以在今天解決一個任務,並在明天面對同一任務,但有新的事實、工具或治理知識。大多數代理系統可以檢索相關文本或回憶先前的對話,但它們缺乏一種原則性的方式來決定何時先前的解決方案仍然有效,何時必須進行調整,以及何時應該被丟棄。我們介紹了 CORTEX(上下文協調與任務經驗重用),這是一個通用的 AI 系統框架,通過一層經過驗證的經驗將專門的代理連接起來。每個事件記錄其任務條件、來源和工具狀態、決定性謂詞、證明痕跡、驗證者和結果。一個元控制器選擇精確重播、檢查調整、新的綜合或升級。接受的事件可以通過挑戰驅動的開發循環轉變為任務模式和程序策略。這為系統提供了一個隱含的能力層,能夠在不改變模型權重的情況下增長。我們為精確重播和來源版本分離形式化了系統合同,並推導出何時重用可以節省計算。受控的雙域實施在 1,000 個合成案例上測試精確重播核心。完整家庭保留測試在八個臨床和政策拆分中對 1,000 個新家庭案例的程序轉移,具有完整的新證據基礎和對無關領域及插入順序擾動的完美不變性。轉移痕跡揭示了執行經過驗證的策略所需的工作。這些結果為通過可重用程序、類型化經驗和發展轉移建立了一條通向通用智能的初步路徑。

FOCUS: Benchmarking Retinal Model Generalization from Foundation Vision Encoders to Multimodal LLMs

2609.33158v1 by David Restrepo, Chenwei Wu, Luis Filipe Nakayama, Miguel L. Martins, Stergios Christodoulidis, Maria Vakalopoulou, Enzo Ferrante

Progress in AI-based retinal image analysis has advanced with foundation models, yet evaluating their reliability remains challenging. Performance reported on a single dataset does not capture how models behave under dataset shift, across clinical definitions, or for different patient subgroups. This limitation is particularly critical in medical imaging analysis, where robustness, calibration, and fairness are essential for safe deployment. We introduce FOCUS (Foundation Ophthalmic Cross-Dataset Understanding under Shift), a cross-dataset benchmark for evaluating retinal fundus models that considers vision-only encoder models (VM), vision-language dual-encoder models (VLM), and multimodal large language models (MLLM). FOCUS harmonizes binary diabetic retinopathy, referable diabetic retinopathy, and glaucomatous optic neuropathy tasks across ten public datasets spanning diverse geographies, acquisition conditions, and label protocols. The benchmark evaluates models through a unified analysis layer that measures ranking performance, calibration, subgroup disparities, and image-quality robustness. We present a large-scale evaluation covering 532 base configurations and 228 MLLM configurations adapted through supervised fine-tuning with low-rank adaptation (LoRA). Results show that no model family consistently dominates across tasks and datasets: general VM encoders achieve the strongest average ranking performance, medical MLLMs are competitive but variable, and dual encoder VLMs benefit substantially from lightweight adaptation. Fine-tuning improves in-domain performance but exhibits heterogeneous transfer to external datasets, particularly in calibration. These findings demonstrate that retinal model evaluation is inherently multidimensional. FOCUS provides a practical framework and public benchmark to assess generalization, reliability, and robustness beyond single-dataset leaderboards

摘要:進展於基於人工智慧的視網膜影像分析已隨著基礎模型的發展而提升,然而評估其可靠性仍然具有挑戰性。單一數據集上報告的性能無法捕捉模型在數據集轉移、臨床定義之間或不同患者子群體中的行為。這一限制在醫學影像分析中特別關鍵,因為穩健性、校準和公平性對於安全部署至關重要。我們引入了FOCUS(Foundation Ophthalmic Cross-Dataset Understanding under Shift),這是一個跨數據集基準,用於評估視網膜眼底模型,考慮了僅視覺編碼器模型(VM)、視覺-語言雙編碼器模型(VLM)和多模態大型語言模型(MLLM)。FOCUS在十個公共數據集上協調二元糖尿病視網膜病變、可參考糖尿病視網膜病變和青光眼性視神經病變任務,這些數據集涵蓋了多樣的地理位置、獲取條件和標籤協議。該基準通過一個統一的分析層評估模型,測量排名性能、校準、子群體差異和影像質量的穩健性。我們呈現了一個涵蓋532個基本配置和228個經過低秩適應(LoRA)監督微調的MLLM配置的大規模評估。結果顯示,沒有任何模型家族在任務和數據集上始終佔據主導地位:一般的VM編碼器實現了最強的平均排名性能,醫學MLLM在競爭中但變化不定,而雙編碼器VLM在輕量適應中受益匪淺。微調改善了內域性能,但在外部數據集上展現出異質的轉移,特別是在校準方面。這些發現表明,視網膜模型評估本質上是多維的。FOCUS提供了一個實用的框架和公共基準,以評估超越單一數據集排行榜的泛化、可靠性和穩健性。

MedRouter: Demystifying Knowledge Differences Across Medical LLMs for Routing-Based Reasoning

2609.33119v1 by Lang Cao, Binghang Lu, Yuhao Shen, Yue Guo

Medical question answering spans diverse specialties and modalities, and individual medical large language models (LLMs) exhibit distinct strengths across tasks and domains. This heterogeneity suggests that combining specialists may enable broader coverage of medical questions than relying on any single model. However, existing LLM routing methods primarily seek to balance answer quality and inference cost, leaving open how to exploit differences in specialist competence to improve medical reasoning. In this paper, we introduce MedRouter, an agentic system that uses an embedding-based multi-label router to select and query specialist LLMs, then passes their responses to a generator to produce the final answer. We further propose SCALE (Specialist Competence-Aware Learning), a two-stage training framework that first trains the Router with specialist correctness supervision and then optimizes its selections through reinforcement learning. The second stage uses a Performance Gain Reward (PGR) that measures how specialist information affects the generator's answer correctness relative to answering without that information. Experiments on eight text-based and multimodal medical QA benchmarks show that MedRouter outperforms the strongest routing baseline by 8% in average accuracy. Our analysis of specialist outputs further reveals distinct strengths and complementary question-level coverage, motivating learned routing to combine these capabilities for more comprehensive medical reasoning.

摘要:醫療問題回答涵蓋多樣的專科和模式,而個別醫療大型語言模型(LLMs)在任務和領域上展現出不同的優勢。這種異質性表明,結合專家可能比依賴任何單一模型更能廣泛覆蓋醫療問題。然而,現有的LLM路由方法主要尋求平衡答案質量和推理成本,尚未探討如何利用專家的能力差異來改善醫療推理。在本文中,我們介紹了MedRouter,一個使用基於嵌入的多標籤路由器來選擇和查詢專家LLMs的代理系統,然後將它們的回應傳遞給生成器以產生最終答案。我們進一步提出了SCALE(專家能力感知學習),這是一個兩階段的訓練框架,首先用專家的正確性監督來訓練路由器,然後通過強化學習來優化其選擇。第二階段使用性能增益獎勵(PGR),衡量專家信息如何影響生成器的答案正確性,相對於在沒有該信息的情況下回答。對八個基於文本和多模態的醫療QA基準的實驗顯示,MedRouter在平均準確率上比最強的路由基線高出8%。我們對專家輸出的分析進一步揭示了不同的優勢和互補的問題級別覆蓋,促使學習路由結合這些能力以實現更全面的醫療推理。

ECG-Scroll: A Long-Horizon, Streaming Benchmark and Agent Environment for Interpretation of Ambulatory Electrocardiograms

2609.33117v1 by Haitao Li, Chenglin Li, Zhengyao Ding, Ziyu Li, Yiheng Mao, Zhengxing Huang

Multimodal large language models (MLLMs) can now interpret a standard ten-second, twelve-lead electrocardiogram (ECG) with clinically grounded, reward-verified reasoning. Real cardiac monitoring is different. Ambulatory (Holter) and telemetry recordings span hours to days and are read as they stream in, and their clinically decisive findings are paroxysmal, brief episodes buried in an otherwise unremarkable trace. Such a recording cannot be held in one context at diagnostic resolution, and its future has not yet happened, so a reader must work online, deciding what to measure now, committing evidence to memory as it passes, and reporting events as they occur. We recast long-duration ECG interpretation as a long-horizon, online (streaming, causal) sequential decision process and introduce ECG-Scroll. As a benchmark, long ambulatory recordings are streamed to an agent chunk by chunk, and it must localize, quantify, and promptly flag paroxysmal events without access to future signal; because the underlying signal is retained, every answer is checkable against objective ground truth, giving rule-based rather than judge-based rewards, and the streaming formulation adds a metric batch evaluation cannot express, the detection latency between an event's onset and the moment the agent records it. As an agent environment, it is a fixed, gym-style interaction layer that exercises three competencies single-glance ECG models never touch: Memory, Tool use through signal-grounded measurement rather than reading pixels, and Planning of what to measure now and when to commit. We release 390 whole-recording instances spanning 2,536 hours of two-lead ambulatory ECG and evaluate a signal-threshold rule agent alongside off-the-shelf LLM agents online, characterizing how they use memory, tools, and planning and where the benchmark's head-room lies.

摘要:多模態大型語言模型(MLLMs)現在可以以臨床為基礎、經過獎勵驗證的推理來解釋標準的十秒鐘、十二導程心電圖(ECG)。真正的心臟監測則有所不同。動態(Holter)和遙測記錄的時間跨度從幾小時到幾天,並且在流入時即時閱讀,其臨床決定性發現是突發的、短暫的事件,埋藏在其他看似無異常的波形中。這樣的記錄無法在診斷解析度下保持在一個上下文中,且其未來尚未發生,因此讀者必須在線工作,決定現在要測量什麼,將證據在經過時記憶,並在事件發生時報告。我們將長時間心電圖解釋重新構建為一個長期的、在線(流式、因果)序列決策過程,並引入ECG-Scroll。作為基準,長時間的動態記錄以塊的形式流式傳輸給代理,代理必須在無法訪問未來信號的情況下定位、量化並迅速標記突發事件;因為基礎信號被保留,每個答案都可以與客觀真實進行檢查,提供基於規則而非基於評判的獎勵,而流式的表述增加了一個度量,批量評估無法表達,即事件開始與代理記錄之間的檢測延遲。作為代理環境,它是一個固定的、健身房風格的互動層,鍛煉三種單次視覺心電圖模型從未接觸的能力:記憶、通過信號基礎測量而非讀取像素的工具使用,以及現在要測量什麼和何時提交的計劃。我們釋放了390個完整記錄實例,涵蓋2536小時的雙導程動態心電圖,並在線評估一個信號閾值規則代理,與現成的LLM代理進行比較,描述它們如何使用記憶、工具和計劃,以及基準的潛力所在。

SemReward-VL: Semantic Reward-Guided Video-Language Adaptation for Developmental Behavior Assessment

2609.33082v1 by De Jiang, Shuo Zhang, Kehong Yuan, Hongen Liao

Developmental screening videos show how children perform specific behaviors, but clinical records usually contain outcomes rather than descriptions of what happened. We present SemReward-VL, which learns to describe item-specific behavior from these outcomes. A vision-language model generates a description, and a frozen language model scores its agreement with the clinical outcome, relevance to the item, abstention on unrelated video-item pairs, and clarity. Group relative policy optimization (GRPO) updates LoRA adapters using this semantic reward. On 13,379 videos covering 41 items, the method improves aggregate accuracy and the number of items with recall above 0.5. Errors remain in temporal direction, duration, and age-specific interpretations of behavior.

摘要:發展性篩檢影片顯示兒童如何執行特定行為,但臨床記錄通常包含結果而非發生的描述。我們提出SemReward-VL,該模型從這些結果中學習描述特定項目的行為。一個視覺-語言模型生成描述,而一個凍結的語言模型評分其與臨床結果的協議、對項目的相關性、對不相關的視頻-項目對的避免,以及清晰度。群體相對政策優化(GRPO)使用這一語義獎勵更新LoRA適配器。在涵蓋41個項目的13,379個視頻上,該方法提高了總體準確性和召回率超過0.5的項目數量。錯誤仍然存在於時間方向、持續時間和年齡特定的行為解釋中。

NutriVision: Ingredient-Conditioned Fusion and Prediction for Single-Image Food Nutrition Estimation

2609.33076v1 by Aman Kumar, Avinash Anand, Chaitanya Lakhchaura, Ashutosh Kumar, Akshita Abrol, Timothy Liu, Zhengkui Wang, Rajiv Ratn Shah

Nutrition estimation is a fundamental task in consumer diet tracking, clinical dietetics, chronic disease management, sports and hospital nutrition, and broader food computing systems. The existing approaches have progressed along two largely separate axes, vision models that rely on calibrated RGB-depth captures and ingredient-aware methods that use textual cues but use limited multimodal fusion. We introduce NutriVision, an end-to-end framework that leverages visual geometry and ingredient semantics to estimate calories, mass, fat content, carbohydrates, and protein from a single RGB image and an optional ingredient list. It obtains the unavailable depth modality using DepthAnything-V3 and encodes ingredient descriptions using CLIP. It integrates three complementary mechanisms: (1) an \emph{Ingredient-Conditioned Frequency-Aligned Fusion Module (IC-FAFM)}, which uses textual guidance to reweight and align RGB-depth frequency components; (2) an \emph{Ingredient-Aware Mask-based Prediction Head (IA-MPH)}, whose gating and channel masks are conditioned on food identity; and (3) modality-specific \emph{Internal Semantic Modeling (ISM)} blocks. On the Nutrition5k dataset, NutriVision achieves a mean PMAE of $\mathbf{13.60\pm0.10\%}$, outperforming our IGSMNet implementation by $0.89$ percentage points and OmniFood8k by $2.90$ percentage points (both $p<0.001$). The module-level ablations identify the ingredient-aware prediction head as the primary architectural contributor, improving mean PMAE by $1.50\pm0.17$ percentage points ($p<0.001$). These results demonstrate that ingredient-conditioned prediction and frequency-aware RGB-depth fusion provide measurable gains for single-image nutrient estimation. More broadly, NutriVision offers a practical route toward nutrition-assessment systems that exploit geometric and semantic cues without requiring specialized depth-sensing hardware

摘要:營養估算是消費者飲食追蹤、臨床營養學、慢性疾病管理、運動及醫院營養以及更廣泛的食品計算系統中的一項基本任務。現有的方法主要沿著兩個相對獨立的方向發展,一是依賴經過校準的RGB-深度捕捉的視覺模型,二是使用文本線索但進行有限的多模態融合的成分感知方法。我們介紹了NutriVision,一個端到端的框架,利用視覺幾何和成分語義從單一RGB圖像和可選的成分列表中估算卡路里、質量、脂肪含量、碳水化合物和蛋白質。它使用DepthAnything-V3獲取缺失的深度模態,並使用CLIP編碼成分描述。它整合了三個互補機制:(1)\emph{成分條件頻率對齊融合模塊(IC-FAFM)},利用文本指導重新加權和對齊RGB-深度頻率分量;(2)\emph{成分感知基於掩碼的預測頭(IA-MPH)},其閘控和通道掩碼依賴於食物身份;以及(3)特定模態的\emph{內部語義建模(ISM)}模塊。在Nutrition5k數據集上,NutriVision達到了$\mathbf{13.60\pm0.10\%}$的平均PMAE,超越了我們的IGSMNet實現$0.89$個百分點和OmniFood8k的$2.90$個百分點(均為$p<0.001$)。模塊級的消融實驗確定成分感知預測頭是主要的架構貢獻者,將平均PMAE提高了$1.50\pm0.17$個百分點($p<0.001$)。這些結果表明,成分條件預測和頻率感知的RGB-深度融合為單圖像營養素估算提供了可測量的增益。更廣泛地說,NutriVision提供了一條實用的路徑,朝向利用幾何和語義線索的營養評估系統,而無需專門的深度感測硬體。

TCMQA: A 38K-Question Traditional Chinese Medicine Benchmark with a Licensed-Practitioner Reference

2609.33014v1 by Tzu-Heng Huang, Jet Lin, Eric Lin

Medical benchmarks for language models are built almost entirely on Western biomedicine. Traditional Chinese Medicine (TCM) is a separate system, with its own diagnostic framework and its own literature, and it remains largely unmeasured. The few TCM evaluations that exist are small, narrow, and rarely paired with a human reference. We present TCMQA, an open benchmark of 38,279 questions from Chinese TCM licensing examinations, paired with 15,151 responses from 101 licensed practitioners. We evaluate 29 instruction-tuned models from 9 families, spanning 0.27B to 14.8B parameters. Accuracy ranges over 59 points, and no model approaches saturation. Pretraining data predicts TCM ability far better than scale: a 12B Western-pretrained model reaches 39.6%, while a Chinese-pretrained model an eighth its size reaches 60.8%. Nine models exceed the practitioner majority vote of 64.9%, the best by 21.8 points, and all nine come from that same Chinese-pretrained family. Yet difficulty does not transfer between models and practitioners: accuracy is flat across practitioner-rated difficulty, item-level agreement is near zero for all 29 models, and on $8.4\%$ of items the practitioners are correct where the leading model is wrong. We release the corpus, the practitioner responses, the harness, and per-item model outputs at https://huggingface.co/datasets/TechTCM/TCMQA.

摘要:醫療基準對於語言模型幾乎完全建立在西方生物醫學之上。傳統中醫(TCM)是一個獨立的系統,擁有自己的診斷框架和文獻,並且仍然在很大程度上未被測量。現有的少數中醫評估都很小且狹窄,且很少與人類參考相配對。我們提出了TCMQA,一個包含38,279個來自中國中醫執業考試問題的開放基準,並配有101名執業者的15,151個回答。我們評估了來自9個家族的29個指令調整模型,參數範圍從0.27B到14.8B。準確率的範圍超過59個點,且沒有模型接近飽和。預訓練數據對中醫能力的預測遠優於規模:一個12B的西方預訓練模型達到39.6%,而一個大小僅為其八分之一的中國預訓練模型達到60.8%。九個模型超過執業者的多數票64.9%,最佳模型超出21.8個點,且這九個模型均來自同一個中國預訓練家族。然而,難度在模型和執業者之間並不轉移:在執業者評定的難度上,準確率持平,29個模型在項目層面的協議接近於零,且在$8.4\%$的項目中,執業者正確而領先模型錯誤。我們在 https://huggingface.co/datasets/TechTCM/TCMQA 釋出語料庫、執業者回應、工具以及每個項目的模型輸出。

Multimodal LLMs Outperform Pathology Foundation Models in Cross-Domain Histological Similarity

2609.32876v2 by Yishu Zhang, Yun Li, Daiwei Zhang

State-of-the-art pathology foundation models, trained on millions of histology tiles, can fail to preserve tissue similarity when comparisons cross slide or institution boundaries. We show that general-purpose multimodal LLMs, without being trained as pathology foundation models, consistently outperform these specialized models in cross-domain histological similarity judgments. Using a relative similarity framework that we release as the MOSAIC (Model Similarity Assessment across Institutions and Cohorts) benchmark, we evaluate 17 models across 6 datasets and find that pathology encoders often rank same-institution, different-disease tiles as more similar than same-disease, different-institution tiles, a clinically dangerous failure mode invisible to standard within-domain evaluations. LLMs appear less susceptible to this failure, likely because they perform semantic visual comparison of morphology and tissue architecture rather than relying on shortcut features tied to acquisition context. Scaling training data does not resolve the problem for pathology encoders, implicating the learning objective rather than data coverage. Our results expose a fundamental robustness gap in current pathology foundation models and establish multimodal LLMs as a viable alternative for cross-institutional retrieval, dataset harmonization, and multi-site quality control. Code and data will be released upon acceptance.

摘要:最先進的病理基礎模型,經過數百萬個組織學切片的訓練,當比較跨越切片或機構邊界時,可能無法維持組織的相似性。我們顯示,通用的多模態 LLM,未經訓練為病理基礎模型,始終在跨領域的組織學相似性判斷中超越這些專門模型。我們使用一個相對相似性框架,並將其發布為 MOSAIC(跨機構和隊列的模型相似性評估)基準,評估 17 個模型在 6 個數據集上的表現,發現病理編碼器經常將同一機構、不同疾病的切片評為比同一疾病、不同機構的切片更相似,這是一種臨床上危險的失敗模式,對標準的領域內評估來說是不可見的。LLM 似乎對這種失敗的敏感性較低,這可能是因為它們執行形態和組織結構的語義視覺比較,而不是依賴於與獲取上下文相關的捷徑特徵。擴大訓練數據並未解決病理編碼器的問題,這表明學習目標而非數據覆蓋是關鍵。我們的結果揭示了當前病理基礎模型中的基本穩健性差距,並確立了多模態 LLM 作為跨機構檢索、數據集協調和多地點質量控制的可行替代方案。代碼和數據將在接受後發布。

Counterfactual Self-Evolving Agents for Evidence-Grounded Reasoning

2609.32870v1 by Xing Han, Yuxin Wang, Chen Chen, Wei Dai, Gautham Krishna Gudur, Shijun Li, Hsing-Huan Chung, Gregory D. Hager, Joydeep Ghosh, Paul Pu Liang, Suchi Saria

Self-play proposer--solver methods improve reasoning by generating tasks and learning from verified solutions. However, for evidence-identifiable tasks, where case-specific evidence and domain knowledge determine a checkable answer, self-play requires generating plausible cases whose answers can be independently verified. We introduce counterfactual self-evolution, which generates counterfactual context for reconsidering the original case. A trainable Proposer constructs targeted evidence edits and describes potential outcome changes with causal explanations. We handcraft an expert-verified counterfactual instruction-tuning dataset to teach the Proposer to generate high-quality counterfactuals across a broad range of action--outcome scenarios. Each counterfactual instruction-tuning example specifies an edit within a defined category and explains its hypothesized causal effect on the decision, teaching the Proposer to reason systematically about what changes and why. We instruction-tune the Proposer on these examples, then formulate a fine-tuning reward that integrates feedback from the Solver and Verifier. Across diverse counterfactual scenarios, this reward favors high-quality counterfactuals and warranted revisions, while penalizing changes that overturn correct decisions. The counterfactual context aims to correct errors and strengthen confidence in correct decisions. Accepted counterfactuals accumulate in memory that supplies in-context evidence to the frozen Solver; the Solver adapts through evolving context rather than weight updates. We apply the framework to clinical reasoning, fact verification, and business reasoning. Our evaluation tracks performance over successive rounds as counterfactual memory grows, including transfer to harder cases. Our method achieves superior results across diverse frontier models.

摘要:自我對弈提議者--解決者方法通過生成任務並從經過驗證的解決方案中學習來改善推理。然而,對於可識別證據的任務,其中案例特定的證據和領域知識決定了可檢查的答案,自我對弈需要生成可以獨立驗證的合理案例及其答案。我們引入了反事實自我演化,該方法生成反事實背景以重新考慮原始案例。一個可訓練的提議者構建針對性的證據編輯,並用因果解釋描述潛在的結果變化。我們精心製作了一個專家驗證的反事實指令調整數據集,以教導提議者在廣泛的行動--結果場景中生成高質量的反事實。每個反事實指令調整示例指定了一個在定義類別內的編輯,並解釋其對決策的假設因果效應,教導提議者系統性地推理什麼變化以及為什麼變化。我們在這些示例上對提議者進行指令調整,然後制定一個微調獎勵,該獎勵整合了解決者和驗證者的反饋。在多樣的反事實場景中,這個獎勵偏好高質量的反事實和合理的修訂,同時懲罰推翻正確決策的變更。反事實背景旨在糾正錯誤並增強對正確決策的信心。被接受的反事實在記憶中累積,為凍結的解決者提供上下文證據;解決者通過演變的背景而不是權重更新來適應。我們將該框架應用於臨床推理、事實驗證和商業推理。我們的評估跟踪隨著反事實記憶增長而進行的多輪性能,包括轉移到更困難的案例。我們的方法在多樣的前沿模型中取得了優越的結果。

FinancialAuditBench: Benchmark Construction under Differential Privacy Using Real-World Priors

2609.32835v1 by Jerry Huang, Sarvesh Babu, Matt Van Buren, Alexander Wang, Pranav Pillai, Arush Jain, James P. Burton, Julia Hockenmaier

As AI agents are becoming widely adopted in the financial services industry, careful measurement is essential to understand where they can be reliably deployed and where oversight and professional review remain necessary. Such measurement, however, is constrained by limited access to proprietary or privacy-sensitive data. Existing benchmarks therefore often rely on publicly available data, human- and/or LLM-authored tasks, or simplified settings. We introduce FinancialAuditBench, a benchmark for evaluating agents on financial statement audit tasks, along with a framework for systematically generating synthetic engagements. Our task generation framework leverages differentially private aggregate statistics from historical audits along with audit expertise contributed through over 1,100 hours of benchmark development and review. FinancialAuditBench consists of 90 tasks spanning workpaper completion and review across six synthetic audit engagements, each containing an average of 179 files. Evaluation on eleven frontier models shows that while agents complete substantial portions of staff-level audit tasks well, they sometimes perform inappropriate procedures or produce incorrect documentation. Beyond financial auditing, our framework offers an approach for systematically generating synthetic tasks for model evaluation and training in privacy-sensitive domains.

摘要:隨著 AI 代理在金融服務行業的廣泛採用,仔細的測量對於理解它們可以可靠部署的地方以及何處仍需監督和專業審查至關重要。然後,這種測量受到對專有或隱私敏感數據的有限訪問的限制。因此,現有的基準通常依賴於公開可用數據、人類和/或 LLM 編寫的任務或簡化的設置。我們介紹了 FinancialAuditBench,一個用於評估代理在財務報表審計任務上的基準,以及一個系統生成合成參與的框架。我們的任務生成框架利用了來自歷史審計的差分隱私聚合統計數據,以及通過超過 1,100 小時的基準開發和審查貢獻的審計專業知識。FinancialAuditBench 包含 90 個任務,涵蓋六個合成審計參與的工作文件完成和審查,每個參與平均包含 179 個文件。對十一個前沿模型的評估顯示,儘管代理能夠很好地完成大量的員工級審計任務,但有時它們會執行不當的程序或產生不正確的文件。除了財務審計,我們的框架還提供了一種系統生成合成任務的方法,用於在隱私敏感領域進行模型評估和訓練。

LLM

Publish Date Title Authors Homepage Code
2026-10-01 One Basis to Animate Them All: Gaussian Blendshape Distillation for Real-Time Avatars Ramazan Fazylov et.al. 2610.02207v1 null
2026-10-01 KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards Pengfei Li et.al. 2610.02206v1 null
2026-10-01 Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents Yen-Jen Wang et.al. 2610.02204v1 null
2026-10-01 ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research Sohyeon Kim et.al. 2610.02202v1 null
2026-10-01 VISTA: A Visual Harness for Reasoning in an Interactive World Qiushi Han et.al. 2610.02200v1 null
2026-10-01 Hierarchical Continuous Diffusion Language Models Hui Ren et.al. 2610.02193v1 null
2026-10-01 DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation Zhengming Yu et.al. 2610.02188v1 null
2026-10-01 Higher-Order Molecular Grammars for Generative and Foundation Models in Chemistry Yiming Huang et.al. 2610.02186v1 null
2026-10-01 SoftServe: A Scalable Quasi-Newton Method for Deep Learning Joohwan Ko et.al. 2610.02182v1 null
2026-10-01 Generative Cinematographer: Composing Camera and Object Motion in 3D Jiahan Zhang et.al. 2610.02180v1 null
2026-10-01 Every Ablation Is a Dose: Counterweights and the Semblance of Self-Repair Areeb Ahmad et.al. 2610.02173v1 null
2026-10-01 AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents Xuan Zhang et.al. 2610.02163v1 null
2026-10-01 DuoMind: Enabling Distributed Multi-Robot Coordination with Semantic Communication Hanchu Zhou et.al. 2610.02161v1 null
2026-10-01 From Knowledge Access to Source Learning: Developing Source-Specific Competence Lucheng Fu et.al. 2610.02150v1 null
2026-10-01 Keyword Harnesses Fail Open: A Cheap Diagnostic Ladder for Tool-Use Claims in Small Language Models Juan S. Santillana et.al. 2610.02142v1 null
2026-10-01 Finetuning with Sampling: SFT Learns Better Than You Think Aayush Karan et.al. 2610.02140v1 null
2026-10-01 MIRTO: a registration-gated, multiverse-tested evaluation protocol for unsupervised anomaly segmentation in brain MRI Negin Kafee Hernashki et.al. 2610.02136v1 null
2026-10-01 Local Support Learning Assaf Ben-Kish et.al. 2610.02126v1 null
2026-10-01 Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows Gabriel Tomitsuka et.al. 2610.02122v1 null
2026-10-01 Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes Sophia Sirko-Galouchenko et.al. 2610.02117v1 null
2026-10-01 A Comparative Explainability Framework for DeBERTa-v3 in Zero-Shot Medical Abstract Classification Javier Diaz Esteban-Herreros et.al. 2610.02116v1 null
2026-10-01 Scalable, Transferable Meta-network for Data Selection Requires a Different Loss (and Why the Obvious Choice is Problematic) Zilin Du et.al. 2610.02092v1 null
2026-10-01 GeoLatent: Geometry-Guided Latent Structuring with Routed Optimization for 3D Reasoning Yakun Zhu et.al. 2610.02091v1 null
2026-10-01 LLM2Jev: LLMs Are Already Jev-Style Decision Models -- When and How to Fine-Tune Them Yinheng Li et.al. 2610.02076v1 null
2026-10-01 Causal Memory Policy: Making Memory Utility Identifiable by Intervening on Retrieval Arman Behnam et.al. 2610.02070v1 null
2026-10-01 External Observers May See More Clearly: Cross-Model Span-Level Hallucination Detection in Large Language Models via Hidden State Probing Kingshuk Gupta et.al. 2610.02066v1 null
2026-10-01 HydroJEV: A one-second, training-free screen for cyber-attack and fault attribution in water distribution networks Tianwei Mu et.al. 2610.02048v1 null
2026-10-01 Typological Alignment of Stack-Based Language Models on Mildly Context-Sensitive Artificial Languages Nadine El-Naggar et.al. 2610.02040v1 null
2026-10-01 CARM: Cancellation-Aware Response Masking for LLM Reinforcement Learning Yafei Zhang et.al. 2610.02039v1 null
2026-10-01 Mimir: Physics-Grounded LLM Agents for Long-Horizon Irrigation Control Yimeng Liu et.al. 2610.02038v1 null
2026-10-01 Global Coherence: When Every Agent Is Right and the Team Is Still Wrong - A Local-to-Global Semantic Foundation for Multi-Agent Collaboration Xin Heng et.al. 2610.02036v1 null
2026-10-01 SPHERE: Adaptive VR Indoor Scene Generation via LLM-Enhanced Spatial Preference Learning and Human-in-the-Loop RL Hyeonmin Lee et.al. 2610.02023v1 null
2026-10-01 Old Ideas, Novel Problems: The Instability of LLM-Based Novelty Evaluation Noy Sternlicht et.al. 2610.02022v1 null
2026-10-01 Task-Adaptive Grounded 3D-Programmers Using 2D VLMs Arman Raayatsanati et.al. 2610.02021v1 null
2026-10-01 Controllable Multi-label Video Safety Detection via Adaptive Tversky Policy Optimization Guangyu Yang et.al. 2610.02019v1 null
2026-10-01 On Language Drift during RLVR Post-Training Michael Sullivan et.al. 2610.02015v1 null
2026-10-01 Atoms to Processes: The Role of Artificial Intelligence and Machine Learning in Chemical Engineering Michael Baldea et.al. 2610.02014v1 null
2026-10-01 Counting Moves, Weighing Voices: Bayesian Dialectical Argumentation for Calibrated Multi-LLM Councils under Persistent Adversaries Ionel Eduard Stan et.al. 2610.02005v1 null
2026-10-01 Mem++: Non-Destructive Memory for Long-Term Organizational LLM Agents Ahmad Yehia et.al. 2610.02002v1 null
2026-10-01 Mingbird: A Local-First Agent Harness Enabling Small Open Models to Complete Real Tasks Hao Wang et.al. 2610.02001v1 null
2026-10-01 Can AI Oversight Be Zero Knowledge? Alessandro Chiesa et.al. 2610.01995v1 null
2026-10-01 Universal Byte-Level Encoding: UTF-8/UTF-16 Routing to Reduce Cross-Script Token-Budget Disparities Hyunsik Kim et.al. 2610.01984v1 null
2026-10-01 Counterfactual Auditing of Bias in Open-Source Large Language Models for Clinical Triage Manar Aljohani et.al. 2610.01963v1 null
2026-10-01 A Hybrid Approach to Malware Detection: Integrating Few-Shot Model-Agnostic Meta-Learning with Autoencoders Emmanuela Andam et.al. 2610.01949v1 null
2026-10-01 Latent JEPA: Abstract Future Prediction for Latent Reasoning in Chemistry Xinjian Zhao et.al. 2610.01947v1 null
2026-10-01 A rubric landscape for evaluating clinical reasoning in large language models: what exists, what is missing, and what needs to be combined Zhangshu Joshua Jiang et.al. 2610.01938v1 null
2026-10-01 Mapping the RAG Landscape: A Four Axis Taxonomy of Efficiency, Defense, Interactivity, and Reasoning Meghana Sunil et.al. 2610.01936v1 null
2026-10-01 Cross-Lingual Alignment for Decoder-Only Models using MoE Routers Lucas Bandarkar et.al. 2610.01921v1 null
2026-10-01 MoLE: Mixture of Latent Experts for Complementary Visual Reasoning Yingcheng Liu et.al. 2610.01917v1 null
2026-10-01 Asynchronous LLM Post-Training: Group-Mass Capping and Convergence Analysis Qijia He et.al. 2610.01896v1 null
2026-10-01 A Structured State Space Sequence Model for Multi-Class Classification of Malware Emmanuela Andam et.al. 2610.01893v1 null
2026-10-01 Selection-Based Structured Reasoning: Toward Efficient Multimodal Search Agents Feiyu Gavin Zhu et.al. 2610.01892v1 null
2026-10-01 Unsupervised Domain Adaptation for Enhanced Radiometer Image Precipitation Estimation using Conditional Flow Matching Victor Enescu et.al. 2610.01890v1 null
2026-10-01 Stochastic Rounding in Low-Precision Transformer Inference: A Variable-Precision Emulation Study of a Small GPT-2 Yohan Chatelain et.al. 2610.01889v1 null
2026-10-01 Flowing Faster to Coordinate: One-Step Online Multi-Agent Flow Policies Zhuoran Li et.al. 2610.01882v1 null
2026-10-01 Where LLMs Fail with Visualization DSLs Chang Han et.al. 2610.01873v1 null
2026-10-01 From Network Intrusion Detection to Blockchain-Backed Endpoint Detection and Response: Mapping the Landscape of Decentralized Detection-and-Response Architectures Yahya Shahsavari et.al. 2610.01872v1 null
2026-10-01 Walking the Embedding Space: Datastore Extraction from Multimodal RAG Maria Carmen Jica et.al. 2610.01871v1 null
2026-10-01 From Isolated Feature to Orbits: Discovering Music Concepts via Multi-SAE Alignment Liwei Lin et.al. 2610.01864v1 null
2026-10-01 AVSD-Scenes: A Dataset for Audio-Visual Description of Urban Scenes Dhanunjaya Varma Devalraju et.al. 2610.01861v1 null
2026-10-01 Detecting Inconsistencies in Model Specifications with LLM-as-Verifier Reasoning Zichen Xie et.al. 2610.01847v1 null
2026-10-01 Beyond Decodability: Do Acoustic Factors Drive Predictions in Speech-Based Alzheimer's Assessment? Serli Kopar et.al. 2610.01846v1 null
2026-10-01 On the Divergence of Accuracy and Mechanism Consistency in Time Series World Models Haochen Zhang et.al. 2610.01842v1 null
2026-10-01 Code Owns the Simulation, Jev Owns the Evaluation Yaodong Yang et.al. 2610.01834v1 null
2026-10-01 Continuous Process-Level Evaluation for Evolving Enterprise AI Agent Skills Ngoc Phuoc An Vo et.al. 2610.01833v1 null
2026-10-01 The Asymptotics of Language Model Alignment with Memory Haricharan Balasundaram et.al. 2610.01828v1 null
2026-10-01 Token Communication-Assisted Collaborative Embodied Artificial Intelligence: Concepts, Framework, and Opportunities Peng Yi et.al. 2610.01826v1 null
2026-10-01 Beyond Linear Concepts: Discovering and Aligning Non-Linear Concept Manifolds in Large Language Models Tido Specht et.al. 2610.01821v1 null
2026-10-01 A Safe Prototype Is Not a Safety Direction: Reference Dependence and Prompt Confounds in Response-Safety Embeddings Sahil Kadadekar et.al. 2610.01801v1 null
2026-10-01 LineupRL: Verifiable Reinforcement Learning for Time Series Captioning via Caption-to-Series Identification Haochen Zhang et.al. 2610.01800v1 null
2026-10-01 iADD: Improving Alignment and Diversity in Diffusion Policy Optimization Ashok Prasad Neupane et.al. 2610.01789v1 null
2026-10-01 VETO: Video Efficient Token Optimization for Vision Language Models Gueter Josmy Faure et.al. 2610.01785v1 null
2026-10-01 Q-Learning for Reachability in MEC-Free MDPs Lu-Chin Chang et.al. 2610.01781v1 null
2026-10-01 CODesign: Consistency from Data to Trajectory in All-Atom Protein Binder Co-Design Yuanle Mo et.al. 2610.01773v1 null
2026-10-01 A Matryoshka Hierarchical RAG for Efficient Multi-Hop Question Answering Gianluca Bonifazi et.al. 2610.01767v1 null
2026-10-01 VideoEvolve: Evolving Agent Harnesses for Video Temporal Grounding Bingjun Luo et.al. 2610.01766v1 null
2026-10-01 TopK-Guided: Adaptive, Budget-Aware Activation Sparsity for Efficient LLM Inference Mukund Agarwalla et.al. 2610.01763v1 null
2026-10-01 SoK: Decentralized Agent Economic Infrastructure Rui Sun et.al. 2610.01756v1 null
2026-10-01 Cog-VADU: A Training-Free Cognitive Reasoning Framework for Video Anomaly Detection and Understanding Mohd Ubaid Wani et.al. 2610.01754v1 null
2026-10-01 Removing spurious minima for planar features by skip connections Jakob Paul Zimmermann et.al. 2610.01728v1 null
2026-10-01 vFedProtoQNAS: Prototype-Guided Personalized Quantum Neural Architecture Search for Virtual Federated Learning Seok Bin Son et.al. 2610.01718v1 null
2026-10-01 CoEvolve: Construct-to-Edit Visual Grounding with Bidirectional State Refinement Dongwei Sun et.al. 2610.01710v1 null
2026-10-01 Task-Oriented Rank Adaptation for Continual Learning in Text Classification Rey Sanchez Lopez et.al. 2610.01702v1 null
2026-10-01 Acmite: Mitigating Gender Bias in LLMs through Concept-Guided Mutual Information Tian Lan et.al. 2610.01696v1 null
2026-10-01 Compound interpretation is based on analogy Tian Shen et.al. 2610.01688v1 null
2026-10-01 Architectural Sampling: Test-Time Scaling via Computational Diversity in Frozen Vision-Language Models Akshit Singh et.al. 2610.01687v1 null
2026-10-01 Iterative Policy Refinement through Semantic Rollout Analysis Feiyu Gavin Zhu et.al. 2610.01652v1 null
2026-10-01 MCIR: A Feature Dependence-Aware Explainability Method with Reliability Guarantees Poushali Sengupta et.al. 2610.01641v1 null
2026-10-01 Not All Error Yields to Scale: Where Scaling Stops in Vision-Language Inference Xinye Zhao et.al. 2610.01640v1 null
2026-10-01 Yo-ByT5: Efficient and High-Fidelity Diacritic Restoration for Yorùbá Ahmad Samuel Gali et.al. 2610.01634v1 null
2026-10-01 What Makes Something Hard(er)? Explaining Question Difficulty in Natural Language Peng Cui et.al. 2610.01627v1 null
2026-10-01 FedLore: Communication and Memory Efficient Federated Learning via Shared Gradient Low-Rank Projection Junkang Liu et.al. 2610.01620v1 null
2026-10-01 Exposing the Cost of Deep Learning Audio Development Constance Douwes et.al. 2610.01619v1 null
2026-10-01 Agents Are Systems, Not Models: Rethinking Agentic Evaluation Luis Wiedmann et.al. 2610.01618v1 null
2026-10-01 Can LLMs Reliably Annotate Bioassay Metadata to Improve Data Readiness? Laura van Weesep et.al. 2610.01616v1 null
2026-10-01 Hob-VL: A Benchmark for Visually Grounded Boolean Reasoning Yuzhou Wang et.al. 2610.01605v1 null
2026-10-01 Permutation-Robust Decision Modeling with Candidate-Independent Block-Causal Attention Guy Amit et.al. 2610.01601v1 null
2026-10-01 Before It Fades: Reinforcing Temporal Representations at Inference Time in VideoLLMs Youngwoo Shin et.al. 2610.01595v1 null
2026-10-01 Which LLM to pick? Online Active Model Selection for Large Language Models Alessandro Turrin et.al. 2610.01592v1 null
2026-10-01 Evaluating Physical Consistency and Plausibility in Generative Scenario Models for Autonomous Driving Manasa Mariam Mammen et.al. 2610.01581v1 null

Abstracts

One Basis to Animate Them All: Gaussian Blendshape Distillation for Real-Time Avatars

2610.02207v1 by Ramazan Fazylov, Stamatis Lefkimmiatis, Ivan Laptev

3D Gaussian avatars support fast rendering, however, their real-time animation is often challenged by the costly neural inference. We address this bottleneck and show that the animation of pretrained avatar models can be closely approximated by a linear combination of identity-independent blendshapes. Building on this finding, we introduce GALA (Gaussian Animation via Linear Approximation), a distillation method that replaces per-frame heavy neural decoding with a shallow coefficient predictor and a linear blend. To improve fidelity and reduce memory requirements, we propose to construct the basis using block-local PCA under a rendering-aware metric and a memory budget. Our method learns a shallow MLP network to predict blendshape coefficients and applies to various animation architectures without retraining original models. We validate GALA by accelerating the inference of three distinct avatar models for 3D animation of facial expressions and full-bodies with clothing dynamics. Across these models, our distillation generalizes to held-out identities and reduces CPU animation cost by up to three orders of magnitude while preserving most of the rendering quality. Excellent results of our method confirm the shared linear structure of learned avatar representations and enable highly efficient and accurate animation at frame rates reaching up to 60fps on mobile devices. Project page: https://ramazan793.github.io/gala/

摘要:3D 高斯虛擬角色支援快速渲染,然而,它們的即時動畫常常受到昂貴的神經推理的挑戰。我們解決了這一瓶頸,並顯示預訓練虛擬角色模型的動畫可以通過身份無關的混合形狀的線性組合來密切近似。基於這一發現,我們引入了 GALA(通過線性近似的高斯動畫),這是一種蒸餾方法,將每幀重的神經解碼替換為淺層係數預測器和線性混合。為了提高保真度並減少內存需求,我們提出在渲染感知度量和內存預算下使用區塊局部 PCA 來構建基底。我們的方法學習一個淺層 MLP 網絡來預測混合形狀係數,並應用於各種動畫架構而無需重新訓練原始模型。我們通過加速三個不同虛擬角色模型的推理來驗證 GALA,這些模型用於 3D 動畫的面部表情和全身帶衣物動態。在這些模型中,我們的蒸餾方法對保留的身份進行了泛化,並將 CPU 動畫成本降低了多達三個數量級,同時保持大部分渲染質量。我們方法的優異結果證實了學習的虛擬角色表示的共享線性結構,並使得在移動設備上以高達 60fps 的幀率進行高效且準確的動畫成為可能。專案頁面:https://ramazan793.github.io/gala/

KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards

2610.02206v1 by Pengfei Li, Naufal Suryanto, Sicheng Zhang, Muzammal Naseer

LLMs are increasingly applied to cybersecurity workflows, where they are expected to translate analysts' intent into tool invocations. However, existing evaluations focus on knowledge-based assessments or end-to-end agentic tasks, and do not directly measure LLMs' ability to generate executable commands for real-world cybersecurity tools. This gap is critical because cybersecurity operations rely on strict command-line interfaces (CLIs), where minor syntax errors, incorrect flag--value bindings, or argument misordering can invalidate execution. We introduce KaliBench, a fine-grained benchmark and dataset for natural-language--to--CLI translation on Kali Linux, comprising 8,504 query--command pairs spanning 1,642 tools across 23 capability dimensions and 5 security phases. KaliBench is constructed via a manuscript-grounded pipeline with deterministic canonicalization and alias-aware evaluation, enabling precise and reproducible assessment of tool selection and argument construction. To ensure both semantic correctness and practical executability, we develop a multi-stage verification pipeline that combines LLM-based validation, sandboxed terminal execution, and human-in-the-loop refinement. Building on these fine-grained, deterministic signals, KaliBench further enables runtime-free verifiable rewards for training. Across three evaluation modes and 24 configurations of general-purpose and security-focused open-weight models, no open-weight model exceeds 42% exact-command accuracy in the unrestricted setting, highlighting the difficulty of accurate CLI-based cybersecurity tool use without explicit tool hints. We further show that supervised fine-tuning and reinforcement learning with verifiable rewards derived from KaliBench significantly improve an 8B model and achieve performance comparable to a 685B MoE model.

摘要:LLMs 正在越來越多地應用於網絡安全工作流程中,它們被期望將分析師的意圖轉化為工具調用。然而,現有的評估主要集中在基於知識的評估或端到端的代理任務上,並未直接測量 LLMs 生成可執行命令以供現實世界網絡安全工具使用的能力。這一差距至關重要,因為網絡安全操作依賴於嚴格的命令行界面 (CLIs),其中微小的語法錯誤、不正確的標誌--值綁定或參數錯序都可能使執行無效。我們介紹了 KaliBench,這是一個針對 Kali Linux 的自然語言到 CLI 翻譯的細粒度基準和數據集,包含 8,504 個查詢--命令對,涵蓋 1,642 種工具,跨越 23 個能力維度和 5 個安全階段。KaliBench 是通過一個基於手稿的管道構建的,具有確定性的標準化和別名感知評估,能夠精確且可重複地評估工具選擇和參數構建。為了確保語義正確性和實際可執行性,我們開發了一個多階段驗證管道,結合了基於 LLM 的驗證、沙盒終端執行和人類參與的精煉。基於這些細粒度的確定性信號,KaliBench 進一步使得無運行時的可驗證獎勵成為訓練的可能。在三種評估模式和 24 種通用及安全專注的開放權重模型配置中,沒有任何開放權重模型在不受限制的設置中超過 42% 的精確命令準確率,突顯了在沒有明確工具提示的情況下準確使用基於 CLI 的網絡安全工具的困難。我們進一步顯示,從 KaliBench 派生的可驗證獎勵的監督微調和強化學習顯著改善了一個 8B 模型,並達到了與 685B MoE 模型相當的性能。

Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents

2610.02204v1 by Yen-Jen Wang, Haozhe Jiang, Shuying Deng, Haoru Xue, Weirui Ye, Rocky Duan, Nika Haghtalab, S. Shankar Sastry, Pieter Abbeel, Haozhi Qi

Building reliable robot capabilities across diverse tasks requires substantial human effort to develop and maintain skills, design rewards, and integrate perception with control. We present Reconstruct, Practice, Go Real (RPG), a framework for autonomous improvement of robot execution systems without updating model weights. RPG identifies manipulation capabilities in an offline dataset and constructs related practice tasks in simulation. During practice, RPG uses execution feedback, privileged simulator state, and available dataset videos to diagnose failures. It develops new reusable symbolic skills, refines existing skills, and revises the system prompt based on these diagnoses. Cross-task evaluation tests individual candidate changes and merged revisions before they are retained for reuse. At test time, a multimodal LLM uses the resulting system prompt and skill library to coordinate perception and robot control. On held-out initializations of 22 manipulation tasks, RPG improves task success from 28.6% after the first practice round to 95.0% after 15 rounds, outperforming all evaluated baselines, including ASPIRE (75.5%) and CaP-Agent0 powered by GPT-6 Astra Pro (60.0%). After a common calibration and hardware-adaptation procedure, the frozen system succeeds in all 30 physical trials, with ten trials on each of three tasks. Project Website: https://rpg-robot.github.io/

摘要:建立可靠的機器人能力以應對多樣化任務需要大量的人力來發展和維護技能、設計獎勵,以及將感知與控制整合在一起。我們提出了重建、練習、實現真實(RPG),這是一個無需更新模型權重的自主改進機器人執行系統的框架。RPG 在離線數據集中識別操作能力,並在模擬中構建相關的練習任務。在練習過程中,RPG 利用執行反饋、特權模擬器狀態和可用的數據集視頻來診斷失敗。它開發新的可重用符號技能,細化現有技能,並根據這些診斷修訂系統提示。跨任務評估測試單個候選變更和合併修訂,然後才將其保留以便重用。在測試時,多模態 LLM 使用生成的系統提示和技能庫來協調感知和機器人控制。在 22 個操作任務的保留初始化中,RPG 將任務成功率從第一次練習回合後的 28.6% 提高到 15 回合後的 95.0%,超越了所有評估的基準,包括 ASPIRE(75.5%)和由 GPT-6 Astra Pro 驅動的 CaP-Agent0(60.0%)。在經過共同的校準和硬體適應程序後,凍結的系統在所有 30 次實體試驗中成功,每個任務進行十次試驗。項目網站:https://rpg-robot.github.io/

ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research

2610.02202v1 by Sohyeon Kim, Yoonho Lee, Bo Liu, Dayoon Ko, Rulin Shao, Seungone Kim, Graham Neubig, Pang Wei Koh, Aakanksha Chowdhery, Akari Asai, Omar Khattab, Yejin Choi, Gunhee Kim, Chelsea Finn

What makes great scientists great? Even as AI systems start to make progress on open problems, scientists remain far ahead of them at sensing which prior idea, buried in an ever-growing archive of research, a new problem needs. To study this skill, we draw on researchers who know firsthand which earlier work advanced their completed projects, with papers serving as pointers to the ideas within. Using our automated pipeline that makes author annotation scalable, we build ScholarCatalyst by having 184 lead authors of 207 recent computer science papers label which candidates did or could have advanced their project, each with a detailed rationale. We introduce a retrieval task with author-provided judgments: given an initial research question, retrieve these papers from only the literature available when the project began. Agentic search does no better than embedding retrieval (0.42 vs. 0.48 Recall@20) despite calling that same retriever as a tool. Even an agent built on Claude Fable 5.1, which may have seen the completed papers during training, reaches only 0.51 R@20. These results highlight the need for new training recipes that equip models with expert intuition for searching broad corpora. We envision ScholarCatalyst as a step toward scientific agents that can take a half-formed idea and point to the prior research it needs.

摘要:什麼使偉大的科學家偉大?即使人工智慧系統開始在開放問題上取得進展,科學家們在感知新問題需要哪個埋藏在不斷增長的研究檔案中的先前想法方面仍然遙遙領先。為了研究這項技能,我們依賴於那些親身了解哪些早期工作推進了他們完成項目的研究者,論文作為指向內部想法的指標。利用我們的自動化流程,使作者標註可擴展,我們通過讓184位207篇近期計算機科學論文的主要作者標註哪些候選者推進了或可能推進了他們的項目,每個標註都有詳細的理由,來構建ScholarCatalyst。我們引入了一個檢索任務,並提供作者的判斷:給定一個初步的研究問題,從項目開始時僅可用的文獻中檢索這些論文。即使將同一檢索工具稱為工具,主動搜索的表現也不如嵌入檢索(0.42對0.48 Recall@20)。即使是基於Claude Fable 5.1構建的代理,可能在訓練期間看過已完成的論文,R@20的結果也僅為0.51。這些結果突顯了需要新的訓練配方,以使模型具備專家直覺來搜索廣泛的語料庫。我們將ScholarCatalyst視為邁向科學代理的一步,這些代理可以將半成型的想法指向所需的先前研究。

VISTA: A Visual Harness for Reasoning in an Interactive World

2610.02200v1 by Qiushi Han, Keya Hu, Linlu Qiu, Cathy Wu, Kaiming He

We show that multimodal models possess strong reasoning abilities and that an appropriate harness can unlock their potential to solve tasks across diverse interactive environments. We introduce VISTA, a visual harness that gives a general-purpose multimodal model long-horizon vision. VISTA allows the model to directly perceive the environment through visual observations and maintains a lossless visual memory that preserves past observations in their original form. The model can actively retrieve these observations and reorganize its visual input as it reasons. On ARC-AGI-3, VISTA improves Claude Opus 5.0's Relative Human Action Efficiency score from 40.68 to a perfect 100.00, with the model completing all 25 public games using 57.4% fewer actions than first-time human participants. VISTA's simple design also allows it to extend naturally to diverse visual environments with minimal adaptation. Across three additional benchmarks covering a diverse range of visual games and puzzles, it substantially outperforms baselines using the same underlying model with minimal harnesses. Our results highlight VISTA's potential as a general-purpose visual harness for advancing multimodal agents in complex visual environments.

摘要:我們展示了多模態模型擁有強大的推理能力,並且適當的工具可以釋放它們在多樣互動環境中解決任務的潛力。我們介紹了 VISTA,一種視覺工具,使通用多模態模型具備長期視覺能力。VISTA 允許模型通過視覺觀察直接感知環境,並保持無損的視覺記憶,保留過去觀察的原始形式。模型可以主動檢索這些觀察並在推理時重新組織其視覺輸入。在 ARC-AGI-3 上,VISTA 將 Claude Opus 5.0 的相對人類行動效率分數從 40.68 提升至完美的 100.00,模型在完成所有 25 個公共遊戲時,使用的行動比首次參與的人類參與者少 57.4%。VISTA 的簡單設計也使其能夠在不同的視覺環境中自然擴展,適應性極小。在涵蓋多樣視覺遊戲和謎題的三個額外基準中,它顯著超越了使用相同基礎模型和最小工具的基準。我們的結果突顯了 VISTA 作為推進多模態代理在複雜視覺環境中的通用視覺工具的潛力。

Hierarchical Continuous Diffusion Language Models

2610.02193v1 by Hui Ren, Zihan Li, Chang Liu, Huidong Liu, Alexander Schwing

Discrete diffusion language models offer a compelling alternative to autoregressive generation for tasks demanding bidirectional reasoning and global constraint satisfaction. Yet they share a structural bottleneck: when decoding in parallel, each token is sampled independently from its marginal, severing the statistical dependencies among the tokens decoded together. Continuous diffusion language models avoid this by denoising a shared continuous state, but their denoiser sees only that state, so nothing ties it to a valid token configuration until it is finally decoded. To address this, we propose Hierarchical Continuous Diffusion Language Models (HC-DLM), which couple discrete token generation with a continuous latent trajectory in a single, principled denoising process, whose training objective is derived from a variational bound on the token likelihood. In contrast to recent methods that attach continuous context to a self-contained discrete chain, HC-DLM makes the latent the only persistent generative state: tokens are read out from it at every step and feed back as a scaffold for the next latent update. On structured reasoning (Sudoku), mathematical planning (Countdown) and language modeling (LM1B), HC-DLM improves over discrete and continuous diffusion baselines at matched model size, in puzzle accuracy on Sudoku and Countdown and in generative perplexity on LM1B. Project page: https://hc-dlm.github.io/.

摘要:離散擴散語言模型為需要雙向推理和全局約束滿足的任務提供了一個引人注目的替代方案,取代自回歸生成。然而,它們共享一個結構瓶頸:在並行解碼時,每個標記都是獨立從其邊際中抽樣的,這切斷了一起解碼的標記之間的統計依賴關係。連續擴散語言模型通過去噪共享的連續狀態來避免這一點,但它們的去噪器僅看到該狀態,因此在最終解碼之前,沒有任何東西將其與有效的標記配置聯繫起來。為了解決這個問題,我們提出了層次連續擴散語言模型(HC-DLM),它將離散標記生成與單一、原則性的去噪過程中的連續潛在軌跡結合起來,其訓練目標源自於標記似然的變分界限。與最近將連續上下文附加到自包含離散鏈的方法相比,HC-DLM使潛在變量成為唯一持久的生成狀態:標記在每一步都從中讀出,並作為下一次潛在更新的支架進行反饋。在結構推理(數獨)、數學規劃(倒計時)和語言建模(LM1B)方面,HC-DLM在匹配模型大小的情況下優於離散和連續擴散基準,在數獨和倒計時的拼圖準確性以及LM1B的生成困惑度上均有所改善。項目頁面:https://hc-dlm.github.io/.

DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation

2610.02188v1 by Zhengming Yu, Junkun Yuan, Haotian Yang, Gordon Guocheng Qian, Yizhi Wang, Angtian Wang, Yiding Yang, Bo Liu, Xin Li, Wenping Wang, Chongyang Ma

Distribution Matching Distillation (DMD) trains a few-step student from the difference between separately estimated target and student scores, so it must keep an auxiliary diffusion model fitted to the student's evolving distribution at extra memory and computation cost. We introduce DMAD, Distribution Matching as Adversarial Distillation, which recasts distribution matching as classification and learns the required log-density ratios directly. Two discriminator heads on a shared backbone distinguish real data and teacher samples from the student's, and linear losses on their logits train the student without auxiliary score fitting. We prove that at the discriminator optimum these losses recover the distribution-matching gradient underlying DMD, through the classical identity linking discriminator logits to log-density ratios. We further introduce gap-based reweighting, which adapts teacher supervision across noise levels from the real-data head's empirical logit gap between real and teacher samples. DMAD reaches a Fréchet Inception Distance (FID) of 1.04 with one-step generation on ImageNet-64x64, 14.47 with four-step SDXL on COCO-10K, and a VBench total score of 85.15 with four-step Wan2.1-T2V-14B, the best values among the compared few-step methods and the multi-step teachers. On MiniMax-H3-33B, our four-step student achieves overall human preference rates of 79.1% over DMD2 and 84.6% over rCM for joint audio-video generation, excluding ties. Our code, models and demos are available at https://yzmblog.github.io/projects/DMAD.

摘要:分佈匹配蒸餾(DMD)從分別估計的目標和學生分數之間的差異訓練出幾步的學生,因此它必須保持一個輔助擴散模型,以適應學生不斷變化的分佈,這需要額外的記憶體和計算成本。我們引入 DMAD,作為對抗蒸餾的分佈匹配,將分佈匹配重新表述為分類,並直接學習所需的對數密度比率。兩個共享主幹的鑑別器頭區分真實數據和教師樣本與學生的樣本,並通過它們的邏輯值進行線性損失訓練學生,而無需輔助分數擬合。我們證明,在鑑別器最佳點,這些損失恢復了 DMD 背後的分佈匹配梯度,這是通過經典的身份將鑑別器邏輯值與對數密度比率聯繫起來。我們進一步引入基於差距的重加權,這根據真實數據頭的真實樣本和教師樣本之間的經驗邏輯差距,調整教師監督在噪聲水平上的適應性。DMAD 在 ImageNet-64x64 上以一步生成達到 1.04 的 Fréchet Inception Distance(FID),在 COCO-10K 上以四步 SDXL 達到 14.47,在四步 Wan2.1-T2V-14B 上達到 85.15 的 VBench 總分,這是與比較的幾步方法和多步教師中的最佳值。在 MiniMax-H3-33B 上,我們的四步學生在聯合音頻-視頻生成中達到了 79.1% 的人類偏好率,超過 DMD2 和 84.6% 超過 rCM,排除平局。我們的代碼、模型和演示可在 https://yzmblog.github.io/projects/DMAD 獲得。

Higher-Order Molecular Grammars for Generative and Foundation Models in Chemistry

2610.02186v1 by Yiming Huang, Yujie Zeng, Vijay Prakash Dwivedi, Simone Foti, Jianmin Wang, Jure Leskovec, Tolga Birdal

Molecular learning models are strongly shaped by their underlying representations. Yet standard sequential and graph formalisms struggle to explicitly encode higher-order topology, such as ring systems and recurring motifs. Existing higher-order representations can capture these structures directly, but they are often computationally demanding and difficult to decode into valid molecules. Here, we introduce Higher-order Grammar Representation (HGR), a principled, topology-aware framework that lifts molecules to combinatorial complexes and parses each complex into a compact sequence of production rules under a context-free higher-order grammar. By serialising higher-order topology into rule sequences, HGR makes these structures directly compatible with standard sequence models, avoiding the computational overhead of explicit higher-order encodings while preserving topological expressiveness. To reduce benchmark bias towards simple ring systems, we construct RingDiv, a ring-enriched benchmark containing 1.18 million molecules, including the curated RingDiv300k subset, and introduce the ring diversity index (RDI) to quantify ring-system coverage. In molecular generation, HGR-based models uniquely combine 100% validity by construction with leading distributional alignment, ranking first in FCD on all five generation benchmarks. In representation learning, HGR-FM achieves the highest mean AUC across seven MoleculeNet benchmarks under both transfer protocols, improving on the strongest baseline by 8.3 and 3.3 AUC points under probing and full fine-tuning, respectively. Collectively, these results establish HGR as an efficient higher-order representation for molecular generation and transferable representation learning.

摘要:分子學習模型受到其基礎表示的強烈影響。然而,標準的序列和圖形形式在明確編碼高階拓撲方面(如環系統和重複圖案)面臨挑戰。現有的高階表示可以直接捕捉這些結構,但它們通常計算需求高且難以解碼為有效的分子。在此,我們介紹高階語法表示(HGR),這是一個原則性、關注拓撲的框架,將分子提升為組合複合體,並將每個複合體解析為在上下文無關的高階語法下的緊湊生成規則序列。通過將高階拓撲序列化為規則序列,HGR使這些結構與標準序列模型直接兼容,避免了明確高階編碼的計算開銷,同時保留了拓撲表達能力。為了減少對簡單環系統的基準偏見,我們構建了RingDiv,這是一個包含118萬個分子的環增強基準,包括精心策劃的RingDiv300k子集,並引入環多樣性指數(RDI)來量化環系統的覆蓋範圍。在分子生成方面,基於HGR的模型獨特地結合了100%的有效性(由構造決定)與領先的分佈對齊,在所有五個生成基準中FCD排名第一。在表示學習方面,HGR-FM在七個MoleculeNet基準中,在兩種轉移協議下實現了最高的平均AUC,相較於最強基線分別提高了8.3和3.3 AUC點(在探測和完全微調下)。綜合這些結果,HGR確立了作為分子生成和可轉移表示學習的高效高階表示。

SoftServe: A Scalable Quasi-Newton Method for Deep Learning

2610.02182v1 by Joohwan Ko, Tetiana Parshakova, Diana Cai, Robert M. Gower

Quasi-Newton (QN) methods have long been among the most effective methods for large-scale unconstrained convex optimization. Two obstacles have limited their use in deep learning: non-convexity and enormous parameter sizes. We introduce SoftServe, a family of QN methods designed to overcome these obstacles without line searches or ad hoc curvature corrections. SoftServe derives positivedefinite curvature estimates from the variational objective of Berglund et al. (2025), even in the presence of negative curvature. We develop diagonal and Kroneckerfactored variants that preserve positive definiteness by construction and scale to massive neural networks. Finally, SoftServe relies on the stable coupled Newton-Schulz iteration for the required matrix operations, replacing costly matrix decompositions with GPU-friendly matrix multiplications. SoftServe excels on problems that are severely ill-conditioned, including tasks such as recurrent networks, deep autoencoders, physics-informed neural networks, and a 136M-parameter physics-informed diffusion model, often achieving lower losses than established baselines including Adam, Muon, and SOAP.

摘要:準牛頓(QN)方法長期以來一直是大規模無約束凸優化中最有效的方法之一。兩個障礙限制了它們在深度學習中的應用:非凸性和巨大的參數規模。我們介紹了 SoftServe,一系列旨在克服這些障礙的 QN 方法,無需進行線搜索或臨時的曲率修正。SoftServe 從 Berglund 等人(2025)的變分目標中推導出正定的曲率估計,即使在存在負曲率的情況下也是如此。我們開發了對角和克羅內克分解變體,這些變體通過構造保持正定性並能擴展到大型神經網絡。最後,SoftServe 依賴於穩定的耦合牛頓-舒爾茨迭代來進行所需的矩陣操作,將成本高昂的矩陣分解替換為適合 GPU 的矩陣乘法。SoftServe 在極度病態的問題上表現出色,包括循環網絡、深度自編碼器、物理知識神經網絡以及一個 136M 參數的物理知識擴散模型等任務,通常實現比包括 Adam、Muon 和 SOAP 在內的既定基準更低的損失。

Generative Cinematographer: Composing Camera and Object Motion in 3D

2610.02180v1 by Jiahan Zhang, Chaohao Yang, Namitha Guruprasad, Vivekjyoti Banerjee, Trong-Tung Nguyen, Alan Yuille, Anand Bhattad

Current controllable video generation systems often rely on 2D motion trajectories or sparse drag signals for object motion. These controls are ambiguous because the same 2D trajectory can correspond to different 3D motions, especially when the camera and objects move simultaneously. We present Generative Cinematographer (GenCine), a system that lifts a single image into an editable 3D scene scaffold where artists jointly author camera and foreground motion. Artists specify a camera path and move selected foreground regions using local 3D motion handles. Several handles can move different parts of a subject independently, providing a piecewise-rigid approximation to non-rigid motion without a physics simulator or category-specific prior. To communicate these controls to a pretrained video model, we project them into guidance maps. These maps record where the controlled regions appear in each frame, assign each handle a fixed color across frames and encode the current 3D positions of its controlled points in the same world coordinate system as the background. This lets us describe object motion relative to the scene even as the camera moves. For training, we recover controls from the motion observed in real videos and use ground-truth geometry and trajectories from synthetic videos. We train a lightweight guidance branch and LoRA adapters on a pretrained Wan model to follow these controls. Our experiments show consistent camera-relative motion, improved geometric consistency under viewpoint changes, and strong controllability across diverse real-world scenes.

摘要:目前可控的視頻生成系統通常依賴於 2D 動作軌跡或稀疏的拖曳信號來控制物體運動。這些控制是模糊的,因為相同的 2D 軌跡可以對應於不同的 3D 動作,特別是在相機和物體同時移動的情況下。我們提出了生成電影製作人(GenCine),這是一個將單一圖像提升為可編輯的 3D 場景框架的系統,藝術家可以共同創作相機和前景運動。藝術家指定相機路徑並使用局部 3D 動作手柄移動選定的前景區域。幾個手柄可以獨立移動主體的不同部分,提供對非剛性運動的分段剛性近似,而無需物理模擬器或特定類別的先驗知識。為了將這些控制傳達給預訓練的視頻模型,我們將它們投影到指導圖中。這些圖記錄了受控區域在每幀中出現的位置,為每個手柄在幀之間分配固定顏色,並以與背景相同的世界坐標系編碼其受控點的當前 3D 位置。這使我們能夠描述相對於場景的物體運動,即使相機在移動。為了訓練,我們從真實視頻中觀察到的運動中恢復控制,並使用合成視頻中的真實幾何和軌跡。我們在預訓練的 Wan 模型上訓練了一個輕量級的指導分支和 LoRA 適配器,以遵循這些控制。我們的實驗顯示出一致的相機相對運動、在視點變化下改進的幾何一致性,以及在多樣的現實場景中強大的可控性。

Every Ablation Is a Dose: Counterweights and the Semblance of Self-Repair

2610.02173v1 by Areeb Ahmad, Pratinav Seth, Vinay Kumar Sankarapu

Ablate a component of a language model, and other components often appear to adjust and compensate. This phenomenon, termed self-repair, has been observed repeatedly, but its mechanism remains unclear. The most systematic study to date concluded that self-repair is noisy and unlikely to have a single explanation. We argue that it has one: a gain already present before any ablation. Any intervention on a causally important component can be viewed as a point on a coordinate axis $λ$, the signed strength of a counterfactual contrast. Hence, conventional ablation methods are uncalibrated points on this axis. We show that the causal repair response for a fine-grained unit $r$ is governed by an affine law, $E_r(λ)=\mathrm{own}_r+γ_rλ$. The slope $γ_r$ is a fixed coefficient that consistently influences the model, with or without ablation, and its sign determines whether the unit counteracts or reinforces the removed signal. On a factual-verdict task across four models from distinct families (Gemma, Qwen, LLaMA, and Mistral), we identify components including MLP neurons, OV neurons, and singular directions that follow this affine law, 68 of 81 downstream directions in all. Moreover, we can anticipate the magnitude of $γ_r$ from the fixed weights. On the IOI circuit of GPT-2 Small, seven of the ten heads the intervention can reach follow the law, and all seven are counterweights. From this perspective, what may appear as self-repair is a counterweight performing its usual operation when the contrastive signal emerges at the core.

摘要:去除語言模型的一個組件時,其他組件往往會調整並補償。這一現象被稱為自我修復,已經多次被觀察到,但其機制仍不清楚。迄今為止,最系統的研究得出結論,自我修復是嘈雜的,並且不太可能有單一的解釋。我們主張它有一個:在任何去除之前已經存在的增益。對於一個因果重要的組件的任何干預可以被視為坐標軸 $λ$ 上的一個點,即反事實對比的有向強度。因此,傳統的去除方法是在這個軸上的未校準點。我們顯示,對於一個細粒度單元 $r$ 的因果修復反應受一個仿射法則支配,$E_r(λ)=\mathrm{own}_r+γ_rλ$。斜率 $γ_r$ 是一個固定的係數,持續影響模型,無論是否去除,其符號決定了該單元是抵消還是增強被移除的信號。在四個不同家族的模型(Gemma、Qwen、LLaMA 和 Mistral)中的事實判決任務中,我們識別出包括 MLP 神經元、OV 神經元和遵循這一仿射法則的單一方向的組件,總共 81 個下游方向中有 68 個。此外,我們可以從固定權重預測 $γ_r$ 的大小。在 GPT-2 Small 的 IOI 電路中,十個頭中有七個可以到達的干預遵循這一法則,並且這七個都是對重。從這個角度來看,可能看起來像自我修復的現象實際上是一個對重在對比信號出現時執行其正常操作。

AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents

2610.02163v1 by Xuan Zhang, Longtao Zheng, Cunxiao Du, Bo An, Xin Dong

Coding agents solve repository-level software engineering tasks through long trajectories of code inspection, search, editing, and testing. As a task progresses, earlier exploration becomes stale, so managing context is more than avoiding overflow: an agent must decide when to compact, what working state to preserve, and how to continue from it. We introduce AutoCompact, which trains a coding agent to make these decisions as part of its policy. To collect training data, we run the base agent on coding tasks and use a judge to review its compaction decisions, summaries, and actions after compaction. Flawed outputs are replaced with corrected ones before being executed in the environment, so each trajectory continues from the corrected decisions. We use these trajectories for supervised fine-tuning, then jointly optimize coding and compaction through reinforcement learning with task-success rewards. Experiments on SWE-bench Verified and SWE-PolyBench Verified show that AutoCompact improves pass rates over the base model by an absolute 9.2\% and 5.0\%, respectively. The improvements hold across all evaluated inference budgets, with a 256K context window that never overflows and with a 16K window whose overflow triggers fallback compaction.

摘要:編碼代理透過長期的代碼檢查、搜索、編輯和測試來解決庫級軟體工程任務。隨著任務的進展,早期的探索變得過時,因此管理上下文不僅僅是避免溢出:代理必須決定何時壓縮、保留什麼工作狀態,以及如何從中繼續。我們介紹了AutoCompact,這是一個訓練編碼代理作為其策略一部分來做出這些決策的系統。為了收集訓練數據,我們在編碼任務上運行基礎代理,並使用評審來審查其壓縮決策、摘要和壓縮後的行動。缺陷輸出在執行環境中之前會被更正的輸出所替代,因此每個軌跡都從更正的決策繼續。我們使用這些軌跡進行監督微調,然後通過強化學習與任務成功獎勵共同優化編碼和壓縮。在SWE-bench Verified和SWE-PolyBench Verified上的實驗顯示,AutoCompact相對於基礎模型的通過率分別提高了9.2\%和5.0\%。這些改進在所有評估的推理預算中均保持有效,使用256K的上下文窗口從不溢出,並且使用16K窗口時其溢出會觸發回退壓縮。

DuoMind: Enabling Distributed Multi-Robot Coordination with Semantic Communication

2610.02161v1 by Hanchu Zhou, Dechen Gao, Hang Wang, Brendan Lynch, Boqi Zhao, Qiyao Ma, Raman Goyal, Junshan Zhang

Vision-language models (VLMs) and vision-language-action models (VLAs) have recently driven rapid progress in general-purpose robots, yet most progress has focused on single-robot settings. Extending these capabilities to multi-robot systems remains challenging because robots must coordinate long-horizon behaviors while maintaining reliable, fine-grained execution. We introduce DuoMind, a distributed hierarchical framework for multi-robot coordination through semantic communication. Each robot uses a VLA-based action model for low-level execution and a VLM-based orchestrator for high-level reasoning and inter-agent coordination. At each planning step, the orchestrator at each robot reasons over the task instruction, local observations, and messages received from other robots. It then generates low-level instructions for the action model and semantic messages for peer robots. This architecture exploits the complementary strengths of pretrained models by combining the semantic reasoning capabilities of VLMs with the precise action-generation capabilities of VLAs. To address the scarcity of benchmarks for multi-robot coordination, we further develop RoboPoly, a benchmark comprising long-horizon manipulation tasks that require coordinated, closed-loop execution under distributed control. Experiments on RoboPoly and RoboTwin demonstrate that DuoMind improves multi-robot task performance, while ablation studies confirm the contributions of hierarchical orchestration and semantic communication. More details are available on our project page.

摘要:視覺語言模型(VLMs)和視覺語言行動模型(VLAs)最近在通用機器人領域推動了快速進展,然而大多數進展集中在單機器人環境中。將這些能力擴展到多機器人系統仍然具有挑戰性,因為機器人必須協調長期行為,同時保持可靠且精確的執行。我們介紹了DuoMind,一個通過語義通信實現多機器人協調的分散式層次框架。每個機器人使用基於VLA的行動模型進行低層次執行,並使用基於VLM的協調者進行高層次推理和代理間協調。在每個規劃步驟中,每個機器人的協調者會對任務指令、當地觀察和來自其他機器人的消息進行推理。然後,它生成行動模型的低層次指令和針對同儕機器人的語義消息。這種架構通過結合VLM的語義推理能力與VLA的精確行動生成能力,充分利用了預訓練模型的互補優勢。為了解決多機器人協調基準的稀缺性,我們進一步開發了RoboPoly,一個包含需要協調、閉環執行的長期操作任務的基準。在RoboPoly和RoboTwin上的實驗表明,DuoMind改善了多機器人任務的表現,而消融研究確認了層次協調和語義通信的貢獻。更多細節可在我們的項目頁面上獲得。

From Knowledge Access to Source Learning: Developing Source-Specific Competence

2610.02150v1 by Lucheng Fu, Kejing Xia, Yiyang Wang, Yiqiao Jin, Jinjin He, Xiyuan Yang, Haoxin Liu, Ye Yu, Haibo Jin, Yijia Xiao, Wenke Lee, B. Aditya Prakash, Haohan Wang

Large language model (LLM) agents increasingly rely on persistent external sources to solve sequences of knowledge-intensive tasks. Existing methods improve how source content is accessed and organized, while agent-memory systems preserve reusable knowledge from prior interactions, but repeated use of the same source is still largely treated as repeated access rather than an opportunity to progressively improve understanding of that source. We study source learning: developing reusable source-specific competence over a persistent authoritative source. We represent this competence with a persistent source model that captures reusable understanding of the source, including how its knowledge is structured, interpreted, and applied. To construct and progressively refine such models, we propose SourceLearn, which combines two complementary learning mechanisms. Self-Directed Source Learning identifies what remains incompletely understood and adaptively revisits the source, while Task-Guided Source Learning uses downstream experience to reveal local representational gaps and recurring needs in how source knowledge should be organized. In both cases, learning signals determine what should be reconsidered, while persistent updates are reconstructed from the authoritative source. Across five benchmarks and three LLM backends, SourceLearn achieves the best performance in 13 of 15 settings, with gains of up to 22.6 points over Hybrid RAG and substantial overall improvements over static source representations and experience-based memory baselines.

摘要:大型語言模型(LLM)代理越來越依賴持久的外部來源來解決一系列知識密集型任務。現有方法改善了如何訪問和組織來源內容,而代理記憶系統則保留了來自先前互動的可重用知識,但對同一來源的重複使用仍然主要被視為重複訪問,而不是逐步改善對該來源理解的機會。我們研究來源學習:在持久的權威來源上發展可重用的來源特定能力。我們用一個持久的來源模型來表示這種能力,該模型捕捉了對來源的可重用理解,包括其知識的結構、解釋和應用方式。為了構建和逐步完善這樣的模型,我們提出了SourceLearn,該模型結合了兩種互補的學習機制。自我導向來源學習識別尚未完全理解的內容並適應性地重新訪問來源,而任務引導來源學習則利用下游經驗揭示如何組織來源知識的局部表徵差距和重複需求。在這兩種情況下,學習信號決定了應該重新考慮的內容,而持久更新則是從權威來源重建的。在五個基準和三個LLM後端中,SourceLearn在15個設置中的13個中實現了最佳性能,與Hybrid RAG相比,增益高達22.6點,並且在靜態來源表示和基於經驗的記憶基準上有顯著的整體改進。

Keyword Harnesses Fail Open: A Cheap Diagnostic Ladder for Tool-Use Claims in Small Language Models

2610.02142v1 by Juan S. Santillana

Keyword-matching benchmarks can credit small models for tool use they never perform. We document such a false positive in a matched-architecture pair of Spanish security language models and propose a ladder of strict, cheap diagnostics. A 661.6M parameter model (approx. 65% code/technical text; no dedicated SFT) and a 1,109M model (web-heavy multi-phase curriculum; 6B-token tool-SFT) share decoder, tokenizer, and special tokens, scoring almost identically on lenient tool-use metrics (B4: 0.660 vs. 0.650). Verbatim-reproduction checks on training examples separate them completely: the 600M emits valid tool calls with generalized arguments on 6/6 examples; the 1B does so on 0/6 across checkpoints. A first-token probe localizes the 1B's failure to a missing prior (prob. $10^{-4}$--$10^{-5}$ on <|tool_call|>), which was erased by its web-heavy training phase. A targeted SFT recipe (diverse corpus, 5x higher learning rate, 2,202 steps, ~3.3 GPU-hours) repairs the 1B using three orders of magnitude fewer tokens than the failed phase. On all 269 corpus rows, valid emission rises from 0.100 to 0.959 (600M: 0.926). On 238 unseen prompts, the repaired 1B passes 0.536 vs. the 600M's 0.428 ($p = 0.004$). Embedding-drift checks show the repair did not move the trigger token's tied embedding (97.7% of the bf16 table remains bit-identical), meaning changes live in the surrounding network. Both models over-trigger, rarely answering negative prompts without a call (0.09 for 600M, 0.17 for repaired 1B). Factorial analyses confirm all repair configurations install the format, though suppression benefits from a diverse corpus remain a hypothesis due to seed sensitivity. This cheap diagnostic ladder costs minutes of CPU time and should gate tool-use claims on small models.

摘要:關鍵字匹配基準可以錯誤地將小型模型的工具使用歸功於它們從未執行過的行為。我們在一對匹配架構的西班牙安全語言模型中記錄了這樣的假陽性,並提出了一個嚴格且便宜的診斷階梯。一個661.6M參數的模型(約65%的代碼/技術文本;沒有專門的SFT)和一個1,109M的模型(以網絡為主的多階段課程;6B-token工具SFT)共享解碼器、分詞器和特殊標記,在寬鬆的工具使用指標上幾乎得分相同(B4: 0.660對0.650)。對訓練範例的逐字重現檢查完全將它們分開:600M在6/6個範例中發出有效的工具調用,並帶有一般化的參數;而1B在各檢查點中則在0/6中發出。首次標記探測將1B的失敗定位於缺失的先前狀態(概率$10^{-4}$--$10^{-5}$在<|tool_call|>上),該狀態在其以網絡為主的訓練階段中被抹去。一個針對性的SFT配方(多樣化語料庫、5倍更高的學習率、2,202步驟、約3.3 GPU小時)使用比失敗階段少三個數量級的標記修復了1B。在所有269個語料行中,有效發射率從0.100上升至0.959(600M: 0.926)。在238個未見的提示中,修復後的1B通過了0.536,而600M則為0.428($p = 0.004$)。嵌入漂移檢查顯示修復並未移動觸發標記的綁定嵌入(97.7%的bf16表保持位元相同),這意味著變化存在於周圍的網絡中。兩個模型都過度觸發,幾乎不在沒有調用的情況下回答負面提示(600M為0.09,修復後的1B為0.17)。因子分析確認所有修復配置都安裝了格式,儘管由於種子敏感性,來自多樣化語料庫的抑制效益仍然是一個假設。這個便宜的診斷階梯僅需幾分鐘的CPU時間,應該限制小型模型的工具使用聲明。

Finetuning with Sampling: SFT Learns Better Than You Think

2610.02140v1 by Aayush Karan, Sitan Chen, Yilun Du

Introducing new capabilities to frontier models has long been the goal of posttraining, which predominantly employs supervised finetuning (SFT) and reinforcement learning (RL) to this end. Conventional wisdom dictates that RL enables strong generalization on new tasks without losing existing capabilities, while SFT is prone to weak generalization and catastrophic forgetting. At the same time, SFT can learn from off-policy expert data, whereas RL must rely on a model's ability to find successful trajectories with repeated sampling. In our work, we seek to leverage the strength of on-policy learning while utilizing the privileged information contained in off-policy data. However, rather than modifying the learning objective to accommodate this data, we instead tailor the data distribution to better suit the learner. We introduce a Markov chain Monte Carlo (MCMC) sampling algorithm that progressively transforms off-policy traces to be more on-policy given a reference model for finetuning. Across tasks like scientific skill acquisition, mathematical reasoning, and open-ended expertise, our sampling algorithm enables SFT to rival prevailing posttraining techniques, often generalizing better and forgetting less than strong on-policy baselines. In addition, the resulting finetuned models exhibit strong distributional performance and are capable of learning beyond sharpening the base model distribution. At a higher level, our approach presents sampling as a model-native operator that shapes data for learnability, offering broader utility as a general-purpose primitive throughout the posttraining stack.

摘要:引入新能力到前沿模型一直是後訓練的目標,這主要通過監督微調(SFT)和強化學習(RL)來實現。傳統智慧認為,RL 能夠在不失去現有能力的情況下,對新任務進行強泛化,而 SFT 則容易導致弱泛化和災難性遺忘。與此同時,SFT 可以從離策略專家數據中學習,而 RL 必須依賴模型的能力來通過重複採樣找到成功的軌跡。在我們的工作中,我們尋求利用在線學習的優勢,同時利用離策略數據中包含的特權信息。然而,我們並不是修改學習目標以適應這些數據,而是調整數據分佈以更好地適應學習者。我們引入了一種馬爾可夫鏈蒙特卡羅(MCMC)採樣算法,該算法逐步將離策略痕跡轉變為更符合在線策略的形式,前提是有一個參考模型進行微調。在科學技能獲得、數學推理和開放式專業知識等任務中,我們的採樣算法使得 SFT 能夠與當前的後訓練技術相抗衡,通常在泛化能力上表現更好,且遺忘程度較低於強在線基準。此外,最終微調的模型展現出強大的分佈性能,並能夠學習超越僅僅是加強基礎模型分佈。在更高的層面上,我們的方法將採樣呈現為一種模型原生操作符,為學習性塑造數據,並在整個後訓練堆棧中提供更廣泛的通用性。

MIRTO: a registration-gated, multiverse-tested evaluation protocol for unsupervised anomaly segmentation in brain MRI

2610.02136v1 by Negin Kafee Hernashki, Soumick Chatterjee

Unsupervised anomaly detection (UAD) methods for brain MRI are ranked by a single score, yet that score rests on choices that are rarely reported: how each anomaly map is aligned with the reference, how and on which data the threshold is set, and which false-positive budget, metric, aggregation and lesion definition are used. We present MIRTO, an evaluation protocol that makes these choices explicit and measures their effect. It gates the geometry of every comparison with a registration check and label-free diagnostics of known power, sets thresholds on validation data alone and reports the false-positive volume actually realised on test, repeats each comparison over 15,552 defensible evaluation pipelines, and attaches paired subject-bootstrap intervals with multiplicity control. Applied to four UAD methods trained on the same healthy data and tested on 312 BraTS 2020 subjects, MIRTO showed that an axis-order mismatch between stored maps and the reference lowered a diffusion model's voxel AUROC from 0.873 to 0.583 whilst barely moving its slice-level AUROC. Within each metric, the method explained at least 0.95 of the variance in voxel AUROC and AUPRC and 0.77 in Dice, but only 0.14 in lesion sensitivity, where the lesion definition and hit criterion dominated. A Dice advantage that was significant at validation thresholds vanished at equal realised false-positive burden, and an exact identity attributes it to threshold transfer. A training-free change to REFLECT's latent aggregation raised Dice at equal burden by 0.052. Nine hypotheses were tested against explicit criteria; because the same cohort served to develop the protocol, all inference is exploratory.

摘要:未監督異常檢測(UAD)方法對於腦部 MRI 的排名是基於單一分數,但該分數依賴於鮮少報告的選擇:每個異常圖與參考的對齊方式、如何以及基於哪些數據設置閾值,以及使用哪種假陽性預算、指標、聚合和病變定義。我們提出了 MIRTO,一種評估協議,使這些選擇變得明確並測量其影響。它通過註冊檢查和無標籤診斷已知功率來限制每次比較的幾何,僅在驗證數據上設置閾值,並報告在測試中實際實現的假陽性體積,重複每次比較超過 15,552 條可辯護的評估管道,並附上成對的主體自助間隔及多重性控制。應用於四種基於相同健康數據訓練並在 312 名 BraTS 2020 受試者上測試的 UAD 方法,MIRTO 顯示存儲圖與參考之間的軸序不匹配使擴散模型的體素 AUROC 從 0.873 降至 0.583,同時幾乎不影響其切片級 AUROC。在每個指標中,該方法解釋了至少 0.95 的體素 AUROC 和 AUPRC 的變異,及 0.77 的 Dice,但在病變敏感性中僅為 0.14,病變定義和命中標準主導了這一結果。在驗證閾值下顯著的 Dice 優勢在相等的實現假陽性負擔時消失,並且一個精確的身份將其歸因於閾值轉移。對 REFLECT 的潛在聚合進行無訓練的變更,在相等負擔下將 Dice 提高了 0.052。針對明確標準測試了九個假設;由於相同的隊列用於開發該協議,所有推斷都是探索性的。

Local Support Learning

2610.02126v1 by Assaf Ben-Kish, Akarsh Kumar, James Glass, Raja Giryes

We explore catastrophic forgetting in the context of large pre-trained models. By considering forgetting as a geometric problem in the input space of each weight matrix, we uncover a natural retention objective under which updates produced by gradient-based optimizers are suboptimal. Following this observation, we propose Local Support Learning (LSL), a general-purpose framework that augments gradient-based training for retention of prior capabilities without access to prior data. During a new learning phase, LSL pairs two components with distinct roles: a standard weight adapter, trained as usual to minimize the loss, and a gating function that enables the adapter only on input activations from its own training distribution, making the update local to that distribution. The key challenge is that this gate must route data from all learning phases while training only on data from the current one. We address this with a gate based on a Gaussian Mixture Model (GMM), whose likelihood decays rapidly away from its training data, giving it a natural tendency to stay closed on data from prior phases. We show that this post-training approach can resolve forgetting in LLMs of up to 7 billion parameters, retaining both pretrained and finetuned capabilities across multiple training phases, while being efficient in memory and compute, robust to hyperparameter choice, and showing scaling potential.

摘要:我們探討在大型預訓練模型背景下的災難性遺忘。通過將遺忘視為每個權重矩陣的輸入空間中的幾何問題,我們揭示了一個自然的保留目標,在該目標下,由梯度基優化器產生的更新是次優的。根據這一觀察,我們提出了本地支持學習(Local Support Learning,LSL),這是一個通用框架,增強了基於梯度的訓練,以保留先前的能力,而無需訪問先前數據。在新的學習階段,LSL 配對了兩個具有不同角色的組件:一個標準的權重適配器,像往常一樣訓練以最小化損失,以及一個閘控函數,僅在來自其自身訓練分佈的輸入激活上啟用適配器,使更新局限於該分佈。關鍵挑戰在於,這個閘必須在訓練僅基於當前數據的同時,路由來自所有學習階段的數據。我們通過基於高斯混合模型(Gaussian Mixture Model,GMM)的閘來解決這一問題,其似然在遠離其訓練數據時迅速衰減,這使其對來自先前階段的數據自然傾向於保持關閉。我們展示了這種後訓練方法可以解決高達70億參數的大型語言模型中的遺忘,保留了多個訓練階段中的預訓練和微調能力,同時在內存和計算上高效,對超參數選擇具有穩健性,並顯示出擴展潛力。

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

2610.02122v1 by Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma

Real-world enterprise data science and analytics workflows require reasoning across dozens of tables, performing statistical analyses, and acting on the results. Established text-to-SQL benchmarks evaluate query generation alone, and audits have found their answer keys frequently wrong. Because real enterprise warehouses are too sensitive to release, these benchmarks are built on public datasets where a business event fits in a single table. We introduce Argo-Bench, an evaluation framework comprising 210 data science and analytics tasks. Drawing on public data, peer-reviewed industry literature, and regulatory filings, we simulate a food delivery platform in New York City at true scale, with 81 million orders in 2024, grounded economics, fraud patterns, and marketplace incentives. We export this world to an ERP warehouse of 235 tables and 7.5 billion rows, modeled on the Oracle E-Business Suite schema. The simulator's ground-truth state is withheld from the warehouse the agent sees, so tasks require reconstructing facts by navigating the warehouse before acting on them. Argo-Bench goes beyond text-to-SQL: the agent files actions such as banning fraudulent accounts, allocating courier incentive budgets, or issuing back pay, and the grader scores each by its consequences in the simulator. Every task has an executable reference solution that demonstrates solvability using only the warehouse. The strongest of 14 frontier and open-weight models scores 95 or higher on only 34.8% of tasks and averages 59.5 points. We hope Argo-Bench drives progress toward agents that understand, navigate, and act within real data environments.

摘要:現實世界的企業數據科學和分析工作流程需要跨越數十個表格進行推理、執行統計分析並根據結果採取行動。現有的文本到SQL基準僅評估查詢生成,而審計發現其答案鍵經常錯誤。由於真正的企業數據倉庫過於敏感而無法釋放,這些基準是基於公共數據集構建的,其中商業事件適合於單個表格。我們介紹了Argo-Bench,一個包含210個數據科學和分析任務的評估框架。基於公共數據、同行評審的行業文獻和監管文件,我們模擬了一個在紐約市的食品配送平台,真實規模為2024年的8100萬個訂單,並考慮了經濟基礎、詐騙模式和市場激勵。我們將這個世界導出到一個擁有235個表格和75億行的ERP數據倉庫,該倉庫基於Oracle E-Business Suite架構進行建模。模擬器的真實狀態對代理所見的數據倉庫是保密的,因此任務需要通過導航數據倉庫來重建事實,然後再對其進行操作。Argo-Bench超越了文本到SQL:代理執行的行動包括禁止詐騙帳戶、分配快遞員激勵預算或發放補發工資,評分者根據模擬器中的結果對每個行動進行評分。每個任務都有一個可執行的參考解決方案,該解決方案僅使用數據倉庫來展示可解性。14個前沿和開放權重模型中最強的模型在僅34.8%的任務上得分95或更高,平均得分為59.5分。我們希望Argo-Bench能推動代理理解、導航並在真實數據環境中行動的進展。

Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes

2610.02117v1 by Sophia Sirko-Galouchenko, Monika Wysoczanska, Andrei Bursuc, Nicolas Thome, Spyros Gidaris

On-policy self-distillation has recently emerged as an effective approach for improving language-model reasoning by supervising students with a frozen or EMA version of themselves that receives privileged information. Its application to multimodal large language models (MLLMs), however, remains largely unexplored. Recent approaches use privileged visual information, such as image crops corresponding to a question, to improve fine-grained perception, but their gains are confined to tasks that benefit from such visual zooming and require either human-annotated grounding data or external teacher models. We introduce a different form of on-policy self-distillation for MLLMs that provides the teacher with textual, spatially grounded guidance identifying the visual elements relevant to a query. We use procedurally generated scenes with automatically available object identities and spatial coordinates, enabling scalable and annotation-free post-training. The teacher uses this spatial guidance to locate and integrate evidence from multiple relevant image regions, while the student learns to reproduce the resulting behavior from the image and question alone. Our approach consistently improves performance on counting, document and chart understanding benchmarks across multiple models. Importantly, although post-training uses only synthetic scenes, the resulting improvements transfer to real-world perception benchmarks, yielding a 3.23-point gain in average performance across CVBench, V*, ZoomBench, BLINK, HR-Bench, and MME-RealWorld. These results show that spatially grounded privileged information can induce broader perceptual capabilities through on-policy self-distillation, enabling substantial synthetic-to-real transfer beyond the task and data distribution used for post-training. Project page: https://github.com/sirkosophia/Where-OPD

摘要:在政策自蒸餾最近出現,作為一種有效的方法來改善語言模型推理,通過用一個凍結或EMA版本的自己來監督學生,該版本接收特權信息。然而,它在多模態大型語言模型(MLLMs)上的應用仍然在很大程度上未被探索。最近的方法使用特權的視覺信息,例如與問題相對應的圖像裁剪,以改善細粒度的感知,但它們的增益僅限於那些受益於這種視覺放大並需要人類標註的基礎數據或外部教師模型的任務。我們為MLLMs引入了一種不同形式的政策自蒸餾,為教師提供文本的、空間上有根據的指導,以識別與查詢相關的視覺元素。我們使用程序生成的場景,具有自動可用的物體身份和空間坐標,實現可擴展且無需標註的後訓練。教師使用這種空間指導來定位並整合來自多個相關圖像區域的證據,而學生則學習僅從圖像和問題中重現結果行為。我們的方法在計數、文檔和圖表理解基準上,始終提高多個模型的性能。重要的是,儘管後訓練僅使用合成場景,但所產生的改進能夠轉移到現實世界的感知基準上,在CVBench、V*、ZoomBench、BLINK、HR-Bench和MME-RealWorld上實現了平均性能提高3.23點。這些結果顯示,空間上有根據的特權信息可以通過政策自蒸餾引發更廣泛的感知能力,使得合成到現實的轉移超越了用於後訓練的任務和數據分佈。項目頁面:https://github.com/sirkosophia/Where-OPD

A Comparative Explainability Framework for DeBERTa-v3 in Zero-Shot Medical Abstract Classification

2610.02116v1 by Javier Diaz Esteban-Herreros, David Muñoz-Valero, Raquel Martínez-España, Jose M. Juarez, Juan Moreno-Garcia

A comparative explainability framework is presented to audit DeBERTa-v3 under zero-shot classification of medical abstracts. The work addresses the disagreement problem in Explainable Artificial Intelligence, where different attribution methods produce divergent explanations for the same input and prediction. A natural language inference engine is implemented over the Medical Abstracts corpus with five enriched hypotheses per diagnostic category and a balanced sample of one thousand texts per class. Five explanation methods are compared: SHAP and LIME as model-agnostic approaches, occlusion and Input x Gradient as deep-learning-specific approaches, and Attention x Gradient as a transformer-specific approach. Explanations are standardized through top-token attribution, and pairwise agreement is quantified using the Jaccard index. High predictive accuracy is achieved across well-defined clinical domains, whereas performance degrades under high semantic ambiguity. Explanatory stability directly mirrors predictive certainty, exhibiting strong convergence in univalent categories and a marked drop under diagnostic uncertainty. Furthermore, qualitative error auditing uncovers three systemic failure mechanisms: lexical hypersensitivity, semantic overlap, and loss of attribution coherence. The results support the combined use of several explanation methods and quantitative agreement metrics when auditing transformer-based models in medical text classification, and suggest prioritizing specific clinical ontologies over broad diagnostic labels.

摘要:比較可解釋性框架被提出以審計 DeBERTa-v3 在醫學摘要的零樣本分類下。這項工作解決了可解釋人工智慧中的不一致問題,即不同的歸因方法對相同的輸入和預測產生不同的解釋。自然語言推理引擎在醫學摘要語料庫上實施,每個診斷類別有五個增強的假設,並且每個類別有一千篇文本的平衡樣本。比較了五種解釋方法:SHAP 和 LIME 作為模型無關的方法,遮蔽和輸入 x 梯度作為深度學習特定的方法,以及注意力 x 梯度作為Transformer特定的方法。通過頂部標記歸因標準化解釋,並使用 Jaccard 指數量化成對一致性。在明確定義的臨床領域中實現了高預測準確性,而在高語義模糊性下性能下降。解釋穩定性直接反映預測確定性,在單值類別中顯示出強烈的收斂,並在診斷不確定性下顯著下降。此外,定性錯誤審計揭示了三種系統性失效機制:詞彙過敏、語義重疊和歸因一致性的喪失。結果支持在醫學文本分類中審計基於Transformer的模型時,結合使用幾種解釋方法和定量一致性指標,並建議優先考慮特定的臨床本體論而非廣泛的診斷標籤。

Scalable, Transferable Meta-network for Data Selection Requires a Different Loss (and Why the Obvious Choice is Problematic)

2610.02092v1 by Zilin Du, Bowen Yang, Boyang Albert Li

Data selection is critical for training large language models on massive and heterogeneous corpora. Meta-learning for Training-data Selection offers a principled alternative to heuristic scoring by learning data weights from a target validation objective, but existing methods face a trade-off between fine-grained valuation and transferability to unseen data. A natural solution is to replace per-sample weights with a selection network. However, we find that directly incorporating such a network into existing MTS objectives leads to unstable optimization and poor generalization, caused by weight suppression and persistent reliance on easy-to-learn features. To address these issues, we propose Transferable Example Scoring and Selection (TESS), a scalable data-selection framework built on a Pointwise Value Matching objective (PVM). Experiments on LLM safety and targeted instruction tuning demonstrate strong transfer across datasets, from subsets to full corpora, and from smaller to larger models.

摘要:資料選擇對於在龐大且異質的語料庫上訓練大型語言模型至關重要。針對訓練數據選擇的元學習提供了一種基於原則的替代方案,通過從目標驗證目標學習數據權重,但現有方法在細緻評價和對未見數據的可轉移性之間面臨權衡。一個自然的解決方案是用選擇網絡來替代每個樣本的權重。然而,我們發現將這樣的網絡直接納入現有的MTS目標會導致不穩定的優化和較差的泛化,這是由於權重抑制和持續依賴易於學習的特徵所造成的。為了解決這些問題,我們提出了可轉移範例評分和選擇(TESS),這是一個基於逐點價值匹配目標(PVM)的可擴展數據選擇框架。在LLM安全性和針對性指令調整的實驗中,顯示出跨數據集的強大轉移能力,從子集到完整語料庫,從較小的模型到較大的模型。

GeoLatent: Geometry-Guided Latent Structuring with Routed Optimization for 3D Reasoning

2610.02091v1 by Yakun Zhu, Yi Bin, Yujuan Ding, Zheng Wang, Pengpeng Zeng, Duo Peng, Jingkuan Song, Heng Tao Shen

Despite progress in vision-language models, 3D spatial reasoning from 2D images remains challenging. Text-based methods describe intermediate geometry with discrete tokens, limiting fidelity for continuous spatial relations. Continuous latents offer richer representations, but a single latent type does not explicitly separate the cues needed across spatial tasks. Decomposed spatial latents address this by representing position, direction, and global geometry separately under geometric supervision. Yet the geometry representation can still collapse toward one dominant direction, and unrestricted attention can leave the latents underused during answer learning. We introduce GeoLatent, combining Common--Residual Geometry Alignment (CR-GEO) with routed optimization to structure the geometry states while promoting latent-mediated answer learning. CR-GEO separates shared from residual teacher geometry; routed optimization jointly trains geometry and language, temporarily directs visual answer learning through the latents, and restores full attention with geometry supervision. In controlled comparisons, CR-GEO raises geometry effective rank from 1.00 to 3.87, while blocking latent readout at the bottleneck lowers direction accuracy from 89.1% to 25.8% on 128 fixed questions. After recovery, the differentiated geometry representation and latent-mediated visual route remain available alongside direct image access. GeoLatent achieves 73.0% on SPAR-Bench and 72.1% on SPBench, outperforming previously reported methods on both.

摘要:儘管在視覺-語言模型方面取得了進展,從2D圖像進行3D空間推理仍然具有挑戰性。基於文本的方法使用離散標記描述中間幾何,限制了對連續空間關係的忠實度。連續潛變量提供了更豐富的表徵,但單一潛變量類型並未明確分離在空間任務中所需的線索。分解的空間潛變量通過在幾何監督下分別表示位置、方向和全局幾何來解決這個問題。然而,幾何表徵仍然可能向一個主導方向收斂,而不受限制的注意力可能在答案學習過程中使潛變量未被充分利用。我們引入了GeoLatent,結合了共同-殘差幾何對齊(CR-GEO)與路由優化,以結構化幾何狀態,同時促進潛變量介導的答案學習。CR-GEO將共享的幾何與殘差教師幾何分開;路由優化共同訓練幾何和語言,暫時通過潛變量引導視覺答案學習,並在幾何監督下恢復完整的注意力。在受控比較中,CR-GEO將幾何有效排名從1.00提高到3.87,而在瓶頸處阻止潛變量讀出則使方向準確率從89.1%降低到25.8%,針對128個固定問題。恢復後,區分的幾何表徵和潛變量介導的視覺路徑仍然可用,並與直接圖像訪問並存。GeoLatent在SPAR-Bench上達到73.0%,在SPBench上達到72.1%,在兩者上均超越了先前報告的方法。

LLM2Jev: LLMs Are Already Jev-Style Decision Models -- When and How to Fine-Tune Them

2610.02076v1 by Yinheng Li, Justin Wagle

Jev-style decision models return categorical probability distributions over predefined options without generating free-form text, enabling software systems to act on their outputs directly. In this work, we investigate the extent to which general-purpose LLMs already possess this capability out of the box, and when fine-tuning is actually necessary. We present LLM2Jev, an architecture-preserving framework that extracts calibrated decisions directly from next-token probabilities over bracketed numeric identifiers. LLM2Jev provides both a training-free inference recipe and a fine-tuning objective that optimizes candidate selection via a tree-factorized listwise loss while anchoring auxiliary predictions to the base model using KL divergence penalties. Evaluating on Qwen3.5-4B and Qwen3-0.6B, we find that modern LLMs are inherently effective decision models: without training, the 4B model matches community Jev-style models built on the same backbone, outperforms letter-logit readouts, supports arbitrary option counts, and natively handles multimodal decisions over images. Fine-tuning provides targeted rather than universal benefits -- substantially improving weaker models and specific tasks (such as many-option intent routing), but offering diminishing returns for strong backbones. Crucially, our KL anchors prevent behavioral degradation in conversational text generation, with LoRA delivering the strongest performance on capable models.

摘要:Jev風格的決策模型返回預定選項的類別概率分佈,而不生成自由格式的文本,使得軟體系統能夠直接根據其輸出進行操作。在這項工作中,我們調查通用LLM在開箱即用的情況下已經具備這種能力的程度,以及何時實際需要進行微調。我們提出了LLM2Jev,一個保留架構的框架,直接從括號中的數字標識符的下一個標記概率中提取經過校準的決策。LLM2Jev提供了一個無需訓練的推理配方和一個微調目標,通過樹狀分解的列表損失優化候選選擇,同時使用KL散度懲罰將輔助預測固定到基礎模型。在Qwen3.5-4B和Qwen3-0.6B上進行評估,我們發現現代LLM本質上是有效的決策模型:在未經訓練的情況下,4B模型與基於相同骨幹構建的社區Jev風格模型相匹配,超越字母邏輯讀出,支持任意選項數量,並原生處理圖像上的多模態決策。微調提供的是針對性的而非普遍的好處——顯著改善較弱的模型和特定任務(例如多選意圖路由),但對強大的骨幹則提供遞減的回報。至關重要的是,我們的KL錨點防止了對話文本生成中的行為退化,LoRA在有能力的模型上提供了最強的性能。

Causal Memory Policy: Making Memory Utility Identifiable by Intervening on Retrieval

2610.02070v1 by Arman Behnam, Binghui Wang

Memory-augmented large language models must decide which memories to retain, and recent systems do so by estimating each memory's effect on task performance. However, these estimates rely entirely on retrieved memories. When a memory is never retrieved, store-level interventions produce identical outcomes, leaving its utility unidentified. This is a retrieval-level positivity violation, invisible to diagnostics that examine only memory operations. We introduce Causal Memory Policy (CMP), a causal framework that restores identification by intervening on retrieval itself, reserving a fixed number of context slots for memories sampled with known propensities. CMP estimates memory utility by self-normalized inverse propensity weighting under a balanced assignment design. We prove the causal factorization of memory utility through retrieval, the unbiasedness and exact variance of the estimator, and the optimal decision rule under irreversible operations. Empirically, identification fails for 54% of required memories on LongMemEval and 67% on LoCoMo, and the failure persists in a deployed memory system. CMP improves discrimination between required and non-required memories from 0.54 to 0.66 AUC. Finally, we show that identified memory utility alone is insufficient for retention decisions: per-query utility reaches 0.78 AUC on the query for which it is estimated, yet no aggregation available to a retention policy predicts a memory's value on unseen queries. Code is available at: https://anonymous.4open.science/r/cmp-release-D0C3/.

摘要:記憶增強的大型語言模型必須決定保留哪些記憶,而最近的系統通過估計每個記憶對任務表現的影響來實現這一點。然而,這些估計完全依賴於檢索到的記憶。當一個記憶從未被檢索時,存儲層級的干預會產生相同的結果,使其效用無法識別。這是一種檢索層級的正向性違反,對於僅檢查記憶操作的診斷來說是不可見的。我們引入了因果記憶策略(Causal Memory Policy, CMP),這是一個因果框架,通過對檢索本身進行干預來恢復識別,為以已知傾向抽樣的記憶保留固定數量的上下文槽位。CMP 通過自我正規化的逆傾向加權來估計記憶效用,並在平衡分配設計下進行。我們證明了通過檢索的記憶效用的因果分解、估計量的無偏性和精確方差,以及在不可逆操作下的最佳決策規則。實證上,在 LongMemEval 上所需記憶的識別失敗率為 54%,而在 LoCoMo 上則為 67%,而且這種失敗在部署的記憶系統中持續存在。CMP 將所需記憶和非所需記憶之間的區分從 0.54 提高到 0.66 AUC。最後,我們顯示僅僅識別的記憶效用對於保留決策是不夠的:每查詢的效用在其估計的查詢上達到 0.78 AUC,然而對於未見查詢,保留策略無法預測記憶的價值。代碼可在以下網址獲得:https://anonymous.4open.science/r/cmp-release-D0C3/.

External Observers May See More Clearly: Cross-Model Span-Level Hallucination Detection in Large Language Models via Hidden State Probing

2610.02066v1 by Kingshuk Gupta, Davide Buscaldi

As Large Language Models (LLMs) increasingly serve as foundational reasoning engines, their tendency to hallucinate remains a critical vulnerability. While recent internal state probes offer a promising alternative to slow external retrieval systems, they largely reduce hallucination detection to a token-wise binary classification task, failing to capture the structured, sequential boundaries of semantic drift. Here, we introduce an internal hidden state framework for fine-grained, span-level hallucination detection. By inspecting layer-wise activation patterns, we attempt to detect the exact hallucination onset and continuation tokens in an LLM generation. Our experiments show that this approach successfully isolates hallucination onsets, achieving substantial improvements in Precision-Recall AUC over random baselines despite extreme class imbalance. Ultimately, we propose a novel cross-model detection framework in which one model observes the internal representations elicited by another model's generation. We find that an external observer can match or exceed a generator's self-detection of its own hallucination onsets, including when the observer is the smaller model, suggesting that self-detection is not the ceiling for onset localisation.

摘要:隨著大型語言模型(LLMs)越來越多地作為基礎推理引擎,它們的幻覺傾向仍然是一個關鍵的脆弱性。雖然最近的內部狀態探測提供了一種有前景的替代方案來取代緩慢的外部檢索系統,但它們在很大程度上將幻覺檢測簡化為一個逐字的二元分類任務,未能捕捉到語義漂移的結構性、序列性邊界。在此,我們介紹了一個內部隱藏狀態框架,用於細粒度的跨度級幻覺檢測。通過檢查層級激活模式,我們試圖檢測LLM生成中的確切幻覺開始和持續標記。我們的實驗表明,這種方法成功地隔離了幻覺的開始,儘管類別極度不平衡,仍在精確度-召回率AUC上實現了相對隨機基準的顯著改善。最終,我們提出了一個新穎的跨模型檢測框架,其中一個模型觀察另一個模型生成所引發的內部表示。我們發現,外部觀察者可以匹配或超越生成器對其自身幻覺開始的自我檢測,包括當觀察者是較小的模型時,這表明自我檢測並不是開始定位的上限。

HydroJEV: A one-second, training-free screen for cyber-attack and fault attribution in water distribution networks

2610.02048v1 by Tianwei Mu, Shengyan Jiang, Mingzhe Yuan, Qing Luo, Min Xiao, Wenhong Wang, Jun Li, Manhong Huang

When a SCADA alarm is raised in a water distribution network, operators must decide quickly whether it reflects a cyberattack, a physical fault, a normal transient or a faulty sensor. Supervised classifiers need labelled incidents that utilities rarely have, and frontier large language models (LLMs) take tens of seconds per decision. We tested whether Jev, a training-free model that returns class probabilities in about one second, can serve as the first tier of this triage. On a four-class cause-attribution benchmark built on the C-Town network in EPANET, Jev was compared with a hand-written rule tree, a supervised classifier and seven cloud LLMs on identical evidence in four sealed, pre-registered rounds. With only a label-free prior correction, Jev matched the rule tree (macro-F1 0.62-0.64 against 0.56-0.61 in distribution) and exceeded the supervised classifier by 0.36-0.42 on event subtypes absent from its labels, in all four rounds, and it outperformed the classifier whenever fewer than about four labelled events per class were available. Jev also decided 20-40 times faster than frontier LLMs. Accepting only benign Jev verdicts confirmed by the rule tree spared an LLM reviewer 35-38% of windows on fresh sealed sets without loss of macro-F1. Transferred unchanged to two further networks, this gated cascade stayed within the non-inferiority margin of its reviewer on all four sets. A fast, training-free screen can therefore take over about a third of the review load in SCADA anomaly triage while preserving the accuracy of deliberate review.

摘要:當水分配網絡中發生 SCADA 警報時,操作員必須迅速決定這是否反映了網絡攻擊、物理故障、正常瞬態或故障傳感器。監督式分類器需要標記的事件,而公用事業公司很少擁有這些事件,前沿的大型語言模型 (LLMs) 每次決策需要幾十秒。我們測試了 Jev,這是一個無需訓練的模型,能在約一秒內返回類別概率,是否可以作為這一分診的第一層。在基於 EPANET 的 C-Town 網絡構建的四類原因歸因基準上,Jev 與手寫規則樹、監督式分類器和七個雲端 LLM 在四輪相同證據的封閉、預註冊回合中進行了比較。僅通過無標籤的先驗修正,Jev 在四輪中與規則樹相匹配(宏 F1 0.62-0.64 對比 0.56-0.61 的分佈),並在缺少其標籤的事件子類型上超過了監督式分類器 0.36-0.42,並且每當每類可用標記事件少於約四個時,它都超越了分類器。Jev 的決策速度也比前沿 LLM 快 20-40 倍。僅接受由規則樹確認的良性 Jev 判決,讓 LLM 審核者在新封閉集上節省了 35-38% 的窗口,而不損失宏 F1。這一閘道級聯在另外兩個網絡上未經改變地轉移,並在所有四組中保持在其審核者的非劣性邊際內。因此,快速、無需訓練的篩選可以接管 SCADA 異常分診中約三分之一的審核負擔,同時保持仔細審核的準確性。

Typological Alignment of Stack-Based Language Models on Mildly Context-Sensitive Artificial Languages

2610.02040v1 by Nadine El-Naggar, Tatsuki Kuribayashi, Ted Briscoe

Some properties of languages, e.g., subject-object-verb (SOV) word order, are more prevalent than others among the thousands of attested natural languages (NLs). Such typological commonality is often attributed to learning biases. Computational simulations, recently with language models (LMs), have facilitated the exploration of this theory. In this paper, we extend existing analyses of the relationship between LMs' learning biases and typological commonality on both data and model sides, focusing on: (i) cross-serial dependencies, the upper limit of attested syntactic complexity, and (ii) stack-based LMs (SLMs), potentially facilitating learning of hierarchical patterns. We first evaluate generalization of SLMs on cross-serial dependencies across diverse artificial languages and confirm that they struggle with such constructions. However, SLMs with limited working memory generalize better suggesting a possible basis for such inductive bias and thus the typological commonality of some word order configurations.

摘要:某些語言的特性,例如主詞-受詞-動詞(SOV)語序,在數千種已證實的自然語言(NLs)中比其他特性更為普遍。這種類型學的共性通常被歸因於學習偏差。計算模擬,最近使用語言模型(LMs),促進了對這一理論的探索。在本文中,我們擴展了現有的分析,研究 LMs 的學習偏差與類型學共性之間的關係,重點關注:(i)交叉序列依賴,已證實的句法複雜性的上限,以及(ii)基於堆疊的 LMs(SLMs),可能促進對層次模式的學習。我們首先評估 SLMs 在各種人工語言中的交叉序列依賴的概括能力,並確認它們在這類結構上存在困難。然而,具有有限工作記憶的 SLMs 的概括能力更佳,這暗示了這種歸納偏差的可能基礎,從而解釋某些語序配置的類型學共性。

CARM: Cancellation-Aware Response Masking for LLM Reinforcement Learning

2610.02039v1 by Yafei Zhang, Songshuo Lu, Sicong Liao, Zhi Chen, Yaohua Tang

Recent years have witnessed the rapid adoption of reinforcement learning (RL) in large language model (LLM) post-training, with substantial gains in mathematical reasoning and code generation. In practical systems, however, policy updates and differences between rollout and training engines can make sampled responses off-policy. Sequence-level masking addresses this mismatch by deciding whether an entire response should contribute to optimization. A common masking rule uses the length-normalized geometric mean of sampled token probability ratios. Its signed log-ratios can cancel across positions, concealing substantial bidirectional policy drift. We propose \emph{Cancellation-Aware Response Masking} (CARM), a sequence-level mask that takes the absolute value of each token log-ratio before averaging, preventing opposing probability changes from canceling. We prove that accepted responses satisfy a joint bound on the fraction of sampled-token ratios outside a prescribed band and their mean log-distance beyond its boundaries. Experiments on mathematical reasoning and code generation show that CARM improves mean@16 averaged over AIME 2024/2025/2026 and BeyondAIME by up to $3.13$ percentage points over geometric-mean masking, and increases average pass@1 across four code benchmarks by $2.88$ points over the strongest evaluated baseline. These findings support CARM as a theoretically grounded and effective method for response-level off-policy control in LLM reinforcement learning.

摘要:近年來,強化學習 (RL) 在大型語言模型 (LLM) 的後訓練中迅速被採用,並在數學推理和程式碼生成方面取得了顯著的進展。然而,在實際系統中,政策更新和回滾與訓練引擎之間的差異可能使得抽樣的回應偏離政策。序列級的遮罩通過決定整個回應是否應該參與優化來解決這種不匹配。一個常見的遮罩規則使用抽樣的標記概率比的長度正規化幾何平均數。其簽名對數比可以在位置之間相互抵消,隱藏了實質性的雙向政策漂移。我們提出了 \emph{Cancellation-Aware Response Masking} (CARM),這是一種序列級遮罩,在平均之前取每個標記對數比的絕對值,防止相對的概率變化相互抵消。我們證明了接受的回應滿足在規定範圍外的抽樣標記比的比例和其均值對數距離的聯合界限。在數學推理和程式碼生成的實驗中顯示,CARM 在 AIME 2024/2025/2026 和 BeyondAIME 的 mean@16 上比幾何平均遮罩提高了最多 $3.13$ 個百分點,並且在四個程式碼基準上平均 pass@1 比最強的評估基準提高了 $2.88$ 分。這些發現支持 CARM 作為一種理論基礎和有效的方法,用於 LLM 強化學習中的回應級偏政策控制。

Mimir: Physics-Grounded LLM Agents for Long-Horizon Irrigation Control

2610.02038v1 by Yimeng Liu, Mi Zhang, Younsuk Dong, Zhichao Cao

Large language model (LLM) agents increasingly combine reasoning, tool use, and action, but most evidence comes from episodic tasks with relatively immediate feedback and reset failures. Long-running physical control operates in a different regime: actions alter future states, errors compound across decisions, and an agent must improve from experience without being allowed to rewrite the physical rules that make execution safe. We study this regime through irrigation, where daily decisions interact with soil-water dynamics over entire growing seasons. We present Mimir, a physics-grounded LLM agent organized around two repair timescales. At the fast timescale, a structured physical interface and deterministic simulator turn an LLM output into a proposal that we numerically check, revise, and subject to bounded deterministic action selection before execution. At the slow timescale, recurrent failure patterns are consolidated into persistent contextual principles that condition future proposals, while the physical model, evaluator, and execution constraints remain immutable. Under a common retrospective evaluator across multiple sites, crops, and years, Mimir attains the lowest reported aggregate control cost among the evaluated references and uses about 51% less irrigation than the historical schedule replay. The ablation study show higher control cost when forward simulation, verified revision, or persistent context is removed; model-scale and model-family studies show no monotonic gain from increasing LLM size. The resulting lesson show that persistent physical agents can combine semantic reasoning with bounded, evidence-driven self-improvement while reserving physical truth and actuator authority for explicit numerical mechanisms.

摘要:大型語言模型(LLM)代理越來越多地結合推理、工具使用和行動,但大多數證據來自於具有相對即時反饋和重置失敗的情境任務。長期運行的物理控制運作在不同的範疇:行動改變未來狀態,錯誤在決策中累積,代理必須在不被允許重寫使執行安全的物理規則的情況下從經驗中改進。我們通過灌溉來研究這個範疇,日常決策與整個生長季節的土壤-水動力學相互作用。我們提出了Mimir,一個基於物理的LLM代理,圍繞兩個修復時間尺度組織。在快速時間尺度上,結構化的物理介面和確定性模擬器將LLM輸出轉換為提案,我們對其進行數值檢查、修訂,並在執行之前進行有界的確定性行動選擇。在慢速時間尺度上,重複的失敗模式被整合成持久的上下文原則,這些原則調節未來的提案,而物理模型、評估者和執行約束保持不變。在多個地點、作物和年份的共同回顧評估者下,Mimir在評估的參考中達到了最低報告的總體控制成本,並使用了比歷史計劃重播少約51%的灌溉量。消融研究顯示,當移除前向模擬、驗證修訂或持久上下文時,控制成本會提高;模型規模和模型家族研究顯示,增加LLM大小並未產生單調增益。最終的教訓顯示,持久的物理代理可以將語義推理與有界的、基於證據的自我改進結合,同時為明確的數值機制保留物理真理和執行機構的權威。

Global Coherence: When Every Agent Is Right and the Team Is Still Wrong - A Local-to-Global Semantic Foundation for Multi-Agent Collaboration

2610.02036v1 by Xin Heng

AI agents can each make locally valid decisions yet jointly produce an invalid result. We call this the global coherence problem: a failure of shared state, not merely of model intelligence. Our Observation-Aliasing Impossibility Theorem gives the exact boundary. A policy can guarantee a valid action exactly when all worlds producing the same observation share an admissible action. If k indistinguishable worlds require pairwise-disjoint actions, the best randomized worst-case success is 1/k; more reasoning, roles, messages, or samples cannot recover the missing distinction. A stronger model can reason better within its context, but it cannot see beyond it. We then give local-to-global runtime semantics X = (H, C, G, F; D): topology H records overlapping scopes; category C governs state-changing actions; groupoid G retains reversible translations; sheaf F tests whether local views glue into one world; and minimal history D keeps only distinctions that alter legal futures. Models propose; the harness owns shared state and governs commit. Nine studies test both the failure and its boundary. On a controlled revision benchmark, the same frontier model scores 40/40 when the deciding event is visible; when it is hidden, tested arms score 12--17/40, consistent with chance (1/3); restoring one authoritative fact returns 40/40. On TeamBench, ordinary teams exceed a shared budget in 5/5 runs, a visible live count leaves 4/5 violations, and commit enforcement leaves 0/5. In tau2-bench Telecom, current-state checks score 0.07 after silent reverts, while the harness scores 1.00. Where a conventional solver already owns the complete relevant state, it ties the harness as predicted. The counterintuitive conclusion is that local intelligence cannot substitute for missing global state.

摘要:AI 代理可以各自做出局部有效的決策,但共同產生無效的結果。我們稱這為全球一致性問題:共享狀態的失敗,而不僅僅是模型智能的失敗。我們的觀察-混淆不可能定理給出了確切的邊界。當所有產生相同觀察的世界共享一個可接受的行動時,政策可以保證一個有效的行動。如果 k 個不可區分的世界需要成對不相交的行動,最佳隨機最壞情況成功率為 1/k;更多的推理、角色、消息或樣本無法恢復缺失的區別。一個更強的模型可以在其上下文中進行更好的推理,但它無法超越這一點。然後我們給出局部到全球的運行時語義 X = (H, C, G, F; D):拓撲 H 記錄重疊的範疇;類別 C 管理狀態改變的行動;群體 G 保留可逆的翻譯;束 F 測試局部視圖是否粘合成一個世界;而最小歷史 D 僅保留改變合法未來的區別。模型提出;鞍具擁有共享狀態並管理提交。九項研究測試了失敗及其邊界。在一個受控的修訂基準上,當決策事件可見時,相同的邊界模型得分 40/40;當它被隱藏時,測試的臂得分 12--17/40,與機會相符 (1/3);恢復一個權威事實返回 40/40。在 TeamBench 上,普通團隊在 5/5 次運行中超過了共享預算,一個可見的實時計數留下 4/5 次違規,而提交執行留下 0/5。在 tau2-bench Telecom 中,當前狀態檢查在靜默回退後得分 0.07,而鞍具得分 1.00。當一個傳統求解器已經擁有完整的相關狀態時,它的表現與預測一致。反直覺的結論是,局部智能無法替代缺失的全球狀態。

SPHERE: Adaptive VR Indoor Scene Generation via LLM-Enhanced Spatial Preference Learning and Human-in-the-Loop RL

2610.02023v1 by Hyeonmin Lee, Zheng Wei, Kyungmin Kwon, Jumin Seo, Jiwon Park, Hayoung Oh

While Large Language Models (LLMs) advance 3D indoor scene synthesis, current pipelines fail to retain user-specific preferences across sessions, making immersive authoring a repetitive and physically fatiguing process. We present SPHERE, an adaptive VR generation framework that transforms isolated synthesis into continuous human-AI co-creation. SPHERE extracts persistent spatial preferences from natural multimodal interactions (speech and controller edits). To ensure geometric resilience against spatial distortions, it abstracts these raw edits into hierarchical constraints modeling both local functional and global topological contexts. Furthermore, a human-in-the-loop reinforcement learning mechanism dynamically updates retrieval policies based on the user's final edited scenes. A mixed-design user study ($N=42$) and an offline ablation demonstrate that SPHERE significantly reduces corrective edits and physical demand, preventing bias toward shallow object-level traits to yield geometrically resilient, profile-aligned layouts. Ultimately, SPHERE demonstrates how capturing demonstrated spatial logic enables controlled spatial adaptation, establishing a reliable, governed human-AI collaboration framework for immersive authoring. Project page and source code will be available at: https://github.com/hyeonmin11/SPHERE

摘要:大型語言模型(LLMs)在三維室內場景合成方面取得了進展,但當前的流程無法在不同會話中保留用戶特定的偏好,使得沉浸式創作成為一個重複且身體疲憊的過程。我們提出了SPHERE,一個自適應虛擬現實生成框架,將孤立的合成轉變為持續的人機協作創作。SPHERE從自然的多模態互動(語音和控制器編輯)中提取持久的空間偏好。為了確保對空間扭曲的幾何韌性,它將這些原始編輯抽象為層次約束,建模本地功能和全局拓撲上下文。此外,一個人機交互的強化學習機制根據用戶最終編輯的場景動態更新檢索策略。一項混合設計的用戶研究($N=42$)和一個離線消融實驗顯示,SPHERE顯著減少了修正編輯和身體需求,防止對淺層物體特徵的偏見,以產生幾何上韌性、與用戶檔案對齊的佈局。最終,SPHERE展示了如何捕捉所展示的空間邏輯實現受控的空間適應,建立了一個可靠的、有規範的人機協作框架以進行沉浸式創作。項目頁面和源代碼將在以下鏈接提供:https://github.com/hyeonmin11/SPHERE

Old Ideas, Novel Problems: The Instability of LLM-Based Novelty Evaluation

2610.02022v1 by Noy Sternlicht, Simra Shahid, Peter Jansen, Daniel S. Weld, Pao Siangliulue, Tom Hope

Automated ideation systems are often evaluated on the novelty of the ideas they produce, and that judgment is increasingly delegated to large language models. Such judges are typically built ad hoc and validated, if at all, on human-authored papers rather than on the generated ideas they are meant to score. So, how do novelty judges perform? Not well. We present a systematic controlled study of novelty evaluation design choices. We first build an evaluation set automatically, mining OpenReview for passages where reviewers explicitly affirm or dispute a paper's originality and keeping only submissions with unanimous agreement at the extremes of their research area; we pair these with ideas from a vanilla LLM generator. Across six judges, we find that small prompt design choices have large consequences; e.g., simply telling the judge that reviewers found one idea novel and the other not can change its verdict on more than half of the identical idea pairs it is shown, shifting pairwise accuracy by over 50 points and occasionally pushing it below chance. The same change helps one judge and hurts another. Retrieval and larger reasoning budgets help little, and two purpose-built novelty evaluators are outperformed by our cheapest prompted baseline. These results raise questions about reported novelty gains of automated ideation systems, and call for robust novelty evaluation methods.

摘要:自動化構思系統通常根據其產出的創新性來進行評估,而這一判斷越來越多地委託給大型語言模型。這些評判者通常是臨時構建的,並且如果有的話,通常是在人工撰寫的論文上進行驗證,而不是在它們所要評分的生成想法上。因此,創新性評判者的表現如何?
表現不佳。我們呈現了一項系統性的控制研究,探討創新性評估設計選擇。我們首先自動構建一個評估集,從 OpenReview 中挖掘出評論者明確肯定或質疑論文原創性的段落,並僅保留在其研究領域極端處獲得一致同意的提交;我們將這些與來自普通 LLM 生成器的想法配對。在六位評判者中,我們發現小的提示設計選擇會產生大的後果;例如,僅僅告訴評判者評論者認為一個想法是新穎的而另一個不是,就可以改變其對超過一半相同想法對的判決,將成對準確率改變超過 50 點,有時甚至將其推至低於隨機機率。同樣的變化對一位評判者有幫助,卻對另一位評判者造成傷害。檢索和更大的推理預算幫助不大,而兩個專門構建的創新性評估者的表現不及我們最便宜的提示基準。這些結果引發了對自動化構思系統報告的創新性增益的質疑,並呼籲建立穩健的創新性評估方法。

Task-Adaptive Grounded 3D-Programmers Using 2D VLMs

2610.02021v1 by Arman Raayatsanati, Sombit Dey, Anna-Maria Halacheva, Jan-Nico Zaech, Luc Van Gool, Danda Pani Paudel

Recent vision-language models (VLMs) exhibit remarkable generalization and reasoning abilities, yet 3D understanding in these models is limited by data scale, training diversity, and reasoning capacity. Instead of naively extending these models into 3D, we take a different approach: we enable powerful 2D VLMs to operate reliably in 3D by introducing 3D grounding and iterative feedback loops with two novel concepts: Canonical Coordinate Framing (CCF) and Task-Adaptive Feedback (TAF). CCF serves as a unified visual representation that anchors both inputs and outputs to a shared Euclidean coordinate system, solving common challenges in 3D grounding such as axis ambiguity, inconsistent metric scale, and floating references. Complementary to this structured framing of the 3D inputs, TAF closes the reasoning loop with task-adaptive dynamic feedback that enables 2D VLMs to perform varied open-vocabulary tasks within their native visual context. Building on this foundation, we introduce 3D-Prog, a 3D understanding, reasoning, and generation framework that jointly employs the capabilities of CCF and TAF together with powerful VLMs. Without requiring any retraining, 3D-Prog performs open-vocabulary 3D understanding, manipulation, and generation across both object-level and scene-level tasks. Our experiments show that the joint use of CCF and TAF transforms 2D VLMs into geometry-aware 3D programmers, achieving consistent, interpretable, and high-quality results across diverse 3D tasks.

摘要:最近的視覺-語言模型(VLMs)展現出卓越的泛化和推理能力,但這些模型在3D理解方面受到數據規模、訓練多樣性和推理能力的限制。與其天真地將這些模型擴展到3D,我們採取了不同的方法:我們通過引入3D基礎和迭代反饋循環,使強大的2D VLMs能夠在3D中可靠運作,並提出了兩個新概念:典範坐標框架(CCF)和任務自適應反饋(TAF)。CCF作為一種統一的視覺表示,將輸入和輸出固定在共享的歐幾里得坐標系中,解決了3D基礎中常見的挑戰,如軸模糊、不一致的度量尺度和浮動參考。TAF則補充了這種3D輸入的結構化框架,通過任務自適應的動態反饋關閉推理循環,使2D VLMs能夠在其本土視覺上下文中執行各種開放詞彙任務。在這一基礎上,我們介紹了3D-Prog,一個3D理解、推理和生成框架,該框架共同利用CCF和TAF的能力以及強大的VLMs。3D-Prog在不需要任何重新訓練的情況下,能夠在物體級和場景級任務中執行開放詞彙的3D理解、操作和生成。我們的實驗表明,CCF和TAF的聯合使用將2D VLMs轉變為幾何感知的3D程序員,在各種3D任務中實現一致、可解釋和高質量的結果。

Controllable Multi-label Video Safety Detection via Adaptive Tversky Policy Optimization

2610.02019v1 by Guangyu Yang, Jingbiao Mei, Mingsheng Sun, Jinghong Chen, Yingtong Bu, Pengda Qin, Da Chen, Bill Byrne

The rapid growth of video-based social media has increased users' exposure to harmful content, creating a need for reliable automated video safety detection. Although recent Vision-Language Models (VLMs) show strong video understanding capabilities, existing harmful video detection systems face two key limitations: they typically reduce safety detection to binary classification, overlooking the inherently multi-label nature of unsafe videos, and they rely on static training objectives that do not support controllable precision-recall trade-offs, though the desired operating point may vary across moderation pipelines and unsafe categories. To address these gaps, we propose Adaptive Tversky Policy Optimization (ATPO), a reinforcement learning framework for Multi-label Video Safety Detection (Multi-VSD). ATPO introduces the Adaptive Tversky Reward (ATR), which dynamically adjusts false-positive and false-negative penalties during training to enable controllable precision-recall trade-offs. Experiments on SafeWatch-Bench and XD-Violence show that ATPO substantially improves multi-label performance, increasing the Jaccard Index from 40.66 to 75.44 on SafeWatch-Bench-Real. Moreover, ATR enables reliable steering of the precision-recall operating point, supporting deployment scenarios with heterogeneous policy requirements. Code and checkpoints are provided at https://bruceyg.github.io/ATPO-project-page/ .

摘要:視頻社交媒體的快速增長增加了用戶接觸有害內容的機會,這創造了對可靠自動視頻安全檢測的需求。儘管最近的視覺-語言模型(VLMs)顯示出強大的視頻理解能力,但現有的有害視頻檢測系統面臨兩個主要限制:它們通常將安全檢測簡化為二元分類,忽視了不安全視頻固有的多標籤特性,並且依賴於靜態訓練目標,這些目標不支持可控的精確度-召回率權衡,儘管所需的操作點可能在不同的審核管道和不安全類別之間變化。為了解決這些問題,我們提出了自適應Tversky政策優化(ATPO),這是一個用於多標籤視頻安全檢測(Multi-VSD)的強化學習框架。ATPO引入了自適應Tversky獎勵(ATR),在訓練過程中動態調整假陽性和假陰性懲罰,以實現可控的精確度-召回率權衡。在SafeWatch-Bench和XD-Violence上的實驗顯示,ATPO顯著提高了多標籤性能,將SafeWatch-Bench-Real上的Jaccard指數從40.66提高到75.44。此外,ATR使得精確度-召回率操作點的可靠調整成為可能,支持具有異質政策需求的部署場景。代碼和檢查點可在https://bruceyg.github.io/ATPO-project-page/ 獲得。

On Language Drift during RLVR Post-Training

2610.02015v1 by Michael Sullivan, Alexander Koller

Recent advances in LLM reasoning models---driven primarily by the paradigm of post-training via reinforcement learning with verifiable reward (RLVR)---have enabled them to accomplish impressively complex tasks. However, in parallel with their rising capabilities, LLMs have increasingly displayed signs of language drift in their chains of thought (CoTs): unusual, non-standard, and seemingly nonsensical language use. Although it is well-documented---and can potentially impair CoT monitorability---the causes of language drift are thus far poorly understood. In this paper, we identify the conditions under which language drift occurs: we prove theoretically that RLVR optimization pressure permits unbounded language drift, while supervised fine-tuning does not. We then show empirically that language drift specifically arises during RLVR on novel reasoning tasks---i.e. when the target behavior cannot be drawn out of the base model. Finally, we prove that it is not possible to constrain language drift without constraining expected reward, suggesting that CoT monitorability cannot be improved without harming performance during RLVR post-training at the frontier.

摘要:最近在大型語言模型(LLM)推理模型方面的進展——主要受到可驗證獎勵的強化學習後訓練(RLVR)範式的驅動——使得它們能夠完成令人印象深刻的複雜任務。然而,隨著它們能力的提升,LLM在其思維鏈(CoTs)中越來越顯示出語言漂移的跡象:不尋常的、非標準的,且似乎毫無意義的語言使用。儘管這一點已被充分記錄——並且可能會影響CoT的可監控性——語言漂移的原因至今仍然了解不深。在本文中,我們確定了語言漂移發生的條件:我們理論上證明,RLVR優化壓力允許無界的語言漂移,而監督式微調則不然。然後,我們實證顯示,語言漂移特別是在RLVR處理新推理任務時出現——即當目標行為無法從基礎模型中引出時。最後,我們證明,若不限制預期獎勵,就無法約束語言漂移,這表明在RLVR後訓練的前沿中,無法改善CoT的可監控性而不損害性能。

Atoms to Processes: The Role of Artificial Intelligence and Machine Learning in Chemical Engineering

2610.02014v1 by Michael Baldea, Linda J. Broadbelt, Marianthi G. Ierapetritou, Akhilesh Jain, Ankur Kumar, Thomas A. Kwan, Fèlix Llovell, Andrew J. Medford, Ilias Mitrai, Joel Paulson, Junyi Qiao, Matthew P. Rivera, Kirti C. Sahu, Lev Sarkisov, Zachary P. Smith, Calvin Tsay, Ching-Mei Wen, Victor M. Zavala, Huacheng Zhang, Dan Zhao

The rapid maturation of artificial intelligence (AI) and machine learning (ML) has catalyzed a profound shift in how chemical engineering problems are formulated, analyzed, and solved. Advances in computing, data availability, and learning algorithms have enabled AI/ML methods to impact applications spanning atomic-scale simulations, materials and catalyst discovery, transport and thermodynamics, separations, process systems engineering, and industrial operations. This article provides a perspective on recent methodological developments and representative applications, emphasizing how AI/ML tools are being integrated with first-principles models to address challenges of predictive accuracy, data scarcity, extrapolation, interpretability, and model lifecycle management. Across domains, a unifying trend is the move away from purely black-box approaches toward hybrid and physics-informed frameworks that explicitly respect conservation laws, thermodynamic consistency, and known structural constraints. These approaches not only improve robustness and reliability, but also enable meaningful human-AI collaboration by providing information at an appropriate level of abstraction for the task and decision context. We conclude that AI and ML are not replacing the core principles of chemical engineering; rather, they are amplifying them. As the field advances toward increasingly autonomous, adaptive, and sustainable systems, the thoughtful integration of AI/ML with first-principles understanding and domain expertise will be essential to realizing their full potential across both research and industrial practice.

摘要:人工智慧(AI)和機器學習(ML)的快速成熟催化了化學工程問題的公式化、分析和解決方式的深刻變化。計算、數據可用性和學習算法的進步使得AI/ML方法能夠影響從原子級模擬、材料和催化劑發現、傳輸和熱力學、分離、過程系統工程到工業運營的應用。本文提供了對近期方法論發展和代表性應用的觀點,強調AI/ML工具如何與第一性原理模型相結合,以應對預測準確性、數據稀缺性、外推、可解釋性和模型生命周期管理的挑戰。在各個領域,一個統一的趨勢是從純粹的黑箱方法轉向混合和物理知識驅動的框架,這些框架明確遵循守恆法則、熱力學一致性和已知結構約束。這些方法不僅提高了穩健性和可靠性,還通過在適當的抽象層次上提供信息來促進有意義的人機協作,以適應任務和決策背景。我們得出結論,AI和ML並不是取代化學工程的核心原則;相反,它們是在放大這些原則。隨著該領域向越來越自主、自適應和可持續的系統邁進,AI/ML與第一性原理理解和領域專業知識的深思熟慮的整合將對實現其在研究和工業實踐中的全部潛力至關重要。

Counting Moves, Weighing Voices: Bayesian Dialectical Argumentation for Calibrated Multi-LLM Councils under Persistent Adversaries

2610.02005v1 by Ionel Eduard Stan, Paolo Napoletano

A multi-LLM \emph{council} lets several large language models (LLMs) deliberate on a question and return an answer together with a confidence estimate. As these systems become increasingly used for reasoning, that confidence should represent a calibrated \emph{probability of being correct}, and the decision should remain robust when some agents are persistently unreliable. Existing \emph{council aggregation} methods fail on both fronts: their confidence estimates measure decisiveness rather than correctness, and they cannot identify or discount persistently unreliable agents. We introduce Bayesian Dialectical Argumentation (BDA), which treats the council's \emph{typed} moves---who proposed, challenged, or conceded which answer---as observations of a classical annotator model with \emph{per-agent} reliabilities. This formulation recasts multi-agent deliberation as a reliability estimation problem, using the deliberation trace to infer agent reliability under persistent adversarial behavior. By weighting evidence according to inferred agent reliability, BDA yields calibrated posterior probabilities over candidate answers while allowing persistently unreliable agents to be inverted rather than merely outvoted. Across binary and multi-class benchmarks, BDA achieves the best calibration among zero-cost council aggregation methods, requiring no additional LLM calls, and improves robustness under persistent adversarial coalitions while remaining competitive in clean settings.

摘要:一個多LLM \emph{委員會} 讓幾個大型語言模型(LLMs)對一個問題進行討論並一起返回答案及信心估計。隨著這些系統在推理中的使用越來越多,這種信心應該代表一個經過校準的 \emph{正確性概率},而且當某些代理持續不可靠時,決策應保持穩健。現有的 \emph{委員會聚合} 方法在這兩方面都失敗:它們的信心估計衡量的是決斷性而非正確性,並且無法識別或排除持續不可靠的代理。我們引入貝葉斯辯證論證(BDA),將委員會的 \emph{類型化} 行動——誰提出、挑戰或讓步於哪個答案——視為具有 \emph{每個代理} 可靠性的經典標註者模型的觀察。這一表述將多代理討論重新定義為一個可靠性估計問題,利用討論痕跡來推斷在持續對抗行為下的代理可靠性。通過根據推斷的代理可靠性加權證據,BDA 產生對候選答案的經過校準的後驗概率,同時允許持續不可靠的代理被反轉,而不僅僅是被投票淘汰。在二元和多類基準測試中,BDA 在零成本委員會聚合方法中實現了最佳的校準,無需額外的 LLM 調用,並在持續對抗聯盟下提高了穩健性,同時在乾淨的環境中保持競爭力。

Mem++: Non-Destructive Memory for Long-Term Organizational LLM Agents

2610.02002v1 by Ahmad Yehia, Aly O. Abdelkareem, Islam Ahmed, Hesham Omran, Khaled Alashmouny, Christian Claudel, Abduallah Mohamed

Large Language Model (LLM) agents now take part in organizational work, where many authors record decisions across documents over months. Because a revised decision arrives as a new document rather than an edit, answering a question requires knowing which version held at a given time. However, most memory systems compress the record at write time. By distilling each document into facts, notes or graph edges, these methods fix what can be answered before any question is asked. To address this, we propose Mem++, a non-destructive memory framework shifting from write-time distillation to read-time selection. Mem++ stores every document whole with its date and author, and it calls no generative model at write time. At read time, it retrieves only documents dated up to the time a question asks about and fuses lexical and semantic rankings. Unlike systems that overwrite older versions, Mem++ keeps them and leaves the choice to the answering model. Evaluations on the organizational benchmark OrgMemBench demonstrate that Mem++ surpasses the strongest memory system baseline by 8.0 to 13.1 points across two answering models. With gpt-4.1-mini, it also achieves the best overall score, 2.6 points above RAG. In addition, Mem++ achieves the best average LLM-judge score on LoCoMo and ranks second on LongMemEval-S, behind only its entity-graph variant. Code for benchmark evaluation is available at https://github.com/AIDAChip-Inc/mem-plus-plus.

摘要:大型語言模型(LLM)代理現在參與組織工作,許多作者在數月內記錄決策於文件中。因為修訂的決策以新文件的形式出現,而不是編輯,因此回答問題需要知道在特定時間持有的是哪個版本。然而,大多數記憶系統在寫入時會壓縮記錄。通過將每個文件提煉成事實、筆記或圖邊,這些方法在任何問題被提出之前固定了可以回答的內容。為了解決這個問題,我們提出了Mem++,這是一個非破壞性的記憶框架,從寫入時的提煉轉向讀取時的選擇。Mem++ 將每個文件完整地存儲,並附上日期和作者,並且在寫入時不調用任何生成模型。在讀取時,它僅檢索在問題詢問時的日期之前的文件,並融合詞彙和語義排名。與覆蓋舊版本的系統不同,Mem++ 保留它們,並將選擇權留給回答模型。在組織基準測試 OrgMemBench 上的評估顯示,Mem++ 在兩個回答模型中超越了最強記憶系統基線 8.0 到 13.1 分。使用 gpt-4.1-mini 時,它還獲得了最佳整體分數,比 RAG 高出 2.6 分。此外,Mem++ 在 LoCoMo 上獲得了最佳平均 LLM-judge 分數,並在 LongMemEval-S 中排名第二,僅次於其實體圖變體。基準評估的代碼可在 https://github.com/AIDAChip-Inc/mem-plus-plus 獲得。

Mingbird: A Local-First Agent Harness Enabling Small Open Models to Complete Real Tasks

2610.02001v1 by Hao Wang, Ting Huang

Small open-weight models (2-9B) run on ordinary laptops, but under cloud-scale agent harnesses they rarely complete real tasks: tool prefill overflows the context, self-correction diverges, tool demonstrations loop, and tasks are silently abandoned. We present evidence, from a controlled single-machine comparison and one third-party benchmark, that a substantial share of these failures is attributable to the harness rather than the model. We introduce Mingbird, a local-first agent harness for Windows and Ollama whose ten mechanisms compensate point-by-point for small-model failure forms, three of them representative: a byte-level net-zero prefill budget, a finish gate that re-reads the task before accepting completion, and signature-level loop detection. On LRAB, a controlled comparison holding machine, models, budgets, and scoring fixed (4 harnesses $\times$ 4 open models (2B-35B) $\times$ 18 real tasks, deterministic artifact scoring), Mingbird reaches 0.886 overall against 0.631 (goose), 0.479 (opencode), and 0.405 (agent-mini), with all 288 cells published; on $τ^2$-bench (278 tasks, three arms, one protocol) it totals 0.856 against 0.791 and 0.737; and a frontier-model probe on the same 18 tasks spans 0.997 to 0.478 across harnesses, with well-formed scaffolds staying within 0.072 of each other. A leave-one-mechanism-out ablation is reported as directional only: same-night replications of the same arm move its mean by up to 0.069, the size of every nominal single-trial delta, and the one batch-matched comparison (full mechanism stack versus text re-read alone) gives the executable completion guards a paired +0.10 across three replications. The evidence carries stated limits: a self-built benchmark, a single machine, and single-trial scoring.

摘要:小型開放權重模型(2-9B)可以在普通筆記型電腦上運行,但在雲端規模的代理環境下,它們很少能完成實際任務:工具預填超出上下文,自我修正偏離,工具演示循環,任務則被默默放棄。我們提供證據,來自受控的單機比較和一個第三方基準,顯示這些失敗的相當一部分是由於代理環境而非模型本身。我們介紹了 Mingbird,一個針對 Windows 和 Ollama 的本地優先代理環境,其十個機制逐點補償小型模型的失敗形式,其中三個具有代表性:字節級的淨零預填預算、一個在接受完成之前重新閱讀任務的完成門,以及簽名級的循環檢測。在 LRAB 上,進行了受控比較,固定了機器、模型和預算(4 個代理環境 × 4 個開放模型(2B-35B) × 18 個真實任務,確定性工件評分),Mingbird 的整體得分為 0.886,相較於 0.631(goose)、0.479(opencode)和 0.405(agent-mini),所有 288 個單元均已發布;在 $τ^2$-bench(278 個任務、三個臂、一個協議)中,其總得分為 0.856,相較於 0.791 和 0.737;而在同 18 個任務上的前沿模型探測中,跨越的得分範圍為 0.997 到 0.478,各代理環境之間的良好結構保持在 0.072 之內。一個去除一個機制的消融實驗僅報告為方向性:同夜對同一臂的重複實驗使其均值變化最多達 0.069,這是每個名義單次試驗的變化量,而唯一的批次匹配比較(完整機制堆疊與僅文本重讀)在三次重複中給可執行完成保護提供了配對的 +0.10。這些證據有明確的限制:自建基準、單一機器和單次試驗評分。

Can AI Oversight Be Zero Knowledge?

2610.01995v1 by Alessandro Chiesa, Ziyi Guan, Burcu Yildiz

AI systems increasingly produce outputs from confidential data, such as a fitness-for-duty assessment from medical records or the predicted properties of a drug candidate from its secret structure. It is important to verify that such outputs are correct without revealing the underlying data. A recent line of work studies verification of AI outputs via interactive proofs and debate for oracle-aided computation, where correctness may depend on an oracle such as human judgment, a physical experiment, or the web. These works focus on verification by a verifier that runs much faster than the computation. However, such efficient verification is impossible for general oracle-aided computation, and these works therefore rely on additional assumptions. We focus instead on privacy: allowing the verifier to run in time polynomial in the computation, we ask whether interactive arguments for oracle-aided computation can be zero knowledge, so that the verifier learns nothing about the confidential data beyond the correctness of the output. We prove that, in general, they cannot. In the random oracle model, there are no zero-knowledge proofs for all oracle-aided computations, even if both the prover and the verifier are allowed to run much longer than the computation itself. The impossibility extends to debate, a canonical model for scalable oversight. On the positive side, we show that if the oracle attaches a cryptographic signature to each of its answers, then every oracle-aided computation can be verified in zero knowledge with an efficient prover and verifier, assuming only collision-resistant hash functions. Beyond privacy, this also gives an alternative approach to scalable oversight that relies neither on an honest opponent, as in debate, nor on the robustness of the computation, as in prior single-prover protocols.

摘要:AI 系統越來越多地從機密數據中產生輸出,例如從醫療記錄中進行的適任性評估或從其秘密結構中預測的藥物候選物的性質。驗證這些輸出是否正確而不透露底層數據是很重要的。最近的一系列研究通過互動證明和辯論研究 AI 輸出的驗證,用於有 oracle 輔助的計算,其中正確性可能依賴於 oracle,例如人類判斷、物理實驗或網絡。這些研究專注於由運行速度遠快於計算的驗證者進行的驗證。然而,對於一般的有 oracle 輔助計算,這樣的高效驗證是不可能的,因此這些研究依賴於額外的假設。我們則專注於隱私:允許驗證者在計算的多項式時間內運行,我們詢問有 oracle 輔助計算的互動論證是否可以是零知識的,以便驗證者不會學到關於機密數據的任何信息,除了輸出的正確性。我們證明,通常情況下,它們是不可能的。在隨機 oracle 模型中,對於所有有 oracle 輔助的計算,沒有零知識證明,即使證明者和驗證者都被允許運行的時間遠超過計算本身。這種不可能性擴展到辯論,這是一個可擴展監督的典型模型。從積極的一面來看,我們展示了如果 oracle 為其每個答案附加加密簽名,那麼每個有 oracle 輔助的計算都可以在零知識中進行驗證,並且有高效的證明者和驗證者,只假設碰撞抗性哈希函數。除了隱私之外,這還提供了一種可擴展監督的替代方法,既不依賴於誠實的對手(如辯論中),也不依賴於計算的穩健性(如以前的單證明者協議中)。

Universal Byte-Level Encoding: UTF-8/UTF-16 Routing to Reduce Cross-Script Token-Budget Disparities

2610.01984v1 by Hyunsik Kim, Youngmoon Jung

Byte-level byte-pair encoding (BBPE) tokenizers are attractive for multilingual large language models (LLMs) because they cover all Unicode text. In UTF-8-based BBPE, however, many scripts start from a higher fallback cost than English: when no learned merges can be applied, a multibyte character requires multiple byte-derived symbols. We call this worst-case pre-merge cost the encoding floor. A higher floor can increase token counts and per-request cost and shrink usable context. Changing the text encoding can reduce this gap, but a single global encoding can make already-efficient English spans more expensive in mixed-script text. We propose Universal Byte-Level Encoding (UBE), a dual-alphabet tokenizer that keeps 1-2-byte UTF-8 characters on the UTF-8 path while routing 3-4-byte UTF-8 characters through UTF-16. This lowers the encoding floor for 3-byte Basic Multilingual Plane (BMP) characters in scripts with high token premiums (token counts relative to English) without raising it for already-efficient spans in mixed-script text. UBE changes only the byte representation presented to byte-pair encoding (BPE); the merge rule remains standard, and exact decoding is preserved. UBE also composes with alternative boundary policies and morphology-based representations. In a Unicode 17 audit, UBE exactly round-trips all Unicode scalar values and all inputs in the official normalization, grapheme-break, and emoji test suites. Across intrinsic evaluations, UBE lowers dispersion in English-normalized token-count ratios, reducing cross-lingual token-budget disparity. In multilingual language model (LM) experiments, UBE matches BBPE's LM quality. In the main multilingual settings, UBE reduces token counts most for high-premium scripts and slightly lowers English token counts, yielding more usable context under fixed token budgets and faster prompt processing in content-matched benchmarks.

摘要:Byte-level byte-pair encoding (BBPE) 令牌器對於多語言大型語言模型 (LLMs) 來說非常有吸引力,因為它們涵蓋了所有 Unicode 文本。然而,在基於 UTF-8 的 BBPE 中,許多字母系統的回退成本高於英語:當無法應用學習到的合併時,多字節字符需要多個字節衍生符號。我們將這種最壞情況的合併前成本稱為編碼底線。較高的底線會增加令牌數量和每次請求的成本,並縮小可用上下文。改變文本編碼可以減少這一差距,但單一的全局編碼可能會使已經高效的英語範圍在混合字母系統文本中變得更昂貴。我們提出了通用字節級編碼 (UBE),這是一種雙字母表的令牌器,保持 1-2 字節的 UTF-8 字符在 UTF-8 路徑上,同時將 3-4 字節的 UTF-8 字符通過 UTF-16 路由。這降低了在高令牌溢價(相對於英語的令牌數量)的字母系統中 3 字節基本多語言平面 (BMP) 字符的編碼底線,而不提高在混合字母系統文本中已經高效的範圍的底線。UBE 只改變呈現給字節對編碼 (BPE) 的字節表示;合併規則保持標準,並保留精確解碼。UBE 還可以與替代邊界策略和基於形態學的表示進行組合。在 Unicode 17 的審計中,UBE 精確地回傳所有 Unicode 標量值以及官方標準化、字形斷裂和表情符號測試套件中的所有輸入。在內部評估中,UBE 降低了英語標準化令牌計數比率的分散性,減少了跨語言的令牌預算差異。在多語言語言模型 (LM) 實驗中,UBE 的 LM 質量與 BBPE 相匹配。在主要的多語言設置中,UBE 對高溢價字母系統的令牌數量減少最多,並稍微降低英語的令牌數量,從而在固定令牌預算下產生更多可用上下文,並在內容匹配基準測試中加快提示處理速度。

Counterfactual Auditing of Bias in Open-Source Large Language Models for Clinical Triage

2610.01963v1 by Manar Aljohani, Brandon Ho, Kenneth McKinley, Dennis Ren, Xuan Wang

Emergency department (ED) triage is a high-stakes prioritization task in which demographic, socioeconomic, and system-context information may improperly influence acuity assignment. Although open-source large language models (LLMs) are increasingly considered for local and privacy-preserving clinical decision support, it remains unclear how counterfactual bias varies across model families, sizes, medical-domain models, and domain-adapted models. We present a comparative counterfactual audit of ten open-source LLMs for pediatric Emergency Severity Index (ESI) prediction. Starting from real and handbook-style clinical vignettes, we construct paired counterfactual variants that change only one injected demographic, socioeconomic, healthcare-access, behavioral, social, or system-context variable while holding the clinical presentation fixed. Models include Qwen2.5-7B, Qwen2.5-14B-Instruct, a QLoRA fine-tuned Qwen2.5-7B, MedGemma variants, MedLLaMA2-7B, GPT-OSS-20B, and GPT-OSS-120B. We measure any counterfactual shift, undertriage, overtriage, shifts greater than one ESI level, mean shift, and mean absolute shift. Counterfactual sensitivity varied substantially and did not consistently decrease with larger model size or medical-domain pretraining. The fine-tuned Qwen2.5-7B showed the lowest overall sensitivity, with a 5.27% any-shift rate and mean absolute shift of 0.0534, versus 16.02% and 0.1706 for the base model. Several larger or medical-domain models showed more significant shifts. Stratified and correlation analyses further revealed clinically important directionality and shared failure patterns hidden by aggregate rates. These findings support counterfactual auditing as a lightweight, clinically interpretable framework for comparing fairness risks in open-source LLMs before clinical deployment.

摘要:急診部(ED)分診是一項高風險的優先排序任務,其中人口統計、社會經濟和系統背景信息可能不當影響急性程度的分配。儘管開源大型語言模型(LLMs)越來越被考慮用於本地和隱私保護的臨床決策支持,但目前尚不清楚反事實偏見在不同模型家族、大小、醫療領域模型和領域適應模型之間的變化情況。我們對十個開源LLM進行了針對兒科緊急嚴重性指數(ESI)預測的比較反事實審計。從真實和手冊風格的臨床小插曲開始,我們構建了配對的反事實變體,僅改變一個注入的人口統計、社會經濟、醫療訪問、行為、社會或系統背景變量,同時保持臨床表現不變。模型包括Qwen2.5-7B、Qwen2.5-14B-Instruct、經過QLoRA微調的Qwen2.5-7B、MedGemma變體、MedLLaMA2-7B、GPT-OSS-20B和GPT-OSS-120B。我們測量任何反事實變化、低估分診、過度分診、超過一個ESI級別的變化、平均變化和平均絕對變化。反事實敏感性變化顯著,且不一致地隨著模型大小或醫療領域預訓練的增大而減少。經過微調的Qwen2.5-7B顯示出最低的整體敏感性,任何變化率為5.27%,平均絕對變化為0.0534,而基礎模型則為16.02%和0.1706。幾個較大或醫療領域模型顯示出更顯著的變化。分層和相關分析進一步揭示了臨床上重要的方向性和由聚合率隱藏的共同失敗模式。這些發現支持反事實審計作為一種輕量級、臨床可解釋的框架,用於在臨床部署前比較開源LLM中的公平風險。

A Hybrid Approach to Malware Detection: Integrating Few-Shot Model-Agnostic Meta-Learning with Autoencoders

2610.01949v1 by Emmanuela Andam, Yasir Abbas Zaidi, Abdelali Hadir, Emmanuel Grant, Naima Kaabouch

Ransomware has emerged as a major cybersecurity threat, with incidents increasing in frequency and impact across critical sectors. These attacks are typically launched through phishing emails, malicious downloads, or exploitation of software vulnerabilities to gain system access. Once inside, the malware encrypts files and demands a ransom, often in cryptocurrency, for the decryption key. Conventional detection methods often struggle with novel or scarce samples, leaving systems vulnerable. To address these challenges, this paper proposes a hybrid deep learning framework that combines an Autoencoder Feature Extractor (AFE) with a Model Agnostic Meta Learning (MAML) classifier for few shot malware detection. The AFE generates compact latent features that reduce noise and dimensionality, while the MAML classifier rapidly adapts to new threats using limited labeled data. Experiments conducted on the Ransomware Dataset 2024 demonstrate the effectiveness of the framework in binary classification tasks. Across one to fifty shot settings, the proposed model consistently achieves high accuracy, F1 score, and Matthews Correlation Coefficient values, maintaining reliable classification even under extreme scarcity. These results highlight the model's robustness and effectiveness in adapting to limited data scenarios, demonstrating the potential of combining feature extraction with meta learning to enhance resilience against malware, particularly in sectors such as healthcare, manufacturing, and public infrastructure, where cyberattacks can cause significant operational and financial disruption.

摘要:勒索病毒已成為一個主要的網絡安全威脅,事件在關鍵行業中的頻率和影響不斷增加。這些攻擊通常通過釣魚電子郵件、惡意下載或利用軟件漏洞來獲得系統訪問權限。一旦進入,惡意軟件會加密文件並要求贖金,通常以加密貨幣的形式支付解密密鑰。傳統的檢測方法在面對新穎或稀少的樣本時往往會掙扎,讓系統變得脆弱。為了應對這些挑戰,本文提出了一個混合深度學習框架,結合了自編碼器特徵提取器(AFE)和模型無關的元學習(MAML)分類器,用於少量樣本的惡意軟件檢測。AFE生成緊湊的潛在特徵,減少噪聲和維度,而MAML分類器則利用有限的標記數據快速適應新威脅。在2024年勒索病毒數據集上進行的實驗證明了該框架在二元分類任務中的有效性。在一到五十次樣本設置中,所提出的模型始終能夠實現高準確率、F1分數和馬修斯相關係數值,即使在極度稀缺的情況下也能保持可靠的分類。這些結果突顯了該模型的穩健性和在有限數據場景中適應的有效性,展示了將特徵提取與元學習相結合以增強對抗惡意軟件的韌性的潛力,特別是在醫療、製造和公共基礎設施等行業中,這些行業的網絡攻擊可能會造成重大的運營和財務中斷。

Latent JEPA: Abstract Future Prediction for Latent Reasoning in Chemistry

2610.01947v1 by Xinjian Zhao, Yaoyao Xu, Xuemin Chen, Xiaozhuang Song, Tianshu Yu

Large language models offer a promising foundation for chemical reasoning, bringing together chemical knowledge and multistep problem solving. Chemical intuition can provide an initial sense of plausible outcomes before the details of a solution are fully worked out. Inspired by how such expectations complement explicit analysis, we study how continuous latent thoughts can be trained to anticipate informative aspects of future solutions without verbalizing every intermediate step. We introduce Latent JEPA, a framework that combines autoregressive learning with joint-embedding prediction of one or more future views. For chemical reasoning, we develop textual and molecular prediction objectives that connect latent thoughts to both subsequent reasoning and molecular outcomes. Experiments on ChemCoTBench show gains in molecular optimization and on several editing and reaction metrics. Representation analyses show that future prediction makes latent thoughts more informative about molecular outcomes and strengthens their correspondence with chemical structure. These findings support abstract future prediction as a learning principle for connecting continuous latent reasoning with scientific outcomes.

摘要:大型語言模型為化學推理提供了一個有前景的基礎,將化學知識和多步驟問題解決結合在一起。化學直覺能在解決方案的細節完全展開之前,提供一種合理結果的初步感知。受到這種期望如何補充明確分析的啟發,我們研究如何訓練連續潛在思維,以預測未來解決方案的資訊性方面,而不需要逐步口頭表達每一個中間步驟。我們介紹了潛在JEPA,一個將自回歸學習與一個或多個未來視圖的聯合嵌入預測相結合的框架。針對化學推理,我們開發了文本和分子預測目標,將潛在思維與後續推理和分子結果連接起來。在ChemCoTBench上的實驗顯示,在分子優化以及幾個編輯和反應指標上都有提升。表徵分析顯示,未來預測使潛在思維對分子結果的資訊性更強,並加強了它們與化學結構的對應性。這些發現支持將抽象的未來預測作為一種學習原則,以連接連續的潛在推理與科學結果。

A rubric landscape for evaluating clinical reasoning in large language models: what exists, what is missing, and what needs to be combined

2610.01938v1 by Zhangshu Joshua Jiang, Zina Ibrahim, James T. Teo

Exam-style accuracy does not establish whether large language models (LLMs) reason well over clinical records. We define clinical reasoning as integrating and updating evidence across time and sources to form, revise and justify a patient's problem representation and a defensible plan. This structured narrative review maps three literatures: medical education assessment instruments, clinical LLM benchmarks published from 2023 onwards, and general-domain methods for evaluating long-form generation. We examine six dimensions: problem representation, temporal synthesis, differential and management reasoning, counterfactual reasoning, calibrated uncertainty, and reasoning faithfulness. Preprints are included and flagged. No single instrument covers all six dimensions. Problem representation and differential or management reasoning are reasonably covered, although reliability varies by instrument and setting. TIMER-Eval targets temporal synthesis, and ER-Reason assesses sequential diagnostic belief updating. Dedicated uncertainty and counterfactual evaluations are emerging, but their applicability to longitudinal free-text reasoning remains limited. Factual completeness is well theorised in general-domain evaluation, with early clinical evidence of important omissions. Faithfulness remains the weakest dimension, with one identified clinical causal-ablation study on multiple-choice questions. Existing tools should be combined through binary rubric items, separate completeness and correctness scores, case-specific importance weighting with non-compensable safety caps, temporal order-consistency checks, and chance-corrected reliability reporting. Further design work is needed for calibrated uncertainty, counterfactual reasoning and faithfulness over longitudinal free-text records. This review provides a design rationale, not a validated instrument.

摘要:考試風格的準確性並不能確定大型語言模型(LLMs)在臨床記錄上是否能夠進行良好的推理。我們將臨床推理定義為整合和更新跨時間和來源的證據,以形成、修訂和辯護病人的問題表述及可辯護的計劃。
這篇結構化的敘述性回顧映射了三個文獻領域:醫學教育評估工具、2023年以來發表的臨床LLM基準,以及評估長篇生成的一般領域方法。我們檢視了六個維度:問題表述、時間綜合、差異和管理推理、反事實推理、校準的不確定性,以及推理的忠實性。預印本已被納入並標記。
沒有單一的工具涵蓋所有六個維度。問題表述以及差異或管理推理的覆蓋相對合理,儘管可靠性因工具和環境而異。TIMER-Eval 針對時間綜合,而 ER-Reason 評估連續的診斷信念更新。專門的不確定性和反事實評估正在出現,但它們對於縱向自由文本推理的適用性仍然有限。事實的完整性在一般領域評估中有良好的理論基礎,並且早期臨床證據顯示出重要的遺漏。忠實性仍然是最薄弱的維度,其中有一項針對多選題的臨床因果消融研究被識別。
現有工具應通過二元評分項目、分開的完整性和正確性分數、特定案例的重要性加權(帶有不可補償的安全上限)、時間順序一致性檢查以及機會修正的可靠性報告進行結合。對於校準的不確定性、反事實推理和縱向自由文本記錄的忠實性,還需要進一步的設計工作。這篇回顧提供了一個設計的理由,而不是一個經過驗證的工具。

Mapping the RAG Landscape: A Four Axis Taxonomy of Efficiency, Defense, Interactivity, and Reasoning

2610.01936v1 by Meghana Sunil, Shravya V, Shravan Venkatraman, Joe Dhanith PR

Large Language Models (LLMs) have demonstrated remarkable fluency across many tasks but remain limited by their static, parameter bound knowledge and their susceptibility to hallucinating information. Retrieval Augmented Generation (RAG) addresses these issues by incorporating external retrieval into the generation process, grounding model outputs in verifiable and up to date sources. While prior surveys primarily focus on core RAG architectures and standard pipelines, recent research explores broader challenges and capabilities that extend beyond these foundational designs. This survey provides a consolidated and structured examination of contemporary RAG developments, organizing the field into a four axis taxonomy: improving retrieval efficiency, strengthening robustness and security, supporting user driven and interactive workflows, and enabling multi step or complex reasoning. We formalize key components of the RAG framework and review methods spanning dense and sparse retrieval, fusion strategies, embedding optimizations, and reinforcement learning based retrieval policies, highlighting how these advances influence practical deployment and system design. We also synthesize evaluation practices, domain specific applications, and architectural variants such as Naive, Advanced, and Modular RAG. Finally, we outline persistent challenges related to retrieval quality, reliability, domain adaptation, scalability, and explainability, and identify opportunities for building RAG systems that are more reliable, adaptable, and transparent.

摘要:大型語言模型(LLMs)在許多任務中展現了卓越的流暢性,但仍然受到靜態的、參數限制的知識以及對虛假信息的易感性的限制。檢索增強生成(RAG)通過將外部檢索納入生成過程來解決這些問題,使模型輸出基於可驗證且最新的來源。雖然之前的調查主要集中在核心RAG架構和標準流程上,但最近的研究探討了超越這些基礎設計的更廣泛挑戰和能力。本調查提供了一個當代RAG發展的綜合和結構化檢視,將該領域組織為四個軸向的分類法:提高檢索效率、加強穩健性和安全性、支持用戶驅動和互動工作流程,以及實現多步驟或複雜推理。我們正式化了RAG框架的關鍵組件,並回顧了涵蓋密集和稀疏檢索、融合策略、嵌入優化和強化學習基於檢索政策的方法,強調這些進展如何影響實際部署和系統設計。我們還綜合了評估實踐、特定領域的應用以及如Naive、Advanced和Modular RAG等架構變體。最後,我們概述了與檢索質量、可靠性、領域適應性、可擴展性和可解釋性相關的持續挑戰,並確定了構建更可靠、可適應和透明的RAG系統的機會。

Cross-Lingual Alignment for Decoder-Only Models using MoE Routers

2610.01921v1 by Lucas Bandarkar, Clark Peng, Ahmed Haj Ahmed, Aditi Khandelwal, Nanyun Peng

Cross-lingual contrastive learning has been a core component of multilingual encoder training, but the ability to explicitly align representations is not possible in decoder-only LLMs because of varying multilingual tokenization. However, a growing amount of research suggests that even in LLMs, higher cross-lingual representational alignment leads to improved cross-lingual transfer. In this paper, we propose a novel approach to reimagine cross-lingual contrastive learning given the architectural constraints of modern LLMs. Rather than applying an auxiliary alignment loss on hidden states, we propose using the outputs of the mixture-of-experts (MoE) routers as the target for alignment. Router outputs lend themselves better to pooling over many tokens, enabling more reliable cross-lingual comparisons at the sequence-level. Controlled continual pre-training experiments on four open-source MoEs show that incorporating this routing loss also aligns the underlying hidden representations across languages. Most importantly, this loss improves multilingual performance on our diverse evaluation suite, demonstrating the potential of cross-lingual MoE router alignment.

摘要:跨語言對比學習一直是多語言編碼器訓練的核心組成部分,但在僅使用解碼器的LLM中,由於多語言標記化的差異,無法明確對齊表示。然而,越來越多的研究表明,即使在LLM中,更高的跨語言表示對齊也能改善跨語言轉移。在本文中,我們提出了一種新穎的方法,重新構想在現代LLM的架構限制下進行跨語言對比學習。我們提議不在隱藏狀態上應用輔助對齊損失,而是使用混合專家(MoE)路由器的輸出作為對齊的目標。路由器的輸出更適合在多個標記上進行池化,從而在序列級別上實現更可靠的跨語言比較。在四個開源MoE上進行的受控持續預訓練實驗顯示,納入這一路由損失也使得不同語言之間的隱藏表示得以對齊。最重要的是,這一損失提高了我們多樣化評估套件上的多語言性能,展示了跨語言MoE路由器對齊的潛力。

MoLE: Mixture of Latent Experts for Complementary Visual Reasoning

2610.01917v1 by Yingcheng Liu, Tianyi Jiang, Yujuan Ding, jiangbo Ai, Xun Jiang, Guoqing Wang, Wei Ye, Yi Bin

Latent visual reasoning equips vision--language models with continuous intermediate states that can process visual evidence without explicit textual reasoning traces or repeated image operations. However, existing methods often allow multiple latent tokens to access the same visual evidence through shared value projections, providing no mechanism for them to extract complementary visual information; simply increasing the latent budget can therefore yield redundant latent representations. We argue that effective latent reasoning should encourage different latent tokens to extract complementary visual information, and thereby act as specialized visual experts. Based on this insight, we propose MoLE, a Mixture of Latent Experts framework that controls both what visual evidence each latent visual expert observes and how it transforms that evidence. MoLE isolates latent visual experts during evidence extraction and uses dedicated latent summary experts to aggregate the complementary representations of latent visual experts. A two-stage training pipeline first forces visual evidence through this latent pathway and then restores direct visual access, requiring neither predefined expert roles nor intermediate visual targets. Across five visual reasoning benchmarks, MoLE achieves an average score of 78.6, outperforming data-matched supervised fine-tuning by 4.9 and the strongest evaluated latent visual reasoning baseline at the same latent budget by 3.6. Representation analyses show lower latent-state similarity and more diverse visual attention, while masking the latent pathway reduces average performance by 9.2. These results demonstrate that specializing latent computation is more effective than merely increasing the number of latent tokens.

摘要:潛在視覺推理使視覺-語言模型具備連續的中間狀態,能夠在沒有明確文本推理痕跡或重複圖像操作的情況下處理視覺證據。然而,現有的方法往往允許多個潛在標記通過共享的值投影訪問相同的視覺證據,這並未提供提取互補視覺信息的機制;因此,僅僅增加潛在預算可能會產生冗餘的潛在表示。我們認為,有效的潛在推理應該鼓勵不同的潛在標記提取互補的視覺信息,從而充當專門的視覺專家。基於這一見解,我們提出了MoLE,一種潛在專家混合框架,控制每個潛在視覺專家觀察的視覺證據及其如何轉化這些證據。MoLE在證據提取過程中隔離潛在視覺專家,並使用專門的潛在摘要專家來聚合潛在視覺專家的互補表示。一個兩階段的訓練流程首先強迫視覺證據通過這一潛在路徑,然後恢復直接的視覺訪問,既不需要預定義的專家角色,也不需要中間視覺目標。在五個視覺推理基準上,MoLE的平均得分為78.6,比數據匹配的監督微調高出4.9,比在相同潛在預算下評估的最強潛在視覺推理基線高出3.6。表示分析顯示潛在狀態相似性較低,視覺注意力更為多樣,而屏蔽潛在路徑則使平均性能降低9.2。這些結果表明,專門化潛在計算比單純增加潛在標記的數量更有效。

Asynchronous LLM Post-Training: Group-Mass Capping and Convergence Analysis

2610.01896v1 by Qijia He, Ruinan Jin, Jun Luo, Shaofeng Zou, Yingbin Liang

Asynchronous reinforcement learning (RL) improves the efficiency of large language model post-training but introduces stale rollouts generated by earlier policies. Theoretical understanding of how this staleness affects convergence and how to mitigate its impact remains limited. We derive a convergence bound for GRPO-style algorithms that explicitly characterizes the tradeoff between the gradient estimator's second moment and bias. For trajectory-level importance-weighted estimators, our analysis shows that once the second moment is uniformly controlled, delay enters the bound through the bias introduced by clipping or rescaling. Guided by this insight, we propose a novel group mass capping GRPO (GMC-GRPO) method, which minimizes a ratio-based bias bound within a class of weighted estimators sharing a common second-moment guarantee. We establish convergence guarantees for asynchronous GMC-GRPO and show that, compared with TIC-GRPO, it improves the threshold dependence of the fourth-order delay term from $O(ε^{-4})$ to $O(ε^{-2})$ as $ε\to0$, where $1+ε$ is the ratio threshold. Under local policy overlap, the delay-dependent term decreases as $G^{-2/5}$ after tuning the step size, where $G$ is the group size. For fixed behavior and current policies, the bias introduced by group rescaling also vanishes as $G\to\infty$, whereas the bias from trajectory-wise clipping can persist. Experiments across Qwen3 models and reasoning benchmarks demonstrate improved robustness to stale rollouts, with GMC-GRPO achieving the best performance among stable baselines under large rollout delays.

摘要:非同步強化學習 (RL) 提升了大型語言模型後訓練的效率,但引入了由早期策略生成的過時回饋。對於這種過時性如何影響收斂以及如何減輕其影響的理論理解仍然有限。我們推導了 GRPO 風格算法的收斂界限,明確描述了梯度估計器的二階矩與偏差之間的權衡。對於軌跡級別的重要性加權估計器,我們的分析顯示,一旦二階矩被均勻控制,延遲通過剪裁或重新縮放引入的偏差進入界限。在這一見解的指導下,我們提出了一種新穎的群體質量上限 GRPO (GMC-GRPO) 方法,該方法在共享共同二階矩保證的加權估計器類別中最小化基於比率的偏差界限。我們為非同步 GMC-GRPO 建立了收斂保證,並顯示與 TIC-GRPO 相比,它改善了四階延遲項的閾值依賴性,從 $O(ε^{-4})$ 提升至 $O(ε^{-2})$ 當 $ε\to0$ 時,其中 $1+ε$ 是比率閾值。在局部策略重疊下,延遲依賴項在調整步長後減少為 $G^{-2/5}$,其中 $G$ 是群體大小。對於固定的行為和當前策略,群體重新縮放引入的偏差在 $G\to\infty$ 時也會消失,而來自軌跡級別剪裁的偏差則可能持續存在。針對 Qwen3 模型和推理基準的實驗顯示對過時回饋的穩健性有所改善,GMC-GRPO 在大型回饋延遲下在穩定基準中達到了最佳性能。

A Structured State Space Sequence Model for Multi-Class Classification of Malware

2610.01893v1 by Emmanuela Andam, Rana Shaaban, Emanuel Grant, Naima Kaabouch

By 2030, Internet of Things (IoT) devices are projected to reach 40 billion, with fast-paced technological advancements in fields such as industry, healthcare, agriculture, automobiles, and building/home automation systems. This expansion has created a large attack surface for cybercrime, as the majority of these devices open the door for cybercriminals to exploit vulnerabilities, as they lack adequate built-in security. Cybercriminals launch malware attacks to compromise systems or steal sensitive data, and once a system is compromised, a ransom is typically demanded for its release. Current cybersecurity measures in place are being outpaced by the rapid growth of the IoT, which is accompanied by a subsequent growth in malware variants being created per day. Recognizing this pitfall, this research examines and proposes a novel approach to malware detection and classification to safeguard devices from further attacks and make IoT systems more robust and secure. The framework proposed utilizes a Structured State Space Sequence (S4) model, which discretizes sequences of malware samples in a sequence and captures long-range dependencies, essentially identifying the "cause" and "effect" hidden within malware execution flow. This study presents two novel contributions: the first empirical application of the S4 model for malware analysis, and a comprehensive comparison of its performance against other deep learning architectures, laying the stepping stone for future research in this new paradigm.

摘要:到2030年,物聯網(IoT)設備預計將達到400億個,隨著工業、醫療保健、農業、汽車以及建築/家庭自動化系統等領域的快速技術進步。這一擴張為網絡犯罪創造了巨大的攻擊面,因為大多數這些設備為網絡犯罪分子利用漏洞提供了機會,因為它們缺乏足夠的內建安全性。網絡犯罪分子發動惡意軟體攻擊以破壞系統或竊取敏感數據,一旦系統被攻破,通常會要求贖金以換取其釋放。目前的網絡安全措施已被物聯網的快速增長所超越,這伴隨著每天創造的惡意軟體變種的增長。認識到這一陷阱,本研究檢視並提出了一種新穎的惡意軟體檢測和分類方法,以保護設備免受進一步攻擊並使物聯網系統更加穩健和安全。所提出的框架利用結構化狀態空間序列(S4)模型,該模型將惡意軟體樣本的序列離散化並捕捉長期依賴性,實質上識別出惡意軟體執行流程中隱藏的“原因”和“結果”。本研究提出了兩項新穎的貢獻:首次將S4模型應用於惡意軟體分析,以及對其性能與其他深度學習架構的全面比較,為未來在這一新範式中的研究奠定了基礎。

Selection-Based Structured Reasoning: Toward Efficient Multimodal Search Agents

2610.01892v1 by Feiyu Gavin Zhu, Xiaoyu Zhu, Jiqi Yang, Rui Yang, Arnab Kumar Mondal, Yancheng Wang, Xinke Deng, Jean Oh, Reid Simmons, Joerg Liebelt, Xiang Kong, Zhongyu Jiang

Multimodal agents commonly generate free-form reasoning before each action. For small models, limited model capacity can result in lengthy reasoning that provides little useful guidance for action generation while incurring substantial inference cost. To address this challenge, we introduce Selection-based Structured Reasoning (SSR), a framework that reformulates reasoning as selection instead of open-ended generation. SSR represents recurring high-level reasoning as pre-specified, reusable natural-language candidates. At each turn, the model selects from these reasoning candidates based on their likelihoods given the current context, without requiring an auxiliary task head. Using pre-specified reasoning traces enables parallel scoring, where teacher-forced prefilling computes token likelihoods concurrently within and across candidates using a shared context KV cache. We evaluate SSR on seven multimodal search benchmarks using 2B and 4B models. Across multiple reinforcement learning objectives and supervised fine-tuning, SSR delivers significant efficiency gains without sacrificing task performance. SSR achieves an average success rate competitive with leading search agents of the same scale, while reducing per-turn reasoning latency by over 90% and total per-question model inference latency by 28-54%. Project page: https://zfy0314.github.io/ssr-webpage/.

摘要:多模態代理通常在每個行動之前生成自由形式的推理。對於小型模型,有限的模型容量可能導致冗長的推理,這對於行動生成提供的有用指導有限,同時產生可觀的推理成本。為了解決這一挑戰,我們引入了基於選擇的結構化推理(SSR),這是一個將推理重新定義為選擇而非開放式生成的框架。SSR將重複的高級推理表示為預先指定的、可重用的自然語言候選者。在每一輪中,模型根據當前上下文的可能性從這些推理候選者中進行選擇,而不需要輔助任務頭。使用預先指定的推理痕跡可以實現並行打分,其中教師強制預填充在共享的上下文KV緩存中同時計算候選者內部和之間的標記可能性。我們在七個多模態搜索基準上使用2B和4B模型評估SSR。在多個強化學習目標和監督微調中,SSR在不犧牲任務性能的情況下實現了顯著的效率提升。SSR的平均成功率與同規模的領先搜索代理競爭,同時將每輪推理延遲減少超過90%,每題模型推理延遲減少28-54%。項目頁面:https://zfy0314.github.io/ssr-webpage/。

Unsupervised Domain Adaptation for Enhanced Radiometer Image Precipitation Estimation using Conditional Flow Matching

2610.01890v1 by Victor Enescu, Assaad Zeghina, Matthieu Meignin, Nicolas Viltard, Cécile Mallet

Deep generative networks have recently achieved unprecedented performance in precise image and video editing using sophisticated textual prompts. However, the effectiveness of such models heavily depends on access to very large supervised and annotated image datasets, which can be very difficult to obtain. This is particularly true for satellite instruments, which very rarely overlap with labelled data, and suffer from domain shifts in the rare occasions they do. In this paper, we investigate the potential of flow matching models for unsupervised domain adaptation of satellite radiometer images. Our main contribution is a novel unsupervised method that achieves precise domain alignment by leveraging parts of the deterministic ordinary differential equations in flow matching models, conditioned on different satellite instruments. A key strength of our approach is its ability to preserve essential information while adapting across any domains since the perturbations are in theory bijective. Extensive experiments conducted on the GPM-Core constellation show the benefit of our conditional domain adaptation, particularly in improving rain precipitation estimation from radiometer imagery.

摘要:深度生成網絡最近在使用複雜文本提示進行精確圖像和視頻編輯方面取得了前所未有的表現。然而,這些模型的有效性在很大程度上依賴於獲取非常大的監督和標註圖像數據集,而這往往非常困難。這一點對於衛星儀器尤其如此,因為它們與標記數據的重疊非常少,並且在少數情況下重疊時會遭受領域轉移。在本文中,我們探討了流匹配模型在衛星輻射計圖像的無監督領域適應中的潛力。我們的主要貢獻是一種新穎的無監督方法,通過利用流匹配模型中確定性常微分方程的部分,實現精確的領域對齊,並以不同的衛星儀器為條件。我們方法的一個關鍵優勢是它在跨越任何領域時能夠保留重要信息,因為擾動在理論上是雙射的。在GPM-Core星座上進行的大量實驗顯示了我們的條件領域適應的好處,特別是在改善來自輻射計圖像的降雨量估算方面。

Stochastic Rounding in Low-Precision Transformer Inference: A Variable-Precision Emulation Study of a Small GPT-2

2610.01889v1 by Yohan Chatelain, Pablo de Oliveira Castro

Should low-precision transformer inference use stochastic rounding (SR) or round-to-nearest (RN)? The answer depends on where in the network you look. We isolate this effect by holding the numerical format fixed and varying only the rounding rule at individual operation sites. To enable experiments at freely chosen precisions, we extend the PRISM vectorized rounding library to arbitrary virtual precision via a variable-precision stochastic rounding (VPSR) algorithm, proving that the rounding decision is evaluated exactly in hardware floating point. We develop two analyses providing complementary insight into this site-level trade-off. First, a probabilistic forward-error bound for linear projections shows that SR's error envelope grows as $O(\sqrt{n} u)$ in reduction length $n$, versus $O(n u)$ for RN, a gap that widens rapidly at low precision and is most pronounced in the long multilayer perceptron (MLP) down-projection. Second, a second-order decomposition of expected cross-entropy loss change at the output softmax into signed drift, drift curvature, and a Fisher-weighted variance penalty reveals why the two sites behave oppositely: MLP noise is predominantly a uniform logit shift to which softmax is invariant, so SR's variance is largely discounted; head noise is non-uniform across the vocabulary and is not. On DistilGPT-2 at $t=6$ significand bits, observations match theory: SR in the MLP raises perplexity to 1.15x the full-precision reference, versus 2.21x for RN. At the language-model head, the ordering reverses because SR introduces non-uniform variance, whereas deterministic RN carries none. In a mixed-precision configuration (MLP output at $t=6$), assigning SR to the MLP and RN to the head brings perplexity within 1.10x of the full-precision reference, a 28% reduction over matched-bit RN.

摘要:低精度Transformer推斷應該使用隨機四捨五入 (SR) 還是四捨五入至最近值 (RN)?答案取決於你在網絡中的哪個位置觀察。我們通過固定數字格式並僅在各個操作位置變化四捨五入規則來隔離這一效應。為了在自由選擇的精度下進行實驗,我們通過變精度隨機四捨五入 (VPSR) 算法擴展了 PRISM 向量化四捨五入庫,以支持任意虛擬精度,證明四捨五入決策在硬體浮點中被精確評估。

我們開發了兩個分析,提供對這一位置級權衡的互補見解。首先,對線性投影的概率前向誤差界限顯示,SR 的誤差範圍隨著減少長度 $n$ 增長為 $O(\sqrt{n} u)$,而 RN 則為 $O(n u)$,這一差距在低精度下迅速擴大,並在長多層感知器 (MLP) 向下投影中最為明顯。其次,將輸出 softmax 的期望交叉熵損失變化進行二階分解為有符號漂移、漂移曲率和費舍爾加權方差懲罰,揭示了為什麼這兩個位置的行為相反:MLP 噪聲主要是一種均勻的 logit 偏移,softmax 對此不變,因此 SR 的方差在很大程度上被折扣;而頭部噪聲在詞彙中是非均勻的。

在 $t=6$ 的 DistilGPT-2 顯著位中,觀察結果與理論相符:MLP 中的 SR 將困惑度提高到全精度參考的 1.15 倍,而 RN 則為 2.21 倍。在語言模型頭部,排序顛倒,因為 SR 引入了非均勻方差,而確定性 RN 則沒有。在混合精度配置中 (MLP 輸出為 $t=6$),將 SR 指派給 MLP,將 RN 指派給頭部,使困惑度在全精度參考的 1.10 倍內,相比於匹配位的 RN 減少了 28%。

Flowing Faster to Coordinate: One-Step Online Multi-Agent Flow Policies

2610.01882v1 by Zhuoran Li, Yunzhan Li, Xun Wang, Yihan Du, Longbo Huang

Multi-agent reinforcement learning (MARL) provides a powerful framework for learning coordinated behaviors through interactions with the environment. Developing MARL policies requires balancing expressive modeling of complex and multimodal action distributions with efficient training and execution. Generative policies, particularly diffusionbased policies, can faithfully capture complex and multimodal behaviors, but costly iterative sampling hinders their scalability in online multi-agent settings. We propose an Online MARL framework via one-step Flow model (OMAF) that combines expressive generative policies with efficient one-step action generation. OMAF employs a Transformer-based flow policy to capture complex coordination behaviors, while its approximate path score surrogate provides a principled route to synchronized flow policy optimization. To enable stable and sampleefficient learning, we further develop a joint optimization scheme coupling softmax Q-value estimation with a joint flow policy objective for coordinated policy learning. By eliminating iterative sampling, OMAF dramatically reduces training overhead without sacrificing policy expressiveness. Extensive experiments across 10 standard tasks from MPE and MAMuJoCo show that OMAF consistently achieves superior performance, with up to 3.4x higher returns and 10.5x sample efficiency improvement compared with baseline methods. These results validate the effectiveness of OMAF as an expressive and computationally efficient one-step flow policy paradigm for online MARL.

摘要:多智能體強化學習(MARL)提供了一個強大的框架,通過與環境的互動來學習協調行為。開發MARL策略需要在複雜和多模態行動分佈的表達建模與高效訓練和執行之間取得平衡。生成策略,特別是基於擴散的策略,能夠真實捕捉複雜和多模態行為,但昂貴的迭代取樣限制了它們在在線多智能體環境中的可擴展性。我們提出了一個通過一步流模型(OMAF)的在線MARL框架,將表達豐富的生成策略與高效的一步行動生成相結合。OMAF採用基於Transformer的流策略來捕捉複雜的協調行為,而其近似路徑分數替代品則提供了一條原則性的路徑以實現同步流策略的優化。為了實現穩定且樣本高效的學習,我們進一步開發了一個聯合優化方案,將softmax Q值估計與協調策略學習的聯合流策略目標相結合。通過消除迭代取樣,OMAF顯著減少了訓練開銷,而不犧牲策略的表達性。在來自MPE和MAMuJoCo的10個標準任務中進行的廣泛實驗顯示,OMAF始終實現了卓越的性能,與基線方法相比,回報提高了高達3.4倍,樣本效率改善了10.5倍。這些結果驗證了OMAF作為一種表達豐富且計算高效的一步流策略範式在在線MARL中的有效性。

Where LLMs Fail with Visualization DSLs

2610.01873v1 by Chang Han, Andrew McNutt, Katherine Isaacs

As LLMs take up the role of authoring charts using visualization domain-specific languages (DSLs), the human constraints that shaped those languages may no longer apply, as what is easy for a person is not necessarily easy for a model. To understand how LLMs might work better with DSLs, we explore where and how they fail with current DSL designs. We evaluate 10 JSON-style visualization DSLs with 41 tasks across 3 LLMs, then assess the generated specifications with JSON and rendering checks, and qualitative coding of failed cases. Analyzing how this specification generation process fails, we identify four recurring failure patterns, link each to specific DSL features, and discuss design considerations for future DSL designs.

摘要:隨著大型語言模型(LLMs)擔任使用可視化領域特定語言(DSLs)創建圖表的角色,塑造這些語言的人類限制可能不再適用,因為對於人類來說簡單的事情不一定對模型來說也簡單。為了了解LLMs如何更好地與DSLs協作,我們探討了它們在當前DSL設計中失敗的地方和方式。我們評估了10個JSON風格的可視化DSL,涵蓋41個任務,並在3個LLMs上進行測試,然後通過JSON和渲染檢查以及失敗案例的定性編碼來評估生成的規範。分析這一規範生成過程的失敗,我們識別出四種重複出現的失敗模式,將每種模式與特定的DSL特徵聯繫起來,並討論未來DSL設計的設計考量。

From Network Intrusion Detection to Blockchain-Backed Endpoint Detection and Response: Mapping the Landscape of Decentralized Detection-and-Response Architectures

2610.01872v1 by Yahya Shahsavari, Sara Rouhani, Kaiwen Zhang

While the literature on blockchain-assisted intrusion detection and prevention systems (IDS/IPS) for Internet of Things (IoT) and Industrial Internet of Things (IIoT) networks is mature, existing systematic reviews suffer from two critical limitations: they overlook the structural shift toward modern Endpoint Detection and Response (EDR) and Extended Detection and Response (XDR) architectures, and they conflate blockchain's distinct functional roles into a single monolithic category. This Systematization of Knowledge (SoK) addresses these gaps by proposing a three-axis taxonomy that classifies proposals by detection-system class (NIDS, HIDS, EDR/XDR), blockchain functional role, and response-automation maturity. Synthesizing research published in high-impact venues between 2019 and 2026, we provide a rigorous gap analysis exposing why a genuine per-endpoint blockchain-anchored response loop remains nearly nonexistent due to latency, deployment, and community mismatches. Furthermore, we evaluate structural, cross-cutting challenges persisting across the literature, including consensus latency on constrained devices, post-quantum cryptographic vulnerability, smart-contract attack surfaces, and the adversarial vulnerability of evolving LLM-based detection engines. Finally, we outline a comprehensive research agenda centered on hybrid on-chain/off-chain orchestration to bridge the gap between decentralized trust and rapid response automation.

摘要:雖然有關區塊鏈輔助的入侵檢測和預防系統(IDS/IPS)在物聯網(IoT)和工業物聯網(IIoT)網絡中的文獻已相當成熟,但現有的系統性評估存在兩個關鍵限制:它們忽視了向現代端點檢測與響應(EDR)和擴展檢測與響應(XDR)架構的結構性轉變,並且將區塊鏈的不同功能角色混淆為一個單一的整體類別。這項知識系統化(SoK)通過提出一個三軸分類法來解決這些空白,該分類法根據檢測系統類別(NIDS、HIDS、EDR/XDR)、區塊鏈功能角色和響應自動化成熟度對提案進行分類。綜合2019年至2026年間在高影響力期刊上發表的研究,我們提供了一個嚴謹的差距分析,揭示了為何真正的每個端點區塊鏈錨定響應循環幾乎不存在,原因在於延遲、部署和社群不匹配。此外,我們評估了文獻中持續存在的結構性、跨領域挑戰,包括在受限設備上的共識延遲、後量子密碼學脆弱性、智能合約攻擊面以及不斷演變的基於LLM的檢測引擎的對抗性脆弱性。最後,我們概述了一個以混合鏈上/鏈下協同為中心的全面研究議程,以彌合去中心化信任與快速響應自動化之間的鴻溝。

Walking the Embedding Space: Datastore Extraction from Multimodal RAG

2610.01871v1 by Maria Carmen Jica, Ali Satvaty, Suzan Verberne, Fatih Turkmen

Multimodal Retrieval-Augmented Generation (MRAG) has emerged as a reliable and cost-effective technique of grounding the generative capabilities of Multimodal Large Language Models (MLLMs) into relevant, up-to-date, external knowledge. Despite presenting several benefits, such as reducing hallucinatory behavior, they also introduce new attack surfaces, including leakage of private information and vulnerabilities against data extraction attacks. In this paper, we introduce $\immrag$, an adaptive and automatic data extraction attack procedure operating in a black box setting against \emph{image-returning} MRAG, a configuration in which the retrieved visual artifact is itself the response. Each query blends an attacker-held shadow image with an image already recovered from the system, and relevance-weighted resampling steers subsequent queries towards regions of the embedding space that still yield novel retrievals. Unlike current extraction attacks that aim to persuade the model towards data leakage by placing a malicious query as a textual prompt, $\immrag$ embeds the malicious instructions inside a user-given input image. We evaluate $\immrag$ on three plausible and distinct real-world scenarios: medical assistant, document-focused helper and general purpose tool. The experiments involve the study of the effectiveness of the attack on multiple CLIP-family retrievers, as well as the impact of various generators. A single 2500-query run reconstructs up to 611 distinct radiology images, 566 document scans and 416 general-purpose images under local-feature correspondence, and reaches up to $5.6\times$ as many distinct datastore items as a non-adaptive baseline. Our results show the urgent need for safeguards specifically designed for multimodal data.

摘要:多模態檢索增強生成(MRAG)已成為將多模態大型語言模型(MLLMs)的生成能力與相關的、最新的外部知識相結合的一種可靠且具成本效益的技術。儘管它提供了幾個好處,例如減少幻覺行為,但它們也引入了新的攻擊面,包括私密信息洩漏和對數據提取攻擊的脆弱性。 在本文中,我們介紹了 $\immrag$,這是一種適應性和自動化的數據提取攻擊程序,針對 \emph{圖像返回} MRAG 在黑箱環境中運作,這是一種檢索的視覺工件本身就是回應的配置。每個查詢將攻擊者持有的影像與系統中已恢復的影像混合,並且相關性加權重採樣引導後續查詢朝向仍能產生新穎檢索的嵌入空間區域。與目前旨在通過將惡意查詢作為文本提示來說服模型進行數據洩漏的提取攻擊不同,$\immrag$ 將惡意指令嵌入用戶提供的輸入影像中。我們在三個合理且不同的現實場景中評估了 $\immrag$:醫療助手、文件專注助手和通用工具。實驗涉及對多個 CLIP 家族檢索器的攻擊有效性以及各種生成器的影響進行研究。一次 2500 次查詢的運行重建了多達 611 幅不同的放射學影像、566 幅文件掃描和 416 幅通用影像,根據局部特徵對應,並達到高達 $5.6\times$ 的不同數據庫項目數量,相較於非適應性基準。我們的結果顯示出對專門為多模態數據設計的安全措施的迫切需求。

From Isolated Feature to Orbits: Discovering Music Concepts via Multi-SAE Alignment

2610.01864v1 by Liwei Lin, Gus Xia

How can we understand what a music foundation model has learned \textit{internally}? Most interpretability approaches, such as probing and Sparse Autoencoders (SAEs), focus on identifying individual features with minimal structural assumptions. We argue that many concepts are better understood as \textit{structured relations} rather than isolated features. This is especially prominent in music, where tonal structures are organized in the space of pitch and time. For example, concepts such as chords or keys are naturally expressed as structured sets (e.g., the 12 transpositions of a chord or the diatonic system within a key), rather than isolated features. In this study, \textbf{we shift from feature identification to structure-based analysis}, asking whether the learned inner representations of music foundation model emerge as organized structures over features. To this end, we introduce a framework that uses pitch transposition as an inductive bias to induce ordered orbits via multi-view SAE alignment. Concretely, we generate pitch-shifted input pairs and align their SAE representations to discover structured groups of pitch-related features. Experimental results show that this approach recovers orbit structures corresponding to chords, keys, and melodic patterns across two state-of-the-art music foundation models, while requiring only minimal grounding (e.g., a few anchor examples) to interpret entire concept families.

摘要:如何理解音樂基礎模型所學到的\textit{內部}知識?大多數可解釋性方法,如探測和稀疏自編碼器(SAEs),專注於識別具有最小結構假設的個別特徵。我們認為,許多概念更應被理解為\textit{結構化關係}而非孤立特徵。這在音樂中尤為明顯,音調結構在音高和時間的空間中組織。例如,和弦或調的概念自然表達為結構化集合(例如,和弦的12個移調或調內的自然音階系統),而不是孤立特徵。在這項研究中,\textbf{我們從特徵識別轉向基於結構的分析},詢問音樂基礎模型學習到的內部表徵是否以結構化形式出現。為此,我們引入一個框架,利用音高移調作為誘導偏見,通過多視角SAE對齊來產生有序的軌道。具體而言,我們生成音高移位的輸入對,並對齊它們的SAE表徵,以發現與音高相關特徵的結構化組。實驗結果顯示,這種方法恢復了對應於和弦、調和旋律模式的軌道結構,並且只需最少的基礎(例如,幾個錨點示例)即可解釋整個概念家族。

AVSD-Scenes: A Dataset for Audio-Visual Description of Urban Scenes

2610.01861v1 by Dhanunjaya Varma Devalraju, Arshdeep Singh, Mark D. Plumbley

Natural language descriptions can provide rich semantic representations of audio-visual urban scenes, yet datasets that jointly describe both auditory and visual information remain limited. In this paper, we introduce AVSD-Scenes, a paired audio-visual scene description dataset for urban environments. The dataset contains 12,291 audio-visual scene descriptions generated from the TAU Urban Audio-Visual Scenes dataset. To construct the dataset, we first generate audio- and visual-based descriptions using Qwen2-Audio-7B and Qwen2.5-VL-7B, respectively. These modality-specific descriptions are then combined using large language models, namely Qwen3-14B, Mistral-Small-3.2-24B-Instruct-2506, and Gemma-3-27B-it, to produce multimodal descriptions that capture complementary information from both modalities. We benchmark AVSD-Scenes using semantic alignment, cross-modal retrieval, scene classification, LLM-as-a-judge evaluation, and human subjective assessment. Results show that multimodal descriptions improve semantic alignment and cross-modal retrieval performance compared with modality-specific descriptions while preserving strong scene-discriminative information. The generated descriptions achieve up to 94.5% accuracy in urban scene classification, while combining audio, visual, and description embeddings further improves accuracy to 95.4%. Furthermore, the descriptions remain highly scene-discriminative even when scene labels are removed from the prompting instructions, indicating that they capture semantic information derived from the audio-visual content rather than merely reflecting label information.

摘要:自然語言描述可以提供音視覺城市場景的豐富語義表示,但同時描述聽覺和視覺信息的數據集仍然有限。在本文中,我們介紹了 AVSD-Scenes,一個針對城市環境的配對音視覺場景描述數據集。該數據集包含 12,291 條從 TAU Urban Audio-Visual Scenes 數據集中生成的音視覺場景描述。為了構建該數據集,我們首先分別使用 Qwen2-Audio-7B 和 Qwen2.5-VL-7B 生成基於音頻和視覺的描述。這些特定於模態的描述然後使用大型語言模型進行結合,即 Qwen3-14B、Mistral-Small-3.2-24B-Instruct-2506 和 Gemma-3-27B-it,以生成捕捉兩種模態互補信息的多模態描述。我們使用語義對齊、跨模態檢索、場景分類、LLM作為評判標準的評估以及人類主觀評估來基準測試 AVSD-Scenes。結果顯示,相較於特定模態的描述,多模態描述改善了語義對齊和跨模態檢索性能,同時保留了強大的場景區分信息。生成的描述在城市場景分類中達到高達 94.5% 的準確率,而將音頻、視覺和描述嵌入結合進一步提高了準確率至 95.4%。此外,即使在提示指令中移除場景標籤,這些描述仍然保持高度的場景區分性,表明它們捕捉到的是源自音視覺內容的語義信息,而不僅僅是反映標籤信息。

Detecting Inconsistencies in Model Specifications with LLM-as-Verifier Reasoning

2610.01847v1 by Zichen Xie, Mrigank Pawagi, Lize Shao, Yang Hu, Wenxi Wang

Model specifications define how large language models (LLMs) should behave, guiding alignment training, inference-time behavior, and evaluation. Yet these specifications may themselves contain defects: two individually reasonable principles may prescribe incompatible behavior when applied to the same situation, leaving no response that satisfies both. Detecting such inconsistencies is challenging. Formalizing natural-language specifications risks losing subtle distinctions, while behavior-based testing cannot reliably distinguish specification defects from differences in model behavior. We introduce VeriSpec, the first approach to directly detect inconsistencies in model specifications by auditing the specification text itself. Our key insight is to preserve the specification in natural language while using an LLM as a verifier. VeriSpec extracts structured, context-aware rules, constructs a topic-guided graph to cluster behaviorally related rules at the same authority level, and applies LLM-as-verifier reasoning to detect inconsistencies. Applying VeriSpec to the OpenAI Model Spec, we extract 405 rules and manually validate five inconsistencies, all reported to its developers, who responded positively and have initiated internal discussions. Compared with five baselines, VeriSpec identifies the most validated inconsistencies, achieves the highest precision (38.5%), and incurs the lowest cost per validated inconsistency ($11.12). These results establish direct specification auditing as a practical complement to behavioral alignment evaluation, catching defects at the source before they shape any model. The code is available at https://github.com/HIPREL-Group/VeriSpec.

摘要:模型規範定義了大型語言模型(LLMs)應該如何運作,指導對齊訓練、推理時的行為和評估。然而,這些規範本身可能包含缺陷:兩個各自合理的原則在應用於相同情境時可能會規定不相容的行為,導致沒有任何回應能同時滿足兩者。檢測這種不一致性是具有挑戰性的。將自然語言規範形式化可能會失去微妙的區別,而基於行為的測試則無法可靠地區分規範缺陷與模型行為的差異。我們引入了VeriSpec,這是第一種通過審核規範文本本身直接檢測模型規範中不一致性的方法。我們的關鍵見解是保留自然語言中的規範,同時使用LLM作為驗證者。VeriSpec提取結構化的、上下文感知的規則,構建主題引導的圖以聚類同一權威級別下行為相關的規則,並應用LLM作為驗證者的推理來檢測不一致性。將VeriSpec應用於OpenAI模型規範,我們提取了405條規則並手動驗證了五個不一致性,所有這些都已報告給其開發者,開發者對此做出了積極回應並已啟動內部討論。與五個基準相比,VeriSpec識別了最多的經過驗證的不一致性,達到了最高的精確度(38.5%),並且每個經過驗證的不一致性的成本最低($11.12)。這些結果確立了直接規範審核作為行為對齊評估的實用補充,在缺陷影響任何模型之前,及時捕捉到缺陷。代碼可在 https://github.com/HIPREL-Group/VeriSpec 獲得。

Beyond Decodability: Do Acoustic Factors Drive Predictions in Speech-Based Alzheimer's Assessment?

2610.01846v1 by Serli Kopar, Alkis Koudounas, Roshan P. Rane, Sam Gijsen, Paula A. Perez-Toro, Kerstin Ritter

Speech-based Alzheimer's disease (AD) assessments increasingly rely on pretrained self-supervised learning (SSL) models that learn acoustic representations directly from raw audio, exposing the model to recording factors. We ask whether such factors are merely encoded in SSL representations or can systematically alter predictions. Using ADReSSo and three large SSL backbones, we apply controlled noise and reverberation interventions to participant-speech-only, non-speech, and full-recording audio. We combine layer-wise linear decoding, input- and representation-space interventions, and geometric alignment analysis to distinguish acoustic decodability from influence on AD prediction. Our results show that controlled acoustic interventions alter AD predictions across all three SSL backbones. Noise, despite showing no significant diagnostic-group difference in the original data, produces the strongest intervention effects. Importantly, these effects are systematically structured relative to the classifier's decision direction, replicate on the held-out test set and reverse when the representation-space intervention direction is reversed. Together, these findings show that high predictive performance and the absence of a significant diagnostic-group difference in a measured acoustic factor are not sufficient for robustness. We argue that intervention-based robustness tests should become standard for trustworthy clinical speech models.

摘要:基於語音的阿茲海默症(AD)評估越來越依賴於預訓練的自我監督學習(SSL)模型,這些模型直接從原始音頻中學習聲學表示,讓模型接觸到錄音因素。我們詢問這些因素是否僅僅被編碼在SSL表示中,或是否可以系統性地改變預測。使用ADReSSo和三個大型SSL骨幹,我們對參與者的語音、非語音和完整錄音音頻應用控制噪音和混響干預。我們結合層級線性解碼、輸入和表示空間干預,以及幾何對齊分析,以區分聲學可解碼性與對AD預測的影響。我們的結果顯示,控制的聲學干預改變了所有三個SSL骨幹的AD預測。儘管在原始數據中未顯示出顯著的診斷組差異,噪音卻產生了最強的干預效果。重要的是,這些效果相對於分類器的決策方向系統性地結構化,在保留的測試集中重現,並在表示空間干預方向反轉時逆轉。總的來說,這些發現表明,高預測性能和在測量的聲學因素中缺乏顯著的診斷組差異並不足以保證穩健性。我們主張,基於干預的穩健性測試應成為可信臨床語音模型的標準。

On the Divergence of Accuracy and Mechanism Consistency in Time Series World Models

2610.01842v1 by Haochen Zhang, Jiaheng Guo, Zhen Xu, Zachary Plotkin, Nicholas Konz, Zhen Tan, Tianlong Chen

A time series world model (TSWM) predicts a controlled system's state from its observed history and planned actions and exogenous inputs. Current approaches build forecasters with actions as covariates, trained and evaluated on prediction error under the executed plan. Yet world models compare unexecuted plans, but their responses to changed plans remain untested. We ask which design choices matter and whether accurate forecasters respond to changed plans as real systems do. We address both with a formalization and benchmark. The formalization separates state, actions and exogenous inputs, distinguishes continuous, mode and event actions, and introduces mechanism consistency, a metric built on declared action-state relations with known directions, such as a vasopressor raising blood pressure: it checks whether shifting an action moves the forecast in the declared direction. The benchmark consolidates eight public datasets with real actions from engineered infrastructure and clinical care, varying prediction space, plan fusion and plan encoding across seven backbones and five seeds. First, a frozen latent prediction space lowers MAE by 9.9% over observation space and gated output fusion lowers it by 12.7% over input concatenation on average, with both improving all eight datasets; temporal plan encoding changes average MAE by at most 2.2%. Second, prediction error and mechanism consistency diverge: the lowest-error configuration is at or below chance in consistency on four of five datasets with declared mechanisms, and no design choice avoids this. Finally, directional supervision, a loss penalizing the wrong-signed part of the response to a shifted action, significantly raises consistency on penalized mechanisms with no change in MAE. Together they give TSWMs a recipe: a frozen latent space and output-side fusion for accuracy, and a training objective for mechanism consistency.

摘要:時間序列世界模型 (TSWM) 從其觀察歷史、計畫行動和外部輸入預測受控系統的狀態。當前的方法建立了以行動為協變數的預測器,這些預測器在執行計畫下的預測誤差上進行訓練和評估。然而,世界模型比較未執行的計畫,但對於變更計畫的反應仍未經測試。我們詢問哪些設計選擇是重要的,以及準確的預測器是否像真實系統一樣對變更計畫做出反應。我們通過形式化和基準來解決這兩個問題。形式化將狀態、行動和外部輸入分開,區分連續、模式和事件行動,並引入機制一致性,這是一種基於已知方向的聲明行動-狀態關係構建的指標,例如一種升壓藥提高血壓:它檢查改變行動是否將預測移動到聲明的方向。基準整合了八個來自工程基礎設施和臨床護理的公共數據集,這些數據集中有真實行動,並在七個骨幹和五個種子中變化預測空間、計畫融合和計畫編碼。首先,凍結的潛在預測空間使 MAE 在觀察空間上降低了 9.9%,而門控輸出融合使其在輸入串接上平均降低了 12.7%,兩者都改善了所有八個數據集;時間計畫編碼的變化使平均 MAE 變化最多為 2.2%。其次,預測誤差和機制一致性出現分歧:在五個具有聲明機制的數據集中,最低誤差配置在一致性上與隨機相同或更低,且沒有任何設計選擇能避免這一點。最後,方向性監督,一種對於對移動行動的反應中錯誤符號部分進行懲罰的損失,顯著提高了懲罰機制的一致性,且 MAE 沒有變化。這些共同為 TSWM 提供了一個配方:凍結的潛在空間和輸出側融合以提高準確性,以及一個針對機制一致性的訓練目標。

Code Owns the Simulation, Jev Owns the Evaluation

2610.01834v1 by Yaodong Yang, Hongyao Tang, Yi Ma, Xingyu Fan, Weixun Wang, Jinpeng Li, Tianpei Yang

Judgment models such as \jev{} return, in a single call and without reasoning text, a probability for each described option. This makes them attractive as an agent's action-selection layer, but it is unclear which decisions they can be trusted with. We test \jev{} on reflection tests, one-shot matrix games, the text game ALFWorld and robot control, and find a sharp boundary. \jev{} succeeds when the right option can be judged from what the input describes, which we call \emph{evaluation}. Specifically, it solves 99\% of the counterintuitive Cognitive Reflection Test questions. However, it fails when the right option depends on \emph{simulation} (i.e., predicting something not in the input), such as the opponent's action or the subgoal that must come first. In games, \jev{} plays suboptimally as if its rational opponent acted at random, because the opponent's action is not given. In ALFWorld, \jev{} favors commands that mention an object or place named in the task description. For example, given the task ``put a clean knife in the drawer'', \jev{} carries an unwashed knife straight to the drawer instead of first washing it at the sink. Surprisingly, many of these failures are not due to a lack of knowledge. Asked separately what the opponent will do, \jev{} usually answers correctly, and it responds well given the opponent's action. It fails when one call must both perform the simulation and evaluate based on it. This suggests letting code make the prediction or simulation. When code supplies it, such as a lookahead in ALFWorld and physics simulation in robot control, \jev{} becomes an expert controller through its general evaluation ability.

摘要:判斷模型如 \jev{} 在單次呼叫中返回每個描述選項的概率,且不需要推理文本。這使得它們作為代理的行動選擇層變得具有吸引力,但尚不清楚它們可以信任哪些決策。我們在反思測試、一回合矩陣遊戲、文本遊戲 ALFWorld 和機器人控制上測試 \jev{},並發現了一個明確的邊界。當正確選項可以從輸入描述中判斷時,我們稱之為 \emph{評估},\jev{} 成功地解決了 99\% 的反直覺認知反思測試問題。然而,當正確選項依賴於 \emph{模擬}(即預測輸入中不存在的事物)時,它則失敗,例如對手的行動或必須先完成的子目標。在遊戲中,\jev{} 表現得次優,彷彿其理性的對手隨機行動,因為對手的行動並未給出。在 ALFWorld 中,\jev{} 偏好提到任務描述中物體或地點的命令。例如,給定任務「將乾淨的刀放入抽屜」,\jev{} 直接將未洗的刀帶到抽屜,而不是先在水槽中清洗它。令人驚訝的是,這些失敗中的許多並非因為缺乏知識。單獨詢問對手會做什麼時,\jev{} 通常能正確回答,並且在給定對手的行動時反應良好。當一次呼叫必須同時執行模擬並基於此進行評估時,它則失敗。這表明應讓代碼進行預測或模擬。當代碼提供這些信息時,例如在 ALFWorld 中的前瞻和在機器人控制中的物理模擬,\jev{} 通過其一般評估能力成為專家控制器。

Continuous Process-Level Evaluation for Evolving Enterprise AI Agent Skills

2610.01833v1 by Ngoc Phuoc An Vo, Aarya Doshi, Vadim Sheinin

Enterprise AI agent skills evolve as tool APIs, models, and specifications change, yet final-output evaluation can miss process-level behavioral drift. We present a continuous evaluation framework combining outcome-level and process-level checks, applied to Revenue and Productivity variants of a Business Value Determination skill in an enterprise Value Aware Resiliency system. The framework independently computes per-run ground truth, materializes reusable template tests, and evaluates tool selection, arguments, execution order, and database integrity through programmatic checks and a narrowly scoped LLM judge. We evaluate 240 trials across two skills, two specification variants, two agent harnesses, and three models. Of 175 trials passing all applicable final numerical checks, 162 (92.6 percent; Wilson 95 percent CI: 87.7-95.6 percent) contained another evaluator-detected deviation. Under a broader seven-check final-state definition, 151 of 164 passing runs (92.1 percent; 95 percent CI: 86.9-95.3 percent) still violated a trajectory check. Dependency attribution reduced a mean of 6.34 failed checks per run to 2.65 roots. Specification sensitivity varied by model and harness, with exploratory bootstrap interaction intervals excluding zero for all three Revenue comparisons and one of three Productivity comparisons. Runtime-resolved templates provided reusable regression coverage across the evaluated configurations; longitudinal validation under actual API evolution remains future work.

摘要:企業AI代理的技能隨著工具API、模型和規範的變化而演變,但最終輸出評估可能會忽略過程層級的行為漂移。我們提出了一個持續評估框架,結合了結果層級和過程層級的檢查,應用於企業價值感知彈性系統中商業價值判定技能的收入和生產力變體。該框架獨立計算每次運行的真實值,實現可重用的模板測試,並通過程式檢查和狹義範圍的LLM評判來評估工具選擇、參數、執行順序和數據庫完整性。我們在兩項技能、兩個規範變體、兩個代理框架和三個模型上評估了240次試驗。在175次通過所有適用的最終數值檢查的試驗中,162次(92.6%;Wilson 95% CI:87.7-95.6%)包含了另一個評估者檢測到的偏差。在更廣泛的七項檢查最終狀態定義下,164次通過的運行中有151次(92.1%;95% CI:86.9-95.3%)仍然違反了軌跡檢查。依賴性歸因將每次運行的平均失敗檢查數從6.34減少到2.65個根源。規範敏感性因模型和框架而異,探索性自助引導交互區間在所有三個收入比較和三個生產力比較中的一個中均不包括零。運行時解析的模板在評估的配置中提供了可重用的回歸覆蓋;在實際API演變下的縱向驗證仍然是未來的工作。

The Asymptotics of Language Model Alignment with Memory

2610.01828v1 by Haricharan Balasundaram, V. Arvind Rameshwar

Language model (LM) alignment broadly aims to perturb a given LM $Q$ into an aligned LM $q$ such that i) the outputs produced by $q$ and $Q$ are 'close' in probability, ii) $q$ has a higher expected reward than $Q$. Two common techniques for LM alignment are: KL-constrained RL, which requires knowledge of the LM distribution and is computationally expensive, and the best-of-$n$ algorithm, which requires only sampling from the LM. The work of Yang et al. established asymptotic closeness between the distributions produced by the two alignment methods for an $m$--length i.i.d. token sequence output by the LM, in the limit as $m$ increases to infinity. However, the i.i.d. assumption is not representative of practical LMs, whose output sequences often have memory. In this paper, we extend the asymptotic closeness result to the case when the $m$--length token sequence outputted by the LM is Markovian. Further, for finite-length output sequences -- particularly, when $m=1$ -- we provide a complete characterization of LM distributions and reward functions for which the KL-divergence between the distributions produced by the two alignment methods is zero -- a question first posed in Yang et al.

摘要:語言模型(LM)對齊的廣泛目標是將給定的 LM $Q$ 轉變為一個對齊的 LM $q$,使得 i) $q$ 和 $Q$ 所產生的輸出在概率上是「接近」的,ii) $q$ 的期望獎勵高於 $Q$。兩種常見的 LM 對齊技術是:KL 約束強化學習,這需要對 LM 分佈的了解並且計算上昂貴,以及最佳的 $n$ 算法,這僅需要從 LM 中進行取樣。Yang 等人的研究確立了在 $m$ 長度的獨立同分佈(i.i.d.)標記序列的情況下,兩種對齊方法所產生的分佈之間的漸近接近性,當 $m$ 增加到無限大時。然而,i.i.d. 假設並不代表實際的 LM,因為它們的輸出序列通常具有記憶性。在本文中,我們將漸近接近性結果擴展到 LM 輸出的 $m$ 長度標記序列為馬爾可夫過程的情況。此外,對於有限長度的輸出序列——特別是當 $m=1$ 時——我們提供了 LM 分佈和獎勵函數的完整特徵描述,對於這些情況,兩種對齊方法所產生的分佈之間的 KL 散度為零——這是一個最初由 Yang 等人提出的問題。

Token Communication-Assisted Collaborative Embodied Artificial Intelligence: Concepts, Framework, and Opportunities

2610.01826v1 by Peng Yi, Ying-Chang Liang

Collaborative embodied artificial intelligence (CEAI) enables multiple physical agents to perceive, reason, and act cooperatively in dynamic environments. Effective communication is essential for CEAI, yet CEAI agents must exchange not only large multimodal observations but also task-relevant insights, intents, and interactive information over long horizons. This article investigates token communication (TokCom) as a native intelligence interface for CEAI, in which tokens serve jointly as compact semantic carriers for communication and fundamental inference units for generative foundation models (GFMs). We first discuss how TokCom supports insight sharing, intent alignment, and interactive control among embodied agents. We then propose a TokCom-assisted CEAI framework driven by a task-adaptive communication protocol. Comprising a compact codebook, syntax rules, and contextual examples, this protocol guides GFM-based transceivers to distill messages into compact tokens and reconstruct them after wireless transmission. A case study on collaborative object transport demonstrates that the proposed TokCom framework substantially reduces the source payload bit consumption while preserving task efficiency and showing robustness under noisy channels. Finally, we outline future research directions.

摘要:協作具身人工智慧(CEAI)使多個物理代理能夠在動態環境中共同感知、推理和行動。有效的溝通對於CEAI至關重要,然而CEAI代理必須交換不僅是大量的多模態觀察,還包括與任務相關的見解、意圖和互動信息,並且這些交流需要在長時間範圍內進行。本文探討了作為CEAI本地智能介面的標記通信(TokCom),其中標記共同作為溝通的緊湊語義載體和生成基礎模型(GFMs)的基本推理單元。我們首先討論了TokCom如何支持具身代理之間的見解共享、意圖對齊和互動控制。接著,我們提出了一個由任務自適應通信協議驅動的TokCom輔助CEAI框架。該協議由緊湊的代碼本、語法規則和上下文示例組成,指導基於GFM的發射接收器將消息提煉為緊湊的標記,並在無線傳輸後重建它們。一個關於協作物體運輸的案例研究顯示,所提出的TokCom框架在保持任務效率和在噪聲通道下顯示穩健性的同時,顯著減少了源負載位元消耗。最後,我們概述了未來的研究方向。

Beyond Linear Concepts: Discovering and Aligning Non-Linear Concept Manifolds in Large Language Models

2610.01821v1 by Tido Specht, Elias Benedict Krey, Nils Neukirch, Nils Strodthoff

Understanding information processing in large language models (LLMs) requires dissecting the geometric organization of their internal token representations. While existing mechanistic interpretability (MI) methods seek to extract concepts, they are constrained by a strong linearity assumption challenged by evidence of non-linear feature manifolds. We move beyond linear concepts by adapting Non-Linear Multi-Dimensional Concept Discovery (NLMCD) from computer vision to token-level LLM activations, modeling concepts as low-dimensional manifolds. To compare concept manifolds across layers and models, we introduce a concept-based alignment (CBA) score, a generalized Rand index that measures geometric proximity without explicit feature matching. Our analysis yields six key findings: (i) a neighboring-layer sanity check shows CBA is more sensitive than PCA- or CKA-based linear baselines; (ii) layer-by-layer alignment matrices reveal two block structures in intermediate and late layers, consistent across models and obscured by linear metrics; (iii) concept composition remains syntax-dominated through most of the network before giving way to increasingly mixed syntactic-semantic concepts in later layers, with increasing output-orientation toward the final layers; (iv) multilingual concept sharing between English and Mandarin is training-dependent rather than universal, strongest in Qwen, weaker in Llama, and absent in GPT-2; (v) inter-model alignment mirrors this structure, with strong correspondence between same-family Qwen models of different scale but weak alignment across model families; and (vi) across Tulu-3 training stages, alignment is highest between adjacent stages, with the largest shift between the base model and SFT, while subsequent preference-alignment stages (DPO, RLVR) leave early layers largely unchanged and RLVR mostly preserves DPO's concepts in late layers.

摘要:理解大型語言模型(LLMs)中的信息處理需要剖析其內部標記表示的幾何組織。雖然現有的機械可解釋性(MI)方法試圖提取概念,但它們受到強線性假設的限制,而這一假設受到非線性特徵流形證據的挑戰。我們通過將非線性多維概念發現(NLMCD)從計算機視覺適應到標記級別的LLM激活,超越線性概念,將概念建模為低維流形。為了比較不同層和模型之間的概念流形,我們引入了一種基於概念的對齊(CBA)分數,這是一種廣義的Rand指數,用於測量幾何接近性而不需要明確的特徵匹配。我們的分析得出了六個關鍵發現:(i)相鄰層的合理性檢查顯示CBA比基於PCA或CKA的線性基準更敏感;(ii)逐層對齊矩陣揭示了中間層和後期層中的兩個區塊結構,這在不同模型中是一致的,但被線性指標所掩蓋;(iii)概念組合在網絡的大部分時間內仍然以語法為主,然後在後期層轉向越來越混合的語法-語義概念,並且對最終層的輸出取向逐漸增加;(iv)英語和普通話之間的多語言概念共享是依賴於訓練的,而不是普遍存在的,在Qwen中最強,在Llama中較弱,而在GPT-2中則不存在;(v)模型間的對齊反映了這一結構,同一家族的Qwen模型之間存在強對應,但不同模型家族之間的對齊較弱;(vi)在Tulu-3的訓練階段中,相鄰階段之間的對齊最高,基礎模型和SFT之間的變化最大,而隨後的偏好對齊階段(DPO、RLVR)使早期層幾乎保持不變,並且RLVR在後期層中大多保留了DPO的概念。

A Safe Prototype Is Not a Safety Direction: Reference Dependence and Prompt Confounds in Response-Safety Embeddings

2610.01801v1 by Sahil Kadadekar

Can response safety be scored by cosine similarity to the mean embedding of known-safe responses? A recent sleeper-agent detector proposes exactly this score, yet the raw positive-centroid rule is not identified: positive observations locate the safe class relative to an encoder origin, but do not determine which direction separates safe from unsafe responses. We audit the rule on two prompt-controlled, human-labeled corpora and one auxiliary jury-labeled source control, using four frozen encoders and prompt-grouped splits. On the human-labeled corpora the safe prototype reaches ROC-AUC 0.457-0.545, with two cells significantly below chance and one above, while an explicit safe-minus-unsafe reference reaches 0.588-0.738 on the same embeddings; on the jury control the prototype is inverted (0.358-0.405) and the reference reaches 0.754-0.793. At validation-calibrated 5% false-safe thresholds, the reference accepts more safe responses on PKU-SafeRLHF (0.153-0.263 versus 0.039-0.061 across encoders) and Aegis (0.189-0.291 versus 0.004-0.045), but not reliably on BeaverTails. A fully unlabeled held-out reference recovers part to most of the referenced ranking, much less when only 5% of the pool is unsafe, whereas 80-634 labeled unsafe responses recover most of it. Prompt-only ablations show that prompt-label composition can inflate uncontrolled evaluations. This is a bounded result about a raw positive centroid, not all one-class methods or safety-specialized guards. A class mean is a location, not necessarily a safety direction; a declared reference with enough unsafe mass identifies orientation.

摘要:可以通過與已知安全回應的平均嵌入的餘弦相似度來評分回應的安全性嗎?最近的潛伏特工檢測器正是提出了這一評分,但原始的正中心規則並未被識別:正觀察將安全類別定位於編碼器原點相對的位置,但並未確定哪個方向將安全回應與不安全回應分開。我們對兩個提示控制的人類標記語料庫和一個輔助陪審團標記的源控制進行了審核,使用四個凍結的編碼器和提示分組拆分。在人類標記的語料庫中,安全原型的ROC-AUC達到0.457-0.545,其中兩個單元顯著低於隨機機率,一個則高於隨機機率,而明確的安全減不安全參考在相同的嵌入上達到0.588-0.738;在陪審團控制中,原型被反轉(0.358-0.405),參考達到0.754-0.793。在驗證校準的5%假安全閾值下,參考在PKU-SafeRLHF上接受了更多的安全回應(0.153-0.263對比0.039-0.061,跨編碼器)和Aegis(0.189-0.291對比0.004-0.045),但在BeaverTails上並不可靠。一個完全未標記的保留參考恢復了部分到大多數的參考排名,當只有5%的池是不安全的時候,恢復的程度要小得多,而80-634個標記的不安全回應則恢復了大部分。僅提示的消融實驗顯示,提示標籤的組合可能會膨脹不受控制的評估。這是一個關於原始正中心的有限結果,而不是所有單類方法或安全專門防護的結果。類的平均值是一個位置,不一定是一個安全方向;一個擁有足夠不安全質量的聲明參考確定了方向。

LineupRL: Verifiable Reinforcement Learning for Time Series Captioning via Caption-to-Series Identification

2610.01800v1 by Haochen Zhang, Laura Yao, Zachary Plotkin, Gengwei Zhang, Tianlong Chen

Time series captioning is a fundamental step in time series understanding and can also serve as the bridge between signal and natural language. Supervised fine-tuning (SFT) relies on a larger model's captions and cannot exceed their quality. Reinforcement learning (RL) can, but its rewards were designed for other modalities and other tasks, and they transfer poorly to open-ended generation in the time series domain. We address this by proposing LineupRL, a reinforcement learning with verifiable rewards (RLVR) pipeline whose reward is caption-to-series identification. The reward model is a frozen large language model (LLM) verifier that reads the generated caption and the candidate time series as raw values, never the chart, and must pick the described time series from multiple distractors. Matching is a far lighter demand on the verifier than writing questions or judging a caption, so an off-the-shelf LLM can supply the reward. Across two captioning benchmarks, and on forecasting and reconstruction where the predictor sees only the caption, LineupRL outperforms SFT and RL baselines on every metric. The 3B vision language model (VLM) trained by LineupRL also outperforms, at 1/24 of the parameters, the 72B VLM whose captions the SFT baseline is distilled from. Our case study shows that LineupRL resists reward hacking, and that the captioner it trains both traces the trend and names the values at key points.

摘要:時間序列標註是時間序列理解中的基本步驟,並且可以作為信號與自然語言之間的橋樑。監督微調(SFT)依賴於更大模型的標註,並且無法超越它們的質量。強化學習(RL)可以做到,但其獎勵是為其他模態和其他任務設計的,並且在時間序列領域的開放式生成中轉移效果不佳。我們通過提出LineupRL來解決這個問題,這是一個具有可驗證獎勵的強化學習(RLVR)管道,其獎勵是標註到序列的識別。獎勵模型是一個凍結的大型語言模型(LLM)驗證器,它將生成的標註和候選時間序列作為原始值進行閱讀,而不是圖表,並且必須從多個干擾項中選擇所描述的時間序列。對驗證器的匹配要求比撰寫問題或評估標註輕得多,因此現成的LLM可以提供獎勵。在兩個標註基準上,以及在預測和重建中,當預測器僅看到標註時,LineupRL在每個指標上都超越了SFT和RL基準。由LineupRL訓練的3B視覺語言模型(VLM)在參數為1/24的情況下,也超越了72B VLM,該模型的標註是從SFT基準中提煉出來的。我們的案例研究顯示LineupRL抵抗獎勵操控,並且它訓練的標註者能夠追蹤趨勢並在關鍵點命名數值。

iADD: Improving Alignment and Diversity in Diffusion Policy Optimization

2610.01789v1 by Ashok Prasad Neupane, Saugat Adhikari, Pramish Paudel, Ajad Chhatkuli, Danda Pani Paudel

Reinforcement learning based post training of diffusion models, such as Denoising Diffusion Policy Optimization (DDPO), optimizes a reverse diffusion process under a reward function. However, current approaches to reward optimizations do so at the cost of diversity and quality. In this paper, we provide better tradeoffs through careful theoretical considerations and method design. We analyze the theoretical framework and mathematically demonstrate that \emph{only-latter timestep} updates of diffusion model may be harmful for diversity contrary to the conclusions presented in a previous work. Additionally, we propose an incremental Feynman-Kac training based on strong theoretical foundations in order to achieve the best-yet alignment-diversity tradeoffs. We perform extensive experiments and compare our method against related diffusion policy optimization approaches in three different tasks and also provide strong ablations for each component, thus validating strong performance gains in both alignment and diversity.

摘要:基於強化學習的擴散模型後訓練,例如去噪擴散策略優化(DDPO),在獎勵函數下優化反向擴散過程。然而,當前的獎勵優化方法以多樣性和質量為代價。在本文中,我們通過仔細的理論考量和方法設計提供了更好的權衡。我們分析了理論框架,並數學上證明擴散模型的\emph{僅後時間步}更新可能對多樣性有害,這與之前工作的結論相反。此外,我們提出了一種基於強大理論基礎的增量費曼-卡克訓練,以實現最佳的對齊-多樣性權衡。我們進行了廣泛的實驗,並在三個不同的任務中將我們的方法與相關的擴散政策優化方法進行比較,還為每個組件提供了強有力的消融實驗,從而驗證了在對齊和多樣性方面的顯著性能提升。

VETO: Video Efficient Token Optimization for Vision Language Models

2610.01785v1 by Gueter Josmy Faure, Hao Ping Wang, Min-Hung Chen, Winston H. Hsu

Processing long videos with Vision-Language Models (VLMs) is bottlenecked by the quadratic cost of visual tokens, making long-form inference prohibitively expensive. While single-axis compression methods mitigate this, they hit a hard efficiency floor because they treat spatial and temporal redundancy independently. We present VETO (Video Efficient Token Optimization for Vision-Language Models), a training-optional plug-in that eliminates this bottleneck through dual-axis compression: (i) an intra-frame compressor that merges semantically similar tokens within each frame via optimal-transport inspired matching, and (ii) an inter-frame compressor that identifies and merges temporally redundant frames. The key design insight is hierarchical ordering: by first compressing spatial dimensions, VETO drastically reduces the cost of subsequent global temporal matching, bypassing the efficiency wall of single-axis approaches, with an advantage that grows with modern fully-fused attention infrastructure. Empirically, VETO achieves up to 45% faster inference (e.g., on LLaVA-OneVision-7B) while preserving or improving accuracy. Under extreme token starvation (10% budget), VETO outperforms VFlowOpt (54.9%), VisionZip (52.6%), and FastV (47.9%) with 55.7% accuracy. We demonstrate universal applicability across LLaVA-OneVision, InternVL-2.5, and LongVA, with zero-shot accuracy preserved or improved in all cases.

摘要:處理長視頻的視覺語言模型(VLMs)受到視覺標記的二次成本限制,使得長格式推理變得過於昂貴。雖然單軸壓縮方法可以緩解這一問題,但由於它們獨立處理空間和時間冗餘,因此達到了一個硬效率底線。我們提出了 VETO(視頻高效標記優化器),這是一個可選的插件,通過雙軸壓縮消除了這一瓶頸:(i)一個幀內壓縮器,通過受最優運輸啟發的匹配合併每幀內語義相似的標記,以及(ii)一個幀間壓縮器,識別並合併時間上冗餘的幀。關鍵的設計見解是分層排序:通過首先壓縮空間維度,VETO 大幅降低了隨後全局時間匹配的成本,繞過了單軸方法的效率牆,並且隨著現代全融合注意力基礎設施的發展,這一優勢不斷增長。實證結果顯示,VETO 在保持或提高準確度的同時,實現了高達 45% 的推理加速(例如,在 LLaVA-OneVision-7B 上)。在極端標記短缺(10% 預算)下,VETO 的準確率為 55.7%,超過了 VFlowOpt(54.9%)、VisionZip(52.6%)和 FastV(47.9%)。我們展示了在 LLaVA-OneVision、InternVL-2.5 和 LongVA 上的普遍適用性,在所有情況下均保持或提高了零樣本準確度。

Q-Learning for Reachability in MEC-Free MDPs

2610.01781v1 by Lu-Chin Chang, Suguman Bansal

Reinforcement learning (RL) for reachability specifications is fundamental to sequential decision-making. Prior work establishes asymptotic convergence to optimal policies, but only through model-based methods that must explicitly estimate the transition probabilities of the underlying Markov Decision Process (MDP). We present Quasar, the first model-free algorithm with asymptotic guarantees for reachability on the fragment of MDPs free of non-terminal maximal end components (MECs), a building block to which every MDP reduces by the standard MEC quotient. Our algorithm follows the classical Q-learning approach, using temporal-difference updates to converge to an optimal policy without ever learning the transition probabilities. The resulting learner reduces the memory footprint from the O(|S|^2|A|) that model-based methods require to O(|S||A|). On the standardized Quantitative Verification Benchmark Set, our algorithm converges to the optimal policy with orders of magnitude fewer samples than the previous model-based state-of-the-art. Together these results are a concrete step toward the practical deployment of reachability learning and, with it, of specification-guided RL.

摘要:強化學習(RL)對於可達性規範是序列決策的基礎。先前的研究建立了對最佳政策的漸近收斂,但僅通過必須明確估計基礎馬爾可夫決策過程(MDP)轉移概率的基於模型的方法。我們提出了Quasar,這是第一個對於不含非終端最大端元組件(MECs)片段的MDP具有漸近保證的無模型算法,這是每個MDP通過標準MEC商減少的構建塊。我們的算法遵循經典的Q學習方法,使用時間差更新來收斂到最佳政策,而無需學習轉移概率。最終的學習器將基於模型的方法所需的O(|S|^2|A|)的內存佔用減少到O(|S||A|)。在標準化的定量驗證基準集上,我們的算法以比先前基於模型的最先進技術少幾個量級的樣本收斂到最佳政策。這些結果共同為可達性學習的實際部署邁出了具體的一步,並隨之推進了以規範為指導的RL。

CODesign: Consistency from Data to Trajectory in All-Atom Protein Binder Co-Design

2610.01773v1 by Yuanle Mo, Bo Qiang, Haitao Lin, Qinghan Wang, Gang Du, Odin Zhang, Pheng Ann Heng

The central challenge in de novo protein design is generating plausible, mutually compatible structures and sequences, such that each designed sequence folds into its intended structure and the structure accommodates that sequence. Compared to typical two-stage design methods, which decouple the modeling of the interdependent modalities, co-design models improve the cross-modal consistency by jointly generating sequences and structures. However, naively generating sequences and structures simultaneously does not ensure their consistency. To address this challenge, we propose CODesign framework. We improve data consistency by generating approximately 105,000 consistency-distilled dimers. We further promote consistency through a multimodal joint flow model that captures the joint distribution of sequences, backbone structures, and local atomic configurations, together with a consistency-aware joint resampling strategy that iteratively refines sequences and side chains. Experiments show that CODesign achieves state-of-the-art performance with the highest in silico success rates on both protein- and ligand-target binder design. Ablation studies also demonstrate our distilled dataset increases performance by 70.9%, which can be further improved by our proposed resampling mechanism with negligible additional computational cost. Code, model weights and the new dataset will be completely open-source.

摘要:中心挑戰在於全新蛋白質設計中生成合理且相互兼容的結構和序列,使得每個設計的序列能夠摺疊成其預期的結構,並且該結構能夠容納該序列。與典型的兩階段設計方法相比,這些方法將相互依賴的模態建模分開,協同設計模型通過共同生成序列和結構來改善跨模態的一致性。然而,天真地同時生成序列和結構並不能確保它們的一致性。為了解決這一挑戰,我們提出了CODesign框架。我們通過生成約105,000個一致性提煉的二聚體來改善數據一致性。我們進一步通過一個多模態聯合流模型來促進一致性,該模型捕捉序列、主鏈結構和局部原子配置的聯合分佈,並結合一個一致性感知的聯合重採樣策略,該策略迭代地精煉序列和側鏈。實驗表明,CODesign在蛋白質和配體靶向結合物設計上達到了最先進的性能,並且在計算機模擬成功率上達到了最高。消融研究也顯示我們的提煉數據集使性能提高了70.9%,而且可以通過我們提出的重採樣機制進一步改善,且額外的計算成本微乎其微。代碼、模型權重和新數據集將完全開源。

A Matryoshka Hierarchical RAG for Efficient Multi-Hop Question Answering

2610.01767v1 by Gianluca Bonifazi, Christopher Buratti, Michele Marchetti, Federica Parlapiano, Giulia Quaglieri, Davide Traini, Domenico Ursino, Luca Virgili

Retrieval-Augmented Generation (RAG) systems for multi-hop Question Answering (QA) must balance retrieval quality with computational cost. This cost is incurred during indexing time, through the use of expensive Knowledge Graphs (KGs) or Large Language Models (LLMs) to generate summaries, or during querying, through iterative LLM-driven retrieval. To reduce it while maintaining retrieval quality, we present MatRAG, a hierarchical framework that combines RAG systems with Matryoshka Representation Learning (MRL). MatRAG addresses both kinds of cost by aligning the semantic hierarchy of a clustering structure with the nested structure of MRL. Specifically, it organizes the corpus of documents into a Directed Acyclic Graph (DAG) of clusters with progressively coarser granularity. Each level is indexed by a lower Matryoshka dimension. MatRAG pairs an iterative, top-down traversal of the DAG with an entity-driven mechanism that controls the hop budget and re-ranks candidates. We evaluated MatRAG on three standard multi-hop QA benchmarks against seven representative baselines. MatRAG outperforms its strongest competitors in terms of retrieval quality; furthermore, it reduces indexing costs by avoiding KG construction and LLM-based summarization, and lowers query-time costs through dimension-aware similarity.

摘要:檢索增強生成(RAG)系統在多跳問題回答(QA)中必須平衡檢索質量與計算成本。這個成本在索引時產生,通過使用昂貴的知識圖譜(KG)或大型語言模型(LLM)來生成摘要,或在查詢時,通過迭代的LLM驅動檢索。為了在保持檢索質量的同時降低成本,我們提出了MatRAG,一個將RAG系統與馬特里奧什卡表示學習(MRL)相結合的分層框架。MatRAG通過將聚類結構的語義層次與MRL的嵌套結構對齊,解決了這兩種成本。具體而言,它將文檔語料庫組織成一個具有逐漸粗糙粒度的有向無環圖(DAG)聚類。每個層級由較低的馬特里奧什卡維度進行索引。MatRAG將DAG的迭代自上而下遍歷與一種驅動實體的機制相結合,該機制控制跳躍預算並重新排名候選者。我們在三個標準的多跳QA基準上評估了MatRAG,並與七個代表性的基準進行比較。MatRAG在檢索質量方面超越了其最強的競爭對手;此外,它通過避免KG構建和基於LLM的摘要來降低索引成本,並通過維度感知相似性來降低查詢時間成本。

VideoEvolve: Evolving Agent Harnesses for Video Temporal Grounding

2610.01766v1 by Bingjun Luo, Yuhuan Fan, Jialin Guo, Siqi Li

Video temporal grounding aims to localize events in videos from natural-language queries. For agents built around frozen video-language models, the harness determines how queries guide temporal predictions and how those predictions are refined. Manually refining these harnesses requires diagnosing grounding failures and coordinating changes to both agent workflows and instructions. We introduce VideoEvolve, a framework that automatically evolves agent harnesses for video temporal grounding. VideoEvolve uses a Cloze-Structured Harness Representation that preserves stage interfaces while leaving agent workflows and instructions open to evolution. Branch-Guided Harness Evolution preserves promising code branches for continued refinement, using execution feedback to guide local edits and validation to determine which improvements are carried forward. Experiments demonstrate improved grounding performance across multiple benchmarks. Component analyses identify instruction refinement as a consistent source of gains, while the benefits of evolved code vary across evaluation settings. Together, these results support automated harness evolution as an effective approach to improving video temporal grounding. Code is available at https://github.com/bingjunluo/VideoEvolve .

摘要:視頻時間定位旨在從自然語言查詢中定位視頻中的事件。對於基於凍結視頻-語言模型構建的代理,這個工具決定了查詢如何指導時間預測以及這些預測如何被精煉。手動精煉這些工具需要診斷定位失敗並協調對代理工作流程和指令的變更。我們介紹了VideoEvolve,一個自動演化代理工具以進行視頻時間定位的框架。VideoEvolve使用一種克洛茲結構工具表示法,保留了階段接口,同時使代理工作流程和指令保持開放以便演化。分支引導的工具演化保留了有前景的代碼分支以便持續精煉,利用執行反饋來指導局部編輯,並通過驗證來確定哪些改進將被保留。實驗顯示在多個基準上改進了定位性能。組件分析確定指令精煉是一個一致的增益來源,而演化代碼的好處在不同的評估設置中有所不同。總體而言,這些結果支持自動化工具演化作為改善視頻時間定位的有效方法。代碼可在 https://github.com/bingjunluo/VideoEvolve 獲得。

TopK-Guided: Adaptive, Budget-Aware Activation Sparsity for Efficient LLM Inference

2610.01763v1 by Mukund Agarwalla, Chih-Jen Lin

Activation sparsity speeds up large language model (LLM) inference by setting unimportant activations to zero so that the corresponding computations can be skipped. Existing training-free methods, however, make different trade-offs: threshold-based methods such as TEAL adapt the sparsity level to each token but do not tightly control the realised sparsity, while TopK-based methods such as WINA enforce a fixed sparsity level but use the same sparsity budget for every token. Both also apply the same budget across transformer blocks, despite large differences in block sensitivity. We introduce TopK-Guided, a training-free method that addresses both limitations by combining bounded token-level sparsity adaptation with sensitivity-aware block-level budget allocation. Across Llama-2 and Llama-3 models, TopK-Guided consistently improves perplexity and downstream accuracy over TEAL and WINA while preserving essentially the same sparsitydependent projection compute as WINA, with the largest gains at high sparsity. Ablations show that both components provide complementary improvements.

摘要:激活稀疏性透過將不重要的激活設置為零來加速大型語言模型(LLM)的推理,以便可以跳過相應的計算。然而,現有的無需訓練的方法則做出了不同的權衡:基於閾值的方法如TEAL將稀疏性水平適應於每個標記,但並未嚴格控制實現的稀疏性,而基於TopK的方法如WINA則強制執行固定的稀疏性水平,但對每個標記使用相同的稀疏性預算。儘管區塊敏感性存在很大差異,但兩者在Transformer區塊上也應用了相同的預算。我們介紹了TopK-Guided,這是一種無需訓練的方法,通過將有界的標記級稀疏性適應與敏感性意識的區塊級預算分配相結合,解決了這兩個限制。在Llama-2和Llama-3模型中,TopK-Guided始終在保持與WINA基本相同的稀疏性依賴投影計算的同時,顯著提高了對TEAL和WINA的困惑度和下游準確性,在高稀疏性下獲得了最大的增益。消融實驗顯示,這兩個組件提供了互補的改進。

SoK: Decentralized Agent Economic Infrastructure

2610.01756v1 by Rui Sun, Xihan Xiong, Qin Wang, Fei Gao, Zelin Li, Zehua Cheng, Jiahao Sun, Zhipeng Wang

Decentralized agent economies increasingly build a single task from protocols that were designed and secured separately. This creates a simple problem: a workflow can look correct at each step and still produce the wrong outcome. For example, a correct escrow may release payment on an authorized approval that provides little evidence that the delivered work actually satisfied the task. We systematize this problem across the full lifecycle of an agent task. Our study organizes security and economic requirements into 17 property families over six stages, with receipt soundness and completeness assessed separately. We examine 12 systems and standards, five reusable mechanism families, and four classical baselines. We introduce guarantee closure, a task-relative criterion for determining whether guarantees established at one stage remain available and constrain the later decisions that depend on them. We apply the criterion to controlled and native workflows, covering 840 matched executions and an exhaustive 11,648-case check over a finite objective-task domain. Our results expose recurring failures between verification and settlement, where conforming work can remain unaccepted or valid evidence can be ignored. Public records and model judgments further distinguish recorded approval from evidence of task conformance, while economic analysis identifies the report, penalty, and shared-error assumptions behind these guarantees. These findings show where end-to-end guarantees fail and what must be repaired to preserve them across the workflow.

摘要:去中心化的代理經濟越來越多地從單獨設計和保護的協議中構建單一任務。這造成了一個簡單的問題:工作流程在每一步看起來都正確,但仍然可能產生錯誤的結果。舉例來說,一個正確的保管協議可能會在授權批准下釋放付款,但該批准幾乎沒有證據表明交付的工作實際上滿足了任務要求。 我們在代理任務的整個生命周期中系統化這個問題。我們的研究將安全和經濟要求組織成17個屬性家族,分為六個階段,並分別評估收據的健全性和完整性。我們檢查了12個系統和標準、五個可重用的機制家族以及四個經典基準。我們引入了保證閉合,這是一個相對於任務的標準,用於確定在一個階段建立的保證是否仍然可用,並約束依賴於它們的後續決策。 我們將該標準應用於受控和本地工作流程,涵蓋840次匹配執行以及在有限的目標任務域內進行的11,648個案例的徹底檢查。我們的結果揭示了驗證與結算之間的重複失敗,在這些情況下,符合要求的工作可能仍然未被接受,或者有效證據可能被忽視。公共記錄和模型判斷進一步區分了記錄的批准與任務符合性的證據,而經濟分析則確定了這些保證背後的報告、懲罰和共享錯誤假設。這些發現顯示了端到端保證失效的地方,以及為了在工作流程中保留這些保證必須修復的內容。

Cog-VADU: A Training-Free Cognitive Reasoning Framework for Video Anomaly Detection and Understanding

2610.01754v1 by Mohd Ubaid Wani, Sara Atito, Josef Kittler, Muhammad Awais

Video Anomaly Detection (VAD) aims to temporally localize abnormal events in videos. Most existing approaches rely on dataset-specific training and curated annotations, limiting generalization in open-set scenarios. Recent zero-shot methods based on Large Vision- Language Models (LVLMs) alleviate this dependency but often lack temporal continuity and structured reasoning. We propose Cog-VADU, a fully training-free framework that reformulates VAD as a sequential cognitive reasoning task. Cog-VADU introduces Chain-of- Anomaly Detection Thought Prompting (CoADTP), which unrolls an LVLM into a recurrent reasoning chain across video segments. By propagating structured rationales over time, the model maintains implicit temporal memory, enabling robust discrimination between com- plex anomalies and high-motion normal activities. To improve reliability, we further design a cross-modal re-ranking stage that aligns textual rationales with visual embeddings, enforcing semantic consistency and temporal coherence for refined and stable predictions. Extensive experiments on multiple public VAD benchmarks demonstrate that Cog-VADU achieves competitive zero-shot performance. Moreover, cross-model evaluations show that CoADTP consistently enhances reasoning-based anomaly detection in a model-agnostic manner, pro- viding interpretable and generalizable anomaly understanding for real-world applications.

摘要:視頻異常檢測(VAD)旨在時間上定位視頻中的異常事件。大多數現有的方法依賴於特定數據集的訓練和精心策劃的標註,限制了在開放集場景中的泛化能力。最近基於大型視覺-語言模型(LVLMs)的零樣本方法減輕了這一依賴,但通常缺乏時間連續性和結構化推理。我們提出了Cog-VADU,一個完全無需訓練的框架,將VAD重新定義為一個序列認知推理任務。Cog-VADU引入了異常檢測思維提示鏈(CoADTP),將LVLM展開為跨視頻片段的遞歸推理鏈。通過隨時間傳播結構化的推理,該模型維持隱式的時間記憶,使其能夠在複雜異常和高運動正常活動之間進行穩健的區分。為了提高可靠性,我們進一步設計了一個跨模態重新排序階段,將文本推理與視覺嵌入對齊,強化語義一致性和時間連貫性,以實現精細和穩定的預測。在多個公共VAD基準上的廣泛實驗表明,Cog-VADU實現了具有競爭力的零樣本性能。此外,跨模型評估顯示,CoADTP始終以模型無關的方式增強基於推理的異常檢測,為現實世界應用提供可解釋和可泛化的異常理解。

Removing spurious minima for planar features by skip connections

2610.01728v1 by Jakob Paul Zimmermann, Moritz Grillo, Andrei Balakin, Georg Loho

Understanding loss landscapes is central to explaining neural-network training, yet their structure remains only partially understood even in simple models. We study the Gaussian population loss of shallow, bias-free ReLU networks in the teacher--student setting. This provides a simple model for studying essential aspects such as feature learning and overparameterization. For teacher networks with positive output weights and planar features, we show that including a learned linear skip removes all spurious local minima with non-negative student output weights once the student network is at least as wide as the teacher network. In contrast, without the skip, we construct a fixed teacher network with positive output weights and only three hidden neurons in input dimension two whose spurious local minima persist at every student width at least three. Thus, a learned linear skip can remove spurious minima that persist under arbitrary overparameterization. Furthermore, we show that a positive output weight student network always learns the subspace spanned by the teacher features: student features at local minima with non-negative student output weights lie in the span of the teacher features. For ReLU networks in two dimensions, even heavily overparameterized student networks have effective width controlled by the teacher width: every critical point with positive student output weights has at most twice as many distinct student feature directions as teacher neurons. Finally, we transfer the benignity result to empirical minima over parameter balls of any prescribed radius, with the required sampling accuracy depending on that radius.

摘要:理解損失景觀對於解釋神經網絡訓練至關重要,然而即使在簡單模型中,它們的結構仍然只有部分被理解。我們研究教師-學生設置中淺層、無偏的ReLU網絡的高斯族群損失。這提供了一個簡單的模型來研究如特徵學習和過度參數化等基本方面。對於具有正輸出權重和平面特徵的教師網絡,我們顯示包含學習的線性跳過可以消除所有具有非負學生輸出權重的虛假局部最小值,只要學生網絡的寬度至少與教師網絡一樣寬。相反,如果不使用跳過,我們構造了一個固定的教師網絡,其具有正輸出權重且在輸入維度為二的情況下只有三個隱藏神經元,這樣的虛假局部最小值在每個學生寬度至少為三的情況下持續存在。因此,學習的線性跳過可以消除在任意過度參數化下持續存在的虛假最小值。此外,我們顯示具有正輸出權重的學生網絡總是學習由教師特徵所跨越的子空間:在具有非負學生輸出權重的局部最小值下,學生特徵位於教師特徵的跨度內。對於二維的ReLU網絡,即使是高度過度參數化的學生網絡,其有效寬度也受到教師寬度的控制:每個具有正學生輸出權重的臨界點最多有教師神經元的兩倍不同學生特徵方向。最後,我們將良性結果轉移到任何指定半徑的參數球上的經驗最小值,所需的取樣精度取決於該半徑。

vFedProtoQNAS: Prototype-Guided Personalized Quantum Neural Architecture Search for Virtual Federated Learning

2610.01718v1 by Seok Bin Son, Samuel Yen-Chi Chen, Soohyun Park, Joongheon Kim

Quantum federated learning (QFL) has emerged as a promising approach for collaboratively training compact quantum neural networks (QNNs) over distributed private data on resource-constrained devices. However, differences in device capabilities make a single shared QNN architecture unsuitable for all clients. While personalized quantum neural architecture search (QNAS) allows each client to select a device-specific QNN, averaging parameters across structurally different QNN architectures mixes semantically inconsistent circuit operations. To address this, prototype-guided personalized QNAS for virtual FL (vFedProtoQNAS) is proposed, where model parameters are never aggregated across clients and federated collaboration is achieved through class-wise prototype sharing. Each client independently searches and trains a client-specific QNN, computes class-wise local prototypes from latent representations, and refines them using global prototypes from the server as federated semantic anchors. Experiments demonstrate that vFedProtoQNAS improves accuracy by 3.70\% over FedAvg and enhances class-consistent representation alignment.

摘要:量子聯邦學習(QFL)已成為一種有前景的方法,用於在資源有限的設備上協作訓練緊湊的量子神經網絡(QNNs),以處理分散的私有數據。然而,設備能力的差異使得單一共享的QNN架構不適合所有客戶端。雖然個性化量子神經架構搜索(QNAS)允許每個客戶端選擇特定於設備的QNN,但在結構上不同的QNN架構之間平均參數會混合語義不一致的電路操作。為了解決這個問題,提出了針對虛擬聯邦學習(vFedProtoQNAS)的原型引導個性化QNAS,在這裡模型參數從不在客戶端之間聚合,聯邦協作是通過類別級原型共享來實現的。每個客戶端獨立搜索和訓練特定於客戶端的QNN,從潛在表示中計算類別級本地原型,並使用來自伺服器的全局原型進行精煉,作為聯邦語義錨點。實驗表明,vFedProtoQNAS的準確率比FedAvg提高了3.70\%,並增強了類別一致的表示對齊。

CoEvolve: Construct-to-Edit Visual Grounding with Bidirectional State Refinement

2610.01710v1 by Dongwei Sun, Yujie Zhang, Bowen Yao, Pei Liu, Jing Yao, Xiangyong Cao

Visual grounding localizes an object described by language with a bounding box. Most multimodal grounding models compress target identification, spatial reasoning, and boundary estimation into one terminal prediction. Free-form rationales make reasoning linguistically explicit but do not necessarily expose measurable, editable spatial states. Intermediate localization errors are therefore difficult to diagnose and correct, allowing incorrect region choices and imprecise boundaries to persist in the final box. We introduce CoEvolve, a construct-to-edit framework that separates grounding into explicit state construction and state editing. Region-Evolution Reinforcement (RER) organizes grounding analysis into a progressive semantic--spatial trajectory, with each reasoning step committing to an explicit candidate region. Bidirectional Denoising Refiner (BDR) treats the reasoning text as fixed semantic context and refines the trajectory's coordinate fields through bidirectional same-position reconstruction. Geometry- and behavior-level objectives provide target geometry and edit-preference signals for consolidating reliable candidates, preserving accurate inputs, or correcting toward annotations. Evaluations cover natural-image and remote-sensing grounding. With a 9B backbone, CoEvolve rivals models up to 241B parameters in grounding accuracy. Under controlled corruption, a single BDR pass improves mean box overlap by over 27 percentage points, demonstrating strong recovery from substantial localization errors. State-source comparisons further support the complementarity of explicit state construction and source-matched editing. The project is at https://sundongwei.github.io/CoEvolve_Project/.

摘要:視覺基礎將語言描述的物體定位於邊界框內。大多數多模態基礎模型將目標識別、空間推理和邊界估計壓縮為一個終端預測。自由形式的推理使推理在語言上變得明確,但不一定揭示可測量、可編輯的空間狀態。因此,中間定位錯誤難以診斷和修正,導致不正確的區域選擇和不精確的邊界在最終框中持續存在。我們介紹了 CoEvolve,一個構建-編輯框架,將基礎分為明確的狀態構建和狀態編輯。區域演化強化(RER)將基礎分析組織成一個漸進的語義-空間軌跡,每一步推理都承諾於一個明確的候選區域。雙向去噪精煉器(BDR)將推理文本視為固定的語義上下文,並通過雙向同位置重建來精煉軌跡的坐標場。幾何和行為層面的目標提供了目標幾何和編輯偏好信號,以鞏固可靠的候選者,保留準確的輸入或朝向註釋進行修正。評估涵蓋自然影像和遙感基礎。在 9B 的主幹下,CoEvolve 在基礎準確性上與高達 241B 參數的模型相媲美。在受控損壞下,單次 BDR 通過提高平均框重疊超過 27 個百分點,顯示出從重大定位錯誤中強有力的恢復。狀態源比較進一步支持明確狀態構建和源匹配編輯的互補性。該項目位於 https://sundongwei.github.io/CoEvolve_Project/。

Task-Oriented Rank Adaptation for Continual Learning in Text Classification

2610.01702v1 by Rey Sanchez Lopez, Eduardo Morales Manzanares, Hugo Jair Escalante

Continual learning (CL) in text classification faces two critical challenges: catastrophic forgetting and negative transfer across sequential tasks. Parameter-Efficient Fine-Tuning (PEFT) methods such as LoRA enable efficient adaptation by learning low-rank updates of the model parameters. However, these compact representations are normally trained in isolation, limiting their reuse across related tasks. We introduce Task-Oriented Rank Adaptation (TORA), a geometric routing framework that leverages the low-rank structure of LoRA adapters to decide whether to transfer knowledge from the most compatible expert (Boosting) or isolate the new task (Shielding) based on structural similarity. Evaluated across 15 diverse text classification benchmarks, TORA consistently avoids harmful routing decisions: compatible tasks exceed their isolated performance while reducing training time, and structurally distant tasks are protected from interference with no loss in accuracy. With a single geometric threshold and no reliance on task identities or predefined sequences, TORA provides a simple and effective approach for dynamic adapter routing in sequential text classification systems.

摘要:持續學習(CL)在文本分類中面臨兩個關鍵挑戰:災難性遺忘和在序列任務中的負轉移。參數高效微調(PEFT)方法如 LoRA 通過學習模型參數的低秩更新來實現高效適應。然而,這些緊湊的表示通常是在孤立的情況下訓練的,限制了它們在相關任務中的重用。我們引入了任務導向秩適應(TORA),這是一個幾何路由框架,利用 LoRA 適配器的低秩結構來決定是從最兼容的專家(提升)轉移知識,還是根據結構相似性隔離新任務(保護)。在 15 個不同的文本分類基準上進行評估,TORA 始終避免有害的路由決策:兼容任務的表現超過其孤立的性能,同時減少訓練時間,而結構上相距較遠的任務則受到保護,沒有準確度損失。TORA 以單一的幾何閾值運作,且不依賴於任務身份或預定序列,為序列文本分類系統中的動態適配器路由提供了一種簡單而有效的方法。

Acmite: Mitigating Gender Bias in LLMs through Concept-Guided Mutual Information

2610.01696v1 by Tian Lan, Xiaoqing Cheng, Han Zhang, Jiang Li

Large language models (LLMs) can reproduce social stereotypes from their training data, motivating extensive research on model debiasing. However, existing methods often rely on explicit biased examples or predefined group-term substitutions, making them sensitive to wording and less effective at capturing stereotype concepts shared across diverse contexts. More importantly, they typically suppress biased outputs without explicitly modeling the statistical dependence between model outputs and the underlying stereotype concepts. We propose Acmite, a lightweight concept-guided framework for targeted and selective debiasing. Acmite represents stereotypes as structured semantic concepts and uses maximal marginal relevance (MMR) to select diverse concepts for debiasing. Inspired by mutual information minimization, it approximates this dependence with token-level KL divergence while preserving task semantics. A lightweight LoRA adapter is trained with the base model frozen and activated at inference time only when the input is sufficiently similar to stereotype-related concepts; otherwise, the original model is used directly. We evaluate Acmite on BBQ, CrowS-Pairs, and StereoSet, and assess general capability preservation on ARC-Challenge, GSM8K, and PIQA. Experiments across three LLMs show that Acmite effectively mitigates gender bias across complementary evaluation formats while maintaining competitive performance on bias-unrelated tasks. Anonymous code and data are available at https://anonymous.4open.science/r/Acmite-18E2/.

摘要:大型語言模型(LLMs)可以從其訓練數據中再現社會刻板印象,這促使了對模型去偏見的廣泛研究。然而,現有的方法通常依賴於明確的偏見示例或預定義的群體術語替代,這使得它們對措辭敏感,並且在捕捉跨多樣背景的刻板印象概念時效果不佳。更重要的是,它們通常會抑制偏見輸出,而不明確建模模型輸出與潛在刻板印象概念之間的統計依賴。我們提出了Acmite,一種輕量級的概念引導框架,用於有針對性和選擇性的去偏見。Acmite將刻板印象表示為結構化的語義概念,並使用最大邊際相關性(MMR)來選擇多樣的概念進行去偏見。受到互信息最小化的啟發,它通過標記級的KL散度來近似這種依賴,同時保留任務語義。一個輕量級的LoRA適配器在基礎模型凍結的情況下進行訓練,並僅在輸入與刻板印象相關概念足夠相似時在推理時啟用;否則,直接使用原始模型。我們在BBQ、CrowS-Pairs和StereoSet上評估Acmite,並在ARC-Challenge、GSM8K和PIQA上評估一般能力的保留。在三個LLM上的實驗顯示,Acmite有效減輕了性別偏見,並在互補評估格式中保持了競爭性能,對於與偏見無關的任務也是如此。匿名代碼和數據可在 https://anonymous.4open.science/r/Acmite-18E2/ 獲得。

Compound interpretation is based on analogy

2610.01688v1 by Tian Shen, Harald Baayen

How compound meanings are best predicted from constituent meanings remains a central question in computational models of lexical semantics. Comparing different computational models provides a way to evaluate alternative accounts of how semantic information is combined during compound comprehension. We propose a new model, the Compound Analogy Model (CAM), that predicts a compound's embedding by adding its constituent embeddings together with the average shift vectors of the two constituents' compound families. The resulting model is parameter-free and exploits local analogical structure in the semantic space. We evaluated CAM against the CAOSS model on Mandarin Chinese compounds. CAM consistently achieved higher prediction accuracy than CAOSS on both training and held-out data, with the exception of three-character compounds, for which analogical generalization is constrained by both small constituent families and a pronounced imbalance in family size between the two constituents. The advantage of CAM remained when evaluation was based on frequency-defined train-test splits that better approximate generalization from familiar to novel compounds. To assess the cognitive plausibility of the two models, we further examined whether model-derived semantic measures predict visual lexical decision latencies for two-character compounds. Predictors derived from CAM provided improved prediction for response latencies compared to predictors derived from the CAOSS model. These findings indicate that compound meaning is better characterized as local analogical generalization than as the application of a learned global linear transformation, and demonstrate that analogical semantic structure provides a cognitively plausible basis for compound comprehension.

摘要:如何從成分意義中最佳預測複合意義仍然是計算語義學模型中的一個核心問題。比較不同的計算模型提供了一種評估替代解釋的方式,這些解釋涉及在複合理解過程中語義信息是如何結合的。我們提出了一個新模型,即複合類比模型(Compound Analogy Model, CAM),它通過將成分嵌入與兩個成分的複合家族的平均位移向量相加來預測複合詞的嵌入。結果模型是無參數的,並利用語義空間中的局部類比結構。我們在普通話的複合詞上將CAM與CAOSS模型進行了評估。CAM在訓練數據和保留數據上始終比CAOSS達到更高的預測準確性,除了三字複合詞,因為類比推廣受到成分家族小和兩個成分之間家族大小明顯不平衡的限制。當評估基於頻率定義的訓練-測試拆分時,CAM的優勢依然存在,這些拆分更好地近似從熟悉到新穎的複合詞的推廣。為了評估這兩個模型的認知合理性,我們進一步檢查了模型衍生的語義度量是否能預測兩字複合詞的視覺詞彙決策延遲。與CAOSS模型衍生的預測因子相比,CAM衍生的預測因子對反應延遲的預測有所改善。這些發現表明,複合意義更好地被描述為局部類比推廣,而不是學習的全局線性變換的應用,並且展示了類比語義結構為複合理解提供了認知上合理的基礎。

Architectural Sampling: Test-Time Scaling via Computational Diversity in Frozen Vision-Language Models

2610.01687v1 by Akshit Singh, Shyam Marjit, Wei Lin, Leonid Karlinsky, M. Jehanzeb Mirza

Test-time scaling often seeks better answers by sampling multiple responses from a frozen model, yet conventional temperature sampling generates every candidate along the same fixed computation path. We introduce architectural sampling, a training-free method that generates candidates through distinct forward computations by reusing selected blocks of decoder layers. Varying the block location and repetition count introduces computational diversity without updating model weights or adding auxiliary parameters. Across five Qwen checkpoints and twelve multimodal benchmarks, architectural sampling improves pass@9 over standard-path temperature sampling by 6.58 percentage points on average at the same nine-candidate budget. Reusing early layers yields the strongest gains, and the improvement in candidate coverage persists even under greedy decoding. The resulting candidates show lower lexical overlap and improve accuracy when used as rollouts for label-free test-time reinforcement learning. These findings extend the benefits of our architectural sampling beyond candidate coverage, demonstrating more effective learning from a model's own outputs.

摘要:測試時的擴展通常透過從凍結模型中抽樣多個回應來尋求更好的答案,然而傳統的溫度抽樣沿著相同的固定計算路徑生成每個候選項。我們引入了架構抽樣,這是一種無需訓練的方法,通過重用選定的解碼器層區塊來通過不同的前向計算生成候選項。變更區塊位置和重複次數引入了計算多樣性,而不必更新模型權重或添加輔助參數。在五個Qwen檢查點和十二個多模態基準測試中,架構抽樣在相同的九個候選預算下,平均提高了相對於標準路徑溫度抽樣6.58個百分點的pass@9。重用早期層獲得了最強的增益,即使在貪婪解碼下,候選覆蓋的改善仍然持續。所生成的候選項顯示出較低的詞彙重疊,並在用作無標籤測試時的強化學習回滾時提高了準確性。這些發現擴展了我們的架構抽樣的好處,不僅限於候選覆蓋,還展示了從模型自身輸出中更有效的學習。

Iterative Policy Refinement through Semantic Rollout Analysis

2610.01652v1 by Feiyu Gavin Zhu, Qi Xu, Zhifei Deng, Zhigang Hua, Luke Simon, Jean Oh, Reid Simmons

Structured policies improve efficiency, robustness, and interpretability in imitation learning by introducing task-specific inductive bias, but existing structure generation methods rely either on extensive human input or on static domain knowledge encoded in LLMs, which may be inconsistent with the expert demonstrations. We propose a closed-loop framework that iteratively refines structured policies using LLM-guided analysis of policy rollouts. By logging rollouts as semantically meaningful tabular data and prompting the LLM to generate diagnostic analysis code, our method identifies suboptimalities in the policy structure and iteratively corrects them without requiring human instruction. Experiments on car racing and door opening tasks show that our approach improves imitation learning performance by up to 15% over zero-shot LLM-generated structures and requires 75% less compute to achieve the same reinforcement learning performance. These results demonstrate that tabular rollout analysis provides an effective feedback signal to align LLM-generated policy structures with expert demonstrations, and we can utilize it to generate good policy structures automatically.

摘要:結構化政策透過引入特定任務的歸納偏見來提升模仿學習的效率、穩健性和可解釋性,但現有的結構生成方法要麼依賴大量的人類輸入,要麼依賴於編碼在大型語言模型(LLMs)中的靜態領域知識,這可能與專家的示範不一致。我們提出了一個閉環框架,通過使用LLM引導的政策展開分析來迭代地改進結構化政策。通過將展開記錄為語義上有意義的表格數據,並提示LLM生成診斷分析代碼,我們的方法識別出政策結構中的次優性,並在不需要人類指導的情況下進行迭代修正。在賽車和開門任務上的實驗顯示,我們的方法在模仿學習性能上比零樣本LLM生成的結構提高了多達15%,並且需要75%更少的計算來達到相同的強化學習性能。這些結果表明,表格展開分析提供了一個有效的反饋信號,以使LLM生成的政策結構與專家示範對齊,我們可以利用它自動生成良好的政策結構。

MCIR: A Feature Dependence-Aware Explainability Method with Reliability Guarantees

2610.01641v1 by Poushali Sengupta, Sabita Maharjan, Frank Eliassen, Shashi Raj Pandey, Yan Zhang

Modern machine-learning models often contain strongly dependent or redundant features, making feature attribution difficult because shared predictive information can be distributed across correlated predictors. Existing methods such as SHAP, LIME, HSIC, MI/CMI, and SAGE may therefore produce unstable rankings under multicollinearity or near-duplicate predictors. We propose the Mutual Correlation Impact Ratio Method (MCIR-M), a dependence-aware global feature-importance approach that quantifies the unique predictive information contributed by each feature beyond a selected dependence neighbourhood. MCIR-M introduces the Mutual Correlation Impact Ratio (MCIR), which conditions each feature on strongly dependent neighbours and computes a normalized ratio of conditional to block-level information. The population score lies in [0,1] and equals zero under exact conditional redundancy. We also introduce a lightweight estimation procedure that computes MCIR using a fraction of the available data and evaluates agreement with full-data explanations. Across controlled synthetic redundancy experiments and the UCI HAR benchmark, MCIR shows dependence-aware ranking behaviour, with its clearest advantage under injected near-duplicate predictors. Comparisons with independent and conditional SHAP, SAGE, HSIC, MI-based scores, and CIR-family baselines are mixed across real-data criteria. Reduced explanation samples lower computational burden in the evaluated configurations, while agreement with full-data explanations is assessed separately through ranking, head-set, and faithfulness diagnostics. Overall, MCIR-M provides a practical dependence-aware diagnostic for global explanation under strong feature dependence.

摘要:現代機器學習模型通常包含強相關或冗餘的特徵,使得特徵歸因變得困難,因為共享的預測信息可能分佈在相關的預測變數之間。現有的方法如SHAP、LIME、HSIC、MI/CMI和SAGE在多重共線性或近乎重複的預測變數下可能因此產生不穩定的排名。我們提出了互相關影響比率方法(MCIR-M),這是一種考慮依賴性的全局特徵重要性方法,量化每個特徵在選定的依賴鄰域之外所貢獻的獨特預測信息。MCIR-M引入了互相關影響比率(MCIR),該比率在強依賴的鄰居上對每個特徵進行條件化,並計算條件信息與區塊級信息的標準化比率。該人口得分位於[0,1]之間,並在精確的條件冗餘下等於零。我們還引入了一種輕量級的估計程序,該程序使用部分可用數據計算MCIR,並評估與全數據解釋的一致性。在受控的合成冗餘實驗和UCI HAR基準測試中,MCIR顯示出考慮依賴性的排名行為,其在注入的近重複預測變數下的優勢最為明顯。與獨立和條件SHAP、SAGE、HSIC、基於MI的得分以及CIR系列基準的比較在真實數據標準下是混合的。在評估的配置中,減少的解釋樣本降低了計算負擔,而與全數據解釋的一致性則通過排名、頭部集和忠實性診斷單獨評估。總體而言,MCIR-M為強特徵依賴下的全局解釋提供了一種實用的考慮依賴性的診斷方法。

Not All Error Yields to Scale: Where Scaling Stops in Vision-Language Inference

2610.01640v1 by Xinye Zhao, Yunkai Dang, Yunchen Wu, Wenbin Li

Vision-language models (VLMs) face a fixed-budget trade-off between processing more visual information for fine-grained perception and using a larger language backbone for complex reasoning. Existing studies do not tell us which combination of backbone size and input resolution to deploy, especially in high-resolution deployments. To address this gap, we propose the Separable Law that describes how VLM performance changes with language backbone size and visual token count. We fit the law to measurements from 26 InternVL and QwenVL models, with language backbone sizes from 1B to 72B, on four high-resolution benchmarks with image sizes from 224 pixels to 8K. We find that the questions responding to scaling can be predicted from the skill they require, while a substantial fraction never responds at all. We also find that the two model families gain similarly from a larger backbone, while their gains from more visual tokens differ sharply. Combined with a cost law, the Separable Law gives a closed-form rule for allocating compute between backbone size and visual tokens. When deployment is limited to available configurations, the law identifies model and image sizes that perform close to the best feasible choice under the same budget. We hope our work offers a principled way to decide how much a model should be allowed to see at high resolution, given what it must reason about.

摘要:視覺語言模型(VLMs)在處理更多視覺信息以實現細緻感知和使用更大語言骨幹以進行複雜推理之間面臨固定預算的權衡。現有研究並未告訴我們在高解析度部署中應該使用哪種骨幹大小和輸入解析度的組合。為了解決這一空白,我們提出了可分離法則,描述了VLM性能如何隨著語言骨幹大小和視覺標記數量的變化而變化。我們將該法則適配於來自26個InternVL和QwenVL模型的測量,這些模型的語言骨幹大小從1B到72B,並在四個高解析度基準上進行測試,圖像大小從224像素到8K。我們發現,對於擴展的問題,其反應可以根據所需的技能進行預測,而相當一部分則根本不會反應。我們還發現,這兩個模型系列在使用更大骨幹時獲益相似,而它們從更多視覺標記中獲得的收益則有明顯差異。結合成本法則,可分離法則提供了一個封閉形式的規則,用於在骨幹大小和視覺標記之間分配計算資源。當部署受限於可用配置時,該法則識別出在相同預算下表現接近最佳可行選擇的模型和圖像大小。我們希望我們的工作能提供一種原則性的方法,以決定模型在高解析度下應該被允許看到多少,考慮到它必須推理的內容。

Yo-ByT5: Efficient and High-Fidelity Diacritic Restoration for Yorùbá

2610.01634v1 by Ahmad Samuel Gali, Shamsuddeen Hassan Muhammad

Yorùbá is a widely spoken tonal language that depends on diacritics to avoid lexical ambiguity. However, it is often written without these diacritics, thereby hindering downstream Natural Language Processing (NLP) tasks. In this paper, we introduce Yo-ByT5, a byte-level Automatic Diacritic Restoration (ADR) model fine-tuned from ByT5-small. We evaluate Yo-ByT5 alongside five publicly released Yorùbá ADR models and one open-weight large language model (LLM) on the YAD benchmark under a consistent protocol. Our results demonstrate that Yo-ByT5 matches the performance of the strongest existing model, mT5-base, with a DER of 10.14% and a CER of 3.48%. Furthermore, it exhibits superior text fidelity despite using approximately half the parameter count of mT5-base. We also release our training code and model outputs, as well as call for the development of a larger, purpose-built benchmark for Yorùbá diacritic restoration.

摘要:Yorùbá 是一種廣泛使用的音調語言,依賴於變音符號以避免詞彙歧義。然而,它經常在沒有這些變音符號的情況下書寫,從而妨礙了下游的自然語言處理 (NLP) 任務。在本文中,我們介紹了 Yo-ByT5,一個從 ByT5-small 微調而來的字節級自動變音符號恢復 (ADR) 模型。我們在 YAD 基準上,根據一致的協議,將 Yo-ByT5 與五個公開發布的 Yorùbá ADR 模型和一個開放權重的大型語言模型 (LLM) 進行評估。我們的結果顯示,Yo-ByT5 的性能與現有最強模型 mT5-base 相當,具有 10.14% 的 DER 和 3.48% 的 CER。此外,儘管使用的參數數量約為 mT5-base 的一半,但它在文本保真度上表現出色。我們還發布了我們的訓練代碼和模型輸出,並呼籲開發一個更大、專門針對 Yorùbá 變音符號恢復的基準。

What Makes Something Hard(er)? Explaining Question Difficulty in Natural Language

2610.01627v1 by Peng Cui, Qiaoyuan Zheng, Rudolf Debelak, Mrinmaya Sachan

Difficulty is one of the most fundamental properties of a question: it determines whether the question can meaningfully discriminate between models of differing ability. Although a variety of methods can now estimate or predict difficulty automatically, they yield only a single descriptive number, with no account of the underlying factors that make a question difficult in the first place. In this work, we propose a data-driven approach that automatically generates and validates natural-language hypotheses explaining what makes one question harder than another. We first estimate each item's difficulty from the responses of a large pool of LLMs using Item Response Theory. We then sample contrasting sets of easy and hard questions and prompt an LLM to propose candidate explanations of the difference, which are subsequently validated and selected on held-out questions. Experimental results across three datasets spanning mathematical, logical, and commonsense reasoning show that our method produces interpretable and predictive hypotheses. On their own, they predict the difficulty of unseen questions competitively with, or better than, advanced black-box difficulty regressors; used as additional features, they further improve those regressors, implying that they discover difficulty signals that existing models fail to capture. Moreover, we demonstrate that editing questions according to a hypothesis can shift their measured difficulty in the expected direction, indicating that the discovered hypotheses are causally valid difficulty factors rather than post-hoc descriptions. Our approach thus turns a purely descriptive difficulty score into actionable statements.

摘要:困難度是問題最基本的特性之一:它決定了問題是否能夠有意義地區分不同能力的模型。雖然現在有多種方法可以自動估計或預測困難度,但它們僅產生一個描述性的數字,並未考慮使問題變得困難的潛在因素。在這項工作中,我們提出了一種數據驅動的方法,自動生成和驗證自然語言假設,解釋為什麼一個問題比另一個問題更難。我們首先使用項目反應理論從大量大型語言模型的回應中估計每個項目的困難度。然後,我們抽取一組對比的簡單和困難問題,並提示一個大型語言模型提出候選解釋這些差異,這些解釋隨後在保留的問題上進行驗證和選擇。跨越數學、邏輯和常識推理的三個數據集的實驗結果顯示,我們的方法產生了可解釋且具有預測性的假設。僅憑這些假設,它們能夠與先進的黑箱困難回歸模型競爭地預測未見問題的困難度;作為額外特徵使用時,它們進一步改善了這些回歸模型,這意味著它們發現了現有模型未能捕捉的困難信號。此外,我們證明根據假設編輯問題可以將其測量的困難度朝預期方向轉變,這表明所發現的假設是因果有效的困難因素,而非事後描述。因此,我們的方法將純粹描述性的困難分數轉化為可行的陳述。

FedLore: Communication and Memory Efficient Federated Learning via Shared Gradient Low-Rank Projection

2610.01620v1 by Junkang Liu

Federated training of foundation models is constrained by client memory and communication costs. LoRA-based methods reduce these costs through low-rank adapters, but their fixed rank budget can limit adaptation. Gradient low-rank optimization offers greater flexibility, yet independently chosen client subspaces create a problem we term \emph{subspace fragmentation}: local projections interact with data heterogeneity to bias aggregated directions, while aggregation can increase update rank and communication cost. Thus, accurate local gradient compression need not preserve global descent. We propose \texttt{FedLore}, which shares a low-rank optimization basis within each round and refreshes it across rounds. The shared basis enables exact aggregation in low-rank coordinates and eliminates the identified projection bias. Subspace refresh allows the accumulated model update to exceed the per-round rank budget. We characterize the aggregation bias and establish an $O(T^{-1/2})$ stationarity bound for the projected-SGD variant under a global-gradient coverage condition and standard smoothness and variance assumptions, with bounded gradient heterogeneity. Experiments on vision and language tasks, including federated pre-training, show that \texttt{FedLore} outperforms the evaluated low-rank adapter baselines and matches or exceeds full-parameter training, while reducing communication and optimizer-state memory.

摘要:聯邦訓練基礎模型受到客戶端記憶體和通信成本的限制。基於LoRA的方法通過低秩適配器降低這些成本,但其固定的秩預算可能限制適應性。梯度低秩優化提供了更大的靈活性,但獨立選擇的客戶端子空間會產生我們稱之為\emph{subspace fragmentation}的問題:局部投影與數據異質性相互作用,偏向於聚合方向,而聚合可能會增加更新秩和通信成本。因此,準確的局部梯度壓縮不必保留全局下降。我們提出\texttt{FedLore},在每一輪中共享低秩優化基礎,並在輪與輪之間進行刷新。共享基礎使得在低秩坐標中進行精確聚合,並消除了已識別的投影偏差。子空間刷新允許累積的模型更新超過每輪的秩預算。我們描述了聚合偏差,並在全局梯度覆蓋條件及標準平滑性和方差假設下,對投影-SGD變體建立了$O(T^{-1/2})$的平穩性界限,並且具有有界的梯度異質性。在視覺和語言任務上的實驗,包括聯邦預訓練,顯示\texttt{FedLore}的表現超過了評估的低秩適配器基準,並且與全參數訓練相匹配或超過,同時減少了通信和優化器狀態記憶體。

Exposing the Cost of Deep Learning Audio Development

2610.01619v1 by Constance Douwes, Paul Magron, Romain Serizel

The environmental impact of deep learning has attracted increasing attention over the past decade. Existing studies mainly focus on the energy and carbon emissions of model training and inference, while the whole development phase is often overlooked. Yet, architecture prototyping and intensive experiments are conducted during this stage, which is highly energy-demanding. In this article, we propose a methodology to estimate these costs, based on activity logs from the Grid5000 shared computing platform used by the LORIA laboratory. As a case-study, we focus on audio projects developed in the Multispeech research team. We evaluate the overall energy cost of four projects, and we compare them to those of training the reported models. Our results show that the energy required for the development phase is 3 to 256 times greater than that required to train the best-performing model alone. These results advocate for a more systematic reporting of energy consumption across the entire life cycle of deep learning-based audio projects.

摘要:深度學習的環境影響在過去十年中引起了越來越多的關注。現有的研究主要集中在模型訓練和推理的能量和碳排放上,而整個開發階段常常被忽視。然而,在這個階段進行架構原型設計和密集實驗,這是非常耗能的。在本文中,我們提出了一種基於LORIA實驗室使用的Grid5000共享計算平台的活動日誌來估算這些成本的方法。作為案例研究,我們專注於Multispeech研究團隊開發的音頻項目。我們評估了四個項目的整體能量成本,並將其與訓練報告模型的能量成本進行比較。我們的結果顯示,開發階段所需的能量是僅訓練最佳性能模型所需能量的3到256倍。這些結果提倡對基於深度學習的音頻項目整個生命週期的能量消耗進行更系統的報告。

Agents Are Systems, Not Models: Rethinking Agentic Evaluation

2610.01618v1 by Luis Wiedmann, Leander Girrbach, Cordelia Schmid, Zeynep Akata

Agent evaluations increasingly go beyond a single success rate, reporting metrics such as cost, consistency, and robustness. Yet they typically treat the agent itself as fixed. In practice, an agent is a configurable system: users decide what to tell it, how long to let it run, and which model to use, and each of these choices can change how well and how consistently it performs. We study these choices on a new benchmark of four scientific tasks, where a coding agent must find and correctly operate a published specialist model. We investigate five parts of the agent's configuration: task information, reasoning, self-verification, time budget, and backbone model. We find substantial run-to-run variability, with approximately 54% of the outcome variance coming from repeating the same configuration rather than changing it. Across configurations, the information provided to the agent has the largest effect, exceeding both time budget and model size, while also reducing cost and improving calibration. Configuration choices also interact: additional time helps only when the agent has sufficient information or a capable enough model to use it. Finally, a trajectory-based taxonomy of agent behavior reveals that prompting an agent to verify its answer has little effect on its verification behavior, whereas providing a dedicated verification tool changes that behavior substantially. These results suggest that agents should be evaluated as configurable systems themselves, and that some desired behaviors are more effectively implemented in the system than requested through prompting. We release the benchmark and more than 18,000 agent trajectories.

摘要:代理評估越來越超越單一的成功率,報告如成本、一致性和穩健性等指標。 然而,它們通常將代理本身視為固定的。 實際上,代理是一個可配置的系統:用戶決定告訴它什麼、讓它運行多久以及使用哪個模型,而這些選擇都會改變它的表現效果和一致性。 我們在一個新的基準上研究這些選擇,該基準包含四個科學任務,其中一個編碼代理必須找到並正確操作一個已發表的專家模型。 我們調查了代理配置的五個部分:任務信息、推理、自我驗證、時間預算和骨幹模型。 我們發現運行之間存在顯著的變異性,大約54%的結果變異來自重複相同的配置,而不是改變它。 在不同配置中,提供給代理的信息具有最大的影響,超過了時間預算和模型大小,同時還降低了成本並改善了校準。 配置選擇之間也存在相互作用:額外的時間僅在代理擁有足夠的信息或足夠能力的模型來使用時才有幫助。 最後,基於軌跡的代理行為分類法顯示,促使代理驗證其答案對其驗證行為幾乎沒有影響,而提供專用的驗證工具則會顯著改變該行為。 這些結果表明,代理應該被評估為可配置的系統本身,而某些期望的行為在系統中實現的效果比通過提示請求更有效。 我們發布了基準和超過18,000條代理軌跡。

Can LLMs Reliably Annotate Bioassay Metadata to Improve Data Readiness?

2610.01616v1 by Laura van Weesep, Riccardo Tedoldi, Jens Sjölund, Hossein Azizpour, Susanne Winiwarter, Ola Engkvist, Jon Paul Janet, Samuel Genheden, Juan Viguera Diez

The emergence of foundation models for molecular property prediction requires a high degree of AI data readiness, including reliable metadata annotation. However, both public repositories and industrial screening databases suffer from missing, inconsistent, or conflated assay annotations. In this work, we quantify the extent of missing annotations in PubChem for the BioAssay Ontology (BAO) assay format and physical detection method fields and investigate whether open-source and proprietary large language models (LLMs) can reliably predict and audit metadata annotations directly from the assay text. In our assessment, we found that the annotation coverage across PubChem's $\sim$2 million bioassays is critically sparse, 36\% lacking an assay format, 89\% a BioAssay type, and >99.9\% any BAO-mapped assay format or detection technology term. This motivates the need for automated test-metadata curation. Using evaluation sets derived from PubChem and ChEMBL, we assess the agreement of seven open-source and proprietary LLMs with existing silver labels. Recall is at least 0.96 for biochemical and cell-based assay formats, with a similar pattern for detection technology, although disagreements increase on under-represented classes. Manual inspection shows that many of these disagreements trace back to inconsistencies between silver sources rather than to LLM error. Moreover, in a qualitative study with a senior industrial curator, LLM-generated evidence prompted the expert to revise some of their own labels, showing LLMs can flag potentially mislabeled assays. Across the study, performance differences between proprietary and open-source models were small. Together, these results suggest LLMs can support the large-scale annotation and auditing of assay metadata, though per-class reliability estimates and targeted human review remain necessary before such labels enter downstream ML pipelines.

摘要:基於分子性質預測的基礎模型的出現需要高度的人工智慧數據準備,包括可靠的元數據註釋。然而,公共資料庫和工業篩選數據庫都存在缺失、不一致或混淆的檢測註釋。在這項工作中,我們量化了PubChem中BioAssay本體(BAO)檢測格式和物理檢測方法字段缺失註釋的程度,並調查開源和專有大型語言模型(LLMs)是否能夠可靠地從檢測文本中直接預測和審核元數據註釋。在我們的評估中,我們發現PubChem約200萬個生物檢測的註釋覆蓋率極其稀疏,36\%缺乏檢測格式,89\%缺乏BioAssay類型,且超過99.9\%缺乏任何BAO映射的檢測格式或檢測技術術語。這促使了自動測試元數據整理的需求。利用從PubChem和ChEMBL衍生的評估集,我們評估了七個開源和專有LLMs與現有銀標籤的一致性。生化和基於細胞的檢測格式的召回率至少為0.96,檢測技術的模式相似,儘管在代表性不足的類別上分歧增加。手動檢查顯示,許多這些分歧源於銀來源之間的不一致,而不是LLM的錯誤。此外,在與一位資深工業策展人的定性研究中,LLM生成的證據促使專家修訂他們的一些標籤,顯示LLMs可以標記潛在錯誤標記的檢測。在整個研究中,專有模型和開源模型之間的性能差異很小。綜合這些結果表明,LLMs可以支持檢測元數據的大規模註釋和審核,儘管在這些標籤進入下游機器學習管道之前,仍然需要每類的可靠性評估和針對性的人工審查。

Hob-VL: A Benchmark for Visually Grounded Boolean Reasoning

2610.01605v1 by Yuzhou Wang, Emile Anand, Ijay Narang

Reliable visual reasoning requires composing multiple visual observations and returning consistent answers to logically equivalent questions. We introduce Hob-VL, a benchmark for visually grounded Boolean reasoning. Hob-VL comprises two tasks: (1) evaluating whether a Boolean rule holds in an image, and (2) identifying the (unique) object satisfying a Boolean description. Hob-VL contains 6,000 human-verified balanced Yes/No questions, each defined by a Boolean combination of ten visual statements, across 1,000 generated scenes and 46 diverse labeled photographs, along with 1,000 object-identification questions over the same photographs. Our question families are deliberately constructed to challenge reasoning through misleading local cues and nested logical operations, and include symbolic and structured natural-language presentations. Across eight model configurations with thinking disabled or minimized, Boolean accuracy ranges from 48.52% to 50.57%, while the identification accuracy reaches at most 43.0%. A thinking-enabled GLM configuration achieves uneven gains while retaining substantial errors and inconsistencies. Hob-VL exposes these failures through executable reference answers and matched evaluations.

摘要:可靠的視覺推理需要組合多個視覺觀察,並對邏輯上等價的問題給出一致的答案。 我們介紹了 Hob-VL,一個針對視覺基礎布林推理的基準。 Hob-VL 包含兩個任務:(1)評估布林規則在圖像中是否成立,以及(2)識別滿足布林描述的(唯一)物體。 Hob-VL 包含 6,000 個經過人工驗證的平衡是/否問題,每個問題由十個視覺陳述的布林組合定義,涵蓋 1,000 個生成場景和 46 張多樣化的標記照片,以及 1,000 個針對相同照片的物體識別問題。我們的問題家族故意構建,以通過誤導性的局部提示和嵌套邏輯操作來挑戰推理,並包括符號和結構化的自然語言呈現。在八種思考被禁用或最小化的模型配置中,布林準確率範圍從 48.52% 到 50.57%,而識別準確率最多達到 43.0%。 一個啟用思考的 GLM 配置在保留相當大的錯誤和不一致性的同時,實現了不均勻的增益。 Hob-VL 通過可執行的參考答案和匹配的評估揭示了這些失敗。

Permutation-Robust Decision Modeling with Candidate-Independent Block-Causal Attention

2610.01601v1 by Guy Amit

Decision models often score a variable-sized set of candidate actions encoded in a single sequence. This setting is increasingly relevant for System 1 components inside generative systems, where candidates may be proposed or ordered differently across runs. Standard causal cross-encoding is expressive, but it can make a candidate's score depend on serialization order rather than on the underlying decision problem. We introduce candidate-independent block-causal attention, which preserves causal computation within the shared context and each candidate while blocking cross-candidate information flow and resetting candidate positions. We compare this architecture with standard causal attention and complementary invariant baselines across Gemma 3 1B, Qwen3 1.7B, and Qwen3 4B backbones. Candidate-independent attention consistently reduces permutation sensitivity while retaining competitive decision quality; ablations indicate that candidate isolation is the primary source of the effect, with position resetting completing the intended symmetry. A larger Qwen3-4B study further examines the behavior of the proposed architecture with substantially more training data. Code is available at the \href{https://github.com/guyAmit/ci-decision-models}{\textcolor{blue}{project repository}}, and the \href{https://huggingface.co/Guy-Amit/qwen3-4b-ci-decision-4096-poc}{\textcolor{blue}{Qwen3-4B model artifact}} is available on Hugging Face.

摘要:決策模型通常對編碼在單一序列中的可變大小候選行動集進行評分。這種設定對於生成系統中的系統1組件越來越相關,其中候選者在不同運行中可能以不同方式被提議或排序。標準的因果交叉編碼具有表達能力,但它可能使候選者的分數依賴於序列化順序,而不是基於底層的決策問題。我們引入了候選者獨立的區塊因果注意力,該方法在共享上下文和每個候選者內部保留因果計算,同時阻止候選者之間的信息流動並重置候選者位置。我們將這種架構與標準因果注意力和互補不變基準進行比較,涵蓋了Gemma 3 1B、Qwen3 1.7B和Qwen3 4B主幹。候選者獨立的注意力在保持競爭性決策質量的同時,始終減少排列敏感性;消融實驗表明,候選者隔離是該效應的主要來源,位置重置則完成了預期的對稱性。一項更大規模的Qwen3-4B研究進一步檢查了所提議架構在大量訓練數據下的行為。代碼可在\href{https://github.com/guyAmit/ci-decision-models}{\textcolor{blue}{項目庫}}獲得,而\href{https://huggingface.co/Guy-Amit/qwen3-4b-ci-decision-4096-poc}{\textcolor{blue}{Qwen3-4B模型文物}}則可在Hugging Face上獲得。

Before It Fades: Reinforcing Temporal Representations at Inference Time in VideoLLMs

2610.01595v1 by Youngwoo Shin, Yusung Ro, Minseo Kim, Junmo Kim

Video Large Language Models (VideoLLMs) receive frames in sequential order and interpret how visual content evolves along the temporal axis, yet temporal reasoning remains a persistent weakness across architectures. Reversing the frame order of a video, a transformation that should invert temporal answers, often leaves the final prediction unchanged. We investigate where this failure originates by defining the temporal divergence vector $τ_l$, the layer-wise representational difference induced by reversing temporal order. Tracking its magnitude across layers reveals a consistent temporal divergence profile where the divergence peaks at intermediate layers and progressively diminishes toward the output. We confirm this peak is specific to temporal reasoning and functionally critical for predictions, establishing that VideoLLMs acquire temporal information at intermediate layers but fail to maintain it to the output. This progressive fading motivates our method, Temporal Activation Injection (TAI), which extracts $τ_l$ at the peak of the profile for each input and reinjects it into subsequent layers following the measured decay. TAI requires no training and consistently improves temporal reasoning across three VideoLLMs and four benchmarks with negligible impact on non-temporal tasks. Code is available at https://github.com/Youngwoo-git/Before-It-Fades.

摘要:視頻大型語言模型(VideoLLMs)以順序方式接收幀並解釋視覺內容如何沿時間軸演變,然而,時間推理在各種架構中仍然是一個持久的弱點。反轉視頻的幀順序,這一轉換應該會顛倒時間答案,但最終預測往往保持不變。我們通過定義時間發散向量 $τ_l$ 來調查這一失敗的來源,這是由反轉時間順序引起的層級表示差異。跟踪其在各層的大小揭示了一個一致的時間發散輪廓,其中發散在中間層達到峰值,並逐漸減少到輸出。我們確認這一峰值是特定於時間推理的,並且對預測至關重要,確立了 VideoLLMs 在中間層獲取時間信息但未能將其保持到輸出的事實。這一漸進的衰減激發了我們的方法——時間激活注入(Temporal Activation Injection, TAI),該方法在每個輸入的輪廓峰值處提取 $τ_l$,並在隨後的層中重新注入,根據測量的衰減進行。TAI 不需要訓練,並在三個 VideoLLMs 和四個基準測試中一致改善時間推理,對非時間任務的影響微乎其微。代碼可在 https://github.com/Youngwoo-git/Before-It-Fades 獲得。

Which LLM to pick? Online Active Model Selection for Large Language Models

2610.01592v1 by Alessandro Turrin, Patrik Okanovic, Torsten Hoefler, Nezihe Merve Gürel

Large Language Models (LLMs) are increasingly applied to process streaming data, with practitioners relying on benchmarks to select the best model even though these signals only approximate real performance. While oracle annotations can provide reliable feedback, they are often costly and difficult to obtain at scale. To address this challenge, we propose ONLINE LLM PICKER, the first framework for active model selection for LLMs in online settings. Given an arbitrary stream of queries and a limited annotation budget, ONLINE LLM PICKER selects the most informative prompts for annotation to identify the best LLM among candidate models. Across multiple tasks including 10 datasets, for over 130 language models, we show that ONLINE LLM PICKER saves annotation cost by up to 71.67% while reliably identifying the best or near-best model for the stream. We also show that using the returned model for sequential generation on unannotated prompts across the stream reduces regret by up to a factor of 2.51x, indicating that ONLINE LLM PICKER can identify the best or near-best model well before processing all streaming prompts.

摘要:大型語言模型(LLMs)越來越多地應用於處理串流數據,實踐者依賴基準來選擇最佳模型,即使這些信號僅能近似實際性能。雖然預言者註釋可以提供可靠的反饋,但它們通常成本高昂且難以大規模獲得。為了解決這一挑戰,我們提出了ONLINE LLM PICKER,這是第一個針對在線環境中LLMs的主動模型選擇框架。給定一個任意的查詢串流和有限的註釋預算,ONLINE LLM PICKER選擇最具信息量的提示進行註釋,以識別候選模型中最佳的LLM。在包括10個數據集的多個任務中,針對超過130個語言模型,我們顯示ONLINE LLM PICKER將註釋成本降低了高達71.67%,同時可靠地識別出串流中的最佳或接近最佳模型。我們還顯示,使用返回的模型在串流中的未註釋提示上進行順序生成,將後悔值降低了高達2.51倍,這表明ONLINE LLM PICKER能夠在處理所有串流提示之前就識別出最佳或接近最佳的模型。

Evaluating Physical Consistency and Plausibility in Generative Scenario Models for Autonomous Driving

2610.01581v1 by Manasa Mariam Mammen, Zafer Kayatas, Stefan Wagner

Generative AI models are increasingly used for scenario generation in autonomous driving. While they can generate realistic-looking scenarios, they often provide limited transparency into learned representations and consistency with real-world vehicle dynamics. This lack of formal assurance limits their use in safety-critical validation and certification workflows. To address this aspect, we introduce a layered evaluation protocol that complements existing methods by assessing models across five layers. The first four layers inspect internal representations and network layers through kinematic alignment, statistical baseline comparison, latent controllability, and activation analysis. The fifth layer evaluates model outputs against vehicle dynamics constraints such as lateral jerk thresholds. We demonstrate the protocol on a Variational Autoencoder (VAE)-based scenario generator. Although standard output-level metrics and visualizations suggest that the generated scenarios are realistic, our protocol provides deeper insight into the extent to which the model's latent space aligns with kinematic features and whether visually plausible trajectories satisfy vehicle-dynamics constraints. We further apply the protocol to additional generative models, demonstrating its applicability beyond the VAE architecture.

摘要:生成式人工智慧模型在自動駕駛的場景生成中被越來越多地使用。雖然它們可以生成看起來現實的場景,但通常對於學習到的表徵和與現實世界車輛動態的一致性提供的透明度有限。這種缺乏正式保證的情況限制了它們在安全關鍵的驗證和認證工作流程中的使用。為了解決這個問題,我們提出了一種分層評估協議,通過在五個層面上評估模型來補充現有方法。前四個層面通過運動學對齊、統計基準比較、潛在可控性和激活分析檢查內部表徵和網絡層。第五個層面則評估模型輸出是否符合車輛動態約束,例如橫向加速度閾值。我們在一個基於變分自編碼器(VAE)的場景生成器上演示了這一協議。雖然標準的輸出級別指標和可視化顯示生成的場景是現實的,但我們的協議提供了更深入的見解,以了解模型的潛在空間在多大程度上與運動學特徵對齊,以及視覺上合理的軌跡是否滿足車輛動態約束。我們進一步將該協議應用於其他生成模型,展示其在超越VAE架構的適用性。

Medical explainable AI

Publish Date Title Authors Homepage Code
2026-10-01 MIRTO: a registration-gated, multiverse-tested evaluation protocol for unsupervised anomaly segmentation in brain MRI Negin Kafee Hernashki et.al. 2610.02136v1 null
2026-10-01 A Comparative Explainability Framework for DeBERTa-v3 in Zero-Shot Medical Abstract Classification Javier Diaz Esteban-Herreros et.al. 2610.02116v1 null
2026-10-01 Can AI Oversight Be Zero Knowledge? Alessandro Chiesa et.al. 2610.01995v1 null
2026-10-01 Mapping the RAG Landscape: A Four Axis Taxonomy of Efficiency, Defense, Interactivity, and Reasoning Meghana Sunil et.al. 2610.01936v1 null
2026-10-01 Removing spurious minima for planar features by skip connections Jakob Paul Zimmermann et.al. 2610.01728v1 null
2026-10-01 MCIR: A Feature Dependence-Aware Explainability Method with Reliability Guarantees Poushali Sengupta et.al. 2610.01641v1 null
2026-10-01 What Makes Something Hard(er)? Explaining Question Difficulty in Natural Language Peng Cui et.al. 2610.01627v1 null
2026-10-01 Measuring the Stability Assumption Behind Action Chunking Aryan Goyal et.al. 2610.01626v1 null
2026-10-01 Exact Distinguishability in Non-Markovian Decision Processes Kabir Murjani et.al. 2610.01527v1 null
2026-10-01 OpenMTB-Audit: Exposing Over-Refusal and Clinical Expert Perspectives in LLM-Based Molecular Tumor Board Safety Evaluation Negin Ashrafi et.al. 2610.01497v1 null
2026-10-01 Multi-Party Backchannel Prediction: a Diagnosis, a Benchmark, and a Ceiling Mohammed Hafsati et.al. 2610.01488v1 null
2026-10-01 Detect, Explain, Interpret: An End-to-End Benchmark for Time Series Anomaly Detection, Explainability and Interpretability Roberto Stanzione et.al. 2610.01168v1 null
2026-10-01 CineMR: Tool-Integrated Vision-Language Reasoning for Quantitative Cardiac MRI Assessment Kunyang Li et.al. 2610.01166v1 null
2026-10-01 What Can Analogy Tell Us About Artificial Consciousness? Keith J. Holyoak et.al. 2610.01002v1 null
2026-09-30 When Reasoning Helps Action: Monitoring and Steering Chain-of-Thought in Vision-Language-Action Policies Sathwik Karnik et.al. 2610.00601v1 null
2026-09-30 Worse Together: How Performance Breaks Down in Multi-User Multi-Agent Teams Sahan Paliskara et.al. 2610.00583v1 null
2026-09-30 No One Architecture Fits All: A Cross-Environment Evaluation of Hierarchical Red Team Agents Ayan Javeed Shaikh et.al. 2610.00557v1 null
2026-09-30 EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights Jiayi Geng et.al. 2610.00492v1 null
2026-09-30 CAS II: Symmetric Partitions as Kolmogorov Models Romie Banerjee et.al. 2609.40290v1 null
2026-09-30 Unlearnable, or Unmeasured? On the Reliability of Difficulty Labels in RLVR Chandak Chakma et.al. 2609.40115v1 null
2026-09-30 What Can Component-Replacement Evidence Establish? A Critical Scoping Review of Local Decisions in LLM Agents Shuyang Zhang et.al. 2609.39989v1 null
2026-09-30 How Does Local Landscape Geometry Evolve in Language Model Pre-Training? Zhanpeng Zhou et.al. 2609.39767v1 null
2026-09-30 Trust Is Not a Score: Runtime Assurance Contracts for High-Risk AI Agents Serhii Zabolotnii et.al. 2609.39717v1 null
2026-09-30 ChronoGraph: Functional 4D Scene Graphs with Vision-Language Models for Interaction Understanding and Grounded Planning Chenyangguang Zhang et.al. 2609.39665v1 null
2026-09-30 Zero-Compute Cross-Lingual Transferability Estimation Using Typological Feature Proxies Dalton Raphael Harmsen et.al. 2609.39640v1 null
2026-09-30 Disentangling Self-Distillation: Measuring and Modeling Acquisition and Retention Luis Zuin et.al. 2609.39494v1 null
2026-09-30 From Speech to Editable Concepts: Probing Emotion Recognition with Concept Bottleneck Models Hezhao Zhang et.al. 2609.39453v1 null
2026-09-30 Towards Trustworthy AI for Glioma Diagnosis: A Task-Aware Evaluation of Uncertainty Quantification Gonzalo Esteban Mosquera Rojas et.al. 2609.39429v1 null
2026-09-30 The Golden Path Hypothesis: Reusable Schedules in Diffusion Caching Dong Wang et.al. 2609.39343v1 null
2026-09-30 A Time-Aware Bag-of-Receptive-Fields for Interpretable Irregular Time Series Classification Francesco Spinnato et.al. 2609.39268v1 null
2026-09-30 When Harnesses Lose the Signal: Causal Evaluation of Recovery in LLM Agents Shuyao Xiao et.al. 2610.00372v1 null
2026-09-30 Loop-Free Inverse Reinforcement Learning via Sequential Value Recovery with Q-Score Matching Yang chen et.al. 2609.38955v1 null
2026-09-30 Reasoning Externalization for Faithful Large Language Model Narratives of Stock Return Predictions Sujung Kim et.al. 2609.38869v1 null
2026-09-30 SkillSeek: Revisiting Agent Skill Retrieval at Marketplace Scale Guanqun Yang et.al. 2609.38822v1 null
2026-09-30 When Reasoning Goes Astray: Attention Dynamics of Uncontrolled Reasoning Yuanhe Zhang et.al. 2609.38817v1 null
2026-09-29 Understanding Off- vs On-Policy Distillation: A Tale of Distinct Training Objectives Qiwei Di et.al. 2609.38666v1 null
2026-09-29 Defining and Categorising Human-AI Interactions in Clinical Trials: A Multidimensional Human-AI Classification Approach Sandra Woolley et.al. 2609.38559v1 null
2026-09-29 Demographic Pluralism: Inference-Time Modeling of Pluralistic Human Preference Distributions Meng-Chen Wu et.al. 2609.38555v1 null
2026-09-29 Does This Action Still Explain the Task? Reverse Scoring for Diffusion Language Model Agents Jiacheng Qiu et.al. 2609.38536v1 null
2026-09-29 Caption-Mediated Perceived-Safety Estimation for Pedestrian Routing Simon Parkinson et.al. 2609.38479v1 null
2026-09-29 Aegis: Generative Gradient Masking for Privacy-Preserving Medical Federated Learning Chaoyu Zhang et.al. 2609.38339v1 null
2026-09-29 Does Local Video Understanding Transfer Across Encounters? The EgoGears Benchmark Yuedong Tan et.al. 2609.37938v1 null
2026-09-29 OmniVCBench: Benchmarking Evidence-Grounded Multimodal Reasoning Towards AI Virtual Cells Manyu Li et.al. 2609.37773v1 null
2026-09-29 Multi-Site Real-World Performance of Commercial AI for Pulmonary and Incidental Pulmonary Embolism Detection Aawez Mansuri et.al. 2609.37750v1 null
2026-09-29 XU-RS: Explaining Credal Width in Random-Set Language Models David Achara et.al. 2609.37594v1 null
2026-09-29 Raw Imagery Impacting Your AI: Should You Care? Adrien Dorise et.al. 2609.38265v1 null
2026-09-29 Selecting The Most Informative Tokens in Natural Language Autoencoders Federico Torrielli et.al. 2609.37040v1 null
2026-09-29 Harness Evolution as Learning: Approximation, Generalization, and Optimization Limits of Self-Improving Personal Agents Zeyu Gan et.al. 2609.36892v1 null
2026-09-29 Predictive Credit: Measuring What Scientific Explanations Add to Experimental Forecasts Jingjie Ning et.al. 2610.00314v1 null
2026-09-28 Engineering Simplicity: Simple Mechanism Interfaces Steer LLM Agents Kehang Zhu et.al. 2609.36365v1 null
2026-09-28 Explainability from Training with Applications to TCR-Epitope Prediction Jiarui Li et.al. 2609.36354v1 null
2026-09-28 ThuRunel: Dynamic Decoupling for Structured Advisory Dialogue Yuyan Chen et.al. 2609.36340v1 null
2026-09-28 FigAct: Turning Scientific Figures into Active Canvases for Explanation Shishi Xiao et.al. 2609.36190v1 null
2026-09-28 An Exact Generate - Transform Decomposition of Small-LLM Team Scaling Across Orchestration Architectures Blaz Bertalanic et.al. 2609.36104v1 null
2026-09-28 One Geometry, Different Outcomes: Readout-Dependent Effects of the Modality Gap in Vision-Language Models Aditya Sharma et.al. 2609.36101v1 null
2026-09-28 Shockingly Simple Self-retrospection Improves Agentic Models Without RL Jonathan Light et.al. 2609.35741v1 null
2026-09-28 Rethinking Circuit Evaluation: Do Circuits Explain Model Errors? Li Zhang et.al. 2609.35686v1 null
2026-09-28 Signatures of semantic search in the activations of large language models Luke Leckie et.al. 2609.35599v2 null
2026-09-28 A decision-support system applied to Law: Reasoning and explainability of the decision Jeremy Bouche-Pillon et.al. 2609.35370v1 null
2026-09-28 Jev thinks "I don't know'', but doesn't say it: Introducing Sys1Cal-v1 Dataset for Probability Calibration Riccardo Porcedda et.al. 2609.35342v1 null
2026-09-28 The Argument and the Letterhead: Source-Position Coherence in AI Evaluation Michele Loi et.al. 2609.35286v1 null
2026-09-28 Towards Reliable AI Data Scientists: Data Agents with Workflow Harnesses Huachi Zhou et.al. 2609.35255v1 null
2026-09-28 Beneath the Tokens: A Performance Engineering Study of Multi-Token Prediction in GPU-Accelerated LLM Inference Suwesh Prasad Sah et.al. 2609.35188v1 null
2026-09-28 Applying Language Models in Clinical Medicine: Recent Trends and Perspectives Erik Aerts et.al. 2609.34780v2 null
2026-09-28 A Unifying Framework of Concept-based Explainable AI with Completeness Guarantees Vojtěch Kůr et.al. 2609.34750v1 null
2026-09-28 From Human Narrative to Harmonic Structure: A Human-Centered Investigation of Algorithmic Music Generation through the Chord Wheel Diagram Josef Pavlíček et.al. 2609.34735v1 null
2026-09-28 Understanding Generalization Requires Universal Induction Aram Ebtekar et.al. 2609.34458v1 null
2026-09-28 Social Circuits behind Multi-agent Echo Chambers Chuiyang Meng et.al. 2609.34444v1 null
2026-09-28 Improving Large Language Models for Code through Runtime Program-State Reasoning Hongwei Li et.al. 2609.34359v1 null
2026-09-28 Dynamical Parameters: An Interpretability Framework for Time-Series Foundation Models Kang Yang et.al. 2609.34316v1 null
2026-09-28 Evo2Team: When Do Evolved Skills Transfer? From Selection to Deployment Renxiang Wang et.al. 2609.34135v1 null
2026-09-28 JET: Judge-Guided Evolution at Test Time for Agent Programs Yao Long Teng et.al. 2609.34126v1 null
2026-09-28 Do World Models Learn Global Understanding? Alexander Detkov et.al. 2609.34058v1 null
2026-09-27 Uncovering shortcut learning in audio classifiers by discovering recurring concepts in temporal explanations Cecilia Bolaños et.al. 2609.34030v1 null
2026-09-27 Safety Reconstructed: Generative Modeling via Masked Diffusion Builds Strong Safety Guardrails Gert Lek et.al. 2609.33634v1 null
2026-09-27 Learning Dynamics of Continual Learning: A Unified View of Data Attribution, Forgetting, and Plasticity Loss Yi Ren et.al. 2609.33620v1 null
2026-09-27 Temporal Graph Learning of Wearable Actigraphy and Sleep Traces for Modelling Adolescent Crystallized Intelligence Md. Tanvir Rahman et.al. 2609.33428v1 null
2026-09-27 Explainable Deep Learning of Resting-State Functional Connectomes Reveals Network Biomarkers of Adolescent Intelligence Md. Tanvir Rahman et.al. 2609.33422v1 null
2026-09-27 Decoupling Token Roles in Autoregressive Pretraining Suqin Yuan et.al. 2609.33405v1 null
2026-09-27 The Error You See Is Not the Error You Made: Progression-aware Reasoning Origin for Reasoning Error Localization Yiguo Wang et.al. 2609.33297v1 null
2026-09-27 CORTEX: A Verified Experience Layer for Generalist Agents Garapati Keerthana et.al. 2609.33260v1 null
2026-09-27 FOCUS: Benchmarking Retinal Model Generalization from Foundation Vision Encoders to Multimodal LLMs David Restrepo et.al. 2609.33158v1 null
2026-09-26 Relative Generalization Invariance of LLM Pretraining Fengzhuo Zhang et.al. 2609.33016v1 null
2026-09-26 DynamicDx: Evaluating Evidence Acquisition in Video-Based Diagnosis Jiahui Li et.al. 2609.32957v1 null
2026-09-26 Counterfactual Self-Evolving Agents for Evidence-Grounded Reasoning Xing Han et.al. 2609.32870v1 null
2026-09-26 FinancialAuditBench: Benchmark Construction under Differential Privacy Using Real-World Priors Jerry Huang et.al. 2609.32835v1 null
2026-09-26 Mandela-Bench: Multimodal Models Remember Canonical Images Instead of Seeing Them Yicheng Bao et.al. 2609.32763v1 null
2026-09-26 What Would Falsify It? A Variable Specific Evidence Standard for Mechanistic Claims About Self Explanation Arshia Eftekhari zadeh et.al. 2609.32670v1 null
2026-09-26 When Users Change Their Minds: Measuring and Repairing Intent Drift in LLM Agents Yanjie Zhang et.al. 2609.32520v1 null
2026-09-26 Explaining Textual Entailment with Lexical Entailments: Using LLMs to Supply Lexical Relations for Formal Proofs Jorryt de Jong et.al. 2609.32491v1 null
2026-09-26 Superposed Inference for Hyperdimensional Computing Quanling Zhao et.al. 2609.32320v1 null
2026-09-26 Why Directly Learning Periodic Trajectories Can Fail Kaixin Zheng et.al. 2609.32254v1 null
2026-09-26 A bilingual AI audiologist built through rubric-guided playbook induction outperforms human audiologists in a blinded evaluation of simulated cases Linkai Li et.al. 2609.32220v1 null
2026-09-26 Evaluating Single and Multi-Omics Based Explainable Artificial Intelligence (MOXAI) for Molecular Subclass Classification of Adult-Type Diffuse Gliomas Md Zahangir Alom et.al. 2609.32190v1 null
2026-09-26 REALM: Regime-Switching, Explainable, and Activation-Induced Linear Models Xiaoran Cheng et.al. 2609.32141v1 null
2026-09-25 Reasoning Concentrates Errors, and Self-Consistency Never Notices Asaad Althoubi et.al. 2609.32035v1 null
2026-09-25 A Safety-Bounded SDC-to-MCP Gateway for Medical AI Agents Bennet Gerlach et.al. 2609.31358v1 null
2026-09-25 DriveHierarchy: A Benchmark for Diagnosing VLM Driving Capabilities from Open-Loop Understanding to Closed-Loop Execution Chengkai Xu et.al. 2609.31814v1 null
2026-09-25 Rethinking Data Quality for AI-Driven Systems: Evidence from Practitioner Interviews Hariharan Gopinath et.al. 2609.31191v1 null
2026-09-25 Bayesian Optimization with Fisher Information Geometry: Gradient Bounds and Trust-Region Methods Saksham Kiroriwal et.al. 2609.31107v1 null

Abstracts

MIRTO: a registration-gated, multiverse-tested evaluation protocol for unsupervised anomaly segmentation in brain MRI

2610.02136v1 by Negin Kafee Hernashki, Soumick Chatterjee

Unsupervised anomaly detection (UAD) methods for brain MRI are ranked by a single score, yet that score rests on choices that are rarely reported: how each anomaly map is aligned with the reference, how and on which data the threshold is set, and which false-positive budget, metric, aggregation and lesion definition are used. We present MIRTO, an evaluation protocol that makes these choices explicit and measures their effect. It gates the geometry of every comparison with a registration check and label-free diagnostics of known power, sets thresholds on validation data alone and reports the false-positive volume actually realised on test, repeats each comparison over 15,552 defensible evaluation pipelines, and attaches paired subject-bootstrap intervals with multiplicity control. Applied to four UAD methods trained on the same healthy data and tested on 312 BraTS 2020 subjects, MIRTO showed that an axis-order mismatch between stored maps and the reference lowered a diffusion model's voxel AUROC from 0.873 to 0.583 whilst barely moving its slice-level AUROC. Within each metric, the method explained at least 0.95 of the variance in voxel AUROC and AUPRC and 0.77 in Dice, but only 0.14 in lesion sensitivity, where the lesion definition and hit criterion dominated. A Dice advantage that was significant at validation thresholds vanished at equal realised false-positive burden, and an exact identity attributes it to threshold transfer. A training-free change to REFLECT's latent aggregation raised Dice at equal burden by 0.052. Nine hypotheses were tested against explicit criteria; because the same cohort served to develop the protocol, all inference is exploratory.

摘要:未監督異常檢測(UAD)方法對於腦部 MRI 的排名是基於單一分數,但該分數依賴於鮮少報告的選擇:每個異常圖與參考的對齊方式、如何以及基於哪些數據設置閾值,以及使用哪種假陽性預算、指標、聚合和病變定義。我們提出了 MIRTO,一種評估協議,使這些選擇變得明確並測量其影響。它通過註冊檢查和無標籤診斷已知功率來限制每次比較的幾何,僅在驗證數據上設置閾值,並報告在測試中實際實現的假陽性體積,重複每次比較超過 15,552 條可辯護的評估管道,並附上成對的主體自助間隔及多重性控制。應用於四種基於相同健康數據訓練並在 312 名 BraTS 2020 受試者上測試的 UAD 方法,MIRTO 顯示存儲圖與參考之間的軸序不匹配使擴散模型的體素 AUROC 從 0.873 降至 0.583,同時幾乎不影響其切片級 AUROC。在每個指標中,該方法解釋了至少 0.95 的體素 AUROC 和 AUPRC 的變異,及 0.77 的 Dice,但在病變敏感性中僅為 0.14,病變定義和命中標準主導了這一結果。在驗證閾值下顯著的 Dice 優勢在相等的實現假陽性負擔時消失,並且一個精確的身份將其歸因於閾值轉移。對 REFLECT 的潛在聚合進行無訓練的變更,在相等負擔下將 Dice 提高了 0.052。針對明確標準測試了九個假設;由於相同的隊列用於開發該協議,所有推斷都是探索性的。

A Comparative Explainability Framework for DeBERTa-v3 in Zero-Shot Medical Abstract Classification

2610.02116v1 by Javier Diaz Esteban-Herreros, David Muñoz-Valero, Raquel Martínez-España, Jose M. Juarez, Juan Moreno-Garcia

A comparative explainability framework is presented to audit DeBERTa-v3 under zero-shot classification of medical abstracts. The work addresses the disagreement problem in Explainable Artificial Intelligence, where different attribution methods produce divergent explanations for the same input and prediction. A natural language inference engine is implemented over the Medical Abstracts corpus with five enriched hypotheses per diagnostic category and a balanced sample of one thousand texts per class. Five explanation methods are compared: SHAP and LIME as model-agnostic approaches, occlusion and Input x Gradient as deep-learning-specific approaches, and Attention x Gradient as a transformer-specific approach. Explanations are standardized through top-token attribution, and pairwise agreement is quantified using the Jaccard index. High predictive accuracy is achieved across well-defined clinical domains, whereas performance degrades under high semantic ambiguity. Explanatory stability directly mirrors predictive certainty, exhibiting strong convergence in univalent categories and a marked drop under diagnostic uncertainty. Furthermore, qualitative error auditing uncovers three systemic failure mechanisms: lexical hypersensitivity, semantic overlap, and loss of attribution coherence. The results support the combined use of several explanation methods and quantitative agreement metrics when auditing transformer-based models in medical text classification, and suggest prioritizing specific clinical ontologies over broad diagnostic labels.

摘要:比較可解釋性框架被提出以審計 DeBERTa-v3 在醫學摘要的零樣本分類下。這項工作解決了可解釋人工智慧中的不一致問題,即不同的歸因方法對相同的輸入和預測產生不同的解釋。自然語言推理引擎在醫學摘要語料庫上實施,每個診斷類別有五個增強的假設,並且每個類別有一千篇文本的平衡樣本。比較了五種解釋方法:SHAP 和 LIME 作為模型無關的方法,遮蔽和輸入 x 梯度作為深度學習特定的方法,以及注意力 x 梯度作為Transformer特定的方法。通過頂部標記歸因標準化解釋,並使用 Jaccard 指數量化成對一致性。在明確定義的臨床領域中實現了高預測準確性,而在高語義模糊性下性能下降。解釋穩定性直接反映預測確定性,在單值類別中顯示出強烈的收斂,並在診斷不確定性下顯著下降。此外,定性錯誤審計揭示了三種系統性失效機制:詞彙過敏、語義重疊和歸因一致性的喪失。結果支持在醫學文本分類中審計基於Transformer的模型時,結合使用幾種解釋方法和定量一致性指標,並建議優先考慮特定的臨床本體論而非廣泛的診斷標籤。

Can AI Oversight Be Zero Knowledge?

2610.01995v1 by Alessandro Chiesa, Ziyi Guan, Burcu Yildiz

AI systems increasingly produce outputs from confidential data, such as a fitness-for-duty assessment from medical records or the predicted properties of a drug candidate from its secret structure. It is important to verify that such outputs are correct without revealing the underlying data. A recent line of work studies verification of AI outputs via interactive proofs and debate for oracle-aided computation, where correctness may depend on an oracle such as human judgment, a physical experiment, or the web. These works focus on verification by a verifier that runs much faster than the computation. However, such efficient verification is impossible for general oracle-aided computation, and these works therefore rely on additional assumptions. We focus instead on privacy: allowing the verifier to run in time polynomial in the computation, we ask whether interactive arguments for oracle-aided computation can be zero knowledge, so that the verifier learns nothing about the confidential data beyond the correctness of the output. We prove that, in general, they cannot. In the random oracle model, there are no zero-knowledge proofs for all oracle-aided computations, even if both the prover and the verifier are allowed to run much longer than the computation itself. The impossibility extends to debate, a canonical model for scalable oversight. On the positive side, we show that if the oracle attaches a cryptographic signature to each of its answers, then every oracle-aided computation can be verified in zero knowledge with an efficient prover and verifier, assuming only collision-resistant hash functions. Beyond privacy, this also gives an alternative approach to scalable oversight that relies neither on an honest opponent, as in debate, nor on the robustness of the computation, as in prior single-prover protocols.

摘要:AI 系統越來越多地從機密數據中產生輸出,例如從醫療記錄中進行的適任性評估或從其秘密結構中預測的藥物候選物的性質。驗證這些輸出是否正確而不透露底層數據是很重要的。最近的一系列研究通過互動證明和辯論研究 AI 輸出的驗證,用於有 oracle 輔助的計算,其中正確性可能依賴於 oracle,例如人類判斷、物理實驗或網絡。這些研究專注於由運行速度遠快於計算的驗證者進行的驗證。然而,對於一般的有 oracle 輔助計算,這樣的高效驗證是不可能的,因此這些研究依賴於額外的假設。我們則專注於隱私:允許驗證者在計算的多項式時間內運行,我們詢問有 oracle 輔助計算的互動論證是否可以是零知識的,以便驗證者不會學到關於機密數據的任何信息,除了輸出的正確性。我們證明,通常情況下,它們是不可能的。在隨機 oracle 模型中,對於所有有 oracle 輔助的計算,沒有零知識證明,即使證明者和驗證者都被允許運行的時間遠超過計算本身。這種不可能性擴展到辯論,這是一個可擴展監督的典型模型。從積極的一面來看,我們展示了如果 oracle 為其每個答案附加加密簽名,那麼每個有 oracle 輔助的計算都可以在零知識中進行驗證,並且有高效的證明者和驗證者,只假設碰撞抗性哈希函數。除了隱私之外,這還提供了一種可擴展監督的替代方法,既不依賴於誠實的對手(如辯論中),也不依賴於計算的穩健性(如以前的單證明者協議中)。

Mapping the RAG Landscape: A Four Axis Taxonomy of Efficiency, Defense, Interactivity, and Reasoning

2610.01936v1 by Meghana Sunil, Shravya V, Shravan Venkatraman, Joe Dhanith PR

Large Language Models (LLMs) have demonstrated remarkable fluency across many tasks but remain limited by their static, parameter bound knowledge and their susceptibility to hallucinating information. Retrieval Augmented Generation (RAG) addresses these issues by incorporating external retrieval into the generation process, grounding model outputs in verifiable and up to date sources. While prior surveys primarily focus on core RAG architectures and standard pipelines, recent research explores broader challenges and capabilities that extend beyond these foundational designs. This survey provides a consolidated and structured examination of contemporary RAG developments, organizing the field into a four axis taxonomy: improving retrieval efficiency, strengthening robustness and security, supporting user driven and interactive workflows, and enabling multi step or complex reasoning. We formalize key components of the RAG framework and review methods spanning dense and sparse retrieval, fusion strategies, embedding optimizations, and reinforcement learning based retrieval policies, highlighting how these advances influence practical deployment and system design. We also synthesize evaluation practices, domain specific applications, and architectural variants such as Naive, Advanced, and Modular RAG. Finally, we outline persistent challenges related to retrieval quality, reliability, domain adaptation, scalability, and explainability, and identify opportunities for building RAG systems that are more reliable, adaptable, and transparent.

摘要:大型語言模型(LLMs)在許多任務中展現了卓越的流暢性,但仍然受到靜態的、參數限制的知識以及對虛假信息的易感性的限制。檢索增強生成(RAG)通過將外部檢索納入生成過程來解決這些問題,使模型輸出基於可驗證且最新的來源。雖然之前的調查主要集中在核心RAG架構和標準流程上,但最近的研究探討了超越這些基礎設計的更廣泛挑戰和能力。本調查提供了一個當代RAG發展的綜合和結構化檢視,將該領域組織為四個軸向的分類法:提高檢索效率、加強穩健性和安全性、支持用戶驅動和互動工作流程,以及實現多步驟或複雜推理。我們正式化了RAG框架的關鍵組件,並回顧了涵蓋密集和稀疏檢索、融合策略、嵌入優化和強化學習基於檢索政策的方法,強調這些進展如何影響實際部署和系統設計。我們還綜合了評估實踐、特定領域的應用以及如Naive、Advanced和Modular RAG等架構變體。最後,我們概述了與檢索質量、可靠性、領域適應性、可擴展性和可解釋性相關的持續挑戰,並確定了構建更可靠、可適應和透明的RAG系統的機會。

Removing spurious minima for planar features by skip connections

2610.01728v1 by Jakob Paul Zimmermann, Moritz Grillo, Andrei Balakin, Georg Loho

Understanding loss landscapes is central to explaining neural-network training, yet their structure remains only partially understood even in simple models. We study the Gaussian population loss of shallow, bias-free ReLU networks in the teacher--student setting. This provides a simple model for studying essential aspects such as feature learning and overparameterization. For teacher networks with positive output weights and planar features, we show that including a learned linear skip removes all spurious local minima with non-negative student output weights once the student network is at least as wide as the teacher network. In contrast, without the skip, we construct a fixed teacher network with positive output weights and only three hidden neurons in input dimension two whose spurious local minima persist at every student width at least three. Thus, a learned linear skip can remove spurious minima that persist under arbitrary overparameterization. Furthermore, we show that a positive output weight student network always learns the subspace spanned by the teacher features: student features at local minima with non-negative student output weights lie in the span of the teacher features. For ReLU networks in two dimensions, even heavily overparameterized student networks have effective width controlled by the teacher width: every critical point with positive student output weights has at most twice as many distinct student feature directions as teacher neurons. Finally, we transfer the benignity result to empirical minima over parameter balls of any prescribed radius, with the required sampling accuracy depending on that radius.

摘要:理解損失景觀對於解釋神經網絡訓練至關重要,然而即使在簡單模型中,它們的結構仍然只有部分被理解。我們研究教師-學生設置中淺層、無偏的ReLU網絡的高斯族群損失。這提供了一個簡單的模型來研究如特徵學習和過度參數化等基本方面。對於具有正輸出權重和平面特徵的教師網絡,我們顯示包含學習的線性跳過可以消除所有具有非負學生輸出權重的虛假局部最小值,只要學生網絡的寬度至少與教師網絡一樣寬。相反,如果不使用跳過,我們構造了一個固定的教師網絡,其具有正輸出權重且在輸入維度為二的情況下只有三個隱藏神經元,這樣的虛假局部最小值在每個學生寬度至少為三的情況下持續存在。因此,學習的線性跳過可以消除在任意過度參數化下持續存在的虛假最小值。此外,我們顯示具有正輸出權重的學生網絡總是學習由教師特徵所跨越的子空間:在具有非負學生輸出權重的局部最小值下,學生特徵位於教師特徵的跨度內。對於二維的ReLU網絡,即使是高度過度參數化的學生網絡,其有效寬度也受到教師寬度的控制:每個具有正學生輸出權重的臨界點最多有教師神經元的兩倍不同學生特徵方向。最後,我們將良性結果轉移到任何指定半徑的參數球上的經驗最小值,所需的取樣精度取決於該半徑。

MCIR: A Feature Dependence-Aware Explainability Method with Reliability Guarantees

2610.01641v1 by Poushali Sengupta, Sabita Maharjan, Frank Eliassen, Shashi Raj Pandey, Yan Zhang

Modern machine-learning models often contain strongly dependent or redundant features, making feature attribution difficult because shared predictive information can be distributed across correlated predictors. Existing methods such as SHAP, LIME, HSIC, MI/CMI, and SAGE may therefore produce unstable rankings under multicollinearity or near-duplicate predictors. We propose the Mutual Correlation Impact Ratio Method (MCIR-M), a dependence-aware global feature-importance approach that quantifies the unique predictive information contributed by each feature beyond a selected dependence neighbourhood. MCIR-M introduces the Mutual Correlation Impact Ratio (MCIR), which conditions each feature on strongly dependent neighbours and computes a normalized ratio of conditional to block-level information. The population score lies in [0,1] and equals zero under exact conditional redundancy. We also introduce a lightweight estimation procedure that computes MCIR using a fraction of the available data and evaluates agreement with full-data explanations. Across controlled synthetic redundancy experiments and the UCI HAR benchmark, MCIR shows dependence-aware ranking behaviour, with its clearest advantage under injected near-duplicate predictors. Comparisons with independent and conditional SHAP, SAGE, HSIC, MI-based scores, and CIR-family baselines are mixed across real-data criteria. Reduced explanation samples lower computational burden in the evaluated configurations, while agreement with full-data explanations is assessed separately through ranking, head-set, and faithfulness diagnostics. Overall, MCIR-M provides a practical dependence-aware diagnostic for global explanation under strong feature dependence.

摘要:現代機器學習模型通常包含強相關或冗餘的特徵,使得特徵歸因變得困難,因為共享的預測信息可能分佈在相關的預測變數之間。現有的方法如SHAP、LIME、HSIC、MI/CMI和SAGE在多重共線性或近乎重複的預測變數下可能因此產生不穩定的排名。我們提出了互相關影響比率方法(MCIR-M),這是一種考慮依賴性的全局特徵重要性方法,量化每個特徵在選定的依賴鄰域之外所貢獻的獨特預測信息。MCIR-M引入了互相關影響比率(MCIR),該比率在強依賴的鄰居上對每個特徵進行條件化,並計算條件信息與區塊級信息的標準化比率。該人口得分位於[0,1]之間,並在精確的條件冗餘下等於零。我們還引入了一種輕量級的估計程序,該程序使用部分可用數據計算MCIR,並評估與全數據解釋的一致性。在受控的合成冗餘實驗和UCI HAR基準測試中,MCIR顯示出考慮依賴性的排名行為,其在注入的近重複預測變數下的優勢最為明顯。與獨立和條件SHAP、SAGE、HSIC、基於MI的得分以及CIR系列基準的比較在真實數據標準下是混合的。在評估的配置中,減少的解釋樣本降低了計算負擔,而與全數據解釋的一致性則通過排名、頭部集和忠實性診斷單獨評估。總體而言,MCIR-M為強特徵依賴下的全局解釋提供了一種實用的考慮依賴性的診斷方法。

What Makes Something Hard(er)? Explaining Question Difficulty in Natural Language

2610.01627v1 by Peng Cui, Qiaoyuan Zheng, Rudolf Debelak, Mrinmaya Sachan

Difficulty is one of the most fundamental properties of a question: it determines whether the question can meaningfully discriminate between models of differing ability. Although a variety of methods can now estimate or predict difficulty automatically, they yield only a single descriptive number, with no account of the underlying factors that make a question difficult in the first place. In this work, we propose a data-driven approach that automatically generates and validates natural-language hypotheses explaining what makes one question harder than another. We first estimate each item's difficulty from the responses of a large pool of LLMs using Item Response Theory. We then sample contrasting sets of easy and hard questions and prompt an LLM to propose candidate explanations of the difference, which are subsequently validated and selected on held-out questions. Experimental results across three datasets spanning mathematical, logical, and commonsense reasoning show that our method produces interpretable and predictive hypotheses. On their own, they predict the difficulty of unseen questions competitively with, or better than, advanced black-box difficulty regressors; used as additional features, they further improve those regressors, implying that they discover difficulty signals that existing models fail to capture. Moreover, we demonstrate that editing questions according to a hypothesis can shift their measured difficulty in the expected direction, indicating that the discovered hypotheses are causally valid difficulty factors rather than post-hoc descriptions. Our approach thus turns a purely descriptive difficulty score into actionable statements.

摘要:困難度是問題最基本的特性之一:它決定了問題是否能夠有意義地區分不同能力的模型。雖然現在有多種方法可以自動估計或預測困難度,但它們僅產生一個描述性的數字,並未考慮使問題變得困難的潛在因素。在這項工作中,我們提出了一種數據驅動的方法,自動生成和驗證自然語言假設,解釋為什麼一個問題比另一個問題更難。我們首先使用項目反應理論從大量大型語言模型的回應中估計每個項目的困難度。然後,我們抽取一組對比的簡單和困難問題,並提示一個大型語言模型提出候選解釋這些差異,這些解釋隨後在保留的問題上進行驗證和選擇。跨越數學、邏輯和常識推理的三個數據集的實驗結果顯示,我們的方法產生了可解釋且具有預測性的假設。僅憑這些假設,它們能夠與先進的黑箱困難回歸模型競爭地預測未見問題的困難度;作為額外特徵使用時,它們進一步改善了這些回歸模型,這意味著它們發現了現有模型未能捕捉的困難信號。此外,我們證明根據假設編輯問題可以將其測量的困難度朝預期方向轉變,這表明所發現的假設是因果有效的困難因素,而非事後描述。因此,我們的方法將純粹描述性的困難分數轉化為可行的陳述。

Measuring the Stability Assumption Behind Action Chunking

2610.01626v1 by Aryan Goyal

Action chunking improves the performance of policies learned by behavioural cloning, and several mechanisms have been proposed to explain why, including temporal consistency, horizon reduction, representation learning, and reduced error compounding. We instead study what happens to an action error once it enters the system. At each state, we inject a small action error and measure how fast it grows or shrinks under two execution regimes: open-loop, where the rest of the chunk is replayed without replanning, and closed-loop, where the policy replans after the perturbation. The fitted rate labels each state as contracting, expanding, or unresolved. Across twelve manipulation tasks from three benchmark suites, we find that confidently stable states are rare, while error amplification is common among states whose propagation rate can be resolved. We further find that the measured propagation rate depends strongly on the fitting horizon: amplification is typically front-loaded, so short windows can overestimate longer-horizon propagation. Finally, we train predictors on these labels and find that a state's open-loop regime can be recovered from camera frames and proprioception alone, while its closed-loop propagation is only partially recoverable because it also depends on how the policy acts after the perturbation. These results suggest that error-compounding arguments alone do not provide a complete account of action chunking: neither passive open-loop dynamics nor policy replanning consistently contracts an injected error, and replanning rarely turns open-loop amplification into confident contraction. This suggests that closed-loop reactivity should be trained explicitly, using perturbation- and tree-coverage-oriented training to expose policies to deviations they must recover from, rather than expected to emerge reliably from standard imitation learning.

摘要:行動分塊改善了通過行為複製學習的政策的表現,並提出了幾種機制來解釋原因,包括時間一致性、視野縮減、表示學習和減少錯誤累積。我們則研究一旦行動錯誤進入系統後會發生什麼。在每個狀態下,我們注入一個小的行動錯誤,並測量在兩種執行模式下它的增長或縮小速度:開環模式,其中其餘的分塊在不重新規劃的情況下重播;以及閉環模式,在擾動後政策重新規劃。擬合速率將每個狀態標記為收縮、擴張或未解決。在三個基準套件的十二個操作任務中,我們發現自信穩定的狀態是稀有的,而錯誤放大在其傳播速率可以解決的狀態中是常見的。我們進一步發現,測量的傳播速率強烈依賴於擬合視野:放大通常是前置的,因此短時間窗口可能會高估長視野的傳播。最後,我們對這些標籤訓練預測器,並發現狀態的開環模式可以僅從攝像頭幀和本體感知中恢復,而其閉環傳播僅部分可恢復,因為它還取決於政策在擾動後的行為。這些結果表明,僅僅依賴錯誤累積的論點並不能完整解釋行動分塊:無論是被動的開環動力學還是政策重新規劃都不一致地收縮注入的錯誤,而重新規劃很少將開環放大轉化為自信的收縮。這表明閉環反應性應該明確進行訓練,使用擾動和樹覆蓋導向的訓練來使政策暴露於必須恢復的偏差,而不是期待從標準模仿學習中可靠地出現。

Exact Distinguishability in Non-Markovian Decision Processes

2610.01527v1 by Kabir Murjani, Nisarg Patel

Non-Markovian environments are often modeled as Regular Decision Processes (RDPs), where dynamics depend on the interaction history through a finite automaton. Existing offline guarantees for RDPs rely on a distinguishability assumption on the behaviour policy but provide no means of verifying it. When the assumption is violated, distinct models may explain the data equally well. We study when data collected under a fixed behaviour policy can distinguish two candidate RDPs. We prove that the posterior odds between observationally equivalent candidates remain equal to the prior odds at every sample size, even when the policy visits every automaton state, and verify both results formally in Lean 4. We then characterize this equivalence exactly and derive PEC, an algorithm that decides it in time linear in the size of the product automaton. The distinguishability assumption of prior work fails on three of our four test environments, and the experiment identified by PEC restores it in each case.

摘要:非馬可夫環境通常被建模為正則決策過程(RDP),其動態依賴於通過有限自動機的互動歷史。現有的RDP離線保證依賴於行為策略的可區分性假設,但未提供驗證該假設的手段。當假設被違反時,不同的模型可能同樣能夠解釋數據。我們研究在固定行為策略下收集的數據何時能區分兩個候選RDP。我們證明觀察上等價的候選者之間的後驗比率在每個樣本大小下保持等於先驗比率,即使該策略訪問了每個自動機狀態,並在Lean 4中正式驗證了這兩個結果。我們然後準確地表徵這種等價性,並推導出PEC,一種在產品自動機大小的線性時間內決定它的算法。先前工作的可區分性假設在我們四個測試環境中的三個失敗,而PEC識別的實驗在每個案例中恢復了它。

OpenMTB-Audit: Exposing Over-Refusal and Clinical Expert Perspectives in LLM-Based Molecular Tumor Board Safety Evaluation

2610.01497v1 by Negin Ashrafi, Jia Luo, Stacey M. Frumm, Roxana Daneshjou

Molecular tumor boards integrate genomic findings, clinical context, and therapeutic evidence to support precision oncology. As AI enters this workflow, a key safety challenge is distinguishing truly unsupported recommendations from evidence-supported options that still require oncologist review because of incomplete information, poor ECOG performance status, or other clinical caveats. We introduce OpenMTB-Audit, an open-source benchmark of 500 synthetic non-small cell lung cancer cases spanning five adversarial error categories and four safety labels: Supported, Partially Supported, Unsupported, and Insufficient Information. Across eight large language model configurations, we identify pervasive over-refusal: all LLM configurations failed to retain the Partially Supported label in 83.3-100% of true Partially Supported cases, achieving high aggregate safety scores through label collapse rather than clinically calibrated reasoning. To address this limitation, we developed MTB-AuditAgent, a deterministic seven-module framework separating evidence verification, missing-information detection, safety classification, and abstention. It reduces over-refusal to 6.7% and achieves 91.2% accuracy (95% CI: 88.6-93.6%). A two-oncologist annotation study found disagreement concentrated at the boundary between information sufficiency and treatment optimization, underscoring the need to preserve clinically meaningful distinctions.

摘要:分子腫瘤委員會整合基因組發現、臨床背景和治療證據,以支持精準腫瘤學。隨著人工智慧進入這一工作流程,一個主要的安全挑戰是區分真正不被支持的建議與仍需腫瘤醫生審查的證據支持選項,因為信息不完整、ECOG表現狀態不佳或其他臨床警告。我們介紹了OpenMTB-Audit,一個開源基準,包含500個合成的非小細胞肺癌案例,涵蓋五個對抗性錯誤類別和四個安全標籤:支持、部分支持、不支持和信息不足。在八種大型語言模型配置中,我們發現普遍的過度拒絕:所有LLM配置在83.3-100%的真實部分支持案例中未能保留部分支持標籤,通過標籤崩潰而非臨床校準推理獲得高整體安全分數。為了解決這一限制,我們開發了MTB-AuditAgent,一個確定性的七模塊框架,將證據驗證、缺失信息檢測、安全分類和放棄分開。它將過度拒絕降低到6.7%,並達到91.2%的準確率(95% CI:88.6-93.6%)。一項由兩位腫瘤醫生進行的標註研究發現,分歧集中在信息充分性和治療優化之間的邊界,強調了保留臨床上有意義的區別的必要性。

Multi-Party Backchannel Prediction: a Diagnosis, a Benchmark, and a Ceiling

2610.01488v1 by Mohammed Hafsati, Ahmed Loughzali

Backchannel prediction has been studied almost entirely in dyadic conversation. We introduce a multi-party benchmark based on the AMI corpus, comprising 682 masked-listener views from 171 meetings, 190 speakers, and 18,697 backchannel events, with a person-disjoint held-out split. A state-of-the-art dyadic model applied zero-shot to meeting audio performs at chance (AUROC 0.499); nevertheless, its frozen acoustic features remain informative: a linear probe reaches 0.704, and retraining the predictor raises performance to 0.751. Retraining reveals a second limitation. Listener conditioning improves prediction for listeners seen during training but not for unseen listeners, and the gap remains under capacity reduction, listener-adversarial training, per-listener adaptation, and oracle lexical conditioning. Adversarial training removes only part of the speaker-identity information, while stronger removal hurts prediction, suggesting that identity is entangled with cues that are useful for backchanneling. A within-model control helps explain this pattern: with the same features and data splits, turn-onset prediction transfers to unseen listeners, while backchannel prediction does not. Backchannel rates also vary about twice as much across individuals as turn-onset rates. Since backchannels occupy only about 1% of frames, frame-level F1 is strongly affected by the base rate. We therefore report AUROC alongside event-F1 on listener-active regions. We release the benchmark and evaluation tools at https://github.com/HafsatiMohammed/bc_multiparty_release.

摘要:回饋通道預測幾乎完全在雙人對話中進行了研究。我們基於AMI語料庫引入了一個多方基準,包含來自171次會議的682個被遮蔽的聆聽者視角、190位講者和18,697個回饋通道事件,並設有一個人員不重疊的保留分割。一個最先進的雙人模型在會議音頻上進行零樣本應用,表現與隨機相當(AUROC 0.499);然而,其凍結的聲學特徵仍然具有信息性:線性探測器達到0.704,並且重新訓練預測器將性能提高到0.751。重新訓練揭示了第二個限制。聆聽者的條件化改善了對在訓練期間出現的聆聽者的預測,但對未見過的聆聽者則沒有改善,並且在容量減少、聆聽者對抗訓練、每位聆聽者的適應和oracle詞彙條件下,這一差距依然存在。對抗訓練僅去除了部分講者身份信息,而更強的去除會損害預測,這表明身份與對回饋通道有用的線索交織在一起。模型內部控制有助於解釋這一模式:在相同的特徵和數據分割下,轉換開始預測能夠轉移到未見過的聆聽者,而回饋通道預測則無法。回饋通道的比率在個體之間的變化大約是轉換開始比率的兩倍。由於回饋通道僅佔約1%的幀,因此幀級F1受到基礎比率的強烈影響。因此,我們在聆聽者活躍區域報告AUROC和事件-F1。我們在https://github.com/HafsatiMohammed/bc_multiparty_release發布了基準和評估工具。

Detect, Explain, Interpret: An End-to-End Benchmark for Time Series Anomaly Detection, Explainability and Interpretability

2610.01168v1 by Roberto Stanzione, Jules Barbe, Magali Parrino, Jérémie Fourmann, Paul Boniol

Time Series Anomaly Detection has received increasing attention, driven by the growing availability of complex time series data. This surge has led to the development of numerous detection methods, as well as a variety of benchmarks aimed at thoroughly evaluating their performance. However, most existing detectors remain largely agnostic to domain context, overlooking explainability and interpretability. One of the main reasons for this gap is that current benchmarks primarily focus on detection accuracy, and only few of them evaluate spatial explainability. Moreover, no benchmark currently provides sufficiently rich semantic annotations to support the generation of human-understandable interpretations of anomalies. To address these limitations, we introduce SHAD (Scality High-dimensional Anomaly Detection benchmark), a fully annotated benchmark composed of 215 multivariate, high-dimensional time series collected from real-world distributed cloud storage systems operated by Scality. The proposed dataset includes rich contextual information, covering three families of anomalies with varying degrees of severity. As further contribution, we provide a foundation for future work by evaluating baseline methods for Detection, Explainability, and Interpretability, covering all stages of a TSAD pipeline. For Detection, we benchmark a wide range of existing anomaly detectors, testing their effectiveness on the proposed real-world dataset. Then, we consider explainability by evaluating whether measuring the contribution of each dimension in the generated anomaly score can provide accurate anomaly attributions. Finally, for interpretability, we investigate the effectiveness of frozen LLM baselines in localizing and interpreting anomalies.

摘要:時間序列異常檢測受到越來越多的關注,這是由於複雜時間序列數據的日益可用性所驅動。這一增長促使了許多檢測方法的發展,以及各種基準的出現,旨在徹底評估它們的性能。然而,大多數現有的檢測器在很大程度上對領域上下文保持無知,忽視了可解釋性和可理解性。這一差距的主要原因之一是目前的基準主要集中在檢測準確性上,只有少數評估空間可解釋性。此外,目前沒有任何基準提供足夠豐富的語義註釋,以支持生成易於人類理解的異常解釋。為了解決這些限制,我們引入了SHAD(Scality高維異常檢測基準),這是一個完全註釋的基準,由215個來自Scality運營的真實分佈雲存儲系統的多變量高維時間序列組成。所提出的數據集包括豐富的上下文信息,涵蓋三類具有不同嚴重程度的異常。作為進一步的貢獻,我們通過評估檢測、可解釋性和可理解性的基線方法,為未來的工作提供了一個基礎,涵蓋了TSAD管道的所有階段。對於檢測,我們基準測試了各種現有的異常檢測器,測試它們在所提出的真實世界數據集上的有效性。然後,我們通過評估在生成的異常分數中測量每個維度的貢獻是否能提供準確的異常歸因來考慮可解釋性。最後,對於可理解性,我們調查了凍結的LLM基線在定位和解釋異常方面的有效性。

CineMR: Tool-Integrated Vision-Language Reasoning for Quantitative Cardiac MRI Assessment

2610.01166v1 by Kunyang Li, Hai Nguyen, Joshua Lowe, Chenguang Zhao, Peace C. Madueme, Mehdi Hedjazi Moghari, Mubarak Shah, Pegah Khosravi, Yuzhang Zhang

Cardiovascular magnetic resonance (CMR), including cine imaging, is a reference standard for the noninvasive assessment of cardiac morphology and ventricular function. Cine CMR interpretation integrates qualitative visual assessment with quantitative measurements of ventricular volumes, ejection fraction, myocardial mass, wall thickness, and regional wall motion. Current medical vision-language models (VLMs) cannot reliably derive quantitative measurements from multidimensional cine images without analysis tools. We present CineMR, a tool-augmented VLM that invokes cardiac image-analysis tools and integrates their outputs into interleaved reasoning for quantitative CMR assessment. We also construct a multi-cohort visual question answering benchmark covering quantitative metric extraction, multiclass diagnosis, and differential diagnosis, together with tools for segmentation, phase selection, volumetry, morphometry, and regional wall motion analysis. CineMR is trained with supervised fine-tuning (SFT) on tool-interaction traces followed by Group Relative Policy Optimization (GRPO) with conditional tool-use rewards. On the multi-cohort cine CMR benchmark, CineMR achieves 35.9% pass@1 and 58.9% pass@4, compared with 1.5% pass@1 for the Qwen3-VL-8B backbone and 0.0% and 7.0% pass@1 for LLaVA-Med v1.5 and MedGemma-4B, respectively. Correct tool invocation reaches 99.8% after GRPO, up from 78.9% after SFT. Live tool outputs improve ventricular measurement accuracy by 20.4--23.7% over direct model predictions, and removing all tools reduces pass@1 from 35.9% to 27.9%. These results highlight the importance of reliable tool use for quantitative cine CMR reasoning and support CineMR as a promising approach for assistive cardiac image assessment. Code, benchmark resources, and model weights are available at https://github.com/AI-MIND-Lab/CineMR.

摘要:心血管磁共振(CMR),包括動態影像,是非侵入性評估心臟形態和心室功能的參考標準。動態 CMR 解釋將定性視覺評估與心室體積、射血分數、心肌質量、壁厚和區域壁運動的定量測量相結合。目前的醫療視覺-語言模型(VLMs)在沒有分析工具的情況下,無法可靠地從多維動態影像中推導出定量測量。我們提出了 CineMR,一種增強工具的 VLM,調用心臟影像分析工具並將其輸出整合到交錯推理中,以進行定量 CMR 評估。我們還構建了一個涵蓋定量指標提取、多類別診斷和鑑別診斷的多隊列視覺問答基準,並提供分割、相位選擇、體積測量、形態測量和區域壁運動分析的工具。CineMR 在工具互動痕跡上進行了監督微調(SFT),隨後使用條件工具使用獎勵進行了群體相對策略優化(GRPO)。在多隊列動態 CMR 基準上,CineMR 的 pass@1 為 35.9%,pass@4 為 58.9%,而 Qwen3-VL-8B 的 pass@1 僅為 1.5%,LLaVA-Med v1.5 和 MedGemma-4B 的 pass@1 分別為 0.0% 和 7.0%。經過 GRPO 正確調用工具的比例達到 99.8%,而 SFT 後為 78.9%。實時工具輸出提高了心室測量的準確性,較直接模型預測提高了 20.4% 至 23.7%,而去除所有工具則使 pass@1 從 35.9% 降至 27.9%。這些結果突顯了可靠工具使用在定量動態 CMR 推理中的重要性,並支持 CineMR 作為輔助心臟影像評估的有前景方法。代碼、基準資源和模型權重可在 https://github.com/AI-MIND-Lab/CineMR 獲得。

What Can Analogy Tell Us About Artificial Consciousness?

2610.01002v1 by Keith J. Holyoak, Martin M. Monti

Who or what is conscious? Because subjective experience is directly accessible only in the first person, judgments about consciousness in other entities depend partly on analogy. Historically, such inferences have focused on nonhuman animals, but advances in artificial intelligence have raised the possibility of conscious AI. Here we develop a causal framework for evaluating such evidential analogies. The key distinction is between similarities in factors plausibly involved in generating consciousness and similarities in downstream behavioural or cognitive effects. Our framework weights source-target similarity by causal relevance while allowing for unknown causes, disabling differences and alternative routes to consciousness. Applied to biological systems, it explains why analogical support generally weakens with increasing causal distance from humans. Applied to contemporary AI, it suggests that behavioural similarity provides only limited evidence for consciousness because relevant causal correspondences remain poorly established. The framework also clarifies what evidence would strengthen claims of artificial consciousness.

摘要:誰或什麼是有意識的?因為主觀經驗僅在第一人稱中直接可得,對其他實體意識的判斷部分依賴於類比。歷史上,這種推斷主要集中在非人類動物上,但人工智慧的進步已經提高了有意識 AI 的可能性。在這裡,我們發展了一個評估這種證據類比的因果框架。關鍵的區別在於可能涉及生成意識的因素之間的相似性,以及下游行為或認知效應之間的相似性。我們的框架根據因果相關性對源目標相似性進行加權,同時考慮未知原因、禁用差異和通往意識的替代路徑。應用於生物系統,它解釋了為什麼類比支持通常隨著與人類的因果距離增加而減弱。應用於當代 AI,它表明行為相似性僅提供有限的意識證據,因為相關的因果對應仍然建立得不夠充分。該框架還闡明了什麼證據可以加強人工意識的主張。

When Reasoning Helps Action: Monitoring and Steering Chain-of-Thought in Vision-Language-Action Policies

2610.00601v1 by Sathwik Karnik, Joseph JR. Lee, Aryaman Gupta, Somil Bansal

Reasoning-enabled VLA policies expose chain-of-thought (CoT) traces that appear to explain and guide their actions, creating a potential interface for runtime safety through reasoning monitoring and correction. In this work, we define and operationalize two evaluation axes for assessing when this interface can improve embodied behavior: correctability, which measures whether unreliable reasoning can be detected and improved during generation, and actionability, which measures whether reasoning corrections produce behaviorally meaningful changes in the intended direction. To enable correctability, we introduce Token-level Reward for Utility-Steered Chain-of-Thought (TRUST), an offline-trained value model that predicts eventual reasoning correctness from partial prefixes and uses these estimates to monitor and selectively steer reasoning generation in frozen VLA policies. On the Alpamayo 1.5 driving VLA, TRUST monitors correctness with 88.9% accuracy and improves reasoning correctness from 75.9% to 90.0%. On a baseline-defined challenging subset in AlpaSim, TRUST reduces collision rate by 30.4% and maximum trajectory error by 11.5% relative to the unsteered policy, outperforming a compute-matched Best-of-4 baseline. On the DeepThinkVLA manipulation VLA, TRUST improves the correctness of grasp-state claims from 69.3% to 90.2% and action-choice claims from 68.8% to 85.9%, yet closed-loop task performance on LIBERO-Plus remains largely unchanged. Empirical analysis reveals intent-consistent behavioral effects in Alpamayo 1.5 but limited effects in DeepThinkVLA, helping interpret these different task-level outcomes. Together, our results show that gains in reasoning correctness do not automatically imply gains in embodied performance, motivating evaluation of correctability and actionability when using CoT as a runtime safety interface.

摘要:推理驅動的 VLA 政策揭示了思考過程(CoT)痕跡,這些痕跡似乎解釋並指導其行動,通過推理監控和修正創造了一個潛在的運行時安全介面。在這項工作中,我們定義並操作化了兩個評估軸,以評估何時這個介面可以改善具身行為:可修正性,衡量在生成過程中是否能檢測到不可靠的推理並加以改進;以及可行性,衡量推理修正是否能產生在預期方向上有意義的行為變化。為了實現可修正性,我們引入了基於效用驅動的思考過程的標記級獎勵(TRUST),這是一個離線訓練的價值模型,能夠從部分前綴預測最終的推理正確性,並利用這些估計來監控和選擇性地引導凍結的 VLA 政策中的推理生成。在 Alpamayo 1.5 驅動的 VLA 上,TRUST 以 88.9% 的準確率監控正確性,並將推理正確性從 75.9% 提高到 90.0%。在 AlpaSim 中的基準定義挑戰子集上,TRUST 相對於未引導政策將碰撞率降低了 30.4%,最大軌跡誤差降低了 11.5%,超越了計算匹配的 Best-of-4 基準。在 DeepThinkVLA 操作 VLA 上,TRUST 將抓取狀態聲明的正確性從 69.3% 提高到 90.2%,將行動選擇聲明的正確性從 68.8% 提高到 85.9%,然而在 LIBERO-Plus 上的閉環任務性能仍然基本保持不變。實證分析顯示在 Alpamayo 1.5 中存在意圖一致的行為效果,但在 DeepThinkVLA 中效果有限,這有助於解釋這些不同的任務級結果。綜合來看,我們的結果表明,推理正確性的提升並不自動意味著具身表現的提升,這促使在使用 CoT 作為運行時安全介面時評估可修正性和可行性。

Worse Together: How Performance Breaks Down in Multi-User Multi-Agent Teams

2610.00583v1 by Sahan Paliskara, Nattaput Namchittai, Andrew Lampinen

People are increasingly delegating tasks to AI agents, and those agents are increasingly encountering other people's agents over shared resources such as a codebase, a calendar, or a budget. When each agent acts for a different user with different goals, coordination often fails, and the group ends up worse off than if a single agent had acted for everyone. We study this multi-user, multi-agent setting across five frontier models and 77 scenarios in four environments: an API key environment in which agents share a compute budget, a clinic in which they share a calendar, a personal assistant environment in which they share a group order or booking, and a merge queue in which they share a release cutoff. In each scenario, we compare a single agent that serves every user (a coordinator) to a team in which each agent serves one user, with and without a communication channel between the agents. Teams deliver worse group outcomes than the coordinator in every environment: without a channel, they completely collapse in two environments, and even with one, coordination overhead creates substantial gaps. For example, in the personal assistant environment, the coordinator fulfills a targeted user request about twice as often as teams. We identify distinct behaviors associated with this poor group-level performance, including stalling as teams grow, overriding each other's actions, and fabricating claims. We find effective but environment-specific mitigations, such as a team lead, explicit procedural instructions, and a platform check that makes an agent read its peers' messages before committing. We will release the API key, clinic, and personal assistant environments as MAMUBench, comprising 74 scenarios for evaluating multi-user, multi-agent coordination.

摘要:人們越來越多地將任務委派給 AI 代理,而這些代理也越來越多地與其他人的代理在共享資源上相遇,例如代碼庫、日曆或預算。當每個代理代表不同的用戶且目標不同時,協調往往失敗,結果小組的情況比由單一代理為所有人行動時更糟。我們研究了這種多用戶、多代理的設定,涵蓋五個前沿模型和四個環境中的 77 種情境:一個 API 金鑰環境,在這裡代理共享計算預算;一個診所,在這裡他們共享日曆;一個個人助理環境,在這裡他們共享團體訂單或預訂;以及一個合併隊列,在這裡他們共享發佈截止時間。在每個情境中,我們將為每個用戶服務的單一代理(協調者)與每個代理服務一位用戶的團隊進行比較,並考慮代理之間是否有通信渠道。團隊在每個環境中提供的群體結果都比協調者差:在沒有渠道的情況下,他們在兩個環境中完全崩潰,即使有一個,協調開銷也會造成相當大的差距。例如,在個人助理環境中,協調者滿足目標用戶請求的頻率約為團隊的兩倍。我們識別出與這種低群體表現相關的不同行為,包括隨著團隊增長而停滯、覆蓋彼此的行動以及捏造聲明。我們發現有效但特定於環境的緩解措施,例如團隊負責人、明確的程序指示,以及一個平台檢查,使代理在提交之前閱讀其同伴的消息。我們將發布 API 金鑰、診所和個人助理環境作為 MAMUBench,包含 74 種情境以評估多用戶、多代理的協調。

No One Architecture Fits All: A Cross-Environment Evaluation of Hierarchical Red Team Agents

2610.00557v1 by Ayan Javeed Shaikh, Arunesh Sinha, Nathaniel D. Bastian, Ankit Shah

Autonomous red team agents increasingly stress-test AI-enabled cyber defenses by planning strategy and executing multistage attacks. Reinforcement learning (RL) and large language models (LLMs) offer complementary mechanisms for the planning and execution such agents require, and prior work has combined them in hybrid hierarchies. Yet a given architecture is typically developed and evaluated within a single environment, leaving open whether an observed advantage reflects a generally stronger decision mechanism or merely alignment with a particular setting. We address this gap with a controlled cross-environment comparison of two homogeneous hierarchical red team architectures: an RL planner with an RL executor (RL+RL) and an LLM planner with an LLM executor (LLM+LLM). We evaluate both against expert autonomous defenders in CybORG CAGE-4 and in Cyberwheel at two network scales, across 18 configurations under one unified disruption metric. We find a pronounced environment-dependent inversion. RL+RL wins the compact, densely rewarded CAGE-4 (78.5% disruption success versus 18.0% for the strongest LLM configuration) and the 100-host Cyberwheel network (81.0% versus 50.5%), while a pretrained cybersecurity LLM agent wins the larger, escalation-gated 1010-host Cyberwheel network (55.0% versus 0.0% for RL). A kill-chain analysis explains the inversion through architecture-specific bottlenecks that aggregate success rates conceal.In the 1010-host Cyberwheel network, RL discovers and compromises hosts but stalls at privilege escalation, whereas in CAGE-4, LLM agents obtain privileged access but rarely convert it into operational impact. These results indicate that conclusions drawn in a single environment may not generalize, and that hybrid planner-executor designs should be motivated by specific failure modes rather than the assumption that one architecture is universally preferable.

摘要:自主紅隊代理人越來越多地通過規劃策略和執行多階段攻擊來壓力測試 AI 驅動的網絡防禦。強化學習 (RL) 和大型語言模型 (LLMs) 提供了這些代理人所需的規劃和執行的互補機制,先前的工作已將它們結合在混合層級中。然而,給定的架構通常是在單一環境中開發和評估的,這使得觀察到的優勢是否反映出一般更強的決策機制,或者僅僅是與特定設置的一致性仍然是個未解之謎。我們通過對兩個同質層級紅隊架構進行受控的跨環境比較來解決這一空白:一個是具有 RL 執行者的 RL 規劃者 (RL+RL),另一個是具有 LLM 執行者的 LLM 規劃者 (LLM+LLM)。我們在 CybORG CAGE-4 和 Cyberwheel 中針對專家自主防禦者評估這兩者,並在兩個網絡規模下,根據一個統一的干擾指標進行 18 種配置的比較。我們發現了一個明顯的環境依賴性反轉。RL+RL 在緊湊、密集獎勵的 CAGE-4 中獲勝(78.5% 的干擾成功率對比最強 LLM 配置的 18.0%),以及 100 主機的 Cyberwheel 網絡(81.0% 對比 50.5%),而一個預訓練的網絡安全 LLM 代理在更大、升級限制的 1010 主機 Cyberwheel 網絡中獲勝(55.0% 對比 RL 的 0.0%)。一項殺鏈分析通過架構特定的瓶頸解釋了這一反轉,這些瓶頸會掩蓋成功率的聚合。在 1010 主機的 Cyberwheel 網絡中,RL 發現並攻陷主機,但在特權升級時停滯不前,而在 CAGE-4 中,LLM 代理獲得特權訪問,但很少將其轉化為操作影響。這些結果表明,在單一環境中得出的結論可能無法推廣,混合規劃者-執行者設計應該基於特定的失敗模式,而不是假設某一架構是普遍可取的。

EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights

2610.00492v1 by Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig

When Isaac Newton discovered the law of gravitation, he did so through an iterative process of analyzing observed data such as planetary patterns, finding the underlying mechanisms by describing patterns in mathematical equations, and refining his theory against the Moon's orbit, revealing the startling insight that the same force governs both falling apples and orbiting planets. Would it be possible for AI agents to make similar discoveries? To measure this ability, we introduce EurekaBench, a cross-domain benchmark that tests AI agents' ability to conduct long-horizon experiments and discover mechanisms that explain observations. We evaluate these mechanisms by the scientific insights that can be derived from them. EurekaBench contains an expert-verified set of 26 long-horizon tasks across neuroscience, computer science, chemistry, astrophysics, geophysics, and plasma physics, with a total of 306 scientific insights that the discovered mechanisms are expected to support. Our evaluation framework tests three axes of scientific discovery: agents' ability to follow known scientific constraints, the predictive accuracy of the discovered mechanisms, and whether these mechanisms yield scientific insights or inform future research. Our results show that current AI agents often overly fixate on predictive accuracy optimization, surpassing human scientists, while falling substantially short in deriving scientific insights.

摘要:當艾薩克·牛頓發現萬有引力定律時,他是通過一個迭代過程來分析觀察到的數據,如行星運行模式,通過在數學方程中描述模式來尋找潛在的機制,並根據月球的軌道來完善他的理論,揭示了驚人的見解:同一種力量支配著掉落的蘋果和運行的行星。人工智慧代理是否有可能做出類似的發現?為了衡量這種能力,我們引入了EurekaBench,一個跨領域的基準,測試人工智慧代理進行長期實驗和發現解釋觀察的機制的能力。我們通過從這些機制中得出的科學見解來評估這些機制。EurekaBench包含一組經專家驗證的26個長期任務,涵蓋神經科學、計算機科學、化學、天體物理學、地球物理學和等離子體物理學,總共有306個預期由發現的機制支持的科學見解。我們的評估框架測試科學發現的三個軸心:代理遵循已知科學約束的能力、發現機制的預測準確性,以及這些機制是否產生科學見解或為未來研究提供信息。我們的結果顯示,當前的人工智慧代理往往過於專注於預測準確性的優化,超越了人類科學家,但在推導科學見解方面則大幅不足。

CAS II: Symmetric Partitions as Kolmogorov Models

2609.40290v1 by Romie Banerjee

In algorithmic statistics a string x is explained by a finite set containing it, and Kolmogorov's structure function records the smallest such model at each level of complexity. Vereshchagin's strong models, those computable from the data by a total algorithm, are essentially the cells of simple partitions. We read a partition of binary strings as a hypothesis, with the cell containing x as its model, and develop algorithmic statistics over symmetric partitions: the orbit partitions of groups acting on strings. The Galois connection between subgroups and partitions gives each ambient group a lattice of symmetric partitions, with canonical certificates, canonical costs, and an algebra of hypotheses. The resulting structure function and symmetric sophistication measure which part of the regularity of x is symmetric. For the full symmetric group every partition is symmetric: cells recover all Kolmogorov models, cells of cheap partitions recover exactly the strong models, and normal and strange strings are characterized by symmetry. For GL(n,2) the cells are exactly the linearly homogeneous sets, so linear symmetry is a restricted model class. For nonzero x, the linear-symmetry structure function lies in a band between the sufficiency line and the trivial bound, and both edges are attained: there are stochastic normal strings whose simple structure is invisible to linear symmetry. We also give coordinates on the space of permutation groups: each group is an element of a Burnside ring (its type) together with a permutation (its placement), and restriction moves refine partitions via the Mackey formula. In these coordinates the collapse for the symmetric group is a statement about placement, a linear hypothesis is determined by its type up to n^2 bits, and the maximal gap theorem shows that any space of symmetry hypotheses small enough to search is small enough to miss simple structure.

摘要:在算法統計中,字符串 x 是由包含它的有限集合來解釋的,而 Kolmogorov 的結構函數記錄了每個複雜度級別下最小的這樣的模型。Vereshchagin 的強模型,即由總算法從數據中可計算出的模型,本質上是簡單劃分的單元。我們將二進制字符串的劃分視為一個假設,包含 x 的單元作為其模型,並在對稱劃分上發展算法統計:作用於字符串的群的軌道劃分。子群與劃分之間的 Galois 連接為每個環境群提供了一個對稱劃分的格,並附有典範證明、典範成本和假設的代數。由此產生的結構函數和對稱複雜度測量 x 的正則性中哪一部分是對稱的。對於完整的對稱群,每個劃分都是對稱的:單元恢復所有 Kolmogorov 模型,廉價劃分的單元正好恢復強模型,而正常和奇怪的字符串則以對稱性為特徵。對於 GL(n,2),單元正好是線性齊次集合,因此線性對稱性是一個受限的模型類。對於非零 x,線性對稱結構函數位於充分性線和微不足道界限之間的帶中,且兩個邊界均可達:存在隨機正常字符串,其簡單結構對線性對稱性是不可見的。我們還給出了置換群空間的坐標:每個群都是一個 Burnside 環的元素(其類型)以及一個置換(其位置),而限制移動通過 Mackey 公式細化劃分。在這些坐標中,對稱群的崩潰是關於位置的陳述,線性假設由其類型決定,最多 n^2 位,最大間隙定理顯示,任何足夠小以進行搜索的對稱假設空間都足夠小以錯過簡單結構。

Unlearnable, or Unmeasured? On the Reliability of Difficulty Labels in RLVR

2609.40115v1 by Chandak Chakma, Syed Nazmus Sakib, Nafiul Haque, Shifat E. Arman

Reinforcement learning with verifiable rewards (RLVR) has become an important approach for improving reasoning during post-training. Recent work suggests that some difficult prompts remain resistant to learning even when they occasionally produce correct solutions. We revisit this unlearnability phenomenon and find that the affected prompts do improve, at roughly one third of the learnable rate, while the difficulty-defined set used to study them is much less reproducible than expected. These difficulty labels are estimated from a limited number of sampled responses. Combining them across seeds can further change which prompts are selected instead of simply reducing measurement noise. We develop a sampling-based framework for quantifying this instability and determining how much evaluation is required for difficulty assignments to reproduce reliably. We also revisit the gradient-similarity evidence proposed to explain unlearnability and show that part of the observed separation arises because difficult prompts provide fewer correct rollouts from which their gradients can be estimated. Matching this sample count weakens the gradient difference but does not remove it. Overall, the slow-learning phenomenon survives our reanalysis, while both the prompts used to define it and the evidence used to explain it require more careful measurement.

摘要:強化學習與可驗證獎勵(RLVR)已成為改善後訓練推理的重要方法。最近的研究表明,即使某些困難的提示偶爾產生正確的解決方案,它們仍然對學習具有抵抗力。我們重新檢視這一不可學習現象,發現受影響的提示確實有所改善,改善速度約為可學習速率的三分之一,而用來研究它們的困難定義集的可重現性遠低於預期。這些困難標籤是從有限數量的樣本反應中估算得出的。跨種子結合它們可能進一步改變所選擇的提示,而不僅僅是減少測量噪音。我們開發了一個基於抽樣的框架來量化這種不穩定性,並確定為了使困難分配可靠地重現需要多少評估。我們還重新檢視了用於解釋不可學習的梯度相似性證據,並顯示觀察到的分離部分源於困難提示提供的正確回饋較少,從中無法估算其梯度。匹配這一樣本數量削弱了梯度差異,但並未消除它。總體而言,緩慢學習現象在我們的重新分析中仍然存在,而用來定義它的提示和用來解釋它的證據都需要更仔細的測量。

What Can Component-Replacement Evidence Establish? A Critical Scoping Review of Local Decisions in LLM Agents

2609.39989v1 by Shuyang Zhang, Jianshuo Chang

Background. A component replacement in a language-model agent changes an execution trajectory, potentially altering later observations, resource use, and recovery opportunities. Different evidence is needed to assess its task-level benefit and the contribution of local decision quality. Methods. This critical scoping review maps 348 studies and examines 90 comparison records: 88 from 40 included studies and two from supplementary studies. Eight purposively selected cases structure the synthesis around the replaced decision, executed conditions, measurement comparability, controls, and remaining explanations. Results. Of 222 studies reporting local decision metrics, 142 also report measured task endpoints and 49 report proxies. These counts identify studies that report both types of measurement, without establishing that the measurements come from matched comparisons. Outcome Monitors reports a package-level completion gain whose attribution to detector quality remains limited; First-chunk selection reports a local improvement assessed against an offline proxy endpoint; Evidence-Carrying Termination reports fewer premature unsupported terminations and completion non-inferiority, without establishing completion superiority. Cross-case analysis identifies three candidate mechanisms involving recovery and disruption, intervention timing, and downstream use. Attribution and deployment depend on the comparison controls, label definitions, and information available to the controller. Conclusions. The review distinguishes the task-level benefit of a component replacement from the contribution of local decision quality and derives eight claim-specific reporting items. Neither online execution nor simultaneous gains in local and task metrics alone establish that better local decisions explain the task-level gain.

摘要:背景。語言模型代理中的組件替換改變了執行軌跡,可能改變後續觀察、資源使用和恢復機會。需要不同的證據來評估其任務層面的好處以及當地決策質量的貢獻。方法。這項關鍵範疇評估回顧映射了348項研究並檢查了90個比較記錄:88個來自40項納入的研究,兩個來自補充研究。八個有目的選擇的案例圍繞被替換的決策、執行條件、測量可比性、控制和剩餘解釋結構化合成。結果。在222項報告當地決策指標的研究中,142項還報告了測量的任務端點,49項報告了代理指標。這些數量識別了報告兩種類型測量的研究,但並未確立這些測量來自匹配比較。結果監控報告了一個包級別的完成增益,其歸因於檢測器質量的限制;第一塊選擇報告了一個相對於離線代理端點評估的當地改進;證據攜帶終止報告了較少的過早無支持終止和完成非劣性,但未確立完成優越性。跨案例分析識別了三個候選機制,涉及恢復和中斷、干預時機和下游使用。歸因和部署取決於比較控制、標籤定義和控制者可用的信息。結論。該評估區分了組件替換的任務層面好處與當地決策質量的貢獻,並推導出八個特定於主張的報告項目。僅僅依賴在線執行或當地和任務指標的同時增益並不能確立更好的當地決策解釋了任務層面的增益。

How Does Local Landscape Geometry Evolve in Language Model Pre-Training?

2609.39767v1 by Zhanpeng Zhou, Yuhan Sun, Bingrui Li, Jinbo Wang, Huaijin Wu, Lei Wu, Junchi Yan

The scale and expense of pre-training language models make efficient hyperparameter tuning essential, yet a principled guidance is still missing. In this work, we analyze language model pre-training dynamics from a local landscape geometry perspective. Our study reveals two distinct phases. In Phase I, sharpness of the local landscape is initially high, leading to instability and loss plateaus under large learning rates (LRs). The landscape shifts from sharp to flatter regions early in training. This dynamic explains the necessity of LR warmup and further suggests that larger peak LRs require proportionally longer warmup periods. In Phase II, the local landscape is governed by the gradient noise scale. Our theory identifies a depth flatness trade-off: high noise from smaller batches widens the loss basin, whereas reduced noise from larger batches deepens it. This theory motivates a dynamic batch-size (BS) scheduler that begins with a small BS and increases it late in training. Together, we provide a unified view of loss landscape evolution, which translates into actionable tuning strategies for large-scale pre-training.

摘要:預訓練語言模型的規模和成本使得高效的超參數調整變得至關重要,但仍然缺乏原則性的指導。在本研究中,我們從局部景觀幾何的角度分析語言模型的預訓練動態。我們的研究揭示了兩個不同的階段。在第一階段,局部景觀的尖銳度最初很高,導致在較大學習率(LRs)下的不穩定性和損失平穩期。隨著訓練的進行,景觀從尖銳轉向較平坦的區域。這一動態解釋了LR預熱的必要性,並進一步表明較大的峰值LR需要相應更長的預熱期。在第二階段,局部景觀受梯度噪聲尺度的影響。我們的理論確定了一個深度平坦度的權衡:來自較小批次的高噪聲擴大了損失盆地,而來自較大批次的低噪聲則使其變深。這一理論促使我們提出了一個動態批次大小(BS)調度器,該調度器在訓練初期從小BS開始,並在訓練後期增加它。總體而言,我們提供了一個損失景觀演變的統一視角,這轉化為大規模預訓練的可操作調整策略。

Trust Is Not a Score: Runtime Assurance Contracts for High-Risk AI Agents

2609.39717v1 by Serhii Zabolotnii

Benchmarks, audits, and agent protocols describe performance, permissions, and repair, but not how observed evidence should change an agent's authority during a consequential task. We call this the assurance-transition gap. We propose a Runtime Assurance Contract (RAC), a policy-level formal schema binding autonomy boundaries, component eligibility, evidence state, transition policy, human-review capacity, and non-compensatory gates. Under RAC, soft metrics may inform routing, whereas a failed or unknown mandatory gate forces retry, switch, escalation, deferral, or stop; aggregate performance cannot authorize action. We define the contract, an evidence record, a permission rule, and five invariants, and illustrate them in clinical, industrial, and judicial failure probes. We then report a deterministic failure-injection study in agentic coding: 280 constructed cases evaluated by a gate conjunction, a score-only rule, and a restricted protocol baseline. At the published example weights and threshold, the score rule admits 80 of 100 block-required injections and all 40 review-required injections. Tuned in hindsight, it matches the conjunction on this corpus. For positive weights, a positive threshold, binary risk signals, zero-signal controls, and an injected case firing each signal alone, we show that exact agreement holds if and only if the threshold does not exceed the smallest weight. A separate set of 18 hand-authored traces checks version-pinned evidence and review transitions against simpler policy variants. In a further prospective synthetic holdout of 24 episodes, two blinded LLM judges assign identical labels to all 72 action attempts; RAC and a separately implemented full stateful baseline both match these labels. These studies test mechanisms on synthetic cases; they establish neither deployed safety nor cross-domain effectiveness.

摘要:基準、審計和代理協議描述了性能、權限和修復,但並未說明在關鍵任務中,觀察到的證據應如何改變代理的權限。我們稱之為保證過渡差距。我們提出了一個運行時保證合約(RAC),這是一個政策層級的正式架構,約束自主邊界、組件資格、證據狀態、過渡政策、人類審查能力和非補償性閘門。在RAC下,軟指標可以用來指導路由,而失敗或未知的強制閘門則強迫重試、切換、升級、延遲或停止;總體性能無法授權行動。我們定義了合約、一個證據記錄、一條許可規則和五個不變量,並在臨床、工業和司法失敗探測中進行了說明。然後,我們報告了一項在代理編碼中的確定性失敗注入研究:280個構建的案例通過閘門聯合、一個僅計分的規則和一個受限的協議基線進行評估。在已發表的示例權重和閾值下,計分規則允許100個區塊所需注入中的80個和所有40個審查所需的注入。事後調整後,它在這個語料庫上與聯合匹配。對於正權重、正閾值、二元風險信號、零信號控制和每個信號單獨觸發的注入案例,我們顯示出精確一致性僅在閾值不超過最小權重時成立。一組18個手工編寫的痕跡檢查版本固定的證據和審查過渡,與更簡單的政策變體進行比較。在進一步的前瞻性合成保留中,24個集數中,兩位盲法LLM評審對所有72次行動嘗試分配了相同的標籤;RAC和一個單獨實施的完整狀態基線都與這些標籤相匹配。這些研究在合成案例上測試機制;它們既未建立已部署的安全性,也未建立跨領域的有效性。

ChronoGraph: Functional 4D Scene Graphs with Vision-Language Models for Interaction Understanding and Grounded Planning

2609.39665v1 by Chenyangguang Zhang, Malgorzata Gwiazda, Guanlong Jiao, Yuanchen Ju, Federico Tombari, Koushil Sreenath, Marc Pollefeys, Sunghwan Hong

Embodied agents must determine where to act, anticipate the resulting scene changes, and interpret observed outcomes to guide subsequent actions. This requires connecting 4D interaction understanding, which explains how past actions changed the scene, with spatially grounded planning, which determines how and where to act toward a goal and anticipates the resulting scene changes. We introduce ChronoGraph, a functional 4D scene graph that links actions on affordance parts to semantic and geometric state changes. By representing observed and anticipated transitions in the same form, it provides a shared basis for understanding and planning. We construct ChronoGraphBench through an automatic data engine that converts human-interaction videos and simulated robot trajectories into graph-annotated questions for training and evaluating Vision-Language Models (VLMs) on both tasks. Using these annotations, we train ChronoGraphVLM by adapting pretrained VLMs in two stages. Graph-as-Chain-of-Thought supervised fine-tuning teaches the models to reconstruct observed transitions and predict future ones as graph traces before answering. Subsequent joint 4D graph reinforcement learning directly rewards graph properties and answer correctness. Experiments across model scales show improvements over the corresponding pretrained baselines and zero-shot transfer to VLM4D. Real-world demonstrations further show that graph-based planning and affordance grounding support mobile manipulation through existing robot skills without additional fine-tuning.

摘要:具身代理必須確定行動的地點,預測隨之而來的場景變化,並解釋觀察到的結果以指導後續行動。這需要將 4D 互動理解(解釋過去的行動如何改變場景)與空間基礎規劃(確定如何以及在哪裡朝著目標行動並預測隨之而來的場景變化)連接起來。我們介紹 ChronoGraph,一個功能性 4D 場景圖,將對可供性部分的行動與語義和幾何狀態變化聯繫起來。通過以相同的形式表示觀察到的和預期的轉變,它為理解和規劃提供了一個共同的基礎。我們通過一個自動數據引擎構建 ChronoGraphBench,該引擎將人類互動視頻和模擬機器人軌跡轉換為帶有圖形標註的問題,以便在兩個任務上訓練和評估視覺-語言模型(VLMs)。利用這些標註,我們通過在兩個階段適應預訓練的 VLMs 來訓練 ChronoGraphVLM。作為思維鏈的圖形監督微調教導模型重建觀察到的轉變並在回答之前預測未來的轉變作為圖形痕跡。隨後的聯合 4D 圖形強化學習直接獎勵圖形屬性和答案的正確性。跨模型規模的實驗顯示出相對於相應的預訓練基線的改進,以及對 VLM4D 的零樣本轉移。現實世界的演示進一步表明,基於圖形的規劃和可供性基礎支持通過現有的機器人技能進行移動操作,而無需額外的微調。

Zero-Compute Cross-Lingual Transferability Estimation Using Typological Feature Proxies

2609.39640v1 by Dalton Raphael Harmsen, Swier Garst, Thomas van Osch, Zarè Palanciyan, Joaquin Vanschoren

Cross-lingual transfer describes how knowledge in a source language benefits a target language. Measuring it quantitatively requires broad multilingual pre-training, as prior work has done with cross-lingual transfer matrices. We ask whether transfer is predictable from freely available typological features, and whether the prominence of high-resource source languages reflects typology or data quality and quantity. We show that typological databases contain cheap and dense signals about cross-lingual transfer. Our typology-only random forest on a 24-language prior-work transfer matrix scores leave-one-language-out $ρ{=}0.705$ and $R^2{=}0.49$, beating a non-typological control at $ρ{=}0.62$, which verifies the ability of typology-only predictions to reconstruct costly measured cross-lingual transfer. The signal survives leave-one-script-out and leave-one-family-out protocols, so script and family confounding do not explain the effect. By decomposing the transfer into a typology term and a resource-and-script bias term, we find the best-source ranking sensitive to this bias. In contrast, typology is not affected by this bias, which makes it a zero-compute screening tool that replaces hundreds of training runs with a model fit. Our code is available \href{https://github.com/dharmsen/typo-x-ling-transfer}{here}.

摘要:跨語言轉移描述了來源語言的知識如何惠及目標語言。定量測量需要廣泛的多語言預訓練,正如先前的工作所做的跨語言轉移矩陣。我們詢問是否可以從自由可用的類型特徵預測轉移,以及高資源來源語言的顯著性是否反映了類型學或數據質量和數量。我們展示了類型學數據庫包含有關跨語言轉移的廉價且密集的信號。我們的僅基於類型學的隨機森林在24語言的先前工作轉移矩陣上的得分為留一語言外 $ρ{=}0.705$ 和 $R^2{=}0.49$,超過了 $ρ{=}0.62$ 的非類型學控制,這證實了僅基於類型學的預測能夠重建昂貴的測量跨語言轉移的能力。該信號在留一腳本外和留一語系外的協議中仍然存在,因此腳本和語系的混淆並不能解釋這一效果。通過將轉移分解為類型學項和資源與腳本偏差項,我們發現最佳來源排名對此偏差敏感。相比之下,類型學不受此偏差影響,這使其成為一種零計算篩選工具,能夠用模型擬合取代數百次訓練運行。我們的代碼可在 \href{https://github.com/dharmsen/typo-x-ling-transfer}{這裡} 獲得。

Disentangling Self-Distillation: Measuring and Modeling Acquisition and Retention

2609.39494v1 by Luis Zuin, Alexis Huet, Dario Rossi, Zied Ben Houidi

Self-distillation with privileged context adapts a language model from demonstrations by letting the model, once conditioned on a reference response, teach its context-free copy token by token. Our taxonomy reveals existing methods differ along three entangled axes: (i) the rollout source (student or teacher), (ii) the teacher coupling (frozen, or an exponential moving average of the student at some coupling rate) and (iii) the KL direction (reverse or forward), yet these axes are usually studied in fixed combinations and have led to conflicting conclusions. We formalize a unifying framework to encompass all self-distillation methods vs classic supervised fine-tuning: we train every combination of the three axes, on Qwen2.5-7B and Ministral-3-3B across ordinary and contradictory tasks, totaling 1,200 adaptation runs, to systematically investigate the impact of the above axes. We propose a controlled model of the same objective to explain the resulting acquisition-retention trade-offs. We find that (i) the rollout source matters mostly where the task contradicts the pretrained behavior: there teacher rollouts raise acquisition well above what student rollouts achieve, with almost no change in retention; (ii) the teacher coupling changes acquisition most, on every task: acquisition rises with the coupling rate, then falls past a task-specific rate; (iii) switching the KL direction costs retention in one model but not the other so which axis to tune first depends on the model. The controlled model reproduces the three trends.

摘要:自我蒸餾與特權上下文透過讓模型在參考回應的條件下,逐步教導其無上下文副本,來適應語言模型。我們的分類法揭示現有方法在三個交織的軸向上有所不同:(i)展開來源(學生或教師),(ii)教師耦合(凍結,或以某種耦合速率的學生指數移動平均)以及(iii)KL方向(反向或正向),然而這些軸通常以固定的組合進行研究,並導致相互矛盾的結論。我們正式化了一個統一框架,以涵蓋所有自我蒸餾方法與經典的監督微調:我們在Qwen2.5-7B和Ministral-3-3B上,針對普通和矛盾任務訓練三個軸的每一種組合,總計1,200次適應運行,以系統性地調查上述軸的影響。我們提出了一個相同目標的受控模型,以解釋所得到的獲取-保留權衡。我們發現:(i)展開來源主要在任務與預訓練行為矛盾時才重要:在這種情況下,教師的展開使獲取遠高於學生的展開,幾乎沒有保留的變化;(ii)教師耦合對每個任務的獲取影響最大:獲取隨著耦合速率上升,然後在特定任務的速率後下降;(iii)切換KL方向在一個模型中會影響保留,但在另一個模型中則不會,因此首先調整哪個軸取決於模型。受控模型重現了這三個趨勢。

From Speech to Editable Concepts: Probing Emotion Recognition with Concept Bottleneck Models

2609.39453v1 by Hezhao Zhang, Thomas Hain

Speech emotion recognition (SER) is the task of assigning emotion labels to utterances. Early systems relied on acoustic features, whereas recent approaches combine multiple modalities, most commonly speech and text. Still, performance remains poor on many datasets. Large language models (LLMs) have therefore attracted interest for SER, as they can process diverse inputs jointly with instructions. However, direct audio input raises questions of explainability. To address similar questions in image classification, concept bottleneck models were introduced. This work adapts concept bottlenecks to SER to examine how individual predictions depend on transcripts, acoustic descriptions and speaker attributes. Experiments test three LLMs on CREMA-D, IEMOCAP and MELD, with concepts extracted by separate tools. On scripted corpora, LLMs are strongly biased towards the transcript in the zero-shot setting, which lowers Macro-F1 from 27.8 to 5.8 on CREMA-D. Fine-tuning removes this bias, and the transcript raises Macro-F1 from 41.8 to 45.1. Removing speech rate changes 48% of Neutral predictions to Disgust on CREMA-D; removing intensity level on MELD changes predictions despite little change in Macro-F1. These findings show that aggregate performance changes alone do not capture the effects of concept removal on individual predictions.

摘要:語音情感識別(SER)是將情感標籤分配給話語的任務。早期的系統依賴於聲學特徵,而最近的方法則結合了多種模態,最常見的是語音和文本。儘管如此,許多數據集上的性能仍然較差。因此,大型語言模型(LLMs)引起了對SER的興趣,因為它們可以與指令共同處理多樣的輸入。然而,直接的音頻輸入引發了可解釋性的問題。為了解決圖像分類中的類似問題,引入了概念瓶頸模型。本研究將概念瓶頸應用於SER,以檢查個別預測如何依賴於文字稿、聲學描述和說話者屬性。實驗測試了三個LLM在CREMA-D、IEMOCAP和MELD上的表現,概念由不同的工具提取。在腳本語料庫中,LLM在零樣本設置中對文字稿有強烈的偏見,這使得CREMA-D上的Macro-F1從27.8降低到5.8。微調消除了這種偏見,文字稿使得Macro-F1從41.8提高到45.1。移除語音速率使CREMA-D上48%的中性預測變為厭惡;在MELD上移除強度水平則改變了預測,儘管Macro-F1幾乎沒有變化。這些發現表明,僅僅改變總體性能並不能捕捉到概念移除對個別預測的影響。

Towards Trustworthy AI for Glioma Diagnosis: A Task-Aware Evaluation of Uncertainty Quantification

2609.39429v1 by Gonzalo Esteban Mosquera Rojas, Sebastian R. van der Voort, Carolin M. Pirkl, Sandeep Kaushik, Marion Smits, Stefan Klein

Uncertainty Quantification (UQ) is a key requirement for trustworthy AI in high-stakes medical image analysis. In this work, we evaluate UQ in a multi-task Deep Learning framework for MRI-based glioma diagnosis that performs tumor segmentation and predicts IDH mutation status, 1p/19q co-deletion status, and tumor grade. Monte Carlo Dropout (MCD) is used for a detailed task-aware analysis of predictive, aleatoric, and epistemic uncertainty. We assess MC sample convergence, calibration, error detection, selective prediction, associations with segmentation performance, and the effect of voxel-wise uncertainty aggregation on case-level reliability. We also compare MCD with Deep Ensembles (DE) and Monte Carlo Deep Ensembles (MCDE), examine interactions between segmentation quality and classification, and evaluate a composite trust score integrating segmentation and classification uncertainty. Across tasks, uncertainty estimates supported meaningful error detection, while calibration depended on the dropout rate, with moderate rates yielding the most reliable probabilities. Uncertainty decomposition provided task-dependent interpretability but did not consistently improve error detection over predictive uncertainty alone. DE and MCDE showed comparable operational utility, with no method consistently dominating across tasks and metrics. The composite trust score did not consistently outperform classification uncertainty for selective prediction. Overall, our results provide a task-aware evaluation strategy and practical guidance for the development of trustworthy AI for glioma diagnosis.

摘要:不確定性量化(UQ)是高風險醫療影像分析中可信賴人工智慧的關鍵要求。在這項工作中,我們評估了一個多任務深度學習框架中的UQ,該框架用於基於MRI的膠質瘤診斷,執行腫瘤分割並預測IDH突變狀態、1p/19q共同缺失狀態和腫瘤等級。使用蒙特卡羅隨機失活(MCD)對預測性、隨機性和認知性不確定性進行詳細的任務感知分析。我們評估了蒙特卡羅樣本收斂性、校準、錯誤檢測、選擇性預測、與分割性能的關聯,以及體素級不確定性聚合對案例級可靠性的影響。我們還將MCD與深度集成(DE)和蒙特卡羅深度集成(MCDE)進行比較,檢查分割質量與分類之間的相互作用,並評估整合分割和分類不確定性的綜合信任分數。在各項任務中,不確定性估計支持有意義的錯誤檢測,而校準則依賴於隨機失活率,中等率產生最可靠的概率。不確定性分解提供了任務依賴的可解釋性,但並未始終改善僅依賴預測不確定性的錯誤檢測。DE和MCDE顯示出可比的操作效用,沒有一種方法在各任務和指標中始終佔優。綜合信任分數在選擇性預測中並未始終優於分類不確定性。總體而言,我們的結果提供了一種任務感知的評估策略和實用指導,旨在為膠質瘤診斷的可信賴人工智慧發展提供支持。

The Golden Path Hypothesis: Reusable Schedules in Diffusion Caching

2609.39343v1 by Dong Wang, Wenwu Tang, Francesco Corti, Yun Cheng, Lothar Thiele, Olga Saukh

Diffusion caching accelerates generation by replacing transformer computation with cached or predicted features at selected denoising steps. We introduce the Golden Path Hypothesis (GPH): under fixed inference conditions, prompt-independent cache schedules can achieve final-output quality comparable to the best prompt-specific schedules across prompts. We investigate the GPH across ten caching methods, four image and video models, and three cache ratios. Prompt-adaptive methods repeatedly select a small number of schedules, and reusing their most frequent schedules on new prompts closely matches the quality of prompt-specific choices. Exhaustive evaluation of 1.4 million schedules on four examples further identifies prompt-independent schedules that remain competitive on unseen prompts. To explain this transfer, we analyze denoising trajectories and the accumulation of caching errors. Latent-state trajectories exhibit similar structures across datasets and seeds, while an exact error decomposition shows that accumulated effects of earlier errors predict final latent-state error better than local approximation errors. This motivates searching for end-to-end schedules using final-output quality. With only a small set of examples, the resulting golden paths transfer across prompts and datasets, and can be tuned to the desired quality objective, including reconstruction fidelity or perceptual similarity.

摘要:擴散快取通過在選定的去噪步驟中用快取或預測的特徵替代Transformer計算來加速生成。我們提出了黃金路徑假設(GPH):在固定的推理條件下,與最佳的提示特定計劃相比,與提示無關的快取計劃可以達到相似的最終輸出質量。我們在十種快取方法、四種圖像和視頻模型以及三種快取比率上研究了GPH。提示自適應方法重複選擇少量計劃,並在新提示上重用其最頻繁的計劃,這與提示特定選擇的質量非常接近。對140萬個計劃在四個示例上的全面評估進一步識別出在未見提示上仍具競爭力的與提示無關的計劃。為了解釋這一轉移,我們分析了去噪軌跡和快取錯誤的累積。潛在狀態軌跡在數據集和種子之間顯示出相似的結構,而精確的錯誤分解顯示,早期錯誤的累積效應比局部近似錯誤更好地預測最終潛在狀態錯誤。這促使我們尋找使用最終輸出質量的端到端計劃。僅用一小組示例,結果的黃金路徑在提示和數據集之間轉移,並可以調整以達到所需的質量目標,包括重建保真度或感知相似性。

A Time-Aware Bag-of-Receptive-Fields for Interpretable Irregular Time Series Classification

2609.39268v1 by Francesco Spinnato

Irregular time series, characterized by non-uniform sampling intervals, missing observations, and variable lengths, are ubiquitous in healthcare, mobility, and environmental monitoring, yet effective and interpretable classifiers for this setting are limited. Existing approaches often rely on imputation, which can obscure the temporal structure of the data, or require complex neural architectures that are opaque and difficult to explain. In this work, we extend the Bag-Of-Receptive-Fields (BORF), a fast, deterministic, and interpretable transform for time series, to the irregular setting. Our key contribution is a time-weighted normalization scheme in which each observation is weighted proportionally to its associated time delta, making pattern extraction sensitive to the actual temporal distribution of samples rather than only their index position. This requires deriving an efficient sliding-window recurrence for the time-weighted standard deviation, preserving the linear time complexity of BORF. We benchmark the resulting method against state-of-the-art irregular time series classifiers on datasets from the PYRREGULAR repository, demonstrating competitive classification performance with the added benefit of human-interpretable explanations.

摘要:不規則時間序列的特徵是取樣間隔不均、缺失觀測值和變化的長度,這在醫療、移動性和環境監測中隨處可見,但在這種情境下有效且可解釋的分類器卻有限。現有的方法通常依賴於插補,這可能會掩蓋數據的時間結構,或需要複雜的神經架構,這些架構不透明且難以解釋。在這項工作中,我們將快速、確定性且可解釋的時間序列變換——感受野包(Bag-Of-Receptive-Fields, BORF)擴展到不規則的情境。我們的主要貢獻是一種時間加權標準化方案,其中每個觀測值的權重與其相關的時間增量成比例,這使得模式提取對樣本的實際時間分佈敏感,而不僅僅是它們的索引位置。這需要導出一種高效的滑動窗口重複計算時間加權標準差,從而保持BORF的線性時間複雜度。我們將所得到的方法與來自PYRREGULAR數據庫的最先進不規則時間序列分類器進行基準測試,展示了具有競爭力的分類性能,並附帶可供人類解釋的解釋。

When Harnesses Lose the Signal: Causal Evaluation of Recovery in LLM Agents

2610.00372v1 by Shuyao Xiao, Shengling Wang, Xuan Chen, Ke Chao, Ming Cui, Feifei Qian, Chaoyang Mei, Fanlin Meng, Ziming Yu, Junxi Yin

Large language model agents rely on external harnesses to pass information between the model and its environment and to recover from execution errors. Yet recovery is usually judged only by average task success. This hides an important tension. The same operation can rescue a failing trajectory or disrupt one that would otherwise succeed. We frame recovery as a causal decision problem. Starting from the same execution state, we compare what happens with and without recovery, separate rescue from harm, and study how the value of recovery changes over time. We then introduce the Causal Intervention Router (CIR), a lightweight policy that uses information available before recovery to decide when intervention is worthwhile. On long-horizon ALFWorld tasks with Qwen3-14B, CIR raises success from 70.33% to 73.33%, a gain of 3.00 percentage points. It leaves all evaluated trajectories with correct observations untouched. Additional controls show that the benefit of recovery cannot be explained solely by the new observation returned by the environment. These results provide a practical way to evaluate recovery and apply it selectively.

摘要:大型語言模型代理依賴外部裝置在模型與其環境之間傳遞信息並從執行錯誤中恢復。然而,恢復通常僅根據平均任務成功率來評估。這隱藏了一個重要的緊張關係。同一操作可以挽救失敗的軌跡,或破壞本來會成功的軌跡。我們將恢復框架設置為一個因果決策問題。從相同的執行狀態開始,我們比較有無恢復的情況下發生的事情,將救援與傷害分開,並研究恢復的價值如何隨時間變化。然後我們介紹了因果干預路由器(CIR),這是一種輕量級策略,利用恢復前可用的信息來決定何時進行干預是值得的。在使用Qwen3-14B的長期ALFWorld任務中,CIR將成功率從70.33%提高到73.33%,增幅為3.00個百分點。它保持所有評估的軌跡的正確觀察不變。額外的控制顯示,恢復的好處不能僅僅用環境返回的新觀察來解釋。這些結果提供了一種實際的方法來評估恢復並選擇性地應用它。

Loop-Free Inverse Reinforcement Learning via Sequential Value Recovery with Q-Score Matching

2609.38955v1 by Yang chen, Yitan Zhang, Michael Witbrock, Shuyue Hu

Inverse Reinforcement Learning (IRL) aims to recover a reward function that explains expert demonstrations. Existing IRL methods typically rely on a bi-level optimization procedure that alternates between reward learning and policy optimization, leading to substantial computational burden and training instability. In this work, we introduce a different route that eliminates policy optimization entirely by leveraging diffusion policies. Our key insight is that a diffusion policy encodes the action-gradient structure of the optimal soft Q function, enabling reward learning to be cast as a sequence of value recovery problems, thereby allowing us to bypass reward-policy loops inherent in prior IRL methods. Specifically, our method proceeds in three stages: (I) recovering the optimal soft Q function via action-gradient matching and estimating the corresponding soft value function (LogSumExp of Q values) in a way inspired by Gumbel regression; (II) calibrating these soft values by inferring a state-dependent offset; (III) extracting the reward by enforcing Bellman consistency. This leads to Loop-Free Inverse Reinforcement Learning (LFIRL), a fully offline algorithm that operates in a simple, loop-free, and sequential manner. LFIRL is simple to implement and significantly improves training efficiency while maintaining strong reward recovery performance. Empirically, across Maze, Franka Kitchen, Adroit Hand Pen, and Push-T benchmarks, LFIRL achieves 2-3x speedup over the fastest baselines, while matching or surpassing state-of-the-art methods in reward recovery quality.

摘要:逆向強化學習(IRL)旨在恢復解釋專家示範的獎勵函數。現有的IRL方法通常依賴於一種雙層優化程序,在獎勵學習和策略優化之間交替進行,這導致了相當大的計算負擔和訓練不穩定性。在這項工作中,我們提出了一種不同的路徑,通過利用擴散策略完全消除了策略優化。我們的關鍵見解是,擴散策略編碼了最佳軟Q函數的行動梯度結構,使得獎勵學習可以被視為一系列價值恢復問題,從而使我們能夠繞過先前IRL方法中固有的獎勵-策略循環。具體而言,我們的方法分為三個階段:(I)通過行動梯度匹配恢復最佳軟Q函數,並以受到Gumbel回歸啟發的方式估計相應的軟值函數(Q值的LogSumExp);(II)通過推斷狀態依賴的偏移來校準這些軟值;(III)通過強制執行Bellman一致性來提取獎勵。這導致了無循環逆向強化學習(LFIRL),這是一種完全離線的算法,以簡單、無循環和順序的方式運行。LFIRL實現簡單,顯著提高了訓練效率,同時保持強大的獎勵恢復性能。在Maze、Franka Kitchen、Adroit Hand Pen和Push-T基準測試中,LFIRL在速度上比最快的基準提高了2-3倍,同時在獎勵恢復質量上達到或超越了最先進的方法。

Reasoning Externalization for Faithful Large Language Model Narratives of Stock Return Predictions

2609.38869v1 by Sujung Kim, Seung Hwan Cho, Sangjin Park, Young-Min Kim

In finance, interpreting machine learning predictions is essential, yet the numerical outputs of explainable AI can be difficult for non-experts to understand. While large language models (LLMs) can translate these outputs into natural language, they may produce errors when inferring numerical changes and feature relations. We propose an LLM narrative framework for cross-sectional stock return prediction that combines temporal Shapley additive explanations (SHAP) evidence with historical regime analogs. Temporal evidence tracks changes in the normalized global SHAP importance of an XGBoost model over six months. Historical analogs are past periods with similar changes in SHAP importance, their model performance and subsequent market returns are provided as comparative context. Using this framework, we conduct a controlled study of progressive reasoning externalization, sequentially providing raw SHAP sequences, deterministic temporal descriptors, and feature relations. Each generated claim is verified against provenance-linked evidence. Across Qwen3, externalizing numerical and relational reasoning improved evidence faithfulness as well as temporal and relational accuracy. Evidence faithfulness increased from 0.696 to 0.996 for Qwen3-32B-Instruct. While historical analogs did not improve structured automatic faithfulness, they received higher human-rated usefulness scores. These results suggest that externalizing verifiable reasoning enhances narrative faithfulness and that historical context adds interpretive value.

摘要:在金融領域,解釋機器學習預測是至關重要的,但可解釋人工智慧的數值輸出對於非專家來說可能難以理解。雖然大型語言模型(LLMs)可以將這些輸出轉換為自然語言,但在推斷數值變化和特徵關係時,它們可能會產生錯誤。我們提出了一個LLM敘事框架,用於橫斷面股票回報預測,該框架結合了時間性Shapley加法解釋(SHAP)證據和歷史制度類比。時間性證據跟蹤XGBoost模型在六個月內的標準化全球SHAP重要性的變化。歷史類比是過去在SHAP重要性上有類似變化的時期,提供其模型表現和隨後市場回報作為比較背景。利用這個框架,我們進行了一項受控研究,對進步推理外化進行了逐步的探討,依次提供原始SHAP序列、確定性時間描述符和特徵關係。每個生成的主張都與來源鏈接的證據進行驗證。在Qwen3上,外化數值和關係推理提高了證據的真實性以及時間和關係的準確性。對於Qwen3-32B-Instruct,證據的真實性從0.696提高到0.996。雖然歷史類比並未改善結構化自動真實性,但它們獲得了更高的人類評價的有用性分數。這些結果表明,外化可驗證的推理增強了敘事的真實性,而歷史背景則增添了解釋價值。

SkillSeek: Revisiting Agent Skill Retrieval at Marketplace Scale

2609.38822v1 by Guanqun Yang, Wenlong Zhang, Tian Shi, Ping Wang

Anthropic's Agent Skills package reusable procedural know-how for an LLM agent into SKILL.md directories, and open-source aggregations have grown past 230,000 skills, making selection rather than authoring the bottleneck. The standing answer in the literature outsources selection to the agent itself: an LLM-mediated retrieval loop that rewrites queries and refines candidates inside the agent's decision loop, paying LLM tokens on every task. We present SkillSeek, an open-source two-stage skill retriever built from the standard IR recipe (a BGE-base bi-encoder feeding a small cross-encoder, exposed over MCP). Across a $4 \times 11$ grid of pool, backbone, and method on the 89-task SkillsBench benchmark, SkillSeek reaches observed parity with the LLM-mediated loop of Liu et al. at essentially no extra cost: plain bm25 alone records a pass rate at or above their refined loop on three of four settings, and a small cross-encoder covers the remaining difference on the fourth. A first-stage recall ceiling explains the pattern, and total per-trial spend drops from USD 51.30 to USD 27.54 (within fifty cents of the no-skill baseline). Under the SkillsBench tasks and OpenHands harness we tested, this positions the standard IR recipe as a strong default for agent-skill retrieval, with LLM-mediated alternatives a natural fit for cases where deterministic methods fall short.

摘要:Anthropic 的 Agent Skills 套件將可重用的程序知識整合到 LLM 代理的 SKILL.md 目錄中,開源聚合已經超過 230,000 種技能,使得選擇而非創作成為瓶頸。文獻中的現有解答將選擇外包給代理本身:一個 LLM 媒介的檢索循環,重寫查詢並在代理的決策循環內精煉候選者,對每個任務支付 LLM 代幣。我們提出了 SkillSeek,一個基於標準 IR 食譜的開源兩階段技能檢索器(使用 BGE 基礎的雙編碼器供應小型交叉編碼器,通過 MCP 暴露)。在 89 任務的 SkillsBench 基準上,SkillSeek 在 $4 \times 11$ 的池、骨幹和方法網格中達到了與 Liu 等人的 LLM 媒介循環的觀察平衡,幾乎沒有額外成本:單純的 bm25 在四個設置中的三個上記錄的通過率達到或超過他們的精煉循環,而小型交叉編碼器則覆蓋了第四個設置的剩餘差異。一階段召回上限解釋了這一模式,並且每次試驗的總支出從 51.30 美元降至 27.54 美元(在無技能基準線的五十美分之內)。在我們測試的 SkillsBench 任務和 OpenHands 環境下,這使得標準 IR 食譜成為代理技能檢索的強大默認選擇,而 LLM 媒介的替代方案則自然適合於確定性方法無法滿足的情況。

When Reasoning Goes Astray: Attention Dynamics of Uncontrolled Reasoning

2609.38817v1 by Yuanhe Zhang, Ziwei Wang, Jie Ren, Haoran Gao, Zhenhong Zhou, Fanyu Meng, Cong Wu, Li Sun, Sen Su

Large reasoning models (LRMs) improve performance on complex tasks through extended reasoning, yet the same process can degenerate into redundant verification and persistent generation loops. Such uncontrolled reasoning increases inference cost and creates risks of resource exhaustion and service degradation. However, existing mitigations largely truncate long outputs or react to surface repetition, and thus fail to distinguish normal thinking from uncontrolled reasoning or explain how benign reasoning degenerates into harmful behavior. In this paper, we operationalize LRM generation as four states and further introduce Reasoning-state Analysis via Dynamic Attention Responses (RADAR), which identifies the current reasoning state in real time and characterizes how effective reflection can develop into uncontrolled generation. Guided by RADAR's analysis, we further realign abnormal attention distributions toward patterns observed in normal requests and examine how this correction affects excessive reflection and persistent looping. Temporal analyses show that uncontrolled reasoning is characterized by attention distributions that deviate from normal generation, with abnormal trends becoming detectable before repetition begins. Correcting these deviations through Attention Realignment consistently reduces looping while largely preserving benign performance. Together, RADAR provide a mechanistic account of how reasoning becomes uncontrolled, offering actionable guidance for identifying critical failure stages and designing targeted runtime interventions.

摘要:大型推理模型(LRMs)透過擴展推理來提高在複雜任務上的表現,然而相同的過程可能會退化為冗餘的驗證和持續的生成循環。這種不受控制的推理增加了推斷成本,並創造了資源耗盡和服務降級的風險。然而,現有的緩解措施主要是截斷長輸出或對表面重複作出反應,因此未能區分正常思考與不受控制的推理,或解釋良性推理如何退化為有害行為。在本文中,我們將LRM生成操作化為四個狀態,並進一步引入動態注意力反應下的推理狀態分析(RADAR),該方法實時識別當前的推理狀態,並描述有效反思如何發展成不受控制的生成。在RADAR的分析指導下,我們進一步重新調整異常的注意力分佈,朝向在正常請求中觀察到的模式,並檢查這一修正如何影響過度反思和持續循環。時間分析顯示,不受控制的推理以偏離正常生成的注意力分佈為特徵,異常趨勢在重複開始之前就變得可檢測。通過注意力重新調整來修正這些偏差,持續減少循環,同時在很大程度上保持良性表現。總的來說,RADAR提供了一個機制性解釋,說明推理如何變得不受控制,並提供可行的指導,以識別關鍵失敗階段並設計針對性的運行時干預。

Understanding Off- vs On-Policy Distillation: A Tale of Distinct Training Objectives

2609.38666v1 by Qiwei Di, Xuheng Li, Kaixuan Ji, Chenggong Zhang, Heyang Zhao, Quanquan Gu

On-policy distillation (OPD) learns from teacher feedback on student-generated responses and has shown promise in reducing forgetting relative to supervised fine-tuning (SFT). However, its benefits and fragility remain incompletely understood. We study sequential distillation from multiple teachers, where the student minimizes its average divergence from the teachers. Forward Kullback--Leibler (KL) divergence yields a weighted arithmetic mixture, while reverse KL yields a normalized weighted geometric aggregate. We develop algorithms that learn these targets under off-policy and on-policy feedback, respectively, establishing logarithmic regret bounds in the tabular setting and extending the analysis to function approximation. By analyzing these aggregation targets, we identify mechanisms that help explain both the benefits and fragility of OPD. Relative to forward KL, reverse KL can better retain a confident expert's preferences under uninformative feedback, but is more sensitive to teachers that assign very low probabilities to correct responses. Its token-level conditionals also reveal a dependence on continuation distributions that can favor incorrect prefixes over long horizons.

摘要:在政策蒸餾 (OPD) 中,學生根據教師的反饋學習生成的回應,並在減少相對於監督微調 (SFT) 的遺忘方面顯示出潛力。 然而,它的好處和脆弱性仍然未完全理解。 我們研究來自多位教師的序列蒸餾,學生最小化與教師的平均差異。 前向 Kullback--Leibler (KL) 散度產生加權算術混合,而反向 KL 則產生標準化的加權幾何聚合。 我們開發了在離政策和在政策反饋下學習這些目標的算法,分別在表格設置中建立對數遺憾界限,並將分析擴展到函數近似。 通過分析這些聚合目標,我們確定了幫助解釋 OPD 的好處和脆弱性的機制。 相對於前向 KL,反向 KL 更能在無信息反饋下保留自信專家的偏好,但對於給正確回應分配非常低概率的教師則更敏感。 它的標記級條件也揭示了對延續分佈的依賴,這可能在長期內偏向於不正確的前綴。

Defining and Categorising Human-AI Interactions in Clinical Trials: A Multidimensional Human-AI Classification Approach

2609.38559v1 by Sandra Woolley, Tim Collins, Khalid Khattak, Illia Chernomorets, Ariane Arevalo, Chris Richardson

This paper examines human-AI interactions (HAIIs) in clinical trials and presents a multidimensional categorisation framework that classifies interactions according to AI tasks, human-AI relationships, interaction configurations and interacting human groups. We define HAII, examine existing taxonomies and extend existing categorisation approaches through this novel multidimensional framework. We purposively sampled 15 clinical trials from a previously reported dataset. Each trial was independently categorised by two human reviewers and six large language model (LLM) classifiers. The proposed categorisation provides a structured method for the consistent identification, comparison and synthesis of human-AI interactions across clinical-trial records. The framework is intended to support more consistent comparison and synthesis of AI-related clinical trials and to make explicit the different forms of human involvement associated with AI interventions. The results demonstrate the potential for LLM-assisted categorisation while indicating the continuing importance of human judgement where trial records are incomplete or ambiguous. The principal contribution is a proposed multidimensional framework that brings together AI tasks, human-AI relationships, interaction configurations and interacting human groups within a single approach designed for clinical-trial records. Its significance lies in its potential to support more systematic identification, comparison and synthesis of how humans and AI interact in clinical trials.

摘要:這篇論文探討了臨床試驗中的人類-人工智慧互動(HAIIs),並提出了一個多維度的分類框架,根據人工智慧任務、人類-人工智慧關係、互動配置和互動人群對互動進行分類。我們定義了HAII,檢視現有的分類法,並通過這個新穎的多維框架擴展現有的分類方法。我們有目的地從先前報告的數據集中抽取了15個臨床試驗。每個試驗由兩位人類評審和六個大型語言模型(LLM)分類器獨立進行分類。所提出的分類提供了一種結構化的方法,用於在臨床試驗記錄中一致地識別、比較和綜合人類-人工智慧互動。該框架旨在支持對人工智慧相關臨床試驗的更一致的比較和綜合,並明確不同形式的人類參與與人工智慧干預相關聯。結果顯示了LLM輔助分類的潛力,同時指出在人類判斷仍然重要的情況下,當試驗記錄不完整或模糊時,仍需依賴人類判斷。主要貢獻是一個提出的多維框架,將人工智慧任務、人類-人工智慧關係、互動配置和互動人群整合在一個針對臨床試驗記錄的單一方法中。其重要性在於它能支持對人類和人工智慧在臨床試驗中互動的更系統的識別、比較和綜合。

Demographic Pluralism: Inference-Time Modeling of Pluralistic Human Preference Distributions

2609.38555v1 by Meng-Chen Wu, Qipin Chen, Ansh Jain, Tess Wood, Zhe Du, Si-Chi Chin

Large language models (LLMs) are increasingly used in culturally sensitive settings, where alignment requires representing diverse preferences within populations. Yet existing methods model populations at coarse demographic or community levels and overlook within-group variation. We introduce Demographic Pluralism, an inference-time framework that estimates population-level opinion distributions without opinion-distribution training data or task-specific fine-tuning by generating multiple perspectives within demographically grounded groups. Across four backbones on GlobalOpinionQA and VITAL, it reduces Jensen-Shannon distance by 8.4%-26.4% over Modular Pluralism. Among weighted, equal-weighted, and inverse-weighted aggregation, equal weighting performs best overall; group-level error also increases with group weight, helping explain weighted aggregation's weaker performance.

摘要:大型語言模型(LLMs)在文化敏感的環境中越來越被使用,這些環境中的對齊需要代表人口中的多樣化偏好。然而,現有的方法在粗略的人口或社區層面建模人口,並忽略了群體內的變異。我們引入了人口多元主義,這是一種推斷時框架,通過在以人口為基礎的群體中生成多個視角,來估計人口層級的意見分佈,而不需要意見分佈的訓練數據或特定任務的微調。在 GlobalOpinionQA 和 VITAL 的四個基礎模型中,它將詹森-香農距離減少了 8.4%-26.4%,相較於模組化多元主義。在加權、等權重和反向加權聚合中,等權重的表現最佳;群體層級的誤差也隨著群體權重的增加而增加,這有助於解釋加權聚合的較弱表現。

Does This Action Still Explain the Task? Reverse Scoring for Diffusion Language Model Agents

2609.38536v1 by Jiacheng Qiu, Christopher E. Mower, Jan Peters, Haitham Bou-Ammar, Matthieu Zimmer

Diffusion-based large language models (dLLMs) promise to break the sequential latency bottleneck of autoregressive agents through parallel decoding, but recent evaluations show this efficiency does not transfer to embodied agentic competence: dLLM-backed agents repeatedly fall into retry loops, re-issuing an action long after it has failed. We give a mechanistic account of this failure and a training-free remedy. We trace the retry loop to the adaptivity of masked decoding: the sampler commits the positions it is most confident about and defers the uncertain ones, and at a failure state the context already offers a confident fill for the deferred decision, i.e. the failed action itself, so the retry is committed without the failure feedback ever being confronted. We model the resulting distortion of the action distribution as a task-blind corruption: contextually salient actions (e.g., the action just taken) receive inflated probability by a factor that depends on the state and the action but not on the task. Under this model, we analyse an invariance proposition: the task-blind factor cancels exactly from the reverse conditional, i.e. the likelihood of the task given the state and a candidate action, which coincides with the task posterior of an idealized uncorrupted model. Masked dLLMs evaluate the reverse conditional natively, unlike autoregressive models, by masking the task tokens and denoising, at the cost of a few parallel passes per candidate. We instantiate the rule as Reflect Reverse and evaluate it on four multi-turn embodied benchmarks, where it improves task success and progression rates over forward-scoring baselines.

摘要:擴散基的大型語言模型 (dLLMs) 承諾通過並行解碼打破自回歸代理的序列延遲瓶頸,但最近的評估顯示這種效率並未轉移到具身代理的能力上:基於 dLLM 的代理反覆陷入重試循環,在動作失敗很久之後重新發出動作。我們對這一失敗給出了機制性解釋和無需訓練的補救措施。我們將重試循環追溯到掩蔽解碼的適應性:採樣器承諾其最有信心的位置,並推遲不確定的位置,而在失敗狀態下,上下文已經為推遲的決策提供了一個自信的填充,即失敗的動作本身,因此重試是在從未面對失敗反饋的情況下進行的。我們將動作分佈的扭曲建模為一種任務盲腐敗:在上下文中顯著的動作(例如,剛剛執行的動作)因為一個取決於狀態和動作但不依賴於任務的因子而獲得了膨脹的概率。在這個模型下,我們分析了一個不變性命題:任務盲因子在反向條件中恰好抵消,即給定狀態和候選動作的任務的可能性,這與理想化的未腐敗模型的任務後驗相吻合。掩蔽 dLLMs 本土評估反向條件,與自回歸模型不同,通過掩蔽任務標記和去噪,在每個候選者的幾次並行通過的成本下。我們將該規則具體化為反射反向,並在四個多輪具身基準上進行評估,在這些基準上,它改善了任務成功率和進展率,超過了正向評分的基準。

Caption-Mediated Perceived-Safety Estimation for Pedestrian Routing

2609.38479v1 by Simon Parkinson, Paloma Liu, Wei Zheng, Mohammadreza Sheikhfathollahi

This paper presents an explainable approach to pedestrian routing, in which perceived safety is estimated from street-level imagery through an explicit natural-language intermediate representation. A vision--language model caption is generated and stored before any scoring is undertaken, and the perceived-risk class is derived entirely from structured features of that stored text, so that every segment score remains inspectable by the user. Nine captioning conditions across five model families are benchmarked against a direct Contrastive Language--Image Pre-training (CLIP) image-embedding baseline under an identical downstream pipeline, and the caption-mediated representation is found to reach parity with the image embedding rather than to trail it. The approach was deployed over 654,115 images covering 36 electoral wards in two locations in Northern England (Manchester and Huddersfield). Independent field validation against 3,669 locally collected ratings of 494 images across 70 participant sessions established agreement that is statistically significant but modest, at $r=0.262$, against a measured noise ceiling of 0.737 imposed by disagreement between raters. A single-use confirmatory test then found that a pipeline 44\% stronger on the supervised benchmark did not produce measurable improvement in the field ($r=0.250$, $p=0.84$), so the benchmark gains did not predict the deployment gains in this case. Routing behaviour varies systematically with journey length. There is negligible change below 1\,km, reaching a median increase of 12.78\% in low-risk route length for a median detour of 2.73\% on journeys of 3 to 6 km.

摘要:這篇論文提出了一種可解釋的行人路徑規劃方法,其中感知安全性是通過明確的自然語言中介表示從街景影像中估算得出的。在進行任何評分之前,生成並存儲一個視覺-語言模型的標題,並且感知風險類別完全源自於該存儲文本的結構特徵,以便每個段落的分數都能被用戶檢查。在相同的下游流程下,對五個模型家族中的九個標題條件進行了基準測試,與直接的對比語言-圖像預訓練(CLIP)圖像嵌入基準進行比較,發現標題中介表示的表現達到了與圖像嵌入相當的水平,而不是落後於它。該方法在英格蘭北部的兩個地點(曼徹斯特和哈德斯菲爾德)涵蓋了654,115張影像,涉及36個選區。對494張影像在70個參與者會議中收集的3,669條本地評分進行的獨立現場驗證顯示,達成的協議在統計上顯著但適度,相關係數為$r=0.262$,而評分者之間的不一致造成的噪音上限為0.737。隨後進行的一次單次確認測試發現,在監督基準上強度提高44\%的流程在現場並未產生可測量的改善($r=0.250$,$p=0.84$),因此在這種情況下,基準增益並未預測部署增益。路徑行為隨著行程長度系統性變化。在1公里以下幾乎沒有變化,對於3到6公里的行程,低風險路徑長度的中位數增加達到12.78\%,而中位數繞行為2.73\%。

Aegis: Generative Gradient Masking for Privacy-Preserving Medical Federated Learning

2609.38339v1 by Chaoyu Zhang, Shanghao Shi, Heng Jin, Ning Wang, Y. Thomas Hou, Wenjing Lou

Federated learning (FL) has become a foundational paradigm for multi-institutional medical AI, allowing hospitals and research centers to jointly train diagnostic models without exchanging patient records. This privacy promise, however, is increasingly contested: a malicious or honest-but-curious server can launch model inversion attacks (MIAs) that reconstruct private patient images directly from shared model updates, and recent scalable, closed-form attacks penetrate even secure aggregation at clinically realistic batch sizes. Existing defenses face an unsatisfactory dilemma. Gradient-perturbation methods such as differential privacy and pruning trade away the diagnostic accuracy on which clinical reliability depends, while cryptographic protocols add system complexity yet still leave updates exposed to these scalable attacks. We propose Aegis, a principled client-side defense that breaks this dilemma without perturbing patient data or modifying the FL protocol. Our key insight is that the success of every known MIA is fundamentally bounded by the local batch size relative to the model's leakage capacity; once this limit is exceeded, distinct samples collide and reconstructions collapse into indistinguishable mixtures. Aegis turns this universal bottleneck into a defense: each client superimposes onto its real update a masking gradient computed on locally synthesized, task-relevant data, deliberately pushing the effective batch beyond the attack's recovery capacity. We complement the design with theoretical convergence guarantees under standard convex assumptions and evaluate Aegis on MNIST, CIFAR-10, and three MedMNIST modalities (chest X-ray, abdominal CT, colon pathology). Aegis neutralizes three state-of-the-art MIAs while preserving model utility and incurring only modest overhead, offering a practical privacy primitive for medical FL.

摘要:聯邦學習(FL)已成為多機構醫療人工智慧的基礎範式,使醫院和研究中心能夠共同訓練診斷模型,而無需交換病歷記錄。 然而,這一隱私承諾正受到越來越多的質疑:一個惡意或誠實但好奇的伺服器可以發動模型反演攻擊(MIA),直接從共享的模型更新中重建私人病人影像,而最近可擴展的封閉形式攻擊甚至能夠穿透臨床現實批次大小的安全聚合。 現有的防禦面臨著不令人滿意的困境。 像差分隱私和修剪這樣的梯度擾動方法犧牲了臨床可靠性所依賴的診斷準確性,而加密協議則增加了系統的複雜性,卻仍然讓更新暴露於這些可擴展的攻擊之下。 我們提出了Aegis,一種原則性的客戶端防禦,打破了這一困境,而不擾動病人數據或修改FL協議。 我們的關鍵見解是,每個已知的MIA的成功在根本上受到相對於模型泄漏能力的本地批次大小的限制;一旦超過這一限制,不同的樣本將發生碰撞,重建將崩潰為無法區分的混合物。 Aegis將這一普遍瓶頸轉化為防禦:每個客戶端在其真實更新上疊加一個基於本地合成的、與任務相關的數據計算出的掩蔽梯度,故意將有效批次推向超過攻擊的恢復能力。 我們在標準凸假設下補充了理論收斂保證,並在MNIST、CIFAR-10和三種MedMNIST模態(胸部X光、腹部CT、結腸病理)上評估了Aegis。 Aegis中和了三種最先進的MIA,同時保留了模型的效用,並僅產生適度的開銷,為醫療FL提供了一種實用的隱私原語。

Does Local Video Understanding Transfer Across Encounters? The EgoGears Benchmark

2609.37938v1 by Yuedong Tan, Lei Qi, Yu Liu, Di Wen, Ruiping Liu, Xiaoye Wang, Yufan Chen, Junwei Zheng, Chengzhi Wu, Chen Zhang, Zhihang Chen, Haiwen Sun, Zongwei Wu, Radu Timofte, Danda Pani Paudel, Kunyu Peng

Embodied systems must make knowledge acquired during one encounter usable in another despite changes in viewpoint, motion, and illumination. Yet aggregate cross-video accuracy conflates failures of local perception with failures to preserve observation identity, establish correspondence, and compose evidence, obscuring whether local video understanding actually transfers. We introduce EgoGears, a complementary single- and multi-video benchmark designed to diagnose this transition. It contains 567 single-video and 1,487 multi-video questions derived from 126 human-collected egocentric recordings covering 39 outdoor routes. Repeated traversals across movement speeds and lighting conditions ground comparisons in shared physical environments; 531 questions require alignment across independent recordings. Single-video questions measure the local visual, spatial, and motion evidence available to a model, while multi-video questions test whether evidence remains bound to the correct observation and can be composed into consistent route relationships. We report 29 single-video and 31 multi-video MLLM configurations across six model families in the main leaderboard. Among the 20 configurations evaluated comparably on both splits, every model performs worse on multi-video questions, with a mean decrease of 22.5 percentage points, and the gap persists when answer format and scoring are held fixed. The gap is not explained simply by additional videos or recording boundaries. The central bottlenecks are observation--evidence binding and ordered route-state tracking. The code and benchmark are publicly available at https://github.com/lei-qi-233/EgoGears.

摘要:具身系統必須使在一次遭遇中獲得的知識能夠在另一個遭遇中使用,儘管視角、運動和照明條件有所改變。然而,綜合跨視頻的準確性將局部感知的失敗與未能保持觀察身份、建立對應關係和組合證據的失敗混淆,模糊了局部視頻理解是否真的能夠轉移。我們介紹了EgoGears,一個補充性的單視頻和多視頻基準,旨在診斷這一過渡。它包含567個單視頻和1,487個多視頻問題,這些問題源自126個人類收集的自我中心錄音,涵蓋39條戶外路線。在不同的運動速度和光照條件下的重複遍歷,使比較基於共享的物理環境;531個問題要求在獨立錄音之間進行對齊。單視頻問題測量模型可用的局部視覺、空間和運動證據,而多視頻問題則測試證據是否仍然與正確的觀察綁定,並且能夠組合成一致的路徑關係。我們在主要排行榜上報告了六個模型系列中的29個單視頻和31個多視頻MLLM配置。在20個在兩個分組中進行可比評估的配置中,每個模型在多視頻問題上的表現都較差,平均下降22.5個百分點,並且當答案格式和評分保持固定時,這一差距仍然存在。這一差距並不能簡單地用額外視頻或錄製邊界來解釋。主要瓶頸是觀察—證據綁定和有序路徑狀態跟踪。代碼和基準可在https://github.com/lei-qi-233/EgoGears上公開獲取。

OmniVCBench: Benchmarking Evidence-Grounded Multimodal Reasoning Towards AI Virtual Cells

2609.37773v1 by Manyu Li, Xunkai Li, Yongfu Xiong, Yi Liu, Rong-Hua Li, Guoren Wang

Artificial Intelligence Virtual Cells (AIVCs) are envisioned as scientific agents that simulate cellular responses, explain underlying mechanisms, and support hypothesis-driven discovery. Existing AIVC benchmarks, however, operate primarily at the simulation layer, motivating complementary evaluation of how models interpret experimental evidence and formulate biological hypotheses. We introduce OmniVCBench, a figure-centric, source-traceable benchmark for the interpretation component of an AIVC. It contains 6,077 curated single- and multi-subfigure question--answer pairs derived from figures and experimental contexts in the scientific literature. Guided by Bloom's taxonomy, we instantiate interpretation-layer counterparts of the AIVC Predict--Explain--Discover agenda through three scientific reasoning tasks. We further introduce AIVC-Judge, a task-conditioned MLLM-as-a-judge framework with category-specific, reference-aware rubrics for evaluating open-ended responses. A complementary Model-Derived Hard-Negative Mining (MDHNM) strategy converts plausible errors observed during model inference into MCQ distractors for lower-cost evaluation. Within the evaluated heterogeneous model pool, MCQ accuracy correlates positively with AIVC-Judge scores, providing a complementary view of performance alongside open-response evaluation. Code and data demo are available at https://anonymous.4open.science/r/OmniVCBench.

摘要:人工智慧虛擬細胞(AIVCs)被設想為模擬細胞反應、解釋潛在機制並支持假說驅動發現的科學代理。然現有的AIVC基準主要在模擬層面運作,這促使我們對模型如何解釋實驗證據和制定生物假說進行補充評估。我們介紹OmniVCBench,一個以圖形為中心、可追溯來源的AIVC解釋組件基準。它包含6,077個經過策劃的單圖和多圖問題--答案對,這些問題來自科學文獻中的圖形和實驗背景。在布魯姆的分類法指導下,我們通過三個科學推理任務實現AIVC預測--解釋--發現議程的解釋層對應。我們進一步介紹AIVC-Judge,一個任務條件的MLLM作為評判框架,具有特定類別的、參考意識的評分標準,用於評估開放式回答。一個補充的模型衍生困難負樣本挖掘(MDHNM)策略將在模型推理過程中觀察到的合理錯誤轉換為多選題的干擾項,以降低評估成本。在評估的異質模型池中,多選題的準確率與AIVC-Judge的分數呈正相關,提供了與開放式回答評估相輔相成的性能視角。代碼和數據演示可在https://anonymous.4open.science/r/OmniVCBench獲得。

Multi-Site Real-World Performance of Commercial AI for Pulmonary and Incidental Pulmonary Embolism Detection

2609.37750v1 by Aawez Mansuri, Mohammadreza Chavoshi, Theodorus Dapamede, Wasif Bala, Beatrice Brown-Mulry, Rohan Isaac, Bardia Khosravi, Hanzhou Li, Frank Li, John T. Moon, Chad Robichaux, Dan I. G. Cohen-Addad, Ninad V. Salastekar, Janice Newsome, Judy W. Gichoya, Hari Trivedi

Pulmonary embolism (PE) is a leading cause of cardiovascular mortality, yet the real-world performance of FDA-cleared AI detection models remains incompletely characterized. We retrospectively evaluated two FDA-cleared AI algorithms from a single commercial platform (Aidoc Medical BriefCase), one for PE triage on dedicated CT pulmonary angiography (CTPA; n = 30,678) and one for incidental PE (iPE) detection on routine contrast-enhanced CTs (n = 37,191), across a 17-facility academic health system. Reference-standard labels were extracted from radiology reports using a validated LLM pipeline (97% accuracy, kappa = 0.94). The PE model achieved 86.8% sensitivity and 99.1% specificity, with sensitivity declining from 99.3% for saddle emboli to 72.9% for subsegmental PE, and from 89.7% for acute to 65.3% for non-acute PE. The iPE model achieved 73.5% sensitivity and 99.8% specificity. Both models demonstrated lower sensitivity than FDA-clearance benchmarks while exceeding cleared specificity, with diminishing performance for peripheral and non-acute emboli mirroring known human reader limitations and underscoring the need for standardized post-market surveillance of AI-enabled medical devices.

摘要:肺栓塞(PE)是心血管死亡的主要原因,但FDA批准的AI檢測模型在實際應用中的表現仍然未完全明確。我們回顧性地評估了來自單一商業平台(Aidoc Medical BriefCase)的兩個FDA批准的AI算法,一個用於專用CT肺動脈造影(CTPA;n = 30,678)的PE分流,另一個用於常規對比增強CT(n = 37,191)的偶然PE(iPE)檢測,涵蓋了17家學術醫療系統。參考標準標籤是通過經驗證的LLM管道(97%準確率,kappa = 0.94)從放射學報告中提取的。PE模型的敏感性達到86.8%,特異性為99.1%,其中敏感性從鞍狀栓塞的99.3%下降到亞段PE的72.9%,從急性PE的89.7%下降到非急性PE的65.3%。iPE模型的敏感性為73.5%,特異性為99.8%。兩個模型的敏感性均低於FDA批准的基準,但特異性超過批准標準,對於周邊和非急性栓塞的表現下降反映了已知的人類讀者限制,並強調了對AI驅動醫療設備標準化市場後監測的需求。

XU-RS: Explaining Credal Width in Random-Set Language Models

2609.37594v1 by David Achara, Maryam Sultana, Alexander D. Rast, Fabio Cuzzolin

Uncertainty estimates tell us how unsure a model is, but not why. Without knowing which parts of an input influences a model's uncertainty, we cannot tell whether that uncertainty score depends on input features that are relevant for the task. We study this problem in randomset classifiers built using pretrained language models. These classifiers assign probability to individual answers and to groups of answers, producing lower and upper probabilities for each answer; The difference between these probabilities, called credal width, is used to represent epistemic uncertainty about an answer arising from limited training data. We propose XU-RS, a framework that attributes an answer's credal width to the input tokens (words or word pieces) supplied to a language model. XU-RS uses Expected Gradients (a standard feature attribution method) to estimate how input tokens contribute to credal width. The proposed framework is evaluated on a MedQA dataset using SmolLM3-3B and Llama-2-7B models, demonstrating that setting the embedding of a token ranked highly by XU-RS to zero (zero-masking) causes larger changes in credal width than zero-masking randomly selected tokens. In addition, we show that normalisation can cause other answer groups to influence an answer's width, reveal how token attribution can mask numerical errors, and provide diagnostic checks to verify whether a token ranked highly by XU-RS meaningfully explains model uncertainty.

摘要:不確定性估計告訴我們模型有多不確定,但並不告訴我們原因。若不知道輸入的哪些部分影響模型的不確定性,我們無法判斷該不確定性分數是否依賴於與任務相關的輸入特徵。 我們在使用預訓練語言模型構建的隨機集分類器中研究這個問題。這些分類器為單個答案和答案組分配概率,為每個答案生成下限和上限概率;這些概率之間的差異稱為信念寬度,用於表示由於訓練數據有限而產生的對答案的認識不確定性。我們提出了XU-RS,一個將答案的信念寬度歸因於提供給語言模型的輸入標記(單詞或詞片)的框架。XU-RS使用期望梯度(標準特徵歸因方法)來估計輸入標記對信念寬度的貢獻。所提出的框架在使用SmolLM3-3B和Llama-2-7B模型的MedQA數據集上進行評估,證明將XU-RS排名較高的標記的嵌入設置為零(零掩蔽)會導致信念寬度的變化比隨機選擇的標記的零掩蔽更大。此外,我們展示了正規化可能導致其他答案組影響答案的寬度,揭示了標記歸因如何掩蓋數值錯誤,並提供診斷檢查以驗證XU-RS排名較高的標記是否有意義地解釋模型的不確定性。

Raw Imagery Impacting Your AI: Should You Care?

2609.38265v1 by Adrien Dorise, Marjorie Bellizzi, Stéphane May

Onboard AI is gaining interest for space applications such as vessel, wildfire, and cloud detection, where real-time processing can improve mission reactivity and reduce downlink needs. However, onboard models may operate on raw or minimally processed imagery rather than on restored ground products. This study evaluates how image degradation affects object detection by varying Signal-to-Noise Ratio (SNR), Modulation Transfer Function (MTF) at Nyquist, and Ground Sampling Distance (GSD). Controlled degradations are applied to Very High Resolution Maxar imagery, and three lightweight detectors, YOLOv5s, YOLOX-S, and NanoDet, are evaluated on the resulting operating points. The results show that the impact of image quality depends on the degradation mechanism, and that increasing degradation does not necessarily lead to a proportional decrease in vessel detection performance. GSD produces the most consistent performance shift, while MTF and SNR effects depend more on the model and resolution. Severe combinations of blur and noise produce the largest losses. These results provide task-level information that can support sensor, processing, and AI trade-offs for future onboard systems.

摘要:在太空應用中,機載人工智慧正受到關注,例如船隻、野火和雲層檢測,其中即時處理可以提高任務反應能力並減少下行鏈路需求。然而,機載模型可能在原始或經過最小處理的影像上運作,而不是在恢復的地面產品上。本研究評估影像退化如何影響物體檢測,通過改變信噪比(SNR)、奈奎斯特的調變傳遞函數(MTF)和地面取樣距離(GSD)。對非常高解析度的Maxar影像施加控制退化,並在結果操作點上評估三個輕量級檢測器,YOLOv5s、YOLOX-S和NanoDet。結果顯示,影像質量的影響取決於退化機制,且增加退化不一定會導致船隻檢測性能成比例下降。GSD產生最一致的性能變化,而MTF和SNR的影響則更多地依賴於模型和解析度。模糊和噪聲的嚴重組合會產生最大的損失。這些結果提供了任務級別的信息,可以支持未來機載系統的傳感器、處理和人工智慧的權衡。

Selecting The Most Informative Tokens in Natural Language Autoencoders

2609.37040v1 by Federico Torrielli, Gianluca Barmina, Andrea Blasi Núñez, Amon Rapp, Luigi Di Caro, Peter Schneider-Kamp, Lukas Galke Poech

Natural language autoencoders translate a language model's internal activations into readable explanations. Explaining every token position is costly. Which positions should an auditor inspect to understand a potential threat? We study this question across $4.7$ million explanations on prompt injection and concealment. We compare signals from model computation with a ranker trained only on chat structure. Chat structure usually selects more relevant explanations than the computational signals, without requiring a model forward pass for position selection. On three of four datasets, explaining just $5\%$ of positions retains nearly all of the success rate from explaining every position, where success means obtaining an explanation about the threat. The benefit varies with the audit task. We also show that pretrained verbalizers recover words that models have learned to conceal through fine-tuning, without additional verbalizer training. These results identify where auditors can concentrate explanation generation and show that useful explanations can extend beyond the model a verbalizer was trained to describe.

摘要:自然語言自動編碼器將語言模型的內部激活轉換為可讀的解釋。解釋每個標記位置是昂貴的。審計員應該檢查哪些位置以理解潛在威脅?我們在 $4.7$ 百萬個關於提示注入和隱藏的解釋中研究了這個問題。我們將模型計算的信號與僅基於聊天結構訓練的排名器進行比較。聊天結構通常選擇比計算信號更相關的解釋,而不需要模型前向傳遞來選擇位置。在四個數據集中有三個中,僅解釋 $5\%$ 的位置幾乎保留了解釋每個位置的所有成功率,其中成功意味著獲得有關威脅的解釋。這一好處隨著審計任務而異。我們還顯示,預訓練的詞語化器能夠恢復模型通過微調學會隱藏的單詞,而不需要額外的詞語化器訓練。這些結果確定了審計員可以集中解釋生成的地方,並顯示有用的解釋可以超越詞語化器所訓練描述的模型。

Harness Evolution as Learning: Approximation, Generalization, and Optimization Limits of Self-Improving Personal Agents

2609.36892v1 by Zeyu Gan, Zixuan Gong, Yong Liu

As the capabilities of large language models (LLMs) continue to advance, increasing attention is turning to how to translate their abilities into useful behavior. Personal agents bring this question into everyday settings, where models are expected to serve individual users and continually adapt to their preferences. With the underlying model held fixed, such adaptation relies on harness engineering: designing and evolving the surrounding layer that manages context, memory, tools, and execution. Despite rapid progress, the factors governing effective harness evolution remain insufficiently understood. To narrow this gap, we investigate three central questions concerning harness architecture, harness scale, and self-evolution algorithms through complementary empirical and theoretical analyses. Empirically, we introduce a preference-oriented benchmark and systematically characterize the capabilities and limitations of personal agents associated with these three dimensions. Theoretically, we formulate harness evolution as a learning problem and explain these phenomena through approximation, generalization, and optimization errors. Analyses of reachable policies, capacity under finite interaction evidence, and biased update dynamics provide theoretical accounts of the observed phenomena. Together, these results offer a unified perspective on the limits of personalization through harness evolution and inform future harness design.

摘要:隨著大型語言模型(LLMs)能力的持續進步,越來越多的關注轉向如何將其能力轉化為有用的行為。個人代理將這個問題帶入日常環境,在這裡模型被期望為個別用戶服務並不斷適應他們的偏好。在基礎模型保持固定的情況下,這種適應依賴於飼養工程:設計和發展管理上下文、記憶、工具和執行的周邊層。儘管進展迅速,但影響有效飼養演變的因素仍然理解不足。為了縮小這一差距,我們通過互補的實證和理論分析,研究有關飼養架構、飼養規模和自我演變算法的三個核心問題。在實證方面,我們引入了一個以偏好為導向的基準,並系統性地描述與這三個維度相關的個人代理的能力和局限性。在理論方面,我們將飼養演變公式化為一個學習問題,並通過近似、泛化和優化誤差解釋這些現象。可達政策、有限互動證據下的容量和偏見更新動態的分析提供了對觀察到的現象的理論解釋。總體而言,這些結果提供了對通過飼養演變實現個性化限制的統一視角,並為未來的飼養設計提供了指導。

Predictive Credit: Measuring What Scientific Explanations Add to Experimental Forecasts

2610.00314v1 by Jingjie Ning, Xueqi Li, Yibo Kong, Dongting Li

Research agents explain planned experiments. We measure predictive credit with paired forecasts sharing an intervention, forecaster, and outcome while varying description, matched explanation, and donor context. Five checks track commitment, delivery, predictive gain, alignment, and known-signal uptake. Across 336 prospective states in controlled learning, 12 Tox21 endpoints, and 24 OpenML tasks, v5's frozen credit decision was inconclusive. Tox21's preregistered ROC AUC interval-score harm test was unmet ($D-M=-.0026$, 95 percent interval [$-.0174$, .0104]); OpenML's joint formation, point-equivalence, and repeatability rule was unmet. Matched point-accuracy gains over description remained unconfirmed, and Tox21/OpenML seed-donor intervals spanned zero. Under requested DeepSeek V4 Pro, matched and donor cards reduced secondary Tox21 drift by 64.5 and 59.1 percent. A DeepSeek V4 Flash replay raised matched point MAE from .01823 to .02020 and missed matched-donor interval-score equivalence. OpenML full-card assignment widened nominal 80 percent intervals by 21 percent, with 49.3 percent coverage versus 51.4 percent for description and content in 66/144 cards. Direct-text Flash delivered all 144 notes without detectable matched point-accuracy gain. A researcher-authored mechanism positive control lowered point MAE by 2.60 percentage points versus description. The protocol measures predictive credit for research-agent benchmarks and scientific forecasting; natural-explanation credit remained unconfirmed at the tested donor resolutions.

摘要:研究代理人解釋計劃的實驗。我們通過共享干預、預測者和結果的配對預測來測量預測信用,同時變化描述、匹配解釋和捐贈者背景。五個檢查跟踪承諾、交付、預測增益、一致性和已知信號的吸收。在336個受控學習的前瞻性狀態、12個Tox21端點和24個OpenML任務中,v5的凍結信用決策結果不確定。Tox21的預註冊ROC AUC區間分數損害測試未達標($D-M=-.0026$,95百分位區間[$-.0174$,.0104]);OpenML的聯合形成、點等價和重複性規則未達標。對於描述的匹配點準確度增益仍未得到確認,Tox21/OpenML種子-捐贈者區間跨越零。在請求的DeepSeek V4 Pro下,匹配和捐贈卡減少了次級Tox21漂移64.5和59.1百分比。DeepSeek V4 Flash重播將匹配點的MAE從.01823提高到.02020,並錯過了匹配-捐贈者區間分數等價。OpenML全卡分配將名義80百分比區間擴大了21百分比,對於66/144張卡片,覆蓋率為49.3百分比,而描述和內容的覆蓋率為51.4百分比。直接文本Flash交付了所有144條備註,沒有檢測到匹配點準確度的增益。一個研究者撰寫的機制正控制降低了點MAE,相對於描述降低了2.60個百分點。該協議測量研究代理基準和科學預測的預測信用;在測試的捐贈者解析度下,自然解釋信用仍未得到確認。

Engineering Simplicity: Simple Mechanism Interfaces Steer LLM Agents

2609.36365v1 by Kehang Zhu, Anand Shah, David Parkes

Can interaction formats and textual scaffolds help large language model (LLM) agents make better decisions, and do better decisions come with better explanations? We study these questions in auctions and matching, multi-agent environments with explicit rules and known optimal strategies. These settings let us vary how a decision problem is presented while retaining a benchmark for evaluating behavior. Drawing on human-motivated theories of simplicity, we compare interfaces that elicit a complete bid or ranking with sequential interfaces that make safe choices easier to identify. We then hold the interaction format fixed and vary reasoning scaffolds and rule descriptions. Across four model families, the ascending auction interface substantially reduces bid deviations. The matching comparison also shows why sequential responses require different error accounting from complete rankings. Laying out payoff contingencies and explaining why truth-telling is safe also improve choices, whereas prompts to plan through matching rounds or form beliefs about opponents worsen play overall. In auctions, these behavioral gains are not accompanied by corresponding improvements in measured verbal indicators of strategic understanding in the agents' short stated plans. Other prompts change those indicators without improving bids. Our findings suggest that human-motivated theories of simplicity can inform the design of decision environments for artificial agents. They also show why scaffolds should be evaluated through realized choices as well as explanations: improvements in one need not appear in the other.

摘要:互動格式和文本支架能否幫助大型語言模型 (LLM) 代理做出更好的決策,而更好的決策是否伴隨著更好的解釋?我們在拍賣和匹配的多代理環境中研究這些問題,這些環境具有明確的規則和已知的最佳策略。這些設置使我們能夠變化決策問題的呈現方式,同時保留評估行為的基準。基於人類動機的簡單性理論,我們比較了引發完整出價或排名的介面與使安全選擇更容易識別的序列介面。我們然後固定互動格式,變化推理支架和規則描述。在四個模型家族中,升序拍賣介面顯著減少了出價偏差。匹配比較也顯示為什麼序列反應需要不同的錯誤計算,與完整排名相比。列出支付條件並解釋為什麼誠實報告是安全的也改善了選擇,而計劃通過匹配輪次或形成對對手的信念的提示則整體上惡化了遊戲。在拍賣中,這些行為上的增益並未伴隨著代理短期陳述計劃中戰略理解的口頭指標的相應改善。其他提示改變了這些指標,但並未改善出價。我們的發現表明,人類動機的簡單性理論可以為人工代理的決策環境設計提供指導。它們還顯示為什麼支架應該通過實現的選擇以及解釋來評估:一方面的改善不一定會在另一方面出現。

Explainability from Training with Applications to TCR-Epitope Prediction

2609.36354v1 by Jiarui Li, Zixiang Yin, Samuel Landry, Zhengming Ding, Ramgopal Mettu

Deep learning models have achieved strong performance in artificial intelligence for science, yet their black-box nature limits our understanding of how they learn scientific tasks. Existing methods for interpretability provide limited insight into how models organize evidence and evolve during learning. We introduce explainability from training (EFT), a model-agnostic paradigm that traces model interpretation during training to explain why models rely on specific features and how they organize these features as predictive evidence. We apply EFT to four state-of-the-art T cell receptor (TCR)-epitope prediction models, TCR-SRIM, TULIP, MixTCRpred, and NetTCR-2.2, spanning post-hoc and interpret-by-design approaches as well as transformers and CNNs. To investigate how structural information affects model explanations, we introduce a benchmark, TCR-XAI2, containing 388 unique experimentally resolved TCR-epitope structures, complemented by structures predicted using AlphaFold3, Boltz-2, TCRModel2, tFold-TCR, and OpenFold3. Using EFT with TCR-XAI2, we demonstrate that (1) CNN and transformer models exhibit distinct learning trajectories; (2) TCR $α$ and $β$ evidence can conflict during learning, limiting the benefits of jointly modeling both chains, while MHC information mitigates this; and (3) real versus predicted structural data for TCR-epitope prediction exhibits distinct TCR and peptide feature preferences as well as differing trajectories of model certainty.

摘要:深度學習模型在科學人工智慧中取得了強大的表現,但其黑箱特性限制了我們對它們如何學習科學任務的理解。現有的可解釋性方法對模型如何組織證據和在學習過程中如何演變提供的見解有限。我們介紹了訓練中的可解釋性(EFT),這是一種與模型無關的範式,追蹤模型在訓練過程中的解釋,以解釋為什麼模型依賴於特定特徵以及它們如何將這些特徵組織為預測證據。我們將EFT應用於四個最先進的T細胞受體(TCR)-表位預測模型,TCR-SRIM、TULIP、MixTCRpred和NetTCR-2.2,涵蓋了事後解釋和設計解釋的方法,以及Transformer和卷積神經網絡。為了研究結構信息如何影響模型解釋,我們引入了一個基準,TCR-XAI2,包含388個獨特的實驗解析TCR-表位結構,並補充了使用AlphaFold3、Boltz-2、TCRModel2、tFold-TCR和OpenFold3預測的結構。使用EFT和TCR-XAI2,我們證明了(1)CNN和Transformer模型顯示出不同的學習軌跡;(2)TCR $α$ 和 $β$ 證據在學習過程中可能會衝突,限制了同時建模這兩條鏈的好處,而MHC信息則減輕了這一點;(3)TCR-表位預測的實際結構數據與預測結構數據顯示出不同的TCR和肽特徵偏好,以及不同的模型確定性軌跡。

ThuRunel: Dynamic Decoupling for Structured Advisory Dialogue

2609.36340v1 by Yuyan Chen

High-stakes advisory domains such as medical aesthetics, legal consultation, and educational planning exhibit a two-phase structure. The early phase requires empathetic elicitation and emotional support, and the late phase requires authoritative specialist judgment. Neither fully automated agents nor human junior consultants adequately address this structure at scale. We formalize the core design challenge as dynamic decoupling, asking how an AI advisory agent should decide what to ask, when to stop, what to resolve autonomously, and what to forward to the specialist. We present ThuRunel, an advisory agent combining a finite-state belief management framework, a chain-of-thought teacher synthesis protocol, and learned generation adapters. Against eleven baselines, ThuRunel achieves consistent improvements in elicitation completeness and specialist brief quality. ThuRunel is publicly deployed as a bilingual web application in which the same decoupling decisions operate from the client's side, grounded in a curated knowledge base that cites its sources in every answer.

摘要:高風險的諮詢領域如醫療美學、法律諮詢和教育規劃展示出雙階段結構。早期階段需要同理心的引導和情感支持,而後期階段則需要權威專家的判斷。無論是完全自動化的代理還是人類初級顧問,都無法在規模上充分應對這一結構。我們將核心設計挑戰形式化為動態解耦,詢問AI諮詢代理應如何決定詢問什麼、何時停止、什麼可以自主解決以及什麼需要轉交給專家。我們提出了ThuRunel,一個結合有限狀態信念管理框架、思維鏈教師綜合協議和學習生成適配器的諮詢代理。在十一個基準測試中,ThuRunel在引導完整性和專家簡報質量上實現了一致的改進。ThuRunel作為一個雙語網絡應用程序公開部署,其中相同的解耦決策從客戶端運作,基於一個策劃的知識庫,並在每個答案中引用其來源。

FigAct: Turning Scientific Figures into Active Canvases for Explanation

2609.36190v1 by Shishi Xiao, Zichao Wang, Alexa Siu, David H. Laidlaw, Jennifer Healey

Scientific figures are designed to communicate information visually, yet MLLMs typically explain them by translating their visual content back into text. This requires readers to manually map the resulting explanations back to the figure. Inspired by how people present visual information, we introduce FigAct, a framework that transforms static scientific figures into question-conditioned visual presentations by acting directly on their existing graphical elements. Like a human presenter, FigAct generates a sequence of short narrations, grounds each narration in the corresponding visual evidence, and applies visual actions to guide the viewer's attention. We develop a hierarchical search strategy for efficient element localization, reducing token usage by approximately 40$\times$. We further train FigAct-8B using three task-specific rewards for grounding accuracy, search efficiency, and rendering quality. We further build a human-verified benchmark from figures in real-world scientific papers to evaluate the ability of MLLMs to generate grounded visual explanations. Our results demonstrate the effectiveness of FigAct and show that treating scientific figures as presentation canvases makes explanations clearer and easier to follow.

摘要:科學圖形旨在以視覺方式傳達信息,但 MLLMs 通常通過將其視覺內容轉換回文本來解釋它們。這要求讀者手動將生成的解釋映射回圖形。受到人們如何呈現視覺信息的啟發,我們介紹了 FigAct,一個通過直接作用於現有圖形元素將靜態科學圖形轉換為基於問題的視覺展示的框架。像人類演講者一樣,FigAct 生成一系列簡短的敘述,將每個敘述與相應的視覺證據相結合,並應用視覺動作來引導觀眾的注意力。我們開發了一種層次搜索策略,以提高元素定位的效率,將標記的使用減少約 40$\times$。我們進一步使用三個特定任務的獎勵來訓練 FigAct-8B,以提高基準準確性、搜索效率和渲染質量。我們還從現實世界的科學論文中的圖形建立了一個經過人工驗證的基準,以評估 MLLMs 生成有根據的視覺解釋的能力。我們的結果證明了 FigAct 的有效性,並顯示將科學圖形視為展示畫布使解釋更清晰、更易於理解。

An Exact Generate - Transform Decomposition of Small-LLM Team Scaling Across Orchestration Architectures

2609.36104v1 by Blaz Bertalanic, Carolina Fortuna

Replacing one LLM agent with a collaborating team can raise accuracy, but whether scaling the team helps, and which architecture to scale, is unclear. Sweeping eight agent orchestration architectures across five instruction-tuned 7-9B models, five short-answer benchmarks, and an executable-code benchmark up to 30 calls, we find that the returns to team scaling are sharply task-dependent: from three to thirty calls accuracy rises by up to 17 points on the two arithmetic word-problem benchmarks (GSM8K, GSMHard) but by at most four on ARC, GPQA, and MMLU, for every architecture, a split the usual task-averaged number conceals. Proposer-Critic captures the arithmetic gains, scaling steepest and, in aggregate, surpassing every other architecture at the largest budget (item-clustered intervals exclude zero), though it ranks among the weakest elsewhere, and no architecture wins across tasks. We explain these trajectories with an exact generate-transform decomposition. Partitioning any workflow into proposal coverage and a downstream transform, any accuracy change splits exactly into an extensive coverage dividend and an intensive transformation change. The decomposition diagnoses each task: arithmetic offers coverage headroom that a critic-guided transform converts, whereas the multiple-choice benchmarks either saturate in coverage or fail to convert it, and on open-ended code generative recovery nearly vanishes so accuracy tracks coverage. At equal call budgets token cost still varies 2.1x. Extra calls therefore create candidate opportunity that only some architectures, on some tasks, convert. Team scaling is a task- and architecture-specific bet, not a uniform lever.

摘要:替換一個 LLM 代理為一個協作團隊可以提高準確性,但擴大團隊是否有幫助,以及擴大哪種架構仍不清楚。對於五個經過指令調整的 7-9B 模型、五個短答案基準以及一個可執行代碼基準進行八種代理協調架構的廣泛測試,最多進行 30 次調用,我們發現團隊擴大的回報明顯依賴於任務:在兩個算術文字問題基準(GSM8K、GSMHard)上,從三次到三十次調用的準確性提高了最多 17 個點,但在 ARC、GPQA 和 MMLU 上最多只提高四個點,這是每種架構的情況,通常的任務平均數隱藏了這一點。提議者-評價者捕捉了算術增益,擴大最陡峭,並且在總體上超越了其他所有架構在最大預算下(項目聚類區間不包括零),儘管在其他地方排名較弱,且沒有任何架構在所有任務中獲勝。
我們用精確的生成-轉換分解來解釋這些軌跡。將任何工作流程劃分為提議覆蓋和下游轉換,任何準確性變化都精確地分為廣泛的覆蓋紅利和密集的轉換變化。這一分解診斷每個任務:算術提供了覆蓋的頭部空間,評價者引導的轉換將其轉換,而多選基準要麼在覆蓋上飽和,要麼未能轉換,並且在開放式代碼生成恢復中幾乎消失,因此準確性跟蹤覆蓋。在相等的調用預算下,令牌成本仍然變化 2.1 倍。因此,額外的調用創造了候選機會,只有一些架構在某些任務上能夠轉換。團隊擴大是一個特定於任務和架構的賭注,而不是一個統一的槓桿。

One Geometry, Different Outcomes: Readout-Dependent Effects of the Modality Gap in Vision-Language Models

2609.36101v1 by Aditya Sharma, Divya Saxena

Contrastive vision-language models learn shared embedding spaces by aligning matched image-text pairs, yet their representations remain separated by a modality gap. Prior work reports divergent effects of modifying this gap: reducing it can improve zero-shot classification and cross-modal alignment, whereas removing gap-related structure can degrade image-text retrieval. In this paper, we provide a unified geometric explanation for these task-dependent effects. Across CLIP and SigLIP encoders, we find that a single dominant direction captures 94.4-99.9% of the squared norm of the image-text mean separation, revealing that the mean-separation component is approximately rank-one. A decomposition of the similarity score then identifies three task-specific roles. In zero-shot classification, query-side fixed gap-offset subtraction is exactly equivalent to an additive class bias. In standard cross-modal retrieval, projecting out the gap direction and renormalising residuals discards candidate-specific norm information, inducing a multiplicative ranking distortion; a geometry-derived exponent tracks the grid-search optimum (Spearman rho = 0.93) and restores performance in some settings, although the gains transfer unevenly. In mixed-modal retrieval, the gap direction sorts candidates by modality; its removal can improve cross-modal ranking, unlike random or non-gap controls. Residual semantic structure after removal defines the limits of the rank-one account. Together, these results explain why gap modification can improve, degrade, or restore performance across downstream settings. By clarifying when and why gap modification changes model behavior, this account provides a principled basis for selecting gap interventions in similarity-based vision-language systems across evaluated downstream tasks.

摘要:對比視覺-語言模型通過對齊匹配的圖像-文本對來學習共享的嵌入空間,但它們的表示仍然受到模態差距的分隔。先前的研究報告了修改這一差距的不同效果:減少它可以改善零樣本分類和跨模態對齊,而去除與差距相關的結構則可能降低圖像-文本檢索的效果。在本文中,我們提供了一個統一的幾何解釋來說明這些依賴於任務的效果。在CLIP和SigLIP編碼器中,我們發現一個主導方向捕捉了94.4-99.9%的圖像-文本平均分離的平方範數,揭示了平均分離成分大約是秩一的。然後,對相似度分數的分解確定了三個特定於任務的角色。在零樣本分類中,查詢端固定的差距偏移減法與加性類別偏差完全等價。在標準的跨模態檢索中,投影出差距方向並重新標準化殘差會丟棄候選特定的範數信息,導致乘法排名失真;一個幾何推導的指數跟踪網格搜索最佳(Spearman rho = 0.93),並在某些設置中恢復性能,儘管增益轉移不均勻。在混合模態檢索中,差距方向按模態對候選進行排序;其去除可以改善跨模態排名,這與隨機或非差距控制不同。去除後的殘餘語義結構定義了秩一解釋的極限。總體而言,這些結果解釋了為什麼差距修改可以改善、降低或恢復下游設置中的性能。通過澄清何時以及為什麼差距修改會改變模型行為,這一解釋為在評估的下游任務中選擇基於相似性的視覺-語言系統中的差距干預提供了原則性的基礎。

Shockingly Simple Self-retrospection Improves Agentic Models Without RL

2609.35741v1 by Jonathan Light, Christopher Zhang Cui, Jeonghye Kim, Roger Creus Castanyer, Emiliano Penaloza, Zhengyan Shi, Alessandro Sordoni, Marc-Alexandre Côté, Xingdi Yuan, Minseon Kim

People learn not only by repeating successful actions, but also by recounting and explaining their experiences, revising their understanding to guide future behavior. Can a language-model agent improve its future actions by training only on explanations of its own experience? We investigate this question by studying Retrospection-Only Fine-Tuning (ROFT), a minimal online procedure designed to isolate the effect of explanation-only training on subsequent behavior. The agent attempts a task, observes available feedback, generates a retrospective explanation, and is fine-tuned with a next-token prediction loss on the explanation tokens alone. The procedure uses neither an external teacher nor a reward-based policy update. In software-engineering experiments with Qwen3.5-4B, ROFT is trained on problems with mixed successful and unsuccessful base-model attempts. On held-out SWE-bench Verified and Pro, it reaches 49.2% and 26.8% solve rates after 20 updates without using a verifier, compared with GRPO's 48.0% and 25.3% after 40 updates in the evaluated runs, and makes faster early progress in training time and sampled attempts. It also learns to solve individual tasks on which all 64 sampled base-model attempts failed, showing that learning can begin without any initially successful trajectories. Behavioral analyses find that ROFT indirectly assigns credit to actions, encouraging good actions and discouraging incorrect ones. Moreover, prompting retrospections to emphasize more direct solutions yields shorter subsequent attempts even without an explicit length penalty. Together, these findings show that learning to explain can also improve learning to do, establishing self-generated retrospections as useful training targets and motivating further study of explanation-to-action transfer.

摘要:人們不僅通過重複成功的行動來學習,還通過敘述和解釋他們的經驗,修正他們的理解以指導未來的行為。語言模型代理能否僅通過對自身經驗的解釋進行訓練來改善其未來的行動?我們通過研究回顧性僅微調(Retrospection-Only Fine-Tuning, ROFT)來探討這個問題,這是一種旨在孤立解釋性訓練對後續行為影響的最小在線程序。代理嘗試一個任務,觀察可用的反饋,生成一個回顧性解釋,並僅對解釋標記進行下一標記預測損失的微調。該程序既不使用外部教師,也不進行基於獎勵的策略更新。在使用 Qwen3.5-4B 的軟體工程實驗中,ROFT 在成功和不成功的基模型嘗試混合的問題上進行訓練。在保留的 SWE-bench Verified 和 Pro 上,它在 20 次更新後達到 49.2% 和 26.8% 的解決率,而不使用驗證器,與 GRPO 在評估運行中 40 次更新後的 48.0% 和 25.3% 相比,並在訓練時間和抽樣嘗試中更快地取得早期進展。它還學會了解決所有 64 次抽樣基模型嘗試失敗的個別任務,顯示學習可以在沒有任何最初成功的軌跡的情況下開始。行為分析發現,ROFT 間接地將功勞分配給行動,鼓勵良好的行動並抑制不正確的行動。此外,促使回顧以強調更直接的解決方案,即使在沒有明確的長度懲罰的情況下,也會產生更短的後續嘗試。這些發現表明,學習解釋也可以改善學習行動,確立自我生成的回顧作為有用的訓練目標,並激勵進一步研究解釋到行動的轉移。

Rethinking Circuit Evaluation: Do Circuits Explain Model Errors?

2609.35686v1 by Li Zhang, Chuqin Geng, Mark Zhang, Chen Yang, Luke Zhang, Haolin Ye, Xujie Si

Mechanistic interpretability (MI) aims to explain a model's behaviour through analyzing its internal computations; circuit-based explanations aim to isolate these computations with compact subnetworks validated by ablating the rest of the model. We show that circuits validated this way may fail to recover the underlying mechanism of the model's behaviour by closely reproducing its successful decisions while failing to account for most of its errors. Such explanations should account for the model's particular errors as well as its successes. We evaluate this requirement by measuring exact answer agreement separately on model successes and failures, across circuit sizes and ablation settings, on IOI, Docstring, and six model-task settings from the Mechanistic Interpretability Benchmark. We discover that many tested circuits closely replicate correct behaviour while missing most of the model's errors. On indirect object identification (IOI) for GPT-2 small, under mean ablation, the manual circuit and tested automated circuits, including one trained against the model's full output distribution, agree with the model on 97.3-99.5% of prompts it answers correctly but only 11.4-41.7% of errors. An IOI case study shows that lost errors are recoverable by restoring omitted attention-heads which raise error reproduction from 14.2% to 75.1% on a separate held-out set with 0.41 percentage point decrease on correct agreement, exceeding matched random extensions and scalar-biased control. Intervention traces show how omitted computations produce specific wrong answers for a reproducible subset of errors. In all, these findings show circuits can preserve task success without adequately explaining model's failures, and support exact error reproduction as a necessary, but not sufficient, test of circuit-based explanations of model behaviour.

摘要:機械解釋性(MI)旨在通過分析模型的內部計算來解釋模型的行為;基於電路的解釋旨在通過消除模型的其餘部分來隔離這些計算,並使用緊湊的子網絡進行驗證。我們顯示,這種方式驗證的電路可能無法恢復模型行為的基本機制,因為它們雖然能夠緊密再現成功的決策,但卻未能考慮到大多數錯誤。這樣的解釋應該考慮到模型的特定錯誤以及它的成功。我們通過在模型的成功和失敗之間,分別測量準確答案的一致性,來評估這一要求,並在不同的電路大小和消融設置下,使用 IOI、Docstring 和來自機械解釋基準的六個模型任務設置。我們發現,許多測試過的電路能夠緊密複製正確的行為,但卻錯過了模型的大多數錯誤。在 GPT-2 small 的間接物體識別(IOI)中,在平均消融下,手動電路和測試過的自動電路,包括一個針對模型的完整輸出分佈進行訓練的電路,對於模型正確回答的提示達到 97.3-99.5% 的一致性,但對於錯誤的僅有 11.4-41.7%。一個 IOI 案例研究顯示,通過恢復省略的注意力頭來恢復遺失的錯誤,將錯誤再現率從 14.2% 提升至 75.1%,在一個獨立的保留集上,正確一致性下降了 0.41 個百分點,超過了匹配的隨機擴展和標量偏置控制。干預痕跡顯示,省略的計算如何為可重現的錯誤子集產生特定的錯誤答案。總的來說,這些發現表明,電路可以保持任務的成功,而未能充分解釋模型的失敗,並支持準確的錯誤再現作為基於電路的模型行為解釋的必要但不充分的測試。

Signatures of semantic search in the activations of large language models

2609.35599v2 by Luke Leckie, Peter M. Todd, Jacob G. Foster

When recalling lists of concepts (e.g., animals) during the semantic fluency task (SFT), both humans and large language models (LLMs) organise their output into clusters of related items (e.g., sea animals) that are punctuated by strategic switches between clusters. In humans, this pattern can be explained by a semantic foraging process, whereby distinct neural and behavioural signatures accompany within-cluster production ("exploit") and between-cluster switching ("explore"). Whether LLMs likewise represent these two search regimes within their internal states is unknown. Here, we apply a range of mechanistic interpretability techniques to provide evidence for this. In Study 1, we use the Jacobian lens (J-lens), which maps intermediate-layer residual-stream representations to token-level activations, to show that concept-level activations predict switching. First, we find that switching coincides with low next-token activations. Moreover, the probability of switching rises as the set of strongest J-lens activations (the J-space) becomes depleted of items from the category currently being produced, analogous to explore-exploit decision-making during patch foraging. We then show that middle-layer J-lens activations of abstract category-related labels (e.g., "water") increase in anticipation of switching into that category. We confirm these representations to causally influence switching by deriving steering vectors that target category switching. In Study 2, we identify generic residual stream directions that are activated during and in anticipation of switching. By steering activations along these directions, we bias increased or decreased rates of switching. Our study extends the semantic foraging framework to artificial intelligences and provides evidence that LLMs maintain distinct representational signatures for exploration and exploitation as they verbalise conceptual information.

摘要:在語意流暢性任務(SFT)中回想概念列表(例如,動物)時,人類和大型語言模型(LLMs)都將其輸出組織成相關項目的集群(例如,海洋動物),並在集群之間進行戰略性切換。在人類中,這種模式可以通過語意覓食過程來解釋,其中不同的神經和行為特徵伴隨著集群內的產出(「利用」)和集群之間的切換(「探索」)。LLMs是否同樣在其內部狀態中表示這兩種搜索模式尚不清楚。在這裡,我們應用一系列機械可解釋性技術來提供證據。在研究 1 中,我們使用雅可比透鏡(J-lens),該透鏡將中間層的殘差流表示映射到標記級別的激活,來顯示概念級別的激活預測切換。首先,我們發現切換與低的下一標記激活相吻合。此外,隨著最強 J-lens 激活集(J-space)中的當前產出類別項目的耗盡,切換的概率上升,類似於在斑塊覓食過程中的探索-利用決策。我們接著顯示,抽象類別相關標籤(例如,「水」)的中層 J-lens 激活在預期切換到該類別之前會增加。我們通過推導針對類別切換的引導向量來確認這些表示對切換的因果影響。在研究 2 中,我們識別在切換過程中及其預期期間被激活的通用殘差流方向。通過沿這些方向引導激活,我們偏向於增加或減少切換的頻率。我們的研究將語意覓食框架擴展到人工智慧,並提供證據表明 LLMs 在口頭表達概念信息時保持探索和利用的不同表徵特徵。

A decision-support system applied to Law: Reasoning and explainability of the decision

2609.35370v1 by Jeremy Bouche-Pillon, Pascale Zarat{é}, Yannick Chevalier, Nathalie Aussenac-Gilles

The emergence of the digital transition brought an increasing need to control the processing of digital information, including in Law Enforcement Agencies (LEAs). At the EU level, in recent years, many regulations have emerged to control data processing and exchange. Texts other than the GDPR, such as the ''Law Enforcement Directive (LED)'', appeared to regulate specifically how Law Enforcement Agencies (LEAs) could process data. A formal representation of these regulations can be part of decision systems that support LEAs in processing data in compliance with the regulations. Although many new formalisms have emerged to represent legal norms and rules, few are provided with a reasoning mechanism. Furthermore, systems used in decision-making processes in critical contexts such as medical diagnoses or legal decisions cannot be fully automated, and the explainability of their results is essential to ensure user confidence in decisions. This explainability aspect, while crucial, is lacking in most modern approaches that rely on machine learning. This paper describes a framework to operate formal rules from regulations, by focusing on explainability of the decision. After describing the general architecture of the proposed decision support framework, the paper showcases how symbolic AI and the SPARQL query language can support legal reasoning. It then describes an algorithm to generate a justification for the reasoning results, and outlines the procedure to be followed when the reasoning does not lead to a satisfactory conclusion. We notably focus on a method based on decision trees to determine what additional information to request from the user.

摘要:數位轉型的出現帶來了對數位資訊處理的日益需求,包括在執法機構(LEAs)中。在歐盟層面上,近年來出現了許多規範來控制數據處理和交換。除了GDPR之外,還出現了如“執法指令(LED)”等文本,專門規範執法機構(LEAs)如何處理數據。這些規範的正式表述可以成為支持執法機構在遵守規範的情況下處理數據的決策系統的一部分。儘管許多新的形式主義已經出現以表達法律規範和規則,但很少有配備推理機制的形式。此外,用於醫療診斷或法律決策等關鍵情境的決策過程中使用的系統不能完全自動化,其結果的可解釋性對於確保用戶對決策的信心至關重要。這一可解釋性方面雖然至關重要,但在大多數依賴機器學習的現代方法中卻缺乏。本文描述了一個運作規範形式規則的框架,重點在於決策的可解釋性。在描述所提議的決策支持框架的一般架構後,本文展示了符號人工智慧和SPARQL查詢語言如何支持法律推理。接著描述了一種生成推理結果的理由的算法,並概述了當推理未能導致令人滿意的結論時應遵循的程序。我們特別關注一種基於決策樹的方法,以確定需要向用戶請求的額外信息。

Jev thinks "I don't know'', but doesn't say it: Introducing Sys1Cal-v1 Dataset for Probability Calibration

2609.35342v1 by Riccardo Porcedda

The appearance of Jev marked the era of System One Models, foundation models that return structured decisions with probability distributions rather than text. Aside from low cost and great speed, Jev's central promise is that these probabilities are calibrated: such claim is not backed by any public test and available external benchmarks evaluate confidence calibration, not whether every returned option probability has the right numerical meaning. To tackle this issue, we introduce Sys1Cal-v1, a dataset of True/False questions about a proposition $A$ for which the exact probability $P(A)$ is known by construction. Each item is queried through the three Jev primitives - Noul, Choice and Score - and evaluated by total variation distance from the ground-truth distribution, which can be used to estimate a soft accuracy of System One Models. We showcase the utility of Sys1Cal-v1 as a benchmark dataset by evaluating Jev and SemIf, an open-source Choice-style baseline. In this work, however, we focus even more deeply on Jev, by studying the calibration of its Score and Choice answers. In particular, we discover a peculiar behaviour that can be explained by assuming that Jev suppresses a third truth value, going beyond True and False. In other words, in \texttt{Choice} answers, $P(A)$ and $P(\neg A)$ are presented as if $P(A)+P(\neg A)=1$, while a term $P(U)\neq0$ is missing in the sum. Recovering $P(U)$ leads to an improvement of median soft accuracy in \texttt{Choice} answers from $0.771$ to $0.978$, suggesting that, even in binary decisions, Jev wants to answer with a third option:``I don't know''.

摘要:Jev的出現標誌著系統一模型的時代,這些基礎模型返回結構化的決策,並以概率分佈而非文本形式呈現。除了低成本和高速度外,Jev的核心承諾是這些概率是經過校準的:這一說法並沒有任何公開測試的支持,且可用的外部基準評估的是置信度校準,而不是每個返回的選項概率是否具有正確的數值意義。為了應對這一問題,我們引入了Sys1Cal-v1,一個關於命題$A$的真/假問題數據集,該命題的確切概率$P(A)$是通過構造已知的。每個項目通過三個Jev原語 - Noul、Choice和Score進行查詢,並通過與真實分佈的總變異距離進行評估,這可以用來估計系統一模型的軟準確性。我們展示了Sys1Cal-v1作為基準數據集的實用性,通過評估Jev和SemIf,一個開源的選擇風格基準。然而,在這項工作中,我們更深入地專注於Jev,研究其Score和Choice答案的校準。特別是,我們發現了一種特殊的行為,可以通過假設Jev抑制第三個真值來解釋,這超越了真和假。換句話說,在\texttt{Choice}答案中,$P(A)$和$P(\neg A)$被呈現得好像$P(A)+P(\neg A)=1$,而在總和中缺少了一項$P(U)\neq0$。恢復$P(U)$使得\texttt{Choice}答案的中位數軟準確性從$0.771$提高到$0.978$,這表明,即使在二元決策中,Jev也希望以第三個選項回答:“我不知道”。

The Argument and the Letterhead: Source-Position Coherence in AI Evaluation

2609.35286v1 by Michele Loi

An argument can be surprising coming from a particular speaker without being a bad argument. Do AI evaluators keep these judgments apart? Two preregistered descriptive studies and a later Jev supplement collected 2,976 usable evaluations of six fixed texts about US AI policy, Germany's debt brake and Swiss nuclear energy. Each text was presented under several source attributions. The key comparison asks whether the gap between two sources changes when the argument changes. On Sol, for example, a national-security argument received mean ratings of 0.359 under CODEPINK and 0.639 under College Republicans; a civil-rights argument received 0.742 and 0.721. A constant preference for one source cannot explain that pattern. Related interactions appeared across topics and recent model configurations, including those with reasoning enabled, while several comparisons yielded small effects. The later European Jev supplement yielded five interactions below the adopted absolute reference of 0.05; its distinct rubric and interrupted collection limit comparison with the chat systems. Some written evaluations explicitly invoked a mismatch between a source and its attributed position. Taken together, the numerical and verbal evidence supports source-position coherence as a plausible explanation, alongside competing accounts involving credibility, authenticity and interpretation of the task. The paper develops this inference through controlled comparisons, reports conditional post hoc p-values in an appendix, and documents the human decisions and delegated checks behind an AI-conducted study.

摘要:一個論點來自特定發言者時可能會令人驚訝,但這並不意味著它是一個糟糕的論點。AI 評估者是否將這些判斷分開?兩項預註冊的描述性研究和後來的 Jev 補充收集了 2,976 個可用的評估,這些評估涉及六篇關於美國 AI 政策、德國的債務制動器和瑞士核能的固定文本。每篇文本在幾個來源歸屬下呈現。關鍵比較在於當論點改變時,兩個來源之間的差距是否會改變。例如,在 Sol 上,國家安全論點在 CODEPINK 下的平均評分為 0.359,而在大學共和黨人下為 0.639;公民權利論點的評分分別為 0.742 和 0.721。對於一個來源的持續偏好無法解釋這種模式。相關的互動出現在不同主題和最近的模型配置中,包括那些啟用推理的配置,而幾個比較則產生了小的效果。後來的歐洲 Jev 補充產生了五個低於採用的絕對參考值 0.05 的互動;其獨特的標準和中斷的收集限制了與聊天系統的比較。一些書面評估明確提到了來源與其歸屬位置之間的不匹配。綜合來看,數字和口頭證據支持來源位置一致性作為一個合理的解釋,並且還有涉及可信度、真實性和任務解釋的競爭說明。本文通過控制比較來發展這一推論,在附錄中報告條件後驗 p 值,並記錄 AI 進行研究背後的人類決策和委派檢查。

Towards Reliable AI Data Scientists: Data Agents with Workflow Harnesses

2609.35255v1 by Huachi Zhou, Yujing Zhang, Jiahe Du, Jiacheng Cai, Zijin Hong, Chuang Zhou, Zheng Yuan, Qinggang Zhang, Qing Li, Xiao Huang

Large language model agents are increasingly deployed for data-intensive work, yet reliable data analysis requires more than general-purpose reasoning and ad hoc tool augmentation. Data Agents, equipped with workflow harnesses, offer a promising paradigm for automating the end-to-end data science lifecycle. This paper examines Data Agents from a harness-centric perspective. First, we introduce a taxonomy of Data Agents and associated data environments, organizing the literature around five functional stages: perception, planning, execution, verification, and repair. Second, we analyze the key technical routes within each stage, identifying 15 distinct approaches ranging from data structure probing to data state reconstruction. Third, we identify four open reliability problems: inactive semantic calibration, missing clarification, missing experience transfer, and the missing verification-repair repository. These problems explain why silent failures can persist even when individual components function correctly, highlighting the need for rigorous workflow harnesses and shared reliability resources. Finally, we summarize the horizontal task families of Data Agents, examine their vertical application settings, and benchmarks for evaluation, while maintaining a companion repository at https://github.com/DEEP-PolyU/Awesome-Data-Agents.

摘要:大型語言模型代理越來越多地被用於數據密集型工作,但可靠的數據分析需要的不僅僅是通用推理和臨時工具增強。數據代理配備了工作流程鞍具,為自動化端到端數據科學生命周期提供了一種有前景的範式。本文從鞍具中心的角度檢視數據代理。首先,我們介紹了一個數據代理及其相關數據環境的分類法,將文獻組織為五個功能階段:感知、規劃、執行、驗證和修復。其次,我們分析了每個階段內的關鍵技術路徑,識別出15種不同的方法,範圍從數據結構探測到數據狀態重建。第三,我們確定了四個開放的可靠性問題:非活動的語義校準、缺失的澄清、缺失的經驗轉移以及缺失的驗證-修復庫。這些問題解釋了為什麼即使單個組件正常運作,靜默失敗仍然會持續存在,突顯了對嚴格工作流程鞍具和共享可靠性資源的需求。最後,我們總結了數據代理的橫向任務家族,檢視其縱向應用設置及評估基準,同時在 https://github.com/DEEP-PolyU/Awesome-Data-Agents 維護一個伴隨的庫。

Beneath the Tokens: A Performance Engineering Study of Multi-Token Prediction in GPU-Accelerated LLM Inference

2609.35188v1 by Suwesh Prasad Sah

Autoregressive large language model inference repeatedly invokes the target model to generate one token at a time, making generation sensitive to GPU memory movement and sequential execution. This study evaluates two-token multi-token prediction (MTP) against autoregressive decoding in a controlled single-request deployment on an NVIDIA A10G GPU. A 360-request benchmark covered plain-text, reasoning-intensive, and tool-calling workloads, while runtime telemetry, Nsight Systems, PyTorch Profiler, and selected Nsight Compute measurements were used to explain the observed performance. MTP increased output throughput by (1.91\times) to (2.19\times) across all prompts and reduced time to first output by 10.0--14.2\%. Median mean acceptance length ranged from 2.370 to 2.595 tokens per verification iteration. Profiling showed that MTP introduced a longer and more complex execution path, including proposal, sampling, attention, gathering, and reduction operations. However, it required 56.4--78.1\% fewer executions of the selected repeating CUDA Graph per generated token. The dominant MTP GEMM kernel was not faster than the dominant autoregressive GEMV kernel, and selected instances of both approached the A10G memory-bandwidth limit. These results show that MTP improved inference through amortization: greater token progress reduced repeated GPU execution sufficiently to outweigh the additional speculative-execution cost.

摘要:自回歸大型語言模型推理重複調用目標模型以一次生成一個標記,這使得生成對 GPU 記憶體移動和順序執行非常敏感。本研究在 NVIDIA A10G GPU 上的受控單請求部署中評估了兩標記多標記預測(MTP)與自回歸解碼的表現。一個 360 請求的基準涵蓋了純文本、推理密集型和工具調用工作負載,同時使用運行時遙測、Nsight Systems、PyTorch Profiler 和選定的 Nsight Compute 測量來解釋觀察到的性能。MTP 在所有提示中將輸出吞吐量提高了 (1.91\times) 到 (2.19\times),並將首次輸出的時間減少了 10.0--14.2\%。中位數平均接受長度在每次驗證迭代中範圍為 2.370 到 2.595 個標記。分析顯示,MTP 引入了更長且更複雜的執行路徑,包括提案、抽樣、注意力、收集和減少操作。然而,它每生成一個標記所需的選定重複 CUDA Graph 的執行次數減少了 56.4--78.1\%。主導的 MTP GEMM 核心並不比主導的自回歸 GEMV 核心更快,且兩者的選定實例均接近 A10G 記憶體帶寬限制。這些結果顯示,MTP 通過攤銷改善了推理:更大的標記進展足以減少重複的 GPU 執行,以抵消額外的推測執行成本。

2609.34780v2 by Erik Aerts

The use and applicability of artificial intelligence (AI) in medical research and clinical practice has received increasing attention in the literature over recent years. The emergence of large language models (LLMs) has expanded discussions in regards to applications of AI within healthcare. While traditional deep learning based AI applications in medicine have often focused on specific and defined tasks, LLMs offer broader capabilities and flexibility in working with available data,. At the same time of writing, the integration of LLMs into medical settings raises important questions regarding their reliability, accuracy, transparency, safety, and appropriate role in a medical setting. This text presents and discusses recent talks and articles concerning the application of LLMs in medicine, with particular emphasis on their potential utility in research and clinical practice. It considers both the opportunities offered by these technologies and the challenges associated with their implementation, aiming to provide a perspective on the current and emerging role of LLMs within the medical field.

摘要:人工智慧(AI)在醫學研究和臨床實踐中的使用和適用性在近年來的文獻中受到越來越多的關注。大型語言模型(LLMs)的出現擴大了關於AI在醫療保健中應用的討論。雖然傳統基於深度學習的AI應用在醫學中往往專注於特定和明確的任務,但LLMs在處理可用數據方面提供了更廣泛的能力和靈活性。撰寫本文的同時,將LLMs整合進醫療環境中引發了有關其可靠性、準確性、透明度、安全性和在醫療環境中適當角色的重要問題。本文呈現並討論了有關LLMs在醫學中應用的最近演講和文章,特別強調它們在研究和臨床實踐中的潛在效用。它考慮了這些技術所提供的機會以及與其實施相關的挑戰,旨在提供對LLMs在醫療領域中當前和新興角色的看法。

A Unifying Framework of Concept-based Explainable AI with Completeness Guarantees

2609.34750v1 by Vojtěch Kůr, Adam Kukučka, Tomáš Brázdil, Vít Musil

Concept-based explanations describe neural network predictions through human-understandable properties of inputs called concepts. The field encompasses approaches that differ in how they define and represent concepts and connect them to model predictions. We introduce a theoretical framework that describes these approaches in a common mathematical language and supports a shared analysis of their properties. For concept discovery, which identifies concepts automatically within a latent space of a trained model, we employ a concept autoencoder view. An encoder extracts concept representations from the model's latent space, and a decoder uses them to reconstruct the original latent representation. The autoencoder's reconstruction error measures how accurately its decoder recovers the original latent representation. We revisit model completeness: how well the concepts can reproduce the model's outputs. We show that model incompleteness of the concepts can be bounded by the autoencoder's reconstruction error. The autoencoder view also provides a common way to define individual concept attributions, which measure each concept's contribution to a prediction. We establish when these attributions sum to the model's prediction, and bound the discrepancy otherwise, thus providing attribution completeness guarantees.

摘要:概念基礎的解釋通過稱為概念的輸入的人類可理解特性來描述神經網絡的預測。這個領域涵蓋了在定義和表示概念以及將其與模型預測連接方面有所不同的方法。我們引入了一個理論框架,該框架用共同的數學語言描述這些方法,並支持對其特性的共享分析。對於概念發現,即在訓練模型的潛在空間中自動識別概念,我們採用了概念自編碼器的視角。編碼器從模型的潛在空間中提取概念表示,解碼器則利用這些表示重建原始的潛在表示。自編碼器的重建誤差衡量其解碼器恢復原始潛在表示的準確性。我們重新審視模型的完整性:這些概念能多好地再現模型的輸出。我們展示了概念的模型不完整性可以被自編碼器的重建誤差所界定。自編碼器的視角還提供了一種共同的方式來定義個別概念的歸因,這些歸因衡量每個概念對預測的貢獻。我們確立了這些歸因何時加總為模型的預測,並在其他情況下界定了差異,從而提供了歸因完整性的保證。

From Human Narrative to Harmonic Structure: A Human-Centered Investigation of Algorithmic Music Generation through the Chord Wheel Diagram

2609.34735v1 by Josef Pavlíček, Petra Pavlíčková, Irena Štrausová

Contemporary AI-based music generation can produce compositions that satisfy formal requirements of tonality and musical coherence. However, whether musical expression can be described by mathematical properties alone remains a fundamental question. Human composers operate within personal and cultural contexts that influence harmonic decisions and deliberate departures from established patterns. This study investigates six narrative-driven popular songs by Bob Dylan, Johnny Cash, and Ritchie Valens. Original human harmonies are compared with outputs of an explainable computational harmonizer operating on the same melodies without access to the original chord progressions. We examine harmonic vocabulary, functional persistence, repetition, non-diatonic events, and tension-resolution patterns using Chord Wheel Diagrams and BPMN-based representations. Results show that high melody-chord compatibility does not necessarily imply preservation of the original human harmonic decision pattern. Some generated harmonizations retain the economical structure of the reference, while others alter harmonic diversity or suppress distinctive events while remaining compatible with the melody. Rather than quantifying artistic quality, the study introduces narrative-conditioned harmonic structure as a complementary perspective for computational music analysis. The findings suggest that generative systems may benefit from modeling not only harmonic correctness, but also structural identity, context, and human compositional intention.

摘要:當代基於人工智慧的音樂生成可以創作滿足音調和音樂一致性正式要求的作品。然而,音樂表達是否僅能用數學特性來描述仍然是一個根本問題。人類作曲家在個人和文化背景中運作,這些背景影響和諧決策以及故意偏離既定模式的選擇。本研究調查了六首由Bob Dylan、Johnny Cash和Ritchie Valens創作的敘事驅動流行歌曲。原始的人類和聲與一個可解釋的計算和聲器的輸出進行比較,該計算和聲器在沒有訪問原始和弦進行的情況下對相同旋律進行操作。我們使用和弦輪圖和基於BPMN的表示法來檢查和聲詞彙、功能持續性、重複、非調性事件和緊張-解決模式。結果顯示,高旋律與和弦的相容性並不一定意味著保留原始人類和聲決策模式。一些生成的和聲保留了參考的經濟結構,而其他則改變了和聲的多樣性或壓制了獨特事件,同時仍與旋律相容。本研究並非量化藝術品質,而是引入敘事條件的和聲結構作為計算音樂分析的補充視角。研究結果表明,生成系統可能受益於不僅建模和聲正確性,還包括結構身份、上下文和人類創作意圖。

Understanding Generalization Requires Universal Induction

2609.34458v1 by Aram Ebtekar, Marcus Hutter, Danica J. Sutherland

Classical statistical theory is insufficient to explain the successes of general-purpose AI models, because it depends on handcrafted inductive biases that it cannot justify. No Free Lunch (NFL) theorems force any learner that beats chance on some environments to underperform on others. We might hope that past experience informs which environments to expect, but NFL applies equally to meta-learning. Thus, any method that makes meaningful predictions necessarily begins with an inductive bias external to the data. Choosing to bias toward short programs yields Solomonoff induction (SI), whose performance is competitive against all computable learners - albeit up to "constants" that become large when comparing against specialized methods that exploit background information. We therefore relativize SI to an information vantage point, biasing toward short programs with access to all preexisting information. This reframes the inductive bias: instead of seeking some absolute notion of simplicity, we favor accessibility with respect to our vantage point. An algorithm can only outpredict the relativized SI to the extent that its code contains additional information about the data, and no algorithm can generate such information. While SI is incomputable and hence not a practical algorithm, it provides a formal optimum for inference in the limit of infinite compute, and there is evidence to suggest that frontier AI systems roughly approximate it. Thus, the only known answer to meta-NFL is rooted in algorithmic information theory, which we should expect to play a fundamental role in explaining the generalization behavior of modern (and future) AI systems.

摘要:古典統計理論不足以解釋通用人工智慧模型的成功,因為它依賴於無法證明的手工製作的歸納偏見。無免費午餐(NFL)定理迫使任何在某些環境中超越隨機的學習者在其他環境中表現不佳。我們可能希望過去的經驗能告訴我們預期哪些環境,但NFL同樣適用於元學習。因此,任何能做出有意義預測的方法必然以一種外部於數據的歸納偏見開始。選擇偏向短程序會產生所羅門諾夫歸納(SI),其性能在所有可計算學習者中具有競爭力——儘管在與利用背景信息的專門方法比較時,這些“常數”會變得很大。因此,我們將SI相對化到一個信息視角,偏向於短程序並訪問所有現有信息。這重新框定了歸納偏見:我們不再尋求某種絕對的簡單性概念,而是更重視相對於我們視角的可及性。一個算法只能在其代碼包含有關數據的額外信息的程度上超越相對化的SI,且沒有任何算法能生成這種信息。雖然SI是不可計算的,因此不是一個實用的算法,但它為在無限計算的極限下的推理提供了一個形式上的最優解,並且有證據表明前沿人工智慧系統大致上近似於它。因此,對於元NFL唯一已知的答案根植於算法信息理論,我們應該預期它在解釋現代(和未來)人工智慧系統的泛化行為中扮演基本角色。

Social Circuits behind Multi-agent Echo Chambers

2609.34444v1 by Chuiyang Meng, Wenlu Yu, Ming Tang, Cheng Li

Language-model agents exchange messages to combine evidence, but their communication can also create echo chambers that reinforce shared errors. However, overall task performance does not explain how a message changes the receiving agent's internal activations and affects its decision. In this work, we introduce Social Circuits, a framework for tracing message effects through receiver activations. We compare the receiver's answers before and after changing a message. Then, we restore selected activations recorded under the original message to determine how much of the message effect these activations reproduce. Based on Social Circuits, we propose Circuit-Guided Deliberation (CGD), which learns to select useful messages using receiver activation changes. We establish when activation replacement preserves receiver decisions and bound the gap between CGD's task performance and the best achievable through message selection. Experiments show that receiver activation changes explain the message effects and guide message selection that improves the task performance. Across three models and four datasets, CGD achieves the highest or joint-highest average accuracy in our main comparisons while generating fewer tokens than multi-agent baselines.

摘要:語言模型代理之間交換訊息以結合證據,但它們的溝通也可能創造回音室,強化共同的錯誤。然而,整體任務表現並不能解釋一條訊息如何改變接收代理的內部激活並影響其決策。在這項工作中,我們引入社會電路(Social Circuits),這是一個追蹤訊息影響的框架,通過接收者的激活來進行分析。我們比較接收者在改變訊息前後的回答。然後,我們恢復在原始訊息下記錄的選定激活,以確定這些激活重現了訊息效果的多少。基於社會電路,我們提出了電路引導的深思(Circuit-Guided Deliberation, CGD),它學會利用接收者激活變化來選擇有用的訊息。我們確定何時激活替換能夠保留接收者的決策,並界定CGD的任務表現與通過訊息選擇所能達到的最佳表現之間的差距。實驗顯示,接收者激活變化解釋了訊息效果,並指導訊息選擇以改善任務表現。在三個模型和四個數據集上,CGD在我們的主要比較中達到了最高或並列最高的平均準確率,同時生成的標記數量少於多代理基準。

Improving Large Language Models for Code through Runtime Program-State Reasoning

2609.34359v1 by Hongwei Li, Spandan Garg, Yufan Huang

Large language models receive limited explicit training in reasoning about runtime program states. We study whether training models to reason about runtime program states improves downstream software-engineering capabilities. We introduce two complementary program-state reasoning tasks. Buggy input-output reasoning requires a model to generate a concrete input that exposes a behavioral difference between a buggy program and a hidden correct implementation and to predict the resulting execution behavior. Precondition-postcondition reasoning requires an agent to symbolically characterize a bug-triggering precondition, predict the expected postcondition, explain their causal connection, and instantiate this reasoning as an executable regression test. By incorporating these two tasks into a staged post-training pipeline, we develop Comet-9B, a 9B language model based on Qwen3.5-9B Base. We evaluate the resulting checkpoints on repository-level patch generation, regression-test generation, and security PoC generation. Adding both program-state reasoning tasks to supervised fine-tuning (SFT) on issue resolution improves success rates by 7.25 percentage points on SWE-bench Pro and 9.70 points on SWT-Bench Verified. Sequential reinforcement learning on the two tasks yields further gains of 7.25, 26.79, and 4.67 percentage points on SWE-bench Pro, SWT-Bench Verified, and CyberGym, respectively. Despite having only 9B parameters, Comet-9B achieves a score comparable to the reported GPT-5.2 result on SWE-bench Pro and matches the reported success rate of a GPT-4o-based agent on SWT-Bench Verified.

摘要:大型語言模型在推理運行時程序狀態方面接受的明確訓練有限。我們研究訓練模型推理運行時程序狀態是否能改善下游軟體工程能力。我們引入了兩個互補的程序狀態推理任務。有缺陷的輸入輸出推理要求模型生成一個具體的輸入,該輸入能揭示有缺陷的程序和隱藏的正確實現之間的行為差異,並預測結果執行行為。前置條件-後置條件推理要求代理符號化地描述觸發錯誤的前置條件,預測預期的後置條件,解釋它們之間的因果關係,並將這一推理實例化為可執行的回歸測試。通過將這兩個任務納入分階段的後訓練流程,我們開發了 Comet-9B,一個基於 Qwen3.5-9B Base 的 9B 語言模型。我們在庫級補丁生成、回歸測試生成和安全 PoC 生成上評估了結果檢查點。將這兩個程序狀態推理任務添加到針對問題解決的監督微調(SFT)中,使成功率在 SWE-bench Pro 上提高了 7.25 個百分點,在 SWT-Bench Verified 上提高了 9.70 個百分點。在這兩個任務上進行的序列強化學習分別在 SWE-bench Pro、SWT-Bench Verified 和 CyberGym 上獲得了進一步的增益,分別為 7.25、26.79 和 4.67 個百分點。儘管只有 9B 參數,Comet-9B 在 SWE-bench Pro 上達到了與報告的 GPT-5.2 結果相當的分數,並且與基於 GPT-4o 的代理在 SWT-Bench Verified 上的報告成功率相匹配。

Dynamical Parameters: An Interpretability Framework for Time-Series Foundation Models

2609.34316v1 by Kang Yang, Gaofeng Dong, Liying Han, Mani Srivastava

This work studies a central gap in interpreting time-series foundation models (TSFMs): a dynamical property may be accessible in a hidden state even when the forecast fails to respond correctly as that property changes. We formalize these properties as Dynamical Parameters, including trend slope, oscillation frequency, and autoregressive dependence. We compare their representation accessibility, measured by recovery from hidden states, with their forecast response, measured by agreement with the expected forecast change. Across nine frozen TSFMs and thirteen laws, 42 of 63 model-parameter cells achieve accessibility above 0.95, whereas their median reference-aligned response relative to the conditional reference is only 0.46. To explain this gap, causal geometry compares the hidden-state change required to produce the reference response with the change induced by the parameter intervention. Directly modifying the hidden state recovers the reference response, but the parameter intervention often moves the state in a different direction. These results show that accessible parameter information need not be expressed in forecasts when input changes miss the required hidden-state direction.

摘要:這項工作研究了解釋時間序列基礎模型(TSFMs)中的一個核心差距:即使預測未能正確響應該特性變化,動態性質也可能在隱藏狀態中可訪問。我們將這些性質形式化為動態參數,包括趨勢斜率、振盪頻率和自回歸依賴性。我們比較它們的表示可訪問性,通過從隱藏狀態的恢復來測量,與它們的預測響應進行比較,後者通過與預期預測變化的一致性來測量。在九個凍結的TSFMs和十三條法則中,63個模型參數單元中有42個達到0.95以上的可訪問性,而它們相對於條件參考的中位數參考對齊響應僅為0.46。為了解釋這一差距,因果幾何比較了產生參考響應所需的隱藏狀態變化與參數干預所引起的變化。直接修改隱藏狀態可以恢復參考響應,但參數干預往往會將狀態移動到不同的方向。這些結果表明,當輸入變化錯過所需的隱藏狀態方向時,可訪問的參數信息不必在預測中表達。

Evo2Team: When Do Evolved Skills Transfer? From Selection to Deployment

2609.34135v1 by Renxiang Wang, Jiaming Cui

A skill bank that helps one multi-agent system may leave another's behavior unchanged. A transferred rule helps only when target agents act on it successfully. We study this path for routing and communication skills in Count-Frequency and AgentsNet, using teams of 4--32 agents and GPT and Qwen model ladders. Source evolution meets a joint quality, cost, model-tier, and confirmation goal in 14 of 16 settings. We then evaluate Evo2Team, which selects, adapts, and confirms source skills for the target team, alongside six frozen selectors across 28 transfer directions. Evo2Team's target-side exploration cost is below that of evolving a new target bank in every direction, even when reused reference evaluations are charged once. Twenty of 28 held-out outcomes meet the positive-transfer criterion, including three saved diagnostic tests. Selection alone does not explain these outcomes: KNN and CORAL choose different banks in two AgentsNet directions but produce identical recorded executions. When Evo2Team changes execution, gains can reach many tasks, as in a Count-Frequency direction that improves 28 of 32 tasks over KNN. Seven positive AgentsNet outcomes save 6.1--14.6\% in deployment cost while using transferred skills on only three to six of fifteen tasks. In five earlier accepted directions, all 22 task records using transferred skills pass three fixed-graph confirmations, but four fail in recorded executions on new graphs. Graphs and model responses change together in this comparison. These results show that skill transfer must be assessed through the actions agents take, the tasks those actions reach, and the quality and cost of the final deployment.

摘要:一個幫助某個多代理系統的技能庫可能不會改變另一個系統的行為。當目標代理成功地執行轉移的規則時,這個規則才會有幫助。我們研究了在 Count-Frequency 和 AgentsNet 中的路由和通信技能,使用 4 至 32 個代理的團隊以及 GPT 和 Qwen 模型梯度。在 16 個設置中,有 14 個達成了源演化的聯合質量、成本、模型層級和確認目標。我們接著評估 Evo2Team,該系統為目標團隊選擇、調整和確認源技能,並在 28 個轉移方向中使用六個固定選擇器。Evo2Team 的目標側探索成本在每個方向上都低於演化一個新的目標庫的成本,即使重用的參考評估只收費一次。28 個保留結果中有 20 個符合正向轉移標準,包括三個保存的診斷測試。僅僅依賴選擇無法解釋這些結果:KNN 和 CORAL 在兩個 AgentsNet 方向中選擇不同的庫,但產生相同的記錄執行。當 Evo2Team 改變執行時,收益可以達到許多任務,例如在一個 Count-Frequency 方向上,KNN 的 32 個任務中有 28 個得到了改善。七個正向的 AgentsNet 結果在只使用轉移技能的十五個任務中的三到六個任務上節省了 6.1% 到 14.6% 的部署成本。在五個早期接受的方向中,所有 22 個使用轉移技能的任務記錄都通過了三個固定圖形的確認,但在新圖形上的記錄執行中有四個失敗。在這次比較中,圖形和模型反應是一起改變的。這些結果顯示,技能轉移必須通過代理所採取的行動、這些行動所達到的任務以及最終部署的質量和成本來進行評估。

JET: Judge-Guided Evolution at Test Time for Agent Programs

2609.34126v1 by Yao Long Teng, Jiayi Cai, Bo An

An agent's executable program governs how it uses tools, processes observations, and responds to failures. Evolving this program at test time can help adaptation, but deciding which changes to retain is difficult when true rewards are unavailable. Execution traces provide evidence of agent behavior, yet interpreting that evidence requires a judge that remains useful as tasks and candidate programs change. We introduce Judge-Guided Evolution at Test Time (JET), which evolves an executable judge on labeled source trajectories, then freezes and transfers it to guide target-side program evolution. The judge supplies scores and diagnostic feedback without target evaluator access or model-weight updates. On unseen WebShop tasks, JET achieves approximately 13% higher mean reward than fixed-rubric guidance when evolution begins from an unevolved program (cold start) and 4% higher when it begins from one already optimized on source tasks (warm start), with a 36% relative improvement in cold-start exact success. An exact-judge control on PushT, where the judge reconstructs the scoring rule from observations, shows that without judge error, program search becomes the bottleneck. Analyses identify useful reward-prediction logic in the evolved code and show that better final selection alone cannot explain the gains. These results support executable judge transfer for program adaptation under evaluator-preserving task shifts.

摘要:一個代理的可執行程序決定了它如何使用工具、處理觀察結果以及對失敗的反應。 在測試時進化這個程序可以幫助適應,但在真正的獎勵不可用時,決定保留哪些變更是困難的。 執行痕跡提供了代理行為的證據,但解釋這些證據需要一個在任務和候選程序變化時仍然有用的評判者。 我們介紹了測試時的評判者引導進化(JET),它在標記的源軌跡上進化一個可執行的評判者,然後將其凍結並轉移到目標端程序進化的指導。 評判者提供分數和診斷反饋,而不需要目標評估者的訪問或模型權重的更新。 在未見過的WebShop任務上,當進化從未進化的程序(冷啟動)開始時,JET的平均獎勵比固定標準指導高出約13%,而當它從已在源任務上優化的程序(熱啟動)開始時,高出4%,冷啟動的精確成功率提高了36%。 在PushT上的精確評判者控制中,評判者從觀察中重建評分規則,顯示在沒有評判者錯誤的情況下,程序搜索成為瓶頸。 分析確定了進化代碼中有用的獎勵預測邏輯,並顯示僅僅更好的最終選擇無法解釋增益。 這些結果支持在保留評估者的任務變化下進行可執行評判者轉移以適應程序。

Do World Models Learn Global Understanding?

2609.34058v1 by Alexander Detkov, Matt Thomson

AI systems often feel brittle and fragmented. A large language model (LLM) may correctly explain a concept but fail to apply it, or follow safety instructions in one context but not another. This behavior suggests a general failure to lift local information to a global understanding. To gain fundamental insight, we frame "understanding" as learning constraints and propagating their consequences. We construct learning tasks on monoid worlds, sets of states connected by action transitions, where observed training transitions and an unseen constraint jointly determine held-out transitions. Measuring generalization tests whether models can learn global constraints from local transitions and propagate their consequences. We consider inverse, commutativity, composition, and periodicity constraints relevant to spatial and semantic structure. Across attention, recurrent, and state-space architectures, next-state training fits the data but fails to propagate non-trivial constraints. Compositional training, which uses identical paths but hides intermediate states from the input, achieves 96% accuracy on inverse, commutativity, and composition constraints across architectures, yields corresponding improvements in geometric generalization of world models trained on embodied environments and relational generalization in Wikidata-finetuned LLMs. How far do models propagate constraints when inferring an unseen fact may depend on first inferring others? We define proof depth d of a held-out transition, measuring the minimum number of inference rounds to infer the transition, and find that model generalization decreases sharply with proof depth. Increasing compositional path length T improves generalization. These results provide a formal way to investigate global understanding in language and world models and demonstrate that compositional training promotes information propagation and integration.

摘要:AI 系統經常感覺脆弱且支離破碎。一個大型語言模型 (LLM) 可能正確解釋一個概念,但無法應用它,或者在一個情境中遵循安全指示,但在另一個情境中卻不然。這種行為表明在將局部信息提升到全球理解方面存在普遍失敗。為了獲得基本見解,我們將「理解」框架化為學習約束並傳播其後果。我們在單元世界上構建學習任務,這些世界是由行動轉換相連的狀態集合,其中觀察到的訓練轉換和未見的約束共同決定了保留的轉換。測量泛化測試模型是否能從局部轉換中學習全球約束並傳播其後果。我們考慮與空間和語義結構相關的逆、交換性、組合性和周期性約束。在注意力、遞歸和狀態空間架構中,下一狀態訓練適合數據,但未能傳播非平凡約束。組合訓練使用相同的路徑,但將中間狀態從輸入中隱藏,在各架構上對逆、交換性和組合性約束達到 96% 的準確率,並在基於具體環境訓練的世界模型中產生相應的幾何泛化改善,以及在經過 Wikidata 微調的 LLM 中的關係泛化。當推斷一個未見的事實可能依賴於首先推斷其他事實時,模型傳播約束的程度有多遠?我們定義保留轉換的證明深度 d,測量推斷該轉換所需的最小推理輪次,並發現模型的泛化隨著證明深度的增加而急劇下降。增加組合路徑長度 T 改善了泛化。這些結果提供了一種正式的方法來研究語言和世界模型中的全球理解,並證明組合訓練促進信息的傳播和整合。

Uncovering shortcut learning in audio classifiers by discovering recurring concepts in temporal explanations

2609.34030v1 by Cecilia Bolaños, Luciana Ferrer, Magdalena Fuentes

Correlations between events in machine learning datasets may result in shortcut learning, where models learn to predict the target event based on the presence of a correlated event. When these correlations are spurious -- arising from data collection artifacts -- models are likely to perform poorly in practice. We propose a pipeline to uncover shortcut learning in audio classifiers by discovering recurring concepts in their temporal explanations. Specifically, we isolate audio segments that explain classifier decisions, caption them with an ensemble of Large Audio-Language Models, and use a Large Language Model to extract recurring concepts. The resulting concepts can be audited by humans to uncover potential shortcut learning. We evaluate our framework using datasets curated from AudioSet Strong, controlling for the presence or absence of spurious correlations. Results show that this approach reliably uncovers learned shortcuts, such as the model relying on the presence of "laughter" to predict "applause".

摘要:事件之間的相關性在機器學習數據集中可能導致捷徑學習,模型學會根據相關事件的存在來預測目標事件。當這些相關性是虛假的——源於數據收集的工藝問題——模型在實際應用中可能表現不佳。我們提出了一個管道來揭示音頻分類器中的捷徑學習,通過發現它們時間解釋中的重複概念。具體而言,我們隔離解釋分類器決策的音頻片段,使用一組大型音頻-語言模型為其標題,並利用大型語言模型提取重複概念。所得到的概念可以由人類進行審核,以揭示潛在的捷徑學習。我們使用從AudioSet Strong整理的數據集來評估我們的框架,控制虛假相關性的存在或缺失。結果顯示,這種方法可靠地揭示了學習到的捷徑,例如模型依賴「笑聲」的存在來預測「掌聲」。

Safety Reconstructed: Generative Modeling via Masked Diffusion Builds Strong Safety Guardrails

2609.33634v1 by Gert Lek, Abele Malan, Chaoyi Zhu, Pin-Yu Chen, Robert Birke, Lydia Chen

Guard models are the last line of defense between a language model and a harmful output, yet their training objective is surprisingly narrow. Existing guards learn to predict a single verdict token from a conversational context, concentrating supervision on a single target. The consequences are structural: models latch onto shortcut features, are overconfident, and remain sensitive to where safety evidence appears in the sequence rather than its role in the full context. We propose a different framing. Rather than predicting a label from text, our LLaDA-Guard asks which label better explains the text: scoring the prompt or response under each label hypothesis and classifying based on their difference. This shifts supervision to every token in the moderated region, forcing the model to account for full content rather than its most discriminative fragments. We instantiate this idea with a masked diffusion language model, fine-tuning LLaDA-8B-Instruct with a class-conditional reconstruction objective using LoRA and requiring no architectural changes beyond the base model. LLaDA-Guard leads on average rank against discriminative baselines trained on stronger backbones across seven held-out safety benchmarks, while exhibiting substantially better confidence calibration (ECE 0.0875 vs. 0.1384 for Qwen3Guard), less over-defense on benign prompts with unsafe-looking cues, and less prompt leakage when moderating responses. Its generative nature further enables token-level risk localization as a natural byproduct, yielding a pipeline for rewriting unsafe prompts into safe equivalents without additional training and achieving a 60.7% average conversion-to-safe rate.

摘要:守衛模型是語言模型與有害輸出之間的最後防線,但它們的訓練目標卻出乎意料地狹窄。現有的守衛學習從對話上下文中預測單一的判決標記,將監督集中在單一目標上。這帶來了結構性的後果:模型依賴於捷徑特徵,過於自信,並對安全證據在序列中出現的位置保持敏感,而不是其在完整上下文中的角色。我們提出了一種不同的框架。我們的LLaDA-Guard不是從文本中預測標籤,而是詢問哪個標籤更好地解釋文本:在每個標籤假設下對提示或回應進行打分,並根據它們的差異進行分類。這將監督轉移到被調節區域中的每個標記,迫使模型考慮完整內容,而不是其最具區分性的片段。我們用一個掩蔽擴散語言模型來實現這個想法,通過使用LoRA的類條件重建目標來微調LLaDA-8B-Instruct,並且不需要超出基礎模型的架構變更。LLaDA-Guard在七個保留的安全基準中,對比在更強的基礎模型上訓練的區分基準,平均排名領先,同時顯示出顯著更好的信心校準(ECE 0.0875對比Qwen3Guard的0.1384),在具有不安全外觀提示的良性提示上過度防禦較少,並在調節回應時提示洩漏較少。其生成特性進一步使得標記級風險定位成為一種自然副產品,產生了一個將不安全提示重寫為安全等價物的管道,無需額外訓練,並實現了60.7%的平均轉換為安全率。

Learning Dynamics of Continual Learning: A Unified View of Data Attribution, Forgetting, and Plasticity Loss

2609.33620v1 by Yi Ren, Wenlong Deng, Guanzhe Hong, Clare Lyle, Yarin Gal

Modern language models are likely to be updated throughout their lifetime rather than trained once and frozen. Each update therefore participates in a recurring cycle: decide which experience to learn from, understand what that update changes, and remain capable of learning from what comes next. We show that these challenges are governed by the same evolving update--behavior interaction. We derive a token- and layer-wise decomposition of how learning from one token changes another prediction. By separating the softmax force, shared readout geometry, and residual connections, it exposes two interaction channels and yields a forward-computable approximation. Following this interaction through time reveals a unified picture of continual adaptation. Positive interaction identifies useful experience; negative interaction produces either concentrated collision or accumulated erosion; over longer horizons, updates reshape the shared geometry mediating future learning signals, reducing their transmission. These predictions lead to effective data selection, mechanism-specific controls for interference, and a readout-based diagnostic of future learnability whose degradation predicts the benefit of restoring the readout. Across models and training regimes, the same local interaction thus explains both what an update changes now and how learning today changes what can be learned tomorrow. This view connects data attribution, forgetting, and plasticity loss as distinct regimes of the same evolving learning dynamics.

摘要:現代語言模型在其生命週期內可能會不斷更新,而不是一次訓練後就凍結。因此,每次更新都參與一個重複的循環:決定從哪種經驗中學習,理解這次更新改變了什麼,並保持能夠從接下來的經驗中學習。我們顯示這些挑戰受到相同演變的更新—行為互動的支配。我們推導出一種基於標記和層的分解,說明從一個標記學習如何改變另一個預測。通過分離softmax力、共享讀取幾何和殘差連接,它揭示了兩個互動通道並產生了一個可前向計算的近似。隨著時間推移,跟隨這種互動揭示了持續適應的統一圖景。正向互動識別有用的經驗;負向互動則產生集中碰撞或累積侵蝕;在更長的時間範圍內,更新重塑了調解未來學習信號的共享幾何,減少了其傳輸。這些預測導致有效的數據選擇、特定機制的干擾控制,以及基於讀取的未來可學習性的診斷,其退化預測了恢復讀取的好處。在模型和訓練體系中,相同的局部互動因此解釋了更新當前改變了什麼,以及今天的學習如何改變明天可以學習的內容。這一觀點將數據歸因、遺忘和可塑性喪失連接為同一演變學習動態的不同範疇。

Temporal Graph Learning of Wearable Actigraphy and Sleep Traces for Modelling Adolescent Crystallized Intelligence

2609.33428v1 by Md. Tanvir Rahman, Nabil Anan Orka, Asaduzzaman Khan, Mohammad Ali Moni

Wearable actigraphy offers a scalable, ecologically valid alternative to episodic clinical assessment. However, predicting continuous adolescent crystallized intelligence ($G_c$) from such traces remains challenging due to irregular device adherence and complex behavioral-environmental interactions. We address this using daily summary data derived from 21-day Fitbit records of 6,091 adolescents in the Adolescent Brain Cognitive Development Study (Release 5.1). We propose SATURN, a Sleep-Activity Temporal Unified Regression Network. It represents participants as 21-node temporal graphs encoding daily behaviors and temporal adjacency. To prevent imputation artifacts, invalid-day edges are dynamically pruned during forward passes. Node embeddings are refined via residual GATv2 layers, aggregated through masked attention pooling, and fused with sociodemographic covariates. Under family-controlled, age-sex-BMI-stratified cross-validation, SATURN achieves $R^2 = 0.2783 \pm 0.0127$, consistently improving upon flattened machine learning (Gradient Boosting, $R^2 = 0.2372$) and sequential deep learning (BiLSTM, $R^2 = 0.2688$) baselines. Explainability analyses identify light activity, metabolic equivalents, and sleep duration as dominant predictors, while Monte Carlo dropout and subgroup analyses confirm equitable performance across sociodemographic strata. Ultimately, SATURN establishes a rigorous computational framework for digital cognitive phenotyping, offering a scalable pathway to complement traditional assessments by highlighting macro-level behavioral anomalies.

摘要:可穿戴行為測量提供了一種可擴展的、生態有效的替代方案,以取代臨床評估的偶發性。然而,從這些數據中預測持續的青少年結晶智力 ($G_c$) 仍然具有挑戰性,因為設備遵從性不規則且行為與環境之間的互動複雜。我們使用來自 6,091 名青少年在青少年大腦認知發展研究(版本 5.1)中,為期 21 天的 Fitbit 記錄所衍生的每日摘要數據來解決這個問題。我們提出了 SATURN,一個睡眠-活動時間統一回歸網絡。它將參與者表示為 21 節點的時間圖,編碼每日行為和時間相鄰性。為了防止插補伪影,在前向傳播過程中動態修剪無效日邊緣。節點嵌入通過殘差 GATv2 層進行精煉,通過遮罩注意力池化進行聚合,並與社會人口學協變量融合。在家庭控制、年齡-性別-BMI 分層的交叉驗證下,SATURN 的 $R^2 = 0.2783 \pm 0.0127$,持續優於扁平化的機器學習(梯度提升,$R^2 = 0.2372$)和序列深度學習(BiLSTM,$R^2 = 0.2688$)基準。可解釋性分析確定輕度活動、代謝當量和睡眠持續時間為主要預測因子,而蒙特卡羅隨機失活和子群分析則確認了在社會人口學層次上表現公平。最終,SATURN 建立了一個嚴謹的計算框架,用於數位認知表型,提供了一條可擴展的途徑,以通過突顯宏觀層面的行為異常來補充傳統評估。

Explainable Deep Learning of Resting-State Functional Connectomes Reveals Network Biomarkers of Adolescent Intelligence

2609.33422v1 by Md. Tanvir Rahman, Nabil Anan Orka, Asaduzzaman Khan, Mohammad Ali Moni

Mapping resting-state brain organization to individual differences in cognitive ability remains a major challenge in population neuroinformatics. Although deep learning enables flexible modeling of brain connectivity, limited interpretability restricts its scientific and clinical utility. To address this objective, we developed an explainable deep learning framework based on sparse projected residual networks to predict fluid, crystallized, and total intelligence from resting-state functional magnetic resonance imaging in 5,285 participants from the Adolescent Brain Cognitive Development study. We incorporated three complementary explainability methods (Integrated Gradients, Gradient Shapley Additive Explanations, and Occlusion) to interpret model behavior. The framework outperformed existing approaches, achieving Pearson correlations of 0.44, 0.58, and 0.56 for fluid, crystallized, and total intelligence, respectively, corresponding to predictive improvements of 6 to 9 percent. All three explainability methods produced near-identical feature rankings (pairwise rank correlations greater than 0.99). Consensus maps revealed a dual-layered functional architecture where primary predictive hubs localized within canonical systems, while the strongest global predictive pathways frequently bypassed these hubs through distributed, long-range relay connections. These findings suggest that intelligence emerges from the interaction between localized computational hubs and distributed communication pathways. Ultimately, these normative network architectures provide clinical reference maps to detect individual deviations, supporting earlier diagnosis, cognitive subtype stratification, and treatment monitoring in atypical neurodevelopment.

摘要:將靜息狀態下的大腦組織映射到個體在認知能力上的差異,仍然是人口神經資訊學中的一大挑戰。雖然深度學習使得大腦連接的靈活建模成為可能,但有限的可解釋性限制了其科學和臨床的實用性。為了達成這一目標,我們開發了一個基於稀疏投影殘差網絡的可解釋深度學習框架,從5,285名來自青少年大腦認知發展研究的參與者的靜息狀態功能性磁共振成像中預測流體智力、結晶智力和總智力。我們結合了三種互補的可解釋性方法(整合梯度、梯度沙普利加法解釋和遮蔽)來解釋模型行為。該框架的表現超過了現有的方法,對流體智力、結晶智力和總智力的皮爾森相關係數分別達到0.44、0.58和0.56,對應的預測改進為6%到9%。所有三種可解釋性方法產生了幾乎相同的特徵排名(成對排名相關係數大於0.99)。共識圖揭示了一種雙層功能架構,其中主要的預測樞紐位於典型系統內,而最強的全球預測通路則經常通過分散的長距離中繼連接繞過這些樞紐。這些發現表明,智力是由局部計算樞紐和分散通信通路之間的互動所產生的。最終,這些規範性網絡架構提供了臨床參考圖,以檢測個體偏差,支持早期診斷、認知亞型分層和在非典型神經發展中的治療監測。

Decoupling Token Roles in Autoregressive Pretraining

2609.33405v1 by Suqin Yuan, Runqi Lin, Kevin Qinghong Lin, Junchi Yu, Lei Feng, Chris Russell, Tongliang Liu

Autoregressive pretraining increasingly draws on heterogeneous data, making it important to understand how a model learns from an individual token. The next-token prediction objective naturally identifies a token's contribution with its own loss. However, each token is not only a prediction target but also context for what follows. Using controlled corruption, we decouple these two roles and find a reversal: making a noisy token easier to predict reduces its damage as a target but increases it as context. The same decoupling helps explain text generated by language models: generation selects each token by its fit to the prefix, while its role as context is never tested against an independently determined continuation, because that continuation is generated to fit it. At known corrupted positions, acting through the context can reduce damage that removing the token's own loss does not. Understanding and controlling what a model learns from a token therefore requires decoupling its roles.

摘要:自回歸預訓練越來越依賴異質數據,因此理解模型如何從單個標記中學習變得重要。下一個標記的預測目標自然將標記的貢獻與其自身的損失相識別。然而,每個標記不僅是預測目標,也是後續內容的上下文。通過使用控制性腐敗,我們將這兩個角色解耦,並發現了一種逆轉:使一個嘈雜的標記更容易預測會減少其作為目標的損害,但會增加其作為上下文的損害。同樣的解耦有助於解釋語言模型生成的文本:生成過程根據每個標記與前綴的契合度進行選擇,而其作為上下文的角色從未與獨立確定的延續進行測試,因為該延續是為了適應它而生成的。在已知的腐敗位置,通過上下文的作用可以減少去除標記自身損失所無法減少的損害。因此,理解和控制模型從標記中學習的內容需要解耦其角色。

The Error You See Is Not the Error You Made: Progression-aware Reasoning Origin for Reasoning Error Localization

2609.33297v1 by Yiguo Wang, Ziyuan Yang, Yi Zou, Dan Lin, Rongsheng Li, Yi Zhang

Verifying multi-step LLM reasoning requires more than determining whether a trace is correct: a useful verifier should identify where the reasoning first goes wrong. However, existing holistic methods provide little positional evidence, while forward sequential verification often treats the first rejected step as the error source. Under error propagation, this assumption can fail, since an earlier mistake may remain locally plausible and become observable only through its downstream consequences. We therefore rethink reasoning verification as a progression-aware error-source localization problem: rather than asking only where a reasoning trace first appears inconsistent, we ask which earlier step best explains how that inconsistency emerges along the trajectory. Based on this view, we propose Progression-aware Reasoning Origin (PRO), a training-free framework for first-error localization. PRO jointly models incoming support from the preceding context and outgoing compatibility with subsequent reasoning, selectively refines regions where these signals disagree, and finally performs detector-conditioned source attribution with intervention-based evidence to distinguish the true error origin from its propagated manifestations. We further formalize the gap between forward rejection and structural exposure, showing why incoming-side evidence alone is insufficient for reliable localization under error propagation. Experiments across open-form, medical, and structured reasoning tasks demonstrate consistent improvements over strong verification baselines, supporting progression-aware source attribution as a more faithful formulation of reasoning verification.

摘要:驗證多步驟 LLM 推理不僅需要確定一個痕跡是否正確:一個有用的驗證器應該能夠識別推理首次出錯的地方。然而,現有的整體方法提供的位置信息有限,而前向序列驗證通常將第一個被拒絕的步驟視為錯誤來源。在錯誤傳播的情況下,這一假設可能會失效,因為早期的錯誤可能在局部上仍然是合理的,並且只有通過其下游後果才能被觀察到。因此,我們重新思考推理驗證,將其視為一個進程感知的錯誤來源定位問題:我們不僅詢問推理痕跡首次出現不一致的地方,而是詢問哪一個早期步驟最能解釋沿著軌跡出現的不一致。基於這一觀點,我們提出了進程感知推理來源(PRO),這是一個無需訓練的首錯定位框架。PRO 共同建模來自前一上下文的支持和與後續推理的兼容性,選擇性地細化這些信號不一致的區域,並最終通過基於干預的證據進行檢測器條件的來源歸因,以區分真實的錯誤來源和其傳播的表現。我們進一步形式化了前向拒絕和結構曝光之間的差距,顯示為什麼僅依賴來自進入側的證據對於在錯誤傳播下的可靠定位是不足夠的。在開放式、醫療和結構化推理任務中的實驗顯示出對強驗證基準的一致改進,支持進程感知來源歸因作為推理驗證的更真實表述。

CORTEX: A Verified Experience Layer for Generalist Agents

2609.33260v1 by Garapati Keerthana, Manik Gupta

An agent can solve a task today and face the same task under new facts, tools, or governing knowledge tomorrow. Most agent systems can retrieve relevant text or recall prior conversations, but they lack a principled way to decide when a previous solution is still valid, when it must be adapted, and when it should be discarded. We introduce CORTEX (Contextual Orchestration and Reuse of Task EXperience), a general AI systems framework that connects specialized agents through an external layer of verified experience. Each episode records its task conditions, source and tool state, decisive predicates, proof trace, verifier, and outcome. A meta-controller chooses exact replay, checked adaptation, fresh synthesis, or escalation. Accepted episodes can become task patterns and procedural strategies through a challenge-driven development loop. This gives the system an implicit competence layer that can grow without changing model weights. We formalize system contracts for exact replay and source-version separation, and derive when reuse saves computation. A controlled two-domain implementation tests the exact-replay core on 1,000 synthetic cases. Complete-family holdouts test procedural transfer on 1,000 new-family cases across eight clinical and policy splits, with complete fresh-evidence grounding and perfect invariance to irrelevant-field and insertion-order perturbations. The transfer trace exposes the work required for verified strategy execution. These results establish an initial path toward general intelligence through reusable procedures, typed experience, and developmental transfer.

摘要:一個代理可以在今天解決一個任務,並在明天面對同一任務,但有新的事實、工具或治理知識。大多數代理系統可以檢索相關文本或回憶先前的對話,但它們缺乏一種原則性的方式來決定何時先前的解決方案仍然有效,何時必須進行調整,以及何時應該被丟棄。我們介紹了 CORTEX(上下文協調與任務經驗重用),這是一個通用的 AI 系統框架,通過一層經過驗證的經驗將專門的代理連接起來。每個事件記錄其任務條件、來源和工具狀態、決定性謂詞、證明痕跡、驗證者和結果。一個元控制器選擇精確重播、檢查調整、新的綜合或升級。接受的事件可以通過挑戰驅動的開發循環轉變為任務模式和程序策略。這為系統提供了一個隱含的能力層,能夠在不改變模型權重的情況下增長。我們為精確重播和來源版本分離形式化了系統合同,並推導出何時重用可以節省計算。受控的雙域實施在 1,000 個合成案例上測試精確重播核心。完整家庭保留測試在八個臨床和政策拆分中對 1,000 個新家庭案例的程序轉移,具有完整的新證據基礎和對無關領域及插入順序擾動的完美不變性。轉移痕跡揭示了執行經過驗證的策略所需的工作。這些結果為通過可重用程序、類型化經驗和發展轉移建立了一條通向通用智能的初步路徑。

FOCUS: Benchmarking Retinal Model Generalization from Foundation Vision Encoders to Multimodal LLMs

2609.33158v1 by David Restrepo, Chenwei Wu, Luis Filipe Nakayama, Miguel L. Martins, Stergios Christodoulidis, Maria Vakalopoulou, Enzo Ferrante

Progress in AI-based retinal image analysis has advanced with foundation models, yet evaluating their reliability remains challenging. Performance reported on a single dataset does not capture how models behave under dataset shift, across clinical definitions, or for different patient subgroups. This limitation is particularly critical in medical imaging analysis, where robustness, calibration, and fairness are essential for safe deployment. We introduce FOCUS (Foundation Ophthalmic Cross-Dataset Understanding under Shift), a cross-dataset benchmark for evaluating retinal fundus models that considers vision-only encoder models (VM), vision-language dual-encoder models (VLM), and multimodal large language models (MLLM). FOCUS harmonizes binary diabetic retinopathy, referable diabetic retinopathy, and glaucomatous optic neuropathy tasks across ten public datasets spanning diverse geographies, acquisition conditions, and label protocols. The benchmark evaluates models through a unified analysis layer that measures ranking performance, calibration, subgroup disparities, and image-quality robustness. We present a large-scale evaluation covering 532 base configurations and 228 MLLM configurations adapted through supervised fine-tuning with low-rank adaptation (LoRA). Results show that no model family consistently dominates across tasks and datasets: general VM encoders achieve the strongest average ranking performance, medical MLLMs are competitive but variable, and dual encoder VLMs benefit substantially from lightweight adaptation. Fine-tuning improves in-domain performance but exhibits heterogeneous transfer to external datasets, particularly in calibration. These findings demonstrate that retinal model evaluation is inherently multidimensional. FOCUS provides a practical framework and public benchmark to assess generalization, reliability, and robustness beyond single-dataset leaderboards

摘要:進展於基於人工智慧的視網膜影像分析已隨著基礎模型的發展而提升,然而評估其可靠性仍然具有挑戰性。單一數據集上報告的性能無法捕捉模型在數據集轉移、臨床定義之間或不同患者子群體中的行為。這一限制在醫學影像分析中特別關鍵,因為穩健性、校準和公平性對於安全部署至關重要。我們引入了FOCUS(Foundation Ophthalmic Cross-Dataset Understanding under Shift),這是一個跨數據集基準,用於評估視網膜眼底模型,考慮了僅視覺編碼器模型(VM)、視覺-語言雙編碼器模型(VLM)和多模態大型語言模型(MLLM)。FOCUS在十個公共數據集上協調二元糖尿病視網膜病變、可參考糖尿病視網膜病變和青光眼性視神經病變任務,這些數據集涵蓋了多樣的地理位置、獲取條件和標籤協議。該基準通過一個統一的分析層評估模型,測量排名性能、校準、子群體差異和影像質量的穩健性。我們呈現了一個涵蓋532個基本配置和228個經過低秩適應(LoRA)監督微調的MLLM配置的大規模評估。結果顯示,沒有任何模型家族在任務和數據集上始終佔據主導地位:一般的VM編碼器實現了最強的平均排名性能,醫學MLLM在競爭中但變化不定,而雙編碼器VLM在輕量適應中受益匪淺。微調改善了內域性能,但在外部數據集上展現出異質的轉移,特別是在校準方面。這些發現表明,視網膜模型評估本質上是多維的。FOCUS提供了一個實用的框架和公共基準,以評估超越單一數據集排行榜的泛化、可靠性和穩健性。

Relative Generalization Invariance of LLM Pretraining

2609.33016v1 by Fengzhuo Zhang, Shuche Wang, Shenggui Li, Tianyu Ruan, Jianliang He, Ivor Tsang, Tianyu Pang, Chao Du, Tianwei Zhang, Zhuoran Yang

Large Language Model (LLM) pretraining performance is jointly shaped by three components of the training triplet: the optimizer, model architecture, and training data stream. However, how these components influence performance in distinct ways remains unclear. We take a first step toward isolating their effects by studying relative generalization. We introduce Relative Generalization Invariance (RGI), the invariance of the validation-loss difference between any two tokens across models. We show that RGI approximately holds across a wide range of optimizers and moderate architectural variations, suggesting that these choices induce an approximately uniform shift in token-wise losses. In contrast, changing the training data stream can substantially alter relative generalization. We further show that RGI cannot be explained by the neural tangent kernel or mean-field regimes alone and prove that it can emerge in an overparameterized quadratic model. Overall, our work identifies RGI as a new phenomenon in LLM pretraining that helps distinguish the effects of optimizers and architectures from those of training data.

摘要:大型語言模型(LLM)預訓練的性能是由訓練三元組的三個組件共同影響的:優化器、模型架構和訓練數據流。然而,這些組件如何以不同方式影響性能仍然不清楚。我們邁出了第一步,通過研究相對泛化來隔離它們的影響。我們引入了相對泛化不變性(RGI),即在不同模型之間任何兩個標記的驗證損失差異的不變性。我們展示了RGI在廣泛的優化器和適度的架構變化中大致成立,這表明這些選擇會在標記損失上引起大致均勻的變化。相反,改變訓練數據流可以顯著改變相對泛化。我們進一步表明,RGI不能僅僅通過神經切線核或均值場範疇來解釋,並證明它可以在過參數化的二次模型中出現。總體而言,我們的工作將RGI確定為LLM預訓練中的一種新現象,幫助區分優化器和架構的影響與訓練數據的影響。

DynamicDx: Evaluating Evidence Acquisition in Video-Based Diagnosis

2609.32957v1 by Jiahui Li, Yutong Guo, Nan Yang, Wenzhan Song, Jin Lu, Fei Dou

Diagnosing a patient from video requires more than recognizing the sign: a vision-language model must turn what it sees into hypotheses, questions and tests. DynamicDx evaluates each step in 71 neurological consultations across 11 sign categories, linking authentic patient videos to confirmed diagnoses and fixed charts built from the same case reports, so that every model queries the same evidence. Across five such models, video improves accuracy by 9.9-22.5 percentage points over blind input, but neither recognition alone nor temporal order explains the gain: the cause is usually missing from the model's video-only differential diagnosis even when the sign is recognized, and shuffling the frames produces no reliable accuracy loss. Instead, a trajectory replay traces most of the gain to the investigation results the video prompts. Evidence acquisition is the bottleneck: supplying the decisive investigations raises accuracy to 73.2-93.0%. Two interventions act on it. A post-trained 4B video describer improves sign descriptions, especially from a short, densely sampled segment, and source-clean literature retrieval expands initial hypotheses; both bring the tests a model orders closer to those the treating clinicians documented and, through them, raise accuracy. For video-based diagnosis, seeing better helps when it leads to asking better.

摘要:診斷患者的視頻需要的不僅僅是識別標誌:視覺-語言模型必須將其所見轉化為假設、問題和測試。DynamicDx 評估了 71 次神經諮詢中的每一步,涵蓋 11 種標誌類別,將真實患者視頻與確認的診斷和基於相同案例報告製作的固定圖表聯繫起來,以便每個模型查詢相同的證據。在這五個模型中,視頻的準確性比盲輸入提高了 9.9-22.5 個百分點,但僅僅依賴識別或時間順序並不能解釋這一增益:即使標誌被識別,模型的視頻僅差異診斷中通常缺少原因,並且打亂幀並不會產生可靠的準確性損失。相反,軌跡重播將大部分增益追溯到視頻促進的調查結果。證據獲取是瓶頸:提供關鍵調查將準確性提高到 73.2-93.0%。有兩個干預措施對此產生影響。一個經過後訓練的 4B 視頻描述器改善了標誌描述,特別是來自短而密集取樣段的描述,而來源清理文獻檢索擴展了初步假設;兩者都使模型所訂購的測試更接近治療臨床醫生記錄的測試,並通過它們提高準確性。對於基於視頻的診斷,當更好的視覺導致更好的提問時,看到更清楚是有幫助的。

Counterfactual Self-Evolving Agents for Evidence-Grounded Reasoning

2609.32870v1 by Xing Han, Yuxin Wang, Chen Chen, Wei Dai, Gautham Krishna Gudur, Shijun Li, Hsing-Huan Chung, Gregory D. Hager, Joydeep Ghosh, Paul Pu Liang, Suchi Saria

Self-play proposer--solver methods improve reasoning by generating tasks and learning from verified solutions. However, for evidence-identifiable tasks, where case-specific evidence and domain knowledge determine a checkable answer, self-play requires generating plausible cases whose answers can be independently verified. We introduce counterfactual self-evolution, which generates counterfactual context for reconsidering the original case. A trainable Proposer constructs targeted evidence edits and describes potential outcome changes with causal explanations. We handcraft an expert-verified counterfactual instruction-tuning dataset to teach the Proposer to generate high-quality counterfactuals across a broad range of action--outcome scenarios. Each counterfactual instruction-tuning example specifies an edit within a defined category and explains its hypothesized causal effect on the decision, teaching the Proposer to reason systematically about what changes and why. We instruction-tune the Proposer on these examples, then formulate a fine-tuning reward that integrates feedback from the Solver and Verifier. Across diverse counterfactual scenarios, this reward favors high-quality counterfactuals and warranted revisions, while penalizing changes that overturn correct decisions. The counterfactual context aims to correct errors and strengthen confidence in correct decisions. Accepted counterfactuals accumulate in memory that supplies in-context evidence to the frozen Solver; the Solver adapts through evolving context rather than weight updates. We apply the framework to clinical reasoning, fact verification, and business reasoning. Our evaluation tracks performance over successive rounds as counterfactual memory grows, including transfer to harder cases. Our method achieves superior results across diverse frontier models.

摘要:自我對弈提議者--解決者方法通過生成任務並從經過驗證的解決方案中學習來改善推理。然而,對於可識別證據的任務,其中案例特定的證據和領域知識決定了可檢查的答案,自我對弈需要生成可以獨立驗證的合理案例及其答案。我們引入了反事實自我演化,該方法生成反事實背景以重新考慮原始案例。一個可訓練的提議者構建針對性的證據編輯,並用因果解釋描述潛在的結果變化。我們精心製作了一個專家驗證的反事實指令調整數據集,以教導提議者在廣泛的行動--結果場景中生成高質量的反事實。每個反事實指令調整示例指定了一個在定義類別內的編輯,並解釋其對決策的假設因果效應,教導提議者系統性地推理什麼變化以及為什麼變化。我們在這些示例上對提議者進行指令調整,然後制定一個微調獎勵,該獎勵整合了解決者和驗證者的反饋。在多樣的反事實場景中,這個獎勵偏好高質量的反事實和合理的修訂,同時懲罰推翻正確決策的變更。反事實背景旨在糾正錯誤並增強對正確決策的信心。被接受的反事實在記憶中累積,為凍結的解決者提供上下文證據;解決者通過演變的背景而不是權重更新來適應。我們將該框架應用於臨床推理、事實驗證和商業推理。我們的評估跟踪隨著反事實記憶增長而進行的多輪性能,包括轉移到更困難的案例。我們的方法在多樣的前沿模型中取得了優越的結果。

FinancialAuditBench: Benchmark Construction under Differential Privacy Using Real-World Priors

2609.32835v1 by Jerry Huang, Sarvesh Babu, Matt Van Buren, Alexander Wang, Pranav Pillai, Arush Jain, James P. Burton, Julia Hockenmaier

As AI agents are becoming widely adopted in the financial services industry, careful measurement is essential to understand where they can be reliably deployed and where oversight and professional review remain necessary. Such measurement, however, is constrained by limited access to proprietary or privacy-sensitive data. Existing benchmarks therefore often rely on publicly available data, human- and/or LLM-authored tasks, or simplified settings. We introduce FinancialAuditBench, a benchmark for evaluating agents on financial statement audit tasks, along with a framework for systematically generating synthetic engagements. Our task generation framework leverages differentially private aggregate statistics from historical audits along with audit expertise contributed through over 1,100 hours of benchmark development and review. FinancialAuditBench consists of 90 tasks spanning workpaper completion and review across six synthetic audit engagements, each containing an average of 179 files. Evaluation on eleven frontier models shows that while agents complete substantial portions of staff-level audit tasks well, they sometimes perform inappropriate procedures or produce incorrect documentation. Beyond financial auditing, our framework offers an approach for systematically generating synthetic tasks for model evaluation and training in privacy-sensitive domains.

摘要:隨著 AI 代理在金融服務行業的廣泛採用,仔細的測量對於理解它們可以可靠部署的地方以及何處仍需監督和專業審查至關重要。然後,這種測量受到對專有或隱私敏感數據的有限訪問的限制。因此,現有的基準通常依賴於公開可用數據、人類和/或 LLM 編寫的任務或簡化的設置。我們介紹了 FinancialAuditBench,一個用於評估代理在財務報表審計任務上的基準,以及一個系統生成合成參與的框架。我們的任務生成框架利用了來自歷史審計的差分隱私聚合統計數據,以及通過超過 1,100 小時的基準開發和審查貢獻的審計專業知識。FinancialAuditBench 包含 90 個任務,涵蓋六個合成審計參與的工作文件完成和審查,每個參與平均包含 179 個文件。對十一個前沿模型的評估顯示,儘管代理能夠很好地完成大量的員工級審計任務,但有時它們會執行不當的程序或產生不正確的文件。除了財務審計,我們的框架還提供了一種系統生成合成任務的方法,用於在隱私敏感領域進行模型評估和訓練。

Mandela-Bench: Multimodal Models Remember Canonical Images Instead of Seeing Them

2609.32763v1 by Yicheng Bao, Zhenkun Gao, Xiahui Guo, Mingqian Yang, Xueheng Li, Bangwei Liu, Mingang Chen, Lijun Li, Xuhong Wang, Xin Tan

Historical photographs and other canonical images can now be edited seamlessly with a single instruction, often leaving no reliable pixel-level trace. In such cases, the only evidence of manipulation may be a fact about what the image depicts. Existing benchmarks instead rely on generator artefacts, image-caption inconsistencies, visual implausibilities, or external references, and therefore do not test whether a model can use its own world knowledge to verify a recognized image. We introduce Mandela-Bench, containing 1,507 edits of canonical images: 1,359 knowledge-only forgeries, each contradicting one verifiable fact, and 148 anchor-free controls that preserve the editing process without introducing a factual contradiction, together with 474 untouched originals. We score not only whether a model detects a forgery, but whether its explanation identifies the inserted entity or the fact being violated. Across 36 multimodal models, from 0.8B parameters to frontier scale, we find a consistent failure mode. When a public figure is removed from a familiar photograph, models still name that person in up to 72.7% of responses. Some models can distinguish the replacement face from the original when shown in isolation, yet still judge the full edited photograph as authentic. Providing the true event and date does not improve knowledge-grounded detection, whereas providing the same information after cropping away the recognizable composition does. Even under explicit verification prompts, only one of the 36 models meets the KGR criterion on at least half of the forged images. These results suggest that the failures cannot be explained by missing knowledge or inadequate perception alone. Instead, they are consistent with recognition biasing verification toward the remembered canonical image rather than the observed edit.

摘要:歷史照片和其他經典圖像現在可以通過單一指令無縫編輯,通常不留下可靠的像素級痕跡。在這種情況下,唯一的操控證據可能是圖像所描繪的事實。現有的基準測試則依賴於生成器產物、圖像標題不一致、視覺不合理性或外部參考,因此並未測試模型是否能夠利用自身的世界知識來驗證已識別的圖像。我們引入了 Mandela-Bench,包含 1,507 個經典圖像的編輯:1,359 個僅知識的偽造,每個都與一個可驗證的事實相矛盾,以及 148 個無錨控件,它們保留了編輯過程而不引入事實矛盾,還有 474 個未觸碰的原始圖像。我們不僅評分模型是否檢測到偽造,還評分其解釋是否識別出插入的實體或被違反的事實。在 36 個多模態模型中,從 0.8B 參數到前沿規模,我們發現了一種一致的失敗模式。當公共人物從熟悉的照片中移除時,模型仍在高達 72.7% 的回應中提到該人。一些模型在單獨顯示替換臉時可以區分與原始臉的不同,但仍然判斷整張編輯過的照片為真實。提供真實事件和日期並未改善基於知識的檢測,而在裁剪掉可識別構圖後提供相同的信息則有改善。即使在明確的驗證提示下,36 個模型中只有一個在至少一半的偽造圖像上達到 KGR 標準。這些結果表明,失敗無法僅用知識缺失或感知不足來解釋。相反,它們與識別偏見將驗證偏向於記憶中的經典圖像而非觀察到的編輯是一致的。

What Would Falsify It? A Variable Specific Evidence Standard for Mechanistic Claims About Self Explanation

2609.32670v1 by Arshia Eftekhari zadeh

When a language model explains an answer it has already given, does it reuse the computation that produced the answer or reconstruct a story from the answer alone? Attribution, transportability and recoverability are each compatible with causal use without establishing it. We propose an evidence standard: pair each positive statistic with a variable specific null that removes the tested variable's identity while matching relevant nuisance dimensions as far as possible, and audit unmatched dimensions. We apply this standard to a known cause. A cue naming a wrong option raises the rate of choosing that option by 64 to 68 percentage points across three models. Explanations mention the cue in 1.8 percent of items or fewer in three of four models tested. Three estimator classes yield favorable statistics, but none establishes causal sensitivity to the cue contrast under its own control in the three-model analysis. In the strongest case, a recovered cue direction reaches $R^2$ of 0.95 and exceeds a geometry matched random direction in all three seeds, while a direction fitted by the same pipeline with cue labels scrambled reproduces 61 to 76 percent of its effect at comparable realized edit magnitude. A fourth model passes one interchange endpoint, but unequal edit magnitudes and a contrast that changes both cue identity and cue-answer agreement limit its interpretation. These experiments leave causal access unresolved. They establish an evidentiary requirement: favorable mechanistic statistics must survive controls for variable identity and nuisance structure. Reusable controls separate generic from identity specific transport effects, fit null directions with scrambled labels, and audit realized intervention magnitudes.

摘要:當一個語言模型解釋它已經給出的答案時,它是重用產生該答案的計算,還是僅僅從答案重建一個故事?歸因、可轉移性和可恢復性在不建立因果關係的情況下各自與因果使用相容。我們提出了一個證據標準:將每個正向統計數據與一個特定於變數的零假設配對,該零假設在盡可能匹配相關的干擾維度的同時去除被測變數的身份,並審核未匹配的維度。
我們將這一標準應用於一個已知的原因。命名錯誤選項的提示使得選擇該選項的比率在三個模型中提高了64到68個百分點。在四個測試的模型中,解釋中提到提示的比例在1.8%或更少。三個估計器類別產生了有利的統計數據,但在三模型分析中,沒有一個在其自身控制下建立對提示對比的因果敏感性。在最強的情況下,恢復的提示方向達到$R^2$為0.95,並在所有三個隨機種子中超過幾何匹配的隨機方向,而由同一管道擬合的提示標籤被打亂的方向在可比較的實現編輯幅度下重現了61到76%的效果。一個第四模型通過了一個互換端點,但不等的編輯幅度以及一個同時改變提示身份和提示-答案一致性的對比限制了其解釋。
這些實驗使因果訪問未得到解決。它們建立了一個證據要求:有利的機制統計必須在變數身份和干擾結構的控制下存活。可重用的控制分離一般的與身份特定的傳輸效應,擬合帶有打亂標籤的零方向,並審核實現的干預幅度。

When Users Change Their Minds: Measuring and Repairing Intent Drift in LLM Agents

2609.32520v1 by Yanjie Zhang, Bowen Cao, Zixin Chen, Yushi Sun

LLM agents often operate over multi-turn interactions in which user intent changes before execution. We study intent drift: the failure mode in which superseded parts of the user's intent continue to influence the final answer or tool action. We introduce IntentFlux, an executable benchmark that converts verifiable tasks into dialogues with controlled intent changes while preserving their original graders. In a 627-case calibration, mean task score falls from 0.476 to 0.384 as dialogues contain more superseded and withdrawn information. Across eight models, the rate of fully correct solutions is significantly lower when the same final task must be recovered from an evolving dialogue rather than given directly in a single turn. We further introduce StateForge, which explicitly maintains the active requirements before generation. On General-Test, it improves mean task score from 0.367 to 0.467. Providing the ground-truth final state improves performance further but still does not recover single-turn performance, indicating that state-estimation errors explain only part of the gap. These results establish intent drift as a measurable multi-turn failure mode and explicit state maintenance as a partial mitigation.

摘要:LLM 代理通常在多輪互動中運作,其中使用者的意圖在執行之前會發生變化。我們研究意圖漂移:這是一種失敗模式,其中被取代的使用者意圖部分繼續影響最終答案或工具行動。我們引入了 IntentFlux,一個可執行的基準,將可驗證的任務轉換為具有控制意圖變化的對話,同時保留其原始評分者。在627個案例的校準中,當對話包含更多被取代和撤回的信息時,平均任務分數從0.476降至0.384。在八個模型中,當必須從一個不斷演變的對話中恢復相同的最終任務,而不是直接在單輪中給出時,完全正確解決方案的比率顯著降低。我們進一步引入了 StateForge,它在生成之前明確維護活動要求。在 General-Test 上,它將平均任務分數從0.367提高到0.467。提供真實的最終狀態進一步改善了性能,但仍然無法恢復單輪性能,這表明狀態估計錯誤僅解釋了部分差距。這些結果確立了意圖漂移作為可測量的多輪失敗模式,以及明確的狀態維護作為部分緩解措施。

Explaining Textual Entailment with Lexical Entailments: Using LLMs to Supply Lexical Relations for Formal Proofs

2609.32491v1 by Jorryt de Jong, Stefan Moraca, Ettore Cesari, Lasha Abzianidze

Large Language Models (LLMs) are highly capable of natural language reasoning and appear to store a great deal of lexical knowledge, but it is still unclear how much of this knowledge they actually use when reasoning, and whether they use it in the right way. On the other hand, logic-based Natural Language Inference (NLI) systems provide transparent and formally grounded reasoning, but they need to be supplied with rich lexical knowledge to prove inferences beyond purely logical ones. In this paper, we evaluate whether LLMs can identify all lexical knowledge needed to solve NLI problems and how much this knowledge contributes to proof search in a logic-based NLI system. Our research focuses exclusively on structured lexical entailments (e.g., chinchilla$\sqsubseteq$small animal) as a proxy for structured explanations for NLI problems with an entailment label. First, we curate a dataset for a new task of explaining sentential entailments with a set of lexical entailments. The dataset is used to intrinsically evaluate LLMs on generating structured lexical explanations. Then, we use NLI as an extrinsic evaluation in a simple neuro-symbolic setting, assessing whether LLMs can supply sufficient lexical relations to LangPro, a natural-logic theorem prover for natural language. The results show that the proposed task remains challenging even for hosted proprietary LLMs, and that their contribution to theorem proving is moderate: generated relations are often only partially sound and may be tailored to the specific NLI problem rather than representing generally valid lexical knowledge.

摘要:大型語言模型 (LLMs) 在自然語言推理方面具有很高的能力,並且似乎儲存了大量的詞彙知識,但目前仍不清楚它們在推理時實際使用了多少這些知識,以及是否以正確的方式使用。另一方面,基於邏輯的自然語言推理 (NLI) 系統提供透明且有正式基礎的推理,但它們需要提供豐富的詞彙知識,以證明超越純邏輯的推論。在本文中,我們評估 LLMs 是否能夠識別解決 NLI 問題所需的所有詞彙知識,以及這些知識對基於邏輯的 NLI 系統中的證明搜索的貢獻程度。我們的研究專注於結構化詞彙推論(例如,chinchilla$\sqsubseteq$small animal),作為具有推論標籤的 NLI 問題的結構化解釋的代理。首先,我們為解釋句子推論的新任務策劃了一個數據集,該數據集包含一組詞彙推論。該數據集用於對 LLMs 在生成結構化詞彙解釋方面進行內部評估。然後,我們在一個簡單的神經符號設置中使用 NLI 作為外部評估,評估 LLMs 是否能夠為 LangPro 提供足夠的詞彙關係,LangPro 是一個用於自然語言的自然邏輯定理證明器。結果顯示,即使對於託管的專有 LLMs,所提出的任務仍然具有挑戰性,並且它們對定理證明的貢獻是適度的:生成的關係往往僅部分有效,並且可能針對特定的 NLI 問題進行調整,而不是代表普遍有效的詞彙知識。

Superposed Inference for Hyperdimensional Computing

2609.32320v1 by Quanling Zhao, Nilesh Prasad Pandey, Ye Tian, Tajana Rosing

Hyperdimensional computing (HDC) is attractive for efficient and robust learning, but conventional inference still encodes every query independently, repeatedly paying the cost of high-dimensional projection. We introduce SupHDC, a new inference paradigm that processes multiple queries through a shared encoding computation. SupHDC assigns lightweight random slot keys, superposes the keyed queries before encoding, and uses slot-specific classifiers to recover their individual predictions. A random-feature kernel view explains why exact recovery of each hypervector is unnecessary: inference only needs to preserve the class evidence that determines the prediction. Across ten datasets, SupHDC achieves 1.39x analytical speedup with no average accuracy loss, and up to 2.08x speedup with only a 2.67 percentage-point mean accuracy loss. On a Raspberry Pi~5, it delivers 2.01x measured wall-clock speedup with a 2.26 percentage-point loss in mean prediction accuracy. SupHDC shows that high-dimensional redundancy can be used not only for robustness, but also as capacity for shared inference.

摘要:超維計算(HDC)因其高效和穩健的學習而受到青睞,但傳統推理仍然獨立編碼每個查詢,重複支付高維投影的成本。我們介紹了SupHDC,一種通過共享編碼計算處理多個查詢的新推理範式。SupHDC分配輕量級隨機槽鍵,在編碼之前對鍵入的查詢進行疊加,並使用槽特定的分類器來恢復它們的個別預測。一個隨機特徵核視角解釋了為什麼不需要精確恢復每個超向量:推理只需要保留決定預測的類別證據。在十個數據集上,SupHDC實現了1.39倍的分析加速,且沒有平均準確度損失,並且在僅有2.67個百分點的平均準確度損失的情況下,最高可達2.08倍的加速。在Raspberry Pi~5上,它實現了2.01倍的測量牆時計加速,並且平均預測準確度損失為2.26個百分點。SupHDC顯示高維冗餘不僅可以用於穩健性,還可以作為共享推理的容量。

Why Directly Learning Periodic Trajectories Can Fail

2609.32254v1 by Kaixin Zheng, Anita Layton

Operator learning of periodic solutions requires deciding how simulation data should be recorded and represented. A natural choice is to integrate long enough for transients to decay and record a window wide enough to contain at least one full period of all trajectories. We find that these conservative choices can make the resulting trajectories difficult to learn, even when the underlying periodic orbits vary regularly with system parameters. Unaligned trajectories generalize poorly even within the training distribution. Phase alignment substantially improves in-distribution generalization, but models trained on a fixed physical-time window still have large errors on trajectories with periods outside the training range. We explain both failures through a common mechanism: frequency differences accumulate over time, so the target phase varies rapidly with the parameters. Predictors that cannot track this variation incur a population MSE floor in both settings; for fixed window prediction, we also derive a per-sample lower bound. We then study one of the simplest representations that escape these floors: learning an aligned, normalized waveform and its period separately. We establish regularity of the decoupled targets under ODE assumptions and show experimentally that this approach avoids both failures in ODE systems and a PDE case study.

摘要:操作學習週期解需要決定如何記錄和表示模擬數據。一個自然的選擇是整合足夠長的時間以使瞬態衰減,並記錄一個足夠寬的窗口以包含所有軌跡的至少一個完整週期。我們發現,這些保守的選擇會使得結果軌跡難以學習,即使基礎的週期軌道隨著系統參數規則變化。未對齊的軌跡即使在訓練分佈內也會泛化不佳。相位對齊顯著改善了分佈內的泛化,但在固定物理時間窗口上訓練的模型在週期超出訓練範圍的軌跡上仍然有較大的誤差。我們通過一個共同機制解釋這兩種失敗:頻率差異隨著時間累積,因此目標相位隨著參數快速變化。無法追蹤這種變化的預測器在這兩種情況下都會產生一個群體均方誤差下限;對於固定窗口預測,我們還推導出每個樣本的下限。我們接著研究一種逃避這些下限的最簡單表示之一:分別學習對齊的、歸一化的波形及其週期。我們在常微分方程假設下建立了解耦目標的規律性,並實驗表明這種方法避免了常微分方程系統中的兩種失敗以及一個偏微分方程案例研究。

A bilingual AI audiologist built through rubric-guided playbook induction outperforms human audiologists in a blinded evaluation of simulated cases

2609.32220v1 by Linkai Li, Changgeng Mo, Hanlin Yu, Congxi Lu, Shangqiguo Wang, Matthew B Fitzgerald, Shan X Wang

Audiology consultation requires structured history-taking, audiometric interpretation and patient-centred communication, yet real-world case material is scarce. We present a bilingual AI audiologist pairing a general-purpose large language model with rubric-guided playbook induction, multimodal audiogram interpretation and retrieval-augmented grounding, without fine-tuning the language-model backbone. Using a 21-item rubric and an AI patient simulator, we induced a 19-rule consultation policy from 73 training cases (43 English, 30 Chinese) and evaluated the system on 58 independent simulated cases (30 Chinese, 28 English) in a pre-specified, source-blinded comparison with 17 practising audiologists. The AI audiologist outperformed human audiologists on every case (58/58; mean paired $Δ$ = +1.35 on a 5-point composite, Cohen's d = 1.84, $P = 4.5 \times 10^{-20}$), on 20 of 21 rubric items and in both languages. Component ablation identified the playbook as the largest contributor, offering a practical route to specialist consultation agents in low-data medical domains.

摘要:聽力學諮詢需要結構化的病史採集、聽力測試解釋和以病人為中心的溝通,但現實世界中的案例材料卻稀缺。我們提出了一個雙語AI聽力學家,將通用的大型語言模型與指導性評分標準的劇本引導、多模態聽力圖解釋和檢索增強的基礎相結合,且不對語言模型的主幹進行微調。使用一個包含21項的評分標準和一個AI病人模擬器,我們從73個訓練案例(43個英文,30個中文)中引導出19條諮詢政策,並在與17名執業聽力學家的預先指定、來源盲測比較中,對58個獨立的模擬案例(30個中文,28個英文)進行了評估。AI聽力學家在每個案例中均超越了人類聽力學家(58/58;平均配對$Δ$ = +1.35,基於5分的綜合評分,Cohen's d = 1.84,$P = 4.5 \times 10^{-20}$),在21項評分標準中的20項以及兩種語言中均表現優異。組件消融識別出劇本是最大的貢獻者,為低數據醫療領域中的專家諮詢代理提供了一條實用的途徑。

Evaluating Single and Multi-Omics Based Explainable Artificial Intelligence (MOXAI) for Molecular Subclass Classification of Adult-Type Diffuse Gliomas

2609.32190v1 by Md Zahangir Alom, Quynh T. Tran, Breuer Alexandar, Brent A. Orr

DNA methylation (DNAM) profiling has emerged as a powerful diagnostic tool for classifying brain and solid tumors. However, existing computational models typically analyze methylation and copy number variation (CNV) data separately, failing to capture the complementary information their integration could provide. Moreover, current classification models lack mechanisms for within-class risk assessment analogous to traditional tumor grading, and no established explainability method can attribute classification decisions to specific genomic loci. In this paper, we present MOXAI (Multi-Omics Based Explainable AI), a deep learning framework that integrates DNA methylation and copy number data from methylation arrays to classify molecular subtypes of adult-type diffuse gliomas, alongside single-modality variants for comparison. Using a cohort from The Cancer Genome Atlas (TCGA), we trained ResNet50, DINOv2, and Graph Attention Network (GAT) models on methylation data alone, copy number data alone, and combined multimodal data. We further developed explainable AI (XAI) methods based on class activation maps (CAMs) and gradient-weighted CAM (Grad-CAM) to identify the specific CpG sites, genes, and chromosomal regions most relevant to each classification decision. The multimodal model achieved up to 92.98% cross-validation accuracy, outperforming models trained on CNV data alone. DINOv2 showed the strongest generalization, reaching 94.25% accuracy (confidence >0.9) on independent validation sets. XAI results aligned with established molecular features of adult-type diffuse glioma subtypes, confirming the biological interpretability of the framework.

摘要:DNA 甲基化 (DNAM) 檔案已成為分類腦部和實體腫瘤的強大診斷工具。然而,現有的計算模型通常分別分析甲基化和拷貝數變異 (CNV) 數據,未能捕捉其整合所能提供的互補信息。此外,當前的分類模型缺乏類內風險評估機制,類似於傳統腫瘤分級,且沒有建立的可解釋性方法能將分類決策歸因於特定的基因組位點。在本文中,我們提出了 MOXAI (基於多組學的可解釋 AI),這是一個深度學習框架,整合了來自甲基化陣列的 DNA 甲基化和拷貝數據,以分類成人型擴散性膠質瘤的分子亞型,並提供單一模態變體以供比較。使用來自癌症基因組圖譜 (TCGA) 的一個隊列,我們僅在甲基化數據、僅在拷貝數據及結合多模態數據上訓練了 ResNet50、DINOv2 和圖注意網絡 (GAT) 模型。我們進一步開發了基於類激活圖 (CAMs) 和梯度加權 CAM (Grad-CAM) 的可解釋 AI (XAI) 方法,以識別與每個分類決策最相關的特定 CpG 位點、基因和染色體區域。多模態模型達到了高達 92.98% 的交叉驗證準確率,超越了僅在 CNV 數據上訓練的模型。DINOv2 展現出最強的泛化能力,在獨立驗證集上達到了 94.25% 的準確率 (信心 >0.9)。XAI 結果與已建立的成人型擴散性膠質瘤亞型的分子特徵一致,確認了該框架的生物學可解釋性。

REALM: Regime-Switching, Explainable, and Activation-Induced Linear Models

2609.32141v1 by Xiaoran Cheng, Sen Na, Jia Li

Deep ReLU networks are piecewise-affine mappings that partition the input space into cells, each characterized by a distinct activation pattern. This structure motivates fitting a local linear model within each cell to preserve predictive accuracy while improving interpretability. The challenge is to identify regimes that are stable, data-adaptive, and easy to explain. We propose REALM, a mixture of linear models whose regimes are induced by neural activation patterns. Because the number of activation cells in a deep neural network (DNN) can grow rapidly with depth, we first distill a deep teacher into a wide, shallow student network (WSSN), then binarize and cluster its hidden-layer activations to define the regimes and fit a linear model within each regime. Since the regimes are discovered from internal structure, the router does not carry the predictive burden. To make regime assignment interpretable, we train a multiclass logistic regression, the explanatory gate, to reproduce the regime assignments. The two-level structure is interpretable at both stages in terms of raw tabular or learned convolutional features: the gate identifies features that determine regime assignments, while the linear models identify features that drive predictions within each regime. We analyze an idealized setting that illustrates a trade-off between partition complexity and stability: as the number of regimes grows, finer partitions can improve approximation but may reduce regime-assignment stability. Experiments on tabular and image datasets show that REALM achieves competitive predictive performance relative to other DNN-guided mixture surrogates and inherently interpretable models while producing stable regime-level explanations.

摘要:深度 ReLU 網絡是分段仿射映射,將輸入空間劃分為各個單元,每個單元都有其獨特的激活模式。這種結構促使我們在每個單元內擬合一個局部線性模型,以保持預測準確性同時提高可解釋性。挑戰在於識別穩定、數據自適應且易於解釋的狀態。我們提出了 REALM,一種由神經激活模式引導的線性模型混合。由於深度神經網絡 (DNN) 中的激活單元數量可能隨著深度迅速增長,我們首先將深層教師網絡提煉成一個寬而淺的學生網絡 (WSSN),然後對其隱藏層激活進行二值化和聚類,以定義狀態並在每個狀態內擬合線性模型。由於這些狀態是從內部結構中發現的,因此路由器不承擔預測負擔。為了使狀態分配可解釋,我們訓練了一個多類別邏輯回歸模型,即解釋閘,以重現狀態分配。這種兩級結構在原始表格或學習的卷積特徵方面在兩個階段都是可解釋的:閘識別決定狀態分配的特徵,而線性模型則識別在每個狀態內驅動預測的特徵。我們分析了一個理想化的設置,展示了劃分複雜性和穩定性之間的權衡:隨著狀態數量的增加,更細的劃分可以改善近似,但可能會降低狀態分配的穩定性。在表格和圖像數據集上的實驗顯示,REALM 相較於其他 DNN 引導的混合代理和內在可解釋的模型,實現了具有競爭力的預測性能,同時產生穩定的狀態級解釋。

Reasoning Concentrates Errors, and Self-Consistency Never Notices

2609.32035v1 by Asaad Althoubi

Self-consistency assumes that independent samples disagree when a model is unsure, so agreement is evidence of correctness. Holding weights fixed and toggling only a reasoning mode, over five benchmarks and 74,944 samples, we show that reasoning concentrates a model's errors: the probability that two independently drawn wrong answers coincide rises in all ten dataset-scale comparisons (p = 0.00098), and in nine of nine after restricting both arms to the problems each gets wrong. Where the answer space is unbounded, reasoning cuts the distinct answers produced to 0.43-0.65 of the non-reasoning count; where it is bounded, both arms hold an identical option set and reasoning concentrates mass on it instead, which no positional prior can explain at fixed weights. The aggregate cost is smaller than the mechanism predicts, because reasoning also shrinks the set of problems where answer diversity can decide anything, in ten of ten cells and by 2.7x; normalized for available headroom, both arms convert a quarter of it in domain. Confidence weighting does not recover what is left. Across 280 method-dataset-model combinations on eight models and five benchmarks, not one beats plain majority voting after correction; weighted voting agrees with it on 98.5% of problem-method pairs and is right 56.3% of the time on the rest; and a signal's direction can invert within fixed weights, with answer log-probability predicting correctness when reasoning is off and error when it is on. A learned six-signal combination gains nothing out of domain. Confidence signals should be evaluated on decisions, not on discrimination.

摘要:自我一致性假設當模型不確定時,獨立樣本會出現不一致,因此一致性是正確性的證據。固定權重並僅切換推理模式,在五個基準和74,944個樣本中,我們顯示推理集中了一個模型的錯誤:兩個獨立抽取的錯誤答案重合的概率在所有十個數據集規模的比較中上升(p = 0.00098),在將兩個臂限制於各自錯誤的問題後,九個中有九個也如此。當答案空間是無界的時候,推理將產生的不同答案減少到非推理計數的0.43-0.65;當它是有界的時候,兩個臂持有相同的選項集,而推理則將質量集中於此,這是固定權重下任何位置先驗無法解釋的。總體成本小於機制預測的,因為推理也縮小了答案多樣性能決定任何事情的問題集,在十個單元中均如此,且縮小幅度為2.7倍;經過可用空間的標準化,兩個臂在領域中轉換了四分之一的空間。信心加權無法恢復剩餘的部分。在280種方法-數據集-模型組合中,涵蓋八個模型和五個基準,經過修正後,沒有一種方法超過普通的多數投票;加權投票在98.5%的問題-方法對上與其一致,並在其餘的情況下正確率為56.3%;而信號的方向可以在固定權重內反轉,當推理關閉時,答案的對數概率預測正確性,而當推理開啟時則預測錯誤性。一個學習到的六信號組合在領域外沒有任何收益。信心信號應該在決策上進行評估,而不是在區分上。

A Safety-Bounded SDC-to-MCP Gateway for Medical AI Agents

2609.31358v1 by Bennet Gerlach, Stefan Fischer

The Model Context Protocol (MCP) provides a common interface through which AI applications discover and use external resources and tools. It allows language-model agents to ground their reasoning in current system state and interact with heterogeneous services. In medical environments, however, exposing device state and action affordances requires deterministic constraints on possible effects. We present an IEEE 11073 Service-Oriented Device Connectivity (SDC)-to-MCP gateway that exposes metrics, alarms, context references, and semantic metadata as read-only resources, while representing selected action affordances as policy-validated dry-run tools. The term safety-bounded denotes a narrow no-execution property: agent-facing requests dispatch no SDC device operation. A Python prototype supports simulated fault and lifecycle experiments, a software-reference protocol path spanning independent Java and Python implementations, deterministic baselines, representation ablations, and multi-model agent evaluation. The results show semantically explicit resource exposure, visible rejection of invalid or outdated state, and preservation of the no-execution boundary across resource, proposal, and authorization paths. Explicit semantic metadata improved conformity to required metric identifiers in structured alarm outputs relative to a generic representation, while retained structured-output failures reveal a distinction between plausible narrative answers and task-compliant machine-readable results.

摘要:模型上下文協議 (MCP) 提供了一個共同的介面,讓 AI 應用程式發現並使用外部資源和工具。它允許語言模型代理根據當前系統狀態進行推理並與異構服務互動。然而,在醫療環境中,暴露設備狀態和行動可行性需要對可能的影響施加確定性的限制。我們提出了一個 IEEE 11073 服務導向設備連接 (SDC) 到 MCP 的閘道,該閘道將指標、警報、上下文參考和語義元數據作為只讀資源暴露,同時將選定的行動可行性表示為經政策驗證的模擬工具。術語安全界限表示一種狹窄的無執行特性:面向代理的請求不會調度任何 SDC 設備操作。一個 Python 原型支持模擬故障和生命週期實驗,涵蓋獨立的 Java 和 Python 實現的軟體參考協議路徑、確定性基準、表示消融和多模型代理評估。結果顯示語義明確的資源暴露、對無效或過時狀態的可見拒絕,以及在資源、提案和授權路徑中保持無執行邊界。明確的語義元數據改善了結構化警報輸出中對所需指標標識符的符合性,相較於一般表示,保留的結構化輸出失敗揭示了合理敘述答案與符合任務的機器可讀結果之間的區別。

DriveHierarchy: A Benchmark for Diagnosing VLM Driving Capabilities from Open-Loop Understanding to Closed-Loop Execution

2609.31814v1 by Chengkai Xu, Jiaqi Liu, Yicheng Guo, Peng Hang, Jian Sun

Evaluating VLM-based autonomous driving remains difficult because driving competence is composite, where a capable system must ground traffic participants and hazards, integrate context across views and time, reason about future evolution, and act appropriately under closed-loop interaction. Existing benchmarks usually assess either open-loop understanding or closed-loop driving but provide limited structure for explaining how these abilities are organized, how they relate, and how they may inform model diagnosis and improvement. We present \textsc{DriveHierarchy}, a hierarchical benchmark that organizes VLM-based autonomous driving into four ranks, spanning perceptual grounding, contextual memory, mental reasoning, and closed-loop execution. To instantiate this hierarchy, we integrate multiple open-source autonomous-driving datasets into a unified open-loop benchmark with 76,798 question-answer pairs over 84,279 frames and develop a closed-loop simulation platform with interactive scenario construction on a real-world road network, from which 100 driving scenarios are curated for embodied evaluation. Experiments on 15 VLMs show that \textsc{DriveHierarchy} captures structured but non-redundant capability variation, relates open-loop understanding to closed-loop driving, and provides a practical basis for diagnosis and benchmark-guided optimization. \textsc{DriveHierarchy} therefore serves as a unified framework for evaluating and improving VLM-based autonomous driving systems. An anonymized project has been released on https://github.com/PerfectXu88/DriveHierarchy

摘要:評估基於 VLM 的自主駕駛仍然困難,因為駕駛能力是複合的,能夠的系統必須能夠定位交通參與者和危險,整合跨視角和時間的上下文,推理未來的演變,並在閉環互動中適當行動。現有的基準通常評估開環理解或閉環駕駛,但對於解釋這些能力如何組織、它們之間的關係,以及它們如何能夠幫助模型診斷和改進,提供的結構有限。我們提出了 \textsc{DriveHierarchy},這是一個將基於 VLM 的自主駕駛組織成四個等級的分層基準,涵蓋感知定位、上下文記憶、心理推理和閉環執行。為了實現這一層級,我們將多個開源自主駕駛數據集整合成一個統一的開環基準,包含 76,798 個問答對,涵蓋 84,279 幀,並開發了一個閉環模擬平台,能夠在現實世界的道路網絡上進行互動場景構建,從中策劃出 100 個駕駛場景以進行具體評估。對 15 個 VLM 的實驗顯示,\textsc{DriveHierarchy} 捕捉了結構化但不冗餘的能力變化,將開環理解與閉環駕駛相關聯,並提供了診斷和基準引導優化的實用基礎。因此,\textsc{DriveHierarchy} 作為評估和改進基於 VLM 的自主駕駛系統的統一框架。已在 https://github.com/PerfectXu88/DriveHierarchy 上發布了一個匿名項目。

Rethinking Data Quality for AI-Driven Systems: Evidence from Practitioner Interviews

2609.31191v1 by Hariharan Gopinath, Jan Bosch, Helena Holmström Olsson

Data quality research has usually treated data as an input that is stored, processed, and validated. In AI-driven software-intensive systems, data also shapes model behavior, evaluation, and lawful use. Empirical evidence remains limited on how practitioners define, assess, and manage quality under these conditions. We interviewed 16 practitioners from nine organizations and analyzed the transcripts using reflexive thematic analysis and developed six themes from participants' accounts. In AI systems, traceability shifted from modular debugging to attributing model behavior, while using models as quality assessors introduced circularity. Agent context and memory became data objects, and synthetic and pseudo-labeled data made authenticity a quality concern. In foundation-model development, lawfulness became a gate for training data, while representativeness was judged through coverage of situations in which the system must behave safely. Prior ML research examines many of these problems separately. Our study provides a practitioner-grounded account of how they are encountered together as an engineering and organizational concern. We also interpret five recurring conditions as helping explain how the themes relate to reduced trust in data and AI outcomes. We synthesize these findings through lifecycle assurance: a conceptual framing focused on producing evidence that data can support a specific AI claim when its influence may be embedded in model behavior, model-based judgments, or agent actions.

摘要:數據質量研究通常將數據視為一種被儲存、處理和驗證的輸入。在以 AI 驅動的軟體密集型系統中,數據也塑造了模型行為、評估和合法使用。在這些條件下,實證證據對於從業者如何定義、評估和管理質量仍然有限。我們訪談了來自九個組織的 16 位從業者,並使用反思主題分析法分析了訪談記錄,從參與者的敘述中發展出六個主題。在 AI 系統中,追溯性從模組調試轉變為歸因於模型行為,而將模型用作質量評估者則引入了循環性。代理上下文和記憶成為數據對象,而合成數據和偽標記數據使得真實性成為一個質量問題。在基礎模型開發中,合法性成為訓練數據的門檻,而代表性則通過系統必須安全行為的情境覆蓋來評判。先前的機器學習研究分別檢視了許多這些問題。我們的研究提供了一個以從業者為基礎的敘述,說明它們如何作為工程和組織問題共同出現。我們還解釋了五個反覆出現的條件,幫助說明這些主題如何與對數據和 AI 結果的信任減少相關。我們通過生命週期保證綜合這些發現:這是一個專注於產生證據的概念框架,證明數據可以支持特定 AI 主張,當其影響可能嵌入在模型行為、基於模型的判斷或代理行動中時。

Bayesian Optimization with Fisher Information Geometry: Gradient Bounds and Trust-Region Methods

2609.31107v1 by Saksham Kiroriwal, Julius Pfrommer, Jürgen Beyerer

We study Bayesian optimization (BO) through the lens of information geometry. Pulling back the Fisher information metric through the surrogate posterior map yields a local sensitivity tensor on the input space, which leads to an upper bound on the gradient of reparameterizable acquisition functions. This view explains vanishing-gradient behavior in high-dimensional BO and provides a common interpretation of heuristics such as RAASP and dimension-scaled lengthscales. Building on this analysis, we propose FITR, a trust-region-based BO method that replaces lengthscale-based scaling by local pullback-Fisher weights. FITR is not restricted to GP kernels with explicit lengthscales. On GP benchmarks with an SE kernel, experiments show competitive performance using FITR. The proposed method also easily generalizes to non-isotropic surrogates, although the gains are more task-dependent in that setting.

摘要:我們通過信息幾何的視角研究貝葉斯優化 (BO)。通過代理後驗映射回推費舍爾信息度量,產生了一個輸入空間上的局部敏感性張量,這導致了可重新參數化獲取函數梯度的上界。這一觀點解釋了高維 BO 中的消失梯度行為,並提供了對 RAASP 和維度縮放長度尺度等啟發式方法的共同解釋。在此分析的基礎上,我們提出了 FITR,一種基於信任區域的 BO 方法,通過局部回推費舍爾權重取代基於長度尺度的縮放。FITR 不僅限於具有明確長度尺度的 GP 核心。在具有 SE 核心的 GP 基準測試中,實驗顯示使用 FITR 的競爭性能。所提出的方法也很容易推廣到非各向同性的代理,儘管在該設置中增益更依賴於任務。