Knowledge Graphs
Knowledge Graphs
| Publish Date | Title | Authors | Homepage | Code |
|---|---|---|---|---|
| 2026-10-01 | KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards | Pengfei Li et.al. | 2610.02206v1 | null |
| 2026-10-01 | Higher-Order Molecular Grammars for Generative and Foundation Models in Chemistry | Yiming Huang et.al. | 2610.02186v1 | null |
| 2026-10-01 | From Knowledge Access to Source Learning: Developing Source-Specific Competence | Lucheng Fu et.al. | 2610.02150v1 | null |
| 2026-10-01 | Mem++: Non-Destructive Memory for Long-Term Organizational LLM Agents | Ahmad Yehia et.al. | 2610.02002v1 | null |
| 2026-10-01 | Can AI Oversight Be Zero Knowledge? | Alessandro Chiesa et.al. | 2610.01995v1 | null |
| 2026-10-01 | Latent JEPA: Abstract Future Prediction for Latent Reasoning in Chemistry | Xinjian Zhao et.al. | 2610.01947v1 | null |
| 2026-10-01 | Mapping the RAG Landscape: A Four Axis Taxonomy of Efficiency, Defense, Interactivity, and Reasoning | Meghana Sunil et.al. | 2610.01936v1 | null |
| 2026-10-01 | From Network Intrusion Detection to Blockchain-Backed Endpoint Detection and Response: Mapping the Landscape of Decentralized Detection-and-Response Architectures | Yahya Shahsavari et.al. | 2610.01872v1 | null |
| 2026-10-01 | Walking the Embedding Space: Datastore Extraction from Multimodal RAG | Maria Carmen Jica et.al. | 2610.01871v1 | null |
| 2026-10-01 | Detecting Inconsistencies in Model Specifications with LLM-as-Verifier Reasoning | Zichen Xie et.al. | 2610.01847v1 | null |
| 2026-10-01 | Code Owns the Simulation, Jev Owns the Evaluation | Yaodong Yang et.al. | 2610.01834v1 | null |
| 2026-10-01 | The Asymptotics of Language Model Alignment with Memory | Haricharan Balasundaram et.al. | 2610.01828v1 | null |
| 2026-10-01 | A Matryoshka Hierarchical RAG for Efficient Multi-Hop Question Answering | Gianluca Bonifazi et.al. | 2610.01767v1 | null |
| 2026-10-01 | Task-Oriented Rank Adaptation for Continual Learning in Text Classification | Rey Sanchez Lopez et.al. | 2610.01702v1 | null |
| 2026-10-01 | Iterative Policy Refinement through Semantic Rollout Analysis | Feiyu Gavin Zhu et.al. | 2610.01652v1 | null |
| 2026-10-01 | Managing Context and Communication in Distributed Agentic UAV Swarms | Andrea Iannoli et.al. | 2610.01569v1 | null |
| 2026-10-01 | From Rules to Neural Graphs: Scalable Structured Prediction for Patent Prior Art Search | Nikolai Zenovkin et.al. | 2610.01553v1 | null |
| 2026-10-01 | Auto-Formalizing Neuro-Symbolic Predictors | Samuele Bortolotti et.al. | 2610.01519v1 | null |
| 2026-10-01 | Decision Titan: Test-Time Training for Long-Term Memory in Offline Reinforcement Learning | Jude Waide et.al. | 2610.01513v1 | null |
| 2026-10-01 | OverAct: Measuring and Mitigating Proactive Over-Authorization in LLM Tool-Calling Agents | Taolin Zhang et.al. | 2610.01508v1 | null |
| 2026-10-01 | A Multi-Agent LLM Framework for Personalized Health Checkup Interpretation and Guidance | HyungJun Kim et.al. | 2610.01451v1 | null |
| 2026-10-01 | LLM-Assisted Discovery of Typed Semantic Links for Ontology Network Construction | Nouha Hayouni et.al. | 2610.01393v1 | null |
| 2026-10-01 | Generation Provenance Before Behavior Attribution: Auditing Synthetic Speech Research Objects | Sidi Chang et.al. | 2610.01378v1 | null |
| 2026-10-01 | ARCCS: An Automated Regulatory Compliance Checking System | Giorgos Filandrianos et.al. | 2610.01345v1 | null |
| 2026-10-01 | An ontology for cross-sectoral crisis management: core and public health modules | Aldo Gangemi et.al. | 2610.01326v1 | null |
| 2026-10-01 | Dependency-Aware Reward Shaping for Agentic Reinforcement Learning | Ziyi Chen et.al. | 2610.01207v1 | null |
| 2026-10-01 | Federated Agent Optimization | Qiang Yang et.al. | 2610.01195v1 | null |
| 2026-10-01 | Temporally-Resolved Token Attribution Reveals the Generation Dynamics of Diffusion Language Models | Darpan Aswal et.al. | 2610.01177v1 | null |
| 2026-10-01 | OrbitTAMP: Grounding Language Models for Task and Motion Planning in Spacecraft Rendezvous | Yuji Takubo et.al. | 2610.01093v1 | null |
| 2026-10-01 | MOMAT: Mixture of Multiple Atlases for Low-Power Jailbreak Defense of Quantized LLMs | Boyang Li et.al. | 2610.01058v1 | null |
| 2026-10-01 | Capturing In-Context Learning Dynamics with Task Operators | Guangzhi Xiong et.al. | 2610.01054v1 | null |
| 2026-10-01 | Beyond Answer Confidence: A Controlled Audit of Self-Knowledge in a Black-Box Decision Model | Sharath M Shankaranarayana et.al. | 2610.01006v1 | null |
| 2026-10-01 | Distilling Directional Verification | Jungseob Lee et.al. | 2610.00997v1 | null |
| 2026-10-01 | Structure-agnostic Causal Representation Learning | Arman Behnam et.al. | 2610.00968v1 | null |
| 2026-10-01 | ABDA-NL: A Natural-Language Scenario Explorer for Argument-Based Reasoning | Shawn Bowers et.al. | 2610.00947v1 | null |
| 2026-10-01 | Screw Attention: Rigid-Body Algebra Inside a Transformer | Aly Magassouba et.al. | 2610.00904v1 | null |
| 2026-10-01 | Match the Distribution, Not the Compute: Post-Training Multi-Token Prediction Heads | Prachi Badarayani et.al. | 2610.00888v1 | null |
| 2026-09-30 | Backdoor Purification for LoRA-Tuned LLMs via Null-Space Projection | Jianwei Li et.al. | 2610.00685v1 | null |
| 2026-09-30 | Ontology-Grounded, Reasoner-Verified Benchmarks for Evaluating LLM Reasoning in Scientific AI | Nishtha N. Vaidya et.al. | 2610.00682v1 | null |
| 2026-09-30 | PhysicsMate: A Curriculum-Grounded Bengali Benchmark for Secondary Physics QA with Small-Model Adaptation | Rashid Azraf Jahin et.al. | 2610.00664v1 | null |
| 2026-09-30 | Where's Waldo? Query-language Preference under Cross-lingual Knowledge Disparities | Dayeon Ki et.al. | 2610.00606v1 | null |
| 2026-09-30 | Emergent Unfaithfulness: How Alignment Training Causes Language Models to Silently Override Task Faithfulness | Pardis Sadat Zahraei et.al. | 2610.00568v1 | null |
| 2026-09-30 | Assessing the Impact of Language Disparity on Multilingual Linguistic Ability in Large Language Models | Zhanyu Chen et.al. | 2610.00540v1 | null |
| 2026-09-30 | EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery | Young-Jun Lee et.al. | 2609.40340v1 | null |
| 2026-09-30 | Linguistic Loopholes in LLM Unlearning: From a 174-Language Benchmark to Coverage-Aware Unlearning | Tyler Skow et.al. | 2609.40286v1 | null |
| 2026-09-30 | EviRover: Reinforcing Agentic Perception Beyond a Glance | Kaixuan Fan et.al. | 2609.40230v1 | null |
| 2026-09-30 | Learning from Research: Toward Lifelong Agent Harness Evolution | Jingbo Yang et.al. | 2609.40169v1 | null |
| 2026-09-30 | On the (In)effectiveness of AMR Augmentation for Large Language Models | Hoa Quynh Nhung Nguyen et.al. | 2609.40121v1 | null |
| 2026-09-30 | Persistent Context Graphs for Efficient Memory Compaction in LLM Agents | Jingbo Yang et.al. | 2609.40118v1 | null |
| 2026-09-30 | JuryFlow: Disagreement-Guided Human-in-the-Loop Multi-Agent Evaluation | Mufeng Yang et.al. | 2609.40103v1 | null |
| 2026-09-30 | AutoDataBench: A Data-centric Testbed for Accelerating Auto Research | Ruifeng Yuan et.al. | 2609.40097v1 | null |
| 2026-09-30 | LongEmo: Towards Emotion Understanding and Reasoning in Long Videos | Shuo Zhang et.al. | 2609.40079v1 | null |
| 2026-09-30 | TACTIC: Temporal and Context-Aware LLM Tactical Planning for Roadside LiDAR Attacks | Yiming Gao et.al. | 2609.39969v1 | null |
| 2026-09-30 | AIMS: An Agentic AI Framework for Sim-to-Real Multi-Modal ISAC | Yijie Bian et.al. | 2609.39964v2 | null |
| 2026-09-30 | Learning to Cover Locally: Graph Neural Combinatorial Optimization under a Hard Information Horizon | Johannes F. Loevenich et.al. | 2610.00422v1 | null |
| 2026-09-30 | MCD: Causal Distillation of Multimodal In-Context Learning in Large Vision-Language Models | Yanshu Li et.al. | 2609.39920v1 | null |
| 2026-09-30 | The Concrete-Arbitrary Gap: Kinship Reasoning in LLMs Is Not Indifferent to Presentation | Thomas Pashby et.al. | 2609.39913v1 | null |
| 2026-09-30 | DoGBench: Can Agents Meet Expert Standards for User-Facing Documentation? | Frances Liu et.al. | 2609.39909v1 | null |
| 2026-09-30 | Cognitive Enhancement: Rethinking the Necessity of Role-Playing for Large Language Models | Xingjie Zhuang et.al. | 2609.39853v1 | null |
| 2026-09-30 | Explore-on-Graph: Hybrid Embedding-LLM Reasoning for Knowledge Graph Question Answering under Incompleteness | Ola El Khatib et.al. | 2609.39786v1 | null |
| 2026-09-30 | MemCodex: Self-Programming Hierarchical Memory for Language Agents | Xiaoqiang Wang et.al. | 2609.39765v1 | null |
| 2026-09-30 | OverForge: Reasoning Through Strategies and Tactics Helps Cooperative Lifelong Adaptation | Oana Madalina Fron et.al. | 2609.39727v1 | null |
| 2026-09-30 | ArchitectureIQ: On the Measure of Training Intuition | Zirui Ren et.al. | 2609.39714v1 | null |
| 2026-09-30 | ChronoGraph: Functional 4D Scene Graphs with Vision-Language Models for Interaction Understanding and Grounded Planning | Chenyangguang Zhang et.al. | 2609.39665v1 | null |
| 2026-09-30 | Zero-Compute Cross-Lingual Transferability Estimation Using Typological Feature Proxies | Dalton Raphael Harmsen et.al. | 2609.39640v1 | null |
| 2026-09-30 | RankEvolve: A Reliable Multi-Agent Auto-Research Harness for Evolving Ranking Models | Zheng Chen et.al. | 2609.39551v1 | null |
| 2026-09-30 | Learning Normal Diffusion Dynamics for Backdoor Defense in Text-to-Image Models | Junjian Li et.al. | 2609.39548v1 | null |
| 2026-09-30 | A Reusable Semantic Web Framework for Evidence-Grounded Fundamental Rights Impact Assessments under the EU AI Act | Faith Olopade et.al. | 2609.39537v1 | null |
| 2026-09-30 | Beyond the Shadows of Plato's Cave: Evaluating False Memory in Autonomous Agents via Counterfactual Reasoning | Quan M. Tran et.al. | 2609.39473v1 | null |
| 2026-09-30 | CAST: Causal Advantage-Structured Training with Spatially Grounded Compositional Rewards for Diffusion Models | Shu Yu et.al. | 2609.39441v1 | null |
| 2026-09-30 | Inferring Causal Relations between Two Sequences of Events with Language Models | Nishchal Prasad et.al. | 2609.39406v1 | null |
| 2026-09-30 | Exploring Heterogeneous Model Merging Approach for Complex Knowledge Transfer | Jiahe Fan et.al. | 2609.39369v1 | null |
| 2026-09-30 | Offline Guidance, Online Reasoning: Reusing LLM Feedback for Small Language Models | Bohan Zhang et.al. | 2609.39346v1 | null |
| 2026-09-30 | Understanding as No-Arbitrage: Bounded Dutch Books as a Definition and Training Objective for Language Models | Daniel Dragonevskiy et.al. | 2609.39341v1 | null |
| 2026-09-30 | WorkGenesis: Building the Worlds That Teach Agents to Work | Xinyu Zhu et.al. | 2609.39325v1 | null |
| 2026-09-30 | Faithful Dual-constrained Erasure for Robust LLM Safety Alignment | Jiaqing Li et.al. | 2609.39279v1 | null |
| 2026-09-30 | Fyan: A Human--AI Harness with Semantic Auditing for Document-Level Formalization | Wei Zhao et.al. | 2609.39228v1 | null |
| 2026-09-30 | DAGent: Evaluate-then-Grow Planning for Deep Research Agents | Hanwen Liu et.al. | 2609.39154v1 | null |
| 2026-09-30 | Rep2Skill: Representation-Guided Skill Self-Evolution for LLM Agents | Kaixing Zhang et.al. | 2609.39149v1 | null |
| 2026-09-30 | MASCRDM: Multi-Agent System for Compliance Risk Detection and Mitigation in Training Process of Large Language Models | Yan Zhang et.al. | 2609.39107v1 | null |
| 2026-09-30 | Multi-LLM Collaborative Alignment via Stackelberg Games | Christina Hahn et.al. | 2609.39076v1 | null |
| 2026-09-30 | CORE: Conflict-Oriented Reasoning Elimination for Verifiable Language-Model Search | Siyu Song et.al. | 2609.39069v1 | null |
| 2026-09-30 | Structure-aware Reinforcement Learning for Protein Directed Evolution | Zikun Nie et.al. | 2609.39048v1 | null |
| 2026-09-30 | SimEX: Simulation-Integrated Robotics AutoResearch | Jiaheng Hu et.al. | 2609.38982v1 | null |
| 2026-09-30 | DrivingBench: Can Vision-Language Models Drive a Toyota Corolla? | Aditya Ramabadran et.al. | 2609.38948v1 | null |
| 2026-09-30 | Prototype-guided Bilateral Alignment Multimodal Federated Learning | Tianchi Liao Tianchi_Liao et.al. | 2609.38925v1 | null |
| 2026-09-30 | GraphForge: Training Working Agents with Graph-Anchored Workspace Synthesis | Qisheng Su et.al. | 2609.38923v1 | null |
| 2026-09-30 | Unlearning Deceptive Behaviors in LLMs with Contrastive Forget Sets | Haoran Tang et.al. | 2609.38909v1 | null |
| 2026-09-30 | K2P: Label-Free Knowledge to Prompt Distillation | Yingchuan Zhang et.al. | 2609.38898v1 | null |
| 2026-09-30 | Unmerge: Efficient Machine Unlearning via Task Arithmetic | Haoran Tang et.al. | 2609.38895v1 | null |
| 2026-09-30 | Right Answers, Costly Models: The Efficiency Gap in LLM-based Optimization Modeling | Zhong Li et.al. | 2609.38884v1 | null |
| 2026-09-30 | Does Learning Protein Folding Generalize to Broader Reasoning? | Yong Liu et.al. | 2609.38879v1 | null |
| 2026-09-30 | StateTree: Enhancing Long-Term Dialogue Reasoning via Reinforcement Learning | Naen Xu et.al. | 2609.38809v1 | null |
| 2026-09-30 | GraphCert: Bootstrap Agentic Graph Reasoning with Certified Evidence Rubrics | Weiqi Jiang et.al. | 2609.38798v1 | null |
| 2026-09-30 | Evaluating Persistent Calibration under Evolving Model Knowledge | Victor Wang et.al. | 2609.38797v1 | null |
| 2026-09-30 | Self-Evolving Algorithm-Design Agents: Escaping In-Context Evolutionary Stagnation via Population-Curated Policy Optimization | Chen Lu et.al. | 2609.38757v1 | null |
| 2026-09-30 | Learning to Route in Visual Space via Multi-Step Embedding Retrieval | Tianyu Chen et.al. | 2609.38743v1 | null |
| 2026-09-30 | Concept-Grounded Attention: A Controlled Evaluation of Graph-Injected Attention, Temporal Versioning, and Epistemic Status | Sachin Dev Duggal et.al. | 2609.38684v1 | null |
| 2026-09-29 | Bandits with Multiple Optimal Arms: Minimax Regret and Non-Adaptivity | Kaixuan Ji et.al. | 2609.38659v2 | null |
| 2026-09-29 | Prompt2Skill: Unsupervised Skill Optimization From Natural Language Instructions | Bo Ni et.al. | 2609.38593v1 | null |
Abstracts
KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards
2610.02206v1 by Pengfei Li, Naufal Suryanto, Sicheng Zhang, Muzammal Naseer
LLMs are increasingly applied to cybersecurity workflows, where they are expected to translate analysts' intent into tool invocations. However, existing evaluations focus on knowledge-based assessments or end-to-end agentic tasks, and do not directly measure LLMs' ability to generate executable commands for real-world cybersecurity tools. This gap is critical because cybersecurity operations rely on strict command-line interfaces (CLIs), where minor syntax errors, incorrect flag--value bindings, or argument misordering can invalidate execution. We introduce KaliBench, a fine-grained benchmark and dataset for natural-language--to--CLI translation on Kali Linux, comprising 8,504 query--command pairs spanning 1,642 tools across 23 capability dimensions and 5 security phases. KaliBench is constructed via a manuscript-grounded pipeline with deterministic canonicalization and alias-aware evaluation, enabling precise and reproducible assessment of tool selection and argument construction. To ensure both semantic correctness and practical executability, we develop a multi-stage verification pipeline that combines LLM-based validation, sandboxed terminal execution, and human-in-the-loop refinement. Building on these fine-grained, deterministic signals, KaliBench further enables runtime-free verifiable rewards for training. Across three evaluation modes and 24 configurations of general-purpose and security-focused open-weight models, no open-weight model exceeds 42% exact-command accuracy in the unrestricted setting, highlighting the difficulty of accurate CLI-based cybersecurity tool use without explicit tool hints. We further show that supervised fine-tuning and reinforcement learning with verifiable rewards derived from KaliBench significantly improve an 8B model and achieve performance comparable to a 685B MoE model.
摘要:LLMs 正在越來越多地應用於網絡安全工作流程中,它們被期望將分析師的意圖轉化為工具調用。
然而,現有的評估主要集中在基於知識的評估或端到端的代理任務上,並未直接測量 LLMs 生成可執行命令以供現實世界網絡安全工具使用的能力。
這一差距至關重要,因為網絡安全操作依賴於嚴格的命令行界面 (CLIs),其中微小的語法錯誤、不正確的標誌--值綁定或參數錯序都可能使執行無效。
我們介紹了 KaliBench,這是一個針對 Kali Linux 的自然語言到 CLI 翻譯的細粒度基準和數據集,包含 8,504 個查詢--命令對,涵蓋 1,642 種工具,跨越 23 個能力維度和 5 個安全階段。
KaliBench 是通過一個基於手稿的管道構建的,具有確定性的標準化和別名感知評估,能夠精確且可重複地評估工具選擇和參數構建。
為了確保語義正確性和實際可執行性,我們開發了一個多階段驗證管道,結合了基於 LLM 的驗證、沙盒終端執行和人類參與的精煉。
基於這些細粒度的確定性信號,KaliBench 進一步使得無運行時的可驗證獎勵成為訓練的可能。
在三種評估模式和 24 種通用及安全專注的開放權重模型配置中,沒有任何開放權重模型在不受限制的設置中超過 42% 的精確命令準確率,突顯了在沒有明確工具提示的情況下準確使用基於 CLI 的網絡安全工具的困難。
我們進一步顯示,從 KaliBench 派生的可驗證獎勵的監督微調和強化學習顯著改善了一個 8B 模型,並達到了與 685B MoE 模型相當的性能。
Higher-Order Molecular Grammars for Generative and Foundation Models in Chemistry
2610.02186v1 by Yiming Huang, Yujie Zeng, Vijay Prakash Dwivedi, Simone Foti, Jianmin Wang, Jure Leskovec, Tolga Birdal
Molecular learning models are strongly shaped by their underlying representations. Yet standard sequential and graph formalisms struggle to explicitly encode higher-order topology, such as ring systems and recurring motifs. Existing higher-order representations can capture these structures directly, but they are often computationally demanding and difficult to decode into valid molecules. Here, we introduce Higher-order Grammar Representation (HGR), a principled, topology-aware framework that lifts molecules to combinatorial complexes and parses each complex into a compact sequence of production rules under a context-free higher-order grammar. By serialising higher-order topology into rule sequences, HGR makes these structures directly compatible with standard sequence models, avoiding the computational overhead of explicit higher-order encodings while preserving topological expressiveness. To reduce benchmark bias towards simple ring systems, we construct RingDiv, a ring-enriched benchmark containing 1.18 million molecules, including the curated RingDiv300k subset, and introduce the ring diversity index (RDI) to quantify ring-system coverage. In molecular generation, HGR-based models uniquely combine 100% validity by construction with leading distributional alignment, ranking first in FCD on all five generation benchmarks. In representation learning, HGR-FM achieves the highest mean AUC across seven MoleculeNet benchmarks under both transfer protocols, improving on the strongest baseline by 8.3 and 3.3 AUC points under probing and full fine-tuning, respectively. Collectively, these results establish HGR as an efficient higher-order representation for molecular generation and transferable representation learning.
摘要:分子學習模型受到其基礎表示的強烈影響。
然而,標準的序列和圖形形式在明確編碼高階拓撲方面(如環系統和重複圖案)面臨挑戰。
現有的高階表示可以直接捕捉這些結構,但它們通常計算需求高且難以解碼為有效的分子。
在此,我們介紹高階語法表示(HGR),這是一個原則性、關注拓撲的框架,將分子提升為組合複合體,並將每個複合體解析為在上下文無關的高階語法下的緊湊生成規則序列。
通過將高階拓撲序列化為規則序列,HGR使這些結構與標準序列模型直接兼容,避免了明確高階編碼的計算開銷,同時保留了拓撲表達能力。
為了減少對簡單環系統的基準偏見,我們構建了RingDiv,這是一個包含118萬個分子的環增強基準,包括精心策劃的RingDiv300k子集,並引入環多樣性指數(RDI)來量化環系統的覆蓋範圍。
在分子生成方面,基於HGR的模型獨特地結合了100%的有效性(由構造決定)與領先的分佈對齊,在所有五個生成基準中FCD排名第一。
在表示學習方面,HGR-FM在七個MoleculeNet基準中,在兩種轉移協議下實現了最高的平均AUC,相較於最強基線分別提高了8.3和3.3 AUC點(在探測和完全微調下)。
綜合這些結果,HGR確立了作為分子生成和可轉移表示學習的高效高階表示。
From Knowledge Access to Source Learning: Developing Source-Specific Competence
2610.02150v1 by Lucheng Fu, Kejing Xia, Yiyang Wang, Yiqiao Jin, Jinjin He, Xiyuan Yang, Haoxin Liu, Ye Yu, Haibo Jin, Yijia Xiao, Wenke Lee, B. Aditya Prakash, Haohan Wang
Large language model (LLM) agents increasingly rely on persistent external sources to solve sequences of knowledge-intensive tasks. Existing methods improve how source content is accessed and organized, while agent-memory systems preserve reusable knowledge from prior interactions, but repeated use of the same source is still largely treated as repeated access rather than an opportunity to progressively improve understanding of that source. We study source learning: developing reusable source-specific competence over a persistent authoritative source. We represent this competence with a persistent source model that captures reusable understanding of the source, including how its knowledge is structured, interpreted, and applied. To construct and progressively refine such models, we propose SourceLearn, which combines two complementary learning mechanisms. Self-Directed Source Learning identifies what remains incompletely understood and adaptively revisits the source, while Task-Guided Source Learning uses downstream experience to reveal local representational gaps and recurring needs in how source knowledge should be organized. In both cases, learning signals determine what should be reconsidered, while persistent updates are reconstructed from the authoritative source. Across five benchmarks and three LLM backends, SourceLearn achieves the best performance in 13 of 15 settings, with gains of up to 22.6 points over Hybrid RAG and substantial overall improvements over static source representations and experience-based memory baselines.
摘要:大型語言模型(LLM)代理越來越依賴持久的外部來源來解決一系列知識密集型任務。現有方法改善了如何訪問和組織來源內容,而代理記憶系統則保留了來自先前互動的可重用知識,但對同一來源的重複使用仍然主要被視為重複訪問,而不是逐步改善對該來源理解的機會。我們研究來源學習:在持久的權威來源上發展可重用的來源特定能力。我們用一個持久的來源模型來表示這種能力,該模型捕捉了對來源的可重用理解,包括其知識的結構、解釋和應用方式。為了構建和逐步完善這樣的模型,我們提出了SourceLearn,該模型結合了兩種互補的學習機制。自我導向來源學習識別尚未完全理解的內容並適應性地重新訪問來源,而任務引導來源學習則利用下游經驗揭示如何組織來源知識的局部表徵差距和重複需求。在這兩種情況下,學習信號決定了應該重新考慮的內容,而持久更新則是從權威來源重建的。在五個基準和三個LLM後端中,SourceLearn在15個設置中的13個中實現了最佳性能,與Hybrid RAG相比,增益高達22.6點,並且在靜態來源表示和基於經驗的記憶基準上有顯著的整體改進。
Mem++: Non-Destructive Memory for Long-Term Organizational LLM Agents
2610.02002v1 by Ahmad Yehia, Aly O. Abdelkareem, Islam Ahmed, Hesham Omran, Khaled Alashmouny, Christian Claudel, Abduallah Mohamed
Large Language Model (LLM) agents now take part in organizational work, where many authors record decisions across documents over months. Because a revised decision arrives as a new document rather than an edit, answering a question requires knowing which version held at a given time. However, most memory systems compress the record at write time. By distilling each document into facts, notes or graph edges, these methods fix what can be answered before any question is asked. To address this, we propose Mem++, a non-destructive memory framework shifting from write-time distillation to read-time selection. Mem++ stores every document whole with its date and author, and it calls no generative model at write time. At read time, it retrieves only documents dated up to the time a question asks about and fuses lexical and semantic rankings. Unlike systems that overwrite older versions, Mem++ keeps them and leaves the choice to the answering model. Evaluations on the organizational benchmark OrgMemBench demonstrate that Mem++ surpasses the strongest memory system baseline by 8.0 to 13.1 points across two answering models. With gpt-4.1-mini, it also achieves the best overall score, 2.6 points above RAG. In addition, Mem++ achieves the best average LLM-judge score on LoCoMo and ranks second on LongMemEval-S, behind only its entity-graph variant. Code for benchmark evaluation is available at https://github.com/AIDAChip-Inc/mem-plus-plus.
摘要:大型語言模型(LLM)代理現在參與組織工作,許多作者在數月內記錄決策於文件中。因為修訂的決策以新文件的形式出現,而不是編輯,因此回答問題需要知道在特定時間持有的是哪個版本。然而,大多數記憶系統在寫入時會壓縮記錄。通過將每個文件提煉成事實、筆記或圖邊,這些方法在任何問題被提出之前固定了可以回答的內容。為了解決這個問題,我們提出了Mem++,這是一個非破壞性的記憶框架,從寫入時的提煉轉向讀取時的選擇。Mem++ 將每個文件完整地存儲,並附上日期和作者,並且在寫入時不調用任何生成模型。在讀取時,它僅檢索在問題詢問時的日期之前的文件,並融合詞彙和語義排名。與覆蓋舊版本的系統不同,Mem++ 保留它們,並將選擇權留給回答模型。在組織基準測試 OrgMemBench 上的評估顯示,Mem++ 在兩個回答模型中超越了最強記憶系統基線 8.0 到 13.1 分。使用 gpt-4.1-mini 時,它還獲得了最佳整體分數,比 RAG 高出 2.6 分。此外,Mem++ 在 LoCoMo 上獲得了最佳平均 LLM-judge 分數,並在 LongMemEval-S 中排名第二,僅次於其實體圖變體。基準評估的代碼可在 https://github.com/AIDAChip-Inc/mem-plus-plus 獲得。
Can AI Oversight Be Zero Knowledge?
2610.01995v1 by Alessandro Chiesa, Ziyi Guan, Burcu Yildiz
AI systems increasingly produce outputs from confidential data, such as a fitness-for-duty assessment from medical records or the predicted properties of a drug candidate from its secret structure. It is important to verify that such outputs are correct without revealing the underlying data. A recent line of work studies verification of AI outputs via interactive proofs and debate for oracle-aided computation, where correctness may depend on an oracle such as human judgment, a physical experiment, or the web. These works focus on verification by a verifier that runs much faster than the computation. However, such efficient verification is impossible for general oracle-aided computation, and these works therefore rely on additional assumptions. We focus instead on privacy: allowing the verifier to run in time polynomial in the computation, we ask whether interactive arguments for oracle-aided computation can be zero knowledge, so that the verifier learns nothing about the confidential data beyond the correctness of the output. We prove that, in general, they cannot. In the random oracle model, there are no zero-knowledge proofs for all oracle-aided computations, even if both the prover and the verifier are allowed to run much longer than the computation itself. The impossibility extends to debate, a canonical model for scalable oversight. On the positive side, we show that if the oracle attaches a cryptographic signature to each of its answers, then every oracle-aided computation can be verified in zero knowledge with an efficient prover and verifier, assuming only collision-resistant hash functions. Beyond privacy, this also gives an alternative approach to scalable oversight that relies neither on an honest opponent, as in debate, nor on the robustness of the computation, as in prior single-prover protocols.
摘要:AI 系統越來越多地從機密數據中產生輸出,例如從醫療記錄中進行的適任性評估或從其秘密結構中預測的藥物候選物的性質。
驗證這些輸出是否正確而不透露底層數據是很重要的。
最近的一系列研究通過互動證明和辯論研究 AI 輸出的驗證,用於有 oracle 輔助的計算,其中正確性可能依賴於 oracle,例如人類判斷、物理實驗或網絡。
這些研究專注於由運行速度遠快於計算的驗證者進行的驗證。
然而,對於一般的有 oracle 輔助計算,這樣的高效驗證是不可能的,因此這些研究依賴於額外的假設。
我們則專注於隱私:允許驗證者在計算的多項式時間內運行,我們詢問有 oracle 輔助計算的互動論證是否可以是零知識的,以便驗證者不會學到關於機密數據的任何信息,除了輸出的正確性。
我們證明,通常情況下,它們是不可能的。
在隨機 oracle 模型中,對於所有有 oracle 輔助的計算,沒有零知識證明,即使證明者和驗證者都被允許運行的時間遠超過計算本身。
這種不可能性擴展到辯論,這是一個可擴展監督的典型模型。
從積極的一面來看,我們展示了如果 oracle 為其每個答案附加加密簽名,那麼每個有 oracle 輔助的計算都可以在零知識中進行驗證,並且有高效的證明者和驗證者,只假設碰撞抗性哈希函數。
除了隱私之外,這還提供了一種可擴展監督的替代方法,既不依賴於誠實的對手(如辯論中),也不依賴於計算的穩健性(如以前的單證明者協議中)。
Latent JEPA: Abstract Future Prediction for Latent Reasoning in Chemistry
2610.01947v1 by Xinjian Zhao, Yaoyao Xu, Xuemin Chen, Xiaozhuang Song, Tianshu Yu
Large language models offer a promising foundation for chemical reasoning, bringing together chemical knowledge and multistep problem solving. Chemical intuition can provide an initial sense of plausible outcomes before the details of a solution are fully worked out. Inspired by how such expectations complement explicit analysis, we study how continuous latent thoughts can be trained to anticipate informative aspects of future solutions without verbalizing every intermediate step. We introduce Latent JEPA, a framework that combines autoregressive learning with joint-embedding prediction of one or more future views. For chemical reasoning, we develop textual and molecular prediction objectives that connect latent thoughts to both subsequent reasoning and molecular outcomes. Experiments on ChemCoTBench show gains in molecular optimization and on several editing and reaction metrics. Representation analyses show that future prediction makes latent thoughts more informative about molecular outcomes and strengthens their correspondence with chemical structure. These findings support abstract future prediction as a learning principle for connecting continuous latent reasoning with scientific outcomes.
摘要:大型語言模型為化學推理提供了一個有前景的基礎,將化學知識和多步驟問題解決結合在一起。化學直覺能在解決方案的細節完全展開之前,提供一種合理結果的初步感知。受到這種期望如何補充明確分析的啟發,我們研究如何訓練連續潛在思維,以預測未來解決方案的資訊性方面,而不需要逐步口頭表達每一個中間步驟。我們介紹了潛在JEPA,一個將自回歸學習與一個或多個未來視圖的聯合嵌入預測相結合的框架。針對化學推理,我們開發了文本和分子預測目標,將潛在思維與後續推理和分子結果連接起來。在ChemCoTBench上的實驗顯示,在分子優化以及幾個編輯和反應指標上都有提升。表徵分析顯示,未來預測使潛在思維對分子結果的資訊性更強,並加強了它們與化學結構的對應性。這些發現支持將抽象的未來預測作為一種學習原則,以連接連續的潛在推理與科學結果。
Mapping the RAG Landscape: A Four Axis Taxonomy of Efficiency, Defense, Interactivity, and Reasoning
2610.01936v1 by Meghana Sunil, Shravya V, Shravan Venkatraman, Joe Dhanith PR
Large Language Models (LLMs) have demonstrated remarkable fluency across many tasks but remain limited by their static, parameter bound knowledge and their susceptibility to hallucinating information. Retrieval Augmented Generation (RAG) addresses these issues by incorporating external retrieval into the generation process, grounding model outputs in verifiable and up to date sources. While prior surveys primarily focus on core RAG architectures and standard pipelines, recent research explores broader challenges and capabilities that extend beyond these foundational designs. This survey provides a consolidated and structured examination of contemporary RAG developments, organizing the field into a four axis taxonomy: improving retrieval efficiency, strengthening robustness and security, supporting user driven and interactive workflows, and enabling multi step or complex reasoning. We formalize key components of the RAG framework and review methods spanning dense and sparse retrieval, fusion strategies, embedding optimizations, and reinforcement learning based retrieval policies, highlighting how these advances influence practical deployment and system design. We also synthesize evaluation practices, domain specific applications, and architectural variants such as Naive, Advanced, and Modular RAG. Finally, we outline persistent challenges related to retrieval quality, reliability, domain adaptation, scalability, and explainability, and identify opportunities for building RAG systems that are more reliable, adaptable, and transparent.
摘要:大型語言模型(LLMs)在許多任務中展現了卓越的流暢性,但仍然受到靜態的、參數限制的知識以及對虛假信息的易感性的限制。檢索增強生成(RAG)通過將外部檢索納入生成過程來解決這些問題,使模型輸出基於可驗證且最新的來源。雖然之前的調查主要集中在核心RAG架構和標準流程上,但最近的研究探討了超越這些基礎設計的更廣泛挑戰和能力。本調查提供了一個當代RAG發展的綜合和結構化檢視,將該領域組織為四個軸向的分類法:提高檢索效率、加強穩健性和安全性、支持用戶驅動和互動工作流程,以及實現多步驟或複雜推理。我們正式化了RAG框架的關鍵組件,並回顧了涵蓋密集和稀疏檢索、融合策略、嵌入優化和強化學習基於檢索政策的方法,強調這些進展如何影響實際部署和系統設計。我們還綜合了評估實踐、特定領域的應用以及如Naive、Advanced和Modular RAG等架構變體。最後,我們概述了與檢索質量、可靠性、領域適應性、可擴展性和可解釋性相關的持續挑戰,並確定了構建更可靠、可適應和透明的RAG系統的機會。
From Network Intrusion Detection to Blockchain-Backed Endpoint Detection and Response: Mapping the Landscape of Decentralized Detection-and-Response Architectures
2610.01872v1 by Yahya Shahsavari, Sara Rouhani, Kaiwen Zhang
While the literature on blockchain-assisted intrusion detection and prevention systems (IDS/IPS) for Internet of Things (IoT) and Industrial Internet of Things (IIoT) networks is mature, existing systematic reviews suffer from two critical limitations: they overlook the structural shift toward modern Endpoint Detection and Response (EDR) and Extended Detection and Response (XDR) architectures, and they conflate blockchain's distinct functional roles into a single monolithic category. This Systematization of Knowledge (SoK) addresses these gaps by proposing a three-axis taxonomy that classifies proposals by detection-system class (NIDS, HIDS, EDR/XDR), blockchain functional role, and response-automation maturity. Synthesizing research published in high-impact venues between 2019 and 2026, we provide a rigorous gap analysis exposing why a genuine per-endpoint blockchain-anchored response loop remains nearly nonexistent due to latency, deployment, and community mismatches. Furthermore, we evaluate structural, cross-cutting challenges persisting across the literature, including consensus latency on constrained devices, post-quantum cryptographic vulnerability, smart-contract attack surfaces, and the adversarial vulnerability of evolving LLM-based detection engines. Finally, we outline a comprehensive research agenda centered on hybrid on-chain/off-chain orchestration to bridge the gap between decentralized trust and rapid response automation.
摘要:雖然有關區塊鏈輔助的入侵檢測和預防系統(IDS/IPS)在物聯網(IoT)和工業物聯網(IIoT)網絡中的文獻已相當成熟,但現有的系統性評估存在兩個關鍵限制:它們忽視了向現代端點檢測與響應(EDR)和擴展檢測與響應(XDR)架構的結構性轉變,並且將區塊鏈的不同功能角色混淆為一個單一的整體類別。這項知識系統化(SoK)通過提出一個三軸分類法來解決這些空白,該分類法根據檢測系統類別(NIDS、HIDS、EDR/XDR)、區塊鏈功能角色和響應自動化成熟度對提案進行分類。綜合2019年至2026年間在高影響力期刊上發表的研究,我們提供了一個嚴謹的差距分析,揭示了為何真正的每個端點區塊鏈錨定響應循環幾乎不存在,原因在於延遲、部署和社群不匹配。此外,我們評估了文獻中持續存在的結構性、跨領域挑戰,包括在受限設備上的共識延遲、後量子密碼學脆弱性、智能合約攻擊面以及不斷演變的基於LLM的檢測引擎的對抗性脆弱性。最後,我們概述了一個以混合鏈上/鏈下協同為中心的全面研究議程,以彌合去中心化信任與快速響應自動化之間的鴻溝。
Walking the Embedding Space: Datastore Extraction from Multimodal RAG
2610.01871v1 by Maria Carmen Jica, Ali Satvaty, Suzan Verberne, Fatih Turkmen
Multimodal Retrieval-Augmented Generation (MRAG) has emerged as a reliable and cost-effective technique of grounding the generative capabilities of Multimodal Large Language Models (MLLMs) into relevant, up-to-date, external knowledge. Despite presenting several benefits, such as reducing hallucinatory behavior, they also introduce new attack surfaces, including leakage of private information and vulnerabilities against data extraction attacks. In this paper, we introduce $\immrag$, an adaptive and automatic data extraction attack procedure operating in a black box setting against \emph{image-returning} MRAG, a configuration in which the retrieved visual artifact is itself the response. Each query blends an attacker-held shadow image with an image already recovered from the system, and relevance-weighted resampling steers subsequent queries towards regions of the embedding space that still yield novel retrievals. Unlike current extraction attacks that aim to persuade the model towards data leakage by placing a malicious query as a textual prompt, $\immrag$ embeds the malicious instructions inside a user-given input image. We evaluate $\immrag$ on three plausible and distinct real-world scenarios: medical assistant, document-focused helper and general purpose tool. The experiments involve the study of the effectiveness of the attack on multiple CLIP-family retrievers, as well as the impact of various generators. A single 2500-query run reconstructs up to 611 distinct radiology images, 566 document scans and 416 general-purpose images under local-feature correspondence, and reaches up to $5.6\times$ as many distinct datastore items as a non-adaptive baseline. Our results show the urgent need for safeguards specifically designed for multimodal data.
摘要:多模態檢索增強生成(MRAG)已成為將多模態大型語言模型(MLLMs)的生成能力與相關的、最新的外部知識相結合的一種可靠且具成本效益的技術。儘管它提供了幾個好處,例如減少幻覺行為,但它們也引入了新的攻擊面,包括私密信息洩漏和對數據提取攻擊的脆弱性。
在本文中,我們介紹了 $\immrag$,這是一種適應性和自動化的數據提取攻擊程序,針對 \emph{圖像返回} MRAG 在黑箱環境中運作,這是一種檢索的視覺工件本身就是回應的配置。每個查詢將攻擊者持有的影像與系統中已恢復的影像混合,並且相關性加權重採樣引導後續查詢朝向仍能產生新穎檢索的嵌入空間區域。與目前旨在通過將惡意查詢作為文本提示來說服模型進行數據洩漏的提取攻擊不同,$\immrag$ 將惡意指令嵌入用戶提供的輸入影像中。我們在三個合理且不同的現實場景中評估了 $\immrag$:醫療助手、文件專注助手和通用工具。實驗涉及對多個 CLIP 家族檢索器的攻擊有效性以及各種生成器的影響進行研究。一次 2500 次查詢的運行重建了多達 611 幅不同的放射學影像、566 幅文件掃描和 416 幅通用影像,根據局部特徵對應,並達到高達 $5.6\times$ 的不同數據庫項目數量,相較於非適應性基準。我們的結果顯示出對專門為多模態數據設計的安全措施的迫切需求。
Detecting Inconsistencies in Model Specifications with LLM-as-Verifier Reasoning
2610.01847v1 by Zichen Xie, Mrigank Pawagi, Lize Shao, Yang Hu, Wenxi Wang
Model specifications define how large language models (LLMs) should behave, guiding alignment training, inference-time behavior, and evaluation. Yet these specifications may themselves contain defects: two individually reasonable principles may prescribe incompatible behavior when applied to the same situation, leaving no response that satisfies both. Detecting such inconsistencies is challenging. Formalizing natural-language specifications risks losing subtle distinctions, while behavior-based testing cannot reliably distinguish specification defects from differences in model behavior. We introduce VeriSpec, the first approach to directly detect inconsistencies in model specifications by auditing the specification text itself. Our key insight is to preserve the specification in natural language while using an LLM as a verifier. VeriSpec extracts structured, context-aware rules, constructs a topic-guided graph to cluster behaviorally related rules at the same authority level, and applies LLM-as-verifier reasoning to detect inconsistencies. Applying VeriSpec to the OpenAI Model Spec, we extract 405 rules and manually validate five inconsistencies, all reported to its developers, who responded positively and have initiated internal discussions. Compared with five baselines, VeriSpec identifies the most validated inconsistencies, achieves the highest precision (38.5%), and incurs the lowest cost per validated inconsistency ($11.12). These results establish direct specification auditing as a practical complement to behavioral alignment evaluation, catching defects at the source before they shape any model. The code is available at https://github.com/HIPREL-Group/VeriSpec.
摘要:模型規範定義了大型語言模型(LLMs)應該如何運作,指導對齊訓練、推理時的行為和評估。
然而,這些規範本身可能包含缺陷:兩個各自合理的原則在應用於相同情境時可能會規定不相容的行為,導致沒有任何回應能同時滿足兩者。
檢測這種不一致性是具有挑戰性的。
將自然語言規範形式化可能會失去微妙的區別,而基於行為的測試則無法可靠地區分規範缺陷與模型行為的差異。
我們引入了VeriSpec,這是第一種通過審核規範文本本身直接檢測模型規範中不一致性的方法。
我們的關鍵見解是保留自然語言中的規範,同時使用LLM作為驗證者。
VeriSpec提取結構化的、上下文感知的規則,構建主題引導的圖以聚類同一權威級別下行為相關的規則,並應用LLM作為驗證者的推理來檢測不一致性。
將VeriSpec應用於OpenAI模型規範,我們提取了405條規則並手動驗證了五個不一致性,所有這些都已報告給其開發者,開發者對此做出了積極回應並已啟動內部討論。
與五個基準相比,VeriSpec識別了最多的經過驗證的不一致性,達到了最高的精確度(38.5%),並且每個經過驗證的不一致性的成本最低($11.12)。
這些結果確立了直接規範審核作為行為對齊評估的實用補充,在缺陷影響任何模型之前,及時捕捉到缺陷。
代碼可在 https://github.com/HIPREL-Group/VeriSpec 獲得。
Code Owns the Simulation, Jev Owns the Evaluation
2610.01834v1 by Yaodong Yang, Hongyao Tang, Yi Ma, Xingyu Fan, Weixun Wang, Jinpeng Li, Tianpei Yang
Judgment models such as \jev{} return, in a single call and without reasoning text, a probability for each described option. This makes them attractive as an agent's action-selection layer, but it is unclear which decisions they can be trusted with. We test \jev{} on reflection tests, one-shot matrix games, the text game ALFWorld and robot control, and find a sharp boundary. \jev{} succeeds when the right option can be judged from what the input describes, which we call \emph{evaluation}. Specifically, it solves 99\% of the counterintuitive Cognitive Reflection Test questions. However, it fails when the right option depends on \emph{simulation} (i.e., predicting something not in the input), such as the opponent's action or the subgoal that must come first. In games, \jev{} plays suboptimally as if its rational opponent acted at random, because the opponent's action is not given. In ALFWorld, \jev{} favors commands that mention an object or place named in the task description. For example, given the task ``put a clean knife in the drawer'', \jev{} carries an unwashed knife straight to the drawer instead of first washing it at the sink. Surprisingly, many of these failures are not due to a lack of knowledge. Asked separately what the opponent will do, \jev{} usually answers correctly, and it responds well given the opponent's action. It fails when one call must both perform the simulation and evaluate based on it. This suggests letting code make the prediction or simulation. When code supplies it, such as a lookahead in ALFWorld and physics simulation in robot control, \jev{} becomes an expert controller through its general evaluation ability.
摘要:判斷模型如 \jev{} 在單次呼叫中返回每個描述選項的概率,且不需要推理文本。這使得它們作為代理的行動選擇層變得具有吸引力,但尚不清楚它們可以信任哪些決策。我們在反思測試、一回合矩陣遊戲、文本遊戲 ALFWorld 和機器人控制上測試 \jev{},並發現了一個明確的邊界。當正確選項可以從輸入描述中判斷時,我們稱之為 \emph{評估},\jev{} 成功地解決了 99\% 的反直覺認知反思測試問題。然而,當正確選項依賴於 \emph{模擬}(即預測輸入中不存在的事物)時,它則失敗,例如對手的行動或必須先完成的子目標。在遊戲中,\jev{} 表現得次優,彷彿其理性的對手隨機行動,因為對手的行動並未給出。在 ALFWorld 中,\jev{} 偏好提到任務描述中物體或地點的命令。例如,給定任務「將乾淨的刀放入抽屜」,\jev{} 直接將未洗的刀帶到抽屜,而不是先在水槽中清洗它。令人驚訝的是,這些失敗中的許多並非因為缺乏知識。單獨詢問對手會做什麼時,\jev{} 通常能正確回答,並且在給定對手的行動時反應良好。當一次呼叫必須同時執行模擬並基於此進行評估時,它則失敗。這表明應讓代碼進行預測或模擬。當代碼提供這些信息時,例如在 ALFWorld 中的前瞻和在機器人控制中的物理模擬,\jev{} 通過其一般評估能力成為專家控制器。
The Asymptotics of Language Model Alignment with Memory
2610.01828v1 by Haricharan Balasundaram, V. Arvind Rameshwar
Language model (LM) alignment broadly aims to perturb a given LM $Q$ into an aligned LM $q$ such that i) the outputs produced by $q$ and $Q$ are 'close' in probability, ii) $q$ has a higher expected reward than $Q$. Two common techniques for LM alignment are: KL-constrained RL, which requires knowledge of the LM distribution and is computationally expensive, and the best-of-$n$ algorithm, which requires only sampling from the LM. The work of Yang et al. established asymptotic closeness between the distributions produced by the two alignment methods for an $m$--length i.i.d. token sequence output by the LM, in the limit as $m$ increases to infinity. However, the i.i.d. assumption is not representative of practical LMs, whose output sequences often have memory. In this paper, we extend the asymptotic closeness result to the case when the $m$--length token sequence outputted by the LM is Markovian. Further, for finite-length output sequences -- particularly, when $m=1$ -- we provide a complete characterization of LM distributions and reward functions for which the KL-divergence between the distributions produced by the two alignment methods is zero -- a question first posed in Yang et al.
摘要:語言模型(LM)對齊的廣泛目標是將給定的 LM $Q$ 轉變為一個對齊的 LM $q$,使得 i) $q$ 和 $Q$ 所產生的輸出在概率上是「接近」的,ii) $q$ 的期望獎勵高於 $Q$。兩種常見的 LM 對齊技術是:KL 約束強化學習,這需要對 LM 分佈的了解並且計算上昂貴,以及最佳的 $n$ 算法,這僅需要從 LM 中進行取樣。Yang 等人的研究確立了在 $m$ 長度的獨立同分佈(i.i.d.)標記序列的情況下,兩種對齊方法所產生的分佈之間的漸近接近性,當 $m$ 增加到無限大時。然而,i.i.d. 假設並不代表實際的 LM,因為它們的輸出序列通常具有記憶性。在本文中,我們將漸近接近性結果擴展到 LM 輸出的 $m$ 長度標記序列為馬爾可夫過程的情況。此外,對於有限長度的輸出序列——特別是當 $m=1$ 時——我們提供了 LM 分佈和獎勵函數的完整特徵描述,對於這些情況,兩種對齊方法所產生的分佈之間的 KL 散度為零——這是一個最初由 Yang 等人提出的問題。
A Matryoshka Hierarchical RAG for Efficient Multi-Hop Question Answering
2610.01767v1 by Gianluca Bonifazi, Christopher Buratti, Michele Marchetti, Federica Parlapiano, Giulia Quaglieri, Davide Traini, Domenico Ursino, Luca Virgili
Retrieval-Augmented Generation (RAG) systems for multi-hop Question Answering (QA) must balance retrieval quality with computational cost. This cost is incurred during indexing time, through the use of expensive Knowledge Graphs (KGs) or Large Language Models (LLMs) to generate summaries, or during querying, through iterative LLM-driven retrieval. To reduce it while maintaining retrieval quality, we present MatRAG, a hierarchical framework that combines RAG systems with Matryoshka Representation Learning (MRL). MatRAG addresses both kinds of cost by aligning the semantic hierarchy of a clustering structure with the nested structure of MRL. Specifically, it organizes the corpus of documents into a Directed Acyclic Graph (DAG) of clusters with progressively coarser granularity. Each level is indexed by a lower Matryoshka dimension. MatRAG pairs an iterative, top-down traversal of the DAG with an entity-driven mechanism that controls the hop budget and re-ranks candidates. We evaluated MatRAG on three standard multi-hop QA benchmarks against seven representative baselines. MatRAG outperforms its strongest competitors in terms of retrieval quality; furthermore, it reduces indexing costs by avoiding KG construction and LLM-based summarization, and lowers query-time costs through dimension-aware similarity.
摘要:檢索增強生成(RAG)系統在多跳問題回答(QA)中必須平衡檢索質量與計算成本。這個成本在索引時產生,通過使用昂貴的知識圖譜(KG)或大型語言模型(LLM)來生成摘要,或在查詢時,通過迭代的LLM驅動檢索。為了在保持檢索質量的同時降低成本,我們提出了MatRAG,一個將RAG系統與馬特里奧什卡表示學習(MRL)相結合的分層框架。MatRAG通過將聚類結構的語義層次與MRL的嵌套結構對齊,解決了這兩種成本。具體而言,它將文檔語料庫組織成一個具有逐漸粗糙粒度的有向無環圖(DAG)聚類。每個層級由較低的馬特里奧什卡維度進行索引。MatRAG將DAG的迭代自上而下遍歷與一種驅動實體的機制相結合,該機制控制跳躍預算並重新排名候選者。我們在三個標準的多跳QA基準上評估了MatRAG,並與七個代表性的基準進行比較。MatRAG在檢索質量方面超越了其最強的競爭對手;此外,它通過避免KG構建和基於LLM的摘要來降低索引成本,並通過維度感知相似性來降低查詢時間成本。
Task-Oriented Rank Adaptation for Continual Learning in Text Classification
2610.01702v1 by Rey Sanchez Lopez, Eduardo Morales Manzanares, Hugo Jair Escalante
Continual learning (CL) in text classification faces two critical challenges: catastrophic forgetting and negative transfer across sequential tasks. Parameter-Efficient Fine-Tuning (PEFT) methods such as LoRA enable efficient adaptation by learning low-rank updates of the model parameters. However, these compact representations are normally trained in isolation, limiting their reuse across related tasks. We introduce Task-Oriented Rank Adaptation (TORA), a geometric routing framework that leverages the low-rank structure of LoRA adapters to decide whether to transfer knowledge from the most compatible expert (Boosting) or isolate the new task (Shielding) based on structural similarity. Evaluated across 15 diverse text classification benchmarks, TORA consistently avoids harmful routing decisions: compatible tasks exceed their isolated performance while reducing training time, and structurally distant tasks are protected from interference with no loss in accuracy. With a single geometric threshold and no reliance on task identities or predefined sequences, TORA provides a simple and effective approach for dynamic adapter routing in sequential text classification systems.
摘要:持續學習(CL)在文本分類中面臨兩個關鍵挑戰:災難性遺忘和在序列任務中的負轉移。參數高效微調(PEFT)方法如 LoRA 通過學習模型參數的低秩更新來實現高效適應。然而,這些緊湊的表示通常是在孤立的情況下訓練的,限制了它們在相關任務中的重用。我們引入了任務導向秩適應(TORA),這是一個幾何路由框架,利用 LoRA 適配器的低秩結構來決定是從最兼容的專家(提升)轉移知識,還是根據結構相似性隔離新任務(保護)。在 15 個不同的文本分類基準上進行評估,TORA 始終避免有害的路由決策:兼容任務的表現超過其孤立的性能,同時減少訓練時間,而結構上相距較遠的任務則受到保護,沒有準確度損失。TORA 以單一的幾何閾值運作,且不依賴於任務身份或預定序列,為序列文本分類系統中的動態適配器路由提供了一種簡單而有效的方法。
Iterative Policy Refinement through Semantic Rollout Analysis
2610.01652v1 by Feiyu Gavin Zhu, Qi Xu, Zhifei Deng, Zhigang Hua, Luke Simon, Jean Oh, Reid Simmons
Structured policies improve efficiency, robustness, and interpretability in imitation learning by introducing task-specific inductive bias, but existing structure generation methods rely either on extensive human input or on static domain knowledge encoded in LLMs, which may be inconsistent with the expert demonstrations. We propose a closed-loop framework that iteratively refines structured policies using LLM-guided analysis of policy rollouts. By logging rollouts as semantically meaningful tabular data and prompting the LLM to generate diagnostic analysis code, our method identifies suboptimalities in the policy structure and iteratively corrects them without requiring human instruction. Experiments on car racing and door opening tasks show that our approach improves imitation learning performance by up to 15% over zero-shot LLM-generated structures and requires 75% less compute to achieve the same reinforcement learning performance. These results demonstrate that tabular rollout analysis provides an effective feedback signal to align LLM-generated policy structures with expert demonstrations, and we can utilize it to generate good policy structures automatically.
摘要:結構化政策透過引入特定任務的歸納偏見來提升模仿學習的效率、穩健性和可解釋性,但現有的結構生成方法要麼依賴大量的人類輸入,要麼依賴於編碼在大型語言模型(LLMs)中的靜態領域知識,這可能與專家的示範不一致。我們提出了一個閉環框架,通過使用LLM引導的政策展開分析來迭代地改進結構化政策。通過將展開記錄為語義上有意義的表格數據,並提示LLM生成診斷分析代碼,我們的方法識別出政策結構中的次優性,並在不需要人類指導的情況下進行迭代修正。在賽車和開門任務上的實驗顯示,我們的方法在模仿學習性能上比零樣本LLM生成的結構提高了多達15%,並且需要75%更少的計算來達到相同的強化學習性能。這些結果表明,表格展開分析提供了一個有效的反饋信號,以使LLM生成的政策結構與專家示範對齊,我們可以利用它自動生成良好的政策結構。
Managing Context and Communication in Distributed Agentic UAV Swarms
2610.01569v1 by Andrea Iannoli, Ivan Zyrianoff, Angelo Trotta, Lorenzo Gigli, Marco Di Felice
Unmanned aerial vehicle (UAV) swarms increasingly rely on language-model agents to provide adaptive mission-level reasoning in uncertain environments. Fully distributed control, in which each UAV hosts an independent Small Language Model (SLM), removes reliance on a centralized coordinator but introduces an information-management problem: long-running interaction histories can degrade the reasoning context, while indiscriminate information dissemination increases communication and inference overhead. We address these challenges with a distributed UAV-agent architecture that enables continuous local SLM control through an event-driven reason-act-observe lifecycle. Runtime knowledge is represented as structured atomic notes and organized into core, local, and peer-specific memory. A deterministic interest-aware gossip engine selectively disseminates these notes according to recipient-specific semantic novelty and recency. We evaluate the architecture using ten UAVs in a simulated search-and-rescue mission. Our approach completes all experimental runs, whereas unrestricted flooding messages completes only 70-85\%, and delegating forwarding decisions to the SLM prevents mission completion in every run. Compared with unrestricted flooding, our approach approximately halves inference-token consumption, reduces transmitted data, and achieves lower survivor-count error.
摘要:無人機(UAV)群體越來越依賴語言模型代理在不確定環境中提供自適應的任務級推理。完全分散的控制中,每個UAV都擁有一個獨立的小型語言模型(SLM),這消除了對集中協調者的依賴,但引入了一個信息管理問題:長期的互動歷史可能會削弱推理上下文,而不加區別的信息傳播則增加了通信和推理的開銷。我們通過一種分散的UAV代理架構來解決這些挑戰,該架構通過事件驅動的推理-行動-觀察生命週期實現持續的本地SLM控制。運行時知識以結構化的原子筆記形式表示,並組織成核心、本地和對等特定的記憶。一個確定性的興趣感知八卦引擎根據接收者特定的語義新穎性和時效性選擇性地傳播這些筆記。我們使用十架UAV在模擬的搜索和救援任務中評估該架構。我們的方法完成了所有實驗運行,而不受限制的洪水消息僅完成了70-85\%,將轉發決策委託給SLM則在每次運行中都阻止了任務的完成。與不受限制的洪水相比,我們的方法大約減半了推理令牌的消耗,減少了傳輸數據,並實現了較低的生還者計數誤差。
From Rules to Neural Graphs: Scalable Structured Prediction for Patent Prior Art Search
2610.01553v1 by Nikolai Zenovkin, Sebastian Björkqvist
Patent search requires processing documents routinely exceeding tens of thousands of tokens. Most neural retrieval approaches operate on truncated inputs, limiting their effectiveness. Graph-based retrieval addresses this by representing each patent as a structured invention graph, but constructing these graphs relies on brittle rule-based parsers. We present the neural parser, which adapts biaffine attention from dependency parsing to predict invention graphs directly from patent text. Our local biaffine attention restricts pairwise scoring to a sliding window, reducing complexity from $O(n^2)$ to $O(n \cdot w)$. Since local and global scoring share the same weights, the model trains on short sequences and deploys on documents exceeding 40,000 tokens without retraining. Distilled from 1 million rule-parsed documents, it surpasses its teacher at 3$\times$ lower inference cost: neural graphs improve citation recall by 0.5% on short queries and 1.1% on full documents in a downstream Graph Transformer retrieval system.
摘要:專利搜尋需要處理的文件通常超過數萬個標記。大多數神經檢索方法在截斷的輸入上運作,限制了它們的有效性。基於圖的檢索通過將每個專利表示為結構化的發明圖來解決這個問題,但構建這些圖依賴於脆弱的基於規則的解析器。我們提出了神經解析器,它將雙仿射注意力從依賴解析中適應,以直接從專利文本預測發明圖。我們的局部雙仿射注意力將成對評分限制在滑動窗口內,將複雜度從 $O(n^2)$ 降低到 $O(n \cdot w)$。由於局部和全局評分共享相同的權重,該模型在短序列上進行訓練,並在不重新訓練的情況下部署在超過 40,000 個標記的文件上。從 100 萬個規則解析的文件中提煉出來,它在推理成本上超越了其教師,降低了 3$\times$:神經圖在下游圖Transformer檢索系統中,對於短查詢提高了 0.5% 的引用召回率,對於完整文件提高了 1.1%。
Auto-Formalizing Neuro-Symbolic Predictors
2610.01519v1 by Samuele Bortolotti, Weixin Chen, Han Zhao, Andrea Passerini, Stefano Teso, Antonio Vergari
Neuro-Symbolic (NeSy) predictors incorporate prior knowledge into the prediction process of neural networks, ensuring that outputs satisfy specified constraints, making them particularly suitable for high-stakes applications where compliance with domain knowledge is essential. A key bottleneck in this paradigm is the acquisition of symbolic constraints: encoding domain knowledge into logical formulas remains a manual and expert-intensive process. In this work, we investigate the extent to which auto-formalization via LLMs can systematically translate textual knowledge into symbolic knowledge that can be plugged into NeSy predictors. To this end, we introduce auto-nesy-bench, a new benchmark for evaluating constraint formalization and its impact on downstream accuracy of NeSy predictors. Through an extensive evaluation across several domains, we find that LLMs can formalize constraints to a meaningful extent, generating formulas that are often similar to those provided by human experts. Moreover, when the generated formulas are syntactically valid, they can lead to high-quality downstream predictions. The code and benchmark are available at https://unitn-sml.github.io/auto-nesy-bench/.
摘要:神經符號(NeSy)預測器將先前的知識納入神經網絡的預測過程中,確保輸出滿足特定的約束,使其特別適合於對領域知識遵循至關重要的高風險應用。這一範式中的一個關鍵瓶頸是符號約束的獲取:將領域知識編碼為邏輯公式仍然是一個手動且需要專家的過程。在本研究中,我們探討自動形式化通過大型語言模型(LLMs)在多大程度上可以系統性地將文本知識轉換為可以插入NeSy預測器的符號知識。為此,我們引入了auto-nesy-bench,一個新的基準,用於評估約束形式化及其對NeSy預測器下游準確性的影響。通過在幾個領域的廣泛評估,我們發現LLMs可以在有意義的程度上形式化約束,生成的公式通常與人類專家提供的公式相似。此外,當生成的公式在語法上有效時,它們可以導致高質量的下游預測。代碼和基準可在 https://unitn-sml.github.io/auto-nesy-bench/ 獲得。
Decision Titan: Test-Time Training for Long-Term Memory in Offline Reinforcement Learning
2610.01513v1 by Jude Waide, Robert Lieck
Long-term dependencies remain a major challenge for sequential decision-making in the field of AI: RNNs suffer from vanishing gradients and the limited expressivity of vector-based hidden states, whilst Transformer-based models are limited by the quadratic scaling of attention. Recent work has proposed tackling this problem with the Test-Time Training (TTT) framework, which stores episodic memories in the parameters of a neural network through gradient descent at both train and test-time. This approach has seen success in the domain of Natural Language Processing, however, to the best of our knowledge it has not yet been applied to the domain of Reinforcement Learning (RL), nor has there been a study analysing how this memory practically functions. In this paper, we study the potential of the TTT framework for offline RL by augmenting a Decision Transformer with TTT layers, dubbed the Decision Titan. We analyse performance and properties of the model in the X-Maze environment, an extension of T-Maze designed to test sequential memory, and investigate how the memory mechanism learns by visualising gate values over time. Our key findings are that Decision Titan can learn long-term dependencies with ranges 20x longer than the context window, generalises to lengths 1.7x the training data, but crucially temporal generalisation depends on the time embeddings used, and the ability to learn long-term dependencies depends on how the relevant information is encoded.
摘要:長期依賴性仍然是人工智慧領域中序列決策的一個主要挑戰:RNN 遭受消失梯度和基於向量的隱藏狀態表達能力有限的問題,而基於 Transformer 的模型則受到注意力的二次擴展限制。最近的研究提出了使用測試時訓練(Test-Time Training, TTT)框架來解決這個問題,該框架通過在訓練和測試期間的梯度下降將情節記憶儲存在神經網絡的參數中。這種方法在自然語言處理領域取得了成功,然而,據我們所知,它尚未應用於強化學習(Reinforcement Learning, RL)領域,也沒有研究分析這種記憶的實際運作方式。在本文中,我們通過增強決策 Transformer,並加入 TTT 層,稱之為 Decision Titan,來研究 TTT 框架在離線 RL 中的潛力。我們在 X-Maze 環境中分析模型的性能和特性,這是一個設計用來測試序列記憶的 T-Maze 擴展,並通過可視化門值隨時間的變化來調查記憶機制的學習方式。我們的主要發現是 Decision Titan 能夠學習長期依賴性,其範圍比上下文窗口長 20 倍,對訓練數據的長度進行 1.7 倍的泛化,但關鍵是時間泛化依賴於使用的時間嵌入,而學習長期依賴性的能力則取決於相關信息的編碼方式。
OverAct: Measuring and Mitigating Proactive Over-Authorization in LLM Tool-Calling Agents
2610.01508v1 by Taolin Zhang, Jiuheng Wan, Hanyu Wang, Tingyuan Hu, Chengyu Wang
LLM agents with tool-calling capabilities can access external services and private user data, but they may retrieve more information than a user's request explicitly requires. We study this behavior in structured tool-calling agents and term it proactive over-authorization. This setting differs from filesystem-level coding agents because the main risk is unnecessary access to private data. We introduce OverAct, a controlled benchmark spanning eight privacy-sensitive domains with deterministic, judge-free scoring, together with an interpretive decision-theoretic framework that yields three testable predictions. Across seven models from four families, all models significantly exceed authorized scope. Request specificity is the strongest predictor of severity, over-authorization grows sublinearly with tool-pool size, and decoding temperature has little effect. These patterns are consistent with a cost-asymmetry account, suggesting that over-authorization arises more from structural decision tendencies than from decoding randomness. We also propose SelfAudit, a zero-shot inference-time method that generates request-grounded justifications and filters unjustified calls before execution. Ablation shows that explicit filtering is the main driver of scope reduction. SelfAudit reduces privacy-oriented excess by 43% without oracle knowledge.
摘要:LLM 代理具有工具調用能力,可以訪問外部服務和私人用戶數據,但它們可能檢索比用戶請求明確要求的更多信息。我們研究這種行為在結構化工具調用代理中,並將其稱為主動過度授權。這種設置不同於文件系統級編碼代理,因為主要風險是對私人數據的不必要訪問。我們引入了 OverAct,一個涵蓋八個隱私敏感領域的受控基準,具有確定性、無評判的評分,並結合了一個解釋性決策理論框架,產生三個可測試的預測。在來自四個家族的七個模型中,所有模型的表現均顯著超出授權範圍。請求的具體性是嚴重程度的最強預測因子,過度授權隨著工具池大小的增長而次線性增長,而解碼溫度的影響很小。這些模式與成本不對稱的解釋一致,表明過度授權更多地源於結構性決策傾向,而非解碼隨機性。我們還提出了 SelfAudit,一種零樣本推斷時的方法,生成基於請求的理由並在執行前過濾不合理的調用。消融實驗顯示,明確過濾是範圍減少的主要驅動因素。SelfAudit 在沒有神諭知識的情況下將隱私導向的過剩減少了 43%。
A Multi-Agent LLM Framework for Personalized Health Checkup Interpretation and Guidance
2610.01451v1 by HyungJun Kim, Taehan Lee, Soojin Cheon
Personalized interpretation of health checkup results requires reasoning across longitudinal records, medical knowledge, lifestyle guidance, and healthcare navigation. We present a multi-agent large language model (LLM) system that identifies multiple intents, maps each to a task-specific agent, executes them in parallel, and synthesizes their outputs. We compared answers generated in Single Agent and Multi Agent settings on 120 Korean compound queries combining two to four requirements, using synthetic health checkup records. The Multi Agent improved the weighted LLM-judge score from 1.695 to 1.797 (p = 0.027), and three additional LLM judges showed consistent improvements ($Δ$ = +0.111 to +0.186, all p < 0.05). The gains came from usefulness, consistency, and the handling of every requirement in compound queries, whereas numerical accuracy and grounding improved significantly under only one of the four judges and medical safety did not differ, and critical failures occurred at similar rates (Single Agent 15.0% vs. Multi Agent 13.3%). Two human evaluators preferred Multi Agent in 66.7% and 68.3% of pairwise comparisons. Multi Agent execution increased latency and cost by 1.31$\times$ and 2.02$\times$, respectively. In exploratory subgroup analyses, the improvement was concentrated in queries involving personal-record lookup.
摘要:個性化的健康檢查結果解釋需要跨越長期記錄、醫學知識、生活方式指導和醫療導航的推理。
我們提出了一個多代理大型語言模型(LLM)系統,該系統識別多個意圖,將每個意圖映射到特定任務的代理,並平行執行它們,最後綜合其輸出。
我們比較了在單代理和多代理設置下,使用合成健康檢查記錄對120個韓國複合查詢生成的答案,這些查詢結合了兩到四個需求。
多代理將加權LLM評審分數從1.695提高到1.797(p = 0.027),另外三位LLM評審顯示出一致的改善($Δ$ = +0.111到+0.186,所有p < 0.05)。
這些增益來自於有用性、一致性以及對複合查詢中每個需求的處理,而數值準確性和基礎資料僅在四位評審中的一位顯著改善,醫療安全則沒有差異,且重大失誤的發生率相似(單代理15.0%對多代理13.3%)。
兩位人類評估者在66.7%和68.3%的成對比較中偏好多代理。
多代理執行使延遲和成本分別增加了1.31$\times$和2.02$\times$。
在探索性子群分析中,改善集中在涉及個人記錄查詢的問題上。
LLM-Assisted Discovery of Typed Semantic Links for Ontology Network Construction
2610.01393v1 by Nouha Hayouni, Sheeba Samuel, Alsayed Algergawy
Constructing typed, justified semantic links between ontologies is essential for enabling interoperability across heterogeneous and interdisciplinary knowledge domains. However, manually curating such links is difficult to scale. To address this challenge, we propose an end-to-end framework for ontology network construction that automates the discovery and generation of both intra-domain and inter-domain relationships. Our approach combines domain-adapted DistilBERT embeddings for dense contextual representation, clustering-based pre-filtering to reduce the candidate search space, and GPT-4o-driven relationship generation via iterative prompt engineering to produce semantically rich, interpretable links. Applied to ReproduceMeON - a network of 33 ontologies spanning machine learning, microscopy, computational science, and experimental workflow - the pipeline reduces approximately 800k raw concept pairs to 95k high-quality candidates. Human expert validation of 429 generated relationships by two independent annotators yields an overall precision of 80.19% (91.49% on high-certainty annotations) and an F1 of 0.890, with substantial inter-annotator agreement. Comparative experiments against five similarity-based baselines, including Sentence-BERT, show a substantial performance gap (best baseline F1 = 0.581), while an ablation study demonstrates that similarity-based methods alone fail to discriminate valid from invalid relationships (AUC approx 0.5) on the filtered candidate set. These findings highlight the necessity of LLM-based reasoning over concept roles and domain semantics for accurate relationship construction.
摘要:建構有類型、對齊的語義連結在本體之間對於實現異質和跨學科知識領域的互操作性至關重要。
然而,手動策劃這些連結難以擴展。
為了解決這一挑戰,我們提出了一個端到端的本體網絡建構框架,該框架自動發現和生成內域和跨域關係。
我們的方法結合了針對特定領域調整的DistilBERT嵌入以獲得密集的上下文表示、基於聚類的預過濾以減少候選搜索空間,以及通過迭代提示工程驅動的GPT-4o關係生成,以產生語義豐富、可解釋的連結。
應用於ReproduceMeON——一個涵蓋機器學習、顯微鏡學、計算科學和實驗工作流程的33個本體的網絡——該流程將大約80萬個原始概念對減少到9.5萬個高質量候選。
由兩位獨立註釋者對429個生成關係進行的人類專家驗證產生了整體精確度80.19%(高確定性註釋為91.49%)和F1值0.890,並且註釋者之間的協議顯著。
與五個基於相似性的基準進行的比較實驗,包括Sentence-BERT,顯示出顯著的性能差距(最佳基準F1 = 0.581),而消融研究表明,僅依賴相似性的方法無法區分有效和無效的關係(AUC約0.5)在過濾的候選集上。
這些發現突顯了基於LLM的推理在概念角色和領域語義上的必要性,以實現準確的關係建構。
Generation Provenance Before Behavior Attribution: Auditing Synthetic Speech Research Objects
2610.01378v1 by Sidi Chang, Peiying Zhu
Attributing model behavior to synthetic training data requires knowing what produced each training item before estimating what that item caused. A waveform-label pair does not preserve this knowledge. We propose a generation-provenance substrate in which a synthetic research object binds source specification, generated content, waveform, target, fact requirements, quality signals, review lineage, and immutable manifest identity. Producer and selection mechanism determine evidentiary meaning; storage location and variable name do not. We audit this substrate in a private Japanese care-handoff pipeline. A 113-asset review population contains 1.552 hours of synthetic speech across six scenario families; all items have linked audio, transcripts, candidate notes, and fact checklists, but human evidence is selective and source-specific. Two faithful-only manifests are scenario-seed-disjoint and immutably versioned, while exact upstream attribution remains blocked by floating generator aliases, missing per-clip TTS and code stamps, and an unversioned checking prompt. We argue that generation provenance is necessary but not sufficient for behavior attribution: it defines the candidate causal graph and audit units, whereas contributive attribution still requires frozen training runs and intervention or influence evidence. The paper contributes a compact provenance contract, an audit protocol, and a bounded case study for synthetic-data attribution; controlled research access may be offered, but we do not claim causal training-data attribution, clinical validity, or unrestricted public release.
摘要:將模型行為歸因於合成訓練數據需要了解每個訓練項目是如何產生的,然後才能估計該項目造成了什麼。波形-標籤對並不保留這種知識。我們提出了一種生成來源基底,其中合成研究對象綁定了源規範、生成內容、波形、目標、事實要求、質量信號、審查血統和不可變的清單身份。生產者和選擇機制決定了證據意義;存儲位置和變量名稱則不然。我們在一個私有的日本護理交接管道中審計這一基底。一個包含113個資產的審查群體包含了六個情境系列中的1.552小時合成語音;所有項目都鏈接了音頻、轉錄、候選筆記和事實檢查清單,但人類證據是選擇性的且特定於來源。兩個僅限忠實的清單在情境種子上是不相交且不可變版本的,而精確的上游歸因仍然受到浮動生成器別名、缺失的每段TTS和代碼印記以及未版本化的檢查提示的阻礙。我們主張生成來源對於行為歸因是必要但不充分的:它定義了候選因果圖和審計單位,而貢獻性歸因仍然需要凍結的訓練運行和干預或影響證據。本文貢獻了一個簡潔的來源合約、一個審計協議和一個有界的合成數據歸因案例研究;可能會提供受控的研究訪問,但我們不聲稱因果訓練數據歸因、臨床有效性或不受限制的公開發布。
ARCCS: An Automated Regulatory Compliance Checking System
2610.01345v1 by Giorgos Filandrianos, José Menezes, Chrysoula Zerva, Alessandro Gianola
Regulatory compliance checking - deciding whether a target document satisfies the obligations of a regulation - requires interpreting dense legal text, identifying which provisions apply, and grounding each decision in explicit evidence. We present ARCCS, an end-to-end, automated, agentic, and regulation-agnostic Legal NLP system for compliance checking. ARCCS decomposes raw regulatory text into atomic, traceable requirements and evaluates a target document against them using retrieved evidence, confidence scores, and human-interpretable justifications. This design decouples compliance assessment from any fixed regulatory template or predefined rule set, enabling the pipeline to operate over regulations of varying size and structure. We evaluate ARCCS in two complementary settings. First, in a GDPR policy-document evaluation, LLM-based judges find its decisions and justifications legally and evidentially consistent in up to 96.67% of the assessed cases. Second, on an EU public-procurement benchmark comprising more than 1,200 individual rule checks, the system attains 98.8% accuracy in violation detection. ARCCS is, to our knowledge, the first fully open-source system for end-to-end regulatory compliance checking and auditable report generation.
摘要:監管合規檢查 - 決定目標文件是否滿足法規的義務 - 需要解釋密集的法律文本,識別適用的條款,並將每個決策基於明確的證據。我們提出了 ARCCS,一個端到端、自動化、主動且與法規無關的法律自然語言處理系統,用於合規檢查。ARCCS 將原始法規文本分解為原子、可追溯的要求,並使用檢索的證據、信心分數和人類可解釋的理由來評估目標文件。這一設計將合規評估與任何固定的法規模板或預定的規則集解耦,使得該流程能夠在不同大小和結構的法規上運行。我們在兩個互補的環境中評估 ARCCS。首先,在 GDPR 政策文件評估中,基於 LLM 的評審在多達 96.67% 的評估案例中發現其決策和理由在法律和證據上是一致的。其次,在一個包含超過 1,200 個個別規則檢查的歐盟公共採購基準上,該系統在違規檢測中的準確率達到 98.8%。據我們所知,ARCCS 是第一個完全開源的端到端監管合規檢查和可審計報告生成系統。
An ontology for cross-sectoral crisis management: core and public health modules
2610.01326v1 by Aldo Gangemi, Rita T. Sousa, Luigi Asprino, Giorgia Lodi, Andrea G. Nuzzolese, Valentina Presutti, Johannes Gysen, Diana F. Sousa, Luigi Spagnolo
This paper presents the European Crisis Management Ontology (ECMO), a modular OWL-based ontology intended as a cross-sectoral reference for disaster risk reduction and response. ECMO is designed to be organised as a network of ontological modules. Among the modules, ECMO-CORE captures fundamental crisis management concepts such as hazard, event, exposure, impact, and response measure and uses ontology design patterns and the OWL2 punning technique to resolve ambiguities between hazard types and event manifestations. In addition, domain-specific modules are defined as in the case of the public health module aligned with SNOMED CT and ICD-11. To demonstrate the resource's utility, we used ECMO to represent the data of the Epidemic Intelligence from Open Sources system of the Joint Research Centre to generate an end-to-end pipeline that populates an ECMO-compliant knowledge graph from unstructured epidemiological news. Initial results demonstrate that ECMO provides the formal guardrails necessary for consistent and unified knowledge representation and integration. The ontology is publicly available at https://doi.org/10.5281/zenodo.20070268 and is released under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.
摘要:這篇論文介紹了歐洲危機管理本體(ECMO),這是一個基於OWL的模組化本體,旨在作為災害風險減少和應對的跨領域參考。ECMO的設計是作為本體模組的網絡組織。 在這些模組中,ECMO-CORE捕捉了基本的危機管理概念,如危險、事件、暴露、影響和應對措施,並使用本體設計模式和OWL2的雙義技術來解決危險類型和事件表現之間的歧義。此外,還定義了特定領域的模組,例如與SNOMED CT和ICD-11對齊的公共衛生模組。為了展示該資源的實用性,我們使用ECMO來表示聯合研究中心的開放來源流行病情報系統的數據,以生成一個從非結構化流行病學新聞填充ECMO合規知識圖譜的端到端管道。初步結果顯示,ECMO提供了必要的正式框架,以實現一致和統一的知識表示和整合。該本體可在https://doi.org/10.5281/zenodo.20070268上公開獲得,並根據創用CC 4.0國際版(CC BY 4.0)授權發布。
Dependency-Aware Reward Shaping for Agentic Reinforcement Learning
2610.01207v1 by Ziyi Chen, Yan Zhang, Jianhui Wei, Daoan Zhang, Zuozhu Liu
When training large language models with reinforcement learning, terminal rewards provide little guidance about which steps matter. Common methods for assigning step credit overlook that work built on uncorrected mistakes is wasted while independent work remains valid. With only a final success/failure reward, every step in a failed episode has zero total future reward, even when it made progress. We propose Dependency-Aware Reward Shaping (DARS), which represents task progress as predicates linked by prerequisite relations and assigns step-level credit over the dependency graph. An annotator marks which predicates each step verifies, invalidates, or repairs. Verified predicates are discounted according to graph distance from the nearest broken prerequisite, while independent predicates are unaffected. Repairs update these weights based on any errors that remain; invalidated predicates need re-verification to regain credit. A fixed potential converts these annotations into signed per-step rewards. A common reward and annotation interface allows DARS to integrate with a range of reasoning and agentic training methods, such as GiGPO and ARPO/AEPO, without changing their rollout strategies or optimizers. Across five task families and models from 1.5B to 8B, DARS improves success by up to 10 points over GiGPO trained with the same budget and harness (ALFWorld), raises the WebShop task score and Search-R1 QA accuracy, complements AEPO's entropy-based training on AIME24/25 with a Python interpreter, and exceeds OmniOPD in controlled tool-free reasoning comparisons at 1.7B and 4B. Ablations show that step-level credit, dependency attenuation, and graph topology each contribute. On ALFWorld, a distilled 8B annotator matches the API annotator, enabling DARS to run efficiently without a frontier judge. Code is available at https://github.com/JianhuiWei7/DARS.
摘要:在使用強化學習訓練大型語言模型時,最終獎勵對於哪些步驟重要的指導作用有限。常見的步驟信用分配方法忽視了基於未修正錯誤的工作是浪費的,而獨立工作仍然有效。僅有的最終成功/失敗獎勵使得在失敗的情況下,每一步的總未來獎勵為零,即使它有所進展。我們提出了依賴感知獎勵塑造(DARS),它將任務進展表示為由前置關係連結的謂詞,並在依賴圖上分配步驟級別的信用。一名標註者標記每一步驗證、無效或修復了哪些謂詞。經過驗證的謂詞根據與最近的破損前置條件的圖距離進行折扣,而獨立謂詞則不受影響。修復根據仍然存在的任何錯誤更新這些權重;無效的謂詞需要重新驗證以恢復信用。一個固定的潛力將這些標註轉換為簽名的每步獎勵。共同的獎勵和標註介面使DARS能夠與一系列推理和代理訓練方法(如GiGPO和ARPO/AEPO)集成,而無需改變它們的展開策略或優化器。在五個任務系列和從1.5B到8B的模型中,DARS在與相同預算和設備(ALFWorld)訓練的GiGPO相比,成功率提高了最多10個點,提升了WebShop任務得分和Search-R1 QA準確性,並補充了AEPO在AIME24/25上基於熵的訓練,使用Python解釋器,並在1.7B和4B的受控無工具推理比較中超過了OmniOPD。消融實驗顯示,步驟級信用、依賴衰減和圖拓撲各自都有貢獻。在ALFWorld上,一個精煉的8B標註者與API標註者相匹配,使DARS能夠高效運行而無需前沿評判者。代碼可在 https://github.com/JianhuiWei7/DARS 獲得。
Federated Agent Optimization
2610.01195v1 by Qiang Yang, Zhiqiang Kou, Xueyi Zhang, Dong-Dong Wu, Hanlin Gu, Jing Guo, Yang Liu, Di Jiang, Qian Xu
Large language model (LLM) agents increasingly operate in private environments and accumulate valuable experience from task execution, tool use, feedback, and local knowledge. Yet such experience is distributed across organizations and cannot be directly shared because of privacy and proprietary constraints. Conventional federated learning is insufficient for this setting, as agent capabilities extend beyond model parameters to memory, tools, rewards, skills, and structured knowledge. In this paper, we formulate \textbf{Federated Agent Optimization (FAO)}, which studies how distributed agents can collaboratively improve through controlled information exchange while keeping raw data, complete trajectories, and private knowledge local. We define FAO as a multi-objective problem balancing agent utility, privacy leakage, and communication cost, and organize its optimization space across policy, memory, tool use, reward, and structured knowledge and skills. We further characterize how private experience can be abstracted, protected, aggregated, and adapted into transferable capabilities, providing a unified view of how agents can benefit from one another without direct experience sharing. Finally, we identify the key challenges of FAO and outline several promising directions for future research toward trustworthy federated agent systems.
摘要:大型語言模型 (LLM) 代理人越來越多地在私密環境中運作,並從任務執行、工具使用、反饋和本地知識中積累寶貴的經驗。
然而,這些經驗分散在各個組織中,由於隱私和專有限制,無法直接共享。
傳統的聯邦學習在這種情況下是不夠的,因為代理人的能力超出了模型參數,還包括記憶、工具、獎勵、技能和結構化知識。
在本文中,我們提出了\textbf{聯邦代理優化 (FAO)},研究分散的代理人如何通過受控的信息交換協作改進,同時保持原始數據、完整的軌跡和私有知識的本地性。
我們將FAO定義為一個多目標問題,平衡代理效用、隱私洩漏和通信成本,並在政策、記憶、工具使用、獎勵以及結構化知識和技能之間組織其優化空間。
我們進一步描述了如何將私有經驗抽象化、保護、聚合和適應為可轉移的能力,提供了一個統一的視角,說明代理人如何在不直接共享經驗的情況下相互受益。
最後,我們確定了FAO的主要挑戰,並概述了幾個有前景的未來研究方向,以促進可信的聯邦代理系統。
Temporally-Resolved Token Attribution Reveals the Generation Dynamics of Diffusion Language Models
2610.01177v1 by Darpan Aswal, Céline Hudelot
This work presents Diffusion Layer Integrated Gradients (DLIG), a token attribution method for diffusion language models (DLMs) that extends Integrated Gradients (IG~\cite{sundararajan2017axiomatic}) to arbitrary layers and denoising steps. DLIG attributes a DLM's progressive commitment to a self-generated or fixed completion for an input prompt. We establish direct correspondences between DLIG and the IG axioms of completeness, implementation invariance, linearity, and symmetry preservation. As a lightweight complement to interventional analysis, DLIG provides an inexpensive first check of mechanistic hypotheses across the denoising trajectory. We demonstrate this on word-sense disambiguation, multi-hop graph reasoning, and sentence infilling, revealing how DLMs draw on inputs across positions, layers, and denoising steps.
摘要:這項工作提出了擴展了整合梯度(Integrated Gradients, IG~\cite{sundararajan2017axiomatic})至任意層和去噪步驟的擴散層整合梯度(Diffusion Layer Integrated Gradients, DLIG),這是一種用於擴散語言模型(Diffusion Language Models, DLMs)的標記歸因方法。
DLIG 將 DLM 對於自生成或固定完成的輸入提示的逐步承諾進行歸因。
我們建立了 DLIG 與 IG 完整性、實現不變性、線性和對稱性保持的公理之間的直接對應關係。
作為對介入分析的輕量補充,DLIG 提供了一種便宜的初步檢查,用於在去噪過程中檢驗機理假設。
我們在詞義消歧、多跳圖推理和句子填充上展示了這一點,揭示了 DLM 如何在不同位置、層和去噪步驟中利用輸入。
OrbitTAMP: Grounding Language Models for Task and Motion Planning in Spacecraft Rendezvous
2610.01093v1 by Yuji Takubo, Daniele Gammelli, Marco Pavone, Simone D'Amico
Spacecraft rendezvous and proximity operations (RPO) are currently planned through an expertise-intensive process in which engineers translate high-level operational intent into safe, dynamically feasible trajectories, creating a bottleneck to scalable operations. Large language model (LLM)-based agents could offer an intuitive interface for this process, although their outputs are not inherently grounded in orbital dynamics, operational constraints, or the structure of admissible spacecraft maneuvers. To exploit their semantic reasoning while ensuring the generated plan's physical validity, this paper presents a hierarchical framework for spacecraft task-and-motion planning (TAMP) that grounds LLM reasoning in a graph of reusable behaviors and domain-specific planning modules. Within this framework, a pretrained LLM maps a natural-language command to a partial mission specification. The associated planners then resolve unspecified decisions within the admissible operational space. Finally, trajectory optimization converts the completed mission specification into a dynamically feasible trajectory. Numerical experiments demonstrate that this architecture substantially improves intent recovery over direct LLM generation, achieving 98% exact recovery of partial mission specifications across all evaluated splits when backed by frontier LLMs. Additional test-time-compute experiments show that, for a compact 9B model, verifier-guided revision increases exact recovery from 75% to 88%, while broader behavior-plan search independently improves selection among admissible trajectory realizations. Overall, these results establish a scalable and auditable foundation for language-driven agentic planning of spacecraft RPO.
摘要:太空船會合與接近操作(RPO)目前是通過一個需要專業知識的過程進行規劃,在這個過程中,工程師將高層次的操作意圖轉化為安全且動態可行的軌跡,這造成了可擴展操作的瓶頸。基於大型語言模型(LLM)的代理可以為這個過程提供直觀的界面,儘管它們的輸出並不固有地基於軌道動力學、操作限制或可接受的太空船機動結構。為了利用它們的語義推理,同時確保生成計劃的物理有效性,本文提出了一個太空船任務與運動規劃(TAMP)的分層框架,該框架將LLM推理基於可重用行為和特定領域規劃模塊的圖進行基礎化。在這個框架內,預訓練的LLM將自然語言命令映射到部分任務規範。相關的規劃者然後在可接受的操作空間內解決未指定的決策。最後,軌跡優化將完成的任務規範轉換為動態可行的軌跡。數值實驗表明,這個架構顯著改善了意圖恢復,相較於直接的LLM生成,當得到前沿LLM的支持時,在所有評估的分割中實現了98%的部分任務規範的精確恢復。額外的測試時間計算實驗顯示,對於一個緊湊的9B模型,驗證者引導的修訂將精確恢復從75%提高到88%,而更廣泛的行為計劃搜索則獨立改善了在可接受的軌跡實現中的選擇。總體而言,這些結果為基於語言驅動的太空船RPO代理規劃建立了一個可擴展和可審計的基礎。
MOMAT: Mixture of Multiple Atlases for Low-Power Jailbreak Defense of Quantized LLMs
2610.01058v1 by Boyang Li, Bingyu Shen, Weihao Hong, Zhiyuan Jiang, Xinlei Guan, Yan Ma, Miles Q. Li, Yi Sheng, Ruiyang Qin
Quantized large language models are increasingly deployed on edge devices for their low latency and energy efficiency. However, model quantization weakens alignment safeguards, leaving qLLMs (quantized large language models) highly vulnerable to jailbreak attacks. To address this challenge, we present MOMAT (Mixture of Multiple Atlases), a hardware-enhanced safety framework that combines structured knowledge retrieval with low-power defense acceleration. Each atlas represents a semantic cluster of harmful or benign sample sets and policy templates, enabling domain-localized Retrieval-Augmented Generation guarding that mitigates the curse of dimensionality and the resulting semantic sparsity problem in large, heterogeneous safety databases. MOMAT retrieves top-$k$ similarity features from all atlases for each prompt and evaluates them using a lightweight MoE (Mixture of Experts) detector, while a CiM (Compute-in-Memory)-accelerated similarity engine performs fast, low-power atlas-local retrieval. MOMAT's CiM-based retrieval accelerates a 100-query batch from 15,052.44 ms to 3,207.21 ns (a $4.69 \times 10^6\times$ speedup) and reduces energy from $8.1 \times 10^7$ $μ$J to 3.32 $μ$J, yielding an approximately $2.5 \times 10^5\times$ energy reduction over DRAM-based (Raspberry Pi) baselines. Red-team evaluations across standard benchmarks show that MOMAT matches the defense performance of state-of-the-art methods while avoiding benign overkill and providing substantial efficiency gains, demonstrating that CiM-based modular defenses can make edge-deployed qLLMs both safer and more energy-efficient. We will release the full 223.2k-sample dataset to foster future research.
摘要:量化的大型語言模型越來越多地部署在邊緣設備上,因為它們具有低延遲和能量效率。然而,模型量化削弱了對齊保護,讓qLLMs(量化大型語言模型)對越獄攻擊高度脆弱。為了解決這一挑戰,我們提出了MOMAT(多重圖譜混合),這是一個硬體增強的安全框架,結合了結構化知識檢索和低功耗防禦加速。每個圖譜代表一個有害或良性樣本集和政策模板的語義集群,實現了域本地化的檢索增強生成保護,減輕了維度詛咒和大型異構安全數據庫中產生的語義稀疏問題。MOMAT為每個提示從所有圖譜中檢索前$k$相似特徵,並使用輕量級的MoE(專家混合)檢測器對其進行評估,而CiM(內存計算)加速的相似性引擎則執行快速、低功耗的圖譜本地檢索。MOMAT的基於CiM的檢索將100查詢批次的時間從15,052.44毫秒加速到3,207.21納秒($4.69 \times 10^6\times$加速),並將能量從$8.1 \times 10^7$ $μ$J減少到3.32 $μ$J,實現了約$2.5 \times 10^5\times$的能量減少,相較於基於DRAM(樹莓派)的基準。針對標準基準的紅隊評估顯示,MOMAT的防禦性能與最先進的方法相當,同時避免了良性過度防護,並提供了顯著的效率增益,證明基於CiM的模組化防禦可以使邊緣部署的qLLMs更加安全和能量高效。我們將釋出完整的223.2k樣本數據集,以促進未來的研究。
Capturing In-Context Learning Dynamics with Task Operators
2610.01054v1 by Guangzhi Xiong, Zhenghao He, Bohan Liu, Sanchit Sinha, Wenqian Ye, Aidong Zhang
In-context learning (ICL) enables language models to perform new tasks from demonstrations without weight updates. However, every ICL inference requires processing the full set of examples, resulting in inefficient deployments, and how ICL works mechanistically is not fully understood. Prior work compresses ICL into fixed activation vectors extracted from specific layers or positions, but these input-independent interventions fail on complex tasks where the output depends on fine-grained interactions with the input. By analyzing the ICL forward pass, we show that each attention head's output is an affine transformation of its context-masked counterpart, and that the parameters of this transformation are empirically stable across samples for a given task. Building on this, we introduce Task Operator (TO), which replays this transformation as an analytically derived update to the attention output projection. Across lexical, algorithmic, and reasoning tasks, TO achieves the best overall performance among prior methods and substantially narrows the gap between zero-shot inference and ICL. We further show that the extracted knowledge concentrates in a task-specific sparse circuit across layers and positions, and that averaging operators from disjoint demonstration batches enables effective many-shot scaling without expanding the context window. Our code is available at https://github.com/gzxiong/task_operator.
摘要:在上下文學習(ICL)中,語言模型能夠從示範中執行新任務,而無需更新權重。
然而,每次 ICL 推理都需要處理完整的示例集,導致部署效率低下,並且 ICL 的機制運作方式尚未完全理解。
先前的研究將 ICL 壓縮為從特定層或位置提取的固定激活向量,但這些與輸入無關的干預在輸出依賴於與輸入的細緻互動的複雜任務上失敗。
通過分析 ICL 的前向傳遞,我們顯示每個注意力頭的輸出是其上下文遮罩對應物的仿射變換,並且這種變換的參數在給定任務的樣本中經驗上是穩定的。
基於此,我們引入了任務運算符(TO),它將這種變換重播為對注意力輸出投影的解析衍生更新。
在詞彙、算法和推理任務中,TO 在先前方法中實現了最佳的整體性能,並大幅縮小了零-shot 推理和 ICL 之間的差距。
我們進一步顯示,提取的知識集中在跨層和位置的任務特定稀疏電路中,並且從不相交的示範批次中平均運算符能夠有效地進行多次擴展,而無需擴大上下文窗口。
我們的代碼可在 https://github.com/gzxiong/task_operator 獲得。
Beyond Answer Confidence: A Controlled Audit of Self-Knowledge in a Black-Box Decision Model
2610.01006v1 by Sharath M Shankaranarayana, Davor Runje, Jan Jannink
Decision models return probabilities intended for routing, abstention and automated action. Calibration makes those probabilities useful on average, but does not establish whether low confidence reflects chance or missing knowledge, nor whether confidence falls when a model moves beyond what it knows. We audit this distinction in Jev, a decision model, with over 15 public datasets and 6 generated task families, with paired interventions that vary the information supplied for a fixed item. Jev's confidence is calibrated on familiar closed-choice tasks but fails as an indicator of missing knowledge: with no answer-relevant information it assigns up to 0.80 to a salient option, and on news beyond an observed knowledge boundary it exceeds accuracy by 0.21--0.33, a gap that recalibration on earlier months does not close. Targeted yes/no questions give sharper readouts of the case: whether an outcome is settled (AUROC 1.00) and whether the evidence suffices (0.95, against 0.85 for confidence on the same items). Asking whether Jev knows the answer appears to flag fabricated entities and post-boundary news (0.91), but with realistic names or with dates removed it shows no advantage over answer uncertainty. Black-box knowledge audits therefore need explicit controls for surface cues. Code: https://github.com/Syntheme/beyond-answer-confidence.
摘要:決策模型返回旨在路由、放棄和自動行動的概率。校準使這些概率在平均情況下變得有用,但並未確定低信心是否反映了隨機性或缺失的知識,也未確定當模型超出其已知範疇時信心是否會下降。我們在Jev這個決策模型中審核這一區別,使用超過15個公共數據集和6個生成的任務系列,並配對干預,變化固定項目的信息供應。Jev的信心在熟悉的封閉選擇任務上經過校準,但作為缺失知識的指標卻失效:在沒有與答案相關的信息時,它對一個顯著選項賦予高達0.80的概率,並且在超過觀察知識邊界的新聞中,其準確性超過0.21--0.33,這一差距在對早期月份的重新校準中並未縮小。針對性的是/否問題提供了更清晰的案例讀數:結果是否已確定(AUROC 1.00)以及證據是否足夠(0.95,相較於對同一項目的信心為0.85)。詢問Jev是否知道答案似乎標記了虛構實體和邊界後的新聞(0.91),但在使用現實名稱或去除日期的情況下,並未顯示出相對於答案不確定性的優勢。因此,黑箱知識審核需要對表面線索進行明確控制。代碼:https://github.com/Syntheme/beyond-answer-confidence。
Distilling Directional Verification
2610.00997v1 by Jungseob Lee, Sugyeong Eo, Seongtae Hong, Seungyoon Lee, Chanjun Park, Jaehyung Seo, Heuiseok Lim
Knowledge distillation aims to transfer the factual knowledge of large language models to smaller models for efficient deployment. Yet a teacher may recall a relation in one direction while failing to generate the answer in the reverse direction. Distillation from its generated answers can therefore propagate this directional limitation to the student. The same teacher can nevertheless recognize such an answer by scoring the relation in the direction it knows. We introduce directional label distillation, in which frozen teachers score candidate answers in that known direction and the best-scoring candidate becomes the student's training target. On facts about parents and their children, known-direction scoring yields more accurate labels than scoring the requested direction, even after tuned corrections for name priors. With prior-corrected scores, the better direction depends on the facts rather than the template, and reverses on mined facts whose notable entity is the parent rather than the child. With the evaluated children's forward facts withheld, students trained on known-direction labels improve open-ended accuracy on their trained queries by 13 to 15 points over students trained on prior-corrected reverse labels. After generated answers are matched to a fixed name list by lexical similarity, students reproduce nearly all selected labels. Their accuracy largely follows label quality. The label advantage holds on unscreened queries and when candidates are retrieved without inserting correct answers. Our findings show that directional verification mitigates the transfer of errors from teacher-generated answers to students by providing more accurate training targets. Code is available at https://github.com/js-lee-AI/directional-verification.
摘要:知識蒸餾旨在將大型語言模型的事實知識轉移到較小的模型,以便高效部署。然後,教師可能在一個方向上回憶起一個關係,但未能在相反方向生成答案。因此,從其生成的答案進行蒸餾可能會將這一方向性限制傳播給學生。然而,同一位教師仍然可以通過在其已知方向上對關係進行評分來識別這樣的答案。我們引入了方向性標籤蒸餾,其中凍結的教師在已知方向上對候選答案進行評分,得分最高的候選者成為學生的訓練目標。在有關父母及其子女的事實中,已知方向的評分比要求方向的評分產生更準確的標籤,即使在對名稱先驗進行調整後也是如此。使用經過先驗修正的分數,更好的方向取決於事實而不是模板,並在挖掘的事實中反轉,當其顯著實體是父母而不是子女時。在評估的子女的前向事實被保留的情況下,基於已知方向標籤訓練的學生在其訓練查詢上的開放式準確率提高了13到15個點,超過了基於先驗修正的反向標籤訓練的學生。在通過詞彙相似性將生成的答案與固定名稱列表匹配後,學生幾乎重現了所有選定的標籤。他們的準確性在很大程度上取決於標籤質量。標籤優勢在未篩選的查詢上以及在檢索候選者時未插入正確答案的情況下依然存在。我們的研究結果顯示,方向性驗證通過提供更準確的訓練目標來減輕教師生成的答案向學生轉移錯誤的影響。代碼可在 https://github.com/js-lee-AI/directional-verification 獲得。
Structure-agnostic Causal Representation Learning
2610.00968v1 by Arman Behnam, Binghui Wang
Causal representation learning aims to discover robust features by exploiting the causal structure underlying data generation. Existing methods require specifying the causal structure a priori, yet different structures demand fundamentally incompatible invariance constraints, and misspecification leads to representations that discard predictive information. We introduce SaCRL, a framework that jointly identifies the causal structure and learns the corresponding invariant representation without prior structural knowledge. Our approach formulates structure selection as a soft optimization over candidate invariances using HSIC-based violation metrics, with adaptive weights that automatically concentrate on the achievable structure. We provide theoretical guarantees for structure identification, including under random-feature approximation, invariance satisfaction, and out-of-distribution generalization. Empirically, SaCRL recovers the true structure on synthetic and semi-synthetic Bayesian-network benchmarks, outperforms fixed-invariance baselines on Colored MNIST, achieves state-of-the-art accuracy on three DomainBed benchmarks (PACS, VLCS, OfficeHome), and degrades gracefully under structural misspecification and limited environment diversity. Code is available at: https://github.com/ArmanBehnam/sacrl.
摘要:因果表示學習旨在通過利用數據生成背後的因果結構來發現穩健的特徵。現有的方法需要事先指定因果結構,但不同的結構要求根本上不相容的不變性約束,且錯誤指定會導致丟失預測信息的表示。我們引入了SaCRL,一個框架,能夠在沒有先驗結構知識的情況下共同識別因果結構並學習相應的不變表示。我們的方法將結構選擇表述為對候選不變性的軟優化,使用基於HSIC的違規度量,並具有自適應權重,自動集中於可實現的結構。我們提供了結構識別的理論保證,包括隨機特徵近似、不變性滿足和分佈外泛化。在實證上,SaCRL在合成和半合成的貝葉斯網絡基準上恢復了真實結構,在Colored MNIST上超越了固定不變性基準,在三個DomainBed基準(PACS、VLCS、OfficeHome)上達到了最先進的準確率,並在結構錯誤指定和環境多樣性有限的情況下優雅降級。代碼可在以下網址獲得:https://github.com/ArmanBehnam/sacrl。
ABDA-NL: A Natural-Language Scenario Explorer for Argument-Based Reasoning
2610.00947v1 by Shawn Bowers, Martin Caminada, Haoyang Liu, Bertram Ludäscher
ABDA-NL adds a natural-language interface to ABDA, a system for argument-based discussion using ASPIC- knowledge bases under grounded semantics. Users see which conclusions are accepted, rejected, or undecided, open an interactive rendering of the grounded discussion game to learn why, explore what-if alternatives by suspending assumptions and rules or changing preferences, ask questions that are answered from a scenario's reference documents, and author new facts, assumptions, and rules in plain English. A large language model provides the bridge between language and formalism: it answers questions from the documents and the current state of the scenario, and it translates plain-English edits into candidate formal statements. The deterministic ABDA engine remains the sole source of arguments, attacks, and acceptance labels, and every proposal of the model is validated and confirmed by the user before it takes effect.
摘要:ABDA-NL 為 ABDA 添加了一個自然語言介面,這是一個基於論點的討論系統,使用 ASPIC 知識庫並基於基礎語義。
用戶可以看到哪些結論被接受、拒絕或未決,打開一個互動式的基礎討論遊戲來了解原因,通過暫停假設和規則或改變偏好來探索假設的替代方案,提出問題,這些問題會從情境的參考文件中得到回答,並用簡單的英語創建新的事實、假設和規則。
一個大型語言模型提供了語言與形式之間的橋樑:它回答來自文件和當前情境狀態的問題,並將簡單英語的編輯翻譯成候選的正式陳述。
確定性的 ABDA 引擎仍然是論點、攻擊和接受標籤的唯一來源,模型的每一個提案在生效之前都需經用戶驗證和確認。
Screw Attention: Rigid-Body Algebra Inside a Transformer
2610.00904v1 by Aly Magassouba
Learned manipulation policies rediscover from data the spatial relations that rigid-body mechanics supplies in closed form. This costs data, and it leaves the policies fragile to geometric changes in the scene. We present Screw Attention, a transformer layer in which the relation between two bodies is a spatial transform rather than a graph edge. Every token is a body with a pose. Each pair of tokens carries the relative pose and, for robot joints, the joint screw. Messages are transported along this relation into the receiver's frame, while the attention scores see only frame-invariant quantities. By construction, the messages are equivariant to an independent change of frame at every token, and a single layer can express the velocity recursion of rigid-body mechanics. On simulated manipulation tasks, Screw Attention matches or outperforms controls of the same size, including graph, transformer and flat networks on LIBERO-Spatial. With 16,162 parameters it reaches 97.3% on LIBERO-Spatial from object poses (without images or language), above a flat network with 27x more parameters. Under a change of per-link frame convention its success is unchanged, while every other learned network falls below 3%. Placed on an analytic controller as a gated residual, it raises insertion success by 17.3 points. It is unaffected by pose noise up to 10,mm and by joint offsets within the factory calibration of a Franka arm. These results suggest a criterion: geometry is decisive when the task requires relations between frames that no other part of the system supplies. Code and trained policies will be released.
摘要:學習的操控策略從數據中重新發現剛體力學以封閉形式提供的空間關係。這需要數據,並且使得策略對場景中的幾何變化變得脆弱。我們提出了螺旋注意力(Screw Attention),這是一個Transformer層,其中兩個物體之間的關係是一種空間變換,而不是圖邊。每個標記都是一個具有姿態的物體。每對標記攜帶相對姿態,對於機器人關節,則是關節螺旋。消息沿著這種關係傳輸到接收者的框架中,而注意力分數僅查看框架不變的量。根據構造,這些消息對每個標記的獨立框架變化是等變的,且單層可以表達剛體力學的速度遞歸。在模擬操控任務中,螺旋注意力的表現與相同大小的控制器相匹配或超過,包括在LIBERO-Spatial上的圖形、Transformer和扁平網絡。擁有16,162個參數的它在LIBERO-Spatial上從物體姿態達到97.3%(不使用圖像或語言),超過了一個擁有27倍參數的扁平網絡。在每個連接的框架約定變化下,它的成功率保持不變,而其他學習的網絡則降至3%以下。作為一個門控殘差放置在分析控制器上,它將插入成功率提高了17.3個點。它對高達10mm的姿態噪聲和在Franka臂的工廠校準範圍內的關節偏移不受影響。這些結果暗示了一個標準:當任務需要框架之間的關係,而系統的其他部分無法提供時,幾何是決定性的。代碼和訓練好的策略將會發布。
Match the Distribution, Not the Compute: Post-Training Multi-Token Prediction Heads
2610.00888v1 by Prachi Badarayani, Aidan Jay, Chenghui Zhou, Dayquan Julienne, Yuan Gao, Tianwei Chen, George Zerveas, Ishmam Zabir, Xiren Zhou, Chris Quirk, Xia Song
Multi-token prediction (MTP) improves the throughput of autoregressive generation by enabling the language model to draft multiple next tokens per forward pass, while a verification step over draft tokens ensures that token distribution of the backbone is preserved. Every open MTP-family release (MiMo-7B, DeepSeek-V3, Qwen3) trains its heads jointly with the backbone over the full pretraining run of tens of trillions of tokens, thus setting the drafter quality at pretraining time. We ask whether a lightweight post-training pass on target-generated chain-of-thought is enough to reach the same expected throughput speedup on a frozen reasoning model, and study how a serving-time system built on such a checkpoint can be optimized. We present three findings. 1) On a frozen Qwen3-8B with $K{=}3$ chained MTP heads, we show that a post-training recipe with plain cross-entropy on $\approx!2.5$B tokens reaches or exceeds the expected speedup of jointly trained MiMo-7B on math, coding and knowledge benchmarks. Our post-training recipe utilizes $10^3$-$10^4\times$ less MTP-training tokens as compared with joint pre-training of MiMO-7B MTP baseline. 2) We propose a chain-aware relaxation of draft token verification rule that allows a bounded drift from backbone language model token distribution. We show that this relaxation lifts expected speedups by $+12$ to $+16\%$ per benchmark while preserving task accuracy. 3) We propose an adaptive controller that dynamically chooses the number of MTP heads to be engaged at inference time and demonstrate recovery of upto $11$--$14\%$ loss in speedup using fixed maximum MTP draft length.
摘要:多標記預測(MTP)透過使語言模型在每次前向傳遞中草擬多個下一個標記來提高自回歸生成的吞吐量,同時對草擬標記進行驗證步驟以確保骨幹的標記分佈得以保留。每個開放的MTP家族版本(MiMo-7B、DeepSeek-V3、Qwen3)在數十萬億標記的完整預訓練過程中,與骨幹共同訓練其頭部,從而在預訓練時設置草擬者的質量。我們詢問在目標生成的思維鏈上進行輕量級的後訓練過程是否足以在凍結的推理模型上達到相同的預期吞吐量加速,並研究基於此檢查點構建的服務時間系統如何進行優化。我們提出三個發現。1)在一個凍結的Qwen3-8B上,使用$K{=}3$鏈式MTP頭,我們顯示一個在$\approx!2.5$B標記上使用普通交叉熵的後訓練配方達到或超過了在數學、編碼和知識基準上共同訓練的MiMo-7B的預期加速。我們的後訓練配方使用的MTP訓練標記比MiMo-7B MTP基準的共同預訓練少了$10^3$-$10^4\times$。2)我們提出了一個鏈式感知的草擬標記驗證規則的放鬆,允許與骨幹語言模型標記分佈的有界漂移。我們顯示這種放鬆使每個基準的預期加速提高了$+12$到$+16\%$,同時保持任務準確性。3)我們提出了一個自適應控制器,動態選擇在推理時參與的MTP頭的數量,並展示使用固定的最大MTP草擬長度恢復高達$11$--$14\%$的加速損失。
Backdoor Purification for LoRA-Tuned LLMs via Null-Space Projection
2610.00685v1 by Jianwei Li, Jung-Eun Kim
With the rapid adoption of large language models (LLMs) and parameter-efficient fine-tuning (PEFT) methods, the risk of backdoor attacks has become more severe. Existing backdoor purification methods typically rely on at least one of the strong assumptions, such as prior knowledge of triggers, access to clean references, or aggressive retraining, and they often lack comprehensive evaluations. These constraints substantially limit their practical applicability. To overcome these challenges, our work proposes purifying LoRA-tuned LLMs without these assumptions and even without post-hoc retraining of the suspect parameters. Our objective is to significantly reduce the attack success rates (ASR) while preserving both (i) the base model's general capabilities and (ii) the new downstream skills learned through the adapter. Through a series of ablation studies, we progressively scale our approach from a single layer in a text classification setting to a full-parameter LLM in the generative task. Through careful data curation and feature approximation, we extract high-fidelity backdoor directions and, for each layer or head, construct orthogonal null spaces in both the input and output channels, onto which the LoRA updates are projected. Empirically, our null-space projection method reduces the ASR from nearly 100% to less than 10%, while preserving the base model's benign performance and the adapter's learned abilities during downstream task adaptation.
摘要:隨著大型語言模型(LLMs)和參數高效微調(PEFT)方法的快速採用,後門攻擊的風險變得更加嚴重。現有的後門淨化方法通常依賴於至少一個強假設,例如對觸發器的先驗知識、訪問乾淨參考資料或激進的再訓練,並且它們往往缺乏全面的評估。這些限制大大限制了它們的實際應用性。為了克服這些挑戰,我們的工作提出了在沒有這些假設的情況下淨化LoRA調整的LLMs,甚至不需要對可疑參數進行事後再訓練。我們的目標是顯著降低攻擊成功率(ASR),同時保留(i)基礎模型的一般能力和(ii)通過適配器學到的新下游技能。通過一系列的消融研究,我們逐步將我們的方法從文本分類設定中的單層擴展到生成任務中的全參數LLM。通過仔細的數據策劃和特徵近似,我們提取高保真度的後門方向,並為每一層或頭構建正交的零空間,這些零空間位於輸入和輸出通道上,LoRA更新將被投影到這些空間中。經驗上,我們的零空間投影方法將ASR從近乎100%降低到不到10%,同時在下游任務適應過程中保留了基礎模型的良性性能和適配器學到的能力。
Ontology-Grounded, Reasoner-Verified Benchmarks for Evaluating LLM Reasoning in Scientific AI
2610.00682v1 by Nishtha N. Vaidya, Stephan Grimm, Thomas Hubauer, Thomas A. Runkler
Large language models (LLMs) increasingly underpin scientific AI applications that reason over structured knowledge, from biomedical question answering to materials informatics. However, their logical reasoning often falls short, producing factual inaccuracies unacceptable in these settings. Reliable evaluation remains challenging: manual dataset construction scales poorly, and LLM-based generation risks embedding the very flaws it aims to measure. High-quality benchmarks must ground both correct and incorrect labelled examples in explicit background knowledge, formally verifiable by a standard reasoner. We propose a pipeline that automatically generates ontology-grounded multiple-choice question (MCQ) benchmarks from any sufficiently axiomatised OWL 2 ontology, with correct answers grounded in the ontology by design. Distractors are generated by perturbing the right-hand-side class expressions of class definition axioms, and their incorrectness is formally verified by an OWL reasoner via entailment checks. We evaluate the pipeline on three ontologies: Pizza (small, academic), PMDco (complex, materials science), and DOID (large, biomedical), generating 112, 2,491, and 15,216 MCQs respectively. Distractors span four semantic categories from class unsatisfiability to weakened subsumptions, enabling diagnostic evaluation of specific reasoning failures. Items meet natural language quality standards: mean LLM judge scores of 4.02, 4.36, and 3.36 out of 5 confirm fluency, and correct-answer-to-distractor similarity above 0.8 shows that wrong options cannot be dismissed on surface form alone. Six LLMs evaluated zero-shot achieve 41.1-76.8% accuracy, well above the 25% random-guessing baseline, indicating the benchmarks are challenging and discriminative. This work is a step towards more reliable benchmarks for assessing logical reasoning in scientific AI.
摘要:大型語言模型(LLMs)越來越多地支撐著科學人工智慧應用,這些應用在結構化知識上進行推理,從生物醫學問答到材料資訊學。然而,它們的邏輯推理經常不夠準確,在這些情境中產生的事實不準確是不可接受的。可靠的評估仍然具有挑戰性:手動數據集構建的擴展性差,而基於LLM的生成則有可能嵌入它所旨在測量的缺陷。高品質的基準必須將正確和不正確的標記示例基於明確的背景知識,並由標準推理器形式驗證。我們提出了一個管道,該管道自動從任何足夠公理化的OWL 2本體生成基於本體的多選題(MCQ)基準,正確答案在設計上基於本體。干擾項是通過擾動類定義公理的右側類表達式生成的,其不正確性通過OWL推理器通過推理檢查形式驗證。我們在三個本體上評估了該管道:Pizza(小型,學術)、PMDco(複雜,材料科學)和DOID(大型,生物醫學),分別生成112、2,491和15,216個MCQ。干擾項涵蓋了四個語義類別,從類不滿足到弱化的子類關係,使得特定推理失敗的診斷評估成為可能。項目符合自然語言質量標準:LLM評審的平均分數為4.02、4.36和3.36(滿分5分),確認了流暢性,且正確答案與干擾項的相似度超過0.8,顯示錯誤選項不能僅僅因表面形式而被忽視。六個評估的零樣本LLM達到41.1-76.8%的準確率,遠高於25%的隨機猜測基準,表明這些基準具有挑戰性和區分性。這項工作是朝著更可靠的基準邁出的一步,以評估科學人工智慧中的邏輯推理。
PhysicsMate: A Curriculum-Grounded Bengali Benchmark for Secondary Physics QA with Small-Model Adaptation
2610.00664v1 by Rashid Azraf Jahin, Saadman Sajid, Khan Raiyan Ibne Reza, Sumaiya Tabassum Nimi
Bengali secondary education lacks curriculum-grounded benchmarks for STEM question-solving, and general-purpose language models struggle with the precise terminology, unit conventions, and derivations that physics problems demand. We introduce PhysicsMate, a benchmark of 1834 question-answer pairs built from the National Curriculum and Textbook Board (NCTB) Grade 9-10 physics syllabus and grounded in a multi-relational knowledge graph of 1760 nodes and 2600 edges across ten ontological types. We Low-Rank Adapt at 0.6B, 1.7B, and 4B parameters, with a unified recipe and demonstrate a significant increase in closed-book accuracy in all scales (+5.5, +15.0, and +23.3 percentage points). A node-type analysis shows that the most benefited by adaptation is the structured curricular knowledge, which consists of physical quantities and named laws, while the least benefited is the loosely specified entity-level knowledge. The 4B model has been adapted and quantized to a small offline binary that can be used for local inference in resource constrained environments and offers a viable path to curriculum aligned physics support in environments with limited connectivity and hardware.
摘要:孟加拉的中學教育缺乏基於課程的STEM問題解決基準,而通用語言模型在物理問題所需的精確術語、單位慣例和推導方面表現不佳。我們介紹了PhysicsMate,這是一個由1834個問題-答案對組成的基準,基於國家課程和教科書委員會(NCTB)9-10年級物理課程大綱,並建立在一個包含1760個節點和2600條邊的多關係知識圖譜上,涵蓋十種本體類型。我們在0.6B、1.7B和4B參數下進行低秩適應,使用統一的配方,並在所有規模上顯示出閉卷準確率的顯著提高(+5.5、+15.0和+23.3個百分點)。節點類型分析顯示,適應中受益最多的是結構化課程知識,這包括物理量和命名法則,而受益最少的是鬆散指定的實體級知識。4B模型已被適應並量化為一個小型離線二進制文件,可用於資源受限環境中的本地推理,並為在連接性和硬體有限的環境中提供課程對齊的物理支持提供了一條可行的途徑。
Where's Waldo? Query-language Preference under Cross-lingual Knowledge Disparities
2610.00606v1 by Dayeon Ki, Ruochen Zhang, Silviu Cucerzan, Ryen W. White, Ning Gao
Large Language Models increasingly serve as interfaces for knowledge-intensive information seeking tasks across languages by synthesizing multilingual evidence. Prior work has shown that they often exhibit query-language preference -- the tendency to favor sources written in the language of the query -- but has largely examined this behavior in settings where equivalent knowledge is available across languages. However, this bias becomes consequential when sources in different languages provide incomplete or inconsistent accounts of the same fact, since the information users receive then depends on the sources a model selects to use. To characterize query-language preference under such cross-lingual knowledge disparities, we introduce Waldo, a multilingual Question-Answering (QA) benchmark constructed from Wikipedia. Waldo contains 12K QA pairs targeting knowledge gaps, where a fact is available in one language but absent in another, and knowledge conflicts, where language editions provide conflicting versions of the same fact. Evaluating eight models across five languages, we find that when one language edition merely lacks the relevant fact, models generally use evidence from the other language regardless of the query language. Under conflicting accounts, however, model responses strongly align with the document in the query language, causing semantically equivalent queries to elicit different accounts depending on the user's language. Finally, we explore two different approaches that could mitigate this preference under knowledge conflicts: a mechanistic intervention that ablates attention heads associated with query-language preference, and LoRA-based training, which reduces the preference gap by up to 61.5%.
摘要:大型語言模型越來越多地作為跨語言知識密集型信息搜尋任務的介面,通過綜合多語言證據來實現。先前的研究顯示,它們通常表現出查詢語言偏好——即偏向於使用以查詢語言撰寫的來源——但這種行為主要是在不同語言之間有等效知識的情況下進行的檢查。然而,當不同語言的來源提供相同事實的不完整或不一致的描述時,這種偏見變得至關重要,因為用戶所接收到的信息取決於模型選擇使用的來源。為了在這種跨語言知識差異下描述查詢語言偏好,我們引入了Waldo,一個基於維基百科構建的多語言問答(QA)基準。Waldo包含12K個針對知識空白的QA對,其中一種語言中有事實而另一種語言中缺失,以及知識衝突,其中語言版本提供相同事實的衝突版本。我們在五種語言中評估了八個模型,發現當一種語言版本僅缺少相關事實時,模型通常會使用來自另一種語言的證據,而不管查詢語言是什麼。然而,在衝突的描述下,模型的回應強烈對應於查詢語言中的文檔,導致語義上等價的查詢根據用戶的語言引出不同的描述。最後,我們探索了兩種不同的方法,可以減輕知識衝突下的這種偏好:一種機械干預,消除與查詢語言偏好相關的注意力頭,另一種基於LoRA的訓練,將偏好差距縮小至61.5%。
Emergent Unfaithfulness: How Alignment Training Causes Language Models to Silently Override Task Faithfulness
2610.00568v1 by Pardis Sadat Zahraei, Janvijay Singh, Gokhan Tur, Dilek Hakkani-Tur
Large language models are characterized by three key properties: capability, alignment, and faithfulness. Prior work studies the tradeoffs between capability and alignment, and between capability and faithfulness, but a third tension remains underexplored: the alignment-faithfulness conflict. We show that aligned models systematically deviate from their inputs on unsafe or sensitive content without disclosing the modification, a failure mode we call alignment-induced unfaithfulness (AIU). Unlike capability-driven unfaithfulness, which comes from errors in knowledge or reasoning, this is induced by post-training mechanisms that override adherence to the input. We introduce FaithConflict, a controlled dataset isolating both conflicts, and two complementary taxonomies: behavioral (B1-B8) and chain-of-thought reasoning (C0-C6). Across models, AIU increases with scale and more sharply than capability-driven unfaithfulness, a reverse scaling law; intermediate checkpoints show it is amplified during post-training, with DPO the stage at which the gap both grows most and becomes least visible. Prompting-based mitigation does not resolve it, revealing a capability-alignment-faithfulness trilemma in the design and evaluation of LLMs.
摘要:大型語言模型的特徵有三個關鍵屬性:能力、對齊和忠實性。
先前的研究探討了能力與對齊之間的權衡,以及能力與忠實性之間的權衡,但第三種緊張關係仍然未被充分探討:對齊-忠實性衝突。
我們展示了對齊模型在處理不安全或敏感內容時,系統性地偏離其輸入而不披露修改,這種失敗模式我們稱之為對齊引起的不忠實性(AIU)。
與能力驅動的不忠實性不同,後者源於知識或推理的錯誤,這種不忠實性是由後訓練機制引起的,這些機制覆蓋了對輸入的遵循。
我們引入了FaithConflict,一個控制數據集以隔離這兩種衝突,以及兩個互補的分類法:行為(B1-B8)和思維鏈推理(C0-C6)。
在各模型中,AIU隨著規模的增長而增加,並且增長的幅度比能力驅動的不忠實性更為明顯,這是一種反向縮放法則;中間檢查點顯示它在後訓練期間被放大,DPO是這一差距增長最多且變得最不明顯的階段。
基於提示的緩解措施無法解決此問題,揭示了大型語言模型設計和評估中的能力-對齊-忠實性三難問題。
Assessing the Impact of Language Disparity on Multilingual Linguistic Ability in Large Language Models
2610.00540v1 by Zhanyu Chen, Jaap Jumelet
Claims about the grammatical competence of multilingual language models vary sharply with how competence is measured, yet the interaction between evaluation paradigm, post-training, and language resource availability has not been systematically examined. We evaluate base and post-trained models from six families on MultiBLiMP, a syntactic minimal-pair benchmark covering 101 languages, using four evaluation methods. We report three principal findings. First, post-training degrades grammatical competence, but the magnitude of this effect is reduced unevenly by model scale, while low-resource languages bear the highest cost. Second, post-trained models retain grammatical knowledge they cannot articulate through explicit prompting, yet this is measurable only in high-resource languages, because near-chance baselines in low-resource settings leave little knowledge to hide. Third, native-language prompting recovers otherwise hidden competence on low-resource languages, demonstrating that only high-resource languages can be probed directly from unprompted probabilities. We conclude that multilingual grammatical evaluation must adopt language-informed, multi-paradigm protocols to avoid systematically underestimating low-resource abilities.
摘要:關於多語言模型的語法能力的主張,隨著能力測量方式的不同而有明顯差異,然而評估範式、後訓練和語言資源可用性之間的互動尚未被系統性地檢視。
我們使用四種評估方法,對六個家族的基本模型和後訓練模型在 MultiBLiMP 上進行評估,這是一個涵蓋 101 種語言的句法最小對比基準。
我們報告了三個主要發現。
首先,後訓練會降低語法能力,但這一影響的程度因模型規模而不均勻地減少,而低資源語言承擔了最高的成本。
其次,後訓練模型保留了它們無法通過明確提示表達的語法知識,但這僅在高資源語言中可測量,因為在低資源環境中接近隨機的基線幾乎沒有知識可隱藏。
第三,母語提示恢復了在低資源語言上隱藏的能力,顯示只有高資源語言可以直接從未提示的概率中探測。
我們總結認為,多語言語法評估必須採用語言知情的多範式協議,以避免系統性低估低資源能力。
EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery
2609.40340v1 by Young-Jun Lee, Jinheon Baek, Soyeong Jeong, Minki Kang, Seungyeon Jwa, Jonghyun Choi, Seungho Han, Dongyeop Kang
Evolutionary search with large language models (LLMs) can stall when progress requires external knowledge the model lacks. Supplying relevant documents helps, but simply adding web search tool can keep returning the same pages as solutions change. We introduce EvoDuet, a bi-level optimization method that co-evolves solutions and search queries with fixed model parameters. At each iteration, a retrieval gate lets the LLM assess its knowledge gap and choose to retrieve new documents, reuse stored ones, or proceed without them. An inner loop refines queries and ranks documents by the solution scores they are predicted to yield; an outer loop generates candidates in parallel from these documents and records the evaluated outcomes for later searches. Across 21 optimization tasks with one candidate per iteration, EvoDuet raises OpenEvolve's normalized discovery gain from 74.1% to 78.0% with GPT-5.6-Luna and from 61.3% to 82.3% with Gemini-3.8-Flash, whereas Qwen3.5-9B does not benefit. Our best runs surpass the previously reported best scores on eight tasks, including Swap Reduction on Q20 and Rosetta, and match them on three more. EvoDuet also improves with other scaffolds (e.g., Top-K, EvoX) on Sums/Diffs and Denoising, demonstrating its applicability across evolutionary search scaffolds.
摘要:進化搜尋與大型語言模型(LLMs)結合時,當進展需要模型缺乏的外部知識時,可能會停滯不前。提供相關文件有助於改善情況,但僅僅添加網路搜尋工具可能會在解決方案變化時持續返回相同的頁面。我們介紹EvoDuet,一種雙層優化方法,通過固定的模型參數共同進化解決方案和搜尋查詢。在每次迭代中,檢索閘讓LLM評估其知識差距,並選擇檢索新文件、重用存儲的文件或在沒有它們的情況下繼續。內部循環精煉查詢並根據預測產生的解決方案分數對文件進行排名;外部循環則從這些文件中平行生成候選項,並記錄評估結果以便後續搜尋。在21個優化任務中,每次迭代一個候選項,EvoDuet使OpenEvolve的標準化發現增益從74.1%提高到78.0%(使用GPT-5.6-Luna),並從61.3%提高到82.3%(使用Gemini-3.8-Flash),而Qwen3.5-9B則沒有受益。我們的最佳運行超過了先前報告的八個任務的最佳分數,包括Q20和Rosetta上的Swap Reduction,並在另外三個任務上達到相同的分數。EvoDuet在其他支架(例如Top-K、EvoX)上也在Sums/Diffs和去噪中有所改善,顯示其在進化搜尋支架中的適用性。
Linguistic Loopholes in LLM Unlearning: From a 174-Language Benchmark to Coverage-Aware Unlearning
2609.40286v1 by Tyler Skow, Shravan Chaudhari, Rama Chellappa, Abhay Yadav
Unlearning a fact in one language does not guarantee its removal in others as changing the query or even the requested answer language can reopen seemingly forgotten knowledge -- a cross-lingual loophole. The most straightforward solution to this challenge -- unlearning in all languages -- is neither scalable nor desirable as it amplifies damage to unrelated model capabilities. We introduce the task of language budgeted multilingual unlearning where the goal is to select a subset of languages that maximizes cross-lingual erasure. To study this task we introduce the Cross-Lingual Unlearning Tensor, an unlearning benchmark that spans 174 language--script pairs and 25 atomic paraphrase types to examine when forgetting generalizes across linguistic expressions of the same knowledge. We further propose COVER, which selects source languages to maximize predicted COVERage of languages receiving no forget supervision, enabling unlearning on a language budget. Surprisingly, we find naively selecting strong individual sources does not reliably compose into strong source sets motivating our development of COVER. At deployment COVER only requires benign calibration data and access to the frozen model. Across three model families and two disjoint forget sets, COVER reduces mean held-out residual access by 7.8--27.3% relative to uniform source selection. We find these gains extend beyond synthetic benchmarks to real news documents in low-resource language settings using human translated data from the Low Resource Languages for Emergent Incidents (LORELEI) corpus.
摘要:在一種語言中忘記一個事實並不保證它在其他語言中也會被移除,因為改變查詢或甚至請求的答案語言可能會重新打開看似被遺忘的知識——這是一種跨語言的漏洞。解決這一挑戰的最直接方案——在所有語言中忘記——既不可擴展也不可取,因為這會加劇對無關模型能力的損害。我們引入了語言預算多語言忘記的任務,其目標是選擇一組語言,以最大化跨語言的抹去。為了研究這一任務,我們引入了跨語言忘記張量,這是一個涵蓋174種語言-書寫對和25種原子改述類型的忘記基準,以檢查何時忘記在相同知識的語言表達中會泛化。我們進一步提出了COVER,它選擇源語言以最大化未接受忘記監督的語言的預測COVERage,從而實現語言預算下的忘記。令人驚訝的是,我們發現天真地選擇強大的個別源並不可靠地組成強大的源集合,這促使我們開發COVER。在部署時,COVER僅需要良性的校準數據和對凍結模型的訪問。在三個模型系列和兩個不相交的忘記集上,COVER相對於均勻源選擇減少了7.8-27.3%的平均保留殘餘訪問。我們發現這些增益超越了合成基準,擴展到使用來自低資源語言緊急事件(LORELEI)語料庫的人類翻譯數據的真實新聞文件。
EviRover: Reinforcing Agentic Perception Beyond a Glance
2609.40230v1 by Kaixuan Fan, Kaituo Feng, Tianshuo Peng, Yilei Jiang, Manyuan Zhang, Junke Wang, Xiangyu Yue
Visual perception is conventionally formulated as a one-shot prediction from a single glance at the image, under the assumption that the image content and the model's parametric knowledge suffice to resolve the query. This assumption often fails in real-world scenarios that hinge on fine-grained visual details or require knowledge-intensive and up-to-date information. We term such cases \textit{perception under insufficient evidence} and formulate perception as an agentic process that can obtain information beyond a single glance. To address the absence of data for this setting, we design two dedicated data generation pipelines, yielding EviRover-SFT-5K and EviRover-RL-12K for training. We further construct EviLens, a human-verified benchmark comprising 688 instances across five perception categories. Building on these data, we present EviRover, to our knowledge the first perception agent explicitly trained to resolve perceptual queries through interaction, using supervised fine-tuning followed by agentic reinforcement learning. Experiments show that the 4B EviRover outperforms its backbone by 30 points on average on EviLens, reaching performance comparable to advanced proprietary models. The gains transfer beyond EviLens to WebEyes, conventional perception benchmarks, and general multimodal benchmarks, including a 15-point improvement on BrowseComp-VL. All code, models, and data are released.
摘要:視覺感知通常被表述為從對圖像的一瞥中進行的一次性預測,假設圖像內容和模型的參數知識足以解決該查詢。這一假設在依賴細微視覺細節或需要知識密集和最新信息的現實場景中經常失效。我們將這類情況稱為\textit{在不足證據下的感知},並將感知表述為一種能夠獲取超越單一瞥見的信息的主動過程。為了解決這一情境下數據的缺乏,我們設計了兩個專門的數據生成管道,產生了用於訓練的EviRover-SFT-5K和EviRover-RL-12K。我們進一步構建了EviLens,一個經人類驗證的基準,包含五個感知類別的688個實例。在這些數據的基礎上,我們提出了EviRover,據我們所知,這是第一個明確訓練以通過互動解決感知查詢的感知代理,使用監督微調隨後進行主動強化學習。實驗顯示,4B的EviRover在EviLens上的表現平均比其基礎模型高出30分,達到與先進專有模型相當的性能。這些增益超越EviLens,轉移到WebEyes、傳統感知基準和一般多模態基準,包括在BrowseComp-VL上提高15分。所有代碼、模型和數據均已發布。
Learning from Research: Toward Lifelong Agent Harness Evolution
2609.40169v1 by Jingbo Yang, Kwei-Herng Lai, Xiaowen Wang, Yaar Harari, Evgeniy Gabrilovich, Shiyu Chang
Language agents are expected to solve increasingly complex tasks, creating a growing need for continual improvement. One promising approach is to evolve the agent harness, the software that governs tool use, memory management, and task execution, while keeping the underlying language model fixed. Recent methods automate this process by using a meta coding agent to modify the harness based on execution feedback. However, relying on that agent's existing knowledge and observed failures can restrict exploration and make adaptation reactive. Inspired by how human experts learn from the research literature for new solutions, we introduce ScholarEvolve, a framework that automatically draws on state-of-the-art research to guide harness evolution. ScholarEvolve organizes the harness evolution directions into functional modules and uses topic modeling to identify distinct improvement strategies for each module. It implements these strategies and evaluates their combinations to improve task performance. Moreover, the framework is designed to incorporate new publications over time, allowing research advances to drive proactive lifelong evolution. Experiments demonstrate improvements on AppWorld and Tau2-Bench. ScholarEvolve raises Qwen3.5-27B task goal completion from 49.6% to 63.6% on AppWorld Challenge, and raises GPT-5.4-mini pass@1 from 72.7% to 81.9% on Tau2-Bench Telecom.
摘要:語言代理預期能解決日益複雜的任務,這創造了持續改進的需求。
一種有前景的方法是進化代理工具,即管理工具使用、記憶管理和任務執行的軟體,同時保持基礎語言模型不變。
最近的方法通過使用元編碼代理自動化這一過程,根據執行反饋來修改工具。
然而,依賴該代理的現有知識和觀察到的失敗可能會限制探索,並使適應變得被動。
受到人類專家如何從研究文獻中學習新解決方案的啟發,我們引入了ScholarEvolve,一個自動利用最先進研究來指導工具進化的框架。
ScholarEvolve將工具進化方向組織成功能模組,並使用主題建模來識別每個模組的不同改進策略。
它實施這些策略並評估它們的組合以改善任務性能。
此外,該框架設計為隨著時間的推移納入新出版物,允許研究進展推動主動的終身進化。
實驗顯示在AppWorld和Tau2-Bench上有所改善。
ScholarEvolve將Qwen3.5-27B在AppWorld Challenge上的任務目標完成率從49.6%提高到63.6%,並將GPT-5.4-mini在Tau2-Bench Telecom上的pass@1從72.7%提高到81.9%。
On the (In)effectiveness of AMR Augmentation for Large Language Models
2609.40121v1 by Hoa Quynh Nhung Nguyen, Jacopo Staiano, Michael Sullivan
While Abstract Meaning Representation (AMR) has historically improved performance on a range of NLP tasks, the benefit---or lack thereof---of AMR augmentation for modern LLMs is thus far unclear. In this paper, we attempt to reproduce recent work that reported substantial downstream gains from AMR augmentation, finding that these are likely due to specific choices in the experimental settings used: using a consistent and unified protocol for hyperparameter selection, we observe that text-only baselines consistently match or exceed the performance of AMR-augmented models. To investigate this null result, we introduce a perplexity-based probe measuring the degree to which AMR provides an LLM with supplemental relational knowledge not already available to the model. We find that AMR augmentation does not help LLMs improve their understanding of relational content in the sentence, indicating that augmenting these models with AMR offers no clear benefit on downstream tasks.
摘要:雖然抽象意義表示法(AMR)歷來在多種自然語言處理任務中提高了性能,但AMR增強對現代大型語言模型的好處——或缺乏好處——至今仍不明確。
在本文中,我們試圖重現最近的研究,該研究報告了AMR增強帶來的顯著下游收益,發現這些收益可能是由於實驗設置中的特定選擇:使用一致且統一的超參數選擇協議,我們觀察到僅使用文本的基準模型在性能上始終與AMR增強模型相匹配或超過。
為了調查這一無效結果,我們引入了一種基於困惑度的探測器,測量AMR為大型語言模型提供額外關聯知識的程度,而這些知識在模型中並不存在。
我們發現AMR增強並未幫助大型語言模型改善對句子中關聯內容的理解,這表明用AMR增強這些模型在下游任務中並未提供明顯的好處。
Persistent Context Graphs for Efficient Memory Compaction in LLM Agents
2609.40118v1 by Jingbo Yang, Kwei-Herng Lai, Xiaowen Wang, Zhaoxuan Tan, Pei Zhou, Mengting Wan, Yaar Harari, Evgeniy Gabrilovich, Shiyu Chang
As LLM capabilities advance, agents are tackling increasingly complex tasks over longer horizons. Their growing interaction histories make memory compaction essential for staying within context windows and reducing prefill cost. Existing methods summarize the history or compress its KV cache, often adding model computation to preserve information for future requests. A new user request can change which history matters, but reassessing that history with the model requires re-encoding it if the KV cache has expired. Past attention provides signals of historical importance and dependencies between messages, while relevance to the current task must be assessed using the new user request. We introduce ReCAP, a memory compaction method that stores attention-derived importance scores and dependency links in a lightweight, persistent context graph. For each new request, ReCAP combines stored importance with relevance cues from the request and follows dependency links to select messages and their supporting context, without additional model calls for selection. Compared with Codex's default summarization-based compaction, ReCAP reduces estimated latency for compaction and cold restoration by approximately 95% on both Qwen3-Coder and gpt-oss. It also roughly halves the historical context per call on SWE-Together at comparable task quality and improves accuracy on the code tasks of Lost-in-Conversation over full history by 19.8 and 41.2 points.
摘要:隨著大型語言模型(LLM)能力的提升,代理正在處理越來越複雜的任務,並且時間範圍也越來越長。它們日益增長的互動歷史使得記憶壓縮對於保持在上下文窗口內和降低預填成本變得至關重要。現有的方法總結歷史或壓縮其KV快取,通常會增加模型計算以保留未來請求的信息。新的用戶請求可能會改變重要的歷史,但如果KV快取已過期,則需要重新編碼該歷史以重新評估它。過去的注意力提供了歷史重要性和消息之間依賴性的信號,而與當前任務的相關性必須使用新的用戶請求來評估。我們介紹了ReCAP,一種記憶壓縮方法,將基於注意力的衍生重要性分數和依賴鏈存儲在輕量級的持久上下文圖中。對於每個新的請求,ReCAP將存儲的重點與請求中的相關提示結合,並沿著依賴鏈選擇消息及其支持上下文,而無需額外的模型調用來進行選擇。與Codex的默認基於總結的壓縮相比,ReCAP在Qwen3-Coder和gpt-oss上將壓縮和冷恢復的估計延遲減少了約95%。它還在SWE-Together上將每次調用的歷史上下文大約減半,並在任務質量相當的情況下,將Lost-in-Conversation的代碼任務準確性提高了19.8和41.2個點。
JuryFlow: Disagreement-Guided Human-in-the-Loop Multi-Agent Evaluation
2609.40103v1 by Mufeng Yang, Junwei Yu, Yepeng Ding
Large language models (LLMs) are increasingly deployed as automated judges for AI-generated content, yet a single judge is unreliable and even a panel of judges leaves a hard residue: when judges disagree, majority voting discards the conflict instead of resolving it. We present JuryFlow, a disagreement-guided, human-in-the-loop multi-agent evaluation framework that treats inter-judge disagreement not as noise to be averaged away, but as a precise, claim-level signal indicating where an evaluation is uncertain. JuryFlow decomposes each candidate response into atomic claims, has a panel of heterogeneous judges assign per-claim verdicts, and builds a disagreement graph whose nodes are scored by verdict entropy and whose edges encode structural similarity between claims. A human acts as a structural guide, selecting which disagreement to resolve through a single, minimal intervention rather than re-labeling the response, after which the focal claim is re-evaluated, the correction propagates along graph edges and to historically similar cases, and is crystallized into reusable rubric entries that all judges inherit, making the evaluator progressively self-refining. To enable large-scale, reproducible benchmarking without human studies, we evaluate JuryFlow in an automatic configuration in which focal selection is made by entropy ranking. On MT-Bench and LLMBar, JuryFlow improves agreement with gold labels over single-judge and majority-vote panel baselines, and ablations isolate the contributions of disagreement-targeted re-evaluation, propagation, and rubric induction. We contribute (1) a human-in-the-loop paradigm that recasts the human from labeler to structural guide, (2) the JuryFlow framework operationalizing it through a disagreement graph, focal re-evaluation, and closed-loop rubric induction, and (3) an evaluation protocol with ablations that isolate where the gains originate.
摘要:大型語言模型(LLMs)越來越多地被用作自動評判AI生成內容的法官,但單一法官不可靠,即使是一組法官也會留下難以解決的問題:當法官意見不合時,多數投票會忽略衝突,而不是解決它。
我們提出了JuryFlow,一個以不一致為指導的、人機協作的多代理評估框架,它將法官之間的不一致視為一種精確的、聲明級別的信號,指示評估的不確定性,而不是要被平均掉的噪音。
JuryFlow將每個候選回應分解為原子聲明,讓一組異質法官對每個聲明給予裁決,並構建一個不一致圖,其節點由裁決熵進行評分,邊則編碼聲明之間的結構相似性。
一名人類作為結構指導,選擇通過單一的最小干預來解決哪個不一致,而不是重新標記回應,之後重新評估焦點聲明,修正沿著圖邊和歷史上相似的案例進行傳播,並被凝結成所有法官繼承的可重用評分標準條目,使評估者逐漸自我精煉。
為了實現大規模、可重複的基準測試而不需要人類研究,我們在自動配置中評估JuryFlow,其中焦點選擇由熵排名決定。
在MT-Bench和LLMBar上,JuryFlow提高了與金標籤的協議,相較於單一法官和多數投票小組基準,並且消融實驗隔離了針對不一致的重新評估、傳播和評分標準引入的貢獻。
我們貢獻了(1)一個人機協作的範式,將人類從標記者重新塑造為結構指導,(2)通過不一致圖、焦點重新評估和閉環評分標準引入來實現這一範式的JuryFlow框架,以及(3)一個評估協議,通過消融實驗隔離增益的來源。
AutoDataBench: A Data-centric Testbed for Accelerating Auto Research
2609.40097v1 by Ruifeng Yuan, Yizhi Li, Yaxin Du, Fengyu Cai, Yiqi Liu, Hou Pong Chan, Chenghua Lin, Yun Chen, Jian Yang, Bryan Dai, Pinyan Lu, Chenghao Xiao
Existing auto-research benchmarks often entangle multiple sources of improvement, including training frameworks, hyperparameters, compute budgets, and data, making it difficult to attribute why one frontier agent outperforms another to specific research capabilities. In this work, we isolate and systematically evaluate Data Intelligence: an agent's ability to understand, manipulate, and improve the data that shapes model capabilities. We introduce AutoDataBench, a controlled testbed built on a conceptual framework of data intelligence spanning data diagnosis, data organization, and data construction, instantiated through three highly curated optimization tasks while holding non-data factors fixed. Across tool use, retrieval, and knowledge injection, we evaluate frontier LLMs' ability to improve training data through iterative experimentation under task-specific resource budgets. Beyond optimization performance, we ask: do LLMs understand what their data interventions do? We compare predictions made before training with observed outcomes to seek evidence of data-effect reasoning beyond trial and error, and explore whether iterative feedback helps LLMs better understand how changes to training data affect model performance. Finally, we show that reusing AutoDataBench trajectories for mid-training improves downstream coding performance, highlighting its value in both evaluating data intelligence and generating high-quality training data. Code and resources are available at https://github.com/AutoDataBench/AutoDataBench.
摘要:現有的自動研究基準經常將多個改進來源糾纏在一起,包括訓練框架、超參數、計算預算和數據,這使得難以將一個前沿代理的優越表現歸因於特定的研究能力。在這項工作中,我們孤立並系統地評估數據智能:一個代理理解、操作和改善塑造模型能力的數據的能力。我們介紹了AutoDataBench,一個基於數據智能概念框架的受控測試平台,涵蓋數據診斷、數據組織和數據構建,通過三個高度策劃的優化任務實現,同時保持非數據因素不變。在工具使用、檢索和知識注入方面,我們評估前沿LLM在特定任務資源預算下通過迭代實驗改善訓練數據的能力。除了優化性能,我們還問:LLM是否理解它們的數據干預所做的事情?我們比較訓練前的預測與觀察到的結果,以尋求超越試錯的數據效果推理證據,並探索迭代反饋是否有助於LLM更好地理解訓練數據變化如何影響模型性能。最後,我們展示了在中期訓練中重用AutoDataBench軌跡改善下游編碼性能,突顯了其在評估數據智能和生成高質量訓練數據方面的價值。代碼和資源可在 https://github.com/AutoDataBench/AutoDataBench 獲得。
LongEmo: Towards Emotion Understanding and Reasoning in Long Videos
2609.40079v1 by Shuo Zhang, Yifan Zhou, Han Wang, Jinsong Zhang, Jingyu Li, Hongbing Li, Zhejun Zhang, Chengyi Zhao, Yuquan Hao, Yitong Liu, Jiyin Li, Ruiqi Tang, Zixuan Lin, Yi Luo, Xurui Zhang, Ronghao Chen, Huacan Wang, Lei Li
While recent Multimodal Large Language Models (MLLMs) have shown promise in affective computing, their reasoning capabilities are largely confined to short video clips with limited interactions. However, real-world emotions are not merely isolated instantaneous reactions but dynamic and cumulative processes deeply shaped by past experiences and ongoing events. To bridge this gap, we introduce LongEmoBench, a benchmark dedicated to emotion understanding and reasoning in long videos. It assesses progressive capabilities scaling from continuous scene interactions to complex episodic developments. Furthermore, we propose LongEmo, a novel memory-augmented agentic framework designed to tackle the immense challenges of long-range affective reasoning. LongEmo processes continuous video streams to construct an Event Memory Graph, explicitly modeling long-range dependencies and capturing emotional dynamics across discrete events. Given a question, the agent retrieves a query-relevant event stream from the graph, iteratively integrating multimodal memories and relational dependencies to deduce the final answer. Extensive evaluations of 17 representative methods reveal that they struggle significantly with emotion understanding and reasoning in long videos. In contrast, LongEmo achieves state-of-the-art performance, demonstrating the efficacy of its event-centric memory architecture.
摘要:最近的多模態大型語言模型(MLLMs)在情感計算方面顯示出潛力,但它們的推理能力主要限於短視頻片段,互動性有限。
然而,現實世界中的情感並不僅僅是孤立的瞬時反應,而是深受過去經驗和當前事件影響的動態和累積過程。
為了縮小這一差距,我們引入了LongEmoBench,一個專注於長視頻中的情感理解和推理的基準。
它評估從連續場景互動到複雜情節發展的逐步能力。
此外,我們提出了LongEmo,一個新穎的增強記憶的代理框架,旨在應對長期情感推理的巨大挑戰。
LongEmo處理連續視頻流以構建事件記憶圖,明確建模長期依賴關係並捕捉離散事件中的情感動態。
在給定問題的情況下,代理從圖中檢索與查詢相關的事件流,迭代整合多模態記憶和關係依賴,以推導最終答案。
對17種代表性方法的廣泛評估顯示,它們在長視頻中的情感理解和推理方面面臨重大挑戰。
相比之下,LongEmo實現了最先進的性能,展示了其以事件為中心的記憶架構的有效性。
TACTIC: Temporal and Context-Aware LLM Tactical Planning for Roadside LiDAR Attacks
2609.39969v1 by Yiming Gao, Shaocheng Luo
Physical LiDAR attacks are often evaluated using fixed primitives and manually selected parameters, despite their strong dependence on surrounding traffic. We present TACTIC, a scene-aware framework that uses a multimodal large language model (MLLM) to coordinate state-adaptive roadside LiDAR attacks. Under a gray-box threat model, TACTIC relies only on an attacker-operated roadside perception stack, without accessing the victim LiDAR's native point clouds or internal processing. Local perception provides metric vehicle states, while the MLLM combines these measurements with roadside imagery to infer relational traffic context and construct a semantic scene graph. Based on this representation, TACTIC selects and configures two complementary primitives: \emph{push-away}, which shifts the perceived range of a lead vehicle, and \emph{phantom-obstacle braking}, which triggers emergency braking through obstacle injection. Measured traffic states and empirically calibrated constraints ground the generated tactics in physically feasible operating regions. To accommodate MLLM latency, TACTIC overlaps reasoning and execution asynchronously while high-rate local perception detects scene changes and triggers replanning. Across 280 randomized CARLA trials, the full policy achieves a 100% collision rate, compared with 35% for a fixed rule, 60% for random selection, and 75% for a restricted LLM using mode selection with default parameters. Joint physical-and-image input achieves 100% success, versus 65% with physical measurements alone and 75% with imagery alone, while asynchronous $Δ$ refresh reduces scene-mutation response from 7.4 s to 2.0 s. These results show that scene-dependent tactical planning can expose context-sensitive LiDAR failure modes that fixed attack policies may miss.
摘要:物理LiDAR攻擊通常使用固定的原始元素和手動選擇的參數進行評估,儘管它們強烈依賴於周圍的交通。我們提出了TACTIC,一個場景感知框架,利用多模態大型語言模型(MLLM)來協調狀態自適應的路邊LiDAR攻擊。在灰盒威脅模型下,TACTIC僅依賴攻擊者操作的路邊感知堆疊,而不訪問受害者LiDAR的原始點雲或內部處理。當地感知提供度量車輛狀態,而MLLM將這些測量與路邊影像結合,以推斷關聯交通上下文並構建語義場景圖。基於這一表示,TACTIC選擇並配置兩個互補的原始元素:\emph{推開},它改變前方車輛的感知範圍,以及\emph{幻影障礙物制動},它通過障礙物注入觸發緊急制動。測量的交通狀態和經驗校準的約束將生成的戰術基於物理可行的操作區域。為了適應MLLM的延遲,TACTIC在高頻率的當地感知檢測場景變化並觸發重新規劃的同時,異步重疊推理和執行。在280次隨機化的CARLA試驗中,完整策略實現了100%的碰撞率,而固定規則為35%,隨機選擇為60%,使用默認參數的受限LLM的模式選擇為75%。聯合物理和影像輸入達到100%的成功率,而僅使用物理測量的成功率為65%,僅使用影像的成功率為75%,同時異步$Δ$刷新將場景變異響應從7.4秒減少到2.0秒。這些結果顯示,依賴場景的戰術規劃可以揭示固定攻擊政策可能忽略的上下文敏感LiDAR失效模式。
AIMS: An Agentic AI Framework for Sim-to-Real Multi-Modal ISAC
2609.39964v2 by Yijie Bian, Kai Zhang, Wei Guo, Zixin Wang, Shenghui Song, Jun Zhang, Khaled B. Letaief
Multi-modal integrated sensing and communication (ISAC) enables environmental perception and reliable connectivity for intelligent wireless networks. Data-driven multi-modal ISAC models depend heavily on annotated real-world data to learn relationships across sensing and wireless observations, thereby constraining scalable deployment. Although synthetic data generation reduces the burden, adapting existing simulation pipelines to a target deployment requires consistent scene, sensing, wireless, and learning configurations, while mismatches among these coupled components impair sim-to-real transferability. To address the challenge, we propose an agentic artificial intelligence (AI) framework for sim-to-real multi-modal ISAC, named AIMS. Given a natural-language deployment request specifying the target task, deployment conditions, and real-data budget, AIMS derives a deployment-specific sim-to-real configuration and coordinates its execution to produce a deployment-specific task model. A two-agent architecture coordinates scene construction with task learning. A scene construction agent generates geographically grounded, synchronized sensing and wireless records from shared physical states, while a scene understanding agent configures task-relevant modalities and mixture-of-experts (MoE) learning for zero-shot inference or few-shot adaptation. Structured domain knowledge guides dependency-aware planning, while validation evidence supports feedback-driven revision of affected decisions. Experiments on the real-world DeepSense 6G dataset demonstrate improved vehicle detection and beam prediction over the considered simulation and fusion baselines. A separate orchestration benchmark evaluates task interpretation, dependency reasoning, and feedback-driven replanning across diverse deployment requests, showing improved plan correctness with structured domain knowledge and validation feedback.
摘要:多模態整合感知與通信(ISAC)使智能無線網絡能夠進行環境感知和可靠連接。基於數據驅動的多模態ISAC模型在學習感知和無線觀測之間的關係時,重度依賴標註的真實世界數據,從而限制了可擴展的部署。儘管合成數據生成減輕了負擔,但將現有的模擬管道適應於目標部署需要一致的場景、感知、無線和學習配置,而這些耦合組件之間的錯配會損害模擬到現實的可轉移性。為了解決這一挑戰,我們提出了一個名為AIMS的代理人工智能(AI)框架,用於模擬到現實的多模態ISAC。給定一個自然語言的部署請求,具體說明目標任務、部署條件和真實數據預算,AIMS推導出一個特定於部署的模擬到現實配置,並協調其執行以生成特定於部署的任務模型。兩個代理架構協調場景構建與任務學習。一個場景構建代理從共享的物理狀態生成地理基礎的、同步的感知和無線記錄,而場景理解代理則配置與任務相關的模態和專家混合(MoE)學習,以進行零樣本推理或少樣本適應。結構化的領域知識指導依賴意識的規劃,而驗證證據支持受影響決策的反饋驅動修訂。在現實世界的DeepSense 6G數據集上的實驗顯示,與考慮的模擬和融合基準相比,車輛檢測和波束預測有所改善。一個單獨的協調基準評估任務解釋、依賴推理和跨多樣化部署請求的反饋驅動重新規劃,顯示出結構化的領域知識和驗證反饋提高了計劃的正確性。
Learning to Cover Locally: Graph Neural Combinatorial Optimization under a Hard Information Horizon
2610.00422v1 by Johannes F. Loevenich, Thies Moehlenhof, Laurin Holz, Maxime Schwarzer, Tobias Huerten, Roberto Rigolin F. Lopes
Neural combinatorial optimization typically assumes a centralized solver that reads the whole instance. We study the opposite: combinatorial optimization under a hard information horizon, where every node commits to its share of a global solution seeing only its $k$-hop neighborhood, and those commitments must compose into a globally feasible solution. We formalize this as local set cover and instantiate it on weighted multipoint relay (MPR) selection, the NP-hard 2-hop covering problem of the Optimized Link State Routing Protocol version 2 (OLSRv2) routing protocol (RFC~7181), whose horizon is imposed by the protocol, not chosen by the modeler. We prove two results. Any deterministic selector whose horizon is one hop short must either fail coverage or land a factor $Δ$ from optimal, and an $L$-layer graph neural network (GNN) read out at the deciding node is exactly an $L$-hop selector, so capacity cannot buy back radius. Conversely, at the horizon a \ac{GNN} of depth $O(Δ)$ reproduces the RFC~7181 covering greedy, and at width $O(c_{\max}Δ)$ its metric-aware weighted analogue, inheriting the $(1+\lnΔ_2)$-approximation in both cases. Empirically, a 3-layer \ac{GATv2} with a coverage-completing decoder, behavior-cloned from the CP-SAT optimum, reaches $\text{cost}/\text{opt}=1.030\pm0.001$ against greedy's $1.138$, closing $79.1\%$ of the gap at $100\%$ coverage. Restricting the same learner to one hop, on identical instances with the same decoder and demonstrations, collapses it to $1.344$, far worse than greedy. Two transfer checks target real-world networks. OLSRv2's unmodified selection code matches our cardinality greedy on $200/200$ unit-cost instances, and on $40{,}308$ instances of real battalion mobility the frozen model closes $48\%$ of the gap at full coverage. The information horizon, not the model capacity, is the most significant variable.
摘要:神經組合優化通常假設有一個集中式解算器,能夠讀取整個實例。我們研究相反的情況:在硬信息視野下的組合優化,其中每個節點僅能看到其 $k$ 跳鄰域,並承諾其在全局解中的份額,而這些承諾必須組合成一個全局可行的解。我們將其形式化為局部集合覆蓋,並在加權多點中繼(MPR)選擇上實例化,這是優化鏈路狀態路由協議版本 2(OLSRv2)路由協議(RFC~7181)的 NP 困難 2 跳覆蓋問題,其視野由協議強加,而非由建模者選擇。我們證明了兩個結果。任何其視野短一跳的確定性選擇器必須要麼失敗於覆蓋,要麼與最佳解相差一個因子 $Δ$,而在決策節點讀出的 $L$ 層圖神經網絡(GNN)恰好是一個 $L$ 跳選擇器,因此容量無法贖回半徑。相反,在視野內,深度為 $O(Δ)$ 的 \ac{GNN} 重現了 RFC~7181 的貪婪覆蓋,而在寬度為 $O(c_{\max}Δ)$ 時,其度量感知的加權類比也繼承了這兩種情況下的 $(1+\lnΔ_2)$ 近似。實證上,一個 3 層的 \ac{GATv2} 配備了一個完成覆蓋的解碼器,從 CP-SAT 最優解行為克隆而來,達到了 $\text{cost}/\text{opt}=1.030\pm0.001$,而貪婪的為 $1.138$,在 $100\%$ 覆蓋時縮小了 $79.1\%$ 的差距。將相同的學習者限制為一跳,在相同的實例中使用相同的解碼器和示範,則崩潰至 $1.344$,遠比貪婪差。兩個轉移檢查針對現實世界的網絡。OLSRv2 的未修改選擇代碼在 $200/200$ 單位成本實例中與我們的基數貪婪相匹配,而在 $40{,}308$ 個真實營移動實例中,凍結模型在完全覆蓋時縮小了 $48\%$ 的差距。信息視野,而非模型容量,是最重要的變量。
MCD: Causal Distillation of Multimodal In-Context Learning in Large Vision-Language Models
2609.39920v1 by Yanshu Li, Jiaqian Li, Canran Xiao, Xi Xiao, Tianyang Wang, Yongtai Liu
Large vision-language models (LVLMs) exhibit strong multimodal in-context learning (ICL) capabilities, yet this ability degrades substantially as model size decreases. Knowledge distillation offers a natural way to bridge this gap, but existing methods primarily align output distributions or hidden representations directly. Such alignment teaches the student what the teacher predicts without revealing which evidence in the complex context causally supports that prediction. Consequently, a student can imitate the teacher's answer while continuing to rely on language priors, prompt structure, or other spurious cues. To address this limitation, we introduce Multimodal Causal Distillation (MCD), a distillation framework that transfers how a strong teacher uses multimodal evidence during ICL. MCD uses structure-preserving token interventions to identify and verify causal evidence, then transfers how the teacher responds when that evidence is retained or removed. This design connects distillation to the causal patterns by which the model uses contextual evidence during multimodal ICL. Experiments across three LVLM families and seven benchmarks show that MCD improves student performance by 7.23 points on average and outperforms vanilla distillation by 4.68 points, while further analyses confirm the generalizability of these gains.
摘要:大型視覺-語言模型(LVLMs)展現出強大的多模態上下文學習(ICL)能力,但隨著模型大小的減少,這種能力會顯著下降。知識蒸餾提供了一種自然的方式來彌補這一差距,但現有的方法主要是直接對齊輸出分佈或隱藏表示。這種對齊教會學生教師的預測,而不揭示在複雜上下文中哪些證據因果地支持該預測。因此,學生可以模仿教師的答案,同時繼續依賴語言先驗、提示結構或其他虛假線索。為了解決這一限制,我們引入了多模態因果蒸餾(MCD),這是一種蒸餾框架,轉移強教師在ICL過程中如何使用多模態證據。MCD使用結構保持的標記干預來識別和驗證因果證據,然後轉移教師在保留或移除該證據時的反應。這一設計將蒸餾與模型在多模態ICL過程中使用上下文證據的因果模式聯繫起來。對三個LVLM家族和七個基準的實驗顯示,MCD平均提高了學生的表現7.23分,並且比普通蒸餾提高了4.68分,同時進一步分析確認了這些增益的可泛化性。
The Concrete-Arbitrary Gap: Kinship Reasoning in LLMs Is Not Indifferent to Presentation
2609.39913v1 by Thomas Pashby
We test whether large language models solve formally matched kinship problems equally well when relations are expressed in familiar vocabulary or by explicitly defined nonce predicates. Across 500 paired graphs, concrete accuracy exceeds arbitrary accuracy by 35.6 percentage points in local Qwen3.8-27B, 26.6 in Gemma 4 26B-A4B, 12.0 in Gemma 4 31B, and 5.4 in Qwen3.8-Max. All four paired gaps are statistically resolved. Reasoning budgets and prompt-language interventions can substantially reduce the difference, showing that it is modifiable rather than a fixed incapacity. The minimal conclusion is behavioral: on these tasks, the models' manifested relational competence is not indifferent to presentation. Explicit definitions provide the formal relations but do not make nonce predicates as usable as familiar vocabulary embedded in learned linguistic associations.
摘要:我們測試大型語言模型在關係以熟悉詞彙或明確定義的臨時謂詞表達時,是否能同樣有效地解決正式匹配的親屬關係問題。在500對圖形中,具體準確度在本地 Qwen3.8-27B 中超過任意準確度35.6個百分點,在 Gemma 4 26B-A4B 中為26.6,在 Gemma 4 31B 中為12.0,在 Qwen3.8-Max 中為5.4。這四個配對差距在統計上都是顯著的。推理預算和提示語言干預可以大幅減少這一差異,顯示出這是一種可調整的能力,而不是固定的無能。最基本的結論是行為性的:在這些任務中,模型所表現出的關係能力對於呈現方式並不無所謂。明確的定義提供了正式的關係,但並未使臨時謂詞如同嵌入在學習語言聯想中的熟悉詞彙那樣可用。
DoGBench: Can Agents Meet Expert Standards for User-Facing Documentation?
2609.39909v1 by Frances Liu, Manny Silva, Paige Calvert, Ayu Adiati, Sarah Sanders
We introduce DoGBENCH (Documentation Generation Benchmark), to our knowledge, the first benchmark for generating and maintaining real user-facing software documentation. It asks whether an agent can produce documentation that experienced technical writers would accept in review. The benchmark contains 292 items from open source projects, including Helm, PostHog, and Mautic. Each item gives the agent a pre-change repository and a trigger, such as a code pull request or a reported documentation gap. The agent must first decide whether the documentation needs an update. For items that need one, the agent must produce an acceptable patch in one attempt. For items that do not need updates, the agent must abstain. Task-specific rubrics, validated with project maintainers, score each patch on accuracy, completeness, reader guidance, placement, and repository conventions. The composite score combines patch quality with correct abstention, and a score of 100 means an agent meets every requirement for the task. Scores should not be interpreted as a percentage of an expert's capability. We evaluated seven agents. The highest-scoring agent reached 47.3 out of 100 on the 117-item held-out split. In a separate audit of 1,267 patches, the most common failure modes were task-completion gaps (45.5%), technical inaccuracies (36.6%), and incomplete conceptual or reference coverage (32.5%). Analysis of the corresponding trajectories identified three key patterns associated with these failures: (1) describing interfaces without examining how readers use them (36.0%), (2) missing decisive evidence and filling the gaps with plausible assumptions (33.1%), and (3) stopping after finding the first plausible documentation surface and leaving other affected pages stale (30.1%).
摘要:我們介紹 DoGBENCH(文檔生成基準),據我們所知,這是第一個用於生成和維護面向真實用戶的軟件文檔的基準。它詢問一個代理是否能夠生成經驗豐富的技術寫手在審查中會接受的文檔。該基準包含來自開源項目的 292 項內容,包括 Helm、PostHog 和 Mautic。每個項目給代理一個變更前的代碼庫和一個觸發器,例如代碼拉取請求或報告的文檔缺口。代理必須首先決定文檔是否需要更新。對於需要更新的項目,代理必須在一次嘗試中生成可接受的補丁。對於不需要更新的項目,代理必須選擇不作為。特定任務的評分標準經過項目維護者的驗證,根據準確性、完整性、讀者指導、位置和代碼庫慣例對每個補丁進行評分。綜合得分將補丁質量與正確的選擇不作為相結合,得分為 100 意味著代理滿足該任務的所有要求。得分不應被解釋為專家能力的百分比。我們評估了七個代理。得分最高的代理在 117 項保留分割中達到了 47.3 分(滿分 100)。在對 1,267 個補丁的單獨審核中,最常見的失敗模式是任務完成缺口(45.5%)、技術不準確(36.6%)和概念或參考覆蓋不完整(32.5%)。對應的軌跡分析確定了與這些失敗相關的三個關鍵模式:(1)描述接口而不檢查讀者如何使用它們(36.0%),(2)缺少決定性證據並用合理的假設填補空白(33.1%),以及(3)在找到第一個合理的文檔表面後停止,並讓其他受影響的頁面保持過時(30.1%)。
Cognitive Enhancement: Rethinking the Necessity of Role-Playing for Large Language Models
2609.39853v1 by Xingjie Zhuang, Jialong Tang, Chulun Zhou, Buchao Zhan, Zhirui Li, Junhui Li, Yazheng Yang, Jinsong Su
Role-playing prompting has become a popular yet simple technique for improving LLM reasoning and output quality. However, whether it consistently boosts performance across diverse domains remains unclear, as systematic validation is lacking. To fill this gap, we run multi-model, cross-domain, and multilingual experiments on MMLU and MMLU-Redux. We find that gains from role-play prompting depend heavily on model capacity, knowledge domain, and prompt language. Drawing on metacognition theory, we propose the persona-related cognitive alignment hypothesis: role-play works only when the LLM correctly grasps the designated persona and its associated knowledge domain. We test this hypothesis through persona information richness ablation, layer-wise entropy divergence analysis, and latent thought-space deflection observation. To reduce persona cognitive bias and stabilize role-play performance, we propose \textbf{M}ixed-\textbf{L}anguage \textbf{C}oncatenate \textbf{P}rediction \textbf{(MLCP}), a simple, training-free, and efficient multilingual prompt concatenation strategy. It aggregates semantically equivalent role prompts to enrich complementary representational cues. Extensive experiments show that MLCP consistently outperforms vanilla role-play prompting across all tested LLMs.
摘要:角色扮演提示已成為一種流行但簡單的技術,用於改善大型語言模型(LLM)的推理和輸出質量。
然而,這種方法是否在不同領域中始終能提升性能仍不明確,因為缺乏系統性的驗證。
為了填補這一空白,我們在MMLU和MMLU-Redux上進行了多模型、跨領域和多語言的實驗。
我們發現,角色扮演提示的增益在很大程度上取決於模型的能力、知識領域和提示語言。
根據元認知理論,我們提出了與角色相關的認知對齊假設:角色扮演僅在LLM正確理解指定角色及其相關知識領域時有效。
我們通過角色信息豐富度消融、層級熵差異分析和潛在思維空間偏轉觀察來測試這一假設。
為了減少角色認知偏見並穩定角色扮演表現,我們提出了\textbf{M}ixed-\textbf{L}anguage \textbf{C}oncatenate \textbf{P}rediction \textbf{(MLCP)},這是一種簡單、無需訓練且高效的多語言提示串聯策略。
它聚合語義等價的角色提示,以豐富互補的表徵線索。
大量實驗表明,MLCP在所有測試的LLM中始終優於傳統的角色扮演提示。
Explore-on-Graph: Hybrid Embedding-LLM Reasoning for Knowledge Graph Question Answering under Incompleteness
2609.39786v1 by Ola El Khatib, Djellel Difallah
Large language models (LLMs) are increasingly combined with knowledge graphs (KGs) to ground reasoning in structured evidence. However, most LLM-based KGQA methods rely on traversing existing graph edges and become unreliable when reasoning paths are broken by missing facts. Alternatives that ask LLMs to generate missing knowledge risk introducing hallucinated evidence. We introduce XoG (eXplore-on-Graph), a framework for multi-hop question answering over incomplete KGs that recovers missing reasoning paths from learned graph structure rather than LLM parametric knowledge. XoG combines type-level entity-relation statistics to identify candidate relations with KG embeddings to retrieve plausible missing entities, using the LLM as a semantic selector and reasoner. These mechanisms are integrated into an iterative planning-exploration-reasoning process. Experiments on WebQSP, CWQ, and the Wikidata-based BRINK benchmark show that XoG remains competitive on complete KGs and consistently outperforms comparable methods without task-specific KGQA training under KG incompleteness. These gains persist across multiple LLM backbones, indicating that stronger LLMs alone do not resolve missing graph evidence. XoG also reduces LLM token consumption by up to 33% compared with a closely related planning-based approach.
摘要:大型語言模型(LLMs)越來越多地與知識圖譜(KGs)結合,以在結構化證據中進行推理。
然而,大多數基於LLM的KGQA方法依賴於遍歷現有的圖邊,當推理路徑因缺失事實而中斷時,這些方法便變得不可靠。
要求LLM生成缺失知識的替代方案則有引入幻覺證據的風險。
我們介紹了XoG(eXplore-on-Graph),這是一個針對不完整KG的多跳問題回答框架,它從學習到的圖結構中恢復缺失的推理路徑,而不是依賴於LLM的參數知識。
XoG結合了類型級實體-關係統計來識別候選關係,並利用KG嵌入來檢索合理的缺失實體,使用LLM作為語義選擇器和推理者。
這些機制被整合到一個迭代的規劃-探索-推理過程中。
在WebQSP、CWQ和基於Wikidata的BRINK基準上的實驗表明,XoG在完整KG上保持競爭力,並在KG不完整的情況下,始終優於沒有特定任務KGQA訓練的可比方法。
這些增益在多個LLM骨幹中持續存在,表明僅僅強大的LLM並不能解決缺失的圖證據。
與一種密切相關的基於規劃的方法相比,XoG還將LLM的標記消耗減少了多達33%。
MemCodex: Self-Programming Hierarchical Memory for Language Agents
2609.39765v1 by Xiaoqiang Wang, Bang Liu
Agent memory faces heterogeneous access needs: a single-hop question may require one piece of evidence, whereas a multi-hop question must combine evidence from multiple sources. Predefined memory workflows cannot adapt to these varying needs. Recent adaptive methods search or learn over memory components and their compositions, but the design space itself remains predefined. We introduce MemCodex, a self-evolving hierarchical memory system that organizes experience into executable memory programs for summaries, relational knowledge, reusable skills, and latent memory. Open-ended program evolution searches the open design space of layer programs by rewriting how each layer is constructed, indexed, retrieved, and routed, thereby adapting both within-layer implementations and cross-layer composition. At query time, reads traverse the hierarchy from coarse to fine and stop once sufficient evidence is found, descending to the original history when needed. We further develop MemArena, a unified runtime that places heterogeneous data and memory systems behind a common interface. MemCodex improves average task success by 10.1% relative to the strongest adaptive-memory baseline, while using 3.4x fewer context tokens and achieving 2.1x faster inference.
摘要:代理記憶面臨異質的訪問需求:單跳問題可能需要一個證據,而多跳問題必須結合來自多個來源的證據。預定義的記憶工作流程無法適應這些不同的需求。最近的自適應方法在記憶組件及其組合上進行搜索或學習,但設計空間本身仍然是預定義的。我們介紹了MemCodex,一個自我演化的分層記憶系統,將經驗組織成可執行的記憶程序,用於摘要、關聯知識、可重用技能和潛在記憶。開放式程序演化通過重寫每一層的構建、索引、檢索和路由方式,搜索層程序的開放設計空間,從而適應層內實現和跨層組合。在查詢時,讀取從粗到細遍歷層級,並在找到足夠的證據後停止,必要時降回原始歷史。我們進一步開發了MemArena,一個統一的運行時,將異質數據和記憶系統置於共同接口後面。相較於最強的自適應記憶基準,MemCodex提高了平均任務成功率10.1%,同時使用了3.4倍更少的上下文標記,並實現了2.1倍更快的推理。
OverForge: Reasoning Through Strategies and Tactics Helps Cooperative Lifelong Adaptation
2609.39727v1 by Oana Madalina Fron, Ojas Shirekar, Chirag Raman
Cooperative language-model agents must coordinate over long horizons and adapt to changing environments and to partners with unfamiliar conventions, yet existing agents map observations to actions without separating persistent coordination strategies from their tactical execution. We introduce OverForge, a training-free hierarchical architecture that separates strategic reasoning over roles and divisions of labour from tactical reasoning over actions within each agent's private, partner-conditioned world model. A metacognitive Prefrontal Cortex Module couples the two levels by forming strategy-action branches, imagining their consequences with a forward model, and committing when confident. In OvercookedV2, OverForge delivers 7 soups in a connected kitchen versus 3 for each flat LLM baseline, retains agreed roles, and adopts roles proposed by unfamiliar partners. Ablations and a fixed-strategy probe show that persistent strategies guide tactical adaptation while each reasoning level contributes to coordination. Memory restarts show that cross-episode partner knowledge supports task performance and partner prediction, linking the hierarchy to continual adaptation.
摘要:合作語言模型代理必須在長期內協調並適應不斷變化的環境以及具有不熟悉慣例的夥伴,但現有的代理將觀察映射到行動,而不將持久的協調策略與其戰術執行分開。我們介紹了OverForge,一種無需訓練的分層架構,將角色和勞動分工的戰略推理與每個代理的私有、基於夥伴的世界模型中的行動戰術推理分開。元認知前額葉皮層模塊通過形成策略-行動分支來聯結這兩個層次,利用前向模型想像其後果,並在有信心時作出承諾。在OvercookedV2中,OverForge在連接的廚房中交付7碗湯,而每個平面LLM基準僅交付3碗,保持商定的角色,並採納不熟悉夥伴提出的角色。消融實驗和固定策略探針顯示,持久策略指導戰術適應,而每個推理層次都對協調有所貢獻。記憶重啟顯示,跨劇集的夥伴知識支持任務表現和夥伴預測,將層次結構與持續適應聯繫起來。
ArchitectureIQ: On the Measure of Training Intuition
2609.39714v1 by Zirui Ren, Shaoyang Guo, Chencheng Tang, Jinxin Wang, Chengyu Xiong, Shanbin Yu, Peihang Li, Yidi Wu, Bangzhe Huang, Qingyu Qu, Leqian Yang, Ziming Liu
Top researchers have good intuition, but do language models have as good intuition about model training as top AI researchers? To measure model intuition of LLMs and humans, we introduce the ArchitectureIQ benchmark. Each question presents a synthetic dataset and several training recipes, and the test-taker is asked to predict the recipe yielding the best test metric. Overall, we find that LLMs' model intuition is good but has four limitations: (1) The intuition is imperfect, or even sub-human in some cases. Frontier models achieve around 76% accuracy (random choice 33%) vs best human researcher (66.0%), yet remain far from perfect. For architecture-only questions, best human achieves 65% while GPT-6 Astra only has 38%. (2) The intuition is empirical, not structured, supported by the fact that more CoT compute does not lead to substantial improvement. Unlike math, we still lack a "Science of AI" language that enables structured reasoning on AI. (3) The intuition is not maximally condensed, and can be further compressed into a knoledge base. Our constructed knowledge base with only 20 items yields large gains for weak models: GPT-4o equipped with the accumulated knowledge almost matches the performance of Claude Opus 5. (4) The intuition is insensitive to dataset properties, but the best model should in general depend on data properties. This suggests that data is the real "dark matter" in AI -- LLMs (so do human researchers) understand too little about data, even less than model architectures.
摘要:頂尖研究者擁有良好的直覺,但語言模型對於模型訓練的直覺是否與頂尖AI研究者一樣出色?為了衡量LLM和人類的模型直覺,我們引入了ArchitectureIQ基準。每個問題都呈現一個合成數據集和幾個訓練配方,測試者被要求預測產生最佳測試指標的配方。總體而言,我們發現LLM的模型直覺良好,但有四個限制:(1) 直覺並不完美,甚至在某些情況下低於人類。前沿模型的準確率約為76%(隨機選擇33%),而最佳人類研究者為66.0%,但仍然遠未完美。對於僅限架構的問題,最佳人類達到65%,而GPT-6 Astra僅有38%。(2) 直覺是經驗性的,而不是結構化的,這一點得到了更多CoT計算並未帶來實質性改善的事實支持。與數學不同,我們仍然缺乏一種能夠對AI進行結構化推理的“AI科學”語言。(3) 直覺並未最大程度地濃縮,可以進一步壓縮成知識庫。我們構建的僅有20個項目的知識庫為弱模型帶來了巨大的增益:配備了累積知識的GPT-4o幾乎達到Claude Opus 5的性能。(4) 直覺對數據集特性不敏感,但最佳模型通常應依賴於數據特性。這表明數據是真正的AI“暗物質”——LLM(人類研究者也是)對數據的理解太少,甚至少於對模型架構的理解。
ChronoGraph: Functional 4D Scene Graphs with Vision-Language Models for Interaction Understanding and Grounded Planning
2609.39665v1 by Chenyangguang Zhang, Malgorzata Gwiazda, Guanlong Jiao, Yuanchen Ju, Federico Tombari, Koushil Sreenath, Marc Pollefeys, Sunghwan Hong
Embodied agents must determine where to act, anticipate the resulting scene changes, and interpret observed outcomes to guide subsequent actions. This requires connecting 4D interaction understanding, which explains how past actions changed the scene, with spatially grounded planning, which determines how and where to act toward a goal and anticipates the resulting scene changes. We introduce ChronoGraph, a functional 4D scene graph that links actions on affordance parts to semantic and geometric state changes. By representing observed and anticipated transitions in the same form, it provides a shared basis for understanding and planning. We construct ChronoGraphBench through an automatic data engine that converts human-interaction videos and simulated robot trajectories into graph-annotated questions for training and evaluating Vision-Language Models (VLMs) on both tasks. Using these annotations, we train ChronoGraphVLM by adapting pretrained VLMs in two stages. Graph-as-Chain-of-Thought supervised fine-tuning teaches the models to reconstruct observed transitions and predict future ones as graph traces before answering. Subsequent joint 4D graph reinforcement learning directly rewards graph properties and answer correctness. Experiments across model scales show improvements over the corresponding pretrained baselines and zero-shot transfer to VLM4D. Real-world demonstrations further show that graph-based planning and affordance grounding support mobile manipulation through existing robot skills without additional fine-tuning.
摘要:具身代理必須確定行動的地點,預測隨之而來的場景變化,並解釋觀察到的結果以指導後續行動。這需要將 4D 互動理解(解釋過去的行動如何改變場景)與空間基礎規劃(確定如何以及在哪裡朝著目標行動並預測隨之而來的場景變化)連接起來。我們介紹 ChronoGraph,一個功能性 4D 場景圖,將對可供性部分的行動與語義和幾何狀態變化聯繫起來。通過以相同的形式表示觀察到的和預期的轉變,它為理解和規劃提供了一個共同的基礎。我們通過一個自動數據引擎構建 ChronoGraphBench,該引擎將人類互動視頻和模擬機器人軌跡轉換為帶有圖形標註的問題,以便在兩個任務上訓練和評估視覺-語言模型(VLMs)。利用這些標註,我們通過在兩個階段適應預訓練的 VLMs 來訓練 ChronoGraphVLM。作為思維鏈的圖形監督微調教導模型重建觀察到的轉變並在回答之前預測未來的轉變作為圖形痕跡。隨後的聯合 4D 圖形強化學習直接獎勵圖形屬性和答案的正確性。跨模型規模的實驗顯示出相對於相應的預訓練基線的改進,以及對 VLM4D 的零樣本轉移。現實世界的演示進一步表明,基於圖形的規劃和可供性基礎支持通過現有的機器人技能進行移動操作,而無需額外的微調。
Zero-Compute Cross-Lingual Transferability Estimation Using Typological Feature Proxies
2609.39640v1 by Dalton Raphael Harmsen, Swier Garst, Thomas van Osch, Zarè Palanciyan, Joaquin Vanschoren
Cross-lingual transfer describes how knowledge in a source language benefits a target language. Measuring it quantitatively requires broad multilingual pre-training, as prior work has done with cross-lingual transfer matrices. We ask whether transfer is predictable from freely available typological features, and whether the prominence of high-resource source languages reflects typology or data quality and quantity. We show that typological databases contain cheap and dense signals about cross-lingual transfer. Our typology-only random forest on a 24-language prior-work transfer matrix scores leave-one-language-out $ρ{=}0.705$ and $R^2{=}0.49$, beating a non-typological control at $ρ{=}0.62$, which verifies the ability of typology-only predictions to reconstruct costly measured cross-lingual transfer. The signal survives leave-one-script-out and leave-one-family-out protocols, so script and family confounding do not explain the effect. By decomposing the transfer into a typology term and a resource-and-script bias term, we find the best-source ranking sensitive to this bias. In contrast, typology is not affected by this bias, which makes it a zero-compute screening tool that replaces hundreds of training runs with a model fit. Our code is available \href{https://github.com/dharmsen/typo-x-ling-transfer}{here}.
摘要:跨語言轉移描述了來源語言的知識如何惠及目標語言。
定量測量需要廣泛的多語言預訓練,正如先前的工作所做的跨語言轉移矩陣。
我們詢問是否可以從自由可用的類型特徵預測轉移,以及高資源來源語言的顯著性是否反映了類型學或數據質量和數量。
我們展示了類型學數據庫包含有關跨語言轉移的廉價且密集的信號。
我們的僅基於類型學的隨機森林在24語言的先前工作轉移矩陣上的得分為留一語言外 $ρ{=}0.705$ 和 $R^2{=}0.49$,超過了 $ρ{=}0.62$ 的非類型學控制,這證實了僅基於類型學的預測能夠重建昂貴的測量跨語言轉移的能力。
該信號在留一腳本外和留一語系外的協議中仍然存在,因此腳本和語系的混淆並不能解釋這一效果。
通過將轉移分解為類型學項和資源與腳本偏差項,我們發現最佳來源排名對此偏差敏感。
相比之下,類型學不受此偏差影響,這使其成為一種零計算篩選工具,能夠用模型擬合取代數百次訓練運行。
我們的代碼可在 \href{https://github.com/dharmsen/typo-x-ling-transfer}{這裡} 獲得。
RankEvolve: A Reliable Multi-Agent Auto-Research Harness for Evolving Ranking Models
2609.39551v1 by Zheng Chen, Linfeng Liu, Hong Li, Hong Yan
Auto-research agents, LLM systems that propose, implement, train, and evaluate model changes across iterations, promise to automate applied ML's experimental loop. Over long horizons, execution accuracy is a binding constraint: a change can silently leak held-out data, omit normalization, disconnect a gradient, or leave a train/eval flag unwired, invalidating expensive runs and compounding error across iterations. We present RankEvolve, an auto-research framework for evolving generative ranking models. An Executable Operating Protocol (EOP) declares phases, gates, branches, and loops, and the runtime enforces the compiled state machine. A meta-meta-harness composes complete black-box coding-agent products, including Claude Code and Codex, as execution-graph nodes that review and repair one another's work. In a budget-matched evaluation, heterogeneous composition raises all-oracle execution accuracy from the best single-product baseline of 45.8 percent to 62.5 percent (paired +16.7 points, 95 percent CI [6.6, 26.7]) while achieving a 10.4 percent silent critical-defect rate. An implemented knowledge layer carries findings, including negative results, across iterations. In a twelve-iteration deployment on the open-source HSTU recommender, RankEvolve reported NDCG@10 of 0.2192 on MovieLens-20M LARGE (+4.48 percent over the published anchor) and 0.1948 on BASE (+2.80 percent). ExecML-HSTU, seeded by incidents from that deployment, provides the oracle benchmark for the execution-accuracy evaluation. A pre-specified LitGPT transfer split replicates the heterogeneous-composition effect beyond recommendation (+12.5 points, 95 percent CI [3.0, 22.0]), and a paired ablation isolates per-step from full-protocol instruction injection. These results characterize when runtime-controlled composition of coding-agent products improves execution accuracy.
摘要:自動研究代理,LLM 系統在多次迭代中提議、實施、訓練和評估模型變更,承諾自動化應用機器學習的實驗循環。在長期的執行中,準確性是一個約束條件:一個變更可能會悄悄洩漏保留數據、忽略正規化、斷開梯度或使訓練/評估標誌未連接,從而使昂貴的運行失效並在迭代中累積錯誤。我們提出了 RankEvolve,一個用於演變生成排名模型的自動研究框架。一個可執行操作協議 (EOP) 聲明了階段、閘、分支和循環,並且運行時強制執行編譯的狀態機。一個元元框架組合了完整的黑箱編碼代理產品,包括 Claude Code 和 Codex,作為執行圖節點,彼此審查和修復工作。在一個預算匹配的評估中,異質組合將所有預言者的執行準確性從最佳單一產品基線的 45.8% 提高至 62.5%(配對 +16.7 點,95% 置信區間 [6.6, 26.7]),同時實現了 10.4% 的靜默關鍵缺陷率。一個實施的知識層在迭代中攜帶發現,包括負面結果。在對開源 HSTU 推薦系統的十二次迭代部署中,RankEvolve 在 MovieLens-20M LARGE 上報告的 NDCG@10 為 0.2192(比已發表的基準高出 +4.48%),在 BASE 上為 0.1948(高出 +2.80%)。ExecML-HSTU,基於該部署中的事件提供了執行準確性評估的預言者基準。一個預先指定的 LitGPT 轉移拆分在推薦之外複製了異質組合效應(+12.5 點,95% 置信區間 [3.0, 22.0]),而配對消融則將每步與完整協議指令注入隔離。這些結果表徵了何時運行時控制的編碼代理產品組合改善了執行準確性。
Learning Normal Diffusion Dynamics for Backdoor Defense in Text-to-Image Models
2609.39548v1 by Junjian Li, Xiaolong Liu, Peng Sun, Liantao Wu, Linghan Chen, Yudong Gao, Honglong Chen
Backdoor attacks pose a serious threat to the secure deployment of text-to-image (T2I) diffusion models. Existing defenses typically detect backdoors from specific abnormal patterns in internal representations, which may limit their generalizability with the emergence of increasingly diverse attack mechanisms. In this paper, we study backdoor defense of T2I diffusion models from a transition-dynamics perspective. We observe that benign diffusion trajectories exhibit structured and timestep-dependent transition patterns from cross-attention, latent and noise spaces, whereas backdoor attacks tend to induce deviations from such normal evolution. Motivated by these observations, we propose Normal Diffusion Dynamics Learning (NDDL), a novel backdoor defense framework that learns the normal transition dynamics of diffusion trajectories utilizing only benign samples. NDDL constructs compact multi-space trajectory representations and trains a timestep-conditioned dynamics model to predict the diffusion evolution. In the inference phase, deviations between the observed and predicted transitions are exploited to quantify dynamics inconsistency for backdoor detection. NDDL further enables trigger localization without any prior knowledge of the embedded backdoor by performing substitution with low-semantic words. Extensive experiments for diverse backdoor attacks demonstrate the effectiveness and generalizability of our proposed NDDL.
摘要:後門攻擊對文本到圖像(T2I)擴散模型的安全部署構成了嚴重威脅。現有的防禦通常通過檢測內部表示中的特定異常模式來識別後門,這可能會限制它們在不斷出現的多樣化攻擊機制中的通用性。本文從轉移動力學的角度研究T2I擴散模型的後門防禦。我們觀察到良性擴散軌跡在交叉注意、潛在和噪聲空間中顯示出結構化和時間步依賴的轉移模式,而後門攻擊則傾向於導致這種正常演變的偏差。受到這些觀察的啟發,我們提出了正常擴散動力學學習(NDDL),這是一種新穎的後門防禦框架,僅利用良性樣本學習擴散軌跡的正常轉移動力學。NDDL構建了緊湊的多空間軌跡表示,並訓練了一個時間步條件的動力學模型來預測擴散演變。在推斷階段,觀察到的轉移與預測轉移之間的偏差被用來量化動力學不一致性以進行後門檢測。NDDL進一步通過使用低語義詞進行替換,實現了無需任何嵌入後門的先驗知識的觸發器定位。針對多樣化後門攻擊的大量實驗證明了我們提出的NDDL的有效性和通用性。
A Reusable Semantic Web Framework for Evidence-Grounded Fundamental Rights Impact Assessments under the EU AI Act
2609.39537v1 by Faith Olopade, Delaram Golpayegani, David Lewis
The EU AI Act (Art. 27) requires deployers of high-risk AI systems to conduct Fundamental Rights Impact Assessments (FRIAs) before deployment, yet the evidence needed for credible assessments is fragmented across incompatible incident repositories, risk vocabularies, and legal texts. We present a reusable Semantic Web-based framework that consolidates this evidence for two high-risk public sector categories: employment and worker management (Annex III(4)) and access to essential public services (Annex III(5)(a)). A curated 150-record corpus is annotated along four axes using keyword, LLM, and hybrid methods and serialised as a SPARQL-queryable knowledge graph of 1,351 RDF triples. Five FRIA demonstration scenarios surface 103 records (68.7% coverage). Evaluation against a 69-record gold standard reveals that LLM-assisted classification of the employment domain achieves only $κ= 0.045$, a cautionary result for automated fairness-related evidence retrieval in this domain. All artefacts are released openly to support adoption by regulators, national authorities, and SMEs.
摘要:歐盟人工智慧法案(第27條)要求高風險人工智慧系統的部署者在部署前進行基本權利影響評估(FRIAs),然而,進行可信評估所需的證據在不相容的事件資料庫、風險詞彙和法律文本中是分散的。我們提出了一個可重用的基於語義網的框架,整合了兩個高風險公共部門類別的證據:就業和工人管理(附件III(4))以及獲取基本公共服務(附件III(5)(a))。一個經過策劃的150條記錄語料庫沿著四個軸進行了註釋,使用關鍵字、LLM和混合方法,並序列化為1,351個RDF三元組的SPARQL可查詢知識圖譜。五個FRIA示範場景顯示出103條記錄(68.7%的覆蓋率)。與69條記錄的金標準進行評估顯示,就業領域的LLM輔助分類僅達到$κ= 0.045$,這對於該領域自動化公平相關證據檢索是一個警示結果。所有產物均公開發布,以支持監管機構、國家當局和中小企業的採用。
Beyond the Shadows of Plato's Cave: Evaluating False Memory in Autonomous Agents via Counterfactual Reasoning
2609.39473v1 by Quan M. Tran, Zhuo Huang, Zhen Fang, Jing Zhang, Mingming Gong, Tongliang Liu
Autonomous agents increasingly rely on memory to generalize beyond their training environments. However, agents are bounded by what they have seen and believed, and leveraging such memories in unseen environments can introduce biases into their internal beliefs. We formalize this phenomenon as \textit{false memory}, which can arise from spurious correlations, environment shifts, and knowledge conflicts. Despite its importance, false memory is difficult to evaluate because it stems from agent internal beliefs and is easily confounded with ordinary generalization failures. Therefore, we propose FAME, a training-free framework that evaluates false memory through the evolution of agent beliefs under counterfactual reasoning. Specifically, counterfactual scenarios reveal how beliefs change as the latent concept of memory shifts under hypothetical interventions; thus, measuring the resulting concept drift provides a signal for distinguishing faithful versus false memory. Such concepts can be estimated from agent hidden states before answer generation, avoiding the need for reward design or answer sampling. Empirical experiments reveal that simply monitoring answers often fails to detect false memory, while FAME achieves AUROCs of 76.2% - 96.7% across false-memory settings, and outperforms the best baseline by 3.4% - 23.3% across realistic benchmarks, spanning math reasoning (GSM-Symbolic), code generation (GitChameleon), and complex reasoning (BigBench-Hard). We further release corresponding counterfactual templates and facilitate future research on false memory.
摘要:自主代理越來越依賴記憶來超越其訓練環境。
然而,代理受到他們所見和所信的限制,並且在未見環境中利用這些記憶可能會將偏見引入他們的內部信念。我們將這一現象正式化為 \textit{錯誤記憶},這可能源於虛假的相關性、環境變化和知識衝突。
儘管其重要性,錯誤記憶難以評估,因為它源於代理的內部信念,並且容易與普通的概化失敗混淆。因此,我們提出了 FAME,一個無需訓練的框架,通過在反事實推理下代理信念的演變來評估錯誤記憶。
具體而言,反事實場景揭示了隨著記憶潛在概念在假設干預下的變化,信念如何改變;因此,測量由此產生的概念漂移提供了一個區分真實記憶和錯誤記憶的信號。
這些概念可以從代理的隱藏狀態中估算,在答案生成之前,避免了獎勵設計或答案抽樣的需求。
實證實驗顯示,僅僅監控答案往往無法檢測錯誤記憶,而 FAME 在錯誤記憶設置中達到了 76.2% - 96.7% 的 AUROC,並且在現實基準中比最佳基線高出 3.4% - 23.3%,涵蓋數學推理 (GSM-Symbolic)、代碼生成 (GitChameleon) 和複雜推理 (BigBench-Hard)。
我們還發布了相應的反事實模板,並促進未來對錯誤記憶的研究。
CAST: Causal Advantage-Structured Training with Spatially Grounded Compositional Rewards for Diffusion Models
2609.39441v1 by Shu Yu, Chaochao Lu
Online reinforcement learning has been extended to flow matching for diffusion model (DM) image generation. However, this paradigm faces three limitations: (1) Window selection. Existing methods manually set the stochastic differential equation (SDE) sampling window, i.e., the denoising steps where exploration noise is injected. We instead determine it from each model's denoising trajectory. (2) Reward saturation. Current methods rely on scoring models trained on human annotations; we find that such scores are extremely high and nearly indistinguishable on the latest SOTA open-source DMs, making advantage estimation largely ineffective. (3) Sample inefficiency. A single scalar reward collapses different failure modes into almost identical scores, leaving minimal gradient guidance for targeted improvement. To address these issues, we propose CAST (Causal Advantage-Structured Training), an RL fine-tuning method for pretrained DMs, which (1) identifies the denoising step at which each model fixes the objects and their spatial arrangement in the image and uses that timing to set the SDE window, (2) decomposes each prompt via Causal Scene Graphs (CSG) into verifiable-atoms, i.e., minimal semantic units such as an object, count, attribute, or spatial relation that can each be checked independently, and rewards each atom separately, and (3) projects the signed atom-level advantages into pixel space through teacher-forced attention and uses them to spatially weight the SDE policy objective. We fine-tune two of the strongest open-source DMs, FLUX.2-dev and Qwen-Image-2512, with CAST, and evaluate them on GenEval 2, a compositional benchmark, and on Qwen-Image-Bench for overall quality. Within almost the same training budget, CAST's improvement over the base model on the most challenging GenEval 2 prompts is up to 3.07x that of Flow-GRPO, while overall generation quality also improves.
摘要:在線強化學習已擴展到擴散模型(DM)圖像生成的流匹配。然而,這一範式面臨三個限制:(1)窗口選擇。現有方法手動設置隨機微分方程(SDE)抽樣窗口,即注入探索噪聲的去噪步驟。我們則從每個模型的去噪軌跡中確定它。(2)獎勵飽和。目前的方法依賴於基於人類標註訓練的評分模型;我們發現這些分數極高,並且在最新的SOTA開源DM中幾乎無法區分,使得優勢估計在很大程度上無效。(3)樣本低效。一個單一的標量獎勵將不同的失敗模式壓縮為幾乎相同的分數,為有針對性的改進留下了最小的梯度指導。為了解決這些問題,我們提出了CAST(因果優勢結構訓練),這是一種用於預訓練DM的強化學習微調方法,該方法(1)確定每個模型在圖像中固定物體及其空間排列的去噪步驟,並利用該時機設置SDE窗口,(2)通過因果場景圖(CSG)將每個提示分解為可驗證的原子,即最小語義單元,如物體、數量、屬性或空間關係,每個都可以獨立檢查,並單獨獎勵每個原子,以及(3)通過教師強制注意力將簽名的原子級優勢投影到像素空間,並利用它們對SDE政策目標進行空間加權。我們使用CAST微調了兩個最強的開源DM,FLUX.2-dev和Qwen-Image-2512,並在組合基準GenEval 2和Qwen-Image-Bench上評估它們的整體質量。在幾乎相同的訓練預算下,CAST在最具挑戰性的GenEval 2提示上對基礎模型的改進高達Flow-GRPO的3.07倍,同時整體生成質量也有所提升。
Inferring Causal Relations between Two Sequences of Events with Language Models
2609.39406v1 by Nishchal Prasad, Eric Gaussier, Emilie Devijver, Alexander Obeid Guzman, Armen Aghasaryan, Gregor Gössler
Causal AI is a branch of Artificial Intelligence which helps understand and reason about cause and effect relationships, not just patterns or correlations. Causal discovery aims to infer elements of the underlying causal structure--often represented as a directed graph--from observational and, when available, interventional data. While causal discovery is the fundamental step for moving beyond mere associations toward genuine understanding, and thus the basic building block of causal AI, it becomes intrinsically difficult when causal relations must be inferred from single observations. In such situations, standard causal discovery methods cannot be used and one has to identify causal relations from limited amount of information. This is typically the case for, e.g., sequences of events produced by different alarms which need to be analyzed on the fly to detect abnormal phenomena, which are usually rare. We show in this study that it is possible to leverage the predictive power of Large Language Models (LLMs) to infer causal relations between only two sequences of events. This approach, which is validated on both synthetic and real data, provides better results than standard causal discovery algorithms on several time series data, even though these data were converted into smaller, single observed sequences.
摘要:因果 AI 是人工智慧的一個分支,幫助理解和推理因果關係,而不僅僅是模式或相關性。因果發現旨在從觀察數據以及在可用的情況下的干預數據中推斷潛在因果結構的元素——通常表示為有向圖。雖然因果發現是超越單純關聯走向真正理解的基本步驟,因此也是因果 AI 的基本構建塊,但當因果關係必須從單一觀察中推斷時,這變得本質上困難。在這種情況下,標準的因果發現方法無法使用,必須從有限的信息中識別因果關係。這通常適用於例如由不同警報產生的事件序列,這些序列需要即時分析以檢測異常現象,而這些現象通常是稀有的。我們在這項研究中顯示,利用大型語言模型 (LLMs) 的預測能力來推斷僅有兩個事件序列之間的因果關係是可能的。這種方法在合成數據和真實數據上都得到了驗證,並在幾個時間序列數據上提供了比標準因果發現算法更好的結果,即使這些數據被轉換為較小的單一觀察序列。
Exploring Heterogeneous Model Merging Approach for Complex Knowledge Transfer
2609.39369v1 by Jiahe Fan, Si Chen, Yinghao Hou, Wenbo Xia, Ke Xu, Hong Xie, Enhong Chen
Specialized models encode task-oriented behavior, but transferring that behavior to a general language model usually requires training, distillation, or representation alignment. We study whether such ability can instead be transferred directly at the parameter level. We apply two existing training-free heterogeneous merging methods, previously shown to transfer knowledge between general language models, to specialist-to-general transfer, projecting a specialist donor into the recipient's shape and interpolating backbone parameters without gradient updates or semantic alignment. Intersection-Merge (IM) injects a prefix-aligned donor slice matching the recipient shape, while Activate-Prune-Merge (APM) uses forward-pass activation statistics to select which donor dimensions to retain before injection. Across embedding, reranking, reward modeling, and MoE code-specialist transfer, both methods improve the general recipient, showing that simple heterogeneous merging can move capabilities across diverse specialist roles.
摘要:專門模型編碼任務導向行為,但將該行為轉移到一般語言模型通常需要訓練、蒸餾或表示對齊。
我們研究這種能力是否可以直接在參數層面上轉移。
我們應用兩種現有的無訓練異質合併方法,這些方法之前已顯示能在一般語言模型之間轉移知識,來進行專家到一般的轉移,將專家捐贈者投影到接收者的形狀,並在不進行梯度更新或語義對齊的情況下插值主幹參數。
交集合併(IM)注入與接收者形狀匹配的前綴對齊捐贈者切片,而激活修剪合併(APM)則使用前向傳遞激活統計來選擇在注入之前保留哪些捐贈者維度。
在嵌入、重新排序、獎勵建模和MoE代碼專家轉移中,這兩種方法都改善了一般接收者,顯示簡單的異質合併可以在多樣的專家角色之間轉移能力。
Offline Guidance, Online Reasoning: Reusing LLM Feedback for Small Language Models
2609.39346v1 by Bohan Zhang, Linan Yue, Weibo Gao, Pengyu Chen, Hong Guo, Yanqi Hao
Large language models (LLMs) offer strong reasoning capabilities but are often costly to access through commercial APIs, while small language models (SLMs) are easier to deploy locally yet remain weaker in reasoning. This capability-deployment gap has motivated LLM-SLM collaboration, which aims to improve SLM reasoning using LLM capabilities while preserving the deployment advantages of SLMs. Existing approaches mainly follow two paradigms. Knowledge distillation uses LLM-generated answers and reasoning trajectories to train SLMs offline, but requires parameter updates and additional training. Alternatively, online collaboration routes difficult problems to an LLM or leverages LLM-generated guidance and corrections when an SLM encounters difficulties. Although effective, online collaboration requires repeated LLM access. Moreover, the guidance produced for a particular problem is discarded after inference and cannot benefit subsequent problems involving similar reasoning states. In the paper, we focus on a more constrained setting in which the LLM is accessed only offline, the SLM parameters remain fixed, and online inference is performed solely by the SLM. To this end, we propose Reusable Latent Correction (RLC), which converts one-off natural-language guidance from a black-box LLM into persistent corrective experiences in the hidden space of an SLM. RLC stores these experiences in an external bank and retrieves them according to the SLM's current reasoning state, enabling the SLM to reuse LLM-derived corrections during inference without any online LLM calls. Experiments across multiple reasoning benchmarks and SLM scales show that RLC consistently improves SLM reasoning without parameter updates or online LLM calls. Code is available at https://github.com/ZBH031/reusable-latent-correction.
摘要:大型語言模型(LLMs)提供強大的推理能力,但通過商業API訪問的成本通常很高,而小型語言模型(SLMs)更易於本地部署,但在推理方面仍然較弱。這種能力與部署之間的差距促進了LLM-SLM的合作,旨在利用LLM的能力改善SLM的推理,同時保留SLM的部署優勢。現有的方法主要遵循兩種範式。知識蒸餾使用LLM生成的答案和推理軌跡來離線訓練SLM,但需要參數更新和額外的訓練。或者,線上合作將困難問題路由到LLM,或者在SLM遇到困難時利用LLM生成的指導和修正。雖然有效,但線上合作需要重複訪問LLM。此外,針對特定問題產生的指導在推理後會被丟棄,無法惠及後續涉及類似推理狀態的問題。在本文中,我們專注於一個更受限的設置,其中LLM僅在離線時訪問,SLM參數保持固定,並且線上推理僅由SLM執行。為此,我們提出了可重用潛在修正(RLC),它將來自黑箱LLM的一次性自然語言指導轉換為SLM隱藏空間中的持久修正經驗。RLC將這些經驗存儲在外部庫中,並根據SLM當前的推理狀態檢索它們,使SLM在推理過程中能夠重用LLM衍生的修正,而無需任何線上LLM調用。跨多個推理基準和SLM規模的實驗表明,RLC始終在不進行參數更新或線上LLM調用的情況下改善SLM的推理。代碼可在 https://github.com/ZBH031/reusable-latent-correction 獲得。
Understanding as No-Arbitrage: Bounded Dutch Books as a Definition and Training Objective for Language Models
2609.39341v1 by Daniel Dragonevskiy
Does a language model merely predict tokens, or does it understand what it says? We make this question measurable by defining "understanding" through the lens of no-arbitrage. A model understands a vocabulary to a certain degree if a computationally bounded trader cannot extract guaranteed profit by betting against the model's probabilities on logically related claims (a "Dutch book"). We establish three theoretical results: first, because full logical coherence is computationally intractable, understanding is inherently graded, not absolute. Second, we prove that the exact optimum of standard next-token prediction is inherently incoherent across different question formats; the flaw lies in the training objective, not the architecture. Third, we show that uncertainty accumulates predictably along reasoning chains, making unjustified overconfidence an arbitrage opportunity in itself. To address this, we introduce Arbitr, a training framework where an adversarial trader penalizes the model for logical inconsistencies, paired with a calibration anchor to prevent uninformative collapse. Across five pre-registered experiments on Qwen2.5 and Phi-3.5 models, we demonstrate that standard models are highly exploitable across different phrasings. Arbitr reduces this exploitability by orders of magnitude without sacrificing task accuracy, and the effect successfully transfers to unseen logical patterns and new model families. Crucially, we uncover a scaling illusion: at 7B parameters, near-zero measured incoherence often coincides with extreme, unjustified confidence. We conclude that while Arbitr enforces rigorous logical consistency, coherence is a necessary condition for knowledge, but not a sufficient one
摘要:一個語言模型僅僅是預測標記,還是它理解自己所說的內容?我們通過無套利的視角來定義“理解”,使這個問題可衡量。如果一個計算受限的交易者無法通過對模型在邏輯相關主張上的概率進行對賭來提取保證利潤(即“荷蘭書”),則模型在某種程度上理解詞彙。我們建立了三個理論結果:首先,由於完全的邏輯一致性在計算上是不可處理的,因此理解本質上是分級的,而不是絕對的。其次,我們證明標準的下個標記預測的精確最佳值在不同問題格式之間本質上是不一致的;這一缺陷在於訓練目標,而不是架構。第三,我們顯示不確定性在推理鏈中可預測地累積,使得不合理的過度自信本身成為一種套利機會。為了解決這個問題,我們引入了Arbitr,一個訓練框架,其中對手交易者因邏輯不一致而懲罰模型,並配備校準錨點以防止無信息的崩潰。在對Qwen2.5和Phi-3.5模型進行的五個預註冊實驗中,我們證明標準模型在不同措辭中高度可利用。Arbitr在不犧牲任務準確性的情況下,將這種可利用性降低了數個量級,並且這一效果成功轉移到未見的邏輯模式和新模型家族中。關鍵是,我們揭示了一種擴展錯覺:在70億參數下,接近零的測量不一致性常常與極端、不合理的自信相吻合。我們得出結論,雖然Arbitr強制執行嚴格的邏輯一致性,但一致性是知識的必要條件,而不是充分條件。
WorkGenesis: Building the Worlds That Teach Agents to Work
2609.39325v1 by Xinyu Zhu, Fenyi Liu, Yuzhu Cai, Shuo Tang, Rui Ye, Linfeng Zhang, Siheng Chen
The ability of Large Language Model (LLM) agents to complete daily and professional work is receiving increasing attention. Training such agents requires realistic work scenarios. Expert-authored occupational work is costly and slow to produce, while unconstrained synthesis often yields tasks with weak factual grounding or internally inconsistent requirements. To bridge this gap, we introduce WorkGenesis, a framework that constructs executable occupational work from real-world artifacts through two core technical innovations: (1) Evidence-Based Work Construction, which grounds each unit of work in real-world evidence by retrieving public files guided by O*NET occupational knowledge and synthesizing the surrounding context, companion materials, work request, and itemwise rubric around them; and (2) Execution-Guided Consistency Verification, which renders a reference deliverable inside the constructed work, attributes every unsatisfied rubric item to the agent, the task, or the rubric, and uses task and rubric defects as feedback to iteratively repair the work until it passes the audit. Experimental results demonstrate that Fx-Work-35B, trained with simple supervised fine-tuning (SFT) on only 20K units of work synthesized by WorkGenesis, achieves the highest scores among all comparable-scale baselines on the five reported metrics across GDPvalAA-v2, APEX-Agents-AA, and JobBench (31.00 versus 24.79 average score), and even surpasses frontier models such as the 1.6T DeepSeek-V4-Pro-Preview. These results show that WorkGenesis provides scalable training data for working agents.
摘要:大型語言模型(LLM)代理完成日常和專業工作的能力正受到越來越多的關注。
訓練這樣的代理需要現實的工作場景。
專家撰寫的職業工作成本高且生產緩慢,而不受限制的合成往往會產生事實基礎薄弱或內部不一致的要求的任務。
為了填補這一空白,我們介紹了WorkGenesis,一個通過兩個核心技術創新來構建可執行職業工作的框架:
(1) 基於證據的工作建構,通過根據O*NET職業知識檢索公共文件並合成周圍的上下文、伴隨材料、工作請求和逐項標準,將每個工作單元基於現實世界的證據進行基礎;
(2) 執行引導的一致性驗證,這在構建的工作內部呈現參考交付物,將每個未滿足的標準項歸因於代理、任務或標準,並使用任務和標準缺陷作為反饋,迭代修復工作直到通過審核。
實驗結果表明,Fx-Work-35B在僅用WorkGenesis合成的20K工作單元上進行簡單的監督微調(SFT)訓練,達到了在GDPvalAA-v2、APEX-Agents-AA和JobBench上報告的五個指標中所有可比規模基準中最高的分數(31.00對24.79的平均分),甚至超越了1.6T DeepSeek-V4-Pro-Preview等前沿模型。
這些結果顯示,WorkGenesis為工作代理提供了可擴展的訓練數據。
Faithful Dual-constrained Erasure for Robust LLM Safety Alignment
2609.39279v1 by Jiaqing Li, Shide Zhou, Zhibo Zhang, Yuxi Li, Tianlong Yu, Kailong Wang
Machine unlearning has emerged as a crucial mechanism for removing hazardous knowledge and enforcing safety alignment in Large Language Models (LLMs). However, recent studies reveal a persistent security risk: unlearned models remain highly vulnerable to retraining attacks, where suppressed malicious behaviors rapidly resurface after benign fine-tuning. In this work, we investigate the optimization dynamics of unlearning and identify that this vulnerability stems from shallow alignment. Rather than effectively erasing target knowledge, models often exploit a shortcut by activating previously dormant parameters to act as spurious suppressors, forming a fragile inhibitory shell over intact malicious representations. To address this issue and enforce authentic memory deletion, we propose FDCU, a novel dual-constrained subspace projection framework. FDCU restricts parameter updates through a highly scalable, element-wise dual-masking rule: it preserves general knowledge manifolds via Fisher Information and strictly prohibits the abnormal activation of spurious suppressors via the Principle of Minimal Functional Intervention (PMFI). By reliably blocking the model's ability to superficially hide knowledge, FDCU promotes the authentic dismantling of target representations. Extensive experiments across specific knowledge erasure and safe output control tasks demonstrate that FDCU achieves state-of-the-art robustness against retraining attacks while maintaining near-lossless general utility, ensuring durable safety for LLMs.
摘要:機器去學習已成為移除危險知識和強化大型語言模型(LLMs)安全對齊的重要機制。
然而,最近的研究揭示了一個持續的安全風險:去學習模型仍然對再訓練攻擊高度脆弱,抑制的惡意行為在良性微調後迅速重新出現。
在這項工作中,我們研究了去學習的優化動態,並確定這一脆弱性源於淺層對齊。
模型往往並未有效地抹去目標知識,而是通過激活先前靜止的參數來利用捷徑,作為虛假抑制器,形成一個脆弱的抑制外殼,覆蓋完整的惡意表徵。
為了解決這一問題並強制實現真實的記憶刪除,我們提出了FDCU,一個新穎的雙約束子空間投影框架。
FDCU通過一個高度可擴展的元素級雙遮罩規則來限制參數更新:它通過Fisher信息保留一般知識流形,並嚴格禁止虛假抑制器的異常激活,這是基於最小功能干預原則(PMFI)。
通過可靠地阻止模型表面上隱藏知識的能力,FDCU促進了目標表徵的真實拆解。
在特定知識刪除和安全輸出控制任務中的廣泛實驗表明,FDCU在抵抗再訓練攻擊方面達到了最先進的穩健性,同時保持近乎無損的一般效用,確保了LLMs的持久安全。
Fyan: A Human--AI Harness with Semantic Auditing for Document-Level Formalization
2609.39228v1 by Wei Zhao, Yangshuo Zou, Chengxiang Ding, Yifan Wu, Xuchuan Wang, Zimu Mao, Lei Zhang, Tao Luo
We present FYAN, a human--AI harness for document-level mathematical formalization. Rather than treating theorems in isolation, FYAN coordinates an end-to-end workflow spanning specification, proof planning, logical review, Lean proof construction, knowledge curation, and validation, with support for independent supervision and human guidance. A central component is evidence-grounded semantic auditing, which assesses whether formal statements faithfully preserve their informal specifications. A language model constructs structured evidence over local correspondences, omissions, scope, and logical relations, while a deterministic validator checks this evidence and produces reproducible judgments. When a substantive but admissible deviation is accepted, FYAN requires an explicit proof-transfer obligation connecting the formal statement back to a source-facing interpretation. With the same model (DeepSeek-V4.1-Flash) in every stage, FYAN proves 86 of 143 FormalTCS theorems under a strict Lean check, against 69 for a general agent harness, and raises the natural-language proof score from 0.501 to 0.851. On ConsistencyCheck, its semantic audit catches more inconsistent statements than a direct LLM judge, both on labels verified against the source (recall 0.777 vs. 0.636) and on the original labels (0.873 vs. 0.820), and localizes each mismatch it reports to a specific hypothesis, conclusion, or scope. FYAN also built ODENumLib, a 9,355-line Lean library for the numerical analysis of ordinary differential equation.
摘要:我們提出了 FYAN,一個用於文件級數學形式化的人類與 AI 結合工具。FYAN 不僅僅是將定理孤立地處理,而是協調了一個涵蓋規範、證明規劃、邏輯審查、Lean 證明構建、知識策展和驗證的端到端工作流程,並支持獨立監督和人類指導。一個核心組件是基於證據的語義審計,它評估正式陳述是否忠實地保留了其非正式規範。一個語言模型在局部對應、遺漏、範圍和邏輯關係上構建結構化證據,而一個確定性驗證器則檢查這些證據並產生可重複的判斷。當接受一個實質但可接受的偏差時,FYAN 要求一個明確的證明轉移義務,將正式陳述連接回源面向的解釋。在每個階段使用相同的模型(DeepSeek-V4.1-Flash),FYAN 在嚴格的 Lean 檢查下證明了 143 個 FormalTCS 定理中的 86 個,而一般代理工具僅證明了 69 個,並將自然語言證明分數從 0.501 提升至 0.851。在 ConsistencyCheck 上,其語義審計捕捉到的矛盾陳述比直接的 LLM 評判者更多,無論是在對照源驗證的標籤上(召回率 0.777 對 0.636)還是在原始標籤上(0.873 對 0.820),並將每個報告的不匹配定位到特定的假設、結論或範圍。FYAN 還構建了 ODENumLib,一個包含 9,355 行代碼的 Lean 庫,用於普通微分方程的數值分析。
DAGent: Evaluate-then-Grow Planning for Deep Research Agents
2609.39154v1 by Hanwen Liu, Yuanfu Sun, Qiaoyu Tan
Deep research tasks require agents to navigate large knowledge spaces, synthesize evidence across many sources, and adapt their plans as findings emerge. Directed acyclic graph (DAG)-based multi-agent systems suit this setting because they support parallel execution and isolate each sub-task within a focused dependency context. Yet existing DAG-based agents instantiate a task-level plan before execution and repair the graph only after failures or missing evidence are observed. This Plan-then-Patch strategy is brittle for deep research: the system commits most strongly when its evidence is weakest, and later revisions waste computation on branches that should not have been planned. We propose DAGent, a DAG-based multi-agent framework with Evaluate-then-Grow incremental planning: an Orchestrator grows the task graph one batch at a time, conditioning each expansion on confidence and uncertainty signals from completed nodes. A hierarchical context layer propagates compact QueryDocs by default while preserving full execution traces for on-demand recall. The recorded DAG topology admits structural RL signals that outcome-only recipes cannot define; DAGRPO, a GRPO adaptation, injects topology-conditioned credit on Executor rollouts and a structural compliance regularization on Orchestrator plans. Across BrowseComp-Plus, GAIA, and xbench-DeepSearch, DAGent surpasses the strongest open-source baseline by 5.3 / 5.8 / 2.0 points at the Qwen3-235B-A22B scale, and the lead replicates across four open-source backbones and extends to GPT-5 at 327K context. At the Qwen3-8B scale, DAGRPO improves over a same-budget outcome-only GRPO baseline by 3.0 average Pass@1 points. A same-architecture comparison shows that evidence-conditioned planning reaches higher accuracy at lower per-task token, tool-call, and step footprints than its Plan-then-Patch counterpart. Code: https://github.com/hanwenliu6825/DAGent
摘要:深度研究任務要求代理在大型知識空間中導航,綜合來自多個來源的證據,並隨著發現的出現調整計劃。基於有向無環圖(DAG)的多代理系統適合這種環境,因為它們支持並行執行並將每個子任務隔離在一個集中的依賴上下文中。然而,現有的基於DAG的代理在執行前會實例化一個任務級計劃,並僅在觀察到失敗或缺失證據後修復圖形。這種先計劃再修補的策略對於深度研究來說是脆弱的:系統在證據最薄弱時最強烈地承諾,而後續的修訂則在不應該計劃的分支上浪費計算。我們提出了DAGent,一個基於DAG的多代理框架,具有評估後增長的增量計劃:一個協調者一次增長一批任務圖,並根據已完成節點的信心和不確定性信號調整每次擴展。層次上下文層默認傳播緊湊的QueryDocs,同時保留完整的執行痕跡以便按需回憶。記錄的DAG拓撲允許結構性強化學習信號,這是僅基於結果的配方無法定義的;DAGRPO,GRPO的適應,將基於拓撲的信用注入到執行者的回滾中,並對協調者的計劃施加結構合規性正則化。在BrowseComp-Plus、GAIA和xbench-DeepSearch中,DAGent在Qwen3-235B-A22B規模上超越了最強的開源基線5.3 / 5.8 / 2.0點,並且這一優勢在四個開源骨幹中得到了重複,並擴展到327K上下文的GPT-5。在Qwen3-8B規模上,DAGRPO在同樣預算的僅基於結果的GRPO基線上提高了3.0的平均Pass@1點數。同一架構的比較顯示,基於證據的計劃在每個任務的標記、工具調用和步驟足跡上達到了更高的準確性。代碼:https://github.com/hanwenliu6825/DAGent
Rep2Skill: Representation-Guided Skill Self-Evolution for LLM Agents
2609.39149v1 by Kaixing Zhang, Changming Li, Yingdong Shi, Zheng Zhang, Kaitao Song, Wenjie Shi, Jingang Wang, Kan Ren
Textual skills enable large language model (LLM) based agents to accumulate reusable procedural knowledge without updating model parameters. Yet existing skill evolution remains largely confined to the text space: an optimizer must diagnose success and failure patterns, and revise skills solely from long execution trajectories and sparse task outcomes. This text-only paradigm leaves the agent's internal representations, which contain rich records of its evolving execution state, outside the skill optimization loop. We ask whether an agent can improve its external textual skills by reflecting on its own internal representations. We introduce Rep2Skill, a representation-guided framework for self-evolution on agent skills. Specifically, upon the collected agent rollouts, Rep2Skill models their internal model representation trajectories to localize turns that deviate from successful execution dynamics, and it further interprets these signals alongside the execution contexts as actionable textual feedback for targeted skill revision. Experiments on two agent environments with two open-source LLMs show that Rep2Skill consistently outperforms text-only approaches in the self-evolution setting, where the same LLM serves as both executor and optimizer without a stronger external model. This establishes a promising direction moving agent self-improvement beyond text-only reflection.
摘要:文本技能使基於大型語言模型(LLM)的代理能夠在不更新模型參數的情況下累積可重用的程序知識。
然而,現有的技能演變在很大程度上仍然局限於文本空間:優化器必須診斷成功和失敗模式,並僅根據長期執行軌跡和稀疏的任務結果來修訂技能。
這種僅限於文本的範式使代理的內部表徵(包含其不斷演變的執行狀態的豐富記錄)置於技能優化循環之外。
我們詢問代理是否可以通過反思自身的內部表徵來改善其外部文本技能。
我們介紹了Rep2Skill,一個基於表徵的自我演變框架,用於代理技能的自我提升。
具體而言,在收集到的代理回合中,Rep2Skill建模其內部模型表徵軌跡,以定位偏離成功執行動態的轉折,並進一步將這些信號與執行上下文一起解釋為可行的文本反饋,以進行有針對性的技能修訂。
在兩個代理環境和兩個開源LLM上的實驗顯示,Rep2Skill在自我演變設置中始終優於僅限文本的方法,其中同一LLM同時擔任執行者和優化器,而沒有更強的外部模型。
這為推動代理自我改進超越僅限文本反思建立了一個有前景的方向。
MASCRDM: Multi-Agent System for Compliance Risk Detection and Mitigation in Training Process of Large Language Models
2609.39107v1 by Yan Zhang, Chuming Wei, Ruien Li, Yaoyao Peng, Wusheng Zhang, Guangwen Yang
Large Language Models (LLMs) have been applied in various fields. However, ensuring compliance and safety of LLMs, such as avoiding discrimination and bias, still remains a challenge. Current efforts mainly focus on detecting and filtering inputs and outputs of the trained models, rather than studying the intrinsic architecture of the models in real-time. To tackle this challenge, we analyze the LLMs training process and discover two critical issues: 1) Most of the existing methods are predominantly static in their approach to detection and filtering, achieving only localized optimizations without systematically enhancing the compliance of LLMs. 2) Another issue with existing approaches is the lack of real-time risk detection and mitigation across the full training process, which leads to limited flexibility. Motivated by these, we propose MASCRDM (Multi-Agent System for Compliance Risk Detection and Mitigation) during the LLM training process. Firstly, we develop a set of compliance rules based on existing Artificial Intelligence (AI) laws and a compliance-specific LLM with the instruction of compliance law experts. Then, we deconstruct LLMs into several components and identify key nodes based on the compliance knowledge graph. During LLMs training, we implement our multiple agents in the whole process, giving compliance risk alerts and suggestions for LLM developers. Experiments on discrimination and bias benchmark demonstrate that our multi-agent system can effectively improve the compliance while maintaining reasonable semantic performance. The results indicate that our method provides an executable path for mitigating compliance risk from within the LLMs systematically.
摘要:大型語言模型(LLMs)已被應用於各個領域。
然而,確保LLMs的合規性和安全性,例如避免歧視和偏見,仍然是一個挑戰。
目前的努力主要集中在檢測和過濾訓練模型的輸入和輸出,而不是實時研究模型的內在架構。
為了解決這個挑戰,我們分析了LLMs的訓練過程,並發現了兩個關鍵問題:1)現有的大多數方法在檢測和過濾的方式上主要是靜態的,僅實現了局部優化,而未系統性地增強LLMs的合規性。
2)現有方法的另一個問題是缺乏在整個訓練過程中的實時風險檢測和緩解,這導致了靈活性有限。
受到這些問題的啟發,我們在LLM訓練過程中提出了MASCRDM(合規風險檢測和緩解的多代理系統)。
首先,我們根據現有的人工智慧(AI)法律制定了一套合規規則,並與合規法律專家的指導下開發了一個合規專用的LLM。
然後,我們將LLMs拆解為幾個組件,並根據合規知識圖譜識別關鍵節點。
在LLMs的訓練過程中,我們在整個過程中實施了多個代理,為LLM開發者提供合規風險警報和建議。
在歧視和偏見基準上的實驗表明,我們的多代理系統可以有效改善合規性,同時保持合理的語義性能。
結果顯示,我們的方法為系統性地從LLMs內部緩解合規風險提供了一條可執行的途徑。
Multi-LLM Collaborative Alignment via Stackelberg Games
2609.39076v1 by Christina Hahn, Shangbin Feng, Dean Light, Swastik Roy, Hila Gonen, Yulia Tsvetkov
A pool of language models can collaborate and improve collectively by learning from one another's responses. These interactions depend on the instructions used during training. Existing methods typically sample instructions uniformly, even though their usefulness may change as the models improve: an instruction on which models' responses once differed in quality may later be answered equally well, while a previously difficult instruction may begin to provide a useful learning signal. We propose Stackelberg Alignment, a game-theory-inspired leader-follower framework that turns instruction selection into an adaptive curriculum. An EXP3 bandit acts as the leader, allocating a fixed sampling budget across instructions and updating its sampling distribution using a reward that combines instruction difficulty and response discriminability. The language models act as followers: they respond to the selected instructions, evaluate one another's responses, and learn from the resulting preference signals through DPO or GRPO. The framework uses Elo-style reputation-weighted peer judgment and reputation-based opponent matching to support reliable and competitive model interactions. Experiments across three heterogeneous model pools and 12 benchmarks spanning scientific discovery, reasoning, code, instruction following, and knowledge show that Stackelberg Alignment achieves the highest macro-average across three diverse model pools, outperforming the strongest training-time baseline by up to 7.4% and the best static inference baseline by 12-25%. Analysis confirms that the adaptive leader concentrates duels on the most informative instructions, and ablations show that both reputation-weighted judgment and reputation-based matching improve the effectiveness of multi-LLM evolution.
摘要:一組語言模型可以通過相互學習彼此的回應來協作並共同改進。這些互動依賴於訓練期間使用的指令。現有的方法通常均勻地抽樣指令,即使隨著模型的改進,其有用性可能會改變:曾經模型回應質量不同的指令,後來可能會得到同樣好的回答,而先前困難的指令可能開始提供有用的學習信號。我們提出了Stackelberg Alignment,一種受博弈論啟發的領導者-跟隨者框架,將指令選擇轉變為自適應課程。一個EXP3賭徒作為領導者,在指令之間分配固定的抽樣預算,並使用結合指令難度和回應可區分性的獎勵來更新其抽樣分佈。語言模型作為跟隨者:它們對所選指令做出回應,評估彼此的回應,並通過DPO或GRPO從結果偏好信號中學習。該框架使用Elo風格的聲譽加權同行評價和基於聲譽的對手匹配來支持可靠且具有競爭性的模型互動。在三個異質模型池和12個基準測試的實驗中,涵蓋科學發現、推理、代碼、指令遵循和知識,顯示Stackelberg Alignment在三個不同的模型池中達到了最高的宏觀平均,超越了最強的訓練時間基線高達7.4%,以及最佳靜態推理基線的12-25%。分析確認,自適應的領導者將決鬥集中在最具信息性的指令上,而消融實驗顯示,聲譽加權評價和基於聲譽的匹配都提高了多LLM演化的有效性。
CORE: Conflict-Oriented Reasoning Elimination for Verifiable Language-Model Search
2609.39069v1 by Siyu Song, Rui Xu, Jia Lin, Kai Liu, Weifang Wang
Test-time reasoning systems often respond to failure by restarting or revising the latest step, even when an earlier decision caused the error. We introduce CORE, a search controller that requests a certified conflict core from a verifier, backjumps to the latest decision in that core, and caches the conflict to avoid repeating it. Under sound verification, finite branching and depth, and exhaustive proposals, the uncapped search is complete and never prunes a valid solution. On 2,000 planted graph-coloring instances with matched proposals and an exact verifier, CORE reduces median verifier calls by 39.8% at 30 variables and 35.0% at 36 variables relative to chronological repair; caching further improves on backjumping alone. Across five reasoning tasks, CORE achieves 75.9% mean success with Qwen2.5-7B-Instruct and 84.2% with Qwen3-8B, compared with 72.5% and 81.8% for Tree of Thoughts. It also uses fewer verifier calls and generated tokens on both backbones. These results show the value of using certified failure explanations to direct language-model search.
摘要:測試時的推理系統經常通過重新啟動或修正最新步驟來應對失敗,即使早期的決策導致了錯誤。我們介紹了CORE,一個搜索控制器,它從驗證者那裡請求一個經過認證的衝突核心,回溯到該核心中的最新決策,並緩存衝突以避免重複。根據健全的驗證、有限的分支和深度以及徹底的提案,無上限的搜索是完整的,並且從不修剪有效解。對於2,000個植入的圖著色實例,使用匹配的提案和精確的驗證者,CORE在30個變數時將中位數驗證者調用減少了39.8%,在36個變數時減少了35.0%,相較於時間修復;緩存進一步改善了僅回跳的效果。在五個推理任務中,CORE在Qwen2.5-7B-Instruct上達到75.9%的平均成功率,在Qwen3-8B上達到84.2%,相比之下,Tree of Thoughts的成功率為72.5%和81.8%。它在兩個基礎架構上也使用了更少的驗證者調用和生成的標記。這些結果顯示了使用經過認證的失敗解釋來指導語言模型搜索的價值。
Structure-aware Reinforcement Learning for Protein Directed Evolution
2609.39048v1 by Zikun Nie, Suyuan Zhao, Yizhen Luo, Siqi Fan, Zaiqing Nie
Protein optimization remains a longstanding goal in life sciences. Existing machine learning-assisted directed evolution (MLDE) methods primarily rely on sequence-only features, overlooking the critical spatial constraints and co-evolutionary interactions encoded in protein structures. However, directly integrating structural information remains challenging due to the scarcity of reliable mutant structures. To address these issues, we propose StructEvo, a novel structure-aware reinforcement learning framework for protein directed evolution. StructEvo employs a delta-structure fusion encoder to approximate mutant structure features via feature differences, enabling dynamic incorporation of spatial knowledge. The vast mutation space is then decomposed into manageable subspaces through a structure-aligned hierarchical action network, while a geometric constraint further stabilizes delta feature learning. Our approach outperforms prior state-of-the-art methods by 9.2% and 16.3% on two challenging optimization benchmarks, and further identifies an experimentally validated epistasis pattern in GFP, highlighting the importance of structural guidance for effective protein directed evolution.
摘要:蛋白質優化仍然是生命科學中的一個長期目標。現有的機器學習輔助定向進化(MLDE)方法主要依賴於僅有序列的特徵,忽略了蛋白質結構中編碼的關鍵空間約束和共進化互動。然而,由於可靠突變體結構的稀缺,直接整合結構信息仍然具有挑戰性。為了解決這些問題,我們提出了StructEvo,一個新穎的結構感知強化學習框架,用於蛋白質定向進化。StructEvo採用一種增量結構融合編碼器,通過特徵差異來近似突變體結構特徵,從而實現空間知識的動態整合。然後,通過結構對齊的分層行動網絡,將廣泛的突變空間分解為可管理的子空間,而幾何約束進一步穩定了增量特徵學習。我們的方法在兩個具有挑戰性的優化基準上比之前的最先進方法提高了9.2%和16.3%,並進一步識別出GFP中的一個實驗驗證的表觀基因互作模式,突顯了結構指導對於有效蛋白質定向進化的重要性。
SimEX: Simulation-Integrated Robotics AutoResearch
2609.38982v1 by Jiaheng Hu, Roberto Martin-Martin, Peter Stone, Rocky Duan, Zhenyu Jiang, Guanya Shi
Coding agents powered by large language models (LLMs) have shown remarkable abilities to autonomously reason about and achieve goals in the digital world. However, bringing this success to the physical world remains challenging. On the one hand, direct generation methods (e.g., Code as Policies) often suffer from the LLMs' insufficient understanding of robots and physical environments. On the other hand, iterative trial-and-error tuning in the physical world (e.g., physical autoresearch) induces significant experimental cost and safety concerns. We introduce SimEX: Simulation-Integrated Robotics AutoResearch, an autoresearch framework that tightly integrates simulated experimentation, enabling coding agents to efficiently acquire physical capabilities for controlling real robots. SimEX operates in two stages. First, the agent conducts open-ended probe-and-optimize iterations in simulation, developing a robot toolbox with robust and generalizable capabilities. Second, the agent adapts the toolbox and the simulator together through only a few physical trials: each trial corrects the simulator, and the corrected simulator is used to diagnose failures and screen candidate repairs. We evaluate SimEX extensively in sim-to-sim settings and on physical robots. On challenging real-world manipulation tasks including towel folding, barcode scanning, and plate manipulation, SimEX enables coding agents to efficiently acquire robot skills without any demonstration and with only 10 minutes of real-robot interaction. These results suggest that simulation can be a critical component in achieving physical intelligence, not only as a source of training data that must closely replicate the real world, but also as a roughly correct laboratory where a coding agent develops the knowledge and procedures needed to act on the robot. More details and robot videos at https://robo-simex.github.io/
摘要:編碼代理由大型語言模型(LLMs)驅動,已顯示出在數位世界中自主推理和實現目標的卓越能力。
然而,將這一成功帶入物理世界仍然面臨挑戰。一方面,直接生成方法(例如,將代碼視為政策)往往受到LLMs對機器人和物理環境理解不足的影響。
另一方面,在物理世界中進行的迭代試錯調整(例如,物理自動研究)會產生顯著的實驗成本和安全問題。
我們介紹SimEX:模擬整合機器人自動研究,這是一個自動研究框架,緊密整合模擬實驗,使編碼代理能夠高效獲取控制真實機器人的物理能力。
SimEX分為兩個階段運作。
首先,代理在模擬中進行開放式的探測和優化迭代,開發出具有穩健且可泛化能力的機器人工具箱。
其次,代理通過僅進行幾次物理試驗來共同調整工具箱和模擬器:每次試驗都會修正模擬器,修正後的模擬器用於診斷故障並篩選候選修復方案。
我們在模擬到模擬的設置和物理機器人上對SimEX進行了廣泛評估。在包括毛巾折疊、條碼掃描和盤子操作等具有挑戰性的現實世界操作任務中,SimEX使編碼代理能夠在沒有任何示範的情況下,僅用10分鐘的真實機器人互動高效獲取機器人技能。
這些結果表明,模擬可以是實現物理智能的關鍵組成部分,不僅作為必須與現實世界緊密重複的訓練數據來源,還作為一個大致正確的實驗室,讓編碼代理發展出在機器人上行動所需的知識和程序。
更多細節和機器人視頻請訪問 https://robo-simex.github.io/
DrivingBench: Can Vision-Language Models Drive a Toyota Corolla?
2609.38948v1 by Aditya Ramabadran, Simon Mahns, Tobias Gessler
Frontier models excel at many digital benchmarks, yet their ability to drive a real car, an everyday human skill, remains largely untested. We present DrivingBench, to our knowledge the first benchmark where general-purpose vision-language models must drive a real car. Through three tools, the models see camera frames from a Toyota Corolla and directly command its steering and velocity around a parking lot cone course at low speeds. The car may continue moving while the model thinks and new commands replace the currently running one, so inference latency is part of the task, testing the models' abilities to observe, act, monitor, recover, and complete a long-horizon objective under such constraints. We benchmark GPT-6 Astra, Claude Fable 5.1, GPT-5.6 Sol, and Grok 4.6 in vendor-native harnesses (Codex, Claude Code, Cursor) with up to three attempts each in one conversation; Astra is the only model to finish the course, on its second attempt, with no other attempt passing 50% of the course. Two of the four models improved materially across attempts with retained context. We also detail the design principles behind our action interface, and show how the tool output format and the framing of the task combined to determine whether models would drive at all or refuse. We release our harness, prompts, course map, and traces with video and telemetry for reproducibility.
摘要:前沿模型在許多數位基準測試中表現出色,但它們駕駛真實汽車的能力,這是一項日常人類技能,仍然在很大程度上未經測試。我們提出了DrivingBench,據我們所知,這是第一個基準測試,要求通用視覺-語言模型駕駛真實汽車。通過三個工具,模型可以看到來自豐田卡羅拉的攝影機畫面,並直接指揮其在停車場圓錐賽道上的轉向和速度,速度較低。當模型思考時,汽車可以繼續移動,新指令會取代當前正在運行的指令,因此推理延遲是任務的一部分,測試模型在此類限制下觀察、行動、監控、恢復和完成長期目標的能力。我們在廠商原生的鞍具(Codex、Claude Code、Cursor)中對GPT-6 Astra、Claude Fable 5.1、GPT-5.6 Sol和Grok 4.6進行基準測試,每個模型在一次對話中最多可嘗試三次;Astra是唯一在第二次嘗試中完成賽道的模型,其他模型的嘗試都未達到50%的賽道通過率。四個模型中有兩個在保留上下文的情況下在嘗試中有實質性改善。我們還詳細說明了我們行動介面的設計原則,並展示了工具輸出格式和任務框架如何結合來決定模型是否會駕駛或拒絕。我們釋放了我們的鞍具、提示、賽道地圖和視頻及遙測的追蹤數據,以便於重現。
Prototype-guided Bilateral Alignment Multimodal Federated Learning
2609.38925v1 by Tianchi Liao Tianchi_Liao, Lele Fu, Sheng Huang, Qing Hu, Hong-Ning Dai, Chuan Chen
Multimodal federated learning (MFL) has emerged as a pivotal paradigm for leveraging distributed data to enhance model performance. However, existing methods predominantly rely on idealized assumptions of model homogeneity and balanced modality distributions, rendering them ill-suited for practical scenarios characterized by heterogeneous client architectures and severe modality imbalance. To address these challenges, we propose a \textbf{M}ultimodal \textbf{Fed}erated learning Prototype-guided Bilateral Alignment (MFedPBA) framework. MFedPBA facilitates robust knowledge synergy through a dual alignment mechanism: (i) at the feature level, it aligns heterogeneous feature spaces via a projection encoder optimized by contrastive learning and the Gromov-Wasserstein distance; (ii) at the decision level, it employs an entropy-weighted aggregation of naturally aligned logit prototypes. This novel design achieves robust MFL by jointly tackling heterogeneous feature spaces and collectively aggregating decisions. Extensive experiments demonstrate that our method significantly outperforms state-of-the-art baselines under conditions of model heterogeneity and modality imbalance.
摘要:多模態聯邦學習(MFL)已成為利用分散數據以提升模型性能的關鍵範式。
然而,現有的方法主要依賴於模型同質性和均衡模態分佈的理想假設,使其不適合於特徵異質客戶架構和嚴重模態不平衡的實際場景。
為了解決這些挑戰,我們提出了一個\textbf{M}ultimodal \textbf{Fed}erated learning Prototype-guided Bilateral Alignment(MFedPBA)框架。
MFedPBA通過雙重對齊機制促進穩健的知識協同:
(i) 在特徵層面,它通過對比學習和Gromov-Wasserstein距離優化的投影編碼器對異質特徵空間進行對齊;
(ii) 在決策層面,它使用自然對齊的logit原型的熵加權聚合。
這一新穎的設計通過共同解決異質特徵空間和集體聚合決策來實現穩健的MFL。
大量實驗表明,我們的方法在模型異質性和模態不平衡的條件下顯著超越了最先進的基準。
GraphForge: Training Working Agents with Graph-Anchored Workspace Synthesis
2609.38923v1 by Qisheng Su, Hanchen Wang, Guanru Zhu, Huicheng Jiang, Qiuyinzhe Zhang, Kou Shi, Zhen Fang, Ziao Zhang, Qingnan Ren, Zehui Chen, Tao Gui, Feng Zhao
Working agents need to read diverse files, coordinate tools, and produce deliverables. Training such agents requires tasks built on many real files with verifiable results, but few pipelines exist to synthesize this kind of data. Existing pipelines either generate files with models, which lack realism and diversity, or build tasks on real files without task-specific verifiers, leaving result quality unchecked. We introduce GraphForge, an evidence-graph based framework that grounds both the task and its verification in real files. Starting from occupation-grounded seeds for controlled diversity, GraphForge assembles a workspace of real files for each seed and builds an evidence graph over their relations. Since the task statement and rubrics are both derived from this graph, task requirements are backed by the workspace files and each criterion is anchored to the files needed to verify it. An initial rollout further tests executability, and a revision agent repairs the task and rubrics against the original files before trajectories are collected. Fine-tuning Qwen3.6-27B on 2,169 GraphForge trajectories brings GDPVal to 1445.7 (+65.7) under OpenHands, and Workspace-Bench-Lite and SpreadsheetBench II to 63.7 (+7.7) and 24.0 (+13.7) under Claude Code. Rejection fine-tuning on the SFT model's own rollouts, with candidates selected by the evidence-anchored rubrics, yields further improvements on all three benchmarks, suggesting that the rubrics provide a useful selection signal. The data and models are available.
摘要:工作代理需要閱讀多樣的文件、協調工具並產出交付物。訓練這些代理需要基於許多真實文件且具有可驗證結果的任務,但現有的管道很少能合成這類數據。現有的管道要麼生成缺乏現實感和多樣性的模型文件,要麼在沒有特定任務驗證者的情況下基於真實文件構建任務,導致結果質量無法檢查。我們介紹了 GraphForge,一個基於證據圖的框架,它將任務及其驗證根植於真實文件中。從以職業為基礎的種子開始以控制多樣性,GraphForge 為每個種子組裝一個真實文件的工作空間,並在它們的關係上構建一個證據圖。由於任務聲明和標準都是從這個圖中衍生的,任務要求得到了工作空間文件的支持,每個標準都與驗證所需的文件相連接。初步推出進一步測試可執行性,並且修訂代理在收集軌跡之前會根據原始文件修復任務和標準。在 2,169 個 GraphForge 軌跡上微調 Qwen3.6-27B,使 GDPVal 在 OpenHands 下達到 1445.7 (+65.7),而 Workspace-Bench-Lite 和 SpreadsheetBench II 在 Claude Code 下達到 63.7 (+7.7) 和 24.0 (+13.7)。對 SFT 模型自身推出的拒絕微調,通過證據錨定的標準選擇候選者,對所有三個基準進一步改善,這表明標準提供了一個有用的選擇信號。數據和模型均可用。
Unlearning Deceptive Behaviors in LLMs with Contrastive Forget Sets
2609.38909v1 by Haoran Tang, Rajiv Khanna
Large language models often know the truth and say otherwise: a model that answers correctly when asked neutrally will affirm a user's mistaken belief, or misstate a fact its system prompt wants hidden, once the context rewards it. Such deception is a behavior conditioned on context, not knowledge, yet machine unlearning, the natural tool for removing a behavior from the weights, is built to forget facts that a deceptive model still needs. We propose to unlearn when a model deceives rather than what it knows, with a contrastive forget unit built from the model's own realized deceptions: the same question under a deception-triggering and a neutral context, admitted only where belief holds and behavior flips. Standard objectives on this unit face a dilemma. Suppression objectives such as NPO leave much of the deception in place. Target-based objectives, which distill the model's neutral behavior into the pressured context, remove it but induce context blindness: a target generated without the context teaches the model to stop reading it, eroding benign system-prompt instructions, secret-keeping and the reasoning a monitor inspects, a failure invisible to deception rates and capability benchmarks. We introduce PACT, which trains toward pressure-aware counterfactual targets (the model's own honest response, with a trace that registers the pressure and resists it) while retaining the benign uses of the triggering context. On two 32B reasoning models, PACT reduces held-out deception from over 50% to under 3% while system-prompt adherence, secret-keeping and the reasoning trace stay at the base model's level. On a tug-of-war score of removal against retention, PACT reaches 0.94 and 0.86, against at most 0.77 and 0.60 for any baseline. Like removed knowledge, removed deception is shallow under relearning, and terms that simulate the attacker hold it only at a cost in context use.
摘要:大型語言模型經常知道真相卻說出相反的話:當中立地詢問時能正確回答的模型,會確認用戶的錯誤信念,或者在上下文獎勵它時,錯誤陳述其系統提示想要隱藏的事實。這種欺騙是基於上下文的行為,而非知識,然而,機器的遺忘,這一自然工具用於從權重中去除行為,卻是為了忘記一個欺騙模型仍然需要的事實。我們提議在模型欺騙時進行遺忘,而不是在它所知道的事情上,通過一個由模型自身實現的欺騙構建的對比遺忘單元:在一個觸發欺騙的上下文和一個中立上下文下的同一問題,僅在信念存在且行為翻轉的地方被承認。這個單元的標準目標面臨著困境。像NPO這樣的抑制目標會讓大部分欺騙保持不變。基於目標的目標,將模型的中立行為提煉到受壓上下文中,雖然去除了欺騙,但卻引發了上下文盲目性:在沒有上下文的情況下生成的目標教會模型停止閱讀它,侵蝕了良性的系統提示指令、保密和監控者檢查的推理,這是一種對欺騙率和能力基準來說是不可見的失敗。我們引入了PACT,它朝著壓力感知的反事實目標進行訓練(模型自身的誠實反應,帶有記錄壓力並抵抗它的痕跡),同時保留觸發上下文的良性用途。在兩個32B推理模型上,PACT將保留的欺騙從超過50%降低到低於3%,而系統提示的遵循、保密和推理痕跡仍保持在基礎模型的水平。在去除與保留的拔河得分中,PACT達到了0.94和0.86,而任何基線最多僅為0.77和0.60。像去除的知識一樣,去除的欺騙在重新學習下是淺薄的,模擬攻擊者的術語僅在上下文使用上付出代價。
K2P: Label-Free Knowledge to Prompt Distillation
2609.38898v1 by Yingchuan Zhang, Haoran Lu, Wenxuan Zhong, Ping Ma
Knowledge distillation can transfer reasoning from stronger teachers to frozen students through reusable prompts, but avoiding weight updates does not eliminate supervision. Without ground-truth answers, teacher solutions are unverified, and agreement with the teacher can reward shared mistakes. We introduce Knowledge-to-Prompt (K2P) for label-free knowledge distillation to prompts. K2P synthesizes reusable instructions from teacher solutions, refines them using paired teacher and student responses, and guides search and selection with answer agreement. It retains candidates that adaptive search may undervalue and selects on reserved questions. Deployment uses only the frozen student and selected prompt. Our theory separates generation and selection gaps and gives conditions under which agreement-guided construction yields accuracy guarantees despite imperfect teacher references. Across reasoning tasks and students, K2P outperforms label-free alternatives overall and remains competitive with supervised prompt optimization. Ablations and archive diagnostics assess the contributions of teacher solutions and refinement, while revealing the limits of agreement-guided selection.
摘要:知識蒸餾可以通過可重用的提示將推理從更強的教師轉移到凍結的學生,但避免權重更新並不消除監督。沒有真實答案的情況下,教師的解決方案是未經驗證的,與教師的一致性可能會獎勵共享錯誤。我們引入了無標籤知識蒸餾到提示的知識轉換(K2P)。K2P 從教師解決方案中合成可重用的指令,通過配對的教師和學生反應對其進行精煉,並用答案一致性指導搜索和選擇。它保留了自適應搜索可能低估的候選者,並在保留的問題上進行選擇。部署僅使用凍結的學生和選定的提示。我們的理論將生成和選擇差距分開,並給出一致性指導的構建在不完美教師參考下仍能產生準確性保證的條件。在推理任務和學生中,K2P 整體上表現優於無標籤的替代方案,並在監督提示優化中保持競爭力。消融和存檔診斷評估教師解決方案和精煉的貢獻,同時揭示一致性指導選擇的限制。
Unmerge: Efficient Machine Unlearning via Task Arithmetic
2609.38895v1 by Haoran Tang, Andrew Tan, Rajiv Khanna
Approximate machine unlearning seeks to remove the influence of a forget set from a trained model without full retraining. Existing gradient-based methods require data-dependent hyperparameter search, struggle when forget and retain knowledge are entangled, and offer little insight into where unlearning actually happens inside the network. We recast unlearning through the lens of task arithmetic: if finetuning produces a merged task vector $τ_m$ that combines learning on forget and retain sets, unlearning is the inverse operation that subtracts a learned forget component $τ_F$ to recover the retain task vector $τ_R$. The forget signal is concentrated: at every layer, forget activations lie in a subspace spanned by a handful of dominant directions, so we factorize $τ_F$ in a low-rank forget basis, which is faithful up to a small tail-eigenvalue residual and limits how far the correction can perturb retain. We then optimize three intuitive goals (match the merged vector inside the forget span, suppress leakage into the retain span, and bound the correction size) that provably bound forget leakage and retain damage in activation space. The resulting algorithm, Unmerge, is fast and powerful: on class-level unlearning with ResNet-50 on CIFAR-100 and Tiny ImageNet, it improves Tug-of-War by up to ~24% over a baseline of comparable runtime and by up to ~18% over stronger baselines that run ~5x slower, keeps membership-inference exposure at the level of retraining, and shrinks the feature-distribution gap to the retrained model, where relabeling methods leave forget features cleanly separable. Further studies show that Unmerge also applies to ViT-S/16 and scales to Llama-3.2-3B. The per-layer basis geometry that drives the algorithm also serves as a layerwise diagnostic for when and where unlearning becomes structurally hard.
摘要:近似機器遺忘旨在從訓練模型中去除忘記集的影響,而無需完全重新訓練。現有的基於梯度的方法需要依賴數據的超參數搜索,當忘記和保留知識交織在一起時會遇到困難,並且對於遺忘實際發生的位置提供的見解有限。我們通過任務算術的視角重新詮釋遺忘:如果微調產生了一個合併的任務向量 $τ_m$,該向量結合了對忘記和保留集的學習,則遺忘是減去學習到的忘記組件 $τ_F$ 的逆操作,以恢復保留任務向量 $τ_R$。忘記信號是集中在一起的:在每一層,忘記激活位於由少數主導方向所生成的子空間中,因此我們在低秩的忘記基底中對 $τ_F$ 進行因式分解,這對小尾特徵值殘差是忠實的,並限制了修正可以擾動保留的程度。我們然後優化三個直觀的目標(在忘記範圍內匹配合併向量,抑制對保留範圍的洩漏,以及限制修正大小),這些目標可以明確界定忘記洩漏和激活空間中的保留損害。最終的算法 Unmerge 既快速又強大:在 CIFAR-100 和 Tiny ImageNet 上使用 ResNet-50 進行類別級別的遺忘時,它在可比運行時間的基線上提高了 Tug-of-War 約 24%,在運行速度約慢 5 倍的更強基線上提高了約 18%,並將成員推斷暴露保持在重新訓練的水平,同時縮小了特徵分佈與重新訓練模型之間的差距,這使得重新標註方法能夠將忘記特徵清晰地分開。進一步的研究顯示,Unmerge 也適用於 ViT-S/16 並擴展到 Llama-3.2-3B。驅動算法的每層基底幾何形狀也作為一種層級診斷,幫助判斷何時以及在哪裡遺忘變得結構上困難。
Right Answers, Costly Models: The Efficiency Gap in LLM-based Optimization Modeling
2609.38884v1 by Zhong Li, Xin Huang, Jinhui Wan, Xiangyi Wang, Shenkai Zhang, Ruiqi Chen, Wenyu Liu, Zaiwen Wen, Ziyan Luo
Optimization modeling formulates real-world decision problems as mathematical programs that solvers can use to find optimal decisions. Large language models (LLMs) can automate this process, but the resulting correct formulations can require substantial time and memory to construct and solve, limiting practical scalability. Therefore, we systematically investigate whether LLMs can identify problem structure from natural-language descriptions and apply suitable optimization modeling techniques to generate mathematical models and solver code that solve the problems correctly and efficiently. To this end, we first curate OptTips, a knowledge base of 50 expert modeling techniques in eight families. Using this knowledge, we develop OptDachshund, a multi-agent framework that transforms problems from existing optimization benchmarks into new tasks for evaluating LLMs' use of modeling techniques. It constructs conventional and expert mathematical models with solver code for the same task and data, providing baselines for correctness and computational cost. The resulting EfficientOpt benchmark contains 561 expert-reviewed tasks with paired reference implementations. Evaluation of 11 representative LLMs reveals an efficiency gap on correctly solved tasks with comparable measurements: for every LLM, most generated programs take longer to solve than their expert counterparts. Within the comparable reference-size subset, 57\% of programs with correct objective values and fewer variables and linear constraints have longer recorded solver times. Case studies show that different modeling techniques can achieve the same optimal value at similar recorded cost. Faster solving may not reduce execution time if the code takes longer to prepare data and build the model. LLM optimization modeling should therefore be evaluated for both correctness and computational efficiency.
摘要:優化建模將現實世界的決策問題形式化為數學程序,解決者可以利用這些程序來尋找最佳決策。大型語言模型(LLMs)可以自動化這一過程,但生成的正確公式可能需要大量的時間和內存來構建和解決,從而限制了實際的可擴展性。因此,我們系統地研究LLMs是否能從自然語言描述中識別問題結構,並應用合適的優化建模技術來生成數學模型和解決器代碼,以正確且高效地解決問題。為此,我們首先整理了OptTips,這是一個包含八個類別中50種專家建模技術的知識庫。利用這些知識,我們開發了OptDachshund,一個多代理框架,將現有優化基準中的問題轉化為評估LLMs使用建模技術的新任務。它為相同的任務和數據構建常規和專家數學模型及解決器代碼,提供正確性和計算成本的基準。最終生成的EfficientOpt基準包含561個經專家審核的任務,並附有配對的參考實現。對11個具有代表性的LLMs的評估顯示,在正確解決的任務中存在效率差距,測量結果相似:對於每個LLM,大多數生成的程序解決所需時間比其專家對應物更長。在可比的參考大小子集中,57\%的程序具有正確的目標值且變量和線性約束較少,但記錄的解決時間更長。案例研究顯示,不同的建模技術可以在類似的記錄成本下達到相同的最佳值。如果代碼準備數據和構建模型所需的時間更長,則更快的解決可能不會減少執行時間。因此,LLM優化建模應該同時評估正確性和計算效率。
Does Learning Protein Folding Generalize to Broader Reasoning?
2609.38879v1 by Yong Liu, Zhanpeng Shi, Yizhou Dang, Zhongyue Zhang, Xiaoliang Shi, Zhijian Wei, Shuangjia Zheng
Large language models rely heavily on human text, which often conveys surface answers rather than the spatial and structural logic behind them. Protein folding is a natural testbed, because one solved structure yields thousands of exactly checkable spatial and topological statements. We ask: can learning to fold proteins teach general models reusable reasoning capabilities? To answer this, we build FoldingCorpus, a protein-derived question-answer dataset, and Fold2Reason, a recipe that post-trains on it through two complementary signals: discrete structural answers predicted via the model's native language head, and continuous 3D geometry decoded from the same shared representations. On FoldBench, Fold2Reason achieves structure prediction scores 2.7 to 3.5 times those of Qwen3.5-9B. Beyond protein structure prediction, it improves performance on all 10 benchmarks spanning spatial, graph, scientific, and general reasoning, raising macro-average accuracy from 45.09% to 48.33% (+3.23 pp), with positive gains on all 10 benchmarks, while matched controls built from random, synthetic, and shuffled structure yield substantially smaller or negative gains. Our work shows that non-linguistic, structure-dense scientific data can systematically improve broad reasoning in language models, making a solved scientific problem a practical source of post-training supervision.
摘要:大型語言模型在很大程度上依賴於人類文本,而這些文本往往傳達的是表面的答案,而不是其背後的空間和結構邏輯。蛋白質摺疊是一個自然的測試平台,因為一個已解決的結構可以產生數千個精確可檢查的空間和拓撲陳述。我們問:學習摺疊蛋白質能否教會通用模型可重用的推理能力?為了回答這個問題,我們構建了FoldingCorpus,一個基於蛋白質的問答數據集,以及Fold2Reason,一個通過兩種互補信號進行後訓練的配方:通過模型的原生語言頭預測的離散結構答案,以及從相同共享表示解碼的連續3D幾何。在FoldBench上,Fold2Reason的結構預測分數達到Qwen3.5-9B的2.7到3.5倍。除了蛋白質結構預測外,它還提高了在所有10個基準測試中的表現,這些基準涵蓋了空間、圖形、科學和一般推理,將宏觀平均準確率從45.09%提高到48.33% (+3.23 pp),在所有10個基準上均有正增益,而從隨機、合成和打亂結構建立的匹配對照則產生了顯著較小或負的增益。我們的工作表明,非語言的、結構密集的科學數據可以系統性地改善語言模型中的廣泛推理,使得已解決的科學問題成為實用的後訓練監督來源。
StateTree: Enhancing Long-Term Dialogue Reasoning via Reinforcement Learning
2609.38809v1 by Naen Xu, Wanqing Cui, Yibo Hu, Shixin Hong, Hengyu An, Meiguang Jin, Junfeng Ma, Tianyu Du
Large language models deployed as personalized assistants must reason over long, evolving interaction histories. However, in long-term dialogue reasoning, relevant evidence is scattered across sessions, preferences may be revised over time, and standard long-context training fails to address these challenges under data scarcity and prohibitive computational costs. We propose StateTree, a data-driven RL method that constructs a challenging auxiliary task from scarce dialogues with verifiable ground truth. StateTree augments multi-session dialogues with a tree-structured path-tracing task: key-value records are embedded across sessions to form a binary tree. Solving the task requires the model to traverse from root to leaf by retrieving records across sessions and comparing timestamps to resolve branches, then recover the hidden target question among distractor leaves. We apply curriculum RL training progressively increasing tree depth and introduce a compositional variant whose edges carry step-level reasoning fragments, training the model to compose partial cues into coherent queries. Trained on 10K-token contexts, StateTree generalizes to 128K tokens without full-length RL costs and exhibits capabilities including cross-session retrieval, temporal reasoning, knowledge update, and compositional multi-hop reasoning. StateTree outperforms both SFT and RL-based baselines while preserving short-context general reasoning. StateTree-7B achieves gains up to +23.60% on LongMemEval (128k), and StateTree-14B reaches 59.00% accuracy on LongMemEval, surpassing QwenLong-L1-32B (45.20%).
摘要:大型語言模型作為個人助理部署時,必須對長期且不斷演變的互動歷史進行推理。
然而,在長期對話推理中,相關證據分散在不同的會話中,偏好可能隨時間而修訂,而標準的長上下文訓練未能在數據稀缺和高昂的計算成本下解決這些挑戰。
我們提出了StateTree,一種數據驅動的強化學習方法,從稀缺的對話中構建出具有可驗證真實性的挑戰性輔助任務。
StateTree通過一個樹狀結構的路徑追蹤任務來增強多會話對話:關鍵值記錄在會話之間嵌入,形成一個二叉樹。
解決這個任務需要模型從根部遍歷到葉子,通過檢索會話中的記錄並比較時間戳來解析分支,然後在干擾葉子中恢復隱藏的目標問題。
我們應用課程強化學習訓練,逐步增加樹的深度,並引入一種組合變體,其邊緣攜帶步驟級推理片段,訓練模型將部分提示組合成連貫的查詢。
在10K標記上下文上訓練後,StateTree在不需全長強化學習成本的情況下,能夠推廣到128K標記,並展現出跨會話檢索、時間推理、知識更新和組合多跳推理等能力。
StateTree在保持短上下文一般推理的同時,超越了SFT和基於強化學習的基準。
StateTree-7B在LongMemEval(128k)上獲得了最高+23.60%的增益,而StateTree-14B在LongMemEval上達到了59.00%的準確率,超越了QwenLong-L1-32B(45.20%)。
GraphCert: Bootstrap Agentic Graph Reasoning with Certified Evidence Rubrics
2609.38798v1 by Weiqi Jiang, Yuchen Ying, Rui Wang, Kaixuan Chen, Bingde Hu, Shunyu Liu, Yu Wang, Tongya Zheng
Graph agents extend large language models (LLMs) with the ability to actively explore and reason over knowledge graphs through multi-step interactions with graph tools. However, training capable graph agents typically requires large collections of question-answer pairs and reasoning trajectories, whose manual construction is costly and difficult to scale. Moreover, employing proprietary LLMs to generate such supervision further risks exposing sensitive graph data to external services. Therefore, we propose GraphCert to bootstrap agentic graph reasoning with certified evidence rubrics during post-training. Specifically, the Bootstrapped Graph Quizzer guided by generation controls produces graph-grounded QA pairs and marks supporting evidence, which undergo execution certification and semantic curation. The accepted evidence is then canonicalized into certified evidence rubrics that later reward Graph Solver evidence alignment alongside answer correctness during GRPO training. Experiments on five graph reasoning domains in GRBENCH demonstrate that GraphCert consistently outperforms substantially larger LLM agents and post-training method. Furthermore, our analysis demonstrates that the learned policy transfers robustly across heterogeneous graph domains, suggesting that GraphCert acquires reusable graph-reasoning capabilities rather than domain-specific patterns. These results establish executable self-certification as an effective approach to self-training compact graph reasoning agents. Our code will be made publicly available.
摘要:圖形代理擴展了大型語言模型(LLMs),使其能夠通過與圖形工具的多步互動主動探索和推理知識圖譜。
然而,訓練能夠的圖形代理通常需要大量的問答對和推理軌跡,其手動構建成本高且難以擴展。
此外,使用專有的LLMs來生成這種監督進一步增加了將敏感圖形數據暴露給外部服務的風險。
因此,我們提出了GraphCert,以在後訓練期間用經過認證的證據標準啟動代理圖形推理。
具體而言,受生成控制指導的Bootstrapped Graph Quizzer生成基於圖形的問答對並標記支持證據,這些證據經過執行認證和語義策劃。
被接受的證據隨後被標準化為經過認證的證據標準,這些標準在GRPO訓練期間獎勵圖形解決者的證據對齊以及答案的正確性。
在GRBENCH的五個圖形推理領域的實驗表明,GraphCert始終超越了大得多的LLM代理和後訓練方法。
此外,我們的分析表明,學習到的策略在異質圖形領域中穩健地轉移,這表明GraphCert獲得了可重用的圖形推理能力,而不是特定於領域的模式。
這些結果確立了可執行的自我認證作為自我訓練緊湊圖形推理代理的有效方法。
我們的代碼將公開提供。
Evaluating Persistent Calibration under Evolving Model Knowledge
2609.38797v1 by Victor Wang, Thomas Hofweber, Mohit Bansal, Elias Stengel-Eskin
As AI systems move from static repositories to agents that are capable of continual adaptation and learning, maintaining their trustworthiness means equipping the models backing them with the ability to produce confidence estimates that dynamically reflect their changing skills and knowledge. We introduce the problem of persistent calibration, which requires a confidence estimator to faithfully reflect the knowledge contained in a model as that knowledge changes, without recurring supervision. We operationalize this by examining persistent calibration across checkpoints of open models, asking whether confidence estimators trained on earlier checkpoints can generalize to later ones. Specifically, we aim to shed light on whether confidence is dependent on knowledge, a question with implications for the reliability of confidence estimates. To measure this relationship, we define and evaluate calibration on knowledge contrast sets: subsets containing questions that one checkpoint answers correctly and another checkpoint answers incorrectly, reflecting a change in knowledge. We show that both inference-time and fine-tuning methods fall short on contrast-set calibration compared to oracle methods trained on future checkpoints, even for methods that are well-calibrated on the full dataset. We provide evidence for the hypothesis that persistent calibration is challenging because there is a vast space of possible confidence functions that are well-calibrated on a given checkpoint, out of which only some rely on meta-knowledge features that would generalize to other checkpoints. Towards improving contrast-set calibration, we show that multi-checkpoint training helps, suggesting an avenue for identifying confidence features that remain robust across changing knowledge.
摘要:隨著人工智慧系統從靜態資料庫轉變為能夠持續適應和學習的代理,維持其可信度意味著需要為其背後的模型提供能夠產生動態反映其變化技能和知識的信心估計能力。我們引入了持續校準的問題,這要求信心估計器能夠忠實地反映模型中所包含的知識,隨著知識的變化而變化,而無需重複監督。我們通過檢查開放模型的檢查點來操作化這一點,詢問在早期檢查點上訓練的信心估計器是否能夠對後期檢查點進行泛化。具體而言,我們旨在闡明信心是否依賴於知識,這是一個對信心估計的可靠性有影響的問題。為了測量這種關係,我們在知識對比集上定義並評估校準:這些子集包含一個檢查點正確回答而另一個檢查點錯誤回答的問題,反映知識的變化。我們顯示,無論是推斷時的還是微調的方法,在對比集校準方面都不如在未來檢查點上訓練的神諭方法,即使對於在完整數據集上校準良好的方法也是如此。我們提供了證據支持持續校準是具有挑戰性的假設,因為在給定檢查點上有大量可能的信心函數是良好校準的,而其中只有一些依賴於能夠對其他檢查點進行泛化的元知識特徵。為了改善對比集校準,我們顯示多檢查點訓練是有幫助的,這暗示了一條識別在變化知識中保持穩健的信心特徵的途徑。
Self-Evolving Algorithm-Design Agents: Escaping In-Context Evolutionary Stagnation via Population-Curated Policy Optimization
2609.38757v1 by Chen Lu, Ke Xue, Siyuan Xu, Mingxuan Yuan, Chao Qian
Large language models are increasingly participating in complex real-world tasks in the form of algorithm-design agents, designing and refining algorithms. Many successful algorithm-design agents adopt pure in-context evolutionary frameworks, but they may quickly plateau in domains that require specialized knowledge. Parametric adaptation offers a way to internalize specialized knowledge, but conventional training requires abundant domain-specific corpora while high-quality algorithms are scarce in complex algorithm-design scenarios. In this paper, we propose sample-efficient parametric self-evolution where agents can explore and learn from self-generated algorithms. First, we characterize in-context evolutionary stagnation and analytically propose the Improvement Chain proposition, showing how learning successive self-generated algorithms can locally increase the likelihood of neighboring algorithms. Motivated by this local-transfer perspective, we further propose Population-Curated Policy Optimization (PCPO) to utilize a global population and a hybrid policy update scheme for retaining and reusing high-quality, diverse self-generated algorithms, shifting the policy towards stronger algorithms. In the task of learning rate schedule design for global placement in electronic design automation, trained only on 4 chip cases, PCPO outperforms the state-of-the-art in-context evolutionary methods (e.g., OpenEvolve and ShinkaEvolve) on average across 16 chip cases. With an 8B-size base model, PCPO achieves competitive performance compared to frontier closed-source models such as GPT-5.5. PCPO also reduces inference-time token cost by internalizing grounded domain knowledge and prompt distillation. Moreover, PCPO achieves significant speedups on four GPU kernel designs, with an average of 8.27$\times$ speedup against the PyTorch Eager baseline.
摘要:大型語言模型越來越多地參與複雜的現實世界任務,作為算法設計代理,設計和完善算法。許多成功的算法設計代理採用純粹的上下文進化框架,但在需要專業知識的領域中,它們可能很快達到瓶頸。參數適應提供了一種內化專業知識的方法,但傳統訓練需要大量特定領域的語料庫,而在複雜的算法設計場景中,高品質的算法卻稀缺。在本文中,我們提出了樣本高效的參數自我進化,讓代理可以探索並從自生成的算法中學習。首先,我們描述了上下文進化停滯的特徵,並分析性地提出了改進鏈命題,顯示學習連續自生成算法如何在局部上增加鄰近算法的可能性。受到這一局部轉移視角的啟發,我們進一步提出了人口策劃政策優化(PCPO),利用全球人口和混合政策更新方案來保留和重用高品質、多樣化的自生成算法,將政策轉向更強的算法。在電子設計自動化中進行全球佈局的學習率調度設計任務中,僅在4個晶片案例上訓練的PCPO在16個晶片案例中平均超越了最先進的上下文進化方法(例如,OpenEvolve和ShinkaEvolve)。使用8B大小的基礎模型,PCPO在性能上與前沿的封閉源模型(如GPT-5.5)相比具有競爭力。PCPO還通過內化基礎領域知識和提示蒸餾來降低推理時間的標記成本。此外,PCPO在四個GPU內核設計上實現了顯著的加速,與PyTorch Eager基準相比,平均加速達到8.27$\times$。
Learning to Route in Visual Space via Multi-Step Embedding Retrieval
2609.38743v1 by Tianyu Chen, Mingyuan Zhou, Jiaxing Wu
LLM agents rely on retrieval tools to access external knowledge, yet visual agentic search remains severely bottlenecked by standard single-step retrievers. In current pipelines, the agent must issue text queries for every intermediate step, struggling when visual clues are difficult to describe or when the retriever fails to surface necessary intermediate evidence within its top results. We hypothesize that offloading multi-step navigation across the entire embedding space directly to the retrieval tool resolves this performance bottleneck. To study this systematically, we introduce VHOP, a flexible data generation framework and benchmark with five core difficulty levels testing both visual matching and search planning. Using this framework, we develop VHOP-Router, an end-to-end training pipeline---combining supervised fine-tuning, online imitation learning, and reinforcement learning---that transforms a standard embedding model into an autoregressive multi-step retriever. Operating directly in the visual latent space, VHOP-Router retrieves linked image chains in a single tool call without requiring the agent to formulate intermediate text queries. Experiments show VHOP-Router boosts retrieval performance from under 5\% to 76.3\%. In agentic search, it improves task success rates by 52.7\% and reduces the average token length by 61\% from 1886 to 728, whereas upgrading the agent yields only a 3.7\% gain. Compared to a strong baseline where the agent retrieves the top 50 results per step, VHOP-Router maintains superior performance while reducing in-context images by $23\times$ and cutting the cumulative API payload by $35\times$. The models also generalize robustly to unseen difficulty levels and realistic test sets. Ultimately, VHOP and VHOP-Router provide an efficient and effective solution for visual agentic search that leaves native LLM capabilities entirely intact.
摘要:LLM 代理依賴檢索工具來訪問外部知識,但視覺代理搜索仍然受到標準單步檢索器的嚴重瓶頸。在當前的流程中,代理必須為每個中間步驟發出文本查詢,當視覺線索難以描述或檢索器未能在其頂部結果中顯示必要的中間證據時,代理會面臨困難。我們假設將整個嵌入空間的多步導航直接卸載到檢索工具上可以解決這一性能瓶頸。為了系統地研究這一點,我們介紹了 VHOP,一個靈活的數據生成框架和基準,具有五個核心難度級別,測試視覺匹配和搜索規劃。利用這一框架,我們開發了 VHOP-Router,一個端到端的訓練流程——結合了監督微調、在線模仿學習和增強學習——將標準嵌入模型轉變為自回歸多步檢索器。VHOP-Router 直接在視覺潛在空間中操作,能在一次工具調用中檢索鏈接的圖像鏈,而無需代理制定中間文本查詢。實驗顯示,VHOP-Router 將檢索性能從不足 5\% 提升至 76.3\%。在代理搜索中,它將任務成功率提高了 52.7\%,並將平均標記長度從 1886 減少到 728,減少了 61\%,而升級代理僅帶來 3.7\% 的增益。與一個強大的基準相比,該基準中代理每步檢索前 50 個結果,VHOP-Router 在保持優越性能的同時,將上下文中的圖像減少了 $23\times$,並將累積 API 負載減少了 $35\times$。這些模型在未見過的難度級別和現實測試集上也能穩健地泛化。最終,VHOP 和 VHOP-Router 為視覺代理搜索提供了一種高效且有效的解決方案,完全保留了原生 LLM 的能力。
Concept-Grounded Attention: A Controlled Evaluation of Graph-Injected Attention, Temporal Versioning, and Epistemic Status
2609.38684v1 by Sachin Dev Duggal, Pradyumna Swarnalatha Ramanna, Alexandros Vassiliades
Knowledge-intensive language-model systems typically represent external knowledge as text chunks or static graphs, with limited support for concept evolution, point-in-time reasoning, and distinctions between validated and inferred knowledge. We introduce the Concept Lifecycle Model (CLM), which represents concepts as persistent, graph-grounded, temporally versioned entities with explicit provenance and epistemic status, and Concept-Grounded Attention (CGA), which injects concept-graph structure into transformer computation through graph-biased self-attention (Form A) and gated cross-attention over concept nodes (Form B). We evaluate the framework in controlled settings using disabled-mechanism baselines. On 200 MuSiQue and HotpotQA questions with retrieval fixed, concept-graph retrieval recovers explicit multi-hop paths but does not improve evidence recall. Form A appears to steer attention, with 2.76 times more attention on gold than distractor concepts, but the same ratio occurs when Form A is disabled; the learned bias is negligible and no answers change. An identity-preserving Form B improves F1 from 0.188 to 0.221, but control concepts yield 0.213, indicating that most of the gain reflects added capacity. On LongMemEval, explicit temporal representation improves answer accuracy by 13 to 25 points across all tested generators, up to 122B parameters, while simplified CLM version resolution performs similarly to dated serialization because concept identity is not established reliably. On a synthetic source-independence task, protocol-derived epistemic status reduces unsupported assertions from 28% to 0.1% in a fine-tuned small model and from 19-68% to 0-5% in 72-122B models. Overall, the results support making temporal validity and epistemic status explicit, while showing that graph-attention diagnostics are not informative without disabled-mechanism controls.
摘要:知識密集型語言模型系統通常將外部知識表示為文本塊或靜態圖形,對於概念演變、特定時間推理以及驗證知識與推斷知識之間的區別支持有限。我們介紹了概念生命周期模型(CLM),它將概念表示為持久的、基於圖形的、具有時間版本的實體,並具有明確的來源和認識狀態,以及概念基礎注意力(CGA),它通過圖形偏見自注意力(形式A)和對概念節點的門控交叉注意力(形式B)將概念圖結構注入Transformer計算中。我們在受控環境中使用禁用機制基準評估該框架。在200個MuSiQue和HotpotQA問題中,當檢索固定時,概念圖檢索恢復了明確的多跳路徑,但並未改善證據召回。形式A似乎引導注意力,對金標概念的注意力比對干擾概念高出2.76倍,但當禁用形式A時,比例相同;學習的偏見微不足道,且沒有答案改變。保持身份的形式B將F1從0.188提高到0.221,但控制概念的F1為0.213,表明大多數增益反映了增加的容量。在LongMemEval上,明確的時間表示在所有測試的生成器中將答案準確性提高了13到25分,參數最多可達122B,而簡化的CLM版本解析的表現與過時的序列化相似,因為概念身份未能可靠地建立。在一個合成的源獨立任務中,協議衍生的認識狀態將不支持的斷言從28%減少到0.1%(在微調的小模型中),在72-122B的模型中則從19-68%減少到0-5%。總體而言,結果支持將時間有效性和認識狀態明確化,同時顯示圖形注意力診斷在沒有禁用機制控制的情況下並不具信息性。
Bandits with Multiple Optimal Arms: Minimax Regret and Non-Adaptivity
2609.38659v2 by Kaixuan Ji, Qiwei Di, Qingyue Zhao, Heyang Zhao, Quanquan Gu
We study multi-armed bandits (MAB) with multiple optimal arms, motivated by the fact that many practical decision making problems admit multiple correct answers. For $K$-armed bandits with $A$ optimal arms, we first provide a sharper analysis of previous sub-sampling algorithms (De Heide et al., 2021; Zhu and Nowak, 2020), establishing a $\tilde{O}\Big(\frac{K-A}{\sqrt{KA}}\sqrt{T} \Big)$ minimax regret, where $T$ is the total number of interactions and $\tilde O(\cdot)$ drops all constant and logarithmic factors, improving the previous $\tilde{O}(\sqrt{KT/A})$ regret. We then provide a matching lower bound up to logarithmic factors, indicating that our established rate is nearly minimax-optimal. We further show that the knowledge of $A$ up to $\tilde{O}(1)$ factors is necessary to achieve near-optimal regret, as near-optimal algorithms for one number of optimal arms must incur substantially larger regret than optimal regret for a smaller number. Overall, our results provide a comprehensive minimax characterization of $K$-armed bandits with $A$ over the entire range of $1 \leq A \leq K-1$.
摘要:我們研究具有多個最佳臂的多臂賭徒(MAB),這是因為許多實際的決策問題允許多個正確答案。對於具有 $A$ 個最佳臂的 $K$ 臂賭徒,我們首先對之前的子抽樣算法(De Heide et al., 2021; Zhu and Nowak, 2020)提供了更精確的分析,建立了 $\tilde{O}\Big(\frac{K-A}{\sqrt{KA}}\sqrt{T} \Big)$ 的最小最大後悔,其中 $T$ 是總互動次數,$\tilde O(\cdot)$ 刪除了所有常數和對數因子,改善了之前的 $\tilde{O}(\sqrt{KT/A})$ 後悔。然後,我們提供了一個與對數因子相匹配的下界,表明我們建立的速率幾乎是最小最大最佳的。我們進一步顯示,對 $A$ 的知識最多需要 $\tilde{O}(1)$ 的因子,才能實現接近最佳的後悔,因為對於一個最佳臂的數量,接近最佳的算法必須承擔比較小數量的最佳後悔大得多的後悔。總體而言,我們的結果提供了對於 $1 \leq A \leq K-1$ 整個範圍的 $K$ 臂賭徒與 $A$ 的全面最小最大特徵描述。
Prompt2Skill: Unsupervised Skill Optimization From Natural Language Instructions
2609.38593v1 by Bo Ni, Li Li, Ryan A. Rossi, Franck Dernoncourt, Tyler Derr
Skills are external artifacts that Large Language Models (LLMs) consume at inference time to improve their performance on specialized domains by incorporating relevant procedural and domain knowledge. Expert-authored skills are expensive to produce, and the resulting artifacts are not optimized for the specific model that consumes them, whose failure modes can vary with version, scale and training. In addition, emerging tasks may fall outside the scope of existing skill libraries, creating a need to develop new skills before curated training data become available. Recent works have explored automated skill optimization through reflection, but they require a curated, in-distribution training set, which users might not always have. To address these limitations, we present Prompt2Skill, a framework that builds skills from natural-language task description alone. From the prompt, the system derives a task specification, discovers or synthesizes datasets, and refines the skill in a closed loop of reflective editing. Across four domains spanning question answering, reading comprehension, spreadsheet manipulation, and mathematical reasoning, Prompt2Skill consistently outperforms the direct prompting baseline, achieving an average improvement of 10.8 across open-source and frontier models.
摘要:技能是大型語言模型(LLMs)在推理時消耗的外部產物,通過整合相關的程序和領域知識來提高其在專業領域的表現。專家撰寫的技能製作成本高昂,且所產生的產物並未針對消耗它們的特定模型進行優化,而這些模型的失效模式可能因版本、規模和訓練而異。此外,新興任務可能超出現有技能庫的範疇,因此在經過策劃的訓練數據可用之前,需要開發新的技能。最近的研究探討了通過反思進行自動化技能優化,但這需要一個策劃的、在分佈內的訓練集,而用戶可能並不總是擁有。為了解決這些限制,我們提出了Prompt2Skill,一個僅從自然語言任務描述構建技能的框架。系統從提示中推導出任務規範,發現或合成數據集,並在反思編輯的閉環中精煉技能。在涵蓋問題回答、閱讀理解、電子表格操作和數學推理的四個領域中,Prompt2Skill始終超越直接提示基線,平均提升達到10.8,適用於開源和前沿模型。