LLM
LLM
| Publish Date | Title | Authors | Homepage | Code |
|---|---|---|---|---|
| 2026-10-01 | One Basis to Animate Them All: Gaussian Blendshape Distillation for Real-Time Avatars | Ramazan Fazylov et.al. | 2610.02207v1 | null |
| 2026-10-01 | KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards | Pengfei Li et.al. | 2610.02206v1 | null |
| 2026-10-01 | Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents | Yen-Jen Wang et.al. | 2610.02204v1 | null |
| 2026-10-01 | ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research | Sohyeon Kim et.al. | 2610.02202v1 | null |
| 2026-10-01 | VISTA: A Visual Harness for Reasoning in an Interactive World | Qiushi Han et.al. | 2610.02200v1 | null |
| 2026-10-01 | Hierarchical Continuous Diffusion Language Models | Hui Ren et.al. | 2610.02193v1 | null |
| 2026-10-01 | DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation | Zhengming Yu et.al. | 2610.02188v1 | null |
| 2026-10-01 | Higher-Order Molecular Grammars for Generative and Foundation Models in Chemistry | Yiming Huang et.al. | 2610.02186v1 | null |
| 2026-10-01 | SoftServe: A Scalable Quasi-Newton Method for Deep Learning | Joohwan Ko et.al. | 2610.02182v1 | null |
| 2026-10-01 | Generative Cinematographer: Composing Camera and Object Motion in 3D | Jiahan Zhang et.al. | 2610.02180v1 | null |
| 2026-10-01 | Every Ablation Is a Dose: Counterweights and the Semblance of Self-Repair | Areeb Ahmad et.al. | 2610.02173v1 | null |
| 2026-10-01 | AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents | Xuan Zhang et.al. | 2610.02163v1 | null |
| 2026-10-01 | DuoMind: Enabling Distributed Multi-Robot Coordination with Semantic Communication | Hanchu Zhou et.al. | 2610.02161v1 | null |
| 2026-10-01 | From Knowledge Access to Source Learning: Developing Source-Specific Competence | Lucheng Fu et.al. | 2610.02150v1 | null |
| 2026-10-01 | Keyword Harnesses Fail Open: A Cheap Diagnostic Ladder for Tool-Use Claims in Small Language Models | Juan S. Santillana et.al. | 2610.02142v1 | null |
| 2026-10-01 | Finetuning with Sampling: SFT Learns Better Than You Think | Aayush Karan et.al. | 2610.02140v1 | null |
| 2026-10-01 | MIRTO: a registration-gated, multiverse-tested evaluation protocol for unsupervised anomaly segmentation in brain MRI | Negin Kafee Hernashki et.al. | 2610.02136v1 | null |
| 2026-10-01 | Local Support Learning | Assaf Ben-Kish et.al. | 2610.02126v1 | null |
| 2026-10-01 | Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows | Gabriel Tomitsuka et.al. | 2610.02122v1 | null |
| 2026-10-01 | Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes | Sophia Sirko-Galouchenko et.al. | 2610.02117v1 | null |
| 2026-10-01 | A Comparative Explainability Framework for DeBERTa-v3 in Zero-Shot Medical Abstract Classification | Javier Diaz Esteban-Herreros et.al. | 2610.02116v1 | null |
| 2026-10-01 | Scalable, Transferable Meta-network for Data Selection Requires a Different Loss (and Why the Obvious Choice is Problematic) | Zilin Du et.al. | 2610.02092v1 | null |
| 2026-10-01 | GeoLatent: Geometry-Guided Latent Structuring with Routed Optimization for 3D Reasoning | Yakun Zhu et.al. | 2610.02091v1 | null |
| 2026-10-01 | LLM2Jev: LLMs Are Already Jev-Style Decision Models -- When and How to Fine-Tune Them | Yinheng Li et.al. | 2610.02076v1 | null |
| 2026-10-01 | Causal Memory Policy: Making Memory Utility Identifiable by Intervening on Retrieval | Arman Behnam et.al. | 2610.02070v1 | null |
| 2026-10-01 | External Observers May See More Clearly: Cross-Model Span-Level Hallucination Detection in Large Language Models via Hidden State Probing | Kingshuk Gupta et.al. | 2610.02066v1 | null |
| 2026-10-01 | HydroJEV: A one-second, training-free screen for cyber-attack and fault attribution in water distribution networks | Tianwei Mu et.al. | 2610.02048v1 | null |
| 2026-10-01 | Typological Alignment of Stack-Based Language Models on Mildly Context-Sensitive Artificial Languages | Nadine El-Naggar et.al. | 2610.02040v1 | null |
| 2026-10-01 | CARM: Cancellation-Aware Response Masking for LLM Reinforcement Learning | Yafei Zhang et.al. | 2610.02039v1 | null |
| 2026-10-01 | Mimir: Physics-Grounded LLM Agents for Long-Horizon Irrigation Control | Yimeng Liu et.al. | 2610.02038v1 | null |
| 2026-10-01 | Global Coherence: When Every Agent Is Right and the Team Is Still Wrong - A Local-to-Global Semantic Foundation for Multi-Agent Collaboration | Xin Heng et.al. | 2610.02036v1 | null |
| 2026-10-01 | SPHERE: Adaptive VR Indoor Scene Generation via LLM-Enhanced Spatial Preference Learning and Human-in-the-Loop RL | Hyeonmin Lee et.al. | 2610.02023v1 | null |
| 2026-10-01 | Old Ideas, Novel Problems: The Instability of LLM-Based Novelty Evaluation | Noy Sternlicht et.al. | 2610.02022v1 | null |
| 2026-10-01 | Task-Adaptive Grounded 3D-Programmers Using 2D VLMs | Arman Raayatsanati et.al. | 2610.02021v1 | null |
| 2026-10-01 | Controllable Multi-label Video Safety Detection via Adaptive Tversky Policy Optimization | Guangyu Yang et.al. | 2610.02019v1 | null |
| 2026-10-01 | On Language Drift during RLVR Post-Training | Michael Sullivan et.al. | 2610.02015v1 | null |
| 2026-10-01 | Atoms to Processes: The Role of Artificial Intelligence and Machine Learning in Chemical Engineering | Michael Baldea et.al. | 2610.02014v1 | null |
| 2026-10-01 | Counting Moves, Weighing Voices: Bayesian Dialectical Argumentation for Calibrated Multi-LLM Councils under Persistent Adversaries | Ionel Eduard Stan et.al. | 2610.02005v1 | null |
| 2026-10-01 | Mem++: Non-Destructive Memory for Long-Term Organizational LLM Agents | Ahmad Yehia et.al. | 2610.02002v1 | null |
| 2026-10-01 | Mingbird: A Local-First Agent Harness Enabling Small Open Models to Complete Real Tasks | Hao Wang et.al. | 2610.02001v1 | null |
| 2026-10-01 | Can AI Oversight Be Zero Knowledge? | Alessandro Chiesa et.al. | 2610.01995v1 | null |
| 2026-10-01 | Universal Byte-Level Encoding: UTF-8/UTF-16 Routing to Reduce Cross-Script Token-Budget Disparities | Hyunsik Kim et.al. | 2610.01984v1 | null |
| 2026-10-01 | Counterfactual Auditing of Bias in Open-Source Large Language Models for Clinical Triage | Manar Aljohani et.al. | 2610.01963v1 | null |
| 2026-10-01 | A Hybrid Approach to Malware Detection: Integrating Few-Shot Model-Agnostic Meta-Learning with Autoencoders | Emmanuela Andam et.al. | 2610.01949v1 | null |
| 2026-10-01 | Latent JEPA: Abstract Future Prediction for Latent Reasoning in Chemistry | Xinjian Zhao et.al. | 2610.01947v1 | null |
| 2026-10-01 | A rubric landscape for evaluating clinical reasoning in large language models: what exists, what is missing, and what needs to be combined | Zhangshu Joshua Jiang et.al. | 2610.01938v1 | null |
| 2026-10-01 | Mapping the RAG Landscape: A Four Axis Taxonomy of Efficiency, Defense, Interactivity, and Reasoning | Meghana Sunil et.al. | 2610.01936v1 | null |
| 2026-10-01 | Cross-Lingual Alignment for Decoder-Only Models using MoE Routers | Lucas Bandarkar et.al. | 2610.01921v1 | null |
| 2026-10-01 | MoLE: Mixture of Latent Experts for Complementary Visual Reasoning | Yingcheng Liu et.al. | 2610.01917v1 | null |
| 2026-10-01 | Asynchronous LLM Post-Training: Group-Mass Capping and Convergence Analysis | Qijia He et.al. | 2610.01896v1 | null |
| 2026-10-01 | A Structured State Space Sequence Model for Multi-Class Classification of Malware | Emmanuela Andam et.al. | 2610.01893v1 | null |
| 2026-10-01 | Selection-Based Structured Reasoning: Toward Efficient Multimodal Search Agents | Feiyu Gavin Zhu et.al. | 2610.01892v1 | null |
| 2026-10-01 | Unsupervised Domain Adaptation for Enhanced Radiometer Image Precipitation Estimation using Conditional Flow Matching | Victor Enescu et.al. | 2610.01890v1 | null |
| 2026-10-01 | Stochastic Rounding in Low-Precision Transformer Inference: A Variable-Precision Emulation Study of a Small GPT-2 | Yohan Chatelain et.al. | 2610.01889v1 | null |
| 2026-10-01 | Flowing Faster to Coordinate: One-Step Online Multi-Agent Flow Policies | Zhuoran Li et.al. | 2610.01882v1 | null |
| 2026-10-01 | Where LLMs Fail with Visualization DSLs | Chang Han et.al. | 2610.01873v1 | null |
| 2026-10-01 | From Network Intrusion Detection to Blockchain-Backed Endpoint Detection and Response: Mapping the Landscape of Decentralized Detection-and-Response Architectures | Yahya Shahsavari et.al. | 2610.01872v1 | null |
| 2026-10-01 | Walking the Embedding Space: Datastore Extraction from Multimodal RAG | Maria Carmen Jica et.al. | 2610.01871v1 | null |
| 2026-10-01 | From Isolated Feature to Orbits: Discovering Music Concepts via Multi-SAE Alignment | Liwei Lin et.al. | 2610.01864v1 | null |
| 2026-10-01 | AVSD-Scenes: A Dataset for Audio-Visual Description of Urban Scenes | Dhanunjaya Varma Devalraju et.al. | 2610.01861v1 | null |
| 2026-10-01 | Detecting Inconsistencies in Model Specifications with LLM-as-Verifier Reasoning | Zichen Xie et.al. | 2610.01847v1 | null |
| 2026-10-01 | Beyond Decodability: Do Acoustic Factors Drive Predictions in Speech-Based Alzheimer's Assessment? | Serli Kopar et.al. | 2610.01846v1 | null |
| 2026-10-01 | On the Divergence of Accuracy and Mechanism Consistency in Time Series World Models | Haochen Zhang et.al. | 2610.01842v1 | null |
| 2026-10-01 | Code Owns the Simulation, Jev Owns the Evaluation | Yaodong Yang et.al. | 2610.01834v1 | null |
| 2026-10-01 | Continuous Process-Level Evaluation for Evolving Enterprise AI Agent Skills | Ngoc Phuoc An Vo et.al. | 2610.01833v1 | null |
| 2026-10-01 | The Asymptotics of Language Model Alignment with Memory | Haricharan Balasundaram et.al. | 2610.01828v1 | null |
| 2026-10-01 | Token Communication-Assisted Collaborative Embodied Artificial Intelligence: Concepts, Framework, and Opportunities | Peng Yi et.al. | 2610.01826v1 | null |
| 2026-10-01 | Beyond Linear Concepts: Discovering and Aligning Non-Linear Concept Manifolds in Large Language Models | Tido Specht et.al. | 2610.01821v1 | null |
| 2026-10-01 | A Safe Prototype Is Not a Safety Direction: Reference Dependence and Prompt Confounds in Response-Safety Embeddings | Sahil Kadadekar et.al. | 2610.01801v1 | null |
| 2026-10-01 | LineupRL: Verifiable Reinforcement Learning for Time Series Captioning via Caption-to-Series Identification | Haochen Zhang et.al. | 2610.01800v1 | null |
| 2026-10-01 | iADD: Improving Alignment and Diversity in Diffusion Policy Optimization | Ashok Prasad Neupane et.al. | 2610.01789v1 | null |
| 2026-10-01 | VETO: Video Efficient Token Optimization for Vision Language Models | Gueter Josmy Faure et.al. | 2610.01785v1 | null |
| 2026-10-01 | Q-Learning for Reachability in MEC-Free MDPs | Lu-Chin Chang et.al. | 2610.01781v1 | null |
| 2026-10-01 | CODesign: Consistency from Data to Trajectory in All-Atom Protein Binder Co-Design | Yuanle Mo et.al. | 2610.01773v1 | null |
| 2026-10-01 | A Matryoshka Hierarchical RAG for Efficient Multi-Hop Question Answering | Gianluca Bonifazi et.al. | 2610.01767v1 | null |
| 2026-10-01 | VideoEvolve: Evolving Agent Harnesses for Video Temporal Grounding | Bingjun Luo et.al. | 2610.01766v1 | null |
| 2026-10-01 | TopK-Guided: Adaptive, Budget-Aware Activation Sparsity for Efficient LLM Inference | Mukund Agarwalla et.al. | 2610.01763v1 | null |
| 2026-10-01 | SoK: Decentralized Agent Economic Infrastructure | Rui Sun et.al. | 2610.01756v1 | null |
| 2026-10-01 | Cog-VADU: A Training-Free Cognitive Reasoning Framework for Video Anomaly Detection and Understanding | Mohd Ubaid Wani et.al. | 2610.01754v1 | null |
| 2026-10-01 | Removing spurious minima for planar features by skip connections | Jakob Paul Zimmermann et.al. | 2610.01728v1 | null |
| 2026-10-01 | vFedProtoQNAS: Prototype-Guided Personalized Quantum Neural Architecture Search for Virtual Federated Learning | Seok Bin Son et.al. | 2610.01718v1 | null |
| 2026-10-01 | CoEvolve: Construct-to-Edit Visual Grounding with Bidirectional State Refinement | Dongwei Sun et.al. | 2610.01710v1 | null |
| 2026-10-01 | Task-Oriented Rank Adaptation for Continual Learning in Text Classification | Rey Sanchez Lopez et.al. | 2610.01702v1 | null |
| 2026-10-01 | Acmite: Mitigating Gender Bias in LLMs through Concept-Guided Mutual Information | Tian Lan et.al. | 2610.01696v1 | null |
| 2026-10-01 | Compound interpretation is based on analogy | Tian Shen et.al. | 2610.01688v1 | null |
| 2026-10-01 | Architectural Sampling: Test-Time Scaling via Computational Diversity in Frozen Vision-Language Models | Akshit Singh et.al. | 2610.01687v1 | null |
| 2026-10-01 | Iterative Policy Refinement through Semantic Rollout Analysis | Feiyu Gavin Zhu et.al. | 2610.01652v1 | null |
| 2026-10-01 | MCIR: A Feature Dependence-Aware Explainability Method with Reliability Guarantees | Poushali Sengupta et.al. | 2610.01641v1 | null |
| 2026-10-01 | Not All Error Yields to Scale: Where Scaling Stops in Vision-Language Inference | Xinye Zhao et.al. | 2610.01640v1 | null |
| 2026-10-01 | Yo-ByT5: Efficient and High-Fidelity Diacritic Restoration for Yorùbá | Ahmad Samuel Gali et.al. | 2610.01634v1 | null |
| 2026-10-01 | What Makes Something Hard(er)? Explaining Question Difficulty in Natural Language | Peng Cui et.al. | 2610.01627v1 | null |
| 2026-10-01 | FedLore: Communication and Memory Efficient Federated Learning via Shared Gradient Low-Rank Projection | Junkang Liu et.al. | 2610.01620v1 | null |
| 2026-10-01 | Exposing the Cost of Deep Learning Audio Development | Constance Douwes et.al. | 2610.01619v1 | null |
| 2026-10-01 | Agents Are Systems, Not Models: Rethinking Agentic Evaluation | Luis Wiedmann et.al. | 2610.01618v1 | null |
| 2026-10-01 | Can LLMs Reliably Annotate Bioassay Metadata to Improve Data Readiness? | Laura van Weesep et.al. | 2610.01616v1 | null |
| 2026-10-01 | Hob-VL: A Benchmark for Visually Grounded Boolean Reasoning | Yuzhou Wang et.al. | 2610.01605v1 | null |
| 2026-10-01 | Permutation-Robust Decision Modeling with Candidate-Independent Block-Causal Attention | Guy Amit et.al. | 2610.01601v1 | null |
| 2026-10-01 | Before It Fades: Reinforcing Temporal Representations at Inference Time in VideoLLMs | Youngwoo Shin et.al. | 2610.01595v1 | null |
| 2026-10-01 | Which LLM to pick? Online Active Model Selection for Large Language Models | Alessandro Turrin et.al. | 2610.01592v1 | null |
| 2026-10-01 | Evaluating Physical Consistency and Plausibility in Generative Scenario Models for Autonomous Driving | Manasa Mariam Mammen et.al. | 2610.01581v1 | null |
Abstracts
One Basis to Animate Them All: Gaussian Blendshape Distillation for Real-Time Avatars
2610.02207v1 by Ramazan Fazylov, Stamatis Lefkimmiatis, Ivan Laptev
3D Gaussian avatars support fast rendering, however, their real-time animation is often challenged by the costly neural inference. We address this bottleneck and show that the animation of pretrained avatar models can be closely approximated by a linear combination of identity-independent blendshapes. Building on this finding, we introduce GALA (Gaussian Animation via Linear Approximation), a distillation method that replaces per-frame heavy neural decoding with a shallow coefficient predictor and a linear blend. To improve fidelity and reduce memory requirements, we propose to construct the basis using block-local PCA under a rendering-aware metric and a memory budget. Our method learns a shallow MLP network to predict blendshape coefficients and applies to various animation architectures without retraining original models. We validate GALA by accelerating the inference of three distinct avatar models for 3D animation of facial expressions and full-bodies with clothing dynamics. Across these models, our distillation generalizes to held-out identities and reduces CPU animation cost by up to three orders of magnitude while preserving most of the rendering quality. Excellent results of our method confirm the shared linear structure of learned avatar representations and enable highly efficient and accurate animation at frame rates reaching up to 60fps on mobile devices. Project page: https://ramazan793.github.io/gala/
摘要:3D 高斯虛擬角色支援快速渲染,然而,它們的即時動畫常常受到昂貴的神經推理的挑戰。
我們解決了這一瓶頸,並顯示預訓練虛擬角色模型的動畫可以通過身份無關的混合形狀的線性組合來密切近似。
基於這一發現,我們引入了 GALA(通過線性近似的高斯動畫),這是一種蒸餾方法,將每幀重的神經解碼替換為淺層係數預測器和線性混合。
為了提高保真度並減少內存需求,我們提出在渲染感知度量和內存預算下使用區塊局部 PCA 來構建基底。
我們的方法學習一個淺層 MLP 網絡來預測混合形狀係數,並應用於各種動畫架構而無需重新訓練原始模型。
我們通過加速三個不同虛擬角色模型的推理來驗證 GALA,這些模型用於 3D 動畫的面部表情和全身帶衣物動態。
在這些模型中,我們的蒸餾方法對保留的身份進行了泛化,並將 CPU 動畫成本降低了多達三個數量級,同時保持大部分渲染質量。
我們方法的優異結果證實了學習的虛擬角色表示的共享線性結構,並使得在移動設備上以高達 60fps 的幀率進行高效且準確的動畫成為可能。
專案頁面:https://ramazan793.github.io/gala/
KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards
2610.02206v1 by Pengfei Li, Naufal Suryanto, Sicheng Zhang, Muzammal Naseer
LLMs are increasingly applied to cybersecurity workflows, where they are expected to translate analysts' intent into tool invocations. However, existing evaluations focus on knowledge-based assessments or end-to-end agentic tasks, and do not directly measure LLMs' ability to generate executable commands for real-world cybersecurity tools. This gap is critical because cybersecurity operations rely on strict command-line interfaces (CLIs), where minor syntax errors, incorrect flag--value bindings, or argument misordering can invalidate execution. We introduce KaliBench, a fine-grained benchmark and dataset for natural-language--to--CLI translation on Kali Linux, comprising 8,504 query--command pairs spanning 1,642 tools across 23 capability dimensions and 5 security phases. KaliBench is constructed via a manuscript-grounded pipeline with deterministic canonicalization and alias-aware evaluation, enabling precise and reproducible assessment of tool selection and argument construction. To ensure both semantic correctness and practical executability, we develop a multi-stage verification pipeline that combines LLM-based validation, sandboxed terminal execution, and human-in-the-loop refinement. Building on these fine-grained, deterministic signals, KaliBench further enables runtime-free verifiable rewards for training. Across three evaluation modes and 24 configurations of general-purpose and security-focused open-weight models, no open-weight model exceeds 42% exact-command accuracy in the unrestricted setting, highlighting the difficulty of accurate CLI-based cybersecurity tool use without explicit tool hints. We further show that supervised fine-tuning and reinforcement learning with verifiable rewards derived from KaliBench significantly improve an 8B model and achieve performance comparable to a 685B MoE model.
摘要:LLMs 正在越來越多地應用於網絡安全工作流程中,它們被期望將分析師的意圖轉化為工具調用。
然而,現有的評估主要集中在基於知識的評估或端到端的代理任務上,並未直接測量 LLMs 生成可執行命令以供現實世界網絡安全工具使用的能力。
這一差距至關重要,因為網絡安全操作依賴於嚴格的命令行界面 (CLIs),其中微小的語法錯誤、不正確的標誌--值綁定或參數錯序都可能使執行無效。
我們介紹了 KaliBench,這是一個針對 Kali Linux 的自然語言到 CLI 翻譯的細粒度基準和數據集,包含 8,504 個查詢--命令對,涵蓋 1,642 種工具,跨越 23 個能力維度和 5 個安全階段。
KaliBench 是通過一個基於手稿的管道構建的,具有確定性的標準化和別名感知評估,能夠精確且可重複地評估工具選擇和參數構建。
為了確保語義正確性和實際可執行性,我們開發了一個多階段驗證管道,結合了基於 LLM 的驗證、沙盒終端執行和人類參與的精煉。
基於這些細粒度的確定性信號,KaliBench 進一步使得無運行時的可驗證獎勵成為訓練的可能。
在三種評估模式和 24 種通用及安全專注的開放權重模型配置中,沒有任何開放權重模型在不受限制的設置中超過 42% 的精確命令準確率,突顯了在沒有明確工具提示的情況下準確使用基於 CLI 的網絡安全工具的困難。
我們進一步顯示,從 KaliBench 派生的可驗證獎勵的監督微調和強化學習顯著改善了一個 8B 模型,並達到了與 685B MoE 模型相當的性能。
Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents
2610.02204v1 by Yen-Jen Wang, Haozhe Jiang, Shuying Deng, Haoru Xue, Weirui Ye, Rocky Duan, Nika Haghtalab, S. Shankar Sastry, Pieter Abbeel, Haozhi Qi
Building reliable robot capabilities across diverse tasks requires substantial human effort to develop and maintain skills, design rewards, and integrate perception with control. We present Reconstruct, Practice, Go Real (RPG), a framework for autonomous improvement of robot execution systems without updating model weights. RPG identifies manipulation capabilities in an offline dataset and constructs related practice tasks in simulation. During practice, RPG uses execution feedback, privileged simulator state, and available dataset videos to diagnose failures. It develops new reusable symbolic skills, refines existing skills, and revises the system prompt based on these diagnoses. Cross-task evaluation tests individual candidate changes and merged revisions before they are retained for reuse. At test time, a multimodal LLM uses the resulting system prompt and skill library to coordinate perception and robot control. On held-out initializations of 22 manipulation tasks, RPG improves task success from 28.6% after the first practice round to 95.0% after 15 rounds, outperforming all evaluated baselines, including ASPIRE (75.5%) and CaP-Agent0 powered by GPT-6 Astra Pro (60.0%). After a common calibration and hardware-adaptation procedure, the frozen system succeeds in all 30 physical trials, with ten trials on each of three tasks. Project Website: https://rpg-robot.github.io/
摘要:建立可靠的機器人能力以應對多樣化任務需要大量的人力來發展和維護技能、設計獎勵,以及將感知與控制整合在一起。
我們提出了重建、練習、實現真實(RPG),這是一個無需更新模型權重的自主改進機器人執行系統的框架。
RPG 在離線數據集中識別操作能力,並在模擬中構建相關的練習任務。
在練習過程中,RPG 利用執行反饋、特權模擬器狀態和可用的數據集視頻來診斷失敗。
它開發新的可重用符號技能,細化現有技能,並根據這些診斷修訂系統提示。
跨任務評估測試單個候選變更和合併修訂,然後才將其保留以便重用。
在測試時,多模態 LLM 使用生成的系統提示和技能庫來協調感知和機器人控制。
在 22 個操作任務的保留初始化中,RPG 將任務成功率從第一次練習回合後的 28.6% 提高到 15 回合後的 95.0%,超越了所有評估的基準,包括 ASPIRE(75.5%)和由 GPT-6 Astra Pro 驅動的 CaP-Agent0(60.0%)。
在經過共同的校準和硬體適應程序後,凍結的系統在所有 30 次實體試驗中成功,每個任務進行十次試驗。
項目網站:https://rpg-robot.github.io/
ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research
2610.02202v1 by Sohyeon Kim, Yoonho Lee, Bo Liu, Dayoon Ko, Rulin Shao, Seungone Kim, Graham Neubig, Pang Wei Koh, Aakanksha Chowdhery, Akari Asai, Omar Khattab, Yejin Choi, Gunhee Kim, Chelsea Finn
What makes great scientists great? Even as AI systems start to make progress on open problems, scientists remain far ahead of them at sensing which prior idea, buried in an ever-growing archive of research, a new problem needs. To study this skill, we draw on researchers who know firsthand which earlier work advanced their completed projects, with papers serving as pointers to the ideas within. Using our automated pipeline that makes author annotation scalable, we build ScholarCatalyst by having 184 lead authors of 207 recent computer science papers label which candidates did or could have advanced their project, each with a detailed rationale. We introduce a retrieval task with author-provided judgments: given an initial research question, retrieve these papers from only the literature available when the project began. Agentic search does no better than embedding retrieval (0.42 vs. 0.48 Recall@20) despite calling that same retriever as a tool. Even an agent built on Claude Fable 5.1, which may have seen the completed papers during training, reaches only 0.51 R@20. These results highlight the need for new training recipes that equip models with expert intuition for searching broad corpora. We envision ScholarCatalyst as a step toward scientific agents that can take a half-formed idea and point to the prior research it needs.
摘要:什麼使偉大的科學家偉大?即使人工智慧系統開始在開放問題上取得進展,科學家們在感知新問題需要哪個埋藏在不斷增長的研究檔案中的先前想法方面仍然遙遙領先。為了研究這項技能,我們依賴於那些親身了解哪些早期工作推進了他們完成項目的研究者,論文作為指向內部想法的指標。利用我們的自動化流程,使作者標註可擴展,我們通過讓184位207篇近期計算機科學論文的主要作者標註哪些候選者推進了或可能推進了他們的項目,每個標註都有詳細的理由,來構建ScholarCatalyst。我們引入了一個檢索任務,並提供作者的判斷:給定一個初步的研究問題,從項目開始時僅可用的文獻中檢索這些論文。即使將同一檢索工具稱為工具,主動搜索的表現也不如嵌入檢索(0.42對0.48 Recall@20)。即使是基於Claude Fable 5.1構建的代理,可能在訓練期間看過已完成的論文,R@20的結果也僅為0.51。這些結果突顯了需要新的訓練配方,以使模型具備專家直覺來搜索廣泛的語料庫。我們將ScholarCatalyst視為邁向科學代理的一步,這些代理可以將半成型的想法指向所需的先前研究。
VISTA: A Visual Harness for Reasoning in an Interactive World
2610.02200v1 by Qiushi Han, Keya Hu, Linlu Qiu, Cathy Wu, Kaiming He
We show that multimodal models possess strong reasoning abilities and that an appropriate harness can unlock their potential to solve tasks across diverse interactive environments. We introduce VISTA, a visual harness that gives a general-purpose multimodal model long-horizon vision. VISTA allows the model to directly perceive the environment through visual observations and maintains a lossless visual memory that preserves past observations in their original form. The model can actively retrieve these observations and reorganize its visual input as it reasons. On ARC-AGI-3, VISTA improves Claude Opus 5.0's Relative Human Action Efficiency score from 40.68 to a perfect 100.00, with the model completing all 25 public games using 57.4% fewer actions than first-time human participants. VISTA's simple design also allows it to extend naturally to diverse visual environments with minimal adaptation. Across three additional benchmarks covering a diverse range of visual games and puzzles, it substantially outperforms baselines using the same underlying model with minimal harnesses. Our results highlight VISTA's potential as a general-purpose visual harness for advancing multimodal agents in complex visual environments.
摘要:我們展示了多模態模型擁有強大的推理能力,並且適當的工具可以釋放它們在多樣互動環境中解決任務的潛力。我們介紹了 VISTA,一種視覺工具,使通用多模態模型具備長期視覺能力。VISTA 允許模型通過視覺觀察直接感知環境,並保持無損的視覺記憶,保留過去觀察的原始形式。模型可以主動檢索這些觀察並在推理時重新組織其視覺輸入。在 ARC-AGI-3 上,VISTA 將 Claude Opus 5.0 的相對人類行動效率分數從 40.68 提升至完美的 100.00,模型在完成所有 25 個公共遊戲時,使用的行動比首次參與的人類參與者少 57.4%。VISTA 的簡單設計也使其能夠在不同的視覺環境中自然擴展,適應性極小。在涵蓋多樣視覺遊戲和謎題的三個額外基準中,它顯著超越了使用相同基礎模型和最小工具的基準。我們的結果突顯了 VISTA 作為推進多模態代理在複雜視覺環境中的通用視覺工具的潛力。
Hierarchical Continuous Diffusion Language Models
2610.02193v1 by Hui Ren, Zihan Li, Chang Liu, Huidong Liu, Alexander Schwing
Discrete diffusion language models offer a compelling alternative to autoregressive generation for tasks demanding bidirectional reasoning and global constraint satisfaction. Yet they share a structural bottleneck: when decoding in parallel, each token is sampled independently from its marginal, severing the statistical dependencies among the tokens decoded together. Continuous diffusion language models avoid this by denoising a shared continuous state, but their denoiser sees only that state, so nothing ties it to a valid token configuration until it is finally decoded. To address this, we propose Hierarchical Continuous Diffusion Language Models (HC-DLM), which couple discrete token generation with a continuous latent trajectory in a single, principled denoising process, whose training objective is derived from a variational bound on the token likelihood. In contrast to recent methods that attach continuous context to a self-contained discrete chain, HC-DLM makes the latent the only persistent generative state: tokens are read out from it at every step and feed back as a scaffold for the next latent update. On structured reasoning (Sudoku), mathematical planning (Countdown) and language modeling (LM1B), HC-DLM improves over discrete and continuous diffusion baselines at matched model size, in puzzle accuracy on Sudoku and Countdown and in generative perplexity on LM1B. Project page: https://hc-dlm.github.io/.
摘要:離散擴散語言模型為需要雙向推理和全局約束滿足的任務提供了一個引人注目的替代方案,取代自回歸生成。
然而,它們共享一個結構瓶頸:在並行解碼時,每個標記都是獨立從其邊際中抽樣的,這切斷了一起解碼的標記之間的統計依賴關係。
連續擴散語言模型通過去噪共享的連續狀態來避免這一點,但它們的去噪器僅看到該狀態,因此在最終解碼之前,沒有任何東西將其與有效的標記配置聯繫起來。
為了解決這個問題,我們提出了層次連續擴散語言模型(HC-DLM),它將離散標記生成與單一、原則性的去噪過程中的連續潛在軌跡結合起來,其訓練目標源自於標記似然的變分界限。
與最近將連續上下文附加到自包含離散鏈的方法相比,HC-DLM使潛在變量成為唯一持久的生成狀態:標記在每一步都從中讀出,並作為下一次潛在更新的支架進行反饋。
在結構推理(數獨)、數學規劃(倒計時)和語言建模(LM1B)方面,HC-DLM在匹配模型大小的情況下優於離散和連續擴散基準,在數獨和倒計時的拼圖準確性以及LM1B的生成困惑度上均有所改善。
項目頁面:https://hc-dlm.github.io/.
DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation
2610.02188v1 by Zhengming Yu, Junkun Yuan, Haotian Yang, Gordon Guocheng Qian, Yizhi Wang, Angtian Wang, Yiding Yang, Bo Liu, Xin Li, Wenping Wang, Chongyang Ma
Distribution Matching Distillation (DMD) trains a few-step student from the difference between separately estimated target and student scores, so it must keep an auxiliary diffusion model fitted to the student's evolving distribution at extra memory and computation cost. We introduce DMAD, Distribution Matching as Adversarial Distillation, which recasts distribution matching as classification and learns the required log-density ratios directly. Two discriminator heads on a shared backbone distinguish real data and teacher samples from the student's, and linear losses on their logits train the student without auxiliary score fitting. We prove that at the discriminator optimum these losses recover the distribution-matching gradient underlying DMD, through the classical identity linking discriminator logits to log-density ratios. We further introduce gap-based reweighting, which adapts teacher supervision across noise levels from the real-data head's empirical logit gap between real and teacher samples. DMAD reaches a Fréchet Inception Distance (FID) of 1.04 with one-step generation on ImageNet-64x64, 14.47 with four-step SDXL on COCO-10K, and a VBench total score of 85.15 with four-step Wan2.1-T2V-14B, the best values among the compared few-step methods and the multi-step teachers. On MiniMax-H3-33B, our four-step student achieves overall human preference rates of 79.1% over DMD2 and 84.6% over rCM for joint audio-video generation, excluding ties. Our code, models and demos are available at https://yzmblog.github.io/projects/DMAD.
摘要:分佈匹配蒸餾(DMD)從分別估計的目標和學生分數之間的差異訓練出幾步的學生,因此它必須保持一個輔助擴散模型,以適應學生不斷變化的分佈,這需要額外的記憶體和計算成本。我們引入 DMAD,作為對抗蒸餾的分佈匹配,將分佈匹配重新表述為分類,並直接學習所需的對數密度比率。兩個共享主幹的鑑別器頭區分真實數據和教師樣本與學生的樣本,並通過它們的邏輯值進行線性損失訓練學生,而無需輔助分數擬合。我們證明,在鑑別器最佳點,這些損失恢復了 DMD 背後的分佈匹配梯度,這是通過經典的身份將鑑別器邏輯值與對數密度比率聯繫起來。我們進一步引入基於差距的重加權,這根據真實數據頭的真實樣本和教師樣本之間的經驗邏輯差距,調整教師監督在噪聲水平上的適應性。DMAD 在 ImageNet-64x64 上以一步生成達到 1.04 的 Fréchet Inception Distance(FID),在 COCO-10K 上以四步 SDXL 達到 14.47,在四步 Wan2.1-T2V-14B 上達到 85.15 的 VBench 總分,這是與比較的幾步方法和多步教師中的最佳值。在 MiniMax-H3-33B 上,我們的四步學生在聯合音頻-視頻生成中達到了 79.1% 的人類偏好率,超過 DMD2 和 84.6% 超過 rCM,排除平局。我們的代碼、模型和演示可在 https://yzmblog.github.io/projects/DMAD 獲得。
Higher-Order Molecular Grammars for Generative and Foundation Models in Chemistry
2610.02186v1 by Yiming Huang, Yujie Zeng, Vijay Prakash Dwivedi, Simone Foti, Jianmin Wang, Jure Leskovec, Tolga Birdal
Molecular learning models are strongly shaped by their underlying representations. Yet standard sequential and graph formalisms struggle to explicitly encode higher-order topology, such as ring systems and recurring motifs. Existing higher-order representations can capture these structures directly, but they are often computationally demanding and difficult to decode into valid molecules. Here, we introduce Higher-order Grammar Representation (HGR), a principled, topology-aware framework that lifts molecules to combinatorial complexes and parses each complex into a compact sequence of production rules under a context-free higher-order grammar. By serialising higher-order topology into rule sequences, HGR makes these structures directly compatible with standard sequence models, avoiding the computational overhead of explicit higher-order encodings while preserving topological expressiveness. To reduce benchmark bias towards simple ring systems, we construct RingDiv, a ring-enriched benchmark containing 1.18 million molecules, including the curated RingDiv300k subset, and introduce the ring diversity index (RDI) to quantify ring-system coverage. In molecular generation, HGR-based models uniquely combine 100% validity by construction with leading distributional alignment, ranking first in FCD on all five generation benchmarks. In representation learning, HGR-FM achieves the highest mean AUC across seven MoleculeNet benchmarks under both transfer protocols, improving on the strongest baseline by 8.3 and 3.3 AUC points under probing and full fine-tuning, respectively. Collectively, these results establish HGR as an efficient higher-order representation for molecular generation and transferable representation learning.
摘要:分子學習模型受到其基礎表示的強烈影響。
然而,標準的序列和圖形形式在明確編碼高階拓撲方面(如環系統和重複圖案)面臨挑戰。
現有的高階表示可以直接捕捉這些結構,但它們通常計算需求高且難以解碼為有效的分子。
在此,我們介紹高階語法表示(HGR),這是一個原則性、關注拓撲的框架,將分子提升為組合複合體,並將每個複合體解析為在上下文無關的高階語法下的緊湊生成規則序列。
通過將高階拓撲序列化為規則序列,HGR使這些結構與標準序列模型直接兼容,避免了明確高階編碼的計算開銷,同時保留了拓撲表達能力。
為了減少對簡單環系統的基準偏見,我們構建了RingDiv,這是一個包含118萬個分子的環增強基準,包括精心策劃的RingDiv300k子集,並引入環多樣性指數(RDI)來量化環系統的覆蓋範圍。
在分子生成方面,基於HGR的模型獨特地結合了100%的有效性(由構造決定)與領先的分佈對齊,在所有五個生成基準中FCD排名第一。
在表示學習方面,HGR-FM在七個MoleculeNet基準中,在兩種轉移協議下實現了最高的平均AUC,相較於最強基線分別提高了8.3和3.3 AUC點(在探測和完全微調下)。
綜合這些結果,HGR確立了作為分子生成和可轉移表示學習的高效高階表示。
SoftServe: A Scalable Quasi-Newton Method for Deep Learning
2610.02182v1 by Joohwan Ko, Tetiana Parshakova, Diana Cai, Robert M. Gower
Quasi-Newton (QN) methods have long been among the most effective methods for large-scale unconstrained convex optimization. Two obstacles have limited their use in deep learning: non-convexity and enormous parameter sizes. We introduce SoftServe, a family of QN methods designed to overcome these obstacles without line searches or ad hoc curvature corrections. SoftServe derives positivedefinite curvature estimates from the variational objective of Berglund et al. (2025), even in the presence of negative curvature. We develop diagonal and Kroneckerfactored variants that preserve positive definiteness by construction and scale to massive neural networks. Finally, SoftServe relies on the stable coupled Newton-Schulz iteration for the required matrix operations, replacing costly matrix decompositions with GPU-friendly matrix multiplications. SoftServe excels on problems that are severely ill-conditioned, including tasks such as recurrent networks, deep autoencoders, physics-informed neural networks, and a 136M-parameter physics-informed diffusion model, often achieving lower losses than established baselines including Adam, Muon, and SOAP.
摘要:準牛頓(QN)方法長期以來一直是大規模無約束凸優化中最有效的方法之一。
兩個障礙限制了它們在深度學習中的應用:非凸性和巨大的參數規模。
我們介紹了 SoftServe,一系列旨在克服這些障礙的 QN 方法,無需進行線搜索或臨時的曲率修正。
SoftServe 從 Berglund 等人(2025)的變分目標中推導出正定的曲率估計,即使在存在負曲率的情況下也是如此。
我們開發了對角和克羅內克分解變體,這些變體通過構造保持正定性並能擴展到大型神經網絡。
最後,SoftServe 依賴於穩定的耦合牛頓-舒爾茨迭代來進行所需的矩陣操作,將成本高昂的矩陣分解替換為適合 GPU 的矩陣乘法。
SoftServe 在極度病態的問題上表現出色,包括循環網絡、深度自編碼器、物理知識神經網絡以及一個 136M 參數的物理知識擴散模型等任務,通常實現比包括 Adam、Muon 和 SOAP 在內的既定基準更低的損失。
Generative Cinematographer: Composing Camera and Object Motion in 3D
2610.02180v1 by Jiahan Zhang, Chaohao Yang, Namitha Guruprasad, Vivekjyoti Banerjee, Trong-Tung Nguyen, Alan Yuille, Anand Bhattad
Current controllable video generation systems often rely on 2D motion trajectories or sparse drag signals for object motion. These controls are ambiguous because the same 2D trajectory can correspond to different 3D motions, especially when the camera and objects move simultaneously. We present Generative Cinematographer (GenCine), a system that lifts a single image into an editable 3D scene scaffold where artists jointly author camera and foreground motion. Artists specify a camera path and move selected foreground regions using local 3D motion handles. Several handles can move different parts of a subject independently, providing a piecewise-rigid approximation to non-rigid motion without a physics simulator or category-specific prior. To communicate these controls to a pretrained video model, we project them into guidance maps. These maps record where the controlled regions appear in each frame, assign each handle a fixed color across frames and encode the current 3D positions of its controlled points in the same world coordinate system as the background. This lets us describe object motion relative to the scene even as the camera moves. For training, we recover controls from the motion observed in real videos and use ground-truth geometry and trajectories from synthetic videos. We train a lightweight guidance branch and LoRA adapters on a pretrained Wan model to follow these controls. Our experiments show consistent camera-relative motion, improved geometric consistency under viewpoint changes, and strong controllability across diverse real-world scenes.
摘要:目前可控的視頻生成系統通常依賴於 2D 動作軌跡或稀疏的拖曳信號來控制物體運動。這些控制是模糊的,因為相同的 2D 軌跡可以對應於不同的 3D 動作,特別是在相機和物體同時移動的情況下。我們提出了生成電影製作人(GenCine),這是一個將單一圖像提升為可編輯的 3D 場景框架的系統,藝術家可以共同創作相機和前景運動。藝術家指定相機路徑並使用局部 3D 動作手柄移動選定的前景區域。幾個手柄可以獨立移動主體的不同部分,提供對非剛性運動的分段剛性近似,而無需物理模擬器或特定類別的先驗知識。為了將這些控制傳達給預訓練的視頻模型,我們將它們投影到指導圖中。這些圖記錄了受控區域在每幀中出現的位置,為每個手柄在幀之間分配固定顏色,並以與背景相同的世界坐標系編碼其受控點的當前 3D 位置。這使我們能夠描述相對於場景的物體運動,即使相機在移動。為了訓練,我們從真實視頻中觀察到的運動中恢復控制,並使用合成視頻中的真實幾何和軌跡。我們在預訓練的 Wan 模型上訓練了一個輕量級的指導分支和 LoRA 適配器,以遵循這些控制。我們的實驗顯示出一致的相機相對運動、在視點變化下改進的幾何一致性,以及在多樣的現實場景中強大的可控性。
Every Ablation Is a Dose: Counterweights and the Semblance of Self-Repair
2610.02173v1 by Areeb Ahmad, Pratinav Seth, Vinay Kumar Sankarapu
Ablate a component of a language model, and other components often appear to adjust and compensate. This phenomenon, termed self-repair, has been observed repeatedly, but its mechanism remains unclear. The most systematic study to date concluded that self-repair is noisy and unlikely to have a single explanation. We argue that it has one: a gain already present before any ablation. Any intervention on a causally important component can be viewed as a point on a coordinate axis $λ$, the signed strength of a counterfactual contrast. Hence, conventional ablation methods are uncalibrated points on this axis. We show that the causal repair response for a fine-grained unit $r$ is governed by an affine law, $E_r(λ)=\mathrm{own}_r+γ_rλ$. The slope $γ_r$ is a fixed coefficient that consistently influences the model, with or without ablation, and its sign determines whether the unit counteracts or reinforces the removed signal. On a factual-verdict task across four models from distinct families (Gemma, Qwen, LLaMA, and Mistral), we identify components including MLP neurons, OV neurons, and singular directions that follow this affine law, 68 of 81 downstream directions in all. Moreover, we can anticipate the magnitude of $γ_r$ from the fixed weights. On the IOI circuit of GPT-2 Small, seven of the ten heads the intervention can reach follow the law, and all seven are counterweights. From this perspective, what may appear as self-repair is a counterweight performing its usual operation when the contrastive signal emerges at the core.
摘要:去除語言模型的一個組件時,其他組件往往會調整並補償。這一現象被稱為自我修復,已經多次被觀察到,但其機制仍不清楚。迄今為止,最系統的研究得出結論,自我修復是嘈雜的,並且不太可能有單一的解釋。我們主張它有一個:在任何去除之前已經存在的增益。對於一個因果重要的組件的任何干預可以被視為坐標軸 $λ$ 上的一個點,即反事實對比的有向強度。因此,傳統的去除方法是在這個軸上的未校準點。我們顯示,對於一個細粒度單元 $r$ 的因果修復反應受一個仿射法則支配,$E_r(λ)=\mathrm{own}_r+γ_rλ$。斜率 $γ_r$ 是一個固定的係數,持續影響模型,無論是否去除,其符號決定了該單元是抵消還是增強被移除的信號。在四個不同家族的模型(Gemma、Qwen、LLaMA 和 Mistral)中的事實判決任務中,我們識別出包括 MLP 神經元、OV 神經元和遵循這一仿射法則的單一方向的組件,總共 81 個下游方向中有 68 個。此外,我們可以從固定權重預測 $γ_r$ 的大小。在 GPT-2 Small 的 IOI 電路中,十個頭中有七個可以到達的干預遵循這一法則,並且這七個都是對重。從這個角度來看,可能看起來像自我修復的現象實際上是一個對重在對比信號出現時執行其正常操作。
AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents
2610.02163v1 by Xuan Zhang, Longtao Zheng, Cunxiao Du, Bo An, Xin Dong
Coding agents solve repository-level software engineering tasks through long trajectories of code inspection, search, editing, and testing. As a task progresses, earlier exploration becomes stale, so managing context is more than avoiding overflow: an agent must decide when to compact, what working state to preserve, and how to continue from it. We introduce AutoCompact, which trains a coding agent to make these decisions as part of its policy. To collect training data, we run the base agent on coding tasks and use a judge to review its compaction decisions, summaries, and actions after compaction. Flawed outputs are replaced with corrected ones before being executed in the environment, so each trajectory continues from the corrected decisions. We use these trajectories for supervised fine-tuning, then jointly optimize coding and compaction through reinforcement learning with task-success rewards. Experiments on SWE-bench Verified and SWE-PolyBench Verified show that AutoCompact improves pass rates over the base model by an absolute 9.2\% and 5.0\%, respectively. The improvements hold across all evaluated inference budgets, with a 256K context window that never overflows and with a 16K window whose overflow triggers fallback compaction.
摘要:編碼代理透過長期的代碼檢查、搜索、編輯和測試來解決庫級軟體工程任務。隨著任務的進展,早期的探索變得過時,因此管理上下文不僅僅是避免溢出:代理必須決定何時壓縮、保留什麼工作狀態,以及如何從中繼續。我們介紹了AutoCompact,這是一個訓練編碼代理作為其策略一部分來做出這些決策的系統。為了收集訓練數據,我們在編碼任務上運行基礎代理,並使用評審來審查其壓縮決策、摘要和壓縮後的行動。缺陷輸出在執行環境中之前會被更正的輸出所替代,因此每個軌跡都從更正的決策繼續。我們使用這些軌跡進行監督微調,然後通過強化學習與任務成功獎勵共同優化編碼和壓縮。在SWE-bench Verified和SWE-PolyBench Verified上的實驗顯示,AutoCompact相對於基礎模型的通過率分別提高了9.2\%和5.0\%。這些改進在所有評估的推理預算中均保持有效,使用256K的上下文窗口從不溢出,並且使用16K窗口時其溢出會觸發回退壓縮。
DuoMind: Enabling Distributed Multi-Robot Coordination with Semantic Communication
2610.02161v1 by Hanchu Zhou, Dechen Gao, Hang Wang, Brendan Lynch, Boqi Zhao, Qiyao Ma, Raman Goyal, Junshan Zhang
Vision-language models (VLMs) and vision-language-action models (VLAs) have recently driven rapid progress in general-purpose robots, yet most progress has focused on single-robot settings. Extending these capabilities to multi-robot systems remains challenging because robots must coordinate long-horizon behaviors while maintaining reliable, fine-grained execution. We introduce DuoMind, a distributed hierarchical framework for multi-robot coordination through semantic communication. Each robot uses a VLA-based action model for low-level execution and a VLM-based orchestrator for high-level reasoning and inter-agent coordination. At each planning step, the orchestrator at each robot reasons over the task instruction, local observations, and messages received from other robots. It then generates low-level instructions for the action model and semantic messages for peer robots. This architecture exploits the complementary strengths of pretrained models by combining the semantic reasoning capabilities of VLMs with the precise action-generation capabilities of VLAs. To address the scarcity of benchmarks for multi-robot coordination, we further develop RoboPoly, a benchmark comprising long-horizon manipulation tasks that require coordinated, closed-loop execution under distributed control. Experiments on RoboPoly and RoboTwin demonstrate that DuoMind improves multi-robot task performance, while ablation studies confirm the contributions of hierarchical orchestration and semantic communication. More details are available on our project page.
摘要:視覺語言模型(VLMs)和視覺語言行動模型(VLAs)最近在通用機器人領域推動了快速進展,然而大多數進展集中在單機器人環境中。將這些能力擴展到多機器人系統仍然具有挑戰性,因為機器人必須協調長期行為,同時保持可靠且精確的執行。我們介紹了DuoMind,一個通過語義通信實現多機器人協調的分散式層次框架。每個機器人使用基於VLA的行動模型進行低層次執行,並使用基於VLM的協調者進行高層次推理和代理間協調。在每個規劃步驟中,每個機器人的協調者會對任務指令、當地觀察和來自其他機器人的消息進行推理。然後,它生成行動模型的低層次指令和針對同儕機器人的語義消息。這種架構通過結合VLM的語義推理能力與VLA的精確行動生成能力,充分利用了預訓練模型的互補優勢。為了解決多機器人協調基準的稀缺性,我們進一步開發了RoboPoly,一個包含需要協調、閉環執行的長期操作任務的基準。在RoboPoly和RoboTwin上的實驗表明,DuoMind改善了多機器人任務的表現,而消融研究確認了層次協調和語義通信的貢獻。更多細節可在我們的項目頁面上獲得。
From Knowledge Access to Source Learning: Developing Source-Specific Competence
2610.02150v1 by Lucheng Fu, Kejing Xia, Yiyang Wang, Yiqiao Jin, Jinjin He, Xiyuan Yang, Haoxin Liu, Ye Yu, Haibo Jin, Yijia Xiao, Wenke Lee, B. Aditya Prakash, Haohan Wang
Large language model (LLM) agents increasingly rely on persistent external sources to solve sequences of knowledge-intensive tasks. Existing methods improve how source content is accessed and organized, while agent-memory systems preserve reusable knowledge from prior interactions, but repeated use of the same source is still largely treated as repeated access rather than an opportunity to progressively improve understanding of that source. We study source learning: developing reusable source-specific competence over a persistent authoritative source. We represent this competence with a persistent source model that captures reusable understanding of the source, including how its knowledge is structured, interpreted, and applied. To construct and progressively refine such models, we propose SourceLearn, which combines two complementary learning mechanisms. Self-Directed Source Learning identifies what remains incompletely understood and adaptively revisits the source, while Task-Guided Source Learning uses downstream experience to reveal local representational gaps and recurring needs in how source knowledge should be organized. In both cases, learning signals determine what should be reconsidered, while persistent updates are reconstructed from the authoritative source. Across five benchmarks and three LLM backends, SourceLearn achieves the best performance in 13 of 15 settings, with gains of up to 22.6 points over Hybrid RAG and substantial overall improvements over static source representations and experience-based memory baselines.
摘要:大型語言模型(LLM)代理越來越依賴持久的外部來源來解決一系列知識密集型任務。現有方法改善了如何訪問和組織來源內容,而代理記憶系統則保留了來自先前互動的可重用知識,但對同一來源的重複使用仍然主要被視為重複訪問,而不是逐步改善對該來源理解的機會。我們研究來源學習:在持久的權威來源上發展可重用的來源特定能力。我們用一個持久的來源模型來表示這種能力,該模型捕捉了對來源的可重用理解,包括其知識的結構、解釋和應用方式。為了構建和逐步完善這樣的模型,我們提出了SourceLearn,該模型結合了兩種互補的學習機制。自我導向來源學習識別尚未完全理解的內容並適應性地重新訪問來源,而任務引導來源學習則利用下游經驗揭示如何組織來源知識的局部表徵差距和重複需求。在這兩種情況下,學習信號決定了應該重新考慮的內容,而持久更新則是從權威來源重建的。在五個基準和三個LLM後端中,SourceLearn在15個設置中的13個中實現了最佳性能,與Hybrid RAG相比,增益高達22.6點,並且在靜態來源表示和基於經驗的記憶基準上有顯著的整體改進。
Keyword Harnesses Fail Open: A Cheap Diagnostic Ladder for Tool-Use Claims in Small Language Models
2610.02142v1 by Juan S. Santillana
Keyword-matching benchmarks can credit small models for tool use they never perform. We document such a false positive in a matched-architecture pair of Spanish security language models and propose a ladder of strict, cheap diagnostics. A 661.6M parameter model (approx. 65% code/technical text; no dedicated SFT) and a 1,109M model (web-heavy multi-phase curriculum; 6B-token tool-SFT) share decoder, tokenizer, and special tokens, scoring almost identically on lenient tool-use metrics (B4: 0.660 vs. 0.650). Verbatim-reproduction checks on training examples separate them completely: the 600M emits valid tool calls with generalized arguments on 6/6 examples; the 1B does so on 0/6 across checkpoints. A first-token probe localizes the 1B's failure to a missing prior (prob. $10^{-4}$--$10^{-5}$ on <|tool_call|>), which was erased by its web-heavy training phase. A targeted SFT recipe (diverse corpus, 5x higher learning rate, 2,202 steps, ~3.3 GPU-hours) repairs the 1B using three orders of magnitude fewer tokens than the failed phase. On all 269 corpus rows, valid emission rises from 0.100 to 0.959 (600M: 0.926). On 238 unseen prompts, the repaired 1B passes 0.536 vs. the 600M's 0.428 ($p = 0.004$). Embedding-drift checks show the repair did not move the trigger token's tied embedding (97.7% of the bf16 table remains bit-identical), meaning changes live in the surrounding network. Both models over-trigger, rarely answering negative prompts without a call (0.09 for 600M, 0.17 for repaired 1B). Factorial analyses confirm all repair configurations install the format, though suppression benefits from a diverse corpus remain a hypothesis due to seed sensitivity. This cheap diagnostic ladder costs minutes of CPU time and should gate tool-use claims on small models.
摘要:關鍵字匹配基準可以錯誤地將小型模型的工具使用歸功於它們從未執行過的行為。我們在一對匹配架構的西班牙安全語言模型中記錄了這樣的假陽性,並提出了一個嚴格且便宜的診斷階梯。一個661.6M參數的模型(約65%的代碼/技術文本;沒有專門的SFT)和一個1,109M的模型(以網絡為主的多階段課程;6B-token工具SFT)共享解碼器、分詞器和特殊標記,在寬鬆的工具使用指標上幾乎得分相同(B4: 0.660對0.650)。
對訓練範例的逐字重現檢查完全將它們分開:600M在6/6個範例中發出有效的工具調用,並帶有一般化的參數;而1B在各檢查點中則在0/6中發出。首次標記探測將1B的失敗定位於缺失的先前狀態(概率$10^{-4}$--$10^{-5}$在<|tool_call|>上),該狀態在其以網絡為主的訓練階段中被抹去。一個針對性的SFT配方(多樣化語料庫、5倍更高的學習率、2,202步驟、約3.3 GPU小時)使用比失敗階段少三個數量級的標記修復了1B。在所有269個語料行中,有效發射率從0.100上升至0.959(600M: 0.926)。在238個未見的提示中,修復後的1B通過了0.536,而600M則為0.428($p = 0.004$)。嵌入漂移檢查顯示修復並未移動觸發標記的綁定嵌入(97.7%的bf16表保持位元相同),這意味著變化存在於周圍的網絡中。
兩個模型都過度觸發,幾乎不在沒有調用的情況下回答負面提示(600M為0.09,修復後的1B為0.17)。因子分析確認所有修復配置都安裝了格式,儘管由於種子敏感性,來自多樣化語料庫的抑制效益仍然是一個假設。這個便宜的診斷階梯僅需幾分鐘的CPU時間,應該限制小型模型的工具使用聲明。
Finetuning with Sampling: SFT Learns Better Than You Think
2610.02140v1 by Aayush Karan, Sitan Chen, Yilun Du
Introducing new capabilities to frontier models has long been the goal of posttraining, which predominantly employs supervised finetuning (SFT) and reinforcement learning (RL) to this end. Conventional wisdom dictates that RL enables strong generalization on new tasks without losing existing capabilities, while SFT is prone to weak generalization and catastrophic forgetting. At the same time, SFT can learn from off-policy expert data, whereas RL must rely on a model's ability to find successful trajectories with repeated sampling. In our work, we seek to leverage the strength of on-policy learning while utilizing the privileged information contained in off-policy data. However, rather than modifying the learning objective to accommodate this data, we instead tailor the data distribution to better suit the learner. We introduce a Markov chain Monte Carlo (MCMC) sampling algorithm that progressively transforms off-policy traces to be more on-policy given a reference model for finetuning. Across tasks like scientific skill acquisition, mathematical reasoning, and open-ended expertise, our sampling algorithm enables SFT to rival prevailing posttraining techniques, often generalizing better and forgetting less than strong on-policy baselines. In addition, the resulting finetuned models exhibit strong distributional performance and are capable of learning beyond sharpening the base model distribution. At a higher level, our approach presents sampling as a model-native operator that shapes data for learnability, offering broader utility as a general-purpose primitive throughout the posttraining stack.
摘要:引入新能力到前沿模型一直是後訓練的目標,這主要通過監督微調(SFT)和強化學習(RL)來實現。傳統智慧認為,RL 能夠在不失去現有能力的情況下,對新任務進行強泛化,而 SFT 則容易導致弱泛化和災難性遺忘。與此同時,SFT 可以從離策略專家數據中學習,而 RL 必須依賴模型的能力來通過重複採樣找到成功的軌跡。在我們的工作中,我們尋求利用在線學習的優勢,同時利用離策略數據中包含的特權信息。然而,我們並不是修改學習目標以適應這些數據,而是調整數據分佈以更好地適應學習者。我們引入了一種馬爾可夫鏈蒙特卡羅(MCMC)採樣算法,該算法逐步將離策略痕跡轉變為更符合在線策略的形式,前提是有一個參考模型進行微調。在科學技能獲得、數學推理和開放式專業知識等任務中,我們的採樣算法使得 SFT 能夠與當前的後訓練技術相抗衡,通常在泛化能力上表現更好,且遺忘程度較低於強在線基準。此外,最終微調的模型展現出強大的分佈性能,並能夠學習超越僅僅是加強基礎模型分佈。在更高的層面上,我們的方法將採樣呈現為一種模型原生操作符,為學習性塑造數據,並在整個後訓練堆棧中提供更廣泛的通用性。
MIRTO: a registration-gated, multiverse-tested evaluation protocol for unsupervised anomaly segmentation in brain MRI
2610.02136v1 by Negin Kafee Hernashki, Soumick Chatterjee
Unsupervised anomaly detection (UAD) methods for brain MRI are ranked by a single score, yet that score rests on choices that are rarely reported: how each anomaly map is aligned with the reference, how and on which data the threshold is set, and which false-positive budget, metric, aggregation and lesion definition are used. We present MIRTO, an evaluation protocol that makes these choices explicit and measures their effect. It gates the geometry of every comparison with a registration check and label-free diagnostics of known power, sets thresholds on validation data alone and reports the false-positive volume actually realised on test, repeats each comparison over 15,552 defensible evaluation pipelines, and attaches paired subject-bootstrap intervals with multiplicity control. Applied to four UAD methods trained on the same healthy data and tested on 312 BraTS 2020 subjects, MIRTO showed that an axis-order mismatch between stored maps and the reference lowered a diffusion model's voxel AUROC from 0.873 to 0.583 whilst barely moving its slice-level AUROC. Within each metric, the method explained at least 0.95 of the variance in voxel AUROC and AUPRC and 0.77 in Dice, but only 0.14 in lesion sensitivity, where the lesion definition and hit criterion dominated. A Dice advantage that was significant at validation thresholds vanished at equal realised false-positive burden, and an exact identity attributes it to threshold transfer. A training-free change to REFLECT's latent aggregation raised Dice at equal burden by 0.052. Nine hypotheses were tested against explicit criteria; because the same cohort served to develop the protocol, all inference is exploratory.
摘要:未監督異常檢測(UAD)方法對於腦部 MRI 的排名是基於單一分數,但該分數依賴於鮮少報告的選擇:每個異常圖與參考的對齊方式、如何以及基於哪些數據設置閾值,以及使用哪種假陽性預算、指標、聚合和病變定義。我們提出了 MIRTO,一種評估協議,使這些選擇變得明確並測量其影響。它通過註冊檢查和無標籤診斷已知功率來限制每次比較的幾何,僅在驗證數據上設置閾值,並報告在測試中實際實現的假陽性體積,重複每次比較超過 15,552 條可辯護的評估管道,並附上成對的主體自助間隔及多重性控制。應用於四種基於相同健康數據訓練並在 312 名 BraTS 2020 受試者上測試的 UAD 方法,MIRTO 顯示存儲圖與參考之間的軸序不匹配使擴散模型的體素 AUROC 從 0.873 降至 0.583,同時幾乎不影響其切片級 AUROC。在每個指標中,該方法解釋了至少 0.95 的體素 AUROC 和 AUPRC 的變異,及 0.77 的 Dice,但在病變敏感性中僅為 0.14,病變定義和命中標準主導了這一結果。在驗證閾值下顯著的 Dice 優勢在相等的實現假陽性負擔時消失,並且一個精確的身份將其歸因於閾值轉移。對 REFLECT 的潛在聚合進行無訓練的變更,在相等負擔下將 Dice 提高了 0.052。針對明確標準測試了九個假設;由於相同的隊列用於開發該協議,所有推斷都是探索性的。
Local Support Learning
2610.02126v1 by Assaf Ben-Kish, Akarsh Kumar, James Glass, Raja Giryes
We explore catastrophic forgetting in the context of large pre-trained models. By considering forgetting as a geometric problem in the input space of each weight matrix, we uncover a natural retention objective under which updates produced by gradient-based optimizers are suboptimal. Following this observation, we propose Local Support Learning (LSL), a general-purpose framework that augments gradient-based training for retention of prior capabilities without access to prior data. During a new learning phase, LSL pairs two components with distinct roles: a standard weight adapter, trained as usual to minimize the loss, and a gating function that enables the adapter only on input activations from its own training distribution, making the update local to that distribution. The key challenge is that this gate must route data from all learning phases while training only on data from the current one. We address this with a gate based on a Gaussian Mixture Model (GMM), whose likelihood decays rapidly away from its training data, giving it a natural tendency to stay closed on data from prior phases. We show that this post-training approach can resolve forgetting in LLMs of up to 7 billion parameters, retaining both pretrained and finetuned capabilities across multiple training phases, while being efficient in memory and compute, robust to hyperparameter choice, and showing scaling potential.
摘要:我們探討在大型預訓練模型背景下的災難性遺忘。通過將遺忘視為每個權重矩陣的輸入空間中的幾何問題,我們揭示了一個自然的保留目標,在該目標下,由梯度基優化器產生的更新是次優的。根據這一觀察,我們提出了本地支持學習(Local Support Learning,LSL),這是一個通用框架,增強了基於梯度的訓練,以保留先前的能力,而無需訪問先前數據。在新的學習階段,LSL 配對了兩個具有不同角色的組件:一個標準的權重適配器,像往常一樣訓練以最小化損失,以及一個閘控函數,僅在來自其自身訓練分佈的輸入激活上啟用適配器,使更新局限於該分佈。關鍵挑戰在於,這個閘必須在訓練僅基於當前數據的同時,路由來自所有學習階段的數據。我們通過基於高斯混合模型(Gaussian Mixture Model,GMM)的閘來解決這一問題,其似然在遠離其訓練數據時迅速衰減,這使其對來自先前階段的數據自然傾向於保持關閉。我們展示了這種後訓練方法可以解決高達70億參數的大型語言模型中的遺忘,保留了多個訓練階段中的預訓練和微調能力,同時在內存和計算上高效,對超參數選擇具有穩健性,並顯示出擴展潛力。
Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
2610.02122v1 by Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma
Real-world enterprise data science and analytics workflows require reasoning across dozens of tables, performing statistical analyses, and acting on the results. Established text-to-SQL benchmarks evaluate query generation alone, and audits have found their answer keys frequently wrong. Because real enterprise warehouses are too sensitive to release, these benchmarks are built on public datasets where a business event fits in a single table. We introduce Argo-Bench, an evaluation framework comprising 210 data science and analytics tasks. Drawing on public data, peer-reviewed industry literature, and regulatory filings, we simulate a food delivery platform in New York City at true scale, with 81 million orders in 2024, grounded economics, fraud patterns, and marketplace incentives. We export this world to an ERP warehouse of 235 tables and 7.5 billion rows, modeled on the Oracle E-Business Suite schema. The simulator's ground-truth state is withheld from the warehouse the agent sees, so tasks require reconstructing facts by navigating the warehouse before acting on them. Argo-Bench goes beyond text-to-SQL: the agent files actions such as banning fraudulent accounts, allocating courier incentive budgets, or issuing back pay, and the grader scores each by its consequences in the simulator. Every task has an executable reference solution that demonstrates solvability using only the warehouse. The strongest of 14 frontier and open-weight models scores 95 or higher on only 34.8% of tasks and averages 59.5 points. We hope Argo-Bench drives progress toward agents that understand, navigate, and act within real data environments.
摘要:現實世界的企業數據科學和分析工作流程需要跨越數十個表格進行推理、執行統計分析並根據結果採取行動。現有的文本到SQL基準僅評估查詢生成,而審計發現其答案鍵經常錯誤。由於真正的企業數據倉庫過於敏感而無法釋放,這些基準是基於公共數據集構建的,其中商業事件適合於單個表格。我們介紹了Argo-Bench,一個包含210個數據科學和分析任務的評估框架。基於公共數據、同行評審的行業文獻和監管文件,我們模擬了一個在紐約市的食品配送平台,真實規模為2024年的8100萬個訂單,並考慮了經濟基礎、詐騙模式和市場激勵。我們將這個世界導出到一個擁有235個表格和75億行的ERP數據倉庫,該倉庫基於Oracle E-Business Suite架構進行建模。模擬器的真實狀態對代理所見的數據倉庫是保密的,因此任務需要通過導航數據倉庫來重建事實,然後再對其進行操作。Argo-Bench超越了文本到SQL:代理執行的行動包括禁止詐騙帳戶、分配快遞員激勵預算或發放補發工資,評分者根據模擬器中的結果對每個行動進行評分。每個任務都有一個可執行的參考解決方案,該解決方案僅使用數據倉庫來展示可解性。14個前沿和開放權重模型中最強的模型在僅34.8%的任務上得分95或更高,平均得分為59.5分。我們希望Argo-Bench能推動代理理解、導航並在真實數據環境中行動的進展。
Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes
2610.02117v1 by Sophia Sirko-Galouchenko, Monika Wysoczanska, Andrei Bursuc, Nicolas Thome, Spyros Gidaris
On-policy self-distillation has recently emerged as an effective approach for improving language-model reasoning by supervising students with a frozen or EMA version of themselves that receives privileged information. Its application to multimodal large language models (MLLMs), however, remains largely unexplored. Recent approaches use privileged visual information, such as image crops corresponding to a question, to improve fine-grained perception, but their gains are confined to tasks that benefit from such visual zooming and require either human-annotated grounding data or external teacher models. We introduce a different form of on-policy self-distillation for MLLMs that provides the teacher with textual, spatially grounded guidance identifying the visual elements relevant to a query. We use procedurally generated scenes with automatically available object identities and spatial coordinates, enabling scalable and annotation-free post-training. The teacher uses this spatial guidance to locate and integrate evidence from multiple relevant image regions, while the student learns to reproduce the resulting behavior from the image and question alone. Our approach consistently improves performance on counting, document and chart understanding benchmarks across multiple models. Importantly, although post-training uses only synthetic scenes, the resulting improvements transfer to real-world perception benchmarks, yielding a 3.23-point gain in average performance across CVBench, V*, ZoomBench, BLINK, HR-Bench, and MME-RealWorld. These results show that spatially grounded privileged information can induce broader perceptual capabilities through on-policy self-distillation, enabling substantial synthetic-to-real transfer beyond the task and data distribution used for post-training. Project page: https://github.com/sirkosophia/Where-OPD
摘要:在政策自蒸餾最近出現,作為一種有效的方法來改善語言模型推理,通過用一個凍結或EMA版本的自己來監督學生,該版本接收特權信息。然而,它在多模態大型語言模型(MLLMs)上的應用仍然在很大程度上未被探索。最近的方法使用特權的視覺信息,例如與問題相對應的圖像裁剪,以改善細粒度的感知,但它們的增益僅限於那些受益於這種視覺放大並需要人類標註的基礎數據或外部教師模型的任務。我們為MLLMs引入了一種不同形式的政策自蒸餾,為教師提供文本的、空間上有根據的指導,以識別與查詢相關的視覺元素。我們使用程序生成的場景,具有自動可用的物體身份和空間坐標,實現可擴展且無需標註的後訓練。教師使用這種空間指導來定位並整合來自多個相關圖像區域的證據,而學生則學習僅從圖像和問題中重現結果行為。我們的方法在計數、文檔和圖表理解基準上,始終提高多個模型的性能。重要的是,儘管後訓練僅使用合成場景,但所產生的改進能夠轉移到現實世界的感知基準上,在CVBench、V*、ZoomBench、BLINK、HR-Bench和MME-RealWorld上實現了平均性能提高3.23點。這些結果顯示,空間上有根據的特權信息可以通過政策自蒸餾引發更廣泛的感知能力,使得合成到現實的轉移超越了用於後訓練的任務和數據分佈。項目頁面:https://github.com/sirkosophia/Where-OPD
A Comparative Explainability Framework for DeBERTa-v3 in Zero-Shot Medical Abstract Classification
2610.02116v1 by Javier Diaz Esteban-Herreros, David Muñoz-Valero, Raquel Martínez-España, Jose M. Juarez, Juan Moreno-Garcia
A comparative explainability framework is presented to audit DeBERTa-v3 under zero-shot classification of medical abstracts. The work addresses the disagreement problem in Explainable Artificial Intelligence, where different attribution methods produce divergent explanations for the same input and prediction. A natural language inference engine is implemented over the Medical Abstracts corpus with five enriched hypotheses per diagnostic category and a balanced sample of one thousand texts per class. Five explanation methods are compared: SHAP and LIME as model-agnostic approaches, occlusion and Input x Gradient as deep-learning-specific approaches, and Attention x Gradient as a transformer-specific approach. Explanations are standardized through top-token attribution, and pairwise agreement is quantified using the Jaccard index. High predictive accuracy is achieved across well-defined clinical domains, whereas performance degrades under high semantic ambiguity. Explanatory stability directly mirrors predictive certainty, exhibiting strong convergence in univalent categories and a marked drop under diagnostic uncertainty. Furthermore, qualitative error auditing uncovers three systemic failure mechanisms: lexical hypersensitivity, semantic overlap, and loss of attribution coherence. The results support the combined use of several explanation methods and quantitative agreement metrics when auditing transformer-based models in medical text classification, and suggest prioritizing specific clinical ontologies over broad diagnostic labels.
摘要:比較可解釋性框架被提出以審計 DeBERTa-v3 在醫學摘要的零樣本分類下。這項工作解決了可解釋人工智慧中的不一致問題,即不同的歸因方法對相同的輸入和預測產生不同的解釋。自然語言推理引擎在醫學摘要語料庫上實施,每個診斷類別有五個增強的假設,並且每個類別有一千篇文本的平衡樣本。比較了五種解釋方法:SHAP 和 LIME 作為模型無關的方法,遮蔽和輸入 x 梯度作為深度學習特定的方法,以及注意力 x 梯度作為Transformer特定的方法。通過頂部標記歸因標準化解釋,並使用 Jaccard 指數量化成對一致性。在明確定義的臨床領域中實現了高預測準確性,而在高語義模糊性下性能下降。解釋穩定性直接反映預測確定性,在單值類別中顯示出強烈的收斂,並在診斷不確定性下顯著下降。此外,定性錯誤審計揭示了三種系統性失效機制:詞彙過敏、語義重疊和歸因一致性的喪失。結果支持在醫學文本分類中審計基於Transformer的模型時,結合使用幾種解釋方法和定量一致性指標,並建議優先考慮特定的臨床本體論而非廣泛的診斷標籤。
Scalable, Transferable Meta-network for Data Selection Requires a Different Loss (and Why the Obvious Choice is Problematic)
2610.02092v1 by Zilin Du, Bowen Yang, Boyang Albert Li
Data selection is critical for training large language models on massive and heterogeneous corpora. Meta-learning for Training-data Selection offers a principled alternative to heuristic scoring by learning data weights from a target validation objective, but existing methods face a trade-off between fine-grained valuation and transferability to unseen data. A natural solution is to replace per-sample weights with a selection network. However, we find that directly incorporating such a network into existing MTS objectives leads to unstable optimization and poor generalization, caused by weight suppression and persistent reliance on easy-to-learn features. To address these issues, we propose Transferable Example Scoring and Selection (TESS), a scalable data-selection framework built on a Pointwise Value Matching objective (PVM). Experiments on LLM safety and targeted instruction tuning demonstrate strong transfer across datasets, from subsets to full corpora, and from smaller to larger models.
摘要:資料選擇對於在龐大且異質的語料庫上訓練大型語言模型至關重要。針對訓練數據選擇的元學習提供了一種基於原則的替代方案,通過從目標驗證目標學習數據權重,但現有方法在細緻評價和對未見數據的可轉移性之間面臨權衡。一個自然的解決方案是用選擇網絡來替代每個樣本的權重。然而,我們發現將這樣的網絡直接納入現有的MTS目標會導致不穩定的優化和較差的泛化,這是由於權重抑制和持續依賴易於學習的特徵所造成的。為了解決這些問題,我們提出了可轉移範例評分和選擇(TESS),這是一個基於逐點價值匹配目標(PVM)的可擴展數據選擇框架。在LLM安全性和針對性指令調整的實驗中,顯示出跨數據集的強大轉移能力,從子集到完整語料庫,從較小的模型到較大的模型。
GeoLatent: Geometry-Guided Latent Structuring with Routed Optimization for 3D Reasoning
2610.02091v1 by Yakun Zhu, Yi Bin, Yujuan Ding, Zheng Wang, Pengpeng Zeng, Duo Peng, Jingkuan Song, Heng Tao Shen
Despite progress in vision-language models, 3D spatial reasoning from 2D images remains challenging. Text-based methods describe intermediate geometry with discrete tokens, limiting fidelity for continuous spatial relations. Continuous latents offer richer representations, but a single latent type does not explicitly separate the cues needed across spatial tasks. Decomposed spatial latents address this by representing position, direction, and global geometry separately under geometric supervision. Yet the geometry representation can still collapse toward one dominant direction, and unrestricted attention can leave the latents underused during answer learning. We introduce GeoLatent, combining Common--Residual Geometry Alignment (CR-GEO) with routed optimization to structure the geometry states while promoting latent-mediated answer learning. CR-GEO separates shared from residual teacher geometry; routed optimization jointly trains geometry and language, temporarily directs visual answer learning through the latents, and restores full attention with geometry supervision. In controlled comparisons, CR-GEO raises geometry effective rank from 1.00 to 3.87, while blocking latent readout at the bottleneck lowers direction accuracy from 89.1% to 25.8% on 128 fixed questions. After recovery, the differentiated geometry representation and latent-mediated visual route remain available alongside direct image access. GeoLatent achieves 73.0% on SPAR-Bench and 72.1% on SPBench, outperforming previously reported methods on both.
摘要:儘管在視覺-語言模型方面取得了進展,從2D圖像進行3D空間推理仍然具有挑戰性。基於文本的方法使用離散標記描述中間幾何,限制了對連續空間關係的忠實度。連續潛變量提供了更豐富的表徵,但單一潛變量類型並未明確分離在空間任務中所需的線索。分解的空間潛變量通過在幾何監督下分別表示位置、方向和全局幾何來解決這個問題。然而,幾何表徵仍然可能向一個主導方向收斂,而不受限制的注意力可能在答案學習過程中使潛變量未被充分利用。我們引入了GeoLatent,結合了共同-殘差幾何對齊(CR-GEO)與路由優化,以結構化幾何狀態,同時促進潛變量介導的答案學習。CR-GEO將共享的幾何與殘差教師幾何分開;路由優化共同訓練幾何和語言,暫時通過潛變量引導視覺答案學習,並在幾何監督下恢復完整的注意力。在受控比較中,CR-GEO將幾何有效排名從1.00提高到3.87,而在瓶頸處阻止潛變量讀出則使方向準確率從89.1%降低到25.8%,針對128個固定問題。恢復後,區分的幾何表徵和潛變量介導的視覺路徑仍然可用,並與直接圖像訪問並存。GeoLatent在SPAR-Bench上達到73.0%,在SPBench上達到72.1%,在兩者上均超越了先前報告的方法。
LLM2Jev: LLMs Are Already Jev-Style Decision Models -- When and How to Fine-Tune Them
2610.02076v1 by Yinheng Li, Justin Wagle
Jev-style decision models return categorical probability distributions over predefined options without generating free-form text, enabling software systems to act on their outputs directly. In this work, we investigate the extent to which general-purpose LLMs already possess this capability out of the box, and when fine-tuning is actually necessary. We present LLM2Jev, an architecture-preserving framework that extracts calibrated decisions directly from next-token probabilities over bracketed numeric identifiers. LLM2Jev provides both a training-free inference recipe and a fine-tuning objective that optimizes candidate selection via a tree-factorized listwise loss while anchoring auxiliary predictions to the base model using KL divergence penalties. Evaluating on Qwen3.5-4B and Qwen3-0.6B, we find that modern LLMs are inherently effective decision models: without training, the 4B model matches community Jev-style models built on the same backbone, outperforms letter-logit readouts, supports arbitrary option counts, and natively handles multimodal decisions over images. Fine-tuning provides targeted rather than universal benefits -- substantially improving weaker models and specific tasks (such as many-option intent routing), but offering diminishing returns for strong backbones. Crucially, our KL anchors prevent behavioral degradation in conversational text generation, with LoRA delivering the strongest performance on capable models.
摘要:Jev風格的決策模型返回預定選項的類別概率分佈,而不生成自由格式的文本,使得軟體系統能夠直接根據其輸出進行操作。
在這項工作中,我們調查通用LLM在開箱即用的情況下已經具備這種能力的程度,以及何時實際需要進行微調。
我們提出了LLM2Jev,一個保留架構的框架,直接從括號中的數字標識符的下一個標記概率中提取經過校準的決策。
LLM2Jev提供了一個無需訓練的推理配方和一個微調目標,通過樹狀分解的列表損失優化候選選擇,同時使用KL散度懲罰將輔助預測固定到基礎模型。
在Qwen3.5-4B和Qwen3-0.6B上進行評估,我們發現現代LLM本質上是有效的決策模型:在未經訓練的情況下,4B模型與基於相同骨幹構建的社區Jev風格模型相匹配,超越字母邏輯讀出,支持任意選項數量,並原生處理圖像上的多模態決策。
微調提供的是針對性的而非普遍的好處——顯著改善較弱的模型和特定任務(例如多選意圖路由),但對強大的骨幹則提供遞減的回報。
至關重要的是,我們的KL錨點防止了對話文本生成中的行為退化,LoRA在有能力的模型上提供了最強的性能。
Causal Memory Policy: Making Memory Utility Identifiable by Intervening on Retrieval
2610.02070v1 by Arman Behnam, Binghui Wang
Memory-augmented large language models must decide which memories to retain, and recent systems do so by estimating each memory's effect on task performance. However, these estimates rely entirely on retrieved memories. When a memory is never retrieved, store-level interventions produce identical outcomes, leaving its utility unidentified. This is a retrieval-level positivity violation, invisible to diagnostics that examine only memory operations. We introduce Causal Memory Policy (CMP), a causal framework that restores identification by intervening on retrieval itself, reserving a fixed number of context slots for memories sampled with known propensities. CMP estimates memory utility by self-normalized inverse propensity weighting under a balanced assignment design. We prove the causal factorization of memory utility through retrieval, the unbiasedness and exact variance of the estimator, and the optimal decision rule under irreversible operations. Empirically, identification fails for 54% of required memories on LongMemEval and 67% on LoCoMo, and the failure persists in a deployed memory system. CMP improves discrimination between required and non-required memories from 0.54 to 0.66 AUC. Finally, we show that identified memory utility alone is insufficient for retention decisions: per-query utility reaches 0.78 AUC on the query for which it is estimated, yet no aggregation available to a retention policy predicts a memory's value on unseen queries. Code is available at: https://anonymous.4open.science/r/cmp-release-D0C3/.
摘要:記憶增強的大型語言模型必須決定保留哪些記憶,而最近的系統通過估計每個記憶對任務表現的影響來實現這一點。
然而,這些估計完全依賴於檢索到的記憶。
當一個記憶從未被檢索時,存儲層級的干預會產生相同的結果,使其效用無法識別。
這是一種檢索層級的正向性違反,對於僅檢查記憶操作的診斷來說是不可見的。
我們引入了因果記憶策略(Causal Memory Policy, CMP),這是一個因果框架,通過對檢索本身進行干預來恢復識別,為以已知傾向抽樣的記憶保留固定數量的上下文槽位。
CMP 通過自我正規化的逆傾向加權來估計記憶效用,並在平衡分配設計下進行。
我們證明了通過檢索的記憶效用的因果分解、估計量的無偏性和精確方差,以及在不可逆操作下的最佳決策規則。
實證上,在 LongMemEval 上所需記憶的識別失敗率為 54%,而在 LoCoMo 上則為 67%,而且這種失敗在部署的記憶系統中持續存在。
CMP 將所需記憶和非所需記憶之間的區分從 0.54 提高到 0.66 AUC。
最後,我們顯示僅僅識別的記憶效用對於保留決策是不夠的:每查詢的效用在其估計的查詢上達到 0.78 AUC,然而對於未見查詢,保留策略無法預測記憶的價值。
代碼可在以下網址獲得:https://anonymous.4open.science/r/cmp-release-D0C3/.
External Observers May See More Clearly: Cross-Model Span-Level Hallucination Detection in Large Language Models via Hidden State Probing
2610.02066v1 by Kingshuk Gupta, Davide Buscaldi
As Large Language Models (LLMs) increasingly serve as foundational reasoning engines, their tendency to hallucinate remains a critical vulnerability. While recent internal state probes offer a promising alternative to slow external retrieval systems, they largely reduce hallucination detection to a token-wise binary classification task, failing to capture the structured, sequential boundaries of semantic drift. Here, we introduce an internal hidden state framework for fine-grained, span-level hallucination detection. By inspecting layer-wise activation patterns, we attempt to detect the exact hallucination onset and continuation tokens in an LLM generation. Our experiments show that this approach successfully isolates hallucination onsets, achieving substantial improvements in Precision-Recall AUC over random baselines despite extreme class imbalance. Ultimately, we propose a novel cross-model detection framework in which one model observes the internal representations elicited by another model's generation. We find that an external observer can match or exceed a generator's self-detection of its own hallucination onsets, including when the observer is the smaller model, suggesting that self-detection is not the ceiling for onset localisation.
摘要:隨著大型語言模型(LLMs)越來越多地作為基礎推理引擎,它們的幻覺傾向仍然是一個關鍵的脆弱性。
雖然最近的內部狀態探測提供了一種有前景的替代方案來取代緩慢的外部檢索系統,但它們在很大程度上將幻覺檢測簡化為一個逐字的二元分類任務,未能捕捉到語義漂移的結構性、序列性邊界。
在此,我們介紹了一個內部隱藏狀態框架,用於細粒度的跨度級幻覺檢測。
通過檢查層級激活模式,我們試圖檢測LLM生成中的確切幻覺開始和持續標記。
我們的實驗表明,這種方法成功地隔離了幻覺的開始,儘管類別極度不平衡,仍在精確度-召回率AUC上實現了相對隨機基準的顯著改善。
最終,我們提出了一個新穎的跨模型檢測框架,其中一個模型觀察另一個模型生成所引發的內部表示。
我們發現,外部觀察者可以匹配或超越生成器對其自身幻覺開始的自我檢測,包括當觀察者是較小的模型時,這表明自我檢測並不是開始定位的上限。
HydroJEV: A one-second, training-free screen for cyber-attack and fault attribution in water distribution networks
2610.02048v1 by Tianwei Mu, Shengyan Jiang, Mingzhe Yuan, Qing Luo, Min Xiao, Wenhong Wang, Jun Li, Manhong Huang
When a SCADA alarm is raised in a water distribution network, operators must decide quickly whether it reflects a cyberattack, a physical fault, a normal transient or a faulty sensor. Supervised classifiers need labelled incidents that utilities rarely have, and frontier large language models (LLMs) take tens of seconds per decision. We tested whether Jev, a training-free model that returns class probabilities in about one second, can serve as the first tier of this triage. On a four-class cause-attribution benchmark built on the C-Town network in EPANET, Jev was compared with a hand-written rule tree, a supervised classifier and seven cloud LLMs on identical evidence in four sealed, pre-registered rounds. With only a label-free prior correction, Jev matched the rule tree (macro-F1 0.62-0.64 against 0.56-0.61 in distribution) and exceeded the supervised classifier by 0.36-0.42 on event subtypes absent from its labels, in all four rounds, and it outperformed the classifier whenever fewer than about four labelled events per class were available. Jev also decided 20-40 times faster than frontier LLMs. Accepting only benign Jev verdicts confirmed by the rule tree spared an LLM reviewer 35-38% of windows on fresh sealed sets without loss of macro-F1. Transferred unchanged to two further networks, this gated cascade stayed within the non-inferiority margin of its reviewer on all four sets. A fast, training-free screen can therefore take over about a third of the review load in SCADA anomaly triage while preserving the accuracy of deliberate review.
摘要:當水分配網絡中發生 SCADA 警報時,操作員必須迅速決定這是否反映了網絡攻擊、物理故障、正常瞬態或故障傳感器。監督式分類器需要標記的事件,而公用事業公司很少擁有這些事件,前沿的大型語言模型 (LLMs) 每次決策需要幾十秒。我們測試了 Jev,這是一個無需訓練的模型,能在約一秒內返回類別概率,是否可以作為這一分診的第一層。在基於 EPANET 的 C-Town 網絡構建的四類原因歸因基準上,Jev 與手寫規則樹、監督式分類器和七個雲端 LLM 在四輪相同證據的封閉、預註冊回合中進行了比較。僅通過無標籤的先驗修正,Jev 在四輪中與規則樹相匹配(宏 F1 0.62-0.64 對比 0.56-0.61 的分佈),並在缺少其標籤的事件子類型上超過了監督式分類器 0.36-0.42,並且每當每類可用標記事件少於約四個時,它都超越了分類器。Jev 的決策速度也比前沿 LLM 快 20-40 倍。僅接受由規則樹確認的良性 Jev 判決,讓 LLM 審核者在新封閉集上節省了 35-38% 的窗口,而不損失宏 F1。這一閘道級聯在另外兩個網絡上未經改變地轉移,並在所有四組中保持在其審核者的非劣性邊際內。因此,快速、無需訓練的篩選可以接管 SCADA 異常分診中約三分之一的審核負擔,同時保持仔細審核的準確性。
Typological Alignment of Stack-Based Language Models on Mildly Context-Sensitive Artificial Languages
2610.02040v1 by Nadine El-Naggar, Tatsuki Kuribayashi, Ted Briscoe
Some properties of languages, e.g., subject-object-verb (SOV) word order, are more prevalent than others among the thousands of attested natural languages (NLs). Such typological commonality is often attributed to learning biases. Computational simulations, recently with language models (LMs), have facilitated the exploration of this theory. In this paper, we extend existing analyses of the relationship between LMs' learning biases and typological commonality on both data and model sides, focusing on: (i) cross-serial dependencies, the upper limit of attested syntactic complexity, and (ii) stack-based LMs (SLMs), potentially facilitating learning of hierarchical patterns. We first evaluate generalization of SLMs on cross-serial dependencies across diverse artificial languages and confirm that they struggle with such constructions. However, SLMs with limited working memory generalize better suggesting a possible basis for such inductive bias and thus the typological commonality of some word order configurations.
摘要:某些語言的特性,例如主詞-受詞-動詞(SOV)語序,在數千種已證實的自然語言(NLs)中比其他特性更為普遍。這種類型學的共性通常被歸因於學習偏差。計算模擬,最近使用語言模型(LMs),促進了對這一理論的探索。在本文中,我們擴展了現有的分析,研究 LMs 的學習偏差與類型學共性之間的關係,重點關注:(i)交叉序列依賴,已證實的句法複雜性的上限,以及(ii)基於堆疊的 LMs(SLMs),可能促進對層次模式的學習。我們首先評估 SLMs 在各種人工語言中的交叉序列依賴的概括能力,並確認它們在這類結構上存在困難。然而,具有有限工作記憶的 SLMs 的概括能力更佳,這暗示了這種歸納偏差的可能基礎,從而解釋某些語序配置的類型學共性。
CARM: Cancellation-Aware Response Masking for LLM Reinforcement Learning
2610.02039v1 by Yafei Zhang, Songshuo Lu, Sicong Liao, Zhi Chen, Yaohua Tang
Recent years have witnessed the rapid adoption of reinforcement learning (RL) in large language model (LLM) post-training, with substantial gains in mathematical reasoning and code generation. In practical systems, however, policy updates and differences between rollout and training engines can make sampled responses off-policy. Sequence-level masking addresses this mismatch by deciding whether an entire response should contribute to optimization. A common masking rule uses the length-normalized geometric mean of sampled token probability ratios. Its signed log-ratios can cancel across positions, concealing substantial bidirectional policy drift. We propose \emph{Cancellation-Aware Response Masking} (CARM), a sequence-level mask that takes the absolute value of each token log-ratio before averaging, preventing opposing probability changes from canceling. We prove that accepted responses satisfy a joint bound on the fraction of sampled-token ratios outside a prescribed band and their mean log-distance beyond its boundaries. Experiments on mathematical reasoning and code generation show that CARM improves mean@16 averaged over AIME 2024/2025/2026 and BeyondAIME by up to $3.13$ percentage points over geometric-mean masking, and increases average pass@1 across four code benchmarks by $2.88$ points over the strongest evaluated baseline. These findings support CARM as a theoretically grounded and effective method for response-level off-policy control in LLM reinforcement learning.
摘要:近年來,強化學習 (RL) 在大型語言模型 (LLM) 的後訓練中迅速被採用,並在數學推理和程式碼生成方面取得了顯著的進展。
然而,在實際系統中,政策更新和回滾與訓練引擎之間的差異可能使得抽樣的回應偏離政策。
序列級的遮罩通過決定整個回應是否應該參與優化來解決這種不匹配。
一個常見的遮罩規則使用抽樣的標記概率比的長度正規化幾何平均數。
其簽名對數比可以在位置之間相互抵消,隱藏了實質性的雙向政策漂移。
我們提出了 \emph{Cancellation-Aware Response Masking} (CARM),這是一種序列級遮罩,在平均之前取每個標記對數比的絕對值,防止相對的概率變化相互抵消。
我們證明了接受的回應滿足在規定範圍外的抽樣標記比的比例和其均值對數距離的聯合界限。
在數學推理和程式碼生成的實驗中顯示,CARM 在 AIME 2024/2025/2026 和 BeyondAIME 的 mean@16 上比幾何平均遮罩提高了最多 $3.13$ 個百分點,並且在四個程式碼基準上平均 pass@1 比最強的評估基準提高了 $2.88$ 分。
這些發現支持 CARM 作為一種理論基礎和有效的方法,用於 LLM 強化學習中的回應級偏政策控制。
Mimir: Physics-Grounded LLM Agents for Long-Horizon Irrigation Control
2610.02038v1 by Yimeng Liu, Mi Zhang, Younsuk Dong, Zhichao Cao
Large language model (LLM) agents increasingly combine reasoning, tool use, and action, but most evidence comes from episodic tasks with relatively immediate feedback and reset failures. Long-running physical control operates in a different regime: actions alter future states, errors compound across decisions, and an agent must improve from experience without being allowed to rewrite the physical rules that make execution safe. We study this regime through irrigation, where daily decisions interact with soil-water dynamics over entire growing seasons. We present Mimir, a physics-grounded LLM agent organized around two repair timescales. At the fast timescale, a structured physical interface and deterministic simulator turn an LLM output into a proposal that we numerically check, revise, and subject to bounded deterministic action selection before execution. At the slow timescale, recurrent failure patterns are consolidated into persistent contextual principles that condition future proposals, while the physical model, evaluator, and execution constraints remain immutable. Under a common retrospective evaluator across multiple sites, crops, and years, Mimir attains the lowest reported aggregate control cost among the evaluated references and uses about 51% less irrigation than the historical schedule replay. The ablation study show higher control cost when forward simulation, verified revision, or persistent context is removed; model-scale and model-family studies show no monotonic gain from increasing LLM size. The resulting lesson show that persistent physical agents can combine semantic reasoning with bounded, evidence-driven self-improvement while reserving physical truth and actuator authority for explicit numerical mechanisms.
摘要:大型語言模型(LLM)代理越來越多地結合推理、工具使用和行動,但大多數證據來自於具有相對即時反饋和重置失敗的情境任務。長期運行的物理控制運作在不同的範疇:行動改變未來狀態,錯誤在決策中累積,代理必須在不被允許重寫使執行安全的物理規則的情況下從經驗中改進。我們通過灌溉來研究這個範疇,日常決策與整個生長季節的土壤-水動力學相互作用。我們提出了Mimir,一個基於物理的LLM代理,圍繞兩個修復時間尺度組織。在快速時間尺度上,結構化的物理介面和確定性模擬器將LLM輸出轉換為提案,我們對其進行數值檢查、修訂,並在執行之前進行有界的確定性行動選擇。在慢速時間尺度上,重複的失敗模式被整合成持久的上下文原則,這些原則調節未來的提案,而物理模型、評估者和執行約束保持不變。在多個地點、作物和年份的共同回顧評估者下,Mimir在評估的參考中達到了最低報告的總體控制成本,並使用了比歷史計劃重播少約51%的灌溉量。消融研究顯示,當移除前向模擬、驗證修訂或持久上下文時,控制成本會提高;模型規模和模型家族研究顯示,增加LLM大小並未產生單調增益。最終的教訓顯示,持久的物理代理可以將語義推理與有界的、基於證據的自我改進結合,同時為明確的數值機制保留物理真理和執行機構的權威。
Global Coherence: When Every Agent Is Right and the Team Is Still Wrong - A Local-to-Global Semantic Foundation for Multi-Agent Collaboration
2610.02036v1 by Xin Heng
AI agents can each make locally valid decisions yet jointly produce an invalid result. We call this the global coherence problem: a failure of shared state, not merely of model intelligence. Our Observation-Aliasing Impossibility Theorem gives the exact boundary. A policy can guarantee a valid action exactly when all worlds producing the same observation share an admissible action. If k indistinguishable worlds require pairwise-disjoint actions, the best randomized worst-case success is 1/k; more reasoning, roles, messages, or samples cannot recover the missing distinction. A stronger model can reason better within its context, but it cannot see beyond it. We then give local-to-global runtime semantics X = (H, C, G, F; D): topology H records overlapping scopes; category C governs state-changing actions; groupoid G retains reversible translations; sheaf F tests whether local views glue into one world; and minimal history D keeps only distinctions that alter legal futures. Models propose; the harness owns shared state and governs commit. Nine studies test both the failure and its boundary. On a controlled revision benchmark, the same frontier model scores 40/40 when the deciding event is visible; when it is hidden, tested arms score 12--17/40, consistent with chance (1/3); restoring one authoritative fact returns 40/40. On TeamBench, ordinary teams exceed a shared budget in 5/5 runs, a visible live count leaves 4/5 violations, and commit enforcement leaves 0/5. In tau2-bench Telecom, current-state checks score 0.07 after silent reverts, while the harness scores 1.00. Where a conventional solver already owns the complete relevant state, it ties the harness as predicted. The counterintuitive conclusion is that local intelligence cannot substitute for missing global state.
摘要:AI 代理可以各自做出局部有效的決策,但共同產生無效的結果。我們稱這為全球一致性問題:共享狀態的失敗,而不僅僅是模型智能的失敗。
我們的觀察-混淆不可能定理給出了確切的邊界。當所有產生相同觀察的世界共享一個可接受的行動時,政策可以保證一個有效的行動。如果 k 個不可區分的世界需要成對不相交的行動,最佳隨機最壞情況成功率為 1/k;更多的推理、角色、消息或樣本無法恢復缺失的區別。一個更強的模型可以在其上下文中進行更好的推理,但它無法超越這一點。
然後我們給出局部到全球的運行時語義 X = (H, C, G, F; D):拓撲 H 記錄重疊的範疇;類別 C 管理狀態改變的行動;群體 G 保留可逆的翻譯;束 F 測試局部視圖是否粘合成一個世界;而最小歷史 D 僅保留改變合法未來的區別。模型提出;鞍具擁有共享狀態並管理提交。
九項研究測試了失敗及其邊界。在一個受控的修訂基準上,當決策事件可見時,相同的邊界模型得分 40/40;當它被隱藏時,測試的臂得分 12--17/40,與機會相符 (1/3);恢復一個權威事實返回 40/40。在 TeamBench 上,普通團隊在 5/5 次運行中超過了共享預算,一個可見的實時計數留下 4/5 次違規,而提交執行留下 0/5。在 tau2-bench Telecom 中,當前狀態檢查在靜默回退後得分 0.07,而鞍具得分 1.00。當一個傳統求解器已經擁有完整的相關狀態時,它的表現與預測一致。反直覺的結論是,局部智能無法替代缺失的全球狀態。
SPHERE: Adaptive VR Indoor Scene Generation via LLM-Enhanced Spatial Preference Learning and Human-in-the-Loop RL
2610.02023v1 by Hyeonmin Lee, Zheng Wei, Kyungmin Kwon, Jumin Seo, Jiwon Park, Hayoung Oh
While Large Language Models (LLMs) advance 3D indoor scene synthesis, current pipelines fail to retain user-specific preferences across sessions, making immersive authoring a repetitive and physically fatiguing process. We present SPHERE, an adaptive VR generation framework that transforms isolated synthesis into continuous human-AI co-creation. SPHERE extracts persistent spatial preferences from natural multimodal interactions (speech and controller edits). To ensure geometric resilience against spatial distortions, it abstracts these raw edits into hierarchical constraints modeling both local functional and global topological contexts. Furthermore, a human-in-the-loop reinforcement learning mechanism dynamically updates retrieval policies based on the user's final edited scenes. A mixed-design user study ($N=42$) and an offline ablation demonstrate that SPHERE significantly reduces corrective edits and physical demand, preventing bias toward shallow object-level traits to yield geometrically resilient, profile-aligned layouts. Ultimately, SPHERE demonstrates how capturing demonstrated spatial logic enables controlled spatial adaptation, establishing a reliable, governed human-AI collaboration framework for immersive authoring. Project page and source code will be available at: https://github.com/hyeonmin11/SPHERE
摘要:大型語言模型(LLMs)在三維室內場景合成方面取得了進展,但當前的流程無法在不同會話中保留用戶特定的偏好,使得沉浸式創作成為一個重複且身體疲憊的過程。
我們提出了SPHERE,一個自適應虛擬現實生成框架,將孤立的合成轉變為持續的人機協作創作。
SPHERE從自然的多模態互動(語音和控制器編輯)中提取持久的空間偏好。
為了確保對空間扭曲的幾何韌性,它將這些原始編輯抽象為層次約束,建模本地功能和全局拓撲上下文。
此外,一個人機交互的強化學習機制根據用戶最終編輯的場景動態更新檢索策略。
一項混合設計的用戶研究($N=42$)和一個離線消融實驗顯示,SPHERE顯著減少了修正編輯和身體需求,防止對淺層物體特徵的偏見,以產生幾何上韌性、與用戶檔案對齊的佈局。
最終,SPHERE展示了如何捕捉所展示的空間邏輯實現受控的空間適應,建立了一個可靠的、有規範的人機協作框架以進行沉浸式創作。
項目頁面和源代碼將在以下鏈接提供:https://github.com/hyeonmin11/SPHERE
Old Ideas, Novel Problems: The Instability of LLM-Based Novelty Evaluation
2610.02022v1 by Noy Sternlicht, Simra Shahid, Peter Jansen, Daniel S. Weld, Pao Siangliulue, Tom Hope
Automated ideation systems are often evaluated on the novelty of the ideas they produce, and that judgment is increasingly delegated to large language models. Such judges are typically built ad hoc and validated, if at all, on human-authored papers rather than on the generated ideas they are meant to score. So, how do novelty judges perform? Not well. We present a systematic controlled study of novelty evaluation design choices. We first build an evaluation set automatically, mining OpenReview for passages where reviewers explicitly affirm or dispute a paper's originality and keeping only submissions with unanimous agreement at the extremes of their research area; we pair these with ideas from a vanilla LLM generator. Across six judges, we find that small prompt design choices have large consequences; e.g., simply telling the judge that reviewers found one idea novel and the other not can change its verdict on more than half of the identical idea pairs it is shown, shifting pairwise accuracy by over 50 points and occasionally pushing it below chance. The same change helps one judge and hurts another. Retrieval and larger reasoning budgets help little, and two purpose-built novelty evaluators are outperformed by our cheapest prompted baseline. These results raise questions about reported novelty gains of automated ideation systems, and call for robust novelty evaluation methods.
摘要:自動化構思系統通常根據其產出的創新性來進行評估,而這一判斷越來越多地委託給大型語言模型。這些評判者通常是臨時構建的,並且如果有的話,通常是在人工撰寫的論文上進行驗證,而不是在它們所要評分的生成想法上。因此,創新性評判者的表現如何?
表現不佳。我們呈現了一項系統性的控制研究,探討創新性評估設計選擇。我們首先自動構建一個評估集,從 OpenReview 中挖掘出評論者明確肯定或質疑論文原創性的段落,並僅保留在其研究領域極端處獲得一致同意的提交;我們將這些與來自普通 LLM 生成器的想法配對。在六位評判者中,我們發現小的提示設計選擇會產生大的後果;例如,僅僅告訴評判者評論者認為一個想法是新穎的而另一個不是,就可以改變其對超過一半相同想法對的判決,將成對準確率改變超過 50 點,有時甚至將其推至低於隨機機率。同樣的變化對一位評判者有幫助,卻對另一位評判者造成傷害。檢索和更大的推理預算幫助不大,而兩個專門構建的創新性評估者的表現不及我們最便宜的提示基準。這些結果引發了對自動化構思系統報告的創新性增益的質疑,並呼籲建立穩健的創新性評估方法。
Task-Adaptive Grounded 3D-Programmers Using 2D VLMs
2610.02021v1 by Arman Raayatsanati, Sombit Dey, Anna-Maria Halacheva, Jan-Nico Zaech, Luc Van Gool, Danda Pani Paudel
Recent vision-language models (VLMs) exhibit remarkable generalization and reasoning abilities, yet 3D understanding in these models is limited by data scale, training diversity, and reasoning capacity. Instead of naively extending these models into 3D, we take a different approach: we enable powerful 2D VLMs to operate reliably in 3D by introducing 3D grounding and iterative feedback loops with two novel concepts: Canonical Coordinate Framing (CCF) and Task-Adaptive Feedback (TAF). CCF serves as a unified visual representation that anchors both inputs and outputs to a shared Euclidean coordinate system, solving common challenges in 3D grounding such as axis ambiguity, inconsistent metric scale, and floating references. Complementary to this structured framing of the 3D inputs, TAF closes the reasoning loop with task-adaptive dynamic feedback that enables 2D VLMs to perform varied open-vocabulary tasks within their native visual context. Building on this foundation, we introduce 3D-Prog, a 3D understanding, reasoning, and generation framework that jointly employs the capabilities of CCF and TAF together with powerful VLMs. Without requiring any retraining, 3D-Prog performs open-vocabulary 3D understanding, manipulation, and generation across both object-level and scene-level tasks. Our experiments show that the joint use of CCF and TAF transforms 2D VLMs into geometry-aware 3D programmers, achieving consistent, interpretable, and high-quality results across diverse 3D tasks.
摘要:最近的視覺-語言模型(VLMs)展現出卓越的泛化和推理能力,但這些模型在3D理解方面受到數據規模、訓練多樣性和推理能力的限制。與其天真地將這些模型擴展到3D,我們採取了不同的方法:我們通過引入3D基礎和迭代反饋循環,使強大的2D VLMs能夠在3D中可靠運作,並提出了兩個新概念:典範坐標框架(CCF)和任務自適應反饋(TAF)。CCF作為一種統一的視覺表示,將輸入和輸出固定在共享的歐幾里得坐標系中,解決了3D基礎中常見的挑戰,如軸模糊、不一致的度量尺度和浮動參考。TAF則補充了這種3D輸入的結構化框架,通過任務自適應的動態反饋關閉推理循環,使2D VLMs能夠在其本土視覺上下文中執行各種開放詞彙任務。
在這一基礎上,我們介紹了3D-Prog,一個3D理解、推理和生成框架,該框架共同利用CCF和TAF的能力以及強大的VLMs。3D-Prog在不需要任何重新訓練的情況下,能夠在物體級和場景級任務中執行開放詞彙的3D理解、操作和生成。我們的實驗表明,CCF和TAF的聯合使用將2D VLMs轉變為幾何感知的3D程序員,在各種3D任務中實現一致、可解釋和高質量的結果。
Controllable Multi-label Video Safety Detection via Adaptive Tversky Policy Optimization
2610.02019v1 by Guangyu Yang, Jingbiao Mei, Mingsheng Sun, Jinghong Chen, Yingtong Bu, Pengda Qin, Da Chen, Bill Byrne
The rapid growth of video-based social media has increased users' exposure to harmful content, creating a need for reliable automated video safety detection. Although recent Vision-Language Models (VLMs) show strong video understanding capabilities, existing harmful video detection systems face two key limitations: they typically reduce safety detection to binary classification, overlooking the inherently multi-label nature of unsafe videos, and they rely on static training objectives that do not support controllable precision-recall trade-offs, though the desired operating point may vary across moderation pipelines and unsafe categories. To address these gaps, we propose Adaptive Tversky Policy Optimization (ATPO), a reinforcement learning framework for Multi-label Video Safety Detection (Multi-VSD). ATPO introduces the Adaptive Tversky Reward (ATR), which dynamically adjusts false-positive and false-negative penalties during training to enable controllable precision-recall trade-offs. Experiments on SafeWatch-Bench and XD-Violence show that ATPO substantially improves multi-label performance, increasing the Jaccard Index from 40.66 to 75.44 on SafeWatch-Bench-Real. Moreover, ATR enables reliable steering of the precision-recall operating point, supporting deployment scenarios with heterogeneous policy requirements. Code and checkpoints are provided at https://bruceyg.github.io/ATPO-project-page/ .
摘要:視頻社交媒體的快速增長增加了用戶接觸有害內容的機會,這創造了對可靠自動視頻安全檢測的需求。儘管最近的視覺-語言模型(VLMs)顯示出強大的視頻理解能力,但現有的有害視頻檢測系統面臨兩個主要限制:它們通常將安全檢測簡化為二元分類,忽視了不安全視頻固有的多標籤特性,並且依賴於靜態訓練目標,這些目標不支持可控的精確度-召回率權衡,儘管所需的操作點可能在不同的審核管道和不安全類別之間變化。為了解決這些問題,我們提出了自適應Tversky政策優化(ATPO),這是一個用於多標籤視頻安全檢測(Multi-VSD)的強化學習框架。ATPO引入了自適應Tversky獎勵(ATR),在訓練過程中動態調整假陽性和假陰性懲罰,以實現可控的精確度-召回率權衡。在SafeWatch-Bench和XD-Violence上的實驗顯示,ATPO顯著提高了多標籤性能,將SafeWatch-Bench-Real上的Jaccard指數從40.66提高到75.44。此外,ATR使得精確度-召回率操作點的可靠調整成為可能,支持具有異質政策需求的部署場景。代碼和檢查點可在https://bruceyg.github.io/ATPO-project-page/ 獲得。
On Language Drift during RLVR Post-Training
2610.02015v1 by Michael Sullivan, Alexander Koller
Recent advances in LLM reasoning models---driven primarily by the paradigm of post-training via reinforcement learning with verifiable reward (RLVR)---have enabled them to accomplish impressively complex tasks. However, in parallel with their rising capabilities, LLMs have increasingly displayed signs of language drift in their chains of thought (CoTs): unusual, non-standard, and seemingly nonsensical language use. Although it is well-documented---and can potentially impair CoT monitorability---the causes of language drift are thus far poorly understood. In this paper, we identify the conditions under which language drift occurs: we prove theoretically that RLVR optimization pressure permits unbounded language drift, while supervised fine-tuning does not. We then show empirically that language drift specifically arises during RLVR on novel reasoning tasks---i.e. when the target behavior cannot be drawn out of the base model. Finally, we prove that it is not possible to constrain language drift without constraining expected reward, suggesting that CoT monitorability cannot be improved without harming performance during RLVR post-training at the frontier.
摘要:最近在大型語言模型(LLM)推理模型方面的進展——主要受到可驗證獎勵的強化學習後訓練(RLVR)範式的驅動——使得它們能夠完成令人印象深刻的複雜任務。
然而,隨著它們能力的提升,LLM在其思維鏈(CoTs)中越來越顯示出語言漂移的跡象:不尋常的、非標準的,且似乎毫無意義的語言使用。
儘管這一點已被充分記錄——並且可能會影響CoT的可監控性——語言漂移的原因至今仍然了解不深。
在本文中,我們確定了語言漂移發生的條件:我們理論上證明,RLVR優化壓力允許無界的語言漂移,而監督式微調則不然。
然後,我們實證顯示,語言漂移特別是在RLVR處理新推理任務時出現——即當目標行為無法從基礎模型中引出時。
最後,我們證明,若不限制預期獎勵,就無法約束語言漂移,這表明在RLVR後訓練的前沿中,無法改善CoT的可監控性而不損害性能。
Atoms to Processes: The Role of Artificial Intelligence and Machine Learning in Chemical Engineering
2610.02014v1 by Michael Baldea, Linda J. Broadbelt, Marianthi G. Ierapetritou, Akhilesh Jain, Ankur Kumar, Thomas A. Kwan, Fèlix Llovell, Andrew J. Medford, Ilias Mitrai, Joel Paulson, Junyi Qiao, Matthew P. Rivera, Kirti C. Sahu, Lev Sarkisov, Zachary P. Smith, Calvin Tsay, Ching-Mei Wen, Victor M. Zavala, Huacheng Zhang, Dan Zhao
The rapid maturation of artificial intelligence (AI) and machine learning (ML) has catalyzed a profound shift in how chemical engineering problems are formulated, analyzed, and solved. Advances in computing, data availability, and learning algorithms have enabled AI/ML methods to impact applications spanning atomic-scale simulations, materials and catalyst discovery, transport and thermodynamics, separations, process systems engineering, and industrial operations. This article provides a perspective on recent methodological developments and representative applications, emphasizing how AI/ML tools are being integrated with first-principles models to address challenges of predictive accuracy, data scarcity, extrapolation, interpretability, and model lifecycle management. Across domains, a unifying trend is the move away from purely black-box approaches toward hybrid and physics-informed frameworks that explicitly respect conservation laws, thermodynamic consistency, and known structural constraints. These approaches not only improve robustness and reliability, but also enable meaningful human-AI collaboration by providing information at an appropriate level of abstraction for the task and decision context. We conclude that AI and ML are not replacing the core principles of chemical engineering; rather, they are amplifying them. As the field advances toward increasingly autonomous, adaptive, and sustainable systems, the thoughtful integration of AI/ML with first-principles understanding and domain expertise will be essential to realizing their full potential across both research and industrial practice.
摘要:人工智慧(AI)和機器學習(ML)的快速成熟催化了化學工程問題的公式化、分析和解決方式的深刻變化。計算、數據可用性和學習算法的進步使得AI/ML方法能夠影響從原子級模擬、材料和催化劑發現、傳輸和熱力學、分離、過程系統工程到工業運營的應用。本文提供了對近期方法論發展和代表性應用的觀點,強調AI/ML工具如何與第一性原理模型相結合,以應對預測準確性、數據稀缺性、外推、可解釋性和模型生命周期管理的挑戰。在各個領域,一個統一的趨勢是從純粹的黑箱方法轉向混合和物理知識驅動的框架,這些框架明確遵循守恆法則、熱力學一致性和已知結構約束。這些方法不僅提高了穩健性和可靠性,還通過在適當的抽象層次上提供信息來促進有意義的人機協作,以適應任務和決策背景。我們得出結論,AI和ML並不是取代化學工程的核心原則;相反,它們是在放大這些原則。隨著該領域向越來越自主、自適應和可持續的系統邁進,AI/ML與第一性原理理解和領域專業知識的深思熟慮的整合將對實現其在研究和工業實踐中的全部潛力至關重要。
Counting Moves, Weighing Voices: Bayesian Dialectical Argumentation for Calibrated Multi-LLM Councils under Persistent Adversaries
2610.02005v1 by Ionel Eduard Stan, Paolo Napoletano
A multi-LLM \emph{council} lets several large language models (LLMs) deliberate on a question and return an answer together with a confidence estimate. As these systems become increasingly used for reasoning, that confidence should represent a calibrated \emph{probability of being correct}, and the decision should remain robust when some agents are persistently unreliable. Existing \emph{council aggregation} methods fail on both fronts: their confidence estimates measure decisiveness rather than correctness, and they cannot identify or discount persistently unreliable agents. We introduce Bayesian Dialectical Argumentation (BDA), which treats the council's \emph{typed} moves---who proposed, challenged, or conceded which answer---as observations of a classical annotator model with \emph{per-agent} reliabilities. This formulation recasts multi-agent deliberation as a reliability estimation problem, using the deliberation trace to infer agent reliability under persistent adversarial behavior. By weighting evidence according to inferred agent reliability, BDA yields calibrated posterior probabilities over candidate answers while allowing persistently unreliable agents to be inverted rather than merely outvoted. Across binary and multi-class benchmarks, BDA achieves the best calibration among zero-cost council aggregation methods, requiring no additional LLM calls, and improves robustness under persistent adversarial coalitions while remaining competitive in clean settings.
摘要:一個多LLM \emph{委員會} 讓幾個大型語言模型(LLMs)對一個問題進行討論並一起返回答案及信心估計。隨著這些系統在推理中的使用越來越多,這種信心應該代表一個經過校準的 \emph{正確性概率},而且當某些代理持續不可靠時,決策應保持穩健。現有的 \emph{委員會聚合} 方法在這兩方面都失敗:它們的信心估計衡量的是決斷性而非正確性,並且無法識別或排除持續不可靠的代理。我們引入貝葉斯辯證論證(BDA),將委員會的 \emph{類型化} 行動——誰提出、挑戰或讓步於哪個答案——視為具有 \emph{每個代理} 可靠性的經典標註者模型的觀察。這一表述將多代理討論重新定義為一個可靠性估計問題,利用討論痕跡來推斷在持續對抗行為下的代理可靠性。通過根據推斷的代理可靠性加權證據,BDA 產生對候選答案的經過校準的後驗概率,同時允許持續不可靠的代理被反轉,而不僅僅是被投票淘汰。在二元和多類基準測試中,BDA 在零成本委員會聚合方法中實現了最佳的校準,無需額外的 LLM 調用,並在持續對抗聯盟下提高了穩健性,同時在乾淨的環境中保持競爭力。
Mem++: Non-Destructive Memory for Long-Term Organizational LLM Agents
2610.02002v1 by Ahmad Yehia, Aly O. Abdelkareem, Islam Ahmed, Hesham Omran, Khaled Alashmouny, Christian Claudel, Abduallah Mohamed
Large Language Model (LLM) agents now take part in organizational work, where many authors record decisions across documents over months. Because a revised decision arrives as a new document rather than an edit, answering a question requires knowing which version held at a given time. However, most memory systems compress the record at write time. By distilling each document into facts, notes or graph edges, these methods fix what can be answered before any question is asked. To address this, we propose Mem++, a non-destructive memory framework shifting from write-time distillation to read-time selection. Mem++ stores every document whole with its date and author, and it calls no generative model at write time. At read time, it retrieves only documents dated up to the time a question asks about and fuses lexical and semantic rankings. Unlike systems that overwrite older versions, Mem++ keeps them and leaves the choice to the answering model. Evaluations on the organizational benchmark OrgMemBench demonstrate that Mem++ surpasses the strongest memory system baseline by 8.0 to 13.1 points across two answering models. With gpt-4.1-mini, it also achieves the best overall score, 2.6 points above RAG. In addition, Mem++ achieves the best average LLM-judge score on LoCoMo and ranks second on LongMemEval-S, behind only its entity-graph variant. Code for benchmark evaluation is available at https://github.com/AIDAChip-Inc/mem-plus-plus.
摘要:大型語言模型(LLM)代理現在參與組織工作,許多作者在數月內記錄決策於文件中。因為修訂的決策以新文件的形式出現,而不是編輯,因此回答問題需要知道在特定時間持有的是哪個版本。然而,大多數記憶系統在寫入時會壓縮記錄。通過將每個文件提煉成事實、筆記或圖邊,這些方法在任何問題被提出之前固定了可以回答的內容。為了解決這個問題,我們提出了Mem++,這是一個非破壞性的記憶框架,從寫入時的提煉轉向讀取時的選擇。Mem++ 將每個文件完整地存儲,並附上日期和作者,並且在寫入時不調用任何生成模型。在讀取時,它僅檢索在問題詢問時的日期之前的文件,並融合詞彙和語義排名。與覆蓋舊版本的系統不同,Mem++ 保留它們,並將選擇權留給回答模型。在組織基準測試 OrgMemBench 上的評估顯示,Mem++ 在兩個回答模型中超越了最強記憶系統基線 8.0 到 13.1 分。使用 gpt-4.1-mini 時,它還獲得了最佳整體分數,比 RAG 高出 2.6 分。此外,Mem++ 在 LoCoMo 上獲得了最佳平均 LLM-judge 分數,並在 LongMemEval-S 中排名第二,僅次於其實體圖變體。基準評估的代碼可在 https://github.com/AIDAChip-Inc/mem-plus-plus 獲得。
Mingbird: A Local-First Agent Harness Enabling Small Open Models to Complete Real Tasks
2610.02001v1 by Hao Wang, Ting Huang
Small open-weight models (2-9B) run on ordinary laptops, but under cloud-scale agent harnesses they rarely complete real tasks: tool prefill overflows the context, self-correction diverges, tool demonstrations loop, and tasks are silently abandoned. We present evidence, from a controlled single-machine comparison and one third-party benchmark, that a substantial share of these failures is attributable to the harness rather than the model. We introduce Mingbird, a local-first agent harness for Windows and Ollama whose ten mechanisms compensate point-by-point for small-model failure forms, three of them representative: a byte-level net-zero prefill budget, a finish gate that re-reads the task before accepting completion, and signature-level loop detection. On LRAB, a controlled comparison holding machine, models, budgets, and scoring fixed (4 harnesses $\times$ 4 open models (2B-35B) $\times$ 18 real tasks, deterministic artifact scoring), Mingbird reaches 0.886 overall against 0.631 (goose), 0.479 (opencode), and 0.405 (agent-mini), with all 288 cells published; on $τ^2$-bench (278 tasks, three arms, one protocol) it totals 0.856 against 0.791 and 0.737; and a frontier-model probe on the same 18 tasks spans 0.997 to 0.478 across harnesses, with well-formed scaffolds staying within 0.072 of each other. A leave-one-mechanism-out ablation is reported as directional only: same-night replications of the same arm move its mean by up to 0.069, the size of every nominal single-trial delta, and the one batch-matched comparison (full mechanism stack versus text re-read alone) gives the executable completion guards a paired +0.10 across three replications. The evidence carries stated limits: a self-built benchmark, a single machine, and single-trial scoring.
摘要:小型開放權重模型(2-9B)可以在普通筆記型電腦上運行,但在雲端規模的代理環境下,它們很少能完成實際任務:工具預填超出上下文,自我修正偏離,工具演示循環,任務則被默默放棄。我們提供證據,來自受控的單機比較和一個第三方基準,顯示這些失敗的相當一部分是由於代理環境而非模型本身。我們介紹了 Mingbird,一個針對 Windows 和 Ollama 的本地優先代理環境,其十個機制逐點補償小型模型的失敗形式,其中三個具有代表性:字節級的淨零預填預算、一個在接受完成之前重新閱讀任務的完成門,以及簽名級的循環檢測。在 LRAB 上,進行了受控比較,固定了機器、模型和預算(4 個代理環境 × 4 個開放模型(2B-35B) × 18 個真實任務,確定性工件評分),Mingbird 的整體得分為 0.886,相較於 0.631(goose)、0.479(opencode)和 0.405(agent-mini),所有 288 個單元均已發布;在 $τ^2$-bench(278 個任務、三個臂、一個協議)中,其總得分為 0.856,相較於 0.791 和 0.737;而在同 18 個任務上的前沿模型探測中,跨越的得分範圍為 0.997 到 0.478,各代理環境之間的良好結構保持在 0.072 之內。一個去除一個機制的消融實驗僅報告為方向性:同夜對同一臂的重複實驗使其均值變化最多達 0.069,這是每個名義單次試驗的變化量,而唯一的批次匹配比較(完整機制堆疊與僅文本重讀)在三次重複中給可執行完成保護提供了配對的 +0.10。這些證據有明確的限制:自建基準、單一機器和單次試驗評分。
Can AI Oversight Be Zero Knowledge?
2610.01995v1 by Alessandro Chiesa, Ziyi Guan, Burcu Yildiz
AI systems increasingly produce outputs from confidential data, such as a fitness-for-duty assessment from medical records or the predicted properties of a drug candidate from its secret structure. It is important to verify that such outputs are correct without revealing the underlying data. A recent line of work studies verification of AI outputs via interactive proofs and debate for oracle-aided computation, where correctness may depend on an oracle such as human judgment, a physical experiment, or the web. These works focus on verification by a verifier that runs much faster than the computation. However, such efficient verification is impossible for general oracle-aided computation, and these works therefore rely on additional assumptions. We focus instead on privacy: allowing the verifier to run in time polynomial in the computation, we ask whether interactive arguments for oracle-aided computation can be zero knowledge, so that the verifier learns nothing about the confidential data beyond the correctness of the output. We prove that, in general, they cannot. In the random oracle model, there are no zero-knowledge proofs for all oracle-aided computations, even if both the prover and the verifier are allowed to run much longer than the computation itself. The impossibility extends to debate, a canonical model for scalable oversight. On the positive side, we show that if the oracle attaches a cryptographic signature to each of its answers, then every oracle-aided computation can be verified in zero knowledge with an efficient prover and verifier, assuming only collision-resistant hash functions. Beyond privacy, this also gives an alternative approach to scalable oversight that relies neither on an honest opponent, as in debate, nor on the robustness of the computation, as in prior single-prover protocols.
摘要:AI 系統越來越多地從機密數據中產生輸出,例如從醫療記錄中進行的適任性評估或從其秘密結構中預測的藥物候選物的性質。
驗證這些輸出是否正確而不透露底層數據是很重要的。
最近的一系列研究通過互動證明和辯論研究 AI 輸出的驗證,用於有 oracle 輔助的計算,其中正確性可能依賴於 oracle,例如人類判斷、物理實驗或網絡。
這些研究專注於由運行速度遠快於計算的驗證者進行的驗證。
然而,對於一般的有 oracle 輔助計算,這樣的高效驗證是不可能的,因此這些研究依賴於額外的假設。
我們則專注於隱私:允許驗證者在計算的多項式時間內運行,我們詢問有 oracle 輔助計算的互動論證是否可以是零知識的,以便驗證者不會學到關於機密數據的任何信息,除了輸出的正確性。
我們證明,通常情況下,它們是不可能的。
在隨機 oracle 模型中,對於所有有 oracle 輔助的計算,沒有零知識證明,即使證明者和驗證者都被允許運行的時間遠超過計算本身。
這種不可能性擴展到辯論,這是一個可擴展監督的典型模型。
從積極的一面來看,我們展示了如果 oracle 為其每個答案附加加密簽名,那麼每個有 oracle 輔助的計算都可以在零知識中進行驗證,並且有高效的證明者和驗證者,只假設碰撞抗性哈希函數。
除了隱私之外,這還提供了一種可擴展監督的替代方法,既不依賴於誠實的對手(如辯論中),也不依賴於計算的穩健性(如以前的單證明者協議中)。
Universal Byte-Level Encoding: UTF-8/UTF-16 Routing to Reduce Cross-Script Token-Budget Disparities
2610.01984v1 by Hyunsik Kim, Youngmoon Jung
Byte-level byte-pair encoding (BBPE) tokenizers are attractive for multilingual large language models (LLMs) because they cover all Unicode text. In UTF-8-based BBPE, however, many scripts start from a higher fallback cost than English: when no learned merges can be applied, a multibyte character requires multiple byte-derived symbols. We call this worst-case pre-merge cost the encoding floor. A higher floor can increase token counts and per-request cost and shrink usable context. Changing the text encoding can reduce this gap, but a single global encoding can make already-efficient English spans more expensive in mixed-script text. We propose Universal Byte-Level Encoding (UBE), a dual-alphabet tokenizer that keeps 1-2-byte UTF-8 characters on the UTF-8 path while routing 3-4-byte UTF-8 characters through UTF-16. This lowers the encoding floor for 3-byte Basic Multilingual Plane (BMP) characters in scripts with high token premiums (token counts relative to English) without raising it for already-efficient spans in mixed-script text. UBE changes only the byte representation presented to byte-pair encoding (BPE); the merge rule remains standard, and exact decoding is preserved. UBE also composes with alternative boundary policies and morphology-based representations. In a Unicode 17 audit, UBE exactly round-trips all Unicode scalar values and all inputs in the official normalization, grapheme-break, and emoji test suites. Across intrinsic evaluations, UBE lowers dispersion in English-normalized token-count ratios, reducing cross-lingual token-budget disparity. In multilingual language model (LM) experiments, UBE matches BBPE's LM quality. In the main multilingual settings, UBE reduces token counts most for high-premium scripts and slightly lowers English token counts, yielding more usable context under fixed token budgets and faster prompt processing in content-matched benchmarks.
摘要:Byte-level byte-pair encoding (BBPE) 令牌器對於多語言大型語言模型 (LLMs) 來說非常有吸引力,因為它們涵蓋了所有 Unicode 文本。
然而,在基於 UTF-8 的 BBPE 中,許多字母系統的回退成本高於英語:當無法應用學習到的合併時,多字節字符需要多個字節衍生符號。
我們將這種最壞情況的合併前成本稱為編碼底線。
較高的底線會增加令牌數量和每次請求的成本,並縮小可用上下文。
改變文本編碼可以減少這一差距,但單一的全局編碼可能會使已經高效的英語範圍在混合字母系統文本中變得更昂貴。
我們提出了通用字節級編碼 (UBE),這是一種雙字母表的令牌器,保持 1-2 字節的 UTF-8 字符在 UTF-8 路徑上,同時將 3-4 字節的 UTF-8 字符通過 UTF-16 路由。
這降低了在高令牌溢價(相對於英語的令牌數量)的字母系統中 3 字節基本多語言平面 (BMP) 字符的編碼底線,而不提高在混合字母系統文本中已經高效的範圍的底線。
UBE 只改變呈現給字節對編碼 (BPE) 的字節表示;合併規則保持標準,並保留精確解碼。
UBE 還可以與替代邊界策略和基於形態學的表示進行組合。
在 Unicode 17 的審計中,UBE 精確地回傳所有 Unicode 標量值以及官方標準化、字形斷裂和表情符號測試套件中的所有輸入。
在內部評估中,UBE 降低了英語標準化令牌計數比率的分散性,減少了跨語言的令牌預算差異。
在多語言語言模型 (LM) 實驗中,UBE 的 LM 質量與 BBPE 相匹配。
在主要的多語言設置中,UBE 對高溢價字母系統的令牌數量減少最多,並稍微降低英語的令牌數量,從而在固定令牌預算下產生更多可用上下文,並在內容匹配基準測試中加快提示處理速度。
Counterfactual Auditing of Bias in Open-Source Large Language Models for Clinical Triage
2610.01963v1 by Manar Aljohani, Brandon Ho, Kenneth McKinley, Dennis Ren, Xuan Wang
Emergency department (ED) triage is a high-stakes prioritization task in which demographic, socioeconomic, and system-context information may improperly influence acuity assignment. Although open-source large language models (LLMs) are increasingly considered for local and privacy-preserving clinical decision support, it remains unclear how counterfactual bias varies across model families, sizes, medical-domain models, and domain-adapted models. We present a comparative counterfactual audit of ten open-source LLMs for pediatric Emergency Severity Index (ESI) prediction. Starting from real and handbook-style clinical vignettes, we construct paired counterfactual variants that change only one injected demographic, socioeconomic, healthcare-access, behavioral, social, or system-context variable while holding the clinical presentation fixed. Models include Qwen2.5-7B, Qwen2.5-14B-Instruct, a QLoRA fine-tuned Qwen2.5-7B, MedGemma variants, MedLLaMA2-7B, GPT-OSS-20B, and GPT-OSS-120B. We measure any counterfactual shift, undertriage, overtriage, shifts greater than one ESI level, mean shift, and mean absolute shift. Counterfactual sensitivity varied substantially and did not consistently decrease with larger model size or medical-domain pretraining. The fine-tuned Qwen2.5-7B showed the lowest overall sensitivity, with a 5.27% any-shift rate and mean absolute shift of 0.0534, versus 16.02% and 0.1706 for the base model. Several larger or medical-domain models showed more significant shifts. Stratified and correlation analyses further revealed clinically important directionality and shared failure patterns hidden by aggregate rates. These findings support counterfactual auditing as a lightweight, clinically interpretable framework for comparing fairness risks in open-source LLMs before clinical deployment.
摘要:急診部(ED)分診是一項高風險的優先排序任務,其中人口統計、社會經濟和系統背景信息可能不當影響急性程度的分配。儘管開源大型語言模型(LLMs)越來越被考慮用於本地和隱私保護的臨床決策支持,但目前尚不清楚反事實偏見在不同模型家族、大小、醫療領域模型和領域適應模型之間的變化情況。我們對十個開源LLM進行了針對兒科緊急嚴重性指數(ESI)預測的比較反事實審計。從真實和手冊風格的臨床小插曲開始,我們構建了配對的反事實變體,僅改變一個注入的人口統計、社會經濟、醫療訪問、行為、社會或系統背景變量,同時保持臨床表現不變。模型包括Qwen2.5-7B、Qwen2.5-14B-Instruct、經過QLoRA微調的Qwen2.5-7B、MedGemma變體、MedLLaMA2-7B、GPT-OSS-20B和GPT-OSS-120B。我們測量任何反事實變化、低估分診、過度分診、超過一個ESI級別的變化、平均變化和平均絕對變化。反事實敏感性變化顯著,且不一致地隨著模型大小或醫療領域預訓練的增大而減少。經過微調的Qwen2.5-7B顯示出最低的整體敏感性,任何變化率為5.27%,平均絕對變化為0.0534,而基礎模型則為16.02%和0.1706。幾個較大或醫療領域模型顯示出更顯著的變化。分層和相關分析進一步揭示了臨床上重要的方向性和由聚合率隱藏的共同失敗模式。這些發現支持反事實審計作為一種輕量級、臨床可解釋的框架,用於在臨床部署前比較開源LLM中的公平風險。
A Hybrid Approach to Malware Detection: Integrating Few-Shot Model-Agnostic Meta-Learning with Autoencoders
2610.01949v1 by Emmanuela Andam, Yasir Abbas Zaidi, Abdelali Hadir, Emmanuel Grant, Naima Kaabouch
Ransomware has emerged as a major cybersecurity threat, with incidents increasing in frequency and impact across critical sectors. These attacks are typically launched through phishing emails, malicious downloads, or exploitation of software vulnerabilities to gain system access. Once inside, the malware encrypts files and demands a ransom, often in cryptocurrency, for the decryption key. Conventional detection methods often struggle with novel or scarce samples, leaving systems vulnerable. To address these challenges, this paper proposes a hybrid deep learning framework that combines an Autoencoder Feature Extractor (AFE) with a Model Agnostic Meta Learning (MAML) classifier for few shot malware detection. The AFE generates compact latent features that reduce noise and dimensionality, while the MAML classifier rapidly adapts to new threats using limited labeled data. Experiments conducted on the Ransomware Dataset 2024 demonstrate the effectiveness of the framework in binary classification tasks. Across one to fifty shot settings, the proposed model consistently achieves high accuracy, F1 score, and Matthews Correlation Coefficient values, maintaining reliable classification even under extreme scarcity. These results highlight the model's robustness and effectiveness in adapting to limited data scenarios, demonstrating the potential of combining feature extraction with meta learning to enhance resilience against malware, particularly in sectors such as healthcare, manufacturing, and public infrastructure, where cyberattacks can cause significant operational and financial disruption.
摘要:勒索病毒已成為一個主要的網絡安全威脅,事件在關鍵行業中的頻率和影響不斷增加。這些攻擊通常通過釣魚電子郵件、惡意下載或利用軟件漏洞來獲得系統訪問權限。一旦進入,惡意軟件會加密文件並要求贖金,通常以加密貨幣的形式支付解密密鑰。傳統的檢測方法在面對新穎或稀少的樣本時往往會掙扎,讓系統變得脆弱。為了應對這些挑戰,本文提出了一個混合深度學習框架,結合了自編碼器特徵提取器(AFE)和模型無關的元學習(MAML)分類器,用於少量樣本的惡意軟件檢測。AFE生成緊湊的潛在特徵,減少噪聲和維度,而MAML分類器則利用有限的標記數據快速適應新威脅。在2024年勒索病毒數據集上進行的實驗證明了該框架在二元分類任務中的有效性。在一到五十次樣本設置中,所提出的模型始終能夠實現高準確率、F1分數和馬修斯相關係數值,即使在極度稀缺的情況下也能保持可靠的分類。這些結果突顯了該模型的穩健性和在有限數據場景中適應的有效性,展示了將特徵提取與元學習相結合以增強對抗惡意軟件的韌性的潛力,特別是在醫療、製造和公共基礎設施等行業中,這些行業的網絡攻擊可能會造成重大的運營和財務中斷。
Latent JEPA: Abstract Future Prediction for Latent Reasoning in Chemistry
2610.01947v1 by Xinjian Zhao, Yaoyao Xu, Xuemin Chen, Xiaozhuang Song, Tianshu Yu
Large language models offer a promising foundation for chemical reasoning, bringing together chemical knowledge and multistep problem solving. Chemical intuition can provide an initial sense of plausible outcomes before the details of a solution are fully worked out. Inspired by how such expectations complement explicit analysis, we study how continuous latent thoughts can be trained to anticipate informative aspects of future solutions without verbalizing every intermediate step. We introduce Latent JEPA, a framework that combines autoregressive learning with joint-embedding prediction of one or more future views. For chemical reasoning, we develop textual and molecular prediction objectives that connect latent thoughts to both subsequent reasoning and molecular outcomes. Experiments on ChemCoTBench show gains in molecular optimization and on several editing and reaction metrics. Representation analyses show that future prediction makes latent thoughts more informative about molecular outcomes and strengthens their correspondence with chemical structure. These findings support abstract future prediction as a learning principle for connecting continuous latent reasoning with scientific outcomes.
摘要:大型語言模型為化學推理提供了一個有前景的基礎,將化學知識和多步驟問題解決結合在一起。化學直覺能在解決方案的細節完全展開之前,提供一種合理結果的初步感知。受到這種期望如何補充明確分析的啟發,我們研究如何訓練連續潛在思維,以預測未來解決方案的資訊性方面,而不需要逐步口頭表達每一個中間步驟。我們介紹了潛在JEPA,一個將自回歸學習與一個或多個未來視圖的聯合嵌入預測相結合的框架。針對化學推理,我們開發了文本和分子預測目標,將潛在思維與後續推理和分子結果連接起來。在ChemCoTBench上的實驗顯示,在分子優化以及幾個編輯和反應指標上都有提升。表徵分析顯示,未來預測使潛在思維對分子結果的資訊性更強,並加強了它們與化學結構的對應性。這些發現支持將抽象的未來預測作為一種學習原則,以連接連續的潛在推理與科學結果。
A rubric landscape for evaluating clinical reasoning in large language models: what exists, what is missing, and what needs to be combined
2610.01938v1 by Zhangshu Joshua Jiang, Zina Ibrahim, James T. Teo
Exam-style accuracy does not establish whether large language models (LLMs) reason well over clinical records. We define clinical reasoning as integrating and updating evidence across time and sources to form, revise and justify a patient's problem representation and a defensible plan. This structured narrative review maps three literatures: medical education assessment instruments, clinical LLM benchmarks published from 2023 onwards, and general-domain methods for evaluating long-form generation. We examine six dimensions: problem representation, temporal synthesis, differential and management reasoning, counterfactual reasoning, calibrated uncertainty, and reasoning faithfulness. Preprints are included and flagged. No single instrument covers all six dimensions. Problem representation and differential or management reasoning are reasonably covered, although reliability varies by instrument and setting. TIMER-Eval targets temporal synthesis, and ER-Reason assesses sequential diagnostic belief updating. Dedicated uncertainty and counterfactual evaluations are emerging, but their applicability to longitudinal free-text reasoning remains limited. Factual completeness is well theorised in general-domain evaluation, with early clinical evidence of important omissions. Faithfulness remains the weakest dimension, with one identified clinical causal-ablation study on multiple-choice questions. Existing tools should be combined through binary rubric items, separate completeness and correctness scores, case-specific importance weighting with non-compensable safety caps, temporal order-consistency checks, and chance-corrected reliability reporting. Further design work is needed for calibrated uncertainty, counterfactual reasoning and faithfulness over longitudinal free-text records. This review provides a design rationale, not a validated instrument.
摘要:考試風格的準確性並不能確定大型語言模型(LLMs)在臨床記錄上是否能夠進行良好的推理。我們將臨床推理定義為整合和更新跨時間和來源的證據,以形成、修訂和辯護病人的問題表述及可辯護的計劃。
這篇結構化的敘述性回顧映射了三個文獻領域:醫學教育評估工具、2023年以來發表的臨床LLM基準,以及評估長篇生成的一般領域方法。我們檢視了六個維度:問題表述、時間綜合、差異和管理推理、反事實推理、校準的不確定性,以及推理的忠實性。預印本已被納入並標記。
沒有單一的工具涵蓋所有六個維度。問題表述以及差異或管理推理的覆蓋相對合理,儘管可靠性因工具和環境而異。TIMER-Eval 針對時間綜合,而 ER-Reason 評估連續的診斷信念更新。專門的不確定性和反事實評估正在出現,但它們對於縱向自由文本推理的適用性仍然有限。事實的完整性在一般領域評估中有良好的理論基礎,並且早期臨床證據顯示出重要的遺漏。忠實性仍然是最薄弱的維度,其中有一項針對多選題的臨床因果消融研究被識別。
現有工具應通過二元評分項目、分開的完整性和正確性分數、特定案例的重要性加權(帶有不可補償的安全上限)、時間順序一致性檢查以及機會修正的可靠性報告進行結合。對於校準的不確定性、反事實推理和縱向自由文本記錄的忠實性,還需要進一步的設計工作。這篇回顧提供了一個設計的理由,而不是一個經過驗證的工具。
Mapping the RAG Landscape: A Four Axis Taxonomy of Efficiency, Defense, Interactivity, and Reasoning
2610.01936v1 by Meghana Sunil, Shravya V, Shravan Venkatraman, Joe Dhanith PR
Large Language Models (LLMs) have demonstrated remarkable fluency across many tasks but remain limited by their static, parameter bound knowledge and their susceptibility to hallucinating information. Retrieval Augmented Generation (RAG) addresses these issues by incorporating external retrieval into the generation process, grounding model outputs in verifiable and up to date sources. While prior surveys primarily focus on core RAG architectures and standard pipelines, recent research explores broader challenges and capabilities that extend beyond these foundational designs. This survey provides a consolidated and structured examination of contemporary RAG developments, organizing the field into a four axis taxonomy: improving retrieval efficiency, strengthening robustness and security, supporting user driven and interactive workflows, and enabling multi step or complex reasoning. We formalize key components of the RAG framework and review methods spanning dense and sparse retrieval, fusion strategies, embedding optimizations, and reinforcement learning based retrieval policies, highlighting how these advances influence practical deployment and system design. We also synthesize evaluation practices, domain specific applications, and architectural variants such as Naive, Advanced, and Modular RAG. Finally, we outline persistent challenges related to retrieval quality, reliability, domain adaptation, scalability, and explainability, and identify opportunities for building RAG systems that are more reliable, adaptable, and transparent.
摘要:大型語言模型(LLMs)在許多任務中展現了卓越的流暢性,但仍然受到靜態的、參數限制的知識以及對虛假信息的易感性的限制。檢索增強生成(RAG)通過將外部檢索納入生成過程來解決這些問題,使模型輸出基於可驗證且最新的來源。雖然之前的調查主要集中在核心RAG架構和標準流程上,但最近的研究探討了超越這些基礎設計的更廣泛挑戰和能力。本調查提供了一個當代RAG發展的綜合和結構化檢視,將該領域組織為四個軸向的分類法:提高檢索效率、加強穩健性和安全性、支持用戶驅動和互動工作流程,以及實現多步驟或複雜推理。我們正式化了RAG框架的關鍵組件,並回顧了涵蓋密集和稀疏檢索、融合策略、嵌入優化和強化學習基於檢索政策的方法,強調這些進展如何影響實際部署和系統設計。我們還綜合了評估實踐、特定領域的應用以及如Naive、Advanced和Modular RAG等架構變體。最後,我們概述了與檢索質量、可靠性、領域適應性、可擴展性和可解釋性相關的持續挑戰,並確定了構建更可靠、可適應和透明的RAG系統的機會。
Cross-Lingual Alignment for Decoder-Only Models using MoE Routers
2610.01921v1 by Lucas Bandarkar, Clark Peng, Ahmed Haj Ahmed, Aditi Khandelwal, Nanyun Peng
Cross-lingual contrastive learning has been a core component of multilingual encoder training, but the ability to explicitly align representations is not possible in decoder-only LLMs because of varying multilingual tokenization. However, a growing amount of research suggests that even in LLMs, higher cross-lingual representational alignment leads to improved cross-lingual transfer. In this paper, we propose a novel approach to reimagine cross-lingual contrastive learning given the architectural constraints of modern LLMs. Rather than applying an auxiliary alignment loss on hidden states, we propose using the outputs of the mixture-of-experts (MoE) routers as the target for alignment. Router outputs lend themselves better to pooling over many tokens, enabling more reliable cross-lingual comparisons at the sequence-level. Controlled continual pre-training experiments on four open-source MoEs show that incorporating this routing loss also aligns the underlying hidden representations across languages. Most importantly, this loss improves multilingual performance on our diverse evaluation suite, demonstrating the potential of cross-lingual MoE router alignment.
摘要:跨語言對比學習一直是多語言編碼器訓練的核心組成部分,但在僅使用解碼器的LLM中,由於多語言標記化的差異,無法明確對齊表示。
然而,越來越多的研究表明,即使在LLM中,更高的跨語言表示對齊也能改善跨語言轉移。
在本文中,我們提出了一種新穎的方法,重新構想在現代LLM的架構限制下進行跨語言對比學習。
我們提議不在隱藏狀態上應用輔助對齊損失,而是使用混合專家(MoE)路由器的輸出作為對齊的目標。
路由器的輸出更適合在多個標記上進行池化,從而在序列級別上實現更可靠的跨語言比較。
在四個開源MoE上進行的受控持續預訓練實驗顯示,納入這一路由損失也使得不同語言之間的隱藏表示得以對齊。
最重要的是,這一損失提高了我們多樣化評估套件上的多語言性能,展示了跨語言MoE路由器對齊的潛力。
MoLE: Mixture of Latent Experts for Complementary Visual Reasoning
2610.01917v1 by Yingcheng Liu, Tianyi Jiang, Yujuan Ding, jiangbo Ai, Xun Jiang, Guoqing Wang, Wei Ye, Yi Bin
Latent visual reasoning equips vision--language models with continuous intermediate states that can process visual evidence without explicit textual reasoning traces or repeated image operations. However, existing methods often allow multiple latent tokens to access the same visual evidence through shared value projections, providing no mechanism for them to extract complementary visual information; simply increasing the latent budget can therefore yield redundant latent representations. We argue that effective latent reasoning should encourage different latent tokens to extract complementary visual information, and thereby act as specialized visual experts. Based on this insight, we propose MoLE, a Mixture of Latent Experts framework that controls both what visual evidence each latent visual expert observes and how it transforms that evidence. MoLE isolates latent visual experts during evidence extraction and uses dedicated latent summary experts to aggregate the complementary representations of latent visual experts. A two-stage training pipeline first forces visual evidence through this latent pathway and then restores direct visual access, requiring neither predefined expert roles nor intermediate visual targets. Across five visual reasoning benchmarks, MoLE achieves an average score of 78.6, outperforming data-matched supervised fine-tuning by 4.9 and the strongest evaluated latent visual reasoning baseline at the same latent budget by 3.6. Representation analyses show lower latent-state similarity and more diverse visual attention, while masking the latent pathway reduces average performance by 9.2. These results demonstrate that specializing latent computation is more effective than merely increasing the number of latent tokens.
摘要:潛在視覺推理使視覺-語言模型具備連續的中間狀態,能夠在沒有明確文本推理痕跡或重複圖像操作的情況下處理視覺證據。
然而,現有的方法往往允許多個潛在標記通過共享的值投影訪問相同的視覺證據,這並未提供提取互補視覺信息的機制;因此,僅僅增加潛在預算可能會產生冗餘的潛在表示。
我們認為,有效的潛在推理應該鼓勵不同的潛在標記提取互補的視覺信息,從而充當專門的視覺專家。
基於這一見解,我們提出了MoLE,一種潛在專家混合框架,控制每個潛在視覺專家觀察的視覺證據及其如何轉化這些證據。
MoLE在證據提取過程中隔離潛在視覺專家,並使用專門的潛在摘要專家來聚合潛在視覺專家的互補表示。
一個兩階段的訓練流程首先強迫視覺證據通過這一潛在路徑,然後恢復直接的視覺訪問,既不需要預定義的專家角色,也不需要中間視覺目標。
在五個視覺推理基準上,MoLE的平均得分為78.6,比數據匹配的監督微調高出4.9,比在相同潛在預算下評估的最強潛在視覺推理基線高出3.6。
表示分析顯示潛在狀態相似性較低,視覺注意力更為多樣,而屏蔽潛在路徑則使平均性能降低9.2。
這些結果表明,專門化潛在計算比單純增加潛在標記的數量更有效。
Asynchronous LLM Post-Training: Group-Mass Capping and Convergence Analysis
2610.01896v1 by Qijia He, Ruinan Jin, Jun Luo, Shaofeng Zou, Yingbin Liang
Asynchronous reinforcement learning (RL) improves the efficiency of large language model post-training but introduces stale rollouts generated by earlier policies. Theoretical understanding of how this staleness affects convergence and how to mitigate its impact remains limited. We derive a convergence bound for GRPO-style algorithms that explicitly characterizes the tradeoff between the gradient estimator's second moment and bias. For trajectory-level importance-weighted estimators, our analysis shows that once the second moment is uniformly controlled, delay enters the bound through the bias introduced by clipping or rescaling. Guided by this insight, we propose a novel group mass capping GRPO (GMC-GRPO) method, which minimizes a ratio-based bias bound within a class of weighted estimators sharing a common second-moment guarantee. We establish convergence guarantees for asynchronous GMC-GRPO and show that, compared with TIC-GRPO, it improves the threshold dependence of the fourth-order delay term from $O(ε^{-4})$ to $O(ε^{-2})$ as $ε\to0$, where $1+ε$ is the ratio threshold. Under local policy overlap, the delay-dependent term decreases as $G^{-2/5}$ after tuning the step size, where $G$ is the group size. For fixed behavior and current policies, the bias introduced by group rescaling also vanishes as $G\to\infty$, whereas the bias from trajectory-wise clipping can persist. Experiments across Qwen3 models and reasoning benchmarks demonstrate improved robustness to stale rollouts, with GMC-GRPO achieving the best performance among stable baselines under large rollout delays.
摘要:非同步強化學習 (RL) 提升了大型語言模型後訓練的效率,但引入了由早期策略生成的過時回饋。對於這種過時性如何影響收斂以及如何減輕其影響的理論理解仍然有限。我們推導了 GRPO 風格算法的收斂界限,明確描述了梯度估計器的二階矩與偏差之間的權衡。對於軌跡級別的重要性加權估計器,我們的分析顯示,一旦二階矩被均勻控制,延遲通過剪裁或重新縮放引入的偏差進入界限。在這一見解的指導下,我們提出了一種新穎的群體質量上限 GRPO (GMC-GRPO) 方法,該方法在共享共同二階矩保證的加權估計器類別中最小化基於比率的偏差界限。我們為非同步 GMC-GRPO 建立了收斂保證,並顯示與 TIC-GRPO 相比,它改善了四階延遲項的閾值依賴性,從 $O(ε^{-4})$ 提升至 $O(ε^{-2})$ 當 $ε\to0$ 時,其中 $1+ε$ 是比率閾值。在局部策略重疊下,延遲依賴項在調整步長後減少為 $G^{-2/5}$,其中 $G$ 是群體大小。對於固定的行為和當前策略,群體重新縮放引入的偏差在 $G\to\infty$ 時也會消失,而來自軌跡級別剪裁的偏差則可能持續存在。針對 Qwen3 模型和推理基準的實驗顯示對過時回饋的穩健性有所改善,GMC-GRPO 在大型回饋延遲下在穩定基準中達到了最佳性能。
A Structured State Space Sequence Model for Multi-Class Classification of Malware
2610.01893v1 by Emmanuela Andam, Rana Shaaban, Emanuel Grant, Naima Kaabouch
By 2030, Internet of Things (IoT) devices are projected to reach 40 billion, with fast-paced technological advancements in fields such as industry, healthcare, agriculture, automobiles, and building/home automation systems. This expansion has created a large attack surface for cybercrime, as the majority of these devices open the door for cybercriminals to exploit vulnerabilities, as they lack adequate built-in security. Cybercriminals launch malware attacks to compromise systems or steal sensitive data, and once a system is compromised, a ransom is typically demanded for its release. Current cybersecurity measures in place are being outpaced by the rapid growth of the IoT, which is accompanied by a subsequent growth in malware variants being created per day. Recognizing this pitfall, this research examines and proposes a novel approach to malware detection and classification to safeguard devices from further attacks and make IoT systems more robust and secure. The framework proposed utilizes a Structured State Space Sequence (S4) model, which discretizes sequences of malware samples in a sequence and captures long-range dependencies, essentially identifying the "cause" and "effect" hidden within malware execution flow. This study presents two novel contributions: the first empirical application of the S4 model for malware analysis, and a comprehensive comparison of its performance against other deep learning architectures, laying the stepping stone for future research in this new paradigm.
摘要:到2030年,物聯網(IoT)設備預計將達到400億個,隨著工業、醫療保健、農業、汽車以及建築/家庭自動化系統等領域的快速技術進步。這一擴張為網絡犯罪創造了巨大的攻擊面,因為大多數這些設備為網絡犯罪分子利用漏洞提供了機會,因為它們缺乏足夠的內建安全性。網絡犯罪分子發動惡意軟體攻擊以破壞系統或竊取敏感數據,一旦系統被攻破,通常會要求贖金以換取其釋放。目前的網絡安全措施已被物聯網的快速增長所超越,這伴隨著每天創造的惡意軟體變種的增長。認識到這一陷阱,本研究檢視並提出了一種新穎的惡意軟體檢測和分類方法,以保護設備免受進一步攻擊並使物聯網系統更加穩健和安全。所提出的框架利用結構化狀態空間序列(S4)模型,該模型將惡意軟體樣本的序列離散化並捕捉長期依賴性,實質上識別出惡意軟體執行流程中隱藏的“原因”和“結果”。本研究提出了兩項新穎的貢獻:首次將S4模型應用於惡意軟體分析,以及對其性能與其他深度學習架構的全面比較,為未來在這一新範式中的研究奠定了基礎。
Selection-Based Structured Reasoning: Toward Efficient Multimodal Search Agents
2610.01892v1 by Feiyu Gavin Zhu, Xiaoyu Zhu, Jiqi Yang, Rui Yang, Arnab Kumar Mondal, Yancheng Wang, Xinke Deng, Jean Oh, Reid Simmons, Joerg Liebelt, Xiang Kong, Zhongyu Jiang
Multimodal agents commonly generate free-form reasoning before each action. For small models, limited model capacity can result in lengthy reasoning that provides little useful guidance for action generation while incurring substantial inference cost. To address this challenge, we introduce Selection-based Structured Reasoning (SSR), a framework that reformulates reasoning as selection instead of open-ended generation. SSR represents recurring high-level reasoning as pre-specified, reusable natural-language candidates. At each turn, the model selects from these reasoning candidates based on their likelihoods given the current context, without requiring an auxiliary task head. Using pre-specified reasoning traces enables parallel scoring, where teacher-forced prefilling computes token likelihoods concurrently within and across candidates using a shared context KV cache. We evaluate SSR on seven multimodal search benchmarks using 2B and 4B models. Across multiple reinforcement learning objectives and supervised fine-tuning, SSR delivers significant efficiency gains without sacrificing task performance. SSR achieves an average success rate competitive with leading search agents of the same scale, while reducing per-turn reasoning latency by over 90% and total per-question model inference latency by 28-54%. Project page: https://zfy0314.github.io/ssr-webpage/.
摘要:多模態代理通常在每個行動之前生成自由形式的推理。對於小型模型,有限的模型容量可能導致冗長的推理,這對於行動生成提供的有用指導有限,同時產生可觀的推理成本。為了解決這一挑戰,我們引入了基於選擇的結構化推理(SSR),這是一個將推理重新定義為選擇而非開放式生成的框架。SSR將重複的高級推理表示為預先指定的、可重用的自然語言候選者。在每一輪中,模型根據當前上下文的可能性從這些推理候選者中進行選擇,而不需要輔助任務頭。使用預先指定的推理痕跡可以實現並行打分,其中教師強制預填充在共享的上下文KV緩存中同時計算候選者內部和之間的標記可能性。我們在七個多模態搜索基準上使用2B和4B模型評估SSR。在多個強化學習目標和監督微調中,SSR在不犧牲任務性能的情況下實現了顯著的效率提升。SSR的平均成功率與同規模的領先搜索代理競爭,同時將每輪推理延遲減少超過90%,每題模型推理延遲減少28-54%。項目頁面:https://zfy0314.github.io/ssr-webpage/。
Unsupervised Domain Adaptation for Enhanced Radiometer Image Precipitation Estimation using Conditional Flow Matching
2610.01890v1 by Victor Enescu, Assaad Zeghina, Matthieu Meignin, Nicolas Viltard, Cécile Mallet
Deep generative networks have recently achieved unprecedented performance in precise image and video editing using sophisticated textual prompts. However, the effectiveness of such models heavily depends on access to very large supervised and annotated image datasets, which can be very difficult to obtain. This is particularly true for satellite instruments, which very rarely overlap with labelled data, and suffer from domain shifts in the rare occasions they do. In this paper, we investigate the potential of flow matching models for unsupervised domain adaptation of satellite radiometer images. Our main contribution is a novel unsupervised method that achieves precise domain alignment by leveraging parts of the deterministic ordinary differential equations in flow matching models, conditioned on different satellite instruments. A key strength of our approach is its ability to preserve essential information while adapting across any domains since the perturbations are in theory bijective. Extensive experiments conducted on the GPM-Core constellation show the benefit of our conditional domain adaptation, particularly in improving rain precipitation estimation from radiometer imagery.
摘要:深度生成網絡最近在使用複雜文本提示進行精確圖像和視頻編輯方面取得了前所未有的表現。
然而,這些模型的有效性在很大程度上依賴於獲取非常大的監督和標註圖像數據集,而這往往非常困難。
這一點對於衛星儀器尤其如此,因為它們與標記數據的重疊非常少,並且在少數情況下重疊時會遭受領域轉移。
在本文中,我們探討了流匹配模型在衛星輻射計圖像的無監督領域適應中的潛力。
我們的主要貢獻是一種新穎的無監督方法,通過利用流匹配模型中確定性常微分方程的部分,實現精確的領域對齊,並以不同的衛星儀器為條件。
我們方法的一個關鍵優勢是它在跨越任何領域時能夠保留重要信息,因為擾動在理論上是雙射的。
在GPM-Core星座上進行的大量實驗顯示了我們的條件領域適應的好處,特別是在改善來自輻射計圖像的降雨量估算方面。
Stochastic Rounding in Low-Precision Transformer Inference: A Variable-Precision Emulation Study of a Small GPT-2
2610.01889v1 by Yohan Chatelain, Pablo de Oliveira Castro
Should low-precision transformer inference use stochastic rounding (SR) or round-to-nearest (RN)? The answer depends on where in the network you look. We isolate this effect by holding the numerical format fixed and varying only the rounding rule at individual operation sites. To enable experiments at freely chosen precisions, we extend the PRISM vectorized rounding library to arbitrary virtual precision via a variable-precision stochastic rounding (VPSR) algorithm, proving that the rounding decision is evaluated exactly in hardware floating point. We develop two analyses providing complementary insight into this site-level trade-off. First, a probabilistic forward-error bound for linear projections shows that SR's error envelope grows as $O(\sqrt{n} u)$ in reduction length $n$, versus $O(n u)$ for RN, a gap that widens rapidly at low precision and is most pronounced in the long multilayer perceptron (MLP) down-projection. Second, a second-order decomposition of expected cross-entropy loss change at the output softmax into signed drift, drift curvature, and a Fisher-weighted variance penalty reveals why the two sites behave oppositely: MLP noise is predominantly a uniform logit shift to which softmax is invariant, so SR's variance is largely discounted; head noise is non-uniform across the vocabulary and is not. On DistilGPT-2 at $t=6$ significand bits, observations match theory: SR in the MLP raises perplexity to 1.15x the full-precision reference, versus 2.21x for RN. At the language-model head, the ordering reverses because SR introduces non-uniform variance, whereas deterministic RN carries none. In a mixed-precision configuration (MLP output at $t=6$), assigning SR to the MLP and RN to the head brings perplexity within 1.10x of the full-precision reference, a 28% reduction over matched-bit RN.
摘要:低精度Transformer推斷應該使用隨機四捨五入 (SR) 還是四捨五入至最近值 (RN)?答案取決於你在網絡中的哪個位置觀察。我們通過固定數字格式並僅在各個操作位置變化四捨五入規則來隔離這一效應。為了在自由選擇的精度下進行實驗,我們通過變精度隨機四捨五入 (VPSR) 算法擴展了 PRISM 向量化四捨五入庫,以支持任意虛擬精度,證明四捨五入決策在硬體浮點中被精確評估。
我們開發了兩個分析,提供對這一位置級權衡的互補見解。首先,對線性投影的概率前向誤差界限顯示,SR 的誤差範圍隨著減少長度 $n$ 增長為 $O(\sqrt{n} u)$,而 RN 則為 $O(n u)$,這一差距在低精度下迅速擴大,並在長多層感知器 (MLP) 向下投影中最為明顯。其次,將輸出 softmax 的期望交叉熵損失變化進行二階分解為有符號漂移、漂移曲率和費舍爾加權方差懲罰,揭示了為什麼這兩個位置的行為相反:MLP 噪聲主要是一種均勻的 logit 偏移,softmax 對此不變,因此 SR 的方差在很大程度上被折扣;而頭部噪聲在詞彙中是非均勻的。
在 $t=6$ 的 DistilGPT-2 顯著位中,觀察結果與理論相符:MLP 中的 SR 將困惑度提高到全精度參考的 1.15 倍,而 RN 則為 2.21 倍。在語言模型頭部,排序顛倒,因為 SR 引入了非均勻方差,而確定性 RN 則沒有。在混合精度配置中 (MLP 輸出為 $t=6$),將 SR 指派給 MLP,將 RN 指派給頭部,使困惑度在全精度參考的 1.10 倍內,相比於匹配位的 RN 減少了 28%。
Flowing Faster to Coordinate: One-Step Online Multi-Agent Flow Policies
2610.01882v1 by Zhuoran Li, Yunzhan Li, Xun Wang, Yihan Du, Longbo Huang
Multi-agent reinforcement learning (MARL) provides a powerful framework for learning coordinated behaviors through interactions with the environment. Developing MARL policies requires balancing expressive modeling of complex and multimodal action distributions with efficient training and execution. Generative policies, particularly diffusionbased policies, can faithfully capture complex and multimodal behaviors, but costly iterative sampling hinders their scalability in online multi-agent settings. We propose an Online MARL framework via one-step Flow model (OMAF) that combines expressive generative policies with efficient one-step action generation. OMAF employs a Transformer-based flow policy to capture complex coordination behaviors, while its approximate path score surrogate provides a principled route to synchronized flow policy optimization. To enable stable and sampleefficient learning, we further develop a joint optimization scheme coupling softmax Q-value estimation with a joint flow policy objective for coordinated policy learning. By eliminating iterative sampling, OMAF dramatically reduces training overhead without sacrificing policy expressiveness. Extensive experiments across 10 standard tasks from MPE and MAMuJoCo show that OMAF consistently achieves superior performance, with up to 3.4x higher returns and 10.5x sample efficiency improvement compared with baseline methods. These results validate the effectiveness of OMAF as an expressive and computationally efficient one-step flow policy paradigm for online MARL.
摘要:多智能體強化學習(MARL)提供了一個強大的框架,通過與環境的互動來學習協調行為。開發MARL策略需要在複雜和多模態行動分佈的表達建模與高效訓練和執行之間取得平衡。生成策略,特別是基於擴散的策略,能夠真實捕捉複雜和多模態行為,但昂貴的迭代取樣限制了它們在在線多智能體環境中的可擴展性。我們提出了一個通過一步流模型(OMAF)的在線MARL框架,將表達豐富的生成策略與高效的一步行動生成相結合。OMAF採用基於Transformer的流策略來捕捉複雜的協調行為,而其近似路徑分數替代品則提供了一條原則性的路徑以實現同步流策略的優化。為了實現穩定且樣本高效的學習,我們進一步開發了一個聯合優化方案,將softmax Q值估計與協調策略學習的聯合流策略目標相結合。通過消除迭代取樣,OMAF顯著減少了訓練開銷,而不犧牲策略的表達性。在來自MPE和MAMuJoCo的10個標準任務中進行的廣泛實驗顯示,OMAF始終實現了卓越的性能,與基線方法相比,回報提高了高達3.4倍,樣本效率改善了10.5倍。這些結果驗證了OMAF作為一種表達豐富且計算高效的一步流策略範式在在線MARL中的有效性。
Where LLMs Fail with Visualization DSLs
2610.01873v1 by Chang Han, Andrew McNutt, Katherine Isaacs
As LLMs take up the role of authoring charts using visualization domain-specific languages (DSLs), the human constraints that shaped those languages may no longer apply, as what is easy for a person is not necessarily easy for a model. To understand how LLMs might work better with DSLs, we explore where and how they fail with current DSL designs. We evaluate 10 JSON-style visualization DSLs with 41 tasks across 3 LLMs, then assess the generated specifications with JSON and rendering checks, and qualitative coding of failed cases. Analyzing how this specification generation process fails, we identify four recurring failure patterns, link each to specific DSL features, and discuss design considerations for future DSL designs.
摘要:隨著大型語言模型(LLMs)擔任使用可視化領域特定語言(DSLs)創建圖表的角色,塑造這些語言的人類限制可能不再適用,因為對於人類來說簡單的事情不一定對模型來說也簡單。為了了解LLMs如何更好地與DSLs協作,我們探討了它們在當前DSL設計中失敗的地方和方式。我們評估了10個JSON風格的可視化DSL,涵蓋41個任務,並在3個LLMs上進行測試,然後通過JSON和渲染檢查以及失敗案例的定性編碼來評估生成的規範。分析這一規範生成過程的失敗,我們識別出四種重複出現的失敗模式,將每種模式與特定的DSL特徵聯繫起來,並討論未來DSL設計的設計考量。
From Network Intrusion Detection to Blockchain-Backed Endpoint Detection and Response: Mapping the Landscape of Decentralized Detection-and-Response Architectures
2610.01872v1 by Yahya Shahsavari, Sara Rouhani, Kaiwen Zhang
While the literature on blockchain-assisted intrusion detection and prevention systems (IDS/IPS) for Internet of Things (IoT) and Industrial Internet of Things (IIoT) networks is mature, existing systematic reviews suffer from two critical limitations: they overlook the structural shift toward modern Endpoint Detection and Response (EDR) and Extended Detection and Response (XDR) architectures, and they conflate blockchain's distinct functional roles into a single monolithic category. This Systematization of Knowledge (SoK) addresses these gaps by proposing a three-axis taxonomy that classifies proposals by detection-system class (NIDS, HIDS, EDR/XDR), blockchain functional role, and response-automation maturity. Synthesizing research published in high-impact venues between 2019 and 2026, we provide a rigorous gap analysis exposing why a genuine per-endpoint blockchain-anchored response loop remains nearly nonexistent due to latency, deployment, and community mismatches. Furthermore, we evaluate structural, cross-cutting challenges persisting across the literature, including consensus latency on constrained devices, post-quantum cryptographic vulnerability, smart-contract attack surfaces, and the adversarial vulnerability of evolving LLM-based detection engines. Finally, we outline a comprehensive research agenda centered on hybrid on-chain/off-chain orchestration to bridge the gap between decentralized trust and rapid response automation.
摘要:雖然有關區塊鏈輔助的入侵檢測和預防系統(IDS/IPS)在物聯網(IoT)和工業物聯網(IIoT)網絡中的文獻已相當成熟,但現有的系統性評估存在兩個關鍵限制:它們忽視了向現代端點檢測與響應(EDR)和擴展檢測與響應(XDR)架構的結構性轉變,並且將區塊鏈的不同功能角色混淆為一個單一的整體類別。這項知識系統化(SoK)通過提出一個三軸分類法來解決這些空白,該分類法根據檢測系統類別(NIDS、HIDS、EDR/XDR)、區塊鏈功能角色和響應自動化成熟度對提案進行分類。綜合2019年至2026年間在高影響力期刊上發表的研究,我們提供了一個嚴謹的差距分析,揭示了為何真正的每個端點區塊鏈錨定響應循環幾乎不存在,原因在於延遲、部署和社群不匹配。此外,我們評估了文獻中持續存在的結構性、跨領域挑戰,包括在受限設備上的共識延遲、後量子密碼學脆弱性、智能合約攻擊面以及不斷演變的基於LLM的檢測引擎的對抗性脆弱性。最後,我們概述了一個以混合鏈上/鏈下協同為中心的全面研究議程,以彌合去中心化信任與快速響應自動化之間的鴻溝。
Walking the Embedding Space: Datastore Extraction from Multimodal RAG
2610.01871v1 by Maria Carmen Jica, Ali Satvaty, Suzan Verberne, Fatih Turkmen
Multimodal Retrieval-Augmented Generation (MRAG) has emerged as a reliable and cost-effective technique of grounding the generative capabilities of Multimodal Large Language Models (MLLMs) into relevant, up-to-date, external knowledge. Despite presenting several benefits, such as reducing hallucinatory behavior, they also introduce new attack surfaces, including leakage of private information and vulnerabilities against data extraction attacks. In this paper, we introduce $\immrag$, an adaptive and automatic data extraction attack procedure operating in a black box setting against \emph{image-returning} MRAG, a configuration in which the retrieved visual artifact is itself the response. Each query blends an attacker-held shadow image with an image already recovered from the system, and relevance-weighted resampling steers subsequent queries towards regions of the embedding space that still yield novel retrievals. Unlike current extraction attacks that aim to persuade the model towards data leakage by placing a malicious query as a textual prompt, $\immrag$ embeds the malicious instructions inside a user-given input image. We evaluate $\immrag$ on three plausible and distinct real-world scenarios: medical assistant, document-focused helper and general purpose tool. The experiments involve the study of the effectiveness of the attack on multiple CLIP-family retrievers, as well as the impact of various generators. A single 2500-query run reconstructs up to 611 distinct radiology images, 566 document scans and 416 general-purpose images under local-feature correspondence, and reaches up to $5.6\times$ as many distinct datastore items as a non-adaptive baseline. Our results show the urgent need for safeguards specifically designed for multimodal data.
摘要:多模態檢索增強生成(MRAG)已成為將多模態大型語言模型(MLLMs)的生成能力與相關的、最新的外部知識相結合的一種可靠且具成本效益的技術。儘管它提供了幾個好處,例如減少幻覺行為,但它們也引入了新的攻擊面,包括私密信息洩漏和對數據提取攻擊的脆弱性。
在本文中,我們介紹了 $\immrag$,這是一種適應性和自動化的數據提取攻擊程序,針對 \emph{圖像返回} MRAG 在黑箱環境中運作,這是一種檢索的視覺工件本身就是回應的配置。每個查詢將攻擊者持有的影像與系統中已恢復的影像混合,並且相關性加權重採樣引導後續查詢朝向仍能產生新穎檢索的嵌入空間區域。與目前旨在通過將惡意查詢作為文本提示來說服模型進行數據洩漏的提取攻擊不同,$\immrag$ 將惡意指令嵌入用戶提供的輸入影像中。我們在三個合理且不同的現實場景中評估了 $\immrag$:醫療助手、文件專注助手和通用工具。實驗涉及對多個 CLIP 家族檢索器的攻擊有效性以及各種生成器的影響進行研究。一次 2500 次查詢的運行重建了多達 611 幅不同的放射學影像、566 幅文件掃描和 416 幅通用影像,根據局部特徵對應,並達到高達 $5.6\times$ 的不同數據庫項目數量,相較於非適應性基準。我們的結果顯示出對專門為多模態數據設計的安全措施的迫切需求。
From Isolated Feature to Orbits: Discovering Music Concepts via Multi-SAE Alignment
2610.01864v1 by Liwei Lin, Gus Xia
How can we understand what a music foundation model has learned \textit{internally}? Most interpretability approaches, such as probing and Sparse Autoencoders (SAEs), focus on identifying individual features with minimal structural assumptions. We argue that many concepts are better understood as \textit{structured relations} rather than isolated features. This is especially prominent in music, where tonal structures are organized in the space of pitch and time. For example, concepts such as chords or keys are naturally expressed as structured sets (e.g., the 12 transpositions of a chord or the diatonic system within a key), rather than isolated features. In this study, \textbf{we shift from feature identification to structure-based analysis}, asking whether the learned inner representations of music foundation model emerge as organized structures over features. To this end, we introduce a framework that uses pitch transposition as an inductive bias to induce ordered orbits via multi-view SAE alignment. Concretely, we generate pitch-shifted input pairs and align their SAE representations to discover structured groups of pitch-related features. Experimental results show that this approach recovers orbit structures corresponding to chords, keys, and melodic patterns across two state-of-the-art music foundation models, while requiring only minimal grounding (e.g., a few anchor examples) to interpret entire concept families.
摘要:如何理解音樂基礎模型所學到的\textit{內部}知識?大多數可解釋性方法,如探測和稀疏自編碼器(SAEs),專注於識別具有最小結構假設的個別特徵。我們認為,許多概念更應被理解為\textit{結構化關係}而非孤立特徵。這在音樂中尤為明顯,音調結構在音高和時間的空間中組織。例如,和弦或調的概念自然表達為結構化集合(例如,和弦的12個移調或調內的自然音階系統),而不是孤立特徵。在這項研究中,\textbf{我們從特徵識別轉向基於結構的分析},詢問音樂基礎模型學習到的內部表徵是否以結構化形式出現。為此,我們引入一個框架,利用音高移調作為誘導偏見,通過多視角SAE對齊來產生有序的軌道。具體而言,我們生成音高移位的輸入對,並對齊它們的SAE表徵,以發現與音高相關特徵的結構化組。實驗結果顯示,這種方法恢復了對應於和弦、調和旋律模式的軌道結構,並且只需最少的基礎(例如,幾個錨點示例)即可解釋整個概念家族。
AVSD-Scenes: A Dataset for Audio-Visual Description of Urban Scenes
2610.01861v1 by Dhanunjaya Varma Devalraju, Arshdeep Singh, Mark D. Plumbley
Natural language descriptions can provide rich semantic representations of audio-visual urban scenes, yet datasets that jointly describe both auditory and visual information remain limited. In this paper, we introduce AVSD-Scenes, a paired audio-visual scene description dataset for urban environments. The dataset contains 12,291 audio-visual scene descriptions generated from the TAU Urban Audio-Visual Scenes dataset. To construct the dataset, we first generate audio- and visual-based descriptions using Qwen2-Audio-7B and Qwen2.5-VL-7B, respectively. These modality-specific descriptions are then combined using large language models, namely Qwen3-14B, Mistral-Small-3.2-24B-Instruct-2506, and Gemma-3-27B-it, to produce multimodal descriptions that capture complementary information from both modalities. We benchmark AVSD-Scenes using semantic alignment, cross-modal retrieval, scene classification, LLM-as-a-judge evaluation, and human subjective assessment. Results show that multimodal descriptions improve semantic alignment and cross-modal retrieval performance compared with modality-specific descriptions while preserving strong scene-discriminative information. The generated descriptions achieve up to 94.5% accuracy in urban scene classification, while combining audio, visual, and description embeddings further improves accuracy to 95.4%. Furthermore, the descriptions remain highly scene-discriminative even when scene labels are removed from the prompting instructions, indicating that they capture semantic information derived from the audio-visual content rather than merely reflecting label information.
摘要:自然語言描述可以提供音視覺城市場景的豐富語義表示,但同時描述聽覺和視覺信息的數據集仍然有限。
在本文中,我們介紹了 AVSD-Scenes,一個針對城市環境的配對音視覺場景描述數據集。
該數據集包含 12,291 條從 TAU Urban Audio-Visual Scenes 數據集中生成的音視覺場景描述。
為了構建該數據集,我們首先分別使用 Qwen2-Audio-7B 和 Qwen2.5-VL-7B 生成基於音頻和視覺的描述。
這些特定於模態的描述然後使用大型語言模型進行結合,即 Qwen3-14B、Mistral-Small-3.2-24B-Instruct-2506 和 Gemma-3-27B-it,以生成捕捉兩種模態互補信息的多模態描述。
我們使用語義對齊、跨模態檢索、場景分類、LLM作為評判標準的評估以及人類主觀評估來基準測試 AVSD-Scenes。
結果顯示,相較於特定模態的描述,多模態描述改善了語義對齊和跨模態檢索性能,同時保留了強大的場景區分信息。
生成的描述在城市場景分類中達到高達 94.5% 的準確率,而將音頻、視覺和描述嵌入結合進一步提高了準確率至 95.4%。
此外,即使在提示指令中移除場景標籤,這些描述仍然保持高度的場景區分性,表明它們捕捉到的是源自音視覺內容的語義信息,而不僅僅是反映標籤信息。
Detecting Inconsistencies in Model Specifications with LLM-as-Verifier Reasoning
2610.01847v1 by Zichen Xie, Mrigank Pawagi, Lize Shao, Yang Hu, Wenxi Wang
Model specifications define how large language models (LLMs) should behave, guiding alignment training, inference-time behavior, and evaluation. Yet these specifications may themselves contain defects: two individually reasonable principles may prescribe incompatible behavior when applied to the same situation, leaving no response that satisfies both. Detecting such inconsistencies is challenging. Formalizing natural-language specifications risks losing subtle distinctions, while behavior-based testing cannot reliably distinguish specification defects from differences in model behavior. We introduce VeriSpec, the first approach to directly detect inconsistencies in model specifications by auditing the specification text itself. Our key insight is to preserve the specification in natural language while using an LLM as a verifier. VeriSpec extracts structured, context-aware rules, constructs a topic-guided graph to cluster behaviorally related rules at the same authority level, and applies LLM-as-verifier reasoning to detect inconsistencies. Applying VeriSpec to the OpenAI Model Spec, we extract 405 rules and manually validate five inconsistencies, all reported to its developers, who responded positively and have initiated internal discussions. Compared with five baselines, VeriSpec identifies the most validated inconsistencies, achieves the highest precision (38.5%), and incurs the lowest cost per validated inconsistency ($11.12). These results establish direct specification auditing as a practical complement to behavioral alignment evaluation, catching defects at the source before they shape any model. The code is available at https://github.com/HIPREL-Group/VeriSpec.
摘要:模型規範定義了大型語言模型(LLMs)應該如何運作,指導對齊訓練、推理時的行為和評估。
然而,這些規範本身可能包含缺陷:兩個各自合理的原則在應用於相同情境時可能會規定不相容的行為,導致沒有任何回應能同時滿足兩者。
檢測這種不一致性是具有挑戰性的。
將自然語言規範形式化可能會失去微妙的區別,而基於行為的測試則無法可靠地區分規範缺陷與模型行為的差異。
我們引入了VeriSpec,這是第一種通過審核規範文本本身直接檢測模型規範中不一致性的方法。
我們的關鍵見解是保留自然語言中的規範,同時使用LLM作為驗證者。
VeriSpec提取結構化的、上下文感知的規則,構建主題引導的圖以聚類同一權威級別下行為相關的規則,並應用LLM作為驗證者的推理來檢測不一致性。
將VeriSpec應用於OpenAI模型規範,我們提取了405條規則並手動驗證了五個不一致性,所有這些都已報告給其開發者,開發者對此做出了積極回應並已啟動內部討論。
與五個基準相比,VeriSpec識別了最多的經過驗證的不一致性,達到了最高的精確度(38.5%),並且每個經過驗證的不一致性的成本最低($11.12)。
這些結果確立了直接規範審核作為行為對齊評估的實用補充,在缺陷影響任何模型之前,及時捕捉到缺陷。
代碼可在 https://github.com/HIPREL-Group/VeriSpec 獲得。
Beyond Decodability: Do Acoustic Factors Drive Predictions in Speech-Based Alzheimer's Assessment?
2610.01846v1 by Serli Kopar, Alkis Koudounas, Roshan P. Rane, Sam Gijsen, Paula A. Perez-Toro, Kerstin Ritter
Speech-based Alzheimer's disease (AD) assessments increasingly rely on pretrained self-supervised learning (SSL) models that learn acoustic representations directly from raw audio, exposing the model to recording factors. We ask whether such factors are merely encoded in SSL representations or can systematically alter predictions. Using ADReSSo and three large SSL backbones, we apply controlled noise and reverberation interventions to participant-speech-only, non-speech, and full-recording audio. We combine layer-wise linear decoding, input- and representation-space interventions, and geometric alignment analysis to distinguish acoustic decodability from influence on AD prediction. Our results show that controlled acoustic interventions alter AD predictions across all three SSL backbones. Noise, despite showing no significant diagnostic-group difference in the original data, produces the strongest intervention effects. Importantly, these effects are systematically structured relative to the classifier's decision direction, replicate on the held-out test set and reverse when the representation-space intervention direction is reversed. Together, these findings show that high predictive performance and the absence of a significant diagnostic-group difference in a measured acoustic factor are not sufficient for robustness. We argue that intervention-based robustness tests should become standard for trustworthy clinical speech models.
摘要:基於語音的阿茲海默症(AD)評估越來越依賴於預訓練的自我監督學習(SSL)模型,這些模型直接從原始音頻中學習聲學表示,讓模型接觸到錄音因素。我們詢問這些因素是否僅僅被編碼在SSL表示中,或是否可以系統性地改變預測。使用ADReSSo和三個大型SSL骨幹,我們對參與者的語音、非語音和完整錄音音頻應用控制噪音和混響干預。我們結合層級線性解碼、輸入和表示空間干預,以及幾何對齊分析,以區分聲學可解碼性與對AD預測的影響。我們的結果顯示,控制的聲學干預改變了所有三個SSL骨幹的AD預測。儘管在原始數據中未顯示出顯著的診斷組差異,噪音卻產生了最強的干預效果。重要的是,這些效果相對於分類器的決策方向系統性地結構化,在保留的測試集中重現,並在表示空間干預方向反轉時逆轉。總的來說,這些發現表明,高預測性能和在測量的聲學因素中缺乏顯著的診斷組差異並不足以保證穩健性。我們主張,基於干預的穩健性測試應成為可信臨床語音模型的標準。
On the Divergence of Accuracy and Mechanism Consistency in Time Series World Models
2610.01842v1 by Haochen Zhang, Jiaheng Guo, Zhen Xu, Zachary Plotkin, Nicholas Konz, Zhen Tan, Tianlong Chen
A time series world model (TSWM) predicts a controlled system's state from its observed history and planned actions and exogenous inputs. Current approaches build forecasters with actions as covariates, trained and evaluated on prediction error under the executed plan. Yet world models compare unexecuted plans, but their responses to changed plans remain untested. We ask which design choices matter and whether accurate forecasters respond to changed plans as real systems do. We address both with a formalization and benchmark. The formalization separates state, actions and exogenous inputs, distinguishes continuous, mode and event actions, and introduces mechanism consistency, a metric built on declared action-state relations with known directions, such as a vasopressor raising blood pressure: it checks whether shifting an action moves the forecast in the declared direction. The benchmark consolidates eight public datasets with real actions from engineered infrastructure and clinical care, varying prediction space, plan fusion and plan encoding across seven backbones and five seeds. First, a frozen latent prediction space lowers MAE by 9.9% over observation space and gated output fusion lowers it by 12.7% over input concatenation on average, with both improving all eight datasets; temporal plan encoding changes average MAE by at most 2.2%. Second, prediction error and mechanism consistency diverge: the lowest-error configuration is at or below chance in consistency on four of five datasets with declared mechanisms, and no design choice avoids this. Finally, directional supervision, a loss penalizing the wrong-signed part of the response to a shifted action, significantly raises consistency on penalized mechanisms with no change in MAE. Together they give TSWMs a recipe: a frozen latent space and output-side fusion for accuracy, and a training objective for mechanism consistency.
摘要:時間序列世界模型 (TSWM) 從其觀察歷史、計畫行動和外部輸入預測受控系統的狀態。當前的方法建立了以行動為協變數的預測器,這些預測器在執行計畫下的預測誤差上進行訓練和評估。然而,世界模型比較未執行的計畫,但對於變更計畫的反應仍未經測試。我們詢問哪些設計選擇是重要的,以及準確的預測器是否像真實系統一樣對變更計畫做出反應。我們通過形式化和基準來解決這兩個問題。形式化將狀態、行動和外部輸入分開,區分連續、模式和事件行動,並引入機制一致性,這是一種基於已知方向的聲明行動-狀態關係構建的指標,例如一種升壓藥提高血壓:它檢查改變行動是否將預測移動到聲明的方向。基準整合了八個來自工程基礎設施和臨床護理的公共數據集,這些數據集中有真實行動,並在七個骨幹和五個種子中變化預測空間、計畫融合和計畫編碼。首先,凍結的潛在預測空間使 MAE 在觀察空間上降低了 9.9%,而門控輸出融合使其在輸入串接上平均降低了 12.7%,兩者都改善了所有八個數據集;時間計畫編碼的變化使平均 MAE 變化最多為 2.2%。其次,預測誤差和機制一致性出現分歧:在五個具有聲明機制的數據集中,最低誤差配置在一致性上與隨機相同或更低,且沒有任何設計選擇能避免這一點。最後,方向性監督,一種對於對移動行動的反應中錯誤符號部分進行懲罰的損失,顯著提高了懲罰機制的一致性,且 MAE 沒有變化。這些共同為 TSWM 提供了一個配方:凍結的潛在空間和輸出側融合以提高準確性,以及一個針對機制一致性的訓練目標。
Code Owns the Simulation, Jev Owns the Evaluation
2610.01834v1 by Yaodong Yang, Hongyao Tang, Yi Ma, Xingyu Fan, Weixun Wang, Jinpeng Li, Tianpei Yang
Judgment models such as \jev{} return, in a single call and without reasoning text, a probability for each described option. This makes them attractive as an agent's action-selection layer, but it is unclear which decisions they can be trusted with. We test \jev{} on reflection tests, one-shot matrix games, the text game ALFWorld and robot control, and find a sharp boundary. \jev{} succeeds when the right option can be judged from what the input describes, which we call \emph{evaluation}. Specifically, it solves 99\% of the counterintuitive Cognitive Reflection Test questions. However, it fails when the right option depends on \emph{simulation} (i.e., predicting something not in the input), such as the opponent's action or the subgoal that must come first. In games, \jev{} plays suboptimally as if its rational opponent acted at random, because the opponent's action is not given. In ALFWorld, \jev{} favors commands that mention an object or place named in the task description. For example, given the task ``put a clean knife in the drawer'', \jev{} carries an unwashed knife straight to the drawer instead of first washing it at the sink. Surprisingly, many of these failures are not due to a lack of knowledge. Asked separately what the opponent will do, \jev{} usually answers correctly, and it responds well given the opponent's action. It fails when one call must both perform the simulation and evaluate based on it. This suggests letting code make the prediction or simulation. When code supplies it, such as a lookahead in ALFWorld and physics simulation in robot control, \jev{} becomes an expert controller through its general evaluation ability.
摘要:判斷模型如 \jev{} 在單次呼叫中返回每個描述選項的概率,且不需要推理文本。這使得它們作為代理的行動選擇層變得具有吸引力,但尚不清楚它們可以信任哪些決策。我們在反思測試、一回合矩陣遊戲、文本遊戲 ALFWorld 和機器人控制上測試 \jev{},並發現了一個明確的邊界。當正確選項可以從輸入描述中判斷時,我們稱之為 \emph{評估},\jev{} 成功地解決了 99\% 的反直覺認知反思測試問題。然而,當正確選項依賴於 \emph{模擬}(即預測輸入中不存在的事物)時,它則失敗,例如對手的行動或必須先完成的子目標。在遊戲中,\jev{} 表現得次優,彷彿其理性的對手隨機行動,因為對手的行動並未給出。在 ALFWorld 中,\jev{} 偏好提到任務描述中物體或地點的命令。例如,給定任務「將乾淨的刀放入抽屜」,\jev{} 直接將未洗的刀帶到抽屜,而不是先在水槽中清洗它。令人驚訝的是,這些失敗中的許多並非因為缺乏知識。單獨詢問對手會做什麼時,\jev{} 通常能正確回答,並且在給定對手的行動時反應良好。當一次呼叫必須同時執行模擬並基於此進行評估時,它則失敗。這表明應讓代碼進行預測或模擬。當代碼提供這些信息時,例如在 ALFWorld 中的前瞻和在機器人控制中的物理模擬,\jev{} 通過其一般評估能力成為專家控制器。
Continuous Process-Level Evaluation for Evolving Enterprise AI Agent Skills
2610.01833v1 by Ngoc Phuoc An Vo, Aarya Doshi, Vadim Sheinin
Enterprise AI agent skills evolve as tool APIs, models, and specifications change, yet final-output evaluation can miss process-level behavioral drift. We present a continuous evaluation framework combining outcome-level and process-level checks, applied to Revenue and Productivity variants of a Business Value Determination skill in an enterprise Value Aware Resiliency system. The framework independently computes per-run ground truth, materializes reusable template tests, and evaluates tool selection, arguments, execution order, and database integrity through programmatic checks and a narrowly scoped LLM judge. We evaluate 240 trials across two skills, two specification variants, two agent harnesses, and three models. Of 175 trials passing all applicable final numerical checks, 162 (92.6 percent; Wilson 95 percent CI: 87.7-95.6 percent) contained another evaluator-detected deviation. Under a broader seven-check final-state definition, 151 of 164 passing runs (92.1 percent; 95 percent CI: 86.9-95.3 percent) still violated a trajectory check. Dependency attribution reduced a mean of 6.34 failed checks per run to 2.65 roots. Specification sensitivity varied by model and harness, with exploratory bootstrap interaction intervals excluding zero for all three Revenue comparisons and one of three Productivity comparisons. Runtime-resolved templates provided reusable regression coverage across the evaluated configurations; longitudinal validation under actual API evolution remains future work.
摘要:企業AI代理的技能隨著工具API、模型和規範的變化而演變,但最終輸出評估可能會忽略過程層級的行為漂移。我們提出了一個持續評估框架,結合了結果層級和過程層級的檢查,應用於企業價值感知彈性系統中商業價值判定技能的收入和生產力變體。該框架獨立計算每次運行的真實值,實現可重用的模板測試,並通過程式檢查和狹義範圍的LLM評判來評估工具選擇、參數、執行順序和數據庫完整性。我們在兩項技能、兩個規範變體、兩個代理框架和三個模型上評估了240次試驗。在175次通過所有適用的最終數值檢查的試驗中,162次(92.6%;Wilson 95% CI:87.7-95.6%)包含了另一個評估者檢測到的偏差。在更廣泛的七項檢查最終狀態定義下,164次通過的運行中有151次(92.1%;95% CI:86.9-95.3%)仍然違反了軌跡檢查。依賴性歸因將每次運行的平均失敗檢查數從6.34減少到2.65個根源。規範敏感性因模型和框架而異,探索性自助引導交互區間在所有三個收入比較和三個生產力比較中的一個中均不包括零。運行時解析的模板在評估的配置中提供了可重用的回歸覆蓋;在實際API演變下的縱向驗證仍然是未來的工作。
The Asymptotics of Language Model Alignment with Memory
2610.01828v1 by Haricharan Balasundaram, V. Arvind Rameshwar
Language model (LM) alignment broadly aims to perturb a given LM $Q$ into an aligned LM $q$ such that i) the outputs produced by $q$ and $Q$ are 'close' in probability, ii) $q$ has a higher expected reward than $Q$. Two common techniques for LM alignment are: KL-constrained RL, which requires knowledge of the LM distribution and is computationally expensive, and the best-of-$n$ algorithm, which requires only sampling from the LM. The work of Yang et al. established asymptotic closeness between the distributions produced by the two alignment methods for an $m$--length i.i.d. token sequence output by the LM, in the limit as $m$ increases to infinity. However, the i.i.d. assumption is not representative of practical LMs, whose output sequences often have memory. In this paper, we extend the asymptotic closeness result to the case when the $m$--length token sequence outputted by the LM is Markovian. Further, for finite-length output sequences -- particularly, when $m=1$ -- we provide a complete characterization of LM distributions and reward functions for which the KL-divergence between the distributions produced by the two alignment methods is zero -- a question first posed in Yang et al.
摘要:語言模型(LM)對齊的廣泛目標是將給定的 LM $Q$ 轉變為一個對齊的 LM $q$,使得 i) $q$ 和 $Q$ 所產生的輸出在概率上是「接近」的,ii) $q$ 的期望獎勵高於 $Q$。兩種常見的 LM 對齊技術是:KL 約束強化學習,這需要對 LM 分佈的了解並且計算上昂貴,以及最佳的 $n$ 算法,這僅需要從 LM 中進行取樣。Yang 等人的研究確立了在 $m$ 長度的獨立同分佈(i.i.d.)標記序列的情況下,兩種對齊方法所產生的分佈之間的漸近接近性,當 $m$ 增加到無限大時。然而,i.i.d. 假設並不代表實際的 LM,因為它們的輸出序列通常具有記憶性。在本文中,我們將漸近接近性結果擴展到 LM 輸出的 $m$ 長度標記序列為馬爾可夫過程的情況。此外,對於有限長度的輸出序列——特別是當 $m=1$ 時——我們提供了 LM 分佈和獎勵函數的完整特徵描述,對於這些情況,兩種對齊方法所產生的分佈之間的 KL 散度為零——這是一個最初由 Yang 等人提出的問題。
Token Communication-Assisted Collaborative Embodied Artificial Intelligence: Concepts, Framework, and Opportunities
2610.01826v1 by Peng Yi, Ying-Chang Liang
Collaborative embodied artificial intelligence (CEAI) enables multiple physical agents to perceive, reason, and act cooperatively in dynamic environments. Effective communication is essential for CEAI, yet CEAI agents must exchange not only large multimodal observations but also task-relevant insights, intents, and interactive information over long horizons. This article investigates token communication (TokCom) as a native intelligence interface for CEAI, in which tokens serve jointly as compact semantic carriers for communication and fundamental inference units for generative foundation models (GFMs). We first discuss how TokCom supports insight sharing, intent alignment, and interactive control among embodied agents. We then propose a TokCom-assisted CEAI framework driven by a task-adaptive communication protocol. Comprising a compact codebook, syntax rules, and contextual examples, this protocol guides GFM-based transceivers to distill messages into compact tokens and reconstruct them after wireless transmission. A case study on collaborative object transport demonstrates that the proposed TokCom framework substantially reduces the source payload bit consumption while preserving task efficiency and showing robustness under noisy channels. Finally, we outline future research directions.
摘要:協作具身人工智慧(CEAI)使多個物理代理能夠在動態環境中共同感知、推理和行動。有效的溝通對於CEAI至關重要,然而CEAI代理必須交換不僅是大量的多模態觀察,還包括與任務相關的見解、意圖和互動信息,並且這些交流需要在長時間範圍內進行。本文探討了作為CEAI本地智能介面的標記通信(TokCom),其中標記共同作為溝通的緊湊語義載體和生成基礎模型(GFMs)的基本推理單元。我們首先討論了TokCom如何支持具身代理之間的見解共享、意圖對齊和互動控制。接著,我們提出了一個由任務自適應通信協議驅動的TokCom輔助CEAI框架。該協議由緊湊的代碼本、語法規則和上下文示例組成,指導基於GFM的發射接收器將消息提煉為緊湊的標記,並在無線傳輸後重建它們。一個關於協作物體運輸的案例研究顯示,所提出的TokCom框架在保持任務效率和在噪聲通道下顯示穩健性的同時,顯著減少了源負載位元消耗。最後,我們概述了未來的研究方向。
Beyond Linear Concepts: Discovering and Aligning Non-Linear Concept Manifolds in Large Language Models
2610.01821v1 by Tido Specht, Elias Benedict Krey, Nils Neukirch, Nils Strodthoff
Understanding information processing in large language models (LLMs) requires dissecting the geometric organization of their internal token representations. While existing mechanistic interpretability (MI) methods seek to extract concepts, they are constrained by a strong linearity assumption challenged by evidence of non-linear feature manifolds. We move beyond linear concepts by adapting Non-Linear Multi-Dimensional Concept Discovery (NLMCD) from computer vision to token-level LLM activations, modeling concepts as low-dimensional manifolds. To compare concept manifolds across layers and models, we introduce a concept-based alignment (CBA) score, a generalized Rand index that measures geometric proximity without explicit feature matching. Our analysis yields six key findings: (i) a neighboring-layer sanity check shows CBA is more sensitive than PCA- or CKA-based linear baselines; (ii) layer-by-layer alignment matrices reveal two block structures in intermediate and late layers, consistent across models and obscured by linear metrics; (iii) concept composition remains syntax-dominated through most of the network before giving way to increasingly mixed syntactic-semantic concepts in later layers, with increasing output-orientation toward the final layers; (iv) multilingual concept sharing between English and Mandarin is training-dependent rather than universal, strongest in Qwen, weaker in Llama, and absent in GPT-2; (v) inter-model alignment mirrors this structure, with strong correspondence between same-family Qwen models of different scale but weak alignment across model families; and (vi) across Tulu-3 training stages, alignment is highest between adjacent stages, with the largest shift between the base model and SFT, while subsequent preference-alignment stages (DPO, RLVR) leave early layers largely unchanged and RLVR mostly preserves DPO's concepts in late layers.
摘要:理解大型語言模型(LLMs)中的信息處理需要剖析其內部標記表示的幾何組織。雖然現有的機械可解釋性(MI)方法試圖提取概念,但它們受到強線性假設的限制,而這一假設受到非線性特徵流形證據的挑戰。我們通過將非線性多維概念發現(NLMCD)從計算機視覺適應到標記級別的LLM激活,超越線性概念,將概念建模為低維流形。為了比較不同層和模型之間的概念流形,我們引入了一種基於概念的對齊(CBA)分數,這是一種廣義的Rand指數,用於測量幾何接近性而不需要明確的特徵匹配。我們的分析得出了六個關鍵發現:(i)相鄰層的合理性檢查顯示CBA比基於PCA或CKA的線性基準更敏感;(ii)逐層對齊矩陣揭示了中間層和後期層中的兩個區塊結構,這在不同模型中是一致的,但被線性指標所掩蓋;(iii)概念組合在網絡的大部分時間內仍然以語法為主,然後在後期層轉向越來越混合的語法-語義概念,並且對最終層的輸出取向逐漸增加;(iv)英語和普通話之間的多語言概念共享是依賴於訓練的,而不是普遍存在的,在Qwen中最強,在Llama中較弱,而在GPT-2中則不存在;(v)模型間的對齊反映了這一結構,同一家族的Qwen模型之間存在強對應,但不同模型家族之間的對齊較弱;(vi)在Tulu-3的訓練階段中,相鄰階段之間的對齊最高,基礎模型和SFT之間的變化最大,而隨後的偏好對齊階段(DPO、RLVR)使早期層幾乎保持不變,並且RLVR在後期層中大多保留了DPO的概念。
A Safe Prototype Is Not a Safety Direction: Reference Dependence and Prompt Confounds in Response-Safety Embeddings
2610.01801v1 by Sahil Kadadekar
Can response safety be scored by cosine similarity to the mean embedding of known-safe responses? A recent sleeper-agent detector proposes exactly this score, yet the raw positive-centroid rule is not identified: positive observations locate the safe class relative to an encoder origin, but do not determine which direction separates safe from unsafe responses. We audit the rule on two prompt-controlled, human-labeled corpora and one auxiliary jury-labeled source control, using four frozen encoders and prompt-grouped splits. On the human-labeled corpora the safe prototype reaches ROC-AUC 0.457-0.545, with two cells significantly below chance and one above, while an explicit safe-minus-unsafe reference reaches 0.588-0.738 on the same embeddings; on the jury control the prototype is inverted (0.358-0.405) and the reference reaches 0.754-0.793. At validation-calibrated 5% false-safe thresholds, the reference accepts more safe responses on PKU-SafeRLHF (0.153-0.263 versus 0.039-0.061 across encoders) and Aegis (0.189-0.291 versus 0.004-0.045), but not reliably on BeaverTails. A fully unlabeled held-out reference recovers part to most of the referenced ranking, much less when only 5% of the pool is unsafe, whereas 80-634 labeled unsafe responses recover most of it. Prompt-only ablations show that prompt-label composition can inflate uncontrolled evaluations. This is a bounded result about a raw positive centroid, not all one-class methods or safety-specialized guards. A class mean is a location, not necessarily a safety direction; a declared reference with enough unsafe mass identifies orientation.
摘要:可以通過與已知安全回應的平均嵌入的餘弦相似度來評分回應的安全性嗎?最近的潛伏特工檢測器正是提出了這一評分,但原始的正中心規則並未被識別:正觀察將安全類別定位於編碼器原點相對的位置,但並未確定哪個方向將安全回應與不安全回應分開。我們對兩個提示控制的人類標記語料庫和一個輔助陪審團標記的源控制進行了審核,使用四個凍結的編碼器和提示分組拆分。在人類標記的語料庫中,安全原型的ROC-AUC達到0.457-0.545,其中兩個單元顯著低於隨機機率,一個則高於隨機機率,而明確的安全減不安全參考在相同的嵌入上達到0.588-0.738;在陪審團控制中,原型被反轉(0.358-0.405),參考達到0.754-0.793。在驗證校準的5%假安全閾值下,參考在PKU-SafeRLHF上接受了更多的安全回應(0.153-0.263對比0.039-0.061,跨編碼器)和Aegis(0.189-0.291對比0.004-0.045),但在BeaverTails上並不可靠。一個完全未標記的保留參考恢復了部分到大多數的參考排名,當只有5%的池是不安全的時候,恢復的程度要小得多,而80-634個標記的不安全回應則恢復了大部分。僅提示的消融實驗顯示,提示標籤的組合可能會膨脹不受控制的評估。這是一個關於原始正中心的有限結果,而不是所有單類方法或安全專門防護的結果。類的平均值是一個位置,不一定是一個安全方向;一個擁有足夠不安全質量的聲明參考確定了方向。
LineupRL: Verifiable Reinforcement Learning for Time Series Captioning via Caption-to-Series Identification
2610.01800v1 by Haochen Zhang, Laura Yao, Zachary Plotkin, Gengwei Zhang, Tianlong Chen
Time series captioning is a fundamental step in time series understanding and can also serve as the bridge between signal and natural language. Supervised fine-tuning (SFT) relies on a larger model's captions and cannot exceed their quality. Reinforcement learning (RL) can, but its rewards were designed for other modalities and other tasks, and they transfer poorly to open-ended generation in the time series domain. We address this by proposing LineupRL, a reinforcement learning with verifiable rewards (RLVR) pipeline whose reward is caption-to-series identification. The reward model is a frozen large language model (LLM) verifier that reads the generated caption and the candidate time series as raw values, never the chart, and must pick the described time series from multiple distractors. Matching is a far lighter demand on the verifier than writing questions or judging a caption, so an off-the-shelf LLM can supply the reward. Across two captioning benchmarks, and on forecasting and reconstruction where the predictor sees only the caption, LineupRL outperforms SFT and RL baselines on every metric. The 3B vision language model (VLM) trained by LineupRL also outperforms, at 1/24 of the parameters, the 72B VLM whose captions the SFT baseline is distilled from. Our case study shows that LineupRL resists reward hacking, and that the captioner it trains both traces the trend and names the values at key points.
摘要:時間序列標註是時間序列理解中的基本步驟,並且可以作為信號與自然語言之間的橋樑。監督微調(SFT)依賴於更大模型的標註,並且無法超越它們的質量。強化學習(RL)可以做到,但其獎勵是為其他模態和其他任務設計的,並且在時間序列領域的開放式生成中轉移效果不佳。我們通過提出LineupRL來解決這個問題,這是一個具有可驗證獎勵的強化學習(RLVR)管道,其獎勵是標註到序列的識別。獎勵模型是一個凍結的大型語言模型(LLM)驗證器,它將生成的標註和候選時間序列作為原始值進行閱讀,而不是圖表,並且必須從多個干擾項中選擇所描述的時間序列。對驗證器的匹配要求比撰寫問題或評估標註輕得多,因此現成的LLM可以提供獎勵。在兩個標註基準上,以及在預測和重建中,當預測器僅看到標註時,LineupRL在每個指標上都超越了SFT和RL基準。由LineupRL訓練的3B視覺語言模型(VLM)在參數為1/24的情況下,也超越了72B VLM,該模型的標註是從SFT基準中提煉出來的。我們的案例研究顯示LineupRL抵抗獎勵操控,並且它訓練的標註者能夠追蹤趨勢並在關鍵點命名數值。
iADD: Improving Alignment and Diversity in Diffusion Policy Optimization
2610.01789v1 by Ashok Prasad Neupane, Saugat Adhikari, Pramish Paudel, Ajad Chhatkuli, Danda Pani Paudel
Reinforcement learning based post training of diffusion models, such as Denoising Diffusion Policy Optimization (DDPO), optimizes a reverse diffusion process under a reward function. However, current approaches to reward optimizations do so at the cost of diversity and quality. In this paper, we provide better tradeoffs through careful theoretical considerations and method design. We analyze the theoretical framework and mathematically demonstrate that \emph{only-latter timestep} updates of diffusion model may be harmful for diversity contrary to the conclusions presented in a previous work. Additionally, we propose an incremental Feynman-Kac training based on strong theoretical foundations in order to achieve the best-yet alignment-diversity tradeoffs. We perform extensive experiments and compare our method against related diffusion policy optimization approaches in three different tasks and also provide strong ablations for each component, thus validating strong performance gains in both alignment and diversity.
摘要:基於強化學習的擴散模型後訓練,例如去噪擴散策略優化(DDPO),在獎勵函數下優化反向擴散過程。
然而,當前的獎勵優化方法以多樣性和質量為代價。
在本文中,我們通過仔細的理論考量和方法設計提供了更好的權衡。
我們分析了理論框架,並數學上證明擴散模型的\emph{僅後時間步}更新可能對多樣性有害,這與之前工作的結論相反。
此外,我們提出了一種基於強大理論基礎的增量費曼-卡克訓練,以實現最佳的對齊-多樣性權衡。
我們進行了廣泛的實驗,並在三個不同的任務中將我們的方法與相關的擴散政策優化方法進行比較,還為每個組件提供了強有力的消融實驗,從而驗證了在對齊和多樣性方面的顯著性能提升。
VETO: Video Efficient Token Optimization for Vision Language Models
2610.01785v1 by Gueter Josmy Faure, Hao Ping Wang, Min-Hung Chen, Winston H. Hsu
Processing long videos with Vision-Language Models (VLMs) is bottlenecked by the quadratic cost of visual tokens, making long-form inference prohibitively expensive. While single-axis compression methods mitigate this, they hit a hard efficiency floor because they treat spatial and temporal redundancy independently. We present VETO (Video Efficient Token Optimization for Vision-Language Models), a training-optional plug-in that eliminates this bottleneck through dual-axis compression: (i) an intra-frame compressor that merges semantically similar tokens within each frame via optimal-transport inspired matching, and (ii) an inter-frame compressor that identifies and merges temporally redundant frames. The key design insight is hierarchical ordering: by first compressing spatial dimensions, VETO drastically reduces the cost of subsequent global temporal matching, bypassing the efficiency wall of single-axis approaches, with an advantage that grows with modern fully-fused attention infrastructure. Empirically, VETO achieves up to 45% faster inference (e.g., on LLaVA-OneVision-7B) while preserving or improving accuracy. Under extreme token starvation (10% budget), VETO outperforms VFlowOpt (54.9%), VisionZip (52.6%), and FastV (47.9%) with 55.7% accuracy. We demonstrate universal applicability across LLaVA-OneVision, InternVL-2.5, and LongVA, with zero-shot accuracy preserved or improved in all cases.
摘要:處理長視頻的視覺語言模型(VLMs)受到視覺標記的二次成本限制,使得長格式推理變得過於昂貴。雖然單軸壓縮方法可以緩解這一問題,但由於它們獨立處理空間和時間冗餘,因此達到了一個硬效率底線。我們提出了 VETO(視頻高效標記優化器),這是一個可選的插件,通過雙軸壓縮消除了這一瓶頸:(i)一個幀內壓縮器,通過受最優運輸啟發的匹配合併每幀內語義相似的標記,以及(ii)一個幀間壓縮器,識別並合併時間上冗餘的幀。關鍵的設計見解是分層排序:通過首先壓縮空間維度,VETO 大幅降低了隨後全局時間匹配的成本,繞過了單軸方法的效率牆,並且隨著現代全融合注意力基礎設施的發展,這一優勢不斷增長。實證結果顯示,VETO 在保持或提高準確度的同時,實現了高達 45% 的推理加速(例如,在 LLaVA-OneVision-7B 上)。在極端標記短缺(10% 預算)下,VETO 的準確率為 55.7%,超過了 VFlowOpt(54.9%)、VisionZip(52.6%)和 FastV(47.9%)。我們展示了在 LLaVA-OneVision、InternVL-2.5 和 LongVA 上的普遍適用性,在所有情況下均保持或提高了零樣本準確度。
Q-Learning for Reachability in MEC-Free MDPs
2610.01781v1 by Lu-Chin Chang, Suguman Bansal
Reinforcement learning (RL) for reachability specifications is fundamental to sequential decision-making. Prior work establishes asymptotic convergence to optimal policies, but only through model-based methods that must explicitly estimate the transition probabilities of the underlying Markov Decision Process (MDP). We present Quasar, the first model-free algorithm with asymptotic guarantees for reachability on the fragment of MDPs free of non-terminal maximal end components (MECs), a building block to which every MDP reduces by the standard MEC quotient. Our algorithm follows the classical Q-learning approach, using temporal-difference updates to converge to an optimal policy without ever learning the transition probabilities. The resulting learner reduces the memory footprint from the O(|S|^2|A|) that model-based methods require to O(|S||A|). On the standardized Quantitative Verification Benchmark Set, our algorithm converges to the optimal policy with orders of magnitude fewer samples than the previous model-based state-of-the-art. Together these results are a concrete step toward the practical deployment of reachability learning and, with it, of specification-guided RL.
摘要:強化學習(RL)對於可達性規範是序列決策的基礎。
先前的研究建立了對最佳政策的漸近收斂,但僅通過必須明確估計基礎馬爾可夫決策過程(MDP)轉移概率的基於模型的方法。
我們提出了Quasar,這是第一個對於不含非終端最大端元組件(MECs)片段的MDP具有漸近保證的無模型算法,這是每個MDP通過標準MEC商減少的構建塊。
我們的算法遵循經典的Q學習方法,使用時間差更新來收斂到最佳政策,而無需學習轉移概率。
最終的學習器將基於模型的方法所需的O(|S|^2|A|)的內存佔用減少到O(|S||A|)。
在標準化的定量驗證基準集上,我們的算法以比先前基於模型的最先進技術少幾個量級的樣本收斂到最佳政策。
這些結果共同為可達性學習的實際部署邁出了具體的一步,並隨之推進了以規範為指導的RL。
CODesign: Consistency from Data to Trajectory in All-Atom Protein Binder Co-Design
2610.01773v1 by Yuanle Mo, Bo Qiang, Haitao Lin, Qinghan Wang, Gang Du, Odin Zhang, Pheng Ann Heng
The central challenge in de novo protein design is generating plausible, mutually compatible structures and sequences, such that each designed sequence folds into its intended structure and the structure accommodates that sequence. Compared to typical two-stage design methods, which decouple the modeling of the interdependent modalities, co-design models improve the cross-modal consistency by jointly generating sequences and structures. However, naively generating sequences and structures simultaneously does not ensure their consistency. To address this challenge, we propose CODesign framework. We improve data consistency by generating approximately 105,000 consistency-distilled dimers. We further promote consistency through a multimodal joint flow model that captures the joint distribution of sequences, backbone structures, and local atomic configurations, together with a consistency-aware joint resampling strategy that iteratively refines sequences and side chains. Experiments show that CODesign achieves state-of-the-art performance with the highest in silico success rates on both protein- and ligand-target binder design. Ablation studies also demonstrate our distilled dataset increases performance by 70.9%, which can be further improved by our proposed resampling mechanism with negligible additional computational cost. Code, model weights and the new dataset will be completely open-source.
摘要:中心挑戰在於全新蛋白質設計中生成合理且相互兼容的結構和序列,使得每個設計的序列能夠摺疊成其預期的結構,並且該結構能夠容納該序列。與典型的兩階段設計方法相比,這些方法將相互依賴的模態建模分開,協同設計模型通過共同生成序列和結構來改善跨模態的一致性。然而,天真地同時生成序列和結構並不能確保它們的一致性。為了解決這一挑戰,我們提出了CODesign框架。我們通過生成約105,000個一致性提煉的二聚體來改善數據一致性。我們進一步通過一個多模態聯合流模型來促進一致性,該模型捕捉序列、主鏈結構和局部原子配置的聯合分佈,並結合一個一致性感知的聯合重採樣策略,該策略迭代地精煉序列和側鏈。實驗表明,CODesign在蛋白質和配體靶向結合物設計上達到了最先進的性能,並且在計算機模擬成功率上達到了最高。消融研究也顯示我們的提煉數據集使性能提高了70.9%,而且可以通過我們提出的重採樣機制進一步改善,且額外的計算成本微乎其微。代碼、模型權重和新數據集將完全開源。
A Matryoshka Hierarchical RAG for Efficient Multi-Hop Question Answering
2610.01767v1 by Gianluca Bonifazi, Christopher Buratti, Michele Marchetti, Federica Parlapiano, Giulia Quaglieri, Davide Traini, Domenico Ursino, Luca Virgili
Retrieval-Augmented Generation (RAG) systems for multi-hop Question Answering (QA) must balance retrieval quality with computational cost. This cost is incurred during indexing time, through the use of expensive Knowledge Graphs (KGs) or Large Language Models (LLMs) to generate summaries, or during querying, through iterative LLM-driven retrieval. To reduce it while maintaining retrieval quality, we present MatRAG, a hierarchical framework that combines RAG systems with Matryoshka Representation Learning (MRL). MatRAG addresses both kinds of cost by aligning the semantic hierarchy of a clustering structure with the nested structure of MRL. Specifically, it organizes the corpus of documents into a Directed Acyclic Graph (DAG) of clusters with progressively coarser granularity. Each level is indexed by a lower Matryoshka dimension. MatRAG pairs an iterative, top-down traversal of the DAG with an entity-driven mechanism that controls the hop budget and re-ranks candidates. We evaluated MatRAG on three standard multi-hop QA benchmarks against seven representative baselines. MatRAG outperforms its strongest competitors in terms of retrieval quality; furthermore, it reduces indexing costs by avoiding KG construction and LLM-based summarization, and lowers query-time costs through dimension-aware similarity.
摘要:檢索增強生成(RAG)系統在多跳問題回答(QA)中必須平衡檢索質量與計算成本。這個成本在索引時產生,通過使用昂貴的知識圖譜(KG)或大型語言模型(LLM)來生成摘要,或在查詢時,通過迭代的LLM驅動檢索。為了在保持檢索質量的同時降低成本,我們提出了MatRAG,一個將RAG系統與馬特里奧什卡表示學習(MRL)相結合的分層框架。MatRAG通過將聚類結構的語義層次與MRL的嵌套結構對齊,解決了這兩種成本。具體而言,它將文檔語料庫組織成一個具有逐漸粗糙粒度的有向無環圖(DAG)聚類。每個層級由較低的馬特里奧什卡維度進行索引。MatRAG將DAG的迭代自上而下遍歷與一種驅動實體的機制相結合,該機制控制跳躍預算並重新排名候選者。我們在三個標準的多跳QA基準上評估了MatRAG,並與七個代表性的基準進行比較。MatRAG在檢索質量方面超越了其最強的競爭對手;此外,它通過避免KG構建和基於LLM的摘要來降低索引成本,並通過維度感知相似性來降低查詢時間成本。
VideoEvolve: Evolving Agent Harnesses for Video Temporal Grounding
2610.01766v1 by Bingjun Luo, Yuhuan Fan, Jialin Guo, Siqi Li
Video temporal grounding aims to localize events in videos from natural-language queries. For agents built around frozen video-language models, the harness determines how queries guide temporal predictions and how those predictions are refined. Manually refining these harnesses requires diagnosing grounding failures and coordinating changes to both agent workflows and instructions. We introduce VideoEvolve, a framework that automatically evolves agent harnesses for video temporal grounding. VideoEvolve uses a Cloze-Structured Harness Representation that preserves stage interfaces while leaving agent workflows and instructions open to evolution. Branch-Guided Harness Evolution preserves promising code branches for continued refinement, using execution feedback to guide local edits and validation to determine which improvements are carried forward. Experiments demonstrate improved grounding performance across multiple benchmarks. Component analyses identify instruction refinement as a consistent source of gains, while the benefits of evolved code vary across evaluation settings. Together, these results support automated harness evolution as an effective approach to improving video temporal grounding. Code is available at https://github.com/bingjunluo/VideoEvolve .
摘要:視頻時間定位旨在從自然語言查詢中定位視頻中的事件。對於基於凍結視頻-語言模型構建的代理,這個工具決定了查詢如何指導時間預測以及這些預測如何被精煉。手動精煉這些工具需要診斷定位失敗並協調對代理工作流程和指令的變更。我們介紹了VideoEvolve,一個自動演化代理工具以進行視頻時間定位的框架。VideoEvolve使用一種克洛茲結構工具表示法,保留了階段接口,同時使代理工作流程和指令保持開放以便演化。分支引導的工具演化保留了有前景的代碼分支以便持續精煉,利用執行反饋來指導局部編輯,並通過驗證來確定哪些改進將被保留。實驗顯示在多個基準上改進了定位性能。組件分析確定指令精煉是一個一致的增益來源,而演化代碼的好處在不同的評估設置中有所不同。總體而言,這些結果支持自動化工具演化作為改善視頻時間定位的有效方法。代碼可在 https://github.com/bingjunluo/VideoEvolve 獲得。
TopK-Guided: Adaptive, Budget-Aware Activation Sparsity for Efficient LLM Inference
2610.01763v1 by Mukund Agarwalla, Chih-Jen Lin
Activation sparsity speeds up large language model (LLM) inference by setting unimportant activations to zero so that the corresponding computations can be skipped. Existing training-free methods, however, make different trade-offs: threshold-based methods such as TEAL adapt the sparsity level to each token but do not tightly control the realised sparsity, while TopK-based methods such as WINA enforce a fixed sparsity level but use the same sparsity budget for every token. Both also apply the same budget across transformer blocks, despite large differences in block sensitivity. We introduce TopK-Guided, a training-free method that addresses both limitations by combining bounded token-level sparsity adaptation with sensitivity-aware block-level budget allocation. Across Llama-2 and Llama-3 models, TopK-Guided consistently improves perplexity and downstream accuracy over TEAL and WINA while preserving essentially the same sparsitydependent projection compute as WINA, with the largest gains at high sparsity. Ablations show that both components provide complementary improvements.
摘要:激活稀疏性透過將不重要的激活設置為零來加速大型語言模型(LLM)的推理,以便可以跳過相應的計算。然而,現有的無需訓練的方法則做出了不同的權衡:基於閾值的方法如TEAL將稀疏性水平適應於每個標記,但並未嚴格控制實現的稀疏性,而基於TopK的方法如WINA則強制執行固定的稀疏性水平,但對每個標記使用相同的稀疏性預算。儘管區塊敏感性存在很大差異,但兩者在Transformer區塊上也應用了相同的預算。我們介紹了TopK-Guided,這是一種無需訓練的方法,通過將有界的標記級稀疏性適應與敏感性意識的區塊級預算分配相結合,解決了這兩個限制。在Llama-2和Llama-3模型中,TopK-Guided始終在保持與WINA基本相同的稀疏性依賴投影計算的同時,顯著提高了對TEAL和WINA的困惑度和下游準確性,在高稀疏性下獲得了最大的增益。消融實驗顯示,這兩個組件提供了互補的改進。
SoK: Decentralized Agent Economic Infrastructure
2610.01756v1 by Rui Sun, Xihan Xiong, Qin Wang, Fei Gao, Zelin Li, Zehua Cheng, Jiahao Sun, Zhipeng Wang
Decentralized agent economies increasingly build a single task from protocols that were designed and secured separately. This creates a simple problem: a workflow can look correct at each step and still produce the wrong outcome. For example, a correct escrow may release payment on an authorized approval that provides little evidence that the delivered work actually satisfied the task. We systematize this problem across the full lifecycle of an agent task. Our study organizes security and economic requirements into 17 property families over six stages, with receipt soundness and completeness assessed separately. We examine 12 systems and standards, five reusable mechanism families, and four classical baselines. We introduce guarantee closure, a task-relative criterion for determining whether guarantees established at one stage remain available and constrain the later decisions that depend on them. We apply the criterion to controlled and native workflows, covering 840 matched executions and an exhaustive 11,648-case check over a finite objective-task domain. Our results expose recurring failures between verification and settlement, where conforming work can remain unaccepted or valid evidence can be ignored. Public records and model judgments further distinguish recorded approval from evidence of task conformance, while economic analysis identifies the report, penalty, and shared-error assumptions behind these guarantees. These findings show where end-to-end guarantees fail and what must be repaired to preserve them across the workflow.
摘要:去中心化的代理經濟越來越多地從單獨設計和保護的協議中構建單一任務。這造成了一個簡單的問題:工作流程在每一步看起來都正確,但仍然可能產生錯誤的結果。舉例來說,一個正確的保管協議可能會在授權批准下釋放付款,但該批准幾乎沒有證據表明交付的工作實際上滿足了任務要求。
我們在代理任務的整個生命周期中系統化這個問題。我們的研究將安全和經濟要求組織成17個屬性家族,分為六個階段,並分別評估收據的健全性和完整性。我們檢查了12個系統和標準、五個可重用的機制家族以及四個經典基準。我們引入了保證閉合,這是一個相對於任務的標準,用於確定在一個階段建立的保證是否仍然可用,並約束依賴於它們的後續決策。
我們將該標準應用於受控和本地工作流程,涵蓋840次匹配執行以及在有限的目標任務域內進行的11,648個案例的徹底檢查。我們的結果揭示了驗證與結算之間的重複失敗,在這些情況下,符合要求的工作可能仍然未被接受,或者有效證據可能被忽視。公共記錄和模型判斷進一步區分了記錄的批准與任務符合性的證據,而經濟分析則確定了這些保證背後的報告、懲罰和共享錯誤假設。這些發現顯示了端到端保證失效的地方,以及為了在工作流程中保留這些保證必須修復的內容。
Cog-VADU: A Training-Free Cognitive Reasoning Framework for Video Anomaly Detection and Understanding
2610.01754v1 by Mohd Ubaid Wani, Sara Atito, Josef Kittler, Muhammad Awais
Video Anomaly Detection (VAD) aims to temporally localize abnormal events in videos. Most existing approaches rely on dataset-specific training and curated annotations, limiting generalization in open-set scenarios. Recent zero-shot methods based on Large Vision- Language Models (LVLMs) alleviate this dependency but often lack temporal continuity and structured reasoning. We propose Cog-VADU, a fully training-free framework that reformulates VAD as a sequential cognitive reasoning task. Cog-VADU introduces Chain-of- Anomaly Detection Thought Prompting (CoADTP), which unrolls an LVLM into a recurrent reasoning chain across video segments. By propagating structured rationales over time, the model maintains implicit temporal memory, enabling robust discrimination between com- plex anomalies and high-motion normal activities. To improve reliability, we further design a cross-modal re-ranking stage that aligns textual rationales with visual embeddings, enforcing semantic consistency and temporal coherence for refined and stable predictions. Extensive experiments on multiple public VAD benchmarks demonstrate that Cog-VADU achieves competitive zero-shot performance. Moreover, cross-model evaluations show that CoADTP consistently enhances reasoning-based anomaly detection in a model-agnostic manner, pro- viding interpretable and generalizable anomaly understanding for real-world applications.
摘要:視頻異常檢測(VAD)旨在時間上定位視頻中的異常事件。大多數現有的方法依賴於特定數據集的訓練和精心策劃的標註,限制了在開放集場景中的泛化能力。最近基於大型視覺-語言模型(LVLMs)的零樣本方法減輕了這一依賴,但通常缺乏時間連續性和結構化推理。我們提出了Cog-VADU,一個完全無需訓練的框架,將VAD重新定義為一個序列認知推理任務。Cog-VADU引入了異常檢測思維提示鏈(CoADTP),將LVLM展開為跨視頻片段的遞歸推理鏈。通過隨時間傳播結構化的推理,該模型維持隱式的時間記憶,使其能夠在複雜異常和高運動正常活動之間進行穩健的區分。為了提高可靠性,我們進一步設計了一個跨模態重新排序階段,將文本推理與視覺嵌入對齊,強化語義一致性和時間連貫性,以實現精細和穩定的預測。在多個公共VAD基準上的廣泛實驗表明,Cog-VADU實現了具有競爭力的零樣本性能。此外,跨模型評估顯示,CoADTP始終以模型無關的方式增強基於推理的異常檢測,為現實世界應用提供可解釋和可泛化的異常理解。
Removing spurious minima for planar features by skip connections
2610.01728v1 by Jakob Paul Zimmermann, Moritz Grillo, Andrei Balakin, Georg Loho
Understanding loss landscapes is central to explaining neural-network training, yet their structure remains only partially understood even in simple models. We study the Gaussian population loss of shallow, bias-free ReLU networks in the teacher--student setting. This provides a simple model for studying essential aspects such as feature learning and overparameterization. For teacher networks with positive output weights and planar features, we show that including a learned linear skip removes all spurious local minima with non-negative student output weights once the student network is at least as wide as the teacher network. In contrast, without the skip, we construct a fixed teacher network with positive output weights and only three hidden neurons in input dimension two whose spurious local minima persist at every student width at least three. Thus, a learned linear skip can remove spurious minima that persist under arbitrary overparameterization. Furthermore, we show that a positive output weight student network always learns the subspace spanned by the teacher features: student features at local minima with non-negative student output weights lie in the span of the teacher features. For ReLU networks in two dimensions, even heavily overparameterized student networks have effective width controlled by the teacher width: every critical point with positive student output weights has at most twice as many distinct student feature directions as teacher neurons. Finally, we transfer the benignity result to empirical minima over parameter balls of any prescribed radius, with the required sampling accuracy depending on that radius.
摘要:理解損失景觀對於解釋神經網絡訓練至關重要,然而即使在簡單模型中,它們的結構仍然只有部分被理解。
我們研究教師-學生設置中淺層、無偏的ReLU網絡的高斯族群損失。
這提供了一個簡單的模型來研究如特徵學習和過度參數化等基本方面。
對於具有正輸出權重和平面特徵的教師網絡,我們顯示包含學習的線性跳過可以消除所有具有非負學生輸出權重的虛假局部最小值,只要學生網絡的寬度至少與教師網絡一樣寬。
相反,如果不使用跳過,我們構造了一個固定的教師網絡,其具有正輸出權重且在輸入維度為二的情況下只有三個隱藏神經元,這樣的虛假局部最小值在每個學生寬度至少為三的情況下持續存在。
因此,學習的線性跳過可以消除在任意過度參數化下持續存在的虛假最小值。
此外,我們顯示具有正輸出權重的學生網絡總是學習由教師特徵所跨越的子空間:在具有非負學生輸出權重的局部最小值下,學生特徵位於教師特徵的跨度內。
對於二維的ReLU網絡,即使是高度過度參數化的學生網絡,其有效寬度也受到教師寬度的控制:每個具有正學生輸出權重的臨界點最多有教師神經元的兩倍不同學生特徵方向。
最後,我們將良性結果轉移到任何指定半徑的參數球上的經驗最小值,所需的取樣精度取決於該半徑。
vFedProtoQNAS: Prototype-Guided Personalized Quantum Neural Architecture Search for Virtual Federated Learning
2610.01718v1 by Seok Bin Son, Samuel Yen-Chi Chen, Soohyun Park, Joongheon Kim
Quantum federated learning (QFL) has emerged as a promising approach for collaboratively training compact quantum neural networks (QNNs) over distributed private data on resource-constrained devices. However, differences in device capabilities make a single shared QNN architecture unsuitable for all clients. While personalized quantum neural architecture search (QNAS) allows each client to select a device-specific QNN, averaging parameters across structurally different QNN architectures mixes semantically inconsistent circuit operations. To address this, prototype-guided personalized QNAS for virtual FL (vFedProtoQNAS) is proposed, where model parameters are never aggregated across clients and federated collaboration is achieved through class-wise prototype sharing. Each client independently searches and trains a client-specific QNN, computes class-wise local prototypes from latent representations, and refines them using global prototypes from the server as federated semantic anchors. Experiments demonstrate that vFedProtoQNAS improves accuracy by 3.70\% over FedAvg and enhances class-consistent representation alignment.
摘要:量子聯邦學習(QFL)已成為一種有前景的方法,用於在資源有限的設備上協作訓練緊湊的量子神經網絡(QNNs),以處理分散的私有數據。
然而,設備能力的差異使得單一共享的QNN架構不適合所有客戶端。
雖然個性化量子神經架構搜索(QNAS)允許每個客戶端選擇特定於設備的QNN,但在結構上不同的QNN架構之間平均參數會混合語義不一致的電路操作。
為了解決這個問題,提出了針對虛擬聯邦學習(vFedProtoQNAS)的原型引導個性化QNAS,在這裡模型參數從不在客戶端之間聚合,聯邦協作是通過類別級原型共享來實現的。
每個客戶端獨立搜索和訓練特定於客戶端的QNN,從潛在表示中計算類別級本地原型,並使用來自伺服器的全局原型進行精煉,作為聯邦語義錨點。
實驗表明,vFedProtoQNAS的準確率比FedAvg提高了3.70\%,並增強了類別一致的表示對齊。
CoEvolve: Construct-to-Edit Visual Grounding with Bidirectional State Refinement
2610.01710v1 by Dongwei Sun, Yujie Zhang, Bowen Yao, Pei Liu, Jing Yao, Xiangyong Cao
Visual grounding localizes an object described by language with a bounding box. Most multimodal grounding models compress target identification, spatial reasoning, and boundary estimation into one terminal prediction. Free-form rationales make reasoning linguistically explicit but do not necessarily expose measurable, editable spatial states. Intermediate localization errors are therefore difficult to diagnose and correct, allowing incorrect region choices and imprecise boundaries to persist in the final box. We introduce CoEvolve, a construct-to-edit framework that separates grounding into explicit state construction and state editing. Region-Evolution Reinforcement (RER) organizes grounding analysis into a progressive semantic--spatial trajectory, with each reasoning step committing to an explicit candidate region. Bidirectional Denoising Refiner (BDR) treats the reasoning text as fixed semantic context and refines the trajectory's coordinate fields through bidirectional same-position reconstruction. Geometry- and behavior-level objectives provide target geometry and edit-preference signals for consolidating reliable candidates, preserving accurate inputs, or correcting toward annotations. Evaluations cover natural-image and remote-sensing grounding. With a 9B backbone, CoEvolve rivals models up to 241B parameters in grounding accuracy. Under controlled corruption, a single BDR pass improves mean box overlap by over 27 percentage points, demonstrating strong recovery from substantial localization errors. State-source comparisons further support the complementarity of explicit state construction and source-matched editing. The project is at https://sundongwei.github.io/CoEvolve_Project/.
摘要:視覺基礎將語言描述的物體定位於邊界框內。大多數多模態基礎模型將目標識別、空間推理和邊界估計壓縮為一個終端預測。自由形式的推理使推理在語言上變得明確,但不一定揭示可測量、可編輯的空間狀態。因此,中間定位錯誤難以診斷和修正,導致不正確的區域選擇和不精確的邊界在最終框中持續存在。我們介紹了 CoEvolve,一個構建-編輯框架,將基礎分為明確的狀態構建和狀態編輯。區域演化強化(RER)將基礎分析組織成一個漸進的語義-空間軌跡,每一步推理都承諾於一個明確的候選區域。雙向去噪精煉器(BDR)將推理文本視為固定的語義上下文,並通過雙向同位置重建來精煉軌跡的坐標場。幾何和行為層面的目標提供了目標幾何和編輯偏好信號,以鞏固可靠的候選者,保留準確的輸入或朝向註釋進行修正。評估涵蓋自然影像和遙感基礎。在 9B 的主幹下,CoEvolve 在基礎準確性上與高達 241B 參數的模型相媲美。在受控損壞下,單次 BDR 通過提高平均框重疊超過 27 個百分點,顯示出從重大定位錯誤中強有力的恢復。狀態源比較進一步支持明確狀態構建和源匹配編輯的互補性。該項目位於 https://sundongwei.github.io/CoEvolve_Project/。
Task-Oriented Rank Adaptation for Continual Learning in Text Classification
2610.01702v1 by Rey Sanchez Lopez, Eduardo Morales Manzanares, Hugo Jair Escalante
Continual learning (CL) in text classification faces two critical challenges: catastrophic forgetting and negative transfer across sequential tasks. Parameter-Efficient Fine-Tuning (PEFT) methods such as LoRA enable efficient adaptation by learning low-rank updates of the model parameters. However, these compact representations are normally trained in isolation, limiting their reuse across related tasks. We introduce Task-Oriented Rank Adaptation (TORA), a geometric routing framework that leverages the low-rank structure of LoRA adapters to decide whether to transfer knowledge from the most compatible expert (Boosting) or isolate the new task (Shielding) based on structural similarity. Evaluated across 15 diverse text classification benchmarks, TORA consistently avoids harmful routing decisions: compatible tasks exceed their isolated performance while reducing training time, and structurally distant tasks are protected from interference with no loss in accuracy. With a single geometric threshold and no reliance on task identities or predefined sequences, TORA provides a simple and effective approach for dynamic adapter routing in sequential text classification systems.
摘要:持續學習(CL)在文本分類中面臨兩個關鍵挑戰:災難性遺忘和在序列任務中的負轉移。參數高效微調(PEFT)方法如 LoRA 通過學習模型參數的低秩更新來實現高效適應。然而,這些緊湊的表示通常是在孤立的情況下訓練的,限制了它們在相關任務中的重用。我們引入了任務導向秩適應(TORA),這是一個幾何路由框架,利用 LoRA 適配器的低秩結構來決定是從最兼容的專家(提升)轉移知識,還是根據結構相似性隔離新任務(保護)。在 15 個不同的文本分類基準上進行評估,TORA 始終避免有害的路由決策:兼容任務的表現超過其孤立的性能,同時減少訓練時間,而結構上相距較遠的任務則受到保護,沒有準確度損失。TORA 以單一的幾何閾值運作,且不依賴於任務身份或預定序列,為序列文本分類系統中的動態適配器路由提供了一種簡單而有效的方法。
Acmite: Mitigating Gender Bias in LLMs through Concept-Guided Mutual Information
2610.01696v1 by Tian Lan, Xiaoqing Cheng, Han Zhang, Jiang Li
Large language models (LLMs) can reproduce social stereotypes from their training data, motivating extensive research on model debiasing. However, existing methods often rely on explicit biased examples or predefined group-term substitutions, making them sensitive to wording and less effective at capturing stereotype concepts shared across diverse contexts. More importantly, they typically suppress biased outputs without explicitly modeling the statistical dependence between model outputs and the underlying stereotype concepts. We propose Acmite, a lightweight concept-guided framework for targeted and selective debiasing. Acmite represents stereotypes as structured semantic concepts and uses maximal marginal relevance (MMR) to select diverse concepts for debiasing. Inspired by mutual information minimization, it approximates this dependence with token-level KL divergence while preserving task semantics. A lightweight LoRA adapter is trained with the base model frozen and activated at inference time only when the input is sufficiently similar to stereotype-related concepts; otherwise, the original model is used directly. We evaluate Acmite on BBQ, CrowS-Pairs, and StereoSet, and assess general capability preservation on ARC-Challenge, GSM8K, and PIQA. Experiments across three LLMs show that Acmite effectively mitigates gender bias across complementary evaluation formats while maintaining competitive performance on bias-unrelated tasks. Anonymous code and data are available at https://anonymous.4open.science/r/Acmite-18E2/.
摘要:大型語言模型(LLMs)可以從其訓練數據中再現社會刻板印象,這促使了對模型去偏見的廣泛研究。
然而,現有的方法通常依賴於明確的偏見示例或預定義的群體術語替代,這使得它們對措辭敏感,並且在捕捉跨多樣背景的刻板印象概念時效果不佳。
更重要的是,它們通常會抑制偏見輸出,而不明確建模模型輸出與潛在刻板印象概念之間的統計依賴。
我們提出了Acmite,一種輕量級的概念引導框架,用於有針對性和選擇性的去偏見。
Acmite將刻板印象表示為結構化的語義概念,並使用最大邊際相關性(MMR)來選擇多樣的概念進行去偏見。
受到互信息最小化的啟發,它通過標記級的KL散度來近似這種依賴,同時保留任務語義。
一個輕量級的LoRA適配器在基礎模型凍結的情況下進行訓練,並僅在輸入與刻板印象相關概念足夠相似時在推理時啟用;否則,直接使用原始模型。
我們在BBQ、CrowS-Pairs和StereoSet上評估Acmite,並在ARC-Challenge、GSM8K和PIQA上評估一般能力的保留。
在三個LLM上的實驗顯示,Acmite有效減輕了性別偏見,並在互補評估格式中保持了競爭性能,對於與偏見無關的任務也是如此。
匿名代碼和數據可在 https://anonymous.4open.science/r/Acmite-18E2/ 獲得。
Compound interpretation is based on analogy
2610.01688v1 by Tian Shen, Harald Baayen
How compound meanings are best predicted from constituent meanings remains a central question in computational models of lexical semantics. Comparing different computational models provides a way to evaluate alternative accounts of how semantic information is combined during compound comprehension. We propose a new model, the Compound Analogy Model (CAM), that predicts a compound's embedding by adding its constituent embeddings together with the average shift vectors of the two constituents' compound families. The resulting model is parameter-free and exploits local analogical structure in the semantic space. We evaluated CAM against the CAOSS model on Mandarin Chinese compounds. CAM consistently achieved higher prediction accuracy than CAOSS on both training and held-out data, with the exception of three-character compounds, for which analogical generalization is constrained by both small constituent families and a pronounced imbalance in family size between the two constituents. The advantage of CAM remained when evaluation was based on frequency-defined train-test splits that better approximate generalization from familiar to novel compounds. To assess the cognitive plausibility of the two models, we further examined whether model-derived semantic measures predict visual lexical decision latencies for two-character compounds. Predictors derived from CAM provided improved prediction for response latencies compared to predictors derived from the CAOSS model. These findings indicate that compound meaning is better characterized as local analogical generalization than as the application of a learned global linear transformation, and demonstrate that analogical semantic structure provides a cognitively plausible basis for compound comprehension.
摘要:如何從成分意義中最佳預測複合意義仍然是計算語義學模型中的一個核心問題。
比較不同的計算模型提供了一種評估替代解釋的方式,這些解釋涉及在複合理解過程中語義信息是如何結合的。
我們提出了一個新模型,即複合類比模型(Compound Analogy Model, CAM),它通過將成分嵌入與兩個成分的複合家族的平均位移向量相加來預測複合詞的嵌入。
結果模型是無參數的,並利用語義空間中的局部類比結構。
我們在普通話的複合詞上將CAM與CAOSS模型進行了評估。
CAM在訓練數據和保留數據上始終比CAOSS達到更高的預測準確性,除了三字複合詞,因為類比推廣受到成分家族小和兩個成分之間家族大小明顯不平衡的限制。
當評估基於頻率定義的訓練-測試拆分時,CAM的優勢依然存在,這些拆分更好地近似從熟悉到新穎的複合詞的推廣。
為了評估這兩個模型的認知合理性,我們進一步檢查了模型衍生的語義度量是否能預測兩字複合詞的視覺詞彙決策延遲。
與CAOSS模型衍生的預測因子相比,CAM衍生的預測因子對反應延遲的預測有所改善。
這些發現表明,複合意義更好地被描述為局部類比推廣,而不是學習的全局線性變換的應用,並且展示了類比語義結構為複合理解提供了認知上合理的基礎。
Architectural Sampling: Test-Time Scaling via Computational Diversity in Frozen Vision-Language Models
2610.01687v1 by Akshit Singh, Shyam Marjit, Wei Lin, Leonid Karlinsky, M. Jehanzeb Mirza
Test-time scaling often seeks better answers by sampling multiple responses from a frozen model, yet conventional temperature sampling generates every candidate along the same fixed computation path. We introduce architectural sampling, a training-free method that generates candidates through distinct forward computations by reusing selected blocks of decoder layers. Varying the block location and repetition count introduces computational diversity without updating model weights or adding auxiliary parameters. Across five Qwen checkpoints and twelve multimodal benchmarks, architectural sampling improves pass@9 over standard-path temperature sampling by 6.58 percentage points on average at the same nine-candidate budget. Reusing early layers yields the strongest gains, and the improvement in candidate coverage persists even under greedy decoding. The resulting candidates show lower lexical overlap and improve accuracy when used as rollouts for label-free test-time reinforcement learning. These findings extend the benefits of our architectural sampling beyond candidate coverage, demonstrating more effective learning from a model's own outputs.
摘要:測試時的擴展通常透過從凍結模型中抽樣多個回應來尋求更好的答案,然而傳統的溫度抽樣沿著相同的固定計算路徑生成每個候選項。
我們引入了架構抽樣,這是一種無需訓練的方法,通過重用選定的解碼器層區塊來通過不同的前向計算生成候選項。
變更區塊位置和重複次數引入了計算多樣性,而不必更新模型權重或添加輔助參數。
在五個Qwen檢查點和十二個多模態基準測試中,架構抽樣在相同的九個候選預算下,平均提高了相對於標準路徑溫度抽樣6.58個百分點的pass@9。
重用早期層獲得了最強的增益,即使在貪婪解碼下,候選覆蓋的改善仍然持續。
所生成的候選項顯示出較低的詞彙重疊,並在用作無標籤測試時的強化學習回滾時提高了準確性。
這些發現擴展了我們的架構抽樣的好處,不僅限於候選覆蓋,還展示了從模型自身輸出中更有效的學習。
Iterative Policy Refinement through Semantic Rollout Analysis
2610.01652v1 by Feiyu Gavin Zhu, Qi Xu, Zhifei Deng, Zhigang Hua, Luke Simon, Jean Oh, Reid Simmons
Structured policies improve efficiency, robustness, and interpretability in imitation learning by introducing task-specific inductive bias, but existing structure generation methods rely either on extensive human input or on static domain knowledge encoded in LLMs, which may be inconsistent with the expert demonstrations. We propose a closed-loop framework that iteratively refines structured policies using LLM-guided analysis of policy rollouts. By logging rollouts as semantically meaningful tabular data and prompting the LLM to generate diagnostic analysis code, our method identifies suboptimalities in the policy structure and iteratively corrects them without requiring human instruction. Experiments on car racing and door opening tasks show that our approach improves imitation learning performance by up to 15% over zero-shot LLM-generated structures and requires 75% less compute to achieve the same reinforcement learning performance. These results demonstrate that tabular rollout analysis provides an effective feedback signal to align LLM-generated policy structures with expert demonstrations, and we can utilize it to generate good policy structures automatically.
摘要:結構化政策透過引入特定任務的歸納偏見來提升模仿學習的效率、穩健性和可解釋性,但現有的結構生成方法要麼依賴大量的人類輸入,要麼依賴於編碼在大型語言模型(LLMs)中的靜態領域知識,這可能與專家的示範不一致。我們提出了一個閉環框架,通過使用LLM引導的政策展開分析來迭代地改進結構化政策。通過將展開記錄為語義上有意義的表格數據,並提示LLM生成診斷分析代碼,我們的方法識別出政策結構中的次優性,並在不需要人類指導的情況下進行迭代修正。在賽車和開門任務上的實驗顯示,我們的方法在模仿學習性能上比零樣本LLM生成的結構提高了多達15%,並且需要75%更少的計算來達到相同的強化學習性能。這些結果表明,表格展開分析提供了一個有效的反饋信號,以使LLM生成的政策結構與專家示範對齊,我們可以利用它自動生成良好的政策結構。
MCIR: A Feature Dependence-Aware Explainability Method with Reliability Guarantees
2610.01641v1 by Poushali Sengupta, Sabita Maharjan, Frank Eliassen, Shashi Raj Pandey, Yan Zhang
Modern machine-learning models often contain strongly dependent or redundant features, making feature attribution difficult because shared predictive information can be distributed across correlated predictors. Existing methods such as SHAP, LIME, HSIC, MI/CMI, and SAGE may therefore produce unstable rankings under multicollinearity or near-duplicate predictors. We propose the Mutual Correlation Impact Ratio Method (MCIR-M), a dependence-aware global feature-importance approach that quantifies the unique predictive information contributed by each feature beyond a selected dependence neighbourhood. MCIR-M introduces the Mutual Correlation Impact Ratio (MCIR), which conditions each feature on strongly dependent neighbours and computes a normalized ratio of conditional to block-level information. The population score lies in [0,1] and equals zero under exact conditional redundancy. We also introduce a lightweight estimation procedure that computes MCIR using a fraction of the available data and evaluates agreement with full-data explanations. Across controlled synthetic redundancy experiments and the UCI HAR benchmark, MCIR shows dependence-aware ranking behaviour, with its clearest advantage under injected near-duplicate predictors. Comparisons with independent and conditional SHAP, SAGE, HSIC, MI-based scores, and CIR-family baselines are mixed across real-data criteria. Reduced explanation samples lower computational burden in the evaluated configurations, while agreement with full-data explanations is assessed separately through ranking, head-set, and faithfulness diagnostics. Overall, MCIR-M provides a practical dependence-aware diagnostic for global explanation under strong feature dependence.
摘要:現代機器學習模型通常包含強相關或冗餘的特徵,使得特徵歸因變得困難,因為共享的預測信息可能分佈在相關的預測變數之間。現有的方法如SHAP、LIME、HSIC、MI/CMI和SAGE在多重共線性或近乎重複的預測變數下可能因此產生不穩定的排名。我們提出了互相關影響比率方法(MCIR-M),這是一種考慮依賴性的全局特徵重要性方法,量化每個特徵在選定的依賴鄰域之外所貢獻的獨特預測信息。MCIR-M引入了互相關影響比率(MCIR),該比率在強依賴的鄰居上對每個特徵進行條件化,並計算條件信息與區塊級信息的標準化比率。該人口得分位於[0,1]之間,並在精確的條件冗餘下等於零。我們還引入了一種輕量級的估計程序,該程序使用部分可用數據計算MCIR,並評估與全數據解釋的一致性。在受控的合成冗餘實驗和UCI HAR基準測試中,MCIR顯示出考慮依賴性的排名行為,其在注入的近重複預測變數下的優勢最為明顯。與獨立和條件SHAP、SAGE、HSIC、基於MI的得分以及CIR系列基準的比較在真實數據標準下是混合的。在評估的配置中,減少的解釋樣本降低了計算負擔,而與全數據解釋的一致性則通過排名、頭部集和忠實性診斷單獨評估。總體而言,MCIR-M為強特徵依賴下的全局解釋提供了一種實用的考慮依賴性的診斷方法。
Not All Error Yields to Scale: Where Scaling Stops in Vision-Language Inference
2610.01640v1 by Xinye Zhao, Yunkai Dang, Yunchen Wu, Wenbin Li
Vision-language models (VLMs) face a fixed-budget trade-off between processing more visual information for fine-grained perception and using a larger language backbone for complex reasoning. Existing studies do not tell us which combination of backbone size and input resolution to deploy, especially in high-resolution deployments. To address this gap, we propose the Separable Law that describes how VLM performance changes with language backbone size and visual token count. We fit the law to measurements from 26 InternVL and QwenVL models, with language backbone sizes from 1B to 72B, on four high-resolution benchmarks with image sizes from 224 pixels to 8K. We find that the questions responding to scaling can be predicted from the skill they require, while a substantial fraction never responds at all. We also find that the two model families gain similarly from a larger backbone, while their gains from more visual tokens differ sharply. Combined with a cost law, the Separable Law gives a closed-form rule for allocating compute between backbone size and visual tokens. When deployment is limited to available configurations, the law identifies model and image sizes that perform close to the best feasible choice under the same budget. We hope our work offers a principled way to decide how much a model should be allowed to see at high resolution, given what it must reason about.
摘要:視覺語言模型(VLMs)在處理更多視覺信息以實現細緻感知和使用更大語言骨幹以進行複雜推理之間面臨固定預算的權衡。現有研究並未告訴我們在高解析度部署中應該使用哪種骨幹大小和輸入解析度的組合。為了解決這一空白,我們提出了可分離法則,描述了VLM性能如何隨著語言骨幹大小和視覺標記數量的變化而變化。我們將該法則適配於來自26個InternVL和QwenVL模型的測量,這些模型的語言骨幹大小從1B到72B,並在四個高解析度基準上進行測試,圖像大小從224像素到8K。我們發現,對於擴展的問題,其反應可以根據所需的技能進行預測,而相當一部分則根本不會反應。我們還發現,這兩個模型系列在使用更大骨幹時獲益相似,而它們從更多視覺標記中獲得的收益則有明顯差異。結合成本法則,可分離法則提供了一個封閉形式的規則,用於在骨幹大小和視覺標記之間分配計算資源。當部署受限於可用配置時,該法則識別出在相同預算下表現接近最佳可行選擇的模型和圖像大小。我們希望我們的工作能提供一種原則性的方法,以決定模型在高解析度下應該被允許看到多少,考慮到它必須推理的內容。
Yo-ByT5: Efficient and High-Fidelity Diacritic Restoration for Yorùbá
2610.01634v1 by Ahmad Samuel Gali, Shamsuddeen Hassan Muhammad
Yorùbá is a widely spoken tonal language that depends on diacritics to avoid lexical ambiguity. However, it is often written without these diacritics, thereby hindering downstream Natural Language Processing (NLP) tasks. In this paper, we introduce Yo-ByT5, a byte-level Automatic Diacritic Restoration (ADR) model fine-tuned from ByT5-small. We evaluate Yo-ByT5 alongside five publicly released Yorùbá ADR models and one open-weight large language model (LLM) on the YAD benchmark under a consistent protocol. Our results demonstrate that Yo-ByT5 matches the performance of the strongest existing model, mT5-base, with a DER of 10.14% and a CER of 3.48%. Furthermore, it exhibits superior text fidelity despite using approximately half the parameter count of mT5-base. We also release our training code and model outputs, as well as call for the development of a larger, purpose-built benchmark for Yorùbá diacritic restoration.
摘要:Yorùbá 是一種廣泛使用的音調語言,依賴於變音符號以避免詞彙歧義。
然而,它經常在沒有這些變音符號的情況下書寫,從而妨礙了下游的自然語言處理 (NLP) 任務。
在本文中,我們介紹了 Yo-ByT5,一個從 ByT5-small 微調而來的字節級自動變音符號恢復 (ADR) 模型。
我們在 YAD 基準上,根據一致的協議,將 Yo-ByT5 與五個公開發布的 Yorùbá ADR 模型和一個開放權重的大型語言模型 (LLM) 進行評估。
我們的結果顯示,Yo-ByT5 的性能與現有最強模型 mT5-base 相當,具有 10.14% 的 DER 和 3.48% 的 CER。
此外,儘管使用的參數數量約為 mT5-base 的一半,但它在文本保真度上表現出色。
我們還發布了我們的訓練代碼和模型輸出,並呼籲開發一個更大、專門針對 Yorùbá 變音符號恢復的基準。
What Makes Something Hard(er)? Explaining Question Difficulty in Natural Language
2610.01627v1 by Peng Cui, Qiaoyuan Zheng, Rudolf Debelak, Mrinmaya Sachan
Difficulty is one of the most fundamental properties of a question: it determines whether the question can meaningfully discriminate between models of differing ability. Although a variety of methods can now estimate or predict difficulty automatically, they yield only a single descriptive number, with no account of the underlying factors that make a question difficult in the first place. In this work, we propose a data-driven approach that automatically generates and validates natural-language hypotheses explaining what makes one question harder than another. We first estimate each item's difficulty from the responses of a large pool of LLMs using Item Response Theory. We then sample contrasting sets of easy and hard questions and prompt an LLM to propose candidate explanations of the difference, which are subsequently validated and selected on held-out questions. Experimental results across three datasets spanning mathematical, logical, and commonsense reasoning show that our method produces interpretable and predictive hypotheses. On their own, they predict the difficulty of unseen questions competitively with, or better than, advanced black-box difficulty regressors; used as additional features, they further improve those regressors, implying that they discover difficulty signals that existing models fail to capture. Moreover, we demonstrate that editing questions according to a hypothesis can shift their measured difficulty in the expected direction, indicating that the discovered hypotheses are causally valid difficulty factors rather than post-hoc descriptions. Our approach thus turns a purely descriptive difficulty score into actionable statements.
摘要:困難度是問題最基本的特性之一:它決定了問題是否能夠有意義地區分不同能力的模型。雖然現在有多種方法可以自動估計或預測困難度,但它們僅產生一個描述性的數字,並未考慮使問題變得困難的潛在因素。在這項工作中,我們提出了一種數據驅動的方法,自動生成和驗證自然語言假設,解釋為什麼一個問題比另一個問題更難。我們首先使用項目反應理論從大量大型語言模型的回應中估計每個項目的困難度。然後,我們抽取一組對比的簡單和困難問題,並提示一個大型語言模型提出候選解釋這些差異,這些解釋隨後在保留的問題上進行驗證和選擇。跨越數學、邏輯和常識推理的三個數據集的實驗結果顯示,我們的方法產生了可解釋且具有預測性的假設。僅憑這些假設,它們能夠與先進的黑箱困難回歸模型競爭地預測未見問題的困難度;作為額外特徵使用時,它們進一步改善了這些回歸模型,這意味著它們發現了現有模型未能捕捉的困難信號。此外,我們證明根據假設編輯問題可以將其測量的困難度朝預期方向轉變,這表明所發現的假設是因果有效的困難因素,而非事後描述。因此,我們的方法將純粹描述性的困難分數轉化為可行的陳述。
FedLore: Communication and Memory Efficient Federated Learning via Shared Gradient Low-Rank Projection
2610.01620v1 by Junkang Liu
Federated training of foundation models is constrained by client memory and communication costs. LoRA-based methods reduce these costs through low-rank adapters, but their fixed rank budget can limit adaptation. Gradient low-rank optimization offers greater flexibility, yet independently chosen client subspaces create a problem we term \emph{subspace fragmentation}: local projections interact with data heterogeneity to bias aggregated directions, while aggregation can increase update rank and communication cost. Thus, accurate local gradient compression need not preserve global descent. We propose \texttt{FedLore}, which shares a low-rank optimization basis within each round and refreshes it across rounds. The shared basis enables exact aggregation in low-rank coordinates and eliminates the identified projection bias. Subspace refresh allows the accumulated model update to exceed the per-round rank budget. We characterize the aggregation bias and establish an $O(T^{-1/2})$ stationarity bound for the projected-SGD variant under a global-gradient coverage condition and standard smoothness and variance assumptions, with bounded gradient heterogeneity. Experiments on vision and language tasks, including federated pre-training, show that \texttt{FedLore} outperforms the evaluated low-rank adapter baselines and matches or exceeds full-parameter training, while reducing communication and optimizer-state memory.
摘要:聯邦訓練基礎模型受到客戶端記憶體和通信成本的限制。基於LoRA的方法通過低秩適配器降低這些成本,但其固定的秩預算可能限制適應性。梯度低秩優化提供了更大的靈活性,但獨立選擇的客戶端子空間會產生我們稱之為\emph{subspace fragmentation}的問題:局部投影與數據異質性相互作用,偏向於聚合方向,而聚合可能會增加更新秩和通信成本。因此,準確的局部梯度壓縮不必保留全局下降。我們提出\texttt{FedLore},在每一輪中共享低秩優化基礎,並在輪與輪之間進行刷新。共享基礎使得在低秩坐標中進行精確聚合,並消除了已識別的投影偏差。子空間刷新允許累積的模型更新超過每輪的秩預算。我們描述了聚合偏差,並在全局梯度覆蓋條件及標準平滑性和方差假設下,對投影-SGD變體建立了$O(T^{-1/2})$的平穩性界限,並且具有有界的梯度異質性。在視覺和語言任務上的實驗,包括聯邦預訓練,顯示\texttt{FedLore}的表現超過了評估的低秩適配器基準,並且與全參數訓練相匹配或超過,同時減少了通信和優化器狀態記憶體。
Exposing the Cost of Deep Learning Audio Development
2610.01619v1 by Constance Douwes, Paul Magron, Romain Serizel
The environmental impact of deep learning has attracted increasing attention over the past decade. Existing studies mainly focus on the energy and carbon emissions of model training and inference, while the whole development phase is often overlooked. Yet, architecture prototyping and intensive experiments are conducted during this stage, which is highly energy-demanding. In this article, we propose a methodology to estimate these costs, based on activity logs from the Grid5000 shared computing platform used by the LORIA laboratory. As a case-study, we focus on audio projects developed in the Multispeech research team. We evaluate the overall energy cost of four projects, and we compare them to those of training the reported models. Our results show that the energy required for the development phase is 3 to 256 times greater than that required to train the best-performing model alone. These results advocate for a more systematic reporting of energy consumption across the entire life cycle of deep learning-based audio projects.
摘要:深度學習的環境影響在過去十年中引起了越來越多的關注。
現有的研究主要集中在模型訓練和推理的能量和碳排放上,而整個開發階段常常被忽視。
然而,在這個階段進行架構原型設計和密集實驗,這是非常耗能的。
在本文中,我們提出了一種基於LORIA實驗室使用的Grid5000共享計算平台的活動日誌來估算這些成本的方法。
作為案例研究,我們專注於Multispeech研究團隊開發的音頻項目。
我們評估了四個項目的整體能量成本,並將其與訓練報告模型的能量成本進行比較。
我們的結果顯示,開發階段所需的能量是僅訓練最佳性能模型所需能量的3到256倍。
這些結果提倡對基於深度學習的音頻項目整個生命週期的能量消耗進行更系統的報告。
Agents Are Systems, Not Models: Rethinking Agentic Evaluation
2610.01618v1 by Luis Wiedmann, Leander Girrbach, Cordelia Schmid, Zeynep Akata
Agent evaluations increasingly go beyond a single success rate, reporting metrics such as cost, consistency, and robustness. Yet they typically treat the agent itself as fixed. In practice, an agent is a configurable system: users decide what to tell it, how long to let it run, and which model to use, and each of these choices can change how well and how consistently it performs. We study these choices on a new benchmark of four scientific tasks, where a coding agent must find and correctly operate a published specialist model. We investigate five parts of the agent's configuration: task information, reasoning, self-verification, time budget, and backbone model. We find substantial run-to-run variability, with approximately 54% of the outcome variance coming from repeating the same configuration rather than changing it. Across configurations, the information provided to the agent has the largest effect, exceeding both time budget and model size, while also reducing cost and improving calibration. Configuration choices also interact: additional time helps only when the agent has sufficient information or a capable enough model to use it. Finally, a trajectory-based taxonomy of agent behavior reveals that prompting an agent to verify its answer has little effect on its verification behavior, whereas providing a dedicated verification tool changes that behavior substantially. These results suggest that agents should be evaluated as configurable systems themselves, and that some desired behaviors are more effectively implemented in the system than requested through prompting. We release the benchmark and more than 18,000 agent trajectories.
摘要:代理評估越來越超越單一的成功率,報告如成本、一致性和穩健性等指標。 然而,它們通常將代理本身視為固定的。 實際上,代理是一個可配置的系統:用戶決定告訴它什麼、讓它運行多久以及使用哪個模型,而這些選擇都會改變它的表現效果和一致性。 我們在一個新的基準上研究這些選擇,該基準包含四個科學任務,其中一個編碼代理必須找到並正確操作一個已發表的專家模型。 我們調查了代理配置的五個部分:任務信息、推理、自我驗證、時間預算和骨幹模型。 我們發現運行之間存在顯著的變異性,大約54%的結果變異來自重複相同的配置,而不是改變它。 在不同配置中,提供給代理的信息具有最大的影響,超過了時間預算和模型大小,同時還降低了成本並改善了校準。 配置選擇之間也存在相互作用:額外的時間僅在代理擁有足夠的信息或足夠能力的模型來使用時才有幫助。 最後,基於軌跡的代理行為分類法顯示,促使代理驗證其答案對其驗證行為幾乎沒有影響,而提供專用的驗證工具則會顯著改變該行為。 這些結果表明,代理應該被評估為可配置的系統本身,而某些期望的行為在系統中實現的效果比通過提示請求更有效。 我們發布了基準和超過18,000條代理軌跡。
Can LLMs Reliably Annotate Bioassay Metadata to Improve Data Readiness?
2610.01616v1 by Laura van Weesep, Riccardo Tedoldi, Jens Sjölund, Hossein Azizpour, Susanne Winiwarter, Ola Engkvist, Jon Paul Janet, Samuel Genheden, Juan Viguera Diez
The emergence of foundation models for molecular property prediction requires a high degree of AI data readiness, including reliable metadata annotation. However, both public repositories and industrial screening databases suffer from missing, inconsistent, or conflated assay annotations. In this work, we quantify the extent of missing annotations in PubChem for the BioAssay Ontology (BAO) assay format and physical detection method fields and investigate whether open-source and proprietary large language models (LLMs) can reliably predict and audit metadata annotations directly from the assay text. In our assessment, we found that the annotation coverage across PubChem's $\sim$2 million bioassays is critically sparse, 36\% lacking an assay format, 89\% a BioAssay type, and >99.9\% any BAO-mapped assay format or detection technology term. This motivates the need for automated test-metadata curation. Using evaluation sets derived from PubChem and ChEMBL, we assess the agreement of seven open-source and proprietary LLMs with existing silver labels. Recall is at least 0.96 for biochemical and cell-based assay formats, with a similar pattern for detection technology, although disagreements increase on under-represented classes. Manual inspection shows that many of these disagreements trace back to inconsistencies between silver sources rather than to LLM error. Moreover, in a qualitative study with a senior industrial curator, LLM-generated evidence prompted the expert to revise some of their own labels, showing LLMs can flag potentially mislabeled assays. Across the study, performance differences between proprietary and open-source models were small. Together, these results suggest LLMs can support the large-scale annotation and auditing of assay metadata, though per-class reliability estimates and targeted human review remain necessary before such labels enter downstream ML pipelines.
摘要:基於分子性質預測的基礎模型的出現需要高度的人工智慧數據準備,包括可靠的元數據註釋。然而,公共資料庫和工業篩選數據庫都存在缺失、不一致或混淆的檢測註釋。在這項工作中,我們量化了PubChem中BioAssay本體(BAO)檢測格式和物理檢測方法字段缺失註釋的程度,並調查開源和專有大型語言模型(LLMs)是否能夠可靠地從檢測文本中直接預測和審核元數據註釋。在我們的評估中,我們發現PubChem約200萬個生物檢測的註釋覆蓋率極其稀疏,36\%缺乏檢測格式,89\%缺乏BioAssay類型,且超過99.9\%缺乏任何BAO映射的檢測格式或檢測技術術語。這促使了自動測試元數據整理的需求。利用從PubChem和ChEMBL衍生的評估集,我們評估了七個開源和專有LLMs與現有銀標籤的一致性。生化和基於細胞的檢測格式的召回率至少為0.96,檢測技術的模式相似,儘管在代表性不足的類別上分歧增加。手動檢查顯示,許多這些分歧源於銀來源之間的不一致,而不是LLM的錯誤。此外,在與一位資深工業策展人的定性研究中,LLM生成的證據促使專家修訂他們的一些標籤,顯示LLMs可以標記潛在錯誤標記的檢測。在整個研究中,專有模型和開源模型之間的性能差異很小。綜合這些結果表明,LLMs可以支持檢測元數據的大規模註釋和審核,儘管在這些標籤進入下游機器學習管道之前,仍然需要每類的可靠性評估和針對性的人工審查。
Hob-VL: A Benchmark for Visually Grounded Boolean Reasoning
2610.01605v1 by Yuzhou Wang, Emile Anand, Ijay Narang
Reliable visual reasoning requires composing multiple visual observations and returning consistent answers to logically equivalent questions. We introduce Hob-VL, a benchmark for visually grounded Boolean reasoning. Hob-VL comprises two tasks: (1) evaluating whether a Boolean rule holds in an image, and (2) identifying the (unique) object satisfying a Boolean description. Hob-VL contains 6,000 human-verified balanced Yes/No questions, each defined by a Boolean combination of ten visual statements, across 1,000 generated scenes and 46 diverse labeled photographs, along with 1,000 object-identification questions over the same photographs. Our question families are deliberately constructed to challenge reasoning through misleading local cues and nested logical operations, and include symbolic and structured natural-language presentations. Across eight model configurations with thinking disabled or minimized, Boolean accuracy ranges from 48.52% to 50.57%, while the identification accuracy reaches at most 43.0%. A thinking-enabled GLM configuration achieves uneven gains while retaining substantial errors and inconsistencies. Hob-VL exposes these failures through executable reference answers and matched evaluations.
摘要:可靠的視覺推理需要組合多個視覺觀察,並對邏輯上等價的問題給出一致的答案。 我們介紹了 Hob-VL,一個針對視覺基礎布林推理的基準。 Hob-VL 包含兩個任務:(1)評估布林規則在圖像中是否成立,以及(2)識別滿足布林描述的(唯一)物體。 Hob-VL 包含 6,000 個經過人工驗證的平衡是/否問題,每個問題由十個視覺陳述的布林組合定義,涵蓋 1,000 個生成場景和 46 張多樣化的標記照片,以及 1,000 個針對相同照片的物體識別問題。我們的問題家族故意構建,以通過誤導性的局部提示和嵌套邏輯操作來挑戰推理,並包括符號和結構化的自然語言呈現。在八種思考被禁用或最小化的模型配置中,布林準確率範圍從 48.52% 到 50.57%,而識別準確率最多達到 43.0%。 一個啟用思考的 GLM 配置在保留相當大的錯誤和不一致性的同時,實現了不均勻的增益。 Hob-VL 通過可執行的參考答案和匹配的評估揭示了這些失敗。
Permutation-Robust Decision Modeling with Candidate-Independent Block-Causal Attention
2610.01601v1 by Guy Amit
Decision models often score a variable-sized set of candidate actions encoded in a single sequence. This setting is increasingly relevant for System 1 components inside generative systems, where candidates may be proposed or ordered differently across runs. Standard causal cross-encoding is expressive, but it can make a candidate's score depend on serialization order rather than on the underlying decision problem. We introduce candidate-independent block-causal attention, which preserves causal computation within the shared context and each candidate while blocking cross-candidate information flow and resetting candidate positions. We compare this architecture with standard causal attention and complementary invariant baselines across Gemma 3 1B, Qwen3 1.7B, and Qwen3 4B backbones. Candidate-independent attention consistently reduces permutation sensitivity while retaining competitive decision quality; ablations indicate that candidate isolation is the primary source of the effect, with position resetting completing the intended symmetry. A larger Qwen3-4B study further examines the behavior of the proposed architecture with substantially more training data. Code is available at the \href{https://github.com/guyAmit/ci-decision-models}{\textcolor{blue}{project repository}}, and the \href{https://huggingface.co/Guy-Amit/qwen3-4b-ci-decision-4096-poc}{\textcolor{blue}{Qwen3-4B model artifact}} is available on Hugging Face.
摘要:決策模型通常對編碼在單一序列中的可變大小候選行動集進行評分。這種設定對於生成系統中的系統1組件越來越相關,其中候選者在不同運行中可能以不同方式被提議或排序。標準的因果交叉編碼具有表達能力,但它可能使候選者的分數依賴於序列化順序,而不是基於底層的決策問題。我們引入了候選者獨立的區塊因果注意力,該方法在共享上下文和每個候選者內部保留因果計算,同時阻止候選者之間的信息流動並重置候選者位置。我們將這種架構與標準因果注意力和互補不變基準進行比較,涵蓋了Gemma 3 1B、Qwen3 1.7B和Qwen3 4B主幹。候選者獨立的注意力在保持競爭性決策質量的同時,始終減少排列敏感性;消融實驗表明,候選者隔離是該效應的主要來源,位置重置則完成了預期的對稱性。一項更大規模的Qwen3-4B研究進一步檢查了所提議架構在大量訓練數據下的行為。代碼可在\href{https://github.com/guyAmit/ci-decision-models}{\textcolor{blue}{項目庫}}獲得,而\href{https://huggingface.co/Guy-Amit/qwen3-4b-ci-decision-4096-poc}{\textcolor{blue}{Qwen3-4B模型文物}}則可在Hugging Face上獲得。
Before It Fades: Reinforcing Temporal Representations at Inference Time in VideoLLMs
2610.01595v1 by Youngwoo Shin, Yusung Ro, Minseo Kim, Junmo Kim
Video Large Language Models (VideoLLMs) receive frames in sequential order and interpret how visual content evolves along the temporal axis, yet temporal reasoning remains a persistent weakness across architectures. Reversing the frame order of a video, a transformation that should invert temporal answers, often leaves the final prediction unchanged. We investigate where this failure originates by defining the temporal divergence vector $τ_l$, the layer-wise representational difference induced by reversing temporal order. Tracking its magnitude across layers reveals a consistent temporal divergence profile where the divergence peaks at intermediate layers and progressively diminishes toward the output. We confirm this peak is specific to temporal reasoning and functionally critical for predictions, establishing that VideoLLMs acquire temporal information at intermediate layers but fail to maintain it to the output. This progressive fading motivates our method, Temporal Activation Injection (TAI), which extracts $τ_l$ at the peak of the profile for each input and reinjects it into subsequent layers following the measured decay. TAI requires no training and consistently improves temporal reasoning across three VideoLLMs and four benchmarks with negligible impact on non-temporal tasks. Code is available at https://github.com/Youngwoo-git/Before-It-Fades.
摘要:視頻大型語言模型(VideoLLMs)以順序方式接收幀並解釋視覺內容如何沿時間軸演變,然而,時間推理在各種架構中仍然是一個持久的弱點。反轉視頻的幀順序,這一轉換應該會顛倒時間答案,但最終預測往往保持不變。我們通過定義時間發散向量 $τ_l$ 來調查這一失敗的來源,這是由反轉時間順序引起的層級表示差異。跟踪其在各層的大小揭示了一個一致的時間發散輪廓,其中發散在中間層達到峰值,並逐漸減少到輸出。我們確認這一峰值是特定於時間推理的,並且對預測至關重要,確立了 VideoLLMs 在中間層獲取時間信息但未能將其保持到輸出的事實。這一漸進的衰減激發了我們的方法——時間激活注入(Temporal Activation Injection, TAI),該方法在每個輸入的輪廓峰值處提取 $τ_l$,並在隨後的層中重新注入,根據測量的衰減進行。TAI 不需要訓練,並在三個 VideoLLMs 和四個基準測試中一致改善時間推理,對非時間任務的影響微乎其微。代碼可在 https://github.com/Youngwoo-git/Before-It-Fades 獲得。
Which LLM to pick? Online Active Model Selection for Large Language Models
2610.01592v1 by Alessandro Turrin, Patrik Okanovic, Torsten Hoefler, Nezihe Merve Gürel
Large Language Models (LLMs) are increasingly applied to process streaming data, with practitioners relying on benchmarks to select the best model even though these signals only approximate real performance. While oracle annotations can provide reliable feedback, they are often costly and difficult to obtain at scale. To address this challenge, we propose ONLINE LLM PICKER, the first framework for active model selection for LLMs in online settings. Given an arbitrary stream of queries and a limited annotation budget, ONLINE LLM PICKER selects the most informative prompts for annotation to identify the best LLM among candidate models. Across multiple tasks including 10 datasets, for over 130 language models, we show that ONLINE LLM PICKER saves annotation cost by up to 71.67% while reliably identifying the best or near-best model for the stream. We also show that using the returned model for sequential generation on unannotated prompts across the stream reduces regret by up to a factor of 2.51x, indicating that ONLINE LLM PICKER can identify the best or near-best model well before processing all streaming prompts.
摘要:大型語言模型(LLMs)越來越多地應用於處理串流數據,實踐者依賴基準來選擇最佳模型,即使這些信號僅能近似實際性能。雖然預言者註釋可以提供可靠的反饋,但它們通常成本高昂且難以大規模獲得。為了解決這一挑戰,我們提出了ONLINE LLM PICKER,這是第一個針對在線環境中LLMs的主動模型選擇框架。給定一個任意的查詢串流和有限的註釋預算,ONLINE LLM PICKER選擇最具信息量的提示進行註釋,以識別候選模型中最佳的LLM。在包括10個數據集的多個任務中,針對超過130個語言模型,我們顯示ONLINE LLM PICKER將註釋成本降低了高達71.67%,同時可靠地識別出串流中的最佳或接近最佳模型。我們還顯示,使用返回的模型在串流中的未註釋提示上進行順序生成,將後悔值降低了高達2.51倍,這表明ONLINE LLM PICKER能夠在處理所有串流提示之前就識別出最佳或接近最佳的模型。
Evaluating Physical Consistency and Plausibility in Generative Scenario Models for Autonomous Driving
2610.01581v1 by Manasa Mariam Mammen, Zafer Kayatas, Stefan Wagner
Generative AI models are increasingly used for scenario generation in autonomous driving. While they can generate realistic-looking scenarios, they often provide limited transparency into learned representations and consistency with real-world vehicle dynamics. This lack of formal assurance limits their use in safety-critical validation and certification workflows. To address this aspect, we introduce a layered evaluation protocol that complements existing methods by assessing models across five layers. The first four layers inspect internal representations and network layers through kinematic alignment, statistical baseline comparison, latent controllability, and activation analysis. The fifth layer evaluates model outputs against vehicle dynamics constraints such as lateral jerk thresholds. We demonstrate the protocol on a Variational Autoencoder (VAE)-based scenario generator. Although standard output-level metrics and visualizations suggest that the generated scenarios are realistic, our protocol provides deeper insight into the extent to which the model's latent space aligns with kinematic features and whether visually plausible trajectories satisfy vehicle-dynamics constraints. We further apply the protocol to additional generative models, demonstrating its applicability beyond the VAE architecture.
摘要:生成式人工智慧模型在自動駕駛的場景生成中被越來越多地使用。
雖然它們可以生成看起來現實的場景,但通常對於學習到的表徵和與現實世界車輛動態的一致性提供的透明度有限。
這種缺乏正式保證的情況限制了它們在安全關鍵的驗證和認證工作流程中的使用。
為了解決這個問題,我們提出了一種分層評估協議,通過在五個層面上評估模型來補充現有方法。
前四個層面通過運動學對齊、統計基準比較、潛在可控性和激活分析檢查內部表徵和網絡層。
第五個層面則評估模型輸出是否符合車輛動態約束,例如橫向加速度閾值。
我們在一個基於變分自編碼器(VAE)的場景生成器上演示了這一協議。
雖然標準的輸出級別指標和可視化顯示生成的場景是現實的,但我們的協議提供了更深入的見解,以了解模型的潛在空間在多大程度上與運動學特徵對齊,以及視覺上合理的軌跡是否滿足車輛動態約束。
我們進一步將該協議應用於其他生成模型,展示其在超越VAE架構的適用性。