Medical explainable AI
Medical explainable AI
| Publish Date | Title | Authors | Homepage | Code |
|---|---|---|---|---|
| 2026-10-01 | MIRTO: a registration-gated, multiverse-tested evaluation protocol for unsupervised anomaly segmentation in brain MRI | Negin Kafee Hernashki et.al. | 2610.02136v1 | null |
| 2026-10-01 | A Comparative Explainability Framework for DeBERTa-v3 in Zero-Shot Medical Abstract Classification | Javier Diaz Esteban-Herreros et.al. | 2610.02116v1 | null |
| 2026-10-01 | Can AI Oversight Be Zero Knowledge? | Alessandro Chiesa et.al. | 2610.01995v1 | null |
| 2026-10-01 | Mapping the RAG Landscape: A Four Axis Taxonomy of Efficiency, Defense, Interactivity, and Reasoning | Meghana Sunil et.al. | 2610.01936v1 | null |
| 2026-10-01 | Removing spurious minima for planar features by skip connections | Jakob Paul Zimmermann et.al. | 2610.01728v1 | null |
| 2026-10-01 | MCIR: A Feature Dependence-Aware Explainability Method with Reliability Guarantees | Poushali Sengupta et.al. | 2610.01641v1 | null |
| 2026-10-01 | What Makes Something Hard(er)? Explaining Question Difficulty in Natural Language | Peng Cui et.al. | 2610.01627v1 | null |
| 2026-10-01 | Measuring the Stability Assumption Behind Action Chunking | Aryan Goyal et.al. | 2610.01626v1 | null |
| 2026-10-01 | Exact Distinguishability in Non-Markovian Decision Processes | Kabir Murjani et.al. | 2610.01527v1 | null |
| 2026-10-01 | OpenMTB-Audit: Exposing Over-Refusal and Clinical Expert Perspectives in LLM-Based Molecular Tumor Board Safety Evaluation | Negin Ashrafi et.al. | 2610.01497v1 | null |
| 2026-10-01 | Multi-Party Backchannel Prediction: a Diagnosis, a Benchmark, and a Ceiling | Mohammed Hafsati et.al. | 2610.01488v1 | null |
| 2026-10-01 | Detect, Explain, Interpret: An End-to-End Benchmark for Time Series Anomaly Detection, Explainability and Interpretability | Roberto Stanzione et.al. | 2610.01168v1 | null |
| 2026-10-01 | CineMR: Tool-Integrated Vision-Language Reasoning for Quantitative Cardiac MRI Assessment | Kunyang Li et.al. | 2610.01166v1 | null |
| 2026-10-01 | What Can Analogy Tell Us About Artificial Consciousness? | Keith J. Holyoak et.al. | 2610.01002v1 | null |
| 2026-09-30 | When Reasoning Helps Action: Monitoring and Steering Chain-of-Thought in Vision-Language-Action Policies | Sathwik Karnik et.al. | 2610.00601v1 | null |
| 2026-09-30 | Worse Together: How Performance Breaks Down in Multi-User Multi-Agent Teams | Sahan Paliskara et.al. | 2610.00583v1 | null |
| 2026-09-30 | No One Architecture Fits All: A Cross-Environment Evaluation of Hierarchical Red Team Agents | Ayan Javeed Shaikh et.al. | 2610.00557v1 | null |
| 2026-09-30 | EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights | Jiayi Geng et.al. | 2610.00492v1 | null |
| 2026-09-30 | CAS II: Symmetric Partitions as Kolmogorov Models | Romie Banerjee et.al. | 2609.40290v1 | null |
| 2026-09-30 | Unlearnable, or Unmeasured? On the Reliability of Difficulty Labels in RLVR | Chandak Chakma et.al. | 2609.40115v1 | null |
| 2026-09-30 | What Can Component-Replacement Evidence Establish? A Critical Scoping Review of Local Decisions in LLM Agents | Shuyang Zhang et.al. | 2609.39989v1 | null |
| 2026-09-30 | How Does Local Landscape Geometry Evolve in Language Model Pre-Training? | Zhanpeng Zhou et.al. | 2609.39767v1 | null |
| 2026-09-30 | Trust Is Not a Score: Runtime Assurance Contracts for High-Risk AI Agents | Serhii Zabolotnii et.al. | 2609.39717v1 | null |
| 2026-09-30 | ChronoGraph: Functional 4D Scene Graphs with Vision-Language Models for Interaction Understanding and Grounded Planning | Chenyangguang Zhang et.al. | 2609.39665v1 | null |
| 2026-09-30 | Zero-Compute Cross-Lingual Transferability Estimation Using Typological Feature Proxies | Dalton Raphael Harmsen et.al. | 2609.39640v1 | null |
| 2026-09-30 | Disentangling Self-Distillation: Measuring and Modeling Acquisition and Retention | Luis Zuin et.al. | 2609.39494v1 | null |
| 2026-09-30 | From Speech to Editable Concepts: Probing Emotion Recognition with Concept Bottleneck Models | Hezhao Zhang et.al. | 2609.39453v1 | null |
| 2026-09-30 | Towards Trustworthy AI for Glioma Diagnosis: A Task-Aware Evaluation of Uncertainty Quantification | Gonzalo Esteban Mosquera Rojas et.al. | 2609.39429v1 | null |
| 2026-09-30 | The Golden Path Hypothesis: Reusable Schedules in Diffusion Caching | Dong Wang et.al. | 2609.39343v1 | null |
| 2026-09-30 | A Time-Aware Bag-of-Receptive-Fields for Interpretable Irregular Time Series Classification | Francesco Spinnato et.al. | 2609.39268v1 | null |
| 2026-09-30 | When Harnesses Lose the Signal: Causal Evaluation of Recovery in LLM Agents | Shuyao Xiao et.al. | 2610.00372v1 | null |
| 2026-09-30 | Loop-Free Inverse Reinforcement Learning via Sequential Value Recovery with Q-Score Matching | Yang chen et.al. | 2609.38955v1 | null |
| 2026-09-30 | Reasoning Externalization for Faithful Large Language Model Narratives of Stock Return Predictions | Sujung Kim et.al. | 2609.38869v1 | null |
| 2026-09-30 | SkillSeek: Revisiting Agent Skill Retrieval at Marketplace Scale | Guanqun Yang et.al. | 2609.38822v1 | null |
| 2026-09-30 | When Reasoning Goes Astray: Attention Dynamics of Uncontrolled Reasoning | Yuanhe Zhang et.al. | 2609.38817v1 | null |
| 2026-09-29 | Understanding Off- vs On-Policy Distillation: A Tale of Distinct Training Objectives | Qiwei Di et.al. | 2609.38666v1 | null |
| 2026-09-29 | Defining and Categorising Human-AI Interactions in Clinical Trials: A Multidimensional Human-AI Classification Approach | Sandra Woolley et.al. | 2609.38559v1 | null |
| 2026-09-29 | Demographic Pluralism: Inference-Time Modeling of Pluralistic Human Preference Distributions | Meng-Chen Wu et.al. | 2609.38555v1 | null |
| 2026-09-29 | Does This Action Still Explain the Task? Reverse Scoring for Diffusion Language Model Agents | Jiacheng Qiu et.al. | 2609.38536v1 | null |
| 2026-09-29 | Caption-Mediated Perceived-Safety Estimation for Pedestrian Routing | Simon Parkinson et.al. | 2609.38479v1 | null |
| 2026-09-29 | Aegis: Generative Gradient Masking for Privacy-Preserving Medical Federated Learning | Chaoyu Zhang et.al. | 2609.38339v1 | null |
| 2026-09-29 | Does Local Video Understanding Transfer Across Encounters? The EgoGears Benchmark | Yuedong Tan et.al. | 2609.37938v1 | null |
| 2026-09-29 | OmniVCBench: Benchmarking Evidence-Grounded Multimodal Reasoning Towards AI Virtual Cells | Manyu Li et.al. | 2609.37773v1 | null |
| 2026-09-29 | Multi-Site Real-World Performance of Commercial AI for Pulmonary and Incidental Pulmonary Embolism Detection | Aawez Mansuri et.al. | 2609.37750v1 | null |
| 2026-09-29 | XU-RS: Explaining Credal Width in Random-Set Language Models | David Achara et.al. | 2609.37594v1 | null |
| 2026-09-29 | Raw Imagery Impacting Your AI: Should You Care? | Adrien Dorise et.al. | 2609.38265v1 | null |
| 2026-09-29 | Selecting The Most Informative Tokens in Natural Language Autoencoders | Federico Torrielli et.al. | 2609.37040v1 | null |
| 2026-09-29 | Harness Evolution as Learning: Approximation, Generalization, and Optimization Limits of Self-Improving Personal Agents | Zeyu Gan et.al. | 2609.36892v1 | null |
| 2026-09-29 | Predictive Credit: Measuring What Scientific Explanations Add to Experimental Forecasts | Jingjie Ning et.al. | 2610.00314v1 | null |
| 2026-09-28 | Engineering Simplicity: Simple Mechanism Interfaces Steer LLM Agents | Kehang Zhu et.al. | 2609.36365v1 | null |
| 2026-09-28 | Explainability from Training with Applications to TCR-Epitope Prediction | Jiarui Li et.al. | 2609.36354v1 | null |
| 2026-09-28 | ThuRunel: Dynamic Decoupling for Structured Advisory Dialogue | Yuyan Chen et.al. | 2609.36340v1 | null |
| 2026-09-28 | FigAct: Turning Scientific Figures into Active Canvases for Explanation | Shishi Xiao et.al. | 2609.36190v1 | null |
| 2026-09-28 | An Exact Generate - Transform Decomposition of Small-LLM Team Scaling Across Orchestration Architectures | Blaz Bertalanic et.al. | 2609.36104v1 | null |
| 2026-09-28 | One Geometry, Different Outcomes: Readout-Dependent Effects of the Modality Gap in Vision-Language Models | Aditya Sharma et.al. | 2609.36101v1 | null |
| 2026-09-28 | Shockingly Simple Self-retrospection Improves Agentic Models Without RL | Jonathan Light et.al. | 2609.35741v1 | null |
| 2026-09-28 | Rethinking Circuit Evaluation: Do Circuits Explain Model Errors? | Li Zhang et.al. | 2609.35686v1 | null |
| 2026-09-28 | Signatures of semantic search in the activations of large language models | Luke Leckie et.al. | 2609.35599v2 | null |
| 2026-09-28 | A decision-support system applied to Law: Reasoning and explainability of the decision | Jeremy Bouche-Pillon et.al. | 2609.35370v1 | null |
| 2026-09-28 | Jev thinks "I don't know'', but doesn't say it: Introducing Sys1Cal-v1 Dataset for Probability Calibration | Riccardo Porcedda et.al. | 2609.35342v1 | null |
| 2026-09-28 | The Argument and the Letterhead: Source-Position Coherence in AI Evaluation | Michele Loi et.al. | 2609.35286v1 | null |
| 2026-09-28 | Towards Reliable AI Data Scientists: Data Agents with Workflow Harnesses | Huachi Zhou et.al. | 2609.35255v1 | null |
| 2026-09-28 | Beneath the Tokens: A Performance Engineering Study of Multi-Token Prediction in GPU-Accelerated LLM Inference | Suwesh Prasad Sah et.al. | 2609.35188v1 | null |
| 2026-09-28 | Applying Language Models in Clinical Medicine: Recent Trends and Perspectives | Erik Aerts et.al. | 2609.34780v2 | null |
| 2026-09-28 | A Unifying Framework of Concept-based Explainable AI with Completeness Guarantees | Vojtěch Kůr et.al. | 2609.34750v1 | null |
| 2026-09-28 | From Human Narrative to Harmonic Structure: A Human-Centered Investigation of Algorithmic Music Generation through the Chord Wheel Diagram | Josef Pavlíček et.al. | 2609.34735v1 | null |
| 2026-09-28 | Understanding Generalization Requires Universal Induction | Aram Ebtekar et.al. | 2609.34458v1 | null |
| 2026-09-28 | Social Circuits behind Multi-agent Echo Chambers | Chuiyang Meng et.al. | 2609.34444v1 | null |
| 2026-09-28 | Improving Large Language Models for Code through Runtime Program-State Reasoning | Hongwei Li et.al. | 2609.34359v1 | null |
| 2026-09-28 | Dynamical Parameters: An Interpretability Framework for Time-Series Foundation Models | Kang Yang et.al. | 2609.34316v1 | null |
| 2026-09-28 | Evo2Team: When Do Evolved Skills Transfer? From Selection to Deployment | Renxiang Wang et.al. | 2609.34135v1 | null |
| 2026-09-28 | JET: Judge-Guided Evolution at Test Time for Agent Programs | Yao Long Teng et.al. | 2609.34126v1 | null |
| 2026-09-28 | Do World Models Learn Global Understanding? | Alexander Detkov et.al. | 2609.34058v1 | null |
| 2026-09-27 | Uncovering shortcut learning in audio classifiers by discovering recurring concepts in temporal explanations | Cecilia Bolaños et.al. | 2609.34030v1 | null |
| 2026-09-27 | Safety Reconstructed: Generative Modeling via Masked Diffusion Builds Strong Safety Guardrails | Gert Lek et.al. | 2609.33634v1 | null |
| 2026-09-27 | Learning Dynamics of Continual Learning: A Unified View of Data Attribution, Forgetting, and Plasticity Loss | Yi Ren et.al. | 2609.33620v1 | null |
| 2026-09-27 | Temporal Graph Learning of Wearable Actigraphy and Sleep Traces for Modelling Adolescent Crystallized Intelligence | Md. Tanvir Rahman et.al. | 2609.33428v1 | null |
| 2026-09-27 | Explainable Deep Learning of Resting-State Functional Connectomes Reveals Network Biomarkers of Adolescent Intelligence | Md. Tanvir Rahman et.al. | 2609.33422v1 | null |
| 2026-09-27 | Decoupling Token Roles in Autoregressive Pretraining | Suqin Yuan et.al. | 2609.33405v1 | null |
| 2026-09-27 | The Error You See Is Not the Error You Made: Progression-aware Reasoning Origin for Reasoning Error Localization | Yiguo Wang et.al. | 2609.33297v1 | null |
| 2026-09-27 | CORTEX: A Verified Experience Layer for Generalist Agents | Garapati Keerthana et.al. | 2609.33260v1 | null |
| 2026-09-27 | FOCUS: Benchmarking Retinal Model Generalization from Foundation Vision Encoders to Multimodal LLMs | David Restrepo et.al. | 2609.33158v1 | null |
| 2026-09-26 | Relative Generalization Invariance of LLM Pretraining | Fengzhuo Zhang et.al. | 2609.33016v1 | null |
| 2026-09-26 | DynamicDx: Evaluating Evidence Acquisition in Video-Based Diagnosis | Jiahui Li et.al. | 2609.32957v1 | null |
| 2026-09-26 | Counterfactual Self-Evolving Agents for Evidence-Grounded Reasoning | Xing Han et.al. | 2609.32870v1 | null |
| 2026-09-26 | FinancialAuditBench: Benchmark Construction under Differential Privacy Using Real-World Priors | Jerry Huang et.al. | 2609.32835v1 | null |
| 2026-09-26 | Mandela-Bench: Multimodal Models Remember Canonical Images Instead of Seeing Them | Yicheng Bao et.al. | 2609.32763v1 | null |
| 2026-09-26 | What Would Falsify It? A Variable Specific Evidence Standard for Mechanistic Claims About Self Explanation | Arshia Eftekhari zadeh et.al. | 2609.32670v1 | null |
| 2026-09-26 | When Users Change Their Minds: Measuring and Repairing Intent Drift in LLM Agents | Yanjie Zhang et.al. | 2609.32520v1 | null |
| 2026-09-26 | Explaining Textual Entailment with Lexical Entailments: Using LLMs to Supply Lexical Relations for Formal Proofs | Jorryt de Jong et.al. | 2609.32491v1 | null |
| 2026-09-26 | Superposed Inference for Hyperdimensional Computing | Quanling Zhao et.al. | 2609.32320v1 | null |
| 2026-09-26 | Why Directly Learning Periodic Trajectories Can Fail | Kaixin Zheng et.al. | 2609.32254v1 | null |
| 2026-09-26 | A bilingual AI audiologist built through rubric-guided playbook induction outperforms human audiologists in a blinded evaluation of simulated cases | Linkai Li et.al. | 2609.32220v1 | null |
| 2026-09-26 | Evaluating Single and Multi-Omics Based Explainable Artificial Intelligence (MOXAI) for Molecular Subclass Classification of Adult-Type Diffuse Gliomas | Md Zahangir Alom et.al. | 2609.32190v1 | null |
| 2026-09-26 | REALM: Regime-Switching, Explainable, and Activation-Induced Linear Models | Xiaoran Cheng et.al. | 2609.32141v1 | null |
| 2026-09-25 | Reasoning Concentrates Errors, and Self-Consistency Never Notices | Asaad Althoubi et.al. | 2609.32035v1 | null |
| 2026-09-25 | A Safety-Bounded SDC-to-MCP Gateway for Medical AI Agents | Bennet Gerlach et.al. | 2609.31358v1 | null |
| 2026-09-25 | DriveHierarchy: A Benchmark for Diagnosing VLM Driving Capabilities from Open-Loop Understanding to Closed-Loop Execution | Chengkai Xu et.al. | 2609.31814v1 | null |
| 2026-09-25 | Rethinking Data Quality for AI-Driven Systems: Evidence from Practitioner Interviews | Hariharan Gopinath et.al. | 2609.31191v1 | null |
| 2026-09-25 | Bayesian Optimization with Fisher Information Geometry: Gradient Bounds and Trust-Region Methods | Saksham Kiroriwal et.al. | 2609.31107v1 | null |
Abstracts
MIRTO: a registration-gated, multiverse-tested evaluation protocol for unsupervised anomaly segmentation in brain MRI
2610.02136v1 by Negin Kafee Hernashki, Soumick Chatterjee
Unsupervised anomaly detection (UAD) methods for brain MRI are ranked by a single score, yet that score rests on choices that are rarely reported: how each anomaly map is aligned with the reference, how and on which data the threshold is set, and which false-positive budget, metric, aggregation and lesion definition are used. We present MIRTO, an evaluation protocol that makes these choices explicit and measures their effect. It gates the geometry of every comparison with a registration check and label-free diagnostics of known power, sets thresholds on validation data alone and reports the false-positive volume actually realised on test, repeats each comparison over 15,552 defensible evaluation pipelines, and attaches paired subject-bootstrap intervals with multiplicity control. Applied to four UAD methods trained on the same healthy data and tested on 312 BraTS 2020 subjects, MIRTO showed that an axis-order mismatch between stored maps and the reference lowered a diffusion model's voxel AUROC from 0.873 to 0.583 whilst barely moving its slice-level AUROC. Within each metric, the method explained at least 0.95 of the variance in voxel AUROC and AUPRC and 0.77 in Dice, but only 0.14 in lesion sensitivity, where the lesion definition and hit criterion dominated. A Dice advantage that was significant at validation thresholds vanished at equal realised false-positive burden, and an exact identity attributes it to threshold transfer. A training-free change to REFLECT's latent aggregation raised Dice at equal burden by 0.052. Nine hypotheses were tested against explicit criteria; because the same cohort served to develop the protocol, all inference is exploratory.
摘要:未監督異常檢測(UAD)方法對於腦部 MRI 的排名是基於單一分數,但該分數依賴於鮮少報告的選擇:每個異常圖與參考的對齊方式、如何以及基於哪些數據設置閾值,以及使用哪種假陽性預算、指標、聚合和病變定義。我們提出了 MIRTO,一種評估協議,使這些選擇變得明確並測量其影響。它通過註冊檢查和無標籤診斷已知功率來限制每次比較的幾何,僅在驗證數據上設置閾值,並報告在測試中實際實現的假陽性體積,重複每次比較超過 15,552 條可辯護的評估管道,並附上成對的主體自助間隔及多重性控制。應用於四種基於相同健康數據訓練並在 312 名 BraTS 2020 受試者上測試的 UAD 方法,MIRTO 顯示存儲圖與參考之間的軸序不匹配使擴散模型的體素 AUROC 從 0.873 降至 0.583,同時幾乎不影響其切片級 AUROC。在每個指標中,該方法解釋了至少 0.95 的體素 AUROC 和 AUPRC 的變異,及 0.77 的 Dice,但在病變敏感性中僅為 0.14,病變定義和命中標準主導了這一結果。在驗證閾值下顯著的 Dice 優勢在相等的實現假陽性負擔時消失,並且一個精確的身份將其歸因於閾值轉移。對 REFLECT 的潛在聚合進行無訓練的變更,在相等負擔下將 Dice 提高了 0.052。針對明確標準測試了九個假設;由於相同的隊列用於開發該協議,所有推斷都是探索性的。
A Comparative Explainability Framework for DeBERTa-v3 in Zero-Shot Medical Abstract Classification
2610.02116v1 by Javier Diaz Esteban-Herreros, David Muñoz-Valero, Raquel Martínez-España, Jose M. Juarez, Juan Moreno-Garcia
A comparative explainability framework is presented to audit DeBERTa-v3 under zero-shot classification of medical abstracts. The work addresses the disagreement problem in Explainable Artificial Intelligence, where different attribution methods produce divergent explanations for the same input and prediction. A natural language inference engine is implemented over the Medical Abstracts corpus with five enriched hypotheses per diagnostic category and a balanced sample of one thousand texts per class. Five explanation methods are compared: SHAP and LIME as model-agnostic approaches, occlusion and Input x Gradient as deep-learning-specific approaches, and Attention x Gradient as a transformer-specific approach. Explanations are standardized through top-token attribution, and pairwise agreement is quantified using the Jaccard index. High predictive accuracy is achieved across well-defined clinical domains, whereas performance degrades under high semantic ambiguity. Explanatory stability directly mirrors predictive certainty, exhibiting strong convergence in univalent categories and a marked drop under diagnostic uncertainty. Furthermore, qualitative error auditing uncovers three systemic failure mechanisms: lexical hypersensitivity, semantic overlap, and loss of attribution coherence. The results support the combined use of several explanation methods and quantitative agreement metrics when auditing transformer-based models in medical text classification, and suggest prioritizing specific clinical ontologies over broad diagnostic labels.
摘要:比較可解釋性框架被提出以審計 DeBERTa-v3 在醫學摘要的零樣本分類下。這項工作解決了可解釋人工智慧中的不一致問題,即不同的歸因方法對相同的輸入和預測產生不同的解釋。自然語言推理引擎在醫學摘要語料庫上實施,每個診斷類別有五個增強的假設,並且每個類別有一千篇文本的平衡樣本。比較了五種解釋方法:SHAP 和 LIME 作為模型無關的方法,遮蔽和輸入 x 梯度作為深度學習特定的方法,以及注意力 x 梯度作為Transformer特定的方法。通過頂部標記歸因標準化解釋,並使用 Jaccard 指數量化成對一致性。在明確定義的臨床領域中實現了高預測準確性,而在高語義模糊性下性能下降。解釋穩定性直接反映預測確定性,在單值類別中顯示出強烈的收斂,並在診斷不確定性下顯著下降。此外,定性錯誤審計揭示了三種系統性失效機制:詞彙過敏、語義重疊和歸因一致性的喪失。結果支持在醫學文本分類中審計基於Transformer的模型時,結合使用幾種解釋方法和定量一致性指標,並建議優先考慮特定的臨床本體論而非廣泛的診斷標籤。
Can AI Oversight Be Zero Knowledge?
2610.01995v1 by Alessandro Chiesa, Ziyi Guan, Burcu Yildiz
AI systems increasingly produce outputs from confidential data, such as a fitness-for-duty assessment from medical records or the predicted properties of a drug candidate from its secret structure. It is important to verify that such outputs are correct without revealing the underlying data. A recent line of work studies verification of AI outputs via interactive proofs and debate for oracle-aided computation, where correctness may depend on an oracle such as human judgment, a physical experiment, or the web. These works focus on verification by a verifier that runs much faster than the computation. However, such efficient verification is impossible for general oracle-aided computation, and these works therefore rely on additional assumptions. We focus instead on privacy: allowing the verifier to run in time polynomial in the computation, we ask whether interactive arguments for oracle-aided computation can be zero knowledge, so that the verifier learns nothing about the confidential data beyond the correctness of the output. We prove that, in general, they cannot. In the random oracle model, there are no zero-knowledge proofs for all oracle-aided computations, even if both the prover and the verifier are allowed to run much longer than the computation itself. The impossibility extends to debate, a canonical model for scalable oversight. On the positive side, we show that if the oracle attaches a cryptographic signature to each of its answers, then every oracle-aided computation can be verified in zero knowledge with an efficient prover and verifier, assuming only collision-resistant hash functions. Beyond privacy, this also gives an alternative approach to scalable oversight that relies neither on an honest opponent, as in debate, nor on the robustness of the computation, as in prior single-prover protocols.
摘要:AI 系統越來越多地從機密數據中產生輸出,例如從醫療記錄中進行的適任性評估或從其秘密結構中預測的藥物候選物的性質。
驗證這些輸出是否正確而不透露底層數據是很重要的。
最近的一系列研究通過互動證明和辯論研究 AI 輸出的驗證,用於有 oracle 輔助的計算,其中正確性可能依賴於 oracle,例如人類判斷、物理實驗或網絡。
這些研究專注於由運行速度遠快於計算的驗證者進行的驗證。
然而,對於一般的有 oracle 輔助計算,這樣的高效驗證是不可能的,因此這些研究依賴於額外的假設。
我們則專注於隱私:允許驗證者在計算的多項式時間內運行,我們詢問有 oracle 輔助計算的互動論證是否可以是零知識的,以便驗證者不會學到關於機密數據的任何信息,除了輸出的正確性。
我們證明,通常情況下,它們是不可能的。
在隨機 oracle 模型中,對於所有有 oracle 輔助的計算,沒有零知識證明,即使證明者和驗證者都被允許運行的時間遠超過計算本身。
這種不可能性擴展到辯論,這是一個可擴展監督的典型模型。
從積極的一面來看,我們展示了如果 oracle 為其每個答案附加加密簽名,那麼每個有 oracle 輔助的計算都可以在零知識中進行驗證,並且有高效的證明者和驗證者,只假設碰撞抗性哈希函數。
除了隱私之外,這還提供了一種可擴展監督的替代方法,既不依賴於誠實的對手(如辯論中),也不依賴於計算的穩健性(如以前的單證明者協議中)。
Mapping the RAG Landscape: A Four Axis Taxonomy of Efficiency, Defense, Interactivity, and Reasoning
2610.01936v1 by Meghana Sunil, Shravya V, Shravan Venkatraman, Joe Dhanith PR
Large Language Models (LLMs) have demonstrated remarkable fluency across many tasks but remain limited by their static, parameter bound knowledge and their susceptibility to hallucinating information. Retrieval Augmented Generation (RAG) addresses these issues by incorporating external retrieval into the generation process, grounding model outputs in verifiable and up to date sources. While prior surveys primarily focus on core RAG architectures and standard pipelines, recent research explores broader challenges and capabilities that extend beyond these foundational designs. This survey provides a consolidated and structured examination of contemporary RAG developments, organizing the field into a four axis taxonomy: improving retrieval efficiency, strengthening robustness and security, supporting user driven and interactive workflows, and enabling multi step or complex reasoning. We formalize key components of the RAG framework and review methods spanning dense and sparse retrieval, fusion strategies, embedding optimizations, and reinforcement learning based retrieval policies, highlighting how these advances influence practical deployment and system design. We also synthesize evaluation practices, domain specific applications, and architectural variants such as Naive, Advanced, and Modular RAG. Finally, we outline persistent challenges related to retrieval quality, reliability, domain adaptation, scalability, and explainability, and identify opportunities for building RAG systems that are more reliable, adaptable, and transparent.
摘要:大型語言模型(LLMs)在許多任務中展現了卓越的流暢性,但仍然受到靜態的、參數限制的知識以及對虛假信息的易感性的限制。檢索增強生成(RAG)通過將外部檢索納入生成過程來解決這些問題,使模型輸出基於可驗證且最新的來源。雖然之前的調查主要集中在核心RAG架構和標準流程上,但最近的研究探討了超越這些基礎設計的更廣泛挑戰和能力。本調查提供了一個當代RAG發展的綜合和結構化檢視,將該領域組織為四個軸向的分類法:提高檢索效率、加強穩健性和安全性、支持用戶驅動和互動工作流程,以及實現多步驟或複雜推理。我們正式化了RAG框架的關鍵組件,並回顧了涵蓋密集和稀疏檢索、融合策略、嵌入優化和強化學習基於檢索政策的方法,強調這些進展如何影響實際部署和系統設計。我們還綜合了評估實踐、特定領域的應用以及如Naive、Advanced和Modular RAG等架構變體。最後,我們概述了與檢索質量、可靠性、領域適應性、可擴展性和可解釋性相關的持續挑戰,並確定了構建更可靠、可適應和透明的RAG系統的機會。
Removing spurious minima for planar features by skip connections
2610.01728v1 by Jakob Paul Zimmermann, Moritz Grillo, Andrei Balakin, Georg Loho
Understanding loss landscapes is central to explaining neural-network training, yet their structure remains only partially understood even in simple models. We study the Gaussian population loss of shallow, bias-free ReLU networks in the teacher--student setting. This provides a simple model for studying essential aspects such as feature learning and overparameterization. For teacher networks with positive output weights and planar features, we show that including a learned linear skip removes all spurious local minima with non-negative student output weights once the student network is at least as wide as the teacher network. In contrast, without the skip, we construct a fixed teacher network with positive output weights and only three hidden neurons in input dimension two whose spurious local minima persist at every student width at least three. Thus, a learned linear skip can remove spurious minima that persist under arbitrary overparameterization. Furthermore, we show that a positive output weight student network always learns the subspace spanned by the teacher features: student features at local minima with non-negative student output weights lie in the span of the teacher features. For ReLU networks in two dimensions, even heavily overparameterized student networks have effective width controlled by the teacher width: every critical point with positive student output weights has at most twice as many distinct student feature directions as teacher neurons. Finally, we transfer the benignity result to empirical minima over parameter balls of any prescribed radius, with the required sampling accuracy depending on that radius.
摘要:理解損失景觀對於解釋神經網絡訓練至關重要,然而即使在簡單模型中,它們的結構仍然只有部分被理解。
我們研究教師-學生設置中淺層、無偏的ReLU網絡的高斯族群損失。
這提供了一個簡單的模型來研究如特徵學習和過度參數化等基本方面。
對於具有正輸出權重和平面特徵的教師網絡,我們顯示包含學習的線性跳過可以消除所有具有非負學生輸出權重的虛假局部最小值,只要學生網絡的寬度至少與教師網絡一樣寬。
相反,如果不使用跳過,我們構造了一個固定的教師網絡,其具有正輸出權重且在輸入維度為二的情況下只有三個隱藏神經元,這樣的虛假局部最小值在每個學生寬度至少為三的情況下持續存在。
因此,學習的線性跳過可以消除在任意過度參數化下持續存在的虛假最小值。
此外,我們顯示具有正輸出權重的學生網絡總是學習由教師特徵所跨越的子空間:在具有非負學生輸出權重的局部最小值下,學生特徵位於教師特徵的跨度內。
對於二維的ReLU網絡,即使是高度過度參數化的學生網絡,其有效寬度也受到教師寬度的控制:每個具有正學生輸出權重的臨界點最多有教師神經元的兩倍不同學生特徵方向。
最後,我們將良性結果轉移到任何指定半徑的參數球上的經驗最小值,所需的取樣精度取決於該半徑。
MCIR: A Feature Dependence-Aware Explainability Method with Reliability Guarantees
2610.01641v1 by Poushali Sengupta, Sabita Maharjan, Frank Eliassen, Shashi Raj Pandey, Yan Zhang
Modern machine-learning models often contain strongly dependent or redundant features, making feature attribution difficult because shared predictive information can be distributed across correlated predictors. Existing methods such as SHAP, LIME, HSIC, MI/CMI, and SAGE may therefore produce unstable rankings under multicollinearity or near-duplicate predictors. We propose the Mutual Correlation Impact Ratio Method (MCIR-M), a dependence-aware global feature-importance approach that quantifies the unique predictive information contributed by each feature beyond a selected dependence neighbourhood. MCIR-M introduces the Mutual Correlation Impact Ratio (MCIR), which conditions each feature on strongly dependent neighbours and computes a normalized ratio of conditional to block-level information. The population score lies in [0,1] and equals zero under exact conditional redundancy. We also introduce a lightweight estimation procedure that computes MCIR using a fraction of the available data and evaluates agreement with full-data explanations. Across controlled synthetic redundancy experiments and the UCI HAR benchmark, MCIR shows dependence-aware ranking behaviour, with its clearest advantage under injected near-duplicate predictors. Comparisons with independent and conditional SHAP, SAGE, HSIC, MI-based scores, and CIR-family baselines are mixed across real-data criteria. Reduced explanation samples lower computational burden in the evaluated configurations, while agreement with full-data explanations is assessed separately through ranking, head-set, and faithfulness diagnostics. Overall, MCIR-M provides a practical dependence-aware diagnostic for global explanation under strong feature dependence.
摘要:現代機器學習模型通常包含強相關或冗餘的特徵,使得特徵歸因變得困難,因為共享的預測信息可能分佈在相關的預測變數之間。現有的方法如SHAP、LIME、HSIC、MI/CMI和SAGE在多重共線性或近乎重複的預測變數下可能因此產生不穩定的排名。我們提出了互相關影響比率方法(MCIR-M),這是一種考慮依賴性的全局特徵重要性方法,量化每個特徵在選定的依賴鄰域之外所貢獻的獨特預測信息。MCIR-M引入了互相關影響比率(MCIR),該比率在強依賴的鄰居上對每個特徵進行條件化,並計算條件信息與區塊級信息的標準化比率。該人口得分位於[0,1]之間,並在精確的條件冗餘下等於零。我們還引入了一種輕量級的估計程序,該程序使用部分可用數據計算MCIR,並評估與全數據解釋的一致性。在受控的合成冗餘實驗和UCI HAR基準測試中,MCIR顯示出考慮依賴性的排名行為,其在注入的近重複預測變數下的優勢最為明顯。與獨立和條件SHAP、SAGE、HSIC、基於MI的得分以及CIR系列基準的比較在真實數據標準下是混合的。在評估的配置中,減少的解釋樣本降低了計算負擔,而與全數據解釋的一致性則通過排名、頭部集和忠實性診斷單獨評估。總體而言,MCIR-M為強特徵依賴下的全局解釋提供了一種實用的考慮依賴性的診斷方法。
What Makes Something Hard(er)? Explaining Question Difficulty in Natural Language
2610.01627v1 by Peng Cui, Qiaoyuan Zheng, Rudolf Debelak, Mrinmaya Sachan
Difficulty is one of the most fundamental properties of a question: it determines whether the question can meaningfully discriminate between models of differing ability. Although a variety of methods can now estimate or predict difficulty automatically, they yield only a single descriptive number, with no account of the underlying factors that make a question difficult in the first place. In this work, we propose a data-driven approach that automatically generates and validates natural-language hypotheses explaining what makes one question harder than another. We first estimate each item's difficulty from the responses of a large pool of LLMs using Item Response Theory. We then sample contrasting sets of easy and hard questions and prompt an LLM to propose candidate explanations of the difference, which are subsequently validated and selected on held-out questions. Experimental results across three datasets spanning mathematical, logical, and commonsense reasoning show that our method produces interpretable and predictive hypotheses. On their own, they predict the difficulty of unseen questions competitively with, or better than, advanced black-box difficulty regressors; used as additional features, they further improve those regressors, implying that they discover difficulty signals that existing models fail to capture. Moreover, we demonstrate that editing questions according to a hypothesis can shift their measured difficulty in the expected direction, indicating that the discovered hypotheses are causally valid difficulty factors rather than post-hoc descriptions. Our approach thus turns a purely descriptive difficulty score into actionable statements.
摘要:困難度是問題最基本的特性之一:它決定了問題是否能夠有意義地區分不同能力的模型。雖然現在有多種方法可以自動估計或預測困難度,但它們僅產生一個描述性的數字,並未考慮使問題變得困難的潛在因素。在這項工作中,我們提出了一種數據驅動的方法,自動生成和驗證自然語言假設,解釋為什麼一個問題比另一個問題更難。我們首先使用項目反應理論從大量大型語言模型的回應中估計每個項目的困難度。然後,我們抽取一組對比的簡單和困難問題,並提示一個大型語言模型提出候選解釋這些差異,這些解釋隨後在保留的問題上進行驗證和選擇。跨越數學、邏輯和常識推理的三個數據集的實驗結果顯示,我們的方法產生了可解釋且具有預測性的假設。僅憑這些假設,它們能夠與先進的黑箱困難回歸模型競爭地預測未見問題的困難度;作為額外特徵使用時,它們進一步改善了這些回歸模型,這意味著它們發現了現有模型未能捕捉的困難信號。此外,我們證明根據假設編輯問題可以將其測量的困難度朝預期方向轉變,這表明所發現的假設是因果有效的困難因素,而非事後描述。因此,我們的方法將純粹描述性的困難分數轉化為可行的陳述。
Measuring the Stability Assumption Behind Action Chunking
2610.01626v1 by Aryan Goyal
Action chunking improves the performance of policies learned by behavioural cloning, and several mechanisms have been proposed to explain why, including temporal consistency, horizon reduction, representation learning, and reduced error compounding. We instead study what happens to an action error once it enters the system. At each state, we inject a small action error and measure how fast it grows or shrinks under two execution regimes: open-loop, where the rest of the chunk is replayed without replanning, and closed-loop, where the policy replans after the perturbation. The fitted rate labels each state as contracting, expanding, or unresolved. Across twelve manipulation tasks from three benchmark suites, we find that confidently stable states are rare, while error amplification is common among states whose propagation rate can be resolved. We further find that the measured propagation rate depends strongly on the fitting horizon: amplification is typically front-loaded, so short windows can overestimate longer-horizon propagation. Finally, we train predictors on these labels and find that a state's open-loop regime can be recovered from camera frames and proprioception alone, while its closed-loop propagation is only partially recoverable because it also depends on how the policy acts after the perturbation. These results suggest that error-compounding arguments alone do not provide a complete account of action chunking: neither passive open-loop dynamics nor policy replanning consistently contracts an injected error, and replanning rarely turns open-loop amplification into confident contraction. This suggests that closed-loop reactivity should be trained explicitly, using perturbation- and tree-coverage-oriented training to expose policies to deviations they must recover from, rather than expected to emerge reliably from standard imitation learning.
摘要:行動分塊改善了通過行為複製學習的政策的表現,並提出了幾種機制來解釋原因,包括時間一致性、視野縮減、表示學習和減少錯誤累積。我們則研究一旦行動錯誤進入系統後會發生什麼。在每個狀態下,我們注入一個小的行動錯誤,並測量在兩種執行模式下它的增長或縮小速度:開環模式,其中其餘的分塊在不重新規劃的情況下重播;以及閉環模式,在擾動後政策重新規劃。擬合速率將每個狀態標記為收縮、擴張或未解決。在三個基準套件的十二個操作任務中,我們發現自信穩定的狀態是稀有的,而錯誤放大在其傳播速率可以解決的狀態中是常見的。我們進一步發現,測量的傳播速率強烈依賴於擬合視野:放大通常是前置的,因此短時間窗口可能會高估長視野的傳播。最後,我們對這些標籤訓練預測器,並發現狀態的開環模式可以僅從攝像頭幀和本體感知中恢復,而其閉環傳播僅部分可恢復,因為它還取決於政策在擾動後的行為。這些結果表明,僅僅依賴錯誤累積的論點並不能完整解釋行動分塊:無論是被動的開環動力學還是政策重新規劃都不一致地收縮注入的錯誤,而重新規劃很少將開環放大轉化為自信的收縮。這表明閉環反應性應該明確進行訓練,使用擾動和樹覆蓋導向的訓練來使政策暴露於必須恢復的偏差,而不是期待從標準模仿學習中可靠地出現。
Exact Distinguishability in Non-Markovian Decision Processes
2610.01527v1 by Kabir Murjani, Nisarg Patel
Non-Markovian environments are often modeled as Regular Decision Processes (RDPs), where dynamics depend on the interaction history through a finite automaton. Existing offline guarantees for RDPs rely on a distinguishability assumption on the behaviour policy but provide no means of verifying it. When the assumption is violated, distinct models may explain the data equally well. We study when data collected under a fixed behaviour policy can distinguish two candidate RDPs. We prove that the posterior odds between observationally equivalent candidates remain equal to the prior odds at every sample size, even when the policy visits every automaton state, and verify both results formally in Lean 4. We then characterize this equivalence exactly and derive PEC, an algorithm that decides it in time linear in the size of the product automaton. The distinguishability assumption of prior work fails on three of our four test environments, and the experiment identified by PEC restores it in each case.
摘要:非馬可夫環境通常被建模為正則決策過程(RDP),其動態依賴於通過有限自動機的互動歷史。現有的RDP離線保證依賴於行為策略的可區分性假設,但未提供驗證該假設的手段。當假設被違反時,不同的模型可能同樣能夠解釋數據。我們研究在固定行為策略下收集的數據何時能區分兩個候選RDP。我們證明觀察上等價的候選者之間的後驗比率在每個樣本大小下保持等於先驗比率,即使該策略訪問了每個自動機狀態,並在Lean 4中正式驗證了這兩個結果。我們然後準確地表徵這種等價性,並推導出PEC,一種在產品自動機大小的線性時間內決定它的算法。先前工作的可區分性假設在我們四個測試環境中的三個失敗,而PEC識別的實驗在每個案例中恢復了它。
OpenMTB-Audit: Exposing Over-Refusal and Clinical Expert Perspectives in LLM-Based Molecular Tumor Board Safety Evaluation
2610.01497v1 by Negin Ashrafi, Jia Luo, Stacey M. Frumm, Roxana Daneshjou
Molecular tumor boards integrate genomic findings, clinical context, and therapeutic evidence to support precision oncology. As AI enters this workflow, a key safety challenge is distinguishing truly unsupported recommendations from evidence-supported options that still require oncologist review because of incomplete information, poor ECOG performance status, or other clinical caveats. We introduce OpenMTB-Audit, an open-source benchmark of 500 synthetic non-small cell lung cancer cases spanning five adversarial error categories and four safety labels: Supported, Partially Supported, Unsupported, and Insufficient Information. Across eight large language model configurations, we identify pervasive over-refusal: all LLM configurations failed to retain the Partially Supported label in 83.3-100% of true Partially Supported cases, achieving high aggregate safety scores through label collapse rather than clinically calibrated reasoning. To address this limitation, we developed MTB-AuditAgent, a deterministic seven-module framework separating evidence verification, missing-information detection, safety classification, and abstention. It reduces over-refusal to 6.7% and achieves 91.2% accuracy (95% CI: 88.6-93.6%). A two-oncologist annotation study found disagreement concentrated at the boundary between information sufficiency and treatment optimization, underscoring the need to preserve clinically meaningful distinctions.
摘要:分子腫瘤委員會整合基因組發現、臨床背景和治療證據,以支持精準腫瘤學。隨著人工智慧進入這一工作流程,一個主要的安全挑戰是區分真正不被支持的建議與仍需腫瘤醫生審查的證據支持選項,因為信息不完整、ECOG表現狀態不佳或其他臨床警告。我們介紹了OpenMTB-Audit,一個開源基準,包含500個合成的非小細胞肺癌案例,涵蓋五個對抗性錯誤類別和四個安全標籤:支持、部分支持、不支持和信息不足。在八種大型語言模型配置中,我們發現普遍的過度拒絕:所有LLM配置在83.3-100%的真實部分支持案例中未能保留部分支持標籤,通過標籤崩潰而非臨床校準推理獲得高整體安全分數。為了解決這一限制,我們開發了MTB-AuditAgent,一個確定性的七模塊框架,將證據驗證、缺失信息檢測、安全分類和放棄分開。它將過度拒絕降低到6.7%,並達到91.2%的準確率(95% CI:88.6-93.6%)。一項由兩位腫瘤醫生進行的標註研究發現,分歧集中在信息充分性和治療優化之間的邊界,強調了保留臨床上有意義的區別的必要性。
Multi-Party Backchannel Prediction: a Diagnosis, a Benchmark, and a Ceiling
2610.01488v1 by Mohammed Hafsati, Ahmed Loughzali
Backchannel prediction has been studied almost entirely in dyadic conversation. We introduce a multi-party benchmark based on the AMI corpus, comprising 682 masked-listener views from 171 meetings, 190 speakers, and 18,697 backchannel events, with a person-disjoint held-out split. A state-of-the-art dyadic model applied zero-shot to meeting audio performs at chance (AUROC 0.499); nevertheless, its frozen acoustic features remain informative: a linear probe reaches 0.704, and retraining the predictor raises performance to 0.751. Retraining reveals a second limitation. Listener conditioning improves prediction for listeners seen during training but not for unseen listeners, and the gap remains under capacity reduction, listener-adversarial training, per-listener adaptation, and oracle lexical conditioning. Adversarial training removes only part of the speaker-identity information, while stronger removal hurts prediction, suggesting that identity is entangled with cues that are useful for backchanneling. A within-model control helps explain this pattern: with the same features and data splits, turn-onset prediction transfers to unseen listeners, while backchannel prediction does not. Backchannel rates also vary about twice as much across individuals as turn-onset rates. Since backchannels occupy only about 1% of frames, frame-level F1 is strongly affected by the base rate. We therefore report AUROC alongside event-F1 on listener-active regions. We release the benchmark and evaluation tools at https://github.com/HafsatiMohammed/bc_multiparty_release.
摘要:回饋通道預測幾乎完全在雙人對話中進行了研究。我們基於AMI語料庫引入了一個多方基準,包含來自171次會議的682個被遮蔽的聆聽者視角、190位講者和18,697個回饋通道事件,並設有一個人員不重疊的保留分割。一個最先進的雙人模型在會議音頻上進行零樣本應用,表現與隨機相當(AUROC 0.499);然而,其凍結的聲學特徵仍然具有信息性:線性探測器達到0.704,並且重新訓練預測器將性能提高到0.751。重新訓練揭示了第二個限制。聆聽者的條件化改善了對在訓練期間出現的聆聽者的預測,但對未見過的聆聽者則沒有改善,並且在容量減少、聆聽者對抗訓練、每位聆聽者的適應和oracle詞彙條件下,這一差距依然存在。對抗訓練僅去除了部分講者身份信息,而更強的去除會損害預測,這表明身份與對回饋通道有用的線索交織在一起。模型內部控制有助於解釋這一模式:在相同的特徵和數據分割下,轉換開始預測能夠轉移到未見過的聆聽者,而回饋通道預測則無法。回饋通道的比率在個體之間的變化大約是轉換開始比率的兩倍。由於回饋通道僅佔約1%的幀,因此幀級F1受到基礎比率的強烈影響。因此,我們在聆聽者活躍區域報告AUROC和事件-F1。我們在https://github.com/HafsatiMohammed/bc_multiparty_release發布了基準和評估工具。
Detect, Explain, Interpret: An End-to-End Benchmark for Time Series Anomaly Detection, Explainability and Interpretability
2610.01168v1 by Roberto Stanzione, Jules Barbe, Magali Parrino, Jérémie Fourmann, Paul Boniol
Time Series Anomaly Detection has received increasing attention, driven by the growing availability of complex time series data. This surge has led to the development of numerous detection methods, as well as a variety of benchmarks aimed at thoroughly evaluating their performance. However, most existing detectors remain largely agnostic to domain context, overlooking explainability and interpretability. One of the main reasons for this gap is that current benchmarks primarily focus on detection accuracy, and only few of them evaluate spatial explainability. Moreover, no benchmark currently provides sufficiently rich semantic annotations to support the generation of human-understandable interpretations of anomalies. To address these limitations, we introduce SHAD (Scality High-dimensional Anomaly Detection benchmark), a fully annotated benchmark composed of 215 multivariate, high-dimensional time series collected from real-world distributed cloud storage systems operated by Scality. The proposed dataset includes rich contextual information, covering three families of anomalies with varying degrees of severity. As further contribution, we provide a foundation for future work by evaluating baseline methods for Detection, Explainability, and Interpretability, covering all stages of a TSAD pipeline. For Detection, we benchmark a wide range of existing anomaly detectors, testing their effectiveness on the proposed real-world dataset. Then, we consider explainability by evaluating whether measuring the contribution of each dimension in the generated anomaly score can provide accurate anomaly attributions. Finally, for interpretability, we investigate the effectiveness of frozen LLM baselines in localizing and interpreting anomalies.
摘要:時間序列異常檢測受到越來越多的關注,這是由於複雜時間序列數據的日益可用性所驅動。這一增長促使了許多檢測方法的發展,以及各種基準的出現,旨在徹底評估它們的性能。然而,大多數現有的檢測器在很大程度上對領域上下文保持無知,忽視了可解釋性和可理解性。這一差距的主要原因之一是目前的基準主要集中在檢測準確性上,只有少數評估空間可解釋性。此外,目前沒有任何基準提供足夠豐富的語義註釋,以支持生成易於人類理解的異常解釋。為了解決這些限制,我們引入了SHAD(Scality高維異常檢測基準),這是一個完全註釋的基準,由215個來自Scality運營的真實分佈雲存儲系統的多變量高維時間序列組成。所提出的數據集包括豐富的上下文信息,涵蓋三類具有不同嚴重程度的異常。作為進一步的貢獻,我們通過評估檢測、可解釋性和可理解性的基線方法,為未來的工作提供了一個基礎,涵蓋了TSAD管道的所有階段。對於檢測,我們基準測試了各種現有的異常檢測器,測試它們在所提出的真實世界數據集上的有效性。然後,我們通過評估在生成的異常分數中測量每個維度的貢獻是否能提供準確的異常歸因來考慮可解釋性。最後,對於可理解性,我們調查了凍結的LLM基線在定位和解釋異常方面的有效性。
CineMR: Tool-Integrated Vision-Language Reasoning for Quantitative Cardiac MRI Assessment
2610.01166v1 by Kunyang Li, Hai Nguyen, Joshua Lowe, Chenguang Zhao, Peace C. Madueme, Mehdi Hedjazi Moghari, Mubarak Shah, Pegah Khosravi, Yuzhang Zhang
Cardiovascular magnetic resonance (CMR), including cine imaging, is a reference standard for the noninvasive assessment of cardiac morphology and ventricular function. Cine CMR interpretation integrates qualitative visual assessment with quantitative measurements of ventricular volumes, ejection fraction, myocardial mass, wall thickness, and regional wall motion. Current medical vision-language models (VLMs) cannot reliably derive quantitative measurements from multidimensional cine images without analysis tools. We present CineMR, a tool-augmented VLM that invokes cardiac image-analysis tools and integrates their outputs into interleaved reasoning for quantitative CMR assessment. We also construct a multi-cohort visual question answering benchmark covering quantitative metric extraction, multiclass diagnosis, and differential diagnosis, together with tools for segmentation, phase selection, volumetry, morphometry, and regional wall motion analysis. CineMR is trained with supervised fine-tuning (SFT) on tool-interaction traces followed by Group Relative Policy Optimization (GRPO) with conditional tool-use rewards. On the multi-cohort cine CMR benchmark, CineMR achieves 35.9% pass@1 and 58.9% pass@4, compared with 1.5% pass@1 for the Qwen3-VL-8B backbone and 0.0% and 7.0% pass@1 for LLaVA-Med v1.5 and MedGemma-4B, respectively. Correct tool invocation reaches 99.8% after GRPO, up from 78.9% after SFT. Live tool outputs improve ventricular measurement accuracy by 20.4--23.7% over direct model predictions, and removing all tools reduces pass@1 from 35.9% to 27.9%. These results highlight the importance of reliable tool use for quantitative cine CMR reasoning and support CineMR as a promising approach for assistive cardiac image assessment. Code, benchmark resources, and model weights are available at https://github.com/AI-MIND-Lab/CineMR.
摘要:心血管磁共振(CMR),包括動態影像,是非侵入性評估心臟形態和心室功能的參考標準。動態 CMR 解釋將定性視覺評估與心室體積、射血分數、心肌質量、壁厚和區域壁運動的定量測量相結合。目前的醫療視覺-語言模型(VLMs)在沒有分析工具的情況下,無法可靠地從多維動態影像中推導出定量測量。我們提出了 CineMR,一種增強工具的 VLM,調用心臟影像分析工具並將其輸出整合到交錯推理中,以進行定量 CMR 評估。我們還構建了一個涵蓋定量指標提取、多類別診斷和鑑別診斷的多隊列視覺問答基準,並提供分割、相位選擇、體積測量、形態測量和區域壁運動分析的工具。CineMR 在工具互動痕跡上進行了監督微調(SFT),隨後使用條件工具使用獎勵進行了群體相對策略優化(GRPO)。在多隊列動態 CMR 基準上,CineMR 的 pass@1 為 35.9%,pass@4 為 58.9%,而 Qwen3-VL-8B 的 pass@1 僅為 1.5%,LLaVA-Med v1.5 和 MedGemma-4B 的 pass@1 分別為 0.0% 和 7.0%。經過 GRPO 正確調用工具的比例達到 99.8%,而 SFT 後為 78.9%。實時工具輸出提高了心室測量的準確性,較直接模型預測提高了 20.4% 至 23.7%,而去除所有工具則使 pass@1 從 35.9% 降至 27.9%。這些結果突顯了可靠工具使用在定量動態 CMR 推理中的重要性,並支持 CineMR 作為輔助心臟影像評估的有前景方法。代碼、基準資源和模型權重可在 https://github.com/AI-MIND-Lab/CineMR 獲得。
What Can Analogy Tell Us About Artificial Consciousness?
2610.01002v1 by Keith J. Holyoak, Martin M. Monti
Who or what is conscious? Because subjective experience is directly accessible only in the first person, judgments about consciousness in other entities depend partly on analogy. Historically, such inferences have focused on nonhuman animals, but advances in artificial intelligence have raised the possibility of conscious AI. Here we develop a causal framework for evaluating such evidential analogies. The key distinction is between similarities in factors plausibly involved in generating consciousness and similarities in downstream behavioural or cognitive effects. Our framework weights source-target similarity by causal relevance while allowing for unknown causes, disabling differences and alternative routes to consciousness. Applied to biological systems, it explains why analogical support generally weakens with increasing causal distance from humans. Applied to contemporary AI, it suggests that behavioural similarity provides only limited evidence for consciousness because relevant causal correspondences remain poorly established. The framework also clarifies what evidence would strengthen claims of artificial consciousness.
摘要:誰或什麼是有意識的?因為主觀經驗僅在第一人稱中直接可得,對其他實體意識的判斷部分依賴於類比。歷史上,這種推斷主要集中在非人類動物上,但人工智慧的進步已經提高了有意識 AI 的可能性。在這裡,我們發展了一個評估這種證據類比的因果框架。關鍵的區別在於可能涉及生成意識的因素之間的相似性,以及下游行為或認知效應之間的相似性。我們的框架根據因果相關性對源目標相似性進行加權,同時考慮未知原因、禁用差異和通往意識的替代路徑。應用於生物系統,它解釋了為什麼類比支持通常隨著與人類的因果距離增加而減弱。應用於當代 AI,它表明行為相似性僅提供有限的意識證據,因為相關的因果對應仍然建立得不夠充分。該框架還闡明了什麼證據可以加強人工意識的主張。
When Reasoning Helps Action: Monitoring and Steering Chain-of-Thought in Vision-Language-Action Policies
2610.00601v1 by Sathwik Karnik, Joseph JR. Lee, Aryaman Gupta, Somil Bansal
Reasoning-enabled VLA policies expose chain-of-thought (CoT) traces that appear to explain and guide their actions, creating a potential interface for runtime safety through reasoning monitoring and correction. In this work, we define and operationalize two evaluation axes for assessing when this interface can improve embodied behavior: correctability, which measures whether unreliable reasoning can be detected and improved during generation, and actionability, which measures whether reasoning corrections produce behaviorally meaningful changes in the intended direction. To enable correctability, we introduce Token-level Reward for Utility-Steered Chain-of-Thought (TRUST), an offline-trained value model that predicts eventual reasoning correctness from partial prefixes and uses these estimates to monitor and selectively steer reasoning generation in frozen VLA policies. On the Alpamayo 1.5 driving VLA, TRUST monitors correctness with 88.9% accuracy and improves reasoning correctness from 75.9% to 90.0%. On a baseline-defined challenging subset in AlpaSim, TRUST reduces collision rate by 30.4% and maximum trajectory error by 11.5% relative to the unsteered policy, outperforming a compute-matched Best-of-4 baseline. On the DeepThinkVLA manipulation VLA, TRUST improves the correctness of grasp-state claims from 69.3% to 90.2% and action-choice claims from 68.8% to 85.9%, yet closed-loop task performance on LIBERO-Plus remains largely unchanged. Empirical analysis reveals intent-consistent behavioral effects in Alpamayo 1.5 but limited effects in DeepThinkVLA, helping interpret these different task-level outcomes. Together, our results show that gains in reasoning correctness do not automatically imply gains in embodied performance, motivating evaluation of correctability and actionability when using CoT as a runtime safety interface.
摘要:推理驅動的 VLA 政策揭示了思考過程(CoT)痕跡,這些痕跡似乎解釋並指導其行動,通過推理監控和修正創造了一個潛在的運行時安全介面。在這項工作中,我們定義並操作化了兩個評估軸,以評估何時這個介面可以改善具身行為:可修正性,衡量在生成過程中是否能檢測到不可靠的推理並加以改進;以及可行性,衡量推理修正是否能產生在預期方向上有意義的行為變化。為了實現可修正性,我們引入了基於效用驅動的思考過程的標記級獎勵(TRUST),這是一個離線訓練的價值模型,能夠從部分前綴預測最終的推理正確性,並利用這些估計來監控和選擇性地引導凍結的 VLA 政策中的推理生成。在 Alpamayo 1.5 驅動的 VLA 上,TRUST 以 88.9% 的準確率監控正確性,並將推理正確性從 75.9% 提高到 90.0%。在 AlpaSim 中的基準定義挑戰子集上,TRUST 相對於未引導政策將碰撞率降低了 30.4%,最大軌跡誤差降低了 11.5%,超越了計算匹配的 Best-of-4 基準。在 DeepThinkVLA 操作 VLA 上,TRUST 將抓取狀態聲明的正確性從 69.3% 提高到 90.2%,將行動選擇聲明的正確性從 68.8% 提高到 85.9%,然而在 LIBERO-Plus 上的閉環任務性能仍然基本保持不變。實證分析顯示在 Alpamayo 1.5 中存在意圖一致的行為效果,但在 DeepThinkVLA 中效果有限,這有助於解釋這些不同的任務級結果。綜合來看,我們的結果表明,推理正確性的提升並不自動意味著具身表現的提升,這促使在使用 CoT 作為運行時安全介面時評估可修正性和可行性。
Worse Together: How Performance Breaks Down in Multi-User Multi-Agent Teams
2610.00583v1 by Sahan Paliskara, Nattaput Namchittai, Andrew Lampinen
People are increasingly delegating tasks to AI agents, and those agents are increasingly encountering other people's agents over shared resources such as a codebase, a calendar, or a budget. When each agent acts for a different user with different goals, coordination often fails, and the group ends up worse off than if a single agent had acted for everyone. We study this multi-user, multi-agent setting across five frontier models and 77 scenarios in four environments: an API key environment in which agents share a compute budget, a clinic in which they share a calendar, a personal assistant environment in which they share a group order or booking, and a merge queue in which they share a release cutoff. In each scenario, we compare a single agent that serves every user (a coordinator) to a team in which each agent serves one user, with and without a communication channel between the agents. Teams deliver worse group outcomes than the coordinator in every environment: without a channel, they completely collapse in two environments, and even with one, coordination overhead creates substantial gaps. For example, in the personal assistant environment, the coordinator fulfills a targeted user request about twice as often as teams. We identify distinct behaviors associated with this poor group-level performance, including stalling as teams grow, overriding each other's actions, and fabricating claims. We find effective but environment-specific mitigations, such as a team lead, explicit procedural instructions, and a platform check that makes an agent read its peers' messages before committing. We will release the API key, clinic, and personal assistant environments as MAMUBench, comprising 74 scenarios for evaluating multi-user, multi-agent coordination.
摘要:人們越來越多地將任務委派給 AI 代理,而這些代理也越來越多地與其他人的代理在共享資源上相遇,例如代碼庫、日曆或預算。當每個代理代表不同的用戶且目標不同時,協調往往失敗,結果小組的情況比由單一代理為所有人行動時更糟。我們研究了這種多用戶、多代理的設定,涵蓋五個前沿模型和四個環境中的 77 種情境:一個 API 金鑰環境,在這裡代理共享計算預算;一個診所,在這裡他們共享日曆;一個個人助理環境,在這裡他們共享團體訂單或預訂;以及一個合併隊列,在這裡他們共享發佈截止時間。在每個情境中,我們將為每個用戶服務的單一代理(協調者)與每個代理服務一位用戶的團隊進行比較,並考慮代理之間是否有通信渠道。團隊在每個環境中提供的群體結果都比協調者差:在沒有渠道的情況下,他們在兩個環境中完全崩潰,即使有一個,協調開銷也會造成相當大的差距。例如,在個人助理環境中,協調者滿足目標用戶請求的頻率約為團隊的兩倍。我們識別出與這種低群體表現相關的不同行為,包括隨著團隊增長而停滯、覆蓋彼此的行動以及捏造聲明。我們發現有效但特定於環境的緩解措施,例如團隊負責人、明確的程序指示,以及一個平台檢查,使代理在提交之前閱讀其同伴的消息。我們將發布 API 金鑰、診所和個人助理環境作為 MAMUBench,包含 74 種情境以評估多用戶、多代理的協調。
No One Architecture Fits All: A Cross-Environment Evaluation of Hierarchical Red Team Agents
2610.00557v1 by Ayan Javeed Shaikh, Arunesh Sinha, Nathaniel D. Bastian, Ankit Shah
Autonomous red team agents increasingly stress-test AI-enabled cyber defenses by planning strategy and executing multistage attacks. Reinforcement learning (RL) and large language models (LLMs) offer complementary mechanisms for the planning and execution such agents require, and prior work has combined them in hybrid hierarchies. Yet a given architecture is typically developed and evaluated within a single environment, leaving open whether an observed advantage reflects a generally stronger decision mechanism or merely alignment with a particular setting. We address this gap with a controlled cross-environment comparison of two homogeneous hierarchical red team architectures: an RL planner with an RL executor (RL+RL) and an LLM planner with an LLM executor (LLM+LLM). We evaluate both against expert autonomous defenders in CybORG CAGE-4 and in Cyberwheel at two network scales, across 18 configurations under one unified disruption metric. We find a pronounced environment-dependent inversion. RL+RL wins the compact, densely rewarded CAGE-4 (78.5% disruption success versus 18.0% for the strongest LLM configuration) and the 100-host Cyberwheel network (81.0% versus 50.5%), while a pretrained cybersecurity LLM agent wins the larger, escalation-gated 1010-host Cyberwheel network (55.0% versus 0.0% for RL). A kill-chain analysis explains the inversion through architecture-specific bottlenecks that aggregate success rates conceal.In the 1010-host Cyberwheel network, RL discovers and compromises hosts but stalls at privilege escalation, whereas in CAGE-4, LLM agents obtain privileged access but rarely convert it into operational impact. These results indicate that conclusions drawn in a single environment may not generalize, and that hybrid planner-executor designs should be motivated by specific failure modes rather than the assumption that one architecture is universally preferable.
摘要:自主紅隊代理人越來越多地通過規劃策略和執行多階段攻擊來壓力測試 AI 驅動的網絡防禦。強化學習 (RL) 和大型語言模型 (LLMs) 提供了這些代理人所需的規劃和執行的互補機制,先前的工作已將它們結合在混合層級中。然而,給定的架構通常是在單一環境中開發和評估的,這使得觀察到的優勢是否反映出一般更強的決策機制,或者僅僅是與特定設置的一致性仍然是個未解之謎。我們通過對兩個同質層級紅隊架構進行受控的跨環境比較來解決這一空白:一個是具有 RL 執行者的 RL 規劃者 (RL+RL),另一個是具有 LLM 執行者的 LLM 規劃者 (LLM+LLM)。我們在 CybORG CAGE-4 和 Cyberwheel 中針對專家自主防禦者評估這兩者,並在兩個網絡規模下,根據一個統一的干擾指標進行 18 種配置的比較。我們發現了一個明顯的環境依賴性反轉。RL+RL 在緊湊、密集獎勵的 CAGE-4 中獲勝(78.5% 的干擾成功率對比最強 LLM 配置的 18.0%),以及 100 主機的 Cyberwheel 網絡(81.0% 對比 50.5%),而一個預訓練的網絡安全 LLM 代理在更大、升級限制的 1010 主機 Cyberwheel 網絡中獲勝(55.0% 對比 RL 的 0.0%)。一項殺鏈分析通過架構特定的瓶頸解釋了這一反轉,這些瓶頸會掩蓋成功率的聚合。在 1010 主機的 Cyberwheel 網絡中,RL 發現並攻陷主機,但在特權升級時停滯不前,而在 CAGE-4 中,LLM 代理獲得特權訪問,但很少將其轉化為操作影響。這些結果表明,在單一環境中得出的結論可能無法推廣,混合規劃者-執行者設計應該基於特定的失敗模式,而不是假設某一架構是普遍可取的。
EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights
2610.00492v1 by Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig
When Isaac Newton discovered the law of gravitation, he did so through an iterative process of analyzing observed data such as planetary patterns, finding the underlying mechanisms by describing patterns in mathematical equations, and refining his theory against the Moon's orbit, revealing the startling insight that the same force governs both falling apples and orbiting planets. Would it be possible for AI agents to make similar discoveries? To measure this ability, we introduce EurekaBench, a cross-domain benchmark that tests AI agents' ability to conduct long-horizon experiments and discover mechanisms that explain observations. We evaluate these mechanisms by the scientific insights that can be derived from them. EurekaBench contains an expert-verified set of 26 long-horizon tasks across neuroscience, computer science, chemistry, astrophysics, geophysics, and plasma physics, with a total of 306 scientific insights that the discovered mechanisms are expected to support. Our evaluation framework tests three axes of scientific discovery: agents' ability to follow known scientific constraints, the predictive accuracy of the discovered mechanisms, and whether these mechanisms yield scientific insights or inform future research. Our results show that current AI agents often overly fixate on predictive accuracy optimization, surpassing human scientists, while falling substantially short in deriving scientific insights.
摘要:當艾薩克·牛頓發現萬有引力定律時,他是通過一個迭代過程來分析觀察到的數據,如行星運行模式,通過在數學方程中描述模式來尋找潛在的機制,並根據月球的軌道來完善他的理論,揭示了驚人的見解:同一種力量支配著掉落的蘋果和運行的行星。人工智慧代理是否有可能做出類似的發現?為了衡量這種能力,我們引入了EurekaBench,一個跨領域的基準,測試人工智慧代理進行長期實驗和發現解釋觀察的機制的能力。我們通過從這些機制中得出的科學見解來評估這些機制。EurekaBench包含一組經專家驗證的26個長期任務,涵蓋神經科學、計算機科學、化學、天體物理學、地球物理學和等離子體物理學,總共有306個預期由發現的機制支持的科學見解。我們的評估框架測試科學發現的三個軸心:代理遵循已知科學約束的能力、發現機制的預測準確性,以及這些機制是否產生科學見解或為未來研究提供信息。我們的結果顯示,當前的人工智慧代理往往過於專注於預測準確性的優化,超越了人類科學家,但在推導科學見解方面則大幅不足。
CAS II: Symmetric Partitions as Kolmogorov Models
2609.40290v1 by Romie Banerjee
In algorithmic statistics a string x is explained by a finite set containing it, and Kolmogorov's structure function records the smallest such model at each level of complexity. Vereshchagin's strong models, those computable from the data by a total algorithm, are essentially the cells of simple partitions. We read a partition of binary strings as a hypothesis, with the cell containing x as its model, and develop algorithmic statistics over symmetric partitions: the orbit partitions of groups acting on strings. The Galois connection between subgroups and partitions gives each ambient group a lattice of symmetric partitions, with canonical certificates, canonical costs, and an algebra of hypotheses. The resulting structure function and symmetric sophistication measure which part of the regularity of x is symmetric. For the full symmetric group every partition is symmetric: cells recover all Kolmogorov models, cells of cheap partitions recover exactly the strong models, and normal and strange strings are characterized by symmetry. For GL(n,2) the cells are exactly the linearly homogeneous sets, so linear symmetry is a restricted model class. For nonzero x, the linear-symmetry structure function lies in a band between the sufficiency line and the trivial bound, and both edges are attained: there are stochastic normal strings whose simple structure is invisible to linear symmetry. We also give coordinates on the space of permutation groups: each group is an element of a Burnside ring (its type) together with a permutation (its placement), and restriction moves refine partitions via the Mackey formula. In these coordinates the collapse for the symmetric group is a statement about placement, a linear hypothesis is determined by its type up to n^2 bits, and the maximal gap theorem shows that any space of symmetry hypotheses small enough to search is small enough to miss simple structure.
摘要:在算法統計中,字符串 x 是由包含它的有限集合來解釋的,而 Kolmogorov 的結構函數記錄了每個複雜度級別下最小的這樣的模型。Vereshchagin 的強模型,即由總算法從數據中可計算出的模型,本質上是簡單劃分的單元。我們將二進制字符串的劃分視為一個假設,包含 x 的單元作為其模型,並在對稱劃分上發展算法統計:作用於字符串的群的軌道劃分。子群與劃分之間的 Galois 連接為每個環境群提供了一個對稱劃分的格,並附有典範證明、典範成本和假設的代數。由此產生的結構函數和對稱複雜度測量 x 的正則性中哪一部分是對稱的。對於完整的對稱群,每個劃分都是對稱的:單元恢復所有 Kolmogorov 模型,廉價劃分的單元正好恢復強模型,而正常和奇怪的字符串則以對稱性為特徵。對於 GL(n,2),單元正好是線性齊次集合,因此線性對稱性是一個受限的模型類。對於非零 x,線性對稱結構函數位於充分性線和微不足道界限之間的帶中,且兩個邊界均可達:存在隨機正常字符串,其簡單結構對線性對稱性是不可見的。我們還給出了置換群空間的坐標:每個群都是一個 Burnside 環的元素(其類型)以及一個置換(其位置),而限制移動通過 Mackey 公式細化劃分。在這些坐標中,對稱群的崩潰是關於位置的陳述,線性假設由其類型決定,最多 n^2 位,最大間隙定理顯示,任何足夠小以進行搜索的對稱假設空間都足夠小以錯過簡單結構。
Unlearnable, or Unmeasured? On the Reliability of Difficulty Labels in RLVR
2609.40115v1 by Chandak Chakma, Syed Nazmus Sakib, Nafiul Haque, Shifat E. Arman
Reinforcement learning with verifiable rewards (RLVR) has become an important approach for improving reasoning during post-training. Recent work suggests that some difficult prompts remain resistant to learning even when they occasionally produce correct solutions. We revisit this unlearnability phenomenon and find that the affected prompts do improve, at roughly one third of the learnable rate, while the difficulty-defined set used to study them is much less reproducible than expected. These difficulty labels are estimated from a limited number of sampled responses. Combining them across seeds can further change which prompts are selected instead of simply reducing measurement noise. We develop a sampling-based framework for quantifying this instability and determining how much evaluation is required for difficulty assignments to reproduce reliably. We also revisit the gradient-similarity evidence proposed to explain unlearnability and show that part of the observed separation arises because difficult prompts provide fewer correct rollouts from which their gradients can be estimated. Matching this sample count weakens the gradient difference but does not remove it. Overall, the slow-learning phenomenon survives our reanalysis, while both the prompts used to define it and the evidence used to explain it require more careful measurement.
摘要:強化學習與可驗證獎勵(RLVR)已成為改善後訓練推理的重要方法。最近的研究表明,即使某些困難的提示偶爾產生正確的解決方案,它們仍然對學習具有抵抗力。我們重新檢視這一不可學習現象,發現受影響的提示確實有所改善,改善速度約為可學習速率的三分之一,而用來研究它們的困難定義集的可重現性遠低於預期。這些困難標籤是從有限數量的樣本反應中估算得出的。跨種子結合它們可能進一步改變所選擇的提示,而不僅僅是減少測量噪音。我們開發了一個基於抽樣的框架來量化這種不穩定性,並確定為了使困難分配可靠地重現需要多少評估。我們還重新檢視了用於解釋不可學習的梯度相似性證據,並顯示觀察到的分離部分源於困難提示提供的正確回饋較少,從中無法估算其梯度。匹配這一樣本數量削弱了梯度差異,但並未消除它。總體而言,緩慢學習現象在我們的重新分析中仍然存在,而用來定義它的提示和用來解釋它的證據都需要更仔細的測量。
What Can Component-Replacement Evidence Establish? A Critical Scoping Review of Local Decisions in LLM Agents
2609.39989v1 by Shuyang Zhang, Jianshuo Chang
Background. A component replacement in a language-model agent changes an execution trajectory, potentially altering later observations, resource use, and recovery opportunities. Different evidence is needed to assess its task-level benefit and the contribution of local decision quality. Methods. This critical scoping review maps 348 studies and examines 90 comparison records: 88 from 40 included studies and two from supplementary studies. Eight purposively selected cases structure the synthesis around the replaced decision, executed conditions, measurement comparability, controls, and remaining explanations. Results. Of 222 studies reporting local decision metrics, 142 also report measured task endpoints and 49 report proxies. These counts identify studies that report both types of measurement, without establishing that the measurements come from matched comparisons. Outcome Monitors reports a package-level completion gain whose attribution to detector quality remains limited; First-chunk selection reports a local improvement assessed against an offline proxy endpoint; Evidence-Carrying Termination reports fewer premature unsupported terminations and completion non-inferiority, without establishing completion superiority. Cross-case analysis identifies three candidate mechanisms involving recovery and disruption, intervention timing, and downstream use. Attribution and deployment depend on the comparison controls, label definitions, and information available to the controller. Conclusions. The review distinguishes the task-level benefit of a component replacement from the contribution of local decision quality and derives eight claim-specific reporting items. Neither online execution nor simultaneous gains in local and task metrics alone establish that better local decisions explain the task-level gain.
摘要:背景。語言模型代理中的組件替換改變了執行軌跡,可能改變後續觀察、資源使用和恢復機會。需要不同的證據來評估其任務層面的好處以及當地決策質量的貢獻。方法。這項關鍵範疇評估回顧映射了348項研究並檢查了90個比較記錄:88個來自40項納入的研究,兩個來自補充研究。八個有目的選擇的案例圍繞被替換的決策、執行條件、測量可比性、控制和剩餘解釋結構化合成。結果。在222項報告當地決策指標的研究中,142項還報告了測量的任務端點,49項報告了代理指標。這些數量識別了報告兩種類型測量的研究,但並未確立這些測量來自匹配比較。結果監控報告了一個包級別的完成增益,其歸因於檢測器質量的限制;第一塊選擇報告了一個相對於離線代理端點評估的當地改進;證據攜帶終止報告了較少的過早無支持終止和完成非劣性,但未確立完成優越性。跨案例分析識別了三個候選機制,涉及恢復和中斷、干預時機和下游使用。歸因和部署取決於比較控制、標籤定義和控制者可用的信息。結論。該評估區分了組件替換的任務層面好處與當地決策質量的貢獻,並推導出八個特定於主張的報告項目。僅僅依賴在線執行或當地和任務指標的同時增益並不能確立更好的當地決策解釋了任務層面的增益。
How Does Local Landscape Geometry Evolve in Language Model Pre-Training?
2609.39767v1 by Zhanpeng Zhou, Yuhan Sun, Bingrui Li, Jinbo Wang, Huaijin Wu, Lei Wu, Junchi Yan
The scale and expense of pre-training language models make efficient hyperparameter tuning essential, yet a principled guidance is still missing. In this work, we analyze language model pre-training dynamics from a local landscape geometry perspective. Our study reveals two distinct phases. In Phase I, sharpness of the local landscape is initially high, leading to instability and loss plateaus under large learning rates (LRs). The landscape shifts from sharp to flatter regions early in training. This dynamic explains the necessity of LR warmup and further suggests that larger peak LRs require proportionally longer warmup periods. In Phase II, the local landscape is governed by the gradient noise scale. Our theory identifies a depth flatness trade-off: high noise from smaller batches widens the loss basin, whereas reduced noise from larger batches deepens it. This theory motivates a dynamic batch-size (BS) scheduler that begins with a small BS and increases it late in training. Together, we provide a unified view of loss landscape evolution, which translates into actionable tuning strategies for large-scale pre-training.
摘要:預訓練語言模型的規模和成本使得高效的超參數調整變得至關重要,但仍然缺乏原則性的指導。在本研究中,我們從局部景觀幾何的角度分析語言模型的預訓練動態。我們的研究揭示了兩個不同的階段。在第一階段,局部景觀的尖銳度最初很高,導致在較大學習率(LRs)下的不穩定性和損失平穩期。隨著訓練的進行,景觀從尖銳轉向較平坦的區域。這一動態解釋了LR預熱的必要性,並進一步表明較大的峰值LR需要相應更長的預熱期。在第二階段,局部景觀受梯度噪聲尺度的影響。我們的理論確定了一個深度平坦度的權衡:來自較小批次的高噪聲擴大了損失盆地,而來自較大批次的低噪聲則使其變深。這一理論促使我們提出了一個動態批次大小(BS)調度器,該調度器在訓練初期從小BS開始,並在訓練後期增加它。總體而言,我們提供了一個損失景觀演變的統一視角,這轉化為大規模預訓練的可操作調整策略。
Trust Is Not a Score: Runtime Assurance Contracts for High-Risk AI Agents
2609.39717v1 by Serhii Zabolotnii
Benchmarks, audits, and agent protocols describe performance, permissions, and repair, but not how observed evidence should change an agent's authority during a consequential task. We call this the assurance-transition gap. We propose a Runtime Assurance Contract (RAC), a policy-level formal schema binding autonomy boundaries, component eligibility, evidence state, transition policy, human-review capacity, and non-compensatory gates. Under RAC, soft metrics may inform routing, whereas a failed or unknown mandatory gate forces retry, switch, escalation, deferral, or stop; aggregate performance cannot authorize action. We define the contract, an evidence record, a permission rule, and five invariants, and illustrate them in clinical, industrial, and judicial failure probes. We then report a deterministic failure-injection study in agentic coding: 280 constructed cases evaluated by a gate conjunction, a score-only rule, and a restricted protocol baseline. At the published example weights and threshold, the score rule admits 80 of 100 block-required injections and all 40 review-required injections. Tuned in hindsight, it matches the conjunction on this corpus. For positive weights, a positive threshold, binary risk signals, zero-signal controls, and an injected case firing each signal alone, we show that exact agreement holds if and only if the threshold does not exceed the smallest weight. A separate set of 18 hand-authored traces checks version-pinned evidence and review transitions against simpler policy variants. In a further prospective synthetic holdout of 24 episodes, two blinded LLM judges assign identical labels to all 72 action attempts; RAC and a separately implemented full stateful baseline both match these labels. These studies test mechanisms on synthetic cases; they establish neither deployed safety nor cross-domain effectiveness.
摘要:基準、審計和代理協議描述了性能、權限和修復,但並未說明在關鍵任務中,觀察到的證據應如何改變代理的權限。我們稱之為保證過渡差距。我們提出了一個運行時保證合約(RAC),這是一個政策層級的正式架構,約束自主邊界、組件資格、證據狀態、過渡政策、人類審查能力和非補償性閘門。在RAC下,軟指標可以用來指導路由,而失敗或未知的強制閘門則強迫重試、切換、升級、延遲或停止;總體性能無法授權行動。我們定義了合約、一個證據記錄、一條許可規則和五個不變量,並在臨床、工業和司法失敗探測中進行了說明。然後,我們報告了一項在代理編碼中的確定性失敗注入研究:280個構建的案例通過閘門聯合、一個僅計分的規則和一個受限的協議基線進行評估。在已發表的示例權重和閾值下,計分規則允許100個區塊所需注入中的80個和所有40個審查所需的注入。事後調整後,它在這個語料庫上與聯合匹配。對於正權重、正閾值、二元風險信號、零信號控制和每個信號單獨觸發的注入案例,我們顯示出精確一致性僅在閾值不超過最小權重時成立。一組18個手工編寫的痕跡檢查版本固定的證據和審查過渡,與更簡單的政策變體進行比較。在進一步的前瞻性合成保留中,24個集數中,兩位盲法LLM評審對所有72次行動嘗試分配了相同的標籤;RAC和一個單獨實施的完整狀態基線都與這些標籤相匹配。這些研究在合成案例上測試機制;它們既未建立已部署的安全性,也未建立跨領域的有效性。
ChronoGraph: Functional 4D Scene Graphs with Vision-Language Models for Interaction Understanding and Grounded Planning
2609.39665v1 by Chenyangguang Zhang, Malgorzata Gwiazda, Guanlong Jiao, Yuanchen Ju, Federico Tombari, Koushil Sreenath, Marc Pollefeys, Sunghwan Hong
Embodied agents must determine where to act, anticipate the resulting scene changes, and interpret observed outcomes to guide subsequent actions. This requires connecting 4D interaction understanding, which explains how past actions changed the scene, with spatially grounded planning, which determines how and where to act toward a goal and anticipates the resulting scene changes. We introduce ChronoGraph, a functional 4D scene graph that links actions on affordance parts to semantic and geometric state changes. By representing observed and anticipated transitions in the same form, it provides a shared basis for understanding and planning. We construct ChronoGraphBench through an automatic data engine that converts human-interaction videos and simulated robot trajectories into graph-annotated questions for training and evaluating Vision-Language Models (VLMs) on both tasks. Using these annotations, we train ChronoGraphVLM by adapting pretrained VLMs in two stages. Graph-as-Chain-of-Thought supervised fine-tuning teaches the models to reconstruct observed transitions and predict future ones as graph traces before answering. Subsequent joint 4D graph reinforcement learning directly rewards graph properties and answer correctness. Experiments across model scales show improvements over the corresponding pretrained baselines and zero-shot transfer to VLM4D. Real-world demonstrations further show that graph-based planning and affordance grounding support mobile manipulation through existing robot skills without additional fine-tuning.
摘要:具身代理必須確定行動的地點,預測隨之而來的場景變化,並解釋觀察到的結果以指導後續行動。這需要將 4D 互動理解(解釋過去的行動如何改變場景)與空間基礎規劃(確定如何以及在哪裡朝著目標行動並預測隨之而來的場景變化)連接起來。我們介紹 ChronoGraph,一個功能性 4D 場景圖,將對可供性部分的行動與語義和幾何狀態變化聯繫起來。通過以相同的形式表示觀察到的和預期的轉變,它為理解和規劃提供了一個共同的基礎。我們通過一個自動數據引擎構建 ChronoGraphBench,該引擎將人類互動視頻和模擬機器人軌跡轉換為帶有圖形標註的問題,以便在兩個任務上訓練和評估視覺-語言模型(VLMs)。利用這些標註,我們通過在兩個階段適應預訓練的 VLMs 來訓練 ChronoGraphVLM。作為思維鏈的圖形監督微調教導模型重建觀察到的轉變並在回答之前預測未來的轉變作為圖形痕跡。隨後的聯合 4D 圖形強化學習直接獎勵圖形屬性和答案的正確性。跨模型規模的實驗顯示出相對於相應的預訓練基線的改進,以及對 VLM4D 的零樣本轉移。現實世界的演示進一步表明,基於圖形的規劃和可供性基礎支持通過現有的機器人技能進行移動操作,而無需額外的微調。
Zero-Compute Cross-Lingual Transferability Estimation Using Typological Feature Proxies
2609.39640v1 by Dalton Raphael Harmsen, Swier Garst, Thomas van Osch, Zarè Palanciyan, Joaquin Vanschoren
Cross-lingual transfer describes how knowledge in a source language benefits a target language. Measuring it quantitatively requires broad multilingual pre-training, as prior work has done with cross-lingual transfer matrices. We ask whether transfer is predictable from freely available typological features, and whether the prominence of high-resource source languages reflects typology or data quality and quantity. We show that typological databases contain cheap and dense signals about cross-lingual transfer. Our typology-only random forest on a 24-language prior-work transfer matrix scores leave-one-language-out $ρ{=}0.705$ and $R^2{=}0.49$, beating a non-typological control at $ρ{=}0.62$, which verifies the ability of typology-only predictions to reconstruct costly measured cross-lingual transfer. The signal survives leave-one-script-out and leave-one-family-out protocols, so script and family confounding do not explain the effect. By decomposing the transfer into a typology term and a resource-and-script bias term, we find the best-source ranking sensitive to this bias. In contrast, typology is not affected by this bias, which makes it a zero-compute screening tool that replaces hundreds of training runs with a model fit. Our code is available \href{https://github.com/dharmsen/typo-x-ling-transfer}{here}.
摘要:跨語言轉移描述了來源語言的知識如何惠及目標語言。
定量測量需要廣泛的多語言預訓練,正如先前的工作所做的跨語言轉移矩陣。
我們詢問是否可以從自由可用的類型特徵預測轉移,以及高資源來源語言的顯著性是否反映了類型學或數據質量和數量。
我們展示了類型學數據庫包含有關跨語言轉移的廉價且密集的信號。
我們的僅基於類型學的隨機森林在24語言的先前工作轉移矩陣上的得分為留一語言外 $ρ{=}0.705$ 和 $R^2{=}0.49$,超過了 $ρ{=}0.62$ 的非類型學控制,這證實了僅基於類型學的預測能夠重建昂貴的測量跨語言轉移的能力。
該信號在留一腳本外和留一語系外的協議中仍然存在,因此腳本和語系的混淆並不能解釋這一效果。
通過將轉移分解為類型學項和資源與腳本偏差項,我們發現最佳來源排名對此偏差敏感。
相比之下,類型學不受此偏差影響,這使其成為一種零計算篩選工具,能夠用模型擬合取代數百次訓練運行。
我們的代碼可在 \href{https://github.com/dharmsen/typo-x-ling-transfer}{這裡} 獲得。
Disentangling Self-Distillation: Measuring and Modeling Acquisition and Retention
2609.39494v1 by Luis Zuin, Alexis Huet, Dario Rossi, Zied Ben Houidi
Self-distillation with privileged context adapts a language model from demonstrations by letting the model, once conditioned on a reference response, teach its context-free copy token by token. Our taxonomy reveals existing methods differ along three entangled axes: (i) the rollout source (student or teacher), (ii) the teacher coupling (frozen, or an exponential moving average of the student at some coupling rate) and (iii) the KL direction (reverse or forward), yet these axes are usually studied in fixed combinations and have led to conflicting conclusions. We formalize a unifying framework to encompass all self-distillation methods vs classic supervised fine-tuning: we train every combination of the three axes, on Qwen2.5-7B and Ministral-3-3B across ordinary and contradictory tasks, totaling 1,200 adaptation runs, to systematically investigate the impact of the above axes. We propose a controlled model of the same objective to explain the resulting acquisition-retention trade-offs. We find that (i) the rollout source matters mostly where the task contradicts the pretrained behavior: there teacher rollouts raise acquisition well above what student rollouts achieve, with almost no change in retention; (ii) the teacher coupling changes acquisition most, on every task: acquisition rises with the coupling rate, then falls past a task-specific rate; (iii) switching the KL direction costs retention in one model but not the other so which axis to tune first depends on the model. The controlled model reproduces the three trends.
摘要:自我蒸餾與特權上下文透過讓模型在參考回應的條件下,逐步教導其無上下文副本,來適應語言模型。我們的分類法揭示現有方法在三個交織的軸向上有所不同:(i)展開來源(學生或教師),(ii)教師耦合(凍結,或以某種耦合速率的學生指數移動平均)以及(iii)KL方向(反向或正向),然而這些軸通常以固定的組合進行研究,並導致相互矛盾的結論。我們正式化了一個統一框架,以涵蓋所有自我蒸餾方法與經典的監督微調:我們在Qwen2.5-7B和Ministral-3-3B上,針對普通和矛盾任務訓練三個軸的每一種組合,總計1,200次適應運行,以系統性地調查上述軸的影響。我們提出了一個相同目標的受控模型,以解釋所得到的獲取-保留權衡。我們發現:(i)展開來源主要在任務與預訓練行為矛盾時才重要:在這種情況下,教師的展開使獲取遠高於學生的展開,幾乎沒有保留的變化;(ii)教師耦合對每個任務的獲取影響最大:獲取隨著耦合速率上升,然後在特定任務的速率後下降;(iii)切換KL方向在一個模型中會影響保留,但在另一個模型中則不會,因此首先調整哪個軸取決於模型。受控模型重現了這三個趨勢。
From Speech to Editable Concepts: Probing Emotion Recognition with Concept Bottleneck Models
2609.39453v1 by Hezhao Zhang, Thomas Hain
Speech emotion recognition (SER) is the task of assigning emotion labels to utterances. Early systems relied on acoustic features, whereas recent approaches combine multiple modalities, most commonly speech and text. Still, performance remains poor on many datasets. Large language models (LLMs) have therefore attracted interest for SER, as they can process diverse inputs jointly with instructions. However, direct audio input raises questions of explainability. To address similar questions in image classification, concept bottleneck models were introduced. This work adapts concept bottlenecks to SER to examine how individual predictions depend on transcripts, acoustic descriptions and speaker attributes. Experiments test three LLMs on CREMA-D, IEMOCAP and MELD, with concepts extracted by separate tools. On scripted corpora, LLMs are strongly biased towards the transcript in the zero-shot setting, which lowers Macro-F1 from 27.8 to 5.8 on CREMA-D. Fine-tuning removes this bias, and the transcript raises Macro-F1 from 41.8 to 45.1. Removing speech rate changes 48% of Neutral predictions to Disgust on CREMA-D; removing intensity level on MELD changes predictions despite little change in Macro-F1. These findings show that aggregate performance changes alone do not capture the effects of concept removal on individual predictions.
摘要:語音情感識別(SER)是將情感標籤分配給話語的任務。早期的系統依賴於聲學特徵,而最近的方法則結合了多種模態,最常見的是語音和文本。儘管如此,許多數據集上的性能仍然較差。因此,大型語言模型(LLMs)引起了對SER的興趣,因為它們可以與指令共同處理多樣的輸入。然而,直接的音頻輸入引發了可解釋性的問題。為了解決圖像分類中的類似問題,引入了概念瓶頸模型。本研究將概念瓶頸應用於SER,以檢查個別預測如何依賴於文字稿、聲學描述和說話者屬性。實驗測試了三個LLM在CREMA-D、IEMOCAP和MELD上的表現,概念由不同的工具提取。在腳本語料庫中,LLM在零樣本設置中對文字稿有強烈的偏見,這使得CREMA-D上的Macro-F1從27.8降低到5.8。微調消除了這種偏見,文字稿使得Macro-F1從41.8提高到45.1。移除語音速率使CREMA-D上48%的中性預測變為厭惡;在MELD上移除強度水平則改變了預測,儘管Macro-F1幾乎沒有變化。這些發現表明,僅僅改變總體性能並不能捕捉到概念移除對個別預測的影響。
Towards Trustworthy AI for Glioma Diagnosis: A Task-Aware Evaluation of Uncertainty Quantification
2609.39429v1 by Gonzalo Esteban Mosquera Rojas, Sebastian R. van der Voort, Carolin M. Pirkl, Sandeep Kaushik, Marion Smits, Stefan Klein
Uncertainty Quantification (UQ) is a key requirement for trustworthy AI in high-stakes medical image analysis. In this work, we evaluate UQ in a multi-task Deep Learning framework for MRI-based glioma diagnosis that performs tumor segmentation and predicts IDH mutation status, 1p/19q co-deletion status, and tumor grade. Monte Carlo Dropout (MCD) is used for a detailed task-aware analysis of predictive, aleatoric, and epistemic uncertainty. We assess MC sample convergence, calibration, error detection, selective prediction, associations with segmentation performance, and the effect of voxel-wise uncertainty aggregation on case-level reliability. We also compare MCD with Deep Ensembles (DE) and Monte Carlo Deep Ensembles (MCDE), examine interactions between segmentation quality and classification, and evaluate a composite trust score integrating segmentation and classification uncertainty. Across tasks, uncertainty estimates supported meaningful error detection, while calibration depended on the dropout rate, with moderate rates yielding the most reliable probabilities. Uncertainty decomposition provided task-dependent interpretability but did not consistently improve error detection over predictive uncertainty alone. DE and MCDE showed comparable operational utility, with no method consistently dominating across tasks and metrics. The composite trust score did not consistently outperform classification uncertainty for selective prediction. Overall, our results provide a task-aware evaluation strategy and practical guidance for the development of trustworthy AI for glioma diagnosis.
摘要:不確定性量化(UQ)是高風險醫療影像分析中可信賴人工智慧的關鍵要求。
在這項工作中,我們評估了一個多任務深度學習框架中的UQ,該框架用於基於MRI的膠質瘤診斷,執行腫瘤分割並預測IDH突變狀態、1p/19q共同缺失狀態和腫瘤等級。
使用蒙特卡羅隨機失活(MCD)對預測性、隨機性和認知性不確定性進行詳細的任務感知分析。
我們評估了蒙特卡羅樣本收斂性、校準、錯誤檢測、選擇性預測、與分割性能的關聯,以及體素級不確定性聚合對案例級可靠性的影響。
我們還將MCD與深度集成(DE)和蒙特卡羅深度集成(MCDE)進行比較,檢查分割質量與分類之間的相互作用,並評估整合分割和分類不確定性的綜合信任分數。
在各項任務中,不確定性估計支持有意義的錯誤檢測,而校準則依賴於隨機失活率,中等率產生最可靠的概率。
不確定性分解提供了任務依賴的可解釋性,但並未始終改善僅依賴預測不確定性的錯誤檢測。
DE和MCDE顯示出可比的操作效用,沒有一種方法在各任務和指標中始終佔優。
綜合信任分數在選擇性預測中並未始終優於分類不確定性。
總體而言,我們的結果提供了一種任務感知的評估策略和實用指導,旨在為膠質瘤診斷的可信賴人工智慧發展提供支持。
The Golden Path Hypothesis: Reusable Schedules in Diffusion Caching
2609.39343v1 by Dong Wang, Wenwu Tang, Francesco Corti, Yun Cheng, Lothar Thiele, Olga Saukh
Diffusion caching accelerates generation by replacing transformer computation with cached or predicted features at selected denoising steps. We introduce the Golden Path Hypothesis (GPH): under fixed inference conditions, prompt-independent cache schedules can achieve final-output quality comparable to the best prompt-specific schedules across prompts. We investigate the GPH across ten caching methods, four image and video models, and three cache ratios. Prompt-adaptive methods repeatedly select a small number of schedules, and reusing their most frequent schedules on new prompts closely matches the quality of prompt-specific choices. Exhaustive evaluation of 1.4 million schedules on four examples further identifies prompt-independent schedules that remain competitive on unseen prompts. To explain this transfer, we analyze denoising trajectories and the accumulation of caching errors. Latent-state trajectories exhibit similar structures across datasets and seeds, while an exact error decomposition shows that accumulated effects of earlier errors predict final latent-state error better than local approximation errors. This motivates searching for end-to-end schedules using final-output quality. With only a small set of examples, the resulting golden paths transfer across prompts and datasets, and can be tuned to the desired quality objective, including reconstruction fidelity or perceptual similarity.
摘要:擴散快取通過在選定的去噪步驟中用快取或預測的特徵替代Transformer計算來加速生成。
我們提出了黃金路徑假設(GPH):在固定的推理條件下,與最佳的提示特定計劃相比,與提示無關的快取計劃可以達到相似的最終輸出質量。
我們在十種快取方法、四種圖像和視頻模型以及三種快取比率上研究了GPH。
提示自適應方法重複選擇少量計劃,並在新提示上重用其最頻繁的計劃,這與提示特定選擇的質量非常接近。
對140萬個計劃在四個示例上的全面評估進一步識別出在未見提示上仍具競爭力的與提示無關的計劃。
為了解釋這一轉移,我們分析了去噪軌跡和快取錯誤的累積。
潛在狀態軌跡在數據集和種子之間顯示出相似的結構,而精確的錯誤分解顯示,早期錯誤的累積效應比局部近似錯誤更好地預測最終潛在狀態錯誤。
這促使我們尋找使用最終輸出質量的端到端計劃。
僅用一小組示例,結果的黃金路徑在提示和數據集之間轉移,並可以調整以達到所需的質量目標,包括重建保真度或感知相似性。
A Time-Aware Bag-of-Receptive-Fields for Interpretable Irregular Time Series Classification
2609.39268v1 by Francesco Spinnato
Irregular time series, characterized by non-uniform sampling intervals, missing observations, and variable lengths, are ubiquitous in healthcare, mobility, and environmental monitoring, yet effective and interpretable classifiers for this setting are limited. Existing approaches often rely on imputation, which can obscure the temporal structure of the data, or require complex neural architectures that are opaque and difficult to explain. In this work, we extend the Bag-Of-Receptive-Fields (BORF), a fast, deterministic, and interpretable transform for time series, to the irregular setting. Our key contribution is a time-weighted normalization scheme in which each observation is weighted proportionally to its associated time delta, making pattern extraction sensitive to the actual temporal distribution of samples rather than only their index position. This requires deriving an efficient sliding-window recurrence for the time-weighted standard deviation, preserving the linear time complexity of BORF. We benchmark the resulting method against state-of-the-art irregular time series classifiers on datasets from the PYRREGULAR repository, demonstrating competitive classification performance with the added benefit of human-interpretable explanations.
摘要:不規則時間序列的特徵是取樣間隔不均、缺失觀測值和變化的長度,這在醫療、移動性和環境監測中隨處可見,但在這種情境下有效且可解釋的分類器卻有限。現有的方法通常依賴於插補,這可能會掩蓋數據的時間結構,或需要複雜的神經架構,這些架構不透明且難以解釋。在這項工作中,我們將快速、確定性且可解釋的時間序列變換——感受野包(Bag-Of-Receptive-Fields, BORF)擴展到不規則的情境。我們的主要貢獻是一種時間加權標準化方案,其中每個觀測值的權重與其相關的時間增量成比例,這使得模式提取對樣本的實際時間分佈敏感,而不僅僅是它們的索引位置。這需要導出一種高效的滑動窗口重複計算時間加權標準差,從而保持BORF的線性時間複雜度。我們將所得到的方法與來自PYRREGULAR數據庫的最先進不規則時間序列分類器進行基準測試,展示了具有競爭力的分類性能,並附帶可供人類解釋的解釋。
When Harnesses Lose the Signal: Causal Evaluation of Recovery in LLM Agents
2610.00372v1 by Shuyao Xiao, Shengling Wang, Xuan Chen, Ke Chao, Ming Cui, Feifei Qian, Chaoyang Mei, Fanlin Meng, Ziming Yu, Junxi Yin
Large language model agents rely on external harnesses to pass information between the model and its environment and to recover from execution errors. Yet recovery is usually judged only by average task success. This hides an important tension. The same operation can rescue a failing trajectory or disrupt one that would otherwise succeed. We frame recovery as a causal decision problem. Starting from the same execution state, we compare what happens with and without recovery, separate rescue from harm, and study how the value of recovery changes over time. We then introduce the Causal Intervention Router (CIR), a lightweight policy that uses information available before recovery to decide when intervention is worthwhile. On long-horizon ALFWorld tasks with Qwen3-14B, CIR raises success from 70.33% to 73.33%, a gain of 3.00 percentage points. It leaves all evaluated trajectories with correct observations untouched. Additional controls show that the benefit of recovery cannot be explained solely by the new observation returned by the environment. These results provide a practical way to evaluate recovery and apply it selectively.
摘要:大型語言模型代理依賴外部裝置在模型與其環境之間傳遞信息並從執行錯誤中恢復。
然而,恢復通常僅根據平均任務成功率來評估。
這隱藏了一個重要的緊張關係。
同一操作可以挽救失敗的軌跡,或破壞本來會成功的軌跡。
我們將恢復框架設置為一個因果決策問題。
從相同的執行狀態開始,我們比較有無恢復的情況下發生的事情,將救援與傷害分開,並研究恢復的價值如何隨時間變化。
然後我們介紹了因果干預路由器(CIR),這是一種輕量級策略,利用恢復前可用的信息來決定何時進行干預是值得的。
在使用Qwen3-14B的長期ALFWorld任務中,CIR將成功率從70.33%提高到73.33%,增幅為3.00個百分點。
它保持所有評估的軌跡的正確觀察不變。
額外的控制顯示,恢復的好處不能僅僅用環境返回的新觀察來解釋。
這些結果提供了一種實際的方法來評估恢復並選擇性地應用它。
Loop-Free Inverse Reinforcement Learning via Sequential Value Recovery with Q-Score Matching
2609.38955v1 by Yang chen, Yitan Zhang, Michael Witbrock, Shuyue Hu
Inverse Reinforcement Learning (IRL) aims to recover a reward function that explains expert demonstrations. Existing IRL methods typically rely on a bi-level optimization procedure that alternates between reward learning and policy optimization, leading to substantial computational burden and training instability. In this work, we introduce a different route that eliminates policy optimization entirely by leveraging diffusion policies. Our key insight is that a diffusion policy encodes the action-gradient structure of the optimal soft Q function, enabling reward learning to be cast as a sequence of value recovery problems, thereby allowing us to bypass reward-policy loops inherent in prior IRL methods. Specifically, our method proceeds in three stages: (I) recovering the optimal soft Q function via action-gradient matching and estimating the corresponding soft value function (LogSumExp of Q values) in a way inspired by Gumbel regression; (II) calibrating these soft values by inferring a state-dependent offset; (III) extracting the reward by enforcing Bellman consistency. This leads to Loop-Free Inverse Reinforcement Learning (LFIRL), a fully offline algorithm that operates in a simple, loop-free, and sequential manner. LFIRL is simple to implement and significantly improves training efficiency while maintaining strong reward recovery performance. Empirically, across Maze, Franka Kitchen, Adroit Hand Pen, and Push-T benchmarks, LFIRL achieves 2-3x speedup over the fastest baselines, while matching or surpassing state-of-the-art methods in reward recovery quality.
摘要:逆向強化學習(IRL)旨在恢復解釋專家示範的獎勵函數。現有的IRL方法通常依賴於一種雙層優化程序,在獎勵學習和策略優化之間交替進行,這導致了相當大的計算負擔和訓練不穩定性。在這項工作中,我們提出了一種不同的路徑,通過利用擴散策略完全消除了策略優化。我們的關鍵見解是,擴散策略編碼了最佳軟Q函數的行動梯度結構,使得獎勵學習可以被視為一系列價值恢復問題,從而使我們能夠繞過先前IRL方法中固有的獎勵-策略循環。具體而言,我們的方法分為三個階段:(I)通過行動梯度匹配恢復最佳軟Q函數,並以受到Gumbel回歸啟發的方式估計相應的軟值函數(Q值的LogSumExp);(II)通過推斷狀態依賴的偏移來校準這些軟值;(III)通過強制執行Bellman一致性來提取獎勵。這導致了無循環逆向強化學習(LFIRL),這是一種完全離線的算法,以簡單、無循環和順序的方式運行。LFIRL實現簡單,顯著提高了訓練效率,同時保持強大的獎勵恢復性能。在Maze、Franka Kitchen、Adroit Hand Pen和Push-T基準測試中,LFIRL在速度上比最快的基準提高了2-3倍,同時在獎勵恢復質量上達到或超越了最先進的方法。
Reasoning Externalization for Faithful Large Language Model Narratives of Stock Return Predictions
2609.38869v1 by Sujung Kim, Seung Hwan Cho, Sangjin Park, Young-Min Kim
In finance, interpreting machine learning predictions is essential, yet the numerical outputs of explainable AI can be difficult for non-experts to understand. While large language models (LLMs) can translate these outputs into natural language, they may produce errors when inferring numerical changes and feature relations. We propose an LLM narrative framework for cross-sectional stock return prediction that combines temporal Shapley additive explanations (SHAP) evidence with historical regime analogs. Temporal evidence tracks changes in the normalized global SHAP importance of an XGBoost model over six months. Historical analogs are past periods with similar changes in SHAP importance, their model performance and subsequent market returns are provided as comparative context. Using this framework, we conduct a controlled study of progressive reasoning externalization, sequentially providing raw SHAP sequences, deterministic temporal descriptors, and feature relations. Each generated claim is verified against provenance-linked evidence. Across Qwen3, externalizing numerical and relational reasoning improved evidence faithfulness as well as temporal and relational accuracy. Evidence faithfulness increased from 0.696 to 0.996 for Qwen3-32B-Instruct. While historical analogs did not improve structured automatic faithfulness, they received higher human-rated usefulness scores. These results suggest that externalizing verifiable reasoning enhances narrative faithfulness and that historical context adds interpretive value.
摘要:在金融領域,解釋機器學習預測是至關重要的,但可解釋人工智慧的數值輸出對於非專家來說可能難以理解。雖然大型語言模型(LLMs)可以將這些輸出轉換為自然語言,但在推斷數值變化和特徵關係時,它們可能會產生錯誤。我們提出了一個LLM敘事框架,用於橫斷面股票回報預測,該框架結合了時間性Shapley加法解釋(SHAP)證據和歷史制度類比。時間性證據跟蹤XGBoost模型在六個月內的標準化全球SHAP重要性的變化。歷史類比是過去在SHAP重要性上有類似變化的時期,提供其模型表現和隨後市場回報作為比較背景。利用這個框架,我們進行了一項受控研究,對進步推理外化進行了逐步的探討,依次提供原始SHAP序列、確定性時間描述符和特徵關係。每個生成的主張都與來源鏈接的證據進行驗證。在Qwen3上,外化數值和關係推理提高了證據的真實性以及時間和關係的準確性。對於Qwen3-32B-Instruct,證據的真實性從0.696提高到0.996。雖然歷史類比並未改善結構化自動真實性,但它們獲得了更高的人類評價的有用性分數。這些結果表明,外化可驗證的推理增強了敘事的真實性,而歷史背景則增添了解釋價值。
SkillSeek: Revisiting Agent Skill Retrieval at Marketplace Scale
2609.38822v1 by Guanqun Yang, Wenlong Zhang, Tian Shi, Ping Wang
Anthropic's Agent Skills package reusable procedural know-how for an LLM agent into SKILL.md directories, and open-source aggregations have grown past 230,000 skills, making selection rather than authoring the bottleneck. The standing answer in the literature outsources selection to the agent itself: an LLM-mediated retrieval loop that rewrites queries and refines candidates inside the agent's decision loop, paying LLM tokens on every task. We present SkillSeek, an open-source two-stage skill retriever built from the standard IR recipe (a BGE-base bi-encoder feeding a small cross-encoder, exposed over MCP). Across a $4 \times 11$ grid of pool, backbone, and method on the 89-task SkillsBench benchmark, SkillSeek reaches observed parity with the LLM-mediated loop of Liu et al. at essentially no extra cost: plain bm25 alone records a pass rate at or above their refined loop on three of four settings, and a small cross-encoder covers the remaining difference on the fourth. A first-stage recall ceiling explains the pattern, and total per-trial spend drops from USD 51.30 to USD 27.54 (within fifty cents of the no-skill baseline). Under the SkillsBench tasks and OpenHands harness we tested, this positions the standard IR recipe as a strong default for agent-skill retrieval, with LLM-mediated alternatives a natural fit for cases where deterministic methods fall short.
摘要:Anthropic 的 Agent Skills 套件將可重用的程序知識整合到 LLM 代理的 SKILL.md 目錄中,開源聚合已經超過 230,000 種技能,使得選擇而非創作成為瓶頸。文獻中的現有解答將選擇外包給代理本身:一個 LLM 媒介的檢索循環,重寫查詢並在代理的決策循環內精煉候選者,對每個任務支付 LLM 代幣。我們提出了 SkillSeek,一個基於標準 IR 食譜的開源兩階段技能檢索器(使用 BGE 基礎的雙編碼器供應小型交叉編碼器,通過 MCP 暴露)。在 89 任務的 SkillsBench 基準上,SkillSeek 在 $4 \times 11$ 的池、骨幹和方法網格中達到了與 Liu 等人的 LLM 媒介循環的觀察平衡,幾乎沒有額外成本:單純的 bm25 在四個設置中的三個上記錄的通過率達到或超過他們的精煉循環,而小型交叉編碼器則覆蓋了第四個設置的剩餘差異。一階段召回上限解釋了這一模式,並且每次試驗的總支出從 51.30 美元降至 27.54 美元(在無技能基準線的五十美分之內)。在我們測試的 SkillsBench 任務和 OpenHands 環境下,這使得標準 IR 食譜成為代理技能檢索的強大默認選擇,而 LLM 媒介的替代方案則自然適合於確定性方法無法滿足的情況。
When Reasoning Goes Astray: Attention Dynamics of Uncontrolled Reasoning
2609.38817v1 by Yuanhe Zhang, Ziwei Wang, Jie Ren, Haoran Gao, Zhenhong Zhou, Fanyu Meng, Cong Wu, Li Sun, Sen Su
Large reasoning models (LRMs) improve performance on complex tasks through extended reasoning, yet the same process can degenerate into redundant verification and persistent generation loops. Such uncontrolled reasoning increases inference cost and creates risks of resource exhaustion and service degradation. However, existing mitigations largely truncate long outputs or react to surface repetition, and thus fail to distinguish normal thinking from uncontrolled reasoning or explain how benign reasoning degenerates into harmful behavior. In this paper, we operationalize LRM generation as four states and further introduce Reasoning-state Analysis via Dynamic Attention Responses (RADAR), which identifies the current reasoning state in real time and characterizes how effective reflection can develop into uncontrolled generation. Guided by RADAR's analysis, we further realign abnormal attention distributions toward patterns observed in normal requests and examine how this correction affects excessive reflection and persistent looping. Temporal analyses show that uncontrolled reasoning is characterized by attention distributions that deviate from normal generation, with abnormal trends becoming detectable before repetition begins. Correcting these deviations through Attention Realignment consistently reduces looping while largely preserving benign performance. Together, RADAR provide a mechanistic account of how reasoning becomes uncontrolled, offering actionable guidance for identifying critical failure stages and designing targeted runtime interventions.
摘要:大型推理模型(LRMs)透過擴展推理來提高在複雜任務上的表現,然而相同的過程可能會退化為冗餘的驗證和持續的生成循環。這種不受控制的推理增加了推斷成本,並創造了資源耗盡和服務降級的風險。然而,現有的緩解措施主要是截斷長輸出或對表面重複作出反應,因此未能區分正常思考與不受控制的推理,或解釋良性推理如何退化為有害行為。在本文中,我們將LRM生成操作化為四個狀態,並進一步引入動態注意力反應下的推理狀態分析(RADAR),該方法實時識別當前的推理狀態,並描述有效反思如何發展成不受控制的生成。在RADAR的分析指導下,我們進一步重新調整異常的注意力分佈,朝向在正常請求中觀察到的模式,並檢查這一修正如何影響過度反思和持續循環。時間分析顯示,不受控制的推理以偏離正常生成的注意力分佈為特徵,異常趨勢在重複開始之前就變得可檢測。通過注意力重新調整來修正這些偏差,持續減少循環,同時在很大程度上保持良性表現。總的來說,RADAR提供了一個機制性解釋,說明推理如何變得不受控制,並提供可行的指導,以識別關鍵失敗階段並設計針對性的運行時干預。
Understanding Off- vs On-Policy Distillation: A Tale of Distinct Training Objectives
2609.38666v1 by Qiwei Di, Xuheng Li, Kaixuan Ji, Chenggong Zhang, Heyang Zhao, Quanquan Gu
On-policy distillation (OPD) learns from teacher feedback on student-generated responses and has shown promise in reducing forgetting relative to supervised fine-tuning (SFT). However, its benefits and fragility remain incompletely understood. We study sequential distillation from multiple teachers, where the student minimizes its average divergence from the teachers. Forward Kullback--Leibler (KL) divergence yields a weighted arithmetic mixture, while reverse KL yields a normalized weighted geometric aggregate. We develop algorithms that learn these targets under off-policy and on-policy feedback, respectively, establishing logarithmic regret bounds in the tabular setting and extending the analysis to function approximation. By analyzing these aggregation targets, we identify mechanisms that help explain both the benefits and fragility of OPD. Relative to forward KL, reverse KL can better retain a confident expert's preferences under uninformative feedback, but is more sensitive to teachers that assign very low probabilities to correct responses. Its token-level conditionals also reveal a dependence on continuation distributions that can favor incorrect prefixes over long horizons.
摘要:在政策蒸餾 (OPD) 中,學生根據教師的反饋學習生成的回應,並在減少相對於監督微調 (SFT) 的遺忘方面顯示出潛力。 然而,它的好處和脆弱性仍然未完全理解。 我們研究來自多位教師的序列蒸餾,學生最小化與教師的平均差異。 前向 Kullback--Leibler (KL) 散度產生加權算術混合,而反向 KL 則產生標準化的加權幾何聚合。 我們開發了在離政策和在政策反饋下學習這些目標的算法,分別在表格設置中建立對數遺憾界限,並將分析擴展到函數近似。 通過分析這些聚合目標,我們確定了幫助解釋 OPD 的好處和脆弱性的機制。 相對於前向 KL,反向 KL 更能在無信息反饋下保留自信專家的偏好,但對於給正確回應分配非常低概率的教師則更敏感。 它的標記級條件也揭示了對延續分佈的依賴,這可能在長期內偏向於不正確的前綴。
Defining and Categorising Human-AI Interactions in Clinical Trials: A Multidimensional Human-AI Classification Approach
2609.38559v1 by Sandra Woolley, Tim Collins, Khalid Khattak, Illia Chernomorets, Ariane Arevalo, Chris Richardson
This paper examines human-AI interactions (HAIIs) in clinical trials and presents a multidimensional categorisation framework that classifies interactions according to AI tasks, human-AI relationships, interaction configurations and interacting human groups. We define HAII, examine existing taxonomies and extend existing categorisation approaches through this novel multidimensional framework. We purposively sampled 15 clinical trials from a previously reported dataset. Each trial was independently categorised by two human reviewers and six large language model (LLM) classifiers. The proposed categorisation provides a structured method for the consistent identification, comparison and synthesis of human-AI interactions across clinical-trial records. The framework is intended to support more consistent comparison and synthesis of AI-related clinical trials and to make explicit the different forms of human involvement associated with AI interventions. The results demonstrate the potential for LLM-assisted categorisation while indicating the continuing importance of human judgement where trial records are incomplete or ambiguous. The principal contribution is a proposed multidimensional framework that brings together AI tasks, human-AI relationships, interaction configurations and interacting human groups within a single approach designed for clinical-trial records. Its significance lies in its potential to support more systematic identification, comparison and synthesis of how humans and AI interact in clinical trials.
摘要:這篇論文探討了臨床試驗中的人類-人工智慧互動(HAIIs),並提出了一個多維度的分類框架,根據人工智慧任務、人類-人工智慧關係、互動配置和互動人群對互動進行分類。
我們定義了HAII,檢視現有的分類法,並通過這個新穎的多維框架擴展現有的分類方法。
我們有目的地從先前報告的數據集中抽取了15個臨床試驗。
每個試驗由兩位人類評審和六個大型語言模型(LLM)分類器獨立進行分類。
所提出的分類提供了一種結構化的方法,用於在臨床試驗記錄中一致地識別、比較和綜合人類-人工智慧互動。
該框架旨在支持對人工智慧相關臨床試驗的更一致的比較和綜合,並明確不同形式的人類參與與人工智慧干預相關聯。
結果顯示了LLM輔助分類的潛力,同時指出在人類判斷仍然重要的情況下,當試驗記錄不完整或模糊時,仍需依賴人類判斷。
主要貢獻是一個提出的多維框架,將人工智慧任務、人類-人工智慧關係、互動配置和互動人群整合在一個針對臨床試驗記錄的單一方法中。
其重要性在於它能支持對人類和人工智慧在臨床試驗中互動的更系統的識別、比較和綜合。
Demographic Pluralism: Inference-Time Modeling of Pluralistic Human Preference Distributions
2609.38555v1 by Meng-Chen Wu, Qipin Chen, Ansh Jain, Tess Wood, Zhe Du, Si-Chi Chin
Large language models (LLMs) are increasingly used in culturally sensitive settings, where alignment requires representing diverse preferences within populations. Yet existing methods model populations at coarse demographic or community levels and overlook within-group variation. We introduce Demographic Pluralism, an inference-time framework that estimates population-level opinion distributions without opinion-distribution training data or task-specific fine-tuning by generating multiple perspectives within demographically grounded groups. Across four backbones on GlobalOpinionQA and VITAL, it reduces Jensen-Shannon distance by 8.4%-26.4% over Modular Pluralism. Among weighted, equal-weighted, and inverse-weighted aggregation, equal weighting performs best overall; group-level error also increases with group weight, helping explain weighted aggregation's weaker performance.
摘要:大型語言模型(LLMs)在文化敏感的環境中越來越被使用,這些環境中的對齊需要代表人口中的多樣化偏好。
然而,現有的方法在粗略的人口或社區層面建模人口,並忽略了群體內的變異。我們引入了人口多元主義,這是一種推斷時框架,通過在以人口為基礎的群體中生成多個視角,來估計人口層級的意見分佈,而不需要意見分佈的訓練數據或特定任務的微調。
在 GlobalOpinionQA 和 VITAL 的四個基礎模型中,它將詹森-香農距離減少了 8.4%-26.4%,相較於模組化多元主義。在加權、等權重和反向加權聚合中,等權重的表現最佳;群體層級的誤差也隨著群體權重的增加而增加,這有助於解釋加權聚合的較弱表現。
Does This Action Still Explain the Task? Reverse Scoring for Diffusion Language Model Agents
2609.38536v1 by Jiacheng Qiu, Christopher E. Mower, Jan Peters, Haitham Bou-Ammar, Matthieu Zimmer
Diffusion-based large language models (dLLMs) promise to break the sequential latency bottleneck of autoregressive agents through parallel decoding, but recent evaluations show this efficiency does not transfer to embodied agentic competence: dLLM-backed agents repeatedly fall into retry loops, re-issuing an action long after it has failed. We give a mechanistic account of this failure and a training-free remedy. We trace the retry loop to the adaptivity of masked decoding: the sampler commits the positions it is most confident about and defers the uncertain ones, and at a failure state the context already offers a confident fill for the deferred decision, i.e. the failed action itself, so the retry is committed without the failure feedback ever being confronted. We model the resulting distortion of the action distribution as a task-blind corruption: contextually salient actions (e.g., the action just taken) receive inflated probability by a factor that depends on the state and the action but not on the task. Under this model, we analyse an invariance proposition: the task-blind factor cancels exactly from the reverse conditional, i.e. the likelihood of the task given the state and a candidate action, which coincides with the task posterior of an idealized uncorrupted model. Masked dLLMs evaluate the reverse conditional natively, unlike autoregressive models, by masking the task tokens and denoising, at the cost of a few parallel passes per candidate. We instantiate the rule as Reflect Reverse and evaluate it on four multi-turn embodied benchmarks, where it improves task success and progression rates over forward-scoring baselines.
摘要:擴散基的大型語言模型 (dLLMs) 承諾通過並行解碼打破自回歸代理的序列延遲瓶頸,但最近的評估顯示這種效率並未轉移到具身代理的能力上:基於 dLLM 的代理反覆陷入重試循環,在動作失敗很久之後重新發出動作。我們對這一失敗給出了機制性解釋和無需訓練的補救措施。我們將重試循環追溯到掩蔽解碼的適應性:採樣器承諾其最有信心的位置,並推遲不確定的位置,而在失敗狀態下,上下文已經為推遲的決策提供了一個自信的填充,即失敗的動作本身,因此重試是在從未面對失敗反饋的情況下進行的。我們將動作分佈的扭曲建模為一種任務盲腐敗:在上下文中顯著的動作(例如,剛剛執行的動作)因為一個取決於狀態和動作但不依賴於任務的因子而獲得了膨脹的概率。在這個模型下,我們分析了一個不變性命題:任務盲因子在反向條件中恰好抵消,即給定狀態和候選動作的任務的可能性,這與理想化的未腐敗模型的任務後驗相吻合。掩蔽 dLLMs 本土評估反向條件,與自回歸模型不同,通過掩蔽任務標記和去噪,在每個候選者的幾次並行通過的成本下。我們將該規則具體化為反射反向,並在四個多輪具身基準上進行評估,在這些基準上,它改善了任務成功率和進展率,超過了正向評分的基準。
Caption-Mediated Perceived-Safety Estimation for Pedestrian Routing
2609.38479v1 by Simon Parkinson, Paloma Liu, Wei Zheng, Mohammadreza Sheikhfathollahi
This paper presents an explainable approach to pedestrian routing, in which perceived safety is estimated from street-level imagery through an explicit natural-language intermediate representation. A vision--language model caption is generated and stored before any scoring is undertaken, and the perceived-risk class is derived entirely from structured features of that stored text, so that every segment score remains inspectable by the user. Nine captioning conditions across five model families are benchmarked against a direct Contrastive Language--Image Pre-training (CLIP) image-embedding baseline under an identical downstream pipeline, and the caption-mediated representation is found to reach parity with the image embedding rather than to trail it. The approach was deployed over 654,115 images covering 36 electoral wards in two locations in Northern England (Manchester and Huddersfield). Independent field validation against 3,669 locally collected ratings of 494 images across 70 participant sessions established agreement that is statistically significant but modest, at $r=0.262$, against a measured noise ceiling of 0.737 imposed by disagreement between raters. A single-use confirmatory test then found that a pipeline 44\% stronger on the supervised benchmark did not produce measurable improvement in the field ($r=0.250$, $p=0.84$), so the benchmark gains did not predict the deployment gains in this case. Routing behaviour varies systematically with journey length. There is negligible change below 1\,km, reaching a median increase of 12.78\% in low-risk route length for a median detour of 2.73\% on journeys of 3 to 6 km.
摘要:這篇論文提出了一種可解釋的行人路徑規劃方法,其中感知安全性是通過明確的自然語言中介表示從街景影像中估算得出的。
在進行任何評分之前,生成並存儲一個視覺-語言模型的標題,並且感知風險類別完全源自於該存儲文本的結構特徵,以便每個段落的分數都能被用戶檢查。
在相同的下游流程下,對五個模型家族中的九個標題條件進行了基準測試,與直接的對比語言-圖像預訓練(CLIP)圖像嵌入基準進行比較,發現標題中介表示的表現達到了與圖像嵌入相當的水平,而不是落後於它。
該方法在英格蘭北部的兩個地點(曼徹斯特和哈德斯菲爾德)涵蓋了654,115張影像,涉及36個選區。
對494張影像在70個參與者會議中收集的3,669條本地評分進行的獨立現場驗證顯示,達成的協議在統計上顯著但適度,相關係數為$r=0.262$,而評分者之間的不一致造成的噪音上限為0.737。
隨後進行的一次單次確認測試發現,在監督基準上強度提高44\%的流程在現場並未產生可測量的改善($r=0.250$,$p=0.84$),因此在這種情況下,基準增益並未預測部署增益。
路徑行為隨著行程長度系統性變化。
在1公里以下幾乎沒有變化,對於3到6公里的行程,低風險路徑長度的中位數增加達到12.78\%,而中位數繞行為2.73\%。
Aegis: Generative Gradient Masking for Privacy-Preserving Medical Federated Learning
2609.38339v1 by Chaoyu Zhang, Shanghao Shi, Heng Jin, Ning Wang, Y. Thomas Hou, Wenjing Lou
Federated learning (FL) has become a foundational paradigm for multi-institutional medical AI, allowing hospitals and research centers to jointly train diagnostic models without exchanging patient records. This privacy promise, however, is increasingly contested: a malicious or honest-but-curious server can launch model inversion attacks (MIAs) that reconstruct private patient images directly from shared model updates, and recent scalable, closed-form attacks penetrate even secure aggregation at clinically realistic batch sizes. Existing defenses face an unsatisfactory dilemma. Gradient-perturbation methods such as differential privacy and pruning trade away the diagnostic accuracy on which clinical reliability depends, while cryptographic protocols add system complexity yet still leave updates exposed to these scalable attacks. We propose Aegis, a principled client-side defense that breaks this dilemma without perturbing patient data or modifying the FL protocol. Our key insight is that the success of every known MIA is fundamentally bounded by the local batch size relative to the model's leakage capacity; once this limit is exceeded, distinct samples collide and reconstructions collapse into indistinguishable mixtures. Aegis turns this universal bottleneck into a defense: each client superimposes onto its real update a masking gradient computed on locally synthesized, task-relevant data, deliberately pushing the effective batch beyond the attack's recovery capacity. We complement the design with theoretical convergence guarantees under standard convex assumptions and evaluate Aegis on MNIST, CIFAR-10, and three MedMNIST modalities (chest X-ray, abdominal CT, colon pathology). Aegis neutralizes three state-of-the-art MIAs while preserving model utility and incurring only modest overhead, offering a practical privacy primitive for medical FL.
摘要:聯邦學習(FL)已成為多機構醫療人工智慧的基礎範式,使醫院和研究中心能夠共同訓練診斷模型,而無需交換病歷記錄。 然而,這一隱私承諾正受到越來越多的質疑:一個惡意或誠實但好奇的伺服器可以發動模型反演攻擊(MIA),直接從共享的模型更新中重建私人病人影像,而最近可擴展的封閉形式攻擊甚至能夠穿透臨床現實批次大小的安全聚合。 現有的防禦面臨著不令人滿意的困境。 像差分隱私和修剪這樣的梯度擾動方法犧牲了臨床可靠性所依賴的診斷準確性,而加密協議則增加了系統的複雜性,卻仍然讓更新暴露於這些可擴展的攻擊之下。 我們提出了Aegis,一種原則性的客戶端防禦,打破了這一困境,而不擾動病人數據或修改FL協議。 我們的關鍵見解是,每個已知的MIA的成功在根本上受到相對於模型泄漏能力的本地批次大小的限制;一旦超過這一限制,不同的樣本將發生碰撞,重建將崩潰為無法區分的混合物。 Aegis將這一普遍瓶頸轉化為防禦:每個客戶端在其真實更新上疊加一個基於本地合成的、與任務相關的數據計算出的掩蔽梯度,故意將有效批次推向超過攻擊的恢復能力。 我們在標準凸假設下補充了理論收斂保證,並在MNIST、CIFAR-10和三種MedMNIST模態(胸部X光、腹部CT、結腸病理)上評估了Aegis。 Aegis中和了三種最先進的MIA,同時保留了模型的效用,並僅產生適度的開銷,為醫療FL提供了一種實用的隱私原語。
Does Local Video Understanding Transfer Across Encounters? The EgoGears Benchmark
2609.37938v1 by Yuedong Tan, Lei Qi, Yu Liu, Di Wen, Ruiping Liu, Xiaoye Wang, Yufan Chen, Junwei Zheng, Chengzhi Wu, Chen Zhang, Zhihang Chen, Haiwen Sun, Zongwei Wu, Radu Timofte, Danda Pani Paudel, Kunyu Peng
Embodied systems must make knowledge acquired during one encounter usable in another despite changes in viewpoint, motion, and illumination. Yet aggregate cross-video accuracy conflates failures of local perception with failures to preserve observation identity, establish correspondence, and compose evidence, obscuring whether local video understanding actually transfers. We introduce EgoGears, a complementary single- and multi-video benchmark designed to diagnose this transition. It contains 567 single-video and 1,487 multi-video questions derived from 126 human-collected egocentric recordings covering 39 outdoor routes. Repeated traversals across movement speeds and lighting conditions ground comparisons in shared physical environments; 531 questions require alignment across independent recordings. Single-video questions measure the local visual, spatial, and motion evidence available to a model, while multi-video questions test whether evidence remains bound to the correct observation and can be composed into consistent route relationships. We report 29 single-video and 31 multi-video MLLM configurations across six model families in the main leaderboard. Among the 20 configurations evaluated comparably on both splits, every model performs worse on multi-video questions, with a mean decrease of 22.5 percentage points, and the gap persists when answer format and scoring are held fixed. The gap is not explained simply by additional videos or recording boundaries. The central bottlenecks are observation--evidence binding and ordered route-state tracking. The code and benchmark are publicly available at https://github.com/lei-qi-233/EgoGears.
摘要:具身系統必須使在一次遭遇中獲得的知識能夠在另一個遭遇中使用,儘管視角、運動和照明條件有所改變。
然而,綜合跨視頻的準確性將局部感知的失敗與未能保持觀察身份、建立對應關係和組合證據的失敗混淆,模糊了局部視頻理解是否真的能夠轉移。
我們介紹了EgoGears,一個補充性的單視頻和多視頻基準,旨在診斷這一過渡。
它包含567個單視頻和1,487個多視頻問題,這些問題源自126個人類收集的自我中心錄音,涵蓋39條戶外路線。
在不同的運動速度和光照條件下的重複遍歷,使比較基於共享的物理環境;531個問題要求在獨立錄音之間進行對齊。
單視頻問題測量模型可用的局部視覺、空間和運動證據,而多視頻問題則測試證據是否仍然與正確的觀察綁定,並且能夠組合成一致的路徑關係。
我們在主要排行榜上報告了六個模型系列中的29個單視頻和31個多視頻MLLM配置。
在20個在兩個分組中進行可比評估的配置中,每個模型在多視頻問題上的表現都較差,平均下降22.5個百分點,並且當答案格式和評分保持固定時,這一差距仍然存在。
這一差距並不能簡單地用額外視頻或錄製邊界來解釋。
主要瓶頸是觀察—證據綁定和有序路徑狀態跟踪。
代碼和基準可在https://github.com/lei-qi-233/EgoGears上公開獲取。
OmniVCBench: Benchmarking Evidence-Grounded Multimodal Reasoning Towards AI Virtual Cells
2609.37773v1 by Manyu Li, Xunkai Li, Yongfu Xiong, Yi Liu, Rong-Hua Li, Guoren Wang
Artificial Intelligence Virtual Cells (AIVCs) are envisioned as scientific agents that simulate cellular responses, explain underlying mechanisms, and support hypothesis-driven discovery. Existing AIVC benchmarks, however, operate primarily at the simulation layer, motivating complementary evaluation of how models interpret experimental evidence and formulate biological hypotheses. We introduce OmniVCBench, a figure-centric, source-traceable benchmark for the interpretation component of an AIVC. It contains 6,077 curated single- and multi-subfigure question--answer pairs derived from figures and experimental contexts in the scientific literature. Guided by Bloom's taxonomy, we instantiate interpretation-layer counterparts of the AIVC Predict--Explain--Discover agenda through three scientific reasoning tasks. We further introduce AIVC-Judge, a task-conditioned MLLM-as-a-judge framework with category-specific, reference-aware rubrics for evaluating open-ended responses. A complementary Model-Derived Hard-Negative Mining (MDHNM) strategy converts plausible errors observed during model inference into MCQ distractors for lower-cost evaluation. Within the evaluated heterogeneous model pool, MCQ accuracy correlates positively with AIVC-Judge scores, providing a complementary view of performance alongside open-response evaluation. Code and data demo are available at https://anonymous.4open.science/r/OmniVCBench.
摘要:人工智慧虛擬細胞(AIVCs)被設想為模擬細胞反應、解釋潛在機制並支持假說驅動發現的科學代理。然現有的AIVC基準主要在模擬層面運作,這促使我們對模型如何解釋實驗證據和制定生物假說進行補充評估。我們介紹OmniVCBench,一個以圖形為中心、可追溯來源的AIVC解釋組件基準。它包含6,077個經過策劃的單圖和多圖問題--答案對,這些問題來自科學文獻中的圖形和實驗背景。在布魯姆的分類法指導下,我們通過三個科學推理任務實現AIVC預測--解釋--發現議程的解釋層對應。我們進一步介紹AIVC-Judge,一個任務條件的MLLM作為評判框架,具有特定類別的、參考意識的評分標準,用於評估開放式回答。一個補充的模型衍生困難負樣本挖掘(MDHNM)策略將在模型推理過程中觀察到的合理錯誤轉換為多選題的干擾項,以降低評估成本。在評估的異質模型池中,多選題的準確率與AIVC-Judge的分數呈正相關,提供了與開放式回答評估相輔相成的性能視角。代碼和數據演示可在https://anonymous.4open.science/r/OmniVCBench獲得。
Multi-Site Real-World Performance of Commercial AI for Pulmonary and Incidental Pulmonary Embolism Detection
2609.37750v1 by Aawez Mansuri, Mohammadreza Chavoshi, Theodorus Dapamede, Wasif Bala, Beatrice Brown-Mulry, Rohan Isaac, Bardia Khosravi, Hanzhou Li, Frank Li, John T. Moon, Chad Robichaux, Dan I. G. Cohen-Addad, Ninad V. Salastekar, Janice Newsome, Judy W. Gichoya, Hari Trivedi
Pulmonary embolism (PE) is a leading cause of cardiovascular mortality, yet the real-world performance of FDA-cleared AI detection models remains incompletely characterized. We retrospectively evaluated two FDA-cleared AI algorithms from a single commercial platform (Aidoc Medical BriefCase), one for PE triage on dedicated CT pulmonary angiography (CTPA; n = 30,678) and one for incidental PE (iPE) detection on routine contrast-enhanced CTs (n = 37,191), across a 17-facility academic health system. Reference-standard labels were extracted from radiology reports using a validated LLM pipeline (97% accuracy, kappa = 0.94). The PE model achieved 86.8% sensitivity and 99.1% specificity, with sensitivity declining from 99.3% for saddle emboli to 72.9% for subsegmental PE, and from 89.7% for acute to 65.3% for non-acute PE. The iPE model achieved 73.5% sensitivity and 99.8% specificity. Both models demonstrated lower sensitivity than FDA-clearance benchmarks while exceeding cleared specificity, with diminishing performance for peripheral and non-acute emboli mirroring known human reader limitations and underscoring the need for standardized post-market surveillance of AI-enabled medical devices.
摘要:肺栓塞(PE)是心血管死亡的主要原因,但FDA批准的AI檢測模型在實際應用中的表現仍然未完全明確。我們回顧性地評估了來自單一商業平台(Aidoc Medical BriefCase)的兩個FDA批准的AI算法,一個用於專用CT肺動脈造影(CTPA;n = 30,678)的PE分流,另一個用於常規對比增強CT(n = 37,191)的偶然PE(iPE)檢測,涵蓋了17家學術醫療系統。參考標準標籤是通過經驗證的LLM管道(97%準確率,kappa = 0.94)從放射學報告中提取的。PE模型的敏感性達到86.8%,特異性為99.1%,其中敏感性從鞍狀栓塞的99.3%下降到亞段PE的72.9%,從急性PE的89.7%下降到非急性PE的65.3%。iPE模型的敏感性為73.5%,特異性為99.8%。兩個模型的敏感性均低於FDA批准的基準,但特異性超過批准標準,對於周邊和非急性栓塞的表現下降反映了已知的人類讀者限制,並強調了對AI驅動醫療設備標準化市場後監測的需求。
XU-RS: Explaining Credal Width in Random-Set Language Models
2609.37594v1 by David Achara, Maryam Sultana, Alexander D. Rast, Fabio Cuzzolin
Uncertainty estimates tell us how unsure a model is, but not why. Without knowing which parts of an input influences a model's uncertainty, we cannot tell whether that uncertainty score depends on input features that are relevant for the task. We study this problem in randomset classifiers built using pretrained language models. These classifiers assign probability to individual answers and to groups of answers, producing lower and upper probabilities for each answer; The difference between these probabilities, called credal width, is used to represent epistemic uncertainty about an answer arising from limited training data. We propose XU-RS, a framework that attributes an answer's credal width to the input tokens (words or word pieces) supplied to a language model. XU-RS uses Expected Gradients (a standard feature attribution method) to estimate how input tokens contribute to credal width. The proposed framework is evaluated on a MedQA dataset using SmolLM3-3B and Llama-2-7B models, demonstrating that setting the embedding of a token ranked highly by XU-RS to zero (zero-masking) causes larger changes in credal width than zero-masking randomly selected tokens. In addition, we show that normalisation can cause other answer groups to influence an answer's width, reveal how token attribution can mask numerical errors, and provide diagnostic checks to verify whether a token ranked highly by XU-RS meaningfully explains model uncertainty.
摘要:不確定性估計告訴我們模型有多不確定,但並不告訴我們原因。若不知道輸入的哪些部分影響模型的不確定性,我們無法判斷該不確定性分數是否依賴於與任務相關的輸入特徵。 我們在使用預訓練語言模型構建的隨機集分類器中研究這個問題。這些分類器為單個答案和答案組分配概率,為每個答案生成下限和上限概率;這些概率之間的差異稱為信念寬度,用於表示由於訓練數據有限而產生的對答案的認識不確定性。我們提出了XU-RS,一個將答案的信念寬度歸因於提供給語言模型的輸入標記(單詞或詞片)的框架。XU-RS使用期望梯度(標準特徵歸因方法)來估計輸入標記對信念寬度的貢獻。所提出的框架在使用SmolLM3-3B和Llama-2-7B模型的MedQA數據集上進行評估,證明將XU-RS排名較高的標記的嵌入設置為零(零掩蔽)會導致信念寬度的變化比隨機選擇的標記的零掩蔽更大。此外,我們展示了正規化可能導致其他答案組影響答案的寬度,揭示了標記歸因如何掩蓋數值錯誤,並提供診斷檢查以驗證XU-RS排名較高的標記是否有意義地解釋模型的不確定性。
Raw Imagery Impacting Your AI: Should You Care?
2609.38265v1 by Adrien Dorise, Marjorie Bellizzi, Stéphane May
Onboard AI is gaining interest for space applications such as vessel, wildfire, and cloud detection, where real-time processing can improve mission reactivity and reduce downlink needs. However, onboard models may operate on raw or minimally processed imagery rather than on restored ground products. This study evaluates how image degradation affects object detection by varying Signal-to-Noise Ratio (SNR), Modulation Transfer Function (MTF) at Nyquist, and Ground Sampling Distance (GSD). Controlled degradations are applied to Very High Resolution Maxar imagery, and three lightweight detectors, YOLOv5s, YOLOX-S, and NanoDet, are evaluated on the resulting operating points. The results show that the impact of image quality depends on the degradation mechanism, and that increasing degradation does not necessarily lead to a proportional decrease in vessel detection performance. GSD produces the most consistent performance shift, while MTF and SNR effects depend more on the model and resolution. Severe combinations of blur and noise produce the largest losses. These results provide task-level information that can support sensor, processing, and AI trade-offs for future onboard systems.
摘要:在太空應用中,機載人工智慧正受到關注,例如船隻、野火和雲層檢測,其中即時處理可以提高任務反應能力並減少下行鏈路需求。
然而,機載模型可能在原始或經過最小處理的影像上運作,而不是在恢復的地面產品上。
本研究評估影像退化如何影響物體檢測,通過改變信噪比(SNR)、奈奎斯特的調變傳遞函數(MTF)和地面取樣距離(GSD)。
對非常高解析度的Maxar影像施加控制退化,並在結果操作點上評估三個輕量級檢測器,YOLOv5s、YOLOX-S和NanoDet。
結果顯示,影像質量的影響取決於退化機制,且增加退化不一定會導致船隻檢測性能成比例下降。
GSD產生最一致的性能變化,而MTF和SNR的影響則更多地依賴於模型和解析度。
模糊和噪聲的嚴重組合會產生最大的損失。
這些結果提供了任務級別的信息,可以支持未來機載系統的傳感器、處理和人工智慧的權衡。
Selecting The Most Informative Tokens in Natural Language Autoencoders
2609.37040v1 by Federico Torrielli, Gianluca Barmina, Andrea Blasi Núñez, Amon Rapp, Luigi Di Caro, Peter Schneider-Kamp, Lukas Galke Poech
Natural language autoencoders translate a language model's internal activations into readable explanations. Explaining every token position is costly. Which positions should an auditor inspect to understand a potential threat? We study this question across $4.7$ million explanations on prompt injection and concealment. We compare signals from model computation with a ranker trained only on chat structure. Chat structure usually selects more relevant explanations than the computational signals, without requiring a model forward pass for position selection. On three of four datasets, explaining just $5\%$ of positions retains nearly all of the success rate from explaining every position, where success means obtaining an explanation about the threat. The benefit varies with the audit task. We also show that pretrained verbalizers recover words that models have learned to conceal through fine-tuning, without additional verbalizer training. These results identify where auditors can concentrate explanation generation and show that useful explanations can extend beyond the model a verbalizer was trained to describe.
摘要:自然語言自動編碼器將語言模型的內部激活轉換為可讀的解釋。解釋每個標記位置是昂貴的。審計員應該檢查哪些位置以理解潛在威脅?我們在 $4.7$ 百萬個關於提示注入和隱藏的解釋中研究了這個問題。我們將模型計算的信號與僅基於聊天結構訓練的排名器進行比較。聊天結構通常選擇比計算信號更相關的解釋,而不需要模型前向傳遞來選擇位置。在四個數據集中有三個中,僅解釋 $5\%$ 的位置幾乎保留了解釋每個位置的所有成功率,其中成功意味著獲得有關威脅的解釋。這一好處隨著審計任務而異。我們還顯示,預訓練的詞語化器能夠恢復模型通過微調學會隱藏的單詞,而不需要額外的詞語化器訓練。這些結果確定了審計員可以集中解釋生成的地方,並顯示有用的解釋可以超越詞語化器所訓練描述的模型。
Harness Evolution as Learning: Approximation, Generalization, and Optimization Limits of Self-Improving Personal Agents
2609.36892v1 by Zeyu Gan, Zixuan Gong, Yong Liu
As the capabilities of large language models (LLMs) continue to advance, increasing attention is turning to how to translate their abilities into useful behavior. Personal agents bring this question into everyday settings, where models are expected to serve individual users and continually adapt to their preferences. With the underlying model held fixed, such adaptation relies on harness engineering: designing and evolving the surrounding layer that manages context, memory, tools, and execution. Despite rapid progress, the factors governing effective harness evolution remain insufficiently understood. To narrow this gap, we investigate three central questions concerning harness architecture, harness scale, and self-evolution algorithms through complementary empirical and theoretical analyses. Empirically, we introduce a preference-oriented benchmark and systematically characterize the capabilities and limitations of personal agents associated with these three dimensions. Theoretically, we formulate harness evolution as a learning problem and explain these phenomena through approximation, generalization, and optimization errors. Analyses of reachable policies, capacity under finite interaction evidence, and biased update dynamics provide theoretical accounts of the observed phenomena. Together, these results offer a unified perspective on the limits of personalization through harness evolution and inform future harness design.
摘要:隨著大型語言模型(LLMs)能力的持續進步,越來越多的關注轉向如何將其能力轉化為有用的行為。個人代理將這個問題帶入日常環境,在這裡模型被期望為個別用戶服務並不斷適應他們的偏好。在基礎模型保持固定的情況下,這種適應依賴於飼養工程:設計和發展管理上下文、記憶、工具和執行的周邊層。儘管進展迅速,但影響有效飼養演變的因素仍然理解不足。為了縮小這一差距,我們通過互補的實證和理論分析,研究有關飼養架構、飼養規模和自我演變算法的三個核心問題。在實證方面,我們引入了一個以偏好為導向的基準,並系統性地描述與這三個維度相關的個人代理的能力和局限性。在理論方面,我們將飼養演變公式化為一個學習問題,並通過近似、泛化和優化誤差解釋這些現象。可達政策、有限互動證據下的容量和偏見更新動態的分析提供了對觀察到的現象的理論解釋。總體而言,這些結果提供了對通過飼養演變實現個性化限制的統一視角,並為未來的飼養設計提供了指導。
Predictive Credit: Measuring What Scientific Explanations Add to Experimental Forecasts
2610.00314v1 by Jingjie Ning, Xueqi Li, Yibo Kong, Dongting Li
Research agents explain planned experiments. We measure predictive credit with paired forecasts sharing an intervention, forecaster, and outcome while varying description, matched explanation, and donor context. Five checks track commitment, delivery, predictive gain, alignment, and known-signal uptake. Across 336 prospective states in controlled learning, 12 Tox21 endpoints, and 24 OpenML tasks, v5's frozen credit decision was inconclusive. Tox21's preregistered ROC AUC interval-score harm test was unmet ($D-M=-.0026$, 95 percent interval [$-.0174$, .0104]); OpenML's joint formation, point-equivalence, and repeatability rule was unmet. Matched point-accuracy gains over description remained unconfirmed, and Tox21/OpenML seed-donor intervals spanned zero. Under requested DeepSeek V4 Pro, matched and donor cards reduced secondary Tox21 drift by 64.5 and 59.1 percent. A DeepSeek V4 Flash replay raised matched point MAE from .01823 to .02020 and missed matched-donor interval-score equivalence. OpenML full-card assignment widened nominal 80 percent intervals by 21 percent, with 49.3 percent coverage versus 51.4 percent for description and content in 66/144 cards. Direct-text Flash delivered all 144 notes without detectable matched point-accuracy gain. A researcher-authored mechanism positive control lowered point MAE by 2.60 percentage points versus description. The protocol measures predictive credit for research-agent benchmarks and scientific forecasting; natural-explanation credit remained unconfirmed at the tested donor resolutions.
摘要:研究代理人解釋計劃的實驗。
我們通過共享干預、預測者和結果的配對預測來測量預測信用,同時變化描述、匹配解釋和捐贈者背景。
五個檢查跟踪承諾、交付、預測增益、一致性和已知信號的吸收。
在336個受控學習的前瞻性狀態、12個Tox21端點和24個OpenML任務中,v5的凍結信用決策結果不確定。
Tox21的預註冊ROC AUC區間分數損害測試未達標($D-M=-.0026$,95百分位區間[$-.0174$,.0104]);OpenML的聯合形成、點等價和重複性規則未達標。
對於描述的匹配點準確度增益仍未得到確認,Tox21/OpenML種子-捐贈者區間跨越零。
在請求的DeepSeek V4 Pro下,匹配和捐贈卡減少了次級Tox21漂移64.5和59.1百分比。
DeepSeek V4 Flash重播將匹配點的MAE從.01823提高到.02020,並錯過了匹配-捐贈者區間分數等價。
OpenML全卡分配將名義80百分比區間擴大了21百分比,對於66/144張卡片,覆蓋率為49.3百分比,而描述和內容的覆蓋率為51.4百分比。
直接文本Flash交付了所有144條備註,沒有檢測到匹配點準確度的增益。
一個研究者撰寫的機制正控制降低了點MAE,相對於描述降低了2.60個百分點。
該協議測量研究代理基準和科學預測的預測信用;在測試的捐贈者解析度下,自然解釋信用仍未得到確認。
Engineering Simplicity: Simple Mechanism Interfaces Steer LLM Agents
2609.36365v1 by Kehang Zhu, Anand Shah, David Parkes
Can interaction formats and textual scaffolds help large language model (LLM) agents make better decisions, and do better decisions come with better explanations? We study these questions in auctions and matching, multi-agent environments with explicit rules and known optimal strategies. These settings let us vary how a decision problem is presented while retaining a benchmark for evaluating behavior. Drawing on human-motivated theories of simplicity, we compare interfaces that elicit a complete bid or ranking with sequential interfaces that make safe choices easier to identify. We then hold the interaction format fixed and vary reasoning scaffolds and rule descriptions. Across four model families, the ascending auction interface substantially reduces bid deviations. The matching comparison also shows why sequential responses require different error accounting from complete rankings. Laying out payoff contingencies and explaining why truth-telling is safe also improve choices, whereas prompts to plan through matching rounds or form beliefs about opponents worsen play overall. In auctions, these behavioral gains are not accompanied by corresponding improvements in measured verbal indicators of strategic understanding in the agents' short stated plans. Other prompts change those indicators without improving bids. Our findings suggest that human-motivated theories of simplicity can inform the design of decision environments for artificial agents. They also show why scaffolds should be evaluated through realized choices as well as explanations: improvements in one need not appear in the other.
摘要:互動格式和文本支架能否幫助大型語言模型 (LLM) 代理做出更好的決策,而更好的決策是否伴隨著更好的解釋?我們在拍賣和匹配的多代理環境中研究這些問題,這些環境具有明確的規則和已知的最佳策略。這些設置使我們能夠變化決策問題的呈現方式,同時保留評估行為的基準。基於人類動機的簡單性理論,我們比較了引發完整出價或排名的介面與使安全選擇更容易識別的序列介面。我們然後固定互動格式,變化推理支架和規則描述。在四個模型家族中,升序拍賣介面顯著減少了出價偏差。匹配比較也顯示為什麼序列反應需要不同的錯誤計算,與完整排名相比。列出支付條件並解釋為什麼誠實報告是安全的也改善了選擇,而計劃通過匹配輪次或形成對對手的信念的提示則整體上惡化了遊戲。在拍賣中,這些行為上的增益並未伴隨著代理短期陳述計劃中戰略理解的口頭指標的相應改善。其他提示改變了這些指標,但並未改善出價。我們的發現表明,人類動機的簡單性理論可以為人工代理的決策環境設計提供指導。它們還顯示為什麼支架應該通過實現的選擇以及解釋來評估:一方面的改善不一定會在另一方面出現。
Explainability from Training with Applications to TCR-Epitope Prediction
2609.36354v1 by Jiarui Li, Zixiang Yin, Samuel Landry, Zhengming Ding, Ramgopal Mettu
Deep learning models have achieved strong performance in artificial intelligence for science, yet their black-box nature limits our understanding of how they learn scientific tasks. Existing methods for interpretability provide limited insight into how models organize evidence and evolve during learning. We introduce explainability from training (EFT), a model-agnostic paradigm that traces model interpretation during training to explain why models rely on specific features and how they organize these features as predictive evidence. We apply EFT to four state-of-the-art T cell receptor (TCR)-epitope prediction models, TCR-SRIM, TULIP, MixTCRpred, and NetTCR-2.2, spanning post-hoc and interpret-by-design approaches as well as transformers and CNNs. To investigate how structural information affects model explanations, we introduce a benchmark, TCR-XAI2, containing 388 unique experimentally resolved TCR-epitope structures, complemented by structures predicted using AlphaFold3, Boltz-2, TCRModel2, tFold-TCR, and OpenFold3. Using EFT with TCR-XAI2, we demonstrate that (1) CNN and transformer models exhibit distinct learning trajectories; (2) TCR $α$ and $β$ evidence can conflict during learning, limiting the benefits of jointly modeling both chains, while MHC information mitigates this; and (3) real versus predicted structural data for TCR-epitope prediction exhibits distinct TCR and peptide feature preferences as well as differing trajectories of model certainty.
摘要:深度學習模型在科學人工智慧中取得了強大的表現,但其黑箱特性限制了我們對它們如何學習科學任務的理解。現有的可解釋性方法對模型如何組織證據和在學習過程中如何演變提供的見解有限。我們介紹了訓練中的可解釋性(EFT),這是一種與模型無關的範式,追蹤模型在訓練過程中的解釋,以解釋為什麼模型依賴於特定特徵以及它們如何將這些特徵組織為預測證據。我們將EFT應用於四個最先進的T細胞受體(TCR)-表位預測模型,TCR-SRIM、TULIP、MixTCRpred和NetTCR-2.2,涵蓋了事後解釋和設計解釋的方法,以及Transformer和卷積神經網絡。為了研究結構信息如何影響模型解釋,我們引入了一個基準,TCR-XAI2,包含388個獨特的實驗解析TCR-表位結構,並補充了使用AlphaFold3、Boltz-2、TCRModel2、tFold-TCR和OpenFold3預測的結構。使用EFT和TCR-XAI2,我們證明了(1)CNN和Transformer模型顯示出不同的學習軌跡;(2)TCR $α$ 和 $β$ 證據在學習過程中可能會衝突,限制了同時建模這兩條鏈的好處,而MHC信息則減輕了這一點;(3)TCR-表位預測的實際結構數據與預測結構數據顯示出不同的TCR和肽特徵偏好,以及不同的模型確定性軌跡。
ThuRunel: Dynamic Decoupling for Structured Advisory Dialogue
2609.36340v1 by Yuyan Chen
High-stakes advisory domains such as medical aesthetics, legal consultation, and educational planning exhibit a two-phase structure. The early phase requires empathetic elicitation and emotional support, and the late phase requires authoritative specialist judgment. Neither fully automated agents nor human junior consultants adequately address this structure at scale. We formalize the core design challenge as dynamic decoupling, asking how an AI advisory agent should decide what to ask, when to stop, what to resolve autonomously, and what to forward to the specialist. We present ThuRunel, an advisory agent combining a finite-state belief management framework, a chain-of-thought teacher synthesis protocol, and learned generation adapters. Against eleven baselines, ThuRunel achieves consistent improvements in elicitation completeness and specialist brief quality. ThuRunel is publicly deployed as a bilingual web application in which the same decoupling decisions operate from the client's side, grounded in a curated knowledge base that cites its sources in every answer.
摘要:高風險的諮詢領域如醫療美學、法律諮詢和教育規劃展示出雙階段結構。
早期階段需要同理心的引導和情感支持,而後期階段則需要權威專家的判斷。
無論是完全自動化的代理還是人類初級顧問,都無法在規模上充分應對這一結構。
我們將核心設計挑戰形式化為動態解耦,詢問AI諮詢代理應如何決定詢問什麼、何時停止、什麼可以自主解決以及什麼需要轉交給專家。
我們提出了ThuRunel,一個結合有限狀態信念管理框架、思維鏈教師綜合協議和學習生成適配器的諮詢代理。
在十一個基準測試中,ThuRunel在引導完整性和專家簡報質量上實現了一致的改進。
ThuRunel作為一個雙語網絡應用程序公開部署,其中相同的解耦決策從客戶端運作,基於一個策劃的知識庫,並在每個答案中引用其來源。
FigAct: Turning Scientific Figures into Active Canvases for Explanation
2609.36190v1 by Shishi Xiao, Zichao Wang, Alexa Siu, David H. Laidlaw, Jennifer Healey
Scientific figures are designed to communicate information visually, yet MLLMs typically explain them by translating their visual content back into text. This requires readers to manually map the resulting explanations back to the figure. Inspired by how people present visual information, we introduce FigAct, a framework that transforms static scientific figures into question-conditioned visual presentations by acting directly on their existing graphical elements. Like a human presenter, FigAct generates a sequence of short narrations, grounds each narration in the corresponding visual evidence, and applies visual actions to guide the viewer's attention. We develop a hierarchical search strategy for efficient element localization, reducing token usage by approximately 40$\times$. We further train FigAct-8B using three task-specific rewards for grounding accuracy, search efficiency, and rendering quality. We further build a human-verified benchmark from figures in real-world scientific papers to evaluate the ability of MLLMs to generate grounded visual explanations. Our results demonstrate the effectiveness of FigAct and show that treating scientific figures as presentation canvases makes explanations clearer and easier to follow.
摘要:科學圖形旨在以視覺方式傳達信息,但 MLLMs 通常通過將其視覺內容轉換回文本來解釋它們。這要求讀者手動將生成的解釋映射回圖形。受到人們如何呈現視覺信息的啟發,我們介紹了 FigAct,一個通過直接作用於現有圖形元素將靜態科學圖形轉換為基於問題的視覺展示的框架。像人類演講者一樣,FigAct 生成一系列簡短的敘述,將每個敘述與相應的視覺證據相結合,並應用視覺動作來引導觀眾的注意力。我們開發了一種層次搜索策略,以提高元素定位的效率,將標記的使用減少約 40$\times$。我們進一步使用三個特定任務的獎勵來訓練 FigAct-8B,以提高基準準確性、搜索效率和渲染質量。我們還從現實世界的科學論文中的圖形建立了一個經過人工驗證的基準,以評估 MLLMs 生成有根據的視覺解釋的能力。我們的結果證明了 FigAct 的有效性,並顯示將科學圖形視為展示畫布使解釋更清晰、更易於理解。
An Exact Generate - Transform Decomposition of Small-LLM Team Scaling Across Orchestration Architectures
2609.36104v1 by Blaz Bertalanic, Carolina Fortuna
Replacing one LLM agent with a collaborating team can raise accuracy, but whether scaling the team helps, and which architecture to scale, is unclear. Sweeping eight agent orchestration architectures across five instruction-tuned 7-9B models, five short-answer benchmarks, and an executable-code benchmark up to 30 calls, we find that the returns to team scaling are sharply task-dependent: from three to thirty calls accuracy rises by up to 17 points on the two arithmetic word-problem benchmarks (GSM8K, GSMHard) but by at most four on ARC, GPQA, and MMLU, for every architecture, a split the usual task-averaged number conceals. Proposer-Critic captures the arithmetic gains, scaling steepest and, in aggregate, surpassing every other architecture at the largest budget (item-clustered intervals exclude zero), though it ranks among the weakest elsewhere, and no architecture wins across tasks. We explain these trajectories with an exact generate-transform decomposition. Partitioning any workflow into proposal coverage and a downstream transform, any accuracy change splits exactly into an extensive coverage dividend and an intensive transformation change. The decomposition diagnoses each task: arithmetic offers coverage headroom that a critic-guided transform converts, whereas the multiple-choice benchmarks either saturate in coverage or fail to convert it, and on open-ended code generative recovery nearly vanishes so accuracy tracks coverage. At equal call budgets token cost still varies 2.1x. Extra calls therefore create candidate opportunity that only some architectures, on some tasks, convert. Team scaling is a task- and architecture-specific bet, not a uniform lever.
摘要:替換一個 LLM 代理為一個協作團隊可以提高準確性,但擴大團隊是否有幫助,以及擴大哪種架構仍不清楚。對於五個經過指令調整的 7-9B 模型、五個短答案基準以及一個可執行代碼基準進行八種代理協調架構的廣泛測試,最多進行 30 次調用,我們發現團隊擴大的回報明顯依賴於任務:在兩個算術文字問題基準(GSM8K、GSMHard)上,從三次到三十次調用的準確性提高了最多 17 個點,但在 ARC、GPQA 和 MMLU 上最多只提高四個點,這是每種架構的情況,通常的任務平均數隱藏了這一點。提議者-評價者捕捉了算術增益,擴大最陡峭,並且在總體上超越了其他所有架構在最大預算下(項目聚類區間不包括零),儘管在其他地方排名較弱,且沒有任何架構在所有任務中獲勝。
我們用精確的生成-轉換分解來解釋這些軌跡。將任何工作流程劃分為提議覆蓋和下游轉換,任何準確性變化都精確地分為廣泛的覆蓋紅利和密集的轉換變化。這一分解診斷每個任務:算術提供了覆蓋的頭部空間,評價者引導的轉換將其轉換,而多選基準要麼在覆蓋上飽和,要麼未能轉換,並且在開放式代碼生成恢復中幾乎消失,因此準確性跟蹤覆蓋。在相等的調用預算下,令牌成本仍然變化 2.1 倍。因此,額外的調用創造了候選機會,只有一些架構在某些任務上能夠轉換。團隊擴大是一個特定於任務和架構的賭注,而不是一個統一的槓桿。
One Geometry, Different Outcomes: Readout-Dependent Effects of the Modality Gap in Vision-Language Models
2609.36101v1 by Aditya Sharma, Divya Saxena
Contrastive vision-language models learn shared embedding spaces by aligning matched image-text pairs, yet their representations remain separated by a modality gap. Prior work reports divergent effects of modifying this gap: reducing it can improve zero-shot classification and cross-modal alignment, whereas removing gap-related structure can degrade image-text retrieval. In this paper, we provide a unified geometric explanation for these task-dependent effects. Across CLIP and SigLIP encoders, we find that a single dominant direction captures 94.4-99.9% of the squared norm of the image-text mean separation, revealing that the mean-separation component is approximately rank-one. A decomposition of the similarity score then identifies three task-specific roles. In zero-shot classification, query-side fixed gap-offset subtraction is exactly equivalent to an additive class bias. In standard cross-modal retrieval, projecting out the gap direction and renormalising residuals discards candidate-specific norm information, inducing a multiplicative ranking distortion; a geometry-derived exponent tracks the grid-search optimum (Spearman rho = 0.93) and restores performance in some settings, although the gains transfer unevenly. In mixed-modal retrieval, the gap direction sorts candidates by modality; its removal can improve cross-modal ranking, unlike random or non-gap controls. Residual semantic structure after removal defines the limits of the rank-one account. Together, these results explain why gap modification can improve, degrade, or restore performance across downstream settings. By clarifying when and why gap modification changes model behavior, this account provides a principled basis for selecting gap interventions in similarity-based vision-language systems across evaluated downstream tasks.
摘要:對比視覺-語言模型通過對齊匹配的圖像-文本對來學習共享的嵌入空間,但它們的表示仍然受到模態差距的分隔。先前的研究報告了修改這一差距的不同效果:減少它可以改善零樣本分類和跨模態對齊,而去除與差距相關的結構則可能降低圖像-文本檢索的效果。在本文中,我們提供了一個統一的幾何解釋來說明這些依賴於任務的效果。在CLIP和SigLIP編碼器中,我們發現一個主導方向捕捉了94.4-99.9%的圖像-文本平均分離的平方範數,揭示了平均分離成分大約是秩一的。然後,對相似度分數的分解確定了三個特定於任務的角色。在零樣本分類中,查詢端固定的差距偏移減法與加性類別偏差完全等價。在標準的跨模態檢索中,投影出差距方向並重新標準化殘差會丟棄候選特定的範數信息,導致乘法排名失真;一個幾何推導的指數跟踪網格搜索最佳(Spearman rho = 0.93),並在某些設置中恢復性能,儘管增益轉移不均勻。在混合模態檢索中,差距方向按模態對候選進行排序;其去除可以改善跨模態排名,這與隨機或非差距控制不同。去除後的殘餘語義結構定義了秩一解釋的極限。總體而言,這些結果解釋了為什麼差距修改可以改善、降低或恢復下游設置中的性能。通過澄清何時以及為什麼差距修改會改變模型行為,這一解釋為在評估的下游任務中選擇基於相似性的視覺-語言系統中的差距干預提供了原則性的基礎。
Shockingly Simple Self-retrospection Improves Agentic Models Without RL
2609.35741v1 by Jonathan Light, Christopher Zhang Cui, Jeonghye Kim, Roger Creus Castanyer, Emiliano Penaloza, Zhengyan Shi, Alessandro Sordoni, Marc-Alexandre Côté, Xingdi Yuan, Minseon Kim
People learn not only by repeating successful actions, but also by recounting and explaining their experiences, revising their understanding to guide future behavior. Can a language-model agent improve its future actions by training only on explanations of its own experience? We investigate this question by studying Retrospection-Only Fine-Tuning (ROFT), a minimal online procedure designed to isolate the effect of explanation-only training on subsequent behavior. The agent attempts a task, observes available feedback, generates a retrospective explanation, and is fine-tuned with a next-token prediction loss on the explanation tokens alone. The procedure uses neither an external teacher nor a reward-based policy update. In software-engineering experiments with Qwen3.5-4B, ROFT is trained on problems with mixed successful and unsuccessful base-model attempts. On held-out SWE-bench Verified and Pro, it reaches 49.2% and 26.8% solve rates after 20 updates without using a verifier, compared with GRPO's 48.0% and 25.3% after 40 updates in the evaluated runs, and makes faster early progress in training time and sampled attempts. It also learns to solve individual tasks on which all 64 sampled base-model attempts failed, showing that learning can begin without any initially successful trajectories. Behavioral analyses find that ROFT indirectly assigns credit to actions, encouraging good actions and discouraging incorrect ones. Moreover, prompting retrospections to emphasize more direct solutions yields shorter subsequent attempts even without an explicit length penalty. Together, these findings show that learning to explain can also improve learning to do, establishing self-generated retrospections as useful training targets and motivating further study of explanation-to-action transfer.
摘要:人們不僅通過重複成功的行動來學習,還通過敘述和解釋他們的經驗,修正他們的理解以指導未來的行為。語言模型代理能否僅通過對自身經驗的解釋進行訓練來改善其未來的行動?我們通過研究回顧性僅微調(Retrospection-Only Fine-Tuning, ROFT)來探討這個問題,這是一種旨在孤立解釋性訓練對後續行為影響的最小在線程序。代理嘗試一個任務,觀察可用的反饋,生成一個回顧性解釋,並僅對解釋標記進行下一標記預測損失的微調。該程序既不使用外部教師,也不進行基於獎勵的策略更新。在使用 Qwen3.5-4B 的軟體工程實驗中,ROFT 在成功和不成功的基模型嘗試混合的問題上進行訓練。在保留的 SWE-bench Verified 和 Pro 上,它在 20 次更新後達到 49.2% 和 26.8% 的解決率,而不使用驗證器,與 GRPO 在評估運行中 40 次更新後的 48.0% 和 25.3% 相比,並在訓練時間和抽樣嘗試中更快地取得早期進展。它還學會了解決所有 64 次抽樣基模型嘗試失敗的個別任務,顯示學習可以在沒有任何最初成功的軌跡的情況下開始。行為分析發現,ROFT 間接地將功勞分配給行動,鼓勵良好的行動並抑制不正確的行動。此外,促使回顧以強調更直接的解決方案,即使在沒有明確的長度懲罰的情況下,也會產生更短的後續嘗試。這些發現表明,學習解釋也可以改善學習行動,確立自我生成的回顧作為有用的訓練目標,並激勵進一步研究解釋到行動的轉移。
Rethinking Circuit Evaluation: Do Circuits Explain Model Errors?
2609.35686v1 by Li Zhang, Chuqin Geng, Mark Zhang, Chen Yang, Luke Zhang, Haolin Ye, Xujie Si
Mechanistic interpretability (MI) aims to explain a model's behaviour through analyzing its internal computations; circuit-based explanations aim to isolate these computations with compact subnetworks validated by ablating the rest of the model. We show that circuits validated this way may fail to recover the underlying mechanism of the model's behaviour by closely reproducing its successful decisions while failing to account for most of its errors. Such explanations should account for the model's particular errors as well as its successes. We evaluate this requirement by measuring exact answer agreement separately on model successes and failures, across circuit sizes and ablation settings, on IOI, Docstring, and six model-task settings from the Mechanistic Interpretability Benchmark. We discover that many tested circuits closely replicate correct behaviour while missing most of the model's errors. On indirect object identification (IOI) for GPT-2 small, under mean ablation, the manual circuit and tested automated circuits, including one trained against the model's full output distribution, agree with the model on 97.3-99.5% of prompts it answers correctly but only 11.4-41.7% of errors. An IOI case study shows that lost errors are recoverable by restoring omitted attention-heads which raise error reproduction from 14.2% to 75.1% on a separate held-out set with 0.41 percentage point decrease on correct agreement, exceeding matched random extensions and scalar-biased control. Intervention traces show how omitted computations produce specific wrong answers for a reproducible subset of errors. In all, these findings show circuits can preserve task success without adequately explaining model's failures, and support exact error reproduction as a necessary, but not sufficient, test of circuit-based explanations of model behaviour.
摘要:機械解釋性(MI)旨在通過分析模型的內部計算來解釋模型的行為;基於電路的解釋旨在通過消除模型的其餘部分來隔離這些計算,並使用緊湊的子網絡進行驗證。我們顯示,這種方式驗證的電路可能無法恢復模型行為的基本機制,因為它們雖然能夠緊密再現成功的決策,但卻未能考慮到大多數錯誤。這樣的解釋應該考慮到模型的特定錯誤以及它的成功。我們通過在模型的成功和失敗之間,分別測量準確答案的一致性,來評估這一要求,並在不同的電路大小和消融設置下,使用 IOI、Docstring 和來自機械解釋基準的六個模型任務設置。我們發現,許多測試過的電路能夠緊密複製正確的行為,但卻錯過了模型的大多數錯誤。在 GPT-2 small 的間接物體識別(IOI)中,在平均消融下,手動電路和測試過的自動電路,包括一個針對模型的完整輸出分佈進行訓練的電路,對於模型正確回答的提示達到 97.3-99.5% 的一致性,但對於錯誤的僅有 11.4-41.7%。一個 IOI 案例研究顯示,通過恢復省略的注意力頭來恢復遺失的錯誤,將錯誤再現率從 14.2% 提升至 75.1%,在一個獨立的保留集上,正確一致性下降了 0.41 個百分點,超過了匹配的隨機擴展和標量偏置控制。干預痕跡顯示,省略的計算如何為可重現的錯誤子集產生特定的錯誤答案。總的來說,這些發現表明,電路可以保持任務的成功,而未能充分解釋模型的失敗,並支持準確的錯誤再現作為基於電路的模型行為解釋的必要但不充分的測試。
Signatures of semantic search in the activations of large language models
2609.35599v2 by Luke Leckie, Peter M. Todd, Jacob G. Foster
When recalling lists of concepts (e.g., animals) during the semantic fluency task (SFT), both humans and large language models (LLMs) organise their output into clusters of related items (e.g., sea animals) that are punctuated by strategic switches between clusters. In humans, this pattern can be explained by a semantic foraging process, whereby distinct neural and behavioural signatures accompany within-cluster production ("exploit") and between-cluster switching ("explore"). Whether LLMs likewise represent these two search regimes within their internal states is unknown. Here, we apply a range of mechanistic interpretability techniques to provide evidence for this. In Study 1, we use the Jacobian lens (J-lens), which maps intermediate-layer residual-stream representations to token-level activations, to show that concept-level activations predict switching. First, we find that switching coincides with low next-token activations. Moreover, the probability of switching rises as the set of strongest J-lens activations (the J-space) becomes depleted of items from the category currently being produced, analogous to explore-exploit decision-making during patch foraging. We then show that middle-layer J-lens activations of abstract category-related labels (e.g., "water") increase in anticipation of switching into that category. We confirm these representations to causally influence switching by deriving steering vectors that target category switching. In Study 2, we identify generic residual stream directions that are activated during and in anticipation of switching. By steering activations along these directions, we bias increased or decreased rates of switching. Our study extends the semantic foraging framework to artificial intelligences and provides evidence that LLMs maintain distinct representational signatures for exploration and exploitation as they verbalise conceptual information.
摘要:在語意流暢性任務(SFT)中回想概念列表(例如,動物)時,人類和大型語言模型(LLMs)都將其輸出組織成相關項目的集群(例如,海洋動物),並在集群之間進行戰略性切換。在人類中,這種模式可以通過語意覓食過程來解釋,其中不同的神經和行為特徵伴隨著集群內的產出(「利用」)和集群之間的切換(「探索」)。LLMs是否同樣在其內部狀態中表示這兩種搜索模式尚不清楚。在這裡,我們應用一系列機械可解釋性技術來提供證據。在研究 1 中,我們使用雅可比透鏡(J-lens),該透鏡將中間層的殘差流表示映射到標記級別的激活,來顯示概念級別的激活預測切換。首先,我們發現切換與低的下一標記激活相吻合。此外,隨著最強 J-lens 激活集(J-space)中的當前產出類別項目的耗盡,切換的概率上升,類似於在斑塊覓食過程中的探索-利用決策。我們接著顯示,抽象類別相關標籤(例如,「水」)的中層 J-lens 激活在預期切換到該類別之前會增加。我們通過推導針對類別切換的引導向量來確認這些表示對切換的因果影響。在研究 2 中,我們識別在切換過程中及其預期期間被激活的通用殘差流方向。通過沿這些方向引導激活,我們偏向於增加或減少切換的頻率。我們的研究將語意覓食框架擴展到人工智慧,並提供證據表明 LLMs 在口頭表達概念信息時保持探索和利用的不同表徵特徵。
A decision-support system applied to Law: Reasoning and explainability of the decision
2609.35370v1 by Jeremy Bouche-Pillon, Pascale Zarat{é}, Yannick Chevalier, Nathalie Aussenac-Gilles
The emergence of the digital transition brought an increasing need to control the processing of digital information, including in Law Enforcement Agencies (LEAs). At the EU level, in recent years, many regulations have emerged to control data processing and exchange. Texts other than the GDPR, such as the ''Law Enforcement Directive (LED)'', appeared to regulate specifically how Law Enforcement Agencies (LEAs) could process data. A formal representation of these regulations can be part of decision systems that support LEAs in processing data in compliance with the regulations. Although many new formalisms have emerged to represent legal norms and rules, few are provided with a reasoning mechanism. Furthermore, systems used in decision-making processes in critical contexts such as medical diagnoses or legal decisions cannot be fully automated, and the explainability of their results is essential to ensure user confidence in decisions. This explainability aspect, while crucial, is lacking in most modern approaches that rely on machine learning. This paper describes a framework to operate formal rules from regulations, by focusing on explainability of the decision. After describing the general architecture of the proposed decision support framework, the paper showcases how symbolic AI and the SPARQL query language can support legal reasoning. It then describes an algorithm to generate a justification for the reasoning results, and outlines the procedure to be followed when the reasoning does not lead to a satisfactory conclusion. We notably focus on a method based on decision trees to determine what additional information to request from the user.
摘要:數位轉型的出現帶來了對數位資訊處理的日益需求,包括在執法機構(LEAs)中。在歐盟層面上,近年來出現了許多規範來控制數據處理和交換。除了GDPR之外,還出現了如“執法指令(LED)”等文本,專門規範執法機構(LEAs)如何處理數據。這些規範的正式表述可以成為支持執法機構在遵守規範的情況下處理數據的決策系統的一部分。儘管許多新的形式主義已經出現以表達法律規範和規則,但很少有配備推理機制的形式。此外,用於醫療診斷或法律決策等關鍵情境的決策過程中使用的系統不能完全自動化,其結果的可解釋性對於確保用戶對決策的信心至關重要。這一可解釋性方面雖然至關重要,但在大多數依賴機器學習的現代方法中卻缺乏。本文描述了一個運作規範形式規則的框架,重點在於決策的可解釋性。在描述所提議的決策支持框架的一般架構後,本文展示了符號人工智慧和SPARQL查詢語言如何支持法律推理。接著描述了一種生成推理結果的理由的算法,並概述了當推理未能導致令人滿意的結論時應遵循的程序。我們特別關注一種基於決策樹的方法,以確定需要向用戶請求的額外信息。
Jev thinks "I don't know'', but doesn't say it: Introducing Sys1Cal-v1 Dataset for Probability Calibration
2609.35342v1 by Riccardo Porcedda
The appearance of Jev marked the era of System One Models, foundation models that return structured decisions with probability distributions rather than text. Aside from low cost and great speed, Jev's central promise is that these probabilities are calibrated: such claim is not backed by any public test and available external benchmarks evaluate confidence calibration, not whether every returned option probability has the right numerical meaning. To tackle this issue, we introduce Sys1Cal-v1, a dataset of True/False questions about a proposition $A$ for which the exact probability $P(A)$ is known by construction. Each item is queried through the three Jev primitives - Noul, Choice and Score - and evaluated by total variation distance from the ground-truth distribution, which can be used to estimate a soft accuracy of System One Models. We showcase the utility of Sys1Cal-v1 as a benchmark dataset by evaluating Jev and SemIf, an open-source Choice-style baseline. In this work, however, we focus even more deeply on Jev, by studying the calibration of its Score and Choice answers. In particular, we discover a peculiar behaviour that can be explained by assuming that Jev suppresses a third truth value, going beyond True and False. In other words, in \texttt{Choice} answers, $P(A)$ and $P(\neg A)$ are presented as if $P(A)+P(\neg A)=1$, while a term $P(U)\neq0$ is missing in the sum. Recovering $P(U)$ leads to an improvement of median soft accuracy in \texttt{Choice} answers from $0.771$ to $0.978$, suggesting that, even in binary decisions, Jev wants to answer with a third option:``I don't know''.
摘要:Jev的出現標誌著系統一模型的時代,這些基礎模型返回結構化的決策,並以概率分佈而非文本形式呈現。除了低成本和高速度外,Jev的核心承諾是這些概率是經過校準的:這一說法並沒有任何公開測試的支持,且可用的外部基準評估的是置信度校準,而不是每個返回的選項概率是否具有正確的數值意義。為了應對這一問題,我們引入了Sys1Cal-v1,一個關於命題$A$的真/假問題數據集,該命題的確切概率$P(A)$是通過構造已知的。每個項目通過三個Jev原語 - Noul、Choice和Score進行查詢,並通過與真實分佈的總變異距離進行評估,這可以用來估計系統一模型的軟準確性。
我們展示了Sys1Cal-v1作為基準數據集的實用性,通過評估Jev和SemIf,一個開源的選擇風格基準。然而,在這項工作中,我們更深入地專注於Jev,研究其Score和Choice答案的校準。特別是,我們發現了一種特殊的行為,可以通過假設Jev抑制第三個真值來解釋,這超越了真和假。換句話說,在\texttt{Choice}答案中,$P(A)$和$P(\neg A)$被呈現得好像$P(A)+P(\neg A)=1$,而在總和中缺少了一項$P(U)\neq0$。恢復$P(U)$使得\texttt{Choice}答案的中位數軟準確性從$0.771$提高到$0.978$,這表明,即使在二元決策中,Jev也希望以第三個選項回答:“我不知道”。
The Argument and the Letterhead: Source-Position Coherence in AI Evaluation
2609.35286v1 by Michele Loi
An argument can be surprising coming from a particular speaker without being a bad argument. Do AI evaluators keep these judgments apart? Two preregistered descriptive studies and a later Jev supplement collected 2,976 usable evaluations of six fixed texts about US AI policy, Germany's debt brake and Swiss nuclear energy. Each text was presented under several source attributions. The key comparison asks whether the gap between two sources changes when the argument changes. On Sol, for example, a national-security argument received mean ratings of 0.359 under CODEPINK and 0.639 under College Republicans; a civil-rights argument received 0.742 and 0.721. A constant preference for one source cannot explain that pattern. Related interactions appeared across topics and recent model configurations, including those with reasoning enabled, while several comparisons yielded small effects. The later European Jev supplement yielded five interactions below the adopted absolute reference of 0.05; its distinct rubric and interrupted collection limit comparison with the chat systems. Some written evaluations explicitly invoked a mismatch between a source and its attributed position. Taken together, the numerical and verbal evidence supports source-position coherence as a plausible explanation, alongside competing accounts involving credibility, authenticity and interpretation of the task. The paper develops this inference through controlled comparisons, reports conditional post hoc p-values in an appendix, and documents the human decisions and delegated checks behind an AI-conducted study.
摘要:一個論點來自特定發言者時可能會令人驚訝,但這並不意味著它是一個糟糕的論點。AI 評估者是否將這些判斷分開?兩項預註冊的描述性研究和後來的 Jev 補充收集了 2,976 個可用的評估,這些評估涉及六篇關於美國 AI 政策、德國的債務制動器和瑞士核能的固定文本。每篇文本在幾個來源歸屬下呈現。關鍵比較在於當論點改變時,兩個來源之間的差距是否會改變。例如,在 Sol 上,國家安全論點在 CODEPINK 下的平均評分為 0.359,而在大學共和黨人下為 0.639;公民權利論點的評分分別為 0.742 和 0.721。對於一個來源的持續偏好無法解釋這種模式。相關的互動出現在不同主題和最近的模型配置中,包括那些啟用推理的配置,而幾個比較則產生了小的效果。後來的歐洲 Jev 補充產生了五個低於採用的絕對參考值 0.05 的互動;其獨特的標準和中斷的收集限制了與聊天系統的比較。一些書面評估明確提到了來源與其歸屬位置之間的不匹配。綜合來看,數字和口頭證據支持來源位置一致性作為一個合理的解釋,並且還有涉及可信度、真實性和任務解釋的競爭說明。本文通過控制比較來發展這一推論,在附錄中報告條件後驗 p 值,並記錄 AI 進行研究背後的人類決策和委派檢查。
Towards Reliable AI Data Scientists: Data Agents with Workflow Harnesses
2609.35255v1 by Huachi Zhou, Yujing Zhang, Jiahe Du, Jiacheng Cai, Zijin Hong, Chuang Zhou, Zheng Yuan, Qinggang Zhang, Qing Li, Xiao Huang
Large language model agents are increasingly deployed for data-intensive work, yet reliable data analysis requires more than general-purpose reasoning and ad hoc tool augmentation. Data Agents, equipped with workflow harnesses, offer a promising paradigm for automating the end-to-end data science lifecycle. This paper examines Data Agents from a harness-centric perspective. First, we introduce a taxonomy of Data Agents and associated data environments, organizing the literature around five functional stages: perception, planning, execution, verification, and repair. Second, we analyze the key technical routes within each stage, identifying 15 distinct approaches ranging from data structure probing to data state reconstruction. Third, we identify four open reliability problems: inactive semantic calibration, missing clarification, missing experience transfer, and the missing verification-repair repository. These problems explain why silent failures can persist even when individual components function correctly, highlighting the need for rigorous workflow harnesses and shared reliability resources. Finally, we summarize the horizontal task families of Data Agents, examine their vertical application settings, and benchmarks for evaluation, while maintaining a companion repository at https://github.com/DEEP-PolyU/Awesome-Data-Agents.
摘要:大型語言模型代理越來越多地被用於數據密集型工作,但可靠的數據分析需要的不僅僅是通用推理和臨時工具增強。
數據代理配備了工作流程鞍具,為自動化端到端數據科學生命周期提供了一種有前景的範式。
本文從鞍具中心的角度檢視數據代理。
首先,我們介紹了一個數據代理及其相關數據環境的分類法,將文獻組織為五個功能階段:感知、規劃、執行、驗證和修復。
其次,我們分析了每個階段內的關鍵技術路徑,識別出15種不同的方法,範圍從數據結構探測到數據狀態重建。
第三,我們確定了四個開放的可靠性問題:非活動的語義校準、缺失的澄清、缺失的經驗轉移以及缺失的驗證-修復庫。
這些問題解釋了為什麼即使單個組件正常運作,靜默失敗仍然會持續存在,突顯了對嚴格工作流程鞍具和共享可靠性資源的需求。
最後,我們總結了數據代理的橫向任務家族,檢視其縱向應用設置及評估基準,同時在 https://github.com/DEEP-PolyU/Awesome-Data-Agents 維護一個伴隨的庫。
Beneath the Tokens: A Performance Engineering Study of Multi-Token Prediction in GPU-Accelerated LLM Inference
2609.35188v1 by Suwesh Prasad Sah
Autoregressive large language model inference repeatedly invokes the target model to generate one token at a time, making generation sensitive to GPU memory movement and sequential execution. This study evaluates two-token multi-token prediction (MTP) against autoregressive decoding in a controlled single-request deployment on an NVIDIA A10G GPU. A 360-request benchmark covered plain-text, reasoning-intensive, and tool-calling workloads, while runtime telemetry, Nsight Systems, PyTorch Profiler, and selected Nsight Compute measurements were used to explain the observed performance. MTP increased output throughput by (1.91\times) to (2.19\times) across all prompts and reduced time to first output by 10.0--14.2\%. Median mean acceptance length ranged from 2.370 to 2.595 tokens per verification iteration. Profiling showed that MTP introduced a longer and more complex execution path, including proposal, sampling, attention, gathering, and reduction operations. However, it required 56.4--78.1\% fewer executions of the selected repeating CUDA Graph per generated token. The dominant MTP GEMM kernel was not faster than the dominant autoregressive GEMV kernel, and selected instances of both approached the A10G memory-bandwidth limit. These results show that MTP improved inference through amortization: greater token progress reduced repeated GPU execution sufficiently to outweigh the additional speculative-execution cost.
摘要:自回歸大型語言模型推理重複調用目標模型以一次生成一個標記,這使得生成對 GPU 記憶體移動和順序執行非常敏感。本研究在 NVIDIA A10G GPU 上的受控單請求部署中評估了兩標記多標記預測(MTP)與自回歸解碼的表現。一個 360 請求的基準涵蓋了純文本、推理密集型和工具調用工作負載,同時使用運行時遙測、Nsight Systems、PyTorch Profiler 和選定的 Nsight Compute 測量來解釋觀察到的性能。MTP 在所有提示中將輸出吞吐量提高了 (1.91\times) 到 (2.19\times),並將首次輸出的時間減少了 10.0--14.2\%。中位數平均接受長度在每次驗證迭代中範圍為 2.370 到 2.595 個標記。分析顯示,MTP 引入了更長且更複雜的執行路徑,包括提案、抽樣、注意力、收集和減少操作。然而,它每生成一個標記所需的選定重複 CUDA Graph 的執行次數減少了 56.4--78.1\%。主導的 MTP GEMM 核心並不比主導的自回歸 GEMV 核心更快,且兩者的選定實例均接近 A10G 記憶體帶寬限制。這些結果顯示,MTP 通過攤銷改善了推理:更大的標記進展足以減少重複的 GPU 執行,以抵消額外的推測執行成本。
Applying Language Models in Clinical Medicine: Recent Trends and Perspectives
2609.34780v2 by Erik Aerts
The use and applicability of artificial intelligence (AI) in medical research and clinical practice has received increasing attention in the literature over recent years. The emergence of large language models (LLMs) has expanded discussions in regards to applications of AI within healthcare. While traditional deep learning based AI applications in medicine have often focused on specific and defined tasks, LLMs offer broader capabilities and flexibility in working with available data,. At the same time of writing, the integration of LLMs into medical settings raises important questions regarding their reliability, accuracy, transparency, safety, and appropriate role in a medical setting. This text presents and discusses recent talks and articles concerning the application of LLMs in medicine, with particular emphasis on their potential utility in research and clinical practice. It considers both the opportunities offered by these technologies and the challenges associated with their implementation, aiming to provide a perspective on the current and emerging role of LLMs within the medical field.
摘要:人工智慧(AI)在醫學研究和臨床實踐中的使用和適用性在近年來的文獻中受到越來越多的關注。大型語言模型(LLMs)的出現擴大了關於AI在醫療保健中應用的討論。雖然傳統基於深度學習的AI應用在醫學中往往專注於特定和明確的任務,但LLMs在處理可用數據方面提供了更廣泛的能力和靈活性。撰寫本文的同時,將LLMs整合進醫療環境中引發了有關其可靠性、準確性、透明度、安全性和在醫療環境中適當角色的重要問題。本文呈現並討論了有關LLMs在醫學中應用的最近演講和文章,特別強調它們在研究和臨床實踐中的潛在效用。它考慮了這些技術所提供的機會以及與其實施相關的挑戰,旨在提供對LLMs在醫療領域中當前和新興角色的看法。
A Unifying Framework of Concept-based Explainable AI with Completeness Guarantees
2609.34750v1 by Vojtěch Kůr, Adam Kukučka, Tomáš Brázdil, Vít Musil
Concept-based explanations describe neural network predictions through human-understandable properties of inputs called concepts. The field encompasses approaches that differ in how they define and represent concepts and connect them to model predictions. We introduce a theoretical framework that describes these approaches in a common mathematical language and supports a shared analysis of their properties. For concept discovery, which identifies concepts automatically within a latent space of a trained model, we employ a concept autoencoder view. An encoder extracts concept representations from the model's latent space, and a decoder uses them to reconstruct the original latent representation. The autoencoder's reconstruction error measures how accurately its decoder recovers the original latent representation. We revisit model completeness: how well the concepts can reproduce the model's outputs. We show that model incompleteness of the concepts can be bounded by the autoencoder's reconstruction error. The autoencoder view also provides a common way to define individual concept attributions, which measure each concept's contribution to a prediction. We establish when these attributions sum to the model's prediction, and bound the discrepancy otherwise, thus providing attribution completeness guarantees.
摘要:概念基礎的解釋通過稱為概念的輸入的人類可理解特性來描述神經網絡的預測。這個領域涵蓋了在定義和表示概念以及將其與模型預測連接方面有所不同的方法。我們引入了一個理論框架,該框架用共同的數學語言描述這些方法,並支持對其特性的共享分析。對於概念發現,即在訓練模型的潛在空間中自動識別概念,我們採用了概念自編碼器的視角。編碼器從模型的潛在空間中提取概念表示,解碼器則利用這些表示重建原始的潛在表示。自編碼器的重建誤差衡量其解碼器恢復原始潛在表示的準確性。我們重新審視模型的完整性:這些概念能多好地再現模型的輸出。我們展示了概念的模型不完整性可以被自編碼器的重建誤差所界定。自編碼器的視角還提供了一種共同的方式來定義個別概念的歸因,這些歸因衡量每個概念對預測的貢獻。我們確立了這些歸因何時加總為模型的預測,並在其他情況下界定了差異,從而提供了歸因完整性的保證。
From Human Narrative to Harmonic Structure: A Human-Centered Investigation of Algorithmic Music Generation through the Chord Wheel Diagram
2609.34735v1 by Josef Pavlíček, Petra Pavlíčková, Irena Štrausová
Contemporary AI-based music generation can produce compositions that satisfy formal requirements of tonality and musical coherence. However, whether musical expression can be described by mathematical properties alone remains a fundamental question. Human composers operate within personal and cultural contexts that influence harmonic decisions and deliberate departures from established patterns. This study investigates six narrative-driven popular songs by Bob Dylan, Johnny Cash, and Ritchie Valens. Original human harmonies are compared with outputs of an explainable computational harmonizer operating on the same melodies without access to the original chord progressions. We examine harmonic vocabulary, functional persistence, repetition, non-diatonic events, and tension-resolution patterns using Chord Wheel Diagrams and BPMN-based representations. Results show that high melody-chord compatibility does not necessarily imply preservation of the original human harmonic decision pattern. Some generated harmonizations retain the economical structure of the reference, while others alter harmonic diversity or suppress distinctive events while remaining compatible with the melody. Rather than quantifying artistic quality, the study introduces narrative-conditioned harmonic structure as a complementary perspective for computational music analysis. The findings suggest that generative systems may benefit from modeling not only harmonic correctness, but also structural identity, context, and human compositional intention.
摘要:當代基於人工智慧的音樂生成可以創作滿足音調和音樂一致性正式要求的作品。
然而,音樂表達是否僅能用數學特性來描述仍然是一個根本問題。
人類作曲家在個人和文化背景中運作,這些背景影響和諧決策以及故意偏離既定模式的選擇。
本研究調查了六首由Bob Dylan、Johnny Cash和Ritchie Valens創作的敘事驅動流行歌曲。
原始的人類和聲與一個可解釋的計算和聲器的輸出進行比較,該計算和聲器在沒有訪問原始和弦進行的情況下對相同旋律進行操作。
我們使用和弦輪圖和基於BPMN的表示法來檢查和聲詞彙、功能持續性、重複、非調性事件和緊張-解決模式。
結果顯示,高旋律與和弦的相容性並不一定意味著保留原始人類和聲決策模式。
一些生成的和聲保留了參考的經濟結構,而其他則改變了和聲的多樣性或壓制了獨特事件,同時仍與旋律相容。
本研究並非量化藝術品質,而是引入敘事條件的和聲結構作為計算音樂分析的補充視角。
研究結果表明,生成系統可能受益於不僅建模和聲正確性,還包括結構身份、上下文和人類創作意圖。
Understanding Generalization Requires Universal Induction
2609.34458v1 by Aram Ebtekar, Marcus Hutter, Danica J. Sutherland
Classical statistical theory is insufficient to explain the successes of general-purpose AI models, because it depends on handcrafted inductive biases that it cannot justify. No Free Lunch (NFL) theorems force any learner that beats chance on some environments to underperform on others. We might hope that past experience informs which environments to expect, but NFL applies equally to meta-learning. Thus, any method that makes meaningful predictions necessarily begins with an inductive bias external to the data. Choosing to bias toward short programs yields Solomonoff induction (SI), whose performance is competitive against all computable learners - albeit up to "constants" that become large when comparing against specialized methods that exploit background information. We therefore relativize SI to an information vantage point, biasing toward short programs with access to all preexisting information. This reframes the inductive bias: instead of seeking some absolute notion of simplicity, we favor accessibility with respect to our vantage point. An algorithm can only outpredict the relativized SI to the extent that its code contains additional information about the data, and no algorithm can generate such information. While SI is incomputable and hence not a practical algorithm, it provides a formal optimum for inference in the limit of infinite compute, and there is evidence to suggest that frontier AI systems roughly approximate it. Thus, the only known answer to meta-NFL is rooted in algorithmic information theory, which we should expect to play a fundamental role in explaining the generalization behavior of modern (and future) AI systems.
摘要:古典統計理論不足以解釋通用人工智慧模型的成功,因為它依賴於無法證明的手工製作的歸納偏見。無免費午餐(NFL)定理迫使任何在某些環境中超越隨機的學習者在其他環境中表現不佳。我們可能希望過去的經驗能告訴我們預期哪些環境,但NFL同樣適用於元學習。因此,任何能做出有意義預測的方法必然以一種外部於數據的歸納偏見開始。選擇偏向短程序會產生所羅門諾夫歸納(SI),其性能在所有可計算學習者中具有競爭力——儘管在與利用背景信息的專門方法比較時,這些“常數”會變得很大。因此,我們將SI相對化到一個信息視角,偏向於短程序並訪問所有現有信息。這重新框定了歸納偏見:我們不再尋求某種絕對的簡單性概念,而是更重視相對於我們視角的可及性。一個算法只能在其代碼包含有關數據的額外信息的程度上超越相對化的SI,且沒有任何算法能生成這種信息。雖然SI是不可計算的,因此不是一個實用的算法,但它為在無限計算的極限下的推理提供了一個形式上的最優解,並且有證據表明前沿人工智慧系統大致上近似於它。因此,對於元NFL唯一已知的答案根植於算法信息理論,我們應該預期它在解釋現代(和未來)人工智慧系統的泛化行為中扮演基本角色。
Social Circuits behind Multi-agent Echo Chambers
2609.34444v1 by Chuiyang Meng, Wenlu Yu, Ming Tang, Cheng Li
Language-model agents exchange messages to combine evidence, but their communication can also create echo chambers that reinforce shared errors. However, overall task performance does not explain how a message changes the receiving agent's internal activations and affects its decision. In this work, we introduce Social Circuits, a framework for tracing message effects through receiver activations. We compare the receiver's answers before and after changing a message. Then, we restore selected activations recorded under the original message to determine how much of the message effect these activations reproduce. Based on Social Circuits, we propose Circuit-Guided Deliberation (CGD), which learns to select useful messages using receiver activation changes. We establish when activation replacement preserves receiver decisions and bound the gap between CGD's task performance and the best achievable through message selection. Experiments show that receiver activation changes explain the message effects and guide message selection that improves the task performance. Across three models and four datasets, CGD achieves the highest or joint-highest average accuracy in our main comparisons while generating fewer tokens than multi-agent baselines.
摘要:語言模型代理之間交換訊息以結合證據,但它們的溝通也可能創造回音室,強化共同的錯誤。然而,整體任務表現並不能解釋一條訊息如何改變接收代理的內部激活並影響其決策。在這項工作中,我們引入社會電路(Social Circuits),這是一個追蹤訊息影響的框架,通過接收者的激活來進行分析。我們比較接收者在改變訊息前後的回答。然後,我們恢復在原始訊息下記錄的選定激活,以確定這些激活重現了訊息效果的多少。基於社會電路,我們提出了電路引導的深思(Circuit-Guided Deliberation, CGD),它學會利用接收者激活變化來選擇有用的訊息。我們確定何時激活替換能夠保留接收者的決策,並界定CGD的任務表現與通過訊息選擇所能達到的最佳表現之間的差距。實驗顯示,接收者激活變化解釋了訊息效果,並指導訊息選擇以改善任務表現。在三個模型和四個數據集上,CGD在我們的主要比較中達到了最高或並列最高的平均準確率,同時生成的標記數量少於多代理基準。
Improving Large Language Models for Code through Runtime Program-State Reasoning
2609.34359v1 by Hongwei Li, Spandan Garg, Yufan Huang
Large language models receive limited explicit training in reasoning about runtime program states. We study whether training models to reason about runtime program states improves downstream software-engineering capabilities. We introduce two complementary program-state reasoning tasks. Buggy input-output reasoning requires a model to generate a concrete input that exposes a behavioral difference between a buggy program and a hidden correct implementation and to predict the resulting execution behavior. Precondition-postcondition reasoning requires an agent to symbolically characterize a bug-triggering precondition, predict the expected postcondition, explain their causal connection, and instantiate this reasoning as an executable regression test. By incorporating these two tasks into a staged post-training pipeline, we develop Comet-9B, a 9B language model based on Qwen3.5-9B Base. We evaluate the resulting checkpoints on repository-level patch generation, regression-test generation, and security PoC generation. Adding both program-state reasoning tasks to supervised fine-tuning (SFT) on issue resolution improves success rates by 7.25 percentage points on SWE-bench Pro and 9.70 points on SWT-Bench Verified. Sequential reinforcement learning on the two tasks yields further gains of 7.25, 26.79, and 4.67 percentage points on SWE-bench Pro, SWT-Bench Verified, and CyberGym, respectively. Despite having only 9B parameters, Comet-9B achieves a score comparable to the reported GPT-5.2 result on SWE-bench Pro and matches the reported success rate of a GPT-4o-based agent on SWT-Bench Verified.
摘要:大型語言模型在推理運行時程序狀態方面接受的明確訓練有限。
我們研究訓練模型推理運行時程序狀態是否能改善下游軟體工程能力。
我們引入了兩個互補的程序狀態推理任務。
有缺陷的輸入輸出推理要求模型生成一個具體的輸入,該輸入能揭示有缺陷的程序和隱藏的正確實現之間的行為差異,並預測結果執行行為。
前置條件-後置條件推理要求代理符號化地描述觸發錯誤的前置條件,預測預期的後置條件,解釋它們之間的因果關係,並將這一推理實例化為可執行的回歸測試。
通過將這兩個任務納入分階段的後訓練流程,我們開發了 Comet-9B,一個基於 Qwen3.5-9B Base 的 9B 語言模型。
我們在庫級補丁生成、回歸測試生成和安全 PoC 生成上評估了結果檢查點。
將這兩個程序狀態推理任務添加到針對問題解決的監督微調(SFT)中,使成功率在 SWE-bench Pro 上提高了 7.25 個百分點,在 SWT-Bench Verified 上提高了 9.70 個百分點。
在這兩個任務上進行的序列強化學習分別在 SWE-bench Pro、SWT-Bench Verified 和 CyberGym 上獲得了進一步的增益,分別為 7.25、26.79 和 4.67 個百分點。
儘管只有 9B 參數,Comet-9B 在 SWE-bench Pro 上達到了與報告的 GPT-5.2 結果相當的分數,並且與基於 GPT-4o 的代理在 SWT-Bench Verified 上的報告成功率相匹配。
Dynamical Parameters: An Interpretability Framework for Time-Series Foundation Models
2609.34316v1 by Kang Yang, Gaofeng Dong, Liying Han, Mani Srivastava
This work studies a central gap in interpreting time-series foundation models (TSFMs): a dynamical property may be accessible in a hidden state even when the forecast fails to respond correctly as that property changes. We formalize these properties as Dynamical Parameters, including trend slope, oscillation frequency, and autoregressive dependence. We compare their representation accessibility, measured by recovery from hidden states, with their forecast response, measured by agreement with the expected forecast change. Across nine frozen TSFMs and thirteen laws, 42 of 63 model-parameter cells achieve accessibility above 0.95, whereas their median reference-aligned response relative to the conditional reference is only 0.46. To explain this gap, causal geometry compares the hidden-state change required to produce the reference response with the change induced by the parameter intervention. Directly modifying the hidden state recovers the reference response, but the parameter intervention often moves the state in a different direction. These results show that accessible parameter information need not be expressed in forecasts when input changes miss the required hidden-state direction.
摘要:這項工作研究了解釋時間序列基礎模型(TSFMs)中的一個核心差距:即使預測未能正確響應該特性變化,動態性質也可能在隱藏狀態中可訪問。我們將這些性質形式化為動態參數,包括趨勢斜率、振盪頻率和自回歸依賴性。我們比較它們的表示可訪問性,通過從隱藏狀態的恢復來測量,與它們的預測響應進行比較,後者通過與預期預測變化的一致性來測量。在九個凍結的TSFMs和十三條法則中,63個模型參數單元中有42個達到0.95以上的可訪問性,而它們相對於條件參考的中位數參考對齊響應僅為0.46。為了解釋這一差距,因果幾何比較了產生參考響應所需的隱藏狀態變化與參數干預所引起的變化。直接修改隱藏狀態可以恢復參考響應,但參數干預往往會將狀態移動到不同的方向。這些結果表明,當輸入變化錯過所需的隱藏狀態方向時,可訪問的參數信息不必在預測中表達。
Evo2Team: When Do Evolved Skills Transfer? From Selection to Deployment
2609.34135v1 by Renxiang Wang, Jiaming Cui
A skill bank that helps one multi-agent system may leave another's behavior unchanged. A transferred rule helps only when target agents act on it successfully. We study this path for routing and communication skills in Count-Frequency and AgentsNet, using teams of 4--32 agents and GPT and Qwen model ladders. Source evolution meets a joint quality, cost, model-tier, and confirmation goal in 14 of 16 settings. We then evaluate Evo2Team, which selects, adapts, and confirms source skills for the target team, alongside six frozen selectors across 28 transfer directions. Evo2Team's target-side exploration cost is below that of evolving a new target bank in every direction, even when reused reference evaluations are charged once. Twenty of 28 held-out outcomes meet the positive-transfer criterion, including three saved diagnostic tests. Selection alone does not explain these outcomes: KNN and CORAL choose different banks in two AgentsNet directions but produce identical recorded executions. When Evo2Team changes execution, gains can reach many tasks, as in a Count-Frequency direction that improves 28 of 32 tasks over KNN. Seven positive AgentsNet outcomes save 6.1--14.6\% in deployment cost while using transferred skills on only three to six of fifteen tasks. In five earlier accepted directions, all 22 task records using transferred skills pass three fixed-graph confirmations, but four fail in recorded executions on new graphs. Graphs and model responses change together in this comparison. These results show that skill transfer must be assessed through the actions agents take, the tasks those actions reach, and the quality and cost of the final deployment.
摘要:一個幫助某個多代理系統的技能庫可能不會改變另一個系統的行為。當目標代理成功地執行轉移的規則時,這個規則才會有幫助。我們研究了在 Count-Frequency 和 AgentsNet 中的路由和通信技能,使用 4 至 32 個代理的團隊以及 GPT 和 Qwen 模型梯度。在 16 個設置中,有 14 個達成了源演化的聯合質量、成本、模型層級和確認目標。我們接著評估 Evo2Team,該系統為目標團隊選擇、調整和確認源技能,並在 28 個轉移方向中使用六個固定選擇器。Evo2Team 的目標側探索成本在每個方向上都低於演化一個新的目標庫的成本,即使重用的參考評估只收費一次。28 個保留結果中有 20 個符合正向轉移標準,包括三個保存的診斷測試。僅僅依賴選擇無法解釋這些結果:KNN 和 CORAL 在兩個 AgentsNet 方向中選擇不同的庫,但產生相同的記錄執行。當 Evo2Team 改變執行時,收益可以達到許多任務,例如在一個 Count-Frequency 方向上,KNN 的 32 個任務中有 28 個得到了改善。七個正向的 AgentsNet 結果在只使用轉移技能的十五個任務中的三到六個任務上節省了 6.1% 到 14.6% 的部署成本。在五個早期接受的方向中,所有 22 個使用轉移技能的任務記錄都通過了三個固定圖形的確認,但在新圖形上的記錄執行中有四個失敗。在這次比較中,圖形和模型反應是一起改變的。這些結果顯示,技能轉移必須通過代理所採取的行動、這些行動所達到的任務以及最終部署的質量和成本來進行評估。
JET: Judge-Guided Evolution at Test Time for Agent Programs
2609.34126v1 by Yao Long Teng, Jiayi Cai, Bo An
An agent's executable program governs how it uses tools, processes observations, and responds to failures. Evolving this program at test time can help adaptation, but deciding which changes to retain is difficult when true rewards are unavailable. Execution traces provide evidence of agent behavior, yet interpreting that evidence requires a judge that remains useful as tasks and candidate programs change. We introduce Judge-Guided Evolution at Test Time (JET), which evolves an executable judge on labeled source trajectories, then freezes and transfers it to guide target-side program evolution. The judge supplies scores and diagnostic feedback without target evaluator access or model-weight updates. On unseen WebShop tasks, JET achieves approximately 13% higher mean reward than fixed-rubric guidance when evolution begins from an unevolved program (cold start) and 4% higher when it begins from one already optimized on source tasks (warm start), with a 36% relative improvement in cold-start exact success. An exact-judge control on PushT, where the judge reconstructs the scoring rule from observations, shows that without judge error, program search becomes the bottleneck. Analyses identify useful reward-prediction logic in the evolved code and show that better final selection alone cannot explain the gains. These results support executable judge transfer for program adaptation under evaluator-preserving task shifts.
摘要:一個代理的可執行程序決定了它如何使用工具、處理觀察結果以及對失敗的反應。 在測試時進化這個程序可以幫助適應,但在真正的獎勵不可用時,決定保留哪些變更是困難的。 執行痕跡提供了代理行為的證據,但解釋這些證據需要一個在任務和候選程序變化時仍然有用的評判者。 我們介紹了測試時的評判者引導進化(JET),它在標記的源軌跡上進化一個可執行的評判者,然後將其凍結並轉移到目標端程序進化的指導。 評判者提供分數和診斷反饋,而不需要目標評估者的訪問或模型權重的更新。 在未見過的WebShop任務上,當進化從未進化的程序(冷啟動)開始時,JET的平均獎勵比固定標準指導高出約13%,而當它從已在源任務上優化的程序(熱啟動)開始時,高出4%,冷啟動的精確成功率提高了36%。 在PushT上的精確評判者控制中,評判者從觀察中重建評分規則,顯示在沒有評判者錯誤的情況下,程序搜索成為瓶頸。 分析確定了進化代碼中有用的獎勵預測邏輯,並顯示僅僅更好的最終選擇無法解釋增益。 這些結果支持在保留評估者的任務變化下進行可執行評判者轉移以適應程序。
Do World Models Learn Global Understanding?
2609.34058v1 by Alexander Detkov, Matt Thomson
AI systems often feel brittle and fragmented. A large language model (LLM) may correctly explain a concept but fail to apply it, or follow safety instructions in one context but not another. This behavior suggests a general failure to lift local information to a global understanding. To gain fundamental insight, we frame "understanding" as learning constraints and propagating their consequences. We construct learning tasks on monoid worlds, sets of states connected by action transitions, where observed training transitions and an unseen constraint jointly determine held-out transitions. Measuring generalization tests whether models can learn global constraints from local transitions and propagate their consequences. We consider inverse, commutativity, composition, and periodicity constraints relevant to spatial and semantic structure. Across attention, recurrent, and state-space architectures, next-state training fits the data but fails to propagate non-trivial constraints. Compositional training, which uses identical paths but hides intermediate states from the input, achieves 96% accuracy on inverse, commutativity, and composition constraints across architectures, yields corresponding improvements in geometric generalization of world models trained on embodied environments and relational generalization in Wikidata-finetuned LLMs. How far do models propagate constraints when inferring an unseen fact may depend on first inferring others? We define proof depth d of a held-out transition, measuring the minimum number of inference rounds to infer the transition, and find that model generalization decreases sharply with proof depth. Increasing compositional path length T improves generalization. These results provide a formal way to investigate global understanding in language and world models and demonstrate that compositional training promotes information propagation and integration.
摘要:AI 系統經常感覺脆弱且支離破碎。一個大型語言模型 (LLM) 可能正確解釋一個概念,但無法應用它,或者在一個情境中遵循安全指示,但在另一個情境中卻不然。這種行為表明在將局部信息提升到全球理解方面存在普遍失敗。為了獲得基本見解,我們將「理解」框架化為學習約束並傳播其後果。我們在單元世界上構建學習任務,這些世界是由行動轉換相連的狀態集合,其中觀察到的訓練轉換和未見的約束共同決定了保留的轉換。測量泛化測試模型是否能從局部轉換中學習全球約束並傳播其後果。我們考慮與空間和語義結構相關的逆、交換性、組合性和周期性約束。在注意力、遞歸和狀態空間架構中,下一狀態訓練適合數據,但未能傳播非平凡約束。組合訓練使用相同的路徑,但將中間狀態從輸入中隱藏,在各架構上對逆、交換性和組合性約束達到 96% 的準確率,並在基於具體環境訓練的世界模型中產生相應的幾何泛化改善,以及在經過 Wikidata 微調的 LLM 中的關係泛化。當推斷一個未見的事實可能依賴於首先推斷其他事實時,模型傳播約束的程度有多遠?我們定義保留轉換的證明深度 d,測量推斷該轉換所需的最小推理輪次,並發現模型的泛化隨著證明深度的增加而急劇下降。增加組合路徑長度 T 改善了泛化。這些結果提供了一種正式的方法來研究語言和世界模型中的全球理解,並證明組合訓練促進信息的傳播和整合。
Uncovering shortcut learning in audio classifiers by discovering recurring concepts in temporal explanations
2609.34030v1 by Cecilia Bolaños, Luciana Ferrer, Magdalena Fuentes
Correlations between events in machine learning datasets may result in shortcut learning, where models learn to predict the target event based on the presence of a correlated event. When these correlations are spurious -- arising from data collection artifacts -- models are likely to perform poorly in practice. We propose a pipeline to uncover shortcut learning in audio classifiers by discovering recurring concepts in their temporal explanations. Specifically, we isolate audio segments that explain classifier decisions, caption them with an ensemble of Large Audio-Language Models, and use a Large Language Model to extract recurring concepts. The resulting concepts can be audited by humans to uncover potential shortcut learning. We evaluate our framework using datasets curated from AudioSet Strong, controlling for the presence or absence of spurious correlations. Results show that this approach reliably uncovers learned shortcuts, such as the model relying on the presence of "laughter" to predict "applause".
摘要:事件之間的相關性在機器學習數據集中可能導致捷徑學習,模型學會根據相關事件的存在來預測目標事件。當這些相關性是虛假的——源於數據收集的工藝問題——模型在實際應用中可能表現不佳。我們提出了一個管道來揭示音頻分類器中的捷徑學習,通過發現它們時間解釋中的重複概念。具體而言,我們隔離解釋分類器決策的音頻片段,使用一組大型音頻-語言模型為其標題,並利用大型語言模型提取重複概念。所得到的概念可以由人類進行審核,以揭示潛在的捷徑學習。我們使用從AudioSet Strong整理的數據集來評估我們的框架,控制虛假相關性的存在或缺失。結果顯示,這種方法可靠地揭示了學習到的捷徑,例如模型依賴「笑聲」的存在來預測「掌聲」。
Safety Reconstructed: Generative Modeling via Masked Diffusion Builds Strong Safety Guardrails
2609.33634v1 by Gert Lek, Abele Malan, Chaoyi Zhu, Pin-Yu Chen, Robert Birke, Lydia Chen
Guard models are the last line of defense between a language model and a harmful output, yet their training objective is surprisingly narrow. Existing guards learn to predict a single verdict token from a conversational context, concentrating supervision on a single target. The consequences are structural: models latch onto shortcut features, are overconfident, and remain sensitive to where safety evidence appears in the sequence rather than its role in the full context. We propose a different framing. Rather than predicting a label from text, our LLaDA-Guard asks which label better explains the text: scoring the prompt or response under each label hypothesis and classifying based on their difference. This shifts supervision to every token in the moderated region, forcing the model to account for full content rather than its most discriminative fragments. We instantiate this idea with a masked diffusion language model, fine-tuning LLaDA-8B-Instruct with a class-conditional reconstruction objective using LoRA and requiring no architectural changes beyond the base model. LLaDA-Guard leads on average rank against discriminative baselines trained on stronger backbones across seven held-out safety benchmarks, while exhibiting substantially better confidence calibration (ECE 0.0875 vs. 0.1384 for Qwen3Guard), less over-defense on benign prompts with unsafe-looking cues, and less prompt leakage when moderating responses. Its generative nature further enables token-level risk localization as a natural byproduct, yielding a pipeline for rewriting unsafe prompts into safe equivalents without additional training and achieving a 60.7% average conversion-to-safe rate.
摘要:守衛模型是語言模型與有害輸出之間的最後防線,但它們的訓練目標卻出乎意料地狹窄。現有的守衛學習從對話上下文中預測單一的判決標記,將監督集中在單一目標上。這帶來了結構性的後果:模型依賴於捷徑特徵,過於自信,並對安全證據在序列中出現的位置保持敏感,而不是其在完整上下文中的角色。我們提出了一種不同的框架。我們的LLaDA-Guard不是從文本中預測標籤,而是詢問哪個標籤更好地解釋文本:在每個標籤假設下對提示或回應進行打分,並根據它們的差異進行分類。這將監督轉移到被調節區域中的每個標記,迫使模型考慮完整內容,而不是其最具區分性的片段。我們用一個掩蔽擴散語言模型來實現這個想法,通過使用LoRA的類條件重建目標來微調LLaDA-8B-Instruct,並且不需要超出基礎模型的架構變更。LLaDA-Guard在七個保留的安全基準中,對比在更強的基礎模型上訓練的區分基準,平均排名領先,同時顯示出顯著更好的信心校準(ECE 0.0875對比Qwen3Guard的0.1384),在具有不安全外觀提示的良性提示上過度防禦較少,並在調節回應時提示洩漏較少。其生成特性進一步使得標記級風險定位成為一種自然副產品,產生了一個將不安全提示重寫為安全等價物的管道,無需額外訓練,並實現了60.7%的平均轉換為安全率。
Learning Dynamics of Continual Learning: A Unified View of Data Attribution, Forgetting, and Plasticity Loss
2609.33620v1 by Yi Ren, Wenlong Deng, Guanzhe Hong, Clare Lyle, Yarin Gal
Modern language models are likely to be updated throughout their lifetime rather than trained once and frozen. Each update therefore participates in a recurring cycle: decide which experience to learn from, understand what that update changes, and remain capable of learning from what comes next. We show that these challenges are governed by the same evolving update--behavior interaction. We derive a token- and layer-wise decomposition of how learning from one token changes another prediction. By separating the softmax force, shared readout geometry, and residual connections, it exposes two interaction channels and yields a forward-computable approximation. Following this interaction through time reveals a unified picture of continual adaptation. Positive interaction identifies useful experience; negative interaction produces either concentrated collision or accumulated erosion; over longer horizons, updates reshape the shared geometry mediating future learning signals, reducing their transmission. These predictions lead to effective data selection, mechanism-specific controls for interference, and a readout-based diagnostic of future learnability whose degradation predicts the benefit of restoring the readout. Across models and training regimes, the same local interaction thus explains both what an update changes now and how learning today changes what can be learned tomorrow. This view connects data attribution, forgetting, and plasticity loss as distinct regimes of the same evolving learning dynamics.
摘要:現代語言模型在其生命週期內可能會不斷更新,而不是一次訓練後就凍結。因此,每次更新都參與一個重複的循環:決定從哪種經驗中學習,理解這次更新改變了什麼,並保持能夠從接下來的經驗中學習。我們顯示這些挑戰受到相同演變的更新—行為互動的支配。我們推導出一種基於標記和層的分解,說明從一個標記學習如何改變另一個預測。通過分離softmax力、共享讀取幾何和殘差連接,它揭示了兩個互動通道並產生了一個可前向計算的近似。隨著時間推移,跟隨這種互動揭示了持續適應的統一圖景。正向互動識別有用的經驗;負向互動則產生集中碰撞或累積侵蝕;在更長的時間範圍內,更新重塑了調解未來學習信號的共享幾何,減少了其傳輸。這些預測導致有效的數據選擇、特定機制的干擾控制,以及基於讀取的未來可學習性的診斷,其退化預測了恢復讀取的好處。在模型和訓練體系中,相同的局部互動因此解釋了更新當前改變了什麼,以及今天的學習如何改變明天可以學習的內容。這一觀點將數據歸因、遺忘和可塑性喪失連接為同一演變學習動態的不同範疇。
Temporal Graph Learning of Wearable Actigraphy and Sleep Traces for Modelling Adolescent Crystallized Intelligence
2609.33428v1 by Md. Tanvir Rahman, Nabil Anan Orka, Asaduzzaman Khan, Mohammad Ali Moni
Wearable actigraphy offers a scalable, ecologically valid alternative to episodic clinical assessment. However, predicting continuous adolescent crystallized intelligence ($G_c$) from such traces remains challenging due to irregular device adherence and complex behavioral-environmental interactions. We address this using daily summary data derived from 21-day Fitbit records of 6,091 adolescents in the Adolescent Brain Cognitive Development Study (Release 5.1). We propose SATURN, a Sleep-Activity Temporal Unified Regression Network. It represents participants as 21-node temporal graphs encoding daily behaviors and temporal adjacency. To prevent imputation artifacts, invalid-day edges are dynamically pruned during forward passes. Node embeddings are refined via residual GATv2 layers, aggregated through masked attention pooling, and fused with sociodemographic covariates. Under family-controlled, age-sex-BMI-stratified cross-validation, SATURN achieves $R^2 = 0.2783 \pm 0.0127$, consistently improving upon flattened machine learning (Gradient Boosting, $R^2 = 0.2372$) and sequential deep learning (BiLSTM, $R^2 = 0.2688$) baselines. Explainability analyses identify light activity, metabolic equivalents, and sleep duration as dominant predictors, while Monte Carlo dropout and subgroup analyses confirm equitable performance across sociodemographic strata. Ultimately, SATURN establishes a rigorous computational framework for digital cognitive phenotyping, offering a scalable pathway to complement traditional assessments by highlighting macro-level behavioral anomalies.
摘要:可穿戴行為測量提供了一種可擴展的、生態有效的替代方案,以取代臨床評估的偶發性。然而,從這些數據中預測持續的青少年結晶智力 ($G_c$) 仍然具有挑戰性,因為設備遵從性不規則且行為與環境之間的互動複雜。我們使用來自 6,091 名青少年在青少年大腦認知發展研究(版本 5.1)中,為期 21 天的 Fitbit 記錄所衍生的每日摘要數據來解決這個問題。我們提出了 SATURN,一個睡眠-活動時間統一回歸網絡。它將參與者表示為 21 節點的時間圖,編碼每日行為和時間相鄰性。為了防止插補伪影,在前向傳播過程中動態修剪無效日邊緣。節點嵌入通過殘差 GATv2 層進行精煉,通過遮罩注意力池化進行聚合,並與社會人口學協變量融合。在家庭控制、年齡-性別-BMI 分層的交叉驗證下,SATURN 的 $R^2 = 0.2783 \pm 0.0127$,持續優於扁平化的機器學習(梯度提升,$R^2 = 0.2372$)和序列深度學習(BiLSTM,$R^2 = 0.2688$)基準。可解釋性分析確定輕度活動、代謝當量和睡眠持續時間為主要預測因子,而蒙特卡羅隨機失活和子群分析則確認了在社會人口學層次上表現公平。最終,SATURN 建立了一個嚴謹的計算框架,用於數位認知表型,提供了一條可擴展的途徑,以通過突顯宏觀層面的行為異常來補充傳統評估。
Explainable Deep Learning of Resting-State Functional Connectomes Reveals Network Biomarkers of Adolescent Intelligence
2609.33422v1 by Md. Tanvir Rahman, Nabil Anan Orka, Asaduzzaman Khan, Mohammad Ali Moni
Mapping resting-state brain organization to individual differences in cognitive ability remains a major challenge in population neuroinformatics. Although deep learning enables flexible modeling of brain connectivity, limited interpretability restricts its scientific and clinical utility. To address this objective, we developed an explainable deep learning framework based on sparse projected residual networks to predict fluid, crystallized, and total intelligence from resting-state functional magnetic resonance imaging in 5,285 participants from the Adolescent Brain Cognitive Development study. We incorporated three complementary explainability methods (Integrated Gradients, Gradient Shapley Additive Explanations, and Occlusion) to interpret model behavior. The framework outperformed existing approaches, achieving Pearson correlations of 0.44, 0.58, and 0.56 for fluid, crystallized, and total intelligence, respectively, corresponding to predictive improvements of 6 to 9 percent. All three explainability methods produced near-identical feature rankings (pairwise rank correlations greater than 0.99). Consensus maps revealed a dual-layered functional architecture where primary predictive hubs localized within canonical systems, while the strongest global predictive pathways frequently bypassed these hubs through distributed, long-range relay connections. These findings suggest that intelligence emerges from the interaction between localized computational hubs and distributed communication pathways. Ultimately, these normative network architectures provide clinical reference maps to detect individual deviations, supporting earlier diagnosis, cognitive subtype stratification, and treatment monitoring in atypical neurodevelopment.
摘要:將靜息狀態下的大腦組織映射到個體在認知能力上的差異,仍然是人口神經資訊學中的一大挑戰。雖然深度學習使得大腦連接的靈活建模成為可能,但有限的可解釋性限制了其科學和臨床的實用性。為了達成這一目標,我們開發了一個基於稀疏投影殘差網絡的可解釋深度學習框架,從5,285名來自青少年大腦認知發展研究的參與者的靜息狀態功能性磁共振成像中預測流體智力、結晶智力和總智力。我們結合了三種互補的可解釋性方法(整合梯度、梯度沙普利加法解釋和遮蔽)來解釋模型行為。該框架的表現超過了現有的方法,對流體智力、結晶智力和總智力的皮爾森相關係數分別達到0.44、0.58和0.56,對應的預測改進為6%到9%。所有三種可解釋性方法產生了幾乎相同的特徵排名(成對排名相關係數大於0.99)。共識圖揭示了一種雙層功能架構,其中主要的預測樞紐位於典型系統內,而最強的全球預測通路則經常通過分散的長距離中繼連接繞過這些樞紐。這些發現表明,智力是由局部計算樞紐和分散通信通路之間的互動所產生的。最終,這些規範性網絡架構提供了臨床參考圖,以檢測個體偏差,支持早期診斷、認知亞型分層和在非典型神經發展中的治療監測。
Decoupling Token Roles in Autoregressive Pretraining
2609.33405v1 by Suqin Yuan, Runqi Lin, Kevin Qinghong Lin, Junchi Yu, Lei Feng, Chris Russell, Tongliang Liu
Autoregressive pretraining increasingly draws on heterogeneous data, making it important to understand how a model learns from an individual token. The next-token prediction objective naturally identifies a token's contribution with its own loss. However, each token is not only a prediction target but also context for what follows. Using controlled corruption, we decouple these two roles and find a reversal: making a noisy token easier to predict reduces its damage as a target but increases it as context. The same decoupling helps explain text generated by language models: generation selects each token by its fit to the prefix, while its role as context is never tested against an independently determined continuation, because that continuation is generated to fit it. At known corrupted positions, acting through the context can reduce damage that removing the token's own loss does not. Understanding and controlling what a model learns from a token therefore requires decoupling its roles.
摘要:自回歸預訓練越來越依賴異質數據,因此理解模型如何從單個標記中學習變得重要。下一個標記的預測目標自然將標記的貢獻與其自身的損失相識別。然而,每個標記不僅是預測目標,也是後續內容的上下文。通過使用控制性腐敗,我們將這兩個角色解耦,並發現了一種逆轉:使一個嘈雜的標記更容易預測會減少其作為目標的損害,但會增加其作為上下文的損害。同樣的解耦有助於解釋語言模型生成的文本:生成過程根據每個標記與前綴的契合度進行選擇,而其作為上下文的角色從未與獨立確定的延續進行測試,因為該延續是為了適應它而生成的。在已知的腐敗位置,通過上下文的作用可以減少去除標記自身損失所無法減少的損害。因此,理解和控制模型從標記中學習的內容需要解耦其角色。
The Error You See Is Not the Error You Made: Progression-aware Reasoning Origin for Reasoning Error Localization
2609.33297v1 by Yiguo Wang, Ziyuan Yang, Yi Zou, Dan Lin, Rongsheng Li, Yi Zhang
Verifying multi-step LLM reasoning requires more than determining whether a trace is correct: a useful verifier should identify where the reasoning first goes wrong. However, existing holistic methods provide little positional evidence, while forward sequential verification often treats the first rejected step as the error source. Under error propagation, this assumption can fail, since an earlier mistake may remain locally plausible and become observable only through its downstream consequences. We therefore rethink reasoning verification as a progression-aware error-source localization problem: rather than asking only where a reasoning trace first appears inconsistent, we ask which earlier step best explains how that inconsistency emerges along the trajectory. Based on this view, we propose Progression-aware Reasoning Origin (PRO), a training-free framework for first-error localization. PRO jointly models incoming support from the preceding context and outgoing compatibility with subsequent reasoning, selectively refines regions where these signals disagree, and finally performs detector-conditioned source attribution with intervention-based evidence to distinguish the true error origin from its propagated manifestations. We further formalize the gap between forward rejection and structural exposure, showing why incoming-side evidence alone is insufficient for reliable localization under error propagation. Experiments across open-form, medical, and structured reasoning tasks demonstrate consistent improvements over strong verification baselines, supporting progression-aware source attribution as a more faithful formulation of reasoning verification.
摘要:驗證多步驟 LLM 推理不僅需要確定一個痕跡是否正確:一個有用的驗證器應該能夠識別推理首次出錯的地方。
然而,現有的整體方法提供的位置信息有限,而前向序列驗證通常將第一個被拒絕的步驟視為錯誤來源。在錯誤傳播的情況下,這一假設可能會失效,因為早期的錯誤可能在局部上仍然是合理的,並且只有通過其下游後果才能被觀察到。
因此,我們重新思考推理驗證,將其視為一個進程感知的錯誤來源定位問題:我們不僅詢問推理痕跡首次出現不一致的地方,而是詢問哪一個早期步驟最能解釋沿著軌跡出現的不一致。
基於這一觀點,我們提出了進程感知推理來源(PRO),這是一個無需訓練的首錯定位框架。
PRO 共同建模來自前一上下文的支持和與後續推理的兼容性,選擇性地細化這些信號不一致的區域,並最終通過基於干預的證據進行檢測器條件的來源歸因,以區分真實的錯誤來源和其傳播的表現。
我們進一步形式化了前向拒絕和結構曝光之間的差距,顯示為什麼僅依賴來自進入側的證據對於在錯誤傳播下的可靠定位是不足夠的。
在開放式、醫療和結構化推理任務中的實驗顯示出對強驗證基準的一致改進,支持進程感知來源歸因作為推理驗證的更真實表述。
CORTEX: A Verified Experience Layer for Generalist Agents
2609.33260v1 by Garapati Keerthana, Manik Gupta
An agent can solve a task today and face the same task under new facts, tools, or governing knowledge tomorrow. Most agent systems can retrieve relevant text or recall prior conversations, but they lack a principled way to decide when a previous solution is still valid, when it must be adapted, and when it should be discarded. We introduce CORTEX (Contextual Orchestration and Reuse of Task EXperience), a general AI systems framework that connects specialized agents through an external layer of verified experience. Each episode records its task conditions, source and tool state, decisive predicates, proof trace, verifier, and outcome. A meta-controller chooses exact replay, checked adaptation, fresh synthesis, or escalation. Accepted episodes can become task patterns and procedural strategies through a challenge-driven development loop. This gives the system an implicit competence layer that can grow without changing model weights. We formalize system contracts for exact replay and source-version separation, and derive when reuse saves computation. A controlled two-domain implementation tests the exact-replay core on 1,000 synthetic cases. Complete-family holdouts test procedural transfer on 1,000 new-family cases across eight clinical and policy splits, with complete fresh-evidence grounding and perfect invariance to irrelevant-field and insertion-order perturbations. The transfer trace exposes the work required for verified strategy execution. These results establish an initial path toward general intelligence through reusable procedures, typed experience, and developmental transfer.
摘要:一個代理可以在今天解決一個任務,並在明天面對同一任務,但有新的事實、工具或治理知識。大多數代理系統可以檢索相關文本或回憶先前的對話,但它們缺乏一種原則性的方式來決定何時先前的解決方案仍然有效,何時必須進行調整,以及何時應該被丟棄。我們介紹了 CORTEX(上下文協調與任務經驗重用),這是一個通用的 AI 系統框架,通過一層經過驗證的經驗將專門的代理連接起來。每個事件記錄其任務條件、來源和工具狀態、決定性謂詞、證明痕跡、驗證者和結果。一個元控制器選擇精確重播、檢查調整、新的綜合或升級。接受的事件可以通過挑戰驅動的開發循環轉變為任務模式和程序策略。這為系統提供了一個隱含的能力層,能夠在不改變模型權重的情況下增長。我們為精確重播和來源版本分離形式化了系統合同,並推導出何時重用可以節省計算。受控的雙域實施在 1,000 個合成案例上測試精確重播核心。完整家庭保留測試在八個臨床和政策拆分中對 1,000 個新家庭案例的程序轉移,具有完整的新證據基礎和對無關領域及插入順序擾動的完美不變性。轉移痕跡揭示了執行經過驗證的策略所需的工作。這些結果為通過可重用程序、類型化經驗和發展轉移建立了一條通向通用智能的初步路徑。
FOCUS: Benchmarking Retinal Model Generalization from Foundation Vision Encoders to Multimodal LLMs
2609.33158v1 by David Restrepo, Chenwei Wu, Luis Filipe Nakayama, Miguel L. Martins, Stergios Christodoulidis, Maria Vakalopoulou, Enzo Ferrante
Progress in AI-based retinal image analysis has advanced with foundation models, yet evaluating their reliability remains challenging. Performance reported on a single dataset does not capture how models behave under dataset shift, across clinical definitions, or for different patient subgroups. This limitation is particularly critical in medical imaging analysis, where robustness, calibration, and fairness are essential for safe deployment. We introduce FOCUS (Foundation Ophthalmic Cross-Dataset Understanding under Shift), a cross-dataset benchmark for evaluating retinal fundus models that considers vision-only encoder models (VM), vision-language dual-encoder models (VLM), and multimodal large language models (MLLM). FOCUS harmonizes binary diabetic retinopathy, referable diabetic retinopathy, and glaucomatous optic neuropathy tasks across ten public datasets spanning diverse geographies, acquisition conditions, and label protocols. The benchmark evaluates models through a unified analysis layer that measures ranking performance, calibration, subgroup disparities, and image-quality robustness. We present a large-scale evaluation covering 532 base configurations and 228 MLLM configurations adapted through supervised fine-tuning with low-rank adaptation (LoRA). Results show that no model family consistently dominates across tasks and datasets: general VM encoders achieve the strongest average ranking performance, medical MLLMs are competitive but variable, and dual encoder VLMs benefit substantially from lightweight adaptation. Fine-tuning improves in-domain performance but exhibits heterogeneous transfer to external datasets, particularly in calibration. These findings demonstrate that retinal model evaluation is inherently multidimensional. FOCUS provides a practical framework and public benchmark to assess generalization, reliability, and robustness beyond single-dataset leaderboards
摘要:進展於基於人工智慧的視網膜影像分析已隨著基礎模型的發展而提升,然而評估其可靠性仍然具有挑戰性。單一數據集上報告的性能無法捕捉模型在數據集轉移、臨床定義之間或不同患者子群體中的行為。這一限制在醫學影像分析中特別關鍵,因為穩健性、校準和公平性對於安全部署至關重要。我們引入了FOCUS(Foundation Ophthalmic Cross-Dataset Understanding under Shift),這是一個跨數據集基準,用於評估視網膜眼底模型,考慮了僅視覺編碼器模型(VM)、視覺-語言雙編碼器模型(VLM)和多模態大型語言模型(MLLM)。FOCUS在十個公共數據集上協調二元糖尿病視網膜病變、可參考糖尿病視網膜病變和青光眼性視神經病變任務,這些數據集涵蓋了多樣的地理位置、獲取條件和標籤協議。該基準通過一個統一的分析層評估模型,測量排名性能、校準、子群體差異和影像質量的穩健性。我們呈現了一個涵蓋532個基本配置和228個經過低秩適應(LoRA)監督微調的MLLM配置的大規模評估。結果顯示,沒有任何模型家族在任務和數據集上始終佔據主導地位:一般的VM編碼器實現了最強的平均排名性能,醫學MLLM在競爭中但變化不定,而雙編碼器VLM在輕量適應中受益匪淺。微調改善了內域性能,但在外部數據集上展現出異質的轉移,特別是在校準方面。這些發現表明,視網膜模型評估本質上是多維的。FOCUS提供了一個實用的框架和公共基準,以評估超越單一數據集排行榜的泛化、可靠性和穩健性。
Relative Generalization Invariance of LLM Pretraining
2609.33016v1 by Fengzhuo Zhang, Shuche Wang, Shenggui Li, Tianyu Ruan, Jianliang He, Ivor Tsang, Tianyu Pang, Chao Du, Tianwei Zhang, Zhuoran Yang
Large Language Model (LLM) pretraining performance is jointly shaped by three components of the training triplet: the optimizer, model architecture, and training data stream. However, how these components influence performance in distinct ways remains unclear. We take a first step toward isolating their effects by studying relative generalization. We introduce Relative Generalization Invariance (RGI), the invariance of the validation-loss difference between any two tokens across models. We show that RGI approximately holds across a wide range of optimizers and moderate architectural variations, suggesting that these choices induce an approximately uniform shift in token-wise losses. In contrast, changing the training data stream can substantially alter relative generalization. We further show that RGI cannot be explained by the neural tangent kernel or mean-field regimes alone and prove that it can emerge in an overparameterized quadratic model. Overall, our work identifies RGI as a new phenomenon in LLM pretraining that helps distinguish the effects of optimizers and architectures from those of training data.
摘要:大型語言模型(LLM)預訓練的性能是由訓練三元組的三個組件共同影響的:優化器、模型架構和訓練數據流。
然而,這些組件如何以不同方式影響性能仍然不清楚。
我們邁出了第一步,通過研究相對泛化來隔離它們的影響。
我們引入了相對泛化不變性(RGI),即在不同模型之間任何兩個標記的驗證損失差異的不變性。
我們展示了RGI在廣泛的優化器和適度的架構變化中大致成立,這表明這些選擇會在標記損失上引起大致均勻的變化。
相反,改變訓練數據流可以顯著改變相對泛化。
我們進一步表明,RGI不能僅僅通過神經切線核或均值場範疇來解釋,並證明它可以在過參數化的二次模型中出現。
總體而言,我們的工作將RGI確定為LLM預訓練中的一種新現象,幫助區分優化器和架構的影響與訓練數據的影響。
DynamicDx: Evaluating Evidence Acquisition in Video-Based Diagnosis
2609.32957v1 by Jiahui Li, Yutong Guo, Nan Yang, Wenzhan Song, Jin Lu, Fei Dou
Diagnosing a patient from video requires more than recognizing the sign: a vision-language model must turn what it sees into hypotheses, questions and tests. DynamicDx evaluates each step in 71 neurological consultations across 11 sign categories, linking authentic patient videos to confirmed diagnoses and fixed charts built from the same case reports, so that every model queries the same evidence. Across five such models, video improves accuracy by 9.9-22.5 percentage points over blind input, but neither recognition alone nor temporal order explains the gain: the cause is usually missing from the model's video-only differential diagnosis even when the sign is recognized, and shuffling the frames produces no reliable accuracy loss. Instead, a trajectory replay traces most of the gain to the investigation results the video prompts. Evidence acquisition is the bottleneck: supplying the decisive investigations raises accuracy to 73.2-93.0%. Two interventions act on it. A post-trained 4B video describer improves sign descriptions, especially from a short, densely sampled segment, and source-clean literature retrieval expands initial hypotheses; both bring the tests a model orders closer to those the treating clinicians documented and, through them, raise accuracy. For video-based diagnosis, seeing better helps when it leads to asking better.
摘要:診斷患者的視頻需要的不僅僅是識別標誌:視覺-語言模型必須將其所見轉化為假設、問題和測試。DynamicDx 評估了 71 次神經諮詢中的每一步,涵蓋 11 種標誌類別,將真實患者視頻與確認的診斷和基於相同案例報告製作的固定圖表聯繫起來,以便每個模型查詢相同的證據。在這五個模型中,視頻的準確性比盲輸入提高了 9.9-22.5 個百分點,但僅僅依賴識別或時間順序並不能解釋這一增益:即使標誌被識別,模型的視頻僅差異診斷中通常缺少原因,並且打亂幀並不會產生可靠的準確性損失。相反,軌跡重播將大部分增益追溯到視頻促進的調查結果。證據獲取是瓶頸:提供關鍵調查將準確性提高到 73.2-93.0%。有兩個干預措施對此產生影響。一個經過後訓練的 4B 視頻描述器改善了標誌描述,特別是來自短而密集取樣段的描述,而來源清理文獻檢索擴展了初步假設;兩者都使模型所訂購的測試更接近治療臨床醫生記錄的測試,並通過它們提高準確性。對於基於視頻的診斷,當更好的視覺導致更好的提問時,看到更清楚是有幫助的。
Counterfactual Self-Evolving Agents for Evidence-Grounded Reasoning
2609.32870v1 by Xing Han, Yuxin Wang, Chen Chen, Wei Dai, Gautham Krishna Gudur, Shijun Li, Hsing-Huan Chung, Gregory D. Hager, Joydeep Ghosh, Paul Pu Liang, Suchi Saria
Self-play proposer--solver methods improve reasoning by generating tasks and learning from verified solutions. However, for evidence-identifiable tasks, where case-specific evidence and domain knowledge determine a checkable answer, self-play requires generating plausible cases whose answers can be independently verified. We introduce counterfactual self-evolution, which generates counterfactual context for reconsidering the original case. A trainable Proposer constructs targeted evidence edits and describes potential outcome changes with causal explanations. We handcraft an expert-verified counterfactual instruction-tuning dataset to teach the Proposer to generate high-quality counterfactuals across a broad range of action--outcome scenarios. Each counterfactual instruction-tuning example specifies an edit within a defined category and explains its hypothesized causal effect on the decision, teaching the Proposer to reason systematically about what changes and why. We instruction-tune the Proposer on these examples, then formulate a fine-tuning reward that integrates feedback from the Solver and Verifier. Across diverse counterfactual scenarios, this reward favors high-quality counterfactuals and warranted revisions, while penalizing changes that overturn correct decisions. The counterfactual context aims to correct errors and strengthen confidence in correct decisions. Accepted counterfactuals accumulate in memory that supplies in-context evidence to the frozen Solver; the Solver adapts through evolving context rather than weight updates. We apply the framework to clinical reasoning, fact verification, and business reasoning. Our evaluation tracks performance over successive rounds as counterfactual memory grows, including transfer to harder cases. Our method achieves superior results across diverse frontier models.
摘要:自我對弈提議者--解決者方法通過生成任務並從經過驗證的解決方案中學習來改善推理。然而,對於可識別證據的任務,其中案例特定的證據和領域知識決定了可檢查的答案,自我對弈需要生成可以獨立驗證的合理案例及其答案。我們引入了反事實自我演化,該方法生成反事實背景以重新考慮原始案例。一個可訓練的提議者構建針對性的證據編輯,並用因果解釋描述潛在的結果變化。我們精心製作了一個專家驗證的反事實指令調整數據集,以教導提議者在廣泛的行動--結果場景中生成高質量的反事實。每個反事實指令調整示例指定了一個在定義類別內的編輯,並解釋其對決策的假設因果效應,教導提議者系統性地推理什麼變化以及為什麼變化。我們在這些示例上對提議者進行指令調整,然後制定一個微調獎勵,該獎勵整合了解決者和驗證者的反饋。在多樣的反事實場景中,這個獎勵偏好高質量的反事實和合理的修訂,同時懲罰推翻正確決策的變更。反事實背景旨在糾正錯誤並增強對正確決策的信心。被接受的反事實在記憶中累積,為凍結的解決者提供上下文證據;解決者通過演變的背景而不是權重更新來適應。我們將該框架應用於臨床推理、事實驗證和商業推理。我們的評估跟踪隨著反事實記憶增長而進行的多輪性能,包括轉移到更困難的案例。我們的方法在多樣的前沿模型中取得了優越的結果。
FinancialAuditBench: Benchmark Construction under Differential Privacy Using Real-World Priors
2609.32835v1 by Jerry Huang, Sarvesh Babu, Matt Van Buren, Alexander Wang, Pranav Pillai, Arush Jain, James P. Burton, Julia Hockenmaier
As AI agents are becoming widely adopted in the financial services industry, careful measurement is essential to understand where they can be reliably deployed and where oversight and professional review remain necessary. Such measurement, however, is constrained by limited access to proprietary or privacy-sensitive data. Existing benchmarks therefore often rely on publicly available data, human- and/or LLM-authored tasks, or simplified settings. We introduce FinancialAuditBench, a benchmark for evaluating agents on financial statement audit tasks, along with a framework for systematically generating synthetic engagements. Our task generation framework leverages differentially private aggregate statistics from historical audits along with audit expertise contributed through over 1,100 hours of benchmark development and review. FinancialAuditBench consists of 90 tasks spanning workpaper completion and review across six synthetic audit engagements, each containing an average of 179 files. Evaluation on eleven frontier models shows that while agents complete substantial portions of staff-level audit tasks well, they sometimes perform inappropriate procedures or produce incorrect documentation. Beyond financial auditing, our framework offers an approach for systematically generating synthetic tasks for model evaluation and training in privacy-sensitive domains.
摘要:隨著 AI 代理在金融服務行業的廣泛採用,仔細的測量對於理解它們可以可靠部署的地方以及何處仍需監督和專業審查至關重要。然後,這種測量受到對專有或隱私敏感數據的有限訪問的限制。因此,現有的基準通常依賴於公開可用數據、人類和/或 LLM 編寫的任務或簡化的設置。我們介紹了 FinancialAuditBench,一個用於評估代理在財務報表審計任務上的基準,以及一個系統生成合成參與的框架。我們的任務生成框架利用了來自歷史審計的差分隱私聚合統計數據,以及通過超過 1,100 小時的基準開發和審查貢獻的審計專業知識。FinancialAuditBench 包含 90 個任務,涵蓋六個合成審計參與的工作文件完成和審查,每個參與平均包含 179 個文件。對十一個前沿模型的評估顯示,儘管代理能夠很好地完成大量的員工級審計任務,但有時它們會執行不當的程序或產生不正確的文件。除了財務審計,我們的框架還提供了一種系統生成合成任務的方法,用於在隱私敏感領域進行模型評估和訓練。
Mandela-Bench: Multimodal Models Remember Canonical Images Instead of Seeing Them
2609.32763v1 by Yicheng Bao, Zhenkun Gao, Xiahui Guo, Mingqian Yang, Xueheng Li, Bangwei Liu, Mingang Chen, Lijun Li, Xuhong Wang, Xin Tan
Historical photographs and other canonical images can now be edited seamlessly with a single instruction, often leaving no reliable pixel-level trace. In such cases, the only evidence of manipulation may be a fact about what the image depicts. Existing benchmarks instead rely on generator artefacts, image-caption inconsistencies, visual implausibilities, or external references, and therefore do not test whether a model can use its own world knowledge to verify a recognized image. We introduce Mandela-Bench, containing 1,507 edits of canonical images: 1,359 knowledge-only forgeries, each contradicting one verifiable fact, and 148 anchor-free controls that preserve the editing process without introducing a factual contradiction, together with 474 untouched originals. We score not only whether a model detects a forgery, but whether its explanation identifies the inserted entity or the fact being violated. Across 36 multimodal models, from 0.8B parameters to frontier scale, we find a consistent failure mode. When a public figure is removed from a familiar photograph, models still name that person in up to 72.7% of responses. Some models can distinguish the replacement face from the original when shown in isolation, yet still judge the full edited photograph as authentic. Providing the true event and date does not improve knowledge-grounded detection, whereas providing the same information after cropping away the recognizable composition does. Even under explicit verification prompts, only one of the 36 models meets the KGR criterion on at least half of the forged images. These results suggest that the failures cannot be explained by missing knowledge or inadequate perception alone. Instead, they are consistent with recognition biasing verification toward the remembered canonical image rather than the observed edit.
摘要:歷史照片和其他經典圖像現在可以通過單一指令無縫編輯,通常不留下可靠的像素級痕跡。在這種情況下,唯一的操控證據可能是圖像所描繪的事實。現有的基準測試則依賴於生成器產物、圖像標題不一致、視覺不合理性或外部參考,因此並未測試模型是否能夠利用自身的世界知識來驗證已識別的圖像。我們引入了 Mandela-Bench,包含 1,507 個經典圖像的編輯:1,359 個僅知識的偽造,每個都與一個可驗證的事實相矛盾,以及 148 個無錨控件,它們保留了編輯過程而不引入事實矛盾,還有 474 個未觸碰的原始圖像。我們不僅評分模型是否檢測到偽造,還評分其解釋是否識別出插入的實體或被違反的事實。在 36 個多模態模型中,從 0.8B 參數到前沿規模,我們發現了一種一致的失敗模式。當公共人物從熟悉的照片中移除時,模型仍在高達 72.7% 的回應中提到該人。一些模型在單獨顯示替換臉時可以區分與原始臉的不同,但仍然判斷整張編輯過的照片為真實。提供真實事件和日期並未改善基於知識的檢測,而在裁剪掉可識別構圖後提供相同的信息則有改善。即使在明確的驗證提示下,36 個模型中只有一個在至少一半的偽造圖像上達到 KGR 標準。這些結果表明,失敗無法僅用知識缺失或感知不足來解釋。相反,它們與識別偏見將驗證偏向於記憶中的經典圖像而非觀察到的編輯是一致的。
What Would Falsify It? A Variable Specific Evidence Standard for Mechanistic Claims About Self Explanation
2609.32670v1 by Arshia Eftekhari zadeh
When a language model explains an answer it has already given, does it reuse the computation that produced the answer or reconstruct a story from the answer alone? Attribution, transportability and recoverability are each compatible with causal use without establishing it. We propose an evidence standard: pair each positive statistic with a variable specific null that removes the tested variable's identity while matching relevant nuisance dimensions as far as possible, and audit unmatched dimensions. We apply this standard to a known cause. A cue naming a wrong option raises the rate of choosing that option by 64 to 68 percentage points across three models. Explanations mention the cue in 1.8 percent of items or fewer in three of four models tested. Three estimator classes yield favorable statistics, but none establishes causal sensitivity to the cue contrast under its own control in the three-model analysis. In the strongest case, a recovered cue direction reaches $R^2$ of 0.95 and exceeds a geometry matched random direction in all three seeds, while a direction fitted by the same pipeline with cue labels scrambled reproduces 61 to 76 percent of its effect at comparable realized edit magnitude. A fourth model passes one interchange endpoint, but unequal edit magnitudes and a contrast that changes both cue identity and cue-answer agreement limit its interpretation. These experiments leave causal access unresolved. They establish an evidentiary requirement: favorable mechanistic statistics must survive controls for variable identity and nuisance structure. Reusable controls separate generic from identity specific transport effects, fit null directions with scrambled labels, and audit realized intervention magnitudes.
摘要:當一個語言模型解釋它已經給出的答案時,它是重用產生該答案的計算,還是僅僅從答案重建一個故事?歸因、可轉移性和可恢復性在不建立因果關係的情況下各自與因果使用相容。我們提出了一個證據標準:將每個正向統計數據與一個特定於變數的零假設配對,該零假設在盡可能匹配相關的干擾維度的同時去除被測變數的身份,並審核未匹配的維度。
我們將這一標準應用於一個已知的原因。命名錯誤選項的提示使得選擇該選項的比率在三個模型中提高了64到68個百分點。在四個測試的模型中,解釋中提到提示的比例在1.8%或更少。三個估計器類別產生了有利的統計數據,但在三模型分析中,沒有一個在其自身控制下建立對提示對比的因果敏感性。在最強的情況下,恢復的提示方向達到$R^2$為0.95,並在所有三個隨機種子中超過幾何匹配的隨機方向,而由同一管道擬合的提示標籤被打亂的方向在可比較的實現編輯幅度下重現了61到76%的效果。一個第四模型通過了一個互換端點,但不等的編輯幅度以及一個同時改變提示身份和提示-答案一致性的對比限制了其解釋。
這些實驗使因果訪問未得到解決。它們建立了一個證據要求:有利的機制統計必須在變數身份和干擾結構的控制下存活。可重用的控制分離一般的與身份特定的傳輸效應,擬合帶有打亂標籤的零方向,並審核實現的干預幅度。
When Users Change Their Minds: Measuring and Repairing Intent Drift in LLM Agents
2609.32520v1 by Yanjie Zhang, Bowen Cao, Zixin Chen, Yushi Sun
LLM agents often operate over multi-turn interactions in which user intent changes before execution. We study intent drift: the failure mode in which superseded parts of the user's intent continue to influence the final answer or tool action. We introduce IntentFlux, an executable benchmark that converts verifiable tasks into dialogues with controlled intent changes while preserving their original graders. In a 627-case calibration, mean task score falls from 0.476 to 0.384 as dialogues contain more superseded and withdrawn information. Across eight models, the rate of fully correct solutions is significantly lower when the same final task must be recovered from an evolving dialogue rather than given directly in a single turn. We further introduce StateForge, which explicitly maintains the active requirements before generation. On General-Test, it improves mean task score from 0.367 to 0.467. Providing the ground-truth final state improves performance further but still does not recover single-turn performance, indicating that state-estimation errors explain only part of the gap. These results establish intent drift as a measurable multi-turn failure mode and explicit state maintenance as a partial mitigation.
摘要:LLM 代理通常在多輪互動中運作,其中使用者的意圖在執行之前會發生變化。
我們研究意圖漂移:這是一種失敗模式,其中被取代的使用者意圖部分繼續影響最終答案或工具行動。
我們引入了 IntentFlux,一個可執行的基準,將可驗證的任務轉換為具有控制意圖變化的對話,同時保留其原始評分者。
在627個案例的校準中,當對話包含更多被取代和撤回的信息時,平均任務分數從0.476降至0.384。
在八個模型中,當必須從一個不斷演變的對話中恢復相同的最終任務,而不是直接在單輪中給出時,完全正確解決方案的比率顯著降低。
我們進一步引入了 StateForge,它在生成之前明確維護活動要求。
在 General-Test 上,它將平均任務分數從0.367提高到0.467。
提供真實的最終狀態進一步改善了性能,但仍然無法恢復單輪性能,這表明狀態估計錯誤僅解釋了部分差距。
這些結果確立了意圖漂移作為可測量的多輪失敗模式,以及明確的狀態維護作為部分緩解措施。
Explaining Textual Entailment with Lexical Entailments: Using LLMs to Supply Lexical Relations for Formal Proofs
2609.32491v1 by Jorryt de Jong, Stefan Moraca, Ettore Cesari, Lasha Abzianidze
Large Language Models (LLMs) are highly capable of natural language reasoning and appear to store a great deal of lexical knowledge, but it is still unclear how much of this knowledge they actually use when reasoning, and whether they use it in the right way. On the other hand, logic-based Natural Language Inference (NLI) systems provide transparent and formally grounded reasoning, but they need to be supplied with rich lexical knowledge to prove inferences beyond purely logical ones. In this paper, we evaluate whether LLMs can identify all lexical knowledge needed to solve NLI problems and how much this knowledge contributes to proof search in a logic-based NLI system. Our research focuses exclusively on structured lexical entailments (e.g., chinchilla$\sqsubseteq$small animal) as a proxy for structured explanations for NLI problems with an entailment label. First, we curate a dataset for a new task of explaining sentential entailments with a set of lexical entailments. The dataset is used to intrinsically evaluate LLMs on generating structured lexical explanations. Then, we use NLI as an extrinsic evaluation in a simple neuro-symbolic setting, assessing whether LLMs can supply sufficient lexical relations to LangPro, a natural-logic theorem prover for natural language. The results show that the proposed task remains challenging even for hosted proprietary LLMs, and that their contribution to theorem proving is moderate: generated relations are often only partially sound and may be tailored to the specific NLI problem rather than representing generally valid lexical knowledge.
摘要:大型語言模型 (LLMs) 在自然語言推理方面具有很高的能力,並且似乎儲存了大量的詞彙知識,但目前仍不清楚它們在推理時實際使用了多少這些知識,以及是否以正確的方式使用。另一方面,基於邏輯的自然語言推理 (NLI) 系統提供透明且有正式基礎的推理,但它們需要提供豐富的詞彙知識,以證明超越純邏輯的推論。在本文中,我們評估 LLMs 是否能夠識別解決 NLI 問題所需的所有詞彙知識,以及這些知識對基於邏輯的 NLI 系統中的證明搜索的貢獻程度。我們的研究專注於結構化詞彙推論(例如,chinchilla$\sqsubseteq$small animal),作為具有推論標籤的 NLI 問題的結構化解釋的代理。首先,我們為解釋句子推論的新任務策劃了一個數據集,該數據集包含一組詞彙推論。該數據集用於對 LLMs 在生成結構化詞彙解釋方面進行內部評估。然後,我們在一個簡單的神經符號設置中使用 NLI 作為外部評估,評估 LLMs 是否能夠為 LangPro 提供足夠的詞彙關係,LangPro 是一個用於自然語言的自然邏輯定理證明器。結果顯示,即使對於託管的專有 LLMs,所提出的任務仍然具有挑戰性,並且它們對定理證明的貢獻是適度的:生成的關係往往僅部分有效,並且可能針對特定的 NLI 問題進行調整,而不是代表普遍有效的詞彙知識。
Superposed Inference for Hyperdimensional Computing
2609.32320v1 by Quanling Zhao, Nilesh Prasad Pandey, Ye Tian, Tajana Rosing
Hyperdimensional computing (HDC) is attractive for efficient and robust learning, but conventional inference still encodes every query independently, repeatedly paying the cost of high-dimensional projection. We introduce SupHDC, a new inference paradigm that processes multiple queries through a shared encoding computation. SupHDC assigns lightweight random slot keys, superposes the keyed queries before encoding, and uses slot-specific classifiers to recover their individual predictions. A random-feature kernel view explains why exact recovery of each hypervector is unnecessary: inference only needs to preserve the class evidence that determines the prediction. Across ten datasets, SupHDC achieves 1.39x analytical speedup with no average accuracy loss, and up to 2.08x speedup with only a 2.67 percentage-point mean accuracy loss. On a Raspberry Pi~5, it delivers 2.01x measured wall-clock speedup with a 2.26 percentage-point loss in mean prediction accuracy. SupHDC shows that high-dimensional redundancy can be used not only for robustness, but also as capacity for shared inference.
摘要:超維計算(HDC)因其高效和穩健的學習而受到青睞,但傳統推理仍然獨立編碼每個查詢,重複支付高維投影的成本。我們介紹了SupHDC,一種通過共享編碼計算處理多個查詢的新推理範式。SupHDC分配輕量級隨機槽鍵,在編碼之前對鍵入的查詢進行疊加,並使用槽特定的分類器來恢復它們的個別預測。一個隨機特徵核視角解釋了為什麼不需要精確恢復每個超向量:推理只需要保留決定預測的類別證據。在十個數據集上,SupHDC實現了1.39倍的分析加速,且沒有平均準確度損失,並且在僅有2.67個百分點的平均準確度損失的情況下,最高可達2.08倍的加速。在Raspberry Pi~5上,它實現了2.01倍的測量牆時計加速,並且平均預測準確度損失為2.26個百分點。SupHDC顯示高維冗餘不僅可以用於穩健性,還可以作為共享推理的容量。
Why Directly Learning Periodic Trajectories Can Fail
2609.32254v1 by Kaixin Zheng, Anita Layton
Operator learning of periodic solutions requires deciding how simulation data should be recorded and represented. A natural choice is to integrate long enough for transients to decay and record a window wide enough to contain at least one full period of all trajectories. We find that these conservative choices can make the resulting trajectories difficult to learn, even when the underlying periodic orbits vary regularly with system parameters. Unaligned trajectories generalize poorly even within the training distribution. Phase alignment substantially improves in-distribution generalization, but models trained on a fixed physical-time window still have large errors on trajectories with periods outside the training range. We explain both failures through a common mechanism: frequency differences accumulate over time, so the target phase varies rapidly with the parameters. Predictors that cannot track this variation incur a population MSE floor in both settings; for fixed window prediction, we also derive a per-sample lower bound. We then study one of the simplest representations that escape these floors: learning an aligned, normalized waveform and its period separately. We establish regularity of the decoupled targets under ODE assumptions and show experimentally that this approach avoids both failures in ODE systems and a PDE case study.
摘要:操作學習週期解需要決定如何記錄和表示模擬數據。一個自然的選擇是整合足夠長的時間以使瞬態衰減,並記錄一個足夠寬的窗口以包含所有軌跡的至少一個完整週期。我們發現,這些保守的選擇會使得結果軌跡難以學習,即使基礎的週期軌道隨著系統參數規則變化。未對齊的軌跡即使在訓練分佈內也會泛化不佳。相位對齊顯著改善了分佈內的泛化,但在固定物理時間窗口上訓練的模型在週期超出訓練範圍的軌跡上仍然有較大的誤差。我們通過一個共同機制解釋這兩種失敗:頻率差異隨著時間累積,因此目標相位隨著參數快速變化。無法追蹤這種變化的預測器在這兩種情況下都會產生一個群體均方誤差下限;對於固定窗口預測,我們還推導出每個樣本的下限。我們接著研究一種逃避這些下限的最簡單表示之一:分別學習對齊的、歸一化的波形及其週期。我們在常微分方程假設下建立了解耦目標的規律性,並實驗表明這種方法避免了常微分方程系統中的兩種失敗以及一個偏微分方程案例研究。
A bilingual AI audiologist built through rubric-guided playbook induction outperforms human audiologists in a blinded evaluation of simulated cases
2609.32220v1 by Linkai Li, Changgeng Mo, Hanlin Yu, Congxi Lu, Shangqiguo Wang, Matthew B Fitzgerald, Shan X Wang
Audiology consultation requires structured history-taking, audiometric interpretation and patient-centred communication, yet real-world case material is scarce. We present a bilingual AI audiologist pairing a general-purpose large language model with rubric-guided playbook induction, multimodal audiogram interpretation and retrieval-augmented grounding, without fine-tuning the language-model backbone. Using a 21-item rubric and an AI patient simulator, we induced a 19-rule consultation policy from 73 training cases (43 English, 30 Chinese) and evaluated the system on 58 independent simulated cases (30 Chinese, 28 English) in a pre-specified, source-blinded comparison with 17 practising audiologists. The AI audiologist outperformed human audiologists on every case (58/58; mean paired $Δ$ = +1.35 on a 5-point composite, Cohen's d = 1.84, $P = 4.5 \times 10^{-20}$), on 20 of 21 rubric items and in both languages. Component ablation identified the playbook as the largest contributor, offering a practical route to specialist consultation agents in low-data medical domains.
摘要:聽力學諮詢需要結構化的病史採集、聽力測試解釋和以病人為中心的溝通,但現實世界中的案例材料卻稀缺。我們提出了一個雙語AI聽力學家,將通用的大型語言模型與指導性評分標準的劇本引導、多模態聽力圖解釋和檢索增強的基礎相結合,且不對語言模型的主幹進行微調。使用一個包含21項的評分標準和一個AI病人模擬器,我們從73個訓練案例(43個英文,30個中文)中引導出19條諮詢政策,並在與17名執業聽力學家的預先指定、來源盲測比較中,對58個獨立的模擬案例(30個中文,28個英文)進行了評估。AI聽力學家在每個案例中均超越了人類聽力學家(58/58;平均配對$Δ$ = +1.35,基於5分的綜合評分,Cohen's d = 1.84,$P = 4.5 \times 10^{-20}$),在21項評分標準中的20項以及兩種語言中均表現優異。組件消融識別出劇本是最大的貢獻者,為低數據醫療領域中的專家諮詢代理提供了一條實用的途徑。
Evaluating Single and Multi-Omics Based Explainable Artificial Intelligence (MOXAI) for Molecular Subclass Classification of Adult-Type Diffuse Gliomas
2609.32190v1 by Md Zahangir Alom, Quynh T. Tran, Breuer Alexandar, Brent A. Orr
DNA methylation (DNAM) profiling has emerged as a powerful diagnostic tool for classifying brain and solid tumors. However, existing computational models typically analyze methylation and copy number variation (CNV) data separately, failing to capture the complementary information their integration could provide. Moreover, current classification models lack mechanisms for within-class risk assessment analogous to traditional tumor grading, and no established explainability method can attribute classification decisions to specific genomic loci. In this paper, we present MOXAI (Multi-Omics Based Explainable AI), a deep learning framework that integrates DNA methylation and copy number data from methylation arrays to classify molecular subtypes of adult-type diffuse gliomas, alongside single-modality variants for comparison. Using a cohort from The Cancer Genome Atlas (TCGA), we trained ResNet50, DINOv2, and Graph Attention Network (GAT) models on methylation data alone, copy number data alone, and combined multimodal data. We further developed explainable AI (XAI) methods based on class activation maps (CAMs) and gradient-weighted CAM (Grad-CAM) to identify the specific CpG sites, genes, and chromosomal regions most relevant to each classification decision. The multimodal model achieved up to 92.98% cross-validation accuracy, outperforming models trained on CNV data alone. DINOv2 showed the strongest generalization, reaching 94.25% accuracy (confidence >0.9) on independent validation sets. XAI results aligned with established molecular features of adult-type diffuse glioma subtypes, confirming the biological interpretability of the framework.
摘要:DNA 甲基化 (DNAM) 檔案已成為分類腦部和實體腫瘤的強大診斷工具。
然而,現有的計算模型通常分別分析甲基化和拷貝數變異 (CNV) 數據,未能捕捉其整合所能提供的互補信息。
此外,當前的分類模型缺乏類內風險評估機制,類似於傳統腫瘤分級,且沒有建立的可解釋性方法能將分類決策歸因於特定的基因組位點。
在本文中,我們提出了 MOXAI (基於多組學的可解釋 AI),這是一個深度學習框架,整合了來自甲基化陣列的 DNA 甲基化和拷貝數據,以分類成人型擴散性膠質瘤的分子亞型,並提供單一模態變體以供比較。
使用來自癌症基因組圖譜 (TCGA) 的一個隊列,我們僅在甲基化數據、僅在拷貝數據及結合多模態數據上訓練了 ResNet50、DINOv2 和圖注意網絡 (GAT) 模型。
我們進一步開發了基於類激活圖 (CAMs) 和梯度加權 CAM (Grad-CAM) 的可解釋 AI (XAI) 方法,以識別與每個分類決策最相關的特定 CpG 位點、基因和染色體區域。
多模態模型達到了高達 92.98% 的交叉驗證準確率,超越了僅在 CNV 數據上訓練的模型。
DINOv2 展現出最強的泛化能力,在獨立驗證集上達到了 94.25% 的準確率 (信心 >0.9)。
XAI 結果與已建立的成人型擴散性膠質瘤亞型的分子特徵一致,確認了該框架的生物學可解釋性。
REALM: Regime-Switching, Explainable, and Activation-Induced Linear Models
2609.32141v1 by Xiaoran Cheng, Sen Na, Jia Li
Deep ReLU networks are piecewise-affine mappings that partition the input space into cells, each characterized by a distinct activation pattern. This structure motivates fitting a local linear model within each cell to preserve predictive accuracy while improving interpretability. The challenge is to identify regimes that are stable, data-adaptive, and easy to explain. We propose REALM, a mixture of linear models whose regimes are induced by neural activation patterns. Because the number of activation cells in a deep neural network (DNN) can grow rapidly with depth, we first distill a deep teacher into a wide, shallow student network (WSSN), then binarize and cluster its hidden-layer activations to define the regimes and fit a linear model within each regime. Since the regimes are discovered from internal structure, the router does not carry the predictive burden. To make regime assignment interpretable, we train a multiclass logistic regression, the explanatory gate, to reproduce the regime assignments. The two-level structure is interpretable at both stages in terms of raw tabular or learned convolutional features: the gate identifies features that determine regime assignments, while the linear models identify features that drive predictions within each regime. We analyze an idealized setting that illustrates a trade-off between partition complexity and stability: as the number of regimes grows, finer partitions can improve approximation but may reduce regime-assignment stability. Experiments on tabular and image datasets show that REALM achieves competitive predictive performance relative to other DNN-guided mixture surrogates and inherently interpretable models while producing stable regime-level explanations.
摘要:深度 ReLU 網絡是分段仿射映射,將輸入空間劃分為各個單元,每個單元都有其獨特的激活模式。這種結構促使我們在每個單元內擬合一個局部線性模型,以保持預測準確性同時提高可解釋性。挑戰在於識別穩定、數據自適應且易於解釋的狀態。我們提出了 REALM,一種由神經激活模式引導的線性模型混合。由於深度神經網絡 (DNN) 中的激活單元數量可能隨著深度迅速增長,我們首先將深層教師網絡提煉成一個寬而淺的學生網絡 (WSSN),然後對其隱藏層激活進行二值化和聚類,以定義狀態並在每個狀態內擬合線性模型。由於這些狀態是從內部結構中發現的,因此路由器不承擔預測負擔。為了使狀態分配可解釋,我們訓練了一個多類別邏輯回歸模型,即解釋閘,以重現狀態分配。這種兩級結構在原始表格或學習的卷積特徵方面在兩個階段都是可解釋的:閘識別決定狀態分配的特徵,而線性模型則識別在每個狀態內驅動預測的特徵。我們分析了一個理想化的設置,展示了劃分複雜性和穩定性之間的權衡:隨著狀態數量的增加,更細的劃分可以改善近似,但可能會降低狀態分配的穩定性。在表格和圖像數據集上的實驗顯示,REALM 相較於其他 DNN 引導的混合代理和內在可解釋的模型,實現了具有競爭力的預測性能,同時產生穩定的狀態級解釋。
Reasoning Concentrates Errors, and Self-Consistency Never Notices
2609.32035v1 by Asaad Althoubi
Self-consistency assumes that independent samples disagree when a model is unsure, so agreement is evidence of correctness. Holding weights fixed and toggling only a reasoning mode, over five benchmarks and 74,944 samples, we show that reasoning concentrates a model's errors: the probability that two independently drawn wrong answers coincide rises in all ten dataset-scale comparisons (p = 0.00098), and in nine of nine after restricting both arms to the problems each gets wrong. Where the answer space is unbounded, reasoning cuts the distinct answers produced to 0.43-0.65 of the non-reasoning count; where it is bounded, both arms hold an identical option set and reasoning concentrates mass on it instead, which no positional prior can explain at fixed weights. The aggregate cost is smaller than the mechanism predicts, because reasoning also shrinks the set of problems where answer diversity can decide anything, in ten of ten cells and by 2.7x; normalized for available headroom, both arms convert a quarter of it in domain. Confidence weighting does not recover what is left. Across 280 method-dataset-model combinations on eight models and five benchmarks, not one beats plain majority voting after correction; weighted voting agrees with it on 98.5% of problem-method pairs and is right 56.3% of the time on the rest; and a signal's direction can invert within fixed weights, with answer log-probability predicting correctness when reasoning is off and error when it is on. A learned six-signal combination gains nothing out of domain. Confidence signals should be evaluated on decisions, not on discrimination.
摘要:自我一致性假設當模型不確定時,獨立樣本會出現不一致,因此一致性是正確性的證據。固定權重並僅切換推理模式,在五個基準和74,944個樣本中,我們顯示推理集中了一個模型的錯誤:兩個獨立抽取的錯誤答案重合的概率在所有十個數據集規模的比較中上升(p = 0.00098),在將兩個臂限制於各自錯誤的問題後,九個中有九個也如此。當答案空間是無界的時候,推理將產生的不同答案減少到非推理計數的0.43-0.65;當它是有界的時候,兩個臂持有相同的選項集,而推理則將質量集中於此,這是固定權重下任何位置先驗無法解釋的。總體成本小於機制預測的,因為推理也縮小了答案多樣性能決定任何事情的問題集,在十個單元中均如此,且縮小幅度為2.7倍;經過可用空間的標準化,兩個臂在領域中轉換了四分之一的空間。信心加權無法恢復剩餘的部分。在280種方法-數據集-模型組合中,涵蓋八個模型和五個基準,經過修正後,沒有一種方法超過普通的多數投票;加權投票在98.5%的問題-方法對上與其一致,並在其餘的情況下正確率為56.3%;而信號的方向可以在固定權重內反轉,當推理關閉時,答案的對數概率預測正確性,而當推理開啟時則預測錯誤性。一個學習到的六信號組合在領域外沒有任何收益。信心信號應該在決策上進行評估,而不是在區分上。
A Safety-Bounded SDC-to-MCP Gateway for Medical AI Agents
2609.31358v1 by Bennet Gerlach, Stefan Fischer
The Model Context Protocol (MCP) provides a common interface through which AI applications discover and use external resources and tools. It allows language-model agents to ground their reasoning in current system state and interact with heterogeneous services. In medical environments, however, exposing device state and action affordances requires deterministic constraints on possible effects. We present an IEEE 11073 Service-Oriented Device Connectivity (SDC)-to-MCP gateway that exposes metrics, alarms, context references, and semantic metadata as read-only resources, while representing selected action affordances as policy-validated dry-run tools. The term safety-bounded denotes a narrow no-execution property: agent-facing requests dispatch no SDC device operation. A Python prototype supports simulated fault and lifecycle experiments, a software-reference protocol path spanning independent Java and Python implementations, deterministic baselines, representation ablations, and multi-model agent evaluation. The results show semantically explicit resource exposure, visible rejection of invalid or outdated state, and preservation of the no-execution boundary across resource, proposal, and authorization paths. Explicit semantic metadata improved conformity to required metric identifiers in structured alarm outputs relative to a generic representation, while retained structured-output failures reveal a distinction between plausible narrative answers and task-compliant machine-readable results.
摘要:模型上下文協議 (MCP) 提供了一個共同的介面,讓 AI 應用程式發現並使用外部資源和工具。
它允許語言模型代理根據當前系統狀態進行推理並與異構服務互動。
然而,在醫療環境中,暴露設備狀態和行動可行性需要對可能的影響施加確定性的限制。
我們提出了一個 IEEE 11073 服務導向設備連接 (SDC) 到 MCP 的閘道,該閘道將指標、警報、上下文參考和語義元數據作為只讀資源暴露,同時將選定的行動可行性表示為經政策驗證的模擬工具。
術語安全界限表示一種狹窄的無執行特性:面向代理的請求不會調度任何 SDC 設備操作。
一個 Python 原型支持模擬故障和生命週期實驗,涵蓋獨立的 Java 和 Python 實現的軟體參考協議路徑、確定性基準、表示消融和多模型代理評估。
結果顯示語義明確的資源暴露、對無效或過時狀態的可見拒絕,以及在資源、提案和授權路徑中保持無執行邊界。
明確的語義元數據改善了結構化警報輸出中對所需指標標識符的符合性,相較於一般表示,保留的結構化輸出失敗揭示了合理敘述答案與符合任務的機器可讀結果之間的區別。
DriveHierarchy: A Benchmark for Diagnosing VLM Driving Capabilities from Open-Loop Understanding to Closed-Loop Execution
2609.31814v1 by Chengkai Xu, Jiaqi Liu, Yicheng Guo, Peng Hang, Jian Sun
Evaluating VLM-based autonomous driving remains difficult because driving competence is composite, where a capable system must ground traffic participants and hazards, integrate context across views and time, reason about future evolution, and act appropriately under closed-loop interaction. Existing benchmarks usually assess either open-loop understanding or closed-loop driving but provide limited structure for explaining how these abilities are organized, how they relate, and how they may inform model diagnosis and improvement. We present \textsc{DriveHierarchy}, a hierarchical benchmark that organizes VLM-based autonomous driving into four ranks, spanning perceptual grounding, contextual memory, mental reasoning, and closed-loop execution. To instantiate this hierarchy, we integrate multiple open-source autonomous-driving datasets into a unified open-loop benchmark with 76,798 question-answer pairs over 84,279 frames and develop a closed-loop simulation platform with interactive scenario construction on a real-world road network, from which 100 driving scenarios are curated for embodied evaluation. Experiments on 15 VLMs show that \textsc{DriveHierarchy} captures structured but non-redundant capability variation, relates open-loop understanding to closed-loop driving, and provides a practical basis for diagnosis and benchmark-guided optimization. \textsc{DriveHierarchy} therefore serves as a unified framework for evaluating and improving VLM-based autonomous driving systems. An anonymized project has been released on https://github.com/PerfectXu88/DriveHierarchy
摘要:評估基於 VLM 的自主駕駛仍然困難,因為駕駛能力是複合的,能夠的系統必須能夠定位交通參與者和危險,整合跨視角和時間的上下文,推理未來的演變,並在閉環互動中適當行動。現有的基準通常評估開環理解或閉環駕駛,但對於解釋這些能力如何組織、它們之間的關係,以及它們如何能夠幫助模型診斷和改進,提供的結構有限。我們提出了 \textsc{DriveHierarchy},這是一個將基於 VLM 的自主駕駛組織成四個等級的分層基準,涵蓋感知定位、上下文記憶、心理推理和閉環執行。為了實現這一層級,我們將多個開源自主駕駛數據集整合成一個統一的開環基準,包含 76,798 個問答對,涵蓋 84,279 幀,並開發了一個閉環模擬平台,能夠在現實世界的道路網絡上進行互動場景構建,從中策劃出 100 個駕駛場景以進行具體評估。對 15 個 VLM 的實驗顯示,\textsc{DriveHierarchy} 捕捉了結構化但不冗餘的能力變化,將開環理解與閉環駕駛相關聯,並提供了診斷和基準引導優化的實用基礎。因此,\textsc{DriveHierarchy} 作為評估和改進基於 VLM 的自主駕駛系統的統一框架。已在 https://github.com/PerfectXu88/DriveHierarchy 上發布了一個匿名項目。
Rethinking Data Quality for AI-Driven Systems: Evidence from Practitioner Interviews
2609.31191v1 by Hariharan Gopinath, Jan Bosch, Helena Holmström Olsson
Data quality research has usually treated data as an input that is stored, processed, and validated. In AI-driven software-intensive systems, data also shapes model behavior, evaluation, and lawful use. Empirical evidence remains limited on how practitioners define, assess, and manage quality under these conditions. We interviewed 16 practitioners from nine organizations and analyzed the transcripts using reflexive thematic analysis and developed six themes from participants' accounts. In AI systems, traceability shifted from modular debugging to attributing model behavior, while using models as quality assessors introduced circularity. Agent context and memory became data objects, and synthetic and pseudo-labeled data made authenticity a quality concern. In foundation-model development, lawfulness became a gate for training data, while representativeness was judged through coverage of situations in which the system must behave safely. Prior ML research examines many of these problems separately. Our study provides a practitioner-grounded account of how they are encountered together as an engineering and organizational concern. We also interpret five recurring conditions as helping explain how the themes relate to reduced trust in data and AI outcomes. We synthesize these findings through lifecycle assurance: a conceptual framing focused on producing evidence that data can support a specific AI claim when its influence may be embedded in model behavior, model-based judgments, or agent actions.
摘要:數據質量研究通常將數據視為一種被儲存、處理和驗證的輸入。在以 AI 驅動的軟體密集型系統中,數據也塑造了模型行為、評估和合法使用。在這些條件下,實證證據對於從業者如何定義、評估和管理質量仍然有限。我們訪談了來自九個組織的 16 位從業者,並使用反思主題分析法分析了訪談記錄,從參與者的敘述中發展出六個主題。在 AI 系統中,追溯性從模組調試轉變為歸因於模型行為,而將模型用作質量評估者則引入了循環性。代理上下文和記憶成為數據對象,而合成數據和偽標記數據使得真實性成為一個質量問題。在基礎模型開發中,合法性成為訓練數據的門檻,而代表性則通過系統必須安全行為的情境覆蓋來評判。先前的機器學習研究分別檢視了許多這些問題。我們的研究提供了一個以從業者為基礎的敘述,說明它們如何作為工程和組織問題共同出現。我們還解釋了五個反覆出現的條件,幫助說明這些主題如何與對數據和 AI 結果的信任減少相關。我們通過生命週期保證綜合這些發現:這是一個專注於產生證據的概念框架,證明數據可以支持特定 AI 主張,當其影響可能嵌入在模型行為、基於模型的判斷或代理行動中時。
Bayesian Optimization with Fisher Information Geometry: Gradient Bounds and Trust-Region Methods
2609.31107v1 by Saksham Kiroriwal, Julius Pfrommer, Jürgen Beyerer
We study Bayesian optimization (BO) through the lens of information geometry. Pulling back the Fisher information metric through the surrogate posterior map yields a local sensitivity tensor on the input space, which leads to an upper bound on the gradient of reparameterizable acquisition functions. This view explains vanishing-gradient behavior in high-dimensional BO and provides a common interpretation of heuristics such as RAASP and dimension-scaled lengthscales. Building on this analysis, we propose FITR, a trust-region-based BO method that replaces lengthscale-based scaling by local pullback-Fisher weights. FITR is not restricted to GP kernels with explicit lengthscales. On GP benchmarks with an SE kernel, experiments show competitive performance using FITR. The proposed method also easily generalizes to non-isotropic surrogates, although the gains are more task-dependent in that setting.
摘要:我們通過信息幾何的視角研究貝葉斯優化 (BO)。通過代理後驗映射回推費舍爾信息度量,產生了一個輸入空間上的局部敏感性張量,這導致了可重新參數化獲取函數梯度的上界。這一觀點解釋了高維 BO 中的消失梯度行為,並提供了對 RAASP 和維度縮放長度尺度等啟發式方法的共同解釋。在此分析的基礎上,我們提出了 FITR,一種基於信任區域的 BO 方法,通過局部回推費舍爾權重取代基於長度尺度的縮放。FITR 不僅限於具有明確長度尺度的 GP 核心。在具有 SE 核心的 GP 基準測試中,實驗顯示使用 FITR 的競爭性能。所提出的方法也很容易推廣到非各向同性的代理,儘管在該設置中增益更依賴於任務。