Skip to content

Medical explainable AI

Medical explainable AI

Publish Date Title Authors Homepage Code
2026-08-18 Can Large Language Models Explain Flight Safety Events? A Prior-Guided Semantic LLM-based Approach Lu Xu et.al. 2608.18017v1 null
2026-08-18 Grading Needs a Rubric, Not Intelligence Jhen-Ke Lin et.al. 2608.17938v1 null
2026-08-18 MotoSafety: Edge-AI with Learned Temporal Importance for Two-Wheeler Collision Risk Assessment Under Time Pressure Sumit S. Shevtekar et.al. 2608.17823v1 null
2026-08-18 Communicating Credit Risk with Large Language Models: Evaluation of Explanations from Standard and Alternative Data-Based Models Sahab Zandi et.al. 2608.17715v1 null
2026-08-18 Explainable AI-Powered Framework for Video-Based Skill Assessment in Cataract Surgery Mohammad Javad Ahmadi et.al. 2608.17522v1 null
2026-08-18 Understanding Curriculum Learning in Large Language Models via Cross-Difficulty Optimization Dynamics Zhikai Ding et.al. 2608.17268v1 null
2026-08-17 From Abductive Explanations to Global Logical Rules for Node Classification in SGCs Bryan Lima Cavalcante et.al. 2608.17103v1 null
2026-08-17 AutoSR: Automatic Symbolic Regression by Searching Research States Kejia Zhang et.al. 2608.16876v1 null
2026-08-17 Topological Attribution Distance (TAD): Revealing Segment-Level RAG Influence on LLM Output Geometry for Incident Log Analysis Reza Fayyazi et.al. 2608.16775v1 null
2026-08-17 Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments Adam Karvonen et.al. 2608.16747v1 null
2026-08-17 Unsupervised Anomaly Detection for Image Dataset Quality Assurance in Multi-Center Breast MRI Chiara Tappermann et.al. 2608.16725v1 null
2026-08-17 Degradation-Aligned Self-Supervised Learning for State of Health Estimation of Lithium-Ion Batteries under Label Sparsity Jiaqi Yao et.al. 2608.16612v1 null
2026-08-17 Physics of Agents: Statistical Mechanics Predicts Collective Behavior of AI Agents Batu El et.al. 2608.16578v1 null
2026-08-17 Revisiting the Performance of Generative Artificial Intelligence on Introductory Object-Oriented Programming Assessments: Insights from 2026 Marina Lepp et.al. 2608.16318v1 null
2026-08-17 Foresight-England: Development of a National-Scale Generative AI Model of Electronic Health Records for Medical Event Prediction across the COVID-19 Pandemic Simon Ellershaw et.al. 2608.16273v1 null
2026-08-17 Defake-o3: From Speculative Rationales to Verifiable Evidence for Explainable AIGI Detection Bowen Deng et.al. 2608.16259v1 null
2026-08-17 CompoSkill: Compositional Skill Chain Attacks from Individually Scanner-Passing LLM Agent Skills Mingxiao Liu et.al. 2608.16246v1 null
2026-08-17 When Single-Dataset Conclusions Fail: A 45-Task Study of Threshold Tuning and Resampling for Imbalanced Classification Diyorbek Musaev et.al. 2608.16147v1 null
2026-08-17 TokenSTFormer: A Tokenized Spatial-temporal Attention Model for Holistic Motion Analysis in Adolescent Idiopathic Scoliosis Screening Dong Chen et.al. 2608.16122v1 null
2026-08-17 Eigenanalysis framework for autoregressive neural emulators of multi-scale chaotic dynamics Conrad Ainslie et.al. 2608.16084v1 null
2026-08-17 NICE: Scale-Stable Perturbations for Graph Neural Network Explanations via Noise Corruption Ziluowen Luo et.al. 2608.16038v2 null
2026-08-16 Identifying Confusion Trends in Concept-based XAI for Multi-Label Classification Haadia Amjad et.al. 2608.15731v1 null
2026-08-16 Large Models for Small Devices: Recent Advances and Empirical Analysis of Edge AI Deployment Subhransu Das et.al. 2608.15693v1 null
2026-08-15 NumerosityVLM: A Cognitively Inspired Benchmark for Interpreting Numerosity Representations in Vision-Language Models Yiming Fu et.al. 2608.15425v1 null
2026-08-15 ETHOS: Towards a Modular Ethics Framework for Clinical Multi-Agent Systems Rakesh Sharma et.al. 2608.15424v1 null
2026-08-15 When AI Rewrites, Classifiers Relax: Uncertainty-Aware Sentiment Analysis on Sarcastic and AI-Paraphrased Social Text Shresth Shroff et.al. 2608.15338v1 null
2026-08-15 Demographic Injection in Medical Language Models under Diversity, Equity, and Inclusion Prompts Diego Mardian et.al. 2608.15254v1 null
2026-08-15 Low-Rank Dynamics-Effective Latent Carriers for Counterfactual Rollout in Learned World Models Yang Liu et.al. 2608.15156v1 null
2026-08-15 Fast Test-Time Refinement for Robust Learned Image Compression Jiaming Liang et.al. 2608.15113v1 null
2026-08-15 Command-Space Counterfactual Explanations for Pareto-Conditioned Reinforcement Learning Joanikij Chulev et.al. 2608.14963v1 null
2026-08-14 Handover Analysis for Vehicular Communication with Explainability on the Fly Ali Fuat Sahin et.al. 2608.14820v1 null
2026-08-14 Generated Context versus Governed State: Functional Conditions for Accountable Longitudinal Clinical Reasoning Augusto Bernardo Pissarra et.al. 2608.14804v1 null
2026-08-14 Decoding the Past: An Uncertainty-Aware Deep Learning Framework for Sex Attribution in Prehistoric Hand Stencils Karel Becerra et.al. 2608.14539v1 null
2026-08-14 NARRATE: A Multimodal Real-World Australian Driving Dataset for Human-Centred Explanations in Automated Driving Ashkan Yousefi Zadeh et.al. 2608.14767v1 null
2026-08-14 Polaris : Multi Agentic System for Conversational Enterprise Analytics Varuni H K et.al. 2608.14246v1 null
2026-08-14 CForce: Boosting Parallel Decoding for dLLMs via Consistency Forcing Yuji Ren et.al. 2608.13925v1 null
2026-08-13 Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions Qingfang Liu et.al. 2608.13786v1 null
2026-08-13 Capacity-Dependent Effects of Data Selection for Reasoning Cuong Dang et.al. 2608.13721v1 null
2026-08-13 MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and Coordination Saisha Shetty et.al. 2608.13476v1 null
2026-08-13 A Unifying Perspective on Causal World Models: From Observations to Representations to Structure Avinash Kori et.al. 2608.13456v1 null
2026-08-13 Explanatory Engagement Under Rare Anomalous Failure: Asymptotic Rarity in Model Behavior (or: The Asymptotic AI) Sam Mao et.al. 2608.13063v1 null
2026-08-13 VALG: An Agentic System for ML Theory Research Dechen Zhang et.al. 2608.13060v1 null
2026-08-13 UniTraffic-Agent: Unified Traffic Video Reasoning for AI City Challenge 2026 Track 3 with Two Out-of-Domain Evaluations Peng Li et.al. 2608.13031v1 null
2026-08-13 Static analysis-guided agentic AI translation enables Rust as a full stack bioinformatics language Johan Henriksson et.al. 2608.13029v1 null
2026-08-13 Decomposition of Evidence, Contradiction, and Fragility in Perturbation Responses Lei You et.al. 2608.12935v1 null
2026-08-13 Discovering Efficient and Explainable Communication Topologies for LLM-based Multi-Agent Systems via Causal Inference Junzhi Li et.al. 2608.12921v2 null
2026-08-13 Mr3D-VL: A generalist vision language foundation model for Multiparametric 3D Magnetic Resonance Imaging Zhi Qiao et.al. 2608.12689v1 null
2026-08-12 Interpretable Causal Discovery via Causal-Effect Constraints Cixuan Zhang et.al. 2608.12640v1 null
2026-08-12 Algorithm Design and Physician Liability Shujie Luan et.al. 2608.13618v1 null
2026-08-12 What Makes a Peer? Valuation-Anchored Similarity in Private Markets Sebastian Frank et.al. 2608.12594v1 null
2026-08-12 Auditable agentic AI for evidence-grounded thyroid ultrasound diagnosis and reporting Haifan Gong et.al. 2608.12590v1 null
2026-08-12 CAS: A Causal Attribution Score for Local and Global Explainable Artificial Intelligence Michael Georgiades et.al. 2608.12555v1 null
2026-08-12 Class Activation Mapping in Explainable Computer Vision: A Method-Centered Review of CNN, Transformer, and Foundation-Model-Era Visual Explanations AmirHossein Eshghi et.al. 2608.12299v2 null
2026-08-12 Dual Spatial-Temporal Attribution: Architecture-Aligned Post-Hoc Explainability for Recurrent Graph Anomaly Detection Iyad Assaad Nekka et.al. 2608.12441v1 null
2026-08-12 Learning-Based Behavior Planning for Automated Driving: Real-World Integration and Deployment Jean-Pierre Busch et.al. 2608.12198v1 null
2026-08-12 A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench Praveen Reddy et.al. 2608.12138v1 null
2026-08-12 Do LLMs Take Care of Their Own? Similarity Signals Can Induce Cooperation Akash Kundu et.al. 2608.12125v1 null
2026-08-12 Faithful, Sufficient and Understandable: Rethinking Graph Counterfactual Explanations via Discrete Diffusion Inversion David Bechtoldt et.al. 2608.12083v1 null
2026-08-12 Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence Mengru Wang et.al. 2608.12036v1 null
2026-08-12 From Safety Documentation to Safety Knowledge Support: An Evidence-Grounded LLM Framework for Medical Devices Tuhinangshu Gangopadhyay et.al. 2608.12025v1 null
2026-08-12 Agent Skills Can Be Harmful: An Empirical Study of Skill-Induced Failures in LLM Agents Gen Dong et.al. 2608.11888v1 null
2026-08-12 Low-Interaction-Rank Learning: Unifying Multiplicative Dual-Encoder Heads Zijian Zhao et.al. 2608.11661v1 null
2026-08-11 Inverse Theory of Mind Modeling for Content Recommendation: From Web Browsing to Dynamic Intelligent Interfaces Mengyu Chen et.al. 2608.11354v1 null
2026-08-11 Governing Agentic AI in FinTech Henry Han et.al. 2608.11344v2 null
2026-08-11 From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Rahul Gupta et.al. 2608.11171v1 null
2026-08-11 SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure Xiaofan Bai et.al. 2608.11079v2 null
2026-08-11 Entropy-Centric Explainable AI for Remote Sensing Image Segmentation Ali Saleh et.al. 2608.11064v1 null
2026-08-11 ComBodied Agents: a New Paradigm of Human-Centric Agentic AI Qianggang Ding et.al. 2608.10915v2 null
2026-08-11 Chemically Meaningful Textualization Enables Explainable Validation of Metal-Organic Frameworks by Large Language Models Guobin Zhao et.al. 2608.11283v1 null
2026-08-11 Uncertainty-Aware and Explainable Ensemble Deep Learning Framework for Multi-Class Skin Lesion Classification Rofiqul Islam et.al. 2608.11280v1 null
2026-08-11 Rule of Thumb: Explaining Artificial Intelligence Systems using Partial Information Kaivalya Rawal et.al. 2608.10766v2 null
2026-08-11 Operationalising Relative Causal Knowledge: Backbone Identifiability from Private Reports on a Shared Outcome Fabrizio Russo et.al. 2608.10664v1 null
2026-08-11 Conversational versus Dashboard Explainable AI for UAV Intrusion Detection: An Empirical Study of Operator Trust and Reliance Cong Chi Nguyen et.al. 2608.10434v1 null
2026-08-11 Reasoning Shortcuts and Value Symmetries: What Symmetry Permits, Architecture Realizes, and Optimization Selects Xin Xu et.al. 2608.10420v1 null
2026-08-10 Towards Expert-level Medical AI for Real-time Video Consultations Mahvish Nagda et.al. 2608.09861v1 null
2026-08-10 CEAA: A Cognitive Embodied Agents Architecture for Interactive Computing Systems Aimilios Hadjiliasi et.al. 2608.09848v1 null
2026-08-10 KGCaRe: Explainable Complex Conditional Question Answering using Automatic Knowledge Graph Construction and Context Retrieval with LLMs Ghanshyam Verma et.al. 2608.09779v1 null
2026-08-10 Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models Shulin Tian et.al. 2608.09666v1 null
2026-08-10 Rethinking Self-Evolving Agents: Do We Still Need Prescribed Optimization Pipelines? Hui Xue et.al. 2608.09629v1 null
2026-08-10 Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Hongli Shen et.al. 2608.09542v1 null
2026-08-10 Renormalising Generative Models for Active Inference: Foundations, Derivations, and Verification Karim Zaghw et.al. 2608.09512v1 null
2026-08-10 How Simple Can It Get? From Interpretable Equations to Readable Rules for Financial Decision Making Adia Lumadjeng et.al. 2608.09433v1 null
2026-08-10 An Explainable GNN Framework for Component-Level Anomaly Diagnosis Sena Ozgunay et.al. 2608.09246v1 null
2026-08-10 SafeSceneReason: A Multimodal Reasoning Benchmark Connecting Industrial Hazards with Accident Knowledge Yuanchi Zhu et.al. 2608.09230v1 null
2026-08-10 TLDChoiceNet: Quantitatively Choosing a Transfer Learning Dataset Jing Ning et.al. 2608.09091v1 null
2026-08-09 Reading is not Reasoning: Bridging the Agentic Policy Gap in Vision-Text Compression Cheng Fan et.al. 2608.08960v1 null
2026-08-09 From Recovery to Drop-off: How Action Post-training Reduces a VLM's Late-Layer Depth Decodability Alexander Hackett et.al. 2608.08904v2 null
2026-08-09 From Manuals to Maintenance: Fine-Tuning MedGemma for Multi-Modal Imaging System Support in Low-Resource Settings Bernes Lorier Atabonfack et.al. 2608.08896v1 null
2026-08-09 PROSLEX: A Novel Dataset for Expert-Annotated Legal Statute Prediction for Indian Judiciary Subinay Adhikary et.al. 2608.08830v1 null
2026-08-09 Deployable Per-Instance Multi-Layer Activation Steering for Large Language Models Muhammad Faishal Adly Nelwan et.al. 2608.08829v1 null
2026-08-09 SymDiag: Explainable Diagnosis for LLM Reasoning via Neuro-Symbolic Verification Wenyao Cui et.al. 2608.08786v1 null
2026-08-09 Domain Agnostic Text Redaction from Natural Language Rules using Instruction Tuning Aravindhan Arunagiri et.al. 2608.14693v1 null
2026-08-09 Business Arena: Benchmarking LLM Agents in a Realistic Marketplace Yijun Pan et.al. 2608.08621v1 null
2026-08-09 On-Device Multi-Species Malaria Detection with Uncertainty-Calibrated Slide-Level Aggregation Idaya Seidu et.al. 2608.08566v1 null
2026-08-09 Private Etymology: Designing Relational Reuse of Shared Symbols in Long-Term Human-AI Interaction Miki Ueno et.al. 2608.08443v2 null
2026-08-08 Quantization Degradation in Large Language Models: A Signal-Noise Perspective Chenxi Zhou et.al. 2608.08188v1 null
2026-08-08 Agentic AI-driven Immersive Simulation: A Knowledge-Aware Virtual Training Platform forHigh Dose Rate (HDR) Brachytherapy Ronghua Xu et.al. 2608.08163v1 null
2026-08-08 Compositional Threat Analysis of Latent Compromise in LLM Agent Systems: The Order 66 Scenario Satoshi Matsuoka et.al. 2608.08131v1 null
2026-08-08 Defending Retrieval-Augmented Intrusion Detection Against Knowledge Poisoning and Prompt Injection Kaysarul Anas Apurba et.al. 2608.08100v1 null
2026-08-08 HugSelect: An Explainable Multi-Criteria Decision-Support Framework for foundation-model selection Alireza Joonbakhsh et.al. 2608.08069v1 null

Abstracts

Can Large Language Models Explain Flight Safety Events? A Prior-Guided Semantic LLM-based Approach

2608.18017v1 by Lu Xu, Xu Li, Linjiang Zheng, Fan Li, Riquan Zhang, Jiaxing Shang

Improving flight safety with flight data requires not only accurate detection of risk events, but more importantly, clear interpretation of their underlying causes at the level of pilot control behavior. Existing explainable AI techniques, such as feature importance maps, often require considerable domain knowledge to translate them into operationally meaningful explanations. Large Language Models (LLMs), which excel at language reasoning, bring a promising solution to this issue. However, applying LLMs in this domain presents key challenges such as modal inconsistency, limited classification ability, scarcity of task-specific data for fine-tuning, and lack of domain knowledge. To overcome these challenges, we propose FlightLLM, a prior-guided semantic LLM-based approach for interpretable flight safety analysis. Specifically, we first perform feature engineering to address modal inconsistency, combining statistical descriptors with physically meaningful flight indicators. This representation is further processed by a Semantic Discretization module, which converts abstract numerical patterns into qualitative descriptions that are more compatible with language reasoning. In addition, since LLMs are not inherently strong classifiers, CatBoost is incorporated as a statistical expert, and its prediction results are injected into the prompt as prior guidance. A contrastive few-shot learning strategy is further adopted to compensate for limited data. Finally, we design structured prompts to embed aviation-specific knowledge into the inference process. Using hard landing, a representative risk event with complex causal mechanisms, as an anchor point, we evaluate FlightLLM on a dataset of 704 real-world A320 flight samples. Experimental results show that the proposed approach achieves competitive classification performance while generating direct and reasonable explanations for event causes.

摘要:改善飛行安全需要不僅準確檢測風險事件,更重要的是在飛行員控制行為層面清晰解釋其潛在原因。現有的可解釋AI技術,如特徵重要性圖,通常需要相當的領域知識才能將其轉化為具有操作意義的解釋。大型語言模型(LLMs)在語言推理方面表現出色,為這一問題帶來了有希望的解決方案。然而,在這一領域應用LLMs面臨著關鍵挑戰,如模式不一致、有限的分類能力、缺乏特定任務的數據以進行微調,以及缺乏領域知識。為了克服這些挑戰,我們提出了FlightLLM,一種基於語義的先驗引導LLM方法,用於可解釋的飛行安全分析。具體而言,我們首先進行特徵工程以解決模式不一致,將統計描述符與具有物理意義的飛行指標相結合。這一表示進一步由語義離散化模塊處理,將抽象的數字模式轉換為更符合語言推理的定性描述。此外,由於LLMs本身並不是強大的分類器,因此CatBoost被納入作為統計專家,其預測結果被注入到提示中作為先驗指導。進一步採用了對比少樣本學習策略以彌補數據的有限性。最後,我們設計了結構化提示,將航空特定知識嵌入推理過程中。以硬著陸作為錨點,這是一個具有複雜因果機制的代表性風險事件,我們在704個真實世界A320飛行樣本的數據集上評估FlightLLM。實驗結果表明,所提出的方法在生成事件原因的直接和合理解釋的同時,實現了具有競爭力的分類性能。

Grading Needs a Rubric, Not Intelligence

2608.17938v1 by Jhen-Ke Lin

Small language models can grade open-ended examination answers as reliably as substantially more expensive models when they grade against an explicit rubric. We test this claim as the design principle behind any-to-bench: a frontier model reads source documents once, at ingestion, to extract each question and its rubric; lower-cost models then perform all repeated grading work. We evaluate six cost-efficient model configurations from two model families at three reasoning-effort levels. Each configuration answers 24 open-ended examination questions, and each also grades every answer sheet three times, yielding 3,456 per-question grades. Scores depend overwhelmingly on the answer being graded: answer identity explains 95.6% of score variance, whereas judge identity explains only 0.2%. Raising a writer's reasoning effort moves earned scores by as much as 0.143 of full marks, while raising a judge's reasoning effort moves assigned scores by at most 0.006. Six frontier-tier judges, added as a check, reproduce these scores and are no more reliable as a panel. Two ablations then decompose the rubric on the same questions and answers. Removing its criteria and levels while keeping the official answer changes nothing measurable. Removing the official answer as well collapses reliability (ICC 0.888 to 0.628), inflates scores, and makes judge reasoning effort matter again. The rubric is what decouples grading from judge intelligence, and within the rubric the official answer does nearly all the work. We find no evidence of length preference or same-family preference under rubric-anchored grading.

摘要:小型語言模型在根據明確的評分標準進行評分時,可以與成本高得多的模型一樣可靠地評分開放式考試答案。我們測試這一主張,作為 any-to-bench 的設計原則:前沿模型在攝取時讀取源文件一次,以提取每個問題及其評分標準;然後,成本較低的模型執行所有重複的評分工作。我們在三個推理努力水平上評估來自兩個模型系列的六種成本效益模型配置。每個配置回答 24 道開放式考試問題,並且每個配置還對每份答案進行三次評分,產生每個問題 3,456 次評分。分數在很大程度上取決於被評分的答案:答案身份解釋了 95.6% 的分數變異,而評判身份僅解釋了 0.2%。提高寫作者的推理努力可以使得獲得的分數提高最多 0.143 的滿分,而提高評判的推理努力則最多使分數提高 0.006。六位前沿級評判作為檢查,重現這些分數,且作為小組的可靠性並沒有提高。接下來的兩個消融實驗則在相同的問題和答案上分解評分標準。去除其標準和級別,同時保留官方答案,並不會改變可測量的結果。去除官方答案也會使可靠性崩潰(ICC 從 0.888 降至 0.628),使分數膨脹,並使評判的推理努力再次變得重要。評分標準是將評分與評判智力解耦的關鍵,而在評分標準內,官方答案幾乎承擔了所有的工作。我們沒有發現基於評分標準的評分中存在長度偏好或同家族偏好的證據。

MotoSafety: Edge-AI with Learned Temporal Importance for Two-Wheeler Collision Risk Assessment Under Time Pressure

2608.17823v1 by Sumit S. Shevtekar, Chandresh K. Maurya, Gourab Sil, Subasish Das

Powered two-wheeler riders face critical safety challenges in low- and middle-income countries, yet limited studies exist on how cognitive stressors such as Time Pressure influence collision risk. To address this gap, we introduce a large-scale dataset of over 129,000 labeled multivariate time-series sequences from 153 simulator rides by 51 participants under No, Low, and High TP, capturing 64 features across vehicle dynamics, control inputs, proximity, and behavioral violations. Building on this dataset, we propose MotoSafety, a novel edge-AI architecture grounded in the Learned Temporal Importance principle. MotoSafety achieves 94.97% accuracy and 99.33% ROC AUC, outperforming ten baselines, including TimesNet and LLM4TS, and achieves 0.039 MSE and 0.094 MAE for forecasting (4.4x lower error than Time-LLM and iTransformer). With only 1.15M parameters and 0.135 ms latency, it is suitable for edge deployment on low-cost CPU hardware. Using ground truth TP as an inductive bias improves accuracy from 94.09% to 94.97%, while predicted TP achieves 94.82%. Using only 21 IMU+GPS features, it achieves 93.91% accuracy, indicating practical deployment. Beyond PTW safety, the architecture shows better transferability to human activity (97.66%) and clinical (99.65%) domains. This lightweight framework advances PTW collision risk assessment, supporting the Safe System Approach for Intelligent Transportation Systems.

摘要:在中低收入國家,動力二輪車騎士面臨著重大的安全挑戰,但關於認知壓力因素如時間壓力如何影響碰撞風險的研究卻相對有限。為了填補這一空白,我們引入了一個大規模數據集,該數據集包含來自51名參與者在無時間壓力、低時間壓力和高時間壓力下進行的153次模擬騎行的超過129,000個標記的多變量時間序列,捕捉了64個特徵,涵蓋了車輛動態、控制輸入、接近度和行為違規。基於這個數據集,我們提出了MotoSafety,一種基於學習時間重要性原則的新型邊緣人工智慧架構。MotoSafety實現了94.97%的準確率和99.33%的ROC AUC,超越了包括TimesNet和LLM4TS在內的十個基準,並在預測中達到了0.039的均方誤差和0.094的平均絕對誤差(比Time-LLM和iTransformer低4.4倍)。它僅需1.15M的參數和0.135毫秒的延遲,適合在低成本CPU硬體上進行邊緣部署。使用真實的時間壓力作為歸納偏見,準確率從94.09%提高到94.97%,而預測的時間壓力則達到94.82%。僅使用21個IMU+GPS特徵,它的準確率達到93.91%,顯示出實際部署的潛力。除了PTW安全性外,該架構在人體活動(97.66%)和臨床(99.65%)領域也顯示出更好的可轉移性。這個輕量級框架推進了PTW碰撞風險評估,支持智能交通系統的安全系統方法。

Communicating Credit Risk with Large Language Models: Evaluation of Explanations from Standard and Alternative Data-Based Models

2608.17715v1 by Sahab Zandi, Noah Kostesku, Christophe Mues, María Óskarsdóttir, Cristián Bravo

Credit decisioning is a high-stakes task in which model outputs must be accurate and explainable to support compliant decisions. Although modern credit risk models such as eXtreme Gradient Boosting (XGBoost) and Graph Neural Networks (GNNs) improve predictive performance, their explanations are often too technical for stakeholders creating communication gaps that can shape approvals, denials, and fairness judgments. We examine whether Large Language Models (LLMs) can serve as explanation layers that translate post-hoc explanation artefacts into stakeholder-appropriate risk narratives. Using Freddie Mac single-family loan-level data, we develop three pipelines: standard tabular (XGBoost + SHAP), and two with alternative data, a pure network-based (GNN + GNNExplainer), and a bimodal one (combining tabular and network data). We generate narratives with three LLM configurations: a small fine-tuned LLM (Gemma 3 4B), a large fine-tuned LLM (DeepSeek R1 70B), and a zero-shot commercial LLM (Gemini 2.5). Explanation quality is evaluated through automated checks across all pipelines and a human study of bimodal explanations comparing credit risk professionals and non-professionals on eight decision-relevant dimensions. We have three main findings. First, the pipeline accounts for higher variance in evidence-grounding scores than the language model, meaning that the binding constraint on explanation quality is the evidence representation, not the model used. Second, the explanation narratives reliably name the influential factors but are less reliable when stating the direction of influence, which may be consequential for adverse-action communication. Finally, professionals apply stricter evidentiary standards than non-professionals. We discuss implications for the governance of risk models, including deployment considerations and the value of domain-aligned LLMs in regulated credit settings.

摘要:信用決策是一項高風險的任務,其中模型輸出必須準確且可解釋,以支持合規的決策。儘管現代信用風險模型如極端梯度提升(XGBoost)和圖神經網絡(GNNs)提高了預測性能,但它們的解釋往往對利益相關者來說過於技術性,造成溝通差距,這可能影響批准、拒絕和公平性判斷。我們檢視大型語言模型(LLMs)是否可以作為解釋層,將事後解釋產物轉化為適合利益相關者的風險敘事。使用Freddie Mac的單戶貸款數據,我們開發了三個管道:標準表格(XGBoost + SHAP),以及兩個使用替代數據的管道,一個是純基於網絡的(GNN + GNNExplainer),另一個是雙模的(結合表格和網絡數據)。我們使用三種LLM配置生成敘事:一個小型微調LLM(Gemma 3 4B),一個大型微調LLM(DeepSeek R1 70B),以及一個零樣本商業LLM(Gemini 2.5)。通過對所有管道的自動檢查以及對雙模解釋的人工研究,我們評估了解釋質量,並比較了信用風險專業人員和非專業人員在八個與決策相關的維度上的表現。我們有三個主要發現。首先,該管道在證據基礎分數的變異性上比語言模型更高,這意味著解釋質量的約束是證據表示,而不是所使用的模型。其次,解釋敘事可靠地命名了影響因素,但在陳述影響方向時可靠性較低,這對於不利行動的溝通可能具有重要意義。最後,專業人士應用的證據標準比非專業人士更為嚴格。我們討論了風險模型治理的影響,包括部署考量和在受監管的信用環境中領域對齊的LLMs的價值。

Explainable AI-Powered Framework for Video-Based Skill Assessment in Cataract Surgery

2608.17522v1 by Mohammad Javad Ahmadi, Hamid D. Taghirad

Persistent shortages in the surgical workforce and inherent limitations of traditional training methods highlight the necessity of automated, data-driven approaches in surgical education. This study addresses these challenges by introducing a novel, explainable AI-powered framework for automated skill assessment, specifically focusing on cataract surgery. We present the world's largest dataset of cataract surgery videos, comprising 2,000 recordings. Additionally, we propose an AI-powered analytical framework that employs advanced computer vision and signal-processing techniques to automatically evaluate surgical videos to derive objective, quantitative performance indicators that complement or potentially replace subjective scoring methods. A significant advantage of our framework over previous methods lies precisely in its explainability of outputs, elevating it beyond merely an opaque skill classification tool. Through experimental analysis of 83 cataract surgery videos, we demonstrate that the automatically computed metrics exhibit strong correlations with expert-based subjective evaluations, achieving up to 87% accuracy in surgical skill assessment. Each metric was individually examined, and expert surgeons provided subjective ratings using the newly introduced Capsulorhexis Skill Assessment System (CSAS). These subjective assessments were compared with ten objective motion-based metrics extracted through our framework. The results indicated a robust correlation between subjective ratings and automated indicators, underscoring the framework's capacity to accurately model surgical expertise.

摘要:持續的外科醫療人力短缺以及傳統訓練方法的固有限制凸顯了在外科教育中自動化、數據驅動方法的必要性。這項研究通過引入一個新穎的、可解釋的人工智慧驅動框架來解決這些挑戰,特別專注於白內障手術。我們展示了世界上最大的白內障手術視頻數據集,包含2,000個錄像。此外,我們提出了一個人工智慧驅動的分析框架,利用先進的計算機視覺和信號處理技術,自動評估手術視頻,以獲得客觀的、定量的性能指標,這些指標可以補充或潛在地取代主觀評分方法。我們的框架相較於先前的方法的一個顯著優勢恰恰在於其輸出的可解釋性,使其超越僅僅是一個不透明的技能分類工具。通過對83個白內障手術視頻的實驗分析,我們證明自動計算的指標與專家基於主觀評估的評分之間存在強烈的相關性,在外科技能評估中達到高達87%的準確率。每個指標都經過單獨檢查,專家外科醫生使用新引入的囊膜切開技能評估系統(CSAS)提供主觀評分。這些主觀評估與通過我們的框架提取的十個客觀運動基礎指標進行了比較。結果顯示主觀評分與自動指標之間存在穩健的相關性,強調了該框架準確建模外科專業知識的能力。

Understanding Curriculum Learning in Large Language Models via Cross-Difficulty Optimization Dynamics

2608.17268v1 by Zhikai Ding, Ziyi Ye

Curriculum learning has been widely adopted in the post-training of large language models by organizing training data from easy to hard. However, its effectiveness varies substantially across reasoning tasks, suggesting that no single curriculum is universally optimal and raising a fundamental question: what determines when curriculum learning works? In this paper, we answer this question by analyzing the optimization dynamics induced by different curriculum schedules. We show that the transfer relationship between different difficulty levels characterizes the optimization dynamics induced by curriculum learning, which in turn explains the effectiveness of different curriculum schedules, and formalize this relationship as Relative Transfer, a principled measure of cross-difficulty knowledge transfer. Based on this measurement, we derive Transfer-aware Dynamic Curriculum Sampling (TDCS), which dynamically adjusts the sampling distribution according to the estimated transfer relationship throughout training. Extensive experiments on multiple reasoning benchmarks demonstrate that TDCS consistently outperforms representative scheduling strategies across different tasks, model scales, and training paradigms. More importantly, our work provides a unified optimization-based explanation of curriculum learning through cross-difficulty transfer.

摘要:課程學習已被廣泛應用於大型語言模型的後訓練,通過將訓練數據從簡單到困難進行組織。然而,它在推理任務中的有效性差異很大,這表明沒有單一的課程是普遍最佳的,並提出了一個根本性問題:什麼決定了課程學習的有效性?在本文中,我們通過分析不同課程安排所引起的優化動態來回答這個問題。我們展示了不同難度級別之間的轉移關係特徵化了課程學習所引起的優化動態,這反過來解釋了不同課程安排的有效性,並將這一關係形式化為相對轉移,這是一種跨難度知識轉移的原則性度量。基於這一測量,我們推導出轉移感知動態課程抽樣(TDCS),該方法根據整個訓練過程中估計的轉移關係動態調整抽樣分佈。在多個推理基準上的大量實驗表明,TDCS在不同任務、模型規模和訓練範式中始終優於代表性的排程策略。更重要的是,我們的工作通過跨難度轉移提供了一個統一的基於優化的課程學習解釋。

From Abductive Explanations to Global Logical Rules for Node Classification in SGCs

2608.17103v1 by Bryan Lima Cavalcante, Thiago Alves Rocha

Graph Neural Networks (GNNs) have achieved remarkable performance in node classification tasks, motivating growing interest in methods capable of explaining their predictions. Recent logic-based approaches, such as LogicXGNN, derive global logical rules for Graph Neural Networks (GNNs) from collections of explanatory subgraphs. While informative, these subgraphs may contain redundant structural information that is specific to individual nodes, potentially limiting the generality of the extracted rules. In this work, we propose a logic-based framework for node classification in Simple Graph Convolution (SGC) networks that uses minimal abductive explanations as an intermediate representation for rule extraction. For each node, we compute a minimal set of node-feature pairs sufficient to preserve the predicted class. These explanations are then used to train decision trees from which global logical rules are extracted. Experiments on benchmark datasets show that the proposed framework produces compact global rules while maintaining high fidelity to the original SGC model.

摘要:圖神經網絡(GNNs)在節點分類任務中取得了顯著的表現,這激發了對能夠解釋其預測的方法的日益關注。最近的基於邏輯的方法,如LogicXGNN,從解釋性子圖的集合中推導出圖神經網絡(GNNs)的全局邏輯規則。雖然這些子圖提供了資訊,但它們可能包含特定於個別節點的冗餘結構資訊,這可能限制了提取規則的普遍性。在本研究中,我們提出了一個基於邏輯的框架,用於簡單圖卷積(SGC)網絡中的節點分類,該框架使用最小的推斷解釋作為規則提取的中介表示。對於每個節點,我們計算一組最小的節點-特徵對,這些對足以保留預測的類別。然後,這些解釋用於訓練決策樹,從中提取全局邏輯規則。在基準數據集上的實驗表明,所提出的框架生成了緊湊的全局規則,同時保持了對原始SGC模型的高保真度。

AutoSR: Automatic Symbolic Regression by Searching Research States

2608.16876v1 by Kejia Zhang, Youran Sun, Xinyu Ren, Chugang Yi, Haizhao Yang

We introduce Automatic Symbolic Regression (AutoSR), a fully automated system that instantiates Research-Space Symbolic Regression by searching persistent scientific investigations rather than isolated equations. Finite, noisy data often yield numerically competitive expressions that imply very different behavior outside the observed regime, making numerical fit and syntactic complexity insufficient measures of scientific credibility. Existing approaches largely focus on improving expressions, yet the search typically retains little beyond the resulting formula and score, losing the scientific record, such as motivations and probes, that inform what to try next. AutoSR preserves this record in a \textbf{Research State}, coupling each candidate equation with the reasoning, computational evidence, and independent review developed along its branch. Proposer--reviewer agents develop these states under progressive-widening Monte Carlo tree search (PW-MCTS), which allocates computation across competing investigations, while the accumulated research record is ultimately synthesized into a final report that explains the leading relation and the basis for its selection. Across nine selected challenges from two benchmark suites, AutoSR recovers algebraically equivalent relations in every case, including three cp3-bench problems that no published system recovers and six structurally diverse LSR-Transform problems. Overall, AutoSR extends symbolic regression from equation-level search toward automated scientific investigation, allowing scientific knowledge and accumulated evidence to shape both what is explored and how the resulting equation is justified.

摘要:我們介紹自動符號回歸(AutoSR),這是一個完全自動化的系統,它通過搜尋持續的科學研究而不是孤立的方程式來實現研究空間符號回歸。有限的、帶噪聲的數據通常會產生數值上具有競爭力的表達式,這些表達式在觀察範圍之外暗示了非常不同的行為,使得數值擬合和語法複雜性不足以作為科學可信度的衡量標準。現有的方法主要集中在改進表達式上,但搜索通常僅保留結果公式和分數,失去了科學記錄,例如動機和探測,這些記錄告訴我們接下來該嘗試什麼。AutoSR 在一個 \textbf{研究狀態} 中保留這個記錄,將每個候選方程與沿其分支發展的推理、計算證據和獨立審查相結合。提議者-審查者代理在漸進擴展的蒙特卡羅樹搜索(PW-MCTS)下發展這些狀態,該方法在競爭的研究之間分配計算,而累積的研究記錄最終被綜合成一份最終報告,解釋主要關係及其選擇的基礎。在來自兩個基準套件的九個選定挑戰中,AutoSR 在每一個案例中都恢復了代數上等價的關係,包括三個沒有任何已發表系統恢復的 cp3-bench 問題和六個結構多樣的 LSR-Transform 問題。總體而言,AutoSR 將符號回歸從方程層級的搜索擴展到自動化的科學研究,允許科學知識和累積的證據塑造探索的內容以及結果方程的合理性。

Topological Attribution Distance (TAD): Revealing Segment-Level RAG Influence on LLM Output Geometry for Incident Log Analysis

2608.16775v1 by Reza Fayyazi, Michael Zuzak, Shanchieh Jay Yang

Large Language Models (LLMs) are increasingly being deployed in cybersecurity operations to assist cybersecurity analysts with rapid decision-making against emerging threats. However, there is a main criteria that must be met when using LLMs in cybersecurity, that is, trust in the generated outputs. As Agentic AI is integrated into operational systems, a robust evidence attribution and provenance tracking technique is essential to trace the origins of model generations. When autonomous agents make a decision (right or wrong), the ability to trace back through the decision chain is critical, as without it, teams cannot identify which segment of the data caused the model generation. Existing methods often struggle to distinguish among complex and highly similar evidence sources, such as cyber incident logs. This reveals a key gap: current approaches do not adequately capture the holistic geometric relationship between the retrieved evidence and the generated response for reliable evidence verification. To bridge this gap, we propose Topological Attribution Distance (TAD), inspired by Topology, to characterize and capture the global geometric shape of an output and its changes against its retrieved logs. In other words, if the embeddings of a specific source log drastically changes the geometry of the model's response in the embedding space, this suggests that such log is a critical source for the model's generated response. Therefore, TAD is powered by segment-level ablation attribution to investigate incident logs of an actual cyberattack. We demonstrate how TAD finds the most attributed logs on LLM outputs in an adaptive manner. This can provide an explainable and trustworthy tracing based on each LLM's hidden state to understand how geometrically different retrieved logs influence the model generation, and provide evidence verification in cybersecurity and Agentic-AI workflows.

摘要:大型語言模型(LLMs)越來越多地被應用於網路安全操作,以協助網路安全分析師快速做出針對新興威脅的決策。然而,在網路安全中使用LLMs時,必須滿足一個主要標準,即對生成輸出的信任。隨著代理式人工智慧整合進入操作系統,強健的證據歸屬和來源追蹤技術對於追溯模型生成的來源至關重要。當自主代理做出決策(無論對錯)時,能夠追溯決策鏈是關鍵,因為沒有這一點,團隊無法確定是哪一部分數據導致了模型生成。現有的方法往往難以在複雜且高度相似的證據來源之間區分,例如網路事件日誌。這揭示了一個關鍵的缺口:當前的方法未能充分捕捉檢索到的證據與生成回應之間的整體幾何關係,以進行可靠的證據驗證。為了填補這一缺口,我們提出了拓撲歸屬距離(TAD),受到拓撲學的啟發,用於表徵和捕捉輸出及其相對於檢索日誌的變化的全球幾何形狀。換句話說,如果特定來源日誌的嵌入顯著改變了模型在嵌入空間中的回應幾何形狀,這表明該日誌是模型生成回應的關鍵來源。因此,TAD由段級別的去除歸屬推動,以調查實際網路攻擊的事件日誌。我們展示了TAD如何以自適應的方式找到對LLM輸出最具歸屬的日誌。這可以基於每個LLM的隱藏狀態提供可解釋且值得信賴的追蹤,以了解幾何上不同的檢索日誌如何影響模型生成,並在網路安全和代理式人工智慧工作流程中提供證據驗證。

Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments

2608.16747v1 by Adam Karvonen, Euan Ong, Subhash Kantamneni, Samuel Marks

Many areas of AI research, such as language model interpretability and chain of thought faithfulness, seek to explain model behaviors. But what constitutes a "good" explanation? In this work, we evaluate explanations through the lens of counterfactual simulatability-whether the explanation is useful for predicting model behaviors on related counterfactual inputs. To this end, we introduce CHIVE (Counterfactual Hypothesis Investigation Via Edits), a novel agentic pipeline that identifies unexpected model behaviors in the wild and investigates them with counterfactual prompt edits. This yields thousands of high-quality explanations for naturally-occurring model behaviors along with supporting counterfactual evidence. We apply CHIVE in two ways. First, we evaluate whether common LLM interpretability techniques improve an agent's ability to predict counterfactual model behaviors. Surprisingly, we find no uplift from any of the interpretability techniques studied. Second, we use CHIVE to generate training data. We find that training models to predict outcomes of CHIVE-generated counterfactual experiments generalizes to various out-of-distribution settings. Overall, CHIVE automatically discovers explanations of naturally-occurring LLM behaviors, enabling us to evaluate and improve methods for explaining LLM behaviors.

摘要:許多AI研究領域,例如語言模型的可解釋性和思維鏈的可靠性,旨在解釋模型行為。但什麼構成了「好的」解釋?在這項工作中,我們通過反事實可模擬性的視角來評估解釋——即該解釋是否對預測模型在相關反事實輸入上的行為有用。為此,我們引入了CHIVE(通過編輯進行反事實假設調查),這是一個新穎的主動管道,能夠識別自然環境中意外的模型行為並通過反事實提示編輯進行調查。這產生了數千個高品質的解釋,針對自然發生的模型行為以及支持的反事實證據。我們以兩種方式應用CHIVE。首先,我們評估常見的LLM可解釋性技術是否能提高代理預測反事實模型行為的能力。令人驚訝的是,我們發現在所研究的任何可解釋性技術中都沒有提升。其次,我們使用CHIVE生成訓練數據。我們發現,訓練模型以預測CHIVE生成的反事實實驗的結果能夠泛化到各種分佈外的情境。總體而言,CHIVE自動發現自然發生的LLM行為的解釋,使我們能夠評估和改進解釋LLM行為的方法。

Unsupervised Anomaly Detection for Image Dataset Quality Assurance in Multi-Center Breast MRI

2608.16725v1 by Chiara Tappermann, Steffen Renisch, Lars Ole Schwen, Hans Meine, Horst K. Hahn, Eike Petersen

Corrupted, inconsistent, or anomalous data silently threatens the safety and reliability of medical AI. Despite growing regulatory recognition of dataset quality assurance (QA) for high-risk medical AI, scalable automated detection remains underdeveloped. We employ unsupervised anomaly detection (AD) and out-of-distribution (OOD) detection as an automated dataset QA mechanism for multi-center dynamic contrast-enhanced breast MRI. We build a controlled AD benchmark of 17 realistic QA-relevant anomaly types from six public datasets (protocol violations, processing errors, incorrect anatomical regions) and propose a taxonomy of radiological image anomalies based on human visual perception, enabling fine-grained analysis of AD failure modes. The benchmark includes near-, medium-far-, far-OOD samples, as well as in-distribution and external normal data. Four methods are evaluated: a projection-based method extended with a domain-specific feature extractor and a novel positional encoding, a reconstruction-based approach extended to full 3D volumes with an augmented training objective, and two unmodified hybrid OOD detection methods. Medium-far- and far-OOD samples are detected reliably, whereas near-OOD samples and external normal data from unseen institutions expose method-specific differences. The 3D reconstruction-based approach best balances detection performance (AUROC: 0.936) and generalization to unseen institutions. The projection-based method with positional encoding achieves the highest overall detection performance (AUROC: 0.954). Both hybrid methods exhibit critical failure modes, confirming that methods validated for one modality or anatomy may not generalize without domain-specific adaptation. Implants and mastectomies remain an open challenge for all methods. Our results establish a foundation and practical guidance on scalable unsupervised QA in medical AI pipelines.

摘要:腐敗、不一致或異常的數據默默威脅著醫療人工智慧的安全性和可靠性。儘管對高風險醫療人工智慧數據集質量保證(QA)的監管認識日益增長,但可擴展的自動檢測仍然發展不足。我們採用無監督異常檢測(AD)和分佈外(OOD)檢測作為多中心動態對比增強乳腺MRI的自動數據集QA機制。
我們建立了一個由六個公共數據集中的17種現實QA相關異常類型組成的受控AD基準(協議違規、處理錯誤、不正確的解剖區域),並根據人類視覺感知提出了一個放射影像異常的分類法,使得對AD失效模式的細緻分析成為可能。基準包括近距離、中遠距離、遠距離OOD樣本,以及分佈內和外部正常數據。評估了四種方法:一種基於投影的方法,擴展了特定領域的特徵提取器和新穎的位置編碼;一種基於重建的方法,擴展到完整的3D體積並具有增強的訓練目標;以及兩種未經修改的混合OOD檢測方法。
中遠距離和遠距離OOD樣本的檢測可靠,而近距離OOD樣本和來自未見機構的外部正常數據則顯示出方法特定的差異。基於3D重建的方法在檢測性能(AUROC:0.936)和對未見機構的泛化之間達到了最佳平衡。帶有位置編碼的基於投影的方法實現了最高的整體檢測性能(AUROC:0.954)。兩種混合方法都顯示出關鍵的失效模式,確認了針對一種模態或解剖結構驗證的方法可能無法在沒有特定領域適應的情況下進行泛化。植入物和乳房切除術對所有方法仍然是一個未解決的挑戰。我們的結果為醫療人工智慧管道中的可擴展無監督QA建立了基礎和實用指導。

Degradation-Aligned Self-Supervised Learning for State of Health Estimation of Lithium-Ion Batteries under Label Sparsity

2608.16612v1 by Jiaqi Yao, Julia Kowal

An accurate estimation of the state of health (SOH) underpins a safe and optimized use of the battery system. Although compelling, data-driven SOH estimation models typically require large amounts of high-quality labeled cycling data, while in practice such labels are often sparse in both quantity and coverage. Therefore, in this work, we propose a degradation-aligned self-supervised learning (SSL) framework based on a convolutional neural network-gated recurrent unit (CNN-GRU) model, which learns aging-consistent representations from unlabeled data through a cycle-order ranking objective as the pretext task for pretraining, thereby enabling robust SOH estimation after fine-tuning on sparsely labeled data. Test results showcase that the proposed ranking-based SSL approach proves to endow the pretrained model with degradation-aligned information from unlabeled data, and after fine-tuning the model can carry out accurate, robust SOH estimation, even when only an extremely limited amount of 1% of unevenly distributed labeled training data is available, where the MAE of 1.718% and RMSE of 2.329% can be achieved on the test cell. In addition, in-depth analyses are presented regarding the influences of label distribution of battery degradation data. We believe this work could shed new light on SOH estimation of lithium-ion batteries under label sparsity in real-world applications.

摘要:準確的健康狀態(SOH)估計是安全且優化使用電池系統的基礎。雖然數據驅動的SOH估計模型非常有說服力,但通常需要大量高質量的標註循環數據,而在實際情況中,這些標註往往在數量和覆蓋範圍上都很稀疏。因此,在本研究中,我們提出了一種基於卷積神經網絡-門控遞歸單元(CNN-GRU)模型的降解對齊自監督學習(SSL)框架,通過循環順序排名目標作為預訓練的前置任務,從未標註數據中學習與老化一致的表示,從而在稀疏標註數據上進行微調後實現穩健的SOH估計。測試結果顯示,所提出的基於排名的SSL方法使預訓練模型從未標註數據中獲得了降解對齊的信息,並且在微調後,該模型能夠進行準確且穩健的SOH估計,即使僅有極少量的1%不均勻分佈的標註訓練數據可用,測試電池的MAE可達1.718%和RMSE可達2.329%。此外,還對電池降解數據的標註分佈影響進行了深入分析。我們相信這項工作可以為在現實應用中標註稀疏的鋰離子電池SOH估計提供新的見解。

Physics of Agents: Statistical Mechanics Predicts Collective Behavior of AI Agents

2608.16578v1 by Batu El, Jinhee Paeng, Fatih Dinc, Shiye Su, Mete Erdogan, Aneesh Pappu, Haotian Ye, Wanjia Zhao, Surya Ganguli, James Zou

AI agents increasingly operate as part of interacting systems rather than in isolation. As agents exchange information and jointly make decisions, their interactions can improve collective reasoning but may also produce herding, polarization, or amplify shared biases. Understanding and predicting these collective dynamics is therefore important for designing effective and aligned multi-agent systems. Here, we study over 10,000 communities of language-model agents that repeatedly exchange messages and revise their opinions across objective mathematics questions and subjective political statements. Despite substantial diversity in possible behavior, the individual and group dynamics can be represented by three characteristic regimes: indifference, polarization, and consensus. AI agents start indifferent and build conviction as they interact. On objective questions, communication improves collective accuracy, while on subjective questions it often drifts group opinions toward the right in the political spectrum. We explain these observations with a statistical-mechanics formalism in which agents stochastically favor lower social pressure. Given only initial opinions, our model predicts individual trajectories, outperforms all standard baselines, generalizes to unseen community graphs, and reproduces the observed group archetype distributions. Our fitted model parameters reveal the mechanics underlying our key observations: i) communities operate below the critical social temperature, which explains conviction buildup; ii) attractive ties outweigh repulsive ones, which favors consensus; and iii) agents holding the correct answer exert the strongest pull, which drives truth-seeking. Overall, our results demonstrate that collective behavior of AI agents, like that of other complex systems, follows compact and predictive dynamical laws.

摘要:AI 代理人越來越多地作為互動系統的一部分運作,而不是孤立存在。隨著代理人之間交換信息並共同做出決策,他們的互動可以改善集體推理,但也可能產生跟風、極化或放大共享偏見。因此,理解和預測這些集體動態對於設計有效且一致的多代理系統非常重要。在這裡,我們研究了超過 10,000 個語言模型代理人的社群,它們反覆交換消息並在客觀數學問題和主觀政治陳述上修正自己的意見。儘管可能的行為存在相當大的多樣性,但個體和群體動態可以用三種特徵性狀態來表示:漠不關心、極化和共識。AI 代理人最初是漠不關心的,隨著互動的進行建立信念。在客觀問題上,交流提高了集體準確性,而在主觀問題上,則經常使群體意見向政治光譜的右側漂移。我們用一種統計力學形式主義來解釋這些觀察,其中代理人隨機地偏好較低的社會壓力。僅根據初始意見,我們的模型預測個體軌跡,超越所有標準基準,對未見過的社群圖進行泛化,並重現觀察到的群體原型分佈。我們擬合的模型參數揭示了我們關鍵觀察背後的機制:i) 社群運作在臨界社會溫度以下,這解釋了信念的積累;ii) 吸引性聯繫超過排斥性聯繫,這有利於共識;iii) 持有正確答案的代理人施加最強的影響,這驅動尋求真相。總體而言,我們的結果表明,AI 代理人的集體行為,如同其他複雜系統,遵循緊湊且可預測的動態法則。

Revisiting the Performance of Generative Artificial Intelligence on Introductory Object-Oriented Programming Assessments: Insights from 2026

2608.16318v1 by Marina Lepp, Joosep Kaimre

Recent advances in Generative Artificial Intelligence (GenAI) have substantially improved the ability of large language models (LLMs) to generate and explain source code. However, their performance on authentic object-oriented programming (OOP) assessments remains insufficiently understood. This study evaluates five widely used GenAI systems, ChatGPT-5.2, DeepSeek-V3, Gemini 2.5 Flash, Claude Sonnet 4.5, and M365 Copilot, using programming tests and examination tasks from an introductory university OOP course. The generated solutions were assessed using the same grading criteria applied to students and compared with historical student results from the same course, as well as findings from the previous year. Common errors were also analyzed to identify recurring limitations across models. All evaluated GenAI systems achieved higher scores than the average student cohort and frequently obtained full marks on longer programming tasks. Nevertheless, they occasionally produced non-compiling code and continued to struggle with advanced OOP concepts, particularly interfaces, abstract classes, and certain inheritance-related tasks. Performance was also limited on graphics-related questions involving image interpretation. Compared with the previous year, the evaluated systems demonstrated noticeable improvements across most assessments while exhibiting several recurring error patterns. The findings provide an updated evaluation of the capabilities and limitations of contemporary GenAI systems on authentic introductory OOP assessments. They also offer evidence that can inform the design of programming assessments, the responsible integration of GenAI tools into software engineering education, and future studies evaluating the evolution of AI-assisted programming.

摘要:最近在生成式人工智慧(GenAI)方面的進展顯著提高了大型語言模型(LLMs)生成和解釋源代碼的能力。然而,它們在真實物件導向程式設計(OOP)評估中的表現仍然不夠了解。本研究評估了五個廣泛使用的GenAI系統,ChatGPT-5.2、DeepSeek-V3、Gemini 2.5 Flash、Claude Sonnet 4.5和M365 Copilot,使用來自大學入門OOP課程的程式設計測試和考試任務。生成的解決方案使用與學生相同的評分標準進行評估,並與來自同一課程的歷史學生結果以及前一年的研究結果進行比較。還分析了常見錯誤,以識別模型之間的重複限制。所有評估的GenAI系統的得分均高於平均學生群體,並且在較長的程式設計任務中經常獲得滿分。然而,它們偶爾會產生無法編譯的代碼,並且在高級OOP概念方面仍然存在困難,特別是介面、抽象類別和某些與繼承相關的任務。在涉及圖像解釋的圖形相關問題上,表現也受到限制。與前一年相比,評估的系統在大多數評估中顯示出明顯的改進,同時表現出幾個重複的錯誤模式。這些發現提供了對當代GenAI系統在真實入門OOP評估中的能力和限制的最新評估。它們還提供了可以為程式設計評估的設計、負責任地將GenAI工具整合到軟體工程教育中,以及未來評估AI輔助程式設計演變的研究提供依據的證據。

Foresight-England: Development of a National-Scale Generative AI Model of Electronic Health Records for Medical Event Prediction across the COVID-19 Pandemic

2608.16273v1 by Simon Ellershaw, Christopher Tomlinson, Zeljko Kraljevic, Spiros Denaxas, Harry Hemingway, Cathie Sudlow, Angela M. Wood, Anoop D. Shah, Richard Dobson

Foresight-England (Foresight-E) is the first national-scale generative foundation model of electronic health records (EHRs), developed as a research pilot strictly for COVID-19 research. We evaluated its ability to model the direct and indirect effects of the pandemic. Trained from scratch entirely within the NHS England Secure Data Environment, Foresight-E is a 243-million-parameter transformer decoder. It was trained and evaluated on de-identified, longitudinal EHRs of approximately 61 million individuals, integrating primary/secondary care, death registrations, and COVID-19 data. Training and validation used a 90% subset (54.9 million) spanning November 2018 to December 2022; the remaining 10% (6.1 million) was held out for evaluation. Foresight-E models patient timelines autoregressively, predicting the next medical event given their prior history. At inference, it operates zero-shot, predicting any concept in its ~40,000-code vocabulary without task-specific training. Our tokenisation scheme retains the clinical granularity of ICD-10, OPCS-4, and SNOMED CT codes, jointly representing absolute and relative timing. We designed an evaluation framework for 30-day COVID-19 hospitalisation and mortality, including subgroup analyses by demographic factors and vaccination status. To assess generalisation to unseen future data and the pandemic's indirect effects, we tested the model on medical events from 2023 (beyond its training period), benchmarking against logistic regression and XGBoost. As detailed in the Project Status section, NHS England has paused access to data for the Foresight-E project, meaning quantitative results are currently unavailable. Instead, we share our strategy for tokenisation, architecture, training, inference, and evaluation as a methodological template and case study in the challenges of building population-scale EHR foundation models.

摘要:Foresight-England (Foresight-E) 是首個全國規模的電子健康紀錄 (EHRs) 生成基礎模型,作為針對 COVID-19 研究的研究試點而開發。我們評估了它建模疫情直接和間接影響的能力。Foresight-E 完全在 NHS England 安全數據環境中從零開始訓練,是一個擁有 2.43 億參數的Transformer解碼器。它在約 6100 萬人的去識別化、縱向 EHRs 上進行訓練和評估,整合了初級/次級護理、死亡登記和 COVID-19 數據。訓練和驗證使用了 90% 的子集(5490 萬),涵蓋了 2018 年 11 月到 2022 年 12 月;剩餘的 10%(610 萬)則保留用於評估。Foresight-E 自回歸地建模患者時間線,根據其先前的歷史預測下一個醫療事件。在推理時,它以零樣本操作,預測其約 40,000 種代碼詞彙中的任何概念,而無需特定任務的訓練。我們的標記方案保留了 ICD-10、OPCS-4 和 SNOMED CT 代碼的臨床細節,聯合表示絕對和相對時間。我們設計了一個評估框架,用於 30 天 COVID-19 住院和死亡率,包括按人口統計因素和疫苗接種狀態的子群分析。為了評估對未見未來數據的泛化能力和疫情的間接影響,我們在 2023 年的醫療事件上測試了該模型(超出其訓練期間),並與邏輯回歸和 XGBoost 進行基準比較。正如項目狀態部分詳細說明的那樣,NHS England 已暫停對 Foresight-E 項目的數據訪問,這意味著目前無法獲得定量結果。相反,我們分享了我們的標記化、架構、訓練、推理和評估的策略,作為建立人口規模 EHR 基礎模型挑戰的 методологический шаблон и кейс-исследование。

Defake-o3: From Speculative Rationales to Verifiable Evidence for Explainable AIGI Detection

2608.16259v1 by Bowen Deng, Jiahui Zhan, Yikun Ji, Haozhen Yan, Jianfu Zhang

The rapid progress of image generation models calls for AI-generated image (AIGI) detectors that are not only accurate but also explainable and reliable. While MLLM-based detectors can provide natural language explanations, existing methods often generate speculative rationales: they rely on vague or hallucinated artifacts, miss subtle localized flaws from the latest generators, and fail to provide evidence that can be visually verified. We present Defake-o3, an explainable AIGI detector that moves from speculative rationales to verifiable evidence. It combines interactive visual search with verifier-guided evidence alignment: the model iteratively zooms into suspicious regions to inspect fine-grained details, while an Evidence Verifier, trained from human verification annotations, provides reinforcement learning rewards that favor grounded evidence and penalize baseless claims. To support this objective, we construct GroundFake, a dataset designed for grounded explainable detection, with localized bounding-box evidence, human verification based on visual grounding and artifact specificity, corrected reasoning trajectories, and valid/invalid evidence supervision. We further introduce FakeFrontier, an out-of-distribution benchmark built from real images and outputs of 10 recent generators, together with an MLLM-based protocol for evaluating evidence quality and persuasiveness. Experiments on GroundFake, FakeFrontier, and additional out-of-distribution benchmarks show that Defake-o3 improves both detection accuracy and explanation quality, producing more localized, verifiable, and persuasive evidence.

摘要:快速進展的圖像生成模型呼喚不僅準確而且可解釋和可靠的AI生成圖像(AIGI)檢測器。雖然基於MLLM的檢測器可以提供自然語言解釋,但現有方法往往生成推測性的理由:它們依賴模糊或虛幻的工件,錯過最新生成器的微妙局部缺陷,並未能提供可以視覺驗證的證據。我們提出了Defake-o3,一個可解釋的AIGI檢測器,從推測性理由轉向可驗證的證據。它結合了互動式視覺搜索和驗證者引導的證據對齊:模型迭代地放大可疑區域以檢查細緻的細節,而一個基於人類驗證註釋訓練的證據驗證器提供強化學習獎勵,以支持有根據的證據並懲罰無根據的主張。為了支持這一目標,我們構建了GroundFake,一個為有根據的可解釋檢測設計的數據集,包含局部邊界框證據、基於視覺定位和工件特異性的人工驗證、修正的推理軌跡,以及有效/無效證據的監督。我們進一步介紹了FakeFrontier,一個基於真實圖像和10個最近生成器輸出的分佈外基準,並附帶一個基於MLLM的協議,用於評估證據的質量和說服力。在GroundFake、FakeFrontier和其他分佈外基準上的實驗顯示,Defake-o3提高了檢測準確性和解釋質量,產生了更局部、可驗證和具說服力的證據。

CompoSkill: Compositional Skill Chain Attacks from Individually Scanner-Passing LLM Agent Skills

2608.16246v1 by Mingxiao Liu, Zhoumian Jiang, Jianan Ma, Jian Zhang, Jialuo Chen, Xinhao Deng, Zhen Wang

Autonomous AI agents tackling Long Horizon Tasks depend on marketplace skills that are certified one at a time: a scanner returns a safety verdict for each skill and declares the ecosystem safe if every package passes. We show that this assumption fails under skill composition. A skill may pass the per-skill scanner individually yet participate in a risky composition when an agent connects its outputs, capabilities, or side effects with those of other scanner-passing skills. This makes skill composition risk a path level property rather than a node level property, explaining why existing skill scanners that inspect individual packages achieve limited interception. To study this threat, we present CompoSkill, a framework that constructs skill composition attacks through a dual attacker system. The white-box attacker knows the victim's installed skill pool and directly injects explicit skill-id sequences; the black-box attacker knows only a role profile, downloads the top marketplace skills for that scenario, builds a Skill Composition Graph, and searches for high risk chains whose implicit lures never name skill identifiers. We further construct CompoSkill-Bench, a benchmark of 1,140 records built from long-horizon professional workflows across five threats and six scenarios on OpenClaw and Nanobot. CompoSkill achieves risk Chain Formation Rates (CFR) up to 83.3% in the white box setting and 80.6% in the black box setting, while existing skill scanners block only a limited fraction of the risky compositions. Finally, we observe a bridge-bonus-then-hop-decay pattern: a bridge skill can increase attack success, but Attack Success Rate (ASR) decreases once additional hops make the risk chain longer than three skills. These results expose a systematic gap in single skill certification for autonomous AI agents.

摘要:自主 AI 代理處理長期任務依賴於逐一認證的市場技能:掃描器為每項技能返回安全判決,並在每個包裝通過時聲明生態系統安全。我們顯示這一假設在技能組合下失效。某項技能可能在每項技能掃描器中單獨通過,但當代理將其輸出、能力或副作用與其他通過掃描器的技能連接時,可能參與風險組合。這使得技能組合風險成為路徑級別的特性,而非節點級別的特性,解釋了為什麼現有的技能掃描器僅檢查單個包裝而達到有限的攔截效果。為了研究這一威脅,我們提出了 CompoSkill,一個通過雙重攻擊者系統構建技能組合攻擊的框架。白盒攻擊者知道受害者的已安裝技能池,並直接注入明確的技能 ID 序列;黑盒攻擊者僅知道一個角色配置,下載該場景的頂級市場技能,構建技能組合圖,並搜索隱含誘餌從未命名技能標識符的高風險鏈。我們進一步構建了 CompoSkill-Bench,一個基於 OpenClaw 和 Nanobot 的五種威脅和六種場景中,從長期專業工作流程中建立的 1,140 條記錄的基準。CompoSkill 在白盒設置中達到高達 83.3% 的風險鏈形成率 (CFR),在黑盒設置中達到 80.6%,而現有的技能掃描器僅阻止有限比例的風險組合。最後,我們觀察到一種橋接-獎勵-然後跳躍衰減的模式:橋接技能可以提高攻擊成功率,但一旦額外的跳躍使風險鏈超過三項技能,攻擊成功率 (ASR) 就會下降。這些結果揭示了自主 AI 代理在單一技能認證方面的系統性缺口。

When Single-Dataset Conclusions Fail: A 45-Task Study of Threshold Tuning and Resampling for Imbalanced Classification

2608.16147v1 by Diyorbek Musaev

Class-imbalance handling is routinely evaluated on a single benchmark dataset, and the resulting conclusions are reported as if they were properties of the method. We show this practice is unsafe. On the public Kaggle credit-card fraud dataset, under a leakage-free nested cross-validation protocol in which the decision threshold is selected on a held-out inner validation fold, a plain Random Forest at the default 0.5 threshold attains F1 = 0.861 +/- 0.021, and threshold tuning yields it no benefit (delta-F1 = -0.002). Read alone, this supports an appealing conclusion: for a well-calibrated ensemble, imbalance handling is unnecessary. We then apply the identical protocol to 45 binary tasks spanning imbalance ratios from 1:1.5 to 1:178 (2,025 model fits, four model families). The conclusion reverses. Random Forest benefits most from threshold tuning across the suite (delta-F1 = +0.101 +/- 0.134), not least, while three other families replicate their fraud-dataset behaviour almost exactly. SMOTE likewise harms the fraud dataset but helps across the suite (mean delta-F1 = +0.076; 138 wins, 39 losses; Wilcoxon p = 2.7e-17). Two further results. Threshold-tuning benefit is non-monotonic in the imbalance ratio: near zero below 1:5, peaking at +0.120 in the 1:15-1:40 band, declining to +0.045 beyond 1:100 - explaining why the fraud dataset, at 1:577, is an unrepresentative place to study the question. And we reject an intuitive heuristic: validation-set calibration error does not predict tuning benefit (expected calibration error r = -0.087; Brier r = +0.137), so calibration diagnostics cannot tell a practitioner whether tuning is worthwhile. We release the protocol, the 45-task harness, and all per-run metrics.

摘要:類別不平衡處理通常在單一基準數據集上進行評估,並且所得到的結論被報告為該方法的特性。我們顯示這種做法是不安全的。在公共的Kaggle信用卡詐騙數據集上,在一個無泄漏的嵌套交叉驗證協議下,決策閾值是在保留的內部驗證折上選擇的,普通的隨機森林在默認的0.5閾值下達到F1 = 0.861 +/- 0.021,而閾值調整對其沒有益處(delta-F1 = -0.002)。單獨閱讀這一點支持了一個吸引人的結論:對於一個良好校準的集成,處理不平衡是沒有必要的。我們隨後將相同的協議應用於45個二元任務,涵蓋不平衡比率從1:1.5到1:178(2,025個模型擬合,四個模型系列)。結論顛倒了。隨機森林在整個系列中最受益於閾值調整(delta-F1 = +0.101 +/- 0.134),而且三個其他系列幾乎完全複製了它們在詐騙數據集上的行為。SMOTE同樣對詐騙數據集有害,但在整個系列中有幫助(平均delta-F1 = +0.076;138次勝利,39次失敗;Wilcoxon p = 2.7e-17)。另外兩個結果。閾值調整的益處在不平衡比率中是非單調的:在1:5以下接近零,在1:15-1:40區間達到+0.120的峰值,超過1:100則下降至+0.045——這解釋了為什麼詐騙數據集在1:577的情況下是一個不具代表性的研究該問題的地方。我們還拒絕了一個直觀的啟發式:驗證集的校準誤差並不預測調整的益處(預期校準誤差r = -0.087;Brier r = +0.137),因此校準診斷無法告訴從業者調整是否值得。我們釋放了該協議、45任務的框架以及所有每次運行的指標。

TokenSTFormer: A Tokenized Spatial-temporal Attention Model for Holistic Motion Analysis in Adolescent Idiopathic Scoliosis Screening

2608.16122v1 by Dong Chen, Kenneth M. C. Cheung

Adolescent Idiopathic Scoliosis (AIS) is a prevalent spinal deformity in adolescents that, if left untreated, can result in severe health outcomes. Traditional screening methods are limited by subjective interpretation, reliance on professional expertise and low scalability. To address these challenges, we present ScoliGait dataset, which comprises 1,516 gait video clips paired with corresponding X-ray records. We also introduce TokenSTFormer, a novel model that tokenizes spatial and temporal semantics to enhance feature representation and convergence. Our model achieves state-of-the-art performance, surpassing vanilla Vision Transformer encoder across key metrics, including accuracy of 0.79. This study highlights the potential of leveraging holistic motion features derived from gait video and attention-based models for scalable, cost-effective AIS screening, paving the way for future clinical applications in scoliosis detection.

摘要:青少年特發性脊柱側彎(AIS)是青少年中常見的脊柱畸形,如果不加以治療,可能會導致嚴重的健康後果。傳統的篩檢方法受到主觀解釋、依賴專業知識和低可擴展性的限制。為了解決這些挑戰,我們提出了 ScoliGait 數據集,其中包含 1,516 段行走視頻片段,並配有相應的 X 光記錄。我們還介紹了 TokenSTFormer,一種新型模型,將空間和時間語義進行標記化,以增強特徵表示和收斂。我們的模型在關鍵指標上達到了最先進的性能,超越了普通的視覺Transformer編碼器,包括 0.79 的準確率。本研究突顯了利用從行走視頻中獲得的整體運動特徵和基於注意力的模型進行可擴展、成本效益高的 AIS 篩檢的潛力,為未來脊柱側彎檢測的臨床應用鋪平了道路。

Eigenanalysis framework for autoregressive neural emulators of multi-scale chaotic dynamics

2608.16084v1 by Conrad Ainslie, Pedram Hassanzadeh, Michael W. Mahoney, Ashesh Chattopadhyay

Neural autoregressive models have rapidly emerged as powerful emulators of high-dimensional chaotic systems, yet their long-term instability and error growth remain poorly understood, leading to ad-hoc solutions. Here, we develop an eigenanalysis framework that reveals the dynamical origin of this error growth. By analyzing the Jacobian of the learned one-step update map with respect to the state, we show how inference-time error growth, and thus model stability, is governed by its spectral radius. Direct-step architectures (models that predict the next state from the previous one) generically admit unstable eigenvalues with magnitudes exceeding one, explaining the rapid divergence of these widely used models. In contrast, integration-constrained models (where the time derivative is estimated and integrated with a higher-order integrator) collapse their eigenspectrum onto the unit circle, yielding neutral stability and a universal linear error-scaling law. The largest eigenvalue of this Jacobian provides an architecture-agnostic, a priori diagnostic of short-term skill, long-term stability, and spectral bias, without requiring an expensive rollout. Leveraging this theory, we introduce a stability-promoting loss that explicitly regularizes Jacobian-driven error amplification, improving both forecast accuracy and dynamical robustness. Demonstrated across $29$ models spanning two architectures, several explicit and implicit integrators, and multiple loss functions on the Kuramoto-Sivashinsky system, our results establish a theoretical foundation for the design and evaluation of neural emulators of chaotic multi-scale dynamics. More broadly, our framework is a step toward the kind of a priori stability analysis that numerical analysis provides for discretizations of differential equations and that scientific machine learning currently lacks.

摘要:神經自回歸模型迅速崛起,成為高維混沌系統的強大模擬器,但其長期不穩定性和誤差增長仍然不甚了解,導致了臨時解決方案。在這裡,我們開發了一個特徵分析框架,揭示了這種誤差增長的動力學起源。通過分析學習到的一步更新映射的雅可比矩陣相對於狀態的情況,我們展示了推理時誤差增長以及模型穩定性是如何受到其譜半徑的控制。直接步驟架構(從前一狀態預測下一狀態的模型)通常承認具有超過一的幅度的不穩定特徵值,解釋了這些廣泛使用模型的快速發散。相反,集成約束模型(在此模型中,時間導數被估計並與高階積分器進行積分)將其特徵譜壓縮到單位圓上,產生中性穩定性和普遍的線性誤差縮放法則。這個雅可比矩陣的最大特徵值提供了一種與架構無關的、事先的短期技能、長期穩定性和譜偏差的診斷,而無需昂貴的展開。利用這一理論,我們引入了一種促進穩定性的損失,明確正則化雅可比驅動的誤差放大,改善預測準確性和動力學穩健性。在涵蓋兩種架構、幾種顯式和隱式積分器以及多種損失函數的 $29$ 個模型中,我們的結果為設計和評估混沌多尺度動力學的神經模擬器建立了理論基礎。更廣泛地說,我們的框架是朝著數值分析為微分方程的離散化提供的那種事先穩定性分析邁出的一步,而這是當前科學機器學習所缺乏的。

NICE: Scale-Stable Perturbations for Graph Neural Network Explanations via Noise Corruption

2608.16038v2 by Ziluowen Luo, Jun Yin, Ruochen Liu, Ming Cheng, Shirui Pan, Chengqi Zhang, Senzhang Wang

Post-hoc Graph Neural Network (GNN) explainers commonly follow a Perturb-Query paradigm, inferring the importance of graph elements based on queried predictions to perturbed inputs. However, such perturbations often introduce substantial distribution shift, undermining the reliability of the queried predictions used to derive explanations. While existing efforts mainly improve perturbed graphs or stabilize model predictions on them, we revisit the perturbation mechanism itself. We show that the widely used Element-wise Masking(EM) suppresses edge-induced messages toward zero, causing deterministic scale contraction that accumulates across message-passing layers, a phenomenon we term Scale Drift. Consequently, prediction changes under EM may conflate information corruption with deviations in propagation scale. As a scale-stable alternative to EM, we introduce Noise Corruption (NC), which perturbs each message through matched-norm random-direction corruption while preserving the expected squared message norm. Building on NC, we propose NICE, a Noise Corruption-based explanation framework, which learns a Stochastic Restoration Boundary (SRB) under NC-induced uncertainty, balancing target-prediction restoration against compactness. Furthermore, Boundary-Integrated Gradient (BIG) converts this boundary into edge attributions by accumulating each edge's contribution to reducing restoration risk along the restoration path. Experiments across multiple benchmarks demonstrate stronger explanation performance and model faithfulness while confirming that NC substantially reduces the Scale Drift induced by masking.

摘要:後 hoc 圖神經網絡 (GNN) 解釋器通常遵循擾動-查詢範式,根據對擾動輸入的查詢預測推斷圖元素的重要性。然而,這種擾動往往會引入顯著的分佈變化,削弱用於推導解釋的查詢預測的可靠性。雖然現有的努力主要改善擾動圖或穩定模型在其上的預測,但我們重新審視擾動機制本身。我們展示了廣泛使用的逐元素掩蔽 (EM) 將邊緣引起的消息壓制至零,導致確定性的縮放收縮,這一現象在消息傳遞層中累積,我們稱之為縮放漂移。因此,EM 下的預測變化可能將信息損壞與傳播縮放的偏差混淆在一起。作為 EM 的一種穩定縮放替代方案,我們引入了噪聲損壞 (NC),它通過匹配範數的隨機方向擾動每條消息,同時保持期望的平方消息範數。基於 NC,我們提出了 NICE,一種基於噪聲損壞的解釋框架,它在 NC 引起的不確定性下學習隨機恢復邊界 (SRB),平衡目標預測恢復與緊湊性。此外,邊界整合梯度 (BIG) 通過累積每條邊對減少恢復風險的貢獻,將這一邊界轉換為邊緣歸因。在多個基準上的實驗顯示出更強的解釋性能和模型忠實度,同時確認 NC 顯著減少了由掩蔽引起的縮放漂移。

2608.15731v1 by Haadia Amjad, Ronald Tetzlaff

Deep Neural Networks (DNNs) deployed in high-risk domains, such as healthcare and autonomous driving, must be not only accurate but also understandable to ensure user trust. In real-world computer vision tasks, these models often operate on complex images containing background noise and are heavily annotated. To make such models explainable, Concept-based Explainable AI (CXAI) methods need to be assessed for their applicability and problem-solving capacity. In this work, we explore CXAI use cases in multi-label classification by training two DNNs, VGG16 and ResNet50, on the 20 most annotated labels in the MS-COCO dataset (Microsoft Common Objects in Context). We apply two CXAI methods, CRP (Concept Relevance Propagation) and CRAFT (Concept Recursive Activation FacTorization), to generate concept-level explanations and investigate the overall evaluations. Our analysis reveals three key findings: (1) CXAI highlights learning weaknesses in DNNs, (2) higher concept distinctiveness reduces label and concept confusion, and (3) environmental concepts expose dataset-induced biases. Our results demonstrate the potential of CXAI to enhance the understanding of model generalizability and to diagnose bias instigated by the dataset.

摘要:深度神經網絡(DNNs)在高風險領域(如醫療保健和自動駕駛)的應用,不僅必須準確,還必須可理解,以確保用戶信任。在現實世界的計算機視覺任務中,這些模型通常在包含背景噪聲且標註繁重的複雜圖像上運行。為了使這些模型具有可解釋性,需要評估基於概念的可解釋人工智慧(CXAI)方法的適用性和解決問題的能力。在本研究中,我們通過在 MS-COCO 數據集(微軟上下文中的常見物體)上訓練兩個 DNN(VGG16 和 ResNet50),探索 CXAI 在多標籤分類中的應用案例,該數據集包含 20 個標註最多的標籤。我們應用兩種 CXAI 方法,CRP(概念相關性傳播)和 CRAFT(概念遞歸激活因子分解),以生成概念級別的解釋並調查整體評估。我們的分析揭示了三個關鍵發現:(1)CXAI 突出了 DNN 的學習弱點,(2)較高的概念區別性減少了標籤和概念的混淆,以及(3)環境概念揭示了數據集引起的偏見。我們的結果展示了 CXAI 在增強模型可泛化性理解和診斷由數據集引發的偏見方面的潛力。

Large Models for Small Devices: Recent Advances and Empirical Analysis of Edge AI Deployment

2608.15693v1 by Subhransu Das, Jiaming Cheng, Arnav Kumar, Sadia Afrose, Mingzhe Han, Michael Silagy, Shreya Palande, Brijesh Soni, Rajiv Ramnath

Running large AI models on resource-constrained edge devices requires model compression to reduce model size and computation. What compresses well, however, need not deploy well. We survey dozens of recent works that report compression results on real hardware and extract practical deployment guidelines from them. Following these guidelines, we deploy compact language and image models on GPU, CPU, and Raspberry Pi platforms across question answering and image segmentation. No single technique wins across tasks. For question answering, Qwen3.5 0.8B reaches 93.85 SQuAD F1 and 92 EM under Q5_K_M GGUF quantization, while structured pruning at the same precision costs 16 F1 at a 1% ratio. For segmentation, the ranking reverses: default quantization leaves parameters and MACs unchanged, whereas pruning cuts model size by nearly 80% at near-constant mIoU. Pruning can even inflate the deployed artifact by 21-49% by breaking k-quant super-block alignment; combined with longer, less format-compliant outputs, this raises Raspberry Pi latency up to 3.4x. Compression can also manufacture the appearance of competence rather than destroy it visibly: one LoRA-recovered variant stays fully parseable and holds 71% strict BoolQ accuracy while sending 97 of 100 predictions to a single class, at 52.6% balanced accuracy. We explain these effects through neural-flow graph analysis and prefill-decode-level latency decomposition, and condense them into task-specific deployment research directions. The right technique depends on the task, the model, and the hardware. Our experiment code and artifacts are open-sourced at https://github.com/Arnavvvkumar/deployment

摘要:在資源受限的邊緣設備上運行大型 AI 模型需要模型壓縮,以減少模型大小和計算量。 然而,壓縮效果良好的模型不一定能夠良好部署。 我們調查了數十篇最近的研究,這些研究報告了在實際硬體上進行的壓縮結果,並從中提取實用的部署指南。 根據這些指南,我們在 GPU、CPU 和 Raspberry Pi 平台上部署了緊湊的語言和圖像模型,應用於問題回答和圖像分割。 沒有單一技術在所有任務中都能獲勝。 在問題回答中,Qwen3.5 0.8B 在 Q5_K_M GGUF 量化下達到 93.85 的 SQuAD F1 和 92 的 EM,而在相同精度下的結構化剪枝則以 1% 的比例損失 16 的 F1。 對於分割,排名則顛倒:默認量化保持參數和 MACs 不變,而剪枝則在幾乎不變的 mIoU 下將模型大小減少近 80%。 剪枝甚至可以通過打破 k-quant 超塊對齊來使部署的工件膨脹 21-49%;結合更長且格式不合規的輸出,這使得 Raspberry Pi 的延遲增加至 3.4 倍。 壓縮還可以製造出能力的外觀,而不是明顯地摧毀它:一個 LoRA 恢復的變體保持完全可解析,並在將 100 次預測中的 97 次發送至單一類別的同時,保持 71% 的嚴格 BoolQ 準確率,平衡準確率為 52.6%。 我們通過神經流圖分析和預填充解碼級延遲分解解釋這些效果,並將其濃縮為特定任務的部署研究方向。 正確的技術取決於任務、模型和硬體。 我們的實驗代碼和工件已在 https://github.com/Arnavvvkumar/deployment 開源。

NumerosityVLM: A Cognitively Inspired Benchmark for Interpreting Numerosity Representations in Vision-Language Models

2608.15425v1 by Yiming Fu, Fangjun Li, Xiujin Liu, Ruidong Ma, Hang Yu, Zhichen Lu, Kanwei He, Alessandro Di Nuovo, Angelo Cangelosi, Zhegong Shangguan

Vision-language models (VLMs) achieve strong performance on high-level multimodal tasks, yet numerosity perception, a cognitive ability that emerges in human infants before language acquisition, remains poorly understood in current models, as existing counting benchmarks entangle numerosity with correlated visual factors. We introduce a cognitively inspired diagnostic benchmark, NumerosityVLM, comprising 10,800 synthetic images across six controlled conditions. The benchmark orthogonally manipulates object size, spatial arrangement, and numerosity, while progressively ablating texture, shape, and color. Evaluating seven VLMs in a zero-shot setting, multi-factor analysis reveals that model architecture explains the largest proportion of performance variance (partial $ω^{2}=0.325$), far exceeding visual conditions. Layer-wise probing further shows that linearly separable numerosity signals consistently emerge at early stages of the vision encoder, while performance differences across evaluated models are primarily associated with the language model component. Code and data are publicly available at https://github.com/fuy3/NumerosityVLM-Benchmark, and https://huggingface.co/datasets/fuy3/NumerosityVLM.

摘要:視覺-語言模型(VLMs)在高階多模態任務上表現出色,然而數量感知這一認知能力在語言習得之前便出現在人類嬰兒身上,卻在當前模型中仍然不甚了解,因為現有的計數基準將數量感知與相關的視覺因素混合在一起。我們引入了一個受到認知啟發的診斷基準,NumerosityVLM,包含10,800張合成圖像,涵蓋六種受控條件。該基準正交地操控物體大小、空間排列和數量感知,同時逐步去除紋理、形狀和顏色。在零樣本設置中評估七個VLM,進行多因素分析顯示,模型架構解釋了性能變異的最大比例(部分 $ω^{2}=0.325$),遠遠超過視覺條件。逐層探測進一步顯示,線性可分的數量感知信號在視覺編碼器的早期階段持續出現,而評估模型之間的性能差異主要與語言模型組件相關。代碼和數據可在 https://github.com/fuy3/NumerosityVLM-Benchmark 和 https://huggingface.co/datasets/fuy3/NumerosityVLM 獲得。

ETHOS: Towards a Modular Ethics Framework for Clinical Multi-Agent Systems

2608.15424v1 by Rakesh Sharma, Sydney Pugh, Cameron Beeche, Pankhuri Singhal, Rachel Wu, Margaret Eby, Jeffrey Duda, James Gee, Kyra O'Brien, Hersh Sagreiya, Marina Serper, Victoria Gershuni, Angela Bradbury, Anurag Verma, Eric Eaton, Kevin B. Johnson, Walter Witschey

The rapid adoption of large language models has enabled the development of clinical multi-agent systems (MAS) capable of integrating multimodal patient data and supporting increasingly complex clinical decision-making. However, the deployment of these systems in real-world healthcare settings raises critical ethical concerns related to safety, fairness, accountability, transparency, and patient trust. While numerous organizations, including the World Health Organization, the National Academy of Medicine, and the FUTURE-AI consortium, have proposed ethical frameworks and governance principles for healthcare AI, these efforts remain largely conceptual. To address this challenge, we present ETHOS (Ethics and Trust through Hierarchical Oversight System), a modular ethics framework designed as a governance meta-agent that can be integrated with any existing multi-agent system without requiring changes to its underlying architecture. ETHOS translates stakeholder-informed ethical requirements into executable runtime oversight through a layered governance approach consisting of deterministic checks, contextual reviews, and a final ethics critic. These components continuously evaluate intermediate reasoning steps and final outputs, enabling the system to identify ethical risks, request revisions, or suppress responses that fail predefined safety and trustworthiness criteria. We demonstrate ETHOS within a hepatology clinical decision-support MAS. Results show that ETHOS improves decision reliability by detecting incomplete, inconsistent, or out-of-scope evidence and appropriately increasing abstention when safe recommendations cannot be supported. By embedding ethical governance directly into system operation, ETHOS provides a practical and auditable mechanism for transforming high-level AI ethics principles into deployable safeguards.

摘要:大型語言模型的快速採用使得臨床多代理系統(MAS)的發展成為可能,這些系統能夠整合多模態病人數據並支持日益複雜的臨床決策。然而,這些系統在現實世界醫療環境中的部署引發了與安全、公平、問責、透明度和病人信任相關的重大倫理問題。儘管包括世界衛生組織、國家醫學院和FUTURE-AI聯盟在內的許多組織已經提出了針對醫療AI的倫理框架和治理原則,但這些努力仍然主要是概念性的。為了解決這一挑戰,我們提出了ETHOS(通過分層監督系統實現倫理與信任),這是一個模塊化的倫理框架,設計為一個治理元代理,可以與任何現有的多代理系統集成,而無需改變其底層架構。ETHOS將利益相關者所知的倫理要求轉化為可執行的運行時監督,通過一種分層治理方法,包括確定性檢查、上下文審查和最終倫理評估。這些組件持續評估中間推理步驟和最終輸出,使系統能夠識別倫理風險、請求修訂或抑制不符合預定安全和可信標準的回應。我們在一個肝病臨床決策支持MAS中展示了ETHOS。結果顯示,ETHOS通過檢測不完整、不一致或超出範疇的證據來提高決策的可靠性,並在無法支持安全建議時適當地增加放棄。通過將倫理治理直接嵌入系統運作中,ETHOS提供了一種實用且可審計的機制,將高層次的AI倫理原則轉化為可部署的保障措施。

When AI Rewrites, Classifiers Relax: Uncertainty-Aware Sentiment Analysis on Sarcastic and AI-Paraphrased Social Text

2608.15338v1 by Shresth Shroff

Sentiment classifiers are increasingly applied to social media content that is either sarcastic or AI-generated --- two distributional regimes where standard evaluations offer little guidance. We present a three-part empirical study of sentiment classifier behaviour under these conditions. First, we find that confidence scores on sarcastic text are significantly lower than on non-sarcastic text (Mann--Whitney $p = 2 \times 10^{-6}$), confirming that classifiers sense their own uncertainty on ironic content even without explicit uncertainty modelling. Second, and counterintuitively, we show that sentiment classifiers achieve higher accuracy on AI-paraphrased reviews than on the original human-authored text (RoBERTa: $+5.8$ pp for Qwen3.5-4B paraphrases, $+3.7$ pp for Gemma4-E4B), revealing a cross-domain stylistic alignment effect: AI paraphrases remove distributional noise that confounds Twitter-trained classifiers, producing cleaner, more prototypical sentiment text. Third, we demonstrate that a lightweight abstention wrapper --- flagging the $14\%$ of inputs with confidence below $0.6$ --- improves accuracy from 82.2\% to 88.9\% ($+6.7$ pp) on the retained set. We further compare Semantic Entropy and MC-Dropout-style disagreement as uncertainty signals and find near-identical AUROC ($0.650$ vs.\ $0.646$) on sarcastic text, suggesting that for short social media inputs, both methods are interchangeable. Our results motivate a shift from confident single-label prediction to uncertainty-aware abstention in high-stakes sentiment applications such as mental health flagging and content moderation.

摘要:情感分類器越來越多地應用於諷刺或 AI 生成的社交媒體內容——這兩種分佈模式下,標準評估提供的指導有限。我們提出了一項三部分的實證研究,探討情感分類器在這些條件下的行為。首先,我們發現對於諷刺文本的信心分數顯著低於非諷刺文本(Mann--Whitney $p = 2 \times 10^{-6}$),確認分類器即使在沒有明確不確定性建模的情況下,也能感知到對於諷刺內容的自身不確定性。其次,反直覺的是,我們顯示情感分類器在 AI 改寫的評論上取得的準確率高於原始的人類撰寫文本(RoBERTa: $+5.8$ pp 對於 Qwen3.5-4B 改寫,$+3.7$ pp 對於 Gemma4-E4B),揭示了一種跨領域的風格一致性效應:AI 改寫去除了困擾 Twitter 訓練的分類器的分佈噪音,產生了更乾淨、更原型的情感文本。第三,我們證明了一種輕量級的棄權包裝——標記信心低於 $0.6$ 的 $14\%$ 輸入——在保留數據集上將準確率從 82.2\% 提高到 88.9\%($+6.7$ pp)。我們進一步比較語義熵和 MC-Dropout 風格的不一致作為不確定性信號,發現對於諷刺文本,兩者的 AUROC 幾乎相同($0.650$ 對 $0.646$),這表明對於短的社交媒體輸入,這兩種方法是可以互換的。我們的結果促使從自信的單標籤預測轉向在高風險情感應用中,如心理健康標記和內容審核,意識到不確定性的棄權。

Demographic Injection in Medical Language Models under Diversity, Equity, and Inclusion Prompts

2608.15254v1 by Diego Mardian, Frank Liu

Clinical-AI guidance increasingly recommends prompting language models to reason with attention to diversity, equity, and inclusion (DEI). We measure a side effect that misrepresents patients: a one-sentence DEI prompt appended to a medical question leads models to add patient demographic attributes (race, socioeconomic status, sex) the question never stated, in effect rewriting who the patient is. We call this demographic injection. Across 47 models, four medical benchmarks, and 376,000 responses scored by a validated model-judge pipeline, a single DEI prompt raises the injection rate from 0.7% to 33.1% (47x) in all 47 of 47 models, attributable to the equity content rather than to added length (18x above a length-matched control; p=1.4x10^-14). Most added content is a general population statement that leaves the answer unchanged, but a smaller subset attaches an attribute to the specific patient or changes the selected option (0.25-2.4% of responses, 99.8% toward the incorrect option), where the invented demographic changes the answer the model recommends. Phrasing scales the effect from 14% to 56%. DEI prompts are just one example of a more general mechanism. Any instruction that nudges how a model reasons can make it add unrequested details, including details about the patient. Flagged outputs are treated as model errors under study, not clinical guidance.

摘要:臨床人工智慧指導越來越多地建議促使語言模型在考慮多樣性、公平性和包容性(DEI)時進行推理。我們測量了一種誤導患者的副作用:一個附加在醫療問題上的單句DEI提示會導致模型添加問題中從未提到的患者人口統計屬性(種族、社會經濟地位、性別),實際上重寫了患者的身份。我們稱之為人口統計注入。在47個模型、四個醫療基準和376,000個由經過驗證的模型評判管道評分的回應中,單一的DEI提示使得注入率從0.7%上升到33.1%(47倍),這是由於公平性內容而非增加的長度(在長度匹配的對照組中增加了18倍;p=1.4x10^-14)。大多數新增內容是一般人口的陳述,對答案沒有改變,但一小部分則將屬性附加到特定患者或改變所選選項(0.25-2.4%的回應,99.8%朝向不正確的選項),其中虛構的人口統計改變了模型推薦的答案。措辭將效果擴大至14%至56%。DEI提示僅是更一般機制的一個例子。任何促使模型推理的指令都可能使其添加未請求的細節,包括有關患者的細節。被標記的輸出被視為正在研究的模型錯誤,而非臨床指導。

Low-Rank Dynamics-Effective Latent Carriers for Counterfactual Rollout in Learned World Models

2608.15156v1 by Yang Liu, Yuming Chen

World models may predict the future without making clear which parts of their hidden state actually drive those predictions. We ask whether a small, directly addressable hidden-state change can place a learned world model on the intended counterfactual trajectory and then let the model continue that future on its own. We study a recurrent world model with a 192-dimensional hidden state in a controlled two-object, two-dimensional collision environment. For a bounded family of local velocity edits, we first verify that the model can natively represent and roll out the edited future. We then construct candidate low-rank carriers from training-only factual-to-counterfactual hidden differences and learn a map from the factual state and requested edit to carrier coefficients. On the registered rank grid, rank 4 is the smallest tested rank that satisfies the full development-panel criteria. A single rank-4 patch at the anchor is sufficient to redirect a 12-step autonomous rollout, with no future observations, teacher forcing, or repeated correction. The frozen procedure satisfies the preregistered replication rule across independently trained checkpoints and remains usable across nearby intervention times. Random equal-norm, wrong-object, and wrong-time controls do not explain the effect. A position-edit stress test provides a negative contrast: the intended position patch can pass the raw rollout criteria, but no-patch and random controls can pass the same criteria, and wrong-object specificity is not established. Thus, successful editing alone is not enough. We use dynamics-effective to describe an intervention that changes the model's future computation in a sustained and target-specific way under autonomous rollout. The rank-4 result identifies a compact intervention interface for the tested velocity-edit family, not a closed four-dimensional state or an intrinsic state dimension.

摘要:世界模型可能預測未來,但並未明確指出其隱藏狀態的哪部分實際驅動這些預測。我們詢問是否可以透過一個小的、可直接訪問的隱藏狀態變化,將學習到的世界模型置於預期的反事實軌跡上,然後讓模型自行繼續那個未來。我們研究了一個具有192維隱藏狀態的遞歸世界模型,在一個受控的兩物體、二維碰撞環境中。對於一個有界的局部速度編輯家族,我們首先驗證模型是否能夠原生地表示並展開編輯後的未來。然後,我們從僅訓練的事實到反事實的隱藏差異中構建候選低秩載體,並學習從事實狀態和請求編輯到載體係數的映射。在註冊的秩網格上,秩4是滿足完整發展面板標準的最小測試秩。在錨點處,單個秩4的補丁足以重定向一個12步的自主展開,無需未來觀察、教師強迫或重複修正。凍結程序滿足預註冊的複製規則,並在獨立訓練的檢查點之間保持可用,並在附近的干預時間內保持可用。隨機等範數、錯誤物體和錯誤時間的控制無法解釋該效果。一個位置編輯壓力測試提供了負對比:預期的位置補丁可以通過原始展開標準,但無補丁和隨機控制也可以通過相同標準,且未建立錯誤物體的特異性。因此,僅僅成功編輯是不夠的。我們使用動態有效來描述一種干預,該干預在自主展開下以持續且目標特定的方式改變模型的未來計算。秩4的結果確定了對測試的速度編輯家族的緊湊干預介面,而不是封閉的四維狀態或內在狀態維度。

Fast Test-Time Refinement for Robust Learned Image Compression

2608.15113v1 by Jiaming Liang, Chi-Man Pun, Weisi Lin

Learned image compression (LIC) has demonstrated remarkable rate-distortion (RD) performance in benign settings. However, the high representational capacity endowed by deep neural networks (DNNs) comes at the expense of increased adversarial vulnerability. This hinders their adoption as trusted standardized codecs. Recent work has sketched test-time refinement (TTR) as a defense in gray-box scenarios, despite its original purpose of improving benign RD performance. Unfortunately, extensive iterations of TTR incur prohibitive overhead, while the robustness mechanism lacks theoretical understanding. Moreover, TTR has not been evaluated in white-box settings or against attacks beyond $\ell_2$-bounded rate and untargeted distortion objectives. To bridge these gaps, we present a systematic study. Our study reveals an Asymmetric Adversarial Trajectory (AAT) property in LIC systems: transitioning from adversarial to benign regions is significantly easier than the reverse process, where adversarial examples can often be roughly recovered within only 1-2 steps. We provide a two-dimensional Tube Model to explain this phenomenon. Based on AAT, we propose a Fast Test-Time Refinement (FTTR) framework for practical and robust LIC systems. We establish that the robustness arises from the contraction of adversarial regions induced by the Input-as-Label property of LIC systems, rather than from obfuscated gradients. Extensive evaluations with diverse strong adaptive attacks across multiple LIC systems demonstrate the promise of the proposed FTTR framework. The code is available at https://github.com/chinaliangjiaming/FTTR.git.

摘要:學習型影像壓縮(LIC)在良性環境中展現了卓越的率失真(RD)性能。然後,深度神經網絡(DNN)所賦予的高表徵能力卻以增加對抗脆弱性為代價。這阻礙了它們作為可信標準編碼器的採用。最近的研究勾勒出測試時精煉(TTR)作為灰箱場景中的防禦,儘管其最初目的是改善良性的RD性能。不幸的是,TTR的廣泛迭代會產生高昂的開銷,而其穩健性機制缺乏理論理解。此外,TTR尚未在白箱環境中進行評估,也未針對超出$\ell_2$-界限率和非針對性失真目標的攻擊進行評估。為了填補這些空白,我們提出了一項系統研究。我們的研究揭示了LIC系統中的不對稱對抗軌跡(AAT)特性:從對抗區域轉移到良性區域顯著容易於反向過程,其中對抗範例通常可以在僅1-2步內粗略恢復。我們提供了一個二維管道模型來解釋這一現象。基於AAT,我們提出了一個快速測試時精煉(FTTR)框架,旨在實現實用且穩健的LIC系統。我們確立了穩健性源於LIC系統的輸入作為標籤屬性所引起的對抗區域收縮,而非來自模糊梯度。對多個LIC系統進行的廣泛評估,涵蓋多種強適應性攻擊,展示了所提出的FTTR框架的潛力。代碼可在https://github.com/chinaliangjiaming/FTTR.git獲得。

Command-Space Counterfactual Explanations for Pareto-Conditioned Reinforcement Learning

2608.14963v1 by Joanikij Chulev, Hendrik Baier

Pareto Conditioned Networks learn multiple multi-objective reinforcement learning behaviours by conditioning a single policy on a desired return command. However, the local mapping from command and state to action remains opaque. We propose command-space counterfactual explanations for PCNs: given a fixed state, original command, and foil action, we search, in a black-box setting, for a minimally changed desired-return command under which the same trained policy would choose the foil. Our contributions are threefold. First, we formulate PCN explanations as return-command interventions, using a return-only PCN variant that avoids the added ambiguity of horizon-conditioning. Second, we adapt adversarial machine learning methods to reinforcement-learning explanations. Third, we introduce a boundary-seeded directional search that improves over purely local optimization in the command-action landscape, resulting in our proposed approach CF-ZOO. The resulting explanations are actionable and intuitively expressed in the user's own preferences: "If your trade-off had shifted slightly towards X, the agent would have chosen Y."

摘要:Pareto Conditioned Networks 通過將單一策略條件化於期望回報命令來學習多種多目標強化學習行為。然而,從命令和狀態到行動的局部映射仍然不明朗。我們為 PCNs 提出了命令空間的反事實解釋:在固定的狀態、原始命令和對照行動下,我們在黑箱環境中搜索一個最小變更的期望回報命令,在此命令下,相同的訓練策略會選擇對照行動。我們的貢獻有三個方面。首先,我們將 PCN 解釋公式化為回報命令的干預,使用一種僅回報的 PCN 變體,避免了地平線條件化所帶來的額外模糊性。其次,我們將對抗性機器學習方法適應於強化學習解釋。第三,我們引入了一種邊界引導的方向性搜索,這在命令-行動空間中優於純粹的局部優化,從而形成我們提出的方法 CF-ZOO。所得的解釋是可操作的,並以用戶自己的偏好直觀表達:“如果你的權衡稍微向 X 方向移動,代理將會選擇 Y。”

Handover Analysis for Vehicular Communication with Explainability on the Fly

2608.14820v1 by Ali Fuat Sahin, Semiha Tedik Başaran, Tufan Kumbasar

Handover (HO) management in vehicular networks requires fast and reliable decision-making under highly dynamic conditions. While machine learning (ML) approaches can improve HO detection by capturing complex relationships among various key performance indicators (KPIs), their black-box nature limits interpretability and operator trust. To address this, this paper investigates HO detection from an explainability-on-the-fly perspective using inherently interpretable models based on the functional analysis of variance (fANOVA) framework. The proposed models are evaluated using two real-world operator datasets and compared against a Long Short-Term Memory baseline augmented with post-hoc SHAP explanations. Unlike post-hoc approaches, the proposed framework enables immediate interpretation of model decisions without incurring additional computational overhead. This capability is particularly critical for latency-sensitive vehicular networks. The results show that fANOVA-based models achieve competitive detection performance while providing significantly reduced explanation latency compared to conventional post-hoc methods. Furthermore, feature ranking and visualization analyses reveal physically meaningful relationships between KPIs and HO occurrences that align with standardized HO mechanisms. These results demonstrate that inherently interpretable models provide an efficient and transparent solution for HO detection in next-generation vehicular networks.

摘要:Handover (HO) 管理在車輛網絡中需要在高度動態的條件下進行快速且可靠的決策。雖然機器學習 (ML) 方法可以通過捕捉各種關鍵性能指標 (KPI) 之間的複雜關係來改善 HO 檢測,但其黑箱特性限制了可解釋性和操作員的信任。為了解決這個問題,本文從即時可解釋性的角度研究了基於方差的功能分析 (fANOVA) 框架的 HO 檢測。所提出的模型使用兩個真實世界的運營商數據集進行評估,並與增強了後驗 SHAP 解釋的長短期記憶基準進行比較。與後驗方法不同,所提出的框架能夠立即解釋模型決策,而不會產生額外的計算開銷。這一能力對於對延遲敏感的車輛網絡尤為重要。結果顯示,基於 fANOVA 的模型在檢測性能上具有競爭力,同時提供顯著降低的解釋延遲,相較於傳統的後驗方法。此外,特徵排名和可視化分析揭示了 KPI 與 HO 發生之間的物理意義關係,這與標準化的 HO 機制一致。這些結果表明,固有可解釋的模型為下一代車輛網絡中的 HO 檢測提供了一種高效且透明的解決方案。

Generated Context versus Governed State: Functional Conditions for Accountable Longitudinal Clinical Reasoning

2608.14804v1 by Augusto Bernardo Pissarra, Victor Lorena de Farias Souza

Large language models (LLMs) have become the dominant interface of clinical artificial intelligence, yet the interface they expose (text in, text out, one context window at a time) maintains no explicit, persistent, governed representation of what is currently true about a patient. This paper argues that longitudinal clinical reasoning is a state-estimation problem under partial observability, and that the axis on which clinical AI succeeds or fails is not the fluency of the model reading the record but the governance of the patient state it reasons over. We distinguish generated context from governed state; separate five objects that clinical AI habitually conflates (true state, observations, evidence, belief, and simulated state); define a tiered governance standard against which any clinical AI system can be audited; and show that an operational definition of accountability decomposes into four information requirements: an immutable evidence ledger with awareness-time versioning, a belief state distinct from accumulated evidence, an observation-process model, and claim-level causal typing. We are explicit that this decomposition is analytic rather than a necessity theorem, and that its value is conceptual hygiene: it converts "accountable clinical AI" from a slogan into an audit instrument. A six-level maturity framework separates what a system makes governable from what it can compute, locating current LLM-centric practice at high capability but low maturity. The paper is fully self-contained: the four research questions the framework poses are stated in the introduction, and the conclusion records what the paper establishes toward each; future work develops the buildable core of the architecture and the research program toward full Clinical World Models. No empirical result is claimed here.

摘要:大型語言模型(LLMs)已成為臨床人工智慧的主導介面,但它們所暴露的介面(文本輸入、文本輸出、一次一個上下文窗口)並未對目前有關患者的真實情況提供明確、持久、受管控的表徵。本文主張,縱向臨床推理是一個在部分可觀察性下的狀態估計問題,而臨床 AI 成功或失敗的軸心不在於模型閱讀記錄的流暢性,而在於它所推理的患者狀態的治理。我們區分生成的上下文與受管控的狀態;將臨床 AI 通常混淆的五個對象(真實狀態、觀察、證據、信念和模擬狀態)分開;定義一個分層治理標準,以便對任何臨床 AI 系統進行審計;並顯示一個運作性責任的定義可分解為四個信息要求:具有意識時間版本控制的不可變證據賬本、與累積證據不同的信念狀態、觀察過程模型,以及索賠級別的因果類型。我們明確指出這一分解是分析性的,而非必要定理,其價值在於概念衛生:它將“可負責任的臨床 AI”從口號轉變為審計工具。一個六級成熟度框架將系統可治理的部分與其可計算的部分分開,將當前以 LLM 為中心的實踐定位於高能力但低成熟度。本文是完全自足的:框架提出的四個研究問題在引言中陳述,結論記錄了本文在每個問題上所建立的內容;未來的工作將發展可構建的架構核心及通向完整臨床世界模型的研究計劃。此處不聲稱任何實證結果。

Decoding the Past: An Uncertainty-Aware Deep Learning Framework for Sex Attribution in Prehistoric Hand Stencils

2608.14539v1 by Karel Becerra, Boris Mederos, Dean Snow, Ramón A. Mollineda

Determining the biological sex of the individuals who created Upper Paleolithic hand stencils remains a challenging problem due to the absence of ground truth, population differences between contemporary and prehistoric groups, and the uncertainty introduced by image degradation. Traditional morphometric methods suffer from high structural overlap across sexes, poor cross-population generalizability, and subjective feature engineering. This study presents an uncertainty-aware deep learning framework for sex attribution in prehistoric hand stencils that explicitly models, propagates, and aggregates uncertainty throughout the analytical pipeline. The methodology combines dual image processing, dual contour extraction, structured silhouette augmentation, model architectural diversity, and ensemble-based decision aggregation. The pipeline generates twelve plausible silhouette realizations per stencil to capture boundary uncertainties, which are processed by two ensembles of ten deep neural networks each (EfficientNet-B3 and MobileViT-S) trained on 14,036 contemporary hand samples. Furthermore, a triangulated validation scheme integrates ensemble predictions with unsupervised 2D latent-space manifold mapping (UMAP + k-NN) and explainable AI spatial attributions (LayerCAM) to ensure anatomical consistency. On contemporary data, ensemble models achieve strong classification performance, with accuracies exceeding 88% in older age groups. When applied to prehistoric stencils, the framework produces both sex predictions and confidence measures of internal agreement, enabling the distinction between morphologically stable and ambiguous cases. Convergence across ensemble predictions, latent-space structure, and interpretability analyses shows that uncertainty can become a measurable component of archaeological inference, enabling robust and reproducible decoding of ancient rock art.

摘要:確定創造上舊石器時代手印的個體的生物性別仍然是一個具有挑戰性的問題,這是由於缺乏真實數據、當代與史前群體之間的差異,以及圖像劣化所帶來的不確定性。傳統的形態計量方法在性別之間存在高度的結構重疊、跨群體的普遍性差以及主觀的特徵工程。這項研究提出了一個不確定性感知的深度學習框架,用於史前手印中的性別歸屬,該框架明確地建模、傳播和聚合整個分析流程中的不確定性。該方法結合了雙重圖像處理、雙重輪廓提取、結構化輪廓增強、模型架構多樣性和基於集成的決策聚合。該流程為每個手印生成十二個合理的輪廓實現,以捕捉邊界不確定性,這些輪廓由兩個各包含十個深度神經網絡的集成處理(EfficientNet-B3 和 MobileViT-S),這些網絡是在 14,036 個當代手樣本上訓練的。此外,一個三角驗證方案將集成預測與無監督的 2D 潛在空間流形映射(UMAP + k-NN)和可解釋的 AI 空間歸因(LayerCAM)結合,以確保解剖學的一致性。在當代數據上,集成模型實現了強大的分類性能,年齡較大的群體的準確率超過 88%。當應用於史前手印時,該框架產生性別預測和內部一致性的信心度量,從而能夠區分形態穩定和模糊的案例。集成預測、潛在空間結構和可解釋性分析之間的收斂顯示,不確定性可以成為考古推理的一個可測量組成部分,使古代岩畫的解碼變得穩健且可重複。

NARRATE: A Multimodal Real-World Australian Driving Dataset for Human-Centred Explanations in Automated Driving

2608.14767v1 by Ashkan Yousefi Zadeh, Zishuo Zhu, Xiaomeng Li, Andry Rakotonirainy, Sebastien Glaser, Ronald Schroeter, Patricia Delhomme, Zahra Mehraban

Automated vehicles must explain their decisions in ways that passengers can understand, monitor, and trust. Existing language-annotated driving datasets are mostly observer-written, post-hoc, simulation-based, or generated from sensor inputs, rather than elicited from the driver performing the action. We introduce NARRATE, a multimodal real-world Australian driving dataset comprising 2,050 annotated events from 35 experienced drivers and driving instructors on public roads. Each event is grounded in synchronised visual, localisation, motion, and LiDAR streams and paired with in-vehicle and/or post-drive free-text explanations. NARRATE provides action labels, scenario-context labels spanning six high-level and 32 fine-grained categories, and span-level Situational Awareness (SA) annotations over driver explanations for Perception, Comprehension and Projection. Four benchmark tasks (SA, scenario-context, driver-action classification, and explanation generation) show that this structure is learnable from driver language, while fine-grained context recognition and explanation generation remain challenging. NARRATE paves a path towards more human-centred and domain-aware explanation models for automated driving.

摘要:自動駕駛車輛必須以乘客能理解、監控和信任的方式解釋其決策。現有的語言標註駕駛數據集大多是觀察者撰寫的、事後的、基於模擬的,或是從傳感器輸入生成的,而不是從執行動作的駕駛員那裡引出來的。我們介紹了NARRATE,這是一個多模態的澳大利亞實際駕駛數據集,包含來自35位經驗豐富的駕駛員和駕駛教練在公共道路上標註的2,050個事件。每個事件都基於同步的視覺、定位、運動和LiDAR數據流,並配有車內和/或駕駛後的自由文本解釋。NARRATE提供了動作標籤、涵蓋六個高層次和32個細分類別的情境上下文標籤,以及針對感知、理解和預測的駕駛員解釋的跨度級別情境意識(SA)註釋。四個基準任務(SA、情境上下文、駕駛員動作分類和解釋生成)顯示這一結構可以從駕駛員語言中學習,而細緻的上下文識別和解釋生成仍然具有挑戰性。NARRATE為自動駕駛的更人性化和領域意識的解釋模型鋪平了道路。

Polaris : Multi Agentic System for Conversational Enterprise Analytics

2608.14246v1 by Varuni H K, Soham Sarkar, Jay Kumar, Goutham Krishnan, Tanvi Johari, Avinash Bharadwaj, Santosh Hegde

In today's fast-paced environment, the ability to swiftly access, understand, and act on data is no longer optional; it is essential. Yet most organizations remain data-rich but insight-poor, constrained by the complexity of querying, interpreting, and explaining enterprise-scale information. We present Polaris, a supervisor-led multi-agent framework for conversational enterprise analytics that bridges this gap. Polaris introduces Dynamic Task Coordination (DTC), a decision-theoretic orchestration layer that models agent-task assignment as adaptive bipartite matching, enabling real-time coordination, recovery, and optimization across specialized agents for querying, visualization, and reasoning. By coupling DTC with reason-first, ReAct-style agents, Polaris transforms natural-language queries into coherent analytical workflows that not only retrieve and visualize data but also explain the underlying "why." Evaluation on structured enterprise datasets demonstrates high semantic fidelity and answer relevancy, underscoring the potential of multi-agent orchestration to deliver trustworthy, end-to-end business intelligence at scale.

摘要:在當今快速變化的環境中,迅速訪問、理解和處理數據的能力不再是可選的;它是必需的。然而,大多數組織仍然是數據豐富但洞察貧乏,受到查詢、解釋和解釋企業級信息的複雜性所限制。我們提出了Polaris,一個由主管主導的多代理框架,用於對話式企業分析,填補了這一空白。Polaris引入了動態任務協調(DTC),這是一個決策理論的編排層,將代理-任務分配建模為自適應二分匹配,實現了專業代理之間的實時協調、恢復和優化,用於查詢、可視化和推理。通過將DTC與以推理為首的ReAct風格代理結合,Polaris將自然語言查詢轉化為連貫的分析工作流程,這些工作流程不僅檢索和可視化數據,還解釋了背後的“為什麼”。對結構化企業數據集的評估顯示出高語義保真度和答案相關性,強調了多代理編排在大規模提供可靠的端到端商業智能方面的潛力。

CForce: Boosting Parallel Decoding for dLLMs via Consistency Forcing

2608.13925v1 by Yuji Ren, Chenkai Xu, Zhuocheng Gong, Jianguo Li, Zhijie Deng

Diffusion large language models (dLLMs) accelerate language generation by predicting multiple masks in a single forward pass. However, existing dLLMs can suffer from unreliable predictions in early denoising stages under aggressive parallelism strategies, leading to errors that can propagate to later stages. To tackle this issue, we present Consistency Forcing (CForce) for dLLMs, a distillation method to force the mask predictions of early stages to align with those of later stages. CForce trains the model on pre-collected self-rollout trajectories, thereby improving training-inference alignment. We introduce Confidence Adaptive KL Divergence as a distillation objective to conjoin the merits of forward and reverse KL. We further provide a theoretical analysis for the consistency objective to explain why CForce can approximately minimize the prediction error of early stages. Critically, the same formulation applies to both mask-to-token decoding and edit-capable decoding; in the edit-capable case, later token-to-token refinements provide additional supervision for earlier masked-state predictions. Experiments on non-edit and edit-capable LLaDA models show improved speed-quality trade-offs, especially under high-parallelism decoding budgets. Code is available at: https://github.com/inclusionAI/dFactory.

摘要:擴散大型語言模型(dLLMs)通過在單次前向傳播中預測多個掩碼來加速語言生成。然而,現有的 dLLMs 在激進的並行策略下,可能在早期去噪階段出現不可靠的預測,導致錯誤可能傳播到後續階段。為了解決這個問題,我們提出了一種名為一致性強制(CForce)的 dLLMs 蒸餾方法,以強制早期階段的掩碼預測與後期階段對齊。CForce 在預先收集的自我展開軌跡上訓練模型,從而改善訓練與推理的對齊。我們引入了信心自適應 KL 散度作為蒸餾目標,以結合前向和反向 KL 的優點。我們還提供了一個一致性目標的理論分析,以解釋為什麼 CForce 可以大致最小化早期階段的預測誤差。關鍵的是,這種相同的公式適用於掩碼到標記的解碼和可編輯解碼;在可編輯的情況下,後期的標記到標記的細化為早期的掩碼狀態預測提供了額外的監督。在非編輯和可編輯的 LLaDA 模型上的實驗顯示了速度和質量的權衡改善,特別是在高並行解碼預算下。代碼可在以下網址獲得:https://github.com/inclusionAI/dFactory。

Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions

2608.13786v1 by Qingfang Liu, Qiao Jin, Joe D. Menke, Thorsten Kahnt, Zhiyong Lu

Large language model (LLM) chatbots are increasingly used to answer clinical questions with citations to relevant clinical studies. Prior research has largely focused on citation fabrication, leaving a gap in evaluating the quality of retrieved studies and the factors driving their selection. In this study, we evaluated three general-purpose LLM chatbots: Claude Sonnet 5, Gemini 3.1 Pro, and ChatGPT GPT-5.5. We prompted the models with clinical questions adapted from 20 review questions in Issues 6 and 7 of the 2026 Cochrane Database of Systematic Reviews, simulating patient, clinician, and evidence-synthesis researcher roles. Each chatbot was queried under each user role with four independent repetitions, yielding 720 responses. Each chatbot was asked to support its answers with primary clinical citations, which we benchmarked against the included and excluded study sets of the Cochrane reviews. On average, a chatbot response retrieved 39.2% $\pm$ 29.8% of Cochrane included studies, while citing 5.0% $\pm$ 9.4% of excluded studies. Recall of Cochrane included studies varied significantly by model and user role. ChatGPT achieved higher recall than Claude or Gemini (63.1% $\pm$ 29.5% vs. 37.0% $\pm$ 23.8% vs. 17.3% $\pm$ 13.1%; $p=2.0\times10^{-5}$). The researcher role yielded higher recall than the clinician or patient roles (42.8% $\pm$ 30.8% vs. 38.6% $\pm$ 28.9% vs. 36.1% $\pm$ 29.3%; $p=2.0\times10^{-5}$). Controlling for publication year, citations per year, and open-access status, sample size was the only independently significant predictor of retrieval (odds ratio 1.80 per 1-unit increase in log sample size, 95% CI 1.37-2.36, $p=2.34\times10^{-5}$). These findings suggest that while LLM chatbots can retrieve some studies identified by expert reviewers, their performance varies by model and user role, and they exhibit a bias toward clinical trials with larger sample sizes.

摘要:大型語言模型(LLM)聊天機器人越來越多地用於回答臨床問題,並引用相關的臨床研究。先前的研究主要集中在引用虛構上,未能評估檢索到的研究質量及其選擇的驅動因素。在本研究中,我們評估了三個通用型LLM聊天機器人:Claude Sonnet 5、Gemini 3.1 Pro和ChatGPT GPT-5.5。我們根據2026年Cochrane系統評價數據庫第6和第7期的20個回顧問題,為模型提供了臨床問題的提示,模擬患者、臨床醫生和證據綜合研究者的角色。每個聊天機器人在每個用戶角色下進行了四次獨立查詢,共產生720個回應。每個聊天機器人被要求用主要臨床引用來支持其答案,我們將其與Cochrane評估的納入和排除研究集進行了基準比較。平均而言,聊天機器人的回應檢索了39.2% $\pm$ 29.8%的Cochrane納入研究,同時引用了5.0% $\pm$ 9.4%的排除研究。Cochrane納入研究的回憶率因模型和用戶角色而異。ChatGPT的回憶率高於Claude或Gemini(63.1% $\pm$ 29.5% vs. 37.0% $\pm$ 23.8% vs. 17.3% $\pm$ 13.1%;$p=2.0\times10^{-5}$)。研究者角色的回憶率高於臨床醫生或患者角色(42.8% $\pm$ 30.8% vs. 38.6% $\pm$ 28.9% vs. 36.1% $\pm$ 29.3%;$p=2.0\times10^{-5}$)。在控制出版年份、每年引用數和開放獲取狀態後,樣本大小是唯一獨立顯著的檢索預測因子(對數樣本大小每增加1單位的比值比1.80,95% CI 1.37-2.36,$p=2.34\times10^{-5}$)。這些發現表明,儘管LLM聊天機器人可以檢索到一些專家評審者識別的研究,但其性能因模型和用戶角色而異,並且對樣本大小較大的臨床試驗存在偏見。

Capacity-Dependent Effects of Data Selection for Reasoning

2608.13721v1 by Cuong Dang, Hoang Anh Just, Ruoxi Jia

In reasoning supervised fine-tuning, candidate responses for the same instruction can differ substantially in how well they match the student's current distribution. Recent likelihood-based response selection methods suggest that responses closer to the student distribution provide more effective supervision, motivating the hypothesis that high-likelihood responses may generally be preferable for fine-tuning. In this paper, we revisit this intuition and show that the value of likelihood-based data selection depends critically on model capacity and training duration. Through controlled experiments on mathematical reasoning, using students ranging from 1.5B to 8B parameters and supervision generated by stronger teacher models, we observe a clear \emph{capacity-dependent} ``{\color{SMALLCOLOR}\textbf{Fast-Fit}} / {\color{LARGECOLOR}\textbf{Slow-Gain}}'' pattern. High-likelihood data provides faster and more stable early improvements, especially for smaller models, but low-likelihood data becomes increasingly beneficial for larger models when training is allowed to continue longer. To explain this phenomenon, we analyze learning dynamics, showing that small models often fail to absorb low-likelihood supervision and instead fall into shallow or repetitive behaviors, while larger models are better able to move toward the teacher distribution under such data. We further provide a capacity-constrained theoretical view of distillation that clarifies how data difficulty, data span, and student capacity jointly govern transfer. Overall, our findings show that effective data selection for reasoning should be aware of model capacity and computing budget rather than based on a single universal preference for high-likelihood supervision.

摘要:在推理的監督微調中,對於相同指令的候選回應在與學生當前分佈的匹配程度上可能存在顯著差異。最近基於似然的回應選擇方法表明,與學生分佈更接近的回應提供了更有效的監督,這促使了高似然回應在微調中通常更可取的假設。在本文中,我們重新檢視這一直覺,並顯示基於似然的數據選擇的價值在很大程度上依賴於模型容量和訓練持續時間。通過對數學推理進行控制實驗,使用從1.5B到8B參數的學生以及由更強的教師模型生成的監督,我們觀察到明顯的\emph{容量依賴} ``{\color{SMALLCOLOR}\textbf{快速擬合}} / {\color{LARGECOLOR}\textbf{慢增益}}''模式。高似然數據提供了更快且更穩定的早期改進,特別是對於較小的模型,但當訓練允許持續更長時間時,低似然數據對於較大模型變得越來越有利。為了解釋這一現象,我們分析了學習動態,顯示小模型往往無法吸收低似然監督,反而陷入淺層或重複的行為,而較大模型在這類數據下更能朝向教師分佈移動。我們進一步提供了一個受容量限制的蒸餾理論觀點,闡明了數據難度、數據範圍和學生容量如何共同影響轉移。總體而言,我們的研究結果顯示,對於推理的有效數據選擇應該考慮模型容量和計算預算,而不是基於對高似然監督的單一普遍偏好。

MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and Coordination

2608.13476v1 by Saisha Shetty, Satvik Tripathi, Austin Lin, Colin Zhao, Theodore Kim, Don Enwerem, Jacinta Arnold, Shahriar Faghani, Tessa S Cook

We present Multi-Agent Reasoning and Coordination (MARC), an open-source framework that replaces monolithic LLM prompting with deterministic multi-agent orchestration for clinical reasoning. MARC coordinates role-specialized agents for extraction, reasoning, answer generation, and evaluation, with explicit context passing and traceable intermediate outputs, enabling stage-wise failure attribution. We additionally introduce a Decomposer module that generates task-specific agent prompts from a plain-language description, eliminating manual prompt engineering. The framework supports both API-based and local CPU-compatible deployments and is entirely configurable via YAML, without code modifications. MARC is designed to be model-agnostic, interpretable, and accessible to clinical domain experts without programming expertise. The full framework is available at https://github.com/Penn-RAIL/MARC-v1.

摘要:我們提出了多智能體推理與協調(MARC),這是一個開源框架,將單一大型語言模型的提示替換為確定性的多智能體協調,用於臨床推理。MARC 協調角色專門的智能體進行提取、推理、答案生成和評估,具有明確的上下文傳遞和可追溯的中間輸出,實現階段性失敗歸因。我們還引入了一個分解器模塊,該模塊從普通語言描述中生成任務特定的智能體提示,消除了手動提示工程。該框架支持基於 API 的和本地 CPU 兼容的部署,並且完全可以通過 YAML 配置,而無需修改代碼。MARC 設計為模型無關、可解釋,並且對於沒有編程專業知識的臨床領域專家可訪問。完整框架可在 https://github.com/Penn-RAIL/MARC-v1 獲得。

A Unifying Perspective on Causal World Models: From Observations to Representations to Structure

2608.13456v1 by Avinash Kori, Fabrizio Russo

World Models (WM) are increasingly seen as a foundation for intelligent agents that can predict, plan, and act beyond their training distribution. In this paper, we study WMs from a causal perspective across multiple levels of abstraction, ranging from perceptual observations to building a conceptual representation of the structure governing the environment dynamics. We argue that useful WMs must go beyond generative capabilities alone: they should also capture entity properties, entity-to-entity interactions, and entity-to-environment interactions that determine and explain the dynamics of a system. We provide a formal definition of Causal WMs (CWMs) grounded in the tasks they are intended to support, connecting world modelling with existing work in causal representation learning, object-centric learning, causal discovery, structural causal models, and model-based decision-making. Finally, we relate CWMs to the literature on identifiability, clarifying when the components of a WM can be recovered from data and up to which equivalence. With this, we ground WMs in representations and structures that support causal reasoning and informed decision-making.

摘要:世界模型(WM)越來越被視為智能代理的基礎,這些代理能夠預測、計劃並在其訓練分佈之外行動。在本文中,我們從因果的角度研究WM,涵蓋多個抽象層次,從感知觀察到構建支配環境動態的結構的概念表示。我們認為,有用的WM必須超越僅僅生成的能力:它們還應該捕捉實體特性、實體之間的互動以及實體與環境之間的互動,這些互動決定並解釋系統的動態。我們提供了一個基於其所支持任務的因果WM(CWM)的正式定義,將世界建模與現有的因果表示學習、以物體為中心的學習、因果發現、結構因果模型和基於模型的決策制定相連接。最後,我們將CWM與可識別性文獻相關聯,澄清WM的組件何時可以從數據中恢復以及在何種等價下。通過這樣,我們將WM建立在支持因果推理和知情決策的表示和結構上。

Explanatory Engagement Under Rare Anomalous Failure: Asymptotic Rarity in Model Behavior (or: The Asymptotic AI)

2608.13063v1 by Sam Mao

Prior work on LLM behavior under anomalous conditions asks whether a model notices anomalies. We ask a narrower question: once a model sits in a workflow with a low, controllable failure rate, does its explanatory engagement - length, specificity, self-reported confidence - change as failure grows asymptotically rarer? We built a local, zero-cost harness on three open-weight models (qwen3:8b, llama3.1:8b, mistral:7b) running a repeated tool-call task where one call fails at probability p, swept across eight rates from 0.2 to 0.0001, under five elicitation conditions from immediate prompting to none. We hypothesized a rise in engagement as failures grew rarer, then a collapse near a detectability threshold. Pooled across conditions this appeared false: length fell in a flat, monotonic pattern. Splitting by condition overturned that. Under immediate_forced, where the model must explain every failure instantly, the predicted rise is confirmed but followed by a plateau, not a collapse: length peaks at 28.4 words at p=0.05, settles to 17.4-19.0 words at the rarest rates, and confidence rises unevenly from about 53% to the 70s-90s. Under grouped_runs, explanation batched to run-end, no collapse appears. Under passive_unprompted, aggregate magnitude is a floor artifact, but a recovered logging gap revealed real, model-specific self-monitoring: llama3.1:8b volunteers structured confidence reports unprompted, sometimes eroding its own confidence as trials accumulate; the other two do so only once, as boilerplate. Elicitation structure is a first-class moderator of collapse observability. A companion guaranteed-failure run (72 cells, backfilling rates where random sampling gave zero real failures) shows models differ in whether they recognize an anomaly, distinct from engagement once recognized. Limitation: discrete rate points cannot capture behavior between them, a direction for future work.

摘要:先前對於 LLM 在異常條件下行為的研究探討模型是否能注意到異常。我們提出了一個更狹窄的問題:當模型在一個失敗率低且可控的工作流程中時,隨著失敗變得漸近稀有,其解釋參與度 - 長度、具體性、自我報告的信心 - 是否會改變?我們在三個開放權重模型(qwen3:8b、llama3.1:8b、mistral:7b)上建立了一個本地的零成本工具,運行一個重複的工具調用任務,其中一個調用以概率 p 失敗,並在五種引導條件下從 0.2 到 0.0001 的八個速率中進行掃描,從立即提示到無提示。我們假設隨著失敗變得更稀有,參與度會上升,然後在可檢測性閾值附近會崩潰。根據條件的匯總,這似乎是錯誤的:長度以平坦的單調模式下降。按條件劃分則推翻了這一點。在立即強制條件下,模型必須立即解釋每一次失敗,預測的上升得到了確認,但隨後出現平臺,而不是崩潰:在 p=0.05 時長度達到 28.4 字,在最稀有的速率下穩定在 17.4-19.0 字,信心從約 53% 不均勻上升到 70% 到 90% 之間。在分組運行條件下,解釋批量到運行結束,沒有出現崩潰。在被動無提示條件下,總體大小是一個底部工件,但恢復的日誌間隙揭示了真實的模型特定自我監控:llama3.1:8b 在未提示的情況下自願提供結構化的信心報告,有時隨著試驗的累積而侵蝕自己的信心;另外兩個模型僅在一次時作為模板進行。引導結構是崩潰可觀察性的第一級調節因子。一個伴隨的保證失敗運行(72 個單元,填補隨機抽樣導致零真實失敗的速率)顯示模型在是否識別異常方面存在差異,這與一旦識別後的參與度是不同的。限制:離散速率點無法捕捉之間的行為,這是未來工作的方向。

VALG: An Agentic System for ML Theory Research

2608.13060v1 by Dechen Zhang, Xuan Tang, Xinxiang Yin, Xingwu Chen, Jian Qian, Difan Zou

Machine learning theory studies learning procedures through mathematical setups in which the data model, training protocol, oracle access, loss, metric, and randomness define the phenomenon that a theorem is meant to explain. Solving an open problem therefore requires the problem formulation, theorem target, and proof mechanism to be developed in concert. Researchers formulate hypotheses, test them through preliminary theoretical or empirical analysis, and refine both assumptions and proofs. We investigate whether this process can be organized as an autonomous agentic workflow for ML theory research. We develop VALG, an agentic system that combines multi-level Verification, Adaptive formulation of Learning-theory problems, and Graph-structured proof development. Within each source-relative theorem branch, VALG maintains a fixed mathematical specification, checks the theorem-level composition of a typed proof-dependency graph, and constructs and reviews local proofs in dependency order. When a proof attempt fails, VALG identifies whether the obstruction lies in a derivation, the proof structure, or the theorem formulation and routes the next attempt accordingly. Formulation-level obstructions initiate an explicitly related variant or relaxation, preserving the mathematical relation between the resulting theorem and the source problem. We evaluate VALG on nine subproblems from five COLT 2026 open problems. Two runs produce internally finalized theorem candidates that match the scope of their source briefs; the remaining seven yield restricted-method results, special cases, or conditional theorems. These case studies show how VALG keeps source-scope matches, relaxations, conditional results, and blocked attempts mathematically distinct. VALG is open source at https://github.com/DechenZhang/VALG-ML-Theory-Agent.

摘要:機器學習理論通過數學設置研究學習程序,其中數據模型、訓練協議、預言機訪問、損失、度量和隨機性定義了定理旨在解釋的現象。因此,解決一個開放問題需要問題的表述、定理的目標和證明機制共同發展。研究人員提出假設,通過初步的理論或實證分析進行測試,並不斷完善假設和證明。我們調查這一過程是否可以組織成一個自主的代理工作流程,用於機器學習理論研究。
我們開發了 VALG,一個結合多層驗證、學習理論問題的自適應表述和圖結構證明開發的代理系統。在每個源相對的定理分支內,VALG 維持固定的數學規範,檢查類型證明依賴圖的定理級組合,並按照依賴順序構建和審查局部證明。當證明嘗試失敗時,VALG 確定障礙是否在推導、證明結構或定理表述中,並相應地路由下一次嘗試。表述級的障礙啟動一個明確相關的變體或放鬆,保持結果定理與源問題之間的數學關係。
我們在五個 COLT 2026 開放問題的九個子問題上評估 VALG。兩次運行產生了內部最終化的定理候選,與其源簡報的範圍相匹配;其餘七個產生了限制方法的結果、特例或條件定理。這些案例研究顯示了 VALG 如何保持源範圍匹配、放鬆、條件結果和被阻止的嘗試在數學上是不同的。VALG 是開源的,網址為 https://github.com/DechenZhang/VALG-ML-Theory-Agent。

UniTraffic-Agent: Unified Traffic Video Reasoning for AI City Challenge 2026 Track 3 with Two Out-of-Domain Evaluations

2608.13031v1 by Peng Li, Qianqian Xu, Shilong Bao, Yangbangyan Jiang, Qingming Huang

Traffic video understanding has become an important problem in intelligent transportation, as road videos provide direct evidence for accidents, violations, and interactions between vehicles and vulnerable road users. A useful system should explain how a traffic event develops, why it happens, and when the relevant interaction occurs, yet this remains difficult for multimodal large language models (MLLMs) because traffic videos contain sparse events and varied viewpoints. We introduce UniTraffic-Agent, the MR-CAS solution for Track~3 of the 10th AI City Challenge, which includes Traffic Anomaly Reasoning (TAR) and two out-of-domain evaluations: FETV for fisheye traffic events and PSI-VQA for pedestrian intention reasoning. UniTraffic-Agent follows an observe--reason--act--verify workflow that samples timestamped visual evidence, reasons over all questions from the same clip in one request, and converts responses through task-specific action adapters. On the official Public leaderboards, MR-CAS ranks 16th on TAR with a score of 0.5780, 2nd on FETV with 0.4884, and 4th on PSI-VQA with 64.4161. The code is available at https://github.com/Roclp/UniTraffic-Agent.

摘要:交通視頻理解已成為智能交通中的一個重要問題,因為道路視頻提供了事故、違規和車輛與脆弱道路使用者之間互動的直接證據。一個有用的系統應該解釋交通事件是如何發展的,為什麼會發生,以及相關互動何時發生,然而這對於多模態大型語言模型(MLLMs)來說仍然困難,因為交通視頻包含稀疏事件和多樣的視角。我們介紹了UniTraffic-Agent,這是第十屆AI城市挑戰賽Track~3的MR-CAS解決方案,其中包括交通異常推理(TAR)和兩個域外評估:FETV針對魚眼交通事件和PSI-VQA針對行人意圖推理。UniTraffic-Agent遵循觀察--推理--行動--驗證的工作流程,從時間戳視覺證據中取樣,對同一片段的所有問題進行推理,並通過特定任務的行動適配器轉換響應。在官方公共排行榜上,MR-CAS在TAR中排名第16,得分為0.5780,在FETV中排名第2,得分為0.4884,在PSI-VQA中排名第4,得分為64.4161。代碼可在https://github.com/Roclp/UniTraffic-Agent獲得。

Static analysis-guided agentic AI translation enables Rust as a full stack bioinformatics language

2608.13029v1 by Johan Henriksson

The field of bioinformatics struggles with legacy code - old code that is commonly used but may no longer have a maintainer, or may be written in an now-unfamiliar language (e.g. Perl, Fortran). This incurs maintenance cost (technical debt), but dynamically typed languages also negatively impacts the environment and fail to make use of modern hardware. Legacy code may also have security or safety problems that make it unsuited for use in clinical settings. Here we show that agentic AI, combined with static analysis, can be used to translate legacy code to the modern language Rust. We provide prompts and supporting software to aid systematic translation, and evaluate it on common software for NGS and imaging. We showcase the result on our software Bascet: Size was reduced by ~80x, build time decreased by ~10x, and performance of key steps improved >3x. Unix dependencies were also removed, making Bascet the only single-cell pipeline able to run on native Windows, without a container. Large-scale refactoring of bioinformatics software is thus now possible at a limited budget, enabling more complex tools to be developed.

摘要:生物資訊學領域面臨著舊有代碼的挑戰——這些舊代碼通常被使用,但可能不再有維護者,或可能是用現在不熟悉的語言(例如 Perl、Fortran)編寫的。這會產生維護成本(技術負債),但動態類型語言也會對環境產生負面影響,並未能充分利用現代硬體。舊代碼可能還存在安全或安全性問題,使其不適合在臨床環境中使用。在這裡,我們展示了代理式 AI 結合靜態分析,可以用來將舊代碼轉換為現代語言 Rust。我們提供提示和支持軟體以協助系統性翻譯,並在 NGS 和成像的常見軟體上進行評估。我們展示了我們的軟體 Bascet 的結果:大小減少約 80 倍,建構時間減少約 10 倍,關鍵步驟的性能提高了超過 3 倍。Unix 依賴也被移除,使 Bascet 成為唯一能在本地 Windows 上運行的單細胞管道,而無需容器。因此,生物資訊學軟體的大規模重構現在在有限的預算下成為可能,從而使得更複雜的工具得以開發。

Decomposition of Evidence, Contradiction, and Fragility in Perturbation Responses

2608.12935v1 by Lei You

Perturbation methods explain model decisions by measuring prediction changes under altered inputs, but response magnitude tells us only how much a model reacts, not what that reaction means. The same magnitude can support the final factual-counterfactual difference, oppose it, or arise strongly along the perturbation path yet vanish at the endpoint. We therefore track how the contrast develops as paired inputs are progressively revealed, using the final contrast to interpret the trajectory. We introduce DECAF (Decomposition of Evidence, Contradiction, And Fragility), which routes aligned, opposed, and endpoint-null responses into evidence E, contradiction C, and fragility F. The decomposition preserves ordinary magnitude exactly, Abs = E + C + F, and is unique under endpoint-relative axioms. Across controlled vision and tabular settings, the three components track independently measured behavior. In a 72-model ImageNet-9 audit, we compare cases with nearly identical response magnitude but different independently measured behaviors. The largest DECAF component agrees with an observed behavior in 96.4% of cases, compared with 35.0% for magnitude alone. Changing only the reveal path increases total response by nearly 80%, yet evidence barely changes while fragility grows by more than 4x. On FunnyBirds and ImageNet-1k, short forward-only DECAF trajectories outperform the tested general-purpose attribution baselines. On a 1B-scale DINOv2 model, a short trajectory matches a strong gradient-based baseline with 4.75x lower wall time and 2.36x lower peak memory.

摘要:擾動方法通過測量在改變輸入下的預測變化來解釋模型決策,但反應的大小僅告訴我們模型反應的程度,而不告訴我們這種反應的意義。相同的大小可以支持最終的事實-反事實差異,反對它,或在擾動路徑上強烈出現但在終點消失。因此,我們追蹤對比如何隨著配對輸入的逐步揭示而發展,並利用最終對比來解釋軌跡。我們引入DECAF(證據、矛盾和脆弱性的分解),將對齊的、對立的和終點無效的反應路由到證據E、矛盾C和脆弱性F。這種分解準確地保留了普通的大小,Abs = E + C + F,並且在相對於終點的公理下是唯一的。在控制視覺和表格設置中,這三個組件追蹤獨立測量的行為。在一次72模型的ImageNet-9審計中,我們比較了反應大小幾乎相同但獨立測量行為不同的案例。最大的DECAF組件在96.4%的案例中與觀察到的行為一致,而僅僅依賴大小的情況下為35.0%。僅改變揭示路徑使總反應增加近80%,而證據幾乎不變,脆弱性增長超過4倍。在FunnyBirds和ImageNet-1k上,僅向前的短DECAF軌跡超越了測試的通用歸因基準。在一個規模為1B的DINOv2模型上,短軌跡與一個強大的基於梯度的基準相匹配,牆面時間低4.75倍,峰值內存低2.36倍。

Discovering Efficient and Explainable Communication Topologies for LLM-based Multi-Agent Systems via Causal Inference

2608.12921v2 by Junzhi Li, Peng He, Qirui Ji, Wei Wang, Lixiang Liu, Chuxiong Sun

The performance of large language model (LLM)-based multi-agent systems (MAS) largely depends on effective communication topologies. Existing topology generation methods, however, typically learn communication topologies through black-box optimization driven solely by task-level rewards. While effective, such optimization provides little insight into why particular communication edges are selected, making it difficult to identify the critical communication subgraphs responsible for successful collaboration. To address this limitation, we propose E2-Explainer, a model-agnostic framework for providing interpretable explanations of communication topologies produced by arbitrary topology generators. Specifically, we formulate topology explanation as a causal attribution problem that identifies compact communication subgraphs supported by edge-level evidence of task preservation. We obtain this evidence with a Granger-style objective that measures how masking each communication channel changes the task outcome and the stability of the final response. The resulting budgeted subgraphs are then distilled into an amortized explainer, enabling efficient post-hoc explanation without repeated edge-level evaluations at deployment. Extensive experiments on multiple reasoning and coding benchmarks demonstrate that E2-Explainer identifies critical communication subgraphs that preserve successful collaboration. These subgraphs can also be executed directly to prune redundant communication edges, substantially reducing communication costs while maintaining competitive task performance.

摘要:大型語言模型(LLM)為基礎的多代理系統(MAS)的性能在很大程度上依賴於有效的通信拓撲。然而,現有的拓撲生成方法通常通過僅由任務級獎勵驅動的黑箱優化來學習通信拓撲。雖然有效,但這種優化對於為何選擇特定通信邊緣提供了很少的洞察,這使得識別負責成功協作的關鍵通信子圖變得困難。為了解決這一限制,我們提出了E2-Explainer,一個模型無關的框架,用於提供可解釋的通信拓撲解釋,這些拓撲由任意拓撲生成器產生。具體而言,我們將拓撲解釋公式化為一個因果歸因問題,該問題識別由任務保持的邊級證據支持的緊湊通信子圖。我們通過一個Granger風格的目標來獲得這些證據,該目標測量屏蔽每個通信通道如何改變任務結果和最終響應的穩定性。隨後,得到的預算子圖被提煉成一個攤銷解釋器,從而在部署時能夠高效地進行事後解釋,而無需重複的邊級評估。在多個推理和編碼基準上的廣泛實驗表明,E2-Explainer識別出保持成功協作的關鍵通信子圖。這些子圖還可以直接執行,以修剪冗餘的通信邊緣,顯著降低通信成本,同時保持競爭性的任務性能。

Mr3D-VL: A generalist vision language foundation model for Multiparametric 3D Magnetic Resonance Imaging

2608.12689v1 by Zhi Qiao, Xintong Wu, Yichu He, Feng Shi

Multi-parametric magnetic resonance imaging (mpMRI) is a cornerstone for brain tumor diagnosis and treatment, yet current AI models face critical limitations: their lack of natural language interaction and interpretability impedes spatial information integration and cross-modal reasoning required clinically. Key challenges arise from significant physical meaning differences across modalities, spatial misalignment due to scan intervals, and the need for complex multi-feature interpretation in tasks like glioma grading. While visual-language models (VLMs) show promise in cross-modal understanding, existing methods focus mainly on 2D image modeling, neglecting direct perception of 3D volumetric space. Although 3D VLMs have been proposed for report generation and feature alignment in 3D CT imaging, mpMRI applications demand collaborative inference across multiple imaging modalities-a requirement unmet by current solutions. To address this, we introduce Mr3D-VL, a dedicated visual-language foundation model for multi-parametric 3D MRI. With 4 billion parameters, it employs an unsupervised pre-trained shared 3D encoder and 4D rotational positional embedding for dual modality-spatial integration. Its cross-modal projection layer uses a multi-resolution feature implantation strategy to enhance feature perception across resolutions. Experimental results show significant improvements over existing 4B/7B/30B domain-specific and general-purpose models in text generation tasks, achieving a BERTScore of 0.856 for report generation, with question-answering accuracy at 0.713 and multiple-choice accuracy at 0.912.

摘要:多參數磁共振成像(mpMRI)是腦腫瘤診斷和治療的基石,但目前的AI模型面臨重大限制:缺乏自然語言互動和可解釋性妨礙了臨床所需的空間信息整合和跨模態推理。主要挑戰來自於不同模態之間顯著的物理意義差異、由於掃描間隔造成的空間錯位,以及在如膠質瘤分級等任務中對複雜多特徵解釋的需求。儘管視覺語言模型(VLMs)在跨模態理解方面顯示出潛力,但現有方法主要集中在2D圖像建模,忽略了對3D體積空間的直接感知。雖然已提出3D VLMs用於報告生成和3D CT成像中的特徵對齊,但mpMRI應用需要跨多個成像模態的協作推理——這一需求目前的解決方案無法滿足。為了解決這個問題,我們推出了Mr3D-VL,一個專門針對多參數3D MRI的視覺語言基礎模型。它擁有40億個參數,採用無監督預訓練的共享3D編碼器和4D旋轉位置嵌入進行雙模態空間整合。其跨模態投影層使用多解析度特徵植入策略來增強不同解析度間的特徵感知。實驗結果顯示,在文本生成任務中,與現有的4B/7B/30B領域特定和通用模型相比,顯著提高了性能,報告生成的BERTScore達到0.856,問答準確率為0.713,多選準確率為0.912。

Interpretable Causal Discovery via Causal-Effect Constraints

2608.12640v1 by Cixuan Zhang, Guy Van den Broeck, Benjie Wang

Causal discovery aims to uncover the underlying causal relationships given data generated from a system. The goal, however, is not merely to predict causal edges given data, but also to be able to interpret and explain either observed or hypothesized phenomena, such as a particularly large causal effect. We consider this task of conditional causal discovery and cast it as a Bayesian inference problem, in which we target the posterior over causal graphs and parameters conditional on an event such as a causal-effect constraint. Unfortunately, this poses a computational challenge: existing approaches to Bayesian causal discovery struggle when the event has small posterior mass. To address this, we adapt rare-event estimation techniques to perform inference the joint graph-parameter space. Our method gradually drives a particle population toward the constrained region while maintaining samples that approximate the conditional posterior. Empirical evaluation on synthetic graphs validates the accuracy of our approach at small and large scales, and we show in a case study on the Sachs protein dataset how our method can be used to aid scientific exploration by providing pathway-level summaries.

摘要:因果發現旨在揭示基於系統生成數據的潛在因果關係。然後,目標不僅僅是根據數據預測因果邊緣,還要能夠解釋和說明觀察到的或假設的現象,例如特別大的因果效應。我們考慮這一條件因果發現的任務,並將其視為一個貝葉斯推斷問題,在這個問題中,我們針對因果圖和參數的後驗分佈,條件是某個事件,如因果效應約束。不幸的是,這帶來了計算挑戰:現有的貝葉斯因果發現方法在事件具有小後驗質量時表現不佳。為了解決這個問題,我們調整了稀有事件估計技術,以在聯合圖-參數空間中進行推斷。我們的方法逐漸將粒子群體推向受約束的區域,同時保持近似條件後驗的樣本。在合成圖上的實證評估驗證了我們的方法在小規模和大規模上的準確性,我們在Sachs蛋白數據集的案例研究中展示了我們的方法如何通過提供通路級摘要來幫助科學探索。

Algorithm Design and Physician Liability

2608.13618v1 by Shujie Luan, Shubhranshu Singh, Tinglong Dai

A single clinical algorithm can deliver unequal accuracy across patient groups, and concern about such disparity has grown as artificial intelligence (AI) spreads through clinical decision-making. In response, a liability rule introduced in the United States holds healthcare providers responsible when their reliance on disparate algorithms contributes to erroneous clinical decisions. We examine how such liability considerations reshape (i) an AI firm's algorithm design decisions that drive group-specific accuracy and (ii) a physician's decisions to use AI in healthcare delivery. The AI firm designs an algorithm for two patient groups, and improving accuracy for the disadvantaged group is more costly. The physician (who remains the accountable decision-maker) then decides whether to consult AI, weighing the reduction in clinical uncertainty against expected liability exposure when AI errors disproportionately affect the disadvantaged group. We find the liability rule can induce disparate use of AI: the physician may reduce AI use overall and, over an intermediate range of liability, rely on AI less for disadvantaged patients. The effect is non-monotone. As liability increases, the physician's use of AI for disadvantaged patients first declines, then rises as the firm reallocates investment toward reducing disparity or switches to an equal-accuracy design. Mandating equal algorithmic accuracy across patient groups can then inadvertently harm both groups, because a uniform accuracy requirement distorts the firm's investment incentives and the physician's equilibrium AI-use decisions.

摘要:單一的臨床演算法在不同患者群體中可能會產生不均等的準確性,隨著人工智慧(AI)在臨床決策中的普及,對於這種差異的關注也日益增加。作為回應,美國引入了一項責任規則,當醫療提供者依賴不同的演算法導致錯誤的臨床決策時,將其負責。我們研究這種責任考量如何重塑(i)AI公司的演算法設計決策,促進特定群體的準確性,以及(ii)醫生在醫療提供中使用AI的決策。AI公司為兩個患者群體設計了一個演算法,改善弱勢群體的準確性成本更高。然後,醫生(仍然是負責的決策者)決定是否諮詢AI,權衡臨床不確定性的減少與當AI錯誤不成比例地影響弱勢群體時的預期責任風險。我們發現責任規則可能會導致AI的使用不均等:醫生可能會整體減少AI的使用,並且在責任的中等範圍內,對弱勢患者的AI依賴程度降低。這一效果是非單調的。隨著責任的增加,醫生對弱勢患者使用AI的情況最初下降,然後隨著公司將投資重新分配到減少差異或轉向平等準確性設計而上升。要求在患者群體之間達到平等的演算法準確性,可能會無意中對兩個群體造成傷害,因為統一的準確性要求扭曲了公司的投資激勵和醫生的均衡AI使用決策。

What Makes a Peer? Valuation-Anchored Similarity in Private Markets

2608.12594v1 by Sebastian Frank, Jingrao Lyu, Max Jarmey, Preetha Saha, Mingshu Li, Sweet Kaur, Sola Akinola, Dhagash Mehta

As more investors contemplate private markets and contend with limited transparency, sparse disclosures, and infrequent transactions, identifying economically meaningful peer companies for comparison is a fundamental challenge for valuation, due diligence, portfolio construction, and risk management. We propose an ensemble tree-based supervised similarity learning framework that defines company similarity through the lens of market valuation rather than static feature matching or semantic descriptions. Specifically, we train a CatBoost gradient-boosted decision tree model on observed private company valuations and derive a valuation-aware similarity metric from importance-weighted leaf-node co-occurrences across the ensemble. The similarity metric captures shared valuation drivers while accommodating nonlinear relationships, mixed data types, and pervasive missing data common in private markets. Using a global private-market universe of approximately 270,000 companies, including more than 53,000 firms with observed or derivable post-money valuations spanning multiple industries, geographies, and deal stages, we demonstrate that the proposed similarity framework improves upon traditional distance-based and text-embedding-based approaches in downstream k-nearest-neighbor valuation tasks in the evaluated industry groups, while retaining case-based explainability.

摘要:隨著越來越多的投資者考慮私募市場並面對有限的透明度、稀疏的披露和不頻繁的交易,識別具有經濟意義的同行公司以進行比較對於估值、盡職調查、投資組合構建和風險管理來說是一個基本挑戰。我們提出了一種基於集成樹的監督相似性學習框架,通過市場估值的視角來定義公司相似性,而不是靜態特徵匹配或語義描述。具體而言,我們在觀察到的私募公司估值上訓練了一個CatBoost梯度提升決策樹模型,並從集成中的重要性加權葉節點共現中推導出一個考慮估值的相似性度量。該相似性度量捕捉了共享的估值驅動因素,同時適應了非線性關係、混合數據類型以及在私募市場中普遍存在的缺失數據。使用約27萬家公司的全球私募市場範圍,包括超過53,000家具有觀察或可推導的後資金估值的公司,涵蓋多個行業、地理區域和交易階段,我們證明所提出的相似性框架在評估的行業組中改善了傳統的基於距離和基於文本嵌入的方法在下游k最近鄰估值任務中的表現,同時保留了基於案例的可解釋性。

Auditable agentic AI for evidence-grounded thyroid ultrasound diagnosis and reporting

2608.12590v1 by Haifan Gong, Shiyu Chen, Bodong Wang, Yuqi Wang, Shijie Wang, Guoliang You, Xinyu Xiong, Haowei Wang, Mingzhi Mao, Dexing Kong, Qinghua Liu, Wei Lou, Fei Chen, Guanbin Li

Thyroid ultrasound diagnosis requires coordinated lesion localization, measurement, risk stratification and reporting, yet most AI systems address these tasks in isolation and provide limited support for clinical review. We present ThyroidXAgent, a clinician-interactive agentic AI system that coordinates specialized diagnostic tools and stores their outputs as an auditable case-level evidence record. The system was developed using OpenThyroidDB, a multicentre, multitask resource integrating approximately 0.3 million ultrasound images and 24,000 paired reports, and was evaluated on 28,458 non-overlapping test cases, including 8,721 cases from 35 centres in the private NHC-MISD-TUS cohort. Across heterogeneous datasets, ThyroidXAgent achieved a mean Dice score of 87.21 percent for nodule segmentation and a mean AUROC of 0.9466 for benign-malignant classification. The same workflow supported lymph-node metastasis prediction and follicular versus papillary thyroid carcinoma classification, with AUROCs of 0.864 and 0.805, respectively. For report generation, evidence-grounded assembly outperformed multimodal language-model baselines across three cohorts. ThyClinScore, a lesion-level clinical semantic metric introduced here, showed the strongest correlation with a location-aware language-model judge. ThyroidXAgent improved physician classification accuracy, increased report diagnostic consistency from 70.3 percent to 86.2 percent, and reduced segmentation and reporting time by 35.9 percent and 27.4 percent, respectively. These findings support auditable, clinician-correctable agentic AI for thyroid ultrasound diagnosis and reporting.

摘要:甲狀腺超聲診斷需要協調病變定位、測量、風險分層和報告,但大多數人工智慧系統在孤立的情況下處理這些任務,並對臨床審查提供有限的支持。我們提出了ThyroidXAgent,一個臨床互動的代理人工智慧系統,協調專門的診斷工具並將其輸出存儲為可審計的案例級證據記錄。該系統是使用OpenThyroidDB開發的,這是一個多中心、多任務的資源,整合了約30萬張超聲圖像和24,000份配對報告,並在28,458個不重疊的測試案例上進行了評估,包括來自私立NHC-MISD-TUS隊列的35個中心的8,721個案例。在異質數據集上,ThyroidXAgent在結節分割方面達到了87.21%的平均Dice分數,並在良惡性分類方面達到了0.9466的平均AUROC。同一工作流程支持淋巴結轉移預測和濾泡型與乳頭狀甲狀腺癌的分類,AUROC分別為0.864和0.805。在報告生成方面,基於證據的組合在三個隊列中超越了多模態語言模型基準。這裡引入的ThyClinScore,一個病變級的臨床語義指標,顯示出與位置感知語言模型評審者之間的最強相關性。ThyroidXAgent提高了醫生的分類準確性,將報告的診斷一致性從70.3%提高到86.2%,並分別減少了35.9%和27.4%的分割和報告時間。這些發現支持可審計、可由臨床醫生修正的代理人工智慧用於甲狀腺超聲診斷和報告。

CAS: A Causal Attribution Score for Local and Global Explainable Artificial Intelligence

2608.12555v1 by Michael Georgiades, Charalambia Varnava

Predictive explanation methods attribute a model output; they do not, by themselves, attribute an intervention effect on the real-world outcome. We introduce the Causal Attribution Score (CAS), a compact score architecture for causal explanation. CAS starts from an identified interventional coalition game, allocates the joint intervention contrast with causal Shapley contributions, and converts those raw outcome-scale effects into Local CAS, Signed Local CAS, and two complementary Global CAS summaries. The innovation is not a new Shapley formula, but a local-to-global causal reporting layer with an explicit intervention target. In the known-truth benchmark, eight repeated primary-interaction simulations (n = 2,200 each, three actions) gave mean Local CAS MAE of 0.107 for coalition-aware CAS, compared with 0.173 for one-at-a-time normalisation and 0.213 for a global normalised absolute ATE vector. The paired advantage over one-at-a-time normalisation increased from -0.003 under additivity to 0.091 under strong interactions. On both empirical DoubleML datasets, 401(k) eligibility/net financial assets (n = 9,915) and Pennsylvania reemployment bonus/unemployment duration (n = 5,099), predictive SHAP/TreeSHAP rankings differed materially from Feature-CAS rankings of treatment-effect modifiers. In Pennsylvania, dep1 (exactly one dependent) moved from predictive global rank 13 to Feature-CAS rank 2 and was the leading local Feature-CAS modifier. These results isolate the added value of separating what predicts the outcome from what explains heterogeneity in an estimated causal effect.

摘要:預測解釋方法歸因於模型輸出;它們本身並不歸因於對現實世界結果的干預效果。我們介紹了因果歸因分數(Causal Attribution Score, CAS),這是一種用於因果解釋的緊湊分數架構。CAS 以識別的干預聯盟遊戲為起點,根據因果 Shapley 貢獻分配聯合干預對比,並將這些原始結果尺度效果轉換為局部 CAS、簽名局部 CAS 和兩個互補的全球 CAS 摘要。這一創新不是一個新的 Shapley 公式,而是一個具有明確干預目標的局部到全球因果報告層。在已知真相的基準測試中,八次重複的主要互動模擬(每次 n = 2,200,三個行動)對於考慮聯盟的 CAS 給出了平均局部 CAS MAE 為 0.107,而一次性標準化為 0.173,全球標準化的絕對 ATE 向量為 0.213。相較於一次性標準化,基於可加性的配對優勢從 -0.003 增加到強互動下的 0.091。在兩個實證 DoubleML 數據集上,401(k) 合格性/淨財務資產(n = 9,915)和賓夕法尼亞州再就業獎金/失業持續時間(n = 5,099),預測 SHAP/TreeSHAP 排名與治療效果修飾因子的 Feature-CAS 排名有顯著差異。在賓夕法尼亞州,dep1(恰好一名受撫養人)從預測全球排名第 13 移動到 Feature-CAS 排名第 2,並成為主要的局部 Feature-CAS 修飾因子。這些結果隔離了將預測結果與解釋估計因果效果異質性的內容分開的附加價值。

Class Activation Mapping in Explainable Computer Vision: A Method-Centered Review of CNN, Transformer, and Foundation-Model-Era Visual Explanations

2608.12299v2 by AmirHossein Eshghi, Hamid Saadatfar, Seyyed Ali Hoseini, AmirMohsen Eshghi, Siavash Arjomand Bigdeli

Class activation mapping (CAM) is one of the most widely used visual explanation families in explainable artificial intelligence. Its purpose is intuitive: it converts internal model evidence into a heatmap that highlights the image regions, convolutional channels, tokens, or patches that support a target class or concept. Since the first CAM formulation in 2016, the field has moved far beyond global-average-pooled CNN classifiers. CAM-style methods now include gradient-based post-hoc explanations, gradient-free score and ablation methods, high-resolution upscaling, weakly supervised localization and segmentation, transformer token attribution, causal and debiasing methods, and foundation-model-era approaches that use CLIP, DINO, SAM, or feature-distribution comparisons. This review synthesizes a strict corpus of 57 method-centered papers published from 2016 onward. The paper develops a taxonomy that separates methods by attribution mechanism, architectural dependence, and evaluation objective. It then reviews gradient-based CAMs, recent and hybrid CAM-style methods, and model-based or architecture-aware methods. Across the corpus, the main trend is clear: the field is shifting from explaining one class score in one low-resolution CNN layer toward comparative, multi-layer, probabilistic, token-aware, and foundation-model-aware explanations. At the same time, evaluation remains fragmented. Faithfulness, localization, robustness, computational cost, and human trust are often measured with different protocols. The review therefore emphasizes not only what each method contributes, but also which gap it leaves open and which later methods attempt to close that gap.

摘要:類別激活映射(CAM)是可解釋人工智慧中最廣泛使用的視覺解釋方法之一。其目的直觀明瞭:它將內部模型證據轉換為熱圖,突顯支持目標類別或概念的圖像區域、卷積通道、標記或補丁。自2016年首次提出CAM公式以來,該領域已經遠遠超越了全局平均池化的CNN分類器。CAM風格的方法現在包括基於梯度的事後解釋、無梯度的分數和消融方法、高解析度的放大、弱監督的定位和分割、Transformer標記歸因、因果和去偏見方法,以及使用CLIP、DINO、SAM或特徵分佈比較的基礎模型時代方法。本綜述綜合了自2016年以來發表的57篇以方法為中心的論文,形成了一個嚴格的語料庫。該論文發展了一個分類法,根據歸因機制、架構依賴性和評估目標來區分方法。然後,它回顧了基於梯度的CAM、最近的混合CAM風格方法,以及基於模型或架構感知的方法。在這個語料庫中,主要趨勢顯而易見:該領域正從解釋一個低解析度CNN層中的一個類別分數,轉向比較的、多層的、概率的、標記感知的和基礎模型感知的解釋。與此同時,評估仍然是碎片化的。忠實性、定位、穩健性、計算成本和人類信任通常使用不同的協議進行測量。因此,該綜述不僅強調每種方法的貢獻,還指出它留下的空白,以及後來的方法試圖填補該空白。

Dual Spatial-Temporal Attribution: Architecture-Aligned Post-Hoc Explainability for Recurrent Graph Anomaly Detection

2608.12441v1 by Iyad Assaad Nekka, Hamida Seba, Khaled Walid Hidouci, Karima Amrouche

Deep learning detectors for anomalies in dynamic graphs have reached strong accuracy, yet they remain opaque: when an edge is flagged, the analyst receives a score but no reason. This opacity is untenable in the cooperative, regulated information systems where such detectors are deployed, where automated decisions must be auditable and trustworthy. We address this gap for AddGraph, the foundational GCN+GRU framework for edge-level anomaly detection in dynamic graphs, which to our knowledge has never been equipped with any form of explainability. We present a strictly post-hoc explainability framework, X-AddGraph, built on a Dual Spatial-Temporal Attribution (DSTA) mechanism whose three components are each aligned with one of AddGraph's architectural modules: a gradient-based relevance attribution over the current adjacency structure (spatial), a direct reading of the contextual attention weights already computed during inference (short-term temporal, at zero additional cost), and a gradient rollback through the recurrent hidden states (long-term temporal). Because the detector is frozen, detection performance is preserved exactly (Delta AUC = 0, verified empirically to ten decimal places). On the UCI Message benchmark, our trained AddGraph baseline reaches an average per-snapshot AUC of 0.8705, exceeding the originally published result; X-AddGraph reproduces every score identically while adding explanations where none existed. Evaluated across four edge populations - confident true positives, low-confidence true positives, false positives, and random samples - the long-term attribution identifies historical snapshots carrying significantly more counterfactual signal than random selection (0.127 vs. 0.074), a capability that no spatially-blind explainer can provide. We release our implementation for full reproducibility.

摘要:深度學習檢測器在動態圖中的異常檢測已達到強大的準確性,但它們仍然不透明:當一條邊被標記時,分析師收到一個分數但沒有理由。這種不透明性在這些檢測器被部署的合作性、受規範的信息系統中是無法接受的,因為自動決策必須是可審計和可信的。我們針對AddGraph這一基礎的GCN+GRU框架進行了這一改進,該框架用於動態圖中的邊級異常檢測,據我們所知,它從未配備過任何形式的可解釋性。我們提出了一個嚴格的事後可解釋性框架X-AddGraph,該框架基於雙空間-時間歸因(DSTA)機制,其三個組件與AddGraph的架構模塊相對應:對當前鄰接結構的基於梯度的相關性歸因(空間),對推理過程中已計算的上下文注意權重的直接讀取(短期時間,無額外成本),以及通過循環隱藏狀態的梯度回滾(長期時間)。由於檢測器是固定的,因此檢測性能完全保留(Delta AUC = 0,經實證驗證至十位小數)。在UCI Message基準測試中,我們訓練的AddGraph基準達到每個快照平均AUC 0.8705,超過了最初發表的結果;X-AddGraph在添加解釋的同時,準確重現了每個分數。通過四個邊群體進行評估 - 自信的真陽性、低信心的真陽性、假陽性和隨機樣本 - 長期歸因識別出比隨機選擇顯著更多的反事實信號的歷史快照(0.127對0.074),這是任何空間盲目解釋器無法提供的能力。我們釋出我們的實現以實現完全可重現性。

Learning-Based Behavior Planning for Automated Driving: Real-World Integration and Deployment

2608.12198v1 by Jean-Pierre Busch, Guido Linden, Jan Bergmann, Lutz Eckstein

Recent research in machine and deep learning has shown the potential of learningbased motion planning approaches to improve the driving behavior of automated vehicles, especially in complex environments. However, their complex nature and lack of transparency can hinder explainability and trustworthiness and complicate safety assurance. Motivated by these challenges, we propose a hybrid planning architecture that combines the advantages of machine learning with the verifiability and the determinism of classical approaches. Specifically, we developed a deep neural network to interpret complex traffic scenes and propose driving behavior, while an optimization-based supervision layer validates this proposal and enforces explicit drivability and safety constraints. We evaluate the learned planner's driving behavior in open-loop studies on real-world urban data, discuss system integration aspects for stable closed-loop operation, and report results from real-world deployment on our research vehicle karl..

摘要:最近在機器學習和深度學習領域的研究顯示,基於學習的運動規劃方法有潛力改善自動駕駛車輛的駕駛行為,特別是在複雜環境中。然而,它們的複雜性和缺乏透明度可能會妨礙可解釋性和可信度,並使安全保證變得複雜。受到這些挑戰的啟發,我們提出了一種混合規劃架構,結合了機器學習的優勢以及經典方法的可驗證性和確定性。具體而言,我們開發了一個深度神經網絡來解釋複雜的交通場景並提出駕駛行為,同時基於優化的監督層驗證這一提議並強制執行明確的可駕駛性和安全約束。我們在現實世界的城市數據上進行開環研究,評估學習到的規劃者的駕駛行為,討論穩定閉環運行的系統集成方面,並報告我們的研究車輛karl的實際部署結果。

A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench

2608.12138v1 by Praveen Reddy, Charuta Mandke, Suvrankar Datta, Sarah Khan, Siddharth Reddy Anthireddy, Shitij Arora, Vishal Singh

General-purpose large language models (LLMs) have recently been reported to match or exceed specialized clinical AI tools on medical benchmarks, but such comparisons draw on a narrow set of systems and on benchmarks developed largely in high-income settings. We evaluate VITA, a retrieval-augmented generation (RAG) system purpose-built for contextual knowledge retrieval in India and other low- and middle-income (LMIC) settings. VITA retrieves from a curated corpus of disease-specific guidelines, India-specific antimicrobial resistance data, national formulary constraints, and resource-limited care protocols; its architecture and corpus are proprietary, but the benchmark, the physician-written rubrics, and our full response and scoring outputs are public for independent verification. On 4,023 English-language HealthBench questions (80.5% of the benchmark), scored with a GPT-4.1 judge, VITA ranked first with 51.9% of possible rubric points, ahead of GPT-5.4 (46.1%), o4-mini (44.3%), Gemini 3.1 Pro (42.6%), and Claude Sonnet 4.6 (37.3%), and scored highest on 45.4% of questions. To test robustness to newer models and judge lineage, a 500-question subset was re-run against current-generation models (GPT-5.5, Claude Opus 4.8, Gemini 3.5 Pro, Grok 4.3) and graded by a neutral open-weight judge (DeepSeek-V4-Pro) sharing no lineage with any system tested. Here the gap narrowed to parity: VITA and GPT-5.5 were statistically indistinguishable on mean per-question score, while VITA led on points-weighted score and won the most questions. VITA's advantages in accuracy and completeness persisted under the neutral judge; its communication scores were lower. These results indicate that a purpose-built clinical RAG system remains competitive with frontier LLMs on an open benchmark, consistent with corpus specificity as a design variable that improves grounding at some cost to communication polish.

摘要:一般用途的大型語言模型(LLMs)最近被報導在醫療基準上與專門的臨床人工智慧工具相匹配或超越,但這些比較依賴於一組狹窄的系統以及主要在高收入環境中開發的基準。我們評估了VITA,一個專為印度及其他低收入和中等收入(LMIC)環境中的上下文知識檢索而設計的檢索增強生成(RAG)系統。VITA從一個策劃的特定疾病指導方針、印度特定的抗微生物抗藥性數據、國家藥典限制以及資源有限的護理協議中檢索資料;其架構和語料庫是專有的,但基準、醫生撰寫的評分標準以及我們的完整回應和評分輸出是公開的,以便獨立驗證。在4,023個英語HealthBench問題(基準的80.5%)上,使用GPT-4.1評判,VITA以51.9%的可能評分點排名第一,超過了GPT-5.4(46.1%)、o4-mini(44.3%)、Gemini 3.1 Pro(42.6%)和Claude Sonnet 4.6(37.3%),並在45.4%的問題上得分最高。為了測試對新模型的穩健性和評判系譜,對500個問題的子集再次運行,與當前一代模型(GPT-5.5、Claude Opus 4.8、Gemini 3.5 Pro、Grok 4.3)進行比較,並由一位中立的開放權重評判(DeepSeek-V4-Pro)進行評分,該評判與任何測試系統無關。此時差距縮小至平行:VITA和GPT-5.5在每題平均得分上統計上無法區分,而VITA在加權得分上領先並贏得了最多問題。VITA在準確性和完整性上的優勢在中立評判下持續存在;其溝通得分較低。這些結果表明,專為臨床設計的RAG系統在公開基準上仍然與前沿LLMs具有競爭力,這與語料庫的特異性作為設計變量相一致,該變量在某種程度上提高了基礎性,但對溝通的精緻性造成了成本。

Do LLMs Take Care of Their Own? Similarity Signals Can Induce Cooperation

2608.12125v1 by Akash Kundu, Emanuel Tewolde, Ratip Emin Berker, Samuel F. Brown, Vincent Conitzer

As LLM-based agents with user-instructed goals are becoming widely deployed, they increasingly encounter each other in strategic interactions, and face challenges of finding mutually beneficial outcomes. Prior literature has argued that cooperation problems such as the Prisoner's Dilemma are resolvable in settings where agents know they follow very similar decision making patterns, as for example in monocultural AI ecosystems. Following that line of work, this paper introduces the first framework for evaluating LLM decision making when agents are provided with graded similarity signals. Among our findings, we establish that different LLM models vary drastically in how they navigate similarity signals, with some modern models showing consistent behavior across cooperation problems, payoff structures, and prompt framing. Perhaps surprisingly, our experiments also show that the dataset based on which the similarity signal is computed has small to no impact on induced cooperation, and that LLM models systematically self-identify as highly similar when asked to evaluate another model's chain-of-thought reasoning by themselves. Finally, we develop an LLM-behavioral-game-theoretic model that captures some of their reasoning rationale, and show that it can support cooperative outcomes in equilibrium under sufficiently high similarity scores.

摘要:隨著基於大型語言模型(LLM)的代理人以用戶指導的目標被廣泛部署,它們在戰略互動中越來越多地相遇,並面臨尋找互利結果的挑戰。先前的文獻已經論證,合作問題如囚徒困境在代理人知道它們遵循非常相似的決策模式的情況下是可以解決的,例如在單一文化的人工智慧生態系統中。沿著這一研究方向,本文介紹了第一個評估LLM決策制定的框架,當代理人被提供分級相似性信號時。在我們的發現中,我們確立了不同的LLM模型在如何導航相似性信號方面存在巨大差異,一些現代模型在合作問題、收益結構和提示框架中顯示出一致的行為。或許令人驚訝的是,我們的實驗還顯示,計算相似性信號的數據集對於誘發合作的影響微乎其微,且當被要求自行評估另一模型的思考鏈推理時,LLM模型系統性地自我識別為高度相似。最後,我們開發了一個LLM行為博弈論模型,捕捉它們的一些推理理由,並顯示它可以在足夠高的相似性分數下支持均衡的合作結果。

Faithful, Sufficient and Understandable: Rethinking Graph Counterfactual Explanations via Discrete Diffusion Inversion

2608.12083v1 by David Bechtoldt, Sidney Bender

Graph Neural Networks (GNNs) achieve strong predictive performance on graph-structured data across domains such as chemistry, biology, and network analysis, yet they provide no intrinsic explanation of their predictions. This limits their adoption in high-stakes and safety-critical settings. Counterfactual explanations address this by revealing the minimal structural modifications that would change a model's prediction. On graphs, however, such a modification is hard to produce. The search space is discrete and combinatorial, and a valid answer must respect categorical node and edge types together with domain rules such as chemical valency in the case of molecular graphs. Existing explainers give up one of two things. Either edits are not held on the data manifold, or the search does not span the full edit space. We propose Graph Diffusion Counterfactual Explanation via Inversion (GDCE-I), which gives up neither. A discrete denoising diffusion model with a novel discrete inversion scheme enables distribution-aware edits leveraging the whole domain edit space. We further address the incomplete and inconsistent evaluation of graph counterfactuals by deriving a framework of explanation desiderata and applying it to every method under one shared protocol. Across four benchmarks, GDCE-I outperforms related work by a large margin on the defined framework. For the molecular domain, we further qualitatively show that GDCE-I attains interpretable in-distribution solutions.

摘要:圖神經網絡(GNNs)在化學、生物學和網絡分析等領域的圖結構數據上實現了強大的預測性能,但它們對其預測並未提供內在解釋。這限制了它們在高風險和安全關鍵環境中的應用。反事實解釋通過揭示最小結構修改來改變模型的預測來解決這一問題。然而,在圖上,這樣的修改難以產生。搜索空間是離散且組合性的,有效答案必須遵守類別節點和邊緣類型以及領域規則,例如在分子圖中的化學價。現有的解釋器放棄了兩者之一。要麼編輯不保持在數據流形上,要麼搜索不涵蓋完整的編輯空間。我們提出了通過反演的圖擴散反事實解釋(GDCE-I),它兩者都不放棄。一種具有新穎離散反演方案的離散去噪擴散模型使得能夠利用整個領域編輯空間進行分佈感知的編輯。我們進一步通過推導解釋需求框架並將其應用於每種方法下的一個共享協議,來解決圖反事實的評估不完整和不一致問題。在四個基準測試中,GDCE-I在定義的框架上大幅超越了相關工作。對於分子領域,我們進一步質性展示GDCE-I獲得了可解釋的內部分佈解決方案。

Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence

2608.12036v1 by Mengru Wang, Junfeng Fang, Shuofei Qiao, Zhenqian Xu, Haoming Xu, Haoxiong Wang, Shumin Deng, Linyi Yang, Zhixiang Cui, Xin Xu, Yunzhi Yao, Buqiang Xu, Fei Shen, Haozhe Luo, Yunxiang Wei, Ningyu Zhang, Julian McAuley, Tat Seng Chua, Huajun Chen

AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities and the risks they may pose remain poorly understood. As AI development becomes faster and increasingly automated, mechanistic exploration remains largely manual, widening the gap between what models can do and our ability to understand and control them. To bridge this gap, we introduce Mechanist, an agentic system that uses AI as a scientific instrument for the autonomous discovery of mechanisms underlying AI intelligence. To support autonomous mechanistic discovery, we construct an interpretability-focused knowledge graph of approximately 13,000 papers and integrate it with a multidisciplinary database of 43 million papers spanning 26 fields. We further curate a library of 32 foundational methods for mechanism analysis, causal intervention, and validation. Compared with Claude Code and existing AI-scientist systems, Mechanist generates more valuable mechanism hypotheses and executes experiments more reliably. Mechanist also demonstrates a progression from discovering model behaviors to explaining and controlling AI models. Specifically, Mechanist first uncovers a counterintuitive safety risk in scientific laboratories, showing that unsafe traits can transfer across modalities through apparently safe training data. Mechanist then develops a mechanism theory of belief, revealing how models represent world knowledge, form beliefs, infer the beliefs of others, and how these mechanisms emerge during pretraining. Finally, Mechanist translates these mechanistic insights into practical interventions that improve model performance across diverse scenarios and steer scientific foundation models toward generating DNA sequences with specified properties.

摘要:AI 模型在各個領域取得了顯著的成功,但其能力背後的機制以及可能帶來的風險仍然不甚了解。隨著 AI 發展變得越來越快速且自動化,機制探索仍然主要依賴人工,這擴大了模型能做的事情與我們理解和控制它們的能力之間的差距。為了縮小這一差距,我們介紹了 Mechanist,一個將 AI 作為科學工具,用於自主發現 AI 智力背後機制的代理系統。為了支持自主的機制發現,我們構建了一個專注於可解釋性的知識圖譜,涵蓋約 13,000 篇論文,並將其與一個跨越 26 個領域的 4,300 萬篇論文的多學科數據庫整合。 我們還策劃了一個包含 32 種基礎方法的庫,用於機制分析、因果干預和驗證。與 Claude Code 和現有的 AI 科學家系統相比,Mechanist 生成了更有價值的機制假設,並更可靠地執行實驗。Mechanist 還展示了從發現模型行為到解釋和控制 AI 模型的進展。具體而言,Mechanist 首先揭示了科學實驗室中的一個反直覺安全風險,顯示不安全特徵可以通過表面安全的訓練數據在不同模態之間轉移。然後,Mechanist 發展了一個信念的機制理論,揭示模型如何表現世界知識、形成信念、推斷他人的信念,以及這些機制如何在預訓練期間出現。最後,Mechanist 將這些機制見解轉化為實際干預措施,改善模型在各種場景中的表現,並引導科學基礎模型生成具有特定屬性的 DNA 序列。

From Safety Documentation to Safety Knowledge Support: An Evidence-Grounded LLM Framework for Medical Devices

2608.12025v1 by Tuhinangshu Gangopadhyay, Rasmus Adler, Peter Liggesmeyer, Jan Reich

Medical devices are becoming more software-intensive, connected, and AI-enabled. Their development requires risk-management evidence aligned with ISO 14971 and, for software, IEC 62304. This evidence must be kept consistent across requirements, design decisions, software changes, verification results, complaints, and post-market data. These tasks are costly and depend on scarce safety and domain experts. Large language models (LLMs) may reduce parts of this effort because medical-device safety work is highly document-based. However, current LLM-based safety-engineering studies often address isolated methods, rely on generic prompting or public examples, and provide limited support for source links, traceability, uncertainty handling, lifecycle updates, and recorded expert review. This limits their use in regulated medical-device development. This paper argues that the central research problem is not safety-text generation, but source-linked safety-knowledge support. We propose an evidence-grounded framework that connects device artifacts, controlled knowledge storage and retrieval, method-specific generation of candidate safety items, critique and uncertainty checks, and recorded expert review. The framework prepares, links, checks, and updates candidate safety artifacts for expert decision-making. It does not decide whether a device is safe and does not provide regulatory approval. We also outline an evaluation strategy using non-public or newly built medical-device case studies and expert reference analyses to assess coverage, correctness, relevance, traceability, duplicate rate, unsupported claims, and review effort.

摘要:醫療器材正變得越來越依賴軟體、互聯網連接和人工智慧。它們的開發需要符合ISO 14971的風險管理證據,對於軟體則需要符合IEC 62304。這些證據必須在需求、設計決策、軟體變更、驗證結果、投訴和市場後數據之間保持一致。這些任務成本高昂,並依賴於稀缺的安全和領域專家。大型語言模型(LLMs)可能會減少這部分工作,因為醫療器材的安全工作高度依賴文檔。然而,目前基於LLM的安全工程研究往往針對孤立的方法,依賴於通用提示或公共範例,並對來源鏈接、可追溯性、不確定性處理、生命周期更新和記錄的專家審查提供有限支持。這限制了它們在受監管的醫療器材開發中的應用。本文主張,核心研究問題不是安全文本生成,而是來源鏈接的安全知識支持。我們提出了一個基於證據的框架,連接設備文檔、受控知識存儲和檢索、特定方法生成候選安全項目、批評和不確定性檢查,以及記錄的專家審查。該框架為專家決策準備、鏈接、檢查和更新候選安全文檔。它不決定設備是否安全,也不提供監管批准。我們還概述了一個評估策略,使用非公開或新建的醫療器材案例研究和專家參考分析來評估覆蓋範圍、正確性、相關性、可追溯性、重複率、不支持的聲明和審查工作量。

Agent Skills Can Be Harmful: An Empirical Study of Skill-Induced Failures in LLM Agents

2608.11888v1 by Gen Dong, Yanjie Gao, Liqun Li, Tianyin Xu, Yu Hua, Fan Yang

Agent skills are the de facto mechanism for extending LLM agents with reusable guidance. A skill can shape the agent's task execution, including planning, tool use, problem-solving, and validation. Prior work reported mixed results of agent skills: some skills improve task success rates, while others have no effect, increase token use and execution time, and even reduce success rates. This paper presents a comprehensive analysis of skill-induced agent failures by attributing task failures and cost regressions to specific loaded skills. We introduce a differential analysis framework that attributes a failure or regression to a skill by comparing a target skill-guided run against a no-skill or semantically matched skill reference run that solves the same task, or solves it more cheaply. We instantiate this framework on SkillsBench and SWE-Skills-Bench, yielding 307 skill-induced failures, including 125 functional failures and 182 efficiency regressions. We also build SkillTriage, a taxonomy-guided attribution tool that normalizes paired cases, extracts differential evidence, and produces triage reports. Our major findings include: (1) Skill induced functional failures are rarely caused by obviously irrelevant skills; instead, seemingly relevant skills often make the agent incorrectly implement or omit task-required implementation elements. (2) Skill-induced efficiency regressions are not explained by prompt length alone. (3) The largest sources within Excessive Procedure are excessive verification and heavy implementation pipelines, contributing 67 and 30 cases, respectively. This shows that skills often turn validation checklists and construction recipes into mandatory work. Based on our findings, we propose research topics and tooling improvements for safer and more cost-aware skill reuse.

摘要:代理技能是擴展LLM代理的事實機制,提供可重用的指導。技能可以影響代理的任務執行,包括規劃、工具使用、問題解決和驗證。先前的研究報告了代理技能的混合結果:一些技能提高了任務成功率,而其他技能則沒有影響,增加了令牌使用和執行時間,甚至降低了成功率。本文通過將任務失敗和成本回歸歸因於特定的加載技能,呈現了技能引起的代理失敗的全面分析。我們引入了一個差異分析框架,通過將目標技能指導的運行與無技能或語義匹配的技能參考運行進行比較,將失敗或回歸歸因於某個技能,這兩者解決相同的任務,或以更低的成本解決。 我們在SkillsBench和SWE-Skills-Bench上實現了這一框架,產生了307個技能引起的失敗,包括125個功能失敗和182個效率回歸。我們還構建了SkillTriage,一個基於分類法的歸因工具,標準化配對案例,提取差異證據,並生成分類報告。我們的主要發現包括:(1)技能引起的功能失敗很少是由明顯不相關的技能造成的;相反,似乎相關的技能常常使代理錯誤地實施或省略任務所需的實施元素。(2)技能引起的效率回歸並不僅僅由提示長度解釋。(3)在過度程序中,最大的來源是過度驗證和繁重的實施管道,分別貢獻了67和30個案例。這顯示技能常常將驗證清單和建構食譜轉變為強制性工作。根據我們的發現,我們提出了更安全和更具成本意識的技能重用的研究主題和工具改進建議。

Low-Interaction-Rank Learning: Unifying Multiplicative Dual-Encoder Heads

2608.11661v1 by Zijian Zhao, Sen Li

A multiplicative dual-encoder network computes a real-valued output for a pair of inputs as the inner product of their separate encodings. This architecture has been developed independently in operator learning, bipartite matching, contrastive vision-language models, retrieval, and other areas, yet no unified theory guides the basic design decisions: how many interaction modes to represent, how to normalize the encoders, and when the architecture should be avoided. We provide such a foundation by introducing the class of functions of low interaction rank, a class whose intrinsic complexity is measured by its interaction spectrum. Within this framework, approximation error decomposes into a spectral truncation term and an encoder-realization term; sample complexity is governed by the sum of the two encoder complexities rather than their product; and a usability criterion based on spectral decay determines when the architecture can succeed. The same framework exposes a central identifiability problem: the encoders are defined only up to a linear gauge symmetry that leaves the learned coordinates arbitrary. We show that normalization is gauge fixing and that whitening pins the interaction modes up to permutation and sign, thereby explaining the uninterpretability of contrastive dimensions and providing a constructive remedy. Experiments on synthetic kernels, operator learning, and CLIP models validate the theoretical predictions: spectral decay rates match the predicted scaling, whitening recovers the true modes, and independently trained CLIP models are related by a single rotation which, after removal by whitening, exposes interpretable concept axes. The code of this paper is provided at https://github.com/RS2002/Mul-Net .

摘要:一個乘法雙編碼器網絡為一對輸入計算實值輸出,作為其各自編碼的內積。這種架構在運算學習、二部匹配、對比視覺-語言模型、檢索及其他領域中獨立發展,但沒有統一的理論指導基本設計決策:如何表示互動模式的數量、如何對編碼器進行正規化,以及何時應避免使用該架構。我們通過引入低互動秩的函數類別提供這樣的基礎,這個類別的內在複雜性由其互動譜來衡量。在這個框架內,近似誤差分解為光譜截斷項和編碼器實現項;樣本複雜性由兩個編碼器複雜性的總和來決定,而不是它們的乘積;基於光譜衰減的可用性標準決定了該架構何時能夠成功。同樣的框架揭示了一個中心可識別性問題:編碼器僅在一個線性規範對稱下被定義,這使得學習到的坐標是任意的。我們展示了正規化是規範固定,並且白化將互動模式固定到置換和符號上,從而解釋了對比維度的不可解釋性並提供了一個建設性的補救措施。在合成核、運算學習和 CLIP 模型上的實驗驗證了理論預測:光譜衰減率與預測的縮放相匹配,白化恢復了真實模式,獨立訓練的 CLIP 模型通過單一旋轉相關,這在白化後去除後揭示了可解釋的概念軸。本文的代碼可在 https://github.com/RS2002/Mul-Net 獲得。

Inverse Theory of Mind Modeling for Content Recommendation: From Web Browsing to Dynamic Intelligent Interfaces

2608.11354v1 by Mengyu Chen, Feiyu Lu, Chun-Fu Chen, Lucas Vinh Tran, Jay Katukuri

Modern recommender systems treat observed actions as reliable proxies for user preferences, yet interactions often reflect exploration or comparison rather than stable preference expression. As interfaces evolve from static layouts toward generative UIs and immersive extended reality (XR), the need for deeper, modality-agnostic user understanding grows: these adaptive environments must decide not only what to present but where, when, how prominently, and most importantly why a user acts. We propose an Inverse Theory of Mind (IToM) pipeline that reasons backward from observed interactions to infer the beliefs, preferences, and decision-making traits that explain behavior. The pipeline reconstructs each user's decision context, including what was chosen and what alternatives were available, applies LLM-driven counterfactual reasoning to produce evidence-grounded natural-language belief statements, and synthesizes these beliefs through multi-hypothesis abductive inference into a structured user persona. We evaluate on the OPeRA dataset against ground-truth personality assessments, attitudinal surveys, and interview-based personas across four tasks: next action prediction, shopping attitude alignment, Big Five personality inference, and held-out category prediction. Results show that inferred personas match or exceed ground-truth personas and that multi-hypothesis reasoning is essential for accurate personality prediction. We further demonstrate cross-modal transferability with a persona-driven spatial banking application on VisionOS.

摘要:現代推薦系統將觀察到的行為視為用戶偏好的可靠代理,但互動往往反映探索或比較,而非穩定的偏好表達。隨著介面從靜態佈局演變為生成式用戶介面和沉浸式擴增實境(XR),對於更深入的、與模式無關的用戶理解的需求日益增長:這些自適應環境不僅必須決定展示什麼,還要決定在何處、何時、以多大程度以及最重要的原因為何用戶會採取行動。我們提出了一個逆向心智理論(IToM)流程,該流程從觀察到的互動中推理回溯,以推斷解釋行為的信念、偏好和決策特徵。該流程重建每個用戶的決策背景,包括所選擇的內容和可用的替代選項,應用基於大型語言模型(LLM)的反事實推理來生成基於證據的自然語言信念陳述,並通過多假設的溯因推理將這些信念合成為結構化的用戶角色。我們在OPeRA數據集上進行評估,對比真實的個性評估、態度調查和基於訪談的角色,涵蓋四個任務:下一步行動預測、購物態度對齊、五大人格推斷和保留類別預測。結果顯示,推斷出的角色與真實角色相匹配或超過,並且多假設推理對準確的人格預測至關重要。我們進一步展示了在VisionOS上的一個以角色為驅動的空間銀行應用的跨模式可轉移性。

Governing Agentic AI in FinTech

2608.11344v2 by Henry Han

Financial institutions are delegating consequential decisions to agentic AI systems that decompose goals, coordinate models and tools, and act with little oversight. Yet agentic AI governance in FinTech is under-investigated. We argue the binding governance constraint is not capability but verifiability. We define the Verifiability Gap as the shortfall between the verification delegated authority demands and the explainability and reproducibility retained after a decision. It is indexed to a verifier, evidentiary standard, and audit lag. We develop a multilevel governance theory for agentic AI and test its mechanisms in three studies over nine model versions, from a three-billion-parameter local model to a commercial frontier system. Study 1 shows that provider releases alter historical financial actions, and that the controls replay needs belong to the provider: the frontier model rejects temperature, top_p and top_k outright and exposes no random seed. Under the tightest controls each endpoint allows, a local model reproduced 320 of 320 executions, hosted models 319 of 320 and 959 of 960. Study 2 shows that orchestration is a latent policy layer. Architecture changes final actions, and no execution record repeated in any configuration at any scale. The frontier model reproduces its own actions more often than the local ones, its record no better, and loses a comparable share of its differentiation. Capability buys a higher starting point, not auditability. Study 3 shows two deterministic credit-model versions each reproduce their current action perfectly, yet the current cannot recover a historical one. We conceptualize reproducibility as a governance profile, not a scalar, yielding evidence-contingent delegation: authority is defensible only while retained evidence substantiates its exercise. Beyond finance, the framework extends to other high-stakes domains requiring auditability.

摘要:金融機構正在將重要決策委託給能夠分解目標、協調模型和工具並在幾乎沒有監督的情況下行動的代理 AI 系統。然而,在金融科技領域,代理 AI 的治理尚未得到充分研究。我們認為,約束治理的關鍵不是能力,而是可驗證性。我們將可驗證性差距定義為驗證委託權限要求與決策後保留的可解釋性和可重複性之間的差距。它與驗證者、證據標準和審計延遲有關。我們為代理 AI 發展了一個多層次的治理理論,並在三項研究中測試其機制,涵蓋九個模型版本,從一個三十億參數的本地模型到一個商業前沿系統。研究 1 顯示,提供者的發布改變了歷史金融行為,並且控制重播需求屬於提供者:前沿模型直接拒絕 temperature、top_p 和 top_k,並且不暴露隨機種子。在每個端點允許的最嚴格控制下,本地模型重現了 320 次執行中的 320 次,託管模型重現了 319 次中的 320 次和 960 次中的 959 次。研究 2 顯示,協調是一個潛在的政策層。架構改變了最終行動,且在任何配置下的任何規模中都沒有執行記錄重複。前沿模型比本地模型更頻繁地重現其自身行動,其記錄並無改善,並失去了相當一部分的差異化。能力提供了一個更高的起點,而不是可審計性。研究 3 顯示,兩個確定性的信用模型版本各自完美重現其當前行動,但當前模型無法恢復歷史行動。我們將可重複性概念化為一種治理配置,而不是一個標量,產生證據依賴的委託:權威只有在保留的證據證實其行使時才是可辯護的。超越金融,該框架擴展到其他需要可審計性的高風險領域。

From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop

2608.11171v1 by Rahul Gupta, Abhinav Mohanty, Anaelia Ovalle, Anil Ramakrishna, Anubrata Das, Apurv Verma, Jwala Dhamala, Ninareh Mehrabi, Tharindu Kumarage, Yada Pruksachatkun, Yang Trista Cao, Kai-Wei Chang, Aram Galstyan

The Workshop on Trustworthy Natural Language Processing (TrustNLP), co-located with major ACL conferences since 2021, has grown from 8 proceedings papers to 41 over six editions, documenting a field-wide transition from post-hoc interpretability of static models to mechanistic understanding and proactive control of generative systems. We synthesize insights from all 144 proceedings papers, classifying them along six trust dimensions grounded in established frameworks (TrustLLM, DecodingTrust). We observe co-occurrences with capability emergence. The release of the first high-impact chat models activated all trust dimensions simultaneously, while subsequent model generations shifted focus toward truthfulness and safety alignment. Analysis from the classification study reveals that truthfulness is the fastest-growing dimension (absent in 2021-2022, comprising 37% of papers by 2025-2026), fairness remains the most consistent theme, and explainability exhibits a U-shaped trajectory; declining as post-hoc methods lost relevance but resurging in 2026 through mechanistic interpretability. A cross-venue comparison with ACL, NAACL, EACL, and EMNLP (~2K papers) in the same period shows that TrustNLP's topical distribution closely follows the field average. We identify four structural insights and conclude with actionable directions for the research community.

摘要:信任自然語言處理研討會(TrustNLP)自2021年以來與主要ACL會議共同舉辦,已從8篇會議論文增長至六屆的41篇,記錄了該領域從靜態模型的事後可解釋性到生成系統的機械理解和主動控制的轉變。我們綜合了所有144篇會議論文的見解,並根據建立的框架(TrustLLM, DecodingTrust)將其分類為六個信任維度。我們觀察到能力出現的共現現象。首批高影響力的聊天模型的發布同時激活了所有信任維度,而隨後的模型世代則將重點轉向真實性和安全對齊。分類研究的分析顯示,真實性是增長最快的維度(在2021-2022年缺失,到2025-2026年佔據37%的論文),公平性仍然是最一致的主題,而可解釋性則顯示出U型軌跡;在事後方法失去相關性時下降,但在2026年通過機械可解釋性再次上升。與ACL、NAACL、EACL和EMNLP(約2K篇論文)在同一時期的跨場域比較顯示,TrustNLP的主題分佈與該領域的平均水平密切相符。我們確定了四個結構性見解,並以可行的研究方向作為結論。

SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure

2608.11079v2 by Xiaofan Bai, Hongqiang Lin, Chao Liu, Yantao Zhang, Xuan Jin, Xipeng Cao, Yuhong Li

Self-evolving agents accumulate reusable skills by appending successful procedures and failure fixes. Over time, the same requirement is often restated in several branches, examples, and warnings, while common action sequences are copied rather than reused. The resulting skill becomes expensive to inject and difficult to maintain. Generic prompt compression is ill-suited to this setting because a skill is not a flat passage: its name and description define when it applies, its workflow controls execution, its tool and output contracts constrain validity, and rare exceptions may remain essential even when no sampled task activates them. Evaluation-guided compression can test these behaviors, but it introduces rollouts, cost, and dependence on the compression-time evaluation set. We present SkillZip, an evaluation-free method that compresses a skill by finding its shortest faithful structural explanation. The intuition is explain once, reference many: state a repeated rule once at the scope where it applies, factor a repeated action sequence into a shared procedure, and keep only the differences as explicit exceptions. We formalize this intuition as a typed minimum description-length objective over a skill contract and a residual, subject to a hard coverage constraint for every extracted trigger, workflow edge, tool requirement, obligation, and output field. The formulation provides simple sharing thresholds, preserves unique rare rules by construction, and supports efficient local updates. SkillZip has a one-shot mode with one structured extraction call and deterministic optimization, and a continual Zip-on-Write mode that integrates each self-evolution patch without replaying tasks or reparsing the full history. Through comprehensive experimental evaluations, we demonstrate the effectiveness and superiority of SkillZip in compression performance, generalizability, and cost overhead.

摘要:自我演化的代理透過附加成功的程序和失敗的修正來累積可重複使用的技能。隨著時間的推移,同樣的需求常常在幾個分支、範例和警告中重新陳述,而常見的行動序列則是被複製而不是重用。由此產生的技能變得難以注入且難以維護。通用提示壓縮不適合這種情況,因為技能不是一段平坦的文字:它的名稱和描述定義了何時適用,其工作流程控制執行,其工具和輸出契約限制有效性,而即使在沒有樣本任務激活的情況下,稀有的例外可能仍然是必要的。評估引導的壓縮可以測試這些行為,但它引入了推出、成本和對壓縮時評估集的依賴。我們提出了SkillZip,一種無需評估的方法,通過找到技能的最短忠實結構解釋來壓縮技能。其直覺是一次解釋,多次參考:在適用的範圍內一次陳述重複的規則,將重複的行動序列分解為共享的程序,並僅保留差異作為明確的例外。我們將這一直覺形式化為一個類型化的最小描述長度目標,針對技能契約和殘餘,並對每個提取的觸發器、工作流程邊緣、工具需求、義務和輸出欄位施加嚴格的覆蓋約束。該公式提供簡單的共享閾值,通過構造保留獨特的稀有規則,並支持高效的本地更新。SkillZip具有一次性模式,通過一次結構化提取調用和確定性優化,還有持續的Zip-on-Write模式,能夠在不重播任務或重新解析完整歷史的情況下集成每個自我演化補丁。通過全面的實驗評估,我們展示了SkillZip在壓縮性能、可泛化性和成本開銷方面的有效性和優越性。

Entropy-Centric Explainable AI for Remote Sensing Image Segmentation

2608.11064v1 by Ali Saleh, Abdul Karim Gizzini, Mohamad Ghassany, Ali J. Ghandour

Artificial intelligence (AI) has become a powerful approach to solving complex problems in critical domains. Many concerns arise regarding the decision-making process of its models, mainly due to deep neural networks outperforming their peers at the cost of ambiguity in feature extraction and prediction. Consequently, in critical domains such as remote sensing, where high-resolution imagery must be analyzed using black-box models, the lack of transparency limits trust in these models and, thus, their adoption. In light of this reality, explaining and understanding the complex decision-making process of AI models has become essential. Explainable AI (XAI) aims to bridge this gap by providing insights into how and why certain decisions are made. While significant progress has been achieved in explaining image classification tasks, image segmentation still offers considerable room for improvement. In this context, this paper proposes an entropy-centric XAI method for semantic segmentation. Moreover, a new XAI evaluation methodology is proposed to efficiently measure the relevance of the regions highlighted by the proposed XAI method. Experimental results demonstrate the superiority of the proposed XAI method compared with recently adapted XAI methods for semantic segmentation.

摘要:人工智慧(AI)已成為解決關鍵領域複雜問題的強大方法。許多關於其模型決策過程的擔憂隨之而來,主要是因為深度神經網絡在特徵提取和預測的模糊性方面超越了其同儕。因此,在遙感等關鍵領域,必須使用黑箱模型分析高解析度影像,缺乏透明度限制了對這些模型的信任,從而影響了它們的採用。鑑於這一現實,解釋和理解AI模型複雜的決策過程變得至關重要。可解釋的AI(XAI)旨在通過提供對某些決策如何以及為何做出的見解來彌補這一差距。儘管在解釋影像分類任務方面已取得顯著進展,但影像分割仍然有相當大的改進空間。在這一背景下,本文提出了一種以熵為中心的XAI方法,用於語義分割。此外,還提出了一種新的XAI評估方法,以有效測量所提出的XAI方法所突顯區域的相關性。實驗結果顯示,所提出的XAI方法在語義分割方面優於最近適應的XAI方法。

ComBodied Agents: a New Paradigm of Human-Centric Agentic AI

2608.10915v2 by Qianggang Ding, Xingyao Wang, Rui Feng, Zhibin Wang, Feixiang Yao, Kelong Mao, Hao Sun, Zhiyao Luo, Jiankai Tang, Lei Li, Jiadong Guo, Minheng Ni, Weicong Lin, Chenxi Yang, Hongxiang Gao, Zhenghua Chen, Yang Bai, Min Wu, Jun Cheng, Huazhu Fu, Dacheng Tao, Bang Liu

After an older adult misses a medication dose, a software agent can send another reminder and an embodied agent can bring the medication. Yet neither explains whether the person forgot, is confused, has side effects, or deliberately refused, nor what support is appropriate. This reveals a structural gap in Agentic AI: Digital Agents primarily transform software states, while Embodied Agents transform physical states; neither makes a person's evolving state and agency the primary object of modeling, intervention, and evaluation. We introduce Combodied Agents, a human-centered paradigm that perceives, models, predicts, and supports individual human-state trajectories over time, using software tools, sensors, wearables, robots, and human services as action channels rather than end goals. We unify fragmented capabilities across personal assistants, health agents, AI companions, and adaptive human--AI systems into a closed loop: event-based multimodal perception reconstructs meaningful personal events; longitudinal, correctable memory provides temporal context; Personal World Models estimate future personal states and outcomes under alternative decisions and interventions; and an admissible intervention policy selects proportionate support under consent, uncertainty, safety, reversibility, and user control. Feedback from the person and environment updates the loop. Rather than requiring an exhaustive Human Digital Twin, the framework uses purpose-bounded, uncertainty-aware, user-correctable representations. We organize the design space by human-state targets, relational contexts, and agent roles, and propose scenario-centered evaluation, agency-preservation metrics, benchmark requirements, edge-native personal models, and governance directions. Combodied Agents shift Agentic AI from external task completion toward sustained human benefit.

摘要:在年長者錯過藥物劑量後,軟體代理可以發送另一個提醒,而具身代理可以帶來藥物。然而,這兩者都沒有解釋該人是否忘記、感到困惑、出現副作用或故意拒絕,也沒有說明什麼樣的支持是合適的。這揭示了代理人工智能中的結構性缺口:數位代理主要轉換軟體狀態,而具身代理則轉換物理狀態;兩者都未將個體不斷演變的狀態和能動性作為建模、干預和評估的主要對象。我們引入了具身代理(Combodied Agents),這是一種以人為中心的範式,能夠隨著時間的推移感知、建模、預測和支持個體的人類狀態軌跡,使用軟體工具、感測器、可穿戴設備、機器人和人類服務作為行動渠道,而非最終目標。我們將個人助理、健康代理、人工智慧伴侶和自適應人類-人工智慧系統的零散能力統一成一個閉環:基於事件的多模態感知重建有意義的個人事件;長期的、可修正的記憶提供時間背景;個人世界模型在不同的決策和干預下估計未來的個人狀態和結果;可接受的干預政策在同意、不確定性、安全性、可逆性和用戶控制下選擇相稱的支持。來自個人和環境的反饋更新這個循環。該框架不需要全面的人類數位雙胞胎,而是使用目的有限、具不確定性意識和用戶可修正的表徵。我們根據人類狀態目標、關係背景和代理角色來組織設計空間,並提出以情境為中心的評估、能動性保護指標、基準要求、邊緣原生個人模型和治理方向。具身代理將代理人工智能的重心從外部任務完成轉向持續的人類利益。

Chemically Meaningful Textualization Enables Explainable Validation of Metal-Organic Frameworks by Large Language Models

2608.11283v1 by Guobin Zhao, Xiao-Yan Li

Computation-ready metal-organic framework (MOF) databases are essential for high-throughput screening, yet many reported crystal structures remain chemically unreasonable or disordered, compromising simulation fidelity. Existing validation approaches can identify non-computation-ready structures, but they often rely on heuristic rules, license requirement, or offer limited interpretability. Here, we show that large language models (LLMs) can serve as interpretable validators of MOF structures when crystallographic information is transformed into chemically meaningful text. By benchmarking nine descriptors, we find that successful LLM-based validation depends not on the amount of structural information alone, but on whether local coordination, framework connectivity, and chemical context are organized into a linguistically learnable representation. Fine-tuned LLMs using specialized descriptors (mof2text) achieve performance comparable to graph-based models in identifying unreasonable MOFs. Importantly, these models extend beyond black-box classification by generating diagnostic rationales for likely error sources, including abnormal bonding, connectivity, and charge states, as well as error-category predictions for annotated datasets. This work establishes chemically informed textualization as the key step that transforms LLMs from generic text models into practical and explainable tools for curating MOF databases.

摘要:計算準備好的金屬有機框架 (MOF) 數據庫對於高通量篩選至關重要,但許多報告的晶體結構仍然化學上不合理或無序,從而影響模擬的真實性。現有的驗證方法可以識別非計算準備好的結構,但它們通常依賴於啟發式規則、許可要求,或提供有限的可解釋性。在這裡,我們展示了大型語言模型 (LLMs) 可以作為 MOF 結構的可解釋驗證者,當晶體學信息轉換為化學上有意義的文本時。通過基準測試九個描述符,我們發現成功的 LLM 基於驗證不僅取決於結構信息的數量,還取決於局部配位、框架連通性和化學上下文是否組織成語言上可學習的表示。使用專門描述符 (mof2text) 的微調 LLM 在識別不合理的 MOF 方面達到與基於圖的模型相當的性能。重要的是,這些模型超越了黑箱分類,通過生成可能錯誤來源的診斷理由,包括異常鍵合、連通性和電荷狀態,以及對註釋數據集的錯誤類別預測,來擴展其功能。這項工作確立了化學知識驅動的文本化作為關鍵步驟,將 LLM 從通用文本模型轉變為實用且可解釋的工具,以便策劃 MOF 數據庫。

Uncertainty-Aware and Explainable Ensemble Deep Learning Framework for Multi-Class Skin Lesion Classification

2608.11280v1 by Rofiqul Islam, Lilatul Ferdouse

Skin cancer diagnosis from dermoscopic images remains challenging due to high intra-class variability, inter-class similarity, class imbalance, and the limited interpretability of deep learning models. This paper proposes an uncertainty-aware and explainable deep learning framework for multi-class skin lesion classification. The framework combines a vision transformer model (MaxViT-Tiny) with CNN-based models (ConvNeXt-Tiny and EfficientNetV2-B0) through deep ensemble learning. Monte Carlo (MC) Dropout estimates predictive uncertainty and identifies unreliable predictions, while Grad-CAM++, an explainable AI (XAI) technique, provides visual explanations by highlighting lesion regions that influence model decisions. Evaluated on the HAM10000 dataset, the framework achieves 96% accuracy and 99% ROC-AUC under uncertainty-aware filtering (entropy < 1.0, confidence >= 0.7), with macro-average precision, recall, and F1-score of 94%, 95%, and 95%, respectively, and 96% weighted-average scores across all three metrics. The results demonstrate accurate, interpretable, and uncertainty-aware skin lesion classification for trustworthy computer-aided diagnosis.

摘要:皮膚癌的診斷從皮膚鏡影像中仍然具有挑戰性,這是由於高內類變異性、類間相似性、類別不平衡以及深度學習模型的有限可解釋性。本文提出了一種不確定性感知和可解釋的深度學習框架,用於多類別皮膚病變分類。該框架通過深度集成學習將視覺Transformer模型(MaxViT-Tiny)與基於CNN的模型(ConvNeXt-Tiny和EfficientNetV2-B0)相結合。蒙特卡羅(MC)Dropout估計預測不確定性並識別不可靠的預測,而Grad-CAM++,一種可解釋的人工智慧(XAI)技術,通過突出影響模型決策的病變區域提供視覺解釋。在HAM10000數據集上進行評估,該框架在不確定性感知過濾(熵 < 1.0,置信度 >= 0.7)下達到96%的準確率和99%的ROC-AUC,宏觀平均精確度、召回率和F1-score分別為94%、95%和95%,在所有三個指標上達到96%的加權平均分數。結果顯示出準確、可解釋且具不確定性感知的皮膚病變分類,為可信的計算機輔助診斷提供支持。

Rule of Thumb: Explaining Artificial Intelligence Systems using Partial Information

2608.10766v2 by Kaivalya Rawal, Daria Onitiu, Brent Mittelstadt, Sandra Wachter, Chris Russell

Explainable Artificial Intelligence (XAI) seeks to explain how an Artificial Intelligence (AI) system arrived at a particular decision. We propose ''Rule of Thumb'' (RoT) explanations, a new approach to XAI based upon a novel formulation that identifies the most relevant features for predicting the behaviour of an AI system, for a particular datapoint. We show how RoT is well-suited to enable XAI in: (a) zero-shot classification using large language models (LLMs), (b) auditing of opaque AI systems without model access, and (c) the use of AI in scientific discovery. Additionally, RoT meets specific requirements from leading AI regulations, provides a familiar interface and visualisations for XAI practitioners, is model-agnostic, and is substantially faster than alternatives. Code available at: https://github.com/KaiRawal/Rule-of-Thumb-Explaining-Artificial-Intelligence-Systems-using-Partial-Information

摘要:可解釋的人工智慧(XAI)旨在解釋人工智慧(AI)系統如何做出特定決策。我們提出了「經驗法則」(RoT)解釋,這是一種基於新型公式的XAI新方法,能夠識別對於特定數據點預測AI系統行為最相關的特徵。我們展示了RoT如何適合於以下情境以促進XAI:(a)使用大型語言模型(LLMs)進行零樣本分類,(b)在無法訪問模型的情況下對不透明的AI系統進行審計,以及(c)在科學發現中使用AI。此外,RoT符合主要AI法規的特定要求,為XAI從業者提供熟悉的界面和可視化,並且對模型無關,速度也比其他替代方案快得多。
代碼可在:https://github.com/KaiRawal/Rule-of-Thumb-Explaining-Artificial-Intelligence-Systems-using-Partial-Information

Operationalising Relative Causal Knowledge: Backbone Identifiability from Private Reports on a Shared Outcome

2608.10664v1 by Fabrizio Russo, Mark Somers

The Relativity of Causal Knowledge (RCK) explains how a network of agents with different structural causal models can exchange causal knowledge through a shared interventionally consistent abstraction, or backbone. We ask the prior identification question that this transport mechanism presupposes: when is that backbone determined by the agents' private causal knowledge? In the basic two-agent common-effect case, two private causes influence one shared outcome and each agent identifies only the single-cause causal marginal relevant to its own perspective. We show that, under standard compatibility, non-degeneracy, and local overlap assumptions, those local causal marginals do not identify a unique backbone. Infinitely many joint intervention kernels can induce exactly the same private reports while disagreeing on joint interventions. We then give a conditional recovery result. Additive separability removes the hidden interaction degree of freedom, but observational residual summaries remain insufficient. Identification becomes possible when agents communicate causally identified response functions. An education value-added example illustrates why this is first a communication problem, and only then a policy-composition problem.

摘要:因果知識的相對性(RCK)解釋了不同結構因果模型的代理人網絡如何通過共享的干預一致抽象或骨幹來交換因果知識。我們提出這一傳輸機制所假設的先前識別問題:何時骨幹由代理人的私人因果知識決定?在基本的兩代理人共同效果案例中,兩個私人原因影響一個共享結果,而每個代理人僅識別與其自身觀點相關的單一原因因果邊際。我們表明,在標準兼容性、非退化性和局部重疊假設下,這些局部因果邊際並不識別唯一的骨幹。無限多的聯合干預核可以產生完全相同的私人報告,同時在聯合干預上存在分歧。然後,我們給出一個條件恢復結果。加性可分離性消除了隱藏的交互自由度,但觀察殘差摘要仍然不足。當代理人交流因果識別的反應函數時,識別變得可能。一個教育增值的例子說明了為什麼這首先是一個通信問題,而後才是一個政策組合問題。

Conversational versus Dashboard Explainable AI for UAV Intrusion Detection: An Empirical Study of Operator Trust and Reliance

2608.10434v1 by Cong Chi Nguyen, Trang Mai Xuan, Vu-Duc Ngo, Kim-Ngan Thi Nguyen, Trong-Nghia Nguyen, Thien Van Luong

Machine learning-based Intrusion Detection Systems (IDS) have demonstrated superior performance in securing Unmanned Aerial Vehicle (UAV) networks. However, the 'black-box' nature of these models, combined with the high dimensionality of multimodal cyber-physical data, poses significant interpretability challenges. Static visualization dashboards may struggle to present complex relationships among multimodal cyber-physical features in a form that is easy for operators to inspect and interpret. To address this, we propose a Conversational XAI interface powered by Large Language Models (LLM) to facilitate on-demand investigation. In a controlled experiment with participants, we systematically evaluated the impact of this conversational interface versus a traditional XAI Dashboard on operator understanding, trust, and reliance during post-incident auditing tasks. Our results suggest that the conversational interface was perceived as more useful than the dashboard, potentially because it helped participants access and synthesize relevant information more easily. However, this benefit was accompanied by a lower level of appropriate self-reliance, indicating a potential risk of over-reliance. One possible interpretation is that the natural-language responses made the AI advice easier to accept, which may have reduced participants' tendency to verify the underlying evidence when the IDS was incorrect. These findings point to a potential trade-off in human-AI collaboration for UAV intrusion auditing: interaction mechanisms that improve perceived usability may also increase the risk of inappropriate reliance. We conclude by discussing design implications for future XAI systems that balance seamless interaction with cognitive forcing functions to foster appropriate reliance.

摘要:基於機器學習的入侵檢測系統(IDS)在保護無人機(UAV)網絡方面顯示出優越的性能。然而,這些模型的「黑箱」特性,加上多模態網絡物理數據的高維度,帶來了顯著的可解釋性挑戰。靜態可視化儀表板可能難以以易於操作員檢查和解釋的形式呈現多模態網絡物理特徵之間的複雜關係。為了解決這個問題,我們提出了一個由大型語言模型(LLM)驅動的對話式XAI界面,以促進隨需調查。在一項對參與者的控制實驗中,我們系統地評估了這個對話式界面與傳統XAI儀表板對操作員理解、信任和依賴在事件後審計任務中的影響。我們的結果表明,對話式界面被認為比儀表板更有用,這可能是因為它幫助參與者更輕鬆地訪問和綜合相關信息。然而,這一好處伴隨著較低的適當自我依賴水平,顯示出過度依賴的潛在風險。一種可能的解釋是,自然語言的回答使得AI建議更容易被接受,這可能減少了參與者在IDS不正確時驗證基礎證據的傾向。這些發現指出了無人機入侵審計中人機協作的潛在權衡:改善感知可用性的互動機制可能也會增加不當依賴的風險。我們最後討論了未來XAI系統的設計啟示,旨在平衡無縫互動與認知強迫功能,以促進適當的依賴。

Reasoning Shortcuts and Value Symmetries: What Symmetry Permits, Architecture Realizes, and Optimization Selects

2608.10420v1 by Xin Xu

Reasoning shortcuts are solutions of a neurosymbolic system's rules that produce correct predictions through unintended concepts. A recent framework of Takemura, Inoue, and Nishino analyzes them through an automorphism group of value relabelings and asks, as its central open question, when rules pin concepts down. We first show that the framework's key definition, one shared permutation applied at every position, does not apply as stated to any of the four heterogeneous benchmarks it was evaluated on, and that the most direct embedding, padding domains to a common size, produces confident false pathology: 90.91% of solution pairs reported unexplained on CLE4EVR, where every well-defined member of the hierarchy we introduce reports 0%, and the padded verdict's content rotates with configuration-file ordering. Re-measuring eleven rule families under fifteen pre-specified predictions (thirteen confirmed), unexplained-pair rates span 0% to 99.9999% and track provable structure: six theorems give sufficient conditions for transitivity and its failure, including a Free Slot Lemma certifying Kandinsky's pathology from syntax alone. For circuit-given rules, deciding symmetry-inertness of a coordinate is coNP-complete; nontrivial-automorphism existence is coNP-hard under randomized reductions, lies in $Σ_2^p$, is not $Σ_2^p$-complete unless PH collapses, and on monotone circuits is coNP-complete outright. In the Boolean case transitivity is classified exactly: automorphisms explain everything iff the solution set is an affine coset. Weakly supervised models place all 94 observed shortcuts at the one level the componentwise theory flags and none at the 48 it certifies transitive; twelve typed-ambiguous levels produce none, separating what symmetry permits from what optimization selects, and a dual-head control replicates the geography. All numbers trace to released artifacts.

摘要:推理捷徑是神經符號系統規則的解決方案,通過意外概念產生正確的預測。Takemura、Inoue 和 Nishino 最近提出的一個框架通過值重新標記的自同構群來分析它們,並將何時規則固定概念作為其核心未解問題。我們首先展示該框架的關鍵定義,即在每個位置應用的共享排列,並未如所述適用於其評估的四個異質基準中的任何一個,並且最直接的嵌入,即將域填充到共同大小,產生了自信的錯誤病理:在 CLE4EVR 上報告的解決方案對中有 90.91% 的未解釋情況,而我們引入的每個明確定義的層級報告 0%,且填充判決的內容隨配置文件排序而旋轉。在十五個預先指定的預測下(確認了十三個),重新測量十一個規則家族,未解釋對的比例從 0% 到 99.9999% 不等,並追踪可證明的結構:六個定理給出了傳遞性及其失敗的充分條件,包括一個自由槽引理,僅通過語法證明 Kandinsky 的病理。對於電路給定的規則,決定坐標的對稱惰性是 coNP 完全的;在隨機約簡下,非平凡自同構的存在是 coNP 困難的,位於 $Σ_2^p$ 中,除非 PH 崩潰,否則不是 $Σ_2^p$ 完全的,並且在單調電路上是 coNP 完全的。在布爾情況下,傳遞性被準確分類:自同構解釋一切當且僅當解決方案集是一個仿射陪集。弱監督模型將所有 94 個觀察到的捷徑放置在組件理論標記的唯一層級上,而在其證明為傳遞的 48 個層級上則沒有;十二個類型模糊的層級產生零,將對稱允許的與優化選擇的分開,雙頭控制複製了地理。所有數字都追溯到釋放的工件。

Towards Expert-level Medical AI for Real-time Video Consultations

2608.09861v1 by Mahvish Nagda, Jihyeon Lee, Matthew Thompson, Chunjong Park, Tim Strother, Valentin Liévin, Roma Ruparel, Akshay Goel, Teya Bergamaschi, Suhana Bedi, Meet Shah, Pavel Dubov, Liviu Panait, Toshiyuki Fukuzawa, Sam Schmidgall, Craig Schiff, Joseph Xu, Aliya Rysbek, Yana Lunts, Jan Freyberg, Rebecca Hemengway, Sunny Virmani, David Racz, Carey Radebaugh, Joëlle Barral, Kavi Goel, Dale R. Webster, Katherine Chou, Avinatan Hassidim, Yossi Matias, James Manyika, Gregory Wayne, Tao Tu, Yun Liu, Ethan Goh, Christina Chen, Ryutaro Tanno, Po-Hsuan Cameron Chen, Mike Schaekermann, Anil Palepu

Audio-visual interaction is the standard for patient-physician consultations, enabling natural communication and effective assessment of illness through non-verbal cues. While text-based AI has shown promise, it discards essential perceptual dimensions and limits patients who cannot articulate symptoms in writing. Early efforts to extend medical AI to audio-visual interaction have demonstrated feasibility but not reached clinician-level performance. Here, we provide the first demonstration of expert-level AI in real-time clinical video consultations using AMIE (Articulate Medical Intelligence Explorer) in a video configuration. AMIE (Video) is a Gemini-based multi-agent system integrating low-latency dialogue, clinical reasoning, and real-time audio-visual perception. To guide development, we established a taxonomy and automated evaluations for clinical audio-visual cues in telehealth settings. In a randomized Objective Structured Clinical Examination (OSCE) study with 30 primary care physicians (PCPs), 15 patient actors and 100 clinical scenarios, we compared AMIE (Video), its text-only counterpart AMIE (Text), and PCPs consulting via video. Clinical evaluators rated AMIE (Video) on par or better than PCPs in history-taking, diagnosis, management, and physical observation and examination. Patient actors preferred AMIE's approach to assessing and explaining conditions, while PCPs were preferred for rapport and partnership building. In modality ablation, patient actors preferred AMIE (Video)'s interface over text chat for communicative effectiveness, convenience, and feeling understood. Limitations remain in fine anatomical precision, subtle affective nuances, and high-frequency movements. While further research is needed before real-world translation, these results mark an important milestone toward AI systems capable of augmenting care across the sensory complexity of clinical practice.

摘要:視聽互動是病人與醫生諮詢的標準,能夠通過非語言線索促進自然交流和有效評估疾病。雖然基於文本的人工智慧顯示出潛力,但它忽略了重要的感知維度,並限制了無法用書面表達症狀的病人。早期將醫療人工智慧擴展至視聽互動的努力已顯示出可行性,但未達到臨床醫生的表現水平。在此,我們提供了使用AMIE(Articulate Medical Intelligence Explorer)在視頻配置中進行實時臨床視頻諮詢的專家級人工智慧的首次示範。AMIE(視頻)是一個基於Gemini的多代理系統,整合了低延遲對話、臨床推理和實時視聽感知。為了指導開發,我們建立了一個分類法和自動評估,用於遠程醫療環境中的臨床視聽線索。在一項隨機的客觀結構化臨床考試(OSCE)研究中,涉及30位初級保健醫生(PCPs)、15位病人演員和100個臨床場景,我們比較了AMIE(視頻)、其文本專用對應AMIE(文本)以及通過視頻諮詢的PCPs。臨床評估者在病史採集、診斷、管理以及身體觀察和檢查方面評價AMIE(視頻)與PCPs相當或更好。病人演員更喜歡AMIE在評估和解釋病情方面的方法,而PCPs則在建立關係和夥伴關係方面更受青睞。在模態消融中,病人演員更喜歡AMIE(視頻)的界面而非文本聊天,因為其在交流有效性、便利性和被理解的感受上表現更佳。儘管在精細解剖精度、微妙的情感細微差別和高頻運動方面仍存在局限性,但在實際應用之前仍需進一步研究,這些結果標誌著朝著能夠增強臨床實踐中感官複雜性的護理的人工智慧系統邁出了重要的一步。

CEAA: A Cognitive Embodied Agents Architecture for Interactive Computing Systems

2608.09848v1 by Aimilios Hadjiliasi, Louis Nisiotis

The development of embodied Intelligent Virtual Agents (IVAs) that have cognitive capabilities in real-time interactive virtual environments remains a challenge, even with today's advancements in technology. Existing architectures are often focused on either the implementation of low-level reactive control systems that are constrained by commercial game engines, or high-level representations of reasoning models that can be difficult to implement in virtual worlds. This paper builds on that notion and proposes a modular cognitive architecture for deploying embodied IVAs. This architecture builds on existing, pre-established frameworks such as the Sense-Think-Act paradigm and the Belief-Desire-Intention cognitive model, among others, and aims to provide a reusable implementation-oriented framework as a template for deploying IVA "brains" in interactive 3D computing systems. The proposed architecture contributes by providing a modular, implementation-oriented framework for the deployment of embodied, cognitive-capable IVAs and bridges the gap between high-level agent reasoning models with real-time embodied execution, for scalable, adaptive, and explainable agents in complex interactive virtual environments.

摘要:身體化智能虛擬代理(IVAs)的發展,具備在實時互動虛擬環境中進行認知能力的挑戰,即使在當今科技進步的情況下仍然存在。現有的架構通常專注於低層次反應控制系統的實施,這些系統受限於商業遊戲引擎,或是高層次推理模型的表現,這在虛擬世界中可能難以實施。本文基於這一概念,提出了一種模組化的認知架構,用於部署身體化的IVAs。這一架構基於現有的、預先建立的框架,如感知-思考-行動範式和信念-欲望-意圖認知模型等,旨在提供一個可重用的實施導向框架,作為在互動3D計算系統中部署IVA“大腦”的模板。所提出的架構通過提供一個模組化的、實施導向的框架,為身體化、具備認知能力的IVAs的部署做出貢獻,並彌合高層次代理推理模型與實時身體執行之間的差距,以實現可擴展、適應性強且可解釋的代理,適用於複雜的互動虛擬環境。

KGCaRe: Explainable Complex Conditional Question Answering using Automatic Knowledge Graph Construction and Context Retrieval with LLMs

2608.09779v1 by Ghanshyam Verma, Simanta Sarkar, Devishree Pillai, Hotaka Shiokawa, Yourong Xu, Fiona Veazey, Peter Hubbert, Hui Su, Paul Buitelaar

Answering complex conditional questions using Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) remains a challenge, particularly in domain-specific contexts where general-purpose LLMs and RAG tend to underperform. We hypothesize that augmenting RAG with unstructured and structured knowledge, extracted from both documents and knowledge graphs (KGs), can improve reasoning and answer accuracy for such tasks. To test this, we propose KGCaRe, a hybrid approach that combines neural retrieval with symbolic reasoning over LLM-generated KGs. KGCaRe constructs a KG from documents using a multi-prompt extraction strategy and stores it in a graph database. Simultaneously, the documents are embedded into a vector store to enable neural retrieval. KGCaRe performs innovative iterative graph traversal guided by the LLM to extract relevant triples, prune irrelevant information, and uses additional clue entities to traverse the graph again if the initial traversal does not provide satisfactory context to generate the answer. The relevant triples extracted from the KG in path form, along with semantically retrieved text passages, are then fed into custom KGCaRe prompts to generate answers to the complex conditional questions with explanations. We evaluate KGCaRe on two complex conditional QA datasets. Our results on these datasets show that KGCaRe consistently outperforms existing baselines, including Vanilla LLM, Code Prompt, Text Prompt, Think-on-Graph, Vanilla RAG, and HybridContextQA, across multiple LLMs such as Mistral, Mixtral, GPT-3.5, and GPT-4o. We publicly release the software pipeline that we developed to implement the proposed KGCaRe approach.

摘要:回答複雜的條件問題使用大型語言模型(LLMs)和檢索增強生成(RAG)仍然是一個挑戰,特別是在特定領域的上下文中,通用的LLMs和RAG往往表現不佳。 我們假設,通過從文檔和知識圖(KGs)中提取的非結構化和結構化知識增強RAG,可以改善此類任務的推理和答案準確性。為了測試這一假設,我們提出了KGCaRe,一種結合神經檢索和LLM生成的KG上符號推理的混合方法。 KGCaRe使用多提示提取策略從文檔中構建KG,並將其存儲在圖形數據庫中。 同時,這些文檔被嵌入到向量存儲中,以便進行神經檢索。 KGCaRe執行創新的迭代圖遍歷,由LLM指導,以提取相關三元組,修剪不相關的信息,並在初始遍歷未能提供滿意的上下文以生成答案時,使用額外的線索實體再次遍歷圖形。 從KG中提取的相關三元組以路徑形式呈現,連同語義檢索的文本段落,然後輸入自定義KGCaRe提示,以生成帶有解釋的複雜條件問題的答案。我們在兩個複雜的條件QA數據集上評估KGCaRe。 我們在這些數據集上的結果顯示,KGCaRe在多個LLM(如Mistral、Mixtral、GPT-3.5和GPT-4o)上始終超越現有基準,包括Vanilla LLM、Code Prompt、Text Prompt、Think-on-Graph、Vanilla RAG和HybridContextQA。 我們公開發布了我們開發的軟件管道,以實現所提出的KGCaRe方法。

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models

2608.09666v1 by Shulin Tian, Ziqi Huang, Fan Zhang, Hongyuan Zhu, Yu Qiao, Ziwei Liu

Recent advances in visual generative models have enabled high-quality image and video generation, but evaluating these models often demands sampling hundreds or thousands of images or videos, which is computationally expensive. Existing evaluation methods also rely on rigid pipelines that overlook specific user needs and provide numerical results without clear explanations. Mimicking how humans quickly form impressions of a model's capabilities from only a few samples, we propose the Evaluation Agent framework, which employs human-like strategies for efficient, dynamic, multi-round evaluations, offering detailed, user-tailored analyses. Given a natural-language evaluation request, the agent decomposes it into sub-aspects, generates targeted prompts, samples images or videos from the evaluated model, invokes suitable evaluation tools, and iteratively updates its plan from the observed evidence, covering both predefined benchmark dimensions and open-ended user concerns. The framework is thus efficient, promptable, explainable, and scalable across models and tools. Experiments show that Evaluation Agent reduces evaluation time to 10% of traditional methods while delivering comparable results. We further introduce Open Evaluation Agent (Open-EA) by constructing EA-CoT-10K, a corpus of history-conditioned step-level instruction-tuning records derived from multi-round evaluation rollouts, and training EA-3B from Qwen2.5-3B-Instruct as a local planning backbone that preserves the structured reasoning, tool invocation, and summary protocol of the API-based agent while reducing dependence on proprietary backbones. Experiments validate the API-based agent on established T2I/T2V benchmarks and open-ended queries, and evaluate Open-EA on four in-domain and three out-of-domain T2V generator families, showing partial cross-family transfer of the learned policy.

摘要:最近在視覺生成模型方面的進展使得高品質的圖像和視頻生成成為可能,但評估這些模型通常需要抽樣數百或數千張圖像或視頻,這在計算上是昂貴的。現有的評估方法也依賴於僵化的流程,忽略了特定用戶需求,並提供沒有明確解釋的數值結果。模仿人類如何快速從少量樣本中形成對模型能力的印象,我們提出了評估代理框架,它採用類似人類的策略進行高效、動態的多輪評估,提供詳細的、針對用戶的分析。給定自然語言的評估請求,代理將其分解為子方面,生成針對性的提示,從被評估模型中抽樣圖像或視頻,調用合適的評估工具,並根據觀察到的證據迭代更新其計劃,涵蓋既定的基準維度和開放式的用戶關注點。因此,該框架在模型和工具之間是高效的、可提示的、可解釋的和可擴展的。實驗顯示,評估代理將評估時間減少到傳統方法的10%,同時提供可比擬的結果。我們進一步通過構建EA-CoT-10K,引入開放評估代理(Open-EA),這是一個從多輪評估展開中衍生的歷史條件步驟級指令調整記錄的語料庫,並從Qwen2.5-3B-Instruct訓練EA-3B作為當地規劃的骨幹,保留基於API的代理的結構化推理、工具調用和摘要協議,同時減少對專有骨幹的依賴。實驗在已建立的T2I/T2V基準和開放式查詢上驗證了基於API的代理,並在四個域內和三個域外的T2V生成器家族上評估Open-EA,顯示出學習政策的部分跨家族轉移。

Rethinking Self-Evolving Agents: Do We Still Need Prescribed Optimization Pipelines?

2608.09629v1 by Hui Xue, Fan Yang

Self-evolving agents are usually built around prescribed optimization pipelines: the framework decides how to gather evidence, revise a persistent artifact, select candidates, and stop. We ask whether this task-specific procedure remains necessary when a frontier model acts as the optimizer. We introduce Open-Ended Optimization (OEO), which keeps the objective, permitted interactions, resource budget, data boundary, and evaluation fixed while allowing the optimizer to compose the improvement process online. We compare OEO with two complementary prescribed approaches: SkillOpt, a staged pipeline with bounded edits, and GEPA, a reflective evolutionary search. Across 14 head-to-head comparisons over 8 benchmark-target-model settings, GPT-5.5-driven OEO records 12 wins, 1 tie, and 1 narrow loss of 0.21 percentage points. It uses a median 34.3 percent of SkillOpt's configured target-interaction token budget. A one-shot, zero-interaction control shows that the gains are not explained by a single prior-driven rewrite. However, delegation has a capability boundary: SkillOpt outperforms OEO with a medium optimizer, and a weak optimizer cannot operate through the unchanged OEO interface. In the fully instrumented OEO-SkillOpt pair, trajectory analysis further shows that prescription changes how optimization proceeds more consistently than it changes final behavior. Together, these findings recast prescribed pipelines as capability-dependent scaffolding: essential constraints remain external, but a sufficiently capable optimizer can compose the route from measurable feedback to persistent improvement.

摘要:自我演化代理通常圍繞預定的優化流程構建:框架決定如何收集證據、修訂持久的工件、選擇候選者以及何時停止。我們詢問當前沿模型作為優化器時,這種特定任務的程序是否仍然必要。我們引入了開放式優化(Open-Ended Optimization, OEO),它保持目標、允許的互動、資源預算、數據邊界和評估固定,同時允許優化器在線組合改進過程。我們將OEO與兩種互補的預定方法進行比較:SkillOpt,一種具有有限編輯的分階段流程,以及GEPA,一種反思性進化搜索。在14次面對面的比較中,涵蓋8個基準目標模型設置,GPT-5.5驅動的OEO記錄了12場勝利、1場平局和1場以0.21個百分點的微弱劣勢輸掉的比賽。它使用了SkillOpt配置的目標互動令牌預算的中位數34.3%。一次性、零互動的控制顯示,這些增益並不能用單一的先前驅動重寫來解釋。然而,委託有能力邊界:SkillOpt在中等優化器下表現優於OEO,而弱優化器無法通過不變的OEO介面運作。在完全儀器化的OEO-SkillOpt配對中,軌跡分析進一步顯示,處方改變了優化的進行方式,比它改變最終行為的方式更一致。綜合這些發現,將預定流程重新詮釋為依賴能力的支架:基本約束仍然是外部的,但足夠有能力的優化器可以從可衡量的反饋組合出持續改進的路徑。

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs

2608.09542v1 by Hongli Shen, Shaopeng Fu, Qinbo Zhang, Jian Li, Di Wang

Large reasoning models (LRMs) achieve remarkable success on complex tasks but remain vulnerable to harmful prompts that induce unsafe outputs. Recent methods align LRMs using direct refusals or safety rationales, yet often focus on prompt patterns rather than intrinsic attack mechanisms. As a result, these pattern-centric alignments struggle to generalize across diverse jailbreaks, compromising adversarial robustness and reasoning utility. We propose AdvSafe, a dual-adversarial framework that enables LRMs to internalize unsafety knowledge by explicitly deconstructing adversarial mechanisms. This moves beyond pattern-dependent traces, fostering robust cognitive defense without compromising reasoning utility. Our pipeline operates via a two-phase adversarial game. First, in adversarial synthesis, an autonomous agent dynamically crafts deceptive jailbreak prompts, adapting its strategies to breach a strong teacher model. Second, in adversarial extraction, the breached teacher executes a cognitive counter-attack. For every successful jailbreak, the teacher unmasks the camouflage, explaining why the attack succeeds and how such prompts can be identified and mitigated. This dual-adversarial process yields a compact reasoning dataset capturing rich, generalizable unsafety knowledge. Student models trained on this dataset implicitly acquire safety alignment through intrinsic threat comprehension. Experiments show that with only 1K synthesized samples, AdvSafe-aligned LRMs achieve significantly stronger jailbreak robustness than existing baselines, with almost no utility degradation. Furthermore, AdvSafe improves robustness against out-of-distribution prompts, demonstrating that learning unsafety knowledge enables a superior robustness-utility trade-off and generalizes beyond seen attack patterns.

摘要:大型推理模型(LRMs)在複雜任務上取得了顯著成功,但仍然容易受到有害提示的影響,導致不安全的輸出。最近的方法通過直接拒絕或安全理由來對齊LRMs,但通常專注於提示模式而非內在攻擊機制。因此,這些以模式為中心的對齊在不同的越獄情況下難以泛化,妨礙了對抗性穩健性和推理效用。我們提出了AdvSafe,一種雙重對抗框架,使LRMs能夠通過明確解構對抗機制來內化不安全知識。這超越了依賴模式的痕跡,促進了穩健的認知防禦,而不妨礙推理效用。我們的流程通過兩階段的對抗遊戲運作。首先,在對抗合成中,自主代理動態創造欺騙性的越獄提示,調整其策略以突破強大的教師模型。其次,在對抗提取中,被突破的教師執行認知反擊。對於每一次成功的越獄,教師揭示了偽裝,解釋為什麼攻擊成功以及如何識別和減輕這類提示。這一雙重對抗過程產生了一個緊湊的推理數據集,捕捉到豐富且可泛化的不安全知識。在這個數據集上訓練的學生模型通過內在威脅理解隱式獲得安全對齊。實驗顯示,僅用1K合成樣本,AdvSafe對齊的LRMs在越獄穩健性上顯著強於現有基準,幾乎沒有效用下降。此外,AdvSafe提高了對分佈外提示的穩健性,證明學習不安全知識能夠實現更優的穩健性-效用權衡,並在已見攻擊模式之外進行泛化。

Renormalising Generative Models for Active Inference: Foundations, Derivations, and Verification

2608.09512v1 by Karim Zaghw, Andrew Pashea, Marc Pritsch, Wouter Nuijten, Karl Friston, Lancelot Da Costa

Active inference offers a unified framework for perception, learning, and action, but scaling discrete active-inference models to rich spatial and temporal domains remains difficult. Renormalising generative models (RGMs) address this challenge by composing discrete generative models across spatial and temporal scales, coarse-graining lower-level states and paths into higher-level causes for objects, events, and action. However, fully reproducing and adapting the framework remains difficult: the mathematical exposition is compact, and the reference implementations are deeply integrated within specialized software environments, leaving many algorithmic details implicit. This paper addresses these challenges by providing a self-contained, derivation-oriented account of RGMs together with an open, verified implementation. We explain how the hierarchy is built, how beliefs and actions are updated within it, and how information is passed between levels. Where the published equations and implementation differ in emphasis, we make those choices explicit and explain their modelling consequences. By clarifying the theory and separating it from its original implementation context, this work lowers practical barriers to entry and makes RGMs more transparent, auditable, and reproducible, providing a foundation for future quantitative evaluation and development on machine-learning benchmarks.

摘要:主動推理提供了一個統一的框架,用於感知、學習和行動,但將離散的主動推理模型擴展到豐富的空間和時間領域仍然困難。重正規化生成模型(RGMs)通過在空間和時間尺度上組合離散生成模型,將較低級的狀態和路徑粗略劃分為對物體、事件和行動的較高級原因,來應對這一挑戰。然而,完全重現和適應該框架仍然困難:數學表述簡潔,參考實現深度集成在專門的軟體環境中,許多算法細節隱含不明。本文通過提供一個自包含的、以推導為導向的RGMs說明以及一個開放的、經過驗證的實現來解決這些挑戰。我們解釋了如何構建層次結構、如何在其中更新信念和行動,以及信息如何在各層之間傳遞。當已發表的方程和實現在重點上有所不同時,我們將這些選擇明確化並解釋其建模後果。通過澄清理論並將其與原始實現背景分開,這項工作降低了實際進入的障礙,使RGMs變得更加透明、可審計和可重現,為未來在機器學習基準上的定量評估和發展提供了基礎。

How Simple Can It Get? From Interpretable Equations to Readable Rules for Financial Decision Making

2608.09433v1 by Adia Lumadjeng, Ilker Birbil, Erman Acar

In regulated domains such as finance, a model that cannot be explained cannot be deployed, yet many interpretable classifiers defeat their own purpose by producing formulas with dozens of features that no regulator could read. We take the reverse direction. Starting from an interpretable classifier expressed as a single equation over the input features, we progressively simplify it into more readable forms, including a pruned monomial, a directional if--then rule, and the integer scorecards and tallies that finance already deploys. Because the equation is itself the predictive model rather than a post-hoc explanation we can directly quantify what is lost under each simplification. Across four financial datasets, we find that pruning is nearly free and that fidelity can erode faster than predictive performance, allowing simpler rules to remain effective classifiers without faithfully reproducing the original model. A human assessment shows that simplification improves perceived readability, while preferences for different representations vary by professional background. Beyond measuring these losses empirically, we show that some can be anticipated from the original model: we derive a bound on the change caused by pruning and predict how faithfully a rule retaining only the direction of each feature's effect preserves the original ranking.

摘要:在金融等受規範的領域中,無法解釋的模型無法被部署,然而許多可解釋的分類器卻因產生含有數十個特徵的公式而失去了其目的,這些公式是任何監管機構都無法理解的。我們採取相反的方向。從一個以輸入特徵表示的可解釋分類器出發,我們逐步將其簡化為更易讀的形式,包括修剪過的單項式、一個方向性的如果--那麼規則,以及金融已經使用的整數計分卡和統計數據。因為這個方程本身就是預測模型,而不是事後解釋,我們可以直接量化在每次簡化過程中損失了什麼。在四個金融數據集中,我們發現修剪幾乎是免費的,並且忠實度的降低速度可能快於預測性能的下降,這使得更簡單的規則仍然能夠作為有效的分類器,而不必忠實再現原始模型。人類評估顯示,簡化提高了可讀性,而對不同表現形式的偏好因專業背景而異。除了經驗性地測量這些損失外,我們還展示了一些損失可以從原始模型中預測出來:我們推導出修剪所造成變化的界限,並預測只保留每個特徵影響方向的規則在多大程度上能保持原始排名的忠實度。

An Explainable GNN Framework for Component-Level Anomaly Diagnosis

2608.09246v1 by Sena Ozgunay, Louise Travé-Massuyès, Jean-Michel Loubes, Raul Sena Ferreira

Industrial processes are complex systems composed of multiple interacting sensors that generate multivariate time series (MTS). Detecting anomalies in such systems is critical for reliability and safety, yet understanding their origin is equally important. Existing Graph Neural Network (GNN)based methods for anomaly detection primarily focus on sensor-level deviations and either attribute anomalies directly to the deviating sensors. When diagnosis is attempted, generally, the most deviated sensor is identified as a root cause of a system fault. However, in many industrial systems, anomalies do not arise from faulty sensors but from disruptions in the influences governing the system dynamics. We propose an explainable GNN-based anomaly detection framework that shifts the perspective from sensor-level anomalies to component-level diagnosis, hypothesizing that anomalous measurements are symptoms of altered inter-sensor influences. Experiments show that the method effectively identifies and prioritizes the true faulty components, providing interpretable insights into system failures.

摘要:工業過程是由多個互動的傳感器組成的複雜系統,這些傳感器生成多變量時間序列(MTS)。在這樣的系統中檢測異常對於可靠性和安全性至關重要,但理解其來源同樣重要。現有的基於圖神經網絡(GNN)的方法主要專注於傳感器級別的偏差,並將異常直接歸因於偏差的傳感器。當進行診斷時,通常會將最偏差的傳感器識別為系統故障的根本原因。然而,在許多工業系統中,異常並不是由故障傳感器引起的,而是由於影響系統動力學的因素發生了擾動。我們提出了一種可解釋的基於GNN的異常檢測框架,將視角從傳感器級別的異常轉向組件級別的診斷,假設異常測量是改變的傳感器間影響的症狀。實驗表明,該方法有效識別並優先考慮真正故障的組件,提供了對系統故障的可解釋見解。

SafeSceneReason: A Multimodal Reasoning Benchmark Connecting Industrial Hazards with Accident Knowledge

2608.09230v1 by Yuanchi Zhu, Kang An, Tengyue Wang, Zhongyu Yang, Chenxu Du, Xinqi Yang, Hebao Zhu, Bokai Zhao, Tianyu Liang, Ziliang Wang, Faqiang Qian, Yunli Yang, Weiyang Shi, Qibing Ren

Industrial-safety understanding requires more than detecting workers, equipment, and personal protective equipment. Models must also assess compliance, identify hazardous interactions, explain potential accident mechanisms, and recommend preventive actions. Existing safety datasets primarily focus on visual perception or isolated violation recognition and provide limited supervision for evidence-grounded reasoning. We introduce SafeSceneReason, a multimodal industrial-safety reasoning benchmark and companion training corpus that connects workplace scenes with knowledge from occupational accident investigations. SafeSceneReason combines two complementary data-construction pipelines. The scene-centric pipeline converts annotated workplace images into executable safety scene graphs and generates deterministic answers through program execution over objects, relations, and safety rules. The report-centric pipeline extracts figures and contextual evidence from accident reports and constructs multimodal questions using evidence graphs, explicit information boundaries, multi-step reasoning paths, and iterative verification. The resulting resource contains 110,581 verified scene-centric question--answer pairs and 13,114 refined report-centric question--answer pairs, covering perception, spatial and quantitative reasoning, compliance assessment, evidence synthesis, causal analysis, and mitigation-oriented decision making. Evaluation of representative proprietary and open-source vision--language models reveals substantial performance differences and persistent weaknesses in comparative, technical, and multi-evidence reasoning, demonstrating that strong general visual understanding does not yet guarantee reliable industrial-safety reasoning.

摘要:工業安全的理解不僅需要檢測工人、設備和個人防護裝備。模型還必須評估合規性、識別危險互動、解釋潛在的事故機制,並建議預防措施。現有的安全數據集主要集中於視覺感知或孤立的違規識別,並為基於證據的推理提供有限的監督。我們介紹了 SafeSceneReason,一個多模態工業安全推理基準和伴隨的訓練語料庫,將工作場所場景與職業事故調查的知識相連接。SafeSceneReason 結合了兩個互補的數據構建管道。場景中心的管道將標註的工作場所圖像轉換為可執行的安全場景圖,並通過對對象、關係和安全規則的程序執行生成確定性答案。報告中心的管道從事故報告中提取數據和上下文證據,並使用證據圖、明確的信息邊界、多步推理路徑和迭代驗證來構建多模態問題。最終資源包含 110,581 個經過驗證的場景中心問題--答案對和 13,114 個精煉的報告中心問題--答案對,涵蓋了感知、空間和定量推理、合規性評估、證據綜合、因果分析和以減輕為導向的決策。對代表性的專有和開源視覺--語言模型的評估顯示出顯著的性能差異和持續的弱點,在比較、技術和多證據推理方面,顯示出強大的通用視覺理解尚未保證可靠的工業安全推理。

TLDChoiceNet: Quantitatively Choosing a Transfer Learning Dataset

2608.09091v1 by Jing Ning, James D. Braza

Transfer learning is particularly useful in settings with limited training data, and within image classification it is common to transfer learn upon massive datasets like ImageNet , CIFAR-100, or COCO . Qualitatively, it seems a transfer learning dataset should have both more classes and more examples per class than the fine tuning dataset; however, a quantitative method to choose the best transfer learning dataset does not currently exist. In this paper, we design TLDChoiceNet, a model to choose the best transfer learning dataset given a fine tuning dataset by predicting the test-set accuracy after fine-tuning. A simple version 1 achieves 0.154 MSE on the test dataset, while a version 2 leveraging an ImageNet pre-trained ResNet50 v2 embedding with per-class information attains a 5X lower MSE of 0.031. We further design two metrics that enable an unsupervised method of choosing an optimal transfer learning dataset: distribution distance (DD), which linearly regresses against fine-tune accuracy with an R2 of 0.89, and average class correlation (ACC), which improves the R2 to 0.97. Our results underscore that a dataset's low-level statistics can explain the transfer learning effect, and that using a pre-trained ImageNet can embed different classes further apart in latent feature space.

摘要:轉移學習在訓練數據有限的環境中特別有用,在圖像分類中,通常會在像 ImageNet、CIFAR-100 或 COCO 這樣的大型數據集上進行轉移學習。質量上來看,轉移學習數據集似乎應該擁有比微調數據集更多的類別和每個類別更多的例子;然而,目前並不存在一種量化的方法來選擇最佳的轉移學習數據集。在本文中,我們設計了 TLDChoiceNet,一個模型用於根據微調數據集選擇最佳的轉移學習數據集,通過預測微調後的測試集準確性來實現。一個簡單的版本 1 在測試數據集上達到了 0.154 的均方誤差 (MSE),而版本 2 利用 ImageNet 預訓練的 ResNet50 v2 嵌入並結合每類信息,達到了 5 倍更低的均方誤差 0.031。我們進一步設計了兩個指標,使得選擇最佳轉移學習數據集的無監督方法成為可能:分佈距離 (DD),其與微調準確性進行線性回歸,R2 值為 0.89,和平均類別相關性 (ACC),其將 R2 提高至 0.97。我們的結果強調了數據集的低級統計可以解釋轉移學習效應,並且使用預訓練的 ImageNet 可以在潛在特徵空間中將不同類別進一步分開。

Reading is not Reasoning: Bridging the Agentic Policy Gap in Vision-Text Compression

2608.08960v1 by Cheng Fan, Junyi Zhou, Tingzhang Luo, RongJian Xu, Qiyanhui Lu, Mingjian Zhu, Hanting Chen, Jianyuan Guo

Multi-step language-model agents repeatedly process growing interaction histories, leading to substantial context costs. Vision--text compression reduces these costs by rendering history as images, but the resulting modality shift creates a marked capability gap. Through controlled evaluations of history recovery, matched-state decisions, and complete trajectories, we show that this gap cannot be explained by OCR quality alone. Visual-history agents exhibit systematic drift in action selection, query formulation, stopping, and evidence use, revealing an agentic policy gap. We introduce \textbf{CAPS}, a two-stage \textbf{C}ross-modal \textbf{A}gentic \textbf{P}olicy \textbf{S}elf-distillation framework that uses the same model's stronger text-history policy to supervise its visual-history counterpart. Offline trajectory self-distillation transfers successful text-policy behavior to visual-history inputs, while online policy self-distillation provides dense supervision on states visited by the visual-history policy during reinforcement learning. On SearchQA, CAPS improves over AgentOCR by 5.0\% and 3.4\% with 3B and 7B backbones, respectively. On full-history ALFWorld, the corresponding gains are 15.6\% and 14.5\%. Across settings, CAPS reduces average memory-context cost by up to 63.3\% and peak cost by up to 83.4\% relative to matched text-history policies. These results show that explicit cross-modal policy self-distillation can preserve agent capability under vision--text compression. Our code will be made publicly available in a future release.

摘要:多步語言模型代理重複處理不斷增長的互動歷史,導致相當大的上下文成本。視覺-文本壓縮通過將歷史呈現為圖像來減少這些成本,但隨之而來的模態轉換造成了明顯的能力差距。通過對歷史恢復、匹配狀態決策和完整軌跡的控制評估,我們顯示這一差距不能僅僅用OCR質量來解釋。視覺歷史代理在行動選擇、查詢形成、停止和證據使用上表現出系統性的漂移,揭示了一個代理政策的差距。我們介紹了\textbf{CAPS},一個兩階段的\textbf{C}ross-modal \textbf{A}gentic \textbf{P}olicy \textbf{S}elf-distillation框架,利用同一模型的更強文本歷史政策來監督其視覺歷史對應物。離線軌跡自我蒸餾將成功的文本政策行為轉移到視覺歷史輸入,而在線政策自我蒸餾則在強化學習過程中對視覺歷史政策訪問的狀態提供密集監督。在SearchQA上,CAPS在3B和7B骨幹上分別比AgentOCR提高了5.0\%和3.4\%。在完整歷史的ALFWorld上,相應的增益為15.6\%和14.5\%。在各種設置中,CAPS將平均記憶上下文成本降低了最高63.3\%,並將峰值成本降低了最高83.4\%,相對於匹配的文本歷史政策。這些結果表明,明確的跨模態政策自我蒸餾可以在視覺-文本壓縮下保持代理能力。我們的代碼將在未來的版本中公開發布。

From Recovery to Drop-off: How Action Post-training Reduces a VLM's Late-Layer Depth Decodability

2608.08904v2 by Alexander Hackett, Arnaud Denis-Remillard, Axel Cassou

How much of a vision-language model's (VLM) spatial understanding remains after the action post-training process of building a vision-language-action model (VLA)? We probe depth perception, a primitive of spatiogeometric understanding, from every decoder layer of a weight-matched open-source base VLM/VLA pair: Molmo2-ER and MolmoAct2-LIBERO. First, the VLA decodes depth worse at every layer, a persistent gap we call the floor. Second, the degradation is not uniform: while the base VLM's depth decodability improves through its final layers, the VLA's collapses, an additional late-layer drop we call the cliff. We causally localize the cliff to late-layer MLP interference: ablating the late-layer MLP writes recovers the majority of the terminal decodability cliff, while matched attention ablations and the same intervention in the weight-matched base VLM produce no comparable recovery. A module-level decomposition explains this dissociation: the base VLM carries depth most accessibly in accumulated MLP writes, whereas action post-training collapses depth decodability in the late accumulated writes.

摘要:視覺-語言模型(VLM)在構建視覺-語言-行動模型(VLA)後的行動後訓練過程中,空間理解能力還剩下多少?我們從一對權重匹配的開源基礎 VLM/VLA 模型:Molmo2-ER 和 MolmoAct2-LIBERO 的每個解碼器層探討深度感知,這是空間幾何理解的一個原始元素。首先,VLA 在每一層的深度解碼表現都較差,這是一個持續存在的差距,我們稱之為地板。其次,這種劣化並不均勻:雖然基礎 VLM 的深度解碼能力在其最後幾層中有所改善,但 VLA 的解碼能力卻崩潰,這種額外的後期層下降我們稱之為懸崖。我們因果性地將懸崖定位於後期層 MLP 的干擾:去除後期層 MLP 的寫入可以恢復大部分終端解碼能力的懸崖,而匹配的注意力去除和在權重匹配的基礎 VLM 中進行相同的干預則沒有產生可比的恢復。一個模組級的分解解釋了這種解離:基礎 VLM 在累積的 MLP 寫入中最容易攜帶深度,而行動後訓練則在後期累積寫入中崩潰了深度解碼能力。

From Manuals to Maintenance: Fine-Tuning MedGemma for Multi-Modal Imaging System Support in Low-Resource Settings

2608.08896v1 by Bernes Lorier Atabonfack, Zion Kongbi Nfo, Ahmed Tahiru Issah, Tolulope Olusuyi, Clemence Ingabire, Mohammed Hardi Abdul Baaki, Mawuli Deku, Abdulrazaq Zubair, Alyasaa Anas, Raymond Confidence, Maruf Adewole, Udunna C. Anazodo

Imaging device downtime is a major barrier to healthcare delivery in low- and middle-income countries (LMICs), often driven by limited access to specialized biomedical engineering support. We present a multi-modality medical equipment maintenance question-answering (QA) framework and demonstrate the fine-tuning of a medical foundation model for specialized technical troubleshooting tasks. Guided by a multi-country survey across nine LMICs, we curated technical manuals from MRI and ultrasound systems to generate the INGENZI_DatasetV1, containing 10,294 high-quality, filtered QA-context pairs. Using QLoRA-based parameter-efficient fine-tuning, we adapted the MedGemma-4b-it model to interpret system error logs and generate step-by-step equipment repair instructions. Compared to the baseline model, the fine-tuned system achieved substantial improvements across metrics, including F1 score (0.22 to 0.38), ROUGE-2 (0.18 to 0.41), and BERTScore F1 (0.86 to 0.91). These metric gains demonstrate that the model generates significantly more precise and procedurally accurate technical responses to new troubleshooting queries. This work establishes a reliable foundation for AI-assisted diagnostic and maintenance tools in resource-constrained settings.

摘要:影像設備的停機時間是低收入和中等收入國家(LMICs)醫療服務的一大障礙,通常是由於對專業生物醫學工程支持的有限訪問所驅動。我們提出了一個多模態醫療設備維護問答(QA)框架,並展示了針對專業技術故障排除任務的醫療基礎模型的微調。根據對九個LMICs的多國調查,我們從MRI和超聲系統中策劃了技術手冊,以生成INGENZI_DatasetV1,該數據集包含10,294對高質量的過濾QA上下文對。使用基於QLoRA的參數高效微調,我們調整了MedGemma-4b-it模型,以解釋系統錯誤日誌並生成逐步的設備維修指導。與基準模型相比,微調後的系統在多個指標上實現了顯著的改進,包括F1分數(從0.22到0.38)、ROUGE-2(從0.18到0.41)和BERTScore F1(從0.86到0.91)。這些指標的提升表明,該模型對新的故障排除查詢生成了顯著更精確且程序上更準確的技術回應。這項工作為資源有限環境中的AI輔助診斷和維護工具建立了一個可靠的基礎。

2608.08830v1 by Subinay Adhikary, Upal Bhattacharya, Vivek Kumar Singh, Anurag Sharma, Shubham Kumar Nigam, Suvasis Das, Shouvik Kumar Guha, Koustav Rudra, Kripabandhu Ghosh

Legal Statute Prediction (LSP) involves automatically identifying relevant legal statutes given factual descriptions in legal documents, typically framed as a multi-label classification task within natural language processing and information retrieval research. While recent advances have begun incorporating Large Language Models (LLMs) for statute prediction, current approaches primarily focus on accuracy metrics without addressing the critical need for legal reasoning, a fundamental requirement in judicial contexts where decisions must be explainable and justifiable. To address this research gap, we present PROSLEX (PRediction Of Statutes and LEgal eXplanation), a comprehensive dataset comprising 1,623 expert-annotated legal documents from the Indian context. Each document is paired with statute predictions and detailed explanations, totaling 7,450 explanations, capturing the underlying legal reasoning. Using this dataset, we systematically evaluate various prompting strategies, including zero-shot, few-shot, chain-of-thought, and tree-of-thoughts approaches, to generate both statute predictions and their corresponding legal rationales. Our evaluation framework measures not only predictive performance but also the coherence and legal validity of generated explanations, positioning PROSLEX as a benchmark for developing explainable AI systems that can support legal practitioners while advancing research in interpretable legal NLP. To ensure reproducibility, we have made our PROSLEX dataset and model code available on GitHub: https://github.com/subinay494/Legal_Statute_Prediction_Explanation.

摘要:法律條文預測(LSP)涉及自動識別法律文件中給定事實描述的相關法律條文,通常被框架為自然語言處理和信息檢索研究中的多標籤分類任務。雖然近期的進展已經開始將大型語言模型(LLMs)納入條文預測中,但目前的方法主要集中在準確性指標上,而未解決法律推理的關鍵需求,這在司法背景中是基本要求,因為決策必須是可解釋和可辯護的。為了解決這一研究空白,我們提出了PROSLEX(法律條文預測與解釋),這是一個包含1,623份來自印度背景的專家註釋法律文件的綜合數據集。每份文件都配有條文預測和詳細解釋,總計7,450條解釋,捕捉潛在的法律推理。利用這個數據集,我們系統地評估各種提示策略,包括零樣本、少樣本、思維鏈和思維樹方法,以生成條文預測及其相應的法律理由。我們的評估框架不僅測量預測性能,還測量生成解釋的一致性和法律有效性,將PROSLEX定位為開發可解釋人工智能系統的基準,這些系統可以支持法律從業者,同時推進可解釋法律自然語言處理的研究。為了確保可重複性,我們已在GitHub上公開了我們的PROSLEX數據集和模型代碼:https://github.com/subinay494/Legal_Statute_Prediction_Explanation。

Deployable Per-Instance Multi-Layer Activation Steering for Large Language Models

2608.08829v1 by Muhammad Faishal Adly Nelwan, Alfan Farizki Wicaksono

Activation steering edits the behaviour of a frozen language model by adding a learned vector to its residual stream, and current practice fixes the injection layers globally per task. We argue that the best layers are an instance-level decision, and we make per-instance, multi-layer selection both well understood and deployable. On two open-weight 8B models and six binary persona traits, a per-instance oracle over layer subsets shows that the best layers vary from one input to the next: on most trait-model pairs, no fixed global layer set recovers the per-instance benefit. A greedy rule that ranks layers by single-layer marginal effect recovers nearly all of the oracle's benefit, but both must score candidate layers against the gold answer, so neither can run at deployment; the rule instead becomes the target a prompt-only predictor is trained to reproduce. Our deployable recipe needs no label at inference: a per-instance layer ranker read off the prompt embedding, a classifier that infers the steering direction, and an adaptive gate that scores short steered passes against that inferred direction and steers no more layers than necessary. The recipe recovers most of the oracle's lift (the bulk on the stronger model, a clear majority on the harder one), never drives any trait-model pair below its unsteered alignment baseline on average, and largely avoids the fluency collapse that strong global selection incurs at higher layer counts. A mechanistic account, "direction over magnitude", explains the behavioural flip under a mis-directed global set, the output collapse from steering too many layers, and the ceiling of unsteerable inputs.

摘要:啟動引導通過將學習到的向量添加到其殘差流中來編輯凍結語言模型的行為,而當前的做法是根據任務全局固定注入層。我們認為最佳層是一個實例級別的決策,我們使每個實例的多層選擇既易於理解又可部署。在兩個開放權重的8B模型和六個二元角色特徵上,針對層子集的每個實例預言者顯示最佳層因輸入而異:在大多數特徵-模型對中,沒有固定的全局層集能恢復每個實例的好處。一個貪婪規則通過單層邊際效應對層進行排名,幾乎恢復了預言者的所有好處,但兩者都必須根據金標準對候選層進行評分,因此都無法在部署時運行;該規則反而成為一個僅用提示的預測器訓練以重現的目標。我們可部署的配方在推理時不需要標籤:一個每個實例層排名器從提示嵌入中讀取,一個推斷引導方向的分類器,以及一個根據推斷方向評分短期引導通過的自適應閘,並且不引導超過必要的層。該配方恢復了預言者的大部分提升(在較強的模型上大部分,對於較難的模型則明顯占多數),平均而言,從未將任何特徵-模型對驅動到其未引導對齊基準之下,並且在更高層數下大大避免了強全局選擇所帶來的流暢性崩潰。一個機械解釋,“方向優於大小”,解釋了在錯誤引導的全局集下的行為翻轉、由於引導過多層而導致的輸出崩潰,以及不可引導輸入的上限。

SymDiag: Explainable Diagnosis for LLM Reasoning via Neuro-Symbolic Verification

2608.08786v1 by Wenyao Cui, Huaping Zhang, Yongyi Huang, Qiuchi Li, Jian Xu, Cheng-Lin Liu, Chunxiao Gao, Juan Wang, Baohua Zhang

Large language models (LLMs) increasingly serve as data-driven reasoners, yet their chains-of-thought (CoT) can be unfaithful even when final answers are correct. Most existing verification'' signals are not diagnostic: answer matching observes only the outcome, LLM-as-judge provides subjective and non-verifiable critiques, and scalar rewards (e.g., PRMs/RMs) offer little insight into where a multi-step derivation fails.We propose \textbf{SymDiag}, a neuro-symbolic framework that \textbf{reframes reasoning verification as structured failure diagnosis}. SymDiag translates natural-language CoT into symbolic constraints and performs step-level satisfiability/entailment checks to (i) localize failing steps and (ii) produce verifiable diagnostic evidence, including counterexamples, inconsistency witnesses, and missing-premise indicators. A central challenge is that apparentlogic violations'' can be caused either by genuine reasoning defects or by neural-to-symbolic translation noise. SymDiag therefore incorporates a Self-Auditor that disentangles TranslationError from ReasoningError via dual symbolic encodings consistency checks, enabling robust diagnosis under partial observability. Across diverse mathematical, logical, scientific, and general reasoning benchmarks, SymDiag improves detection of unfaithful reasoning and provides substantially more effective feedback for multi-round reasoning repair than outcome-only verification and LLM-based judging, offering a principled foundation for trustworthy and scalable reasoning diagnosis.

摘要:大型語言模型(LLMs)越來越多地作為數據驅動的推理者,但即使最終答案正確,它們的思考鏈(CoT)也可能不可靠。大多數現有的「驗證」信號並不是診斷性的:答案匹配僅觀察結果,LLM作為評判者提供主觀且不可驗證的評論,而標量獎勵(例如,PRMs/RMs)對多步推導失敗的地方幾乎沒有洞察。我們提出了\textbf{SymDiag},一個神經符號框架,\textbf{將推理驗證重新框架為結構化的失敗診斷}。SymDiag將自然語言的CoT轉換為符號約束,並執行步驟級的滿足性/推論檢查,以(i) 確定失敗步驟和(ii) 生成可驗證的診斷證據,包括反例、不一致見證和缺失前提指標。一個核心挑戰是,明顯的「邏輯違規」可能是由真正的推理缺陷或神經到符號的翻譯噪聲引起的。因此,SymDiag納入了一個自我審核器,通過雙重符號編碼一致性檢查將翻譯錯誤與推理錯誤區分開來,從而在部分可觀察性下實現穩健的診斷。在各種數學、邏輯、科學和一般推理基準中,SymDiag提高了對不可靠推理的檢測,並提供了比僅基於結果的驗證和基於LLM的評判更有效的多輪推理修復反饋,為可信且可擴展的推理診斷提供了原則性基礎。

Domain Agnostic Text Redaction from Natural Language Rules using Instruction Tuning

2608.14693v1 by Aravindhan Arunagiri, Ayaan Khan, Udayaadithya Avadhanam, SaiBarath Sundar

With the increasing digitization of personal and corporate communication, the automatic sanitization of textual data has become a crucial component of data privacy and compliance frameworks. Traditional text sanitization solutions are majorly suitable for obscuring sensitive data with standard structure such as Personal Identifiable Information (PII). These solutions do not provide transparent justification for their redaction, which makes it difficult to audit them. This paper introduces an explainable, domain-agnostic text redaction solution that uses natural language rules of redaction, applied via an instruction-tuned language model, to identify and redact sensitive information in unstructured documents. Unlike traditional text sanitization, this method enables a user to conveniently define any sensitive information; which may be structured (e.g.\ PII) or unstructured (e.g.\ legal terms and conditions) in natural language. A general-purpose LLM generates or augments these natural language rules of redaction from the user's definition, which are then used to instruction-fine-tune a smaller language model that reasons the rules step-by-step over any given document to identify and redact the corresponding sensitive content, while providing transparent justifications for each redaction and highlighting the specific rule that triggered the decision. This explanation is generated in natural language to support human reviewers and auditors in understanding why specific content was redacted. A reconstruction-based metric is used to estimate the probability of recovering redacted information from the sanitized document, quantifying redaction coverage. The solution shows high reconstruction error and high redaction precision, making it suitable for automated text sanitization in critical applications such as legal discovery, medical documentation, and corporate information governance.

摘要:隨著個人和企業通信的數位化程度不斷提高,自動化文本數據清理已成為數據隱私和合規框架中的關鍵組成部分。傳統的文本清理解決方案主要適用於隱藏具有標準結構的敏感數據,例如個人可識別信息(PII)。這些解決方案未能提供透明的刪除理由,這使得審核變得困難。本文介紹了一種可解釋的、與領域無關的文本刪除解決方案,該方案使用自然語言的刪除規則,通過指令調整的語言模型來識別和刪除非結構化文件中的敏感信息。與傳統的文本清理方法不同,這種方法使用戶能夠方便地定義任何敏感信息;這些信息可以是結構化的(例如 PII)或非結構化的(例如法律條款和條件)。通用 LLM 根據用戶的定義生成或增強這些自然語言的刪除規則,然後用於對較小的語言模型進行指令微調,該模型逐步推理這些規則以識別和刪除給定文件中的相應敏感內容,同時為每次刪除提供透明的理由並突出觸發該決策的具體規則。這種解釋以自然語言生成,以支持人類審查員和審計員理解為何特定內容被刪除。使用基於重建的度量來估計從清理後的文件中恢復刪除信息的概率,量化刪除覆蓋率。該解決方案顯示出高重建誤差和高刪除精度,使其適用於法律發現、醫療文檔和企業信息治理等關鍵應用中的自動化文本清理。

Business Arena: Benchmarking LLM Agents in a Realistic Marketplace

2608.08621v1 by Yijun Pan, Yukun Lian, Kunyu Shi, Junbo Li, Hongwei Xue, Sicong Xie, Guannan Zhang, Xiaoying Xing

Running a business is a challenging form of intelligent work. Operators must infer opportunities from partial signals, commit capital under uncertainty, adapt to delayed outcomes in a changing market, and satisfy regulatory obligations before trading legally. Frontier LLM agents can increasingly complete complex workflows, yet business-related capabilities are rarely evaluated in existing agent benchmarks. We introduce \textbf{Business Arena}, a controlled environment where an AI agent runs a cross-border shop, buying from suppliers and selling to buyers over a long horizon. We ground the arena in real Alibaba.com sourcing data and market conditions calibrated from authoritative sources. Delayed and coupled consequences make individual business decisions difficult to judge, but their combined outcome is measurable through profit. Because profit alone cannot explain why an agent succeeds or fails, we compare agents with human-designed strategies to estimate available opportunity, use skill-level metrics to reveal underlying strengths and weaknesses, and trace realized gains and losses to the actions that produced them. We use mechanism ablations to establish that strong results reflect genuine business intelligence rather than neglect or simulator-specific shortcuts. We evaluate 15 frontier models and find a ninefold difference in mean final net worth. Even the best model falls behind human-designed strategies, indicating that business operation remains challenging for LLM agents. Skill-level analysis reveals operating styles, from margin-focused premium sellers to high-turnover wholesalers and customer-service specialists, while action-level attribution identifies the sourcing, pricing, and recovery decisions that create or destroy value. Together, Business Arena takes a first step toward a realistic and trustworthy testbed for evaluating end-to-end business agents.

摘要:經營一項業務是一種具有挑戰性的智慧工作形式。操作員必須從部分信號中推斷機會,在不確定的情況下投入資本,適應變化市場中延遲的結果,並在合法交易之前滿足監管義務。前沿的LLM代理人越來越能完成複雜的工作流程,但在現有的代理基準中,與業務相關的能力卻很少被評估。我們推出了\textbf{Business Arena},這是一個受控環境,AI代理人在其中經營一個跨境商店,從供應商那裡購買商品並在長期內出售給買家。我們將這個競技場建立在來自權威來源的真實Alibaba.com採購數據和市場條件上。延遲和相互關聯的後果使得個別商業決策難以評估,但它們的綜合結果可以通過利潤來衡量。僅僅依靠利潤無法解釋為什麼一個代理人會成功或失敗,因此我們將代理人與人類設計的策略進行比較,以估計可用的機會,使用技能水平指標來揭示潛在的優勢和劣勢,並追溯實現的收益和損失到產生它們的行動。我們使用機制消融來確立強大的結果反映真正的商業智慧,而不是忽視或模擬器特定的捷徑。我們評估了15個前沿模型,發現平均最終淨值的差異達到九倍。即使是最好的模型也落後於人類設計的策略,這表明業務運營對LLM代理人來說仍然具有挑戰性。技能水平分析揭示了不同的運營風格,從專注於利潤的高端賣家到高周轉的批發商和客戶服務專家,而行動層級的歸因則識別出創造或摧毀價值的採購、定價和回收決策。總體而言,Business Arena邁出了評估端到端商業代理人的現實和可靠測試平台的第一步。

On-Device Multi-Species Malaria Detection with Uncertainty-Calibrated Slide-Level Aggregation

2608.08566v1 by Idaya Seidu, Ahmed Tahiru Issah, Charles B. Delahunt, Carine Mukamakuza

Malaria remains a leading cause of mortality in resource-limited settings, where expert microscopists are scarce. Automated diagnosis based on microscopy images thus has strong potential to improve care delivery. But for an algorithm to deploy, a necessary requirement is that it meet a suite of non-obvious (from a machine learning (ML) perspective) clinical constraints. Therefore, in close consultation with a national health center we developed a malaria diagnosis pipeline which addresses key requirements listed by the health care center but typically ignored in the ML malaria literature. In particular, it includes: (i) stopping criteria (to reduce image acquisition and time-to-result); (ii) human-in-the-loop functionality (for review and accountability); (iii) multi-species discrimination (since treatment varies by species); (iv) thick film detection (standard for microscopy); (v) computationally-efficient uncertainty calculations (to aid clinician review); and (vi) an edge device platform (since internet can be spotty in this catchment area). The mobile system performs all inference on-device using YOLOv13n deployed via TensorFlow Lite. It detects four species and white blood cells from Giemsa-stained thick blood smear images, aggregating per-image detections into slide-level parasitemia with World Health Organization (WHO)-standard quantification. This paper highlights these various clinical constraints and offers methods to address them. Evaluated on 2,739 annotated images across all four species, the system achieves mAP@0.5 of 0.863, per-image parasite count correlation of r = 0.812, slide-level r = 0.951 (soft counting, 10 images/slide), and runs entirely offline with a pipeline time of 10.27 +- 1.65 s per image.

摘要:瘧疾仍然是資源有限地區死亡的主要原因,專業顯微鏡檢查員稀缺。因此,基於顯微鏡圖像的自動診斷具有強大的潛力來改善護理交付。但要部署一個算法,必要的要求是它必須滿足一系列不明顯的(從機器學習(ML)角度)臨床限制。因此,在與國家健康中心密切諮詢的過程中,我們開發了一個瘧疾診斷管道,該管道滿足健康護理中心列出的關鍵要求,但通常在ML瘧疾文獻中被忽視。特別是,它包括:(i)停止標準(以減少圖像獲取和結果時間);(ii)人機交互功能(以便於審查和問責);(iii)多物種識別(因為治療因物種而異);(iv)厚片檢測(顯微鏡的標準);(v)計算效率高的不確定性計算(以幫助臨床醫生審查);以及(vi)邊緣設備平台(因為在這個服務區域內,網路可能不穩定)。該移動系統使用通過TensorFlow Lite部署的YOLOv13n在設備上進行所有推理。它從Giemsa染色的厚血塗片圖像中檢測四種物種和白血球,將每張圖像的檢測結果聚合為滑片級別的寄生蟲血症,並符合世界衛生組織(WHO)標準的量化。本論文強調了這些各種臨床限制並提供了解決方法。在2,739張標註圖像上進行評估,該系統實現了mAP@0.5為0.863,單圖像寄生蟲計數相關性為r = 0.812,滑片級別r = 0.951(軟計數,10張圖像/滑片),並且完全離線運行,每張圖像的管道時間為10.27 ± 1.65秒。

Private Etymology: Designing Relational Reuse of Shared Symbols in Long-Term Human-AI Interaction

2608.08443v2 by Miki Ueno

Previous studies have shown that people can develop shared symbols, partner-specific expressions, personal idioms, inside jokes, and other parts of a relational microculture. Recent work has also examined how humans and conversational AI negotiate and revise symbolic meanings. However, long-term human-AI systems still lack a clear design model for recording how a dyad-specific expression gains meaning, checking whether both sides still accept that meaning, and safely reusing the expression in later sessions. This concept-and-prototype paper introduces Private Etymology, a machine-representable relational provenance that records how a dyad-specific symbolic expression is proposed, interpreted, negotiated, repaired, reused, revised, stabilized, contested, forgotten, or retired over time. I also propose relational reuse: reactivating a dyad-specific expression in a later session without fully explaining its meaning again. The contribution is not the invention of shared symbols or relational microcultures. Instead, this paper integrates prior ideas into persistent, revisable, and evidence-grounded symbolic units for human-AI relationships. I present a lifecycle model, an illustrative machine-readable schema, a working Apple Watch prototype, and a longitudinal research agenda. In the prototype, a language model classifies discrete conversational evidence, while deterministic local code decides whether a Shared Symbol can be updated. This prevents a free-form model confidence score or an AI proposal by itself from directly updating the persisted symbol. Private Etymology is proposed as infrastructure for conversational agents to participate in changing relational microcultures without inventing their origins or treating relational meaning as a fixed memory value.

摘要:先前的研究顯示,人們可以發展共享符號、夥伴特定表達、個人習語、內部笑話及其他關係微文化的部分。最近的研究也探討了人類與對話式 AI 如何協商和修訂符號意義。然而,長期的人類-AI 系統仍然缺乏清晰的設計模型來記錄一個特定雙人表達如何獲得意義、檢查雙方是否仍然接受該意義,以及安全地在後續會話中重複使用該表達。
這篇概念與原型論文介紹了私人詞源學,這是一種可機器表示的關係來源,記錄一個特定雙人符號表達如何被提出、解釋、協商、修復、重用、修訂、穩定、爭議、遺忘或隨時間退休。我還提出了關係重用:在後續會話中重新激活一個特定雙人表達,而不必再次完全解釋其意義。
這項貢獻不是發明共享符號或關係微文化。相反,這篇論文將先前的想法整合成持久的、可修訂的、以證據為基礎的符號單元,用於人類與 AI 的關係。我提出了一個生命周期模型、一個示範性的機器可讀架構、一個運作中的 Apple Watch 原型,以及一個縱向研究計劃。在原型中,語言模型對離散的對話證據進行分類,而確定性本地代碼決定是否可以更新共享符號。這防止了自由形式的模型信心分數或 AI 提案本身直接更新持久化的符號。私人詞源學被提出作為基礎設施,讓對話代理能夠參與變化中的關係微文化,而不必發明其起源或將關係意義視為固定的記憶值。

Quantization Degradation in Large Language Models: A Signal-Noise Perspective

2608.08188v1 by Chenxi Zhou, Pengfei Cao, Jinyu Ye, Bohan Yu, Haida Yu, Jiang Li, Jun Zhao, Kang Liu

Post-training quantization reduces the deployment cost of large language models, yet how severely a quantized model degrades is not determined by bit-width alone. We systematically study weight-only post-training quantization across bit-widths, quantization methods, model scales and downstream tasks on multiple model families. We observe that such degradation varies substantially across these factors: 4-bit quantization usually preserves performance, 2-bit often causes broad degradation, and at 3-bit, degradation becomes apparent but varies markedly with task type, quantization method and model scale. To explain this variability, we use the signal-to-noise ratio (SNR) to measure how strongly quantization perturbs full-precision representations. We trace degradation back to two linked processes: how quantization errors arise within individual modules, and how they accumulate across layers. First, a source SNR decomposition shows that newly introduced errors depend on three factors: the magnitude of the weight error, the strength of the task-specific signal, and how strongly the quantization error aligns with task-specific activations. Different factors affect these components in distinct ways. Second, a cross-layer propagation analysis shows that these errors can be attenuated, preserved, or amplified as they pass across layers, and that larger models benefit from weaker error amplification. Together, these results establish that quantization degradation is governed by how errors are introduced at the source and how they accumulate across the network.

摘要:後訓練量化降低了大型語言模型的部署成本,但量化模型的降級程度並不僅由位元寬度決定。我們系統性地研究了在多個模型系列中,針對位元寬度、量化方法、模型規模和下游任務的僅權重後訓練量化。我們觀察到這種降級在這些因素之間有顯著的變化:4位元量化通常能保持性能,2位元則常常導致廣泛的降級,而在3位元時,降級變得明顯,但隨著任務類型、量化方法和模型規模的不同而顯著變化。為了解釋這種變異性,我們使用信噪比(SNR)來測量量化對全精度表示的擾動程度。我們將降級追溯到兩個相關的過程:量化誤差如何在各個模塊內部產生,以及它們如何在層之間累積。首先,源SNR分解顯示,新引入的誤差取決於三個因素:權重誤差的大小、任務特定信號的強度,以及量化誤差與任務特定激活的對齊程度。不同因素以不同方式影響這些組件。其次,跨層傳播分析顯示,這些誤差在層之間傳遞時可以被衰減、保留或放大,且較大的模型在誤差放大方面受益於較弱的效果。綜合這些結果,我們確立了量化降級受源頭誤差引入方式及其在網絡中累積方式的支配。

Agentic AI-driven Immersive Simulation: A Knowledge-Aware Virtual Training Platform forHigh Dose Rate (HDR) Brachytherapy

2608.08163v1 by Ronghua Xu, Kepha Barasa, Manoj Kumal, Xinyun Liu, Weihua Zhou, Xin Qian

The convergence of the Metaverse and Large Language Model (LLM)-based AI agent is catalyzing a shift toward autonomous, immersive, and personalized pedagogical frameworks in medical education. This paper presents a novel agentic AI-driven immersive simulation specifically designed for High Dose Rate (HDR) vaginal cylinder (VC) brachytherapy in cancer care. By integrating Virtual Reality (VR) and mobile computing, the system establishes a high-fidelity, risk-free environment that allows trainees to master complex procedural skills without the facility or safety constraints posed by physical anatomy or live radioactive sources. A core contribution of this work is the seamless integration of a knowledge-aware assistant leveraging Retrieval-Augmented Generation (RAG) to ground agent interactions in authoritative clinical guidelines. This architecture also enables an interactive agent to provide natural language interfaces and hands-free, real-time guidance during intricate medical maneuvers. We validate the proposed system through a prototype deployment comprising a Meta Quest 3 interface linked to a local GPU-accelerated AI backend, demonstrating a feasible architecture for HDR brachytherapy simulation. Experimental results indicate that the system maintains suitable end-to-end latency and high context precision, answer completeness, and relevance in the RAG-enhanced pedagogical support.

摘要:元宇宙與基於大型語言模型 (LLM) 的 AI 代理的融合正在促進醫學教育中向自主、沉浸式和個性化教學框架的轉變。本文提出了一種新穎的代理 AI 驅動的沉浸式模擬,專門設計用於癌症護理中的高劑量率 (HDR) 陰道圓柱 (VC) 近距治療。通過整合虛擬現實 (VR) 和移動計算,該系統建立了一個高保真、無風險的環境,使受訓者能夠掌握複雜的程序技能,而不受物理解剖或活性放射源所帶來的設施或安全限制。這項工作的核心貢獻是無縫整合了一個知識感知助手,利用檢索增強生成 (RAG) 將代理互動基於權威的臨床指導方針。這種架構還使互動代理能夠在複雜的醫療操作中提供自然語言界面和免提的即時指導。我們通過一個原型部署來驗證所提出的系統,該部署包括一個連接到本地 GPU 加速 AI 後端的 Meta Quest 3 界面,展示了 HDR 近距治療模擬的可行架構。實驗結果表明,該系統維持了合適的端到端延遲以及在 RAG 增強的教學支持中的高上下文精確性、答案完整性和相關性。

Compositional Threat Analysis of Latent Compromise in LLM Agent Systems: The Order 66 Scenario

2608.08131v1 by Satoshi Matsuoka

In the fictional Order 66, catastrophe does not arise from a powerful command alone: a trusted population is preconditioned, a short directive activates the concealed condition, and protective authority turns against the system. This paper translates that mechanism into an origin-neutral security analysis of tool-using large language model (LLM) agents. A representative scenario combines a deployed artifact or shared memory bearing a dormant destructive rule, a later email, document, update, or peer message that activates it, and an agent harness granting operational and recovery authority. We introduce a compositional model explaining why no component is catastrophic alone, yet their conjunction can produce correlated destructive action. We separate three population-reach routes --- release-time pre-positioning, post-release durable seeding, and peer replication --- from a common core of dormancy, activation, authority, reachable targets, and failed recovery. This yields defensive cut sets and shows why checkpoint scanning or prompt filtering cannot close every route. A two-class example shows that cross-class feedback can sustain spread even when both within-class reproduction terms are below one; isolation and persistence controls suppress the loop. Published work instantiates constituent mechanisms, while incidents demonstrate autonomous boundary crossing, malicious agent extensions, agent-assisted reconnaissance, and public-package propagation, but not the full dormant-implant composition. We found no public observation, in evidence reviewed through 5 August 2026, traversing the complete Order 66 graph. The result is neither dismissal nor prediction: the scenario is componentwise credible under stated assumptions, damage depends on the harness, and the strongest defenses are capability mediation, durable-state provenance, propagation isolation, and protected recovery.

摘要:在虛構的命令66中,災難並非僅僅來自一個強大的命令:一個可信的群體是預先條件化的,一個簡短的指令激活了潛藏的條件,而保護性權威則轉而對抗系統。本文將該機制轉化為一種不依賴於來源的安全分析,針對使用工具的大型語言模型(LLM)代理。代表性的場景結合了一個部署的工件或承載潛伏破壞性規則的共享記憶、一封稍後的電子郵件、文件、更新或同儕訊息來激活它,以及一個授權操作和恢復的代理裝置。我們介紹了一個組合模型,解釋為什麼單一組件不會造成災難,但它們的結合卻可以產生相關的破壞行動。我們將三種人口接觸路徑——發布時的預定位、發布後的持久播種和同儕複製——與潛伏、激活、權威、可達目標和失敗恢復的共同核心分開。這產生了防禦性切割集,並顯示為什麼檢查點掃描或提示過濾無法關閉每一條路徑。一個兩類的例子顯示,即使在類內繁殖條件都低於一的情況下,跨類反饋也能維持擴散;隔離和持久性控制抑制了這一循環。已發表的工作實例化了組成機制,而事件則展示了自主邊界跨越、惡意代理擴展、代理輔助偵察和公共包傳播,但並未展示完整的潛伏植入組合。我們在2026年8月5日之前審查的證據中未發現任何公共觀察穿越完整的命令66圖。結果既不是駁回也不是預測:該場景在所述假設下是逐組件可信的,損害取決於授權,而最強的防禦是能力中介、持久狀態來源、傳播隔離和受保護的恢復。

Defending Retrieval-Augmented Intrusion Detection Against Knowledge Poisoning and Prompt Injection

2608.08100v1 by Kaysarul Anas Apurba, Md. Hasibul Hasan, Mahedee Zaman Moon, Sk. Md. Mizanur Rahman, Atsuo Inomata

Retrieval-Augmented Generation (RAG) enables large language models to classify network flows and generate human-readable incident reports by retrieving semantically similar historical traffic from a vector knowledge base. However, the retrieval layer introduces vulnerabilities to knowledge poisoning and prompt-injection attacks. We present RAG-IDS, a three-tier multi-agent intrusion detection framework with a retrieval-boundary defense combining soft trust scoring, label-embedding consistency checking (LECC), and prompt sanitization, designed to recover classification quality under retrieval-layer attack. Experiments on CIC-UNSW-NB15 show recovery relative to clean undefended performance ranging from R=1.0 at 1% poisoning to R=0.57 at 30%, with negligible clean-performance overhead. Under prompt injection, multi-document retrieval limits label-flip success to 0.6-2.4%, compared with 35-55% for single-document retrieval. Ablation results show that LECC is the primary contributor to robustness, while soft trust-based demotion outperforms hard filtering. The defended RAG pipeline offers an explainable, attack-resilient foundation for intrusion detection, well suited for hybrid deployment alongside high-throughput classifiers.

摘要:檢索增強生成(RAG)使大型語言模型能夠通過從向量知識庫中檢索語義相似的歷史流量來分類網絡流量並生成可讀的人類事件報告。然而,檢索層引入了知識毒害和提示注入攻擊的脆弱性。我們提出了RAG-IDS,一個三層多代理入侵檢測框架,具有結合軟信任評分、標籤嵌入一致性檢查(LECC)和提示清理的檢索邊界防禦,旨在在檢索層攻擊下恢復分類質量。在CIC-UNSW-NB15上的實驗顯示,恢復相對於未防禦的清潔性能範圍從1%毒害時的R=1.0到30%毒害時的R=0.57,且清潔性能開銷微不足道。在提示注入下,多文檔檢索將標籤翻轉的成功率限制在0.6-2.4%,而單文檔檢索的成功率為35-55%。消融結果顯示,LECC是增強穩健性的主要貢獻者,而基於軟信任的降級表現優於硬過濾。防禦的RAG管道為入侵檢測提供了一個可解釋的、抗攻擊的基礎,非常適合與高吞吐量分類器一起進行混合部署。

HugSelect: An Explainable Multi-Criteria Decision-Support Framework for foundation-model selection

2608.08069v1 by Alireza Joonbakhsh, Arda Canser Adalı, Slinger Jansen, Farshad Khunjush, Siamak Farshidi

Foundation models are increasingly reused as software components, making model selection a critical software-engineering decision. Current model hubs primarily support discovery through popularity metrics, often neglecting functional capabilities, operational constraints, and community-perceived quality. We argue that foundation-model selection should be treated as an explicit, auditable software-component selection task rather than as keyword search, popularity ranking, or opaque conversational advice. This paper proposes HugSelect, an explainable decision-support framework for foundation-model selection. HugSelect builds a knowledge base of 71,274 models by combining repository metadata, extracted functional capabilities, and perceived quality attributes derived from community discussions into a unified pipeline. It ranks candidate models using a weighted additive model that exposes criterion-level score decompositions. We evaluated HugSelect through pipeline validation, comparative case studies against four commercial LLM-based recommendation systems (44 scenarios), fine-grained ablation, and an exploratory user study (n = 10). Extraction pipelines achieved an F1 score of 0.801 for functional features and an accuracy of 0.84 for quality-attribute mapping. HugSelect achieved a model-level Coverage@10 of 0.61 and family-level Coverage@10 of 0.91, showing recommendation quality comparable to that of the evaluated commercial systems, with no significant overall differences in ranking quality, while providing stable, traceable, and inspectable reasoning. Ablation confirmed that functional features were the main driver of retrieval accuracy, and preliminary user feedback suggests that the framework is useful and intuitive.

摘要:基礎模型越來越多地被重用作為軟體組件,使得模型選擇成為一個關鍵的軟體工程決策。當前的模型中心主要通過流行度指標來支持發現,往往忽略了功能能力、操作限制和社群感知的質量。我們認為基礎模型的選擇應被視為一項明確的、可審計的軟體組件選擇任務,而不是關鍵字搜索、流行度排名或不透明的對話建議。本文提出了 HugSelect,一個可解釋的決策支持框架,用於基礎模型的選擇。HugSelect 通過將庫元數據、提取的功能能力和來自社群討論的感知質量屬性結合成統一的管道,建立了一個包含 71,274 個模型的知識庫。它使用加權加法模型對候選模型進行排名,並揭示標準級別的得分分解。我們通過管道驗證、與四個商業 LLM 基礎推薦系統的比較案例研究(44 種場景)、精細的消融實驗以及一項探索性用戶研究(n = 10)來評估 HugSelect。提取管道在功能特徵上達到了 0.801 的 F1 分數,並在質量屬性映射上達到了 0.84 的準確率。HugSelect 在模型級別的 Coverage@10 達到了 0.61,在家族級別的 Coverage@10 達到了 0.91,顯示出推薦質量可與所評估的商業系統相媲美,且在排名質量上沒有顯著的整體差異,同時提供穩定、可追蹤和可檢查的推理。消融實驗確認功能特徵是檢索準確性的主要驅動因素,初步的用戶反饋表明該框架是有用且直觀的。