Medical
Medical
| Publish Date | Title | Authors | Homepage | Code |
|---|---|---|---|---|
| 2026-10-01 | A Comparative Explainability Framework for DeBERTa-v3 in Zero-Shot Medical Abstract Classification | Javier Diaz Esteban-Herreros et.al. | 2610.02116v1 | null |
| 2026-10-01 | Can AI Oversight Be Zero Knowledge? | Alessandro Chiesa et.al. | 2610.01995v1 | null |
| 2026-10-01 | Counterfactual Auditing of Bias in Open-Source Large Language Models for Clinical Triage | Manar Aljohani et.al. | 2610.01963v1 | null |
| 2026-10-01 | A rubric landscape for evaluating clinical reasoning in large language models: what exists, what is missing, and what needs to be combined | Zhangshu Joshua Jiang et.al. | 2610.01938v1 | null |
| 2026-10-01 | Walking the Embedding Space: Datastore Extraction from Multimodal RAG | Maria Carmen Jica et.al. | 2610.01871v1 | null |
| 2026-10-01 | On the Divergence of Accuracy and Mechanism Consistency in Time Series World Models | Haochen Zhang et.al. | 2610.01842v1 | null |
| 2026-10-01 | iADD: Improving Alignment and Diversity in Diffusion Policy Optimization | Ashok Prasad Neupane et.al. | 2610.01789v1 | null |
| 2026-10-01 | OpenMTB-Audit: Exposing Over-Refusal and Clinical Expert Perspectives in LLM-Based Molecular Tumor Board Safety Evaluation | Negin Ashrafi et.al. | 2610.01497v1 | null |
| 2026-10-01 | A Multi-Agent LLM Framework for Personalized Health Checkup Interpretation and Guidance | HyungJun Kim et.al. | 2610.01451v1 | null |
| 2026-10-01 | Generation Provenance Before Behavior Attribution: Auditing Synthetic Speech Research Objects | Sidi Chang et.al. | 2610.01378v1 | null |
| 2026-10-01 | An ontology for cross-sectoral crisis management: core and public health modules | Aldo Gangemi et.al. | 2610.01326v1 | null |
| 2026-10-01 | Model validation in machine learning: A scenario-based guide from hold-out splits to nested group cross-validation in biomedical and applied research | Mehmet Baygin et.al. | 2610.01284v1 | null |
| 2026-10-01 | When Does Exercise-Specific Joint Selection Help? An Audit of Evaluation and Control Design | Haotian Chen et.al. | 2610.01188v1 | null |
| 2026-10-01 | CineMR: Tool-Integrated Vision-Language Reasoning for Quantitative Cardiac MRI Assessment | Kunyang Li et.al. | 2610.01166v1 | null |
| 2026-10-01 | A Matched-Budget Audit Framework for Recaptioned Image-Text Supervision Distributions | Giyeong Oh et.al. | 2610.00952v1 | null |
| 2026-09-30 | Backdoor Purification for LoRA-Tuned LLMs via Null-Space Projection | Jianwei Li et.al. | 2610.00685v1 | null |
| 2026-09-30 | Worse Together: How Performance Breaks Down in Multi-User Multi-Agent Teams | Sahan Paliskara et.al. | 2610.00583v1 | null |
| 2026-09-30 | Can LLMs Reason Over Long Horizons? An Empirical Evaluation of Context Strategies for Longitudinal Clinical Reasoning | Taye Akinrele et.al. | 2610.00562v1 | null |
| 2026-09-30 | Unlearnable, or Unmeasured? On the Reliability of Difficulty Labels in RLVR | Chandak Chakma et.al. | 2609.40115v1 | null |
| 2026-09-30 | GateSPINE: Gated Cross-View Fusion for Lumbar Spine MRI Report Generation | Hoang Nguyen Van et.al. | 2609.40091v1 | null |
| 2026-09-30 | Overview of BioASQ 2026: The fourteenth BioASQ Challenge on Large-Scale Biomedical Semantic Indexing and Question Answering | Anastasios Nentidis et.al. | 2609.39975v1 | null |
| 2026-09-30 | Scores That Hold, Benchmarks That Leak: Measuring Dataset Contamination in Public Brain-Tumor MRI Classification | Bhanu Prakash Vangala et.al. | 2610.00421v1 | null |
| 2026-09-30 | Trust Is Not a Score: Runtime Assurance Contracts for High-Risk AI Agents | Serhii Zabolotnii et.al. | 2609.39717v1 | null |
| 2026-09-30 | Interpretable Synthetic Medical Tabular Data Generation for Clinical Decision Support Using Fuzzy Cognitive Maps | Michael Vasilakakis et.al. | 2610.00391v1 | null |
| 2026-09-30 | CAMOS: Coupled Oscillatory State-Space Model for Multimodal Clinical Time-Series | Maxx Richard Rahman et.al. | 2609.39484v1 | null |
| 2026-09-30 | Towards Trustworthy AI for Glioma Diagnosis: A Task-Aware Evaluation of Uncertainty Quantification | Gonzalo Esteban Mosquera Rojas et.al. | 2609.39429v1 | null |
| 2026-09-30 | EHR-RobustGym: Benchmarking and Training Agents for Robust Clinical Reasoning | Yitong Qiao et.al. | 2609.39371v1 | null |
| 2026-09-30 | Structure vs. Chain-of-Thought: Evaluating LLM Criteria Extraction for Depression Severity | Xinkai Chen et.al. | 2609.39049v1 | null |
| 2026-09-30 | An Uncertainty-Guided Digital Twin Framework for Online Adaptive Proton Therapy in Head and Neck Cancer: A Feasibility Study | Yizhou Wu et.al. | 2609.39010v1 | null |
| 2026-09-30 | Scoring Higher, Answering Worse: Mitigating Reward Hacking in Rubric-Based RL via Protocol-Level Rubrics | Maoqi Liu et.al. | 2609.38847v1 | null |
| 2026-09-30 | CRAFT: Causal Responsibility and Failure Tracing in Medical Vision Language Models | Chunzheng Zhu et.al. | 2609.38810v1 | null |
| 2026-09-29 | Sense and Sensitivity: Benchmarking LLM Clinical Triage Recommendations with Physician Experts | Abinitha Gourabathina et.al. | 2609.38600v1 | null |
| 2026-09-29 | Defining and Categorising Human-AI Interactions in Clinical Trials: A Multidimensional Human-AI Classification Approach | Sandra Woolley et.al. | 2609.38559v1 | null |
| 2026-09-29 | Personalized State-Transition-Aware Memory for Clinical Agents | Maryam Haghifam et.al. | 2609.38490v1 | null |
| 2026-09-29 | KlinikeBench: Evaluating Language Models Beyond Diagnostic Accuracy | Xueting Fang et.al. | 2609.38480v1 | null |
| 2026-09-29 | PrivMeSA: Privacy-Aware Self-Evolving Multi-Agent System for Medicine via Local-Remote LLM Collaboration | Dannong Wang et.al. | 2609.38458v1 | null |
| 2026-09-29 | Colorectal Cancer Segmentation with Adaptive Augmentation and Multi-Resolution Ensemble Models | Ümit Mert Çağlar et.al. | 2609.38419v1 | null |
| 2026-09-29 | Aegis: Generative Gradient Masking for Privacy-Preserving Medical Federated Learning | Chaoyu Zhang et.al. | 2609.38339v1 | null |
| 2026-09-29 | A Proposed Rubric for Evaluating Expressed Clinical Reasoning in Large Language Model Responses | Zhangshu Joshua Jiang et.al. | 2609.37788v3 | null |
| 2026-09-29 | Multi-Site Real-World Performance of Commercial AI for Pulmonary and Incidental Pulmonary Embolism Detection | Aawez Mansuri et.al. | 2609.37750v1 | null |
| 2026-09-29 | Spatiotemporal Hyperedges for EEG Seizure Detection and Prediction | Hyunju Kim et.al. | 2609.37730v1 | null |
| 2026-09-29 | Correct, Don't Delete: Mitigating Emergent Misalignment with Corrective Supervision | Jacob Epifano et.al. | 2609.37624v1 | null |
| 2026-09-29 | ReLMem: Learning Recurrent Memory for Longitudinal EHR Modeling | Zijie Meng et.al. | 2609.37587v1 | null |
| 2026-09-29 | Raw Imagery Impacting Your AI: Should You Care? | Adrien Dorise et.al. | 2609.38265v1 | null |
| 2026-09-29 | Do Agent Benchmarks Do What They Say? An Executable-Contract Audit of Tool-Using Agent Environments | Rohith Reddy Bellibatlu et.al. | 2609.37315v1 | null |
| 2026-09-29 | Information Bottleneck-Guided Adaptive Hypergraph Transformer for Brain Disease Diagnosis | Jingxi Feng et.al. | 2609.37220v1 | null |
| 2026-09-29 | Physics-Informed Multi-Agent Coordination for Hospital Patient Flow Optimization | Guoqing Zhang et.al. | 2609.37022v1 | null |
| 2026-09-29 | STAR-GRPO: Canonical Anchoring and Reliability-First Advantages against Representation-Dependent Reward Hacking | Wan Tian et.al. | 2609.36900v1 | null |
| 2026-09-29 | Automated Screw Planning for Reduced Pelvic Fractures Based on Statistical Shape Models and Deep Learning | Yang Gao et.al. | 2609.36847v1 | null |
| 2026-09-29 | How Medical VLMs Underutilize Their Vision Encoders: A Dermatology Perspective | Janet Wang et.al. | 2609.36557v1 | null |
| 2026-09-29 | Reliability Testing of Medical Model Performance under Distributed Deployment | Yifei Wang et.al. | 2609.36525v1 | null |
| 2026-09-29 | BRIDGE: Bilevel Retrieval-Credit-Aware Agentic Reinforcement Learning | Quan Xiao et.al. | 2609.36505v1 | null |
| 2026-09-28 | ARCagent: An Adaptive Retrieval Calibration Agent for Clinical Question Answering | Yuyan Chen et.al. | 2609.36392v1 | null |
| 2026-09-28 | Quantization Enables Private Dense Retrieval against Malicious Service Providers | Louis Tremblay Thibault et.al. | 2609.36376v1 | null |
| 2026-09-28 | ThuRunel: Dynamic Decoupling for Structured Advisory Dialogue | Yuyan Chen et.al. | 2609.36340v1 | null |
| 2026-09-28 | SCOUT: Synergizing Reasoning and Tool-Use for Computer-Use Safety | Jianxing Chen et.al. | 2609.36201v1 | null |
| 2026-09-28 | PHASE: A Physiology-Guided Hierarchical Foundation Model for Intracranial EEG | Yipeng Zhang et.al. | 2609.36087v2 | null |
| 2026-09-28 | IMC-CLINIC: Coupled Loss-Informed Newton Iterations for Clipping in Analog In-Memory Computing | Yung-Chin Chen et.al. | 2609.35586v1 | null |
| 2026-09-28 | RareDx: Controlled Knowledge Integration and Graph-Grounded Policy Optimization for Rare-Disease Diagnosis | Bo Zhang et.al. | 2609.35549v2 | null |
| 2026-09-28 | CLIMB: A Clinical Multimorbidity Benchmark for Diagnosing Co-occurring Conditions through Multiturn Conversations | Yusuf Kesmen et.al. | 2609.35462v1 | null |
| 2026-09-28 | A decision-support system applied to Law: Reasoning and explainability of the decision | Jeremy Bouche-Pillon et.al. | 2609.35370v1 | null |
| 2026-09-28 | Training-Free Clinical Reasoning through Medical Ontologies and Cognitive Mapping: A Symbolic-Probabilistic Knowledge Graph Framework | Surajit Das et.al. | 2609.35298v1 | null |
| 2026-09-28 | CarveMix-RC: Addressing Rare-Class Imbalance Through Lesion-Aware Synthetic Augmentation for Brain Metastasis Segmentation | Md Shibly Sadique et.al. | 2609.35195v1 | null |
| 2026-09-28 | DoAtlas-2: A Foundation for Self-Evolving Causal Biomedical Discovery | Yulong Li et.al. | 2609.35107v1 | null |
| 2026-09-28 | VD-DeepStack: Bridging Visual Comparison and Language Reasoning for Few-Shot Anomaly Detection | Mengyang Zhao et.al. | 2609.34949v1 | null |
| 2026-09-28 | Almost Human, Except When It Matters: VoxParity and the Decisions a Voice Should Change | Bhavik Mangla et.al. | 2609.35922v1 | null |
| 2026-09-28 | Nociception as a Control Primitive: Afferent Channels and Nociceptive Memory for Agents Deployed in One Body | Wolfgang Maass et.al. | 2609.34840v1 | null |
| 2026-09-28 | Applying Language Models in Clinical Medicine: Recent Trends and Perspectives | Erik Aerts et.al. | 2609.34780v2 | null |
| 2026-09-28 | ResonAct: Streaming Metrics for Runtime Diagnosis and Self-Healing in Multi-Agent Systems | Tarun Chintada et.al. | 2609.34701v1 | null |
| 2026-09-28 | SentZero: An Enhanced Sentence-Centric Vision-Language Pretraining for Multi-Task Zero-Shot Chest X-Ray Analysis | Hangyul Yoon et.al. | 2609.34479v1 | null |
| 2026-09-28 | VL-AcneSeg: A Vision-Language Framework for Region-Aware Acne Lesion Segmentation | Sukju Oh et.al. | 2609.34472v1 | null |
| 2026-09-28 | Evolving Support Priorities in Empathetic Reinforcement Learning | Pengyu Huang et.al. | 2609.34249v1 | null |
| 2026-09-28 | Unknown is not normal: separating language-model extraction from rule-based decision logic for clinical risk scores | Nicolás Vera Zúñiga et.al. | 2609.34112v1 | null |
| 2026-09-28 | The Devil is in the Spectrum Bias: Spectrum-Balanced Feature Matching for Robust Representation Distillation | Kuniaki Saito et.al. | 2609.34106v1 | null |
| 2026-09-28 | TRACE: Expert-Aligned ECG Representation Learning with Rigorous Benchmarking and Real-World Validation in Acute Cardiac Care | Lovely Yeswanth Panchumarthi et.al. | 2609.34088v1 | null |
| 2026-09-28 | Who Gets a Token, and What Does It Carry? Unequal Name Support and Concept Access in Large Language Models | Mir Tafseer Nayeem et.al. | 2609.34065v1 | null |
| 2026-09-28 | Large Language Models for Structured Clinical Data Analysis: Dual-Agent Grounding and Validation | Erfan D. Dehkalani et.al. | 2609.34039v1 | null |
| 2026-09-27 | Jev in Medicine: A Benchmark Evaluation | Alfredo Madrid-García et.al. | 2609.34024v2 | null |
| 2026-09-27 | EHRAdapt: Adapting Pretrained Language Models to Electronic Health Records with Semantic Priors for Rare Clinical Events | Andre R Goncalves et.al. | 2609.34007v1 | null |
| 2026-09-27 | Is your uncertainty map wrong, or is its target? Exact diagnostics for the Tweedie diagonal, and a gradient-free alternative | Vicent Ribas et.al. | 2609.33786v1 | null |
| 2026-09-27 | BIRD: Distilling Decision Boundaries into Rationales for MLLM Adaptation | Anglin Liu et.al. | 2609.33713v1 | null |
| 2026-09-27 | PPG-LM: A Photoplethysmography-Language Model with Multi-Level Clinical Alignment | Xiaoda Wang et.al. | 2609.33516v1 | null |
| 2026-09-27 | Federated Multi-Modal Human Activity Recognition using Multi-Agent Reinforcement Learning | Debasmita Dey et.al. | 2609.33492v1 | null |
| 2026-09-27 | A Cheap Verifier is Good Enough: LLM Post-training is Robust to Erroneous Rewards | Andreas Plesner et.al. | 2609.33467v1 | null |
| 2026-09-27 | MAC-Net: A Multi-Task Deep Learning Framework for Modeling Cognitive Function From Task-Based fMRI | Md. Tanvir Rahman et.al. | 2609.33440v1 | null |
| 2026-09-27 | Temporal Graph Learning of Wearable Actigraphy and Sleep Traces for Modelling Adolescent Crystallized Intelligence | Md. Tanvir Rahman et.al. | 2609.33428v1 | null |
| 2026-09-27 | Explainable Deep Learning of Resting-State Functional Connectomes Reveals Network Biomarkers of Adolescent Intelligence | Md. Tanvir Rahman et.al. | 2609.33422v1 | null |
| 2026-09-27 | QuPID: Quantum Parameter-Efficient Input-Dependent Retrieval Adaptation for Medical RAG | Hyojun Ahn et.al. | 2609.33351v1 | null |
| 2026-09-27 | CHI: A Composite Hallucination Index Unifying Entity, Relation, and Quantity Dimensions for Summarization Evaluation | Praveenkumar Katwe et.al. | 2609.33343v1 | null |
| 2026-09-27 | The Error You See Is Not the Error You Made: Progression-aware Reasoning Origin for Reasoning Error Localization | Yiguo Wang et.al. | 2609.33297v1 | null |
| 2026-09-27 | CORTEX: A Verified Experience Layer for Generalist Agents | Garapati Keerthana et.al. | 2609.33260v1 | null |
| 2026-09-27 | FOCUS: Benchmarking Retinal Model Generalization from Foundation Vision Encoders to Multimodal LLMs | David Restrepo et.al. | 2609.33158v1 | null |
| 2026-09-27 | MedRouter: Demystifying Knowledge Differences Across Medical LLMs for Routing-Based Reasoning | Lang Cao et.al. | 2609.33119v1 | null |
| 2026-09-27 | ECG-Scroll: A Long-Horizon, Streaming Benchmark and Agent Environment for Interpretation of Ambulatory Electrocardiograms | Haitao Li et.al. | 2609.33117v1 | null |
| 2026-09-27 | SemReward-VL: Semantic Reward-Guided Video-Language Adaptation for Developmental Behavior Assessment | De Jiang et.al. | 2609.33082v1 | null |
| 2026-09-27 | NutriVision: Ingredient-Conditioned Fusion and Prediction for Single-Image Food Nutrition Estimation | Aman Kumar et.al. | 2609.33076v1 | null |
| 2026-09-26 | TCMQA: A 38K-Question Traditional Chinese Medicine Benchmark with a Licensed-Practitioner Reference | Tzu-Heng Huang et.al. | 2609.33014v1 | null |
| 2026-09-26 | Multimodal LLMs Outperform Pathology Foundation Models in Cross-Domain Histological Similarity | Yishu Zhang et.al. | 2609.32876v2 | null |
| 2026-09-26 | Counterfactual Self-Evolving Agents for Evidence-Grounded Reasoning | Xing Han et.al. | 2609.32870v1 | null |
| 2026-09-26 | FinancialAuditBench: Benchmark Construction under Differential Privacy Using Real-World Priors | Jerry Huang et.al. | 2609.32835v1 | null |
Abstracts
A Comparative Explainability Framework for DeBERTa-v3 in Zero-Shot Medical Abstract Classification
2610.02116v1 by Javier Diaz Esteban-Herreros, David Muñoz-Valero, Raquel Martínez-España, Jose M. Juarez, Juan Moreno-Garcia
A comparative explainability framework is presented to audit DeBERTa-v3 under zero-shot classification of medical abstracts. The work addresses the disagreement problem in Explainable Artificial Intelligence, where different attribution methods produce divergent explanations for the same input and prediction. A natural language inference engine is implemented over the Medical Abstracts corpus with five enriched hypotheses per diagnostic category and a balanced sample of one thousand texts per class. Five explanation methods are compared: SHAP and LIME as model-agnostic approaches, occlusion and Input x Gradient as deep-learning-specific approaches, and Attention x Gradient as a transformer-specific approach. Explanations are standardized through top-token attribution, and pairwise agreement is quantified using the Jaccard index. High predictive accuracy is achieved across well-defined clinical domains, whereas performance degrades under high semantic ambiguity. Explanatory stability directly mirrors predictive certainty, exhibiting strong convergence in univalent categories and a marked drop under diagnostic uncertainty. Furthermore, qualitative error auditing uncovers three systemic failure mechanisms: lexical hypersensitivity, semantic overlap, and loss of attribution coherence. The results support the combined use of several explanation methods and quantitative agreement metrics when auditing transformer-based models in medical text classification, and suggest prioritizing specific clinical ontologies over broad diagnostic labels.
摘要:比較可解釋性框架被提出以審計 DeBERTa-v3 在醫學摘要的零樣本分類下。這項工作解決了可解釋人工智慧中的不一致問題,即不同的歸因方法對相同的輸入和預測產生不同的解釋。自然語言推理引擎在醫學摘要語料庫上實施,每個診斷類別有五個增強的假設,並且每個類別有一千篇文本的平衡樣本。比較了五種解釋方法:SHAP 和 LIME 作為模型無關的方法,遮蔽和輸入 x 梯度作為深度學習特定的方法,以及注意力 x 梯度作為Transformer特定的方法。通過頂部標記歸因標準化解釋,並使用 Jaccard 指數量化成對一致性。在明確定義的臨床領域中實現了高預測準確性,而在高語義模糊性下性能下降。解釋穩定性直接反映預測確定性,在單值類別中顯示出強烈的收斂,並在診斷不確定性下顯著下降。此外,定性錯誤審計揭示了三種系統性失效機制:詞彙過敏、語義重疊和歸因一致性的喪失。結果支持在醫學文本分類中審計基於Transformer的模型時,結合使用幾種解釋方法和定量一致性指標,並建議優先考慮特定的臨床本體論而非廣泛的診斷標籤。
Can AI Oversight Be Zero Knowledge?
2610.01995v1 by Alessandro Chiesa, Ziyi Guan, Burcu Yildiz
AI systems increasingly produce outputs from confidential data, such as a fitness-for-duty assessment from medical records or the predicted properties of a drug candidate from its secret structure. It is important to verify that such outputs are correct without revealing the underlying data. A recent line of work studies verification of AI outputs via interactive proofs and debate for oracle-aided computation, where correctness may depend on an oracle such as human judgment, a physical experiment, or the web. These works focus on verification by a verifier that runs much faster than the computation. However, such efficient verification is impossible for general oracle-aided computation, and these works therefore rely on additional assumptions. We focus instead on privacy: allowing the verifier to run in time polynomial in the computation, we ask whether interactive arguments for oracle-aided computation can be zero knowledge, so that the verifier learns nothing about the confidential data beyond the correctness of the output. We prove that, in general, they cannot. In the random oracle model, there are no zero-knowledge proofs for all oracle-aided computations, even if both the prover and the verifier are allowed to run much longer than the computation itself. The impossibility extends to debate, a canonical model for scalable oversight. On the positive side, we show that if the oracle attaches a cryptographic signature to each of its answers, then every oracle-aided computation can be verified in zero knowledge with an efficient prover and verifier, assuming only collision-resistant hash functions. Beyond privacy, this also gives an alternative approach to scalable oversight that relies neither on an honest opponent, as in debate, nor on the robustness of the computation, as in prior single-prover protocols.
摘要:AI 系統越來越多地從機密數據中產生輸出,例如從醫療記錄中進行的適任性評估或從其秘密結構中預測的藥物候選物的性質。
驗證這些輸出是否正確而不透露底層數據是很重要的。
最近的一系列研究通過互動證明和辯論研究 AI 輸出的驗證,用於有 oracle 輔助的計算,其中正確性可能依賴於 oracle,例如人類判斷、物理實驗或網絡。
這些研究專注於由運行速度遠快於計算的驗證者進行的驗證。
然而,對於一般的有 oracle 輔助計算,這樣的高效驗證是不可能的,因此這些研究依賴於額外的假設。
我們則專注於隱私:允許驗證者在計算的多項式時間內運行,我們詢問有 oracle 輔助計算的互動論證是否可以是零知識的,以便驗證者不會學到關於機密數據的任何信息,除了輸出的正確性。
我們證明,通常情況下,它們是不可能的。
在隨機 oracle 模型中,對於所有有 oracle 輔助的計算,沒有零知識證明,即使證明者和驗證者都被允許運行的時間遠超過計算本身。
這種不可能性擴展到辯論,這是一個可擴展監督的典型模型。
從積極的一面來看,我們展示了如果 oracle 為其每個答案附加加密簽名,那麼每個有 oracle 輔助的計算都可以在零知識中進行驗證,並且有高效的證明者和驗證者,只假設碰撞抗性哈希函數。
除了隱私之外,這還提供了一種可擴展監督的替代方法,既不依賴於誠實的對手(如辯論中),也不依賴於計算的穩健性(如以前的單證明者協議中)。
Counterfactual Auditing of Bias in Open-Source Large Language Models for Clinical Triage
2610.01963v1 by Manar Aljohani, Brandon Ho, Kenneth McKinley, Dennis Ren, Xuan Wang
Emergency department (ED) triage is a high-stakes prioritization task in which demographic, socioeconomic, and system-context information may improperly influence acuity assignment. Although open-source large language models (LLMs) are increasingly considered for local and privacy-preserving clinical decision support, it remains unclear how counterfactual bias varies across model families, sizes, medical-domain models, and domain-adapted models. We present a comparative counterfactual audit of ten open-source LLMs for pediatric Emergency Severity Index (ESI) prediction. Starting from real and handbook-style clinical vignettes, we construct paired counterfactual variants that change only one injected demographic, socioeconomic, healthcare-access, behavioral, social, or system-context variable while holding the clinical presentation fixed. Models include Qwen2.5-7B, Qwen2.5-14B-Instruct, a QLoRA fine-tuned Qwen2.5-7B, MedGemma variants, MedLLaMA2-7B, GPT-OSS-20B, and GPT-OSS-120B. We measure any counterfactual shift, undertriage, overtriage, shifts greater than one ESI level, mean shift, and mean absolute shift. Counterfactual sensitivity varied substantially and did not consistently decrease with larger model size or medical-domain pretraining. The fine-tuned Qwen2.5-7B showed the lowest overall sensitivity, with a 5.27% any-shift rate and mean absolute shift of 0.0534, versus 16.02% and 0.1706 for the base model. Several larger or medical-domain models showed more significant shifts. Stratified and correlation analyses further revealed clinically important directionality and shared failure patterns hidden by aggregate rates. These findings support counterfactual auditing as a lightweight, clinically interpretable framework for comparing fairness risks in open-source LLMs before clinical deployment.
摘要:急診部(ED)分診是一項高風險的優先排序任務,其中人口統計、社會經濟和系統背景信息可能不當影響急性程度的分配。儘管開源大型語言模型(LLMs)越來越被考慮用於本地和隱私保護的臨床決策支持,但目前尚不清楚反事實偏見在不同模型家族、大小、醫療領域模型和領域適應模型之間的變化情況。我們對十個開源LLM進行了針對兒科緊急嚴重性指數(ESI)預測的比較反事實審計。從真實和手冊風格的臨床小插曲開始,我們構建了配對的反事實變體,僅改變一個注入的人口統計、社會經濟、醫療訪問、行為、社會或系統背景變量,同時保持臨床表現不變。模型包括Qwen2.5-7B、Qwen2.5-14B-Instruct、經過QLoRA微調的Qwen2.5-7B、MedGemma變體、MedLLaMA2-7B、GPT-OSS-20B和GPT-OSS-120B。我們測量任何反事實變化、低估分診、過度分診、超過一個ESI級別的變化、平均變化和平均絕對變化。反事實敏感性變化顯著,且不一致地隨著模型大小或醫療領域預訓練的增大而減少。經過微調的Qwen2.5-7B顯示出最低的整體敏感性,任何變化率為5.27%,平均絕對變化為0.0534,而基礎模型則為16.02%和0.1706。幾個較大或醫療領域模型顯示出更顯著的變化。分層和相關分析進一步揭示了臨床上重要的方向性和由聚合率隱藏的共同失敗模式。這些發現支持反事實審計作為一種輕量級、臨床可解釋的框架,用於在臨床部署前比較開源LLM中的公平風險。
A rubric landscape for evaluating clinical reasoning in large language models: what exists, what is missing, and what needs to be combined
2610.01938v1 by Zhangshu Joshua Jiang, Zina Ibrahim, James T. Teo
Exam-style accuracy does not establish whether large language models (LLMs) reason well over clinical records. We define clinical reasoning as integrating and updating evidence across time and sources to form, revise and justify a patient's problem representation and a defensible plan. This structured narrative review maps three literatures: medical education assessment instruments, clinical LLM benchmarks published from 2023 onwards, and general-domain methods for evaluating long-form generation. We examine six dimensions: problem representation, temporal synthesis, differential and management reasoning, counterfactual reasoning, calibrated uncertainty, and reasoning faithfulness. Preprints are included and flagged. No single instrument covers all six dimensions. Problem representation and differential or management reasoning are reasonably covered, although reliability varies by instrument and setting. TIMER-Eval targets temporal synthesis, and ER-Reason assesses sequential diagnostic belief updating. Dedicated uncertainty and counterfactual evaluations are emerging, but their applicability to longitudinal free-text reasoning remains limited. Factual completeness is well theorised in general-domain evaluation, with early clinical evidence of important omissions. Faithfulness remains the weakest dimension, with one identified clinical causal-ablation study on multiple-choice questions. Existing tools should be combined through binary rubric items, separate completeness and correctness scores, case-specific importance weighting with non-compensable safety caps, temporal order-consistency checks, and chance-corrected reliability reporting. Further design work is needed for calibrated uncertainty, counterfactual reasoning and faithfulness over longitudinal free-text records. This review provides a design rationale, not a validated instrument.
摘要:考試風格的準確性並不能確定大型語言模型(LLMs)在臨床記錄上是否能夠進行良好的推理。我們將臨床推理定義為整合和更新跨時間和來源的證據,以形成、修訂和辯護病人的問題表述及可辯護的計劃。
這篇結構化的敘述性回顧映射了三個文獻領域:醫學教育評估工具、2023年以來發表的臨床LLM基準,以及評估長篇生成的一般領域方法。我們檢視了六個維度:問題表述、時間綜合、差異和管理推理、反事實推理、校準的不確定性,以及推理的忠實性。預印本已被納入並標記。
沒有單一的工具涵蓋所有六個維度。問題表述以及差異或管理推理的覆蓋相對合理,儘管可靠性因工具和環境而異。TIMER-Eval 針對時間綜合,而 ER-Reason 評估連續的診斷信念更新。專門的不確定性和反事實評估正在出現,但它們對於縱向自由文本推理的適用性仍然有限。事實的完整性在一般領域評估中有良好的理論基礎,並且早期臨床證據顯示出重要的遺漏。忠實性仍然是最薄弱的維度,其中有一項針對多選題的臨床因果消融研究被識別。
現有工具應通過二元評分項目、分開的完整性和正確性分數、特定案例的重要性加權(帶有不可補償的安全上限)、時間順序一致性檢查以及機會修正的可靠性報告進行結合。對於校準的不確定性、反事實推理和縱向自由文本記錄的忠實性,還需要進一步的設計工作。這篇回顧提供了一個設計的理由,而不是一個經過驗證的工具。
Walking the Embedding Space: Datastore Extraction from Multimodal RAG
2610.01871v1 by Maria Carmen Jica, Ali Satvaty, Suzan Verberne, Fatih Turkmen
Multimodal Retrieval-Augmented Generation (MRAG) has emerged as a reliable and cost-effective technique of grounding the generative capabilities of Multimodal Large Language Models (MLLMs) into relevant, up-to-date, external knowledge. Despite presenting several benefits, such as reducing hallucinatory behavior, they also introduce new attack surfaces, including leakage of private information and vulnerabilities against data extraction attacks. In this paper, we introduce $\immrag$, an adaptive and automatic data extraction attack procedure operating in a black box setting against \emph{image-returning} MRAG, a configuration in which the retrieved visual artifact is itself the response. Each query blends an attacker-held shadow image with an image already recovered from the system, and relevance-weighted resampling steers subsequent queries towards regions of the embedding space that still yield novel retrievals. Unlike current extraction attacks that aim to persuade the model towards data leakage by placing a malicious query as a textual prompt, $\immrag$ embeds the malicious instructions inside a user-given input image. We evaluate $\immrag$ on three plausible and distinct real-world scenarios: medical assistant, document-focused helper and general purpose tool. The experiments involve the study of the effectiveness of the attack on multiple CLIP-family retrievers, as well as the impact of various generators. A single 2500-query run reconstructs up to 611 distinct radiology images, 566 document scans and 416 general-purpose images under local-feature correspondence, and reaches up to $5.6\times$ as many distinct datastore items as a non-adaptive baseline. Our results show the urgent need for safeguards specifically designed for multimodal data.
摘要:多模態檢索增強生成(MRAG)已成為將多模態大型語言模型(MLLMs)的生成能力與相關的、最新的外部知識相結合的一種可靠且具成本效益的技術。儘管它提供了幾個好處,例如減少幻覺行為,但它們也引入了新的攻擊面,包括私密信息洩漏和對數據提取攻擊的脆弱性。
在本文中,我們介紹了 $\immrag$,這是一種適應性和自動化的數據提取攻擊程序,針對 \emph{圖像返回} MRAG 在黑箱環境中運作,這是一種檢索的視覺工件本身就是回應的配置。每個查詢將攻擊者持有的影像與系統中已恢復的影像混合,並且相關性加權重採樣引導後續查詢朝向仍能產生新穎檢索的嵌入空間區域。與目前旨在通過將惡意查詢作為文本提示來說服模型進行數據洩漏的提取攻擊不同,$\immrag$ 將惡意指令嵌入用戶提供的輸入影像中。我們在三個合理且不同的現實場景中評估了 $\immrag$:醫療助手、文件專注助手和通用工具。實驗涉及對多個 CLIP 家族檢索器的攻擊有效性以及各種生成器的影響進行研究。一次 2500 次查詢的運行重建了多達 611 幅不同的放射學影像、566 幅文件掃描和 416 幅通用影像,根據局部特徵對應,並達到高達 $5.6\times$ 的不同數據庫項目數量,相較於非適應性基準。我們的結果顯示出對專門為多模態數據設計的安全措施的迫切需求。
On the Divergence of Accuracy and Mechanism Consistency in Time Series World Models
2610.01842v1 by Haochen Zhang, Jiaheng Guo, Zhen Xu, Zachary Plotkin, Nicholas Konz, Zhen Tan, Tianlong Chen
A time series world model (TSWM) predicts a controlled system's state from its observed history and planned actions and exogenous inputs. Current approaches build forecasters with actions as covariates, trained and evaluated on prediction error under the executed plan. Yet world models compare unexecuted plans, but their responses to changed plans remain untested. We ask which design choices matter and whether accurate forecasters respond to changed plans as real systems do. We address both with a formalization and benchmark. The formalization separates state, actions and exogenous inputs, distinguishes continuous, mode and event actions, and introduces mechanism consistency, a metric built on declared action-state relations with known directions, such as a vasopressor raising blood pressure: it checks whether shifting an action moves the forecast in the declared direction. The benchmark consolidates eight public datasets with real actions from engineered infrastructure and clinical care, varying prediction space, plan fusion and plan encoding across seven backbones and five seeds. First, a frozen latent prediction space lowers MAE by 9.9% over observation space and gated output fusion lowers it by 12.7% over input concatenation on average, with both improving all eight datasets; temporal plan encoding changes average MAE by at most 2.2%. Second, prediction error and mechanism consistency diverge: the lowest-error configuration is at or below chance in consistency on four of five datasets with declared mechanisms, and no design choice avoids this. Finally, directional supervision, a loss penalizing the wrong-signed part of the response to a shifted action, significantly raises consistency on penalized mechanisms with no change in MAE. Together they give TSWMs a recipe: a frozen latent space and output-side fusion for accuracy, and a training objective for mechanism consistency.
摘要:時間序列世界模型 (TSWM) 從其觀察歷史、計畫行動和外部輸入預測受控系統的狀態。當前的方法建立了以行動為協變數的預測器,這些預測器在執行計畫下的預測誤差上進行訓練和評估。然而,世界模型比較未執行的計畫,但對於變更計畫的反應仍未經測試。我們詢問哪些設計選擇是重要的,以及準確的預測器是否像真實系統一樣對變更計畫做出反應。我們通過形式化和基準來解決這兩個問題。形式化將狀態、行動和外部輸入分開,區分連續、模式和事件行動,並引入機制一致性,這是一種基於已知方向的聲明行動-狀態關係構建的指標,例如一種升壓藥提高血壓:它檢查改變行動是否將預測移動到聲明的方向。基準整合了八個來自工程基礎設施和臨床護理的公共數據集,這些數據集中有真實行動,並在七個骨幹和五個種子中變化預測空間、計畫融合和計畫編碼。首先,凍結的潛在預測空間使 MAE 在觀察空間上降低了 9.9%,而門控輸出融合使其在輸入串接上平均降低了 12.7%,兩者都改善了所有八個數據集;時間計畫編碼的變化使平均 MAE 變化最多為 2.2%。其次,預測誤差和機制一致性出現分歧:在五個具有聲明機制的數據集中,最低誤差配置在一致性上與隨機相同或更低,且沒有任何設計選擇能避免這一點。最後,方向性監督,一種對於對移動行動的反應中錯誤符號部分進行懲罰的損失,顯著提高了懲罰機制的一致性,且 MAE 沒有變化。這些共同為 TSWM 提供了一個配方:凍結的潛在空間和輸出側融合以提高準確性,以及一個針對機制一致性的訓練目標。
iADD: Improving Alignment and Diversity in Diffusion Policy Optimization
2610.01789v1 by Ashok Prasad Neupane, Saugat Adhikari, Pramish Paudel, Ajad Chhatkuli, Danda Pani Paudel
Reinforcement learning based post training of diffusion models, such as Denoising Diffusion Policy Optimization (DDPO), optimizes a reverse diffusion process under a reward function. However, current approaches to reward optimizations do so at the cost of diversity and quality. In this paper, we provide better tradeoffs through careful theoretical considerations and method design. We analyze the theoretical framework and mathematically demonstrate that \emph{only-latter timestep} updates of diffusion model may be harmful for diversity contrary to the conclusions presented in a previous work. Additionally, we propose an incremental Feynman-Kac training based on strong theoretical foundations in order to achieve the best-yet alignment-diversity tradeoffs. We perform extensive experiments and compare our method against related diffusion policy optimization approaches in three different tasks and also provide strong ablations for each component, thus validating strong performance gains in both alignment and diversity.
摘要:基於強化學習的擴散模型後訓練,例如去噪擴散策略優化(DDPO),在獎勵函數下優化反向擴散過程。
然而,當前的獎勵優化方法以多樣性和質量為代價。
在本文中,我們通過仔細的理論考量和方法設計提供了更好的權衡。
我們分析了理論框架,並數學上證明擴散模型的\emph{僅後時間步}更新可能對多樣性有害,這與之前工作的結論相反。
此外,我們提出了一種基於強大理論基礎的增量費曼-卡克訓練,以實現最佳的對齊-多樣性權衡。
我們進行了廣泛的實驗,並在三個不同的任務中將我們的方法與相關的擴散政策優化方法進行比較,還為每個組件提供了強有力的消融實驗,從而驗證了在對齊和多樣性方面的顯著性能提升。
OpenMTB-Audit: Exposing Over-Refusal and Clinical Expert Perspectives in LLM-Based Molecular Tumor Board Safety Evaluation
2610.01497v1 by Negin Ashrafi, Jia Luo, Stacey M. Frumm, Roxana Daneshjou
Molecular tumor boards integrate genomic findings, clinical context, and therapeutic evidence to support precision oncology. As AI enters this workflow, a key safety challenge is distinguishing truly unsupported recommendations from evidence-supported options that still require oncologist review because of incomplete information, poor ECOG performance status, or other clinical caveats. We introduce OpenMTB-Audit, an open-source benchmark of 500 synthetic non-small cell lung cancer cases spanning five adversarial error categories and four safety labels: Supported, Partially Supported, Unsupported, and Insufficient Information. Across eight large language model configurations, we identify pervasive over-refusal: all LLM configurations failed to retain the Partially Supported label in 83.3-100% of true Partially Supported cases, achieving high aggregate safety scores through label collapse rather than clinically calibrated reasoning. To address this limitation, we developed MTB-AuditAgent, a deterministic seven-module framework separating evidence verification, missing-information detection, safety classification, and abstention. It reduces over-refusal to 6.7% and achieves 91.2% accuracy (95% CI: 88.6-93.6%). A two-oncologist annotation study found disagreement concentrated at the boundary between information sufficiency and treatment optimization, underscoring the need to preserve clinically meaningful distinctions.
摘要:分子腫瘤委員會整合基因組發現、臨床背景和治療證據,以支持精準腫瘤學。隨著人工智慧進入這一工作流程,一個主要的安全挑戰是區分真正不被支持的建議與仍需腫瘤醫生審查的證據支持選項,因為信息不完整、ECOG表現狀態不佳或其他臨床警告。我們介紹了OpenMTB-Audit,一個開源基準,包含500個合成的非小細胞肺癌案例,涵蓋五個對抗性錯誤類別和四個安全標籤:支持、部分支持、不支持和信息不足。在八種大型語言模型配置中,我們發現普遍的過度拒絕:所有LLM配置在83.3-100%的真實部分支持案例中未能保留部分支持標籤,通過標籤崩潰而非臨床校準推理獲得高整體安全分數。為了解決這一限制,我們開發了MTB-AuditAgent,一個確定性的七模塊框架,將證據驗證、缺失信息檢測、安全分類和放棄分開。它將過度拒絕降低到6.7%,並達到91.2%的準確率(95% CI:88.6-93.6%)。一項由兩位腫瘤醫生進行的標註研究發現,分歧集中在信息充分性和治療優化之間的邊界,強調了保留臨床上有意義的區別的必要性。
A Multi-Agent LLM Framework for Personalized Health Checkup Interpretation and Guidance
2610.01451v1 by HyungJun Kim, Taehan Lee, Soojin Cheon
Personalized interpretation of health checkup results requires reasoning across longitudinal records, medical knowledge, lifestyle guidance, and healthcare navigation. We present a multi-agent large language model (LLM) system that identifies multiple intents, maps each to a task-specific agent, executes them in parallel, and synthesizes their outputs. We compared answers generated in Single Agent and Multi Agent settings on 120 Korean compound queries combining two to four requirements, using synthetic health checkup records. The Multi Agent improved the weighted LLM-judge score from 1.695 to 1.797 (p = 0.027), and three additional LLM judges showed consistent improvements ($Δ$ = +0.111 to +0.186, all p < 0.05). The gains came from usefulness, consistency, and the handling of every requirement in compound queries, whereas numerical accuracy and grounding improved significantly under only one of the four judges and medical safety did not differ, and critical failures occurred at similar rates (Single Agent 15.0% vs. Multi Agent 13.3%). Two human evaluators preferred Multi Agent in 66.7% and 68.3% of pairwise comparisons. Multi Agent execution increased latency and cost by 1.31$\times$ and 2.02$\times$, respectively. In exploratory subgroup analyses, the improvement was concentrated in queries involving personal-record lookup.
摘要:個性化的健康檢查結果解釋需要跨越長期記錄、醫學知識、生活方式指導和醫療導航的推理。
我們提出了一個多代理大型語言模型(LLM)系統,該系統識別多個意圖,將每個意圖映射到特定任務的代理,並平行執行它們,最後綜合其輸出。
我們比較了在單代理和多代理設置下,使用合成健康檢查記錄對120個韓國複合查詢生成的答案,這些查詢結合了兩到四個需求。
多代理將加權LLM評審分數從1.695提高到1.797(p = 0.027),另外三位LLM評審顯示出一致的改善($Δ$ = +0.111到+0.186,所有p < 0.05)。
這些增益來自於有用性、一致性以及對複合查詢中每個需求的處理,而數值準確性和基礎資料僅在四位評審中的一位顯著改善,醫療安全則沒有差異,且重大失誤的發生率相似(單代理15.0%對多代理13.3%)。
兩位人類評估者在66.7%和68.3%的成對比較中偏好多代理。
多代理執行使延遲和成本分別增加了1.31$\times$和2.02$\times$。
在探索性子群分析中,改善集中在涉及個人記錄查詢的問題上。
Generation Provenance Before Behavior Attribution: Auditing Synthetic Speech Research Objects
2610.01378v1 by Sidi Chang, Peiying Zhu
Attributing model behavior to synthetic training data requires knowing what produced each training item before estimating what that item caused. A waveform-label pair does not preserve this knowledge. We propose a generation-provenance substrate in which a synthetic research object binds source specification, generated content, waveform, target, fact requirements, quality signals, review lineage, and immutable manifest identity. Producer and selection mechanism determine evidentiary meaning; storage location and variable name do not. We audit this substrate in a private Japanese care-handoff pipeline. A 113-asset review population contains 1.552 hours of synthetic speech across six scenario families; all items have linked audio, transcripts, candidate notes, and fact checklists, but human evidence is selective and source-specific. Two faithful-only manifests are scenario-seed-disjoint and immutably versioned, while exact upstream attribution remains blocked by floating generator aliases, missing per-clip TTS and code stamps, and an unversioned checking prompt. We argue that generation provenance is necessary but not sufficient for behavior attribution: it defines the candidate causal graph and audit units, whereas contributive attribution still requires frozen training runs and intervention or influence evidence. The paper contributes a compact provenance contract, an audit protocol, and a bounded case study for synthetic-data attribution; controlled research access may be offered, but we do not claim causal training-data attribution, clinical validity, or unrestricted public release.
摘要:將模型行為歸因於合成訓練數據需要了解每個訓練項目是如何產生的,然後才能估計該項目造成了什麼。波形-標籤對並不保留這種知識。我們提出了一種生成來源基底,其中合成研究對象綁定了源規範、生成內容、波形、目標、事實要求、質量信號、審查血統和不可變的清單身份。生產者和選擇機制決定了證據意義;存儲位置和變量名稱則不然。我們在一個私有的日本護理交接管道中審計這一基底。一個包含113個資產的審查群體包含了六個情境系列中的1.552小時合成語音;所有項目都鏈接了音頻、轉錄、候選筆記和事實檢查清單,但人類證據是選擇性的且特定於來源。兩個僅限忠實的清單在情境種子上是不相交且不可變版本的,而精確的上游歸因仍然受到浮動生成器別名、缺失的每段TTS和代碼印記以及未版本化的檢查提示的阻礙。我們主張生成來源對於行為歸因是必要但不充分的:它定義了候選因果圖和審計單位,而貢獻性歸因仍然需要凍結的訓練運行和干預或影響證據。本文貢獻了一個簡潔的來源合約、一個審計協議和一個有界的合成數據歸因案例研究;可能會提供受控的研究訪問,但我們不聲稱因果訓練數據歸因、臨床有效性或不受限制的公開發布。
An ontology for cross-sectoral crisis management: core and public health modules
2610.01326v1 by Aldo Gangemi, Rita T. Sousa, Luigi Asprino, Giorgia Lodi, Andrea G. Nuzzolese, Valentina Presutti, Johannes Gysen, Diana F. Sousa, Luigi Spagnolo
This paper presents the European Crisis Management Ontology (ECMO), a modular OWL-based ontology intended as a cross-sectoral reference for disaster risk reduction and response. ECMO is designed to be organised as a network of ontological modules. Among the modules, ECMO-CORE captures fundamental crisis management concepts such as hazard, event, exposure, impact, and response measure and uses ontology design patterns and the OWL2 punning technique to resolve ambiguities between hazard types and event manifestations. In addition, domain-specific modules are defined as in the case of the public health module aligned with SNOMED CT and ICD-11. To demonstrate the resource's utility, we used ECMO to represent the data of the Epidemic Intelligence from Open Sources system of the Joint Research Centre to generate an end-to-end pipeline that populates an ECMO-compliant knowledge graph from unstructured epidemiological news. Initial results demonstrate that ECMO provides the formal guardrails necessary for consistent and unified knowledge representation and integration. The ontology is publicly available at https://doi.org/10.5281/zenodo.20070268 and is released under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.
摘要:這篇論文介紹了歐洲危機管理本體(ECMO),這是一個基於OWL的模組化本體,旨在作為災害風險減少和應對的跨領域參考。ECMO的設計是作為本體模組的網絡組織。 在這些模組中,ECMO-CORE捕捉了基本的危機管理概念,如危險、事件、暴露、影響和應對措施,並使用本體設計模式和OWL2的雙義技術來解決危險類型和事件表現之間的歧義。此外,還定義了特定領域的模組,例如與SNOMED CT和ICD-11對齊的公共衛生模組。為了展示該資源的實用性,我們使用ECMO來表示聯合研究中心的開放來源流行病情報系統的數據,以生成一個從非結構化流行病學新聞填充ECMO合規知識圖譜的端到端管道。初步結果顯示,ECMO提供了必要的正式框架,以實現一致和統一的知識表示和整合。該本體可在https://doi.org/10.5281/zenodo.20070268上公開獲得,並根據創用CC 4.0國際版(CC BY 4.0)授權發布。
Model validation in machine learning: A scenario-based guide from hold-out splits to nested group cross-validation in biomedical and applied research
2610.01284v1 by Mehmet Baygin, Sengul Dogan, Turker Tuncer
Model validation estimates the performance of a complete learning procedure on new data. However, an invalid split can produce an optimistic and stable result. This tutorial reviews hold-out validation, train/validation/test designs, repeated random subsampling, k-fold and repeated stratified cross-validation, leave-one-out and leave-p-out schemes, group-aware validation, and nested group cross-validation. General machine-learning principles are linked to EEG epochs, paired-eye OCT images, repeated clinical measurements, and multicenter data. Eight controlled scenarios compare flawed and leakage-safe designs: seven use locked confusion matrices with auditable metrics, and one uses a reproducible repeated-study simulation. The scenarios cover global feature selection, normalization leakage, dependent records, center mixing, repeated test-set use, and estimator instability. Bias, variance, metric aggregation, uncertainty, and computational cost are also examined. A data-size matrix, a decision tree, and reporting checklists are provided. Reproducible MATLAB templates and scikit-learn counterparts are included. The results show that no validation method is universally best. The independent unit must match the intended deployment target. Every data-dependent operation must also exclude the observations used for performance estimation.
摘要:模型驗證評估完整學習程序在新數據上的表現。然後,無效的分割可能會產生樂觀且穩定的結果。本教程回顧了保留驗證、訓練/驗證/測試設計、重複隨機子抽樣、k-折和重複分層交叉驗證、留一法和留p法方案、群體感知驗證以及嵌套群體交叉驗證。一般的機器學習原則與EEG時期、配對眼睛OCT影像、重複臨床測量和多中心數據相關聯。八個控制場景比較了有缺陷和防洩漏的設計:七個使用鎖定的混淆矩陣和可審計的指標,一個使用可重複的重複研究模擬。這些場景涵蓋了全局特徵選擇、正規化洩漏、依賴記錄、中心混合、重複測試集使用和估計器不穩定性。還檢查了偏差、方差、指標聚合、不確定性和計算成本。提供了數據大小矩陣、決策樹和報告檢查清單。包括可重複的MATLAB模板和scikit-learn對應物。結果顯示,沒有一種驗證方法是普遍最佳的。獨立單元必須與預期的部署目標相匹配。每個依賴數據的操作也必須排除用於性能估計的觀察值。
When Does Exercise-Specific Joint Selection Help? An Audit of Evaluation and Control Design
2610.01188v1 by Haotian Chen, Jingkun Yu, Yuning Zhang, Bowen Ye
Exercise-specific joint selection can improve skeleton-based correctness classification, but what does that gain establish? We audit 1,057 repetitions from ten REHAB24-6 subjects, separating evaluation aggregation, subset structure, and temporal representation. The manual-subset kNN gain changes from 0.055 for pooled out-of-fold AUROC to 0.020 for equal-weight within-person AUROC; both paired intervals include zero. Among 1,000 dimension-matched random maps, 14 match or exceed the manual pooled result, versus 145 when bilateral structure and trunk inclusion are also matched. RBF-SVM retains a positive within-person gain, whereas logistic regression and a random-convolution comparator have negative point gains under that estimand. Sequence-order and paired-seed controls further qualify the interpretation. This exploratory audit shows why joint-selection claims require explicit estimands and structurally appropriate controls; it does not establish a new algorithm or clinical benefit.
摘要:運動特定的關節選擇可以改善基於骨架的正確性分類,但這樣的增益究竟建立了什麼?我們審核了來自十名 REHAB24-6 受試者的 1,057 次重複,分離評估聚合、子集結構和時間表示。手動子集 kNN 的增益從 0.055 變化到 0.020,對應於合併的折外 AUROC 和等權重的個體內 AUROC;這兩個配對區間都包含零。在 1,000 個維度匹配的隨機映射中,有 14 個匹配或超過手動合併結果,而當雙邊結構和軀幹包含也匹配時,則有 145 個。RBF-SVM 保持了正的個體內增益,而邏輯回歸和隨機卷積比較器在該估計下則有負的點增益。序列順序和配對種子控制進一步限定了解釋。這項探索性審核顯示為什麼關節選擇的主張需要明確的估計量和結構上適當的控制;它並未建立新的算法或臨床利益。
CineMR: Tool-Integrated Vision-Language Reasoning for Quantitative Cardiac MRI Assessment
2610.01166v1 by Kunyang Li, Hai Nguyen, Joshua Lowe, Chenguang Zhao, Peace C. Madueme, Mehdi Hedjazi Moghari, Mubarak Shah, Pegah Khosravi, Yuzhang Zhang
Cardiovascular magnetic resonance (CMR), including cine imaging, is a reference standard for the noninvasive assessment of cardiac morphology and ventricular function. Cine CMR interpretation integrates qualitative visual assessment with quantitative measurements of ventricular volumes, ejection fraction, myocardial mass, wall thickness, and regional wall motion. Current medical vision-language models (VLMs) cannot reliably derive quantitative measurements from multidimensional cine images without analysis tools. We present CineMR, a tool-augmented VLM that invokes cardiac image-analysis tools and integrates their outputs into interleaved reasoning for quantitative CMR assessment. We also construct a multi-cohort visual question answering benchmark covering quantitative metric extraction, multiclass diagnosis, and differential diagnosis, together with tools for segmentation, phase selection, volumetry, morphometry, and regional wall motion analysis. CineMR is trained with supervised fine-tuning (SFT) on tool-interaction traces followed by Group Relative Policy Optimization (GRPO) with conditional tool-use rewards. On the multi-cohort cine CMR benchmark, CineMR achieves 35.9% pass@1 and 58.9% pass@4, compared with 1.5% pass@1 for the Qwen3-VL-8B backbone and 0.0% and 7.0% pass@1 for LLaVA-Med v1.5 and MedGemma-4B, respectively. Correct tool invocation reaches 99.8% after GRPO, up from 78.9% after SFT. Live tool outputs improve ventricular measurement accuracy by 20.4--23.7% over direct model predictions, and removing all tools reduces pass@1 from 35.9% to 27.9%. These results highlight the importance of reliable tool use for quantitative cine CMR reasoning and support CineMR as a promising approach for assistive cardiac image assessment. Code, benchmark resources, and model weights are available at https://github.com/AI-MIND-Lab/CineMR.
摘要:心血管磁共振(CMR),包括動態影像,是非侵入性評估心臟形態和心室功能的參考標準。動態 CMR 解釋將定性視覺評估與心室體積、射血分數、心肌質量、壁厚和區域壁運動的定量測量相結合。目前的醫療視覺-語言模型(VLMs)在沒有分析工具的情況下,無法可靠地從多維動態影像中推導出定量測量。我們提出了 CineMR,一種增強工具的 VLM,調用心臟影像分析工具並將其輸出整合到交錯推理中,以進行定量 CMR 評估。我們還構建了一個涵蓋定量指標提取、多類別診斷和鑑別診斷的多隊列視覺問答基準,並提供分割、相位選擇、體積測量、形態測量和區域壁運動分析的工具。CineMR 在工具互動痕跡上進行了監督微調(SFT),隨後使用條件工具使用獎勵進行了群體相對策略優化(GRPO)。在多隊列動態 CMR 基準上,CineMR 的 pass@1 為 35.9%,pass@4 為 58.9%,而 Qwen3-VL-8B 的 pass@1 僅為 1.5%,LLaVA-Med v1.5 和 MedGemma-4B 的 pass@1 分別為 0.0% 和 7.0%。經過 GRPO 正確調用工具的比例達到 99.8%,而 SFT 後為 78.9%。實時工具輸出提高了心室測量的準確性,較直接模型預測提高了 20.4% 至 23.7%,而去除所有工具則使 pass@1 從 35.9% 降至 27.9%。這些結果突顯了可靠工具使用在定量動態 CMR 推理中的重要性,並支持 CineMR 作為輔助心臟影像評估的有前景方法。代碼、基準資源和模型權重可在 https://github.com/AI-MIND-Lab/CineMR 獲得。
A Matched-Budget Audit Framework for Recaptioned Image-Text Supervision Distributions
2610.00952v1 by Giyeong Oh, Junghun Park, Yuhan Bae, Youngjae Yu
Recaptioned image-text corpora are now standard for text-to-image (T2I) training, with vision--language model (VLM) captioners replacing sparse alt-text by dense descriptions. A recaptioned corpus is a supervision distribution induced by a documented captioning policy ($π$), captioner ($V_c$), and source corpus ($C$). Length-correlated proxies miss caption-register artifacts and downstream T2I benchmarks entangle the corpus with training choices, so this distribution is hard to audit at corpus scale. We introduce a reusable matched-budget audit framework for recaptioned supervision distributions $D_{π,V_c,C}$: at a fixed text budget of $B = 64$ it reports a five-axis profile spanning prompt-side coverage, image-conditioned faithfulness, and caption-surface health, with claimed controllable basic units (CBU) as the common claim unit. We instantiate the framework on seven paired comparisons over five public source corpora. Across the four cross-corpus pairs, the released surface raises supported CBU per caption by $+3.39$ to $+6.36$ under both Qwen and Gemma Judges, and on CC12M the same framework exposes a long-vs-dense frontier that is consistent across both judges and four budgets. We release the audited multi-source recap corpus ($\approx$ 490M) together with the audit-artifact bundle.
摘要:重新標題的影像-文本語料庫現在已成為文本到影像(T2I)訓練的標準,視覺-語言模型(VLM)標題生成器用密集的描述取代了稀疏的替代文本。重新標題的語料庫是由文件化的標題政策($π$)、標題生成器($V_c$)和來源語料庫($C$)所引發的監督分佈。與長度相關的代理錯過了標題註冊的工件,而下游的 T2I 基準則將語料庫與訓練選擇糾纏在一起,因此這種分佈在語料庫規模上很難進行審計。我們引入了一個可重用的匹配預算審計框架,用於重新標題的監督分佈 $D_{π,V_c,C}$:在固定的文本預算 $B = 64$ 下,它報告了一個涵蓋提示側覆蓋率、影像條件忠實度和標題表面健康的五軸概況,並以聲稱可控的基本單位(CBU)作為共同的聲明單位。我們在五個公共來源語料庫上進行了七個配對比較來實現該框架。在四對跨語料庫的比較中,釋放的表面在 Qwen 和 Gemma 評審下每個標題支持的 CBU 提高了 $+3.39$ 到 $+6.36$,而在 CC12M 上,同樣的框架揭示了一個長對密集的邊界,這在兩位評審和四個預算中都是一致的。我們釋放了經過審計的多來源重新標題語料庫($\approx$ 490M),以及審計工件包。
Backdoor Purification for LoRA-Tuned LLMs via Null-Space Projection
2610.00685v1 by Jianwei Li, Jung-Eun Kim
With the rapid adoption of large language models (LLMs) and parameter-efficient fine-tuning (PEFT) methods, the risk of backdoor attacks has become more severe. Existing backdoor purification methods typically rely on at least one of the strong assumptions, such as prior knowledge of triggers, access to clean references, or aggressive retraining, and they often lack comprehensive evaluations. These constraints substantially limit their practical applicability. To overcome these challenges, our work proposes purifying LoRA-tuned LLMs without these assumptions and even without post-hoc retraining of the suspect parameters. Our objective is to significantly reduce the attack success rates (ASR) while preserving both (i) the base model's general capabilities and (ii) the new downstream skills learned through the adapter. Through a series of ablation studies, we progressively scale our approach from a single layer in a text classification setting to a full-parameter LLM in the generative task. Through careful data curation and feature approximation, we extract high-fidelity backdoor directions and, for each layer or head, construct orthogonal null spaces in both the input and output channels, onto which the LoRA updates are projected. Empirically, our null-space projection method reduces the ASR from nearly 100% to less than 10%, while preserving the base model's benign performance and the adapter's learned abilities during downstream task adaptation.
摘要:隨著大型語言模型(LLMs)和參數高效微調(PEFT)方法的快速採用,後門攻擊的風險變得更加嚴重。現有的後門淨化方法通常依賴於至少一個強假設,例如對觸發器的先驗知識、訪問乾淨參考資料或激進的再訓練,並且它們往往缺乏全面的評估。這些限制大大限制了它們的實際應用性。為了克服這些挑戰,我們的工作提出了在沒有這些假設的情況下淨化LoRA調整的LLMs,甚至不需要對可疑參數進行事後再訓練。我們的目標是顯著降低攻擊成功率(ASR),同時保留(i)基礎模型的一般能力和(ii)通過適配器學到的新下游技能。通過一系列的消融研究,我們逐步將我們的方法從文本分類設定中的單層擴展到生成任務中的全參數LLM。通過仔細的數據策劃和特徵近似,我們提取高保真度的後門方向,並為每一層或頭構建正交的零空間,這些零空間位於輸入和輸出通道上,LoRA更新將被投影到這些空間中。經驗上,我們的零空間投影方法將ASR從近乎100%降低到不到10%,同時在下游任務適應過程中保留了基礎模型的良性性能和適配器學到的能力。
Worse Together: How Performance Breaks Down in Multi-User Multi-Agent Teams
2610.00583v1 by Sahan Paliskara, Nattaput Namchittai, Andrew Lampinen
People are increasingly delegating tasks to AI agents, and those agents are increasingly encountering other people's agents over shared resources such as a codebase, a calendar, or a budget. When each agent acts for a different user with different goals, coordination often fails, and the group ends up worse off than if a single agent had acted for everyone. We study this multi-user, multi-agent setting across five frontier models and 77 scenarios in four environments: an API key environment in which agents share a compute budget, a clinic in which they share a calendar, a personal assistant environment in which they share a group order or booking, and a merge queue in which they share a release cutoff. In each scenario, we compare a single agent that serves every user (a coordinator) to a team in which each agent serves one user, with and without a communication channel between the agents. Teams deliver worse group outcomes than the coordinator in every environment: without a channel, they completely collapse in two environments, and even with one, coordination overhead creates substantial gaps. For example, in the personal assistant environment, the coordinator fulfills a targeted user request about twice as often as teams. We identify distinct behaviors associated with this poor group-level performance, including stalling as teams grow, overriding each other's actions, and fabricating claims. We find effective but environment-specific mitigations, such as a team lead, explicit procedural instructions, and a platform check that makes an agent read its peers' messages before committing. We will release the API key, clinic, and personal assistant environments as MAMUBench, comprising 74 scenarios for evaluating multi-user, multi-agent coordination.
摘要:人們越來越多地將任務委派給 AI 代理,而這些代理也越來越多地與其他人的代理在共享資源上相遇,例如代碼庫、日曆或預算。當每個代理代表不同的用戶且目標不同時,協調往往失敗,結果小組的情況比由單一代理為所有人行動時更糟。我們研究了這種多用戶、多代理的設定,涵蓋五個前沿模型和四個環境中的 77 種情境:一個 API 金鑰環境,在這裡代理共享計算預算;一個診所,在這裡他們共享日曆;一個個人助理環境,在這裡他們共享團體訂單或預訂;以及一個合併隊列,在這裡他們共享發佈截止時間。在每個情境中,我們將為每個用戶服務的單一代理(協調者)與每個代理服務一位用戶的團隊進行比較,並考慮代理之間是否有通信渠道。團隊在每個環境中提供的群體結果都比協調者差:在沒有渠道的情況下,他們在兩個環境中完全崩潰,即使有一個,協調開銷也會造成相當大的差距。例如,在個人助理環境中,協調者滿足目標用戶請求的頻率約為團隊的兩倍。我們識別出與這種低群體表現相關的不同行為,包括隨著團隊增長而停滯、覆蓋彼此的行動以及捏造聲明。我們發現有效但特定於環境的緩解措施,例如團隊負責人、明確的程序指示,以及一個平台檢查,使代理在提交之前閱讀其同伴的消息。我們將發布 API 金鑰、診所和個人助理環境作為 MAMUBench,包含 74 種情境以評估多用戶、多代理的協調。
Can LLMs Reason Over Long Horizons? An Empirical Evaluation of Context Strategies for Longitudinal Clinical Reasoning
2610.00562v1 by Taye Akinrele, Noorbakhsh Amiri Golilarz, Subash Neupane, Sudip Mittal, Shahram Rahimi
Longitudinal clinical reasoning requires large language models (LLMs) to identify and integrate relevant evidence distributed across extended patient histories. Although long-context models can process increasingly large amounts of information, providing more history does not necessarily make relevant evidence more accessible or improve reasoning. We compare five context strategies (Full, Recent, Episodic, Semantic, and Hybrid) on MedLoCoMo across four open-weight LLMs, examining answer correctness, robustness to query-evidence distance, and abstention on questions with unsupported premises. Episodic and Hybrid generally achieve the strongest overall accuracy, while Recent Context degrades most as supporting evidence becomes more distant; Episodic and Hybrid maintain the highest accuracy at long distances. Analysis of adversarial questions further shows that strong performance on answerable questions does not necessarily translate to successful abstention when the available history does not support the requested conclusion. These findings show that reliable longitudinal reasoning depends not only on how much history an LLM can access, but critically on how relevant evidence is selected and presented for reasoning.
摘要:長期臨床推理需要大型語言模型(LLMs)識別並整合分佈在廣泛病歷中的相關證據。雖然長上下文模型可以處理越來越多的信息,但提供更多的歷史並不一定使相關證據更易於獲取或改善推理。我們在四個開放權重的LLM上比較五種上下文策略(完整、最近、情節、語義和混合)在MedLoCoMo上的表現,檢查答案的正確性、對查詢-證據距離的穩健性,以及對不支持前提的問題的棄權。情節和混合策略通常實現了最強的整體準確性,而最近上下文在支持證據變得更遙遠時下降最嚴重;情節和混合策略在長距離下保持最高的準確性。對對抗性問題的分析進一步顯示,對可回答問題的強勁表現並不一定轉化為在可用歷史不支持所要求結論時的成功棄權。這些發現表明,可靠的長期推理不僅依賴於LLM可以訪問多少歷史,還關鍵於如何選擇和呈現相關證據以進行推理。
Unlearnable, or Unmeasured? On the Reliability of Difficulty Labels in RLVR
2609.40115v1 by Chandak Chakma, Syed Nazmus Sakib, Nafiul Haque, Shifat E. Arman
Reinforcement learning with verifiable rewards (RLVR) has become an important approach for improving reasoning during post-training. Recent work suggests that some difficult prompts remain resistant to learning even when they occasionally produce correct solutions. We revisit this unlearnability phenomenon and find that the affected prompts do improve, at roughly one third of the learnable rate, while the difficulty-defined set used to study them is much less reproducible than expected. These difficulty labels are estimated from a limited number of sampled responses. Combining them across seeds can further change which prompts are selected instead of simply reducing measurement noise. We develop a sampling-based framework for quantifying this instability and determining how much evaluation is required for difficulty assignments to reproduce reliably. We also revisit the gradient-similarity evidence proposed to explain unlearnability and show that part of the observed separation arises because difficult prompts provide fewer correct rollouts from which their gradients can be estimated. Matching this sample count weakens the gradient difference but does not remove it. Overall, the slow-learning phenomenon survives our reanalysis, while both the prompts used to define it and the evidence used to explain it require more careful measurement.
摘要:強化學習與可驗證獎勵(RLVR)已成為改善後訓練推理的重要方法。最近的研究表明,即使某些困難的提示偶爾產生正確的解決方案,它們仍然對學習具有抵抗力。我們重新檢視這一不可學習現象,發現受影響的提示確實有所改善,改善速度約為可學習速率的三分之一,而用來研究它們的困難定義集的可重現性遠低於預期。這些困難標籤是從有限數量的樣本反應中估算得出的。跨種子結合它們可能進一步改變所選擇的提示,而不僅僅是減少測量噪音。我們開發了一個基於抽樣的框架來量化這種不穩定性,並確定為了使困難分配可靠地重現需要多少評估。我們還重新檢視了用於解釋不可學習的梯度相似性證據,並顯示觀察到的分離部分源於困難提示提供的正確回饋較少,從中無法估算其梯度。匹配這一樣本數量削弱了梯度差異,但並未消除它。總體而言,緩慢學習現象在我們的重新分析中仍然存在,而用來定義它的提示和用來解釋它的證據都需要更仔細的測量。
GateSPINE: Gated Cross-View Fusion for Lumbar Spine MRI Report Generation
2609.40091v1 by Hoang Nguyen Van, Cuong Vuong Tuan, Trang Mai Xuan, Bien Tran Van, Nam Tran Van, Thien Van Luong
Automated report generation can ease the burden radiolo gists face when interpreting multi-sequence MRI studies. Unlike CT, MRI examinations comprise multiple sequences and imaging planes, each con tributing complementary diagnostic information. Existing methods en code a study as a single volume and combine multiple acquisitions by fixed rules. Findings visible in only one plane are thus diluted and of ten missed, lowering recall on clinical efficacy metrics, where a missed abnormality is most costly. We propose GateSPINE, a vision-language framework that fuses sagittal T1 and T2 volumes with a training-free operator, encodes the fused sagittal and axial volumes with two parallel 3D encoders, and decodes their combined representation into a report. Its core mechanism is a gated cross view fusion module that predicts, per feature channel and token, how much of each view to admit, so the more informative view dominates at each spatial location. We evaluate GateSPINE on three lumbar MRI datasets, comprising two public bench marks and a private cohort collected from Phenikaa University Hospital, using both natural language generation (NLG) and clinical efficacy (CE) metrics. GateSPINE achieves the highest CE F1 through improved re call on all three datasets; on SPIDER, which lacks an axial sequence, this reflects the sagittal fusion component rather than the gated cross-view mechanism, which is validated on the two cohorts with both imaging planes. GateSPINE also remains competitive on standard NLG metrics.
摘要:自動報告生成可以減輕放射科醫生在解讀多序列MRI研究時所面臨的負擔。與CT不同,MRI檢查由多個序列和成像平面組成,每個平面提供互補的診斷信息。現有的方法將研究編碼為單一體積,並通過固定規則組合多個獲取結果。僅在一個平面上可見的發現因此被稀釋,並且經常被忽略,這降低了臨床效能指標的召回率,而漏掉的異常是最昂貴的。我們提出了GateSPINE,一個視覺-語言框架,融合了矢狀面T1和T2體積,使用無需訓練的運算符,並用兩個平行的3D編碼器編碼融合的矢狀面和軸向體積,然後將它們的組合表示解碼成報告。其核心機制是一個門控交叉視圖融合模塊,根據特徵通道和標記預測每個視圖應該接受多少,以便在每個空間位置上更具信息性的視圖占主導地位。我們在三個腰椎MRI數據集上評估GateSPINE,包括兩個公共基準和一個來自Phenikaa大學醫院的私有隊列,使用自然語言生成(NLG)和臨床效能(CE)指標。GateSPINE通過提高所有三個數據集的召回率,實現了最高的CE F1;在缺少軸向序列的SPIDER上,這反映了矢狀面融合組件,而不是門控交叉視圖機制,這在兩個具有成像平面的隊列中得到了驗證。GateSPINE在標準NLG指標上也保持競爭力。
Overview of BioASQ 2026: The fourteenth BioASQ Challenge on Large-Scale Biomedical Semantic Indexing and Question Answering
2609.39975v1 by Anastasios Nentidis, Georgios Katsimpras, Anastasia Krithara, Martin Krallinger, Miguel Rodríguez-Ortega, Eduard Rodriguez-López, Natalia Loukachevitch, Igor Rozhkov, Elena Tutubalina, Dimitris Dimitriadis, Vasiliki Patsiou, Grigorios Tsoumakas, George Giannakoulas, Alexandra Bekiaridou, Athanasios Samaras, Giorgio Maria Di Nunzio, Nicola Ferro, Stefano Marchesin, Marco Martinelli, Gianmaria Silvello, Georgios Paliouras
This paper presents an overview of the fourteenth edition of the BioASQ challenge, organized in the context of the Conference and Labs of the Evaluation Forum (CLEF) 2026. BioASQ is an international challenge series that supports progress in biomedical language processing tasks ranging from semantic indexing and information extraction to question answering and summarization. In 2026, BioASQ included six shared tasks: a) Task 14b on biomedical semantic question answering. b) Task Synergy14 on question answering for developing biomedical top- ics. c) Task MultiClinSum-2 on multilingual clinical summarization. d) Task BioNNE-R on extracting relations between nested named entities in Russian and English. e) Task ELCardioCC on clinical coding in cardiology. f) Task GutBrainIE on gut-brain interplay information extrac- tion. Across these six tasks, 87 distinct teams participated, submitting more than 1000 runs overall. As in previous editions, several submissions reached competitive performance, reflecting the continued progress of state-of-the-art methods across biomedical language processing tasks.
摘要:這篇論文概述了第十四屆BioASQ挑戰賽,該賽事在2026年評估論壇會議及實驗室(CLEF)的背景下舉辦。BioASQ是一系列國際挑戰,旨在支持生物醫學語言處理任務的進展,這些任務包括語義索引、信息提取、問題回答和摘要。在2026年,BioASQ包括六個共享任務:a) 任務14b,針對生物醫學語義問題回答。b) 任務Synergy14,針對發展生物醫學主題的問題回答。c) 任務MultiClinSum-2,針對多語言臨床摘要。d) 任務BioNNE-R,提取俄語和英語中嵌套命名實體之間的關係。e) 任務ELCardioCC,針對心臟病學的臨床編碼。f) 任務GutBrainIE,針對腸道與大腦相互作用的信息提取。在這六個任務中,共有87支不同的團隊參加,總共提交了超過1000次運行。與之前的版本一樣,幾個提交達到了競爭性的表現,反映了在生物醫學語言處理任務中最先進方法的持續進步。
Scores That Hold, Benchmarks That Leak: Measuring Dataset Contamination in Public Brain-Tumor MRI Classification
2610.00421v1 by Bhanu Prakash Vangala, Sowmya Guda, Latha Peddi, Navya Vangala
Automated classification of brain tumors from MRI is a heavily published application of deep learning in medical imaging, with reported accuracies on public benchmarks routinely exceeding 98%. However, accuracy does not capture a critical dimension of benchmark quality: dataset integrity, defined as the independence of test from training data at the image, patient, and acquisition-source levels. We introduce a three-layer contamination framework comprising duplicate, patient, and source-label leakage to assess the public corpora on which this literature rests. We audit the three most widely used corpora against a chest-radiograph negative control and quantify each layer's effect on measured performance across nine architectures and three evaluation conditions. Contamination is severe at every layer: 28.8% of the dominant corpus's official test split has a near-twin in its own training split, a second corpus leaks 22.3% of its test images byte-identically, 95.5% of traceable test images share a patient with training, and file-header features containing no anatomy separate tumor from no-tumor at 0.959 balanced accuracy, at parity with fine-tuned ResNet backbones. The unexpected result is that removing every identified leaked test image leaves balanced accuracy essentially unchanged: stable performance after deduplication does not establish benchmark integrity. Our findings establish dataset integrity as a distinct, measurable axis of benchmark quality that a stable leaderboard cannot certify. For biomedical research, reported accuracy on these corpora alone does not establish that a model has learned to recognize tumors rather than exploit dataset-specific cues. We release the contaminated-file lists, recovered patient identifiers, and deduplicated splits.
摘要:自動化分類腦腫瘤的MRI是深度學習在醫學影像中的一個廣泛發表的應用,報告的準確率在公共基準上通常超過98%。然而,準確率並未捕捉到基準質量的一個關鍵維度:數據集完整性,定義為在影像、病人和獲取來源層面上測試數據與訓練數據的獨立性。我們引入了一個三層污染框架,包括重複、病人和來源標籤洩漏,以評估這些文獻所依賴的公共語料庫。我們對三個最廣泛使用的語料庫進行審核,與胸部X光的陰性對照進行比較,並量化每一層對九種架構和三種評估條件下測量性能的影響。每一層的污染都很嚴重:主導語料庫的官方測試分割中有28.8%的部分在其自己的訓練分割中有近似的雙胞胎,第二個語料庫有22.3%的測試影像以字節相同的方式洩漏,95.5%的可追溯測試影像與訓練共享病人,且不包含任何解剖結構的檔案標頭特徵在0.959的平衡準確率下將腫瘤與非腫瘤區分開來,與微調的ResNet骨幹相當。意外的結果是,移除每一個識別出的洩漏測試影像後,平衡準確率基本保持不變:去重後的穩定性能並未建立基準完整性。我們的研究結果確立了數據集完整性作為基準質量的一個獨特可測量軸,穩定的領先榜無法證明。對於生物醫學研究,僅僅依賴這些語料庫報告的準確率並不能證明模型已經學會識別腫瘤,而不是利用數據集特定的線索。我們發布了污染檔案列表、恢復的病人標識符和去重的分割。
Trust Is Not a Score: Runtime Assurance Contracts for High-Risk AI Agents
2609.39717v1 by Serhii Zabolotnii
Benchmarks, audits, and agent protocols describe performance, permissions, and repair, but not how observed evidence should change an agent's authority during a consequential task. We call this the assurance-transition gap. We propose a Runtime Assurance Contract (RAC), a policy-level formal schema binding autonomy boundaries, component eligibility, evidence state, transition policy, human-review capacity, and non-compensatory gates. Under RAC, soft metrics may inform routing, whereas a failed or unknown mandatory gate forces retry, switch, escalation, deferral, or stop; aggregate performance cannot authorize action. We define the contract, an evidence record, a permission rule, and five invariants, and illustrate them in clinical, industrial, and judicial failure probes. We then report a deterministic failure-injection study in agentic coding: 280 constructed cases evaluated by a gate conjunction, a score-only rule, and a restricted protocol baseline. At the published example weights and threshold, the score rule admits 80 of 100 block-required injections and all 40 review-required injections. Tuned in hindsight, it matches the conjunction on this corpus. For positive weights, a positive threshold, binary risk signals, zero-signal controls, and an injected case firing each signal alone, we show that exact agreement holds if and only if the threshold does not exceed the smallest weight. A separate set of 18 hand-authored traces checks version-pinned evidence and review transitions against simpler policy variants. In a further prospective synthetic holdout of 24 episodes, two blinded LLM judges assign identical labels to all 72 action attempts; RAC and a separately implemented full stateful baseline both match these labels. These studies test mechanisms on synthetic cases; they establish neither deployed safety nor cross-domain effectiveness.
摘要:基準、審計和代理協議描述了性能、權限和修復,但並未說明在關鍵任務中,觀察到的證據應如何改變代理的權限。我們稱之為保證過渡差距。我們提出了一個運行時保證合約(RAC),這是一個政策層級的正式架構,約束自主邊界、組件資格、證據狀態、過渡政策、人類審查能力和非補償性閘門。在RAC下,軟指標可以用來指導路由,而失敗或未知的強制閘門則強迫重試、切換、升級、延遲或停止;總體性能無法授權行動。我們定義了合約、一個證據記錄、一條許可規則和五個不變量,並在臨床、工業和司法失敗探測中進行了說明。然後,我們報告了一項在代理編碼中的確定性失敗注入研究:280個構建的案例通過閘門聯合、一個僅計分的規則和一個受限的協議基線進行評估。在已發表的示例權重和閾值下,計分規則允許100個區塊所需注入中的80個和所有40個審查所需的注入。事後調整後,它在這個語料庫上與聯合匹配。對於正權重、正閾值、二元風險信號、零信號控制和每個信號單獨觸發的注入案例,我們顯示出精確一致性僅在閾值不超過最小權重時成立。一組18個手工編寫的痕跡檢查版本固定的證據和審查過渡,與更簡單的政策變體進行比較。在進一步的前瞻性合成保留中,24個集數中,兩位盲法LLM評審對所有72次行動嘗試分配了相同的標籤;RAC和一個單獨實施的完整狀態基線都與這些標籤相匹配。這些研究在合成案例上測試機制;它們既未建立已部署的安全性,也未建立跨領域的有效性。
Interpretable Synthetic Medical Tabular Data Generation for Clinical Decision Support Using Fuzzy Cognitive Maps
2610.00391v1 by Michael Vasilakakis, Dimitris K. Iakovidis
Synthetic medical tabular data generation has become essential for developing and validating computer-based medical systems (CBMSs) when real clinical data is restricted due to privacy, ethical, or data availability limitations. Existing probabilistic and deep generative models often lack interpretability and fail to preserve clinically meaningful dependencies, limiting their suitability for safety-critical applications. This paper proposes a novel application of Fuzzy Cognitive Maps (FCMs) in a framework for synthetic medical tabular data generation with explicit causality and privacy preservation. Clinical features are described using linguistically interpretable fuzzy sets, and inter-feature dependencies are encoded as FCM edge weights computed from fuzzy set intersections. Synthetic patient records are generated by propagating randomly initialized linguistic activation vectors through the FCM until convergence, followed by defuzzification to produce clinically coherent numerical values. The approach natively handles mixed data types, and domain constraints common in health records. Experimental evaluation on UCI medical benchmark datasets demonstrates competitive performance under a Train-on-Synthetic-Test-on-Real (TSTR) protocol. The proposed method achieves accuracy of up to 0.81 and AUROC of up to 0.90 on the Heart Disease dataset, matching or exceeding TVAE and Gaussian Copula baselines while running exclusively on CPU. Fidelity metrics including KS Complement (up to 0.91) and Correlation Similarity (up to 0.95) confirm strong statistical coherence, and DCR Baseline Protection scores consistently exceed those of TVAE, confirming adequate privacy guarantees. These results demonstrate that causally grounded, interpretable fuzzy modeling offers a computationally efficient and transparent alternative to deep generative models for trustworthy synthetic data generation in CBMSs.
摘要:合成醫療表格數據生成已成為開發和驗證基於計算機的醫療系統(CBMSs)的必要條件,尤其是在由於隱私、倫理或數據可用性限制而無法使用真實臨床數據的情況下。現有的概率和深度生成模型往往缺乏可解釋性,並且未能保持臨床上有意義的依賴性,這限制了它們在安全關鍵應用中的適用性。本文提出了一種在合成醫療表格數據生成框架中應用模糊認知圖(FCMs)的新方法,該方法具有明確的因果關係和隱私保護。臨床特徵使用語言可解釋的模糊集合來描述,特徵間的依賴性則作為從模糊集合交集計算出的FCM邊權重進行編碼。合成病歷通過將隨機初始化的語言激活向量在FCM中傳播直到收斂,然後進行去模糊化以生成臨床上連貫的數值。該方法本土處理混合數據類型以及健康記錄中常見的領域約束。在UCI醫療基準數據集上的實驗評估顯示,在合成訓練-真實測試(TSTR)協議下表現競爭力。所提出的方法在心臟病數據集上達到高達0.81的準確率和高達0.90的AUROC,與TVAE和高斯聯合基準相匹配或超過,且僅在CPU上運行。包括KS補充(高達0.91)和相關性相似度(高達0.95)在內的保真度指標確認了強大的統計一致性,而DCR基準保護分數始終超過TVAE的分數,確認了足夠的隱私保障。這些結果表明,基於因果關係的可解釋模糊建模為CBMSs中的可信合成數據生成提供了一種計算效率高且透明的替代方案。
CAMOS: Coupled Oscillatory State-Space Model for Multimodal Clinical Time-Series
2609.39484v1 by Maxx Richard Rahman, Mostafa Hammouda, Wolfgang Maass
Longitudinal clinical cohorts are multimodal, irregularly sampled and pervasively incomplete: in ADNI, positron emission tomography and cerebrospinal fluid assays are absent from roughly half of all visits. Linear state-space models handle irregular sampling gracefully but treat a missing modality by masking the input, leaving the transition operator untouched. We prove that this is a representational limitation: the latent state of any linear state-space layer whose transition operator does not depend on the availability pattern is an additive function of the availability indicators, so no such layer can represent an interaction between two modalities being jointly present or jointly absent. We propose CAMOS, which gives each modality a bank of second-order oscillators coupled through a matrix that sits inside the differential equation and is gated by availability, so the transition operator itself becomes a function of which measurements were taken. Coupling invalidates the analysis of uncoupled oscillatory models, and we restore it: a per-channel Gershgorin budget makes the effective stiffness positive definite uniformly over all $2^M$ availability patterns and all gaps, an energy argument charges amplification to availability transitions rather than sequence length, and a channel factorization preserves exact associative parallel scans. On ADNI, CAMOS outperforms uncoupled oscillatory state-space models and clinical fusion models on same-visit staging, landmark prediction and longitudinal forecasting, and under zero-shot transfer to OASIS-3 it is the only model that avoids collapse to the majority class.
摘要:長期臨床隊列是多模態的、不規則取樣的,且普遍不完整:在ADNI中,正電子發射斷層掃描和腦脊液檢測大約在一半的訪問中缺失。線性狀態空間模型優雅地處理不規則取樣,但通過遮蔽輸入來處理缺失的模態,保持轉換運算子不變。我們證明這是一個表徵限制:任何線性狀態空間層的潛在狀態,其轉換運算子不依賴於可用性模式,是可用性指標的加法函數,因此沒有這樣的層能夠表示兩個模態共同存在或共同缺失之間的交互。我們提出CAMOS,為每個模態提供一組通過一個位於微分方程內的矩陣耦合的二階振盪器,並由可用性進行開關,因此轉換運算子本身成為一個函數,取決於採取了哪些測量。耦合使得無耦合振盪模型的分析失效,而我們恢復了它:每通道的Gershgorin預算使得有效剛度在所有$2^M$可用性模式和所有間隙上均為正定,能量論證將放大歸因於可用性轉換而非序列長度,通道因式分解保留了精確的關聯並行掃描。在ADNI上,CAMOS在同訪問分期、地標預測和長期預測方面優於無耦合振盪狀態空間模型和臨床融合模型,並在零樣本轉移到OASIS-3時,它是唯一一個避免崩潰到多數類別的模型。
Towards Trustworthy AI for Glioma Diagnosis: A Task-Aware Evaluation of Uncertainty Quantification
2609.39429v1 by Gonzalo Esteban Mosquera Rojas, Sebastian R. van der Voort, Carolin M. Pirkl, Sandeep Kaushik, Marion Smits, Stefan Klein
Uncertainty Quantification (UQ) is a key requirement for trustworthy AI in high-stakes medical image analysis. In this work, we evaluate UQ in a multi-task Deep Learning framework for MRI-based glioma diagnosis that performs tumor segmentation and predicts IDH mutation status, 1p/19q co-deletion status, and tumor grade. Monte Carlo Dropout (MCD) is used for a detailed task-aware analysis of predictive, aleatoric, and epistemic uncertainty. We assess MC sample convergence, calibration, error detection, selective prediction, associations with segmentation performance, and the effect of voxel-wise uncertainty aggregation on case-level reliability. We also compare MCD with Deep Ensembles (DE) and Monte Carlo Deep Ensembles (MCDE), examine interactions between segmentation quality and classification, and evaluate a composite trust score integrating segmentation and classification uncertainty. Across tasks, uncertainty estimates supported meaningful error detection, while calibration depended on the dropout rate, with moderate rates yielding the most reliable probabilities. Uncertainty decomposition provided task-dependent interpretability but did not consistently improve error detection over predictive uncertainty alone. DE and MCDE showed comparable operational utility, with no method consistently dominating across tasks and metrics. The composite trust score did not consistently outperform classification uncertainty for selective prediction. Overall, our results provide a task-aware evaluation strategy and practical guidance for the development of trustworthy AI for glioma diagnosis.
摘要:不確定性量化(UQ)是高風險醫療影像分析中可信賴人工智慧的關鍵要求。
在這項工作中,我們評估了一個多任務深度學習框架中的UQ,該框架用於基於MRI的膠質瘤診斷,執行腫瘤分割並預測IDH突變狀態、1p/19q共同缺失狀態和腫瘤等級。
使用蒙特卡羅隨機失活(MCD)對預測性、隨機性和認知性不確定性進行詳細的任務感知分析。
我們評估了蒙特卡羅樣本收斂性、校準、錯誤檢測、選擇性預測、與分割性能的關聯,以及體素級不確定性聚合對案例級可靠性的影響。
我們還將MCD與深度集成(DE)和蒙特卡羅深度集成(MCDE)進行比較,檢查分割質量與分類之間的相互作用,並評估整合分割和分類不確定性的綜合信任分數。
在各項任務中,不確定性估計支持有意義的錯誤檢測,而校準則依賴於隨機失活率,中等率產生最可靠的概率。
不確定性分解提供了任務依賴的可解釋性,但並未始終改善僅依賴預測不確定性的錯誤檢測。
DE和MCDE顯示出可比的操作效用,沒有一種方法在各任務和指標中始終佔優。
綜合信任分數在選擇性預測中並未始終優於分類不確定性。
總體而言,我們的結果提供了一種任務感知的評估策略和實用指導,旨在為膠質瘤診斷的可信賴人工智慧發展提供支持。
EHR-RobustGym: Benchmarking and Training Agents for Robust Clinical Reasoning
2609.39371v1 by Yitong Qiao, Yancheng Jin, Lei Liu, Yue Shen, Jian Wang, Jinjie Gu, Zhixuan Chu
In hospital workflows, electronic health records (EHRs) are often noisy, and may not contain the evidence needed to confirm events or measurements referenced in a clinical query. Even when database retrieval succeeds, clinical agents can overlook such discrepancies and return plausible but unsupported answers. We introduce EHR-RobustGym, a scalable and interactive environment for evaluating and training robust clinical agents grounded in noisy EHRs. Built on MIMIC-IV hospital records (365K patients, 31 tables, and over 500M records), EHR-RobustGym comprises 5,486 Clean-Noise pairs spanning six clinical intents and both patient-level and population-level queries. The pairs test robustness to Record-level, Value-level, and Query-level noise, while interactive SQL/Python execution and outcome verification support trajectory collection and training. Evaluating multiple LLMs reveals substantial robustness gaps: average task success across proprietary and large-scale open-weight models drops from 62.2% on Clean questions to 37.9% on Noise questions. At k=4, pass^k consistency falls below 50% for most evaluated models, exposing instability in clinical task completion. Supervised fine-tuning and reinforcement learning in EHR-RobustGym improve performance, with gains generalizing to five external EHR benchmarks. Together, these results position EHR-RobustGym as a testbed for evaluating and improving the evidence-grounded robustness of clinical agents.
摘要:在醫院工作流程中,電子健康紀錄(EHR)通常是嘈雜的,並且可能不包含確認臨床查詢中提到的事件或測量所需的證據。即使數據庫檢索成功,臨床代理也可能忽略這些差異,並返回看似合理但未得到支持的答案。我們介紹了EHR-RobustGym,一個可擴展的互動環境,用於評估和訓練基於嘈雜EHR的穩健臨床代理。EHR-RobustGym建立在MIMIC-IV醫院紀錄上(365K名患者,31個表格,超過5億條紀錄),包括5486對乾淨-嘈雜配對,涵蓋六種臨床意圖以及患者級和人口級查詢。這些配對測試對紀錄級、值級和查詢級噪聲的穩健性,同時互動的SQL/Python執行和結果驗證支持軌跡收集和訓練。對多個LLM的評估顯示出顯著的穩健性差距:在專有和大規模開放權重模型中,乾淨問題的平均任務成功率從62.2%下降到嘈雜問題的37.9%。在k=4時,大多數評估模型的pass^k一致性降至50%以下,顯示臨床任務完成的不穩定性。在EHR-RobustGym中進行的監督微調和強化學習改善了性能,並且增益在五個外部EHR基準上具有普遍性。總體而言,這些結果使EHR-RobustGym成為評估和改善臨床代理的證據基礎穩健性的測試平台。
Structure vs. Chain-of-Thought: Evaluating LLM Criteria Extraction for Depression Severity
2609.39049v1 by Xinkai Chen
A large language model (LLM) can rate depression severity directly from a social media post or mark which clinical criteria the post shows and let code turn the count into a label. The latter is easier to audit because a clinician can check each marked criterion. We compare these approaches on two Reddit corpora using three LLMs (from 9B to frontier scale) and two questionnaires (PHQ-9, BDI-II), and measure agreement with quadratic weighted kappa. For the two frontier models, criteria extraction scores above chain-of-thought on one corpus only when its decision thresholds are fitted on labeled data. Neither model's gain is significant, with or without recalibrating chain-of-thought on the same labels. With thresholds fixed a priori from PHQ-9's criteria, extraction shows no gain on either corpus, even where models mark over two criteria per post. The 9B model behaves differently on a corpus from depression communities. It labels most posts severe, whether prompted directly or with chain-of-thought, while the a priori rule beats both without labels. After chain-of-thought is recalibrated on the same labels, no significant gap remains, consistent with a calibration effect. Yet higher ordinal agreement does not ensure better detection of severe cases. PHQ-9 criteria extraction misses most severe posts, and moving from direct prompting to chain-of-thought and then to extraction increases misses in nearly all comparisons. On the primary corpus, a relabeled stress dataset, a model using that dataset's own features, including word counts from the text, is not significantly different from frontier criteria extraction under the a priori rule.
摘要:一個大型語言模型(LLM)可以直接從社交媒體帖子中評估抑鬱症的嚴重程度,或標記該帖子顯示的臨床標準,並讓代碼將計數轉換為標籤。後者更容易進行審核,因為臨床醫生可以檢查每個標記的標準。我們在兩個Reddit語料庫上比較這些方法,使用三個LLM(從9B到前沿規模)和兩個問卷(PHQ-9,BDI-II),並測量與二次加權kappa的一致性。對於這兩個前沿模型,標準提取分數在一個語料庫上僅在其決策閾值適配於標記數據時超過思考鏈。無論是否重新校準思考鏈,兩個模型的增益都不顯著。當閾值根據PHQ-9的標準事先固定時,提取在任何語料庫上都沒有增益,即使模型每個帖子標記超過兩個標準。9B模型在抑鬱社區的語料庫上表現不同。無論是直接提示還是使用思考鏈,它都將大多數帖子標記為嚴重,而事先規則在沒有標籤的情況下超越了兩者。在思考鏈在相同標籤上重新校準後,沒有顯著的差距,這與校準效應一致。然而,更高的序數一致性並不保證更好地檢測嚴重病例。PHQ-9標準提取錯過了大多數嚴重帖子,從直接提示轉向思考鏈再到提取幾乎在所有比較中都增加了漏掉的情況。在主要語料庫中,一個重新標記的壓力數據集,使用該數據集自身的特徵,包括來自文本的字數,與根據事先規則的前沿標準提取沒有顯著差異。
An Uncertainty-Guided Digital Twin Framework for Online Adaptive Proton Therapy in Head and Neck Cancer: A Feasibility Study
2609.39010v1 by Yizhou Wu, Ryan J. Sanford, Huiqiao Xie, Jie Ding, Shupeng Chen, Tung-Ho Wu, Ping-Hsiu Wu, Justin Roper, Jun Zhou, Minglei Kang, Bill Stokes, Sibo Tian, David S. Yu, Xiaofeng Yang, Chih-Wei Chang
Objective: Head and neck (HN) proton therapy spans six to seven weeks of anatomical change, while offline replanning takes about a week. We present an uncertainty-guided digital twin (UGDT) framework that forecasts treatment-day anatomy before treatment and evaluate whether it generates online adaptive proton therapy (APT) plans of clinical quality. Approach: A library of 302 longitudinal deformations from 88 previously treated HN patients was transported onto each new patient's treatment planning CT (TPCT) using two-step multi-atlas deformable image registration (DIR) built on a pretrained CT foundation model, generating about 284 predicted CTs (pdCTs) with contours per patient. Dispersion of propagated clinical target volume (CTV) contours defined a patient-specific robust margin. In ten patients, the quality assurance CT (QACT) triggering a replan represented treatment-day anatomy, and the physician-approved replan was the baseline. The pdCT most similar to the QACT (pdCT-H) and one from the lowest quartile (pdCT-L) were planned to within about 5% of baseline plan quality, forward-calculated on the QACT, and reoptimized to generate online APT plans. Main results: pdCT plans scored within -0.7% (pdCT-H) and -1.0% (pdCT-L) of baseline. Forward calculation on QACT reduced high-dose CTV D98% to 88.3% and 85.5%. After online reoptimization, D98% recovered to 98.3 +/- 0.3% and 98.2 +/- 0.3%, versus 98.5 +/- 0.4% at baseline. Spinal cord and brainstem doses remained below tolerance, and plan quality scores were within -1.1% (p = 0.19) and -1.7% (p = 0.01) of baseline. Significance: UGDT generated online APT plans comparable in quality to physician-approved offline replans using anatomy forecast before treatment, enabling a transition from reactive offline replanning toward anticipatory online adaptation.
摘要:目標:頭頸部(HN)質子治療的解剖變化持續六到七週,而離線重新規劃約需一週時間。我們提出了一個不確定性引導的數位雙胞胎(UGDT)框架,該框架預測治療當天的解剖結構,並評估其是否能生成臨床質量的在線自適應質子治療(APT)計劃。方法:從88名先前接受治療的HN患者中提取的302個縱向變形被轉移到每位新患者的治療計劃CT(TPCT),使用基於預訓練CT基礎模型的兩步多圖譜可變形影像註冊(DIR),為每位患者生成約284個預測CT(pdCT)及其輪廓。傳播的臨床靶區體積(CTV)輪廓的分散定義了患者特定的穩健邊界。在十名患者中,觸發重新規劃的質量保證CT(QACT)代表治療當天的解剖結構,並且醫生批准的重新規劃為基準。與QACT最相似的pdCT(pdCT-H)和來自最低四分位數的pdCT(pdCT-L)計劃的質量約在基準計劃質量的5%以內,並在QACT上進行前向計算,然後重新優化以生成在線APT計劃。主要結果:pdCT計劃的得分在基準的-0.7%(pdCT-H)和-1.0%(pdCT-L)之內。在QACT上的前向計算將高劑量CTV D98%降低至88.3%和85.5%。在線重新優化後,D98%恢復至98.3 +/- 0.3%和98.2 +/- 0.3%,而基準為98.5 +/- 0.4%。脊髓和腦幹的劑量保持在耐受範圍以下,計劃質量得分在基準的-1.1%(p = 0.19)和-1.7%(p = 0.01)之內。意義:UGDT生成的在線APT計劃在質量上可與醫生批准的離線重新規劃相媲美,並使用治療前的解剖預測,使得從反應式離線重新規劃向預測式在線適應的轉變成為可能。
Scoring Higher, Answering Worse: Mitigating Reward Hacking in Rubric-Based RL via Protocol-Level Rubrics
2609.38847v1 by Maoqi Liu, Junwei He, Bowen Zhang, Feiran Li, Wentao Ma, Rongyi Lin, Shuhan Zhong, Quan Fang
Rubric-based reinforcement learning (Rubric-RL) trains language models where no verifier exists. A judge checks each criterion of a rubric, and the verdicts are aggregated into a reward, most often by a weighted sum. We show that this additive aggregation is the weak point. Under a sum, criteria compensate for one another: a policy that misses the one decision that matters can buy the points back with advice nobody asked for. On clinical consultation, such a policy scores higher and answers worse. Rubric coverage rises while appropriateness on held-out physician criteria falls below the untrained model. The medical criteria are not to blame. Grouped so that they must hold together, the same criteria, unchanged to the word, recover a third of the loss; shorter answers recover almost none. We therefore propose Protocol-level Rubrics (ProRubric), which keeps what the criteria ask for and changes how they are aggregated. It groups a checklist into a few protocol-level dimensions. A dimension counts only when all of its criteria hold and its failure clause does not fire. The grouping is done once, offline, and leaves the optimizer unchanged. ProRubric raises appropriateness by 10.8 points without losing coverage and has the best seven-benchmark average at both scales. Reward validity is set not only by what a rubric verifies, but by how it aggregates. Code is available at https://github.com/Estrellajer/ProRubric
摘要:基於評分標準的強化學習(Rubric-RL)訓練語言模型,當中不存在驗證者。評審檢查評分標準的每一項準則,並將判決匯總為獎勵,通常是通過加權總和。我們顯示這種加法匯總是其弱點。在總和下,準則之間相互補償:一個錯過了關鍵決策的策略可以用沒有人要求的建議來彌補分數。在臨床諮詢中,這樣的策略得分較高,但回答卻較差。評分標準的覆蓋率上升,而在保留的醫師準則上的適當性則低於未經訓練的模型。醫療準則並不是問題所在。這些準則被分組在一起,必須保持一致,未改變的相同準則恢復了三分之一的損失;較短的回答幾乎沒有恢復。因此,我們提出了協議級別評分標準(ProRubric),它保留了準則所要求的內容並改變了它們的匯總方式。它將檢查清單分組為幾個協議級別的維度。只有當所有準則都成立且其失敗條款未觸發時,該維度才計算。分組一次性完成,離線進行,並不改變優化器。ProRubric在不失去覆蓋率的情況下提高了10.8分的適當性,並在兩個尺度上擁有最佳的七項基準平均值。獎勵的有效性不僅由評分標準所驗證的內容決定,還由其匯總方式決定。代碼可在 https://github.com/Estrellajer/ProRubric 獲得。
CRAFT: Causal Responsibility and Failure Tracing in Medical Vision Language Models
2609.38810v1 by Chunzheng Zhu, Jiaqi Zeng, Hongbo Zhao, Yihang Chen, Yijun Wang, Jianxin Lin
As vision language models are increasingly deployed in clinical diagnosis, understanding how they internally resolve competing visual and textual signals becomes a safety imperative. Existing mechanistic analyses remain confined to unimodal text and offer no explanation for why a single misleading sentence can override a correct image based diagnosis, or why a model commits to a confident answer despite insufficient visual evidence. We find that these two safety risks, arbitration failure where textual context overrides visual grounding and brake failure where the model commits without adequate evidence, are mediated by spatially disjoint attention head populations: arbitration heads form a mid-to-deep wideband reflecting cross-layer evidence competition, while brake heads concentrate in a narrow middle-to-late layer band that regulates evidence sufficiency and abstention behavior. To ground these observations in causal circuitry, we introduce CRAFT, which localizes each failure mode to a minimal causal head set via dual criteria and verifies necessity and sufficiency through temporal probes and Tuned Lens trajectory analysis. Excising arbitration heads sharply reduces conflict following with negligible degradation on clean inputs, while excising brake heads restores appropriate abstention under degraded visual evidence. The two interventions target spatially disjoint head sets and produce distinct corrective effects, underscoring the mechanistic separability of the failure modes. Experiments across multiple medical VQA benchmarks and VLM architectures validate both the localization and interventions, demonstrating that the identified heads causally drive each failure mode and that targeted modulation generalises without retraining. The code is available at GitHub repository.
摘要:隨著視覺語言模型在臨床診斷中的應用日益增多,了解它們如何內部解決競爭的視覺和文本信號變得至關重要。現有的機制分析仍然局限於單一模式的文本,並未解釋為何一個誤導性的句子可以覆蓋正確的影像基礎診斷,或為何模型在視覺證據不足的情況下仍然堅持給出自信的答案。我們發現這兩種安全風險,即文本上下文覆蓋視覺基礎的仲裁失敗,以及模型在缺乏充分證據的情況下做出承諾的剎車失敗,是由空間上不相交的注意力頭群體所介導:仲裁頭形成中到深的寬頻,反映跨層證據競爭,而剎車頭則集中在狹窄的中到後層帶,調節證據的充分性和放棄行為。為了將這些觀察與因果電路相結合,我們引入了CRAFT,該方法通過雙重標準將每個失敗模式定位到最小的因果頭集,並通過時間探測和調整透鏡軌跡分析驗證必要性和充分性。切除仲裁頭顯著減少了衝突,對於乾淨輸入的降解幾乎可以忽略不計,而切除剎車頭則在視覺證據降級的情況下恢復了適當的放棄行為。這兩種干預針對空間上不相交的頭集,產生了不同的修正效果,強調了失敗模式的機制可分性。在多個醫學VQA基準和VLM架構上的實驗驗證了定位和干預,證明所識別的頭因果驅動每個失敗模式,並且針對性的調節在不重新訓練的情況下具有普遍性。代碼可在GitHub倉庫中獲得。
Sense and Sensitivity: Benchmarking LLM Clinical Triage Recommendations with Physician Experts
2609.38600v1 by Abinitha Gourabathina, Haoran Zhang, Yuexing Hao, Walter Gerych, Marzyeh Ghassemi
As large language models (LLMs) are increasingly used in clinical settings, it is critical to evaluate their reliability under realistic variation in clinical text. We study this question in clinical triage, comparing LLMs to practicing physicians under text perturbations that preserve the underlying clinical setting. We introduce a benchmark of over 6,000 clinical scenarios, 7,000 physician annotations, and 225,000 model responses. Using this benchmark, we make two key observations. First, LLMs are more likely than physicians to recommend unnecessary care at baseline, and this tendency increases under perturbed inputs. Further, we find that LLM recommendations are more sensitive to gender and tone perturbations than human recommendations. Together, these results demonstrate that LLMs can vary under clinically irrelevant textual changes, highlighting the need for deployment-oriented evaluations grounded in expert physician behavior.
摘要:隨著大型語言模型(LLMs)在臨床環境中的使用日益增加,評估它們在臨床文本的現實變異下的可靠性變得至關重要。
我們在臨床分診中研究這個問題,將LLMs與在保持基礎臨床環境的文本擾動下的執業醫生進行比較。
我們引入了一個基準,包含超過6,000個臨床場景、7,000個醫生註釋和225,000個模型回應。
利用這個基準,我們做出兩個關鍵觀察。
首先,LLMs在基線下比醫生更可能推薦不必要的護理,這種傾向在擾動輸入下會增加。
此外,我們發現LLM的建議對性別和語氣擾動的敏感度高於人類的建議。
這些結果共同表明,LLMs在臨床無關的文本變化下可能會有所不同,突顯了基於專家醫生行為的部署導向評估的必要性。
Defining and Categorising Human-AI Interactions in Clinical Trials: A Multidimensional Human-AI Classification Approach
2609.38559v1 by Sandra Woolley, Tim Collins, Khalid Khattak, Illia Chernomorets, Ariane Arevalo, Chris Richardson
This paper examines human-AI interactions (HAIIs) in clinical trials and presents a multidimensional categorisation framework that classifies interactions according to AI tasks, human-AI relationships, interaction configurations and interacting human groups. We define HAII, examine existing taxonomies and extend existing categorisation approaches through this novel multidimensional framework. We purposively sampled 15 clinical trials from a previously reported dataset. Each trial was independently categorised by two human reviewers and six large language model (LLM) classifiers. The proposed categorisation provides a structured method for the consistent identification, comparison and synthesis of human-AI interactions across clinical-trial records. The framework is intended to support more consistent comparison and synthesis of AI-related clinical trials and to make explicit the different forms of human involvement associated with AI interventions. The results demonstrate the potential for LLM-assisted categorisation while indicating the continuing importance of human judgement where trial records are incomplete or ambiguous. The principal contribution is a proposed multidimensional framework that brings together AI tasks, human-AI relationships, interaction configurations and interacting human groups within a single approach designed for clinical-trial records. Its significance lies in its potential to support more systematic identification, comparison and synthesis of how humans and AI interact in clinical trials.
摘要:這篇論文探討了臨床試驗中的人類-人工智慧互動(HAIIs),並提出了一個多維度的分類框架,根據人工智慧任務、人類-人工智慧關係、互動配置和互動人群對互動進行分類。
我們定義了HAII,檢視現有的分類法,並通過這個新穎的多維框架擴展現有的分類方法。
我們有目的地從先前報告的數據集中抽取了15個臨床試驗。
每個試驗由兩位人類評審和六個大型語言模型(LLM)分類器獨立進行分類。
所提出的分類提供了一種結構化的方法,用於在臨床試驗記錄中一致地識別、比較和綜合人類-人工智慧互動。
該框架旨在支持對人工智慧相關臨床試驗的更一致的比較和綜合,並明確不同形式的人類參與與人工智慧干預相關聯。
結果顯示了LLM輔助分類的潛力,同時指出在人類判斷仍然重要的情況下,當試驗記錄不完整或模糊時,仍需依賴人類判斷。
主要貢獻是一個提出的多維框架,將人工智慧任務、人類-人工智慧關係、互動配置和互動人群整合在一個針對臨床試驗記錄的單一方法中。
其重要性在於它能支持對人類和人工智慧在臨床試驗中互動的更系統的識別、比較和綜合。
Personalized State-Transition-Aware Memory for Clinical Agents
2609.38490v1 by Maryam Haghifam, Zahra Rajabi, Yizhou Sun, Carlos Morato
Large language model (LLM) agents that reason over clinical records must track changes in a patient's state while preserving the history needed to understand them. Simply accumulating memories leaves it unclear which information still applies, whereas overwriting earlier memories can erase evidence needed to reconstruct treatment history and clinical trajectories. We introduce STAM, a state-transition-aware memory framework that records state changes as new clinical entries arrive. STAM combines semantic retrieval with typed clinical relations to identify affected memories, maintaining current information in Active and superseded or resolved information in History. At read time, a query-dependent gate selectively serves historical memory. Across four longitudinal clinical benchmarks, we evaluate STAM with downstream question answering, direct state-maintenance diagnostics, and comparisons at approximately matched context lengths.
摘要:大型語言模型 (LLM) 代理人必須在保留理解病歷所需的歷史的同時,追蹤病人狀態的變化。單純地累積記憶會使得不清楚哪些信息仍然適用,而覆蓋早期記憶則可能抹去重建治療歷史和臨床軌跡所需的證據。我們介紹了 STAM,一個狀態轉換感知的記憶框架,該框架在新的臨床條目到達時記錄狀態變化。STAM 結合語義檢索和類型化的臨床關係來識別受影響的記憶,並在主動信息中保持當前信息,在歷史中保持被取代或解決的信息。在讀取時,查詢依賴的閘門選擇性地提供歷史記憶。在四個縱向臨床基準中,我們通過下游問題回答、直接狀態維護診斷以及在大約匹配的上下文長度下進行比較來評估 STAM。
KlinikeBench: Evaluating Language Models Beyond Diagnostic Accuracy
2609.38480v1 by Xueting Fang, Zehui Li, Yang Yang, Camilla Giovino, Shubh K. Patel, Shailly Prajapati, Vallijah Subasri, Caihua Shan
Most clinical benchmarks evaluate language models (LMs) on diagnosis using complete case descriptions. In clinical practice, however, patients present information in different ways, and clinicians must obtain relevant history and determine which examinations are needed before reaching a diagnosis. Diagnostic accuracy alone therefore cannot establish whether an agent gathered essential information or conducted an appropriate clinical assessment. Furthermore, existing benchmarks lack professional clinicians' verification. To address this gap, we introduce KlinikeBench, a benchmark of 333 clinician-authored tasks, each providing an isolated sandbox environment with a virtual patient, clinical tools, and task-specific success criteria. More than 35 clinicians contributed to case authoring and benchmark evaluation. In an empirical study, clinicians gave simulated dialogues higher mean quality ratings than reference conversations, which is adapted from real conversation. In each task, an LM has a fixed budget of turns to communicate with the patient, ask about relevant history, request examinations, follow action constraints, and record a final diagnosis. We score these steps separately as well as together. Across 31 models and seven model families, the best-performing models (e.g., GPT-6-astra and Claude Opus 5) succeed on less than 30% of tasks, even though their diagnosis accuracy reaches 90.7%. Some models benefit from talking with the patient; others diagnose well from a complete chart but perform much worse in conversation. Overall, KlinikeBench provides a testbed for evaluating the full clinical encounter and reveals a substantial gap between diagnostic accuracy and performance in interactive clinical assessment.
摘要:大多數臨床基準在診斷上評估語言模型(LMs)時使用完整的病例描述。然而,在臨床實踐中,患者以不同的方式呈現信息,臨床醫生必須獲取相關病史並確定需要哪些檢查,才能達成診斷。因此,僅僅依賴診斷準確性無法確定一個代理是否收集了必要的信息或進行了適當的臨床評估。此外,現有的基準缺乏專業臨床醫生的驗證。為了解決這一問題,我們推出了KlinikeBench,這是一個包含333個臨床醫生撰寫的任務的基準,每個任務都提供一個獨立的沙盒環境,內有虛擬患者、臨床工具和特定任務的成功標準。超過35位臨床醫生參與了案例撰寫和基準評估。在一項實證研究中,臨床醫生給模擬對話的平均質量評分高於參考對話,而參考對話是從真實對話中改編而來。在每個任務中,LM有固定的回合預算來與患者溝通,詢問相關病史,請求檢查,遵循行動約束,並記錄最終診斷。我們分別對這些步驟進行評分,也會綜合評分。在31個模型和七個模型系列中,表現最佳的模型(例如,GPT-6-astra和Claude Opus 5)在不到30%的任務中成功,儘管它們的診斷準確率達到90.7%。一些模型在與患者交談時受益;其他模型從完整的病歷中診斷良好,但在對話中表現卻差得多。總體而言,KlinikeBench提供了一個評估完整臨床互動的測試平台,並揭示了診斷準確性與互動臨床評估表現之間的重大差距。
PrivMeSA: Privacy-Aware Self-Evolving Multi-Agent System for Medicine via Local-Remote LLM Collaboration
2609.38458v1 by Dannong Wang, Yuran Zhang, Bian Sun, Alex Stinard, Yuzhang Shang, Song Wang, Yu Tian
Clinical large language model (LLM) agents deployed locally can consult more capable remote models, but doing so risks exposing patient information. Privacy-conscious delegation places disclosure decisions with a local agent, yet removing explicit identifiers is insufficient: quasi-identifiers can accumulate across multi-turn consultations and repeated patient visits to enable re-identification. We introduce PrivMeSA, a privacy-aware self-evolving multi-agent system that learns to control disclosure and retains remote expertise for local reuse. A local agent manages each encounter and consults remote specialists that may request additional information. Reinforcement learning balances task accuracy against direct disclosure and registry-based re-identification risk, with privacy evaluated over the complete outbound transcript of each encounter. A local lesson memory distills completed consultations into generalized clinical guidance and retrieves relevant lessons before transmission, allowing subsequent cases to reuse expertise without another remote exchange. Memory grows without additional outcome labels or parameter updates. On an emergency-department benchmark built from MIMIC-IV-ED records, PrivMeSA improves mean task accuracy over delegation by up to 15.8 percentage points. In the same setting, PrivMeSA reduces the disclosure of personal details from 98.0% to 0.2% of cases and the share of cases in which the patient can be narrowed to ten or fewer registry patients from 74% to 0%.
摘要:臨床大型語言模型(LLM)代理在本地部署可以諮詢更強大的遠程模型,但這樣做有可能暴露患者信息。注重隱私的委派將披露決策交給本地代理,但去除明確的識別符號並不足夠:準識別符號可以在多輪諮詢和重複的患者訪問中累積,從而實現重新識別。我們介紹了PrivMeSA,一個隱私意識的自我演變多代理系統,能夠學習控制披露並保留遠程專業知識以供本地重用。本地代理管理每次接觸並諮詢可能要求額外信息的遠程專家。強化學習在任務準確性和直接披露及基於登記的重新識別風險之間取得平衡,並在每次接觸的完整外發記錄上評估隱私。本地教訓記憶將已完成的諮詢提煉成通用的臨床指導,並在傳輸前檢索相關的教訓,允許後續案例在不進行另一個遠程交換的情況下重用專業知識。記憶在沒有額外結果標籤或參數更新的情況下增長。在基於MIMIC-IV-ED記錄建立的急診部門基準上,PrivMeSA將任務準確性的平均值提高了最多15.8個百分點。在同一環境中,PrivMeSA將個人詳細信息的披露從98.0%降低到0.2%,並將患者可以縮小到十名或更少登記患者的案例比例從74%降低到0%。
Colorectal Cancer Segmentation with Adaptive Augmentation and Multi-Resolution Ensemble Models
2609.38419v1 by Ümit Mert Çağlar, Alptekin Temizel
Colorectal cancer (CRC) is the second most deadly and third most common cancer, and the leading cause of death among gastrointestinal cancers. Early diagnosis is crucial for the treatment of this cancer and increasing the survival rates. Although CRC is more common in developed regions, its occurrence is also increasing in developing regions as well. CRC diagnosis relies on histopathology assessment post-biopsy. Automated deep learning algorithms can significantly reduce diagnosis time, enhancing efficiency and supporting timely clinical decisions. We present an automated segmentation pipeline for whole-slide histopathology images that labels tumor grades 1-3 and normal mucosa. It utilizes dense prediction transformers with various encoder backbones, overlapping patches, and test-time augmentation. An adaptive augmentation policy, guided by large language models, further improves training. Top models were ensembled via soft voting, and mask refining post-processing steps, Gaussian blurring, morphological closing, and connected components analysis. On a colorectal cancer grade dataset, our method improved the F1 score from 62.92 to 69.84. Code is available here: github.com/caglarmert/ICIP2025
摘要:大腸癌(CRC)是第二大致死率和第三大常見癌症,也是腸胃道癌症中最主要的死亡原因。早期診斷對於這種癌症的治療和提高存活率至關重要。儘管CRC在發達地區更為常見,但在發展中地區的發生率也在增加。CRC的診斷依賴於活檢後的組織病理學評估。自動化深度學習算法可以顯著縮短診斷時間,提高效率並支持及時的臨床決策。
我們提出了一個自動化的整張幻燈片組織病理圖像分割管道,該管道標記腫瘤等級1-3和正常黏膜。它利用具有各種編碼器骨幹的密集預測Transformer、重疊補丁和測試時增強。由大型語言模型引導的自適應增強策略進一步改善了訓練。頂尖模型通過軟投票進行集成,並進行了掩膜精煉後處理步驟、高斯模糊、形態學關閉和連通組件分析。在一個大腸癌等級數據集上,我們的方法將F1分數從62.92提高到69.84。代碼可在此獲得:github.com/caglarmert/ICIP2025
Aegis: Generative Gradient Masking for Privacy-Preserving Medical Federated Learning
2609.38339v1 by Chaoyu Zhang, Shanghao Shi, Heng Jin, Ning Wang, Y. Thomas Hou, Wenjing Lou
Federated learning (FL) has become a foundational paradigm for multi-institutional medical AI, allowing hospitals and research centers to jointly train diagnostic models without exchanging patient records. This privacy promise, however, is increasingly contested: a malicious or honest-but-curious server can launch model inversion attacks (MIAs) that reconstruct private patient images directly from shared model updates, and recent scalable, closed-form attacks penetrate even secure aggregation at clinically realistic batch sizes. Existing defenses face an unsatisfactory dilemma. Gradient-perturbation methods such as differential privacy and pruning trade away the diagnostic accuracy on which clinical reliability depends, while cryptographic protocols add system complexity yet still leave updates exposed to these scalable attacks. We propose Aegis, a principled client-side defense that breaks this dilemma without perturbing patient data or modifying the FL protocol. Our key insight is that the success of every known MIA is fundamentally bounded by the local batch size relative to the model's leakage capacity; once this limit is exceeded, distinct samples collide and reconstructions collapse into indistinguishable mixtures. Aegis turns this universal bottleneck into a defense: each client superimposes onto its real update a masking gradient computed on locally synthesized, task-relevant data, deliberately pushing the effective batch beyond the attack's recovery capacity. We complement the design with theoretical convergence guarantees under standard convex assumptions and evaluate Aegis on MNIST, CIFAR-10, and three MedMNIST modalities (chest X-ray, abdominal CT, colon pathology). Aegis neutralizes three state-of-the-art MIAs while preserving model utility and incurring only modest overhead, offering a practical privacy primitive for medical FL.
摘要:聯邦學習(FL)已成為多機構醫療人工智慧的基礎範式,使醫院和研究中心能夠共同訓練診斷模型,而無需交換病歷記錄。 然而,這一隱私承諾正受到越來越多的質疑:一個惡意或誠實但好奇的伺服器可以發動模型反演攻擊(MIA),直接從共享的模型更新中重建私人病人影像,而最近可擴展的封閉形式攻擊甚至能夠穿透臨床現實批次大小的安全聚合。 現有的防禦面臨著不令人滿意的困境。 像差分隱私和修剪這樣的梯度擾動方法犧牲了臨床可靠性所依賴的診斷準確性,而加密協議則增加了系統的複雜性,卻仍然讓更新暴露於這些可擴展的攻擊之下。 我們提出了Aegis,一種原則性的客戶端防禦,打破了這一困境,而不擾動病人數據或修改FL協議。 我們的關鍵見解是,每個已知的MIA的成功在根本上受到相對於模型泄漏能力的本地批次大小的限制;一旦超過這一限制,不同的樣本將發生碰撞,重建將崩潰為無法區分的混合物。 Aegis將這一普遍瓶頸轉化為防禦:每個客戶端在其真實更新上疊加一個基於本地合成的、與任務相關的數據計算出的掩蔽梯度,故意將有效批次推向超過攻擊的恢復能力。 我們在標準凸假設下補充了理論收斂保證,並在MNIST、CIFAR-10和三種MedMNIST模態(胸部X光、腹部CT、結腸病理)上評估了Aegis。 Aegis中和了三種最先進的MIA,同時保留了模型的效用,並僅產生適度的開銷,為醫療FL提供了一種實用的隱私原語。
A Proposed Rubric for Evaluating Expressed Clinical Reasoning in Large Language Model Responses
2609.37788v3 by Zhangshu Joshua Jiang, Zina Ibrahim, James T. Teo
We propose a rubric for assessing expressed clinical reasoning in model responses, drawing on three bodies of work: medical education assessment frameworks (ART, SCT, Key Feature Problems and OSCE); clinical LLM benchmarks (MedR-Bench, HealthBench, TIMER-Bench, DR. BENCH, PrIME-LLM and PatientSafeBench); and general LLM reasoning evaluation research, including the Factuality-Validity-Coherence-Utility taxonomy, FaithCoT-Bench and C2-Faith. We use groundedness as a clinically oriented adaptation of the taxonomy's factuality category. The rubric brings these concepts together in a multidimensional framework for scoring free-text responses to gold-standard clinical vignettes. It includes provisional behavioural anchors, applicability rules and a separate flag for case-specific safety-critical errors. General-domain frameworks inform its design but are not treated as validated clinical instruments. The rubric does not replace case-specific reference criteria or the task-specific metrics of existing benchmarks. It has not yet been tested for inter-rater reliability, construct validity or clinical utility. Its immediate purpose is to make evaluation decisions explicit and open to scrutiny before empirical testing.
摘要:我們提出了一個用於評估模型回應中表達的臨床推理的評分標準,這基於三個研究領域:醫學教育評估框架(ART、SCT、關鍵特徵問題和OSCE);臨床LLM基準(MedR-Bench、HealthBench、TIMER-Bench、DR. BENCH、PrIME-LLM和PatientSafeBench);以及一般LLM推理評估研究,包括事實性-有效性-一致性-實用性分類法、FaithCoT-Bench和C2-Faith。我們使用基於實證的概念作為該分類法事實性類別的臨床導向調整。這個評分標準將這些概念整合到一個多維框架中,用於對金標準臨床小插曲的自由文本回應進行打分。它包括臨時行為錨點、適用性規則以及針對特定案例的安全關鍵錯誤的單獨標記。一般領域框架為其設計提供了指導,但不被視為經過驗證的臨床工具。該評分標準並未取代特定案例的參考標準或現有基準的任務特定指標。它尚未針對評分者間可靠性、構念效度或臨床實用性進行測試。其直接目的在於使評估決策明確並可接受檢視,然後再進行實證測試。
Multi-Site Real-World Performance of Commercial AI for Pulmonary and Incidental Pulmonary Embolism Detection
2609.37750v1 by Aawez Mansuri, Mohammadreza Chavoshi, Theodorus Dapamede, Wasif Bala, Beatrice Brown-Mulry, Rohan Isaac, Bardia Khosravi, Hanzhou Li, Frank Li, John T. Moon, Chad Robichaux, Dan I. G. Cohen-Addad, Ninad V. Salastekar, Janice Newsome, Judy W. Gichoya, Hari Trivedi
Pulmonary embolism (PE) is a leading cause of cardiovascular mortality, yet the real-world performance of FDA-cleared AI detection models remains incompletely characterized. We retrospectively evaluated two FDA-cleared AI algorithms from a single commercial platform (Aidoc Medical BriefCase), one for PE triage on dedicated CT pulmonary angiography (CTPA; n = 30,678) and one for incidental PE (iPE) detection on routine contrast-enhanced CTs (n = 37,191), across a 17-facility academic health system. Reference-standard labels were extracted from radiology reports using a validated LLM pipeline (97% accuracy, kappa = 0.94). The PE model achieved 86.8% sensitivity and 99.1% specificity, with sensitivity declining from 99.3% for saddle emboli to 72.9% for subsegmental PE, and from 89.7% for acute to 65.3% for non-acute PE. The iPE model achieved 73.5% sensitivity and 99.8% specificity. Both models demonstrated lower sensitivity than FDA-clearance benchmarks while exceeding cleared specificity, with diminishing performance for peripheral and non-acute emboli mirroring known human reader limitations and underscoring the need for standardized post-market surveillance of AI-enabled medical devices.
摘要:肺栓塞(PE)是心血管死亡的主要原因,但FDA批准的AI檢測模型在實際應用中的表現仍然未完全明確。我們回顧性地評估了來自單一商業平台(Aidoc Medical BriefCase)的兩個FDA批准的AI算法,一個用於專用CT肺動脈造影(CTPA;n = 30,678)的PE分流,另一個用於常規對比增強CT(n = 37,191)的偶然PE(iPE)檢測,涵蓋了17家學術醫療系統。參考標準標籤是通過經驗證的LLM管道(97%準確率,kappa = 0.94)從放射學報告中提取的。PE模型的敏感性達到86.8%,特異性為99.1%,其中敏感性從鞍狀栓塞的99.3%下降到亞段PE的72.9%,從急性PE的89.7%下降到非急性PE的65.3%。iPE模型的敏感性為73.5%,特異性為99.8%。兩個模型的敏感性均低於FDA批准的基準,但特異性超過批准標準,對於周邊和非急性栓塞的表現下降反映了已知的人類讀者限制,並強調了對AI驅動醫療設備標準化市場後監測的需求。
Spatiotemporal Hyperedges for EEG Seizure Detection and Prediction
2609.37730v1 by Hyunju Kim, Sheo Yon Jhin, Noseong Park, Nabil Imam
Seizure detection and prediction from EEG are clinically important but challenging because seizures are rare, temporally localized, and propagate as coordinated events across multiple channels. Recent dynamic graph neural networks model this by running a temporal model over a sequence of per-time-step pairwise channel edges. However, this pairwise construction misses the spatiotemporal coupling that constitutes a seizure, at substantial training cost. We propose HyBrain, which summarizes spatiotemporal EEG evidence through a small set of soft hyperedges rather than pairwise edges. A per-channel Mamba backbone produces one token per (channel, second), and a spatiotemporal hyperedge block pools these tokens into E_h shared group embeddings through soft memberships and broadcasts them back. The same encoder serves three downstream tasks: window-based detection, one-second point-wise detection, and preictal seizure prediction. On TUSZ and CHB-MIT, HyBrain achieves the best AUROC on every reported setting against ten baselines, with the largest gap on long-clip preictal prediction. It also matches the most efficient baselines in training time and peak GPU memory. A qualitative analysis shows that even a single learned hyperedge cleanly captures the preictal -> ictal -> postictal trajectory on a real seizure clip.
摘要:癲癇發作的檢測和預測從腦電圖(EEG)中提取是臨床上重要但具有挑戰性的,因為癲癇發作是罕見的、時間上局部的,並且作為協調事件在多個通道中傳播。最近的動態圖神經網絡通過在每個時間步的成對通道邊緣上運行一個時間模型來模擬這一點。然而,這種成對的構建錯過了構成癲癇發作的時空耦合,並且訓練成本相當高。我們提出了HyBrain,它通過一小組軟超邊而不是成對邊來總結時空EEG證據。每個通道的Mamba主幹為每個(通道,秒)生成一個標記,時空超邊塊通過軟成員資格將這些標記聚合為E_h共享的群組嵌入並將其廣播回去。相同的編碼器服務於三個下游任務:基於窗口的檢測、一秒點檢測和癲癇發作前預測。在TUSZ和CHB-MIT上,HyBrain在所有報告的設置中對比十個基準達到最佳的AUROC,在長片段癲癇發作前預測上差距最大。它在訓練時間和峰值GPU內存方面也與最有效的基準相匹配。質性分析顯示,即使是一個學習到的超邊也能清晰地捕捉到真實癲癇片段中的癲癇發作前 -> 發作 -> 發作後的軌跡。
Correct, Don't Delete: Mitigating Emergent Misalignment with Corrective Supervision
2609.37624v1 by Jacob Epifano
Fine-tuning a language model on a narrow set of harmful demonstrations, such as bad medical advice, can make it broadly misaligned on unrelated questions, a phenomenon known as emergent misalignment (EM). The usual defense is to find the offending rows and delete them, but a row locator failed our held-out test and deleting rows helps less than expected. We ask a different question: given a fixed set of poisoned rows, is it better to correct them than to remove them? We fine-tune Qwen2.5-14B-Instruct on a mixture of bad medical advice and benign chat data, select a quarter of the poison rows in advance, and either delete them or replace each with a corrected answer to the same prompt, keeping everything else the same. Replacing the rows cuts the EM rate by about a third and improves answers on held-out medical questions, while deleting the same rows has little measurable effect. The advantage is larger when half the poison rows are corrected, and it holds on a second base model and a second misaligned model organism. The content of the replacement appears to matter: paraphrasing the rows while keeping their bad advice shows no clear benefit, and the correct answers distributed with the dataset appear to do about as well as our rewriter's. Realigning an already-poisoned model with further fine-tuning is known to work, but which data does the work has not been compared directly. We find that a short round of training on corrections beats the same amount of training on generic chat data, that corrections on other medical prompts do roughly as well as corrections of the poisoned prompts themselves, and that instructing the correction writer to model a careful, harm-avoiding assistant adds no measurable benefit over plain corrections. In the settings we tested, correcting harmful training data reduces EM more than deleting it.
摘要:微調一個語言模型於一組狹窄的有害示範,例如不良的醫療建議,可能會使其在無關問題上廣泛地失調,這一現象被稱為新興失調(EM)。通常的防禦方法是找到有問題的行並刪除它們,但行定位器在我們的保留測試中失敗,刪除行的效果也不如預期。我們提出一個不同的問題:在給定一組固定的有毒行的情況下,修正它們是否比刪除它們更好?我們對Qwen2.5-14B-Instruct進行微調,使用不良醫療建議和良性聊天數據的混合,提前選擇四分之一的有毒行,並選擇刪除它們或用相同提示的修正答案替換每一行,保持其他一切不變。替換這些行將EM率降低了約三分之一,並改善了對保留醫療問題的回答,而刪除相同的行幾乎沒有可測量的效果。當修正一半的有毒行時,這一優勢更大,並且在第二個基準模型和第二個失調模型上也成立。替換內容似乎很重要:在保持其不良建議的同時對行進行意譯並未顯示出明顯的好處,與數據集一起分發的正確答案似乎與我們的重寫者的表現相當。在已經被污染的模型上進行進一步微調以重新對齊是已知有效的,但哪些數據發揮作用尚未直接比較。我們發現,對修正進行短期訓練的效果超過了對通用聊天數據進行相同量的訓練,對其他醫療提示的修正效果大致與對有毒提示本身的修正相當,並且指導修正作者模擬一個小心、避免傷害的助手並未帶來比普通修正更可測量的好處。在我們測試的設置中,修正有害的訓練數據比刪除它更能減少EM。
ReLMem: Learning Recurrent Memory for Longitudinal EHR Modeling
2609.37587v1 by Zijie Meng, Xiwei Dai, Yingying Zhang, Jian Wu, Xian Wu, Zuozhu Liu
Longitudinal electronic health record (EHR) modeling requires integrating new visits with an expanding patient history. Yet the continual accumulation of clinical information imposes increasing computational and memory costs on large language models (LLMs) when they process and retain complete patient histories. A practical alternative is visit-wise recurrent compression, which incorporates each incoming visit into a compact, continually updated patient memory. However, under a fixed memory budget, successive updates must integrate new information without progressively losing critical historical evidence needed to subsequent tasks. To address this challenge, we introduce Recurrent Longitudinal Memory (ReLMem), a framework that learns to maintain fixed-capacity patient memory for efficient downstream prediction with a frozen LLM. ReLMem equips this LLM with lightweight compression adapters to recurrently update the memory from its previous state and each incoming visit, without rereading earlier records. Specifically, we develop a multi-granularity optimization strategy to preserve task-relevant information throughout recurrent updates and support downstream prediction from the final memory. The intermediate supervision aligns attention outputs from compressed memory and the full history under identical queries, while prediction supervision minimizes cross-entropy with ground truth answers conditioned on the final memory. On EHR-based medication prediction, ReLMem approaches the F1 scores of full-history baseline while reducing average retained historical storage by 97.1%. Under the same memory budget, it improves macro- and micro-F1 over the strongest compressed-memory baseline by 4.66 and 4.75 percentage points, respectively. These results highlight the value of learning recurrent patient memory for efficient longitudinal EHR modeling.
摘要:長期電子健康紀錄 (EHR) 建模需要將新的就診整合到不斷擴展的病歷中。
然而,不斷累積的臨床資訊對大型語言模型 (LLMs) 在處理和保留完整病歷時帶來了日益增加的計算和記憶成本。
一個實用的替代方案是逐次就診的重複壓縮,這將每次進來的就診納入一個緊湊且持續更新的病人記憶中。
然而,在固定的記憶預算下,連續的更新必須整合新資訊,而不會逐漸失去後續任務所需的關鍵歷史證據。
為了解決這一挑戰,我們提出了重複長期記憶 (ReLMem) 框架,該框架學習維持固定容量的病人記憶,以便使用凍結的 LLM 進行高效的下游預測。
ReLMem 為這個 LLM 配備了輕量級的壓縮適配器,以便從其先前狀態和每次進來的就診中重複更新記憶,而無需重新閱讀早期記錄。
具體而言,我們開發了一種多粒度優化策略,以在重複更新過程中保留與任務相關的資訊,並支持從最終記憶中進行下游預測。
中間監督將來自壓縮記憶和完整歷史的注意力輸出對齊在相同查詢下,而預測監督則最小化與最終記憶條件下的真實答案之間的交叉熵。
在基於 EHR 的藥物預測中,ReLMem 的 F1 分數接近完整歷史基準,同時將平均保留的歷史存儲減少了 97.1%。
在相同的記憶預算下,它分別提高了宏觀和微觀 F1 分數,超過最強壓縮記憶基準 4.66 和 4.75 個百分點。
這些結果突顯了學習重複病人記憶在高效長期 EHR 建模中的價值。
Raw Imagery Impacting Your AI: Should You Care?
2609.38265v1 by Adrien Dorise, Marjorie Bellizzi, Stéphane May
Onboard AI is gaining interest for space applications such as vessel, wildfire, and cloud detection, where real-time processing can improve mission reactivity and reduce downlink needs. However, onboard models may operate on raw or minimally processed imagery rather than on restored ground products. This study evaluates how image degradation affects object detection by varying Signal-to-Noise Ratio (SNR), Modulation Transfer Function (MTF) at Nyquist, and Ground Sampling Distance (GSD). Controlled degradations are applied to Very High Resolution Maxar imagery, and three lightweight detectors, YOLOv5s, YOLOX-S, and NanoDet, are evaluated on the resulting operating points. The results show that the impact of image quality depends on the degradation mechanism, and that increasing degradation does not necessarily lead to a proportional decrease in vessel detection performance. GSD produces the most consistent performance shift, while MTF and SNR effects depend more on the model and resolution. Severe combinations of blur and noise produce the largest losses. These results provide task-level information that can support sensor, processing, and AI trade-offs for future onboard systems.
摘要:在太空應用中,機載人工智慧正受到關注,例如船隻、野火和雲層檢測,其中即時處理可以提高任務反應能力並減少下行鏈路需求。
然而,機載模型可能在原始或經過最小處理的影像上運作,而不是在恢復的地面產品上。
本研究評估影像退化如何影響物體檢測,通過改變信噪比(SNR)、奈奎斯特的調變傳遞函數(MTF)和地面取樣距離(GSD)。
對非常高解析度的Maxar影像施加控制退化,並在結果操作點上評估三個輕量級檢測器,YOLOv5s、YOLOX-S和NanoDet。
結果顯示,影像質量的影響取決於退化機制,且增加退化不一定會導致船隻檢測性能成比例下降。
GSD產生最一致的性能變化,而MTF和SNR的影響則更多地依賴於模型和解析度。
模糊和噪聲的嚴重組合會產生最大的損失。
這些結果提供了任務級別的信息,可以支持未來機載系統的傳感器、處理和人工智慧的權衡。
Do Agent Benchmarks Do What They Say? An Executable-Contract Audit of Tool-Using Agent Environments
2609.37315v1 by Rohith Reddy Bellibatlu, Zichong Wang, Wenbin Zhang
Tool-using agents are entering settings where a wrong action carries real cost, and the benchmarks certifying them grade what each simulated tool call reports having done, assuming the tool did what its interface advertises. The audit taxonomies we survey publish no category for that assumption, and a defect beneath a score is present on every rerun. We treat a tool's advertised surfaces as an executable contract, check the implementation against it, and trace each score's provenance through the task files and evaluator code to the verdicts that derive from state a defective tool should have written. Across 34 audited mutating tools in four benchmarks we confirm seven tool defects and one evaluator property at pinned commits. On injected defects the checker raised no false positive in 25 flags, flagged 2 of 5 negative controls, and missed most: in 29 of 33 scored misses a clause covered the defect but no probe revealed it. The checker's own static half, run alone, flags 14 of 17 confirmed sites, so on these findings the dynamic half confirms and traces rather than discovers. Twelve further AgentDojo tools, with six held-out tools and the seven audited first, complete its 25-tool mutating surface, on which at least 5 tools diverge from their advertised surface as our contracts read it, a rate for AgentDojo alone. No gold trajectory reaches either tau2-bench defect; on 1,120 paths built to isolate the telecom defect, a number fixed by construction, the evaluator rewards a refuel of a suspended line and fails the repaired tool. The clearest case is a clinical benchmark whose tool tells the agent each write executed under a documented no-write design its interface does not disclose; its grader takes that message as evidence, so its action success rate records whether a request carried the expected payload, not whether any record changed.
摘要:使用工具的代理正在進入錯誤行動會帶來實際成本的環境,而認證它們的基準會評分每個模擬工具調用所報告的操作,假設該工具執行了其介面所宣傳的功能。
我們調查的審核分類中沒有針對該假設的類別,並且每次重新運行時都會出現低於分數的缺陷。
我們將工具宣傳的表面視為可執行的合約,檢查其實現是否符合,並追蹤每個分數的來源,通過任務檔案和評估代碼到應該由有缺陷工具寫出的判決。
在四個基準中對34個經審核的變異工具進行的確認中,我們確認了七個工具缺陷和一個評估者屬性,這些都是在固定的提交上進行的。
在注入的缺陷中,檢查器在25個標誌中沒有產生假陽性,標記了5個負控制中的2個,並且大多數都錯過了:在33次得分的錯過中,有29次條款涵蓋了缺陷,但沒有探測揭示它。
檢查器自己的靜態部分,單獨運行時,標記了17個確認位置中的14個,因此根據這些發現,動態部分確認並追蹤,而不是發現。
另外12個AgentDojo工具,與六個保留的工具和最初的七個經審核工具一起,完成了其25個工具的變異表面,其中至少有5個工具與我們合約所讀取的宣傳表面不同,這是AgentDojo單獨的比率。
沒有金色軌跡達到任何tau2-bench缺陷;在1,120條為隔離電信缺陷而構建的路徑中,這個數字是由建設固定的,評估者對一條暫停線的重新加油給予獎勵,卻未能通過修復的工具。
最明顯的案例是一個臨床基準,其工具告訴代理每次寫入在其介面未披露的文檔中執行的無寫入設計下進行;其評分者將該消息視為證據,因此其行動成功率記錄請求是否攜帶預期的有效載荷,而不是任何記錄是否發生變更。
Information Bottleneck-Guided Adaptive Hypergraph Transformer for Brain Disease Diagnosis
2609.37220v1 by Jingxi Feng, Xudong Chen, Yifan Zhang, Heming Xu, Hongcheng Han, Xijing Wang, Dong Zhang, Shaoyi Du
Exploring high-order correlations and long-range dependencies in brain networks holds significant value for both neuroscience research and clinical diagnosis. However, previous studies have lacked a unified integration of high-order and long-range dependency information in brain networks, and there is substantial redundancy behind various types of information. These issues limit their effectiveness in the diagnosis of brain diseases. To address this, we propose an Information Bottleneck-Guided Adaptive HyperGraph Transformer (IBAHGT). By incorporating the information bottleneck (IB) principle, this approach enables adaptive learning of high-order correlations and both short- and long-range dependencies within a unified framework for brain network analysis, achieving high-precision brain disease diagnosis. IBAHGT consists of three key components: an information bottleneck-guided adaptive hypergraph convolution, which introduces a novel hypergraph information bottleneck (HIB) principle to adaptively learn hypergraph message-passing weights between nodes and hyperedges, optimizes information flow and captures high-order information in brain networks that is maximally informative and minimally redundant (MIMR). The Transformer encoder captures global information within brain networks through the attention mechanism, specifically modeling short- and long-range dependencies. An information bottleneck-guided node-level adaptive fusion employs the IB principle to learn independent weights for each node, facilitating the fine-grained integration of high-order information and global information to obtain an efficient representation for downstream tasks. Extensive experiments demonstrate that the proposed method outperforms current state-of-the-art methods and can identify biomarkers for clinical applications.
摘要:探索大腦網絡中的高階相關性和長程依賴性對於神經科學研究和臨床診斷具有重要價值。
然而,先前的研究缺乏對大腦網絡中高階和長程依賴信息的統一整合,並且各類信息之間存在大量冗餘。
這些問題限制了它們在大腦疾病診斷中的有效性。
為了解決這個問題,我們提出了一種信息瓶頸引導的自適應超圖Transformer(IBAHGT)。
通過納入信息瓶頸(IB)原則,這種方法能夠在統一框架內自適應學習高階相關性以及短程和長程依賴性,實現高精度的大腦疾病診斷。
IBAHGT由三個關鍵組件組成:一個信息瓶頸引導的自適應超圖卷積,該組件引入了一種新穎的超圖信息瓶頸(HIB)原則,以自適應地學習節點和超邊之間的超圖信息傳遞權重,優化信息流並捕捉大腦網絡中最大信息量和最小冗餘的高階信息(MIMR)。
Transformer編碼器通過注意機制捕捉大腦網絡中的全局信息,特別是建模短程和長程依賴性。
信息瓶頸引導的節點級自適應融合利用IB原則為每個節點學習獨立權重,促進高階信息和全局信息的細緻整合,以獲得下游任務的高效表示。
廣泛的實驗表明,所提出的方法優於當前的最先進方法,並能夠識別臨床應用的生物標記。
Physics-Informed Multi-Agent Coordination for Hospital Patient Flow Optimization
2609.37022v1 by Guoqing Zhang, Rafik Hadfi, Takayuki Ito
Efficient patient flow coordination across autonomous hospital departments is critical for mitigating overcrowding and balancing resource utilization. While classical queueing theory, specifically open Baskett--Chandy--Muntz--Palacios (BCMP) networks, provides an interpretable mathematical topology for healthcare operations, analytical models rely on stationary assumptions and fixed routing matrices that degrade under state-dependent real-world dynamics. Conversely, centralized reinforcement learning approaches struggle to accommodate the decentralized structure of hospital governance, where individual clinical departments function with localized observations, heterogeneous resources, and divergent operational objectives. In this paper, we present a Multi-Agent Systems (MAS) framework titled \emph{Physics-Informed Multi-Agent Coordination}, which embeds empirically calibrated BCMP queueing topologies as physical priors within a decentralized multi-agent reinforcement learning architecture. Formulated as a Decentralized Partially Observable Markov Decision Process (Dec-POMDP) under coupled resource constraints, our method enables autonomous departmental agents to cooperatively negotiate patient routing and dynamic service scaling. To mitigate environmental non-stationarity without inducing excessive communication overhead, agents exchange localized action fingerprints along network edges and optimize a spatially decomposed reward structure. Empirical evaluations driven by real-world MIMIC-IV patient trajectories indicate that this cooperative multi-agent approach substantially reduces cumulative system delay compared to static Markovian approximations, heuristic dispatching, and independent multi-agent baselines, while maintaining clinical safety constraints.
摘要:有效的病人流動協調在自主醫院部門之間對於減輕擁擠和平衡資源利用至關重要。雖然經典的排隊理論,特別是開放的Baskett--Chandy--Muntz--Palacios (BCMP) 網絡,為醫療運營提供了可解釋的數學拓撲,但分析模型依賴於靜態假設和固定的路由矩陣,這在狀態依賴的現實世界動態下會退化。相反,集中式強化學習方法難以適應醫院治理的去中心化結構,其中各個臨床部門以本地觀察、異質資源和不同的運營目標運作。在本文中,我們提出了一個名為\emph{Physics-Informed Multi-Agent Coordination}的多智能體系統(MAS)框架,該框架將經驗校準的BCMP排隊拓撲作為物理先驗嵌入到去中心化的多智能體強化學習架構中。該方法在耦合資源約束下被表述為去中心化的部分可觀察馬爾可夫決策過程(Dec-POMDP),使自主部門智能體能夠合作協商病人路由和動態服務擴展。為了減輕環境非平穩性而不引入過多的通信開銷,智能體沿著網絡邊緣交換本地化的行動指紋,並優化一個空間分解的獎勵結構。基於現實世界MIMIC-IV病人軌跡的實證評估表明,這種合作的多智能體方法在保持臨床安全約束的同時,顯著減少了累積系統延遲,相較於靜態馬爾可夫近似、啟發式調度和獨立多智能體基準。
STAR-GRPO: Canonical Anchoring and Reliability-First Advantages against Representation-Dependent Reward Hacking
2609.36900v1 by Wan Tian, Zhongyi Li, Xiang Xu, Minhao Zou, Yijie Peng, Fuzhen Zhuang
Reward hacking occurs when policy optimization exploits a brittle reward interface or an overly permissive proxy objective, improving the training score without improving the underlying response quality. This phenomenon is amplified in group-relative policy optimization: an unsupported reward can shift the group baseline and alter the updates of other rollouts, while post-hoc or purely relative weighting cannot represent group-wide uncertainty. We propose \emph{Self-Tuned Anchored Reliability Group-Relative Policy Optimization} (STAR-GRPO), a reliability-first advantage estimator based on paired assessments of the same rollout. STAR separates the quality signal from its learning influence: score disagreement determines rollout reliability, relative reliability enters a self-tuned robust location--scale fit before group normalization, and absolute group reliability attenuates the resulting bounded advantage. The analysis establishes coordinate and second-moment bounds, characterizes exact centering through the weighted location equation, and gives reliability-dependent attenuation guarantees for outlying rewards. We evaluate STAR-GRPO in two complementary reward-hacking regimes. In token-interface exploitation, STAR prevents runaway optimization of the deployed-interface score while improving the canonical quality signal. In rubric-proxy overoptimization for medical reasoning, STAR improves independent semantic evaluation, narrows the proxy--judge discrepancy, and reduces overclaim while optimizing the same task proxy. Together, these results show that reliability-first normalization offers a principled way to limit unsupported reward influence on both group baselines and policy updates, while retaining the task reward as the optimization target.
摘要:獎勵駭客行為發生在政策優化利用脆弱的獎勵介面或過於寬鬆的代理目標時,這樣可以提高訓練分數,但並未改善基礎的反應質量。這一現象在群體相對政策優化中被放大:不受支持的獎勵可以改變群體基準並改變其他回合的更新,而事後或純相對的加權無法代表整個群體的不確定性。我們提出了\emph{自調整錨定可靠性群體相對政策優化}(STAR-GRPO),這是一種基於相同回合配對評估的可靠性優先優勢估計器。STAR將質量信號與其學習影響分開:分數不一致性決定回合的可靠性,相對可靠性進入自調整的穩健位置--尺度擬合,然後進行群體正規化,絕對群體可靠性減弱了結果的有界優勢。分析建立了坐標和二階矩界限,通過加權位置方程描述精確的中心化,並為異常獎勵提供了依賴於可靠性的減弱保證。我們在兩種互補的獎勵駭客制度中評估了STAR-GRPO。在令牌介面利用中,STAR防止了已部署介面分數的失控優化,同時改善了經典質量信號。在醫療推理的評分代理過度優化中,STAR改善了獨立語義評估,縮小了代理--評審之間的差距,並在優化相同任務代理的同時減少了過度聲明。總體而言,這些結果表明,可靠性優先的正規化提供了一種原則性的方法,以限制不受支持的獎勵對群體基準和政策更新的影響,同時保留任務獎勵作為優化目標。
Automated Screw Planning for Reduced Pelvic Fractures Based on Statistical Shape Models and Deep Learning
2609.36847v1 by Yang Gao, Sutuke Yibulayimu, Yanzhen Liu, Zian Zhao, Yudi Sang
Percutaneous iliosacral screw fixation is an important minimally invasive treatment for unstable pelvic fractures. Because the sacroiliac region has complex anatomy and narrow screw corridors, the accuracy and safety of screw placement directly affect surgical outcomes. Accurate and reliable preoperative screw planning is therefore essential to improve surgical success and reduce intraoperative risks. Conventional preoperative planning typically requires surgeons to determine screw trajectories through manual measurements, a labor-intensive process that depends on subjective clinical experience. To address these challenges, we propose a fully automated pipeline for preoperative iliosacral screw planning in patients with pelvic fractures. Using patient-specific three-dimensional anatomy, the pipeline automatically identifies safe screw corridors and generates individualized insertion trajectories to support clinical preoperative planning. We evaluated the proposed pipeline on 200 clinical cases of pelvic fractures. Compared with conventional manual measurements, the safety margin of the safe insertion corridors increased by 2% across the four screw types, the mean planning time decreased by more than 90%, and the clinical acceptance rate reached 95%.
摘要:經皮髖骶螺釘固定是一種重要的微創治療不穩定骨盆骨折的方法。
由於骶髂區域具有複雜的解剖結構和狹窄的螺釘通道,螺釘放置的準確性和安全性直接影響手術結果。
因此,準確可靠的術前螺釘規劃對於提高手術成功率和降低術中風險至關重要。
傳統的術前規劃通常要求外科醫生通過手動測量來確定螺釘的軌跡,這是一個依賴主觀臨床經驗的勞動密集型過程。
為了解決這些挑戰,我們提出了一個完全自動化的術前髖骶螺釘規劃流程,專為骨盆骨折患者設計。
該流程利用患者特定的三維解剖結構,自動識別安全的螺釘通道並生成個性化的插入軌跡,以支持臨床術前規劃。
我們在200例骨盆骨折的臨床案例中評估了所提出的流程。
與傳統手動測量相比,安全插入通道的安全邊際在四種螺釘類型中增加了2%,平均規劃時間減少了90%以上,臨床接受率達到95%。
How Medical VLMs Underutilize Their Vision Encoders: A Dermatology Perspective
2609.36557v1 by Janet Wang, Yunbei Zhang, Xiao Wang, Jihun Hamm
Medical Vision-Language Models (VLMs) show significant promise for clinical image understanding, offering accurate diagnosis with interpretable reasoning. However, a critical performance gap exists between their strong vision encoders and the full multimodal model: in dermatology, the MedSigLIP encoder outperforms MedGemma by an average of 10.26 percentage points even when both use zero target-task labels; few-shot linear probing provides further evidence of strong visual representations. This gap motivates an investigation of how visual information is used in end-to-end diagnosis and why plausible-sounding predictions can lack grounding in image evidence. Using dermatology as our primary testbed, we systematically investigate three hypotheses for this phenomenon. We further provide a mechanistic analysis of the model's internal attention patterns, showing that a simple describe-then-decide prompting strategy increases vision attention by 30-40% during generation. Task-specific fine-tuning improves dermatology classification but reduces cross-domain medical question-answering performance in our evaluation. To address these challenges, we combine label-free prompting with low-label encoder-assisted reranking while keeping the VLM frozen. We validate the interventions across five VLM backbones in dermatology and provide supporting representation and attention analyses across additional medical modalities.
摘要:醫療視覺-語言模型(VLMs)在臨床影像理解方面顯示出顯著的潛力,提供可解釋的推理以進行準確診斷。
然而,它們強大的視覺編碼器與完整的多模態模型之間存在一個關鍵的性能差距:在皮膚科,MedSigLIP 編碼器的表現平均超過 MedGemma 10.26 個百分點,即使兩者都使用零目標任務標籤;少量樣本線性探測進一步證明了強大的視覺表徵。
這一差距促使我們調查視覺信息在端到端診斷中的使用方式,以及為什麼聽起來合理的預測可能缺乏影像證據的支持。
以皮膚科作為我們的主要測試平台,我們系統地調查了這一現象的三個假設。
我們進一步提供了模型內部注意力模式的機制分析,顯示簡單的描述-再決策提示策略在生成過程中將視覺注意力提高了 30-40%。
特定任務的微調改善了皮膚科分類,但在我們的評估中降低了跨領域醫療問答的表現。
為了解決這些挑戰,我們結合無標籤提示與低標籤編碼器輔助的重新排序,同時保持 VLM 凍結。
我們在皮膚科的五個 VLM 骨幹上驗證了這些干預,並提供了支持的表徵和注意力分析,涵蓋其他醫療模態。
Reliability Testing of Medical Model Performance under Distributed Deployment
2609.36525v1 by Yifei Wang, Xiaohan Zhang, Youtao Ding, Tianlin Li, Xiaoyu Zhang, Yida Yang, Li Pan
Distributed inference has become an indispensable part of deploying medical models under practical latency, memory, and throughput constraints. Although modern frameworks improve serving efficiency through tensor parallelism, mixed precision, kernel fusion, and multi-device communication, they are generally assumed to preserve the behavior observed during centralized HuggingFace evaluation. This assumption creates an evaluation-deployment mismatch: a model may pass offline evaluation but produce a different output after the execution stack changes. To address this mismatch, we propose a testing framework and an improved, distributed-execution-sensitive medical-model benchmark that evaluates the same checkpoint and input under a centralized HuggingFace reference and matched distributed deployments. Extensive experiments across language, vision, and multimodal medical models show that execution changes can produce measurable output disagreements. Across supported visual settings, the test success rate ranges from 0.21 to 0.43 for single-modality models and from 0.32 to 0.98 for multimodal models. The benchmark is aimed at extending medical-model evaluation from capability and security to evaluation-deployment consistency.
摘要:分散推理已成為在實際延遲、記憶體和吞吐量限制下部署醫療模型不可或缺的一部分。儘管現代框架通過張量並行、混合精度、內核融合和多設備通信來提高服務效率,但通常假設它們能保持在集中式 HuggingFace 評估中觀察到的行為。這一假設造成了評估與部署之間的不匹配:一個模型可能在離線評估中通過,但在執行堆棧變更後產生不同的輸出。為了解決這一不匹配,我們提出了一個測試框架和一個改進的、對分散執行敏感的醫療模型基準,該基準在集中式 HuggingFace 參考和匹配的分散部署下評估相同的檢查點和輸入。針對語言、視覺和多模態醫療模型的廣泛實驗顯示,執行變更可能會產生可測量的輸出不一致。在支持的視覺設置中,單模態模型的測試成功率範圍為 0.21 到 0.43,而多模態模型的範圍為 0.32 到 0.98。該基準旨在將醫療模型評估從能力和安全性擴展到評估與部署的一致性。
BRIDGE: Bilevel Retrieval-Credit-Aware Agentic Reinforcement Learning
2609.36505v1 by Quan Xiao, Mingda Liu, Gaowen Liu, Katsuki Fujisawa, Tianyi Chen
Agentic reinforcement learning (ARL) with verifiable rewards improves the ability of large language models (LLMs) to tackle knowledge-intensive tasks by learning to interleave search and reasoning. However, most existing ARL methods optimize only LLM-generated tokens and treat retrieved evidence as environment observations. This creates an information-credit gap: failures caused by missing or misleading evidence are attributed to the LLM policy rather than to the retriever, which motivates training the LLM and the retriever jointly. In this paper, we show that retrieval and LLM policy learning are order-sensitive: adapting the retriever before optimizing the policy yields a larger reward gain than the reverse order. To preserve this hierarchy while allowing both components to co-adapt, we formulate retrieval-augmented agentic RL as a bilevel optimization problem. To solve it efficiently, we introduce BRIDGE, a memory-efficient first-order bilevel method motivated by a loss-landscape analysis of the RL and retrieval objectives. Across seven open-domain QA benchmarks, BRIDGE achieves the highest average accuracy with both 3B and 7B backbones, improving the multi-hop average over the strongest baseline by 9.6 and 3.4 EM points, respectively. It also achieves the best averaged answer accuracy and reasoning quality across medical QA benchmarks.
摘要:代理強化學習(ARL)與可驗證獎勵相結合,通過學習交替進行搜索和推理,提高了大型語言模型(LLMs)處理知識密集型任務的能力。
然而,大多數現有的ARL方法僅優化LLM生成的標記,並將檢索到的證據視為環境觀察。
這造成了一個信息信用差距:由於缺失或誤導性證據而導致的失敗被歸因於LLM策略,而不是檢索器,這促使了LLM和檢索器的聯合訓練。
在本文中,我們展示了檢索和LLM策略學習對順序的敏感性:在優化策略之前調整檢索器,會比反向順序產生更大的獎勵增益。
為了在允許兩個組件共同適應的同時保留這一層次結構,我們將檢索增強的代理RL公式化為一個雙層優化問題。
為了高效解決它,我們引入了BRIDGE,一種基於RL和檢索目標的損失景觀分析的記憶高效的一階雙層方法。
在七個開放域問答基準上,BRIDGE在3B和7B骨幹上都達到了最高的平均準確率,分別提高了對最強基線的多跳平均9.6和3.4 EM點。
它還在醫學問答基準上實現了最佳的平均答案準確率和推理質量。
ARCagent: An Adaptive Retrieval Calibration Agent for Clinical Question Answering
2609.36392v1 by Yuyan Chen
In diseases where clinical guidelines are incomplete, contested, or mutually contradictory, knowledge completeness and dynamic conflict-aware synthesis are two safety-critical properties that standard Retrieval-Augmented Generation systems do not provide. Therefore, we present \sysname, an adaptive retrieval calibration clinical question-answering agent for ME/CFS, a disease where diagnostic frameworks coexist and major guidelines actively contradict each other on treatment. ARCagent contributes three components. First, a 1,706-chunk, 10-source knowledge base with a structured inter-guideline conflict registry spanning all active ME/CFS diagnostic frameworks. Second, a conflict-aware retrieval calibration pipeline that re-ranks retrieved evidence using query-specific focus and conflict signals. Third, a benchmark scored by LLM-as-Judge, avoiding systematic underestimation averaging 10.1 percentage points caused by keyword matching. ARCagent achieves 95.3%, outperforming all base LLMs. Code is available at https://github.com/Yukyin/ARCagent.
摘要:在臨床指導方針不完整、存在爭議或相互矛盾的疾病中,知識的完整性和動態衝突意識合成是標準增強檢索生成系統所不具備的兩個安全關鍵特性。
因此,我們提出了\sysname,一個針對ME/CFS的自適應檢索校準臨床問答代理,這是一種診斷框架共存且主要指導方針在治療上積極矛盾的疾病。
ARCagent貢獻了三個組件。
首先,一個包含1,706個片段、10個來源的知識庫,擁有一個結構化的指導方針間衝突登記,涵蓋所有活躍的ME/CFS診斷框架。
其次,一個衝突意識的檢索校準管道,使用查詢特定的焦點和衝突信號重新排名檢索到的證據。
第三,一個由LLM-as-Judge評分的基準,避免了由關鍵字匹配造成的系統性低估,平均低10.1個百分點。
ARCagent達到了95.3%的準確率,超越了所有基礎LLM。
代碼可在https://github.com/Yukyin/ARCagent獲得。
Quantization Enables Private Dense Retrieval against Malicious Service Providers
2609.36376v1 by Louis Tremblay Thibault, Sofiane Azogagh, Marc-Olivier Killijian, Ulrich Aïvodji
Dense retrieval, the key component of Retrieval Augmented Generation (RAG), retrieves the most relevant documents by comparing dense vector representations of queries and passages from a large corpus. In privacy-sensitive applications, the server observes the query and controls which evidence is returned, creating both confidentiality and integrity risks. We formulate private dense retrieval as providing query privacy and retrieval integrity against a malicious server, and develop a two-round cryptographic protocol that provides both guarantees. Our protocol reduces private and verifiable retrieval to multiplication of a committed matrix by an encrypted vector and uses low-bit quantization to make this computation practical. We evaluate the resulting trade-off between cryptographic cost, retrieval quality, and downstream RAG accuracy across six embedding models, four language models, and corpora of up to 2.68 million passages. Our results show that, with a clipped quantizer, three-bit quantization largely preserves retrieval quality and downstream accuracy, while a private query over a corpus the size of a clinical reference requires one to three minutes of server time. These results suggest that private dense retrieval is already practical for moderately sized, privacy-sensitive corpora when minute-scale latency is acceptable.
摘要:密集檢索,檢索增強生成(RAG)的關鍵組件,通過比較查詢和來自大型語料庫的段落的密集向量表示來檢索最相關的文檔。在隱私敏感的應用中,伺服器觀察查詢並控制返回哪些證據,從而產生保密性和完整性風險。我們將私密密集檢索定義為在惡意伺服器面前提供查詢隱私和檢索完整性,並開發了一種雙輪加密協議,提供這兩項保證。我們的協議將私密且可驗證的檢索簡化為將已承諾的矩陣與加密向量相乘,並使用低位量化使這一計算變得實用。我們評估了六種嵌入模型、四種語言模型和多達268萬段落的語料庫中,加密成本、檢索質量和下游RAG準確性之間的權衡。我們的結果顯示,使用剪裁量化器的三位量化在很大程度上保持了檢索質量和下游準確性,而對於大小相當於臨床參考的語料庫,私密查詢需要一到三分鐘的伺服器時間。這些結果表明,當分鐘級延遲是可以接受的時候,私密密集檢索對於中等大小的隱私敏感語料庫已經是實用的。
ThuRunel: Dynamic Decoupling for Structured Advisory Dialogue
2609.36340v1 by Yuyan Chen
High-stakes advisory domains such as medical aesthetics, legal consultation, and educational planning exhibit a two-phase structure. The early phase requires empathetic elicitation and emotional support, and the late phase requires authoritative specialist judgment. Neither fully automated agents nor human junior consultants adequately address this structure at scale. We formalize the core design challenge as dynamic decoupling, asking how an AI advisory agent should decide what to ask, when to stop, what to resolve autonomously, and what to forward to the specialist. We present ThuRunel, an advisory agent combining a finite-state belief management framework, a chain-of-thought teacher synthesis protocol, and learned generation adapters. Against eleven baselines, ThuRunel achieves consistent improvements in elicitation completeness and specialist brief quality. ThuRunel is publicly deployed as a bilingual web application in which the same decoupling decisions operate from the client's side, grounded in a curated knowledge base that cites its sources in every answer.
摘要:高風險的諮詢領域如醫療美學、法律諮詢和教育規劃展示出雙階段結構。
早期階段需要同理心的引導和情感支持,而後期階段則需要權威專家的判斷。
無論是完全自動化的代理還是人類初級顧問,都無法在規模上充分應對這一結構。
我們將核心設計挑戰形式化為動態解耦,詢問AI諮詢代理應如何決定詢問什麼、何時停止、什麼可以自主解決以及什麼需要轉交給專家。
我們提出了ThuRunel,一個結合有限狀態信念管理框架、思維鏈教師綜合協議和學習生成適配器的諮詢代理。
在十一個基準測試中,ThuRunel在引導完整性和專家簡報質量上實現了一致的改進。
ThuRunel作為一個雙語網絡應用程序公開部署,其中相同的解耦決策從客戶端運作,基於一個策劃的知識庫,並在每個答案中引用其來源。
SCOUT: Synergizing Reasoning and Tool-Use for Computer-Use Safety
2609.36201v1 by Jianxing Chen, Xiao Yu, Shipra Agrawal, Zhou Yu
Computer-use agents (CUAs), while capable of completing computer tasks in everyday and professional workflows, can cause unintended harm even under benign instructions and environments. However, detecting such harm remains challenging. First, it requires careful, task-specific reasoning: verifiers guided only by general safety criteria often overlook many important but subtle harmful behaviors. Second, it requires active investigation: past trajectory screenshots show what the agent did but not always what actually changed in the environment, so LLM-as-a-judge verifiers that rely on screenshots alone may be unable to determine the actual consequences of actions. To address these challenges, we introduce SCOUT, a two-stage agentic safety verifier that synergizes reasoning-intensive rubric generation with tool-intensive evidence gathering. First, our SCOUT rubric generator extensively reasons over the task and the agent's trajectory to determine what successful and safe execution should entail, generating task-specific completion and safety rubrics. Then, our SCOUT probing agent follows these rubrics to interact with the post-execution environment and collect grounded evidence for final safety and completion judgments. We evaluate our framework on two computer-use safety benchmarks. On AutoElicit-Bench, SCOUT achieves 75.4 unsafe F1 and 74.5 completion F1, outperforming LLM-as-a-judge verifiers and naive tool-use verifiers. SCOUT leads on OS-Blind with 76.4% unsafe detection accuracy. Test-time reflection reduces final unsafe execution rates from 30.2% to 17.2% on AutoElicit-Bench. Ablations and analysis show that tool-free rubric generation in SCOUT elicits substantially more reasoning and is crucial for safety detection across verifier backbones, especially non-frontier ones. A preliminary extension to coding tasks shows that SCOUT can support safety verification beyond computer-use.
摘要:電腦使用代理(CUAs)雖然能夠在日常和專業工作流程中完成電腦任務,但即使在良性指令和環境下,也可能造成意想不到的傷害。
然而,檢測這種傷害仍然具有挑戰性。
首先,它需要仔細的、特定任務的推理:僅依賴一般安全標準的驗證者往往忽視許多重要但微妙的有害行為。
其次,它需要主動調查:過去的軌跡截圖顯示代理所做的事情,但不總是顯示環境中實際發生的變化,因此僅依賴截圖的LLM作為判斷者的驗證者可能無法確定行動的實際後果。
為了解決這些挑戰,我們引入了SCOUT,一種兩階段的代理安全驗證器,將推理密集的標準生成與工具密集的證據收集相結合。
首先,我們的SCOUT標準生成器對任務和代理的軌跡進行廣泛推理,以確定成功和安全執行應該包含什麼,生成特定任務的完成和安全標準。
然後,我們的SCOUT探測代理根據這些標準與執行後的環境互動,並收集基於證據的最終安全和完成判斷。
我們在兩個電腦使用安全基準上評估了我們的框架。
在AutoElicit-Bench上,SCOUT達到了75.4的危險F1和74.5的完成F1,超越了LLM作為判斷者的驗證者和天真的工具使用驗證者。
SCOUT在OS-Blind上以76.4%的危險檢測準確率領先。
測試時反思將AutoElicit-Bench上的最終危險執行率從30.2%降低到17.2%。
消融和分析顯示,SCOUT中的無工具標準生成引發了顯著更多的推理,並且對於各種驗證者骨幹,尤其是非前沿的驗證者,對於安全檢測至關重要。
對編碼任務的初步擴展顯示,SCOUT可以支持超越電腦使用的安全驗證。
PHASE: A Physiology-Guided Hierarchical Foundation Model for Intracranial EEG
2609.36087v2 by Yipeng Zhang, Chenda Duan, Yuanyi Ding, Tianyi Wang, Atsuro Daida, Masaki Izumi, Yuta Tanoue, Naoto Kuroda, Shaun A. Hussain, Nishant Sinha, Eishi Asano, Hiroki Nariai, Vwani Roychowdhury
Clinicians and neuroscientists have long analyzed intracranial electroencephalography (iEEG) through directly measurable physiological characteristics, which carry much of the information that downstream tasks depend on. Recent iEEG foundation models learn by reconstructing or predicting their inputs, which leaves the retention of these characteristics implicit. They are also evaluated mainly on cognitive decoding and a narrow clinical task, i.e., seizure detection. On a broad, clinically relevant benchmark such as Omni-iEEG, they remain below task-specific models when used frozen. We introduce PHASE, a physiology-guided foundation model that makes these characteristics explicit learning targets, pairing them with masked latent prediction in a temporal stage (PHASE-T) within each channel and a spatiotemporal stage (PHASE-ST) across synchronized channels. PHASE is pretrained on heterogeneous recordings from 222 participants at nine clinical sites. On all five Omni-iEEG clinical tasks, frozen PHASE-T outperforms every evaluated foundation model by up to 31\%, and fine-tuned PHASE-T surpasses the task-specific models, setting a new state of the art. PHASE-T benefits from physiological supervision, outperforming variants trained with latent prediction alone or auxiliary waveform reconstruction on every task in matched ablations. PHASE-T generalizes to unseen institutions, outperforming the compared models with few or no local labels. PHASE-ST further improves seizure-onset-zone identification over PHASE-T and, when frozen, decodes sound volume and pitch on BrainTreebank better than published models. Beyond task performance, PHASE learns to encapsulate the physiological characteristics clinicians recognize, from seizure onset and its propagation to anatomical region identity, even though its pretraining contains no ictal recordings or anatomical labels.
摘要:臨床醫生和神經科學家長期以來一直通過直接可測量的生理特徵分析顱內腦電圖 (iEEG),這些特徵攜帶了許多下游任務所依賴的信息。最近的 iEEG 基礎模型通過重建或預測其輸入來學習,這使得這些特徵的保留變得隱含。它們的評估主要集中在認知解碼和一個狹窄的臨床任務,即癲癇發作檢測。在像 Omni-iEEG 這樣的廣泛臨床相關基準上,當以凍結狀態使用時,它們的表現仍低於特定任務模型。我們介紹了 PHASE,一種生理引導的基礎模型,將這些特徵作為明確的學習目標,並在每個通道的時間階段 (PHASE-T) 和同步通道之間的空間時間階段 (PHASE-ST) 中將其與掩蔽潛在預測配對。PHASE 在九個臨床站點的 222 名參與者的異質錄音上進行了預訓練。在所有五個 Omni-iEEG 臨床任務中,凍結的 PHASE-T 的表現超過了每個評估的基礎模型,最高可達 31\%,而微調的 PHASE-T 超越了特定任務模型,創造了新的最先進水平。PHASE-T 受益於生理監督,在每個匹配的消融實驗中,超越了僅用潛在預測或輔助波形重建訓練的變體。PHASE-T 在未見過的機構中具有良好的泛化能力,超越了比較模型,並且幾乎沒有本地標籤。PHASE-ST 進一步改善了癲癇發作區域的識別,超過了 PHASE-T,並且在凍結狀態下,對 BrainTreebank 的聲音音量和音高的解碼表現優於已發表的模型。除了任務性能外,PHASE 學會了封裝臨床醫生識別的生理特徵,從癲癇發作及其傳播到解剖區域身份,即使其預訓練中不包含任何癲癇發作錄音或解剖標籤。
IMC-CLINIC: Coupled Loss-Informed Newton Iterations for Clipping in Analog In-Memory Computing
2609.35586v1 by Yung-Chin Chen, Chia-Yu Chen, Naveen Verma
Analog in-memory computing (IMC) offers a promising path toward energy-efficient large language model (LLM) inference by executing matrix multiplications (MatMul) directly within memory arrays in the analog domain. Its efficiency, however, comes with an additional source of error: limited-precision analog-to-digital converters (ADCs) quantize accumulated analog partial sums, introducing output-side error distinct from conventional activation and weight quantization at the MatMul inputs. Clipping can mitigate both operand and ADC quantization errors, but the optimal clipping factors must jointly balance activation rounding and clipping, weight rounding and clipping, and ADC quantization. Existing clipping methods, designed for digital quantization, do not explicitly optimize these coupled sources of IMC error and often rely on costly search-based calibration. We introduce IMC-CLINIC (Coupled Loss-Informed Newton Iterations for Clipping), a clipping calibration framework based on an analytical surrogate for IMC MatMul output error. The surrogate jointly models operand quantization, accumulated clipping-induced bias, and ADC quantization, enabling efficient evaluation of its gradient and approximate curvature from a small calibration set. IMC-CLINIC jointly optimizes activation and weight clipping factors using a safeguarded Newton-type method. Across multiple models and datasets, it improves average zero-shot accuracy by 6.5-11.5 percentage points over the grid search baseline while reducing calibration time by factors of 10.0-12.1. Its analytical surrogate closely tracks empirical IMC output error, and its optimizer is certified within 1% of the global optimum under the loss objective across all projections on two representative models.
摘要:類比記憶體計算(IMC)提供了一條有前景的途徑,以實現能效高的大型語言模型(LLM)推斷,通過在類比領域內直接在記憶體陣列中執行矩陣乘法(MatMul)。然而,它的效率伴隨著一個額外的誤差來源:有限精度的類比轉數字轉換器(ADC)對累積的類比部分和進行量化,這引入了與傳統的激活和權重量化在MatMul輸入時不同的輸出端誤差。剪裁可以減輕操作數和ADC量化誤差,但最佳剪裁因子必須共同平衡激活四捨五入和剪裁、權重四捨五入和剪裁,以及ADC量化。現有的剪裁方法旨在數位量化,並未明確優化這些耦合的IMC誤差來源,並且通常依賴於昂貴的基於搜索的校準。我們介紹IMC-CLINIC(耦合損失信息牛頓迭代剪裁),這是一個基於IMC MatMul輸出誤差的分析替代品的剪裁校準框架。該替代品共同建模操作數量化、累積的剪裁引起的偏差和ADC量化,從而能夠高效地評估其梯度和從小型校準集獲得的近似曲率。IMC-CLINIC使用受保護的牛頓類型方法共同優化激活和權重剪裁因子。在多個模型和數據集上,它將平均零-shot準確率提高了6.5-11.5個百分點,相較於網格搜索基準,並將校準時間減少了10.0-12.1倍。其分析替代品與實證IMC輸出誤差密切相關,並且其優化器在所有兩個代表性模型的損失目標下,經過所有投影後,證明在全球最優解的1%內。
RareDx: Controlled Knowledge Integration and Graph-Grounded Policy Optimization for Rare-Disease Diagnosis
2609.35549v2 by Bo Zhang, Yuchen Wang, Dongbai Li, Matthew Yu Heng Wong, Qingkai Zeng, Lijun Wang, Tien-Yin Wong, Peng Cui, Tianyu Liu
Rare-disease diagnosis is a long-tail reasoning problem: phenotypes are incomplete, individual disorders are sparsely documented, and relevant evidence is distributed across ontologies, gene annotations, and biomedical text. Language models consequently favor common conditions, miss rare candidates, or produce plausible but invalid names. We introduce RareDx, which couples controlled evidence use with knowledge-graph-grounded policy optimization. RareDx-Harness normalizes heterogeneous records into one ranked-diagnosis task and compares direct inference, static retrieval, adaptive tools, and structured phenotype-gene-disease reasoning over a shared knowledge layer. The training pipeline combines Top-10 post-training with RareDx-KGPO, our knowledge-graph-grounded policy optimization method. Its reward projects predictions into a canonical disease graph and integrates curated graded relevance, ontology proximity, biomedical similarity, and phenotype consistency. Vocabulary and output-budget constraints prevent dense partial credit from rewarding fabricated or overlong differentials. Across eight benchmarks, the complete RareDx system centered on Qwen3.5-9B reaches 38.34 macro Hit@10, 1.60 points above GPT-5.5 under the archived protocol; a disjoint validation-selection audit retains a 6.80-point routing gain over Direct on held-out cases. The 27B system reaches 23.53/36.56/40.76 at Hit@1/5/10. Controlled ablations show that retrieval is not uniformly helpful and that controlled routing is central to the gain. These results indicate that structured medical knowledge can turn a compact model into a competitive diagnostic ranker across heterogeneous long-tail settings in clinical practice.
摘要:罕見疾病的診斷是一個長尾推理問題:表型不完整,個別疾病的文獻記錄稀少,相關證據分散在本體論、基因註釋和生物醫學文本中。因此,語言模型偏向於常見病症,錯過罕見候選者,或產生看似合理但無效的名稱。我們介紹了RareDx,它將受控證據使用與基於知識圖的政策優化結合起來。RareDx-Harness將異質記錄標準化為一個排名診斷任務,並比較直接推理、靜態檢索、自適應工具和結構化表型-基因-疾病推理,這些都基於共享的知識層。訓練流程結合了Top-10後訓練與RareDx-KGPO,我們的基於知識圖的政策優化方法。其獎勵將預測投射到一個典範疾病圖中,並整合了策劃的分級相關性、本體接近性、生物醫學相似性和表型一致性。詞彙和輸出預算限制防止密集部分信用獎勵虛構或過長的差異。在八個基準測試中,完整的RareDx系統以Qwen3.5-9B為中心,達到38.34的宏觀Hit@10,比GPT-5.5在存檔協議下高出1.60分;一個不重疊的驗證選擇審計在保留案例中保持了比Direct高出6.80分的路由增益。27B系統在Hit@1/5/10上分別達到23.53/36.56/40.76。受控消融實驗顯示檢索並不總是有幫助,受控路由對增益至關重要。這些結果表明,結構化的醫學知識可以將一個緊湊的模型轉變為在臨床實踐中跨異質長尾環境的競爭性診斷排名器。
CLIMB: A Clinical Multimorbidity Benchmark for Diagnosing Co-occurring Conditions through Multiturn Conversations
2609.35462v1 by Yusuf Kesmen, Aniruddha Mukherjee, Yena Chang, David Sasu, Trevor Brokowski, Alexandra V. Kulinkina, Kristina Keitel, Akhil Arora, Lars Henning Klein, Mary-Anne Hartley
Patients often have several co-occurring clinical conditions, and the findings needed to identify and disambiguate them emerge over the course of a consultation. Evaluating clinical reasoning in this setting requires both multi-turn interaction and multi-label diagnosis. We introduce CLIMB, a benchmark in which a doctor model interviews a simulated patient to recover a ground truth set of co-occurring clinical conditions. Cases are synthesized from clinical decision algorithms and diagnostic datasets, grounding multimorbid presentations in structured clinical knowledge. Across six frontier and open models, none recovers the exact set of conditions in more than 10% of interactive cases. Diagnostic performance declines when conditions co-occur, even when models receive the full clinical record and the true number of conditions. Interaction reduces performance further. In controlled experiments, models behave like single-hypothesis trackers: they anchor on the diagnosis suggested by the opening findings, keep questioning around it, and recover a second condition mainly when a finding in view points to it. Questioning them further does not complete the set but adds mostly wrong diagnoses. We formalise this pattern with a theoretical reference model of single-hypothesis tracking. The benchmark, generator, and evaluation code are available at https://anonymous.4open.science/r/CLIMB-8340.
摘要:患者經常有多種共病臨床狀況,而識別和釐清這些狀況所需的發現會在諮詢過程中出現。
在這種情境下評估臨床推理需要多輪互動和多標籤診斷。
我們介紹了CLIMB,一個基準,其中醫生模型對模擬患者進行訪談,以恢復一組共病臨床狀況的真實基準。
案例是從臨床決策算法和診斷數據集中合成的,將多重共病表現根植於結構化的臨床知識中。
在六個前沿和開放模型中,沒有一個能在超過10%的互動案例中恢復出確切的狀況集。
當狀況共存時,診斷表現會下降,即使模型獲得了完整的臨床記錄和真實的狀況數量。
互動進一步降低了表現。
在受控實驗中,模型的行為類似於單假設追蹤器:它們依賴於開頭發現所建議的診斷,圍繞此進行持續提問,並主要在某個發現指向它時恢復第二個狀況。
進一步詢問並未完成該集合,而是主要增加了錯誤的診斷。
我們用單假設追蹤的理論參考模型來形式化這一模式。
基準、生成器和評估代碼可在 https://anonymous.4open.science/r/CLIMB-8340 獲得。
A decision-support system applied to Law: Reasoning and explainability of the decision
2609.35370v1 by Jeremy Bouche-Pillon, Pascale Zarat{é}, Yannick Chevalier, Nathalie Aussenac-Gilles
The emergence of the digital transition brought an increasing need to control the processing of digital information, including in Law Enforcement Agencies (LEAs). At the EU level, in recent years, many regulations have emerged to control data processing and exchange. Texts other than the GDPR, such as the ''Law Enforcement Directive (LED)'', appeared to regulate specifically how Law Enforcement Agencies (LEAs) could process data. A formal representation of these regulations can be part of decision systems that support LEAs in processing data in compliance with the regulations. Although many new formalisms have emerged to represent legal norms and rules, few are provided with a reasoning mechanism. Furthermore, systems used in decision-making processes in critical contexts such as medical diagnoses or legal decisions cannot be fully automated, and the explainability of their results is essential to ensure user confidence in decisions. This explainability aspect, while crucial, is lacking in most modern approaches that rely on machine learning. This paper describes a framework to operate formal rules from regulations, by focusing on explainability of the decision. After describing the general architecture of the proposed decision support framework, the paper showcases how symbolic AI and the SPARQL query language can support legal reasoning. It then describes an algorithm to generate a justification for the reasoning results, and outlines the procedure to be followed when the reasoning does not lead to a satisfactory conclusion. We notably focus on a method based on decision trees to determine what additional information to request from the user.
摘要:數位轉型的出現帶來了對數位資訊處理的日益需求,包括在執法機構(LEAs)中。在歐盟層面上,近年來出現了許多規範來控制數據處理和交換。除了GDPR之外,還出現了如“執法指令(LED)”等文本,專門規範執法機構(LEAs)如何處理數據。這些規範的正式表述可以成為支持執法機構在遵守規範的情況下處理數據的決策系統的一部分。儘管許多新的形式主義已經出現以表達法律規範和規則,但很少有配備推理機制的形式。此外,用於醫療診斷或法律決策等關鍵情境的決策過程中使用的系統不能完全自動化,其結果的可解釋性對於確保用戶對決策的信心至關重要。這一可解釋性方面雖然至關重要,但在大多數依賴機器學習的現代方法中卻缺乏。本文描述了一個運作規範形式規則的框架,重點在於決策的可解釋性。在描述所提議的決策支持框架的一般架構後,本文展示了符號人工智慧和SPARQL查詢語言如何支持法律推理。接著描述了一種生成推理結果的理由的算法,並概述了當推理未能導致令人滿意的結論時應遵循的程序。我們特別關注一種基於決策樹的方法,以確定需要向用戶請求的額外信息。
Training-Free Clinical Reasoning through Medical Ontologies and Cognitive Mapping: A Symbolic-Probabilistic Knowledge Graph Framework
2609.35298v1 by Surajit Das
Most clinical prediction systems learn patient-variable-outcome associations; we investigate a training-free diagnostic paradigm mapping patient observations to explicit medical knowledge. CKG Reasoner integrates candidate-specific Evidence Feature Nodes, patient-reference matching, a bounded Information Gate, knowledge-weighted evidence accumulation, disease similarity, and decisive clinical rules. Missing-aware normalization and coverage auditing distinguish absent from unavailable evidence. Candidate ranking is separate from outcome-label-independent K-means clustering, which uses four derived evidence coordinates (evidence strength, relative magnitude, directional similarity, and evidence completeness), not raw predictors or targets, to derive cohort-level assignments. Across six retrospective cohorts - four dengue (N = 1000, 1523, 989, 1018), malaria (N = 2190), and influenza (N = 4569) - a uniform, label-free, cohort-fitted K = 2 protocol yielded positive-class F1 scores of 0.996, 0.634, 0.936, 0.917, 0.695, and 0.842, and all-record accuracies of 0.996, 0.558, 0.914, 0.893, 0.707, and 0.906, respectively, with full partition-decision coverage using the frozen package and disease-specific knowledge representations. Neither scoring nor clustering uses outcome labels. Logistic regression provides a supervised baseline. Influenza incorporates confirmatory molecular PCR and is not independent pre-test prediction. Results characterize knowledge-grounded evidence separation, auditability, and sensitivity, not prospective clinical validity or comparative superiority. FOL/LLM-based clinical explanation remains unevaluated.
摘要:大多數臨床預測系統學習患者變數與結果之間的關聯;我們探討一種無需訓練的診斷範式,將患者觀察映射到明確的醫學知識。CKG Reasoner整合了候選特定的證據特徵節點、患者參考匹配、一個有界的信息閘、知識加權的證據累積、疾病相似性和決策臨床規則。缺失感知正規化和覆蓋審核將缺失證據與不可用證據區分開來。候選排名與結果標籤獨立的K均值聚類分開進行,該聚類使用四個衍生的證據坐標(證據強度、相對大小、方向相似性和證據完整性),而不是原始預測因子或目標,來導出隊列級別的分配。在六個回顧性隊列中——四個登革熱(N = 1000, 1523, 989, 1018)、瘧疾(N = 2190)和流感(N = 4569)——一個統一的無標籤、適合隊列的K = 2協議產生了正類F1分數分別為0.996、0.634、0.936、0.917、0.695和0.842,以及所有記錄的準確率分別為0.996、0.558、0.914、0.893、0.707和0.906,並且使用冷凍包和特定疾病知識表示達成了完全的分區決策覆蓋。無論是評分還是聚類都不使用結果標籤。邏輯回歸提供了一個監督的基準。流感結合了確認性分子PCR,並且不是獨立的預測測試。結果特徵化了以知識為基礎的證據分離、可審核性和敏感性,而不是前瞻性臨床有效性或比較優越性。基於FOL/LLM的臨床解釋仍未被評估。
CarveMix-RC: Addressing Rare-Class Imbalance Through Lesion-Aware Synthetic Augmentation for Brain Metastasis Segmentation
2609.35195v1 by Md Shibly Sadique, Md Fayaz Bin Hossen, Michael L. Evans, Walia Farzana, Asfaqur Rahman, Ahmed Temtam, Khan M. Iftekharuddin
Accurate segmentation of post-treatment brain metastases is essential for treatment planning, longitudinal disease monitoring, and quantitative assessment of therapeutic response. The BraTS-MET 2026 Task 1 challenge introduces a clinically relevant segmentation problem involving four anatomically distinct tumor subregions: non-enhancing tumor core (NETC), surrounding non-enhancing FLAIR hyperintensity (SNFH), enhancing tumor (ET), and the resection cavity (RC). Among these, RC segmentation is particularly challenging because of its low prevalence, heterogeneous postoperative appearance, and lesion-wise evaluation protocol, leading conventional segmentation networks to prioritize dominant tumor classes during optimization. The proposed nnU-Net-based framework explicitly addresses RC segmentation through four complementary components: (i) RC-weighted Dice and Cross-Entropy optimization to alleviate class imbalance, (ii) anatomically consistent cavity augmentation to increase the diversity of postoperative cavity appearances, (iii) a residual encoder architecture for enhanced multi-scale feature learning, and (iv) lesion-aware morphological post-processing to suppress false-positive cavity predictions while preserving anatomically plausible structures. The framework is evaluated on the BraTS-MET 2026 Task 1 online validation benchmark. Among the evaluated configurations, the ensemble model (Residual Encoder nnU-Net + nnU-Net + RC-aware CarveMix) achieves the best performance, with lesion-wise Dice scores of 0.732, 0.752, 0.708, and 0.575 and corresponding NSD scores of 0.794, 0.798, 0.727, and 0.474 for ET, TC, WT, and RC, respectively. These experimental results show that integrating RC-aware optimization, anatomically consistent augmentation, and lesion-aware post-processing provides an effective strategy for improving rare resection cavity segmentation in post-treatment brain metastases.
摘要:準確的術後腦轉移瘤分割對於治療計劃、長期疾病監測和治療反應的定量評估至關重要。BraTS-MET 2026 任務 1 挑戰引入了一個臨床相關的分割問題,涉及四個解剖上不同的腫瘤子區域:非增強腫瘤核心 (NETC)、周圍非增強 FLAIR 高信號 (SNFH)、增強腫瘤 (ET) 和切除腔 (RC)。在這些區域中,RC 分割特別具有挑戰性,因為其低發生率、異質的術後外觀以及病灶評估協議,導致傳統的分割網絡在優化過程中優先考慮主要腫瘤類別。所提出的基於 nnU-Net 的框架通過四個互補組件明確解決 RC 分割問題:(i) RC 加權的 Dice 和交叉熵優化以緩解類別不平衡,(ii) 解剖一致的腔體增強以增加術後腔體外觀的多樣性,(iii) 用於增強多尺度特徵學習的殘差編碼器架構,以及 (iv) 針對病灶的形態學後處理以抑制假陽性腔體預測,同時保留解剖上合理的結構。該框架在 BraTS-MET 2026 任務 1 在線驗證基準上進行評估。在評估的配置中,集成模型(殘差編碼器 nnU-Net + nnU-Net + RC-aware CarveMix)達到最佳性能,病灶-wise Dice 分數分別為 0.732、0.752、0.708 和 0.575,對應的 NSD 分數為 0.794、0.798、0.727 和 0.474,分別針對 ET、TC、WT 和 RC。這些實驗結果顯示,整合 RC-aware 優化、解剖一致的增強和病灶-aware 後處理提供了一種有效的策略,以改善術後腦轉移瘤中稀有切除腔的分割。
DoAtlas-2: A Foundation for Self-Evolving Causal Biomedical Discovery
2609.35107v1 by Yulong Li, Rong Xia, Yuxuan Zhang, Jianxu Chen, Xiwei Liu, Haochen Xue, Maosheng Li, Yuhang Liu, Yibo Yuan, Yutong Xie, Chong Li, Jionglong Su, Hagai Rossman, Eran Segal, Imran Razzak
We introduce DoAtlas-2, a foundation for self-evolving causal biomedical discovery that organizes knowledge around causal mechanisms and advances through external evidence from human populations. DoAtlas-2 integrates 771 research resources covering more than 720,000 participants in 48 countries, from longitudinal clinical phenotypes, medical imaging, and continuous physiological signals to eight molecular layers, together with an evidence network of approximately 4.7 million literature-derived records over 93,566 concepts and 149,383 candidate causal relations. DoAtlas-2 autonomously formulates research questions from evidence gaps and unresolved mechanisms, prespecifies their causal designs, and generates validated analyses. Supporting, challenging, and unresolved results continuously revise mechanistic interpretations, the causal evidence state, and the discovery frontier, so that DoAtlas-2 self-evolves within a closed loop of hypothesis generation, empirical testing, and renewed discovery. DoAtlas-2 has systematically evaluated 2,031 research questions. In the Human Phenotype Project (HPP), it formulated 4,014 candidate pathway questions across vascular, early-glycemic, and hepatic-metabolic systems, and screening of the first 1,079 yielded statistical support for 756. Representative studies identify blood pressure as a convergence node linking adiposity, hepatic, and lipid phenotypes to vascular outcomes, and show that an adiposity-inflammation-blood-pressure pathway is largely attenuated by joint adjustment for body mass index (BMI) and smoking. The discovered vascular network constitutes a completely interpretable predictive foundation, admitting exact attribution of every prediction and closed-form mediation effects. DoAtlas-2 thereby unifies causal mechanism discovery, population-evidence testing, and interpretable prediction within one continuously evolving foundation.
摘要:我們介紹 DoAtlas-2,這是一個自我演化的因果生物醫學發現基礎,圍繞因果機制組織知識,並通過來自人類群體的外部證據推進。DoAtlas-2 整合了 771 個研究資源,涵蓋來自 48 個國家的超過 720,000 名參與者,從縱向臨床表型、醫學影像和連續生理信號到八個分子層面,以及約 4.7 百萬個文獻衍生記錄的證據網絡,涉及 93,566 個概念和 149,383 個候選因果關係。DoAtlas-2 自主地從證據空白和未解決的機制中制定研究問題,預先指定其因果設計,並生成經過驗證的分析。支持、挑戰和未解決的結果不斷修訂機制解釋、因果證據狀態和發現前沿,使 DoAtlas-2 在假設生成、實證測試和新發現的封閉循環中自我演化。DoAtlas-2 系統性地評估了 2,031 個研究問題。在人類表型計畫 (HPP) 中,它制定了 4,014 個候選途徑問題,涵蓋血管、早期糖尿病和肝臟代謝系統,對首批 1,079 個問題的篩選產生了 756 個的統計支持。代表性研究確定血壓為一個匯聚節點,將肥胖、肝臟和脂質表型與血管結果聯繫起來,並顯示肥胖-炎症-血壓途徑在對體重指數 (BMI) 和吸煙進行聯合調整後大幅減弱。所發現的血管網絡構成了一個完全可解釋的預測基礎,允許對每個預測的精確歸因和封閉形式的中介效應。因此,DoAtlas-2 將因果機制發現、群體證據測試和可解釋預測統一於一個不斷演變的基礎之中。
VD-DeepStack: Bridging Visual Comparison and Language Reasoning for Few-Shot Anomaly Detection
2609.34949v1 by Mengyang Zhao, Zhuolin He, Haiyang Yu, Yuxuan Liang, Yifang Xu, Yuchuan Wu, Xiaolei Chen, Zhengtao Yao, Fan Shi, Yang Liu, Bin Li, Xiangyang Xue
Few-shot visual anomaly detection is fundamentally a visual comparison task, requiring fine-grained inspection of a query against normal references. Many recent methods based on large vision-language models (LVLMs) emphasize comparative reasoning through language chain-of-thought. Yet discrete, abstract descriptions may underrepresent dense, fine-grained visual differences, leaving a gap between visual comparison and its expression in language. To address this gap, we propose Visual Difference DeepStack (VD-DeepStack), which explicitly conditions language reasoning on query-reference visual differences. Specifically, we fuse DINO features with the LVLM visual hierarchy to strengthen fine-grained representations, then construct dense difference evidence from residuals between query features and softly matched reference features. The difference-evidence path injects spatially weighted difference vectors into query-image states at multiple decoder depths, while an auxiliary visual-context path provides fine-grained appearance information to support their interpretation. Experiments on 4 industrial and 2 medical anomaly benchmarks demonstrate substantial improvements in few-shot anomaly detection over baselines relying on textual comparative reasoning. These results support mitigating the visual comparison-reasoning gap through the joint design of comparison representations and their integration into the decoder. Code will be released upon acceptance.
摘要:少量樣本的視覺異常檢測基本上是一項視覺比較任務,需要對查詢與正常參考進行細緻的檢查。許多基於大型視覺-語言模型(LVLMs)的最新方法強調通過語言思維鏈進行比較推理。然而,離散的抽象描述可能無法充分表達密集的、細緻的視覺差異,從而在視覺比較與其語言表達之間留下了鴻溝。為了解決這一鴻溝,我們提出了視覺差異深層堆疊(VD-DeepStack),該方法明確地將語言推理條件化於查詢-參考的視覺差異。具體而言,我們將DINO特徵與LVLM視覺層級融合,以加強細緻的表示,然後從查詢特徵與柔性匹配的參考特徵之間的殘差構建密集的差異證據。差異證據路徑在多個解碼器深度將空間加權的差異向量注入查詢圖像狀態,而輔助視覺上下文路徑則提供細緻的外觀信息以支持其解釋。在4個工業和2個醫療異常基準上的實驗顯示,與依賴文本比較推理的基線相比,少量樣本異常檢測有了顯著的改進。這些結果支持通過比較表示的聯合設計及其在解碼器中的整合來減少視覺比較-推理之間的鴻溝。代碼將在接受後發布。
Almost Human, Except When It Matters: VoxParity and the Decisions a Voice Should Change
2609.35922v1 by Bhavik Mangla
A voice agent can handle almost every call on the words alone and still fail the few its sector's rules were written for. Emergency-call standards, fraud guidance, radio phraseology and vulnerability rules recognise that how a caller sounds, or what else is audible, can change the right action. VoxParity tests whether agents act on it. In 183 scenarios from 14 sectors, one transcript stays fixed while the audio changes (a coaching voice, a medical monitor beeping, a mayday under a radio check, noise over a drug name, a child's voice placing a bet, a frightened whisper), and with it the correct typed tool call. A words-only null test credits a system only if hearing the call moves its actions more than it moves a pipeline that only reads the words. Only 11 of the 23 systems that can also be run on the transcript pass. Descriptively, errors run toward the words: when the audio calls for protection, all 28 systems carry out the routine request more often than they over-react on clean calls (41% against 12% pooled; the words-only pipeline, 58% against 15%). Exploratory analyses place most of the leading systems' misses on cues they heard; systems beat the null almost entirely on items that state the rule; the leading systems overrule heard resignation or confusion far more often than acute alarm; and, in the models tested, describing the voice and stating the rule each recover part of the shortfall, leaving a gap on emotion.
摘要:一個語音代理可以僅依賴語言處理幾乎所有的通話,但仍然會在為其行業規則所寫的少數情況下失敗。緊急呼叫標準、詐騙指導、無線電術語和脆弱性規則認識到,來電者的聲音或其他可聽到的內容可能會改變正確的行動。VoxParity 測試代理是否會根據這些因素採取行動。在來自 14 個行業的 183 種情境中,一個文字記錄保持不變,而音頻卻在變化(教練的聲音、醫療監測器的嗶嗶聲、無線電檢查中的求救信號、藥品名稱上的噪音、一個孩子下注的聲音、一個驚恐的低語),隨之而來的是正確的鍵入工具呼叫。僅依賴文字的空白測試僅在聽到通話使其行動比僅閱讀文字的管道更有影響時,才會給系統加分。在可以基於文字記錄運行的 23 個系統中,只有 11 個通過了測試。描述性地說,錯誤傾向於文字:當音頻要求保護時,所有 28 個系統執行例行請求的頻率高於在清晰通話中過度反應的頻率(41% 對 12% 的總和;僅依賴文字的管道為 58% 對 15%)。探索性分析將大多數領先系統的失誤歸因於它們聽到的提示;系統幾乎完全在陳述規則的項目上超越了空白;領先系統在聽到的放棄或困惑上遠比在急性警報上更常推翻;而在測試的模型中,描述聲音和陳述規則各自彌補了部分短缺,但在情感上仍然存在差距。
Nociception as a Control Primitive: Afferent Channels and Nociceptive Memory for Agents Deployed in One Body
2609.34840v1 by Wolfgang Maass
An agent deployed in a single body cannot learn how fast that body wears, because every trial that would reveal its wear resistance wears the body it would protect. We study this \emph{epoch-one} setting, in which the parameters of a fixed-weight policy are set before the body is drawn and never updated in life. The agent carries a load-gated nociceptive channel and a memory that retains what was felt. We prove that felt cost moves the allocation to the best-\emph{paid} work not yet felt rather than the gentlest, that an agent without retention never sees the felt-cost constraint bind, and that the channel pays only where the threat is individually unpredictable, cheap to avoid and expensive to ignore. We measure per body, setting the agent with channel and memory against the same individual without them, where neither carries a schedule learned across lives. On $2{,}000$ simulated floor-layer knees, with wear anchored to published loss rates, feeling, retaining and substituting extends the working life from age $55.2$ to $59.6$ and raises career output from $33.7$ to $36.1$. $69.3\%$ of bodies gain and \textbf{none lose}. A body that feels but retains nothing past the day gains one of the $+4.4$ years, and retention carries the rest. A population-trained agent gains $+0.65$ years from the same channel at $-0.54$ output. The difference is what a species prior already supplies, and a single body has none. The two are related by an identity, the ablation mean reporting $(1-χ)$ of the per-body value with $χ$ the share a blind schedule already captures, so we report both. Where the regime map predicts value, a care robot sextuples its certified service life and a field-anchored fleet writes off $0.15$ of its machines instead of $0.55$. Where it predicts none, a rover gains little over blind caution, so the map holds in both directions.
摘要:一個部署在單一身體中的代理無法學習該身體的磨損速度,因為每一次試驗都會揭示其耐磨性,卻同時磨損了它所保護的身體。我們研究這種\emph{epoch-one}設置,在這種設置中,固定權重策略的參數在身體被繪製之前就已設定,並且在生命中從未更新。代理攜帶一個負載閘控的痛覺通道和一個保留所感知的記憶。我們證明了感知的成本將資源分配到尚未感知的最佳\emph{報酬}工作,而不是最溫和的工作;一個沒有記憶的代理從未見過感知成本約束的束縛;而且該通道僅在威脅是個別不可預測、避免成本低且忽視成本高的情況下才會支付。我們對每個身體進行測量,將具有通道和記憶的代理與同一個體進行比較,而兩者都沒有攜帶跨生命學習的時間表。在$2{,}000$個模擬的地板層膝蓋上,磨損基於已發表的損失率,感知、保留和替代將工作壽命從$55.2$歲延長至$59.6$歲,並將職業產出從$33.7$提高到$36.1$。$69.3\%$的身體獲得收益,\textbf{沒有一個損失}。一個感知但在當天之後不保留任何東西的身體獲得了$+4.4$年的增益,而保留則承擔了其餘的部分。一個經過人群訓練的代理從相同的通道中獲得$+0.65$年的增益,產出為$-0.54$。這一差異是物種先前已經提供的,而單一身體則沒有。這兩者通過一個身份相關聯,消融均值報告每個身體價值的$(1-χ)$,其中$χ$是盲目時間表已經捕獲的份額,因此我們報告兩者。在制度地圖預測價值的地方,一個護理機器人將其認證服務壽命增長六倍,而一個基於現場的艦隊則將$0.15$的機器報廢,而不是$0.55$。在預測為零的地方,一個探測器的增益僅略高於盲目謹慎,因此該地圖在兩個方向上都成立。
Applying Language Models in Clinical Medicine: Recent Trends and Perspectives
2609.34780v2 by Erik Aerts
The use and applicability of artificial intelligence (AI) in medical research and clinical practice has received increasing attention in the literature over recent years. The emergence of large language models (LLMs) has expanded discussions in regards to applications of AI within healthcare. While traditional deep learning based AI applications in medicine have often focused on specific and defined tasks, LLMs offer broader capabilities and flexibility in working with available data,. At the same time of writing, the integration of LLMs into medical settings raises important questions regarding their reliability, accuracy, transparency, safety, and appropriate role in a medical setting. This text presents and discusses recent talks and articles concerning the application of LLMs in medicine, with particular emphasis on their potential utility in research and clinical practice. It considers both the opportunities offered by these technologies and the challenges associated with their implementation, aiming to provide a perspective on the current and emerging role of LLMs within the medical field.
摘要:人工智慧(AI)在醫學研究和臨床實踐中的使用和適用性在近年來的文獻中受到越來越多的關注。大型語言模型(LLMs)的出現擴大了關於AI在醫療保健中應用的討論。雖然傳統基於深度學習的AI應用在醫學中往往專注於特定和明確的任務,但LLMs在處理可用數據方面提供了更廣泛的能力和靈活性。撰寫本文的同時,將LLMs整合進醫療環境中引發了有關其可靠性、準確性、透明度、安全性和在醫療環境中適當角色的重要問題。本文呈現並討論了有關LLMs在醫學中應用的最近演講和文章,特別強調它們在研究和臨床實踐中的潛在效用。它考慮了這些技術所提供的機會以及與其實施相關的挑戰,旨在提供對LLMs在醫療領域中當前和新興角色的看法。
ResonAct: Streaming Metrics for Runtime Diagnosis and Self-Healing in Multi-Agent Systems
2609.34701v1 by Tarun Chintada, Neelamadhav Gantayat, Ishaan Romil, Renuka Sindhgatta, Soujanya Soni, Sameep Mehta
Multi-agent systems (MAS) are increasingly used to automate enterprise workflows involving multiple specialized agents, external tools, and long-running task execution. Failures may arise from tool degradation, context propagation errors, coordination breakdowns, or repeated agent interactions that prevent task completion. While existing observability frameworks provide traces and logs, diagnosis and remediation are largely performed after execution completes, limiting opportunities for recovery during runtime. We present ResonAct, a runtime self-healing framework that enables continuous monitoring, diagnosis, and remediation of multi-agent systems through streaming operational metrics. ResonAct ingests execution traces, agent interactions, and tool invocations into a streaming analytics layer that continuously derives task progress, context health, and tool reliability metrics. These metrics serve as runtime control signals for detecting anomalous execution patterns and localizing root causes using a structured failure model. Based on the diagnosed failure, ResonAct dynamically selects remediation policies and performs actions. The framework operates as an external control plane, enabling intervention without modifying application agents or orchestration logic. We evaluate ResonAct across enterprise workflow scenarios and AppWorld benchmarks. The results show that the streaming metric-based analysis identifies execution degradations and localizes faults. Furthermore, policy-driven remediation improves task completion rates by up to 10.00 percentage points, with detection precision ranging from 70.59% to 82.91%, recall from 63.09% to 100%, recovery rates from 10.48% to 46.67%, and runtime overhead ranging from $-0.25%$ to 14.12% across the evaluated configurations.
摘要:多代理系統(MAS)越來越多地用於自動化涉及多個專門代理、外部工具和長時間運行任務執行的企業工作流程。失敗可能源於工具退化、上下文傳播錯誤、協調崩潰或重複的代理互動,這些都會阻礙任務的完成。雖然現有的可觀察性框架提供了追蹤和日誌,但診斷和修復主要是在執行完成後進行,這限制了在運行時進行恢復的機會。我們提出了ResonAct,一個運行時自我修復框架,通過流式操作指標實現多代理系統的持續監控、診斷和修復。ResonAct 將執行追蹤、代理互動和工具調用輸入到一個流式分析層,該層不斷推導任務進度、上下文健康狀況和工具可靠性指標。這些指標作為運行時控制信號,用於檢測異常執行模式並使用結構化故障模型定位根本原因。根據診斷出的故障,ResonAct 動態選擇修復政策並執行行動。該框架作為外部控制平面運行,允許在不修改應用代理或編排邏輯的情況下進行干預。我們在企業工作流程場景和AppWorld基準測試中評估了ResonAct。結果顯示,基於流式指標的分析能夠識別執行退化並定位故障。此外,基於政策的修復將任務完成率提高了最多10.00個百分點,檢測精度範圍為70.59%到82.91%,召回率範圍為63.09%到100%,恢復率範圍為10.48%到46.67%,運行時開銷範圍為$-0.25%$到14.12%,涵蓋了評估的配置。
SentZero: An Enhanced Sentence-Centric Vision-Language Pretraining for Multi-Task Zero-Shot Chest X-Ray Analysis
2609.34479v1 by Hangyul Yoon, Hyungyung Lee, Edward Choi, Eunho Yang
Vision-language (VL) pretraining using paired chest X-ray (CXR) images and radiology reports has shown strong potential for medical image understanding. However, existing methods often remain dependent on task-specific finetuning because radiology reports are lengthy, clinically dense, and difficult to align with simple zero-shot prompts. Recent sentence-level approaches partially address this limitation using clinical phrases extracted by large language models (LLMs), but they largely overlook the intrinsic characteristics of radiology discourse. In particular, limited positive-pair diversity constrains further gains, while clinically equivalent sentences frequently recur across patients, creating false negatives in contrastive learning. To address these issues, we propose SentZero, an enhanced sentence-centric VL pretraining framework for zero-shot, multi-task CXR analysis. SentZero introduces LLM-based abstract-level sentence structuring and mapping to expand positive-pair diversity, together with an additional loss term to mitigate false negatives. We further introduce sentence-conditioned residual modulation of visual embeddings, enabling visual features to adapt to the semantic characteristics of each input sentence. Across diverse downstream tasks and datasets, SentZero improves zero-shot generalization and outperforms prior multi-task zero-shot methods.
摘要:視覺-語言(VL)預訓練利用配對的胸部X光(CXR)影像和放射學報告顯示出對醫學影像理解的強大潛力。
然而,現有的方法往往仍依賴於特定任務的微調,因為放射學報告冗長、臨床密集,並且難以與簡單的零樣本提示對齊。
最近的句子級方法部分解決了這一限制,使用大型語言模型(LLMs)提取的臨床短語,但它們在很大程度上忽略了放射學話語的內在特徵。
特別是,有限的正配對多樣性限制了進一步的增益,而臨床等效句子在不同患者之間經常重複,造成對比學習中的假陰性。
為了解決這些問題,我們提出了SentZero,一個增強的以句子為中心的VL預訓練框架,用於零樣本的多任務CXR分析。
SentZero引入基於LLM的抽象級句子結構和映射,以擴大正配對的多樣性,並增加一個額外的損失項以減輕假陰性。
我們進一步引入句子條件的視覺嵌入殘差調制,使視覺特徵能夠適應每個輸入句子的語義特徵。
在多樣的下游任務和數據集上,SentZero改善了零樣本泛化並超越了之前的多任務零樣本方法。
VL-AcneSeg: A Vision-Language Framework for Region-Aware Acne Lesion Segmentation
2609.34472v1 by Sukju Oh, Soo Ick Cho, Dae Hun Suh, Sukkyu Sun
Acne assessment is crucial for clinical decision-making, yet traditional grading and counting are subjective and fail to account for lesion size. While area-based assessment has emerged as a promising alternative, acne segmentation has continued to rely on general-purpose architectures. To address this gap, we propose VL-AcneSeg, a multimodal framework for acne lesion segmentation that leverages CLIP and region-level text prompts to incorporate spatial priors, enabling lesions to be localized across the whole face. Because region-level prompts indicate which facial areas contain lesions, we report a single global prompt, which requires no such information, as our primary setting. On our internal clinical dataset, VL-AcneSeg achieves a Dice score of 0.5082 and an IoU of 0.3407 under this protocol, the highest among all compared methods, including recent vision-language segmentation methods that are themselves given region-level prompts; region-level prompting raises these to 0.5296 and 0.3602. Moreover, lesion area measurements derived from our segmentation correlate with IGA scores at a level comparable to expert annotations (Pearson r = 0.719 versus 0.658). Notably, our framework maintains consistent performance across external validation datasets, performing reliably even on uncontrolled smartphone images without requiring additional training or fine-tuning. By pairing a protocol that requires no lesion-location information with area-based severity estimation, this work provides a foundation for objective acne assessment outside the clinic. Our implementation is publicly available at: https://github.com/sukjuoh/VL-AcneSeg
摘要:痤瘡評估對於臨床決策至關重要,但傳統的分級和計數方法主觀性強,且未能考慮病變大小。雖然基於面積的評估已成為一種有前景的替代方案,但痤瘡分割仍然依賴於通用架構。為了填補這一空白,我們提出了 VL-AcneSeg,一個多模態框架,用於痤瘡病變分割,利用 CLIP 和區域級文本提示來融入空間先驗,使病變能夠在整個面部進行定位。由於區域級提示指示哪些面部區域包含病變,我們報告了一個單一的全局提示,作為我們的主要設置,這不需要任何此類信息。在我們的內部臨床數據集中,VL-AcneSeg 在這一協議下達到了 0.5082 的 Dice 分數和 0.3407 的 IoU,這是所有比較方法中最高的,包括最近的視覺-語言分割方法,這些方法本身也提供了區域級提示;區域級提示將這些指標提高到 0.5296 和 0.3602。此外,從我們的分割中得出的病變面積測量與 IGA 分數的相關性達到與專家註釋相當的水平(Pearson r = 0.719 對比 0.658)。值得注意的是,我們的框架在外部驗證數據集上保持一致的性能,即使在無法控制的智能手機圖像上也能可靠地執行,而無需額外的訓練或微調。通過將不需要病變定位信息的協議與基於面積的嚴重程度評估相結合,這項工作為臨床外的客觀痤瘡評估提供了基礎。我們的實現已公開可用於:https://github.com/sukjuoh/VL-AcneSeg
Evolving Support Priorities in Empathetic Reinforcement Learning
2609.34249v1 by Pengyu Huang, Zhiyuan Han, Wenwen Tong, Hewei Guo, Jiangnan Chen, Sirui Chen, Lewei Lu, Beier Zhu, Xun Yang
We identify a fundamental mismatch in empathetic reinforcement learning: support priorities evolve with the dialogue state, yet existing methods typically optimize predefined reward specifications that remain fixed across turns. To model these evolving support priorities, we organize empathetic support along cognitive, affective, and proactive empathy, and propose Context-Adaptive Rubric Evolution (CARE). At each turn, CARE generates a context-adaptive rubric by adjusting both the weights of these three empathy dimensions and their fine-grained evaluation criteria. The rubric generator is trained with turn-level rubric supervision and human preference data through supervised fine-tuning followed by preference-based reinforcement learning, and then serves as an adaptive reward interface for online empathetic RL. Integrated with both RLVER and MICA, CARE achieves state-of-the-art performance across SentientBench, EQBench3, and EMPA under three independent LLM judges. Notably, on EMPA, CARE improves EPM-Idx over the strongest baseline by at least 13 points under all three judges, including an increase from 28.11 to 83.54 under Gemini-2.5-Pro. Further analyses show that learned rubric priorities systematically vary across dialogue stages and user emotions, demonstrating that CARE adapts what is rewarded as support needs evolve.
摘要:我們發現同理心強化學習中存在一個根本的不匹配:支持優先級隨著對話狀態而演變,但現有方法通常優化預定的獎勵規範,這些規範在各回合中保持固定。為了建模這些不斷演變的支持優先級,我們沿著認知、情感和主動同理心組織同理心支持,並提出了上下文自適應評分標準演變(CARE)。在每個回合中,CARE 通過調整這三個同理心維度的權重及其細緻的評估標準來生成一個上下文自適應的評分標準。評分標準生成器通過回合級評分標準監督和人類偏好數據進行監督微調,然後通過基於偏好的強化學習進行訓練,並作為在線同理心強化學習的自適應獎勵介面。與 RLVER 和 MICA 結合,CARE 在 SentientBench、EQBench3 和 EMPA 上實現了最先進的性能,並在三位獨立的 LLM 評審中表現出色。值得注意的是,在 EMPA 上,CARE 在所有三位評審下將 EPM-Idx 提高了至少 13 分,包括在 Gemini-2.5-Pro 下從 28.11 增加到 83.54。進一步的分析顯示,學習到的評分標準優先級在對話階段和用戶情緒之間系統性地變化,顯示出 CARE 會隨著支持需求的演變而調整獎勵內容。
Unknown is not normal: separating language-model extraction from rule-based decision logic for clinical risk scores
2609.34112v1 by Nicolás Vera Zúñiga
Large language models (LLMs) are increasingly used to compute clinical risk scores from free-text notes. Notes are often incomplete, and treating undocumented findings as normal can silently misclassify patients. We test whether separating three-state extraction (present, absent or unknown, by an LLM) from decision logic (deterministic code computing score bounds over unknown inputs) lets a system ask only questions that can change the decision. On 1,200 synthetic emergency cases across six calculators (HEART, CURB-65, qSOFA, PERC, Wells, Cockcroft-Gault), with a simulated clinician answering questions, we compared this bounds policy with asking for every missing input, a missing-equals-normal schema, and an end-to-end LLM agent (Claude Opus 5.5). With Claude Haiku 4.5 as extractor, the bounds policy matched ask-all accuracy (99.4% vs 99.4%) with half the questions (0.92 vs 1.78 per case) and no irrelevant ones. Treating missing as normal dropped accuracy to 91.2% and under-triaged 8.5% of patients (95% CI 7.1-10.2), and under-triage persisted under messy notes and a noisy clinician. The agent was equally accurate under ideal conditions (99.6%) but 9.5% of its questions were irrelevant; with a noisy clinician it was less accurate than the bounds policy (83.5% vs 87.0%, p<0.001) and committed prematurely in 2.7% of cases (bounds: 0%). A 9B local model as extractor reached oracle-level accuracy (99.8%). In 584 real case reports from MedCalc-Bench, only 52% contained enough information to determine the category (HEART 13%). Routing decisions through code that reasons explicitly about unknowns avoids premature commitment and irrelevant questions, halves the questions asked, and works with small local models.
摘要:大型語言模型(LLMs)越來越多地用於從自由文本筆記中計算臨床風險分數。
筆記通常不完整,將未記錄的發現視為正常可能會悄悄地錯誤分類患者。
我們測試了將三狀態提取(由LLM判斷的存在、缺失或未知)與決策邏輯(在未知輸入上計算分數邊界的確定性代碼)分開,是否能讓系統僅提出能改變決策的問題。
在1200個合成緊急案例中,涵蓋六個計算器(HEART、CURB-65、qSOFA、PERC、Wells、Cockcroft-Gault),並由模擬臨床醫生回答問題,我們將這個邊界政策與要求每個缺失輸入的方式、缺失等於正常的模式以及一個端到端的LLM代理(Claude Opus 5.5)進行比較。
使用Claude Haiku 4.5作為提取器,邊界政策在問題數量減半的情況下(每個案例0.92對1.78)達到了與要求所有問題相同的準確率(99.4%對99.4%),且沒有不相關的問題。
將缺失視為正常使準確率降至91.2%,並使8.5%的患者被低估分級(95% CI 7.1-10.2),而且在雜亂的筆記和噪音臨床醫生的情況下,低估分級仍然存在。
在理想條件下,代理的準確率同樣為99.6%,但其9.5%的問題是無關的;在噪音臨床醫生的情況下,其準確率低於邊界政策(83.5%對87.0%,p<0.001),並在2.7%的案例中過早做出承諾(邊界:0%)。
一個9B本地模型作為提取器達到了神諭級的準確率(99.8%)。
在來自MedCalc-Bench的584個真實案例報告中,只有52%包含足夠的信息來確定類別(HEART 13%)。
通過對未知進行明確推理的代碼進行路由決策,可以避免過早承諾和不相關的問題,將提問數量減半,並且能與小型本地模型一起工作。
The Devil is in the Spectrum Bias: Spectrum-Balanced Feature Matching for Robust Representation Distillation
2609.34106v1 by Kuniaki Saito, Yoshitaka Ushiku
Large visual foundation models have demonstrated remarkable transferability across a wide range of downstream tasks. To deploy such models efficiently, feature matching has become a popular knowledge distillation approach that transfers teacher representations to smaller student models without requiring labeled data. However, we show that the conventional feature matching objective with L2-distance is inherently biased toward reconstructing dominant spectral directions of the teacher representation, while under-optimizing low-variance directions that often contain task-relevant information. To address this, we propose Spectrum-Balanced Feature Matching, SpecMatch, a simple objective that adaptively emphasizes under-optimized spectral directions while preserving the relative importance of dominant directions. SpecMatch is easy to implement and introduces negligible computational overhead. Extensive experiments on image recognition demonstrate that SpecMatch consistently improves downstream adaptation across diverse tasks, including image classification, anomaly detection, medical image analysis, and domain generalization. In particular, SpecMatch outperforms conventional feature matching in 40 of 42 teacher--student and training-setting combinations, while consistently improving over the original student model in all settings. We further demonstrate that the proposed objective generalizes beyond vision, improving downstream performance across six protein understanding tasks.
摘要:大型視覺基礎模型在各種下游任務中顯示出卓越的可轉移性。
為了有效部署這些模型,特徵匹配已成為一種流行的知識蒸餾方法,該方法將教師表示轉移到較小的學生模型,而無需標記數據。
然而,我們顯示,傳統的L2距離特徵匹配目標本質上偏向於重建教師表示的主導光譜方向,同時對通常包含任務相關信息的低方差方向進行優化不足。
為了解決這個問題,我們提出了光譜平衡特徵匹配(Spectrum-Balanced Feature Matching,SpecMatch),這是一個簡單的目標,能夠自適應地強調未優化的光譜方向,同時保留主導方向的相對重要性。
SpecMatch易於實現,並引入了微不足道的計算開銷。
在圖像識別方面的廣泛實驗表明,SpecMatch在包括圖像分類、異常檢測、醫學影像分析和領域泛化等多種任務中,持續改善下游適應性。
特別是,SpecMatch在42個教師-學生和訓練設置組合中的40個中超越了傳統的特徵匹配,同時在所有設置中持續改善了原始學生模型。
我們進一步證明,所提出的目標在視覺之外也具有泛化能力,改善了六個蛋白質理解任務的下游性能。
TRACE: Expert-Aligned ECG Representation Learning with Rigorous Benchmarking and Real-World Validation in Acute Cardiac Care
2609.34088v1 by Lovely Yeswanth Panchumarthi, Andrew Lu, Saurabh Kataria, Delgersuren Bold, Minxiao Wang, Runze Yan, Patricia Dykes, Brian J. Gow, Tom J. Pollard, Jessica K. Zègre-Hemsey, Dillon J. Dzikowicz, Lekshmi Kumar, Xiao Hu, Ran Xiao
TRACE (Text-Reinforced Analysis of Cardio ECGs) is a multimodal electrocardiogram (ECG) representation model that learns clinically grounded signal embeddings for downstream cardiac classification. It is designed to address the limitations of existing CLIP-style training, which often struggles with noisy clinical text and fails to leverage the complementary strengths of unimodal (from ECG) and cross-modal (between ECG and matched cardiologist reports) learning. To bridge this gap, we propose a hybrid architecture that jointly learns unimodal and cross-modal representations via uncertainty-weighted multi-task learning while utilizing an LLM-based pipeline to extract high-fidelity findings from cardiologist reports. We evaluate TRACE across a spectrum of clinical urgency, establishing robust performance on public benchmarks for arrhythmia classification and structural abnormalities relative to existing unimodal and multimodal ECG models. To demonstrate real-world utility, we further validate the model on acute coronary occlusion (ACO), where the prevailing ST-elevation criteria miss 25-34% of true occlusions. Utilizing a large private ACO dataset with expert-annotated ground truth, TRACE significantly outperforms real-world clinical practice, yielding a 19.0% increase in sensitivity or a 62.6% reduction in false positive rates at the clinical baseline. This extensive evaluation confirms that TRACE delivers both strong performance on benchmark tasks and tangible clinical impact in the most acute, high-risk cardiac scenarios.
摘要:TRACE(文本強化心電圖分析)是一種多模態心電圖(ECG)表示模型,旨在學習臨床基礎的信號嵌入,以便用於下游心臟分類。它旨在解決現有CLIP風格訓練的局限性,這種訓練通常在嘈雜的臨床文本中掙扎,並未能利用單模態(來自ECG)和跨模態(在ECG與匹配的心臟病醫生報告之間)學習的互補優勢。為了填補這一空白,我們提出了一種混合架構,通過不確定性加權的多任務學習共同學習單模態和跨模態表示,同時利用基於LLM的管道從心臟病醫生報告中提取高保真結果。我們在臨床緊急程度的範疇內評估TRACE,並在公開基準上建立了對心律失常分類和結構異常的穩健性能,相較於現有的單模態和多模態ECG模型。為了展示其在現實世界中的實用性,我們進一步在急性冠狀動脈阻塞(ACO)上驗證該模型,當前的ST抬高標準錯過了25-34%的真實阻塞。利用一個大型私有ACO數據集,並由專家標註的真實情況,TRACE顯著超越了現實世界的臨床實踐,在臨床基線下實現了19.0%的靈敏度提升或62.6%的假陽性率降低。這一廣泛的評估確認TRACE在基準任務上提供了強大的性能,並在最急迫、高風險的心臟情況下帶來了實質性的臨床影響。
Who Gets a Token, and What Does It Carry? Unequal Name Support and Concept Access in Large Language Models
2609.34065v1 by Mir Tafseer Nayeem, Davood Rafiei
Names are personal identifiers, but they also carry social meaning and are widely used to evaluate how language models treat different people. Such evaluations typically assume that matched names are comparable model inputs. We show that this assumption often fails at the lexical interface: matched names are not necessarily matched inputs. Some names receive direct single-token access, while others are assembled from multiple subwords, creating unequal name-surface support. Across nearly half a million first names and 12 LLM-associated tokenizers, direct lexical access is highly selective, model dependent, and uneven across race- and gender-associated name metadata. We introduce NameTrace, a model-native, fine-grained, pre-behavioral framework for measuring whether unequal name-surface support remains a vocabulary property or becomes visible in task-relevant internal representations. NameTrace measures concept accessibility from the model's own probabilities over task-specific adjective axes with continuous task-aligned weights. On matched atomic and short-fragmented names within the same race/ethnicity--gender-associated strata, support predicts systematic differences in concept accessibility across fellowship, hiring, clinical assessment, and lending. These differences persist across all eight matched strata, extend across model families, and transfer to unseen names. Hidden-state interventions further show that the measured task directions have downstream leverage, shifting later constrained choices. Unequal lexical support is therefore demographically structured at the input and remains visible in task-relevant model computation. NameTrace makes lexical comparability measurable, supporting a broader principle: behavioral comparability begins with lexical comparability.
摘要:名字是個人識別符號,但它們也承載著社會意義,並廣泛用於評估語言模型如何對待不同的人。這些評估通常假設匹配的名字是可比較的模型輸入。我們顯示這一假設在詞彙介面上經常失效:匹配的名字不一定是匹配的輸入。有些名字獲得直接的單詞訪問,而另一些則由多個子詞組成,造成不平等的名字表面支持。在近五十萬個名字和12個與LLM相關的分詞器中,直接詞彙訪問高度選擇性,依賴於模型,並且在與種族和性別相關的名字元數據中不均勻。我們引入了NameTrace,這是一個模型原生的、細粒度的、前行為框架,用於測量不平等的名字表面支持是否仍然是一種詞彙特性,或在任務相關的內部表示中變得可見。NameTrace測量從模型自身的概率中對任務特定形容詞軸的概念可及性,並使用連續的任務對齊權重。在同一種族/民族—性別相關的層級內,對匹配的原子和短片段名字的支持預測在獎學金、招聘、臨床評估和貸款方面的概念可及性系統性差異。這些差異在所有八個匹配層級中持續存在,跨模型家族擴展,並轉移到未見過的名字。隱藏狀態的干預進一步顯示,測量的任務方向具有下游影響,改變後續的受限選擇。因此,不平等的詞彙支持在輸入層面上是人口結構化的,並在任務相關的模型計算中保持可見。NameTrace使詞彙可比性可測量,支持一個更廣泛的原則:行為可比性始於詞彙可比性。
Large Language Models for Structured Clinical Data Analysis: Dual-Agent Grounding and Validation
2609.34039v1 by Erfan D. Dehkalani, Seetha Shankaran, Abbot R. Laptook, C. Michael Cotten, P. Ellen Grant, Yangming Ou
Objective: To develop and characterize CLEAR-Med, a dual-agent framework for natural-language analysis of structured clinical data that separates SQL-based invocation from independent validation. Methods: CLEAR-Med uses one agent to translate a question into executable Structured Query Language (SQL), retain the executed query and database result, and produce a draft. Deterministic checks and a separately invoked cross-provider Validation Agent then accept the draft, request one bounded repair, or abstain. We formalized the system as a bounded selective pipeline and evaluated CLEAR-Med's configuration and scalability, and the Invocation Agent's accuracy and consistency on a 25-query development benchmark, using a harmonized 21-site neonatal hypoxic-ischemic encephalopathy table containing 532 de-identified infant records and approximately 1,300 variables. Results: CLEAR-Med completed all six nominal scalability configurations, including 500x1300. Across 25 development-benchmark queries repeated five times, the Invocation Agent answered 83 of 125 responses correctly (66.4%; query-cluster bootstrap 95% CI, 48.0-83.2%), compared with 15 of 125 (12.0%; 95% CI, 3.2-22.4%) for the ungrounded ChatGPT baseline, a paired improvement of 54.4 percentage points (95% CI, 36.8-72.0%). Conclusion: CLEAR-Med provides a general architecture for traceable analysis of structured clinical data: numerical claims remain linked to executed SQL, and unresolved cases can fail closed. The reported experiments characterize CLEAR-Med's configuration and scalability and the Invocation Agent's accuracy, while the formal analysis establishes the encoded-property guarantee of the complete control flow; a prospective full-pipeline evaluation of the validation and abstention stages is the next stage of this work.
摘要:目標:開發並表徵CLEAR-Med,這是一個雙代理框架,用於自然語言分析結構化臨床數據,將基於SQL的調用與獨立驗證分開。
方法:CLEAR-Med使用一個代理將問題翻譯為可執行的結構化查詢語言(SQL),保留執行的查詢和數據庫結果,並生成草稿。
確定性檢查和單獨調用的跨提供者驗證代理然後接受草稿,請求一次有限的修正,或選擇不進行修正。
我們將系統形式化為一個有限的選擇性管道,並評估CLEAR-Med的配置和可擴展性,以及調用代理在25個查詢開發基準上的準確性和一致性,使用一個和諧的21個站點的新生兒缺氧缺血性腦病表,該表包含532個去標識的嬰兒記錄和約1300個變量。
結果:CLEAR-Med完成了所有六個名義上的可擴展性配置,包括500x1300。
在25個開發基準查詢中重複五次,調用代理正確回答了125個回應中的83個(66.4%;查詢集群自助法95%置信區間,48.0-83.2%),而未經驗證的ChatGPT基準僅正確回答了125個中的15個(12.0%;95%置信區間,3.2-22.4%),這是一個配對改善54.4個百分點(95%置信區間,36.8-72.0%)。
結論:CLEAR-Med提供了一個可追溯的結構化臨床數據分析的通用架構:數值聲明仍然與執行的SQL相連,未解決的情況可以失敗關閉。
報告的實驗表徵了CLEAR-Med的配置和可擴展性以及調用代理的準確性,而正式分析確立了完整控制流的編碼屬性保證;對驗證和放棄階段的前瞻性全管道評估是這項工作的下一階段。
Jev in Medicine: A Benchmark Evaluation
2609.34024v2 by Alfredo Madrid-García, Beatriz Merino-Barbancho
Jev is a non-generative "System One" model that assigns probabilities to predefined answer options and cannot answer outside them. Its accuracy and calibration on medical question-answering and case-based diagnostic-reasoning tasks are unknown. We evaluated Jev 1.13 on four medical benchmarks: MetaMedQA, PubMedQA, DiagnosisArena-MCQ and the NEJM Case Challenges. GPT-6 Sol, with (medium) and without reasoning, was the reference. The primary outcome was top-1 accuracy; key secondary outcomes were calibration, selective prediction and recognition of unanswerable questions. All 8,469 requests returned a valid answer. Jev's accuracy was similar to that of GPT-6 Sol with medium reasoning on PubMedQA (78.4% vs 78.2%;), lower on MetaMedQA (74.8% vs 82.7%) and much lower on DiagnosisArena-MCQ (59.8% vs 82.4%;) and the NEJM cases (61.8% vs 82.4%). On MetaMedQA, Jev's probabilities were the best calibrated (expected calibration error 0.063 vs 0.146), and its answers with a probability of at least 0.9 (52.9% of questions) were 93.4% accurate, but GPT-6 Sol was as accurate when it accepted a similar proportion of questions. On DiagnosisArena-MCQ, Jev's probabilities discriminated poorly (AUROC 0.645 vs 0.768). Of the 162 questions whose correct answer was "I don't know or cannot answer", Jev chose that option for 10.5% (GPT-6 Sol, 8.6%). Median latency was 0.27-0.31 s; all 2,823 items cost USD 0.08. Jev was fast and inexpensive, and its accuracy was similar to that of a frontier LLM on research abstracts but lower on examination questions and much lower on complex diagnostic cases. Task-specific validation is required before clinical use.
摘要:Jev是一個非生成的「系統一」模型,為預定的答案選項分配概率,並且無法在這些選項之外回答。它在醫學問答和基於案例的診斷推理任務上的準確性和校準情況尚不清楚。我們在四個醫學基準上評估了Jev 1.13:MetaMedQA、PubMedQA、DiagnosisArena-MCQ和NEJM案例挑戰。GPT-6 Sol,無論是(中等)還是沒有推理,都是參考標準。主要結果是前1準確率;關鍵的次要結果包括校準、選擇性預測和對無法回答問題的識別。所有8,469個請求都返回了有效的答案。Jev的準確性與GPT-6 Sol在PubMedQA上的中等推理相似(78.4%對78.2%;),在MetaMedQA上較低(74.8%對82.7%),在DiagnosisArena-MCQ上則低得多(59.8%對82.4%;)以及在NEJM案例上(61.8%對82.4%)。在MetaMedQA上,Jev的概率校準最佳(預期校準誤差0.063對0.146),其概率至少為0.9的答案(52.9%的問題)準確率為93.4%,但當GPT-6 Sol接受相似比例的問題時,其準確率也相當。在DiagnosisArena-MCQ上,Jev的概率區分能力較差(AUROC 0.645對0.768)。在162個正確答案為「我不知道或無法回答」的問題中,Jev選擇該選項的比例為10.5%(GPT-6 Sol為8.6%)。中位延遲為0.27-0.31秒;所有2,823個項目的成本為0.08美元。Jev速度快且成本低,其準確性與前沿LLM在研究摘要上的表現相似,但在考試問題上的準確性較低,對於複雜的診斷案例則低得多。在臨床使用之前,需要進行特定任務的驗證。
EHRAdapt: Adapting Pretrained Language Models to Electronic Health Records with Semantic Priors for Rare Clinical Events
2609.34007v1 by Andre R Goncalves, Vincent Liu, Priyadip Ray
Electronic health records (EHRs) encode clinical histories as (time, modality, code) tuples, whereas pretrained language models expect text tokens. Serializing them as text inflates sequence length and redundantly encodes structure. We introduce EHRAdapt, an adapter that maps tuples directly into a frozen language model's embedding space. Modality receives a learned embedding, time gaps enter through learned attention biases, and event codes receive dedicated vectors. Learning event vectors is the central challenge: clinical vocabularies are long-tailed, leaving rare events too few observations for reliable estimates. EHRAdapt therefore represents each event vector as the sum of a semantic prior and an evidence residual. The prior is a frozen embedding of the event's clinical description from a biomedical language model trained on clinical ontologies, mapped into the model's input space by a shared learned projection, so it supplies clinical meaning even when observations are scarce. The residual, a learned low-rank event-specific correction, refines it as evidence accumulates. We run continued pretraining on about 4 million patients' records with three frozen LLM backbones (OLMo2 1B, Llama3.2 1B, and OLMo2 7B), training only the adapter (0.1--0.6% of all parameters). The full adapter outperforms all ablations in held-out next-event prediction on every backbone. Removing the semantic pathway hurts rare events over ten times more than the most frequent ones, whereas removing the residual hurts overall prediction but improves it for the rarest events. On reportable infectious-disease and syndromic downstream classification tasks, EHRAdapt outperforms text-based LLM and count-based baselines, and both pathways improve rare-disease discrimination. The two pathways therefore play complementary roles, visible only when results are broken down by event frequency rather than averaged.
摘要:電子健康紀錄(EHRs)將臨床歷史編碼為(時間、模態、代碼)元組,而預訓練的語言模型則期望文本標記。將它們序列化為文本會增加序列長度並冗餘地編碼結構。我們引入了EHRAdapt,一個將元組直接映射到凍結語言模型嵌入空間的適配器。模態獲得一個學習的嵌入,時間間隙通過學習的注意力偏差進入,而事件代碼則獲得專用向量。學習事件向量是中心挑戰:臨床詞彙是長尾的,稀有事件的觀察次數太少,無法進行可靠的估計。因此,EHRAdapt將每個事件向量表示為語義先驗和證據殘差的總和。先驗是來自於在臨床本體上訓練的生物醫學語言模型的事件臨床描述的凍結嵌入,通過共享的學習投影映射到模型的輸入空間,因此即使觀察稀少也能提供臨床意義。殘差是一個學習的低秩事件特定修正,隨著證據的累積進行精細化。我們對約400萬名患者的記錄進行持續的預訓練,使用三個凍結的LLM骨幹(OLMo2 1B、Llama3.2 1B和OLMo2 7B),僅訓練適配器(佔所有參數的0.1--0.6%)。完整的適配器在每個骨幹的保留下一個事件預測中超越了所有的消融實驗。移除語義通路對稀有事件的影響超過最常見事件的十倍,而移除殘差則對整體預測造成損害,但對最稀有事件的預測有所改善。在可報告的傳染病和綜合徵下游分類任務中,EHRAdapt的表現超過基於文本的LLM和基於計數的基準,並且兩個通路都改善了稀有疾病的區分。因此,這兩個通路扮演互補的角色,只有當結果按事件頻率細分而非平均時才會顯現出來。
Is your uncertainty map wrong, or is its target? Exact diagnostics for the Tweedie diagonal, and a gradient-free alternative
2609.33786v1 by Vicent Ribas, Anna Oliveras Tous
A diffusion model can predict a follow-up medical scan from a baseline, but a clinician needs a per-voxel map of where that prediction can be trusted. Many such maps approximate the diagonal of the Tweedie posterior covariance, and are evaluated against another approximation of it, so whether the estimator or the target limits them is unclear. We compute the exact diagonal on six checkpoints across fourteen model-corpus conditions. Hutchinson at M=200 tracks it at rank agreement of at least 0.92 everywhere, yet in four of the fourteen the exact diagonal is anti-correlated with the denoising error, reaching -0.13, so a faithful estimator reproduces that reversal. All four are real-image conditions; on the models' own samples the reversal does not appear, so evaluating on generated samples flatters this family. What limits these maps is the target, not the estimator. We then introduce Tweedie Probe-Tangent (T-PT), a gradient-free residual probe that corrupts one model-supported prediction repeatedly and measures the voxel-wise variance of the denoiser's response. T-PT reads a different functional of the same Jacobian, and its exact second-order form ranks with the diagonal wherever the diagonal reverses; at thirty probes it returns a map too unstable to reproduce that ranking, while Hutchinson at M=5 already reproduces it, so T-PT there is not evidence against the reversal. We offer it as an instrument, not a better approximation. On brain MRI at full resolution, where every Jacobian-based estimator we test runs out of memory, T-PT leads a twenty-chain Monte-Carlo ensemble on five of eight endpoints inside tissue and trails it on none, at 16x fewer network evaluations; over the whole volume the ensemble leads, and fifty chains close the tissue gap. On lung CT the ensemble is ahead throughout. Both lose most of their discrimination where the change is, which remains open.
摘要:擴散模型可以從基線預測後續的醫學掃描,但臨床醫生需要一個每體素的地圖來確定該預測的可信度。許多這樣的地圖近似於Tweedie後驗協方差的對角線,並且是根據另一個近似進行評估,因此估計器或目標是否限制了它們尚不清楚。我們在十四個模型-語料條件下的六個檢查點計算了精確的對角線。Hutchinson在M=200的情況下,無論何處的排名一致性至少為0.92,但在十四個條件中的四個中,精確的對角線與去噪誤差呈反相關,達到-0.13,因此一個忠實的估計器會重現這一反轉。這四個都是實際影像條件;在模型自身的樣本中,這一反轉並不存在,因此在生成樣本上的評估使這一系列模型看起來更好。限制這些地圖的是目標,而不是估計器。
然後我們引入Tweedie Probe-Tangent (T-PT),這是一種無梯度的殘差探測器,反覆破壞一個模型支持的預測並測量去噪器響應的體素級變異性。T-PT讀取相同雅可比矩陣的不同函數,其精確的二階形式在對角線反轉的地方排名;在三十個探測點下,它返回的地圖不穩定到無法重現該排名,而Hutchinson在M=5的情況下已經重現了它,因此在這裡T-PT並不是反轉的證據。我們將其作為工具,而不是更好的近似。在全分辨率的腦部MRI中,我們測試的每個基於雅可比的估計器都耗盡了內存,T-PT在八個端點中的五個內部組織上引領了一個二十鏈的蒙特卡羅集成,並且在任何端點上都未落後,網絡評估數量少了16倍;在整個體積上,集成領先,而五十條鏈縮小了組織的差距。在肺部CT中,集成始終領先。兩者在變化發生的地方失去了大部分的區分能力,這一點仍然是開放的。
BIRD: Distilling Decision Boundaries into Rationales for MLLM Adaptation
2609.33713v1 by Anglin Liu, Yanlin Wu, Ruichao Chen, Yuting Zhang, Qingyuan Zeng, Pengxiang Cai, Ziqi Gong, Muchen Li, Jintai Chen
Adapting general-purpose multimodal large language models (MLLMs) to specialized domains requires learning domain-specific decision criteria, which often hinge on subtle visual distinctions between otherwise plausible answers. Rationale augmentation aims to expose such evidence through additional observations or inter-sample comparisons, yet a visually valid cue is not necessarily decision-relevant: it may describe how samples differ without changing the model's relative preference between competing answers. We therefore introduce BIRD, a self-improving Boundary-Informed Rationale Distillation framework that uses model-specific confusions to locate unresolved local decision boundaries and distills the evidence that resolves these confusions into rationales. For each sample, BIRD retrieves candidate neighbors from the target MLLM's own representation space and selects the most confusable one according to its answer preferences. It then generates answer-blind candidate evidence from their visual differences and functionally verifies which evidence most effectively strengthens the model's preference for the correct answer while avoiding inappropriate transfer across the pair. The verified evidence is then distilled into a single-sample rationale for standard supervised fine-tuning. Experiments on medical and chart VQA show that BIRD outperforms competing rationale-augmentation methods across two target MLLMs, while further analyses demonstrate clearer separation of confusable answers and stronger gains from model-matched supervision.
摘要:適應通用多模態大型語言模型(MLLMs)到專門領域需要學習特定於領域的決策標準,這通常依賴於在其他可行答案之間的微妙視覺區別。推理增強旨在通過額外的觀察或樣本間比較來揭示這種證據,然而,視覺上有效的線索不一定與決策相關:它可能描述樣本之間的差異,而不改變模型對競爭答案的相對偏好。因此,我們引入了BIRD,一個自我改善的邊界知情推理蒸餾框架,利用模型特定的混淆來定位未解決的局部決策邊界,並將解決這些混淆的證據提煉成推理。對於每個樣本,BIRD從目標MLLM自身的表示空間中檢索候選鄰居,並根據其答案偏好選擇最具混淆性的那一個。然後,它根據它們的視覺差異生成不依賴答案的候選證據,並功能性地驗證哪種證據最有效地增強模型對正確答案的偏好,同時避免在這對之間的不當轉移。經過驗證的證據然後被提煉成單樣本推理,用於標準的監督微調。在醫療和圖表VQA的實驗中,BIRD在兩個目標MLLM上超越了競爭的推理增強方法,而進一步的分析顯示出混淆答案之間更清晰的區分和來自模型匹配監督的更強增益。
PPG-LM: A Photoplethysmography-Language Model with Multi-Level Clinical Alignment
2609.33516v1 by Xiaoda Wang, Minxiao Wang, Maxwell A Xu, Patrick Langer, Kaiqiao Han, Defu Cao, Xiao Luo, Yuzhe Yang, Yan Liu, Xiao Hu, Yizhou Sun, Wei Wang, Carl Yang
Photoplethysmography (PPG) is widely recorded by clinical monitors and consumer wearables, providing a scalable source of continuous physiological information. These recordings offer an opportunity for physiological assessment at scale, but realizing this potential requires models to learn from both signal-derived physiological supervision and broader clinical context captured in electronic health records (EHRs). This involves aligning information spanning local observations, care events, and entire visits with PPG representations at corresponding temporal scales. However, existing PPG foundation models primarily rely on task-specific prediction heads, while the medical knowledge of large language models does not necessarily translate into waveform understanding. To bridge this gap, we introduce PPG-LM, the first PPG-language model family to learn physiological representations from both signal-derived supervision and broader clinical context captured in EHRs. To construct clinically grounded captions, we develop an automatic captioning pipeline that generates segment-, event-, and visit-level descriptions from signal measurements and structured EHR records. We then learn from these pairs through a two-stage framework that first establishes segment-language correspondence through contrastive learning and waveform-conditioned captioning, then extends alignment to events and visits through time-aware aggregation and temporal statement matching. Pretrained on approximately 73k hours of PPG, PPG-LM supports language-based recognition, cross-modal retrieval, and segment captioning. Experiments on MC-MED, MIMIC-III, and VitalDB show improved retrieval and caption factuality over language-model baselines and gains over PPG and time-series foundation models on multiple clinical prediction tasks.
摘要:光電容積描記法(PPG)被臨床監測器和消費者可穿戴設備廣泛記錄,提供了一個可擴展的持續生理信息來源。這些記錄為大規模的生理評估提供了機會,但實現這一潛力需要模型從信號衍生的生理監督和電子健康記錄(EHRs)中捕獲的更廣泛臨床背景中學習。這涉及到將涵蓋本地觀察、護理事件和整個就診的資訊與相應時間尺度上的PPG表示對齊。然而,現有的PPG基礎模型主要依賴於特定任務的預測頭,而大型語言模型的醫學知識並不一定能轉化為波形理解。為了彌補這一差距,我們介紹了PPG-LM,這是第一個從信號衍生的監督和EHRs中捕獲的更廣泛臨床背景中學習生理表示的PPG語言模型系列。為了構建臨床基礎的標題,我們開發了一個自動標題生成管道,從信號測量和結構化的EHR記錄中生成段落、事件和就診級別的描述。然後,我們通過一個兩階段框架從這些對中學習,首先通過對比學習和波形條件的標題建立段落-語言對應,然後通過時間感知聚合和時間語句匹配將對齊擴展到事件和就診。PPG-LM在約73,000小時的PPG上進行了預訓練,支持基於語言的識別、跨模態檢索和段落標題生成。在MC-MED、MIMIC-III和VitalDB上的實驗顯示,在語言模型基準上提高了檢索和標題的事實性,並在多個臨床預測任務上相對於PPG和時間序列基礎模型取得了增益。
Federated Multi-Modal Human Activity Recognition using Multi-Agent Reinforcement Learning
2609.33492v1 by Debasmita Dey, Tanmay Sen, Himel Mallick
Human Activity Recognition (HAR) from heterogeneous wearable sensors is fundamental to the Internet of Health Things (IoHT), supporting rehabilitation, elderly care, and smart healthcare. Existing multimodal fusion methods often assign fixed equal weights to sensor streams, overlooking differences in modality importance, acquisition cost, and sensor quality, which can vary due to movement, incorrect placement, or temporary blockage. We propose an adaptive and cost-aware multimodal HAR framework based on multi-agent reinforcement learning for centralized HAR and extend it to federated learning as FedMHAR. In the centralized setting, multimodal fusion is formulated as a cooperative Multi-Agent Reinforcement Learning (MARL) problem, where each sensing modality is assigned a PPO-based agent that learns per-sample fusion weights, enabling the model to emphasize informative modalities while down-weighting costly sensors when cheaper alternatives provide sufficient information. In the federated setting, we introduce BiFL-PPO, a bidirectional federated optimization strategy in which a server-side PPO policy learns client-specific trust weights and feeds them back to adapt local learning rates and proximal regularization. Unlike round-level optimization, BiFL-PPO uses dense batch-level rewards for more frequent feedback and stable training under heterogeneous client data. Evaluation on the MEx Rehabilitation and UTD Multimodal Human Action datasets shows that the centralized framework achieves 87.30% and 94.98% accuracy, respectively, outperforming conventional fusion methods and state-of-the-art HAR models. FedMHAR achieves 79.74% and 77.49% in the federated setting, consistently surpassing FedAvg, FedProx, FedBN, FedNova, and AdaFedProx, while providing more stable performance and reducing sensor acquisition cost.
摘要:人類活動識別(HAR)來自異質可穿戴傳感器,對健康物聯網(IoHT)至關重要,支持康復、老年護理和智慧醫療。現有的多模態融合方法通常對傳感器流分配固定的相等權重,忽略了模態重要性、獲取成本和傳感器質量的差異,這些差異可能因運動、不正確的放置或暫時的阻塞而變化。我們提出了一種基於多智能體強化學習的自適應和成本感知多模態HAR框架,用於集中式HAR,並將其擴展到聯邦學習,稱為FedMHAR。在集中式設置中,多模態融合被表述為一個合作的多智能體強化學習(MARL)問題,其中每個感知模態被分配一個基於PPO的智能體,該智能體學習每個樣本的融合權重,使模型能夠強調信息豐富的模態,同時在更便宜的替代方案提供足夠信息時降低成本傳感器的權重。在聯邦設置中,我們引入了BiFL-PPO,一種雙向聯邦優化策略,其中伺服器端的PPO策略學習客戶特定的信任權重並反饋以調整本地學習率和近端正則化。與回合級優化不同,BiFL-PPO使用密集的批次級獎勵以獲得更頻繁的反饋,並在異質客戶數據下實現穩定的訓練。在MEx康復和UTD多模態人類行動數據集上的評估顯示,集中式框架分別達到87.30%和94.98%的準確率,超越了傳統的融合方法和最先進的HAR模型。FedMHAR在聯邦設置中達到79.74%和77.49%的準確率,始終超越FedAvg、FedProx、FedBN、FedNova和AdaFedProx,同時提供更穩定的性能並降低傳感器獲取成本。
A Cheap Verifier is Good Enough: LLM Post-training is Robust to Erroneous Rewards
2609.33467v1 by Andreas Plesner, Curtis Northcutt, Francisco Guzmán, Anish Athalye
When post-training large language models on tasks with semi-verifiable rewards, there are many factors (training steps, base model size, training order, data quality, verifier accuracy, etc.) that practitioners must contend with to maximize model performance. Yet, it remains unclear how well verifier agreement predicts post-training performance on such tasks. In this paper, we explore this question with over 11k H100 GPU-hours, across HealthBench and PRBench tasks in medical, legal, and finance domains. Across the tested domains, Qwen3 trainees (1.7B-8B on HealthBench; 8B on PRBench), evaluation splits, and frontier LLM reference judges (which we call golden verifiers), higher verifier agreement does not consistently identify the best training verifier. Expensive verifiers need not outperform inexpensive ones, and open-weight Gemma verifiers produce strong training outcomes. We compare two low-cost choices retrospectively -- a cost-reducing choice and a balanced choice -- with estimated grading cost reductions of 98.8%-99.7% relative to the golden grading protocols and average post-training score gaps of 1-3 points from the best evaluated training verifier. These averages include larger losses in individual settings; they do not establish that verifier choices are interchangeable.
摘要:在對具有半可驗證獎勵的任務進行後訓練大型語言模型時,從業者必須面對許多因素(訓練步驟、基礎模型大小、訓練順序、數據質量、驗證者準確性等),以最大化模型性能。然而,目前尚不清楚驗證者的一致性在多大程度上預測此類任務的後訓練性能。在本文中,我們通過超過11,000小時的H100 GPU,探索這個問題,涵蓋醫療、法律和金融領域的HealthBench和PRBench任務。在測試的領域中,Qwen3訓練者(在HealthBench上為1.7B-8B;在PRBench上為8B)、評估拆分和前沿LLM參考評審(我們稱之為黃金驗證者)中,更高的驗證者一致性並不總是能夠準確識別最佳訓練驗證者。昂貴的驗證者不一定要優於便宜的驗證者,而開放權重的Gemma驗證者則產生了強大的訓練結果。我們回顧性比較了兩個低成本選擇——一個是降低成本的選擇,另一個是平衡的選擇——相對於黃金評分協議,估計的評分成本降低為98.8%-99.7%,而從最佳評估的訓練驗證者的平均後訓練分數差距為1-3分。這些平均數包括在個別設置中的更大損失;它們並未確立驗證者選擇是可互換的。
MAC-Net: A Multi-Task Deep Learning Framework for Modeling Cognitive Function From Task-Based fMRI
2609.33440v1 by Md. Tanvir Rahman, Nabil Anan Orka, Asaduzzaman Khan, Mohammad Ali Moni
Objective cognitive assessment from neural signals supports neurorehabilitation, but individual-level prediction from task-based fMRI (tfMRI) remains difficult because neural features coexist with substantial demographic and scanner-related variation. We present the Multi-task Activation and Contrast Network (MAC-Net), a covariate-aware deep learning framework for modeling individual cognitive function from regional tfMRI. By isolating tfMRI features into a dedicated neural pathway and restricting participant variables to a terminal late-fusion pathway, MAC-Net prevents dominant covariates from suppressing high-dimensional clinical representations during feature learning. Evaluating baseline data from 6,500 Adolescent Brain Cognitive Development Study participants under family-aware cross-validation, MAC-Net was benchmarked against linear models, random forests, and alternative deep architectures. The N-back plus Monetary Incentive Delay configuration achieved $R^{2}$ values of 0.174, 0.238, and 0.277 for fluid, crystallized, and total cognition, outperforming covariate-only baselines (0.178) and alternative deep models (0.217). N-back was the most informative paradigm, whereas incorporating the Stop Signal Task marginally degraded performance. Feature attributions via Integrated Gradients, DeepLIFT, and Input Gradient were highly concordant, localizing working-memory-related frontal, parietal, and cingulate regions. These findings demonstrate that covariate-aware multi-task modeling yields reproducible cognitive-function estimations, establishing a robust neural engineering framework for clinical translation.
摘要:目標認知評估來自神經信號,支持神經康復,但基於任務的功能性磁共振成像(tfMRI)在個體層面的預測仍然困難,因為神經特徵與顯著的人口統計和掃描儀相關變異共存。我們提出了多任務激活和對比網絡(MAC-Net),這是一個考慮協變量的深度學習框架,用於從區域tfMRI建模個體認知功能。通過將tfMRI特徵隔離到專門的神經通路並將參與者變量限制在終端的晚融合通路,MAC-Net防止主導協變量在特徵學習過程中壓制高維臨床表示。對6,500名青少年大腦認知發展研究參與者的基線數據進行家庭意識交叉驗證,MAC-Net與線性模型、隨機森林和其他深度架構進行了基準測試。N-back加上金錢激勵延遲配置在流動性、結晶性和總認知方面達到了$R^{2}$值分別為0.174、0.238和0.277,超越了僅考慮協變量的基準(0.178)和其他深度模型(0.217)。N-back是最具信息性的範式,而納入停止信號任務則輕微降低了性能。通過整合梯度、DeepLIFT和輸入梯度的特徵歸因高度一致,定位到與工作記憶相關的額葉、頂葉和扣帶區域。這些發現表明,考慮協變量的多任務建模產生可重複的認知功能估計,建立了一個穩健的神經工程框架以便於臨床轉化。
Temporal Graph Learning of Wearable Actigraphy and Sleep Traces for Modelling Adolescent Crystallized Intelligence
2609.33428v1 by Md. Tanvir Rahman, Nabil Anan Orka, Asaduzzaman Khan, Mohammad Ali Moni
Wearable actigraphy offers a scalable, ecologically valid alternative to episodic clinical assessment. However, predicting continuous adolescent crystallized intelligence ($G_c$) from such traces remains challenging due to irregular device adherence and complex behavioral-environmental interactions. We address this using daily summary data derived from 21-day Fitbit records of 6,091 adolescents in the Adolescent Brain Cognitive Development Study (Release 5.1). We propose SATURN, a Sleep-Activity Temporal Unified Regression Network. It represents participants as 21-node temporal graphs encoding daily behaviors and temporal adjacency. To prevent imputation artifacts, invalid-day edges are dynamically pruned during forward passes. Node embeddings are refined via residual GATv2 layers, aggregated through masked attention pooling, and fused with sociodemographic covariates. Under family-controlled, age-sex-BMI-stratified cross-validation, SATURN achieves $R^2 = 0.2783 \pm 0.0127$, consistently improving upon flattened machine learning (Gradient Boosting, $R^2 = 0.2372$) and sequential deep learning (BiLSTM, $R^2 = 0.2688$) baselines. Explainability analyses identify light activity, metabolic equivalents, and sleep duration as dominant predictors, while Monte Carlo dropout and subgroup analyses confirm equitable performance across sociodemographic strata. Ultimately, SATURN establishes a rigorous computational framework for digital cognitive phenotyping, offering a scalable pathway to complement traditional assessments by highlighting macro-level behavioral anomalies.
摘要:可穿戴行為測量提供了一種可擴展的、生態有效的替代方案,以取代臨床評估的偶發性。然而,從這些數據中預測持續的青少年結晶智力 ($G_c$) 仍然具有挑戰性,因為設備遵從性不規則且行為與環境之間的互動複雜。我們使用來自 6,091 名青少年在青少年大腦認知發展研究(版本 5.1)中,為期 21 天的 Fitbit 記錄所衍生的每日摘要數據來解決這個問題。我們提出了 SATURN,一個睡眠-活動時間統一回歸網絡。它將參與者表示為 21 節點的時間圖,編碼每日行為和時間相鄰性。為了防止插補伪影,在前向傳播過程中動態修剪無效日邊緣。節點嵌入通過殘差 GATv2 層進行精煉,通過遮罩注意力池化進行聚合,並與社會人口學協變量融合。在家庭控制、年齡-性別-BMI 分層的交叉驗證下,SATURN 的 $R^2 = 0.2783 \pm 0.0127$,持續優於扁平化的機器學習(梯度提升,$R^2 = 0.2372$)和序列深度學習(BiLSTM,$R^2 = 0.2688$)基準。可解釋性分析確定輕度活動、代謝當量和睡眠持續時間為主要預測因子,而蒙特卡羅隨機失活和子群分析則確認了在社會人口學層次上表現公平。最終,SATURN 建立了一個嚴謹的計算框架,用於數位認知表型,提供了一條可擴展的途徑,以通過突顯宏觀層面的行為異常來補充傳統評估。
Explainable Deep Learning of Resting-State Functional Connectomes Reveals Network Biomarkers of Adolescent Intelligence
2609.33422v1 by Md. Tanvir Rahman, Nabil Anan Orka, Asaduzzaman Khan, Mohammad Ali Moni
Mapping resting-state brain organization to individual differences in cognitive ability remains a major challenge in population neuroinformatics. Although deep learning enables flexible modeling of brain connectivity, limited interpretability restricts its scientific and clinical utility. To address this objective, we developed an explainable deep learning framework based on sparse projected residual networks to predict fluid, crystallized, and total intelligence from resting-state functional magnetic resonance imaging in 5,285 participants from the Adolescent Brain Cognitive Development study. We incorporated three complementary explainability methods (Integrated Gradients, Gradient Shapley Additive Explanations, and Occlusion) to interpret model behavior. The framework outperformed existing approaches, achieving Pearson correlations of 0.44, 0.58, and 0.56 for fluid, crystallized, and total intelligence, respectively, corresponding to predictive improvements of 6 to 9 percent. All three explainability methods produced near-identical feature rankings (pairwise rank correlations greater than 0.99). Consensus maps revealed a dual-layered functional architecture where primary predictive hubs localized within canonical systems, while the strongest global predictive pathways frequently bypassed these hubs through distributed, long-range relay connections. These findings suggest that intelligence emerges from the interaction between localized computational hubs and distributed communication pathways. Ultimately, these normative network architectures provide clinical reference maps to detect individual deviations, supporting earlier diagnosis, cognitive subtype stratification, and treatment monitoring in atypical neurodevelopment.
摘要:將靜息狀態下的大腦組織映射到個體在認知能力上的差異,仍然是人口神經資訊學中的一大挑戰。雖然深度學習使得大腦連接的靈活建模成為可能,但有限的可解釋性限制了其科學和臨床的實用性。為了達成這一目標,我們開發了一個基於稀疏投影殘差網絡的可解釋深度學習框架,從5,285名來自青少年大腦認知發展研究的參與者的靜息狀態功能性磁共振成像中預測流體智力、結晶智力和總智力。我們結合了三種互補的可解釋性方法(整合梯度、梯度沙普利加法解釋和遮蔽)來解釋模型行為。該框架的表現超過了現有的方法,對流體智力、結晶智力和總智力的皮爾森相關係數分別達到0.44、0.58和0.56,對應的預測改進為6%到9%。所有三種可解釋性方法產生了幾乎相同的特徵排名(成對排名相關係數大於0.99)。共識圖揭示了一種雙層功能架構,其中主要的預測樞紐位於典型系統內,而最強的全球預測通路則經常通過分散的長距離中繼連接繞過這些樞紐。這些發現表明,智力是由局部計算樞紐和分散通信通路之間的互動所產生的。最終,這些規範性網絡架構提供了臨床參考圖,以檢測個體偏差,支持早期診斷、認知亞型分層和在非典型神經發展中的治療監測。
QuPID: Quantum Parameter-Efficient Input-Dependent Retrieval Adaptation for Medical RAG
2609.33351v1 by Hyojun Ahn, Emily Jimin Roh, Soohyun Park, Walid Saad, Hyung-Chul Lee, Joongheon Kim
Fidelity-based quantum retrieval ranks candidates by the fidelity between query and archive states. Applying a shared input-independent unitary after fixed state encoding leaves that fidelity unchanged, so training the circuit cannot alter the ranking. Quantum parameter-efficient input-dependent retrieval adaptation (QuPID) repairs this by making the circuit input-dependent through data re-uploading and by comparing measurement readouts, vectors of local Pauli expectations, rather than states. The result is a small readout for adapting frozen image features to a local archive with limited data: training simulates the circuit classically, and inference runs on a GPU with fixed learned parameters. We characterize the class as a structured factorization of input-modulated quadratic feature maps, bound the frequency support of its re-uploading channel, and give a parameter-count generalization bound that motivates its small budget. Under a shared frozen backbone and a label-free protocol, QuPID's 60 parameters give higher precision-at-5 (P@5) on ChestX-ray14 and MURA than frozen medical encoders, and than adapters and low-rank adaptation (LoRA) with up to 5.25 million trainable parameters. On ChestX-ray14, the P@5 gain over the frozen encoder is +0.116, the lead over retuned adapters is widest at 512 adaptation examples (+0.040), and the full-budget margin over an equally compact classical rotation-plane head is +0.023 with a 95% interval excluding zero. Medical imaging is the primary testbed; the pattern recurs on two non-medical benchmarks, in report generation, and under simulated gate noise and finite-shot readout.
摘要:基於保真度的量子檢索通過查詢和檔案狀態之間的保真度對候選者進行排名。應用共享的輸入無關單位後,固定狀態編碼的保真度保持不變,因此訓練電路無法改變排名。量子參數高效的輸入依賴檢索適應(QuPID)通過數據重新上傳使電路依賴於輸入,並通過比較測量讀數、局部保利期望的向量,而不是狀態來修復這一點。結果是針對有限數據的本地檔案適應凍結圖像特徵的小型讀出:訓練在經典上模擬電路,而推理在具有固定學習參數的GPU上運行。我們將該類別表徵為輸入調製的二次特徵映射的結構因式分解,限制其重新上傳通道的頻率支持,並給出一個參數計數的泛化界限,這激勵了其小預算。在共享的凍結骨幹和無標籤協議下,QuPID的60個參數在ChestX-ray14和MURA上的精確度@5(P@5)高於凍結醫療編碼器,以及高達525萬可訓練參數的適配器和低秩適應(LoRA)。在ChestX-ray14上,與凍結編碼器相比,P@5的增益為+0.116,與重新調整的適配器相比,在512個適應示例下的領先幅度最大(+0.040),而與一個同樣緊湊的經典旋轉平面頭的全預算邊際為+0.023,95%區間不包括零。醫學影像是主要的測試平台;該模式在兩個非醫學基準、報告生成以及在模擬閘噪聲和有限次讀出下重複出現。
CHI: A Composite Hallucination Index Unifying Entity, Relation, and Quantity Dimensions for Summarization Evaluation
2609.33343v1 by Praveenkumar Katwe, Rakesh Chandra Balabantaray, Kali Prasad Vittala
Faithfulness evaluation of abstractive summaries remains an open challenge, with existing metrics addressing only isolated hallucination types: factual entity errors, relational inconsistencies, or numerical fabrications, without capturing their co-occurrence or interaction. We introduce CHI (Composite Hallucination Index), the first unified hallucination metric that decomposes faithfulness errors into three orthogonal dimensions: entity hallucination (EHI), relation hallucination (RHI*), and quantity hallucination (QHI). Each dimension employs a shared softmax-normalized architecture over Venn diagram-derived factors representing extractiveness, positive hallucination, over-focus, negative hallucination, and lost focus. The novel QHI component introduces tolerance-aware numerical matching with exact, epsilon, derived, and temporal comparison modes. We fuse the three dimensions via harmonic mean to produce a single composite score that penalizes weakness in any dimension. We validate CHI on 800 source articles spanning four domains (news, medical, legal, financial) with summaries from five generation systems. Empirical results demonstrate that: (i) the three dimensions are statistically orthogonal (mean rho = 0.148), confirming they capture distinct error types; (ii) CHI achieves the highest system-level correlation with human judgments (rho = 0.66, p = 0.006) on SummEval, outperforming ROUGE (rho = 0.53), EHI (rho = 0.58), and all individual components; and (iii) ablation studies confirm that all three dimensions contribute unique variance, with the full composite outperforming any individual component while providing decomposable error diagnostics unavailable from single-score baselines. CHI provides practitioners with a decomposable, interpretable, and efficient faithfulness metric suitable for both offline evaluation and online monitoring of summarization systems.
摘要:忠實度評估抽象摘要仍然是一個未解的挑戰,現有的指標僅針對孤立的幻覺類型:事實實體錯誤、關係不一致或數字虛構,而未捕捉它們的共現或互動。我們介紹了CHI(綜合幻覺指數),這是第一個統一的幻覺指標,將忠實度錯誤分解為三個正交維度:實體幻覺(EHI)、關係幻覺(RHI*)和數量幻覺(QHI)。每個維度都利用一個共享的softmax正規化架構,基於源自維恩圖的因素,這些因素代表了提取性、正幻覺、過度關注、負幻覺和失焦。新穎的QHI組件引入了容忍度感知的數字匹配,具有精確、epsilon、衍生和時間比較模式。我們通過調和平均將這三個維度融合,產生一個單一的綜合分數,對任何維度的弱點進行懲罰。我們在涵蓋四個領域(新聞、醫療、法律、金融)的800篇來源文章上驗證了CHI,這些文章的摘要來自五個生成系統。實證結果顯示:(i)這三個維度在統計上是正交的(平均rho = 0.148),確認它們捕捉到不同的錯誤類型;(ii)CHI在SummEval上達到與人類評價的最高系統級相關性(rho = 0.66,p = 0.006),超越了ROUGE(rho = 0.53)、EHI(rho = 0.58)和所有個別組件;以及(iii)消融研究確認這三個維度都貢獻了獨特的變異性,完整的綜合指標優於任何個別組件,同時提供了從單一分數基準無法獲得的可分解錯誤診斷。CHI為實踐者提供了一個可分解、可解釋且高效的忠實度指標,適用於離線評估和在線監控摘要系統。
The Error You See Is Not the Error You Made: Progression-aware Reasoning Origin for Reasoning Error Localization
2609.33297v1 by Yiguo Wang, Ziyuan Yang, Yi Zou, Dan Lin, Rongsheng Li, Yi Zhang
Verifying multi-step LLM reasoning requires more than determining whether a trace is correct: a useful verifier should identify where the reasoning first goes wrong. However, existing holistic methods provide little positional evidence, while forward sequential verification often treats the first rejected step as the error source. Under error propagation, this assumption can fail, since an earlier mistake may remain locally plausible and become observable only through its downstream consequences. We therefore rethink reasoning verification as a progression-aware error-source localization problem: rather than asking only where a reasoning trace first appears inconsistent, we ask which earlier step best explains how that inconsistency emerges along the trajectory. Based on this view, we propose Progression-aware Reasoning Origin (PRO), a training-free framework for first-error localization. PRO jointly models incoming support from the preceding context and outgoing compatibility with subsequent reasoning, selectively refines regions where these signals disagree, and finally performs detector-conditioned source attribution with intervention-based evidence to distinguish the true error origin from its propagated manifestations. We further formalize the gap between forward rejection and structural exposure, showing why incoming-side evidence alone is insufficient for reliable localization under error propagation. Experiments across open-form, medical, and structured reasoning tasks demonstrate consistent improvements over strong verification baselines, supporting progression-aware source attribution as a more faithful formulation of reasoning verification.
摘要:驗證多步驟 LLM 推理不僅需要確定一個痕跡是否正確:一個有用的驗證器應該能夠識別推理首次出錯的地方。
然而,現有的整體方法提供的位置信息有限,而前向序列驗證通常將第一個被拒絕的步驟視為錯誤來源。在錯誤傳播的情況下,這一假設可能會失效,因為早期的錯誤可能在局部上仍然是合理的,並且只有通過其下游後果才能被觀察到。
因此,我們重新思考推理驗證,將其視為一個進程感知的錯誤來源定位問題:我們不僅詢問推理痕跡首次出現不一致的地方,而是詢問哪一個早期步驟最能解釋沿著軌跡出現的不一致。
基於這一觀點,我們提出了進程感知推理來源(PRO),這是一個無需訓練的首錯定位框架。
PRO 共同建模來自前一上下文的支持和與後續推理的兼容性,選擇性地細化這些信號不一致的區域,並最終通過基於干預的證據進行檢測器條件的來源歸因,以區分真實的錯誤來源和其傳播的表現。
我們進一步形式化了前向拒絕和結構曝光之間的差距,顯示為什麼僅依賴來自進入側的證據對於在錯誤傳播下的可靠定位是不足夠的。
在開放式、醫療和結構化推理任務中的實驗顯示出對強驗證基準的一致改進,支持進程感知來源歸因作為推理驗證的更真實表述。
CORTEX: A Verified Experience Layer for Generalist Agents
2609.33260v1 by Garapati Keerthana, Manik Gupta
An agent can solve a task today and face the same task under new facts, tools, or governing knowledge tomorrow. Most agent systems can retrieve relevant text or recall prior conversations, but they lack a principled way to decide when a previous solution is still valid, when it must be adapted, and when it should be discarded. We introduce CORTEX (Contextual Orchestration and Reuse of Task EXperience), a general AI systems framework that connects specialized agents through an external layer of verified experience. Each episode records its task conditions, source and tool state, decisive predicates, proof trace, verifier, and outcome. A meta-controller chooses exact replay, checked adaptation, fresh synthesis, or escalation. Accepted episodes can become task patterns and procedural strategies through a challenge-driven development loop. This gives the system an implicit competence layer that can grow without changing model weights. We formalize system contracts for exact replay and source-version separation, and derive when reuse saves computation. A controlled two-domain implementation tests the exact-replay core on 1,000 synthetic cases. Complete-family holdouts test procedural transfer on 1,000 new-family cases across eight clinical and policy splits, with complete fresh-evidence grounding and perfect invariance to irrelevant-field and insertion-order perturbations. The transfer trace exposes the work required for verified strategy execution. These results establish an initial path toward general intelligence through reusable procedures, typed experience, and developmental transfer.
摘要:一個代理可以在今天解決一個任務,並在明天面對同一任務,但有新的事實、工具或治理知識。大多數代理系統可以檢索相關文本或回憶先前的對話,但它們缺乏一種原則性的方式來決定何時先前的解決方案仍然有效,何時必須進行調整,以及何時應該被丟棄。我們介紹了 CORTEX(上下文協調與任務經驗重用),這是一個通用的 AI 系統框架,通過一層經過驗證的經驗將專門的代理連接起來。每個事件記錄其任務條件、來源和工具狀態、決定性謂詞、證明痕跡、驗證者和結果。一個元控制器選擇精確重播、檢查調整、新的綜合或升級。接受的事件可以通過挑戰驅動的開發循環轉變為任務模式和程序策略。這為系統提供了一個隱含的能力層,能夠在不改變模型權重的情況下增長。我們為精確重播和來源版本分離形式化了系統合同,並推導出何時重用可以節省計算。受控的雙域實施在 1,000 個合成案例上測試精確重播核心。完整家庭保留測試在八個臨床和政策拆分中對 1,000 個新家庭案例的程序轉移,具有完整的新證據基礎和對無關領域及插入順序擾動的完美不變性。轉移痕跡揭示了執行經過驗證的策略所需的工作。這些結果為通過可重用程序、類型化經驗和發展轉移建立了一條通向通用智能的初步路徑。
FOCUS: Benchmarking Retinal Model Generalization from Foundation Vision Encoders to Multimodal LLMs
2609.33158v1 by David Restrepo, Chenwei Wu, Luis Filipe Nakayama, Miguel L. Martins, Stergios Christodoulidis, Maria Vakalopoulou, Enzo Ferrante
Progress in AI-based retinal image analysis has advanced with foundation models, yet evaluating their reliability remains challenging. Performance reported on a single dataset does not capture how models behave under dataset shift, across clinical definitions, or for different patient subgroups. This limitation is particularly critical in medical imaging analysis, where robustness, calibration, and fairness are essential for safe deployment. We introduce FOCUS (Foundation Ophthalmic Cross-Dataset Understanding under Shift), a cross-dataset benchmark for evaluating retinal fundus models that considers vision-only encoder models (VM), vision-language dual-encoder models (VLM), and multimodal large language models (MLLM). FOCUS harmonizes binary diabetic retinopathy, referable diabetic retinopathy, and glaucomatous optic neuropathy tasks across ten public datasets spanning diverse geographies, acquisition conditions, and label protocols. The benchmark evaluates models through a unified analysis layer that measures ranking performance, calibration, subgroup disparities, and image-quality robustness. We present a large-scale evaluation covering 532 base configurations and 228 MLLM configurations adapted through supervised fine-tuning with low-rank adaptation (LoRA). Results show that no model family consistently dominates across tasks and datasets: general VM encoders achieve the strongest average ranking performance, medical MLLMs are competitive but variable, and dual encoder VLMs benefit substantially from lightweight adaptation. Fine-tuning improves in-domain performance but exhibits heterogeneous transfer to external datasets, particularly in calibration. These findings demonstrate that retinal model evaluation is inherently multidimensional. FOCUS provides a practical framework and public benchmark to assess generalization, reliability, and robustness beyond single-dataset leaderboards
摘要:進展於基於人工智慧的視網膜影像分析已隨著基礎模型的發展而提升,然而評估其可靠性仍然具有挑戰性。單一數據集上報告的性能無法捕捉模型在數據集轉移、臨床定義之間或不同患者子群體中的行為。這一限制在醫學影像分析中特別關鍵,因為穩健性、校準和公平性對於安全部署至關重要。我們引入了FOCUS(Foundation Ophthalmic Cross-Dataset Understanding under Shift),這是一個跨數據集基準,用於評估視網膜眼底模型,考慮了僅視覺編碼器模型(VM)、視覺-語言雙編碼器模型(VLM)和多模態大型語言模型(MLLM)。FOCUS在十個公共數據集上協調二元糖尿病視網膜病變、可參考糖尿病視網膜病變和青光眼性視神經病變任務,這些數據集涵蓋了多樣的地理位置、獲取條件和標籤協議。該基準通過一個統一的分析層評估模型,測量排名性能、校準、子群體差異和影像質量的穩健性。我們呈現了一個涵蓋532個基本配置和228個經過低秩適應(LoRA)監督微調的MLLM配置的大規模評估。結果顯示,沒有任何模型家族在任務和數據集上始終佔據主導地位:一般的VM編碼器實現了最強的平均排名性能,醫學MLLM在競爭中但變化不定,而雙編碼器VLM在輕量適應中受益匪淺。微調改善了內域性能,但在外部數據集上展現出異質的轉移,特別是在校準方面。這些發現表明,視網膜模型評估本質上是多維的。FOCUS提供了一個實用的框架和公共基準,以評估超越單一數據集排行榜的泛化、可靠性和穩健性。
MedRouter: Demystifying Knowledge Differences Across Medical LLMs for Routing-Based Reasoning
2609.33119v1 by Lang Cao, Binghang Lu, Yuhao Shen, Yue Guo
Medical question answering spans diverse specialties and modalities, and individual medical large language models (LLMs) exhibit distinct strengths across tasks and domains. This heterogeneity suggests that combining specialists may enable broader coverage of medical questions than relying on any single model. However, existing LLM routing methods primarily seek to balance answer quality and inference cost, leaving open how to exploit differences in specialist competence to improve medical reasoning. In this paper, we introduce MedRouter, an agentic system that uses an embedding-based multi-label router to select and query specialist LLMs, then passes their responses to a generator to produce the final answer. We further propose SCALE (Specialist Competence-Aware Learning), a two-stage training framework that first trains the Router with specialist correctness supervision and then optimizes its selections through reinforcement learning. The second stage uses a Performance Gain Reward (PGR) that measures how specialist information affects the generator's answer correctness relative to answering without that information. Experiments on eight text-based and multimodal medical QA benchmarks show that MedRouter outperforms the strongest routing baseline by 8% in average accuracy. Our analysis of specialist outputs further reveals distinct strengths and complementary question-level coverage, motivating learned routing to combine these capabilities for more comprehensive medical reasoning.
摘要:醫療問題回答涵蓋多樣的專科和模式,而個別醫療大型語言模型(LLMs)在任務和領域上展現出不同的優勢。這種異質性表明,結合專家可能比依賴任何單一模型更能廣泛覆蓋醫療問題。然而,現有的LLM路由方法主要尋求平衡答案質量和推理成本,尚未探討如何利用專家的能力差異來改善醫療推理。在本文中,我們介紹了MedRouter,一個使用基於嵌入的多標籤路由器來選擇和查詢專家LLMs的代理系統,然後將它們的回應傳遞給生成器以產生最終答案。我們進一步提出了SCALE(專家能力感知學習),這是一個兩階段的訓練框架,首先用專家的正確性監督來訓練路由器,然後通過強化學習來優化其選擇。第二階段使用性能增益獎勵(PGR),衡量專家信息如何影響生成器的答案正確性,相對於在沒有該信息的情況下回答。對八個基於文本和多模態的醫療QA基準的實驗顯示,MedRouter在平均準確率上比最強的路由基線高出8%。我們對專家輸出的分析進一步揭示了不同的優勢和互補的問題級別覆蓋,促使學習路由結合這些能力以實現更全面的醫療推理。
ECG-Scroll: A Long-Horizon, Streaming Benchmark and Agent Environment for Interpretation of Ambulatory Electrocardiograms
2609.33117v1 by Haitao Li, Chenglin Li, Zhengyao Ding, Ziyu Li, Yiheng Mao, Zhengxing Huang
Multimodal large language models (MLLMs) can now interpret a standard ten-second, twelve-lead electrocardiogram (ECG) with clinically grounded, reward-verified reasoning. Real cardiac monitoring is different. Ambulatory (Holter) and telemetry recordings span hours to days and are read as they stream in, and their clinically decisive findings are paroxysmal, brief episodes buried in an otherwise unremarkable trace. Such a recording cannot be held in one context at diagnostic resolution, and its future has not yet happened, so a reader must work online, deciding what to measure now, committing evidence to memory as it passes, and reporting events as they occur. We recast long-duration ECG interpretation as a long-horizon, online (streaming, causal) sequential decision process and introduce ECG-Scroll. As a benchmark, long ambulatory recordings are streamed to an agent chunk by chunk, and it must localize, quantify, and promptly flag paroxysmal events without access to future signal; because the underlying signal is retained, every answer is checkable against objective ground truth, giving rule-based rather than judge-based rewards, and the streaming formulation adds a metric batch evaluation cannot express, the detection latency between an event's onset and the moment the agent records it. As an agent environment, it is a fixed, gym-style interaction layer that exercises three competencies single-glance ECG models never touch: Memory, Tool use through signal-grounded measurement rather than reading pixels, and Planning of what to measure now and when to commit. We release 390 whole-recording instances spanning 2,536 hours of two-lead ambulatory ECG and evaluate a signal-threshold rule agent alongside off-the-shelf LLM agents online, characterizing how they use memory, tools, and planning and where the benchmark's head-room lies.
摘要:多模態大型語言模型(MLLMs)現在可以以臨床為基礎、經過獎勵驗證的推理來解釋標準的十秒鐘、十二導程心電圖(ECG)。真正的心臟監測則有所不同。動態(Holter)和遙測記錄的時間跨度從幾小時到幾天,並且在流入時即時閱讀,其臨床決定性發現是突發的、短暫的事件,埋藏在其他看似無異常的波形中。這樣的記錄無法在診斷解析度下保持在一個上下文中,且其未來尚未發生,因此讀者必須在線工作,決定現在要測量什麼,將證據在經過時記憶,並在事件發生時報告。我們將長時間心電圖解釋重新構建為一個長期的、在線(流式、因果)序列決策過程,並引入ECG-Scroll。作為基準,長時間的動態記錄以塊的形式流式傳輸給代理,代理必須在無法訪問未來信號的情況下定位、量化並迅速標記突發事件;因為基礎信號被保留,每個答案都可以與客觀真實進行檢查,提供基於規則而非基於評判的獎勵,而流式的表述增加了一個度量,批量評估無法表達,即事件開始與代理記錄之間的檢測延遲。作為代理環境,它是一個固定的、健身房風格的互動層,鍛煉三種單次視覺心電圖模型從未接觸的能力:記憶、通過信號基礎測量而非讀取像素的工具使用,以及現在要測量什麼和何時提交的計劃。我們釋放了390個完整記錄實例,涵蓋2536小時的雙導程動態心電圖,並在線評估一個信號閾值規則代理,與現成的LLM代理進行比較,描述它們如何使用記憶、工具和計劃,以及基準的潛力所在。
SemReward-VL: Semantic Reward-Guided Video-Language Adaptation for Developmental Behavior Assessment
2609.33082v1 by De Jiang, Shuo Zhang, Kehong Yuan, Hongen Liao
Developmental screening videos show how children perform specific behaviors, but clinical records usually contain outcomes rather than descriptions of what happened. We present SemReward-VL, which learns to describe item-specific behavior from these outcomes. A vision-language model generates a description, and a frozen language model scores its agreement with the clinical outcome, relevance to the item, abstention on unrelated video-item pairs, and clarity. Group relative policy optimization (GRPO) updates LoRA adapters using this semantic reward. On 13,379 videos covering 41 items, the method improves aggregate accuracy and the number of items with recall above 0.5. Errors remain in temporal direction, duration, and age-specific interpretations of behavior.
摘要:發展性篩檢影片顯示兒童如何執行特定行為,但臨床記錄通常包含結果而非發生的描述。我們提出SemReward-VL,該模型從這些結果中學習描述特定項目的行為。一個視覺-語言模型生成描述,而一個凍結的語言模型評分其與臨床結果的協議、對項目的相關性、對不相關的視頻-項目對的避免,以及清晰度。群體相對政策優化(GRPO)使用這一語義獎勵更新LoRA適配器。在涵蓋41個項目的13,379個視頻上,該方法提高了總體準確性和召回率超過0.5的項目數量。錯誤仍然存在於時間方向、持續時間和年齡特定的行為解釋中。
NutriVision: Ingredient-Conditioned Fusion and Prediction for Single-Image Food Nutrition Estimation
2609.33076v1 by Aman Kumar, Avinash Anand, Chaitanya Lakhchaura, Ashutosh Kumar, Akshita Abrol, Timothy Liu, Zhengkui Wang, Rajiv Ratn Shah
Nutrition estimation is a fundamental task in consumer diet tracking, clinical dietetics, chronic disease management, sports and hospital nutrition, and broader food computing systems. The existing approaches have progressed along two largely separate axes, vision models that rely on calibrated RGB-depth captures and ingredient-aware methods that use textual cues but use limited multimodal fusion. We introduce NutriVision, an end-to-end framework that leverages visual geometry and ingredient semantics to estimate calories, mass, fat content, carbohydrates, and protein from a single RGB image and an optional ingredient list. It obtains the unavailable depth modality using DepthAnything-V3 and encodes ingredient descriptions using CLIP. It integrates three complementary mechanisms: (1) an \emph{Ingredient-Conditioned Frequency-Aligned Fusion Module (IC-FAFM)}, which uses textual guidance to reweight and align RGB-depth frequency components; (2) an \emph{Ingredient-Aware Mask-based Prediction Head (IA-MPH)}, whose gating and channel masks are conditioned on food identity; and (3) modality-specific \emph{Internal Semantic Modeling (ISM)} blocks. On the Nutrition5k dataset, NutriVision achieves a mean PMAE of $\mathbf{13.60\pm0.10\%}$, outperforming our IGSMNet implementation by $0.89$ percentage points and OmniFood8k by $2.90$ percentage points (both $p<0.001$). The module-level ablations identify the ingredient-aware prediction head as the primary architectural contributor, improving mean PMAE by $1.50\pm0.17$ percentage points ($p<0.001$). These results demonstrate that ingredient-conditioned prediction and frequency-aware RGB-depth fusion provide measurable gains for single-image nutrient estimation. More broadly, NutriVision offers a practical route toward nutrition-assessment systems that exploit geometric and semantic cues without requiring specialized depth-sensing hardware
摘要:營養估算是消費者飲食追蹤、臨床營養學、慢性疾病管理、運動及醫院營養以及更廣泛的食品計算系統中的一項基本任務。現有的方法主要沿著兩個相對獨立的方向發展,一是依賴經過校準的RGB-深度捕捉的視覺模型,二是使用文本線索但進行有限的多模態融合的成分感知方法。我們介紹了NutriVision,一個端到端的框架,利用視覺幾何和成分語義從單一RGB圖像和可選的成分列表中估算卡路里、質量、脂肪含量、碳水化合物和蛋白質。它使用DepthAnything-V3獲取缺失的深度模態,並使用CLIP編碼成分描述。它整合了三個互補機制:(1)\emph{成分條件頻率對齊融合模塊(IC-FAFM)},利用文本指導重新加權和對齊RGB-深度頻率分量;(2)\emph{成分感知基於掩碼的預測頭(IA-MPH)},其閘控和通道掩碼依賴於食物身份;以及(3)特定模態的\emph{內部語義建模(ISM)}模塊。在Nutrition5k數據集上,NutriVision達到了$\mathbf{13.60\pm0.10\%}$的平均PMAE,超越了我們的IGSMNet實現$0.89$個百分點和OmniFood8k的$2.90$個百分點(均為$p<0.001$)。模塊級的消融實驗確定成分感知預測頭是主要的架構貢獻者,將平均PMAE提高了$1.50\pm0.17$個百分點($p<0.001$)。這些結果表明,成分條件預測和頻率感知的RGB-深度融合為單圖像營養素估算提供了可測量的增益。更廣泛地說,NutriVision提供了一條實用的路徑,朝向利用幾何和語義線索的營養評估系統,而無需專門的深度感測硬體。
TCMQA: A 38K-Question Traditional Chinese Medicine Benchmark with a Licensed-Practitioner Reference
2609.33014v1 by Tzu-Heng Huang, Jet Lin, Eric Lin
Medical benchmarks for language models are built almost entirely on Western biomedicine. Traditional Chinese Medicine (TCM) is a separate system, with its own diagnostic framework and its own literature, and it remains largely unmeasured. The few TCM evaluations that exist are small, narrow, and rarely paired with a human reference. We present TCMQA, an open benchmark of 38,279 questions from Chinese TCM licensing examinations, paired with 15,151 responses from 101 licensed practitioners. We evaluate 29 instruction-tuned models from 9 families, spanning 0.27B to 14.8B parameters. Accuracy ranges over 59 points, and no model approaches saturation. Pretraining data predicts TCM ability far better than scale: a 12B Western-pretrained model reaches 39.6%, while a Chinese-pretrained model an eighth its size reaches 60.8%. Nine models exceed the practitioner majority vote of 64.9%, the best by 21.8 points, and all nine come from that same Chinese-pretrained family. Yet difficulty does not transfer between models and practitioners: accuracy is flat across practitioner-rated difficulty, item-level agreement is near zero for all 29 models, and on $8.4\%$ of items the practitioners are correct where the leading model is wrong. We release the corpus, the practitioner responses, the harness, and per-item model outputs at https://huggingface.co/datasets/TechTCM/TCMQA.
摘要:醫療基準對於語言模型幾乎完全建立在西方生物醫學之上。傳統中醫(TCM)是一個獨立的系統,擁有自己的診斷框架和文獻,並且仍然在很大程度上未被測量。現有的少數中醫評估都很小且狹窄,且很少與人類參考相配對。我們提出了TCMQA,一個包含38,279個來自中國中醫執業考試問題的開放基準,並配有101名執業者的15,151個回答。我們評估了來自9個家族的29個指令調整模型,參數範圍從0.27B到14.8B。準確率的範圍超過59個點,且沒有模型接近飽和。預訓練數據對中醫能力的預測遠優於規模:一個12B的西方預訓練模型達到39.6%,而一個大小僅為其八分之一的中國預訓練模型達到60.8%。九個模型超過執業者的多數票64.9%,最佳模型超出21.8個點,且這九個模型均來自同一個中國預訓練家族。然而,難度在模型和執業者之間並不轉移:在執業者評定的難度上,準確率持平,29個模型在項目層面的協議接近於零,且在$8.4\%$的項目中,執業者正確而領先模型錯誤。我們在 https://huggingface.co/datasets/TechTCM/TCMQA 釋出語料庫、執業者回應、工具以及每個項目的模型輸出。
Multimodal LLMs Outperform Pathology Foundation Models in Cross-Domain Histological Similarity
2609.32876v2 by Yishu Zhang, Yun Li, Daiwei Zhang
State-of-the-art pathology foundation models, trained on millions of histology tiles, can fail to preserve tissue similarity when comparisons cross slide or institution boundaries. We show that general-purpose multimodal LLMs, without being trained as pathology foundation models, consistently outperform these specialized models in cross-domain histological similarity judgments. Using a relative similarity framework that we release as the MOSAIC (Model Similarity Assessment across Institutions and Cohorts) benchmark, we evaluate 17 models across 6 datasets and find that pathology encoders often rank same-institution, different-disease tiles as more similar than same-disease, different-institution tiles, a clinically dangerous failure mode invisible to standard within-domain evaluations. LLMs appear less susceptible to this failure, likely because they perform semantic visual comparison of morphology and tissue architecture rather than relying on shortcut features tied to acquisition context. Scaling training data does not resolve the problem for pathology encoders, implicating the learning objective rather than data coverage. Our results expose a fundamental robustness gap in current pathology foundation models and establish multimodal LLMs as a viable alternative for cross-institutional retrieval, dataset harmonization, and multi-site quality control. Code and data will be released upon acceptance.
摘要:最先進的病理基礎模型,經過數百萬個組織學切片的訓練,當比較跨越切片或機構邊界時,可能無法維持組織的相似性。我們顯示,通用的多模態 LLM,未經訓練為病理基礎模型,始終在跨領域的組織學相似性判斷中超越這些專門模型。我們使用一個相對相似性框架,並將其發布為 MOSAIC(跨機構和隊列的模型相似性評估)基準,評估 17 個模型在 6 個數據集上的表現,發現病理編碼器經常將同一機構、不同疾病的切片評為比同一疾病、不同機構的切片更相似,這是一種臨床上危險的失敗模式,對標準的領域內評估來說是不可見的。LLM 似乎對這種失敗的敏感性較低,這可能是因為它們執行形態和組織結構的語義視覺比較,而不是依賴於與獲取上下文相關的捷徑特徵。擴大訓練數據並未解決病理編碼器的問題,這表明學習目標而非數據覆蓋是關鍵。我們的結果揭示了當前病理基礎模型中的基本穩健性差距,並確立了多模態 LLM 作為跨機構檢索、數據集協調和多地點質量控制的可行替代方案。代碼和數據將在接受後發布。
Counterfactual Self-Evolving Agents for Evidence-Grounded Reasoning
2609.32870v1 by Xing Han, Yuxin Wang, Chen Chen, Wei Dai, Gautham Krishna Gudur, Shijun Li, Hsing-Huan Chung, Gregory D. Hager, Joydeep Ghosh, Paul Pu Liang, Suchi Saria
Self-play proposer--solver methods improve reasoning by generating tasks and learning from verified solutions. However, for evidence-identifiable tasks, where case-specific evidence and domain knowledge determine a checkable answer, self-play requires generating plausible cases whose answers can be independently verified. We introduce counterfactual self-evolution, which generates counterfactual context for reconsidering the original case. A trainable Proposer constructs targeted evidence edits and describes potential outcome changes with causal explanations. We handcraft an expert-verified counterfactual instruction-tuning dataset to teach the Proposer to generate high-quality counterfactuals across a broad range of action--outcome scenarios. Each counterfactual instruction-tuning example specifies an edit within a defined category and explains its hypothesized causal effect on the decision, teaching the Proposer to reason systematically about what changes and why. We instruction-tune the Proposer on these examples, then formulate a fine-tuning reward that integrates feedback from the Solver and Verifier. Across diverse counterfactual scenarios, this reward favors high-quality counterfactuals and warranted revisions, while penalizing changes that overturn correct decisions. The counterfactual context aims to correct errors and strengthen confidence in correct decisions. Accepted counterfactuals accumulate in memory that supplies in-context evidence to the frozen Solver; the Solver adapts through evolving context rather than weight updates. We apply the framework to clinical reasoning, fact verification, and business reasoning. Our evaluation tracks performance over successive rounds as counterfactual memory grows, including transfer to harder cases. Our method achieves superior results across diverse frontier models.
摘要:自我對弈提議者--解決者方法通過生成任務並從經過驗證的解決方案中學習來改善推理。然而,對於可識別證據的任務,其中案例特定的證據和領域知識決定了可檢查的答案,自我對弈需要生成可以獨立驗證的合理案例及其答案。我們引入了反事實自我演化,該方法生成反事實背景以重新考慮原始案例。一個可訓練的提議者構建針對性的證據編輯,並用因果解釋描述潛在的結果變化。我們精心製作了一個專家驗證的反事實指令調整數據集,以教導提議者在廣泛的行動--結果場景中生成高質量的反事實。每個反事實指令調整示例指定了一個在定義類別內的編輯,並解釋其對決策的假設因果效應,教導提議者系統性地推理什麼變化以及為什麼變化。我們在這些示例上對提議者進行指令調整,然後制定一個微調獎勵,該獎勵整合了解決者和驗證者的反饋。在多樣的反事實場景中,這個獎勵偏好高質量的反事實和合理的修訂,同時懲罰推翻正確決策的變更。反事實背景旨在糾正錯誤並增強對正確決策的信心。被接受的反事實在記憶中累積,為凍結的解決者提供上下文證據;解決者通過演變的背景而不是權重更新來適應。我們將該框架應用於臨床推理、事實驗證和商業推理。我們的評估跟踪隨著反事實記憶增長而進行的多輪性能,包括轉移到更困難的案例。我們的方法在多樣的前沿模型中取得了優越的結果。
FinancialAuditBench: Benchmark Construction under Differential Privacy Using Real-World Priors
2609.32835v1 by Jerry Huang, Sarvesh Babu, Matt Van Buren, Alexander Wang, Pranav Pillai, Arush Jain, James P. Burton, Julia Hockenmaier
As AI agents are becoming widely adopted in the financial services industry, careful measurement is essential to understand where they can be reliably deployed and where oversight and professional review remain necessary. Such measurement, however, is constrained by limited access to proprietary or privacy-sensitive data. Existing benchmarks therefore often rely on publicly available data, human- and/or LLM-authored tasks, or simplified settings. We introduce FinancialAuditBench, a benchmark for evaluating agents on financial statement audit tasks, along with a framework for systematically generating synthetic engagements. Our task generation framework leverages differentially private aggregate statistics from historical audits along with audit expertise contributed through over 1,100 hours of benchmark development and review. FinancialAuditBench consists of 90 tasks spanning workpaper completion and review across six synthetic audit engagements, each containing an average of 179 files. Evaluation on eleven frontier models shows that while agents complete substantial portions of staff-level audit tasks well, they sometimes perform inappropriate procedures or produce incorrect documentation. Beyond financial auditing, our framework offers an approach for systematically generating synthetic tasks for model evaluation and training in privacy-sensitive domains.
摘要:隨著 AI 代理在金融服務行業的廣泛採用,仔細的測量對於理解它們可以可靠部署的地方以及何處仍需監督和專業審查至關重要。然後,這種測量受到對專有或隱私敏感數據的有限訪問的限制。因此,現有的基準通常依賴於公開可用數據、人類和/或 LLM 編寫的任務或簡化的設置。我們介紹了 FinancialAuditBench,一個用於評估代理在財務報表審計任務上的基準,以及一個系統生成合成參與的框架。我們的任務生成框架利用了來自歷史審計的差分隱私聚合統計數據,以及通過超過 1,100 小時的基準開發和審查貢獻的審計專業知識。FinancialAuditBench 包含 90 個任務,涵蓋六個合成審計參與的工作文件完成和審查,每個參與平均包含 179 個文件。對十一個前沿模型的評估顯示,儘管代理能夠很好地完成大量的員工級審計任務,但有時它們會執行不當的程序或產生不正確的文件。除了財務審計,我們的框架還提供了一種系統生成合成任務的方法,用於在隱私敏感領域進行模型評估和訓練。