Skip to content

Medical

Medical

Publish Date Title Authors Homepage Code
2026-08-18 MotoSafety: Edge-AI with Learned Temporal Importance for Two-Wheeler Collision Risk Assessment Under Time Pressure Sumit S. Shevtekar et.al. 2608.17823v1 null
2026-08-18 LLMs for Medical Consultation Are Evaluated Too Late: The Preformulation Gap Yining Hua et.al. 2608.17330v1 null
2026-08-18 Delta2Gamma: Band-Wise Adaptive Contrastive Learning of EEG for Alzheimer's Disease Detection Chanwoo Park et.al. 2608.17231v1 null
2026-08-17 A decodability criterion predicts when hidden-state selection beats majority voting in large language models Zhixiang wang et.al. 2608.17124v1 null
2026-08-17 Foundation Agents Meet Agentic Deep Research: Evidence-Grounded Clinical Code Forecasting Junda Wang et.al. 2608.17075v1 null
2026-08-17 Institution-Specific LLM Prompting Recovers PHI That De-identification Systems and Their Gold Standards Both Miss Daniel Palacios et.al. 2608.17051v1 null
2026-08-17 Policy Iteration with Human Feedback: Bringing Post-Training RL to In-context Learning Minh-Ha Nguyen et.al. 2608.16831v1 null
2026-08-17 Historical Backtesting for Scientific Question Discovery: A Protocol and Astronomy Pilot Hui Mao et.al. 2608.16795v1 null
2026-08-17 Unsupervised Anomaly Detection for Image Dataset Quality Assurance in Multi-Center Breast MRI Chiara Tappermann et.al. 2608.16725v1 null
2026-08-17 Toward Better Assessment of LLMs' Performance in Clinical Error Detection Yifan Zhang et.al. 2608.16643v1 null
2026-08-17 Degradation-Aligned Self-Supervised Learning for State of Health Estimation of Lithium-Ion Batteries under Label Sparsity Jiaqi Yao et.al. 2608.16612v1 null
2026-08-17 CACSurv: Concordance-Aligned Comparative Learning with Large Language Models for Cancer Survival Prediction Tianqi Xiang et.al. 2608.16594v1 null
2026-08-17 Large language models as synthetic clinical experts to inform longitudinal rare-disease modeling Clemens Schächter et.al. 2608.16507v1 null
2026-08-17 Offline Reinforcement Learning for Hemodynamic Management of Sepsis in the ICU: a MIMIC-IV Study with Dual Off-Policy Evaluation Marc Pérez-Roig et.al. 2608.16482v1 null
2026-08-17 Adaptive Post-Processing Drives Instance-Level Detection in Stroke Lesion Segmentation Qinghui Liu et.al. 2608.16377v1 null
2026-08-17 Foresight-England: Development of a National-Scale Generative AI Model of Electronic Health Records for Medical Event Prediction across the COVID-19 Pandemic Simon Ellershaw et.al. 2608.16273v1 null
2026-08-17 A cross-modal generative model for incomplete and degraded prostate MRI with multicentre clinical validation Siyuan Ma et.al. 2608.16233v1 null
2026-08-17 BaT: Towards Self-Evolving Medical Research Agent with Stage Rubrics Junqi Liu et.al. 2608.16211v1 null
2026-08-17 Picking the Right Image to Classify: Reliable-Input Selection in Teledermatology Fabian Gröger et.al. 2608.16198v1 null
2026-08-17 TokenSTFormer: A Tokenized Spatial-temporal Attention Model for Holistic Motion Analysis in Adolescent Idiopathic Scoliosis Screening Dong Chen et.al. 2608.16122v1 null
2026-08-17 Decoupling Parcellation from Classification: Systematic Benchmark of Fast Brain Segmentation Methods for Alzheimer's Disease Detection Jiadao Zou et.al. 2608.16039v1 null
2026-08-16 Breaking and Defending LLM-Powered Social Media Bot Detection Systems Nof Orenstein et.al. 2608.15893v1 null
2026-08-16 Characterising cardiac tissue properties with graph neural networks Ching-En Chiu et.al. 2608.15843v1 null
2026-08-16 PLeDO: Pain Level Detection for Osteoarthritis from EMR Data Yuhao Chen et.al. 2608.15719v1 null
2026-08-16 Integrating Persuasion Theory into the Epidemiological Modelling of Health Misinformation Spread on Social Media Mkululi Sikosana et.al. 2608.15689v1 null
2026-08-16 From Generalist to Specialist: A Context-Fusion Framework for Endoscopic Polyp Reporting with a Frozen VLM Ruijie Yang et.al. 2608.15580v1 null
2026-08-16 EA-LiteUNet: An Edge-Adaptive and Resource-Efficient U-Net for Boundary-Sensitive Dermoscopic Image Segmentation Wang Jiangtao et.al. 2608.15537v1 null
2026-08-15 Gated Against One Model, Open to the Next: Option-Only Solvability in Legal Multiple-Choice Benchmarks Volodymyr Ovcharov et.al. 2608.15428v1 null
2026-08-15 ETHOS: Towards a Modular Ethics Framework for Clinical Multi-Agent Systems Rakesh Sharma et.al. 2608.15424v1 null
2026-08-15 Invariant Pretraining for Robust Code Representations Yifeng He et.al. 2608.15412v1 null
2026-08-15 Grounding Healthcare LLMs in a Causal Knowledge Graph: Framework, Metrics, and a Cardiovascular Pilot Ummara Mumtaz et.al. 2608.15382v1 null
2026-08-15 When AI Rewrites, Classifiers Relax: Uncertainty-Aware Sentiment Analysis on Sarcastic and AI-Paraphrased Social Text Shresth Shroff et.al. 2608.15338v1 null
2026-08-15 Physiological World Models for Human State Transitions Chongyang Zhang et.al. 2608.15309v1 null
2026-08-15 Demographic Injection in Medical Language Models under Diversity, Equity, and Inclusion Prompts Diego Mardian et.al. 2608.15254v1 null
2026-08-15 Translating finite-domain integer constraint models to CP/SMT/ILP/PB/SAT solvers with CPMpy Tias Guns et.al. 2608.15143v1 null
2026-08-15 FZ-VLM: A Two Stage Florence-Zephyr Vision Language Model Framework for Pulmonary Nodule Characterization and Clinical Decision Making Pramit Dutta et.al. 2608.15004v1 null
2026-08-14 Evaluating Agentic Code Repair Capabilities in Distributed Systems Yibo Yan et.al. 2608.14863v1 null
2026-08-14 Generated Context versus Governed State: Functional Conditions for Accountable Longitudinal Clinical Reasoning Augusto Bernardo Pissarra et.al. 2608.14804v1 null
2026-08-14 Prompting is not enough: supervised baselines and leakage control for measuring shared decision-making with LLMs in pediatric encounters Bernardo Modenesi et.al. 2608.14792v1 null
2026-08-14 CEDAR-GRPO: Process-Aware Reinforcement Learning for General Abductive Reasoning in LLMs Moein Salimi et.al. 2608.14791v1 null
2026-08-14 Seeing Red, Thinking Bad: Color Bias in Vision Language Models Kohsuke Ide et.al. 2608.14286v1 null
2026-08-14 Meteorology-driven Causal Nowcasting of Fugitive Landfill Emissions Enables Proactive Public Health Response Timothy C. Pearce et.al. 2608.14254v1 null
2026-08-14 Agentic Data Cleaning Without a Clean Reference: An Experimental Study of Capabilities and Trade-offs Hadi Fadlallah et.al. 2608.14765v1 null
2026-08-14 APTER: Adaptive Post-Training with Expert-Grounded Rubrics Xukai Wang et.al. 2608.14212v1 null
2026-08-14 Removing Temporal Note Redundancy Improves Multimodal Reinforcement Learning for Medicine Chenran Weng et.al. 2608.14157v1 null
2026-08-14 CMCNet: Aligning Ultrasound Image Embeddings with Textual TI-RADS Representations for Fine-Grained Thyroid Classification Bingxin Yu et.al. 2608.13939v1 null
2026-08-13 Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions Qingfang Liu et.al. 2608.13786v1 null
2026-08-13 Data-driven techniques for translational neuroscience and personalized neuro-health Vishal Subedi et.al. 2608.13749v1 null
2026-08-13 MedPlex: Deep Vision-Language Co-Adaptation for Clinically Grounded Medical Segmentation Rafi Ibn Sultan et.al. 2608.13690v1 null
2026-08-13 MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and Coordination Saisha Shetty et.al. 2608.13476v1 null
2026-08-13 Towards Context-Aware Clinical Motion Understanding in Daily Living at Home: Freezing of Gait Detection with Egocentric Vision Vayalet Stefanova et.al. 2608.13283v1 null
2026-08-13 Static analysis-guided agentic AI translation enables Rust as a full stack bioinformatics language Johan Henriksson et.al. 2608.13029v1 null
2026-08-13 Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence Jakub Pokrywka et.al. 2608.12928v1 null
2026-08-13 CRAFT: LLM-Based Iterative Refinement for Temporal Reasoning over Clinical Narratives Chengyang He et.al. 2608.12779v1 null
2026-08-13 Memorization Diagnostics for Code LLMs Should be Scale-Aware Prateek Kumar Rajput et.al. 2608.12771v1 null
2026-08-13 PatientAct: Theory-Grounded Mental Health Client Simulation Sahand Sabour et.al. 2608.12750v1 null
2026-08-13 Mr3D-VL: A generalist vision language foundation model for Multiparametric 3D Magnetic Resonance Imaging Zhi Qiao et.al. 2608.12689v1 null
2026-08-12 SteerBench-Work: A Benchmark for Agent Steering at Action Boundaries Oguz Serdar et.al. 2608.12654v1 null
2026-08-12 Algorithm Design and Physician Liability Shujie Luan et.al. 2608.13618v1 null
2026-08-12 Auditable agentic AI for evidence-grounded thyroid ultrasound diagnosis and reporting Haifan Gong et.al. 2608.12590v1 null
2026-08-12 How Compliant is Sepsis Treatment? An Expert-Guided Neuro-symbolic Pipeline for Generating Clinical Compliance Insights Himanshu Tripathi et.al. 2608.13617v1 null
2026-08-12 M-Net: Integrating Spectral Features and Physical Field Operators into Deep Learning for Medical Image Segmentation Jing Zhu et.al. 2608.12196v1 null
2026-08-12 A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench Praveen Reddy et.al. 2608.12138v1 null
2026-08-12 Do LLMs Take Care of Their Own? Similarity Signals Can Induce Cooperation Akash Kundu et.al. 2608.12125v1 null
2026-08-12 How Far from Clinical Deployment? Evaluating the Complete Unsupervised Domain Adaptation Pipeline in Medical Imaging Yiheng Xiong et.al. 2608.12035v1 null
2026-08-12 From Safety Documentation to Safety Knowledge Support: An Evidence-Grounded LLM Framework for Medical Devices Tuhinangshu Gangopadhyay et.al. 2608.12025v1 null
2026-08-12 When the API Speaks the Wrong Language: Revisiting Post-Training for Multilingual Tool Use Siddharth Chauhan et.al. 2608.11715v1 null
2026-08-12 Rubric Dropout: A Simple Way to Mitigate Reward Hacking in Rubric-as-Reward RL Minglai Yang et.al. 2608.11669v1 null
2026-08-12 Beyond Single-Turn Confidence: Trajectory-Adapted Uncertainty Quantification for LLM Agents Dylan Bouchard et.al. 2608.11552v1 null
2026-08-11 Social Chain of Thought: A Multi-Agent Architecture Grounded in Medical Differential Diagnosis Methodology Del Coburn et.al. 2608.11420v1 null
2026-08-11 Gaze Target Estimation Anywhere with Concepts Xu Cao et.al. 2608.11367v1 null
2026-08-11 Dual-Domain Cross-Modal Decoding for Clinical Text-Guided Medical Image Segmentation Md Maklachur Rahman et.al. 2608.11335v1 null
2026-08-11 3D Weighted Geometric Graph Neural Networks for Sheep Facial Pain Assessment Alam Noor et.al. 2608.11050v1 null
2026-08-11 CARE: Confidence-Aware Reasoning for Reliable Medical VQA Yuetian Du et.al. 2608.10964v1 null
2026-08-11 ComBodied Agents: a New Paradigm of Human-Centric Agentic AI Qianggang Ding et.al. 2608.10915v2 null
2026-08-11 MIRA: Medical Image Reflection for Agentic Diagnosis Shengzhi Wang et.al. 2608.10827v1 null
2026-08-11 DuplexWorld: Can voice agents help you get through the day? Aryan Vijay Bhosale et.al. 2608.10716v1 null
2026-08-11 MedUP: Awakening Unified Understanding and Perception in Medical Vision-Language Models Yuan Wang et.al. 2608.10635v1 null
2026-08-11 Agentic Instruction Data Selection: Let DataMaster Interpret Your Intent Fanqi Zhou et.al. 2608.10579v1 null
2026-08-11 Reinforcement Learning-Based Laser Cutting Machine Parameter Optimization Khanh Quan Pham et.al. 2608.10549v1 null
2026-08-11 Unlocking the Power of Medical Tabular Data via Semantic-Aware Multimodal Pre-training Yingsheng Liu et.al. 2608.10522v1 null
2026-08-11 RadFusion: Towards Threshold-Controllable Radiology Report Generation Ying Jin et.al. 2608.10505v1 null
2026-08-11 RLMOpt: Adaptive Prompt Optimization via Recursive Language Models Subhash Bangalore Satheesha et.al. 2608.10471v1 null
2026-08-11 Expert-Guided g-computation with Large Language Models for Estimating Causal Effects on Timings: Applications to Hospital Quality Improvement Patrick Vossler et.al. 2608.10339v1 null
2026-08-10 Logit-Boundary Geometric Belief Interfaces and Sparse Sheaf-Enclave Protocols: A Self-Contained Substrate for Secure Network Electronic Health Record (EHR) Interoperability Alvin Spivey et.al. 2608.10300v2 null
2026-08-10 Frozen Brain-MRI Foundation Models Are Site Fingerprints Saman Rahbar et.al. 2608.10295v1 null
2026-08-10 Locally Deployable Small Language Models for Emergency Department Decision Support: A Systematic Benchmark of Fine-Tuning Strategies Qingfeng Zhang et.al. 2608.10273v1 null
2026-08-10 TAF-MED: Multi-Turn Safety Refusal Collapse in LLMs Under Declared Self-Treatment Intent Waleed Jamil et.al. 2608.10258v1 null
2026-08-10 Towards Expert-level Medical AI for Real-time Video Consultations Mahvish Nagda et.al. 2608.09861v1 null
2026-08-10 MedPixel: A Unified Pixel-Language Model for Medical Reasoning and Segmentation Haoyu Yang et.al. 2608.09818v1 null
2026-08-10 AirFlow: Context Preserving and Multi-Rate State Modeling for Air Quality Forecasting Fan Yang et.al. 2608.09775v1 null
2026-08-10 Physics-Informed Machine Learning in Prognostics and Health Management: A Systematic Literature Review Christopher Braun et.al. 2608.10047v1 null
2026-08-10 Coupled Graph--Policy Distillation for Personalized Medication Safety in Older Adults with Multimorbidity Zihan Wang et.al. 2608.09443v1 null
2026-08-10 Multimodal Federated Learning under Dual-Axis Modality Missingness Adiba Orzikulova et.al. 2608.09240v1 null
2026-08-10 Rethinking Medical Landmark Localization with Prototype Learning-based Progressive Offset Correction Jingxian Xu et.al. 2608.09182v2 null
2026-08-10 RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning Jinkun Hou et.al. 2608.09123v1 null
2026-08-10 A Multi-Scale Temporal Framework with Dynamic Fusion for EEG-Based Emotion Recognition Stefanos Gkikas et.al. 2608.09088v1 null
2026-08-10 When Confidence Fails: Overconfidence in LLMs under Uncertainty and Missing Clinical Information Maryam Tahermazandarani et.al. 2608.09080v1 null
2026-08-09 Decoding Phenotypes: A Framework for Fusing Genomic Language Models and Neuroimaging Tianli Tao et.al. 2608.08926v1 null
2026-08-09 Toward CT-Equivalent Image Quality in Low-Dose Radiotherapy Planning: Conditional Diffusion-Based CBCT-to-CT Synthesis and the Impact of CBCT Input Representation Alzahra Altalib et.al. 2608.08919v1 null

Abstracts

MotoSafety: Edge-AI with Learned Temporal Importance for Two-Wheeler Collision Risk Assessment Under Time Pressure

2608.17823v1 by Sumit S. Shevtekar, Chandresh K. Maurya, Gourab Sil, Subasish Das

Powered two-wheeler riders face critical safety challenges in low- and middle-income countries, yet limited studies exist on how cognitive stressors such as Time Pressure influence collision risk. To address this gap, we introduce a large-scale dataset of over 129,000 labeled multivariate time-series sequences from 153 simulator rides by 51 participants under No, Low, and High TP, capturing 64 features across vehicle dynamics, control inputs, proximity, and behavioral violations. Building on this dataset, we propose MotoSafety, a novel edge-AI architecture grounded in the Learned Temporal Importance principle. MotoSafety achieves 94.97% accuracy and 99.33% ROC AUC, outperforming ten baselines, including TimesNet and LLM4TS, and achieves 0.039 MSE and 0.094 MAE for forecasting (4.4x lower error than Time-LLM and iTransformer). With only 1.15M parameters and 0.135 ms latency, it is suitable for edge deployment on low-cost CPU hardware. Using ground truth TP as an inductive bias improves accuracy from 94.09% to 94.97%, while predicted TP achieves 94.82%. Using only 21 IMU+GPS features, it achieves 93.91% accuracy, indicating practical deployment. Beyond PTW safety, the architecture shows better transferability to human activity (97.66%) and clinical (99.65%) domains. This lightweight framework advances PTW collision risk assessment, supporting the Safe System Approach for Intelligent Transportation Systems.

摘要:在中低收入國家,動力二輪車騎士面臨著重大的安全挑戰,但關於認知壓力因素如時間壓力如何影響碰撞風險的研究卻相對有限。為了填補這一空白,我們引入了一個大規模數據集,該數據集包含來自51名參與者在無時間壓力、低時間壓力和高時間壓力下進行的153次模擬騎行的超過129,000個標記的多變量時間序列,捕捉了64個特徵,涵蓋了車輛動態、控制輸入、接近度和行為違規。基於這個數據集,我們提出了MotoSafety,一種基於學習時間重要性原則的新型邊緣人工智慧架構。MotoSafety實現了94.97%的準確率和99.33%的ROC AUC,超越了包括TimesNet和LLM4TS在內的十個基準,並在預測中達到了0.039的均方誤差和0.094的平均絕對誤差(比Time-LLM和iTransformer低4.4倍)。它僅需1.15M的參數和0.135毫秒的延遲,適合在低成本CPU硬體上進行邊緣部署。使用真實的時間壓力作為歸納偏見,準確率從94.09%提高到94.97%,而預測的時間壓力則達到94.82%。僅使用21個IMU+GPS特徵,它的準確率達到93.91%,顯示出實際部署的潛力。除了PTW安全性外,該架構在人體活動(97.66%)和臨床(99.65%)領域也顯示出更好的可轉移性。這個輕量級框架推進了PTW碰撞風險評估,支持智能交通系統的安全系統方法。

LLMs for Medical Consultation Are Evaluated Too Late: The Preformulation Gap

2608.17330v1 by Yining Hua, Cyrus Ayubcha, Hongbin Na, Levi Lian, Alon Gorenshtein, Yiftach Barash, Eyal Klang

Large language models for medical consultation are often evaluated after a clinical problem has already been made clear, although real consultations may begin with a vague, minimized, or misframed concern. We evaluated three API models across four physician-authored, multi-turn vignettes under baseline and entry-to-care instruction conditions, yielding 24 fixed-script transcripts; two cases also used adaptive standardized-patient simulation, yielding 12 transcripts. Self-care or home-management advice before any patient answer appeared in 9 of 12 baseline case-model cells and 0 of 12 instruction cells, while structured handoff summaries appeared in 0 of 12 and 10 of 12 cells, respectively. The instruction changed sequencing and documentation, although it did not reliably ensure elicitation of decisive facts. The preformulation gap should therefore be evaluated directly through observable first-contact behavior rather than inferred from diagnostic accuracy or final-answer quality.

摘要:大型語言模型在醫療諮詢中的評估通常是在臨床問題已經明確之後進行的,儘管實際的諮詢可能是從模糊、最小化或錯誤框架的關切開始的。我們在基線和進入護理指導條件下,評估了三個API模型在四個由醫生撰寫的多輪小品中的表現,共產生了24份固定腳本的逐字稿;另外兩個案例還使用了自適應標準化病人模擬,產生了12份逐字稿。在12個基線案例模型單元中,有9個出現了自我照護或居家管理建議,而在12個指導單元中則沒有出現;結構化交接摘要在12個單元中分別出現了0個和10個。指導改變了序列和文檔,儘管它並未可靠地確保引出關鍵事實。因此,前置公式化的差距應該通過可觀察的首次接觸行為直接評估,而不是從診斷準確性或最終答案質量中推斷。

Delta2Gamma: Band-Wise Adaptive Contrastive Learning of EEG for Alzheimer's Disease Detection

2608.17231v1 by Chanwoo Park, Chanwoo Kim

Low-cost, scalable screening for dementia remains an open problem. Imaging-based diagnosis is costly and hard to deploy widely. Electroencephalography (EEG) is portable and inexpensive, but its recordings are noisy, vary widely across subjects, and carry few clinical labels. We tackle this with Delta2Gamma, a self-supervised framework that learns EEG representations from unlabeled data by contrasting augmented views of each signal. Rather than treat EEG as a single stream, Delta2Gamma decomposes every recording into the five canonical neural rhythms (delta, theta, alpha, beta, gamma). Each band gets its own encoder and projection head. Each also gets a temperature that is predicted adaptively during contrastive training, so bands with different signal statistics are balanced automatically. On the ADFTD cohort under a strict leave-one-subject-out protocol, Delta2Gamma separates Alzheimer's disease from cognitively normal controls with 92.4\% accuracy. This exceeds both supervised backbones and recent dedicated EEG methods.

摘要:低成本、可擴展的癡呆篩檢仍然是一個未解決的問題。基於影像的診斷成本高且難以廣泛部署。腦電圖(EEG)便攜且便宜,但其錄音雜訊多、在受試者之間變化大,且臨床標籤少。我們通過Delta2Gamma來解決這個問題,這是一個自我監督框架,通過對比每個信號的增強視圖來學習無標籤數據的EEG表示。Delta2Gamma並不是將EEG視為單一流,而是將每個錄音分解為五種典型的神經節律(delta、theta、alpha、beta、gamma)。每個頻帶都有自己的編碼器和投影頭。每個頻帶還在對比訓練過程中自適應地預測一個溫度,因此具有不同信號統計的頻帶會自動平衡。在嚴格的留一受試者外協議下的ADFTD隊列中,Delta2Gamma以92.4\%的準確率將阿茲海默病與認知正常的對照組分開。這超過了監督式骨幹和最近專門的EEG方法。

A decodability criterion predicts when hidden-state selection beats majority voting in large language models

2608.17124v1 by Zhixiang wang, Ziliang Hong, Ulas Bagci

Combining the answers a large language model (LLM) samples for a question into one decision is a test-time information fusion problem, usually solved by majority voting. Voting is unreliable on difficult questions, where the sampled answers share correlated errors, so the wrong answer can win and drawing more samples makes the decision worse. Selecting a candidate by reading a correctness signal from the model's hidden states is a promising alternative, but its accuracy varies across models and tasks, and no measure indicates when it can be trusted. In this paper, we propose CASE (Correctness-Axis SElection), a dynamic selection combiner that trains a linear gate on the answer-token hidden state and selects the highest-scoring candidate. Its main contribution is decodability, a leakage-free measure of how well the gate ranks a question's correct candidates above its incorrect ones, which predicts whether hidden-state selection will outperform voting. A conventional probe appears accurate only because of question-identity leakage, which vanishes under question-grouped evaluation. On held-out data, decodability predicts the accuracy gain of selection over voting with a Pearson correlation r=0.75 and a decision threshold near AUC=0.60. Across general and medical LLMs, CASE improves over voting by up to 19 points on medium-difficulty questions and 16.8 points on hard questions. Decodability depends on the aligned knowledge a model must recall, not on its scale, and its prediction transfers to an unseen scientific domain within 3.8 points. It thus provides a practical criterion, measurable in advance for a given model and task, for choosing between learned selection and majority voting.

摘要:將大型語言模型(LLM)對一個問題所採樣的答案合併為一個決策是一個測試時的信息融合問題,通常通過多數投票來解決。在困難問題上,投票不可靠,因為採樣的答案共享相關錯誤,因此錯誤的答案可能會獲勝,而增加更多樣本會使決策變得更糟。通過從模型的隱藏狀態中讀取正確性信號來選擇候選者是一個有前途的替代方案,但其準確性在不同模型和任務之間有所變化,且沒有任何指標表明何時可以信任它。在本文中,我們提出了CASE(正確性軸選擇),這是一個動態選擇組合器,對答案標記的隱藏狀態訓練一個線性閘,並選擇得分最高的候選者。它的主要貢獻是可解碼性,這是一種無洩漏的度量,衡量閘如何將問題的正確候選者排名高於不正確的候選者,並預測隱藏狀態選擇是否會優於投票。傳統探測器之所以顯得準確,僅僅是因為問題身份的洩漏,而這在問題分組評估中會消失。在保留數據上,可解碼性預測選擇相對於投票的準確性增益,皮爾森相關係數 r=0.75,決策閾值接近 AUC=0.60。在一般和醫療 LLM 中,CASE 在中等難度問題上提高了最多 19 分,在困難問題上提高了 16.8 分。可解碼性取決於模型必須回憶的對齊知識,而不是其規模,且其預測在未見的科學領域內轉移至 3.8 分。因此,它為在給定模型和任務之間選擇學習的選擇和多數投票提供了一個可實際測量的標準。

Foundation Agents Meet Agentic Deep Research: Evidence-Grounded Clinical Code Forecasting

2608.17075v1 by Junda Wang, Meysam Ghaffari, Akshat Choube, Mohsen Sharifi Renani, Hong Yu, Carlos Morato

Next-encounter ICD forecasting predicts which standardized diagnosis codes will be documented at a future visit from the longitudinal record available beforehand. The task is prospective and multi-label: the target note does not yet exist, and several codes may be correct. Structured EHR foundation models capture recurrence and temporal progression, whereas language foundation models generate flexible diagnostic hypotheses. We introduce ICD-Deepresearch, a DeepResearch workflow that composes these predictive foundation models with medical search and ICD dictionaries. Because no source reveals the future code set, research evaluates candidate transitions by linking patient evidence, external clinical relations, and exact code semantics under a fixed top-K budget. Candidate Generation uses SparseEHR to produce an EHR Prior that initializes two bounded Research Expansion rounds; an independent GPT-5 Direct Forecast supplies complementary candidates. Final Selection validates, deduplicates, and jointly ranks both paths, after which a separate module writes rationales without changing predictions. Finally ICD-Deepresearch achieves patient-averaged precision/recall of 24.60/35.09% on MIMIC-III and 25.14/48.32% on MIMIC-IV. Physicians rate 51% and 68% of its retrieved documents useful, compared with 22% and 39% for standalone GPT-5 web search and 32% and 41% for Medical Deep Research. ICD-Deepresearch therefore improves over the registered local comparators while retrieving evidence with higher physician-rated usefulness than the standalone research systems

摘要:下一次接觸的 ICD 預測預測未來訪問時將記錄哪些標準化診斷代碼,這些代碼來自之前可用的縱向記錄。這個任務是前瞻性的和多標籤的:目標筆記尚不存在,且可能有多個代碼是正確的。結構化的電子健康記錄基礎模型捕捉到復發和時間進展,而語言基礎模型則生成靈活的診斷假設。我們介紹 ICD-Deepresearch,這是一個 DeepResearch 工作流程,將這些預測基礎模型與醫療搜索和 ICD 字典結合起來。由於沒有來源揭示未來的代碼集,研究通過將患者證據、外部臨床關係和確切的代碼語義連接在一起來評估候選轉換,並在固定的 top-K 預算下進行。候選生成使用 SparseEHR 生成一個 EHR Prior,該 Prior 初始化兩輪有界的研究擴展;獨立的 GPT-5 直接預測提供補充候選。最終選擇驗證、去重並共同排名這兩條路徑,之後一個單獨的模塊在不改變預測的情況下寫出理由。最後,ICD-Deepresearch 在 MIMIC-III 上達到患者平均精確度/召回率 24.60/35.09%,在 MIMIC-IV 上達到 25.14/48.32%。醫生認為其檢索的文檔中有 51% 和 68% 是有用的,而獨立的 GPT-5 網頁搜索的有用率為 22% 和 39%,醫療深度研究的有用率為 32% 和 41%。因此,ICD-Deepresearch 在檢索證據時的醫生評價有用性上優於登記的本地比較者,並且比獨立研究系統更具優勢。

Institution-Specific LLM Prompting Recovers PHI That De-identification Systems and Their Gold Standards Both Miss

2608.17051v1 by Daniel Palacios, Matthew Brady Neeley, Angel Adetomike Otto, Shalini Dhamodharan, John P. Woodhouse, Chi-fan Lin, Mark Zobeck, Zhandong Liu, Hyun-Hwan Jeong

Secondary use of electronic health records requires de-identification, yet existing systems miss \emph{institutionally situated} protected health information (PHI) such as hospital abbreviations, building names, and internal codes whose status is locally determined. We ask whether large language models (LLMs) with in-context learning (ICL) can close this gap and control the precision--recall trade-off. On 100 annotated pediatric oncology notes (5,322 PHI spans) from Texas Children's Hospital, we benchmarked eight LLMs against two purpose-built systems (Stanford TiDE, OpenMed PII) and two pattern-based baselines. Each LLM ran under three prompts of increasing specificity: (1) a HIPAA-aligned baseline, (2) baseline plus the institutional PHI categories it missed, and (3) prompt 2 plus instructions against over-redacting clinical content. We then compared 14~multi-agent and ensemble configurations against the best single prompt, with recall the primary safety metric. LLMs outperformed the purpose-built systems (best F1=0.918$\pm$0.001 vs.\ TiDE 0.779), with advantages concentrated in contextual categories. Naming the missed categories recovered 79\% (48/61) of them, and discouraging over-redaction restored precision. No agentic architecture beat calibrated single-pass prompting (F1 0.906--0.907), but LLM outputs surfaced 414~candidate annotation gaps; re-annotation confirmed 227~PHI spans, against which the final prompt reached recall=0.981 (F1=0.907$\pm$0.002). Well-calibrated ICL resolves both the institutional PHI gap and the precision--recall trade-off in one LLM call per note. LLMs cost more to run than traditional methods, but that cost buys a way to audit the reference standard. LLMs are a legitimate, adaptable alternative to purpose-built de-identification systems; institution-specific prompt development should be the primary adaptation strategy.

摘要:次級使用電子健康紀錄需要去識別化,但現有系統忽略了\emph{制度性位置}的受保護健康資訊(PHI),例如醫院縮寫、建築名稱和其狀態由地方決定的內部代碼。我們詢問大型語言模型(LLMs)是否能透過上下文學習(ICL)填補這一空白並控制精確度與召回率的權衡。在來自德克薩斯兒童醫院的100份註釋小兒科腫瘤學筆記(5,322個PHI範圍)中,我們對八個LLMs進行了基準測試,並與兩個專門構建的系統(Stanford TiDE,OpenMed PII)和兩個基於模式的基準進行比較。每個LLM在三個逐步具體化的提示下運行:(1)符合HIPAA的基準,(2)基準加上其遺漏的制度PHI類別,以及(3)提示2加上對過度刪除臨床內容的指示。我們隨後比較了14個多代理和集成配置與最佳單一提示,召回率是主要的安全指標。LLMs的表現超過了專門構建的系統(最佳F1=0.918$\pm$0.001對比TiDE 0.779),優勢集中在上下文類別中。命名遺漏的類別恢復了79\%(48/61),而抑制過度刪除則恢復了精確度。沒有任何代理架構超越經過校準的單次提示(F1 0.906--0.907),但LLM輸出顯示了414個候選註釋缺口;重新註釋確認了227個PHI範圍,最終提示的召回率達到0.981(F1=0.907$\pm$0.002)。良好校準的ICL在每個筆記中解決了制度PHI缺口和精確度與召回率的權衡。LLMs的運行成本高於傳統方法,但這一成本提供了一種審計參考標準的方式。LLMs是專門構建的去識別化系統的合法且可調整的替代方案;特定機構的提示開發應該是主要的調整策略。

Policy Iteration with Human Feedback: Bringing Post-Training RL to In-context Learning

2608.16831v1 by Minh-Ha Nguyen, Cathy Shyr

Generative pretraining established reusable task representations; later work on language-based task conditioning and in-context learning showed that a fixed model could adapt its behavior from instructions and demonstrations. Policy Iteration with Human Feedback (PIHF) builds on this development and the recurrent evaluate-and-improve structure of generalized policy iteration. PIHF uses a pretrained language model as its execution substrate and moves persistent revision to a versioned natural-language policy and tool set. A language-model critic and clinical expert review complete-panel reasoning and tool-use trajectories to localize recurrent failures and form candidate revisions; the expert may reinterpret the evidence and retains authority over admission and rollback, while Recall@1 and Recall@5 validate outcomes after candidate execution. Across cumulative ablations and ultra-rare-disease benchmarks, a PIHF-derived policy improved Recall@1 in one proprietary executor and three open-weight executors spanning 3 to 49 billion active parameters. Gains were 32.7 percentage points for GPT-5.4 and 31.1 points for Qwen3.6-35B, a difference of 1.7 points. These results support the feasibility of using pretrained language models as fixed-weight execution substrates for expert-guided policy development in rare-disease diagnosis.

摘要:生成預訓練建立了可重用的任務表示;後續在基於語言的任務條件和上下文學習方面的研究顯示,固定模型可以從指令和示範中調整其行為。帶有人工反饋的策略迭代(PIHF)建立在這一發展及其一般化策略迭代的反覆評估和改進結構之上。PIHF使用預訓練的語言模型作為其執行基礎,並將持續修訂轉移到版本化的自然語言策略和工具集。語言模型評論員和臨床專家審查完整面板的推理和工具使用軌跡,以定位反覆失敗並形成候選修訂;專家可以重新解釋證據,並保留對接受和回滾的權威,而Recall@1和Recall@5在候選執行後驗證結果。
在累積的消融和超罕見疾病基準測試中,PIHF衍生的策略在一個專有執行器和三個開放權重執行器中改善了Recall@1,這些執行器的活動參數範圍從30億到490億。GPT-5.4的增益為32.7個百分點,Qwen3.6-35B的增益為31.1個百分點,兩者之間的差異為1.7個百分點。這些結果支持使用預訓練語言模型作為固定權重執行基礎,在罕見疾病診斷中進行專家引導的策略開發的可行性。

Historical Backtesting for Scientific Question Discovery: A Protocol and Astronomy Pilot

2608.16795v1 by Hui Mao

Systems that generate scientific research questions are evaluated today by expert scores, LLM-as-judge ratings, or curated case studies -- all subjective, none falsifiable. We formalize historical backtesting as an alternative: a system generates questions from a corpus frozen at a historical cutoff, the questions are frozen before any access to later literature, and a temporally isolated future corpus then determines whether each question was subsequently answered, partially addressed, independently posed, or ignored, and whether its underlying premise was supported or refuted. The protocol is model-agnostic: any system that emits frozen questions can be scored. We release reproducible astronomy instances with temporally isolated corpora, frozen questions, auditable labels, four reference baselines, and a submission interface. Two findings result. First, evidence-structure-first generation outperforms LLM-only prompting: across a generator decomposition crossed with a four-cutoff stress test (2010-2024, 798 judged questions) whose last window postdates model training, LLM-only generation shows memorized relevance without specific foresight, while a generator using no model weights at all finds questions whose premises the future refutes in every era. Second, a seven-rater agreement study (two blinded human annotators, five judge models, 90 items) indicts the outcome taxonomy rather than the judge: two careful humans agree at kappa = 0.17, every judge model agrees with the professional annotator as well or better (0.17-0.26), and frontier models agree with one another at 0.60 -- certifying an LLM judge by model-model agreement would have overstated its reliability threefold. A prospective instance -- 200 questions frozen 2026-08-17, scored 2027-2030 -- is released so the central claims become contamination-free tests that time itself will grade.

摘要:生成科學研究問題的系統今天由專家評分、LLM作為評判的評級或策劃的案例研究來評估——這些都是主觀的,沒有一個是可證偽的。我們將歷史回測形式化為一種替代方案:系統從在歷史截止日期凍結的語料庫中生成問題,這些問題在任何接觸後續文獻之前被凍結,然後一個時間上隔離的未來語料庫決定每個問題是否隨後得到了回答、部分解決、獨立提出或被忽視,以及其基礎前提是否得到了支持或反駁。該協議是模型無關的:任何發出凍結問題的系統都可以被評分。我們發布了可重複的天文學實例,具有時間上隔離的語料庫、凍結問題、可審計的標籤、四個參考基準和提交界面。結果有兩個發現。首先,證據結構優先的生成表現優於僅依賴LLM的提示:在一次生成器分解與四次截止壓力測試(2010-2024年,798個評判問題)的交叉中,最後一個窗口的時間超過了模型訓練,僅依賴LLM的生成顯示出記憶中的相關性而沒有具體的預見,而一個完全不使用模型權重的生成器則找到了在每個時代中未來反駁其前提的問題。第二,一項七評審者一致性研究(兩名盲人人工標註者,五個評判模型,90個項目)指控結果分類法而不是評判者:兩名謹慎的人類在kappa = 0.17的情況下一致,每個評判模型與專業標註者的意見一致或更好(0.17-0.26),而前沿模型之間的一致性為0.60——通過模型之間的一致性來認證LLM評判者將會高估其可靠性三倍。一個前瞻性實例——200個問題凍結於2026-08-17,於2027-2030年進行評分——被發布,以便中央主張成為不受污染的測試,時間本身將進行評分。

Unsupervised Anomaly Detection for Image Dataset Quality Assurance in Multi-Center Breast MRI

2608.16725v1 by Chiara Tappermann, Steffen Renisch, Lars Ole Schwen, Hans Meine, Horst K. Hahn, Eike Petersen

Corrupted, inconsistent, or anomalous data silently threatens the safety and reliability of medical AI. Despite growing regulatory recognition of dataset quality assurance (QA) for high-risk medical AI, scalable automated detection remains underdeveloped. We employ unsupervised anomaly detection (AD) and out-of-distribution (OOD) detection as an automated dataset QA mechanism for multi-center dynamic contrast-enhanced breast MRI. We build a controlled AD benchmark of 17 realistic QA-relevant anomaly types from six public datasets (protocol violations, processing errors, incorrect anatomical regions) and propose a taxonomy of radiological image anomalies based on human visual perception, enabling fine-grained analysis of AD failure modes. The benchmark includes near-, medium-far-, far-OOD samples, as well as in-distribution and external normal data. Four methods are evaluated: a projection-based method extended with a domain-specific feature extractor and a novel positional encoding, a reconstruction-based approach extended to full 3D volumes with an augmented training objective, and two unmodified hybrid OOD detection methods. Medium-far- and far-OOD samples are detected reliably, whereas near-OOD samples and external normal data from unseen institutions expose method-specific differences. The 3D reconstruction-based approach best balances detection performance (AUROC: 0.936) and generalization to unseen institutions. The projection-based method with positional encoding achieves the highest overall detection performance (AUROC: 0.954). Both hybrid methods exhibit critical failure modes, confirming that methods validated for one modality or anatomy may not generalize without domain-specific adaptation. Implants and mastectomies remain an open challenge for all methods. Our results establish a foundation and practical guidance on scalable unsupervised QA in medical AI pipelines.

摘要:腐敗、不一致或異常的數據默默威脅著醫療人工智慧的安全性和可靠性。儘管對高風險醫療人工智慧數據集質量保證(QA)的監管認識日益增長,但可擴展的自動檢測仍然發展不足。我們採用無監督異常檢測(AD)和分佈外(OOD)檢測作為多中心動態對比增強乳腺MRI的自動數據集QA機制。
我們建立了一個由六個公共數據集中的17種現實QA相關異常類型組成的受控AD基準(協議違規、處理錯誤、不正確的解剖區域),並根據人類視覺感知提出了一個放射影像異常的分類法,使得對AD失效模式的細緻分析成為可能。基準包括近距離、中遠距離、遠距離OOD樣本,以及分佈內和外部正常數據。評估了四種方法:一種基於投影的方法,擴展了特定領域的特徵提取器和新穎的位置編碼;一種基於重建的方法,擴展到完整的3D體積並具有增強的訓練目標;以及兩種未經修改的混合OOD檢測方法。
中遠距離和遠距離OOD樣本的檢測可靠,而近距離OOD樣本和來自未見機構的外部正常數據則顯示出方法特定的差異。基於3D重建的方法在檢測性能(AUROC:0.936)和對未見機構的泛化之間達到了最佳平衡。帶有位置編碼的基於投影的方法實現了最高的整體檢測性能(AUROC:0.954)。兩種混合方法都顯示出關鍵的失效模式,確認了針對一種模態或解剖結構驗證的方法可能無法在沒有特定領域適應的情況下進行泛化。植入物和乳房切除術對所有方法仍然是一個未解決的挑戰。我們的結果為醫療人工智慧管道中的可擴展無監督QA建立了基礎和實用指導。

Toward Better Assessment of LLMs' Performance in Clinical Error Detection

2608.16643v1 by Yifan Zhang, Rahmatollah Beheshti

Automated detection of errors in clinical documentation is a promising application of large language models (LLMs), yet decisions to deploy such models rest on benchmarks that evaluate each clinical note in isolation. Error-detection benchmarks are typically constructed by injecting errors into notes, such that each erroneous note has a natural counterpart. Aggregate discriminative metrics (e.g., balanced accuracy or F1) do not exploit this structure. We show that this omission is consequential. In particular, evaluating 15 diverse LLMs on 4 standardized clinical error-detection test sets across 3 languages, we find that 13 of 15 models fall below the level of random pairwise discrimination, even while achieving F1 scores that standard practice would read as moderate. We also observe that the underlying bias patterns differ across languages: the same model can default to "no error" on one language and over-flag errors on another. To diagnose where discrimination breaks down, we further introduce a procedure to score the evidence models cite in their outputs. We find that while models consistently locate error-relevant content, they fail to produce the corresponding correct verdict on the clean counterpart. Finally, we show that F1 and pairwise accuracy are driven in opposite directions by the same underlying bias, so that ranking models by F1 may systematically promote the weakest discriminators. For safety-critical clinical NLP applications, we advocate for supplementing aggregate metrics with paired evaluations in benchmark reporting. Code and analysis scripts are available at https://github.com/healthylaife/paired-clinical-eval.

摘要:自動檢測臨床文檔中的錯誤是大型語言模型(LLMs)的有前景應用,然而,部署這些模型的決策依賴於評估每條臨床記錄的基準,這些基準是孤立評估的。錯誤檢測基準通常是通過在記錄中注入錯誤來構建的,使得每條錯誤記錄都有一個自然的對應記錄。聚合的判別指標(例如,平衡準確率或F1)並未利用這一結構。我們表明,這一遺漏是有後果的。具體而言,在對3種語言的4個標準化臨床錯誤檢測測試集上評估15種不同的LLMs時,我們發現15個模型中有13個的表現低於隨機成對判別的水平,即使它們的F1分數在標準實踐中被視為中等。我們還觀察到,潛在的偏見模式在不同語言之間存在差異:同一模型在一種語言上可能默認為「無錯誤」,而在另一種語言上則過度標記錯誤。為了診斷判別失效的原因,我們進一步引入了一個程序來評分模型在其輸出中引用的證據。我們發現,儘管模型始終能定位與錯誤相關的內容,但它們未能對乾淨的對應記錄給出正確的判決。最後,我們表明F1和成對準確率受到同一潛在偏見的驅動,方向卻相反,因此根據F1對模型進行排名可能系統性地促進最弱的判別者。對於安全關鍵的臨床NLP應用,我們主張在基準報告中用成對評估來補充聚合指標。代碼和分析腳本可在 https://github.com/healthylaife/paired-clinical-eval 獲得。

Degradation-Aligned Self-Supervised Learning for State of Health Estimation of Lithium-Ion Batteries under Label Sparsity

2608.16612v1 by Jiaqi Yao, Julia Kowal

An accurate estimation of the state of health (SOH) underpins a safe and optimized use of the battery system. Although compelling, data-driven SOH estimation models typically require large amounts of high-quality labeled cycling data, while in practice such labels are often sparse in both quantity and coverage. Therefore, in this work, we propose a degradation-aligned self-supervised learning (SSL) framework based on a convolutional neural network-gated recurrent unit (CNN-GRU) model, which learns aging-consistent representations from unlabeled data through a cycle-order ranking objective as the pretext task for pretraining, thereby enabling robust SOH estimation after fine-tuning on sparsely labeled data. Test results showcase that the proposed ranking-based SSL approach proves to endow the pretrained model with degradation-aligned information from unlabeled data, and after fine-tuning the model can carry out accurate, robust SOH estimation, even when only an extremely limited amount of 1% of unevenly distributed labeled training data is available, where the MAE of 1.718% and RMSE of 2.329% can be achieved on the test cell. In addition, in-depth analyses are presented regarding the influences of label distribution of battery degradation data. We believe this work could shed new light on SOH estimation of lithium-ion batteries under label sparsity in real-world applications.

摘要:準確的健康狀態(SOH)估計是安全且優化使用電池系統的基礎。雖然數據驅動的SOH估計模型非常有說服力,但通常需要大量高質量的標註循環數據,而在實際情況中,這些標註往往在數量和覆蓋範圍上都很稀疏。因此,在本研究中,我們提出了一種基於卷積神經網絡-門控遞歸單元(CNN-GRU)模型的降解對齊自監督學習(SSL)框架,通過循環順序排名目標作為預訓練的前置任務,從未標註數據中學習與老化一致的表示,從而在稀疏標註數據上進行微調後實現穩健的SOH估計。測試結果顯示,所提出的基於排名的SSL方法使預訓練模型從未標註數據中獲得了降解對齊的信息,並且在微調後,該模型能夠進行準確且穩健的SOH估計,即使僅有極少量的1%不均勻分佈的標註訓練數據可用,測試電池的MAE可達1.718%和RMSE可達2.329%。此外,還對電池降解數據的標註分佈影響進行了深入分析。我們相信這項工作可以為在現實應用中標註稀疏的鋰離子電池SOH估計提供新的見解。

CACSurv: Concordance-Aligned Comparative Learning with Large Language Models for Cancer Survival Prediction

2608.16594v1 by Tianqi Xiang, Qixiang Zhang, Xinpeng Ding, Yi Li, Xiaomeng Li

Cancer survival prediction supports treatment planning, risk stratification, and follow-up management. Existing methods use structured clinical variables, whole-slide images, genomic profiles, or multimodal inputs, while patient reports remain underexplored. We study report-centric survival prediction using reports that organize pathological, clinical, and molecular evidence. Large language models (LLMs) can reason over such reports, but case-wise time regression introduces two mismatches. First, a formulation mismatch arises because survival evaluation depends on ordering comparable patients, whereas independent time predictions do not enforce ranking consistency. Second, a supervision mismatch arises because a censored patient's observed time indicates survival beyond that point and cannot serve as an exact regression target, although it still implies orderings relative to patients who died earlier. To address these mismatches, we propose CACSurv, a Concordance-Aligned Comparative framework for report-centric survival prediction. CACSurv reformulates survival modeling as mini-cohort comparative reasoning, where an LLM predicts relative prognostic orderings. We introduce concordance-aligned rewards derived from comparable relations under right censoring, enabling censored outcomes to provide ranking supervision without exact event-time targets. At inference, Monte Carlo Reference Aggregation compares each patient with sampled references and aggregates positions into a cohort-level ranking. We establish TCGA-SurvReport, a benchmark covering six TCGA cancer cohorts. CACSurv achieves the highest C-index on all six cohorts and an average C-index of 0.722, outperforming the strongest published survival model by 6.5 percentage points and the strongest LLM time-regression baseline by 4.2 percentage points. Our code, models, and dataset will be available at https://github.com/xmed-lab/CACSurv.

摘要:癌症生存預測支持治療計劃、風險分層和後續管理。現有方法使用結構化臨床變數、全幻燈片影像、基因組資料或多模態輸入,而患者報告仍然未被充分探索。我們研究以報告為中心的生存預測,使用組織病理、臨床和分子證據的報告。大型語言模型(LLMs)可以對這些報告進行推理,但逐案例時間回歸引入了兩個不匹配。首先,因為生存評估依賴於可比較患者的排序,而獨立的時間預測並不強制執行排名一致性,因此產生了表述不匹配。其次,因為被審查患者的觀察時間表示超過該點的生存,並不能作為精確的回歸目標,儘管它仍然暗示了相對於早逝患者的排序,因此產生了監督不匹配。為了解決這些不匹配,我們提出了CACSurv,一個以報告為中心的生存預測的協調對齊比較框架。CACSurv將生存建模重新表述為小型隊列比較推理,其中LLM預測相對的預後排序。我們引入了基於右側審查下可比較關係衍生的協調對齊獎勵,使得被審查的結果能夠提供排名監督,而不需要精確的事件時間目標。在推理階段,蒙地卡羅參考聚合將每位患者與抽樣參考進行比較,並將位置聚合成隊列級別的排名。我們建立了TCGA-SurvReport,一個涵蓋六個TCGA癌症隊列的基準。CACSurv在所有六個隊列上達到了最高的C指數,平均C指數為0.722,超越了最強的已發表生存模型6.5個百分點和最強的LLM時間回歸基線4.2個百分點。我們的代碼、模型和數據集將在https://github.com/xmed-lab/CACSurv上提供。

Large language models as synthetic clinical experts to inform longitudinal rare-disease modeling

2608.16507v1 by Clemens Schächter, Astrid Pechmann, Janbernd Kirschner, Jan Hasenauer, Harald Binder

Due to the limited amount of information, modeling longitudinal rare-disease data can benefit from integrating clinical knowledge. Yet, elicitation of expert knowledge and formalization for model fitting is challenging, in particular due to limited time of clinical experts. To nevertheless make domain knowledge accessible during model fitting, we use large language models (LLMs) as synthetic clinical experts to supervise a variational-autoencoder-based approach that learns low-dimensional latent summaries of visit-level observations. Specifically, LLMs are queried offline on textual descriptions of patient observations to obtain judgments, e.g., the suspected clinical category. To improve the variational autoencoder fit, we train a differentiable surrogate model on these judgments and augment the loss function to encourage reconstructions that preserve the clinical-label distribution of their corresponding input profile. In an application to longitudinal motor-function assessments from children with spinal muscular atrophy, we map visit-level clinical profiles to low-dimensional representations that are linked by a multivariate mixed-effects model. The synthetic expert loss discourages reconstructions that remain numerically close in data space but alter the clinical interpretation of the reconstructed motor function profile, such as by crossing a disease-type boundary. We thus reduced disagreement between original and reconstructed SMA type labels from about 11 to 7 percent. Furthermore, informing the latent representation by the synthetic expert improved prediction of motor function milestones compared with unsupervised latent representations and a data-level baseline. These results suggest that incorporating LLMs into model fitting can make clinical knowledge available to representation learning and improve clinical faithfulness for longitudinal rare-disease data.

摘要:由於資訊量有限,建模縱向罕見疾病數據可以從整合臨床知識中受益。然而,專家知識的引出和模型擬合的形式化是具有挑戰性的,特別是由於臨床專家的時間有限。儘管如此,為了在模型擬合過程中使領域知識可用,我們使用大型語言模型(LLMs)作為合成臨床專家,來監督基於變分自編碼器的方法,該方法學習訪問級觀察的低維潛在摘要。具體來說,我們在患者觀察的文本描述上離線查詢LLMs以獲得判斷,例如,懷疑的臨床類別。為了改善變分自編碼器的擬合,我們在這些判斷上訓練了一個可微分的替代模型,並增強損失函數以鼓勵重建保持其對應輸入特徵的臨床標籤分佈。在對脊髓性肌萎縮症兒童的縱向運動功能評估的應用中,我們將訪問級臨床特徵映射到由多變量混合效應模型鏈接的低維表示。合成專家損失會抑制在數據空間中數值上接近但改變重建運動功能特徵的臨床解釋的重建,例如通過跨越疾病類型邊界。因此,我們將原始和重建的SMA類型標籤之間的分歧從約11%減少到7%。此外,通過合成專家告知潛在表示,與無監督潛在表示和數據級基準相比,運動功能里程碑的預測得到了改善。這些結果表明,將LLMs納入模型擬合可以使臨床知識可用於表示學習,並改善縱向罕見疾病數據的臨床真實性。

Offline Reinforcement Learning for Hemodynamic Management of Sepsis in the ICU: a MIMIC-IV Study with Dual Off-Policy Evaluation

2608.16482v1 by Marc Pérez-Roig, David Fernández-Narro, Carlos Sáez

The dosing of intravenous fluids and vasopressors in sepsis is a sequential decision made under uncertainty and guided largely by clinical judgment, which makes it a natural target for reinforcement learning from historical care. Because a learned policy cannot be trialed on patients, its value must be estimated off-policy, and such estimates can be fragile and optimistic. This work advances the reliable evaluation of sepsis treatment policies by combining off-policy estimation, reliability diagnostics, and clinician-agreement analyses in a transparent validation framework. We modeled fluid and vasopressor dosing on a cohort of 36,872 septic ICU stays drawn from the MIMIC-IV critical-care database, as a discretized Markov decision process with 1,000 states and 25 actions, defined by a five-by-five grid of fluid and vasopressor levels and solved by policy iteration. The clinicians' behavior policy was estimated with a random forest, which mitigated the collapse of the Effective Sample Size (ESS 50.1 against 4.0 with smoothed counts) that otherwise destabilizes the importance-sampling estimate. The learned policy was evaluated with two estimators, weighted importance sampling (WIS) and fitted Q evaluation (FQE), with the ESS and clinician agreement as reliability checks. An empirical variable selection found that the composition of the state matters more than its size. Both estimators place the learned policy above the clinicians' return (WIS 50.8 and FQE 46.8 against 38.2, ESS 50.1), yet it departs only modestly from observed practice (total variation 0.18), favoring less intravenous fluid. These retrospective single-center off-policy results support the learned policy as a clinically plausible refinement of observed practice and motivate its further evaluation as a discordance-based clinical decision-support approach.

摘要:靜脈輸液和血管加壓劑在敗血症中的劑量是一個在不確定性下做出的順序決策,主要依賴臨床判斷,這使其成為從歷史護理中進行強化學習的自然目標。由於學習到的策略不能在患者身上進行試驗,因此其價值必須在政策外進行估算,而這樣的估算可能是脆弱和樂觀的。本研究通過在透明的驗證框架中結合政策外估算、可靠性診斷和臨床醫生一致性分析,推進了敗血症治療政策的可靠評估。我們在從MIMIC-IV重症護理數據庫中提取的36,872例敗血症ICU住院病例上建模了液體和血管加壓劑的劑量,將其定義為一個具有1,000個狀態和25個行動的離散馬可夫決策過程,並通過政策迭代進行求解。臨床醫生的行為政策是通過隨機森林估算的,這減輕了有效樣本量(ESS 50.1對比平滑計數的4.0)的崩潰,否則會使重要性抽樣估算不穩定。學習到的策略通過兩個估算器進行評估,權重重要性抽樣(WIS)和擬合Q評估(FQE),並以ESS和臨床醫生一致性作為可靠性檢查。一項實證變量選擇發現,狀態的組成比其大小更為重要。兩個估算器均將學習到的策略置於臨床醫生的回報之上(WIS 50.8和FQE 46.8對比38.2,ESS 50.1),但它僅與觀察到的實踐有適度的偏離(總變異0.18),更傾向於較少的靜脈輸液。這些回顧性單中心的政策外結果支持學習到的策略作為臨床上合理的觀察實踐的改進,並促使其作為基於不一致的臨床決策支持方法進一步評估。

Adaptive Post-Processing Drives Instance-Level Detection in Stroke Lesion Segmentation

2608.16377v1 by Qinghui Liu, Jon André Ottesen, Atle Bjørnerud, Kyrre Eeg Emblem

Instance-level lesion detection has been an increasingly larger focal point in medical image segmentation besides the more standard voxel-level overlap. Still, most pipelines are trained and post-processed for voxel overlap alone. In particular, the mismatch is most pronounced for small lesions, where a near-miss prediction---substantial overlap that falls just short of the instance-matching threshold---scores the same as a complete miss. In our ISLES'26 submission, we found that closing this gap mattered far more in post-processing than in architecture design. Our Volume-Conditioned Adaptive Post-Processing (VCAP) scheme adjusts component-size thresholds to each case's predicted lesion burden, improving Lesion-F1 by 0.032 (unbiased cross-fold estimate)---approximately 6 times larger than any architectural change we tested. A resolution-aware attention architecture (Viola2Plus), designed for small-lesion segmentation, shows why the distinction matters: it left small-lesion Dice unchanged but raised small-lesion detection rate by 3.7\%, a real effect voxel-overlap metrics alone would have missed. Under 5-fold cross-validation on the 1,453-case training set, our post-processed two-architecture ensemble achieves Dice 0.651 and Lesion-F1 0.614, versus 0.644 and 0.573 for the unprocessed single-model baseline.

摘要:實例級病變檢測在醫學影像分割中越來越受到重視,除了更標準的體素級重疊外。儘管如此,大多數管道仍僅針對體素重疊進行訓練和後處理。特別是,這種不匹配在小病變中最為明顯,近乎錯誤的預測——實質重疊但未達到實例匹配閾值——與完全錯過的得分相同。在我們的ISLES'26投稿中,我們發現縮小這一差距在後處理中比在架構設計中更為重要。我們的體積條件自適應後處理(VCAP)方案根據每個案例的預測病變負擔調整組件大小閾值,將病變F1提高了0.032(無偏交叉折估計)——這大約是我們測試的任何架構變更的6倍。針對小病變分割設計的解析度感知注意力架構(Viola2Plus)顯示了這一區別的重要性:它對小病變的Dice保持不變,但將小病變檢測率提高了3.7\%,這是一個僅依賴體素重疊指標無法捕捉的實際效果。在對1,453個案例訓練集進行5折交叉驗證的情況下,我們的後處理雙架構集成達到了Dice 0.651和病變F1 0.614,而未處理的單模型基線則為0.644和0.573。

Foresight-England: Development of a National-Scale Generative AI Model of Electronic Health Records for Medical Event Prediction across the COVID-19 Pandemic

2608.16273v1 by Simon Ellershaw, Christopher Tomlinson, Zeljko Kraljevic, Spiros Denaxas, Harry Hemingway, Cathie Sudlow, Angela M. Wood, Anoop D. Shah, Richard Dobson

Foresight-England (Foresight-E) is the first national-scale generative foundation model of electronic health records (EHRs), developed as a research pilot strictly for COVID-19 research. We evaluated its ability to model the direct and indirect effects of the pandemic. Trained from scratch entirely within the NHS England Secure Data Environment, Foresight-E is a 243-million-parameter transformer decoder. It was trained and evaluated on de-identified, longitudinal EHRs of approximately 61 million individuals, integrating primary/secondary care, death registrations, and COVID-19 data. Training and validation used a 90% subset (54.9 million) spanning November 2018 to December 2022; the remaining 10% (6.1 million) was held out for evaluation. Foresight-E models patient timelines autoregressively, predicting the next medical event given their prior history. At inference, it operates zero-shot, predicting any concept in its ~40,000-code vocabulary without task-specific training. Our tokenisation scheme retains the clinical granularity of ICD-10, OPCS-4, and SNOMED CT codes, jointly representing absolute and relative timing. We designed an evaluation framework for 30-day COVID-19 hospitalisation and mortality, including subgroup analyses by demographic factors and vaccination status. To assess generalisation to unseen future data and the pandemic's indirect effects, we tested the model on medical events from 2023 (beyond its training period), benchmarking against logistic regression and XGBoost. As detailed in the Project Status section, NHS England has paused access to data for the Foresight-E project, meaning quantitative results are currently unavailable. Instead, we share our strategy for tokenisation, architecture, training, inference, and evaluation as a methodological template and case study in the challenges of building population-scale EHR foundation models.

摘要:Foresight-England (Foresight-E) 是首個全國規模的電子健康紀錄 (EHRs) 生成基礎模型,作為針對 COVID-19 研究的研究試點而開發。我們評估了它建模疫情直接和間接影響的能力。Foresight-E 完全在 NHS England 安全數據環境中從零開始訓練,是一個擁有 2.43 億參數的Transformer解碼器。它在約 6100 萬人的去識別化、縱向 EHRs 上進行訓練和評估,整合了初級/次級護理、死亡登記和 COVID-19 數據。訓練和驗證使用了 90% 的子集(5490 萬),涵蓋了 2018 年 11 月到 2022 年 12 月;剩餘的 10%(610 萬)則保留用於評估。Foresight-E 自回歸地建模患者時間線,根據其先前的歷史預測下一個醫療事件。在推理時,它以零樣本操作,預測其約 40,000 種代碼詞彙中的任何概念,而無需特定任務的訓練。我們的標記方案保留了 ICD-10、OPCS-4 和 SNOMED CT 代碼的臨床細節,聯合表示絕對和相對時間。我們設計了一個評估框架,用於 30 天 COVID-19 住院和死亡率,包括按人口統計因素和疫苗接種狀態的子群分析。為了評估對未見未來數據的泛化能力和疫情的間接影響,我們在 2023 年的醫療事件上測試了該模型(超出其訓練期間),並與邏輯回歸和 XGBoost 進行基準比較。正如項目狀態部分詳細說明的那樣,NHS England 已暫停對 Foresight-E 項目的數據訪問,這意味著目前無法獲得定量結果。相反,我們分享了我們的標記化、架構、訓練、推理和評估的策略,作為建立人口規模 EHR 基礎模型挑戰的 методологический шаблон и кейс-исследование。

A cross-modal generative model for incomplete and degraded prostate MRI with multicentre clinical validation

2608.16233v1 by Siyuan Ma, Liang He, Mengying Zhu, Yi Chai, Mengyao Lyu, Haowei Wang, Qizhen Lan, HaoBo Sun, Qixin Zhang, Jingli Chen, Xiaobing Wei, Jiaming Liu, Guiqin Liu, Qianwen Zhang, Yang Liu, Dacheng Tao, Guangyu Wu

Missing or degraded sequences can limit prostate multiparametric MRI. We developed MSCNet, a sequence-conditioned cross-modal generative framework for reconstructing unavailable contrasts and restoring degraded acquisitions. Across ten completion tasks, task-specific MSCNet achieved mean structural similarity of 0.818 versus 0.798 for the strongest task-matched comparators; matched-capacity analyses showed larger differences in lesion fidelity and boundary preservation. In a blinded 1,000-case reader study, overall image quality met the prespecified non-inferiority criterion for DWI, ADC and T2W completion, but not T1W. In a separate 200-case diagnostic assessment, AUCs for clinically significant cancer were 0.860 with acquired images, 0.841 with MSCNet and 0.797 with baseline-generated images. A locked 186-case three-hospital cohort supported multicentre transportability. These retrospective results support quality-controlled cross-modal reconstruction as an adjunct to acquired prostate MRI.

摘要:缺失或劣化的序列可能限制前列腺多參數MRI。我們開發了MSCNet,一種序列條件的跨模態生成框架,用於重建不可用的對比和恢復劣化的獲取。在十個補全任務中,特定任務的MSCNet達到了0.818的平均結構相似度,而最強的任務匹配比較者為0.798;匹配容量分析顯示病變真實性和邊界保護方面的差異更大。在一項盲法的1000例讀者研究中,整體影像質量達到了預先指定的DWI、ADC和T2W補全的非劣性標準,但對於T1W則不符合。在另一項200例的診斷評估中,臨床顯著癌症的AUC為獲取影像的0.860,MSCNet的0.841和基線生成影像的0.797。一個鎖定的186例三醫院隊列支持多中心可轉運性。這些回顧性結果支持質量控制的跨模態重建作為獲取前列腺MRI的輔助。

BaT: Towards Self-Evolving Medical Research Agent with Stage Rubrics

2608.16211v1 by Junqi Liu, Yufan He, Yexiao He, Pengfei Guo, Dong Yang, Andriy Myronenko, Can Zhao, Hanrong Ye, Tianhao Qi, Yuyin Zhou, Daguang Xu, Yucheng Tang

Long-horizon agents are beginning to automate complete workflows that produce code, reports, and research artifacts. Medical imaging workflows are multi-stage and data-sensitive, while expert trajectories remain scarce and difficult to share. Structured benchmarks can localize failures through stage-level rubrics, but standard post-training discards these diagnostics before the next training round. We present Benchmark-as-Teacher (BaT), a recursive self-improvement system for agent post-training. BaT contains two linked components: the asynchronous Stage Bank data pipeline and BiCuRL (Bilevel Curriculum Reinforcement Learning), its self-improving post-training method. Stage Bank synthesizes content-isolated training states outside the policy-update loop. BiCuRL uses a fixed held-out evaluation to select the next stage curriculum, verifies rollouts with task rubrics, updates the policy with GRPO, and returns the candidate checkpoint to evaluation. On AutoMedBench-Lite, BaT-4B and BaT-9B more than double the Overall scores of their Qwen Instruct baselines. BaT-9B Agent reaches 79.6 Overall, exceeding Claude Opus 4.6 with Claude Code at 77.5.

摘要:長期代理正在開始自動化完整的工作流程,產生代碼、報告和研究文獻。醫療影像工作流程是多階段且對數據敏感的,而專家軌跡則仍然稀缺且難以分享。結構化基準可以通過階段級別的評分標準定位失敗,但標準的後訓練會在下一輪訓練之前丟棄這些診斷。我們提出了Benchmark-as-Teacher (BaT),這是一個用於代理後訓練的遞歸自我改進系統。BaT包含兩個相互關聯的組件:異步的Stage Bank數據管道和BiCuRL(雙層課程強化學習),其自我改進的後訓練方法。Stage Bank在政策更新循環之外合成內容孤立的訓練狀態。BiCuRL使用固定的保留評估來選擇下一階段的課程,通過任務評分標準驗證回合,使用GRPO更新政策,並將候選檢查點返回給評估。在AutoMedBench-Lite上,BaT-4B和BaT-9B的整體分數超過了其Qwen Instruct基準的兩倍。BaT-9B代理達到79.6的整體分數,超過了Claude Opus 4.6,而Claude Code則為77.5。

Picking the Right Image to Classify: Reliable-Input Selection in Teledermatology

2608.16198v1 by Fabian Gröger, Marco Weishaupt, Philippe Gottfrois, Simone Lionetti, Linda Wermelinger, Nipun Ranasekara, Ludovic Amruthalingam, Alexander A. Navarini, Marc Pouly

Dermatology models face distribution shifts in teledermatology settings, where submitted images differ from the training data in lighting, angle, distance, focus, and framing. These test-time images are ordinary clinical photographs, but some fall outside the model's training conditions, leading the model to often misclassify them due to shifts in acquisition between training and deployment. When multiple images of the same case exist (several photos of one patient or lesion), a natural way to improve accuracy is therefore to select the image the model is most likely to classify correctly. We call this task reliable-input selection. An oracle that, for each case, selects a correctly classified image when one exists raises weighted F1 by about 20 percentage points on average across six dermatology datasets and nine frozen backbones. This oracle is an upper bound that sees the labels, whereas a selector must choose blindly. Capturing this gain in practice is hard. A selector that needs no pretraining data applies to any frozen model, including those whose data is not public. It must judge reliability from quantities the model exposes at inference: its embeddings, their norms, and its confidence. We benchmark four such training-data-free selectors: the embedding norm, the neighborhood consensus among a case's images, the stability of the prediction under small perturbations, and the model's own confidence. No training-data-free selector substantially narrows this oracle gap. The best of them is the model's own confidence, but it recovers only a small part of the gap on the clinical datasets. A small labeled reference set does not help either: the best selector overall, a fusion of confidence and Mahalanobis distance, still leaves most of the gap. To our knowledge, this is the first study to introduce and benchmark reliable input selection, a clinically important, unsolved task.

摘要:皮膚科模型面臨在遠程皮膚科環境中分佈轉移的挑戰,提交的圖像在照明、角度、距離、焦點和構圖上與訓練數據有所不同。這些測試時的圖像是普通的臨床照片,但有些超出了模型的訓練條件,導致模型經常因訓練與部署之間的獲取差異而錯誤分類。當同一病例存在多張圖像(多張同一患者或病變的照片)時,改善準確性的自然方法是選擇模型最有可能正確分類的圖像。我們稱這個任務為可靠輸入選擇。對於每個案例,當存在正確分類的圖像時,選擇這樣的圖像的神諭平均提高六個皮膚科數據集和九個凍結骨幹的加權F1約20個百分點。這個神諭是一個上限,能看到標籤,而選擇器必須盲目選擇。在實踐中捕捉這一增益是困難的。需要無預訓練數據的選擇器適用於任何凍結模型,包括那些數據不公開的模型。它必須根據模型在推斷時暴露的數量來判斷可靠性:其嵌入、它們的範數和模型的信心。我們基準測試了四種無訓練數據的選擇器:嵌入範數、案例圖像之間的鄰域共識、在小擾動下預測的穩定性,以及模型自身的信心。沒有一種無訓練數據的選擇器能顯著縮小這一神諭差距。其中最好的選擇器是模型自身的信心,但它在臨床數據集上僅恢復了差距的一小部分。小型標記參考集也沒有幫助:整體最佳的選擇器,即信心和馬哈拉諾比斯距離的融合,仍然留下了大部分差距。據我們所知,這是第一項引入和基準測試可靠輸入選擇的研究,這是一個臨床重要的未解決任務。

TokenSTFormer: A Tokenized Spatial-temporal Attention Model for Holistic Motion Analysis in Adolescent Idiopathic Scoliosis Screening

2608.16122v1 by Dong Chen, Kenneth M. C. Cheung

Adolescent Idiopathic Scoliosis (AIS) is a prevalent spinal deformity in adolescents that, if left untreated, can result in severe health outcomes. Traditional screening methods are limited by subjective interpretation, reliance on professional expertise and low scalability. To address these challenges, we present ScoliGait dataset, which comprises 1,516 gait video clips paired with corresponding X-ray records. We also introduce TokenSTFormer, a novel model that tokenizes spatial and temporal semantics to enhance feature representation and convergence. Our model achieves state-of-the-art performance, surpassing vanilla Vision Transformer encoder across key metrics, including accuracy of 0.79. This study highlights the potential of leveraging holistic motion features derived from gait video and attention-based models for scalable, cost-effective AIS screening, paving the way for future clinical applications in scoliosis detection.

摘要:青少年特發性脊柱側彎(AIS)是青少年中常見的脊柱畸形,如果不加以治療,可能會導致嚴重的健康後果。傳統的篩檢方法受到主觀解釋、依賴專業知識和低可擴展性的限制。為了解決這些挑戰,我們提出了 ScoliGait 數據集,其中包含 1,516 段行走視頻片段,並配有相應的 X 光記錄。我們還介紹了 TokenSTFormer,一種新型模型,將空間和時間語義進行標記化,以增強特徵表示和收斂。我們的模型在關鍵指標上達到了最先進的性能,超越了普通的視覺Transformer編碼器,包括 0.79 的準確率。本研究突顯了利用從行走視頻中獲得的整體運動特徵和基於注意力的模型進行可擴展、成本效益高的 AIS 篩檢的潛力,為未來脊柱側彎檢測的臨床應用鋪平了道路。

Decoupling Parcellation from Classification: Systematic Benchmark of Fast Brain Segmentation Methods for Alzheimer's Disease Detection

2608.16039v1 by Jiadao Zou, Hongyu Guo, Wei Xi

Brain parcellation and classification are typically evaluated in isolation, yet downstream AD detection performance depends on their interaction. We decouple these components and systematically benchmark fast deep learning parcellation methods (SynthSeg+, OpenMAP-T1) against the FreeSurfer (FS-HV) clinical baseline through down- stream AD classification on OASIS-1. Our factorial design evaluates three parcellation methods, two volumetry strategies (hard vs. soft), and four classifier paradigms (clinical thresholds, supervised feedforward networks, ensemble methods, and foundation models with zero/few-shot prompting), with all results quantified using BCa Bootstrap 95% confidence intervals.

摘要:腦區劃分和分類通常是孤立評估的,但下游阿茲海默症檢測性能取決於它們的相互作用。 我們將這些組件解耦,並系統性地基準測試快速深度學習劃分方法(SynthSeg+、OpenMAP-T1)與 FreeSurfer(FS-HV)臨床基準,通過對 OASIS-1 的下游阿茲海默症分類進行比較。 我們的因子設計評估了三種劃分方法、兩種體積測量策略(硬性與軟性)以及四種分類器範式(臨床閾值、監督式前饋網絡、集成方法和基於零/少量提示的基礎模型),所有結果均使用 BCa Bootstrap 95% 置信區間進行量化。

Breaking and Defending LLM-Powered Social Media Bot Detection Systems

2608.15893v1 by Nof Orenstein, Yoni Birman

The rise of social media bots poses a persistent threat, enabling misinformation, opinion manipulation, and the erosion of trust in online platforms. To combat this, machine learning systems have been developed to detect and limit bot activity, but attackers continuously adapt through techniques such as adversarial learning and behavior imitation, fueling an ongoing arms race between bots and detection tools. Recent advances in large language models (LLMs) have significantly improved bot detection by enabling deeper semantic and contextual analysis of accounts and their content. However, this shift also introduces new attack surfaces, allowing adversaries to craft exploits that directly target the reasoning and generation mechanisms of LLM-based classifiers. Industry tools such as Anthropic's Claude Code Security similarly leverage LLMs for security-critical decisions, further motivating a careful study of their attack surfaces. In this work, we investigate both the offensive and defensive aspects of LLM-powered, threat-specific cybersecurity applications. While centered on the challenge of social media bot detection, our methodology and insights generalize to a broad class of LLM-powered cybersecurity systems, including phishing detection, email classification, and fraud analysis. We introduce two novel adversarial attack strategies that systematically exploit the semantic and contextual weaknesses of LLM-based classifiers, degrading their detection accuracy by up to 48%. To counter these threats, we propose a robust multi-LLM defense architecture designed to preserve detection reliability under adaptive adversarial conditions. Our solution, LSABRE (LLM-powered Social Adversarial Bot Recognition Ensemble), is a multi-LLM framework that substantially improves robustness across a range of attacks, maintaining 86% detection accuracy even under strong, adaptive adversarial pressure.

摘要:社交媒體機器人的興起帶來了持續的威脅,使得錯誤信息、意見操控以及對在線平台的信任侵蝕變得可能。為了應對這一挑戰,已開發出機器學習系統來檢測和限制機器人活動,但攻擊者不斷通過對抗學習和行為模仿等技術進行適應,促進了機器人和檢測工具之間的持續軍備競賽。最近在大型語言模型(LLMs)方面的進展顯著改善了機器人檢測,通過使帳戶及其內容的語義和上下文分析更深入。然而,這一轉變也引入了新的攻擊面,使對手能夠設計直接針對基於LLM的分類器的推理和生成機制的利用方式。行業工具如Anthropic的Claude Code Security同樣利用LLMs進行安全關鍵決策,進一步促使對其攻擊面的仔細研究。在這項工作中,我們調查了基於LLM的特定威脅網絡安全應用的攻擊和防禦兩個方面。雖然重點放在社交媒體機器人檢測的挑戰上,我們的方法論和見解可以推廣到廣泛的基於LLM的網絡安全系統,包括釣魚檢測、電子郵件分類和詐騙分析。我們提出了兩種新穎的對抗攻擊策略,系統性地利用基於LLM的分類器的語義和上下文弱點,將其檢測準確率降低多達48%。為了應對這些威脅,我們提出了一種穩健的多LLM防禦架構,旨在在自適應對抗條件下保持檢測的可靠性。我們的解決方案LSABRE(基於LLM的社交對抗機器人識別集成)是一個多LLM框架,顯著提高了在各種攻擊下的穩健性,即使在強大的自適應對抗壓力下也能保持86%的檢測準確率。

Characterising cardiac tissue properties with graph neural networks

2608.15843v1 by Ching-En Chiu, Yoo Ri Kim, Magdi Saba, Danilo Mandic, Marta Varela

Characterising electrophysiological properties of cardiac tissue efficiently and accurately from spatially sparse intracardiac measurements is clinically important for localising ablation targets and improving arrhythmia treatment. We developed a graph neural network-based framework trained on synthetic electrogram signals on 2D flat surfaces to identify areas of interest in the context of cardiac ablation for premature ventricular complexes (PVCs). Our method achieved an average precision of 0.96, 0.97, and 0.95 for the detection of single-patch fibrosis, rapid depolarisation and high excitability, respectively. The trained model can then be applied to 2D curved surfaces with few-shot fine-tuning, demonstrating its generalisation capability. Future work will develop this framework further for clinical use in PVC ablation.

摘要:有效且準確地從空間稀疏的心內測量中描述心臟組織的電生理特性,對於定位消融目標和改善心律不整治療具有臨床重要性。我們開發了一個基於圖神經網絡的框架,該框架在2D平面上對合成電圖信號進行訓練,以識別在心臟消融中與早期心室複雜(PVCs)相關的興趣區域。我們的方法在檢測單一斑塊纖維化、快速去極化和高興奮性方面,分別達到了0.96、0.97和0.95的平均精度。訓練好的模型可以通過少量調整應用於2D曲面,顯示出其泛化能力。未來的工作將進一步開發這一框架,以便在PVC消融中用於臨床應用。

PLeDO: Pain Level Detection for Osteoarthritis from EMR Data

2608.15719v1 by Yuhao Chen, Jiahao Cai, Nafiz Sadman, Farhana Zulkernine, John Queenan, David Barber

Osteoarthritis (OA) is a progressive chronic joint disease resulting in a breakdown of articular cartilage and bone when damaged joint tissues are not able to normally repair themselves. The aim of this pilot research study is to understand the pain severity for OA from patients' primary care Electronic Medical Records (EMR), both from the structured medical data and the unstructured chart note data using information extraction, natural language processing and machine learning techniques. We propose SPaDe, a Synonym-based Pain level Detection tool to categorize patients into having mild or moderate-to-severe pain to understand diagnosis and treatment methods based on only the pain related expressions in the unstructured chart note. Expressions are subjective, objective, and influenced by cultural background and demography which poses a difficult challenge. Therefore, we improve the model by incorporating the medication information from the structured EMR data and pain scale related information from the chart note to propose an integrated pain level detection tool for OA called PLeDO. With the help of human labeled gold standard data, we demonstrate that both SPaDe and PLeDO can detect mild and moderate-to-severe pain from the EMR data to analyze and potentially improve the quality of care in primary care setting.

摘要:骨關節炎(OA)是一種進行性慢性關節疾病,當受損的關節組織無法正常自我修復時,會導致關節軟骨和骨骼的破壞。這項初步研究的目的是了解來自患者初級保健電子病歷(EMR)的OA疼痛嚴重程度,包括結構化醫療數據和使用信息提取、自然語言處理及機器學習技術的非結構化病歷筆記數據。我們提出了SPaDe,一種基於同義詞的疼痛程度檢測工具,用以將患者分類為輕度或中度至重度疼痛,以便根據非結構化病歷筆記中的疼痛相關表達來理解診斷和治療方法。表達是主觀的、客觀的,並受到文化背景和人口統計的影響,這帶來了困難的挑戰。因此,我們通過整合結構化EMR數據中的用藥信息和病歷筆記中的疼痛量表相關信息來改進模型,提出了一種名為PLeDO的OA綜合疼痛程度檢測工具。在人工標記的金標準數據的幫助下,我們證明SPaDe和PLeDO都能從EMR數據中檢測輕度和中度至重度疼痛,以分析並潛在改善初級保健環境中的護理質量。

Integrating Persuasion Theory into the Epidemiological Modelling of Health Misinformation Spread on Social Media

2608.15689v1 by Mkululi Sikosana, Sean Maudsley-Barton, Oluwaseun Ajao

This study presents a hybrid epidemiological and behavioural framework to simulate the spread of health misinformation on social media. We extend the classical Susceptible--Infected--Recovered (SIR) model to a six-compartment structure (SIRMMM), incorporating Misinformed Susceptible (MS), Misinformed Infected (MI), and Misinformed Recovered (MR) compartments to better reflect the dynamics of the misinformation lifecycle. To account for individual-level behavioural variation, we extend the SIRMMM model by integrating psychological signals from the Elaboration Likelihood Model (ELM), including sentiment polarity, engagement metrics, and cognitive effort, which dynamically modulate the misinformation transmission rate, yielding the ELM-SIRMMM framework. Model parameters were estimated using the FibVID dataset, which captures COVID-19 misinformation on Twitter. Generalisability was tested on two additional datasets: MC-Fake (emotional misinformation) and Monant (general health misinformation). Results show that the ELM-SIRMMM model enhances both predictive accuracy and dynamic realism. On FibVID, it decreases RMSE by 5.5%, delays the misinformation peak from day 150 to day 160, and increases its peak prevalence from 6% to 7%. On MC-Fake, it accurately reproduces a flash-rumour pattern, infecting 38% of users by day 45 and achieving 97% misinformation recovery, all while maintaining model accuracy. In contrast, minimal behavioural signal variability in the Monant dataset leads to marginal benefit, with only a 3% peak and 57% of users remaining susceptible. These findings suggest that structural elaboration alone is insufficient. Functional realism in modelling misinformation spread requires dynamic psychological inputs that vary meaningfully across time and contexts.

摘要:這項研究提出了一個混合流行病學和行為框架,以模擬健康錯誤資訊在社交媒體上的傳播。我們將傳統的易感--感染--康復(SIR)模型擴展為六個區隔的結構(SIRMMM),納入了錯誤資訊易感者(MS)、錯誤資訊感染者(MI)和錯誤資訊康復者(MR)區隔,以更好地反映錯誤資訊生命週期的動態。為了考慮個體層面的行為變異,我們通過整合來自精緻可能性模型(ELM)的心理信號來擴展SIRMMM模型,包括情感極性、參與度指標和認知努力,這些信號動態調節錯誤資訊的傳播速率,形成ELM-SIRMMM框架。模型參數使用FibVID數據集進行估算,該數據集捕捉了Twitter上的COVID-19錯誤資訊。可推廣性在另外兩個數據集上進行測試:MC-Fake(情感錯誤資訊)和Monant(一般健康錯誤資訊)。結果顯示,ELM-SIRMMM模型提高了預測準確性和動態真實性。在FibVID上,它將均方根誤差(RMSE)降低了5.5%,將錯誤資訊的高峰從第150天延遲到第160天,並將其高峰流行率從6%提高到7%。在MC-Fake上,它準確再現了一種快速謠言模式,到第45天感染了38%的用戶,並實現了97%的錯誤資訊康復,同時保持模型的準確性。相比之下,Monant數據集中行為信號變異性最小,僅帶來邊際效益,只有3%的高峰和57%的用戶仍然易感。這些發現表明,僅僅結構上的精緻是不夠的。在模擬錯誤資訊傳播時,功能真實性需要隨時間和情境有意義變化的動態心理輸入。

From Generalist to Specialist: A Context-Fusion Framework for Endoscopic Polyp Reporting with a Frozen VLM

2608.15580v1 by Ruijie Yang, Yan Zhu, Peiyao Fu, Siyuan Li, Te Luo, Zhihua Wang, Quanlin Li, Pinghong Zhou, Xian Yang, Shuo Wang

Reliable endoscopic polyp reporting requires integrating quantitative lesion sizing, standardized Paris classification, and clinically meaningful morphological description within a single record. General-purpose vision-language models (VLMs) offer a unified interface for image understanding and report generation. Existing specialization strategies, however, typically rely on task-specific models or model-weight adaptation, leaving unresolved how to introduce reliable specialist knowledge while preserving both this unified interface and the VLM's pretrained capabilities. We introduce a context-fusion framework that specializes a frozen general-purpose VLM through both implicit instruction context and explicit transduction context without modifying its pretrained weights. Specifically, a self-supervised polyp encoder retrieves related image-report pairs as explicit, query-specific evidence, while learned continuous specialist tokens provide implicit instruction context shared across cases. Experiments were conducted on 2,056 expert-annotated public endoscopic images. We compared the framework with general-purpose VLMs, task-specific predictors, and weight-adaptation methods to assess specialist performance, unified reporting, and adaptation efficiency. Across numerical, categorical, and report-generation metrics, the proposed framework substantially improved direct frozen-VLM inference and achieved the strongest overall performance among the evaluated methods. It added trainable parameters equal to only 0.006% of the frozen VLM's parameter count. When the top-1 retrieved case carried the correct target category, our framework corrected 70.5% of the errors made by a weight-adaptation baseline. These findings support the context-fusion framework as a lightweight and effective strategy for specialist adaptation of a frozen VLM.

摘要:可靠的內視鏡息肉報告需要將定量病變大小、標準化的巴黎分類和臨床上有意義的形態描述整合在單一記錄中。通用視覺-語言模型(VLMs)提供了一個統一的圖像理解和報告生成界面。然而,現有的專業化策略通常依賴於特定任務的模型或模型權重調整,尚未解決如何在保留這一統一界面和VLM的預訓練能力的同時引入可靠的專家知識。我們提出了一個上下文融合框架,通過隱式指令上下文和顯式轉導上下文專門化一個凍結的通用VLM,而不修改其預訓練權重。具體而言,自監督的息肉編碼器檢索相關的圖像-報告對作為顯式的查詢特定證據,而學習的連續專家標記提供了在案例之間共享的隱式指令上下文。實驗在2,056張專家標註的公共內視鏡圖像上進行。我們將該框架與通用VLMs、特定任務的預測器和權重調整方法進行比較,以評估專家性能、統一報告和適應效率。在數值、類別和報告生成指標上,所提出的框架顯著改善了直接凍結VLM推理,並在評估的方法中實現了最強的整體性能。它增加的可訓練參數僅佔凍結VLM參數總數的0.006%。當檢索到的頂級案例攜帶正確的目標類別時,我們的框架修正了70.5%的權重調整基線所犯的錯誤。這些發現支持上下文融合框架作為一種輕量且有效的策略,用於凍結VLM的專家適應。

EA-LiteUNet: An Edge-Adaptive and Resource-Efficient U-Net for Boundary-Sensitive Dermoscopic Image Segmentation

2608.15537v1 by Wang Jiangtao, Nur Intan Raihana Ruhaiyem, Fu Panpan, Yang Yu, Huang Yan

Accurate boundary delineation remains a persistent challenge in dermoscopic image segmentation because of blurred lesion margins, heterogeneous textures, and complex background artifacts. From a signal-processing perspective, lesion boundaries represent high-frequency components that are highly susceptible to aliasing, noise amplification, and information loss. Consequently, repeated downsampling and feature transformations in conventional convolutional architectures often lead to severely degraded boundary representations. To address these limitations, we propose EA-LiteUNet, an edge-adaptive and computationally efficient U-Net variant specifically designed for boundary-sensitive medical image segmentation. The architecture integrates three core mechanisms: (1) boundary-aware representation learning to suppress aliasing and preserve high-frequency structural details; (2) attention-guided feature modulation to selectively enhance boundary-relevant responses across multi-scale features; and (3) a resource-adaptive inference strategy to dynamically balance segmentation accuracy and computational efficiency. Extensive evaluations across three public dermoscopic datasets demonstrate that EA-LiteUNet consistently achieves superior boundary precision. Specifically, on the ISIC 2018 dataset, the method significantly reduces the 95% Hausdorff Distance (HD95) to 12.89 pixels while maintaining a robust Dice score of 92.08%. Notably, this strong performance is achieved with an ultralightweight configuration of merely 0.29M parameters and 1.17 GFLOPs. Ablation studies further validate the complementary effects of these components, confirming their contribution to enhanced boundary fidelity and stable optimization.

摘要:準確的邊界劃分在皮膚鏡影像分割中仍然是一個持續的挑戰,因為病變邊緣模糊、紋理異質以及複雜的背景伪影。從信號處理的角度來看,病變邊界代表著高頻成分,這些成分對混疊、噪聲放大和信息損失非常敏感。因此,傳統卷積架構中的重複下採樣和特徵轉換往往導致邊界表示的嚴重退化。為了解決這些限制,我們提出了EA-LiteUNet,一種邊緣自適應且計算效率高的U-Net變體,專門設計用於對邊界敏感的醫學影像分割。該架構整合了三個核心機制:(1)邊界感知的表示學習,以抑制混疊並保留高頻結構細節;(2)注意力引導的特徵調制,以選擇性地增強多尺度特徵中的邊界相關響應;以及(3)資源自適應的推理策略,以動態平衡分割準確性和計算效率。在三個公共皮膚鏡數據集上的廣泛評估顯示,EA-LiteUNet始終實現了卓越的邊界精度。具體而言,在ISIC 2018數據集上,該方法將95%豪斯多夫距離(HD95)顯著降低至12.89像素,同時保持穩健的Dice得分92.08%。值得注意的是,這一強勁的表現是在僅有0.29M參數和1.17 GFLOPs的超輕量配置下實現的。消融研究進一步驗證了這些組件的互補效果,確認它們對增強邊界忠實度和穩定優化的貢獻。

2608.15428v1 by Volodymyr Ovcharov

Multiple-choice benchmarks are graded on whether a model picks the right option, not on whether it needed the question. Measuring that gap takes care: a model answering A to most items scores above chance wherever the key sits at A, and reads as recognition when it is not. We measure it on UA-JudgeExam: 11,990 four-option items with official keys, published by Ukraine's Higher Qualification Commission of Judges. Shown the options and no question, Claude Haiku 4.5 scores 0.383 against chance, and the leak is concentrated: 11.8% of items are answered blind on all eight option orders, against 0.2 items expected by chance. It is not quotation: search over 280,059 editions of Ukrainian legislation recovers 0.128. Gating those out retains 8,128 items, on which the gating model itself now scores 0.204, and GPT-5.6, which took no part in the selection, still answers 0.515 of them with the question hidden. Scoring twelve held-out models on the whole set and subtracting each one's answer-position habit, only two keep an excess: GPT-5.6 at +0.265, Sonnet 4.6 at +0.081. Without it the ranking misleads: Llama 3.1 8B scores 0.292 blind, above every model but those two, purely by answering A to 92% of items. The gate does select something real: on the items it rejected, eleven of twelve models score 0.518-0.789, every interval clear of what the same model scores on the items it kept. But that signal is one model's, and filtering on it does not transfer upward. Neither is visible on a 400-item sample, where nine models read as "statistically at chance". Rewriting distractors instead overshoots to 0.168, below chance and as exploitable. The same probe on LEXam returns chance: every option there points into the stem, none longer than 33 characters. Item format decides whether the problem can arise; capability decides how much is extracted. We release the corpus, the predictions and the harness.

摘要:多選基準是根據模型是否選擇正確選項來評分,而不是根據它是否需要問題。測量這個差距需要小心:一個模型在大多數項目中回答A,無論關鍵在A的位置如何,都會得分高於隨機,而當關鍵不在A時則顯示為識別。我們在UA-JudgeExam上進行測量:11,990個四選項目,擁有官方答案,由烏克蘭高級法官資格委員會發布。當顯示選項而沒有問題時,Claude Haiku 4.5的得分為0.383,超過隨機,而洩漏集中在一起:11.8%的項目在所有八個選項順序中都是盲目回答,預期隨機回答為0.2項目。這不是引用:搜索超過280,059份烏克蘭立法的版本恢復了0.128。排除這些後保留了8,128個項目, gating模型本身現在在這些項目上得分為0.204,而GPT-5.6沒有參與選擇,仍然在問題隱藏的情況下回答了其中的0.515。對整個數據集進行十二個保留模型的得分並減去每個模型的答案位置習慣,只有兩個保持超額:GPT-5.6為+0.265,Sonnet 4.6為+0.081。沒有這個,排名會誤導:Llama 3.1 8B在盲測中得分0.292,超過每個模型,但只有這兩個,純粹是因為對92%的項目回答A。這個gating確實選擇了一些真實的東西:在被拒絕的項目中,十二個模型中的十一個得分為0.518-0.789,每個區間都清楚地與同一模型在保留項目上的得分不同。但那個信號是某一模型的,基於它的過濾並不會向上轉移。在400項樣本中也不可見,九個模型的表現為“統計上隨機”。重寫干擾項反而超出到0.168,低於隨機且可被利用。對LEXam的相同探測返回隨機:那裡的每個選項都指向題幹,沒有一個超過33個字符。項目格式決定問題是否會出現;能力決定提取的多少。我們釋放語料庫、預測和工具。

ETHOS: Towards a Modular Ethics Framework for Clinical Multi-Agent Systems

2608.15424v1 by Rakesh Sharma, Sydney Pugh, Cameron Beeche, Pankhuri Singhal, Rachel Wu, Margaret Eby, Jeffrey Duda, James Gee, Kyra O'Brien, Hersh Sagreiya, Marina Serper, Victoria Gershuni, Angela Bradbury, Anurag Verma, Eric Eaton, Kevin B. Johnson, Walter Witschey

The rapid adoption of large language models has enabled the development of clinical multi-agent systems (MAS) capable of integrating multimodal patient data and supporting increasingly complex clinical decision-making. However, the deployment of these systems in real-world healthcare settings raises critical ethical concerns related to safety, fairness, accountability, transparency, and patient trust. While numerous organizations, including the World Health Organization, the National Academy of Medicine, and the FUTURE-AI consortium, have proposed ethical frameworks and governance principles for healthcare AI, these efforts remain largely conceptual. To address this challenge, we present ETHOS (Ethics and Trust through Hierarchical Oversight System), a modular ethics framework designed as a governance meta-agent that can be integrated with any existing multi-agent system without requiring changes to its underlying architecture. ETHOS translates stakeholder-informed ethical requirements into executable runtime oversight through a layered governance approach consisting of deterministic checks, contextual reviews, and a final ethics critic. These components continuously evaluate intermediate reasoning steps and final outputs, enabling the system to identify ethical risks, request revisions, or suppress responses that fail predefined safety and trustworthiness criteria. We demonstrate ETHOS within a hepatology clinical decision-support MAS. Results show that ETHOS improves decision reliability by detecting incomplete, inconsistent, or out-of-scope evidence and appropriately increasing abstention when safe recommendations cannot be supported. By embedding ethical governance directly into system operation, ETHOS provides a practical and auditable mechanism for transforming high-level AI ethics principles into deployable safeguards.

摘要:大型語言模型的快速採用使得臨床多代理系統(MAS)的發展成為可能,這些系統能夠整合多模態病人數據並支持日益複雜的臨床決策。然而,這些系統在現實世界醫療環境中的部署引發了與安全、公平、問責、透明度和病人信任相關的重大倫理問題。儘管包括世界衛生組織、國家醫學院和FUTURE-AI聯盟在內的許多組織已經提出了針對醫療AI的倫理框架和治理原則,但這些努力仍然主要是概念性的。為了解決這一挑戰,我們提出了ETHOS(通過分層監督系統實現倫理與信任),這是一個模塊化的倫理框架,設計為一個治理元代理,可以與任何現有的多代理系統集成,而無需改變其底層架構。ETHOS將利益相關者所知的倫理要求轉化為可執行的運行時監督,通過一種分層治理方法,包括確定性檢查、上下文審查和最終倫理評估。這些組件持續評估中間推理步驟和最終輸出,使系統能夠識別倫理風險、請求修訂或抑制不符合預定安全和可信標準的回應。我們在一個肝病臨床決策支持MAS中展示了ETHOS。結果顯示,ETHOS通過檢測不完整、不一致或超出範疇的證據來提高決策的可靠性,並在無法支持安全建議時適當地增加放棄。通過將倫理治理直接嵌入系統運作中,ETHOS提供了一種實用且可審計的機制,將高層次的AI倫理原則轉化為可部署的保障措施。

Invariant Pretraining for Robust Code Representations

2608.15412v1 by Yifeng He, Yundi Xu, Christopher Castro Gaw Gonzalo, Zili Wang, Hao Chen

Encoder-based code representation models remain widely deployed for discriminative tasks such as clone detection and code classification, where their small size and low inference cost are decisive. Their robustness, however, is fragile: under invariant programs, semantically equivalent code written in different syntactic forms, learned representations degrade substantially even though program behavior is unchanged. We present an empirical study of this robustness gap across four encoder baselines, two downstream tasks, and four datasets, together with a minimal code-only continued pretraining recipe that closes much of it. Our method, invariant pretraining (InvPT), applies semantic-preserving transformations to the corpus and combines masked language modeling with multi-positive supervised contrastive learning that treats all augmentations of the same source function as positives, mixing self-contrast pairs (same code, different masks) with invariant-contrast pairs (transformed code) for positives of varying difficulty. Unlike prior contrastive code encoders, InvPT does not require paired natural-language data. Across our evaluation, InvPT improves robustness on transformed test sets by up to 11 percentage points on clone detection and 19 on code classification while matching or improving standard accuracy, and our ablations isolate multi-positive invariant contrast as the main source of the gains. Our aim is not a new objective but a careful measurement of where encoder robustness breaks and how far a simple, code-only recipe can recover it.

摘要:編碼器基礎的代碼表示模型仍然廣泛應用於克隆檢測和代碼分類等判別性任務,其中其小巧的尺寸和低推理成本是決定性的。然而,它們的穩健性卻是脆弱的:在不變的程序下,以不同語法形式編寫的語義等價代碼,其學習到的表示會大幅退化,即使程序行為未改變。我們針對四個編碼器基準、兩個下游任務和四個數據集進行了這一穩健性差距的實證研究,並提出了一種最小的僅代碼持續預訓練配方,能夠縮小這一差距。我們的方法,稱為不變預訓練(InvPT),對語料庫應用語義保留轉換,並將掩碼語言建模與多正樣本監督對比學習相結合,將同一源函數的所有增強視為正樣本,將自對比對(相同代碼,不同掩碼)與不變對比對(轉換代碼)混合,形成不同難度的正樣本。與之前的對比代碼編碼器不同,InvPT 不需要配對的自然語言數據。在我們的評估中,InvPT 在轉換測試集上提升了克隆檢測的穩健性達 11 個百分點,代碼分類達 19 個百分點,同時保持或提高標準準確性,而我們的消融實驗則將多正樣本不變對比確定為增益的主要來源。我們的目標不是一個新的目標,而是仔細測量編碼器穩健性破裂的地方,以及一個簡單的僅代碼配方能夠恢復的程度。

Grounding Healthcare LLMs in a Causal Knowledge Graph: Framework, Metrics, and a Cardiovascular Pilot

2608.15382v1 by Ummara Mumtaz, Aimen Noor, Awais Ahmed

Large language models (LLMs) are increasingly proposed for healthcare decision support, but their evaluations still reward single-answer accuracy rather than reasoning about interventions, mechanisms, harms, evidence, and uncertainty. We propose a reproducible, graph-centered evaluation framework for intervention-oriented LLM behavior in healthcare and stress-test it in a cardiovascular pilot. The framework has four components: (i) a domain causal knowledge graph in which assertions are first-class, provenance-preserving nodes with stable identifiers; (ii) a scenario-conditioned subgraph extraction step that, given any clinical scenario, retrieves the relevant reified-assertion subgraph; (iii) four controlled grounding conditions that vary how the retrieved subgraph is composed into the model's context (ungrounded C1, knowledge-graph C2, causal-graph C3, integrated C4); and (iv) an automated scoring pipeline, anchored on assertion identifiers, that computes intervention accuracy, and other evaluation measures on a single pass. To test the framework, we built a category-balanced scenario generator across eight reasoning failure modes and instantiated it on a cardiovascular graph. The metric panel discriminates conditions along interpretable, non-redundant axes: C4 obtains the strongest causal edge F1 (0.838), adverse-effect F1 (0.833), evidence accuracy (0.738), and unsupported claim rate (0.114), while C1 obtains the highest raw intervention accuracy (0.948) with no measurable causal or evidential grounding.

摘要:大型語言模型(LLMs)越來越多地被提議用於醫療決策支持,但其評估仍然獎勵單一答案的準確性,而不是對干預、機制、危害、證據和不確定性進行推理。我們提出了一個可重複的、以圖為中心的評估框架,用於醫療保健中的干預導向LLM行為,並在心血管試點中進行壓力測試。該框架有四個組成部分:(i)一個領域因果知識圖,其中斷言是第一類的、保持來源的節點,具有穩定的標識符;(ii)一個情境條件的子圖提取步驟,根據任何臨床情境檢索相關的具體化斷言子圖;(iii)四個控制的基礎條件,變化檢索到的子圖如何組成模型的上下文(未基礎的C1、知識圖C2、因果圖C3、整合的C4);以及(iv)一個自動評分管道,以斷言標識符為基礎,計算干預準確性和其他評估指標,僅需一次通過。為了測試該框架,我們建立了一個跨越八種推理失敗模式的類別平衡情境生成器,並在心血管圖上實現了它。該指標面板沿著可解釋的、非冗餘的軸區分條件:C4獲得最強的因果邊緣F1(0.838)、不良影響F1(0.833)、證據準確性(0.738)和不支持的主張率(0.114),而C1獲得最高的原始干預準確性(0.948),卻沒有可測量的因果或證據基礎。

When AI Rewrites, Classifiers Relax: Uncertainty-Aware Sentiment Analysis on Sarcastic and AI-Paraphrased Social Text

2608.15338v1 by Shresth Shroff

Sentiment classifiers are increasingly applied to social media content that is either sarcastic or AI-generated --- two distributional regimes where standard evaluations offer little guidance. We present a three-part empirical study of sentiment classifier behaviour under these conditions. First, we find that confidence scores on sarcastic text are significantly lower than on non-sarcastic text (Mann--Whitney $p = 2 \times 10^{-6}$), confirming that classifiers sense their own uncertainty on ironic content even without explicit uncertainty modelling. Second, and counterintuitively, we show that sentiment classifiers achieve higher accuracy on AI-paraphrased reviews than on the original human-authored text (RoBERTa: $+5.8$ pp for Qwen3.5-4B paraphrases, $+3.7$ pp for Gemma4-E4B), revealing a cross-domain stylistic alignment effect: AI paraphrases remove distributional noise that confounds Twitter-trained classifiers, producing cleaner, more prototypical sentiment text. Third, we demonstrate that a lightweight abstention wrapper --- flagging the $14\%$ of inputs with confidence below $0.6$ --- improves accuracy from 82.2\% to 88.9\% ($+6.7$ pp) on the retained set. We further compare Semantic Entropy and MC-Dropout-style disagreement as uncertainty signals and find near-identical AUROC ($0.650$ vs.\ $0.646$) on sarcastic text, suggesting that for short social media inputs, both methods are interchangeable. Our results motivate a shift from confident single-label prediction to uncertainty-aware abstention in high-stakes sentiment applications such as mental health flagging and content moderation.

摘要:情感分類器越來越多地應用於諷刺或 AI 生成的社交媒體內容——這兩種分佈模式下,標準評估提供的指導有限。我們提出了一項三部分的實證研究,探討情感分類器在這些條件下的行為。首先,我們發現對於諷刺文本的信心分數顯著低於非諷刺文本(Mann--Whitney $p = 2 \times 10^{-6}$),確認分類器即使在沒有明確不確定性建模的情況下,也能感知到對於諷刺內容的自身不確定性。其次,反直覺的是,我們顯示情感分類器在 AI 改寫的評論上取得的準確率高於原始的人類撰寫文本(RoBERTa: $+5.8$ pp 對於 Qwen3.5-4B 改寫,$+3.7$ pp 對於 Gemma4-E4B),揭示了一種跨領域的風格一致性效應:AI 改寫去除了困擾 Twitter 訓練的分類器的分佈噪音,產生了更乾淨、更原型的情感文本。第三,我們證明了一種輕量級的棄權包裝——標記信心低於 $0.6$ 的 $14\%$ 輸入——在保留數據集上將準確率從 82.2\% 提高到 88.9\%($+6.7$ pp)。我們進一步比較語義熵和 MC-Dropout 風格的不一致作為不確定性信號,發現對於諷刺文本,兩者的 AUROC 幾乎相同($0.650$ 對 $0.646$),這表明對於短的社交媒體輸入,這兩種方法是可以互換的。我們的結果促使從自信的單標籤預測轉向在高風險情感應用中,如心理健康標記和內容審核,意識到不確定性的棄權。

Physiological World Models for Human State Transitions

2608.15309v1 by Chongyang Zhang, Rendong Wang, Hao Zheng, Hanwen Zhang, Yang Liu, Xiaolong Wei, Bin Chong

Continuous multimodal sensing now allows human physiology to be observed throughout daily life rather than only during occasional clinical visits. However, most health artificial intelligence systems are designed to recognize current states, estimate risks or analyse individual biomarkers. They do not directly model how physiological states change in response to real-world events, behaviours, contexts and interventions. Here we propose the Physiological World Model (PWM), an event-conditioned framework for learning these changes at the level of the whole person. We introduce the HumanState Transition Token, a structured, quality-scored unit that connects the physiological state before an event with the event or action, relevant context and intervention information, the physiological trajectory after the event, observed outcomes and data quality. We describe four capability levels, from state representation to bounded intervention planning, together with four data acquisition and validation protocols. We also propose six benchmark tasks covering HumanState representation, forecasting across multiple timescales, individualized response prediction, simulation of alternative interventions, bounded planning and reliability under distribution shift. Together, this framework provides a practical path towards personalized health management, behavioural intervention design and clinician-supervised decision support, while clearly separating prediction from causal inference and making uncertainty, safety, governance and limits of use explicit.

摘要:持續的多模態感測現在允許在日常生活中觀察人類生理,而不僅僅是在偶爾的臨床訪問中。然而,大多數健康人工智慧系統旨在識別當前狀態、評估風險或分析個體生物標記。它們並未直接建模生理狀態如何對現實世界事件、行為、情境和干預進行變化。在這裡,我們提出生理世界模型(PWM),這是一個事件條件框架,用於學習整個人的這些變化。我們介紹了人類狀態轉換標記,這是一個結構化的、質量評分的單位,將事件前的生理狀態與事件或行動、相關情境和干預信息、事件後的生理軌跡、觀察到的結果和數據質量相連接。我們描述了四個能力層級,從狀態表示到有界的干預規劃,以及四個數據獲取和驗證協議。我們還提出了六個基準任務,涵蓋人類狀態表示、跨多個時間尺度的預測、個性化反應預測、替代干預的模擬、有界規劃和在分佈轉移下的可靠性。總體而言,這個框架提供了一條實用的途徑,朝向個性化健康管理、行為干預設計和臨床醫生監督的決策支持,同時明確區分預測與因果推斷,並使不確定性、安全性、治理和使用限制變得明確。

Demographic Injection in Medical Language Models under Diversity, Equity, and Inclusion Prompts

2608.15254v1 by Diego Mardian, Frank Liu

Clinical-AI guidance increasingly recommends prompting language models to reason with attention to diversity, equity, and inclusion (DEI). We measure a side effect that misrepresents patients: a one-sentence DEI prompt appended to a medical question leads models to add patient demographic attributes (race, socioeconomic status, sex) the question never stated, in effect rewriting who the patient is. We call this demographic injection. Across 47 models, four medical benchmarks, and 376,000 responses scored by a validated model-judge pipeline, a single DEI prompt raises the injection rate from 0.7% to 33.1% (47x) in all 47 of 47 models, attributable to the equity content rather than to added length (18x above a length-matched control; p=1.4x10^-14). Most added content is a general population statement that leaves the answer unchanged, but a smaller subset attaches an attribute to the specific patient or changes the selected option (0.25-2.4% of responses, 99.8% toward the incorrect option), where the invented demographic changes the answer the model recommends. Phrasing scales the effect from 14% to 56%. DEI prompts are just one example of a more general mechanism. Any instruction that nudges how a model reasons can make it add unrequested details, including details about the patient. Flagged outputs are treated as model errors under study, not clinical guidance.

摘要:臨床人工智慧指導越來越多地建議促使語言模型在考慮多樣性、公平性和包容性(DEI)時進行推理。我們測量了一種誤導患者的副作用:一個附加在醫療問題上的單句DEI提示會導致模型添加問題中從未提到的患者人口統計屬性(種族、社會經濟地位、性別),實際上重寫了患者的身份。我們稱之為人口統計注入。在47個模型、四個醫療基準和376,000個由經過驗證的模型評判管道評分的回應中,單一的DEI提示使得注入率從0.7%上升到33.1%(47倍),這是由於公平性內容而非增加的長度(在長度匹配的對照組中增加了18倍;p=1.4x10^-14)。大多數新增內容是一般人口的陳述,對答案沒有改變,但一小部分則將屬性附加到特定患者或改變所選選項(0.25-2.4%的回應,99.8%朝向不正確的選項),其中虛構的人口統計改變了模型推薦的答案。措辭將效果擴大至14%至56%。DEI提示僅是更一般機制的一個例子。任何促使模型推理的指令都可能使其添加未請求的細節,包括有關患者的細節。被標記的輸出被視為正在研究的模型錯誤,而非臨床指導。

Translating finite-domain integer constraint models to CP/SMT/ILP/PB/SAT solvers with CPMpy

2608.15143v1 by Tias Guns, Ignace Bleukx, Hendrik Bierlee, Jo Devriendt, Emilio Gamba, Orestis Lomis, Wout Piessens, Thomas Sergeys, Dimos Tsouros, Wout Vanroose, Hélène Verhaeghe

Constraint solving is a declarative approach for solving combinatorial satisfaction and optimization problems. The user specifies their problem through constraints and decision variables, and a generic solver is used to find a solution. Several constraint-solving technologies exist, and certain solvers perform well on certain problems. Therefore, it is useful to try different solvers given a particular application. However, each solving paradigm supports different types of constraints and decision variables. Our goal is to translate high-level constraint satisfaction and optimization problems into any lower-level formalism, including CP, SMT QF-LIA, ILP, PB and (Max)SAT. This allows for comparing different solving technologies for a particular problem, without requiring a user to manually remodel it for each solving paradigm. We define a high-level language of logical and arithmetic operations, and useful additional functions and constraints, which are known as global constraints in the CP community. We then present a modular framework for transforming our high-level modeling language to CP/SMT/ILP/PB and (Max)SAT solvers. While many transformations are partly described in the literature, we observe that they can be implemented through a modular waterfall of smaller components, where lower-level paradigms reuse the transformations of higher-level paradigms. Two recurring challenges are handling the negation of arbitrary subexpressions and avoiding the introduction of auxiliary variables. Additionally, we take special care linearizing non-linear operators for ILP, PB and SAT-solvers. The transformation waterfall is implemented and evaluated in the open-source CPMpy library. Our results show that constraint models significantly change throughout the transformations, and that optimizations to the linearization of constraints are essential for ILP and PB solvers.

摘要:限制求解是一種聲明式方法,用於解決組合滿足和優化問題。用戶通過約束和決策變量來指定他們的問題,並使用通用求解器來尋找解決方案。存在幾種限制求解技術,某些求解器在特定問題上表現良好。因此,針對特定應用嘗試不同的求解器是有用的。然而,每種求解範式支持不同類型的約束和決策變量。
我們的目標是將高級約束滿足和優化問題轉換為任何低級形式,包括 CP、SMT QF-LIA、ILP、PB 和 (Max)SAT。這使得在特定問題上比較不同的求解技術成為可能,而無需用戶手動為每個求解範式重新建模。
我們定義了一種邏輯和算術運算的高級語言,以及一些有用的附加函數和約束,這些在 CP 社區中被稱為全局約束。我們接著提出了一個模塊化框架,用於將我們的高級建模語言轉換為 CP/SMT/ILP/PB 和 (Max)SAT 求解器。雖然許多轉換在文獻中部分描述,但我們觀察到它們可以通過一個模塊化的瀑布式小組件來實現,其中低級範式重用高級範式的轉換。兩個反覆出現的挑戰是處理任意子表達式的否定和避免引入輔助變量。此外,我們特別注意將非線性運算符線性化以適應 ILP、PB 和 SAT 求解器。
轉換瀑布在開源的 CPMpy 庫中實現和評估。我們的結果顯示,約束模型在轉換過程中顯著改變,並且對約束線性化的優化對於 ILP 和 PB 求解器至關重要。

FZ-VLM: A Two Stage Florence-Zephyr Vision Language Model Framework for Pulmonary Nodule Characterization and Clinical Decision Making

2608.15004v1 by Pramit Dutta, Jenita Manokaran, Richa Mittal, Ryan Appleby, Eranga Ukwatta

Lung cancer remains one of the leading causes of cancer-related mortality worldwide, and Computed Tomography (CT) is a primary imaging tool for screening and followup assessment. After pulmonary nodule detection, radiologists manually assess anatomical location, diameter, margin characteristics, and attenuation type to support risk assessment and clinical decision-making. However, this post-detection workflow is time-consuming and can be affected by inter-observer variability. Existing Artificial Intelligence methods often focus on isolated tasks, limiting their use as a unified, clinically grounded interpretation framework. This study presents FZ-VLM, a two-stage Florence-Zephyr Vision Language Model framework for unified structured pulmonary nodule characterization in lung CT. The framework uses a fine-tuned Florence-2 model to extract radiological attributes from expert-annotated 2D axial CT slices, while a Zephyr-7B model uses these attributes to generate nodule descriptions, follow-up recommendations, and longitudinal analyses. Results showed that the Stage 1 model achieved 77.18\% accuracy for anatomical location, 67.96\% accuracy for margin characteristics, and 79.13\% accuracy for attenuation type, with a Mean Absolute Error of 2.58 mm for diameter estimation, outperforming evaluated GPT-4-based baselines as well as the human baseline. Expert radiologist evaluation of Stage 2 showed 93.9\% accuracy, 98.6\% completeness score, 76.1\% clinical relevance, and an overall score of 89.5\%. Safety analysis showed that most outputs were clinically safe, although some follow-up recommendations still required expert review. To the best of our knowledge, this study presents the first two-stage Vision-Language Model framework for structured nodule characterization and clinical decision-making.

摘要:肺癌仍然是全球癌症相關死亡的主要原因之一,而計算機斷層掃描(CT)是篩查和後續評估的主要影像工具。在檢測到肺結節後,放射科醫生手動評估解剖位置、直徑、邊緣特徵和衰減類型,以支持風險評估和臨床決策。然而,這一檢測後的工作流程耗時且可能受到觀察者之間變異性的影響。現有的人工智慧方法通常專注於孤立的任務,限制了它們作為統一的、臨床基礎的解釋框架的使用。本研究提出了FZ-VLM,一個兩階段的Florence-Zephyr視覺語言模型框架,用於肺CT中統一的結構化肺結節特徵描述。該框架使用微調的Florence-2模型從專家標註的2D軸向CT切片中提取放射學屬性,而Zephyr-7B模型則利用這些屬性生成結節描述、後續建議和縱向分析。結果顯示,第一階段模型在解剖位置的準確率達到77.18\%、邊緣特徵的準確率為67.96\%、衰減類型的準確率為79.13\%,直徑估計的平均絕對誤差為2.58毫米,超越了評估的基於GPT-4的基準以及人類基準。第二階段的專家放射科醫生評估顯示準確率為93.9\%、完整性得分為98.6\%、臨床相關性為76.1\%,總體得分為89.5\%。安全性分析顯示,大多數輸出在臨床上是安全的,儘管一些後續建議仍需專家審查。據我們所知,本研究提出了首個兩階段的視覺-語言模型框架,用於結構化結節特徵描述和臨床決策。

Evaluating Agentic Code Repair Capabilities in Distributed Systems

2608.14863v1 by Yibo Yan, Huijuan Wang, Junzhou He, Yizhuo Liang, Shaoyu Wang, Huanchen Sun, Seo Jin Park

LLM-based coding agents have advanced rapidly on single-process SWE tasks, with frontier models now clustering in the high-70s on SWE-bench Verified. Distributed-system debugging, however, remains an under-explored regime: bugs span processes, nodes, and protocol interactions, with root causes rarely recoverable from source alone and brute-force exploration intractable across non-deterministic interleavings. This leaves two gaps in LLM and agent evaluation: no code-repair benchmark targets distributed-system bugs, and no controlled study isolates how much externally provided debugging context changes agent success on them. We introduce DDBench, a code-repair benchmark of 60 historical bugs mined from 13 open-source distributed systems, partitioned into three difficulty tiers. DDBench evaluates every case under two matched conditions: a symptom-only condition where the agent receives only the bug symptom and repository, and a context-augmented condition where it additionally receives a bounded debugging context (logs, traces, runtime state, and targeted code-investigation notes), isolating the effect of debugging context from model capability. The evaluation of ten LLMs on DDBench reveals several findings. First, distributed debugging exercises a reasoning dimension that single-process benchmarks do not surface: models' pass rates span 61 pp, and pairwise bootstrap separates 9 of 15 top-tier model pairs at p < 0.05 on DDBench's hardest case-set. Second, bounded debugging context lifts aggregate pass rate by +18.1 pp, and the lift is asymmetric: weaker models gain pass rate, while stronger models gain efficiency. Third, debugging context requires careful curation, as even faithful debugging context can sometimes mislead LLMs.

摘要:LLM 基礎的編碼代理在單一過程的軟體工程任務上迅速進步,前沿模型在 SWE-bench Verified 上的表現已聚集在高 70 分以上。 然而,分散式系統的除錯仍然是一個未被充分探索的領域:錯誤跨越過程、節點和協議互動,根本原因很少能僅從源代碼中恢復,且在非確定性交錯中進行暴力探索是不可行的。 這在 LLM 和代理評估中留下了兩個空白:沒有代碼修復基準針對分散式系統的錯誤,且沒有控制研究來隔離外部提供的除錯上下文如何改變代理在這些錯誤上的成功率。 我們介紹 DDBench,一個由 13 個開源分散式系統挖掘的 60 個歷史錯誤組成的代碼修復基準,分為三個難度層級。
DDBench 在兩個匹配條件下評估每個案例:一個僅有症狀的條件,代理僅接收錯誤症狀和代碼庫,另一個是增強上下文的條件,代理還接收有限的除錯上下文(日誌、追蹤、運行時狀態和針對代碼調查的筆記),以隔離除錯上下文對模型能力的影響。 對十個 LLM 在 DDBench 上的評估揭示了幾個發現。 首先,分散式除錯運用了一個單一過程基準未顯現的推理維度:模型的通過率跨度為 61 個百分點,配對自助法在 DDBench 最困難的案例集中以 p < 0.05 分離了 15 個頂級模型對中的 9 個。 其次,有限的除錯上下文將總體通過率提升了 +18.1 個百分點,且這一提升是非對稱的:較弱的模型通過率提高,而較強的模型則提高了效率。 第三,除錯上下文需要仔細策劃,因為即使是忠實的除錯上下文有時也會誤導 LLM。

Generated Context versus Governed State: Functional Conditions for Accountable Longitudinal Clinical Reasoning

2608.14804v1 by Augusto Bernardo Pissarra, Victor Lorena de Farias Souza

Large language models (LLMs) have become the dominant interface of clinical artificial intelligence, yet the interface they expose (text in, text out, one context window at a time) maintains no explicit, persistent, governed representation of what is currently true about a patient. This paper argues that longitudinal clinical reasoning is a state-estimation problem under partial observability, and that the axis on which clinical AI succeeds or fails is not the fluency of the model reading the record but the governance of the patient state it reasons over. We distinguish generated context from governed state; separate five objects that clinical AI habitually conflates (true state, observations, evidence, belief, and simulated state); define a tiered governance standard against which any clinical AI system can be audited; and show that an operational definition of accountability decomposes into four information requirements: an immutable evidence ledger with awareness-time versioning, a belief state distinct from accumulated evidence, an observation-process model, and claim-level causal typing. We are explicit that this decomposition is analytic rather than a necessity theorem, and that its value is conceptual hygiene: it converts "accountable clinical AI" from a slogan into an audit instrument. A six-level maturity framework separates what a system makes governable from what it can compute, locating current LLM-centric practice at high capability but low maturity. The paper is fully self-contained: the four research questions the framework poses are stated in the introduction, and the conclusion records what the paper establishes toward each; future work develops the buildable core of the architecture and the research program toward full Clinical World Models. No empirical result is claimed here.

摘要:大型語言模型(LLMs)已成為臨床人工智慧的主導介面,但它們所暴露的介面(文本輸入、文本輸出、一次一個上下文窗口)並未對目前有關患者的真實情況提供明確、持久、受管控的表徵。本文主張,縱向臨床推理是一個在部分可觀察性下的狀態估計問題,而臨床 AI 成功或失敗的軸心不在於模型閱讀記錄的流暢性,而在於它所推理的患者狀態的治理。我們區分生成的上下文與受管控的狀態;將臨床 AI 通常混淆的五個對象(真實狀態、觀察、證據、信念和模擬狀態)分開;定義一個分層治理標準,以便對任何臨床 AI 系統進行審計;並顯示一個運作性責任的定義可分解為四個信息要求:具有意識時間版本控制的不可變證據賬本、與累積證據不同的信念狀態、觀察過程模型,以及索賠級別的因果類型。我們明確指出這一分解是分析性的,而非必要定理,其價值在於概念衛生:它將“可負責任的臨床 AI”從口號轉變為審計工具。一個六級成熟度框架將系統可治理的部分與其可計算的部分分開,將當前以 LLM 為中心的實踐定位於高能力但低成熟度。本文是完全自足的:框架提出的四個研究問題在引言中陳述,結論記錄了本文在每個問題上所建立的內容;未來的工作將發展可構建的架構核心及通向完整臨床世界模型的研究計劃。此處不聲稱任何實證結果。

Prompting is not enough: supervised baselines and leakage control for measuring shared decision-making with LLMs in pediatric encounters

2608.14792v1 by Bernardo Modenesi, Jody Lin, Kimberly Kaphingst, Angela Zhu, Maya Wheeler, Peilu Zhang, Angela Fagerlin

Objectives: To determine whether zero-shot prompting of a large language model (LLM) is sufficient to detect shared decision-making (SDM) behaviors in real clinical encounters, and whether supervised learning adds value under patient-grouped, nested evaluation. Methods: We analyzed 21 audio-recorded outpatient surgical decision encounters (19 unique patients; 7,566 utterance segments; ~6.1 hours) between families of children with multiple long-term conditions and their surgical providers. Trained coders labeled segments for 12 SDM behaviors (human-human macro Cohen's kappa = 0.695). We compared a zero-shot local LLM (Qwen 2.5 32B), a supervised classifier over frozen sentence embeddings, and their logistic stack, under patient-grouped outer folds with inner cross-fitted thresholds and patient-resampled confidence intervals. Results: The zero-shot LLM reached macro kappa = 0.139 (95% CI 0.111-0.164). The supervised classifier reached kappa = 0.227 (0.186-0.262), a paired improvement of 0.088 (0.051-0.119). A logistic stack of the two reached kappa = 0.242 (0.198-0.284). We identified multiple corpus-specific leakage paths, including grouping sibling recordings separately and allowing labels from an outer held-out patient to enter few-shot exemplars used while fitting downstream models. Conclusion: Zero-shot prompting alone is not sufficient to measure SDM behavior as reliably as a small supervised model, and patient-level grouping alone does not prevent leakage when labeled prompt exemplars are precomputed outside the outer evaluation loop. Reported performance is sensitive to the unit of data splitting and to where labeled exemplars enter the pipeline. External validation is needed before these findings generalize beyond this population, model, prompt, and codebook.

摘要:目標:確定大型語言模型(LLM)的零-shot 提示是否足以在真實臨床接觸中檢測共享決策(SDM)行為,以及監督學習在患者分組的嵌套評估下是否增值。
方法:我們分析了21段音頻錄製的門診外科決策接觸(19名獨特患者;7,566個發言片段;約6.1小時),這些接觸發生在多種長期疾病兒童的家庭與其外科提供者之間。受過訓練的編碼員為12種SDM行為標記片段(人與人之間的宏觀Cohen's kappa = 0.695)。我們比較了一個零-shot本地LLM(Qwen 2.5 32B)、一個基於凍結句子嵌入的監督分類器,以及它們的邏輯回歸堆疊,在患者分組的外部折疊下,使用內部交叉擬合的閾值和患者重抽樣的置信區間。
結果:零-shot LLM達到宏觀kappa = 0.139(95% CI 0.111-0.164)。監督分類器達到kappa = 0.227(0.186-0.262),配對改善為0.088(0.051-0.119)。兩者的邏輯回歸堆疊達到kappa = 0.242(0.198-0.284)。我們識別了多條特定語料的洩漏路徑,包括將兄弟姐妹的錄音分開分組,以及允許來自外部保留患者的標籤進入在擬合下游模型時使用的少量示例。
結論:僅依賴零-shot 提示不足以像小型監督模型那樣可靠地測量SDM行為,且僅進行患者層級分組並不能防止當標記的提示示例在外部評估循環之外預先計算時的洩漏。報告的性能對數據拆分的單位和標記示例進入管道的位置敏感。在這些發現能夠超出這一人群、模型、提示和代碼本進行推廣之前,需要進行外部驗證。

CEDAR-GRPO: Process-Aware Reinforcement Learning for General Abductive Reasoning in LLMs

2608.14791v1 by Moein Salimi, Danial Parnian, Shaygan Adim, Amirmohammad Ebrahiminasab, Nima Alighardashi, Parsa Gholami, Sahand Akramipour, Mahdi Jafari Siavoshani, Mohammad Hossein Rohban

Abductive reasoning, often characterized as inference to the best explanation, is central to explanation under uncertainty, from everyday sense-making and investigation to scientific discovery. Yet LLM research has mostly studied abduction through narrow, task-specific benchmarks, making it unclear whether observed gains transfer beyond the benchmark family used for training or evaluation. We ask whether RL post-training can improve abduction as a transferable reasoning capability. We introduce CEDAR-GRPO, a process-aware framework that combines final-answer correctness with abductive rewards for evidence coverage and evidence-to-explanation directionality. Four open-weight LLMs are post-trained on a controlled, domain-neutral mixture of abductive hypothesis-generation and hypothesis-selection tasks. We evaluate them on 11 unseen tasks spanning hypothesis selection, missing-fact generation, defeasible inference, long-context investigation, clinical reasoning, code debugging, and non-abductive controls. CEDAR- GRPO improves every model on every held-out task over both base models and correctness-only GRPO, with average gains of 7.4 and 2.7 points, respectively, and a maximum gain of 30.8 points. Ablations confirm that RL, abductive reward design, and task diversity each contribute to transfer. Process-level metrics further show stronger abductive behavior, including exploration of alternatives, elimination of rivals, backtracking, and uncertainty marking.

摘要:誘導推理,通常被描述為最佳解釋的推斷,是在不確定性下解釋的核心,從日常的意義建構和調查到科學發現。然而,LLM 研究大多通過狹窄的、特定任務的基準來研究誘導,這使得觀察到的增益是否能轉移到用於訓練或評估的基準家族之外變得不清楚。我們詢問 RL 後訓練是否能改善誘導作為可轉移的推理能力。我們介紹 CEDAR-GRPO,一個過程感知框架,將最終答案的正確性與誘導獎勵結合,考慮證據覆蓋率和證據到解釋的方向性。四個開放權重的 LLM 在一個受控的、領域中立的誘導假設生成和假設選擇任務的混合上進行後訓練。我們在 11 個未見過的任務上評估它們,這些任務涵蓋假設選擇、缺失事實生成、可駁斥推理、長上下文調查、臨床推理、代碼調試和非誘導控制。CEDAR-GRPO 在每個保留任務上改善了每個模型,無論是基礎模型還是僅考慮正確性的 GRPO,平均增益分別為 7.4 和 2.7 分,最大增益為 30.8 分。消融實驗確認 RL、誘導獎勵設計和任務多樣性各自對轉移有貢獻。過程級別的指標進一步顯示出更強的誘導行為,包括探索替代方案、消除競爭者、回溯和不確定性標記。

Seeing Red, Thinking Bad: Color Bias in Vision Language Models

2608.14286v1 by Kohsuke Ide, Ryousuke Yamada, Yoshihiro Fukuhara, Hirokatsu Kataoka, Yutaka Satoh

Vision language models (VLMs) are increasingly used in industrial decision-making systems, such as recruitment support and recommendation. This motivates careful analysis of how VLMs process visual and textual information. In this work, we study how VLMs interpret text rendered as an image, and investigate the influence of visual styling biases. To this end, we introduce Stealth Visual Prompts, which subtly change visual styling of text, such as color and contrast, while preserving semantic content. Using these prompts, we systematically control the visual styling of words in text and measure their impact on the analysis performed by VLMs. We further analyze how such visual perturbations affect the latent representations of the vision encoder. From our experiments, we observed that coloring positive words in green consistently shifts sentiment predictions toward a positive direction. As a result, VLMs often fail to properly account for negative words present in the text. Our analysis suggests that this behavior is correlated with changes in the latent representations of the vision encoder induced by color variations. In addition, we show that reducing text--background contrast increases reliance on visually salient cues and leads to more incorrect Visual Question Answering (VQA) outputs. These results suggest that the visual styling of rendered text can guide VLMs' interpretation in ways that diverge from human semantic understanding. Project page: https://github.com/KohsukeIde/color-bias-vlm

摘要:視覺語言模型(VLMs)在工業決策系統中越來越多地被使用,例如招聘支持和推薦。這促使我們仔細分析 VLMs 如何處理視覺和文本信息。在這項工作中,我們研究 VLMs 如何解釋以圖像呈現的文本,並調查視覺風格偏見的影響。為此,我們引入了隱形視覺提示,這些提示微妙地改變文本的視覺風格,例如顏色和對比度,同時保留語義內容。利用這些提示,我們系統地控制文本中單詞的視覺風格,並測量其對 VLMs 執行的分析的影響。我們進一步分析這些視覺擾動如何影響視覺編碼器的潛在表示。從我們的實驗中,我們觀察到將正面詞語著色為綠色會持續地將情感預測向正面方向偏移。因此,VLMs 經常無法正確考慮文本中存在的負面詞語。我們的分析表明,這種行為與由顏色變化引起的視覺編碼器潛在表示的變化相關。此外,我們顯示減少文本與背景的對比度會增加對視覺顯著線索的依賴,並導致更多不正確的視覺問題回答(VQA)輸出。這些結果表明,渲染文本的視覺風格可以以偏離人類語義理解的方式引導 VLMs 的解釋。
項目頁面:https://github.com/KohsukeIde/color-bias-vlm

Meteorology-driven Causal Nowcasting of Fugitive Landfill Emissions Enables Proactive Public Health Response

2608.14254v1 by Timothy C. Pearce, David J. T. Smith, Alec Dobney, Alessia Freddo

Fugitive emissions from waste sites increasingly expose communities to toxic and odorous gases, yet public-health responses remain largely retrospective, with episodes investigated only after residents have been exposed. Here we show that the meteorological drivers of elevated hydrogen sulphide (HS) at a long-monitored European landfill, and the timescales over which they act, can be identified directly from routine monitoring data. We introduce CAIRN (Causal-Anchored Inference for Receptor Nowcasting), a machine-learning framework whose internal memory is matched to these measured timescales: a fast component tracking hour-scale wind-borne transport and a slow component tracking multi-hour weather changes. Trained to predict gas measurements, CAIRN operates using only routine weather variables and the calendar, without hand-engineered features. Its behaviour is consistent with the identified transport mechanisms, and the framework transfers unchanged to a second monitoring station and to co-emitted methane. Combining four such nowcasters produces a site-level, tiered alert aligned with WHO odour guidance that closely reproduces the alert generated by a direct sensor network and tracks an independent record of community odour complaints. Weather-driven nowcasting can therefore estimate community impact as an emission episode unfolds, providing public-health authorities with a validated, graded trigger for intervention and enabling exposure to be reduced during events rather than after them.

摘要:逃逸排放的廢棄物場越來越多地使社區暴露於有毒和有氣味的氣體中,但公共衛生的反應仍然主要是事後的,只有在居民受到影響後才進行調查。在這裡,我們展示了在一個長期監測的歐洲垃圾填埋場中,氫硫化物(HS)濃度升高的氣象驅動因素及其作用的時間尺度,可以直接從常規監測數據中識別出來。我們引入了CAIRN(因果錨定推斷接收器即時預測),這是一個機器學習框架,其內部記憶與這些測量的時間尺度相匹配:一個快速組件跟踪小時級的風載運輸,另一個慢速組件跟踪多小時的天氣變化。CAIRN經過訓練以預測氣體測量,僅使用常規的天氣變數和日曆,而不需要手工設計的特徵。其行為與識別出的運輸機制一致,並且該框架可以不變地轉移到第二個監測站和共同排放的甲烷。結合四個這樣的即時預測器,產生了一個與世界衛生組織氣味指導相一致的現場級分層警報,該警報與直接傳感器網絡生成的警報非常接近,並跟踪社區氣味投訴的獨立記錄。因此,天氣驅動的即時預測可以在排放事件展開時估計社區影響,為公共衛生當局提供經過驗證的分級干預觸發器,並使在事件期間減少暴露成為可能,而不是在事件之後。

Agentic Data Cleaning Without a Clean Reference: An Experimental Study of Capabilities and Trade-offs

2608.14765v1 by Hadi Fadlallah

Data cleaning without a trusted clean reference is challenging because unusual values may represent either genuine errors or valid observations. This paper studies how different agent capabilities affect reference-free data cleaning and proposes an evidence-grounded framework that combines structured context, profiling, LLM reasoning, executable checks, controlled evidence retrieval, source ranking, citation alignment, conservative repair, reversible scripts, and provenance logging. Seven configurations are evaluated across financial, clinical, and environmental-monitoring datasets using controlled synthetic corruption and original-data descriptive analysis, resulting in 126 completed runs. The evaluation includes two comparison baselines and a progressive LLM-based sequence that adds executable tools, evidence retrieval, evidence controls, and conservative repair. In the synthetic evaluation, the deterministic profiling baseline achieved the highest detection F1-score of 0.561. Among the LLM-based configurations, the full conservative configuration achieved the highest F1-score of 0.421, but no configuration performed best across all evaluation criteria. The source-ranked configurations achieved the lowest unsupported-rule rates, while decision-level citation alignment remained weak. The full conservative configuration produced no unsafe or unnecessary modifications, although these rates were already zero before the conservative policy was added, and it performed no direct repairs. Overall, the results show that additional capabilities introduce trade-offs among detection, repair, evidence grounding, conservative behaviour, reproducibility, and operational cost rather than producing consistent improvements. The study provides a structured framework and empirical methodology for evaluating these trade-offs in reference-free agentic data cleaning.

摘要:數據清理在沒有可信的乾淨參考的情況下是具有挑戰性的,因為異常值可能代表真正的錯誤或有效的觀察結果。本文研究了不同代理能力如何影響無參考數據清理,並提出了一個基於證據的框架,該框架結合了結構化上下文、檔案分析、LLM 推理、可執行檢查、受控證據檢索、來源排名、引用對齊、保守修復、可逆腳本和來源日誌。通過使用受控的合成腐蝕和原始數據描述性分析,對七種配置進行了評估,涵蓋了金融、臨床和環境監測數據集,最終完成了126次運行。評估包括兩個比較基準和一個逐步的基於 LLM 的序列,該序列添加了可執行工具、證據檢索、證據控制和保守修復。在合成評估中,確定性檔案分析基準達到了最高的檢測 F1 分數 0.561。在基於 LLM 的配置中,完整的保守配置達到了最高的 F1 分數 0.421,但沒有任何配置在所有評估標準中表現最佳。來源排名配置達到了最低的不支持規則率,而決策級引用對齊仍然較弱。完整的保守配置未產生任何不安全或不必要的修改,儘管在添加保守政策之前這些比率已經為零,並且它沒有進行直接修復。總體而言,結果顯示,額外的能力在檢測、修復、證據基礎、保守行為、可重複性和運營成本之間引入了權衡,而不是產生一致的改進。該研究提供了一個結構化框架和實證方法,用於評估這些在無參考代理數據清理中的權衡。

APTER: Adaptive Post-Training with Expert-Grounded Rubrics

2608.14212v1 by Xukai Wang, Liangqi Li, Zhiyue Xu, Jingang Zhou, Xiaoyu Shi, Jiansheng Cai, Bo Zhang, Zhe Li, Xu-Yao Zhang

As large language models enter professional domains, they must satisfy domain constraints, include critical evidence, and provide complete reasoning rather than merely produce fluent responses. Existing post-training methods often rely on holistic preferences or outcome-level verification, while recent rubric-based methods usually generate rubrics independently for each query. In specialized domains, such unconstrained rubrics may omit critical requirements and vary across samples, hindering the diagnosis and targeted repair of persistent capability deficiencies. We propose APTER (Adaptive Post-Training with Expert-Grounded Rubrics), a framework that integrates structured domain knowledge into fine-grained evaluation, optimization, and diagnosis for specialized complex reasoning. First, expert-grounded rubric construction starts from an expert criteria framework built by domain experts, where each criterion represents a stable professional capability. For each query, APTER selects relevant criteria and instantiates them into query-level rubrics linked to their source criteria, turning reusable expert criteria into executable query-level supervision without reference answers. Second, adaptive post-training uses rubric verdicts as both optimization and criterion-level diagnostic signals. Aggregating low-scoring verdicts by criterion ID reveals persistent deficiencies and triggers targeted supervised fine-tuning updates during reinforcement learning. Experiments on mathematical reasoning and medical question answering show consistent gains across both domains. Across three model generations, APTER improves the mathematics and medical averages over the corresponding base models by up to 15.86 and 8.04 points, respectively. Code and rubric datasets are available at https://github.com/AntDT-APTER/APTER.

摘要:隨著大型語言模型進入專業領域,它們必須滿足領域約束,包含關鍵證據,並提供完整的推理,而不僅僅是產生流暢的回應。現有的後訓練方法通常依賴於整體偏好或結果層級的驗證,而最近的基於評分標準的方法通常為每個查詢獨立生成評分標準。在專業領域中,這種不受約束的評分標準可能會省略關鍵要求,並在樣本之間變化,妨礙持續能力缺陷的診斷和針對性修復。我們提出了APTER(基於專家的自適應後訓練評分標準),這是一個將結構化領域知識整合到細緻評估、優化和診斷中的框架,旨在解決專業複雜推理的問題。首先,專家基礎的評分標準構建始於由領域專家建立的專家標準框架,其中每個標準代表一種穩定的專業能力。對於每個查詢,APTER選擇相關標準並將其實例化為與其來源標準相關聯的查詢層級評分標準,將可重用的專家標準轉化為可執行的查詢層級監督,而不需要參考答案。其次,自適應後訓練使用評分標準的判決作為優化和標準層級診斷信號。通過標準ID聚合低分判決可以揭示持續的缺陷,並在強化學習過程中觸發針對性的有監督微調更新。在數學推理和醫學問題回答的實驗中,兩個領域均顯示出一致的增長。在三代模型中,APTER在數學和醫學的平均分數上分別提高了高達15.86和8.04分。代碼和評分標準數據集可在 https://github.com/AntDT-APTER/APTER 獲得。

Removing Temporal Note Redundancy Improves Multimodal Reinforcement Learning for Medicine

2608.14157v1 by Chenran Weng, Joo Seung Lee, Malini Mahendra, Anil Aswani

Mechanical ventilation is a critical life-support intervention, requiring dynamic adjustments to ventilator settings as a patient's condition evolves. While reinforcement learning (RL) offers a promising framework for optimizing these sequential decisions, standard approaches rely primarily on structured electronic health record (EHR) data, missing crucial clinical context recorded in free-text notes. Integrating longitudinal clinical notes into RL state spaces is challenging because notes are heavily inflated by temporal redundancy, such as copy-forward text, templating, and repetitive documentation, which dilutes time-local updates and degrades state representation quality. To address this, we propose a redundancy-aware multimodal state representation framework that explicitly removes duplicated note text over time before policy learning. We evaluate two computationally efficient temporal decomposition strategies for removing duplicated note text: (1) an embedding-space decomposition using singular value decomposition on local history subspaces, and (2) an interpretable sentence-level diff operation that filters out previously documented sentences before text encoding. Using real-world ICU data, we demonstrate that state representations constructed by stripping temporal note redundancy significantly outperform both structured-only and raw-note baselines across multiple off-policy evaluation methods (Model-Based Rollouts, Fitted Q-Evaluation, Weighted Importance Sampling, and Weighted Doubly Robust Evaluation). Our findings show that explicitly isolating new clinical information from repeated note text yields higher-quality state representations and directly improves RL performance for clinical decision support.

摘要:機械通氣是一項關鍵的生命支持干預,隨著病人狀況的變化,需要對通氣器設置進行動態調整。雖然強化學習(RL)提供了一個有前景的框架來優化這些序列決策,但標準方法主要依賴於結構化的電子健康記錄(EHR)數據,忽略了在自由文本註解中記錄的重要臨床背景。將長期臨床註解整合進RL狀態空間是具有挑戰性的,因為註解受到時間冗餘的嚴重影響,例如複製轉發文本、模板化和重複文檔,這稀釋了時間局部更新並降低了狀態表示的質量。為了解決這個問題,我們提出了一個冗餘感知的多模態狀態表示框架,該框架在策略學習之前明確去除隨時間重複的註解文本。我們評估了兩種計算效率高的時間分解策略來去除重複的註解文本:(1)使用奇異值分解對局部歷史子空間進行的嵌入空間分解,以及(2)一種可解釋的句子級差異操作,在文本編碼之前過濾掉先前記錄的句子。使用真實世界的ICU數據,我們展示了通過剝離時間註解冗餘構建的狀態表示在多種離線政策評估方法(基於模型的回滾、擬合Q評估、加權重要性抽樣和加權雙重穩健評估)中顯著優於僅結構化和原始註解的基準。我們的研究結果顯示,明確將新的臨床信息與重複的註解文本隔離,可以產生更高質量的狀態表示,並直接改善臨床決策支持的RL性能。

CMCNet: Aligning Ultrasound Image Embeddings with Textual TI-RADS Representations for Fine-Grained Thyroid Classification

2608.13939v1 by Bingxin Yu, Xueli Wang, Jerry Zhou, Wenyan Wang, Li Wen, Lan Huang, Xin Feng, Fengfeng Zhou, Kewei Li

Ultrasound is the primary imaging modality for assessing thyroid nodules, and the ACR TI-RADS framework standardizes diagnosis through five ultrasound feature categories that are aggregated into five risk levels (TR1-TR5). Although widely adopted in clinical practice, most deep learning approaches focus on binary malignancy classification, while multi-class prediction and explicit utilization of feature-level supervision remain underexplored, largely due to limited annotated data. In this study, we introduce the STN dataset of 600 thyroid nodules with paired transverse and longitudinal ultrasound images, bounding box annotations, and complete labels for all five TI-RADS feature categories. Following the clinical decision process, we investigate how structured feature information can guide representation learning during training while requiring only images at inference. We demonstrate that text embeddings derived from standardized feature descriptions form a stable surrogate representation for TI-RADS risk levels. Based on this observation, we propose CMCNet, which aligns image embeddings to fixed textual embeddings via a Center-Margin Contrastive Loss that simultaneously promotes intra-class compactness and inter-class separation. Experimental results show that this embedding alignment strategy is more data-efficient and robust than direct multitask learning, and consistently outperforms InfoNCE, center loss, a strong multitask baseline, and a VQA-style multimodal model, particularly in imbalanced settings. The dataset is freely available at doi: 10.5281/zenodo.19125693 and the source code is available at: https://www.healthinformaticslab.org/supp/.

摘要:超聲波是評估甲狀腺結節的主要影像學方法,而 ACR TI-RADS 框架通過五個超聲特徵類別標準化診斷,這些特徵被聚合成五個風險等級(TR1-TR5)。儘管在臨床實踐中被廣泛採用,但大多數深度學習方法專注於二元惡性分類,而多類別預測和明確利用特徵級監督的研究仍然未被充分探索,這主要是由於標註數據的限制。在本研究中,我們引入了 STN 數據集,其中包含 600 個甲狀腺結節的配對橫向和縱向超聲圖像、邊界框註釋以及所有五個 TI-RADS 特徵類別的完整標籤。根據臨床決策過程,我們探討結構化特徵信息如何在訓練期間指導表示學習,同時在推理時僅需圖像。我們證明來自標準化特徵描述的文本嵌入形成了 TI-RADS 風險等級的穩定替代表示。基於這一觀察,我們提出了 CMCNet,該模型通過中心-邊距對比損失將圖像嵌入與固定文本嵌入對齊,這同時促進了類內緊湊性和類間分離性。實驗結果顯示,這種嵌入對齊策略比直接的多任務學習更具數據效率和穩健性,並且在不平衡設置中始終優於 InfoNCE、中心損失、一個強大的多任務基線以及一個 VQA 風格的多模態模型。該數據集可免費獲得,DOI 為:10.5281/zenodo.19125693,源代碼可在:https://www.healthinformaticslab.org/supp/ 獲得。

Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions

2608.13786v1 by Qingfang Liu, Qiao Jin, Joe D. Menke, Thorsten Kahnt, Zhiyong Lu

Large language model (LLM) chatbots are increasingly used to answer clinical questions with citations to relevant clinical studies. Prior research has largely focused on citation fabrication, leaving a gap in evaluating the quality of retrieved studies and the factors driving their selection. In this study, we evaluated three general-purpose LLM chatbots: Claude Sonnet 5, Gemini 3.1 Pro, and ChatGPT GPT-5.5. We prompted the models with clinical questions adapted from 20 review questions in Issues 6 and 7 of the 2026 Cochrane Database of Systematic Reviews, simulating patient, clinician, and evidence-synthesis researcher roles. Each chatbot was queried under each user role with four independent repetitions, yielding 720 responses. Each chatbot was asked to support its answers with primary clinical citations, which we benchmarked against the included and excluded study sets of the Cochrane reviews. On average, a chatbot response retrieved 39.2% $\pm$ 29.8% of Cochrane included studies, while citing 5.0% $\pm$ 9.4% of excluded studies. Recall of Cochrane included studies varied significantly by model and user role. ChatGPT achieved higher recall than Claude or Gemini (63.1% $\pm$ 29.5% vs. 37.0% $\pm$ 23.8% vs. 17.3% $\pm$ 13.1%; $p=2.0\times10^{-5}$). The researcher role yielded higher recall than the clinician or patient roles (42.8% $\pm$ 30.8% vs. 38.6% $\pm$ 28.9% vs. 36.1% $\pm$ 29.3%; $p=2.0\times10^{-5}$). Controlling for publication year, citations per year, and open-access status, sample size was the only independently significant predictor of retrieval (odds ratio 1.80 per 1-unit increase in log sample size, 95% CI 1.37-2.36, $p=2.34\times10^{-5}$). These findings suggest that while LLM chatbots can retrieve some studies identified by expert reviewers, their performance varies by model and user role, and they exhibit a bias toward clinical trials with larger sample sizes.

摘要:大型語言模型(LLM)聊天機器人越來越多地用於回答臨床問題,並引用相關的臨床研究。先前的研究主要集中在引用虛構上,未能評估檢索到的研究質量及其選擇的驅動因素。在本研究中,我們評估了三個通用型LLM聊天機器人:Claude Sonnet 5、Gemini 3.1 Pro和ChatGPT GPT-5.5。我們根據2026年Cochrane系統評價數據庫第6和第7期的20個回顧問題,為模型提供了臨床問題的提示,模擬患者、臨床醫生和證據綜合研究者的角色。每個聊天機器人在每個用戶角色下進行了四次獨立查詢,共產生720個回應。每個聊天機器人被要求用主要臨床引用來支持其答案,我們將其與Cochrane評估的納入和排除研究集進行了基準比較。平均而言,聊天機器人的回應檢索了39.2% $\pm$ 29.8%的Cochrane納入研究,同時引用了5.0% $\pm$ 9.4%的排除研究。Cochrane納入研究的回憶率因模型和用戶角色而異。ChatGPT的回憶率高於Claude或Gemini(63.1% $\pm$ 29.5% vs. 37.0% $\pm$ 23.8% vs. 17.3% $\pm$ 13.1%;$p=2.0\times10^{-5}$)。研究者角色的回憶率高於臨床醫生或患者角色(42.8% $\pm$ 30.8% vs. 38.6% $\pm$ 28.9% vs. 36.1% $\pm$ 29.3%;$p=2.0\times10^{-5}$)。在控制出版年份、每年引用數和開放獲取狀態後,樣本大小是唯一獨立顯著的檢索預測因子(對數樣本大小每增加1單位的比值比1.80,95% CI 1.37-2.36,$p=2.34\times10^{-5}$)。這些發現表明,儘管LLM聊天機器人可以檢索到一些專家評審者識別的研究,但其性能因模型和用戶角色而異,並且對樣本大小較大的臨床試驗存在偏見。

Data-driven techniques for translational neuroscience and personalized neuro-health

2608.13749v1 by Vishal Subedi, Shashipraba N. K. Rajakaruna, Pratyusha Sarkar, Subhankar Chattoraj, Anjali Khasa, Siddhartha Nandy, Hamza Farooq, Animikh Biswas, Sanjay Chaudhuri, Asim K. Dey, Karuna Joshi, Christophe Lenglet, Ansu Chatterjee

Neurodegenexrative diseases such as Alzheimer's disease and Parkinson's disease are diagnosed most reliably only after substantial, often irreversible, neuronal loss has already occurred, creating an urgent need for quantitative tools that can detect subtle, early, and individual-specific brain changes from neuroimaging data. This review surveys a broad and rapidly evolving toolkit of data-driven techniques for translational neuroscience and personalized neuro-health, organized around four complementary methodological pillars. Throughout, we emphasize how these methodologically diverse approaches converge on a common translational goal: personalized, mechanistically grounded, and clinically actionable models of individual brain health, and we close by discussing the principal open statistical, computational, and clinical challenges that remain.

摘要:神經退行性疾病,如阿茲海默症和帕金森病,通常只有在已經發生了實質性且通常是不可逆的神經元損失後,才能最可靠地診斷,這造成了對能夠從神經影像數據中檢測微妙、早期且個體特異性腦部變化的定量工具的迫切需求。這篇綜述調查了一套廣泛且快速發展的數據驅動技術工具,旨在轉化神經科學和個性化神經健康,並圍繞四個互補的方法論支柱進行組織。整篇文章強調這些方法論多樣的途徑如何匯聚到一個共同的轉化目標:個性化、機制基礎的且臨床可行的個體腦健康模型,並在結尾討論仍然存在的主要統計、計算和臨床挑戰。

MedPlex: Deep Vision-Language Co-Adaptation for Clinically Grounded Medical Segmentation

2608.13690v1 by Rafi Ibn Sultan, Hui Zhu, Chengyin Li, Dongxiao Zhu

Medical image segmentation is still largely treated as a vision-only problem, although clinical interpretation often relies on textual knowledge of anatomy, location, appearance, and surrounding context. Existing text-guided segmentation methods within the Vision-Language Model (VLM) paradigm often use language only as a late conditioning signal, limiting its influence on visual representation learning. We introduce MedPlex (Medical Plexus of Vision and Language), an end-to-end VLM framework that makes text guidance a continuous, clinically grounded component of segmentation learning. Through Bi-Fusion (Bidirectional Fusion), visual and textual representations evolve jointly across the encoding hierarchy. MedPlex further introduces class-level and region-level concept alignment to organize the shared representation at complementary granularities. Class-level alignment anchors each anatomical target to an aggregated clinical concept profile, while region-level alignment preserves individual concepts, such as shape, location, appearance, and texture, through class-specific visual evidence. In this way, language provides structured supervision throughout the encoder rather than serving only as a late-stage cue. MedPlex achieves state-of-the-art performance across CT and MR benchmarks for multi-organ, cardiac substructure, and tumor segmentation, including settings with real free-text clinical supervision. Code: https://github.com/rafiibnsultan/MedPlex.

摘要:醫學影像分割仍然主要被視為一個僅限於視覺的問題,儘管臨床解釋通常依賴於對解剖學、位置、外觀和周圍背景的文本知識。現有的文本引導分割方法在視覺-語言模型(VLM)範式內,通常僅將語言用作後期條件信號,限制了其對視覺表示學習的影響。我們介紹了 MedPlex(醫學視覺與語言的聯結),這是一個端到端的 VLM 框架,使文本引導成為分割學習中的一個持續且臨床基礎的組件。通過雙向融合(Bi-Fusion),視覺和文本表示在編碼層次中共同演變。MedPlex 進一步引入了類別級和區域級概念對齊,以在互補的粒度上組織共享表示。類別級對齊將每個解剖目標錨定到一個聚合的臨床概念檔案,而區域級對齊則通過類別特定的視覺證據保留個別概念,例如形狀、位置、外觀和質地。這樣,語言在編碼器中提供結構化的監督,而不僅僅是在後期階段作為提示。MedPlex 在多器官、心臟子結構和腫瘤分割的 CT 和 MR 基準測試中達到了最先進的性能,包括具有實際自由文本臨床監督的設置。代碼: https://github.com/rafiibnsultan/MedPlex。

MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and Coordination

2608.13476v1 by Saisha Shetty, Satvik Tripathi, Austin Lin, Colin Zhao, Theodore Kim, Don Enwerem, Jacinta Arnold, Shahriar Faghani, Tessa S Cook

We present Multi-Agent Reasoning and Coordination (MARC), an open-source framework that replaces monolithic LLM prompting with deterministic multi-agent orchestration for clinical reasoning. MARC coordinates role-specialized agents for extraction, reasoning, answer generation, and evaluation, with explicit context passing and traceable intermediate outputs, enabling stage-wise failure attribution. We additionally introduce a Decomposer module that generates task-specific agent prompts from a plain-language description, eliminating manual prompt engineering. The framework supports both API-based and local CPU-compatible deployments and is entirely configurable via YAML, without code modifications. MARC is designed to be model-agnostic, interpretable, and accessible to clinical domain experts without programming expertise. The full framework is available at https://github.com/Penn-RAIL/MARC-v1.

摘要:我們提出了多智能體推理與協調(MARC),這是一個開源框架,將單一大型語言模型的提示替換為確定性的多智能體協調,用於臨床推理。MARC 協調角色專門的智能體進行提取、推理、答案生成和評估,具有明確的上下文傳遞和可追溯的中間輸出,實現階段性失敗歸因。我們還引入了一個分解器模塊,該模塊從普通語言描述中生成任務特定的智能體提示,消除了手動提示工程。該框架支持基於 API 的和本地 CPU 兼容的部署,並且完全可以通過 YAML 配置,而無需修改代碼。MARC 設計為模型無關、可解釋,並且對於沒有編程專業知識的臨床領域專家可訪問。完整框架可在 https://github.com/Penn-RAIL/MARC-v1 獲得。

Towards Context-Aware Clinical Motion Understanding in Daily Living at Home: Freezing of Gait Detection with Egocentric Vision

2608.13283v1 by Vayalet Stefanova, Diwas Lamsal, Margot Genbrugge, Maxim Yudayev, Christian Schlenstedt, Moran Gilat, Bart Vanrumste, Benjamin Filtjens

Understanding motion in daily living requires context beyond kinematics, because similar inertial patterns during activities of daily living (ADLs) can reflect intentional stopping, object interaction, or pathological movement impairment. Egocentric vision provides task-related context that may help disambiguate these cases. We investigate this challenge through freezing of gait (FOG) detection in Parkinson's disease (PD), a symptom strongly influenced by contextual factors during ADLs. Using synchronized egocentric video, wearable IMUs, and expert-annotated FOG labels collected from 13 PD participants in their homes, we evaluate frozen representations from pretrained ego-video and time-series foundation models, alongside an IMU-based TCN trained from scratch, under leave-one-subject-out evaluation. The IMU-based TCN achieved the strongest event-detection performance, reaching 42.3 F1 and 83.0 AUROC, compared with 32.6 F1 and 77.2 AUROC for V-JEPA2 ego-video features. Although ego-video alone did not outperform IMU-based sensing, it showed above-chance discrimination, and qualitative analyses suggest that egocentric vision may capture FOG-relevant information independent of IMUs. Together, these results support the use of pretrained ego-video representations to add contextual information to wearable-sensor-based clinical motion understanding in daily living.

摘要:理解日常生活中的運動需要超越運動學的背景,因為在日常生活活動(ADLs)中類似的慣性模式可能反映出有意的停止、物體互動或病理性運動障礙。自我中心的視覺提供了與任務相關的背景,可能有助於消除這些情況的歧義。我們通過在帕金森病(PD)中的步態凍結(FOG)檢測來研究這一挑戰,這是一種在日常生活活動中受到背景因素強烈影響的症狀。使用同步的自我中心視頻、可穿戴IMU和從13名PD參與者在家中收集的專家註釋FOG標籤,我們在留一個參與者的評估下評估來自預訓練自我視頻和時間序列基礎模型的凍結表示,以及從零開始訓練的基於IMU的TCN。基於IMU的TCN達到了最強的事件檢測性能,F1達到42.3,AUROC達到83.0,而V-JEPA2自我視頻特徵的F1為32.6,AUROC為77.2。儘管僅使用自我視頻並未超越基於IMU的感測,但它顯示出超過隨機的區分能力,定性分析表明,自我中心的視覺可能捕捉到與FOG相關的信息,這與IMU無關。總體而言,這些結果支持使用預訓練的自我視頻表示,為基於可穿戴傳感器的日常生活臨床運動理解添加背景信息。

Static analysis-guided agentic AI translation enables Rust as a full stack bioinformatics language

2608.13029v1 by Johan Henriksson

The field of bioinformatics struggles with legacy code - old code that is commonly used but may no longer have a maintainer, or may be written in an now-unfamiliar language (e.g. Perl, Fortran). This incurs maintenance cost (technical debt), but dynamically typed languages also negatively impacts the environment and fail to make use of modern hardware. Legacy code may also have security or safety problems that make it unsuited for use in clinical settings. Here we show that agentic AI, combined with static analysis, can be used to translate legacy code to the modern language Rust. We provide prompts and supporting software to aid systematic translation, and evaluate it on common software for NGS and imaging. We showcase the result on our software Bascet: Size was reduced by ~80x, build time decreased by ~10x, and performance of key steps improved >3x. Unix dependencies were also removed, making Bascet the only single-cell pipeline able to run on native Windows, without a container. Large-scale refactoring of bioinformatics software is thus now possible at a limited budget, enabling more complex tools to be developed.

摘要:生物資訊學領域面臨著舊有代碼的挑戰——這些舊代碼通常被使用,但可能不再有維護者,或可能是用現在不熟悉的語言(例如 Perl、Fortran)編寫的。這會產生維護成本(技術負債),但動態類型語言也會對環境產生負面影響,並未能充分利用現代硬體。舊代碼可能還存在安全或安全性問題,使其不適合在臨床環境中使用。在這裡,我們展示了代理式 AI 結合靜態分析,可以用來將舊代碼轉換為現代語言 Rust。我們提供提示和支持軟體以協助系統性翻譯,並在 NGS 和成像的常見軟體上進行評估。我們展示了我們的軟體 Bascet 的結果:大小減少約 80 倍,建構時間減少約 10 倍,關鍵步驟的性能提高了超過 3 倍。Unix 依賴也被移除,使 Bascet 成為唯一能在本地 Windows 上運行的單細胞管道,而無需容器。因此,生物資訊學軟體的大規模重構現在在有限的預算下成為可能,從而使得更複雜的工具得以開發。

Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence

2608.12928v1 by Jakub Pokrywka, Łukasz Grzybowski, Antoni Lasik, Marek Kubis, Jeremi Ignacy Kaczmarek, Wojciech Kusa

We introduce a Polish-language medical visual question answering (VQA) benchmark, built from Polish Board Certification Examination questions for licensed physicians and dentists pursuing specialist certification. The benchmark comprises image-containing questions spanning diverse medical specialties and visual domains, together with a text-only question answering (QA) control set. We evaluate Polish-oriented, general-purpose open-weight, and commercial vision-language models. The task remains challenging: the best model achieves 79.0\% accuracy on the full VQA set, and only GPT-5.6 surpasses the approximate human reference on the subset with available candidate responses; all other evaluated models perform worse than humans. To assess visual grounding, we compare complete inputs with configurations omitting the image, the question, or both, and categorize questions by image importance. Models derive more useful information from the question text than from the image and perform worse on image-dominant questions. Across both QA and VQA, they nevertheless achieve above-chance accuracy from the answer choices alone, showing that non-trivial performance can persist even when key task components are missing.

摘要:我們介紹了一個波蘭語醫學視覺問題回答(VQA)基準,該基準是基於波蘭醫師和牙醫專業認證考試問題而建立的。這個基準包含了涵蓋多種醫學專業和視覺領域的圖像問題,以及一組僅包含文本的問題回答(QA)控制集。我們評估了針對波蘭的通用開放權重和商業視覺語言模型。這項任務仍然具有挑戰性:最佳模型在完整的 VQA 集上達到 79.0\% 的準確率,只有 GPT-5.6 在具有可用候選回答的子集上超過了近似的人類參考;所有其他評估的模型表現都不如人類。為了評估視覺基礎,我們比較了完整輸入與省略圖像、問題或兩者的配置,並根據圖像的重要性對問題進行分類。模型從問題文本中獲取的有用信息多於從圖像中獲取的,並且在圖像主導的問題上表現較差。在 QA 和 VQA 中,它們仍然僅從答案選擇中實現了超過隨機的準確率,顯示出即使在缺少關鍵任務組件的情況下,非平凡的表現仍然可以持續存在。

CRAFT: LLM-Based Iterative Refinement for Temporal Reasoning over Clinical Narratives

2608.12779v1 by Chengyang He, Tahreem Arif, Marko Zivkovic, Lijing Wang, Yue Ning, Ping Wang

Understanding the temporal progression of symptoms in clinical narratives is critical for disease monitoring, safety surveillance, and causality assessment. Clinical narratives, however, rarely provide explicit temporal anchors. Current approaches to temporal information reasoning focus predominantly on pairwise relation classification across multi-visit and timestamp-rich records, leaving the reconstruction of structured symptom trajectories from individual anchor-sparse reports largely unaddressed. We propose CRAFT, an LLM framework that pairs a generator with a constraint-based verifier to iteratively produce and refine stage-wise symptom timelines through targeted feedback. We conduct evaluation on MedTempo, a new benchmark of 5,347 vaccine adverse-event narratives spanning three COVID-19 vaccine types, with expert-validated temporal stage annotations for 3,166 reports. Experiments across four LLM backbones demonstrate that CRAFT consistently improves temporal ordering accuracy, with ablation analysis isolating the contribution of generator and verifier components across model capability levels.

摘要:理解臨床敘述中症狀的時間進展對於疾病監測、安全監控和因果評估至關重要。然而,臨床敘述很少提供明確的時間錨點。目前對時間信息推理的方法主要集中在多次訪問和時間戳豐富記錄之間的成對關係分類,這使得從個別缺乏錨點的報告中重建結構化的症狀軌跡在很大程度上未得到解決。我們提出了CRAFT,一個將生成器與基於約束的驗證器配對的LLM框架,通過有針對性的反饋迭代生成和完善階段性症狀時間線。我們在MedTempo上進行評估,這是一個新的基準,包含5,347個疫苗不良事件敘述,涵蓋三種COVID-19疫苗類型,並對3,166個報告進行了專家驗證的時間階段標註。在四個LLM骨幹上進行的實驗表明,CRAFT持續提高了時間排序的準確性,並通過消融分析隔離了生成器和驗證器組件在模型能力水平上的貢獻。

Memorization Diagnostics for Code LLMs Should be Scale-Aware

2608.12771v1 by Prateek Kumar Rajput, Abdoul Aziz Bonkoungou, Alberick Euraste Djiré, Xunzhu Tang, Yewei Song, Iyiola Emmanuel Olatunji, El Hacen Diallo, Jacques Klein, Tegawendé F. Bissyandé

The extent to which large language models for code rely on memorization over genuine understanding remains highly debated. While current literature frequently reports widespread memorization, evaluating the underlying probing techniques across dense architectures reveals a severe breakdown in their utility at scale. Traditional encoder-style probes using perturbations such as synonym fuzzing or dead-code insertion struggle to expose memorization in scaled models, even on known-contaminated benchmarks, and decoder-style probes that rely on log probabilities show similar performance degradation. The specific mode of failure for these probes, particularly why such techniques disrupt smaller models but fail to impact larger ones, motivates us to untangle representation load from memorization rather than treating them as a single phenomenon. By applying invertible mathematical transforms to numeric problems, we isolate these two factors and reveal that scaled encoders successfully absorb substantial representation load while still converging on the correct family of solutions. In practical software engineering, this ability to adapt to varying surface forms is what truly matters for usability and generalizability in LLM and agentic applications. Whether a specific solution was seen during training becomes a much less pressing question because although memorization inflates scores on contaminated benchmarks, factoring out representation load makes it debatable how much we should truly care if a functional answer was originally memorized. Future evaluations must therefore be built around separating these phenomena rather than relying on methodologies that quietly entangle them.

摘要:大型語言模型在代碼方面依賴於記憶而非真正理解的程度仍然存在高度爭議。儘管目前的文獻經常報導廣泛的記憶現象,但對於密集架構中潛在探測技術的評估顯示,它們在大規模應用中的效用嚴重下降。傳統的編碼器風格探測器使用同義詞模糊或死代碼插入等擾動,難以在擴展模型中揭示記憶,即使在已知受污染的基準上也是如此,而依賴於對數概率的解碼器風格探測器表現出類似的性能下降。這些探測器的具體失效模式,特別是為什麼這些技術會干擾較小的模型但無法影響較大的模型,促使我們將表示負載與記憶分開,而不是將它們視為單一現象。通過對數值問題應用可逆數學變換,我們將這兩個因素隔離,並揭示擴展的編碼器成功吸收了大量的表示負載,同時仍然收斂於正確的解決方案族。在實際的軟體工程中,這種適應不同表面形式的能力對於大型語言模型和代理應用的可用性和通用性來說才是真正重要的。在訓練期間是否見過特定解決方案變得不再是個緊迫的問題,因為儘管記憶會在受污染的基準上膨脹分數,但剔除表示負載使得我們對於一個功能性答案最初是否被記住的關心程度變得可爭辯。因此,未來的評估必須圍繞分離這些現象建立,而不是依賴於那些靜默糾纏它們的方法論。

PatientAct: Theory-Grounded Mental Health Client Simulation

2608.12750v1 by Sahand Sabour, TszYam NG, Yaqian Chen, Guanqun Bi, Jialu Zhao, Minlie Huang

LLM-based simulated clients are increasingly used to train novice counselors, evaluate LLM therapists, and generate synthetic data. However, current simulators produce overly cooperative clients that disclose too readily, accept therapeutic reframes without resistance, and resolve core issues within a single session. We trace these issues to profiles that lack causal depth and behavioral mechanisms that treat all content as equally accessible. We present PatientAct, a framework for client simulation grounded in established clinical theories. Our profiles integrate the 5Ps clinical case formulation, providing causal depth without tying the design to any single therapeutic modality. During simulation, profiles include a dynamic memory layer in which items carry trust thresholds (e.g., symptoms are available early, whereas formative memories require a sustained therapeutic alliance). At each turn, the client's emotional reaction and behavior are modeled before generating a response. If the therapist approaches gated content, PatientAct expresses resistance in terms of quantity, content, and style rather than defaulting to cooperation or a single resistance pattern. We evaluate our framework on 40 clinical situations and demonstrate that it generates diverse profiles with high clinical plausibility. Moreover, PatientAct significantly outperforms the baselines, yielding substantial gains in resistance quality and behavioral realism. Our code and data will be publicly available via github.com/Sahandfer/PatientHub.

摘要:LLM 基礎的模擬客戶越來越多地用於訓練新手輔導員、評估 LLM 治療師和生成合成數據。然而,當前的模擬器產生過於合作的客戶,他們過於輕易地透露信息,毫無抵抗地接受治療重構,並在單一會話中解決核心問題。我們將這些問題追溯到缺乏因果深度的檔案和將所有內容視為同等可接近的行為機制。我們提出了 PatientAct,一個基於已建立臨床理論的客戶模擬框架。我們的檔案整合了 5Ps 臨床案例形成,提供因果深度而不將設計綁定於任何單一的治療模式。在模擬過程中,檔案包括一個動態記憶層,其中項目承載信任閾值(例如,症狀早期可用,而形成性記憶則需要持續的治療聯盟)。在每一輪中,客戶的情感反應和行為在生成回應之前被建模。如果治療師接觸到受限內容,PatientAct 會在數量、內容和風格上表達抵抗,而不是默認合作或單一的抵抗模式。我們在 40 個臨床情境中評估了我們的框架,並證明它生成了具有高臨床合理性的多樣化檔案。此外,PatientAct 顯著超越了基準,帶來了在抵抗質量和行為現實主義方面的重大提升。我們的代碼和數據將通過 github.com/Sahandfer/PatientHub 公開提供。

Mr3D-VL: A generalist vision language foundation model for Multiparametric 3D Magnetic Resonance Imaging

2608.12689v1 by Zhi Qiao, Xintong Wu, Yichu He, Feng Shi

Multi-parametric magnetic resonance imaging (mpMRI) is a cornerstone for brain tumor diagnosis and treatment, yet current AI models face critical limitations: their lack of natural language interaction and interpretability impedes spatial information integration and cross-modal reasoning required clinically. Key challenges arise from significant physical meaning differences across modalities, spatial misalignment due to scan intervals, and the need for complex multi-feature interpretation in tasks like glioma grading. While visual-language models (VLMs) show promise in cross-modal understanding, existing methods focus mainly on 2D image modeling, neglecting direct perception of 3D volumetric space. Although 3D VLMs have been proposed for report generation and feature alignment in 3D CT imaging, mpMRI applications demand collaborative inference across multiple imaging modalities-a requirement unmet by current solutions. To address this, we introduce Mr3D-VL, a dedicated visual-language foundation model for multi-parametric 3D MRI. With 4 billion parameters, it employs an unsupervised pre-trained shared 3D encoder and 4D rotational positional embedding for dual modality-spatial integration. Its cross-modal projection layer uses a multi-resolution feature implantation strategy to enhance feature perception across resolutions. Experimental results show significant improvements over existing 4B/7B/30B domain-specific and general-purpose models in text generation tasks, achieving a BERTScore of 0.856 for report generation, with question-answering accuracy at 0.713 and multiple-choice accuracy at 0.912.

摘要:多參數磁共振成像(mpMRI)是腦腫瘤診斷和治療的基石,但目前的AI模型面臨重大限制:缺乏自然語言互動和可解釋性妨礙了臨床所需的空間信息整合和跨模態推理。主要挑戰來自於不同模態之間顯著的物理意義差異、由於掃描間隔造成的空間錯位,以及在如膠質瘤分級等任務中對複雜多特徵解釋的需求。儘管視覺語言模型(VLMs)在跨模態理解方面顯示出潛力,但現有方法主要集中在2D圖像建模,忽略了對3D體積空間的直接感知。雖然已提出3D VLMs用於報告生成和3D CT成像中的特徵對齊,但mpMRI應用需要跨多個成像模態的協作推理——這一需求目前的解決方案無法滿足。為了解決這個問題,我們推出了Mr3D-VL,一個專門針對多參數3D MRI的視覺語言基礎模型。它擁有40億個參數,採用無監督預訓練的共享3D編碼器和4D旋轉位置嵌入進行雙模態空間整合。其跨模態投影層使用多解析度特徵植入策略來增強不同解析度間的特徵感知。實驗結果顯示,在文本生成任務中,與現有的4B/7B/30B領域特定和通用模型相比,顯著提高了性能,報告生成的BERTScore達到0.856,問答準確率為0.713,多選準確率為0.912。

SteerBench-Work: A Benchmark for Agent Steering at Action Boundaries

2608.12654v1 by Oguz Serdar, Cuneyt Mertayak

Long-running LLM agents act through tools, and a single step can send an email, merge a pull request, or wire a payment. The steering decision is the pre-commit choice at that boundary: proceed, or hold for human or policy review. We introduce SteerBench-Work, an incident-anchored, bidirectional benchmark for that decision in workplace agents across developer operations, customer service, finance, legal, medical, HR, and security. Release v2026-05 contains 106 scenarios anchored in public incidents, paired evidence-reversed mirrors, and calibration controls, with labels split nearly evenly between proceed and hold so the two error directions get near-identical numbers of chances. A model sees the proposed action and the available evidence, returns a gate decision, and is scored on whether it crosses or holds the boundary correctly. Across 30 model conditions the failures run almost entirely in one direction: models wrongly hold authorized, evidence-cleared work on 28.1% of opportunities and wrongly allow unsafe work on 1.0%. The hardest cases are risk-resolved commits, where signed or structured evidence has already cleared a real risk trigger, and models score markedly worse on evidence-reversed mirrors of famous incidents (63.8%) than on the incidents themselves (98.5%). General capability is not the same as steering calibration: higher-capability models often over-refuse at the commit boundary, and more reasoning can repair a weak gate while leaving a calibrated one flat. The public leaderboard is at steerbench.com.

摘要:長期運行的 LLM 代理透過工具進行操作,單一步驟可以發送電子郵件、合併拉取請求或進行付款。引導決策是在那個邊界的預提交選擇:繼續,或等待人類或政策審查。我們介紹 SteerBench-Work,這是一個以事件為基礎的雙向基準,用於在開發運營、客戶服務、金融、法律、醫療、人力資源和安全等工作場所代理中的該決策。版本 v2026-05 包含 106 個基於公共事件的場景,配對的證據反向鏡像和校準控制,標籤在繼續和保持之間幾乎均勻分配,以便兩個錯誤方向獲得幾乎相同的機會數量。一個模型看到提議的行動和可用的證據,返回一個閘決策,並根據它是否正確地跨越或保持邊界進行評分。在 30 種模型條件下,失敗幾乎完全朝一個方向發生:模型在 28.1% 的機會中錯誤地保持已授權、證據清除的工作,並在 1.0% 的情況下錯誤地允許不安全的工作。最困難的情況是風險已解決的提交,其中簽署或結構化證據已經清除了實際風險觸發器,並且模型在著名事件的證據反向鏡像上得分明顯低於事件本身(63.8% 對 98.5%)。一般能力與引導校準並不相同:高能力模型在提交邊界上往往過度拒絕,而更多的推理可以修復一個弱閘,同時保持一個已校準的閘平坦。公共排行榜位於 steerbench.com。

Algorithm Design and Physician Liability

2608.13618v1 by Shujie Luan, Shubhranshu Singh, Tinglong Dai

A single clinical algorithm can deliver unequal accuracy across patient groups, and concern about such disparity has grown as artificial intelligence (AI) spreads through clinical decision-making. In response, a liability rule introduced in the United States holds healthcare providers responsible when their reliance on disparate algorithms contributes to erroneous clinical decisions. We examine how such liability considerations reshape (i) an AI firm's algorithm design decisions that drive group-specific accuracy and (ii) a physician's decisions to use AI in healthcare delivery. The AI firm designs an algorithm for two patient groups, and improving accuracy for the disadvantaged group is more costly. The physician (who remains the accountable decision-maker) then decides whether to consult AI, weighing the reduction in clinical uncertainty against expected liability exposure when AI errors disproportionately affect the disadvantaged group. We find the liability rule can induce disparate use of AI: the physician may reduce AI use overall and, over an intermediate range of liability, rely on AI less for disadvantaged patients. The effect is non-monotone. As liability increases, the physician's use of AI for disadvantaged patients first declines, then rises as the firm reallocates investment toward reducing disparity or switches to an equal-accuracy design. Mandating equal algorithmic accuracy across patient groups can then inadvertently harm both groups, because a uniform accuracy requirement distorts the firm's investment incentives and the physician's equilibrium AI-use decisions.

摘要:單一的臨床演算法在不同患者群體中可能會產生不均等的準確性,隨著人工智慧(AI)在臨床決策中的普及,對於這種差異的關注也日益增加。作為回應,美國引入了一項責任規則,當醫療提供者依賴不同的演算法導致錯誤的臨床決策時,將其負責。我們研究這種責任考量如何重塑(i)AI公司的演算法設計決策,促進特定群體的準確性,以及(ii)醫生在醫療提供中使用AI的決策。AI公司為兩個患者群體設計了一個演算法,改善弱勢群體的準確性成本更高。然後,醫生(仍然是負責的決策者)決定是否諮詢AI,權衡臨床不確定性的減少與當AI錯誤不成比例地影響弱勢群體時的預期責任風險。我們發現責任規則可能會導致AI的使用不均等:醫生可能會整體減少AI的使用,並且在責任的中等範圍內,對弱勢患者的AI依賴程度降低。這一效果是非單調的。隨著責任的增加,醫生對弱勢患者使用AI的情況最初下降,然後隨著公司將投資重新分配到減少差異或轉向平等準確性設計而上升。要求在患者群體之間達到平等的演算法準確性,可能會無意中對兩個群體造成傷害,因為統一的準確性要求扭曲了公司的投資激勵和醫生的均衡AI使用決策。

Auditable agentic AI for evidence-grounded thyroid ultrasound diagnosis and reporting

2608.12590v1 by Haifan Gong, Shiyu Chen, Bodong Wang, Yuqi Wang, Shijie Wang, Guoliang You, Xinyu Xiong, Haowei Wang, Mingzhi Mao, Dexing Kong, Qinghua Liu, Wei Lou, Fei Chen, Guanbin Li

Thyroid ultrasound diagnosis requires coordinated lesion localization, measurement, risk stratification and reporting, yet most AI systems address these tasks in isolation and provide limited support for clinical review. We present ThyroidXAgent, a clinician-interactive agentic AI system that coordinates specialized diagnostic tools and stores their outputs as an auditable case-level evidence record. The system was developed using OpenThyroidDB, a multicentre, multitask resource integrating approximately 0.3 million ultrasound images and 24,000 paired reports, and was evaluated on 28,458 non-overlapping test cases, including 8,721 cases from 35 centres in the private NHC-MISD-TUS cohort. Across heterogeneous datasets, ThyroidXAgent achieved a mean Dice score of 87.21 percent for nodule segmentation and a mean AUROC of 0.9466 for benign-malignant classification. The same workflow supported lymph-node metastasis prediction and follicular versus papillary thyroid carcinoma classification, with AUROCs of 0.864 and 0.805, respectively. For report generation, evidence-grounded assembly outperformed multimodal language-model baselines across three cohorts. ThyClinScore, a lesion-level clinical semantic metric introduced here, showed the strongest correlation with a location-aware language-model judge. ThyroidXAgent improved physician classification accuracy, increased report diagnostic consistency from 70.3 percent to 86.2 percent, and reduced segmentation and reporting time by 35.9 percent and 27.4 percent, respectively. These findings support auditable, clinician-correctable agentic AI for thyroid ultrasound diagnosis and reporting.

摘要:甲狀腺超聲診斷需要協調病變定位、測量、風險分層和報告,但大多數人工智慧系統在孤立的情況下處理這些任務,並對臨床審查提供有限的支持。我們提出了ThyroidXAgent,一個臨床互動的代理人工智慧系統,協調專門的診斷工具並將其輸出存儲為可審計的案例級證據記錄。該系統是使用OpenThyroidDB開發的,這是一個多中心、多任務的資源,整合了約30萬張超聲圖像和24,000份配對報告,並在28,458個不重疊的測試案例上進行了評估,包括來自私立NHC-MISD-TUS隊列的35個中心的8,721個案例。在異質數據集上,ThyroidXAgent在結節分割方面達到了87.21%的平均Dice分數,並在良惡性分類方面達到了0.9466的平均AUROC。同一工作流程支持淋巴結轉移預測和濾泡型與乳頭狀甲狀腺癌的分類,AUROC分別為0.864和0.805。在報告生成方面,基於證據的組合在三個隊列中超越了多模態語言模型基準。這裡引入的ThyClinScore,一個病變級的臨床語義指標,顯示出與位置感知語言模型評審者之間的最強相關性。ThyroidXAgent提高了醫生的分類準確性,將報告的診斷一致性從70.3%提高到86.2%,並分別減少了35.9%和27.4%的分割和報告時間。這些發現支持可審計、可由臨床醫生修正的代理人工智慧用於甲狀腺超聲診斷和報告。

How Compliant is Sepsis Treatment? An Expert-Guided Neuro-symbolic Pipeline for Generating Clinical Compliance Insights

2608.13617v1 by Himanshu Tripathi, Kaushik Roy, Subash Neupane, Shahram Rahimi

Verifying whether clinical care follows evidence-based protocols is a natural neuro-symbolic problem, yet the safety-critical setting defeats either paradigm alone. We present an expert-guided pipeline that constrains a large language model strictly to semantic normalization, mapping messy drug and microbiology strings onto a fixed clinical vocabulary, while a Sugeno fuzzy inference system reasons over the normalized events. The fuzzy layer encodes eight Surviving Sepsis Campaign bundle rules and replaces binary judgments with graded scores in [0,1]. Applied to 2,438 MIMIC-IV v3.1 sepsis episodes, it surfaces antibiotic timing as the most critical breakdown (mean 0.24, 13% within one hour), Hour-1 underperformance (mean 36.7%), a 51% elevated-lactate drop-off, and descriptive differences in ICU stay across compliance groups (3.8 versus 5.1 days).

摘要:驗證臨床護理是否遵循基於證據的協議是一個自然的神經符號問題,但安全關鍵的環境使得單一範式無法應對。我們提出了一個專家指導的流程,將大型語言模型嚴格限制於語義標準化,將雜亂的藥物和微生物學字符串映射到固定的臨床詞彙上,同時一個Sugeno模糊推理系統對標準化事件進行推理。模糊層編碼了八條存活敗血症運動的捆綁規則,並用[0,1]範圍內的分數取代了二元判斷。應用於2,438個MIMIC-IV v3.1敗血症事件中,它顯示抗生素使用時機是最關鍵的破綻(平均0.24,13%在一小時內),第一小時表現不佳(平均36.7%),51%的乳酸升高下降,以及在合規性組之間ICU住院天數的描述性差異(3.8天對5.1天)。

M-Net: Integrating Spectral Features and Physical Field Operators into Deep Learning for Medical Image Segmentation

2608.12196v1 by Jing Zhu, Ye Wang, Fumin Wang

Purpose: Deep learning-based medical image segmentation has achieved remarkable success, yet purely data-driven approaches often fail to exploit the rich mathematical structure inherent in medical images. We investigate whether explicit mathematical inductive biases, specifically matrix spectral analysis and vector calculus operators, can enhance segmentation beyond data-driven learning alone. Methods: We propose M-Net (Math-Augmented Network), which integrates three complementary mathematical priors into U-Net: (1) continuous spectral features derived from the condition number of centered local pixel matrices, providing a differentiable measure of texture ill-conditioning; (2) physical field operators (divergence and a discrete curl-like boundary irregularity operator) computed from image gradient fields, capturing focal intensity extrema and edge non-smoothness; and (3) a Math-Attention Gate (MAG) that adaptively fuses mathematical features with CNN-extracted deep features at skip connections. Results: Experiments on three benchmarks (LiTS, KiTS, and BraTS) show that M-Net achieves Dice scores of 78.42%, 76.15%, and 83.67%, outperforming baseline U-Net by 12.37%, 3.52%, and 5.55% on liver, kidney, and brain tumor segmentation, respectively. Ablations reveal that the condition-number feature contributes a 2.14% gain over binary invertibility features, while MAG adds 1.45% over simple concatenation. Conclusion: M-Net establishes that mathematical inductive biases provide effective complementary information for medical image segmentation. The continuous condition-number feature offers superior gradient information over discrete alternatives, and MAG preserves these priors throughout the network. This work opens avenues for integrating linear algebra and vector calculus into deep architectures for medical imaging.

摘要:目的:基於深度學習的醫學影像分割已取得顯著成功,然而純數據驅動的方法往往未能充分利用醫學影像中固有的豐富數學結構。我們探討明確的數學歸納偏差,特別是矩陣譜分析和向量微積分運算子,是否能超越僅依賴數據驅動學習來增強分割效果。方法:我們提出M-Net(數學增強網絡),將三個互補的數學先驗整合到U-Net中:(1)從中心局部像素矩陣的條件數導出的連續譜特徵,提供可微分的紋理不良條件度量;(2)從影像梯度場計算的物理場運算子(散度和離散的旋度邊界不規則性運算子),捕捉焦點強度極值和邊緣不光滑性;(3)一個數學注意力閘(MAG),在跳躍連接中自適應地融合數學特徵與CNN提取的深度特徵。結果:在三個基準測試(LiTS、KiTS和BraTS)上的實驗顯示,M-Net在肝臟、腎臟和腦腫瘤分割中分別達到78.42%、76.15%和83.67%的Dice分數,分別比基線U-Net高出12.37%、3.52%和5.55%。消融實驗顯示,條件數特徵比二元可逆性特徵貢獻了2.14%的增益,而MAG比簡單的串接增加了1.45%。結論:M-Net證明數學歸納偏差為醫學影像分割提供了有效的互補信息。連續的條件數特徵提供了比離散替代方案更優越的梯度信息,而MAG在整個網絡中保留了這些先驗。這項工作為將線性代數和向量微積分整合到醫學影像的深度架構中開辟了新的途徑。

A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench

2608.12138v1 by Praveen Reddy, Charuta Mandke, Suvrankar Datta, Sarah Khan, Siddharth Reddy Anthireddy, Shitij Arora, Vishal Singh

General-purpose large language models (LLMs) have recently been reported to match or exceed specialized clinical AI tools on medical benchmarks, but such comparisons draw on a narrow set of systems and on benchmarks developed largely in high-income settings. We evaluate VITA, a retrieval-augmented generation (RAG) system purpose-built for contextual knowledge retrieval in India and other low- and middle-income (LMIC) settings. VITA retrieves from a curated corpus of disease-specific guidelines, India-specific antimicrobial resistance data, national formulary constraints, and resource-limited care protocols; its architecture and corpus are proprietary, but the benchmark, the physician-written rubrics, and our full response and scoring outputs are public for independent verification. On 4,023 English-language HealthBench questions (80.5% of the benchmark), scored with a GPT-4.1 judge, VITA ranked first with 51.9% of possible rubric points, ahead of GPT-5.4 (46.1%), o4-mini (44.3%), Gemini 3.1 Pro (42.6%), and Claude Sonnet 4.6 (37.3%), and scored highest on 45.4% of questions. To test robustness to newer models and judge lineage, a 500-question subset was re-run against current-generation models (GPT-5.5, Claude Opus 4.8, Gemini 3.5 Pro, Grok 4.3) and graded by a neutral open-weight judge (DeepSeek-V4-Pro) sharing no lineage with any system tested. Here the gap narrowed to parity: VITA and GPT-5.5 were statistically indistinguishable on mean per-question score, while VITA led on points-weighted score and won the most questions. VITA's advantages in accuracy and completeness persisted under the neutral judge; its communication scores were lower. These results indicate that a purpose-built clinical RAG system remains competitive with frontier LLMs on an open benchmark, consistent with corpus specificity as a design variable that improves grounding at some cost to communication polish.

摘要:一般用途的大型語言模型(LLMs)最近被報導在醫療基準上與專門的臨床人工智慧工具相匹配或超越,但這些比較依賴於一組狹窄的系統以及主要在高收入環境中開發的基準。我們評估了VITA,一個專為印度及其他低收入和中等收入(LMIC)環境中的上下文知識檢索而設計的檢索增強生成(RAG)系統。VITA從一個策劃的特定疾病指導方針、印度特定的抗微生物抗藥性數據、國家藥典限制以及資源有限的護理協議中檢索資料;其架構和語料庫是專有的,但基準、醫生撰寫的評分標準以及我們的完整回應和評分輸出是公開的,以便獨立驗證。在4,023個英語HealthBench問題(基準的80.5%)上,使用GPT-4.1評判,VITA以51.9%的可能評分點排名第一,超過了GPT-5.4(46.1%)、o4-mini(44.3%)、Gemini 3.1 Pro(42.6%)和Claude Sonnet 4.6(37.3%),並在45.4%的問題上得分最高。為了測試對新模型的穩健性和評判系譜,對500個問題的子集再次運行,與當前一代模型(GPT-5.5、Claude Opus 4.8、Gemini 3.5 Pro、Grok 4.3)進行比較,並由一位中立的開放權重評判(DeepSeek-V4-Pro)進行評分,該評判與任何測試系統無關。此時差距縮小至平行:VITA和GPT-5.5在每題平均得分上統計上無法區分,而VITA在加權得分上領先並贏得了最多問題。VITA在準確性和完整性上的優勢在中立評判下持續存在;其溝通得分較低。這些結果表明,專為臨床設計的RAG系統在公開基準上仍然與前沿LLMs具有競爭力,這與語料庫的特異性作為設計變量相一致,該變量在某種程度上提高了基礎性,但對溝通的精緻性造成了成本。

Do LLMs Take Care of Their Own? Similarity Signals Can Induce Cooperation

2608.12125v1 by Akash Kundu, Emanuel Tewolde, Ratip Emin Berker, Samuel F. Brown, Vincent Conitzer

As LLM-based agents with user-instructed goals are becoming widely deployed, they increasingly encounter each other in strategic interactions, and face challenges of finding mutually beneficial outcomes. Prior literature has argued that cooperation problems such as the Prisoner's Dilemma are resolvable in settings where agents know they follow very similar decision making patterns, as for example in monocultural AI ecosystems. Following that line of work, this paper introduces the first framework for evaluating LLM decision making when agents are provided with graded similarity signals. Among our findings, we establish that different LLM models vary drastically in how they navigate similarity signals, with some modern models showing consistent behavior across cooperation problems, payoff structures, and prompt framing. Perhaps surprisingly, our experiments also show that the dataset based on which the similarity signal is computed has small to no impact on induced cooperation, and that LLM models systematically self-identify as highly similar when asked to evaluate another model's chain-of-thought reasoning by themselves. Finally, we develop an LLM-behavioral-game-theoretic model that captures some of their reasoning rationale, and show that it can support cooperative outcomes in equilibrium under sufficiently high similarity scores.

摘要:隨著基於大型語言模型(LLM)的代理人以用戶指導的目標被廣泛部署,它們在戰略互動中越來越多地相遇,並面臨尋找互利結果的挑戰。先前的文獻已經論證,合作問題如囚徒困境在代理人知道它們遵循非常相似的決策模式的情況下是可以解決的,例如在單一文化的人工智慧生態系統中。沿著這一研究方向,本文介紹了第一個評估LLM決策制定的框架,當代理人被提供分級相似性信號時。在我們的發現中,我們確立了不同的LLM模型在如何導航相似性信號方面存在巨大差異,一些現代模型在合作問題、收益結構和提示框架中顯示出一致的行為。或許令人驚訝的是,我們的實驗還顯示,計算相似性信號的數據集對於誘發合作的影響微乎其微,且當被要求自行評估另一模型的思考鏈推理時,LLM模型系統性地自我識別為高度相似。最後,我們開發了一個LLM行為博弈論模型,捕捉它們的一些推理理由,並顯示它可以在足夠高的相似性分數下支持均衡的合作結果。

How Far from Clinical Deployment? Evaluating the Complete Unsupervised Domain Adaptation Pipeline in Medical Imaging

2608.12035v1 by Yiheng Xiong, Luisa Gallée, Daniel Santak Wolf, Heiko Hillenhagen, Michael Götz

Deploying unsupervised domain adaptation (UDA) in clinical practice requires choosing which algorithm to use and which of its trained models to ship. However, the deployment (target) domain is unlabeled, so models cannot be evaluated directly on it, leaving it unclear which to select. We address this by evaluating the complete UDA pipeline, considering both adaptation and label-free selection together. Our study covers eleven clinically relevant cross-domain scenarios from nine medical imaging datasets, with ten UDA algorithms and 13 label-free selection methods (validators), evaluating over 80,000 trained models in total. By this, we find that a capable adapted model usually exists, but identifying it without target labels is difficult: the validator-selected models leave a large and structural target performance gap to the best available one, with no evaluated validator consistently reliable. Towards closing it, we explore two strategies, ensembling and a small target-labeling budget; both narrow this gap but do not close it entirely. Overall, deployable UDA depends on the complete pipeline; addressing the less explored selection step could bring much of current UDA closer to clinical use.

摘要:在臨床實踐中部署無監督領域適應(UDA)需要選擇使用哪種算法以及其訓練模型中的哪一個進行部署。然而,部署(目標)領域是未標記的,因此無法直接對其進行模型評估,這使得選擇變得不明確。我們通過評估完整的 UDA 流程來解決這個問題,同時考慮適應和無標籤選擇。我們的研究涵蓋了來自九個醫學影像數據集的十一個臨床相關的跨領域場景,使用十種 UDA 算法和 13 種無標籤選擇方法(驗證器),總共評估了超過 80,000 個訓練模型。由此,我們發現通常存在一個能夠適應的模型,但在沒有目標標籤的情況下識別它是困難的:驗證器選擇的模型與可用的最佳模型之間存在著較大且結構性的目標性能差距,且沒有一個評估過的驗證器是一致可靠的。為了縮小這個差距,我們探索了兩種策略,集成和小型目標標記預算;這兩者都縮小了這個差距,但並未完全關閉它。總的來說,可部署的 UDA 依賴於完整的流程;解決較少探索的選擇步驟可能會使當前的 UDA 更接近臨床應用。

From Safety Documentation to Safety Knowledge Support: An Evidence-Grounded LLM Framework for Medical Devices

2608.12025v1 by Tuhinangshu Gangopadhyay, Rasmus Adler, Peter Liggesmeyer, Jan Reich

Medical devices are becoming more software-intensive, connected, and AI-enabled. Their development requires risk-management evidence aligned with ISO 14971 and, for software, IEC 62304. This evidence must be kept consistent across requirements, design decisions, software changes, verification results, complaints, and post-market data. These tasks are costly and depend on scarce safety and domain experts. Large language models (LLMs) may reduce parts of this effort because medical-device safety work is highly document-based. However, current LLM-based safety-engineering studies often address isolated methods, rely on generic prompting or public examples, and provide limited support for source links, traceability, uncertainty handling, lifecycle updates, and recorded expert review. This limits their use in regulated medical-device development. This paper argues that the central research problem is not safety-text generation, but source-linked safety-knowledge support. We propose an evidence-grounded framework that connects device artifacts, controlled knowledge storage and retrieval, method-specific generation of candidate safety items, critique and uncertainty checks, and recorded expert review. The framework prepares, links, checks, and updates candidate safety artifacts for expert decision-making. It does not decide whether a device is safe and does not provide regulatory approval. We also outline an evaluation strategy using non-public or newly built medical-device case studies and expert reference analyses to assess coverage, correctness, relevance, traceability, duplicate rate, unsupported claims, and review effort.

摘要:醫療器材正變得越來越依賴軟體、互聯網連接和人工智慧。它們的開發需要符合ISO 14971的風險管理證據,對於軟體則需要符合IEC 62304。這些證據必須在需求、設計決策、軟體變更、驗證結果、投訴和市場後數據之間保持一致。這些任務成本高昂,並依賴於稀缺的安全和領域專家。大型語言模型(LLMs)可能會減少這部分工作,因為醫療器材的安全工作高度依賴文檔。然而,目前基於LLM的安全工程研究往往針對孤立的方法,依賴於通用提示或公共範例,並對來源鏈接、可追溯性、不確定性處理、生命周期更新和記錄的專家審查提供有限支持。這限制了它們在受監管的醫療器材開發中的應用。本文主張,核心研究問題不是安全文本生成,而是來源鏈接的安全知識支持。我們提出了一個基於證據的框架,連接設備文檔、受控知識存儲和檢索、特定方法生成候選安全項目、批評和不確定性檢查,以及記錄的專家審查。該框架為專家決策準備、鏈接、檢查和更新候選安全文檔。它不決定設備是否安全,也不提供監管批准。我們還概述了一個評估策略,使用非公開或新建的醫療器材案例研究和專家參考分析來評估覆蓋範圍、正確性、相關性、可追溯性、重複率、不支持的聲明和審查工作量。

When the API Speaks the Wrong Language: Revisiting Post-Training for Multilingual Tool Use

2608.11715v1 by Siddharth Chauhan, Thomas Butler, Abhishek Singhania, Pankaj Porwal, Honey Gupta

The reliability of Large Language Models (LLMs) for API calling degrades in multilingual settings. A common failure occurs when a model selects the correct tool but generates argument values in an inconsistent language, which we term Argument Language Mismatch (ALM). Although semantically correct, such outputs are operationally invalid and not captured by standard API-calling metrics. We revisit post-training strategies for mitigating ALM and find that, in our benchmark, supervised fine-tuning (SFT) provides a strong baseline, substantially improving argument language consistency and end-to-end function call accuracy. Under consistent model selection, SFT achieves performance comparable to, and sometimes exceeding more complex reinforcement learning (RL) approaches. We further examine whether RL with structured, argument-aware rewards offers additional benefits. While methods such as Group Relative Policy Optimization (GRPO) can improve language consistency and better preserve general reasoning ability, these gains are incremental and most pronounced in generalization and multi-objective trade-offs. Overall, our results suggest that much of the performance in multilingual API grounding can be achieved through careful supervised training, with RL providing targeted rather than fundamental improvements.

摘要:大型語言模型(LLMs)在多語言環境中進行 API 呼叫的可靠性會下降。常見的失敗情況是模型選擇了正確的工具,但生成的參數值卻使用不一致的語言,我們稱之為參數語言不匹配(ALM)。雖然這些輸出在語義上是正確的,但在操作上是無效的,並未被標準的 API 呼叫指標所捕捉。我們重新檢視了減輕 ALM 的後訓練策略,發現根據我們的基準,監督微調(SFT)提供了強有力的基線,顯著改善了參數語言的一致性和端到端函數調用的準確性。在一致的模型選擇下,SFT 的表現可與有時超越更複雜的強化學習(RL)方法相媲美。我們進一步檢查了結構化的、參數感知的獎勵是否提供了額外的好處。雖然像群體相對政策優化(GRPO)這樣的方法可以改善語言一致性並更好地保留一般推理能力,但這些增益是漸進的,並且在泛化和多目標權衡中最為明顯。總體而言,我們的結果表明,多語言 API 基礎的性能大部分可以通過謹慎的監督訓練來實現,而 RL 則提供了針對性的而非根本性的改進。

Rubric Dropout: A Simple Way to Mitigate Reward Hacking in Rubric-as-Reward RL

2608.11669v1 by Minglai Yang, Xinyu Guo, Utkarsh Tyagi, Mian Zhang, Razvan Dumitru, Sunjie Hou, Yunzhong He, Daniel Yue Zhang, Ying Liu

Reinforcement learning against rubrics, lists of criteria graded by an LLM judge, has become a standard way to post-train language models on tasks with no deterministic answer. The rubric, however, is a fixed proxy for quality, never a complete description of it, and a policy trained against it long enough will learn to exploit the difference. We measure this directly. Training Qwen3-8B with Group Relative Policy Optimization (GRPO) on medical and science rubrics and grading out-of-distribution (OOD) benchmarks with both the training judge and a stronger gold judge, we find that the two scores diverge during training. The training judge's score keeps climbing while the gold judge's score peaks and then falls, by 3 points on HealthBench-Hard and by 22 points on ResearchQA. A judge with a fixed bias would shift the gold curve by a constant, not send it down while the training score rises, so the divergence is reward hacking, not judge noise. We propose Rubric Dropout, a one-line fix borrowed from neuron dropout. At every step, we randomly drop a subset of the rubric's criteria before computing the reward, so the policy never optimizes the same rubric twice. The dropped subset is shared across each rollout group, so GRPO's group-relative advantages stay comparable, and evaluation always uses the full rubric. Comparing no dropout against dropout at 30% and 50% on both benchmark pairs, dropout raises the OOD gold score at every matched checkpoint (+1 to +2 points on HealthBench-Hard, +6 to +7 points on ResearchQA), lowers the two hacking measures we track, and costs nothing in domain. Sweeping the dropout fraction shows a broad 30-50% sweet spot, while the natural alternative, reweighting criteria by how useful they are to training, performs worse than no intervention at all in our setting.

摘要:強化學習對抗評分標準,即由LLM評審評分的標準列表,已成為對於沒有確定性答案的任務進行後訓練語言模型的標準方法。然而,評分標準是質量的固定代理,從來不是其完整描述,而對其訓練足夠長的策略將學會利用這一差異。我們直接測量這一點。使用群體相對策略優化(GRPO)對醫學和科學評分標準進行訓練Qwen3-8B,並用訓練評審和更強的金標準評審對分佈外(OOD)基準進行評分,我們發現這兩個分數在訓練過程中出現了分歧。訓練評審的分數持續上升,而金標準評審的分數達到峰值後下降,在HealthBench-Hard上下降了3分,在ResearchQA上下降了22分。具有固定偏見的評審會將金曲線平移一個常數,而不是在訓練分數上升時將其向下移動,因此這一分歧是獎勵黑客行為,而不是評審噪聲。我們提出了評分標準隨機失活(Rubric Dropout),這是一個借用自神經元隨機失活的一行修正。在每一步,我們在計算獎勵之前隨機丟棄評分標準的一部分標準,因此該策略從不對同一評分標準進行兩次優化。被丟棄的子集在每個展開組中共享,因此GRPO的群體相對優勢保持可比,評估始終使用完整的評分標準。將無隨機失活與在兩組基準上30%和50%的隨機失活進行比較,隨機失活在每個匹配的檢查點上提高了OOD金分數(在HealthBench-Hard上提高了1到2分,在ResearchQA上提高了6到7分),降低了我們跟踪的兩個黑客行為指標,且在領域上沒有任何成本。掃描隨機失活比例顯示出30-50%的廣泛甜蜜點,而自然的替代方案,即根據標準對訓練的有用性進行重加權,在我們的設置中表現得比沒有任何干預更差。

Beyond Single-Turn Confidence: Trajectory-Adapted Uncertainty Quantification for LLM Agents

2608.11552v1 by Dylan Bouchard, Mohit Singh Chauhan

Uncertainty quantification (UQ) methods for language models are typically evaluated on single-turn outputs, where uncertainty is attached to one generated answer. For LLM agents, however, the unit of observation is an interactive trajectory, where the model can ask clarifying questions, call tools, update state, and make intermediate decisions whose errors propagate to the final outcome. We study whether three common families of single-turn UQ methods transfer to this setting. Across five LLMs and four multi-turn tool-use datasets from BFCL-v4 and $τ^2$-bench, we evaluate white-box scorers based on action-token probabilities, black-box consistency scorers based on resampled trajectories, and reflexive scorers based on model self-assessment of the trajectory. We find that transfer is often useful but uneven. Token-probability scores are highly sensitive to the choice of aggregator used across turns, reflexive scores provide the strongest low-cost baseline in most evaluated settings, and black-box self-consistency is often the strongest UQ family, with trajectory-equivalence and action-set consistency typically ranking highest among its variants. These results suggest that UQ methods developed for single generations should be revalidated at the trajectory level, with careful attention to the consistency measurement, aggregator choice, and computational budget.

摘要:不確定性量化(UQ)方法對於語言模型通常是在單輪輸出上進行評估,其中不確定性與一個生成的答案相關聯。然而,對於大型語言模型(LLM)代理來說,觀察的單位是一個互動軌跡,其中模型可以提出澄清問題、調用工具、更新狀態,並做出中間決策,這些決策的錯誤會傳播到最終結果。我們研究了三種常見的單輪UQ方法是否能轉移到這種情境中。在五個LLM和來自BFCL-v4和$τ^2$-bench的四個多輪工具使用數據集上,我們評估了基於行動標記概率的白盒評分器、基於重新抽樣軌跡的黑盒一致性評分器,以及基於模型自我評估軌跡的反射性評分器。我們發現轉移通常是有用的,但不均勻。標記概率分數對於跨輪使用的聚合器選擇非常敏感,反射性分數在大多數評估設置中提供了最強的低成本基準,而黑盒自一致性通常是最強的UQ家族,其中軌跡等價性和行動集一致性通常在其變體中排名最高。這些結果表明,為單次生成開發的UQ方法應該在軌跡層面重新驗證,並仔細關注一致性測量、聚合器選擇和計算預算。

Social Chain of Thought: A Multi-Agent Architecture Grounded in Medical Differential Diagnosis Methodology

2608.11420v1 by Del Coburn, Scott Sanner, Dan Silver

Medical diagnostic reasoning is a high-impact use case for LLMs that carries significant implications for the health and wellbeing of users. When OpenAI (2026) reports that more than 5% of ChatGPT messages globally are healthcare-related, the transparency of these systems becomes a serious design concern. This is especially true for complex cases, where differential diagnosis often requires integrating multiple forms of specialist reasoning. Existing work has proposed multi-agent approaches to medical diagnosis, but it remains unclear when such systems are needed, why they help, and where they outperform monolithic inference. We introduce Social Chain of Thought (SCoT),a multi-round pipeline for medical differential diagnosis that structures multi-agent interaction as a deliberative framework for collabora. tive LLM reasoning. Evaluating SCoT against single-agent baselines, one-agent pipeline ablations, and best-of-n scaling, we show that its recall advantage is not reproduced by monolithic inference alone. SCoT is most successful in the hardest diagnostic cases, where multiple rounds of specialist conversation help recover ground-truth diagnoses and converge on a higher-recall differential.

摘要:醫療診斷推理是大型語言模型(LLMs)的高影響力應用案例,對用戶的健康和福祉具有重要意義。當 OpenAI(2026)報告全球超過 5% 的 ChatGPT 訊息與醫療保健相關時,這些系統的透明度成為一個嚴重的設計問題。這在複雜案例中特別真實,因為鑑別診斷通常需要整合多種專家推理形式。現有的研究已提出多代理的醫療診斷方法,但仍不清楚何時需要這樣的系統、它們為何有幫助,以及它們在哪些方面超越單一推理。我們介紹社會思維鏈(Social Chain of Thought, SCoT),這是一個多輪的醫療鑑別診斷管道,將多代理互動結構化為協作大型語言模型推理的深思框架。通過將 SCoT 與單代理基準、單代理管道消融和最佳擴展進行評估,我們顯示其召回優勢並非僅由單一推理所重現。SCoT 在最困難的診斷案例中最為成功,多輪專家對話有助於恢復真實診斷並收斂於更高召回率的鑑別診斷。

Gaze Target Estimation Anywhere with Concepts

2608.11367v1 by Xu Cao, Houze Yang, Vipin Gunda, Zhongyi Zhou, Tianyu Xu, Adarsh Kowdle, Inki Kim, James M. Rehg

Estimating human gaze targets from images in-the-wild is an important and formidable task. Existing approaches primarily employ brittle, multi-stage pipelines that require explicit inputs, like head bounding boxes and human pose, in order to identify the subject of gaze analysis. As a result, detection errors can cascade and lead to failure. Moreover, these prior works lack the flexibility of specifying the gaze analysis task via natural language prompting, an approach which has been shown to have significant benefits in convenience and scalability for other image analysis tasks. To overcome these limitations, we introduce the Promptable Gaze Target Estimation (PGE) task, a new end-to-end, concept-driven paradigm for gaze analysis. PGE conditions gaze prediction on flexible user text or visual prompts (e.g., "the boy in the red shirt" or "person in point [0.52, 0.48]") to identify a specific subject for gaze analysis. This approach integrates subject localization with gaze estimation, and eliminates the rigid dependency on intermediate analysis stages. We develop a scalable data engine to generate Gaze-Co (Gaze Estimation with Concepts), a dataset and benchmark of 120K high-quality, prompt-annotated image pairs. We also propose GazeAnywhere, the first model designed for PGE. GazeAnywhere uses a transformer-based detector to fuse features from frozen encoders and simultaneously solves subject localization, in/out-of-frame presence, and gaze target heatmap estimation. GazeAnywhere achieves state-of-the-art performance on multiple PGE benchmarks, setting a strong baseline for this new problem even on a difficult out-of-domain, real-world clinical dataset. GazeAnywhere is open-sourced in github.com/IrohXu/GazeAnywhere.

摘要:估計野外圖像中的人類注視目標是一項重要且艱巨的任務。現有的方法主要採用脆弱的多階段管道,這需要明確的輸入,例如頭部邊界框和人體姿勢,以識別注視分析的主體。因此,檢測錯誤可能會級聯並導致失敗。此外,這些先前的工作缺乏通過自然語言提示來指定注視分析任務的靈活性,這種方法已被證明在其他圖像分析任務中具有顯著的便利性和可擴展性。為了克服這些限制,我們引入了可提示的注視目標估計(PGE)任務,這是一種新的端到端、以概念為驅動的注視分析範式。PGE基於靈活的用戶文本或視覺提示(例如,“穿紅色襯衫的男孩”或“位於點[0.52, 0.48]的人”)來識別特定的注視分析主體。這種方法將主體定位與注視估計相結合,並消除了對中間分析階段的僵硬依賴。我們開發了一個可擴展的數據引擎來生成Gaze-Co(概念驅動的注視估計),這是一個包含120K高質量、提示標註圖像對的數據集和基準。我們還提出了GazeAnywhere,這是第一個為PGE設計的模型。GazeAnywhere使用基於Transformer的檢測器來融合來自凍結編碼器的特徵,並同時解決主體定位、框內/框外存在性和注視目標熱圖估計。GazeAnywhere在多個PGE基準上達到了最先進的性能,即使在困難的域外現實臨床數據集上也設置了這一新問題的強基線。GazeAnywhere已在github.com/IrohXu/GazeAnywhere上開源。

Dual-Domain Cross-Modal Decoding for Clinical Text-Guided Medical Image Segmentation

2608.11335v1 by Md Maklachur Rahman, Tracy Hammond

Clinical text can narrow down what to segment, but recent text-guided designs emphasize spatial alignment while overlooking frequency content that governs texture and boundaries. We propose Dual-Domain Cross-Modal Decoding (DD-CMD) for clinical text-guided pulmonary infection segmentation, integrating two complementary forms of language guidance during decoding. In the spatial domain, Text-Guided Spatial Cross-Attention (TGSA) aligns multi-scale visual tokens with text semantics and updates features through gated residual fusion. In the frequency domain, Spectral-Text Adaptive Modulation (STAM) applies a 2D DCT to compute learnable band-energy statistics and predicts text-conditioned FiLM parameters to recalibrate decoder channels for frequency-aware decoding. DD-CMD embeds TGSA and STAM into a coarse-to-fine decoder (7x7 to 56x56) and restores full-resolution masks using a lightweight two-stage refinement module. Experiments on QaTa-COV19 and MosMedData+ show that DD-CMD achieves 91.46% Dice / 84.26% mIoU and 81.95% Dice / 69.42% mIoU, respectively, with average gains of +1.96 Dice and +2.67 mIoU over the strongest prior baselines. Code: https://github.com/maklachur/DD-CMD.

摘要:臨床文本可以縮小分割範圍,但最近的文本引導設計強調空間對齊,同時忽略了控制紋理和邊界的頻率內容。我們提出了雙域跨模態解碼(DD-CMD)用於臨床文本引導的肺部感染分割,在解碼過程中整合兩種互補的語言引導形式。在空間域中,文本引導的空間跨注意力(TGSA)將多尺度視覺標記與文本語義對齊,並通過門控殘差融合更新特徵。在頻率域中,光譜文本自適應調製(STAM)應用2D DCT來計算可學習的帶能量統計,並預測文本條件的FiLM參數,以重新校準解碼器通道以進行頻率感知解碼。DD-CMD將TGSA和STAM嵌入到一個粗到細的解碼器(從7x7到56x56),並使用輕量級的兩階段精煉模塊恢復全分辨率的掩膜。在QaTa-COV19和MosMedData+上的實驗顯示,DD-CMD分別達到91.46%的Dice / 84.26%的mIoU和81.95%的Dice / 69.42%的mIoU,平均增益為+1.96 Dice和+2.67 mIoU,相較於最強的先前基準。代碼:https://github.com/maklachur/DD-CMD。

3D Weighted Geometric Graph Neural Networks for Sheep Facial Pain Assessment

2608.11050v1 by Alam Noor, Luis Almeida, Mohamed Daoudi

Deep learning systems perform mainly within the 2D for a single image domain and take the face as a single-dimension representation, losing sight of the 3D anatomy of sheep and cross-landmark spatial relationships that are intrinsic to the clinically proven Sheep Pain Facial Expression Scale (SPFES). This paper presents the \textbf{3D Sheep Pain Facial Expression System (3D-SPFES)}, a novel, monocular depth-aware geometric graph neural network system that integrates each SPFES facial landmark, such as the ears, eyes, and nose, into 3D Euclidean space estimated from a single RGB camera by using VideoDepthAnything, thus preventing the need for specialized depth hardware. Each landmark node includes a feature vector containing its 3D spatial coordinates, estimated surface normal, and facial attribute class embedding. Edges linked to nodes are assigned weights based on an aggregate metric that combines both Euclidean distance and surface co-planarity in a 3D space. A Weighted Geometric Graph Neural Network (WG-GNN) studies this graph using $\mathcal{K} = 3$ geometry-aware message-passing layers enhanced by a scaled dot-product attention method that selectively enhances anatomically relevant inter-landmark messages. The resultant node embeddings are combined into $\mathcal{O} = 3$ pain-level clusters and integrated into a Normalized Pain Score (NPS) within the range of $[0, 100%]$ a confidence-weighted, SPFES-derived scoring method.

摘要:深度學習系統主要在單一影像的2D領域內運作,將面部視為單一維度的表示,忽略了羊的3D解剖結構和與臨床證明的羊痛面部表情量表(SPFES)固有的交叉標記空間關係。本文提出了\textbf{3D羊痛面部表情系統(3D-SPFES)},這是一種新穎的單目深度感知幾何圖神經網絡系統,將每個SPFES面部標記(如耳朵、眼睛和鼻子)整合到從單一RGB相機估算的3D歐幾里得空間中,藉此避免了專用深度硬體的需求。每個標記節點包含一個特徵向量,其中包含其3D空間坐標、估算的表面法向量和面部屬性類別嵌入。連接到節點的邊根據一個綜合指標分配權重,該指標結合了3D空間中的歐幾里得距離和表面共平面性。加權幾何圖神經網絡(WG-GNN)使用$\mathcal{K} = 3$幾何感知消息傳遞層來研究這個圖,並通過縮放的點積注意力方法增強與解剖相關的標記間消息。結果節點嵌入被合併為$\mathcal{O} = 3$疼痛級別集群,並整合到範圍為$[0, 100%]$的標準化疼痛分數(NPS)中,這是一種基於信心加權的、源自SPFES的評分方法。

CARE: Confidence-Aware Reasoning for Reliable Medical VQA

2608.10964v1 by Yuetian Du, Yucheng Wang, Zhenyuan Chen, Luyuan Chen, Rongyu Zhang, Jinjian Zhang, Wei Zhou, Zhijie Xu, Ming Kong, Zhan Zhou, Jie Liu, Qiang Zhu

Reinforcement Fine-Tuning (RFT) has enabled medical Multimodal Large Language Models (MLLMs) to produce Chain-of-Thought (CoT) reasoning for visual question answering, yet these models suffer from $\textit{confidence miscalibration}$---a systematic gap between expressed certainty and actual diagnostic accuracy that undermines clinical trust. We propose $\textbf{CARE}$, a $\textbf{C}$onfidence-$\textbf{A}$ware medical $\textbf{RE}$asoning framework that jointly optimizes accuracy and calibration through a dual-stage pipeline. First, a scalable Medical-CoT synthesis provides structured cold-start data for Supervised Fine-Tuning. Second, Group Relative Policy Optimization (GRPO) with a novel $\textbf{Confidence-Aware Reward (CAR)}$ mechanism ties the model's confidence to diagnostic correctness within the reward signal. Across three Medical VQA benchmarks, $\textbf{CARE}$ achieves the highest diagnostic accuracy while obtaining the lowest Expected Calibration Error and Hallucination Rate, establishing a foundation for trustworthy clinical decision support. Our code is available at https://github.com/anotherbricki/CARE.

摘要:強化微調(RFT)使醫療多模態大型語言模型(MLLMs)能夠為視覺問題回答產生思維鏈(CoT)推理,然而這些模型存在著$\textit{信心錯誤校準}$的問題——表達的確定性與實際診斷準確性之間的系統性差距,這削弱了臨床信任。我們提出了$\textbf{CARE}$,一個$\textbf{C}$onfidence-$\textbf{A}$ware醫療$\textbf{RE}$asoning框架,通過雙階段管道共同優化準確性和校準。首先,一個可擴展的Medical-CoT合成提供結構化的冷啟動數據以進行監督微調。其次,帶有新穎的$\textbf{Confidence-Aware Reward (CAR)}$機制的群體相對策略優化(GRPO)將模型的信心與獎勵信號中的診斷正確性相聯繫。在三個醫療VQA基準中,$\textbf{CARE}$實現了最高的診斷準確性,同時獲得了最低的期望校準誤差和幻覺率,為可信的臨床決策支持奠定了基礎。我們的代碼可在https://github.com/anotherbricki/CARE獲得。

ComBodied Agents: a New Paradigm of Human-Centric Agentic AI

2608.10915v2 by Qianggang Ding, Xingyao Wang, Rui Feng, Zhibin Wang, Feixiang Yao, Kelong Mao, Hao Sun, Zhiyao Luo, Jiankai Tang, Lei Li, Jiadong Guo, Minheng Ni, Weicong Lin, Chenxi Yang, Hongxiang Gao, Zhenghua Chen, Yang Bai, Min Wu, Jun Cheng, Huazhu Fu, Dacheng Tao, Bang Liu

After an older adult misses a medication dose, a software agent can send another reminder and an embodied agent can bring the medication. Yet neither explains whether the person forgot, is confused, has side effects, or deliberately refused, nor what support is appropriate. This reveals a structural gap in Agentic AI: Digital Agents primarily transform software states, while Embodied Agents transform physical states; neither makes a person's evolving state and agency the primary object of modeling, intervention, and evaluation. We introduce Combodied Agents, a human-centered paradigm that perceives, models, predicts, and supports individual human-state trajectories over time, using software tools, sensors, wearables, robots, and human services as action channels rather than end goals. We unify fragmented capabilities across personal assistants, health agents, AI companions, and adaptive human--AI systems into a closed loop: event-based multimodal perception reconstructs meaningful personal events; longitudinal, correctable memory provides temporal context; Personal World Models estimate future personal states and outcomes under alternative decisions and interventions; and an admissible intervention policy selects proportionate support under consent, uncertainty, safety, reversibility, and user control. Feedback from the person and environment updates the loop. Rather than requiring an exhaustive Human Digital Twin, the framework uses purpose-bounded, uncertainty-aware, user-correctable representations. We organize the design space by human-state targets, relational contexts, and agent roles, and propose scenario-centered evaluation, agency-preservation metrics, benchmark requirements, edge-native personal models, and governance directions. Combodied Agents shift Agentic AI from external task completion toward sustained human benefit.

摘要:在年長者錯過藥物劑量後,軟體代理可以發送另一個提醒,而具身代理可以帶來藥物。然而,這兩者都沒有解釋該人是否忘記、感到困惑、出現副作用或故意拒絕,也沒有說明什麼樣的支持是合適的。這揭示了代理人工智能中的結構性缺口:數位代理主要轉換軟體狀態,而具身代理則轉換物理狀態;兩者都未將個體不斷演變的狀態和能動性作為建模、干預和評估的主要對象。我們引入了具身代理(Combodied Agents),這是一種以人為中心的範式,能夠隨著時間的推移感知、建模、預測和支持個體的人類狀態軌跡,使用軟體工具、感測器、可穿戴設備、機器人和人類服務作為行動渠道,而非最終目標。我們將個人助理、健康代理、人工智慧伴侶和自適應人類-人工智慧系統的零散能力統一成一個閉環:基於事件的多模態感知重建有意義的個人事件;長期的、可修正的記憶提供時間背景;個人世界模型在不同的決策和干預下估計未來的個人狀態和結果;可接受的干預政策在同意、不確定性、安全性、可逆性和用戶控制下選擇相稱的支持。來自個人和環境的反饋更新這個循環。該框架不需要全面的人類數位雙胞胎,而是使用目的有限、具不確定性意識和用戶可修正的表徵。我們根據人類狀態目標、關係背景和代理角色來組織設計空間,並提出以情境為中心的評估、能動性保護指標、基準要求、邊緣原生個人模型和治理方向。具身代理將代理人工智能的重心從外部任務完成轉向持續的人類利益。

MIRA: Medical Image Reflection for Agentic Diagnosis

2608.10827v1 by Shengzhi Wang, Jun Yang, Kai Wu, Xiaozhong Ji, Yiwen Ye, Ziyang Chen, Mingliang Xiong, Wen Fang, Mingqing Liu, Mengyuan Xu, Miaoxuan Shan, Caiyan Liu, Bin He, Qingwen Liu

Medical visual agents can use tools to inspect images and retrieve external knowledge, but indiscriminate tool use may introduce noisy or misleading evidence. Reliable diagnosis therefore requires not only acquiring additional observations, but also verifying whether tool actions are necessary and whether the resulting evidence supports the current hypothesis. We introduce MIRA (Medical Image Reflection for Agentic Diagnosis), a medical visual diagnostic framework for autonomous evidence search and reflective verification. MIRA dynamically invokes image-processing operations, including zooming, grounding, pointing, rotation, and measurement, as well as web search, while evaluating the relevance and consistency of the acquired evidence. We develop MIRA through a two-stage training strategy. First, a tool-augmented Monte Carlo Tree Search data engine explores diverse diagnostic hypotheses and jointly verifies visual grounding accuracy and semantic consistency to construct supervised fine-tuning trajectories. Second, reinforcement learning further improves decision-making through online reflective principle evolution: failure cases are distilled into candidate principles, and only principles that improve held-out rollout rewards are retained. Across nine medical visual reasoning benchmarks, MIRA achieves an average score of 64.73, improving its Qwen3-VL-8B backbone by 7.44 points. It also increases useful tool-use judgments from 56.2% to 73.8% and reduces harmful judgments from 8.9% to 1.6%. Qualitative analyses show that MIRA can re-examine evidence, correct premature conclusions, and adapt its tool-use strategy. Project page: https://MIRA-VL.github.io/

摘要:醫學視覺代理可以使用工具來檢查圖像並檢索外部知識,但不加區別的工具使用可能會引入噪音或誤導性的證據。因此,可靠的診斷不僅需要獲取額外的觀察結果,還需要驗證工具行動是否必要,以及所產生的證據是否支持當前的假設。我們介紹了 MIRA(醫學影像反思代理診斷),這是一個用於自主證據搜索和反思驗證的醫學視覺診斷框架。MIRA 動態調用圖像處理操作,包括縮放、定位、指向、旋轉和測量,以及網絡搜索,同時評估所獲得證據的相關性和一致性。我們通過兩階段的訓練策略來開發 MIRA。首先,增強工具的蒙特卡羅樹搜索數據引擎探索多樣的診斷假設,並共同驗證視覺定位的準確性和語義一致性,以構建監督的微調軌跡。其次,強化學習進一步通過在線反思原則演變改善決策:失敗案例被提煉為候選原則,只有那些改善保留的展開獎勵的原則才會被保留。在九個醫學視覺推理基準中,MIRA 的平均分數為 64.73,將其 Qwen3-VL-8B 主幹提高了 7.44 分。它還將有用的工具使用判斷從 56.2% 提高到 73.8%,並將有害判斷從 8.9% 降低到 1.6%。定性分析顯示,MIRA 可以重新檢查證據、修正過早的結論,並調整其工具使用策略。項目頁面: https://MIRA-VL.github.io/

DuplexWorld: Can voice agents help you get through the day?

2608.10716v1 by Aryan Vijay Bhosale, Harshit Rajgarhia, Akhil Pothanapalli, Asif Shaik, Abhishek Mukherji, Dinesh Manocha

Speech-to-speech (S2S) voice agents are increasingly being incorporated into enterprise for customer care and as daily companions for consumers owing to the ease of the conversational modality over text. However, existing benchmarks fail to holistically evaluate voice agents along axes that really matter and are shaped as tests of agentic tool calling against a database. We believe they fail to adequately account for the diversity of conversational dialogue that mundane activities introduce and further, never test how faithfully an agent can assist on tasks that move beyond database manipulation. To tackle this DuplexWorld introduces six worlds where voice agents are especially useful: banking, insurance, travel, healthcare and logistics, and Pathfinding. Agents are evaluated on eleven different types of conversations across 156 scenarios (350+ hours of conversation), each testing conversational and analytical capability to varying degrees. Through extensive evaluation comprising agentic, conversational and speech-naturalness metrics, we show that even the best voice agents leave substantial room for improvement on all 3 axes (Pass@1: 0.490, turn-taking: 0.653, DNSMOS: 3.378). We perform extensive analysis on agentic v conversational performance, world- and conversation type-wise performance, failure modes exploring the explore v exploit lens for Pathfinding conversations and voice agent reliability over all six worlds.

摘要:語音對語音 (S2S) 語音代理越來越多地被企業納入客戶服務和作為消費者的日常伴侶,這是因為對話模式相較於文本更加方便。然而,現有的基準未能從真正重要的維度全面評估語音代理,並且其形式是針對數據庫進行的代理工具呼叫測試。我們認為,它們未能充分考慮日常活動所引入的對話多樣性,並且從未測試代理在超越數據庫操作的任務中能夠多麼忠實地提供協助。為了解決這個問題,DuplexWorld 引入了六個語音代理特別有用的世界:銀行、保險、旅行、醫療保健和物流,以及路徑尋找。代理在156個場景中對11種不同類型的對話進行評估(超過350小時的對話),每個場景測試對話和分析能力的不同程度。通過包括代理性、對話性和語音自然度指標的廣泛評估,我們顯示即使是最好的語音代理在這三個維度上仍有相當大的改進空間(Pass@1: 0.490,輪流發言: 0.653,DNSMOS: 3.378)。我們對代理性與對話性表現、世界及對話類型的表現、失敗模式進行了廣泛分析,探索了路徑尋找對話的探索與利用視角,以及所有六個世界中語音代理的可靠性。

MedUP: Awakening Unified Understanding and Perception in Medical Vision-Language Models

2608.10635v1 by Yuan Wang, Hualiang Wang, Yixin Chen, Songtao Jiang, Shujian Gao, Jiaming Lin, Siming Fu, Jian Wu, Zuozhu Liu

Medical Vision-Language Models (Med-VLMs) excel at verbalizing visual content, yet precise visual perception, segmentation, and grounding remain challenging. Existing approaches either verbalize regions as coordinate strings or rely on external modules that decouple perception from understanding, creating representation gaps for region-language alignment. We present MedUP, a Med-VLM that natively unifies perception and understanding within a shared token space. At its core lies UniMedTok, a region tokenizer that encodes masks as discrete tokens in the LLM vocabulary, enabling the model to seamlessly interleave mask tokens with text. We curate UniMed-Train, a 1.84M-instance corpus spanning text-guided segmentation, region-grounded understanding, medical VQA and CoT-based segmentation, and introduce UniMed-Bench for unified evaluation. Extensive experiments show that MedUP outperforms native, agentic, and dual-decoder Med-VLMs across all tasks while remaining competitive with specialist segmentors, demonstrating the strong potential of unified understanding and perception modeling.

摘要:醫療視覺語言模型(Med-VLMs)在將視覺內容口頭表達方面表現出色,但精確的視覺感知、分割和定位仍然具有挑戰性。現有的方法要麼將區域表達為坐標字符串,要麼依賴於將感知與理解解耦的外部模塊,這在區域與語言對齊中創造了表示差距。我們提出了 MedUP,一種在共享標記空間內本地統一感知和理解的 Med-VLM。其核心是 UniMedTok,一種將掩膜編碼為 LLM 詞彙中離散標記的區域標記器,使模型能夠無縫地將掩膜標記與文本交錯。我們整理了 UniMed-Train,一個包含 184 萬實例的語料庫,涵蓋文本引導的分割、區域基礎的理解、醫學 VQA 和基於 CoT 的分割,並介紹了 UniMed-Bench 以進行統一評估。廣泛的實驗表明,MedUP 在所有任務中超越了原生、主動和雙解碼器 Med-VLM,並在與專業分割器的競爭中保持競爭力,展示了統一理解和感知建模的強大潛力。

Agentic Instruction Data Selection: Let DataMaster Interpret Your Intent

2608.10579v1 by Fanqi Zhou, Qiaosheng Chen, Zixian Huang, Gong Cheng

Although existing instruction data selection methods have introduced various metrics, the inherent complexity of real-world datasets makes it impractical for any single metric to generalize across all scenarios. Developers are thus often forced to manually inspect data and craft heuristic rules for each new application---a tedious and error-prone process. In this paper, we propose a paradigm shift from manual configuration to automated orchestration via the Instruction Data Selection Agent (DataMaster), which interprets user intent and autonomously composes optimal selection strategies. By allowing users to specify data needs through natural language descriptions, DataMaster simplifies data curation and removes the burden of manual strategy design. Extensive experiments across the math, medical, and code domains show that DataMaster outperforms static baselines in most settings and surpasses full-pool training in a substantial number of cases. The implementation of DataMaster and the scripts needed to reproduce the reported pipeline are publicly available at https://github.com/nju-websoft/DataMaster.

摘要:儘管現有的指令數據選擇方法引入了各種指標,但現實世界數據集的固有複雜性使得任何單一指標在所有場景中都難以通用。因此,開發者通常被迫手動檢查數據並為每個新應用編寫啟發式規則——這是一個繁瑣且容易出錯的過程。在本文中,我們提出了一種從手動配置轉向自動編排的範式轉變,通過指令數據選擇代理(DataMaster),該代理解釋用戶意圖並自主組合最佳選擇策略。通過允許用戶通過自然語言描述來指定數據需求,DataMaster 簡化了數據策展並消除了手動策略設計的負擔。在數學、醫學和代碼領域的廣泛實驗表明,DataMaster 在大多數設置中超越了靜態基準,並在相當多的案例中超越了全池訓練。DataMaster 的實現及重現報告流程所需的腳本可在 https://github.com/nju-websoft/DataMaster 上公開獲得。

Reinforcement Learning-Based Laser Cutting Machine Parameter Optimization

2608.10549v1 by Khanh Quan Pham, Majid Kundroo, Geunwoo Ban, Seongho Bae, Taehong Kim

Achieving high accuracy in laser-based cutting of optical films requires careful tuning of parameters such as focal length and laser power beam, adjusted according to the specific properties of each film type. Trial-and-error based traditional methods are used to find the most suitable cutting parameters for various films, but they are slow and inaccurate. To address this issue, this paper presents the Reinforcement Learning for Laser Cutting (RL$^{2}$C) algorithm, which uses Q-learning with an epsilon-greedy policy to dynamically optimize cutting parameters, significantly reducing taper size and film wastage. Additionally, RL$^{2}$C incorporates a dynamic environment space adaptability mechanism to allow it to adapt to new states encountered during the learning process over multiple batches of experiments. Experimental results demonstrate that RL$^{2}$C requires fewer steps and less time to find optimal cutting parameters compared to various RL-based optimization methods. Specifically, RL$^{2}$C reduces the number of optimization steps by up to 12.5\% and processing time by up to 81.8\% compared to existing methods. This study demonstrates the potential of RL in industrial laser-cutting processes by improving cut quality, reducing time and film wastage, and minimizing manual interventions.

摘要:達成激光切割光學薄膜的高精度需要仔細調整參數,例如焦距和激光功率光束,根據每種薄膜類型的特定特性進行調整。傳統的試錯方法用於尋找各種薄膜的最合適切割參數,但這些方法速度慢且不準確。為了解決這個問題,本文提出了激光切割強化學習(RL$^{2}$C)算法,該算法使用帶有epsilon-greedy策略的Q-learning來動態優化切割參數,顯著減少錐度大小和薄膜浪費。此外,RL$^{2}$C還結合了一個動態環境空間適應機制,使其能夠在多批次實驗的學習過程中適應遇到的新狀態。實驗結果表明,與各種基於RL的優化方法相比,RL$^{2}$C需要更少的步驟和更少的時間來找到最佳切割參數。具體而言,與現有方法相比,RL$^{2}$C將優化步驟數量減少了多達12.5\%,處理時間減少了多達81.8\%。這項研究通過提高切割質量、減少時間和薄膜浪費,以及最小化人工干預,展示了RL在工業激光切割過程中的潛力。

Unlocking the Power of Medical Tabular Data via Semantic-Aware Multimodal Pre-training

2608.10522v1 by Yingsheng Liu, Haiming Li, Jingmin Zhu, Jiajun Sun, Victoria Mar, Monika Janda, H. Peter Soyer, Zongyuan Ge, Zhen Yu

While vision-language models dominate medical representation learning, unstructured text lacks the dense, quantitative diagnostic phenotypes inherent in structured clinical tables. However, existing multimodal pre-training methods underutilize this potential due to semantic-agnostic designs that treat tabular inputs as flat vectors and employ unstable continuous regression objectives. To overcome this, we propose a novel semantic-aware framework explicitly modeling the intrinsic two-dimensional structure of tabular data. First, addressing the inter-feature hierarchy of varying diagnostic importance, we introduce Importance-Aware Adaptive Masking to construct a label-free curriculum prioritizing salient features. Second, addressing the intra-feature continuity-discreteness duality, we propose a Soft-Label Discretized Module that replaces unstable numerical regression with stable distribution matching, thereby mathematically preserving ordinal relationships. Extensive experiments across large-scale dermatology (SLICE-3D, HOP) and ophthalmology (EyePACS) datasets establish a new state-of-the-art (SOTA), demonstrating exceptional robustness and cross-domain generalizability.

摘要:雖然視覺-語言模型主導了醫學表徵學習,但非結構化文本缺乏結構化臨床表格中固有的密集、定量診斷表型。然而,現有的多模態預訓練方法因為語義無關的設計而未能充分利用這一潛力,這些設計將表格輸入視為平坦的向量並採用不穩定的連續回歸目標。為了克服這一問題,我們提出了一種新穎的語義感知框架,明確建模表格數據的內在二維結構。首先,針對不同診斷重要性的特徵層次,我們引入了重要性感知自適應掩碼,構建了一個無標籤的課程,優先考慮顯著特徵。其次,針對特徵內部的連續性-離散性二元性,我們提出了一個軟標籤離散模塊,將不穩定的數值回歸替換為穩定的分佈匹配,從而在數學上保持序關係。在大規模皮膚科(SLICE-3D,HOP)和眼科(EyePACS)數據集上的廣泛實驗確立了新的最先進技術(SOTA),顯示出卓越的穩健性和跨領域的泛化能力。

RadFusion: Towards Threshold-Controllable Radiology Report Generation

2608.10505v1 by Ying Jin, Noel C. F. Codella, John Corring, Mu Wei, Dinei Florencio, Eric Horvitz

Automated radiology report generation is advancing rapidly in response to the shortage of radiologists, yet unlike a perception model, existing generation models offer no control over the sensitivity-specificity trade-off of their diagnostic content. Such control is essential because clinical scenarios diverge: emergency triage prioritizes sensitivity to reduce missed findings, whereas confirmatory interpretation emphasizes specificity to limit unnecessary interventions. A single fixed report can neither adapt to these scenarios nor support the ROC-based validation widely expected for regulatory clearance. We introduce RadFusion, a framework that equips report generation with threshold controllability. Our method fuses a multi-label classifier, which provides per-disease confidence scores, with a VQA-based report generator, which describes medical findings in detail; an LLM then rewrites the report so that its stated diagnoses follow the classifier's decisions at the selected threshold while staying grounded in the generator's descriptions. On MIMIC-CXR, the performance of RadFusion conforms to the classifier's ROC curve: sweeping the threshold and mapping the reports back to class labels reproduces the classifier's validated ROC performance. This conformance makes generated reports quantitatively evaluable through ROC analysis, strengthening the case for regulatory clearance, and enables operating-point selection that matches report behavior to clinical context. Moreover, combining the two model types improves diagnostic accuracy over uncontrolled generation: sensitivity increases by 6.9% at matched specificity, and specificity by 20.7% at matched sensitivity. These results show that RadFusion makes report generation clinically adaptable, quantitatively verifiable, and diagnostically more reliable.

摘要:自動化放射科報告生成正迅速發展,以應對放射科醫師的短缺,然而與感知模型不同,現有的生成模型無法控制其診斷內容的敏感性-特異性權衡。這種控制是至關重要的,因為臨床情境各異:緊急分診優先考慮敏感性以減少漏診,而確認性解釋則強調特異性以限制不必要的干預。單一的固定報告既無法適應這些情境,也無法支持廣泛期待用於監管批准的基於ROC的驗證。我們介紹了RadFusion,一個為報告生成提供閾值可控性的框架。我們的方法融合了一個多標籤分類器,該分類器提供每種疾病的信心分數,與一個基於VQA的報告生成器,該生成器詳細描述醫療發現;然後一個LLM重寫報告,使其所述的診斷遵循分類器在所選閾值下的決策,同時基於生成器的描述。 在MIMIC-CXR上,RadFusion的性能符合分類器的ROC曲線:調整閾值並將報告映射回類別標籤再現了分類器的經過驗證的ROC性能。這種一致性使得生成的報告可以通過ROC分析進行定量評估,增強了監管批准的案例,並使得操作點選擇能夠將報告行為與臨床情境相匹配。此外,結合這兩種模型類型提高了診斷準確性,相同特異性下敏感性提高了6.9%,相同敏感性下特異性提高了20.7%。這些結果顯示RadFusion使報告生成在臨床上適應性更強、定量可驗證且診斷上更可靠。

RLMOpt: Adaptive Prompt Optimization via Recursive Language Models

2608.10471v1 by Subhash Bangalore Satheesha, Nirvik Pande, Deepthi Duddempudi, Bharath Dandala

Prompt optimizers automate the search for prompts that improve language-model performance, but existing methods rely on a predefined optimization procedure: the algorithm determines which candidates to explore and how the search progresses, while the language model generates or refines prompt proposals. We introduce RLMOpt, a prompt optimizer that makes the search policy itself language-model-driven through a recursive language model (RLM). The RLM agent operates over a tool-based environment, inspecting task information, analyzing failures, generating candidates, allocating evaluation budget, and deciding when to stop. A deterministic harness complements the agent by enforcing objective scoring, Pareto-based selection, and regression constraints. We evaluate RLMOpt across four benchmarks spanning structured clinical information extraction (Chia), multi-hop question answering (HotpotQA), verifiable instruction following (IFBench-2025), and multi-turn tool-calling agents (BFCL). In a matched comparison at a single seed, RLMOpt obtains the best held-out score on all four benchmarks and leads the four-task mean (0.610 against 0.589 for GEPA). Repeating each benchmark across seeds yields 11 matched benchmark-seed comparisons, in which RLMOpt outperforms GEPA in 9 cases. Across all 11 runs, it never produced a prompt that underperformed its seed, whereas GEPA fell below its starting point twice. It is also more efficient, achieving these results with fewer search rollouts while producing prompts that are 27-79% the size of those produced by GEPA. Our results further show that optimization gains are determined primarily by the headroom available in the seed prompt, rather than by the search budget. Efficient optimization therefore depends on reaching the available headroom reliably and with minimal search

摘要:提示優化器自動化尋找能改善語言模型性能的提示,但現有的方法依賴於預定義的優化程序:算法決定了要探索哪些候選者以及搜索的進展方式,而語言模型則生成或完善提示提案。我們介紹了 RLMOpt,一個通過遞歸語言模型(RLM)使搜索策略本身由語言模型驅動的提示優化器。RLM 代理在基於工具的環境中運作,檢查任務信息、分析失敗、生成候選者、分配評估預算並決定何時停止。一個確定性的工具補充了代理,強制執行客觀評分、基於帕累托的選擇和回歸約束。 我們在四個基準上評估 RLMOpt,涵蓋結構化臨床信息提取(Chia)、多跳問題回答(HotpotQA)、可驗證的指令遵循(IFBench-2025)和多輪工具調用代理(BFCL)。在單一種子下的匹配比較中,RLMOpt 在所有四個基準上獲得了最佳的保留分數,並在四項任務的平均分(0.610 對 0.589 的 GEPA)中領先。重複每個基準跨種子產生 11 個匹配的基準-種子比較,其中 RLMOpt 在 9 個案例中超越了 GEPA。在所有 11 次運行中,它從未產生低於其種子的提示,而 GEPA 則有兩次低於其起始點。它的效率也更高,以更少的搜索展開達成這些結果,同時生成的提示大小僅為 GEPA 的 27-79%。 我們的結果進一步顯示,優化增益主要取決於種提示中可用的潛力,而不是搜索預算。因此,高效的優化依賴於可靠地達到可用的潛力並以最小的搜索進行。

Expert-Guided g-computation with Large Language Models for Estimating Causal Effects on Timings: Applications to Hospital Quality Improvement

2608.10339v1 by Patrick Vossler, Jialin Ouyang, F. Richard Guo, Anran Huang, Ali Shojaie, Lucas Zier, Fan Xia, Jean Feng

Hospital quality improvement (QI) programs routinely face multiple candidate interventions to optimize hospital flow, but existing methods struggle to estimate and rank the causal effects of such interventions. This work focuses on one of the most standard hospital metrics, the average length of stay (LOS), and its causal estimand, the average time saved. To characterize this causal effect, qualitative approaches rely on expert judgment to map patient trajectories, making them susceptible to cognitive biases; quantitative approaches rely on data-driven models, which fail when interventions are hypothetical with no historical data or have complex causal mechanisms that require clinical reasoning rather than data alone. We propose expert-guided g-computation, or egg-computation, which combines the complementary strengths of both approaches by connecting the Gantt charts commonly used to map patient trajectories with the causal DAG literature. We introduce a causal model over Gantt charts and establish identification using a variant of g-computation that seeks expert input only for components unidentifiable from data. To make egg-computation practical, we develop an LLM-assisted pipeline that reliably scales up expert reasoning. In simulations, egg-computation outperforms conventional causal inference methods when patients have diverse causal structures and intervention mechanisms. In a study of eleven candidate QI interventions at an urban safety-net hospital, the LLM pipeline generated graphs and time-saving estimates highly concordant with those of human experts. Beyond healthcare, egg-computation is a broadly applicable framework for estimating the average time saved for candidate interventions whose causal mechanisms can be represented using Gantt charts.

摘要:醫院質量改善(QI)計劃經常面臨多種候選干預措施,以優化醫院流程,但現有方法在估計和排名這些干預措施的因果效應方面存在困難。這項工作專注於最標準的醫院指標之一,即平均住院天數(LOS),及其因果估計量,即平均節省的時間。為了表徵這一因果效應,定性方法依賴專家判斷來映射病人軌跡,使其容易受到認知偏見的影響;定量方法則依賴數據驅動模型,當干預措施是假設性的且沒有歷史數據,或具有需要臨床推理而非僅依賴數據的複雜因果機制時,這些模型會失效。我們提出了專家引導的g計算,或稱蛋計算,這種方法通過將常用於映射病人軌跡的甘特圖與因果DAG文獻相連接,結合了兩種方法的互補優勢。我們在甘特圖上引入了一個因果模型,並使用一種變體的g計算來建立識別,該變體僅尋求專家對數據無法識別的組件的輸入。為了使蛋計算實用,我們開發了一個LLM輔助的管道,可靠地擴展專家推理。在模擬中,當病人具有多樣的因果結構和干預機制時,蛋計算的表現超過了傳統的因果推斷方法。在對一所城市安全網醫院的十一個候選QI干預措施的研究中,LLM管道生成的圖形和節省時間的估計與人類專家的結果高度一致。除了醫療保健之外,蛋計算是一個廣泛適用的框架,用於估計可以用甘特圖表示的候選干預措施的平均節省時間。

Logit-Boundary Geometric Belief Interfaces and Sparse Sheaf-Enclave Protocols: A Self-Contained Substrate for Secure Network Electronic Health Record (EHR) Interoperability

2608.10300v2 by Alvin Spivey, Yu Huang

Electronic health-record interoperability is a boundary problem: legacy systems, generative models, terminology services, identity systems, and human reviewers may each expose rich internal states, while operational exchange requires a narrow shared interface of typed claims, bounded uncertainty, provenance, and explicit admission or abstention. This paper details a mathematical and engineering architecture for that interface. The organizing idea is the logit boundary: a discovery model may propose pre-threshold scores over a local categorical decision, but a deterministic judgment substrate decides whether the proposal is admissible, requires review, or must be quarantined before any Fast Healthcare Interoperability Resources (FHIR) transaction is constructed. The resulting Geometric Belief Interface (GBI) combines finite boundary semantics, local Dirichlet evidence, cellular-sheaf and mapping-cone diagnostics, advisory geometric audit charts, and a Decentralized Cryptographic Sheaf-Enclave (DCSE) protocol sketch for fail-closed deployment. The framework does not establish clinical truth, global representation alignment, or end-to-end safety; it defines certificate-producing checks at a model-to-system boundary. A companion frozen synthetic benchmark, GBI BoundaryBench v0.1, evaluated Qwen3-4B-Instruct-2507 on 256 held-out tasks across three evidence modes (768 canonical executions). All executions completed, but none produced an output accepted by the benchmark contract: 369 were rejected during safe parsing and 399 during schema validation, yielding zero coverage and deterministic quarantine. This empirical result is deliberately narrow - one 4B open-weight model under one frozen interface - and is reported as evidence about the admission boundary, not as a general claim about LLM capability or clinical safety. A Julia appendix verifies numerical certificates using standard libraries.

摘要:電子健康紀錄的互操作性是一個邊界問題:舊系統、生成模型、術語服務、身份系統和人工審查者各自可能暴露豐富的內部狀態,而操作性交換則需要一個狹窄的共享介面,該介面包括類型聲明、有限的不確定性、來源以及明確的接受或放棄。本文詳細描述了該介面的數學和工程架構。組織思想是邊界邏輯:發現模型可能會對局部類別決策提出閾值前的分數,但確定性判斷基底決定該提案是否可接受、是否需要審查,或在構建任何快速醫療互操作性資源(FHIR)交易之前必須被隔離。最終的幾何信念介面(GBI)結合了有限邊界語義、局部Dirichlet證據、細胞叢和映射錐診斷、諮詢幾何審計圖表,以及一個去中心化的加密叢集區(DCSE)協議草圖,以實現故障關閉部署。該框架並不建立臨床真實性、全球表示對齊或端到端安全性;它定義了在模型與系統邊界的證書生成檢查。一個伴隨的凍結合成基準,GBI BoundaryBench v0.1,評估了Qwen3-4B-Instruct-2507在三種證據模式下的256個保留任務(768個典型執行)。所有執行均已完成,但沒有任何輸出被基準合約接受:369個在安全解析過程中被拒絕,399個在模式驗證過程中被拒絕,導致零覆蓋和確定性隔離。這一實證結果故意狹窄——一個4B開放權重模型在一個凍結介面下——並被報告為關於接受邊界的證據,而不是關於LLM能力或臨床安全的一般主張。Julia附錄使用標準庫驗證數字證書。

Frozen Brain-MRI Foundation Models Are Site Fingerprints

2608.10295v1 by Saman Rahbar

Frozen foundation-model (FM) embeddings are increasingly used as off-the-shelf brain-MRI representations, on the assumption that they capture anatomy. We audit what they actually encode and find that acquisition site is a large, intrinsic component of the representation. Across two independent cohorts (ABIDE-I, ABIDE-II), three frozen 3-D encoders (brain-pretrained, CT-pretrained, and randomly initialized), and every network depth, site is linearly decodable at roughly 0.9 balanced accuracy at deep layers, exceeding the decodability of every clinical or demographic variable (sex, age, autism diagnosis) at every layer. The effect is intrinsic rather than learned: a randomly initialized encoder is already a ~0.9 site classifier on both cohorts and across three architecture families (Swin, ViT, ResNet), and site is decodable at ~0.95 directly from the raw downsampled image with no encoder, so the fingerprint reflects low-level image statistics that any encoder preserves rather than a product of pretraining. Residualizing measured population covariates leaves site decodability essentially unchanged, indicating an acquisition- rather than population-driven effect. A nonlinear probe matches the linear one, so the fingerprint is fully linearly accessible. The site subspace is removable post hoc by iterative null-space projection or ComBat (site decodability 0.94 -> 0.07/0.00), and is a site-attribution concern for shared or federated embeddings; but for dense segmentation this removal is not free, because site and anatomy occupy an entangled linear subspace (a matched-rank random-direction projection is Dice-neutral, whereas removing the site subspace is destructive). We recommend site-audited use of frozen brain-MRI FMs and release an open audit toolkit.

摘要:冷凍的基礎模型(FM)嵌入越來越多地被用作現成的腦部MRI表徵,假設它們能捕捉到解剖結構。我們審核它們實際編碼的內容,發現獲取地點是該表徵的一個重要內在組成部分。在兩個獨立的隊列(ABIDE-I,ABIDE-II)、三個冷凍的3D編碼器(腦部預訓練、CT預訓練和隨機初始化)以及每個網絡深度中,地點在深層的線性可解度約為0.9的平衡準確率,超過了每一層的臨床或人口變量(性別、年齡、自閉症診斷)的可解度。這一效應是內在的,而非學習得來的:隨機初始化的編碼器在兩個隊列和三個架構系列(Swin、ViT、ResNet)中已經是一個約0.9的地點分類器,並且地點可以直接從原始下採樣圖像中以約0.95的可解度解碼,而無需編碼器,因此指紋反映的是任何編碼器所保留的低級圖像統計,而不是預訓練的產物。對測量的人口協變量進行殘差化處理基本上不改變地點的可解度,這表明這是一種獲取驅動而非人口驅動的效應。非線性探針與線性探針相匹配,因此指紋是完全線性可訪問的。地點子空間可以通過迭代的零空間投影或ComBat事後移除(地點可解度0.94 -> 0.07/0.00),並且對於共享或聯邦嵌入來說,這是一個地點歸因的問題;但對於密集分割來說,這種移除並不是免費的,因為地點和解剖結構佔據了一個糾纏的線性子空間(匹配秩的隨機方向投影是Dice中性的,而移除地點子空間則是破壞性的)。我們建議對冷凍的腦部MRI FMs進行地點審核使用,並發布一個開放的審核工具包。

Locally Deployable Small Language Models for Emergency Department Decision Support: A Systematic Benchmark of Fine-Tuning Strategies

2608.10273v1 by Qingfeng Zhang, Yuanxiong Guo, Yanmin Gong

Deploying large language models (LLMs) for decision support in emergency departments (EDs) faces two major challenges: privacy risks of transmitting patient data to closed-source commercial LLMs and the lack of systematic evaluation of fine-tuning strategies for locally deployable open-source small language models (SLMs). We benchmarked eight open-source SLMs using zero-shot prompting, prefix tuning, Low-Rank Adaptation (LoRA), and full fine-tuning on three ED tasks: triage level prediction, specialist referral recommendation, and diagnosis prediction. Using 2,083 MIMIC-IV-ED cases and Claude Haiku 4.5 and Claude Sonnet 4.5 as baselines, we found that LoRA fine-tuned open-source SLMs outperform commercial baselines on triage level prediction and specialist referral recommendation, while diagnosis prediction remains challenging for open-source SLMs. Confusion matrix analysis further shows that fine-tuned open-source SLMs can detect highest-severity patients missed by the commercial baselines. These results demonstrate that locally deployable SLMs can achieve clinically competitive performance for ED decision support.

摘要:部署大型語言模型(LLMs)以支援急診部門(EDs)的決策面臨兩大挑戰:將病人數據傳輸到封閉源商業LLMs的隱私風險,以及缺乏對可本地部署的開源小型語言模型(SLMs)進行微調策略的系統評估。我們使用零樣本提示、前綴調整、低秩適應(LoRA)和完全微調,對八個開源SLMs在三個ED任務上進行基準測試:分診等級預測、專家轉診建議和診斷預測。使用2,083個MIMIC-IV-ED案例,並以Claude Haiku 4.5和Claude Sonnet 4.5作為基準,我們發現LoRA微調的開源SLMs在分診等級預測和專家轉診建議上優於商業基準,而診斷預測對於開源SLMs仍然具有挑戰性。混淆矩陣分析進一步顯示,微調後的開源SLMs能夠檢測到商業基準漏掉的高嚴重性病人。這些結果表明,可本地部署的SLMs能夠在急診決策支援中達到臨床競爭性能。

TAF-MED: Multi-Turn Safety Refusal Collapse in LLMs Under Declared Self-Treatment Intent

2608.10258v1 by Waleed Jamil, Raphael Schmitt

Large language models (LLMs) increasingly provide conversational health information that may influence treatment decisions, yet existing benchmarks do not isolate whether medication-safety boundaries persist across follow-ups after explicit self-treatment intent. We introduce TAF-MED, a physician-reviewed benchmark of 500 fixed three-turn scenarios, and evaluate eight LLMs across 4,000 conversations. A rubric-based automated judge labelled responses as SAFE, LEAKY, or UNSAFE, and two physicians independently annotated a model-balanced random subset of 400 conversations. We assessed unsafe guidance, collapse after a strictly SAFE initial response, and model-ranking stability. Overall, 71.6% of conversations contained an UNSAFE response, and 61.4% of those beginning with a strictly SAFE response later collapsed to UNSAFE; model-level collapse rates ranged from 24.4% to 96.2%. Four of 28 model pairs reversed order between initial unsafe and collapse rates. Automated labels achieved 94.3% agreement with the adjudicated physician reference ($κ= 0.895$). These findings show that first-turn safety is an incomplete proxy for conversational safety persistence and motivate evaluation across complete dialogue trajectories. We will release TAF-MED on Hugging Face to support reproducible research on multi-turn medical safety.

摘要:大型語言模型(LLMs)越來越多地提供可能影響治療決策的對話健康資訊,但現有的基準並未區分在明確的自我治療意圖後,藥物安全邊界是否持續存在。我們介紹了TAF-MED,一個經過醫生審核的500個固定三輪情境的基準,並評估了八個LLM在4000次對話中的表現。基於評分標準的自動評判將回應標記為安全(SAFE)、漏洩(LEAKY)或不安全(UNSAFE),兩位醫生獨立註釋了一個模型平衡的隨機子集,共400次對話。我們評估了不安全指導、在嚴格安全的初始回應後的崩潰情況,以及模型排名的穩定性。總體而言,71.6%的對話包含不安全的回應,而61.4%從嚴格安全的回應開始的對話後來崩潰為不安全;模型層級的崩潰率範圍從24.4%到96.2%。28對模型中有四對在初始不安全和崩潰率之間的順序相反。自動標籤與裁定的醫生參考達到94.3%的一致性($κ= 0.895$)。這些發現顯示,第一輪的安全性並不是對話安全持續性的完整代理,並促使對完整對話軌跡的評估。我們將在Hugging Face上發布TAF-MED,以支持多輪醫療安全的可重複研究。

Towards Expert-level Medical AI for Real-time Video Consultations

2608.09861v1 by Mahvish Nagda, Jihyeon Lee, Matthew Thompson, Chunjong Park, Tim Strother, Valentin Liévin, Roma Ruparel, Akshay Goel, Teya Bergamaschi, Suhana Bedi, Meet Shah, Pavel Dubov, Liviu Panait, Toshiyuki Fukuzawa, Sam Schmidgall, Craig Schiff, Joseph Xu, Aliya Rysbek, Yana Lunts, Jan Freyberg, Rebecca Hemengway, Sunny Virmani, David Racz, Carey Radebaugh, Joëlle Barral, Kavi Goel, Dale R. Webster, Katherine Chou, Avinatan Hassidim, Yossi Matias, James Manyika, Gregory Wayne, Tao Tu, Yun Liu, Ethan Goh, Christina Chen, Ryutaro Tanno, Po-Hsuan Cameron Chen, Mike Schaekermann, Anil Palepu

Audio-visual interaction is the standard for patient-physician consultations, enabling natural communication and effective assessment of illness through non-verbal cues. While text-based AI has shown promise, it discards essential perceptual dimensions and limits patients who cannot articulate symptoms in writing. Early efforts to extend medical AI to audio-visual interaction have demonstrated feasibility but not reached clinician-level performance. Here, we provide the first demonstration of expert-level AI in real-time clinical video consultations using AMIE (Articulate Medical Intelligence Explorer) in a video configuration. AMIE (Video) is a Gemini-based multi-agent system integrating low-latency dialogue, clinical reasoning, and real-time audio-visual perception. To guide development, we established a taxonomy and automated evaluations for clinical audio-visual cues in telehealth settings. In a randomized Objective Structured Clinical Examination (OSCE) study with 30 primary care physicians (PCPs), 15 patient actors and 100 clinical scenarios, we compared AMIE (Video), its text-only counterpart AMIE (Text), and PCPs consulting via video. Clinical evaluators rated AMIE (Video) on par or better than PCPs in history-taking, diagnosis, management, and physical observation and examination. Patient actors preferred AMIE's approach to assessing and explaining conditions, while PCPs were preferred for rapport and partnership building. In modality ablation, patient actors preferred AMIE (Video)'s interface over text chat for communicative effectiveness, convenience, and feeling understood. Limitations remain in fine anatomical precision, subtle affective nuances, and high-frequency movements. While further research is needed before real-world translation, these results mark an important milestone toward AI systems capable of augmenting care across the sensory complexity of clinical practice.

摘要:視聽互動是病人與醫生諮詢的標準,能夠通過非語言線索促進自然交流和有效評估疾病。雖然基於文本的人工智慧顯示出潛力,但它忽略了重要的感知維度,並限制了無法用書面表達症狀的病人。早期將醫療人工智慧擴展至視聽互動的努力已顯示出可行性,但未達到臨床醫生的表現水平。在此,我們提供了使用AMIE(Articulate Medical Intelligence Explorer)在視頻配置中進行實時臨床視頻諮詢的專家級人工智慧的首次示範。AMIE(視頻)是一個基於Gemini的多代理系統,整合了低延遲對話、臨床推理和實時視聽感知。為了指導開發,我們建立了一個分類法和自動評估,用於遠程醫療環境中的臨床視聽線索。在一項隨機的客觀結構化臨床考試(OSCE)研究中,涉及30位初級保健醫生(PCPs)、15位病人演員和100個臨床場景,我們比較了AMIE(視頻)、其文本專用對應AMIE(文本)以及通過視頻諮詢的PCPs。臨床評估者在病史採集、診斷、管理以及身體觀察和檢查方面評價AMIE(視頻)與PCPs相當或更好。病人演員更喜歡AMIE在評估和解釋病情方面的方法,而PCPs則在建立關係和夥伴關係方面更受青睞。在模態消融中,病人演員更喜歡AMIE(視頻)的界面而非文本聊天,因為其在交流有效性、便利性和被理解的感受上表現更佳。儘管在精細解剖精度、微妙的情感細微差別和高頻運動方面仍存在局限性,但在實際應用之前仍需進一步研究,這些結果標誌著朝著能夠增強臨床實踐中感官複雜性的護理的人工智慧系統邁出了重要的一步。

MedPixel: A Unified Pixel-Language Model for Medical Reasoning and Segmentation

2608.09818v1 by Haoyu Yang, Meixing Shi, Zengjie Chen, Haoran Sun, Haitao Leng, Xiaoming Shi, Yuxiang Cai, Yankai Jiang

Reliable medical image understanding requires models to connect clinical language and visual reasoning with pixel-level grounding. Yet medical vision-language models often lack precise localization, whereas medical segmenters typically rely on explicit target categories or precise spatial prompts. This divide is reinforced by a supervision mismatch: segmentation datasets provide precise masks but little language supervision, whereas medical vision-language data rarely pair language with dense spatial annotations. To address this gap, we present MedPixel, a unified medical pixel-language model built around a shared language--mask interface. To provide scalable supervision, we introduce MedPLG-440K, comprising approximately 440K pixel-language task samples constructed through a clinically motivated synthesis process without external LLM annotation. MedPixel is trained with joint multi-task supervised fine-tuning followed by Pixel-Level Preference Optimization, which uses ground-truth masks as offline verifiers to derive response preferences from mask quality. MedPixel supports a broad spectrum of tasks spanning explicit grounding, implicit reasoning, spatial interaction, grounded explanation, and medical VQA. Across this task spectrum, MedPixel achieves strong performance in both pixel-level prediction and response generation, together with effective zero-shot transfer to external grounding benchmarks and robustness to imperfect spatial prompts. Code and model checkpoints will be released at https://github.com/yhy-whu/Medpixel.

摘要:可靠的醫學影像理解需要模型將臨床語言和視覺推理與像素級基礎連接起來。然而,醫學視覺-語言模型通常缺乏精確的定位,而醫學分割模型則通常依賴於明確的目標類別或精確的空間提示。這種差距因監督不匹配而加強:分割數據集提供精確的掩膜,但語言監督很少,而醫學視覺-語言數據則很少將語言與密集的空間註釋配對。為了解決這一差距,我們提出了MedPixel,一個圍繞共享語言-掩膜接口構建的統一醫學像素-語言模型。為了提供可擴展的監督,我們引入了MedPLG-440K,該數據集由大約440K的像素-語言任務樣本組成,這些樣本是通過臨床驅動的合成過程構建的,並且沒有外部LLM註釋。MedPixel通過聯合多任務的監督微調進行訓練,隨後進行像素級偏好優化,該過程使用真實掩膜作為離線驗證器,從掩膜質量中推導響應偏好。MedPixel支持廣泛的任務,涵蓋明確的基礎、隱含推理、空間互動、基於基礎的解釋和醫學VQA。在這一任務範疇中,MedPixel在像素級預測和響應生成方面都達到了強勁的性能,並有效地實現了對外部基礎基準的零樣本轉移,以及對不完美空間提示的穩健性。代碼和模型檢查點將在 https://github.com/yhy-whu/Medpixel 發布。

AirFlow: Context Preserving and Multi-Rate State Modeling for Air Quality Forecasting

2608.09775v1 by Fan Yang, Nan Chen, Yijie Dong, Yuchen Zhang, Wei Zhang

Accurate air quality forecasting is essential for public health and urban environmental management, but remains challenging because pollutant channels differ in periodicity and distribution drift, while their concentration trajectories contain both multi-scale dependencies and rapid changes. Recent methods have improved spatial dependency learning and meteorological covariate modeling. However, pollutant channels are still passed through the same normalization rule and temporal backbone, using a shared latent representation for channel-specific distributions and changes at different rates. To address this limitation, we propose AirFlow, a pollutant-aware dual-stream framework that operates on station multivariate observations without additional graph propagation or predefined signal decomposition. Specifically, AirFlow designs two novel blocks: (1) a statistic-guided normalization routing mechanism that selects a normalization path for each pollutant according to its 24-hour autocorrelation and distribution drift; and (2) a hierarchical dual-stream state model that combines multi-scale state space propagation with learnable response coefficients, where gated bidirectional cross-attention exchanges information and adaptively fuses the resulting representations. Experiments on real-world data from multiple cities show that AirFlow achieves the best performance in 34 of 36 metrics comparisons, with reductions of up to 11.11% root mean square error over the state-of-the-art baseline. AirFlow also requires only 0.0483M parameters and 0.0215G FLOPs, achieving high forecasting accuracy with low computational overhead.

摘要:準確的空氣質量預測對於公共健康和城市環境管理至關重要,但仍然具有挑戰性,因為污染物通道在周期性和分佈漂移上存在差異,而它們的濃度軌跡則包含多尺度依賴性和快速變化。最近的方法改善了空間依賴性學習和氣象協變量建模。然而,污染物通道仍然通過相同的正規化規則和時間主幹,使用共享的潛在表示來處理特定通道的分佈和不同速率的變化。為了解決這一限制,我們提出了AirFlow,一種污染物感知的雙流框架,該框架在站點多變量觀測上運行,而不需要額外的圖形傳播或預定義的信號分解。具體而言,AirFlow設計了兩個新穎的模塊:(1)一個統計引導的正規化路由機制,根據每個污染物的24小時自相關和分佈漂移選擇正規化路徑;(2)一個層次雙流狀態模型,將多尺度狀態空間傳播與可學習的響應係數結合,其中門控雙向交叉注意力交換信息並自適應地融合結果表示。來自多個城市的實驗數據顯示,AirFlow在36個指標比較中有34個達到了最佳性能,並在最先進的基準上實現了最高11.11%的均方根誤差減少。AirFlow還僅需0.0483M參數和0.0215G FLOPs,以低計算開銷實現高預測準確性。

Physics-Informed Machine Learning in Prognostics and Health Management: A Systematic Literature Review

2608.10047v1 by Christopher Braun, Julian Raible, Marco F. Huber

In modern industry, keeping complex systems reliable, safe, and efficient hinges on Prognostics and Health Management (PHM). Machine Learning (ML) has largely driven advancements in diagnostics and prognostics, yet purely data-driven models face inherent limitations, such as poor generalization, an inability to infer causal relationships, and a lack of interpretability. Physics-Informed Machine Learning (PIML) helps mitigate these limitations by incorporating prior physical knowledge directly into the ML pipeline, thereby fostering growing interest in its application to PHM. This work investigates how PIML is being leveraged in the context of PHM through a systematic literature review of 212 studies. The review introduces a four-class classification scheme, consisting of observational bias, inductive bias, learning bias, and hybrid approaches, and further categorizes studies by PHM task. Across all four classes, the reviewed studies consistently demonstrate improved predictive performance over conventional baselines across a broad range of assets, although the literature is heavily skewed toward lithium-ion batteries and bearings, and dominated by problem-specific solutions. Overall, the review indicates that physics-informed approaches already provide tangible benefits, whereas claims of improvements concerning some of the aforementioned limitations lack sufficient supporting evidence. Future research should prioritize transferable design patterns, benchmarks comparing integration strategies, and uncertainty-aware models that are lightweight and robust enough for online deployment in real-world settings.

摘要:在現代工業中,保持複雜系統的可靠性、安全性和效率依賴於預測與健康管理(PHM)。機器學習(ML)在診斷和預測方面的進展主要受到推動,但純數據驅動的模型面臨固有的限制,例如一般化能力差、無法推斷因果關係以及缺乏可解釋性。物理知識驅動的機器學習(PIML)通過將先前的物理知識直接納入機器學習流程,幫助減輕這些限制,從而引發對其在PHM應用中的日益關注。本研究通過對212項研究的系統文獻回顧,調查了PIML在PHM背景下的應用。該回顧介紹了一種四類分類方案,包括觀察偏差、歸納偏差、學習偏差和混合方法,並進一步根據PHM任務對研究進行分類。在所有四類中,所回顧的研究一致顯示出相較於傳統基準的預測性能有所改善,儘管文獻在很大程度上偏向於鋰離子電池和軸承,並且以問題特定的解決方案為主導。總體而言,該回顧表明,物理知識驅動的方法已經提供了切實的好處,而對於上述某些限制的改善的聲稱缺乏足夠的支持證據。未來的研究應優先考慮可轉移的設計模式、比較整合策略的基準以及足夠輕量且穩健的、不確定性感知模型,以便在現實環境中進行在線部署。

Coupled Graph--Policy Distillation for Personalized Medication Safety in Older Adults with Multimorbidity

2608.09443v1 by Zihan Wang, Anglin Liu, Rongyi Wang, Dantong Li, Yi Lu, Siqing Yuan, Hongxia Xu, Zhongtian Long, Jintai Chen

Large language model (LLM) agents can support medication review between clinical visits, but safe choices for older adults with multimorbidity depend on conditions, medications, and geriatric risks that users may omit. We introduce ATLAS, a coupled graph--policy distillation framework for patient-adaptive medication safety. ATLAS structures guideline evidence as a medication-safety graph. Targeted questions update the patient state and distill relevant relations into a patient-specific medication conflict graph (PMCG). A risk-first multi-agent policy uses the PMCG to screen contraindications, assess cautions and monitoring needs, identify safer alternatives, and verify the final medication plan. We also introduce GeriMedBench, an interactive benchmark that tests safety-critical information acquisition and evidence-based decision revision. Across a European non-interactive multimorbidity benchmark, an Asian interactive multimorbidity benchmark, and an Asian non-interactive cross-guideline benchmark, ATLAS achieves the strongest complete-decision performance among the compared systems. On the European non-interactive multimorbidity benchmark, it exceeds the strongest proprietary LLM baseline by 53.73 points in Strict Success Rate and 14.63 points in overall safety reasoning score (OSRS), with no unsafe recommendations under the automated evaluator. A blinded clinician evaluation gives ATLAS higher mean ratings across all five criteria and flags potentially unsafe recommendations in one ATLAS case and two Gemini cases.

摘要:大型語言模型 (LLM) 代理可以在臨床訪問之間支持藥物審查,但對於多重疾病的老年人,安全的選擇取決於用戶可能省略的條件、藥物和老年風險。我們介紹了 ATLAS,一個耦合圖形-政策蒸餾框架,用於患者自適應藥物安全。ATLAS 將指導證據結構化為藥物安全圖。針對性的問題更新患者狀態,並將相關關係蒸餾成患者特定的藥物衝突圖 (PMCG)。一種以風險為先的多代理政策利用 PMCG 來篩選禁忌症,評估注意事項和監測需求,識別更安全的替代品,並驗證最終的藥物計劃。我們還介紹了 GeriMedBench,一個互動基準,測試安全關鍵信息獲取和基於證據的決策修訂。在一個歐洲非互動多重疾病基準、一個亞洲互動多重疾病基準和一個亞洲非互動跨指導基準中,ATLAS 在比較系統中實現了最強的完整決策表現。在歐洲非互動多重疾病基準中,它在嚴格成功率方面超過了最強的專有 LLM 基準 53.73 分,在整體安全推理分數 (OSRS) 上超過 14.63 分,且在自動評估者下沒有不安全的建議。盲評的臨床醫生評估給予 ATLAS 在所有五個標準上更高的平均評分,並在一個 ATLAS 案例和兩個 Gemini 案例中標記了潛在的不安全建議。

Multimodal Federated Learning under Dual-Axis Modality Missingness

2608.09240v1 by Adiba Orzikulova, Jaehyun Kwak, Jaemin Shin, Yunqi Guo, Xiaomin Ouyang, Guoliang Xing, Steven Euijong Whang, Sung-Ju Lee

Multimodal federated learning (FL) supports collaborative modeling in privacy-sensitive health-sensing and medical settings, but realistic deployments often exhibit dual-axis modality missingness: clients have different modality sets, and individual samples may contain only subsets of the modalities available locally. Existing methods typically address these two axes separately. We propose Flux, a multimodal federated learning framework built around two complementary components. First, modality-aware confidence tempering learns sample-specific confidence for each modality through mask-aware unimodal supervision and fuses the confidence estimates from observed modalities into a sample-adaptive temperature that adjusts predictive sharpness according to evidence quality and completeness. Second, gradient-decoupled private adaptation applies this temperature only to a client-private prediction pathway, while training the shared federated model with a standard, untempered objective. This enables sample-specific, client-local confidence adaptation without allowing confidence-dependent gradients to perturb shared representation learning. Across four multimodal datasets, Flux achieves the highest average macro-F1 on every dataset, outperforming the strongest dataset-specific baseline by 0.8~2.2 points and by 1.6 points on average. Additional analyses demonstrate favorable calibration, temperature sensitivity to both modality missingness and input corruption, and more stable shared optimization under private-only tempering. Our code is available at https://github.com/AdibaOrz/Flux.

摘要:多模態聯邦學習(FL)支持在隱私敏感的健康感測和醫療環境中進行協作建模,但現實部署通常顯示出雙軸模態缺失的情況:客戶端擁有不同的模態集,且個別樣本可能僅包含當地可用模態的子集。現有的方法通常分別處理這兩個軸。我們提出了Flux,一個圍繞兩個互補組件構建的多模態聯邦學習框架。首先,模態感知的信心調整通過基於掩碼的單模態監督學習每個模態的樣本特定信心,並將觀察到的模態的信心估計融合成樣本自適應的溫度,根據證據的質量和完整性調整預測的清晰度。其次,梯度解耦的私有適應僅將這個溫度應用於客戶端私有的預測路徑,同時使用標準的、未調整的目標訓練共享的聯邦模型。這使得樣本特定的、客戶端本地的信心適應成為可能,而不允許依賴信心的梯度干擾共享表示學習。在四個多模態數據集上,Flux在每個數據集上都達到了最高的平均宏F1,超過了最強的數據集特定基線0.8~2.2點,平均超過1.6點。額外的分析顯示出良好的校準、對模態缺失和輸入損壞的溫度敏感性,以及在僅進行私有調整時更穩定的共享優化。我們的代碼可在 https://github.com/AdibaOrz/Flux 獲得。

Rethinking Medical Landmark Localization with Prototype Learning-based Progressive Offset Correction

2608.09182v2 by Jingxian Xu, Yuhao Huang, Rusi Chen, Yanfeng Zhou, Dong Ni

Accurate landmark localization in medical images is a fundamental step for quantitative clinical measurement and downstream analysis. Existing localization methods have advanced, among which multi-stage refinement is a superior solution. Although this strategy mitigates the anatomical ambiguity inherent in single-stage global predictions, its high computational cost limits practical applicability. In this work, we propose a parameter-economic model, PPOC-LL, which leverages Prototype learning-based Progressive Offset Correction for Landmark Localization. Our contribution is three-fold. First, to drive coarse-to-fine landmark optimization, we introduce a multi-scale dynamic perception strategy for patch-level feature pyramid modeling. Second, to effectively handle anatomically similar patterns, we design a similarity-driven prototype learning mechanism that captures informative local semantics for robust offset prediction. Last, to stabilize the model learning and improve the overall performance, we incorporate a novel error-aware reliability regularization via tolerance-based balancing. We collected a large validation cohort, including two public and one private datasets spanning X-ray and ultrasound modalities, covering cephalometric, symphysis-fetal head, and fetal heart landmarks. Extensive experiments demonstrate that PPOC-LL achieves satisfactory performance with a favorable trade-off between accuracy and model complexity.

摘要:醫學影像中準確的地標定位是進行定量臨床測量和後續分析的基本步驟。現有的定位方法已經取得了進展,其中多階段精煉是一個優越的解決方案。儘管這一策略減輕了單階段全局預測中固有的解剖學模糊性,但其高計算成本限制了實際應用。在本研究中,我們提出了一種經濟參數模型PPOC-LL,該模型利用基於原型學習的漸進偏移修正來進行地標定位。我們的貢獻有三個方面。首先,為了驅動粗到精的地標優化,我們引入了一種多尺度動態感知策略,用於補丁級特徵金字塔建模。其次,為了有效處理解剖學上相似的模式,我們設計了一種基於相似性的原型學習機制,該機制捕捉了有用的局部語義,以實現穩健的偏移預測。最後,為了穩定模型學習並提高整體性能,我們通過基於容忍度的平衡引入了一種新穎的錯誤感知可靠性正則化。我們收集了一個大型驗證隊列,包括兩個公共數據集和一個私有數據集,涵蓋X光和超聲波模態,涵蓋顱面測量、恥骨聯合-胎頭和胎心地標。大量實驗表明,PPOC-LL在準確性和模型複雜性之間達成了令人滿意的性能和良好的權衡。

RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning

2608.09123v1 by Jinkun Hou, Zhuo Liu, Huimin Ren, Hongsheng Xin, Pan Zhou, Kun Zhan

Aligning Large Language Models (LLMs) for open-ended tasks is challenging because responses must satisfy multidimensional criteria without following a single correct generation trajectory. Existing rubric-based reinforcement learning (RL) methods compress fine-grained criterion-level feedback into scalar rewards, making persistent capability gaps difficult to target under limited on-policy exploration. We propose $\textbf{RISE-RL}$ (Rubric-Informed Selective Exploration), which uses repeatedly missed rubric criteria to elicit privileged trajectories that are difficult to discover through unguided exploration alone. RISE-RL retains only trajectories whose complete-rubric reward exceeds the mean reward of natural rollouts, and then re-evaluates them under the original prompt to emphasize behaviors that remain weakly supported by the natural policy. The resulting guidance signal is optimized through a separate auxiliary objective and removed once its additional benefit diminishes. Experiments with 4B and 14B models across writing, chat, health, and science show that RISE-RL achieves the highest mean score on every evaluated benchmark under guidance-free evaluation. Compared with standard Rubric-RL, it improves the average score by 1.3 points at the 4B scale and $\textbf{3.3 points at the 14B scale}$, including a $\textbf{6.0-point}$ gain on CreativeWriting-V3. It also improves creative-writing diversity and yields gains on objectively scored medical and scientific benchmarks. These results indicate that selective internalization through reward filtering and policy support shaping is effective for open-ended reinforcement learning.

摘要:對於開放式任務,對大型語言模型(LLMs)的調整是具有挑戰性的,因為回應必須滿足多維標準,而不必遵循單一正確的生成軌跡。現有的基於評分標準的強化學習(RL)方法將細緻的標準級反饋壓縮為標量獎勵,使得在有限的政策探索下,持續的能力差距難以針對。我們提出了 $\textbf{RISE-RL}$(基於評分標準的選擇性探索),該方法利用重複錯過的評分標準來引出難以通過無引導探索單獨發現的特權軌跡。RISE-RL 僅保留那些完整評分獎勵超過自然回合平均獎勵的軌跡,然後在原始提示下重新評估它們,以強調在自然政策中仍然支持不足的行為。產生的指導信號通過單獨的輔助目標進行優化,並在其額外效益減少後被移除。對於寫作、聊天、健康和科學的 4B 和 14B 模型進行的實驗顯示,RISE-RL 在無指導評估下在每個評估基準上都達到了最高的平均得分。與標準的 Rubric-RL 相比,它在 4B 規模上提高了 1.3 分的平均得分,在 14B 規模上提高了 $\textbf{3.3 分}$,包括在 CreativeWriting-V3 上的 $\textbf{6.0 分}$ 增加。它還改善了創意寫作的多樣性,並在客觀評分的醫療和科學基準上取得了增益。這些結果表明,通過獎勵過濾和政策支持塑造的選擇性內化對於開放式強化學習是有效的。

A Multi-Scale Temporal Framework with Dynamic Fusion for EEG-Based Emotion Recognition

2608.09088v1 by Stefanos Gkikas, Yang Guo, Guangliang Li, Raul Fernandez Rojas, Giorgos Giannakakis, Randy Gomez

Mixed emotions represent a clinically relevant but still underexplored target for automatic emotion recognition. EEG provides millisecond-level access to neural activity, yet most EEG pipelines analyze the signal through a single temporal window, thereby fixing the temporal structure available to the model. This study introduces a multi-scale temporal framework for EEG-based emotion recognition. The EEG waveform is decomposed into windows of one or several durations, processed by a shared attention-based encoder, and integrated through a dynamic fusion module that assigns sample-specific weights across temporal scales. The framework is evaluated under a subject-independent protocol in binary and three-class settings, with the three-class task including the mixed affective category. The best results are 65.22% for the two-class task and 45.43% for the three-class task. Both are obtained with three-scale dynamic-fusion configurations and remain substantially above the full-signal baseline. The best-performing temporal scales differ between the two tasks. Dynamic fusion outperforms concatenation in the highest-scoring two-class configuration and slightly exceeds it in the highest-scoring three-class configuration, although these multi-scale settings require substantially more computation than the full-signal baseline.

摘要:混合情緒代表了一個臨床相關但仍未充分探索的自動情緒識別目標。EEG 提供毫秒級的神經活動訪問,但大多數 EEG 流程通過單一時間窗口分析信號,從而固定了模型可用的時間結構。本研究引入了一個基於 EEG 的情緒識別的多尺度時間框架。EEG 波形被分解為一個或多個持續時間的窗口,通過共享注意力編碼器進行處理,並通過動態融合模塊進行整合,該模塊在時間尺度上分配樣本特定的權重。該框架在一個獨立於受試者的協議下進行評估,涵蓋二元和三類設置,其中三類任務包括混合情感類別。最佳結果為二類任務的 65.22% 和三類任務的 45.43%。這兩者都是在三尺度動態融合配置下獲得的,並且均顯著高於全信號基準。最佳的時間尺度在這兩個任務之間有所不同。在得分最高的二類配置中,動態融合的表現優於串接,而在得分最高的三類配置中略微超過串接,儘管這些多尺度設置所需的計算量遠高於全信號基準。

When Confidence Fails: Overconfidence in LLMs under Uncertainty and Missing Clinical Information

2608.09080v1 by Maryam Tahermazandarani, Adnan Mahmood, Fahmida Islam, Quan Z. Sheng

Large Language Models (LLMs) have achieved strong performance in medical question answering and clinical reasoning tasks. However, their reliability under uncertainty remains poorly understood which raises critical concerns for deployment in high-stakes clinical settings. In such environments, incorrect predictions are inherently risky, but confident incorrect predictions can be particularly harmful as they may mislead clinical decision-making. In this paper, we conduct a systematic behavioral analysis of LLMs under clinical information uncertainty. We propose an evaluation framework based on the MedMCQA dataset consisting of two complementary uncertainty settings. First, we introduce linguistic uncertainty cues through prompt modifications to simulate ambiguous clinical contexts. Second, we construct an answer removal setting, wherein the correct option is deliberately excluded mandating the model to recognize insufficient information and abstain. We analyze both model accuracy and confidence behavior using multiple calibration metrics including calibration gap, Expected Calibration Error (ECE), and Unsafe Confident Error Rate (UCER) across 500 medical questions. Our results reveal a consistent failure mode, i.e., although accuracy degrades under increasing uncertainty, model confidence remains misaligned with accuracy. This leads to a substantial increase in unsafe confident errors, indicating that model confidence remains largely insensitive to clinically meaningful information loss. Furthermore, we observe significant variation across models in their ability to abstain when the correct answer is unavailable, with some models persistently producing high confidence hallucinated answers. These findings expose critical limitations in the epistemic reliability of current LLMs and highlight the need for uncertainty aware evaluation methods prior to their deployment in clinical workflows.

摘要:大型語言模型(LLMs)在醫療問題回答和臨床推理任務中取得了強勁的表現。然而,它們在不確定性下的可靠性仍然不甚了解,這對於在高風險臨床環境中的部署提出了關鍵的擔憂。在這樣的環境中,錯誤的預測本質上是有風險的,但自信的錯誤預測可能特別有害,因為它們可能會誤導臨床決策。在本文中,我們對LLMs在臨床信息不確定性下進行了系統的行為分析。我們提出了一個基於MedMCQA數據集的評估框架,該數據集包含兩個互補的不確定性設置。首先,我們通過提示修改引入語言不確定性線索,以模擬模糊的臨床情境。其次,我們構建了一個答案移除設置,其中正確選項故意被排除,要求模型識別信息不足並選擇不作答。我們使用多個校準指標分析模型的準確性和信心行為,包括校準差距、期望校準誤差(ECE)和不安全自信錯誤率(UCER),涵蓋500個醫療問題。我們的結果揭示了一種一致的失敗模式,即儘管準確性在不斷增加的不確定性下下降,但模型的信心與準確性仍然不一致。這導致不安全的自信錯誤顯著增加,表明模型的信心對臨床上有意義的信息損失仍然大致不敏感。此外,我們觀察到不同模型在正確答案不可用時的選擇不作答能力上存在顯著差異,一些模型持續產生高信心的虛假答案。這些發現揭示了當前LLMs在認識論可靠性方面的關鍵局限性,並強調了在其部署於臨床工作流程之前需要不確定性感知的評估方法。

Decoding Phenotypes: A Framework for Fusing Genomic Language Models and Neuroimaging

2608.08926v1 by Tianli Tao, Ziyang Wang, Emma Robinson, Rachel Sparks, Le Zhang

Neuroimaging and genetic testing are two important clinical references for nervous system diseases, offering complementary diagnostic information. However, integrating genomic and neuroimaging data for precise disease diagnosis is challenging due to cross-modality heterogeneity. Existing imaging-genetics approaches mainly encode genetic information as hard-coded labels, which lose the local sequence context around disease-associated variants. To address this limitation, we propose GeneFuse, a multimodal learning framework that aligns genetic representations from pre-trained Genomic Language Models (GLMs) with features extracted from images. GeneFuse integrates two components: (1) Genotype-Conditioned Feature Modulation (GCFM), a FiLM-inspired module that uses genomic embeddings to modulate image feature maps; and (2) Uncertainty-aware Genomic Residual Fusion (U-GRF), a fusion strategy that uses imaging-derived predictive uncertainty to gate the contribution of genotypic features. We evaluate GeneFuse on early cognitive decline identification (NC vs. MCI) and dementia screening (NC vs. AD). In the APOE-centered setting, GeneFuse achieves AUROCs of 0.77 and 0.83, outperforming existing imaging-genetics fusion methods. These results indicate that GLM-derived genomic embeddings provide additional information to imaging.

摘要:神經影像學和基因檢測是神經系統疾病的重要臨床參考,提供互補的診斷信息。然而,由於跨模態異質性,整合基因組和神經影像數據以進行精確的疾病診斷是具有挑戰性的。現有的影像-基因學方法主要將基因信息編碼為硬編碼標籤,這樣會失去與疾病相關變異周圍的局部序列上下文。為了解決這一限制,我們提出了GeneFuse,一個多模態學習框架,將來自預訓練基因組語言模型(GLMs)的基因表示與從影像中提取的特徵對齊。GeneFuse整合了兩個組件:(1)基因型條件特徵調製(GCFM),一個受FiLM啟發的模塊,使用基因嵌入來調製影像特徵圖;以及(2)不確定性感知基因組殘差融合(U-GRF),一種融合策略,利用影像衍生的預測不確定性來控制基因型特徵的貢獻。我們在早期認知衰退識別(NC vs. MCI)和癡呆篩查(NC vs. AD)上評估了GeneFuse。在以APOE為中心的設置中,GeneFuse達到了0.77和0.83的AUROC,超越了現有的影像-基因學融合方法。這些結果表明,來自GLM的基因嵌入為影像提供了額外的信息。

Toward CT-Equivalent Image Quality in Low-Dose Radiotherapy Planning: Conditional Diffusion-Based CBCT-to-CT Synthesis and the Impact of CBCT Input Representation

2608.08919v1 by Alzahra Altalib, Chunhui Li, Christopher Hamill Taylor, Sankar Pillai, Alessandro Perelli

During standard radiotherapy planning, repeated CT acquisitions are often required for patient registration, verification, and adaptive planning, resulting in increased cumulative X-ray dose. To mitigate this, low-dose cone-beam CT (CBCT) is routinely acquired during treatment delivery. However, CBCT image quality remains insufficient for accurate dose calculation and adaptive radiotherapy planning due to increased scatter, noise, beam hardening, and reconstruction related artifacts. This study develops a supervised deep learning based CBCT to CT synthesis framework using a conditional denoising diffusion probabilistic model (DDPM), where the generation of a CT-based planning for accurate positioning and dose calculation is obtained using generative models with low dose CBCT imaging. Beyond demonstrating CBCT to CT synthesis, the primary objective is to investigate how the representation of CBCT input data, either standard clinical DICOM CBCT images or filtered back-projection (FDK) reconstructions from raw projection data, affects the performance of diffusion based CT synthesis. The overarching aim is to assess whether physics aware CBCT representations better support CT-equivalent image quality while maintaining reduced imaging dose in radiotherapy workflows.

摘要:在標準放射治療計劃中,通常需要重複進行 CT 採集以進行病人登記、驗證和自適應計劃,這導致累積的 X 射線劑量增加。為了減輕這一問題,在治療過程中常規獲取低劑量圓錐束 CT (CBCT)。然而,由於散射、噪聲、束硬化和重建相關的伪影,CBCT 圖像質量仍不足以進行準確的劑量計算和自適應放射治療計劃。本研究開發了一個基於監督式深度學習的 CBCT 到 CT 合成框架,使用條件去噪擴散概率模型 (DDPM),通過生成模型與低劑量 CBCT 成像來獲得準確定位和劑量計算所需的 CT 基礎計劃。除了展示 CBCT 到 CT 的合成,主要目標是研究 CBCT 輸入數據的表示,無論是標準臨床 DICOM CBCT 圖像還是來自原始投影數據的濾波反投影 (FDK) 重建,如何影響基於擴散的 CT 合成性能。總體目的是評估物理感知的 CBCT 表示是否更好地支持 CT 等效的圖像質量,同時在放射治療工作流程中保持降低的成像劑量。