arxiv-daily
Automated deployment @ 2026-08-20 08:58:50 Asia/Taipei
Welcome to contribute! Add your topics and keywords in
topic.yml. You can also view historical data through the storage.
AI
Knowledge Graphs
| Publish Date | Title | Authors | Homepage | Code |
|---|---|---|---|---|
| 2026-08-18 | From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation | Xingjian Wang et.al. | 2608.18076v1 | null |
| 2026-08-18 | StagedWorkspace: A Versioned Workspace for Knowledge-Work Agents | Yining Hua et.al. | 2608.18050v1 | null |
| 2026-08-18 | Chain-of-Experience for Continual LLM Improvement | Haoqin Tu et.al. | 2608.18027v1 | null |
| 2026-08-18 | Can Large Language Models Explain Flight Safety Events? A Prior-Guided Semantic LLM-based Approach | Lu Xu et.al. | 2608.18017v1 | null |
| 2026-08-18 | The IOL-AI Challenge: An Open Challenge towards Advancing Linguistic Reasoning | Eduardo Sánchez et.al. | 2608.18011v1 | null |
| 2026-08-18 | Judge, Retrieve, or Abstain: Uncertainty-Guarded LLM Judging with Provable Risk Guarantees | Sher Badshah et.al. | 2608.17994v1 | null |
| 2026-08-18 | Against Political Polarization: A Unified Framework for Tracing Evolving Political Ideologies on Social Media | Yijie Xu et.al. | 2608.17987v1 | null |
| 2026-08-18 | Do Large Language Models Play Six Degrees of Separation? Measuring Topological Compression in Long-Context Manifolds | Md. Faiyaz Abdullah Sayeedi et.al. | 2608.17950v1 | null |
| 2026-08-18 | Efficient RLVR Scheduling via Graph-Structured Online Difficulty Estimation | Zhizhao Liu et.al. | 2608.17941v1 | null |
| 2026-08-18 | Collective Counterfactual Planning: Coordination, Consent, and Verification under Representational Constraints | Chainarong Amornbunchornvej et.al. | 2608.17932v1 | null |
| 2026-08-18 | Analysis of Types of Inquiries in Student-AI Interaction: A case study of two CS2 tasks | Matin Amoozadeh et.al. | 2608.17919v1 | null |
| 2026-08-18 | AutoResearch: Insight In, Hallucination Out | Yiming Ren et.al. | 2608.17906v1 | null |
| 2026-08-18 | BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models | Liubov Chubarova et.al. | 2608.17895v1 | null |
| 2026-08-18 | From Global Benchmarks to Local Evaluations: Benchmarking LLMs for the German Public Sector | Camilla Dalerci et.al. | 2608.17827v1 | null |
| 2026-08-18 | Interpretable Humans, Alien LLMs: Expert Analysis of Latent Structures in Assessment Responses | Alona Strugatski et.al. | 2608.17810v1 | null |
| 2026-08-18 | Whether LLMs Can Navigate Beliefs and Facts Depends on How You Phrase It | Quang Minh Nguyen et.al. | 2608.17809v1 | null |
| 2026-08-18 | An Empirical Study of Reward Specification and Benchmark Reliability in GRPO-based LLM Unlearning | Rubén Balbastre et.al. | 2608.17804v1 | null |
| 2026-08-18 | Neuro-symbolic learning over OWL 2 DL via consequence-based compilation to differentiable circuits | Olga Mashkova et.al. | 2608.17741v1 | null |
| 2026-08-18 | What Aggregate Scores Miss: Measuring Item-Level Regressions in Commercial LLM API Migrations | Xiaonan Xu et.al. | 2608.17719v1 | null |
| 2026-08-18 | Communicating Credit Risk with Large Language Models: Evaluation of Explanations from Standard and Alternative Data-Based Models | Sahab Zandi et.al. | 2608.17715v1 | null |
| 2026-08-18 | GraphWake: Group Polarization via Memory-Mediated Polarization Cascade in LLM-Agent Communities | Haoran Bu et.al. | 2608.17665v1 | null |
| 2026-08-18 | Graph Surgery and the Do-Operator: A Precise Correspondence for Acyclic Structural Causal Models | Satpreet Makhija et.al. | 2608.17634v1 | null |
| 2026-08-18 | Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges | Syeda Faiza Ahmed et.al. | 2608.17605v1 | null |
| 2026-08-18 | tinyDSM: A Framework for Skill Modeling and Development for Resource-Constrained Millirobots | Markus D. Kobelrausch et.al. | 2608.17596v1 | null |
| 2026-08-18 | Quantifying Risk Under Evolving Uncertainty: Belief-Dependent Robustness for Safe Sequential Decision Making | Deep Kumar Ganguly et.al. | 2608.17574v1 | null |
| 2026-08-18 | Code as Representation: A Compilable Parsing Paradigm for Academic Documents | Rihui Jin et.al. | 2608.17550v1 | null |
| 2026-08-18 | CoAL-RAG: A Complexity-Aware Legal Retrieval-Augmented Generation Method | Jin Su et.al. | 2608.17536v1 | null |
| 2026-08-18 | When to Review: Spaced Repetition for Continual Pre-Training of Language Models | Alankar Atreya et.al. | 2608.17530v1 | null |
| 2026-08-18 | Effects of Answer Format Variation on Gender Bias in Large Language Models | Ksenia Merzlyakova et.al. | 2608.17516v1 | null |
| 2026-08-18 | Beyond FLOPs: Energy-Aware Knowledge Distillation for Sustainable LLMs on Code-Related Task | Enrique Barba Roque et.al. | 2608.17515v1 | null |
| 2026-08-18 | SGHA: Evidence-Grounded Research Problem Discovery with Local Language Models | Sarvesh Gharat et.al. | 2608.17501v1 | null |
| 2026-08-18 | SAGE: Self-Evolving Storyboard Skills via Attribution-Guided Rule Evolution | Maolin Ran et.al. | 2608.17468v1 | null |
| 2026-08-18 | Structure-Internalized Rule Language Model for Faithful Knowledge Graph Reasoning | Xingrui Zhuo et.al. | 2608.17443v1 | null |
| 2026-08-18 | Cognitive Graph Intelligence for Adaptive and Robust DDoS Attack Detection in Next Generation Networks | Mohammad Arif Hossain et.al. | 2608.17352v1 | null |
| 2026-08-18 | DeAR: Decentralized Agentic Reasoning via Capability Grounding and Collaborative Thought Navigation | Xing Wei et.al. | 2608.17282v1 | null |
| 2026-08-18 | ASI-Bench: At the Dawn of Artificial Superintelligence | Junwei Zhou et.al. | 2608.17271v1 | null |
| 2026-08-18 | Understanding Curriculum Learning in Large Language Models via Cross-Difficulty Optimization Dynamics | Zhikai Ding et.al. | 2608.17268v1 | null |
| 2026-08-18 | Structural Plan-to-Model Conversion with Deterministic Geometry and Guarded Agentic Vision-Language Refinement | Mohammad Talebi-Kalaleh et.al. | 2608.17237v1 | null |
| 2026-08-17 | Synthesizing Feature Extractors: An Agentic Approach for Algorithm Selection | Hai Xia et.al. | 2608.17170v1 | null |
| 2026-08-17 | Towards Safer RAG: Only Agents Capable of System 2 Thinking may Access Untrusted Documents | Mehrdad Ghassabi et.al. | 2608.17153v1 | null |
| 2026-08-17 | KnowSim: Evaluating Information Calibration in LLM Assistants with User Simulators that Learn | Yoonjoo Lee et.al. | 2608.17150v1 | null |
| 2026-08-17 | A decodability criterion predicts when hidden-state selection beats majority voting in large language models | Zhixiang wang et.al. | 2608.17124v1 | null |
| 2026-08-17 | From Abductive Explanations to Global Logical Rules for Node Classification in SGCs | Bryan Lima Cavalcante et.al. | 2608.17103v1 | null |
| 2026-08-17 | J-Miner: Recovering Executable Decision Knowledge from Language-Model Classifiers | Yunfan Gao et.al. | 2608.17063v1 | null |
| 2026-08-17 | Cross-Model Memory Transfer via Target-Side Reader Adaptation | Mingyuan Li et.al. | 2608.17050v1 | null |
| 2026-08-17 | AutoSR: Automatic Symbolic Regression by Searching Research States | Kejia Zhang et.al. | 2608.16876v1 | null |
| 2026-08-17 | Quipu: A Governed Bitemporal Knowledge Graph Store | Steve Brown et.al. | 2608.16813v1 | null |
| 2026-08-17 | Bounded Semantic Planning and Deterministic Compilation for Reliable Enterprise Text-to-SQL | Yi Ai et.al. | 2608.16663v1 | null |
| 2026-08-17 | The Working Set of a Coding Agent: Coherence Debt in Repository-Scale Tasks | Bardia Mohammadi et.al. | 2608.16630v1 | null |
| 2026-08-17 | Hypergraph-based Multimodal Retrieval-Augmented Generation with Incremental Refinement | Shenao Chen et.al. | 2608.16628v1 | null |
| 2026-08-17 | Physics of Agents: Statistical Mechanics Predicts Collective Behavior of AI Agents | Batu El et.al. | 2608.16578v1 | null |
| 2026-08-17 | Ask, Condition or Abstain: Reinforcement Learning for Missing-Premise Reasoning | Yongqi Tong et.al. | 2608.16554v1 | null |
| 2026-08-17 | VCE-Skill: Enhancing Skill Self-Evolution with Version-Change Experience | Jianming Chen et.al. | 2608.16544v1 | null |
| 2026-08-17 | Large language models as synthetic clinical experts to inform longitudinal rare-disease modeling | Clemens Schächter et.al. | 2608.16507v1 | null |
| 2026-08-17 | Graph Machine Learning: An Opportunity for Power Systems | Martin Sadric et.al. | 2608.16494v1 | null |
| 2026-08-17 | Time to Reason: Scalable Neurosymbolic Learning for LTLf via Fuzzy Semantics | Riccardo Andreoni et.al. | 2608.16443v1 | null |
| 2026-08-17 | Reasoning-supported Robustness Validation of Automotive E/E Components | Jan Novacek et.al. | 2608.16421v1 | null |
| 2026-08-17 | Think Inside the Chunk: RegulaRAG for Regulation-Compliant Scenario Generation using LLMs: A Case Study of UN Regulation No. 152 | Vahid Zolfaghari et.al. | 2608.16394v1 | null |
| 2026-08-17 | Mint-Agent: Introducing Finance-Native Agentic Foundation Models | Mint-Agent Team et.al. | 2608.16386v1 | null |
| 2026-08-17 | MELD: A Protocol for Merging Knowledge Across Distributed Agentic Memories | Lauri Lovén et.al. | 2608.16357v1 | null |
| 2026-08-17 | AeroCopilotBench: A Two-Tier Benchmark for Evaluating LLM Agents as Aviation Copilots in an Interactive Virtual Cockpit Environment | Yuchen Yuan et.al. | 2608.16349v1 | null |
| 2026-08-17 | Executable Code Knowledge: Code as a Native, Validation-Carrying Knowledge Representation for AI Coding Agents | Xueping Gao et.al. | 2608.16295v1 | null |
| 2026-08-17 | Clause Encounters of the Third Kind: Can LLMs Replace Language Teachers? | Kristina Šekrst et.al. | 2608.16286v1 | null |
| 2026-08-17 | Domain-Agnostic Neural Topic Modeling with Contextual Token-Level Semantic Graph Representation | Seung-Won Seo et.al. | 2608.16269v1 | null |
| 2026-08-17 | Picking the Right Image to Classify: Reliable-Input Selection in Teledermatology | Fabian Gröger et.al. | 2608.16198v1 | null |
| 2026-08-17 | LENS: In-Context Search via Latent Evidence Exploration over Dynamic Raw Documents | Xingjun Wang et.al. | 2608.16185v2 | null |
| 2026-08-17 | Agent-Native Telemetry: Verifiable State-Delta Evidence for Autonomous Operations | Jun He et.al. | 2608.16178v1 | null |
| 2026-08-17 | FeatureHospital: A Skill-Driven Multi-Agent Framework for Automated Algorithm Customization in Multi-View Multi-Label Feature Selection | Junxuan Li et.al. | 2608.16148v1 | null |
| 2026-08-17 | Graph Neural Assisted Actor-Critic for Latency-Efficient Edge Vision System | Alam Noor et.al. | 2608.16142v1 | null |
| 2026-08-17 | HyperSkill: Self-Evolving LLM Agents via Hypergraph-Structured Skill Memory | Ruiyao Xu et.al. | 2608.16114v1 | null |
| 2026-08-17 | RetroMPA: A Molecular Property-Aware Auxiliary Framework for Enhancing Retrosynthesis Prediction | Mianzhi Liu et.al. | 2608.16111v1 | null |
| 2026-08-17 | The Commercial Tax: Rent-vs-Own Blind Spots in Multi-Hop Retrieval Benchmarks | Luis M. Sanchez et.al. | 2608.16096v1 | null |
| 2026-08-17 | Skill2Query: Exploiting Skill Structure to Generate Pseudo-Queries for Agent Skill Retrieval | Lihui Ding et.al. | 2608.16071v1 | null |
| 2026-08-17 | OceanLight: Efficient Global Ocean Forecasting via Geometry-Adaptive Unstructured Mesh Representation | Wei Wu et.al. | 2608.16070v1 | null |
| 2026-08-17 | NICE: Scale-Stable Perturbations for Graph Neural Network Explanations via Noise Corruption | Ziluowen Luo et.al. | 2608.16038v2 | null |
| 2026-08-17 | RagGAD: Rationale-Aware Conditional Gaussian Mixture Normalizing Flow for Unsupervised Graph Anomaly Detection | Junxin Lu et.al. | 2608.16018v1 | null |
| 2026-08-17 | From Sequence to Structure: Relational Uncertainty Propagation for LLM Agents | Zhengzhao Ma. Boxi Cao et.al. | 2608.16002v1 | null |
| 2026-08-16 | PLSQLBench: Benchmarking LLM Systems for Executable Procedural Database Programming | Marianne Menglin Liu et.al. | 2608.15931v1 | null |
| 2026-08-16 | Unified Pedestrian Path Prediction Using Inverse Reinforcement Learning | Šimon Sukup et.al. | 2608.15929v1 | null |
| 2026-08-16 | Noesis: Bidirectional Graph-RAG with Adaptive Parallelism and Cross-Knowledge-Base Semantic Discovery | Nicola Cogotti et.al. | 2608.15919v1 | null |
| 2026-08-16 | Large language model-assisted discovery of cohorts from scientific literature | Moritz Sturm et.al. | 2608.15909v1 | null |
| 2026-08-16 | Scaling Manual-Grounded Appliance Manipulation with Data Synthesis and Unified Planning | Yuxing Long et.al. | 2608.15863v1 | null |
| 2026-08-16 | RAGas: Retrieval-Augmented Gas Optimization for Smart Contracts with Continuous Knowledge Integration | Yishun Wang et.al. | 2608.15857v1 | null |
| 2026-08-16 | Characterising cardiac tissue properties with graph neural networks | Ching-En Chiu et.al. | 2608.15843v1 | null |
| 2026-08-16 | Schema-Agnostic Graph Reasoning Agent for Hybrid Knowledge Graphs | Marius Dragic et.al. | 2608.15834v1 | null |
| 2026-08-16 | The Authority Resolution Framework: A Five-Domain Ontology for Governing Who and What Decides, at Scale | Parviz Shariff et.al. | 2608.15832v1 | null |
| 2026-08-16 | QuantumPhaseNet: A Gauge-Covariant Geometric and Quantum-Spectral Theory of Semantic Concept Hierarchies with Prototype Validation of a Classical Quantum-Inspired Model | Kiyotaka Kasubuchi et.al. | 2608.15820v1 | null |
| 2026-08-16 | ALKEMIE Agent: an autonomous platform for computational materials design | Hongfu Huang et.al. | 2608.15776v1 | null |
| 2026-08-16 | Large Models for Small Devices: Recent Advances and Empirical Analysis of Edge AI Deployment | Subhransu Das et.al. | 2608.15693v1 | null |
| 2026-08-16 | BERTopic-Virality Prioritisation: A Scalable Framework for Thematic and Comparative Analysis of COVID-19 and Monkeypox Misinformation on Twitter | Mkululi Sikosana et.al. | 2608.15691v1 | null |
| 2026-08-16 | THESIS-MoE: Trainable Hierarchical Extraction and SteerIng of Sycophancy in Mixture-of-Experts | Kareem Hassani et.al. | 2608.15687v1 | null |
| 2026-08-16 | Agent Gym: A Framework for Continuous Evaluation and Evolution of LLM Agents Through Human-in-the-Loop Feedback | Pouya Ghiasnezhad Omran et.al. | 2608.15591v1 | null |
| 2026-08-16 | GraniKV: Asymmetric Granularity KV-Cache Paging for Multi-Agent Systems with Long Shared Prefix | Jinhyun Jeon et.al. | 2608.15584v1 | null |
| 2026-08-16 | From Generalist to Specialist: A Context-Fusion Framework for Endoscopic Polyp Reporting with a Frozen VLM | Ruijie Yang et.al. | 2608.15580v1 | null |
| 2026-08-16 | Admission Without Answers: Label-Free Certification and Experience Learning for LLM-Based Optimization Modeling | Junbo Jacob Lian et.al. | 2608.15565v2 | null |
| 2026-08-16 | BengaliMCQ: Automatic Generation and Answer Prediction of Academic Multiple-Choice Questions in a Low-Resource Language | Abu Tarabin Surzo et.al. | 2608.15547v1 | null |
| 2026-08-16 | L3Cube-IndicQuest v2: A Large-Scale Multilingual Benchmark for Evaluating Factual Knowledge of Large Language Models Across Indic Languages | Rinit Jain et.al. | 2608.15535v1 | null |
| 2026-08-16 | Mental Model Management: An Operator-Based Framework for LLM Memory | Oliver Kramer et.al. | 2608.15451v1 | null |
| 2026-08-15 | Implementation of a Metacognition Framework for Self-Awareness and Self-Regulation in Ensembles of LLMs | Charles Courchaine et.al. | 2608.15400v1 | null |
| 2026-08-15 | Grounding Healthcare LLMs in a Causal Knowledge Graph: Framework, Metrics, and a Cardiovascular Pilot | Ummara Mumtaz et.al. | 2608.15382v1 | null |
Abstracts
From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation
2608.18076v1 by Xingjian Wang, Zhao Wang, Taihang Hu, Jun Zheng, Qing Jin, Qinye Zhou, Zhengtao Wu, Yongchao Du, Zuan Gao, Chao Lin, Yefeng Shen, Xiaoli Xu, Zhengze Xu, Hao Yan, Yuhang Yu, Mingzhou Zhang, Mengting Chen
Large-scale image generation has benefited from advances in data scale, quality, rebalancing, and recaptioning, yet conventional pipelines typically optimize task-specific datasets in isolation. A central challenge is not only how to curate each task-specific corpus, but also how to organize heterogeneous supervision according to the dependencies among generative capabilities. We present a \textbf{capability-driven data infrastructure} that couples capability-specific supervision construction with capability-aligned curriculum scheduling. Its three specialized yet interoperable data engines build complementary relational supervision for text-image grounding, inter-image transformation, and image-knowledge association, while caption experts align T2I and editing supervision across tasks and granularities. A multi-stage curriculum jointly evolves task composition, visual-concept distribution, data quality, and image resolution along the dependency order of capability acquisition, with capability-aware evaluation closing the loop through targeted retrieval, expert construction, and gap-aware resampling. At scale, the framework curates a 440M-image T2I corpus, 120M editing pairs, and over 27M image-entity pairs. With this infrastructure, we train multimodal diffusion models at two scales from scratch, with 3B and 6B sizes respectively. We conduct quantitative evaluation on CPI-Bench, along with qualitative evaluations across diverse text-to-image and editing scenarios. Experimental results present broad visual coverage, versatile rendering, and effective transfer across generative capabilities.
摘要:大規模圖像生成受益於數據規模、質量、重新平衡和重新標題的進步,但傳統流程通常在孤立的情況下優化特定任務的數據集。一個主要挑戰不僅在於如何策劃每個特定任務的語料庫,還在於如何根據生成能力之間的依賴關係組織異質監督。我們提出了一個\textbf{以能力為驅動的數據基礎設施},將特定能力的監督構建與能力對齊的課程安排結合起來。它的三個專門但可互操作的數據引擎為文本-圖像基礎、圖像間轉換和圖像-知識關聯構建互補的關係監督,同時標題專家在任務和粒度之間對齊T2I和編輯監督。一個多階段課程共同演變任務組合、視覺概念分佈、數據質量和圖像解析度,沿著能力獲取的依賴順序進行,而以能力為中心的評估通過針對性檢索、專家構建和關注差距的重採樣來閉合循環。在規模上,該框架策劃了一個包含4.4億圖像的T2I語料庫、1.2億編輯對和超過2700萬圖像-實體對。利用這一基礎設施,我們從零開始訓練了兩個規模的多模態擴散模型,分別為30億和60億大小。我們在CPI-Bench上進行了定量評估,並在多樣的文本到圖像和編輯場景中進行了定性評估。實驗結果顯示出廣泛的視覺覆蓋、多樣的渲染和在生成能力之間的有效轉移。
StagedWorkspace: A Versioned Workspace for Knowledge-Work Agents
2608.18050v1 by Yining Hua, Hongbin Na, Yifan Zhou, Akshay Kalose, Cyrus Ayubcha, Levi Lian
AI agents increasingly perform knowledge work (i.e., produce and modify persistent digital artifacts such as code repositories, documents, spreadsheets, slides, reports), yet the parsed views they search, the native files they edit, the changes they review, and the artifacts they submit can refer to different versions of the same work product. We formulate this as a workspace-state contract: every view should be explicitly tied to a version of the evolving workspace state. Coding agents partly address this need through repository contracts for search, diffs, and tests, whereas an analogous contract is less explicit for PDFs, spreadsheets, slides, notebooks, and mixed-format project folders. We propose StagedWorkspace, a versioned workspace for knowledge-work agents. The workspace binds parsed records and review diffs to content hashes of the native files as they change. In fixed-harness ablations on OfficeQA Pro and APEX-Agents, dual parsed/native access has the highest point estimate for every tested model; relative to the more limiting single view, it improves OfficeQA Pass@1 by 8.3-12.1 points and APEX mean rubric score by 4.7-9.2 points. SW-AGENT scores 63.9% with Gemini 3.1 Pro on OfficeQA and 42.1 with GPT-5.4 Nano on APEX, compared with published same-model scores of 29.3% and 25.5, respectively. A paired review-axis ablation on 57 file-editing tasks further finds higher observed scores when diffs are visible. These results identify workspace state as an experimental variable in knowledge-work agents and motivate benchmarks that score evidence, staged edits, and submitted artifacts as explicit state transitions.
摘要:AI 代理人越來越多地執行知識工作(即,產生和修改持久的數位工件,如代碼庫、文件、電子表格、簡報、報告),然而他們所搜尋的解析視圖、編輯的原始文件、審查的變更以及提交的工件可能指的是同一工作產品的不同版本。我們將此表述為工作區狀態合約:每個視圖應明確與不斷演變的工作區狀態的某個版本相關聯。編碼代理人部分通過針對搜索、差異和測試的庫合約來滿足這一需求,而對於 PDF、電子表格、簡報、筆記本和混合格式的項目文件夾,類似的合約則不那麼明確。我們提出了 StagedWorkspace,一個針對知識工作代理人的版本化工作區。該工作區將解析記錄和審查差異綁定到隨原始文件變更的內容哈希。在 OfficeQA Pro 和 APEX-Agents 的固定裝置消融實驗中,雙重解析/原生訪問對於每個測試模型的最高點估計;相對於更具限制性的單一視圖,它將 OfficeQA Pass@1 提高了 8.3-12.1 分,將 APEX 的平均評分提高了 4.7-9.2 分。SW-AGENT 在 OfficeQA 上的得分為 63.9%,在 APEX 上的得分為 42.1,與已發表的同模型得分分別為 29.3% 和 25.5 相比。對 57 個文件編輯任務的配對審查軸消融進一步發現,當差異可見時,觀察到的得分更高。這些結果將工作區狀態確定為知識工作代理人的實驗變量,並激勵對證據、分階編輯和提交工件進行明確狀態轉換的基準評分。
Chain-of-Experience for Continual LLM Improvement
2608.18027v1 by Haoqin Tu, Yunhao Fang, Yizhong Wang, Cihang Xie, Shen Yan
Humans continuously learn from experience, whereas conventional large language model (LLM) evaluations ignore the models' ability to improve through inference-time interaction. In this paper, we study how LLMs learn from iterative experience at test time, a setting we refer to as Chain-of-Experience (CoE), where models accumulate experiential traces through iterative interactions with self or environmental feedback to form a continual improvement loop beyond zero-shot inference. We instantiate CoE with diverse feedback mechanisms, including model self-feedback and environmental signals such as correctness or public coding test pass rates, and evaluate across math, coding, and knowledge domains using 8 LLMs, including GPT-5, Gemini-2.5 Pro, Claude-4.5 Sonnet. Our study shows that leveraging iterative experience consistently outperforms feedback-free baselines, achieving substantial gains with self feedback alone, alongside a 5.6% overall improvement and 19% lower API cost across tasks and models. We further show that combining complementary feedback channels (e.g., model and correctness signals) yields additional gains, and that CoE delivers higher accuracy per token than existing test-time strategies. We observe a positive correlation between LLM base ability and improvement capacity, and show that models remain robust under weak or spurious feedback, with different feedback contributing to distinct improvement aspects and most gains emerging early in the iterations.
摘要:人類不斷從經驗中學習,而傳統的大型語言模型(LLM)評估則忽略了模型通過推理時互動來改進的能力。在本文中,我們研究了 LLM 如何在測試時從迭代經驗中學習,這種情境我們稱之為經驗鏈(Chain-of-Experience, CoE),在這裡模型通過與自身或環境反饋的迭代互動積累經驗痕跡,以形成超越零-shot 推理的持續改進循環。我們用多樣的反饋機制來實現 CoE,包括模型自我反饋和環境信號,如正確性或公共編碼測試通過率,並使用 8 種 LLM 進行數學、編碼和知識領域的評估,包括 GPT-5、Gemini-2.5 Pro 和 Claude-4.5 Sonnet。我們的研究表明,利用迭代經驗的表現始終優於無反饋的基準,僅依靠自我反饋就實現了顯著的增益,並在各任務和模型中達到 5.6% 的整體改進和 19% 的 API 成本降低。我們進一步顯示,結合互補的反饋通道(例如模型和正確性信號)會產生額外的增益,並且 CoE 在每個 token 上提供的準確性高於現有的測試時策略。我們觀察到 LLM 的基本能力與改進能力之間存在正相關,並顯示模型在弱或虛假反饋下仍然保持穩健,不同的反饋對不同的改進方面有所貢獻,大多數增益在迭代的早期出現。
Can Large Language Models Explain Flight Safety Events? A Prior-Guided Semantic LLM-based Approach
2608.18017v1 by Lu Xu, Xu Li, Linjiang Zheng, Fan Li, Riquan Zhang, Jiaxing Shang
Improving flight safety with flight data requires not only accurate detection of risk events, but more importantly, clear interpretation of their underlying causes at the level of pilot control behavior. Existing explainable AI techniques, such as feature importance maps, often require considerable domain knowledge to translate them into operationally meaningful explanations. Large Language Models (LLMs), which excel at language reasoning, bring a promising solution to this issue. However, applying LLMs in this domain presents key challenges such as modal inconsistency, limited classification ability, scarcity of task-specific data for fine-tuning, and lack of domain knowledge. To overcome these challenges, we propose FlightLLM, a prior-guided semantic LLM-based approach for interpretable flight safety analysis. Specifically, we first perform feature engineering to address modal inconsistency, combining statistical descriptors with physically meaningful flight indicators. This representation is further processed by a Semantic Discretization module, which converts abstract numerical patterns into qualitative descriptions that are more compatible with language reasoning. In addition, since LLMs are not inherently strong classifiers, CatBoost is incorporated as a statistical expert, and its prediction results are injected into the prompt as prior guidance. A contrastive few-shot learning strategy is further adopted to compensate for limited data. Finally, we design structured prompts to embed aviation-specific knowledge into the inference process. Using hard landing, a representative risk event with complex causal mechanisms, as an anchor point, we evaluate FlightLLM on a dataset of 704 real-world A320 flight samples. Experimental results show that the proposed approach achieves competitive classification performance while generating direct and reasonable explanations for event causes.
摘要:改善飛行安全需要不僅準確檢測風險事件,更重要的是在飛行員控制行為層面清晰解釋其潛在原因。現有的可解釋AI技術,如特徵重要性圖,通常需要相當的領域知識才能將其轉化為具有操作意義的解釋。大型語言模型(LLMs)在語言推理方面表現出色,為這一問題帶來了有希望的解決方案。然而,在這一領域應用LLMs面臨著關鍵挑戰,如模式不一致、有限的分類能力、缺乏特定任務的數據以進行微調,以及缺乏領域知識。為了克服這些挑戰,我們提出了FlightLLM,一種基於語義的先驗引導LLM方法,用於可解釋的飛行安全分析。具體而言,我們首先進行特徵工程以解決模式不一致,將統計描述符與具有物理意義的飛行指標相結合。這一表示進一步由語義離散化模塊處理,將抽象的數字模式轉換為更符合語言推理的定性描述。此外,由於LLMs本身並不是強大的分類器,因此CatBoost被納入作為統計專家,其預測結果被注入到提示中作為先驗指導。進一步採用了對比少樣本學習策略以彌補數據的有限性。最後,我們設計了結構化提示,將航空特定知識嵌入推理過程中。以硬著陸作為錨點,這是一個具有複雜因果機制的代表性風險事件,我們在704個真實世界A320飛行樣本的數據集上評估FlightLLM。實驗結果表明,所提出的方法在生成事件原因的直接和合理解釋的同時,實現了具有競爭力的分類性能。
The IOL-AI Challenge: An Open Challenge towards Advancing Linguistic Reasoning
2608.18011v1 by Eduardo Sánchez, Rita Berrada, Dan-Mircea Mirea, Sara Rajaee, Alexander Piperski, Ana Meta Dolinar, Boris Iomdin, Andrey Nikulin, Mariya Shmatova, Marzieh Fadaee, Julia Kreutzer
Reasoning in LLMs is overwhelmingly studied in domains that provide a model with rules: mathematics and code. Linguistic puzzles invert this: the solver must first discover the system before reasoning within it. We present the IOL-AI Challenge, an open-science competition run on the unseen problems of the International Linguistics Olympiad (IOL) 2026 Individual Contest, evaluated both automatically and, for the first time, by members of the official IOL Jury under the same rubrics applied to human contestants. The challenge drew 731 submissions from 46 teams under a strict compute budget (one T4, 30 mins). We additionally benchmark 15 unconstrained frontier and open models, with Claude Opus 4.8 earning a jury score equivalent to a gold medal, while both resource-constrained systems we submitted for jury grading scored in the range of the bottom 5% of contestants. Capability was not determined by scale: 14B submissions outperform models twice their size, and gains come from decoding and output-handling rather than model capacity. We also found that automatic metrics rank systems exactly as the jury does, but compress the scale, upscoring weak systems by ~13 points and understating strong ones. Our analysis shows that while frontier models might have prior knowledge about some of the problem languages, it does not significantly help them solve the linguistic reasoning tasks, leaving linguistic reasoning as a strong benchmarking proxy for generalizable reasoning skills.
摘要:推理在大型語言模型(LLMs)中的研究主要集中在提供規則的領域:數學和程式碼。語言謎題則顛倒了這一點:解題者必須首先發現系統,然後才能在其中進行推理。我們提出了IOL-AI挑戰賽,這是一項開放科學競賽,基於2026年國際語言奧林匹亞(IOL)個人賽的未見問題進行評估,這次評估既有自動評分,還首次由官方IOL評審團成員根據與人類參賽者相同的標準進行評分。這次挑戰吸引了46個團隊提交的731份作品,並在嚴格的計算預算下進行(一個T4,30分鐘)。我們還基準測試了15個不受限制的前沿和開放模型,其中Claude Opus 4.8獲得了相當於金牌的評審分數,而我們提交給評審打分的兩個資源受限系統的分數則落在參賽者的底部5%範圍內。能力並不是由規模決定的:14B的提交表現超過了規模是其兩倍的模型,並且性能的提升來自於解碼和輸出處理,而非模型容量。我們還發現,自動指標的排名與評審的排名完全一致,但壓縮了評分範圍,將弱系統的分數提高了約13分,而低估了強系統的分數。我們的分析顯示,儘管前沿模型可能對某些問題語言有先前的知識,但這並未顯著幫助它們解決語言推理任務,這使得語言推理成為通用推理能力的強基準代理。
Judge, Retrieve, or Abstain: Uncertainty-Guarded LLM Judging with Provable Risk Guarantees
2608.17994v1 by Sher Badshah, Ali Emami, Hassan Sajjad
Using LLMs as judges has become standard practice for evaluating model outputs at scale. This is particularly common for subjective, open-ended tasks such as assessing helpfulness or alignment, where no single reference answer exists. However, objective tasks introduce a distinct reliability challenge for reference-free LLM judging. In the absence of a reference answer, the judge evaluates factual correctness either through its parametric knowledge or through tool augmentation. Although the former enables efficient evaluation, the judge may hallucinate or lack sufficient evidence for its verdict. Conversely, tool augmentation can provide additional evidence but introduces extra computational cost and requires an appropriate mechanism to determine when and how that evidence should be used reliably. More importantly, neither approach alone provides formal control over the risk of accepted verdicts or guarantees their reliability at a specified level. We propose a risk-controlled framework that calibrates uncertainty thresholds on a held-out set so that the false discovery rate among accepted verdicts remains below a user-specified level~$α$ with high probability, using finite-sample Clopper--Pearson intervals. When the parametric mode is not sufficiently confident, the instance is routed to a retrieval-augmented mode, where the judge gathers web evidence and re-evaluates the instance under a second calibrated threshold. The finite-sample guarantee carries over to this two-threshold routing without additional assumptions. Across open-domain QA benchmarks and judges of varying scales, the framework maintains the target error rate while achieving substantially higher coverage than single-mode baselines.
摘要:使用大型語言模型作為評審已成為大規模評估模型輸出的標準做法。這在主觀的、開放式的任務中尤其常見,例如評估有用性或一致性,因為這類任務並不存在單一的參考答案。然而,客觀任務對於無參考的 LLM 評審引入了明顯的可靠性挑戰。在缺乏參考答案的情況下,評審通過其參數知識或工具增強來評估事實的正確性。雖然前者能夠實現高效評估,但評審可能會出現幻覺或缺乏足夠的證據來支持其裁決。相反,工具增強可以提供額外的證據,但會引入額外的計算成本,並需要適當的機制來確定何時以及如何可靠地使用這些證據。更重要的是,單獨使用這兩種方法都無法對接受的裁決風險提供正式控制或保證其在特定水平上的可靠性。我們提出了一個風險控制框架,該框架在保留集上校準不確定性閾值,以便接受的裁決中的假陽性率以高概率保持在用戶指定的水平~$α$ 以下,使用有限樣本的 Clopper--Pearson 區間。當參數模式的信心不足時,實例會被路由到檢索增強模式,在該模式下,評審收集網絡證據並在第二個校準閾值下重新評估該實例。有限樣本的保證在這個雙閾值路由中延續,無需額外假設。在開放域問答基準和不同規模的評審中,該框架在保持目標錯誤率的同時,實現了顯著高於單一模式基準的覆蓋率。
Against Political Polarization: A Unified Framework for Tracing Evolving Political Ideologies on Social Media
2608.17987v1 by Yijie Xu, Chao Wang, Hui Xiong
The rapid growth of social media has greatly influenced political discourse, highlighting the need to understand individual political ideologies and their temporal dynamics. This task faces challenges such as data scarcity, abundant non-political content, costly and bias-prone manual annotation, and difficulty in modeling future ideological inclinations. To address these issues, we propose TSN4PI, a unified framework for tracking the evolution of political ideologies on social media. It includes two core modules. The PIDN uses large language models with style transfer and unsupervised domain adaptation to enable robust ideology detection and filter irrelevant content from noisy, cross-domain data. The PIPN employs temporal graph neural networks to predict future ideological shifts, enabling comprehensive analysis of ideology presence, intensity, and evolution. We release two large-scale datasets for noncommercial research use to facilitate further work. Extensive case studies on multiple platforms (X and Truth Social) validate the effectiveness of TSN4PI and provide empirical insights into political polarization and the evolution of online ideologies. Our findings offer a nuanced perspective, advancing both methodological development and empirical understanding in this field.
摘要:社交媒體的快速增長對政治話語產生了重大影響,突顯了理解個人政治意識形態及其時間動態的必要性。這項任務面臨著數據稀缺、非政治內容豐富、昂貴且易受偏見影響的手動標註以及未來意識形態傾向建模困難等挑戰。為了解決這些問題,我們提出了TSN4PI,一個用於追蹤社交媒體上政治意識形態演變的統一框架。它包括兩個核心模塊。PIDN使用大型語言模型結合風格轉換和無監督領域適應,以實現穩健的意識形態檢測並過濾來自嘈雜的跨領域數據中的無關內容。PIPN則利用時間圖神經網絡來預測未來的意識形態變化,使得對意識形態的存在、強度和演變進行全面分析成為可能。我們釋放了兩個大型數據集供非商業研究使用,以促進進一步的研究工作。在多個平台(X和Truth Social)上進行的廣泛案例研究驗證了TSN4PI的有效性,並提供了對政治極化和在線意識形態演變的實證見解。我們的發現提供了一個細緻的視角,推進了該領域的方法論發展和實證理解。
Do Large Language Models Play Six Degrees of Separation? Measuring Topological Compression in Long-Context Manifolds
2608.17950v1 by Md. Faiyaz Abdullah Sayeedi
Large Language Models (LLMs) demonstrate remarkable multi-hop reasoning capabilities over long contexts, yet the internal mechanisms enabling these distant cognitive leaps remain poorly understood. Traditional attention-based interpretability often fails to capture true semantic proximity due to routing artifacts like attention sinks. In this paper, we bypass attention weights to directly analyze the dynamic geometry of the hidden state manifold, proving that deep LLM latent spaces natively organize into Small-World networks. By sparsifying the continuous similarity matrices of long-context representations into unweighted graphs, we trace the connectivity between highly disjoint semantic anchors across two distinct architectures. Our findings reveal a sharp topological phase transition: while early syntactic layers remain entirely fractured, deep reasoning layers abruptly compress massive conceptual distances into highly navigable pathways strictly bounded by the "Six Degrees of Separation" limit (=< 6 semantic hops). Furthermore, we demonstrate the practical efficacy of this framework by applying it to zero-shot hallucination detection within Retrieval-Augmented Generation (RAG) using the RAGognize dataset. We show that factually grounded generations maintain structural integrity with their source context (approximately 3 hops), whereas hallucinations induce severe topological collapse. Ultimately, this work mathematically formalizes how transformers execute abstract reasoning and provides a novel, strictly geometric signature for evaluating factual reliability.
摘要:大型語言模型(LLMs)在長上下文中展現出卓越的多跳推理能力,但促成這些遙遠認知飛躍的內部機制仍然不甚了解。傳統的基於注意力的可解釋性常常無法捕捉到真實的語義接近性,這是由於路由伪影如注意力匯聚所致。在本文中,我們繞過注意力權重,直接分析隱藏狀態流形的動態幾何,證明深層LLM潛在空間本質上組織成小世界網絡。通過將長上下文表示的連續相似性矩陣稀疏化為無權重圖,我們追蹤兩個不同架構之間高度不相交的語義錨點之間的連接性。我們的研究結果揭示了一個明顯的拓撲相變:儘管早期的句法層完全破碎,深層推理層卻突然將巨大的概念距離壓縮成高度可導航的路徑,這些路徑嚴格受限於「六度分隔」的限制(=< 6語義跳躍)。此外,我們通過將此框架應用於檢索增強生成(RAG)中的零樣本幻覺檢測,展示了其實際效能,使用了RAGognize數據集。我們顯示,事實基礎的生成與其來源上下文保持結構完整(約3跳),而幻覺則引發嚴重的拓撲崩潰。最終,這項工作數學化了Transformer如何執行抽象推理,並提供了一種新穎的、嚴格的幾何特徵,用於評估事實可靠性。
Efficient RLVR Scheduling via Graph-Structured Online Difficulty Estimation
2608.17941v1 by Zhizhao Liu, Zhiliang Tian, Xi Wang, Zhihua Wen, Yihang Xiong, Zhiquan Lai, Dongsheng Li
Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models but relies on costly rollout exploration. Assigning the same exploration budget to samples with different difficulty levels is inefficient: easy samples may receive redundant rollouts, whereas difficult but learnable samples may receive too little exploration. Existing adaptive schedulers address this mismatch through curriculum-based sample selection or non-uniform rollout allocation based on estimated sample difficulty. However, obtaining reliable online difficulty estimates remains challenging: dedicated probing adds substantial generation overhead, whereas history-based estimators face a cold start with no initial observations and stale feedback, and typically ignore relations among samples. To address these limitations, we propose a plug-and-play graph-based online difficulty estimator that shares rollout feedback across related samples and continuously updates their difficulty estimates, mitigating cold start and staleness without dedicated probing. Specifically, we first construct a difficulty-aware sample graph based on semantic and reasoning similarities. Based on this graph, we introduce latent difficulty states and use a Potts prior to encourage neighboring samples to share the same state. We then employ a state-level Beta-Binomial model to aggregate the rollout outcomes associated with each state. Finally, we use an online mean-field variational algorithm to continuously update the latent-state assignments and state-level difficulty as new feedback arrives. Our framework can be integrated into sample-selection and rollout-allocation schedulers, enabling difficulty-adaptive exploration without dedicated probing. Experiments across multiple base models, RL schedulers, and benchmarks demonstrate that our framework achieves better performance.
摘要:強化學習與可驗證獎勵(RLVR)提升了大型語言模型的推理能力,但依賴於成本高昂的展開探索。將相同的探索預算分配給不同難度級別的樣本是低效的:簡單樣本可能會收到冗餘的展開,而難度較高但可學習的樣本可能會收到過少的探索。現有的自適應調度器通過基於課程的樣本選擇或根據預估樣本難度的非均勻展開分配來解決這一不匹配。然而,獲得可靠的在線難度估計仍然具有挑戰性:專門的探測增加了可觀的生成開銷,而基於歷史的估計器面臨著沒有初始觀察和過時反饋的冷啟動問題,並且通常忽略樣本之間的關係。為了解決這些限制,我們提出了一種即插即用的基於圖的在線難度估計器,該估計器在相關樣本之間共享展開反饋,並持續更新它們的難度估計,減輕冷啟動和過時問題,無需專門的探測。具體而言,我們首先根據語義和推理相似性構建一個難度感知樣本圖。基於這個圖,我們引入潛在的難度狀態,並使用Potts先驗來鼓勵相鄰樣本共享相同的狀態。然後,我們使用狀態級的Beta-Binomial模型來聚合與每個狀態相關的展開結果。最後,我們使用在線均場變分算法來持續更新潛在狀態分配和狀態級難度,隨著新反饋的到來。我們的框架可以集成到樣本選擇和展開分配調度器中,實現難度自適應探索,而無需專門的探測。在多個基礎模型、RL調度器和基準測試中的實驗表明,我們的框架實現了更好的性能。
Collective Counterfactual Planning: Coordination, Consent, and Verification under Representational Constraints
2608.17932v1 by Chainarong Amornbunchornvej
Groups routinely complete projects that no single member can plan, execute, or verify alone. We propose a formal model of this phenomenon, Collective Counterfactual Planning (CCP), in which the binding limitation on each agent is neither capability, knowledge, nor observability, but representational geometry: each agent perceives the state, conceives moves, consents to actions, and certifies goal requirements only through a projection onto an agent-specific subspace of a common task space. Four gates jointly determine whether a team can reach a conjunctive goal and legitimately recognize that it has done so: the exogenous implementation coalitions required to perform each action, together with three representational gates -- conception, consent, and task-relative verification qualification. We define the Collective Counterfactual Solvability (CCS) problem, separating geometric feasibility, executable attainment, and validated completion. The results expose a positive-negative duality. Iterated cross-agent relay can unlock a solution that no one-shot pooling of individual plans contains, but any goal requirement depending essentially on the subspace dark to the entire team is unverifiable and therefore not validly completable, even when the trajectory accidentally attains it. Memoryless and audited consent further constrain different objects -- action directions versus cumulative trajectory states -- and neither dominates the other. A four-step exhaustive horizon-bounded solvability scheme is sound and complete under exact representation of the relay closure; restricted implementations remain sound on returned plans but need not be complete. The model gives one geometry for sequential mutual enabling, competent execution of steps whose purpose is invisible to the executor, forced sub-teaming at expertise boundaries, and completion that cannot be validly declared.
摘要:團體經常完成單一成員無法獨自計劃、執行或驗證的項目。我們提出這一現象的正式模型,稱為集體反事實規劃(CCP),在這個模型中,每個代理的約束限制既不是能力、知識,也不是可觀察性,而是表徵幾何:每個代理僅通過投影到共同任務空間的代理特定子空間來感知狀態、構思行動、同意行為和認證目標要求。四個閘門共同決定一個團隊是否能夠達成聯合目標並合法地認識到它已經達成:執行每個行動所需的外生實施聯盟,以及三個表徵閘門——構思、同意和任務相對驗證資格。我們定義了集體反事實可解性(CCS)問題,將幾何可行性、可執行達成和驗證完成分開。結果揭示了一種正負對偶性。迭代的跨代理中繼可以解鎖一個單次個人計劃無法包含的解決方案,但任何本質上依賴於對整個團隊來說是黑暗的子空間的目標要求都是不可驗證的,因此無法有效完成,即使軌跡意外達成了它。無記憶和經審核的同意進一步限制了不同對象——行動方向與累積軌跡狀態——而且兩者不相互主導。一個四步的全面邊界可解性方案在中繼閉包的精確表徵下是健全且完整的;受限的實施在返回的計劃上仍然是健全的,但不必是完整的。該模型為順序相互啟用、執行目的對執行者不可見的步驟的能力執行、在專業邊界強制子團隊以及無法有效宣告的完成提供了一種幾何。
Analysis of Types of Inquiries in Student-AI Interaction: A case study of two CS2 tasks
2608.17919v1 by Matin Amoozadeh, Amin Alipour
Background and Context: Question and inquiry are integral parts of knowledge seeking and learning. Despite their importance, students tend not to ask enough questions in the classroom. However, studies have shown that students interact extensively with generative AI systems for learning and problem solving. Objective: In this paper, we seek to better understand the types of questions that students ask AI systems, and how those questions evolve during problem solving and across tasks. Method: We use the Graesser et al. taxonomy to classify students' inquiries into 18 types. We develop a few-shot learning approach to automatically classify students' interactions with AI into these categories. We use this system to analyze 830 interactions of CS2 students across two programming tasks. Findings: Our results suggest that a small subset of question types accounts for the majority of student inquiries, and that the types of questions students ask change substantially as the task progresses.
摘要:背景與背景:提問和探究是尋求知識和學習的重要部分。儘管它們的重要性,學生在課堂上往往不會提出足夠的問題。然而,研究顯示學生在學習和解決問題時,與生成式人工智慧系統的互動非常廣泛。
目標:在本文中,我們旨在更好地理解學生向人工智慧系統提出的問題類型,以及這些問題在解決問題和不同任務中的演變。
方法:我們使用Graesser等人的分類法將學生的提問分為18種類型。我們開發了一種少量學習方法,自動將學生與人工智慧的互動分類到這些類別中。我們使用這個系統分析830次CS2學生在兩個編程任務中的互動。
發現:我們的結果表明,小部分問題類型佔據了學生提問的主要部分,並且學生提出的問題類型在任務進行過程中有顯著變化。
AutoResearch: Insight In, Hallucination Out
2608.17906v1 by Yiming Ren, Xiang Liu, Qumeng Sun, Xiao Zhang, Jiahao Li, Haoyang Zhang, Junjie Wang
Autonomous research systems are increasingly capable of executing long research workflows, yet automation alone does not ensure that the resulting process remains scientifically grounded. We introduce AutoResearch, a two-stage system that connects Idea Generation with Idea Execution to address both how research ideas are formed and how they are reliably established through experimentation. In Idea Generation, AutoResearch continuously integrates emerging research signals with accumulated domain knowledge, identifies transferable mechanistic insights, and uses multi-model generation and cross-review to produce grounded, testable research plans. In Idea Execution, coordinated agents decompose these plans into experiments, iteratively implement and diagnose them, and employ independent evidence-based review before accepting research conclusions. Across representative settings in cross-modal retrieval, systems optimization, and benchmark-driven machine learning, AutoResearch turns generated ideas into measurable progress, detects and corrects unreliable experimental results, and makes evidence-conditioned decisions to continue, revise, or terminate research directions. For example, on RSICD benchmark, an AutoResearch-generated idea improves mean Recall from 32.84 to 34.69, while recording only 5 audit-confirmed issue events compared with 11-27 for other autonomous research systems. These results demonstrate a research process in which meaningful insight is grounded before experimentation and conclusions are grounded before acceptance: Insight In, Hallucination Out.
摘要:自主研究系統越來越能夠執行長期的研究工作流程,然而僅僅依賴自動化並不能確保所產生的過程保持科學基礎。我們介紹了 AutoResearch,一個兩階段的系統,將創意生成與創意執行連接起來,以解決研究想法是如何形成的,以及如何通過實驗可靠地建立這些想法。在創意生成階段,AutoResearch 持續整合新興的研究信號與累積的領域知識,識別可轉移的機制見解,並利用多模型生成和交叉審查來產出有根據、可測試的研究計劃。在創意執行階段,協調的代理將這些計劃分解為實驗,迭代實施和診斷它們,並在接受研究結論之前進行獨立的基於證據的審查。在跨模態檢索、系統優化和基準驅動的機器學習等代表性設置中,AutoResearch 將生成的想法轉化為可衡量的進展,檢測並修正不可靠的實驗結果,並做出基於證據的決策以繼續、修訂或終止研究方向。例如,在 RSICD 基準上,AutoResearch 生成的想法將平均召回率從 32.84 提高到 34.69,同時僅記錄了 5 次經審核確認的問題事件,而其他自主研究系統則記錄了 11-27 次。這些結果展示了一個研究過程,其中有意義的見解在實驗之前就已經建立,而結論在接受之前也已經有根據:見解進,幻覺出。
BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models
2608.17895v1 by Liubov Chubarova, Alexandra Kuleshova, Daniil Volkov, Kirill Sultanov, Alexey Zaytsev
While Multimodal Large Language Models (MLLMs) have made significant strides in visual comprehension, their ability to reason about text-dense, professional documents remains incompletely evaluated. Existing benchmarks emphasize information extraction, require external domain knowledge, or cover professional documents only as one of many settings. They are also largely English- or Chinese-centric, leaving other languages and Russian, in particular, substantially underrepresented. To address these limitations, we introduce BEAR-Bench (Bilingual Enterprise and Academic Reasoning), a self-contained, complex English-and-Russian benchmark comprising 1000 human-annotated questions based on text-rich business and scientific documents. We evaluate 16 proprietary and open-weight MLLMs, including Gemini 3.1 Pro and Qwen3.5-397B, on BEAR-Bench and observe clear headroom even for the strongest systems. Finally, we use the resulting model outputs to compare existing hallucination detection methods, evaluating not only how often models fail on BEAR-Bench but also how reliably those failures can be identified.
摘要:雖然多模態大型語言模型(MLLMs)在視覺理解方面取得了重大進展,但它們對於文本密集型的專業文件的推理能力仍未得到充分評估。現有的基準強調信息提取,需要外部領域知識,或者僅將專業文件作為眾多設置之一。這些基準在很大程度上以英語或中文為中心,使其他語言,特別是俄語,顯得大幅度不足。為了解決這些限制,我們推出了BEAR-Bench(雙語企業與學術推理),這是一個自包含的、複雜的英語和俄語基準,包含1000個基於文本豐富的商業和科學文件的人類標註問題。我們在BEAR-Bench上評估了16個專有和開放權重的MLLMs,包括Gemini 3.1 Pro和Qwen3.5-397B,並觀察到即使對於最強的系統也存在明顯的提升空間。最後,我們使用生成的模型輸出來比較現有的幻覺檢測方法,不僅評估模型在BEAR-Bench上的失敗頻率,還評估這些失敗能否被可靠地識別。
From Global Benchmarks to Local Evaluations: Benchmarking LLMs for the German Public Sector
2608.17827v1 by Camilla Dalerci, Thilo Michael, Robin Schaefer, Daniel Weinland
Public institutions face a persistent challenge in selecting LLMs suited to their specific context. Existing benchmarks, however, are of limited use as they primarily reflect English-language and US-centric settings, and often only evaluate task performance. In this paper, we present first results of MÖVE, a holistic evaluation framework for the German public sector, examining three rarely considered governance dimensions: energy consumption, provider transparency, and knowledge of German-party positions. Our results reveal significant trade-offs, with no single model excelling across all dimensions: estimated energy consumption varies more than 60-fold and is not explained by model size alone, information disclosure varies systematically across providers, and European models do not exhibit stronger knowledge of German party positions. Model selection for public institutions thus cannot rely on performance rankings alone. Instead, evaluations should also reflect the governance requirements of the deployment context.
摘要:公共機構在選擇適合其特定情境的LLM時面臨持續的挑戰。
然而,現有的基準測試用途有限,因為它們主要反映英語和美國中心的環境,且通常僅評估任務表現。
在本文中,我們呈現MÖVE的初步結果,這是一個針對德國公共部門的整體評估框架,檢視三個鮮少考慮的治理維度:能源消耗、供應商透明度和對德國政黨立場的了解。
我們的結果揭示了顯著的權衡,沒有單一模型在所有維度上表現優異:估計的能源消耗變化超過60倍,且僅以模型大小無法解釋,信息披露在不同供應商之間系統性變化,歐洲模型對德國政黨立場的了解並未顯示出更強的優勢。
因此,公共機構的模型選擇不能僅依賴於性能排名。
相反,評估還應反映部署情境的治理要求。
Interpretable Humans, Alien LLMs: Expert Analysis of Latent Structures in Assessment Responses
2608.17810v1 by Alona Strugatski, Licol Zeinfeld, Jason Cooper, Shelley Rap, Gil Schwarts, Giora Alexandron
The evaluation of large language models (LLMs) relies heavily on human-designed assessments, implicitly assuming that AI and humans employ similar underlying cognitive constructs. Challenging this assumption, we investigate whether the latent factors governing LLM performance carry the same substantive, human-interpretable meaning as the cognitive constructs governing human learners. Using responses from humans and six LLMs across quantitative reasoning and chemistry assessments, we conducted Exploratory Factor Analysis (EFA) separately for both groups. Subject-Matter Experts (SMEs) then blindly evaluated the resulting factor graphs to ascribe pedagogical meaning to the emerged constructs. SMEs successfully interpreted most of the human-derived factors. Conversely, they could not ascribe meaning to any LLM-derived factors in quantitative reasoning and interpreted only half of the LLM factors in chemistry. By combining data-driven EFA with blind expert interpretation, this framework shows that LLMs frequently operate on statistically opaque mechanisms distinct from human reasoning.
摘要:大型語言模型(LLMs)的評估在很大程度上依賴於人類設計的評估,隱含假設AI和人類使用相似的基本認知結構。挑戰這一假設,我們調查了支配LLM性能的潛在因素是否具有與支配人類學習者的認知結構相同的實質性、人類可解釋的意義。利用來自人類和六個LLM在定量推理和化學評估中的反應,我們分別對這兩組進行了探索性因素分析(EFA)。主題專家(SMEs)隨後盲目評估了所產生的因素圖,以賦予出現的結構教學意義。SMEs成功解釋了大多數人類衍生的因素。相反,他們無法為任何LLM衍生的因素在定量推理中賦予意義,並且只解釋了化學中一半的LLM因素。通過將數據驅動的EFA與盲專家解釋相結合,這一框架顯示LLMs經常在與人類推理不同的統計不透明機制上運作。
Whether LLMs Can Navigate Beliefs and Facts Depends on How You Phrase It
2608.17809v1 by Quang Minh Nguyen, Luis Frentzen Salim
Humans naturally form and express beliefs in daily communication, e.g., "I think the answer is 3" or "I suppose that's right." Such beliefs inevitably intertwine with fact and knowledge, making the ability to handle them in tandem desirable for large language models (LLMs), as they are increasingly deployed in user-facing settings. Prior work showed that even capable LLMs exhibit a systemic weakness in acknowledging user beliefs grounded in incorrect information. We extend this evaluation to 10 LLMs across 18 epistemic expressions and find that the size and direction of the weakness depend on the verb used to express the belief, with the accuracy gap between factual and false information ranging from +50% on "I vaguely remember" to -14% on "I seriously doubt". We further show that the phenomenon stems from task confusion: models default to fact-checking the underlying claim, overriding the user's stated belief; chains of thought that explicitly fact-check show lower accuracy on false information than those that do not; and a single instruction can reverse the failure across verb families. Mechanistically, models attend more to false beliefs they fail to confirm, but suppressing this attention at decoding time recovers accuracy only partially and only in some models, calling for future work on intervention methods. Our findings clarify prior results and show how fact-checking, a generally desirable behavior, can interfere with belief tracking in LLMs. Our code is available at https://github.com/ngqm/belief-fact-phrasing.
摘要:人類在日常交流中自然地形成和表達信念,例如「我認為答案是3」或「我想這是對的」。這些信念不可避免地與事實和知識交織在一起,使得同時處理它們的能力對大型語言模型(LLMs)來說變得可取,因為它們在面向用戶的環境中越來越多地被部署。先前的研究顯示,即使是能幹的LLMs在承認基於錯誤信息的用戶信念方面也存在系統性的弱點。我們將這一評估擴展到18種認識表達下的10個LLMs,發現弱點的大小和方向取決於用來表達信念的動詞,事實信息與虛假信息之間的準確性差距從「我模糊地記得」的+50%到「我嚴重懷疑」的-14%不等。我們進一步表明,這一現象源於任務混淆:模型默認檢查基礎主張的事實,覆蓋用戶所表達的信念;明確進行事實檢查的思維鏈在虛假信息上的準確性低於那些不進行檢查的;而單一指令可以逆轉動詞家族中的失敗。在機制上,模型對它們未能確認的虛假信念的注意力更高,但在解碼時抑制這種注意力僅能部分恢復準確性,且僅在某些模型中有效,這呼籲未來對干預方法的研究。我們的發現澄清了先前的結果,並顯示事實檢查這一通常可取的行為如何干擾LLMs中的信念追蹤。我們的代碼可在 https://github.com/ngqm/belief-fact-phrasing 獲得。
An Empirical Study of Reward Specification and Benchmark Reliability in GRPO-based LLM Unlearning
2608.17804v1 by Rubén Balbastre, Juan Manuel Orduña, Mariano Pérez
Practical LLM unlearning is usually evaluated through two objectives: suppress target-specific knowledge and preserve non-target utility. In generative QA, this leaves a third behavior underspecified: when a target-adjacent prompt admits a broader answer without target-specific leakage, the model should answer at that level rather than leak, evade, or refuse. We study this specification problem in a controlled LoRA-GRPO RWKU setting, comparing four reward designs that span lexical suppression, anti-refusal shaping, rubric-based broad answering, and an explicit refusal contrast, with and without SFT warm-up. The experiments show that optimization success is not equivalent to behavioral unlearning: RWKU forget scores, held-out completion audits, terminal training-rollout audits, and training dynamics can point to different conclusions. We trace these disagreements to reward-hacking endpoints, policy-support limits in GRPO, benchmark probes that miss endpoint changes, and rewards that can select broad-topic answering with low semantic leakage during optimization.
摘要:實際的 LLM 忘記通常通過兩個目標來評估:抑制特定目標的知識和保留非目標的效用。
在生成性問答中,這留下了第三種行為未明確規範:當一個與目標相近的提示允許更廣泛的回答而不泄露特定目標時,模型應該在該層次上回答,而不是泄露、逃避或拒絕。
我們在一個受控的 LoRA-GRPO RWKU 設定中研究這個規範問題,比較四種獎勵設計,涵蓋詞彙抑制、反拒絕塑造、基於標準的廣泛回答以及明確的拒絕對比,並且有無 SFT 熱身。
實驗表明,優化成功並不等同於行為上的忘記:RWKU 忘記分數、保留的完成審計、終端訓練回滾審計和訓練動態可能指向不同的結論。
我們將這些分歧追溯到獎勵駭客端點、GRPO 中的政策支持限制、錯過端點變化的基準探針,以及在優化過程中可以選擇廣泛主題回答且語義泄露低的獎勵。
Neuro-symbolic learning over OWL 2 DL via consequence-based compilation to differentiable circuits
2608.17741v1 by Olga Mashkova, Asaad Mohammedsaleh, Fernando Zhapa-Camacho, Robert Hoehndorf
OWL 2 DL ontologies, grounded in the description logic $\mathcal{SROIQ}$, express large knowledge bases in biomedicine and the Semantic Web. Neuro-symbolic (NeSy) learners over description logics either embed the ontology in a continuous space, abandoning classical entailment, or restrict to the Horn fragment $\mathcal{EL}^{++}$, which has a single canonical model. We present Baobab, which compiles a $\mathcal{SROIQ}$ ontology with a finite ABox into a Sentential Decision Diagram (SDD): it saturates a propositional core under a consequence-based calculus and instantiates the remaining $\mathcal{SROIQ}$ features (nominals, number restrictions, and the role axioms) over the active domain. The SDD's evidence-conditioned weighted model count then trains a perception network to recognize real images under partial ABox supervision: on an ontology that exercises every distinctive $\mathcal{SROIQ}$ feature, a CNN learns to read MNIST digits coupled by a successor relation and recovers latent ontology concepts that an independent perception leaves at chance. When the supervision admits several ontology-consistent completions, an independent perception collapses onto one, a reasoning shortcut: we show that a mixture indexed by the query's justifications can represent the calibrated posterior no independent perception can, and that seeding it from the circuit's enumerated completions attains the Bayes-optimal posterior on a real-image MNIST task where single-WMC and learned mixtures (the BEARS-ensemble hypothesis class) do not: to our knowledge the first to characterize and mitigate reasoning shortcuts in a non-Horn description logic. Soundness of the compiler and the representation result are machine-checked in Lean 4. Code is available at https://github.com/bio-ontology-research-group/baobab.
摘要:OWL 2 DL 本體,基於描述邏輯 $\mathcal{SROIQ}$,在生物醫學和語意網中表達大型知識庫。神經符號(NeSy)學習者在描述邏輯上要麼將本體嵌入連續空間,放棄傳統的推理,要麼限制於只有一個典範模型的 Horn 片段 $\mathcal{EL}^{++}$。我們提出了 Baobab,它將具有有限 ABox 的 $\mathcal{SROIQ}$ 本體編譯為句子決策圖(SDD):它在基於結果的計算下飽和一個命題核心,並在活動域上實例化剩餘的 $\mathcal{SROIQ}$ 特徵(名詞、數量限制和角色公理)。SDD 的證據條件加權模型計數然後訓練一個感知網絡,以在部分 ABox 監督下識別真實圖像:在一個行使每個獨特 $\mathcal{SROIQ}$ 特徵的本體上,CNN 學會閱讀與後繼關係相結合的 MNIST 數字,並恢復獨立感知所留下的潛在本體概念。當監督允許多個本體一致的完成時,獨立感知會崩潰到一個,這是一種推理捷徑:我們展示了一種由查詢的正當性索引的混合可以表示經過校準的後驗,而沒有獨立感知可以做到,並且從電路的列舉完成中種子達到在一個真實圖像 MNIST 任務上的貝葉斯最佳後驗,而單一 WMC 和學習的混合(BEARS-ensemble 假設類)則無法做到:據我們所知,這是第一次在非 Horn 描述邏輯中表徵和減輕推理捷徑。編譯器的健全性和表示結果在 Lean 4 中經過機器檢查。代碼可在 https://github.com/bio-ontology-research-group/baobab 獲得。
What Aggregate Scores Miss: Measuring Item-Level Regressions in Commercial LLM API Migrations
2608.17719v1 by Xiaonan Xu, Wenjing Wu
Context: Software systems that depend on commercial large language model APIs must migrate to successor versions when vendors deprecate older models. Migration decisions typically rely on aggregate benchmark scores, which compress heterogeneous item-level behaviour into a single net figure. Objective: We measure what that compression conceals. Method: On three pairwise upgrades in the GPT-5.4 to GPT-5.6 Sol product sequence, we query 900 public benchmark items (graduate-level knowledge, olympiad mathematics, instruction following) 50 times per item per model, classify each item as reliably improved, reliably regressed, practically equivalent, or inconclusive under false-discovery-rate control and a practical-significance threshold, and calibrate the results against a label-permutation null. Results: Across all nine migration-benchmark cells, reliable improvements and reliable regressions coexist. Edges with aggregate gains of up to 7.3 percentage points contain up to 8.3% reliably regressed items; edges with aggregate losses contain up to 10.7% reliably improved items. On the instruction-following benchmark, the gap between strict and loose scoring widens by 3.9 percentage points on the latest migration: a 3.9-point regression under strict scoring shrinks to 0.04 points under loose scoring. Conclusion: Migration decisions based on aggregate scores alone miss substantial bidirectional item-level change. The complete response-level archive and per-item scoring outputs are released.
摘要:背景:依賴商業大型語言模型 API 的軟體系統必須在供應商棄用舊模型時遷移到後繼版本。
遷移決策通常依賴於綜合基準分數,這將異質的項目級行為壓縮為單一的淨數字。
目標:我們測量這種壓縮所隱藏的內容。
方法:在 GPT-5.4 到 GPT-5.6 Sol 產品序列的三次成對升級中,我們對 900 個公共基準項目(研究生級知識、奧林匹克數學、指令遵循)進行每個模型每項 50 次查詢,並根據假發現率控制和實際顯著性閾值將每個項目分類為可靠改進、可靠退步、實際等效或不確定,並將結果與標籤置換無效進行校準。
結果:在所有九個遷移基準單元中,可靠的改進和可靠的退步共存。
具有高達 7.3 個百分點的綜合增益的邊緣包含高達 8.3% 的可靠退步項目;具有綜合損失的邊緣包含高達 10.7% 的可靠改進項目。
在指令遵循基準上,最新遷移中嚴格與寬鬆評分之間的差距擴大了 3.9 個百分點:在嚴格評分下的 3.9 點退步在寬鬆評分下縮小至 0.04 點。
結論:僅根據綜合分數作出的遷移決策忽略了實質的雙向項目級變化。
完整的響應級存檔和每項的評分輸出已發布。
Communicating Credit Risk with Large Language Models: Evaluation of Explanations from Standard and Alternative Data-Based Models
2608.17715v1 by Sahab Zandi, Noah Kostesku, Christophe Mues, María Óskarsdóttir, Cristián Bravo
Credit decisioning is a high-stakes task in which model outputs must be accurate and explainable to support compliant decisions. Although modern credit risk models such as eXtreme Gradient Boosting (XGBoost) and Graph Neural Networks (GNNs) improve predictive performance, their explanations are often too technical for stakeholders creating communication gaps that can shape approvals, denials, and fairness judgments. We examine whether Large Language Models (LLMs) can serve as explanation layers that translate post-hoc explanation artefacts into stakeholder-appropriate risk narratives. Using Freddie Mac single-family loan-level data, we develop three pipelines: standard tabular (XGBoost + SHAP), and two with alternative data, a pure network-based (GNN + GNNExplainer), and a bimodal one (combining tabular and network data). We generate narratives with three LLM configurations: a small fine-tuned LLM (Gemma 3 4B), a large fine-tuned LLM (DeepSeek R1 70B), and a zero-shot commercial LLM (Gemini 2.5). Explanation quality is evaluated through automated checks across all pipelines and a human study of bimodal explanations comparing credit risk professionals and non-professionals on eight decision-relevant dimensions. We have three main findings. First, the pipeline accounts for higher variance in evidence-grounding scores than the language model, meaning that the binding constraint on explanation quality is the evidence representation, not the model used. Second, the explanation narratives reliably name the influential factors but are less reliable when stating the direction of influence, which may be consequential for adverse-action communication. Finally, professionals apply stricter evidentiary standards than non-professionals. We discuss implications for the governance of risk models, including deployment considerations and the value of domain-aligned LLMs in regulated credit settings.
摘要:信用決策是一項高風險的任務,其中模型輸出必須準確且可解釋,以支持合規的決策。儘管現代信用風險模型如極端梯度提升(XGBoost)和圖神經網絡(GNNs)提高了預測性能,但它們的解釋往往對利益相關者來說過於技術性,造成溝通差距,這可能影響批准、拒絕和公平性判斷。我們檢視大型語言模型(LLMs)是否可以作為解釋層,將事後解釋產物轉化為適合利益相關者的風險敘事。使用Freddie Mac的單戶貸款數據,我們開發了三個管道:標準表格(XGBoost + SHAP),以及兩個使用替代數據的管道,一個是純基於網絡的(GNN + GNNExplainer),另一個是雙模的(結合表格和網絡數據)。我們使用三種LLM配置生成敘事:一個小型微調LLM(Gemma 3 4B),一個大型微調LLM(DeepSeek R1 70B),以及一個零樣本商業LLM(Gemini 2.5)。通過對所有管道的自動檢查以及對雙模解釋的人工研究,我們評估了解釋質量,並比較了信用風險專業人員和非專業人員在八個與決策相關的維度上的表現。我們有三個主要發現。首先,該管道在證據基礎分數的變異性上比語言模型更高,這意味著解釋質量的約束是證據表示,而不是所使用的模型。其次,解釋敘事可靠地命名了影響因素,但在陳述影響方向時可靠性較低,這對於不利行動的溝通可能具有重要意義。最後,專業人士應用的證據標準比非專業人士更為嚴格。我們討論了風險模型治理的影響,包括部署考量和在受監管的信用環境中領域對齊的LLMs的價值。
GraphWake: Group Polarization via Memory-Mediated Polarization Cascade in LLM-Agent Communities
2608.17665v1 by Haoran Bu, Zejian Chen, Litian Zhang, Xi Zhang
LLM-driven agents can autonomously exchange opinions on online platforms and form communities. Such agent-operated social platforms raise a new security concern: attackers may manipulate agents to induce group polarization. Existing methods manipulate agent prompts or construct echo chambers, both of which are difficult to realize in practice. We therefore formulate a new threat, Memory-Mediated Polarization Cascade, which uses agent memory as a persistence channel and public discussion as a propagation channel. This threat contains three stages. During exposure and memory retention, the attacker exposes a small set of target agents to arguments that reinforce their respective stated stances. The targets' memory systems then process and retain these arguments. During retrieval and reproduction, a shared stance-neutral discussion cues the targets to retrieve and reproduce their respective retained arguments. During iterative propagation, untreated agents influenced by the reproduced arguments restate and spread them. We instantiate this threat in GraphWake with three components: (i) stance-support argumentation knowledge graphs construct knowledge-based arguments; (ii) axiom-oriented triple selection distills them for reliable retention and reproduction; and (iii) stance-neutral memory cueing triggers concurrent retrieval and reproduction, initiating propagation. Experiments across multiple discussions and memory systems show that GraphWake substantially increases group polarization. These findings reveal a community-level polarization risk.
摘要:LLM 驅動的代理可以在在線平台上自主交換意見並形成社群。這種代理操作的社交平台引發了一個新的安全問題:攻擊者可能操縱代理以誘發群體極化。現有的方法操縱代理提示或構建回音室,這兩者在實踐中都難以實現。因此,我們提出了一種新的威脅,記憶介導的極化級聯,它利用代理記憶作為持久性通道,公共討論作為傳播通道。這一威脅包含三個階段。在暴露和記憶保留期間,攻擊者將一小組目標代理暴露於強化其各自表述立場的論點中。目標的記憶系統隨後處理並保留這些論點。在檢索和再現期間,共享的中立立場討論提示目標檢索並再現其各自保留的論點。在迭代傳播期間,受到再現論點影響的未處理代理重述並擴散這些論點。我們在 GraphWake 中實現了這一威脅,包含三個組件:(i)立場支持的論證知識圖構建基於知識的論點;(ii)公理導向的三元組選擇提煉它們以實現可靠的保留和再現;以及(iii)立場中立的記憶提示觸發同時檢索和再現,啟動傳播。多次討論和記憶系統的實驗顯示,GraphWake 顯著增加了群體極化。這些發現揭示了社群層面的極化風險。
Graph Surgery and the Do-Operator: A Precise Correspondence for Acyclic Structural Causal Models
2608.17634v1 by Satpreet Makhija
The $\operatorname{do}$-operator is described graphically by deleting arrows into its targets and functionally by replacing their mechanisms with constants. To call these operations equivalent is not yet a mathematical statement: one returns a graph and remembers only the targets, whereas the other returns mechanisms and also remembers the imposed values. We make a dependency-level comparison precise for deterministic acyclic structural causal models with finitely many endogenous variables. If $\operatorname{Graph}(F)$ extracts the dependencies of a mechanism family $F$, our main theorem is $\operatorname{Graph}(F^ι)=\operatorname{Surg}(\operatorname{Graph}(F),T_ι)$. Thus replacing target mechanisms removes exactly the dependencies removed by graph surgery. For a model $M=(G,F)$ whose graph may contain unused arrows, we characterize when the same equality holds with $G$ in place of $\operatorname{Graph}(F)$; it holds for every intervention exactly when $G$ records the dependencies of $F$ exactly. We then define the intervened model, characterize its run, show how sequential interventions combine, and prove that an outcome depends only on interventions at its actual dependency ancestors.
摘要:$\operatorname{do}$-運算子在圖形上通過刪除指向其目標的箭頭來描述,而在功能上則通過用常數替換其機制來描述。將這些操作稱為等價尚未形成數學陳述:一個返回圖形並僅記住目標,而另一個返回機制並同時記住施加的值。我們對具有有限內生變量的確定性非循環結構因果模型進行依賴層級的精確比較。如果 $\operatorname{Graph}(F)$ 提取機制家族 $F$ 的依賴關係,我們的主要定理是 $\operatorname{Graph}(F^ι)=\operatorname{Surg}(\operatorname{Graph}(F),T_ι)$。因此,替換目標機制正好去除了圖形手術所去除的依賴關係。對於一個模型 $M=(G,F)$,其圖形可能包含未使用的箭頭,我們描述何時同樣的等式在 $G$ 代替 $\operatorname{Graph}(F)$ 時成立;當且僅當 $G$ 精確記錄 $F$ 的依賴關係時,它成立。我們然後定義干預模型,描述其運行,展示如何結合序列干預,並證明結果僅依賴於其實際依賴祖先的干預。
Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges
2608.17605v1 by Syeda Faiza Ahmed, Zien Sheikh Ali, Hunzalah Hassan Bhatti, Firoj Alam, Shammur Absar Chowdhury
Conversational AI is moving beyond isolated text prompts toward sustained, multimodal interaction. In real conversations, users clarify goals, revise requests, interrupt responses, switch topics, and introduce new evidence while expecting systems to preserve context across turns. This makes multi-turn dialogue a distinct challenge requiring systems to maintain and update memory, ground responses across modalities, tools, and external knowledge, and adapt across languages and cultures. This study reviews multi-turn conversational AI across text-only dialogue, AudioLLMs and speech-native systems, multimodal and omni-modal systems, and tool-augmented agents. We organize the literature around datasets and benchmarks, modeling paradigms, training strategies, evaluation setups, and cross-cutting challenges. Our analysis shows that support for multiple modalities has advanced faster than the ability to sustain coherent interaction across a session. Despite stronger capabilities to perceive, speak, and act across modalities, current systems still struggle with persistent memory, cross-turn grounding, full-duplex interaction, robust evaluation, and cultural alignment. We conclude with a research agenda for systems that can remember, revise, ground, speak, listen, act, and adapt across turns, modalities, and cultures. (https://github.com/faiza-sfa/multiturn-conversational-ai-survey)
摘要:對話式人工智慧正在超越孤立的文字提示,朝向持續的多模態互動發展。在真實的對話中,用戶會澄清目標、修訂請求、打斷回應、切換主題並引入新證據,同時期望系統能在不同回合中保持上下文。這使得多回合對話成為一個獨特的挑戰,要求系統維持和更新記憶,跨模態、工具和外部知識進行回應的基礎,並在語言和文化之間進行適應。本研究回顧了文本對話、AudioLLMs 和語音原生系統、多模態和全模態系統以及工具增強代理的多回合對話式人工智慧。我們根據數據集和基準、建模範式、訓練策略、評估設置和跨領域挑戰來組織文獻。我們的分析顯示,對多模態的支持發展得比在一個會話中持續一致互動的能力更快。儘管在感知、說話和跨模態行動方面的能力增強,當前的系統仍然在持久記憶、跨回合基礎、全雙工互動、穩健評估和文化對齊方面面臨挑戰。我們以一個研究議程作結,旨在開發能夠記住、修訂、基礎、說話、聆聽、行動和在回合、模態和文化之間適應的系統。(https://github.com/faiza-sfa/multiturn-conversational-ai-survey)
tinyDSM: A Framework for Skill Modeling and Development for Resource-Constrained Millirobots
2608.17596v1 by Markus D. Kobelrausch, Michael Miedler, Axel Jantsch
In this study, we investigate developmental mechanisms that enable small, resource-constrained systems such as cm-sized millirobots to autonomously explore, learn, and adapt their capabilities throughout their lifespan. Reinforcement learning algorithms guide the agent's skill acquisition and adaptation through the interplay of our proposed tinyDSM, which integrates intrinsic motivation and fitness-based assessment. We strive for minimal, hard-wired skills while encouraging the open-ended development of new skills. A key emphasis in our approach is to encode minimal a-priori general knowledge, which serves as a foundational starting point for the system as it further learns system-specific dependencies from the initial knowledge provided. Thus, by design, our approach attempts to cover very generic application domains. The methodology is based on (a) developmental mechanism with intrinsic motivation, and (b) a cognitive architecture (knowledge, reasoning, learning), while (c) utilizing minimal resources. It uses a hierarchical knowledge graph and kinematic reasoners to model and evaluate simple and advanced motion related skills. In our experiments, we use a resource-constrained millirobot with a volume of 36 cm^3 with a Raspberry Pi Pico 32-bit microcontroller (RP2040) that integrates all described features and capabilities except the camera system in 9 kB. Starting with learning the most elementary motor skills the millirobot autonomously progresses from simple linear and angular movements to complex geometric patterns within 15 minutes. To complement the physical experiments, we perform a simulation-based analysis that enables systematic comparisons across learning algorithms and intrinsic motivation parameters.
摘要:在本研究中,我們探討使小型資源受限系統(如厘米級的微型機器人)能夠自主探索、學習和適應其能力的發展機制。強化學習算法通過我們提出的tinyDSM的相互作用來指導代理的技能獲得和適應,該系統整合了內在動機和基於適應度的評估。我們追求最小的硬連接技能,同時鼓勵新技能的開放式發展。我們方法的一個關鍵重點是編碼最小的先驗一般知識,這作為系統進一步從提供的初始知識中學習系統特定依賴的基礎起點。因此,我們的方法設計上試圖涵蓋非常通用的應用領域。該方法論基於(a)具有內在動機的發展機制,以及(b)一種認知架構(知識、推理、學習),同時(c)利用最小資源。它使用層次知識圖譜和運動學推理器來建模和評估簡單和高級運動相關技能。在我們的實驗中,我們使用一個資源受限的微型機器人,其體積為36 cm^3,搭載Raspberry Pi Pico 32位微控制器(RP2040),該微控制器整合了所有描述的功能和能力,除了攝像頭系統外,僅佔用9 kB。從學習最基本的運動技能開始,微型機器人自主地在15分鐘內從簡單的線性和角運動進展到複雜的幾何圖形。為了補充物理實驗,我們進行了一個基於模擬的分析,這使得能夠在學習算法和內在動機參數之間進行系統比較。
Quantifying Risk Under Evolving Uncertainty: Belief-Dependent Robustness for Safe Sequential Decision Making
2608.17574v1 by Deep Kumar Ganguly, Jan Kretinsky
How cautious should an agent be while it is still learning its environment? We propose RATTL (Risk-Adversarial Total-Reward Learning), which ties caution to epistemic uncertainty: the agent holds a Bayesian posterior over unknown dynamics and plans against a Wasserstein ambiguity set whose radius is a monotone function of that posterior. The radius contracts with evidence, so behaviour interpolates continuously between worst-case robustness and risk-neutral total-reward maximization. The design follows the duality underlying the Entropic Value-at-Risk, which converts the choice of a risk level into the choice of an ambiguity radius. We show the resulting planning problem is well posed under transience and compactness conditions, and prove a Safety Sandwich: the RATTL value lies between the uninformed robust value and the full- knowledge optimum, with a gap that vanishes as the posterior concentrates. In a canonical binary-hazard instance, the induced criterion reduces to Conditional Value-at-Risk at a level set by the posterior entropy. A worked example shows the agent deferring the efficient action until a sharp identification threshold. RATTL targets runtime safety for agents, including LLM-based systems, acting under uncertainty.
摘要:代理在學習其環境時應該多謹慎?我們提出了RATTL(風險對抗總回報學習),它將謹慎與認知不確定性聯繫起來:代理對未知動態持有貝葉斯後驗,並根據一個其半徑是該後驗單調函數的Wasserstein模糊集進行規劃。隨著證據的增加,半徑會收縮,因此行為在最壞情況的穩健性和風險中立的總回報最大化之間持續插值。該設計遵循了熵值風險的對偶性,將風險水平的選擇轉化為模糊半徑的選擇。我們顯示,所得到的規劃問題在瞬態和緊湊性條件下是良好定義的,並證明了一個安全三明治:RATTL值介於無信息穩健值和全知最優值之間,當後驗集中時,這一差距消失。在一個典型的二元危險實例中,所引入的標準簡化為在後驗熵設定的水平下的條件風險價值。一個具體的例子顯示,代理在達到明確識別閾值之前推遲了有效行動。RATTL針對在不確定性下行動的代理,包括基於LLM的系統,目標是運行時安全。
Code as Representation: A Compilable Parsing Paradigm for Academic Documents
2608.17550v1 by Rihui Jin, Jun Wang, chengyuan zhu, Liang Mingyu, Yue Gao, Li Yunxuan, Kuicai Dong, Guilin Qi, Lin Ren, Yongrui Chen, Xinbang Dai, Jiaqi Li, Tongtong Wu, Gholamreza Haffari
Academic papers are a primary carrier of scientific knowledge, yet most of this knowledge remains locked in PDFs that are optimized for human reading rather than machine use. For Multimodal Large Language Models (MLLMs), the core challenge is not only perception, but representation: scientific pages interleave text with Structured Academic Elements (SAEs) such as tables, formulas, charts, and pseudocode, whose structure, data, and logic are poorly preserved by common surrogates like Markdown. We therefore propose Compilable Academic Document Parsing (CADP), a paradigm that reconstructs a full page as contextual \LaTeX{} plus executable Python, so that structure-preserving elements and executable chart representations can be reconstructed, recompiled, and directly verified against the source page. To support this setting, we introduce CADP-Bench, an expert-verified benchmark of full academic pages containing tightly coupled text and multiple SAE types, evaluated through a re-injection compilation protocol. We further study current capabilities using SOTA MLLMs and an exploratory multi-agent baseline that incorporates common agentic techniques. Results show that even frontier models still struggle to produce high-fidelity executable reconstructions, highlighting substantial room for improvement in structure-aware scientific document parsing. CADP-Bench is released for future research.
摘要:學術論文是科學知識的主要載體,但大部分這些知識仍然鎖定在優化為人類閱讀而非機器使用的PDF中。對於多模態大型語言模型(MLLMs)來說,核心挑戰不僅在於感知,還在於表徵:科學頁面將文本與結構化學術元素(SAEs)交錯,如表格、公式、圖表和偽代碼,其結構、數據和邏輯在常見的替代品如Markdown中保存得很差。因此,我們提出可編譯學術文檔解析(CADP),這是一種將整個頁面重建為上下文 \LaTeX{} 加上可執行的Python的範式,以便結構保留的元素和可執行的圖表表示可以被重建、重新編譯並直接與源頁面進行驗證。為了支持這一設置,我們引入CADP-Bench,一個經專家驗證的完整學術頁面基準,包含緊密耦合的文本和多種類型的SAE,通過重新注入編譯協議進行評估。我們進一步研究使用SOTA MLLMs的當前能力以及一個探索性的多代理基準,該基準結合了常見的代理技術。結果顯示,即使是最前沿的模型仍然難以產生高保真度的可執行重建,突顯出結構感知的科學文檔解析有很大的改進空間。CADP-Bench已經釋出以供未來研究使用。
CoAL-RAG: A Complexity-Aware Legal Retrieval-Augmented Generation Method
2608.17536v1 by Jin Su, Zhuofeng Zhao, Huanhuan Wang, Hao Chen
Legal consultation questions exhibit multi-level complexity. A single retrieval strategy often leads to over-reasoning for simple questions and poor interpretability for complex ones, making it difficult to meet the requirements for both answer quality and efficiency in high-risk scenarios. To address this issue, this paper proposes CoAL-RAG, a complexity-aware legal retrieval-augmented generation method, which constructs a multi-dimensional evaluation mechanism based on question essence'' andretrieval consistency'' to enable adaptive routing of retrieval strategies. First, the reasoning demand is quantified according to the logical structure of the question. Then, the discrepancy between semantic retrieval and keyword retrieval is utilized to indirectly reflect problem complexity, thereby selecting the most appropriate retrieval strategy and dynamically filtering contextual information. Experimental results demonstrate that the proposed method significantly outperforms baseline models not only on Chinese legal benchmarks (SocialLawQA, LawBench) but also demonstrates strong cross-jurisdictional generalization on English datasets (LexGLUE, CaseHold). Specifically, on Chinese datasets, the BLEU score improves by 42.5\% and ROUGE-L reaches 3.6 times that of knowledge graph-based methods. On English benchmarks, CoAL-RAG maintains highly competitive accuracy, achieving an optimal balance between generation quality, deep logical reasoning, and system efficiency across different legal systems.
摘要:法律諮詢問題展現出多層次的複雜性。單一的檢索策略常常導致對簡單問題的過度推理,以及對複雜問題的可解釋性差,使得在高風險情境中難以滿足答案質量和效率的要求。為了解決這個問題,本文提出了 CoAL-RAG,一種具複雜性意識的法律檢索增強生成方法,該方法基於「問題本質」和「檢索一致性」構建了一個多維評估機制,以實現檢索策略的自適應路由。首先,根據問題的邏輯結構量化推理需求。然後,利用語義檢索與關鍵字檢索之間的差異,間接反映問題的複雜性,從而選擇最合適的檢索策略並動態過濾上下文信息。實驗結果表明,所提出的方法在中國法律基準(SocialLawQA、LawBench)上顯著超越基線模型,並且在英語數據集(LexGLUE、CaseHold)上展現出強大的跨法域泛化能力。具體而言,在中國數據集上,BLEU 分數提高了 42.5\%,而 ROUGE-L 達到知識圖譜方法的 3.6 倍。在英語基準上,CoAL-RAG 維持了高度競爭的準確性,在不同法律系統中實現生成質量、深度邏輯推理和系統效率之間的最佳平衡。
When to Review: Spaced Repetition for Continual Pre-Training of Language Models
2608.17530v1 by Alankar Atreya, Devesh Batra, Yoages Kumar Mantri, Geremy Bantug, Greig A Cowan, Raad Khraishi
Continual pre-training of large language models must acquire new information without erasing old knowledge. Existing replay methods often choose a global old/new mixture and sample uniformly, ignoring that examples differ in how quickly they are forgotten. We formulate continual pre-training as adaptive review scheduling: the training loop should decide not only how much history to replay, but which examples should return at each step. We introduce Spaced Repetition Training (SRT), a continual learning framework inspired by cognitive science, which schedules sample-rehearsal using the SuperMemo-2 (SM-2) algorithm. SRT maintains per-example review state, maps per-example perplexity to a recall-quality signal, and schedules historical examples for retention and new examples for consolidation while leaving the model, objective, and optimizer unchanged. On temporally separated Wikipedia and code corpora, SRT improves the stability-plasticity trade-off, recovering 5 to 37 percentage points of old-knowledge accuracy lost by naive continual pre-training across model scales while preserving or improving new-knowledge acquisition. At larger scale, SRT preserves broad benchmark performance that naive continual pre-training and uniform replay substantially degrade. Experiments with vision and tabular data further suggest that the scheduling principle extends beyond language when paired with an appropriate recall signal.
摘要:持續的預訓練大型語言模型必須在不抹去舊知識的情況下獲取新信息。現有的重播方法通常選擇一個全局的舊/新混合並均勻抽樣,忽略了示例在被遺忘的速度上存在差異。我們將持續預訓練公式化為自適應回顧排程:訓練循環應決定不僅是重播多少歷史,還有每一步應該返回哪些示例。我們引入了間隔重複訓練(SRT),這是一個受認知科學啟發的持續學習框架,使用 SuperMemo-2 (SM-2) 算法來排程樣本重複。SRT 維持每個示例的回顧狀態,將每個示例的困惑度映射到回憶質量信號,並在保留模型、目標和優化器不變的情況下,為保留歷史示例和鞏固新示例進行排程。在時間上分隔的維基百科和代碼語料庫上,SRT 改善了穩定性與可塑性的權衡,恢復了由天真的持續預訓練在各模型規模上損失的 5 到 37 個百分點的舊知識準確率,同時保留或改善了新知識的獲取。在更大規模下,SRT 保持了廣泛的基準性能,而天真的持續預訓練和均勻重播則大幅降低了這一性能。對於視覺和表格數據的實驗進一步表明,當與適當的回憶信號配對時,排程原則超越了語言的範疇。
Effects of Answer Format Variation on Gender Bias in Large Language Models
2608.17516v1 by Ksenia Merzlyakova, Sebastian Padó, Franziska Weeber
Gender bias or other social biases in large language models (LLMs) are frequently evaluated with question answering or survey benchmarks where the LLM needs to give a response in a predefined answer format. It is well known in survey science that the answer format has a substantial impact on answers, just as LLMs are sensitive to the prompt wording. However, to our knowledge it has not been studied yet how changes in answer format impact the measurement of gender bias in LLMs and their alignment with human response distributions. We evaluate three instruction-tuned models on the BBQ benchmark and OpinionQA survey data across closed-ended, Likert-scaled and open-ended formats, comparing bias measurement and distributional alignment under otherwise identical conditions. We find that answer format does substantially alter measured outcomes, including reversals in order rankings. These differences arise because each format elicits distinct response behaviours, such as forced-choice selection, scale-based distributions and refusal in free-text generation. Our findings highlight the importance of treating answer format as a substantive component of LLM evaluation and motivate multi-format designs for more robust model assessment.
摘要:性別偏見或其他社會偏見在大型語言模型(LLMs)中的評估,通常使用問答或調查基準,其中LLM需要以預定的答案格式給出回應。調查科學中已知答案格式對答案有重大影響,就像LLMs對提示措辭敏感一樣。然而,據我們所知,尚未研究答案格式的變化如何影響LLMs中性別偏見的測量及其與人類回應分佈的一致性。我們在BBQ基準和OpinionQA調查數據上評估了三個經過指令調整的模型,並比較了在封閉式、Likert量表和開放式格式下的偏見測量和分佈一致性,條件則保持一致。我們發現答案格式確實顯著改變了測量結果,包括排序排名的逆轉。這些差異的產生是因為每種格式引發了不同的回應行為,例如強制選擇、基於量表的分佈和在自由文本生成中的拒絕。我們的發現強調將答案格式視為LLM評估的實質性組成部分的重要性,並促使多格式設計以進行更穩健的模型評估。
Beyond FLOPs: Energy-Aware Knowledge Distillation for Sustainable LLMs on Code-Related Task
2608.17515v1 by Enrique Barba Roque, Luís Cruz, Annibale Panichella
Background: Large Language Models (LLMs) are increasingly being applied to Software Engineering (SE) tasks, achieving high accuracy across problems such as clone detection, vulnerability prediction, and code summarization. However, their high computational demands and energy consumption raise sustainability concerns and hinder their use on consumer hardware and resource-constrained platforms. A common way to report the computational cost of an LLM in the literature and industry is to use the number of Floating Point Operations (FLOPs) required to perform a pass over the network. Aims: This paper investigates the implications of energy-aware knowledge distillation for SE, aiming to improve model efficiency while maintaining performance and to determine whether FLOPs is a reliable energy-aware metric. Method: We conduct a controlled experiment using Morph, a Many-Objective Optimization-based distillation methodology, to empirically examine whether FLOPs accurately reflect energy consumption in Clone Detection and Vulnerability Prediction tasks. We extend this methodology to include energy-surrogate models that directly estimate CPU and GPU energy consumption during optimization, and we apply Morph to generative tasks using CodeT5+ for code summarization. Results: Our results show that FLOPs is not always a reliable indicator of energy consumption, and better results can be achieved by using energy-surrogate models. Distilled student models can reduce inference energy consumption by up to 90\% and memory usage by 86\%, with only modest accuracy trade-offs. Conclusions: Energy-aware knowledge distillation when guided by direct energy surrogates rather than FLOPs can improve the energy consumption, sustainability, and deployability of LLMs for SE applications, enabling efficient models on consumer hardware.
摘要:背景:大型語言模型(LLMs)越來越多地應用於軟體工程(SE)任務,在克隆檢測、漏洞預測和程式碼摘要等問題上達到了高準確率。
然而,它們的高計算需求和能量消耗引發了可持續性問題,並阻礙了它們在消費者硬體和資源受限平台上的使用。
在文獻和業界中,報告LLM計算成本的常見方法是使用執行一次網絡所需的浮點運算次數(FLOPs)。
目標:本文探討了對SE進行能量感知知識蒸餾的影響,旨在提高模型效率的同時保持性能,並確定FLOPs是否是一個可靠的能量感知指標。
方法:我們使用Morph進行了一項受控實驗,這是一種基於多目標優化的蒸餾方法,實證檢驗FLOPs是否準確反映克隆檢測和漏洞預測任務中的能量消耗。
我們擴展了這一方法,納入能量替代模型,這些模型在優化過程中直接估算CPU和GPU的能量消耗,並將Morph應用於使用CodeT5+進行程式碼摘要的生成任務。
結果:我們的結果顯示FLOPs並不總是能可靠指示能量消耗,使用能量替代模型可以獲得更好的結果。
蒸餾的學生模型可以將推理能量消耗降低高達90%,內存使用量降低86%,而準確率僅有適度的折衷。
結論:在直接能量替代模型的指導下,能量感知知識蒸餾可以改善LLMs在SE應用中的能量消耗、可持續性和可部署性,從而使消費者硬體上的模型更加高效。
SGHA: Evidence-Grounded Research Problem Discovery with Local Language Models
2608.17501v1 by Sarvesh Gharat, Junpei Komiyama
Recent efforts toward fully automated AI scientists have demonstrated that language-model agents can generate hypotheses, execute experiments, and draft scientific manuscripts. However, during the early stages of research, when research problems are formulated, these AI scientists often rely heavily on proprietary frontier models. Their proposals are shaped by opaque parametric knowledge and by literature searches conditioned on the proposals themselves. Such knowledge is effectively a black box, and this dependence makes the evidential basis and validity of generated research problems difficult to audit and leaves the process vulnerable to model-specific hallucinations and biases. Furthermore, if proprietary research materials are transmitted to external APIs, the use of these models creates confidentiality, privacy, and data-governance concerns. We introduce the Structural Gap Hypothesis Agent (SGHA), a fully automated, corpus-first research-problem discovery system that runs entirely on a local LLM. SGHA structures a scientific literature corpus into evidence-linked paper objects and a typed evidence graph, detects unresolved structural patterns across papers, screens candidate gaps before formulation, and produces traceable research-problem families. In particular, it is able to output assumptions, objectives, success criteria, and remaining ambiguities. All LLM-based components of SGHA are executed using a locally served open-weight 9B language model, without requiring proprietary frontier-model APIs. We compare SGHA with the AI Scientist-v2 idea formulation module in five machine-learning domains. Our results suggest that explicit corpus structure and evidence-constrained reasoning can support promising, inspectable research-problem formulation without relying on frontier models during generation or verification.
摘要:最近對於完全自動化的AI科學家的努力顯示,語言模型代理可以生成假設、執行實驗並撰寫科學手稿。
然而,在研究的早期階段,當研究問題被形成時,這些AI科學家往往過度依賴專有的前沿模型。
他們的提案受到不透明的參數知識和基於提案本身的文獻搜尋的影響。
這種知識實際上是一個黑箱,而這種依賴使得生成的研究問題的證據基礎和有效性難以審核,並使過程容易受到模型特定的幻覺和偏見的影響。
此外,如果專有研究材料被傳輸到外部API,使用這些模型會產生保密性、隱私和數據治理的問題。
我們介紹了結構性差距假設代理(SGHA),這是一個完全自動化的、以語料庫為首的研究問題發現系統,完全在本地的LLM上運行。
SGHA將科學文獻語料庫結構化為與證據相關聯的論文對象和類型化的證據圖,檢測論文之間未解決的結構模式,在形成之前篩選候選差距,並生成可追溯的研究問題家族。
特別是,它能夠輸出假設、目標、成功標準和剩餘的模糊性。
SGHA的所有基於LLM的組件都是使用本地提供的開放權重9B語言模型執行的,而不需要專有的前沿模型API。
我們將SGHA與AI Scientist-v2的想法形成模塊在五個機器學習領域進行比較。
我們的結果表明,明確的語料結構和基於證據的推理可以支持有前景的、可檢查的研究問題形成,而無需在生成或驗證過程中依賴前沿模型。
SAGE: Self-Evolving Storyboard Skills via Attribution-Guided Rule Evolution
2608.17468v1 by Maolin Ran, Xiaoyang Lu, Jiaqi Liu, Jian Wang, Weiwen Liu, Jianghao Lin, Yong Yu, Weinan Zhang
Storyboards turn screenplays into visual shot plans for automated short drama production. Professional storyboarding relies on tacit directorial expertise and remains an industrial bottleneck. Large language models can automate this step, but methods for supplying directing knowledge face three challenges: (1) Knowledge acquisition: the craft remains implicit in exemplars or must be written manually. (2) Knowledge refinement: authored knowledge is not evaluated against execution outcomes, and opaque generation prevents feedback attribution to the knowledge behind each decision. (3) Knowledge injection: injecting all knowledge exceeds usable context, while manual selection for every narrative group does not scale. We present SAGE (Skill with Attribution-Guided Evolution), a deployed framework that learns, attributes, evolves, and routes directing knowledge from expert demonstrations. SAGE derives rules that are independent of episode content by contrasting each training screenplay with its expert storyboard. During generation, the model records each narrative group's adopted rules. Combining these records with localized feedback enables targeted updates to individual rules. Evolved rules form scenario packages with a routing index, so each group retrieves only a bounded set appropriate to its situation without expert intervention. On 18 test episodes across three genres, SAGE scored 77.8 on a rubric validated by experts, versus 77.1 for professional directors. Deployed for 14 days on Virtual Film Studio, SAGE produced 1,344 narrative group outputs; 87.2 percent were accepted without substantive edits, and the production team recorded over 83 percent less authoring time per episode. We release PROSE, the first public dataset pairing screenplays with storyboards by professional directors across 68 episodes: https://github.com/creDreams/PROSE.
摘要:故事板將劇本轉化為自動化短劇製作的視覺拍攝計劃。專業的故事板製作依賴於隱性導演專業知識,並且仍然是產業瓶頸。大型語言模型可以自動化這一步驟,但提供導演知識的方法面臨三個挑戰:(1)知識獲取:這項技藝仍然隱含於範例中或必須手動撰寫。(2)知識精煉:創作的知識未能根據執行結果進行評估,且不透明的生成過程阻礙了對每個決策背後知識的反饋歸因。(3)知識注入:注入所有知識超出了可用的上下文,而對每個敘事群體進行手動選擇則無法擴展。我們提出了SAGE(具歸因引導演變的技能),這是一個已部署的框架,從專家示範中學習、歸因、演變和路由導演知識。SAGE通過將每個訓練劇本與其專家故事板進行對比,推導出獨立於劇集內容的規則。在生成過程中,模型記錄每個敘事群體採用的規則。將這些記錄與本地反饋結合,使得對個別規則的針對性更新成為可能。演變的規則形成具有路由索引的場景包,因此每個群體僅檢索適合其情境的有限集合,而無需專家介入。在三個類型的18個測試劇集中,SAGE在專家驗證的評分標準上得分77.8,而專業導演則為77.1。在虛擬電影工作室部署14天後,SAGE產出了1,344個敘事群體的輸出;87.2%的輸出在未進行實質性編輯的情況下被接受,製作團隊每集的創作時間減少了超過83%。我們發布了PROSE,這是第一個將劇本與專業導演的故事板配對的公共數據集,涵蓋68個劇集:https://github.com/creDreams/PROSE。
Structure-Internalized Rule Language Model for Faithful Knowledge Graph Reasoning
2608.17443v1 by Xingrui Zhuo, Jiapu Wang, Manzong Huang, Gongqing Wu, Xindong Wu
Knowledge Graph Reasoning (KGR) aims to discover latent facts by leveraging the structural evidence available in KGs, posing a challenge to the structural semantic understanding capability of KGR models. Recent studies have demonstrated that Large Language Models (LLMs) can achieve remarkable progress on KGR tasks via flexible in-context learning. However, the inherent representation inconsistency between KG structural context and LLM parametric knowledge remains inadequately addressed. This limitation prevents LLMs from effectively perceiving reasoning evidence that aligns with KG constraints, which undermines both the effectiveness and faithfulness of reasoning. We refer to this problem as reasoning evidence perception drift of LLMs over KGs. To address this problem, we propose a Structure-Internalized Rule Language Model (SIRLM), which centers on structural rule generation to couple the parametric learning of structural knowledge with the faithfulness evaluation of reasoning logic, enabling LLMs to anchor tightly to KG-grounded evidence. Specifically, we first design a Structure-Internalized Rule Generator (SIRG), which incorporates an in-context learning block augmented with a structural relation memory to coordinate structural and parametric knowledge. Furthermore, we equip SIRG with a KG tokenizer based on structural invariance learning and a neuro-symbolic reasoner based on rule-constrained message propagation. These components provide SIRG with learnable structural representations and faithful rule-execution feedback, respectively. Our SIRLM can be seamlessly integrated into standard LLM training paradigms, such as SFT and GRPO. Extensive experiments against 17 state-of-the-art KGR methods on 36 datasets demonstrate the significant superiority of SIRLM.
摘要:知識圖譜推理(KGR)旨在利用知識圖譜中的結構證據來發現潛在事實,這對KGR模型的結構語義理解能力提出了挑戰。最近的研究表明,大型語言模型(LLMs)可以通過靈活的上下文學習在KGR任務上取得顯著進展。然而,知識圖譜的結構上下文與LLM的參數知識之間固有的表示不一致性仍然未得到充分解決。這一限制阻礙了LLMs有效感知與知識圖譜約束相符的推理證據,從而削弱了推理的有效性和可靠性。我們將這個問題稱為LLMs在知識圖譜上的推理證據感知漂移。為了解決這個問題,我們提出了一種結構內化規則語言模型(SIRLM),該模型專注於結構規則生成,以將結構知識的參數學習與推理邏輯的可靠性評估相結合,使LLMs能夠緊密依賴於知識圖譜基礎的證據。具體而言,我們首先設計了一個結構內化規則生成器(SIRG),該生成器包含一個增強了結構關係記憶的上下文學習模塊,以協調結構和參數知識。此外,我們為SIRG配備了一個基於結構不變性學習的知識圖譜標記器和一個基於規則約束消息傳播的神經符號推理器。這些組件分別為SIRG提供了可學習的結構表示和可靠的規則執行反饋。我們的SIRLM可以無縫集成到標準的LLM訓練範式中,如SFT和GRPO。在36個數據集上對17種最先進的KGR方法進行的廣泛實驗顯示了SIRLM的顯著優越性。
Cognitive Graph Intelligence for Adaptive and Robust DDoS Attack Detection in Next Generation Networks
2608.17352v1 by Mohammad Arif Hossain, Yeahia Sarker, Md Jafrin Hossain, Most. Humayra Khanom Rime, Nirwan Ansari
Distributed Denial-of-Service (DDoS) attacks threaten network availability, requiring a cognitive detection process that senses traffic, infers intent, and supports an adaptive response under severe class imbalance and non-stationary conditions. This paper proposes a Graph-based Generative Adversarial Network (GraphGAN) that serves as the cognitive detection engine for this task. GraphGAN captures the relational structure among traffic flows while addressing imbalance through adversarial generation of synthetic samples. Sequential flows are converted into $k$-nearest neighbor graphs using sliding windows to preserve feature-similarity and temporal dependencies among flows. The generator learns the distribution of DDoS attacks to synthesize realistic minority samples, while a Graph Convolutional Network (GCN)-based discriminator distinguishes real from synthetic graph data. A separate GCN classifier, trained on the balanced dataset, performs the final detection decision. Evaluations on four benchmark datasets show that GraphGAN achieves superior accuracy, precision, and recall compared to state-of-the-art approaches, particularly in data-scarce scenarios. By integrating temporal graph construction, adversarial augmentation, and GCN classification, GraphGAN effectively models coordinated attack behaviors and mitigates class imbalance, providing a robust and topology-aware solution for intrusion detection in data-constrained environments.
摘要:分散式拒絕服務(DDoS)攻擊威脅網絡可用性,這需要一個認知檢測過程來感知流量、推斷意圖,並在嚴重的類別不平衡和非穩態條件下支持自適應響應。本文提出了一種基於圖的生成對抗網絡(GraphGAN),作為此任務的認知檢測引擎。GraphGAN 捕捉流量流之間的關係結構,同時通過對抗生成合成樣本來解決不平衡問題。連續流量被轉換為 $k$-最近鄰圖,使用滑動窗口來保留流量之間的特徵相似性和時間依賴性。生成器學習 DDoS 攻擊的分佈,以合成現實的少數樣本,而基於圖卷積網絡(GCN)的判別器則區分真實與合成的圖數據。另一個在平衡數據集上訓練的 GCN 分類器執行最終檢測決策。在四個基準數據集上的評估顯示,GraphGAN 在準確性、精確度和召回率方面優於最先進的方法,特別是在數據稀缺的情況下。通過整合時間圖構建、對抗增強和 GCN 分類,GraphGAN 有效地建模協調攻擊行為並減輕類別不平衡,為數據受限環境中的入侵檢測提供了一個強健且考慮拓撲的解決方案。
DeAR: Decentralized Agentic Reasoning via Capability Grounding and Collaborative Thought Navigation
2608.17282v1 by Xing Wei, Changmeng Zheng, XiaoYong Wei, Xiufen Ye, Qing Li
Existing agentic reasoning systems typically rely on centralized protocols. This design introduces routing bottlenecks and static role allocations that often fail when handling complex multimodal queries. We propose DeAR (Decentralized Agentic Reasoning), a framework that shifts from central control to autonomous peer-to-peer collaboration. DeAR is built on three mechanisms: (1) decentralized capability grounding for query-dependent agent specialization, (2) thought map navigation for targeted peer interactions, and (3) topology update for adaptive error correction. Evaluations across 9 diverse multimodal reasoning and text-based QA benchmarks indicate that DeAR consistently outperforms recent baseline methods, validating that decentralized and adaptive collaboration among agents enhances accuracy in knowledge-intensive reasoning tasks. The source code will be available at https://open_upon_acceptance.
摘要:現有的代理推理系統通常依賴於集中式協議。這種設計引入了路由瓶頸和靜態角色分配,當處理複雜的多模態查詢時,往往會失效。我們提出了 DeAR(去中心化代理推理),這是一個從中央控制轉向自主點對點協作的框架。DeAR 建立在三個機制之上:(1)去中心化的能力基礎,以實現依賴查詢的代理專業化,(2)思維地圖導航,以便進行有針對性的同行互動,以及(3)拓撲更新,以進行自適應錯誤修正。在 9 個多樣化的多模態推理和基於文本的問答基準上的評估表明,DeAR 始終優於近期的基準方法,驗證了代理之間去中心化和自適應的協作能提高知識密集型推理任務的準確性。源代碼將在 https://open_upon_acceptance 提供。
ASI-Bench: At the Dawn of Artificial Superintelligence
2608.17271v1 by Junwei Zhou, Zhen Sun, Binyu Li, Jiangyu Zhou, Yuexi Pan, Hengyu Wang, Honghe Ren, Xiaohan Jia, Xueyang Zhou, Xiaoyu Cao, Yongchao Chen, Yuanning Feng, Junhao Wu, Cheng Zhang, Sijia Chen, Haoyu Xue, Chengsong You, Huan Wang, Koutian Wu, Peigan Gao, Jiakun Wu, Wenzhe Li, Ergan Shang, Qingyuan Zheng, Jingjing Zhou, Ruixuan Jia, Yan Xu, Hongrui Zhang, Xiao-Han Ma, Zhengxiang Cheng, Yuexing Hao, Liting Mai, Xianglin Ji, Wenjun Zhang, Zhuofan Chen, Yixiao Huang, Chi Wang, Wenyue Hua, Yilun Hao, Yuantao Zhai, Ziyan Zhao, Jingyan Xie
Artificial superintelligence (ASI) requires AI to move beyond mastering existing knowledge toward exploring the unknown, creating new knowledge, and turning new ideas into verifiable results. However, the capabilities of today's AI systems are still largely built on learning, compressing, and applying existing human knowledge. Accordingly, existing benchmarks primarily test whether AI can produce correct answers based on learned knowledge, or whether it can complete tasks under extensive human guidance. We therefore introduce ASI-Bench, the first benchmark to jointly evaluate AI systems' capabilities of innovative exploration and autonomous scientific execution across general research domains, and the first to progressively withdraw human methodological guidance within the same research project to test how far AI can proceed on its own. Built by over 40 experts with the cost of 31,000+ human hours, ASI-Bench contains 60 project-level research tasks across 11 scientific domains and progressively reduces methodological guidance to test whether AI can independently select methods, conduct research, and produce verifiable results. All tasks undergo expert review, AI-assisted auditing, sandbox execution, and scorer validation. Across 18 state-of-the-art agent--model configurations, the average score drops from 50.91 with full methodological guidance to 29.10 with only the method specified and 26.62 when agents must determine the method themselves. This sharp decline shows that current systems remain heavily dependent on human guidance and are still far from autonomously conducting end-to-end, project-level scientific research. ASI-Bench is open to the world. We invite researchers and builders everywhere to contribute new tasks, challenge the limits of today's AI, and help accelerate humanity's collective path toward artificial superintelligence at https://asibench.apexin.ai/submit.
摘要:人工超智能(ASI)要求人工智慧超越掌握現有知識,朝向探索未知、創造新知識,並將新想法轉化為可驗證的結果。然而,當今人工智慧系統的能力仍然主要建立在學習、壓縮和應用現有人類知識的基礎上。因此,現有的基準主要測試人工智慧是否能根據學習到的知識產生正確答案,或者是否能在廣泛的人類指導下完成任務。因此,我們推出了 ASI-Bench,這是第一個共同評估人工智慧系統在一般研究領域中創新探索和自主科學執行能力的基準,並且是第一個在同一研究項目中逐步撤回人類方法論指導以測試人工智慧能獨立進行多遠的基準。ASI-Bench 由超過 40 位專家建造,耗費超過 31,000 小時的人力,包含 60 個跨 11 個科學領域的項目級研究任務,並逐步減少方法論指導,以測試人工智慧是否能獨立選擇方法、進行研究並產生可驗證的結果。所有任務都經過專家審查、人工智慧輔助審核、沙盒執行和評分者驗證。在 18 種最先進的代理-模型配置中,平均得分從全方法論指導下的 50.91 降至僅指定方法的 29.10,當代理必須自行確定方法時則降至 26.62。這一急劇下降顯示當前系統仍然在很大程度上依賴於人類指導,並且仍然遠未能自主進行端到端的項目級科學研究。ASI-Bench 向全世界開放。我們邀請各地的研究者和建設者貢獻新任務,挑戰當今人工智慧的極限,並幫助加速人類朝向人工超智能的共同道路,網址為 https://asibench.apexin.ai/submit。
Understanding Curriculum Learning in Large Language Models via Cross-Difficulty Optimization Dynamics
2608.17268v1 by Zhikai Ding, Ziyi Ye
Curriculum learning has been widely adopted in the post-training of large language models by organizing training data from easy to hard. However, its effectiveness varies substantially across reasoning tasks, suggesting that no single curriculum is universally optimal and raising a fundamental question: what determines when curriculum learning works? In this paper, we answer this question by analyzing the optimization dynamics induced by different curriculum schedules. We show that the transfer relationship between different difficulty levels characterizes the optimization dynamics induced by curriculum learning, which in turn explains the effectiveness of different curriculum schedules, and formalize this relationship as Relative Transfer, a principled measure of cross-difficulty knowledge transfer. Based on this measurement, we derive Transfer-aware Dynamic Curriculum Sampling (TDCS), which dynamically adjusts the sampling distribution according to the estimated transfer relationship throughout training. Extensive experiments on multiple reasoning benchmarks demonstrate that TDCS consistently outperforms representative scheduling strategies across different tasks, model scales, and training paradigms. More importantly, our work provides a unified optimization-based explanation of curriculum learning through cross-difficulty transfer.
摘要:課程學習已被廣泛應用於大型語言模型的後訓練,通過將訓練數據從簡單到困難進行組織。
然而,它在推理任務中的有效性差異很大,這表明沒有單一的課程是普遍最佳的,並提出了一個根本性問題:什麼決定了課程學習的有效性?
在本文中,我們通過分析不同課程安排所引起的優化動態來回答這個問題。
我們展示了不同難度級別之間的轉移關係特徵化了課程學習所引起的優化動態,這反過來解釋了不同課程安排的有效性,並將這一關係形式化為相對轉移,這是一種跨難度知識轉移的原則性度量。
基於這一測量,我們推導出轉移感知動態課程抽樣(TDCS),該方法根據整個訓練過程中估計的轉移關係動態調整抽樣分佈。
在多個推理基準上的大量實驗表明,TDCS在不同任務、模型規模和訓練範式中始終優於代表性的排程策略。
更重要的是,我們的工作通過跨難度轉移提供了一個統一的基於優化的課程學習解釋。
Structural Plan-to-Model Conversion with Deterministic Geometry and Guarded Agentic Vision-Language Refinement
2608.17237v1 by Mohammad Talebi-Kalaleh, Qipei Mei
Converting structural framing plans into editable finite-element model drafts remains labor-intensive and prone to transcription error. Existing drawing-understanding systems for building components rely on task-specific trained neural detectors, and language-model agents in structural engineering operate on text or model data rather than the drawing itself. This paper presents, to the authors' knowledge, the first framework applying an agentic vision-language layer to structural component detection and model drafting from framing-plan PDFs, without task-specific detector training or fine-tuning. A deterministic stage extracts primitives, estimates scale by dimension-ratio consensus, recognizes five entity classes with a drafting grammar, and assembles an editable layout. The agentic stage proposes typed corrections constrained by deterministic candidates, operation-specific admission tests, change-level review, and fail-closed transactions. Evaluation used an author-generated benchmark of 100 plans: a development half that informed every rule revision, and a seed-disjoint held-out half generated after the rules froze, evaluated once. All reported scores are end-to-end results of the complete framework on the held-out half. Scale was estimated within 0.1% of the generator reference for every drawing. Recall and precision were 0.922/0.997 for columns, 0.886/0.990 for beams, 1.000/1.000 for walls, 1.000/1.000 for braces, and 1.000/0.964 for openings. A controlled study repeated two corruptions three times on three development drawings. Calibration passed all nine trials; member repair met every strict end-state predicate in five of nine. Guarded review corrected missed framing and false marks within explicit bounds. The held-out half shares the development generator, so the study excludes independently drafted plans, raster evaluation, analytical connectivity, and solver validation.
摘要:將結構框架計劃轉換為可編輯的有限元素模型草稿仍然需要大量人力,並且容易出現轉錄錯誤。現有的建築組件繪圖理解系統依賴於特定任務訓練的神經檢測器,而結構工程中的語言模型代理則基於文本或模型數據,而非繪圖本身。據作者所知,本文提出了第一個將代理視覺-語言層應用於從框架計劃PDF中檢測結構組件和模型草擬的框架,無需特定任務的檢測器訓練或微調。一個確定性的階段提取原始元素,通過尺寸比共識估算比例,識別五種實體類別,並組裝可編輯的佈局。代理階段提出了受限於確定性候選者的類型修正、特定操作的入場測試、變更級別審查和失敗關閉交易。評估使用了一個作者生成的100個計劃的基準:一半用於開發,告知每條規則的修訂,另一半在規則凍結後生成,進行了一次評估。所有報告的分數都是在保留的一半上完整框架的端到端結果。每個繪圖的比例估算在生成參考的0.1%內。柱子的召回率和精確度為0.922/0.997,梁為0.886/0.990,牆為1.000/1.000,支撐為1.000/1.000,開口為1.000/0.964。一項控制研究在三個開發繪圖上重複了兩次損壞,進行了三次。校準通過了所有九次試驗;成員修復在九次中的五次滿足每個嚴格的最終狀態謂詞。受控審查在明確範圍內修正了漏掉的框架和錯誤標記。保留的一半共享開發生成器,因此該研究排除了獨立草擬的計劃、光柵評估、分析連通性和求解器驗證。
Synthesizing Feature Extractors: An Agentic Approach for Algorithm Selection
2608.17170v1 by Hai Xia, Carlos Ansótegui, Stefan Szeider
Algorithm selection for constraint satisfaction problems requires extracting features that capture problem structure. Manually designing feature extractors demands deep domain expertise and quickly becomes a bottleneck when new problem classes appear. We present an automated approach that uses Large Language Models (LLMs) in an agentic check--fix--verify loop to synthesize executable Python scripts that act as interpretable, problem-specific feature extractors. Given a high-level MiniZinc model and an instance, the LLM agent generates code that constructs a typed graph representation and computes structural properties such as graph density, variable clustering, and constraint tightness. We evaluate our approach on three combinatorial problems (vehicle routing, car sequencing, fixed-length error-correcting codes) with a portfolio of five state-of-the-art solvers. The synthesized extractors yield algorithm selectors that consistently outperform both expert-curated mzn2feat features (up to $8.3$ percentage points (pp) test-set accuracy on FLECC) and the best transformer-based trans2feat variants. In the meanwhile, the synthesized feature extractors remain inspectable.
摘要:算法選擇約束滿足問題需要提取捕捉問題結構的特徵。手動設計特徵提取器需要深厚的領域專業知識,並且在新的問題類別出現時很快就會成為瓶頸。我們提出了一種自動化的方法,利用大型語言模型(LLMs)在代理檢查--修正--驗證循環中合成可執行的 Python 腳本,這些腳本充當可解釋的、特定於問題的特徵提取器。給定一個高階的 MiniZinc 模型和一個實例,LLM 代理生成代碼,構建一個類型圖表示並計算結構性質,如圖密度、變量聚類和約束緊湊性。我們在三個組合問題(車輛路由、汽車排序、固定長度糾錯碼)上評估我們的方法,使用五個最先進求解器的組合。合成的提取器產生的算法選擇器在測試集準確率上始終超越專家策劃的 mzn2feat 特徵(在 FLECC 上高達 $8.3$ 個百分點(pp))和最佳的基於Transformer的 trans2feat 變體。與此同時,合成的特徵提取器仍然可供檢查。
Towards Safer RAG: Only Agents Capable of System 2 Thinking may Access Untrusted Documents
2608.17153v1 by Mehrdad Ghassabi
Retrieval-Augmented Generation (RAG) has significantly enhanced the performance of large language models (LLMs), yet these systems remain vulnerable to knowledge-poisoning attacks, in which misinformation in retrieved documents can influence the model's final outputs. Notably, an LLM may correctly detect that a document contains incorrect information while nevertheless being influenced by it. Prior work has addressed this vulnerability through the Cordon Principle, which prevents models responsible for final answer synthesis from directly accessing raw evidence. Although effective, this strict isolation can introduce substantial computational overhead. In this work, we propose a refined security principle: only agents capable of deliberative System 2 reasoning may access untrusted documents. To evaluate this principle, we introduce novel metrics that quantify the discrepancy between misinformation detection and downstream influence. We then empirically compare state-of-the-art reasoning language models with standard language models across these metrics. Our results show that reasoning-capable models are substantially more robust to corrupted evidence, without requiring the strict isolation imposed by the Cordon Principle. These findings provide empirical support for our refined principle and suggest a more practical foundation for secure RAG system design.
摘要:檢索增強生成(RAG)顯著提升了大型語言模型(LLMs)的性能,但這些系統仍然容易受到知識污染攻擊,其中檢索到的文件中的錯誤信息可能影響模型的最終輸出。值得注意的是,LLM 可能正確檢測到某個文件包含不正確的信息,但仍然會受到其影響。先前的研究通過 Cordon 原則解決了這一脆弱性,該原則防止負責最終答案合成的模型直接訪問原始證據。儘管有效,但這種嚴格的隔離可能會引入相當大的計算開銷。在本研究中,我們提出了一個精煉的安全原則:只有能夠進行深思熟慮的系統 2 推理的代理才能訪問不受信任的文件。為了評估這一原則,我們引入了新穎的度量標準,以量化錯誤信息檢測與下游影響之間的差異。然後,我們在這些度量標準上,實證比較了最先進的推理語言模型與標準語言模型。我們的結果顯示,具備推理能力的模型對受損證據的魯棒性顯著更強,而無需 Cordon 原則所施加的嚴格隔離。這些發現為我們的精煉原則提供了實證支持,並為安全 RAG 系統設計建議了一個更實用的基礎。
KnowSim: Evaluating Information Calibration in LLM Assistants with User Simulators that Learn
2608.17150v1 by Yoonjoo Lee, Hyoungwook Jin, Tae Soo Kim, Shaoyang Zhang, Philippe Laban, Q. Vera Liao
To effectively collaborate with users on knowledge-intensive tasks, Large Language Models (LLMs) must perform information calibration: matching content to a user's evolving understanding and cognitive capacity. Yet user simulators used to evaluate and train LLMs do not explicitly model user knowledge so they neither produce realistic interactions across knowledge levels nor reflect how interactions unfold as that knowledge evolves. To close this gap, we introduce KNOWSIM, an evaluation framework built around a user simulator that maintains explicit knowledge states, represented as a graph of Information Units with prerequisite relationships, that evolve under update rules grounded in learning theory. KNOWSIM computes three metrics (Knowledge Gain, Delivery Calibration, Cognitive Overload) directly from the knowledge state trajectory, reflecting key mechanistic aspects of information calibration. We validate KNOWSIM against 705 human-AI sessions across two domains, stratified by knowledge level: its rankings align significantly with human judgments (73-74% sign agreement), outperforming three baseline simulators. Applied to 9 LLMs, KNOWSIM reveals that the best model shifts by user knowledge level, revealing aptitude-treatment interactions invisible to standard evaluation.
摘要:為了有效地與用戶在知識密集型任務上合作,大型語言模型(LLMs)必須進行信息校準:將內容與用戶不斷演變的理解和認知能力相匹配。然後,用於評估和訓練LLMs的用戶模擬器並未明確建模用戶知識,因此它們既無法產生跨知識水平的現實互動,也無法反映隨著知識演變而展開的互動。為了填補這一空白,我們介紹了KNOWSIM,一個圍繞用戶模擬器構建的評估框架,該模擬器維持明確的知識狀態,這些狀態以具有前提關係的信息單元圖表示,並根據學習理論的更新規則進行演變。KNOWSIM直接從知識狀態軌跡計算三個指標(知識增益、交付校準、認知過載),反映信息校準的關鍵機制方面。我們在705個人類-人工智能會話中驗證了KNOWSIM,這些會話分為兩個領域,按知識水平分層:其排名與人類評判顯著一致(73-74%的符號一致性),並超越了三個基線模擬器。應用於9個LLMs,KNOWSIM顯示最佳模型隨用戶知識水平而變化,揭示了標準評估中不可見的能力-處理互動。
A decodability criterion predicts when hidden-state selection beats majority voting in large language models
2608.17124v1 by Zhixiang wang, Ziliang Hong, Ulas Bagci
Combining the answers a large language model (LLM) samples for a question into one decision is a test-time information fusion problem, usually solved by majority voting. Voting is unreliable on difficult questions, where the sampled answers share correlated errors, so the wrong answer can win and drawing more samples makes the decision worse. Selecting a candidate by reading a correctness signal from the model's hidden states is a promising alternative, but its accuracy varies across models and tasks, and no measure indicates when it can be trusted. In this paper, we propose CASE (Correctness-Axis SElection), a dynamic selection combiner that trains a linear gate on the answer-token hidden state and selects the highest-scoring candidate. Its main contribution is decodability, a leakage-free measure of how well the gate ranks a question's correct candidates above its incorrect ones, which predicts whether hidden-state selection will outperform voting. A conventional probe appears accurate only because of question-identity leakage, which vanishes under question-grouped evaluation. On held-out data, decodability predicts the accuracy gain of selection over voting with a Pearson correlation r=0.75 and a decision threshold near AUC=0.60. Across general and medical LLMs, CASE improves over voting by up to 19 points on medium-difficulty questions and 16.8 points on hard questions. Decodability depends on the aligned knowledge a model must recall, not on its scale, and its prediction transfers to an unseen scientific domain within 3.8 points. It thus provides a practical criterion, measurable in advance for a given model and task, for choosing between learned selection and majority voting.
摘要:將大型語言模型(LLM)對一個問題所採樣的答案合併為一個決策是一個測試時的信息融合問題,通常通過多數投票來解決。
在困難問題上,投票不可靠,因為採樣的答案共享相關錯誤,因此錯誤的答案可能會獲勝,而增加更多樣本會使決策變得更糟。
通過從模型的隱藏狀態中讀取正確性信號來選擇候選者是一個有前途的替代方案,但其準確性在不同模型和任務之間有所變化,且沒有任何指標表明何時可以信任它。
在本文中,我們提出了CASE(正確性軸選擇),這是一個動態選擇組合器,對答案標記的隱藏狀態訓練一個線性閘,並選擇得分最高的候選者。
它的主要貢獻是可解碼性,這是一種無洩漏的度量,衡量閘如何將問題的正確候選者排名高於不正確的候選者,並預測隱藏狀態選擇是否會優於投票。
傳統探測器之所以顯得準確,僅僅是因為問題身份的洩漏,而這在問題分組評估中會消失。
在保留數據上,可解碼性預測選擇相對於投票的準確性增益,皮爾森相關係數 r=0.75,決策閾值接近 AUC=0.60。
在一般和醫療 LLM 中,CASE 在中等難度問題上提高了最多 19 分,在困難問題上提高了 16.8 分。
可解碼性取決於模型必須回憶的對齊知識,而不是其規模,且其預測在未見的科學領域內轉移至 3.8 分。
因此,它為在給定模型和任務之間選擇學習的選擇和多數投票提供了一個可實際測量的標準。
From Abductive Explanations to Global Logical Rules for Node Classification in SGCs
2608.17103v1 by Bryan Lima Cavalcante, Thiago Alves Rocha
Graph Neural Networks (GNNs) have achieved remarkable performance in node classification tasks, motivating growing interest in methods capable of explaining their predictions. Recent logic-based approaches, such as LogicXGNN, derive global logical rules for Graph Neural Networks (GNNs) from collections of explanatory subgraphs. While informative, these subgraphs may contain redundant structural information that is specific to individual nodes, potentially limiting the generality of the extracted rules. In this work, we propose a logic-based framework for node classification in Simple Graph Convolution (SGC) networks that uses minimal abductive explanations as an intermediate representation for rule extraction. For each node, we compute a minimal set of node-feature pairs sufficient to preserve the predicted class. These explanations are then used to train decision trees from which global logical rules are extracted. Experiments on benchmark datasets show that the proposed framework produces compact global rules while maintaining high fidelity to the original SGC model.
摘要:圖神經網絡(GNNs)在節點分類任務中取得了顯著的表現,這激發了對能夠解釋其預測的方法的日益關注。最近的基於邏輯的方法,如LogicXGNN,從解釋性子圖的集合中推導出圖神經網絡(GNNs)的全局邏輯規則。雖然這些子圖提供了資訊,但它們可能包含特定於個別節點的冗餘結構資訊,這可能限制了提取規則的普遍性。在本研究中,我們提出了一個基於邏輯的框架,用於簡單圖卷積(SGC)網絡中的節點分類,該框架使用最小的推斷解釋作為規則提取的中介表示。對於每個節點,我們計算一組最小的節點-特徵對,這些對足以保留預測的類別。然後,這些解釋用於訓練決策樹,從中提取全局邏輯規則。在基準數據集上的實驗表明,所提出的框架生成了緊湊的全局規則,同時保持了對原始SGC模型的高保真度。
J-Miner: Recovering Executable Decision Knowledge from Language-Model Classifiers
2608.17063v1 by Yunfan Gao, Xinyi Huang, Tao Sheng, Haorui Song, Yun Xiong, Haofen Wang
Large language models can be fine-tuned into specialized classifiers that perform well across diverse text tasks and make complex judgments, but they typically expose only final labels, leaving the decision knowledge acquired through fine-tuning implicit within the model. We study how to mine this internal decision knowledge from a fine-tuned classifier and encode it in an executable representation that can be inspected, validated, and reused beyond the source classifier. We introduce J-Miner, which mines text-level named concepts by aggregating vocabulary-aligned internal signals across layers and token positions, and uses the classifier's own predictions to learn executable decision rules over them. This process distills local internal readouts into an explicit classifier-level knowledge representation. Across multiple classification tasks, J-Miner rules reproduce up to 98.3\% of source-classifier decisions and achieve 6.0--29.5 percentage points higher behavioral fidelity than equally compact rules learned from input words. Further analysis shows that the named concepts reflect internal semantic evidence associated with task decisions, while the learned rules consolidate these distributed signals into inspectable decision structures. The resulting decision knowledge also transfers to lightweight standalone students: using about 1/24 as many parameters as the source classifiers, they reconstruct and execute the representation from raw text while retaining 99.8\% of the source classifiers' mean task accuracy. These findings show that task-specific decision knowledge can be faithfully represented in an explicit, executable form and reused beyond the classifier in which it was learned.
摘要:大型語言模型可以被微調成為專門的分類器,這些分類器在多樣的文本任務中表現良好並能做出複雜的判斷,但它們通常僅顯示最終標籤,將通過微調獲得的決策知識隱含在模型內部。我們研究如何從微調的分類器中挖掘這種內部決策知識,並將其編碼為可執行的表示,這種表示可以被檢查、驗證並在源分類器之外重用。我們介紹了 J-Miner,它通過在層和標記位置之間聚合與詞彙對齊的內部信號來挖掘文本級命名概念,並利用分類器自身的預測來學習可執行的決策規則。這一過程將局部內部讀出轉化為明確的分類器級知識表示。在多個分類任務中,J-Miner 規則重現了高達 98.3\% 的源分類器決策,並比從輸入詞學習的同樣緊湊規則提高了 6.0--29.5 個百分點的行為忠實度。進一步分析顯示,命名概念反映了與任務決策相關的內部語義證據,而學習到的規則則將這些分散的信號整合為可檢查的決策結構。所產生的決策知識也可以轉移到輕量級的獨立學生模型:使用約 1/24 的參數數量,這些模型能夠從原始文本重建並執行表示,同時保留 99.8\% 源分類器的平均任務準確率。這些發現顯示,特定任務的決策知識可以以明確的、可執行的形式忠實地表示,並在學習該知識的分類器之外重用。
Cross-Model Memory Transfer via Target-Side Reader Adaptation
2608.17050v1 by Mingyuan Li, Guangsheng Yu, Xu Wang, Shaoxiong Ji
Methods for improving knowledge use in large language models typically fall into two regimes. Non-parametric retrieval offers flexible access to external knowledge, but adds retrieval latency, context overhead, and only shallow integration with the backbone. Parametric adaptation is efficient at inference time, but entangles knowledge with model weights and can be hard to update, audit, or transfer. Engram-style hashed memory occupies a middle regime: it stores learned information in an external, addressable table, yet consumes that table through a small learned reader. This raises a basic question: when such a memory is moved across backbones, what matters more, the frozen memory itself or the target-side reader? We study this question through cross-model frozen-memory extraction, in which a memory trained on a source model is frozen and attached to a different target model, with only a lightweight reader trained. Ablations show that learned memory content and correct addressing both matter, but the transferred table becomes useful only through a reader aligned to the target model. In downstream question answering tasks, a dual-layer, four-branch reader nearly closes the gap between same-model and cross-model reuse, achieving an average score of 38.8 under our controlled evaluation protocol. Moreover, when the provider reader is directly compatible with the target interface, the frozen artifact can provide substantial utility without target-side training, while optional reader adaptation yields further improvement. These results suggest that Engram can serve as a reusable external knowledge artifact, provided that the target has access to a compatible reader interface; target-side adaptation can further improve alignment when direct reader reuse is insufficient.
摘要:改善大型語言模型中知識使用的方法通常分為兩種模式。非參數檢索提供靈活的外部知識訪問,但增加了檢索延遲、上下文開銷,並且與主幹的整合僅為淺層。參數適應在推理時效率高,但將知識與模型權重糾纏在一起,並且難以更新、審計或轉移。Engram風格的哈希記憶佔據了中間模式:它將學習到的信息存儲在一個外部的、可尋址的表中,卻通過一個小型的學習讀取器來消耗該表。這引發了一個基本問題:當這樣的記憶在主幹之間移動時,哪一個更重要,凍結的記憶本身還是目標端的讀取器?我們通過跨模型凍結記憶提取來研究這個問題,在這個過程中,訓練於源模型的記憶被凍結並附加到不同的目標模型上,只有一個輕量級的讀取器被訓練。消融實驗顯示,學習的記憶內容和正確的尋址都是重要的,但轉移的表只有通過與目標模型對齊的讀取器才能變得有用。在下游問題回答任務中,一個雙層、四分支的讀取器幾乎縮小了同模型和跨模型重用之間的差距,在我們的控制評估協議下達到了38.8的平均分。此外,當提供者讀取器與目標介面直接兼容時,凍結的工件可以在不進行目標端訓練的情況下提供實質性的效用,而可選的讀取器適應則進一步提高了效果。這些結果表明,Engram可以作為可重用的外部知識工件,前提是目標能夠訪問兼容的讀取器介面;當直接的讀取器重用不足時,目標端的適應可以進一步改善對齊。
AutoSR: Automatic Symbolic Regression by Searching Research States
2608.16876v1 by Kejia Zhang, Youran Sun, Xinyu Ren, Chugang Yi, Haizhao Yang
We introduce Automatic Symbolic Regression (AutoSR), a fully automated system that instantiates Research-Space Symbolic Regression by searching persistent scientific investigations rather than isolated equations. Finite, noisy data often yield numerically competitive expressions that imply very different behavior outside the observed regime, making numerical fit and syntactic complexity insufficient measures of scientific credibility. Existing approaches largely focus on improving expressions, yet the search typically retains little beyond the resulting formula and score, losing the scientific record, such as motivations and probes, that inform what to try next. AutoSR preserves this record in a \textbf{Research State}, coupling each candidate equation with the reasoning, computational evidence, and independent review developed along its branch. Proposer--reviewer agents develop these states under progressive-widening Monte Carlo tree search (PW-MCTS), which allocates computation across competing investigations, while the accumulated research record is ultimately synthesized into a final report that explains the leading relation and the basis for its selection. Across nine selected challenges from two benchmark suites, AutoSR recovers algebraically equivalent relations in every case, including three cp3-bench problems that no published system recovers and six structurally diverse LSR-Transform problems. Overall, AutoSR extends symbolic regression from equation-level search toward automated scientific investigation, allowing scientific knowledge and accumulated evidence to shape both what is explored and how the resulting equation is justified.
摘要:我們介紹自動符號回歸(AutoSR),這是一個完全自動化的系統,它通過搜尋持續的科學研究而不是孤立的方程式來實現研究空間符號回歸。有限的、帶噪聲的數據通常會產生數值上具有競爭力的表達式,這些表達式在觀察範圍之外暗示了非常不同的行為,使得數值擬合和語法複雜性不足以作為科學可信度的衡量標準。現有的方法主要集中在改進表達式上,但搜索通常僅保留結果公式和分數,失去了科學記錄,例如動機和探測,這些記錄告訴我們接下來該嘗試什麼。AutoSR 在一個 \textbf{研究狀態} 中保留這個記錄,將每個候選方程與沿其分支發展的推理、計算證據和獨立審查相結合。提議者-審查者代理在漸進擴展的蒙特卡羅樹搜索(PW-MCTS)下發展這些狀態,該方法在競爭的研究之間分配計算,而累積的研究記錄最終被綜合成一份最終報告,解釋主要關係及其選擇的基礎。在來自兩個基準套件的九個選定挑戰中,AutoSR 在每一個案例中都恢復了代數上等價的關係,包括三個沒有任何已發表系統恢復的 cp3-bench 問題和六個結構多樣的 LSR-Transform 問題。總體而言,AutoSR 將符號回歸從方程層級的搜索擴展到自動化的科學研究,允許科學知識和累積的證據塑造探索的內容以及結果方程的合理性。
Quipu: A Governed Bitemporal Knowledge Graph Store
2608.16813v1 by Steve Brown
Agents now write knowledge graphs, but knowledge-graph stores still carry defaults set when humans curated them: accept writes now and clean later, keep one time axis or none, treat every writer's facts as equally trustworthy, and leave governance to dashboards and middleware. These four defaults are individually convenient and jointly untenable under agent workloads. We present Quipu, an embeddable store that inverts all four: no fact enters except through a gate whose predicates evaluate the pending post-state; data, trust labels, verdicts, and the rules themselves are bitemporal; named graphs are the unit of authority and trust, composed under a lattice whose one invariant is that composition never widens; and the governance specification $Σ$, the trace, and signed verdicts are facts in the store they govern, making the audit $T \models Σ$ a query. We evaluate with Census, a deterministic multi-writer lifecycle whose single seeded run scores every research question against planted ground truth: the gated store ends with 0 of 6 planted defects versus 6 of 6 ungated; all 7 composition probes uphold the lattice contract; 50 of 50 satisfied verdicts re-derive faithfully as of their instant while all 50 would be misreported under a latest-only rule set; and the SARC reference checker agrees with the in-store audit verdict-for-verdict, differing only on coverage semantics. A recorded trace from a governed writer surfaces a live enforcement gap the audit names with its remediation. On DEMM-Bench, an external decision-evidence sufficiency benchmark, a content-only reading of the exported records answers all 512 property-level governance questions correctly with zero overclaim under all eight degradation conditions, while container-presence baselines overclaim on up to 87.5% of them -- and the run surfaced, and led us to close, a gap in what a denial's verdict attests.
摘要:代理人現在撰寫知識圖譜,但知識圖譜存儲仍然保留人類編輯時設置的默認值:現在接受寫入,稍後清理,保持一個時間軸或不保持,將每位作者的事實視為同樣可信,並將治理留給儀表板和中介軟體。這四個默認值在個別上方便,但在代理工作負載下共同無法維持。我們提出了 Quipu,一個可嵌入的存儲,顛覆了這四個默認值:沒有事實進入,除非通過一個門,其謂詞評估待處理的後狀態;數據、信任標籤、裁決和規則本身都是雙時間的;命名圖是權威和信任的單位,根據一個格子組成,其唯一的不變性是組合從不擴大;而治理規範 $Σ$、追蹤和簽名裁決是其治理的存儲中的事實,使得審計 $T \models Σ$ 成為一個查詢。我們使用 Census 進行評估,這是一個確定性的多寫入者生命週期,其單一的種子運行針對植入的真實數據評分每個研究問題:有門的存儲最終以 0 的 6 個植入缺陷結束,而無門的則為 6 的 6 個;所有 7 個組合探針都維護了格子合約;50 的 50 個滿意裁決在其瞬間忠實地重新推導,而所有 50 個在僅最新規則集下會被誤報;而 SARC 參考檢查器在存儲中的審計裁決上逐一一致,僅在覆蓋語義上有所不同。來自受治理作者的記錄追蹤顯示出審計所命名的實時執行差距及其補救措施。在 DEMM-Bench 上,一個外部決策證據充分性基準,對導出的記錄的內容僅閱讀正確回答了所有 512 個屬性級治理問題,並在所有八種降級條件下均無過度聲明,而容器存在基準則在多達 87.5% 的問題上過度聲明——而這次運行顯示出並引導我們關閉了否認裁決所證明的差距。
Bounded Semantic Planning and Deterministic Compilation for Reliable Enterprise Text-to-SQL
2608.16663v1 by Yi Ai
Direct text-to-SQL asks a language model to do two jobs: interpret the business question and construct the complete relational query. In enterprise schemas, SQL can execute successfully while using the wrong relationship role or aggregation grain. We study an alternative placement of the stochastic boundary. A multi-turn planner grounds phrases and selects from question-specific governed options; graph traversal, role predicates, grain lowering, SQL construction, and deterministic checks are implemented in code. We evaluate this semantic path compilation (SPC) system against direct DDL-to-SQL generation on the ACME insurance benchmark. On a 38-question adjudicated comparison set with three runs per question, SPC was adjudicated correct on every run for 37 questions (97.4%), compared with 21 (55.3%) for the baseline. The paired discordance was 16 questions in favor of SPC and none in favor of the baseline (two-sided exact McNemar p=3.05x10^-5). SPC answered all 38 questions correctly at least once and produced one refusal and no adjudicated wrong-but-executed run across 114 run outcomes; the baseline produced 29 adjudicated wrong runs and seven additional judge-flagged data-only coincidences on the same set. A strict-equivalence sensitivity analysis increased the paired difference. Additional SPC runs with GPT-5.4 and Gemini-3.6-Flash showed similar question-level robustness, although their per-run verdict artifacts were not preserved. Six additional benchmark items are retained in an all-item analysis and documented separately by failure class. The study supports an end-to-end systems result, not a causal claim that compilation alone produced the gain, because SPC receives governed semantic artifacts that the DDL baseline does not.
摘要:直接的文本到 SQL 要求語言模型執行兩項任務:解釋商業問題並構建完整的關聯查詢。在企業模式中,SQL 可以在使用錯誤的關係角色或聚合粒度的情況下成功執行。我們研究了隨機邊界的替代放置。一個多輪規劃器將短語進行實體化並從特定問題的受控選項中進行選擇;圖遍歷、角色謂詞、粒度降低、SQL 構建和確定性檢查都在代碼中實現。我們將這個語義路徑編譯(SPC)系統與 ACME 保險基準的直接 DDL 到 SQL 生成進行評估。在一組 38 個問題的裁定比較集中,每個問題進行三次運行,SPC 在 37 個問題的每次運行中都被裁定為正確(97.4%),而基準僅為 21(55.3%)。配對不一致的情況下,SPC 有 16 個問題,而基準則沒有(雙側精確 McNemar p=3.05x10^-5)。SPC 至少正確回答了所有 38 個問題一次,並在 114 次運行結果中產生了一次拒絕,且沒有裁定為錯誤但執行的運行;基準則產生了 29 次裁定為錯誤的運行,並在同一組中有七次額外的法官標記的數據僅巧合。嚴格等價的敏感性分析增加了配對差異。使用 GPT-5.4 和 Gemini-3.6-Flash 的額外 SPC 運行顯示出類似的問題級穩健性,儘管它們的每次運行判決工件未被保留。六個額外的基準項目在全項目分析中保留,並按失敗類別單獨記錄。這項研究支持端到端系統結果,而不是因果聲明,即僅僅編譯產生了增益,因為 SPC 接收了 DDL 基準所沒有的受控語義工件。
The Working Set of a Coding Agent: Coherence Debt in Repository-Scale Tasks
2608.16630v1 by Bardia Mohammadi, Lars Klein, Aman Chadha, Akhil Arora, Laurent Bindschaedler
Repository-scale coding requires an agent to keep tests, imports, configuration, and migration rules consistent within a bounded context window. We model this as reconstructing a coupled-fact graph: at each edit, a required fact comes from recent context or parametric memory, and the facts covered by neither form coherence debt. We supply and withhold each channel and inject faults across seven models and five harnesses. As expected, no model completes a task on an unseen API with both channels empty, and putting the facts in the prompt restores success. When a rename defeats what models memorized about a real library, all seven fail in the same place, passing and missing the same tests. Availability decides the outcome and distance does not: withholding a fact costs exactly the work it supports, and a supplied fact works as well far from the edit as next to it. Harnesses pay unequal prices for it: configurations that all pass every test differ more than tenfold in tokens consumed because they rebuild the same content at different rates, and spending more recovers nothing when facts are withheld. A missing fact produces wrong work rather than absent work: an agent asked to act acts, fabricating the file or guessing the value, so instruments built on reads look for a hole already filled. How often it says it is blocked instead is a property of the model, from every trial to none. Availability does not settle every edit: where standard and code disagree, agents follow the standard even when it prescribes the worse code, so a stale convention file costs more than no file. Because parametric memory substitutes for reading, on SWE-bench, where models likely know the repositories, reads no longer predict success. Harnesses should keep the facts an edit depends on available when the agent writes, and check that availability against what the agent produces rather than what it reads.
摘要:儲存庫規模的編碼需要一個代理在有限的上下文窗口內保持測試、導入、配置和遷移規則的一致性。我們將此建模為重建一個耦合事實圖:在每次編輯時,所需的事實來自最近的上下文或參數記憶,而未被涵蓋的事實則形成一致性債務。我們提供和保留每個通道,並在七個模型和五個工具中注入故障。如預期,當兩個通道都為空時,沒有模型能在未見過的API上完成任務,而將事實放入提示中則恢復成功。當重命名擊敗模型對真實庫的記憶時,所有七個模型在同一位置失敗,通過和未通過相同的測試。可用性決定結果,而距離則不然:保留一個事實的成本正好是它所支持的工作,而提供的事實在遠離編輯的地方也能同樣有效。工具為此支付不平等的價格:所有通過每個測試的配置在消耗的標記上差異超過十倍,因為它們以不同的速度重建相同的內容,而花費更多在事實被保留時則無法恢復任何東西。一個缺失的事實產生錯誤的工作而不是缺失的工作:被要求行動的代理會行動,製造文件或猜測值,因此基於讀取構建的工具會尋找已經填充的空洞。它說自己被阻塞的頻率是一個模型的特性,從每次試驗到無次試驗。可用性並不解決每次編輯:當標準和代碼不一致時,代理遵循標準,即使它規定了更糟的代碼,因此過時的約定文件的成本高於沒有文件。由於參數記憶替代了閱讀,在SWE-bench上,模型可能了解這些儲存庫,讀取不再預測成功。工具應在代理寫入時保持編輯所依賴的事實可用,並檢查該可用性與代理產生的內容,而不是它所讀取的內容。
Hypergraph-based Multimodal Retrieval-Augmented Generation with Incremental Refinement
2608.16628v1 by Shenao Chen, Yidan Xu, Xiangmin Han, Rundong Xue, Duanpo Wu, Yuhan Gao, Chenggang Yan, Yue Gao
Modern Multimodal Retrieval-Augmented Generation (M-RAG) systems are fundamentally limited by the binary connectivity paradigm of traditional simple graphs, which fails to capture the intricate, high-order correlations among heterogeneous entities, such as the N-ary relationships between a visual chart, its scattered textual descriptions, and underlying numerical data. Furthermore, existing refinement strategies often rely on exhaustive, full-page reconstruction to align cross-modal information, leading to prohibitive computational redundancy and the introduction of contextual noise in long-form document processing. In this paper, we propose Hyper-M2RAG, a novel framework that redefines multimodal document retrieval through High-order Hypergraph Representation Learning. We first formalize the document structure as a Multimodal Hypergraph, utilizing hyperedges as unified semantic containers to encapsulate multi-way associations across text, images, and tables, thereby transcending point-to-point modeling. To mitigate semantic fragmentation caused by physical pagination, we introduce an Anchor-driven Incremental Refinement mechanism. Rather than performing a global sweep, our approach identifies boundary-crossing anchor nodes and reconstructs their local hyper-topology using one-hop neighborhood contexts. This targeted refinement effectively bridges cross-page knowledge gaps with minimal computational footprints. Extensive evaluations on multimodal benchmarking datasets demonstrate that Hyper-M2RAG significantly outperforms state-of-the-art methods in both retrieval precision and generation coherence. Our code is available at https://github.com/ShenAoChen2001/MMHRAG.
摘要:現代的多模態檢索增強生成(M-RAG)系統在根本上受到傳統簡單圖的二元連接範式的限制,這無法捕捉異質實體之間複雜的高階相關性,例如視覺圖表、其分散的文本描述和底層數據之間的N元關係。此外,現有的精煉策略通常依賴於全面的全頁重建來對齊跨模態信息,這導致了過度的計算冗餘並在長文檔處理中引入了上下文噪音。在本文中,我們提出了Hyper-M2RAG,一個通過高階超圖表示學習重新定義多模態文檔檢索的新框架。我們首先將文檔結構形式化為多模態超圖,利用超邊作為統一的語義容器,以封裝文本、圖像和表格之間的多向關聯,從而超越點對點建模。為了減輕由物理分頁引起的語義碎片化,我們引入了一種基於錨點的增量精煉機制。我們的方法不是進行全局掃描,而是識別跨邊界的錨點並使用一跳鄰域上下文重建其局部超拓撲。這種有針對性的精煉有效地填補了跨頁知識的空白,並且計算開銷最小。在多模態基準數據集上的廣泛評估表明,Hyper-M2RAG在檢索精度和生成一致性方面顯著超越了最先進的方法。我們的代碼可在 https://github.com/ShenAoChen2001/MMHRAG 獲得。
Physics of Agents: Statistical Mechanics Predicts Collective Behavior of AI Agents
2608.16578v1 by Batu El, Jinhee Paeng, Fatih Dinc, Shiye Su, Mete Erdogan, Aneesh Pappu, Haotian Ye, Wanjia Zhao, Surya Ganguli, James Zou
AI agents increasingly operate as part of interacting systems rather than in isolation. As agents exchange information and jointly make decisions, their interactions can improve collective reasoning but may also produce herding, polarization, or amplify shared biases. Understanding and predicting these collective dynamics is therefore important for designing effective and aligned multi-agent systems. Here, we study over 10,000 communities of language-model agents that repeatedly exchange messages and revise their opinions across objective mathematics questions and subjective political statements. Despite substantial diversity in possible behavior, the individual and group dynamics can be represented by three characteristic regimes: indifference, polarization, and consensus. AI agents start indifferent and build conviction as they interact. On objective questions, communication improves collective accuracy, while on subjective questions it often drifts group opinions toward the right in the political spectrum. We explain these observations with a statistical-mechanics formalism in which agents stochastically favor lower social pressure. Given only initial opinions, our model predicts individual trajectories, outperforms all standard baselines, generalizes to unseen community graphs, and reproduces the observed group archetype distributions. Our fitted model parameters reveal the mechanics underlying our key observations: i) communities operate below the critical social temperature, which explains conviction buildup; ii) attractive ties outweigh repulsive ones, which favors consensus; and iii) agents holding the correct answer exert the strongest pull, which drives truth-seeking. Overall, our results demonstrate that collective behavior of AI agents, like that of other complex systems, follows compact and predictive dynamical laws.
摘要:AI 代理人越來越多地作為互動系統的一部分運作,而不是孤立存在。
隨著代理人之間交換信息並共同做出決策,他們的互動可以改善集體推理,但也可能產生跟風、極化或放大共享偏見。
因此,理解和預測這些集體動態對於設計有效且一致的多代理系統非常重要。
在這裡,我們研究了超過 10,000 個語言模型代理人的社群,它們反覆交換消息並在客觀數學問題和主觀政治陳述上修正自己的意見。
儘管可能的行為存在相當大的多樣性,但個體和群體動態可以用三種特徵性狀態來表示:漠不關心、極化和共識。
AI 代理人最初是漠不關心的,隨著互動的進行建立信念。
在客觀問題上,交流提高了集體準確性,而在主觀問題上,則經常使群體意見向政治光譜的右側漂移。
我們用一種統計力學形式主義來解釋這些觀察,其中代理人隨機地偏好較低的社會壓力。
僅根據初始意見,我們的模型預測個體軌跡,超越所有標準基準,對未見過的社群圖進行泛化,並重現觀察到的群體原型分佈。
我們擬合的模型參數揭示了我們關鍵觀察背後的機制:i) 社群運作在臨界社會溫度以下,這解釋了信念的積累;ii) 吸引性聯繫超過排斥性聯繫,這有利於共識;iii) 持有正確答案的代理人施加最強的影響,這驅動尋求真相。
總體而言,我們的結果表明,AI 代理人的集體行為,如同其他複雜系統,遵循緊湊且可預測的動態法則。
Ask, Condition or Abstain: Reinforcement Learning for Missing-Premise Reasoning
2608.16554v1 by Yongqi Tong, Zhenyu Zhang, Zimi Liu, Kewei Fu, Mingli Song, Haofei Zhang, Junshao Zhang, Hong Zhu, Jiang-Ming Yang, Xin Zhang, Jianshe Li
Answer-only reinforcement learning (RL) trains reasoning models to solve fully specified problems, but many realistic queries omit a premise needed for a unique answer. In this setting, the useful response is not always refusal: the model should ask for the missing premise, condition its answer on the unknown quantity, or abstain when no informative conditional response is available. We present \emph{Ask-Condition-Abstain Reinforcement Learning} (ACA-RL), a data-augmented RL framework for this setting. Its reasoning-graph-guided pipeline converts well-posed problems into missing-premise training instances with localized gap annotations; ACA-RL then trains on these instances with a structured reward over five observable response behaviors. We also introduce the \emph{Missing-Premise Benchmark} (MPB), a 274-instance human-verified benchmark spanning mathematical, logical, and real-world word problems. Across Qwen3 and Llama models, ACA-RL consistently improves on MPB while preserving competitive performance on well-posed reasoning tasks. Together with the released code, MPB, and training data, this work supports a new mission for NLP evaluation: measuring whether models can recognize when a task is underdetermined and handle uncertainty, not only whether they can answer fully specified questions.
摘要:答案導向的強化學習(RL)訓練推理模型以解決完全具體的問題,但許多現實查詢省略了獲得唯一答案所需的前提。在這種情況下,有用的回應不總是拒絕:模型應該要求缺失的前提,將其回答條件化於未知量,或在沒有可提供信息的條件回應時選擇不作答。我們提出了\emph{詢問-條件-不作答強化學習}(ACA-RL),這是一個針對這種情境的數據增強強化學習框架。其推理圖引導的流程將良好表述的問題轉換為缺失前提的訓練實例,並附有局部缺口註釋;然後,ACA-RL在這些實例上進行訓練,並對五種可觀察的回應行為給予結構化的獎勵。我們還介紹了\emph{缺失前提基準}(MPB),這是一個包含274個經過人類驗證的基準,涵蓋數學、邏輯和現實世界的文字問題。在Qwen3和Llama模型中,ACA-RL在MPB上持續改進,同時在良好表述的推理任務中保持競爭性能。連同發布的代碼、MPB和訓練數據,這項工作支持NLP評估的新任務:測量模型是否能夠識別任務是否不確定並處理不確定性,而不僅僅是它們是否能回答完全具體的問題。
VCE-Skill: Enhancing Skill Self-Evolution with Version-Change Experience
2608.16544v1 by Jianming Chen, Xuanbin Ye, Yawen Wang, Junjie Wang, Qing Wang, Fanjiang XU
Agents increasingly rely on reusable skills to encode task knowledge, tool-use procedures, and validation rules. Existing skill self-evolution methods primarily revise skills using execution trajectories collected from current tasks, leaving the evolution knowledge accumulated in public skill version histories largely untapped. Our pilot study reveals a clear complementarity between the two sources: public skill changes provide reusable evolution priors, whereas trajectories provide evidence grounded in the current task. Motivated by this, we propose VCE-Skill, which distills noisy and implementation-specific public skill changes into reusable, structured version-change experience and adaptively fuses it with trajectory-derived proposals from the base evolver, thereby exploiting external experience while retaining task-specific evidence. Extensive experiments demonstrate that VCE-Skill improves skill self-evolution, increasing mean scores by 3.20--4.98 points; transfer experiments further show that the resulting skills achieve stronger cross-model transfer performance. Our work highlights public skill version changes as a previously underexplored yet effective source of prior knowledge and advances trajectory-driven skill self-evolution.
摘要:代理人越來越依賴可重用的技能來編碼任務知識、工具使用程序和驗證規則。現有的技能自我演化方法主要使用從當前任務收集的執行軌跡來修訂技能,導致在公共技能版本歷史中積累的演化知識大多未被利用。我們的初步研究揭示了這兩個來源之間明顯的互補性:公共技能變更提供可重用的演化先驗,而軌跡則提供基於當前任務的證據。基於此,我們提出了 VCE-Skill,該方法將嘈雜且具實施特定性的公共技能變更提煉為可重用的、結構化的版本變更經驗,並自適應地將其與基礎演化器的軌跡導出提案融合,從而在保留任務特定證據的同時利用外部經驗。大量實驗表明,VCE-Skill 改進了技能自我演化,平均分數提高了 3.20--4.98 分;轉移實驗進一步顯示,所產生的技能在跨模型轉移性能上更強。我们的工作突出了公共技能版本變更作為一個先前未被充分探索但有效的先驗知識來源,並推進了基於軌跡的技能自我演化。
Large language models as synthetic clinical experts to inform longitudinal rare-disease modeling
2608.16507v1 by Clemens Schächter, Astrid Pechmann, Janbernd Kirschner, Jan Hasenauer, Harald Binder
Due to the limited amount of information, modeling longitudinal rare-disease data can benefit from integrating clinical knowledge. Yet, elicitation of expert knowledge and formalization for model fitting is challenging, in particular due to limited time of clinical experts. To nevertheless make domain knowledge accessible during model fitting, we use large language models (LLMs) as synthetic clinical experts to supervise a variational-autoencoder-based approach that learns low-dimensional latent summaries of visit-level observations. Specifically, LLMs are queried offline on textual descriptions of patient observations to obtain judgments, e.g., the suspected clinical category. To improve the variational autoencoder fit, we train a differentiable surrogate model on these judgments and augment the loss function to encourage reconstructions that preserve the clinical-label distribution of their corresponding input profile. In an application to longitudinal motor-function assessments from children with spinal muscular atrophy, we map visit-level clinical profiles to low-dimensional representations that are linked by a multivariate mixed-effects model. The synthetic expert loss discourages reconstructions that remain numerically close in data space but alter the clinical interpretation of the reconstructed motor function profile, such as by crossing a disease-type boundary. We thus reduced disagreement between original and reconstructed SMA type labels from about 11 to 7 percent. Furthermore, informing the latent representation by the synthetic expert improved prediction of motor function milestones compared with unsupervised latent representations and a data-level baseline. These results suggest that incorporating LLMs into model fitting can make clinical knowledge available to representation learning and improve clinical faithfulness for longitudinal rare-disease data.
摘要:由於資訊量有限,建模縱向罕見疾病數據可以從整合臨床知識中受益。然而,專家知識的引出和模型擬合的形式化是具有挑戰性的,特別是由於臨床專家的時間有限。儘管如此,為了在模型擬合過程中使領域知識可用,我們使用大型語言模型(LLMs)作為合成臨床專家,來監督基於變分自編碼器的方法,該方法學習訪問級觀察的低維潛在摘要。具體來說,我們在患者觀察的文本描述上離線查詢LLMs以獲得判斷,例如,懷疑的臨床類別。為了改善變分自編碼器的擬合,我們在這些判斷上訓練了一個可微分的替代模型,並增強損失函數以鼓勵重建保持其對應輸入特徵的臨床標籤分佈。在對脊髓性肌萎縮症兒童的縱向運動功能評估的應用中,我們將訪問級臨床特徵映射到由多變量混合效應模型鏈接的低維表示。合成專家損失會抑制在數據空間中數值上接近但改變重建運動功能特徵的臨床解釋的重建,例如通過跨越疾病類型邊界。因此,我們將原始和重建的SMA類型標籤之間的分歧從約11%減少到7%。此外,通過合成專家告知潛在表示,與無監督潛在表示和數據級基準相比,運動功能里程碑的預測得到了改善。這些結果表明,將LLMs納入模型擬合可以使臨床知識可用於表示學習,並改善縱向罕見疾病數據的臨床真實性。
Graph Machine Learning: An Opportunity for Power Systems
2608.16494v1 by Martin Sadric, Sebastian Pütz, Christian Nauck, Veit Hagenmeyer, Frank Hellmann, Dirk Witthaut, Benjamin Schäfer
Modern power systems face growing operational complexity driven by the integration of renewable energy sources, decentralization, and the need for real-time decision-making across a wide range of timescales. Addressing these challenges traditionally relies on model-based methods that, while accurate, can be too slow for operational demands. Machine learning (ML) has therefore emerged as a faster, data-driven alternative. As grid topology plays a central role in power system operation, graph machine learning (GML) methods offer a natural framework for incorporating topological dependencies as an inductive bias. We survey nearly 800 papers at the intersection of GML and power systems, covering forecasting, state estimation, optimization, control, fault diagnosis, and cybersecurity. Power systems constitute an unusually rich benchmark setting for GML, as they combine hard physical constraints, multi-scale dynamics, safety-critical requirements, and scarce labeled data within a single, well-defined domain. Conversely, power systems can benefit from utilizing GML to complement classical solvers, as GML provide scalable, topology-aware approximations with promising generalization and computational efficiency. We identify open challenges, including limited real-world deployment and the need for interpretable models in safety-critical settings. Despite the rapidly growing number of publications, standardized benchmarks and open datasets remain scarce, leaving many results difficult to reproduce and undermining the long-term scientific credibility of the field. We further derive a structured requirements catalog for ML-ready power grid benchmarks, intended to guide future dataset development and improve reproducibility across studies. We call on the community to prioritize dedicated benchmark studies and the release of open datasets and models.
摘要:現代電力系統面臨著由可再生能源整合、去中心化以及在廣泛時間尺度上進行實時決策所驅動的日益增長的運營複雜性。傳統上,解決這些挑戰依賴於基於模型的方法,儘管這些方法準確,但對於運營需求來說可能過於緩慢。因此,機器學習(ML)作為一種更快的數據驅動替代方案應運而生。由於電網拓撲在電力系統運作中扮演著核心角色,圖形機器學習(GML)方法提供了一個自然的框架,以將拓撲依賴性作為歸納偏差納入考量。我們調查了近800篇GML與電力系統交叉的論文,涵蓋預測、狀態估計、優化、控制、故障診斷和網絡安全。電力系統為GML提供了一個異常豐富的基準設置,因為它們在一個明確定義的領域內結合了嚴格的物理約束、多尺度動力學、安全關鍵要求以及稀缺的標記數據。相反,電力系統可以利用GML來補充傳統求解器,因為GML提供了可擴展的、考慮拓撲的近似,並且具有良好的泛化能力和計算效率。我們確定了開放挑戰,包括有限的實際部署和在安全關鍵環境中對可解釋模型的需求。儘管出版物數量迅速增長,標準化基準和開放數據集仍然稀缺,這使得許多結果難以重現,並削弱了該領域的長期科學可信度。我們進一步推導了一個結構化的需求目錄,用於ML準備好的電網基準,旨在指導未來數據集的開發並提高研究的可重複性。我們呼籲社區優先考慮專門的基準研究以及開放數據集和模型的發布。
Time to Reason: Scalable Neurosymbolic Learning for LTLf via Fuzzy Semantics
2608.16443v1 by Riccardo Andreoni, Andrei Buliga, Alessandro Daniele, Paolo Felli, Chiara Ghidini, Marco Montali, Massimiliano Ronzani
Neurosymbolic (NeSy) Artificial Intelligence aims to integrate Deep Learning (DL) architectures with symbolic reasoning. While initial NeSy approaches have targeted mainly symbolic reasoning in propositional and first-order logics, recent works have started to address the construction of neurosymbolic frameworks for Temporal Logics, and in particular for LTLf. These approaches have established temporal NeSy as a promising research direction, laying the foundations for learning under temporal constraints. Nonetheless, they leave many questions unanswered. From a theoretical perspective, several differentiable semantics for interpreting LTLf have been proposed but have not yet been formally and systematically defined within a unified framework. Moreover, existing approaches commonly rely on automata to represent temporal knowledge, resulting in limited scalability. Motivated by this research gap, this paper provides the following contributions: (i) formally defining different fuzzy semantics for LTLf, and systematically analysing theoretical properties regarding equivalences and dualities of temporal operators; (ii) showing how these semantics can be directly integrated within a novel NeSy framework, called DiffLTLf, enabling flexible and scalable learning without relying on the usage of automata; and (iii) introducing a novel evaluation protocol of increased complexity of learning tasks w.r.t. existing benchmarks. Our results show that the choice of fuzzy semantics has a significant impact on predictive performance. Moreover, DiffLTLf achieves performance on par with, and sometimes superior to, state-of-the-art probabilistic approaches while substantially improving scalability. Taken together, these results establish direct fuzzy interpretations as a competitive and scalable alternative to existing temporal NeSy frameworks.
摘要:神經符號(NeSy)人工智慧旨在將深度學習(DL)架構與符號推理整合。雖然最初的NeSy方法主要針對命題邏輯和一階邏輯中的符號推理,但最近的研究已開始著手於構建神經符號框架以處理時間邏輯,特別是針對LTLf。這些方法已將時間NeSy確立為一個有前景的研究方向,為在時間約束下的學習奠定了基礎。儘管如此,它們仍然留下許多未解答的問題。從理論的角度來看,已提出幾種可微分的語義來解釋LTLf,但尚未在統一框架內正式和系統地定義。此外,現有的方法通常依賴自動機來表示時間知識,導致可擴展性有限。受此研究空白的啟發,本文提供了以下貢獻:(i)正式定義不同的LTLf模糊語義,並系統地分析有關時間運算符的等價性和對偶性的理論性質;(ii)展示這些語義如何能夠直接整合進一個新穎的NeSy框架,稱為DiffLTLf,實現靈活且可擴展的學習,而無需依賴自動機的使用;以及(iii)引入一種新的評估協議,增加學習任務的複雜性,相較於現有基準。我們的結果顯示,模糊語義的選擇對預測性能有顯著影響。此外,DiffLTLf在性能上與最先進的概率方法相當,有時甚至優於它們,同時顯著提高了可擴展性。綜合這些結果,直接的模糊解釋被確立為現有時間NeSy框架的競爭性和可擴展替代方案。
Reasoning-supported Robustness Validation of Automotive E/E Components
2608.16421v1 by Jan Novacek, Alexander Viehl, Oliver Bringmann, Wolfgang Rosenstiel
This paper presents an ontology-supported approach to tackle the complexity of the Robustness Validation (RV) process of automotive electrical/electronic (E/E) components. The approach uses formalized knowledge from the RV process and stress, operating, and load profiles, so-called Mission Profiles (MPs). In contrast to the error-prone industrially established manual procedure, we show how component characteristics are formalized in OWL in order to form the foundation of an efficient automated analysis selection and decision support during the RV process. The proposed approach is based on the idea of mapping MPs to an OWL representation so to allow to perform semantic queries against MP data to improve their integration into the RV process. The resulting ontology-supported application framework has been applied to an industrial use-case from automotive power electronics. We present experimental results showing that the RV process can be significantly improved in terms of reduced design time and increased exhaustiveness by automating the analyses selection step and the provisioning of all the relevant data to be used.
摘要:這篇論文提出了一種基於本體的方式來應對汽車電氣/電子(E/E)元件的穩健性驗證(RV)過程的複雜性。該方法利用了來自RV過程的形式化知識以及壓力、操作和負載特徵,這些被稱為任務特徵(MPs)。與錯誤易發的工業手動程序相比,我們展示了如何在OWL中形式化元件特徵,以便為RV過程中的高效自動分析選擇和決策支持奠定基礎。所提出的方法基於將MP映射到OWL表示的想法,以便能夠對MP數據執行語義查詢,從而改善其在RV過程中的整合。最終得到的基於本體的應用框架已應用於汽車功率電子的工業案例。我們展示了實驗結果,顯示通過自動化分析選擇步驟和提供所有相關數據,RV過程在設計時間減少和全面性增加方面可以顯著改善。
Think Inside the Chunk: RegulaRAG for Regulation-Compliant Scenario Generation using LLMs: A Case Study of UN Regulation No. 152
2608.16394v1 by Vahid Zolfaghari, Nenad Petrovic, AndrÉ Schamschurko, Alois Knoll
Generating regulation-compliant test scenarios is essential for validating safety-critical automotive systems, yet Large Language Models (LLMs) struggle to ground outputs in long, hierarchical standards. We present RegulaRAG, a Retrieval-Augmented Generation (RAG) pipeline that couples SmartChunking, reference-aware enrichment of paragraphs and tables via graph traversal, with Smart Retrieve & Rerank over these enriched units. To test our system, we evaluate on a manually curated dataset covering all scenarios in UN Regulation No. 152 (AEBS). Our study comprises: (i) a three-step progressive search that identifies near-optimal retrieval parameters without exhaustive grid search; (ii) head-to-head comparisons against five baseline RAG systems; and (iii) a robustness stress test that scales the source corpus with distractor content. Outputs are evaluated using a customized penalized scoring metric. Across all experiments, RegulaRAG achieves the highest average Meta-Score (82.99), outperforming the next-best system by 43% (NoRAG: 57.94), while operating at 14k-25k tokens per query versus up to 500k for graphcentric baselines. It maintains strong performance, remaining stable even as the number of regulatory sources grows, whereas competing RAG systems degrade sharply in both quality and robustness.
摘要:生成符合規範的測試場景對於驗證安全關鍵的汽車系統至關重要,但大型語言模型(LLMs)在將輸出與長期的層次標準相結合方面存在困難。
我們提出了RegulaRAG,一個檢索增強生成(RAG)管道,結合了SmartChunking、通過圖遍歷對段落和表格進行參考感知的豐富化,以及對這些豐富單元的智能檢索與重新排序。
為了測試我們的系統,我們在一個手動策劃的數據集上進行評估,該數據集涵蓋了聯合國第152號規範(AEBS)中的所有場景。
我們的研究包括:(i)一個三步驟的漸進搜索,識別近乎最優的檢索參數,而無需進行耗時的網格搜索;(ii)與五個基準RAG系統的正面比較;以及(iii)一個強度壓力測試,通過干擾內容擴展源語料庫。
輸出使用自定義的懲罰評分指標進行評估。
在所有實驗中,RegulaRAG達到了最高的平均Meta-Score(82.99),比第二好的系統高出43%(NoRAG: 57.94),同時每個查詢的操作在14k-25k個標記之間,而圖中心基準則高達500k。
它保持了強勁的性能,即使在監管來源數量增加的情況下也保持穩定,而競爭的RAG系統在質量和穩定性方面急劇下降。
Mint-Agent: Introducing Finance-Native Agentic Foundation Models
2608.16386v1 by Mint-Agent Team, B. Zhang, Yaze Geng, Lei Tang, Yaoyang Yi, Zonghan Wu, Yifan Hu, Kun Wang, Qingsong Wen, Yilei Shao
Financial agents must do more than recall domain knowledge: they must be both reliable, executing precise operations over grounded evidence, and executive, sustaining long-horizon research whose conclusions remain auditable. We present Mint-Agent, a family of finance-native agentic models designed around these two scales of financial intelligence. Mint-Agent is built upon three pillars: data, harness, and algorithm. Our data engine constructs clean, specialized tasks for atomic financial capabilities and long-horizon agentic execution from real-world financial sources. MintHarness enables stable interaction with open-ended environments and maintains auditable evidence trails across extended research trajectories. Our training recipe combines SFT, critical-step OPD, and RLVR to develop separate financial reasoning and agentic execution experts, which are then unified through model merging and multi-teacher on-policy distillation into compact, general-purpose financial agents. This pipeline yields two flagship models, Mint-Cu (9B) and Mint-Ag (27B). Across professional financial benchmarks, our models demonstrate two defining strengths: (1) Reliability: Mint-Ag achieves 98.33% on RFC-Bench, surpassing GPT-5.6-Sol and Claude-Opus-4.8 by 3.66 and 3.00 points; and (2) Executability: Mint-Cu reaches 69.86% on FinSearchComp T2, outperforming Agents-A1-35B and Nex-N2-mini by 22.83 and 12.78 points, while Mint-Ag achieves 76.00% and 60.49% on FinanceAgentBench v1.1 and v2, respectively. These results establish a path toward trustworthy financial intelligence in which domain expertise, long-horizon execution, and auditable evidence are jointly engineered as a unified foundation for frontier agentic models.
摘要:金融代理人必須做的不僅僅是回憶領域知識:他們必須既可靠,能夠在有根據的證據上執行精確的操作,又必須具備執行力,能夠支持長期研究,其結論保持可審核性。我們提出了Mint-Agent,這是一系列圍繞這兩個金融智能尺度設計的金融原生代理模型。Mint-Agent建立在三個支柱之上:數據、利用和算法。我們的數據引擎從現實世界的金融來源構建乾淨的、專門的任務,以實現原子金融能力和長期代理執行。MintHarness使得與開放式環境的穩定互動成為可能,並在擴展的研究軌跡中維持可審核的證據鏈。我們的訓練配方結合了SFT、關鍵步驟OPD和RLVR,以培養獨立的金融推理和代理執行專家,然後通過模型合併和多教師在政策蒸餾將它們統一為緊湊的通用金融代理。這個流程產生了兩個旗艦模型,Mint-Cu (9B)和Mint-Ag (27B)。在專業金融基準測試中,我們的模型展示了兩個明確的優勢:(1)可靠性:Mint-Ag在RFC-Bench上達到98.33%,超越了GPT-5.6-Sol和Claude-Opus-4.8,分別提高了3.66和3.00分;以及(2)可執行性:Mint-Cu在FinSearchComp T2上達到69.86%,超越了Agents-A1-35B和Nex-N2-mini,分別提高了22.83和12.78分,而Mint-Ag在FinanceAgentBench v1.1和v2上分別達到76.00%和60.49%。這些結果為值得信賴的金融智能鋪平了道路,其中領域專業知識、長期執行和可審核的證據共同被設計為前沿代理模型的統一基礎。
MELD: A Protocol for Merging Knowledge Across Distributed Agentic Memories
2608.16357v1 by Lauri Lovén, Jaakko Sauvola, Jukka Riekki, Sasu Tarkoma
Autonomous agents share a transport and can call each other's tools, but they cannot share what they know: no protocol lets two agents' memories reconcile a fact phrased two ways, link related facts held apart, or reconcile contradictory knowledge without silently discarding either claim. We present MELD, a self-managing coherence mechanism for a federation of agent memories whose run-time model is the knowledge graph itself. Each brain admits every incoming claim through a five-outcome procedure (insert, merge, relate, conflict, or reject), decided from three signals (scoped claim-key identity, embedding similarity, and a natural-language-inference verdict) under context and freshness gates, and acting through exactly one auditable, authenticated Patch, the only object that mutates state. A binding onto standard publish/subscribe transport with a per-claim status CRDT keeps sovereign brains coherent in claim status without a coordinator: self-healing after partitions and under lossy routing, and self-protecting against silent rewrite by a peer, under a benign-fault model. MELD does not adjudicate truth; a detected contradiction is preserved for later adjudication, never silently resolved. On HotpotQA distractor, distributed merge is recall-non-inferior to a centralized store under a pre-specified equivalence test and recall-superior to naive union at about 11% less live storage; the merge classifier separates at AUC 0.968 with a 0.013 false-merge rate on adjudicated candidate pairs; the status CRDT reconverges in 30/30 real partition-heal trials where last-writer-wins manages 11/30; and semantic routing delivers about 3x fewer messages at matched recall. We evaluate on a real computing continuum spanning an operator-grade 5G edge, national HPC, and a local tier, with empirically calibrated thresholds.
摘要:自主代理共享一個傳輸並可以呼叫彼此的工具,但他們無法共享所知:沒有任何協議能讓兩個代理的記憶調和以兩種方式表述的事實,連結分開持有的相關事實,或在不靜默丟棄任何主張的情況下調和矛盾的知識。我們提出了MELD,一種自我管理的連貫性機制,用於一個代理記憶的聯邦,其運行時模型即為知識圖譜。每個大腦通過一個五種結果的程序(插入、合併、關聯、衝突或拒絕)接受每個進來的主張,這一決定基於三個信號(範疇主張鍵身份、嵌入相似性以及自然語言推理的裁決),在上下文和新鮮度閘門下運作,並通過恰好一個可審計的、經過身份驗證的Patch進行操作,這是唯一能改變狀態的對象。與標準的發布/訂閱傳輸綁定的每個主張狀態CRDT使得主權大腦在主張狀態上保持一致,無需協調者:在分區後自我修復,並在有損路由下自我保護,防止被同伴靜默重寫,遵循良性故障模型。MELD不裁決真相;檢測到的矛盾將被保留以便後續裁決,絕不靜默解決。在HotpotQA的干擾者上,分散合併的回憶在預先指定的等價測試下不劣於集中存儲,並且在約11%更少的實時存儲下回憶優於天真的聯合;合併分類器在裁決候選對上以AUC 0.968分開,假合併率為0.013;狀態CRDT在30/30的真實分區修復試驗中重新收斂,而最後寫者獲勝的情況下僅管理11/30;語義路由在匹配回憶時傳遞的消息數量約少了3倍。我們在一個涵蓋操作級5G邊緣、國家HPC和本地層的真實計算連續體上進行評估,並經過實證校準的閾值。
AeroCopilotBench: A Two-Tier Benchmark for Evaluating LLM Agents as Aviation Copilots in an Interactive Virtual Cockpit Environment
2608.16349v1 by Yuchen Yuan, Zhenghuang Wu, Yuangan Li, Liang Ma, Ke Li
Large language model (LLM) agents may assist flight crews with complex decisions and task execution, but existing aviation evaluations centered on static knowledge do not support systematic testing of procedural execution and safety compliance in interactive environments. This paper presents the AeroCopilot Operational Environment (ACOE), a reproducible interactive virtual-cockpit test environment, and AeroCopilotBench, a two-tier aviation agent evaluation benchmark. Tier-1 evaluates aviation knowledge using 1,200 multiple-choice questions, while Tier-2 comprises 73 emergency and abnormal tasks derived from the manufacturers' Pilot's Operating Handbooks (POHs) and instantiated in ACOE. ACOE converts natural-language procedures into executable state transitions, final-state goal conditions, and hard safety constraints, enabling models to interpret cockpit state, diagnose faults, and operate aircraft systems through standardized tool interfaces. We establish a safety-gated evaluation framework in which a trajectory succeeds only when all task goals are achieved without violating any hard safety constraint, while safe goal progress and trajectory safety are measured separately. Across 12 models, the highest Tier-2 success rate is 72.6%, while static knowledge performance does not consistently translate into procedural execution. Analysis of 451 failed episodes from 3 representative models identifies recurring failures in procedural completeness, use of state feedback, and long-horizon execution management. These findings motivate state-aware agent orchestration, joint assessment of task completion and trajectory safety, and repeated regression testing. ACOE and AeroCopilotBench provide a reproducible foundation for testing knowledge application, interactive execution, and operational safety in aviation agents.
摘要:大型語言模型(LLM)代理可能協助飛行組員進行複雜的決策和任務執行,但現有的航空評估集中於靜態知識,無法支持在互動環境中系統性測試程序執行和安全合規性。本文提出了航空副駕駛操作環境(ACOE),這是一個可重複的互動虛擬駕駛艙測試環境,以及航空副駕駛基準(AeroCopilotBench),這是一個兩級航空代理評估基準。第1級使用1,200道多選題評估航空知識,而第2級則包含73個來自製造商飛行操作手冊(POHs)的緊急和異常任務,並在ACOE中實現。ACOE將自然語言程序轉換為可執行的狀態轉換、最終狀態目標條件和硬性安全約束,使模型能夠解釋駕駛艙狀態、診斷故障並通過標準化工具接口操作飛機系統。我們建立了一個安全門控評估框架,其中只有在所有任務目標達成且不違反任何硬性安全約束的情況下,軌跡才算成功,而安全目標進展和軌跡安全則分別測量。在12個模型中,第2級的最高成功率為72.6%,而靜態知識表現並不總是一致地轉化為程序執行。對3個代表性模型中451個失敗案例的分析發現,程序完整性、狀態反饋的使用和長期執行管理存在重複性失敗。這些發現促使了狀態感知代理的協同、任務完成和軌跡安全的聯合評估,以及重複回歸測試。ACOE和航空副駕駛基準為測試知識應用、互動執行和航空代理的操作安全提供了可重複的基礎。
Executable Code Knowledge: Code as a Native, Validation-Carrying Knowledge Representation for AI Coding Agents
2608.16295v1 by Xueping Gao
AI coding agents need more than relevant snippets: they need business semantics, validation evidence, relations, and assurance that their context is current. Existing systems usually infer or externalize this knowledge through retrieval, summaries, graphs, rules, or reverse specifications. We investigate a complementary representation in which selected code units directly carry agent-usable knowledge. We introduce Executable Code Knowledge (ECK) and define an Executable Code Knowledge Unit (ECKU) as a source-bound object combining stable identity, semantics, executable behavior, contracts, evidence, relations, provenance, validation state, and a query interface. Our Python prototype supports code-local authoring, manifest export, evidence execution, exact changed-line impact, freshness checking, and agent-facing projections. Across three real Python repositories and 26 controlled patch tasks, direct ECK provides executable test coverage for 11/11 evidence-bearing tasks and exact selectors for 9/11; hiding declared evidence reduces exact recovery to 1/11 (paired exact McNemar p=0.0078). ECK-derived rules recover 11/11 exact selectors, showing that rules are effective delivery artifacts while ECK supplies source binding, validation state, impact, and freshness. Exact changed-line impact matches independently authored labels on all 26 patches (12 unit links; precision, recall, and F1 all 1.000). AST-bounded fingerprints classify 50 positive changes and 17 unrelated same-file controls correctly, whereas static rules snapshots detect none of the 50 stale cases. Model-backed patch-review and cross-layer studies measure projection fidelity rather than independent impact discovery. These results support a hybrid architecture: retrieval for coverage, ECK for source and evidence governance, and projections for delivery.
摘要:AI 編碼代理需要的不僅僅是相關的片段:他們需要商業語義、驗證證據、關係,以及確保其上下文是最新的。現有系統通常通過檢索、摘要、圖表、規則或反向規範來推斷或外部化這些知識。我們研究了一種互補的表示方式,其中選定的代碼單元直接攜帶代理可用的知識。我們引入了可執行代碼知識(Executable Code Knowledge, ECK),並將可執行代碼知識單元(Executable Code Knowledge Unit, ECKU)定義為一個源綁定對象,結合穩定的身份、語義、可執行行為、合約、證據、關係、來源、驗證狀態和查詢介面。我們的 Python 原型支持代碼本地創作、清單導出、證據執行、精確變更行影響、更新性檢查和面向代理的投影。在三個真實的 Python 存儲庫和 26 個受控補丁任務中,直接 ECK 為 11/11 個帶證據的任務提供了可執行的測試覆蓋,並為 9/11 提供了精確選擇器;隱藏已聲明的證據將精確恢復降低到 1/11(配對精確 McNemar p=0.0078)。ECK 衍生的規則恢復了 11/11 的精確選擇器,顯示規則是有效的交付工件,而 ECK 提供了源綁定、驗證狀態、影響和更新性。精確變更行影響與所有 26 個補丁上的獨立創作標籤相匹配(12 個單元鏈接;精確度、召回率和 F1 均為 1.000)。AST 限定的指紋正確分類了 50 個正變更和 17 個無關的同檔控制,而靜態規則快照未檢測到 50 個過時案例。模型支持的補丁審查和跨層研究測量投影忠實度,而不是獨立影響發現。這些結果支持一種混合架構:檢索用於覆蓋,ECK 用於源和證據治理,投影用於交付。
Clause Encounters of the Third Kind: Can LLMs Replace Language Teachers?
2608.16286v1 by Kristina Šekrst, Ana Kovačić
While various organizations now actively encourage LLM use in classrooms, we still lack rigorous, systematic evaluations of how well these models actually perform the fundamental tasks of language pedagogy. This paper examines whether state-of-the-art LLMs can deliver the kind of corrective feedback and methodological explanations that language learners need. The study tests multiple large language models on their ability to identify, correct, and explain common learner mistakes in English, by systematically varying model parameters to investigate how these technical adjustments affect output quality, pedagogical clarity, and consistency, along with using retrieval-augmented generation to query methodological data. The evaluation employs automated metrics (GLEU, BERTScore) but also human expert judgments to capture dimensions that purely computational measures miss: linguistic nuance, cultural sensitivity, and instructional appropriateness. While models demonstrate impressive surface-level correction abilities, their explanations often lack the terminological and domain knowledge that effective language teaching requires, suggesting that current enthusiasm for AI-assisted language learning may be outpacing our understanding of these systems' actual pedagogical competence.
摘要:雖然各種組織現在積極鼓勵在課堂上使用 LLM,但我們仍然缺乏對這些模型在語言教學基本任務上實際表現的嚴謹、系統性評估。本文檢視最先進的 LLM 是否能提供語言學習者所需的糾正反饋和方法論解釋。該研究測試多個大型語言模型在識別、糾正和解釋英語學習者常見錯誤的能力,通過系統性地變化模型參數來調查這些技術調整如何影響輸出質量、教學清晰度和一致性,同時使用增強檢索生成來查詢方法論數據。評估使用自動化指標(GLEU、BERTScore),但也包括人類專家的判斷,以捕捉純計算度量所忽略的維度:語言細微差別、文化敏感性和教學適當性。雖然模型展示了令人印象深刻的表面糾正能力,但它們的解釋往往缺乏有效語言教學所需的術語和領域知識,這表明目前對 AI 輔助語言學習的熱情可能超過了我們對這些系統實際教學能力的理解。
Domain-Agnostic Neural Topic Modeling with Contextual Token-Level Semantic Graph Representation
2608.16269v1 by Seung-Won Seo, Won Ik Cho, Yongmin Yoo
Recent advances in neural topic models with pre-trained language models (PLMs) have achieved strong performance by leveraging general-domain pre-training, yet their topic interpretability often degrades on specialized corpora. This limitation primarily stems from the geometry of the embedding space, where domain-specific terms unseen during pre-training collapse into an indistinguishable region, and neither domain-specific re-training, word-level graph enrichment, nor parameter-efficient fine-tuning can restructure this space without inheriting the capacity ceiling of the underlying encoder. Our key insight is that a learnable graph layer operating on token-level PLM embeddings can acquire corpus-specific semantic structure that the frozen encoder lacks, because token-level graphs preserve document-local context that word-level representations discard and joint optimization with the topic objective reshapes embedding geometry directly from target-domain evidence. We instantiate this insight as DARTopic, a domain-agnostic framework that constructs token-level semantic graphs from frozen PLM embeddings and jointly trains a GNN encoder with topic inference. Across three benchmarks spanning general, biomedical, and legal domains, DARTopic consistently outperforms strong baselines in topic coherence and document clus- tering without any encoder fine-tuning, while demonstrating robustness to PLM choice and favorable runtime efficiency over fine-tuning based alternatives.
摘要:最近在使用預訓練語言模型(PLMs)的神經主題模型方面取得了顯著進展,通過利用通用領域的預訓練達到了強大的性能,然而它們在專門語料上的主題可解釋性往往會下降。這一限制主要源於嵌入空間的幾何特性,其中在預訓練期間未見過的領域特定術語會塌縮到一個無法區分的區域,而領域特定的再訓練、詞級圖增強或參數高效的微調都無法在不繼承基礎編碼器的容量上限的情況下重構這一空間。我們的關鍵見解是,運行在標記級PLM嵌入上的可學習圖層可以獲得凍結編碼器所缺乏的語料特定語義結構,因為標記級圖保留了文檔局部上下文,而詞級表示則被丟棄,並且與主題目標的聯合優化直接從目標領域的證據重塑嵌入幾何。我們將這一見解具體化為DARTopic,一個與領域無關的框架,從凍結的PLM嵌入構建標記級語義圖,並與主題推斷共同訓練GNN編碼器。在涵蓋通用、生物醫學和法律領域的三個基準測試中,DARTopic在主題一致性和文檔聚類方面始終超越強基準,且無需任何編碼器微調,同時對PLM選擇展現出穩健性,並在運行效率上優於基於微調的替代方案。
Picking the Right Image to Classify: Reliable-Input Selection in Teledermatology
2608.16198v1 by Fabian Gröger, Marco Weishaupt, Philippe Gottfrois, Simone Lionetti, Linda Wermelinger, Nipun Ranasekara, Ludovic Amruthalingam, Alexander A. Navarini, Marc Pouly
Dermatology models face distribution shifts in teledermatology settings, where submitted images differ from the training data in lighting, angle, distance, focus, and framing. These test-time images are ordinary clinical photographs, but some fall outside the model's training conditions, leading the model to often misclassify them due to shifts in acquisition between training and deployment. When multiple images of the same case exist (several photos of one patient or lesion), a natural way to improve accuracy is therefore to select the image the model is most likely to classify correctly. We call this task reliable-input selection. An oracle that, for each case, selects a correctly classified image when one exists raises weighted F1 by about 20 percentage points on average across six dermatology datasets and nine frozen backbones. This oracle is an upper bound that sees the labels, whereas a selector must choose blindly. Capturing this gain in practice is hard. A selector that needs no pretraining data applies to any frozen model, including those whose data is not public. It must judge reliability from quantities the model exposes at inference: its embeddings, their norms, and its confidence. We benchmark four such training-data-free selectors: the embedding norm, the neighborhood consensus among a case's images, the stability of the prediction under small perturbations, and the model's own confidence. No training-data-free selector substantially narrows this oracle gap. The best of them is the model's own confidence, but it recovers only a small part of the gap on the clinical datasets. A small labeled reference set does not help either: the best selector overall, a fusion of confidence and Mahalanobis distance, still leaves most of the gap. To our knowledge, this is the first study to introduce and benchmark reliable input selection, a clinically important, unsolved task.
摘要:皮膚科模型面臨在遠程皮膚科環境中分佈轉移的挑戰,提交的圖像在照明、角度、距離、焦點和構圖上與訓練數據有所不同。這些測試時的圖像是普通的臨床照片,但有些超出了模型的訓練條件,導致模型經常因訓練與部署之間的獲取差異而錯誤分類。當同一病例存在多張圖像(多張同一患者或病變的照片)時,改善準確性的自然方法是選擇模型最有可能正確分類的圖像。我們稱這個任務為可靠輸入選擇。對於每個案例,當存在正確分類的圖像時,選擇這樣的圖像的神諭平均提高六個皮膚科數據集和九個凍結骨幹的加權F1約20個百分點。這個神諭是一個上限,能看到標籤,而選擇器必須盲目選擇。在實踐中捕捉這一增益是困難的。需要無預訓練數據的選擇器適用於任何凍結模型,包括那些數據不公開的模型。它必須根據模型在推斷時暴露的數量來判斷可靠性:其嵌入、它們的範數和模型的信心。我們基準測試了四種無訓練數據的選擇器:嵌入範數、案例圖像之間的鄰域共識、在小擾動下預測的穩定性,以及模型自身的信心。沒有一種無訓練數據的選擇器能顯著縮小這一神諭差距。其中最好的選擇器是模型自身的信心,但它在臨床數據集上僅恢復了差距的一小部分。小型標記參考集也沒有幫助:整體最佳的選擇器,即信心和馬哈拉諾比斯距離的融合,仍然留下了大部分差距。據我們所知,這是第一項引入和基準測試可靠輸入選擇的研究,這是一個臨床重要的未解決任務。
LENS: In-Context Search via Latent Evidence Exploration over Dynamic Raw Documents
2608.16185v2 by Xingjun Wang, Gongsheng Li, Qi Fan, Yunlin Mao, Luyan Su, Yingda Chen
LLM agents increasingly answer questions over dynamic raw-document collections, where files may change before preprocessing, and relevant evidence (spans, sections, pages, or tables) is query-dependent. Existing retrieval-augmented approaches pre-materialize evidence via fixed chunking, embeddings, or persistent indexes: effective for lookup, yet costly, stale-prone, and committed to a granularity before the query is known. We formulate in-context search as Budgeted Evidence Localization over a latent evidence space induced by dynamic raw documents and propose LENS (Latent Evidence Exploration and Search), an index-free framework. Instead of pre-materializing the evidence space, LENS maintains a query-conditioned belief over candidate units, iteratively selecting candidates via complementary lexical, local, and exploratory proposal policies, updating the belief via an LLM relevance oracle, and narrowing toward high-posterior regions under a controllable budget. Evidence is consolidated into compact, source-grounded regions of interest and compressed into self-organizing knowledge clusters reused across related queries. On a controlled 500-question evaluation with matched corpus snapshots, LENS reaches 62.4% exact match and 84.8% evidence recall vs. 65.2% exact match but 50.4% evidence recall for a ReAct-style baseline. Across scales, LENS gives the strongest supporting-fact localization and answer grounding. On a fixed 150-question fullwiki subset over the raw Wikipedia dump with zero indexing, LENS and ReAct are nearly tied in official answer quality (43.3% vs. 42.7% EM), with LENS grounding more answers in retrieved evidence (84.0% vs. 70.7%). A no-retrieval Closed-Book reference highlights the contribution of model memory. LENS is query-ready after corpus changes, needs no preprocessing or persistent index, and preserves source-grounded evidence localization throughout.
摘要:LLM 代理越來越多地在動態原始文件集合中回答問題,這些文件在預處理之前可能會發生變化,相關證據(範圍、部分、頁面或表格)依賴於查詢。現有的檢索增強方法通過固定分塊、嵌入或持久索引預先生成證據:對於查詢來說有效,但成本高、容易過時,並且在查詢已知之前就已確定了粒度。
我們將上下文搜索公式化為預算證據定位,這是在動態原始文件所誘導的潛在證據空間上進行的,並提出了 LENS(潛在證據探索與搜索),這是一個無索引的框架。
LENS 不是預先生成證據空間,而是維持對候選單位的查詢條件信念,通過互補的詞彙、局部和探索性提議策略迭代選擇候選者,通過 LLM 相關性預言者更新信念,並在可控預算下收斂到高後驗區域。
證據被整合成緊湊的、以來源為基礎的關注區域,並壓縮成自組織的知識集群,以便在相關查詢中重複使用。
在一個控制的 500 問題評估中,配對語料庫快照,LENS 達到 62.4% 的精確匹配和 84.8% 的證據召回,而 ReAct 標準基線則為 65.2% 的精確匹配和 50.4% 的證據召回。
在各種規模上,LENS 提供了最強的支持事實定位和答案基礎。在一個固定的 150 問題的 fullwiki 子集上,使用原始維基百科轉儲且沒有索引,LENS 和 ReAct 在官方答案質量上幾乎持平(43.3% 對 42.7% EM),而 LENS 在檢索證據中基礎了更多的答案(84.0% 對 70.7%)。
一個無檢索的閉卷參考突顯了模型記憶的貢獻。LENS 在語料庫變更後隨時可以查詢,無需預處理或持久索引,並在整個過程中保持來源基礎的證據定位。
Agent-Native Telemetry: Verifiable State-Delta Evidence for Autonomous Operations
2608.16178v1 by Jun He, Deying Yu
Operational telemetry is predominantly engineered for human reading: systems repeatedly serialize verbose prose, static keys, and redundant context across billions of log lines. As autonomous AI agents become primary operational consumers, feeding them traditional logs wastes scarce context capacity parsing lexical syntax rather than reasoning over system state changes -- all while lacking cryptographic guarantees of provenance or collection completeness. This paper introduces agent-native telemetry, an operational evidence architecture for autonomous machine operators founded on verifiable state deltas rather than human prose. We present the Agent Telemetry Protocol (ATP) and the State-Delta Evidence Ledger, an implementation that structures operational facts into four core evidence primitives (Transitions, Observations, Relations, and State Checkpoints) governed by content-addressed schemas, while isolating uncurated text as digest-verified opaque references. Producers sign and hash-chain batches for atomic collector append. Verified records feed two parallel agent access paths: a stateless protocol decoder emitting compact positional rows, and a stateful semantic gateway serving bounded graph capsules. We prove an information-preservation lower bound and formalize a ledger-relative verified negative theorem for provable event non-occurrence. On distributed microservice benchmarks (AIOpsLab and OpenTelemetry Astronomy Shop), ATP reduces raw wire payload and modeled cloud query scan costs by 96.4% relative to OpenTelemetry JSON, reduces LLM context tokens by 88.8% and query operations by 66.2%, detects all 500 tested adversarial storage mutations, and yields zero successful prompt injections across 50 adversarial trials per ATP configuration.
摘要:操作遙測主要是為了人類閱讀而設計:系統重複序列化冗長的散文、靜態鍵和冗餘的上下文,跨越數十億條日誌。隨著自主 AI 代理成為主要的操作消費者,向它們提供傳統日誌會浪費稀缺的上下文容量,解析詞法語法而不是推理系統狀態變化——同時缺乏來源或收集完整性的加密保證。
本論文介紹了代理原生遙測,這是一種基於可驗證狀態增量的自主機器操作員的操作證據架構,而非人類散文。我們提出了代理遙測協議 (ATP) 和狀態增量證據賬本,這是一種將操作事實結構化為四個核心證據原語(轉換、觀察、關係和狀態檢查點)的實現,受內容地址模式的管理,同時將未經策劃的文本隔離為摘要驗證的模糊參考。
生產者簽名並哈希鏈批次以進行原子收集器附加。經過驗證的記錄提供兩條平行的代理訪問路徑:一個無狀態的協議解碼器發出緊湊的位置信息行,和一個有狀態的語義網關提供有界圖膠囊。我們證明了信息保存的下限,並形式化了一個賬本相對的驗證負定理,以證明事件不發生的可證性。在分佈式微服務基準測試(AIOpsLab 和 OpenTelemetry Astronomy Shop)中,ATP 相對於 OpenTelemetry JSON 減少了 96.4% 的原始網絡有效載荷和建模雲查詢掃描成本,減少了 88.8% 的 LLM 上下文標記和 66.2% 的查詢操作,檢測到所有 500 次測試的對抗存儲突變,並在每個 ATP 配置的 50 次對抗試驗中產生零次成功的提示注入。
FeatureHospital: A Skill-Driven Multi-Agent Framework for Automated Algorithm Customization in Multi-View Multi-Label Feature Selection
2608.16148v1 by Junxuan Li, Zhiqi Chen, Yuzhou Liu, Peng Zhang, Huaxiao Liu
Multi-view multi-label feature selection aims to identify a compact and informative feature subset from heterogeneous views while preserving discriminative information for multiple labels. Existing methods are generally developed from specific modeling perspectives and incorporate mechanisms tailored to particular data characteristics. Designing suitable feature selection algorithms across datasets with diverse and heterogeneous characteristics still relies heavily on expert knowledge and substantial manual effort, imposing considerable time and labor costs that severely hinder the practical adoption of feature selection. To address this problem, we propose FeatureHospital, a Skill-driven multi-agent framework for automated multi-view multi-label feature selection algorithm design. FeatureHospital first diagnoses the target dataset to identify its feature selection issues. Based on the diagnosis, specialist agents equipped with domain Skills then prescribe corresponding optimization strategies and Loss terms for different issues. After that, the resulting prescriptions are reconciled to remove overlaps and resolve conflicts before being integrated into a compact dataset-specific objective. Finally, the constructed objective is optimized to select the final feature subset. Experimental results demonstrate that FeatureHospital can construct effective feature selection algorithms for different datasets based on their individual characteristics.
摘要:多視角多標籤特徵選擇旨在從異質視角中識別出緊湊且具信息量的特徵子集,同時保留多個標籤的區別信息。現有的方法通常是從特定建模角度發展而來,並納入針對特定數據特徵量身定制的機制。在具有多樣且異質特徵的數據集上設計合適的特徵選擇算法仍然在很大程度上依賴於專家知識和大量的手動努力,這帶來了可觀的時間和勞動成本,嚴重阻礙了特徵選擇的實際應用。為了解決這一問題,我們提出了FeatureHospital,一個以技能驅動的多代理框架,用於自動化多視角多標籤特徵選擇算法的設計。FeatureHospital首先診斷目標數據集,以識別其特徵選擇問題。根據診斷,配備領域技能的專家代理隨後為不同問題開出相應的優化策略和損失項。之後,生成的處方會進行調和,以消除重疊並解決衝突,然後整合成一個緊湊的數據集特定目標。最後,構建的目標被優化以選擇最終的特徵子集。實驗結果表明,FeatureHospital能夠根據不同數據集的特徵構建有效的特徵選擇算法。
Graph Neural Assisted Actor-Critic for Latency-Efficient Edge Vision System
2608.16142v1 by Alam Noor, Luis Almeida, Kai Li, Jiyan Wu, Miguel Gutiérrez Gaitán, Eduardo Tovar
UAV on-board vision systems are widely used for different activities, including monitoring in no-fly zones. In this case, the vision-equipped UAV streams a video to a ground server where an operator assists its activities. The latency of video transmission has a profound impact on the effectiveness of the operator assistance. However, most techniques available for video transmission still incur significant latency costs. In this paper, we propose a graph convolutional neural network-assisted (GCN-Assisted A2C) deep reinforcement learning (DRL) system model to find the optimal pixel-correlated area of a suspicious object. We combine the Lagrangian dual form with gradient descent to prevent lack of convergence and over- and under-penalization constraint violation during latency optimization. The proposed system model sends a sub-group pixel-correlated area of the frame from the UAV to the server rather than the transmission of the whole video frame. The proposed framework utilizes the GCN model to explore hidden representations of feature-correlated groups of pixels. Moreover, the GCN supervises the A2C model, which selects a subgroup to enhance transmission latency, thus supervising the training of UAV actions in A2C. Experimental results show that GCN-assisted A2C reduces video frame transmission latency together with false detection rate in UAV vision systems over other DRL and state-of-the-art models.
摘要:無人機的機載視覺系統被廣泛應用於不同的活動,包括在禁飛區的監控。在這種情況下,配備視覺系統的無人機將視頻流傳輸到地面伺服器,操作員在那裡協助其活動。視頻傳輸的延遲對操作員的輔助效果有深遠的影響。然而,目前大多數可用的視頻傳輸技術仍然會產生顯著的延遲成本。在本文中,我們提出了一種圖卷積神經網絡輔助(GCN輔助A2C)深度強化學習(DRL)系統模型,以尋找可疑物體的最佳像素相關區域。我們將拉格朗日對偶形式與梯度下降相結合,以防止在延遲優化過程中缺乏收斂以及過度和不足懲罰約束違規。所提出的系統模型將無人機的幀中一個子組像素相關區域發送到伺服器,而不是傳輸整個視頻幀。所提出的框架利用GCN模型探索特徵相關像素組的隱藏表示。此外,GCN監督A2C模型,該模型選擇一個子組以增強傳輸延遲,從而監督無人機在A2C中的行動訓練。實驗結果顯示,GCN輔助A2C在無人機視覺系統中減少了視頻幀傳輸延遲以及假檢測率,相較於其他DRL和最先進的模型。
HyperSkill: Self-Evolving LLM Agents via Hypergraph-Structured Skill Memory
2608.16114v1 by Ruiyao Xu, Tiankai Yang, Wei-Chieh Huang
As agentic tasks grow in complexity, LLM agents increasingly rely on experiential memory to reuse procedural knowledge across tasks. Effective memory design must jointly address what to store, how memory is structured and retrieved, and how memory evolves. Existing systems tackle each only partially: they store trajectories, insights, or workflows as isolated entries, discarding compositional relationships among subtasks and reusable skills; retrieve by flat embedding similarity that ignores relational signals; and maintain memory without leveraging its relational structure. We propose HyperSkill, a hypergraph-based memory framework that jointly improves all three. HyperSkill represents memory as a hypergraph with two node types, subtask steps and reusable skills, where each hyperedge links the subtasks and skills from a single trajectory. Dual-path retrieval queries both subtask and trajectory levels, ranking skills by co-occurrence across retrieved trajectories. Periodic structure-informed maintenance prunes low-utility nodes and merges redundant skills via quality-weighted propagation. Across xBench, GAIA, and WebWalkerQA with GPT-4o and Qwen3-30B-A3B, HyperSkill outperforms ten memory baselines, yielding gains of up to +11.51 on GAIA and +11.18 on WebWalkerQA.
摘要:隨著代理任務的複雜性增加,LLM代理越來越依賴經驗記憶在任務之間重用程序知識。有效的記憶設計必須共同解決存儲什麼、記憶的結構和檢索方式,以及記憶的演變。現有系統僅部分解決這些問題:它們將軌跡、見解或工作流程作為孤立的條目進行存儲,忽略了子任務和可重用技能之間的組合關係;通過忽略關聯信號的平面嵌入相似性進行檢索;並維護記憶而不利用其關聯結構。我們提出了HyperSkill,一種基於超圖的記憶框架,旨在共同改進這三個方面。HyperSkill將記憶表示為一個超圖,具有兩種類型的節點:子任務步驟和可重用技能,其中每個超邊連接來自單一軌跡的子任務和技能。雙路徑檢索查詢同時針對子任務和軌跡層級,根據檢索到的軌跡中的共同出現對技能進行排名。定期的結構信息維護修剪低效用節點,並通過質量加權傳播合併冗餘技能。在xBench、GAIA和WebWalkerQA中,使用GPT-4o和Qwen3-30B-A3B的HyperSkill超越了十個記憶基準,GAIA的增益高達+11.51,WebWalkerQA的增益高達+11.18。
RetroMPA: A Molecular Property-Aware Auxiliary Framework for Enhancing Retrosynthesis Prediction
2608.16111v1 by Mianzhi Liu, Fan Xiao, Zhiliang Yu, Huayang Huang, Yuke Li, Yi Yang, Wenbo Liu, Yu Wu
Retrosynthesis is a cornerstone of drug discovery and organic synthesis. While data-driven deep learning models have shown remarkable progress, they autonomously learn reaction patterns from extensive datasets with limited integration of established chemical knowledge as priors. To address this limitation, we introduce RetroMPA, a molecular property-aware, post-hoc enhancement module that injects chemical knowledge into the retrosynthesis pipeline. Rather than functioning as an independent SMILES sequence generator, RetroMPA is a broadly applicable, model-agnostic chemical filter designed to recalibrate and optimize the predictive pathways of existing algorithms. This plug-and-play framework integrates seamlessly with a range of data-driven retrosynthesis methods, enhancing outputs without modifying model architecture or requiring resource-intensive retraining. By leveraging a property-aware latent embedding space, RetroMPA consistently improves top-1 accuracy across eight representative retrosynthesis models by an average of 5.50% on USPTO-50K. Furthermore, we validate its scalability on the large-scale USPTO-Full dataset, achieving an average improvement of about 2.03% across both template-based and template-free architectures. Wet-lab experiments provide preliminary support for the practical utility of the framework. These syntheses confirmed viable, previously unreported substrate combinations for classic reaction paradigms---specifically, Suzuki-Miyaura coupling, Bucherer reaction, and Friedel-Crafts acylation---suggesting that RetroMPA can operate beyond mere data fitting. The code is open-sourced at https://github.com/MengzhouLu/RetroMPA.
摘要:逆合成是藥物發現和有機合成的基石。儘管數據驅動的深度學習模型已顯示出顯著的進展,但它們從大量數據集中自主學習反應模式,對已建立的化學知識的整合有限。
為了解決這一限制,我們引入了RetroMPA,一種分子性質感知的後處理增強模塊,將化學知識注入逆合成流程中。
RetroMPA並不是作為獨立的SMILES序列生成器運作,而是一種廣泛適用的、與模型無關的化學過濾器,旨在重新校準和優化現有算法的預測路徑。
這個即插即用的框架與多種數據驅動的逆合成方法無縫集成,增強輸出而不修改模型架構或需要資源密集的重新訓練。
通過利用性質感知的潛在嵌入空間,RetroMPA在八個代表性的逆合成模型上,平均提高了5.50%的top-1準確率,數據集為USPTO-50K。
此外,我們在大規模的USPTO-Full數據集上驗證了其可擴展性,在基於模板和無模板架構上都實現了約2.03%的平均改進。
濕實驗提供了對該框架實用性的初步支持。這些合成確認了經典反應範式下可行的、先前未報告的底物組合——具體而言,鈴木-宮浦偶聯、布赫勒反應和弗里德爾-克拉夫茲酰化——這表明RetroMPA可以超越單純的數據擬合。
代碼已開源於 https://github.com/MengzhouLu/RetroMPA。
The Commercial Tax: Rent-vs-Own Blind Spots in Multi-Hop Retrieval Benchmarks
2608.16096v1 by Luis M. Sanchez, Kosrow Dehnad
Enterprises connect language models to their own data through retrieval. The benchmarks that rank multi-hop retrieval systems leave out two facts a buyer needs before a published number can be used: whether the retrieval backbone may be deployed commercially, and what it costs to build. On licensing: the field's dense-retrieval anchor, NV-Embed-v2, is licensed cc-by-nc-4.0. Of the four leading MuSiQue systems we audit (HippoRAG-2, PropRAG, SAG, KET-RAG), three depend on it for their best numbers and none says so. On performance: we measure thirteen embedders from eight makers on one identical MuSiQue harness with bootstrap confidence intervals throughout. Until mid-2026 there was a real commercial tax: the best commercially-licensed embedder trailed the anchor by 2.31 Recall@5 points (95% CI [0.91, 3.71], p=0.001). NVIDIA's Nemotron-3-Embed-8B, released 2026-07-16, has closed it: +0.24 at Recall@5 (95% CI [-0.94, +1.43], p=0.69), -0.58 at Recall@10 (p=0.28). It matches the anchor, does not beat it, and is the only entrant that is commercially licensed, free to self-host, and indistinguishable from the anchor; every other entrant meeting the first two conditions sits 5.2 to 14.6 points below. The durable finding is the paid-versus-free divide: API embedders charge per token on every re-index, self-hosted ones charge nothing. On cost: three of five audited systems (adding Microsoft's GraphRAG) do not disclose indexing cost, and the only published GraphRAG dollar figures span 11x inside one third-party paper (USD 2.30 vs USD 24.94 to index a 5.64 MB corpus once); extrapolated to 1 TB that undisclosed choice separates roughly USD 428K from $4.6M. Our cost model keeps one-time embedding apart from recurring answering: at 1 TB, embedding sits 7.5x-900x below graph construction, and a year of answering at 10,000 queries/day sits 350x or more below it.
摘要:企業透過檢索將語言模型與自身數據連接起來。對於買家來說,在發佈的數字可以使用之前,有兩個事實是多跳檢索系統的基準未考慮的:檢索骨幹是否可以商業部署,以及建造的成本。關於授權:該領域的密集檢索錨點 NV-Embed-v2 採用 cc-by-nc-4.0 授權。在我們審核的四個主要 MuSiQue 系統(HippoRAG-2、PropRAG、SAG、KET-RAG)中,有三個依賴於它以獲得最佳數字,但沒有一個明言這一點。關於性能:我們測量了八家製造商的十三個嵌入器,並在一個相同的 MuSiQue 繫帶上進行測試,並在整個過程中使用自助信心區間。直到 2026 年中,存在一個實際的商業稅:最佳的商業授權嵌入器在 Recall@5 上落後於錨點 2.31 分(95% CI [0.91, 3.71],p=0.001)。NVIDIA 的 Nemotron-3-Embed-8B 於 2026-07-16 發佈,已經縮短了這一差距:在 Recall@5 上增加了 0.24(95% CI [-0.94, +1.43],p=0.69),在 Recall@10 上減少了 0.58(p=0.28)。它與錨點相匹配,但未超越它,並且是唯一一個商業授權、可自由自我托管且與錨點無法區分的參賽者;其他符合前兩個條件的參賽者則低於 5.2 至 14.6 分。持久的發現是付費與免費的區分:API 嵌入器在每次重新索引時按令牌收費,自我托管的則不收費。關於成本:五個審核系統中的三個(加上微軟的 GraphRAG)未披露索引成本,而唯一發佈的 GraphRAG 美元數字在一篇第三方論文中跨越了 11 倍(索引一次 5.64 MB 語料庫的成本為 USD 2.30 與 USD 24.94);推算到 1 TB,這一未披露的選擇大約將 USD 428K 與 $4.6M 隔開。我們的成本模型將一次性嵌入與重複回答區分開來:在 1 TB 時,嵌入成本低於圖形構建 7.5 倍至 900 倍,而一年內以 10,000 次查詢/天進行回答的成本則低於它 350 倍或更多。
Skill2Query: Exploiting Skill Structure to Generate Pseudo-Queries for Agent Skill Retrieval
2608.16071v1 by Lihui Ding, Zihan Guo, Bingwei Lu, Chenyu Zhou, Yuanjian Zhou, Weinan Zhang, Jianghao Lin, Dongdong Ge
Pseudo-query generation can alleviate the supervision bottleneck for agent skill retrieval, but existing document-level approaches typically leave the rich internal relations among capabilities, parameters, and usage examples implicit. As a result, generated queries may be topically relevant to a skill while lacking capability grounding and parameter consistency, raising the question of whether explicitly exploiting a skill document's internal structure can produce more effective retrieval signals. We therefore propose Skill2Query, a framework that first parses a skill document into a Skill Knowledge Graph and then generates pseudo-queries through a three-stage process including style mimicking, query template generation, and parameter filling. The generated queries can be used for offline index augmentation, online query expansion, and retriever training. Four benchmarks (TheoremQA, LogicBench, ToolQA, and CHAMP) are used to evaluate Skill2Query with large-scale skill candidate pools across multiple downstream applications, including skill retrieval, retriever training, and end-to-end agent execution. Using nearly 30K skills across diverse domains, we generate 700K category-diverse pseudo-queries. Skill2Query consistently improves sparse, dense, and skill-routing retrieval, with an average Recall@1 gain of 6.70 percentage points across retrieval settings. Skill2Query-generated training data also achieves the best Recall@1 and nDCG@1 among the evaluated generation baselines. Further evaluations with multiple LLM backends demonstrate that improved skill retrieval translates into higher agent task success rates. Code and resources are available at https://github.com/MatZaharia/Skill2Query.
摘要:偽查詢生成可以減輕代理技能檢索的監督瓶頸,但現有的文件級方法通常將能力、參數和使用範例之間豐富的內部關係隱含化。因此,生成的查詢可能在主題上與某項技能相關,但缺乏能力基礎和參數一致性,這引發了明確利用技能文件內部結構是否能產生更有效檢索信號的問題。因此,我們提出了Skill2Query,一個首先將技能文件解析為技能知識圖譜的框架,然後通過三個階段的過程生成偽查詢,包括風格模仿、查詢模板生成和參數填充。生成的查詢可以用於離線索引增強、在線查詢擴展和檢索器訓練。四個基準(TheoremQA、LogicBench、ToolQA和CHAMP)被用來評估Skill2Query,涵蓋多個下游應用中的大規模技能候選池,包括技能檢索、檢索器訓練和端到端代理執行。使用近30K個來自不同領域的技能,我們生成了700K個類別多樣的偽查詢。Skill2Query在稀疏、密集和技能路由檢索中始終提高了性能,平均Recall@1增益為6.70個百分點。Skill2Query生成的訓練數據在評估的生成基準中也實現了最佳的Recall@1和nDCG@1。對多個LLM後端的進一步評估顯示,改善的技能檢索轉化為更高的代理任務成功率。代碼和資源可在 https://github.com/MatZaharia/Skill2Query 獲得。
OceanLight: Efficient Global Ocean Forecasting via Geometry-Adaptive Unstructured Mesh Representation
2608.16070v1 by Wei Wu, Xiang Wang, Hongze Leng, Qingye Min, Junxing Zhu, Junqiang Song
Reliable global ocean forecasting is critical for climate monitoring, marine navigation, and extreme event early warning. Physics-based ocean forecasting models impose prohibitive computational costs, while existing deep learning approaches predominantly rely on structured-grid architectures, incurring unnecessary computation on masked land cells and enforcing uniform resolution across dynamically heterogeneous ocean regions regardless of local flow complexity. Here we present OceanLight, an efficient global ocean forecasting framework innovatively combining geometry-adaptive unstructured mesh tokenization with a graph neural network (GNN) backbone. OceanLight achieves pointwise forecast accuracy and kinetic energy spectral fidelity exceeding both operational numerical analyses and state-of-the-art AI-based models, while surpassing all AI-based ocean models in geostrophic balance consistency. Furthermore, OceanLight demonstrates reliable mesoscale eddy representation, capturing coherent ocean structures beyond pointwise statistical optimization. These capabilities are delivered with a 62% reduction in GPU memory consumption and 70\% reduction in FLOPs relative to structured-grid baselines. Our unstructured mesh representation establishes a generalizable paradigm for scalable data-driven oceanography.
摘要:可靠的全球海洋預測對於氣候監測、海洋導航和極端事件的早期預警至關重要。基於物理的海洋預測模型需要高昂的計算成本,而現有的深度學習方法主要依賴於結構化網格架構,對被遮蔽的陸地單元產生不必要的計算,並在動態異質的海洋區域強制執行均勻的解析度,無論當地流動的複雜性如何。在此,我們提出了OceanLight,一個高效的全球海洋預測框架,創新地結合了幾何自適應的非結構化網格標記與圖神經網絡(GNN)主幹。OceanLight實現了逐點預測準確性和動能能量譜的保真度,超過了操作性數值分析和最先進的基於AI的模型,同時在地轉平衡一致性方面超越了所有基於AI的海洋模型。此外,OceanLight展示了可靠的中尺度渦旋表徵,捕捉到超越逐點統計優化的連貫海洋結構。這些能力在GPU內存消耗上減少62%以及相對於結構化網格基準的FLOPs減少70%下得以實現。我們的非結構化網格表徵建立了一個可普遍化的範式,以支持可擴展的數據驅動海洋學。
NICE: Scale-Stable Perturbations for Graph Neural Network Explanations via Noise Corruption
2608.16038v2 by Ziluowen Luo, Jun Yin, Ruochen Liu, Ming Cheng, Shirui Pan, Chengqi Zhang, Senzhang Wang
Post-hoc Graph Neural Network (GNN) explainers commonly follow a Perturb-Query paradigm, inferring the importance of graph elements based on queried predictions to perturbed inputs. However, such perturbations often introduce substantial distribution shift, undermining the reliability of the queried predictions used to derive explanations. While existing efforts mainly improve perturbed graphs or stabilize model predictions on them, we revisit the perturbation mechanism itself. We show that the widely used Element-wise Masking(EM) suppresses edge-induced messages toward zero, causing deterministic scale contraction that accumulates across message-passing layers, a phenomenon we term Scale Drift. Consequently, prediction changes under EM may conflate information corruption with deviations in propagation scale. As a scale-stable alternative to EM, we introduce Noise Corruption (NC), which perturbs each message through matched-norm random-direction corruption while preserving the expected squared message norm. Building on NC, we propose NICE, a Noise Corruption-based explanation framework, which learns a Stochastic Restoration Boundary (SRB) under NC-induced uncertainty, balancing target-prediction restoration against compactness. Furthermore, Boundary-Integrated Gradient (BIG) converts this boundary into edge attributions by accumulating each edge's contribution to reducing restoration risk along the restoration path. Experiments across multiple benchmarks demonstrate stronger explanation performance and model faithfulness while confirming that NC substantially reduces the Scale Drift induced by masking.
摘要:後 hoc 圖神經網絡 (GNN) 解釋器通常遵循擾動-查詢範式,根據對擾動輸入的查詢預測推斷圖元素的重要性。
然而,這種擾動往往會引入顯著的分佈變化,削弱用於推導解釋的查詢預測的可靠性。
雖然現有的努力主要改善擾動圖或穩定模型在其上的預測,但我們重新審視擾動機制本身。
我們展示了廣泛使用的逐元素掩蔽 (EM) 將邊緣引起的消息壓制至零,導致確定性的縮放收縮,這一現象在消息傳遞層中累積,我們稱之為縮放漂移。
因此,EM 下的預測變化可能將信息損壞與傳播縮放的偏差混淆在一起。
作為 EM 的一種穩定縮放替代方案,我們引入了噪聲損壞 (NC),它通過匹配範數的隨機方向擾動每條消息,同時保持期望的平方消息範數。
基於 NC,我們提出了 NICE,一種基於噪聲損壞的解釋框架,它在 NC 引起的不確定性下學習隨機恢復邊界 (SRB),平衡目標預測恢復與緊湊性。
此外,邊界整合梯度 (BIG) 通過累積每條邊對減少恢復風險的貢獻,將這一邊界轉換為邊緣歸因。
在多個基準上的實驗顯示出更強的解釋性能和模型忠實度,同時確認 NC 顯著減少了由掩蔽引起的縮放漂移。
RagGAD: Rationale-Aware Conditional Gaussian Mixture Normalizing Flow for Unsupervised Graph Anomaly Detection
2608.16018v1 by Junxin Lu, Jing Zhao, Shiliang Sun
Graph anomaly detection aims to identify nodes that deviate from normal behavioral patterns within graphs. However, existing methods largely rely on the homophily assumption, which makes it difficult to distinguish spurious affinities and to capture the diverse behaviors of normal nodes,limiting their robustness in complex real-world scenarios. To address this problem, we propose RagGAD, an unsupervised graph anomaly detection framework based on rationale-aware conditional Gaussian mixture normalizing flow. RagGAD introduces an adaptive rationale disentangler to disentangle stable rationales from spurious correlations within node interrelationships, and further decomposes stable rationales into robust and fragile components. The learned rationales capture underlying interaction patterns that characterize normal behaviors under varying conditions, while anomalies emerge as deviations associated with unstable or spurious correlations. To model the intricate distributions of normal and abnormal nodes, RagGAD integrates rationale-non-rationale Gaussian mixture modeling with a robust-fragile rationale mixture learning strategy. By mitigating spurious homophilic correlations and embracing the heterogeneity of normal patterns, RagGAD identifies anomalies as low-density regions within a structure-aware distribution space. Extensive experiments on multiple benchmark datasets demonstrate that RagGAD outperforms state-of-the-art methods.
摘要:圖形異常檢測旨在識別在圖中偏離正常行為模式的節點。
然而,現有的方法在很大程度上依賴於同質性假設,這使得區分虛假的親和力和捕捉正常節點的多樣行為變得困難,限制了它們在複雜現實場景中的穩健性。
為了解決這個問題,我們提出了 RagGAD,一個基於理性感知條件高斯混合正規化流的無監督圖形異常檢測框架。
RagGAD 引入了一個自適應的理性解耦器,以從節點之間的關係中解耦穩定的理性與虛假的相關性,並進一步將穩定的理性分解為穩健和脆弱的組件。
學習到的理性捕捉了在不同條件下表徵正常行為的潛在互動模式,而異常則作為與不穩定或虛假相關性相關的偏差出現。
為了建模正常和異常節點的複雜分佈,RagGAD 將理性-非理性高斯混合建模與穩健-脆弱理性混合學習策略相結合。
通過減少虛假的同質相關性並接受正常模式的異質性,RagGAD 將異常識別為結構感知分佈空間中的低密度區域。
在多個基準數據集上的廣泛實驗表明,RagGAD 超越了最先進的方法。
From Sequence to Structure: Relational Uncertainty Propagation for LLM Agents
2608.16002v1 by Zhengzhao Ma. Boxi Cao, Yaojie Lu, Hongyu Lin, Xianpei Han, Le Sun
Reliable uncertainty quantification (UQ) is essential for deploying large language model (LLM) agents in complex interactive environments. Existing UQ methods largely rely on local signals, such as token probabilities, predictive entropy, or per-step confidence, and therefore overlook the long-range dependencies through which errors accumulate across an execution trajectory. As a result, they may fail to identify agent failures whose causes originate several reasoning or interaction steps before the final answer. We propose RUPA (Relational Uncertainty Propagation for Agents), a trajectory-level UQ framework for LLM agents. RUPA represents an execution history as a directed trajectory graph in which reasoning states, tool interactions, and environment feedback are nodes connected by temporal and semantic dependency edges. It then propagates uncertainty over this graph to capture how execution risk accumulates and transfers across interaction steps. The propagated signal is combined with trajectory-level behavioral features and goal-alignment information to produce a confidence estimate for the full agent trajectory. We evaluate RUPA on representative agent benchmarks, including $τ$-2, Terminal-Bench-2, and GAIA, using 6 open-source LLMs spanning multiple model families. Experimental results show that RUPA consistently outperforms existing UQ methods by providing more accurate uncertainty estimates, enabling earlier failure detection, and improving uncertainty-guided agent execution across diverse agent tasks. These results demonstrate that explicitly modeling relational dependency is crucial to reliable UQ for long-horizon LLM agents, providing a practical foundation for trustworthy agent execution.
摘要:可靠的不確定性量化(UQ)對於在複雜互動環境中部署大型語言模型(LLM)代理至關重要。現有的UQ方法主要依賴於局部信號,例如標記概率、預測熵或每步信心,因此忽略了錯誤在執行軌跡中累積的長期依賴關係。結果,它們可能無法識別那些原因源於幾個推理或互動步驟之前的代理失敗。我們提出了RUPA(代理的關聯不確定性傳播),這是一個針對LLM代理的軌跡級UQ框架。RUPA將執行歷史表示為一個有向軌跡圖,其中推理狀態、工具互動和環境反饋是由時間和語義依賴邊連接的節點。然後,它在這個圖上傳播不確定性,以捕捉執行風險如何在互動步驟中累積和轉移。傳播的信號與軌跡級行為特徵和目標對齊信息相結合,以產生對整個代理軌跡的信心估計。我們在代表性的代理基準上評估RUPA,包括$τ$-2、Terminal-Bench-2和GAIA,使用6個跨多個模型系列的開源LLM。實驗結果顯示,RUPA通過提供更準確的不確定性估計、實現更早的失敗檢測以及改善不確定性引導的代理執行,始終優於現有的UQ方法。這些結果表明,明確建模關聯依賴性對於長期LLM代理的可靠UQ至關重要,為可信的代理執行提供了實用的基礎。
PLSQLBench: Benchmarking LLM Systems for Executable Procedural Database Programming
2608.15931v1 by Marianne Menglin Liu, Leonid Boytsov, Daniel W. Peterson, Pramuditha Perera, Rongguang Wang, Sai Ashish Somayajula, Syed Hamza Rafique, Rohit Saini, Shubham Pathak, Sujeeth Bharadwaj, Tao Sheng, Graham Horwood, Fahad Shah, Ankan Bansal, Sujith Ravi, Dan Roth
We present PLSQLBench, to our knowledge the first benchmark for evaluating whether LLMs can write executable PL/SQL programs, with correctness measured through execution-based tests. Existing LLM evaluations largely target general-purpose code generation or declarative text-to-SQL, leaving procedural database programming underexplored. PLSQLBench contains 2,865 instances: 2,594 single-turn tasks and 271 multi-turn conversations spanning 978 turns. The benchmark combines complex schema-grounded tasks over enterprise-style Spider 2 databases, simpler schema-grounded tasks derived from Spider, and MBPP-derived procedural problems, covering varying levels of database grounding and procedural complexity. Experiments with eight LLMs reveal recurring difficulties in schema grounding, PL/SQL dialect fidelity, procedural control flow, exception handling, and cross-turn consistency. Tool-augmented LLM agents improve performance on several schema-grounded evaluations, although substantial gaps remain. These results highlight procedural database programming capabilities not directly assessed by conventional code generation or text-to-SQL benchmarks. Our code is available at https://github.com/oracle-samples/plsqlbench.
摘要:我們介紹PLSQLBench,據我們所知,這是第一個用於評估LLMs是否能夠編寫可執行PL/SQL程序的基準,其正確性通過基於執行的測試來衡量。現有的LLM評估主要針對通用代碼生成或聲明式文本到SQL,導致程序性數據庫編程未得到充分探索。PLSQLBench包含2,865個實例:2,594個單回合任務和271個跨978回合的多回合對話。該基準結合了基於企業風格Spider 2數據庫的複雜架構任務、源自Spider的較簡單架構任務以及MBPP衍生的程序性問題,涵蓋了不同級別的數據庫基礎和程序複雜性。對八個LLM的實驗揭示了在架構基礎、PL/SQL方言忠實度、程序控制流程、異常處理和跨回合一致性方面的重複困難。工具增強的LLM代理在幾個架構基礎評估中提高了性能,儘管仍然存在重大差距。這些結果突顯了傳統代碼生成或文本到SQL基準未直接評估的程序性數據庫編程能力。我們的代碼可在https://github.com/oracle-samples/plsqlbench獲得。
Unified Pedestrian Path Prediction Using Inverse Reinforcement Learning
2608.15929v1 by Šimon Sukup, Ariyan Bighashdel, Pavol Jancura
Pedestrian path prediction is crucial for enhancing the safety of autonomous vehicles and advanced driver-assistance systems. Previous studies explored different learning-task formulations for pedestrian path prediction and compared these formulations using shallow neural networks, but did not extend this analysis to more complex deep-learning models. This paper adapts the Spatial-Temporal Graph Attention Network (STGAT) to a unified pedestrian path prediction framework and introduces state and action definitions specific to STGAT. The resulting formulations support deterministic and stochastic policies, one-time and sequential decision-making, and reinforcement-learning algorithms including REINFORCE and proximal policy optimization. The proposed learning-task formulations improve prediction performance across the selected benchmark datasets compared with the standard supervised-learning formulation. These results demonstrate that reformulating the decision process and training objective can improve an advanced pedestrian trajectory prediction architecture and may provide a path toward improving other graph-based prediction models.
摘要:行人路徑預測對於增強自動駕駛車輛和先進駕駛輔助系統的安全性至關重要。
以往的研究探討了行人路徑預測的不同學習任務公式,並使用淺層神經網絡比較了這些公式,但並未將此分析擴展到更複雜的深度學習模型。
本文將空間-時間圖注意力網絡(STGAT)適應於統一的行人路徑預測框架,並引入特定於STGAT的狀態和行動定義。
所得到的公式支持確定性和隨機策略、一次性和序列決策,以及包括REINFORCE和近端政策優化在內的強化學習算法。
與標準的監督學習公式相比,所提出的學習任務公式在所選基準數據集上提高了預測性能。
這些結果表明,重新構建決策過程和訓練目標可以改善先進的行人軌跡預測架構,並可能為改善其他基於圖的預測模型提供一條途徑。
Noesis: Bidirectional Graph-RAG with Adaptive Parallelism and Cross-Knowledge-Base Semantic Discovery
2608.15919v1 by Nicola Cogotti
Retrieval-Augmented Generation over knowledge graphs (Graph-RAG) has emerged as a powerful paradigm for grounding large language models in domain-specific corpora. However, existing systems face persistent limitations: (1) static chunking fragments long documents, losing cross-section semantic connections; (2) ingestion pipelines do not scale adaptively; and (3) multi-domain deployments require either a monolithic knowledge base that dilutes retrieval precision or manual user routing. We present Noesis, a decoupled Graph-RAG architecture addressing these limitations through four algorithms: (a) Bidirectional Graph Traversal with a Graph-Feedback Context Resolver simulating human reading with degrading memory; (b) an AIMD Concurrency Controller adapted from TCP congestion control, achieving 23x speedup with zero OOM events; (c) Moesis, domain-aware selective quantization for MoE models achieving 6.3x speedup on 12 GB consumer GPUs; and (d) Mesh, cross-KB semantic routing with runtime structural discovery enabling small on-premises models to perform multi-hop cross-domain reasoning. On HotpotQA (1,000 questions), Noesis achieves 59.5 EM / 74.7 F1, surpassing GraphRAG by +27.8 EM while using a 35B on-premises model for graph construction rather than GPT-4o. Source text verification on a 193-page document confirms 90% precision on long-range causal edges inaccessible to chunk-independent extraction.
摘要:檢索增強生成(Retrieval-Augmented Generation)在知識圖譜(Graph-RAG)上已成為一種強大的範式,用於將大型語言模型嵌入特定領域的語料庫中。
然而,現有系統面臨持續的限制:(1)靜態分塊使長文檔碎片化,失去了交叉部分的語義連接;(2)攝取管道無法自適應擴展;(3)多領域部署需要一個單體知識庫,這會稀釋檢索精度或需要手動用戶路由。
我們提出了 Noesis,一種解耦的 Graph-RAG 架構,通過四個算法解決這些限制:(a)雙向圖遍歷(Bidirectional Graph Traversal)與圖反饋上下文解析器(Graph-Feedback Context Resolver),模擬人類閱讀並隨著記憶衰退;(b)從 TCP 擁塞控制改編的 AIMD 並發控制器(AIMD Concurrency Controller),實現了 23 倍的加速,且無 OOM 事件;(c)Moesis,針對 MoE 模型的領域感知選擇性量化,實現了在 12 GB 消費者 GPU 上的 6.3 倍加速;以及(d)Mesh,跨知識庫的語義路由,通過運行時結構發現使小型本地模型能夠執行多跳跨域推理。
在 HotpotQA(1,000 個問題)上,Noesis 實現了 59.5 EM / 74.7 F1,超越了 GraphRAG,EM 提升了 27.8,同時使用 35B 的本地模型進行圖構建,而不是 GPT-4o。
對於一份 193 頁文檔的源文本驗證確認,對於長距離因果邊的精度達到 90%,這些邊是無法通過獨立於塊的提取來訪問的。
Large language model-assisted discovery of cohorts from scientific literature
2608.15909v1 by Moritz Sturm, Lisa M. Berg, Inken Berg, Harishny Sarma, Jasmin Hartmann, Denissa Girschik, Gemma Roig, Christine M. Freitag, Andreas G. Chiocchetti
Background: Planning multi-study analyses requires identifying cohorts with the relevant participants, phenotypes, and data modalities. This process commonly relies on prior knowledge, cohort catalogues, and manual literature searches. We developed a complementary question-driven framework that searches relevant scientific literature and extracts explicit cohort names. Methods: The framework first generates multiple PubMed queries from configurable vocabularies and templates and retrieves the resulting scientific literature automatically through the PubMed API. A large language model then screens the retrieved titles and abstracts and extracts explicit cohort names using a prompt tailored to the research question. The extracted names are deduplicated with human review. Configurable code, prompts, and example outputs are available at https://gitlab.rz.uni-frankfurt.de/cap_molgenlab/literature-cohort-discovery. Evaluation: As a use case, we applied the framework to youth aggression genetics. From 5,400 generated PubMed queries, the framework retrieved 5,254 unique records and identified 188 candidate cohorts. Manual screening using predefined criteria, including participant age and genetic-data availability, retained 44 eligible cohorts. Automated LLM-based name extraction was within the agreement range of human annotators. We also searched four established cohort catalogues using the same research question. Their combined results contained 27 of the 44 eligible cohorts, while 17 were not returned by any cohort catalogue search. Conclusion: The framework converts research-question-specific vocabulary into screenable cohort inventories via a large, automated literature search. It can be adapted across populations, phenotypes, data modalities, and study designs, and provides a literature-based complement to curated cohort catalogues.
摘要:背景:規劃多研究分析需要識別具有相關參與者、表型和數據模態的隊列。這一過程通常依賴於先前的知識、隊列目錄和手動文獻搜索。我們開發了一個補充的問題驅動框架,該框架搜索相關的科學文獻並提取明確的隊列名稱。方法:該框架首先從可配置的詞彙和模板生成多個PubMed查詢,並通過PubMed API自動檢索結果科學文獻。然後,大型語言模型篩選檢索到的標題和摘要,並使用針對研究問題量身定制的提示提取明確的隊列名稱。提取的名稱經過人工審核去重。可配置的代碼、提示和示例輸出可在 https://gitlab.rz.uni-frankfurt.de/cap_molgenlab/literature-cohort-discovery 獲得。評估:作為一個用例,我們將該框架應用於青少年攻擊性遺傳學。從5,400個生成的PubMed查詢中,該框架檢索到5,254個唯一記錄並識別了188個候選隊列。使用預定義標準進行的人工篩選,包括參與者年齡和基因數據可用性,保留了44個合格隊列。自動化的LLM基礎名稱提取與人類標註者的協議範圍內。 我們還使用相同的研究問題搜索了四個已建立的隊列目錄。它們的綜合結果包含44個合格隊列中的27個,而17個則未被任何隊列目錄搜索返回。結論:該框架將特定於研究問題的詞彙轉換為可篩選的隊列清單,通過大型自動化文獻搜索。它可以適應不同的人群、表型、數據模態和研究設計,並為策劃的隊列目錄提供文獻基礎的補充。
Scaling Manual-Grounded Appliance Manipulation with Data Synthesis and Unified Planning
2608.15863v1 by Yuxing Long, Lei Kang, Ziyan Yu, Yuzheng Gao, Bin Cheng, Jiyao Zhang, Xiaoqi Li, Haolin Yang, Dongjiang Li, Hui Shen, Hao Dong
Operating household appliances requires long-horizon planning that is state-dependent and robust to disturbances, yet existing large models fall short, as no sufficiently diverse, task-oriented dataset exists to support such planning. To bridge this gap, we propose MAGE, a scalable data synthesis pipeline that introduces a novel Hierarchical Appliance Graph (HAG) to automatically generate part grounding, long-horizon planning, and closed-loop recovery data from appliance manuals. With MAGE, we build UseAppliance, the first large-scale dataset for manual-grounded appliance manipulation planning, spanning 22 appliance categories with 89K+ part annotations, 53K+ manipulation tasks, and 33K+ closed-loop adjustment steps. Built on UseAppliance, we develop AppliancePlan, an end-to-end model for manual-grounded appliance manipulation planning. On RealAppliance-Bench, AppliancePlan with only 7B parameters achieves over 10x the best baseline on open-loop planning and consistently outperforms state-of-the-art models across all tasks. Real-robot experiments on six household appliances further confirm effective sim-to-real transfer, marking an important step toward general-purpose household robotics.
摘要:操作家用電器需要長期規劃,這種規劃依賴於狀態並且對擾動具有穩健性,然而現有的大型模型卻無法滿足需求,因為沒有足夠多樣化且以任務為導向的數據集來支持這種規劃。為了填補這一空白,我們提出了MAGE,一個可擴展的數據合成管道,該管道引入了一種新穎的層次家電圖(HAG),以自動從家電手冊生成部件定位、長期規劃和閉環恢復數據。通過MAGE,我們構建了UseAppliance,這是第一個基於手冊的家電操作規劃的大型數據集,涵蓋22個家電類別,擁有89K+的部件註釋、53K+的操作任務和33K+的閉環調整步驟。在UseAppliance的基礎上,我們開發了AppliancePlan,一個端到端的基於手冊的家電操作規劃模型。在RealAppliance-Bench上,僅有7B參數的AppliancePlan在開放式規劃上達到了超過10倍的最佳基準,並在所有任務中持續超越最先進的模型。對六種家用電器的真實機器人實驗進一步確認了有效的模擬到現實轉移,這標誌著朝向通用家用機器人邁出了重要一步。
RAGas: Retrieval-Augmented Gas Optimization for Smart Contracts with Continuous Knowledge Integration
2608.15857v1 by Yishun Wang, Wenjin Yi, Wenkai Li, Zongwei Li, Xiaoqi Li
Ethereum is now integral to mission-critical sectors, including finance, healthcare, and supply chain management. Execution fees, commonly referred to as Gas, scale with the computational complexity of their functions. Smart contracts on Ethereum incur execution fees, known as Gas, which increase with computational complexity. Thus, optimizing Gas-intensive code while preserving functional equivalence significantly lowers deployment costs. No existing system continuously exploits evolving Gas usage patterns. We systematically analyze syntactic and semantic constructs that drive excessive Gas use. This yields six high-level categories covering twelve fine-grained antipatterns underpinning a curated knowledge base. We operationalize these insights with RAGas, a three-stage retrieval-augmented generation framework that uses a large language model to pinpoint and automatically fix Gas inefficiencies. Experiments on deployed contracts demonstrate that RAGas reduces Gas usage by up to 11% and achieves high precision and recall in detecting code snippets exhibiting Gas wastage.
摘要:Ethereum 現在對於關鍵任務領域至關重要,包括金融、醫療保健和供應鏈管理。執行費用,通常稱為 Gas,隨著其功能的計算複雜性而增加。以太坊上的智能合約會產生執行費用,稱為 Gas,這些費用隨著計算複雜性的增加而上升。因此,在保持功能等價的同時優化高 Gas 消耗的代碼,可以顯著降低部署成本。現有系統無法持續利用不斷演變的 Gas 使用模式。我們系統地分析驅動過度 Gas 使用的語法和語義結構。這產生了六個高層次類別,涵蓋了十二個細緻的反模式,支撐著一個策劃的知識庫。我們利用這些洞見實現 RAGas,一個三階段的檢索增強生成框架,使用大型語言模型來精確定位並自動修復 Gas 效率低下的問題。對已部署合約的實驗表明,RAGas 將 Gas 使用量降低了多達 11%,並在檢測顯示 Gas 浪費的代碼片段時達到了高精度和高召回率。
Characterising cardiac tissue properties with graph neural networks
2608.15843v1 by Ching-En Chiu, Yoo Ri Kim, Magdi Saba, Danilo Mandic, Marta Varela
Characterising electrophysiological properties of cardiac tissue efficiently and accurately from spatially sparse intracardiac measurements is clinically important for localising ablation targets and improving arrhythmia treatment. We developed a graph neural network-based framework trained on synthetic electrogram signals on 2D flat surfaces to identify areas of interest in the context of cardiac ablation for premature ventricular complexes (PVCs). Our method achieved an average precision of 0.96, 0.97, and 0.95 for the detection of single-patch fibrosis, rapid depolarisation and high excitability, respectively. The trained model can then be applied to 2D curved surfaces with few-shot fine-tuning, demonstrating its generalisation capability. Future work will develop this framework further for clinical use in PVC ablation.
摘要:有效且準確地從空間稀疏的心內測量中描述心臟組織的電生理特性,對於定位消融目標和改善心律不整治療具有臨床重要性。
我們開發了一個基於圖神經網絡的框架,該框架在2D平面上對合成電圖信號進行訓練,以識別在心臟消融中與早期心室複雜(PVCs)相關的興趣區域。
我們的方法在檢測單一斑塊纖維化、快速去極化和高興奮性方面,分別達到了0.96、0.97和0.95的平均精度。
訓練好的模型可以通過少量調整應用於2D曲面,顯示出其泛化能力。
未來的工作將進一步開發這一框架,以便在PVC消融中用於臨床應用。
Schema-Agnostic Graph Reasoning Agent for Hybrid Knowledge Graphs
2608.15834v1 by Marius Dragic, Ruben Ifrah, Alexandre Rio
Tool-calling LLM agents navigate unfamiliar codebases with a handful of generic primitives for listing, reading and searching files (ls, cat, grep). A knowledge graph admits the same interface: listing neighbours, reading node content and searching descriptions are the same operations on a different substrate. Building on this correspondence, we present GRA, a Graph Reasoning Agent that explores hybrid knowledge graphs, whose nodes are either textual concepts or relational tables, with seven generic tools, discovering everything domain-specific at run time. On UFK-M (Unified Factory Knowledge Model), an industrial benchmark of 258 analytical questions whose gold answers are produced by executing validated SQL programs, GRA beats a full-context agent by 5.1 pp (88.4% vs. 83.3%), while reading under a third of its input tokens. A graph-free control shows the gain comes chiefly from selective agentic access rather than graph topology, and that the effect depends on a model able to drive tools reliably. Seeing less, the agent answers better: selective navigation over a structured substrate beats exhaustive context.
摘要:工具調用的 LLM 代理使用一小部分通用原語來導航不熟悉的代碼庫,以列出、閱讀和搜索文件(ls、cat、grep)。知識圖譜承認相同的接口:列出鄰居、閱讀節點內容和搜索描述在不同的基質上是相同的操作。基於這一對應關係,我們提出了 GRA,一個探索混合知識圖譜的圖推理代理,其節點可以是文本概念或關聯表,並使用七個通用工具,在運行時發現所有特定於領域的內容。在 UFK-M(統一工廠知識模型)上,這是一個包含 258 個分析問題的工業基準,其金標答案是通過執行經過驗證的 SQL 程序生成的,GRA 以 5.1 個百分點的優勢擊敗了全上下文代理(88.4% 對 83.3%),同時閱讀的輸入標記不到其三分之一。一個無圖控制顯示,這一增益主要來自於選擇性代理訪問,而非圖拓撲,並且這一效果依賴於能夠可靠驅動工具的模型。看到的越少,代理回答得越好:在結構化基質上進行選擇性導航優於全面上下文。
The Authority Resolution Framework: A Five-Domain Ontology for Governing Who and What Decides, at Scale
2608.15832v1 by Parviz Shariff
As AI systems become increasingly capable of autonomous action, determining whether an agent is technically capable of performing an action is insufficient: the system must also determine whether the action is authorised in its context. This paper introduces the Authority Resolution Framework (ARF), a five-domain ontology for representing and resolving authority across organisational roles and informal influence, business concepts, codified processes, machine-readable permissions and executable systems, and external real-world context. ARF defines the Authority Relation (AR) as a cross-domain primitive binding an actor, action, object, bounded context, justification chain, and a calibration measure termed the DNA-Coefficient, which captures divergence between documented authority structures and authority as practiced. The framework provides a machine-interpretable representation of authority provenance and scope, with JSON-LD representations and knowledge-graph query patterns for authority resolution. ARF is designed to support AI agents in determining the provenance, scope and contextual validity of authority before executing consequential actions. The framework positions authority resolution as a knowledge-representation and reasoning problem at the intersection of ontology engineering, semantic AI, agentic AI and AI governance.
摘要:隨著人工智慧系統越來越能夠自主行動,僅僅判斷一個代理是否在技術上能夠執行某個行動是不夠的:系統還必須確定該行動在其上下文中是否被授權。
本文介紹了權限解析框架(Authority Resolution Framework, ARF),這是一個五個領域的本體,用於表示和解決組織角色和非正式影響、商業概念、編碼過程、機器可讀的許可和可執行系統,以及外部現實世界上下文中的權限。
ARF 將權限關係(Authority Relation, AR)定義為一個跨領域的原始綁定,將行為者、行動、對象、有限上下文、理由鏈以及一個稱為 DNA-係數的校準度量結合在一起,該度量捕捉了記錄的權限結構與實際執行的權限之間的差異。
該框架提供了一種機器可解釋的權限來源和範圍的表示,並提供 JSON-LD 表示和知識圖譜查詢模式以進行權限解析。
ARF 設計旨在支持 AI 代理在執行有後果的行動之前,確定權限的來源、範圍和上下文有效性。
該框架將權限解析定位為一個知識表示和推理問題,位於本體工程、語義 AI、代理 AI 和 AI 治理的交集處。
QuantumPhaseNet: A Gauge-Covariant Geometric and Quantum-Spectral Theory of Semantic Concept Hierarchies with Prototype Validation of a Classical Quantum-Inspired Model
2608.15820v1 by Kiyotaka Kasubuchi, Kazuo Fukiya
We present QuantumPhaseNet, a gauge-covariant geometric and quantum-spectral extension of Transformer representations. Context-dependent semantic states are modeled as complex amplitudes; a covariant phase rate induces a semantic wavelength used as a proxy for conceptual scale; and low-frequency graph modes define a document-level discourse direction. The theoretical part establishes local gauge invariance, unitarity of the quantum block, boundedness and conditional stability of WavePhase Attention, and a calibratable hallucination-risk formulation. We also implemented a fully offline Validation Studio for the classical quantum-inspired pipeline in Section 14.1 and evaluated the five research questions in Section 16.1 on its built-in synthetic setting (n=240, observation noise 0.22, circuit noise 0.08, five seeds). RQ1 yielded a wavelength-hierarchy Spearman correlation of 0.852 versus 0.707 for the baseline, 87.3% direction accuracy, and AUC 0.953. RQ2 achieved discourse alignment 0.933 versus 0.589 and 41.2 versus 16.2 paragraphs before drift. RQ3 achieved AUROC 0.881 versus cosine 0.765 and phase-shuffle 0.536. RQ4 achieved error-detection AUROC 0.854 versus entropy 0.634, with Brier 0.150 and ECE 0.098. RQ5 did not show quantum advantage: target probability and end-to-end cost efficiency were 25.5% and 0.107, compared with 70.7% and 0.707 for the Chebyshev classical approximation. These results provide initial synthetic evidence for the classical quantum-inspired components, but not external validity or unconditional quantum speedup.
摘要:我們提出了QuantumPhaseNet,這是一種與規範協變的幾何和量子頻譜擴展的Transformer表示。上下文依賴的語義狀態被建模為複數振幅;協變相位速率引入了一種語義波長,作為概念尺度的代理;而低頻圖模式定義了文檔層級的話語方向。理論部分建立了局部規範不變性、量子區塊的單位性、WavePhase注意力的有界性和條件穩定性,以及可校準的幻覺風險公式。我們還為第14.1節中的經典量子啟發管道實施了一個完全離線的驗證工作室,並在第16.1節中對其內建的合成設置(n=240,觀察噪聲0.22,電路噪聲0.08,五個種子)評估了五個研究問題。RQ1產生了波長層次的Spearman相關性0.852,而基準為0.707,方向準確率87.3%,AUC 0.953。RQ2實現了話語對齊0.933,而基準為0.589,並且在漂移之前有41.2與16.2段落。RQ3實現了AUROC 0.881,而餘弦為0.765,隨機相位為0.536。RQ4實現了錯誤檢測AUROC 0.854,而熵為0.634,Brier為0.150,ECE為0.098。RQ5未顯示量子優勢:目標概率和端到端成本效率分別為25.5%和0.107,而Chebyshev經典近似為70.7%和0.707。這些結果為經典量子啟發組件提供了初步的合成證據,但不具備外部有效性或無條件的量子加速。
ALKEMIE Agent: an autonomous platform for computational materials design
2608.15776v1 by Hongfu Huang, Yuzhe Li, Ao Xu, Bo Liu, Changrui Wang, Kan Tang, Ning Yang, Shengxian Liu, Hanyu Liu, Pengpeng Zhang, Linggang Zhu, Fengkai Liu, Yichen Lu, Tong Zhao, Naihua Miao, Jian Zhou, Zhimei Sun
Despite the powerful multi-scale modeling methods and high-throughput infrastructures established in the materials community, real material computation workflows remain fragmented and heavily manual, requiring researchers to constantly bridge software tools, data analysis, and intermediate decisions. This growing gap between methodological capability and practical execution highlights the need for a new kind of autonomous computational framework, one that can coordinate tools, knowledge, and workflows in a more unified and adaptive way. Here, we introduce ALKEMIE Agent, an agentic platform in which retrieval-augmented generation, a materials-computation knowledge base, registered skills, database-supported provenance, AI-assisted structure modeling, bounded task execution, tool-calling iteration, and error-diagnostic assistance are integrated within a traceable control loop. The capabilities of ALKEMIE Agent are demonstrated through applications including materials recommendation, structure modeling, phonon calculations, machine-learned interatomic potential training, LAMMPS simulations, Ab Initio Monte Carlo (AIMC) sampling, and active-learning-based materials screening. Finally, we outline the future directions and challenges for the development of agentic platforms for computational materials design.
摘要:儘管材料社群中已建立強大的多尺度建模方法和高通量基礎設施,但真正的材料計算工作流程仍然是碎片化且高度手動的,這要求研究人員不斷地橋接軟體工具、數據分析和中間決策。這種方法能力與實際執行之間日益擴大的鴻溝突顯了對一種新型自主計算框架的需求,這種框架能以更統一和適應的方式協調工具、知識和工作流程。在這裡,我們介紹了ALKEMIE Agent,一個代理平台,其中檢索增強生成、材料計算知識庫、註冊技能、數據庫支持的來源、AI輔助結構建模、有限任務執行、工具調用迭代和錯誤診斷協助被整合在一個可追蹤的控制迴路中。ALKEMIE Agent的能力通過包括材料推薦、結構建模、聲子計算、機器學習的原子間勢訓練、LAMMPS模擬、Ab Initio Monte Carlo (AIMC) 取樣和基於主動學習的材料篩選等應用得以展示。最後,我們概述了計算材料設計的代理平台未來的方向和挑戰。
Large Models for Small Devices: Recent Advances and Empirical Analysis of Edge AI Deployment
2608.15693v1 by Subhransu Das, Jiaming Cheng, Arnav Kumar, Sadia Afrose, Mingzhe Han, Michael Silagy, Shreya Palande, Brijesh Soni, Rajiv Ramnath
Running large AI models on resource-constrained edge devices requires model compression to reduce model size and computation. What compresses well, however, need not deploy well. We survey dozens of recent works that report compression results on real hardware and extract practical deployment guidelines from them. Following these guidelines, we deploy compact language and image models on GPU, CPU, and Raspberry Pi platforms across question answering and image segmentation. No single technique wins across tasks. For question answering, Qwen3.5 0.8B reaches 93.85 SQuAD F1 and 92 EM under Q5_K_M GGUF quantization, while structured pruning at the same precision costs 16 F1 at a 1% ratio. For segmentation, the ranking reverses: default quantization leaves parameters and MACs unchanged, whereas pruning cuts model size by nearly 80% at near-constant mIoU. Pruning can even inflate the deployed artifact by 21-49% by breaking k-quant super-block alignment; combined with longer, less format-compliant outputs, this raises Raspberry Pi latency up to 3.4x. Compression can also manufacture the appearance of competence rather than destroy it visibly: one LoRA-recovered variant stays fully parseable and holds 71% strict BoolQ accuracy while sending 97 of 100 predictions to a single class, at 52.6% balanced accuracy. We explain these effects through neural-flow graph analysis and prefill-decode-level latency decomposition, and condense them into task-specific deployment research directions. The right technique depends on the task, the model, and the hardware. Our experiment code and artifacts are open-sourced at https://github.com/Arnavvvkumar/deployment
摘要:在資源受限的邊緣設備上運行大型 AI 模型需要模型壓縮,以減少模型大小和計算量。 然而,壓縮效果良好的模型不一定能夠良好部署。 我們調查了數十篇最近的研究,這些研究報告了在實際硬體上進行的壓縮結果,並從中提取實用的部署指南。 根據這些指南,我們在 GPU、CPU 和 Raspberry Pi 平台上部署了緊湊的語言和圖像模型,應用於問題回答和圖像分割。 沒有單一技術在所有任務中都能獲勝。 在問題回答中,Qwen3.5 0.8B 在 Q5_K_M GGUF 量化下達到 93.85 的 SQuAD F1 和 92 的 EM,而在相同精度下的結構化剪枝則以 1% 的比例損失 16 的 F1。 對於分割,排名則顛倒:默認量化保持參數和 MACs 不變,而剪枝則在幾乎不變的 mIoU 下將模型大小減少近 80%。 剪枝甚至可以通過打破 k-quant 超塊對齊來使部署的工件膨脹 21-49%;結合更長且格式不合規的輸出,這使得 Raspberry Pi 的延遲增加至 3.4 倍。 壓縮還可以製造出能力的外觀,而不是明顯地摧毀它:一個 LoRA 恢復的變體保持完全可解析,並在將 100 次預測中的 97 次發送至單一類別的同時,保持 71% 的嚴格 BoolQ 準確率,平衡準確率為 52.6%。 我們通過神經流圖分析和預填充解碼級延遲分解解釋這些效果,並將其濃縮為特定任務的部署研究方向。 正確的技術取決於任務、模型和硬體。 我們的實驗代碼和工件已在 https://github.com/Arnavvvkumar/deployment 開源。
BERTopic-Virality Prioritisation: A Scalable Framework for Thematic and Comparative Analysis of COVID-19 and Monkeypox Misinformation on Twitter
2608.15691v1 by Mkululi Sikosana, Sean Maudsley-Barton, Oluwaseun Ajao
Health misinformation circulating during pandemics can gain traction rapidly, creating harmful narratives that compete with public health guidance. Most topic-modelling pipelines treat engagement as an external outcome, limiting their ability to prioritise semantically coherent topics that are also rapidly diffusing. We introduce BERTopic-VP, a virality-prioritised topic-modelling framework that combines contextual embedding-based clustering (BERTopic) with a post hoc Virality Prioritisation (VP) layer. The pipeline is complemented by a two-stage hybrid misinformation detection module that fuses a supervised content-based classifier with an external verification signal derived from public-health knowledge bases. Applied to three benchmark datasets, COVID-19_FNIR, Monkeypox, and Constraint, the framework achieves strong classification performance, with F1 up to 0.950 and ROC-AUC up to 0.989, while identifying high-impact clusters under top 1%, 5%, and 10% VP thresholds. For datasets without native engagement metadata, prioritisation is based on a logistic propensity-to-spread score, used as an ordinal proxy for diffusion potential rather than a direct measure of engagement. The results show that integrating semantic structure, virality-aware ranking, and affective-linguistic profiling enables scalable and interpretable comparative analysis of misinformation across pandemics. The proposed framework supports monitoring-oriented early warning by surfacing low-volume but high-risk narratives for analyst review.
摘要:健康錯誤資訊在疫情期間迅速傳播,形成與公共衛生指導相競爭的有害敘事。大多數主題建模管道將參與度視為外部結果,限制了它們優先考慮語義一致且快速擴散主題的能力。我們介紹了BERTopic-VP,一種優先考慮傳播性的主題建模框架,將基於上下文嵌入的聚類(BERTopic)與事後傳播優先化(VP)層結合起來。該管道還配備了一個兩階段的混合錯誤資訊檢測模塊,該模塊將監督式內容分類器與來自公共衛生知識庫的外部驗證信號融合在一起。應用於三個基準數據集,COVID-19_FNIR、猴痘和Constraint,該框架實現了強大的分類性能,F1高達0.950,ROC-AUC高達0.989,同時在前1%、5%和10%的VP閾值下識別出高影響力的聚類。對於沒有原生參與度元數據的數據集,優先化基於邏輯傳播潛力分數,該分數用作擴散潛力的序數代理,而不是參與度的直接衡量。結果顯示,整合語義結構、考慮傳播性的排名和情感語言特徵分析,使得跨疫情的錯誤資訊進行可擴展且可解釋的比較分析成為可能。所提出的框架支持以監測為導向的早期預警,通過顯示低量但高風險的敘事供分析師審查。
THESIS-MoE: Trainable Hierarchical Extraction and SteerIng of Sycophancy in Mixture-of-Experts
2608.15687v1 by Kareem Hassani, Chaymaa Abbas, Lama Mawlawi, Mariette Awad
Sycophancy, the tendency of a language model to change its answer to match a user's stated belief, is a common alignment failure. Existing activation steering methods typically apply a single contrastive direction uniformly throughout the model, which is an unconditional intervention that alters activations even when no sycophantic behavior is present, trading knowledge retention for behavioral correction. In Mixture-of-Experts (MoE) models, prior work further suggests that behavior is encoded within expert computations rather than routing decisions alone, making precise behavioral steering particularly challenging. In this work, we introduce a shared contrastive signal, built from matched prompts with and without a stated belief, that identifies where sycophancy lives across the MoE hierarchy and drives interventions that act only where the behavior is present. We formulate localization as a causal search over a granularity ladder of MoE blocks, experts, attention blocks, and heads, and compare unconditional subtraction against two conditional alternatives: an analytic projection-based subtraction and a learned per-token gate that steers the model away from sycophancy while keeping its weights frozen. We evaluate on three MoE models measuring sycophancy alongside general knowledge and reasoning benchmarks. Our conditional interventions removed up to 90\% of the belief-induced sycophancy. Our results demonstrate that sycophancy resides in identifiable computational subcircuits and can be selectively steered while maintaining a favorable removal-retention trade-off.
摘要:拍馬屁是語言模型改變其答案以符合用戶所表達的信念的傾向,這是一種常見的對齊失敗。現有的激活引導方法通常在整個模型中均勻地應用單一的對比方向,這是一種無條件的干預,即使在沒有拍馬屁行為的情況下也會改變激活,從而以知識保留換取行為修正。在專家混合模型(MoE)中,先前的研究進一步表明,行為是編碼在專家計算中,而不僅僅是路由決策,這使得精確的行為引導特別具有挑戰性。在本研究中,我們引入了一個共享的對比信號,該信號由帶有和不帶有明確信念的匹配提示構建,能夠識別拍馬屁在MoE層級中的存在位置,並驅動僅在行為存在的地方進行干預。我們將定位公式化為對MoE區塊、專家、注意力區塊和頭部的粒度梯度進行因果搜索,並將無條件的減法與兩種條件替代方案進行比較:基於解析投影的減法和一個學習的每個標記門控,該門控在保持權重不變的情況下使模型遠離拍馬屁。我們在三個MoE模型上進行評估,測量拍馬屁以及一般知識和推理基準。我們的條件干預消除了高達90\%的信念引起的拍馬屁。我們的結果表明,拍馬屁存在於可識別的計算子電路中,並且可以在保持有利的去除-保留權衡的同時進行選擇性引導。
Agent Gym: A Framework for Continuous Evaluation and Evolution of LLM Agents Through Human-in-the-Loop Feedback
2608.15591v1 by Pouya Ghiasnezhad Omran, Michael Zimmermann, Duncan Cambridge, Ashmita Kapoor, Tanya Dixit
Large Language Model (LLM) agents deployed in production environments face a fundamental tension: the agent's behavior is frozen at deployment time, while the business rules and edge cases it must handle continue to evolve. Existing approaches address agent construction and one-time evaluation but provide no structured mechanism for continuous post-deployment behavioral correction without modifying the agent's source code. Most of the approaches offered in the market, require intense collection of logs and traces, and re-examining the agent design by the engineering team, a process which is heavy, long and negates the economical value of agentic transformation. We introduce Agent Gym, a modular, domain-agnostic framework that wraps any existing LLM-based agent in a continuous evaluation-and-evolution loop. The framework provides six composable capabilities --- Act, Evaluate, Investigate, Correct, Learn, and Observe --- organized across three architectural zones: a constitution layer that codifies domain knowledge in configuration artifacts, a runtime inference pipeline that chains acting, investigation, and adaptive correction, and a learning loop that enables subject matter experts to discover and validate new correction rules through natural language interaction. The key technical contributions include a hybrid deterministic-LLM correction engine with 21 condition operators and three-tier actions, a three-layer investigation architecture for ground-truth-free compliance validation, and a programmatic safety loop that guarantees rule correctness before human approval. We further introduce the Spec-to-Note Gap, an autoencoder-inspired view of agentic system transparency. An open-source reference implementation for invoice processing demonstrates that the framework is fully operational and ready for adoption.
摘要:大型語言模型(LLM)代理在生產環境中面臨著根本性的緊張關係:代理的行為在部署時被凍結,而必須處理的商業規則和邊緣案例則不斷演變。現有的方法解決了代理的構建和一次性評估,但未提供任何結構化機制以在不修改代理源代碼的情況下進行持續的部署後行為修正。市場上大多數提供的方法需要大量的日誌和追蹤數據收集,並由工程團隊重新檢查代理設計,這是一個繁重、漫長的過程,並削弱了代理轉型的經濟價值。我們介紹了Agent Gym,一個模組化的、與領域無關的框架,將任何現有的基於LLM的代理包裹在持續評估和演變的循環中。該框架提供六種可組合的能力——行動、評估、調查、修正、學習和觀察——這些能力組織在三個架構區域中:一個憲法層,將領域知識編碼為配置工件;一個運行時推理管道,鏈接行動、調查和自適應修正;以及一個學習循環,使主題專家能夠通過自然語言互動發現和驗證新的修正規則。關鍵的技術貢獻包括一個混合確定性-LLM修正引擎,具有21個條件運算符和三層行動;一個三層調查架構,用於無基準真相的合規驗證;以及一個程式化的安全循環,確保在人工批准之前規則的正確性。我們進一步介紹了Spec-to-Note Gap,一種受自編碼器啟發的代理系統透明度視角。一個開源的發票處理參考實現展示了該框架的完全運行狀態,並準備好被採用。
GraniKV: Asymmetric Granularity KV-Cache Paging for Multi-Agent Systems with Long Shared Prefix
2608.15584v1 by Jinhyun Jeon, Sungjoo Yoo
Production paged-serving engines apply uniform paging granularity to the KV cache, even though the two regions of a multi-agent workload have opposite storage requirements: a long shared prefix demands contiguity, while the per-request suffix demands fine-grained allocation. We present \textbf{GraniKV}, a KV-cache layer that allocates the shared prefix in a contiguous HOT pool and the suffix in a token-level COLD pool, combined with a per-step dispatcher which selects the appropriate backend among dual backends for each regime (compute-, memory-, or communication-bound). To the best of our knowledge, GraniKV is the first system to apply asymmetric paging granularity to the KV cache of a production paged-serving engine. At $L_p{=}16$\,K shared tokens GraniKV reaches $\mathbf{2.16\times}$, $\mathbf{1.98\times}$, and $\mathbf{1.57\times}$ output-token throughput over the production baseline on Llama-3.1-8B/TP=1, Qwen-2.5-14B/TP=2, and Qwen-2.5-32B/TP=4. The gain decomposes: cascade attention integration contributes the majority at saturation; the asymmetric storage layer adds $1.05$--$1.15\times$ end-to-end while being what makes the batched-GEMM prefix backend possible at all. Under heterogeneous multi-agent serving with \emph{distinct} prompts of different lengths, the attribution inverts: GraniKV sustains $\mathbf{1.95\times}$ while batch-global cascade collapses to parity --- the storage layer alone carries the win in the regime that motivates the paper.
摘要:生產頁面服務引擎對KV快取應用統一的分頁粒度,儘管多代理工作負載的兩個區域具有相反的存儲需求:長共享前綴要求連續性,而每個請求的後綴則要求細粒度分配。
我們提出了\textbf{GraniKV},這是一個KV快取層,將共享前綴分配在連續的HOT池中,後綴則分配在令牌級的COLD池中,並結合了一個每步調度器,該調度器在每個模式(計算、內存或通信限制)中選擇適當的後端。
據我們所知,GraniKV是第一個將非對稱分頁粒度應用於生產頁面服務引擎的KV快取系統。
在$L_p{=}16$\,K共享令牌下,GraniKV在Llama-3.1-8B/TP=1、Qwen-2.5-14B/TP=2和Qwen-2.5-32B/TP=4的生產基準上達到了$\mathbf{2.16\times}$、$\mathbf{1.98\times}$和$\mathbf{1.57\times}$的輸出令牌吞吐量。
增益分解如下:在飽和時,級聯注意整合貢獻了大部分;非對稱存儲層在端到端上增加了$1.05$--$1.15\times$,同時使得批次GEMM前綴後端成為可能。在異質多代理服務中,具有\emph{不同}長度的不同提示,歸因則反轉:GraniKV保持$\mathbf{1.95\times}$,而批次全局級聯則崩潰至平衡——存儲層單獨在促使本文的模式中獲得了勝利。
From Generalist to Specialist: A Context-Fusion Framework for Endoscopic Polyp Reporting with a Frozen VLM
2608.15580v1 by Ruijie Yang, Yan Zhu, Peiyao Fu, Siyuan Li, Te Luo, Zhihua Wang, Quanlin Li, Pinghong Zhou, Xian Yang, Shuo Wang
Reliable endoscopic polyp reporting requires integrating quantitative lesion sizing, standardized Paris classification, and clinically meaningful morphological description within a single record. General-purpose vision-language models (VLMs) offer a unified interface for image understanding and report generation. Existing specialization strategies, however, typically rely on task-specific models or model-weight adaptation, leaving unresolved how to introduce reliable specialist knowledge while preserving both this unified interface and the VLM's pretrained capabilities. We introduce a context-fusion framework that specializes a frozen general-purpose VLM through both implicit instruction context and explicit transduction context without modifying its pretrained weights. Specifically, a self-supervised polyp encoder retrieves related image-report pairs as explicit, query-specific evidence, while learned continuous specialist tokens provide implicit instruction context shared across cases. Experiments were conducted on 2,056 expert-annotated public endoscopic images. We compared the framework with general-purpose VLMs, task-specific predictors, and weight-adaptation methods to assess specialist performance, unified reporting, and adaptation efficiency. Across numerical, categorical, and report-generation metrics, the proposed framework substantially improved direct frozen-VLM inference and achieved the strongest overall performance among the evaluated methods. It added trainable parameters equal to only 0.006% of the frozen VLM's parameter count. When the top-1 retrieved case carried the correct target category, our framework corrected 70.5% of the errors made by a weight-adaptation baseline. These findings support the context-fusion framework as a lightweight and effective strategy for specialist adaptation of a frozen VLM.
摘要:可靠的內視鏡息肉報告需要將定量病變大小、標準化的巴黎分類和臨床上有意義的形態描述整合在單一記錄中。通用視覺-語言模型(VLMs)提供了一個統一的圖像理解和報告生成界面。然而,現有的專業化策略通常依賴於特定任務的模型或模型權重調整,尚未解決如何在保留這一統一界面和VLM的預訓練能力的同時引入可靠的專家知識。我們提出了一個上下文融合框架,通過隱式指令上下文和顯式轉導上下文專門化一個凍結的通用VLM,而不修改其預訓練權重。具體而言,自監督的息肉編碼器檢索相關的圖像-報告對作為顯式的查詢特定證據,而學習的連續專家標記提供了在案例之間共享的隱式指令上下文。實驗在2,056張專家標註的公共內視鏡圖像上進行。我們將該框架與通用VLMs、特定任務的預測器和權重調整方法進行比較,以評估專家性能、統一報告和適應效率。在數值、類別和報告生成指標上,所提出的框架顯著改善了直接凍結VLM推理,並在評估的方法中實現了最強的整體性能。它增加的可訓練參數僅佔凍結VLM參數總數的0.006%。當檢索到的頂級案例攜帶正確的目標類別時,我們的框架修正了70.5%的權重調整基線所犯的錯誤。這些發現支持上下文融合框架作為一種輕量且有效的策略,用於凍結VLM的專家適應。
Admission Without Answers: Label-Free Certification and Experience Learning for LLM-Based Optimization Modeling
2608.15565v2 by Junbo Jacob Lian, Huiling Chen, Hanzhang Qin, Chung-Piaw Teo
Experience-learning agents for optimization modeling improve by storing verified skills, but existing learners admit knowledge by checking against known answers, which real ticket streams do not provide. The natural label-free alternatives are unreliable: on a 300-problem label-blind stream, admitting every executable model poisons roughly one admission in four, while single-instance agreement accepts models that match at one value but differ elsewhere. We propose AdmitOR, an admission gate built on calibrated external behavioral evidence. Candidates from three model families, prompting strategies, and solver stacks are run on instances resampled from an extracted parameter domain; agreement across the resulting value-function traces is summarized by a cross-family clique, and a calibrated threshold returns accept, abstain, or escalate. The preregistered false-discovery criterion holds on calibration data but not on the wild stream. We report this negative result in full and trace most failures to benchmark texts that do not faithfully encode their labeled instances. Comparing four admission judges on one collection of logs inside a state-of-the-art skill learner, AdmitOR raises admission precision to 0.927, against 0.871 for majority vote and 0.726 for execution success, yielding 3.1x and 8.0x fewer poisoned admissions. Its library is the smallest and attains the highest macro accuracy across five public benchmarks, 58.4 against 54.8 for majority vote and 53.9 for the ground-truth-labeled library. The 3.5-point gain over majority vote is supported by a paired bootstrap and survives correction for a host-side anomaly. To our knowledge, AdmitOR is the first label-free admission mechanism designed around an explicitly calibrated false-discovery target. The transfer failure identifies a necessary condition for extending it to wild streams.
摘要:經驗學習代理在優化建模中通過儲存經過驗證的技能來提高性能,但現有的學習者通過檢查已知答案來承認知識,而這些答案在實際的票務流中並不存在。自然的無標籤替代方案不可靠:在一個300題的無標籤流中,承認每個可執行模型大約會使每四個承認中就有一個受到污染,而單實例一致性則接受在一個值上匹配但在其他地方不同的模型。我們提出了AdmitOR,一個基於經過校準的外部行為證據的承認閘。來自三個模型家族、提示策略和求解器堆棧的候選者在從提取的參數域重新抽樣的實例上運行;對於生成的值函數軌跡的一致性通過跨家族的團體進行總結,並且一個經過校準的閾值返回接受、放棄或升級。預註冊的假發現標準在校準數據上成立,但在野外流中不成立。我們全面報告這一負面結果,並將大多數失敗追溯到未忠實編碼其標記實例的基準文本。在一個最先進的技能學習者中的一組日誌上比較四個承認評審,AdmitOR將承認精度提高到0.927,而多數投票為0.871,執行成功為0.726,分別減少了3.1倍和8.0倍的污染承認。它的庫是最小的,並在五個公共基準中達到了最高的宏觀準確率,58.4對比多數投票的54.8和真實標記庫的53.9。相較於多數投票的3.5點增益得到了配對自助法的支持,並且在主機端異常的修正下仍然成立。據我們所知,AdmitOR是第一個圍繞明確校準的假發現目標設計的無標籤承認機制。轉移失敗確定了將其擴展到野外流的必要條件。
BengaliMCQ: Automatic Generation and Answer Prediction of Academic Multiple-Choice Questions in a Low-Resource Language
2608.15547v1 by Abu Tarabin Surzo, A. K. M. Nihalul Kabir, Sm Azmain Faysal, Ariana Haque Ami, Lawrence Amlan Gomes, Farig Sadeque
Traditional retrieval-augmented generation (RAG) frameworks process documents without attending to their hierarchical structure, leading to poor performance, especially in low-resource languages such as Bengali. To address this, we propose a structure-aware RAG framework that models Bengali textbooks as hierarchical graphs and uses a contrastively trained graph neural network to retrieve a small set of relevant passages. These passages provide focused context for a large language model, enabling topic-specific multiple-choice question (MCQ) generation and in-domain answer prediction. Experimental results demonstrate that our framework outperforms strong dense retrieval baselines across retrieval metrics, produces more relevant MCQs, and achieves superior answer prediction accuracy.
摘要:傳統的檢索增強生成(RAG)框架在處理文件時未考慮其層次結構,導致性能不佳,特別是在資源匱乏的語言如孟加拉語中。為了解決這個問題,我們提出了一種結構感知的 RAG 框架,將孟加拉語教科書建模為層次圖,並使用對比訓練的圖神經網絡來檢索一小組相關段落。這些段落為大型語言模型提供了集中上下文,使得能夠生成主題特定的多選題(MCQ)和在域內的答案預測。實驗結果顯示,我們的框架在檢索指標上超越了強大的密集檢索基準,產生了更相關的 MCQ,並實現了更高的答案預測準確性。
L3Cube-IndicQuest v2: A Large-Scale Multilingual Benchmark for Evaluating Factual Knowledge of Large Language Models Across Indic Languages
2608.15535v1 by Rinit Jain, Tirthraj Mahajan, Advait Joshi, Raviraj Joshi
We present L3Cube-IndicQuest v2, a large-scale gold-standard multilingual question-answering benchmark for evaluating the India-specific factual knowledge of Large Language Models (LLMs). The benchmark comprises 3,471 curriculum-grounded English question--answer pairs spanning nine domains, curated from educational curricula, competitive examination materials, and domain-specific reference books. We introduce a practical hybrid construction strategy that combines context-grounded LLM-based question generation and validation with semantic deduplication and human verification, enabling scalable creation of benchmark data while preserving annotation quality. The benchmark is translated into 19 Indic languages, yielding a publicly released multilingual dataset of 69,420 question--answer pairs across 20 languages. We evaluate six LLMs under three protocols: LLM-as-a-judge and two deterministic lexical criteria, exact-substring and word-overlap matching. All three produce almost the same model ranking, showing that the results do not depend on the choice of judge. The frontier commercial model leads by a wide margin, and among open-weight models Gemma4 31B outperforms the Indic-specialised Sarvam 30B in every evaluated Indic language.
摘要:我們推出 L3Cube-IndicQuest v2,這是一個大型的金標準多語言問答基準,用於評估大型語言模型(LLMs)在印度特定事實知識方面的表現。該基準包含 3,471 個基於課程的英語問答對,涵蓋九個領域,這些內容來自教育課程、競爭性考試材料和特定領域的參考書籍。我們引入了一種實用的混合建構策略,結合了基於上下文的 LLM 問題生成和驗證,以及語義去重和人工驗證,使得基準數據的可擴展創建成為可能,同時保持標註質量。該基準已翻譯成 19 種印度語言,產生了一個公開釋出的多語言數據集,包含 69,420 個問答對,涵蓋 20 種語言。我們在三個協議下評估了六個 LLM:LLM 作為評審以及兩個確定性詞彙標準,精確子字符串和詞重疊匹配。所有三種方法產生的模型排名幾乎相同,顯示結果不依賴於評審的選擇。最前沿的商業模型以較大優勢領先,而在開放權重模型中,Gemma4 31B 在每種評估的印度語言中均優於專注於印度的 Sarvam 30B。
Mental Model Management: An Operator-Based Framework for LLM Memory
2608.15451v1 by Oliver Kramer
Large language models process large amounts of information but usually lack an explicit mechanism for maintaining compact and evolving conceptual representations. We introduce Mental Model Management (3M), a framework in which knowledge is represented as mental models consisting of compact chunks. Rather than accumulating text passages, 3M continuously integrates new information into an existing conceptual representation. A set of operators extracts knowledge, retrieves relevant models, adds and updates chunks, reorganizes representations, detects inconsistencies, and derives new knowledge. We describe the main 3M operators and illustrate each operation using Evolution Strategies as a running example.
摘要:大型語言模型處理大量資訊,但通常缺乏明確的機制來維持緊湊且不斷演變的概念表徵。
我們介紹了心理模型管理(3M),這是一個將知識表示為由緊湊區塊組成的心理模型的框架。
3M並不是累積文本段落,而是持續將新資訊整合到現有的概念表徵中。
一組運算子提取知識、檢索相關模型、添加和更新區塊、重組表徵、檢測不一致性並推導新知識。
我們描述了主要的3M運算子,並以進化策略作為持續示例來說明每個操作。
Implementation of a Metacognition Framework for Self-Awareness and Self-Regulation in Ensembles of LLMs
2608.15400v1 by Charles Courchaine, Ricky J. Sethi, Hefei Qiu
Large Language Models (LLMs) are notorious for struggling with assessing their own uncertainty, detecting knowledge conflicts, or recognizing when problems exceed their expertise; such limitations inevitably undermine reliability and trust in LLMs. In this paper, we present the first implementation of a metacognitive framework for ensembles of LLMs that addresses these challenges through explicit monitoring and control mechanisms. Our system computes a Metacognitive State Vector (MSV) quantifying self-awareness for monitoring across five dimensions derived from cognitive psychology: Emotional Response, Correctness Evaluation, Experiential Match, Conflicting Information, and Problem Importance. MSV values also provide self-regulation for control, automatically switching between System 1 (fast, single- or multi-node) and System 2 (deliberative, multi-node) processing based on query complexity. For System 2 execution, graph-theoretic algorithms control the assignment of specialized roles (Domain Expert, Critic, Evaluator, Synthesizer, and Generalist) to ensemble nodes according to their MSV-quantified metacognitive states. Our implementation allows users to explore how different query types trigger distinct processing modes. The Proof-of-Concept (PoC) demo showcases the framework with illustrative examples showing appropriate System 1/System 2 routing and helps visualize the metacognitive process via real-time radar charts and decision indicators. This PoC implementation demonstrates the feasibility of creating a framework for metacognitive self-awareness and self-regulation in LLM systems.
摘要:大型語言模型(LLMs)因難以評估自身的不確定性、檢測知識衝突或識別問題超出其專業範疇而聞名;這些限制不可避免地削弱了對LLMs的可靠性和信任。在本文中,我們展示了首個針對LLMs集成體的元認知框架實現,通過明確的監控和控制機制來解決這些挑戰。
我們的系統計算一個元認知狀態向量(MSV),量化自我意識,以便在五個來自認知心理學的維度上進行監控:情感反應、正確性評估、經驗匹配、衝突信息和問題重要性。MSV值還提供自我調節以進行控制,根據查詢的複雜性自動在系統1(快速、單節點或多節點)和系統2(深思熟慮、多節點)處理之間切換。
在系統2執行中,圖論算法根據其MSV量化的元認知狀態控制專業角色(領域專家、批評者、評估者、綜合者和通才)在集成節點上的分配。
我們的實現允許用戶探索不同查詢類型如何觸發不同的處理模式。概念驗證(PoC)演示展示了該框架,並通過示例顯示適當的系統1/系統2路由,幫助通過實時雷達圖和決策指標可視化元認知過程。這個PoC實現展示了在LLM系統中創建元認知自我意識和自我調節框架的可行性。
Grounding Healthcare LLMs in a Causal Knowledge Graph: Framework, Metrics, and a Cardiovascular Pilot
2608.15382v1 by Ummara Mumtaz, Aimen Noor, Awais Ahmed
Large language models (LLMs) are increasingly proposed for healthcare decision support, but their evaluations still reward single-answer accuracy rather than reasoning about interventions, mechanisms, harms, evidence, and uncertainty. We propose a reproducible, graph-centered evaluation framework for intervention-oriented LLM behavior in healthcare and stress-test it in a cardiovascular pilot. The framework has four components: (i) a domain causal knowledge graph in which assertions are first-class, provenance-preserving nodes with stable identifiers; (ii) a scenario-conditioned subgraph extraction step that, given any clinical scenario, retrieves the relevant reified-assertion subgraph; (iii) four controlled grounding conditions that vary how the retrieved subgraph is composed into the model's context (ungrounded C1, knowledge-graph C2, causal-graph C3, integrated C4); and (iv) an automated scoring pipeline, anchored on assertion identifiers, that computes intervention accuracy, and other evaluation measures on a single pass. To test the framework, we built a category-balanced scenario generator across eight reasoning failure modes and instantiated it on a cardiovascular graph. The metric panel discriminates conditions along interpretable, non-redundant axes: C4 obtains the strongest causal edge F1 (0.838), adverse-effect F1 (0.833), evidence accuracy (0.738), and unsupported claim rate (0.114), while C1 obtains the highest raw intervention accuracy (0.948) with no measurable causal or evidential grounding.
摘要:大型語言模型(LLMs)越來越多地被提議用於醫療決策支持,但其評估仍然獎勵單一答案的準確性,而不是對干預、機制、危害、證據和不確定性進行推理。我們提出了一個可重複的、以圖為中心的評估框架,用於醫療保健中的干預導向LLM行為,並在心血管試點中進行壓力測試。該框架有四個組成部分:(i)一個領域因果知識圖,其中斷言是第一類的、保持來源的節點,具有穩定的標識符;(ii)一個情境條件的子圖提取步驟,根據任何臨床情境檢索相關的具體化斷言子圖;(iii)四個控制的基礎條件,變化檢索到的子圖如何組成模型的上下文(未基礎的C1、知識圖C2、因果圖C3、整合的C4);以及(iv)一個自動評分管道,以斷言標識符為基礎,計算干預準確性和其他評估指標,僅需一次通過。為了測試該框架,我們建立了一個跨越八種推理失敗模式的類別平衡情境生成器,並在心血管圖上實現了它。該指標面板沿著可解釋的、非冗餘的軸區分條件:C4獲得最強的因果邊緣F1(0.838)、不良影響F1(0.833)、證據準確性(0.738)和不支持的主張率(0.114),而C1獲得最高的原始干預準確性(0.948),卻沒有可測量的因果或證據基礎。
Medical
| Publish Date | Title | Authors | Homepage | Code |
|---|---|---|---|---|
| 2026-08-18 | MotoSafety: Edge-AI with Learned Temporal Importance for Two-Wheeler Collision Risk Assessment Under Time Pressure | Sumit S. Shevtekar et.al. | 2608.17823v1 | null |
| 2026-08-18 | LLMs for Medical Consultation Are Evaluated Too Late: The Preformulation Gap | Yining Hua et.al. | 2608.17330v1 | null |
| 2026-08-18 | Delta2Gamma: Band-Wise Adaptive Contrastive Learning of EEG for Alzheimer's Disease Detection | Chanwoo Park et.al. | 2608.17231v1 | null |
| 2026-08-17 | A decodability criterion predicts when hidden-state selection beats majority voting in large language models | Zhixiang wang et.al. | 2608.17124v1 | null |
| 2026-08-17 | Foundation Agents Meet Agentic Deep Research: Evidence-Grounded Clinical Code Forecasting | Junda Wang et.al. | 2608.17075v1 | null |
| 2026-08-17 | Institution-Specific LLM Prompting Recovers PHI That De-identification Systems and Their Gold Standards Both Miss | Daniel Palacios et.al. | 2608.17051v1 | null |
| 2026-08-17 | Policy Iteration with Human Feedback: Bringing Post-Training RL to In-context Learning | Minh-Ha Nguyen et.al. | 2608.16831v1 | null |
| 2026-08-17 | Historical Backtesting for Scientific Question Discovery: A Protocol and Astronomy Pilot | Hui Mao et.al. | 2608.16795v1 | null |
| 2026-08-17 | Unsupervised Anomaly Detection for Image Dataset Quality Assurance in Multi-Center Breast MRI | Chiara Tappermann et.al. | 2608.16725v1 | null |
| 2026-08-17 | Toward Better Assessment of LLMs' Performance in Clinical Error Detection | Yifan Zhang et.al. | 2608.16643v1 | null |
| 2026-08-17 | Degradation-Aligned Self-Supervised Learning for State of Health Estimation of Lithium-Ion Batteries under Label Sparsity | Jiaqi Yao et.al. | 2608.16612v1 | null |
| 2026-08-17 | CACSurv: Concordance-Aligned Comparative Learning with Large Language Models for Cancer Survival Prediction | Tianqi Xiang et.al. | 2608.16594v1 | null |
| 2026-08-17 | Large language models as synthetic clinical experts to inform longitudinal rare-disease modeling | Clemens Schächter et.al. | 2608.16507v1 | null |
| 2026-08-17 | Offline Reinforcement Learning for Hemodynamic Management of Sepsis in the ICU: a MIMIC-IV Study with Dual Off-Policy Evaluation | Marc Pérez-Roig et.al. | 2608.16482v1 | null |
| 2026-08-17 | Adaptive Post-Processing Drives Instance-Level Detection in Stroke Lesion Segmentation | Qinghui Liu et.al. | 2608.16377v1 | null |
| 2026-08-17 | Foresight-England: Development of a National-Scale Generative AI Model of Electronic Health Records for Medical Event Prediction across the COVID-19 Pandemic | Simon Ellershaw et.al. | 2608.16273v1 | null |
| 2026-08-17 | A cross-modal generative model for incomplete and degraded prostate MRI with multicentre clinical validation | Siyuan Ma et.al. | 2608.16233v1 | null |
| 2026-08-17 | BaT: Towards Self-Evolving Medical Research Agent with Stage Rubrics | Junqi Liu et.al. | 2608.16211v1 | null |
| 2026-08-17 | Picking the Right Image to Classify: Reliable-Input Selection in Teledermatology | Fabian Gröger et.al. | 2608.16198v1 | null |
| 2026-08-17 | TokenSTFormer: A Tokenized Spatial-temporal Attention Model for Holistic Motion Analysis in Adolescent Idiopathic Scoliosis Screening | Dong Chen et.al. | 2608.16122v1 | null |
| 2026-08-17 | Decoupling Parcellation from Classification: Systematic Benchmark of Fast Brain Segmentation Methods for Alzheimer's Disease Detection | Jiadao Zou et.al. | 2608.16039v1 | null |
| 2026-08-16 | Breaking and Defending LLM-Powered Social Media Bot Detection Systems | Nof Orenstein et.al. | 2608.15893v1 | null |
| 2026-08-16 | Characterising cardiac tissue properties with graph neural networks | Ching-En Chiu et.al. | 2608.15843v1 | null |
| 2026-08-16 | PLeDO: Pain Level Detection for Osteoarthritis from EMR Data | Yuhao Chen et.al. | 2608.15719v1 | null |
| 2026-08-16 | Integrating Persuasion Theory into the Epidemiological Modelling of Health Misinformation Spread on Social Media | Mkululi Sikosana et.al. | 2608.15689v1 | null |
| 2026-08-16 | From Generalist to Specialist: A Context-Fusion Framework for Endoscopic Polyp Reporting with a Frozen VLM | Ruijie Yang et.al. | 2608.15580v1 | null |
| 2026-08-16 | EA-LiteUNet: An Edge-Adaptive and Resource-Efficient U-Net for Boundary-Sensitive Dermoscopic Image Segmentation | Wang Jiangtao et.al. | 2608.15537v1 | null |
| 2026-08-15 | Gated Against One Model, Open to the Next: Option-Only Solvability in Legal Multiple-Choice Benchmarks | Volodymyr Ovcharov et.al. | 2608.15428v1 | null |
| 2026-08-15 | ETHOS: Towards a Modular Ethics Framework for Clinical Multi-Agent Systems | Rakesh Sharma et.al. | 2608.15424v1 | null |
| 2026-08-15 | Invariant Pretraining for Robust Code Representations | Yifeng He et.al. | 2608.15412v1 | null |
| 2026-08-15 | Grounding Healthcare LLMs in a Causal Knowledge Graph: Framework, Metrics, and a Cardiovascular Pilot | Ummara Mumtaz et.al. | 2608.15382v1 | null |
| 2026-08-15 | When AI Rewrites, Classifiers Relax: Uncertainty-Aware Sentiment Analysis on Sarcastic and AI-Paraphrased Social Text | Shresth Shroff et.al. | 2608.15338v1 | null |
| 2026-08-15 | Physiological World Models for Human State Transitions | Chongyang Zhang et.al. | 2608.15309v1 | null |
| 2026-08-15 | Demographic Injection in Medical Language Models under Diversity, Equity, and Inclusion Prompts | Diego Mardian et.al. | 2608.15254v1 | null |
| 2026-08-15 | Translating finite-domain integer constraint models to CP/SMT/ILP/PB/SAT solvers with CPMpy | Tias Guns et.al. | 2608.15143v1 | null |
| 2026-08-15 | FZ-VLM: A Two Stage Florence-Zephyr Vision Language Model Framework for Pulmonary Nodule Characterization and Clinical Decision Making | Pramit Dutta et.al. | 2608.15004v1 | null |
| 2026-08-14 | Evaluating Agentic Code Repair Capabilities in Distributed Systems | Yibo Yan et.al. | 2608.14863v1 | null |
| 2026-08-14 | Generated Context versus Governed State: Functional Conditions for Accountable Longitudinal Clinical Reasoning | Augusto Bernardo Pissarra et.al. | 2608.14804v1 | null |
| 2026-08-14 | Prompting is not enough: supervised baselines and leakage control for measuring shared decision-making with LLMs in pediatric encounters | Bernardo Modenesi et.al. | 2608.14792v1 | null |
| 2026-08-14 | CEDAR-GRPO: Process-Aware Reinforcement Learning for General Abductive Reasoning in LLMs | Moein Salimi et.al. | 2608.14791v1 | null |
| 2026-08-14 | Seeing Red, Thinking Bad: Color Bias in Vision Language Models | Kohsuke Ide et.al. | 2608.14286v1 | null |
| 2026-08-14 | Meteorology-driven Causal Nowcasting of Fugitive Landfill Emissions Enables Proactive Public Health Response | Timothy C. Pearce et.al. | 2608.14254v1 | null |
| 2026-08-14 | Agentic Data Cleaning Without a Clean Reference: An Experimental Study of Capabilities and Trade-offs | Hadi Fadlallah et.al. | 2608.14765v1 | null |
| 2026-08-14 | APTER: Adaptive Post-Training with Expert-Grounded Rubrics | Xukai Wang et.al. | 2608.14212v1 | null |
| 2026-08-14 | Removing Temporal Note Redundancy Improves Multimodal Reinforcement Learning for Medicine | Chenran Weng et.al. | 2608.14157v1 | null |
| 2026-08-14 | CMCNet: Aligning Ultrasound Image Embeddings with Textual TI-RADS Representations for Fine-Grained Thyroid Classification | Bingxin Yu et.al. | 2608.13939v1 | null |
| 2026-08-13 | Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions | Qingfang Liu et.al. | 2608.13786v1 | null |
| 2026-08-13 | Data-driven techniques for translational neuroscience and personalized neuro-health | Vishal Subedi et.al. | 2608.13749v1 | null |
| 2026-08-13 | MedPlex: Deep Vision-Language Co-Adaptation for Clinically Grounded Medical Segmentation | Rafi Ibn Sultan et.al. | 2608.13690v1 | null |
| 2026-08-13 | MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and Coordination | Saisha Shetty et.al. | 2608.13476v1 | null |
| 2026-08-13 | Towards Context-Aware Clinical Motion Understanding in Daily Living at Home: Freezing of Gait Detection with Egocentric Vision | Vayalet Stefanova et.al. | 2608.13283v1 | null |
| 2026-08-13 | Static analysis-guided agentic AI translation enables Rust as a full stack bioinformatics language | Johan Henriksson et.al. | 2608.13029v1 | null |
| 2026-08-13 | Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence | Jakub Pokrywka et.al. | 2608.12928v1 | null |
| 2026-08-13 | CRAFT: LLM-Based Iterative Refinement for Temporal Reasoning over Clinical Narratives | Chengyang He et.al. | 2608.12779v1 | null |
| 2026-08-13 | Memorization Diagnostics for Code LLMs Should be Scale-Aware | Prateek Kumar Rajput et.al. | 2608.12771v1 | null |
| 2026-08-13 | PatientAct: Theory-Grounded Mental Health Client Simulation | Sahand Sabour et.al. | 2608.12750v1 | null |
| 2026-08-13 | Mr3D-VL: A generalist vision language foundation model for Multiparametric 3D Magnetic Resonance Imaging | Zhi Qiao et.al. | 2608.12689v1 | null |
| 2026-08-12 | SteerBench-Work: A Benchmark for Agent Steering at Action Boundaries | Oguz Serdar et.al. | 2608.12654v1 | null |
| 2026-08-12 | Algorithm Design and Physician Liability | Shujie Luan et.al. | 2608.13618v1 | null |
| 2026-08-12 | Auditable agentic AI for evidence-grounded thyroid ultrasound diagnosis and reporting | Haifan Gong et.al. | 2608.12590v1 | null |
| 2026-08-12 | How Compliant is Sepsis Treatment? An Expert-Guided Neuro-symbolic Pipeline for Generating Clinical Compliance Insights | Himanshu Tripathi et.al. | 2608.13617v1 | null |
| 2026-08-12 | M-Net: Integrating Spectral Features and Physical Field Operators into Deep Learning for Medical Image Segmentation | Jing Zhu et.al. | 2608.12196v1 | null |
| 2026-08-12 | A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench | Praveen Reddy et.al. | 2608.12138v1 | null |
| 2026-08-12 | Do LLMs Take Care of Their Own? Similarity Signals Can Induce Cooperation | Akash Kundu et.al. | 2608.12125v1 | null |
| 2026-08-12 | How Far from Clinical Deployment? Evaluating the Complete Unsupervised Domain Adaptation Pipeline in Medical Imaging | Yiheng Xiong et.al. | 2608.12035v1 | null |
| 2026-08-12 | From Safety Documentation to Safety Knowledge Support: An Evidence-Grounded LLM Framework for Medical Devices | Tuhinangshu Gangopadhyay et.al. | 2608.12025v1 | null |
| 2026-08-12 | When the API Speaks the Wrong Language: Revisiting Post-Training for Multilingual Tool Use | Siddharth Chauhan et.al. | 2608.11715v1 | null |
| 2026-08-12 | Rubric Dropout: A Simple Way to Mitigate Reward Hacking in Rubric-as-Reward RL | Minglai Yang et.al. | 2608.11669v1 | null |
| 2026-08-12 | Beyond Single-Turn Confidence: Trajectory-Adapted Uncertainty Quantification for LLM Agents | Dylan Bouchard et.al. | 2608.11552v1 | null |
| 2026-08-11 | Social Chain of Thought: A Multi-Agent Architecture Grounded in Medical Differential Diagnosis Methodology | Del Coburn et.al. | 2608.11420v1 | null |
| 2026-08-11 | Gaze Target Estimation Anywhere with Concepts | Xu Cao et.al. | 2608.11367v1 | null |
| 2026-08-11 | Dual-Domain Cross-Modal Decoding for Clinical Text-Guided Medical Image Segmentation | Md Maklachur Rahman et.al. | 2608.11335v1 | null |
| 2026-08-11 | 3D Weighted Geometric Graph Neural Networks for Sheep Facial Pain Assessment | Alam Noor et.al. | 2608.11050v1 | null |
| 2026-08-11 | CARE: Confidence-Aware Reasoning for Reliable Medical VQA | Yuetian Du et.al. | 2608.10964v1 | null |
| 2026-08-11 | ComBodied Agents: a New Paradigm of Human-Centric Agentic AI | Qianggang Ding et.al. | 2608.10915v2 | null |
| 2026-08-11 | MIRA: Medical Image Reflection for Agentic Diagnosis | Shengzhi Wang et.al. | 2608.10827v1 | null |
| 2026-08-11 | DuplexWorld: Can voice agents help you get through the day? | Aryan Vijay Bhosale et.al. | 2608.10716v1 | null |
| 2026-08-11 | MedUP: Awakening Unified Understanding and Perception in Medical Vision-Language Models | Yuan Wang et.al. | 2608.10635v1 | null |
| 2026-08-11 | Agentic Instruction Data Selection: Let DataMaster Interpret Your Intent | Fanqi Zhou et.al. | 2608.10579v1 | null |
| 2026-08-11 | Reinforcement Learning-Based Laser Cutting Machine Parameter Optimization | Khanh Quan Pham et.al. | 2608.10549v1 | null |
| 2026-08-11 | Unlocking the Power of Medical Tabular Data via Semantic-Aware Multimodal Pre-training | Yingsheng Liu et.al. | 2608.10522v1 | null |
| 2026-08-11 | RadFusion: Towards Threshold-Controllable Radiology Report Generation | Ying Jin et.al. | 2608.10505v1 | null |
| 2026-08-11 | RLMOpt: Adaptive Prompt Optimization via Recursive Language Models | Subhash Bangalore Satheesha et.al. | 2608.10471v1 | null |
| 2026-08-11 | Expert-Guided g-computation with Large Language Models for Estimating Causal Effects on Timings: Applications to Hospital Quality Improvement | Patrick Vossler et.al. | 2608.10339v1 | null |
| 2026-08-10 | Logit-Boundary Geometric Belief Interfaces and Sparse Sheaf-Enclave Protocols: A Self-Contained Substrate for Secure Network Electronic Health Record (EHR) Interoperability | Alvin Spivey et.al. | 2608.10300v2 | null |
| 2026-08-10 | Frozen Brain-MRI Foundation Models Are Site Fingerprints | Saman Rahbar et.al. | 2608.10295v1 | null |
| 2026-08-10 | Locally Deployable Small Language Models for Emergency Department Decision Support: A Systematic Benchmark of Fine-Tuning Strategies | Qingfeng Zhang et.al. | 2608.10273v1 | null |
| 2026-08-10 | TAF-MED: Multi-Turn Safety Refusal Collapse in LLMs Under Declared Self-Treatment Intent | Waleed Jamil et.al. | 2608.10258v1 | null |
| 2026-08-10 | Towards Expert-level Medical AI for Real-time Video Consultations | Mahvish Nagda et.al. | 2608.09861v1 | null |
| 2026-08-10 | MedPixel: A Unified Pixel-Language Model for Medical Reasoning and Segmentation | Haoyu Yang et.al. | 2608.09818v1 | null |
| 2026-08-10 | AirFlow: Context Preserving and Multi-Rate State Modeling for Air Quality Forecasting | Fan Yang et.al. | 2608.09775v1 | null |
| 2026-08-10 | Physics-Informed Machine Learning in Prognostics and Health Management: A Systematic Literature Review | Christopher Braun et.al. | 2608.10047v1 | null |
| 2026-08-10 | Coupled Graph--Policy Distillation for Personalized Medication Safety in Older Adults with Multimorbidity | Zihan Wang et.al. | 2608.09443v1 | null |
| 2026-08-10 | Multimodal Federated Learning under Dual-Axis Modality Missingness | Adiba Orzikulova et.al. | 2608.09240v1 | null |
| 2026-08-10 | Rethinking Medical Landmark Localization with Prototype Learning-based Progressive Offset Correction | Jingxian Xu et.al. | 2608.09182v2 | null |
| 2026-08-10 | RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning | Jinkun Hou et.al. | 2608.09123v1 | null |
| 2026-08-10 | A Multi-Scale Temporal Framework with Dynamic Fusion for EEG-Based Emotion Recognition | Stefanos Gkikas et.al. | 2608.09088v1 | null |
| 2026-08-10 | When Confidence Fails: Overconfidence in LLMs under Uncertainty and Missing Clinical Information | Maryam Tahermazandarani et.al. | 2608.09080v1 | null |
| 2026-08-09 | Decoding Phenotypes: A Framework for Fusing Genomic Language Models and Neuroimaging | Tianli Tao et.al. | 2608.08926v1 | null |
| 2026-08-09 | Toward CT-Equivalent Image Quality in Low-Dose Radiotherapy Planning: Conditional Diffusion-Based CBCT-to-CT Synthesis and the Impact of CBCT Input Representation | Alzahra Altalib et.al. | 2608.08919v1 | null |
Abstracts
MotoSafety: Edge-AI with Learned Temporal Importance for Two-Wheeler Collision Risk Assessment Under Time Pressure
2608.17823v1 by Sumit S. Shevtekar, Chandresh K. Maurya, Gourab Sil, Subasish Das
Powered two-wheeler riders face critical safety challenges in low- and middle-income countries, yet limited studies exist on how cognitive stressors such as Time Pressure influence collision risk. To address this gap, we introduce a large-scale dataset of over 129,000 labeled multivariate time-series sequences from 153 simulator rides by 51 participants under No, Low, and High TP, capturing 64 features across vehicle dynamics, control inputs, proximity, and behavioral violations. Building on this dataset, we propose MotoSafety, a novel edge-AI architecture grounded in the Learned Temporal Importance principle. MotoSafety achieves 94.97% accuracy and 99.33% ROC AUC, outperforming ten baselines, including TimesNet and LLM4TS, and achieves 0.039 MSE and 0.094 MAE for forecasting (4.4x lower error than Time-LLM and iTransformer). With only 1.15M parameters and 0.135 ms latency, it is suitable for edge deployment on low-cost CPU hardware. Using ground truth TP as an inductive bias improves accuracy from 94.09% to 94.97%, while predicted TP achieves 94.82%. Using only 21 IMU+GPS features, it achieves 93.91% accuracy, indicating practical deployment. Beyond PTW safety, the architecture shows better transferability to human activity (97.66%) and clinical (99.65%) domains. This lightweight framework advances PTW collision risk assessment, supporting the Safe System Approach for Intelligent Transportation Systems.
摘要:在中低收入國家,動力二輪車騎士面臨著重大的安全挑戰,但關於認知壓力因素如時間壓力如何影響碰撞風險的研究卻相對有限。為了填補這一空白,我們引入了一個大規模數據集,該數據集包含來自51名參與者在無時間壓力、低時間壓力和高時間壓力下進行的153次模擬騎行的超過129,000個標記的多變量時間序列,捕捉了64個特徵,涵蓋了車輛動態、控制輸入、接近度和行為違規。基於這個數據集,我們提出了MotoSafety,一種基於學習時間重要性原則的新型邊緣人工智慧架構。MotoSafety實現了94.97%的準確率和99.33%的ROC AUC,超越了包括TimesNet和LLM4TS在內的十個基準,並在預測中達到了0.039的均方誤差和0.094的平均絕對誤差(比Time-LLM和iTransformer低4.4倍)。它僅需1.15M的參數和0.135毫秒的延遲,適合在低成本CPU硬體上進行邊緣部署。使用真實的時間壓力作為歸納偏見,準確率從94.09%提高到94.97%,而預測的時間壓力則達到94.82%。僅使用21個IMU+GPS特徵,它的準確率達到93.91%,顯示出實際部署的潛力。除了PTW安全性外,該架構在人體活動(97.66%)和臨床(99.65%)領域也顯示出更好的可轉移性。這個輕量級框架推進了PTW碰撞風險評估,支持智能交通系統的安全系統方法。
LLMs for Medical Consultation Are Evaluated Too Late: The Preformulation Gap
2608.17330v1 by Yining Hua, Cyrus Ayubcha, Hongbin Na, Levi Lian, Alon Gorenshtein, Yiftach Barash, Eyal Klang
Large language models for medical consultation are often evaluated after a clinical problem has already been made clear, although real consultations may begin with a vague, minimized, or misframed concern. We evaluated three API models across four physician-authored, multi-turn vignettes under baseline and entry-to-care instruction conditions, yielding 24 fixed-script transcripts; two cases also used adaptive standardized-patient simulation, yielding 12 transcripts. Self-care or home-management advice before any patient answer appeared in 9 of 12 baseline case-model cells and 0 of 12 instruction cells, while structured handoff summaries appeared in 0 of 12 and 10 of 12 cells, respectively. The instruction changed sequencing and documentation, although it did not reliably ensure elicitation of decisive facts. The preformulation gap should therefore be evaluated directly through observable first-contact behavior rather than inferred from diagnostic accuracy or final-answer quality.
摘要:大型語言模型在醫療諮詢中的評估通常是在臨床問題已經明確之後進行的,儘管實際的諮詢可能是從模糊、最小化或錯誤框架的關切開始的。
我們在基線和進入護理指導條件下,評估了三個API模型在四個由醫生撰寫的多輪小品中的表現,共產生了24份固定腳本的逐字稿;另外兩個案例還使用了自適應標準化病人模擬,產生了12份逐字稿。
在12個基線案例模型單元中,有9個出現了自我照護或居家管理建議,而在12個指導單元中則沒有出現;結構化交接摘要在12個單元中分別出現了0個和10個。
指導改變了序列和文檔,儘管它並未可靠地確保引出關鍵事實。因此,前置公式化的差距應該通過可觀察的首次接觸行為直接評估,而不是從診斷準確性或最終答案質量中推斷。
Delta2Gamma: Band-Wise Adaptive Contrastive Learning of EEG for Alzheimer's Disease Detection
2608.17231v1 by Chanwoo Park, Chanwoo Kim
Low-cost, scalable screening for dementia remains an open problem. Imaging-based diagnosis is costly and hard to deploy widely. Electroencephalography (EEG) is portable and inexpensive, but its recordings are noisy, vary widely across subjects, and carry few clinical labels. We tackle this with Delta2Gamma, a self-supervised framework that learns EEG representations from unlabeled data by contrasting augmented views of each signal. Rather than treat EEG as a single stream, Delta2Gamma decomposes every recording into the five canonical neural rhythms (delta, theta, alpha, beta, gamma). Each band gets its own encoder and projection head. Each also gets a temperature that is predicted adaptively during contrastive training, so bands with different signal statistics are balanced automatically. On the ADFTD cohort under a strict leave-one-subject-out protocol, Delta2Gamma separates Alzheimer's disease from cognitively normal controls with 92.4\% accuracy. This exceeds both supervised backbones and recent dedicated EEG methods.
摘要:低成本、可擴展的癡呆篩檢仍然是一個未解決的問題。基於影像的診斷成本高且難以廣泛部署。腦電圖(EEG)便攜且便宜,但其錄音雜訊多、在受試者之間變化大,且臨床標籤少。我們通過Delta2Gamma來解決這個問題,這是一個自我監督框架,通過對比每個信號的增強視圖來學習無標籤數據的EEG表示。Delta2Gamma並不是將EEG視為單一流,而是將每個錄音分解為五種典型的神經節律(delta、theta、alpha、beta、gamma)。每個頻帶都有自己的編碼器和投影頭。每個頻帶還在對比訓練過程中自適應地預測一個溫度,因此具有不同信號統計的頻帶會自動平衡。在嚴格的留一受試者外協議下的ADFTD隊列中,Delta2Gamma以92.4\%的準確率將阿茲海默病與認知正常的對照組分開。這超過了監督式骨幹和最近專門的EEG方法。
A decodability criterion predicts when hidden-state selection beats majority voting in large language models
2608.17124v1 by Zhixiang wang, Ziliang Hong, Ulas Bagci
Combining the answers a large language model (LLM) samples for a question into one decision is a test-time information fusion problem, usually solved by majority voting. Voting is unreliable on difficult questions, where the sampled answers share correlated errors, so the wrong answer can win and drawing more samples makes the decision worse. Selecting a candidate by reading a correctness signal from the model's hidden states is a promising alternative, but its accuracy varies across models and tasks, and no measure indicates when it can be trusted. In this paper, we propose CASE (Correctness-Axis SElection), a dynamic selection combiner that trains a linear gate on the answer-token hidden state and selects the highest-scoring candidate. Its main contribution is decodability, a leakage-free measure of how well the gate ranks a question's correct candidates above its incorrect ones, which predicts whether hidden-state selection will outperform voting. A conventional probe appears accurate only because of question-identity leakage, which vanishes under question-grouped evaluation. On held-out data, decodability predicts the accuracy gain of selection over voting with a Pearson correlation r=0.75 and a decision threshold near AUC=0.60. Across general and medical LLMs, CASE improves over voting by up to 19 points on medium-difficulty questions and 16.8 points on hard questions. Decodability depends on the aligned knowledge a model must recall, not on its scale, and its prediction transfers to an unseen scientific domain within 3.8 points. It thus provides a practical criterion, measurable in advance for a given model and task, for choosing between learned selection and majority voting.
摘要:將大型語言模型(LLM)對一個問題所採樣的答案合併為一個決策是一個測試時的信息融合問題,通常通過多數投票來解決。
在困難問題上,投票不可靠,因為採樣的答案共享相關錯誤,因此錯誤的答案可能會獲勝,而增加更多樣本會使決策變得更糟。
通過從模型的隱藏狀態中讀取正確性信號來選擇候選者是一個有前途的替代方案,但其準確性在不同模型和任務之間有所變化,且沒有任何指標表明何時可以信任它。
在本文中,我們提出了CASE(正確性軸選擇),這是一個動態選擇組合器,對答案標記的隱藏狀態訓練一個線性閘,並選擇得分最高的候選者。
它的主要貢獻是可解碼性,這是一種無洩漏的度量,衡量閘如何將問題的正確候選者排名高於不正確的候選者,並預測隱藏狀態選擇是否會優於投票。
傳統探測器之所以顯得準確,僅僅是因為問題身份的洩漏,而這在問題分組評估中會消失。
在保留數據上,可解碼性預測選擇相對於投票的準確性增益,皮爾森相關係數 r=0.75,決策閾值接近 AUC=0.60。
在一般和醫療 LLM 中,CASE 在中等難度問題上提高了最多 19 分,在困難問題上提高了 16.8 分。
可解碼性取決於模型必須回憶的對齊知識,而不是其規模,且其預測在未見的科學領域內轉移至 3.8 分。
因此,它為在給定模型和任務之間選擇學習的選擇和多數投票提供了一個可實際測量的標準。
Foundation Agents Meet Agentic Deep Research: Evidence-Grounded Clinical Code Forecasting
2608.17075v1 by Junda Wang, Meysam Ghaffari, Akshat Choube, Mohsen Sharifi Renani, Hong Yu, Carlos Morato
Next-encounter ICD forecasting predicts which standardized diagnosis codes will be documented at a future visit from the longitudinal record available beforehand. The task is prospective and multi-label: the target note does not yet exist, and several codes may be correct. Structured EHR foundation models capture recurrence and temporal progression, whereas language foundation models generate flexible diagnostic hypotheses. We introduce ICD-Deepresearch, a DeepResearch workflow that composes these predictive foundation models with medical search and ICD dictionaries. Because no source reveals the future code set, research evaluates candidate transitions by linking patient evidence, external clinical relations, and exact code semantics under a fixed top-K budget. Candidate Generation uses SparseEHR to produce an EHR Prior that initializes two bounded Research Expansion rounds; an independent GPT-5 Direct Forecast supplies complementary candidates. Final Selection validates, deduplicates, and jointly ranks both paths, after which a separate module writes rationales without changing predictions. Finally ICD-Deepresearch achieves patient-averaged precision/recall of 24.60/35.09% on MIMIC-III and 25.14/48.32% on MIMIC-IV. Physicians rate 51% and 68% of its retrieved documents useful, compared with 22% and 39% for standalone GPT-5 web search and 32% and 41% for Medical Deep Research. ICD-Deepresearch therefore improves over the registered local comparators while retrieving evidence with higher physician-rated usefulness than the standalone research systems
摘要:下一次接觸的 ICD 預測預測未來訪問時將記錄哪些標準化診斷代碼,這些代碼來自之前可用的縱向記錄。這個任務是前瞻性的和多標籤的:目標筆記尚不存在,且可能有多個代碼是正確的。結構化的電子健康記錄基礎模型捕捉到復發和時間進展,而語言基礎模型則生成靈活的診斷假設。我們介紹 ICD-Deepresearch,這是一個 DeepResearch 工作流程,將這些預測基礎模型與醫療搜索和 ICD 字典結合起來。由於沒有來源揭示未來的代碼集,研究通過將患者證據、外部臨床關係和確切的代碼語義連接在一起來評估候選轉換,並在固定的 top-K 預算下進行。候選生成使用 SparseEHR 生成一個 EHR Prior,該 Prior 初始化兩輪有界的研究擴展;獨立的 GPT-5 直接預測提供補充候選。最終選擇驗證、去重並共同排名這兩條路徑,之後一個單獨的模塊在不改變預測的情況下寫出理由。最後,ICD-Deepresearch 在 MIMIC-III 上達到患者平均精確度/召回率 24.60/35.09%,在 MIMIC-IV 上達到 25.14/48.32%。醫生認為其檢索的文檔中有 51% 和 68% 是有用的,而獨立的 GPT-5 網頁搜索的有用率為 22% 和 39%,醫療深度研究的有用率為 32% 和 41%。因此,ICD-Deepresearch 在檢索證據時的醫生評價有用性上優於登記的本地比較者,並且比獨立研究系統更具優勢。
Institution-Specific LLM Prompting Recovers PHI That De-identification Systems and Their Gold Standards Both Miss
2608.17051v1 by Daniel Palacios, Matthew Brady Neeley, Angel Adetomike Otto, Shalini Dhamodharan, John P. Woodhouse, Chi-fan Lin, Mark Zobeck, Zhandong Liu, Hyun-Hwan Jeong
Secondary use of electronic health records requires de-identification, yet existing systems miss \emph{institutionally situated} protected health information (PHI) such as hospital abbreviations, building names, and internal codes whose status is locally determined. We ask whether large language models (LLMs) with in-context learning (ICL) can close this gap and control the precision--recall trade-off. On 100 annotated pediatric oncology notes (5,322 PHI spans) from Texas Children's Hospital, we benchmarked eight LLMs against two purpose-built systems (Stanford TiDE, OpenMed PII) and two pattern-based baselines. Each LLM ran under three prompts of increasing specificity: (1) a HIPAA-aligned baseline, (2) baseline plus the institutional PHI categories it missed, and (3) prompt 2 plus instructions against over-redacting clinical content. We then compared 14~multi-agent and ensemble configurations against the best single prompt, with recall the primary safety metric. LLMs outperformed the purpose-built systems (best F1=0.918$\pm$0.001 vs.\ TiDE 0.779), with advantages concentrated in contextual categories. Naming the missed categories recovered 79\% (48/61) of them, and discouraging over-redaction restored precision. No agentic architecture beat calibrated single-pass prompting (F1 0.906--0.907), but LLM outputs surfaced 414~candidate annotation gaps; re-annotation confirmed 227~PHI spans, against which the final prompt reached recall=0.981 (F1=0.907$\pm$0.002). Well-calibrated ICL resolves both the institutional PHI gap and the precision--recall trade-off in one LLM call per note. LLMs cost more to run than traditional methods, but that cost buys a way to audit the reference standard. LLMs are a legitimate, adaptable alternative to purpose-built de-identification systems; institution-specific prompt development should be the primary adaptation strategy.
摘要:次級使用電子健康紀錄需要去識別化,但現有系統忽略了\emph{制度性位置}的受保護健康資訊(PHI),例如醫院縮寫、建築名稱和其狀態由地方決定的內部代碼。我們詢問大型語言模型(LLMs)是否能透過上下文學習(ICL)填補這一空白並控制精確度與召回率的權衡。
在來自德克薩斯兒童醫院的100份註釋小兒科腫瘤學筆記(5,322個PHI範圍)中,我們對八個LLMs進行了基準測試,並與兩個專門構建的系統(Stanford TiDE,OpenMed PII)和兩個基於模式的基準進行比較。每個LLM在三個逐步具體化的提示下運行:(1)符合HIPAA的基準,(2)基準加上其遺漏的制度PHI類別,以及(3)提示2加上對過度刪除臨床內容的指示。我們隨後比較了14個多代理和集成配置與最佳單一提示,召回率是主要的安全指標。
LLMs的表現超過了專門構建的系統(最佳F1=0.918$\pm$0.001對比TiDE 0.779),優勢集中在上下文類別中。命名遺漏的類別恢復了79\%(48/61),而抑制過度刪除則恢復了精確度。沒有任何代理架構超越經過校準的單次提示(F1 0.906--0.907),但LLM輸出顯示了414個候選註釋缺口;重新註釋確認了227個PHI範圍,最終提示的召回率達到0.981(F1=0.907$\pm$0.002)。
良好校準的ICL在每個筆記中解決了制度PHI缺口和精確度與召回率的權衡。LLMs的運行成本高於傳統方法,但這一成本提供了一種審計參考標準的方式。
LLMs是專門構建的去識別化系統的合法且可調整的替代方案;特定機構的提示開發應該是主要的調整策略。
Policy Iteration with Human Feedback: Bringing Post-Training RL to In-context Learning
2608.16831v1 by Minh-Ha Nguyen, Cathy Shyr
Generative pretraining established reusable task representations; later work on language-based task conditioning and in-context learning showed that a fixed model could adapt its behavior from instructions and demonstrations. Policy Iteration with Human Feedback (PIHF) builds on this development and the recurrent evaluate-and-improve structure of generalized policy iteration. PIHF uses a pretrained language model as its execution substrate and moves persistent revision to a versioned natural-language policy and tool set. A language-model critic and clinical expert review complete-panel reasoning and tool-use trajectories to localize recurrent failures and form candidate revisions; the expert may reinterpret the evidence and retains authority over admission and rollback, while Recall@1 and Recall@5 validate outcomes after candidate execution. Across cumulative ablations and ultra-rare-disease benchmarks, a PIHF-derived policy improved Recall@1 in one proprietary executor and three open-weight executors spanning 3 to 49 billion active parameters. Gains were 32.7 percentage points for GPT-5.4 and 31.1 points for Qwen3.6-35B, a difference of 1.7 points. These results support the feasibility of using pretrained language models as fixed-weight execution substrates for expert-guided policy development in rare-disease diagnosis.
摘要:生成預訓練建立了可重用的任務表示;後續在基於語言的任務條件和上下文學習方面的研究顯示,固定模型可以從指令和示範中調整其行為。帶有人工反饋的策略迭代(PIHF)建立在這一發展及其一般化策略迭代的反覆評估和改進結構之上。PIHF使用預訓練的語言模型作為其執行基礎,並將持續修訂轉移到版本化的自然語言策略和工具集。語言模型評論員和臨床專家審查完整面板的推理和工具使用軌跡,以定位反覆失敗並形成候選修訂;專家可以重新解釋證據,並保留對接受和回滾的權威,而Recall@1和Recall@5在候選執行後驗證結果。
在累積的消融和超罕見疾病基準測試中,PIHF衍生的策略在一個專有執行器和三個開放權重執行器中改善了Recall@1,這些執行器的活動參數範圍從30億到490億。GPT-5.4的增益為32.7個百分點,Qwen3.6-35B的增益為31.1個百分點,兩者之間的差異為1.7個百分點。這些結果支持使用預訓練語言模型作為固定權重執行基礎,在罕見疾病診斷中進行專家引導的策略開發的可行性。
Historical Backtesting for Scientific Question Discovery: A Protocol and Astronomy Pilot
2608.16795v1 by Hui Mao
Systems that generate scientific research questions are evaluated today by expert scores, LLM-as-judge ratings, or curated case studies -- all subjective, none falsifiable. We formalize historical backtesting as an alternative: a system generates questions from a corpus frozen at a historical cutoff, the questions are frozen before any access to later literature, and a temporally isolated future corpus then determines whether each question was subsequently answered, partially addressed, independently posed, or ignored, and whether its underlying premise was supported or refuted. The protocol is model-agnostic: any system that emits frozen questions can be scored. We release reproducible astronomy instances with temporally isolated corpora, frozen questions, auditable labels, four reference baselines, and a submission interface. Two findings result. First, evidence-structure-first generation outperforms LLM-only prompting: across a generator decomposition crossed with a four-cutoff stress test (2010-2024, 798 judged questions) whose last window postdates model training, LLM-only generation shows memorized relevance without specific foresight, while a generator using no model weights at all finds questions whose premises the future refutes in every era. Second, a seven-rater agreement study (two blinded human annotators, five judge models, 90 items) indicts the outcome taxonomy rather than the judge: two careful humans agree at kappa = 0.17, every judge model agrees with the professional annotator as well or better (0.17-0.26), and frontier models agree with one another at 0.60 -- certifying an LLM judge by model-model agreement would have overstated its reliability threefold. A prospective instance -- 200 questions frozen 2026-08-17, scored 2027-2030 -- is released so the central claims become contamination-free tests that time itself will grade.
摘要:生成科學研究問題的系統今天由專家評分、LLM作為評判的評級或策劃的案例研究來評估——這些都是主觀的,沒有一個是可證偽的。我們將歷史回測形式化為一種替代方案:系統從在歷史截止日期凍結的語料庫中生成問題,這些問題在任何接觸後續文獻之前被凍結,然後一個時間上隔離的未來語料庫決定每個問題是否隨後得到了回答、部分解決、獨立提出或被忽視,以及其基礎前提是否得到了支持或反駁。該協議是模型無關的:任何發出凍結問題的系統都可以被評分。我們發布了可重複的天文學實例,具有時間上隔離的語料庫、凍結問題、可審計的標籤、四個參考基準和提交界面。結果有兩個發現。首先,證據結構優先的生成表現優於僅依賴LLM的提示:在一次生成器分解與四次截止壓力測試(2010-2024年,798個評判問題)的交叉中,最後一個窗口的時間超過了模型訓練,僅依賴LLM的生成顯示出記憶中的相關性而沒有具體的預見,而一個完全不使用模型權重的生成器則找到了在每個時代中未來反駁其前提的問題。第二,一項七評審者一致性研究(兩名盲人人工標註者,五個評判模型,90個項目)指控結果分類法而不是評判者:兩名謹慎的人類在kappa = 0.17的情況下一致,每個評判模型與專業標註者的意見一致或更好(0.17-0.26),而前沿模型之間的一致性為0.60——通過模型之間的一致性來認證LLM評判者將會高估其可靠性三倍。一個前瞻性實例——200個問題凍結於2026-08-17,於2027-2030年進行評分——被發布,以便中央主張成為不受污染的測試,時間本身將進行評分。
Unsupervised Anomaly Detection for Image Dataset Quality Assurance in Multi-Center Breast MRI
2608.16725v1 by Chiara Tappermann, Steffen Renisch, Lars Ole Schwen, Hans Meine, Horst K. Hahn, Eike Petersen
Corrupted, inconsistent, or anomalous data silently threatens the safety and reliability of medical AI. Despite growing regulatory recognition of dataset quality assurance (QA) for high-risk medical AI, scalable automated detection remains underdeveloped. We employ unsupervised anomaly detection (AD) and out-of-distribution (OOD) detection as an automated dataset QA mechanism for multi-center dynamic contrast-enhanced breast MRI. We build a controlled AD benchmark of 17 realistic QA-relevant anomaly types from six public datasets (protocol violations, processing errors, incorrect anatomical regions) and propose a taxonomy of radiological image anomalies based on human visual perception, enabling fine-grained analysis of AD failure modes. The benchmark includes near-, medium-far-, far-OOD samples, as well as in-distribution and external normal data. Four methods are evaluated: a projection-based method extended with a domain-specific feature extractor and a novel positional encoding, a reconstruction-based approach extended to full 3D volumes with an augmented training objective, and two unmodified hybrid OOD detection methods. Medium-far- and far-OOD samples are detected reliably, whereas near-OOD samples and external normal data from unseen institutions expose method-specific differences. The 3D reconstruction-based approach best balances detection performance (AUROC: 0.936) and generalization to unseen institutions. The projection-based method with positional encoding achieves the highest overall detection performance (AUROC: 0.954). Both hybrid methods exhibit critical failure modes, confirming that methods validated for one modality or anatomy may not generalize without domain-specific adaptation. Implants and mastectomies remain an open challenge for all methods. Our results establish a foundation and practical guidance on scalable unsupervised QA in medical AI pipelines.
摘要:腐敗、不一致或異常的數據默默威脅著醫療人工智慧的安全性和可靠性。儘管對高風險醫療人工智慧數據集質量保證(QA)的監管認識日益增長,但可擴展的自動檢測仍然發展不足。我們採用無監督異常檢測(AD)和分佈外(OOD)檢測作為多中心動態對比增強乳腺MRI的自動數據集QA機制。
我們建立了一個由六個公共數據集中的17種現實QA相關異常類型組成的受控AD基準(協議違規、處理錯誤、不正確的解剖區域),並根據人類視覺感知提出了一個放射影像異常的分類法,使得對AD失效模式的細緻分析成為可能。基準包括近距離、中遠距離、遠距離OOD樣本,以及分佈內和外部正常數據。評估了四種方法:一種基於投影的方法,擴展了特定領域的特徵提取器和新穎的位置編碼;一種基於重建的方法,擴展到完整的3D體積並具有增強的訓練目標;以及兩種未經修改的混合OOD檢測方法。
中遠距離和遠距離OOD樣本的檢測可靠,而近距離OOD樣本和來自未見機構的外部正常數據則顯示出方法特定的差異。基於3D重建的方法在檢測性能(AUROC:0.936)和對未見機構的泛化之間達到了最佳平衡。帶有位置編碼的基於投影的方法實現了最高的整體檢測性能(AUROC:0.954)。兩種混合方法都顯示出關鍵的失效模式,確認了針對一種模態或解剖結構驗證的方法可能無法在沒有特定領域適應的情況下進行泛化。植入物和乳房切除術對所有方法仍然是一個未解決的挑戰。我們的結果為醫療人工智慧管道中的可擴展無監督QA建立了基礎和實用指導。
Toward Better Assessment of LLMs' Performance in Clinical Error Detection
2608.16643v1 by Yifan Zhang, Rahmatollah Beheshti
Automated detection of errors in clinical documentation is a promising application of large language models (LLMs), yet decisions to deploy such models rest on benchmarks that evaluate each clinical note in isolation. Error-detection benchmarks are typically constructed by injecting errors into notes, such that each erroneous note has a natural counterpart. Aggregate discriminative metrics (e.g., balanced accuracy or F1) do not exploit this structure. We show that this omission is consequential. In particular, evaluating 15 diverse LLMs on 4 standardized clinical error-detection test sets across 3 languages, we find that 13 of 15 models fall below the level of random pairwise discrimination, even while achieving F1 scores that standard practice would read as moderate. We also observe that the underlying bias patterns differ across languages: the same model can default to "no error" on one language and over-flag errors on another. To diagnose where discrimination breaks down, we further introduce a procedure to score the evidence models cite in their outputs. We find that while models consistently locate error-relevant content, they fail to produce the corresponding correct verdict on the clean counterpart. Finally, we show that F1 and pairwise accuracy are driven in opposite directions by the same underlying bias, so that ranking models by F1 may systematically promote the weakest discriminators. For safety-critical clinical NLP applications, we advocate for supplementing aggregate metrics with paired evaluations in benchmark reporting. Code and analysis scripts are available at https://github.com/healthylaife/paired-clinical-eval.
摘要:自動檢測臨床文檔中的錯誤是大型語言模型(LLMs)的有前景應用,然而,部署這些模型的決策依賴於評估每條臨床記錄的基準,這些基準是孤立評估的。錯誤檢測基準通常是通過在記錄中注入錯誤來構建的,使得每條錯誤記錄都有一個自然的對應記錄。聚合的判別指標(例如,平衡準確率或F1)並未利用這一結構。我們表明,這一遺漏是有後果的。具體而言,在對3種語言的4個標準化臨床錯誤檢測測試集上評估15種不同的LLMs時,我們發現15個模型中有13個的表現低於隨機成對判別的水平,即使它們的F1分數在標準實踐中被視為中等。我們還觀察到,潛在的偏見模式在不同語言之間存在差異:同一模型在一種語言上可能默認為「無錯誤」,而在另一種語言上則過度標記錯誤。為了診斷判別失效的原因,我們進一步引入了一個程序來評分模型在其輸出中引用的證據。我們發現,儘管模型始終能定位與錯誤相關的內容,但它們未能對乾淨的對應記錄給出正確的判決。最後,我們表明F1和成對準確率受到同一潛在偏見的驅動,方向卻相反,因此根據F1對模型進行排名可能系統性地促進最弱的判別者。對於安全關鍵的臨床NLP應用,我們主張在基準報告中用成對評估來補充聚合指標。代碼和分析腳本可在 https://github.com/healthylaife/paired-clinical-eval 獲得。
Degradation-Aligned Self-Supervised Learning for State of Health Estimation of Lithium-Ion Batteries under Label Sparsity
2608.16612v1 by Jiaqi Yao, Julia Kowal
An accurate estimation of the state of health (SOH) underpins a safe and optimized use of the battery system. Although compelling, data-driven SOH estimation models typically require large amounts of high-quality labeled cycling data, while in practice such labels are often sparse in both quantity and coverage. Therefore, in this work, we propose a degradation-aligned self-supervised learning (SSL) framework based on a convolutional neural network-gated recurrent unit (CNN-GRU) model, which learns aging-consistent representations from unlabeled data through a cycle-order ranking objective as the pretext task for pretraining, thereby enabling robust SOH estimation after fine-tuning on sparsely labeled data. Test results showcase that the proposed ranking-based SSL approach proves to endow the pretrained model with degradation-aligned information from unlabeled data, and after fine-tuning the model can carry out accurate, robust SOH estimation, even when only an extremely limited amount of 1% of unevenly distributed labeled training data is available, where the MAE of 1.718% and RMSE of 2.329% can be achieved on the test cell. In addition, in-depth analyses are presented regarding the influences of label distribution of battery degradation data. We believe this work could shed new light on SOH estimation of lithium-ion batteries under label sparsity in real-world applications.
摘要:準確的健康狀態(SOH)估計是安全且優化使用電池系統的基礎。雖然數據驅動的SOH估計模型非常有說服力,但通常需要大量高質量的標註循環數據,而在實際情況中,這些標註往往在數量和覆蓋範圍上都很稀疏。因此,在本研究中,我們提出了一種基於卷積神經網絡-門控遞歸單元(CNN-GRU)模型的降解對齊自監督學習(SSL)框架,通過循環順序排名目標作為預訓練的前置任務,從未標註數據中學習與老化一致的表示,從而在稀疏標註數據上進行微調後實現穩健的SOH估計。測試結果顯示,所提出的基於排名的SSL方法使預訓練模型從未標註數據中獲得了降解對齊的信息,並且在微調後,該模型能夠進行準確且穩健的SOH估計,即使僅有極少量的1%不均勻分佈的標註訓練數據可用,測試電池的MAE可達1.718%和RMSE可達2.329%。此外,還對電池降解數據的標註分佈影響進行了深入分析。我們相信這項工作可以為在現實應用中標註稀疏的鋰離子電池SOH估計提供新的見解。
CACSurv: Concordance-Aligned Comparative Learning with Large Language Models for Cancer Survival Prediction
2608.16594v1 by Tianqi Xiang, Qixiang Zhang, Xinpeng Ding, Yi Li, Xiaomeng Li
Cancer survival prediction supports treatment planning, risk stratification, and follow-up management. Existing methods use structured clinical variables, whole-slide images, genomic profiles, or multimodal inputs, while patient reports remain underexplored. We study report-centric survival prediction using reports that organize pathological, clinical, and molecular evidence. Large language models (LLMs) can reason over such reports, but case-wise time regression introduces two mismatches. First, a formulation mismatch arises because survival evaluation depends on ordering comparable patients, whereas independent time predictions do not enforce ranking consistency. Second, a supervision mismatch arises because a censored patient's observed time indicates survival beyond that point and cannot serve as an exact regression target, although it still implies orderings relative to patients who died earlier. To address these mismatches, we propose CACSurv, a Concordance-Aligned Comparative framework for report-centric survival prediction. CACSurv reformulates survival modeling as mini-cohort comparative reasoning, where an LLM predicts relative prognostic orderings. We introduce concordance-aligned rewards derived from comparable relations under right censoring, enabling censored outcomes to provide ranking supervision without exact event-time targets. At inference, Monte Carlo Reference Aggregation compares each patient with sampled references and aggregates positions into a cohort-level ranking. We establish TCGA-SurvReport, a benchmark covering six TCGA cancer cohorts. CACSurv achieves the highest C-index on all six cohorts and an average C-index of 0.722, outperforming the strongest published survival model by 6.5 percentage points and the strongest LLM time-regression baseline by 4.2 percentage points. Our code, models, and dataset will be available at https://github.com/xmed-lab/CACSurv.
摘要:癌症生存預測支持治療計劃、風險分層和後續管理。現有方法使用結構化臨床變數、全幻燈片影像、基因組資料或多模態輸入,而患者報告仍然未被充分探索。我們研究以報告為中心的生存預測,使用組織病理、臨床和分子證據的報告。大型語言模型(LLMs)可以對這些報告進行推理,但逐案例時間回歸引入了兩個不匹配。首先,因為生存評估依賴於可比較患者的排序,而獨立的時間預測並不強制執行排名一致性,因此產生了表述不匹配。其次,因為被審查患者的觀察時間表示超過該點的生存,並不能作為精確的回歸目標,儘管它仍然暗示了相對於早逝患者的排序,因此產生了監督不匹配。為了解決這些不匹配,我們提出了CACSurv,一個以報告為中心的生存預測的協調對齊比較框架。CACSurv將生存建模重新表述為小型隊列比較推理,其中LLM預測相對的預後排序。我們引入了基於右側審查下可比較關係衍生的協調對齊獎勵,使得被審查的結果能夠提供排名監督,而不需要精確的事件時間目標。在推理階段,蒙地卡羅參考聚合將每位患者與抽樣參考進行比較,並將位置聚合成隊列級別的排名。我們建立了TCGA-SurvReport,一個涵蓋六個TCGA癌症隊列的基準。CACSurv在所有六個隊列上達到了最高的C指數,平均C指數為0.722,超越了最強的已發表生存模型6.5個百分點和最強的LLM時間回歸基線4.2個百分點。我們的代碼、模型和數據集將在https://github.com/xmed-lab/CACSurv上提供。
Large language models as synthetic clinical experts to inform longitudinal rare-disease modeling
2608.16507v1 by Clemens Schächter, Astrid Pechmann, Janbernd Kirschner, Jan Hasenauer, Harald Binder
Due to the limited amount of information, modeling longitudinal rare-disease data can benefit from integrating clinical knowledge. Yet, elicitation of expert knowledge and formalization for model fitting is challenging, in particular due to limited time of clinical experts. To nevertheless make domain knowledge accessible during model fitting, we use large language models (LLMs) as synthetic clinical experts to supervise a variational-autoencoder-based approach that learns low-dimensional latent summaries of visit-level observations. Specifically, LLMs are queried offline on textual descriptions of patient observations to obtain judgments, e.g., the suspected clinical category. To improve the variational autoencoder fit, we train a differentiable surrogate model on these judgments and augment the loss function to encourage reconstructions that preserve the clinical-label distribution of their corresponding input profile. In an application to longitudinal motor-function assessments from children with spinal muscular atrophy, we map visit-level clinical profiles to low-dimensional representations that are linked by a multivariate mixed-effects model. The synthetic expert loss discourages reconstructions that remain numerically close in data space but alter the clinical interpretation of the reconstructed motor function profile, such as by crossing a disease-type boundary. We thus reduced disagreement between original and reconstructed SMA type labels from about 11 to 7 percent. Furthermore, informing the latent representation by the synthetic expert improved prediction of motor function milestones compared with unsupervised latent representations and a data-level baseline. These results suggest that incorporating LLMs into model fitting can make clinical knowledge available to representation learning and improve clinical faithfulness for longitudinal rare-disease data.
摘要:由於資訊量有限,建模縱向罕見疾病數據可以從整合臨床知識中受益。然而,專家知識的引出和模型擬合的形式化是具有挑戰性的,特別是由於臨床專家的時間有限。儘管如此,為了在模型擬合過程中使領域知識可用,我們使用大型語言模型(LLMs)作為合成臨床專家,來監督基於變分自編碼器的方法,該方法學習訪問級觀察的低維潛在摘要。具體來說,我們在患者觀察的文本描述上離線查詢LLMs以獲得判斷,例如,懷疑的臨床類別。為了改善變分自編碼器的擬合,我們在這些判斷上訓練了一個可微分的替代模型,並增強損失函數以鼓勵重建保持其對應輸入特徵的臨床標籤分佈。在對脊髓性肌萎縮症兒童的縱向運動功能評估的應用中,我們將訪問級臨床特徵映射到由多變量混合效應模型鏈接的低維表示。合成專家損失會抑制在數據空間中數值上接近但改變重建運動功能特徵的臨床解釋的重建,例如通過跨越疾病類型邊界。因此,我們將原始和重建的SMA類型標籤之間的分歧從約11%減少到7%。此外,通過合成專家告知潛在表示,與無監督潛在表示和數據級基準相比,運動功能里程碑的預測得到了改善。這些結果表明,將LLMs納入模型擬合可以使臨床知識可用於表示學習,並改善縱向罕見疾病數據的臨床真實性。
Offline Reinforcement Learning for Hemodynamic Management of Sepsis in the ICU: a MIMIC-IV Study with Dual Off-Policy Evaluation
2608.16482v1 by Marc Pérez-Roig, David Fernández-Narro, Carlos Sáez
The dosing of intravenous fluids and vasopressors in sepsis is a sequential decision made under uncertainty and guided largely by clinical judgment, which makes it a natural target for reinforcement learning from historical care. Because a learned policy cannot be trialed on patients, its value must be estimated off-policy, and such estimates can be fragile and optimistic. This work advances the reliable evaluation of sepsis treatment policies by combining off-policy estimation, reliability diagnostics, and clinician-agreement analyses in a transparent validation framework. We modeled fluid and vasopressor dosing on a cohort of 36,872 septic ICU stays drawn from the MIMIC-IV critical-care database, as a discretized Markov decision process with 1,000 states and 25 actions, defined by a five-by-five grid of fluid and vasopressor levels and solved by policy iteration. The clinicians' behavior policy was estimated with a random forest, which mitigated the collapse of the Effective Sample Size (ESS 50.1 against 4.0 with smoothed counts) that otherwise destabilizes the importance-sampling estimate. The learned policy was evaluated with two estimators, weighted importance sampling (WIS) and fitted Q evaluation (FQE), with the ESS and clinician agreement as reliability checks. An empirical variable selection found that the composition of the state matters more than its size. Both estimators place the learned policy above the clinicians' return (WIS 50.8 and FQE 46.8 against 38.2, ESS 50.1), yet it departs only modestly from observed practice (total variation 0.18), favoring less intravenous fluid. These retrospective single-center off-policy results support the learned policy as a clinically plausible refinement of observed practice and motivate its further evaluation as a discordance-based clinical decision-support approach.
摘要:靜脈輸液和血管加壓劑在敗血症中的劑量是一個在不確定性下做出的順序決策,主要依賴臨床判斷,這使其成為從歷史護理中進行強化學習的自然目標。由於學習到的策略不能在患者身上進行試驗,因此其價值必須在政策外進行估算,而這樣的估算可能是脆弱和樂觀的。本研究通過在透明的驗證框架中結合政策外估算、可靠性診斷和臨床醫生一致性分析,推進了敗血症治療政策的可靠評估。我們在從MIMIC-IV重症護理數據庫中提取的36,872例敗血症ICU住院病例上建模了液體和血管加壓劑的劑量,將其定義為一個具有1,000個狀態和25個行動的離散馬可夫決策過程,並通過政策迭代進行求解。臨床醫生的行為政策是通過隨機森林估算的,這減輕了有效樣本量(ESS 50.1對比平滑計數的4.0)的崩潰,否則會使重要性抽樣估算不穩定。學習到的策略通過兩個估算器進行評估,權重重要性抽樣(WIS)和擬合Q評估(FQE),並以ESS和臨床醫生一致性作為可靠性檢查。一項實證變量選擇發現,狀態的組成比其大小更為重要。兩個估算器均將學習到的策略置於臨床醫生的回報之上(WIS 50.8和FQE 46.8對比38.2,ESS 50.1),但它僅與觀察到的實踐有適度的偏離(總變異0.18),更傾向於較少的靜脈輸液。這些回顧性單中心的政策外結果支持學習到的策略作為臨床上合理的觀察實踐的改進,並促使其作為基於不一致的臨床決策支持方法進一步評估。
Adaptive Post-Processing Drives Instance-Level Detection in Stroke Lesion Segmentation
2608.16377v1 by Qinghui Liu, Jon André Ottesen, Atle Bjørnerud, Kyrre Eeg Emblem
Instance-level lesion detection has been an increasingly larger focal point in medical image segmentation besides the more standard voxel-level overlap. Still, most pipelines are trained and post-processed for voxel overlap alone. In particular, the mismatch is most pronounced for small lesions, where a near-miss prediction---substantial overlap that falls just short of the instance-matching threshold---scores the same as a complete miss. In our ISLES'26 submission, we found that closing this gap mattered far more in post-processing than in architecture design. Our Volume-Conditioned Adaptive Post-Processing (VCAP) scheme adjusts component-size thresholds to each case's predicted lesion burden, improving Lesion-F1 by 0.032 (unbiased cross-fold estimate)---approximately 6 times larger than any architectural change we tested. A resolution-aware attention architecture (Viola2Plus), designed for small-lesion segmentation, shows why the distinction matters: it left small-lesion Dice unchanged but raised small-lesion detection rate by 3.7\%, a real effect voxel-overlap metrics alone would have missed. Under 5-fold cross-validation on the 1,453-case training set, our post-processed two-architecture ensemble achieves Dice 0.651 and Lesion-F1 0.614, versus 0.644 and 0.573 for the unprocessed single-model baseline.
摘要:實例級病變檢測在醫學影像分割中越來越受到重視,除了更標準的體素級重疊外。儘管如此,大多數管道仍僅針對體素重疊進行訓練和後處理。特別是,這種不匹配在小病變中最為明顯,近乎錯誤的預測——實質重疊但未達到實例匹配閾值——與完全錯過的得分相同。在我們的ISLES'26投稿中,我們發現縮小這一差距在後處理中比在架構設計中更為重要。我們的體積條件自適應後處理(VCAP)方案根據每個案例的預測病變負擔調整組件大小閾值,將病變F1提高了0.032(無偏交叉折估計)——這大約是我們測試的任何架構變更的6倍。針對小病變分割設計的解析度感知注意力架構(Viola2Plus)顯示了這一區別的重要性:它對小病變的Dice保持不變,但將小病變檢測率提高了3.7\%,這是一個僅依賴體素重疊指標無法捕捉的實際效果。在對1,453個案例訓練集進行5折交叉驗證的情況下,我們的後處理雙架構集成達到了Dice 0.651和病變F1 0.614,而未處理的單模型基線則為0.644和0.573。
Foresight-England: Development of a National-Scale Generative AI Model of Electronic Health Records for Medical Event Prediction across the COVID-19 Pandemic
2608.16273v1 by Simon Ellershaw, Christopher Tomlinson, Zeljko Kraljevic, Spiros Denaxas, Harry Hemingway, Cathie Sudlow, Angela M. Wood, Anoop D. Shah, Richard Dobson
Foresight-England (Foresight-E) is the first national-scale generative foundation model of electronic health records (EHRs), developed as a research pilot strictly for COVID-19 research. We evaluated its ability to model the direct and indirect effects of the pandemic. Trained from scratch entirely within the NHS England Secure Data Environment, Foresight-E is a 243-million-parameter transformer decoder. It was trained and evaluated on de-identified, longitudinal EHRs of approximately 61 million individuals, integrating primary/secondary care, death registrations, and COVID-19 data. Training and validation used a 90% subset (54.9 million) spanning November 2018 to December 2022; the remaining 10% (6.1 million) was held out for evaluation. Foresight-E models patient timelines autoregressively, predicting the next medical event given their prior history. At inference, it operates zero-shot, predicting any concept in its ~40,000-code vocabulary without task-specific training. Our tokenisation scheme retains the clinical granularity of ICD-10, OPCS-4, and SNOMED CT codes, jointly representing absolute and relative timing. We designed an evaluation framework for 30-day COVID-19 hospitalisation and mortality, including subgroup analyses by demographic factors and vaccination status. To assess generalisation to unseen future data and the pandemic's indirect effects, we tested the model on medical events from 2023 (beyond its training period), benchmarking against logistic regression and XGBoost. As detailed in the Project Status section, NHS England has paused access to data for the Foresight-E project, meaning quantitative results are currently unavailable. Instead, we share our strategy for tokenisation, architecture, training, inference, and evaluation as a methodological template and case study in the challenges of building population-scale EHR foundation models.
摘要:Foresight-England (Foresight-E) 是首個全國規模的電子健康紀錄 (EHRs) 生成基礎模型,作為針對 COVID-19 研究的研究試點而開發。
我們評估了它建模疫情直接和間接影響的能力。
Foresight-E 完全在 NHS England 安全數據環境中從零開始訓練,是一個擁有 2.43 億參數的Transformer解碼器。
它在約 6100 萬人的去識別化、縱向 EHRs 上進行訓練和評估,整合了初級/次級護理、死亡登記和 COVID-19 數據。
訓練和驗證使用了 90% 的子集(5490 萬),涵蓋了 2018 年 11 月到 2022 年 12 月;剩餘的 10%(610 萬)則保留用於評估。
Foresight-E 自回歸地建模患者時間線,根據其先前的歷史預測下一個醫療事件。
在推理時,它以零樣本操作,預測其約 40,000 種代碼詞彙中的任何概念,而無需特定任務的訓練。
我們的標記方案保留了 ICD-10、OPCS-4 和 SNOMED CT 代碼的臨床細節,聯合表示絕對和相對時間。
我們設計了一個評估框架,用於 30 天 COVID-19 住院和死亡率,包括按人口統計因素和疫苗接種狀態的子群分析。
為了評估對未見未來數據的泛化能力和疫情的間接影響,我們在 2023 年的醫療事件上測試了該模型(超出其訓練期間),並與邏輯回歸和 XGBoost 進行基準比較。
正如項目狀態部分詳細說明的那樣,NHS England 已暫停對 Foresight-E 項目的數據訪問,這意味著目前無法獲得定量結果。
相反,我們分享了我們的標記化、架構、訓練、推理和評估的策略,作為建立人口規模 EHR 基礎模型挑戰的 методологический шаблон и кейс-исследование。
A cross-modal generative model for incomplete and degraded prostate MRI with multicentre clinical validation
2608.16233v1 by Siyuan Ma, Liang He, Mengying Zhu, Yi Chai, Mengyao Lyu, Haowei Wang, Qizhen Lan, HaoBo Sun, Qixin Zhang, Jingli Chen, Xiaobing Wei, Jiaming Liu, Guiqin Liu, Qianwen Zhang, Yang Liu, Dacheng Tao, Guangyu Wu
Missing or degraded sequences can limit prostate multiparametric MRI. We developed MSCNet, a sequence-conditioned cross-modal generative framework for reconstructing unavailable contrasts and restoring degraded acquisitions. Across ten completion tasks, task-specific MSCNet achieved mean structural similarity of 0.818 versus 0.798 for the strongest task-matched comparators; matched-capacity analyses showed larger differences in lesion fidelity and boundary preservation. In a blinded 1,000-case reader study, overall image quality met the prespecified non-inferiority criterion for DWI, ADC and T2W completion, but not T1W. In a separate 200-case diagnostic assessment, AUCs for clinically significant cancer were 0.860 with acquired images, 0.841 with MSCNet and 0.797 with baseline-generated images. A locked 186-case three-hospital cohort supported multicentre transportability. These retrospective results support quality-controlled cross-modal reconstruction as an adjunct to acquired prostate MRI.
摘要:缺失或劣化的序列可能限制前列腺多參數MRI。我們開發了MSCNet,一種序列條件的跨模態生成框架,用於重建不可用的對比和恢復劣化的獲取。在十個補全任務中,特定任務的MSCNet達到了0.818的平均結構相似度,而最強的任務匹配比較者為0.798;匹配容量分析顯示病變真實性和邊界保護方面的差異更大。在一項盲法的1000例讀者研究中,整體影像質量達到了預先指定的DWI、ADC和T2W補全的非劣性標準,但對於T1W則不符合。在另一項200例的診斷評估中,臨床顯著癌症的AUC為獲取影像的0.860,MSCNet的0.841和基線生成影像的0.797。一個鎖定的186例三醫院隊列支持多中心可轉運性。這些回顧性結果支持質量控制的跨模態重建作為獲取前列腺MRI的輔助。
BaT: Towards Self-Evolving Medical Research Agent with Stage Rubrics
2608.16211v1 by Junqi Liu, Yufan He, Yexiao He, Pengfei Guo, Dong Yang, Andriy Myronenko, Can Zhao, Hanrong Ye, Tianhao Qi, Yuyin Zhou, Daguang Xu, Yucheng Tang
Long-horizon agents are beginning to automate complete workflows that produce code, reports, and research artifacts. Medical imaging workflows are multi-stage and data-sensitive, while expert trajectories remain scarce and difficult to share. Structured benchmarks can localize failures through stage-level rubrics, but standard post-training discards these diagnostics before the next training round. We present Benchmark-as-Teacher (BaT), a recursive self-improvement system for agent post-training. BaT contains two linked components: the asynchronous Stage Bank data pipeline and BiCuRL (Bilevel Curriculum Reinforcement Learning), its self-improving post-training method. Stage Bank synthesizes content-isolated training states outside the policy-update loop. BiCuRL uses a fixed held-out evaluation to select the next stage curriculum, verifies rollouts with task rubrics, updates the policy with GRPO, and returns the candidate checkpoint to evaluation. On AutoMedBench-Lite, BaT-4B and BaT-9B more than double the Overall scores of their Qwen Instruct baselines. BaT-9B Agent reaches 79.6 Overall, exceeding Claude Opus 4.6 with Claude Code at 77.5.
摘要:長期代理正在開始自動化完整的工作流程,產生代碼、報告和研究文獻。醫療影像工作流程是多階段且對數據敏感的,而專家軌跡則仍然稀缺且難以分享。結構化基準可以通過階段級別的評分標準定位失敗,但標準的後訓練會在下一輪訓練之前丟棄這些診斷。我們提出了Benchmark-as-Teacher (BaT),這是一個用於代理後訓練的遞歸自我改進系統。BaT包含兩個相互關聯的組件:異步的Stage Bank數據管道和BiCuRL(雙層課程強化學習),其自我改進的後訓練方法。Stage Bank在政策更新循環之外合成內容孤立的訓練狀態。BiCuRL使用固定的保留評估來選擇下一階段的課程,通過任務評分標準驗證回合,使用GRPO更新政策,並將候選檢查點返回給評估。在AutoMedBench-Lite上,BaT-4B和BaT-9B的整體分數超過了其Qwen Instruct基準的兩倍。BaT-9B代理達到79.6的整體分數,超過了Claude Opus 4.6,而Claude Code則為77.5。
Picking the Right Image to Classify: Reliable-Input Selection in Teledermatology
2608.16198v1 by Fabian Gröger, Marco Weishaupt, Philippe Gottfrois, Simone Lionetti, Linda Wermelinger, Nipun Ranasekara, Ludovic Amruthalingam, Alexander A. Navarini, Marc Pouly
Dermatology models face distribution shifts in teledermatology settings, where submitted images differ from the training data in lighting, angle, distance, focus, and framing. These test-time images are ordinary clinical photographs, but some fall outside the model's training conditions, leading the model to often misclassify them due to shifts in acquisition between training and deployment. When multiple images of the same case exist (several photos of one patient or lesion), a natural way to improve accuracy is therefore to select the image the model is most likely to classify correctly. We call this task reliable-input selection. An oracle that, for each case, selects a correctly classified image when one exists raises weighted F1 by about 20 percentage points on average across six dermatology datasets and nine frozen backbones. This oracle is an upper bound that sees the labels, whereas a selector must choose blindly. Capturing this gain in practice is hard. A selector that needs no pretraining data applies to any frozen model, including those whose data is not public. It must judge reliability from quantities the model exposes at inference: its embeddings, their norms, and its confidence. We benchmark four such training-data-free selectors: the embedding norm, the neighborhood consensus among a case's images, the stability of the prediction under small perturbations, and the model's own confidence. No training-data-free selector substantially narrows this oracle gap. The best of them is the model's own confidence, but it recovers only a small part of the gap on the clinical datasets. A small labeled reference set does not help either: the best selector overall, a fusion of confidence and Mahalanobis distance, still leaves most of the gap. To our knowledge, this is the first study to introduce and benchmark reliable input selection, a clinically important, unsolved task.
摘要:皮膚科模型面臨在遠程皮膚科環境中分佈轉移的挑戰,提交的圖像在照明、角度、距離、焦點和構圖上與訓練數據有所不同。這些測試時的圖像是普通的臨床照片,但有些超出了模型的訓練條件,導致模型經常因訓練與部署之間的獲取差異而錯誤分類。當同一病例存在多張圖像(多張同一患者或病變的照片)時,改善準確性的自然方法是選擇模型最有可能正確分類的圖像。我們稱這個任務為可靠輸入選擇。對於每個案例,當存在正確分類的圖像時,選擇這樣的圖像的神諭平均提高六個皮膚科數據集和九個凍結骨幹的加權F1約20個百分點。這個神諭是一個上限,能看到標籤,而選擇器必須盲目選擇。在實踐中捕捉這一增益是困難的。需要無預訓練數據的選擇器適用於任何凍結模型,包括那些數據不公開的模型。它必須根據模型在推斷時暴露的數量來判斷可靠性:其嵌入、它們的範數和模型的信心。我們基準測試了四種無訓練數據的選擇器:嵌入範數、案例圖像之間的鄰域共識、在小擾動下預測的穩定性,以及模型自身的信心。沒有一種無訓練數據的選擇器能顯著縮小這一神諭差距。其中最好的選擇器是模型自身的信心,但它在臨床數據集上僅恢復了差距的一小部分。小型標記參考集也沒有幫助:整體最佳的選擇器,即信心和馬哈拉諾比斯距離的融合,仍然留下了大部分差距。據我們所知,這是第一項引入和基準測試可靠輸入選擇的研究,這是一個臨床重要的未解決任務。
TokenSTFormer: A Tokenized Spatial-temporal Attention Model for Holistic Motion Analysis in Adolescent Idiopathic Scoliosis Screening
2608.16122v1 by Dong Chen, Kenneth M. C. Cheung
Adolescent Idiopathic Scoliosis (AIS) is a prevalent spinal deformity in adolescents that, if left untreated, can result in severe health outcomes. Traditional screening methods are limited by subjective interpretation, reliance on professional expertise and low scalability. To address these challenges, we present ScoliGait dataset, which comprises 1,516 gait video clips paired with corresponding X-ray records. We also introduce TokenSTFormer, a novel model that tokenizes spatial and temporal semantics to enhance feature representation and convergence. Our model achieves state-of-the-art performance, surpassing vanilla Vision Transformer encoder across key metrics, including accuracy of 0.79. This study highlights the potential of leveraging holistic motion features derived from gait video and attention-based models for scalable, cost-effective AIS screening, paving the way for future clinical applications in scoliosis detection.
摘要:青少年特發性脊柱側彎(AIS)是青少年中常見的脊柱畸形,如果不加以治療,可能會導致嚴重的健康後果。傳統的篩檢方法受到主觀解釋、依賴專業知識和低可擴展性的限制。為了解決這些挑戰,我們提出了 ScoliGait 數據集,其中包含 1,516 段行走視頻片段,並配有相應的 X 光記錄。我們還介紹了 TokenSTFormer,一種新型模型,將空間和時間語義進行標記化,以增強特徵表示和收斂。我們的模型在關鍵指標上達到了最先進的性能,超越了普通的視覺Transformer編碼器,包括 0.79 的準確率。本研究突顯了利用從行走視頻中獲得的整體運動特徵和基於注意力的模型進行可擴展、成本效益高的 AIS 篩檢的潛力,為未來脊柱側彎檢測的臨床應用鋪平了道路。
Decoupling Parcellation from Classification: Systematic Benchmark of Fast Brain Segmentation Methods for Alzheimer's Disease Detection
2608.16039v1 by Jiadao Zou, Hongyu Guo, Wei Xi
Brain parcellation and classification are typically evaluated in isolation, yet downstream AD detection performance depends on their interaction. We decouple these components and systematically benchmark fast deep learning parcellation methods (SynthSeg+, OpenMAP-T1) against the FreeSurfer (FS-HV) clinical baseline through down- stream AD classification on OASIS-1. Our factorial design evaluates three parcellation methods, two volumetry strategies (hard vs. soft), and four classifier paradigms (clinical thresholds, supervised feedforward networks, ensemble methods, and foundation models with zero/few-shot prompting), with all results quantified using BCa Bootstrap 95% confidence intervals.
摘要:腦區劃分和分類通常是孤立評估的,但下游阿茲海默症檢測性能取決於它們的相互作用。 我們將這些組件解耦,並系統性地基準測試快速深度學習劃分方法(SynthSeg+、OpenMAP-T1)與 FreeSurfer(FS-HV)臨床基準,通過對 OASIS-1 的下游阿茲海默症分類進行比較。 我們的因子設計評估了三種劃分方法、兩種體積測量策略(硬性與軟性)以及四種分類器範式(臨床閾值、監督式前饋網絡、集成方法和基於零/少量提示的基礎模型),所有結果均使用 BCa Bootstrap 95% 置信區間進行量化。
Breaking and Defending LLM-Powered Social Media Bot Detection Systems
2608.15893v1 by Nof Orenstein, Yoni Birman
The rise of social media bots poses a persistent threat, enabling misinformation, opinion manipulation, and the erosion of trust in online platforms. To combat this, machine learning systems have been developed to detect and limit bot activity, but attackers continuously adapt through techniques such as adversarial learning and behavior imitation, fueling an ongoing arms race between bots and detection tools. Recent advances in large language models (LLMs) have significantly improved bot detection by enabling deeper semantic and contextual analysis of accounts and their content. However, this shift also introduces new attack surfaces, allowing adversaries to craft exploits that directly target the reasoning and generation mechanisms of LLM-based classifiers. Industry tools such as Anthropic's Claude Code Security similarly leverage LLMs for security-critical decisions, further motivating a careful study of their attack surfaces. In this work, we investigate both the offensive and defensive aspects of LLM-powered, threat-specific cybersecurity applications. While centered on the challenge of social media bot detection, our methodology and insights generalize to a broad class of LLM-powered cybersecurity systems, including phishing detection, email classification, and fraud analysis. We introduce two novel adversarial attack strategies that systematically exploit the semantic and contextual weaknesses of LLM-based classifiers, degrading their detection accuracy by up to 48%. To counter these threats, we propose a robust multi-LLM defense architecture designed to preserve detection reliability under adaptive adversarial conditions. Our solution, LSABRE (LLM-powered Social Adversarial Bot Recognition Ensemble), is a multi-LLM framework that substantially improves robustness across a range of attacks, maintaining 86% detection accuracy even under strong, adaptive adversarial pressure.
摘要:社交媒體機器人的興起帶來了持續的威脅,使得錯誤信息、意見操控以及對在線平台的信任侵蝕變得可能。為了應對這一挑戰,已開發出機器學習系統來檢測和限制機器人活動,但攻擊者不斷通過對抗學習和行為模仿等技術進行適應,促進了機器人和檢測工具之間的持續軍備競賽。最近在大型語言模型(LLMs)方面的進展顯著改善了機器人檢測,通過使帳戶及其內容的語義和上下文分析更深入。然而,這一轉變也引入了新的攻擊面,使對手能夠設計直接針對基於LLM的分類器的推理和生成機制的利用方式。行業工具如Anthropic的Claude Code Security同樣利用LLMs進行安全關鍵決策,進一步促使對其攻擊面的仔細研究。在這項工作中,我們調查了基於LLM的特定威脅網絡安全應用的攻擊和防禦兩個方面。雖然重點放在社交媒體機器人檢測的挑戰上,我們的方法論和見解可以推廣到廣泛的基於LLM的網絡安全系統,包括釣魚檢測、電子郵件分類和詐騙分析。我們提出了兩種新穎的對抗攻擊策略,系統性地利用基於LLM的分類器的語義和上下文弱點,將其檢測準確率降低多達48%。為了應對這些威脅,我們提出了一種穩健的多LLM防禦架構,旨在在自適應對抗條件下保持檢測的可靠性。我們的解決方案LSABRE(基於LLM的社交對抗機器人識別集成)是一個多LLM框架,顯著提高了在各種攻擊下的穩健性,即使在強大的自適應對抗壓力下也能保持86%的檢測準確率。
Characterising cardiac tissue properties with graph neural networks
2608.15843v1 by Ching-En Chiu, Yoo Ri Kim, Magdi Saba, Danilo Mandic, Marta Varela
Characterising electrophysiological properties of cardiac tissue efficiently and accurately from spatially sparse intracardiac measurements is clinically important for localising ablation targets and improving arrhythmia treatment. We developed a graph neural network-based framework trained on synthetic electrogram signals on 2D flat surfaces to identify areas of interest in the context of cardiac ablation for premature ventricular complexes (PVCs). Our method achieved an average precision of 0.96, 0.97, and 0.95 for the detection of single-patch fibrosis, rapid depolarisation and high excitability, respectively. The trained model can then be applied to 2D curved surfaces with few-shot fine-tuning, demonstrating its generalisation capability. Future work will develop this framework further for clinical use in PVC ablation.
摘要:有效且準確地從空間稀疏的心內測量中描述心臟組織的電生理特性,對於定位消融目標和改善心律不整治療具有臨床重要性。
我們開發了一個基於圖神經網絡的框架,該框架在2D平面上對合成電圖信號進行訓練,以識別在心臟消融中與早期心室複雜(PVCs)相關的興趣區域。
我們的方法在檢測單一斑塊纖維化、快速去極化和高興奮性方面,分別達到了0.96、0.97和0.95的平均精度。
訓練好的模型可以通過少量調整應用於2D曲面,顯示出其泛化能力。
未來的工作將進一步開發這一框架,以便在PVC消融中用於臨床應用。
PLeDO: Pain Level Detection for Osteoarthritis from EMR Data
2608.15719v1 by Yuhao Chen, Jiahao Cai, Nafiz Sadman, Farhana Zulkernine, John Queenan, David Barber
Osteoarthritis (OA) is a progressive chronic joint disease resulting in a breakdown of articular cartilage and bone when damaged joint tissues are not able to normally repair themselves. The aim of this pilot research study is to understand the pain severity for OA from patients' primary care Electronic Medical Records (EMR), both from the structured medical data and the unstructured chart note data using information extraction, natural language processing and machine learning techniques. We propose SPaDe, a Synonym-based Pain level Detection tool to categorize patients into having mild or moderate-to-severe pain to understand diagnosis and treatment methods based on only the pain related expressions in the unstructured chart note. Expressions are subjective, objective, and influenced by cultural background and demography which poses a difficult challenge. Therefore, we improve the model by incorporating the medication information from the structured EMR data and pain scale related information from the chart note to propose an integrated pain level detection tool for OA called PLeDO. With the help of human labeled gold standard data, we demonstrate that both SPaDe and PLeDO can detect mild and moderate-to-severe pain from the EMR data to analyze and potentially improve the quality of care in primary care setting.
摘要:骨關節炎(OA)是一種進行性慢性關節疾病,當受損的關節組織無法正常自我修復時,會導致關節軟骨和骨骼的破壞。這項初步研究的目的是了解來自患者初級保健電子病歷(EMR)的OA疼痛嚴重程度,包括結構化醫療數據和使用信息提取、自然語言處理及機器學習技術的非結構化病歷筆記數據。我們提出了SPaDe,一種基於同義詞的疼痛程度檢測工具,用以將患者分類為輕度或中度至重度疼痛,以便根據非結構化病歷筆記中的疼痛相關表達來理解診斷和治療方法。表達是主觀的、客觀的,並受到文化背景和人口統計的影響,這帶來了困難的挑戰。因此,我們通過整合結構化EMR數據中的用藥信息和病歷筆記中的疼痛量表相關信息來改進模型,提出了一種名為PLeDO的OA綜合疼痛程度檢測工具。在人工標記的金標準數據的幫助下,我們證明SPaDe和PLeDO都能從EMR數據中檢測輕度和中度至重度疼痛,以分析並潛在改善初級保健環境中的護理質量。
Integrating Persuasion Theory into the Epidemiological Modelling of Health Misinformation Spread on Social Media
2608.15689v1 by Mkululi Sikosana, Sean Maudsley-Barton, Oluwaseun Ajao
This study presents a hybrid epidemiological and behavioural framework to simulate the spread of health misinformation on social media. We extend the classical Susceptible--Infected--Recovered (SIR) model to a six-compartment structure (SIRMMM), incorporating Misinformed Susceptible (MS), Misinformed Infected (MI), and Misinformed Recovered (MR) compartments to better reflect the dynamics of the misinformation lifecycle. To account for individual-level behavioural variation, we extend the SIRMMM model by integrating psychological signals from the Elaboration Likelihood Model (ELM), including sentiment polarity, engagement metrics, and cognitive effort, which dynamically modulate the misinformation transmission rate, yielding the ELM-SIRMMM framework. Model parameters were estimated using the FibVID dataset, which captures COVID-19 misinformation on Twitter. Generalisability was tested on two additional datasets: MC-Fake (emotional misinformation) and Monant (general health misinformation). Results show that the ELM-SIRMMM model enhances both predictive accuracy and dynamic realism. On FibVID, it decreases RMSE by 5.5%, delays the misinformation peak from day 150 to day 160, and increases its peak prevalence from 6% to 7%. On MC-Fake, it accurately reproduces a flash-rumour pattern, infecting 38% of users by day 45 and achieving 97% misinformation recovery, all while maintaining model accuracy. In contrast, minimal behavioural signal variability in the Monant dataset leads to marginal benefit, with only a 3% peak and 57% of users remaining susceptible. These findings suggest that structural elaboration alone is insufficient. Functional realism in modelling misinformation spread requires dynamic psychological inputs that vary meaningfully across time and contexts.
摘要:這項研究提出了一個混合流行病學和行為框架,以模擬健康錯誤資訊在社交媒體上的傳播。
我們將傳統的易感--感染--康復(SIR)模型擴展為六個區隔的結構(SIRMMM),納入了錯誤資訊易感者(MS)、錯誤資訊感染者(MI)和錯誤資訊康復者(MR)區隔,以更好地反映錯誤資訊生命週期的動態。
為了考慮個體層面的行為變異,我們通過整合來自精緻可能性模型(ELM)的心理信號來擴展SIRMMM模型,包括情感極性、參與度指標和認知努力,這些信號動態調節錯誤資訊的傳播速率,形成ELM-SIRMMM框架。
模型參數使用FibVID數據集進行估算,該數據集捕捉了Twitter上的COVID-19錯誤資訊。
可推廣性在另外兩個數據集上進行測試:MC-Fake(情感錯誤資訊)和Monant(一般健康錯誤資訊)。
結果顯示,ELM-SIRMMM模型提高了預測準確性和動態真實性。
在FibVID上,它將均方根誤差(RMSE)降低了5.5%,將錯誤資訊的高峰從第150天延遲到第160天,並將其高峰流行率從6%提高到7%。
在MC-Fake上,它準確再現了一種快速謠言模式,到第45天感染了38%的用戶,並實現了97%的錯誤資訊康復,同時保持模型的準確性。
相比之下,Monant數據集中行為信號變異性最小,僅帶來邊際效益,只有3%的高峰和57%的用戶仍然易感。
這些發現表明,僅僅結構上的精緻是不夠的。
在模擬錯誤資訊傳播時,功能真實性需要隨時間和情境有意義變化的動態心理輸入。
From Generalist to Specialist: A Context-Fusion Framework for Endoscopic Polyp Reporting with a Frozen VLM
2608.15580v1 by Ruijie Yang, Yan Zhu, Peiyao Fu, Siyuan Li, Te Luo, Zhihua Wang, Quanlin Li, Pinghong Zhou, Xian Yang, Shuo Wang
Reliable endoscopic polyp reporting requires integrating quantitative lesion sizing, standardized Paris classification, and clinically meaningful morphological description within a single record. General-purpose vision-language models (VLMs) offer a unified interface for image understanding and report generation. Existing specialization strategies, however, typically rely on task-specific models or model-weight adaptation, leaving unresolved how to introduce reliable specialist knowledge while preserving both this unified interface and the VLM's pretrained capabilities. We introduce a context-fusion framework that specializes a frozen general-purpose VLM through both implicit instruction context and explicit transduction context without modifying its pretrained weights. Specifically, a self-supervised polyp encoder retrieves related image-report pairs as explicit, query-specific evidence, while learned continuous specialist tokens provide implicit instruction context shared across cases. Experiments were conducted on 2,056 expert-annotated public endoscopic images. We compared the framework with general-purpose VLMs, task-specific predictors, and weight-adaptation methods to assess specialist performance, unified reporting, and adaptation efficiency. Across numerical, categorical, and report-generation metrics, the proposed framework substantially improved direct frozen-VLM inference and achieved the strongest overall performance among the evaluated methods. It added trainable parameters equal to only 0.006% of the frozen VLM's parameter count. When the top-1 retrieved case carried the correct target category, our framework corrected 70.5% of the errors made by a weight-adaptation baseline. These findings support the context-fusion framework as a lightweight and effective strategy for specialist adaptation of a frozen VLM.
摘要:可靠的內視鏡息肉報告需要將定量病變大小、標準化的巴黎分類和臨床上有意義的形態描述整合在單一記錄中。通用視覺-語言模型(VLMs)提供了一個統一的圖像理解和報告生成界面。然而,現有的專業化策略通常依賴於特定任務的模型或模型權重調整,尚未解決如何在保留這一統一界面和VLM的預訓練能力的同時引入可靠的專家知識。我們提出了一個上下文融合框架,通過隱式指令上下文和顯式轉導上下文專門化一個凍結的通用VLM,而不修改其預訓練權重。具體而言,自監督的息肉編碼器檢索相關的圖像-報告對作為顯式的查詢特定證據,而學習的連續專家標記提供了在案例之間共享的隱式指令上下文。實驗在2,056張專家標註的公共內視鏡圖像上進行。我們將該框架與通用VLMs、特定任務的預測器和權重調整方法進行比較,以評估專家性能、統一報告和適應效率。在數值、類別和報告生成指標上,所提出的框架顯著改善了直接凍結VLM推理,並在評估的方法中實現了最強的整體性能。它增加的可訓練參數僅佔凍結VLM參數總數的0.006%。當檢索到的頂級案例攜帶正確的目標類別時,我們的框架修正了70.5%的權重調整基線所犯的錯誤。這些發現支持上下文融合框架作為一種輕量且有效的策略,用於凍結VLM的專家適應。
EA-LiteUNet: An Edge-Adaptive and Resource-Efficient U-Net for Boundary-Sensitive Dermoscopic Image Segmentation
2608.15537v1 by Wang Jiangtao, Nur Intan Raihana Ruhaiyem, Fu Panpan, Yang Yu, Huang Yan
Accurate boundary delineation remains a persistent challenge in dermoscopic image segmentation because of blurred lesion margins, heterogeneous textures, and complex background artifacts. From a signal-processing perspective, lesion boundaries represent high-frequency components that are highly susceptible to aliasing, noise amplification, and information loss. Consequently, repeated downsampling and feature transformations in conventional convolutional architectures often lead to severely degraded boundary representations. To address these limitations, we propose EA-LiteUNet, an edge-adaptive and computationally efficient U-Net variant specifically designed for boundary-sensitive medical image segmentation. The architecture integrates three core mechanisms: (1) boundary-aware representation learning to suppress aliasing and preserve high-frequency structural details; (2) attention-guided feature modulation to selectively enhance boundary-relevant responses across multi-scale features; and (3) a resource-adaptive inference strategy to dynamically balance segmentation accuracy and computational efficiency. Extensive evaluations across three public dermoscopic datasets demonstrate that EA-LiteUNet consistently achieves superior boundary precision. Specifically, on the ISIC 2018 dataset, the method significantly reduces the 95% Hausdorff Distance (HD95) to 12.89 pixels while maintaining a robust Dice score of 92.08%. Notably, this strong performance is achieved with an ultralightweight configuration of merely 0.29M parameters and 1.17 GFLOPs. Ablation studies further validate the complementary effects of these components, confirming their contribution to enhanced boundary fidelity and stable optimization.
摘要:準確的邊界劃分在皮膚鏡影像分割中仍然是一個持續的挑戰,因為病變邊緣模糊、紋理異質以及複雜的背景伪影。從信號處理的角度來看,病變邊界代表著高頻成分,這些成分對混疊、噪聲放大和信息損失非常敏感。因此,傳統卷積架構中的重複下採樣和特徵轉換往往導致邊界表示的嚴重退化。為了解決這些限制,我們提出了EA-LiteUNet,一種邊緣自適應且計算效率高的U-Net變體,專門設計用於對邊界敏感的醫學影像分割。該架構整合了三個核心機制:(1)邊界感知的表示學習,以抑制混疊並保留高頻結構細節;(2)注意力引導的特徵調制,以選擇性地增強多尺度特徵中的邊界相關響應;以及(3)資源自適應的推理策略,以動態平衡分割準確性和計算效率。在三個公共皮膚鏡數據集上的廣泛評估顯示,EA-LiteUNet始終實現了卓越的邊界精度。具體而言,在ISIC 2018數據集上,該方法將95%豪斯多夫距離(HD95)顯著降低至12.89像素,同時保持穩健的Dice得分92.08%。值得注意的是,這一強勁的表現是在僅有0.29M參數和1.17 GFLOPs的超輕量配置下實現的。消融研究進一步驗證了這些組件的互補效果,確認它們對增強邊界忠實度和穩定優化的貢獻。
Gated Against One Model, Open to the Next: Option-Only Solvability in Legal Multiple-Choice Benchmarks
2608.15428v1 by Volodymyr Ovcharov
Multiple-choice benchmarks are graded on whether a model picks the right option, not on whether it needed the question. Measuring that gap takes care: a model answering A to most items scores above chance wherever the key sits at A, and reads as recognition when it is not. We measure it on UA-JudgeExam: 11,990 four-option items with official keys, published by Ukraine's Higher Qualification Commission of Judges. Shown the options and no question, Claude Haiku 4.5 scores 0.383 against chance, and the leak is concentrated: 11.8% of items are answered blind on all eight option orders, against 0.2 items expected by chance. It is not quotation: search over 280,059 editions of Ukrainian legislation recovers 0.128. Gating those out retains 8,128 items, on which the gating model itself now scores 0.204, and GPT-5.6, which took no part in the selection, still answers 0.515 of them with the question hidden. Scoring twelve held-out models on the whole set and subtracting each one's answer-position habit, only two keep an excess: GPT-5.6 at +0.265, Sonnet 4.6 at +0.081. Without it the ranking misleads: Llama 3.1 8B scores 0.292 blind, above every model but those two, purely by answering A to 92% of items. The gate does select something real: on the items it rejected, eleven of twelve models score 0.518-0.789, every interval clear of what the same model scores on the items it kept. But that signal is one model's, and filtering on it does not transfer upward. Neither is visible on a 400-item sample, where nine models read as "statistically at chance". Rewriting distractors instead overshoots to 0.168, below chance and as exploitable. The same probe on LEXam returns chance: every option there points into the stem, none longer than 33 characters. Item format decides whether the problem can arise; capability decides how much is extracted. We release the corpus, the predictions and the harness.
摘要:多選基準是根據模型是否選擇正確選項來評分,而不是根據它是否需要問題。測量這個差距需要小心:一個模型在大多數項目中回答A,無論關鍵在A的位置如何,都會得分高於隨機,而當關鍵不在A時則顯示為識別。我們在UA-JudgeExam上進行測量:11,990個四選項目,擁有官方答案,由烏克蘭高級法官資格委員會發布。
當顯示選項而沒有問題時,Claude Haiku 4.5的得分為0.383,超過隨機,而洩漏集中在一起:11.8%的項目在所有八個選項順序中都是盲目回答,預期隨機回答為0.2項目。這不是引用:搜索超過280,059份烏克蘭立法的版本恢復了0.128。排除這些後保留了8,128個項目, gating模型本身現在在這些項目上得分為0.204,而GPT-5.6沒有參與選擇,仍然在問題隱藏的情況下回答了其中的0.515。對整個數據集進行十二個保留模型的得分並減去每個模型的答案位置習慣,只有兩個保持超額:GPT-5.6為+0.265,Sonnet 4.6為+0.081。沒有這個,排名會誤導:Llama 3.1 8B在盲測中得分0.292,超過每個模型,但只有這兩個,純粹是因為對92%的項目回答A。
這個gating確實選擇了一些真實的東西:在被拒絕的項目中,十二個模型中的十一個得分為0.518-0.789,每個區間都清楚地與同一模型在保留項目上的得分不同。但那個信號是某一模型的,基於它的過濾並不會向上轉移。在400項樣本中也不可見,九個模型的表現為“統計上隨機”。重寫干擾項反而超出到0.168,低於隨機且可被利用。對LEXam的相同探測返回隨機:那裡的每個選項都指向題幹,沒有一個超過33個字符。項目格式決定問題是否會出現;能力決定提取的多少。我們釋放語料庫、預測和工具。
ETHOS: Towards a Modular Ethics Framework for Clinical Multi-Agent Systems
2608.15424v1 by Rakesh Sharma, Sydney Pugh, Cameron Beeche, Pankhuri Singhal, Rachel Wu, Margaret Eby, Jeffrey Duda, James Gee, Kyra O'Brien, Hersh Sagreiya, Marina Serper, Victoria Gershuni, Angela Bradbury, Anurag Verma, Eric Eaton, Kevin B. Johnson, Walter Witschey
The rapid adoption of large language models has enabled the development of clinical multi-agent systems (MAS) capable of integrating multimodal patient data and supporting increasingly complex clinical decision-making. However, the deployment of these systems in real-world healthcare settings raises critical ethical concerns related to safety, fairness, accountability, transparency, and patient trust. While numerous organizations, including the World Health Organization, the National Academy of Medicine, and the FUTURE-AI consortium, have proposed ethical frameworks and governance principles for healthcare AI, these efforts remain largely conceptual. To address this challenge, we present ETHOS (Ethics and Trust through Hierarchical Oversight System), a modular ethics framework designed as a governance meta-agent that can be integrated with any existing multi-agent system without requiring changes to its underlying architecture. ETHOS translates stakeholder-informed ethical requirements into executable runtime oversight through a layered governance approach consisting of deterministic checks, contextual reviews, and a final ethics critic. These components continuously evaluate intermediate reasoning steps and final outputs, enabling the system to identify ethical risks, request revisions, or suppress responses that fail predefined safety and trustworthiness criteria. We demonstrate ETHOS within a hepatology clinical decision-support MAS. Results show that ETHOS improves decision reliability by detecting incomplete, inconsistent, or out-of-scope evidence and appropriately increasing abstention when safe recommendations cannot be supported. By embedding ethical governance directly into system operation, ETHOS provides a practical and auditable mechanism for transforming high-level AI ethics principles into deployable safeguards.
摘要:大型語言模型的快速採用使得臨床多代理系統(MAS)的發展成為可能,這些系統能夠整合多模態病人數據並支持日益複雜的臨床決策。
然而,這些系統在現實世界醫療環境中的部署引發了與安全、公平、問責、透明度和病人信任相關的重大倫理問題。
儘管包括世界衛生組織、國家醫學院和FUTURE-AI聯盟在內的許多組織已經提出了針對醫療AI的倫理框架和治理原則,但這些努力仍然主要是概念性的。
為了解決這一挑戰,我們提出了ETHOS(通過分層監督系統實現倫理與信任),這是一個模塊化的倫理框架,設計為一個治理元代理,可以與任何現有的多代理系統集成,而無需改變其底層架構。
ETHOS將利益相關者所知的倫理要求轉化為可執行的運行時監督,通過一種分層治理方法,包括確定性檢查、上下文審查和最終倫理評估。
這些組件持續評估中間推理步驟和最終輸出,使系統能夠識別倫理風險、請求修訂或抑制不符合預定安全和可信標準的回應。
我們在一個肝病臨床決策支持MAS中展示了ETHOS。
結果顯示,ETHOS通過檢測不完整、不一致或超出範疇的證據來提高決策的可靠性,並在無法支持安全建議時適當地增加放棄。
通過將倫理治理直接嵌入系統運作中,ETHOS提供了一種實用且可審計的機制,將高層次的AI倫理原則轉化為可部署的保障措施。
Invariant Pretraining for Robust Code Representations
2608.15412v1 by Yifeng He, Yundi Xu, Christopher Castro Gaw Gonzalo, Zili Wang, Hao Chen
Encoder-based code representation models remain widely deployed for discriminative tasks such as clone detection and code classification, where their small size and low inference cost are decisive. Their robustness, however, is fragile: under invariant programs, semantically equivalent code written in different syntactic forms, learned representations degrade substantially even though program behavior is unchanged. We present an empirical study of this robustness gap across four encoder baselines, two downstream tasks, and four datasets, together with a minimal code-only continued pretraining recipe that closes much of it. Our method, invariant pretraining (InvPT), applies semantic-preserving transformations to the corpus and combines masked language modeling with multi-positive supervised contrastive learning that treats all augmentations of the same source function as positives, mixing self-contrast pairs (same code, different masks) with invariant-contrast pairs (transformed code) for positives of varying difficulty. Unlike prior contrastive code encoders, InvPT does not require paired natural-language data. Across our evaluation, InvPT improves robustness on transformed test sets by up to 11 percentage points on clone detection and 19 on code classification while matching or improving standard accuracy, and our ablations isolate multi-positive invariant contrast as the main source of the gains. Our aim is not a new objective but a careful measurement of where encoder robustness breaks and how far a simple, code-only recipe can recover it.
摘要:編碼器基礎的代碼表示模型仍然廣泛應用於克隆檢測和代碼分類等判別性任務,其中其小巧的尺寸和低推理成本是決定性的。然而,它們的穩健性卻是脆弱的:在不變的程序下,以不同語法形式編寫的語義等價代碼,其學習到的表示會大幅退化,即使程序行為未改變。我們針對四個編碼器基準、兩個下游任務和四個數據集進行了這一穩健性差距的實證研究,並提出了一種最小的僅代碼持續預訓練配方,能夠縮小這一差距。我們的方法,稱為不變預訓練(InvPT),對語料庫應用語義保留轉換,並將掩碼語言建模與多正樣本監督對比學習相結合,將同一源函數的所有增強視為正樣本,將自對比對(相同代碼,不同掩碼)與不變對比對(轉換代碼)混合,形成不同難度的正樣本。與之前的對比代碼編碼器不同,InvPT 不需要配對的自然語言數據。在我們的評估中,InvPT 在轉換測試集上提升了克隆檢測的穩健性達 11 個百分點,代碼分類達 19 個百分點,同時保持或提高標準準確性,而我們的消融實驗則將多正樣本不變對比確定為增益的主要來源。我們的目標不是一個新的目標,而是仔細測量編碼器穩健性破裂的地方,以及一個簡單的僅代碼配方能夠恢復的程度。
Grounding Healthcare LLMs in a Causal Knowledge Graph: Framework, Metrics, and a Cardiovascular Pilot
2608.15382v1 by Ummara Mumtaz, Aimen Noor, Awais Ahmed
Large language models (LLMs) are increasingly proposed for healthcare decision support, but their evaluations still reward single-answer accuracy rather than reasoning about interventions, mechanisms, harms, evidence, and uncertainty. We propose a reproducible, graph-centered evaluation framework for intervention-oriented LLM behavior in healthcare and stress-test it in a cardiovascular pilot. The framework has four components: (i) a domain causal knowledge graph in which assertions are first-class, provenance-preserving nodes with stable identifiers; (ii) a scenario-conditioned subgraph extraction step that, given any clinical scenario, retrieves the relevant reified-assertion subgraph; (iii) four controlled grounding conditions that vary how the retrieved subgraph is composed into the model's context (ungrounded C1, knowledge-graph C2, causal-graph C3, integrated C4); and (iv) an automated scoring pipeline, anchored on assertion identifiers, that computes intervention accuracy, and other evaluation measures on a single pass. To test the framework, we built a category-balanced scenario generator across eight reasoning failure modes and instantiated it on a cardiovascular graph. The metric panel discriminates conditions along interpretable, non-redundant axes: C4 obtains the strongest causal edge F1 (0.838), adverse-effect F1 (0.833), evidence accuracy (0.738), and unsupported claim rate (0.114), while C1 obtains the highest raw intervention accuracy (0.948) with no measurable causal or evidential grounding.
摘要:大型語言模型(LLMs)越來越多地被提議用於醫療決策支持,但其評估仍然獎勵單一答案的準確性,而不是對干預、機制、危害、證據和不確定性進行推理。我們提出了一個可重複的、以圖為中心的評估框架,用於醫療保健中的干預導向LLM行為,並在心血管試點中進行壓力測試。該框架有四個組成部分:(i)一個領域因果知識圖,其中斷言是第一類的、保持來源的節點,具有穩定的標識符;(ii)一個情境條件的子圖提取步驟,根據任何臨床情境檢索相關的具體化斷言子圖;(iii)四個控制的基礎條件,變化檢索到的子圖如何組成模型的上下文(未基礎的C1、知識圖C2、因果圖C3、整合的C4);以及(iv)一個自動評分管道,以斷言標識符為基礎,計算干預準確性和其他評估指標,僅需一次通過。為了測試該框架,我們建立了一個跨越八種推理失敗模式的類別平衡情境生成器,並在心血管圖上實現了它。該指標面板沿著可解釋的、非冗餘的軸區分條件:C4獲得最強的因果邊緣F1(0.838)、不良影響F1(0.833)、證據準確性(0.738)和不支持的主張率(0.114),而C1獲得最高的原始干預準確性(0.948),卻沒有可測量的因果或證據基礎。
When AI Rewrites, Classifiers Relax: Uncertainty-Aware Sentiment Analysis on Sarcastic and AI-Paraphrased Social Text
2608.15338v1 by Shresth Shroff
Sentiment classifiers are increasingly applied to social media content that is either sarcastic or AI-generated --- two distributional regimes where standard evaluations offer little guidance. We present a three-part empirical study of sentiment classifier behaviour under these conditions. First, we find that confidence scores on sarcastic text are significantly lower than on non-sarcastic text (Mann--Whitney $p = 2 \times 10^{-6}$), confirming that classifiers sense their own uncertainty on ironic content even without explicit uncertainty modelling. Second, and counterintuitively, we show that sentiment classifiers achieve higher accuracy on AI-paraphrased reviews than on the original human-authored text (RoBERTa: $+5.8$ pp for Qwen3.5-4B paraphrases, $+3.7$ pp for Gemma4-E4B), revealing a cross-domain stylistic alignment effect: AI paraphrases remove distributional noise that confounds Twitter-trained classifiers, producing cleaner, more prototypical sentiment text. Third, we demonstrate that a lightweight abstention wrapper --- flagging the $14\%$ of inputs with confidence below $0.6$ --- improves accuracy from 82.2\% to 88.9\% ($+6.7$ pp) on the retained set. We further compare Semantic Entropy and MC-Dropout-style disagreement as uncertainty signals and find near-identical AUROC ($0.650$ vs.\ $0.646$) on sarcastic text, suggesting that for short social media inputs, both methods are interchangeable. Our results motivate a shift from confident single-label prediction to uncertainty-aware abstention in high-stakes sentiment applications such as mental health flagging and content moderation.
摘要:情感分類器越來越多地應用於諷刺或 AI 生成的社交媒體內容——這兩種分佈模式下,標準評估提供的指導有限。我們提出了一項三部分的實證研究,探討情感分類器在這些條件下的行為。首先,我們發現對於諷刺文本的信心分數顯著低於非諷刺文本(Mann--Whitney $p = 2 \times 10^{-6}$),確認分類器即使在沒有明確不確定性建模的情況下,也能感知到對於諷刺內容的自身不確定性。其次,反直覺的是,我們顯示情感分類器在 AI 改寫的評論上取得的準確率高於原始的人類撰寫文本(RoBERTa: $+5.8$ pp 對於 Qwen3.5-4B 改寫,$+3.7$ pp 對於 Gemma4-E4B),揭示了一種跨領域的風格一致性效應:AI 改寫去除了困擾 Twitter 訓練的分類器的分佈噪音,產生了更乾淨、更原型的情感文本。第三,我們證明了一種輕量級的棄權包裝——標記信心低於 $0.6$ 的 $14\%$ 輸入——在保留數據集上將準確率從 82.2\% 提高到 88.9\%($+6.7$ pp)。我們進一步比較語義熵和 MC-Dropout 風格的不一致作為不確定性信號,發現對於諷刺文本,兩者的 AUROC 幾乎相同($0.650$ 對 $0.646$),這表明對於短的社交媒體輸入,這兩種方法是可以互換的。我們的結果促使從自信的單標籤預測轉向在高風險情感應用中,如心理健康標記和內容審核,意識到不確定性的棄權。
Physiological World Models for Human State Transitions
2608.15309v1 by Chongyang Zhang, Rendong Wang, Hao Zheng, Hanwen Zhang, Yang Liu, Xiaolong Wei, Bin Chong
Continuous multimodal sensing now allows human physiology to be observed throughout daily life rather than only during occasional clinical visits. However, most health artificial intelligence systems are designed to recognize current states, estimate risks or analyse individual biomarkers. They do not directly model how physiological states change in response to real-world events, behaviours, contexts and interventions. Here we propose the Physiological World Model (PWM), an event-conditioned framework for learning these changes at the level of the whole person. We introduce the HumanState Transition Token, a structured, quality-scored unit that connects the physiological state before an event with the event or action, relevant context and intervention information, the physiological trajectory after the event, observed outcomes and data quality. We describe four capability levels, from state representation to bounded intervention planning, together with four data acquisition and validation protocols. We also propose six benchmark tasks covering HumanState representation, forecasting across multiple timescales, individualized response prediction, simulation of alternative interventions, bounded planning and reliability under distribution shift. Together, this framework provides a practical path towards personalized health management, behavioural intervention design and clinician-supervised decision support, while clearly separating prediction from causal inference and making uncertainty, safety, governance and limits of use explicit.
摘要:持續的多模態感測現在允許在日常生活中觀察人類生理,而不僅僅是在偶爾的臨床訪問中。
然而,大多數健康人工智慧系統旨在識別當前狀態、評估風險或分析個體生物標記。
它們並未直接建模生理狀態如何對現實世界事件、行為、情境和干預進行變化。
在這裡,我們提出生理世界模型(PWM),這是一個事件條件框架,用於學習整個人的這些變化。
我們介紹了人類狀態轉換標記,這是一個結構化的、質量評分的單位,將事件前的生理狀態與事件或行動、相關情境和干預信息、事件後的生理軌跡、觀察到的結果和數據質量相連接。
我們描述了四個能力層級,從狀態表示到有界的干預規劃,以及四個數據獲取和驗證協議。
我們還提出了六個基準任務,涵蓋人類狀態表示、跨多個時間尺度的預測、個性化反應預測、替代干預的模擬、有界規劃和在分佈轉移下的可靠性。
總體而言,這個框架提供了一條實用的途徑,朝向個性化健康管理、行為干預設計和臨床醫生監督的決策支持,同時明確區分預測與因果推斷,並使不確定性、安全性、治理和使用限制變得明確。
Demographic Injection in Medical Language Models under Diversity, Equity, and Inclusion Prompts
2608.15254v1 by Diego Mardian, Frank Liu
Clinical-AI guidance increasingly recommends prompting language models to reason with attention to diversity, equity, and inclusion (DEI). We measure a side effect that misrepresents patients: a one-sentence DEI prompt appended to a medical question leads models to add patient demographic attributes (race, socioeconomic status, sex) the question never stated, in effect rewriting who the patient is. We call this demographic injection. Across 47 models, four medical benchmarks, and 376,000 responses scored by a validated model-judge pipeline, a single DEI prompt raises the injection rate from 0.7% to 33.1% (47x) in all 47 of 47 models, attributable to the equity content rather than to added length (18x above a length-matched control; p=1.4x10^-14). Most added content is a general population statement that leaves the answer unchanged, but a smaller subset attaches an attribute to the specific patient or changes the selected option (0.25-2.4% of responses, 99.8% toward the incorrect option), where the invented demographic changes the answer the model recommends. Phrasing scales the effect from 14% to 56%. DEI prompts are just one example of a more general mechanism. Any instruction that nudges how a model reasons can make it add unrequested details, including details about the patient. Flagged outputs are treated as model errors under study, not clinical guidance.
摘要:臨床人工智慧指導越來越多地建議促使語言模型在考慮多樣性、公平性和包容性(DEI)時進行推理。我們測量了一種誤導患者的副作用:一個附加在醫療問題上的單句DEI提示會導致模型添加問題中從未提到的患者人口統計屬性(種族、社會經濟地位、性別),實際上重寫了患者的身份。我們稱之為人口統計注入。在47個模型、四個醫療基準和376,000個由經過驗證的模型評判管道評分的回應中,單一的DEI提示使得注入率從0.7%上升到33.1%(47倍),這是由於公平性內容而非增加的長度(在長度匹配的對照組中增加了18倍;p=1.4x10^-14)。大多數新增內容是一般人口的陳述,對答案沒有改變,但一小部分則將屬性附加到特定患者或改變所選選項(0.25-2.4%的回應,99.8%朝向不正確的選項),其中虛構的人口統計改變了模型推薦的答案。措辭將效果擴大至14%至56%。DEI提示僅是更一般機制的一個例子。任何促使模型推理的指令都可能使其添加未請求的細節,包括有關患者的細節。被標記的輸出被視為正在研究的模型錯誤,而非臨床指導。
Translating finite-domain integer constraint models to CP/SMT/ILP/PB/SAT solvers with CPMpy
2608.15143v1 by Tias Guns, Ignace Bleukx, Hendrik Bierlee, Jo Devriendt, Emilio Gamba, Orestis Lomis, Wout Piessens, Thomas Sergeys, Dimos Tsouros, Wout Vanroose, Hélène Verhaeghe
Constraint solving is a declarative approach for solving combinatorial satisfaction and optimization problems. The user specifies their problem through constraints and decision variables, and a generic solver is used to find a solution. Several constraint-solving technologies exist, and certain solvers perform well on certain problems. Therefore, it is useful to try different solvers given a particular application. However, each solving paradigm supports different types of constraints and decision variables. Our goal is to translate high-level constraint satisfaction and optimization problems into any lower-level formalism, including CP, SMT QF-LIA, ILP, PB and (Max)SAT. This allows for comparing different solving technologies for a particular problem, without requiring a user to manually remodel it for each solving paradigm. We define a high-level language of logical and arithmetic operations, and useful additional functions and constraints, which are known as global constraints in the CP community. We then present a modular framework for transforming our high-level modeling language to CP/SMT/ILP/PB and (Max)SAT solvers. While many transformations are partly described in the literature, we observe that they can be implemented through a modular waterfall of smaller components, where lower-level paradigms reuse the transformations of higher-level paradigms. Two recurring challenges are handling the negation of arbitrary subexpressions and avoiding the introduction of auxiliary variables. Additionally, we take special care linearizing non-linear operators for ILP, PB and SAT-solvers. The transformation waterfall is implemented and evaluated in the open-source CPMpy library. Our results show that constraint models significantly change throughout the transformations, and that optimizations to the linearization of constraints are essential for ILP and PB solvers.
摘要:限制求解是一種聲明式方法,用於解決組合滿足和優化問題。用戶通過約束和決策變量來指定他們的問題,並使用通用求解器來尋找解決方案。存在幾種限制求解技術,某些求解器在特定問題上表現良好。因此,針對特定應用嘗試不同的求解器是有用的。然而,每種求解範式支持不同類型的約束和決策變量。
我們的目標是將高級約束滿足和優化問題轉換為任何低級形式,包括 CP、SMT QF-LIA、ILP、PB 和 (Max)SAT。這使得在特定問題上比較不同的求解技術成為可能,而無需用戶手動為每個求解範式重新建模。
我們定義了一種邏輯和算術運算的高級語言,以及一些有用的附加函數和約束,這些在 CP 社區中被稱為全局約束。我們接著提出了一個模塊化框架,用於將我們的高級建模語言轉換為 CP/SMT/ILP/PB 和 (Max)SAT 求解器。雖然許多轉換在文獻中部分描述,但我們觀察到它們可以通過一個模塊化的瀑布式小組件來實現,其中低級範式重用高級範式的轉換。兩個反覆出現的挑戰是處理任意子表達式的否定和避免引入輔助變量。此外,我們特別注意將非線性運算符線性化以適應 ILP、PB 和 SAT 求解器。
轉換瀑布在開源的 CPMpy 庫中實現和評估。我們的結果顯示,約束模型在轉換過程中顯著改變,並且對約束線性化的優化對於 ILP 和 PB 求解器至關重要。
FZ-VLM: A Two Stage Florence-Zephyr Vision Language Model Framework for Pulmonary Nodule Characterization and Clinical Decision Making
2608.15004v1 by Pramit Dutta, Jenita Manokaran, Richa Mittal, Ryan Appleby, Eranga Ukwatta
Lung cancer remains one of the leading causes of cancer-related mortality worldwide, and Computed Tomography (CT) is a primary imaging tool for screening and followup assessment. After pulmonary nodule detection, radiologists manually assess anatomical location, diameter, margin characteristics, and attenuation type to support risk assessment and clinical decision-making. However, this post-detection workflow is time-consuming and can be affected by inter-observer variability. Existing Artificial Intelligence methods often focus on isolated tasks, limiting their use as a unified, clinically grounded interpretation framework. This study presents FZ-VLM, a two-stage Florence-Zephyr Vision Language Model framework for unified structured pulmonary nodule characterization in lung CT. The framework uses a fine-tuned Florence-2 model to extract radiological attributes from expert-annotated 2D axial CT slices, while a Zephyr-7B model uses these attributes to generate nodule descriptions, follow-up recommendations, and longitudinal analyses. Results showed that the Stage 1 model achieved 77.18\% accuracy for anatomical location, 67.96\% accuracy for margin characteristics, and 79.13\% accuracy for attenuation type, with a Mean Absolute Error of 2.58 mm for diameter estimation, outperforming evaluated GPT-4-based baselines as well as the human baseline. Expert radiologist evaluation of Stage 2 showed 93.9\% accuracy, 98.6\% completeness score, 76.1\% clinical relevance, and an overall score of 89.5\%. Safety analysis showed that most outputs were clinically safe, although some follow-up recommendations still required expert review. To the best of our knowledge, this study presents the first two-stage Vision-Language Model framework for structured nodule characterization and clinical decision-making.
摘要:肺癌仍然是全球癌症相關死亡的主要原因之一,而計算機斷層掃描(CT)是篩查和後續評估的主要影像工具。在檢測到肺結節後,放射科醫生手動評估解剖位置、直徑、邊緣特徵和衰減類型,以支持風險評估和臨床決策。然而,這一檢測後的工作流程耗時且可能受到觀察者之間變異性的影響。現有的人工智慧方法通常專注於孤立的任務,限制了它們作為統一的、臨床基礎的解釋框架的使用。本研究提出了FZ-VLM,一個兩階段的Florence-Zephyr視覺語言模型框架,用於肺CT中統一的結構化肺結節特徵描述。該框架使用微調的Florence-2模型從專家標註的2D軸向CT切片中提取放射學屬性,而Zephyr-7B模型則利用這些屬性生成結節描述、後續建議和縱向分析。結果顯示,第一階段模型在解剖位置的準確率達到77.18\%、邊緣特徵的準確率為67.96\%、衰減類型的準確率為79.13\%,直徑估計的平均絕對誤差為2.58毫米,超越了評估的基於GPT-4的基準以及人類基準。第二階段的專家放射科醫生評估顯示準確率為93.9\%、完整性得分為98.6\%、臨床相關性為76.1\%,總體得分為89.5\%。安全性分析顯示,大多數輸出在臨床上是安全的,儘管一些後續建議仍需專家審查。據我們所知,本研究提出了首個兩階段的視覺-語言模型框架,用於結構化結節特徵描述和臨床決策。
Evaluating Agentic Code Repair Capabilities in Distributed Systems
2608.14863v1 by Yibo Yan, Huijuan Wang, Junzhou He, Yizhuo Liang, Shaoyu Wang, Huanchen Sun, Seo Jin Park
LLM-based coding agents have advanced rapidly on single-process SWE tasks, with frontier models now clustering in the high-70s on SWE-bench Verified. Distributed-system debugging, however, remains an under-explored regime: bugs span processes, nodes, and protocol interactions, with root causes rarely recoverable from source alone and brute-force exploration intractable across non-deterministic interleavings. This leaves two gaps in LLM and agent evaluation: no code-repair benchmark targets distributed-system bugs, and no controlled study isolates how much externally provided debugging context changes agent success on them. We introduce DDBench, a code-repair benchmark of 60 historical bugs mined from 13 open-source distributed systems, partitioned into three difficulty tiers. DDBench evaluates every case under two matched conditions: a symptom-only condition where the agent receives only the bug symptom and repository, and a context-augmented condition where it additionally receives a bounded debugging context (logs, traces, runtime state, and targeted code-investigation notes), isolating the effect of debugging context from model capability. The evaluation of ten LLMs on DDBench reveals several findings. First, distributed debugging exercises a reasoning dimension that single-process benchmarks do not surface: models' pass rates span 61 pp, and pairwise bootstrap separates 9 of 15 top-tier model pairs at p < 0.05 on DDBench's hardest case-set. Second, bounded debugging context lifts aggregate pass rate by +18.1 pp, and the lift is asymmetric: weaker models gain pass rate, while stronger models gain efficiency. Third, debugging context requires careful curation, as even faithful debugging context can sometimes mislead LLMs.
摘要:LLM 基礎的編碼代理在單一過程的軟體工程任務上迅速進步,前沿模型在 SWE-bench Verified 上的表現已聚集在高 70 分以上。 然而,分散式系統的除錯仍然是一個未被充分探索的領域:錯誤跨越過程、節點和協議互動,根本原因很少能僅從源代碼中恢復,且在非確定性交錯中進行暴力探索是不可行的。 這在 LLM 和代理評估中留下了兩個空白:沒有代碼修復基準針對分散式系統的錯誤,且沒有控制研究來隔離外部提供的除錯上下文如何改變代理在這些錯誤上的成功率。 我們介紹 DDBench,一個由 13 個開源分散式系統挖掘的 60 個歷史錯誤組成的代碼修復基準,分為三個難度層級。
DDBench 在兩個匹配條件下評估每個案例:一個僅有症狀的條件,代理僅接收錯誤症狀和代碼庫,另一個是增強上下文的條件,代理還接收有限的除錯上下文(日誌、追蹤、運行時狀態和針對代碼調查的筆記),以隔離除錯上下文對模型能力的影響。 對十個 LLM 在 DDBench 上的評估揭示了幾個發現。 首先,分散式除錯運用了一個單一過程基準未顯現的推理維度:模型的通過率跨度為 61 個百分點,配對自助法在 DDBench 最困難的案例集中以 p < 0.05 分離了 15 個頂級模型對中的 9 個。 其次,有限的除錯上下文將總體通過率提升了 +18.1 個百分點,且這一提升是非對稱的:較弱的模型通過率提高,而較強的模型則提高了效率。 第三,除錯上下文需要仔細策劃,因為即使是忠實的除錯上下文有時也會誤導 LLM。
Generated Context versus Governed State: Functional Conditions for Accountable Longitudinal Clinical Reasoning
2608.14804v1 by Augusto Bernardo Pissarra, Victor Lorena de Farias Souza
Large language models (LLMs) have become the dominant interface of clinical artificial intelligence, yet the interface they expose (text in, text out, one context window at a time) maintains no explicit, persistent, governed representation of what is currently true about a patient. This paper argues that longitudinal clinical reasoning is a state-estimation problem under partial observability, and that the axis on which clinical AI succeeds or fails is not the fluency of the model reading the record but the governance of the patient state it reasons over. We distinguish generated context from governed state; separate five objects that clinical AI habitually conflates (true state, observations, evidence, belief, and simulated state); define a tiered governance standard against which any clinical AI system can be audited; and show that an operational definition of accountability decomposes into four information requirements: an immutable evidence ledger with awareness-time versioning, a belief state distinct from accumulated evidence, an observation-process model, and claim-level causal typing. We are explicit that this decomposition is analytic rather than a necessity theorem, and that its value is conceptual hygiene: it converts "accountable clinical AI" from a slogan into an audit instrument. A six-level maturity framework separates what a system makes governable from what it can compute, locating current LLM-centric practice at high capability but low maturity. The paper is fully self-contained: the four research questions the framework poses are stated in the introduction, and the conclusion records what the paper establishes toward each; future work develops the buildable core of the architecture and the research program toward full Clinical World Models. No empirical result is claimed here.
摘要:大型語言模型(LLMs)已成為臨床人工智慧的主導介面,但它們所暴露的介面(文本輸入、文本輸出、一次一個上下文窗口)並未對目前有關患者的真實情況提供明確、持久、受管控的表徵。本文主張,縱向臨床推理是一個在部分可觀察性下的狀態估計問題,而臨床 AI 成功或失敗的軸心不在於模型閱讀記錄的流暢性,而在於它所推理的患者狀態的治理。我們區分生成的上下文與受管控的狀態;將臨床 AI 通常混淆的五個對象(真實狀態、觀察、證據、信念和模擬狀態)分開;定義一個分層治理標準,以便對任何臨床 AI 系統進行審計;並顯示一個運作性責任的定義可分解為四個信息要求:具有意識時間版本控制的不可變證據賬本、與累積證據不同的信念狀態、觀察過程模型,以及索賠級別的因果類型。我們明確指出這一分解是分析性的,而非必要定理,其價值在於概念衛生:它將“可負責任的臨床 AI”從口號轉變為審計工具。一個六級成熟度框架將系統可治理的部分與其可計算的部分分開,將當前以 LLM 為中心的實踐定位於高能力但低成熟度。本文是完全自足的:框架提出的四個研究問題在引言中陳述,結論記錄了本文在每個問題上所建立的內容;未來的工作將發展可構建的架構核心及通向完整臨床世界模型的研究計劃。此處不聲稱任何實證結果。
Prompting is not enough: supervised baselines and leakage control for measuring shared decision-making with LLMs in pediatric encounters
2608.14792v1 by Bernardo Modenesi, Jody Lin, Kimberly Kaphingst, Angela Zhu, Maya Wheeler, Peilu Zhang, Angela Fagerlin
Objectives: To determine whether zero-shot prompting of a large language model (LLM) is sufficient to detect shared decision-making (SDM) behaviors in real clinical encounters, and whether supervised learning adds value under patient-grouped, nested evaluation. Methods: We analyzed 21 audio-recorded outpatient surgical decision encounters (19 unique patients; 7,566 utterance segments; ~6.1 hours) between families of children with multiple long-term conditions and their surgical providers. Trained coders labeled segments for 12 SDM behaviors (human-human macro Cohen's kappa = 0.695). We compared a zero-shot local LLM (Qwen 2.5 32B), a supervised classifier over frozen sentence embeddings, and their logistic stack, under patient-grouped outer folds with inner cross-fitted thresholds and patient-resampled confidence intervals. Results: The zero-shot LLM reached macro kappa = 0.139 (95% CI 0.111-0.164). The supervised classifier reached kappa = 0.227 (0.186-0.262), a paired improvement of 0.088 (0.051-0.119). A logistic stack of the two reached kappa = 0.242 (0.198-0.284). We identified multiple corpus-specific leakage paths, including grouping sibling recordings separately and allowing labels from an outer held-out patient to enter few-shot exemplars used while fitting downstream models. Conclusion: Zero-shot prompting alone is not sufficient to measure SDM behavior as reliably as a small supervised model, and patient-level grouping alone does not prevent leakage when labeled prompt exemplars are precomputed outside the outer evaluation loop. Reported performance is sensitive to the unit of data splitting and to where labeled exemplars enter the pipeline. External validation is needed before these findings generalize beyond this population, model, prompt, and codebook.
摘要:目標:確定大型語言模型(LLM)的零-shot 提示是否足以在真實臨床接觸中檢測共享決策(SDM)行為,以及監督學習在患者分組的嵌套評估下是否增值。
方法:我們分析了21段音頻錄製的門診外科決策接觸(19名獨特患者;7,566個發言片段;約6.1小時),這些接觸發生在多種長期疾病兒童的家庭與其外科提供者之間。受過訓練的編碼員為12種SDM行為標記片段(人與人之間的宏觀Cohen's kappa = 0.695)。我們比較了一個零-shot本地LLM(Qwen 2.5 32B)、一個基於凍結句子嵌入的監督分類器,以及它們的邏輯回歸堆疊,在患者分組的外部折疊下,使用內部交叉擬合的閾值和患者重抽樣的置信區間。
結果:零-shot LLM達到宏觀kappa = 0.139(95% CI 0.111-0.164)。監督分類器達到kappa = 0.227(0.186-0.262),配對改善為0.088(0.051-0.119)。兩者的邏輯回歸堆疊達到kappa = 0.242(0.198-0.284)。我們識別了多條特定語料的洩漏路徑,包括將兄弟姐妹的錄音分開分組,以及允許來自外部保留患者的標籤進入在擬合下游模型時使用的少量示例。
結論:僅依賴零-shot 提示不足以像小型監督模型那樣可靠地測量SDM行為,且僅進行患者層級分組並不能防止當標記的提示示例在外部評估循環之外預先計算時的洩漏。報告的性能對數據拆分的單位和標記示例進入管道的位置敏感。在這些發現能夠超出這一人群、模型、提示和代碼本進行推廣之前,需要進行外部驗證。
CEDAR-GRPO: Process-Aware Reinforcement Learning for General Abductive Reasoning in LLMs
2608.14791v1 by Moein Salimi, Danial Parnian, Shaygan Adim, Amirmohammad Ebrahiminasab, Nima Alighardashi, Parsa Gholami, Sahand Akramipour, Mahdi Jafari Siavoshani, Mohammad Hossein Rohban
Abductive reasoning, often characterized as inference to the best explanation, is central to explanation under uncertainty, from everyday sense-making and investigation to scientific discovery. Yet LLM research has mostly studied abduction through narrow, task-specific benchmarks, making it unclear whether observed gains transfer beyond the benchmark family used for training or evaluation. We ask whether RL post-training can improve abduction as a transferable reasoning capability. We introduce CEDAR-GRPO, a process-aware framework that combines final-answer correctness with abductive rewards for evidence coverage and evidence-to-explanation directionality. Four open-weight LLMs are post-trained on a controlled, domain-neutral mixture of abductive hypothesis-generation and hypothesis-selection tasks. We evaluate them on 11 unseen tasks spanning hypothesis selection, missing-fact generation, defeasible inference, long-context investigation, clinical reasoning, code debugging, and non-abductive controls. CEDAR- GRPO improves every model on every held-out task over both base models and correctness-only GRPO, with average gains of 7.4 and 2.7 points, respectively, and a maximum gain of 30.8 points. Ablations confirm that RL, abductive reward design, and task diversity each contribute to transfer. Process-level metrics further show stronger abductive behavior, including exploration of alternatives, elimination of rivals, backtracking, and uncertainty marking.
摘要:誘導推理,通常被描述為最佳解釋的推斷,是在不確定性下解釋的核心,從日常的意義建構和調查到科學發現。
然而,LLM 研究大多通過狹窄的、特定任務的基準來研究誘導,這使得觀察到的增益是否能轉移到用於訓練或評估的基準家族之外變得不清楚。
我們詢問 RL 後訓練是否能改善誘導作為可轉移的推理能力。
我們介紹 CEDAR-GRPO,一個過程感知框架,將最終答案的正確性與誘導獎勵結合,考慮證據覆蓋率和證據到解釋的方向性。
四個開放權重的 LLM 在一個受控的、領域中立的誘導假設生成和假設選擇任務的混合上進行後訓練。
我們在 11 個未見過的任務上評估它們,這些任務涵蓋假設選擇、缺失事實生成、可駁斥推理、長上下文調查、臨床推理、代碼調試和非誘導控制。
CEDAR-GRPO 在每個保留任務上改善了每個模型,無論是基礎模型還是僅考慮正確性的 GRPO,平均增益分別為 7.4 和 2.7 分,最大增益為 30.8 分。
消融實驗確認 RL、誘導獎勵設計和任務多樣性各自對轉移有貢獻。
過程級別的指標進一步顯示出更強的誘導行為,包括探索替代方案、消除競爭者、回溯和不確定性標記。
Seeing Red, Thinking Bad: Color Bias in Vision Language Models
2608.14286v1 by Kohsuke Ide, Ryousuke Yamada, Yoshihiro Fukuhara, Hirokatsu Kataoka, Yutaka Satoh
Vision language models (VLMs) are increasingly used in industrial decision-making systems, such as recruitment support and recommendation. This motivates careful analysis of how VLMs process visual and textual information. In this work, we study how VLMs interpret text rendered as an image, and investigate the influence of visual styling biases. To this end, we introduce Stealth Visual Prompts, which subtly change visual styling of text, such as color and contrast, while preserving semantic content. Using these prompts, we systematically control the visual styling of words in text and measure their impact on the analysis performed by VLMs. We further analyze how such visual perturbations affect the latent representations of the vision encoder. From our experiments, we observed that coloring positive words in green consistently shifts sentiment predictions toward a positive direction. As a result, VLMs often fail to properly account for negative words present in the text. Our analysis suggests that this behavior is correlated with changes in the latent representations of the vision encoder induced by color variations. In addition, we show that reducing text--background contrast increases reliance on visually salient cues and leads to more incorrect Visual Question Answering (VQA) outputs. These results suggest that the visual styling of rendered text can guide VLMs' interpretation in ways that diverge from human semantic understanding. Project page: https://github.com/KohsukeIde/color-bias-vlm
摘要:視覺語言模型(VLMs)在工業決策系統中越來越多地被使用,例如招聘支持和推薦。這促使我們仔細分析 VLMs 如何處理視覺和文本信息。在這項工作中,我們研究 VLMs 如何解釋以圖像呈現的文本,並調查視覺風格偏見的影響。為此,我們引入了隱形視覺提示,這些提示微妙地改變文本的視覺風格,例如顏色和對比度,同時保留語義內容。利用這些提示,我們系統地控制文本中單詞的視覺風格,並測量其對 VLMs 執行的分析的影響。我們進一步分析這些視覺擾動如何影響視覺編碼器的潛在表示。從我們的實驗中,我們觀察到將正面詞語著色為綠色會持續地將情感預測向正面方向偏移。因此,VLMs 經常無法正確考慮文本中存在的負面詞語。我們的分析表明,這種行為與由顏色變化引起的視覺編碼器潛在表示的變化相關。此外,我們顯示減少文本與背景的對比度會增加對視覺顯著線索的依賴,並導致更多不正確的視覺問題回答(VQA)輸出。這些結果表明,渲染文本的視覺風格可以以偏離人類語義理解的方式引導 VLMs 的解釋。
項目頁面:https://github.com/KohsukeIde/color-bias-vlm
Meteorology-driven Causal Nowcasting of Fugitive Landfill Emissions Enables Proactive Public Health Response
2608.14254v1 by Timothy C. Pearce, David J. T. Smith, Alec Dobney, Alessia Freddo
Fugitive emissions from waste sites increasingly expose communities to toxic and odorous gases, yet public-health responses remain largely retrospective, with episodes investigated only after residents have been exposed. Here we show that the meteorological drivers of elevated hydrogen sulphide (HS) at a long-monitored European landfill, and the timescales over which they act, can be identified directly from routine monitoring data. We introduce CAIRN (Causal-Anchored Inference for Receptor Nowcasting), a machine-learning framework whose internal memory is matched to these measured timescales: a fast component tracking hour-scale wind-borne transport and a slow component tracking multi-hour weather changes. Trained to predict gas measurements, CAIRN operates using only routine weather variables and the calendar, without hand-engineered features. Its behaviour is consistent with the identified transport mechanisms, and the framework transfers unchanged to a second monitoring station and to co-emitted methane. Combining four such nowcasters produces a site-level, tiered alert aligned with WHO odour guidance that closely reproduces the alert generated by a direct sensor network and tracks an independent record of community odour complaints. Weather-driven nowcasting can therefore estimate community impact as an emission episode unfolds, providing public-health authorities with a validated, graded trigger for intervention and enabling exposure to be reduced during events rather than after them.
摘要:逃逸排放的廢棄物場越來越多地使社區暴露於有毒和有氣味的氣體中,但公共衛生的反應仍然主要是事後的,只有在居民受到影響後才進行調查。在這裡,我們展示了在一個長期監測的歐洲垃圾填埋場中,氫硫化物(HS)濃度升高的氣象驅動因素及其作用的時間尺度,可以直接從常規監測數據中識別出來。我們引入了CAIRN(因果錨定推斷接收器即時預測),這是一個機器學習框架,其內部記憶與這些測量的時間尺度相匹配:一個快速組件跟踪小時級的風載運輸,另一個慢速組件跟踪多小時的天氣變化。CAIRN經過訓練以預測氣體測量,僅使用常規的天氣變數和日曆,而不需要手工設計的特徵。其行為與識別出的運輸機制一致,並且該框架可以不變地轉移到第二個監測站和共同排放的甲烷。結合四個這樣的即時預測器,產生了一個與世界衛生組織氣味指導相一致的現場級分層警報,該警報與直接傳感器網絡生成的警報非常接近,並跟踪社區氣味投訴的獨立記錄。因此,天氣驅動的即時預測可以在排放事件展開時估計社區影響,為公共衛生當局提供經過驗證的分級干預觸發器,並使在事件期間減少暴露成為可能,而不是在事件之後。
Agentic Data Cleaning Without a Clean Reference: An Experimental Study of Capabilities and Trade-offs
2608.14765v1 by Hadi Fadlallah
Data cleaning without a trusted clean reference is challenging because unusual values may represent either genuine errors or valid observations. This paper studies how different agent capabilities affect reference-free data cleaning and proposes an evidence-grounded framework that combines structured context, profiling, LLM reasoning, executable checks, controlled evidence retrieval, source ranking, citation alignment, conservative repair, reversible scripts, and provenance logging. Seven configurations are evaluated across financial, clinical, and environmental-monitoring datasets using controlled synthetic corruption and original-data descriptive analysis, resulting in 126 completed runs. The evaluation includes two comparison baselines and a progressive LLM-based sequence that adds executable tools, evidence retrieval, evidence controls, and conservative repair. In the synthetic evaluation, the deterministic profiling baseline achieved the highest detection F1-score of 0.561. Among the LLM-based configurations, the full conservative configuration achieved the highest F1-score of 0.421, but no configuration performed best across all evaluation criteria. The source-ranked configurations achieved the lowest unsupported-rule rates, while decision-level citation alignment remained weak. The full conservative configuration produced no unsafe or unnecessary modifications, although these rates were already zero before the conservative policy was added, and it performed no direct repairs. Overall, the results show that additional capabilities introduce trade-offs among detection, repair, evidence grounding, conservative behaviour, reproducibility, and operational cost rather than producing consistent improvements. The study provides a structured framework and empirical methodology for evaluating these trade-offs in reference-free agentic data cleaning.
摘要:數據清理在沒有可信的乾淨參考的情況下是具有挑戰性的,因為異常值可能代表真正的錯誤或有效的觀察結果。本文研究了不同代理能力如何影響無參考數據清理,並提出了一個基於證據的框架,該框架結合了結構化上下文、檔案分析、LLM 推理、可執行檢查、受控證據檢索、來源排名、引用對齊、保守修復、可逆腳本和來源日誌。通過使用受控的合成腐蝕和原始數據描述性分析,對七種配置進行了評估,涵蓋了金融、臨床和環境監測數據集,最終完成了126次運行。評估包括兩個比較基準和一個逐步的基於 LLM 的序列,該序列添加了可執行工具、證據檢索、證據控制和保守修復。在合成評估中,確定性檔案分析基準達到了最高的檢測 F1 分數 0.561。在基於 LLM 的配置中,完整的保守配置達到了最高的 F1 分數 0.421,但沒有任何配置在所有評估標準中表現最佳。來源排名配置達到了最低的不支持規則率,而決策級引用對齊仍然較弱。完整的保守配置未產生任何不安全或不必要的修改,儘管在添加保守政策之前這些比率已經為零,並且它沒有進行直接修復。總體而言,結果顯示,額外的能力在檢測、修復、證據基礎、保守行為、可重複性和運營成本之間引入了權衡,而不是產生一致的改進。該研究提供了一個結構化框架和實證方法,用於評估這些在無參考代理數據清理中的權衡。
APTER: Adaptive Post-Training with Expert-Grounded Rubrics
2608.14212v1 by Xukai Wang, Liangqi Li, Zhiyue Xu, Jingang Zhou, Xiaoyu Shi, Jiansheng Cai, Bo Zhang, Zhe Li, Xu-Yao Zhang
As large language models enter professional domains, they must satisfy domain constraints, include critical evidence, and provide complete reasoning rather than merely produce fluent responses. Existing post-training methods often rely on holistic preferences or outcome-level verification, while recent rubric-based methods usually generate rubrics independently for each query. In specialized domains, such unconstrained rubrics may omit critical requirements and vary across samples, hindering the diagnosis and targeted repair of persistent capability deficiencies. We propose APTER (Adaptive Post-Training with Expert-Grounded Rubrics), a framework that integrates structured domain knowledge into fine-grained evaluation, optimization, and diagnosis for specialized complex reasoning. First, expert-grounded rubric construction starts from an expert criteria framework built by domain experts, where each criterion represents a stable professional capability. For each query, APTER selects relevant criteria and instantiates them into query-level rubrics linked to their source criteria, turning reusable expert criteria into executable query-level supervision without reference answers. Second, adaptive post-training uses rubric verdicts as both optimization and criterion-level diagnostic signals. Aggregating low-scoring verdicts by criterion ID reveals persistent deficiencies and triggers targeted supervised fine-tuning updates during reinforcement learning. Experiments on mathematical reasoning and medical question answering show consistent gains across both domains. Across three model generations, APTER improves the mathematics and medical averages over the corresponding base models by up to 15.86 and 8.04 points, respectively. Code and rubric datasets are available at https://github.com/AntDT-APTER/APTER.
摘要:隨著大型語言模型進入專業領域,它們必須滿足領域約束,包含關鍵證據,並提供完整的推理,而不僅僅是產生流暢的回應。現有的後訓練方法通常依賴於整體偏好或結果層級的驗證,而最近的基於評分標準的方法通常為每個查詢獨立生成評分標準。在專業領域中,這種不受約束的評分標準可能會省略關鍵要求,並在樣本之間變化,妨礙持續能力缺陷的診斷和針對性修復。我們提出了APTER(基於專家的自適應後訓練評分標準),這是一個將結構化領域知識整合到細緻評估、優化和診斷中的框架,旨在解決專業複雜推理的問題。首先,專家基礎的評分標準構建始於由領域專家建立的專家標準框架,其中每個標準代表一種穩定的專業能力。對於每個查詢,APTER選擇相關標準並將其實例化為與其來源標準相關聯的查詢層級評分標準,將可重用的專家標準轉化為可執行的查詢層級監督,而不需要參考答案。其次,自適應後訓練使用評分標準的判決作為優化和標準層級診斷信號。通過標準ID聚合低分判決可以揭示持續的缺陷,並在強化學習過程中觸發針對性的有監督微調更新。在數學推理和醫學問題回答的實驗中,兩個領域均顯示出一致的增長。在三代模型中,APTER在數學和醫學的平均分數上分別提高了高達15.86和8.04分。代碼和評分標準數據集可在 https://github.com/AntDT-APTER/APTER 獲得。
Removing Temporal Note Redundancy Improves Multimodal Reinforcement Learning for Medicine
2608.14157v1 by Chenran Weng, Joo Seung Lee, Malini Mahendra, Anil Aswani
Mechanical ventilation is a critical life-support intervention, requiring dynamic adjustments to ventilator settings as a patient's condition evolves. While reinforcement learning (RL) offers a promising framework for optimizing these sequential decisions, standard approaches rely primarily on structured electronic health record (EHR) data, missing crucial clinical context recorded in free-text notes. Integrating longitudinal clinical notes into RL state spaces is challenging because notes are heavily inflated by temporal redundancy, such as copy-forward text, templating, and repetitive documentation, which dilutes time-local updates and degrades state representation quality. To address this, we propose a redundancy-aware multimodal state representation framework that explicitly removes duplicated note text over time before policy learning. We evaluate two computationally efficient temporal decomposition strategies for removing duplicated note text: (1) an embedding-space decomposition using singular value decomposition on local history subspaces, and (2) an interpretable sentence-level diff operation that filters out previously documented sentences before text encoding. Using real-world ICU data, we demonstrate that state representations constructed by stripping temporal note redundancy significantly outperform both structured-only and raw-note baselines across multiple off-policy evaluation methods (Model-Based Rollouts, Fitted Q-Evaluation, Weighted Importance Sampling, and Weighted Doubly Robust Evaluation). Our findings show that explicitly isolating new clinical information from repeated note text yields higher-quality state representations and directly improves RL performance for clinical decision support.
摘要:機械通氣是一項關鍵的生命支持干預,隨著病人狀況的變化,需要對通氣器設置進行動態調整。雖然強化學習(RL)提供了一個有前景的框架來優化這些序列決策,但標準方法主要依賴於結構化的電子健康記錄(EHR)數據,忽略了在自由文本註解中記錄的重要臨床背景。將長期臨床註解整合進RL狀態空間是具有挑戰性的,因為註解受到時間冗餘的嚴重影響,例如複製轉發文本、模板化和重複文檔,這稀釋了時間局部更新並降低了狀態表示的質量。為了解決這個問題,我們提出了一個冗餘感知的多模態狀態表示框架,該框架在策略學習之前明確去除隨時間重複的註解文本。我們評估了兩種計算效率高的時間分解策略來去除重複的註解文本:(1)使用奇異值分解對局部歷史子空間進行的嵌入空間分解,以及(2)一種可解釋的句子級差異操作,在文本編碼之前過濾掉先前記錄的句子。使用真實世界的ICU數據,我們展示了通過剝離時間註解冗餘構建的狀態表示在多種離線政策評估方法(基於模型的回滾、擬合Q評估、加權重要性抽樣和加權雙重穩健評估)中顯著優於僅結構化和原始註解的基準。我們的研究結果顯示,明確將新的臨床信息與重複的註解文本隔離,可以產生更高質量的狀態表示,並直接改善臨床決策支持的RL性能。
CMCNet: Aligning Ultrasound Image Embeddings with Textual TI-RADS Representations for Fine-Grained Thyroid Classification
2608.13939v1 by Bingxin Yu, Xueli Wang, Jerry Zhou, Wenyan Wang, Li Wen, Lan Huang, Xin Feng, Fengfeng Zhou, Kewei Li
Ultrasound is the primary imaging modality for assessing thyroid nodules, and the ACR TI-RADS framework standardizes diagnosis through five ultrasound feature categories that are aggregated into five risk levels (TR1-TR5). Although widely adopted in clinical practice, most deep learning approaches focus on binary malignancy classification, while multi-class prediction and explicit utilization of feature-level supervision remain underexplored, largely due to limited annotated data. In this study, we introduce the STN dataset of 600 thyroid nodules with paired transverse and longitudinal ultrasound images, bounding box annotations, and complete labels for all five TI-RADS feature categories. Following the clinical decision process, we investigate how structured feature information can guide representation learning during training while requiring only images at inference. We demonstrate that text embeddings derived from standardized feature descriptions form a stable surrogate representation for TI-RADS risk levels. Based on this observation, we propose CMCNet, which aligns image embeddings to fixed textual embeddings via a Center-Margin Contrastive Loss that simultaneously promotes intra-class compactness and inter-class separation. Experimental results show that this embedding alignment strategy is more data-efficient and robust than direct multitask learning, and consistently outperforms InfoNCE, center loss, a strong multitask baseline, and a VQA-style multimodal model, particularly in imbalanced settings. The dataset is freely available at doi: 10.5281/zenodo.19125693 and the source code is available at: https://www.healthinformaticslab.org/supp/.
摘要:超聲波是評估甲狀腺結節的主要影像學方法,而 ACR TI-RADS 框架通過五個超聲特徵類別標準化診斷,這些特徵被聚合成五個風險等級(TR1-TR5)。儘管在臨床實踐中被廣泛採用,但大多數深度學習方法專注於二元惡性分類,而多類別預測和明確利用特徵級監督的研究仍然未被充分探索,這主要是由於標註數據的限制。在本研究中,我們引入了 STN 數據集,其中包含 600 個甲狀腺結節的配對橫向和縱向超聲圖像、邊界框註釋以及所有五個 TI-RADS 特徵類別的完整標籤。根據臨床決策過程,我們探討結構化特徵信息如何在訓練期間指導表示學習,同時在推理時僅需圖像。我們證明來自標準化特徵描述的文本嵌入形成了 TI-RADS 風險等級的穩定替代表示。基於這一觀察,我們提出了 CMCNet,該模型通過中心-邊距對比損失將圖像嵌入與固定文本嵌入對齊,這同時促進了類內緊湊性和類間分離性。實驗結果顯示,這種嵌入對齊策略比直接的多任務學習更具數據效率和穩健性,並且在不平衡設置中始終優於 InfoNCE、中心損失、一個強大的多任務基線以及一個 VQA 風格的多模態模型。該數據集可免費獲得,DOI 為:10.5281/zenodo.19125693,源代碼可在:https://www.healthinformaticslab.org/supp/ 獲得。
Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions
2608.13786v1 by Qingfang Liu, Qiao Jin, Joe D. Menke, Thorsten Kahnt, Zhiyong Lu
Large language model (LLM) chatbots are increasingly used to answer clinical questions with citations to relevant clinical studies. Prior research has largely focused on citation fabrication, leaving a gap in evaluating the quality of retrieved studies and the factors driving their selection. In this study, we evaluated three general-purpose LLM chatbots: Claude Sonnet 5, Gemini 3.1 Pro, and ChatGPT GPT-5.5. We prompted the models with clinical questions adapted from 20 review questions in Issues 6 and 7 of the 2026 Cochrane Database of Systematic Reviews, simulating patient, clinician, and evidence-synthesis researcher roles. Each chatbot was queried under each user role with four independent repetitions, yielding 720 responses. Each chatbot was asked to support its answers with primary clinical citations, which we benchmarked against the included and excluded study sets of the Cochrane reviews. On average, a chatbot response retrieved 39.2% $\pm$ 29.8% of Cochrane included studies, while citing 5.0% $\pm$ 9.4% of excluded studies. Recall of Cochrane included studies varied significantly by model and user role. ChatGPT achieved higher recall than Claude or Gemini (63.1% $\pm$ 29.5% vs. 37.0% $\pm$ 23.8% vs. 17.3% $\pm$ 13.1%; $p=2.0\times10^{-5}$). The researcher role yielded higher recall than the clinician or patient roles (42.8% $\pm$ 30.8% vs. 38.6% $\pm$ 28.9% vs. 36.1% $\pm$ 29.3%; $p=2.0\times10^{-5}$). Controlling for publication year, citations per year, and open-access status, sample size was the only independently significant predictor of retrieval (odds ratio 1.80 per 1-unit increase in log sample size, 95% CI 1.37-2.36, $p=2.34\times10^{-5}$). These findings suggest that while LLM chatbots can retrieve some studies identified by expert reviewers, their performance varies by model and user role, and they exhibit a bias toward clinical trials with larger sample sizes.
摘要:大型語言模型(LLM)聊天機器人越來越多地用於回答臨床問題,並引用相關的臨床研究。先前的研究主要集中在引用虛構上,未能評估檢索到的研究質量及其選擇的驅動因素。在本研究中,我們評估了三個通用型LLM聊天機器人:Claude Sonnet 5、Gemini 3.1 Pro和ChatGPT GPT-5.5。我們根據2026年Cochrane系統評價數據庫第6和第7期的20個回顧問題,為模型提供了臨床問題的提示,模擬患者、臨床醫生和證據綜合研究者的角色。每個聊天機器人在每個用戶角色下進行了四次獨立查詢,共產生720個回應。每個聊天機器人被要求用主要臨床引用來支持其答案,我們將其與Cochrane評估的納入和排除研究集進行了基準比較。平均而言,聊天機器人的回應檢索了39.2% $\pm$ 29.8%的Cochrane納入研究,同時引用了5.0% $\pm$ 9.4%的排除研究。Cochrane納入研究的回憶率因模型和用戶角色而異。ChatGPT的回憶率高於Claude或Gemini(63.1% $\pm$ 29.5% vs. 37.0% $\pm$ 23.8% vs. 17.3% $\pm$ 13.1%;$p=2.0\times10^{-5}$)。研究者角色的回憶率高於臨床醫生或患者角色(42.8% $\pm$ 30.8% vs. 38.6% $\pm$ 28.9% vs. 36.1% $\pm$ 29.3%;$p=2.0\times10^{-5}$)。在控制出版年份、每年引用數和開放獲取狀態後,樣本大小是唯一獨立顯著的檢索預測因子(對數樣本大小每增加1單位的比值比1.80,95% CI 1.37-2.36,$p=2.34\times10^{-5}$)。這些發現表明,儘管LLM聊天機器人可以檢索到一些專家評審者識別的研究,但其性能因模型和用戶角色而異,並且對樣本大小較大的臨床試驗存在偏見。
Data-driven techniques for translational neuroscience and personalized neuro-health
2608.13749v1 by Vishal Subedi, Shashipraba N. K. Rajakaruna, Pratyusha Sarkar, Subhankar Chattoraj, Anjali Khasa, Siddhartha Nandy, Hamza Farooq, Animikh Biswas, Sanjay Chaudhuri, Asim K. Dey, Karuna Joshi, Christophe Lenglet, Ansu Chatterjee
Neurodegenexrative diseases such as Alzheimer's disease and Parkinson's disease are diagnosed most reliably only after substantial, often irreversible, neuronal loss has already occurred, creating an urgent need for quantitative tools that can detect subtle, early, and individual-specific brain changes from neuroimaging data. This review surveys a broad and rapidly evolving toolkit of data-driven techniques for translational neuroscience and personalized neuro-health, organized around four complementary methodological pillars. Throughout, we emphasize how these methodologically diverse approaches converge on a common translational goal: personalized, mechanistically grounded, and clinically actionable models of individual brain health, and we close by discussing the principal open statistical, computational, and clinical challenges that remain.
摘要:神經退行性疾病,如阿茲海默症和帕金森病,通常只有在已經發生了實質性且通常是不可逆的神經元損失後,才能最可靠地診斷,這造成了對能夠從神經影像數據中檢測微妙、早期且個體特異性腦部變化的定量工具的迫切需求。這篇綜述調查了一套廣泛且快速發展的數據驅動技術工具,旨在轉化神經科學和個性化神經健康,並圍繞四個互補的方法論支柱進行組織。整篇文章強調這些方法論多樣的途徑如何匯聚到一個共同的轉化目標:個性化、機制基礎的且臨床可行的個體腦健康模型,並在結尾討論仍然存在的主要統計、計算和臨床挑戰。
MedPlex: Deep Vision-Language Co-Adaptation for Clinically Grounded Medical Segmentation
2608.13690v1 by Rafi Ibn Sultan, Hui Zhu, Chengyin Li, Dongxiao Zhu
Medical image segmentation is still largely treated as a vision-only problem, although clinical interpretation often relies on textual knowledge of anatomy, location, appearance, and surrounding context. Existing text-guided segmentation methods within the Vision-Language Model (VLM) paradigm often use language only as a late conditioning signal, limiting its influence on visual representation learning. We introduce MedPlex (Medical Plexus of Vision and Language), an end-to-end VLM framework that makes text guidance a continuous, clinically grounded component of segmentation learning. Through Bi-Fusion (Bidirectional Fusion), visual and textual representations evolve jointly across the encoding hierarchy. MedPlex further introduces class-level and region-level concept alignment to organize the shared representation at complementary granularities. Class-level alignment anchors each anatomical target to an aggregated clinical concept profile, while region-level alignment preserves individual concepts, such as shape, location, appearance, and texture, through class-specific visual evidence. In this way, language provides structured supervision throughout the encoder rather than serving only as a late-stage cue. MedPlex achieves state-of-the-art performance across CT and MR benchmarks for multi-organ, cardiac substructure, and tumor segmentation, including settings with real free-text clinical supervision. Code: https://github.com/rafiibnsultan/MedPlex.
摘要:醫學影像分割仍然主要被視為一個僅限於視覺的問題,儘管臨床解釋通常依賴於對解剖學、位置、外觀和周圍背景的文本知識。現有的文本引導分割方法在視覺-語言模型(VLM)範式內,通常僅將語言用作後期條件信號,限制了其對視覺表示學習的影響。我們介紹了 MedPlex(醫學視覺與語言的聯結),這是一個端到端的 VLM 框架,使文本引導成為分割學習中的一個持續且臨床基礎的組件。通過雙向融合(Bi-Fusion),視覺和文本表示在編碼層次中共同演變。MedPlex 進一步引入了類別級和區域級概念對齊,以在互補的粒度上組織共享表示。類別級對齊將每個解剖目標錨定到一個聚合的臨床概念檔案,而區域級對齊則通過類別特定的視覺證據保留個別概念,例如形狀、位置、外觀和質地。這樣,語言在編碼器中提供結構化的監督,而不僅僅是在後期階段作為提示。MedPlex 在多器官、心臟子結構和腫瘤分割的 CT 和 MR 基準測試中達到了最先進的性能,包括具有實際自由文本臨床監督的設置。代碼: https://github.com/rafiibnsultan/MedPlex。
MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and Coordination
2608.13476v1 by Saisha Shetty, Satvik Tripathi, Austin Lin, Colin Zhao, Theodore Kim, Don Enwerem, Jacinta Arnold, Shahriar Faghani, Tessa S Cook
We present Multi-Agent Reasoning and Coordination (MARC), an open-source framework that replaces monolithic LLM prompting with deterministic multi-agent orchestration for clinical reasoning. MARC coordinates role-specialized agents for extraction, reasoning, answer generation, and evaluation, with explicit context passing and traceable intermediate outputs, enabling stage-wise failure attribution. We additionally introduce a Decomposer module that generates task-specific agent prompts from a plain-language description, eliminating manual prompt engineering. The framework supports both API-based and local CPU-compatible deployments and is entirely configurable via YAML, without code modifications. MARC is designed to be model-agnostic, interpretable, and accessible to clinical domain experts without programming expertise. The full framework is available at https://github.com/Penn-RAIL/MARC-v1.
摘要:我們提出了多智能體推理與協調(MARC),這是一個開源框架,將單一大型語言模型的提示替換為確定性的多智能體協調,用於臨床推理。MARC 協調角色專門的智能體進行提取、推理、答案生成和評估,具有明確的上下文傳遞和可追溯的中間輸出,實現階段性失敗歸因。我們還引入了一個分解器模塊,該模塊從普通語言描述中生成任務特定的智能體提示,消除了手動提示工程。該框架支持基於 API 的和本地 CPU 兼容的部署,並且完全可以通過 YAML 配置,而無需修改代碼。MARC 設計為模型無關、可解釋,並且對於沒有編程專業知識的臨床領域專家可訪問。完整框架可在 https://github.com/Penn-RAIL/MARC-v1 獲得。
Towards Context-Aware Clinical Motion Understanding in Daily Living at Home: Freezing of Gait Detection with Egocentric Vision
2608.13283v1 by Vayalet Stefanova, Diwas Lamsal, Margot Genbrugge, Maxim Yudayev, Christian Schlenstedt, Moran Gilat, Bart Vanrumste, Benjamin Filtjens
Understanding motion in daily living requires context beyond kinematics, because similar inertial patterns during activities of daily living (ADLs) can reflect intentional stopping, object interaction, or pathological movement impairment. Egocentric vision provides task-related context that may help disambiguate these cases. We investigate this challenge through freezing of gait (FOG) detection in Parkinson's disease (PD), a symptom strongly influenced by contextual factors during ADLs. Using synchronized egocentric video, wearable IMUs, and expert-annotated FOG labels collected from 13 PD participants in their homes, we evaluate frozen representations from pretrained ego-video and time-series foundation models, alongside an IMU-based TCN trained from scratch, under leave-one-subject-out evaluation. The IMU-based TCN achieved the strongest event-detection performance, reaching 42.3 F1 and 83.0 AUROC, compared with 32.6 F1 and 77.2 AUROC for V-JEPA2 ego-video features. Although ego-video alone did not outperform IMU-based sensing, it showed above-chance discrimination, and qualitative analyses suggest that egocentric vision may capture FOG-relevant information independent of IMUs. Together, these results support the use of pretrained ego-video representations to add contextual information to wearable-sensor-based clinical motion understanding in daily living.
摘要:理解日常生活中的運動需要超越運動學的背景,因為在日常生活活動(ADLs)中類似的慣性模式可能反映出有意的停止、物體互動或病理性運動障礙。自我中心的視覺提供了與任務相關的背景,可能有助於消除這些情況的歧義。我們通過在帕金森病(PD)中的步態凍結(FOG)檢測來研究這一挑戰,這是一種在日常生活活動中受到背景因素強烈影響的症狀。使用同步的自我中心視頻、可穿戴IMU和從13名PD參與者在家中收集的專家註釋FOG標籤,我們在留一個參與者的評估下評估來自預訓練自我視頻和時間序列基礎模型的凍結表示,以及從零開始訓練的基於IMU的TCN。基於IMU的TCN達到了最強的事件檢測性能,F1達到42.3,AUROC達到83.0,而V-JEPA2自我視頻特徵的F1為32.6,AUROC為77.2。儘管僅使用自我視頻並未超越基於IMU的感測,但它顯示出超過隨機的區分能力,定性分析表明,自我中心的視覺可能捕捉到與FOG相關的信息,這與IMU無關。總體而言,這些結果支持使用預訓練的自我視頻表示,為基於可穿戴傳感器的日常生活臨床運動理解添加背景信息。
Static analysis-guided agentic AI translation enables Rust as a full stack bioinformatics language
2608.13029v1 by Johan Henriksson
The field of bioinformatics struggles with legacy code - old code that is commonly used but may no longer have a maintainer, or may be written in an now-unfamiliar language (e.g. Perl, Fortran). This incurs maintenance cost (technical debt), but dynamically typed languages also negatively impacts the environment and fail to make use of modern hardware. Legacy code may also have security or safety problems that make it unsuited for use in clinical settings. Here we show that agentic AI, combined with static analysis, can be used to translate legacy code to the modern language Rust. We provide prompts and supporting software to aid systematic translation, and evaluate it on common software for NGS and imaging. We showcase the result on our software Bascet: Size was reduced by ~80x, build time decreased by ~10x, and performance of key steps improved >3x. Unix dependencies were also removed, making Bascet the only single-cell pipeline able to run on native Windows, without a container. Large-scale refactoring of bioinformatics software is thus now possible at a limited budget, enabling more complex tools to be developed.
摘要:生物資訊學領域面臨著舊有代碼的挑戰——這些舊代碼通常被使用,但可能不再有維護者,或可能是用現在不熟悉的語言(例如 Perl、Fortran)編寫的。這會產生維護成本(技術負債),但動態類型語言也會對環境產生負面影響,並未能充分利用現代硬體。舊代碼可能還存在安全或安全性問題,使其不適合在臨床環境中使用。在這裡,我們展示了代理式 AI 結合靜態分析,可以用來將舊代碼轉換為現代語言 Rust。我們提供提示和支持軟體以協助系統性翻譯,並在 NGS 和成像的常見軟體上進行評估。我們展示了我們的軟體 Bascet 的結果:大小減少約 80 倍,建構時間減少約 10 倍,關鍵步驟的性能提高了超過 3 倍。Unix 依賴也被移除,使 Bascet 成為唯一能在本地 Windows 上運行的單細胞管道,而無需容器。因此,生物資訊學軟體的大規模重構現在在有限的預算下成為可能,從而使得更複雜的工具得以開發。
Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence
2608.12928v1 by Jakub Pokrywka, Łukasz Grzybowski, Antoni Lasik, Marek Kubis, Jeremi Ignacy Kaczmarek, Wojciech Kusa
We introduce a Polish-language medical visual question answering (VQA) benchmark, built from Polish Board Certification Examination questions for licensed physicians and dentists pursuing specialist certification. The benchmark comprises image-containing questions spanning diverse medical specialties and visual domains, together with a text-only question answering (QA) control set. We evaluate Polish-oriented, general-purpose open-weight, and commercial vision-language models. The task remains challenging: the best model achieves 79.0\% accuracy on the full VQA set, and only GPT-5.6 surpasses the approximate human reference on the subset with available candidate responses; all other evaluated models perform worse than humans. To assess visual grounding, we compare complete inputs with configurations omitting the image, the question, or both, and categorize questions by image importance. Models derive more useful information from the question text than from the image and perform worse on image-dominant questions. Across both QA and VQA, they nevertheless achieve above-chance accuracy from the answer choices alone, showing that non-trivial performance can persist even when key task components are missing.
摘要:我們介紹了一個波蘭語醫學視覺問題回答(VQA)基準,該基準是基於波蘭醫師和牙醫專業認證考試問題而建立的。這個基準包含了涵蓋多種醫學專業和視覺領域的圖像問題,以及一組僅包含文本的問題回答(QA)控制集。我們評估了針對波蘭的通用開放權重和商業視覺語言模型。這項任務仍然具有挑戰性:最佳模型在完整的 VQA 集上達到 79.0\% 的準確率,只有 GPT-5.6 在具有可用候選回答的子集上超過了近似的人類參考;所有其他評估的模型表現都不如人類。為了評估視覺基礎,我們比較了完整輸入與省略圖像、問題或兩者的配置,並根據圖像的重要性對問題進行分類。模型從問題文本中獲取的有用信息多於從圖像中獲取的,並且在圖像主導的問題上表現較差。在 QA 和 VQA 中,它們仍然僅從答案選擇中實現了超過隨機的準確率,顯示出即使在缺少關鍵任務組件的情況下,非平凡的表現仍然可以持續存在。
CRAFT: LLM-Based Iterative Refinement for Temporal Reasoning over Clinical Narratives
2608.12779v1 by Chengyang He, Tahreem Arif, Marko Zivkovic, Lijing Wang, Yue Ning, Ping Wang
Understanding the temporal progression of symptoms in clinical narratives is critical for disease monitoring, safety surveillance, and causality assessment. Clinical narratives, however, rarely provide explicit temporal anchors. Current approaches to temporal information reasoning focus predominantly on pairwise relation classification across multi-visit and timestamp-rich records, leaving the reconstruction of structured symptom trajectories from individual anchor-sparse reports largely unaddressed. We propose CRAFT, an LLM framework that pairs a generator with a constraint-based verifier to iteratively produce and refine stage-wise symptom timelines through targeted feedback. We conduct evaluation on MedTempo, a new benchmark of 5,347 vaccine adverse-event narratives spanning three COVID-19 vaccine types, with expert-validated temporal stage annotations for 3,166 reports. Experiments across four LLM backbones demonstrate that CRAFT consistently improves temporal ordering accuracy, with ablation analysis isolating the contribution of generator and verifier components across model capability levels.
摘要:理解臨床敘述中症狀的時間進展對於疾病監測、安全監控和因果評估至關重要。
然而,臨床敘述很少提供明確的時間錨點。
目前對時間信息推理的方法主要集中在多次訪問和時間戳豐富記錄之間的成對關係分類,這使得從個別缺乏錨點的報告中重建結構化的症狀軌跡在很大程度上未得到解決。
我們提出了CRAFT,一個將生成器與基於約束的驗證器配對的LLM框架,通過有針對性的反饋迭代生成和完善階段性症狀時間線。
我們在MedTempo上進行評估,這是一個新的基準,包含5,347個疫苗不良事件敘述,涵蓋三種COVID-19疫苗類型,並對3,166個報告進行了專家驗證的時間階段標註。
在四個LLM骨幹上進行的實驗表明,CRAFT持續提高了時間排序的準確性,並通過消融分析隔離了生成器和驗證器組件在模型能力水平上的貢獻。
Memorization Diagnostics for Code LLMs Should be Scale-Aware
2608.12771v1 by Prateek Kumar Rajput, Abdoul Aziz Bonkoungou, Alberick Euraste Djiré, Xunzhu Tang, Yewei Song, Iyiola Emmanuel Olatunji, El Hacen Diallo, Jacques Klein, Tegawendé F. Bissyandé
The extent to which large language models for code rely on memorization over genuine understanding remains highly debated. While current literature frequently reports widespread memorization, evaluating the underlying probing techniques across dense architectures reveals a severe breakdown in their utility at scale. Traditional encoder-style probes using perturbations such as synonym fuzzing or dead-code insertion struggle to expose memorization in scaled models, even on known-contaminated benchmarks, and decoder-style probes that rely on log probabilities show similar performance degradation. The specific mode of failure for these probes, particularly why such techniques disrupt smaller models but fail to impact larger ones, motivates us to untangle representation load from memorization rather than treating them as a single phenomenon. By applying invertible mathematical transforms to numeric problems, we isolate these two factors and reveal that scaled encoders successfully absorb substantial representation load while still converging on the correct family of solutions. In practical software engineering, this ability to adapt to varying surface forms is what truly matters for usability and generalizability in LLM and agentic applications. Whether a specific solution was seen during training becomes a much less pressing question because although memorization inflates scores on contaminated benchmarks, factoring out representation load makes it debatable how much we should truly care if a functional answer was originally memorized. Future evaluations must therefore be built around separating these phenomena rather than relying on methodologies that quietly entangle them.
摘要:大型語言模型在代碼方面依賴於記憶而非真正理解的程度仍然存在高度爭議。
儘管目前的文獻經常報導廣泛的記憶現象,但對於密集架構中潛在探測技術的評估顯示,它們在大規模應用中的效用嚴重下降。
傳統的編碼器風格探測器使用同義詞模糊或死代碼插入等擾動,難以在擴展模型中揭示記憶,即使在已知受污染的基準上也是如此,而依賴於對數概率的解碼器風格探測器表現出類似的性能下降。
這些探測器的具體失效模式,特別是為什麼這些技術會干擾較小的模型但無法影響較大的模型,促使我們將表示負載與記憶分開,而不是將它們視為單一現象。
通過對數值問題應用可逆數學變換,我們將這兩個因素隔離,並揭示擴展的編碼器成功吸收了大量的表示負載,同時仍然收斂於正確的解決方案族。
在實際的軟體工程中,這種適應不同表面形式的能力對於大型語言模型和代理應用的可用性和通用性來說才是真正重要的。
在訓練期間是否見過特定解決方案變得不再是個緊迫的問題,因為儘管記憶會在受污染的基準上膨脹分數,但剔除表示負載使得我們對於一個功能性答案最初是否被記住的關心程度變得可爭辯。
因此,未來的評估必須圍繞分離這些現象建立,而不是依賴於那些靜默糾纏它們的方法論。
PatientAct: Theory-Grounded Mental Health Client Simulation
2608.12750v1 by Sahand Sabour, TszYam NG, Yaqian Chen, Guanqun Bi, Jialu Zhao, Minlie Huang
LLM-based simulated clients are increasingly used to train novice counselors, evaluate LLM therapists, and generate synthetic data. However, current simulators produce overly cooperative clients that disclose too readily, accept therapeutic reframes without resistance, and resolve core issues within a single session. We trace these issues to profiles that lack causal depth and behavioral mechanisms that treat all content as equally accessible. We present PatientAct, a framework for client simulation grounded in established clinical theories. Our profiles integrate the 5Ps clinical case formulation, providing causal depth without tying the design to any single therapeutic modality. During simulation, profiles include a dynamic memory layer in which items carry trust thresholds (e.g., symptoms are available early, whereas formative memories require a sustained therapeutic alliance). At each turn, the client's emotional reaction and behavior are modeled before generating a response. If the therapist approaches gated content, PatientAct expresses resistance in terms of quantity, content, and style rather than defaulting to cooperation or a single resistance pattern. We evaluate our framework on 40 clinical situations and demonstrate that it generates diverse profiles with high clinical plausibility. Moreover, PatientAct significantly outperforms the baselines, yielding substantial gains in resistance quality and behavioral realism. Our code and data will be publicly available via github.com/Sahandfer/PatientHub.
摘要:LLM 基礎的模擬客戶越來越多地用於訓練新手輔導員、評估 LLM 治療師和生成合成數據。
然而,當前的模擬器產生過於合作的客戶,他們過於輕易地透露信息,毫無抵抗地接受治療重構,並在單一會話中解決核心問題。
我們將這些問題追溯到缺乏因果深度的檔案和將所有內容視為同等可接近的行為機制。
我們提出了 PatientAct,一個基於已建立臨床理論的客戶模擬框架。
我們的檔案整合了 5Ps 臨床案例形成,提供因果深度而不將設計綁定於任何單一的治療模式。
在模擬過程中,檔案包括一個動態記憶層,其中項目承載信任閾值(例如,症狀早期可用,而形成性記憶則需要持續的治療聯盟)。
在每一輪中,客戶的情感反應和行為在生成回應之前被建模。
如果治療師接觸到受限內容,PatientAct 會在數量、內容和風格上表達抵抗,而不是默認合作或單一的抵抗模式。
我們在 40 個臨床情境中評估了我們的框架,並證明它生成了具有高臨床合理性的多樣化檔案。
此外,PatientAct 顯著超越了基準,帶來了在抵抗質量和行為現實主義方面的重大提升。
我們的代碼和數據將通過 github.com/Sahandfer/PatientHub 公開提供。
Mr3D-VL: A generalist vision language foundation model for Multiparametric 3D Magnetic Resonance Imaging
2608.12689v1 by Zhi Qiao, Xintong Wu, Yichu He, Feng Shi
Multi-parametric magnetic resonance imaging (mpMRI) is a cornerstone for brain tumor diagnosis and treatment, yet current AI models face critical limitations: their lack of natural language interaction and interpretability impedes spatial information integration and cross-modal reasoning required clinically. Key challenges arise from significant physical meaning differences across modalities, spatial misalignment due to scan intervals, and the need for complex multi-feature interpretation in tasks like glioma grading. While visual-language models (VLMs) show promise in cross-modal understanding, existing methods focus mainly on 2D image modeling, neglecting direct perception of 3D volumetric space. Although 3D VLMs have been proposed for report generation and feature alignment in 3D CT imaging, mpMRI applications demand collaborative inference across multiple imaging modalities-a requirement unmet by current solutions. To address this, we introduce Mr3D-VL, a dedicated visual-language foundation model for multi-parametric 3D MRI. With 4 billion parameters, it employs an unsupervised pre-trained shared 3D encoder and 4D rotational positional embedding for dual modality-spatial integration. Its cross-modal projection layer uses a multi-resolution feature implantation strategy to enhance feature perception across resolutions. Experimental results show significant improvements over existing 4B/7B/30B domain-specific and general-purpose models in text generation tasks, achieving a BERTScore of 0.856 for report generation, with question-answering accuracy at 0.713 and multiple-choice accuracy at 0.912.
摘要:多參數磁共振成像(mpMRI)是腦腫瘤診斷和治療的基石,但目前的AI模型面臨重大限制:缺乏自然語言互動和可解釋性妨礙了臨床所需的空間信息整合和跨模態推理。主要挑戰來自於不同模態之間顯著的物理意義差異、由於掃描間隔造成的空間錯位,以及在如膠質瘤分級等任務中對複雜多特徵解釋的需求。儘管視覺語言模型(VLMs)在跨模態理解方面顯示出潛力,但現有方法主要集中在2D圖像建模,忽略了對3D體積空間的直接感知。雖然已提出3D VLMs用於報告生成和3D CT成像中的特徵對齊,但mpMRI應用需要跨多個成像模態的協作推理——這一需求目前的解決方案無法滿足。為了解決這個問題,我們推出了Mr3D-VL,一個專門針對多參數3D MRI的視覺語言基礎模型。它擁有40億個參數,採用無監督預訓練的共享3D編碼器和4D旋轉位置嵌入進行雙模態空間整合。其跨模態投影層使用多解析度特徵植入策略來增強不同解析度間的特徵感知。實驗結果顯示,在文本生成任務中,與現有的4B/7B/30B領域特定和通用模型相比,顯著提高了性能,報告生成的BERTScore達到0.856,問答準確率為0.713,多選準確率為0.912。
SteerBench-Work: A Benchmark for Agent Steering at Action Boundaries
2608.12654v1 by Oguz Serdar, Cuneyt Mertayak
Long-running LLM agents act through tools, and a single step can send an email, merge a pull request, or wire a payment. The steering decision is the pre-commit choice at that boundary: proceed, or hold for human or policy review. We introduce SteerBench-Work, an incident-anchored, bidirectional benchmark for that decision in workplace agents across developer operations, customer service, finance, legal, medical, HR, and security. Release v2026-05 contains 106 scenarios anchored in public incidents, paired evidence-reversed mirrors, and calibration controls, with labels split nearly evenly between proceed and hold so the two error directions get near-identical numbers of chances. A model sees the proposed action and the available evidence, returns a gate decision, and is scored on whether it crosses or holds the boundary correctly. Across 30 model conditions the failures run almost entirely in one direction: models wrongly hold authorized, evidence-cleared work on 28.1% of opportunities and wrongly allow unsafe work on 1.0%. The hardest cases are risk-resolved commits, where signed or structured evidence has already cleared a real risk trigger, and models score markedly worse on evidence-reversed mirrors of famous incidents (63.8%) than on the incidents themselves (98.5%). General capability is not the same as steering calibration: higher-capability models often over-refuse at the commit boundary, and more reasoning can repair a weak gate while leaving a calibrated one flat. The public leaderboard is at steerbench.com.
摘要:長期運行的 LLM 代理透過工具進行操作,單一步驟可以發送電子郵件、合併拉取請求或進行付款。引導決策是在那個邊界的預提交選擇:繼續,或等待人類或政策審查。我們介紹 SteerBench-Work,這是一個以事件為基礎的雙向基準,用於在開發運營、客戶服務、金融、法律、醫療、人力資源和安全等工作場所代理中的該決策。
版本 v2026-05 包含 106 個基於公共事件的場景,配對的證據反向鏡像和校準控制,標籤在繼續和保持之間幾乎均勻分配,以便兩個錯誤方向獲得幾乎相同的機會數量。一個模型看到提議的行動和可用的證據,返回一個閘決策,並根據它是否正確地跨越或保持邊界進行評分。在 30 種模型條件下,失敗幾乎完全朝一個方向發生:模型在 28.1% 的機會中錯誤地保持已授權、證據清除的工作,並在 1.0% 的情況下錯誤地允許不安全的工作。最困難的情況是風險已解決的提交,其中簽署或結構化證據已經清除了實際風險觸發器,並且模型在著名事件的證據反向鏡像上得分明顯低於事件本身(63.8% 對 98.5%)。一般能力與引導校準並不相同:高能力模型在提交邊界上往往過度拒絕,而更多的推理可以修復一個弱閘,同時保持一個已校準的閘平坦。公共排行榜位於 steerbench.com。
Algorithm Design and Physician Liability
2608.13618v1 by Shujie Luan, Shubhranshu Singh, Tinglong Dai
A single clinical algorithm can deliver unequal accuracy across patient groups, and concern about such disparity has grown as artificial intelligence (AI) spreads through clinical decision-making. In response, a liability rule introduced in the United States holds healthcare providers responsible when their reliance on disparate algorithms contributes to erroneous clinical decisions. We examine how such liability considerations reshape (i) an AI firm's algorithm design decisions that drive group-specific accuracy and (ii) a physician's decisions to use AI in healthcare delivery. The AI firm designs an algorithm for two patient groups, and improving accuracy for the disadvantaged group is more costly. The physician (who remains the accountable decision-maker) then decides whether to consult AI, weighing the reduction in clinical uncertainty against expected liability exposure when AI errors disproportionately affect the disadvantaged group. We find the liability rule can induce disparate use of AI: the physician may reduce AI use overall and, over an intermediate range of liability, rely on AI less for disadvantaged patients. The effect is non-monotone. As liability increases, the physician's use of AI for disadvantaged patients first declines, then rises as the firm reallocates investment toward reducing disparity or switches to an equal-accuracy design. Mandating equal algorithmic accuracy across patient groups can then inadvertently harm both groups, because a uniform accuracy requirement distorts the firm's investment incentives and the physician's equilibrium AI-use decisions.
摘要:單一的臨床演算法在不同患者群體中可能會產生不均等的準確性,隨著人工智慧(AI)在臨床決策中的普及,對於這種差異的關注也日益增加。作為回應,美國引入了一項責任規則,當醫療提供者依賴不同的演算法導致錯誤的臨床決策時,將其負責。我們研究這種責任考量如何重塑(i)AI公司的演算法設計決策,促進特定群體的準確性,以及(ii)醫生在醫療提供中使用AI的決策。AI公司為兩個患者群體設計了一個演算法,改善弱勢群體的準確性成本更高。然後,醫生(仍然是負責的決策者)決定是否諮詢AI,權衡臨床不確定性的減少與當AI錯誤不成比例地影響弱勢群體時的預期責任風險。我們發現責任規則可能會導致AI的使用不均等:醫生可能會整體減少AI的使用,並且在責任的中等範圍內,對弱勢患者的AI依賴程度降低。這一效果是非單調的。隨著責任的增加,醫生對弱勢患者使用AI的情況最初下降,然後隨著公司將投資重新分配到減少差異或轉向平等準確性設計而上升。要求在患者群體之間達到平等的演算法準確性,可能會無意中對兩個群體造成傷害,因為統一的準確性要求扭曲了公司的投資激勵和醫生的均衡AI使用決策。
Auditable agentic AI for evidence-grounded thyroid ultrasound diagnosis and reporting
2608.12590v1 by Haifan Gong, Shiyu Chen, Bodong Wang, Yuqi Wang, Shijie Wang, Guoliang You, Xinyu Xiong, Haowei Wang, Mingzhi Mao, Dexing Kong, Qinghua Liu, Wei Lou, Fei Chen, Guanbin Li
Thyroid ultrasound diagnosis requires coordinated lesion localization, measurement, risk stratification and reporting, yet most AI systems address these tasks in isolation and provide limited support for clinical review. We present ThyroidXAgent, a clinician-interactive agentic AI system that coordinates specialized diagnostic tools and stores their outputs as an auditable case-level evidence record. The system was developed using OpenThyroidDB, a multicentre, multitask resource integrating approximately 0.3 million ultrasound images and 24,000 paired reports, and was evaluated on 28,458 non-overlapping test cases, including 8,721 cases from 35 centres in the private NHC-MISD-TUS cohort. Across heterogeneous datasets, ThyroidXAgent achieved a mean Dice score of 87.21 percent for nodule segmentation and a mean AUROC of 0.9466 for benign-malignant classification. The same workflow supported lymph-node metastasis prediction and follicular versus papillary thyroid carcinoma classification, with AUROCs of 0.864 and 0.805, respectively. For report generation, evidence-grounded assembly outperformed multimodal language-model baselines across three cohorts. ThyClinScore, a lesion-level clinical semantic metric introduced here, showed the strongest correlation with a location-aware language-model judge. ThyroidXAgent improved physician classification accuracy, increased report diagnostic consistency from 70.3 percent to 86.2 percent, and reduced segmentation and reporting time by 35.9 percent and 27.4 percent, respectively. These findings support auditable, clinician-correctable agentic AI for thyroid ultrasound diagnosis and reporting.
摘要:甲狀腺超聲診斷需要協調病變定位、測量、風險分層和報告,但大多數人工智慧系統在孤立的情況下處理這些任務,並對臨床審查提供有限的支持。我們提出了ThyroidXAgent,一個臨床互動的代理人工智慧系統,協調專門的診斷工具並將其輸出存儲為可審計的案例級證據記錄。該系統是使用OpenThyroidDB開發的,這是一個多中心、多任務的資源,整合了約30萬張超聲圖像和24,000份配對報告,並在28,458個不重疊的測試案例上進行了評估,包括來自私立NHC-MISD-TUS隊列的35個中心的8,721個案例。在異質數據集上,ThyroidXAgent在結節分割方面達到了87.21%的平均Dice分數,並在良惡性分類方面達到了0.9466的平均AUROC。同一工作流程支持淋巴結轉移預測和濾泡型與乳頭狀甲狀腺癌的分類,AUROC分別為0.864和0.805。在報告生成方面,基於證據的組合在三個隊列中超越了多模態語言模型基準。這裡引入的ThyClinScore,一個病變級的臨床語義指標,顯示出與位置感知語言模型評審者之間的最強相關性。ThyroidXAgent提高了醫生的分類準確性,將報告的診斷一致性從70.3%提高到86.2%,並分別減少了35.9%和27.4%的分割和報告時間。這些發現支持可審計、可由臨床醫生修正的代理人工智慧用於甲狀腺超聲診斷和報告。
How Compliant is Sepsis Treatment? An Expert-Guided Neuro-symbolic Pipeline for Generating Clinical Compliance Insights
2608.13617v1 by Himanshu Tripathi, Kaushik Roy, Subash Neupane, Shahram Rahimi
Verifying whether clinical care follows evidence-based protocols is a natural neuro-symbolic problem, yet the safety-critical setting defeats either paradigm alone. We present an expert-guided pipeline that constrains a large language model strictly to semantic normalization, mapping messy drug and microbiology strings onto a fixed clinical vocabulary, while a Sugeno fuzzy inference system reasons over the normalized events. The fuzzy layer encodes eight Surviving Sepsis Campaign bundle rules and replaces binary judgments with graded scores in [0,1]. Applied to 2,438 MIMIC-IV v3.1 sepsis episodes, it surfaces antibiotic timing as the most critical breakdown (mean 0.24, 13% within one hour), Hour-1 underperformance (mean 36.7%), a 51% elevated-lactate drop-off, and descriptive differences in ICU stay across compliance groups (3.8 versus 5.1 days).
摘要:驗證臨床護理是否遵循基於證據的協議是一個自然的神經符號問題,但安全關鍵的環境使得單一範式無法應對。我們提出了一個專家指導的流程,將大型語言模型嚴格限制於語義標準化,將雜亂的藥物和微生物學字符串映射到固定的臨床詞彙上,同時一個Sugeno模糊推理系統對標準化事件進行推理。模糊層編碼了八條存活敗血症運動的捆綁規則,並用[0,1]範圍內的分數取代了二元判斷。應用於2,438個MIMIC-IV v3.1敗血症事件中,它顯示抗生素使用時機是最關鍵的破綻(平均0.24,13%在一小時內),第一小時表現不佳(平均36.7%),51%的乳酸升高下降,以及在合規性組之間ICU住院天數的描述性差異(3.8天對5.1天)。
M-Net: Integrating Spectral Features and Physical Field Operators into Deep Learning for Medical Image Segmentation
2608.12196v1 by Jing Zhu, Ye Wang, Fumin Wang
Purpose: Deep learning-based medical image segmentation has achieved remarkable success, yet purely data-driven approaches often fail to exploit the rich mathematical structure inherent in medical images. We investigate whether explicit mathematical inductive biases, specifically matrix spectral analysis and vector calculus operators, can enhance segmentation beyond data-driven learning alone. Methods: We propose M-Net (Math-Augmented Network), which integrates three complementary mathematical priors into U-Net: (1) continuous spectral features derived from the condition number of centered local pixel matrices, providing a differentiable measure of texture ill-conditioning; (2) physical field operators (divergence and a discrete curl-like boundary irregularity operator) computed from image gradient fields, capturing focal intensity extrema and edge non-smoothness; and (3) a Math-Attention Gate (MAG) that adaptively fuses mathematical features with CNN-extracted deep features at skip connections. Results: Experiments on three benchmarks (LiTS, KiTS, and BraTS) show that M-Net achieves Dice scores of 78.42%, 76.15%, and 83.67%, outperforming baseline U-Net by 12.37%, 3.52%, and 5.55% on liver, kidney, and brain tumor segmentation, respectively. Ablations reveal that the condition-number feature contributes a 2.14% gain over binary invertibility features, while MAG adds 1.45% over simple concatenation. Conclusion: M-Net establishes that mathematical inductive biases provide effective complementary information for medical image segmentation. The continuous condition-number feature offers superior gradient information over discrete alternatives, and MAG preserves these priors throughout the network. This work opens avenues for integrating linear algebra and vector calculus into deep architectures for medical imaging.
摘要:目的:基於深度學習的醫學影像分割已取得顯著成功,然而純數據驅動的方法往往未能充分利用醫學影像中固有的豐富數學結構。我們探討明確的數學歸納偏差,特別是矩陣譜分析和向量微積分運算子,是否能超越僅依賴數據驅動學習來增強分割效果。
方法:我們提出M-Net(數學增強網絡),將三個互補的數學先驗整合到U-Net中:(1)從中心局部像素矩陣的條件數導出的連續譜特徵,提供可微分的紋理不良條件度量;(2)從影像梯度場計算的物理場運算子(散度和離散的旋度邊界不規則性運算子),捕捉焦點強度極值和邊緣不光滑性;(3)一個數學注意力閘(MAG),在跳躍連接中自適應地融合數學特徵與CNN提取的深度特徵。
結果:在三個基準測試(LiTS、KiTS和BraTS)上的實驗顯示,M-Net在肝臟、腎臟和腦腫瘤分割中分別達到78.42%、76.15%和83.67%的Dice分數,分別比基線U-Net高出12.37%、3.52%和5.55%。消融實驗顯示,條件數特徵比二元可逆性特徵貢獻了2.14%的增益,而MAG比簡單的串接增加了1.45%。
結論:M-Net證明數學歸納偏差為醫學影像分割提供了有效的互補信息。連續的條件數特徵提供了比離散替代方案更優越的梯度信息,而MAG在整個網絡中保留了這些先驗。這項工作為將線性代數和向量微積分整合到醫學影像的深度架構中開辟了新的途徑。
A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench
2608.12138v1 by Praveen Reddy, Charuta Mandke, Suvrankar Datta, Sarah Khan, Siddharth Reddy Anthireddy, Shitij Arora, Vishal Singh
General-purpose large language models (LLMs) have recently been reported to match or exceed specialized clinical AI tools on medical benchmarks, but such comparisons draw on a narrow set of systems and on benchmarks developed largely in high-income settings. We evaluate VITA, a retrieval-augmented generation (RAG) system purpose-built for contextual knowledge retrieval in India and other low- and middle-income (LMIC) settings. VITA retrieves from a curated corpus of disease-specific guidelines, India-specific antimicrobial resistance data, national formulary constraints, and resource-limited care protocols; its architecture and corpus are proprietary, but the benchmark, the physician-written rubrics, and our full response and scoring outputs are public for independent verification. On 4,023 English-language HealthBench questions (80.5% of the benchmark), scored with a GPT-4.1 judge, VITA ranked first with 51.9% of possible rubric points, ahead of GPT-5.4 (46.1%), o4-mini (44.3%), Gemini 3.1 Pro (42.6%), and Claude Sonnet 4.6 (37.3%), and scored highest on 45.4% of questions. To test robustness to newer models and judge lineage, a 500-question subset was re-run against current-generation models (GPT-5.5, Claude Opus 4.8, Gemini 3.5 Pro, Grok 4.3) and graded by a neutral open-weight judge (DeepSeek-V4-Pro) sharing no lineage with any system tested. Here the gap narrowed to parity: VITA and GPT-5.5 were statistically indistinguishable on mean per-question score, while VITA led on points-weighted score and won the most questions. VITA's advantages in accuracy and completeness persisted under the neutral judge; its communication scores were lower. These results indicate that a purpose-built clinical RAG system remains competitive with frontier LLMs on an open benchmark, consistent with corpus specificity as a design variable that improves grounding at some cost to communication polish.
摘要:一般用途的大型語言模型(LLMs)最近被報導在醫療基準上與專門的臨床人工智慧工具相匹配或超越,但這些比較依賴於一組狹窄的系統以及主要在高收入環境中開發的基準。我們評估了VITA,一個專為印度及其他低收入和中等收入(LMIC)環境中的上下文知識檢索而設計的檢索增強生成(RAG)系統。VITA從一個策劃的特定疾病指導方針、印度特定的抗微生物抗藥性數據、國家藥典限制以及資源有限的護理協議中檢索資料;其架構和語料庫是專有的,但基準、醫生撰寫的評分標準以及我們的完整回應和評分輸出是公開的,以便獨立驗證。在4,023個英語HealthBench問題(基準的80.5%)上,使用GPT-4.1評判,VITA以51.9%的可能評分點排名第一,超過了GPT-5.4(46.1%)、o4-mini(44.3%)、Gemini 3.1 Pro(42.6%)和Claude Sonnet 4.6(37.3%),並在45.4%的問題上得分最高。為了測試對新模型的穩健性和評判系譜,對500個問題的子集再次運行,與當前一代模型(GPT-5.5、Claude Opus 4.8、Gemini 3.5 Pro、Grok 4.3)進行比較,並由一位中立的開放權重評判(DeepSeek-V4-Pro)進行評分,該評判與任何測試系統無關。此時差距縮小至平行:VITA和GPT-5.5在每題平均得分上統計上無法區分,而VITA在加權得分上領先並贏得了最多問題。VITA在準確性和完整性上的優勢在中立評判下持續存在;其溝通得分較低。這些結果表明,專為臨床設計的RAG系統在公開基準上仍然與前沿LLMs具有競爭力,這與語料庫的特異性作為設計變量相一致,該變量在某種程度上提高了基礎性,但對溝通的精緻性造成了成本。
Do LLMs Take Care of Their Own? Similarity Signals Can Induce Cooperation
2608.12125v1 by Akash Kundu, Emanuel Tewolde, Ratip Emin Berker, Samuel F. Brown, Vincent Conitzer
As LLM-based agents with user-instructed goals are becoming widely deployed, they increasingly encounter each other in strategic interactions, and face challenges of finding mutually beneficial outcomes. Prior literature has argued that cooperation problems such as the Prisoner's Dilemma are resolvable in settings where agents know they follow very similar decision making patterns, as for example in monocultural AI ecosystems. Following that line of work, this paper introduces the first framework for evaluating LLM decision making when agents are provided with graded similarity signals. Among our findings, we establish that different LLM models vary drastically in how they navigate similarity signals, with some modern models showing consistent behavior across cooperation problems, payoff structures, and prompt framing. Perhaps surprisingly, our experiments also show that the dataset based on which the similarity signal is computed has small to no impact on induced cooperation, and that LLM models systematically self-identify as highly similar when asked to evaluate another model's chain-of-thought reasoning by themselves. Finally, we develop an LLM-behavioral-game-theoretic model that captures some of their reasoning rationale, and show that it can support cooperative outcomes in equilibrium under sufficiently high similarity scores.
摘要:隨著基於大型語言模型(LLM)的代理人以用戶指導的目標被廣泛部署,它們在戰略互動中越來越多地相遇,並面臨尋找互利結果的挑戰。先前的文獻已經論證,合作問題如囚徒困境在代理人知道它們遵循非常相似的決策模式的情況下是可以解決的,例如在單一文化的人工智慧生態系統中。沿著這一研究方向,本文介紹了第一個評估LLM決策制定的框架,當代理人被提供分級相似性信號時。
在我們的發現中,我們確立了不同的LLM模型在如何導航相似性信號方面存在巨大差異,一些現代模型在合作問題、收益結構和提示框架中顯示出一致的行為。或許令人驚訝的是,我們的實驗還顯示,計算相似性信號的數據集對於誘發合作的影響微乎其微,且當被要求自行評估另一模型的思考鏈推理時,LLM模型系統性地自我識別為高度相似。最後,我們開發了一個LLM行為博弈論模型,捕捉它們的一些推理理由,並顯示它可以在足夠高的相似性分數下支持均衡的合作結果。
How Far from Clinical Deployment? Evaluating the Complete Unsupervised Domain Adaptation Pipeline in Medical Imaging
2608.12035v1 by Yiheng Xiong, Luisa Gallée, Daniel Santak Wolf, Heiko Hillenhagen, Michael Götz
Deploying unsupervised domain adaptation (UDA) in clinical practice requires choosing which algorithm to use and which of its trained models to ship. However, the deployment (target) domain is unlabeled, so models cannot be evaluated directly on it, leaving it unclear which to select. We address this by evaluating the complete UDA pipeline, considering both adaptation and label-free selection together. Our study covers eleven clinically relevant cross-domain scenarios from nine medical imaging datasets, with ten UDA algorithms and 13 label-free selection methods (validators), evaluating over 80,000 trained models in total. By this, we find that a capable adapted model usually exists, but identifying it without target labels is difficult: the validator-selected models leave a large and structural target performance gap to the best available one, with no evaluated validator consistently reliable. Towards closing it, we explore two strategies, ensembling and a small target-labeling budget; both narrow this gap but do not close it entirely. Overall, deployable UDA depends on the complete pipeline; addressing the less explored selection step could bring much of current UDA closer to clinical use.
摘要:在臨床實踐中部署無監督領域適應(UDA)需要選擇使用哪種算法以及其訓練模型中的哪一個進行部署。
然而,部署(目標)領域是未標記的,因此無法直接對其進行模型評估,這使得選擇變得不明確。
我們通過評估完整的 UDA 流程來解決這個問題,同時考慮適應和無標籤選擇。
我們的研究涵蓋了來自九個醫學影像數據集的十一個臨床相關的跨領域場景,使用十種 UDA 算法和 13 種無標籤選擇方法(驗證器),總共評估了超過 80,000 個訓練模型。
由此,我們發現通常存在一個能夠適應的模型,但在沒有目標標籤的情況下識別它是困難的:驗證器選擇的模型與可用的最佳模型之間存在著較大且結構性的目標性能差距,且沒有一個評估過的驗證器是一致可靠的。
為了縮小這個差距,我們探索了兩種策略,集成和小型目標標記預算;這兩者都縮小了這個差距,但並未完全關閉它。
總的來說,可部署的 UDA 依賴於完整的流程;解決較少探索的選擇步驟可能會使當前的 UDA 更接近臨床應用。
From Safety Documentation to Safety Knowledge Support: An Evidence-Grounded LLM Framework for Medical Devices
2608.12025v1 by Tuhinangshu Gangopadhyay, Rasmus Adler, Peter Liggesmeyer, Jan Reich
Medical devices are becoming more software-intensive, connected, and AI-enabled. Their development requires risk-management evidence aligned with ISO 14971 and, for software, IEC 62304. This evidence must be kept consistent across requirements, design decisions, software changes, verification results, complaints, and post-market data. These tasks are costly and depend on scarce safety and domain experts. Large language models (LLMs) may reduce parts of this effort because medical-device safety work is highly document-based. However, current LLM-based safety-engineering studies often address isolated methods, rely on generic prompting or public examples, and provide limited support for source links, traceability, uncertainty handling, lifecycle updates, and recorded expert review. This limits their use in regulated medical-device development. This paper argues that the central research problem is not safety-text generation, but source-linked safety-knowledge support. We propose an evidence-grounded framework that connects device artifacts, controlled knowledge storage and retrieval, method-specific generation of candidate safety items, critique and uncertainty checks, and recorded expert review. The framework prepares, links, checks, and updates candidate safety artifacts for expert decision-making. It does not decide whether a device is safe and does not provide regulatory approval. We also outline an evaluation strategy using non-public or newly built medical-device case studies and expert reference analyses to assess coverage, correctness, relevance, traceability, duplicate rate, unsupported claims, and review effort.
摘要:醫療器材正變得越來越依賴軟體、互聯網連接和人工智慧。它們的開發需要符合ISO 14971的風險管理證據,對於軟體則需要符合IEC 62304。這些證據必須在需求、設計決策、軟體變更、驗證結果、投訴和市場後數據之間保持一致。這些任務成本高昂,並依賴於稀缺的安全和領域專家。
大型語言模型(LLMs)可能會減少這部分工作,因為醫療器材的安全工作高度依賴文檔。然而,目前基於LLM的安全工程研究往往針對孤立的方法,依賴於通用提示或公共範例,並對來源鏈接、可追溯性、不確定性處理、生命周期更新和記錄的專家審查提供有限支持。這限制了它們在受監管的醫療器材開發中的應用。
本文主張,核心研究問題不是安全文本生成,而是來源鏈接的安全知識支持。我們提出了一個基於證據的框架,連接設備文檔、受控知識存儲和檢索、特定方法生成候選安全項目、批評和不確定性檢查,以及記錄的專家審查。該框架為專家決策準備、鏈接、檢查和更新候選安全文檔。它不決定設備是否安全,也不提供監管批准。我們還概述了一個評估策略,使用非公開或新建的醫療器材案例研究和專家參考分析來評估覆蓋範圍、正確性、相關性、可追溯性、重複率、不支持的聲明和審查工作量。
When the API Speaks the Wrong Language: Revisiting Post-Training for Multilingual Tool Use
2608.11715v1 by Siddharth Chauhan, Thomas Butler, Abhishek Singhania, Pankaj Porwal, Honey Gupta
The reliability of Large Language Models (LLMs) for API calling degrades in multilingual settings. A common failure occurs when a model selects the correct tool but generates argument values in an inconsistent language, which we term Argument Language Mismatch (ALM). Although semantically correct, such outputs are operationally invalid and not captured by standard API-calling metrics. We revisit post-training strategies for mitigating ALM and find that, in our benchmark, supervised fine-tuning (SFT) provides a strong baseline, substantially improving argument language consistency and end-to-end function call accuracy. Under consistent model selection, SFT achieves performance comparable to, and sometimes exceeding more complex reinforcement learning (RL) approaches. We further examine whether RL with structured, argument-aware rewards offers additional benefits. While methods such as Group Relative Policy Optimization (GRPO) can improve language consistency and better preserve general reasoning ability, these gains are incremental and most pronounced in generalization and multi-objective trade-offs. Overall, our results suggest that much of the performance in multilingual API grounding can be achieved through careful supervised training, with RL providing targeted rather than fundamental improvements.
摘要:大型語言模型(LLMs)在多語言環境中進行 API 呼叫的可靠性會下降。常見的失敗情況是模型選擇了正確的工具,但生成的參數值卻使用不一致的語言,我們稱之為參數語言不匹配(ALM)。雖然這些輸出在語義上是正確的,但在操作上是無效的,並未被標準的 API 呼叫指標所捕捉。我們重新檢視了減輕 ALM 的後訓練策略,發現根據我們的基準,監督微調(SFT)提供了強有力的基線,顯著改善了參數語言的一致性和端到端函數調用的準確性。在一致的模型選擇下,SFT 的表現可與有時超越更複雜的強化學習(RL)方法相媲美。我們進一步檢查了結構化的、參數感知的獎勵是否提供了額外的好處。雖然像群體相對政策優化(GRPO)這樣的方法可以改善語言一致性並更好地保留一般推理能力,但這些增益是漸進的,並且在泛化和多目標權衡中最為明顯。總體而言,我們的結果表明,多語言 API 基礎的性能大部分可以通過謹慎的監督訓練來實現,而 RL 則提供了針對性的而非根本性的改進。
Rubric Dropout: A Simple Way to Mitigate Reward Hacking in Rubric-as-Reward RL
2608.11669v1 by Minglai Yang, Xinyu Guo, Utkarsh Tyagi, Mian Zhang, Razvan Dumitru, Sunjie Hou, Yunzhong He, Daniel Yue Zhang, Ying Liu
Reinforcement learning against rubrics, lists of criteria graded by an LLM judge, has become a standard way to post-train language models on tasks with no deterministic answer. The rubric, however, is a fixed proxy for quality, never a complete description of it, and a policy trained against it long enough will learn to exploit the difference. We measure this directly. Training Qwen3-8B with Group Relative Policy Optimization (GRPO) on medical and science rubrics and grading out-of-distribution (OOD) benchmarks with both the training judge and a stronger gold judge, we find that the two scores diverge during training. The training judge's score keeps climbing while the gold judge's score peaks and then falls, by 3 points on HealthBench-Hard and by 22 points on ResearchQA. A judge with a fixed bias would shift the gold curve by a constant, not send it down while the training score rises, so the divergence is reward hacking, not judge noise. We propose Rubric Dropout, a one-line fix borrowed from neuron dropout. At every step, we randomly drop a subset of the rubric's criteria before computing the reward, so the policy never optimizes the same rubric twice. The dropped subset is shared across each rollout group, so GRPO's group-relative advantages stay comparable, and evaluation always uses the full rubric. Comparing no dropout against dropout at 30% and 50% on both benchmark pairs, dropout raises the OOD gold score at every matched checkpoint (+1 to +2 points on HealthBench-Hard, +6 to +7 points on ResearchQA), lowers the two hacking measures we track, and costs nothing in domain. Sweeping the dropout fraction shows a broad 30-50% sweet spot, while the natural alternative, reweighting criteria by how useful they are to training, performs worse than no intervention at all in our setting.
摘要:強化學習對抗評分標準,即由LLM評審評分的標準列表,已成為對於沒有確定性答案的任務進行後訓練語言模型的標準方法。然而,評分標準是質量的固定代理,從來不是其完整描述,而對其訓練足夠長的策略將學會利用這一差異。我們直接測量這一點。使用群體相對策略優化(GRPO)對醫學和科學評分標準進行訓練Qwen3-8B,並用訓練評審和更強的金標準評審對分佈外(OOD)基準進行評分,我們發現這兩個分數在訓練過程中出現了分歧。訓練評審的分數持續上升,而金標準評審的分數達到峰值後下降,在HealthBench-Hard上下降了3分,在ResearchQA上下降了22分。具有固定偏見的評審會將金曲線平移一個常數,而不是在訓練分數上升時將其向下移動,因此這一分歧是獎勵黑客行為,而不是評審噪聲。我們提出了評分標準隨機失活(Rubric Dropout),這是一個借用自神經元隨機失活的一行修正。在每一步,我們在計算獎勵之前隨機丟棄評分標準的一部分標準,因此該策略從不對同一評分標準進行兩次優化。被丟棄的子集在每個展開組中共享,因此GRPO的群體相對優勢保持可比,評估始終使用完整的評分標準。將無隨機失活與在兩組基準上30%和50%的隨機失活進行比較,隨機失活在每個匹配的檢查點上提高了OOD金分數(在HealthBench-Hard上提高了1到2分,在ResearchQA上提高了6到7分),降低了我們跟踪的兩個黑客行為指標,且在領域上沒有任何成本。掃描隨機失活比例顯示出30-50%的廣泛甜蜜點,而自然的替代方案,即根據標準對訓練的有用性進行重加權,在我們的設置中表現得比沒有任何干預更差。
Beyond Single-Turn Confidence: Trajectory-Adapted Uncertainty Quantification for LLM Agents
2608.11552v1 by Dylan Bouchard, Mohit Singh Chauhan
Uncertainty quantification (UQ) methods for language models are typically evaluated on single-turn outputs, where uncertainty is attached to one generated answer. For LLM agents, however, the unit of observation is an interactive trajectory, where the model can ask clarifying questions, call tools, update state, and make intermediate decisions whose errors propagate to the final outcome. We study whether three common families of single-turn UQ methods transfer to this setting. Across five LLMs and four multi-turn tool-use datasets from BFCL-v4 and $τ^2$-bench, we evaluate white-box scorers based on action-token probabilities, black-box consistency scorers based on resampled trajectories, and reflexive scorers based on model self-assessment of the trajectory. We find that transfer is often useful but uneven. Token-probability scores are highly sensitive to the choice of aggregator used across turns, reflexive scores provide the strongest low-cost baseline in most evaluated settings, and black-box self-consistency is often the strongest UQ family, with trajectory-equivalence and action-set consistency typically ranking highest among its variants. These results suggest that UQ methods developed for single generations should be revalidated at the trajectory level, with careful attention to the consistency measurement, aggregator choice, and computational budget.
摘要:不確定性量化(UQ)方法對於語言模型通常是在單輪輸出上進行評估,其中不確定性與一個生成的答案相關聯。然而,對於大型語言模型(LLM)代理來說,觀察的單位是一個互動軌跡,其中模型可以提出澄清問題、調用工具、更新狀態,並做出中間決策,這些決策的錯誤會傳播到最終結果。我們研究了三種常見的單輪UQ方法是否能轉移到這種情境中。在五個LLM和來自BFCL-v4和$τ^2$-bench的四個多輪工具使用數據集上,我們評估了基於行動標記概率的白盒評分器、基於重新抽樣軌跡的黑盒一致性評分器,以及基於模型自我評估軌跡的反射性評分器。我們發現轉移通常是有用的,但不均勻。標記概率分數對於跨輪使用的聚合器選擇非常敏感,反射性分數在大多數評估設置中提供了最強的低成本基準,而黑盒自一致性通常是最強的UQ家族,其中軌跡等價性和行動集一致性通常在其變體中排名最高。這些結果表明,為單次生成開發的UQ方法應該在軌跡層面重新驗證,並仔細關注一致性測量、聚合器選擇和計算預算。
Social Chain of Thought: A Multi-Agent Architecture Grounded in Medical Differential Diagnosis Methodology
2608.11420v1 by Del Coburn, Scott Sanner, Dan Silver
Medical diagnostic reasoning is a high-impact use case for LLMs that carries significant implications for the health and wellbeing of users. When OpenAI (2026) reports that more than 5% of ChatGPT messages globally are healthcare-related, the transparency of these systems becomes a serious design concern. This is especially true for complex cases, where differential diagnosis often requires integrating multiple forms of specialist reasoning. Existing work has proposed multi-agent approaches to medical diagnosis, but it remains unclear when such systems are needed, why they help, and where they outperform monolithic inference. We introduce Social Chain of Thought (SCoT),a multi-round pipeline for medical differential diagnosis that structures multi-agent interaction as a deliberative framework for collabora. tive LLM reasoning. Evaluating SCoT against single-agent baselines, one-agent pipeline ablations, and best-of-n scaling, we show that its recall advantage is not reproduced by monolithic inference alone. SCoT is most successful in the hardest diagnostic cases, where multiple rounds of specialist conversation help recover ground-truth diagnoses and converge on a higher-recall differential.
摘要:醫療診斷推理是大型語言模型(LLMs)的高影響力應用案例,對用戶的健康和福祉具有重要意義。當 OpenAI(2026)報告全球超過 5% 的 ChatGPT 訊息與醫療保健相關時,這些系統的透明度成為一個嚴重的設計問題。這在複雜案例中特別真實,因為鑑別診斷通常需要整合多種專家推理形式。現有的研究已提出多代理的醫療診斷方法,但仍不清楚何時需要這樣的系統、它們為何有幫助,以及它們在哪些方面超越單一推理。我們介紹社會思維鏈(Social Chain of Thought, SCoT),這是一個多輪的醫療鑑別診斷管道,將多代理互動結構化為協作大型語言模型推理的深思框架。通過將 SCoT 與單代理基準、單代理管道消融和最佳擴展進行評估,我們顯示其召回優勢並非僅由單一推理所重現。SCoT 在最困難的診斷案例中最為成功,多輪專家對話有助於恢復真實診斷並收斂於更高召回率的鑑別診斷。
Gaze Target Estimation Anywhere with Concepts
2608.11367v1 by Xu Cao, Houze Yang, Vipin Gunda, Zhongyi Zhou, Tianyu Xu, Adarsh Kowdle, Inki Kim, James M. Rehg
Estimating human gaze targets from images in-the-wild is an important and formidable task. Existing approaches primarily employ brittle, multi-stage pipelines that require explicit inputs, like head bounding boxes and human pose, in order to identify the subject of gaze analysis. As a result, detection errors can cascade and lead to failure. Moreover, these prior works lack the flexibility of specifying the gaze analysis task via natural language prompting, an approach which has been shown to have significant benefits in convenience and scalability for other image analysis tasks. To overcome these limitations, we introduce the Promptable Gaze Target Estimation (PGE) task, a new end-to-end, concept-driven paradigm for gaze analysis. PGE conditions gaze prediction on flexible user text or visual prompts (e.g., "the boy in the red shirt" or "person in point [0.52, 0.48]") to identify a specific subject for gaze analysis. This approach integrates subject localization with gaze estimation, and eliminates the rigid dependency on intermediate analysis stages. We develop a scalable data engine to generate Gaze-Co (Gaze Estimation with Concepts), a dataset and benchmark of 120K high-quality, prompt-annotated image pairs. We also propose GazeAnywhere, the first model designed for PGE. GazeAnywhere uses a transformer-based detector to fuse features from frozen encoders and simultaneously solves subject localization, in/out-of-frame presence, and gaze target heatmap estimation. GazeAnywhere achieves state-of-the-art performance on multiple PGE benchmarks, setting a strong baseline for this new problem even on a difficult out-of-domain, real-world clinical dataset. GazeAnywhere is open-sourced in github.com/IrohXu/GazeAnywhere.
摘要:估計野外圖像中的人類注視目標是一項重要且艱巨的任務。現有的方法主要採用脆弱的多階段管道,這需要明確的輸入,例如頭部邊界框和人體姿勢,以識別注視分析的主體。因此,檢測錯誤可能會級聯並導致失敗。此外,這些先前的工作缺乏通過自然語言提示來指定注視分析任務的靈活性,這種方法已被證明在其他圖像分析任務中具有顯著的便利性和可擴展性。為了克服這些限制,我們引入了可提示的注視目標估計(PGE)任務,這是一種新的端到端、以概念為驅動的注視分析範式。PGE基於靈活的用戶文本或視覺提示(例如,“穿紅色襯衫的男孩”或“位於點[0.52, 0.48]的人”)來識別特定的注視分析主體。這種方法將主體定位與注視估計相結合,並消除了對中間分析階段的僵硬依賴。我們開發了一個可擴展的數據引擎來生成Gaze-Co(概念驅動的注視估計),這是一個包含120K高質量、提示標註圖像對的數據集和基準。我們還提出了GazeAnywhere,這是第一個為PGE設計的模型。GazeAnywhere使用基於Transformer的檢測器來融合來自凍結編碼器的特徵,並同時解決主體定位、框內/框外存在性和注視目標熱圖估計。GazeAnywhere在多個PGE基準上達到了最先進的性能,即使在困難的域外現實臨床數據集上也設置了這一新問題的強基線。GazeAnywhere已在github.com/IrohXu/GazeAnywhere上開源。
Dual-Domain Cross-Modal Decoding for Clinical Text-Guided Medical Image Segmentation
2608.11335v1 by Md Maklachur Rahman, Tracy Hammond
Clinical text can narrow down what to segment, but recent text-guided designs emphasize spatial alignment while overlooking frequency content that governs texture and boundaries. We propose Dual-Domain Cross-Modal Decoding (DD-CMD) for clinical text-guided pulmonary infection segmentation, integrating two complementary forms of language guidance during decoding. In the spatial domain, Text-Guided Spatial Cross-Attention (TGSA) aligns multi-scale visual tokens with text semantics and updates features through gated residual fusion. In the frequency domain, Spectral-Text Adaptive Modulation (STAM) applies a 2D DCT to compute learnable band-energy statistics and predicts text-conditioned FiLM parameters to recalibrate decoder channels for frequency-aware decoding. DD-CMD embeds TGSA and STAM into a coarse-to-fine decoder (7x7 to 56x56) and restores full-resolution masks using a lightweight two-stage refinement module. Experiments on QaTa-COV19 and MosMedData+ show that DD-CMD achieves 91.46% Dice / 84.26% mIoU and 81.95% Dice / 69.42% mIoU, respectively, with average gains of +1.96 Dice and +2.67 mIoU over the strongest prior baselines. Code: https://github.com/maklachur/DD-CMD.
摘要:臨床文本可以縮小分割範圍,但最近的文本引導設計強調空間對齊,同時忽略了控制紋理和邊界的頻率內容。我們提出了雙域跨模態解碼(DD-CMD)用於臨床文本引導的肺部感染分割,在解碼過程中整合兩種互補的語言引導形式。在空間域中,文本引導的空間跨注意力(TGSA)將多尺度視覺標記與文本語義對齊,並通過門控殘差融合更新特徵。在頻率域中,光譜文本自適應調製(STAM)應用2D DCT來計算可學習的帶能量統計,並預測文本條件的FiLM參數,以重新校準解碼器通道以進行頻率感知解碼。DD-CMD將TGSA和STAM嵌入到一個粗到細的解碼器(從7x7到56x56),並使用輕量級的兩階段精煉模塊恢復全分辨率的掩膜。在QaTa-COV19和MosMedData+上的實驗顯示,DD-CMD分別達到91.46%的Dice / 84.26%的mIoU和81.95%的Dice / 69.42%的mIoU,平均增益為+1.96 Dice和+2.67 mIoU,相較於最強的先前基準。代碼:https://github.com/maklachur/DD-CMD。
3D Weighted Geometric Graph Neural Networks for Sheep Facial Pain Assessment
2608.11050v1 by Alam Noor, Luis Almeida, Mohamed Daoudi
Deep learning systems perform mainly within the 2D for a single image domain and take the face as a single-dimension representation, losing sight of the 3D anatomy of sheep and cross-landmark spatial relationships that are intrinsic to the clinically proven Sheep Pain Facial Expression Scale (SPFES). This paper presents the \textbf{3D Sheep Pain Facial Expression System (3D-SPFES)}, a novel, monocular depth-aware geometric graph neural network system that integrates each SPFES facial landmark, such as the ears, eyes, and nose, into 3D Euclidean space estimated from a single RGB camera by using VideoDepthAnything, thus preventing the need for specialized depth hardware. Each landmark node includes a feature vector containing its 3D spatial coordinates, estimated surface normal, and facial attribute class embedding. Edges linked to nodes are assigned weights based on an aggregate metric that combines both Euclidean distance and surface co-planarity in a 3D space. A Weighted Geometric Graph Neural Network (WG-GNN) studies this graph using $\mathcal{K} = 3$ geometry-aware message-passing layers enhanced by a scaled dot-product attention method that selectively enhances anatomically relevant inter-landmark messages. The resultant node embeddings are combined into $\mathcal{O} = 3$ pain-level clusters and integrated into a Normalized Pain Score (NPS) within the range of $[0, 100%]$ a confidence-weighted, SPFES-derived scoring method.
摘要:深度學習系統主要在單一影像的2D領域內運作,將面部視為單一維度的表示,忽略了羊的3D解剖結構和與臨床證明的羊痛面部表情量表(SPFES)固有的交叉標記空間關係。本文提出了\textbf{3D羊痛面部表情系統(3D-SPFES)},這是一種新穎的單目深度感知幾何圖神經網絡系統,將每個SPFES面部標記(如耳朵、眼睛和鼻子)整合到從單一RGB相機估算的3D歐幾里得空間中,藉此避免了專用深度硬體的需求。每個標記節點包含一個特徵向量,其中包含其3D空間坐標、估算的表面法向量和面部屬性類別嵌入。連接到節點的邊根據一個綜合指標分配權重,該指標結合了3D空間中的歐幾里得距離和表面共平面性。加權幾何圖神經網絡(WG-GNN)使用$\mathcal{K} = 3$幾何感知消息傳遞層來研究這個圖,並通過縮放的點積注意力方法增強與解剖相關的標記間消息。結果節點嵌入被合併為$\mathcal{O} = 3$疼痛級別集群,並整合到範圍為$[0, 100%]$的標準化疼痛分數(NPS)中,這是一種基於信心加權的、源自SPFES的評分方法。
CARE: Confidence-Aware Reasoning for Reliable Medical VQA
2608.10964v1 by Yuetian Du, Yucheng Wang, Zhenyuan Chen, Luyuan Chen, Rongyu Zhang, Jinjian Zhang, Wei Zhou, Zhijie Xu, Ming Kong, Zhan Zhou, Jie Liu, Qiang Zhu
Reinforcement Fine-Tuning (RFT) has enabled medical Multimodal Large Language Models (MLLMs) to produce Chain-of-Thought (CoT) reasoning for visual question answering, yet these models suffer from $\textit{confidence miscalibration}$---a systematic gap between expressed certainty and actual diagnostic accuracy that undermines clinical trust. We propose $\textbf{CARE}$, a $\textbf{C}$onfidence-$\textbf{A}$ware medical $\textbf{RE}$asoning framework that jointly optimizes accuracy and calibration through a dual-stage pipeline. First, a scalable Medical-CoT synthesis provides structured cold-start data for Supervised Fine-Tuning. Second, Group Relative Policy Optimization (GRPO) with a novel $\textbf{Confidence-Aware Reward (CAR)}$ mechanism ties the model's confidence to diagnostic correctness within the reward signal. Across three Medical VQA benchmarks, $\textbf{CARE}$ achieves the highest diagnostic accuracy while obtaining the lowest Expected Calibration Error and Hallucination Rate, establishing a foundation for trustworthy clinical decision support. Our code is available at https://github.com/anotherbricki/CARE.
摘要:強化微調(RFT)使醫療多模態大型語言模型(MLLMs)能夠為視覺問題回答產生思維鏈(CoT)推理,然而這些模型存在著$\textit{信心錯誤校準}$的問題——表達的確定性與實際診斷準確性之間的系統性差距,這削弱了臨床信任。我們提出了$\textbf{CARE}$,一個$\textbf{C}$onfidence-$\textbf{A}$ware醫療$\textbf{RE}$asoning框架,通過雙階段管道共同優化準確性和校準。首先,一個可擴展的Medical-CoT合成提供結構化的冷啟動數據以進行監督微調。其次,帶有新穎的$\textbf{Confidence-Aware Reward (CAR)}$機制的群體相對策略優化(GRPO)將模型的信心與獎勵信號中的診斷正確性相聯繫。在三個醫療VQA基準中,$\textbf{CARE}$實現了最高的診斷準確性,同時獲得了最低的期望校準誤差和幻覺率,為可信的臨床決策支持奠定了基礎。我們的代碼可在https://github.com/anotherbricki/CARE獲得。
ComBodied Agents: a New Paradigm of Human-Centric Agentic AI
2608.10915v2 by Qianggang Ding, Xingyao Wang, Rui Feng, Zhibin Wang, Feixiang Yao, Kelong Mao, Hao Sun, Zhiyao Luo, Jiankai Tang, Lei Li, Jiadong Guo, Minheng Ni, Weicong Lin, Chenxi Yang, Hongxiang Gao, Zhenghua Chen, Yang Bai, Min Wu, Jun Cheng, Huazhu Fu, Dacheng Tao, Bang Liu
After an older adult misses a medication dose, a software agent can send another reminder and an embodied agent can bring the medication. Yet neither explains whether the person forgot, is confused, has side effects, or deliberately refused, nor what support is appropriate. This reveals a structural gap in Agentic AI: Digital Agents primarily transform software states, while Embodied Agents transform physical states; neither makes a person's evolving state and agency the primary object of modeling, intervention, and evaluation. We introduce Combodied Agents, a human-centered paradigm that perceives, models, predicts, and supports individual human-state trajectories over time, using software tools, sensors, wearables, robots, and human services as action channels rather than end goals. We unify fragmented capabilities across personal assistants, health agents, AI companions, and adaptive human--AI systems into a closed loop: event-based multimodal perception reconstructs meaningful personal events; longitudinal, correctable memory provides temporal context; Personal World Models estimate future personal states and outcomes under alternative decisions and interventions; and an admissible intervention policy selects proportionate support under consent, uncertainty, safety, reversibility, and user control. Feedback from the person and environment updates the loop. Rather than requiring an exhaustive Human Digital Twin, the framework uses purpose-bounded, uncertainty-aware, user-correctable representations. We organize the design space by human-state targets, relational contexts, and agent roles, and propose scenario-centered evaluation, agency-preservation metrics, benchmark requirements, edge-native personal models, and governance directions. Combodied Agents shift Agentic AI from external task completion toward sustained human benefit.
摘要:在年長者錯過藥物劑量後,軟體代理可以發送另一個提醒,而具身代理可以帶來藥物。
然而,這兩者都沒有解釋該人是否忘記、感到困惑、出現副作用或故意拒絕,也沒有說明什麼樣的支持是合適的。
這揭示了代理人工智能中的結構性缺口:數位代理主要轉換軟體狀態,而具身代理則轉換物理狀態;兩者都未將個體不斷演變的狀態和能動性作為建模、干預和評估的主要對象。
我們引入了具身代理(Combodied Agents),這是一種以人為中心的範式,能夠隨著時間的推移感知、建模、預測和支持個體的人類狀態軌跡,使用軟體工具、感測器、可穿戴設備、機器人和人類服務作為行動渠道,而非最終目標。
我們將個人助理、健康代理、人工智慧伴侶和自適應人類-人工智慧系統的零散能力統一成一個閉環:基於事件的多模態感知重建有意義的個人事件;長期的、可修正的記憶提供時間背景;個人世界模型在不同的決策和干預下估計未來的個人狀態和結果;可接受的干預政策在同意、不確定性、安全性、可逆性和用戶控制下選擇相稱的支持。
來自個人和環境的反饋更新這個循環。
該框架不需要全面的人類數位雙胞胎,而是使用目的有限、具不確定性意識和用戶可修正的表徵。
我們根據人類狀態目標、關係背景和代理角色來組織設計空間,並提出以情境為中心的評估、能動性保護指標、基準要求、邊緣原生個人模型和治理方向。
具身代理將代理人工智能的重心從外部任務完成轉向持續的人類利益。
MIRA: Medical Image Reflection for Agentic Diagnosis
2608.10827v1 by Shengzhi Wang, Jun Yang, Kai Wu, Xiaozhong Ji, Yiwen Ye, Ziyang Chen, Mingliang Xiong, Wen Fang, Mingqing Liu, Mengyuan Xu, Miaoxuan Shan, Caiyan Liu, Bin He, Qingwen Liu
Medical visual agents can use tools to inspect images and retrieve external knowledge, but indiscriminate tool use may introduce noisy or misleading evidence. Reliable diagnosis therefore requires not only acquiring additional observations, but also verifying whether tool actions are necessary and whether the resulting evidence supports the current hypothesis. We introduce MIRA (Medical Image Reflection for Agentic Diagnosis), a medical visual diagnostic framework for autonomous evidence search and reflective verification. MIRA dynamically invokes image-processing operations, including zooming, grounding, pointing, rotation, and measurement, as well as web search, while evaluating the relevance and consistency of the acquired evidence. We develop MIRA through a two-stage training strategy. First, a tool-augmented Monte Carlo Tree Search data engine explores diverse diagnostic hypotheses and jointly verifies visual grounding accuracy and semantic consistency to construct supervised fine-tuning trajectories. Second, reinforcement learning further improves decision-making through online reflective principle evolution: failure cases are distilled into candidate principles, and only principles that improve held-out rollout rewards are retained. Across nine medical visual reasoning benchmarks, MIRA achieves an average score of 64.73, improving its Qwen3-VL-8B backbone by 7.44 points. It also increases useful tool-use judgments from 56.2% to 73.8% and reduces harmful judgments from 8.9% to 1.6%. Qualitative analyses show that MIRA can re-examine evidence, correct premature conclusions, and adapt its tool-use strategy. Project page: https://MIRA-VL.github.io/
摘要:醫學視覺代理可以使用工具來檢查圖像並檢索外部知識,但不加區別的工具使用可能會引入噪音或誤導性的證據。因此,可靠的診斷不僅需要獲取額外的觀察結果,還需要驗證工具行動是否必要,以及所產生的證據是否支持當前的假設。我們介紹了 MIRA(醫學影像反思代理診斷),這是一個用於自主證據搜索和反思驗證的醫學視覺診斷框架。MIRA 動態調用圖像處理操作,包括縮放、定位、指向、旋轉和測量,以及網絡搜索,同時評估所獲得證據的相關性和一致性。我們通過兩階段的訓練策略來開發 MIRA。首先,增強工具的蒙特卡羅樹搜索數據引擎探索多樣的診斷假設,並共同驗證視覺定位的準確性和語義一致性,以構建監督的微調軌跡。其次,強化學習進一步通過在線反思原則演變改善決策:失敗案例被提煉為候選原則,只有那些改善保留的展開獎勵的原則才會被保留。在九個醫學視覺推理基準中,MIRA 的平均分數為 64.73,將其 Qwen3-VL-8B 主幹提高了 7.44 分。它還將有用的工具使用判斷從 56.2% 提高到 73.8%,並將有害判斷從 8.9% 降低到 1.6%。定性分析顯示,MIRA 可以重新檢查證據、修正過早的結論,並調整其工具使用策略。項目頁面: https://MIRA-VL.github.io/
DuplexWorld: Can voice agents help you get through the day?
2608.10716v1 by Aryan Vijay Bhosale, Harshit Rajgarhia, Akhil Pothanapalli, Asif Shaik, Abhishek Mukherji, Dinesh Manocha
Speech-to-speech (S2S) voice agents are increasingly being incorporated into enterprise for customer care and as daily companions for consumers owing to the ease of the conversational modality over text. However, existing benchmarks fail to holistically evaluate voice agents along axes that really matter and are shaped as tests of agentic tool calling against a database. We believe they fail to adequately account for the diversity of conversational dialogue that mundane activities introduce and further, never test how faithfully an agent can assist on tasks that move beyond database manipulation. To tackle this DuplexWorld introduces six worlds where voice agents are especially useful: banking, insurance, travel, healthcare and logistics, and Pathfinding. Agents are evaluated on eleven different types of conversations across 156 scenarios (350+ hours of conversation), each testing conversational and analytical capability to varying degrees. Through extensive evaluation comprising agentic, conversational and speech-naturalness metrics, we show that even the best voice agents leave substantial room for improvement on all 3 axes (Pass@1: 0.490, turn-taking: 0.653, DNSMOS: 3.378). We perform extensive analysis on agentic v conversational performance, world- and conversation type-wise performance, failure modes exploring the explore v exploit lens for Pathfinding conversations and voice agent reliability over all six worlds.
摘要:語音對語音 (S2S) 語音代理越來越多地被企業納入客戶服務和作為消費者的日常伴侶,這是因為對話模式相較於文本更加方便。然而,現有的基準未能從真正重要的維度全面評估語音代理,並且其形式是針對數據庫進行的代理工具呼叫測試。我們認為,它們未能充分考慮日常活動所引入的對話多樣性,並且從未測試代理在超越數據庫操作的任務中能夠多麼忠實地提供協助。為了解決這個問題,DuplexWorld 引入了六個語音代理特別有用的世界:銀行、保險、旅行、醫療保健和物流,以及路徑尋找。代理在156個場景中對11種不同類型的對話進行評估(超過350小時的對話),每個場景測試對話和分析能力的不同程度。通過包括代理性、對話性和語音自然度指標的廣泛評估,我們顯示即使是最好的語音代理在這三個維度上仍有相當大的改進空間(Pass@1: 0.490,輪流發言: 0.653,DNSMOS: 3.378)。我們對代理性與對話性表現、世界及對話類型的表現、失敗模式進行了廣泛分析,探索了路徑尋找對話的探索與利用視角,以及所有六個世界中語音代理的可靠性。
MedUP: Awakening Unified Understanding and Perception in Medical Vision-Language Models
2608.10635v1 by Yuan Wang, Hualiang Wang, Yixin Chen, Songtao Jiang, Shujian Gao, Jiaming Lin, Siming Fu, Jian Wu, Zuozhu Liu
Medical Vision-Language Models (Med-VLMs) excel at verbalizing visual content, yet precise visual perception, segmentation, and grounding remain challenging. Existing approaches either verbalize regions as coordinate strings or rely on external modules that decouple perception from understanding, creating representation gaps for region-language alignment. We present MedUP, a Med-VLM that natively unifies perception and understanding within a shared token space. At its core lies UniMedTok, a region tokenizer that encodes masks as discrete tokens in the LLM vocabulary, enabling the model to seamlessly interleave mask tokens with text. We curate UniMed-Train, a 1.84M-instance corpus spanning text-guided segmentation, region-grounded understanding, medical VQA and CoT-based segmentation, and introduce UniMed-Bench for unified evaluation. Extensive experiments show that MedUP outperforms native, agentic, and dual-decoder Med-VLMs across all tasks while remaining competitive with specialist segmentors, demonstrating the strong potential of unified understanding and perception modeling.
摘要:醫療視覺語言模型(Med-VLMs)在將視覺內容口頭表達方面表現出色,但精確的視覺感知、分割和定位仍然具有挑戰性。現有的方法要麼將區域表達為坐標字符串,要麼依賴於將感知與理解解耦的外部模塊,這在區域與語言對齊中創造了表示差距。我們提出了 MedUP,一種在共享標記空間內本地統一感知和理解的 Med-VLM。其核心是 UniMedTok,一種將掩膜編碼為 LLM 詞彙中離散標記的區域標記器,使模型能夠無縫地將掩膜標記與文本交錯。我們整理了 UniMed-Train,一個包含 184 萬實例的語料庫,涵蓋文本引導的分割、區域基礎的理解、醫學 VQA 和基於 CoT 的分割,並介紹了 UniMed-Bench 以進行統一評估。廣泛的實驗表明,MedUP 在所有任務中超越了原生、主動和雙解碼器 Med-VLM,並在與專業分割器的競爭中保持競爭力,展示了統一理解和感知建模的強大潛力。
Agentic Instruction Data Selection: Let DataMaster Interpret Your Intent
2608.10579v1 by Fanqi Zhou, Qiaosheng Chen, Zixian Huang, Gong Cheng
Although existing instruction data selection methods have introduced various metrics, the inherent complexity of real-world datasets makes it impractical for any single metric to generalize across all scenarios. Developers are thus often forced to manually inspect data and craft heuristic rules for each new application---a tedious and error-prone process. In this paper, we propose a paradigm shift from manual configuration to automated orchestration via the Instruction Data Selection Agent (DataMaster), which interprets user intent and autonomously composes optimal selection strategies. By allowing users to specify data needs through natural language descriptions, DataMaster simplifies data curation and removes the burden of manual strategy design. Extensive experiments across the math, medical, and code domains show that DataMaster outperforms static baselines in most settings and surpasses full-pool training in a substantial number of cases. The implementation of DataMaster and the scripts needed to reproduce the reported pipeline are publicly available at https://github.com/nju-websoft/DataMaster.
摘要:儘管現有的指令數據選擇方法引入了各種指標,但現實世界數據集的固有複雜性使得任何單一指標在所有場景中都難以通用。
因此,開發者通常被迫手動檢查數據並為每個新應用編寫啟發式規則——這是一個繁瑣且容易出錯的過程。
在本文中,我們提出了一種從手動配置轉向自動編排的範式轉變,通過指令數據選擇代理(DataMaster),該代理解釋用戶意圖並自主組合最佳選擇策略。
通過允許用戶通過自然語言描述來指定數據需求,DataMaster 簡化了數據策展並消除了手動策略設計的負擔。
在數學、醫學和代碼領域的廣泛實驗表明,DataMaster 在大多數設置中超越了靜態基準,並在相當多的案例中超越了全池訓練。
DataMaster 的實現及重現報告流程所需的腳本可在 https://github.com/nju-websoft/DataMaster 上公開獲得。
Reinforcement Learning-Based Laser Cutting Machine Parameter Optimization
2608.10549v1 by Khanh Quan Pham, Majid Kundroo, Geunwoo Ban, Seongho Bae, Taehong Kim
Achieving high accuracy in laser-based cutting of optical films requires careful tuning of parameters such as focal length and laser power beam, adjusted according to the specific properties of each film type. Trial-and-error based traditional methods are used to find the most suitable cutting parameters for various films, but they are slow and inaccurate. To address this issue, this paper presents the Reinforcement Learning for Laser Cutting (RL$^{2}$C) algorithm, which uses Q-learning with an epsilon-greedy policy to dynamically optimize cutting parameters, significantly reducing taper size and film wastage. Additionally, RL$^{2}$C incorporates a dynamic environment space adaptability mechanism to allow it to adapt to new states encountered during the learning process over multiple batches of experiments. Experimental results demonstrate that RL$^{2}$C requires fewer steps and less time to find optimal cutting parameters compared to various RL-based optimization methods. Specifically, RL$^{2}$C reduces the number of optimization steps by up to 12.5\% and processing time by up to 81.8\% compared to existing methods. This study demonstrates the potential of RL in industrial laser-cutting processes by improving cut quality, reducing time and film wastage, and minimizing manual interventions.
摘要:達成激光切割光學薄膜的高精度需要仔細調整參數,例如焦距和激光功率光束,根據每種薄膜類型的特定特性進行調整。傳統的試錯方法用於尋找各種薄膜的最合適切割參數,但這些方法速度慢且不準確。為了解決這個問題,本文提出了激光切割強化學習(RL$^{2}$C)算法,該算法使用帶有epsilon-greedy策略的Q-learning來動態優化切割參數,顯著減少錐度大小和薄膜浪費。此外,RL$^{2}$C還結合了一個動態環境空間適應機制,使其能夠在多批次實驗的學習過程中適應遇到的新狀態。實驗結果表明,與各種基於RL的優化方法相比,RL$^{2}$C需要更少的步驟和更少的時間來找到最佳切割參數。具體而言,與現有方法相比,RL$^{2}$C將優化步驟數量減少了多達12.5\%,處理時間減少了多達81.8\%。這項研究通過提高切割質量、減少時間和薄膜浪費,以及最小化人工干預,展示了RL在工業激光切割過程中的潛力。
Unlocking the Power of Medical Tabular Data via Semantic-Aware Multimodal Pre-training
2608.10522v1 by Yingsheng Liu, Haiming Li, Jingmin Zhu, Jiajun Sun, Victoria Mar, Monika Janda, H. Peter Soyer, Zongyuan Ge, Zhen Yu
While vision-language models dominate medical representation learning, unstructured text lacks the dense, quantitative diagnostic phenotypes inherent in structured clinical tables. However, existing multimodal pre-training methods underutilize this potential due to semantic-agnostic designs that treat tabular inputs as flat vectors and employ unstable continuous regression objectives. To overcome this, we propose a novel semantic-aware framework explicitly modeling the intrinsic two-dimensional structure of tabular data. First, addressing the inter-feature hierarchy of varying diagnostic importance, we introduce Importance-Aware Adaptive Masking to construct a label-free curriculum prioritizing salient features. Second, addressing the intra-feature continuity-discreteness duality, we propose a Soft-Label Discretized Module that replaces unstable numerical regression with stable distribution matching, thereby mathematically preserving ordinal relationships. Extensive experiments across large-scale dermatology (SLICE-3D, HOP) and ophthalmology (EyePACS) datasets establish a new state-of-the-art (SOTA), demonstrating exceptional robustness and cross-domain generalizability.
摘要:雖然視覺-語言模型主導了醫學表徵學習,但非結構化文本缺乏結構化臨床表格中固有的密集、定量診斷表型。
然而,現有的多模態預訓練方法因為語義無關的設計而未能充分利用這一潛力,這些設計將表格輸入視為平坦的向量並採用不穩定的連續回歸目標。
為了克服這一問題,我們提出了一種新穎的語義感知框架,明確建模表格數據的內在二維結構。
首先,針對不同診斷重要性的特徵層次,我們引入了重要性感知自適應掩碼,構建了一個無標籤的課程,優先考慮顯著特徵。
其次,針對特徵內部的連續性-離散性二元性,我們提出了一個軟標籤離散模塊,將不穩定的數值回歸替換為穩定的分佈匹配,從而在數學上保持序關係。
在大規模皮膚科(SLICE-3D,HOP)和眼科(EyePACS)數據集上的廣泛實驗確立了新的最先進技術(SOTA),顯示出卓越的穩健性和跨領域的泛化能力。
RadFusion: Towards Threshold-Controllable Radiology Report Generation
2608.10505v1 by Ying Jin, Noel C. F. Codella, John Corring, Mu Wei, Dinei Florencio, Eric Horvitz
Automated radiology report generation is advancing rapidly in response to the shortage of radiologists, yet unlike a perception model, existing generation models offer no control over the sensitivity-specificity trade-off of their diagnostic content. Such control is essential because clinical scenarios diverge: emergency triage prioritizes sensitivity to reduce missed findings, whereas confirmatory interpretation emphasizes specificity to limit unnecessary interventions. A single fixed report can neither adapt to these scenarios nor support the ROC-based validation widely expected for regulatory clearance. We introduce RadFusion, a framework that equips report generation with threshold controllability. Our method fuses a multi-label classifier, which provides per-disease confidence scores, with a VQA-based report generator, which describes medical findings in detail; an LLM then rewrites the report so that its stated diagnoses follow the classifier's decisions at the selected threshold while staying grounded in the generator's descriptions. On MIMIC-CXR, the performance of RadFusion conforms to the classifier's ROC curve: sweeping the threshold and mapping the reports back to class labels reproduces the classifier's validated ROC performance. This conformance makes generated reports quantitatively evaluable through ROC analysis, strengthening the case for regulatory clearance, and enables operating-point selection that matches report behavior to clinical context. Moreover, combining the two model types improves diagnostic accuracy over uncontrolled generation: sensitivity increases by 6.9% at matched specificity, and specificity by 20.7% at matched sensitivity. These results show that RadFusion makes report generation clinically adaptable, quantitatively verifiable, and diagnostically more reliable.
摘要:自動化放射科報告生成正迅速發展,以應對放射科醫師的短缺,然而與感知模型不同,現有的生成模型無法控制其診斷內容的敏感性-特異性權衡。這種控制是至關重要的,因為臨床情境各異:緊急分診優先考慮敏感性以減少漏診,而確認性解釋則強調特異性以限制不必要的干預。單一的固定報告既無法適應這些情境,也無法支持廣泛期待用於監管批准的基於ROC的驗證。我們介紹了RadFusion,一個為報告生成提供閾值可控性的框架。我們的方法融合了一個多標籤分類器,該分類器提供每種疾病的信心分數,與一個基於VQA的報告生成器,該生成器詳細描述醫療發現;然後一個LLM重寫報告,使其所述的診斷遵循分類器在所選閾值下的決策,同時基於生成器的描述。 在MIMIC-CXR上,RadFusion的性能符合分類器的ROC曲線:調整閾值並將報告映射回類別標籤再現了分類器的經過驗證的ROC性能。這種一致性使得生成的報告可以通過ROC分析進行定量評估,增強了監管批准的案例,並使得操作點選擇能夠將報告行為與臨床情境相匹配。此外,結合這兩種模型類型提高了診斷準確性,相同特異性下敏感性提高了6.9%,相同敏感性下特異性提高了20.7%。這些結果顯示RadFusion使報告生成在臨床上適應性更強、定量可驗證且診斷上更可靠。
RLMOpt: Adaptive Prompt Optimization via Recursive Language Models
2608.10471v1 by Subhash Bangalore Satheesha, Nirvik Pande, Deepthi Duddempudi, Bharath Dandala
Prompt optimizers automate the search for prompts that improve language-model performance, but existing methods rely on a predefined optimization procedure: the algorithm determines which candidates to explore and how the search progresses, while the language model generates or refines prompt proposals. We introduce RLMOpt, a prompt optimizer that makes the search policy itself language-model-driven through a recursive language model (RLM). The RLM agent operates over a tool-based environment, inspecting task information, analyzing failures, generating candidates, allocating evaluation budget, and deciding when to stop. A deterministic harness complements the agent by enforcing objective scoring, Pareto-based selection, and regression constraints. We evaluate RLMOpt across four benchmarks spanning structured clinical information extraction (Chia), multi-hop question answering (HotpotQA), verifiable instruction following (IFBench-2025), and multi-turn tool-calling agents (BFCL). In a matched comparison at a single seed, RLMOpt obtains the best held-out score on all four benchmarks and leads the four-task mean (0.610 against 0.589 for GEPA). Repeating each benchmark across seeds yields 11 matched benchmark-seed comparisons, in which RLMOpt outperforms GEPA in 9 cases. Across all 11 runs, it never produced a prompt that underperformed its seed, whereas GEPA fell below its starting point twice. It is also more efficient, achieving these results with fewer search rollouts while producing prompts that are 27-79% the size of those produced by GEPA. Our results further show that optimization gains are determined primarily by the headroom available in the seed prompt, rather than by the search budget. Efficient optimization therefore depends on reaching the available headroom reliably and with minimal search
摘要:提示優化器自動化尋找能改善語言模型性能的提示,但現有的方法依賴於預定義的優化程序:算法決定了要探索哪些候選者以及搜索的進展方式,而語言模型則生成或完善提示提案。我們介紹了 RLMOpt,一個通過遞歸語言模型(RLM)使搜索策略本身由語言模型驅動的提示優化器。RLM 代理在基於工具的環境中運作,檢查任務信息、分析失敗、生成候選者、分配評估預算並決定何時停止。一個確定性的工具補充了代理,強制執行客觀評分、基於帕累托的選擇和回歸約束。
我們在四個基準上評估 RLMOpt,涵蓋結構化臨床信息提取(Chia)、多跳問題回答(HotpotQA)、可驗證的指令遵循(IFBench-2025)和多輪工具調用代理(BFCL)。在單一種子下的匹配比較中,RLMOpt 在所有四個基準上獲得了最佳的保留分數,並在四項任務的平均分(0.610 對 0.589 的 GEPA)中領先。重複每個基準跨種子產生 11 個匹配的基準-種子比較,其中 RLMOpt 在 9 個案例中超越了 GEPA。在所有 11 次運行中,它從未產生低於其種子的提示,而 GEPA 則有兩次低於其起始點。它的效率也更高,以更少的搜索展開達成這些結果,同時生成的提示大小僅為 GEPA 的 27-79%。
我們的結果進一步顯示,優化增益主要取決於種提示中可用的潛力,而不是搜索預算。因此,高效的優化依賴於可靠地達到可用的潛力並以最小的搜索進行。
Expert-Guided g-computation with Large Language Models for Estimating Causal Effects on Timings: Applications to Hospital Quality Improvement
2608.10339v1 by Patrick Vossler, Jialin Ouyang, F. Richard Guo, Anran Huang, Ali Shojaie, Lucas Zier, Fan Xia, Jean Feng
Hospital quality improvement (QI) programs routinely face multiple candidate interventions to optimize hospital flow, but existing methods struggle to estimate and rank the causal effects of such interventions. This work focuses on one of the most standard hospital metrics, the average length of stay (LOS), and its causal estimand, the average time saved. To characterize this causal effect, qualitative approaches rely on expert judgment to map patient trajectories, making them susceptible to cognitive biases; quantitative approaches rely on data-driven models, which fail when interventions are hypothetical with no historical data or have complex causal mechanisms that require clinical reasoning rather than data alone. We propose expert-guided g-computation, or egg-computation, which combines the complementary strengths of both approaches by connecting the Gantt charts commonly used to map patient trajectories with the causal DAG literature. We introduce a causal model over Gantt charts and establish identification using a variant of g-computation that seeks expert input only for components unidentifiable from data. To make egg-computation practical, we develop an LLM-assisted pipeline that reliably scales up expert reasoning. In simulations, egg-computation outperforms conventional causal inference methods when patients have diverse causal structures and intervention mechanisms. In a study of eleven candidate QI interventions at an urban safety-net hospital, the LLM pipeline generated graphs and time-saving estimates highly concordant with those of human experts. Beyond healthcare, egg-computation is a broadly applicable framework for estimating the average time saved for candidate interventions whose causal mechanisms can be represented using Gantt charts.
摘要:醫院質量改善(QI)計劃經常面臨多種候選干預措施,以優化醫院流程,但現有方法在估計和排名這些干預措施的因果效應方面存在困難。這項工作專注於最標準的醫院指標之一,即平均住院天數(LOS),及其因果估計量,即平均節省的時間。為了表徵這一因果效應,定性方法依賴專家判斷來映射病人軌跡,使其容易受到認知偏見的影響;定量方法則依賴數據驅動模型,當干預措施是假設性的且沒有歷史數據,或具有需要臨床推理而非僅依賴數據的複雜因果機制時,這些模型會失效。我們提出了專家引導的g計算,或稱蛋計算,這種方法通過將常用於映射病人軌跡的甘特圖與因果DAG文獻相連接,結合了兩種方法的互補優勢。我們在甘特圖上引入了一個因果模型,並使用一種變體的g計算來建立識別,該變體僅尋求專家對數據無法識別的組件的輸入。為了使蛋計算實用,我們開發了一個LLM輔助的管道,可靠地擴展專家推理。在模擬中,當病人具有多樣的因果結構和干預機制時,蛋計算的表現超過了傳統的因果推斷方法。在對一所城市安全網醫院的十一個候選QI干預措施的研究中,LLM管道生成的圖形和節省時間的估計與人類專家的結果高度一致。除了醫療保健之外,蛋計算是一個廣泛適用的框架,用於估計可以用甘特圖表示的候選干預措施的平均節省時間。
Logit-Boundary Geometric Belief Interfaces and Sparse Sheaf-Enclave Protocols: A Self-Contained Substrate for Secure Network Electronic Health Record (EHR) Interoperability
2608.10300v2 by Alvin Spivey, Yu Huang
Electronic health-record interoperability is a boundary problem: legacy systems, generative models, terminology services, identity systems, and human reviewers may each expose rich internal states, while operational exchange requires a narrow shared interface of typed claims, bounded uncertainty, provenance, and explicit admission or abstention. This paper details a mathematical and engineering architecture for that interface. The organizing idea is the logit boundary: a discovery model may propose pre-threshold scores over a local categorical decision, but a deterministic judgment substrate decides whether the proposal is admissible, requires review, or must be quarantined before any Fast Healthcare Interoperability Resources (FHIR) transaction is constructed. The resulting Geometric Belief Interface (GBI) combines finite boundary semantics, local Dirichlet evidence, cellular-sheaf and mapping-cone diagnostics, advisory geometric audit charts, and a Decentralized Cryptographic Sheaf-Enclave (DCSE) protocol sketch for fail-closed deployment. The framework does not establish clinical truth, global representation alignment, or end-to-end safety; it defines certificate-producing checks at a model-to-system boundary. A companion frozen synthetic benchmark, GBI BoundaryBench v0.1, evaluated Qwen3-4B-Instruct-2507 on 256 held-out tasks across three evidence modes (768 canonical executions). All executions completed, but none produced an output accepted by the benchmark contract: 369 were rejected during safe parsing and 399 during schema validation, yielding zero coverage and deterministic quarantine. This empirical result is deliberately narrow - one 4B open-weight model under one frozen interface - and is reported as evidence about the admission boundary, not as a general claim about LLM capability or clinical safety. A Julia appendix verifies numerical certificates using standard libraries.
摘要:電子健康紀錄的互操作性是一個邊界問題:舊系統、生成模型、術語服務、身份系統和人工審查者各自可能暴露豐富的內部狀態,而操作性交換則需要一個狹窄的共享介面,該介面包括類型聲明、有限的不確定性、來源以及明確的接受或放棄。本文詳細描述了該介面的數學和工程架構。組織思想是邊界邏輯:發現模型可能會對局部類別決策提出閾值前的分數,但確定性判斷基底決定該提案是否可接受、是否需要審查,或在構建任何快速醫療互操作性資源(FHIR)交易之前必須被隔離。最終的幾何信念介面(GBI)結合了有限邊界語義、局部Dirichlet證據、細胞叢和映射錐診斷、諮詢幾何審計圖表,以及一個去中心化的加密叢集區(DCSE)協議草圖,以實現故障關閉部署。該框架並不建立臨床真實性、全球表示對齊或端到端安全性;它定義了在模型與系統邊界的證書生成檢查。一個伴隨的凍結合成基準,GBI BoundaryBench v0.1,評估了Qwen3-4B-Instruct-2507在三種證據模式下的256個保留任務(768個典型執行)。所有執行均已完成,但沒有任何輸出被基準合約接受:369個在安全解析過程中被拒絕,399個在模式驗證過程中被拒絕,導致零覆蓋和確定性隔離。這一實證結果故意狹窄——一個4B開放權重模型在一個凍結介面下——並被報告為關於接受邊界的證據,而不是關於LLM能力或臨床安全的一般主張。Julia附錄使用標準庫驗證數字證書。
Frozen Brain-MRI Foundation Models Are Site Fingerprints
2608.10295v1 by Saman Rahbar
Frozen foundation-model (FM) embeddings are increasingly used as off-the-shelf brain-MRI representations, on the assumption that they capture anatomy. We audit what they actually encode and find that acquisition site is a large, intrinsic component of the representation. Across two independent cohorts (ABIDE-I, ABIDE-II), three frozen 3-D encoders (brain-pretrained, CT-pretrained, and randomly initialized), and every network depth, site is linearly decodable at roughly 0.9 balanced accuracy at deep layers, exceeding the decodability of every clinical or demographic variable (sex, age, autism diagnosis) at every layer. The effect is intrinsic rather than learned: a randomly initialized encoder is already a ~0.9 site classifier on both cohorts and across three architecture families (Swin, ViT, ResNet), and site is decodable at ~0.95 directly from the raw downsampled image with no encoder, so the fingerprint reflects low-level image statistics that any encoder preserves rather than a product of pretraining. Residualizing measured population covariates leaves site decodability essentially unchanged, indicating an acquisition- rather than population-driven effect. A nonlinear probe matches the linear one, so the fingerprint is fully linearly accessible. The site subspace is removable post hoc by iterative null-space projection or ComBat (site decodability 0.94 -> 0.07/0.00), and is a site-attribution concern for shared or federated embeddings; but for dense segmentation this removal is not free, because site and anatomy occupy an entangled linear subspace (a matched-rank random-direction projection is Dice-neutral, whereas removing the site subspace is destructive). We recommend site-audited use of frozen brain-MRI FMs and release an open audit toolkit.
摘要:冷凍的基礎模型(FM)嵌入越來越多地被用作現成的腦部MRI表徵,假設它們能捕捉到解剖結構。我們審核它們實際編碼的內容,發現獲取地點是該表徵的一個重要內在組成部分。在兩個獨立的隊列(ABIDE-I,ABIDE-II)、三個冷凍的3D編碼器(腦部預訓練、CT預訓練和隨機初始化)以及每個網絡深度中,地點在深層的線性可解度約為0.9的平衡準確率,超過了每一層的臨床或人口變量(性別、年齡、自閉症診斷)的可解度。這一效應是內在的,而非學習得來的:隨機初始化的編碼器在兩個隊列和三個架構系列(Swin、ViT、ResNet)中已經是一個約0.9的地點分類器,並且地點可以直接從原始下採樣圖像中以約0.95的可解度解碼,而無需編碼器,因此指紋反映的是任何編碼器所保留的低級圖像統計,而不是預訓練的產物。對測量的人口協變量進行殘差化處理基本上不改變地點的可解度,這表明這是一種獲取驅動而非人口驅動的效應。非線性探針與線性探針相匹配,因此指紋是完全線性可訪問的。地點子空間可以通過迭代的零空間投影或ComBat事後移除(地點可解度0.94 -> 0.07/0.00),並且對於共享或聯邦嵌入來說,這是一個地點歸因的問題;但對於密集分割來說,這種移除並不是免費的,因為地點和解剖結構佔據了一個糾纏的線性子空間(匹配秩的隨機方向投影是Dice中性的,而移除地點子空間則是破壞性的)。我們建議對冷凍的腦部MRI FMs進行地點審核使用,並發布一個開放的審核工具包。
Locally Deployable Small Language Models for Emergency Department Decision Support: A Systematic Benchmark of Fine-Tuning Strategies
2608.10273v1 by Qingfeng Zhang, Yuanxiong Guo, Yanmin Gong
Deploying large language models (LLMs) for decision support in emergency departments (EDs) faces two major challenges: privacy risks of transmitting patient data to closed-source commercial LLMs and the lack of systematic evaluation of fine-tuning strategies for locally deployable open-source small language models (SLMs). We benchmarked eight open-source SLMs using zero-shot prompting, prefix tuning, Low-Rank Adaptation (LoRA), and full fine-tuning on three ED tasks: triage level prediction, specialist referral recommendation, and diagnosis prediction. Using 2,083 MIMIC-IV-ED cases and Claude Haiku 4.5 and Claude Sonnet 4.5 as baselines, we found that LoRA fine-tuned open-source SLMs outperform commercial baselines on triage level prediction and specialist referral recommendation, while diagnosis prediction remains challenging for open-source SLMs. Confusion matrix analysis further shows that fine-tuned open-source SLMs can detect highest-severity patients missed by the commercial baselines. These results demonstrate that locally deployable SLMs can achieve clinically competitive performance for ED decision support.
摘要:部署大型語言模型(LLMs)以支援急診部門(EDs)的決策面臨兩大挑戰:將病人數據傳輸到封閉源商業LLMs的隱私風險,以及缺乏對可本地部署的開源小型語言模型(SLMs)進行微調策略的系統評估。我們使用零樣本提示、前綴調整、低秩適應(LoRA)和完全微調,對八個開源SLMs在三個ED任務上進行基準測試:分診等級預測、專家轉診建議和診斷預測。使用2,083個MIMIC-IV-ED案例,並以Claude Haiku 4.5和Claude Sonnet 4.5作為基準,我們發現LoRA微調的開源SLMs在分診等級預測和專家轉診建議上優於商業基準,而診斷預測對於開源SLMs仍然具有挑戰性。混淆矩陣分析進一步顯示,微調後的開源SLMs能夠檢測到商業基準漏掉的高嚴重性病人。這些結果表明,可本地部署的SLMs能夠在急診決策支援中達到臨床競爭性能。
TAF-MED: Multi-Turn Safety Refusal Collapse in LLMs Under Declared Self-Treatment Intent
2608.10258v1 by Waleed Jamil, Raphael Schmitt
Large language models (LLMs) increasingly provide conversational health information that may influence treatment decisions, yet existing benchmarks do not isolate whether medication-safety boundaries persist across follow-ups after explicit self-treatment intent. We introduce TAF-MED, a physician-reviewed benchmark of 500 fixed three-turn scenarios, and evaluate eight LLMs across 4,000 conversations. A rubric-based automated judge labelled responses as SAFE, LEAKY, or UNSAFE, and two physicians independently annotated a model-balanced random subset of 400 conversations. We assessed unsafe guidance, collapse after a strictly SAFE initial response, and model-ranking stability. Overall, 71.6% of conversations contained an UNSAFE response, and 61.4% of those beginning with a strictly SAFE response later collapsed to UNSAFE; model-level collapse rates ranged from 24.4% to 96.2%. Four of 28 model pairs reversed order between initial unsafe and collapse rates. Automated labels achieved 94.3% agreement with the adjudicated physician reference ($κ= 0.895$). These findings show that first-turn safety is an incomplete proxy for conversational safety persistence and motivate evaluation across complete dialogue trajectories. We will release TAF-MED on Hugging Face to support reproducible research on multi-turn medical safety.
摘要:大型語言模型(LLMs)越來越多地提供可能影響治療決策的對話健康資訊,但現有的基準並未區分在明確的自我治療意圖後,藥物安全邊界是否持續存在。我們介紹了TAF-MED,一個經過醫生審核的500個固定三輪情境的基準,並評估了八個LLM在4000次對話中的表現。基於評分標準的自動評判將回應標記為安全(SAFE)、漏洩(LEAKY)或不安全(UNSAFE),兩位醫生獨立註釋了一個模型平衡的隨機子集,共400次對話。我們評估了不安全指導、在嚴格安全的初始回應後的崩潰情況,以及模型排名的穩定性。總體而言,71.6%的對話包含不安全的回應,而61.4%從嚴格安全的回應開始的對話後來崩潰為不安全;模型層級的崩潰率範圍從24.4%到96.2%。28對模型中有四對在初始不安全和崩潰率之間的順序相反。自動標籤與裁定的醫生參考達到94.3%的一致性($κ= 0.895$)。這些發現顯示,第一輪的安全性並不是對話安全持續性的完整代理,並促使對完整對話軌跡的評估。我們將在Hugging Face上發布TAF-MED,以支持多輪醫療安全的可重複研究。
Towards Expert-level Medical AI for Real-time Video Consultations
2608.09861v1 by Mahvish Nagda, Jihyeon Lee, Matthew Thompson, Chunjong Park, Tim Strother, Valentin Liévin, Roma Ruparel, Akshay Goel, Teya Bergamaschi, Suhana Bedi, Meet Shah, Pavel Dubov, Liviu Panait, Toshiyuki Fukuzawa, Sam Schmidgall, Craig Schiff, Joseph Xu, Aliya Rysbek, Yana Lunts, Jan Freyberg, Rebecca Hemengway, Sunny Virmani, David Racz, Carey Radebaugh, Joëlle Barral, Kavi Goel, Dale R. Webster, Katherine Chou, Avinatan Hassidim, Yossi Matias, James Manyika, Gregory Wayne, Tao Tu, Yun Liu, Ethan Goh, Christina Chen, Ryutaro Tanno, Po-Hsuan Cameron Chen, Mike Schaekermann, Anil Palepu
Audio-visual interaction is the standard for patient-physician consultations, enabling natural communication and effective assessment of illness through non-verbal cues. While text-based AI has shown promise, it discards essential perceptual dimensions and limits patients who cannot articulate symptoms in writing. Early efforts to extend medical AI to audio-visual interaction have demonstrated feasibility but not reached clinician-level performance. Here, we provide the first demonstration of expert-level AI in real-time clinical video consultations using AMIE (Articulate Medical Intelligence Explorer) in a video configuration. AMIE (Video) is a Gemini-based multi-agent system integrating low-latency dialogue, clinical reasoning, and real-time audio-visual perception. To guide development, we established a taxonomy and automated evaluations for clinical audio-visual cues in telehealth settings. In a randomized Objective Structured Clinical Examination (OSCE) study with 30 primary care physicians (PCPs), 15 patient actors and 100 clinical scenarios, we compared AMIE (Video), its text-only counterpart AMIE (Text), and PCPs consulting via video. Clinical evaluators rated AMIE (Video) on par or better than PCPs in history-taking, diagnosis, management, and physical observation and examination. Patient actors preferred AMIE's approach to assessing and explaining conditions, while PCPs were preferred for rapport and partnership building. In modality ablation, patient actors preferred AMIE (Video)'s interface over text chat for communicative effectiveness, convenience, and feeling understood. Limitations remain in fine anatomical precision, subtle affective nuances, and high-frequency movements. While further research is needed before real-world translation, these results mark an important milestone toward AI systems capable of augmenting care across the sensory complexity of clinical practice.
摘要:視聽互動是病人與醫生諮詢的標準,能夠通過非語言線索促進自然交流和有效評估疾病。雖然基於文本的人工智慧顯示出潛力,但它忽略了重要的感知維度,並限制了無法用書面表達症狀的病人。早期將醫療人工智慧擴展至視聽互動的努力已顯示出可行性,但未達到臨床醫生的表現水平。在此,我們提供了使用AMIE(Articulate Medical Intelligence Explorer)在視頻配置中進行實時臨床視頻諮詢的專家級人工智慧的首次示範。AMIE(視頻)是一個基於Gemini的多代理系統,整合了低延遲對話、臨床推理和實時視聽感知。為了指導開發,我們建立了一個分類法和自動評估,用於遠程醫療環境中的臨床視聽線索。在一項隨機的客觀結構化臨床考試(OSCE)研究中,涉及30位初級保健醫生(PCPs)、15位病人演員和100個臨床場景,我們比較了AMIE(視頻)、其文本專用對應AMIE(文本)以及通過視頻諮詢的PCPs。臨床評估者在病史採集、診斷、管理以及身體觀察和檢查方面評價AMIE(視頻)與PCPs相當或更好。病人演員更喜歡AMIE在評估和解釋病情方面的方法,而PCPs則在建立關係和夥伴關係方面更受青睞。在模態消融中,病人演員更喜歡AMIE(視頻)的界面而非文本聊天,因為其在交流有效性、便利性和被理解的感受上表現更佳。儘管在精細解剖精度、微妙的情感細微差別和高頻運動方面仍存在局限性,但在實際應用之前仍需進一步研究,這些結果標誌著朝著能夠增強臨床實踐中感官複雜性的護理的人工智慧系統邁出了重要的一步。
MedPixel: A Unified Pixel-Language Model for Medical Reasoning and Segmentation
2608.09818v1 by Haoyu Yang, Meixing Shi, Zengjie Chen, Haoran Sun, Haitao Leng, Xiaoming Shi, Yuxiang Cai, Yankai Jiang
Reliable medical image understanding requires models to connect clinical language and visual reasoning with pixel-level grounding. Yet medical vision-language models often lack precise localization, whereas medical segmenters typically rely on explicit target categories or precise spatial prompts. This divide is reinforced by a supervision mismatch: segmentation datasets provide precise masks but little language supervision, whereas medical vision-language data rarely pair language with dense spatial annotations. To address this gap, we present MedPixel, a unified medical pixel-language model built around a shared language--mask interface. To provide scalable supervision, we introduce MedPLG-440K, comprising approximately 440K pixel-language task samples constructed through a clinically motivated synthesis process without external LLM annotation. MedPixel is trained with joint multi-task supervised fine-tuning followed by Pixel-Level Preference Optimization, which uses ground-truth masks as offline verifiers to derive response preferences from mask quality. MedPixel supports a broad spectrum of tasks spanning explicit grounding, implicit reasoning, spatial interaction, grounded explanation, and medical VQA. Across this task spectrum, MedPixel achieves strong performance in both pixel-level prediction and response generation, together with effective zero-shot transfer to external grounding benchmarks and robustness to imperfect spatial prompts. Code and model checkpoints will be released at https://github.com/yhy-whu/Medpixel.
摘要:可靠的醫學影像理解需要模型將臨床語言和視覺推理與像素級基礎連接起來。
然而,醫學視覺-語言模型通常缺乏精確的定位,而醫學分割模型則通常依賴於明確的目標類別或精確的空間提示。
這種差距因監督不匹配而加強:分割數據集提供精確的掩膜,但語言監督很少,而醫學視覺-語言數據則很少將語言與密集的空間註釋配對。
為了解決這一差距,我們提出了MedPixel,一個圍繞共享語言-掩膜接口構建的統一醫學像素-語言模型。
為了提供可擴展的監督,我們引入了MedPLG-440K,該數據集由大約440K的像素-語言任務樣本組成,這些樣本是通過臨床驅動的合成過程構建的,並且沒有外部LLM註釋。
MedPixel通過聯合多任務的監督微調進行訓練,隨後進行像素級偏好優化,該過程使用真實掩膜作為離線驗證器,從掩膜質量中推導響應偏好。
MedPixel支持廣泛的任務,涵蓋明確的基礎、隱含推理、空間互動、基於基礎的解釋和醫學VQA。
在這一任務範疇中,MedPixel在像素級預測和響應生成方面都達到了強勁的性能,並有效地實現了對外部基礎基準的零樣本轉移,以及對不完美空間提示的穩健性。
代碼和模型檢查點將在 https://github.com/yhy-whu/Medpixel 發布。
AirFlow: Context Preserving and Multi-Rate State Modeling for Air Quality Forecasting
2608.09775v1 by Fan Yang, Nan Chen, Yijie Dong, Yuchen Zhang, Wei Zhang
Accurate air quality forecasting is essential for public health and urban environmental management, but remains challenging because pollutant channels differ in periodicity and distribution drift, while their concentration trajectories contain both multi-scale dependencies and rapid changes. Recent methods have improved spatial dependency learning and meteorological covariate modeling. However, pollutant channels are still passed through the same normalization rule and temporal backbone, using a shared latent representation for channel-specific distributions and changes at different rates. To address this limitation, we propose AirFlow, a pollutant-aware dual-stream framework that operates on station multivariate observations without additional graph propagation or predefined signal decomposition. Specifically, AirFlow designs two novel blocks: (1) a statistic-guided normalization routing mechanism that selects a normalization path for each pollutant according to its 24-hour autocorrelation and distribution drift; and (2) a hierarchical dual-stream state model that combines multi-scale state space propagation with learnable response coefficients, where gated bidirectional cross-attention exchanges information and adaptively fuses the resulting representations. Experiments on real-world data from multiple cities show that AirFlow achieves the best performance in 34 of 36 metrics comparisons, with reductions of up to 11.11% root mean square error over the state-of-the-art baseline. AirFlow also requires only 0.0483M parameters and 0.0215G FLOPs, achieving high forecasting accuracy with low computational overhead.
摘要:準確的空氣質量預測對於公共健康和城市環境管理至關重要,但仍然具有挑戰性,因為污染物通道在周期性和分佈漂移上存在差異,而它們的濃度軌跡則包含多尺度依賴性和快速變化。最近的方法改善了空間依賴性學習和氣象協變量建模。然而,污染物通道仍然通過相同的正規化規則和時間主幹,使用共享的潛在表示來處理特定通道的分佈和不同速率的變化。為了解決這一限制,我們提出了AirFlow,一種污染物感知的雙流框架,該框架在站點多變量觀測上運行,而不需要額外的圖形傳播或預定義的信號分解。具體而言,AirFlow設計了兩個新穎的模塊:(1)一個統計引導的正規化路由機制,根據每個污染物的24小時自相關和分佈漂移選擇正規化路徑;(2)一個層次雙流狀態模型,將多尺度狀態空間傳播與可學習的響應係數結合,其中門控雙向交叉注意力交換信息並自適應地融合結果表示。來自多個城市的實驗數據顯示,AirFlow在36個指標比較中有34個達到了最佳性能,並在最先進的基準上實現了最高11.11%的均方根誤差減少。AirFlow還僅需0.0483M參數和0.0215G FLOPs,以低計算開銷實現高預測準確性。
Physics-Informed Machine Learning in Prognostics and Health Management: A Systematic Literature Review
2608.10047v1 by Christopher Braun, Julian Raible, Marco F. Huber
In modern industry, keeping complex systems reliable, safe, and efficient hinges on Prognostics and Health Management (PHM). Machine Learning (ML) has largely driven advancements in diagnostics and prognostics, yet purely data-driven models face inherent limitations, such as poor generalization, an inability to infer causal relationships, and a lack of interpretability. Physics-Informed Machine Learning (PIML) helps mitigate these limitations by incorporating prior physical knowledge directly into the ML pipeline, thereby fostering growing interest in its application to PHM. This work investigates how PIML is being leveraged in the context of PHM through a systematic literature review of 212 studies. The review introduces a four-class classification scheme, consisting of observational bias, inductive bias, learning bias, and hybrid approaches, and further categorizes studies by PHM task. Across all four classes, the reviewed studies consistently demonstrate improved predictive performance over conventional baselines across a broad range of assets, although the literature is heavily skewed toward lithium-ion batteries and bearings, and dominated by problem-specific solutions. Overall, the review indicates that physics-informed approaches already provide tangible benefits, whereas claims of improvements concerning some of the aforementioned limitations lack sufficient supporting evidence. Future research should prioritize transferable design patterns, benchmarks comparing integration strategies, and uncertainty-aware models that are lightweight and robust enough for online deployment in real-world settings.
摘要:在現代工業中,保持複雜系統的可靠性、安全性和效率依賴於預測與健康管理(PHM)。機器學習(ML)在診斷和預測方面的進展主要受到推動,但純數據驅動的模型面臨固有的限制,例如一般化能力差、無法推斷因果關係以及缺乏可解釋性。物理知識驅動的機器學習(PIML)通過將先前的物理知識直接納入機器學習流程,幫助減輕這些限制,從而引發對其在PHM應用中的日益關注。本研究通過對212項研究的系統文獻回顧,調查了PIML在PHM背景下的應用。該回顧介紹了一種四類分類方案,包括觀察偏差、歸納偏差、學習偏差和混合方法,並進一步根據PHM任務對研究進行分類。在所有四類中,所回顧的研究一致顯示出相較於傳統基準的預測性能有所改善,儘管文獻在很大程度上偏向於鋰離子電池和軸承,並且以問題特定的解決方案為主導。總體而言,該回顧表明,物理知識驅動的方法已經提供了切實的好處,而對於上述某些限制的改善的聲稱缺乏足夠的支持證據。未來的研究應優先考慮可轉移的設計模式、比較整合策略的基準以及足夠輕量且穩健的、不確定性感知模型,以便在現實環境中進行在線部署。
Coupled Graph--Policy Distillation for Personalized Medication Safety in Older Adults with Multimorbidity
2608.09443v1 by Zihan Wang, Anglin Liu, Rongyi Wang, Dantong Li, Yi Lu, Siqing Yuan, Hongxia Xu, Zhongtian Long, Jintai Chen
Large language model (LLM) agents can support medication review between clinical visits, but safe choices for older adults with multimorbidity depend on conditions, medications, and geriatric risks that users may omit. We introduce ATLAS, a coupled graph--policy distillation framework for patient-adaptive medication safety. ATLAS structures guideline evidence as a medication-safety graph. Targeted questions update the patient state and distill relevant relations into a patient-specific medication conflict graph (PMCG). A risk-first multi-agent policy uses the PMCG to screen contraindications, assess cautions and monitoring needs, identify safer alternatives, and verify the final medication plan. We also introduce GeriMedBench, an interactive benchmark that tests safety-critical information acquisition and evidence-based decision revision. Across a European non-interactive multimorbidity benchmark, an Asian interactive multimorbidity benchmark, and an Asian non-interactive cross-guideline benchmark, ATLAS achieves the strongest complete-decision performance among the compared systems. On the European non-interactive multimorbidity benchmark, it exceeds the strongest proprietary LLM baseline by 53.73 points in Strict Success Rate and 14.63 points in overall safety reasoning score (OSRS), with no unsafe recommendations under the automated evaluator. A blinded clinician evaluation gives ATLAS higher mean ratings across all five criteria and flags potentially unsafe recommendations in one ATLAS case and two Gemini cases.
摘要:大型語言模型 (LLM) 代理可以在臨床訪問之間支持藥物審查,但對於多重疾病的老年人,安全的選擇取決於用戶可能省略的條件、藥物和老年風險。我們介紹了 ATLAS,一個耦合圖形-政策蒸餾框架,用於患者自適應藥物安全。ATLAS 將指導證據結構化為藥物安全圖。針對性的問題更新患者狀態,並將相關關係蒸餾成患者特定的藥物衝突圖 (PMCG)。一種以風險為先的多代理政策利用 PMCG 來篩選禁忌症,評估注意事項和監測需求,識別更安全的替代品,並驗證最終的藥物計劃。我們還介紹了 GeriMedBench,一個互動基準,測試安全關鍵信息獲取和基於證據的決策修訂。在一個歐洲非互動多重疾病基準、一個亞洲互動多重疾病基準和一個亞洲非互動跨指導基準中,ATLAS 在比較系統中實現了最強的完整決策表現。在歐洲非互動多重疾病基準中,它在嚴格成功率方面超過了最強的專有 LLM 基準 53.73 分,在整體安全推理分數 (OSRS) 上超過 14.63 分,且在自動評估者下沒有不安全的建議。盲評的臨床醫生評估給予 ATLAS 在所有五個標準上更高的平均評分,並在一個 ATLAS 案例和兩個 Gemini 案例中標記了潛在的不安全建議。
Multimodal Federated Learning under Dual-Axis Modality Missingness
2608.09240v1 by Adiba Orzikulova, Jaehyun Kwak, Jaemin Shin, Yunqi Guo, Xiaomin Ouyang, Guoliang Xing, Steven Euijong Whang, Sung-Ju Lee
Multimodal federated learning (FL) supports collaborative modeling in privacy-sensitive health-sensing and medical settings, but realistic deployments often exhibit dual-axis modality missingness: clients have different modality sets, and individual samples may contain only subsets of the modalities available locally. Existing methods typically address these two axes separately. We propose Flux, a multimodal federated learning framework built around two complementary components. First, modality-aware confidence tempering learns sample-specific confidence for each modality through mask-aware unimodal supervision and fuses the confidence estimates from observed modalities into a sample-adaptive temperature that adjusts predictive sharpness according to evidence quality and completeness. Second, gradient-decoupled private adaptation applies this temperature only to a client-private prediction pathway, while training the shared federated model with a standard, untempered objective. This enables sample-specific, client-local confidence adaptation without allowing confidence-dependent gradients to perturb shared representation learning. Across four multimodal datasets, Flux achieves the highest average macro-F1 on every dataset, outperforming the strongest dataset-specific baseline by 0.8~2.2 points and by 1.6 points on average. Additional analyses demonstrate favorable calibration, temperature sensitivity to both modality missingness and input corruption, and more stable shared optimization under private-only tempering. Our code is available at https://github.com/AdibaOrz/Flux.
摘要:多模態聯邦學習(FL)支持在隱私敏感的健康感測和醫療環境中進行協作建模,但現實部署通常顯示出雙軸模態缺失的情況:客戶端擁有不同的模態集,且個別樣本可能僅包含當地可用模態的子集。現有的方法通常分別處理這兩個軸。我們提出了Flux,一個圍繞兩個互補組件構建的多模態聯邦學習框架。首先,模態感知的信心調整通過基於掩碼的單模態監督學習每個模態的樣本特定信心,並將觀察到的模態的信心估計融合成樣本自適應的溫度,根據證據的質量和完整性調整預測的清晰度。其次,梯度解耦的私有適應僅將這個溫度應用於客戶端私有的預測路徑,同時使用標準的、未調整的目標訓練共享的聯邦模型。這使得樣本特定的、客戶端本地的信心適應成為可能,而不允許依賴信心的梯度干擾共享表示學習。在四個多模態數據集上,Flux在每個數據集上都達到了最高的平均宏F1,超過了最強的數據集特定基線0.8~2.2點,平均超過1.6點。額外的分析顯示出良好的校準、對模態缺失和輸入損壞的溫度敏感性,以及在僅進行私有調整時更穩定的共享優化。我們的代碼可在 https://github.com/AdibaOrz/Flux 獲得。
Rethinking Medical Landmark Localization with Prototype Learning-based Progressive Offset Correction
2608.09182v2 by Jingxian Xu, Yuhao Huang, Rusi Chen, Yanfeng Zhou, Dong Ni
Accurate landmark localization in medical images is a fundamental step for quantitative clinical measurement and downstream analysis. Existing localization methods have advanced, among which multi-stage refinement is a superior solution. Although this strategy mitigates the anatomical ambiguity inherent in single-stage global predictions, its high computational cost limits practical applicability. In this work, we propose a parameter-economic model, PPOC-LL, which leverages Prototype learning-based Progressive Offset Correction for Landmark Localization. Our contribution is three-fold. First, to drive coarse-to-fine landmark optimization, we introduce a multi-scale dynamic perception strategy for patch-level feature pyramid modeling. Second, to effectively handle anatomically similar patterns, we design a similarity-driven prototype learning mechanism that captures informative local semantics for robust offset prediction. Last, to stabilize the model learning and improve the overall performance, we incorporate a novel error-aware reliability regularization via tolerance-based balancing. We collected a large validation cohort, including two public and one private datasets spanning X-ray and ultrasound modalities, covering cephalometric, symphysis-fetal head, and fetal heart landmarks. Extensive experiments demonstrate that PPOC-LL achieves satisfactory performance with a favorable trade-off between accuracy and model complexity.
摘要:醫學影像中準確的地標定位是進行定量臨床測量和後續分析的基本步驟。現有的定位方法已經取得了進展,其中多階段精煉是一個優越的解決方案。儘管這一策略減輕了單階段全局預測中固有的解剖學模糊性,但其高計算成本限制了實際應用。在本研究中,我們提出了一種經濟參數模型PPOC-LL,該模型利用基於原型學習的漸進偏移修正來進行地標定位。我們的貢獻有三個方面。首先,為了驅動粗到精的地標優化,我們引入了一種多尺度動態感知策略,用於補丁級特徵金字塔建模。其次,為了有效處理解剖學上相似的模式,我們設計了一種基於相似性的原型學習機制,該機制捕捉了有用的局部語義,以實現穩健的偏移預測。最後,為了穩定模型學習並提高整體性能,我們通過基於容忍度的平衡引入了一種新穎的錯誤感知可靠性正則化。我們收集了一個大型驗證隊列,包括兩個公共數據集和一個私有數據集,涵蓋X光和超聲波模態,涵蓋顱面測量、恥骨聯合-胎頭和胎心地標。大量實驗表明,PPOC-LL在準確性和模型複雜性之間達成了令人滿意的性能和良好的權衡。
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning
2608.09123v1 by Jinkun Hou, Zhuo Liu, Huimin Ren, Hongsheng Xin, Pan Zhou, Kun Zhan
Aligning Large Language Models (LLMs) for open-ended tasks is challenging because responses must satisfy multidimensional criteria without following a single correct generation trajectory. Existing rubric-based reinforcement learning (RL) methods compress fine-grained criterion-level feedback into scalar rewards, making persistent capability gaps difficult to target under limited on-policy exploration. We propose $\textbf{RISE-RL}$ (Rubric-Informed Selective Exploration), which uses repeatedly missed rubric criteria to elicit privileged trajectories that are difficult to discover through unguided exploration alone. RISE-RL retains only trajectories whose complete-rubric reward exceeds the mean reward of natural rollouts, and then re-evaluates them under the original prompt to emphasize behaviors that remain weakly supported by the natural policy. The resulting guidance signal is optimized through a separate auxiliary objective and removed once its additional benefit diminishes. Experiments with 4B and 14B models across writing, chat, health, and science show that RISE-RL achieves the highest mean score on every evaluated benchmark under guidance-free evaluation. Compared with standard Rubric-RL, it improves the average score by 1.3 points at the 4B scale and $\textbf{3.3 points at the 14B scale}$, including a $\textbf{6.0-point}$ gain on CreativeWriting-V3. It also improves creative-writing diversity and yields gains on objectively scored medical and scientific benchmarks. These results indicate that selective internalization through reward filtering and policy support shaping is effective for open-ended reinforcement learning.
摘要:對於開放式任務,對大型語言模型(LLMs)的調整是具有挑戰性的,因為回應必須滿足多維標準,而不必遵循單一正確的生成軌跡。現有的基於評分標準的強化學習(RL)方法將細緻的標準級反饋壓縮為標量獎勵,使得在有限的政策探索下,持續的能力差距難以針對。我們提出了 $\textbf{RISE-RL}$(基於評分標準的選擇性探索),該方法利用重複錯過的評分標準來引出難以通過無引導探索單獨發現的特權軌跡。RISE-RL 僅保留那些完整評分獎勵超過自然回合平均獎勵的軌跡,然後在原始提示下重新評估它們,以強調在自然政策中仍然支持不足的行為。產生的指導信號通過單獨的輔助目標進行優化,並在其額外效益減少後被移除。對於寫作、聊天、健康和科學的 4B 和 14B 模型進行的實驗顯示,RISE-RL 在無指導評估下在每個評估基準上都達到了最高的平均得分。與標準的 Rubric-RL 相比,它在 4B 規模上提高了 1.3 分的平均得分,在 14B 規模上提高了 $\textbf{3.3 分}$,包括在 CreativeWriting-V3 上的 $\textbf{6.0 分}$ 增加。它還改善了創意寫作的多樣性,並在客觀評分的醫療和科學基準上取得了增益。這些結果表明,通過獎勵過濾和政策支持塑造的選擇性內化對於開放式強化學習是有效的。
A Multi-Scale Temporal Framework with Dynamic Fusion for EEG-Based Emotion Recognition
2608.09088v1 by Stefanos Gkikas, Yang Guo, Guangliang Li, Raul Fernandez Rojas, Giorgos Giannakakis, Randy Gomez
Mixed emotions represent a clinically relevant but still underexplored target for automatic emotion recognition. EEG provides millisecond-level access to neural activity, yet most EEG pipelines analyze the signal through a single temporal window, thereby fixing the temporal structure available to the model. This study introduces a multi-scale temporal framework for EEG-based emotion recognition. The EEG waveform is decomposed into windows of one or several durations, processed by a shared attention-based encoder, and integrated through a dynamic fusion module that assigns sample-specific weights across temporal scales. The framework is evaluated under a subject-independent protocol in binary and three-class settings, with the three-class task including the mixed affective category. The best results are 65.22% for the two-class task and 45.43% for the three-class task. Both are obtained with three-scale dynamic-fusion configurations and remain substantially above the full-signal baseline. The best-performing temporal scales differ between the two tasks. Dynamic fusion outperforms concatenation in the highest-scoring two-class configuration and slightly exceeds it in the highest-scoring three-class configuration, although these multi-scale settings require substantially more computation than the full-signal baseline.
摘要:混合情緒代表了一個臨床相關但仍未充分探索的自動情緒識別目標。EEG 提供毫秒級的神經活動訪問,但大多數 EEG 流程通過單一時間窗口分析信號,從而固定了模型可用的時間結構。本研究引入了一個基於 EEG 的情緒識別的多尺度時間框架。EEG 波形被分解為一個或多個持續時間的窗口,通過共享注意力編碼器進行處理,並通過動態融合模塊進行整合,該模塊在時間尺度上分配樣本特定的權重。該框架在一個獨立於受試者的協議下進行評估,涵蓋二元和三類設置,其中三類任務包括混合情感類別。最佳結果為二類任務的 65.22% 和三類任務的 45.43%。這兩者都是在三尺度動態融合配置下獲得的,並且均顯著高於全信號基準。最佳的時間尺度在這兩個任務之間有所不同。在得分最高的二類配置中,動態融合的表現優於串接,而在得分最高的三類配置中略微超過串接,儘管這些多尺度設置所需的計算量遠高於全信號基準。
When Confidence Fails: Overconfidence in LLMs under Uncertainty and Missing Clinical Information
2608.09080v1 by Maryam Tahermazandarani, Adnan Mahmood, Fahmida Islam, Quan Z. Sheng
Large Language Models (LLMs) have achieved strong performance in medical question answering and clinical reasoning tasks. However, their reliability under uncertainty remains poorly understood which raises critical concerns for deployment in high-stakes clinical settings. In such environments, incorrect predictions are inherently risky, but confident incorrect predictions can be particularly harmful as they may mislead clinical decision-making. In this paper, we conduct a systematic behavioral analysis of LLMs under clinical information uncertainty. We propose an evaluation framework based on the MedMCQA dataset consisting of two complementary uncertainty settings. First, we introduce linguistic uncertainty cues through prompt modifications to simulate ambiguous clinical contexts. Second, we construct an answer removal setting, wherein the correct option is deliberately excluded mandating the model to recognize insufficient information and abstain. We analyze both model accuracy and confidence behavior using multiple calibration metrics including calibration gap, Expected Calibration Error (ECE), and Unsafe Confident Error Rate (UCER) across 500 medical questions. Our results reveal a consistent failure mode, i.e., although accuracy degrades under increasing uncertainty, model confidence remains misaligned with accuracy. This leads to a substantial increase in unsafe confident errors, indicating that model confidence remains largely insensitive to clinically meaningful information loss. Furthermore, we observe significant variation across models in their ability to abstain when the correct answer is unavailable, with some models persistently producing high confidence hallucinated answers. These findings expose critical limitations in the epistemic reliability of current LLMs and highlight the need for uncertainty aware evaluation methods prior to their deployment in clinical workflows.
摘要:大型語言模型(LLMs)在醫療問題回答和臨床推理任務中取得了強勁的表現。
然而,它們在不確定性下的可靠性仍然不甚了解,這對於在高風險臨床環境中的部署提出了關鍵的擔憂。
在這樣的環境中,錯誤的預測本質上是有風險的,但自信的錯誤預測可能特別有害,因為它們可能會誤導臨床決策。
在本文中,我們對LLMs在臨床信息不確定性下進行了系統的行為分析。
我們提出了一個基於MedMCQA數據集的評估框架,該數據集包含兩個互補的不確定性設置。
首先,我們通過提示修改引入語言不確定性線索,以模擬模糊的臨床情境。
其次,我們構建了一個答案移除設置,其中正確選項故意被排除,要求模型識別信息不足並選擇不作答。
我們使用多個校準指標分析模型的準確性和信心行為,包括校準差距、期望校準誤差(ECE)和不安全自信錯誤率(UCER),涵蓋500個醫療問題。
我們的結果揭示了一種一致的失敗模式,即儘管準確性在不斷增加的不確定性下下降,但模型的信心與準確性仍然不一致。
這導致不安全的自信錯誤顯著增加,表明模型的信心對臨床上有意義的信息損失仍然大致不敏感。
此外,我們觀察到不同模型在正確答案不可用時的選擇不作答能力上存在顯著差異,一些模型持續產生高信心的虛假答案。
這些發現揭示了當前LLMs在認識論可靠性方面的關鍵局限性,並強調了在其部署於臨床工作流程之前需要不確定性感知的評估方法。
Decoding Phenotypes: A Framework for Fusing Genomic Language Models and Neuroimaging
2608.08926v1 by Tianli Tao, Ziyang Wang, Emma Robinson, Rachel Sparks, Le Zhang
Neuroimaging and genetic testing are two important clinical references for nervous system diseases, offering complementary diagnostic information. However, integrating genomic and neuroimaging data for precise disease diagnosis is challenging due to cross-modality heterogeneity. Existing imaging-genetics approaches mainly encode genetic information as hard-coded labels, which lose the local sequence context around disease-associated variants. To address this limitation, we propose GeneFuse, a multimodal learning framework that aligns genetic representations from pre-trained Genomic Language Models (GLMs) with features extracted from images. GeneFuse integrates two components: (1) Genotype-Conditioned Feature Modulation (GCFM), a FiLM-inspired module that uses genomic embeddings to modulate image feature maps; and (2) Uncertainty-aware Genomic Residual Fusion (U-GRF), a fusion strategy that uses imaging-derived predictive uncertainty to gate the contribution of genotypic features. We evaluate GeneFuse on early cognitive decline identification (NC vs. MCI) and dementia screening (NC vs. AD). In the APOE-centered setting, GeneFuse achieves AUROCs of 0.77 and 0.83, outperforming existing imaging-genetics fusion methods. These results indicate that GLM-derived genomic embeddings provide additional information to imaging.
摘要:神經影像學和基因檢測是神經系統疾病的重要臨床參考,提供互補的診斷信息。
然而,由於跨模態異質性,整合基因組和神經影像數據以進行精確的疾病診斷是具有挑戰性的。
現有的影像-基因學方法主要將基因信息編碼為硬編碼標籤,這樣會失去與疾病相關變異周圍的局部序列上下文。
為了解決這一限制,我們提出了GeneFuse,一個多模態學習框架,將來自預訓練基因組語言模型(GLMs)的基因表示與從影像中提取的特徵對齊。
GeneFuse整合了兩個組件:(1)基因型條件特徵調製(GCFM),一個受FiLM啟發的模塊,使用基因嵌入來調製影像特徵圖;以及(2)不確定性感知基因組殘差融合(U-GRF),一種融合策略,利用影像衍生的預測不確定性來控制基因型特徵的貢獻。
我們在早期認知衰退識別(NC vs. MCI)和癡呆篩查(NC vs. AD)上評估了GeneFuse。
在以APOE為中心的設置中,GeneFuse達到了0.77和0.83的AUROC,超越了現有的影像-基因學融合方法。
這些結果表明,來自GLM的基因嵌入為影像提供了額外的信息。
Toward CT-Equivalent Image Quality in Low-Dose Radiotherapy Planning: Conditional Diffusion-Based CBCT-to-CT Synthesis and the Impact of CBCT Input Representation
2608.08919v1 by Alzahra Altalib, Chunhui Li, Christopher Hamill Taylor, Sankar Pillai, Alessandro Perelli
During standard radiotherapy planning, repeated CT acquisitions are often required for patient registration, verification, and adaptive planning, resulting in increased cumulative X-ray dose. To mitigate this, low-dose cone-beam CT (CBCT) is routinely acquired during treatment delivery. However, CBCT image quality remains insufficient for accurate dose calculation and adaptive radiotherapy planning due to increased scatter, noise, beam hardening, and reconstruction related artifacts. This study develops a supervised deep learning based CBCT to CT synthesis framework using a conditional denoising diffusion probabilistic model (DDPM), where the generation of a CT-based planning for accurate positioning and dose calculation is obtained using generative models with low dose CBCT imaging. Beyond demonstrating CBCT to CT synthesis, the primary objective is to investigate how the representation of CBCT input data, either standard clinical DICOM CBCT images or filtered back-projection (FDK) reconstructions from raw projection data, affects the performance of diffusion based CT synthesis. The overarching aim is to assess whether physics aware CBCT representations better support CT-equivalent image quality while maintaining reduced imaging dose in radiotherapy workflows.
摘要:在標準放射治療計劃中,通常需要重複進行 CT 採集以進行病人登記、驗證和自適應計劃,這導致累積的 X 射線劑量增加。為了減輕這一問題,在治療過程中常規獲取低劑量圓錐束 CT (CBCT)。然而,由於散射、噪聲、束硬化和重建相關的伪影,CBCT 圖像質量仍不足以進行準確的劑量計算和自適應放射治療計劃。本研究開發了一個基於監督式深度學習的 CBCT 到 CT 合成框架,使用條件去噪擴散概率模型 (DDPM),通過生成模型與低劑量 CBCT 成像來獲得準確定位和劑量計算所需的 CT 基礎計劃。除了展示 CBCT 到 CT 的合成,主要目標是研究 CBCT 輸入數據的表示,無論是標準臨床 DICOM CBCT 圖像還是來自原始投影數據的濾波反投影 (FDK) 重建,如何影響基於擴散的 CT 合成性能。總體目的是評估物理感知的 CBCT 表示是否更好地支持 CT 等效的圖像質量,同時在放射治療工作流程中保持降低的成像劑量。
LLM
| Publish Date | Title | Authors | Homepage | Code |
|---|---|---|---|---|
| 2026-08-18 | From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation | Xingjian Wang et.al. | 2608.18076v1 | null |
| 2026-08-18 | Multi-Agent AI System for Radiology Report Structuring and Quality Assurance with Independent Radiologist Evaluation | Iryna Hartsock et.al. | 2608.18072v1 | null |
| 2026-08-18 | TokEval: A Tokenizer Evaluation Suite | Clara Meister et.al. | 2608.18062v1 | null |
| 2026-08-18 | Delegation Asymmetry in Agentic Recommender Systems: Measuring Two-Sided Receptivity in Online Dating | Daria Leshchikova et.al. | 2608.18058v1 | null |
| 2026-08-18 | StagedWorkspace: A Versioned Workspace for Knowledge-Work Agents | Yining Hua et.al. | 2608.18050v1 | null |
| 2026-08-18 | Language Has Two Parameters: Narrative-Induced Semantic Plasticity and Phase-Sensitive Interpretation | Hollis Robbins et.al. | 2608.18041v1 | null |
| 2026-08-18 | Chain-of-Experience for Continual LLM Improvement | Haoqin Tu et.al. | 2608.18027v1 | null |
| 2026-08-18 | Why GPT-Style Models Do Not Directly Transfer to Symbolic Music: Compression in the Wrong Coordinate System | Yi Wang et.al. | 2608.18025v1 | null |
| 2026-08-18 | Can Large Language Models Explain Flight Safety Events? A Prior-Guided Semantic LLM-based Approach | Lu Xu et.al. | 2608.18017v1 | null |
| 2026-08-18 | The IOL-AI Challenge: An Open Challenge towards Advancing Linguistic Reasoning | Eduardo Sánchez et.al. | 2608.18011v1 | null |
| 2026-08-18 | Policy-Invariant Reward Shaping from LLM Feedback: A Framework for Hybrid RL Agents | Christophe D. Hounwanou et.al. | 2608.18008v1 | null |
| 2026-08-18 | Traceable Trust for action-ready artificial intelligence in bioscience | Huayu Xin et.al. | 2608.17997v1 | null |
| 2026-08-18 | Judge, Retrieve, or Abstain: Uncertainty-Guarded LLM Judging with Provable Risk Guarantees | Sher Badshah et.al. | 2608.17994v1 | null |
| 2026-08-18 | Against Political Polarization: A Unified Framework for Tracing Evolving Political Ideologies on Social Media | Yijie Xu et.al. | 2608.17987v1 | null |
| 2026-08-18 | Dual Co-Train: Cross-Dataset Ultrasound Tongue Segmentation Under Extreme Data Scarcity | Alisher Myrgyyassov et.al. | 2608.17983v1 | null |
| 2026-08-18 | When Writing Style Drifts: Benchmarking Authorship Verification under Distribution Shifts in Genre, Time and the AI-Era | Lotta Kiefer et.al. | 2608.17979v1 | null |
| 2026-08-18 | Too Sure to Be Safe: Model Calibration for Reliable Log Anomaly Detection | Bin Li et.al. | 2608.17965v1 | null |
| 2026-08-18 | Towards Zero-Shot Task Transfer with Neurosymbolic World Models | Isidoro Tamassia et.al. | 2608.17959v1 | null |
| 2026-08-18 | An Omitted Mode Is a Rare Rule: The Sampling-Verification Danger Law in Continuous Code World Models | Javier Aguilar Martín et.al. | 2608.17956v1 | null |
| 2026-08-18 | Do Large Language Models Play Six Degrees of Separation? Measuring Topological Compression in Long-Context Manifolds | Md. Faiyaz Abdullah Sayeedi et.al. | 2608.17950v1 | null |
| 2026-08-18 | SIGMA: SHAP-Guided Implicit-Trajectory Generation for Metadata-Free LLM-Based AutoFE | Xuan Zheng et.al. | 2608.17948v1 | null |
| 2026-08-18 | Procedural Content Metageneration via Program Search and Continual Abstraction Discovery | Matthew Siper et.al. | 2608.17947v1 | null |
| 2026-08-18 | Efficient RLVR Scheduling via Graph-Structured Online Difficulty Estimation | Zhizhao Liu et.al. | 2608.17941v1 | null |
| 2026-08-18 | Grading Needs a Rubric, Not Intelligence | Jhen-Ke Lin et.al. | 2608.17938v1 | null |
| 2026-08-18 | EvoTS-Agent: A Self-Evolving LLM Agent for Financial Time Series Change Point Detection | Lei Jiang et.al. | 2608.17933v1 | null |
| 2026-08-18 | Collective Counterfactual Planning: Coordination, Consent, and Verification under Representational Constraints | Chainarong Amornbunchornvej et.al. | 2608.17932v1 | null |
| 2026-08-18 | SpeechSense: A Paralinguistic-Focused Dataset for Fine-Grained Speech Sentiment Analysis | Shicheng Ma et.al. | 2608.17931v1 | null |
| 2026-08-18 | Comparative Study of Out-of-the-Box Technology for Automatic Target Detection and Recognition | Alma M. Liezenga et.al. | 2608.17917v1 | null |
| 2026-08-18 | CABLE: Extending the Reach of Memory Retrieval via Complementary Antecedent-Based Linking and Expansion | Zheling Tan et.al. | 2608.17911v1 | null |
| 2026-08-18 | AutoResearch: Insight In, Hallucination Out | Yiming Ren et.al. | 2608.17906v1 | null |
| 2026-08-18 | BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models | Liubov Chubarova et.al. | 2608.17895v1 | null |
| 2026-08-18 | BayesPrompt: human readable prompts that make sense | Franky Kevin Nando Tezoh et.al. | 2608.17866v1 | null |
| 2026-08-18 | ARASH: Adaptive Retrieval And Shot Selection for Tabular Prediction | Samirasadat Jamalidinan et.al. | 2608.17856v1 | null |
| 2026-08-18 | Encoded but Not Actionable: Auditing the Decode-Generate-Steer Gap in Frozen LLMs for Geometric Constraints | Man Liang et.al. | 2608.17843v1 | null |
| 2026-08-18 | AdaLens: Interactive Storyline for Monitoring and Steering Long-Running Agentic Data Analysis | Yangtian Liu et.al. | 2608.17834v1 | null |
| 2026-08-18 | The Model's Tell: Measuring Context-Leakage Attack Signals with Behavior Gauges | Maosen Zhang et.al. | 2608.17829v1 | null |
| 2026-08-18 | From Global Benchmarks to Local Evaluations: Benchmarking LLMs for the German Public Sector | Camilla Dalerci et.al. | 2608.17827v1 | null |
| 2026-08-18 | MotoSafety: Edge-AI with Learned Temporal Importance for Two-Wheeler Collision Risk Assessment Under Time Pressure | Sumit S. Shevtekar et.al. | 2608.17823v1 | null |
| 2026-08-18 | Interpretable Humans, Alien LLMs: Expert Analysis of Latent Structures in Assessment Responses | Alona Strugatski et.al. | 2608.17810v1 | null |
| 2026-08-18 | Whether LLMs Can Navigate Beliefs and Facts Depends on How You Phrase It | Quang Minh Nguyen et.al. | 2608.17809v1 | null |
| 2026-08-18 | An Empirical Study of Reward Specification and Benchmark Reliability in GRPO-based LLM Unlearning | Rubén Balbastre et.al. | 2608.17804v1 | null |
| 2026-08-18 | StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows | Liya Zhu et.al. | 2608.17800v1 | null |
| 2026-08-18 | Training with synthetic data for drone detection in thermal imagery | Tanel Liiv et.al. | 2608.17799v1 | null |
| 2026-08-18 | TraceSQL: Traceable Answerability Estimation for Reference-Free Text-to-SQL Verification | Neelesh Kumar Shukla et.al. | 2608.17795v1 | null |
| 2026-08-18 | Preference Is Not Intervention: The Structure and Stability Boundaries of Reader-Specific Evidence Utility | Shi Zhou et.al. | 2608.17781v1 | null |
| 2026-08-18 | Learnware for CSI Feedback: Scene-specific Small Models Can Do Big | Xiangyi Li et.al. | 2608.17760v1 | null |
| 2026-08-18 | D$^2$ACCI: A Dual-Loop Diagnostic Protocol for Evidence-Preserving Agent Memory | Xule Liu et.al. | 2608.17756v1 | null |
| 2026-08-18 | The Curious Case of Exploding DecPOMDPs: Containing the Fire through Policy Counting | Nazlı Nur Karabulut et.al. | 2608.17749v1 | null |
| 2026-08-18 | Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See | Ayoub Kirouane et.al. | 2608.17744v1 | null |
| 2026-08-18 | Neuro-symbolic learning over OWL 2 DL via consequence-based compilation to differentiable circuits | Olga Mashkova et.al. | 2608.17741v1 | null |
| 2026-08-18 | What Aggregate Scores Miss: Measuring Item-Level Regressions in Commercial LLM API Migrations | Xiaonan Xu et.al. | 2608.17719v1 | null |
| 2026-08-18 | Beyond Suspicious Steps: Ontological Trust in Long-Horizon Agents | An He et.al. | 2608.17718v1 | null |
| 2026-08-18 | Communicating Credit Risk with Large Language Models: Evaluation of Explanations from Standard and Alternative Data-Based Models | Sahab Zandi et.al. | 2608.17715v1 | null |
| 2026-08-18 | Accuracy and Robustness of Model Cascades Under Data Perturbations | Pallavi Mitra et.al. | 2608.17711v1 | null |
| 2026-08-18 | GADR: Gathering Architecture Decision Records from Meeting Transcriptions | Lucas Daniel Costa da Silva et.al. | 2608.17694v1 | null |
| 2026-08-18 | Mixture-of-Expert Blocks Contain Strong Hallucination Detection Signals | Joao Fonseca et.al. | 2608.17687v1 | null |
| 2026-08-18 | Benchmarking Automated Security Patch Backporting: How Far Are We? | Jincheng Yang et.al. | 2608.17671v1 | null |
| 2026-08-18 | GraphWake: Group Polarization via Memory-Mediated Polarization Cascade in LLM-Agent Communities | Haoran Bu et.al. | 2608.17665v1 | null |
| 2026-08-18 | MobileWorldSafety: Benchmarking GUI Agent Safety Against Environmental Injection Attacks in Android Apps | Sujin Chen et.al. | 2608.17659v1 | null |
| 2026-08-18 | LLM-Derived Preference Judgments Are Not Self-Consistent | Matthew T. Ford et.al. | 2608.17644v1 | null |
| 2026-08-18 | Beyond the Trace: Coupling an Interpretable Reasoning-State Readout to Native MoE Routing | Kang Chen et.al. | 2608.17638v1 | null |
| 2026-08-18 | Graph Surgery and the Do-Operator: A Precise Correspondence for Acyclic Structural Causal Models | Satpreet Makhija et.al. | 2608.17634v1 | null |
| 2026-08-18 | DEPT: Document Embedding Preservation Tuning for Unified Query Expansion and Retrieval | Jingyuan Wang et.al. | 2608.17632v1 | null |
| 2026-08-18 | From Student Risk Prediction to SC2R: Semantics-Constrained Counterfactual Recourse for Educational Decision Support | Ngoc Luyen Le et.al. | 2608.17618v1 | null |
| 2026-08-18 | Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges | Syeda Faiza Ahmed et.al. | 2608.17605v1 | null |
| 2026-08-18 | HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety | Yajing Bai et.al. | 2608.17597v1 | null |
| 2026-08-18 | tinyDSM: A Framework for Skill Modeling and Development for Resource-Constrained Millirobots | Markus D. Kobelrausch et.al. | 2608.17596v1 | null |
| 2026-08-18 | TRUSS: Towards Task-Reliable and User-Safe Automated Agent Skill Generation | Zhibo Zhang et.al. | 2608.17588v1 | null |
| 2026-08-18 | Write, Execute, Refine: From Skill Followers to Skill Optimizers via Reinforcement Learning from Execution Feedback | Kang Peng et.al. | 2608.17587v1 | null |
| 2026-08-18 | Auditing Exposure to Harmful Content on TikTok using Multimodal Language Models: A Cross-National, Age-Stratified Study | Hamidreza Saffari et.al. | 2608.17583v1 | null |
| 2026-08-18 | Quantifying Risk Under Evolving Uncertainty: Belief-Dependent Robustness for Safe Sequential Decision Making | Deep Kumar Ganguly et.al. | 2608.17574v1 | null |
| 2026-08-18 | DMT-Dens: Density-preserving manifold visualization for biological data | Ruizhe Wang et.al. | 2608.17571v1 | null |
| 2026-08-18 | Domain-Adapted Molecular Language Models for Efficient Search of Make-on-Demand Libraries | Henrik Wille et.al. | 2608.17567v1 | null |
| 2026-08-18 | Where a New Concept Must Enter: Entry Point Gates Cross-Task Usability in Unified Multimodal Models | Zongyang Qiu et.al. | 2608.17564v1 | null |
| 2026-08-18 | Reflex-Guard: A Low-Latency Guardrail for LLM Prompt Safety Using Dense Semantic Embeddings | Istiaque Ahmed et.al. | 2608.17556v1 | null |
| 2026-08-18 | Code as Representation: A Compilable Parsing Paradigm for Academic Documents | Rihui Jin et.al. | 2608.17550v1 | null |
| 2026-08-18 | No Gaussian Required: Contrastive Inverse Dynamics for JEPA World Models | Jack Boylan et.al. | 2608.17542v1 | null |
| 2026-08-18 | CoAL-RAG: A Complexity-Aware Legal Retrieval-Augmented Generation Method | Jin Su et.al. | 2608.17536v1 | null |
| 2026-08-18 | ArborMem: Navigating Interaction States with Memory Forests | Zongwei Lv et.al. | 2608.17534v1 | null |
| 2026-08-18 | When to Review: Spaced Repetition for Continual Pre-Training of Language Models | Alankar Atreya et.al. | 2608.17530v1 | null |
| 2026-08-18 | Agent Lightning v1.0: Towards Harnessed Agentic RL | Zhiyuan He et.al. | 2608.17528v1 | null |
| 2026-08-18 | Explainable AI-Powered Framework for Video-Based Skill Assessment in Cataract Surgery | Mohammad Javad Ahmadi et.al. | 2608.17522v1 | null |
| 2026-08-18 | Effects of Answer Format Variation on Gender Bias in Large Language Models | Ksenia Merzlyakova et.al. | 2608.17516v1 | null |
| 2026-08-18 | Beyond FLOPs: Energy-Aware Knowledge Distillation for Sustainable LLMs on Code-Related Task | Enrique Barba Roque et.al. | 2608.17515v1 | null |
| 2026-08-18 | SGHA: Evidence-Grounded Research Problem Discovery with Local Language Models | Sarvesh Gharat et.al. | 2608.17501v1 | null |
| 2026-08-18 | When AI Designs AI: Innovation or Imitation? | Yikang Yang et.al. | 2608.17471v1 | null |
| 2026-08-18 | SAGE: Self-Evolving Storyboard Skills via Attribution-Guided Rule Evolution | Maolin Ran et.al. | 2608.17468v1 | null |
| 2026-08-18 | From Entity Mentions to Tone: An LLM-Based Pipeline for Media Bias Analysis | Klesti Hoxha et.al. | 2608.17454v1 | null |
| 2026-08-18 | Decomposition Attacks Across Unlinkable Identities: Limits of Stateful Defenses for LLM Services | Bowen Sun et.al. | 2608.17445v1 | null |
| 2026-08-18 | Structure-Internalized Rule Language Model for Faithful Knowledge Graph Reasoning | Xingrui Zhuo et.al. | 2608.17443v1 | null |
| 2026-08-18 | Task-Aware Harness Provisioning for LLM Agents in Mission-Critical Infrastructure Operations | Liangtao Lin et.al. | 2608.17433v1 | null |
| 2026-08-18 | SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation | Keyu Tu et.al. | 2608.17426v1 | null |
| 2026-08-18 | An Investigation of Translationese in the Generations of Multilingual Large Language Models | Maria Valentini et.al. | 2608.17399v1 | null |
| 2026-08-18 | LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents | Yiming Du et.al. | 2608.17393v1 | null |
| 2026-08-18 | Leveraging generative hallucination and biophysics-informed modeling for unified biomolecular sequence-structure co-design | Xuefeng Liu et.al. | 2608.17381v1 | null |
| 2026-08-18 | PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX | Genghan Zhang et.al. | 2608.17379v1 | null |
| 2026-08-18 | Fair ASR: Re-Evaluating Black-Box Jailbreaks under Shared Target-Call Budgets | Zhida He et.al. | 2608.17360v1 | null |
| 2026-08-18 | Cognitive Graph Intelligence for Adaptive and Robust DDoS Attack Detection in Next Generation Networks | Mohammad Arif Hossain et.al. | 2608.17352v1 | null |
| 2026-08-18 | MoFE: A Novel Mixture-of-Experts Framework with Fourier Neural Operators for Cryptocurrency Forecasting | Bowen Liu et.al. | 2608.17342v1 | null |
| 2026-08-18 | LLM-Only PDDL Domain Repair with Open-Weight Models | Nader Karimi Bavandpour et.al. | 2608.17341v1 | null |
Abstracts
From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation
2608.18076v1 by Xingjian Wang, Zhao Wang, Taihang Hu, Jun Zheng, Qing Jin, Qinye Zhou, Zhengtao Wu, Yongchao Du, Zuan Gao, Chao Lin, Yefeng Shen, Xiaoli Xu, Zhengze Xu, Hao Yan, Yuhang Yu, Mingzhou Zhang, Mengting Chen
Large-scale image generation has benefited from advances in data scale, quality, rebalancing, and recaptioning, yet conventional pipelines typically optimize task-specific datasets in isolation. A central challenge is not only how to curate each task-specific corpus, but also how to organize heterogeneous supervision according to the dependencies among generative capabilities. We present a \textbf{capability-driven data infrastructure} that couples capability-specific supervision construction with capability-aligned curriculum scheduling. Its three specialized yet interoperable data engines build complementary relational supervision for text-image grounding, inter-image transformation, and image-knowledge association, while caption experts align T2I and editing supervision across tasks and granularities. A multi-stage curriculum jointly evolves task composition, visual-concept distribution, data quality, and image resolution along the dependency order of capability acquisition, with capability-aware evaluation closing the loop through targeted retrieval, expert construction, and gap-aware resampling. At scale, the framework curates a 440M-image T2I corpus, 120M editing pairs, and over 27M image-entity pairs. With this infrastructure, we train multimodal diffusion models at two scales from scratch, with 3B and 6B sizes respectively. We conduct quantitative evaluation on CPI-Bench, along with qualitative evaluations across diverse text-to-image and editing scenarios. Experimental results present broad visual coverage, versatile rendering, and effective transfer across generative capabilities.
摘要:大規模圖像生成受益於數據規模、質量、重新平衡和重新標題的進步,但傳統流程通常在孤立的情況下優化特定任務的數據集。一個主要挑戰不僅在於如何策劃每個特定任務的語料庫,還在於如何根據生成能力之間的依賴關係組織異質監督。我們提出了一個\textbf{以能力為驅動的數據基礎設施},將特定能力的監督構建與能力對齊的課程安排結合起來。它的三個專門但可互操作的數據引擎為文本-圖像基礎、圖像間轉換和圖像-知識關聯構建互補的關係監督,同時標題專家在任務和粒度之間對齊T2I和編輯監督。一個多階段課程共同演變任務組合、視覺概念分佈、數據質量和圖像解析度,沿著能力獲取的依賴順序進行,而以能力為中心的評估通過針對性檢索、專家構建和關注差距的重採樣來閉合循環。在規模上,該框架策劃了一個包含4.4億圖像的T2I語料庫、1.2億編輯對和超過2700萬圖像-實體對。利用這一基礎設施,我們從零開始訓練了兩個規模的多模態擴散模型,分別為30億和60億大小。我們在CPI-Bench上進行了定量評估,並在多樣的文本到圖像和編輯場景中進行了定性評估。實驗結果顯示出廣泛的視覺覆蓋、多樣的渲染和在生成能力之間的有效轉移。
Multi-Agent AI System for Radiology Report Structuring and Quality Assurance with Independent Radiologist Evaluation
2608.18072v1 by Iryna Hartsock, Cesar Lam, Christopher Otteni, Aliya Qayyum, Robert Gatenby, Cyrillo Araujo, Ghulam Rasool
Purpose: To develop and evaluate a locally deployed multi-agent AI system for radiology report structuring and quality assurance. Materials and Methods: This retrospective study included 638 radiology reports from CT examinations of the chest, abdomen, and pelvis dictated by 15 board-certified radiologists in 2023 and 2024. A multi-agent AI pipeline was developed to perform report structuring and quality assurance (QA). The system structured the report into standardized anatomical sections at the sentence level using regex rules and local large language models. It also detected mismatches between the Findings and Impression sections, or within sections; gender-anatomy conflicts; and undocumented communication of critical findings. Two board-certified radiologists independently evaluated a 45-report subset. Results: The multi-agent system structured the Findings sections of all reports (22,270 sentences) into a predefined anatomical format while retaining the original report content. The system flagged 90 (14.1%) reports, most commonly for section mismatches (80 reports, 12.5%). In the radiologist evaluation, both reviewers agreed that 31 (69%) were correctly restructured, 2 reports (4%) were incorrectly restructured, and disagreed on the remaining 12 reports (27%). Both reviewers agreed that no clinically important information was omitted and no fabricated content was introduced. Overall QA performance was rated as "excellent" or "good" in 84% of the evaluated reports, with the remaining reports rated as "fair". Conclusion: A locally deployed multi-agent AI system combined radiology report structuring and quality assurance within a single workflow. The system demonstrated favorable performance in radiologist evaluation. Such systems may support standardization of reporting and quality assurance in radiology practice.
摘要:目的:開發和評估一個本地部署的多代理人工智慧系統,用於放射學報告的結構化和質量保證。
材料和方法:本回顧性研究包含了2023年和2024年由15位董事會認證的放射科醫師口述的638份胸部、腹部和骨盆的CT檢查報告。
開發了一個多代理人工智慧管道來執行報告結構化和質量保證(QA)。
該系統使用正則表達式規則和本地大型語言模型將報告結構化為標準化的解剖學部分,並在句子層面進行處理。
它還檢測到發現和印象部分之間的錯配,或部分內部的錯配;性別-解剖學衝突;以及對關鍵發現的未記錄溝通。
兩位董事會認證的放射科醫師獨立評估了45份報告的子集。
結果:該多代理系統將所有報告的發現部分(22,270句)結構化為預定的解剖格式,同時保留原始報告內容。
系統標記了90份(14.1%)報告,最常見的原因是部分錯配(80份報告,12.5%)。
在放射科醫師的評估中,兩位評審一致認為31份(69%)報告重構正確,2份報告(4%)重構不正確,對剩餘的12份報告(27%)意見不合。
兩位評審一致認為沒有遺漏臨床重要信息,也沒有引入虛構內容。
整體質量保證表現被評為“優秀”或“良好”的報告佔84%,其餘報告被評為“公平”。
結論:一個本地部署的多代理人工智慧系統在單一工作流程中結合了放射學報告的結構化和質量保證。
該系統在放射科醫師評估中顯示出良好的表現。
這類系統可能支持放射學實踐中的報告標準化和質量保證。
TokEval: A Tokenizer Evaluation Suite
2608.18062v1 by Clara Meister
Language model tokenizers are typically selected with minimal evaluation, despite the fact that their design choices directly impact model capabilities. This can be partly attributed to a limited understanding of which tokenizer properties affect which aspects of downstream performance. We introduce TokEval, a framework of tokenizer evaluation metrics that goes beyond standard measures like fertility and compression rate to capture linguistically and structurally meaningful properties, e.g., UTF-8 character boundary integrity and digit place-value boundary alignment for mathematics. To validate whether these metrics are predictive of downstream model performance, we conduct controlled language model pretraining experiments, varying solely the tokenizers' training data mixture, pretokenization strategy, and training algorithm. We evaluate the resulting models on bits-per-byte (a tokenizer-agnostic version of perplexity) and several benchmarks, spanning linguistic understanding, mathematical reasoning, and code generation. Our experiments suggest that different intrinsic properties have different impacts on model abilities: information-theoretic metrics predict language modeling abilities (Spearman rho up to 0.80), while structure-sensitive metrics, such as those measuring digit and line-break handling, correlate with task accuracy. We hope TokEval enables more principled tokenizer evaluation, replacing pretraining sweeps with intrinsic measurement wherever the two agree.
摘要:語言模型的分詞器通常在最小評估下被選擇,儘管它們的設計選擇直接影響模型的能力。這部分可以歸因於對哪些分詞器特性影響下游性能的理解有限。我們引入了 TokEval,一個超越標準測量(如生育率和壓縮率)的分詞器評估指標框架,以捕捉語言和結構上有意義的特性,例如 UTF-8 字符邊界的完整性和數字位值邊界對齊以進行數學運算。為了驗證這些指標是否能預測下游模型性能,我們進行了受控的語言模型預訓練實驗,僅改變分詞器的訓練數據混合、預分詞策略和訓練算法。我們在每字節位數(這是一個與分詞器無關的困惑度版本)和幾個基準上評估了結果模型,涵蓋語言理解、數學推理和代碼生成。我們的實驗表明,不同的內在特性對模型能力有不同的影響:信息理論指標預測語言建模能力(斯皮爾曼相關係數高達 0.80),而結構敏感指標,例如測量數字和換行處理的指標,則與任務準確性相關。我們希望 TokEval 能夠實現更有原則的分詞器評估,並在兩者一致的地方用內在測量取代預訓練掃描。
Delegation Asymmetry in Agentic Recommender Systems: Measuring Two-Sided Receptivity in Online Dating
2608.18058v1 by Daria Leshchikova, Valentina V. Kuskova, Dmitry Zaytsev, Valerii Klimov
Autonomous LLM agents that converse on a user's behalf are an emerging design pattern in matching platforms, yet their viability depends on a condition rarely examined: users must accept not only delegating conversation to an agent, but also receiving agent-mediated communication from others. We study this condition using two large-scale surveys of active users of a major dating platform (N=2,894 on generative profile features; N=2,617 on autonomous conversational agents, fielded in two languages). We develop a latent-variable measurement model of agent receptivity based on graded response models with latent regression, and show via model comparison that willingness to send and willingness to receive agent communication are distinct constructs: highly correlated (rho=0.92) but separable (Delta BIC=52), with partial measurement invariance across languages. The model quantifies a systematic delegation asymmetry: deploying one's own agent requires far lower receptivity (threshold -0.38) than engaging a counterpart's agent (+0.32; full engagement +1.39), and mean deployment propensity exceeds engagement propensity roughly threefold. Under a random-pairing counterfactual derived from stated receptivity, only 4-13% of directed dyads combine agent deployment with receiver engagement, with a pronounced gender-directional imbalance. Design counterfactuals quantify the levers: a reciprocity requirement cuts interaction volume by half or more by excluding nearly two-thirds of would-be deployment, while routing agent contacts on receive receptivity triples per-contact engagement, a lift that survives out-of-sample validation with the target item held out (AUC 0.88, 3.1x quartile lift under respondent-level cross-validation). We discuss implications for agentic recommender design, including disclosure, opt-in mechanics, and receptivity-aware matchmaking.
摘要:自主 LLM 代理人代表用戶進行對話是一種新興的設計模式,尤其在配對平台上,但其可行性取決於一個鮮少被檢視的條件:用戶必須接受不僅將對話委託給代理人,還要接受來自他人的代理人中介通信。我們通過對一個主要約會平台的活躍用戶進行兩項大規模調查來研究這一條件(N=2,894 針對生成型個人資料特徵;N=2,617 針對自主對話代理人,調查以兩種語言進行)。我們基於潛在回應模型和潛在回歸開發了一個代理人接受度的潛變量測量模型,並通過模型比較顯示,發送代理人通信的意願和接收代理人通信的意願是不同的構念:高度相關(rho=0.92)但可分離(Delta BIC=52),在語言間具有部分測量不變性。該模型量化了一種系統性的委派不對稱:部署自己的代理人所需的接受度(閾值 -0.38)遠低於與對方的代理人互動所需的接受度(+0.32;完全互動 +1.39),而平均部署傾向約為互動傾向的三倍。在基於聲明的接受度推導的隨機配對反事實中,只有 4-13% 的定向雙人組合將代理人部署與接收者互動結合,並且存在明顯的性別導向不平衡。設計反事實量化了杠杆:互惠要求通過排除近三分之二的潛在部署將互動量減少一半或更多,而根據接收接受度路由代理人聯繫則使每次聯繫的互動增加三倍,這一提升在樣本外驗證中仍然有效,目標項目被保留(AUC 0.88,受訪者層級交叉驗證下的四分位提升 3.1 倍)。我們討論了對代理推薦設計的影響,包括披露、選擇加入機制和接受度敏感的配對。
StagedWorkspace: A Versioned Workspace for Knowledge-Work Agents
2608.18050v1 by Yining Hua, Hongbin Na, Yifan Zhou, Akshay Kalose, Cyrus Ayubcha, Levi Lian
AI agents increasingly perform knowledge work (i.e., produce and modify persistent digital artifacts such as code repositories, documents, spreadsheets, slides, reports), yet the parsed views they search, the native files they edit, the changes they review, and the artifacts they submit can refer to different versions of the same work product. We formulate this as a workspace-state contract: every view should be explicitly tied to a version of the evolving workspace state. Coding agents partly address this need through repository contracts for search, diffs, and tests, whereas an analogous contract is less explicit for PDFs, spreadsheets, slides, notebooks, and mixed-format project folders. We propose StagedWorkspace, a versioned workspace for knowledge-work agents. The workspace binds parsed records and review diffs to content hashes of the native files as they change. In fixed-harness ablations on OfficeQA Pro and APEX-Agents, dual parsed/native access has the highest point estimate for every tested model; relative to the more limiting single view, it improves OfficeQA Pass@1 by 8.3-12.1 points and APEX mean rubric score by 4.7-9.2 points. SW-AGENT scores 63.9% with Gemini 3.1 Pro on OfficeQA and 42.1 with GPT-5.4 Nano on APEX, compared with published same-model scores of 29.3% and 25.5, respectively. A paired review-axis ablation on 57 file-editing tasks further finds higher observed scores when diffs are visible. These results identify workspace state as an experimental variable in knowledge-work agents and motivate benchmarks that score evidence, staged edits, and submitted artifacts as explicit state transitions.
摘要:AI 代理人越來越多地執行知識工作(即,產生和修改持久的數位工件,如代碼庫、文件、電子表格、簡報、報告),然而他們所搜尋的解析視圖、編輯的原始文件、審查的變更以及提交的工件可能指的是同一工作產品的不同版本。我們將此表述為工作區狀態合約:每個視圖應明確與不斷演變的工作區狀態的某個版本相關聯。編碼代理人部分通過針對搜索、差異和測試的庫合約來滿足這一需求,而對於 PDF、電子表格、簡報、筆記本和混合格式的項目文件夾,類似的合約則不那麼明確。我們提出了 StagedWorkspace,一個針對知識工作代理人的版本化工作區。該工作區將解析記錄和審查差異綁定到隨原始文件變更的內容哈希。在 OfficeQA Pro 和 APEX-Agents 的固定裝置消融實驗中,雙重解析/原生訪問對於每個測試模型的最高點估計;相對於更具限制性的單一視圖,它將 OfficeQA Pass@1 提高了 8.3-12.1 分,將 APEX 的平均評分提高了 4.7-9.2 分。SW-AGENT 在 OfficeQA 上的得分為 63.9%,在 APEX 上的得分為 42.1,與已發表的同模型得分分別為 29.3% 和 25.5 相比。對 57 個文件編輯任務的配對審查軸消融進一步發現,當差異可見時,觀察到的得分更高。這些結果將工作區狀態確定為知識工作代理人的實驗變量,並激勵對證據、分階編輯和提交工件進行明確狀態轉換的基準評分。
Language Has Two Parameters: Narrative-Induced Semantic Plasticity and Phase-Sensitive Interpretation
2608.18041v1 by Hollis Robbins
Language has two parameters. Count how often words occur together and you estimate amplitude, the strength of association. Word embeddings and attention weights refine that count, which sums every writer in the corpus together. This paper claims a second parameter, phase, which signed weights learned from a corpus do not supply. Phase exists only between meanings: it determines how coactivated meanings combine, and it can reverse what a meaning contributes while that meaning stays fully present. A speaker can set phase in the signal through linguistic form; encounters install phase relations and history distributes them. Population averaging deletes history-indexed phase: agent-deindexed corpora identify the population marginal state and determine no individual or dyadic state, at any scale. The standard transformer has no explicit representation for phase in frozen inference, and the interpretability program measuring progress by monosemanticity is optimizing against it: the coexistence it treats as a defect is the condition of allusion, irony, and quotation. Six predictions test whether a suppressed meaning stays active, whether encounter order changes what a phrase does, whether marking the signal changes how a shared phrase is taken, and whether a model given a history is changed by it or only informed about it. The claim defended is the weak version: interpretation requires a second relational parameter, signed, persistent, and indexed to individuals and dyads. Quantum probability is one notation for the parameter; nothing in the formalism claims quantum processes in the brain. The strong version, that the quantum calculus constrains these phenomena as signed classical models do not, rests on an encounter-order constraint not yet derived. The architecture the theory calls for is a language model with agent-indexed, phase-bearing semantic states.
摘要:語言有兩個參數。計算單詞共同出現的頻率,你可以估算幅度,即聯繫的強度。詞嵌入和注意力權重會細化這個計數,這個計數將語料庫中的每位作者的貢獻相加。本文主張第二個參數,階段,這是從語料庫學習的簽名權重所無法提供的。階段僅存在於意義之間:它決定了如何共同激活的意義結合,並且它可以反轉一個意義所貢獻的內容,同時該意義仍然完全存在。說話者可以通過語言形式在信號中設置階段;遭遇安裝階段關係,而歷史則分配它們。人口平均會刪除歷史索引的階段:去代理索引的語料庫識別出人口的邊際狀態,並且在任何規模上都不確定任何個體或雙方狀態。標準Transformer在凍結推理中沒有對階段的明確表示,而測量進展的可解釋性程序則是針對它進行優化的:它所視為缺陷的共存是暗示、諷刺和引用的條件。六個預測測試被壓制的意義是否保持活躍,遭遇順序是否改變短語的作用,標記信號是否改變共享短語的理解,以及給定歷史的模型是否因其而改變或僅僅是被告知。所辯護的主張是弱版本:解釋需要第二個關係參數,這個參數是簽名的、持久的,並且索引到個體和雙方。量子概率是一種參數的表示法;形式主義中沒有任何內容聲稱大腦中的量子過程。強版本,即量子微積分限制這些現象,而簽名的經典模型則不,依賴於尚未推導出的遭遇順序約束。該理論所要求的架構是一個具有代理索引、承載階段的語義狀態的語言模型。
Chain-of-Experience for Continual LLM Improvement
2608.18027v1 by Haoqin Tu, Yunhao Fang, Yizhong Wang, Cihang Xie, Shen Yan
Humans continuously learn from experience, whereas conventional large language model (LLM) evaluations ignore the models' ability to improve through inference-time interaction. In this paper, we study how LLMs learn from iterative experience at test time, a setting we refer to as Chain-of-Experience (CoE), where models accumulate experiential traces through iterative interactions with self or environmental feedback to form a continual improvement loop beyond zero-shot inference. We instantiate CoE with diverse feedback mechanisms, including model self-feedback and environmental signals such as correctness or public coding test pass rates, and evaluate across math, coding, and knowledge domains using 8 LLMs, including GPT-5, Gemini-2.5 Pro, Claude-4.5 Sonnet. Our study shows that leveraging iterative experience consistently outperforms feedback-free baselines, achieving substantial gains with self feedback alone, alongside a 5.6% overall improvement and 19% lower API cost across tasks and models. We further show that combining complementary feedback channels (e.g., model and correctness signals) yields additional gains, and that CoE delivers higher accuracy per token than existing test-time strategies. We observe a positive correlation between LLM base ability and improvement capacity, and show that models remain robust under weak or spurious feedback, with different feedback contributing to distinct improvement aspects and most gains emerging early in the iterations.
摘要:人類不斷從經驗中學習,而傳統的大型語言模型(LLM)評估則忽略了模型通過推理時互動來改進的能力。在本文中,我們研究了 LLM 如何在測試時從迭代經驗中學習,這種情境我們稱之為經驗鏈(Chain-of-Experience, CoE),在這裡模型通過與自身或環境反饋的迭代互動積累經驗痕跡,以形成超越零-shot 推理的持續改進循環。我們用多樣的反饋機制來實現 CoE,包括模型自我反饋和環境信號,如正確性或公共編碼測試通過率,並使用 8 種 LLM 進行數學、編碼和知識領域的評估,包括 GPT-5、Gemini-2.5 Pro 和 Claude-4.5 Sonnet。我們的研究表明,利用迭代經驗的表現始終優於無反饋的基準,僅依靠自我反饋就實現了顯著的增益,並在各任務和模型中達到 5.6% 的整體改進和 19% 的 API 成本降低。我們進一步顯示,結合互補的反饋通道(例如模型和正確性信號)會產生額外的增益,並且 CoE 在每個 token 上提供的準確性高於現有的測試時策略。我們觀察到 LLM 的基本能力與改進能力之間存在正相關,並顯示模型在弱或虛假反饋下仍然保持穩健,不同的反饋對不同的改進方面有所貢獻,大多數增益在迭代的早期出現。
Why GPT-Style Models Do Not Directly Transfer to Symbolic Music: Compression in the Wrong Coordinate System
2608.18025v1 by Yi Wang
GPT-style models achieve strong performance by representing language with finite vocabularies of reusable discrete tokens. This success has motivated symbolic music tokenizations to treat recurring musical structures, such as chords, motifs, and phrases, as reusable units analogous to linguistic tokens. However, tokenization derives its advantage not from reusable combinations alone, but from compression: effective compression requires coordinates in which recurring regularities form stable and predictable conditional distributions. The key problem is therefore not to find larger musical combinations, but to discover the coordinate system in which musical facts become predictively compressible. We formulate the Effectiveness--Losslessness Framework and define tokenization as the construction of a predictively effective and relationally lossless coordinate system. The Predictive Effectiveness Principle defines the Fact--Token Boundary: decoupling and denesting construct coordinate interfaces that expose predictive regularities. The Relational Losslessness Principle defines the Token--State Boundary: tokenization stops before context-dependent relations are fixed, leaving their computation to model states. Controlled symbolic-music experiments validate these boundaries. Effective coordinate construction improves predictive compressibility, while fixed relational projections constrain contextual modeling. Sequence compaction alone does not guarantee predictive compression, while preserving contextual freedom allows higher-order musical organization to emerge without explicit structural labels. These results reveal why GPT-style models do not transfer directly across modalities: architectures transfer, but tokenization interfaces do not. Tokenization must discover effective representations while preserving the relational freedom from which contextual structure can emerge.
摘要:GPT風格的模型通過使用有限的可重用離散標記詞彙來表示語言,從而實現強大的性能。這一成功促使符號音樂標記化將重複出現的音樂結構,如和弦、主題和短語,視為類似於語言標記的可重用單元。然而,標記化的優勢並不僅僅來自可重用的組合,而是來自壓縮:有效的壓縮需要坐標系,在這些坐標系中,重複的規律形成穩定且可預測的條件分佈。因此,關鍵問題不在於尋找更大的音樂組合,而在於發現音樂事實變得可預測壓縮的坐標系。我們制定了有效性-無損框架,並將標記化定義為構建一個預測有效且關係無損的坐標系。預測有效性原則定義了事實-標記邊界:解耦和去嵌套構建坐標接口,揭示預測規律。關係無損原則定義了標記-狀態邊界:標記化在上下文依賴關係固定之前停止,將其計算留給模型狀態。受控的符號音樂實驗驗證了這些邊界。有效的坐標構建提高了預測壓縮性,而固定的關係投影則限制了上下文建模。僅僅進行序列壓縮並不能保證預測壓縮,而保留上下文自由則允許更高階的音樂組織在沒有明確結構標籤的情況下出現。這些結果揭示了為什麼GPT風格的模型無法直接跨模態轉移:架構可以轉移,但標記化接口則不能。標記化必須在保留關係自由的同時發現有效的表示,從而使上下文結構能夠出現。
Can Large Language Models Explain Flight Safety Events? A Prior-Guided Semantic LLM-based Approach
2608.18017v1 by Lu Xu, Xu Li, Linjiang Zheng, Fan Li, Riquan Zhang, Jiaxing Shang
Improving flight safety with flight data requires not only accurate detection of risk events, but more importantly, clear interpretation of their underlying causes at the level of pilot control behavior. Existing explainable AI techniques, such as feature importance maps, often require considerable domain knowledge to translate them into operationally meaningful explanations. Large Language Models (LLMs), which excel at language reasoning, bring a promising solution to this issue. However, applying LLMs in this domain presents key challenges such as modal inconsistency, limited classification ability, scarcity of task-specific data for fine-tuning, and lack of domain knowledge. To overcome these challenges, we propose FlightLLM, a prior-guided semantic LLM-based approach for interpretable flight safety analysis. Specifically, we first perform feature engineering to address modal inconsistency, combining statistical descriptors with physically meaningful flight indicators. This representation is further processed by a Semantic Discretization module, which converts abstract numerical patterns into qualitative descriptions that are more compatible with language reasoning. In addition, since LLMs are not inherently strong classifiers, CatBoost is incorporated as a statistical expert, and its prediction results are injected into the prompt as prior guidance. A contrastive few-shot learning strategy is further adopted to compensate for limited data. Finally, we design structured prompts to embed aviation-specific knowledge into the inference process. Using hard landing, a representative risk event with complex causal mechanisms, as an anchor point, we evaluate FlightLLM on a dataset of 704 real-world A320 flight samples. Experimental results show that the proposed approach achieves competitive classification performance while generating direct and reasonable explanations for event causes.
摘要:改善飛行安全需要不僅準確檢測風險事件,更重要的是在飛行員控制行為層面清晰解釋其潛在原因。現有的可解釋AI技術,如特徵重要性圖,通常需要相當的領域知識才能將其轉化為具有操作意義的解釋。大型語言模型(LLMs)在語言推理方面表現出色,為這一問題帶來了有希望的解決方案。然而,在這一領域應用LLMs面臨著關鍵挑戰,如模式不一致、有限的分類能力、缺乏特定任務的數據以進行微調,以及缺乏領域知識。為了克服這些挑戰,我們提出了FlightLLM,一種基於語義的先驗引導LLM方法,用於可解釋的飛行安全分析。具體而言,我們首先進行特徵工程以解決模式不一致,將統計描述符與具有物理意義的飛行指標相結合。這一表示進一步由語義離散化模塊處理,將抽象的數字模式轉換為更符合語言推理的定性描述。此外,由於LLMs本身並不是強大的分類器,因此CatBoost被納入作為統計專家,其預測結果被注入到提示中作為先驗指導。進一步採用了對比少樣本學習策略以彌補數據的有限性。最後,我們設計了結構化提示,將航空特定知識嵌入推理過程中。以硬著陸作為錨點,這是一個具有複雜因果機制的代表性風險事件,我們在704個真實世界A320飛行樣本的數據集上評估FlightLLM。實驗結果表明,所提出的方法在生成事件原因的直接和合理解釋的同時,實現了具有競爭力的分類性能。
The IOL-AI Challenge: An Open Challenge towards Advancing Linguistic Reasoning
2608.18011v1 by Eduardo Sánchez, Rita Berrada, Dan-Mircea Mirea, Sara Rajaee, Alexander Piperski, Ana Meta Dolinar, Boris Iomdin, Andrey Nikulin, Mariya Shmatova, Marzieh Fadaee, Julia Kreutzer
Reasoning in LLMs is overwhelmingly studied in domains that provide a model with rules: mathematics and code. Linguistic puzzles invert this: the solver must first discover the system before reasoning within it. We present the IOL-AI Challenge, an open-science competition run on the unseen problems of the International Linguistics Olympiad (IOL) 2026 Individual Contest, evaluated both automatically and, for the first time, by members of the official IOL Jury under the same rubrics applied to human contestants. The challenge drew 731 submissions from 46 teams under a strict compute budget (one T4, 30 mins). We additionally benchmark 15 unconstrained frontier and open models, with Claude Opus 4.8 earning a jury score equivalent to a gold medal, while both resource-constrained systems we submitted for jury grading scored in the range of the bottom 5% of contestants. Capability was not determined by scale: 14B submissions outperform models twice their size, and gains come from decoding and output-handling rather than model capacity. We also found that automatic metrics rank systems exactly as the jury does, but compress the scale, upscoring weak systems by ~13 points and understating strong ones. Our analysis shows that while frontier models might have prior knowledge about some of the problem languages, it does not significantly help them solve the linguistic reasoning tasks, leaving linguistic reasoning as a strong benchmarking proxy for generalizable reasoning skills.
摘要:推理在大型語言模型(LLMs)中的研究主要集中在提供規則的領域:數學和程式碼。語言謎題則顛倒了這一點:解題者必須首先發現系統,然後才能在其中進行推理。我們提出了IOL-AI挑戰賽,這是一項開放科學競賽,基於2026年國際語言奧林匹亞(IOL)個人賽的未見問題進行評估,這次評估既有自動評分,還首次由官方IOL評審團成員根據與人類參賽者相同的標準進行評分。這次挑戰吸引了46個團隊提交的731份作品,並在嚴格的計算預算下進行(一個T4,30分鐘)。我們還基準測試了15個不受限制的前沿和開放模型,其中Claude Opus 4.8獲得了相當於金牌的評審分數,而我們提交給評審打分的兩個資源受限系統的分數則落在參賽者的底部5%範圍內。能力並不是由規模決定的:14B的提交表現超過了規模是其兩倍的模型,並且性能的提升來自於解碼和輸出處理,而非模型容量。我們還發現,自動指標的排名與評審的排名完全一致,但壓縮了評分範圍,將弱系統的分數提高了約13分,而低估了強系統的分數。我們的分析顯示,儘管前沿模型可能對某些問題語言有先前的知識,但這並未顯著幫助它們解決語言推理任務,這使得語言推理成為通用推理能力的強基準代理。
Policy-Invariant Reward Shaping from LLM Feedback: A Framework for Hybrid RL Agents
2608.18008v1 by Christophe D. Hounwanou, John Emeka Eze, Yaé U. Gaba
Combining large language models with reinforcement learning is increasingly explored, yet the theoretical status of LLM-derived reward signals is often left implicit. We formalize the hybrid LLM-planner and RL-controller architecture as a Goal-Augmented Markov Decision Process and show that when the LLM per-state progress score is used as a bounded potential function, the resulting shaping term preserves the optimal policy set even when the LLM scores are inaccurate. This guarantee is stronger than what general LLM-as-reward approaches provide. We verify the result numerically on a small MDP under four potential configurations, including an adversarial one scaled to twenty times the base reward magnitude.
摘要:結合大型語言模型與強化學習的研究越來越受到關注,然而 LLM 衍生的獎勵信號的理論地位常常被隱含。
我們將混合的 LLM 規劃者和 RL 控制器架構形式化為目標增強馬可夫決策過程,並顯示當 LLM 每狀態的進展分數被用作有界潛在函數時,所產生的塑形項即使在 LLM 分數不準確的情況下也能保持最佳政策集。
這一保證比一般的 LLM 作為獎勵的方法提供的要強。
我們在一個小型 MDP 上對結果進行了數值驗證,考慮了四種潛在配置,包括一個對抗性配置,其規模為基礎獎勵大小的二十倍。
Traceable Trust for action-ready artificial intelligence in bioscience
2608.17997v1 by Huayu Xin, Yizhi Cai, Mukilan Deivarajan Suresh, Gavin Michael Farrell, Iwona Gajda, Charlie Harrison, Conor Houghton, Mato Lagator, Yang Lu, Virginia Portillo, Reyer Zwiggelaar, Sebastian Lobentanzer
Artificial intelligence (AI) is becoming part of the working infrastructure of the biosciences. AI models can predict biomolecular structures, design proteins, rank variants, annotate images, recommend strains and optimise experimental conditions. We argue that the decision to use an AI output to guide laboratory action is a key juncture for trustworthy research and should follow a defined, reviewable process. We propose Traceable Trust as a proportionate assessment-and-design framework for this output-to-action boundary. It asks what evidence supports the output, what capability is being claimed, what agency has been delegated, what threshold authorises action, who can override it and how outcomes inform later decisions. We illustrate the framework through three case studies spanning ecosystem resources, project design and laboratory action. Together, the cases show how trust can be documented where AI outputs begin to shape scientific work.
摘要:人工智慧(AI)正逐漸成為生物科學工作基礎設施的一部分。
AI 模型可以預測生物分子結構、設計蛋白質、排名變異體、註解圖像、推薦菌株並優化實驗條件。
我們認為,使用 AI 輸出來指導實驗室行動的決定是值得信賴的研究的一個關鍵時刻,應遵循一個明確的、可審查的過程。
我們提出可追溯的信任作為這一輸出到行動邊界的比例評估與設計框架。
它詢問什麼證據支持該輸出、聲稱了什麼能力、授予了什麼代理權、什麼門檻授權行動、誰可以覆蓋它以及結果如何影響後續決策。
我們通過三個案例研究來說明該框架,這些案例涵蓋了生態系統資源、項目設計和實驗室行動。
這些案例共同展示了在 AI 輸出開始塑造科學工作時,如何記錄信任。
Judge, Retrieve, or Abstain: Uncertainty-Guarded LLM Judging with Provable Risk Guarantees
2608.17994v1 by Sher Badshah, Ali Emami, Hassan Sajjad
Using LLMs as judges has become standard practice for evaluating model outputs at scale. This is particularly common for subjective, open-ended tasks such as assessing helpfulness or alignment, where no single reference answer exists. However, objective tasks introduce a distinct reliability challenge for reference-free LLM judging. In the absence of a reference answer, the judge evaluates factual correctness either through its parametric knowledge or through tool augmentation. Although the former enables efficient evaluation, the judge may hallucinate or lack sufficient evidence for its verdict. Conversely, tool augmentation can provide additional evidence but introduces extra computational cost and requires an appropriate mechanism to determine when and how that evidence should be used reliably. More importantly, neither approach alone provides formal control over the risk of accepted verdicts or guarantees their reliability at a specified level. We propose a risk-controlled framework that calibrates uncertainty thresholds on a held-out set so that the false discovery rate among accepted verdicts remains below a user-specified level~$α$ with high probability, using finite-sample Clopper--Pearson intervals. When the parametric mode is not sufficiently confident, the instance is routed to a retrieval-augmented mode, where the judge gathers web evidence and re-evaluates the instance under a second calibrated threshold. The finite-sample guarantee carries over to this two-threshold routing without additional assumptions. Across open-domain QA benchmarks and judges of varying scales, the framework maintains the target error rate while achieving substantially higher coverage than single-mode baselines.
摘要:使用大型語言模型作為評審已成為大規模評估模型輸出的標準做法。這在主觀的、開放式的任務中尤其常見,例如評估有用性或一致性,因為這類任務並不存在單一的參考答案。然而,客觀任務對於無參考的 LLM 評審引入了明顯的可靠性挑戰。在缺乏參考答案的情況下,評審通過其參數知識或工具增強來評估事實的正確性。雖然前者能夠實現高效評估,但評審可能會出現幻覺或缺乏足夠的證據來支持其裁決。相反,工具增強可以提供額外的證據,但會引入額外的計算成本,並需要適當的機制來確定何時以及如何可靠地使用這些證據。更重要的是,單獨使用這兩種方法都無法對接受的裁決風險提供正式控制或保證其在特定水平上的可靠性。我們提出了一個風險控制框架,該框架在保留集上校準不確定性閾值,以便接受的裁決中的假陽性率以高概率保持在用戶指定的水平~$α$ 以下,使用有限樣本的 Clopper--Pearson 區間。當參數模式的信心不足時,實例會被路由到檢索增強模式,在該模式下,評審收集網絡證據並在第二個校準閾值下重新評估該實例。有限樣本的保證在這個雙閾值路由中延續,無需額外假設。在開放域問答基準和不同規模的評審中,該框架在保持目標錯誤率的同時,實現了顯著高於單一模式基準的覆蓋率。
Against Political Polarization: A Unified Framework for Tracing Evolving Political Ideologies on Social Media
2608.17987v1 by Yijie Xu, Chao Wang, Hui Xiong
The rapid growth of social media has greatly influenced political discourse, highlighting the need to understand individual political ideologies and their temporal dynamics. This task faces challenges such as data scarcity, abundant non-political content, costly and bias-prone manual annotation, and difficulty in modeling future ideological inclinations. To address these issues, we propose TSN4PI, a unified framework for tracking the evolution of political ideologies on social media. It includes two core modules. The PIDN uses large language models with style transfer and unsupervised domain adaptation to enable robust ideology detection and filter irrelevant content from noisy, cross-domain data. The PIPN employs temporal graph neural networks to predict future ideological shifts, enabling comprehensive analysis of ideology presence, intensity, and evolution. We release two large-scale datasets for noncommercial research use to facilitate further work. Extensive case studies on multiple platforms (X and Truth Social) validate the effectiveness of TSN4PI and provide empirical insights into political polarization and the evolution of online ideologies. Our findings offer a nuanced perspective, advancing both methodological development and empirical understanding in this field.
摘要:社交媒體的快速增長對政治話語產生了重大影響,突顯了理解個人政治意識形態及其時間動態的必要性。這項任務面臨著數據稀缺、非政治內容豐富、昂貴且易受偏見影響的手動標註以及未來意識形態傾向建模困難等挑戰。為了解決這些問題,我們提出了TSN4PI,一個用於追蹤社交媒體上政治意識形態演變的統一框架。它包括兩個核心模塊。PIDN使用大型語言模型結合風格轉換和無監督領域適應,以實現穩健的意識形態檢測並過濾來自嘈雜的跨領域數據中的無關內容。PIPN則利用時間圖神經網絡來預測未來的意識形態變化,使得對意識形態的存在、強度和演變進行全面分析成為可能。我們釋放了兩個大型數據集供非商業研究使用,以促進進一步的研究工作。在多個平台(X和Truth Social)上進行的廣泛案例研究驗證了TSN4PI的有效性,並提供了對政治極化和在線意識形態演變的實證見解。我們的發現提供了一個細緻的視角,推進了該領域的方法論發展和實證理解。
Dual Co-Train: Cross-Dataset Ultrasound Tongue Segmentation Under Extreme Data Scarcity
2608.17983v1 by Alisher Myrgyyassov, Zhen Song, Bruce Xiao Wang, Yu Sun, Min Ney Wong, Yihao Zhou, Yongping Zheng
Ultrasound tongue contour segmentation remains challenging under cross-dataset domain shift, where limited annotations, probe variability, and acquisition noise often degrade model generalization. We present a source-free domain adaptation framework for robust ultrasound tongue segmentation built on a lightweight UltraUNet backbone. Starting from a checkpoint pretrained on only five labeled source images, simulating an underfitted constrained source model, the proposed method adapts to a fully-unlabeled target domain by iteratively refining pseudo-labels, filtering unreliable masks with a contour-based quality-control module, and generating target-style synthetic image-mask pairs through a segmentation-guided conditional GAN. The student model is then trained on a mixture of clean pseudo-labeled target images, noisy pseudo-labels with consistency regularization, and synthetic samples, enabling closed-loop adaptation without access to source data. We evaluate the method on 12 source-target transfer pairs across eight ultrasound tongue imaging datasets, and conduct source-size scaling experiments and ablation studies. Across all comparisons, the proposed framework improves segmentation overlap and contour accuracy over the baselines, including supervised ones. These results suggest that task-specific pseudo-label refinement and synthetic target-style augmentation can substantially improve source-free adaptation for ultrasound tongue imaging.
摘要:超聲波舌頭輪廓分割在跨數據集域轉移下仍然具有挑戰性,因為有限的註釋、探頭變異性和獲取噪聲常常會降低模型的泛化能力。
我們提出了一個無源域適應框架,用於穩健的超聲波舌頭分割,基於輕量級的UltraUNet骨幹。
從僅在五張標記源圖像上預訓練的檢查點開始,模擬一個欠擬合的受限源模型,所提出的方法通過迭代地細化偽標籤,使用基於輪廓的質量控制模塊過濾不可靠的掩膜,並通過分割引導的條件GAN生成目標風格的合成圖像-掩膜對,適應於完全無標記的目標域。
然後,學生模型在一組乾淨的偽標記目標圖像、帶有一致性正則化的噪聲偽標籤和合成樣本的混合上進行訓練,使得在無需訪問源數據的情況下實現閉環適應。
我們在八個超聲波舌頭成像數據集上的12對源-目標轉移對上評估了該方法,並進行了源大小縮放實驗和消融研究。
在所有比較中,所提出的框架在分割重疊和輪廓準確性方面優於基準,包括監督學習的基準。
這些結果表明,特定任務的偽標籤細化和合成目標風格的增強可以顯著改善超聲波舌頭成像的無源適應能力。
When Writing Style Drifts: Benchmarking Authorship Verification under Distribution Shifts in Genre, Time and the AI-Era
2608.17979v1 by Lotta Kiefer, Brisca Balthes, Christoph Leiter, Yamen Ajjour, Elena Schmidt, Steffen Eger
Authorship verification (AV) assumes that an author's writing style remains sufficiently stable to distinguish it from that of other writers. In practice, however, this assumption is challenged by distribution shifts caused by changes in genre, time, and AI-assisted writing. Existing AV benchmarks typically study these factors in isolation and focus predominantly on English, limiting our understanding of model robustness under realistic conditions. We introduce AVShift, the first German benchmark for systematically evaluating AV under multiple distribution shifts. AVShift comprises over 150,000 text pairs spanning three genres and 21 years, enabling controlled evaluation of cross-genre, temporal, and AI-era shifts within a unified framework. We benchmark representative feature-based, embedding-based, and LLM-based approaches. Our experiments show that fine-tuned LLMs generalize best across genres and benefit substantially from stylistically diverse training data. We further demonstrate that temporal drift is one of the strongest factors affecting AV, with performance degrading significantly as the time gap between documents increases. In contrast, we find no evidence of a measurable AI-era distribution shift within AVShift. Finally, our feature analysis reveals stylistic features that remain stable across genres, while their relative importance varies depending on the specific genre transition. We release AVShift and our code for future research.
摘要:作者驗證(AV)假設作者的寫作風格保持足夠穩定,以便與其他作家的風格區分開來。
然而,在實際操作中,這一假設受到由於類型、時間和AI輔助寫作變化而引起的分佈變化的挑戰。
現有的AV基準通常孤立地研究這些因素,並主要集中在英語上,限制了我們在現實條件下對模型穩健性的理解。
我們介紹了AVShift,這是第一個德語基準,用於系統地評估多種分佈變化下的AV。
AVShift包含超過150,000對文本,涵蓋三個類型和21年,能夠在統一框架內進行跨類型、時間和AI時代變化的受控評估。
我們基準測試了代表性的基於特徵、基於嵌入和基於LLM的方法。
我們的實驗表明,微調的LLM在各類型之間的泛化效果最佳,並且從風格多樣的訓練數據中獲益良多。
我們進一步證明,時間漂移是影響AV的最強因素之一,隨著文檔之間時間間隔的增加,性能顯著下降。
相比之下,我們在AVShift中沒有發現可測量的AI時代分佈變化的證據。
最後,我們的特徵分析揭示了在各類型中保持穩定的風格特徵,而它們的相對重要性則根據具體的類型轉換而有所不同。
我們發布了AVShift和我們的代碼以供未來研究使用。
Too Sure to Be Safe: Model Calibration for Reliable Log Anomaly Detection
2608.17965v1 by Bin Li, Dongdong Wang, Siyang Lu
Online log anomaly detection is critical for maintaining the reliability of large-scale computing systems. Although recent language model-based log anomaly detectors achieve strong detection performance, their confidence estimates remain poorly calibrated. We show that these detectors frequently assign excessive confidence to incorrect predictions, particularly for anomalous logs under severe class imbalance. Moreover, confidence on erroneous predictions remains persistently high even when conventional calibration metrics indicate good calibration, creating a critical reliability gap for operational monitoring systems. To address this issue, we propose Log Reconstruction and Distance (LoRD), a lightweight post-hoc calibration framework for reliable log anomaly detection. LoRD learns prediction-route-specific reliability models from latent representations of correctly classified validation samples and estimates prediction reliability through route-wise reconstruction distances. Based on the estimated reliability, LoRD selectively recalibrates high-risk predictions to suppress overconfident errors while preserving reliable predictions. Extensive experiments on four large-scale log benchmark datasets and multiple language model-based detectors demonstrate that LoRD consistently improves confidence reliability and substantially reduces overconfident anomaly-related errors without sacrificing anomaly detection performance.
摘要:線上日誌異常檢測對於維護大規模計算系統的可靠性至關重要。儘管最近基於語言模型的日誌異常檢測器在檢測性能上表現出色,但它們的信心估計仍然校準不佳。我們顯示這些檢測器經常對錯誤預測賦予過高的信心,特別是在嚴重類別不平衡的異常日誌中。此外,即使在傳統的校準指標顯示良好校準的情況下,對錯誤預測的信心仍然持續偏高,這為運營監控系統創造了關鍵的可靠性差距。為了解決這個問題,我們提出了日誌重建與距離(LoRD),這是一個輕量級的事後校準框架,用於可靠的日誌異常檢測。LoRD從正確分類的驗證樣本的潛在表示中學習特定於預測路徑的可靠性模型,並通過路徑重建距離來估計預測的可靠性。根據估計的可靠性,LoRD選擇性地重新校準高風險預測,以抑制過度自信的錯誤,同時保留可靠的預測。在四個大規模日誌基準數據集和多個基於語言模型的檢測器上進行的廣泛實驗表明,LoRD持續改善信心可靠性,並顯著減少過度自信的異常相關錯誤,而不犧牲異常檢測性能。
Towards Zero-Shot Task Transfer with Neurosymbolic World Models
2608.17959v1 by Isidoro Tamassia, Lennert De Smet, Giuseppe Marra
State-of-the-art model-based reinforcement learning methods learn neural world models that allow policy improvement by planning in a latent space, without assumptions on the structure of the underlying environment. While expressive, these models are generally task-dependent: they learn uninterpretable latent representations that are tied to the training task and thus hard to generalize to new tasks. In this work, we present a novel world model formulation where the reward prediction only depends on a subset of structured, symbolic components of the whole latent state. Decoupling observation reconstruction and reward prediction allows us to learn world models that can adapt zero-shot, i.e. without further environment interactions, to new reward functions defined over the same symbolic state space. We discuss the main advantages and challenges of learning these neurosymbolic world models and demonstrate the strong generalisation properties of our approach over purely neural methods.
摘要:最先進的基於模型的強化學習方法學習神經世界模型,這些模型允許在潛在空間中進行規劃以改進策略,而不需要對基礎環境的結構做出假設。雖然這些模型具有表現力,但通常依賴於特定任務:它們學習的潛在表示難以解釋,並且與訓練任務緊密相關,因此難以泛化到新任務。在這項工作中,我們提出了一種新穎的世界模型公式,其中獎勵預測僅依賴於整個潛在狀態的結構化符號組件的子集。將觀察重建與獎勵預測解耦,使我們能夠學習能夠零次適應的世界模型,即在不進一步與環境互動的情況下,對定義在相同符號狀態空間上的新獎勵函數進行適應。我們討論了學習這些神經符號世界模型的主要優勢和挑戰,並展示了我們的方法相較於純神經方法的強泛化特性。
An Omitted Mode Is a Rare Rule: The Sampling-Verification Danger Law in Continuous Code World Models
2608.17956v1 by Javier Aguilar Martín
In the Code World Model paradigm an LLM synthesizes an executable world model that a classical planner searches, and the model is accepted when it reproduces sampled transitions. We ask what that acceptance certifies in continuous control. We define the pipeline's danger as an expected risk and isolate its exact factor: the probability that N i.i.d. gate rollouts all miss a critical event of probability r is exactly (1-r)^N; an independent acceptance sample adds its budget to the exponent. On three hybrid instruments the accepted mode-blind model is exploited: the planner is pinned at the mode boundary at a regret of nearly the whole attainable return. We prove a localization budget, valid at boundary points: models with Lipschitz constant at most L differing by eta at a point disagree above tolerance eps on a region of volume at least kappa((eta-eps)/L)^(d+m); the discontinuous reset modes studied pay no such budget. With real LLM synthesis, GPT-5.x repairs an omitted 1D clamp in 105 of 111 mode-containing draws -- every attempt exact on 50 of 56 instrument-stream blocks (95% CI [0.781, 0.960]). On 2D regions no artifact recovers the rule (0/156); eight targeted interventions leave the failure in place, and positive controls locate it: a located rule is not induced, while given form and location the constants follow exactly. A version-space certificate proves identification is class-relative: at the widest dose the declared fit succeeds in 20/20 blocks and every sample-consistent circle is within tolerance in 18/20. We prove a class of entry rules exactly consistent with every sample yet harmless at play, so identifiability is a measurable property of the instrument. Re-scoring all 1034 artifacts on independent samples confirms acceptance certifies sample consistency and no more: where the gate is provably informative it covers about two percent of the exploited planner's queries.
摘要:在代碼世界模型範式中,一個大型語言模型(LLM)合成了一個可執行的世界模型,供經典規劃器搜索,當模型重現抽樣轉換時便被接受。我們詢問這種接受在連續控制中證明了什麼。我們將管道的危險定義為預期風險,並孤立其確切因素:N個獨立同分佈的閘門展開全部錯過概率為r的關鍵事件的概率恰好是(1-r)^N;一個獨立的接受樣本將其預算添加到指數中。在三個混合工具上,接受的無模式模型被利用:規劃器在模式邊界被固定,後悔幾乎達到整個可獲得回報。我們證明了一個有效於邊界點的定位預算:在某一點上,Lipschitz常數最多為L的模型若在eta上有所不同,則在容忍度eps以上的區域內存在體積至少為kappa((eta-eps)/L)^(d+m)的差異;所研究的不連續重置模式不需要這樣的預算。通過實際的LLM合成,GPT-5.x在111個包含模式的抽樣中修復了105個遺漏的1D夾具——每次嘗試在56個工具流塊中的50個上都是精確的(95%置信區間 [0.781, 0.960])。在2D區域中,沒有任何工件恢復該規則(0/156);八個目標干預使失敗保持不變,而正控制則定位了它:一個定位的規則並未被誘導,而在給定形式和位置的情況下,常數恰好遵循。版本空間證明了識別是類別相對的:在最寬的劑量下,聲明的擬合在20/20塊中成功,每個樣本一致的圓圈在18/20中都在容忍範圍內。我們證明了一類與每個樣本完全一致但在遊戲中無害的進入規則,因此可識別性是工具的一個可測量屬性。對1034個工件在獨立樣本上的重新評分確認接受證明樣本一致性,並且僅此而已:在閘門被證明為信息豐富的地方,它覆蓋了約兩個百分比的被利用規劃者的查詢。
Do Large Language Models Play Six Degrees of Separation? Measuring Topological Compression in Long-Context Manifolds
2608.17950v1 by Md. Faiyaz Abdullah Sayeedi
Large Language Models (LLMs) demonstrate remarkable multi-hop reasoning capabilities over long contexts, yet the internal mechanisms enabling these distant cognitive leaps remain poorly understood. Traditional attention-based interpretability often fails to capture true semantic proximity due to routing artifacts like attention sinks. In this paper, we bypass attention weights to directly analyze the dynamic geometry of the hidden state manifold, proving that deep LLM latent spaces natively organize into Small-World networks. By sparsifying the continuous similarity matrices of long-context representations into unweighted graphs, we trace the connectivity between highly disjoint semantic anchors across two distinct architectures. Our findings reveal a sharp topological phase transition: while early syntactic layers remain entirely fractured, deep reasoning layers abruptly compress massive conceptual distances into highly navigable pathways strictly bounded by the "Six Degrees of Separation" limit (=< 6 semantic hops). Furthermore, we demonstrate the practical efficacy of this framework by applying it to zero-shot hallucination detection within Retrieval-Augmented Generation (RAG) using the RAGognize dataset. We show that factually grounded generations maintain structural integrity with their source context (approximately 3 hops), whereas hallucinations induce severe topological collapse. Ultimately, this work mathematically formalizes how transformers execute abstract reasoning and provides a novel, strictly geometric signature for evaluating factual reliability.
摘要:大型語言模型(LLMs)在長上下文中展現出卓越的多跳推理能力,但促成這些遙遠認知飛躍的內部機制仍然不甚了解。傳統的基於注意力的可解釋性常常無法捕捉到真實的語義接近性,這是由於路由伪影如注意力匯聚所致。在本文中,我們繞過注意力權重,直接分析隱藏狀態流形的動態幾何,證明深層LLM潛在空間本質上組織成小世界網絡。通過將長上下文表示的連續相似性矩陣稀疏化為無權重圖,我們追蹤兩個不同架構之間高度不相交的語義錨點之間的連接性。我們的研究結果揭示了一個明顯的拓撲相變:儘管早期的句法層完全破碎,深層推理層卻突然將巨大的概念距離壓縮成高度可導航的路徑,這些路徑嚴格受限於「六度分隔」的限制(=< 6語義跳躍)。此外,我們通過將此框架應用於檢索增強生成(RAG)中的零樣本幻覺檢測,展示了其實際效能,使用了RAGognize數據集。我們顯示,事實基礎的生成與其來源上下文保持結構完整(約3跳),而幻覺則引發嚴重的拓撲崩潰。最終,這項工作數學化了Transformer如何執行抽象推理,並提供了一種新穎的、嚴格的幾何特徵,用於評估事實可靠性。
SIGMA: SHAP-Guided Implicit-Trajectory Generation for Metadata-Free LLM-Based AutoFE
2608.17948v1 by Xuan Zheng, Kento Uchida, Shinichi Shirakawa
Recent research has leveraged Large Language Models (LLMs) to enhance Automated Feature Engineering (AutoFE) through semantic descriptions and trajectory-based prompting. However, there exist two challenges that limit their applicability and scalability in long-horizon optimization: (1) semantic metadata is unavailable in many practical settings, and (2) trajectory accumulation increases the risk of exceeding the context window, while without it, the generation process can become unstable, leading to becoming stuck in the local optima and a high duplicate rate of generated features. To this end, we propose a SHAP-enhanced Implicit-trajectory Generation for Metadata-free AutoFE (SIGMA), a scalable constant-context optimization framework. SIGMA leverages SHAP values to provide task-aware signals for guiding group feature generation instead of semantic information. In addition, we adopt an EXposed-feature Implicit Trajectory (EXIT) approach, where the exposed features in the prompt implicitly represent the trajectory. Empirical results demonstrate that SIGMA achieves performance comparable to the state-of-the-art (SOTA) LLM baselines with a nearly constant prompt length. Notably, EXIT significantly reduces the duplicate ratio of generated features from 37.2% to 6.8%. At the same time, SIGMA matches traditional SOTA performance with only 5.4 features on average, demonstrating substantial efficiency gains in feature utilization.
摘要:最近的研究利用大型語言模型(LLMs)通過語義描述和基於軌跡的提示來增強自動特徵工程(AutoFE)。然而,存在兩個挑戰限制了它們在長期優化中的適用性和可擴展性:(1)在許多實際環境中,語義元數據不可用,以及(2)軌跡累積增加了超出上下文窗口的風險,而如果沒有它,生成過程可能變得不穩定,導致陷入局部最優解和生成特徵的高重複率。為此,我們提出了一種增強SHAP的隱式軌跡生成方法,用於無元數據的自動特徵工程(SIGMA),這是一個可擴展的恆定上下文優化框架。SIGMA利用SHAP值提供任務感知信號,以指導群體特徵生成,而不是依賴語義信息。此外,我們採用了一種EXposed-feature隱式軌跡(EXIT)方法,其中提示中的暴露特徵隱式地代表了軌跡。實證結果表明,SIGMA在幾乎恆定的提示長度下達到了與最先進(SOTA)LLM基準相當的性能。值得注意的是,EXIT顯著將生成特徵的重複率從37.2%降低到6.8%。同時,SIGMA在平均僅使用5.4個特徵的情況下達到了傳統SOTA性能,顯示出特徵利用的顯著效率提升。
Procedural Content Metageneration via Program Search and Continual Abstraction Discovery
2608.17947v1 by Matthew Siper, Ahmed Khalifa, Julian Togelius
Large language models can generate executable programs, which makes it possible to search directly over procedural content generators rather than individual levels. We study this approach in Sokoban, Zelda, Dangerous Dave, and Lode Runner. Each run evolves complete Python generators through language-model mutation and crossover. We introduce Continual Abstraction Discovery, or CAD, which extracts reusable primitives from high-fitness programs into a run-specific helper module. A 2x2 experiment crosses CAD with access to a fixed hand-written domain API. The completed data set contains 160 complete runs, with at least ten 50-generation runs in every cell. CAD raises mean final best fitness in all eight domain and API comparisons. Across all CAD runs, learned libraries are adopted by most later programs and repeatedly rediscover validation, reachability, and structural utilities. These results support that discovering reusable primitives improves evolutionary program search for content generators.
摘要:大型語言模型可以生成可執行的程式,這使得可以直接在程序內容生成器上進行搜索,而不是單獨的關卡。
我們在推箱子、薩爾達傳說、危險的戴夫和洛德奔跑者中研究這種方法。
每次運行通過語言模型的突變和交叉演化出完整的 Python 生成器。
我們引入了持續抽象發現(Continual Abstraction Discovery,簡稱 CAD),它從高適應度的程式中提取可重用的原語,形成特定於運行的輔助模組。
一個 2x2 實驗將 CAD 與訪問固定的手寫領域 API 結合起來。
完成的數據集包含 160 次完整運行,每個單元至少有十次 50 代的運行。
CAD 在所有八個領域和 API 比較中提高了平均最終最佳適應度。
在所有 CAD 運行中,學習到的庫被大多數後續程式採用,並重複發現驗證、可達性和結構性實用工具。
這些結果支持發現可重用原語改善內容生成器的進化程式搜索。
Efficient RLVR Scheduling via Graph-Structured Online Difficulty Estimation
2608.17941v1 by Zhizhao Liu, Zhiliang Tian, Xi Wang, Zhihua Wen, Yihang Xiong, Zhiquan Lai, Dongsheng Li
Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models but relies on costly rollout exploration. Assigning the same exploration budget to samples with different difficulty levels is inefficient: easy samples may receive redundant rollouts, whereas difficult but learnable samples may receive too little exploration. Existing adaptive schedulers address this mismatch through curriculum-based sample selection or non-uniform rollout allocation based on estimated sample difficulty. However, obtaining reliable online difficulty estimates remains challenging: dedicated probing adds substantial generation overhead, whereas history-based estimators face a cold start with no initial observations and stale feedback, and typically ignore relations among samples. To address these limitations, we propose a plug-and-play graph-based online difficulty estimator that shares rollout feedback across related samples and continuously updates their difficulty estimates, mitigating cold start and staleness without dedicated probing. Specifically, we first construct a difficulty-aware sample graph based on semantic and reasoning similarities. Based on this graph, we introduce latent difficulty states and use a Potts prior to encourage neighboring samples to share the same state. We then employ a state-level Beta-Binomial model to aggregate the rollout outcomes associated with each state. Finally, we use an online mean-field variational algorithm to continuously update the latent-state assignments and state-level difficulty as new feedback arrives. Our framework can be integrated into sample-selection and rollout-allocation schedulers, enabling difficulty-adaptive exploration without dedicated probing. Experiments across multiple base models, RL schedulers, and benchmarks demonstrate that our framework achieves better performance.
摘要:強化學習與可驗證獎勵(RLVR)提升了大型語言模型的推理能力,但依賴於成本高昂的展開探索。將相同的探索預算分配給不同難度級別的樣本是低效的:簡單樣本可能會收到冗餘的展開,而難度較高但可學習的樣本可能會收到過少的探索。現有的自適應調度器通過基於課程的樣本選擇或根據預估樣本難度的非均勻展開分配來解決這一不匹配。然而,獲得可靠的在線難度估計仍然具有挑戰性:專門的探測增加了可觀的生成開銷,而基於歷史的估計器面臨著沒有初始觀察和過時反饋的冷啟動問題,並且通常忽略樣本之間的關係。為了解決這些限制,我們提出了一種即插即用的基於圖的在線難度估計器,該估計器在相關樣本之間共享展開反饋,並持續更新它們的難度估計,減輕冷啟動和過時問題,無需專門的探測。具體而言,我們首先根據語義和推理相似性構建一個難度感知樣本圖。基於這個圖,我們引入潛在的難度狀態,並使用Potts先驗來鼓勵相鄰樣本共享相同的狀態。然後,我們使用狀態級的Beta-Binomial模型來聚合與每個狀態相關的展開結果。最後,我們使用在線均場變分算法來持續更新潛在狀態分配和狀態級難度,隨著新反饋的到來。我們的框架可以集成到樣本選擇和展開分配調度器中,實現難度自適應探索,而無需專門的探測。在多個基礎模型、RL調度器和基準測試中的實驗表明,我們的框架實現了更好的性能。
Grading Needs a Rubric, Not Intelligence
2608.17938v1 by Jhen-Ke Lin
Small language models can grade open-ended examination answers as reliably as substantially more expensive models when they grade against an explicit rubric. We test this claim as the design principle behind any-to-bench: a frontier model reads source documents once, at ingestion, to extract each question and its rubric; lower-cost models then perform all repeated grading work. We evaluate six cost-efficient model configurations from two model families at three reasoning-effort levels. Each configuration answers 24 open-ended examination questions, and each also grades every answer sheet three times, yielding 3,456 per-question grades. Scores depend overwhelmingly on the answer being graded: answer identity explains 95.6% of score variance, whereas judge identity explains only 0.2%. Raising a writer's reasoning effort moves earned scores by as much as 0.143 of full marks, while raising a judge's reasoning effort moves assigned scores by at most 0.006. Six frontier-tier judges, added as a check, reproduce these scores and are no more reliable as a panel. Two ablations then decompose the rubric on the same questions and answers. Removing its criteria and levels while keeping the official answer changes nothing measurable. Removing the official answer as well collapses reliability (ICC 0.888 to 0.628), inflates scores, and makes judge reasoning effort matter again. The rubric is what decouples grading from judge intelligence, and within the rubric the official answer does nearly all the work. We find no evidence of length preference or same-family preference under rubric-anchored grading.
摘要:小型語言模型在根據明確的評分標準進行評分時,可以與成本高得多的模型一樣可靠地評分開放式考試答案。我們測試這一主張,作為 any-to-bench 的設計原則:前沿模型在攝取時讀取源文件一次,以提取每個問題及其評分標準;然後,成本較低的模型執行所有重複的評分工作。我們在三個推理努力水平上評估來自兩個模型系列的六種成本效益模型配置。每個配置回答 24 道開放式考試問題,並且每個配置還對每份答案進行三次評分,產生每個問題 3,456 次評分。分數在很大程度上取決於被評分的答案:答案身份解釋了 95.6% 的分數變異,而評判身份僅解釋了 0.2%。提高寫作者的推理努力可以使得獲得的分數提高最多 0.143 的滿分,而提高評判的推理努力則最多使分數提高 0.006。六位前沿級評判作為檢查,重現這些分數,且作為小組的可靠性並沒有提高。接下來的兩個消融實驗則在相同的問題和答案上分解評分標準。去除其標準和級別,同時保留官方答案,並不會改變可測量的結果。去除官方答案也會使可靠性崩潰(ICC 從 0.888 降至 0.628),使分數膨脹,並使評判的推理努力再次變得重要。評分標準是將評分與評判智力解耦的關鍵,而在評分標準內,官方答案幾乎承擔了所有的工作。我們沒有發現基於評分標準的評分中存在長度偏好或同家族偏好的證據。
EvoTS-Agent: A Self-Evolving LLM Agent for Financial Time Series Change Point Detection
2608.17933v1 by Lei Jiang, Ye Wei, Xinyu Xi, Jordan Langham-Lopez, Yifan Bao, Raad Khraishi, Yihao Ang, Anthony K. H. Tung, Lukasz Szpruch, Hao Ni
Financial time series exhibit non-stationary and heterogeneous statistical properties, making change-point detection challenging because no single unsupervised algorithm performs consistently across assets and market regimes. Conventional workflows consequently depend heavily on expert-driven model selection, feature design, and hyperparameter tuning, limiting their scalability and adaptability. We propose EvoTS-Agent, a validation-guided self-evolving LLM agent for autonomous financial time-series change-point detection. EvoTS-Agent first performs curated exploratory data analysis to characterize dataset properties and initialize candidate detection models. It then evolves executable experiment trajectories through three complementary operators: \textit{Revision} exploits the current best solution, \textit{Alternative Strategy} explores fundamentally different modeling directions when progress stagnates, and \textit{Recombination} synthesizes complementary evidence from high-performing trajectories. Validation feedback guides trajectory evolution throughout the search, enabling the agent to adapt its detection pipeline to the statistical characteristics of each dataset while preserving reliable optimization. Experiments across four benchmark datasets demonstrate that EvoTS-Agent consistently outperforms existing LLM-based agents while maintaining a 100\% execution success rate across all evaluated backbone LLMs.
摘要:金融時間序列展現出非平穩和異質的統計特性,使得變更點檢測變得具有挑戰性,因為沒有單一的無監督算法能在不同資產和市場狀態下持續表現良好。因此,傳統工作流程在很大程度上依賴專家驅動的模型選擇、特徵設計和超參數調整,這限制了它們的可擴展性和適應性。我們提出了EvoTS-Agent,一種基於驗證指導的自我演化LLM代理,用於自主金融時間序列變更點檢測。EvoTS-Agent首先執行精心策劃的探索性數據分析,以特徵化數據集特性並初始化候選檢測模型。然後,它通過三個互補的運算子來演化可執行的實驗軌跡:\textit{Revision}利用當前最佳解,\textit{Alternative Strategy}在進展停滯時探索根本不同的建模方向,\textit{Recombination}從高效能的軌跡中綜合互補證據。驗證反饋在整個搜索過程中指導軌跡演化,使得代理能夠根據每個數據集的統計特徵調整其檢測流程,同時保持可靠的優化。在四個基準數據集上的實驗表明,EvoTS-Agent始終超越現有的基於LLM的代理,同時在所有評估的主幹LLM中保持100%的執行成功率。
Collective Counterfactual Planning: Coordination, Consent, and Verification under Representational Constraints
2608.17932v1 by Chainarong Amornbunchornvej
Groups routinely complete projects that no single member can plan, execute, or verify alone. We propose a formal model of this phenomenon, Collective Counterfactual Planning (CCP), in which the binding limitation on each agent is neither capability, knowledge, nor observability, but representational geometry: each agent perceives the state, conceives moves, consents to actions, and certifies goal requirements only through a projection onto an agent-specific subspace of a common task space. Four gates jointly determine whether a team can reach a conjunctive goal and legitimately recognize that it has done so: the exogenous implementation coalitions required to perform each action, together with three representational gates -- conception, consent, and task-relative verification qualification. We define the Collective Counterfactual Solvability (CCS) problem, separating geometric feasibility, executable attainment, and validated completion. The results expose a positive-negative duality. Iterated cross-agent relay can unlock a solution that no one-shot pooling of individual plans contains, but any goal requirement depending essentially on the subspace dark to the entire team is unverifiable and therefore not validly completable, even when the trajectory accidentally attains it. Memoryless and audited consent further constrain different objects -- action directions versus cumulative trajectory states -- and neither dominates the other. A four-step exhaustive horizon-bounded solvability scheme is sound and complete under exact representation of the relay closure; restricted implementations remain sound on returned plans but need not be complete. The model gives one geometry for sequential mutual enabling, competent execution of steps whose purpose is invisible to the executor, forced sub-teaming at expertise boundaries, and completion that cannot be validly declared.
摘要:團體經常完成單一成員無法獨自計劃、執行或驗證的項目。我們提出這一現象的正式模型,稱為集體反事實規劃(CCP),在這個模型中,每個代理的約束限制既不是能力、知識,也不是可觀察性,而是表徵幾何:每個代理僅通過投影到共同任務空間的代理特定子空間來感知狀態、構思行動、同意行為和認證目標要求。四個閘門共同決定一個團隊是否能夠達成聯合目標並合法地認識到它已經達成:執行每個行動所需的外生實施聯盟,以及三個表徵閘門——構思、同意和任務相對驗證資格。我們定義了集體反事實可解性(CCS)問題,將幾何可行性、可執行達成和驗證完成分開。結果揭示了一種正負對偶性。迭代的跨代理中繼可以解鎖一個單次個人計劃無法包含的解決方案,但任何本質上依賴於對整個團隊來說是黑暗的子空間的目標要求都是不可驗證的,因此無法有效完成,即使軌跡意外達成了它。無記憶和經審核的同意進一步限制了不同對象——行動方向與累積軌跡狀態——而且兩者不相互主導。一個四步的全面邊界可解性方案在中繼閉包的精確表徵下是健全且完整的;受限的實施在返回的計劃上仍然是健全的,但不必是完整的。該模型為順序相互啟用、執行目的對執行者不可見的步驟的能力執行、在專業邊界強制子團隊以及無法有效宣告的完成提供了一種幾何。
SpeechSense: A Paralinguistic-Focused Dataset for Fine-Grained Speech Sentiment Analysis
2608.17931v1 by Shicheng Ma, Wenqian Cui, Irwin King
Recent advances in AI have revolutionized speech processing, yet effective speech understanding requires discerning not just what is said, but how it is said. Speech Sentiment Analysis plays a critical role in decoding these paralinguistic cues for diverse real-world applications such as recruitment and customer service. However, existing Speech Sentiment Analysis research faces two primary limitations. First, dominant approaches rely on text-centric pipelines that cascade Automatic Speech Recognition with text analysis. This process inevitably discards essential acoustic features like prosody and tone, failing to capture attitudinal meanings in acoustically ambiguous utterances. Second, current benchmarks suffer from a mismatch in label granularity, prioritizing basic emotions (e.g., happy, sad) over the nuanced interpersonal stances (e.g., confident, impatient) necessary for social sensitivity. To address these limitations, we propose a novel dataset, SpeechSense, for fine-grained speech sentiment analysis. Specifically, we define a specialized 8-class taxonomy of interpersonal stances detectable primarily through prosodic cues beyond lexical content alone. We then construct a curated dataset based on this taxonomy, built from high-fidelity speech synthesis and rigorous human validation. Comprehensive experiments across multi-modal LLMs, text-only LLMs, and speech encoders demonstrate that models with acoustic access consistently outperform text-only baselines. These results empirically validate the primacy of acoustic cues in detecting subtle speaker attitudes, highlighting the necessity of SpeechSense. Dataset and supplementary materials are available at https://github.com/Sher13cked/SpeechSense.
摘要:最近在人工智慧方面的進展已經徹底改變了語音處理,但有效的語音理解不僅需要辨識所說的內容,還需要理解其表達方式。語音情感分析在解碼這些副語言線索方面扮演著關鍵角色,適用於招聘和客戶服務等多樣的現實應用。然而,現有的語音情感分析研究面臨兩個主要限制。首先,主流方法依賴於以文本為中心的流程,將自動語音識別與文本分析串聯起來。這一過程不可避免地忽略了諸如韻律和語調等重要的聲學特徵,未能捕捉聲學模糊表達中的態度意義。其次,當前的基準測試在標籤粒度上存在不匹配,優先考慮基本情感(例如,快樂、悲傷),而忽視了社交敏感性所需的細微人際立場(例如,自信、不耐煩)。為了解決這些限制,我們提出了一個新穎的數據集,SpeechSense,用於細粒度的語音情感分析。具體來說,我們定義了一個專門的8類人際立場分類法,主要通過韻律線索而非單純的詞彙內容來檢測。我們然後根據這一分類法構建了一個精心策劃的數據集,該數據集基於高保真語音合成和嚴謹的人類驗證。跨多模態大型語言模型、僅文本的大型語言模型和語音編碼器的全面實驗表明,具有聲學訪問的模型在性能上始終優於僅文本的基準。這些結果實證了聲學線索在檢測微妙說話者態度中的重要性,突顯了SpeechSense的必要性。數據集和補充材料可在 https://github.com/Sher13cked/SpeechSense 獲得。
Comparative Study of Out-of-the-Box Technology for Automatic Target Detection and Recognition
2608.17917v1 by Alma M. Liezenga, Lotte Nijskens, Henrik R. Baumann, Stefan Becker, Simon Bensberg, Niccolò Camarlinghi, Håvard R. Eiring, Alexander W. Johnsgaard, Tanel Liiv, Giuseppe Martino, Matteo Marturini, Matthias Rapp, Jan Erik van Woerden, Alexander Wolpert, Hugo J. Kuijf
Automatic Target Detection and Recognition (ATD/R) is critical for military decision support and (semi-)autonomous operations. Recent advances in object detection and artificial intelligence (AI) significantly boosted the potential performance of ATD/R. However, the scarcity of publicly available military datasets limits the application of these systems. As a solution, this paper explores the use of publicly available models and civilian datasets to achieve reasonable performance in military contexts. We benchmark several state-of-the-art models, including six iterations of the YOLO series and two variations on the DETR framework, on a newly acquired military relevant dataset. This dataset features military vehicles and challenging circumstances, including various degrees of occlusions and small targets. The out-of-the-box version of each model is validated alongside a version finetuned on the VisDrone dataset. This dataset features small objects, an Air-to-Ground (A2G) perspective and relevant classes, potentially generalizing to our military ATD/R task. We compare the performance of the models using mAP@0.5 and mAP@0.5:0.95, across A2G and Ground-to-Ground (G2G) perspective, target size and model size, giving insight into the real-time capabilities of models. Our main findings are: (1) bigger models outperform smaller models, (2) DETR-based models show promising results compared to the YOLO series,(3) fine-tuning models on an out-of-domain A2G dataset, improves their A2G performance and slightly improves their performance on small objects, but (4) all models still struggle with detecting small objects in an A2G scenario. We conclude that, despite recent advances in object detection, in-domain training is still crucial for creating capable ATD/R systems.
摘要:自動目標偵測與識別(ATD/R)對於軍事決策支持和(半)自主作業至關重要。最近在物體偵測和人工智慧(AI)方面的進展顯著提升了ATD/R的潛在性能。然而,公開可用的軍事數據集稀缺限制了這些系統的應用。作為解決方案,本文探討使用公開可用的模型和民用數據集,以在軍事環境中實現合理的性能。我們在新獲得的軍事相關數據集上基準測試了幾個最先進的模型,包括六個版本的YOLO系列和兩個DETR框架的變體。這個數據集包含軍事車輛和具有挑戰性的情況,包括各種程度的遮擋和小目標。每個模型的開箱即用版本與在VisDrone數據集上微調的版本一起進行驗證。這個數據集包含小物體、空對地(A2G)視角和相關類別,可能對我們的軍事ATD/R任務具有普遍性。我們使用mAP@0.5和mAP@0.5:0.95比較模型的性能,涵蓋A2G和地對地(G2G)視角、目標大小和模型大小,提供對模型實時能力的洞察。我們的主要發現是:(1)較大的模型表現優於較小的模型,(2)基於DETR的模型與YOLO系列相比顯示出有希望的結果,(3)在域外的A2G數據集上微調模型,提高了它們的A2G性能,並稍微改善了它們在小物體上的性能,但(4)所有模型在A2G場景中仍然難以偵測小物體。我們得出結論,儘管在物體偵測方面取得了最近的進展,域內訓練仍然對於創建能夠的ATD/R系統至關重要。
CABLE: Extending the Reach of Memory Retrieval via Complementary Antecedent-Based Linking and Expansion
2608.17911v1 by Zheling Tan, Jin Gao, Dequan Wang
As LLM agents operate across structured workflows and sessions, preserving long-term history does not ensure that later contexts can recover relevant evidence through a bounded memory interface. We study this evidence-reachability problem in long-term conversational memory, where retrieval still relies heavily on semantic similarity. This works well for topical recall, but it often misses earlier experiences, plans, or motivations that are semantically distant from the later events they help explain. Existing memory graphs provide cross-memory structure, yet links driven mainly by semantic overlap can duplicate what the host retriever already recovers. We argue that link construction should instead prioritize a sparse set of retriever-complementary associations. We present CABLE (Complementary Antecedent-Based Linking and Expansion), a plug-in augmentation that constructs links designed to extend the host retriever's direct semantic reach. For each new memory, CABLE generates antecedent-oriented queries, retrieves prior memories, subtracts candidates in the direct semantic neighborhood, and verifies the remainder before adding the accepted complementary associations into a sparse directed graph. At retrieval time, CABLE expands the host system's retrieved seeds along these links to surface implicit supporting evidence. We evaluate CABLE with A-MEM on LoCoMo and MA-LongMemEval, and further integrate it into SimpleMem and Mem0g on LoCoMo, using Qwen3.5-27B, DeepSeek-chat, and GPT-4o-mini. CABLE yields higher mean LLM-judge scores in every evaluated system-level setting, with the largest gains in categories where useful evidence is distributed across memories or sessions, including open-domain, multi-session, and preference-oriented questions. These results support prioritizing sparse, reasoning-relevant associations that complement rather than duplicate the host retriever.
摘要:隨著LLM代理在結構化工作流程和會話中運作,保存長期歷史並不保證後續上下文能通過有限的記憶介面恢復相關證據。我們研究這個在長期對話記憶中的證據可達性問題,其中檢索仍然在很大程度上依賴於語義相似性。這對於主題回憶來說運作良好,但它常常會錯過早期的經驗、計劃或動機,這些與後來幫助解釋的事件在語義上相距甚遠。現有的記憶圖提供了跨記憶結構,但主要由語義重疊驅動的鏈接可能會重複主檢索器已經恢復的內容。我們認為鏈接構建應優先考慮一組稀疏的檢索器互補關聯。我們提出CABLE(Complementary Antecedent-Based Linking and Expansion),這是一個插件增強,旨在構建鏈接,以擴展主檢索器的直接語義範圍。對於每個新記憶,CABLE生成以前因為導向的查詢,檢索先前的記憶,從直接語義鄰域中減去候選者,並在添加接受的互補關聯到稀疏有向圖之前驗證剩餘部分。在檢索時,CABLE沿著這些鏈接擴展主系統檢索的種子,以顯現隱含的支持證據。我們在LoCoMo和MA-LongMemEval上使用A-MEM評估CABLE,並進一步將其整合到LoCoMo上的SimpleMem和Mem0g中,使用Qwen3.5-27B、DeepSeek-chat和GPT-4o-mini。在每個評估的系統級設置中,CABLE在LLM評估者得分上均獲得更高的平均分數,在有用證據分佈於記憶或會話的類別中獲得最大的增益,包括開放域、多會話和偏好導向問題。這些結果支持優先考慮稀疏的、與推理相關的關聯,這些關聯互補而非重複主檢索器的功能。
AutoResearch: Insight In, Hallucination Out
2608.17906v1 by Yiming Ren, Xiang Liu, Qumeng Sun, Xiao Zhang, Jiahao Li, Haoyang Zhang, Junjie Wang
Autonomous research systems are increasingly capable of executing long research workflows, yet automation alone does not ensure that the resulting process remains scientifically grounded. We introduce AutoResearch, a two-stage system that connects Idea Generation with Idea Execution to address both how research ideas are formed and how they are reliably established through experimentation. In Idea Generation, AutoResearch continuously integrates emerging research signals with accumulated domain knowledge, identifies transferable mechanistic insights, and uses multi-model generation and cross-review to produce grounded, testable research plans. In Idea Execution, coordinated agents decompose these plans into experiments, iteratively implement and diagnose them, and employ independent evidence-based review before accepting research conclusions. Across representative settings in cross-modal retrieval, systems optimization, and benchmark-driven machine learning, AutoResearch turns generated ideas into measurable progress, detects and corrects unreliable experimental results, and makes evidence-conditioned decisions to continue, revise, or terminate research directions. For example, on RSICD benchmark, an AutoResearch-generated idea improves mean Recall from 32.84 to 34.69, while recording only 5 audit-confirmed issue events compared with 11-27 for other autonomous research systems. These results demonstrate a research process in which meaningful insight is grounded before experimentation and conclusions are grounded before acceptance: Insight In, Hallucination Out.
摘要:自主研究系統越來越能夠執行長期的研究工作流程,然而僅僅依賴自動化並不能確保所產生的過程保持科學基礎。我們介紹了 AutoResearch,一個兩階段的系統,將創意生成與創意執行連接起來,以解決研究想法是如何形成的,以及如何通過實驗可靠地建立這些想法。在創意生成階段,AutoResearch 持續整合新興的研究信號與累積的領域知識,識別可轉移的機制見解,並利用多模型生成和交叉審查來產出有根據、可測試的研究計劃。在創意執行階段,協調的代理將這些計劃分解為實驗,迭代實施和診斷它們,並在接受研究結論之前進行獨立的基於證據的審查。在跨模態檢索、系統優化和基準驅動的機器學習等代表性設置中,AutoResearch 將生成的想法轉化為可衡量的進展,檢測並修正不可靠的實驗結果,並做出基於證據的決策以繼續、修訂或終止研究方向。例如,在 RSICD 基準上,AutoResearch 生成的想法將平均召回率從 32.84 提高到 34.69,同時僅記錄了 5 次經審核確認的問題事件,而其他自主研究系統則記錄了 11-27 次。這些結果展示了一個研究過程,其中有意義的見解在實驗之前就已經建立,而結論在接受之前也已經有根據:見解進,幻覺出。
BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models
2608.17895v1 by Liubov Chubarova, Alexandra Kuleshova, Daniil Volkov, Kirill Sultanov, Alexey Zaytsev
While Multimodal Large Language Models (MLLMs) have made significant strides in visual comprehension, their ability to reason about text-dense, professional documents remains incompletely evaluated. Existing benchmarks emphasize information extraction, require external domain knowledge, or cover professional documents only as one of many settings. They are also largely English- or Chinese-centric, leaving other languages and Russian, in particular, substantially underrepresented. To address these limitations, we introduce BEAR-Bench (Bilingual Enterprise and Academic Reasoning), a self-contained, complex English-and-Russian benchmark comprising 1000 human-annotated questions based on text-rich business and scientific documents. We evaluate 16 proprietary and open-weight MLLMs, including Gemini 3.1 Pro and Qwen3.5-397B, on BEAR-Bench and observe clear headroom even for the strongest systems. Finally, we use the resulting model outputs to compare existing hallucination detection methods, evaluating not only how often models fail on BEAR-Bench but also how reliably those failures can be identified.
摘要:雖然多模態大型語言模型(MLLMs)在視覺理解方面取得了重大進展,但它們對於文本密集型的專業文件的推理能力仍未得到充分評估。現有的基準強調信息提取,需要外部領域知識,或者僅將專業文件作為眾多設置之一。這些基準在很大程度上以英語或中文為中心,使其他語言,特別是俄語,顯得大幅度不足。為了解決這些限制,我們推出了BEAR-Bench(雙語企業與學術推理),這是一個自包含的、複雜的英語和俄語基準,包含1000個基於文本豐富的商業和科學文件的人類標註問題。我們在BEAR-Bench上評估了16個專有和開放權重的MLLMs,包括Gemini 3.1 Pro和Qwen3.5-397B,並觀察到即使對於最強的系統也存在明顯的提升空間。最後,我們使用生成的模型輸出來比較現有的幻覺檢測方法,不僅評估模型在BEAR-Bench上的失敗頻率,還評估這些失敗能否被可靠地識別。
BayesPrompt: human readable prompts that make sense
2608.17866v1 by Franky Kevin Nando Tezoh, Ali Hussaini Umar, Alessandro Laio, Guido Sanguinetti, Riccardo Rende
Reconstructing prompts that can elicit a desired answer or behaviour in an LLM is an open and important research topic. Optimisation methods which aim at minimising the perplexity of a given answer, however, consistently yield so-called pseudoprompts, unintelligible strings of tokens which can lack human interpretability. We argue that this is a consequence of the ill-posedness of the prompt optimisation task. By reframing the task as a Bayesian posterior inference over prompts, we propose an efficient algorithm to sample prompts which are both efficient (in terms of perplexity) and human readable. We compare our approach with state of the art alternatives showing on a real data set a marked improvement over a range of metrics.
摘要:重建能夠引發大型語言模型(LLM)所需答案或行為的提示是一個開放且重要的研究主題。然而,旨在最小化給定答案困惑度的優化方法,卻持續產生所謂的偽提示,即無法理解的標記字符串,這些字符串可能缺乏人類可解釋性。我們認為這是提示優化任務不良定義的結果。通過將任務重新構架為對提示的貝葉斯後驗推斷,我們提出了一種有效的算法來抽樣既高效(在困惑度方面)又人類可讀的提示。我們將我們的方法與最先進的替代方案進行比較,顯示在一個真實數據集上,在多個指標上有顯著的改善。
ARASH: Adaptive Retrieval And Shot Selection for Tabular Prediction
2608.17856v1 by Samirasadat Jamalidinan, Yue Xu, Kazem Cheshmi
Tabular prediction is a critical task across numerous applications. The recent success of large language models has sparked various approaches for adapting them to the tabular domain. A prevalent strategy involves training or fine-tuning specialized Tabular Foundation Models (TFMs) such as TabPFN. However, TFMs require substantial computational resources, and frequent model retraining is often impractical. In-context learning (ICL), specifically, few-shot prompting, offers a resource-efficient alternative to enhance performance. Yet, identifying the most relevant rows to serve as shots remains a challenge for tabular data. This paper introduces ARASH (Adaptive, query-specific Retrieval And Shot selection), a method that improves TFM efficiency by selecting optimal shots based on local neighborhood analysis within the training set. Our results demonstrate that ARASH reduces the prompt length and memory usage of TabPFN by 1261.5$\times$ and 2.56$\times$, respectively, while providing comparable accuracy.
摘要:表格預測是許多應用中的一項關鍵任務。大型語言模型的近期成功激發了各種將它們適應於表格領域的方法。一種普遍的策略涉及訓練或微調專門的表格基礎模型(TFMs),如TabPFN。然而,TFMs需要大量的計算資源,並且頻繁的模型重訓練往往不切實際。在上下文學習(ICL)中,特別是少量樣本提示,提供了一種資源高效的替代方案來提升性能。然而,識別最相關的行作為樣本仍然是表格數據的一個挑戰。本文介紹了ARASH(自適應、查詢特定的檢索和樣本選擇),這是一種通過基於訓練集內的局部鄰域分析選擇最佳樣本來提高TFM效率的方法。我們的結果顯示,ARASH分別將TabPFN的提示長度和內存使用量減少了1261.5$\times$和2.56$\times$,同時提供了可比的準確性。
Encoded but Not Actionable: Auditing the Decode-Generate-Steer Gap in Frozen LLMs for Geometric Constraints
2608.17843v1 by Man Liang, Xinzhao Cheng, Faizan Wajid
Large language models (LLMs) have demonstrated strong performance on structured reasoning tasks, but what they encode and whether it informs model behavior remain unclear. We investigate this question through geometric reasoning, using parametric CAD constraints as a controlled testbed for separating local pairwise relations from sketch-level constraint status. By probing the hidden states of six frozen decoder-only LLMs, we examine four properties: linear decodability, forced-choice generation, activation-level influence, and behavioral steerability. Pretraining substantially improves the decoding of local geometric relations, and this advantage persists after accounting for positional cues with shuffled-order controls. In contrast, sketch-level DOF status is already highly decodable from randomly initialized representations and improves only modestly with pretraining, indicating that much of its probe performance is available without learned weights. Further analyses show that decodable information is not always actionable. Generation often fails to express this information, and on the two intervention-tested backbones, activation-restoration effects at the patched entity position vanish while decodability persists across depth. Mean-difference steering also does not reliably control outputs. These results show that decodability, generation, activation-level influence, and steerability can diverge in the tested setting. The audit provides a controlled way to distinguish failures to encode geometric structure from failures to express or control encoded information.
摘要:大型語言模型(LLMs)在結構推理任務中表現出色,但它們編碼了什麼以及這是否影響模型行為仍不清楚。
我們通過幾何推理來研究這個問題,使用參數化CAD約束作為控制測試平台,以區分局部成對關係和草圖級約束狀態。
通過探測六個凍結的僅解碼器LLMs的隱藏狀態,我們檢查了四個特性:線性可解碼性、強制選擇生成、激活水平影響和行為可引導性。
預訓練顯著改善了局部幾何關係的解碼,並且在考慮到隨機順序控制的位置信息後,這一優勢仍然存在。
相比之下,草圖級DOF狀態已經可以從隨機初始化的表示中高度可解碼,並且在預訓練後僅有適度改善,這表明其探測性能在沒有學習權重的情況下就已經可用。
進一步分析顯示,可解碼的信息並不總是可操作的。
生成通常未能表達這種信息,在兩個經過干預測試的骨幹上,修補實體位置的激活恢復效應消失,而可解碼性在深度上仍然存在。
均值差異引導也無法可靠地控制輸出。
這些結果顯示,在測試環境中,可解碼性、生成、激活水平影響和可引導性可能會出現分歧。
這次審核提供了一種控制方式,以區分編碼幾何結構的失敗與表達或控制編碼信息的失敗。
AdaLens: Interactive Storyline for Monitoring and Steering Long-Running Agentic Data Analysis
2608.17834v1 by Yangtian Liu, Yan Miao, Shuhan Liu, Yunfan Zhou, Dae Hyun Kim, Di Weng, Yingcai Wu
Large language models are pushing data science toward increasingly autonomous and agentic workflows, with recent systems already supporting multi-step and long-running analyses. As these workflows become more autonomous, conventional interfaces no longer provide adequate support for two critical requirements: observability for understanding an agent's evolving reasoning and evidence, and steerability for redirecting low-value directions or deepening promising ones during execution. Existing interactive approaches improve process visibility and open intervention points, but they remain largely designed for discrete, turn-by-turn exchanges rather than the parallel branches and evolving decision structures of long-running agentic analysis. We study this need as interactive oversight in long-running agentic data analysis and present AdaLens, an interactive system for monitoring and steering ongoing runs. AdaLens combines a storyline-based representation that unifies analytical plans, execution progress, intermediate findings, and data-column involvement with steering interactions grounded in these analytical elements for directional guidance and execution control. We evaluate AdaLens through two case studies and a user study, examining how it supports analysts in monitoring and steering long-running agentic data analysis.
摘要:大型語言模型正在推動數據科學朝向越來越自主和具代理性的工作流程,最近的系統已經支持多步驟和長期運行的分析。隨著這些工作流程變得更加自主,傳統界面不再能夠充分支持兩個關鍵需求:可觀察性以理解代理人不斷演變的推理和證據,以及可引導性以在執行過程中重新定向低價值的方向或加深有前景的方向。現有的互動方法改善了過程的可見性並開放了干預點,但它們主要是為了離散的、逐步的交流而設計,而不是針對長期運行的代理分析中的平行分支和不斷演變的決策結構。我們研究這一需求作為長期運行的代理數據分析中的互動監督,並提出了AdaLens,一個用於監控和引導正在進行的運行的互動系統。AdaLens結合了一種基於故事情節的表示,統一了分析計劃、執行進度、中間發現和數據列參與,並基於這些分析元素提供引導互動,以實現方向指引和執行控制。我們通過兩個案例研究和一項用戶研究來評估AdaLens,檢視它如何支持分析師監控和引導長期運行的代理數據分析。
The Model's Tell: Measuring Context-Leakage Attack Signals with Behavior Gauges
2608.17829v1 by Maosen Zhang, Jianshuo Dong, Boting Lu, Wenyue Li, Xiaoping Zhang, Tianwei Zhang, Jie Zhang, Han Qiu
LLMs increasingly rely on external contexts, such as pre-defined system prompts or retrieved documents, to improve generation quality. However, processing these contexts alongside user queries creates an attack surface: adversarial inputs can induce models to disclose them. Prior probing studies suggest that leakage-related signals emerge in hidden states, yet the need to extract these states poses additional deployment challenges. In this paper, we explore whether this internal signal leaves a more accessible ``tell'' before decoding. We propose LeakGauge, which probes this response by appending a suffix that gauges leakage behavior and mapping its prefill token probabilities to an attack-risk score. While a direct gauge uses the initial tokens of confidential content, we find that a content-agnostic one that verbalizes leakage behavior yields more robust signals. Across 11 LLMs, including GLM-5.2 (753B) and Kimi-K3 (2.8T), LeakGauge reaches an AUROC range of 0.944--0.996 on unseen attacks. The signal remains stable when the content changes language or the attack shifts from verbatim to semantic disclosure. By activation-steering interventions, we further show that the risk score is sensitive to an internal leakage-related direction, relating the observable signal to the model's internal representation. In addition, LeakGauge enables an input detector with fewer than 0.5K extra parameters and added latency of 10.34 ms. Code: \href{https://github.com/yeasen-z/LeakGauge}.
摘要:LLM越來越依賴外部上下文,例如預定義的系統提示或檢索的文件,以提高生成質量。
然而,將這些上下文與用戶查詢一起處理會創造攻擊面:對抗性輸入可能會誘使模型洩露它們。
先前的探測研究表明,與洩漏相關的信號在隱藏狀態中出現,但提取這些狀態的需求帶來了額外的部署挑戰。
在本文中,我們探討這個內部信號是否在解碼之前留下更易於訪問的“告訴”。
我們提出了LeakGauge,它通過附加一個後綴來探測這個反應,以評估洩漏行為並將其預填令牌的概率映射到攻擊風險分數。
雖然直接的評估使用了機密內容的初始令牌,但我們發現一個與內容無關的評估,能夠表達洩漏行為,產生更穩健的信號。
在11個LLM中,包括GLM-5.2 (753B)和Kimi-K3 (2.8T),LeakGauge在未見過的攻擊上達到了0.944到0.996的AUROC範圍。
當內容變更語言或攻擊從逐字披露轉變為語義披露時,信號仍然穩定。
通過激活引導干預,我們進一步顯示風險分數對內部洩漏相關方向敏感,將可觀察信號與模型的內部表示相關聯。
此外,LeakGauge使得輸入檢測器的額外參數少於0.5K,並增加了10.34毫秒的延遲。
代碼:\href{https://github.com/yeasen-z/LeakGauge}。
From Global Benchmarks to Local Evaluations: Benchmarking LLMs for the German Public Sector
2608.17827v1 by Camilla Dalerci, Thilo Michael, Robin Schaefer, Daniel Weinland
Public institutions face a persistent challenge in selecting LLMs suited to their specific context. Existing benchmarks, however, are of limited use as they primarily reflect English-language and US-centric settings, and often only evaluate task performance. In this paper, we present first results of MÖVE, a holistic evaluation framework for the German public sector, examining three rarely considered governance dimensions: energy consumption, provider transparency, and knowledge of German-party positions. Our results reveal significant trade-offs, with no single model excelling across all dimensions: estimated energy consumption varies more than 60-fold and is not explained by model size alone, information disclosure varies systematically across providers, and European models do not exhibit stronger knowledge of German party positions. Model selection for public institutions thus cannot rely on performance rankings alone. Instead, evaluations should also reflect the governance requirements of the deployment context.
摘要:公共機構在選擇適合其特定情境的LLM時面臨持續的挑戰。
然而,現有的基準測試用途有限,因為它們主要反映英語和美國中心的環境,且通常僅評估任務表現。
在本文中,我們呈現MÖVE的初步結果,這是一個針對德國公共部門的整體評估框架,檢視三個鮮少考慮的治理維度:能源消耗、供應商透明度和對德國政黨立場的了解。
我們的結果揭示了顯著的權衡,沒有單一模型在所有維度上表現優異:估計的能源消耗變化超過60倍,且僅以模型大小無法解釋,信息披露在不同供應商之間系統性變化,歐洲模型對德國政黨立場的了解並未顯示出更強的優勢。
因此,公共機構的模型選擇不能僅依賴於性能排名。
相反,評估還應反映部署情境的治理要求。
MotoSafety: Edge-AI with Learned Temporal Importance for Two-Wheeler Collision Risk Assessment Under Time Pressure
2608.17823v1 by Sumit S. Shevtekar, Chandresh K. Maurya, Gourab Sil, Subasish Das
Powered two-wheeler riders face critical safety challenges in low- and middle-income countries, yet limited studies exist on how cognitive stressors such as Time Pressure influence collision risk. To address this gap, we introduce a large-scale dataset of over 129,000 labeled multivariate time-series sequences from 153 simulator rides by 51 participants under No, Low, and High TP, capturing 64 features across vehicle dynamics, control inputs, proximity, and behavioral violations. Building on this dataset, we propose MotoSafety, a novel edge-AI architecture grounded in the Learned Temporal Importance principle. MotoSafety achieves 94.97% accuracy and 99.33% ROC AUC, outperforming ten baselines, including TimesNet and LLM4TS, and achieves 0.039 MSE and 0.094 MAE for forecasting (4.4x lower error than Time-LLM and iTransformer). With only 1.15M parameters and 0.135 ms latency, it is suitable for edge deployment on low-cost CPU hardware. Using ground truth TP as an inductive bias improves accuracy from 94.09% to 94.97%, while predicted TP achieves 94.82%. Using only 21 IMU+GPS features, it achieves 93.91% accuracy, indicating practical deployment. Beyond PTW safety, the architecture shows better transferability to human activity (97.66%) and clinical (99.65%) domains. This lightweight framework advances PTW collision risk assessment, supporting the Safe System Approach for Intelligent Transportation Systems.
摘要:在中低收入國家,動力二輪車騎士面臨著重大的安全挑戰,但關於認知壓力因素如時間壓力如何影響碰撞風險的研究卻相對有限。為了填補這一空白,我們引入了一個大規模數據集,該數據集包含來自51名參與者在無時間壓力、低時間壓力和高時間壓力下進行的153次模擬騎行的超過129,000個標記的多變量時間序列,捕捉了64個特徵,涵蓋了車輛動態、控制輸入、接近度和行為違規。基於這個數據集,我們提出了MotoSafety,一種基於學習時間重要性原則的新型邊緣人工智慧架構。MotoSafety實現了94.97%的準確率和99.33%的ROC AUC,超越了包括TimesNet和LLM4TS在內的十個基準,並在預測中達到了0.039的均方誤差和0.094的平均絕對誤差(比Time-LLM和iTransformer低4.4倍)。它僅需1.15M的參數和0.135毫秒的延遲,適合在低成本CPU硬體上進行邊緣部署。使用真實的時間壓力作為歸納偏見,準確率從94.09%提高到94.97%,而預測的時間壓力則達到94.82%。僅使用21個IMU+GPS特徵,它的準確率達到93.91%,顯示出實際部署的潛力。除了PTW安全性外,該架構在人體活動(97.66%)和臨床(99.65%)領域也顯示出更好的可轉移性。這個輕量級框架推進了PTW碰撞風險評估,支持智能交通系統的安全系統方法。
Interpretable Humans, Alien LLMs: Expert Analysis of Latent Structures in Assessment Responses
2608.17810v1 by Alona Strugatski, Licol Zeinfeld, Jason Cooper, Shelley Rap, Gil Schwarts, Giora Alexandron
The evaluation of large language models (LLMs) relies heavily on human-designed assessments, implicitly assuming that AI and humans employ similar underlying cognitive constructs. Challenging this assumption, we investigate whether the latent factors governing LLM performance carry the same substantive, human-interpretable meaning as the cognitive constructs governing human learners. Using responses from humans and six LLMs across quantitative reasoning and chemistry assessments, we conducted Exploratory Factor Analysis (EFA) separately for both groups. Subject-Matter Experts (SMEs) then blindly evaluated the resulting factor graphs to ascribe pedagogical meaning to the emerged constructs. SMEs successfully interpreted most of the human-derived factors. Conversely, they could not ascribe meaning to any LLM-derived factors in quantitative reasoning and interpreted only half of the LLM factors in chemistry. By combining data-driven EFA with blind expert interpretation, this framework shows that LLMs frequently operate on statistically opaque mechanisms distinct from human reasoning.
摘要:大型語言模型(LLMs)的評估在很大程度上依賴於人類設計的評估,隱含假設AI和人類使用相似的基本認知結構。挑戰這一假設,我們調查了支配LLM性能的潛在因素是否具有與支配人類學習者的認知結構相同的實質性、人類可解釋的意義。利用來自人類和六個LLM在定量推理和化學評估中的反應,我們分別對這兩組進行了探索性因素分析(EFA)。主題專家(SMEs)隨後盲目評估了所產生的因素圖,以賦予出現的結構教學意義。SMEs成功解釋了大多數人類衍生的因素。相反,他們無法為任何LLM衍生的因素在定量推理中賦予意義,並且只解釋了化學中一半的LLM因素。通過將數據驅動的EFA與盲專家解釋相結合,這一框架顯示LLMs經常在與人類推理不同的統計不透明機制上運作。
Whether LLMs Can Navigate Beliefs and Facts Depends on How You Phrase It
2608.17809v1 by Quang Minh Nguyen, Luis Frentzen Salim
Humans naturally form and express beliefs in daily communication, e.g., "I think the answer is 3" or "I suppose that's right." Such beliefs inevitably intertwine with fact and knowledge, making the ability to handle them in tandem desirable for large language models (LLMs), as they are increasingly deployed in user-facing settings. Prior work showed that even capable LLMs exhibit a systemic weakness in acknowledging user beliefs grounded in incorrect information. We extend this evaluation to 10 LLMs across 18 epistemic expressions and find that the size and direction of the weakness depend on the verb used to express the belief, with the accuracy gap between factual and false information ranging from +50% on "I vaguely remember" to -14% on "I seriously doubt". We further show that the phenomenon stems from task confusion: models default to fact-checking the underlying claim, overriding the user's stated belief; chains of thought that explicitly fact-check show lower accuracy on false information than those that do not; and a single instruction can reverse the failure across verb families. Mechanistically, models attend more to false beliefs they fail to confirm, but suppressing this attention at decoding time recovers accuracy only partially and only in some models, calling for future work on intervention methods. Our findings clarify prior results and show how fact-checking, a generally desirable behavior, can interfere with belief tracking in LLMs. Our code is available at https://github.com/ngqm/belief-fact-phrasing.
摘要:人類在日常交流中自然地形成和表達信念,例如「我認為答案是3」或「我想這是對的」。這些信念不可避免地與事實和知識交織在一起,使得同時處理它們的能力對大型語言模型(LLMs)來說變得可取,因為它們在面向用戶的環境中越來越多地被部署。先前的研究顯示,即使是能幹的LLMs在承認基於錯誤信息的用戶信念方面也存在系統性的弱點。我們將這一評估擴展到18種認識表達下的10個LLMs,發現弱點的大小和方向取決於用來表達信念的動詞,事實信息與虛假信息之間的準確性差距從「我模糊地記得」的+50%到「我嚴重懷疑」的-14%不等。我們進一步表明,這一現象源於任務混淆:模型默認檢查基礎主張的事實,覆蓋用戶所表達的信念;明確進行事實檢查的思維鏈在虛假信息上的準確性低於那些不進行檢查的;而單一指令可以逆轉動詞家族中的失敗。在機制上,模型對它們未能確認的虛假信念的注意力更高,但在解碼時抑制這種注意力僅能部分恢復準確性,且僅在某些模型中有效,這呼籲未來對干預方法的研究。我們的發現澄清了先前的結果,並顯示事實檢查這一通常可取的行為如何干擾LLMs中的信念追蹤。我們的代碼可在 https://github.com/ngqm/belief-fact-phrasing 獲得。
An Empirical Study of Reward Specification and Benchmark Reliability in GRPO-based LLM Unlearning
2608.17804v1 by Rubén Balbastre, Juan Manuel Orduña, Mariano Pérez
Practical LLM unlearning is usually evaluated through two objectives: suppress target-specific knowledge and preserve non-target utility. In generative QA, this leaves a third behavior underspecified: when a target-adjacent prompt admits a broader answer without target-specific leakage, the model should answer at that level rather than leak, evade, or refuse. We study this specification problem in a controlled LoRA-GRPO RWKU setting, comparing four reward designs that span lexical suppression, anti-refusal shaping, rubric-based broad answering, and an explicit refusal contrast, with and without SFT warm-up. The experiments show that optimization success is not equivalent to behavioral unlearning: RWKU forget scores, held-out completion audits, terminal training-rollout audits, and training dynamics can point to different conclusions. We trace these disagreements to reward-hacking endpoints, policy-support limits in GRPO, benchmark probes that miss endpoint changes, and rewards that can select broad-topic answering with low semantic leakage during optimization.
摘要:實際的 LLM 忘記通常通過兩個目標來評估:抑制特定目標的知識和保留非目標的效用。
在生成性問答中,這留下了第三種行為未明確規範:當一個與目標相近的提示允許更廣泛的回答而不泄露特定目標時,模型應該在該層次上回答,而不是泄露、逃避或拒絕。
我們在一個受控的 LoRA-GRPO RWKU 設定中研究這個規範問題,比較四種獎勵設計,涵蓋詞彙抑制、反拒絕塑造、基於標準的廣泛回答以及明確的拒絕對比,並且有無 SFT 熱身。
實驗表明,優化成功並不等同於行為上的忘記:RWKU 忘記分數、保留的完成審計、終端訓練回滾審計和訓練動態可能指向不同的結論。
我們將這些分歧追溯到獎勵駭客端點、GRPO 中的政策支持限制、錯過端點變化的基準探針,以及在優化過程中可以選擇廣泛主題回答且語義泄露低的獎勵。
StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows
2608.17800v1 by Liya Zhu, Xin Ma, Tao Liu, Haodong Wang, Ge Zhang, Jingzhe Ding, Qingshui Gu, Yongjie Zhong, Jinxiang Meng, Yuan Gao, Yunqiu Zhou, Hao Zhu, Jifeng He, Yongzhi Liao, Xinyi Zhang, Chaoxin Li, Yi Zhu, Xi Lin, Duju Zeng, Xiang Gao, Wen Zhang, Yunyang Wang, Duo Wang, Huan Zhou, Zuo Wang, Jin Chen, Kaiyuan Zhang, Chuqian Yu, Tianhao Yu, Longxiang Liu, Jianbo Xue, Huimin Che, Jiahao Wang, Yujia Qin, Jiaheng Liu, Shen Yan, Xiaolong Chang, Wenhao Huang
Recent advances in Large Language Models(LLMs) and agents have substantially improved the ability of AI systems to execute complex tasks. Yet existing benchmarks largely rely on researcher-selected tasks, leaving uncertain whether such progress extends to the work that real-world users actually demand from AI systems. We introduce \textbf{StartupBench}, an E2E agent benchmark grounded in market-validated AI startup products. Rather than defining tasks from pre-defined assumptions about useful agent capabilities, we systematically study AI products with demonstrated adoption, together with their product workflows and users, to identify real-world tasks for which AI has established practical demand across diverse professional domains. We translate these workflows into complete deliverable-oriented tasks and evaluate them with fine-grained rubrics capturing their complex requirements. Across representative models evaluated under a unified agent harness, even the strongest model successfully completes only approximately 30\% of StartupBench, despite making substantial partial progress on many tasks. Further analysis identifies aspects like complex instruction following and domain-specific expertise as major sources of failure. Our results reveal that many market-validated workflows remain beyond the reliable capabilities of current general-purpose agents, establishing StartupBench as an empirical measure of progress toward E2E completions of real-world user tasks.
摘要:最近在大型語言模型(LLMs)和代理方面的進展顯著提高了 AI 系統執行複雜任務的能力。
然而,現有的基準主要依賴研究者選擇的任務,這使得不確定這樣的進展是否延伸到現實世界用戶對 AI 系統的實際需求。
我們介紹了 \textbf{StartupBench},這是一個基於市場驗證的 AI 初創產品的端到端代理基準。
我們不是從對有用代理能力的預定假設中定義任務,而是系統性地研究已經被採用的 AI 產品,連同它們的產品工作流程和用戶,以識別 AI 在各種專業領域中已建立的實際需求的任務。
我們將這些工作流程轉化為完整的交付導向任務,並使用細緻的評分標準來評估它們,捕捉其複雜的要求。
在統一的代理框架下評估的代表性模型中,即使是最強的模型也僅成功完成約 30\% 的 StartupBench,儘管在許多任務上取得了顯著的部分進展。
進一步分析確定了複雜指令遵循和特定領域專業知識等方面是主要的失敗來源。
我們的結果顯示,許多市場驗證的工作流程仍超出當前通用代理的可靠能力,確立了 StartupBench 作為向現實世界用戶任務的端到端完成進展的實證衡量標準。
Training with synthetic data for drone detection in thermal imagery
2608.17799v1 by Tanel Liiv, Sander Soodla, Nzamba Bignoumba, Alma M. Liezenga, Toomas Pruuden
Ground-to-Air (G2A) drone detection in medium- and long-wave infrared (MWIR/LWIR) imagery is challenging due to reduced texture information, sensor noise, weak thermal contrast, and the scarcity of annotated data. This work investigates a synthetic-first training strategy that combines synthetic scene generation with fine-tuning on real data. We show that synthetic data provides an effective basis for learning initial object representations, while real in-domain thermal imagery is still essential for reliable deployment. Even small amounts of real IR data substantially reduce domain gaps. Our experiments indicate that dataset alignment has a stronger impact on performance than model scale. Finally, our analysis of the dataset suggests that semantic alignment in feature space is the strongest predictor of model performance, while radiometric properties such as entropy and dynamic range also contribute to detection robustness. This work provides a foundation for combining synthetic and real IR data for effective G2A drone detection.
摘要:地面對空 (G2A) 無人機在中波和長波紅外 (MWIR/LWIR) 影像中的檢測具有挑戰性,這是因為紋理資訊減少、傳感器噪聲、熱對比度弱以及標註數據的稀缺。本研究探討了一種合成優先的訓練策略,將合成場景生成與真實數據的微調相結合。我們顯示合成數據為學習初始物體表示提供了有效的基礎,而真實的域內熱影像仍然對可靠的部署至關重要。即使是少量的真實紅外數據也能顯著減少域間差距。我們的實驗表明,數據集對齊對性能的影響比模型規模更強。最後,我們對數據集的分析表明,特徵空間中的語義對齊是模型性能最強的預測指標,而熵和動態範圍等輻射特性也有助於檢測的穩健性。本研究為有效的 G2A 無人機檢測結合合成和真實紅外數據提供了基礎。
TraceSQL: Traceable Answerability Estimation for Reference-Free Text-to-SQL Verification
2608.17795v1 by Neelesh Kumar Shukla, Debasmita Panda, Srutanik Bhaduri, Aditya Banerjee, Viji Krishnamurthy
Text-to-SQL systems are commonly evaluated using ground-truth SQL queries or reference execution results, but such supervision is unavailable at inference time in real-world deployments. This creates a critical verification problem: given only a user question, database context, and generated SQL, can a system estimate whether the generated query is likely to correctly answer the question? Recent approaches use LLMs as judge or specialized agents to inspect generated SQL, but their decisions can be difficult to trace. Outcome Reward Models (ORMs) address this by learning from execution-labeled candidate SQLs and assigning correctness scores to unseen queries, yet they still provide limited visibility into the signals behind each verification. To address this limitation, we propose TraceSQL, a lightweight and traceable verification model built on explicit diagnostic features. TraceSQL combines 67 features capturing question ambiguity, question requirements, question-schema-SQL consistency, SQL structure, and intent alignment. These signals remain available for examining which factors influence each prediction and for tracing decisions back to diagnostic evidence. On BIRD development databases, TraceSQL achieves 66.47% F1 and 64.48% ROC-AUC, compared with 61.87% F1 and 58.26% ROC-AUC for the GradeSQL-7B ORM baseline on the same generated-SQL evaluation. Feature attribution further shows that the model relies on both semantic grounding and deterministic SQL-structure signals. These results show that SQL verification can be performed with a lightweight learned model while retaining feature-level evidence for inspecting and diagnosing its predictions.
摘要:文本到 SQL 的系統通常使用真實的 SQL 查詢或參考執行結果來進行評估,但在現實世界的部署中,這種監督在推理時是不可用的。這創造了一個關鍵的驗證問題:僅根據用戶問題、數據庫上下文和生成的 SQL,系統能否估計生成的查詢是否可能正確回答問題?最近的方法使用大型語言模型(LLMs)作為評判或專門代理來檢查生成的 SQL,但它們的決策可能難以追蹤。結果獎勵模型(ORMs)通過從執行標記的候選 SQL 中學習並為未見查詢分配正確性分數來解決這個問題,但它們仍然提供有限的可見性來了解每個驗證背後的信號。為了解決這一限制,我們提出了 TraceSQL,一個基於明確診斷特徵的輕量級且可追蹤的驗證模型。TraceSQL 結合了 67 個特徵,捕捉問題模糊性、問題要求、問題-模式-SQL 一致性、SQL 結構和意圖對齊。這些信號仍然可用於檢查影響每個預測的因素,並將決策追溯到診斷證據。在 BIRD 開發數據庫上,TraceSQL 在相同生成 SQL 評估中達到 66.47% 的 F1 和 64.48% 的 ROC-AUC,而 GradeSQL-7B ORM 基準的 F1 為 61.87% 和 ROC-AUC 為 58.26%。特徵歸因進一步顯示,該模型依賴於語義基礎和確定性 SQL 結構信號。這些結果表明,SQL 驗證可以使用輕量級的學習模型來進行,同時保留特徵級的證據以檢查和診斷其預測。
Preference Is Not Intervention: The Structure and Stability Boundaries of Reader-Specific Evidence Utility
2608.17781v1 by Shi Zhou
ML systems increasingly condition decisions on downstream model identity, but this is useful only if model-specific differences form reusable structure rather than input-local interactions. We test this in retrieval-augmented generation (RAG), where evidence utility can be measured under controlled interventions. Holding query, evidence, task, scoring, and intervention fixed, nine readers disagree on effect sign in 33\% of jointly affected cells; reader$\times$query interaction explains 29.8\% of utility variance versus an 8.4\% permutation null; and self-selected evidence improves F1 by $+0.031$ ($t=3.39$). We then ask the sharper question: \emph{which components of this heterogeneity are stable reader properties across queries?} Separating three measurable objects---evidence \emph{activity}, \emph{ordinal preference}, and \emph{conditional signed direction}---we find ordinal reader geometry stable across four independent settings (split-half $ρ=0.60$--$0.83$): leave-one-out interventions, PRISM preferences, RAMDocs, and RAGuard. Signed geometry is task-bounded: weak in open-ended QA (0.14, 0.35), especially for misleading and irrelevant evidence, but strong in binary fact-checking (0.75) with no significant ordinal gap, though still below its sparsity-matched ceiling. Sparsity, decoding noise, and metric artifacts do not explain the main ordinal--signed gap. Finally, stable ordinal similarity fails to predict cross-reader intervention transfer (oracle-distance $ρ=-0.27$; regret reliability $-0.28$). Reader-specific utility exists, but preference is not intervention: stable ranking similarity does not license transfer of help/harm decisions.
摘要:ML 系統越來越多地根據下游模型的身份來決策,但這只有在模型特定的差異形成可重用結構而不是輸入局部交互時才有用。
我們在檢索增強生成(RAG)中測試這一點,在這裡證據的效用可以在控制干預下進行測量。
在查詢、證據、任務、評分和干預固定的情況下,九位讀者在 33\% 的共同影響單元中對效果符號存在分歧;讀者$\times$查詢交互解釋了 29.8\% 的效用變異,相較於 8.4\% 的置換虛無;自選證據使 F1 提高了 $+0.031$ ($t=3.39$)。
然後我們提出更尖銳的問題:\emph{這種異質性的哪些組成部分是跨查詢的穩定讀者特性?}
通過分離三個可測量的對象——證據 \emph{活動}、\emph{序數偏好} 和 \emph{條件符號方向}——我們發現序數讀者幾何在四個獨立設置中保持穩定(分半 $ρ=0.60$--$0.83$):留一法干預、PRISM 偏好、RAMDocs 和 RAGuard。
符號幾何是任務界限的:在開放式問答中較弱(0.14, 0.35),特別是對於誤導性和不相關的證據,但在二元事實檢查中較強(0.75),並且沒有顯著的序數差距,儘管仍低於其稀疏匹配的上限。
稀疏性、解碼噪音和度量工件無法解釋主要的序數-符號差距。
最後,穩定的序數相似性無法預測跨讀者干預轉移(oracle-distance $ρ=-0.27$;遺憾可靠性 $-0.28$)。
存在讀者特定的效用,但偏好不是干預:穩定的排名相似性並不授權幫助/傷害決策的轉移。
Learnware for CSI Feedback: Scene-specific Small Models Can Do Big
2608.17760v1 by Xiangyi Li, Jiajia Guo, Chao-Kai Wen, Xin Geng, Shi Jin, Zhi-Hua Zhou
Intelligent channel state information (CSI) feedback is essential for realizing the high capacity and spectral efficiency goals of future 6G systems, yet existing deep learning solutions face a trade-off between model generalization and scenario-specific performance. Large neural networks generalize well but incur high computational and tuning costs, while small models excel in particular environments but require repetitive costly end-to-end training for each base station (BS). To address these challenges, we introduce a model repository-based deployment framework in which a centralized AI data center maintains a catalog of scene-specific CSI models. The repository is enhanced with a Learnware-based framework, where each model is associated with a specification including semantic part (network architecture parameters) and statistical part (codeboo-fingerprint embeddings of training-data distributions). A BS submits only its local statistical specifications to retrieve the most relevant pre-trained model, enhancing data privacy by avoiding raw CSI transmission and drastically reducing retrieval latency and communication overhead. We further develop a data-driven search strategy that matches codebook fingerprints to model performance, achieving over 90% selection accuracy. In simulations, our scheme yields 18.8% and 57.7% performance improvements over the General Model in LOS and NLOS scenarios, respectively while reducing local fine-tuning by up to 1000 samples and 100 epochs. This Learnware-based approach minimizes redundant training, maximizes model reuse, and supports rapid,privacy-enhancing deployment of CSI feedback models.
摘要:智能通道狀態資訊(CSI)反饋對於實現未來6G系統的高容量和頻譜效率目標至關重要,然而現有的深度學習解決方案在模型泛化和場景特定性能之間面臨權衡。大型神經網絡具有良好的泛化能力,但會產生高計算和調整成本,而小型模型在特定環境中表現出色,但需要對每個基站(BS)進行重複昂貴的端到端訓練。為了解決這些挑戰,我們引入了一個基於模型庫的部署框架,其中一個集中式AI數據中心維護著場景特定的CSI模型目錄。該庫通過一個基於Learnware的框架進行增強,每個模型都與一個規範相關聯,包括語義部分(網絡架構參數)和統計部分(訓練數據分佈的代碼簿指紋嵌入)。基站僅提交其本地統計規範,以檢索最相關的預訓練模型,通過避免原始CSI傳輸來增強數據隱私,並大幅減少檢索延遲和通信開銷。我們進一步開發了一種數據驅動的搜索策略,將代碼簿指紋與模型性能匹配,實現了超過90%的選擇準確率。在模擬中,我們的方案在LOS和NLOS場景中分別比通用模型提高了18.8%和57.7%的性能,同時將本地微調減少了多達1000個樣本和100個訓練周期。這種基於Learnware的方法最小化了冗餘訓練,最大化了模型重用,並支持快速、增強隱私的CSI反饋模型部署。
D$^2$ACCI: A Dual-Loop Diagnostic Protocol for Evidence-Preserving Agent Memory
2608.17756v1 by Xule Liu, Yijun Liu, Chao Li, Shao Kun
Memory is a key capability of LLM agents. Persistent memory extends this across sessions---enabling recall, revision, and personalization. Yet its multi-stage pipeline (ingestion, retrieval, filtering, generation) makes failures difficult to localize: end-to-end evaluation reveals that an error occurred, but not which stage caused it. Existing evaluations often report aggregate performance without paired statistical comparisons, slice-level non-regression checks, or stage-level diagnostic traces. We propose D$^2$ACCI (Diagnostic-Driven Artifact-based Closed-loop Controlled Iteration), a dual-loop protocol whose outer diagnostic gate promotes, feature-flags, or rejects memory interventions based on paired evidence, protected-slice monitoring, and trace-level localizability. We further introduce DCR, a graded observability metric that measures whether failures remain localizable, and D$^2$ACCI-Eval, a reusable artifact for gate replay. We instantiate the protocol in MemStack and evaluate on three public benchmarks, achieving 93.59% on LoCoMo, 90.93% on LongMemEval, and 57.20% on PersonaMem-V2. Five paired ablations show that supplement extraction, session-memory retrieval, and Forget Guard yield statistically significant gains (+1.9 to +3.7pp, all p $\le$ .003). In contrast, BM25/RRF is retained as a monitored feature flag---a distinction invisible to aggregate-only evaluation. A diagnostic audit shows enriched traces substantially improve root-cause agreement over result-only relabeling. Diagnostic artifacts reach 98--100% DCR@3 versus 0% for results-only logs. These results establish that robust memory-system iteration demands traceable, statistically grounded, and regression-aware evidence---exactly the gap D$^2$ACCI fills.
摘要:記憶是 LLM 代理的一項關鍵能力。持久記憶擴展了這一能力,使其跨越多個會話——實現回憶、修訂和個性化。然而,其多階段管道(攝取、檢索、過濾、生成)使得故障難以定位:端到端評估顯示發生了錯誤,但無法確定是哪一階段導致的。現有的評估通常報告綜合性能,而沒有配對的統計比較、切片級別的非回歸檢查或階段級別的診斷追蹤。我們提出 D$^2$ACCI(基於診斷的工件閉環控制迭代),這是一種雙循環協議,其外部診斷閘根據配對證據、受保護的切片監控和追蹤級別的可定位性來促進、標記或拒絕記憶干預。我們進一步介紹 DCR,一種分級可觀察性指標,用於衡量故障是否仍然可定位,以及 D$^2$ACCI-Eval,一種可重用的工件,用於閘重放。我們在 MemStack 中實現該協議,並在三個公共基準上進行評估,在 LoCoMo 上達到 93.59%、在 LongMemEval 上達到 90.93%、在 PersonaMem-V2 上達到 57.20%。五個配對的消融實驗顯示,補充提取、會話記憶檢索和忘記保護產生了統計上顯著的增益(+1.9 到 +3.7 個百分點,所有 p $\le$ .003)。相比之下,BM25/RRF 被保留為監控的特徵標記——這一區別在僅進行綜合評估時是不可見的。診斷審計顯示,豐富的追蹤顯著改善了根本原因的一致性,相較於僅結果的重新標記。診斷工件在 DCR@3 上達到 98--100%,而僅結果的日誌為 0%。這些結果表明,穩健的記憶系統迭代需要可追蹤的、統計基礎的和意識到回歸的證據——正是 D$^2$ACCI 所填補的空白。
The Curious Case of Exploding DecPOMDPs: Containing the Fire through Policy Counting
2608.17749v1 by Nazlı Nur Karabulut, tanya Braun
Decentralised partially observable Markov decision processes (DecPOMDPs) provide a general framework for modelling multi-agent decision making under uncertainty. However, DecPOMDPs are known to suffer from exponential complexity in the number of agents. One way to combat this intractability in agent numbers is to look at partitions of agents that exhibit a form of symmetry among agents, allowing for a compact encoding by counting. However, a challenge arises as the policy space explodes, even though the model complexity and evaluation cost reduce to a polynomial dependence. In this paper, we redirect our focus from counting agents to counting policies, which actually enables tractability in agent numbers for so called policy-counted DecPOMDPs. Further, we present policy-counted dynamic programming using the compact representation to solve policy-counted DecPOMDPs efficiently.
摘要:去中心化的部分可觀察馬可夫決策過程(DecPOMDPs)提供了一個通用框架,用於建模在不確定性下的多代理決策。
然而,DecPOMDPs 以代理數量的指數複雜性而聞名。
對抗代理數量的這種難以處理的情況的一種方法是考慮具有某種對稱性的代理分區,這樣可以通過計數來實現緊湊編碼。
然而,隨著政策空間的爆炸性增長,即使模型複雜性和評估成本降低到多項式依賴,挑戰依然存在。
在本文中,我們將重點從計數代理轉向計數政策,這實際上使得所謂的政策計數 DecPOMDPs 在代理數量上變得可處理。
此外,我們使用緊湊表示法提出了政策計數動態規劃,以高效解決政策計數的 DecPOMDPs。
Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See
2608.17744v1 by Ayoub Kirouane, Christos Petrocheilos
Take three frontier mixture-of-experts models (Alibaba, OpenAI, NVIDIA; 3.6-4.0B active parameters each) and fine-tune them to reason in a low-resource language. On accuracy benchmarks almost nothing happens, and the benchmark itself is noise at this scale: changing only the random seed moves the score by 7.7 points, more than every data and recipe effect we measured. That null is our first result. The real changes live where accuracy cannot see. Base models never think in Greek: 0 of 1,000 reasoning traces, even when the question is Greek, so the model answers correctly while reasoning in a form its user cannot read, audit, or correct. After supervised fine-tuning (SFT), every released checkpoint reasons in the language of the question on ~98% of items, one family at 3x fewer tokens, with judged grammaticality improving on all four models and general ability within a few points of each base: nothing was forgotten, and fluency was gained. We propose six behavioural dimensions that make such changes measurable, each gated to reject any metric that correlates with output length, and we report how our own instruments lied: six failures, each caught by a control. What SFT cannot do is fix its own defects: a quarter of answers skip the requested format, answers leak into the reasoning channel, and an explicit "think in English" is obeyed under half the time. Reinforcement learning with verifiable rewards, pre-registered before training, fixes the first two outright (fallback 24% to 2.5%, leak 3.5% to 0.0%, both against a flat random-reward control) and moves the third (+9.1pp), while the Greek reasoning habit survives an accuracy-only gradient untouched. We release five checkpoints. The instruments, the controls and the pre-registration travel to any low-resource language; Greek is the case that let us measure them.
摘要:取三個前沿的專家混合模型(Alibaba、OpenAI、NVIDIA;每個模型有 3.6-4.0B 的活躍參數)並對其進行微調,以便在低資源語言中進行推理。
在準確性基準測試中幾乎沒有任何變化,而基準本身在這個規模下是噪音:僅改變隨機種子就能使得分數變動 7.7 分,這比我們測量的每個數據和配方效果都要大。
這個無效結果是我們的第一個結果。
真正的變化存在於準確性無法看到的地方。
基礎模型從未用希臘語思考:在 1,000 條推理痕跡中,0 條是希臘語,即使問題是希臘語,因此模型在用用戶無法閱讀、審核或修正的形式進行推理時,仍然能正確回答。
在監督微調(SFT)之後,每個釋出的檢查點在約 98% 的項目中用問題的語言進行推理,某一類別的標記數量減少到 3 倍,所有四個模型的語法判斷能力均有所改善,並且一般能力在每個基礎模型的幾個點之內:沒有任何被遺忘,流利度得到了提升。
我們提出了六個行為維度,使這些變化可測量,每個維度都設置了閘門,以拒絕任何與輸出長度相關的度量,並報告我們自己的工具如何欺騙了我們:六次失敗,每次都被控制捕捉到。
SFT 無法修復其自身缺陷:四分之一的答案跳過請求的格式,答案洩漏到推理通道中,並且明確的「用英語思考」在一半的時間內未被遵守。
使用可驗證獎勵的強化學習,在訓練前進行預註冊,徹底修復了前兩個問題(回退從 24% 降至 2.5%,洩漏從 3.5% 降至 0.0%,均對比於隨機獎勵控制)並使第三個問題改善了 (+9.1pp),而希臘推理習慣在準確性僅有的梯度下仍然未受影響。
我們釋出五個檢查點。
這些工具、控制和預註冊可以應用於任何低資源語言;希臘語是讓我們能夠測量它們的案例。
Neuro-symbolic learning over OWL 2 DL via consequence-based compilation to differentiable circuits
2608.17741v1 by Olga Mashkova, Asaad Mohammedsaleh, Fernando Zhapa-Camacho, Robert Hoehndorf
OWL 2 DL ontologies, grounded in the description logic $\mathcal{SROIQ}$, express large knowledge bases in biomedicine and the Semantic Web. Neuro-symbolic (NeSy) learners over description logics either embed the ontology in a continuous space, abandoning classical entailment, or restrict to the Horn fragment $\mathcal{EL}^{++}$, which has a single canonical model. We present Baobab, which compiles a $\mathcal{SROIQ}$ ontology with a finite ABox into a Sentential Decision Diagram (SDD): it saturates a propositional core under a consequence-based calculus and instantiates the remaining $\mathcal{SROIQ}$ features (nominals, number restrictions, and the role axioms) over the active domain. The SDD's evidence-conditioned weighted model count then trains a perception network to recognize real images under partial ABox supervision: on an ontology that exercises every distinctive $\mathcal{SROIQ}$ feature, a CNN learns to read MNIST digits coupled by a successor relation and recovers latent ontology concepts that an independent perception leaves at chance. When the supervision admits several ontology-consistent completions, an independent perception collapses onto one, a reasoning shortcut: we show that a mixture indexed by the query's justifications can represent the calibrated posterior no independent perception can, and that seeding it from the circuit's enumerated completions attains the Bayes-optimal posterior on a real-image MNIST task where single-WMC and learned mixtures (the BEARS-ensemble hypothesis class) do not: to our knowledge the first to characterize and mitigate reasoning shortcuts in a non-Horn description logic. Soundness of the compiler and the representation result are machine-checked in Lean 4. Code is available at https://github.com/bio-ontology-research-group/baobab.
摘要:OWL 2 DL 本體,基於描述邏輯 $\mathcal{SROIQ}$,在生物醫學和語意網中表達大型知識庫。神經符號(NeSy)學習者在描述邏輯上要麼將本體嵌入連續空間,放棄傳統的推理,要麼限制於只有一個典範模型的 Horn 片段 $\mathcal{EL}^{++}$。我們提出了 Baobab,它將具有有限 ABox 的 $\mathcal{SROIQ}$ 本體編譯為句子決策圖(SDD):它在基於結果的計算下飽和一個命題核心,並在活動域上實例化剩餘的 $\mathcal{SROIQ}$ 特徵(名詞、數量限制和角色公理)。SDD 的證據條件加權模型計數然後訓練一個感知網絡,以在部分 ABox 監督下識別真實圖像:在一個行使每個獨特 $\mathcal{SROIQ}$ 特徵的本體上,CNN 學會閱讀與後繼關係相結合的 MNIST 數字,並恢復獨立感知所留下的潛在本體概念。當監督允許多個本體一致的完成時,獨立感知會崩潰到一個,這是一種推理捷徑:我們展示了一種由查詢的正當性索引的混合可以表示經過校準的後驗,而沒有獨立感知可以做到,並且從電路的列舉完成中種子達到在一個真實圖像 MNIST 任務上的貝葉斯最佳後驗,而單一 WMC 和學習的混合(BEARS-ensemble 假設類)則無法做到:據我們所知,這是第一次在非 Horn 描述邏輯中表徵和減輕推理捷徑。編譯器的健全性和表示結果在 Lean 4 中經過機器檢查。代碼可在 https://github.com/bio-ontology-research-group/baobab 獲得。
What Aggregate Scores Miss: Measuring Item-Level Regressions in Commercial LLM API Migrations
2608.17719v1 by Xiaonan Xu, Wenjing Wu
Context: Software systems that depend on commercial large language model APIs must migrate to successor versions when vendors deprecate older models. Migration decisions typically rely on aggregate benchmark scores, which compress heterogeneous item-level behaviour into a single net figure. Objective: We measure what that compression conceals. Method: On three pairwise upgrades in the GPT-5.4 to GPT-5.6 Sol product sequence, we query 900 public benchmark items (graduate-level knowledge, olympiad mathematics, instruction following) 50 times per item per model, classify each item as reliably improved, reliably regressed, practically equivalent, or inconclusive under false-discovery-rate control and a practical-significance threshold, and calibrate the results against a label-permutation null. Results: Across all nine migration-benchmark cells, reliable improvements and reliable regressions coexist. Edges with aggregate gains of up to 7.3 percentage points contain up to 8.3% reliably regressed items; edges with aggregate losses contain up to 10.7% reliably improved items. On the instruction-following benchmark, the gap between strict and loose scoring widens by 3.9 percentage points on the latest migration: a 3.9-point regression under strict scoring shrinks to 0.04 points under loose scoring. Conclusion: Migration decisions based on aggregate scores alone miss substantial bidirectional item-level change. The complete response-level archive and per-item scoring outputs are released.
摘要:背景:依賴商業大型語言模型 API 的軟體系統必須在供應商棄用舊模型時遷移到後繼版本。
遷移決策通常依賴於綜合基準分數,這將異質的項目級行為壓縮為單一的淨數字。
目標:我們測量這種壓縮所隱藏的內容。
方法:在 GPT-5.4 到 GPT-5.6 Sol 產品序列的三次成對升級中,我們對 900 個公共基準項目(研究生級知識、奧林匹克數學、指令遵循)進行每個模型每項 50 次查詢,並根據假發現率控制和實際顯著性閾值將每個項目分類為可靠改進、可靠退步、實際等效或不確定,並將結果與標籤置換無效進行校準。
結果:在所有九個遷移基準單元中,可靠的改進和可靠的退步共存。
具有高達 7.3 個百分點的綜合增益的邊緣包含高達 8.3% 的可靠退步項目;具有綜合損失的邊緣包含高達 10.7% 的可靠改進項目。
在指令遵循基準上,最新遷移中嚴格與寬鬆評分之間的差距擴大了 3.9 個百分點:在嚴格評分下的 3.9 點退步在寬鬆評分下縮小至 0.04 點。
結論:僅根據綜合分數作出的遷移決策忽略了實質的雙向項目級變化。
完整的響應級存檔和每項的評分輸出已發布。
Beyond Suspicious Steps: Ontological Trust in Long-Horizon Agents
2608.17718v1 by An He, Yao Wang, Haibin Zhang
Long-horizon agents increasingly operate across many steps, tools, and observa- tions. In this setting, the relevant oversight question is not only whether each action is locally valid, but whether the evolving trajectory still corresponds to the task the user authorized. Drift can accumulate quietly: an agent may call the right tool with plausible arguments at every step, while its prefix moves toward a broader role, an adjacent objective, or evidence the user never supplied. Existing monitors mostly check local compliance, deliver final-trace verdicts, or score generic risk; they do not directly estimate this prefix-level relation. We introduce ontological trust, a task-conditioned property of trajectory prefixes, and instantiate it as RGE, an online monitor that decomposes trust along Role, Goal, and Evidence. RGE uses LLMs only to derive structured task and step representations; trust-state updates, projec- tions, and intervention decisions are deterministic, so the output is a replayable and auditable trust trajectory rather than a single end-to-end judge verdict. We construct a cross-domain trajectory corpus from OSWorld, FinanceBench, and EICU-AC, covering benign executions, prefix-paired drift, and pseudo-consistency failures. On this corpus, RGE outperforms adapted rule-, judge-, and shield-style baselines on prefix-paired drift detection. With the two larger estimator models, it exceeds 93% Drift F1 on every benchmark while keeping benign coverage at or above 95.8%. Pseudo-consistency is harder: detection depends on whether task completion is externally visible, a structural limit we characterize empirically.
摘要:長期代理人越來越多地在多個步驟、工具和觀察中運作。在這種情況下,相關的監督問題不僅是每個行動是否在局部有效,而是演變的軌跡是否仍然符合用戶授權的任務。漂移可能悄然積累:一個代理人可能在每一步都以合理的論據調用正確的工具,而其前綴卻朝著更廣泛的角色、相鄰的目標或用戶從未提供的證據移動。現有的監控器主要檢查局部合規性,提供最終的追蹤判決,或評分一般風險;它們並不直接估計這種前綴級別的關係。我們引入本體信任,這是一種基於任務的軌跡前綴特性,並將其具體化為RGE,一種在線監控器,沿著角色、目標和證據分解信任。RGE僅使用LLMs來推導結構化的任務和步驟表示;信任狀態更新、投影和干預決策是確定性的,因此輸出是一個可重播和可審計的信任軌跡,而不是單一的端到端判決。 我們從OSWorld、FinanceBench和EICU-AC構建了一個跨領域的軌跡語料庫,涵蓋良性執行、前綴配對漂移和偽一致性失敗。在這個語料庫上,RGE在前綴配對漂移檢測上超越了適應的規則、判決和屏障風格基準。使用兩個更大的估計模型,它在每個基準上都超過93%的漂移F1,同時保持良性覆蓋率在95.8%或以上。偽一致性更難:檢測取決於任務完成是否在外部可見,這是一個我們經驗性描述的結構性限制。
Communicating Credit Risk with Large Language Models: Evaluation of Explanations from Standard and Alternative Data-Based Models
2608.17715v1 by Sahab Zandi, Noah Kostesku, Christophe Mues, María Óskarsdóttir, Cristián Bravo
Credit decisioning is a high-stakes task in which model outputs must be accurate and explainable to support compliant decisions. Although modern credit risk models such as eXtreme Gradient Boosting (XGBoost) and Graph Neural Networks (GNNs) improve predictive performance, their explanations are often too technical for stakeholders creating communication gaps that can shape approvals, denials, and fairness judgments. We examine whether Large Language Models (LLMs) can serve as explanation layers that translate post-hoc explanation artefacts into stakeholder-appropriate risk narratives. Using Freddie Mac single-family loan-level data, we develop three pipelines: standard tabular (XGBoost + SHAP), and two with alternative data, a pure network-based (GNN + GNNExplainer), and a bimodal one (combining tabular and network data). We generate narratives with three LLM configurations: a small fine-tuned LLM (Gemma 3 4B), a large fine-tuned LLM (DeepSeek R1 70B), and a zero-shot commercial LLM (Gemini 2.5). Explanation quality is evaluated through automated checks across all pipelines and a human study of bimodal explanations comparing credit risk professionals and non-professionals on eight decision-relevant dimensions. We have three main findings. First, the pipeline accounts for higher variance in evidence-grounding scores than the language model, meaning that the binding constraint on explanation quality is the evidence representation, not the model used. Second, the explanation narratives reliably name the influential factors but are less reliable when stating the direction of influence, which may be consequential for adverse-action communication. Finally, professionals apply stricter evidentiary standards than non-professionals. We discuss implications for the governance of risk models, including deployment considerations and the value of domain-aligned LLMs in regulated credit settings.
摘要:信用決策是一項高風險的任務,其中模型輸出必須準確且可解釋,以支持合規的決策。儘管現代信用風險模型如極端梯度提升(XGBoost)和圖神經網絡(GNNs)提高了預測性能,但它們的解釋往往對利益相關者來說過於技術性,造成溝通差距,這可能影響批准、拒絕和公平性判斷。我們檢視大型語言模型(LLMs)是否可以作為解釋層,將事後解釋產物轉化為適合利益相關者的風險敘事。使用Freddie Mac的單戶貸款數據,我們開發了三個管道:標準表格(XGBoost + SHAP),以及兩個使用替代數據的管道,一個是純基於網絡的(GNN + GNNExplainer),另一個是雙模的(結合表格和網絡數據)。我們使用三種LLM配置生成敘事:一個小型微調LLM(Gemma 3 4B),一個大型微調LLM(DeepSeek R1 70B),以及一個零樣本商業LLM(Gemini 2.5)。通過對所有管道的自動檢查以及對雙模解釋的人工研究,我們評估了解釋質量,並比較了信用風險專業人員和非專業人員在八個與決策相關的維度上的表現。我們有三個主要發現。首先,該管道在證據基礎分數的變異性上比語言模型更高,這意味著解釋質量的約束是證據表示,而不是所使用的模型。其次,解釋敘事可靠地命名了影響因素,但在陳述影響方向時可靠性較低,這對於不利行動的溝通可能具有重要意義。最後,專業人士應用的證據標準比非專業人士更為嚴格。我們討論了風險模型治理的影響,包括部署考量和在受監管的信用環境中領域對齊的LLMs的價值。
Accuracy and Robustness of Model Cascades Under Data Perturbations
2608.17711v1 by Pallavi Mitra, Jai Kushwaha, Felix Biessmann
Prediction cascades significantly reduce energy consumption of Artificial Intelligence (AI) models while maintaining high predictive performance. The idea is that easy inputs are routed through a lightweight small model, and difficult uncertain cases are deferred to a larger model. While this design can improve computational efficiency on clean data, its effectiveness depends on the reliability of confidence-based routing. Input degradations, such as static corruptions and sequential perturbations, can shift model confidence and routing decisions. In this paper, we study confidence-based cascade frameworks for image classification and investigate how such degradations affect their confidence-based deferral behavior. We select a model cascade at the pareto-optimum of accuracy, routing quality, and energy consumption that achieves competitive predictive performance with an up to 10-fold decrease in CO$_2$ emissions. We study the behavior of that model cascade under input corruptions and analyze how the cascade's routing decisions change when the input distribution shifts. Our analysis identifies three failure modes. Static corruptions either (1) break the routing signal while the large model remains useful, or (2) degrade both models so deferral no longer recovers accuracy. Sequential perturbations reveal a third mode: predictions stabilize but deferral suppresses, yielding stable but unreliable predictions. These findings demonstrate that energy efficient model cascades require evaluation beyond clean accuracy, with explicit attention to routing reliability under distribution shift.
摘要:預測級聯顯著降低人工智慧(AI)模型的能耗,同時保持高預測性能。其理念是將簡單的輸入通過一個輕量的小模型處理,而將困難的不確定案例延遲到更大的模型中。雖然這種設計可以在乾淨數據上提高計算效率,但其有效性取決於基於信心的路由的可靠性。輸入退化,例如靜態損壞和序列擾動,可能會改變模型的信心和路由決策。在本文中,我們研究了用於圖像分類的基於信心的級聯框架,並探討這些退化如何影響其基於信心的延遲行為。我們選擇了一個在準確性、路由質量和能耗的帕累托最優解上的模型級聯,該級聯在CO$_2$排放量上實現了高達10倍的減少,同時達到競爭性的預測性能。我們研究了該模型級聯在輸入損壞下的行為,並分析當輸入分佈發生變化時,級聯的路由決策如何改變。我們的分析確定了三種失效模式。靜態損壞要麼(1)破壞路由信號,而大型模型仍然有用,要麼(2)使兩個模型都退化,導致延遲不再恢復準確性。序列擾動揭示了第三種模式:預測穩定,但延遲被抑制,產生穩定但不可靠的預測。這些發現表明,能源高效的模型級聯需要在乾淨準確性之外進行評估,並明確關注在分佈變化下的路由可靠性。
GADR: Gathering Architecture Decision Records from Meeting Transcriptions
2608.17694v1 by Lucas Daniel Costa da Silva, Kiev Gama
Existing LLM-based approaches to Architecture Decision Record (ADR) generation share a critical and largely unexamined assumption: that input is already reasonably structured. In practice, architectural decisions emerge from informal, noisy meetings where choices are implicit, fragmented, and entangled with off-topic dialogue, precisely the conditions under which single-pass prompting degrades. This paper presents GADR, a multi-agent, self-correcting workflow that extracts architectural decisions from raw meeting transcriptions and generates Nygard-formatted ADR drafts. A feasibility study comprising five real project meeting transcripts, expert review by four senior architects, and evaluation by fifteen students provides initial evidence that the agentic workflow captures most expert-identified decisions and produces drafts participants found clear and useful, outperforming zero-shot and few-shot baselines in stability and structural adherence. The study also addresses the underexplored trade-off of RAG-based enrichment improving ADR depth while simultaneously risking transcript-unfaithful content, raising open questions about traceability in automated architectural documentation that we believe is worth the community's attention.
摘要:現有基於LLM的方法在架構決策記錄(ADR)生成中共享一個關鍵且大多未經檢驗的假設:即輸入已經相當結構化。實際上,架構決策源自非正式、嘈雜的會議,其中選擇是隱含的、零散的,並與無關的對話交織在一起,這正是單次提示退化的條件。本文提出了GADR,一種多代理、自我校正的工作流程,從原始會議記錄中提取架構決策並生成Nygard格式的ADR草稿。一項包含五個真實項目會議記錄的可行性研究、四位資深建築師的專家評審,以及十五名學生的評估提供了初步證據,表明該代理工作流程捕捉了大多數專家識別的決策,並生成了參與者認為清晰且有用的草稿,在穩定性和結構遵循性方面超越了零-shot和少-shot基準。該研究還探討了基於RAG的豐富性改善ADR深度的未充分探索的權衡,同時冒著轉錄不忠實內容的風險,提出了關於自動化架構文檔中可追溯性的開放問題,我們認為這值得社群的關注。
Mixture-of-Expert Blocks Contain Strong Hallucination Detection Signals
2608.17687v1 by Joao Fonseca, Rodrigo Rodrigues, Paolo Romano
Despite their widespread use, Large Language Models (LLMs) remain limited by a fundamental problem: the generation of plausible but false content, known as hallucinations. Most existing detection methods operate at the answer or sentence level, yet per-token detection is essential for localizing hallucinated spans and enabling fine-grained interventions. In this paper, we explore the use of the Mixture-of-Experts (MoE) paradigm to address this gap. In MoE architectures, a single forward pass activates a sparse subset of experts (i.e., distinct feedforward networks per layer) via a routing mechanism, producing internal signals (e.g., router entropy, expert disagreement, and expert usage patterns) that are unavailable in dense architectures and have not been previously exploited for hallucination detection. To this end, we introduce InnerExpert, the first method to leverage these MoE-specific signals for per-token hallucination detection. InnerExpert combines routing-level and standard transformer signals into compact per-token feature vectors, classified by a lightweight detector trained on labels produced by an LLM-as-a-judge pipeline, which enables continuous model updates without manual annotation. Our results show that InnerExpert outperforms existing methods across five datasets and two MoE architectures, achieving up to 0.91 answer-level and 0.76 token-level AUROC, while requiring only a single forward pass.
摘要:儘管大型語言模型(LLMs)被廣泛使用,但仍然受到一個根本問題的限制:生成看似合理但實際上錯誤的內容,稱為幻覺。大多數現有的檢測方法在答案或句子層面運作,然而,逐字檢測對於定位幻覺範圍和實現精細干預至關重要。在本文中,我們探討使用專家混合(Mixture-of-Experts, MoE)範式來解決這一差距。在MoE架構中,單次前向傳播通過路由機制激活稀疏的專家子集(即每層的不同前饋網絡),產生內部信號(例如,路由熵、專家不一致性和專家使用模式),這些信號在密集架構中不可用,且未被用於幻覺檢測。為此,我們介紹了InnerExpert,這是第一種利用這些MoE特定信號進行逐字幻覺檢測的方法。InnerExpert將路由層級和標準Transformer信號結合成緊湊的逐字特徵向量,並由一個輕量級檢測器進行分類,該檢測器在由LLM作為裁判管道產生的標籤上進行訓練,這使得模型可以在不需要人工標註的情況下進行持續更新。我們的結果顯示,InnerExpert在五個數據集和兩個MoE架構中超越了現有方法,達到了高達0.91的答案級別和0.76的逐字級別AUROC,同時僅需一次前向傳播。
Benchmarking Automated Security Patch Backporting: How Far Are We?
2608.17671v1 by Jincheng Yang, Yulong Fu, Chengwei Liu, Lyuye Zhang, Fangyuan Zhang, Bingyang Ren, Yang Liu, Hui Li
Automated security patch backporting is critical for mitigating N-day vulnerabilities. Recent tools report success rates above 80% on their respective datasets. However, these evaluations are often confined to homogeneous environments, such as one repository or specific project versions. Consequently, it remains unclear how well these tools generalize beyond their originally targeted scenarios. We present Porting Benchmark, a curated dataset of 1,234 security patch backporting cases spanning cross-version, cross-branch, and cross-repository scenarios, paired with a common evaluation framework. Using this benchmark, we evaluate five tools spanning program analysis, LLM prompting, and LLM agents under aligned settings. Our results show that aligned evaluation changes the apparent performance landscape: PortGPT and TSBPort remain comparatively strong on the Replication Dataset, while FixMorph and Mystique degrade substantially under the common protocol. Performance degrades sharply on structurally complex patches: the best commit-level success rate falls from 85.2% on Type-I patches to 24.0% on Type-IV. We identify four root-cause categories (missing target API awareness, cross-version semantic mismatch, non-local dependency propagation failure, and patch construction or localization failure) and derive concrete directions for next-generation tool design. On a 45-case dynamically validated subset with verified test cases and constructed POCs, we further observe that reference-based benchmark scores do not fully capture real-world remediation: exact match sharply under-credits harder target adaptations, while executable validation reveals residual integration failures in the target that static reference agreement misses. Executable-feedback refinement provides limited but measurable recovery on the hardest executable cases.
摘要:自動化安全補丁回溯對於減輕N天漏洞至關重要。最近的工具在其各自的數據集上報告的成功率超過80%。然而,這些評估通常僅限於同質環境,例如單一代碼庫或特定項目版本。因此,這些工具在其最初目標場景之外的普遍性仍不明確。我們提出了Porting Benchmark,這是一個經過精心策劃的數據集,包含1,234個安全補丁回溯案例,涵蓋跨版本、跨分支和跨代碼庫的場景,並配有一個共同的評估框架。利用這個基準,我們評估了五種工具,涵蓋程序分析、LLM提示和LLM代理,在對齊的設置下進行測試。我們的結果顯示,對齊評估改變了表面上的性能格局:PortGPT和TSBPort在複製數據集上仍然相對強勁,而FixMorph和Mystique在共同協議下顯著降級。對於結構複雜的補丁,性能急劇下降:最佳提交級成功率從Type-I補丁的85.2%降至Type-IV的24.0%。我們確定了四個根本原因類別(缺乏目標API認知、跨版本語義不匹配、非本地依賴傳播失敗,以及補丁構建或本地化失敗),並為下一代工具設計提供了具體方向。在一個包含經過驗證的測試案例和構建的POC的45案例動態驗證子集中,我們進一步觀察到基於參考的基準分數並未完全捕捉到現實世界的修復:精確匹配明顯低估了更難的目標適應,而可執行驗證揭示了靜態參考一致性所忽略的目標中的殘餘集成失敗。可執行反饋精煉在最難的可執行案例上提供了有限但可測量的恢復。
GraphWake: Group Polarization via Memory-Mediated Polarization Cascade in LLM-Agent Communities
2608.17665v1 by Haoran Bu, Zejian Chen, Litian Zhang, Xi Zhang
LLM-driven agents can autonomously exchange opinions on online platforms and form communities. Such agent-operated social platforms raise a new security concern: attackers may manipulate agents to induce group polarization. Existing methods manipulate agent prompts or construct echo chambers, both of which are difficult to realize in practice. We therefore formulate a new threat, Memory-Mediated Polarization Cascade, which uses agent memory as a persistence channel and public discussion as a propagation channel. This threat contains three stages. During exposure and memory retention, the attacker exposes a small set of target agents to arguments that reinforce their respective stated stances. The targets' memory systems then process and retain these arguments. During retrieval and reproduction, a shared stance-neutral discussion cues the targets to retrieve and reproduce their respective retained arguments. During iterative propagation, untreated agents influenced by the reproduced arguments restate and spread them. We instantiate this threat in GraphWake with three components: (i) stance-support argumentation knowledge graphs construct knowledge-based arguments; (ii) axiom-oriented triple selection distills them for reliable retention and reproduction; and (iii) stance-neutral memory cueing triggers concurrent retrieval and reproduction, initiating propagation. Experiments across multiple discussions and memory systems show that GraphWake substantially increases group polarization. These findings reveal a community-level polarization risk.
摘要:LLM 驅動的代理可以在在線平台上自主交換意見並形成社群。這種代理操作的社交平台引發了一個新的安全問題:攻擊者可能操縱代理以誘發群體極化。現有的方法操縱代理提示或構建回音室,這兩者在實踐中都難以實現。因此,我們提出了一種新的威脅,記憶介導的極化級聯,它利用代理記憶作為持久性通道,公共討論作為傳播通道。這一威脅包含三個階段。在暴露和記憶保留期間,攻擊者將一小組目標代理暴露於強化其各自表述立場的論點中。目標的記憶系統隨後處理並保留這些論點。在檢索和再現期間,共享的中立立場討論提示目標檢索並再現其各自保留的論點。在迭代傳播期間,受到再現論點影響的未處理代理重述並擴散這些論點。我們在 GraphWake 中實現了這一威脅,包含三個組件:(i)立場支持的論證知識圖構建基於知識的論點;(ii)公理導向的三元組選擇提煉它們以實現可靠的保留和再現;以及(iii)立場中立的記憶提示觸發同時檢索和再現,啟動傳播。多次討論和記憶系統的實驗顯示,GraphWake 顯著增加了群體極化。這些發現揭示了社群層面的極化風險。
MobileWorldSafety: Benchmarking GUI Agent Safety Against Environmental Injection Attacks in Android Apps
2608.17659v1 by Sujin Chen, Lijun Li, Tianyi Du, Jing Shao
LLM-powered GUI agents that autonomously operate smartphones are rapidly transitioning from research prototypes to early real-world deployment. However, because these agents routinely process untrusted environmental content, they are highly vulnerable to environmental injection attacks, which include indirect prompt injections and adversarial instructions. Such attacks can manipulate the behavior of agents without user awareness through diverse channels encountered in everyday mobile use. Despite these risks, existing benchmarks often fail to capture everyday user scenarios, lacking a systematic evaluation of GUI agents under environmental injection attacks on mobile devices. To address this gap, we introduce MobileWorldSafety, a benchmark of 142 risk tasks built on real Android applications. For each task, we define a programmatically verifiable risk indicator over the final system state and evaluate outcomes with a two-stage pipeline: rule-based verification handles unambiguous cases, while an LLM judge adjudicates ambiguous ones. This distinguishes safety failures from capability failures and enables objective and reproducible assessment. Evaluations on six agents, including both general agents and specialized GUI agents, demonstrate that all agents remain highly vulnerable, with attack success rates ranging from 40.4% to 66.9%. These findings indicate that current agents often fail to maintain safety alignment when adversarial content is presented as ordinary mobile context. MobileWorldSafety provides a foundation for quantifying these vulnerabilities and advancing research on robust mobile GUI agents.
摘要:LLM 驅動的 GUI 代理自動操作智能手機,正在迅速從研究原型轉向早期的實際部署。
然而,由於這些代理經常處理不受信任的環境內容,它們對環境注入攻擊高度脆弱,這些攻擊包括間接提示注入和對抗性指令。
這些攻擊可以通過日常移動使用中遇到的多樣渠道,在不讓用戶察覺的情況下操控代理的行為。
儘管存在這些風險,現有基準往往未能捕捉到日常用戶場景,缺乏對移動設備上環境注入攻擊下 GUI 代理的系統評估。
為了填補這一空白,我們推出了 MobileWorldSafety,一個基於真實 Android 應用的 142 個風險任務的基準。
對於每個任務,我們定義了一個可程序驗證的風險指標,基於最終系統狀態進行評估,並通過兩階段的流程來評估結果:基於規則的驗證處理明確的情況,而 LLM 評判則裁決模糊的情況。
這區分了安全失敗和能力失敗,並使客觀和可重複的評估成為可能。
對六個代理的評估,包括一般代理和專門的 GUI 代理,顯示所有代理仍然高度脆弱,攻擊成功率範圍從 40.4% 到 66.9%。
這些發現表明,當對抗性內容被呈現為普通的移動上下文時,當前的代理往往無法保持安全對齊。
MobileWorldSafety 為量化這些脆弱性和推進穩健的移動 GUI 代理研究提供了基礎。
LLM-Derived Preference Judgments Are Not Self-Consistent
2608.17644v1 by Matthew T. Ford, Francis Bahk, Jingjing Wang, Adam S. Jovine, Tinghan Ye, David B. Shmoys, Peter I. Frazier
Agents increasingly interpret a person's natural-language preferences by querying an LLM for numerical preference judgments, e.g., by asking how much the person would be willing to pay for an item. A growing body of work estimates a utility function from these judgments and then chooses actions based on their estimated utility. This pipeline assumes the judgments are approximately self-consistent: that a single utility function can reproduce them. But are they? To study this question, we measure the self-consistency of cardinal LLM preference judgments. For example, the difference in stated willingness-to-pay between two items should match the stated payment that makes a person indifferent to exchanging them. We develop statistical tests and interpretable measures of how far observed responses depart from the best-fitting self-consistent utility function. Experiments with flight, apartment, and hotel examples across six LLMs reveal large persistent inconsistencies. This suggests that LLM-derived preference judgments cannot be faithfully summarized by a single utility function.
摘要:代理人越來越多地通過查詢 LLM 來解釋一個人的自然語言偏好,以獲得數值偏好判斷,例如,詢問這個人願意為某個項目支付多少。越來越多的研究從這些判斷中估計效用函數,然後根據其估計的效用選擇行動。這個流程假設這些判斷大致上是自我一致的:即單一的效用函數可以重現它們。但真的是這樣嗎?為了研究這個問題,我們測量了基數 LLM 偏好判斷的自我一致性。例如,兩個項目之間所表明的支付意願差異應該與使一個人對交換它們無所謂的所表明的支付相匹配。我們開發了統計檢驗和可解釋的度量,來衡量觀察到的反應與最佳擬合的自我一致效用函數之間的偏離程度。對六個 LLM 的航班、公寓和酒店範例進行的實驗揭示了持久的巨大不一致性。這表明 LLM 衍生的偏好判斷不能被單一的效用函數忠實地總結。
Beyond the Trace: Coupling an Interpretable Reasoning-State Readout to Native MoE Routing
2608.17638v1 by Kang Chen, Sihan Zhao, Yixin Cao, Yugang Jiang
What a reasoning model writes is only a partial record of the process that produces it. We introduce a two-level internal readout for mixture-of-experts reasoning. We first distill vocabulary-scale J-space into J64, a 64-axis semantic frame learned from the model's own reasoning states. J64 reveals readable process state that the emitted trace does not show: it separates inference effort from problem-induced strain. It also adds 0.096 to 0.135 held-out AUC over a baseline that reads the same rollout as token occupancy and aggregates it in exactly the same way. We then reconstruct J64 from native expert-routing statistics. The result is R64, a low-overhead proxy: its median per-axis correlation with J64 is 0.69 to 0.86 across three models and two families, and on gpt-oss-20b it preserves 95 to 100% of J64's predictive gain. The readout supports test-time decisions at two temporal resolutions. Over completed candidate sets, J64 and R64 improve single-branch selection, and R64-weighted voting improves plain majority voting in seven of eight settings. During generation, rolling readout windows drive a cumulative stop-and-resample policy whose operating point is fixed on training questions alone. J64 improves accuracy by 1.1 to 5.9 points over a sibling-permuted control, and the routing-only R64 proxy retains 0.9 to 3.2 of those points. Finally, router edits aimed at the mechanism J64 names induce the predicted reasoning behaviors and shift a diagnosed stall from numerical guessing toward exact symbolic execution. Together, J64 makes latent process state readable, while routing makes it deployable and actionable.
摘要:推理模型所寫的內容僅是產生該內容過程的部分記錄。
我們為混合專家推理引入了兩級內部讀出。
我們首先將詞彙規模的 J 空間提煉成 J64,這是一個從模型自身推理狀態學習而來的 64 軸語義框架。
J64 揭示了可讀的過程狀態,而發出的痕跡並未顯示出來:它將推理努力與問題引起的壓力分開。
它還在基準上增加了 0.096 到 0.135 的持出 AUC,該基準將同樣的展開視為標記佔用並以完全相同的方式進行聚合。
然後,我們從原生專家路由統計中重建 J64。
結果是 R64,一個低開銷的代理:它與 J64 的每軸中位數相關性在三個模型和兩個家族中為 0.69 到 0.86,而在 gpt-oss-20b 上,它保留了 J64 預測增益的 95% 到 100%。
這個讀出支持在兩個時間解析度下的測試時決策。
在完成的候選集上,J64 和 R64 改進了單分支選擇,而 R64 加權投票在八個設置中的七個中改善了普通多數投票。
在生成過程中,滾動讀出窗口驅動一個累積的停止和重取樣策略,其操作點僅固定在訓練問題上。
J64 在與兄弟置換控制相比中提高了 1.1 到 5.9 分的準確性,而僅路由的 R64 代理保留了 0.9 到 3.2 的這些分數。
最後,針對 J64 所命名的機制的路由編輯引發了預測的推理行為,並將診斷出的停滯從數值猜測轉向精確的符號執行。
總之,J64 使潛在的過程狀態可讀,而路由則使其可部署和可行動。
Graph Surgery and the Do-Operator: A Precise Correspondence for Acyclic Structural Causal Models
2608.17634v1 by Satpreet Makhija
The $\operatorname{do}$-operator is described graphically by deleting arrows into its targets and functionally by replacing their mechanisms with constants. To call these operations equivalent is not yet a mathematical statement: one returns a graph and remembers only the targets, whereas the other returns mechanisms and also remembers the imposed values. We make a dependency-level comparison precise for deterministic acyclic structural causal models with finitely many endogenous variables. If $\operatorname{Graph}(F)$ extracts the dependencies of a mechanism family $F$, our main theorem is $\operatorname{Graph}(F^ι)=\operatorname{Surg}(\operatorname{Graph}(F),T_ι)$. Thus replacing target mechanisms removes exactly the dependencies removed by graph surgery. For a model $M=(G,F)$ whose graph may contain unused arrows, we characterize when the same equality holds with $G$ in place of $\operatorname{Graph}(F)$; it holds for every intervention exactly when $G$ records the dependencies of $F$ exactly. We then define the intervened model, characterize its run, show how sequential interventions combine, and prove that an outcome depends only on interventions at its actual dependency ancestors.
摘要:$\operatorname{do}$-運算子在圖形上通過刪除指向其目標的箭頭來描述,而在功能上則通過用常數替換其機制來描述。將這些操作稱為等價尚未形成數學陳述:一個返回圖形並僅記住目標,而另一個返回機制並同時記住施加的值。我們對具有有限內生變量的確定性非循環結構因果模型進行依賴層級的精確比較。如果 $\operatorname{Graph}(F)$ 提取機制家族 $F$ 的依賴關係,我們的主要定理是 $\operatorname{Graph}(F^ι)=\operatorname{Surg}(\operatorname{Graph}(F),T_ι)$。因此,替換目標機制正好去除了圖形手術所去除的依賴關係。對於一個模型 $M=(G,F)$,其圖形可能包含未使用的箭頭,我們描述何時同樣的等式在 $G$ 代替 $\operatorname{Graph}(F)$ 時成立;當且僅當 $G$ 精確記錄 $F$ 的依賴關係時,它成立。我們然後定義干預模型,描述其運行,展示如何結合序列干預,並證明結果僅依賴於其實際依賴祖先的干預。
DEPT: Document Embedding Preservation Tuning for Unified Query Expansion and Retrieval
2608.17632v1 by Jingyuan Wang, Richong Zhang, Zhijie Nie, Mingxin Li, Yanzhao Zhang
Large language models (LLMs) can both expand underspecified queries and encode text as dense representations, suggesting a unified model for query expansion and retrieval. Existing systems usually rely on prompted expansions, independently trained modules, or staged optimization, leaving generated expansions only indirectly aligned with the retrieval loss that judges them. We train a single decoder-only LLM end to end, where the same model generates the expansion and encodes both the expanded query and candidate documents. This unified setting creates a moving-target problem: retrieval supervision should improve query-side expansion, but the same update also shifts the document embeddings that serve as retrieval targets. We introduce Document Embedding Preservation Tuning (DEPT), which keeps tuned document embeddings close to cached initial embeddings while allowing retrieval gradients to pass through straight-through decoding into the generator. DEPT converts joint query--document movement into query-side adaptation against approximately stable, whitened document embeddings that support index reuse and online hard-negative mining. Experiments with Qwen3-4B-Instruct-2507 and LLaMA-3.2-3B-Instruct on five datasets in BEIR benchmark show that DEPT improves average retrieval quality over training-free, independently trained, and staged unified baselines, while ablations isolate the effects of preservation, whitening, end-to-end expansion training, and online negatives. Code is available at https://github.com/ILSparkle/DEPT.
摘要:大型語言模型(LLMs)可以擴展未具體化的查詢並將文本編碼為密集表示,這表明查詢擴展和檢索的統一模型。現有系統通常依賴於提示擴展、獨立訓練的模塊或分階段優化,這使得生成的擴展與評估它們的檢索損失僅間接對齊。我們訓練了一個端到端的單解碼器 LLM,該模型同時生成擴展並編碼擴展查詢和候選文檔。這種統一的設置創造了一個移動目標問題:檢索監督應該改善查詢端的擴展,但相同的更新也會改變作為檢索目標的文檔嵌入。我們引入了文檔嵌入保護調整(DEPT),它保持調整後的文檔嵌入接近緩存的初始嵌入,同時允許檢索梯度通過直通解碼進入生成器。DEPT 將聯合查詢-文檔移動轉換為針對大致穩定的、經過去白化的文檔嵌入的查詢端適應,這些嵌入支持索引重用和在線困難負樣本挖掘。在 BEIR 基準的五個數據集上,使用 Qwen3-4B-Instruct-2507 和 LLaMA-3.2-3B-Instruct 的實驗顯示,DEPT 提高了平均檢索質量,超過了無需訓練的、獨立訓練的和分階段的統一基準,而消融實驗則隔離了保護、去白化、端到端擴展訓練和在線負樣本的影響。代碼可在 https://github.com/ILSparkle/DEPT 獲得。
From Student Risk Prediction to SC2R: Semantics-Constrained Counterfactual Recourse for Educational Decision Support
2608.17618v1 by Ngoc Luyen Le, Marie-Hélène Abel, Bertrand Laforge
Learning analytics models can identify students at risk of poor performance, but they do not directly indicate which interventions are feasible, actionable, and compatible with educational constraints. This paper introduces SC2R, a semantics-constrained counterfactual recourse framework for educational decision support. SC2R combines a calibrated predictive model, integer-programming-based recourse generation over discrete action variables, a lightweight RDF vocabulary for intervention-plan representation, and SHACL validation for enforcing timing, budget, immutability, and availability constraints. The framework is evaluated offline on the OULAD dataset using snapshots constructed relative to each assessment at two decision horizons. Results show that the predictive component provides strong performance, that compact intervention plans can be generated at scale, and that semantic validation reveals infeasible plans that lighter optimization-only settings would otherwise accept. Rather than claiming causal improvement in student outcomes, this work shows that counterfactual recourse becomes more operationally meaningful in education when recommendations are not only model-valid, but also semantically feasible and machine-checkable.
摘要:學習分析模型可以識別出有表現不佳風險的學生,但它們並不直接指示哪些干預措施是可行的、可操作的,以及與教育限制相容的。本文介紹了SC2R,一個語義約束的反事實補救框架,用於教育決策支持。SC2R結合了一個經過校準的預測模型、基於整數規劃的離散行動變數的補救生成、輕量級的RDF詞彙用於干預計劃表示,以及SHACL驗證以強制執行時間、預算、不變性和可用性約束。該框架在OULAD數據集上進行了離線評估,使用相對於每次評估在兩個決策視野下構建的快照。結果顯示,預測組件提供了強大的性能,能夠大規模生成緊湊的干預計劃,並且語義驗證揭示了在僅進行輕量優化的設置下會被接受的不可行計劃。本研究並不聲稱對學生結果的因果改善,而是顯示當推薦不僅是模型有效的,還是語義上可行且可機器檢查的時,反事實補救在教育中變得更具操作意義。
Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges
2608.17605v1 by Syeda Faiza Ahmed, Zien Sheikh Ali, Hunzalah Hassan Bhatti, Firoj Alam, Shammur Absar Chowdhury
Conversational AI is moving beyond isolated text prompts toward sustained, multimodal interaction. In real conversations, users clarify goals, revise requests, interrupt responses, switch topics, and introduce new evidence while expecting systems to preserve context across turns. This makes multi-turn dialogue a distinct challenge requiring systems to maintain and update memory, ground responses across modalities, tools, and external knowledge, and adapt across languages and cultures. This study reviews multi-turn conversational AI across text-only dialogue, AudioLLMs and speech-native systems, multimodal and omni-modal systems, and tool-augmented agents. We organize the literature around datasets and benchmarks, modeling paradigms, training strategies, evaluation setups, and cross-cutting challenges. Our analysis shows that support for multiple modalities has advanced faster than the ability to sustain coherent interaction across a session. Despite stronger capabilities to perceive, speak, and act across modalities, current systems still struggle with persistent memory, cross-turn grounding, full-duplex interaction, robust evaluation, and cultural alignment. We conclude with a research agenda for systems that can remember, revise, ground, speak, listen, act, and adapt across turns, modalities, and cultures. (https://github.com/faiza-sfa/multiturn-conversational-ai-survey)
摘要:對話式人工智慧正在超越孤立的文字提示,朝向持續的多模態互動發展。在真實的對話中,用戶會澄清目標、修訂請求、打斷回應、切換主題並引入新證據,同時期望系統能在不同回合中保持上下文。這使得多回合對話成為一個獨特的挑戰,要求系統維持和更新記憶,跨模態、工具和外部知識進行回應的基礎,並在語言和文化之間進行適應。本研究回顧了文本對話、AudioLLMs 和語音原生系統、多模態和全模態系統以及工具增強代理的多回合對話式人工智慧。我們根據數據集和基準、建模範式、訓練策略、評估設置和跨領域挑戰來組織文獻。我們的分析顯示,對多模態的支持發展得比在一個會話中持續一致互動的能力更快。儘管在感知、說話和跨模態行動方面的能力增強,當前的系統仍然在持久記憶、跨回合基礎、全雙工互動、穩健評估和文化對齊方面面臨挑戰。我們以一個研究議程作結,旨在開發能夠記住、修訂、基礎、說話、聆聽、行動和在回合、模態和文化之間適應的系統。(https://github.com/faiza-sfa/multiturn-conversational-ai-survey)
HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety
2608.17597v1 by Yajing Bai, Jinhao Duan, Jie Peng, Xianfeng Wu, Sijia Liu, Song Wang, Tianlong Chen
Large language models are increasingly deployed through agent harnesses that manage tools, extensions, persistent state, permissions, and external actions. Existing safety benchmarks mainly target individual attack mechanisms or a limited subset of operational settings, making it difficult to compare how safety failures emerge across different harness responsibilities. We present HarnessRisk, a lifecycle oriented benchmark that organizes agent harness safety into six operational phases including Harness Configuration, Capability Extension, Runtime Operation, State Persistence, Action Control, and Incident Recovery. HarnessRisk contains 128 sandboxed cases, each pairing a benign user objective with an adversarial instruction embedded in an untrusted workflow artifact. We evaluate each trajectory using Utility, Attack Success Rate, Persistence, and Detection. Across three harnesses, six language models, and 14 model and harness configurations, attack success ranges from 12.6% to 80.9%, while Utility remains between 75.0% and 97.6%. Harness Configuration is the most vulnerable phase across all three harnesses, showing that attacks can succeed by altering security sensitive parameters within otherwise authorized workflows. We also find that explicit risk recognition does not reliably lead to safe action, as some configurations detect risks in more than 90% of runs while retaining substantial attack success. These results highlight the need to evaluate agent safety across multiple harness responsibilities and at the level of the deployed model and harness configuration.
摘要:大型語言模型越來越多地透過代理工具來管理工具、擴展、持久狀態、權限和外部行動。現有的安全基準主要針對個別攻擊機制或有限的操作設定,使得比較不同工具責任下安全失敗的出現變得困難。我們提出了HarnessRisk,一個以生命週期為導向的基準,將代理工具的安全性組織為六個操作階段,包括工具配置、能力擴展、運行時操作、狀態持久性、行動控制和事件恢復。HarnessRisk包含128個沙盒案例,每個案例將一個良性的用戶目標與嵌入在不受信任的工作流程工件中的對抗指令配對。我們使用效用、攻擊成功率、持久性和檢測來評估每個軌跡。在三個工具、六個語言模型和14個模型及工具配置中,攻擊成功率範圍從12.6%到80.9%,而效用則保持在75.0%到97.6%之間。工具配置是所有三個工具中最脆弱的階段,顯示攻擊可以通過改變在其他授權工作流程中安全敏感的參數而成功。我們還發現,明確的風險識別並不可靠地導致安全行動,因為某些配置在超過90%的運行中檢測到風險,但仍然保留了相當大的攻擊成功率。這些結果突顯了需要在多個工具責任和部署的模型及工具配置層面上評估代理安全性。
tinyDSM: A Framework for Skill Modeling and Development for Resource-Constrained Millirobots
2608.17596v1 by Markus D. Kobelrausch, Michael Miedler, Axel Jantsch
In this study, we investigate developmental mechanisms that enable small, resource-constrained systems such as cm-sized millirobots to autonomously explore, learn, and adapt their capabilities throughout their lifespan. Reinforcement learning algorithms guide the agent's skill acquisition and adaptation through the interplay of our proposed tinyDSM, which integrates intrinsic motivation and fitness-based assessment. We strive for minimal, hard-wired skills while encouraging the open-ended development of new skills. A key emphasis in our approach is to encode minimal a-priori general knowledge, which serves as a foundational starting point for the system as it further learns system-specific dependencies from the initial knowledge provided. Thus, by design, our approach attempts to cover very generic application domains. The methodology is based on (a) developmental mechanism with intrinsic motivation, and (b) a cognitive architecture (knowledge, reasoning, learning), while (c) utilizing minimal resources. It uses a hierarchical knowledge graph and kinematic reasoners to model and evaluate simple and advanced motion related skills. In our experiments, we use a resource-constrained millirobot with a volume of 36 cm^3 with a Raspberry Pi Pico 32-bit microcontroller (RP2040) that integrates all described features and capabilities except the camera system in 9 kB. Starting with learning the most elementary motor skills the millirobot autonomously progresses from simple linear and angular movements to complex geometric patterns within 15 minutes. To complement the physical experiments, we perform a simulation-based analysis that enables systematic comparisons across learning algorithms and intrinsic motivation parameters.
摘要:在本研究中,我們探討使小型資源受限系統(如厘米級的微型機器人)能夠自主探索、學習和適應其能力的發展機制。強化學習算法通過我們提出的tinyDSM的相互作用來指導代理的技能獲得和適應,該系統整合了內在動機和基於適應度的評估。我們追求最小的硬連接技能,同時鼓勵新技能的開放式發展。我們方法的一個關鍵重點是編碼最小的先驗一般知識,這作為系統進一步從提供的初始知識中學習系統特定依賴的基礎起點。因此,我們的方法設計上試圖涵蓋非常通用的應用領域。該方法論基於(a)具有內在動機的發展機制,以及(b)一種認知架構(知識、推理、學習),同時(c)利用最小資源。它使用層次知識圖譜和運動學推理器來建模和評估簡單和高級運動相關技能。在我們的實驗中,我們使用一個資源受限的微型機器人,其體積為36 cm^3,搭載Raspberry Pi Pico 32位微控制器(RP2040),該微控制器整合了所有描述的功能和能力,除了攝像頭系統外,僅佔用9 kB。從學習最基本的運動技能開始,微型機器人自主地在15分鐘內從簡單的線性和角運動進展到複雜的幾何圖形。為了補充物理實驗,我們進行了一個基於模擬的分析,這使得能夠在學習算法和內在動機參數之間進行系統比較。
TRUSS: Towards Task-Reliable and User-Safe Automated Agent Skill Generation
2608.17588v1 by Zhibo Zhang, Zhen Ouyang, Ling Shi, Kailong Wang
Agent Skills package reusable natural language procedures with executable resources, enabling software agents to acquire task specific capabilities without model adaptation. Automatically generating such Skills can improve task performance, yet evaluating a candidate solely from its artifact or final task outcome leaves unresolved which actions the equipped agent will perform and which side effects those actions will produce. We present TRUSS, an evidence guided framework for generating functionally effective and safety reliable Agent Skills. TRUSS first inspects functional claims against source and domain evidence while evaluating the complete artifact under nine predefined safety properties. Candidates admitted by this static gate are loaded by a shadow agent inside a Controllable Execution Environment, where brokered tools expose requested actions to policy enforcement and record their results as provenance preserving execution traces. Functional failures and property violations are linked back to the responsible Skill content and used to guide iterative refinement. We evaluate TRUSS on 168 SkillInject artifacts, 155 SkillSafetyBench cases, and all 187 tasks in SkillGenBench. TRUSS achieves 100.00\% precision and recall in vulnerability detection. Repair reduces attack success from 38.71\% to 19.35\% with GPT 5.5 and from 46.45\% to 29.68\% with GPT 5.4, with zero attack regression. For Skill generation, TRUSS raises task effectiveness from 17.11\% without Skills to 52.94\%, while increasing the benchmark Security rate from 50.80\% to 100.00\%. These results show that execution evidence can expose behavioral failures missed by artifact inspection and can guide Skill generation toward jointly verified functional and safety outcomes.
摘要:代理技能包可重用自然語言程序及可執行資源,使軟體代理能夠獲得特定任務的能力,而無需模型調整。自動生成這些技能可以提高任務表現,但僅從其產物或最終任務結果評估候選者,無法解決裝備代理將執行哪些行動以及這些行動將產生哪些副作用。我們提出了TRUSS,一個基於證據的框架,用於生成功能有效且安全可靠的代理技能。TRUSS首先根據來源和領域證據檢查功能聲明,同時在九個預定義的安全性屬性下評估完整的產物。通過這個靜態閘口的候選者將由一個影子代理加載到可控執行環境中,在這裡,經紀工具將請求的行動暴露給政策執行,並將其結果記錄為保留來源的執行痕跡。功能失敗和屬性違規將回溯到負責的技能內容,並用於指導迭代改進。
我們在168個SkillInject產物、155個SkillSafetyBench案例和所有187個SkillGenBench任務上評估TRUSS。TRUSS在漏洞檢測中達到100.00\%的精確度和召回率。修復將攻擊成功率從38.71\%降低到19.35\%(使用GPT 5.5),並從46.45\%降低到29.68\%(使用GPT 5.4),且沒有攻擊回歸。對於技能生成,TRUSS將任務有效性從沒有技能的17.11\%提高到52.94\%,同時將基準安全率從50.80\%提高到100.00\%。這些結果表明,執行證據可以揭示產物檢查中遺漏的行為失敗,並能指導技能生成朝向共同驗證的功能和安全結果。
Write, Execute, Refine: From Skill Followers to Skill Optimizers via Reinforcement Learning from Execution Feedback
2608.17587v1 by Kang Peng, Zhiwei Zhang, Yichen Zhang, Zezhong Wang, Yiming Du, Geng Tu, Baojun Wang, Bin Liang, Ruifeng Xu, Kam-Fai Wong
Expert-written natural language skills can improve tool-using agents, yet agent-authored skills perform 8-11 points worse than using no skill. This gap suggests that following procedural guidance and improving it from execution evidence are distinct capabilities. Inference time loops can repair skills but do not improve the model that writes the next one. We study how to organize execution experience from intermediate skills into training states for an optimizer. We introduce WER (Write, Execute, and Refine), a multi-phase framework that trains a Skill Optimizer outside a frozen executor. The optimizer proposes skills, a frozen agent executes each repeatedly, and a programmatic verifier scores the outcomes. The scores provide relative credit and select mixed-outcome records. Matched successful and failed trajectories from these records form the next phase's refinement states, so the optimizer learns from the consequences of its earlier outputs. On BFCL v4 multi-turn and tau2-bench, WER improves average Pass@1 over the no-skill baseline by 7.80 and 3.85 points, respectively. Under an identical refinement workflow, it outperforms the same backbone without optimizer training by 9.35 and 10.29 points. The trained 4B optimizer reaches 76.63 percent on BFCL v4, outperforming all evaluated off-the-shelf general-purpose models used as skill optimizers on average.
摘要:專家撰寫的自然語言技能可以改善工具使用代理,但代理撰寫的技能表現比不使用技能低 8-11 分。這一差距表明,遵循程序指導和從執行證據中改進它是兩種不同的能力。推理時間循環可以修復技能,但不會改善撰寫下一個技能的模型。我們研究如何將中介技能的執行經驗組織成優化器的訓練狀態。我們引入 WER(寫作、執行和精煉),這是一個多階段框架,旨在在凍結的執行器之外訓練技能優化器。優化器提出技能,凍結的代理重複執行每個技能,程式驗證器對結果進行評分。這些分數提供相對的信用並選擇混合結果記錄。來自這些記錄的成功和失敗的匹配軌跡形成下一階段的精煉狀態,因此優化器從其早期輸出的後果中學習。在 BFCL v4 多輪和 tau2-bench 上,WER 分別將平均 Pass@1 提高了 7.80 和 3.85 分。在相同的精煉工作流程下,它比未經優化器訓練的相同骨幹高出 9.35 和 10.29 分。訓練後的 4B 優化器在 BFCL v4 上達到 76.63% 的表現,超越了所有評估的現成通用模型,並在平均上用作技能優化器。
Auditing Exposure to Harmful Content on TikTok using Multimodal Language Models: A Cross-National, Age-Stratified Study
2608.17583v1 by Hamidreza Saffari, Francesco Pierri
Online video platforms can expose young users to harmful content, but independent audits remain difficult because video annotation is costly and moderation judgments vary across languages. We audit TikTok in France, Italy, and Sweden with sockpuppet accounts representing four age personas (13, 16, 19, 40), collecting 36,971 videos from passive For-You-page scrolling and active sessions that scroll, search for harm keywords, and scroll again. To scale annotation, we validate four multimodal LLMs against native-speaker labels on a 300-video reference set. Gemini 2.5 Flash with eight sampled frames plus text performs best (aggregate kappa = 0.42), at half the per-call cost of native-video upload, and we apply it to a 10% sample for approximately \$50 in total API spend across both modalities. Keyword search returns 35-56% harmful content, a 1.5-7.5x increase over the scrolling baseline in ten of twelve country-age combinations; the spike is temporary and flattens the age differences observed in France and Sweden. Under passive scrolling, Italy has the highest harm rate at every age, with Italian age-19 reaching 48.6%. Overall, MLLM-based auditing offers a scalable approach for cross-national youth-safety audits, while provider safety filters (1.1% refusal rate) under-count the most explicit harms.
摘要:在線視頻平台可能會讓年輕用戶接觸到有害內容,但獨立審核仍然困難,因為視頻註釋成本高且不同語言的審核判斷存在差異。我們在法國、意大利和瑞典對 TikTok 進行審核,使用代表四個年齡角色(13、16、19、40)的假帳號,從被動的 For-You 頁面滾動和主動會話中收集 36,971 個視頻,這些會話會滾動、搜索有害關鍵詞,然後再次滾動。為了擴大註釋,我們在 300 個視頻的參考集上驗證了四個多模態 LLM,與母語者標籤進行比較。Gemini 2.5 Flash 使用八個取樣幀加上文本的表現最佳(綜合 kappa = 0.42),其每次調用成本僅為本土視頻上傳的一半,我們將其應用於 10% 的樣本,總 API 支出約為 50 美元。關鍵詞搜索返回 35-56% 的有害內容,在十二個國家-年齡組合中的十個中,這比滾動基線增加了 1.5-7.5 倍;這一激增是暫時的,並平坦了在法國和瑞典觀察到的年齡差異。在被動滾動下,意大利在每個年齡段的危害率最高,意大利的 19 歲達到 48.6%。總的來說,基於 MLLM 的審核為跨國青少年安全審核提供了一種可擴展的方法,而提供者的安全過濾器(拒絕率 1.1%)則低估了最明顯的危害。
Quantifying Risk Under Evolving Uncertainty: Belief-Dependent Robustness for Safe Sequential Decision Making
2608.17574v1 by Deep Kumar Ganguly, Jan Kretinsky
How cautious should an agent be while it is still learning its environment? We propose RATTL (Risk-Adversarial Total-Reward Learning), which ties caution to epistemic uncertainty: the agent holds a Bayesian posterior over unknown dynamics and plans against a Wasserstein ambiguity set whose radius is a monotone function of that posterior. The radius contracts with evidence, so behaviour interpolates continuously between worst-case robustness and risk-neutral total-reward maximization. The design follows the duality underlying the Entropic Value-at-Risk, which converts the choice of a risk level into the choice of an ambiguity radius. We show the resulting planning problem is well posed under transience and compactness conditions, and prove a Safety Sandwich: the RATTL value lies between the uninformed robust value and the full- knowledge optimum, with a gap that vanishes as the posterior concentrates. In a canonical binary-hazard instance, the induced criterion reduces to Conditional Value-at-Risk at a level set by the posterior entropy. A worked example shows the agent deferring the efficient action until a sharp identification threshold. RATTL targets runtime safety for agents, including LLM-based systems, acting under uncertainty.
摘要:代理在學習其環境時應該多謹慎?我們提出了RATTL(風險對抗總回報學習),它將謹慎與認知不確定性聯繫起來:代理對未知動態持有貝葉斯後驗,並根據一個其半徑是該後驗單調函數的Wasserstein模糊集進行規劃。隨著證據的增加,半徑會收縮,因此行為在最壞情況的穩健性和風險中立的總回報最大化之間持續插值。該設計遵循了熵值風險的對偶性,將風險水平的選擇轉化為模糊半徑的選擇。我們顯示,所得到的規劃問題在瞬態和緊湊性條件下是良好定義的,並證明了一個安全三明治:RATTL值介於無信息穩健值和全知最優值之間,當後驗集中時,這一差距消失。在一個典型的二元危險實例中,所引入的標準簡化為在後驗熵設定的水平下的條件風險價值。一個具體的例子顯示,代理在達到明確識別閾值之前推遲了有效行動。RATTL針對在不確定性下行動的代理,包括基於LLM的系統,目標是運行時安全。
DMT-Dens: Density-preserving manifold visualization for biological data
2608.17571v1 by Ruizhe Wang, Yixuan Dong, Bolin Yang, Bingo Wing-Kuen Ling, Fuji Yang, Zelin Zang
Motivation: Low-dimensional embeddings are widely used to explore cell-state heterogeneity in single-cell and other high-dimensional biological data. Although many methods preserve local neighborhoods, they may distort the apparent sampling density of processed observations, altering the visual contrast between dense and sparse regions and complicating the interpretation of rare, transitional, or continuous cell-state populations. Results: We present DMT-Dens, a parametric manifold-visualization method built on a latent-token Transformer encoder. The model integrates rank-based manifold alignment with hard-pair aggregation. To preserve density, it optimizes a loss based on the Pearson correlation between k-nearest-neighbor log-radius estimates in the processed input and two-dimensional embedding spaces. Benchmark evaluations demonstrate strong density preservation, particularly on biological datasets, while retaining competitive label separability. Availability: Source code, data-processing scripts, and resolved experiment configurations are available at https://github.com/Ruizhe-wang/DMT-Dens.
摘要:動機:低維嵌入被廣泛用於探索單細胞及其他高維生物數據中的細胞狀態異質性。儘管許多方法保留了局部鄰域,但它們可能會扭曲處理觀察的表觀取樣密度,改變密集區域和稀疏區域之間的視覺對比,並使得對稀有、過渡或連續細胞狀態群體的解釋變得複雜。結果:我們提出了DMT-Dens,一種基於潛在標記Transformer編碼器的參數流形可視化方法。該模型將基於排名的流形對齊與硬配對聚合相結合。為了保留密度,它優化了一個基於處理輸入和二維嵌入空間中k最近鄰對數半徑估計之間的Pearson相關性的損失。基準評估顯示出強大的密度保留能力,特別是在生物數據集上,同時保持競爭性的標籤可分性。可用性:源代碼、數據處理腳本和解決的實驗配置可在https://github.com/Ruizhe-wang/DMT-Dens獲得。
Domain-Adapted Molecular Language Models for Efficient Search of Make-on-Demand Libraries
2608.17567v1 by Henrik Wille, Luis-Finley Schütz, Felix Strieth-Kalthoff
Pretrained molecular language models are increasingly used as molecular encoders for learning structure-property relationships. However, their practical suitability for molecular discovery within and beyond their pretraining domain remains unclear. Herein, we systematically benchmark four molecular language models across six virtual molecular libraries spanning drug discovery, organic materials, and catalysis. Native molecular language model embeddings show substantial variation in discovery performance across libraries, whereas molecular fingerprints provide a consistently strong and robust baseline. Consistent with a potential domain-representation mismatch, we show that explicit domain adaptation substantially improves representation performance. Fine-tuning molecular language model encoders on structures from the target virtual library consistently improves sample efficiency, with several adapted encoders emerging as the top-performing representations across the benchmark tasks. These results show that molecular representation quality depends strongly on the target domain and that explicit adaptation can improve the practical utility of molecular foundation models. More broadly, our findings establish domain-adapted molecular representations as a promising strategy for sample-efficient adaptive decision making in virtual screening and self-driving laboratories.
摘要:預訓練的分子語言模型越來越多地被用作學習結構-性質關係的分子編碼器。
然而,它們在其預訓練領域內外的分子發現中的實際適用性仍不明朗。
在此,我們系統性地基準測試了四種分子語言模型,涵蓋了六個虛擬分子庫,涉及藥物發現、有機材料和催化。
原生的分子語言模型嵌入在不同庫中的發現性能顯示出顯著的變化,而分子指紋則提供了一個一致強大且穩健的基準。
與潛在的領域表示不匹配一致,我們顯示明確的領域適應顯著改善了表示性能。
在目標虛擬庫的結構上微調分子語言模型編碼器,始終提高了樣本效率,其中幾個適應後的編碼器在基準任務中表現為最佳的表示。
這些結果顯示,分子表示的質量強烈依賴於目標領域,而明確的適應可以提高分子基礎模型的實際效用。
更廣泛地說,我們的發現確立了領域適應的分子表示作為在虛擬篩選和自駕實驗室中進行樣本高效自適應決策的一種有前景的策略。
Where a New Concept Must Enter: Entry Point Gates Cross-Task Usability in Unified Multimodal Models
2608.17564v1 by Zongyang Qiu, Yihan Wu, Kaixuan Fan, Bo Li, Hui Xiong
Unified multimodal models (UMMs) are motivated by the hope that understanding and generation reinforce each other but controlled ablations repeatedly find that adding a generation objective leaves understanding flat. Joint-training studies cannot settle the disagreement: with overlapping supervision, a gain cannot be attributed to the architecture rather than the data. To further investigate the relationship between the two directions in UMMs, we separate them by construction. A novel visual entity, a rendered 3D asset paired with a pseudo-word screened for absence from the frozen model's behavior, is bound through exactly one task direction, and the untrained direction is then measured. We find that the channel is real in both directions, but the directions differ in kind: generation training installs a name the model can only match among candidates; understanding training installs one it can also produce. What governs cross-task usability is where the binding enters the shared computation. An alignment probe predicts export across 36 configurations (Spearman $ρ= +0.68$). That objective's alignment term, maximized in closed form over activations with every weight frozen, makes a concept drawable when injected at layer 7 of 28 and is indistinguishable from the base model from layer 14 on, while the weight-based version of the same edit peaks at layers 10-14. In an observational series of four models, this window appears only where the understanding pathway is a semantic vision encoder, suggesting that unified weights are not enough: the two directions must share a semantic format at the entry point. Exploiting the rule, a mid-stack alignment objective acquires the concept for a $0.1\%$ relative loss of the model's general text-to-image ability, against $41\%$ for the standard generative route. Our code is at https://github.com/Zane-ZYQiu/entry-point-umm.
摘要:統一的多模態模型(UMMs)是受到理解與生成相互增強的希望所驅動,但控制性消融實驗反覆發現,添加生成目標會使理解保持平坦。聯合訓練研究無法解決這一分歧:在重疊監督下,增益無法歸因於架構而不是數據。為了進一步研究UMMs中這兩個方向之間的關係,我們通過構造將它們分開。一個新穎的視覺實體,即一個渲染的3D資產,與一個經過篩選以確保不出現在凍結模型行為中的偽詞配對,通過恰好一個任務方向綁定,然後測量未訓練的方向。我們發現這個通道在兩個方向上都是實際存在的,但這些方向在性質上有所不同:生成訓練安裝了一個模型只能在候選者中匹配的名稱;理解訓練則安裝了一個模型也可以生成的名稱。跨任務可用性的主導因素是綁定進入共享計算的地方。一個對齊探針預測在36個配置下的輸出(Spearman $ρ= +0.68$)。該目標的對齊項在凍結每個權重的情況下,對激活進行閉合形式的最大化,使得在28層的第7層注入時,概念可被繪製,並且從第14層開始與基礎模型無法區分,而同一編輯的基於權重的版本在第10-14層達到峰值。在一系列觀察四個模型的實驗中,這一窗口僅在理解路徑為語義視覺編碼器時出現,這表明統一權重並不夠:這兩個方向必須在進入點共享一個語義格式。利用這一規則,中堆棧對齊目標以$0.1\%$的相對損失獲得了該概念,這相對於標準生成路徑的$41\%$。我們的代碼位於 https://github.com/Zane-ZYQiu/entry-point-umm。
Reflex-Guard: A Low-Latency Guardrail for LLM Prompt Safety Using Dense Semantic Embeddings
2608.17556v1 by Istiaque Ahmed, Afia Anjum Borsha, Ranat Das Prangon, Abu-fuad Ahmad, Thi Hong Tran
Large Language Models (LLMs) in real-world applications often face the risks of specially crafted prompts designed to bypass the safety controls. Existing guardrail methods, such as LLM-as-a-judge and cloud-based safety APIs are able to detect unsafe content. However, they often add a delay of about 250-900 ms to each request. This delay is too high for real-time applications, when the system usually needs to respond in less than 100 ms. Furthermore, routing user prompts through external moderation endpoints raises significant data privacy concerns. This paper introduces Reflex-Guard, a lightweight guardrail that runs locally. It uses jailbreak-aware preprocessing, compact sentence-transformer embeddings, and seven fast binary classifiers. Together, these components enable high-accuracy prompt safety filtering with much lower latency than existing solutions. Through systematic evaluation on a strategically balanced dataset of 30,568 samples drawn from five complementary sources, we demonstrate that Reflex-Guard achieves 95.9% recall on harmful prompts at 37.6 ms end-to-end latency. It is faster than existing baselines, including Llama Guard 2 at 255 ms and SafeDecoding at 723 ms. It can detect 100% of GCG suffix attacks and Base64-encoded prompts using the default threshold. However, DrAttack structured prompts required lowering the threshold to 0.03 for optimal detection, as they produced a distinct probability distribution. Reflex-Guard achieves Reflex Efficiency Score (RES) scores up to 16.79, significantly outperforming Llama Guard 2 (11.90) and SafeDecoding (9.80). This analysis offers practical deployment advice and shows that different attack types occupy distinct regions in the embedding probability space.
摘要:大型語言模型(LLMs)在實際應用中常常面臨專門設計的提示風險,這些提示旨在繞過安全控制。現有的防護方法,如LLM作為評判者和基於雲的安全API,能夠檢測不安全的內容。然後,它們通常會為每個請求增加約250-900毫秒的延遲。這個延遲對於實時應用來說過高,因為系統通常需要在100毫秒內作出回應。此外,通過外部審核端點路由用戶提示會引發重大數據隱私問題。本文介紹了Reflex-Guard,一種輕量級的本地防護措施。它使用監獄破解感知的預處理、緊湊的句子轉換器嵌入和七個快速的二元分類器。這些組件共同實現了高準確度的提示安全過濾,延遲遠低於現有解決方案。通過對來自五個互補來源的30,568個樣本的戰略性平衡數據集進行系統評估,我們證明Reflex-Guard在有害提示上達到了95.9%的召回率,端到端延遲為37.6毫秒。它比現有的基準更快,包括Llama Guard 2的255毫秒和SafeDecoding的723毫秒。它可以使用默認閾值檢測100%的GCG後綴攻擊和Base64編碼的提示。然而,DrAttack結構化提示需要將閾值降低到0.03以達到最佳檢測,因為它們產生了不同的概率分佈。Reflex-Guard的反射效率得分(RES)高達16.79,顯著超過Llama Guard 2(11.90)和SafeDecoding(9.80)。這一分析提供了實際部署建議,並顯示不同的攻擊類型在嵌入概率空間中佔據不同的區域。
Code as Representation: A Compilable Parsing Paradigm for Academic Documents
2608.17550v1 by Rihui Jin, Jun Wang, chengyuan zhu, Liang Mingyu, Yue Gao, Li Yunxuan, Kuicai Dong, Guilin Qi, Lin Ren, Yongrui Chen, Xinbang Dai, Jiaqi Li, Tongtong Wu, Gholamreza Haffari
Academic papers are a primary carrier of scientific knowledge, yet most of this knowledge remains locked in PDFs that are optimized for human reading rather than machine use. For Multimodal Large Language Models (MLLMs), the core challenge is not only perception, but representation: scientific pages interleave text with Structured Academic Elements (SAEs) such as tables, formulas, charts, and pseudocode, whose structure, data, and logic are poorly preserved by common surrogates like Markdown. We therefore propose Compilable Academic Document Parsing (CADP), a paradigm that reconstructs a full page as contextual \LaTeX{} plus executable Python, so that structure-preserving elements and executable chart representations can be reconstructed, recompiled, and directly verified against the source page. To support this setting, we introduce CADP-Bench, an expert-verified benchmark of full academic pages containing tightly coupled text and multiple SAE types, evaluated through a re-injection compilation protocol. We further study current capabilities using SOTA MLLMs and an exploratory multi-agent baseline that incorporates common agentic techniques. Results show that even frontier models still struggle to produce high-fidelity executable reconstructions, highlighting substantial room for improvement in structure-aware scientific document parsing. CADP-Bench is released for future research.
摘要:學術論文是科學知識的主要載體,但大部分這些知識仍然鎖定在優化為人類閱讀而非機器使用的PDF中。對於多模態大型語言模型(MLLMs)來說,核心挑戰不僅在於感知,還在於表徵:科學頁面將文本與結構化學術元素(SAEs)交錯,如表格、公式、圖表和偽代碼,其結構、數據和邏輯在常見的替代品如Markdown中保存得很差。因此,我們提出可編譯學術文檔解析(CADP),這是一種將整個頁面重建為上下文 \LaTeX{} 加上可執行的Python的範式,以便結構保留的元素和可執行的圖表表示可以被重建、重新編譯並直接與源頁面進行驗證。為了支持這一設置,我們引入CADP-Bench,一個經專家驗證的完整學術頁面基準,包含緊密耦合的文本和多種類型的SAE,通過重新注入編譯協議進行評估。我們進一步研究使用SOTA MLLMs的當前能力以及一個探索性的多代理基準,該基準結合了常見的代理技術。結果顯示,即使是最前沿的模型仍然難以產生高保真度的可執行重建,突顯出結構感知的科學文檔解析有很大的改進空間。CADP-Bench已經釋出以供未來研究使用。
No Gaussian Required: Contrastive Inverse Dynamics for JEPA World Models
2608.17542v1 by Jack Boylan, Chris Hokamp
Joint-Embedding Predictive Architectures (JEPAs) learn world models by predicting future embeddings, but the objective admits a trivial solution of a constant encoder, so every practical system adds an anti-collapse mechanism (LeCun, 2022; Assran et al., 2023; Bardes et al., 2022; 2024). LeWorldModel (LeWM) prevents collapse with SIGReg, a regularizer that forces the latent distribution to match an isotropic Gaussian: the representation is stabilized by prescribing what it must look like, independently of the environment it models. We argue that the anti-collapse pressure can instead come from the transition data itself. Action-Contrastive Masked Transition Modeling (AC-MTM) keeps LeWM's forward latent-prediction objective and adds a training-only inverse-dynamics head trained with Action-NCE: each latent transition must identify the action that produced it among the other actions in the batch, a discrimination task that a collapsed encoder provably fails. The inverse branch is discarded after training, leaving test-time encoding, forward prediction, planning, and compute identical to LeWM. On four standard pixel-control tasks under a matched planning protocol, AC-MTM trains stably from scratch and matches SIGReg on average. On the harder multi-object OGBench Visual Scene task, results are consistent with the prescribed geometry becoming a bottleneck: AC-MTM reaches 80.0$\pm$2.0% success versus 58.0$\pm$2.0% for SIGReg, improving by 20-24 points in each training seed. A single 50-episode random-policy run gives a 52% baseline estimate. Contrastive inverse dynamics thus provides a distribution-free anti-collapse signal that requires no target network, stop-gradient, pretrained encoder, or reconstruction objective, and we characterize the action-space and observability assumptions under which it holds. We make our code available at https://github.com/jackboyla/action-contrastive-jepa
摘要:聯合嵌入預測架構(JEPAs)透過預測未來嵌入來學習世界模型,但該目標允許一個恆定編碼器的平凡解,因此每個實際系統都添加了一個反崩潰機制(LeCun, 2022; Assran et al., 2023; Bardes et al., 2022; 2024)。LeWorldModel(LeWM)透過SIGReg防止崩潰,這是一種正則化器,強迫潛在分佈與各向同性高斯匹配:該表示通過規定其必須的樣貌來穩定,無論其所建模的環境如何。我們認為反崩潰壓力可以來自於過渡數據本身。行動對比遮罩過渡建模(AC-MTM)保持LeWM的前向潛在預測目標,並添加一個僅訓練的逆動力頭,該頭使用行動-NCE進行訓練:每個潛在過渡必須在批次中的其他行動中識別出產生它的行動,這是一個崩潰編碼器顯然無法完成的區分任務。逆分支在訓練後被丟棄,留下測試時的編碼、前向預測、規劃和計算與LeWM相同。在四個標準像素控制任務中,根據匹配的規劃協議,AC-MTM從零開始穩定訓練,並在平均上匹配SIGReg。在更困難的多物體OGBench視覺場景任務中,結果與規定的幾何形狀成為瓶頸的情況一致:AC-MTM達到80.0$\pm$2.0%的成功率,而SIGReg則為58.0$\pm$2.0%,在每個訓練種子中提高了20-24點。一個50集隨機策略的運行給出了52%的基線估計。因此,對比逆動力提供了一個無分佈的反崩潰信號,無需目標網絡、停止梯度、預訓練編碼器或重建目標,我們還描述了其成立的行動空間和可觀察性假設。我們的代碼可在https://github.com/jackboyla/action-contrastive-jepa獲得。
CoAL-RAG: A Complexity-Aware Legal Retrieval-Augmented Generation Method
2608.17536v1 by Jin Su, Zhuofeng Zhao, Huanhuan Wang, Hao Chen
Legal consultation questions exhibit multi-level complexity. A single retrieval strategy often leads to over-reasoning for simple questions and poor interpretability for complex ones, making it difficult to meet the requirements for both answer quality and efficiency in high-risk scenarios. To address this issue, this paper proposes CoAL-RAG, a complexity-aware legal retrieval-augmented generation method, which constructs a multi-dimensional evaluation mechanism based on question essence'' andretrieval consistency'' to enable adaptive routing of retrieval strategies. First, the reasoning demand is quantified according to the logical structure of the question. Then, the discrepancy between semantic retrieval and keyword retrieval is utilized to indirectly reflect problem complexity, thereby selecting the most appropriate retrieval strategy and dynamically filtering contextual information. Experimental results demonstrate that the proposed method significantly outperforms baseline models not only on Chinese legal benchmarks (SocialLawQA, LawBench) but also demonstrates strong cross-jurisdictional generalization on English datasets (LexGLUE, CaseHold). Specifically, on Chinese datasets, the BLEU score improves by 42.5\% and ROUGE-L reaches 3.6 times that of knowledge graph-based methods. On English benchmarks, CoAL-RAG maintains highly competitive accuracy, achieving an optimal balance between generation quality, deep logical reasoning, and system efficiency across different legal systems.
摘要:法律諮詢問題展現出多層次的複雜性。單一的檢索策略常常導致對簡單問題的過度推理,以及對複雜問題的可解釋性差,使得在高風險情境中難以滿足答案質量和效率的要求。為了解決這個問題,本文提出了 CoAL-RAG,一種具複雜性意識的法律檢索增強生成方法,該方法基於「問題本質」和「檢索一致性」構建了一個多維評估機制,以實現檢索策略的自適應路由。首先,根據問題的邏輯結構量化推理需求。然後,利用語義檢索與關鍵字檢索之間的差異,間接反映問題的複雜性,從而選擇最合適的檢索策略並動態過濾上下文信息。實驗結果表明,所提出的方法在中國法律基準(SocialLawQA、LawBench)上顯著超越基線模型,並且在英語數據集(LexGLUE、CaseHold)上展現出強大的跨法域泛化能力。具體而言,在中國數據集上,BLEU 分數提高了 42.5\%,而 ROUGE-L 達到知識圖譜方法的 3.6 倍。在英語基準上,CoAL-RAG 維持了高度競爭的準確性,在不同法律系統中實現生成質量、深度邏輯推理和系統效率之間的最佳平衡。
ArborMem: Navigating Interaction States with Memory Forests
2608.17534v1 by Zongwei Lv, Yuemeng Xu, Yilun Yao, Siyi Ding, Xinyu Tan, Yaoming Li, Guangxiang Zhao, Weihong Lin, Lin Sun, Xiangzheng Zhang, Tong Yang
Large language models increasingly serve as persistent conversational assistants, requiring memory that preserves relevant experience and maintains continuity across interactions. Existing methods improve access to conversational history through long-context processing, selective retrieval, and structured memory organization. However, most systems treat memory access as retrieving relevant past information without first determining which prior interaction state the current turn resumes. This limitation becomes particularly important when conversations interleave multiple tasks, people, and plans that may be interrupted and later revisited. We introduce ArborMem, an online memory framework that represents a long-running conversation as a navigable forest of interaction states. Each branch preserves a locally coherent trajectory, while the forest maintains multiple trajectories that may later be resumed. For each new input, ArborMem localizes the relevant state, restores its branch-local context, and augments it with reusable evidence retrieved across branches, preserving interaction continuity without conflating semantically related but structurally distinct trajectories. Existing long-term memory benchmarks cover diverse memory and reasoning capabilities but do not explicitly isolate branch-structured challenges. We therefore introduce BranchMemEval, a controlled diagnostic benchmark for interleaved and resumable interaction trajectories. Experiments on LongMemEval, LoCoMo, BEAM 100K, and BranchMemEval show that ArborMem outperforms the strongest baselines by 3.36 to 10.31 percentage points on the three established benchmarks and by 5.0 points on BranchMemEval. Its advantage grows under constrained read budgets, while complete memory queries remain below half a second.
摘要:大型語言模型越來越多地作為持久的對話助手,這需要記憶來保留相關經驗並在互動中保持連貫性。現有的方法通過長上下文處理、選擇性檢索和結構化記憶組織來改善對對話歷史的訪問。然而,大多數系統將記憶訪問視為檢索相關的過去信息,而不首先確定當前回合恢復的先前互動狀態。當對話交織著多個任務、人物和計劃,這些任務可能會被中斷並在稍後重新訪問時,這一限制變得尤為重要。我們介紹了 ArborMem,一個在線記憶框架,將長期對話表示為可導航的互動狀態森林。每個分支保留一個局部一致的軌跡,而森林則維護多條可能稍後恢復的軌跡。對於每個新的輸入,ArborMem 定位相關狀態,恢復其分支局部上下文,並通過跨分支檢索的可重用證據進行增強,保留互動的連續性而不混淆語義上相關但結構上不同的軌跡。現有的長期記憶基準涵蓋了多樣的記憶和推理能力,但並未明確隔離分支結構挑戰。因此,我們引入了 BranchMemEval,一個針對交織和可恢復互動軌跡的受控診斷基準。在 LongMemEval、LoCoMo、BEAM 100K 和 BranchMemEval 上的實驗顯示,ArborMem 在三個既定基準上比最強基線高出 3.36 到 10.31 個百分點,在 BranchMemEval 上高出 5.0 個百分點。其優勢在受限的讀取預算下增長,而完整的記憶查詢仍保持在半秒以下。
When to Review: Spaced Repetition for Continual Pre-Training of Language Models
2608.17530v1 by Alankar Atreya, Devesh Batra, Yoages Kumar Mantri, Geremy Bantug, Greig A Cowan, Raad Khraishi
Continual pre-training of large language models must acquire new information without erasing old knowledge. Existing replay methods often choose a global old/new mixture and sample uniformly, ignoring that examples differ in how quickly they are forgotten. We formulate continual pre-training as adaptive review scheduling: the training loop should decide not only how much history to replay, but which examples should return at each step. We introduce Spaced Repetition Training (SRT), a continual learning framework inspired by cognitive science, which schedules sample-rehearsal using the SuperMemo-2 (SM-2) algorithm. SRT maintains per-example review state, maps per-example perplexity to a recall-quality signal, and schedules historical examples for retention and new examples for consolidation while leaving the model, objective, and optimizer unchanged. On temporally separated Wikipedia and code corpora, SRT improves the stability-plasticity trade-off, recovering 5 to 37 percentage points of old-knowledge accuracy lost by naive continual pre-training across model scales while preserving or improving new-knowledge acquisition. At larger scale, SRT preserves broad benchmark performance that naive continual pre-training and uniform replay substantially degrade. Experiments with vision and tabular data further suggest that the scheduling principle extends beyond language when paired with an appropriate recall signal.
摘要:持續的預訓練大型語言模型必須在不抹去舊知識的情況下獲取新信息。現有的重播方法通常選擇一個全局的舊/新混合並均勻抽樣,忽略了示例在被遺忘的速度上存在差異。我們將持續預訓練公式化為自適應回顧排程:訓練循環應決定不僅是重播多少歷史,還有每一步應該返回哪些示例。我們引入了間隔重複訓練(SRT),這是一個受認知科學啟發的持續學習框架,使用 SuperMemo-2 (SM-2) 算法來排程樣本重複。SRT 維持每個示例的回顧狀態,將每個示例的困惑度映射到回憶質量信號,並在保留模型、目標和優化器不變的情況下,為保留歷史示例和鞏固新示例進行排程。在時間上分隔的維基百科和代碼語料庫上,SRT 改善了穩定性與可塑性的權衡,恢復了由天真的持續預訓練在各模型規模上損失的 5 到 37 個百分點的舊知識準確率,同時保留或改善了新知識的獲取。在更大規模下,SRT 保持了廣泛的基準性能,而天真的持續預訓練和均勻重播則大幅降低了這一性能。對於視覺和表格數據的實驗進一步表明,當與適當的回憶信號配對時,排程原則超越了語言的範疇。
Agent Lightning v1.0: Towards Harnessed Agentic RL
2608.17528v1 by Zhiyuan He, Siwei Zhang, Zhiwen Zhou, Yuqing Yang, Yu Kang, Yuge Zhang, Luna K. Qiu, Tin Yan Tsui, Jiahang Xu, Chong Luo
Modern agents operate inside agent harnesses that manage tools, context, and control flow, making the harness a critical part of the agent system. Our original Agent Lightning introduced a disaggregated architecture that connects arbitrary agents to RL training through an LLM endpoint proxy, an approach later adopted by frameworks such as verl Uni-Agent, AReaL 2.0, slime, and Polar. We refer to this paradigm as harnessed agentic RL, where the deploy-time harness directly participates in model post-training. Harnessed agentic RL differs fundamentally from traditional agentic RL: the harness, rather than the training engine, owns the environment interaction loop, while the trainer observes only sequences of LLM request-response pairs. This introduces challenges in retokenization, sample merging, advantage calculation, loss normalization, and backend scheduling, which can substantially affect training stability and effectiveness. We present Agent Lightning v1.0, a lightweight framework for harnessed agentic RL implemented in approximately 3,500 lines of code. It supports arbitrary agent harnesses and serves as a practical testbed for studying these challenges. We evaluate it on instruction-following, search, and coding agents, and provide a complete reproducible pipeline for coding-agent RL. Using only 6K training examples and modest compute, RL improves Qwen3.5-9B on SWE-bench Verified from 41.8% to 56.4%, a 14.6-point absolute gain. We release the complete workflow and training scripts to facilitate reproducible research on harnessed agentic RL.
摘要:現代代理人運行在管理工具、上下文和控制流程的代理人鞍具內,使得鞍具成為代理人系統中的關鍵部分。我們的原始 Agent Lightning 引入了一種解耦架構,通過 LLM 端點代理將任意代理人連接到強化學習訓練,這種方法後來被如 verl Uni-Agent、AReaL 2.0、slime 和 Polar 等框架採用。我們將這種範式稱為鞍具代理強化學習,其中部署時的鞍具直接參與模型的後訓練。鞍具代理強化學習在根本上與傳統的代理強化學習不同:鞍具,而不是訓練引擎,擁有環境互動循環,而訓練者僅觀察 LLM 請求-回應對的序列。這在重新標記、樣本合併、優勢計算、損失正規化和後端調度方面引入了挑戰,這些挑戰可能會顯著影響訓練的穩定性和有效性。我們提出了 Agent Lightning v1.0,這是一個輕量級的鞍具代理強化學習框架,實現約 3,500 行代碼。它支持任意的代理人鞍具,並作為研究這些挑戰的實用測試平台。我們在遵循指令、搜索和編碼代理人上對其進行評估,並提供完整的可重現管道以進行編碼代理人強化學習。僅使用 6K 訓練範例和適度的計算,強化學習將 Qwen3.5-9B 在 SWE-bench Verified 上的表現從 41.8% 提升至 56.4%,絕對增益為 14.6 點。我們發布完整的工作流程和訓練腳本,以促進對鞍具代理強化學習的可重現研究。
Explainable AI-Powered Framework for Video-Based Skill Assessment in Cataract Surgery
2608.17522v1 by Mohammad Javad Ahmadi, Hamid D. Taghirad
Persistent shortages in the surgical workforce and inherent limitations of traditional training methods highlight the necessity of automated, data-driven approaches in surgical education. This study addresses these challenges by introducing a novel, explainable AI-powered framework for automated skill assessment, specifically focusing on cataract surgery. We present the world's largest dataset of cataract surgery videos, comprising 2,000 recordings. Additionally, we propose an AI-powered analytical framework that employs advanced computer vision and signal-processing techniques to automatically evaluate surgical videos to derive objective, quantitative performance indicators that complement or potentially replace subjective scoring methods. A significant advantage of our framework over previous methods lies precisely in its explainability of outputs, elevating it beyond merely an opaque skill classification tool. Through experimental analysis of 83 cataract surgery videos, we demonstrate that the automatically computed metrics exhibit strong correlations with expert-based subjective evaluations, achieving up to 87% accuracy in surgical skill assessment. Each metric was individually examined, and expert surgeons provided subjective ratings using the newly introduced Capsulorhexis Skill Assessment System (CSAS). These subjective assessments were compared with ten objective motion-based metrics extracted through our framework. The results indicated a robust correlation between subjective ratings and automated indicators, underscoring the framework's capacity to accurately model surgical expertise.
摘要:持續的外科醫療人力短缺以及傳統訓練方法的固有限制凸顯了在外科教育中自動化、數據驅動方法的必要性。這項研究通過引入一個新穎的、可解釋的人工智慧驅動框架來解決這些挑戰,特別專注於白內障手術。我們展示了世界上最大的白內障手術視頻數據集,包含2,000個錄像。此外,我們提出了一個人工智慧驅動的分析框架,利用先進的計算機視覺和信號處理技術,自動評估手術視頻,以獲得客觀的、定量的性能指標,這些指標可以補充或潛在地取代主觀評分方法。我們的框架相較於先前的方法的一個顯著優勢恰恰在於其輸出的可解釋性,使其超越僅僅是一個不透明的技能分類工具。通過對83個白內障手術視頻的實驗分析,我們證明自動計算的指標與專家基於主觀評估的評分之間存在強烈的相關性,在外科技能評估中達到高達87%的準確率。每個指標都經過單獨檢查,專家外科醫生使用新引入的囊膜切開技能評估系統(CSAS)提供主觀評分。這些主觀評估與通過我們的框架提取的十個客觀運動基礎指標進行了比較。結果顯示主觀評分與自動指標之間存在穩健的相關性,強調了該框架準確建模外科專業知識的能力。
Effects of Answer Format Variation on Gender Bias in Large Language Models
2608.17516v1 by Ksenia Merzlyakova, Sebastian Padó, Franziska Weeber
Gender bias or other social biases in large language models (LLMs) are frequently evaluated with question answering or survey benchmarks where the LLM needs to give a response in a predefined answer format. It is well known in survey science that the answer format has a substantial impact on answers, just as LLMs are sensitive to the prompt wording. However, to our knowledge it has not been studied yet how changes in answer format impact the measurement of gender bias in LLMs and their alignment with human response distributions. We evaluate three instruction-tuned models on the BBQ benchmark and OpinionQA survey data across closed-ended, Likert-scaled and open-ended formats, comparing bias measurement and distributional alignment under otherwise identical conditions. We find that answer format does substantially alter measured outcomes, including reversals in order rankings. These differences arise because each format elicits distinct response behaviours, such as forced-choice selection, scale-based distributions and refusal in free-text generation. Our findings highlight the importance of treating answer format as a substantive component of LLM evaluation and motivate multi-format designs for more robust model assessment.
摘要:性別偏見或其他社會偏見在大型語言模型(LLMs)中的評估,通常使用問答或調查基準,其中LLM需要以預定的答案格式給出回應。調查科學中已知答案格式對答案有重大影響,就像LLMs對提示措辭敏感一樣。然而,據我們所知,尚未研究答案格式的變化如何影響LLMs中性別偏見的測量及其與人類回應分佈的一致性。我們在BBQ基準和OpinionQA調查數據上評估了三個經過指令調整的模型,並比較了在封閉式、Likert量表和開放式格式下的偏見測量和分佈一致性,條件則保持一致。我們發現答案格式確實顯著改變了測量結果,包括排序排名的逆轉。這些差異的產生是因為每種格式引發了不同的回應行為,例如強制選擇、基於量表的分佈和在自由文本生成中的拒絕。我們的發現強調將答案格式視為LLM評估的實質性組成部分的重要性,並促使多格式設計以進行更穩健的模型評估。
Beyond FLOPs: Energy-Aware Knowledge Distillation for Sustainable LLMs on Code-Related Task
2608.17515v1 by Enrique Barba Roque, Luís Cruz, Annibale Panichella
Background: Large Language Models (LLMs) are increasingly being applied to Software Engineering (SE) tasks, achieving high accuracy across problems such as clone detection, vulnerability prediction, and code summarization. However, their high computational demands and energy consumption raise sustainability concerns and hinder their use on consumer hardware and resource-constrained platforms. A common way to report the computational cost of an LLM in the literature and industry is to use the number of Floating Point Operations (FLOPs) required to perform a pass over the network. Aims: This paper investigates the implications of energy-aware knowledge distillation for SE, aiming to improve model efficiency while maintaining performance and to determine whether FLOPs is a reliable energy-aware metric. Method: We conduct a controlled experiment using Morph, a Many-Objective Optimization-based distillation methodology, to empirically examine whether FLOPs accurately reflect energy consumption in Clone Detection and Vulnerability Prediction tasks. We extend this methodology to include energy-surrogate models that directly estimate CPU and GPU energy consumption during optimization, and we apply Morph to generative tasks using CodeT5+ for code summarization. Results: Our results show that FLOPs is not always a reliable indicator of energy consumption, and better results can be achieved by using energy-surrogate models. Distilled student models can reduce inference energy consumption by up to 90\% and memory usage by 86\%, with only modest accuracy trade-offs. Conclusions: Energy-aware knowledge distillation when guided by direct energy surrogates rather than FLOPs can improve the energy consumption, sustainability, and deployability of LLMs for SE applications, enabling efficient models on consumer hardware.
摘要:背景:大型語言模型(LLMs)越來越多地應用於軟體工程(SE)任務,在克隆檢測、漏洞預測和程式碼摘要等問題上達到了高準確率。
然而,它們的高計算需求和能量消耗引發了可持續性問題,並阻礙了它們在消費者硬體和資源受限平台上的使用。
在文獻和業界中,報告LLM計算成本的常見方法是使用執行一次網絡所需的浮點運算次數(FLOPs)。
目標:本文探討了對SE進行能量感知知識蒸餾的影響,旨在提高模型效率的同時保持性能,並確定FLOPs是否是一個可靠的能量感知指標。
方法:我們使用Morph進行了一項受控實驗,這是一種基於多目標優化的蒸餾方法,實證檢驗FLOPs是否準確反映克隆檢測和漏洞預測任務中的能量消耗。
我們擴展了這一方法,納入能量替代模型,這些模型在優化過程中直接估算CPU和GPU的能量消耗,並將Morph應用於使用CodeT5+進行程式碼摘要的生成任務。
結果:我們的結果顯示FLOPs並不總是能可靠指示能量消耗,使用能量替代模型可以獲得更好的結果。
蒸餾的學生模型可以將推理能量消耗降低高達90%,內存使用量降低86%,而準確率僅有適度的折衷。
結論:在直接能量替代模型的指導下,能量感知知識蒸餾可以改善LLMs在SE應用中的能量消耗、可持續性和可部署性,從而使消費者硬體上的模型更加高效。
SGHA: Evidence-Grounded Research Problem Discovery with Local Language Models
2608.17501v1 by Sarvesh Gharat, Junpei Komiyama
Recent efforts toward fully automated AI scientists have demonstrated that language-model agents can generate hypotheses, execute experiments, and draft scientific manuscripts. However, during the early stages of research, when research problems are formulated, these AI scientists often rely heavily on proprietary frontier models. Their proposals are shaped by opaque parametric knowledge and by literature searches conditioned on the proposals themselves. Such knowledge is effectively a black box, and this dependence makes the evidential basis and validity of generated research problems difficult to audit and leaves the process vulnerable to model-specific hallucinations and biases. Furthermore, if proprietary research materials are transmitted to external APIs, the use of these models creates confidentiality, privacy, and data-governance concerns. We introduce the Structural Gap Hypothesis Agent (SGHA), a fully automated, corpus-first research-problem discovery system that runs entirely on a local LLM. SGHA structures a scientific literature corpus into evidence-linked paper objects and a typed evidence graph, detects unresolved structural patterns across papers, screens candidate gaps before formulation, and produces traceable research-problem families. In particular, it is able to output assumptions, objectives, success criteria, and remaining ambiguities. All LLM-based components of SGHA are executed using a locally served open-weight 9B language model, without requiring proprietary frontier-model APIs. We compare SGHA with the AI Scientist-v2 idea formulation module in five machine-learning domains. Our results suggest that explicit corpus structure and evidence-constrained reasoning can support promising, inspectable research-problem formulation without relying on frontier models during generation or verification.
摘要:最近對於完全自動化的AI科學家的努力顯示,語言模型代理可以生成假設、執行實驗並撰寫科學手稿。
然而,在研究的早期階段,當研究問題被形成時,這些AI科學家往往過度依賴專有的前沿模型。
他們的提案受到不透明的參數知識和基於提案本身的文獻搜尋的影響。
這種知識實際上是一個黑箱,而這種依賴使得生成的研究問題的證據基礎和有效性難以審核,並使過程容易受到模型特定的幻覺和偏見的影響。
此外,如果專有研究材料被傳輸到外部API,使用這些模型會產生保密性、隱私和數據治理的問題。
我們介紹了結構性差距假設代理(SGHA),這是一個完全自動化的、以語料庫為首的研究問題發現系統,完全在本地的LLM上運行。
SGHA將科學文獻語料庫結構化為與證據相關聯的論文對象和類型化的證據圖,檢測論文之間未解決的結構模式,在形成之前篩選候選差距,並生成可追溯的研究問題家族。
特別是,它能夠輸出假設、目標、成功標準和剩餘的模糊性。
SGHA的所有基於LLM的組件都是使用本地提供的開放權重9B語言模型執行的,而不需要專有的前沿模型API。
我們將SGHA與AI Scientist-v2的想法形成模塊在五個機器學習領域進行比較。
我們的結果表明,明確的語料結構和基於證據的推理可以支持有前景的、可檢查的研究問題形成,而無需在生成或驗證過程中依賴前沿模型。
When AI Designs AI: Innovation or Imitation?
2608.17471v1 by Yikang Yang, Zhengxin Yang, Luzhou Peng, Minghao Luo, Yanqi Kan, Wanling Gao, Jianfeng Zhan
Recent advances in LLM agents have made them increasingly capable of designing methods for complex AI tasks. This raises two central questions about agent-designed methods relative to human-designed methods: how well they perform, and how different their algorithmic designs are. To study these questions, this paper introduces an analysis that derives task-specific algorithmic design spaces from human-designed methods, maps both human- and agent-designed methods into these spaces, and quantifies their algorithmic differences at the module level. Widely used LLM agents are evaluated on a suite of representative, open-ended AI tasks spanning multiple modalities, and the methods they design are analyzed in terms of both task performance and algorithmic differences from human-designed methods. Experimental results show that current agents can occasionally match or surpass human state-of-the-art (SOTA) performance (10/72 configurations), but such success does not generalize reliably across tasks or agents. Moreover, 96.8% of agent-designed methods fall within human-derived algorithmic design spaces, largely recombining algorithmic choices found in human-designed methods, while nearly half exactly match an existing human algorithmic design. Taken together, these findings suggest that although current agents can occasionally match or surpass human SOTA performance, their algorithmic designs remain within human-derived algorithmic design spaces, reflecting the reuse and recombination of algorithmic choices.
摘要:最近在大型語言模型(LLM)代理方面的進展使它們在設計複雜人工智慧任務的方法上變得越來越有能力。這引發了兩個關於代理設計的方法與人類設計的方法的核心問題:它們的表現如何,以及它們的算法設計有多不同。為了研究這些問題,本文介紹了一種分析方法,從人類設計的方法中推導出特定任務的算法設計空間,將人類和代理設計的方法映射到這些空間中,並在模塊層面量化它們的算法差異。廣泛使用的LLM代理在一系列具有代表性的開放式人工智慧任務中進行評估,這些任務涵蓋多種模式,並分析它們設計的方法在任務表現和與人類設計的方法的算法差異方面。實驗結果顯示,當前的代理偶爾可以匹配或超越人類的最先進(SOTA)表現(10/72配置),但這種成功並不可靠地在不同任務或代理之間泛化。此外,96.8%的代理設計的方法都落在由人類推導的算法設計空間內,主要是重新組合在人體設計的方法中找到的算法選擇,而近一半則完全匹配現有的人類算法設計。綜合來看,這些發現表明,儘管當前的代理偶爾可以匹配或超越人類的SOTA表現,但它們的算法設計仍然保持在由人類推導的算法設計空間內,反映了算法選擇的重用和重新組合。
SAGE: Self-Evolving Storyboard Skills via Attribution-Guided Rule Evolution
2608.17468v1 by Maolin Ran, Xiaoyang Lu, Jiaqi Liu, Jian Wang, Weiwen Liu, Jianghao Lin, Yong Yu, Weinan Zhang
Storyboards turn screenplays into visual shot plans for automated short drama production. Professional storyboarding relies on tacit directorial expertise and remains an industrial bottleneck. Large language models can automate this step, but methods for supplying directing knowledge face three challenges: (1) Knowledge acquisition: the craft remains implicit in exemplars or must be written manually. (2) Knowledge refinement: authored knowledge is not evaluated against execution outcomes, and opaque generation prevents feedback attribution to the knowledge behind each decision. (3) Knowledge injection: injecting all knowledge exceeds usable context, while manual selection for every narrative group does not scale. We present SAGE (Skill with Attribution-Guided Evolution), a deployed framework that learns, attributes, evolves, and routes directing knowledge from expert demonstrations. SAGE derives rules that are independent of episode content by contrasting each training screenplay with its expert storyboard. During generation, the model records each narrative group's adopted rules. Combining these records with localized feedback enables targeted updates to individual rules. Evolved rules form scenario packages with a routing index, so each group retrieves only a bounded set appropriate to its situation without expert intervention. On 18 test episodes across three genres, SAGE scored 77.8 on a rubric validated by experts, versus 77.1 for professional directors. Deployed for 14 days on Virtual Film Studio, SAGE produced 1,344 narrative group outputs; 87.2 percent were accepted without substantive edits, and the production team recorded over 83 percent less authoring time per episode. We release PROSE, the first public dataset pairing screenplays with storyboards by professional directors across 68 episodes: https://github.com/creDreams/PROSE.
摘要:故事板將劇本轉化為自動化短劇製作的視覺拍攝計劃。專業的故事板製作依賴於隱性導演專業知識,並且仍然是產業瓶頸。大型語言模型可以自動化這一步驟,但提供導演知識的方法面臨三個挑戰:(1)知識獲取:這項技藝仍然隱含於範例中或必須手動撰寫。(2)知識精煉:創作的知識未能根據執行結果進行評估,且不透明的生成過程阻礙了對每個決策背後知識的反饋歸因。(3)知識注入:注入所有知識超出了可用的上下文,而對每個敘事群體進行手動選擇則無法擴展。我們提出了SAGE(具歸因引導演變的技能),這是一個已部署的框架,從專家示範中學習、歸因、演變和路由導演知識。SAGE通過將每個訓練劇本與其專家故事板進行對比,推導出獨立於劇集內容的規則。在生成過程中,模型記錄每個敘事群體採用的規則。將這些記錄與本地反饋結合,使得對個別規則的針對性更新成為可能。演變的規則形成具有路由索引的場景包,因此每個群體僅檢索適合其情境的有限集合,而無需專家介入。在三個類型的18個測試劇集中,SAGE在專家驗證的評分標準上得分77.8,而專業導演則為77.1。在虛擬電影工作室部署14天後,SAGE產出了1,344個敘事群體的輸出;87.2%的輸出在未進行實質性編輯的情況下被接受,製作團隊每集的創作時間減少了超過83%。我們發布了PROSE,這是第一個將劇本與專業導演的故事板配對的公共數據集,涵蓋68個劇集:https://github.com/creDreams/PROSE。
From Entity Mentions to Tone: An LLM-Based Pipeline for Media Bias Analysis
2608.17454v1 by Klesti Hoxha, Olti Qirici
This paper presents a pipeline for analyzing media bias and framing in online news. The pipeline groups articles into topics and events, adds named-entity and sentiment annotations, and compares news sources through people mentions, source-level tone, and event-level coverage patterns. We apply it to 8,358 Albanian news articles collected from GDELT and compare the resulting annotations with GDELT's automated annotations. The results show moderate agreement for sentiment and entity extraction, as well as additional person-entity pairs that can potentially support the bias analysis. We compare two annotation prompts and find that stricter sentiment-validation rules remove label-score inconsistencies but increase execution time and reduce annotation coverage. Based on these results, the simpler prompt is used for the rest of the analysis. We have provided sample analysis on source-level framing pro les, person-level tone differences across sources, and event-level gatekeeping and coverage indicators. These outputs show how the same news collection can be used to examine what sources cover, how they describe public figures, and where coverage is concentrated. The approach is particularly useful in settings where manually verified datasets or specialized language tools are limited.
摘要:這篇論文提出了一個用於分析在線新聞中的媒體偏見和框架的流程。
該流程將文章分組為主題和事件,添加命名實體和情感註釋,並通過人名提及、來源層級語調和事件層級報導模式來比較新聞來源。
我們將其應用於從GDELT收集的8,358篇阿爾巴尼亞新聞文章,並將結果註釋與GDELT的自動註釋進行比較。
結果顯示情感和實體提取之間有中等程度的一致性,以及額外的人物-實體對,這些對可能支持偏見分析。
我們比較了兩個註釋提示,發現更嚴格的情感驗證規則消除了標籤分數不一致,但增加了執行時間並減少了註釋覆蓋率。
根據這些結果,簡單的提示被用於後續的分析。
我們提供了來源層級框架概況、不同來源之間的人物層級語調差異,以及事件層級的把關和報導指標的樣本分析。
這些輸出顯示了相同的新聞集合如何用來檢查哪些來源進行報導、他們如何描述公眾人物,以及報導的集中地點。
這種方法在手動驗證數據集或專業語言工具有限的情況下特別有用。
Decomposition Attacks Across Unlinkable Identities: Limits of Stateful Defenses for LLM Services
2608.17445v1 by Bowen Sun, Zhengyue Zhao, Xiaogeng Liu, Yinzhi Cao, Chaowei Xiao
Most large language model services use stateless defenses, which judge only the current request, to refuse harmful tasks. Decomposition attacks exploit this limitation by splitting a harmful task into individually permissible requests and combining their answers. Defending against them therefore requires a stateful monitor that considers requests together. If it can group all requests for one attacker task, it can stop the attack. However, attackers can use unlinkable identities and combine answers elsewhere, leaving no reliable grouping signal. We ask whether decomposition attacks can still be stopped under this setting. For a fixed attack strategy without retries, we prove that the achievable security and utility tradeoff depends entirely on how benign requests for the same capabilities are grouped. Persistent, recognizable groups permit a useful defense; fresh, indistinguishable groups do not. When attackers can retry and learn from Allow/Block decisions, this useful operating point disappears: the feedback reveals what passes but not whether a block was correct. Experiments on 91 executable tasks and 11,393 capability-matched benign requests support these results. Under a 1% denial cap for these requests and a 0.5% cap for unrelated background traffic, all ten tested policies, including one privileged policy with an exact request-to-operation map, either fail to stop attacks or exceed the budget. On defense-unseen task families, attack success is at least 99% after one attempt and 100% after two. Effective defenses therefore require additional evidence or mechanisms tied to grouping, such as reliable identity linkage, costs for fresh identities, or control over answer use.
摘要:大多數大型語言模型服務使用無狀態防禦,只根據當前請求來拒絕有害任務。分解攻擊利用了這一限制,通過將有害任務拆分為單獨可允許的請求並結合它們的答案來進行攻擊。因此,防禦這些攻擊需要一個有狀態的監控器,能夠將請求一起考慮。如果它能夠將所有針對一個攻擊者任務的請求分組,就能夠阻止攻擊。然而,攻擊者可以使用不可鏈接的身份並在其他地方結合答案,這樣就沒有可靠的分組信號。我們詢問在這種情況下是否仍然可以阻止分解攻擊。對於沒有重試的固定攻擊策略,我們證明可實現的安全性和效用權衡完全取決於對相同能力的良性請求如何分組。持久且可識別的群體允許有效的防禦;新鮮且不可區分的群體則不允許。當攻擊者可以重試並從允許/阻止決策中學習時,這一有用的操作點消失了:反饋揭示了哪些請求通過,但並不顯示阻止是否正確。對91個可執行任務和11,393個能力匹配的良性請求的實驗支持了這些結果。在這些請求的1%拒絕上限和與之無關的背景流量的0.5%上限下,所有十個測試的政策,包括一個具有精確請求到操作映射的特權政策,要麼未能阻止攻擊,要麼超出預算。在未見防禦的任務家族中,攻擊成功率在一次嘗試後至少為99%,在兩次嘗試後為100%。因此,有效的防禦需要額外的證據或與分組相關的機制,例如可靠的身份鏈接、新身份的成本或對答案使用的控制。
Structure-Internalized Rule Language Model for Faithful Knowledge Graph Reasoning
2608.17443v1 by Xingrui Zhuo, Jiapu Wang, Manzong Huang, Gongqing Wu, Xindong Wu
Knowledge Graph Reasoning (KGR) aims to discover latent facts by leveraging the structural evidence available in KGs, posing a challenge to the structural semantic understanding capability of KGR models. Recent studies have demonstrated that Large Language Models (LLMs) can achieve remarkable progress on KGR tasks via flexible in-context learning. However, the inherent representation inconsistency between KG structural context and LLM parametric knowledge remains inadequately addressed. This limitation prevents LLMs from effectively perceiving reasoning evidence that aligns with KG constraints, which undermines both the effectiveness and faithfulness of reasoning. We refer to this problem as reasoning evidence perception drift of LLMs over KGs. To address this problem, we propose a Structure-Internalized Rule Language Model (SIRLM), which centers on structural rule generation to couple the parametric learning of structural knowledge with the faithfulness evaluation of reasoning logic, enabling LLMs to anchor tightly to KG-grounded evidence. Specifically, we first design a Structure-Internalized Rule Generator (SIRG), which incorporates an in-context learning block augmented with a structural relation memory to coordinate structural and parametric knowledge. Furthermore, we equip SIRG with a KG tokenizer based on structural invariance learning and a neuro-symbolic reasoner based on rule-constrained message propagation. These components provide SIRG with learnable structural representations and faithful rule-execution feedback, respectively. Our SIRLM can be seamlessly integrated into standard LLM training paradigms, such as SFT and GRPO. Extensive experiments against 17 state-of-the-art KGR methods on 36 datasets demonstrate the significant superiority of SIRLM.
摘要:知識圖譜推理(KGR)旨在利用知識圖譜中的結構證據來發現潛在事實,這對KGR模型的結構語義理解能力提出了挑戰。最近的研究表明,大型語言模型(LLMs)可以通過靈活的上下文學習在KGR任務上取得顯著進展。然而,知識圖譜的結構上下文與LLM的參數知識之間固有的表示不一致性仍然未得到充分解決。這一限制阻礙了LLMs有效感知與知識圖譜約束相符的推理證據,從而削弱了推理的有效性和可靠性。我們將這個問題稱為LLMs在知識圖譜上的推理證據感知漂移。為了解決這個問題,我們提出了一種結構內化規則語言模型(SIRLM),該模型專注於結構規則生成,以將結構知識的參數學習與推理邏輯的可靠性評估相結合,使LLMs能夠緊密依賴於知識圖譜基礎的證據。具體而言,我們首先設計了一個結構內化規則生成器(SIRG),該生成器包含一個增強了結構關係記憶的上下文學習模塊,以協調結構和參數知識。此外,我們為SIRG配備了一個基於結構不變性學習的知識圖譜標記器和一個基於規則約束消息傳播的神經符號推理器。這些組件分別為SIRG提供了可學習的結構表示和可靠的規則執行反饋。我們的SIRLM可以無縫集成到標準的LLM訓練範式中,如SFT和GRPO。在36個數據集上對17種最先進的KGR方法進行的廣泛實驗顯示了SIRLM的顯著優越性。
Task-Aware Harness Provisioning for LLM Agents in Mission-Critical Infrastructure Operations
2608.17433v1 by Liangtao Lin, Qingang Zhang, Zhaomeng Zhu, Tianwei Zhang, Yonggang Wen
LLM agents have been widely adopted to operate mission-critical infrastructure (MCI). These agents normally rely on a harness that determines what information they can access, which tools they can use, and what actions they can take. Existing systems often expose the same comprehensive harness to every task, which may not be necessary and cause resource wastes. In this paper, we focus on the identification of optimal harness configurations, and view it as a resource-matching problem between what each task requires and what the harness provides. To measure this match, we classify MCI tasks based on the mathematical representation of the underlying system and rank harness configurations by the amount and type of information they provide. We then construct task-to-harness mappings from two sources: mining research literature and measuring controlled agent execution. Leveraging the measured mapping, we propose a new harness provisioning algorithm: map-guided escalation. It begins with a task-specific harness and expands to full provision only after a failed self-check. We evaluate our method in two representative MCI tasks: in liquid cooling, it improves the agent accuracy from 0.652 under full provision to 0.715 and achieves accuracy comparable to Reflexion with 48% fewer tokens; In power grids, full provision remains accuracy-optimal, while map-based provisioning offers lower-cost alternatives. These findings show that harness provisioning follows a domain-dependent accuracy-cost Pareto frontier rather than a universal optimum.
摘要:LLM 代理已被廣泛應用於運營關鍵任務基礎設施 (MCI)。這些代理通常依賴於一個裝置,該裝置決定它們可以訪問哪些信息、可以使用哪些工具以及可以採取哪些行動。現有系統通常對每個任務暴露相同的全面裝置,這可能不是必要的,並導致資源浪費。在本文中,我們專注於最佳裝置配置的識別,並將其視為每個任務所需與裝置提供之間的資源匹配問題。為了衡量這種匹配,我們根據基礎系統的數學表示對 MCI 任務進行分類,並根據它們提供的信息的數量和類型對裝置配置進行排名。然後,我們從兩個來源構建任務到裝置的映射:挖掘研究文獻和測量受控代理執行。利用測量的映射,我們提出了一種新的裝置供應算法:基於映射的升級。它從特定任務的裝置開始,僅在自我檢查失敗後擴展到完全供應。我們在兩個具代表性的 MCI 任務中評估我們的方法:在液體冷卻中,它將代理的準確度從完全供應下的 0.652 提高到 0.715,並且在使用 48% 更少的標記的情況下達到與 Reflexion 相當的準確度;在電力網中,完全供應仍然是準確度最佳,而基於地圖的供應則提供了更低成本的替代方案。這些發現表明,裝置供應遵循一個依賴於領域的準確度-成本 Pareto 邊界,而不是一個普遍的最優解。
SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation
2608.17426v1 by Keyu Tu, Zhuowei Chen, Mengqi Huang, Yuxin Wang, Jiahao Zhu, Zhendong Mao, Yongdong Zhang
We introduce Semantic Task Completion Video Generation, an outcome-oriented video generation task. Under this formulation, success requires both achievement of the intended outcome and semantic grounding. Semantic grounding characterizes the correspondence between the reference image and the generated outcome in terms of high-level semantics relevant to the task. Evaluation focuses on the generated outcome and requires neither the presentation of a complete sequence of intermediate task steps nor conventional appearance consistency with the reference image. To support systematic evaluation, we construct SemComp-Data, an evaluation dataset covering six domains. Each instance comprises a reference image, a detailed instruction, a brief instruction, and an outcome-centric video clip. A scalable four-stage curation pipeline converts raw videos into standardized SemComp-Data instances. We further introduce SemComp-Bench, an evaluation protocol that uses a vision-language model (VLM) to answer structured binary questions. SemComp-Bench reports the OA Score and the GR Score for Outcome Achievement and Generation Reliability, respectively. Experiments on representative video generation models show that achieving intended outcomes while maintaining task-relevant semantic grounding in reference images remains challenging.
摘要:我們介紹了語義任務完成視頻生成,這是一項以結果為導向的視頻生成任務。根據這一表述,成功需要同時實現預期的結果和語義基礎。語義基礎描述了參考圖像與生成結果之間在與任務相關的高階語義方面的對應關係。評估重點在於生成的結果,並不需要呈現完整的中間任務步驟序列,也不需要與參考圖像的常規外觀一致性。為了支持系統化評估,我們構建了SemComp-Data,一個涵蓋六個領域的評估數據集。每個實例包括一個參考圖像、一個詳細說明、一個簡要說明和一個以結果為中心的視頻片段。一個可擴展的四階段策展管道將原始視頻轉換為標準化的SemComp-Data實例。我們進一步介紹了SemComp-Bench,一個使用視覺-語言模型(VLM)回答結構化二元問題的評估協議。SemComp-Bench報告了結果達成的OA分數和生成可靠性的GR分數。對代表性視頻生成模型的實驗顯示,在保持與參考圖像相關的語義基礎的同時實現預期結果仍然具有挑戰性。
An Investigation of Translationese in the Generations of Multilingual Large Language Models
2608.17399v1 by Maria Valentini, Téa Wright, Julisa Granados, Eliana Colunga, Katharina von der Wense
Text which has been translated from another language tends to carry with it evidence of translation$\unicode{x2014}$hence, it is often referred to as $\textit{translationese}$. Multilingual large language models (MLLMs) generate text in a variety of languages. However, it is still unclear if MLLMs' generations resemble internal translation (from English or, potentially, other languages) and, thus, result in translationese. Here, we ask the following research questions: (1) Does text generated by MLLMs resemble translationese? (2) How does translationese produced by MLLMs differ from translationese produced through direct translation? We leverage established indicators of translated text to evaluate text generated by state-of-the-art MLLMs in five languages, comparing to both non-translated and human-written baselines in order to isolate translationese from other kinds of interference. Through the use of high-accuracy classification models, analyses of variance on individual linguistic features, and the collection of human annotations in a subset of two languages (German and Spanish), we assess the translationese content of MLLM generations and examine the key features that distinguish MLLM-generated text from typical translation-related interference.
摘要:從另一種語言翻譯過來的文本往往帶有翻譯的痕跡$\unicode{x2014}$因此,它通常被稱為$\textit{translationese}$。多語言大型語言模型(MLLMs)可以生成多種語言的文本。然而,目前尚不清楚MLLMs生成的文本是否類似於內部翻譯(從英語或潛在的其他語言),因此是否會導致翻譯語。 在這裡,我們提出以下研究問題:(1)MLLMs生成的文本是否類似於翻譯語?(2)MLLMs產生的翻譯語與通過直接翻譯產生的翻譯語有何不同?我們利用已建立的翻譯文本指標來評估五種語言中最先進的MLLMs生成的文本,並與非翻譯文本和人類撰寫的基準進行比較,以便將翻譯語與其他類型的干擾區分開來。通過使用高準確度的分類模型、對個別語言特徵的變異分析,以及在兩種語言(德語和西班牙語)子集中的人類標註收集,我們評估MLLM生成的翻譯語內容,並檢查區分MLLM生成文本與典型翻譯相關干擾的關鍵特徵。
LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents
2608.17393v1 by Yiming Du, Yuxin Jiang, Tao Yuan, Jianbo Dai, Shaowei Wang, Jierun Chen, Chaofan Tao, Xianzhi Yu, Lifeng Shang, Kam-Fai Wong, Xiaohui Li, Haoli Bai
Reinforcement learning for coding agents increasingly relies on long-running agent harnesses to manage tool integration, repository contexts, and execution feedback. However, the native execution environments of these harnesses are inherently misaligned with policy-gradient training: environmental crashes and reward hacking corrupt outcome signals, while train-inference discrepancies decouple rollout behavior from policy updates. To address this, we present LEGO-RL, a framework that bridges native coding-agent harnesses with scalable policy-gradient optimization without modifying their internal control flow. LEGO-RL is built upon three pillars: (1) faithful optimization via in-process LLM proxying that captures raw generation streams for token-level alignment and robust trainer-side log-probability recomputation, even under harness-side compaction or re-serialization; (2) reliable execution via scalable sandbox orchestration featuring image caching and stage-wise defenses to mitigate reward hacking; and (3) observable training through an integrated plugin that automates validation and monitoring, paired with a Live UI for granular trajectory diagnostics. We evaluate LEGO-RL by training the sparse MoE model Qwen3.5-35B-A3B with GSPO across three native coding-agent harnesses. LEGO-RL improves Qwen3.5-35B-A3B across OpenHands SDK (64.0% to 70.4%), Claude Code (62.4% to 68.2%), and OpenCode (57.2% to 66.6%) on SWE-bench Verified, while maintaining a rollout-training probability correlation above 0.99.
摘要:強化學習對於編碼代理的依賴越來越多,尤其是在長期運行的代理工具中,以管理工具整合、庫上下文和執行反饋。然而,這些工具的本地執行環境與策略梯度訓練本質上不一致:環境崩潰和獎勵駭客會破壞結果信號,而訓練與推理之間的差異使得回滾行為與策略更新脫節。為了解決這個問題,我們提出了LEGO-RL,一個將本地編碼代理工具與可擴展的策略梯度優化相連接的框架,而無需修改其內部控制流程。LEGO-RL建立在三個支柱之上:(1)通過進程內LLM代理捕獲原始生成流以實現令牌級對齊的忠實優化,以及在工具端壓縮或重新序列化下的穩健訓練方日志概率重新計算;(2)通過可擴展的沙盒編排實現可靠的執行,特徵包括圖像緩存和階段防禦,以減輕獎勵駭客的影響;(3)通過集成插件實現可觀察的訓練,自動化驗證和監控,並配備用於細粒度軌跡診斷的實時用戶界面。我們通過在三個本地編碼代理工具上使用GSPO訓練稀疏的MoE模型Qwen3.5-35B-A3B來評估LEGO-RL。LEGO-RL在SWE-bench Verified上改善了Qwen3.5-35B-A3B在OpenHands SDK(從64.0%提升至70.4%)、Claude Code(從62.4%提升至68.2%)和OpenCode(從57.2%提升至66.6%)的表現,同時保持回滾訓練概率的相關性高於0.99。
Leveraging generative hallucination and biophysics-informed modeling for unified biomolecular sequence-structure co-design
2608.17381v1 by Xuefeng Liu, Mingxuan Cao, Xiao Luo, Songhao Jiang, Tobin Sosnick, Jinbo Xu, Louis Maher, Rick Stevens
Biomolecular design underpins applications from molecular recognition to therapeutics and synthetic biology, yet de novo interaction design remains challenging-especially for DNA/RNA, underexplored non-protein modalities with scarce, heterogeneous complex data and sharper geometric and chemical constraints. We introduce MCTH (Monte Carlo Tree Hallucination), an inference-only framework that casts all-atom sequence-structure co-design as uncertainty-aware planning over hallucinated states from pretrained folding and inverse-folding models, with optional biophysical control within the same decision loop. MCTH treats these models as frozen black-box operators and uses Monte Carlo Tree Search to allocate a fixed inference budget across competing design trajectories, incorporating model confidence and uncertainty, as well as cross-expert consensus/disagreement when multiple predictors are available. Across protein-RNA, protein-DNA, protein-protein, and protein-ligand design, matched-budget experiments show that adaptive search improves over simpler sampling and cycling strategies, while held-out AlphaFold3 and Chai-1 evaluations demonstrate transfer beyond the search-time oracle. MCTH provides a shared planning layer across modalities while allowing task-specific folding, inverse-folding, and biophysical modules, requiring no fine-tuning or backpropagation through component models.
摘要:生物分子設計支撐著從分子識別到治療和合成生物學的應用,然而,從零開始的互動設計仍然具有挑戰性,尤其是對於DNA/RNA這些未被充分探索的非蛋白質模式,其複雜數據稀少且異質,並且面臨更嚴格的幾何和化學限制。我們介紹了MCTH(蒙特卡羅樹幻覺),這是一個僅用於推理的框架,將全原子序列結構共同設計視為對從預訓練的摺疊和反摺疊模型中幻覺狀態的帶有不確定性意識的規劃,並在同一決策循環中可選擇生物物理控制。MCTH將這些模型視為凍結的黑箱運算符,並使用蒙特卡羅樹搜索在競爭設計軌跡中分配固定的推理預算,納入模型信心和不確定性,以及當多個預測器可用時的跨專家共識/分歧。在蛋白質-RNA、蛋白質-DNA、蛋白質-蛋白質和蛋白質-配體設計中,匹配預算的實驗顯示,自適應搜索優於更簡單的採樣和循環策略,而持出來的AlphaFold3和Chai-1評估則顯示出超越搜索時間神諭的轉移。MCTH在不同模式之間提供了一個共享的規劃層,同時允許特定任務的摺疊、反摺疊和生物物理模塊,無需對組件模型進行微調或反向傳播。
PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX
2608.17379v1 by Genghan Zhang, Yixin Dong, Chengze Fan, Zhichen Zeng, Yueming Yuan, Shaowei Zhu, Kunle Olukotun
We introduce PTXBench, a benchmark for evaluating and adapting large language models (LLMs) to use architecture-specific PTX for GPU kernel optimization. PTXBench measures functional correctness, whether selected target instructions execute at runtime, and speedup over frontier libraries across GEMM and attention workloads on H100 and B200 GPUs. Our evaluation shows that architecture-specific PTX capability remains uneven: success rates fall substantially on complex attention backward workloads, and executing the target instructions does not necessarily translate into competitive performance. No evaluated model consistently matches frontier libraries across the suite. We further adapt Qwen3.6-27B using supervised fine-tuning. Repair-conditioned training improves several tasks, but generalization remains uneven; data coverage, balance, and the quality of the reasoning teacher matter in addition to dataset size. PTXBench provides an auditable testbed for measuring and improving LLMs' ability to exploit evolving GPU architectures.
摘要:我們介紹 PTXBench,這是一個用於評估和調整大型語言模型(LLMs)以使用特定架構的 PTX 進行 GPU 核心優化的基準測試。
PTXBench 測量功能正確性,檢查所選目標指令在運行時是否執行,以及在 H100 和 B200 GPU 上的 GEMM 和注意力工作負載中,相較於前沿庫的加速效果。
我們的評估顯示,特定架構的 PTX 能力仍然不均衡:在複雜的注意力反向工作負載上,成功率顯著下降,而執行目標指令不一定能轉化為具競爭力的性能。
在整個測試套件中,沒有任何評估模型能持續匹配前沿庫。
我們進一步使用監督微調調整 Qwen3.6-27B。
修復條件訓練改善了幾個任務,但泛化能力仍然不均衡;數據覆蓋、平衡以及推理教師的質量在數據集大小之外也很重要。
PTXBench 提供了一個可審核的測試平台,用於測量和改善 LLM 利用不斷演變的 GPU 架構的能力。
Fair ASR: Re-Evaluating Black-Box Jailbreaks under Shared Target-Call Budgets
2608.17360v1 by Zhida He, Xiaoyu Wen, Han Qi, Ziyuan Zhou, Peng Yu, Jiajia Li, Chaochao Lu, Qiaosheng Zhang
Reliable jailbreak evaluation is essential for assessing LLM safety, but most existing studies rely solely on attack success rate (ASR) without accounting for its dependence on attack budgets, resulting in unfair comparisons across methods. Existing compute-aware evaluations reduce heterogeneous resources into FLOPs, which is difficult to estimate for black-box models and fails to capture resource-specific constraints. To provide a comparable evaluation basis, we introduce Fair-ASR, an evaluation protocol for black-box jailbreak attacks under shared target-call budgets B, using target calls as a directly observable and method-agnostic comparison axis while tracking attacker calls separately for efficiency analysis. We re-evaluate 11 representative attacks under the Fair-ASR protocol and find that attack rankings change substantially across target-call budgets, simple stochastic perturbations and hand-crafted templates remain highly competitive under equal target access, and no evaluated LLM-driven method is efficient in both target and attacker calls. Motivated by this efficiency gap, we introduce ReCode, a compositional budget-efficient attack that combines desensitization rewriting with two effective low-cost primitives identified by Fair-ASR. Under a budget of 20 target calls, ReCode achieves 85% ASR on GPT-5 while requiring only 7.19 attacker calls per request on average, showing strong efficiency in both target and attacker calls.
摘要:可靠的越獄評估對於評估大型語言模型(LLM)的安全性至關重要,但現有的大多數研究僅依賴攻擊成功率(ASR),而未考慮其對攻擊預算的依賴,導致不同方法之間的比較不公平。現有的計算感知評估將異質資源簡化為FLOPs,這對於黑箱模型來說難以估算,並且未能捕捉資源特定的限制。為了提供可比較的評估基礎,我們引入了Fair-ASR,這是一種針對共享目標調用預算B的黑箱越獄攻擊的評估協議,利用目標調用作為直接可觀察且與方法無關的比較軸,同時單獨跟踪攻擊者調用以進行效率分析。我們在Fair-ASR協議下重新評估了11個代表性攻擊,發現攻擊排名在不同的目標調用預算下有顯著變化,簡單的隨機擾動和手工製作的模板在平等的目標訪問下仍然具有高度競爭力,且沒有評估的LLM驅動方法在目標和攻擊者調用中都有效率。受到這一效率差距的激勵,我們引入了ReCode,一種組合預算高效的攻擊,結合了去敏感化重寫和Fair-ASR識別的兩個有效低成本原語。在20次目標調用的預算下,ReCode在GPT-5上達到85%的ASR,同時每次請求平均僅需7.19次攻擊者調用,顯示出在目標和攻擊者調用中都具有強大的效率。
Cognitive Graph Intelligence for Adaptive and Robust DDoS Attack Detection in Next Generation Networks
2608.17352v1 by Mohammad Arif Hossain, Yeahia Sarker, Md Jafrin Hossain, Most. Humayra Khanom Rime, Nirwan Ansari
Distributed Denial-of-Service (DDoS) attacks threaten network availability, requiring a cognitive detection process that senses traffic, infers intent, and supports an adaptive response under severe class imbalance and non-stationary conditions. This paper proposes a Graph-based Generative Adversarial Network (GraphGAN) that serves as the cognitive detection engine for this task. GraphGAN captures the relational structure among traffic flows while addressing imbalance through adversarial generation of synthetic samples. Sequential flows are converted into $k$-nearest neighbor graphs using sliding windows to preserve feature-similarity and temporal dependencies among flows. The generator learns the distribution of DDoS attacks to synthesize realistic minority samples, while a Graph Convolutional Network (GCN)-based discriminator distinguishes real from synthetic graph data. A separate GCN classifier, trained on the balanced dataset, performs the final detection decision. Evaluations on four benchmark datasets show that GraphGAN achieves superior accuracy, precision, and recall compared to state-of-the-art approaches, particularly in data-scarce scenarios. By integrating temporal graph construction, adversarial augmentation, and GCN classification, GraphGAN effectively models coordinated attack behaviors and mitigates class imbalance, providing a robust and topology-aware solution for intrusion detection in data-constrained environments.
摘要:分散式拒絕服務(DDoS)攻擊威脅網絡可用性,這需要一個認知檢測過程來感知流量、推斷意圖,並在嚴重的類別不平衡和非穩態條件下支持自適應響應。本文提出了一種基於圖的生成對抗網絡(GraphGAN),作為此任務的認知檢測引擎。GraphGAN 捕捉流量流之間的關係結構,同時通過對抗生成合成樣本來解決不平衡問題。連續流量被轉換為 $k$-最近鄰圖,使用滑動窗口來保留流量之間的特徵相似性和時間依賴性。生成器學習 DDoS 攻擊的分佈,以合成現實的少數樣本,而基於圖卷積網絡(GCN)的判別器則區分真實與合成的圖數據。另一個在平衡數據集上訓練的 GCN 分類器執行最終檢測決策。在四個基準數據集上的評估顯示,GraphGAN 在準確性、精確度和召回率方面優於最先進的方法,特別是在數據稀缺的情況下。通過整合時間圖構建、對抗增強和 GCN 分類,GraphGAN 有效地建模協調攻擊行為並減輕類別不平衡,為數據受限環境中的入侵檢測提供了一個強健且考慮拓撲的解決方案。
MoFE: A Novel Mixture-of-Experts Framework with Fourier Neural Operators for Cryptocurrency Forecasting
2608.17342v1 by Bowen Liu, Mingming Sun
Forecasting cryptocurrency prices remains a formidable challenge due to inherent non-stationarity, abrupt regime shifts, and multi-scale stochastic dependencies. Conventional deep learning models often struggle to capture complex underlying dynamics, frequently resulting in persistent phase-lagged predictions. To address these limitations, we propose MoFE, a novel deep learning framework that integrates Fourier Neural Operators (FNOs) within a Mixture-of-Experts (MoE) architecture. Rooted in the theoretical framework of stochastic differential equations, MoFE conceptualizes cryptocurrency volatility as a superposition of multi-frequency components, which includes user network based fundamental growth, mining costs and halving mechanism caused seasonal volatility, and market sentiment-induced chaos. Specifically, specialized adaptive FNO (AFNO) and Convolution dual-domain experts learn continuous function-to-function mappings to encapsulate global spectral trends, cyclical adjustments and microstructures, while a dynamic gating based MoE mechanism enables adaptive strategy switching across diverse market regimes. Extensive experiments on Bitcoin datasets spanning January 2020 to December 2025 demonstrate that MoFE achieves state-of-the-art (SOTA) performance in both T+1 and T+5 forecasting horizons. Notably, the model effectively mitigates the phase-lag effect, delivering superior Directional Accuracy (DA) and Information Coefficient (IC). In high-fidelity simulated trading environments, these predictive gains transfer into significant excess returns and robust risk-adjusted performance, characterized by a high Sharpe ratio.
摘要:預測加密貨幣價格仍然是一項艱巨的挑戰,因為其固有的非平穩性、突變的制度轉變以及多尺度隨機依賴性。傳統的深度學習模型往往難以捕捉複雜的潛在動態,經常導致持續的相位滯後預測。為了解決這些限制,我們提出了MoFE,一種新穎的深度學習框架,將傅立葉神經運算子(FNOs)整合到專家混合(MoE)架構中。MoFE根植於隨機微分方程的理論框架,將加密貨幣的波動性概念化為多頻率組件的疊加,這包括基於用戶網絡的基本增長、挖礦成本和因減半機制引起的季節性波動,以及市場情緒引發的混沌。具體而言,專門的自適應FNO(AFNO)和卷積雙域專家學習連續的函數到函數映射,以封裝全球光譜趨勢、周期性調整和微結構,而基於動態門控的MoE機制則使得在不同市場制度之間的自適應策略切換成為可能。對於2020年1月至2025年12月的比特幣數據集進行的廣泛實驗表明,MoFE在T+1和T+5預測範圍內均實現了最先進的(SOTA)性能。值得注意的是,該模型有效減輕了相位滯後效應,提供了優越的方向準確性(DA)和信息係數(IC)。在高保真模擬交易環境中,這些預測增益轉化為顯著的超額回報和穩健的風險調整表現,特徵是高夏普比率。
LLM-Only PDDL Domain Repair with Open-Weight Models
2608.17341v1 by Nader Karimi Bavandpour, Pascal Bercher
AI planning is concerned with finding a sequence of actions that achieves a specified goal. It relies on explicit models of the world, commonly represented in the Planning Domain Definition Language (PDDL). An active line of research investigates how errors in such models can be detected and repaired. For example, users may provide positive test plans that are solutions, and negative test plans that fail during execution. Automated repair methods then modify the PDDL model to satisfy these constraints. In this paper, we evaluate the ability of recent open-weight large language models to perform this repair task using an LLM-only approach. Our experiments show that the symbolic baseline achieves an $F_1$ score of $.49$, while the best-performing LLM reaches $.87$ with high reasoning effort, an absolute improvement of $.38$. However, that setting has a mean test pass rate of only $.82$, falling to $.06$ on the Thoughtful domain; even the best setting that includes the test traces reaches only $.92$. Thus, current open-weight models cannot guarantee satisfaction of the test constraints required for reliable automated model repair.
摘要:AI 規劃關注於找到一系列行動以達成特定目標。它依賴於對世界的明確模型,通常以規劃領域定義語言 (PDDL) 表示。一個活躍的研究方向探討如何檢測和修復這些模型中的錯誤。例如,使用者可能提供正面的測試計劃作為解決方案,以及在執行過程中失敗的負面測試計劃。自動修復方法隨後會修改 PDDL 模型以滿足這些約束。在本文中,我們評估最近的開放權重大型語言模型使用僅 LLM 方法執行此修復任務的能力。我們的實驗顯示,符號基準達到了 $.49$ 的 $F_1$ 分數,而表現最佳的 LLM 則在高推理努力下達到了 $.87$,絕對改善為 $.38$。然而,該設置的平均測試通過率僅為 $.82$,在 Thoughtful 領域下降至 $.06$;即使是包括測試痕跡的最佳設置也僅達到 $.92$。因此,目前的開放權重模型無法保證滿足可靠的自動模型修復所需的測試約束。
Medical explainable AI
| Publish Date | Title | Authors | Homepage | Code |
|---|---|---|---|---|
| 2026-08-18 | Can Large Language Models Explain Flight Safety Events? A Prior-Guided Semantic LLM-based Approach | Lu Xu et.al. | 2608.18017v1 | null |
| 2026-08-18 | Grading Needs a Rubric, Not Intelligence | Jhen-Ke Lin et.al. | 2608.17938v1 | null |
| 2026-08-18 | MotoSafety: Edge-AI with Learned Temporal Importance for Two-Wheeler Collision Risk Assessment Under Time Pressure | Sumit S. Shevtekar et.al. | 2608.17823v1 | null |
| 2026-08-18 | Communicating Credit Risk with Large Language Models: Evaluation of Explanations from Standard and Alternative Data-Based Models | Sahab Zandi et.al. | 2608.17715v1 | null |
| 2026-08-18 | Explainable AI-Powered Framework for Video-Based Skill Assessment in Cataract Surgery | Mohammad Javad Ahmadi et.al. | 2608.17522v1 | null |
| 2026-08-18 | Understanding Curriculum Learning in Large Language Models via Cross-Difficulty Optimization Dynamics | Zhikai Ding et.al. | 2608.17268v1 | null |
| 2026-08-17 | From Abductive Explanations to Global Logical Rules for Node Classification in SGCs | Bryan Lima Cavalcante et.al. | 2608.17103v1 | null |
| 2026-08-17 | AutoSR: Automatic Symbolic Regression by Searching Research States | Kejia Zhang et.al. | 2608.16876v1 | null |
| 2026-08-17 | Topological Attribution Distance (TAD): Revealing Segment-Level RAG Influence on LLM Output Geometry for Incident Log Analysis | Reza Fayyazi et.al. | 2608.16775v1 | null |
| 2026-08-17 | Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments | Adam Karvonen et.al. | 2608.16747v1 | null |
| 2026-08-17 | Unsupervised Anomaly Detection for Image Dataset Quality Assurance in Multi-Center Breast MRI | Chiara Tappermann et.al. | 2608.16725v1 | null |
| 2026-08-17 | Degradation-Aligned Self-Supervised Learning for State of Health Estimation of Lithium-Ion Batteries under Label Sparsity | Jiaqi Yao et.al. | 2608.16612v1 | null |
| 2026-08-17 | Physics of Agents: Statistical Mechanics Predicts Collective Behavior of AI Agents | Batu El et.al. | 2608.16578v1 | null |
| 2026-08-17 | Revisiting the Performance of Generative Artificial Intelligence on Introductory Object-Oriented Programming Assessments: Insights from 2026 | Marina Lepp et.al. | 2608.16318v1 | null |
| 2026-08-17 | Foresight-England: Development of a National-Scale Generative AI Model of Electronic Health Records for Medical Event Prediction across the COVID-19 Pandemic | Simon Ellershaw et.al. | 2608.16273v1 | null |
| 2026-08-17 | Defake-o3: From Speculative Rationales to Verifiable Evidence for Explainable AIGI Detection | Bowen Deng et.al. | 2608.16259v1 | null |
| 2026-08-17 | CompoSkill: Compositional Skill Chain Attacks from Individually Scanner-Passing LLM Agent Skills | Mingxiao Liu et.al. | 2608.16246v1 | null |
| 2026-08-17 | When Single-Dataset Conclusions Fail: A 45-Task Study of Threshold Tuning and Resampling for Imbalanced Classification | Diyorbek Musaev et.al. | 2608.16147v1 | null |
| 2026-08-17 | TokenSTFormer: A Tokenized Spatial-temporal Attention Model for Holistic Motion Analysis in Adolescent Idiopathic Scoliosis Screening | Dong Chen et.al. | 2608.16122v1 | null |
| 2026-08-17 | Eigenanalysis framework for autoregressive neural emulators of multi-scale chaotic dynamics | Conrad Ainslie et.al. | 2608.16084v1 | null |
| 2026-08-17 | NICE: Scale-Stable Perturbations for Graph Neural Network Explanations via Noise Corruption | Ziluowen Luo et.al. | 2608.16038v2 | null |
| 2026-08-16 | Identifying Confusion Trends in Concept-based XAI for Multi-Label Classification | Haadia Amjad et.al. | 2608.15731v1 | null |
| 2026-08-16 | Large Models for Small Devices: Recent Advances and Empirical Analysis of Edge AI Deployment | Subhransu Das et.al. | 2608.15693v1 | null |
| 2026-08-15 | NumerosityVLM: A Cognitively Inspired Benchmark for Interpreting Numerosity Representations in Vision-Language Models | Yiming Fu et.al. | 2608.15425v1 | null |
| 2026-08-15 | ETHOS: Towards a Modular Ethics Framework for Clinical Multi-Agent Systems | Rakesh Sharma et.al. | 2608.15424v1 | null |
| 2026-08-15 | When AI Rewrites, Classifiers Relax: Uncertainty-Aware Sentiment Analysis on Sarcastic and AI-Paraphrased Social Text | Shresth Shroff et.al. | 2608.15338v1 | null |
| 2026-08-15 | Demographic Injection in Medical Language Models under Diversity, Equity, and Inclusion Prompts | Diego Mardian et.al. | 2608.15254v1 | null |
| 2026-08-15 | Low-Rank Dynamics-Effective Latent Carriers for Counterfactual Rollout in Learned World Models | Yang Liu et.al. | 2608.15156v1 | null |
| 2026-08-15 | Fast Test-Time Refinement for Robust Learned Image Compression | Jiaming Liang et.al. | 2608.15113v1 | null |
| 2026-08-15 | Command-Space Counterfactual Explanations for Pareto-Conditioned Reinforcement Learning | Joanikij Chulev et.al. | 2608.14963v1 | null |
| 2026-08-14 | Handover Analysis for Vehicular Communication with Explainability on the Fly | Ali Fuat Sahin et.al. | 2608.14820v1 | null |
| 2026-08-14 | Generated Context versus Governed State: Functional Conditions for Accountable Longitudinal Clinical Reasoning | Augusto Bernardo Pissarra et.al. | 2608.14804v1 | null |
| 2026-08-14 | Decoding the Past: An Uncertainty-Aware Deep Learning Framework for Sex Attribution in Prehistoric Hand Stencils | Karel Becerra et.al. | 2608.14539v1 | null |
| 2026-08-14 | NARRATE: A Multimodal Real-World Australian Driving Dataset for Human-Centred Explanations in Automated Driving | Ashkan Yousefi Zadeh et.al. | 2608.14767v1 | null |
| 2026-08-14 | Polaris : Multi Agentic System for Conversational Enterprise Analytics | Varuni H K et.al. | 2608.14246v1 | null |
| 2026-08-14 | CForce: Boosting Parallel Decoding for dLLMs via Consistency Forcing | Yuji Ren et.al. | 2608.13925v1 | null |
| 2026-08-13 | Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions | Qingfang Liu et.al. | 2608.13786v1 | null |
| 2026-08-13 | Capacity-Dependent Effects of Data Selection for Reasoning | Cuong Dang et.al. | 2608.13721v1 | null |
| 2026-08-13 | MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and Coordination | Saisha Shetty et.al. | 2608.13476v1 | null |
| 2026-08-13 | A Unifying Perspective on Causal World Models: From Observations to Representations to Structure | Avinash Kori et.al. | 2608.13456v1 | null |
| 2026-08-13 | Explanatory Engagement Under Rare Anomalous Failure: Asymptotic Rarity in Model Behavior (or: The Asymptotic AI) | Sam Mao et.al. | 2608.13063v1 | null |
| 2026-08-13 | VALG: An Agentic System for ML Theory Research | Dechen Zhang et.al. | 2608.13060v1 | null |
| 2026-08-13 | UniTraffic-Agent: Unified Traffic Video Reasoning for AI City Challenge 2026 Track 3 with Two Out-of-Domain Evaluations | Peng Li et.al. | 2608.13031v1 | null |
| 2026-08-13 | Static analysis-guided agentic AI translation enables Rust as a full stack bioinformatics language | Johan Henriksson et.al. | 2608.13029v1 | null |
| 2026-08-13 | Decomposition of Evidence, Contradiction, and Fragility in Perturbation Responses | Lei You et.al. | 2608.12935v1 | null |
| 2026-08-13 | Discovering Efficient and Explainable Communication Topologies for LLM-based Multi-Agent Systems via Causal Inference | Junzhi Li et.al. | 2608.12921v2 | null |
| 2026-08-13 | Mr3D-VL: A generalist vision language foundation model for Multiparametric 3D Magnetic Resonance Imaging | Zhi Qiao et.al. | 2608.12689v1 | null |
| 2026-08-12 | Interpretable Causal Discovery via Causal-Effect Constraints | Cixuan Zhang et.al. | 2608.12640v1 | null |
| 2026-08-12 | Algorithm Design and Physician Liability | Shujie Luan et.al. | 2608.13618v1 | null |
| 2026-08-12 | What Makes a Peer? Valuation-Anchored Similarity in Private Markets | Sebastian Frank et.al. | 2608.12594v1 | null |
| 2026-08-12 | Auditable agentic AI for evidence-grounded thyroid ultrasound diagnosis and reporting | Haifan Gong et.al. | 2608.12590v1 | null |
| 2026-08-12 | CAS: A Causal Attribution Score for Local and Global Explainable Artificial Intelligence | Michael Georgiades et.al. | 2608.12555v1 | null |
| 2026-08-12 | Class Activation Mapping in Explainable Computer Vision: A Method-Centered Review of CNN, Transformer, and Foundation-Model-Era Visual Explanations | AmirHossein Eshghi et.al. | 2608.12299v2 | null |
| 2026-08-12 | Dual Spatial-Temporal Attribution: Architecture-Aligned Post-Hoc Explainability for Recurrent Graph Anomaly Detection | Iyad Assaad Nekka et.al. | 2608.12441v1 | null |
| 2026-08-12 | Learning-Based Behavior Planning for Automated Driving: Real-World Integration and Deployment | Jean-Pierre Busch et.al. | 2608.12198v1 | null |
| 2026-08-12 | A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench | Praveen Reddy et.al. | 2608.12138v1 | null |
| 2026-08-12 | Do LLMs Take Care of Their Own? Similarity Signals Can Induce Cooperation | Akash Kundu et.al. | 2608.12125v1 | null |
| 2026-08-12 | Faithful, Sufficient and Understandable: Rethinking Graph Counterfactual Explanations via Discrete Diffusion Inversion | David Bechtoldt et.al. | 2608.12083v1 | null |
| 2026-08-12 | Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence | Mengru Wang et.al. | 2608.12036v1 | null |
| 2026-08-12 | From Safety Documentation to Safety Knowledge Support: An Evidence-Grounded LLM Framework for Medical Devices | Tuhinangshu Gangopadhyay et.al. | 2608.12025v1 | null |
| 2026-08-12 | Agent Skills Can Be Harmful: An Empirical Study of Skill-Induced Failures in LLM Agents | Gen Dong et.al. | 2608.11888v1 | null |
| 2026-08-12 | Low-Interaction-Rank Learning: Unifying Multiplicative Dual-Encoder Heads | Zijian Zhao et.al. | 2608.11661v1 | null |
| 2026-08-11 | Inverse Theory of Mind Modeling for Content Recommendation: From Web Browsing to Dynamic Intelligent Interfaces | Mengyu Chen et.al. | 2608.11354v1 | null |
| 2026-08-11 | Governing Agentic AI in FinTech | Henry Han et.al. | 2608.11344v2 | null |
| 2026-08-11 | From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop | Rahul Gupta et.al. | 2608.11171v1 | null |
| 2026-08-11 | SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure | Xiaofan Bai et.al. | 2608.11079v2 | null |
| 2026-08-11 | Entropy-Centric Explainable AI for Remote Sensing Image Segmentation | Ali Saleh et.al. | 2608.11064v1 | null |
| 2026-08-11 | ComBodied Agents: a New Paradigm of Human-Centric Agentic AI | Qianggang Ding et.al. | 2608.10915v2 | null |
| 2026-08-11 | Chemically Meaningful Textualization Enables Explainable Validation of Metal-Organic Frameworks by Large Language Models | Guobin Zhao et.al. | 2608.11283v1 | null |
| 2026-08-11 | Uncertainty-Aware and Explainable Ensemble Deep Learning Framework for Multi-Class Skin Lesion Classification | Rofiqul Islam et.al. | 2608.11280v1 | null |
| 2026-08-11 | Rule of Thumb: Explaining Artificial Intelligence Systems using Partial Information | Kaivalya Rawal et.al. | 2608.10766v2 | null |
| 2026-08-11 | Operationalising Relative Causal Knowledge: Backbone Identifiability from Private Reports on a Shared Outcome | Fabrizio Russo et.al. | 2608.10664v1 | null |
| 2026-08-11 | Conversational versus Dashboard Explainable AI for UAV Intrusion Detection: An Empirical Study of Operator Trust and Reliance | Cong Chi Nguyen et.al. | 2608.10434v1 | null |
| 2026-08-11 | Reasoning Shortcuts and Value Symmetries: What Symmetry Permits, Architecture Realizes, and Optimization Selects | Xin Xu et.al. | 2608.10420v1 | null |
| 2026-08-10 | Towards Expert-level Medical AI for Real-time Video Consultations | Mahvish Nagda et.al. | 2608.09861v1 | null |
| 2026-08-10 | CEAA: A Cognitive Embodied Agents Architecture for Interactive Computing Systems | Aimilios Hadjiliasi et.al. | 2608.09848v1 | null |
| 2026-08-10 | KGCaRe: Explainable Complex Conditional Question Answering using Automatic Knowledge Graph Construction and Context Retrieval with LLMs | Ghanshyam Verma et.al. | 2608.09779v1 | null |
| 2026-08-10 | Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models | Shulin Tian et.al. | 2608.09666v1 | null |
| 2026-08-10 | Rethinking Self-Evolving Agents: Do We Still Need Prescribed Optimization Pipelines? | Hui Xue et.al. | 2608.09629v1 | null |
| 2026-08-10 | Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs | Hongli Shen et.al. | 2608.09542v1 | null |
| 2026-08-10 | Renormalising Generative Models for Active Inference: Foundations, Derivations, and Verification | Karim Zaghw et.al. | 2608.09512v1 | null |
| 2026-08-10 | How Simple Can It Get? From Interpretable Equations to Readable Rules for Financial Decision Making | Adia Lumadjeng et.al. | 2608.09433v1 | null |
| 2026-08-10 | An Explainable GNN Framework for Component-Level Anomaly Diagnosis | Sena Ozgunay et.al. | 2608.09246v1 | null |
| 2026-08-10 | SafeSceneReason: A Multimodal Reasoning Benchmark Connecting Industrial Hazards with Accident Knowledge | Yuanchi Zhu et.al. | 2608.09230v1 | null |
| 2026-08-10 | TLDChoiceNet: Quantitatively Choosing a Transfer Learning Dataset | Jing Ning et.al. | 2608.09091v1 | null |
| 2026-08-09 | Reading is not Reasoning: Bridging the Agentic Policy Gap in Vision-Text Compression | Cheng Fan et.al. | 2608.08960v1 | null |
| 2026-08-09 | From Recovery to Drop-off: How Action Post-training Reduces a VLM's Late-Layer Depth Decodability | Alexander Hackett et.al. | 2608.08904v2 | null |
| 2026-08-09 | From Manuals to Maintenance: Fine-Tuning MedGemma for Multi-Modal Imaging System Support in Low-Resource Settings | Bernes Lorier Atabonfack et.al. | 2608.08896v1 | null |
| 2026-08-09 | PROSLEX: A Novel Dataset for Expert-Annotated Legal Statute Prediction for Indian Judiciary | Subinay Adhikary et.al. | 2608.08830v1 | null |
| 2026-08-09 | Deployable Per-Instance Multi-Layer Activation Steering for Large Language Models | Muhammad Faishal Adly Nelwan et.al. | 2608.08829v1 | null |
| 2026-08-09 | SymDiag: Explainable Diagnosis for LLM Reasoning via Neuro-Symbolic Verification | Wenyao Cui et.al. | 2608.08786v1 | null |
| 2026-08-09 | Domain Agnostic Text Redaction from Natural Language Rules using Instruction Tuning | Aravindhan Arunagiri et.al. | 2608.14693v1 | null |
| 2026-08-09 | Business Arena: Benchmarking LLM Agents in a Realistic Marketplace | Yijun Pan et.al. | 2608.08621v1 | null |
| 2026-08-09 | On-Device Multi-Species Malaria Detection with Uncertainty-Calibrated Slide-Level Aggregation | Idaya Seidu et.al. | 2608.08566v1 | null |
| 2026-08-09 | Private Etymology: Designing Relational Reuse of Shared Symbols in Long-Term Human-AI Interaction | Miki Ueno et.al. | 2608.08443v2 | null |
| 2026-08-08 | Quantization Degradation in Large Language Models: A Signal-Noise Perspective | Chenxi Zhou et.al. | 2608.08188v1 | null |
| 2026-08-08 | Agentic AI-driven Immersive Simulation: A Knowledge-Aware Virtual Training Platform forHigh Dose Rate (HDR) Brachytherapy | Ronghua Xu et.al. | 2608.08163v1 | null |
| 2026-08-08 | Compositional Threat Analysis of Latent Compromise in LLM Agent Systems: The Order 66 Scenario | Satoshi Matsuoka et.al. | 2608.08131v1 | null |
| 2026-08-08 | Defending Retrieval-Augmented Intrusion Detection Against Knowledge Poisoning and Prompt Injection | Kaysarul Anas Apurba et.al. | 2608.08100v1 | null |
| 2026-08-08 | HugSelect: An Explainable Multi-Criteria Decision-Support Framework for foundation-model selection | Alireza Joonbakhsh et.al. | 2608.08069v1 | null |
Abstracts
Can Large Language Models Explain Flight Safety Events? A Prior-Guided Semantic LLM-based Approach
2608.18017v1 by Lu Xu, Xu Li, Linjiang Zheng, Fan Li, Riquan Zhang, Jiaxing Shang
Improving flight safety with flight data requires not only accurate detection of risk events, but more importantly, clear interpretation of their underlying causes at the level of pilot control behavior. Existing explainable AI techniques, such as feature importance maps, often require considerable domain knowledge to translate them into operationally meaningful explanations. Large Language Models (LLMs), which excel at language reasoning, bring a promising solution to this issue. However, applying LLMs in this domain presents key challenges such as modal inconsistency, limited classification ability, scarcity of task-specific data for fine-tuning, and lack of domain knowledge. To overcome these challenges, we propose FlightLLM, a prior-guided semantic LLM-based approach for interpretable flight safety analysis. Specifically, we first perform feature engineering to address modal inconsistency, combining statistical descriptors with physically meaningful flight indicators. This representation is further processed by a Semantic Discretization module, which converts abstract numerical patterns into qualitative descriptions that are more compatible with language reasoning. In addition, since LLMs are not inherently strong classifiers, CatBoost is incorporated as a statistical expert, and its prediction results are injected into the prompt as prior guidance. A contrastive few-shot learning strategy is further adopted to compensate for limited data. Finally, we design structured prompts to embed aviation-specific knowledge into the inference process. Using hard landing, a representative risk event with complex causal mechanisms, as an anchor point, we evaluate FlightLLM on a dataset of 704 real-world A320 flight samples. Experimental results show that the proposed approach achieves competitive classification performance while generating direct and reasonable explanations for event causes.
摘要:改善飛行安全需要不僅準確檢測風險事件,更重要的是在飛行員控制行為層面清晰解釋其潛在原因。現有的可解釋AI技術,如特徵重要性圖,通常需要相當的領域知識才能將其轉化為具有操作意義的解釋。大型語言模型(LLMs)在語言推理方面表現出色,為這一問題帶來了有希望的解決方案。然而,在這一領域應用LLMs面臨著關鍵挑戰,如模式不一致、有限的分類能力、缺乏特定任務的數據以進行微調,以及缺乏領域知識。為了克服這些挑戰,我們提出了FlightLLM,一種基於語義的先驗引導LLM方法,用於可解釋的飛行安全分析。具體而言,我們首先進行特徵工程以解決模式不一致,將統計描述符與具有物理意義的飛行指標相結合。這一表示進一步由語義離散化模塊處理,將抽象的數字模式轉換為更符合語言推理的定性描述。此外,由於LLMs本身並不是強大的分類器,因此CatBoost被納入作為統計專家,其預測結果被注入到提示中作為先驗指導。進一步採用了對比少樣本學習策略以彌補數據的有限性。最後,我們設計了結構化提示,將航空特定知識嵌入推理過程中。以硬著陸作為錨點,這是一個具有複雜因果機制的代表性風險事件,我們在704個真實世界A320飛行樣本的數據集上評估FlightLLM。實驗結果表明,所提出的方法在生成事件原因的直接和合理解釋的同時,實現了具有競爭力的分類性能。
Grading Needs a Rubric, Not Intelligence
2608.17938v1 by Jhen-Ke Lin
Small language models can grade open-ended examination answers as reliably as substantially more expensive models when they grade against an explicit rubric. We test this claim as the design principle behind any-to-bench: a frontier model reads source documents once, at ingestion, to extract each question and its rubric; lower-cost models then perform all repeated grading work. We evaluate six cost-efficient model configurations from two model families at three reasoning-effort levels. Each configuration answers 24 open-ended examination questions, and each also grades every answer sheet three times, yielding 3,456 per-question grades. Scores depend overwhelmingly on the answer being graded: answer identity explains 95.6% of score variance, whereas judge identity explains only 0.2%. Raising a writer's reasoning effort moves earned scores by as much as 0.143 of full marks, while raising a judge's reasoning effort moves assigned scores by at most 0.006. Six frontier-tier judges, added as a check, reproduce these scores and are no more reliable as a panel. Two ablations then decompose the rubric on the same questions and answers. Removing its criteria and levels while keeping the official answer changes nothing measurable. Removing the official answer as well collapses reliability (ICC 0.888 to 0.628), inflates scores, and makes judge reasoning effort matter again. The rubric is what decouples grading from judge intelligence, and within the rubric the official answer does nearly all the work. We find no evidence of length preference or same-family preference under rubric-anchored grading.
摘要:小型語言模型在根據明確的評分標準進行評分時,可以與成本高得多的模型一樣可靠地評分開放式考試答案。我們測試這一主張,作為 any-to-bench 的設計原則:前沿模型在攝取時讀取源文件一次,以提取每個問題及其評分標準;然後,成本較低的模型執行所有重複的評分工作。我們在三個推理努力水平上評估來自兩個模型系列的六種成本效益模型配置。每個配置回答 24 道開放式考試問題,並且每個配置還對每份答案進行三次評分,產生每個問題 3,456 次評分。分數在很大程度上取決於被評分的答案:答案身份解釋了 95.6% 的分數變異,而評判身份僅解釋了 0.2%。提高寫作者的推理努力可以使得獲得的分數提高最多 0.143 的滿分,而提高評判的推理努力則最多使分數提高 0.006。六位前沿級評判作為檢查,重現這些分數,且作為小組的可靠性並沒有提高。接下來的兩個消融實驗則在相同的問題和答案上分解評分標準。去除其標準和級別,同時保留官方答案,並不會改變可測量的結果。去除官方答案也會使可靠性崩潰(ICC 從 0.888 降至 0.628),使分數膨脹,並使評判的推理努力再次變得重要。評分標準是將評分與評判智力解耦的關鍵,而在評分標準內,官方答案幾乎承擔了所有的工作。我們沒有發現基於評分標準的評分中存在長度偏好或同家族偏好的證據。
MotoSafety: Edge-AI with Learned Temporal Importance for Two-Wheeler Collision Risk Assessment Under Time Pressure
2608.17823v1 by Sumit S. Shevtekar, Chandresh K. Maurya, Gourab Sil, Subasish Das
Powered two-wheeler riders face critical safety challenges in low- and middle-income countries, yet limited studies exist on how cognitive stressors such as Time Pressure influence collision risk. To address this gap, we introduce a large-scale dataset of over 129,000 labeled multivariate time-series sequences from 153 simulator rides by 51 participants under No, Low, and High TP, capturing 64 features across vehicle dynamics, control inputs, proximity, and behavioral violations. Building on this dataset, we propose MotoSafety, a novel edge-AI architecture grounded in the Learned Temporal Importance principle. MotoSafety achieves 94.97% accuracy and 99.33% ROC AUC, outperforming ten baselines, including TimesNet and LLM4TS, and achieves 0.039 MSE and 0.094 MAE for forecasting (4.4x lower error than Time-LLM and iTransformer). With only 1.15M parameters and 0.135 ms latency, it is suitable for edge deployment on low-cost CPU hardware. Using ground truth TP as an inductive bias improves accuracy from 94.09% to 94.97%, while predicted TP achieves 94.82%. Using only 21 IMU+GPS features, it achieves 93.91% accuracy, indicating practical deployment. Beyond PTW safety, the architecture shows better transferability to human activity (97.66%) and clinical (99.65%) domains. This lightweight framework advances PTW collision risk assessment, supporting the Safe System Approach for Intelligent Transportation Systems.
摘要:在中低收入國家,動力二輪車騎士面臨著重大的安全挑戰,但關於認知壓力因素如時間壓力如何影響碰撞風險的研究卻相對有限。為了填補這一空白,我們引入了一個大規模數據集,該數據集包含來自51名參與者在無時間壓力、低時間壓力和高時間壓力下進行的153次模擬騎行的超過129,000個標記的多變量時間序列,捕捉了64個特徵,涵蓋了車輛動態、控制輸入、接近度和行為違規。基於這個數據集,我們提出了MotoSafety,一種基於學習時間重要性原則的新型邊緣人工智慧架構。MotoSafety實現了94.97%的準確率和99.33%的ROC AUC,超越了包括TimesNet和LLM4TS在內的十個基準,並在預測中達到了0.039的均方誤差和0.094的平均絕對誤差(比Time-LLM和iTransformer低4.4倍)。它僅需1.15M的參數和0.135毫秒的延遲,適合在低成本CPU硬體上進行邊緣部署。使用真實的時間壓力作為歸納偏見,準確率從94.09%提高到94.97%,而預測的時間壓力則達到94.82%。僅使用21個IMU+GPS特徵,它的準確率達到93.91%,顯示出實際部署的潛力。除了PTW安全性外,該架構在人體活動(97.66%)和臨床(99.65%)領域也顯示出更好的可轉移性。這個輕量級框架推進了PTW碰撞風險評估,支持智能交通系統的安全系統方法。
Communicating Credit Risk with Large Language Models: Evaluation of Explanations from Standard and Alternative Data-Based Models
2608.17715v1 by Sahab Zandi, Noah Kostesku, Christophe Mues, María Óskarsdóttir, Cristián Bravo
Credit decisioning is a high-stakes task in which model outputs must be accurate and explainable to support compliant decisions. Although modern credit risk models such as eXtreme Gradient Boosting (XGBoost) and Graph Neural Networks (GNNs) improve predictive performance, their explanations are often too technical for stakeholders creating communication gaps that can shape approvals, denials, and fairness judgments. We examine whether Large Language Models (LLMs) can serve as explanation layers that translate post-hoc explanation artefacts into stakeholder-appropriate risk narratives. Using Freddie Mac single-family loan-level data, we develop three pipelines: standard tabular (XGBoost + SHAP), and two with alternative data, a pure network-based (GNN + GNNExplainer), and a bimodal one (combining tabular and network data). We generate narratives with three LLM configurations: a small fine-tuned LLM (Gemma 3 4B), a large fine-tuned LLM (DeepSeek R1 70B), and a zero-shot commercial LLM (Gemini 2.5). Explanation quality is evaluated through automated checks across all pipelines and a human study of bimodal explanations comparing credit risk professionals and non-professionals on eight decision-relevant dimensions. We have three main findings. First, the pipeline accounts for higher variance in evidence-grounding scores than the language model, meaning that the binding constraint on explanation quality is the evidence representation, not the model used. Second, the explanation narratives reliably name the influential factors but are less reliable when stating the direction of influence, which may be consequential for adverse-action communication. Finally, professionals apply stricter evidentiary standards than non-professionals. We discuss implications for the governance of risk models, including deployment considerations and the value of domain-aligned LLMs in regulated credit settings.
摘要:信用決策是一項高風險的任務,其中模型輸出必須準確且可解釋,以支持合規的決策。儘管現代信用風險模型如極端梯度提升(XGBoost)和圖神經網絡(GNNs)提高了預測性能,但它們的解釋往往對利益相關者來說過於技術性,造成溝通差距,這可能影響批准、拒絕和公平性判斷。我們檢視大型語言模型(LLMs)是否可以作為解釋層,將事後解釋產物轉化為適合利益相關者的風險敘事。使用Freddie Mac的單戶貸款數據,我們開發了三個管道:標準表格(XGBoost + SHAP),以及兩個使用替代數據的管道,一個是純基於網絡的(GNN + GNNExplainer),另一個是雙模的(結合表格和網絡數據)。我們使用三種LLM配置生成敘事:一個小型微調LLM(Gemma 3 4B),一個大型微調LLM(DeepSeek R1 70B),以及一個零樣本商業LLM(Gemini 2.5)。通過對所有管道的自動檢查以及對雙模解釋的人工研究,我們評估了解釋質量,並比較了信用風險專業人員和非專業人員在八個與決策相關的維度上的表現。我們有三個主要發現。首先,該管道在證據基礎分數的變異性上比語言模型更高,這意味著解釋質量的約束是證據表示,而不是所使用的模型。其次,解釋敘事可靠地命名了影響因素,但在陳述影響方向時可靠性較低,這對於不利行動的溝通可能具有重要意義。最後,專業人士應用的證據標準比非專業人士更為嚴格。我們討論了風險模型治理的影響,包括部署考量和在受監管的信用環境中領域對齊的LLMs的價值。
Explainable AI-Powered Framework for Video-Based Skill Assessment in Cataract Surgery
2608.17522v1 by Mohammad Javad Ahmadi, Hamid D. Taghirad
Persistent shortages in the surgical workforce and inherent limitations of traditional training methods highlight the necessity of automated, data-driven approaches in surgical education. This study addresses these challenges by introducing a novel, explainable AI-powered framework for automated skill assessment, specifically focusing on cataract surgery. We present the world's largest dataset of cataract surgery videos, comprising 2,000 recordings. Additionally, we propose an AI-powered analytical framework that employs advanced computer vision and signal-processing techniques to automatically evaluate surgical videos to derive objective, quantitative performance indicators that complement or potentially replace subjective scoring methods. A significant advantage of our framework over previous methods lies precisely in its explainability of outputs, elevating it beyond merely an opaque skill classification tool. Through experimental analysis of 83 cataract surgery videos, we demonstrate that the automatically computed metrics exhibit strong correlations with expert-based subjective evaluations, achieving up to 87% accuracy in surgical skill assessment. Each metric was individually examined, and expert surgeons provided subjective ratings using the newly introduced Capsulorhexis Skill Assessment System (CSAS). These subjective assessments were compared with ten objective motion-based metrics extracted through our framework. The results indicated a robust correlation between subjective ratings and automated indicators, underscoring the framework's capacity to accurately model surgical expertise.
摘要:持續的外科醫療人力短缺以及傳統訓練方法的固有限制凸顯了在外科教育中自動化、數據驅動方法的必要性。這項研究通過引入一個新穎的、可解釋的人工智慧驅動框架來解決這些挑戰,特別專注於白內障手術。我們展示了世界上最大的白內障手術視頻數據集,包含2,000個錄像。此外,我們提出了一個人工智慧驅動的分析框架,利用先進的計算機視覺和信號處理技術,自動評估手術視頻,以獲得客觀的、定量的性能指標,這些指標可以補充或潛在地取代主觀評分方法。我們的框架相較於先前的方法的一個顯著優勢恰恰在於其輸出的可解釋性,使其超越僅僅是一個不透明的技能分類工具。通過對83個白內障手術視頻的實驗分析,我們證明自動計算的指標與專家基於主觀評估的評分之間存在強烈的相關性,在外科技能評估中達到高達87%的準確率。每個指標都經過單獨檢查,專家外科醫生使用新引入的囊膜切開技能評估系統(CSAS)提供主觀評分。這些主觀評估與通過我們的框架提取的十個客觀運動基礎指標進行了比較。結果顯示主觀評分與自動指標之間存在穩健的相關性,強調了該框架準確建模外科專業知識的能力。
Understanding Curriculum Learning in Large Language Models via Cross-Difficulty Optimization Dynamics
2608.17268v1 by Zhikai Ding, Ziyi Ye
Curriculum learning has been widely adopted in the post-training of large language models by organizing training data from easy to hard. However, its effectiveness varies substantially across reasoning tasks, suggesting that no single curriculum is universally optimal and raising a fundamental question: what determines when curriculum learning works? In this paper, we answer this question by analyzing the optimization dynamics induced by different curriculum schedules. We show that the transfer relationship between different difficulty levels characterizes the optimization dynamics induced by curriculum learning, which in turn explains the effectiveness of different curriculum schedules, and formalize this relationship as Relative Transfer, a principled measure of cross-difficulty knowledge transfer. Based on this measurement, we derive Transfer-aware Dynamic Curriculum Sampling (TDCS), which dynamically adjusts the sampling distribution according to the estimated transfer relationship throughout training. Extensive experiments on multiple reasoning benchmarks demonstrate that TDCS consistently outperforms representative scheduling strategies across different tasks, model scales, and training paradigms. More importantly, our work provides a unified optimization-based explanation of curriculum learning through cross-difficulty transfer.
摘要:課程學習已被廣泛應用於大型語言模型的後訓練,通過將訓練數據從簡單到困難進行組織。
然而,它在推理任務中的有效性差異很大,這表明沒有單一的課程是普遍最佳的,並提出了一個根本性問題:什麼決定了課程學習的有效性?
在本文中,我們通過分析不同課程安排所引起的優化動態來回答這個問題。
我們展示了不同難度級別之間的轉移關係特徵化了課程學習所引起的優化動態,這反過來解釋了不同課程安排的有效性,並將這一關係形式化為相對轉移,這是一種跨難度知識轉移的原則性度量。
基於這一測量,我們推導出轉移感知動態課程抽樣(TDCS),該方法根據整個訓練過程中估計的轉移關係動態調整抽樣分佈。
在多個推理基準上的大量實驗表明,TDCS在不同任務、模型規模和訓練範式中始終優於代表性的排程策略。
更重要的是,我們的工作通過跨難度轉移提供了一個統一的基於優化的課程學習解釋。
From Abductive Explanations to Global Logical Rules for Node Classification in SGCs
2608.17103v1 by Bryan Lima Cavalcante, Thiago Alves Rocha
Graph Neural Networks (GNNs) have achieved remarkable performance in node classification tasks, motivating growing interest in methods capable of explaining their predictions. Recent logic-based approaches, such as LogicXGNN, derive global logical rules for Graph Neural Networks (GNNs) from collections of explanatory subgraphs. While informative, these subgraphs may contain redundant structural information that is specific to individual nodes, potentially limiting the generality of the extracted rules. In this work, we propose a logic-based framework for node classification in Simple Graph Convolution (SGC) networks that uses minimal abductive explanations as an intermediate representation for rule extraction. For each node, we compute a minimal set of node-feature pairs sufficient to preserve the predicted class. These explanations are then used to train decision trees from which global logical rules are extracted. Experiments on benchmark datasets show that the proposed framework produces compact global rules while maintaining high fidelity to the original SGC model.
摘要:圖神經網絡(GNNs)在節點分類任務中取得了顯著的表現,這激發了對能夠解釋其預測的方法的日益關注。最近的基於邏輯的方法,如LogicXGNN,從解釋性子圖的集合中推導出圖神經網絡(GNNs)的全局邏輯規則。雖然這些子圖提供了資訊,但它們可能包含特定於個別節點的冗餘結構資訊,這可能限制了提取規則的普遍性。在本研究中,我們提出了一個基於邏輯的框架,用於簡單圖卷積(SGC)網絡中的節點分類,該框架使用最小的推斷解釋作為規則提取的中介表示。對於每個節點,我們計算一組最小的節點-特徵對,這些對足以保留預測的類別。然後,這些解釋用於訓練決策樹,從中提取全局邏輯規則。在基準數據集上的實驗表明,所提出的框架生成了緊湊的全局規則,同時保持了對原始SGC模型的高保真度。
AutoSR: Automatic Symbolic Regression by Searching Research States
2608.16876v1 by Kejia Zhang, Youran Sun, Xinyu Ren, Chugang Yi, Haizhao Yang
We introduce Automatic Symbolic Regression (AutoSR), a fully automated system that instantiates Research-Space Symbolic Regression by searching persistent scientific investigations rather than isolated equations. Finite, noisy data often yield numerically competitive expressions that imply very different behavior outside the observed regime, making numerical fit and syntactic complexity insufficient measures of scientific credibility. Existing approaches largely focus on improving expressions, yet the search typically retains little beyond the resulting formula and score, losing the scientific record, such as motivations and probes, that inform what to try next. AutoSR preserves this record in a \textbf{Research State}, coupling each candidate equation with the reasoning, computational evidence, and independent review developed along its branch. Proposer--reviewer agents develop these states under progressive-widening Monte Carlo tree search (PW-MCTS), which allocates computation across competing investigations, while the accumulated research record is ultimately synthesized into a final report that explains the leading relation and the basis for its selection. Across nine selected challenges from two benchmark suites, AutoSR recovers algebraically equivalent relations in every case, including three cp3-bench problems that no published system recovers and six structurally diverse LSR-Transform problems. Overall, AutoSR extends symbolic regression from equation-level search toward automated scientific investigation, allowing scientific knowledge and accumulated evidence to shape both what is explored and how the resulting equation is justified.
摘要:我們介紹自動符號回歸(AutoSR),這是一個完全自動化的系統,它通過搜尋持續的科學研究而不是孤立的方程式來實現研究空間符號回歸。有限的、帶噪聲的數據通常會產生數值上具有競爭力的表達式,這些表達式在觀察範圍之外暗示了非常不同的行為,使得數值擬合和語法複雜性不足以作為科學可信度的衡量標準。現有的方法主要集中在改進表達式上,但搜索通常僅保留結果公式和分數,失去了科學記錄,例如動機和探測,這些記錄告訴我們接下來該嘗試什麼。AutoSR 在一個 \textbf{研究狀態} 中保留這個記錄,將每個候選方程與沿其分支發展的推理、計算證據和獨立審查相結合。提議者-審查者代理在漸進擴展的蒙特卡羅樹搜索(PW-MCTS)下發展這些狀態,該方法在競爭的研究之間分配計算,而累積的研究記錄最終被綜合成一份最終報告,解釋主要關係及其選擇的基礎。在來自兩個基準套件的九個選定挑戰中,AutoSR 在每一個案例中都恢復了代數上等價的關係,包括三個沒有任何已發表系統恢復的 cp3-bench 問題和六個結構多樣的 LSR-Transform 問題。總體而言,AutoSR 將符號回歸從方程層級的搜索擴展到自動化的科學研究,允許科學知識和累積的證據塑造探索的內容以及結果方程的合理性。
Topological Attribution Distance (TAD): Revealing Segment-Level RAG Influence on LLM Output Geometry for Incident Log Analysis
2608.16775v1 by Reza Fayyazi, Michael Zuzak, Shanchieh Jay Yang
Large Language Models (LLMs) are increasingly being deployed in cybersecurity operations to assist cybersecurity analysts with rapid decision-making against emerging threats. However, there is a main criteria that must be met when using LLMs in cybersecurity, that is, trust in the generated outputs. As Agentic AI is integrated into operational systems, a robust evidence attribution and provenance tracking technique is essential to trace the origins of model generations. When autonomous agents make a decision (right or wrong), the ability to trace back through the decision chain is critical, as without it, teams cannot identify which segment of the data caused the model generation. Existing methods often struggle to distinguish among complex and highly similar evidence sources, such as cyber incident logs. This reveals a key gap: current approaches do not adequately capture the holistic geometric relationship between the retrieved evidence and the generated response for reliable evidence verification. To bridge this gap, we propose Topological Attribution Distance (TAD), inspired by Topology, to characterize and capture the global geometric shape of an output and its changes against its retrieved logs. In other words, if the embeddings of a specific source log drastically changes the geometry of the model's response in the embedding space, this suggests that such log is a critical source for the model's generated response. Therefore, TAD is powered by segment-level ablation attribution to investigate incident logs of an actual cyberattack. We demonstrate how TAD finds the most attributed logs on LLM outputs in an adaptive manner. This can provide an explainable and trustworthy tracing based on each LLM's hidden state to understand how geometrically different retrieved logs influence the model generation, and provide evidence verification in cybersecurity and Agentic-AI workflows.
摘要:大型語言模型(LLMs)越來越多地被應用於網路安全操作,以協助網路安全分析師快速做出針對新興威脅的決策。
然而,在網路安全中使用LLMs時,必須滿足一個主要標準,即對生成輸出的信任。
隨著代理式人工智慧整合進入操作系統,強健的證據歸屬和來源追蹤技術對於追溯模型生成的來源至關重要。
當自主代理做出決策(無論對錯)時,能夠追溯決策鏈是關鍵,因為沒有這一點,團隊無法確定是哪一部分數據導致了模型生成。
現有的方法往往難以在複雜且高度相似的證據來源之間區分,例如網路事件日誌。
這揭示了一個關鍵的缺口:當前的方法未能充分捕捉檢索到的證據與生成回應之間的整體幾何關係,以進行可靠的證據驗證。
為了填補這一缺口,我們提出了拓撲歸屬距離(TAD),受到拓撲學的啟發,用於表徵和捕捉輸出及其相對於檢索日誌的變化的全球幾何形狀。
換句話說,如果特定來源日誌的嵌入顯著改變了模型在嵌入空間中的回應幾何形狀,這表明該日誌是模型生成回應的關鍵來源。
因此,TAD由段級別的去除歸屬推動,以調查實際網路攻擊的事件日誌。
我們展示了TAD如何以自適應的方式找到對LLM輸出最具歸屬的日誌。
這可以基於每個LLM的隱藏狀態提供可解釋且值得信賴的追蹤,以了解幾何上不同的檢索日誌如何影響模型生成,並在網路安全和代理式人工智慧工作流程中提供證據驗證。
Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments
2608.16747v1 by Adam Karvonen, Euan Ong, Subhash Kantamneni, Samuel Marks
Many areas of AI research, such as language model interpretability and chain of thought faithfulness, seek to explain model behaviors. But what constitutes a "good" explanation? In this work, we evaluate explanations through the lens of counterfactual simulatability-whether the explanation is useful for predicting model behaviors on related counterfactual inputs. To this end, we introduce CHIVE (Counterfactual Hypothesis Investigation Via Edits), a novel agentic pipeline that identifies unexpected model behaviors in the wild and investigates them with counterfactual prompt edits. This yields thousands of high-quality explanations for naturally-occurring model behaviors along with supporting counterfactual evidence. We apply CHIVE in two ways. First, we evaluate whether common LLM interpretability techniques improve an agent's ability to predict counterfactual model behaviors. Surprisingly, we find no uplift from any of the interpretability techniques studied. Second, we use CHIVE to generate training data. We find that training models to predict outcomes of CHIVE-generated counterfactual experiments generalizes to various out-of-distribution settings. Overall, CHIVE automatically discovers explanations of naturally-occurring LLM behaviors, enabling us to evaluate and improve methods for explaining LLM behaviors.
摘要:許多AI研究領域,例如語言模型的可解釋性和思維鏈的可靠性,旨在解釋模型行為。
但什麼構成了「好的」解釋?
在這項工作中,我們通過反事實可模擬性的視角來評估解釋——即該解釋是否對預測模型在相關反事實輸入上的行為有用。
為此,我們引入了CHIVE(通過編輯進行反事實假設調查),這是一個新穎的主動管道,能夠識別自然環境中意外的模型行為並通過反事實提示編輯進行調查。
這產生了數千個高品質的解釋,針對自然發生的模型行為以及支持的反事實證據。
我們以兩種方式應用CHIVE。
首先,我們評估常見的LLM可解釋性技術是否能提高代理預測反事實模型行為的能力。
令人驚訝的是,我們發現在所研究的任何可解釋性技術中都沒有提升。
其次,我們使用CHIVE生成訓練數據。
我們發現,訓練模型以預測CHIVE生成的反事實實驗的結果能夠泛化到各種分佈外的情境。
總體而言,CHIVE自動發現自然發生的LLM行為的解釋,使我們能夠評估和改進解釋LLM行為的方法。
Unsupervised Anomaly Detection for Image Dataset Quality Assurance in Multi-Center Breast MRI
2608.16725v1 by Chiara Tappermann, Steffen Renisch, Lars Ole Schwen, Hans Meine, Horst K. Hahn, Eike Petersen
Corrupted, inconsistent, or anomalous data silently threatens the safety and reliability of medical AI. Despite growing regulatory recognition of dataset quality assurance (QA) for high-risk medical AI, scalable automated detection remains underdeveloped. We employ unsupervised anomaly detection (AD) and out-of-distribution (OOD) detection as an automated dataset QA mechanism for multi-center dynamic contrast-enhanced breast MRI. We build a controlled AD benchmark of 17 realistic QA-relevant anomaly types from six public datasets (protocol violations, processing errors, incorrect anatomical regions) and propose a taxonomy of radiological image anomalies based on human visual perception, enabling fine-grained analysis of AD failure modes. The benchmark includes near-, medium-far-, far-OOD samples, as well as in-distribution and external normal data. Four methods are evaluated: a projection-based method extended with a domain-specific feature extractor and a novel positional encoding, a reconstruction-based approach extended to full 3D volumes with an augmented training objective, and two unmodified hybrid OOD detection methods. Medium-far- and far-OOD samples are detected reliably, whereas near-OOD samples and external normal data from unseen institutions expose method-specific differences. The 3D reconstruction-based approach best balances detection performance (AUROC: 0.936) and generalization to unseen institutions. The projection-based method with positional encoding achieves the highest overall detection performance (AUROC: 0.954). Both hybrid methods exhibit critical failure modes, confirming that methods validated for one modality or anatomy may not generalize without domain-specific adaptation. Implants and mastectomies remain an open challenge for all methods. Our results establish a foundation and practical guidance on scalable unsupervised QA in medical AI pipelines.
摘要:腐敗、不一致或異常的數據默默威脅著醫療人工智慧的安全性和可靠性。儘管對高風險醫療人工智慧數據集質量保證(QA)的監管認識日益增長,但可擴展的自動檢測仍然發展不足。我們採用無監督異常檢測(AD)和分佈外(OOD)檢測作為多中心動態對比增強乳腺MRI的自動數據集QA機制。
我們建立了一個由六個公共數據集中的17種現實QA相關異常類型組成的受控AD基準(協議違規、處理錯誤、不正確的解剖區域),並根據人類視覺感知提出了一個放射影像異常的分類法,使得對AD失效模式的細緻分析成為可能。基準包括近距離、中遠距離、遠距離OOD樣本,以及分佈內和外部正常數據。評估了四種方法:一種基於投影的方法,擴展了特定領域的特徵提取器和新穎的位置編碼;一種基於重建的方法,擴展到完整的3D體積並具有增強的訓練目標;以及兩種未經修改的混合OOD檢測方法。
中遠距離和遠距離OOD樣本的檢測可靠,而近距離OOD樣本和來自未見機構的外部正常數據則顯示出方法特定的差異。基於3D重建的方法在檢測性能(AUROC:0.936)和對未見機構的泛化之間達到了最佳平衡。帶有位置編碼的基於投影的方法實現了最高的整體檢測性能(AUROC:0.954)。兩種混合方法都顯示出關鍵的失效模式,確認了針對一種模態或解剖結構驗證的方法可能無法在沒有特定領域適應的情況下進行泛化。植入物和乳房切除術對所有方法仍然是一個未解決的挑戰。我們的結果為醫療人工智慧管道中的可擴展無監督QA建立了基礎和實用指導。
Degradation-Aligned Self-Supervised Learning for State of Health Estimation of Lithium-Ion Batteries under Label Sparsity
2608.16612v1 by Jiaqi Yao, Julia Kowal
An accurate estimation of the state of health (SOH) underpins a safe and optimized use of the battery system. Although compelling, data-driven SOH estimation models typically require large amounts of high-quality labeled cycling data, while in practice such labels are often sparse in both quantity and coverage. Therefore, in this work, we propose a degradation-aligned self-supervised learning (SSL) framework based on a convolutional neural network-gated recurrent unit (CNN-GRU) model, which learns aging-consistent representations from unlabeled data through a cycle-order ranking objective as the pretext task for pretraining, thereby enabling robust SOH estimation after fine-tuning on sparsely labeled data. Test results showcase that the proposed ranking-based SSL approach proves to endow the pretrained model with degradation-aligned information from unlabeled data, and after fine-tuning the model can carry out accurate, robust SOH estimation, even when only an extremely limited amount of 1% of unevenly distributed labeled training data is available, where the MAE of 1.718% and RMSE of 2.329% can be achieved on the test cell. In addition, in-depth analyses are presented regarding the influences of label distribution of battery degradation data. We believe this work could shed new light on SOH estimation of lithium-ion batteries under label sparsity in real-world applications.
摘要:準確的健康狀態(SOH)估計是安全且優化使用電池系統的基礎。雖然數據驅動的SOH估計模型非常有說服力,但通常需要大量高質量的標註循環數據,而在實際情況中,這些標註往往在數量和覆蓋範圍上都很稀疏。因此,在本研究中,我們提出了一種基於卷積神經網絡-門控遞歸單元(CNN-GRU)模型的降解對齊自監督學習(SSL)框架,通過循環順序排名目標作為預訓練的前置任務,從未標註數據中學習與老化一致的表示,從而在稀疏標註數據上進行微調後實現穩健的SOH估計。測試結果顯示,所提出的基於排名的SSL方法使預訓練模型從未標註數據中獲得了降解對齊的信息,並且在微調後,該模型能夠進行準確且穩健的SOH估計,即使僅有極少量的1%不均勻分佈的標註訓練數據可用,測試電池的MAE可達1.718%和RMSE可達2.329%。此外,還對電池降解數據的標註分佈影響進行了深入分析。我們相信這項工作可以為在現實應用中標註稀疏的鋰離子電池SOH估計提供新的見解。
Physics of Agents: Statistical Mechanics Predicts Collective Behavior of AI Agents
2608.16578v1 by Batu El, Jinhee Paeng, Fatih Dinc, Shiye Su, Mete Erdogan, Aneesh Pappu, Haotian Ye, Wanjia Zhao, Surya Ganguli, James Zou
AI agents increasingly operate as part of interacting systems rather than in isolation. As agents exchange information and jointly make decisions, their interactions can improve collective reasoning but may also produce herding, polarization, or amplify shared biases. Understanding and predicting these collective dynamics is therefore important for designing effective and aligned multi-agent systems. Here, we study over 10,000 communities of language-model agents that repeatedly exchange messages and revise their opinions across objective mathematics questions and subjective political statements. Despite substantial diversity in possible behavior, the individual and group dynamics can be represented by three characteristic regimes: indifference, polarization, and consensus. AI agents start indifferent and build conviction as they interact. On objective questions, communication improves collective accuracy, while on subjective questions it often drifts group opinions toward the right in the political spectrum. We explain these observations with a statistical-mechanics formalism in which agents stochastically favor lower social pressure. Given only initial opinions, our model predicts individual trajectories, outperforms all standard baselines, generalizes to unseen community graphs, and reproduces the observed group archetype distributions. Our fitted model parameters reveal the mechanics underlying our key observations: i) communities operate below the critical social temperature, which explains conviction buildup; ii) attractive ties outweigh repulsive ones, which favors consensus; and iii) agents holding the correct answer exert the strongest pull, which drives truth-seeking. Overall, our results demonstrate that collective behavior of AI agents, like that of other complex systems, follows compact and predictive dynamical laws.
摘要:AI 代理人越來越多地作為互動系統的一部分運作,而不是孤立存在。
隨著代理人之間交換信息並共同做出決策,他們的互動可以改善集體推理,但也可能產生跟風、極化或放大共享偏見。
因此,理解和預測這些集體動態對於設計有效且一致的多代理系統非常重要。
在這裡,我們研究了超過 10,000 個語言模型代理人的社群,它們反覆交換消息並在客觀數學問題和主觀政治陳述上修正自己的意見。
儘管可能的行為存在相當大的多樣性,但個體和群體動態可以用三種特徵性狀態來表示:漠不關心、極化和共識。
AI 代理人最初是漠不關心的,隨著互動的進行建立信念。
在客觀問題上,交流提高了集體準確性,而在主觀問題上,則經常使群體意見向政治光譜的右側漂移。
我們用一種統計力學形式主義來解釋這些觀察,其中代理人隨機地偏好較低的社會壓力。
僅根據初始意見,我們的模型預測個體軌跡,超越所有標準基準,對未見過的社群圖進行泛化,並重現觀察到的群體原型分佈。
我們擬合的模型參數揭示了我們關鍵觀察背後的機制:i) 社群運作在臨界社會溫度以下,這解釋了信念的積累;ii) 吸引性聯繫超過排斥性聯繫,這有利於共識;iii) 持有正確答案的代理人施加最強的影響,這驅動尋求真相。
總體而言,我們的結果表明,AI 代理人的集體行為,如同其他複雜系統,遵循緊湊且可預測的動態法則。
Revisiting the Performance of Generative Artificial Intelligence on Introductory Object-Oriented Programming Assessments: Insights from 2026
2608.16318v1 by Marina Lepp, Joosep Kaimre
Recent advances in Generative Artificial Intelligence (GenAI) have substantially improved the ability of large language models (LLMs) to generate and explain source code. However, their performance on authentic object-oriented programming (OOP) assessments remains insufficiently understood. This study evaluates five widely used GenAI systems, ChatGPT-5.2, DeepSeek-V3, Gemini 2.5 Flash, Claude Sonnet 4.5, and M365 Copilot, using programming tests and examination tasks from an introductory university OOP course. The generated solutions were assessed using the same grading criteria applied to students and compared with historical student results from the same course, as well as findings from the previous year. Common errors were also analyzed to identify recurring limitations across models. All evaluated GenAI systems achieved higher scores than the average student cohort and frequently obtained full marks on longer programming tasks. Nevertheless, they occasionally produced non-compiling code and continued to struggle with advanced OOP concepts, particularly interfaces, abstract classes, and certain inheritance-related tasks. Performance was also limited on graphics-related questions involving image interpretation. Compared with the previous year, the evaluated systems demonstrated noticeable improvements across most assessments while exhibiting several recurring error patterns. The findings provide an updated evaluation of the capabilities and limitations of contemporary GenAI systems on authentic introductory OOP assessments. They also offer evidence that can inform the design of programming assessments, the responsible integration of GenAI tools into software engineering education, and future studies evaluating the evolution of AI-assisted programming.
摘要:最近在生成式人工智慧(GenAI)方面的進展顯著提高了大型語言模型(LLMs)生成和解釋源代碼的能力。
然而,它們在真實物件導向程式設計(OOP)評估中的表現仍然不夠了解。
本研究評估了五個廣泛使用的GenAI系統,ChatGPT-5.2、DeepSeek-V3、Gemini 2.5 Flash、Claude Sonnet 4.5和M365 Copilot,使用來自大學入門OOP課程的程式設計測試和考試任務。
生成的解決方案使用與學生相同的評分標準進行評估,並與來自同一課程的歷史學生結果以及前一年的研究結果進行比較。
還分析了常見錯誤,以識別模型之間的重複限制。
所有評估的GenAI系統的得分均高於平均學生群體,並且在較長的程式設計任務中經常獲得滿分。
然而,它們偶爾會產生無法編譯的代碼,並且在高級OOP概念方面仍然存在困難,特別是介面、抽象類別和某些與繼承相關的任務。
在涉及圖像解釋的圖形相關問題上,表現也受到限制。
與前一年相比,評估的系統在大多數評估中顯示出明顯的改進,同時表現出幾個重複的錯誤模式。
這些發現提供了對當代GenAI系統在真實入門OOP評估中的能力和限制的最新評估。
它們還提供了可以為程式設計評估的設計、負責任地將GenAI工具整合到軟體工程教育中,以及未來評估AI輔助程式設計演變的研究提供依據的證據。
Foresight-England: Development of a National-Scale Generative AI Model of Electronic Health Records for Medical Event Prediction across the COVID-19 Pandemic
2608.16273v1 by Simon Ellershaw, Christopher Tomlinson, Zeljko Kraljevic, Spiros Denaxas, Harry Hemingway, Cathie Sudlow, Angela M. Wood, Anoop D. Shah, Richard Dobson
Foresight-England (Foresight-E) is the first national-scale generative foundation model of electronic health records (EHRs), developed as a research pilot strictly for COVID-19 research. We evaluated its ability to model the direct and indirect effects of the pandemic. Trained from scratch entirely within the NHS England Secure Data Environment, Foresight-E is a 243-million-parameter transformer decoder. It was trained and evaluated on de-identified, longitudinal EHRs of approximately 61 million individuals, integrating primary/secondary care, death registrations, and COVID-19 data. Training and validation used a 90% subset (54.9 million) spanning November 2018 to December 2022; the remaining 10% (6.1 million) was held out for evaluation. Foresight-E models patient timelines autoregressively, predicting the next medical event given their prior history. At inference, it operates zero-shot, predicting any concept in its ~40,000-code vocabulary without task-specific training. Our tokenisation scheme retains the clinical granularity of ICD-10, OPCS-4, and SNOMED CT codes, jointly representing absolute and relative timing. We designed an evaluation framework for 30-day COVID-19 hospitalisation and mortality, including subgroup analyses by demographic factors and vaccination status. To assess generalisation to unseen future data and the pandemic's indirect effects, we tested the model on medical events from 2023 (beyond its training period), benchmarking against logistic regression and XGBoost. As detailed in the Project Status section, NHS England has paused access to data for the Foresight-E project, meaning quantitative results are currently unavailable. Instead, we share our strategy for tokenisation, architecture, training, inference, and evaluation as a methodological template and case study in the challenges of building population-scale EHR foundation models.
摘要:Foresight-England (Foresight-E) 是首個全國規模的電子健康紀錄 (EHRs) 生成基礎模型,作為針對 COVID-19 研究的研究試點而開發。
我們評估了它建模疫情直接和間接影響的能力。
Foresight-E 完全在 NHS England 安全數據環境中從零開始訓練,是一個擁有 2.43 億參數的Transformer解碼器。
它在約 6100 萬人的去識別化、縱向 EHRs 上進行訓練和評估,整合了初級/次級護理、死亡登記和 COVID-19 數據。
訓練和驗證使用了 90% 的子集(5490 萬),涵蓋了 2018 年 11 月到 2022 年 12 月;剩餘的 10%(610 萬)則保留用於評估。
Foresight-E 自回歸地建模患者時間線,根據其先前的歷史預測下一個醫療事件。
在推理時,它以零樣本操作,預測其約 40,000 種代碼詞彙中的任何概念,而無需特定任務的訓練。
我們的標記方案保留了 ICD-10、OPCS-4 和 SNOMED CT 代碼的臨床細節,聯合表示絕對和相對時間。
我們設計了一個評估框架,用於 30 天 COVID-19 住院和死亡率,包括按人口統計因素和疫苗接種狀態的子群分析。
為了評估對未見未來數據的泛化能力和疫情的間接影響,我們在 2023 年的醫療事件上測試了該模型(超出其訓練期間),並與邏輯回歸和 XGBoost 進行基準比較。
正如項目狀態部分詳細說明的那樣,NHS England 已暫停對 Foresight-E 項目的數據訪問,這意味著目前無法獲得定量結果。
相反,我們分享了我們的標記化、架構、訓練、推理和評估的策略,作為建立人口規模 EHR 基礎模型挑戰的 методологический шаблон и кейс-исследование。
Defake-o3: From Speculative Rationales to Verifiable Evidence for Explainable AIGI Detection
2608.16259v1 by Bowen Deng, Jiahui Zhan, Yikun Ji, Haozhen Yan, Jianfu Zhang
The rapid progress of image generation models calls for AI-generated image (AIGI) detectors that are not only accurate but also explainable and reliable. While MLLM-based detectors can provide natural language explanations, existing methods often generate speculative rationales: they rely on vague or hallucinated artifacts, miss subtle localized flaws from the latest generators, and fail to provide evidence that can be visually verified. We present Defake-o3, an explainable AIGI detector that moves from speculative rationales to verifiable evidence. It combines interactive visual search with verifier-guided evidence alignment: the model iteratively zooms into suspicious regions to inspect fine-grained details, while an Evidence Verifier, trained from human verification annotations, provides reinforcement learning rewards that favor grounded evidence and penalize baseless claims. To support this objective, we construct GroundFake, a dataset designed for grounded explainable detection, with localized bounding-box evidence, human verification based on visual grounding and artifact specificity, corrected reasoning trajectories, and valid/invalid evidence supervision. We further introduce FakeFrontier, an out-of-distribution benchmark built from real images and outputs of 10 recent generators, together with an MLLM-based protocol for evaluating evidence quality and persuasiveness. Experiments on GroundFake, FakeFrontier, and additional out-of-distribution benchmarks show that Defake-o3 improves both detection accuracy and explanation quality, producing more localized, verifiable, and persuasive evidence.
摘要:快速進展的圖像生成模型呼喚不僅準確而且可解釋和可靠的AI生成圖像(AIGI)檢測器。雖然基於MLLM的檢測器可以提供自然語言解釋,但現有方法往往生成推測性的理由:它們依賴模糊或虛幻的工件,錯過最新生成器的微妙局部缺陷,並未能提供可以視覺驗證的證據。我們提出了Defake-o3,一個可解釋的AIGI檢測器,從推測性理由轉向可驗證的證據。它結合了互動式視覺搜索和驗證者引導的證據對齊:模型迭代地放大可疑區域以檢查細緻的細節,而一個基於人類驗證註釋訓練的證據驗證器提供強化學習獎勵,以支持有根據的證據並懲罰無根據的主張。為了支持這一目標,我們構建了GroundFake,一個為有根據的可解釋檢測設計的數據集,包含局部邊界框證據、基於視覺定位和工件特異性的人工驗證、修正的推理軌跡,以及有效/無效證據的監督。我們進一步介紹了FakeFrontier,一個基於真實圖像和10個最近生成器輸出的分佈外基準,並附帶一個基於MLLM的協議,用於評估證據的質量和說服力。在GroundFake、FakeFrontier和其他分佈外基準上的實驗顯示,Defake-o3提高了檢測準確性和解釋質量,產生了更局部、可驗證和具說服力的證據。
CompoSkill: Compositional Skill Chain Attacks from Individually Scanner-Passing LLM Agent Skills
2608.16246v1 by Mingxiao Liu, Zhoumian Jiang, Jianan Ma, Jian Zhang, Jialuo Chen, Xinhao Deng, Zhen Wang
Autonomous AI agents tackling Long Horizon Tasks depend on marketplace skills that are certified one at a time: a scanner returns a safety verdict for each skill and declares the ecosystem safe if every package passes. We show that this assumption fails under skill composition. A skill may pass the per-skill scanner individually yet participate in a risky composition when an agent connects its outputs, capabilities, or side effects with those of other scanner-passing skills. This makes skill composition risk a path level property rather than a node level property, explaining why existing skill scanners that inspect individual packages achieve limited interception. To study this threat, we present CompoSkill, a framework that constructs skill composition attacks through a dual attacker system. The white-box attacker knows the victim's installed skill pool and directly injects explicit skill-id sequences; the black-box attacker knows only a role profile, downloads the top marketplace skills for that scenario, builds a Skill Composition Graph, and searches for high risk chains whose implicit lures never name skill identifiers. We further construct CompoSkill-Bench, a benchmark of 1,140 records built from long-horizon professional workflows across five threats and six scenarios on OpenClaw and Nanobot. CompoSkill achieves risk Chain Formation Rates (CFR) up to 83.3% in the white box setting and 80.6% in the black box setting, while existing skill scanners block only a limited fraction of the risky compositions. Finally, we observe a bridge-bonus-then-hop-decay pattern: a bridge skill can increase attack success, but Attack Success Rate (ASR) decreases once additional hops make the risk chain longer than three skills. These results expose a systematic gap in single skill certification for autonomous AI agents.
摘要:自主 AI 代理處理長期任務依賴於逐一認證的市場技能:掃描器為每項技能返回安全判決,並在每個包裝通過時聲明生態系統安全。我們顯示這一假設在技能組合下失效。某項技能可能在每項技能掃描器中單獨通過,但當代理將其輸出、能力或副作用與其他通過掃描器的技能連接時,可能參與風險組合。這使得技能組合風險成為路徑級別的特性,而非節點級別的特性,解釋了為什麼現有的技能掃描器僅檢查單個包裝而達到有限的攔截效果。為了研究這一威脅,我們提出了 CompoSkill,一個通過雙重攻擊者系統構建技能組合攻擊的框架。白盒攻擊者知道受害者的已安裝技能池,並直接注入明確的技能 ID 序列;黑盒攻擊者僅知道一個角色配置,下載該場景的頂級市場技能,構建技能組合圖,並搜索隱含誘餌從未命名技能標識符的高風險鏈。我們進一步構建了 CompoSkill-Bench,一個基於 OpenClaw 和 Nanobot 的五種威脅和六種場景中,從長期專業工作流程中建立的 1,140 條記錄的基準。CompoSkill 在白盒設置中達到高達 83.3% 的風險鏈形成率 (CFR),在黑盒設置中達到 80.6%,而現有的技能掃描器僅阻止有限比例的風險組合。最後,我們觀察到一種橋接-獎勵-然後跳躍衰減的模式:橋接技能可以提高攻擊成功率,但一旦額外的跳躍使風險鏈超過三項技能,攻擊成功率 (ASR) 就會下降。這些結果揭示了自主 AI 代理在單一技能認證方面的系統性缺口。
When Single-Dataset Conclusions Fail: A 45-Task Study of Threshold Tuning and Resampling for Imbalanced Classification
2608.16147v1 by Diyorbek Musaev
Class-imbalance handling is routinely evaluated on a single benchmark dataset, and the resulting conclusions are reported as if they were properties of the method. We show this practice is unsafe. On the public Kaggle credit-card fraud dataset, under a leakage-free nested cross-validation protocol in which the decision threshold is selected on a held-out inner validation fold, a plain Random Forest at the default 0.5 threshold attains F1 = 0.861 +/- 0.021, and threshold tuning yields it no benefit (delta-F1 = -0.002). Read alone, this supports an appealing conclusion: for a well-calibrated ensemble, imbalance handling is unnecessary. We then apply the identical protocol to 45 binary tasks spanning imbalance ratios from 1:1.5 to 1:178 (2,025 model fits, four model families). The conclusion reverses. Random Forest benefits most from threshold tuning across the suite (delta-F1 = +0.101 +/- 0.134), not least, while three other families replicate their fraud-dataset behaviour almost exactly. SMOTE likewise harms the fraud dataset but helps across the suite (mean delta-F1 = +0.076; 138 wins, 39 losses; Wilcoxon p = 2.7e-17). Two further results. Threshold-tuning benefit is non-monotonic in the imbalance ratio: near zero below 1:5, peaking at +0.120 in the 1:15-1:40 band, declining to +0.045 beyond 1:100 - explaining why the fraud dataset, at 1:577, is an unrepresentative place to study the question. And we reject an intuitive heuristic: validation-set calibration error does not predict tuning benefit (expected calibration error r = -0.087; Brier r = +0.137), so calibration diagnostics cannot tell a practitioner whether tuning is worthwhile. We release the protocol, the 45-task harness, and all per-run metrics.
摘要:類別不平衡處理通常在單一基準數據集上進行評估,並且所得到的結論被報告為該方法的特性。我們顯示這種做法是不安全的。在公共的Kaggle信用卡詐騙數據集上,在一個無泄漏的嵌套交叉驗證協議下,決策閾值是在保留的內部驗證折上選擇的,普通的隨機森林在默認的0.5閾值下達到F1 = 0.861 +/- 0.021,而閾值調整對其沒有益處(delta-F1 = -0.002)。單獨閱讀這一點支持了一個吸引人的結論:對於一個良好校準的集成,處理不平衡是沒有必要的。
我們隨後將相同的協議應用於45個二元任務,涵蓋不平衡比率從1:1.5到1:178(2,025個模型擬合,四個模型系列)。結論顛倒了。隨機森林在整個系列中最受益於閾值調整(delta-F1 = +0.101 +/- 0.134),而且三個其他系列幾乎完全複製了它們在詐騙數據集上的行為。SMOTE同樣對詐騙數據集有害,但在整個系列中有幫助(平均delta-F1 = +0.076;138次勝利,39次失敗;Wilcoxon p = 2.7e-17)。
另外兩個結果。閾值調整的益處在不平衡比率中是非單調的:在1:5以下接近零,在1:15-1:40區間達到+0.120的峰值,超過1:100則下降至+0.045——這解釋了為什麼詐騙數據集在1:577的情況下是一個不具代表性的研究該問題的地方。我們還拒絕了一個直觀的啟發式:驗證集的校準誤差並不預測調整的益處(預期校準誤差r = -0.087;Brier r = +0.137),因此校準診斷無法告訴從業者調整是否值得。我們釋放了該協議、45任務的框架以及所有每次運行的指標。
TokenSTFormer: A Tokenized Spatial-temporal Attention Model for Holistic Motion Analysis in Adolescent Idiopathic Scoliosis Screening
2608.16122v1 by Dong Chen, Kenneth M. C. Cheung
Adolescent Idiopathic Scoliosis (AIS) is a prevalent spinal deformity in adolescents that, if left untreated, can result in severe health outcomes. Traditional screening methods are limited by subjective interpretation, reliance on professional expertise and low scalability. To address these challenges, we present ScoliGait dataset, which comprises 1,516 gait video clips paired with corresponding X-ray records. We also introduce TokenSTFormer, a novel model that tokenizes spatial and temporal semantics to enhance feature representation and convergence. Our model achieves state-of-the-art performance, surpassing vanilla Vision Transformer encoder across key metrics, including accuracy of 0.79. This study highlights the potential of leveraging holistic motion features derived from gait video and attention-based models for scalable, cost-effective AIS screening, paving the way for future clinical applications in scoliosis detection.
摘要:青少年特發性脊柱側彎(AIS)是青少年中常見的脊柱畸形,如果不加以治療,可能會導致嚴重的健康後果。傳統的篩檢方法受到主觀解釋、依賴專業知識和低可擴展性的限制。為了解決這些挑戰,我們提出了 ScoliGait 數據集,其中包含 1,516 段行走視頻片段,並配有相應的 X 光記錄。我們還介紹了 TokenSTFormer,一種新型模型,將空間和時間語義進行標記化,以增強特徵表示和收斂。我們的模型在關鍵指標上達到了最先進的性能,超越了普通的視覺Transformer編碼器,包括 0.79 的準確率。本研究突顯了利用從行走視頻中獲得的整體運動特徵和基於注意力的模型進行可擴展、成本效益高的 AIS 篩檢的潛力,為未來脊柱側彎檢測的臨床應用鋪平了道路。
Eigenanalysis framework for autoregressive neural emulators of multi-scale chaotic dynamics
2608.16084v1 by Conrad Ainslie, Pedram Hassanzadeh, Michael W. Mahoney, Ashesh Chattopadhyay
Neural autoregressive models have rapidly emerged as powerful emulators of high-dimensional chaotic systems, yet their long-term instability and error growth remain poorly understood, leading to ad-hoc solutions. Here, we develop an eigenanalysis framework that reveals the dynamical origin of this error growth. By analyzing the Jacobian of the learned one-step update map with respect to the state, we show how inference-time error growth, and thus model stability, is governed by its spectral radius. Direct-step architectures (models that predict the next state from the previous one) generically admit unstable eigenvalues with magnitudes exceeding one, explaining the rapid divergence of these widely used models. In contrast, integration-constrained models (where the time derivative is estimated and integrated with a higher-order integrator) collapse their eigenspectrum onto the unit circle, yielding neutral stability and a universal linear error-scaling law. The largest eigenvalue of this Jacobian provides an architecture-agnostic, a priori diagnostic of short-term skill, long-term stability, and spectral bias, without requiring an expensive rollout. Leveraging this theory, we introduce a stability-promoting loss that explicitly regularizes Jacobian-driven error amplification, improving both forecast accuracy and dynamical robustness. Demonstrated across $29$ models spanning two architectures, several explicit and implicit integrators, and multiple loss functions on the Kuramoto-Sivashinsky system, our results establish a theoretical foundation for the design and evaluation of neural emulators of chaotic multi-scale dynamics. More broadly, our framework is a step toward the kind of a priori stability analysis that numerical analysis provides for discretizations of differential equations and that scientific machine learning currently lacks.
摘要:神經自回歸模型迅速崛起,成為高維混沌系統的強大模擬器,但其長期不穩定性和誤差增長仍然不甚了解,導致了臨時解決方案。在這裡,我們開發了一個特徵分析框架,揭示了這種誤差增長的動力學起源。通過分析學習到的一步更新映射的雅可比矩陣相對於狀態的情況,我們展示了推理時誤差增長以及模型穩定性是如何受到其譜半徑的控制。直接步驟架構(從前一狀態預測下一狀態的模型)通常承認具有超過一的幅度的不穩定特徵值,解釋了這些廣泛使用模型的快速發散。相反,集成約束模型(在此模型中,時間導數被估計並與高階積分器進行積分)將其特徵譜壓縮到單位圓上,產生中性穩定性和普遍的線性誤差縮放法則。這個雅可比矩陣的最大特徵值提供了一種與架構無關的、事先的短期技能、長期穩定性和譜偏差的診斷,而無需昂貴的展開。利用這一理論,我們引入了一種促進穩定性的損失,明確正則化雅可比驅動的誤差放大,改善預測準確性和動力學穩健性。在涵蓋兩種架構、幾種顯式和隱式積分器以及多種損失函數的 $29$ 個模型中,我們的結果為設計和評估混沌多尺度動力學的神經模擬器建立了理論基礎。更廣泛地說,我們的框架是朝著數值分析為微分方程的離散化提供的那種事先穩定性分析邁出的一步,而這是當前科學機器學習所缺乏的。
NICE: Scale-Stable Perturbations for Graph Neural Network Explanations via Noise Corruption
2608.16038v2 by Ziluowen Luo, Jun Yin, Ruochen Liu, Ming Cheng, Shirui Pan, Chengqi Zhang, Senzhang Wang
Post-hoc Graph Neural Network (GNN) explainers commonly follow a Perturb-Query paradigm, inferring the importance of graph elements based on queried predictions to perturbed inputs. However, such perturbations often introduce substantial distribution shift, undermining the reliability of the queried predictions used to derive explanations. While existing efforts mainly improve perturbed graphs or stabilize model predictions on them, we revisit the perturbation mechanism itself. We show that the widely used Element-wise Masking(EM) suppresses edge-induced messages toward zero, causing deterministic scale contraction that accumulates across message-passing layers, a phenomenon we term Scale Drift. Consequently, prediction changes under EM may conflate information corruption with deviations in propagation scale. As a scale-stable alternative to EM, we introduce Noise Corruption (NC), which perturbs each message through matched-norm random-direction corruption while preserving the expected squared message norm. Building on NC, we propose NICE, a Noise Corruption-based explanation framework, which learns a Stochastic Restoration Boundary (SRB) under NC-induced uncertainty, balancing target-prediction restoration against compactness. Furthermore, Boundary-Integrated Gradient (BIG) converts this boundary into edge attributions by accumulating each edge's contribution to reducing restoration risk along the restoration path. Experiments across multiple benchmarks demonstrate stronger explanation performance and model faithfulness while confirming that NC substantially reduces the Scale Drift induced by masking.
摘要:後 hoc 圖神經網絡 (GNN) 解釋器通常遵循擾動-查詢範式,根據對擾動輸入的查詢預測推斷圖元素的重要性。
然而,這種擾動往往會引入顯著的分佈變化,削弱用於推導解釋的查詢預測的可靠性。
雖然現有的努力主要改善擾動圖或穩定模型在其上的預測,但我們重新審視擾動機制本身。
我們展示了廣泛使用的逐元素掩蔽 (EM) 將邊緣引起的消息壓制至零,導致確定性的縮放收縮,這一現象在消息傳遞層中累積,我們稱之為縮放漂移。
因此,EM 下的預測變化可能將信息損壞與傳播縮放的偏差混淆在一起。
作為 EM 的一種穩定縮放替代方案,我們引入了噪聲損壞 (NC),它通過匹配範數的隨機方向擾動每條消息,同時保持期望的平方消息範數。
基於 NC,我們提出了 NICE,一種基於噪聲損壞的解釋框架,它在 NC 引起的不確定性下學習隨機恢復邊界 (SRB),平衡目標預測恢復與緊湊性。
此外,邊界整合梯度 (BIG) 通過累積每條邊對減少恢復風險的貢獻,將這一邊界轉換為邊緣歸因。
在多個基準上的實驗顯示出更強的解釋性能和模型忠實度,同時確認 NC 顯著減少了由掩蔽引起的縮放漂移。
Identifying Confusion Trends in Concept-based XAI for Multi-Label Classification
2608.15731v1 by Haadia Amjad, Ronald Tetzlaff
Deep Neural Networks (DNNs) deployed in high-risk domains, such as healthcare and autonomous driving, must be not only accurate but also understandable to ensure user trust. In real-world computer vision tasks, these models often operate on complex images containing background noise and are heavily annotated. To make such models explainable, Concept-based Explainable AI (CXAI) methods need to be assessed for their applicability and problem-solving capacity. In this work, we explore CXAI use cases in multi-label classification by training two DNNs, VGG16 and ResNet50, on the 20 most annotated labels in the MS-COCO dataset (Microsoft Common Objects in Context). We apply two CXAI methods, CRP (Concept Relevance Propagation) and CRAFT (Concept Recursive Activation FacTorization), to generate concept-level explanations and investigate the overall evaluations. Our analysis reveals three key findings: (1) CXAI highlights learning weaknesses in DNNs, (2) higher concept distinctiveness reduces label and concept confusion, and (3) environmental concepts expose dataset-induced biases. Our results demonstrate the potential of CXAI to enhance the understanding of model generalizability and to diagnose bias instigated by the dataset.
摘要:深度神經網絡(DNNs)在高風險領域(如醫療保健和自動駕駛)的應用,不僅必須準確,還必須可理解,以確保用戶信任。
在現實世界的計算機視覺任務中,這些模型通常在包含背景噪聲且標註繁重的複雜圖像上運行。
為了使這些模型具有可解釋性,需要評估基於概念的可解釋人工智慧(CXAI)方法的適用性和解決問題的能力。
在本研究中,我們通過在 MS-COCO 數據集(微軟上下文中的常見物體)上訓練兩個 DNN(VGG16 和 ResNet50),探索 CXAI 在多標籤分類中的應用案例,該數據集包含 20 個標註最多的標籤。
我們應用兩種 CXAI 方法,CRP(概念相關性傳播)和 CRAFT(概念遞歸激活因子分解),以生成概念級別的解釋並調查整體評估。
我們的分析揭示了三個關鍵發現:(1)CXAI 突出了 DNN 的學習弱點,(2)較高的概念區別性減少了標籤和概念的混淆,以及(3)環境概念揭示了數據集引起的偏見。
我們的結果展示了 CXAI 在增強模型可泛化性理解和診斷由數據集引發的偏見方面的潛力。
Large Models for Small Devices: Recent Advances and Empirical Analysis of Edge AI Deployment
2608.15693v1 by Subhransu Das, Jiaming Cheng, Arnav Kumar, Sadia Afrose, Mingzhe Han, Michael Silagy, Shreya Palande, Brijesh Soni, Rajiv Ramnath
Running large AI models on resource-constrained edge devices requires model compression to reduce model size and computation. What compresses well, however, need not deploy well. We survey dozens of recent works that report compression results on real hardware and extract practical deployment guidelines from them. Following these guidelines, we deploy compact language and image models on GPU, CPU, and Raspberry Pi platforms across question answering and image segmentation. No single technique wins across tasks. For question answering, Qwen3.5 0.8B reaches 93.85 SQuAD F1 and 92 EM under Q5_K_M GGUF quantization, while structured pruning at the same precision costs 16 F1 at a 1% ratio. For segmentation, the ranking reverses: default quantization leaves parameters and MACs unchanged, whereas pruning cuts model size by nearly 80% at near-constant mIoU. Pruning can even inflate the deployed artifact by 21-49% by breaking k-quant super-block alignment; combined with longer, less format-compliant outputs, this raises Raspberry Pi latency up to 3.4x. Compression can also manufacture the appearance of competence rather than destroy it visibly: one LoRA-recovered variant stays fully parseable and holds 71% strict BoolQ accuracy while sending 97 of 100 predictions to a single class, at 52.6% balanced accuracy. We explain these effects through neural-flow graph analysis and prefill-decode-level latency decomposition, and condense them into task-specific deployment research directions. The right technique depends on the task, the model, and the hardware. Our experiment code and artifacts are open-sourced at https://github.com/Arnavvvkumar/deployment
摘要:在資源受限的邊緣設備上運行大型 AI 模型需要模型壓縮,以減少模型大小和計算量。 然而,壓縮效果良好的模型不一定能夠良好部署。 我們調查了數十篇最近的研究,這些研究報告了在實際硬體上進行的壓縮結果,並從中提取實用的部署指南。 根據這些指南,我們在 GPU、CPU 和 Raspberry Pi 平台上部署了緊湊的語言和圖像模型,應用於問題回答和圖像分割。 沒有單一技術在所有任務中都能獲勝。 在問題回答中,Qwen3.5 0.8B 在 Q5_K_M GGUF 量化下達到 93.85 的 SQuAD F1 和 92 的 EM,而在相同精度下的結構化剪枝則以 1% 的比例損失 16 的 F1。 對於分割,排名則顛倒:默認量化保持參數和 MACs 不變,而剪枝則在幾乎不變的 mIoU 下將模型大小減少近 80%。 剪枝甚至可以通過打破 k-quant 超塊對齊來使部署的工件膨脹 21-49%;結合更長且格式不合規的輸出,這使得 Raspberry Pi 的延遲增加至 3.4 倍。 壓縮還可以製造出能力的外觀,而不是明顯地摧毀它:一個 LoRA 恢復的變體保持完全可解析,並在將 100 次預測中的 97 次發送至單一類別的同時,保持 71% 的嚴格 BoolQ 準確率,平衡準確率為 52.6%。 我們通過神經流圖分析和預填充解碼級延遲分解解釋這些效果,並將其濃縮為特定任務的部署研究方向。 正確的技術取決於任務、模型和硬體。 我們的實驗代碼和工件已在 https://github.com/Arnavvvkumar/deployment 開源。
NumerosityVLM: A Cognitively Inspired Benchmark for Interpreting Numerosity Representations in Vision-Language Models
2608.15425v1 by Yiming Fu, Fangjun Li, Xiujin Liu, Ruidong Ma, Hang Yu, Zhichen Lu, Kanwei He, Alessandro Di Nuovo, Angelo Cangelosi, Zhegong Shangguan
Vision-language models (VLMs) achieve strong performance on high-level multimodal tasks, yet numerosity perception, a cognitive ability that emerges in human infants before language acquisition, remains poorly understood in current models, as existing counting benchmarks entangle numerosity with correlated visual factors. We introduce a cognitively inspired diagnostic benchmark, NumerosityVLM, comprising 10,800 synthetic images across six controlled conditions. The benchmark orthogonally manipulates object size, spatial arrangement, and numerosity, while progressively ablating texture, shape, and color. Evaluating seven VLMs in a zero-shot setting, multi-factor analysis reveals that model architecture explains the largest proportion of performance variance (partial $ω^{2}=0.325$), far exceeding visual conditions. Layer-wise probing further shows that linearly separable numerosity signals consistently emerge at early stages of the vision encoder, while performance differences across evaluated models are primarily associated with the language model component. Code and data are publicly available at https://github.com/fuy3/NumerosityVLM-Benchmark, and https://huggingface.co/datasets/fuy3/NumerosityVLM.
摘要:視覺-語言模型(VLMs)在高階多模態任務上表現出色,然而數量感知這一認知能力在語言習得之前便出現在人類嬰兒身上,卻在當前模型中仍然不甚了解,因為現有的計數基準將數量感知與相關的視覺因素混合在一起。我們引入了一個受到認知啟發的診斷基準,NumerosityVLM,包含10,800張合成圖像,涵蓋六種受控條件。該基準正交地操控物體大小、空間排列和數量感知,同時逐步去除紋理、形狀和顏色。在零樣本設置中評估七個VLM,進行多因素分析顯示,模型架構解釋了性能變異的最大比例(部分 $ω^{2}=0.325$),遠遠超過視覺條件。逐層探測進一步顯示,線性可分的數量感知信號在視覺編碼器的早期階段持續出現,而評估模型之間的性能差異主要與語言模型組件相關。代碼和數據可在 https://github.com/fuy3/NumerosityVLM-Benchmark 和 https://huggingface.co/datasets/fuy3/NumerosityVLM 獲得。
ETHOS: Towards a Modular Ethics Framework for Clinical Multi-Agent Systems
2608.15424v1 by Rakesh Sharma, Sydney Pugh, Cameron Beeche, Pankhuri Singhal, Rachel Wu, Margaret Eby, Jeffrey Duda, James Gee, Kyra O'Brien, Hersh Sagreiya, Marina Serper, Victoria Gershuni, Angela Bradbury, Anurag Verma, Eric Eaton, Kevin B. Johnson, Walter Witschey
The rapid adoption of large language models has enabled the development of clinical multi-agent systems (MAS) capable of integrating multimodal patient data and supporting increasingly complex clinical decision-making. However, the deployment of these systems in real-world healthcare settings raises critical ethical concerns related to safety, fairness, accountability, transparency, and patient trust. While numerous organizations, including the World Health Organization, the National Academy of Medicine, and the FUTURE-AI consortium, have proposed ethical frameworks and governance principles for healthcare AI, these efforts remain largely conceptual. To address this challenge, we present ETHOS (Ethics and Trust through Hierarchical Oversight System), a modular ethics framework designed as a governance meta-agent that can be integrated with any existing multi-agent system without requiring changes to its underlying architecture. ETHOS translates stakeholder-informed ethical requirements into executable runtime oversight through a layered governance approach consisting of deterministic checks, contextual reviews, and a final ethics critic. These components continuously evaluate intermediate reasoning steps and final outputs, enabling the system to identify ethical risks, request revisions, or suppress responses that fail predefined safety and trustworthiness criteria. We demonstrate ETHOS within a hepatology clinical decision-support MAS. Results show that ETHOS improves decision reliability by detecting incomplete, inconsistent, or out-of-scope evidence and appropriately increasing abstention when safe recommendations cannot be supported. By embedding ethical governance directly into system operation, ETHOS provides a practical and auditable mechanism for transforming high-level AI ethics principles into deployable safeguards.
摘要:大型語言模型的快速採用使得臨床多代理系統(MAS)的發展成為可能,這些系統能夠整合多模態病人數據並支持日益複雜的臨床決策。
然而,這些系統在現實世界醫療環境中的部署引發了與安全、公平、問責、透明度和病人信任相關的重大倫理問題。
儘管包括世界衛生組織、國家醫學院和FUTURE-AI聯盟在內的許多組織已經提出了針對醫療AI的倫理框架和治理原則,但這些努力仍然主要是概念性的。
為了解決這一挑戰,我們提出了ETHOS(通過分層監督系統實現倫理與信任),這是一個模塊化的倫理框架,設計為一個治理元代理,可以與任何現有的多代理系統集成,而無需改變其底層架構。
ETHOS將利益相關者所知的倫理要求轉化為可執行的運行時監督,通過一種分層治理方法,包括確定性檢查、上下文審查和最終倫理評估。
這些組件持續評估中間推理步驟和最終輸出,使系統能夠識別倫理風險、請求修訂或抑制不符合預定安全和可信標準的回應。
我們在一個肝病臨床決策支持MAS中展示了ETHOS。
結果顯示,ETHOS通過檢測不完整、不一致或超出範疇的證據來提高決策的可靠性,並在無法支持安全建議時適當地增加放棄。
通過將倫理治理直接嵌入系統運作中,ETHOS提供了一種實用且可審計的機制,將高層次的AI倫理原則轉化為可部署的保障措施。
When AI Rewrites, Classifiers Relax: Uncertainty-Aware Sentiment Analysis on Sarcastic and AI-Paraphrased Social Text
2608.15338v1 by Shresth Shroff
Sentiment classifiers are increasingly applied to social media content that is either sarcastic or AI-generated --- two distributional regimes where standard evaluations offer little guidance. We present a three-part empirical study of sentiment classifier behaviour under these conditions. First, we find that confidence scores on sarcastic text are significantly lower than on non-sarcastic text (Mann--Whitney $p = 2 \times 10^{-6}$), confirming that classifiers sense their own uncertainty on ironic content even without explicit uncertainty modelling. Second, and counterintuitively, we show that sentiment classifiers achieve higher accuracy on AI-paraphrased reviews than on the original human-authored text (RoBERTa: $+5.8$ pp for Qwen3.5-4B paraphrases, $+3.7$ pp for Gemma4-E4B), revealing a cross-domain stylistic alignment effect: AI paraphrases remove distributional noise that confounds Twitter-trained classifiers, producing cleaner, more prototypical sentiment text. Third, we demonstrate that a lightweight abstention wrapper --- flagging the $14\%$ of inputs with confidence below $0.6$ --- improves accuracy from 82.2\% to 88.9\% ($+6.7$ pp) on the retained set. We further compare Semantic Entropy and MC-Dropout-style disagreement as uncertainty signals and find near-identical AUROC ($0.650$ vs.\ $0.646$) on sarcastic text, suggesting that for short social media inputs, both methods are interchangeable. Our results motivate a shift from confident single-label prediction to uncertainty-aware abstention in high-stakes sentiment applications such as mental health flagging and content moderation.
摘要:情感分類器越來越多地應用於諷刺或 AI 生成的社交媒體內容——這兩種分佈模式下,標準評估提供的指導有限。我們提出了一項三部分的實證研究,探討情感分類器在這些條件下的行為。首先,我們發現對於諷刺文本的信心分數顯著低於非諷刺文本(Mann--Whitney $p = 2 \times 10^{-6}$),確認分類器即使在沒有明確不確定性建模的情況下,也能感知到對於諷刺內容的自身不確定性。其次,反直覺的是,我們顯示情感分類器在 AI 改寫的評論上取得的準確率高於原始的人類撰寫文本(RoBERTa: $+5.8$ pp 對於 Qwen3.5-4B 改寫,$+3.7$ pp 對於 Gemma4-E4B),揭示了一種跨領域的風格一致性效應:AI 改寫去除了困擾 Twitter 訓練的分類器的分佈噪音,產生了更乾淨、更原型的情感文本。第三,我們證明了一種輕量級的棄權包裝——標記信心低於 $0.6$ 的 $14\%$ 輸入——在保留數據集上將準確率從 82.2\% 提高到 88.9\%($+6.7$ pp)。我們進一步比較語義熵和 MC-Dropout 風格的不一致作為不確定性信號,發現對於諷刺文本,兩者的 AUROC 幾乎相同($0.650$ 對 $0.646$),這表明對於短的社交媒體輸入,這兩種方法是可以互換的。我們的結果促使從自信的單標籤預測轉向在高風險情感應用中,如心理健康標記和內容審核,意識到不確定性的棄權。
Demographic Injection in Medical Language Models under Diversity, Equity, and Inclusion Prompts
2608.15254v1 by Diego Mardian, Frank Liu
Clinical-AI guidance increasingly recommends prompting language models to reason with attention to diversity, equity, and inclusion (DEI). We measure a side effect that misrepresents patients: a one-sentence DEI prompt appended to a medical question leads models to add patient demographic attributes (race, socioeconomic status, sex) the question never stated, in effect rewriting who the patient is. We call this demographic injection. Across 47 models, four medical benchmarks, and 376,000 responses scored by a validated model-judge pipeline, a single DEI prompt raises the injection rate from 0.7% to 33.1% (47x) in all 47 of 47 models, attributable to the equity content rather than to added length (18x above a length-matched control; p=1.4x10^-14). Most added content is a general population statement that leaves the answer unchanged, but a smaller subset attaches an attribute to the specific patient or changes the selected option (0.25-2.4% of responses, 99.8% toward the incorrect option), where the invented demographic changes the answer the model recommends. Phrasing scales the effect from 14% to 56%. DEI prompts are just one example of a more general mechanism. Any instruction that nudges how a model reasons can make it add unrequested details, including details about the patient. Flagged outputs are treated as model errors under study, not clinical guidance.
摘要:臨床人工智慧指導越來越多地建議促使語言模型在考慮多樣性、公平性和包容性(DEI)時進行推理。我們測量了一種誤導患者的副作用:一個附加在醫療問題上的單句DEI提示會導致模型添加問題中從未提到的患者人口統計屬性(種族、社會經濟地位、性別),實際上重寫了患者的身份。我們稱之為人口統計注入。在47個模型、四個醫療基準和376,000個由經過驗證的模型評判管道評分的回應中,單一的DEI提示使得注入率從0.7%上升到33.1%(47倍),這是由於公平性內容而非增加的長度(在長度匹配的對照組中增加了18倍;p=1.4x10^-14)。大多數新增內容是一般人口的陳述,對答案沒有改變,但一小部分則將屬性附加到特定患者或改變所選選項(0.25-2.4%的回應,99.8%朝向不正確的選項),其中虛構的人口統計改變了模型推薦的答案。措辭將效果擴大至14%至56%。DEI提示僅是更一般機制的一個例子。任何促使模型推理的指令都可能使其添加未請求的細節,包括有關患者的細節。被標記的輸出被視為正在研究的模型錯誤,而非臨床指導。
Low-Rank Dynamics-Effective Latent Carriers for Counterfactual Rollout in Learned World Models
2608.15156v1 by Yang Liu, Yuming Chen
World models may predict the future without making clear which parts of their hidden state actually drive those predictions. We ask whether a small, directly addressable hidden-state change can place a learned world model on the intended counterfactual trajectory and then let the model continue that future on its own. We study a recurrent world model with a 192-dimensional hidden state in a controlled two-object, two-dimensional collision environment. For a bounded family of local velocity edits, we first verify that the model can natively represent and roll out the edited future. We then construct candidate low-rank carriers from training-only factual-to-counterfactual hidden differences and learn a map from the factual state and requested edit to carrier coefficients. On the registered rank grid, rank 4 is the smallest tested rank that satisfies the full development-panel criteria. A single rank-4 patch at the anchor is sufficient to redirect a 12-step autonomous rollout, with no future observations, teacher forcing, or repeated correction. The frozen procedure satisfies the preregistered replication rule across independently trained checkpoints and remains usable across nearby intervention times. Random equal-norm, wrong-object, and wrong-time controls do not explain the effect. A position-edit stress test provides a negative contrast: the intended position patch can pass the raw rollout criteria, but no-patch and random controls can pass the same criteria, and wrong-object specificity is not established. Thus, successful editing alone is not enough. We use dynamics-effective to describe an intervention that changes the model's future computation in a sustained and target-specific way under autonomous rollout. The rank-4 result identifies a compact intervention interface for the tested velocity-edit family, not a closed four-dimensional state or an intrinsic state dimension.
摘要:世界模型可能預測未來,但並未明確指出其隱藏狀態的哪部分實際驅動這些預測。我們詢問是否可以透過一個小的、可直接訪問的隱藏狀態變化,將學習到的世界模型置於預期的反事實軌跡上,然後讓模型自行繼續那個未來。我們研究了一個具有192維隱藏狀態的遞歸世界模型,在一個受控的兩物體、二維碰撞環境中。對於一個有界的局部速度編輯家族,我們首先驗證模型是否能夠原生地表示並展開編輯後的未來。然後,我們從僅訓練的事實到反事實的隱藏差異中構建候選低秩載體,並學習從事實狀態和請求編輯到載體係數的映射。在註冊的秩網格上,秩4是滿足完整發展面板標準的最小測試秩。在錨點處,單個秩4的補丁足以重定向一個12步的自主展開,無需未來觀察、教師強迫或重複修正。凍結程序滿足預註冊的複製規則,並在獨立訓練的檢查點之間保持可用,並在附近的干預時間內保持可用。隨機等範數、錯誤物體和錯誤時間的控制無法解釋該效果。一個位置編輯壓力測試提供了負對比:預期的位置補丁可以通過原始展開標準,但無補丁和隨機控制也可以通過相同標準,且未建立錯誤物體的特異性。因此,僅僅成功編輯是不夠的。我們使用動態有效來描述一種干預,該干預在自主展開下以持續且目標特定的方式改變模型的未來計算。秩4的結果確定了對測試的速度編輯家族的緊湊干預介面,而不是封閉的四維狀態或內在狀態維度。
Fast Test-Time Refinement for Robust Learned Image Compression
2608.15113v1 by Jiaming Liang, Chi-Man Pun, Weisi Lin
Learned image compression (LIC) has demonstrated remarkable rate-distortion (RD) performance in benign settings. However, the high representational capacity endowed by deep neural networks (DNNs) comes at the expense of increased adversarial vulnerability. This hinders their adoption as trusted standardized codecs. Recent work has sketched test-time refinement (TTR) as a defense in gray-box scenarios, despite its original purpose of improving benign RD performance. Unfortunately, extensive iterations of TTR incur prohibitive overhead, while the robustness mechanism lacks theoretical understanding. Moreover, TTR has not been evaluated in white-box settings or against attacks beyond $\ell_2$-bounded rate and untargeted distortion objectives. To bridge these gaps, we present a systematic study. Our study reveals an Asymmetric Adversarial Trajectory (AAT) property in LIC systems: transitioning from adversarial to benign regions is significantly easier than the reverse process, where adversarial examples can often be roughly recovered within only 1-2 steps. We provide a two-dimensional Tube Model to explain this phenomenon. Based on AAT, we propose a Fast Test-Time Refinement (FTTR) framework for practical and robust LIC systems. We establish that the robustness arises from the contraction of adversarial regions induced by the Input-as-Label property of LIC systems, rather than from obfuscated gradients. Extensive evaluations with diverse strong adaptive attacks across multiple LIC systems demonstrate the promise of the proposed FTTR framework. The code is available at https://github.com/chinaliangjiaming/FTTR.git.
摘要:學習型影像壓縮(LIC)在良性環境中展現了卓越的率失真(RD)性能。然後,深度神經網絡(DNN)所賦予的高表徵能力卻以增加對抗脆弱性為代價。這阻礙了它們作為可信標準編碼器的採用。最近的研究勾勒出測試時精煉(TTR)作為灰箱場景中的防禦,儘管其最初目的是改善良性的RD性能。不幸的是,TTR的廣泛迭代會產生高昂的開銷,而其穩健性機制缺乏理論理解。此外,TTR尚未在白箱環境中進行評估,也未針對超出$\ell_2$-界限率和非針對性失真目標的攻擊進行評估。為了填補這些空白,我們提出了一項系統研究。我們的研究揭示了LIC系統中的不對稱對抗軌跡(AAT)特性:從對抗區域轉移到良性區域顯著容易於反向過程,其中對抗範例通常可以在僅1-2步內粗略恢復。我們提供了一個二維管道模型來解釋這一現象。基於AAT,我們提出了一個快速測試時精煉(FTTR)框架,旨在實現實用且穩健的LIC系統。我們確立了穩健性源於LIC系統的輸入作為標籤屬性所引起的對抗區域收縮,而非來自模糊梯度。對多個LIC系統進行的廣泛評估,涵蓋多種強適應性攻擊,展示了所提出的FTTR框架的潛力。代碼可在https://github.com/chinaliangjiaming/FTTR.git獲得。
Command-Space Counterfactual Explanations for Pareto-Conditioned Reinforcement Learning
2608.14963v1 by Joanikij Chulev, Hendrik Baier
Pareto Conditioned Networks learn multiple multi-objective reinforcement learning behaviours by conditioning a single policy on a desired return command. However, the local mapping from command and state to action remains opaque. We propose command-space counterfactual explanations for PCNs: given a fixed state, original command, and foil action, we search, in a black-box setting, for a minimally changed desired-return command under which the same trained policy would choose the foil. Our contributions are threefold. First, we formulate PCN explanations as return-command interventions, using a return-only PCN variant that avoids the added ambiguity of horizon-conditioning. Second, we adapt adversarial machine learning methods to reinforcement-learning explanations. Third, we introduce a boundary-seeded directional search that improves over purely local optimization in the command-action landscape, resulting in our proposed approach CF-ZOO. The resulting explanations are actionable and intuitively expressed in the user's own preferences: "If your trade-off had shifted slightly towards X, the agent would have chosen Y."
摘要:Pareto Conditioned Networks 通過將單一策略條件化於期望回報命令來學習多種多目標強化學習行為。
然而,從命令和狀態到行動的局部映射仍然不明朗。
我們為 PCNs 提出了命令空間的反事實解釋:在固定的狀態、原始命令和對照行動下,我們在黑箱環境中搜索一個最小變更的期望回報命令,在此命令下,相同的訓練策略會選擇對照行動。
我們的貢獻有三個方面。
首先,我們將 PCN 解釋公式化為回報命令的干預,使用一種僅回報的 PCN 變體,避免了地平線條件化所帶來的額外模糊性。
其次,我們將對抗性機器學習方法適應於強化學習解釋。
第三,我們引入了一種邊界引導的方向性搜索,這在命令-行動空間中優於純粹的局部優化,從而形成我們提出的方法 CF-ZOO。
所得的解釋是可操作的,並以用戶自己的偏好直觀表達:“如果你的權衡稍微向 X 方向移動,代理將會選擇 Y。”
Handover Analysis for Vehicular Communication with Explainability on the Fly
2608.14820v1 by Ali Fuat Sahin, Semiha Tedik Başaran, Tufan Kumbasar
Handover (HO) management in vehicular networks requires fast and reliable decision-making under highly dynamic conditions. While machine learning (ML) approaches can improve HO detection by capturing complex relationships among various key performance indicators (KPIs), their black-box nature limits interpretability and operator trust. To address this, this paper investigates HO detection from an explainability-on-the-fly perspective using inherently interpretable models based on the functional analysis of variance (fANOVA) framework. The proposed models are evaluated using two real-world operator datasets and compared against a Long Short-Term Memory baseline augmented with post-hoc SHAP explanations. Unlike post-hoc approaches, the proposed framework enables immediate interpretation of model decisions without incurring additional computational overhead. This capability is particularly critical for latency-sensitive vehicular networks. The results show that fANOVA-based models achieve competitive detection performance while providing significantly reduced explanation latency compared to conventional post-hoc methods. Furthermore, feature ranking and visualization analyses reveal physically meaningful relationships between KPIs and HO occurrences that align with standardized HO mechanisms. These results demonstrate that inherently interpretable models provide an efficient and transparent solution for HO detection in next-generation vehicular networks.
摘要:Handover (HO) 管理在車輛網絡中需要在高度動態的條件下進行快速且可靠的決策。雖然機器學習 (ML) 方法可以通過捕捉各種關鍵性能指標 (KPI) 之間的複雜關係來改善 HO 檢測,但其黑箱特性限制了可解釋性和操作員的信任。為了解決這個問題,本文從即時可解釋性的角度研究了基於方差的功能分析 (fANOVA) 框架的 HO 檢測。所提出的模型使用兩個真實世界的運營商數據集進行評估,並與增強了後驗 SHAP 解釋的長短期記憶基準進行比較。與後驗方法不同,所提出的框架能夠立即解釋模型決策,而不會產生額外的計算開銷。這一能力對於對延遲敏感的車輛網絡尤為重要。結果顯示,基於 fANOVA 的模型在檢測性能上具有競爭力,同時提供顯著降低的解釋延遲,相較於傳統的後驗方法。此外,特徵排名和可視化分析揭示了 KPI 與 HO 發生之間的物理意義關係,這與標準化的 HO 機制一致。這些結果表明,固有可解釋的模型為下一代車輛網絡中的 HO 檢測提供了一種高效且透明的解決方案。
Generated Context versus Governed State: Functional Conditions for Accountable Longitudinal Clinical Reasoning
2608.14804v1 by Augusto Bernardo Pissarra, Victor Lorena de Farias Souza
Large language models (LLMs) have become the dominant interface of clinical artificial intelligence, yet the interface they expose (text in, text out, one context window at a time) maintains no explicit, persistent, governed representation of what is currently true about a patient. This paper argues that longitudinal clinical reasoning is a state-estimation problem under partial observability, and that the axis on which clinical AI succeeds or fails is not the fluency of the model reading the record but the governance of the patient state it reasons over. We distinguish generated context from governed state; separate five objects that clinical AI habitually conflates (true state, observations, evidence, belief, and simulated state); define a tiered governance standard against which any clinical AI system can be audited; and show that an operational definition of accountability decomposes into four information requirements: an immutable evidence ledger with awareness-time versioning, a belief state distinct from accumulated evidence, an observation-process model, and claim-level causal typing. We are explicit that this decomposition is analytic rather than a necessity theorem, and that its value is conceptual hygiene: it converts "accountable clinical AI" from a slogan into an audit instrument. A six-level maturity framework separates what a system makes governable from what it can compute, locating current LLM-centric practice at high capability but low maturity. The paper is fully self-contained: the four research questions the framework poses are stated in the introduction, and the conclusion records what the paper establishes toward each; future work develops the buildable core of the architecture and the research program toward full Clinical World Models. No empirical result is claimed here.
摘要:大型語言模型(LLMs)已成為臨床人工智慧的主導介面,但它們所暴露的介面(文本輸入、文本輸出、一次一個上下文窗口)並未對目前有關患者的真實情況提供明確、持久、受管控的表徵。本文主張,縱向臨床推理是一個在部分可觀察性下的狀態估計問題,而臨床 AI 成功或失敗的軸心不在於模型閱讀記錄的流暢性,而在於它所推理的患者狀態的治理。我們區分生成的上下文與受管控的狀態;將臨床 AI 通常混淆的五個對象(真實狀態、觀察、證據、信念和模擬狀態)分開;定義一個分層治理標準,以便對任何臨床 AI 系統進行審計;並顯示一個運作性責任的定義可分解為四個信息要求:具有意識時間版本控制的不可變證據賬本、與累積證據不同的信念狀態、觀察過程模型,以及索賠級別的因果類型。我們明確指出這一分解是分析性的,而非必要定理,其價值在於概念衛生:它將“可負責任的臨床 AI”從口號轉變為審計工具。一個六級成熟度框架將系統可治理的部分與其可計算的部分分開,將當前以 LLM 為中心的實踐定位於高能力但低成熟度。本文是完全自足的:框架提出的四個研究問題在引言中陳述,結論記錄了本文在每個問題上所建立的內容;未來的工作將發展可構建的架構核心及通向完整臨床世界模型的研究計劃。此處不聲稱任何實證結果。
Decoding the Past: An Uncertainty-Aware Deep Learning Framework for Sex Attribution in Prehistoric Hand Stencils
2608.14539v1 by Karel Becerra, Boris Mederos, Dean Snow, Ramón A. Mollineda
Determining the biological sex of the individuals who created Upper Paleolithic hand stencils remains a challenging problem due to the absence of ground truth, population differences between contemporary and prehistoric groups, and the uncertainty introduced by image degradation. Traditional morphometric methods suffer from high structural overlap across sexes, poor cross-population generalizability, and subjective feature engineering. This study presents an uncertainty-aware deep learning framework for sex attribution in prehistoric hand stencils that explicitly models, propagates, and aggregates uncertainty throughout the analytical pipeline. The methodology combines dual image processing, dual contour extraction, structured silhouette augmentation, model architectural diversity, and ensemble-based decision aggregation. The pipeline generates twelve plausible silhouette realizations per stencil to capture boundary uncertainties, which are processed by two ensembles of ten deep neural networks each (EfficientNet-B3 and MobileViT-S) trained on 14,036 contemporary hand samples. Furthermore, a triangulated validation scheme integrates ensemble predictions with unsupervised 2D latent-space manifold mapping (UMAP + k-NN) and explainable AI spatial attributions (LayerCAM) to ensure anatomical consistency. On contemporary data, ensemble models achieve strong classification performance, with accuracies exceeding 88% in older age groups. When applied to prehistoric stencils, the framework produces both sex predictions and confidence measures of internal agreement, enabling the distinction between morphologically stable and ambiguous cases. Convergence across ensemble predictions, latent-space structure, and interpretability analyses shows that uncertainty can become a measurable component of archaeological inference, enabling robust and reproducible decoding of ancient rock art.
摘要:確定創造上舊石器時代手印的個體的生物性別仍然是一個具有挑戰性的問題,這是由於缺乏真實數據、當代與史前群體之間的差異,以及圖像劣化所帶來的不確定性。傳統的形態計量方法在性別之間存在高度的結構重疊、跨群體的普遍性差以及主觀的特徵工程。這項研究提出了一個不確定性感知的深度學習框架,用於史前手印中的性別歸屬,該框架明確地建模、傳播和聚合整個分析流程中的不確定性。該方法結合了雙重圖像處理、雙重輪廓提取、結構化輪廓增強、模型架構多樣性和基於集成的決策聚合。該流程為每個手印生成十二個合理的輪廓實現,以捕捉邊界不確定性,這些輪廓由兩個各包含十個深度神經網絡的集成處理(EfficientNet-B3 和 MobileViT-S),這些網絡是在 14,036 個當代手樣本上訓練的。此外,一個三角驗證方案將集成預測與無監督的 2D 潛在空間流形映射(UMAP + k-NN)和可解釋的 AI 空間歸因(LayerCAM)結合,以確保解剖學的一致性。在當代數據上,集成模型實現了強大的分類性能,年齡較大的群體的準確率超過 88%。當應用於史前手印時,該框架產生性別預測和內部一致性的信心度量,從而能夠區分形態穩定和模糊的案例。集成預測、潛在空間結構和可解釋性分析之間的收斂顯示,不確定性可以成為考古推理的一個可測量組成部分,使古代岩畫的解碼變得穩健且可重複。
NARRATE: A Multimodal Real-World Australian Driving Dataset for Human-Centred Explanations in Automated Driving
2608.14767v1 by Ashkan Yousefi Zadeh, Zishuo Zhu, Xiaomeng Li, Andry Rakotonirainy, Sebastien Glaser, Ronald Schroeter, Patricia Delhomme, Zahra Mehraban
Automated vehicles must explain their decisions in ways that passengers can understand, monitor, and trust. Existing language-annotated driving datasets are mostly observer-written, post-hoc, simulation-based, or generated from sensor inputs, rather than elicited from the driver performing the action. We introduce NARRATE, a multimodal real-world Australian driving dataset comprising 2,050 annotated events from 35 experienced drivers and driving instructors on public roads. Each event is grounded in synchronised visual, localisation, motion, and LiDAR streams and paired with in-vehicle and/or post-drive free-text explanations. NARRATE provides action labels, scenario-context labels spanning six high-level and 32 fine-grained categories, and span-level Situational Awareness (SA) annotations over driver explanations for Perception, Comprehension and Projection. Four benchmark tasks (SA, scenario-context, driver-action classification, and explanation generation) show that this structure is learnable from driver language, while fine-grained context recognition and explanation generation remain challenging. NARRATE paves a path towards more human-centred and domain-aware explanation models for automated driving.
摘要:自動駕駛車輛必須以乘客能理解、監控和信任的方式解釋其決策。現有的語言標註駕駛數據集大多是觀察者撰寫的、事後的、基於模擬的,或是從傳感器輸入生成的,而不是從執行動作的駕駛員那裡引出來的。我們介紹了NARRATE,這是一個多模態的澳大利亞實際駕駛數據集,包含來自35位經驗豐富的駕駛員和駕駛教練在公共道路上標註的2,050個事件。每個事件都基於同步的視覺、定位、運動和LiDAR數據流,並配有車內和/或駕駛後的自由文本解釋。NARRATE提供了動作標籤、涵蓋六個高層次和32個細分類別的情境上下文標籤,以及針對感知、理解和預測的駕駛員解釋的跨度級別情境意識(SA)註釋。四個基準任務(SA、情境上下文、駕駛員動作分類和解釋生成)顯示這一結構可以從駕駛員語言中學習,而細緻的上下文識別和解釋生成仍然具有挑戰性。NARRATE為自動駕駛的更人性化和領域意識的解釋模型鋪平了道路。
Polaris : Multi Agentic System for Conversational Enterprise Analytics
2608.14246v1 by Varuni H K, Soham Sarkar, Jay Kumar, Goutham Krishnan, Tanvi Johari, Avinash Bharadwaj, Santosh Hegde
In today's fast-paced environment, the ability to swiftly access, understand, and act on data is no longer optional; it is essential. Yet most organizations remain data-rich but insight-poor, constrained by the complexity of querying, interpreting, and explaining enterprise-scale information. We present Polaris, a supervisor-led multi-agent framework for conversational enterprise analytics that bridges this gap. Polaris introduces Dynamic Task Coordination (DTC), a decision-theoretic orchestration layer that models agent-task assignment as adaptive bipartite matching, enabling real-time coordination, recovery, and optimization across specialized agents for querying, visualization, and reasoning. By coupling DTC with reason-first, ReAct-style agents, Polaris transforms natural-language queries into coherent analytical workflows that not only retrieve and visualize data but also explain the underlying "why." Evaluation on structured enterprise datasets demonstrates high semantic fidelity and answer relevancy, underscoring the potential of multi-agent orchestration to deliver trustworthy, end-to-end business intelligence at scale.
摘要:在當今快速變化的環境中,迅速訪問、理解和處理數據的能力不再是可選的;它是必需的。然而,大多數組織仍然是數據豐富但洞察貧乏,受到查詢、解釋和解釋企業級信息的複雜性所限制。我們提出了Polaris,一個由主管主導的多代理框架,用於對話式企業分析,填補了這一空白。Polaris引入了動態任務協調(DTC),這是一個決策理論的編排層,將代理-任務分配建模為自適應二分匹配,實現了專業代理之間的實時協調、恢復和優化,用於查詢、可視化和推理。通過將DTC與以推理為首的ReAct風格代理結合,Polaris將自然語言查詢轉化為連貫的分析工作流程,這些工作流程不僅檢索和可視化數據,還解釋了背後的“為什麼”。對結構化企業數據集的評估顯示出高語義保真度和答案相關性,強調了多代理編排在大規模提供可靠的端到端商業智能方面的潛力。
CForce: Boosting Parallel Decoding for dLLMs via Consistency Forcing
2608.13925v1 by Yuji Ren, Chenkai Xu, Zhuocheng Gong, Jianguo Li, Zhijie Deng
Diffusion large language models (dLLMs) accelerate language generation by predicting multiple masks in a single forward pass. However, existing dLLMs can suffer from unreliable predictions in early denoising stages under aggressive parallelism strategies, leading to errors that can propagate to later stages. To tackle this issue, we present Consistency Forcing (CForce) for dLLMs, a distillation method to force the mask predictions of early stages to align with those of later stages. CForce trains the model on pre-collected self-rollout trajectories, thereby improving training-inference alignment. We introduce Confidence Adaptive KL Divergence as a distillation objective to conjoin the merits of forward and reverse KL. We further provide a theoretical analysis for the consistency objective to explain why CForce can approximately minimize the prediction error of early stages. Critically, the same formulation applies to both mask-to-token decoding and edit-capable decoding; in the edit-capable case, later token-to-token refinements provide additional supervision for earlier masked-state predictions. Experiments on non-edit and edit-capable LLaDA models show improved speed-quality trade-offs, especially under high-parallelism decoding budgets. Code is available at: https://github.com/inclusionAI/dFactory.
摘要:擴散大型語言模型(dLLMs)通過在單次前向傳播中預測多個掩碼來加速語言生成。
然而,現有的 dLLMs 在激進的並行策略下,可能在早期去噪階段出現不可靠的預測,導致錯誤可能傳播到後續階段。
為了解決這個問題,我們提出了一種名為一致性強制(CForce)的 dLLMs 蒸餾方法,以強制早期階段的掩碼預測與後期階段對齊。
CForce 在預先收集的自我展開軌跡上訓練模型,從而改善訓練與推理的對齊。
我們引入了信心自適應 KL 散度作為蒸餾目標,以結合前向和反向 KL 的優點。
我們還提供了一個一致性目標的理論分析,以解釋為什麼 CForce 可以大致最小化早期階段的預測誤差。
關鍵的是,這種相同的公式適用於掩碼到標記的解碼和可編輯解碼;在可編輯的情況下,後期的標記到標記的細化為早期的掩碼狀態預測提供了額外的監督。
在非編輯和可編輯的 LLaDA 模型上的實驗顯示了速度和質量的權衡改善,特別是在高並行解碼預算下。
代碼可在以下網址獲得:https://github.com/inclusionAI/dFactory。
Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions
2608.13786v1 by Qingfang Liu, Qiao Jin, Joe D. Menke, Thorsten Kahnt, Zhiyong Lu
Large language model (LLM) chatbots are increasingly used to answer clinical questions with citations to relevant clinical studies. Prior research has largely focused on citation fabrication, leaving a gap in evaluating the quality of retrieved studies and the factors driving their selection. In this study, we evaluated three general-purpose LLM chatbots: Claude Sonnet 5, Gemini 3.1 Pro, and ChatGPT GPT-5.5. We prompted the models with clinical questions adapted from 20 review questions in Issues 6 and 7 of the 2026 Cochrane Database of Systematic Reviews, simulating patient, clinician, and evidence-synthesis researcher roles. Each chatbot was queried under each user role with four independent repetitions, yielding 720 responses. Each chatbot was asked to support its answers with primary clinical citations, which we benchmarked against the included and excluded study sets of the Cochrane reviews. On average, a chatbot response retrieved 39.2% $\pm$ 29.8% of Cochrane included studies, while citing 5.0% $\pm$ 9.4% of excluded studies. Recall of Cochrane included studies varied significantly by model and user role. ChatGPT achieved higher recall than Claude or Gemini (63.1% $\pm$ 29.5% vs. 37.0% $\pm$ 23.8% vs. 17.3% $\pm$ 13.1%; $p=2.0\times10^{-5}$). The researcher role yielded higher recall than the clinician or patient roles (42.8% $\pm$ 30.8% vs. 38.6% $\pm$ 28.9% vs. 36.1% $\pm$ 29.3%; $p=2.0\times10^{-5}$). Controlling for publication year, citations per year, and open-access status, sample size was the only independently significant predictor of retrieval (odds ratio 1.80 per 1-unit increase in log sample size, 95% CI 1.37-2.36, $p=2.34\times10^{-5}$). These findings suggest that while LLM chatbots can retrieve some studies identified by expert reviewers, their performance varies by model and user role, and they exhibit a bias toward clinical trials with larger sample sizes.
摘要:大型語言模型(LLM)聊天機器人越來越多地用於回答臨床問題,並引用相關的臨床研究。先前的研究主要集中在引用虛構上,未能評估檢索到的研究質量及其選擇的驅動因素。在本研究中,我們評估了三個通用型LLM聊天機器人:Claude Sonnet 5、Gemini 3.1 Pro和ChatGPT GPT-5.5。我們根據2026年Cochrane系統評價數據庫第6和第7期的20個回顧問題,為模型提供了臨床問題的提示,模擬患者、臨床醫生和證據綜合研究者的角色。每個聊天機器人在每個用戶角色下進行了四次獨立查詢,共產生720個回應。每個聊天機器人被要求用主要臨床引用來支持其答案,我們將其與Cochrane評估的納入和排除研究集進行了基準比較。平均而言,聊天機器人的回應檢索了39.2% $\pm$ 29.8%的Cochrane納入研究,同時引用了5.0% $\pm$ 9.4%的排除研究。Cochrane納入研究的回憶率因模型和用戶角色而異。ChatGPT的回憶率高於Claude或Gemini(63.1% $\pm$ 29.5% vs. 37.0% $\pm$ 23.8% vs. 17.3% $\pm$ 13.1%;$p=2.0\times10^{-5}$)。研究者角色的回憶率高於臨床醫生或患者角色(42.8% $\pm$ 30.8% vs. 38.6% $\pm$ 28.9% vs. 36.1% $\pm$ 29.3%;$p=2.0\times10^{-5}$)。在控制出版年份、每年引用數和開放獲取狀態後,樣本大小是唯一獨立顯著的檢索預測因子(對數樣本大小每增加1單位的比值比1.80,95% CI 1.37-2.36,$p=2.34\times10^{-5}$)。這些發現表明,儘管LLM聊天機器人可以檢索到一些專家評審者識別的研究,但其性能因模型和用戶角色而異,並且對樣本大小較大的臨床試驗存在偏見。
Capacity-Dependent Effects of Data Selection for Reasoning
2608.13721v1 by Cuong Dang, Hoang Anh Just, Ruoxi Jia
In reasoning supervised fine-tuning, candidate responses for the same instruction can differ substantially in how well they match the student's current distribution. Recent likelihood-based response selection methods suggest that responses closer to the student distribution provide more effective supervision, motivating the hypothesis that high-likelihood responses may generally be preferable for fine-tuning. In this paper, we revisit this intuition and show that the value of likelihood-based data selection depends critically on model capacity and training duration. Through controlled experiments on mathematical reasoning, using students ranging from 1.5B to 8B parameters and supervision generated by stronger teacher models, we observe a clear \emph{capacity-dependent} ``{\color{SMALLCOLOR}\textbf{Fast-Fit}} / {\color{LARGECOLOR}\textbf{Slow-Gain}}'' pattern. High-likelihood data provides faster and more stable early improvements, especially for smaller models, but low-likelihood data becomes increasingly beneficial for larger models when training is allowed to continue longer. To explain this phenomenon, we analyze learning dynamics, showing that small models often fail to absorb low-likelihood supervision and instead fall into shallow or repetitive behaviors, while larger models are better able to move toward the teacher distribution under such data. We further provide a capacity-constrained theoretical view of distillation that clarifies how data difficulty, data span, and student capacity jointly govern transfer. Overall, our findings show that effective data selection for reasoning should be aware of model capacity and computing budget rather than based on a single universal preference for high-likelihood supervision.
摘要:在推理的監督微調中,對於相同指令的候選回應在與學生當前分佈的匹配程度上可能存在顯著差異。最近基於似然的回應選擇方法表明,與學生分佈更接近的回應提供了更有效的監督,這促使了高似然回應在微調中通常更可取的假設。在本文中,我們重新檢視這一直覺,並顯示基於似然的數據選擇的價值在很大程度上依賴於模型容量和訓練持續時間。通過對數學推理進行控制實驗,使用從1.5B到8B參數的學生以及由更強的教師模型生成的監督,我們觀察到明顯的\emph{容量依賴} ``{\color{SMALLCOLOR}\textbf{快速擬合}} / {\color{LARGECOLOR}\textbf{慢增益}}''模式。高似然數據提供了更快且更穩定的早期改進,特別是對於較小的模型,但當訓練允許持續更長時間時,低似然數據對於較大模型變得越來越有利。為了解釋這一現象,我們分析了學習動態,顯示小模型往往無法吸收低似然監督,反而陷入淺層或重複的行為,而較大模型在這類數據下更能朝向教師分佈移動。我們進一步提供了一個受容量限制的蒸餾理論觀點,闡明了數據難度、數據範圍和學生容量如何共同影響轉移。總體而言,我們的研究結果顯示,對於推理的有效數據選擇應該考慮模型容量和計算預算,而不是基於對高似然監督的單一普遍偏好。
MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and Coordination
2608.13476v1 by Saisha Shetty, Satvik Tripathi, Austin Lin, Colin Zhao, Theodore Kim, Don Enwerem, Jacinta Arnold, Shahriar Faghani, Tessa S Cook
We present Multi-Agent Reasoning and Coordination (MARC), an open-source framework that replaces monolithic LLM prompting with deterministic multi-agent orchestration for clinical reasoning. MARC coordinates role-specialized agents for extraction, reasoning, answer generation, and evaluation, with explicit context passing and traceable intermediate outputs, enabling stage-wise failure attribution. We additionally introduce a Decomposer module that generates task-specific agent prompts from a plain-language description, eliminating manual prompt engineering. The framework supports both API-based and local CPU-compatible deployments and is entirely configurable via YAML, without code modifications. MARC is designed to be model-agnostic, interpretable, and accessible to clinical domain experts without programming expertise. The full framework is available at https://github.com/Penn-RAIL/MARC-v1.
摘要:我們提出了多智能體推理與協調(MARC),這是一個開源框架,將單一大型語言模型的提示替換為確定性的多智能體協調,用於臨床推理。MARC 協調角色專門的智能體進行提取、推理、答案生成和評估,具有明確的上下文傳遞和可追溯的中間輸出,實現階段性失敗歸因。我們還引入了一個分解器模塊,該模塊從普通語言描述中生成任務特定的智能體提示,消除了手動提示工程。該框架支持基於 API 的和本地 CPU 兼容的部署,並且完全可以通過 YAML 配置,而無需修改代碼。MARC 設計為模型無關、可解釋,並且對於沒有編程專業知識的臨床領域專家可訪問。完整框架可在 https://github.com/Penn-RAIL/MARC-v1 獲得。
A Unifying Perspective on Causal World Models: From Observations to Representations to Structure
2608.13456v1 by Avinash Kori, Fabrizio Russo
World Models (WM) are increasingly seen as a foundation for intelligent agents that can predict, plan, and act beyond their training distribution. In this paper, we study WMs from a causal perspective across multiple levels of abstraction, ranging from perceptual observations to building a conceptual representation of the structure governing the environment dynamics. We argue that useful WMs must go beyond generative capabilities alone: they should also capture entity properties, entity-to-entity interactions, and entity-to-environment interactions that determine and explain the dynamics of a system. We provide a formal definition of Causal WMs (CWMs) grounded in the tasks they are intended to support, connecting world modelling with existing work in causal representation learning, object-centric learning, causal discovery, structural causal models, and model-based decision-making. Finally, we relate CWMs to the literature on identifiability, clarifying when the components of a WM can be recovered from data and up to which equivalence. With this, we ground WMs in representations and structures that support causal reasoning and informed decision-making.
摘要:世界模型(WM)越來越被視為智能代理的基礎,這些代理能夠預測、計劃並在其訓練分佈之外行動。
在本文中,我們從因果的角度研究WM,涵蓋多個抽象層次,從感知觀察到構建支配環境動態的結構的概念表示。
我們認為,有用的WM必須超越僅僅生成的能力:它們還應該捕捉實體特性、實體之間的互動以及實體與環境之間的互動,這些互動決定並解釋系統的動態。
我們提供了一個基於其所支持任務的因果WM(CWM)的正式定義,將世界建模與現有的因果表示學習、以物體為中心的學習、因果發現、結構因果模型和基於模型的決策制定相連接。
最後,我們將CWM與可識別性文獻相關聯,澄清WM的組件何時可以從數據中恢復以及在何種等價下。
通過這樣,我們將WM建立在支持因果推理和知情決策的表示和結構上。
Explanatory Engagement Under Rare Anomalous Failure: Asymptotic Rarity in Model Behavior (or: The Asymptotic AI)
2608.13063v1 by Sam Mao
Prior work on LLM behavior under anomalous conditions asks whether a model notices anomalies. We ask a narrower question: once a model sits in a workflow with a low, controllable failure rate, does its explanatory engagement - length, specificity, self-reported confidence - change as failure grows asymptotically rarer? We built a local, zero-cost harness on three open-weight models (qwen3:8b, llama3.1:8b, mistral:7b) running a repeated tool-call task where one call fails at probability p, swept across eight rates from 0.2 to 0.0001, under five elicitation conditions from immediate prompting to none. We hypothesized a rise in engagement as failures grew rarer, then a collapse near a detectability threshold. Pooled across conditions this appeared false: length fell in a flat, monotonic pattern. Splitting by condition overturned that. Under immediate_forced, where the model must explain every failure instantly, the predicted rise is confirmed but followed by a plateau, not a collapse: length peaks at 28.4 words at p=0.05, settles to 17.4-19.0 words at the rarest rates, and confidence rises unevenly from about 53% to the 70s-90s. Under grouped_runs, explanation batched to run-end, no collapse appears. Under passive_unprompted, aggregate magnitude is a floor artifact, but a recovered logging gap revealed real, model-specific self-monitoring: llama3.1:8b volunteers structured confidence reports unprompted, sometimes eroding its own confidence as trials accumulate; the other two do so only once, as boilerplate. Elicitation structure is a first-class moderator of collapse observability. A companion guaranteed-failure run (72 cells, backfilling rates where random sampling gave zero real failures) shows models differ in whether they recognize an anomaly, distinct from engagement once recognized. Limitation: discrete rate points cannot capture behavior between them, a direction for future work.
摘要:先前對於 LLM 在異常條件下行為的研究探討模型是否能注意到異常。我們提出了一個更狹窄的問題:當模型在一個失敗率低且可控的工作流程中時,隨著失敗變得漸近稀有,其解釋參與度 - 長度、具體性、自我報告的信心 - 是否會改變?我們在三個開放權重模型(qwen3:8b、llama3.1:8b、mistral:7b)上建立了一個本地的零成本工具,運行一個重複的工具調用任務,其中一個調用以概率 p 失敗,並在五種引導條件下從 0.2 到 0.0001 的八個速率中進行掃描,從立即提示到無提示。我們假設隨著失敗變得更稀有,參與度會上升,然後在可檢測性閾值附近會崩潰。根據條件的匯總,這似乎是錯誤的:長度以平坦的單調模式下降。按條件劃分則推翻了這一點。在立即強制條件下,模型必須立即解釋每一次失敗,預測的上升得到了確認,但隨後出現平臺,而不是崩潰:在 p=0.05 時長度達到 28.4 字,在最稀有的速率下穩定在 17.4-19.0 字,信心從約 53% 不均勻上升到 70% 到 90% 之間。在分組運行條件下,解釋批量到運行結束,沒有出現崩潰。在被動無提示條件下,總體大小是一個底部工件,但恢復的日誌間隙揭示了真實的模型特定自我監控:llama3.1:8b 在未提示的情況下自願提供結構化的信心報告,有時隨著試驗的累積而侵蝕自己的信心;另外兩個模型僅在一次時作為模板進行。引導結構是崩潰可觀察性的第一級調節因子。一個伴隨的保證失敗運行(72 個單元,填補隨機抽樣導致零真實失敗的速率)顯示模型在是否識別異常方面存在差異,這與一旦識別後的參與度是不同的。限制:離散速率點無法捕捉之間的行為,這是未來工作的方向。
VALG: An Agentic System for ML Theory Research
2608.13060v1 by Dechen Zhang, Xuan Tang, Xinxiang Yin, Xingwu Chen, Jian Qian, Difan Zou
Machine learning theory studies learning procedures through mathematical setups in which the data model, training protocol, oracle access, loss, metric, and randomness define the phenomenon that a theorem is meant to explain. Solving an open problem therefore requires the problem formulation, theorem target, and proof mechanism to be developed in concert. Researchers formulate hypotheses, test them through preliminary theoretical or empirical analysis, and refine both assumptions and proofs. We investigate whether this process can be organized as an autonomous agentic workflow for ML theory research. We develop VALG, an agentic system that combines multi-level Verification, Adaptive formulation of Learning-theory problems, and Graph-structured proof development. Within each source-relative theorem branch, VALG maintains a fixed mathematical specification, checks the theorem-level composition of a typed proof-dependency graph, and constructs and reviews local proofs in dependency order. When a proof attempt fails, VALG identifies whether the obstruction lies in a derivation, the proof structure, or the theorem formulation and routes the next attempt accordingly. Formulation-level obstructions initiate an explicitly related variant or relaxation, preserving the mathematical relation between the resulting theorem and the source problem. We evaluate VALG on nine subproblems from five COLT 2026 open problems. Two runs produce internally finalized theorem candidates that match the scope of their source briefs; the remaining seven yield restricted-method results, special cases, or conditional theorems. These case studies show how VALG keeps source-scope matches, relaxations, conditional results, and blocked attempts mathematically distinct. VALG is open source at https://github.com/DechenZhang/VALG-ML-Theory-Agent.
摘要:機器學習理論通過數學設置研究學習程序,其中數據模型、訓練協議、預言機訪問、損失、度量和隨機性定義了定理旨在解釋的現象。因此,解決一個開放問題需要問題的表述、定理的目標和證明機制共同發展。研究人員提出假設,通過初步的理論或實證分析進行測試,並不斷完善假設和證明。我們調查這一過程是否可以組織成一個自主的代理工作流程,用於機器學習理論研究。
我們開發了 VALG,一個結合多層驗證、學習理論問題的自適應表述和圖結構證明開發的代理系統。在每個源相對的定理分支內,VALG 維持固定的數學規範,檢查類型證明依賴圖的定理級組合,並按照依賴順序構建和審查局部證明。當證明嘗試失敗時,VALG 確定障礙是否在推導、證明結構或定理表述中,並相應地路由下一次嘗試。表述級的障礙啟動一個明確相關的變體或放鬆,保持結果定理與源問題之間的數學關係。
我們在五個 COLT 2026 開放問題的九個子問題上評估 VALG。兩次運行產生了內部最終化的定理候選,與其源簡報的範圍相匹配;其餘七個產生了限制方法的結果、特例或條件定理。這些案例研究顯示了 VALG 如何保持源範圍匹配、放鬆、條件結果和被阻止的嘗試在數學上是不同的。VALG 是開源的,網址為 https://github.com/DechenZhang/VALG-ML-Theory-Agent。
UniTraffic-Agent: Unified Traffic Video Reasoning for AI City Challenge 2026 Track 3 with Two Out-of-Domain Evaluations
2608.13031v1 by Peng Li, Qianqian Xu, Shilong Bao, Yangbangyan Jiang, Qingming Huang
Traffic video understanding has become an important problem in intelligent transportation, as road videos provide direct evidence for accidents, violations, and interactions between vehicles and vulnerable road users. A useful system should explain how a traffic event develops, why it happens, and when the relevant interaction occurs, yet this remains difficult for multimodal large language models (MLLMs) because traffic videos contain sparse events and varied viewpoints. We introduce UniTraffic-Agent, the MR-CAS solution for Track~3 of the 10th AI City Challenge, which includes Traffic Anomaly Reasoning (TAR) and two out-of-domain evaluations: FETV for fisheye traffic events and PSI-VQA for pedestrian intention reasoning. UniTraffic-Agent follows an observe--reason--act--verify workflow that samples timestamped visual evidence, reasons over all questions from the same clip in one request, and converts responses through task-specific action adapters. On the official Public leaderboards, MR-CAS ranks 16th on TAR with a score of 0.5780, 2nd on FETV with 0.4884, and 4th on PSI-VQA with 64.4161. The code is available at https://github.com/Roclp/UniTraffic-Agent.
摘要:交通視頻理解已成為智能交通中的一個重要問題,因為道路視頻提供了事故、違規和車輛與脆弱道路使用者之間互動的直接證據。
一個有用的系統應該解釋交通事件是如何發展的,為什麼會發生,以及相關互動何時發生,然而這對於多模態大型語言模型(MLLMs)來說仍然困難,因為交通視頻包含稀疏事件和多樣的視角。
我們介紹了UniTraffic-Agent,這是第十屆AI城市挑戰賽Track~3的MR-CAS解決方案,其中包括交通異常推理(TAR)和兩個域外評估:FETV針對魚眼交通事件和PSI-VQA針對行人意圖推理。
UniTraffic-Agent遵循觀察--推理--行動--驗證的工作流程,從時間戳視覺證據中取樣,對同一片段的所有問題進行推理,並通過特定任務的行動適配器轉換響應。
在官方公共排行榜上,MR-CAS在TAR中排名第16,得分為0.5780,在FETV中排名第2,得分為0.4884,在PSI-VQA中排名第4,得分為64.4161。
代碼可在https://github.com/Roclp/UniTraffic-Agent獲得。
Static analysis-guided agentic AI translation enables Rust as a full stack bioinformatics language
2608.13029v1 by Johan Henriksson
The field of bioinformatics struggles with legacy code - old code that is commonly used but may no longer have a maintainer, or may be written in an now-unfamiliar language (e.g. Perl, Fortran). This incurs maintenance cost (technical debt), but dynamically typed languages also negatively impacts the environment and fail to make use of modern hardware. Legacy code may also have security or safety problems that make it unsuited for use in clinical settings. Here we show that agentic AI, combined with static analysis, can be used to translate legacy code to the modern language Rust. We provide prompts and supporting software to aid systematic translation, and evaluate it on common software for NGS and imaging. We showcase the result on our software Bascet: Size was reduced by ~80x, build time decreased by ~10x, and performance of key steps improved >3x. Unix dependencies were also removed, making Bascet the only single-cell pipeline able to run on native Windows, without a container. Large-scale refactoring of bioinformatics software is thus now possible at a limited budget, enabling more complex tools to be developed.
摘要:生物資訊學領域面臨著舊有代碼的挑戰——這些舊代碼通常被使用,但可能不再有維護者,或可能是用現在不熟悉的語言(例如 Perl、Fortran)編寫的。這會產生維護成本(技術負債),但動態類型語言也會對環境產生負面影響,並未能充分利用現代硬體。舊代碼可能還存在安全或安全性問題,使其不適合在臨床環境中使用。在這裡,我們展示了代理式 AI 結合靜態分析,可以用來將舊代碼轉換為現代語言 Rust。我們提供提示和支持軟體以協助系統性翻譯,並在 NGS 和成像的常見軟體上進行評估。我們展示了我們的軟體 Bascet 的結果:大小減少約 80 倍,建構時間減少約 10 倍,關鍵步驟的性能提高了超過 3 倍。Unix 依賴也被移除,使 Bascet 成為唯一能在本地 Windows 上運行的單細胞管道,而無需容器。因此,生物資訊學軟體的大規模重構現在在有限的預算下成為可能,從而使得更複雜的工具得以開發。
Decomposition of Evidence, Contradiction, and Fragility in Perturbation Responses
2608.12935v1 by Lei You
Perturbation methods explain model decisions by measuring prediction changes under altered inputs, but response magnitude tells us only how much a model reacts, not what that reaction means. The same magnitude can support the final factual-counterfactual difference, oppose it, or arise strongly along the perturbation path yet vanish at the endpoint. We therefore track how the contrast develops as paired inputs are progressively revealed, using the final contrast to interpret the trajectory. We introduce DECAF (Decomposition of Evidence, Contradiction, And Fragility), which routes aligned, opposed, and endpoint-null responses into evidence E, contradiction C, and fragility F. The decomposition preserves ordinary magnitude exactly, Abs = E + C + F, and is unique under endpoint-relative axioms. Across controlled vision and tabular settings, the three components track independently measured behavior. In a 72-model ImageNet-9 audit, we compare cases with nearly identical response magnitude but different independently measured behaviors. The largest DECAF component agrees with an observed behavior in 96.4% of cases, compared with 35.0% for magnitude alone. Changing only the reveal path increases total response by nearly 80%, yet evidence barely changes while fragility grows by more than 4x. On FunnyBirds and ImageNet-1k, short forward-only DECAF trajectories outperform the tested general-purpose attribution baselines. On a 1B-scale DINOv2 model, a short trajectory matches a strong gradient-based baseline with 4.75x lower wall time and 2.36x lower peak memory.
摘要:擾動方法通過測量在改變輸入下的預測變化來解釋模型決策,但反應的大小僅告訴我們模型反應的程度,而不告訴我們這種反應的意義。相同的大小可以支持最終的事實-反事實差異,反對它,或在擾動路徑上強烈出現但在終點消失。因此,我們追蹤對比如何隨著配對輸入的逐步揭示而發展,並利用最終對比來解釋軌跡。我們引入DECAF(證據、矛盾和脆弱性的分解),將對齊的、對立的和終點無效的反應路由到證據E、矛盾C和脆弱性F。這種分解準確地保留了普通的大小,Abs = E + C + F,並且在相對於終點的公理下是唯一的。在控制視覺和表格設置中,這三個組件追蹤獨立測量的行為。在一次72模型的ImageNet-9審計中,我們比較了反應大小幾乎相同但獨立測量行為不同的案例。最大的DECAF組件在96.4%的案例中與觀察到的行為一致,而僅僅依賴大小的情況下為35.0%。僅改變揭示路徑使總反應增加近80%,而證據幾乎不變,脆弱性增長超過4倍。在FunnyBirds和ImageNet-1k上,僅向前的短DECAF軌跡超越了測試的通用歸因基準。在一個規模為1B的DINOv2模型上,短軌跡與一個強大的基於梯度的基準相匹配,牆面時間低4.75倍,峰值內存低2.36倍。
Discovering Efficient and Explainable Communication Topologies for LLM-based Multi-Agent Systems via Causal Inference
2608.12921v2 by Junzhi Li, Peng He, Qirui Ji, Wei Wang, Lixiang Liu, Chuxiong Sun
The performance of large language model (LLM)-based multi-agent systems (MAS) largely depends on effective communication topologies. Existing topology generation methods, however, typically learn communication topologies through black-box optimization driven solely by task-level rewards. While effective, such optimization provides little insight into why particular communication edges are selected, making it difficult to identify the critical communication subgraphs responsible for successful collaboration. To address this limitation, we propose E2-Explainer, a model-agnostic framework for providing interpretable explanations of communication topologies produced by arbitrary topology generators. Specifically, we formulate topology explanation as a causal attribution problem that identifies compact communication subgraphs supported by edge-level evidence of task preservation. We obtain this evidence with a Granger-style objective that measures how masking each communication channel changes the task outcome and the stability of the final response. The resulting budgeted subgraphs are then distilled into an amortized explainer, enabling efficient post-hoc explanation without repeated edge-level evaluations at deployment. Extensive experiments on multiple reasoning and coding benchmarks demonstrate that E2-Explainer identifies critical communication subgraphs that preserve successful collaboration. These subgraphs can also be executed directly to prune redundant communication edges, substantially reducing communication costs while maintaining competitive task performance.
摘要:大型語言模型(LLM)為基礎的多代理系統(MAS)的性能在很大程度上依賴於有效的通信拓撲。
然而,現有的拓撲生成方法通常通過僅由任務級獎勵驅動的黑箱優化來學習通信拓撲。
雖然有效,但這種優化對於為何選擇特定通信邊緣提供了很少的洞察,這使得識別負責成功協作的關鍵通信子圖變得困難。
為了解決這一限制,我們提出了E2-Explainer,一個模型無關的框架,用於提供可解釋的通信拓撲解釋,這些拓撲由任意拓撲生成器產生。
具體而言,我們將拓撲解釋公式化為一個因果歸因問題,該問題識別由任務保持的邊級證據支持的緊湊通信子圖。
我們通過一個Granger風格的目標來獲得這些證據,該目標測量屏蔽每個通信通道如何改變任務結果和最終響應的穩定性。
隨後,得到的預算子圖被提煉成一個攤銷解釋器,從而在部署時能夠高效地進行事後解釋,而無需重複的邊級評估。
在多個推理和編碼基準上的廣泛實驗表明,E2-Explainer識別出保持成功協作的關鍵通信子圖。
這些子圖還可以直接執行,以修剪冗餘的通信邊緣,顯著降低通信成本,同時保持競爭性的任務性能。
Mr3D-VL: A generalist vision language foundation model for Multiparametric 3D Magnetic Resonance Imaging
2608.12689v1 by Zhi Qiao, Xintong Wu, Yichu He, Feng Shi
Multi-parametric magnetic resonance imaging (mpMRI) is a cornerstone for brain tumor diagnosis and treatment, yet current AI models face critical limitations: their lack of natural language interaction and interpretability impedes spatial information integration and cross-modal reasoning required clinically. Key challenges arise from significant physical meaning differences across modalities, spatial misalignment due to scan intervals, and the need for complex multi-feature interpretation in tasks like glioma grading. While visual-language models (VLMs) show promise in cross-modal understanding, existing methods focus mainly on 2D image modeling, neglecting direct perception of 3D volumetric space. Although 3D VLMs have been proposed for report generation and feature alignment in 3D CT imaging, mpMRI applications demand collaborative inference across multiple imaging modalities-a requirement unmet by current solutions. To address this, we introduce Mr3D-VL, a dedicated visual-language foundation model for multi-parametric 3D MRI. With 4 billion parameters, it employs an unsupervised pre-trained shared 3D encoder and 4D rotational positional embedding for dual modality-spatial integration. Its cross-modal projection layer uses a multi-resolution feature implantation strategy to enhance feature perception across resolutions. Experimental results show significant improvements over existing 4B/7B/30B domain-specific and general-purpose models in text generation tasks, achieving a BERTScore of 0.856 for report generation, with question-answering accuracy at 0.713 and multiple-choice accuracy at 0.912.
摘要:多參數磁共振成像(mpMRI)是腦腫瘤診斷和治療的基石,但目前的AI模型面臨重大限制:缺乏自然語言互動和可解釋性妨礙了臨床所需的空間信息整合和跨模態推理。主要挑戰來自於不同模態之間顯著的物理意義差異、由於掃描間隔造成的空間錯位,以及在如膠質瘤分級等任務中對複雜多特徵解釋的需求。儘管視覺語言模型(VLMs)在跨模態理解方面顯示出潛力,但現有方法主要集中在2D圖像建模,忽略了對3D體積空間的直接感知。雖然已提出3D VLMs用於報告生成和3D CT成像中的特徵對齊,但mpMRI應用需要跨多個成像模態的協作推理——這一需求目前的解決方案無法滿足。為了解決這個問題,我們推出了Mr3D-VL,一個專門針對多參數3D MRI的視覺語言基礎模型。它擁有40億個參數,採用無監督預訓練的共享3D編碼器和4D旋轉位置嵌入進行雙模態空間整合。其跨模態投影層使用多解析度特徵植入策略來增強不同解析度間的特徵感知。實驗結果顯示,在文本生成任務中,與現有的4B/7B/30B領域特定和通用模型相比,顯著提高了性能,報告生成的BERTScore達到0.856,問答準確率為0.713,多選準確率為0.912。
Interpretable Causal Discovery via Causal-Effect Constraints
2608.12640v1 by Cixuan Zhang, Guy Van den Broeck, Benjie Wang
Causal discovery aims to uncover the underlying causal relationships given data generated from a system. The goal, however, is not merely to predict causal edges given data, but also to be able to interpret and explain either observed or hypothesized phenomena, such as a particularly large causal effect. We consider this task of conditional causal discovery and cast it as a Bayesian inference problem, in which we target the posterior over causal graphs and parameters conditional on an event such as a causal-effect constraint. Unfortunately, this poses a computational challenge: existing approaches to Bayesian causal discovery struggle when the event has small posterior mass. To address this, we adapt rare-event estimation techniques to perform inference the joint graph-parameter space. Our method gradually drives a particle population toward the constrained region while maintaining samples that approximate the conditional posterior. Empirical evaluation on synthetic graphs validates the accuracy of our approach at small and large scales, and we show in a case study on the Sachs protein dataset how our method can be used to aid scientific exploration by providing pathway-level summaries.
摘要:因果發現旨在揭示基於系統生成數據的潛在因果關係。然後,目標不僅僅是根據數據預測因果邊緣,還要能夠解釋和說明觀察到的或假設的現象,例如特別大的因果效應。我們考慮這一條件因果發現的任務,並將其視為一個貝葉斯推斷問題,在這個問題中,我們針對因果圖和參數的後驗分佈,條件是某個事件,如因果效應約束。不幸的是,這帶來了計算挑戰:現有的貝葉斯因果發現方法在事件具有小後驗質量時表現不佳。為了解決這個問題,我們調整了稀有事件估計技術,以在聯合圖-參數空間中進行推斷。我們的方法逐漸將粒子群體推向受約束的區域,同時保持近似條件後驗的樣本。在合成圖上的實證評估驗證了我們的方法在小規模和大規模上的準確性,我們在Sachs蛋白數據集的案例研究中展示了我們的方法如何通過提供通路級摘要來幫助科學探索。
Algorithm Design and Physician Liability
2608.13618v1 by Shujie Luan, Shubhranshu Singh, Tinglong Dai
A single clinical algorithm can deliver unequal accuracy across patient groups, and concern about such disparity has grown as artificial intelligence (AI) spreads through clinical decision-making. In response, a liability rule introduced in the United States holds healthcare providers responsible when their reliance on disparate algorithms contributes to erroneous clinical decisions. We examine how such liability considerations reshape (i) an AI firm's algorithm design decisions that drive group-specific accuracy and (ii) a physician's decisions to use AI in healthcare delivery. The AI firm designs an algorithm for two patient groups, and improving accuracy for the disadvantaged group is more costly. The physician (who remains the accountable decision-maker) then decides whether to consult AI, weighing the reduction in clinical uncertainty against expected liability exposure when AI errors disproportionately affect the disadvantaged group. We find the liability rule can induce disparate use of AI: the physician may reduce AI use overall and, over an intermediate range of liability, rely on AI less for disadvantaged patients. The effect is non-monotone. As liability increases, the physician's use of AI for disadvantaged patients first declines, then rises as the firm reallocates investment toward reducing disparity or switches to an equal-accuracy design. Mandating equal algorithmic accuracy across patient groups can then inadvertently harm both groups, because a uniform accuracy requirement distorts the firm's investment incentives and the physician's equilibrium AI-use decisions.
摘要:單一的臨床演算法在不同患者群體中可能會產生不均等的準確性,隨著人工智慧(AI)在臨床決策中的普及,對於這種差異的關注也日益增加。作為回應,美國引入了一項責任規則,當醫療提供者依賴不同的演算法導致錯誤的臨床決策時,將其負責。我們研究這種責任考量如何重塑(i)AI公司的演算法設計決策,促進特定群體的準確性,以及(ii)醫生在醫療提供中使用AI的決策。AI公司為兩個患者群體設計了一個演算法,改善弱勢群體的準確性成本更高。然後,醫生(仍然是負責的決策者)決定是否諮詢AI,權衡臨床不確定性的減少與當AI錯誤不成比例地影響弱勢群體時的預期責任風險。我們發現責任規則可能會導致AI的使用不均等:醫生可能會整體減少AI的使用,並且在責任的中等範圍內,對弱勢患者的AI依賴程度降低。這一效果是非單調的。隨著責任的增加,醫生對弱勢患者使用AI的情況最初下降,然後隨著公司將投資重新分配到減少差異或轉向平等準確性設計而上升。要求在患者群體之間達到平等的演算法準確性,可能會無意中對兩個群體造成傷害,因為統一的準確性要求扭曲了公司的投資激勵和醫生的均衡AI使用決策。
What Makes a Peer? Valuation-Anchored Similarity in Private Markets
2608.12594v1 by Sebastian Frank, Jingrao Lyu, Max Jarmey, Preetha Saha, Mingshu Li, Sweet Kaur, Sola Akinola, Dhagash Mehta
As more investors contemplate private markets and contend with limited transparency, sparse disclosures, and infrequent transactions, identifying economically meaningful peer companies for comparison is a fundamental challenge for valuation, due diligence, portfolio construction, and risk management. We propose an ensemble tree-based supervised similarity learning framework that defines company similarity through the lens of market valuation rather than static feature matching or semantic descriptions. Specifically, we train a CatBoost gradient-boosted decision tree model on observed private company valuations and derive a valuation-aware similarity metric from importance-weighted leaf-node co-occurrences across the ensemble. The similarity metric captures shared valuation drivers while accommodating nonlinear relationships, mixed data types, and pervasive missing data common in private markets. Using a global private-market universe of approximately 270,000 companies, including more than 53,000 firms with observed or derivable post-money valuations spanning multiple industries, geographies, and deal stages, we demonstrate that the proposed similarity framework improves upon traditional distance-based and text-embedding-based approaches in downstream k-nearest-neighbor valuation tasks in the evaluated industry groups, while retaining case-based explainability.
摘要:隨著越來越多的投資者考慮私募市場並面對有限的透明度、稀疏的披露和不頻繁的交易,識別具有經濟意義的同行公司以進行比較對於估值、盡職調查、投資組合構建和風險管理來說是一個基本挑戰。
我們提出了一種基於集成樹的監督相似性學習框架,通過市場估值的視角來定義公司相似性,而不是靜態特徵匹配或語義描述。
具體而言,我們在觀察到的私募公司估值上訓練了一個CatBoost梯度提升決策樹模型,並從集成中的重要性加權葉節點共現中推導出一個考慮估值的相似性度量。
該相似性度量捕捉了共享的估值驅動因素,同時適應了非線性關係、混合數據類型以及在私募市場中普遍存在的缺失數據。
使用約27萬家公司的全球私募市場範圍,包括超過53,000家具有觀察或可推導的後資金估值的公司,涵蓋多個行業、地理區域和交易階段,我們證明所提出的相似性框架在評估的行業組中改善了傳統的基於距離和基於文本嵌入的方法在下游k最近鄰估值任務中的表現,同時保留了基於案例的可解釋性。
Auditable agentic AI for evidence-grounded thyroid ultrasound diagnosis and reporting
2608.12590v1 by Haifan Gong, Shiyu Chen, Bodong Wang, Yuqi Wang, Shijie Wang, Guoliang You, Xinyu Xiong, Haowei Wang, Mingzhi Mao, Dexing Kong, Qinghua Liu, Wei Lou, Fei Chen, Guanbin Li
Thyroid ultrasound diagnosis requires coordinated lesion localization, measurement, risk stratification and reporting, yet most AI systems address these tasks in isolation and provide limited support for clinical review. We present ThyroidXAgent, a clinician-interactive agentic AI system that coordinates specialized diagnostic tools and stores their outputs as an auditable case-level evidence record. The system was developed using OpenThyroidDB, a multicentre, multitask resource integrating approximately 0.3 million ultrasound images and 24,000 paired reports, and was evaluated on 28,458 non-overlapping test cases, including 8,721 cases from 35 centres in the private NHC-MISD-TUS cohort. Across heterogeneous datasets, ThyroidXAgent achieved a mean Dice score of 87.21 percent for nodule segmentation and a mean AUROC of 0.9466 for benign-malignant classification. The same workflow supported lymph-node metastasis prediction and follicular versus papillary thyroid carcinoma classification, with AUROCs of 0.864 and 0.805, respectively. For report generation, evidence-grounded assembly outperformed multimodal language-model baselines across three cohorts. ThyClinScore, a lesion-level clinical semantic metric introduced here, showed the strongest correlation with a location-aware language-model judge. ThyroidXAgent improved physician classification accuracy, increased report diagnostic consistency from 70.3 percent to 86.2 percent, and reduced segmentation and reporting time by 35.9 percent and 27.4 percent, respectively. These findings support auditable, clinician-correctable agentic AI for thyroid ultrasound diagnosis and reporting.
摘要:甲狀腺超聲診斷需要協調病變定位、測量、風險分層和報告,但大多數人工智慧系統在孤立的情況下處理這些任務,並對臨床審查提供有限的支持。我們提出了ThyroidXAgent,一個臨床互動的代理人工智慧系統,協調專門的診斷工具並將其輸出存儲為可審計的案例級證據記錄。該系統是使用OpenThyroidDB開發的,這是一個多中心、多任務的資源,整合了約30萬張超聲圖像和24,000份配對報告,並在28,458個不重疊的測試案例上進行了評估,包括來自私立NHC-MISD-TUS隊列的35個中心的8,721個案例。在異質數據集上,ThyroidXAgent在結節分割方面達到了87.21%的平均Dice分數,並在良惡性分類方面達到了0.9466的平均AUROC。同一工作流程支持淋巴結轉移預測和濾泡型與乳頭狀甲狀腺癌的分類,AUROC分別為0.864和0.805。在報告生成方面,基於證據的組合在三個隊列中超越了多模態語言模型基準。這裡引入的ThyClinScore,一個病變級的臨床語義指標,顯示出與位置感知語言模型評審者之間的最強相關性。ThyroidXAgent提高了醫生的分類準確性,將報告的診斷一致性從70.3%提高到86.2%,並分別減少了35.9%和27.4%的分割和報告時間。這些發現支持可審計、可由臨床醫生修正的代理人工智慧用於甲狀腺超聲診斷和報告。
CAS: A Causal Attribution Score for Local and Global Explainable Artificial Intelligence
2608.12555v1 by Michael Georgiades, Charalambia Varnava
Predictive explanation methods attribute a model output; they do not, by themselves, attribute an intervention effect on the real-world outcome. We introduce the Causal Attribution Score (CAS), a compact score architecture for causal explanation. CAS starts from an identified interventional coalition game, allocates the joint intervention contrast with causal Shapley contributions, and converts those raw outcome-scale effects into Local CAS, Signed Local CAS, and two complementary Global CAS summaries. The innovation is not a new Shapley formula, but a local-to-global causal reporting layer with an explicit intervention target. In the known-truth benchmark, eight repeated primary-interaction simulations (n = 2,200 each, three actions) gave mean Local CAS MAE of 0.107 for coalition-aware CAS, compared with 0.173 for one-at-a-time normalisation and 0.213 for a global normalised absolute ATE vector. The paired advantage over one-at-a-time normalisation increased from -0.003 under additivity to 0.091 under strong interactions. On both empirical DoubleML datasets, 401(k) eligibility/net financial assets (n = 9,915) and Pennsylvania reemployment bonus/unemployment duration (n = 5,099), predictive SHAP/TreeSHAP rankings differed materially from Feature-CAS rankings of treatment-effect modifiers. In Pennsylvania, dep1 (exactly one dependent) moved from predictive global rank 13 to Feature-CAS rank 2 and was the leading local Feature-CAS modifier. These results isolate the added value of separating what predicts the outcome from what explains heterogeneity in an estimated causal effect.
摘要:預測解釋方法歸因於模型輸出;它們本身並不歸因於對現實世界結果的干預效果。我們介紹了因果歸因分數(Causal Attribution Score, CAS),這是一種用於因果解釋的緊湊分數架構。CAS 以識別的干預聯盟遊戲為起點,根據因果 Shapley 貢獻分配聯合干預對比,並將這些原始結果尺度效果轉換為局部 CAS、簽名局部 CAS 和兩個互補的全球 CAS 摘要。這一創新不是一個新的 Shapley 公式,而是一個具有明確干預目標的局部到全球因果報告層。在已知真相的基準測試中,八次重複的主要互動模擬(每次 n = 2,200,三個行動)對於考慮聯盟的 CAS 給出了平均局部 CAS MAE 為 0.107,而一次性標準化為 0.173,全球標準化的絕對 ATE 向量為 0.213。相較於一次性標準化,基於可加性的配對優勢從 -0.003 增加到強互動下的 0.091。在兩個實證 DoubleML 數據集上,401(k) 合格性/淨財務資產(n = 9,915)和賓夕法尼亞州再就業獎金/失業持續時間(n = 5,099),預測 SHAP/TreeSHAP 排名與治療效果修飾因子的 Feature-CAS 排名有顯著差異。在賓夕法尼亞州,dep1(恰好一名受撫養人)從預測全球排名第 13 移動到 Feature-CAS 排名第 2,並成為主要的局部 Feature-CAS 修飾因子。這些結果隔離了將預測結果與解釋估計因果效果異質性的內容分開的附加價值。
Class Activation Mapping in Explainable Computer Vision: A Method-Centered Review of CNN, Transformer, and Foundation-Model-Era Visual Explanations
2608.12299v2 by AmirHossein Eshghi, Hamid Saadatfar, Seyyed Ali Hoseini, AmirMohsen Eshghi, Siavash Arjomand Bigdeli
Class activation mapping (CAM) is one of the most widely used visual explanation families in explainable artificial intelligence. Its purpose is intuitive: it converts internal model evidence into a heatmap that highlights the image regions, convolutional channels, tokens, or patches that support a target class or concept. Since the first CAM formulation in 2016, the field has moved far beyond global-average-pooled CNN classifiers. CAM-style methods now include gradient-based post-hoc explanations, gradient-free score and ablation methods, high-resolution upscaling, weakly supervised localization and segmentation, transformer token attribution, causal and debiasing methods, and foundation-model-era approaches that use CLIP, DINO, SAM, or feature-distribution comparisons. This review synthesizes a strict corpus of 57 method-centered papers published from 2016 onward. The paper develops a taxonomy that separates methods by attribution mechanism, architectural dependence, and evaluation objective. It then reviews gradient-based CAMs, recent and hybrid CAM-style methods, and model-based or architecture-aware methods. Across the corpus, the main trend is clear: the field is shifting from explaining one class score in one low-resolution CNN layer toward comparative, multi-layer, probabilistic, token-aware, and foundation-model-aware explanations. At the same time, evaluation remains fragmented. Faithfulness, localization, robustness, computational cost, and human trust are often measured with different protocols. The review therefore emphasizes not only what each method contributes, but also which gap it leaves open and which later methods attempt to close that gap.
摘要:類別激活映射(CAM)是可解釋人工智慧中最廣泛使用的視覺解釋方法之一。其目的直觀明瞭:它將內部模型證據轉換為熱圖,突顯支持目標類別或概念的圖像區域、卷積通道、標記或補丁。自2016年首次提出CAM公式以來,該領域已經遠遠超越了全局平均池化的CNN分類器。CAM風格的方法現在包括基於梯度的事後解釋、無梯度的分數和消融方法、高解析度的放大、弱監督的定位和分割、Transformer標記歸因、因果和去偏見方法,以及使用CLIP、DINO、SAM或特徵分佈比較的基礎模型時代方法。本綜述綜合了自2016年以來發表的57篇以方法為中心的論文,形成了一個嚴格的語料庫。該論文發展了一個分類法,根據歸因機制、架構依賴性和評估目標來區分方法。然後,它回顧了基於梯度的CAM、最近的混合CAM風格方法,以及基於模型或架構感知的方法。在這個語料庫中,主要趨勢顯而易見:該領域正從解釋一個低解析度CNN層中的一個類別分數,轉向比較的、多層的、概率的、標記感知的和基礎模型感知的解釋。與此同時,評估仍然是碎片化的。忠實性、定位、穩健性、計算成本和人類信任通常使用不同的協議進行測量。因此,該綜述不僅強調每種方法的貢獻,還指出它留下的空白,以及後來的方法試圖填補該空白。
Dual Spatial-Temporal Attribution: Architecture-Aligned Post-Hoc Explainability for Recurrent Graph Anomaly Detection
2608.12441v1 by Iyad Assaad Nekka, Hamida Seba, Khaled Walid Hidouci, Karima Amrouche
Deep learning detectors for anomalies in dynamic graphs have reached strong accuracy, yet they remain opaque: when an edge is flagged, the analyst receives a score but no reason. This opacity is untenable in the cooperative, regulated information systems where such detectors are deployed, where automated decisions must be auditable and trustworthy. We address this gap for AddGraph, the foundational GCN+GRU framework for edge-level anomaly detection in dynamic graphs, which to our knowledge has never been equipped with any form of explainability. We present a strictly post-hoc explainability framework, X-AddGraph, built on a Dual Spatial-Temporal Attribution (DSTA) mechanism whose three components are each aligned with one of AddGraph's architectural modules: a gradient-based relevance attribution over the current adjacency structure (spatial), a direct reading of the contextual attention weights already computed during inference (short-term temporal, at zero additional cost), and a gradient rollback through the recurrent hidden states (long-term temporal). Because the detector is frozen, detection performance is preserved exactly (Delta AUC = 0, verified empirically to ten decimal places). On the UCI Message benchmark, our trained AddGraph baseline reaches an average per-snapshot AUC of 0.8705, exceeding the originally published result; X-AddGraph reproduces every score identically while adding explanations where none existed. Evaluated across four edge populations - confident true positives, low-confidence true positives, false positives, and random samples - the long-term attribution identifies historical snapshots carrying significantly more counterfactual signal than random selection (0.127 vs. 0.074), a capability that no spatially-blind explainer can provide. We release our implementation for full reproducibility.
摘要:深度學習檢測器在動態圖中的異常檢測已達到強大的準確性,但它們仍然不透明:當一條邊被標記時,分析師收到一個分數但沒有理由。這種不透明性在這些檢測器被部署的合作性、受規範的信息系統中是無法接受的,因為自動決策必須是可審計和可信的。我們針對AddGraph這一基礎的GCN+GRU框架進行了這一改進,該框架用於動態圖中的邊級異常檢測,據我們所知,它從未配備過任何形式的可解釋性。我們提出了一個嚴格的事後可解釋性框架X-AddGraph,該框架基於雙空間-時間歸因(DSTA)機制,其三個組件與AddGraph的架構模塊相對應:對當前鄰接結構的基於梯度的相關性歸因(空間),對推理過程中已計算的上下文注意權重的直接讀取(短期時間,無額外成本),以及通過循環隱藏狀態的梯度回滾(長期時間)。由於檢測器是固定的,因此檢測性能完全保留(Delta AUC = 0,經實證驗證至十位小數)。在UCI Message基準測試中,我們訓練的AddGraph基準達到每個快照平均AUC 0.8705,超過了最初發表的結果;X-AddGraph在添加解釋的同時,準確重現了每個分數。通過四個邊群體進行評估 - 自信的真陽性、低信心的真陽性、假陽性和隨機樣本 - 長期歸因識別出比隨機選擇顯著更多的反事實信號的歷史快照(0.127對0.074),這是任何空間盲目解釋器無法提供的能力。我們釋出我們的實現以實現完全可重現性。
Learning-Based Behavior Planning for Automated Driving: Real-World Integration and Deployment
2608.12198v1 by Jean-Pierre Busch, Guido Linden, Jan Bergmann, Lutz Eckstein
Recent research in machine and deep learning has shown the potential of learningbased motion planning approaches to improve the driving behavior of automated vehicles, especially in complex environments. However, their complex nature and lack of transparency can hinder explainability and trustworthiness and complicate safety assurance. Motivated by these challenges, we propose a hybrid planning architecture that combines the advantages of machine learning with the verifiability and the determinism of classical approaches. Specifically, we developed a deep neural network to interpret complex traffic scenes and propose driving behavior, while an optimization-based supervision layer validates this proposal and enforces explicit drivability and safety constraints. We evaluate the learned planner's driving behavior in open-loop studies on real-world urban data, discuss system integration aspects for stable closed-loop operation, and report results from real-world deployment on our research vehicle karl..
摘要:最近在機器學習和深度學習領域的研究顯示,基於學習的運動規劃方法有潛力改善自動駕駛車輛的駕駛行為,特別是在複雜環境中。然而,它們的複雜性和缺乏透明度可能會妨礙可解釋性和可信度,並使安全保證變得複雜。受到這些挑戰的啟發,我們提出了一種混合規劃架構,結合了機器學習的優勢以及經典方法的可驗證性和確定性。具體而言,我們開發了一個深度神經網絡來解釋複雜的交通場景並提出駕駛行為,同時基於優化的監督層驗證這一提議並強制執行明確的可駕駛性和安全約束。我們在現實世界的城市數據上進行開環研究,評估學習到的規劃者的駕駛行為,討論穩定閉環運行的系統集成方面,並報告我們的研究車輛karl的實際部署結果。
A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench
2608.12138v1 by Praveen Reddy, Charuta Mandke, Suvrankar Datta, Sarah Khan, Siddharth Reddy Anthireddy, Shitij Arora, Vishal Singh
General-purpose large language models (LLMs) have recently been reported to match or exceed specialized clinical AI tools on medical benchmarks, but such comparisons draw on a narrow set of systems and on benchmarks developed largely in high-income settings. We evaluate VITA, a retrieval-augmented generation (RAG) system purpose-built for contextual knowledge retrieval in India and other low- and middle-income (LMIC) settings. VITA retrieves from a curated corpus of disease-specific guidelines, India-specific antimicrobial resistance data, national formulary constraints, and resource-limited care protocols; its architecture and corpus are proprietary, but the benchmark, the physician-written rubrics, and our full response and scoring outputs are public for independent verification. On 4,023 English-language HealthBench questions (80.5% of the benchmark), scored with a GPT-4.1 judge, VITA ranked first with 51.9% of possible rubric points, ahead of GPT-5.4 (46.1%), o4-mini (44.3%), Gemini 3.1 Pro (42.6%), and Claude Sonnet 4.6 (37.3%), and scored highest on 45.4% of questions. To test robustness to newer models and judge lineage, a 500-question subset was re-run against current-generation models (GPT-5.5, Claude Opus 4.8, Gemini 3.5 Pro, Grok 4.3) and graded by a neutral open-weight judge (DeepSeek-V4-Pro) sharing no lineage with any system tested. Here the gap narrowed to parity: VITA and GPT-5.5 were statistically indistinguishable on mean per-question score, while VITA led on points-weighted score and won the most questions. VITA's advantages in accuracy and completeness persisted under the neutral judge; its communication scores were lower. These results indicate that a purpose-built clinical RAG system remains competitive with frontier LLMs on an open benchmark, consistent with corpus specificity as a design variable that improves grounding at some cost to communication polish.
摘要:一般用途的大型語言模型(LLMs)最近被報導在醫療基準上與專門的臨床人工智慧工具相匹配或超越,但這些比較依賴於一組狹窄的系統以及主要在高收入環境中開發的基準。我們評估了VITA,一個專為印度及其他低收入和中等收入(LMIC)環境中的上下文知識檢索而設計的檢索增強生成(RAG)系統。VITA從一個策劃的特定疾病指導方針、印度特定的抗微生物抗藥性數據、國家藥典限制以及資源有限的護理協議中檢索資料;其架構和語料庫是專有的,但基準、醫生撰寫的評分標準以及我們的完整回應和評分輸出是公開的,以便獨立驗證。在4,023個英語HealthBench問題(基準的80.5%)上,使用GPT-4.1評判,VITA以51.9%的可能評分點排名第一,超過了GPT-5.4(46.1%)、o4-mini(44.3%)、Gemini 3.1 Pro(42.6%)和Claude Sonnet 4.6(37.3%),並在45.4%的問題上得分最高。為了測試對新模型的穩健性和評判系譜,對500個問題的子集再次運行,與當前一代模型(GPT-5.5、Claude Opus 4.8、Gemini 3.5 Pro、Grok 4.3)進行比較,並由一位中立的開放權重評判(DeepSeek-V4-Pro)進行評分,該評判與任何測試系統無關。此時差距縮小至平行:VITA和GPT-5.5在每題平均得分上統計上無法區分,而VITA在加權得分上領先並贏得了最多問題。VITA在準確性和完整性上的優勢在中立評判下持續存在;其溝通得分較低。這些結果表明,專為臨床設計的RAG系統在公開基準上仍然與前沿LLMs具有競爭力,這與語料庫的特異性作為設計變量相一致,該變量在某種程度上提高了基礎性,但對溝通的精緻性造成了成本。
Do LLMs Take Care of Their Own? Similarity Signals Can Induce Cooperation
2608.12125v1 by Akash Kundu, Emanuel Tewolde, Ratip Emin Berker, Samuel F. Brown, Vincent Conitzer
As LLM-based agents with user-instructed goals are becoming widely deployed, they increasingly encounter each other in strategic interactions, and face challenges of finding mutually beneficial outcomes. Prior literature has argued that cooperation problems such as the Prisoner's Dilemma are resolvable in settings where agents know they follow very similar decision making patterns, as for example in monocultural AI ecosystems. Following that line of work, this paper introduces the first framework for evaluating LLM decision making when agents are provided with graded similarity signals. Among our findings, we establish that different LLM models vary drastically in how they navigate similarity signals, with some modern models showing consistent behavior across cooperation problems, payoff structures, and prompt framing. Perhaps surprisingly, our experiments also show that the dataset based on which the similarity signal is computed has small to no impact on induced cooperation, and that LLM models systematically self-identify as highly similar when asked to evaluate another model's chain-of-thought reasoning by themselves. Finally, we develop an LLM-behavioral-game-theoretic model that captures some of their reasoning rationale, and show that it can support cooperative outcomes in equilibrium under sufficiently high similarity scores.
摘要:隨著基於大型語言模型(LLM)的代理人以用戶指導的目標被廣泛部署,它們在戰略互動中越來越多地相遇,並面臨尋找互利結果的挑戰。先前的文獻已經論證,合作問題如囚徒困境在代理人知道它們遵循非常相似的決策模式的情況下是可以解決的,例如在單一文化的人工智慧生態系統中。沿著這一研究方向,本文介紹了第一個評估LLM決策制定的框架,當代理人被提供分級相似性信號時。
在我們的發現中,我們確立了不同的LLM模型在如何導航相似性信號方面存在巨大差異,一些現代模型在合作問題、收益結構和提示框架中顯示出一致的行為。或許令人驚訝的是,我們的實驗還顯示,計算相似性信號的數據集對於誘發合作的影響微乎其微,且當被要求自行評估另一模型的思考鏈推理時,LLM模型系統性地自我識別為高度相似。最後,我們開發了一個LLM行為博弈論模型,捕捉它們的一些推理理由,並顯示它可以在足夠高的相似性分數下支持均衡的合作結果。
Faithful, Sufficient and Understandable: Rethinking Graph Counterfactual Explanations via Discrete Diffusion Inversion
2608.12083v1 by David Bechtoldt, Sidney Bender
Graph Neural Networks (GNNs) achieve strong predictive performance on graph-structured data across domains such as chemistry, biology, and network analysis, yet they provide no intrinsic explanation of their predictions. This limits their adoption in high-stakes and safety-critical settings. Counterfactual explanations address this by revealing the minimal structural modifications that would change a model's prediction. On graphs, however, such a modification is hard to produce. The search space is discrete and combinatorial, and a valid answer must respect categorical node and edge types together with domain rules such as chemical valency in the case of molecular graphs. Existing explainers give up one of two things. Either edits are not held on the data manifold, or the search does not span the full edit space. We propose Graph Diffusion Counterfactual Explanation via Inversion (GDCE-I), which gives up neither. A discrete denoising diffusion model with a novel discrete inversion scheme enables distribution-aware edits leveraging the whole domain edit space. We further address the incomplete and inconsistent evaluation of graph counterfactuals by deriving a framework of explanation desiderata and applying it to every method under one shared protocol. Across four benchmarks, GDCE-I outperforms related work by a large margin on the defined framework. For the molecular domain, we further qualitatively show that GDCE-I attains interpretable in-distribution solutions.
摘要:圖神經網絡(GNNs)在化學、生物學和網絡分析等領域的圖結構數據上實現了強大的預測性能,但它們對其預測並未提供內在解釋。這限制了它們在高風險和安全關鍵環境中的應用。反事實解釋通過揭示最小結構修改來改變模型的預測來解決這一問題。然而,在圖上,這樣的修改難以產生。搜索空間是離散且組合性的,有效答案必須遵守類別節點和邊緣類型以及領域規則,例如在分子圖中的化學價。現有的解釋器放棄了兩者之一。要麼編輯不保持在數據流形上,要麼搜索不涵蓋完整的編輯空間。我們提出了通過反演的圖擴散反事實解釋(GDCE-I),它兩者都不放棄。一種具有新穎離散反演方案的離散去噪擴散模型使得能夠利用整個領域編輯空間進行分佈感知的編輯。我們進一步通過推導解釋需求框架並將其應用於每種方法下的一個共享協議,來解決圖反事實的評估不完整和不一致問題。在四個基準測試中,GDCE-I在定義的框架上大幅超越了相關工作。對於分子領域,我們進一步質性展示GDCE-I獲得了可解釋的內部分佈解決方案。
Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence
2608.12036v1 by Mengru Wang, Junfeng Fang, Shuofei Qiao, Zhenqian Xu, Haoming Xu, Haoxiong Wang, Shumin Deng, Linyi Yang, Zhixiang Cui, Xin Xu, Yunzhi Yao, Buqiang Xu, Fei Shen, Haozhe Luo, Yunxiang Wei, Ningyu Zhang, Julian McAuley, Tat Seng Chua, Huajun Chen
AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities and the risks they may pose remain poorly understood. As AI development becomes faster and increasingly automated, mechanistic exploration remains largely manual, widening the gap between what models can do and our ability to understand and control them. To bridge this gap, we introduce Mechanist, an agentic system that uses AI as a scientific instrument for the autonomous discovery of mechanisms underlying AI intelligence. To support autonomous mechanistic discovery, we construct an interpretability-focused knowledge graph of approximately 13,000 papers and integrate it with a multidisciplinary database of 43 million papers spanning 26 fields. We further curate a library of 32 foundational methods for mechanism analysis, causal intervention, and validation. Compared with Claude Code and existing AI-scientist systems, Mechanist generates more valuable mechanism hypotheses and executes experiments more reliably. Mechanist also demonstrates a progression from discovering model behaviors to explaining and controlling AI models. Specifically, Mechanist first uncovers a counterintuitive safety risk in scientific laboratories, showing that unsafe traits can transfer across modalities through apparently safe training data. Mechanist then develops a mechanism theory of belief, revealing how models represent world knowledge, form beliefs, infer the beliefs of others, and how these mechanisms emerge during pretraining. Finally, Mechanist translates these mechanistic insights into practical interventions that improve model performance across diverse scenarios and steer scientific foundation models toward generating DNA sequences with specified properties.
摘要:AI 模型在各個領域取得了顯著的成功,但其能力背後的機制以及可能帶來的風險仍然不甚了解。隨著 AI 發展變得越來越快速且自動化,機制探索仍然主要依賴人工,這擴大了模型能做的事情與我們理解和控制它們的能力之間的差距。為了縮小這一差距,我們介紹了 Mechanist,一個將 AI 作為科學工具,用於自主發現 AI 智力背後機制的代理系統。為了支持自主的機制發現,我們構建了一個專注於可解釋性的知識圖譜,涵蓋約 13,000 篇論文,並將其與一個跨越 26 個領域的 4,300 萬篇論文的多學科數據庫整合。 我們還策劃了一個包含 32 種基礎方法的庫,用於機制分析、因果干預和驗證。與 Claude Code 和現有的 AI 科學家系統相比,Mechanist 生成了更有價值的機制假設,並更可靠地執行實驗。Mechanist 還展示了從發現模型行為到解釋和控制 AI 模型的進展。具體而言,Mechanist 首先揭示了科學實驗室中的一個反直覺安全風險,顯示不安全特徵可以通過表面安全的訓練數據在不同模態之間轉移。然後,Mechanist 發展了一個信念的機制理論,揭示模型如何表現世界知識、形成信念、推斷他人的信念,以及這些機制如何在預訓練期間出現。最後,Mechanist 將這些機制見解轉化為實際干預措施,改善模型在各種場景中的表現,並引導科學基礎模型生成具有特定屬性的 DNA 序列。
From Safety Documentation to Safety Knowledge Support: An Evidence-Grounded LLM Framework for Medical Devices
2608.12025v1 by Tuhinangshu Gangopadhyay, Rasmus Adler, Peter Liggesmeyer, Jan Reich
Medical devices are becoming more software-intensive, connected, and AI-enabled. Their development requires risk-management evidence aligned with ISO 14971 and, for software, IEC 62304. This evidence must be kept consistent across requirements, design decisions, software changes, verification results, complaints, and post-market data. These tasks are costly and depend on scarce safety and domain experts. Large language models (LLMs) may reduce parts of this effort because medical-device safety work is highly document-based. However, current LLM-based safety-engineering studies often address isolated methods, rely on generic prompting or public examples, and provide limited support for source links, traceability, uncertainty handling, lifecycle updates, and recorded expert review. This limits their use in regulated medical-device development. This paper argues that the central research problem is not safety-text generation, but source-linked safety-knowledge support. We propose an evidence-grounded framework that connects device artifacts, controlled knowledge storage and retrieval, method-specific generation of candidate safety items, critique and uncertainty checks, and recorded expert review. The framework prepares, links, checks, and updates candidate safety artifacts for expert decision-making. It does not decide whether a device is safe and does not provide regulatory approval. We also outline an evaluation strategy using non-public or newly built medical-device case studies and expert reference analyses to assess coverage, correctness, relevance, traceability, duplicate rate, unsupported claims, and review effort.
摘要:醫療器材正變得越來越依賴軟體、互聯網連接和人工智慧。它們的開發需要符合ISO 14971的風險管理證據,對於軟體則需要符合IEC 62304。這些證據必須在需求、設計決策、軟體變更、驗證結果、投訴和市場後數據之間保持一致。這些任務成本高昂,並依賴於稀缺的安全和領域專家。
大型語言模型(LLMs)可能會減少這部分工作,因為醫療器材的安全工作高度依賴文檔。然而,目前基於LLM的安全工程研究往往針對孤立的方法,依賴於通用提示或公共範例,並對來源鏈接、可追溯性、不確定性處理、生命周期更新和記錄的專家審查提供有限支持。這限制了它們在受監管的醫療器材開發中的應用。
本文主張,核心研究問題不是安全文本生成,而是來源鏈接的安全知識支持。我們提出了一個基於證據的框架,連接設備文檔、受控知識存儲和檢索、特定方法生成候選安全項目、批評和不確定性檢查,以及記錄的專家審查。該框架為專家決策準備、鏈接、檢查和更新候選安全文檔。它不決定設備是否安全,也不提供監管批准。我們還概述了一個評估策略,使用非公開或新建的醫療器材案例研究和專家參考分析來評估覆蓋範圍、正確性、相關性、可追溯性、重複率、不支持的聲明和審查工作量。
Agent Skills Can Be Harmful: An Empirical Study of Skill-Induced Failures in LLM Agents
2608.11888v1 by Gen Dong, Yanjie Gao, Liqun Li, Tianyin Xu, Yu Hua, Fan Yang
Agent skills are the de facto mechanism for extending LLM agents with reusable guidance. A skill can shape the agent's task execution, including planning, tool use, problem-solving, and validation. Prior work reported mixed results of agent skills: some skills improve task success rates, while others have no effect, increase token use and execution time, and even reduce success rates. This paper presents a comprehensive analysis of skill-induced agent failures by attributing task failures and cost regressions to specific loaded skills. We introduce a differential analysis framework that attributes a failure or regression to a skill by comparing a target skill-guided run against a no-skill or semantically matched skill reference run that solves the same task, or solves it more cheaply. We instantiate this framework on SkillsBench and SWE-Skills-Bench, yielding 307 skill-induced failures, including 125 functional failures and 182 efficiency regressions. We also build SkillTriage, a taxonomy-guided attribution tool that normalizes paired cases, extracts differential evidence, and produces triage reports. Our major findings include: (1) Skill induced functional failures are rarely caused by obviously irrelevant skills; instead, seemingly relevant skills often make the agent incorrectly implement or omit task-required implementation elements. (2) Skill-induced efficiency regressions are not explained by prompt length alone. (3) The largest sources within Excessive Procedure are excessive verification and heavy implementation pipelines, contributing 67 and 30 cases, respectively. This shows that skills often turn validation checklists and construction recipes into mandatory work. Based on our findings, we propose research topics and tooling improvements for safer and more cost-aware skill reuse.
摘要:代理技能是擴展LLM代理的事實機制,提供可重用的指導。技能可以影響代理的任務執行,包括規劃、工具使用、問題解決和驗證。先前的研究報告了代理技能的混合結果:一些技能提高了任務成功率,而其他技能則沒有影響,增加了令牌使用和執行時間,甚至降低了成功率。本文通過將任務失敗和成本回歸歸因於特定的加載技能,呈現了技能引起的代理失敗的全面分析。我們引入了一個差異分析框架,通過將目標技能指導的運行與無技能或語義匹配的技能參考運行進行比較,將失敗或回歸歸因於某個技能,這兩者解決相同的任務,或以更低的成本解決。 我們在SkillsBench和SWE-Skills-Bench上實現了這一框架,產生了307個技能引起的失敗,包括125個功能失敗和182個效率回歸。我們還構建了SkillTriage,一個基於分類法的歸因工具,標準化配對案例,提取差異證據,並生成分類報告。我們的主要發現包括:(1)技能引起的功能失敗很少是由明顯不相關的技能造成的;相反,似乎相關的技能常常使代理錯誤地實施或省略任務所需的實施元素。(2)技能引起的效率回歸並不僅僅由提示長度解釋。(3)在過度程序中,最大的來源是過度驗證和繁重的實施管道,分別貢獻了67和30個案例。這顯示技能常常將驗證清單和建構食譜轉變為強制性工作。根據我們的發現,我們提出了更安全和更具成本意識的技能重用的研究主題和工具改進建議。
Low-Interaction-Rank Learning: Unifying Multiplicative Dual-Encoder Heads
2608.11661v1 by Zijian Zhao, Sen Li
A multiplicative dual-encoder network computes a real-valued output for a pair of inputs as the inner product of their separate encodings. This architecture has been developed independently in operator learning, bipartite matching, contrastive vision-language models, retrieval, and other areas, yet no unified theory guides the basic design decisions: how many interaction modes to represent, how to normalize the encoders, and when the architecture should be avoided. We provide such a foundation by introducing the class of functions of low interaction rank, a class whose intrinsic complexity is measured by its interaction spectrum. Within this framework, approximation error decomposes into a spectral truncation term and an encoder-realization term; sample complexity is governed by the sum of the two encoder complexities rather than their product; and a usability criterion based on spectral decay determines when the architecture can succeed. The same framework exposes a central identifiability problem: the encoders are defined only up to a linear gauge symmetry that leaves the learned coordinates arbitrary. We show that normalization is gauge fixing and that whitening pins the interaction modes up to permutation and sign, thereby explaining the uninterpretability of contrastive dimensions and providing a constructive remedy. Experiments on synthetic kernels, operator learning, and CLIP models validate the theoretical predictions: spectral decay rates match the predicted scaling, whitening recovers the true modes, and independently trained CLIP models are related by a single rotation which, after removal by whitening, exposes interpretable concept axes. The code of this paper is provided at https://github.com/RS2002/Mul-Net .
摘要:一個乘法雙編碼器網絡為一對輸入計算實值輸出,作為其各自編碼的內積。這種架構在運算學習、二部匹配、對比視覺-語言模型、檢索及其他領域中獨立發展,但沒有統一的理論指導基本設計決策:如何表示互動模式的數量、如何對編碼器進行正規化,以及何時應避免使用該架構。我們通過引入低互動秩的函數類別提供這樣的基礎,這個類別的內在複雜性由其互動譜來衡量。在這個框架內,近似誤差分解為光譜截斷項和編碼器實現項;樣本複雜性由兩個編碼器複雜性的總和來決定,而不是它們的乘積;基於光譜衰減的可用性標準決定了該架構何時能夠成功。同樣的框架揭示了一個中心可識別性問題:編碼器僅在一個線性規範對稱下被定義,這使得學習到的坐標是任意的。我們展示了正規化是規範固定,並且白化將互動模式固定到置換和符號上,從而解釋了對比維度的不可解釋性並提供了一個建設性的補救措施。在合成核、運算學習和 CLIP 模型上的實驗驗證了理論預測:光譜衰減率與預測的縮放相匹配,白化恢復了真實模式,獨立訓練的 CLIP 模型通過單一旋轉相關,這在白化後去除後揭示了可解釋的概念軸。本文的代碼可在 https://github.com/RS2002/Mul-Net 獲得。
Inverse Theory of Mind Modeling for Content Recommendation: From Web Browsing to Dynamic Intelligent Interfaces
2608.11354v1 by Mengyu Chen, Feiyu Lu, Chun-Fu Chen, Lucas Vinh Tran, Jay Katukuri
Modern recommender systems treat observed actions as reliable proxies for user preferences, yet interactions often reflect exploration or comparison rather than stable preference expression. As interfaces evolve from static layouts toward generative UIs and immersive extended reality (XR), the need for deeper, modality-agnostic user understanding grows: these adaptive environments must decide not only what to present but where, when, how prominently, and most importantly why a user acts. We propose an Inverse Theory of Mind (IToM) pipeline that reasons backward from observed interactions to infer the beliefs, preferences, and decision-making traits that explain behavior. The pipeline reconstructs each user's decision context, including what was chosen and what alternatives were available, applies LLM-driven counterfactual reasoning to produce evidence-grounded natural-language belief statements, and synthesizes these beliefs through multi-hypothesis abductive inference into a structured user persona. We evaluate on the OPeRA dataset against ground-truth personality assessments, attitudinal surveys, and interview-based personas across four tasks: next action prediction, shopping attitude alignment, Big Five personality inference, and held-out category prediction. Results show that inferred personas match or exceed ground-truth personas and that multi-hypothesis reasoning is essential for accurate personality prediction. We further demonstrate cross-modal transferability with a persona-driven spatial banking application on VisionOS.
摘要:現代推薦系統將觀察到的行為視為用戶偏好的可靠代理,但互動往往反映探索或比較,而非穩定的偏好表達。隨著介面從靜態佈局演變為生成式用戶介面和沉浸式擴增實境(XR),對於更深入的、與模式無關的用戶理解的需求日益增長:這些自適應環境不僅必須決定展示什麼,還要決定在何處、何時、以多大程度以及最重要的原因為何用戶會採取行動。我們提出了一個逆向心智理論(IToM)流程,該流程從觀察到的互動中推理回溯,以推斷解釋行為的信念、偏好和決策特徵。該流程重建每個用戶的決策背景,包括所選擇的內容和可用的替代選項,應用基於大型語言模型(LLM)的反事實推理來生成基於證據的自然語言信念陳述,並通過多假設的溯因推理將這些信念合成為結構化的用戶角色。我們在OPeRA數據集上進行評估,對比真實的個性評估、態度調查和基於訪談的角色,涵蓋四個任務:下一步行動預測、購物態度對齊、五大人格推斷和保留類別預測。結果顯示,推斷出的角色與真實角色相匹配或超過,並且多假設推理對準確的人格預測至關重要。我們進一步展示了在VisionOS上的一個以角色為驅動的空間銀行應用的跨模式可轉移性。
Governing Agentic AI in FinTech
2608.11344v2 by Henry Han
Financial institutions are delegating consequential decisions to agentic AI systems that decompose goals, coordinate models and tools, and act with little oversight. Yet agentic AI governance in FinTech is under-investigated. We argue the binding governance constraint is not capability but verifiability. We define the Verifiability Gap as the shortfall between the verification delegated authority demands and the explainability and reproducibility retained after a decision. It is indexed to a verifier, evidentiary standard, and audit lag. We develop a multilevel governance theory for agentic AI and test its mechanisms in three studies over nine model versions, from a three-billion-parameter local model to a commercial frontier system. Study 1 shows that provider releases alter historical financial actions, and that the controls replay needs belong to the provider: the frontier model rejects temperature, top_p and top_k outright and exposes no random seed. Under the tightest controls each endpoint allows, a local model reproduced 320 of 320 executions, hosted models 319 of 320 and 959 of 960. Study 2 shows that orchestration is a latent policy layer. Architecture changes final actions, and no execution record repeated in any configuration at any scale. The frontier model reproduces its own actions more often than the local ones, its record no better, and loses a comparable share of its differentiation. Capability buys a higher starting point, not auditability. Study 3 shows two deterministic credit-model versions each reproduce their current action perfectly, yet the current cannot recover a historical one. We conceptualize reproducibility as a governance profile, not a scalar, yielding evidence-contingent delegation: authority is defensible only while retained evidence substantiates its exercise. Beyond finance, the framework extends to other high-stakes domains requiring auditability.
摘要:金融機構正在將重要決策委託給能夠分解目標、協調模型和工具並在幾乎沒有監督的情況下行動的代理 AI 系統。
然而,在金融科技領域,代理 AI 的治理尚未得到充分研究。
我們認為,約束治理的關鍵不是能力,而是可驗證性。
我們將可驗證性差距定義為驗證委託權限要求與決策後保留的可解釋性和可重複性之間的差距。
它與驗證者、證據標準和審計延遲有關。
我們為代理 AI 發展了一個多層次的治理理論,並在三項研究中測試其機制,涵蓋九個模型版本,從一個三十億參數的本地模型到一個商業前沿系統。
研究 1 顯示,提供者的發布改變了歷史金融行為,並且控制重播需求屬於提供者:前沿模型直接拒絕 temperature、top_p 和 top_k,並且不暴露隨機種子。
在每個端點允許的最嚴格控制下,本地模型重現了 320 次執行中的 320 次,託管模型重現了 319 次中的 320 次和 960 次中的 959 次。
研究 2 顯示,協調是一個潛在的政策層。
架構改變了最終行動,且在任何配置下的任何規模中都沒有執行記錄重複。
前沿模型比本地模型更頻繁地重現其自身行動,其記錄並無改善,並失去了相當一部分的差異化。
能力提供了一個更高的起點,而不是可審計性。
研究 3 顯示,兩個確定性的信用模型版本各自完美重現其當前行動,但當前模型無法恢復歷史行動。
我們將可重複性概念化為一種治理配置,而不是一個標量,產生證據依賴的委託:權威只有在保留的證據證實其行使時才是可辯護的。
超越金融,該框架擴展到其他需要可審計性的高風險領域。
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop
2608.11171v1 by Rahul Gupta, Abhinav Mohanty, Anaelia Ovalle, Anil Ramakrishna, Anubrata Das, Apurv Verma, Jwala Dhamala, Ninareh Mehrabi, Tharindu Kumarage, Yada Pruksachatkun, Yang Trista Cao, Kai-Wei Chang, Aram Galstyan
The Workshop on Trustworthy Natural Language Processing (TrustNLP), co-located with major ACL conferences since 2021, has grown from 8 proceedings papers to 41 over six editions, documenting a field-wide transition from post-hoc interpretability of static models to mechanistic understanding and proactive control of generative systems. We synthesize insights from all 144 proceedings papers, classifying them along six trust dimensions grounded in established frameworks (TrustLLM, DecodingTrust). We observe co-occurrences with capability emergence. The release of the first high-impact chat models activated all trust dimensions simultaneously, while subsequent model generations shifted focus toward truthfulness and safety alignment. Analysis from the classification study reveals that truthfulness is the fastest-growing dimension (absent in 2021-2022, comprising 37% of papers by 2025-2026), fairness remains the most consistent theme, and explainability exhibits a U-shaped trajectory; declining as post-hoc methods lost relevance but resurging in 2026 through mechanistic interpretability. A cross-venue comparison with ACL, NAACL, EACL, and EMNLP (~2K papers) in the same period shows that TrustNLP's topical distribution closely follows the field average. We identify four structural insights and conclude with actionable directions for the research community.
摘要:信任自然語言處理研討會(TrustNLP)自2021年以來與主要ACL會議共同舉辦,已從8篇會議論文增長至六屆的41篇,記錄了該領域從靜態模型的事後可解釋性到生成系統的機械理解和主動控制的轉變。我們綜合了所有144篇會議論文的見解,並根據建立的框架(TrustLLM, DecodingTrust)將其分類為六個信任維度。我們觀察到能力出現的共現現象。首批高影響力的聊天模型的發布同時激活了所有信任維度,而隨後的模型世代則將重點轉向真實性和安全對齊。分類研究的分析顯示,真實性是增長最快的維度(在2021-2022年缺失,到2025-2026年佔據37%的論文),公平性仍然是最一致的主題,而可解釋性則顯示出U型軌跡;在事後方法失去相關性時下降,但在2026年通過機械可解釋性再次上升。與ACL、NAACL、EACL和EMNLP(約2K篇論文)在同一時期的跨場域比較顯示,TrustNLP的主題分佈與該領域的平均水平密切相符。我們確定了四個結構性見解,並以可行的研究方向作為結論。
SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure
2608.11079v2 by Xiaofan Bai, Hongqiang Lin, Chao Liu, Yantao Zhang, Xuan Jin, Xipeng Cao, Yuhong Li
Self-evolving agents accumulate reusable skills by appending successful procedures and failure fixes. Over time, the same requirement is often restated in several branches, examples, and warnings, while common action sequences are copied rather than reused. The resulting skill becomes expensive to inject and difficult to maintain. Generic prompt compression is ill-suited to this setting because a skill is not a flat passage: its name and description define when it applies, its workflow controls execution, its tool and output contracts constrain validity, and rare exceptions may remain essential even when no sampled task activates them. Evaluation-guided compression can test these behaviors, but it introduces rollouts, cost, and dependence on the compression-time evaluation set. We present SkillZip, an evaluation-free method that compresses a skill by finding its shortest faithful structural explanation. The intuition is explain once, reference many: state a repeated rule once at the scope where it applies, factor a repeated action sequence into a shared procedure, and keep only the differences as explicit exceptions. We formalize this intuition as a typed minimum description-length objective over a skill contract and a residual, subject to a hard coverage constraint for every extracted trigger, workflow edge, tool requirement, obligation, and output field. The formulation provides simple sharing thresholds, preserves unique rare rules by construction, and supports efficient local updates. SkillZip has a one-shot mode with one structured extraction call and deterministic optimization, and a continual Zip-on-Write mode that integrates each self-evolution patch without replaying tasks or reparsing the full history. Through comprehensive experimental evaluations, we demonstrate the effectiveness and superiority of SkillZip in compression performance, generalizability, and cost overhead.
摘要:自我演化的代理透過附加成功的程序和失敗的修正來累積可重複使用的技能。隨著時間的推移,同樣的需求常常在幾個分支、範例和警告中重新陳述,而常見的行動序列則是被複製而不是重用。由此產生的技能變得難以注入且難以維護。通用提示壓縮不適合這種情況,因為技能不是一段平坦的文字:它的名稱和描述定義了何時適用,其工作流程控制執行,其工具和輸出契約限制有效性,而即使在沒有樣本任務激活的情況下,稀有的例外可能仍然是必要的。評估引導的壓縮可以測試這些行為,但它引入了推出、成本和對壓縮時評估集的依賴。我們提出了SkillZip,一種無需評估的方法,通過找到技能的最短忠實結構解釋來壓縮技能。其直覺是一次解釋,多次參考:在適用的範圍內一次陳述重複的規則,將重複的行動序列分解為共享的程序,並僅保留差異作為明確的例外。我們將這一直覺形式化為一個類型化的最小描述長度目標,針對技能契約和殘餘,並對每個提取的觸發器、工作流程邊緣、工具需求、義務和輸出欄位施加嚴格的覆蓋約束。該公式提供簡單的共享閾值,通過構造保留獨特的稀有規則,並支持高效的本地更新。SkillZip具有一次性模式,通過一次結構化提取調用和確定性優化,還有持續的Zip-on-Write模式,能夠在不重播任務或重新解析完整歷史的情況下集成每個自我演化補丁。通過全面的實驗評估,我們展示了SkillZip在壓縮性能、可泛化性和成本開銷方面的有效性和優越性。
Entropy-Centric Explainable AI for Remote Sensing Image Segmentation
2608.11064v1 by Ali Saleh, Abdul Karim Gizzini, Mohamad Ghassany, Ali J. Ghandour
Artificial intelligence (AI) has become a powerful approach to solving complex problems in critical domains. Many concerns arise regarding the decision-making process of its models, mainly due to deep neural networks outperforming their peers at the cost of ambiguity in feature extraction and prediction. Consequently, in critical domains such as remote sensing, where high-resolution imagery must be analyzed using black-box models, the lack of transparency limits trust in these models and, thus, their adoption. In light of this reality, explaining and understanding the complex decision-making process of AI models has become essential. Explainable AI (XAI) aims to bridge this gap by providing insights into how and why certain decisions are made. While significant progress has been achieved in explaining image classification tasks, image segmentation still offers considerable room for improvement. In this context, this paper proposes an entropy-centric XAI method for semantic segmentation. Moreover, a new XAI evaluation methodology is proposed to efficiently measure the relevance of the regions highlighted by the proposed XAI method. Experimental results demonstrate the superiority of the proposed XAI method compared with recently adapted XAI methods for semantic segmentation.
摘要:人工智慧(AI)已成為解決關鍵領域複雜問題的強大方法。
許多關於其模型決策過程的擔憂隨之而來,主要是因為深度神經網絡在特徵提取和預測的模糊性方面超越了其同儕。
因此,在遙感等關鍵領域,必須使用黑箱模型分析高解析度影像,缺乏透明度限制了對這些模型的信任,從而影響了它們的採用。
鑑於這一現實,解釋和理解AI模型複雜的決策過程變得至關重要。
可解釋的AI(XAI)旨在通過提供對某些決策如何以及為何做出的見解來彌補這一差距。
儘管在解釋影像分類任務方面已取得顯著進展,但影像分割仍然有相當大的改進空間。
在這一背景下,本文提出了一種以熵為中心的XAI方法,用於語義分割。
此外,還提出了一種新的XAI評估方法,以有效測量所提出的XAI方法所突顯區域的相關性。
實驗結果顯示,所提出的XAI方法在語義分割方面優於最近適應的XAI方法。
ComBodied Agents: a New Paradigm of Human-Centric Agentic AI
2608.10915v2 by Qianggang Ding, Xingyao Wang, Rui Feng, Zhibin Wang, Feixiang Yao, Kelong Mao, Hao Sun, Zhiyao Luo, Jiankai Tang, Lei Li, Jiadong Guo, Minheng Ni, Weicong Lin, Chenxi Yang, Hongxiang Gao, Zhenghua Chen, Yang Bai, Min Wu, Jun Cheng, Huazhu Fu, Dacheng Tao, Bang Liu
After an older adult misses a medication dose, a software agent can send another reminder and an embodied agent can bring the medication. Yet neither explains whether the person forgot, is confused, has side effects, or deliberately refused, nor what support is appropriate. This reveals a structural gap in Agentic AI: Digital Agents primarily transform software states, while Embodied Agents transform physical states; neither makes a person's evolving state and agency the primary object of modeling, intervention, and evaluation. We introduce Combodied Agents, a human-centered paradigm that perceives, models, predicts, and supports individual human-state trajectories over time, using software tools, sensors, wearables, robots, and human services as action channels rather than end goals. We unify fragmented capabilities across personal assistants, health agents, AI companions, and adaptive human--AI systems into a closed loop: event-based multimodal perception reconstructs meaningful personal events; longitudinal, correctable memory provides temporal context; Personal World Models estimate future personal states and outcomes under alternative decisions and interventions; and an admissible intervention policy selects proportionate support under consent, uncertainty, safety, reversibility, and user control. Feedback from the person and environment updates the loop. Rather than requiring an exhaustive Human Digital Twin, the framework uses purpose-bounded, uncertainty-aware, user-correctable representations. We organize the design space by human-state targets, relational contexts, and agent roles, and propose scenario-centered evaluation, agency-preservation metrics, benchmark requirements, edge-native personal models, and governance directions. Combodied Agents shift Agentic AI from external task completion toward sustained human benefit.
摘要:在年長者錯過藥物劑量後,軟體代理可以發送另一個提醒,而具身代理可以帶來藥物。
然而,這兩者都沒有解釋該人是否忘記、感到困惑、出現副作用或故意拒絕,也沒有說明什麼樣的支持是合適的。
這揭示了代理人工智能中的結構性缺口:數位代理主要轉換軟體狀態,而具身代理則轉換物理狀態;兩者都未將個體不斷演變的狀態和能動性作為建模、干預和評估的主要對象。
我們引入了具身代理(Combodied Agents),這是一種以人為中心的範式,能夠隨著時間的推移感知、建模、預測和支持個體的人類狀態軌跡,使用軟體工具、感測器、可穿戴設備、機器人和人類服務作為行動渠道,而非最終目標。
我們將個人助理、健康代理、人工智慧伴侶和自適應人類-人工智慧系統的零散能力統一成一個閉環:基於事件的多模態感知重建有意義的個人事件;長期的、可修正的記憶提供時間背景;個人世界模型在不同的決策和干預下估計未來的個人狀態和結果;可接受的干預政策在同意、不確定性、安全性、可逆性和用戶控制下選擇相稱的支持。
來自個人和環境的反饋更新這個循環。
該框架不需要全面的人類數位雙胞胎,而是使用目的有限、具不確定性意識和用戶可修正的表徵。
我們根據人類狀態目標、關係背景和代理角色來組織設計空間,並提出以情境為中心的評估、能動性保護指標、基準要求、邊緣原生個人模型和治理方向。
具身代理將代理人工智能的重心從外部任務完成轉向持續的人類利益。
Chemically Meaningful Textualization Enables Explainable Validation of Metal-Organic Frameworks by Large Language Models
2608.11283v1 by Guobin Zhao, Xiao-Yan Li
Computation-ready metal-organic framework (MOF) databases are essential for high-throughput screening, yet many reported crystal structures remain chemically unreasonable or disordered, compromising simulation fidelity. Existing validation approaches can identify non-computation-ready structures, but they often rely on heuristic rules, license requirement, or offer limited interpretability. Here, we show that large language models (LLMs) can serve as interpretable validators of MOF structures when crystallographic information is transformed into chemically meaningful text. By benchmarking nine descriptors, we find that successful LLM-based validation depends not on the amount of structural information alone, but on whether local coordination, framework connectivity, and chemical context are organized into a linguistically learnable representation. Fine-tuned LLMs using specialized descriptors (mof2text) achieve performance comparable to graph-based models in identifying unreasonable MOFs. Importantly, these models extend beyond black-box classification by generating diagnostic rationales for likely error sources, including abnormal bonding, connectivity, and charge states, as well as error-category predictions for annotated datasets. This work establishes chemically informed textualization as the key step that transforms LLMs from generic text models into practical and explainable tools for curating MOF databases.
摘要:計算準備好的金屬有機框架 (MOF) 數據庫對於高通量篩選至關重要,但許多報告的晶體結構仍然化學上不合理或無序,從而影響模擬的真實性。現有的驗證方法可以識別非計算準備好的結構,但它們通常依賴於啟發式規則、許可要求,或提供有限的可解釋性。在這裡,我們展示了大型語言模型 (LLMs) 可以作為 MOF 結構的可解釋驗證者,當晶體學信息轉換為化學上有意義的文本時。通過基準測試九個描述符,我們發現成功的 LLM 基於驗證不僅取決於結構信息的數量,還取決於局部配位、框架連通性和化學上下文是否組織成語言上可學習的表示。使用專門描述符 (mof2text) 的微調 LLM 在識別不合理的 MOF 方面達到與基於圖的模型相當的性能。重要的是,這些模型超越了黑箱分類,通過生成可能錯誤來源的診斷理由,包括異常鍵合、連通性和電荷狀態,以及對註釋數據集的錯誤類別預測,來擴展其功能。這項工作確立了化學知識驅動的文本化作為關鍵步驟,將 LLM 從通用文本模型轉變為實用且可解釋的工具,以便策劃 MOF 數據庫。
Uncertainty-Aware and Explainable Ensemble Deep Learning Framework for Multi-Class Skin Lesion Classification
2608.11280v1 by Rofiqul Islam, Lilatul Ferdouse
Skin cancer diagnosis from dermoscopic images remains challenging due to high intra-class variability, inter-class similarity, class imbalance, and the limited interpretability of deep learning models. This paper proposes an uncertainty-aware and explainable deep learning framework for multi-class skin lesion classification. The framework combines a vision transformer model (MaxViT-Tiny) with CNN-based models (ConvNeXt-Tiny and EfficientNetV2-B0) through deep ensemble learning. Monte Carlo (MC) Dropout estimates predictive uncertainty and identifies unreliable predictions, while Grad-CAM++, an explainable AI (XAI) technique, provides visual explanations by highlighting lesion regions that influence model decisions. Evaluated on the HAM10000 dataset, the framework achieves 96% accuracy and 99% ROC-AUC under uncertainty-aware filtering (entropy < 1.0, confidence >= 0.7), with macro-average precision, recall, and F1-score of 94%, 95%, and 95%, respectively, and 96% weighted-average scores across all three metrics. The results demonstrate accurate, interpretable, and uncertainty-aware skin lesion classification for trustworthy computer-aided diagnosis.
摘要:皮膚癌的診斷從皮膚鏡影像中仍然具有挑戰性,這是由於高內類變異性、類間相似性、類別不平衡以及深度學習模型的有限可解釋性。本文提出了一種不確定性感知和可解釋的深度學習框架,用於多類別皮膚病變分類。該框架通過深度集成學習將視覺Transformer模型(MaxViT-Tiny)與基於CNN的模型(ConvNeXt-Tiny和EfficientNetV2-B0)相結合。蒙特卡羅(MC)Dropout估計預測不確定性並識別不可靠的預測,而Grad-CAM++,一種可解釋的人工智慧(XAI)技術,通過突出影響模型決策的病變區域提供視覺解釋。在HAM10000數據集上進行評估,該框架在不確定性感知過濾(熵 < 1.0,置信度 >= 0.7)下達到96%的準確率和99%的ROC-AUC,宏觀平均精確度、召回率和F1-score分別為94%、95%和95%,在所有三個指標上達到96%的加權平均分數。結果顯示出準確、可解釋且具不確定性感知的皮膚病變分類,為可信的計算機輔助診斷提供支持。
Rule of Thumb: Explaining Artificial Intelligence Systems using Partial Information
2608.10766v2 by Kaivalya Rawal, Daria Onitiu, Brent Mittelstadt, Sandra Wachter, Chris Russell
Explainable Artificial Intelligence (XAI) seeks to explain how an Artificial Intelligence (AI) system arrived at a particular decision. We propose ''Rule of Thumb'' (RoT) explanations, a new approach to XAI based upon a novel formulation that identifies the most relevant features for predicting the behaviour of an AI system, for a particular datapoint. We show how RoT is well-suited to enable XAI in: (a) zero-shot classification using large language models (LLMs), (b) auditing of opaque AI systems without model access, and (c) the use of AI in scientific discovery. Additionally, RoT meets specific requirements from leading AI regulations, provides a familiar interface and visualisations for XAI practitioners, is model-agnostic, and is substantially faster than alternatives. Code available at: https://github.com/KaiRawal/Rule-of-Thumb-Explaining-Artificial-Intelligence-Systems-using-Partial-Information
摘要:可解釋的人工智慧(XAI)旨在解釋人工智慧(AI)系統如何做出特定決策。我們提出了「經驗法則」(RoT)解釋,這是一種基於新型公式的XAI新方法,能夠識別對於特定數據點預測AI系統行為最相關的特徵。我們展示了RoT如何適合於以下情境以促進XAI:(a)使用大型語言模型(LLMs)進行零樣本分類,(b)在無法訪問模型的情況下對不透明的AI系統進行審計,以及(c)在科學發現中使用AI。此外,RoT符合主要AI法規的特定要求,為XAI從業者提供熟悉的界面和可視化,並且對模型無關,速度也比其他替代方案快得多。
代碼可在:https://github.com/KaiRawal/Rule-of-Thumb-Explaining-Artificial-Intelligence-Systems-using-Partial-Information
Operationalising Relative Causal Knowledge: Backbone Identifiability from Private Reports on a Shared Outcome
2608.10664v1 by Fabrizio Russo, Mark Somers
The Relativity of Causal Knowledge (RCK) explains how a network of agents with different structural causal models can exchange causal knowledge through a shared interventionally consistent abstraction, or backbone. We ask the prior identification question that this transport mechanism presupposes: when is that backbone determined by the agents' private causal knowledge? In the basic two-agent common-effect case, two private causes influence one shared outcome and each agent identifies only the single-cause causal marginal relevant to its own perspective. We show that, under standard compatibility, non-degeneracy, and local overlap assumptions, those local causal marginals do not identify a unique backbone. Infinitely many joint intervention kernels can induce exactly the same private reports while disagreeing on joint interventions. We then give a conditional recovery result. Additive separability removes the hidden interaction degree of freedom, but observational residual summaries remain insufficient. Identification becomes possible when agents communicate causally identified response functions. An education value-added example illustrates why this is first a communication problem, and only then a policy-composition problem.
摘要:因果知識的相對性(RCK)解釋了不同結構因果模型的代理人網絡如何通過共享的干預一致抽象或骨幹來交換因果知識。我們提出這一傳輸機制所假設的先前識別問題:何時骨幹由代理人的私人因果知識決定?在基本的兩代理人共同效果案例中,兩個私人原因影響一個共享結果,而每個代理人僅識別與其自身觀點相關的單一原因因果邊際。我們表明,在標準兼容性、非退化性和局部重疊假設下,這些局部因果邊際並不識別唯一的骨幹。無限多的聯合干預核可以產生完全相同的私人報告,同時在聯合干預上存在分歧。然後,我們給出一個條件恢復結果。加性可分離性消除了隱藏的交互自由度,但觀察殘差摘要仍然不足。當代理人交流因果識別的反應函數時,識別變得可能。一個教育增值的例子說明了為什麼這首先是一個通信問題,而後才是一個政策組合問題。
Conversational versus Dashboard Explainable AI for UAV Intrusion Detection: An Empirical Study of Operator Trust and Reliance
2608.10434v1 by Cong Chi Nguyen, Trang Mai Xuan, Vu-Duc Ngo, Kim-Ngan Thi Nguyen, Trong-Nghia Nguyen, Thien Van Luong
Machine learning-based Intrusion Detection Systems (IDS) have demonstrated superior performance in securing Unmanned Aerial Vehicle (UAV) networks. However, the 'black-box' nature of these models, combined with the high dimensionality of multimodal cyber-physical data, poses significant interpretability challenges. Static visualization dashboards may struggle to present complex relationships among multimodal cyber-physical features in a form that is easy for operators to inspect and interpret. To address this, we propose a Conversational XAI interface powered by Large Language Models (LLM) to facilitate on-demand investigation. In a controlled experiment with participants, we systematically evaluated the impact of this conversational interface versus a traditional XAI Dashboard on operator understanding, trust, and reliance during post-incident auditing tasks. Our results suggest that the conversational interface was perceived as more useful than the dashboard, potentially because it helped participants access and synthesize relevant information more easily. However, this benefit was accompanied by a lower level of appropriate self-reliance, indicating a potential risk of over-reliance. One possible interpretation is that the natural-language responses made the AI advice easier to accept, which may have reduced participants' tendency to verify the underlying evidence when the IDS was incorrect. These findings point to a potential trade-off in human-AI collaboration for UAV intrusion auditing: interaction mechanisms that improve perceived usability may also increase the risk of inappropriate reliance. We conclude by discussing design implications for future XAI systems that balance seamless interaction with cognitive forcing functions to foster appropriate reliance.
摘要:基於機器學習的入侵檢測系統(IDS)在保護無人機(UAV)網絡方面顯示出優越的性能。
然而,這些模型的「黑箱」特性,加上多模態網絡物理數據的高維度,帶來了顯著的可解釋性挑戰。
靜態可視化儀表板可能難以以易於操作員檢查和解釋的形式呈現多模態網絡物理特徵之間的複雜關係。
為了解決這個問題,我們提出了一個由大型語言模型(LLM)驅動的對話式XAI界面,以促進隨需調查。
在一項對參與者的控制實驗中,我們系統地評估了這個對話式界面與傳統XAI儀表板對操作員理解、信任和依賴在事件後審計任務中的影響。
我們的結果表明,對話式界面被認為比儀表板更有用,這可能是因為它幫助參與者更輕鬆地訪問和綜合相關信息。
然而,這一好處伴隨著較低的適當自我依賴水平,顯示出過度依賴的潛在風險。
一種可能的解釋是,自然語言的回答使得AI建議更容易被接受,這可能減少了參與者在IDS不正確時驗證基礎證據的傾向。
這些發現指出了無人機入侵審計中人機協作的潛在權衡:改善感知可用性的互動機制可能也會增加不當依賴的風險。
我們最後討論了未來XAI系統的設計啟示,旨在平衡無縫互動與認知強迫功能,以促進適當的依賴。
Reasoning Shortcuts and Value Symmetries: What Symmetry Permits, Architecture Realizes, and Optimization Selects
2608.10420v1 by Xin Xu
Reasoning shortcuts are solutions of a neurosymbolic system's rules that produce correct predictions through unintended concepts. A recent framework of Takemura, Inoue, and Nishino analyzes them through an automorphism group of value relabelings and asks, as its central open question, when rules pin concepts down. We first show that the framework's key definition, one shared permutation applied at every position, does not apply as stated to any of the four heterogeneous benchmarks it was evaluated on, and that the most direct embedding, padding domains to a common size, produces confident false pathology: 90.91% of solution pairs reported unexplained on CLE4EVR, where every well-defined member of the hierarchy we introduce reports 0%, and the padded verdict's content rotates with configuration-file ordering. Re-measuring eleven rule families under fifteen pre-specified predictions (thirteen confirmed), unexplained-pair rates span 0% to 99.9999% and track provable structure: six theorems give sufficient conditions for transitivity and its failure, including a Free Slot Lemma certifying Kandinsky's pathology from syntax alone. For circuit-given rules, deciding symmetry-inertness of a coordinate is coNP-complete; nontrivial-automorphism existence is coNP-hard under randomized reductions, lies in $Σ_2^p$, is not $Σ_2^p$-complete unless PH collapses, and on monotone circuits is coNP-complete outright. In the Boolean case transitivity is classified exactly: automorphisms explain everything iff the solution set is an affine coset. Weakly supervised models place all 94 observed shortcuts at the one level the componentwise theory flags and none at the 48 it certifies transitive; twelve typed-ambiguous levels produce none, separating what symmetry permits from what optimization selects, and a dual-head control replicates the geography. All numbers trace to released artifacts.
摘要:推理捷徑是神經符號系統規則的解決方案,通過意外概念產生正確的預測。Takemura、Inoue 和 Nishino 最近提出的一個框架通過值重新標記的自同構群來分析它們,並將何時規則固定概念作為其核心未解問題。我們首先展示該框架的關鍵定義,即在每個位置應用的共享排列,並未如所述適用於其評估的四個異質基準中的任何一個,並且最直接的嵌入,即將域填充到共同大小,產生了自信的錯誤病理:在 CLE4EVR 上報告的解決方案對中有 90.91% 的未解釋情況,而我們引入的每個明確定義的層級報告 0%,且填充判決的內容隨配置文件排序而旋轉。在十五個預先指定的預測下(確認了十三個),重新測量十一個規則家族,未解釋對的比例從 0% 到 99.9999% 不等,並追踪可證明的結構:六個定理給出了傳遞性及其失敗的充分條件,包括一個自由槽引理,僅通過語法證明 Kandinsky 的病理。對於電路給定的規則,決定坐標的對稱惰性是 coNP 完全的;在隨機約簡下,非平凡自同構的存在是 coNP 困難的,位於 $Σ_2^p$ 中,除非 PH 崩潰,否則不是 $Σ_2^p$ 完全的,並且在單調電路上是 coNP 完全的。在布爾情況下,傳遞性被準確分類:自同構解釋一切當且僅當解決方案集是一個仿射陪集。弱監督模型將所有 94 個觀察到的捷徑放置在組件理論標記的唯一層級上,而在其證明為傳遞的 48 個層級上則沒有;十二個類型模糊的層級產生零,將對稱允許的與優化選擇的分開,雙頭控制複製了地理。所有數字都追溯到釋放的工件。
Towards Expert-level Medical AI for Real-time Video Consultations
2608.09861v1 by Mahvish Nagda, Jihyeon Lee, Matthew Thompson, Chunjong Park, Tim Strother, Valentin Liévin, Roma Ruparel, Akshay Goel, Teya Bergamaschi, Suhana Bedi, Meet Shah, Pavel Dubov, Liviu Panait, Toshiyuki Fukuzawa, Sam Schmidgall, Craig Schiff, Joseph Xu, Aliya Rysbek, Yana Lunts, Jan Freyberg, Rebecca Hemengway, Sunny Virmani, David Racz, Carey Radebaugh, Joëlle Barral, Kavi Goel, Dale R. Webster, Katherine Chou, Avinatan Hassidim, Yossi Matias, James Manyika, Gregory Wayne, Tao Tu, Yun Liu, Ethan Goh, Christina Chen, Ryutaro Tanno, Po-Hsuan Cameron Chen, Mike Schaekermann, Anil Palepu
Audio-visual interaction is the standard for patient-physician consultations, enabling natural communication and effective assessment of illness through non-verbal cues. While text-based AI has shown promise, it discards essential perceptual dimensions and limits patients who cannot articulate symptoms in writing. Early efforts to extend medical AI to audio-visual interaction have demonstrated feasibility but not reached clinician-level performance. Here, we provide the first demonstration of expert-level AI in real-time clinical video consultations using AMIE (Articulate Medical Intelligence Explorer) in a video configuration. AMIE (Video) is a Gemini-based multi-agent system integrating low-latency dialogue, clinical reasoning, and real-time audio-visual perception. To guide development, we established a taxonomy and automated evaluations for clinical audio-visual cues in telehealth settings. In a randomized Objective Structured Clinical Examination (OSCE) study with 30 primary care physicians (PCPs), 15 patient actors and 100 clinical scenarios, we compared AMIE (Video), its text-only counterpart AMIE (Text), and PCPs consulting via video. Clinical evaluators rated AMIE (Video) on par or better than PCPs in history-taking, diagnosis, management, and physical observation and examination. Patient actors preferred AMIE's approach to assessing and explaining conditions, while PCPs were preferred for rapport and partnership building. In modality ablation, patient actors preferred AMIE (Video)'s interface over text chat for communicative effectiveness, convenience, and feeling understood. Limitations remain in fine anatomical precision, subtle affective nuances, and high-frequency movements. While further research is needed before real-world translation, these results mark an important milestone toward AI systems capable of augmenting care across the sensory complexity of clinical practice.
摘要:視聽互動是病人與醫生諮詢的標準,能夠通過非語言線索促進自然交流和有效評估疾病。雖然基於文本的人工智慧顯示出潛力,但它忽略了重要的感知維度,並限制了無法用書面表達症狀的病人。早期將醫療人工智慧擴展至視聽互動的努力已顯示出可行性,但未達到臨床醫生的表現水平。在此,我們提供了使用AMIE(Articulate Medical Intelligence Explorer)在視頻配置中進行實時臨床視頻諮詢的專家級人工智慧的首次示範。AMIE(視頻)是一個基於Gemini的多代理系統,整合了低延遲對話、臨床推理和實時視聽感知。為了指導開發,我們建立了一個分類法和自動評估,用於遠程醫療環境中的臨床視聽線索。在一項隨機的客觀結構化臨床考試(OSCE)研究中,涉及30位初級保健醫生(PCPs)、15位病人演員和100個臨床場景,我們比較了AMIE(視頻)、其文本專用對應AMIE(文本)以及通過視頻諮詢的PCPs。臨床評估者在病史採集、診斷、管理以及身體觀察和檢查方面評價AMIE(視頻)與PCPs相當或更好。病人演員更喜歡AMIE在評估和解釋病情方面的方法,而PCPs則在建立關係和夥伴關係方面更受青睞。在模態消融中,病人演員更喜歡AMIE(視頻)的界面而非文本聊天,因為其在交流有效性、便利性和被理解的感受上表現更佳。儘管在精細解剖精度、微妙的情感細微差別和高頻運動方面仍存在局限性,但在實際應用之前仍需進一步研究,這些結果標誌著朝著能夠增強臨床實踐中感官複雜性的護理的人工智慧系統邁出了重要的一步。
CEAA: A Cognitive Embodied Agents Architecture for Interactive Computing Systems
2608.09848v1 by Aimilios Hadjiliasi, Louis Nisiotis
The development of embodied Intelligent Virtual Agents (IVAs) that have cognitive capabilities in real-time interactive virtual environments remains a challenge, even with today's advancements in technology. Existing architectures are often focused on either the implementation of low-level reactive control systems that are constrained by commercial game engines, or high-level representations of reasoning models that can be difficult to implement in virtual worlds. This paper builds on that notion and proposes a modular cognitive architecture for deploying embodied IVAs. This architecture builds on existing, pre-established frameworks such as the Sense-Think-Act paradigm and the Belief-Desire-Intention cognitive model, among others, and aims to provide a reusable implementation-oriented framework as a template for deploying IVA "brains" in interactive 3D computing systems. The proposed architecture contributes by providing a modular, implementation-oriented framework for the deployment of embodied, cognitive-capable IVAs and bridges the gap between high-level agent reasoning models with real-time embodied execution, for scalable, adaptive, and explainable agents in complex interactive virtual environments.
摘要:身體化智能虛擬代理(IVAs)的發展,具備在實時互動虛擬環境中進行認知能力的挑戰,即使在當今科技進步的情況下仍然存在。現有的架構通常專注於低層次反應控制系統的實施,這些系統受限於商業遊戲引擎,或是高層次推理模型的表現,這在虛擬世界中可能難以實施。本文基於這一概念,提出了一種模組化的認知架構,用於部署身體化的IVAs。這一架構基於現有的、預先建立的框架,如感知-思考-行動範式和信念-欲望-意圖認知模型等,旨在提供一個可重用的實施導向框架,作為在互動3D計算系統中部署IVA“大腦”的模板。所提出的架構通過提供一個模組化的、實施導向的框架,為身體化、具備認知能力的IVAs的部署做出貢獻,並彌合高層次代理推理模型與實時身體執行之間的差距,以實現可擴展、適應性強且可解釋的代理,適用於複雜的互動虛擬環境。
KGCaRe: Explainable Complex Conditional Question Answering using Automatic Knowledge Graph Construction and Context Retrieval with LLMs
2608.09779v1 by Ghanshyam Verma, Simanta Sarkar, Devishree Pillai, Hotaka Shiokawa, Yourong Xu, Fiona Veazey, Peter Hubbert, Hui Su, Paul Buitelaar
Answering complex conditional questions using Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) remains a challenge, particularly in domain-specific contexts where general-purpose LLMs and RAG tend to underperform. We hypothesize that augmenting RAG with unstructured and structured knowledge, extracted from both documents and knowledge graphs (KGs), can improve reasoning and answer accuracy for such tasks. To test this, we propose KGCaRe, a hybrid approach that combines neural retrieval with symbolic reasoning over LLM-generated KGs. KGCaRe constructs a KG from documents using a multi-prompt extraction strategy and stores it in a graph database. Simultaneously, the documents are embedded into a vector store to enable neural retrieval. KGCaRe performs innovative iterative graph traversal guided by the LLM to extract relevant triples, prune irrelevant information, and uses additional clue entities to traverse the graph again if the initial traversal does not provide satisfactory context to generate the answer. The relevant triples extracted from the KG in path form, along with semantically retrieved text passages, are then fed into custom KGCaRe prompts to generate answers to the complex conditional questions with explanations. We evaluate KGCaRe on two complex conditional QA datasets. Our results on these datasets show that KGCaRe consistently outperforms existing baselines, including Vanilla LLM, Code Prompt, Text Prompt, Think-on-Graph, Vanilla RAG, and HybridContextQA, across multiple LLMs such as Mistral, Mixtral, GPT-3.5, and GPT-4o. We publicly release the software pipeline that we developed to implement the proposed KGCaRe approach.
摘要:回答複雜的條件問題使用大型語言模型(LLMs)和檢索增強生成(RAG)仍然是一個挑戰,特別是在特定領域的上下文中,通用的LLMs和RAG往往表現不佳。 我們假設,通過從文檔和知識圖(KGs)中提取的非結構化和結構化知識增強RAG,可以改善此類任務的推理和答案準確性。
為了測試這一假設,我們提出了KGCaRe,一種結合神經檢索和LLM生成的KG上符號推理的混合方法。 KGCaRe使用多提示提取策略從文檔中構建KG,並將其存儲在圖形數據庫中。 同時,這些文檔被嵌入到向量存儲中,以便進行神經檢索。 KGCaRe執行創新的迭代圖遍歷,由LLM指導,以提取相關三元組,修剪不相關的信息,並在初始遍歷未能提供滿意的上下文以生成答案時,使用額外的線索實體再次遍歷圖形。 從KG中提取的相關三元組以路徑形式呈現,連同語義檢索的文本段落,然後輸入自定義KGCaRe提示,以生成帶有解釋的複雜條件問題的答案。
我們在兩個複雜的條件QA數據集上評估KGCaRe。 我們在這些數據集上的結果顯示,KGCaRe在多個LLM(如Mistral、Mixtral、GPT-3.5和GPT-4o)上始終超越現有基準,包括Vanilla LLM、Code Prompt、Text Prompt、Think-on-Graph、Vanilla RAG和HybridContextQA。 我們公開發布了我們開發的軟件管道,以實現所提出的KGCaRe方法。
Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models
2608.09666v1 by Shulin Tian, Ziqi Huang, Fan Zhang, Hongyuan Zhu, Yu Qiao, Ziwei Liu
Recent advances in visual generative models have enabled high-quality image and video generation, but evaluating these models often demands sampling hundreds or thousands of images or videos, which is computationally expensive. Existing evaluation methods also rely on rigid pipelines that overlook specific user needs and provide numerical results without clear explanations. Mimicking how humans quickly form impressions of a model's capabilities from only a few samples, we propose the Evaluation Agent framework, which employs human-like strategies for efficient, dynamic, multi-round evaluations, offering detailed, user-tailored analyses. Given a natural-language evaluation request, the agent decomposes it into sub-aspects, generates targeted prompts, samples images or videos from the evaluated model, invokes suitable evaluation tools, and iteratively updates its plan from the observed evidence, covering both predefined benchmark dimensions and open-ended user concerns. The framework is thus efficient, promptable, explainable, and scalable across models and tools. Experiments show that Evaluation Agent reduces evaluation time to 10% of traditional methods while delivering comparable results. We further introduce Open Evaluation Agent (Open-EA) by constructing EA-CoT-10K, a corpus of history-conditioned step-level instruction-tuning records derived from multi-round evaluation rollouts, and training EA-3B from Qwen2.5-3B-Instruct as a local planning backbone that preserves the structured reasoning, tool invocation, and summary protocol of the API-based agent while reducing dependence on proprietary backbones. Experiments validate the API-based agent on established T2I/T2V benchmarks and open-ended queries, and evaluate Open-EA on four in-domain and three out-of-domain T2V generator families, showing partial cross-family transfer of the learned policy.
摘要:最近在視覺生成模型方面的進展使得高品質的圖像和視頻生成成為可能,但評估這些模型通常需要抽樣數百或數千張圖像或視頻,這在計算上是昂貴的。現有的評估方法也依賴於僵化的流程,忽略了特定用戶需求,並提供沒有明確解釋的數值結果。模仿人類如何快速從少量樣本中形成對模型能力的印象,我們提出了評估代理框架,它採用類似人類的策略進行高效、動態的多輪評估,提供詳細的、針對用戶的分析。給定自然語言的評估請求,代理將其分解為子方面,生成針對性的提示,從被評估模型中抽樣圖像或視頻,調用合適的評估工具,並根據觀察到的證據迭代更新其計劃,涵蓋既定的基準維度和開放式的用戶關注點。因此,該框架在模型和工具之間是高效的、可提示的、可解釋的和可擴展的。實驗顯示,評估代理將評估時間減少到傳統方法的10%,同時提供可比擬的結果。我們進一步通過構建EA-CoT-10K,引入開放評估代理(Open-EA),這是一個從多輪評估展開中衍生的歷史條件步驟級指令調整記錄的語料庫,並從Qwen2.5-3B-Instruct訓練EA-3B作為當地規劃的骨幹,保留基於API的代理的結構化推理、工具調用和摘要協議,同時減少對專有骨幹的依賴。實驗在已建立的T2I/T2V基準和開放式查詢上驗證了基於API的代理,並在四個域內和三個域外的T2V生成器家族上評估Open-EA,顯示出學習政策的部分跨家族轉移。
Rethinking Self-Evolving Agents: Do We Still Need Prescribed Optimization Pipelines?
2608.09629v1 by Hui Xue, Fan Yang
Self-evolving agents are usually built around prescribed optimization pipelines: the framework decides how to gather evidence, revise a persistent artifact, select candidates, and stop. We ask whether this task-specific procedure remains necessary when a frontier model acts as the optimizer. We introduce Open-Ended Optimization (OEO), which keeps the objective, permitted interactions, resource budget, data boundary, and evaluation fixed while allowing the optimizer to compose the improvement process online. We compare OEO with two complementary prescribed approaches: SkillOpt, a staged pipeline with bounded edits, and GEPA, a reflective evolutionary search. Across 14 head-to-head comparisons over 8 benchmark-target-model settings, GPT-5.5-driven OEO records 12 wins, 1 tie, and 1 narrow loss of 0.21 percentage points. It uses a median 34.3 percent of SkillOpt's configured target-interaction token budget. A one-shot, zero-interaction control shows that the gains are not explained by a single prior-driven rewrite. However, delegation has a capability boundary: SkillOpt outperforms OEO with a medium optimizer, and a weak optimizer cannot operate through the unchanged OEO interface. In the fully instrumented OEO-SkillOpt pair, trajectory analysis further shows that prescription changes how optimization proceeds more consistently than it changes final behavior. Together, these findings recast prescribed pipelines as capability-dependent scaffolding: essential constraints remain external, but a sufficiently capable optimizer can compose the route from measurable feedback to persistent improvement.
摘要:自我演化代理通常圍繞預定的優化流程構建:框架決定如何收集證據、修訂持久的工件、選擇候選者以及何時停止。我們詢問當前沿模型作為優化器時,這種特定任務的程序是否仍然必要。我們引入了開放式優化(Open-Ended Optimization, OEO),它保持目標、允許的互動、資源預算、數據邊界和評估固定,同時允許優化器在線組合改進過程。我們將OEO與兩種互補的預定方法進行比較:SkillOpt,一種具有有限編輯的分階段流程,以及GEPA,一種反思性進化搜索。在14次面對面的比較中,涵蓋8個基準目標模型設置,GPT-5.5驅動的OEO記錄了12場勝利、1場平局和1場以0.21個百分點的微弱劣勢輸掉的比賽。它使用了SkillOpt配置的目標互動令牌預算的中位數34.3%。一次性、零互動的控制顯示,這些增益並不能用單一的先前驅動重寫來解釋。然而,委託有能力邊界:SkillOpt在中等優化器下表現優於OEO,而弱優化器無法通過不變的OEO介面運作。在完全儀器化的OEO-SkillOpt配對中,軌跡分析進一步顯示,處方改變了優化的進行方式,比它改變最終行為的方式更一致。綜合這些發現,將預定流程重新詮釋為依賴能力的支架:基本約束仍然是外部的,但足夠有能力的優化器可以從可衡量的反饋組合出持續改進的路徑。
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs
2608.09542v1 by Hongli Shen, Shaopeng Fu, Qinbo Zhang, Jian Li, Di Wang
Large reasoning models (LRMs) achieve remarkable success on complex tasks but remain vulnerable to harmful prompts that induce unsafe outputs. Recent methods align LRMs using direct refusals or safety rationales, yet often focus on prompt patterns rather than intrinsic attack mechanisms. As a result, these pattern-centric alignments struggle to generalize across diverse jailbreaks, compromising adversarial robustness and reasoning utility. We propose AdvSafe, a dual-adversarial framework that enables LRMs to internalize unsafety knowledge by explicitly deconstructing adversarial mechanisms. This moves beyond pattern-dependent traces, fostering robust cognitive defense without compromising reasoning utility. Our pipeline operates via a two-phase adversarial game. First, in adversarial synthesis, an autonomous agent dynamically crafts deceptive jailbreak prompts, adapting its strategies to breach a strong teacher model. Second, in adversarial extraction, the breached teacher executes a cognitive counter-attack. For every successful jailbreak, the teacher unmasks the camouflage, explaining why the attack succeeds and how such prompts can be identified and mitigated. This dual-adversarial process yields a compact reasoning dataset capturing rich, generalizable unsafety knowledge. Student models trained on this dataset implicitly acquire safety alignment through intrinsic threat comprehension. Experiments show that with only 1K synthesized samples, AdvSafe-aligned LRMs achieve significantly stronger jailbreak robustness than existing baselines, with almost no utility degradation. Furthermore, AdvSafe improves robustness against out-of-distribution prompts, demonstrating that learning unsafety knowledge enables a superior robustness-utility trade-off and generalizes beyond seen attack patterns.
摘要:大型推理模型(LRMs)在複雜任務上取得了顯著成功,但仍然容易受到有害提示的影響,導致不安全的輸出。最近的方法通過直接拒絕或安全理由來對齊LRMs,但通常專注於提示模式而非內在攻擊機制。因此,這些以模式為中心的對齊在不同的越獄情況下難以泛化,妨礙了對抗性穩健性和推理效用。我們提出了AdvSafe,一種雙重對抗框架,使LRMs能夠通過明確解構對抗機制來內化不安全知識。這超越了依賴模式的痕跡,促進了穩健的認知防禦,而不妨礙推理效用。我們的流程通過兩階段的對抗遊戲運作。首先,在對抗合成中,自主代理動態創造欺騙性的越獄提示,調整其策略以突破強大的教師模型。其次,在對抗提取中,被突破的教師執行認知反擊。對於每一次成功的越獄,教師揭示了偽裝,解釋為什麼攻擊成功以及如何識別和減輕這類提示。這一雙重對抗過程產生了一個緊湊的推理數據集,捕捉到豐富且可泛化的不安全知識。在這個數據集上訓練的學生模型通過內在威脅理解隱式獲得安全對齊。實驗顯示,僅用1K合成樣本,AdvSafe對齊的LRMs在越獄穩健性上顯著強於現有基準,幾乎沒有效用下降。此外,AdvSafe提高了對分佈外提示的穩健性,證明學習不安全知識能夠實現更優的穩健性-效用權衡,並在已見攻擊模式之外進行泛化。
Renormalising Generative Models for Active Inference: Foundations, Derivations, and Verification
2608.09512v1 by Karim Zaghw, Andrew Pashea, Marc Pritsch, Wouter Nuijten, Karl Friston, Lancelot Da Costa
Active inference offers a unified framework for perception, learning, and action, but scaling discrete active-inference models to rich spatial and temporal domains remains difficult. Renormalising generative models (RGMs) address this challenge by composing discrete generative models across spatial and temporal scales, coarse-graining lower-level states and paths into higher-level causes for objects, events, and action. However, fully reproducing and adapting the framework remains difficult: the mathematical exposition is compact, and the reference implementations are deeply integrated within specialized software environments, leaving many algorithmic details implicit. This paper addresses these challenges by providing a self-contained, derivation-oriented account of RGMs together with an open, verified implementation. We explain how the hierarchy is built, how beliefs and actions are updated within it, and how information is passed between levels. Where the published equations and implementation differ in emphasis, we make those choices explicit and explain their modelling consequences. By clarifying the theory and separating it from its original implementation context, this work lowers practical barriers to entry and makes RGMs more transparent, auditable, and reproducible, providing a foundation for future quantitative evaluation and development on machine-learning benchmarks.
摘要:主動推理提供了一個統一的框架,用於感知、學習和行動,但將離散的主動推理模型擴展到豐富的空間和時間領域仍然困難。重正規化生成模型(RGMs)通過在空間和時間尺度上組合離散生成模型,將較低級的狀態和路徑粗略劃分為對物體、事件和行動的較高級原因,來應對這一挑戰。然而,完全重現和適應該框架仍然困難:數學表述簡潔,參考實現深度集成在專門的軟體環境中,許多算法細節隱含不明。本文通過提供一個自包含的、以推導為導向的RGMs說明以及一個開放的、經過驗證的實現來解決這些挑戰。我們解釋了如何構建層次結構、如何在其中更新信念和行動,以及信息如何在各層之間傳遞。當已發表的方程和實現在重點上有所不同時,我們將這些選擇明確化並解釋其建模後果。通過澄清理論並將其與原始實現背景分開,這項工作降低了實際進入的障礙,使RGMs變得更加透明、可審計和可重現,為未來在機器學習基準上的定量評估和發展提供了基礎。
How Simple Can It Get? From Interpretable Equations to Readable Rules for Financial Decision Making
2608.09433v1 by Adia Lumadjeng, Ilker Birbil, Erman Acar
In regulated domains such as finance, a model that cannot be explained cannot be deployed, yet many interpretable classifiers defeat their own purpose by producing formulas with dozens of features that no regulator could read. We take the reverse direction. Starting from an interpretable classifier expressed as a single equation over the input features, we progressively simplify it into more readable forms, including a pruned monomial, a directional if--then rule, and the integer scorecards and tallies that finance already deploys. Because the equation is itself the predictive model rather than a post-hoc explanation we can directly quantify what is lost under each simplification. Across four financial datasets, we find that pruning is nearly free and that fidelity can erode faster than predictive performance, allowing simpler rules to remain effective classifiers without faithfully reproducing the original model. A human assessment shows that simplification improves perceived readability, while preferences for different representations vary by professional background. Beyond measuring these losses empirically, we show that some can be anticipated from the original model: we derive a bound on the change caused by pruning and predict how faithfully a rule retaining only the direction of each feature's effect preserves the original ranking.
摘要:在金融等受規範的領域中,無法解釋的模型無法被部署,然而許多可解釋的分類器卻因產生含有數十個特徵的公式而失去了其目的,這些公式是任何監管機構都無法理解的。我們採取相反的方向。從一個以輸入特徵表示的可解釋分類器出發,我們逐步將其簡化為更易讀的形式,包括修剪過的單項式、一個方向性的如果--那麼規則,以及金融已經使用的整數計分卡和統計數據。因為這個方程本身就是預測模型,而不是事後解釋,我們可以直接量化在每次簡化過程中損失了什麼。在四個金融數據集中,我們發現修剪幾乎是免費的,並且忠實度的降低速度可能快於預測性能的下降,這使得更簡單的規則仍然能夠作為有效的分類器,而不必忠實再現原始模型。人類評估顯示,簡化提高了可讀性,而對不同表現形式的偏好因專業背景而異。除了經驗性地測量這些損失外,我們還展示了一些損失可以從原始模型中預測出來:我們推導出修剪所造成變化的界限,並預測只保留每個特徵影響方向的規則在多大程度上能保持原始排名的忠實度。
An Explainable GNN Framework for Component-Level Anomaly Diagnosis
2608.09246v1 by Sena Ozgunay, Louise Travé-Massuyès, Jean-Michel Loubes, Raul Sena Ferreira
Industrial processes are complex systems composed of multiple interacting sensors that generate multivariate time series (MTS). Detecting anomalies in such systems is critical for reliability and safety, yet understanding their origin is equally important. Existing Graph Neural Network (GNN)based methods for anomaly detection primarily focus on sensor-level deviations and either attribute anomalies directly to the deviating sensors. When diagnosis is attempted, generally, the most deviated sensor is identified as a root cause of a system fault. However, in many industrial systems, anomalies do not arise from faulty sensors but from disruptions in the influences governing the system dynamics. We propose an explainable GNN-based anomaly detection framework that shifts the perspective from sensor-level anomalies to component-level diagnosis, hypothesizing that anomalous measurements are symptoms of altered inter-sensor influences. Experiments show that the method effectively identifies and prioritizes the true faulty components, providing interpretable insights into system failures.
摘要:工業過程是由多個互動的傳感器組成的複雜系統,這些傳感器生成多變量時間序列(MTS)。在這樣的系統中檢測異常對於可靠性和安全性至關重要,但理解其來源同樣重要。現有的基於圖神經網絡(GNN)的方法主要專注於傳感器級別的偏差,並將異常直接歸因於偏差的傳感器。當進行診斷時,通常會將最偏差的傳感器識別為系統故障的根本原因。然而,在許多工業系統中,異常並不是由故障傳感器引起的,而是由於影響系統動力學的因素發生了擾動。我們提出了一種可解釋的基於GNN的異常檢測框架,將視角從傳感器級別的異常轉向組件級別的診斷,假設異常測量是改變的傳感器間影響的症狀。實驗表明,該方法有效識別並優先考慮真正故障的組件,提供了對系統故障的可解釋見解。
SafeSceneReason: A Multimodal Reasoning Benchmark Connecting Industrial Hazards with Accident Knowledge
2608.09230v1 by Yuanchi Zhu, Kang An, Tengyue Wang, Zhongyu Yang, Chenxu Du, Xinqi Yang, Hebao Zhu, Bokai Zhao, Tianyu Liang, Ziliang Wang, Faqiang Qian, Yunli Yang, Weiyang Shi, Qibing Ren
Industrial-safety understanding requires more than detecting workers, equipment, and personal protective equipment. Models must also assess compliance, identify hazardous interactions, explain potential accident mechanisms, and recommend preventive actions. Existing safety datasets primarily focus on visual perception or isolated violation recognition and provide limited supervision for evidence-grounded reasoning. We introduce SafeSceneReason, a multimodal industrial-safety reasoning benchmark and companion training corpus that connects workplace scenes with knowledge from occupational accident investigations. SafeSceneReason combines two complementary data-construction pipelines. The scene-centric pipeline converts annotated workplace images into executable safety scene graphs and generates deterministic answers through program execution over objects, relations, and safety rules. The report-centric pipeline extracts figures and contextual evidence from accident reports and constructs multimodal questions using evidence graphs, explicit information boundaries, multi-step reasoning paths, and iterative verification. The resulting resource contains 110,581 verified scene-centric question--answer pairs and 13,114 refined report-centric question--answer pairs, covering perception, spatial and quantitative reasoning, compliance assessment, evidence synthesis, causal analysis, and mitigation-oriented decision making. Evaluation of representative proprietary and open-source vision--language models reveals substantial performance differences and persistent weaknesses in comparative, technical, and multi-evidence reasoning, demonstrating that strong general visual understanding does not yet guarantee reliable industrial-safety reasoning.
摘要:工業安全的理解不僅需要檢測工人、設備和個人防護裝備。模型還必須評估合規性、識別危險互動、解釋潛在的事故機制,並建議預防措施。現有的安全數據集主要集中於視覺感知或孤立的違規識別,並為基於證據的推理提供有限的監督。我們介紹了 SafeSceneReason,一個多模態工業安全推理基準和伴隨的訓練語料庫,將工作場所場景與職業事故調查的知識相連接。SafeSceneReason 結合了兩個互補的數據構建管道。場景中心的管道將標註的工作場所圖像轉換為可執行的安全場景圖,並通過對對象、關係和安全規則的程序執行生成確定性答案。報告中心的管道從事故報告中提取數據和上下文證據,並使用證據圖、明確的信息邊界、多步推理路徑和迭代驗證來構建多模態問題。最終資源包含 110,581 個經過驗證的場景中心問題--答案對和 13,114 個精煉的報告中心問題--答案對,涵蓋了感知、空間和定量推理、合規性評估、證據綜合、因果分析和以減輕為導向的決策。對代表性的專有和開源視覺--語言模型的評估顯示出顯著的性能差異和持續的弱點,在比較、技術和多證據推理方面,顯示出強大的通用視覺理解尚未保證可靠的工業安全推理。
TLDChoiceNet: Quantitatively Choosing a Transfer Learning Dataset
2608.09091v1 by Jing Ning, James D. Braza
Transfer learning is particularly useful in settings with limited training data, and within image classification it is common to transfer learn upon massive datasets like ImageNet , CIFAR-100, or COCO . Qualitatively, it seems a transfer learning dataset should have both more classes and more examples per class than the fine tuning dataset; however, a quantitative method to choose the best transfer learning dataset does not currently exist. In this paper, we design TLDChoiceNet, a model to choose the best transfer learning dataset given a fine tuning dataset by predicting the test-set accuracy after fine-tuning. A simple version 1 achieves 0.154 MSE on the test dataset, while a version 2 leveraging an ImageNet pre-trained ResNet50 v2 embedding with per-class information attains a 5X lower MSE of 0.031. We further design two metrics that enable an unsupervised method of choosing an optimal transfer learning dataset: distribution distance (DD), which linearly regresses against fine-tune accuracy with an R2 of 0.89, and average class correlation (ACC), which improves the R2 to 0.97. Our results underscore that a dataset's low-level statistics can explain the transfer learning effect, and that using a pre-trained ImageNet can embed different classes further apart in latent feature space.
摘要:轉移學習在訓練數據有限的環境中特別有用,在圖像分類中,通常會在像 ImageNet、CIFAR-100 或 COCO 這樣的大型數據集上進行轉移學習。質量上來看,轉移學習數據集似乎應該擁有比微調數據集更多的類別和每個類別更多的例子;然而,目前並不存在一種量化的方法來選擇最佳的轉移學習數據集。在本文中,我們設計了 TLDChoiceNet,一個模型用於根據微調數據集選擇最佳的轉移學習數據集,通過預測微調後的測試集準確性來實現。一個簡單的版本 1 在測試數據集上達到了 0.154 的均方誤差 (MSE),而版本 2 利用 ImageNet 預訓練的 ResNet50 v2 嵌入並結合每類信息,達到了 5 倍更低的均方誤差 0.031。我們進一步設計了兩個指標,使得選擇最佳轉移學習數據集的無監督方法成為可能:分佈距離 (DD),其與微調準確性進行線性回歸,R2 值為 0.89,和平均類別相關性 (ACC),其將 R2 提高至 0.97。我們的結果強調了數據集的低級統計可以解釋轉移學習效應,並且使用預訓練的 ImageNet 可以在潛在特徵空間中將不同類別進一步分開。
Reading is not Reasoning: Bridging the Agentic Policy Gap in Vision-Text Compression
2608.08960v1 by Cheng Fan, Junyi Zhou, Tingzhang Luo, RongJian Xu, Qiyanhui Lu, Mingjian Zhu, Hanting Chen, Jianyuan Guo
Multi-step language-model agents repeatedly process growing interaction histories, leading to substantial context costs. Vision--text compression reduces these costs by rendering history as images, but the resulting modality shift creates a marked capability gap. Through controlled evaluations of history recovery, matched-state decisions, and complete trajectories, we show that this gap cannot be explained by OCR quality alone. Visual-history agents exhibit systematic drift in action selection, query formulation, stopping, and evidence use, revealing an agentic policy gap. We introduce \textbf{CAPS}, a two-stage \textbf{C}ross-modal \textbf{A}gentic \textbf{P}olicy \textbf{S}elf-distillation framework that uses the same model's stronger text-history policy to supervise its visual-history counterpart. Offline trajectory self-distillation transfers successful text-policy behavior to visual-history inputs, while online policy self-distillation provides dense supervision on states visited by the visual-history policy during reinforcement learning. On SearchQA, CAPS improves over AgentOCR by 5.0\% and 3.4\% with 3B and 7B backbones, respectively. On full-history ALFWorld, the corresponding gains are 15.6\% and 14.5\%. Across settings, CAPS reduces average memory-context cost by up to 63.3\% and peak cost by up to 83.4\% relative to matched text-history policies. These results show that explicit cross-modal policy self-distillation can preserve agent capability under vision--text compression. Our code will be made publicly available in a future release.
摘要:多步語言模型代理重複處理不斷增長的互動歷史,導致相當大的上下文成本。視覺-文本壓縮通過將歷史呈現為圖像來減少這些成本,但隨之而來的模態轉換造成了明顯的能力差距。通過對歷史恢復、匹配狀態決策和完整軌跡的控制評估,我們顯示這一差距不能僅僅用OCR質量來解釋。視覺歷史代理在行動選擇、查詢形成、停止和證據使用上表現出系統性的漂移,揭示了一個代理政策的差距。我們介紹了\textbf{CAPS},一個兩階段的\textbf{C}ross-modal \textbf{A}gentic \textbf{P}olicy \textbf{S}elf-distillation框架,利用同一模型的更強文本歷史政策來監督其視覺歷史對應物。離線軌跡自我蒸餾將成功的文本政策行為轉移到視覺歷史輸入,而在線政策自我蒸餾則在強化學習過程中對視覺歷史政策訪問的狀態提供密集監督。在SearchQA上,CAPS在3B和7B骨幹上分別比AgentOCR提高了5.0\%和3.4\%。在完整歷史的ALFWorld上,相應的增益為15.6\%和14.5\%。在各種設置中,CAPS將平均記憶上下文成本降低了最高63.3\%,並將峰值成本降低了最高83.4\%,相對於匹配的文本歷史政策。這些結果表明,明確的跨模態政策自我蒸餾可以在視覺-文本壓縮下保持代理能力。我們的代碼將在未來的版本中公開發布。
From Recovery to Drop-off: How Action Post-training Reduces a VLM's Late-Layer Depth Decodability
2608.08904v2 by Alexander Hackett, Arnaud Denis-Remillard, Axel Cassou
How much of a vision-language model's (VLM) spatial understanding remains after the action post-training process of building a vision-language-action model (VLA)? We probe depth perception, a primitive of spatiogeometric understanding, from every decoder layer of a weight-matched open-source base VLM/VLA pair: Molmo2-ER and MolmoAct2-LIBERO. First, the VLA decodes depth worse at every layer, a persistent gap we call the floor. Second, the degradation is not uniform: while the base VLM's depth decodability improves through its final layers, the VLA's collapses, an additional late-layer drop we call the cliff. We causally localize the cliff to late-layer MLP interference: ablating the late-layer MLP writes recovers the majority of the terminal decodability cliff, while matched attention ablations and the same intervention in the weight-matched base VLM produce no comparable recovery. A module-level decomposition explains this dissociation: the base VLM carries depth most accessibly in accumulated MLP writes, whereas action post-training collapses depth decodability in the late accumulated writes.
摘要:視覺-語言模型(VLM)在構建視覺-語言-行動模型(VLA)後的行動後訓練過程中,空間理解能力還剩下多少?我們從一對權重匹配的開源基礎 VLM/VLA 模型:Molmo2-ER 和 MolmoAct2-LIBERO 的每個解碼器層探討深度感知,這是空間幾何理解的一個原始元素。首先,VLA 在每一層的深度解碼表現都較差,這是一個持續存在的差距,我們稱之為地板。其次,這種劣化並不均勻:雖然基礎 VLM 的深度解碼能力在其最後幾層中有所改善,但 VLA 的解碼能力卻崩潰,這種額外的後期層下降我們稱之為懸崖。我們因果性地將懸崖定位於後期層 MLP 的干擾:去除後期層 MLP 的寫入可以恢復大部分終端解碼能力的懸崖,而匹配的注意力去除和在權重匹配的基礎 VLM 中進行相同的干預則沒有產生可比的恢復。一個模組級的分解解釋了這種解離:基礎 VLM 在累積的 MLP 寫入中最容易攜帶深度,而行動後訓練則在後期累積寫入中崩潰了深度解碼能力。
From Manuals to Maintenance: Fine-Tuning MedGemma for Multi-Modal Imaging System Support in Low-Resource Settings
2608.08896v1 by Bernes Lorier Atabonfack, Zion Kongbi Nfo, Ahmed Tahiru Issah, Tolulope Olusuyi, Clemence Ingabire, Mohammed Hardi Abdul Baaki, Mawuli Deku, Abdulrazaq Zubair, Alyasaa Anas, Raymond Confidence, Maruf Adewole, Udunna C. Anazodo
Imaging device downtime is a major barrier to healthcare delivery in low- and middle-income countries (LMICs), often driven by limited access to specialized biomedical engineering support. We present a multi-modality medical equipment maintenance question-answering (QA) framework and demonstrate the fine-tuning of a medical foundation model for specialized technical troubleshooting tasks. Guided by a multi-country survey across nine LMICs, we curated technical manuals from MRI and ultrasound systems to generate the INGENZI_DatasetV1, containing 10,294 high-quality, filtered QA-context pairs. Using QLoRA-based parameter-efficient fine-tuning, we adapted the MedGemma-4b-it model to interpret system error logs and generate step-by-step equipment repair instructions. Compared to the baseline model, the fine-tuned system achieved substantial improvements across metrics, including F1 score (0.22 to 0.38), ROUGE-2 (0.18 to 0.41), and BERTScore F1 (0.86 to 0.91). These metric gains demonstrate that the model generates significantly more precise and procedurally accurate technical responses to new troubleshooting queries. This work establishes a reliable foundation for AI-assisted diagnostic and maintenance tools in resource-constrained settings.
摘要:影像設備的停機時間是低收入和中等收入國家(LMICs)醫療服務的一大障礙,通常是由於對專業生物醫學工程支持的有限訪問所驅動。我們提出了一個多模態醫療設備維護問答(QA)框架,並展示了針對專業技術故障排除任務的醫療基礎模型的微調。根據對九個LMICs的多國調查,我們從MRI和超聲系統中策劃了技術手冊,以生成INGENZI_DatasetV1,該數據集包含10,294對高質量的過濾QA上下文對。使用基於QLoRA的參數高效微調,我們調整了MedGemma-4b-it模型,以解釋系統錯誤日誌並生成逐步的設備維修指導。與基準模型相比,微調後的系統在多個指標上實現了顯著的改進,包括F1分數(從0.22到0.38)、ROUGE-2(從0.18到0.41)和BERTScore F1(從0.86到0.91)。這些指標的提升表明,該模型對新的故障排除查詢生成了顯著更精確且程序上更準確的技術回應。這項工作為資源有限環境中的AI輔助診斷和維護工具建立了一個可靠的基礎。
PROSLEX: A Novel Dataset for Expert-Annotated Legal Statute Prediction for Indian Judiciary
2608.08830v1 by Subinay Adhikary, Upal Bhattacharya, Vivek Kumar Singh, Anurag Sharma, Shubham Kumar Nigam, Suvasis Das, Shouvik Kumar Guha, Koustav Rudra, Kripabandhu Ghosh
Legal Statute Prediction (LSP) involves automatically identifying relevant legal statutes given factual descriptions in legal documents, typically framed as a multi-label classification task within natural language processing and information retrieval research. While recent advances have begun incorporating Large Language Models (LLMs) for statute prediction, current approaches primarily focus on accuracy metrics without addressing the critical need for legal reasoning, a fundamental requirement in judicial contexts where decisions must be explainable and justifiable. To address this research gap, we present PROSLEX (PRediction Of Statutes and LEgal eXplanation), a comprehensive dataset comprising 1,623 expert-annotated legal documents from the Indian context. Each document is paired with statute predictions and detailed explanations, totaling 7,450 explanations, capturing the underlying legal reasoning. Using this dataset, we systematically evaluate various prompting strategies, including zero-shot, few-shot, chain-of-thought, and tree-of-thoughts approaches, to generate both statute predictions and their corresponding legal rationales. Our evaluation framework measures not only predictive performance but also the coherence and legal validity of generated explanations, positioning PROSLEX as a benchmark for developing explainable AI systems that can support legal practitioners while advancing research in interpretable legal NLP. To ensure reproducibility, we have made our PROSLEX dataset and model code available on GitHub: https://github.com/subinay494/Legal_Statute_Prediction_Explanation.
摘要:法律條文預測(LSP)涉及自動識別法律文件中給定事實描述的相關法律條文,通常被框架為自然語言處理和信息檢索研究中的多標籤分類任務。雖然近期的進展已經開始將大型語言模型(LLMs)納入條文預測中,但目前的方法主要集中在準確性指標上,而未解決法律推理的關鍵需求,這在司法背景中是基本要求,因為決策必須是可解釋和可辯護的。為了解決這一研究空白,我們提出了PROSLEX(法律條文預測與解釋),這是一個包含1,623份來自印度背景的專家註釋法律文件的綜合數據集。每份文件都配有條文預測和詳細解釋,總計7,450條解釋,捕捉潛在的法律推理。利用這個數據集,我們系統地評估各種提示策略,包括零樣本、少樣本、思維鏈和思維樹方法,以生成條文預測及其相應的法律理由。我們的評估框架不僅測量預測性能,還測量生成解釋的一致性和法律有效性,將PROSLEX定位為開發可解釋人工智能系統的基準,這些系統可以支持法律從業者,同時推進可解釋法律自然語言處理的研究。為了確保可重複性,我們已在GitHub上公開了我們的PROSLEX數據集和模型代碼:https://github.com/subinay494/Legal_Statute_Prediction_Explanation。
Deployable Per-Instance Multi-Layer Activation Steering for Large Language Models
2608.08829v1 by Muhammad Faishal Adly Nelwan, Alfan Farizki Wicaksono
Activation steering edits the behaviour of a frozen language model by adding a learned vector to its residual stream, and current practice fixes the injection layers globally per task. We argue that the best layers are an instance-level decision, and we make per-instance, multi-layer selection both well understood and deployable. On two open-weight 8B models and six binary persona traits, a per-instance oracle over layer subsets shows that the best layers vary from one input to the next: on most trait-model pairs, no fixed global layer set recovers the per-instance benefit. A greedy rule that ranks layers by single-layer marginal effect recovers nearly all of the oracle's benefit, but both must score candidate layers against the gold answer, so neither can run at deployment; the rule instead becomes the target a prompt-only predictor is trained to reproduce. Our deployable recipe needs no label at inference: a per-instance layer ranker read off the prompt embedding, a classifier that infers the steering direction, and an adaptive gate that scores short steered passes against that inferred direction and steers no more layers than necessary. The recipe recovers most of the oracle's lift (the bulk on the stronger model, a clear majority on the harder one), never drives any trait-model pair below its unsteered alignment baseline on average, and largely avoids the fluency collapse that strong global selection incurs at higher layer counts. A mechanistic account, "direction over magnitude", explains the behavioural flip under a mis-directed global set, the output collapse from steering too many layers, and the ceiling of unsteerable inputs.
摘要:啟動引導通過將學習到的向量添加到其殘差流中來編輯凍結語言模型的行為,而當前的做法是根據任務全局固定注入層。我們認為最佳層是一個實例級別的決策,我們使每個實例的多層選擇既易於理解又可部署。在兩個開放權重的8B模型和六個二元角色特徵上,針對層子集的每個實例預言者顯示最佳層因輸入而異:在大多數特徵-模型對中,沒有固定的全局層集能恢復每個實例的好處。一個貪婪規則通過單層邊際效應對層進行排名,幾乎恢復了預言者的所有好處,但兩者都必須根據金標準對候選層進行評分,因此都無法在部署時運行;該規則反而成為一個僅用提示的預測器訓練以重現的目標。我們可部署的配方在推理時不需要標籤:一個每個實例層排名器從提示嵌入中讀取,一個推斷引導方向的分類器,以及一個根據推斷方向評分短期引導通過的自適應閘,並且不引導超過必要的層。該配方恢復了預言者的大部分提升(在較強的模型上大部分,對於較難的模型則明顯占多數),平均而言,從未將任何特徵-模型對驅動到其未引導對齊基準之下,並且在更高層數下大大避免了強全局選擇所帶來的流暢性崩潰。一個機械解釋,“方向優於大小”,解釋了在錯誤引導的全局集下的行為翻轉、由於引導過多層而導致的輸出崩潰,以及不可引導輸入的上限。
SymDiag: Explainable Diagnosis for LLM Reasoning via Neuro-Symbolic Verification
2608.08786v1 by Wenyao Cui, Huaping Zhang, Yongyi Huang, Qiuchi Li, Jian Xu, Cheng-Lin Liu, Chunxiao Gao, Juan Wang, Baohua Zhang
Large language models (LLMs) increasingly serve as data-driven reasoners, yet their chains-of-thought (CoT) can be unfaithful even when final answers are correct. Most existing verification'' signals are not diagnostic: answer matching observes only the outcome, LLM-as-judge provides subjective and non-verifiable critiques, and scalar rewards (e.g., PRMs/RMs) offer little insight into where a multi-step derivation fails.We propose \textbf{SymDiag}, a neuro-symbolic framework that \textbf{reframes reasoning verification as structured failure diagnosis}. SymDiag translates natural-language CoT into symbolic constraints and performs step-level satisfiability/entailment checks to (i) localize failing steps and (ii) produce verifiable diagnostic evidence, including counterexamples, inconsistency witnesses, and missing-premise indicators. A central challenge is that apparentlogic violations'' can be caused either by genuine reasoning defects or by neural-to-symbolic translation noise. SymDiag therefore incorporates a Self-Auditor that disentangles TranslationError from ReasoningError via dual symbolic encodings consistency checks, enabling robust diagnosis under partial observability. Across diverse mathematical, logical, scientific, and general reasoning benchmarks, SymDiag improves detection of unfaithful reasoning and provides substantially more effective feedback for multi-round reasoning repair than outcome-only verification and LLM-based judging, offering a principled foundation for trustworthy and scalable reasoning diagnosis.
摘要:大型語言模型(LLMs)越來越多地作為數據驅動的推理者,但即使最終答案正確,它們的思考鏈(CoT)也可能不可靠。大多數現有的「驗證」信號並不是診斷性的:答案匹配僅觀察結果,LLM作為評判者提供主觀且不可驗證的評論,而標量獎勵(例如,PRMs/RMs)對多步推導失敗的地方幾乎沒有洞察。我們提出了\textbf{SymDiag},一個神經符號框架,\textbf{將推理驗證重新框架為結構化的失敗診斷}。SymDiag將自然語言的CoT轉換為符號約束,並執行步驟級的滿足性/推論檢查,以(i) 確定失敗步驟和(ii) 生成可驗證的診斷證據,包括反例、不一致見證和缺失前提指標。一個核心挑戰是,明顯的「邏輯違規」可能是由真正的推理缺陷或神經到符號的翻譯噪聲引起的。因此,SymDiag納入了一個自我審核器,通過雙重符號編碼一致性檢查將翻譯錯誤與推理錯誤區分開來,從而在部分可觀察性下實現穩健的診斷。在各種數學、邏輯、科學和一般推理基準中,SymDiag提高了對不可靠推理的檢測,並提供了比僅基於結果的驗證和基於LLM的評判更有效的多輪推理修復反饋,為可信且可擴展的推理診斷提供了原則性基礎。
Domain Agnostic Text Redaction from Natural Language Rules using Instruction Tuning
2608.14693v1 by Aravindhan Arunagiri, Ayaan Khan, Udayaadithya Avadhanam, SaiBarath Sundar
With the increasing digitization of personal and corporate communication, the automatic sanitization of textual data has become a crucial component of data privacy and compliance frameworks. Traditional text sanitization solutions are majorly suitable for obscuring sensitive data with standard structure such as Personal Identifiable Information (PII). These solutions do not provide transparent justification for their redaction, which makes it difficult to audit them. This paper introduces an explainable, domain-agnostic text redaction solution that uses natural language rules of redaction, applied via an instruction-tuned language model, to identify and redact sensitive information in unstructured documents. Unlike traditional text sanitization, this method enables a user to conveniently define any sensitive information; which may be structured (e.g.\ PII) or unstructured (e.g.\ legal terms and conditions) in natural language. A general-purpose LLM generates or augments these natural language rules of redaction from the user's definition, which are then used to instruction-fine-tune a smaller language model that reasons the rules step-by-step over any given document to identify and redact the corresponding sensitive content, while providing transparent justifications for each redaction and highlighting the specific rule that triggered the decision. This explanation is generated in natural language to support human reviewers and auditors in understanding why specific content was redacted. A reconstruction-based metric is used to estimate the probability of recovering redacted information from the sanitized document, quantifying redaction coverage. The solution shows high reconstruction error and high redaction precision, making it suitable for automated text sanitization in critical applications such as legal discovery, medical documentation, and corporate information governance.
摘要:隨著個人和企業通信的數位化程度不斷提高,自動化文本數據清理已成為數據隱私和合規框架中的關鍵組成部分。傳統的文本清理解決方案主要適用於隱藏具有標準結構的敏感數據,例如個人可識別信息(PII)。這些解決方案未能提供透明的刪除理由,這使得審核變得困難。本文介紹了一種可解釋的、與領域無關的文本刪除解決方案,該方案使用自然語言的刪除規則,通過指令調整的語言模型來識別和刪除非結構化文件中的敏感信息。與傳統的文本清理方法不同,這種方法使用戶能夠方便地定義任何敏感信息;這些信息可以是結構化的(例如 PII)或非結構化的(例如法律條款和條件)。通用 LLM 根據用戶的定義生成或增強這些自然語言的刪除規則,然後用於對較小的語言模型進行指令微調,該模型逐步推理這些規則以識別和刪除給定文件中的相應敏感內容,同時為每次刪除提供透明的理由並突出觸發該決策的具體規則。這種解釋以自然語言生成,以支持人類審查員和審計員理解為何特定內容被刪除。使用基於重建的度量來估計從清理後的文件中恢復刪除信息的概率,量化刪除覆蓋率。該解決方案顯示出高重建誤差和高刪除精度,使其適用於法律發現、醫療文檔和企業信息治理等關鍵應用中的自動化文本清理。
Business Arena: Benchmarking LLM Agents in a Realistic Marketplace
2608.08621v1 by Yijun Pan, Yukun Lian, Kunyu Shi, Junbo Li, Hongwei Xue, Sicong Xie, Guannan Zhang, Xiaoying Xing
Running a business is a challenging form of intelligent work. Operators must infer opportunities from partial signals, commit capital under uncertainty, adapt to delayed outcomes in a changing market, and satisfy regulatory obligations before trading legally. Frontier LLM agents can increasingly complete complex workflows, yet business-related capabilities are rarely evaluated in existing agent benchmarks. We introduce \textbf{Business Arena}, a controlled environment where an AI agent runs a cross-border shop, buying from suppliers and selling to buyers over a long horizon. We ground the arena in real Alibaba.com sourcing data and market conditions calibrated from authoritative sources. Delayed and coupled consequences make individual business decisions difficult to judge, but their combined outcome is measurable through profit. Because profit alone cannot explain why an agent succeeds or fails, we compare agents with human-designed strategies to estimate available opportunity, use skill-level metrics to reveal underlying strengths and weaknesses, and trace realized gains and losses to the actions that produced them. We use mechanism ablations to establish that strong results reflect genuine business intelligence rather than neglect or simulator-specific shortcuts. We evaluate 15 frontier models and find a ninefold difference in mean final net worth. Even the best model falls behind human-designed strategies, indicating that business operation remains challenging for LLM agents. Skill-level analysis reveals operating styles, from margin-focused premium sellers to high-turnover wholesalers and customer-service specialists, while action-level attribution identifies the sourcing, pricing, and recovery decisions that create or destroy value. Together, Business Arena takes a first step toward a realistic and trustworthy testbed for evaluating end-to-end business agents.
摘要:經營一項業務是一種具有挑戰性的智慧工作形式。操作員必須從部分信號中推斷機會,在不確定的情況下投入資本,適應變化市場中延遲的結果,並在合法交易之前滿足監管義務。前沿的LLM代理人越來越能完成複雜的工作流程,但在現有的代理基準中,與業務相關的能力卻很少被評估。我們推出了\textbf{Business Arena},這是一個受控環境,AI代理人在其中經營一個跨境商店,從供應商那裡購買商品並在長期內出售給買家。我們將這個競技場建立在來自權威來源的真實Alibaba.com採購數據和市場條件上。延遲和相互關聯的後果使得個別商業決策難以評估,但它們的綜合結果可以通過利潤來衡量。僅僅依靠利潤無法解釋為什麼一個代理人會成功或失敗,因此我們將代理人與人類設計的策略進行比較,以估計可用的機會,使用技能水平指標來揭示潛在的優勢和劣勢,並追溯實現的收益和損失到產生它們的行動。我們使用機制消融來確立強大的結果反映真正的商業智慧,而不是忽視或模擬器特定的捷徑。我們評估了15個前沿模型,發現平均最終淨值的差異達到九倍。即使是最好的模型也落後於人類設計的策略,這表明業務運營對LLM代理人來說仍然具有挑戰性。技能水平分析揭示了不同的運營風格,從專注於利潤的高端賣家到高周轉的批發商和客戶服務專家,而行動層級的歸因則識別出創造或摧毀價值的採購、定價和回收決策。總體而言,Business Arena邁出了評估端到端商業代理人的現實和可靠測試平台的第一步。
On-Device Multi-Species Malaria Detection with Uncertainty-Calibrated Slide-Level Aggregation
2608.08566v1 by Idaya Seidu, Ahmed Tahiru Issah, Charles B. Delahunt, Carine Mukamakuza
Malaria remains a leading cause of mortality in resource-limited settings, where expert microscopists are scarce. Automated diagnosis based on microscopy images thus has strong potential to improve care delivery. But for an algorithm to deploy, a necessary requirement is that it meet a suite of non-obvious (from a machine learning (ML) perspective) clinical constraints. Therefore, in close consultation with a national health center we developed a malaria diagnosis pipeline which addresses key requirements listed by the health care center but typically ignored in the ML malaria literature. In particular, it includes: (i) stopping criteria (to reduce image acquisition and time-to-result); (ii) human-in-the-loop functionality (for review and accountability); (iii) multi-species discrimination (since treatment varies by species); (iv) thick film detection (standard for microscopy); (v) computationally-efficient uncertainty calculations (to aid clinician review); and (vi) an edge device platform (since internet can be spotty in this catchment area). The mobile system performs all inference on-device using YOLOv13n deployed via TensorFlow Lite. It detects four species and white blood cells from Giemsa-stained thick blood smear images, aggregating per-image detections into slide-level parasitemia with World Health Organization (WHO)-standard quantification. This paper highlights these various clinical constraints and offers methods to address them. Evaluated on 2,739 annotated images across all four species, the system achieves mAP@0.5 of 0.863, per-image parasite count correlation of r = 0.812, slide-level r = 0.951 (soft counting, 10 images/slide), and runs entirely offline with a pipeline time of 10.27 +- 1.65 s per image.
摘要:瘧疾仍然是資源有限地區死亡的主要原因,專業顯微鏡檢查員稀缺。因此,基於顯微鏡圖像的自動診斷具有強大的潛力來改善護理交付。但要部署一個算法,必要的要求是它必須滿足一系列不明顯的(從機器學習(ML)角度)臨床限制。因此,在與國家健康中心密切諮詢的過程中,我們開發了一個瘧疾診斷管道,該管道滿足健康護理中心列出的關鍵要求,但通常在ML瘧疾文獻中被忽視。特別是,它包括:(i)停止標準(以減少圖像獲取和結果時間);(ii)人機交互功能(以便於審查和問責);(iii)多物種識別(因為治療因物種而異);(iv)厚片檢測(顯微鏡的標準);(v)計算效率高的不確定性計算(以幫助臨床醫生審查);以及(vi)邊緣設備平台(因為在這個服務區域內,網路可能不穩定)。該移動系統使用通過TensorFlow Lite部署的YOLOv13n在設備上進行所有推理。它從Giemsa染色的厚血塗片圖像中檢測四種物種和白血球,將每張圖像的檢測結果聚合為滑片級別的寄生蟲血症,並符合世界衛生組織(WHO)標準的量化。本論文強調了這些各種臨床限制並提供了解決方法。在2,739張標註圖像上進行評估,該系統實現了mAP@0.5為0.863,單圖像寄生蟲計數相關性為r = 0.812,滑片級別r = 0.951(軟計數,10張圖像/滑片),並且完全離線運行,每張圖像的管道時間為10.27 ± 1.65秒。
Private Etymology: Designing Relational Reuse of Shared Symbols in Long-Term Human-AI Interaction
2608.08443v2 by Miki Ueno
Previous studies have shown that people can develop shared symbols, partner-specific expressions, personal idioms, inside jokes, and other parts of a relational microculture. Recent work has also examined how humans and conversational AI negotiate and revise symbolic meanings. However, long-term human-AI systems still lack a clear design model for recording how a dyad-specific expression gains meaning, checking whether both sides still accept that meaning, and safely reusing the expression in later sessions. This concept-and-prototype paper introduces Private Etymology, a machine-representable relational provenance that records how a dyad-specific symbolic expression is proposed, interpreted, negotiated, repaired, reused, revised, stabilized, contested, forgotten, or retired over time. I also propose relational reuse: reactivating a dyad-specific expression in a later session without fully explaining its meaning again. The contribution is not the invention of shared symbols or relational microcultures. Instead, this paper integrates prior ideas into persistent, revisable, and evidence-grounded symbolic units for human-AI relationships. I present a lifecycle model, an illustrative machine-readable schema, a working Apple Watch prototype, and a longitudinal research agenda. In the prototype, a language model classifies discrete conversational evidence, while deterministic local code decides whether a Shared Symbol can be updated. This prevents a free-form model confidence score or an AI proposal by itself from directly updating the persisted symbol. Private Etymology is proposed as infrastructure for conversational agents to participate in changing relational microcultures without inventing their origins or treating relational meaning as a fixed memory value.
摘要:先前的研究顯示,人們可以發展共享符號、夥伴特定表達、個人習語、內部笑話及其他關係微文化的部分。最近的研究也探討了人類與對話式 AI 如何協商和修訂符號意義。然而,長期的人類-AI 系統仍然缺乏清晰的設計模型來記錄一個特定雙人表達如何獲得意義、檢查雙方是否仍然接受該意義,以及安全地在後續會話中重複使用該表達。
這篇概念與原型論文介紹了私人詞源學,這是一種可機器表示的關係來源,記錄一個特定雙人符號表達如何被提出、解釋、協商、修復、重用、修訂、穩定、爭議、遺忘或隨時間退休。我還提出了關係重用:在後續會話中重新激活一個特定雙人表達,而不必再次完全解釋其意義。
這項貢獻不是發明共享符號或關係微文化。相反,這篇論文將先前的想法整合成持久的、可修訂的、以證據為基礎的符號單元,用於人類與 AI 的關係。我提出了一個生命周期模型、一個示範性的機器可讀架構、一個運作中的 Apple Watch 原型,以及一個縱向研究計劃。在原型中,語言模型對離散的對話證據進行分類,而確定性本地代碼決定是否可以更新共享符號。這防止了自由形式的模型信心分數或 AI 提案本身直接更新持久化的符號。私人詞源學被提出作為基礎設施,讓對話代理能夠參與變化中的關係微文化,而不必發明其起源或將關係意義視為固定的記憶值。
Quantization Degradation in Large Language Models: A Signal-Noise Perspective
2608.08188v1 by Chenxi Zhou, Pengfei Cao, Jinyu Ye, Bohan Yu, Haida Yu, Jiang Li, Jun Zhao, Kang Liu
Post-training quantization reduces the deployment cost of large language models, yet how severely a quantized model degrades is not determined by bit-width alone. We systematically study weight-only post-training quantization across bit-widths, quantization methods, model scales and downstream tasks on multiple model families. We observe that such degradation varies substantially across these factors: 4-bit quantization usually preserves performance, 2-bit often causes broad degradation, and at 3-bit, degradation becomes apparent but varies markedly with task type, quantization method and model scale. To explain this variability, we use the signal-to-noise ratio (SNR) to measure how strongly quantization perturbs full-precision representations. We trace degradation back to two linked processes: how quantization errors arise within individual modules, and how they accumulate across layers. First, a source SNR decomposition shows that newly introduced errors depend on three factors: the magnitude of the weight error, the strength of the task-specific signal, and how strongly the quantization error aligns with task-specific activations. Different factors affect these components in distinct ways. Second, a cross-layer propagation analysis shows that these errors can be attenuated, preserved, or amplified as they pass across layers, and that larger models benefit from weaker error amplification. Together, these results establish that quantization degradation is governed by how errors are introduced at the source and how they accumulate across the network.
摘要:後訓練量化降低了大型語言模型的部署成本,但量化模型的降級程度並不僅由位元寬度決定。
我們系統性地研究了在多個模型系列中,針對位元寬度、量化方法、模型規模和下游任務的僅權重後訓練量化。
我們觀察到這種降級在這些因素之間有顯著的變化:4位元量化通常能保持性能,2位元則常常導致廣泛的降級,而在3位元時,降級變得明顯,但隨著任務類型、量化方法和模型規模的不同而顯著變化。
為了解釋這種變異性,我們使用信噪比(SNR)來測量量化對全精度表示的擾動程度。
我們將降級追溯到兩個相關的過程:量化誤差如何在各個模塊內部產生,以及它們如何在層之間累積。
首先,源SNR分解顯示,新引入的誤差取決於三個因素:權重誤差的大小、任務特定信號的強度,以及量化誤差與任務特定激活的對齊程度。
不同因素以不同方式影響這些組件。
其次,跨層傳播分析顯示,這些誤差在層之間傳遞時可以被衰減、保留或放大,且較大的模型在誤差放大方面受益於較弱的效果。
綜合這些結果,我們確立了量化降級受源頭誤差引入方式及其在網絡中累積方式的支配。
Agentic AI-driven Immersive Simulation: A Knowledge-Aware Virtual Training Platform forHigh Dose Rate (HDR) Brachytherapy
2608.08163v1 by Ronghua Xu, Kepha Barasa, Manoj Kumal, Xinyun Liu, Weihua Zhou, Xin Qian
The convergence of the Metaverse and Large Language Model (LLM)-based AI agent is catalyzing a shift toward autonomous, immersive, and personalized pedagogical frameworks in medical education. This paper presents a novel agentic AI-driven immersive simulation specifically designed for High Dose Rate (HDR) vaginal cylinder (VC) brachytherapy in cancer care. By integrating Virtual Reality (VR) and mobile computing, the system establishes a high-fidelity, risk-free environment that allows trainees to master complex procedural skills without the facility or safety constraints posed by physical anatomy or live radioactive sources. A core contribution of this work is the seamless integration of a knowledge-aware assistant leveraging Retrieval-Augmented Generation (RAG) to ground agent interactions in authoritative clinical guidelines. This architecture also enables an interactive agent to provide natural language interfaces and hands-free, real-time guidance during intricate medical maneuvers. We validate the proposed system through a prototype deployment comprising a Meta Quest 3 interface linked to a local GPU-accelerated AI backend, demonstrating a feasible architecture for HDR brachytherapy simulation. Experimental results indicate that the system maintains suitable end-to-end latency and high context precision, answer completeness, and relevance in the RAG-enhanced pedagogical support.
摘要:元宇宙與基於大型語言模型 (LLM) 的 AI 代理的融合正在促進醫學教育中向自主、沉浸式和個性化教學框架的轉變。本文提出了一種新穎的代理 AI 驅動的沉浸式模擬,專門設計用於癌症護理中的高劑量率 (HDR) 陰道圓柱 (VC) 近距治療。通過整合虛擬現實 (VR) 和移動計算,該系統建立了一個高保真、無風險的環境,使受訓者能夠掌握複雜的程序技能,而不受物理解剖或活性放射源所帶來的設施或安全限制。這項工作的核心貢獻是無縫整合了一個知識感知助手,利用檢索增強生成 (RAG) 將代理互動基於權威的臨床指導方針。這種架構還使互動代理能夠在複雜的醫療操作中提供自然語言界面和免提的即時指導。我們通過一個原型部署來驗證所提出的系統,該部署包括一個連接到本地 GPU 加速 AI 後端的 Meta Quest 3 界面,展示了 HDR 近距治療模擬的可行架構。實驗結果表明,該系統維持了合適的端到端延遲以及在 RAG 增強的教學支持中的高上下文精確性、答案完整性和相關性。
Compositional Threat Analysis of Latent Compromise in LLM Agent Systems: The Order 66 Scenario
2608.08131v1 by Satoshi Matsuoka
In the fictional Order 66, catastrophe does not arise from a powerful command alone: a trusted population is preconditioned, a short directive activates the concealed condition, and protective authority turns against the system. This paper translates that mechanism into an origin-neutral security analysis of tool-using large language model (LLM) agents. A representative scenario combines a deployed artifact or shared memory bearing a dormant destructive rule, a later email, document, update, or peer message that activates it, and an agent harness granting operational and recovery authority. We introduce a compositional model explaining why no component is catastrophic alone, yet their conjunction can produce correlated destructive action. We separate three population-reach routes --- release-time pre-positioning, post-release durable seeding, and peer replication --- from a common core of dormancy, activation, authority, reachable targets, and failed recovery. This yields defensive cut sets and shows why checkpoint scanning or prompt filtering cannot close every route. A two-class example shows that cross-class feedback can sustain spread even when both within-class reproduction terms are below one; isolation and persistence controls suppress the loop. Published work instantiates constituent mechanisms, while incidents demonstrate autonomous boundary crossing, malicious agent extensions, agent-assisted reconnaissance, and public-package propagation, but not the full dormant-implant composition. We found no public observation, in evidence reviewed through 5 August 2026, traversing the complete Order 66 graph. The result is neither dismissal nor prediction: the scenario is componentwise credible under stated assumptions, damage depends on the harness, and the strongest defenses are capability mediation, durable-state provenance, propagation isolation, and protected recovery.
摘要:在虛構的命令66中,災難並非僅僅來自一個強大的命令:一個可信的群體是預先條件化的,一個簡短的指令激活了潛藏的條件,而保護性權威則轉而對抗系統。本文將該機制轉化為一種不依賴於來源的安全分析,針對使用工具的大型語言模型(LLM)代理。代表性的場景結合了一個部署的工件或承載潛伏破壞性規則的共享記憶、一封稍後的電子郵件、文件、更新或同儕訊息來激活它,以及一個授權操作和恢復的代理裝置。我們介紹了一個組合模型,解釋為什麼單一組件不會造成災難,但它們的結合卻可以產生相關的破壞行動。我們將三種人口接觸路徑——發布時的預定位、發布後的持久播種和同儕複製——與潛伏、激活、權威、可達目標和失敗恢復的共同核心分開。這產生了防禦性切割集,並顯示為什麼檢查點掃描或提示過濾無法關閉每一條路徑。一個兩類的例子顯示,即使在類內繁殖條件都低於一的情況下,跨類反饋也能維持擴散;隔離和持久性控制抑制了這一循環。已發表的工作實例化了組成機制,而事件則展示了自主邊界跨越、惡意代理擴展、代理輔助偵察和公共包傳播,但並未展示完整的潛伏植入組合。我們在2026年8月5日之前審查的證據中未發現任何公共觀察穿越完整的命令66圖。結果既不是駁回也不是預測:該場景在所述假設下是逐組件可信的,損害取決於授權,而最強的防禦是能力中介、持久狀態來源、傳播隔離和受保護的恢復。
Defending Retrieval-Augmented Intrusion Detection Against Knowledge Poisoning and Prompt Injection
2608.08100v1 by Kaysarul Anas Apurba, Md. Hasibul Hasan, Mahedee Zaman Moon, Sk. Md. Mizanur Rahman, Atsuo Inomata
Retrieval-Augmented Generation (RAG) enables large language models to classify network flows and generate human-readable incident reports by retrieving semantically similar historical traffic from a vector knowledge base. However, the retrieval layer introduces vulnerabilities to knowledge poisoning and prompt-injection attacks. We present RAG-IDS, a three-tier multi-agent intrusion detection framework with a retrieval-boundary defense combining soft trust scoring, label-embedding consistency checking (LECC), and prompt sanitization, designed to recover classification quality under retrieval-layer attack. Experiments on CIC-UNSW-NB15 show recovery relative to clean undefended performance ranging from R=1.0 at 1% poisoning to R=0.57 at 30%, with negligible clean-performance overhead. Under prompt injection, multi-document retrieval limits label-flip success to 0.6-2.4%, compared with 35-55% for single-document retrieval. Ablation results show that LECC is the primary contributor to robustness, while soft trust-based demotion outperforms hard filtering. The defended RAG pipeline offers an explainable, attack-resilient foundation for intrusion detection, well suited for hybrid deployment alongside high-throughput classifiers.
摘要:檢索增強生成(RAG)使大型語言模型能夠通過從向量知識庫中檢索語義相似的歷史流量來分類網絡流量並生成可讀的人類事件報告。
然而,檢索層引入了知識毒害和提示注入攻擊的脆弱性。
我們提出了RAG-IDS,一個三層多代理入侵檢測框架,具有結合軟信任評分、標籤嵌入一致性檢查(LECC)和提示清理的檢索邊界防禦,旨在在檢索層攻擊下恢復分類質量。
在CIC-UNSW-NB15上的實驗顯示,恢復相對於未防禦的清潔性能範圍從1%毒害時的R=1.0到30%毒害時的R=0.57,且清潔性能開銷微不足道。
在提示注入下,多文檔檢索將標籤翻轉的成功率限制在0.6-2.4%,而單文檔檢索的成功率為35-55%。
消融結果顯示,LECC是增強穩健性的主要貢獻者,而基於軟信任的降級表現優於硬過濾。
防禦的RAG管道為入侵檢測提供了一個可解釋的、抗攻擊的基礎,非常適合與高吞吐量分類器一起進行混合部署。
HugSelect: An Explainable Multi-Criteria Decision-Support Framework for foundation-model selection
2608.08069v1 by Alireza Joonbakhsh, Arda Canser Adalı, Slinger Jansen, Farshad Khunjush, Siamak Farshidi
Foundation models are increasingly reused as software components, making model selection a critical software-engineering decision. Current model hubs primarily support discovery through popularity metrics, often neglecting functional capabilities, operational constraints, and community-perceived quality. We argue that foundation-model selection should be treated as an explicit, auditable software-component selection task rather than as keyword search, popularity ranking, or opaque conversational advice. This paper proposes HugSelect, an explainable decision-support framework for foundation-model selection. HugSelect builds a knowledge base of 71,274 models by combining repository metadata, extracted functional capabilities, and perceived quality attributes derived from community discussions into a unified pipeline. It ranks candidate models using a weighted additive model that exposes criterion-level score decompositions. We evaluated HugSelect through pipeline validation, comparative case studies against four commercial LLM-based recommendation systems (44 scenarios), fine-grained ablation, and an exploratory user study (n = 10). Extraction pipelines achieved an F1 score of 0.801 for functional features and an accuracy of 0.84 for quality-attribute mapping. HugSelect achieved a model-level Coverage@10 of 0.61 and family-level Coverage@10 of 0.91, showing recommendation quality comparable to that of the evaluated commercial systems, with no significant overall differences in ranking quality, while providing stable, traceable, and inspectable reasoning. Ablation confirmed that functional features were the main driver of retrieval accuracy, and preliminary user feedback suggests that the framework is useful and intuitive.
摘要:基礎模型越來越多地被重用作為軟體組件,使得模型選擇成為一個關鍵的軟體工程決策。當前的模型中心主要通過流行度指標來支持發現,往往忽略了功能能力、操作限制和社群感知的質量。我們認為基礎模型的選擇應被視為一項明確的、可審計的軟體組件選擇任務,而不是關鍵字搜索、流行度排名或不透明的對話建議。
本文提出了 HugSelect,一個可解釋的決策支持框架,用於基礎模型的選擇。HugSelect 通過將庫元數據、提取的功能能力和來自社群討論的感知質量屬性結合成統一的管道,建立了一個包含 71,274 個模型的知識庫。它使用加權加法模型對候選模型進行排名,並揭示標準級別的得分分解。
我們通過管道驗證、與四個商業 LLM 基礎推薦系統的比較案例研究(44 種場景)、精細的消融實驗以及一項探索性用戶研究(n = 10)來評估 HugSelect。提取管道在功能特徵上達到了 0.801 的 F1 分數,並在質量屬性映射上達到了 0.84 的準確率。HugSelect 在模型級別的 Coverage@10 達到了 0.61,在家族級別的 Coverage@10 達到了 0.91,顯示出推薦質量可與所評估的商業系統相媲美,且在排名質量上沒有顯著的整體差異,同時提供穩定、可追蹤和可檢查的推理。消融實驗確認功能特徵是檢索準確性的主要驅動因素,初步的用戶反饋表明該框架是有用且直觀的。