Knowledge Graphs
Knowledge Graphs
| Publish Date | Title | Authors | Homepage | Code |
|---|---|---|---|---|
| 2026-08-18 | From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation | Xingjian Wang et.al. | 2608.18076v1 | null |
| 2026-08-18 | StagedWorkspace: A Versioned Workspace for Knowledge-Work Agents | Yining Hua et.al. | 2608.18050v1 | null |
| 2026-08-18 | Chain-of-Experience for Continual LLM Improvement | Haoqin Tu et.al. | 2608.18027v1 | null |
| 2026-08-18 | Can Large Language Models Explain Flight Safety Events? A Prior-Guided Semantic LLM-based Approach | Lu Xu et.al. | 2608.18017v1 | null |
| 2026-08-18 | The IOL-AI Challenge: An Open Challenge towards Advancing Linguistic Reasoning | Eduardo Sánchez et.al. | 2608.18011v1 | null |
| 2026-08-18 | Judge, Retrieve, or Abstain: Uncertainty-Guarded LLM Judging with Provable Risk Guarantees | Sher Badshah et.al. | 2608.17994v1 | null |
| 2026-08-18 | Against Political Polarization: A Unified Framework for Tracing Evolving Political Ideologies on Social Media | Yijie Xu et.al. | 2608.17987v1 | null |
| 2026-08-18 | Do Large Language Models Play Six Degrees of Separation? Measuring Topological Compression in Long-Context Manifolds | Md. Faiyaz Abdullah Sayeedi et.al. | 2608.17950v1 | null |
| 2026-08-18 | Efficient RLVR Scheduling via Graph-Structured Online Difficulty Estimation | Zhizhao Liu et.al. | 2608.17941v1 | null |
| 2026-08-18 | Collective Counterfactual Planning: Coordination, Consent, and Verification under Representational Constraints | Chainarong Amornbunchornvej et.al. | 2608.17932v1 | null |
| 2026-08-18 | Analysis of Types of Inquiries in Student-AI Interaction: A case study of two CS2 tasks | Matin Amoozadeh et.al. | 2608.17919v1 | null |
| 2026-08-18 | AutoResearch: Insight In, Hallucination Out | Yiming Ren et.al. | 2608.17906v1 | null |
| 2026-08-18 | BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models | Liubov Chubarova et.al. | 2608.17895v1 | null |
| 2026-08-18 | From Global Benchmarks to Local Evaluations: Benchmarking LLMs for the German Public Sector | Camilla Dalerci et.al. | 2608.17827v1 | null |
| 2026-08-18 | Interpretable Humans, Alien LLMs: Expert Analysis of Latent Structures in Assessment Responses | Alona Strugatski et.al. | 2608.17810v1 | null |
| 2026-08-18 | Whether LLMs Can Navigate Beliefs and Facts Depends on How You Phrase It | Quang Minh Nguyen et.al. | 2608.17809v1 | null |
| 2026-08-18 | An Empirical Study of Reward Specification and Benchmark Reliability in GRPO-based LLM Unlearning | Rubén Balbastre et.al. | 2608.17804v1 | null |
| 2026-08-18 | Neuro-symbolic learning over OWL 2 DL via consequence-based compilation to differentiable circuits | Olga Mashkova et.al. | 2608.17741v1 | null |
| 2026-08-18 | What Aggregate Scores Miss: Measuring Item-Level Regressions in Commercial LLM API Migrations | Xiaonan Xu et.al. | 2608.17719v1 | null |
| 2026-08-18 | Communicating Credit Risk with Large Language Models: Evaluation of Explanations from Standard and Alternative Data-Based Models | Sahab Zandi et.al. | 2608.17715v1 | null |
| 2026-08-18 | GraphWake: Group Polarization via Memory-Mediated Polarization Cascade in LLM-Agent Communities | Haoran Bu et.al. | 2608.17665v1 | null |
| 2026-08-18 | Graph Surgery and the Do-Operator: A Precise Correspondence for Acyclic Structural Causal Models | Satpreet Makhija et.al. | 2608.17634v1 | null |
| 2026-08-18 | Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges | Syeda Faiza Ahmed et.al. | 2608.17605v1 | null |
| 2026-08-18 | tinyDSM: A Framework for Skill Modeling and Development for Resource-Constrained Millirobots | Markus D. Kobelrausch et.al. | 2608.17596v1 | null |
| 2026-08-18 | Quantifying Risk Under Evolving Uncertainty: Belief-Dependent Robustness for Safe Sequential Decision Making | Deep Kumar Ganguly et.al. | 2608.17574v1 | null |
| 2026-08-18 | Code as Representation: A Compilable Parsing Paradigm for Academic Documents | Rihui Jin et.al. | 2608.17550v1 | null |
| 2026-08-18 | CoAL-RAG: A Complexity-Aware Legal Retrieval-Augmented Generation Method | Jin Su et.al. | 2608.17536v1 | null |
| 2026-08-18 | When to Review: Spaced Repetition for Continual Pre-Training of Language Models | Alankar Atreya et.al. | 2608.17530v1 | null |
| 2026-08-18 | Effects of Answer Format Variation on Gender Bias in Large Language Models | Ksenia Merzlyakova et.al. | 2608.17516v1 | null |
| 2026-08-18 | Beyond FLOPs: Energy-Aware Knowledge Distillation for Sustainable LLMs on Code-Related Task | Enrique Barba Roque et.al. | 2608.17515v1 | null |
| 2026-08-18 | SGHA: Evidence-Grounded Research Problem Discovery with Local Language Models | Sarvesh Gharat et.al. | 2608.17501v1 | null |
| 2026-08-18 | SAGE: Self-Evolving Storyboard Skills via Attribution-Guided Rule Evolution | Maolin Ran et.al. | 2608.17468v1 | null |
| 2026-08-18 | Structure-Internalized Rule Language Model for Faithful Knowledge Graph Reasoning | Xingrui Zhuo et.al. | 2608.17443v1 | null |
| 2026-08-18 | Cognitive Graph Intelligence for Adaptive and Robust DDoS Attack Detection in Next Generation Networks | Mohammad Arif Hossain et.al. | 2608.17352v1 | null |
| 2026-08-18 | DeAR: Decentralized Agentic Reasoning via Capability Grounding and Collaborative Thought Navigation | Xing Wei et.al. | 2608.17282v1 | null |
| 2026-08-18 | ASI-Bench: At the Dawn of Artificial Superintelligence | Junwei Zhou et.al. | 2608.17271v1 | null |
| 2026-08-18 | Understanding Curriculum Learning in Large Language Models via Cross-Difficulty Optimization Dynamics | Zhikai Ding et.al. | 2608.17268v1 | null |
| 2026-08-18 | Structural Plan-to-Model Conversion with Deterministic Geometry and Guarded Agentic Vision-Language Refinement | Mohammad Talebi-Kalaleh et.al. | 2608.17237v1 | null |
| 2026-08-17 | Synthesizing Feature Extractors: An Agentic Approach for Algorithm Selection | Hai Xia et.al. | 2608.17170v1 | null |
| 2026-08-17 | Towards Safer RAG: Only Agents Capable of System 2 Thinking may Access Untrusted Documents | Mehrdad Ghassabi et.al. | 2608.17153v1 | null |
| 2026-08-17 | KnowSim: Evaluating Information Calibration in LLM Assistants with User Simulators that Learn | Yoonjoo Lee et.al. | 2608.17150v1 | null |
| 2026-08-17 | A decodability criterion predicts when hidden-state selection beats majority voting in large language models | Zhixiang wang et.al. | 2608.17124v1 | null |
| 2026-08-17 | From Abductive Explanations to Global Logical Rules for Node Classification in SGCs | Bryan Lima Cavalcante et.al. | 2608.17103v1 | null |
| 2026-08-17 | J-Miner: Recovering Executable Decision Knowledge from Language-Model Classifiers | Yunfan Gao et.al. | 2608.17063v1 | null |
| 2026-08-17 | Cross-Model Memory Transfer via Target-Side Reader Adaptation | Mingyuan Li et.al. | 2608.17050v1 | null |
| 2026-08-17 | AutoSR: Automatic Symbolic Regression by Searching Research States | Kejia Zhang et.al. | 2608.16876v1 | null |
| 2026-08-17 | Quipu: A Governed Bitemporal Knowledge Graph Store | Steve Brown et.al. | 2608.16813v1 | null |
| 2026-08-17 | Bounded Semantic Planning and Deterministic Compilation for Reliable Enterprise Text-to-SQL | Yi Ai et.al. | 2608.16663v1 | null |
| 2026-08-17 | The Working Set of a Coding Agent: Coherence Debt in Repository-Scale Tasks | Bardia Mohammadi et.al. | 2608.16630v1 | null |
| 2026-08-17 | Hypergraph-based Multimodal Retrieval-Augmented Generation with Incremental Refinement | Shenao Chen et.al. | 2608.16628v1 | null |
| 2026-08-17 | Physics of Agents: Statistical Mechanics Predicts Collective Behavior of AI Agents | Batu El et.al. | 2608.16578v1 | null |
| 2026-08-17 | Ask, Condition or Abstain: Reinforcement Learning for Missing-Premise Reasoning | Yongqi Tong et.al. | 2608.16554v1 | null |
| 2026-08-17 | VCE-Skill: Enhancing Skill Self-Evolution with Version-Change Experience | Jianming Chen et.al. | 2608.16544v1 | null |
| 2026-08-17 | Large language models as synthetic clinical experts to inform longitudinal rare-disease modeling | Clemens Schächter et.al. | 2608.16507v1 | null |
| 2026-08-17 | Graph Machine Learning: An Opportunity for Power Systems | Martin Sadric et.al. | 2608.16494v1 | null |
| 2026-08-17 | Time to Reason: Scalable Neurosymbolic Learning for LTLf via Fuzzy Semantics | Riccardo Andreoni et.al. | 2608.16443v1 | null |
| 2026-08-17 | Reasoning-supported Robustness Validation of Automotive E/E Components | Jan Novacek et.al. | 2608.16421v1 | null |
| 2026-08-17 | Think Inside the Chunk: RegulaRAG for Regulation-Compliant Scenario Generation using LLMs: A Case Study of UN Regulation No. 152 | Vahid Zolfaghari et.al. | 2608.16394v1 | null |
| 2026-08-17 | Mint-Agent: Introducing Finance-Native Agentic Foundation Models | Mint-Agent Team et.al. | 2608.16386v1 | null |
| 2026-08-17 | MELD: A Protocol for Merging Knowledge Across Distributed Agentic Memories | Lauri Lovén et.al. | 2608.16357v1 | null |
| 2026-08-17 | AeroCopilotBench: A Two-Tier Benchmark for Evaluating LLM Agents as Aviation Copilots in an Interactive Virtual Cockpit Environment | Yuchen Yuan et.al. | 2608.16349v1 | null |
| 2026-08-17 | Executable Code Knowledge: Code as a Native, Validation-Carrying Knowledge Representation for AI Coding Agents | Xueping Gao et.al. | 2608.16295v1 | null |
| 2026-08-17 | Clause Encounters of the Third Kind: Can LLMs Replace Language Teachers? | Kristina Šekrst et.al. | 2608.16286v1 | null |
| 2026-08-17 | Domain-Agnostic Neural Topic Modeling with Contextual Token-Level Semantic Graph Representation | Seung-Won Seo et.al. | 2608.16269v1 | null |
| 2026-08-17 | Picking the Right Image to Classify: Reliable-Input Selection in Teledermatology | Fabian Gröger et.al. | 2608.16198v1 | null |
| 2026-08-17 | LENS: In-Context Search via Latent Evidence Exploration over Dynamic Raw Documents | Xingjun Wang et.al. | 2608.16185v2 | null |
| 2026-08-17 | Agent-Native Telemetry: Verifiable State-Delta Evidence for Autonomous Operations | Jun He et.al. | 2608.16178v1 | null |
| 2026-08-17 | FeatureHospital: A Skill-Driven Multi-Agent Framework for Automated Algorithm Customization in Multi-View Multi-Label Feature Selection | Junxuan Li et.al. | 2608.16148v1 | null |
| 2026-08-17 | Graph Neural Assisted Actor-Critic for Latency-Efficient Edge Vision System | Alam Noor et.al. | 2608.16142v1 | null |
| 2026-08-17 | HyperSkill: Self-Evolving LLM Agents via Hypergraph-Structured Skill Memory | Ruiyao Xu et.al. | 2608.16114v1 | null |
| 2026-08-17 | RetroMPA: A Molecular Property-Aware Auxiliary Framework for Enhancing Retrosynthesis Prediction | Mianzhi Liu et.al. | 2608.16111v1 | null |
| 2026-08-17 | The Commercial Tax: Rent-vs-Own Blind Spots in Multi-Hop Retrieval Benchmarks | Luis M. Sanchez et.al. | 2608.16096v1 | null |
| 2026-08-17 | Skill2Query: Exploiting Skill Structure to Generate Pseudo-Queries for Agent Skill Retrieval | Lihui Ding et.al. | 2608.16071v1 | null |
| 2026-08-17 | OceanLight: Efficient Global Ocean Forecasting via Geometry-Adaptive Unstructured Mesh Representation | Wei Wu et.al. | 2608.16070v1 | null |
| 2026-08-17 | NICE: Scale-Stable Perturbations for Graph Neural Network Explanations via Noise Corruption | Ziluowen Luo et.al. | 2608.16038v2 | null |
| 2026-08-17 | RagGAD: Rationale-Aware Conditional Gaussian Mixture Normalizing Flow for Unsupervised Graph Anomaly Detection | Junxin Lu et.al. | 2608.16018v1 | null |
| 2026-08-17 | From Sequence to Structure: Relational Uncertainty Propagation for LLM Agents | Zhengzhao Ma. Boxi Cao et.al. | 2608.16002v1 | null |
| 2026-08-16 | PLSQLBench: Benchmarking LLM Systems for Executable Procedural Database Programming | Marianne Menglin Liu et.al. | 2608.15931v1 | null |
| 2026-08-16 | Unified Pedestrian Path Prediction Using Inverse Reinforcement Learning | Šimon Sukup et.al. | 2608.15929v1 | null |
| 2026-08-16 | Noesis: Bidirectional Graph-RAG with Adaptive Parallelism and Cross-Knowledge-Base Semantic Discovery | Nicola Cogotti et.al. | 2608.15919v1 | null |
| 2026-08-16 | Large language model-assisted discovery of cohorts from scientific literature | Moritz Sturm et.al. | 2608.15909v1 | null |
| 2026-08-16 | Scaling Manual-Grounded Appliance Manipulation with Data Synthesis and Unified Planning | Yuxing Long et.al. | 2608.15863v1 | null |
| 2026-08-16 | RAGas: Retrieval-Augmented Gas Optimization for Smart Contracts with Continuous Knowledge Integration | Yishun Wang et.al. | 2608.15857v1 | null |
| 2026-08-16 | Characterising cardiac tissue properties with graph neural networks | Ching-En Chiu et.al. | 2608.15843v1 | null |
| 2026-08-16 | Schema-Agnostic Graph Reasoning Agent for Hybrid Knowledge Graphs | Marius Dragic et.al. | 2608.15834v1 | null |
| 2026-08-16 | The Authority Resolution Framework: A Five-Domain Ontology for Governing Who and What Decides, at Scale | Parviz Shariff et.al. | 2608.15832v1 | null |
| 2026-08-16 | QuantumPhaseNet: A Gauge-Covariant Geometric and Quantum-Spectral Theory of Semantic Concept Hierarchies with Prototype Validation of a Classical Quantum-Inspired Model | Kiyotaka Kasubuchi et.al. | 2608.15820v1 | null |
| 2026-08-16 | ALKEMIE Agent: an autonomous platform for computational materials design | Hongfu Huang et.al. | 2608.15776v1 | null |
| 2026-08-16 | Large Models for Small Devices: Recent Advances and Empirical Analysis of Edge AI Deployment | Subhransu Das et.al. | 2608.15693v1 | null |
| 2026-08-16 | BERTopic-Virality Prioritisation: A Scalable Framework for Thematic and Comparative Analysis of COVID-19 and Monkeypox Misinformation on Twitter | Mkululi Sikosana et.al. | 2608.15691v1 | null |
| 2026-08-16 | THESIS-MoE: Trainable Hierarchical Extraction and SteerIng of Sycophancy in Mixture-of-Experts | Kareem Hassani et.al. | 2608.15687v1 | null |
| 2026-08-16 | Agent Gym: A Framework for Continuous Evaluation and Evolution of LLM Agents Through Human-in-the-Loop Feedback | Pouya Ghiasnezhad Omran et.al. | 2608.15591v1 | null |
| 2026-08-16 | GraniKV: Asymmetric Granularity KV-Cache Paging for Multi-Agent Systems with Long Shared Prefix | Jinhyun Jeon et.al. | 2608.15584v1 | null |
| 2026-08-16 | From Generalist to Specialist: A Context-Fusion Framework for Endoscopic Polyp Reporting with a Frozen VLM | Ruijie Yang et.al. | 2608.15580v1 | null |
| 2026-08-16 | Admission Without Answers: Label-Free Certification and Experience Learning for LLM-Based Optimization Modeling | Junbo Jacob Lian et.al. | 2608.15565v2 | null |
| 2026-08-16 | BengaliMCQ: Automatic Generation and Answer Prediction of Academic Multiple-Choice Questions in a Low-Resource Language | Abu Tarabin Surzo et.al. | 2608.15547v1 | null |
| 2026-08-16 | L3Cube-IndicQuest v2: A Large-Scale Multilingual Benchmark for Evaluating Factual Knowledge of Large Language Models Across Indic Languages | Rinit Jain et.al. | 2608.15535v1 | null |
| 2026-08-16 | Mental Model Management: An Operator-Based Framework for LLM Memory | Oliver Kramer et.al. | 2608.15451v1 | null |
| 2026-08-15 | Implementation of a Metacognition Framework for Self-Awareness and Self-Regulation in Ensembles of LLMs | Charles Courchaine et.al. | 2608.15400v1 | null |
| 2026-08-15 | Grounding Healthcare LLMs in a Causal Knowledge Graph: Framework, Metrics, and a Cardiovascular Pilot | Ummara Mumtaz et.al. | 2608.15382v1 | null |
Abstracts
From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation
2608.18076v1 by Xingjian Wang, Zhao Wang, Taihang Hu, Jun Zheng, Qing Jin, Qinye Zhou, Zhengtao Wu, Yongchao Du, Zuan Gao, Chao Lin, Yefeng Shen, Xiaoli Xu, Zhengze Xu, Hao Yan, Yuhang Yu, Mingzhou Zhang, Mengting Chen
Large-scale image generation has benefited from advances in data scale, quality, rebalancing, and recaptioning, yet conventional pipelines typically optimize task-specific datasets in isolation. A central challenge is not only how to curate each task-specific corpus, but also how to organize heterogeneous supervision according to the dependencies among generative capabilities. We present a \textbf{capability-driven data infrastructure} that couples capability-specific supervision construction with capability-aligned curriculum scheduling. Its three specialized yet interoperable data engines build complementary relational supervision for text-image grounding, inter-image transformation, and image-knowledge association, while caption experts align T2I and editing supervision across tasks and granularities. A multi-stage curriculum jointly evolves task composition, visual-concept distribution, data quality, and image resolution along the dependency order of capability acquisition, with capability-aware evaluation closing the loop through targeted retrieval, expert construction, and gap-aware resampling. At scale, the framework curates a 440M-image T2I corpus, 120M editing pairs, and over 27M image-entity pairs. With this infrastructure, we train multimodal diffusion models at two scales from scratch, with 3B and 6B sizes respectively. We conduct quantitative evaluation on CPI-Bench, along with qualitative evaluations across diverse text-to-image and editing scenarios. Experimental results present broad visual coverage, versatile rendering, and effective transfer across generative capabilities.
摘要:大規模圖像生成受益於數據規模、質量、重新平衡和重新標題的進步,但傳統流程通常在孤立的情況下優化特定任務的數據集。一個主要挑戰不僅在於如何策劃每個特定任務的語料庫,還在於如何根據生成能力之間的依賴關係組織異質監督。我們提出了一個\textbf{以能力為驅動的數據基礎設施},將特定能力的監督構建與能力對齊的課程安排結合起來。它的三個專門但可互操作的數據引擎為文本-圖像基礎、圖像間轉換和圖像-知識關聯構建互補的關係監督,同時標題專家在任務和粒度之間對齊T2I和編輯監督。一個多階段課程共同演變任務組合、視覺概念分佈、數據質量和圖像解析度,沿著能力獲取的依賴順序進行,而以能力為中心的評估通過針對性檢索、專家構建和關注差距的重採樣來閉合循環。在規模上,該框架策劃了一個包含4.4億圖像的T2I語料庫、1.2億編輯對和超過2700萬圖像-實體對。利用這一基礎設施,我們從零開始訓練了兩個規模的多模態擴散模型,分別為30億和60億大小。我們在CPI-Bench上進行了定量評估,並在多樣的文本到圖像和編輯場景中進行了定性評估。實驗結果顯示出廣泛的視覺覆蓋、多樣的渲染和在生成能力之間的有效轉移。
StagedWorkspace: A Versioned Workspace for Knowledge-Work Agents
2608.18050v1 by Yining Hua, Hongbin Na, Yifan Zhou, Akshay Kalose, Cyrus Ayubcha, Levi Lian
AI agents increasingly perform knowledge work (i.e., produce and modify persistent digital artifacts such as code repositories, documents, spreadsheets, slides, reports), yet the parsed views they search, the native files they edit, the changes they review, and the artifacts they submit can refer to different versions of the same work product. We formulate this as a workspace-state contract: every view should be explicitly tied to a version of the evolving workspace state. Coding agents partly address this need through repository contracts for search, diffs, and tests, whereas an analogous contract is less explicit for PDFs, spreadsheets, slides, notebooks, and mixed-format project folders. We propose StagedWorkspace, a versioned workspace for knowledge-work agents. The workspace binds parsed records and review diffs to content hashes of the native files as they change. In fixed-harness ablations on OfficeQA Pro and APEX-Agents, dual parsed/native access has the highest point estimate for every tested model; relative to the more limiting single view, it improves OfficeQA Pass@1 by 8.3-12.1 points and APEX mean rubric score by 4.7-9.2 points. SW-AGENT scores 63.9% with Gemini 3.1 Pro on OfficeQA and 42.1 with GPT-5.4 Nano on APEX, compared with published same-model scores of 29.3% and 25.5, respectively. A paired review-axis ablation on 57 file-editing tasks further finds higher observed scores when diffs are visible. These results identify workspace state as an experimental variable in knowledge-work agents and motivate benchmarks that score evidence, staged edits, and submitted artifacts as explicit state transitions.
摘要:AI 代理人越來越多地執行知識工作(即,產生和修改持久的數位工件,如代碼庫、文件、電子表格、簡報、報告),然而他們所搜尋的解析視圖、編輯的原始文件、審查的變更以及提交的工件可能指的是同一工作產品的不同版本。我們將此表述為工作區狀態合約:每個視圖應明確與不斷演變的工作區狀態的某個版本相關聯。編碼代理人部分通過針對搜索、差異和測試的庫合約來滿足這一需求,而對於 PDF、電子表格、簡報、筆記本和混合格式的項目文件夾,類似的合約則不那麼明確。我們提出了 StagedWorkspace,一個針對知識工作代理人的版本化工作區。該工作區將解析記錄和審查差異綁定到隨原始文件變更的內容哈希。在 OfficeQA Pro 和 APEX-Agents 的固定裝置消融實驗中,雙重解析/原生訪問對於每個測試模型的最高點估計;相對於更具限制性的單一視圖,它將 OfficeQA Pass@1 提高了 8.3-12.1 分,將 APEX 的平均評分提高了 4.7-9.2 分。SW-AGENT 在 OfficeQA 上的得分為 63.9%,在 APEX 上的得分為 42.1,與已發表的同模型得分分別為 29.3% 和 25.5 相比。對 57 個文件編輯任務的配對審查軸消融進一步發現,當差異可見時,觀察到的得分更高。這些結果將工作區狀態確定為知識工作代理人的實驗變量,並激勵對證據、分階編輯和提交工件進行明確狀態轉換的基準評分。
Chain-of-Experience for Continual LLM Improvement
2608.18027v1 by Haoqin Tu, Yunhao Fang, Yizhong Wang, Cihang Xie, Shen Yan
Humans continuously learn from experience, whereas conventional large language model (LLM) evaluations ignore the models' ability to improve through inference-time interaction. In this paper, we study how LLMs learn from iterative experience at test time, a setting we refer to as Chain-of-Experience (CoE), where models accumulate experiential traces through iterative interactions with self or environmental feedback to form a continual improvement loop beyond zero-shot inference. We instantiate CoE with diverse feedback mechanisms, including model self-feedback and environmental signals such as correctness or public coding test pass rates, and evaluate across math, coding, and knowledge domains using 8 LLMs, including GPT-5, Gemini-2.5 Pro, Claude-4.5 Sonnet. Our study shows that leveraging iterative experience consistently outperforms feedback-free baselines, achieving substantial gains with self feedback alone, alongside a 5.6% overall improvement and 19% lower API cost across tasks and models. We further show that combining complementary feedback channels (e.g., model and correctness signals) yields additional gains, and that CoE delivers higher accuracy per token than existing test-time strategies. We observe a positive correlation between LLM base ability and improvement capacity, and show that models remain robust under weak or spurious feedback, with different feedback contributing to distinct improvement aspects and most gains emerging early in the iterations.
摘要:人類不斷從經驗中學習,而傳統的大型語言模型(LLM)評估則忽略了模型通過推理時互動來改進的能力。在本文中,我們研究了 LLM 如何在測試時從迭代經驗中學習,這種情境我們稱之為經驗鏈(Chain-of-Experience, CoE),在這裡模型通過與自身或環境反饋的迭代互動積累經驗痕跡,以形成超越零-shot 推理的持續改進循環。我們用多樣的反饋機制來實現 CoE,包括模型自我反饋和環境信號,如正確性或公共編碼測試通過率,並使用 8 種 LLM 進行數學、編碼和知識領域的評估,包括 GPT-5、Gemini-2.5 Pro 和 Claude-4.5 Sonnet。我們的研究表明,利用迭代經驗的表現始終優於無反饋的基準,僅依靠自我反饋就實現了顯著的增益,並在各任務和模型中達到 5.6% 的整體改進和 19% 的 API 成本降低。我們進一步顯示,結合互補的反饋通道(例如模型和正確性信號)會產生額外的增益,並且 CoE 在每個 token 上提供的準確性高於現有的測試時策略。我們觀察到 LLM 的基本能力與改進能力之間存在正相關,並顯示模型在弱或虛假反饋下仍然保持穩健,不同的反饋對不同的改進方面有所貢獻,大多數增益在迭代的早期出現。
Can Large Language Models Explain Flight Safety Events? A Prior-Guided Semantic LLM-based Approach
2608.18017v1 by Lu Xu, Xu Li, Linjiang Zheng, Fan Li, Riquan Zhang, Jiaxing Shang
Improving flight safety with flight data requires not only accurate detection of risk events, but more importantly, clear interpretation of their underlying causes at the level of pilot control behavior. Existing explainable AI techniques, such as feature importance maps, often require considerable domain knowledge to translate them into operationally meaningful explanations. Large Language Models (LLMs), which excel at language reasoning, bring a promising solution to this issue. However, applying LLMs in this domain presents key challenges such as modal inconsistency, limited classification ability, scarcity of task-specific data for fine-tuning, and lack of domain knowledge. To overcome these challenges, we propose FlightLLM, a prior-guided semantic LLM-based approach for interpretable flight safety analysis. Specifically, we first perform feature engineering to address modal inconsistency, combining statistical descriptors with physically meaningful flight indicators. This representation is further processed by a Semantic Discretization module, which converts abstract numerical patterns into qualitative descriptions that are more compatible with language reasoning. In addition, since LLMs are not inherently strong classifiers, CatBoost is incorporated as a statistical expert, and its prediction results are injected into the prompt as prior guidance. A contrastive few-shot learning strategy is further adopted to compensate for limited data. Finally, we design structured prompts to embed aviation-specific knowledge into the inference process. Using hard landing, a representative risk event with complex causal mechanisms, as an anchor point, we evaluate FlightLLM on a dataset of 704 real-world A320 flight samples. Experimental results show that the proposed approach achieves competitive classification performance while generating direct and reasonable explanations for event causes.
摘要:改善飛行安全需要不僅準確檢測風險事件,更重要的是在飛行員控制行為層面清晰解釋其潛在原因。現有的可解釋AI技術,如特徵重要性圖,通常需要相當的領域知識才能將其轉化為具有操作意義的解釋。大型語言模型(LLMs)在語言推理方面表現出色,為這一問題帶來了有希望的解決方案。然而,在這一領域應用LLMs面臨著關鍵挑戰,如模式不一致、有限的分類能力、缺乏特定任務的數據以進行微調,以及缺乏領域知識。為了克服這些挑戰,我們提出了FlightLLM,一種基於語義的先驗引導LLM方法,用於可解釋的飛行安全分析。具體而言,我們首先進行特徵工程以解決模式不一致,將統計描述符與具有物理意義的飛行指標相結合。這一表示進一步由語義離散化模塊處理,將抽象的數字模式轉換為更符合語言推理的定性描述。此外,由於LLMs本身並不是強大的分類器,因此CatBoost被納入作為統計專家,其預測結果被注入到提示中作為先驗指導。進一步採用了對比少樣本學習策略以彌補數據的有限性。最後,我們設計了結構化提示,將航空特定知識嵌入推理過程中。以硬著陸作為錨點,這是一個具有複雜因果機制的代表性風險事件,我們在704個真實世界A320飛行樣本的數據集上評估FlightLLM。實驗結果表明,所提出的方法在生成事件原因的直接和合理解釋的同時,實現了具有競爭力的分類性能。
The IOL-AI Challenge: An Open Challenge towards Advancing Linguistic Reasoning
2608.18011v1 by Eduardo Sánchez, Rita Berrada, Dan-Mircea Mirea, Sara Rajaee, Alexander Piperski, Ana Meta Dolinar, Boris Iomdin, Andrey Nikulin, Mariya Shmatova, Marzieh Fadaee, Julia Kreutzer
Reasoning in LLMs is overwhelmingly studied in domains that provide a model with rules: mathematics and code. Linguistic puzzles invert this: the solver must first discover the system before reasoning within it. We present the IOL-AI Challenge, an open-science competition run on the unseen problems of the International Linguistics Olympiad (IOL) 2026 Individual Contest, evaluated both automatically and, for the first time, by members of the official IOL Jury under the same rubrics applied to human contestants. The challenge drew 731 submissions from 46 teams under a strict compute budget (one T4, 30 mins). We additionally benchmark 15 unconstrained frontier and open models, with Claude Opus 4.8 earning a jury score equivalent to a gold medal, while both resource-constrained systems we submitted for jury grading scored in the range of the bottom 5% of contestants. Capability was not determined by scale: 14B submissions outperform models twice their size, and gains come from decoding and output-handling rather than model capacity. We also found that automatic metrics rank systems exactly as the jury does, but compress the scale, upscoring weak systems by ~13 points and understating strong ones. Our analysis shows that while frontier models might have prior knowledge about some of the problem languages, it does not significantly help them solve the linguistic reasoning tasks, leaving linguistic reasoning as a strong benchmarking proxy for generalizable reasoning skills.
摘要:推理在大型語言模型(LLMs)中的研究主要集中在提供規則的領域:數學和程式碼。語言謎題則顛倒了這一點:解題者必須首先發現系統,然後才能在其中進行推理。我們提出了IOL-AI挑戰賽,這是一項開放科學競賽,基於2026年國際語言奧林匹亞(IOL)個人賽的未見問題進行評估,這次評估既有自動評分,還首次由官方IOL評審團成員根據與人類參賽者相同的標準進行評分。這次挑戰吸引了46個團隊提交的731份作品,並在嚴格的計算預算下進行(一個T4,30分鐘)。我們還基準測試了15個不受限制的前沿和開放模型,其中Claude Opus 4.8獲得了相當於金牌的評審分數,而我們提交給評審打分的兩個資源受限系統的分數則落在參賽者的底部5%範圍內。能力並不是由規模決定的:14B的提交表現超過了規模是其兩倍的模型,並且性能的提升來自於解碼和輸出處理,而非模型容量。我們還發現,自動指標的排名與評審的排名完全一致,但壓縮了評分範圍,將弱系統的分數提高了約13分,而低估了強系統的分數。我們的分析顯示,儘管前沿模型可能對某些問題語言有先前的知識,但這並未顯著幫助它們解決語言推理任務,這使得語言推理成為通用推理能力的強基準代理。
Judge, Retrieve, or Abstain: Uncertainty-Guarded LLM Judging with Provable Risk Guarantees
2608.17994v1 by Sher Badshah, Ali Emami, Hassan Sajjad
Using LLMs as judges has become standard practice for evaluating model outputs at scale. This is particularly common for subjective, open-ended tasks such as assessing helpfulness or alignment, where no single reference answer exists. However, objective tasks introduce a distinct reliability challenge for reference-free LLM judging. In the absence of a reference answer, the judge evaluates factual correctness either through its parametric knowledge or through tool augmentation. Although the former enables efficient evaluation, the judge may hallucinate or lack sufficient evidence for its verdict. Conversely, tool augmentation can provide additional evidence but introduces extra computational cost and requires an appropriate mechanism to determine when and how that evidence should be used reliably. More importantly, neither approach alone provides formal control over the risk of accepted verdicts or guarantees their reliability at a specified level. We propose a risk-controlled framework that calibrates uncertainty thresholds on a held-out set so that the false discovery rate among accepted verdicts remains below a user-specified level~$α$ with high probability, using finite-sample Clopper--Pearson intervals. When the parametric mode is not sufficiently confident, the instance is routed to a retrieval-augmented mode, where the judge gathers web evidence and re-evaluates the instance under a second calibrated threshold. The finite-sample guarantee carries over to this two-threshold routing without additional assumptions. Across open-domain QA benchmarks and judges of varying scales, the framework maintains the target error rate while achieving substantially higher coverage than single-mode baselines.
摘要:使用大型語言模型作為評審已成為大規模評估模型輸出的標準做法。這在主觀的、開放式的任務中尤其常見,例如評估有用性或一致性,因為這類任務並不存在單一的參考答案。然而,客觀任務對於無參考的 LLM 評審引入了明顯的可靠性挑戰。在缺乏參考答案的情況下,評審通過其參數知識或工具增強來評估事實的正確性。雖然前者能夠實現高效評估,但評審可能會出現幻覺或缺乏足夠的證據來支持其裁決。相反,工具增強可以提供額外的證據,但會引入額外的計算成本,並需要適當的機制來確定何時以及如何可靠地使用這些證據。更重要的是,單獨使用這兩種方法都無法對接受的裁決風險提供正式控制或保證其在特定水平上的可靠性。我們提出了一個風險控制框架,該框架在保留集上校準不確定性閾值,以便接受的裁決中的假陽性率以高概率保持在用戶指定的水平~$α$ 以下,使用有限樣本的 Clopper--Pearson 區間。當參數模式的信心不足時,實例會被路由到檢索增強模式,在該模式下,評審收集網絡證據並在第二個校準閾值下重新評估該實例。有限樣本的保證在這個雙閾值路由中延續,無需額外假設。在開放域問答基準和不同規模的評審中,該框架在保持目標錯誤率的同時,實現了顯著高於單一模式基準的覆蓋率。
Against Political Polarization: A Unified Framework for Tracing Evolving Political Ideologies on Social Media
2608.17987v1 by Yijie Xu, Chao Wang, Hui Xiong
The rapid growth of social media has greatly influenced political discourse, highlighting the need to understand individual political ideologies and their temporal dynamics. This task faces challenges such as data scarcity, abundant non-political content, costly and bias-prone manual annotation, and difficulty in modeling future ideological inclinations. To address these issues, we propose TSN4PI, a unified framework for tracking the evolution of political ideologies on social media. It includes two core modules. The PIDN uses large language models with style transfer and unsupervised domain adaptation to enable robust ideology detection and filter irrelevant content from noisy, cross-domain data. The PIPN employs temporal graph neural networks to predict future ideological shifts, enabling comprehensive analysis of ideology presence, intensity, and evolution. We release two large-scale datasets for noncommercial research use to facilitate further work. Extensive case studies on multiple platforms (X and Truth Social) validate the effectiveness of TSN4PI and provide empirical insights into political polarization and the evolution of online ideologies. Our findings offer a nuanced perspective, advancing both methodological development and empirical understanding in this field.
摘要:社交媒體的快速增長對政治話語產生了重大影響,突顯了理解個人政治意識形態及其時間動態的必要性。這項任務面臨著數據稀缺、非政治內容豐富、昂貴且易受偏見影響的手動標註以及未來意識形態傾向建模困難等挑戰。為了解決這些問題,我們提出了TSN4PI,一個用於追蹤社交媒體上政治意識形態演變的統一框架。它包括兩個核心模塊。PIDN使用大型語言模型結合風格轉換和無監督領域適應,以實現穩健的意識形態檢測並過濾來自嘈雜的跨領域數據中的無關內容。PIPN則利用時間圖神經網絡來預測未來的意識形態變化,使得對意識形態的存在、強度和演變進行全面分析成為可能。我們釋放了兩個大型數據集供非商業研究使用,以促進進一步的研究工作。在多個平台(X和Truth Social)上進行的廣泛案例研究驗證了TSN4PI的有效性,並提供了對政治極化和在線意識形態演變的實證見解。我們的發現提供了一個細緻的視角,推進了該領域的方法論發展和實證理解。
Do Large Language Models Play Six Degrees of Separation? Measuring Topological Compression in Long-Context Manifolds
2608.17950v1 by Md. Faiyaz Abdullah Sayeedi
Large Language Models (LLMs) demonstrate remarkable multi-hop reasoning capabilities over long contexts, yet the internal mechanisms enabling these distant cognitive leaps remain poorly understood. Traditional attention-based interpretability often fails to capture true semantic proximity due to routing artifacts like attention sinks. In this paper, we bypass attention weights to directly analyze the dynamic geometry of the hidden state manifold, proving that deep LLM latent spaces natively organize into Small-World networks. By sparsifying the continuous similarity matrices of long-context representations into unweighted graphs, we trace the connectivity between highly disjoint semantic anchors across two distinct architectures. Our findings reveal a sharp topological phase transition: while early syntactic layers remain entirely fractured, deep reasoning layers abruptly compress massive conceptual distances into highly navigable pathways strictly bounded by the "Six Degrees of Separation" limit (=< 6 semantic hops). Furthermore, we demonstrate the practical efficacy of this framework by applying it to zero-shot hallucination detection within Retrieval-Augmented Generation (RAG) using the RAGognize dataset. We show that factually grounded generations maintain structural integrity with their source context (approximately 3 hops), whereas hallucinations induce severe topological collapse. Ultimately, this work mathematically formalizes how transformers execute abstract reasoning and provides a novel, strictly geometric signature for evaluating factual reliability.
摘要:大型語言模型(LLMs)在長上下文中展現出卓越的多跳推理能力,但促成這些遙遠認知飛躍的內部機制仍然不甚了解。傳統的基於注意力的可解釋性常常無法捕捉到真實的語義接近性,這是由於路由伪影如注意力匯聚所致。在本文中,我們繞過注意力權重,直接分析隱藏狀態流形的動態幾何,證明深層LLM潛在空間本質上組織成小世界網絡。通過將長上下文表示的連續相似性矩陣稀疏化為無權重圖,我們追蹤兩個不同架構之間高度不相交的語義錨點之間的連接性。我們的研究結果揭示了一個明顯的拓撲相變:儘管早期的句法層完全破碎,深層推理層卻突然將巨大的概念距離壓縮成高度可導航的路徑,這些路徑嚴格受限於「六度分隔」的限制(=< 6語義跳躍)。此外,我們通過將此框架應用於檢索增強生成(RAG)中的零樣本幻覺檢測,展示了其實際效能,使用了RAGognize數據集。我們顯示,事實基礎的生成與其來源上下文保持結構完整(約3跳),而幻覺則引發嚴重的拓撲崩潰。最終,這項工作數學化了Transformer如何執行抽象推理,並提供了一種新穎的、嚴格的幾何特徵,用於評估事實可靠性。
Efficient RLVR Scheduling via Graph-Structured Online Difficulty Estimation
2608.17941v1 by Zhizhao Liu, Zhiliang Tian, Xi Wang, Zhihua Wen, Yihang Xiong, Zhiquan Lai, Dongsheng Li
Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models but relies on costly rollout exploration. Assigning the same exploration budget to samples with different difficulty levels is inefficient: easy samples may receive redundant rollouts, whereas difficult but learnable samples may receive too little exploration. Existing adaptive schedulers address this mismatch through curriculum-based sample selection or non-uniform rollout allocation based on estimated sample difficulty. However, obtaining reliable online difficulty estimates remains challenging: dedicated probing adds substantial generation overhead, whereas history-based estimators face a cold start with no initial observations and stale feedback, and typically ignore relations among samples. To address these limitations, we propose a plug-and-play graph-based online difficulty estimator that shares rollout feedback across related samples and continuously updates their difficulty estimates, mitigating cold start and staleness without dedicated probing. Specifically, we first construct a difficulty-aware sample graph based on semantic and reasoning similarities. Based on this graph, we introduce latent difficulty states and use a Potts prior to encourage neighboring samples to share the same state. We then employ a state-level Beta-Binomial model to aggregate the rollout outcomes associated with each state. Finally, we use an online mean-field variational algorithm to continuously update the latent-state assignments and state-level difficulty as new feedback arrives. Our framework can be integrated into sample-selection and rollout-allocation schedulers, enabling difficulty-adaptive exploration without dedicated probing. Experiments across multiple base models, RL schedulers, and benchmarks demonstrate that our framework achieves better performance.
摘要:強化學習與可驗證獎勵(RLVR)提升了大型語言模型的推理能力,但依賴於成本高昂的展開探索。將相同的探索預算分配給不同難度級別的樣本是低效的:簡單樣本可能會收到冗餘的展開,而難度較高但可學習的樣本可能會收到過少的探索。現有的自適應調度器通過基於課程的樣本選擇或根據預估樣本難度的非均勻展開分配來解決這一不匹配。然而,獲得可靠的在線難度估計仍然具有挑戰性:專門的探測增加了可觀的生成開銷,而基於歷史的估計器面臨著沒有初始觀察和過時反饋的冷啟動問題,並且通常忽略樣本之間的關係。為了解決這些限制,我們提出了一種即插即用的基於圖的在線難度估計器,該估計器在相關樣本之間共享展開反饋,並持續更新它們的難度估計,減輕冷啟動和過時問題,無需專門的探測。具體而言,我們首先根據語義和推理相似性構建一個難度感知樣本圖。基於這個圖,我們引入潛在的難度狀態,並使用Potts先驗來鼓勵相鄰樣本共享相同的狀態。然後,我們使用狀態級的Beta-Binomial模型來聚合與每個狀態相關的展開結果。最後,我們使用在線均場變分算法來持續更新潛在狀態分配和狀態級難度,隨著新反饋的到來。我們的框架可以集成到樣本選擇和展開分配調度器中,實現難度自適應探索,而無需專門的探測。在多個基礎模型、RL調度器和基準測試中的實驗表明,我們的框架實現了更好的性能。
Collective Counterfactual Planning: Coordination, Consent, and Verification under Representational Constraints
2608.17932v1 by Chainarong Amornbunchornvej
Groups routinely complete projects that no single member can plan, execute, or verify alone. We propose a formal model of this phenomenon, Collective Counterfactual Planning (CCP), in which the binding limitation on each agent is neither capability, knowledge, nor observability, but representational geometry: each agent perceives the state, conceives moves, consents to actions, and certifies goal requirements only through a projection onto an agent-specific subspace of a common task space. Four gates jointly determine whether a team can reach a conjunctive goal and legitimately recognize that it has done so: the exogenous implementation coalitions required to perform each action, together with three representational gates -- conception, consent, and task-relative verification qualification. We define the Collective Counterfactual Solvability (CCS) problem, separating geometric feasibility, executable attainment, and validated completion. The results expose a positive-negative duality. Iterated cross-agent relay can unlock a solution that no one-shot pooling of individual plans contains, but any goal requirement depending essentially on the subspace dark to the entire team is unverifiable and therefore not validly completable, even when the trajectory accidentally attains it. Memoryless and audited consent further constrain different objects -- action directions versus cumulative trajectory states -- and neither dominates the other. A four-step exhaustive horizon-bounded solvability scheme is sound and complete under exact representation of the relay closure; restricted implementations remain sound on returned plans but need not be complete. The model gives one geometry for sequential mutual enabling, competent execution of steps whose purpose is invisible to the executor, forced sub-teaming at expertise boundaries, and completion that cannot be validly declared.
摘要:團體經常完成單一成員無法獨自計劃、執行或驗證的項目。我們提出這一現象的正式模型,稱為集體反事實規劃(CCP),在這個模型中,每個代理的約束限制既不是能力、知識,也不是可觀察性,而是表徵幾何:每個代理僅通過投影到共同任務空間的代理特定子空間來感知狀態、構思行動、同意行為和認證目標要求。四個閘門共同決定一個團隊是否能夠達成聯合目標並合法地認識到它已經達成:執行每個行動所需的外生實施聯盟,以及三個表徵閘門——構思、同意和任務相對驗證資格。我們定義了集體反事實可解性(CCS)問題,將幾何可行性、可執行達成和驗證完成分開。結果揭示了一種正負對偶性。迭代的跨代理中繼可以解鎖一個單次個人計劃無法包含的解決方案,但任何本質上依賴於對整個團隊來說是黑暗的子空間的目標要求都是不可驗證的,因此無法有效完成,即使軌跡意外達成了它。無記憶和經審核的同意進一步限制了不同對象——行動方向與累積軌跡狀態——而且兩者不相互主導。一個四步的全面邊界可解性方案在中繼閉包的精確表徵下是健全且完整的;受限的實施在返回的計劃上仍然是健全的,但不必是完整的。該模型為順序相互啟用、執行目的對執行者不可見的步驟的能力執行、在專業邊界強制子團隊以及無法有效宣告的完成提供了一種幾何。
Analysis of Types of Inquiries in Student-AI Interaction: A case study of two CS2 tasks
2608.17919v1 by Matin Amoozadeh, Amin Alipour
Background and Context: Question and inquiry are integral parts of knowledge seeking and learning. Despite their importance, students tend not to ask enough questions in the classroom. However, studies have shown that students interact extensively with generative AI systems for learning and problem solving. Objective: In this paper, we seek to better understand the types of questions that students ask AI systems, and how those questions evolve during problem solving and across tasks. Method: We use the Graesser et al. taxonomy to classify students' inquiries into 18 types. We develop a few-shot learning approach to automatically classify students' interactions with AI into these categories. We use this system to analyze 830 interactions of CS2 students across two programming tasks. Findings: Our results suggest that a small subset of question types accounts for the majority of student inquiries, and that the types of questions students ask change substantially as the task progresses.
摘要:背景與背景:提問和探究是尋求知識和學習的重要部分。儘管它們的重要性,學生在課堂上往往不會提出足夠的問題。然而,研究顯示學生在學習和解決問題時,與生成式人工智慧系統的互動非常廣泛。
目標:在本文中,我們旨在更好地理解學生向人工智慧系統提出的問題類型,以及這些問題在解決問題和不同任務中的演變。
方法:我們使用Graesser等人的分類法將學生的提問分為18種類型。我們開發了一種少量學習方法,自動將學生與人工智慧的互動分類到這些類別中。我們使用這個系統分析830次CS2學生在兩個編程任務中的互動。
發現:我們的結果表明,小部分問題類型佔據了學生提問的主要部分,並且學生提出的問題類型在任務進行過程中有顯著變化。
AutoResearch: Insight In, Hallucination Out
2608.17906v1 by Yiming Ren, Xiang Liu, Qumeng Sun, Xiao Zhang, Jiahao Li, Haoyang Zhang, Junjie Wang
Autonomous research systems are increasingly capable of executing long research workflows, yet automation alone does not ensure that the resulting process remains scientifically grounded. We introduce AutoResearch, a two-stage system that connects Idea Generation with Idea Execution to address both how research ideas are formed and how they are reliably established through experimentation. In Idea Generation, AutoResearch continuously integrates emerging research signals with accumulated domain knowledge, identifies transferable mechanistic insights, and uses multi-model generation and cross-review to produce grounded, testable research plans. In Idea Execution, coordinated agents decompose these plans into experiments, iteratively implement and diagnose them, and employ independent evidence-based review before accepting research conclusions. Across representative settings in cross-modal retrieval, systems optimization, and benchmark-driven machine learning, AutoResearch turns generated ideas into measurable progress, detects and corrects unreliable experimental results, and makes evidence-conditioned decisions to continue, revise, or terminate research directions. For example, on RSICD benchmark, an AutoResearch-generated idea improves mean Recall from 32.84 to 34.69, while recording only 5 audit-confirmed issue events compared with 11-27 for other autonomous research systems. These results demonstrate a research process in which meaningful insight is grounded before experimentation and conclusions are grounded before acceptance: Insight In, Hallucination Out.
摘要:自主研究系統越來越能夠執行長期的研究工作流程,然而僅僅依賴自動化並不能確保所產生的過程保持科學基礎。我們介紹了 AutoResearch,一個兩階段的系統,將創意生成與創意執行連接起來,以解決研究想法是如何形成的,以及如何通過實驗可靠地建立這些想法。在創意生成階段,AutoResearch 持續整合新興的研究信號與累積的領域知識,識別可轉移的機制見解,並利用多模型生成和交叉審查來產出有根據、可測試的研究計劃。在創意執行階段,協調的代理將這些計劃分解為實驗,迭代實施和診斷它們,並在接受研究結論之前進行獨立的基於證據的審查。在跨模態檢索、系統優化和基準驅動的機器學習等代表性設置中,AutoResearch 將生成的想法轉化為可衡量的進展,檢測並修正不可靠的實驗結果,並做出基於證據的決策以繼續、修訂或終止研究方向。例如,在 RSICD 基準上,AutoResearch 生成的想法將平均召回率從 32.84 提高到 34.69,同時僅記錄了 5 次經審核確認的問題事件,而其他自主研究系統則記錄了 11-27 次。這些結果展示了一個研究過程,其中有意義的見解在實驗之前就已經建立,而結論在接受之前也已經有根據:見解進,幻覺出。
BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models
2608.17895v1 by Liubov Chubarova, Alexandra Kuleshova, Daniil Volkov, Kirill Sultanov, Alexey Zaytsev
While Multimodal Large Language Models (MLLMs) have made significant strides in visual comprehension, their ability to reason about text-dense, professional documents remains incompletely evaluated. Existing benchmarks emphasize information extraction, require external domain knowledge, or cover professional documents only as one of many settings. They are also largely English- or Chinese-centric, leaving other languages and Russian, in particular, substantially underrepresented. To address these limitations, we introduce BEAR-Bench (Bilingual Enterprise and Academic Reasoning), a self-contained, complex English-and-Russian benchmark comprising 1000 human-annotated questions based on text-rich business and scientific documents. We evaluate 16 proprietary and open-weight MLLMs, including Gemini 3.1 Pro and Qwen3.5-397B, on BEAR-Bench and observe clear headroom even for the strongest systems. Finally, we use the resulting model outputs to compare existing hallucination detection methods, evaluating not only how often models fail on BEAR-Bench but also how reliably those failures can be identified.
摘要:雖然多模態大型語言模型(MLLMs)在視覺理解方面取得了重大進展,但它們對於文本密集型的專業文件的推理能力仍未得到充分評估。現有的基準強調信息提取,需要外部領域知識,或者僅將專業文件作為眾多設置之一。這些基準在很大程度上以英語或中文為中心,使其他語言,特別是俄語,顯得大幅度不足。為了解決這些限制,我們推出了BEAR-Bench(雙語企業與學術推理),這是一個自包含的、複雜的英語和俄語基準,包含1000個基於文本豐富的商業和科學文件的人類標註問題。我們在BEAR-Bench上評估了16個專有和開放權重的MLLMs,包括Gemini 3.1 Pro和Qwen3.5-397B,並觀察到即使對於最強的系統也存在明顯的提升空間。最後,我們使用生成的模型輸出來比較現有的幻覺檢測方法,不僅評估模型在BEAR-Bench上的失敗頻率,還評估這些失敗能否被可靠地識別。
From Global Benchmarks to Local Evaluations: Benchmarking LLMs for the German Public Sector
2608.17827v1 by Camilla Dalerci, Thilo Michael, Robin Schaefer, Daniel Weinland
Public institutions face a persistent challenge in selecting LLMs suited to their specific context. Existing benchmarks, however, are of limited use as they primarily reflect English-language and US-centric settings, and often only evaluate task performance. In this paper, we present first results of MÖVE, a holistic evaluation framework for the German public sector, examining three rarely considered governance dimensions: energy consumption, provider transparency, and knowledge of German-party positions. Our results reveal significant trade-offs, with no single model excelling across all dimensions: estimated energy consumption varies more than 60-fold and is not explained by model size alone, information disclosure varies systematically across providers, and European models do not exhibit stronger knowledge of German party positions. Model selection for public institutions thus cannot rely on performance rankings alone. Instead, evaluations should also reflect the governance requirements of the deployment context.
摘要:公共機構在選擇適合其特定情境的LLM時面臨持續的挑戰。
然而,現有的基準測試用途有限,因為它們主要反映英語和美國中心的環境,且通常僅評估任務表現。
在本文中,我們呈現MÖVE的初步結果,這是一個針對德國公共部門的整體評估框架,檢視三個鮮少考慮的治理維度:能源消耗、供應商透明度和對德國政黨立場的了解。
我們的結果揭示了顯著的權衡,沒有單一模型在所有維度上表現優異:估計的能源消耗變化超過60倍,且僅以模型大小無法解釋,信息披露在不同供應商之間系統性變化,歐洲模型對德國政黨立場的了解並未顯示出更強的優勢。
因此,公共機構的模型選擇不能僅依賴於性能排名。
相反,評估還應反映部署情境的治理要求。
Interpretable Humans, Alien LLMs: Expert Analysis of Latent Structures in Assessment Responses
2608.17810v1 by Alona Strugatski, Licol Zeinfeld, Jason Cooper, Shelley Rap, Gil Schwarts, Giora Alexandron
The evaluation of large language models (LLMs) relies heavily on human-designed assessments, implicitly assuming that AI and humans employ similar underlying cognitive constructs. Challenging this assumption, we investigate whether the latent factors governing LLM performance carry the same substantive, human-interpretable meaning as the cognitive constructs governing human learners. Using responses from humans and six LLMs across quantitative reasoning and chemistry assessments, we conducted Exploratory Factor Analysis (EFA) separately for both groups. Subject-Matter Experts (SMEs) then blindly evaluated the resulting factor graphs to ascribe pedagogical meaning to the emerged constructs. SMEs successfully interpreted most of the human-derived factors. Conversely, they could not ascribe meaning to any LLM-derived factors in quantitative reasoning and interpreted only half of the LLM factors in chemistry. By combining data-driven EFA with blind expert interpretation, this framework shows that LLMs frequently operate on statistically opaque mechanisms distinct from human reasoning.
摘要:大型語言模型(LLMs)的評估在很大程度上依賴於人類設計的評估,隱含假設AI和人類使用相似的基本認知結構。挑戰這一假設,我們調查了支配LLM性能的潛在因素是否具有與支配人類學習者的認知結構相同的實質性、人類可解釋的意義。利用來自人類和六個LLM在定量推理和化學評估中的反應,我們分別對這兩組進行了探索性因素分析(EFA)。主題專家(SMEs)隨後盲目評估了所產生的因素圖,以賦予出現的結構教學意義。SMEs成功解釋了大多數人類衍生的因素。相反,他們無法為任何LLM衍生的因素在定量推理中賦予意義,並且只解釋了化學中一半的LLM因素。通過將數據驅動的EFA與盲專家解釋相結合,這一框架顯示LLMs經常在與人類推理不同的統計不透明機制上運作。
Whether LLMs Can Navigate Beliefs and Facts Depends on How You Phrase It
2608.17809v1 by Quang Minh Nguyen, Luis Frentzen Salim
Humans naturally form and express beliefs in daily communication, e.g., "I think the answer is 3" or "I suppose that's right." Such beliefs inevitably intertwine with fact and knowledge, making the ability to handle them in tandem desirable for large language models (LLMs), as they are increasingly deployed in user-facing settings. Prior work showed that even capable LLMs exhibit a systemic weakness in acknowledging user beliefs grounded in incorrect information. We extend this evaluation to 10 LLMs across 18 epistemic expressions and find that the size and direction of the weakness depend on the verb used to express the belief, with the accuracy gap between factual and false information ranging from +50% on "I vaguely remember" to -14% on "I seriously doubt". We further show that the phenomenon stems from task confusion: models default to fact-checking the underlying claim, overriding the user's stated belief; chains of thought that explicitly fact-check show lower accuracy on false information than those that do not; and a single instruction can reverse the failure across verb families. Mechanistically, models attend more to false beliefs they fail to confirm, but suppressing this attention at decoding time recovers accuracy only partially and only in some models, calling for future work on intervention methods. Our findings clarify prior results and show how fact-checking, a generally desirable behavior, can interfere with belief tracking in LLMs. Our code is available at https://github.com/ngqm/belief-fact-phrasing.
摘要:人類在日常交流中自然地形成和表達信念,例如「我認為答案是3」或「我想這是對的」。這些信念不可避免地與事實和知識交織在一起,使得同時處理它們的能力對大型語言模型(LLMs)來說變得可取,因為它們在面向用戶的環境中越來越多地被部署。先前的研究顯示,即使是能幹的LLMs在承認基於錯誤信息的用戶信念方面也存在系統性的弱點。我們將這一評估擴展到18種認識表達下的10個LLMs,發現弱點的大小和方向取決於用來表達信念的動詞,事實信息與虛假信息之間的準確性差距從「我模糊地記得」的+50%到「我嚴重懷疑」的-14%不等。我們進一步表明,這一現象源於任務混淆:模型默認檢查基礎主張的事實,覆蓋用戶所表達的信念;明確進行事實檢查的思維鏈在虛假信息上的準確性低於那些不進行檢查的;而單一指令可以逆轉動詞家族中的失敗。在機制上,模型對它們未能確認的虛假信念的注意力更高,但在解碼時抑制這種注意力僅能部分恢復準確性,且僅在某些模型中有效,這呼籲未來對干預方法的研究。我們的發現澄清了先前的結果,並顯示事實檢查這一通常可取的行為如何干擾LLMs中的信念追蹤。我們的代碼可在 https://github.com/ngqm/belief-fact-phrasing 獲得。
An Empirical Study of Reward Specification and Benchmark Reliability in GRPO-based LLM Unlearning
2608.17804v1 by Rubén Balbastre, Juan Manuel Orduña, Mariano Pérez
Practical LLM unlearning is usually evaluated through two objectives: suppress target-specific knowledge and preserve non-target utility. In generative QA, this leaves a third behavior underspecified: when a target-adjacent prompt admits a broader answer without target-specific leakage, the model should answer at that level rather than leak, evade, or refuse. We study this specification problem in a controlled LoRA-GRPO RWKU setting, comparing four reward designs that span lexical suppression, anti-refusal shaping, rubric-based broad answering, and an explicit refusal contrast, with and without SFT warm-up. The experiments show that optimization success is not equivalent to behavioral unlearning: RWKU forget scores, held-out completion audits, terminal training-rollout audits, and training dynamics can point to different conclusions. We trace these disagreements to reward-hacking endpoints, policy-support limits in GRPO, benchmark probes that miss endpoint changes, and rewards that can select broad-topic answering with low semantic leakage during optimization.
摘要:實際的 LLM 忘記通常通過兩個目標來評估:抑制特定目標的知識和保留非目標的效用。
在生成性問答中,這留下了第三種行為未明確規範:當一個與目標相近的提示允許更廣泛的回答而不泄露特定目標時,模型應該在該層次上回答,而不是泄露、逃避或拒絕。
我們在一個受控的 LoRA-GRPO RWKU 設定中研究這個規範問題,比較四種獎勵設計,涵蓋詞彙抑制、反拒絕塑造、基於標準的廣泛回答以及明確的拒絕對比,並且有無 SFT 熱身。
實驗表明,優化成功並不等同於行為上的忘記:RWKU 忘記分數、保留的完成審計、終端訓練回滾審計和訓練動態可能指向不同的結論。
我們將這些分歧追溯到獎勵駭客端點、GRPO 中的政策支持限制、錯過端點變化的基準探針,以及在優化過程中可以選擇廣泛主題回答且語義泄露低的獎勵。
Neuro-symbolic learning over OWL 2 DL via consequence-based compilation to differentiable circuits
2608.17741v1 by Olga Mashkova, Asaad Mohammedsaleh, Fernando Zhapa-Camacho, Robert Hoehndorf
OWL 2 DL ontologies, grounded in the description logic $\mathcal{SROIQ}$, express large knowledge bases in biomedicine and the Semantic Web. Neuro-symbolic (NeSy) learners over description logics either embed the ontology in a continuous space, abandoning classical entailment, or restrict to the Horn fragment $\mathcal{EL}^{++}$, which has a single canonical model. We present Baobab, which compiles a $\mathcal{SROIQ}$ ontology with a finite ABox into a Sentential Decision Diagram (SDD): it saturates a propositional core under a consequence-based calculus and instantiates the remaining $\mathcal{SROIQ}$ features (nominals, number restrictions, and the role axioms) over the active domain. The SDD's evidence-conditioned weighted model count then trains a perception network to recognize real images under partial ABox supervision: on an ontology that exercises every distinctive $\mathcal{SROIQ}$ feature, a CNN learns to read MNIST digits coupled by a successor relation and recovers latent ontology concepts that an independent perception leaves at chance. When the supervision admits several ontology-consistent completions, an independent perception collapses onto one, a reasoning shortcut: we show that a mixture indexed by the query's justifications can represent the calibrated posterior no independent perception can, and that seeding it from the circuit's enumerated completions attains the Bayes-optimal posterior on a real-image MNIST task where single-WMC and learned mixtures (the BEARS-ensemble hypothesis class) do not: to our knowledge the first to characterize and mitigate reasoning shortcuts in a non-Horn description logic. Soundness of the compiler and the representation result are machine-checked in Lean 4. Code is available at https://github.com/bio-ontology-research-group/baobab.
摘要:OWL 2 DL 本體,基於描述邏輯 $\mathcal{SROIQ}$,在生物醫學和語意網中表達大型知識庫。神經符號(NeSy)學習者在描述邏輯上要麼將本體嵌入連續空間,放棄傳統的推理,要麼限制於只有一個典範模型的 Horn 片段 $\mathcal{EL}^{++}$。我們提出了 Baobab,它將具有有限 ABox 的 $\mathcal{SROIQ}$ 本體編譯為句子決策圖(SDD):它在基於結果的計算下飽和一個命題核心,並在活動域上實例化剩餘的 $\mathcal{SROIQ}$ 特徵(名詞、數量限制和角色公理)。SDD 的證據條件加權模型計數然後訓練一個感知網絡,以在部分 ABox 監督下識別真實圖像:在一個行使每個獨特 $\mathcal{SROIQ}$ 特徵的本體上,CNN 學會閱讀與後繼關係相結合的 MNIST 數字,並恢復獨立感知所留下的潛在本體概念。當監督允許多個本體一致的完成時,獨立感知會崩潰到一個,這是一種推理捷徑:我們展示了一種由查詢的正當性索引的混合可以表示經過校準的後驗,而沒有獨立感知可以做到,並且從電路的列舉完成中種子達到在一個真實圖像 MNIST 任務上的貝葉斯最佳後驗,而單一 WMC 和學習的混合(BEARS-ensemble 假設類)則無法做到:據我們所知,這是第一次在非 Horn 描述邏輯中表徵和減輕推理捷徑。編譯器的健全性和表示結果在 Lean 4 中經過機器檢查。代碼可在 https://github.com/bio-ontology-research-group/baobab 獲得。
What Aggregate Scores Miss: Measuring Item-Level Regressions in Commercial LLM API Migrations
2608.17719v1 by Xiaonan Xu, Wenjing Wu
Context: Software systems that depend on commercial large language model APIs must migrate to successor versions when vendors deprecate older models. Migration decisions typically rely on aggregate benchmark scores, which compress heterogeneous item-level behaviour into a single net figure. Objective: We measure what that compression conceals. Method: On three pairwise upgrades in the GPT-5.4 to GPT-5.6 Sol product sequence, we query 900 public benchmark items (graduate-level knowledge, olympiad mathematics, instruction following) 50 times per item per model, classify each item as reliably improved, reliably regressed, practically equivalent, or inconclusive under false-discovery-rate control and a practical-significance threshold, and calibrate the results against a label-permutation null. Results: Across all nine migration-benchmark cells, reliable improvements and reliable regressions coexist. Edges with aggregate gains of up to 7.3 percentage points contain up to 8.3% reliably regressed items; edges with aggregate losses contain up to 10.7% reliably improved items. On the instruction-following benchmark, the gap between strict and loose scoring widens by 3.9 percentage points on the latest migration: a 3.9-point regression under strict scoring shrinks to 0.04 points under loose scoring. Conclusion: Migration decisions based on aggregate scores alone miss substantial bidirectional item-level change. The complete response-level archive and per-item scoring outputs are released.
摘要:背景:依賴商業大型語言模型 API 的軟體系統必須在供應商棄用舊模型時遷移到後繼版本。
遷移決策通常依賴於綜合基準分數,這將異質的項目級行為壓縮為單一的淨數字。
目標:我們測量這種壓縮所隱藏的內容。
方法:在 GPT-5.4 到 GPT-5.6 Sol 產品序列的三次成對升級中,我們對 900 個公共基準項目(研究生級知識、奧林匹克數學、指令遵循)進行每個模型每項 50 次查詢,並根據假發現率控制和實際顯著性閾值將每個項目分類為可靠改進、可靠退步、實際等效或不確定,並將結果與標籤置換無效進行校準。
結果:在所有九個遷移基準單元中,可靠的改進和可靠的退步共存。
具有高達 7.3 個百分點的綜合增益的邊緣包含高達 8.3% 的可靠退步項目;具有綜合損失的邊緣包含高達 10.7% 的可靠改進項目。
在指令遵循基準上,最新遷移中嚴格與寬鬆評分之間的差距擴大了 3.9 個百分點:在嚴格評分下的 3.9 點退步在寬鬆評分下縮小至 0.04 點。
結論:僅根據綜合分數作出的遷移決策忽略了實質的雙向項目級變化。
完整的響應級存檔和每項的評分輸出已發布。
Communicating Credit Risk with Large Language Models: Evaluation of Explanations from Standard and Alternative Data-Based Models
2608.17715v1 by Sahab Zandi, Noah Kostesku, Christophe Mues, María Óskarsdóttir, Cristián Bravo
Credit decisioning is a high-stakes task in which model outputs must be accurate and explainable to support compliant decisions. Although modern credit risk models such as eXtreme Gradient Boosting (XGBoost) and Graph Neural Networks (GNNs) improve predictive performance, their explanations are often too technical for stakeholders creating communication gaps that can shape approvals, denials, and fairness judgments. We examine whether Large Language Models (LLMs) can serve as explanation layers that translate post-hoc explanation artefacts into stakeholder-appropriate risk narratives. Using Freddie Mac single-family loan-level data, we develop three pipelines: standard tabular (XGBoost + SHAP), and two with alternative data, a pure network-based (GNN + GNNExplainer), and a bimodal one (combining tabular and network data). We generate narratives with three LLM configurations: a small fine-tuned LLM (Gemma 3 4B), a large fine-tuned LLM (DeepSeek R1 70B), and a zero-shot commercial LLM (Gemini 2.5). Explanation quality is evaluated through automated checks across all pipelines and a human study of bimodal explanations comparing credit risk professionals and non-professionals on eight decision-relevant dimensions. We have three main findings. First, the pipeline accounts for higher variance in evidence-grounding scores than the language model, meaning that the binding constraint on explanation quality is the evidence representation, not the model used. Second, the explanation narratives reliably name the influential factors but are less reliable when stating the direction of influence, which may be consequential for adverse-action communication. Finally, professionals apply stricter evidentiary standards than non-professionals. We discuss implications for the governance of risk models, including deployment considerations and the value of domain-aligned LLMs in regulated credit settings.
摘要:信用決策是一項高風險的任務,其中模型輸出必須準確且可解釋,以支持合規的決策。儘管現代信用風險模型如極端梯度提升(XGBoost)和圖神經網絡(GNNs)提高了預測性能,但它們的解釋往往對利益相關者來說過於技術性,造成溝通差距,這可能影響批准、拒絕和公平性判斷。我們檢視大型語言模型(LLMs)是否可以作為解釋層,將事後解釋產物轉化為適合利益相關者的風險敘事。使用Freddie Mac的單戶貸款數據,我們開發了三個管道:標準表格(XGBoost + SHAP),以及兩個使用替代數據的管道,一個是純基於網絡的(GNN + GNNExplainer),另一個是雙模的(結合表格和網絡數據)。我們使用三種LLM配置生成敘事:一個小型微調LLM(Gemma 3 4B),一個大型微調LLM(DeepSeek R1 70B),以及一個零樣本商業LLM(Gemini 2.5)。通過對所有管道的自動檢查以及對雙模解釋的人工研究,我們評估了解釋質量,並比較了信用風險專業人員和非專業人員在八個與決策相關的維度上的表現。我們有三個主要發現。首先,該管道在證據基礎分數的變異性上比語言模型更高,這意味著解釋質量的約束是證據表示,而不是所使用的模型。其次,解釋敘事可靠地命名了影響因素,但在陳述影響方向時可靠性較低,這對於不利行動的溝通可能具有重要意義。最後,專業人士應用的證據標準比非專業人士更為嚴格。我們討論了風險模型治理的影響,包括部署考量和在受監管的信用環境中領域對齊的LLMs的價值。
GraphWake: Group Polarization via Memory-Mediated Polarization Cascade in LLM-Agent Communities
2608.17665v1 by Haoran Bu, Zejian Chen, Litian Zhang, Xi Zhang
LLM-driven agents can autonomously exchange opinions on online platforms and form communities. Such agent-operated social platforms raise a new security concern: attackers may manipulate agents to induce group polarization. Existing methods manipulate agent prompts or construct echo chambers, both of which are difficult to realize in practice. We therefore formulate a new threat, Memory-Mediated Polarization Cascade, which uses agent memory as a persistence channel and public discussion as a propagation channel. This threat contains three stages. During exposure and memory retention, the attacker exposes a small set of target agents to arguments that reinforce their respective stated stances. The targets' memory systems then process and retain these arguments. During retrieval and reproduction, a shared stance-neutral discussion cues the targets to retrieve and reproduce their respective retained arguments. During iterative propagation, untreated agents influenced by the reproduced arguments restate and spread them. We instantiate this threat in GraphWake with three components: (i) stance-support argumentation knowledge graphs construct knowledge-based arguments; (ii) axiom-oriented triple selection distills them for reliable retention and reproduction; and (iii) stance-neutral memory cueing triggers concurrent retrieval and reproduction, initiating propagation. Experiments across multiple discussions and memory systems show that GraphWake substantially increases group polarization. These findings reveal a community-level polarization risk.
摘要:LLM 驅動的代理可以在在線平台上自主交換意見並形成社群。這種代理操作的社交平台引發了一個新的安全問題:攻擊者可能操縱代理以誘發群體極化。現有的方法操縱代理提示或構建回音室,這兩者在實踐中都難以實現。因此,我們提出了一種新的威脅,記憶介導的極化級聯,它利用代理記憶作為持久性通道,公共討論作為傳播通道。這一威脅包含三個階段。在暴露和記憶保留期間,攻擊者將一小組目標代理暴露於強化其各自表述立場的論點中。目標的記憶系統隨後處理並保留這些論點。在檢索和再現期間,共享的中立立場討論提示目標檢索並再現其各自保留的論點。在迭代傳播期間,受到再現論點影響的未處理代理重述並擴散這些論點。我們在 GraphWake 中實現了這一威脅,包含三個組件:(i)立場支持的論證知識圖構建基於知識的論點;(ii)公理導向的三元組選擇提煉它們以實現可靠的保留和再現;以及(iii)立場中立的記憶提示觸發同時檢索和再現,啟動傳播。多次討論和記憶系統的實驗顯示,GraphWake 顯著增加了群體極化。這些發現揭示了社群層面的極化風險。
Graph Surgery and the Do-Operator: A Precise Correspondence for Acyclic Structural Causal Models
2608.17634v1 by Satpreet Makhija
The $\operatorname{do}$-operator is described graphically by deleting arrows into its targets and functionally by replacing their mechanisms with constants. To call these operations equivalent is not yet a mathematical statement: one returns a graph and remembers only the targets, whereas the other returns mechanisms and also remembers the imposed values. We make a dependency-level comparison precise for deterministic acyclic structural causal models with finitely many endogenous variables. If $\operatorname{Graph}(F)$ extracts the dependencies of a mechanism family $F$, our main theorem is $\operatorname{Graph}(F^ι)=\operatorname{Surg}(\operatorname{Graph}(F),T_ι)$. Thus replacing target mechanisms removes exactly the dependencies removed by graph surgery. For a model $M=(G,F)$ whose graph may contain unused arrows, we characterize when the same equality holds with $G$ in place of $\operatorname{Graph}(F)$; it holds for every intervention exactly when $G$ records the dependencies of $F$ exactly. We then define the intervened model, characterize its run, show how sequential interventions combine, and prove that an outcome depends only on interventions at its actual dependency ancestors.
摘要:$\operatorname{do}$-運算子在圖形上通過刪除指向其目標的箭頭來描述,而在功能上則通過用常數替換其機制來描述。將這些操作稱為等價尚未形成數學陳述:一個返回圖形並僅記住目標,而另一個返回機制並同時記住施加的值。我們對具有有限內生變量的確定性非循環結構因果模型進行依賴層級的精確比較。如果 $\operatorname{Graph}(F)$ 提取機制家族 $F$ 的依賴關係,我們的主要定理是 $\operatorname{Graph}(F^ι)=\operatorname{Surg}(\operatorname{Graph}(F),T_ι)$。因此,替換目標機制正好去除了圖形手術所去除的依賴關係。對於一個模型 $M=(G,F)$,其圖形可能包含未使用的箭頭,我們描述何時同樣的等式在 $G$ 代替 $\operatorname{Graph}(F)$ 時成立;當且僅當 $G$ 精確記錄 $F$ 的依賴關係時,它成立。我們然後定義干預模型,描述其運行,展示如何結合序列干預,並證明結果僅依賴於其實際依賴祖先的干預。
Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges
2608.17605v1 by Syeda Faiza Ahmed, Zien Sheikh Ali, Hunzalah Hassan Bhatti, Firoj Alam, Shammur Absar Chowdhury
Conversational AI is moving beyond isolated text prompts toward sustained, multimodal interaction. In real conversations, users clarify goals, revise requests, interrupt responses, switch topics, and introduce new evidence while expecting systems to preserve context across turns. This makes multi-turn dialogue a distinct challenge requiring systems to maintain and update memory, ground responses across modalities, tools, and external knowledge, and adapt across languages and cultures. This study reviews multi-turn conversational AI across text-only dialogue, AudioLLMs and speech-native systems, multimodal and omni-modal systems, and tool-augmented agents. We organize the literature around datasets and benchmarks, modeling paradigms, training strategies, evaluation setups, and cross-cutting challenges. Our analysis shows that support for multiple modalities has advanced faster than the ability to sustain coherent interaction across a session. Despite stronger capabilities to perceive, speak, and act across modalities, current systems still struggle with persistent memory, cross-turn grounding, full-duplex interaction, robust evaluation, and cultural alignment. We conclude with a research agenda for systems that can remember, revise, ground, speak, listen, act, and adapt across turns, modalities, and cultures. (https://github.com/faiza-sfa/multiturn-conversational-ai-survey)
摘要:對話式人工智慧正在超越孤立的文字提示,朝向持續的多模態互動發展。在真實的對話中,用戶會澄清目標、修訂請求、打斷回應、切換主題並引入新證據,同時期望系統能在不同回合中保持上下文。這使得多回合對話成為一個獨特的挑戰,要求系統維持和更新記憶,跨模態、工具和外部知識進行回應的基礎,並在語言和文化之間進行適應。本研究回顧了文本對話、AudioLLMs 和語音原生系統、多模態和全模態系統以及工具增強代理的多回合對話式人工智慧。我們根據數據集和基準、建模範式、訓練策略、評估設置和跨領域挑戰來組織文獻。我們的分析顯示,對多模態的支持發展得比在一個會話中持續一致互動的能力更快。儘管在感知、說話和跨模態行動方面的能力增強,當前的系統仍然在持久記憶、跨回合基礎、全雙工互動、穩健評估和文化對齊方面面臨挑戰。我們以一個研究議程作結,旨在開發能夠記住、修訂、基礎、說話、聆聽、行動和在回合、模態和文化之間適應的系統。(https://github.com/faiza-sfa/multiturn-conversational-ai-survey)
tinyDSM: A Framework for Skill Modeling and Development for Resource-Constrained Millirobots
2608.17596v1 by Markus D. Kobelrausch, Michael Miedler, Axel Jantsch
In this study, we investigate developmental mechanisms that enable small, resource-constrained systems such as cm-sized millirobots to autonomously explore, learn, and adapt their capabilities throughout their lifespan. Reinforcement learning algorithms guide the agent's skill acquisition and adaptation through the interplay of our proposed tinyDSM, which integrates intrinsic motivation and fitness-based assessment. We strive for minimal, hard-wired skills while encouraging the open-ended development of new skills. A key emphasis in our approach is to encode minimal a-priori general knowledge, which serves as a foundational starting point for the system as it further learns system-specific dependencies from the initial knowledge provided. Thus, by design, our approach attempts to cover very generic application domains. The methodology is based on (a) developmental mechanism with intrinsic motivation, and (b) a cognitive architecture (knowledge, reasoning, learning), while (c) utilizing minimal resources. It uses a hierarchical knowledge graph and kinematic reasoners to model and evaluate simple and advanced motion related skills. In our experiments, we use a resource-constrained millirobot with a volume of 36 cm^3 with a Raspberry Pi Pico 32-bit microcontroller (RP2040) that integrates all described features and capabilities except the camera system in 9 kB. Starting with learning the most elementary motor skills the millirobot autonomously progresses from simple linear and angular movements to complex geometric patterns within 15 minutes. To complement the physical experiments, we perform a simulation-based analysis that enables systematic comparisons across learning algorithms and intrinsic motivation parameters.
摘要:在本研究中,我們探討使小型資源受限系統(如厘米級的微型機器人)能夠自主探索、學習和適應其能力的發展機制。強化學習算法通過我們提出的tinyDSM的相互作用來指導代理的技能獲得和適應,該系統整合了內在動機和基於適應度的評估。我們追求最小的硬連接技能,同時鼓勵新技能的開放式發展。我們方法的一個關鍵重點是編碼最小的先驗一般知識,這作為系統進一步從提供的初始知識中學習系統特定依賴的基礎起點。因此,我們的方法設計上試圖涵蓋非常通用的應用領域。該方法論基於(a)具有內在動機的發展機制,以及(b)一種認知架構(知識、推理、學習),同時(c)利用最小資源。它使用層次知識圖譜和運動學推理器來建模和評估簡單和高級運動相關技能。在我們的實驗中,我們使用一個資源受限的微型機器人,其體積為36 cm^3,搭載Raspberry Pi Pico 32位微控制器(RP2040),該微控制器整合了所有描述的功能和能力,除了攝像頭系統外,僅佔用9 kB。從學習最基本的運動技能開始,微型機器人自主地在15分鐘內從簡單的線性和角運動進展到複雜的幾何圖形。為了補充物理實驗,我們進行了一個基於模擬的分析,這使得能夠在學習算法和內在動機參數之間進行系統比較。
Quantifying Risk Under Evolving Uncertainty: Belief-Dependent Robustness for Safe Sequential Decision Making
2608.17574v1 by Deep Kumar Ganguly, Jan Kretinsky
How cautious should an agent be while it is still learning its environment? We propose RATTL (Risk-Adversarial Total-Reward Learning), which ties caution to epistemic uncertainty: the agent holds a Bayesian posterior over unknown dynamics and plans against a Wasserstein ambiguity set whose radius is a monotone function of that posterior. The radius contracts with evidence, so behaviour interpolates continuously between worst-case robustness and risk-neutral total-reward maximization. The design follows the duality underlying the Entropic Value-at-Risk, which converts the choice of a risk level into the choice of an ambiguity radius. We show the resulting planning problem is well posed under transience and compactness conditions, and prove a Safety Sandwich: the RATTL value lies between the uninformed robust value and the full- knowledge optimum, with a gap that vanishes as the posterior concentrates. In a canonical binary-hazard instance, the induced criterion reduces to Conditional Value-at-Risk at a level set by the posterior entropy. A worked example shows the agent deferring the efficient action until a sharp identification threshold. RATTL targets runtime safety for agents, including LLM-based systems, acting under uncertainty.
摘要:代理在學習其環境時應該多謹慎?我們提出了RATTL(風險對抗總回報學習),它將謹慎與認知不確定性聯繫起來:代理對未知動態持有貝葉斯後驗,並根據一個其半徑是該後驗單調函數的Wasserstein模糊集進行規劃。隨著證據的增加,半徑會收縮,因此行為在最壞情況的穩健性和風險中立的總回報最大化之間持續插值。該設計遵循了熵值風險的對偶性,將風險水平的選擇轉化為模糊半徑的選擇。我們顯示,所得到的規劃問題在瞬態和緊湊性條件下是良好定義的,並證明了一個安全三明治:RATTL值介於無信息穩健值和全知最優值之間,當後驗集中時,這一差距消失。在一個典型的二元危險實例中,所引入的標準簡化為在後驗熵設定的水平下的條件風險價值。一個具體的例子顯示,代理在達到明確識別閾值之前推遲了有效行動。RATTL針對在不確定性下行動的代理,包括基於LLM的系統,目標是運行時安全。
Code as Representation: A Compilable Parsing Paradigm for Academic Documents
2608.17550v1 by Rihui Jin, Jun Wang, chengyuan zhu, Liang Mingyu, Yue Gao, Li Yunxuan, Kuicai Dong, Guilin Qi, Lin Ren, Yongrui Chen, Xinbang Dai, Jiaqi Li, Tongtong Wu, Gholamreza Haffari
Academic papers are a primary carrier of scientific knowledge, yet most of this knowledge remains locked in PDFs that are optimized for human reading rather than machine use. For Multimodal Large Language Models (MLLMs), the core challenge is not only perception, but representation: scientific pages interleave text with Structured Academic Elements (SAEs) such as tables, formulas, charts, and pseudocode, whose structure, data, and logic are poorly preserved by common surrogates like Markdown. We therefore propose Compilable Academic Document Parsing (CADP), a paradigm that reconstructs a full page as contextual \LaTeX{} plus executable Python, so that structure-preserving elements and executable chart representations can be reconstructed, recompiled, and directly verified against the source page. To support this setting, we introduce CADP-Bench, an expert-verified benchmark of full academic pages containing tightly coupled text and multiple SAE types, evaluated through a re-injection compilation protocol. We further study current capabilities using SOTA MLLMs and an exploratory multi-agent baseline that incorporates common agentic techniques. Results show that even frontier models still struggle to produce high-fidelity executable reconstructions, highlighting substantial room for improvement in structure-aware scientific document parsing. CADP-Bench is released for future research.
摘要:學術論文是科學知識的主要載體,但大部分這些知識仍然鎖定在優化為人類閱讀而非機器使用的PDF中。對於多模態大型語言模型(MLLMs)來說,核心挑戰不僅在於感知,還在於表徵:科學頁面將文本與結構化學術元素(SAEs)交錯,如表格、公式、圖表和偽代碼,其結構、數據和邏輯在常見的替代品如Markdown中保存得很差。因此,我們提出可編譯學術文檔解析(CADP),這是一種將整個頁面重建為上下文 \LaTeX{} 加上可執行的Python的範式,以便結構保留的元素和可執行的圖表表示可以被重建、重新編譯並直接與源頁面進行驗證。為了支持這一設置,我們引入CADP-Bench,一個經專家驗證的完整學術頁面基準,包含緊密耦合的文本和多種類型的SAE,通過重新注入編譯協議進行評估。我們進一步研究使用SOTA MLLMs的當前能力以及一個探索性的多代理基準,該基準結合了常見的代理技術。結果顯示,即使是最前沿的模型仍然難以產生高保真度的可執行重建,突顯出結構感知的科學文檔解析有很大的改進空間。CADP-Bench已經釋出以供未來研究使用。
CoAL-RAG: A Complexity-Aware Legal Retrieval-Augmented Generation Method
2608.17536v1 by Jin Su, Zhuofeng Zhao, Huanhuan Wang, Hao Chen
Legal consultation questions exhibit multi-level complexity. A single retrieval strategy often leads to over-reasoning for simple questions and poor interpretability for complex ones, making it difficult to meet the requirements for both answer quality and efficiency in high-risk scenarios. To address this issue, this paper proposes CoAL-RAG, a complexity-aware legal retrieval-augmented generation method, which constructs a multi-dimensional evaluation mechanism based on question essence'' andretrieval consistency'' to enable adaptive routing of retrieval strategies. First, the reasoning demand is quantified according to the logical structure of the question. Then, the discrepancy between semantic retrieval and keyword retrieval is utilized to indirectly reflect problem complexity, thereby selecting the most appropriate retrieval strategy and dynamically filtering contextual information. Experimental results demonstrate that the proposed method significantly outperforms baseline models not only on Chinese legal benchmarks (SocialLawQA, LawBench) but also demonstrates strong cross-jurisdictional generalization on English datasets (LexGLUE, CaseHold). Specifically, on Chinese datasets, the BLEU score improves by 42.5\% and ROUGE-L reaches 3.6 times that of knowledge graph-based methods. On English benchmarks, CoAL-RAG maintains highly competitive accuracy, achieving an optimal balance between generation quality, deep logical reasoning, and system efficiency across different legal systems.
摘要:法律諮詢問題展現出多層次的複雜性。單一的檢索策略常常導致對簡單問題的過度推理,以及對複雜問題的可解釋性差,使得在高風險情境中難以滿足答案質量和效率的要求。為了解決這個問題,本文提出了 CoAL-RAG,一種具複雜性意識的法律檢索增強生成方法,該方法基於「問題本質」和「檢索一致性」構建了一個多維評估機制,以實現檢索策略的自適應路由。首先,根據問題的邏輯結構量化推理需求。然後,利用語義檢索與關鍵字檢索之間的差異,間接反映問題的複雜性,從而選擇最合適的檢索策略並動態過濾上下文信息。實驗結果表明,所提出的方法在中國法律基準(SocialLawQA、LawBench)上顯著超越基線模型,並且在英語數據集(LexGLUE、CaseHold)上展現出強大的跨法域泛化能力。具體而言,在中國數據集上,BLEU 分數提高了 42.5\%,而 ROUGE-L 達到知識圖譜方法的 3.6 倍。在英語基準上,CoAL-RAG 維持了高度競爭的準確性,在不同法律系統中實現生成質量、深度邏輯推理和系統效率之間的最佳平衡。
When to Review: Spaced Repetition for Continual Pre-Training of Language Models
2608.17530v1 by Alankar Atreya, Devesh Batra, Yoages Kumar Mantri, Geremy Bantug, Greig A Cowan, Raad Khraishi
Continual pre-training of large language models must acquire new information without erasing old knowledge. Existing replay methods often choose a global old/new mixture and sample uniformly, ignoring that examples differ in how quickly they are forgotten. We formulate continual pre-training as adaptive review scheduling: the training loop should decide not only how much history to replay, but which examples should return at each step. We introduce Spaced Repetition Training (SRT), a continual learning framework inspired by cognitive science, which schedules sample-rehearsal using the SuperMemo-2 (SM-2) algorithm. SRT maintains per-example review state, maps per-example perplexity to a recall-quality signal, and schedules historical examples for retention and new examples for consolidation while leaving the model, objective, and optimizer unchanged. On temporally separated Wikipedia and code corpora, SRT improves the stability-plasticity trade-off, recovering 5 to 37 percentage points of old-knowledge accuracy lost by naive continual pre-training across model scales while preserving or improving new-knowledge acquisition. At larger scale, SRT preserves broad benchmark performance that naive continual pre-training and uniform replay substantially degrade. Experiments with vision and tabular data further suggest that the scheduling principle extends beyond language when paired with an appropriate recall signal.
摘要:持續的預訓練大型語言模型必須在不抹去舊知識的情況下獲取新信息。現有的重播方法通常選擇一個全局的舊/新混合並均勻抽樣,忽略了示例在被遺忘的速度上存在差異。我們將持續預訓練公式化為自適應回顧排程:訓練循環應決定不僅是重播多少歷史,還有每一步應該返回哪些示例。我們引入了間隔重複訓練(SRT),這是一個受認知科學啟發的持續學習框架,使用 SuperMemo-2 (SM-2) 算法來排程樣本重複。SRT 維持每個示例的回顧狀態,將每個示例的困惑度映射到回憶質量信號,並在保留模型、目標和優化器不變的情況下,為保留歷史示例和鞏固新示例進行排程。在時間上分隔的維基百科和代碼語料庫上,SRT 改善了穩定性與可塑性的權衡,恢復了由天真的持續預訓練在各模型規模上損失的 5 到 37 個百分點的舊知識準確率,同時保留或改善了新知識的獲取。在更大規模下,SRT 保持了廣泛的基準性能,而天真的持續預訓練和均勻重播則大幅降低了這一性能。對於視覺和表格數據的實驗進一步表明,當與適當的回憶信號配對時,排程原則超越了語言的範疇。
Effects of Answer Format Variation on Gender Bias in Large Language Models
2608.17516v1 by Ksenia Merzlyakova, Sebastian Padó, Franziska Weeber
Gender bias or other social biases in large language models (LLMs) are frequently evaluated with question answering or survey benchmarks where the LLM needs to give a response in a predefined answer format. It is well known in survey science that the answer format has a substantial impact on answers, just as LLMs are sensitive to the prompt wording. However, to our knowledge it has not been studied yet how changes in answer format impact the measurement of gender bias in LLMs and their alignment with human response distributions. We evaluate three instruction-tuned models on the BBQ benchmark and OpinionQA survey data across closed-ended, Likert-scaled and open-ended formats, comparing bias measurement and distributional alignment under otherwise identical conditions. We find that answer format does substantially alter measured outcomes, including reversals in order rankings. These differences arise because each format elicits distinct response behaviours, such as forced-choice selection, scale-based distributions and refusal in free-text generation. Our findings highlight the importance of treating answer format as a substantive component of LLM evaluation and motivate multi-format designs for more robust model assessment.
摘要:性別偏見或其他社會偏見在大型語言模型(LLMs)中的評估,通常使用問答或調查基準,其中LLM需要以預定的答案格式給出回應。調查科學中已知答案格式對答案有重大影響,就像LLMs對提示措辭敏感一樣。然而,據我們所知,尚未研究答案格式的變化如何影響LLMs中性別偏見的測量及其與人類回應分佈的一致性。我們在BBQ基準和OpinionQA調查數據上評估了三個經過指令調整的模型,並比較了在封閉式、Likert量表和開放式格式下的偏見測量和分佈一致性,條件則保持一致。我們發現答案格式確實顯著改變了測量結果,包括排序排名的逆轉。這些差異的產生是因為每種格式引發了不同的回應行為,例如強制選擇、基於量表的分佈和在自由文本生成中的拒絕。我們的發現強調將答案格式視為LLM評估的實質性組成部分的重要性,並促使多格式設計以進行更穩健的模型評估。
Beyond FLOPs: Energy-Aware Knowledge Distillation for Sustainable LLMs on Code-Related Task
2608.17515v1 by Enrique Barba Roque, Luís Cruz, Annibale Panichella
Background: Large Language Models (LLMs) are increasingly being applied to Software Engineering (SE) tasks, achieving high accuracy across problems such as clone detection, vulnerability prediction, and code summarization. However, their high computational demands and energy consumption raise sustainability concerns and hinder their use on consumer hardware and resource-constrained platforms. A common way to report the computational cost of an LLM in the literature and industry is to use the number of Floating Point Operations (FLOPs) required to perform a pass over the network. Aims: This paper investigates the implications of energy-aware knowledge distillation for SE, aiming to improve model efficiency while maintaining performance and to determine whether FLOPs is a reliable energy-aware metric. Method: We conduct a controlled experiment using Morph, a Many-Objective Optimization-based distillation methodology, to empirically examine whether FLOPs accurately reflect energy consumption in Clone Detection and Vulnerability Prediction tasks. We extend this methodology to include energy-surrogate models that directly estimate CPU and GPU energy consumption during optimization, and we apply Morph to generative tasks using CodeT5+ for code summarization. Results: Our results show that FLOPs is not always a reliable indicator of energy consumption, and better results can be achieved by using energy-surrogate models. Distilled student models can reduce inference energy consumption by up to 90\% and memory usage by 86\%, with only modest accuracy trade-offs. Conclusions: Energy-aware knowledge distillation when guided by direct energy surrogates rather than FLOPs can improve the energy consumption, sustainability, and deployability of LLMs for SE applications, enabling efficient models on consumer hardware.
摘要:背景:大型語言模型(LLMs)越來越多地應用於軟體工程(SE)任務,在克隆檢測、漏洞預測和程式碼摘要等問題上達到了高準確率。
然而,它們的高計算需求和能量消耗引發了可持續性問題,並阻礙了它們在消費者硬體和資源受限平台上的使用。
在文獻和業界中,報告LLM計算成本的常見方法是使用執行一次網絡所需的浮點運算次數(FLOPs)。
目標:本文探討了對SE進行能量感知知識蒸餾的影響,旨在提高模型效率的同時保持性能,並確定FLOPs是否是一個可靠的能量感知指標。
方法:我們使用Morph進行了一項受控實驗,這是一種基於多目標優化的蒸餾方法,實證檢驗FLOPs是否準確反映克隆檢測和漏洞預測任務中的能量消耗。
我們擴展了這一方法,納入能量替代模型,這些模型在優化過程中直接估算CPU和GPU的能量消耗,並將Morph應用於使用CodeT5+進行程式碼摘要的生成任務。
結果:我們的結果顯示FLOPs並不總是能可靠指示能量消耗,使用能量替代模型可以獲得更好的結果。
蒸餾的學生模型可以將推理能量消耗降低高達90%,內存使用量降低86%,而準確率僅有適度的折衷。
結論:在直接能量替代模型的指導下,能量感知知識蒸餾可以改善LLMs在SE應用中的能量消耗、可持續性和可部署性,從而使消費者硬體上的模型更加高效。
SGHA: Evidence-Grounded Research Problem Discovery with Local Language Models
2608.17501v1 by Sarvesh Gharat, Junpei Komiyama
Recent efforts toward fully automated AI scientists have demonstrated that language-model agents can generate hypotheses, execute experiments, and draft scientific manuscripts. However, during the early stages of research, when research problems are formulated, these AI scientists often rely heavily on proprietary frontier models. Their proposals are shaped by opaque parametric knowledge and by literature searches conditioned on the proposals themselves. Such knowledge is effectively a black box, and this dependence makes the evidential basis and validity of generated research problems difficult to audit and leaves the process vulnerable to model-specific hallucinations and biases. Furthermore, if proprietary research materials are transmitted to external APIs, the use of these models creates confidentiality, privacy, and data-governance concerns. We introduce the Structural Gap Hypothesis Agent (SGHA), a fully automated, corpus-first research-problem discovery system that runs entirely on a local LLM. SGHA structures a scientific literature corpus into evidence-linked paper objects and a typed evidence graph, detects unresolved structural patterns across papers, screens candidate gaps before formulation, and produces traceable research-problem families. In particular, it is able to output assumptions, objectives, success criteria, and remaining ambiguities. All LLM-based components of SGHA are executed using a locally served open-weight 9B language model, without requiring proprietary frontier-model APIs. We compare SGHA with the AI Scientist-v2 idea formulation module in five machine-learning domains. Our results suggest that explicit corpus structure and evidence-constrained reasoning can support promising, inspectable research-problem formulation without relying on frontier models during generation or verification.
摘要:最近對於完全自動化的AI科學家的努力顯示,語言模型代理可以生成假設、執行實驗並撰寫科學手稿。
然而,在研究的早期階段,當研究問題被形成時,這些AI科學家往往過度依賴專有的前沿模型。
他們的提案受到不透明的參數知識和基於提案本身的文獻搜尋的影響。
這種知識實際上是一個黑箱,而這種依賴使得生成的研究問題的證據基礎和有效性難以審核,並使過程容易受到模型特定的幻覺和偏見的影響。
此外,如果專有研究材料被傳輸到外部API,使用這些模型會產生保密性、隱私和數據治理的問題。
我們介紹了結構性差距假設代理(SGHA),這是一個完全自動化的、以語料庫為首的研究問題發現系統,完全在本地的LLM上運行。
SGHA將科學文獻語料庫結構化為與證據相關聯的論文對象和類型化的證據圖,檢測論文之間未解決的結構模式,在形成之前篩選候選差距,並生成可追溯的研究問題家族。
特別是,它能夠輸出假設、目標、成功標準和剩餘的模糊性。
SGHA的所有基於LLM的組件都是使用本地提供的開放權重9B語言模型執行的,而不需要專有的前沿模型API。
我們將SGHA與AI Scientist-v2的想法形成模塊在五個機器學習領域進行比較。
我們的結果表明,明確的語料結構和基於證據的推理可以支持有前景的、可檢查的研究問題形成,而無需在生成或驗證過程中依賴前沿模型。
SAGE: Self-Evolving Storyboard Skills via Attribution-Guided Rule Evolution
2608.17468v1 by Maolin Ran, Xiaoyang Lu, Jiaqi Liu, Jian Wang, Weiwen Liu, Jianghao Lin, Yong Yu, Weinan Zhang
Storyboards turn screenplays into visual shot plans for automated short drama production. Professional storyboarding relies on tacit directorial expertise and remains an industrial bottleneck. Large language models can automate this step, but methods for supplying directing knowledge face three challenges: (1) Knowledge acquisition: the craft remains implicit in exemplars or must be written manually. (2) Knowledge refinement: authored knowledge is not evaluated against execution outcomes, and opaque generation prevents feedback attribution to the knowledge behind each decision. (3) Knowledge injection: injecting all knowledge exceeds usable context, while manual selection for every narrative group does not scale. We present SAGE (Skill with Attribution-Guided Evolution), a deployed framework that learns, attributes, evolves, and routes directing knowledge from expert demonstrations. SAGE derives rules that are independent of episode content by contrasting each training screenplay with its expert storyboard. During generation, the model records each narrative group's adopted rules. Combining these records with localized feedback enables targeted updates to individual rules. Evolved rules form scenario packages with a routing index, so each group retrieves only a bounded set appropriate to its situation without expert intervention. On 18 test episodes across three genres, SAGE scored 77.8 on a rubric validated by experts, versus 77.1 for professional directors. Deployed for 14 days on Virtual Film Studio, SAGE produced 1,344 narrative group outputs; 87.2 percent were accepted without substantive edits, and the production team recorded over 83 percent less authoring time per episode. We release PROSE, the first public dataset pairing screenplays with storyboards by professional directors across 68 episodes: https://github.com/creDreams/PROSE.
摘要:故事板將劇本轉化為自動化短劇製作的視覺拍攝計劃。專業的故事板製作依賴於隱性導演專業知識,並且仍然是產業瓶頸。大型語言模型可以自動化這一步驟,但提供導演知識的方法面臨三個挑戰:(1)知識獲取:這項技藝仍然隱含於範例中或必須手動撰寫。(2)知識精煉:創作的知識未能根據執行結果進行評估,且不透明的生成過程阻礙了對每個決策背後知識的反饋歸因。(3)知識注入:注入所有知識超出了可用的上下文,而對每個敘事群體進行手動選擇則無法擴展。我們提出了SAGE(具歸因引導演變的技能),這是一個已部署的框架,從專家示範中學習、歸因、演變和路由導演知識。SAGE通過將每個訓練劇本與其專家故事板進行對比,推導出獨立於劇集內容的規則。在生成過程中,模型記錄每個敘事群體採用的規則。將這些記錄與本地反饋結合,使得對個別規則的針對性更新成為可能。演變的規則形成具有路由索引的場景包,因此每個群體僅檢索適合其情境的有限集合,而無需專家介入。在三個類型的18個測試劇集中,SAGE在專家驗證的評分標準上得分77.8,而專業導演則為77.1。在虛擬電影工作室部署14天後,SAGE產出了1,344個敘事群體的輸出;87.2%的輸出在未進行實質性編輯的情況下被接受,製作團隊每集的創作時間減少了超過83%。我們發布了PROSE,這是第一個將劇本與專業導演的故事板配對的公共數據集,涵蓋68個劇集:https://github.com/creDreams/PROSE。
Structure-Internalized Rule Language Model for Faithful Knowledge Graph Reasoning
2608.17443v1 by Xingrui Zhuo, Jiapu Wang, Manzong Huang, Gongqing Wu, Xindong Wu
Knowledge Graph Reasoning (KGR) aims to discover latent facts by leveraging the structural evidence available in KGs, posing a challenge to the structural semantic understanding capability of KGR models. Recent studies have demonstrated that Large Language Models (LLMs) can achieve remarkable progress on KGR tasks via flexible in-context learning. However, the inherent representation inconsistency between KG structural context and LLM parametric knowledge remains inadequately addressed. This limitation prevents LLMs from effectively perceiving reasoning evidence that aligns with KG constraints, which undermines both the effectiveness and faithfulness of reasoning. We refer to this problem as reasoning evidence perception drift of LLMs over KGs. To address this problem, we propose a Structure-Internalized Rule Language Model (SIRLM), which centers on structural rule generation to couple the parametric learning of structural knowledge with the faithfulness evaluation of reasoning logic, enabling LLMs to anchor tightly to KG-grounded evidence. Specifically, we first design a Structure-Internalized Rule Generator (SIRG), which incorporates an in-context learning block augmented with a structural relation memory to coordinate structural and parametric knowledge. Furthermore, we equip SIRG with a KG tokenizer based on structural invariance learning and a neuro-symbolic reasoner based on rule-constrained message propagation. These components provide SIRG with learnable structural representations and faithful rule-execution feedback, respectively. Our SIRLM can be seamlessly integrated into standard LLM training paradigms, such as SFT and GRPO. Extensive experiments against 17 state-of-the-art KGR methods on 36 datasets demonstrate the significant superiority of SIRLM.
摘要:知識圖譜推理(KGR)旨在利用知識圖譜中的結構證據來發現潛在事實,這對KGR模型的結構語義理解能力提出了挑戰。最近的研究表明,大型語言模型(LLMs)可以通過靈活的上下文學習在KGR任務上取得顯著進展。然而,知識圖譜的結構上下文與LLM的參數知識之間固有的表示不一致性仍然未得到充分解決。這一限制阻礙了LLMs有效感知與知識圖譜約束相符的推理證據,從而削弱了推理的有效性和可靠性。我們將這個問題稱為LLMs在知識圖譜上的推理證據感知漂移。為了解決這個問題,我們提出了一種結構內化規則語言模型(SIRLM),該模型專注於結構規則生成,以將結構知識的參數學習與推理邏輯的可靠性評估相結合,使LLMs能夠緊密依賴於知識圖譜基礎的證據。具體而言,我們首先設計了一個結構內化規則生成器(SIRG),該生成器包含一個增強了結構關係記憶的上下文學習模塊,以協調結構和參數知識。此外,我們為SIRG配備了一個基於結構不變性學習的知識圖譜標記器和一個基於規則約束消息傳播的神經符號推理器。這些組件分別為SIRG提供了可學習的結構表示和可靠的規則執行反饋。我們的SIRLM可以無縫集成到標準的LLM訓練範式中,如SFT和GRPO。在36個數據集上對17種最先進的KGR方法進行的廣泛實驗顯示了SIRLM的顯著優越性。
Cognitive Graph Intelligence for Adaptive and Robust DDoS Attack Detection in Next Generation Networks
2608.17352v1 by Mohammad Arif Hossain, Yeahia Sarker, Md Jafrin Hossain, Most. Humayra Khanom Rime, Nirwan Ansari
Distributed Denial-of-Service (DDoS) attacks threaten network availability, requiring a cognitive detection process that senses traffic, infers intent, and supports an adaptive response under severe class imbalance and non-stationary conditions. This paper proposes a Graph-based Generative Adversarial Network (GraphGAN) that serves as the cognitive detection engine for this task. GraphGAN captures the relational structure among traffic flows while addressing imbalance through adversarial generation of synthetic samples. Sequential flows are converted into $k$-nearest neighbor graphs using sliding windows to preserve feature-similarity and temporal dependencies among flows. The generator learns the distribution of DDoS attacks to synthesize realistic minority samples, while a Graph Convolutional Network (GCN)-based discriminator distinguishes real from synthetic graph data. A separate GCN classifier, trained on the balanced dataset, performs the final detection decision. Evaluations on four benchmark datasets show that GraphGAN achieves superior accuracy, precision, and recall compared to state-of-the-art approaches, particularly in data-scarce scenarios. By integrating temporal graph construction, adversarial augmentation, and GCN classification, GraphGAN effectively models coordinated attack behaviors and mitigates class imbalance, providing a robust and topology-aware solution for intrusion detection in data-constrained environments.
摘要:分散式拒絕服務(DDoS)攻擊威脅網絡可用性,這需要一個認知檢測過程來感知流量、推斷意圖,並在嚴重的類別不平衡和非穩態條件下支持自適應響應。本文提出了一種基於圖的生成對抗網絡(GraphGAN),作為此任務的認知檢測引擎。GraphGAN 捕捉流量流之間的關係結構,同時通過對抗生成合成樣本來解決不平衡問題。連續流量被轉換為 $k$-最近鄰圖,使用滑動窗口來保留流量之間的特徵相似性和時間依賴性。生成器學習 DDoS 攻擊的分佈,以合成現實的少數樣本,而基於圖卷積網絡(GCN)的判別器則區分真實與合成的圖數據。另一個在平衡數據集上訓練的 GCN 分類器執行最終檢測決策。在四個基準數據集上的評估顯示,GraphGAN 在準確性、精確度和召回率方面優於最先進的方法,特別是在數據稀缺的情況下。通過整合時間圖構建、對抗增強和 GCN 分類,GraphGAN 有效地建模協調攻擊行為並減輕類別不平衡,為數據受限環境中的入侵檢測提供了一個強健且考慮拓撲的解決方案。
DeAR: Decentralized Agentic Reasoning via Capability Grounding and Collaborative Thought Navigation
2608.17282v1 by Xing Wei, Changmeng Zheng, XiaoYong Wei, Xiufen Ye, Qing Li
Existing agentic reasoning systems typically rely on centralized protocols. This design introduces routing bottlenecks and static role allocations that often fail when handling complex multimodal queries. We propose DeAR (Decentralized Agentic Reasoning), a framework that shifts from central control to autonomous peer-to-peer collaboration. DeAR is built on three mechanisms: (1) decentralized capability grounding for query-dependent agent specialization, (2) thought map navigation for targeted peer interactions, and (3) topology update for adaptive error correction. Evaluations across 9 diverse multimodal reasoning and text-based QA benchmarks indicate that DeAR consistently outperforms recent baseline methods, validating that decentralized and adaptive collaboration among agents enhances accuracy in knowledge-intensive reasoning tasks. The source code will be available at https://open_upon_acceptance.
摘要:現有的代理推理系統通常依賴於集中式協議。這種設計引入了路由瓶頸和靜態角色分配,當處理複雜的多模態查詢時,往往會失效。我們提出了 DeAR(去中心化代理推理),這是一個從中央控制轉向自主點對點協作的框架。DeAR 建立在三個機制之上:(1)去中心化的能力基礎,以實現依賴查詢的代理專業化,(2)思維地圖導航,以便進行有針對性的同行互動,以及(3)拓撲更新,以進行自適應錯誤修正。在 9 個多樣化的多模態推理和基於文本的問答基準上的評估表明,DeAR 始終優於近期的基準方法,驗證了代理之間去中心化和自適應的協作能提高知識密集型推理任務的準確性。源代碼將在 https://open_upon_acceptance 提供。
ASI-Bench: At the Dawn of Artificial Superintelligence
2608.17271v1 by Junwei Zhou, Zhen Sun, Binyu Li, Jiangyu Zhou, Yuexi Pan, Hengyu Wang, Honghe Ren, Xiaohan Jia, Xueyang Zhou, Xiaoyu Cao, Yongchao Chen, Yuanning Feng, Junhao Wu, Cheng Zhang, Sijia Chen, Haoyu Xue, Chengsong You, Huan Wang, Koutian Wu, Peigan Gao, Jiakun Wu, Wenzhe Li, Ergan Shang, Qingyuan Zheng, Jingjing Zhou, Ruixuan Jia, Yan Xu, Hongrui Zhang, Xiao-Han Ma, Zhengxiang Cheng, Yuexing Hao, Liting Mai, Xianglin Ji, Wenjun Zhang, Zhuofan Chen, Yixiao Huang, Chi Wang, Wenyue Hua, Yilun Hao, Yuantao Zhai, Ziyan Zhao, Jingyan Xie
Artificial superintelligence (ASI) requires AI to move beyond mastering existing knowledge toward exploring the unknown, creating new knowledge, and turning new ideas into verifiable results. However, the capabilities of today's AI systems are still largely built on learning, compressing, and applying existing human knowledge. Accordingly, existing benchmarks primarily test whether AI can produce correct answers based on learned knowledge, or whether it can complete tasks under extensive human guidance. We therefore introduce ASI-Bench, the first benchmark to jointly evaluate AI systems' capabilities of innovative exploration and autonomous scientific execution across general research domains, and the first to progressively withdraw human methodological guidance within the same research project to test how far AI can proceed on its own. Built by over 40 experts with the cost of 31,000+ human hours, ASI-Bench contains 60 project-level research tasks across 11 scientific domains and progressively reduces methodological guidance to test whether AI can independently select methods, conduct research, and produce verifiable results. All tasks undergo expert review, AI-assisted auditing, sandbox execution, and scorer validation. Across 18 state-of-the-art agent--model configurations, the average score drops from 50.91 with full methodological guidance to 29.10 with only the method specified and 26.62 when agents must determine the method themselves. This sharp decline shows that current systems remain heavily dependent on human guidance and are still far from autonomously conducting end-to-end, project-level scientific research. ASI-Bench is open to the world. We invite researchers and builders everywhere to contribute new tasks, challenge the limits of today's AI, and help accelerate humanity's collective path toward artificial superintelligence at https://asibench.apexin.ai/submit.
摘要:人工超智能(ASI)要求人工智慧超越掌握現有知識,朝向探索未知、創造新知識,並將新想法轉化為可驗證的結果。然而,當今人工智慧系統的能力仍然主要建立在學習、壓縮和應用現有人類知識的基礎上。因此,現有的基準主要測試人工智慧是否能根據學習到的知識產生正確答案,或者是否能在廣泛的人類指導下完成任務。因此,我們推出了 ASI-Bench,這是第一個共同評估人工智慧系統在一般研究領域中創新探索和自主科學執行能力的基準,並且是第一個在同一研究項目中逐步撤回人類方法論指導以測試人工智慧能獨立進行多遠的基準。ASI-Bench 由超過 40 位專家建造,耗費超過 31,000 小時的人力,包含 60 個跨 11 個科學領域的項目級研究任務,並逐步減少方法論指導,以測試人工智慧是否能獨立選擇方法、進行研究並產生可驗證的結果。所有任務都經過專家審查、人工智慧輔助審核、沙盒執行和評分者驗證。在 18 種最先進的代理-模型配置中,平均得分從全方法論指導下的 50.91 降至僅指定方法的 29.10,當代理必須自行確定方法時則降至 26.62。這一急劇下降顯示當前系統仍然在很大程度上依賴於人類指導,並且仍然遠未能自主進行端到端的項目級科學研究。ASI-Bench 向全世界開放。我們邀請各地的研究者和建設者貢獻新任務,挑戰當今人工智慧的極限,並幫助加速人類朝向人工超智能的共同道路,網址為 https://asibench.apexin.ai/submit。
Understanding Curriculum Learning in Large Language Models via Cross-Difficulty Optimization Dynamics
2608.17268v1 by Zhikai Ding, Ziyi Ye
Curriculum learning has been widely adopted in the post-training of large language models by organizing training data from easy to hard. However, its effectiveness varies substantially across reasoning tasks, suggesting that no single curriculum is universally optimal and raising a fundamental question: what determines when curriculum learning works? In this paper, we answer this question by analyzing the optimization dynamics induced by different curriculum schedules. We show that the transfer relationship between different difficulty levels characterizes the optimization dynamics induced by curriculum learning, which in turn explains the effectiveness of different curriculum schedules, and formalize this relationship as Relative Transfer, a principled measure of cross-difficulty knowledge transfer. Based on this measurement, we derive Transfer-aware Dynamic Curriculum Sampling (TDCS), which dynamically adjusts the sampling distribution according to the estimated transfer relationship throughout training. Extensive experiments on multiple reasoning benchmarks demonstrate that TDCS consistently outperforms representative scheduling strategies across different tasks, model scales, and training paradigms. More importantly, our work provides a unified optimization-based explanation of curriculum learning through cross-difficulty transfer.
摘要:課程學習已被廣泛應用於大型語言模型的後訓練,通過將訓練數據從簡單到困難進行組織。
然而,它在推理任務中的有效性差異很大,這表明沒有單一的課程是普遍最佳的,並提出了一個根本性問題:什麼決定了課程學習的有效性?
在本文中,我們通過分析不同課程安排所引起的優化動態來回答這個問題。
我們展示了不同難度級別之間的轉移關係特徵化了課程學習所引起的優化動態,這反過來解釋了不同課程安排的有效性,並將這一關係形式化為相對轉移,這是一種跨難度知識轉移的原則性度量。
基於這一測量,我們推導出轉移感知動態課程抽樣(TDCS),該方法根據整個訓練過程中估計的轉移關係動態調整抽樣分佈。
在多個推理基準上的大量實驗表明,TDCS在不同任務、模型規模和訓練範式中始終優於代表性的排程策略。
更重要的是,我們的工作通過跨難度轉移提供了一個統一的基於優化的課程學習解釋。
Structural Plan-to-Model Conversion with Deterministic Geometry and Guarded Agentic Vision-Language Refinement
2608.17237v1 by Mohammad Talebi-Kalaleh, Qipei Mei
Converting structural framing plans into editable finite-element model drafts remains labor-intensive and prone to transcription error. Existing drawing-understanding systems for building components rely on task-specific trained neural detectors, and language-model agents in structural engineering operate on text or model data rather than the drawing itself. This paper presents, to the authors' knowledge, the first framework applying an agentic vision-language layer to structural component detection and model drafting from framing-plan PDFs, without task-specific detector training or fine-tuning. A deterministic stage extracts primitives, estimates scale by dimension-ratio consensus, recognizes five entity classes with a drafting grammar, and assembles an editable layout. The agentic stage proposes typed corrections constrained by deterministic candidates, operation-specific admission tests, change-level review, and fail-closed transactions. Evaluation used an author-generated benchmark of 100 plans: a development half that informed every rule revision, and a seed-disjoint held-out half generated after the rules froze, evaluated once. All reported scores are end-to-end results of the complete framework on the held-out half. Scale was estimated within 0.1% of the generator reference for every drawing. Recall and precision were 0.922/0.997 for columns, 0.886/0.990 for beams, 1.000/1.000 for walls, 1.000/1.000 for braces, and 1.000/0.964 for openings. A controlled study repeated two corruptions three times on three development drawings. Calibration passed all nine trials; member repair met every strict end-state predicate in five of nine. Guarded review corrected missed framing and false marks within explicit bounds. The held-out half shares the development generator, so the study excludes independently drafted plans, raster evaluation, analytical connectivity, and solver validation.
摘要:將結構框架計劃轉換為可編輯的有限元素模型草稿仍然需要大量人力,並且容易出現轉錄錯誤。現有的建築組件繪圖理解系統依賴於特定任務訓練的神經檢測器,而結構工程中的語言模型代理則基於文本或模型數據,而非繪圖本身。據作者所知,本文提出了第一個將代理視覺-語言層應用於從框架計劃PDF中檢測結構組件和模型草擬的框架,無需特定任務的檢測器訓練或微調。一個確定性的階段提取原始元素,通過尺寸比共識估算比例,識別五種實體類別,並組裝可編輯的佈局。代理階段提出了受限於確定性候選者的類型修正、特定操作的入場測試、變更級別審查和失敗關閉交易。評估使用了一個作者生成的100個計劃的基準:一半用於開發,告知每條規則的修訂,另一半在規則凍結後生成,進行了一次評估。所有報告的分數都是在保留的一半上完整框架的端到端結果。每個繪圖的比例估算在生成參考的0.1%內。柱子的召回率和精確度為0.922/0.997,梁為0.886/0.990,牆為1.000/1.000,支撐為1.000/1.000,開口為1.000/0.964。一項控制研究在三個開發繪圖上重複了兩次損壞,進行了三次。校準通過了所有九次試驗;成員修復在九次中的五次滿足每個嚴格的最終狀態謂詞。受控審查在明確範圍內修正了漏掉的框架和錯誤標記。保留的一半共享開發生成器,因此該研究排除了獨立草擬的計劃、光柵評估、分析連通性和求解器驗證。
Synthesizing Feature Extractors: An Agentic Approach for Algorithm Selection
2608.17170v1 by Hai Xia, Carlos Ansótegui, Stefan Szeider
Algorithm selection for constraint satisfaction problems requires extracting features that capture problem structure. Manually designing feature extractors demands deep domain expertise and quickly becomes a bottleneck when new problem classes appear. We present an automated approach that uses Large Language Models (LLMs) in an agentic check--fix--verify loop to synthesize executable Python scripts that act as interpretable, problem-specific feature extractors. Given a high-level MiniZinc model and an instance, the LLM agent generates code that constructs a typed graph representation and computes structural properties such as graph density, variable clustering, and constraint tightness. We evaluate our approach on three combinatorial problems (vehicle routing, car sequencing, fixed-length error-correcting codes) with a portfolio of five state-of-the-art solvers. The synthesized extractors yield algorithm selectors that consistently outperform both expert-curated mzn2feat features (up to $8.3$ percentage points (pp) test-set accuracy on FLECC) and the best transformer-based trans2feat variants. In the meanwhile, the synthesized feature extractors remain inspectable.
摘要:算法選擇約束滿足問題需要提取捕捉問題結構的特徵。手動設計特徵提取器需要深厚的領域專業知識,並且在新的問題類別出現時很快就會成為瓶頸。我們提出了一種自動化的方法,利用大型語言模型(LLMs)在代理檢查--修正--驗證循環中合成可執行的 Python 腳本,這些腳本充當可解釋的、特定於問題的特徵提取器。給定一個高階的 MiniZinc 模型和一個實例,LLM 代理生成代碼,構建一個類型圖表示並計算結構性質,如圖密度、變量聚類和約束緊湊性。我們在三個組合問題(車輛路由、汽車排序、固定長度糾錯碼)上評估我們的方法,使用五個最先進求解器的組合。合成的提取器產生的算法選擇器在測試集準確率上始終超越專家策劃的 mzn2feat 特徵(在 FLECC 上高達 $8.3$ 個百分點(pp))和最佳的基於Transformer的 trans2feat 變體。與此同時,合成的特徵提取器仍然可供檢查。
Towards Safer RAG: Only Agents Capable of System 2 Thinking may Access Untrusted Documents
2608.17153v1 by Mehrdad Ghassabi
Retrieval-Augmented Generation (RAG) has significantly enhanced the performance of large language models (LLMs), yet these systems remain vulnerable to knowledge-poisoning attacks, in which misinformation in retrieved documents can influence the model's final outputs. Notably, an LLM may correctly detect that a document contains incorrect information while nevertheless being influenced by it. Prior work has addressed this vulnerability through the Cordon Principle, which prevents models responsible for final answer synthesis from directly accessing raw evidence. Although effective, this strict isolation can introduce substantial computational overhead. In this work, we propose a refined security principle: only agents capable of deliberative System 2 reasoning may access untrusted documents. To evaluate this principle, we introduce novel metrics that quantify the discrepancy between misinformation detection and downstream influence. We then empirically compare state-of-the-art reasoning language models with standard language models across these metrics. Our results show that reasoning-capable models are substantially more robust to corrupted evidence, without requiring the strict isolation imposed by the Cordon Principle. These findings provide empirical support for our refined principle and suggest a more practical foundation for secure RAG system design.
摘要:檢索增強生成(RAG)顯著提升了大型語言模型(LLMs)的性能,但這些系統仍然容易受到知識污染攻擊,其中檢索到的文件中的錯誤信息可能影響模型的最終輸出。值得注意的是,LLM 可能正確檢測到某個文件包含不正確的信息,但仍然會受到其影響。先前的研究通過 Cordon 原則解決了這一脆弱性,該原則防止負責最終答案合成的模型直接訪問原始證據。儘管有效,但這種嚴格的隔離可能會引入相當大的計算開銷。在本研究中,我們提出了一個精煉的安全原則:只有能夠進行深思熟慮的系統 2 推理的代理才能訪問不受信任的文件。為了評估這一原則,我們引入了新穎的度量標準,以量化錯誤信息檢測與下游影響之間的差異。然後,我們在這些度量標準上,實證比較了最先進的推理語言模型與標準語言模型。我們的結果顯示,具備推理能力的模型對受損證據的魯棒性顯著更強,而無需 Cordon 原則所施加的嚴格隔離。這些發現為我們的精煉原則提供了實證支持,並為安全 RAG 系統設計建議了一個更實用的基礎。
KnowSim: Evaluating Information Calibration in LLM Assistants with User Simulators that Learn
2608.17150v1 by Yoonjoo Lee, Hyoungwook Jin, Tae Soo Kim, Shaoyang Zhang, Philippe Laban, Q. Vera Liao
To effectively collaborate with users on knowledge-intensive tasks, Large Language Models (LLMs) must perform information calibration: matching content to a user's evolving understanding and cognitive capacity. Yet user simulators used to evaluate and train LLMs do not explicitly model user knowledge so they neither produce realistic interactions across knowledge levels nor reflect how interactions unfold as that knowledge evolves. To close this gap, we introduce KNOWSIM, an evaluation framework built around a user simulator that maintains explicit knowledge states, represented as a graph of Information Units with prerequisite relationships, that evolve under update rules grounded in learning theory. KNOWSIM computes three metrics (Knowledge Gain, Delivery Calibration, Cognitive Overload) directly from the knowledge state trajectory, reflecting key mechanistic aspects of information calibration. We validate KNOWSIM against 705 human-AI sessions across two domains, stratified by knowledge level: its rankings align significantly with human judgments (73-74% sign agreement), outperforming three baseline simulators. Applied to 9 LLMs, KNOWSIM reveals that the best model shifts by user knowledge level, revealing aptitude-treatment interactions invisible to standard evaluation.
摘要:為了有效地與用戶在知識密集型任務上合作,大型語言模型(LLMs)必須進行信息校準:將內容與用戶不斷演變的理解和認知能力相匹配。然後,用於評估和訓練LLMs的用戶模擬器並未明確建模用戶知識,因此它們既無法產生跨知識水平的現實互動,也無法反映隨著知識演變而展開的互動。為了填補這一空白,我們介紹了KNOWSIM,一個圍繞用戶模擬器構建的評估框架,該模擬器維持明確的知識狀態,這些狀態以具有前提關係的信息單元圖表示,並根據學習理論的更新規則進行演變。KNOWSIM直接從知識狀態軌跡計算三個指標(知識增益、交付校準、認知過載),反映信息校準的關鍵機制方面。我們在705個人類-人工智能會話中驗證了KNOWSIM,這些會話分為兩個領域,按知識水平分層:其排名與人類評判顯著一致(73-74%的符號一致性),並超越了三個基線模擬器。應用於9個LLMs,KNOWSIM顯示最佳模型隨用戶知識水平而變化,揭示了標準評估中不可見的能力-處理互動。
A decodability criterion predicts when hidden-state selection beats majority voting in large language models
2608.17124v1 by Zhixiang wang, Ziliang Hong, Ulas Bagci
Combining the answers a large language model (LLM) samples for a question into one decision is a test-time information fusion problem, usually solved by majority voting. Voting is unreliable on difficult questions, where the sampled answers share correlated errors, so the wrong answer can win and drawing more samples makes the decision worse. Selecting a candidate by reading a correctness signal from the model's hidden states is a promising alternative, but its accuracy varies across models and tasks, and no measure indicates when it can be trusted. In this paper, we propose CASE (Correctness-Axis SElection), a dynamic selection combiner that trains a linear gate on the answer-token hidden state and selects the highest-scoring candidate. Its main contribution is decodability, a leakage-free measure of how well the gate ranks a question's correct candidates above its incorrect ones, which predicts whether hidden-state selection will outperform voting. A conventional probe appears accurate only because of question-identity leakage, which vanishes under question-grouped evaluation. On held-out data, decodability predicts the accuracy gain of selection over voting with a Pearson correlation r=0.75 and a decision threshold near AUC=0.60. Across general and medical LLMs, CASE improves over voting by up to 19 points on medium-difficulty questions and 16.8 points on hard questions. Decodability depends on the aligned knowledge a model must recall, not on its scale, and its prediction transfers to an unseen scientific domain within 3.8 points. It thus provides a practical criterion, measurable in advance for a given model and task, for choosing between learned selection and majority voting.
摘要:將大型語言模型(LLM)對一個問題所採樣的答案合併為一個決策是一個測試時的信息融合問題,通常通過多數投票來解決。
在困難問題上,投票不可靠,因為採樣的答案共享相關錯誤,因此錯誤的答案可能會獲勝,而增加更多樣本會使決策變得更糟。
通過從模型的隱藏狀態中讀取正確性信號來選擇候選者是一個有前途的替代方案,但其準確性在不同模型和任務之間有所變化,且沒有任何指標表明何時可以信任它。
在本文中,我們提出了CASE(正確性軸選擇),這是一個動態選擇組合器,對答案標記的隱藏狀態訓練一個線性閘,並選擇得分最高的候選者。
它的主要貢獻是可解碼性,這是一種無洩漏的度量,衡量閘如何將問題的正確候選者排名高於不正確的候選者,並預測隱藏狀態選擇是否會優於投票。
傳統探測器之所以顯得準確,僅僅是因為問題身份的洩漏,而這在問題分組評估中會消失。
在保留數據上,可解碼性預測選擇相對於投票的準確性增益,皮爾森相關係數 r=0.75,決策閾值接近 AUC=0.60。
在一般和醫療 LLM 中,CASE 在中等難度問題上提高了最多 19 分,在困難問題上提高了 16.8 分。
可解碼性取決於模型必須回憶的對齊知識,而不是其規模,且其預測在未見的科學領域內轉移至 3.8 分。
因此,它為在給定模型和任務之間選擇學習的選擇和多數投票提供了一個可實際測量的標準。
From Abductive Explanations to Global Logical Rules for Node Classification in SGCs
2608.17103v1 by Bryan Lima Cavalcante, Thiago Alves Rocha
Graph Neural Networks (GNNs) have achieved remarkable performance in node classification tasks, motivating growing interest in methods capable of explaining their predictions. Recent logic-based approaches, such as LogicXGNN, derive global logical rules for Graph Neural Networks (GNNs) from collections of explanatory subgraphs. While informative, these subgraphs may contain redundant structural information that is specific to individual nodes, potentially limiting the generality of the extracted rules. In this work, we propose a logic-based framework for node classification in Simple Graph Convolution (SGC) networks that uses minimal abductive explanations as an intermediate representation for rule extraction. For each node, we compute a minimal set of node-feature pairs sufficient to preserve the predicted class. These explanations are then used to train decision trees from which global logical rules are extracted. Experiments on benchmark datasets show that the proposed framework produces compact global rules while maintaining high fidelity to the original SGC model.
摘要:圖神經網絡(GNNs)在節點分類任務中取得了顯著的表現,這激發了對能夠解釋其預測的方法的日益關注。最近的基於邏輯的方法,如LogicXGNN,從解釋性子圖的集合中推導出圖神經網絡(GNNs)的全局邏輯規則。雖然這些子圖提供了資訊,但它們可能包含特定於個別節點的冗餘結構資訊,這可能限制了提取規則的普遍性。在本研究中,我們提出了一個基於邏輯的框架,用於簡單圖卷積(SGC)網絡中的節點分類,該框架使用最小的推斷解釋作為規則提取的中介表示。對於每個節點,我們計算一組最小的節點-特徵對,這些對足以保留預測的類別。然後,這些解釋用於訓練決策樹,從中提取全局邏輯規則。在基準數據集上的實驗表明,所提出的框架生成了緊湊的全局規則,同時保持了對原始SGC模型的高保真度。
J-Miner: Recovering Executable Decision Knowledge from Language-Model Classifiers
2608.17063v1 by Yunfan Gao, Xinyi Huang, Tao Sheng, Haorui Song, Yun Xiong, Haofen Wang
Large language models can be fine-tuned into specialized classifiers that perform well across diverse text tasks and make complex judgments, but they typically expose only final labels, leaving the decision knowledge acquired through fine-tuning implicit within the model. We study how to mine this internal decision knowledge from a fine-tuned classifier and encode it in an executable representation that can be inspected, validated, and reused beyond the source classifier. We introduce J-Miner, which mines text-level named concepts by aggregating vocabulary-aligned internal signals across layers and token positions, and uses the classifier's own predictions to learn executable decision rules over them. This process distills local internal readouts into an explicit classifier-level knowledge representation. Across multiple classification tasks, J-Miner rules reproduce up to 98.3\% of source-classifier decisions and achieve 6.0--29.5 percentage points higher behavioral fidelity than equally compact rules learned from input words. Further analysis shows that the named concepts reflect internal semantic evidence associated with task decisions, while the learned rules consolidate these distributed signals into inspectable decision structures. The resulting decision knowledge also transfers to lightweight standalone students: using about 1/24 as many parameters as the source classifiers, they reconstruct and execute the representation from raw text while retaining 99.8\% of the source classifiers' mean task accuracy. These findings show that task-specific decision knowledge can be faithfully represented in an explicit, executable form and reused beyond the classifier in which it was learned.
摘要:大型語言模型可以被微調成為專門的分類器,這些分類器在多樣的文本任務中表現良好並能做出複雜的判斷,但它們通常僅顯示最終標籤,將通過微調獲得的決策知識隱含在模型內部。我們研究如何從微調的分類器中挖掘這種內部決策知識,並將其編碼為可執行的表示,這種表示可以被檢查、驗證並在源分類器之外重用。我們介紹了 J-Miner,它通過在層和標記位置之間聚合與詞彙對齊的內部信號來挖掘文本級命名概念,並利用分類器自身的預測來學習可執行的決策規則。這一過程將局部內部讀出轉化為明確的分類器級知識表示。在多個分類任務中,J-Miner 規則重現了高達 98.3\% 的源分類器決策,並比從輸入詞學習的同樣緊湊規則提高了 6.0--29.5 個百分點的行為忠實度。進一步分析顯示,命名概念反映了與任務決策相關的內部語義證據,而學習到的規則則將這些分散的信號整合為可檢查的決策結構。所產生的決策知識也可以轉移到輕量級的獨立學生模型:使用約 1/24 的參數數量,這些模型能夠從原始文本重建並執行表示,同時保留 99.8\% 源分類器的平均任務準確率。這些發現顯示,特定任務的決策知識可以以明確的、可執行的形式忠實地表示,並在學習該知識的分類器之外重用。
Cross-Model Memory Transfer via Target-Side Reader Adaptation
2608.17050v1 by Mingyuan Li, Guangsheng Yu, Xu Wang, Shaoxiong Ji
Methods for improving knowledge use in large language models typically fall into two regimes. Non-parametric retrieval offers flexible access to external knowledge, but adds retrieval latency, context overhead, and only shallow integration with the backbone. Parametric adaptation is efficient at inference time, but entangles knowledge with model weights and can be hard to update, audit, or transfer. Engram-style hashed memory occupies a middle regime: it stores learned information in an external, addressable table, yet consumes that table through a small learned reader. This raises a basic question: when such a memory is moved across backbones, what matters more, the frozen memory itself or the target-side reader? We study this question through cross-model frozen-memory extraction, in which a memory trained on a source model is frozen and attached to a different target model, with only a lightweight reader trained. Ablations show that learned memory content and correct addressing both matter, but the transferred table becomes useful only through a reader aligned to the target model. In downstream question answering tasks, a dual-layer, four-branch reader nearly closes the gap between same-model and cross-model reuse, achieving an average score of 38.8 under our controlled evaluation protocol. Moreover, when the provider reader is directly compatible with the target interface, the frozen artifact can provide substantial utility without target-side training, while optional reader adaptation yields further improvement. These results suggest that Engram can serve as a reusable external knowledge artifact, provided that the target has access to a compatible reader interface; target-side adaptation can further improve alignment when direct reader reuse is insufficient.
摘要:改善大型語言模型中知識使用的方法通常分為兩種模式。非參數檢索提供靈活的外部知識訪問,但增加了檢索延遲、上下文開銷,並且與主幹的整合僅為淺層。參數適應在推理時效率高,但將知識與模型權重糾纏在一起,並且難以更新、審計或轉移。Engram風格的哈希記憶佔據了中間模式:它將學習到的信息存儲在一個外部的、可尋址的表中,卻通過一個小型的學習讀取器來消耗該表。這引發了一個基本問題:當這樣的記憶在主幹之間移動時,哪一個更重要,凍結的記憶本身還是目標端的讀取器?我們通過跨模型凍結記憶提取來研究這個問題,在這個過程中,訓練於源模型的記憶被凍結並附加到不同的目標模型上,只有一個輕量級的讀取器被訓練。消融實驗顯示,學習的記憶內容和正確的尋址都是重要的,但轉移的表只有通過與目標模型對齊的讀取器才能變得有用。在下游問題回答任務中,一個雙層、四分支的讀取器幾乎縮小了同模型和跨模型重用之間的差距,在我們的控制評估協議下達到了38.8的平均分。此外,當提供者讀取器與目標介面直接兼容時,凍結的工件可以在不進行目標端訓練的情況下提供實質性的效用,而可選的讀取器適應則進一步提高了效果。這些結果表明,Engram可以作為可重用的外部知識工件,前提是目標能夠訪問兼容的讀取器介面;當直接的讀取器重用不足時,目標端的適應可以進一步改善對齊。
AutoSR: Automatic Symbolic Regression by Searching Research States
2608.16876v1 by Kejia Zhang, Youran Sun, Xinyu Ren, Chugang Yi, Haizhao Yang
We introduce Automatic Symbolic Regression (AutoSR), a fully automated system that instantiates Research-Space Symbolic Regression by searching persistent scientific investigations rather than isolated equations. Finite, noisy data often yield numerically competitive expressions that imply very different behavior outside the observed regime, making numerical fit and syntactic complexity insufficient measures of scientific credibility. Existing approaches largely focus on improving expressions, yet the search typically retains little beyond the resulting formula and score, losing the scientific record, such as motivations and probes, that inform what to try next. AutoSR preserves this record in a \textbf{Research State}, coupling each candidate equation with the reasoning, computational evidence, and independent review developed along its branch. Proposer--reviewer agents develop these states under progressive-widening Monte Carlo tree search (PW-MCTS), which allocates computation across competing investigations, while the accumulated research record is ultimately synthesized into a final report that explains the leading relation and the basis for its selection. Across nine selected challenges from two benchmark suites, AutoSR recovers algebraically equivalent relations in every case, including three cp3-bench problems that no published system recovers and six structurally diverse LSR-Transform problems. Overall, AutoSR extends symbolic regression from equation-level search toward automated scientific investigation, allowing scientific knowledge and accumulated evidence to shape both what is explored and how the resulting equation is justified.
摘要:我們介紹自動符號回歸(AutoSR),這是一個完全自動化的系統,它通過搜尋持續的科學研究而不是孤立的方程式來實現研究空間符號回歸。有限的、帶噪聲的數據通常會產生數值上具有競爭力的表達式,這些表達式在觀察範圍之外暗示了非常不同的行為,使得數值擬合和語法複雜性不足以作為科學可信度的衡量標準。現有的方法主要集中在改進表達式上,但搜索通常僅保留結果公式和分數,失去了科學記錄,例如動機和探測,這些記錄告訴我們接下來該嘗試什麼。AutoSR 在一個 \textbf{研究狀態} 中保留這個記錄,將每個候選方程與沿其分支發展的推理、計算證據和獨立審查相結合。提議者-審查者代理在漸進擴展的蒙特卡羅樹搜索(PW-MCTS)下發展這些狀態,該方法在競爭的研究之間分配計算,而累積的研究記錄最終被綜合成一份最終報告,解釋主要關係及其選擇的基礎。在來自兩個基準套件的九個選定挑戰中,AutoSR 在每一個案例中都恢復了代數上等價的關係,包括三個沒有任何已發表系統恢復的 cp3-bench 問題和六個結構多樣的 LSR-Transform 問題。總體而言,AutoSR 將符號回歸從方程層級的搜索擴展到自動化的科學研究,允許科學知識和累積的證據塑造探索的內容以及結果方程的合理性。
Quipu: A Governed Bitemporal Knowledge Graph Store
2608.16813v1 by Steve Brown
Agents now write knowledge graphs, but knowledge-graph stores still carry defaults set when humans curated them: accept writes now and clean later, keep one time axis or none, treat every writer's facts as equally trustworthy, and leave governance to dashboards and middleware. These four defaults are individually convenient and jointly untenable under agent workloads. We present Quipu, an embeddable store that inverts all four: no fact enters except through a gate whose predicates evaluate the pending post-state; data, trust labels, verdicts, and the rules themselves are bitemporal; named graphs are the unit of authority and trust, composed under a lattice whose one invariant is that composition never widens; and the governance specification $Σ$, the trace, and signed verdicts are facts in the store they govern, making the audit $T \models Σ$ a query. We evaluate with Census, a deterministic multi-writer lifecycle whose single seeded run scores every research question against planted ground truth: the gated store ends with 0 of 6 planted defects versus 6 of 6 ungated; all 7 composition probes uphold the lattice contract; 50 of 50 satisfied verdicts re-derive faithfully as of their instant while all 50 would be misreported under a latest-only rule set; and the SARC reference checker agrees with the in-store audit verdict-for-verdict, differing only on coverage semantics. A recorded trace from a governed writer surfaces a live enforcement gap the audit names with its remediation. On DEMM-Bench, an external decision-evidence sufficiency benchmark, a content-only reading of the exported records answers all 512 property-level governance questions correctly with zero overclaim under all eight degradation conditions, while container-presence baselines overclaim on up to 87.5% of them -- and the run surfaced, and led us to close, a gap in what a denial's verdict attests.
摘要:代理人現在撰寫知識圖譜,但知識圖譜存儲仍然保留人類編輯時設置的默認值:現在接受寫入,稍後清理,保持一個時間軸或不保持,將每位作者的事實視為同樣可信,並將治理留給儀表板和中介軟體。這四個默認值在個別上方便,但在代理工作負載下共同無法維持。我們提出了 Quipu,一個可嵌入的存儲,顛覆了這四個默認值:沒有事實進入,除非通過一個門,其謂詞評估待處理的後狀態;數據、信任標籤、裁決和規則本身都是雙時間的;命名圖是權威和信任的單位,根據一個格子組成,其唯一的不變性是組合從不擴大;而治理規範 $Σ$、追蹤和簽名裁決是其治理的存儲中的事實,使得審計 $T \models Σ$ 成為一個查詢。我們使用 Census 進行評估,這是一個確定性的多寫入者生命週期,其單一的種子運行針對植入的真實數據評分每個研究問題:有門的存儲最終以 0 的 6 個植入缺陷結束,而無門的則為 6 的 6 個;所有 7 個組合探針都維護了格子合約;50 的 50 個滿意裁決在其瞬間忠實地重新推導,而所有 50 個在僅最新規則集下會被誤報;而 SARC 參考檢查器在存儲中的審計裁決上逐一一致,僅在覆蓋語義上有所不同。來自受治理作者的記錄追蹤顯示出審計所命名的實時執行差距及其補救措施。在 DEMM-Bench 上,一個外部決策證據充分性基準,對導出的記錄的內容僅閱讀正確回答了所有 512 個屬性級治理問題,並在所有八種降級條件下均無過度聲明,而容器存在基準則在多達 87.5% 的問題上過度聲明——而這次運行顯示出並引導我們關閉了否認裁決所證明的差距。
Bounded Semantic Planning and Deterministic Compilation for Reliable Enterprise Text-to-SQL
2608.16663v1 by Yi Ai
Direct text-to-SQL asks a language model to do two jobs: interpret the business question and construct the complete relational query. In enterprise schemas, SQL can execute successfully while using the wrong relationship role or aggregation grain. We study an alternative placement of the stochastic boundary. A multi-turn planner grounds phrases and selects from question-specific governed options; graph traversal, role predicates, grain lowering, SQL construction, and deterministic checks are implemented in code. We evaluate this semantic path compilation (SPC) system against direct DDL-to-SQL generation on the ACME insurance benchmark. On a 38-question adjudicated comparison set with three runs per question, SPC was adjudicated correct on every run for 37 questions (97.4%), compared with 21 (55.3%) for the baseline. The paired discordance was 16 questions in favor of SPC and none in favor of the baseline (two-sided exact McNemar p=3.05x10^-5). SPC answered all 38 questions correctly at least once and produced one refusal and no adjudicated wrong-but-executed run across 114 run outcomes; the baseline produced 29 adjudicated wrong runs and seven additional judge-flagged data-only coincidences on the same set. A strict-equivalence sensitivity analysis increased the paired difference. Additional SPC runs with GPT-5.4 and Gemini-3.6-Flash showed similar question-level robustness, although their per-run verdict artifacts were not preserved. Six additional benchmark items are retained in an all-item analysis and documented separately by failure class. The study supports an end-to-end systems result, not a causal claim that compilation alone produced the gain, because SPC receives governed semantic artifacts that the DDL baseline does not.
摘要:直接的文本到 SQL 要求語言模型執行兩項任務:解釋商業問題並構建完整的關聯查詢。在企業模式中,SQL 可以在使用錯誤的關係角色或聚合粒度的情況下成功執行。我們研究了隨機邊界的替代放置。一個多輪規劃器將短語進行實體化並從特定問題的受控選項中進行選擇;圖遍歷、角色謂詞、粒度降低、SQL 構建和確定性檢查都在代碼中實現。我們將這個語義路徑編譯(SPC)系統與 ACME 保險基準的直接 DDL 到 SQL 生成進行評估。在一組 38 個問題的裁定比較集中,每個問題進行三次運行,SPC 在 37 個問題的每次運行中都被裁定為正確(97.4%),而基準僅為 21(55.3%)。配對不一致的情況下,SPC 有 16 個問題,而基準則沒有(雙側精確 McNemar p=3.05x10^-5)。SPC 至少正確回答了所有 38 個問題一次,並在 114 次運行結果中產生了一次拒絕,且沒有裁定為錯誤但執行的運行;基準則產生了 29 次裁定為錯誤的運行,並在同一組中有七次額外的法官標記的數據僅巧合。嚴格等價的敏感性分析增加了配對差異。使用 GPT-5.4 和 Gemini-3.6-Flash 的額外 SPC 運行顯示出類似的問題級穩健性,儘管它們的每次運行判決工件未被保留。六個額外的基準項目在全項目分析中保留,並按失敗類別單獨記錄。這項研究支持端到端系統結果,而不是因果聲明,即僅僅編譯產生了增益,因為 SPC 接收了 DDL 基準所沒有的受控語義工件。
The Working Set of a Coding Agent: Coherence Debt in Repository-Scale Tasks
2608.16630v1 by Bardia Mohammadi, Lars Klein, Aman Chadha, Akhil Arora, Laurent Bindschaedler
Repository-scale coding requires an agent to keep tests, imports, configuration, and migration rules consistent within a bounded context window. We model this as reconstructing a coupled-fact graph: at each edit, a required fact comes from recent context or parametric memory, and the facts covered by neither form coherence debt. We supply and withhold each channel and inject faults across seven models and five harnesses. As expected, no model completes a task on an unseen API with both channels empty, and putting the facts in the prompt restores success. When a rename defeats what models memorized about a real library, all seven fail in the same place, passing and missing the same tests. Availability decides the outcome and distance does not: withholding a fact costs exactly the work it supports, and a supplied fact works as well far from the edit as next to it. Harnesses pay unequal prices for it: configurations that all pass every test differ more than tenfold in tokens consumed because they rebuild the same content at different rates, and spending more recovers nothing when facts are withheld. A missing fact produces wrong work rather than absent work: an agent asked to act acts, fabricating the file or guessing the value, so instruments built on reads look for a hole already filled. How often it says it is blocked instead is a property of the model, from every trial to none. Availability does not settle every edit: where standard and code disagree, agents follow the standard even when it prescribes the worse code, so a stale convention file costs more than no file. Because parametric memory substitutes for reading, on SWE-bench, where models likely know the repositories, reads no longer predict success. Harnesses should keep the facts an edit depends on available when the agent writes, and check that availability against what the agent produces rather than what it reads.
摘要:儲存庫規模的編碼需要一個代理在有限的上下文窗口內保持測試、導入、配置和遷移規則的一致性。我們將此建模為重建一個耦合事實圖:在每次編輯時,所需的事實來自最近的上下文或參數記憶,而未被涵蓋的事實則形成一致性債務。我們提供和保留每個通道,並在七個模型和五個工具中注入故障。如預期,當兩個通道都為空時,沒有模型能在未見過的API上完成任務,而將事實放入提示中則恢復成功。當重命名擊敗模型對真實庫的記憶時,所有七個模型在同一位置失敗,通過和未通過相同的測試。可用性決定結果,而距離則不然:保留一個事實的成本正好是它所支持的工作,而提供的事實在遠離編輯的地方也能同樣有效。工具為此支付不平等的價格:所有通過每個測試的配置在消耗的標記上差異超過十倍,因為它們以不同的速度重建相同的內容,而花費更多在事實被保留時則無法恢復任何東西。一個缺失的事實產生錯誤的工作而不是缺失的工作:被要求行動的代理會行動,製造文件或猜測值,因此基於讀取構建的工具會尋找已經填充的空洞。它說自己被阻塞的頻率是一個模型的特性,從每次試驗到無次試驗。可用性並不解決每次編輯:當標準和代碼不一致時,代理遵循標準,即使它規定了更糟的代碼,因此過時的約定文件的成本高於沒有文件。由於參數記憶替代了閱讀,在SWE-bench上,模型可能了解這些儲存庫,讀取不再預測成功。工具應在代理寫入時保持編輯所依賴的事實可用,並檢查該可用性與代理產生的內容,而不是它所讀取的內容。
Hypergraph-based Multimodal Retrieval-Augmented Generation with Incremental Refinement
2608.16628v1 by Shenao Chen, Yidan Xu, Xiangmin Han, Rundong Xue, Duanpo Wu, Yuhan Gao, Chenggang Yan, Yue Gao
Modern Multimodal Retrieval-Augmented Generation (M-RAG) systems are fundamentally limited by the binary connectivity paradigm of traditional simple graphs, which fails to capture the intricate, high-order correlations among heterogeneous entities, such as the N-ary relationships between a visual chart, its scattered textual descriptions, and underlying numerical data. Furthermore, existing refinement strategies often rely on exhaustive, full-page reconstruction to align cross-modal information, leading to prohibitive computational redundancy and the introduction of contextual noise in long-form document processing. In this paper, we propose Hyper-M2RAG, a novel framework that redefines multimodal document retrieval through High-order Hypergraph Representation Learning. We first formalize the document structure as a Multimodal Hypergraph, utilizing hyperedges as unified semantic containers to encapsulate multi-way associations across text, images, and tables, thereby transcending point-to-point modeling. To mitigate semantic fragmentation caused by physical pagination, we introduce an Anchor-driven Incremental Refinement mechanism. Rather than performing a global sweep, our approach identifies boundary-crossing anchor nodes and reconstructs their local hyper-topology using one-hop neighborhood contexts. This targeted refinement effectively bridges cross-page knowledge gaps with minimal computational footprints. Extensive evaluations on multimodal benchmarking datasets demonstrate that Hyper-M2RAG significantly outperforms state-of-the-art methods in both retrieval precision and generation coherence. Our code is available at https://github.com/ShenAoChen2001/MMHRAG.
摘要:現代的多模態檢索增強生成(M-RAG)系統在根本上受到傳統簡單圖的二元連接範式的限制,這無法捕捉異質實體之間複雜的高階相關性,例如視覺圖表、其分散的文本描述和底層數據之間的N元關係。此外,現有的精煉策略通常依賴於全面的全頁重建來對齊跨模態信息,這導致了過度的計算冗餘並在長文檔處理中引入了上下文噪音。在本文中,我們提出了Hyper-M2RAG,一個通過高階超圖表示學習重新定義多模態文檔檢索的新框架。我們首先將文檔結構形式化為多模態超圖,利用超邊作為統一的語義容器,以封裝文本、圖像和表格之間的多向關聯,從而超越點對點建模。為了減輕由物理分頁引起的語義碎片化,我們引入了一種基於錨點的增量精煉機制。我們的方法不是進行全局掃描,而是識別跨邊界的錨點並使用一跳鄰域上下文重建其局部超拓撲。這種有針對性的精煉有效地填補了跨頁知識的空白,並且計算開銷最小。在多模態基準數據集上的廣泛評估表明,Hyper-M2RAG在檢索精度和生成一致性方面顯著超越了最先進的方法。我們的代碼可在 https://github.com/ShenAoChen2001/MMHRAG 獲得。
Physics of Agents: Statistical Mechanics Predicts Collective Behavior of AI Agents
2608.16578v1 by Batu El, Jinhee Paeng, Fatih Dinc, Shiye Su, Mete Erdogan, Aneesh Pappu, Haotian Ye, Wanjia Zhao, Surya Ganguli, James Zou
AI agents increasingly operate as part of interacting systems rather than in isolation. As agents exchange information and jointly make decisions, their interactions can improve collective reasoning but may also produce herding, polarization, or amplify shared biases. Understanding and predicting these collective dynamics is therefore important for designing effective and aligned multi-agent systems. Here, we study over 10,000 communities of language-model agents that repeatedly exchange messages and revise their opinions across objective mathematics questions and subjective political statements. Despite substantial diversity in possible behavior, the individual and group dynamics can be represented by three characteristic regimes: indifference, polarization, and consensus. AI agents start indifferent and build conviction as they interact. On objective questions, communication improves collective accuracy, while on subjective questions it often drifts group opinions toward the right in the political spectrum. We explain these observations with a statistical-mechanics formalism in which agents stochastically favor lower social pressure. Given only initial opinions, our model predicts individual trajectories, outperforms all standard baselines, generalizes to unseen community graphs, and reproduces the observed group archetype distributions. Our fitted model parameters reveal the mechanics underlying our key observations: i) communities operate below the critical social temperature, which explains conviction buildup; ii) attractive ties outweigh repulsive ones, which favors consensus; and iii) agents holding the correct answer exert the strongest pull, which drives truth-seeking. Overall, our results demonstrate that collective behavior of AI agents, like that of other complex systems, follows compact and predictive dynamical laws.
摘要:AI 代理人越來越多地作為互動系統的一部分運作,而不是孤立存在。
隨著代理人之間交換信息並共同做出決策,他們的互動可以改善集體推理,但也可能產生跟風、極化或放大共享偏見。
因此,理解和預測這些集體動態對於設計有效且一致的多代理系統非常重要。
在這裡,我們研究了超過 10,000 個語言模型代理人的社群,它們反覆交換消息並在客觀數學問題和主觀政治陳述上修正自己的意見。
儘管可能的行為存在相當大的多樣性,但個體和群體動態可以用三種特徵性狀態來表示:漠不關心、極化和共識。
AI 代理人最初是漠不關心的,隨著互動的進行建立信念。
在客觀問題上,交流提高了集體準確性,而在主觀問題上,則經常使群體意見向政治光譜的右側漂移。
我們用一種統計力學形式主義來解釋這些觀察,其中代理人隨機地偏好較低的社會壓力。
僅根據初始意見,我們的模型預測個體軌跡,超越所有標準基準,對未見過的社群圖進行泛化,並重現觀察到的群體原型分佈。
我們擬合的模型參數揭示了我們關鍵觀察背後的機制:i) 社群運作在臨界社會溫度以下,這解釋了信念的積累;ii) 吸引性聯繫超過排斥性聯繫,這有利於共識;iii) 持有正確答案的代理人施加最強的影響,這驅動尋求真相。
總體而言,我們的結果表明,AI 代理人的集體行為,如同其他複雜系統,遵循緊湊且可預測的動態法則。
Ask, Condition or Abstain: Reinforcement Learning for Missing-Premise Reasoning
2608.16554v1 by Yongqi Tong, Zhenyu Zhang, Zimi Liu, Kewei Fu, Mingli Song, Haofei Zhang, Junshao Zhang, Hong Zhu, Jiang-Ming Yang, Xin Zhang, Jianshe Li
Answer-only reinforcement learning (RL) trains reasoning models to solve fully specified problems, but many realistic queries omit a premise needed for a unique answer. In this setting, the useful response is not always refusal: the model should ask for the missing premise, condition its answer on the unknown quantity, or abstain when no informative conditional response is available. We present \emph{Ask-Condition-Abstain Reinforcement Learning} (ACA-RL), a data-augmented RL framework for this setting. Its reasoning-graph-guided pipeline converts well-posed problems into missing-premise training instances with localized gap annotations; ACA-RL then trains on these instances with a structured reward over five observable response behaviors. We also introduce the \emph{Missing-Premise Benchmark} (MPB), a 274-instance human-verified benchmark spanning mathematical, logical, and real-world word problems. Across Qwen3 and Llama models, ACA-RL consistently improves on MPB while preserving competitive performance on well-posed reasoning tasks. Together with the released code, MPB, and training data, this work supports a new mission for NLP evaluation: measuring whether models can recognize when a task is underdetermined and handle uncertainty, not only whether they can answer fully specified questions.
摘要:答案導向的強化學習(RL)訓練推理模型以解決完全具體的問題,但許多現實查詢省略了獲得唯一答案所需的前提。在這種情況下,有用的回應不總是拒絕:模型應該要求缺失的前提,將其回答條件化於未知量,或在沒有可提供信息的條件回應時選擇不作答。我們提出了\emph{詢問-條件-不作答強化學習}(ACA-RL),這是一個針對這種情境的數據增強強化學習框架。其推理圖引導的流程將良好表述的問題轉換為缺失前提的訓練實例,並附有局部缺口註釋;然後,ACA-RL在這些實例上進行訓練,並對五種可觀察的回應行為給予結構化的獎勵。我們還介紹了\emph{缺失前提基準}(MPB),這是一個包含274個經過人類驗證的基準,涵蓋數學、邏輯和現實世界的文字問題。在Qwen3和Llama模型中,ACA-RL在MPB上持續改進,同時在良好表述的推理任務中保持競爭性能。連同發布的代碼、MPB和訓練數據,這項工作支持NLP評估的新任務:測量模型是否能夠識別任務是否不確定並處理不確定性,而不僅僅是它們是否能回答完全具體的問題。
VCE-Skill: Enhancing Skill Self-Evolution with Version-Change Experience
2608.16544v1 by Jianming Chen, Xuanbin Ye, Yawen Wang, Junjie Wang, Qing Wang, Fanjiang XU
Agents increasingly rely on reusable skills to encode task knowledge, tool-use procedures, and validation rules. Existing skill self-evolution methods primarily revise skills using execution trajectories collected from current tasks, leaving the evolution knowledge accumulated in public skill version histories largely untapped. Our pilot study reveals a clear complementarity between the two sources: public skill changes provide reusable evolution priors, whereas trajectories provide evidence grounded in the current task. Motivated by this, we propose VCE-Skill, which distills noisy and implementation-specific public skill changes into reusable, structured version-change experience and adaptively fuses it with trajectory-derived proposals from the base evolver, thereby exploiting external experience while retaining task-specific evidence. Extensive experiments demonstrate that VCE-Skill improves skill self-evolution, increasing mean scores by 3.20--4.98 points; transfer experiments further show that the resulting skills achieve stronger cross-model transfer performance. Our work highlights public skill version changes as a previously underexplored yet effective source of prior knowledge and advances trajectory-driven skill self-evolution.
摘要:代理人越來越依賴可重用的技能來編碼任務知識、工具使用程序和驗證規則。現有的技能自我演化方法主要使用從當前任務收集的執行軌跡來修訂技能,導致在公共技能版本歷史中積累的演化知識大多未被利用。我們的初步研究揭示了這兩個來源之間明顯的互補性:公共技能變更提供可重用的演化先驗,而軌跡則提供基於當前任務的證據。基於此,我們提出了 VCE-Skill,該方法將嘈雜且具實施特定性的公共技能變更提煉為可重用的、結構化的版本變更經驗,並自適應地將其與基礎演化器的軌跡導出提案融合,從而在保留任務特定證據的同時利用外部經驗。大量實驗表明,VCE-Skill 改進了技能自我演化,平均分數提高了 3.20--4.98 分;轉移實驗進一步顯示,所產生的技能在跨模型轉移性能上更強。我们的工作突出了公共技能版本變更作為一個先前未被充分探索但有效的先驗知識來源,並推進了基於軌跡的技能自我演化。
Large language models as synthetic clinical experts to inform longitudinal rare-disease modeling
2608.16507v1 by Clemens Schächter, Astrid Pechmann, Janbernd Kirschner, Jan Hasenauer, Harald Binder
Due to the limited amount of information, modeling longitudinal rare-disease data can benefit from integrating clinical knowledge. Yet, elicitation of expert knowledge and formalization for model fitting is challenging, in particular due to limited time of clinical experts. To nevertheless make domain knowledge accessible during model fitting, we use large language models (LLMs) as synthetic clinical experts to supervise a variational-autoencoder-based approach that learns low-dimensional latent summaries of visit-level observations. Specifically, LLMs are queried offline on textual descriptions of patient observations to obtain judgments, e.g., the suspected clinical category. To improve the variational autoencoder fit, we train a differentiable surrogate model on these judgments and augment the loss function to encourage reconstructions that preserve the clinical-label distribution of their corresponding input profile. In an application to longitudinal motor-function assessments from children with spinal muscular atrophy, we map visit-level clinical profiles to low-dimensional representations that are linked by a multivariate mixed-effects model. The synthetic expert loss discourages reconstructions that remain numerically close in data space but alter the clinical interpretation of the reconstructed motor function profile, such as by crossing a disease-type boundary. We thus reduced disagreement between original and reconstructed SMA type labels from about 11 to 7 percent. Furthermore, informing the latent representation by the synthetic expert improved prediction of motor function milestones compared with unsupervised latent representations and a data-level baseline. These results suggest that incorporating LLMs into model fitting can make clinical knowledge available to representation learning and improve clinical faithfulness for longitudinal rare-disease data.
摘要:由於資訊量有限,建模縱向罕見疾病數據可以從整合臨床知識中受益。然而,專家知識的引出和模型擬合的形式化是具有挑戰性的,特別是由於臨床專家的時間有限。儘管如此,為了在模型擬合過程中使領域知識可用,我們使用大型語言模型(LLMs)作為合成臨床專家,來監督基於變分自編碼器的方法,該方法學習訪問級觀察的低維潛在摘要。具體來說,我們在患者觀察的文本描述上離線查詢LLMs以獲得判斷,例如,懷疑的臨床類別。為了改善變分自編碼器的擬合,我們在這些判斷上訓練了一個可微分的替代模型,並增強損失函數以鼓勵重建保持其對應輸入特徵的臨床標籤分佈。在對脊髓性肌萎縮症兒童的縱向運動功能評估的應用中,我們將訪問級臨床特徵映射到由多變量混合效應模型鏈接的低維表示。合成專家損失會抑制在數據空間中數值上接近但改變重建運動功能特徵的臨床解釋的重建,例如通過跨越疾病類型邊界。因此,我們將原始和重建的SMA類型標籤之間的分歧從約11%減少到7%。此外,通過合成專家告知潛在表示,與無監督潛在表示和數據級基準相比,運動功能里程碑的預測得到了改善。這些結果表明,將LLMs納入模型擬合可以使臨床知識可用於表示學習,並改善縱向罕見疾病數據的臨床真實性。
Graph Machine Learning: An Opportunity for Power Systems
2608.16494v1 by Martin Sadric, Sebastian Pütz, Christian Nauck, Veit Hagenmeyer, Frank Hellmann, Dirk Witthaut, Benjamin Schäfer
Modern power systems face growing operational complexity driven by the integration of renewable energy sources, decentralization, and the need for real-time decision-making across a wide range of timescales. Addressing these challenges traditionally relies on model-based methods that, while accurate, can be too slow for operational demands. Machine learning (ML) has therefore emerged as a faster, data-driven alternative. As grid topology plays a central role in power system operation, graph machine learning (GML) methods offer a natural framework for incorporating topological dependencies as an inductive bias. We survey nearly 800 papers at the intersection of GML and power systems, covering forecasting, state estimation, optimization, control, fault diagnosis, and cybersecurity. Power systems constitute an unusually rich benchmark setting for GML, as they combine hard physical constraints, multi-scale dynamics, safety-critical requirements, and scarce labeled data within a single, well-defined domain. Conversely, power systems can benefit from utilizing GML to complement classical solvers, as GML provide scalable, topology-aware approximations with promising generalization and computational efficiency. We identify open challenges, including limited real-world deployment and the need for interpretable models in safety-critical settings. Despite the rapidly growing number of publications, standardized benchmarks and open datasets remain scarce, leaving many results difficult to reproduce and undermining the long-term scientific credibility of the field. We further derive a structured requirements catalog for ML-ready power grid benchmarks, intended to guide future dataset development and improve reproducibility across studies. We call on the community to prioritize dedicated benchmark studies and the release of open datasets and models.
摘要:現代電力系統面臨著由可再生能源整合、去中心化以及在廣泛時間尺度上進行實時決策所驅動的日益增長的運營複雜性。傳統上,解決這些挑戰依賴於基於模型的方法,儘管這些方法準確,但對於運營需求來說可能過於緩慢。因此,機器學習(ML)作為一種更快的數據驅動替代方案應運而生。由於電網拓撲在電力系統運作中扮演著核心角色,圖形機器學習(GML)方法提供了一個自然的框架,以將拓撲依賴性作為歸納偏差納入考量。我們調查了近800篇GML與電力系統交叉的論文,涵蓋預測、狀態估計、優化、控制、故障診斷和網絡安全。電力系統為GML提供了一個異常豐富的基準設置,因為它們在一個明確定義的領域內結合了嚴格的物理約束、多尺度動力學、安全關鍵要求以及稀缺的標記數據。相反,電力系統可以利用GML來補充傳統求解器,因為GML提供了可擴展的、考慮拓撲的近似,並且具有良好的泛化能力和計算效率。我們確定了開放挑戰,包括有限的實際部署和在安全關鍵環境中對可解釋模型的需求。儘管出版物數量迅速增長,標準化基準和開放數據集仍然稀缺,這使得許多結果難以重現,並削弱了該領域的長期科學可信度。我們進一步推導了一個結構化的需求目錄,用於ML準備好的電網基準,旨在指導未來數據集的開發並提高研究的可重複性。我們呼籲社區優先考慮專門的基準研究以及開放數據集和模型的發布。
Time to Reason: Scalable Neurosymbolic Learning for LTLf via Fuzzy Semantics
2608.16443v1 by Riccardo Andreoni, Andrei Buliga, Alessandro Daniele, Paolo Felli, Chiara Ghidini, Marco Montali, Massimiliano Ronzani
Neurosymbolic (NeSy) Artificial Intelligence aims to integrate Deep Learning (DL) architectures with symbolic reasoning. While initial NeSy approaches have targeted mainly symbolic reasoning in propositional and first-order logics, recent works have started to address the construction of neurosymbolic frameworks for Temporal Logics, and in particular for LTLf. These approaches have established temporal NeSy as a promising research direction, laying the foundations for learning under temporal constraints. Nonetheless, they leave many questions unanswered. From a theoretical perspective, several differentiable semantics for interpreting LTLf have been proposed but have not yet been formally and systematically defined within a unified framework. Moreover, existing approaches commonly rely on automata to represent temporal knowledge, resulting in limited scalability. Motivated by this research gap, this paper provides the following contributions: (i) formally defining different fuzzy semantics for LTLf, and systematically analysing theoretical properties regarding equivalences and dualities of temporal operators; (ii) showing how these semantics can be directly integrated within a novel NeSy framework, called DiffLTLf, enabling flexible and scalable learning without relying on the usage of automata; and (iii) introducing a novel evaluation protocol of increased complexity of learning tasks w.r.t. existing benchmarks. Our results show that the choice of fuzzy semantics has a significant impact on predictive performance. Moreover, DiffLTLf achieves performance on par with, and sometimes superior to, state-of-the-art probabilistic approaches while substantially improving scalability. Taken together, these results establish direct fuzzy interpretations as a competitive and scalable alternative to existing temporal NeSy frameworks.
摘要:神經符號(NeSy)人工智慧旨在將深度學習(DL)架構與符號推理整合。雖然最初的NeSy方法主要針對命題邏輯和一階邏輯中的符號推理,但最近的研究已開始著手於構建神經符號框架以處理時間邏輯,特別是針對LTLf。這些方法已將時間NeSy確立為一個有前景的研究方向,為在時間約束下的學習奠定了基礎。儘管如此,它們仍然留下許多未解答的問題。從理論的角度來看,已提出幾種可微分的語義來解釋LTLf,但尚未在統一框架內正式和系統地定義。此外,現有的方法通常依賴自動機來表示時間知識,導致可擴展性有限。受此研究空白的啟發,本文提供了以下貢獻:(i)正式定義不同的LTLf模糊語義,並系統地分析有關時間運算符的等價性和對偶性的理論性質;(ii)展示這些語義如何能夠直接整合進一個新穎的NeSy框架,稱為DiffLTLf,實現靈活且可擴展的學習,而無需依賴自動機的使用;以及(iii)引入一種新的評估協議,增加學習任務的複雜性,相較於現有基準。我們的結果顯示,模糊語義的選擇對預測性能有顯著影響。此外,DiffLTLf在性能上與最先進的概率方法相當,有時甚至優於它們,同時顯著提高了可擴展性。綜合這些結果,直接的模糊解釋被確立為現有時間NeSy框架的競爭性和可擴展替代方案。
Reasoning-supported Robustness Validation of Automotive E/E Components
2608.16421v1 by Jan Novacek, Alexander Viehl, Oliver Bringmann, Wolfgang Rosenstiel
This paper presents an ontology-supported approach to tackle the complexity of the Robustness Validation (RV) process of automotive electrical/electronic (E/E) components. The approach uses formalized knowledge from the RV process and stress, operating, and load profiles, so-called Mission Profiles (MPs). In contrast to the error-prone industrially established manual procedure, we show how component characteristics are formalized in OWL in order to form the foundation of an efficient automated analysis selection and decision support during the RV process. The proposed approach is based on the idea of mapping MPs to an OWL representation so to allow to perform semantic queries against MP data to improve their integration into the RV process. The resulting ontology-supported application framework has been applied to an industrial use-case from automotive power electronics. We present experimental results showing that the RV process can be significantly improved in terms of reduced design time and increased exhaustiveness by automating the analyses selection step and the provisioning of all the relevant data to be used.
摘要:這篇論文提出了一種基於本體的方式來應對汽車電氣/電子(E/E)元件的穩健性驗證(RV)過程的複雜性。該方法利用了來自RV過程的形式化知識以及壓力、操作和負載特徵,這些被稱為任務特徵(MPs)。與錯誤易發的工業手動程序相比,我們展示了如何在OWL中形式化元件特徵,以便為RV過程中的高效自動分析選擇和決策支持奠定基礎。所提出的方法基於將MP映射到OWL表示的想法,以便能夠對MP數據執行語義查詢,從而改善其在RV過程中的整合。最終得到的基於本體的應用框架已應用於汽車功率電子的工業案例。我們展示了實驗結果,顯示通過自動化分析選擇步驟和提供所有相關數據,RV過程在設計時間減少和全面性增加方面可以顯著改善。
Think Inside the Chunk: RegulaRAG for Regulation-Compliant Scenario Generation using LLMs: A Case Study of UN Regulation No. 152
2608.16394v1 by Vahid Zolfaghari, Nenad Petrovic, AndrÉ Schamschurko, Alois Knoll
Generating regulation-compliant test scenarios is essential for validating safety-critical automotive systems, yet Large Language Models (LLMs) struggle to ground outputs in long, hierarchical standards. We present RegulaRAG, a Retrieval-Augmented Generation (RAG) pipeline that couples SmartChunking, reference-aware enrichment of paragraphs and tables via graph traversal, with Smart Retrieve & Rerank over these enriched units. To test our system, we evaluate on a manually curated dataset covering all scenarios in UN Regulation No. 152 (AEBS). Our study comprises: (i) a three-step progressive search that identifies near-optimal retrieval parameters without exhaustive grid search; (ii) head-to-head comparisons against five baseline RAG systems; and (iii) a robustness stress test that scales the source corpus with distractor content. Outputs are evaluated using a customized penalized scoring metric. Across all experiments, RegulaRAG achieves the highest average Meta-Score (82.99), outperforming the next-best system by 43% (NoRAG: 57.94), while operating at 14k-25k tokens per query versus up to 500k for graphcentric baselines. It maintains strong performance, remaining stable even as the number of regulatory sources grows, whereas competing RAG systems degrade sharply in both quality and robustness.
摘要:生成符合規範的測試場景對於驗證安全關鍵的汽車系統至關重要,但大型語言模型(LLMs)在將輸出與長期的層次標準相結合方面存在困難。
我們提出了RegulaRAG,一個檢索增強生成(RAG)管道,結合了SmartChunking、通過圖遍歷對段落和表格進行參考感知的豐富化,以及對這些豐富單元的智能檢索與重新排序。
為了測試我們的系統,我們在一個手動策劃的數據集上進行評估,該數據集涵蓋了聯合國第152號規範(AEBS)中的所有場景。
我們的研究包括:(i)一個三步驟的漸進搜索,識別近乎最優的檢索參數,而無需進行耗時的網格搜索;(ii)與五個基準RAG系統的正面比較;以及(iii)一個強度壓力測試,通過干擾內容擴展源語料庫。
輸出使用自定義的懲罰評分指標進行評估。
在所有實驗中,RegulaRAG達到了最高的平均Meta-Score(82.99),比第二好的系統高出43%(NoRAG: 57.94),同時每個查詢的操作在14k-25k個標記之間,而圖中心基準則高達500k。
它保持了強勁的性能,即使在監管來源數量增加的情況下也保持穩定,而競爭的RAG系統在質量和穩定性方面急劇下降。
Mint-Agent: Introducing Finance-Native Agentic Foundation Models
2608.16386v1 by Mint-Agent Team, B. Zhang, Yaze Geng, Lei Tang, Yaoyang Yi, Zonghan Wu, Yifan Hu, Kun Wang, Qingsong Wen, Yilei Shao
Financial agents must do more than recall domain knowledge: they must be both reliable, executing precise operations over grounded evidence, and executive, sustaining long-horizon research whose conclusions remain auditable. We present Mint-Agent, a family of finance-native agentic models designed around these two scales of financial intelligence. Mint-Agent is built upon three pillars: data, harness, and algorithm. Our data engine constructs clean, specialized tasks for atomic financial capabilities and long-horizon agentic execution from real-world financial sources. MintHarness enables stable interaction with open-ended environments and maintains auditable evidence trails across extended research trajectories. Our training recipe combines SFT, critical-step OPD, and RLVR to develop separate financial reasoning and agentic execution experts, which are then unified through model merging and multi-teacher on-policy distillation into compact, general-purpose financial agents. This pipeline yields two flagship models, Mint-Cu (9B) and Mint-Ag (27B). Across professional financial benchmarks, our models demonstrate two defining strengths: (1) Reliability: Mint-Ag achieves 98.33% on RFC-Bench, surpassing GPT-5.6-Sol and Claude-Opus-4.8 by 3.66 and 3.00 points; and (2) Executability: Mint-Cu reaches 69.86% on FinSearchComp T2, outperforming Agents-A1-35B and Nex-N2-mini by 22.83 and 12.78 points, while Mint-Ag achieves 76.00% and 60.49% on FinanceAgentBench v1.1 and v2, respectively. These results establish a path toward trustworthy financial intelligence in which domain expertise, long-horizon execution, and auditable evidence are jointly engineered as a unified foundation for frontier agentic models.
摘要:金融代理人必須做的不僅僅是回憶領域知識:他們必須既可靠,能夠在有根據的證據上執行精確的操作,又必須具備執行力,能夠支持長期研究,其結論保持可審核性。我們提出了Mint-Agent,這是一系列圍繞這兩個金融智能尺度設計的金融原生代理模型。Mint-Agent建立在三個支柱之上:數據、利用和算法。我們的數據引擎從現實世界的金融來源構建乾淨的、專門的任務,以實現原子金融能力和長期代理執行。MintHarness使得與開放式環境的穩定互動成為可能,並在擴展的研究軌跡中維持可審核的證據鏈。我們的訓練配方結合了SFT、關鍵步驟OPD和RLVR,以培養獨立的金融推理和代理執行專家,然後通過模型合併和多教師在政策蒸餾將它們統一為緊湊的通用金融代理。這個流程產生了兩個旗艦模型,Mint-Cu (9B)和Mint-Ag (27B)。在專業金融基準測試中,我們的模型展示了兩個明確的優勢:(1)可靠性:Mint-Ag在RFC-Bench上達到98.33%,超越了GPT-5.6-Sol和Claude-Opus-4.8,分別提高了3.66和3.00分;以及(2)可執行性:Mint-Cu在FinSearchComp T2上達到69.86%,超越了Agents-A1-35B和Nex-N2-mini,分別提高了22.83和12.78分,而Mint-Ag在FinanceAgentBench v1.1和v2上分別達到76.00%和60.49%。這些結果為值得信賴的金融智能鋪平了道路,其中領域專業知識、長期執行和可審核的證據共同被設計為前沿代理模型的統一基礎。
MELD: A Protocol for Merging Knowledge Across Distributed Agentic Memories
2608.16357v1 by Lauri Lovén, Jaakko Sauvola, Jukka Riekki, Sasu Tarkoma
Autonomous agents share a transport and can call each other's tools, but they cannot share what they know: no protocol lets two agents' memories reconcile a fact phrased two ways, link related facts held apart, or reconcile contradictory knowledge without silently discarding either claim. We present MELD, a self-managing coherence mechanism for a federation of agent memories whose run-time model is the knowledge graph itself. Each brain admits every incoming claim through a five-outcome procedure (insert, merge, relate, conflict, or reject), decided from three signals (scoped claim-key identity, embedding similarity, and a natural-language-inference verdict) under context and freshness gates, and acting through exactly one auditable, authenticated Patch, the only object that mutates state. A binding onto standard publish/subscribe transport with a per-claim status CRDT keeps sovereign brains coherent in claim status without a coordinator: self-healing after partitions and under lossy routing, and self-protecting against silent rewrite by a peer, under a benign-fault model. MELD does not adjudicate truth; a detected contradiction is preserved for later adjudication, never silently resolved. On HotpotQA distractor, distributed merge is recall-non-inferior to a centralized store under a pre-specified equivalence test and recall-superior to naive union at about 11% less live storage; the merge classifier separates at AUC 0.968 with a 0.013 false-merge rate on adjudicated candidate pairs; the status CRDT reconverges in 30/30 real partition-heal trials where last-writer-wins manages 11/30; and semantic routing delivers about 3x fewer messages at matched recall. We evaluate on a real computing continuum spanning an operator-grade 5G edge, national HPC, and a local tier, with empirically calibrated thresholds.
摘要:自主代理共享一個傳輸並可以呼叫彼此的工具,但他們無法共享所知:沒有任何協議能讓兩個代理的記憶調和以兩種方式表述的事實,連結分開持有的相關事實,或在不靜默丟棄任何主張的情況下調和矛盾的知識。我們提出了MELD,一種自我管理的連貫性機制,用於一個代理記憶的聯邦,其運行時模型即為知識圖譜。每個大腦通過一個五種結果的程序(插入、合併、關聯、衝突或拒絕)接受每個進來的主張,這一決定基於三個信號(範疇主張鍵身份、嵌入相似性以及自然語言推理的裁決),在上下文和新鮮度閘門下運作,並通過恰好一個可審計的、經過身份驗證的Patch進行操作,這是唯一能改變狀態的對象。與標準的發布/訂閱傳輸綁定的每個主張狀態CRDT使得主權大腦在主張狀態上保持一致,無需協調者:在分區後自我修復,並在有損路由下自我保護,防止被同伴靜默重寫,遵循良性故障模型。MELD不裁決真相;檢測到的矛盾將被保留以便後續裁決,絕不靜默解決。在HotpotQA的干擾者上,分散合併的回憶在預先指定的等價測試下不劣於集中存儲,並且在約11%更少的實時存儲下回憶優於天真的聯合;合併分類器在裁決候選對上以AUC 0.968分開,假合併率為0.013;狀態CRDT在30/30的真實分區修復試驗中重新收斂,而最後寫者獲勝的情況下僅管理11/30;語義路由在匹配回憶時傳遞的消息數量約少了3倍。我們在一個涵蓋操作級5G邊緣、國家HPC和本地層的真實計算連續體上進行評估,並經過實證校準的閾值。
AeroCopilotBench: A Two-Tier Benchmark for Evaluating LLM Agents as Aviation Copilots in an Interactive Virtual Cockpit Environment
2608.16349v1 by Yuchen Yuan, Zhenghuang Wu, Yuangan Li, Liang Ma, Ke Li
Large language model (LLM) agents may assist flight crews with complex decisions and task execution, but existing aviation evaluations centered on static knowledge do not support systematic testing of procedural execution and safety compliance in interactive environments. This paper presents the AeroCopilot Operational Environment (ACOE), a reproducible interactive virtual-cockpit test environment, and AeroCopilotBench, a two-tier aviation agent evaluation benchmark. Tier-1 evaluates aviation knowledge using 1,200 multiple-choice questions, while Tier-2 comprises 73 emergency and abnormal tasks derived from the manufacturers' Pilot's Operating Handbooks (POHs) and instantiated in ACOE. ACOE converts natural-language procedures into executable state transitions, final-state goal conditions, and hard safety constraints, enabling models to interpret cockpit state, diagnose faults, and operate aircraft systems through standardized tool interfaces. We establish a safety-gated evaluation framework in which a trajectory succeeds only when all task goals are achieved without violating any hard safety constraint, while safe goal progress and trajectory safety are measured separately. Across 12 models, the highest Tier-2 success rate is 72.6%, while static knowledge performance does not consistently translate into procedural execution. Analysis of 451 failed episodes from 3 representative models identifies recurring failures in procedural completeness, use of state feedback, and long-horizon execution management. These findings motivate state-aware agent orchestration, joint assessment of task completion and trajectory safety, and repeated regression testing. ACOE and AeroCopilotBench provide a reproducible foundation for testing knowledge application, interactive execution, and operational safety in aviation agents.
摘要:大型語言模型(LLM)代理可能協助飛行組員進行複雜的決策和任務執行,但現有的航空評估集中於靜態知識,無法支持在互動環境中系統性測試程序執行和安全合規性。本文提出了航空副駕駛操作環境(ACOE),這是一個可重複的互動虛擬駕駛艙測試環境,以及航空副駕駛基準(AeroCopilotBench),這是一個兩級航空代理評估基準。第1級使用1,200道多選題評估航空知識,而第2級則包含73個來自製造商飛行操作手冊(POHs)的緊急和異常任務,並在ACOE中實現。ACOE將自然語言程序轉換為可執行的狀態轉換、最終狀態目標條件和硬性安全約束,使模型能夠解釋駕駛艙狀態、診斷故障並通過標準化工具接口操作飛機系統。我們建立了一個安全門控評估框架,其中只有在所有任務目標達成且不違反任何硬性安全約束的情況下,軌跡才算成功,而安全目標進展和軌跡安全則分別測量。在12個模型中,第2級的最高成功率為72.6%,而靜態知識表現並不總是一致地轉化為程序執行。對3個代表性模型中451個失敗案例的分析發現,程序完整性、狀態反饋的使用和長期執行管理存在重複性失敗。這些發現促使了狀態感知代理的協同、任務完成和軌跡安全的聯合評估,以及重複回歸測試。ACOE和航空副駕駛基準為測試知識應用、互動執行和航空代理的操作安全提供了可重複的基礎。
Executable Code Knowledge: Code as a Native, Validation-Carrying Knowledge Representation for AI Coding Agents
2608.16295v1 by Xueping Gao
AI coding agents need more than relevant snippets: they need business semantics, validation evidence, relations, and assurance that their context is current. Existing systems usually infer or externalize this knowledge through retrieval, summaries, graphs, rules, or reverse specifications. We investigate a complementary representation in which selected code units directly carry agent-usable knowledge. We introduce Executable Code Knowledge (ECK) and define an Executable Code Knowledge Unit (ECKU) as a source-bound object combining stable identity, semantics, executable behavior, contracts, evidence, relations, provenance, validation state, and a query interface. Our Python prototype supports code-local authoring, manifest export, evidence execution, exact changed-line impact, freshness checking, and agent-facing projections. Across three real Python repositories and 26 controlled patch tasks, direct ECK provides executable test coverage for 11/11 evidence-bearing tasks and exact selectors for 9/11; hiding declared evidence reduces exact recovery to 1/11 (paired exact McNemar p=0.0078). ECK-derived rules recover 11/11 exact selectors, showing that rules are effective delivery artifacts while ECK supplies source binding, validation state, impact, and freshness. Exact changed-line impact matches independently authored labels on all 26 patches (12 unit links; precision, recall, and F1 all 1.000). AST-bounded fingerprints classify 50 positive changes and 17 unrelated same-file controls correctly, whereas static rules snapshots detect none of the 50 stale cases. Model-backed patch-review and cross-layer studies measure projection fidelity rather than independent impact discovery. These results support a hybrid architecture: retrieval for coverage, ECK for source and evidence governance, and projections for delivery.
摘要:AI 編碼代理需要的不僅僅是相關的片段:他們需要商業語義、驗證證據、關係,以及確保其上下文是最新的。現有系統通常通過檢索、摘要、圖表、規則或反向規範來推斷或外部化這些知識。我們研究了一種互補的表示方式,其中選定的代碼單元直接攜帶代理可用的知識。我們引入了可執行代碼知識(Executable Code Knowledge, ECK),並將可執行代碼知識單元(Executable Code Knowledge Unit, ECKU)定義為一個源綁定對象,結合穩定的身份、語義、可執行行為、合約、證據、關係、來源、驗證狀態和查詢介面。我們的 Python 原型支持代碼本地創作、清單導出、證據執行、精確變更行影響、更新性檢查和面向代理的投影。在三個真實的 Python 存儲庫和 26 個受控補丁任務中,直接 ECK 為 11/11 個帶證據的任務提供了可執行的測試覆蓋,並為 9/11 提供了精確選擇器;隱藏已聲明的證據將精確恢復降低到 1/11(配對精確 McNemar p=0.0078)。ECK 衍生的規則恢復了 11/11 的精確選擇器,顯示規則是有效的交付工件,而 ECK 提供了源綁定、驗證狀態、影響和更新性。精確變更行影響與所有 26 個補丁上的獨立創作標籤相匹配(12 個單元鏈接;精確度、召回率和 F1 均為 1.000)。AST 限定的指紋正確分類了 50 個正變更和 17 個無關的同檔控制,而靜態規則快照未檢測到 50 個過時案例。模型支持的補丁審查和跨層研究測量投影忠實度,而不是獨立影響發現。這些結果支持一種混合架構:檢索用於覆蓋,ECK 用於源和證據治理,投影用於交付。
Clause Encounters of the Third Kind: Can LLMs Replace Language Teachers?
2608.16286v1 by Kristina Šekrst, Ana Kovačić
While various organizations now actively encourage LLM use in classrooms, we still lack rigorous, systematic evaluations of how well these models actually perform the fundamental tasks of language pedagogy. This paper examines whether state-of-the-art LLMs can deliver the kind of corrective feedback and methodological explanations that language learners need. The study tests multiple large language models on their ability to identify, correct, and explain common learner mistakes in English, by systematically varying model parameters to investigate how these technical adjustments affect output quality, pedagogical clarity, and consistency, along with using retrieval-augmented generation to query methodological data. The evaluation employs automated metrics (GLEU, BERTScore) but also human expert judgments to capture dimensions that purely computational measures miss: linguistic nuance, cultural sensitivity, and instructional appropriateness. While models demonstrate impressive surface-level correction abilities, their explanations often lack the terminological and domain knowledge that effective language teaching requires, suggesting that current enthusiasm for AI-assisted language learning may be outpacing our understanding of these systems' actual pedagogical competence.
摘要:雖然各種組織現在積極鼓勵在課堂上使用 LLM,但我們仍然缺乏對這些模型在語言教學基本任務上實際表現的嚴謹、系統性評估。本文檢視最先進的 LLM 是否能提供語言學習者所需的糾正反饋和方法論解釋。該研究測試多個大型語言模型在識別、糾正和解釋英語學習者常見錯誤的能力,通過系統性地變化模型參數來調查這些技術調整如何影響輸出質量、教學清晰度和一致性,同時使用增強檢索生成來查詢方法論數據。評估使用自動化指標(GLEU、BERTScore),但也包括人類專家的判斷,以捕捉純計算度量所忽略的維度:語言細微差別、文化敏感性和教學適當性。雖然模型展示了令人印象深刻的表面糾正能力,但它們的解釋往往缺乏有效語言教學所需的術語和領域知識,這表明目前對 AI 輔助語言學習的熱情可能超過了我們對這些系統實際教學能力的理解。
Domain-Agnostic Neural Topic Modeling with Contextual Token-Level Semantic Graph Representation
2608.16269v1 by Seung-Won Seo, Won Ik Cho, Yongmin Yoo
Recent advances in neural topic models with pre-trained language models (PLMs) have achieved strong performance by leveraging general-domain pre-training, yet their topic interpretability often degrades on specialized corpora. This limitation primarily stems from the geometry of the embedding space, where domain-specific terms unseen during pre-training collapse into an indistinguishable region, and neither domain-specific re-training, word-level graph enrichment, nor parameter-efficient fine-tuning can restructure this space without inheriting the capacity ceiling of the underlying encoder. Our key insight is that a learnable graph layer operating on token-level PLM embeddings can acquire corpus-specific semantic structure that the frozen encoder lacks, because token-level graphs preserve document-local context that word-level representations discard and joint optimization with the topic objective reshapes embedding geometry directly from target-domain evidence. We instantiate this insight as DARTopic, a domain-agnostic framework that constructs token-level semantic graphs from frozen PLM embeddings and jointly trains a GNN encoder with topic inference. Across three benchmarks spanning general, biomedical, and legal domains, DARTopic consistently outperforms strong baselines in topic coherence and document clus- tering without any encoder fine-tuning, while demonstrating robustness to PLM choice and favorable runtime efficiency over fine-tuning based alternatives.
摘要:最近在使用預訓練語言模型(PLMs)的神經主題模型方面取得了顯著進展,通過利用通用領域的預訓練達到了強大的性能,然而它們在專門語料上的主題可解釋性往往會下降。這一限制主要源於嵌入空間的幾何特性,其中在預訓練期間未見過的領域特定術語會塌縮到一個無法區分的區域,而領域特定的再訓練、詞級圖增強或參數高效的微調都無法在不繼承基礎編碼器的容量上限的情況下重構這一空間。我們的關鍵見解是,運行在標記級PLM嵌入上的可學習圖層可以獲得凍結編碼器所缺乏的語料特定語義結構,因為標記級圖保留了文檔局部上下文,而詞級表示則被丟棄,並且與主題目標的聯合優化直接從目標領域的證據重塑嵌入幾何。我們將這一見解具體化為DARTopic,一個與領域無關的框架,從凍結的PLM嵌入構建標記級語義圖,並與主題推斷共同訓練GNN編碼器。在涵蓋通用、生物醫學和法律領域的三個基準測試中,DARTopic在主題一致性和文檔聚類方面始終超越強基準,且無需任何編碼器微調,同時對PLM選擇展現出穩健性,並在運行效率上優於基於微調的替代方案。
Picking the Right Image to Classify: Reliable-Input Selection in Teledermatology
2608.16198v1 by Fabian Gröger, Marco Weishaupt, Philippe Gottfrois, Simone Lionetti, Linda Wermelinger, Nipun Ranasekara, Ludovic Amruthalingam, Alexander A. Navarini, Marc Pouly
Dermatology models face distribution shifts in teledermatology settings, where submitted images differ from the training data in lighting, angle, distance, focus, and framing. These test-time images are ordinary clinical photographs, but some fall outside the model's training conditions, leading the model to often misclassify them due to shifts in acquisition between training and deployment. When multiple images of the same case exist (several photos of one patient or lesion), a natural way to improve accuracy is therefore to select the image the model is most likely to classify correctly. We call this task reliable-input selection. An oracle that, for each case, selects a correctly classified image when one exists raises weighted F1 by about 20 percentage points on average across six dermatology datasets and nine frozen backbones. This oracle is an upper bound that sees the labels, whereas a selector must choose blindly. Capturing this gain in practice is hard. A selector that needs no pretraining data applies to any frozen model, including those whose data is not public. It must judge reliability from quantities the model exposes at inference: its embeddings, their norms, and its confidence. We benchmark four such training-data-free selectors: the embedding norm, the neighborhood consensus among a case's images, the stability of the prediction under small perturbations, and the model's own confidence. No training-data-free selector substantially narrows this oracle gap. The best of them is the model's own confidence, but it recovers only a small part of the gap on the clinical datasets. A small labeled reference set does not help either: the best selector overall, a fusion of confidence and Mahalanobis distance, still leaves most of the gap. To our knowledge, this is the first study to introduce and benchmark reliable input selection, a clinically important, unsolved task.
摘要:皮膚科模型面臨在遠程皮膚科環境中分佈轉移的挑戰,提交的圖像在照明、角度、距離、焦點和構圖上與訓練數據有所不同。這些測試時的圖像是普通的臨床照片,但有些超出了模型的訓練條件,導致模型經常因訓練與部署之間的獲取差異而錯誤分類。當同一病例存在多張圖像(多張同一患者或病變的照片)時,改善準確性的自然方法是選擇模型最有可能正確分類的圖像。我們稱這個任務為可靠輸入選擇。對於每個案例,當存在正確分類的圖像時,選擇這樣的圖像的神諭平均提高六個皮膚科數據集和九個凍結骨幹的加權F1約20個百分點。這個神諭是一個上限,能看到標籤,而選擇器必須盲目選擇。在實踐中捕捉這一增益是困難的。需要無預訓練數據的選擇器適用於任何凍結模型,包括那些數據不公開的模型。它必須根據模型在推斷時暴露的數量來判斷可靠性:其嵌入、它們的範數和模型的信心。我們基準測試了四種無訓練數據的選擇器:嵌入範數、案例圖像之間的鄰域共識、在小擾動下預測的穩定性,以及模型自身的信心。沒有一種無訓練數據的選擇器能顯著縮小這一神諭差距。其中最好的選擇器是模型自身的信心,但它在臨床數據集上僅恢復了差距的一小部分。小型標記參考集也沒有幫助:整體最佳的選擇器,即信心和馬哈拉諾比斯距離的融合,仍然留下了大部分差距。據我們所知,這是第一項引入和基準測試可靠輸入選擇的研究,這是一個臨床重要的未解決任務。
LENS: In-Context Search via Latent Evidence Exploration over Dynamic Raw Documents
2608.16185v2 by Xingjun Wang, Gongsheng Li, Qi Fan, Yunlin Mao, Luyan Su, Yingda Chen
LLM agents increasingly answer questions over dynamic raw-document collections, where files may change before preprocessing, and relevant evidence (spans, sections, pages, or tables) is query-dependent. Existing retrieval-augmented approaches pre-materialize evidence via fixed chunking, embeddings, or persistent indexes: effective for lookup, yet costly, stale-prone, and committed to a granularity before the query is known. We formulate in-context search as Budgeted Evidence Localization over a latent evidence space induced by dynamic raw documents and propose LENS (Latent Evidence Exploration and Search), an index-free framework. Instead of pre-materializing the evidence space, LENS maintains a query-conditioned belief over candidate units, iteratively selecting candidates via complementary lexical, local, and exploratory proposal policies, updating the belief via an LLM relevance oracle, and narrowing toward high-posterior regions under a controllable budget. Evidence is consolidated into compact, source-grounded regions of interest and compressed into self-organizing knowledge clusters reused across related queries. On a controlled 500-question evaluation with matched corpus snapshots, LENS reaches 62.4% exact match and 84.8% evidence recall vs. 65.2% exact match but 50.4% evidence recall for a ReAct-style baseline. Across scales, LENS gives the strongest supporting-fact localization and answer grounding. On a fixed 150-question fullwiki subset over the raw Wikipedia dump with zero indexing, LENS and ReAct are nearly tied in official answer quality (43.3% vs. 42.7% EM), with LENS grounding more answers in retrieved evidence (84.0% vs. 70.7%). A no-retrieval Closed-Book reference highlights the contribution of model memory. LENS is query-ready after corpus changes, needs no preprocessing or persistent index, and preserves source-grounded evidence localization throughout.
摘要:LLM 代理越來越多地在動態原始文件集合中回答問題,這些文件在預處理之前可能會發生變化,相關證據(範圍、部分、頁面或表格)依賴於查詢。現有的檢索增強方法通過固定分塊、嵌入或持久索引預先生成證據:對於查詢來說有效,但成本高、容易過時,並且在查詢已知之前就已確定了粒度。
我們將上下文搜索公式化為預算證據定位,這是在動態原始文件所誘導的潛在證據空間上進行的,並提出了 LENS(潛在證據探索與搜索),這是一個無索引的框架。
LENS 不是預先生成證據空間,而是維持對候選單位的查詢條件信念,通過互補的詞彙、局部和探索性提議策略迭代選擇候選者,通過 LLM 相關性預言者更新信念,並在可控預算下收斂到高後驗區域。
證據被整合成緊湊的、以來源為基礎的關注區域,並壓縮成自組織的知識集群,以便在相關查詢中重複使用。
在一個控制的 500 問題評估中,配對語料庫快照,LENS 達到 62.4% 的精確匹配和 84.8% 的證據召回,而 ReAct 標準基線則為 65.2% 的精確匹配和 50.4% 的證據召回。
在各種規模上,LENS 提供了最強的支持事實定位和答案基礎。在一個固定的 150 問題的 fullwiki 子集上,使用原始維基百科轉儲且沒有索引,LENS 和 ReAct 在官方答案質量上幾乎持平(43.3% 對 42.7% EM),而 LENS 在檢索證據中基礎了更多的答案(84.0% 對 70.7%)。
一個無檢索的閉卷參考突顯了模型記憶的貢獻。LENS 在語料庫變更後隨時可以查詢,無需預處理或持久索引,並在整個過程中保持來源基礎的證據定位。
Agent-Native Telemetry: Verifiable State-Delta Evidence for Autonomous Operations
2608.16178v1 by Jun He, Deying Yu
Operational telemetry is predominantly engineered for human reading: systems repeatedly serialize verbose prose, static keys, and redundant context across billions of log lines. As autonomous AI agents become primary operational consumers, feeding them traditional logs wastes scarce context capacity parsing lexical syntax rather than reasoning over system state changes -- all while lacking cryptographic guarantees of provenance or collection completeness. This paper introduces agent-native telemetry, an operational evidence architecture for autonomous machine operators founded on verifiable state deltas rather than human prose. We present the Agent Telemetry Protocol (ATP) and the State-Delta Evidence Ledger, an implementation that structures operational facts into four core evidence primitives (Transitions, Observations, Relations, and State Checkpoints) governed by content-addressed schemas, while isolating uncurated text as digest-verified opaque references. Producers sign and hash-chain batches for atomic collector append. Verified records feed two parallel agent access paths: a stateless protocol decoder emitting compact positional rows, and a stateful semantic gateway serving bounded graph capsules. We prove an information-preservation lower bound and formalize a ledger-relative verified negative theorem for provable event non-occurrence. On distributed microservice benchmarks (AIOpsLab and OpenTelemetry Astronomy Shop), ATP reduces raw wire payload and modeled cloud query scan costs by 96.4% relative to OpenTelemetry JSON, reduces LLM context tokens by 88.8% and query operations by 66.2%, detects all 500 tested adversarial storage mutations, and yields zero successful prompt injections across 50 adversarial trials per ATP configuration.
摘要:操作遙測主要是為了人類閱讀而設計:系統重複序列化冗長的散文、靜態鍵和冗餘的上下文,跨越數十億條日誌。隨著自主 AI 代理成為主要的操作消費者,向它們提供傳統日誌會浪費稀缺的上下文容量,解析詞法語法而不是推理系統狀態變化——同時缺乏來源或收集完整性的加密保證。
本論文介紹了代理原生遙測,這是一種基於可驗證狀態增量的自主機器操作員的操作證據架構,而非人類散文。我們提出了代理遙測協議 (ATP) 和狀態增量證據賬本,這是一種將操作事實結構化為四個核心證據原語(轉換、觀察、關係和狀態檢查點)的實現,受內容地址模式的管理,同時將未經策劃的文本隔離為摘要驗證的模糊參考。
生產者簽名並哈希鏈批次以進行原子收集器附加。經過驗證的記錄提供兩條平行的代理訪問路徑:一個無狀態的協議解碼器發出緊湊的位置信息行,和一個有狀態的語義網關提供有界圖膠囊。我們證明了信息保存的下限,並形式化了一個賬本相對的驗證負定理,以證明事件不發生的可證性。在分佈式微服務基準測試(AIOpsLab 和 OpenTelemetry Astronomy Shop)中,ATP 相對於 OpenTelemetry JSON 減少了 96.4% 的原始網絡有效載荷和建模雲查詢掃描成本,減少了 88.8% 的 LLM 上下文標記和 66.2% 的查詢操作,檢測到所有 500 次測試的對抗存儲突變,並在每個 ATP 配置的 50 次對抗試驗中產生零次成功的提示注入。
FeatureHospital: A Skill-Driven Multi-Agent Framework for Automated Algorithm Customization in Multi-View Multi-Label Feature Selection
2608.16148v1 by Junxuan Li, Zhiqi Chen, Yuzhou Liu, Peng Zhang, Huaxiao Liu
Multi-view multi-label feature selection aims to identify a compact and informative feature subset from heterogeneous views while preserving discriminative information for multiple labels. Existing methods are generally developed from specific modeling perspectives and incorporate mechanisms tailored to particular data characteristics. Designing suitable feature selection algorithms across datasets with diverse and heterogeneous characteristics still relies heavily on expert knowledge and substantial manual effort, imposing considerable time and labor costs that severely hinder the practical adoption of feature selection. To address this problem, we propose FeatureHospital, a Skill-driven multi-agent framework for automated multi-view multi-label feature selection algorithm design. FeatureHospital first diagnoses the target dataset to identify its feature selection issues. Based on the diagnosis, specialist agents equipped with domain Skills then prescribe corresponding optimization strategies and Loss terms for different issues. After that, the resulting prescriptions are reconciled to remove overlaps and resolve conflicts before being integrated into a compact dataset-specific objective. Finally, the constructed objective is optimized to select the final feature subset. Experimental results demonstrate that FeatureHospital can construct effective feature selection algorithms for different datasets based on their individual characteristics.
摘要:多視角多標籤特徵選擇旨在從異質視角中識別出緊湊且具信息量的特徵子集,同時保留多個標籤的區別信息。現有的方法通常是從特定建模角度發展而來,並納入針對特定數據特徵量身定制的機制。在具有多樣且異質特徵的數據集上設計合適的特徵選擇算法仍然在很大程度上依賴於專家知識和大量的手動努力,這帶來了可觀的時間和勞動成本,嚴重阻礙了特徵選擇的實際應用。為了解決這一問題,我們提出了FeatureHospital,一個以技能驅動的多代理框架,用於自動化多視角多標籤特徵選擇算法的設計。FeatureHospital首先診斷目標數據集,以識別其特徵選擇問題。根據診斷,配備領域技能的專家代理隨後為不同問題開出相應的優化策略和損失項。之後,生成的處方會進行調和,以消除重疊並解決衝突,然後整合成一個緊湊的數據集特定目標。最後,構建的目標被優化以選擇最終的特徵子集。實驗結果表明,FeatureHospital能夠根據不同數據集的特徵構建有效的特徵選擇算法。
Graph Neural Assisted Actor-Critic for Latency-Efficient Edge Vision System
2608.16142v1 by Alam Noor, Luis Almeida, Kai Li, Jiyan Wu, Miguel Gutiérrez Gaitán, Eduardo Tovar
UAV on-board vision systems are widely used for different activities, including monitoring in no-fly zones. In this case, the vision-equipped UAV streams a video to a ground server where an operator assists its activities. The latency of video transmission has a profound impact on the effectiveness of the operator assistance. However, most techniques available for video transmission still incur significant latency costs. In this paper, we propose a graph convolutional neural network-assisted (GCN-Assisted A2C) deep reinforcement learning (DRL) system model to find the optimal pixel-correlated area of a suspicious object. We combine the Lagrangian dual form with gradient descent to prevent lack of convergence and over- and under-penalization constraint violation during latency optimization. The proposed system model sends a sub-group pixel-correlated area of the frame from the UAV to the server rather than the transmission of the whole video frame. The proposed framework utilizes the GCN model to explore hidden representations of feature-correlated groups of pixels. Moreover, the GCN supervises the A2C model, which selects a subgroup to enhance transmission latency, thus supervising the training of UAV actions in A2C. Experimental results show that GCN-assisted A2C reduces video frame transmission latency together with false detection rate in UAV vision systems over other DRL and state-of-the-art models.
摘要:無人機的機載視覺系統被廣泛應用於不同的活動,包括在禁飛區的監控。在這種情況下,配備視覺系統的無人機將視頻流傳輸到地面伺服器,操作員在那裡協助其活動。視頻傳輸的延遲對操作員的輔助效果有深遠的影響。然而,目前大多數可用的視頻傳輸技術仍然會產生顯著的延遲成本。在本文中,我們提出了一種圖卷積神經網絡輔助(GCN輔助A2C)深度強化學習(DRL)系統模型,以尋找可疑物體的最佳像素相關區域。我們將拉格朗日對偶形式與梯度下降相結合,以防止在延遲優化過程中缺乏收斂以及過度和不足懲罰約束違規。所提出的系統模型將無人機的幀中一個子組像素相關區域發送到伺服器,而不是傳輸整個視頻幀。所提出的框架利用GCN模型探索特徵相關像素組的隱藏表示。此外,GCN監督A2C模型,該模型選擇一個子組以增強傳輸延遲,從而監督無人機在A2C中的行動訓練。實驗結果顯示,GCN輔助A2C在無人機視覺系統中減少了視頻幀傳輸延遲以及假檢測率,相較於其他DRL和最先進的模型。
HyperSkill: Self-Evolving LLM Agents via Hypergraph-Structured Skill Memory
2608.16114v1 by Ruiyao Xu, Tiankai Yang, Wei-Chieh Huang
As agentic tasks grow in complexity, LLM agents increasingly rely on experiential memory to reuse procedural knowledge across tasks. Effective memory design must jointly address what to store, how memory is structured and retrieved, and how memory evolves. Existing systems tackle each only partially: they store trajectories, insights, or workflows as isolated entries, discarding compositional relationships among subtasks and reusable skills; retrieve by flat embedding similarity that ignores relational signals; and maintain memory without leveraging its relational structure. We propose HyperSkill, a hypergraph-based memory framework that jointly improves all three. HyperSkill represents memory as a hypergraph with two node types, subtask steps and reusable skills, where each hyperedge links the subtasks and skills from a single trajectory. Dual-path retrieval queries both subtask and trajectory levels, ranking skills by co-occurrence across retrieved trajectories. Periodic structure-informed maintenance prunes low-utility nodes and merges redundant skills via quality-weighted propagation. Across xBench, GAIA, and WebWalkerQA with GPT-4o and Qwen3-30B-A3B, HyperSkill outperforms ten memory baselines, yielding gains of up to +11.51 on GAIA and +11.18 on WebWalkerQA.
摘要:隨著代理任務的複雜性增加,LLM代理越來越依賴經驗記憶在任務之間重用程序知識。有效的記憶設計必須共同解決存儲什麼、記憶的結構和檢索方式,以及記憶的演變。現有系統僅部分解決這些問題:它們將軌跡、見解或工作流程作為孤立的條目進行存儲,忽略了子任務和可重用技能之間的組合關係;通過忽略關聯信號的平面嵌入相似性進行檢索;並維護記憶而不利用其關聯結構。我們提出了HyperSkill,一種基於超圖的記憶框架,旨在共同改進這三個方面。HyperSkill將記憶表示為一個超圖,具有兩種類型的節點:子任務步驟和可重用技能,其中每個超邊連接來自單一軌跡的子任務和技能。雙路徑檢索查詢同時針對子任務和軌跡層級,根據檢索到的軌跡中的共同出現對技能進行排名。定期的結構信息維護修剪低效用節點,並通過質量加權傳播合併冗餘技能。在xBench、GAIA和WebWalkerQA中,使用GPT-4o和Qwen3-30B-A3B的HyperSkill超越了十個記憶基準,GAIA的增益高達+11.51,WebWalkerQA的增益高達+11.18。
RetroMPA: A Molecular Property-Aware Auxiliary Framework for Enhancing Retrosynthesis Prediction
2608.16111v1 by Mianzhi Liu, Fan Xiao, Zhiliang Yu, Huayang Huang, Yuke Li, Yi Yang, Wenbo Liu, Yu Wu
Retrosynthesis is a cornerstone of drug discovery and organic synthesis. While data-driven deep learning models have shown remarkable progress, they autonomously learn reaction patterns from extensive datasets with limited integration of established chemical knowledge as priors. To address this limitation, we introduce RetroMPA, a molecular property-aware, post-hoc enhancement module that injects chemical knowledge into the retrosynthesis pipeline. Rather than functioning as an independent SMILES sequence generator, RetroMPA is a broadly applicable, model-agnostic chemical filter designed to recalibrate and optimize the predictive pathways of existing algorithms. This plug-and-play framework integrates seamlessly with a range of data-driven retrosynthesis methods, enhancing outputs without modifying model architecture or requiring resource-intensive retraining. By leveraging a property-aware latent embedding space, RetroMPA consistently improves top-1 accuracy across eight representative retrosynthesis models by an average of 5.50% on USPTO-50K. Furthermore, we validate its scalability on the large-scale USPTO-Full dataset, achieving an average improvement of about 2.03% across both template-based and template-free architectures. Wet-lab experiments provide preliminary support for the practical utility of the framework. These syntheses confirmed viable, previously unreported substrate combinations for classic reaction paradigms---specifically, Suzuki-Miyaura coupling, Bucherer reaction, and Friedel-Crafts acylation---suggesting that RetroMPA can operate beyond mere data fitting. The code is open-sourced at https://github.com/MengzhouLu/RetroMPA.
摘要:逆合成是藥物發現和有機合成的基石。儘管數據驅動的深度學習模型已顯示出顯著的進展,但它們從大量數據集中自主學習反應模式,對已建立的化學知識的整合有限。
為了解決這一限制,我們引入了RetroMPA,一種分子性質感知的後處理增強模塊,將化學知識注入逆合成流程中。
RetroMPA並不是作為獨立的SMILES序列生成器運作,而是一種廣泛適用的、與模型無關的化學過濾器,旨在重新校準和優化現有算法的預測路徑。
這個即插即用的框架與多種數據驅動的逆合成方法無縫集成,增強輸出而不修改模型架構或需要資源密集的重新訓練。
通過利用性質感知的潛在嵌入空間,RetroMPA在八個代表性的逆合成模型上,平均提高了5.50%的top-1準確率,數據集為USPTO-50K。
此外,我們在大規模的USPTO-Full數據集上驗證了其可擴展性,在基於模板和無模板架構上都實現了約2.03%的平均改進。
濕實驗提供了對該框架實用性的初步支持。這些合成確認了經典反應範式下可行的、先前未報告的底物組合——具體而言,鈴木-宮浦偶聯、布赫勒反應和弗里德爾-克拉夫茲酰化——這表明RetroMPA可以超越單純的數據擬合。
代碼已開源於 https://github.com/MengzhouLu/RetroMPA。
The Commercial Tax: Rent-vs-Own Blind Spots in Multi-Hop Retrieval Benchmarks
2608.16096v1 by Luis M. Sanchez, Kosrow Dehnad
Enterprises connect language models to their own data through retrieval. The benchmarks that rank multi-hop retrieval systems leave out two facts a buyer needs before a published number can be used: whether the retrieval backbone may be deployed commercially, and what it costs to build. On licensing: the field's dense-retrieval anchor, NV-Embed-v2, is licensed cc-by-nc-4.0. Of the four leading MuSiQue systems we audit (HippoRAG-2, PropRAG, SAG, KET-RAG), three depend on it for their best numbers and none says so. On performance: we measure thirteen embedders from eight makers on one identical MuSiQue harness with bootstrap confidence intervals throughout. Until mid-2026 there was a real commercial tax: the best commercially-licensed embedder trailed the anchor by 2.31 Recall@5 points (95% CI [0.91, 3.71], p=0.001). NVIDIA's Nemotron-3-Embed-8B, released 2026-07-16, has closed it: +0.24 at Recall@5 (95% CI [-0.94, +1.43], p=0.69), -0.58 at Recall@10 (p=0.28). It matches the anchor, does not beat it, and is the only entrant that is commercially licensed, free to self-host, and indistinguishable from the anchor; every other entrant meeting the first two conditions sits 5.2 to 14.6 points below. The durable finding is the paid-versus-free divide: API embedders charge per token on every re-index, self-hosted ones charge nothing. On cost: three of five audited systems (adding Microsoft's GraphRAG) do not disclose indexing cost, and the only published GraphRAG dollar figures span 11x inside one third-party paper (USD 2.30 vs USD 24.94 to index a 5.64 MB corpus once); extrapolated to 1 TB that undisclosed choice separates roughly USD 428K from $4.6M. Our cost model keeps one-time embedding apart from recurring answering: at 1 TB, embedding sits 7.5x-900x below graph construction, and a year of answering at 10,000 queries/day sits 350x or more below it.
摘要:企業透過檢索將語言模型與自身數據連接起來。對於買家來說,在發佈的數字可以使用之前,有兩個事實是多跳檢索系統的基準未考慮的:檢索骨幹是否可以商業部署,以及建造的成本。關於授權:該領域的密集檢索錨點 NV-Embed-v2 採用 cc-by-nc-4.0 授權。在我們審核的四個主要 MuSiQue 系統(HippoRAG-2、PropRAG、SAG、KET-RAG)中,有三個依賴於它以獲得最佳數字,但沒有一個明言這一點。關於性能:我們測量了八家製造商的十三個嵌入器,並在一個相同的 MuSiQue 繫帶上進行測試,並在整個過程中使用自助信心區間。直到 2026 年中,存在一個實際的商業稅:最佳的商業授權嵌入器在 Recall@5 上落後於錨點 2.31 分(95% CI [0.91, 3.71],p=0.001)。NVIDIA 的 Nemotron-3-Embed-8B 於 2026-07-16 發佈,已經縮短了這一差距:在 Recall@5 上增加了 0.24(95% CI [-0.94, +1.43],p=0.69),在 Recall@10 上減少了 0.58(p=0.28)。它與錨點相匹配,但未超越它,並且是唯一一個商業授權、可自由自我托管且與錨點無法區分的參賽者;其他符合前兩個條件的參賽者則低於 5.2 至 14.6 分。持久的發現是付費與免費的區分:API 嵌入器在每次重新索引時按令牌收費,自我托管的則不收費。關於成本:五個審核系統中的三個(加上微軟的 GraphRAG)未披露索引成本,而唯一發佈的 GraphRAG 美元數字在一篇第三方論文中跨越了 11 倍(索引一次 5.64 MB 語料庫的成本為 USD 2.30 與 USD 24.94);推算到 1 TB,這一未披露的選擇大約將 USD 428K 與 $4.6M 隔開。我們的成本模型將一次性嵌入與重複回答區分開來:在 1 TB 時,嵌入成本低於圖形構建 7.5 倍至 900 倍,而一年內以 10,000 次查詢/天進行回答的成本則低於它 350 倍或更多。
Skill2Query: Exploiting Skill Structure to Generate Pseudo-Queries for Agent Skill Retrieval
2608.16071v1 by Lihui Ding, Zihan Guo, Bingwei Lu, Chenyu Zhou, Yuanjian Zhou, Weinan Zhang, Jianghao Lin, Dongdong Ge
Pseudo-query generation can alleviate the supervision bottleneck for agent skill retrieval, but existing document-level approaches typically leave the rich internal relations among capabilities, parameters, and usage examples implicit. As a result, generated queries may be topically relevant to a skill while lacking capability grounding and parameter consistency, raising the question of whether explicitly exploiting a skill document's internal structure can produce more effective retrieval signals. We therefore propose Skill2Query, a framework that first parses a skill document into a Skill Knowledge Graph and then generates pseudo-queries through a three-stage process including style mimicking, query template generation, and parameter filling. The generated queries can be used for offline index augmentation, online query expansion, and retriever training. Four benchmarks (TheoremQA, LogicBench, ToolQA, and CHAMP) are used to evaluate Skill2Query with large-scale skill candidate pools across multiple downstream applications, including skill retrieval, retriever training, and end-to-end agent execution. Using nearly 30K skills across diverse domains, we generate 700K category-diverse pseudo-queries. Skill2Query consistently improves sparse, dense, and skill-routing retrieval, with an average Recall@1 gain of 6.70 percentage points across retrieval settings. Skill2Query-generated training data also achieves the best Recall@1 and nDCG@1 among the evaluated generation baselines. Further evaluations with multiple LLM backends demonstrate that improved skill retrieval translates into higher agent task success rates. Code and resources are available at https://github.com/MatZaharia/Skill2Query.
摘要:偽查詢生成可以減輕代理技能檢索的監督瓶頸,但現有的文件級方法通常將能力、參數和使用範例之間豐富的內部關係隱含化。因此,生成的查詢可能在主題上與某項技能相關,但缺乏能力基礎和參數一致性,這引發了明確利用技能文件內部結構是否能產生更有效檢索信號的問題。因此,我們提出了Skill2Query,一個首先將技能文件解析為技能知識圖譜的框架,然後通過三個階段的過程生成偽查詢,包括風格模仿、查詢模板生成和參數填充。生成的查詢可以用於離線索引增強、在線查詢擴展和檢索器訓練。四個基準(TheoremQA、LogicBench、ToolQA和CHAMP)被用來評估Skill2Query,涵蓋多個下游應用中的大規模技能候選池,包括技能檢索、檢索器訓練和端到端代理執行。使用近30K個來自不同領域的技能,我們生成了700K個類別多樣的偽查詢。Skill2Query在稀疏、密集和技能路由檢索中始終提高了性能,平均Recall@1增益為6.70個百分點。Skill2Query生成的訓練數據在評估的生成基準中也實現了最佳的Recall@1和nDCG@1。對多個LLM後端的進一步評估顯示,改善的技能檢索轉化為更高的代理任務成功率。代碼和資源可在 https://github.com/MatZaharia/Skill2Query 獲得。
OceanLight: Efficient Global Ocean Forecasting via Geometry-Adaptive Unstructured Mesh Representation
2608.16070v1 by Wei Wu, Xiang Wang, Hongze Leng, Qingye Min, Junxing Zhu, Junqiang Song
Reliable global ocean forecasting is critical for climate monitoring, marine navigation, and extreme event early warning. Physics-based ocean forecasting models impose prohibitive computational costs, while existing deep learning approaches predominantly rely on structured-grid architectures, incurring unnecessary computation on masked land cells and enforcing uniform resolution across dynamically heterogeneous ocean regions regardless of local flow complexity. Here we present OceanLight, an efficient global ocean forecasting framework innovatively combining geometry-adaptive unstructured mesh tokenization with a graph neural network (GNN) backbone. OceanLight achieves pointwise forecast accuracy and kinetic energy spectral fidelity exceeding both operational numerical analyses and state-of-the-art AI-based models, while surpassing all AI-based ocean models in geostrophic balance consistency. Furthermore, OceanLight demonstrates reliable mesoscale eddy representation, capturing coherent ocean structures beyond pointwise statistical optimization. These capabilities are delivered with a 62% reduction in GPU memory consumption and 70\% reduction in FLOPs relative to structured-grid baselines. Our unstructured mesh representation establishes a generalizable paradigm for scalable data-driven oceanography.
摘要:可靠的全球海洋預測對於氣候監測、海洋導航和極端事件的早期預警至關重要。基於物理的海洋預測模型需要高昂的計算成本,而現有的深度學習方法主要依賴於結構化網格架構,對被遮蔽的陸地單元產生不必要的計算,並在動態異質的海洋區域強制執行均勻的解析度,無論當地流動的複雜性如何。在此,我們提出了OceanLight,一個高效的全球海洋預測框架,創新地結合了幾何自適應的非結構化網格標記與圖神經網絡(GNN)主幹。OceanLight實現了逐點預測準確性和動能能量譜的保真度,超過了操作性數值分析和最先進的基於AI的模型,同時在地轉平衡一致性方面超越了所有基於AI的海洋模型。此外,OceanLight展示了可靠的中尺度渦旋表徵,捕捉到超越逐點統計優化的連貫海洋結構。這些能力在GPU內存消耗上減少62%以及相對於結構化網格基準的FLOPs減少70%下得以實現。我們的非結構化網格表徵建立了一個可普遍化的範式,以支持可擴展的數據驅動海洋學。
NICE: Scale-Stable Perturbations for Graph Neural Network Explanations via Noise Corruption
2608.16038v2 by Ziluowen Luo, Jun Yin, Ruochen Liu, Ming Cheng, Shirui Pan, Chengqi Zhang, Senzhang Wang
Post-hoc Graph Neural Network (GNN) explainers commonly follow a Perturb-Query paradigm, inferring the importance of graph elements based on queried predictions to perturbed inputs. However, such perturbations often introduce substantial distribution shift, undermining the reliability of the queried predictions used to derive explanations. While existing efforts mainly improve perturbed graphs or stabilize model predictions on them, we revisit the perturbation mechanism itself. We show that the widely used Element-wise Masking(EM) suppresses edge-induced messages toward zero, causing deterministic scale contraction that accumulates across message-passing layers, a phenomenon we term Scale Drift. Consequently, prediction changes under EM may conflate information corruption with deviations in propagation scale. As a scale-stable alternative to EM, we introduce Noise Corruption (NC), which perturbs each message through matched-norm random-direction corruption while preserving the expected squared message norm. Building on NC, we propose NICE, a Noise Corruption-based explanation framework, which learns a Stochastic Restoration Boundary (SRB) under NC-induced uncertainty, balancing target-prediction restoration against compactness. Furthermore, Boundary-Integrated Gradient (BIG) converts this boundary into edge attributions by accumulating each edge's contribution to reducing restoration risk along the restoration path. Experiments across multiple benchmarks demonstrate stronger explanation performance and model faithfulness while confirming that NC substantially reduces the Scale Drift induced by masking.
摘要:後 hoc 圖神經網絡 (GNN) 解釋器通常遵循擾動-查詢範式,根據對擾動輸入的查詢預測推斷圖元素的重要性。
然而,這種擾動往往會引入顯著的分佈變化,削弱用於推導解釋的查詢預測的可靠性。
雖然現有的努力主要改善擾動圖或穩定模型在其上的預測,但我們重新審視擾動機制本身。
我們展示了廣泛使用的逐元素掩蔽 (EM) 將邊緣引起的消息壓制至零,導致確定性的縮放收縮,這一現象在消息傳遞層中累積,我們稱之為縮放漂移。
因此,EM 下的預測變化可能將信息損壞與傳播縮放的偏差混淆在一起。
作為 EM 的一種穩定縮放替代方案,我們引入了噪聲損壞 (NC),它通過匹配範數的隨機方向擾動每條消息,同時保持期望的平方消息範數。
基於 NC,我們提出了 NICE,一種基於噪聲損壞的解釋框架,它在 NC 引起的不確定性下學習隨機恢復邊界 (SRB),平衡目標預測恢復與緊湊性。
此外,邊界整合梯度 (BIG) 通過累積每條邊對減少恢復風險的貢獻,將這一邊界轉換為邊緣歸因。
在多個基準上的實驗顯示出更強的解釋性能和模型忠實度,同時確認 NC 顯著減少了由掩蔽引起的縮放漂移。
RagGAD: Rationale-Aware Conditional Gaussian Mixture Normalizing Flow for Unsupervised Graph Anomaly Detection
2608.16018v1 by Junxin Lu, Jing Zhao, Shiliang Sun
Graph anomaly detection aims to identify nodes that deviate from normal behavioral patterns within graphs. However, existing methods largely rely on the homophily assumption, which makes it difficult to distinguish spurious affinities and to capture the diverse behaviors of normal nodes,limiting their robustness in complex real-world scenarios. To address this problem, we propose RagGAD, an unsupervised graph anomaly detection framework based on rationale-aware conditional Gaussian mixture normalizing flow. RagGAD introduces an adaptive rationale disentangler to disentangle stable rationales from spurious correlations within node interrelationships, and further decomposes stable rationales into robust and fragile components. The learned rationales capture underlying interaction patterns that characterize normal behaviors under varying conditions, while anomalies emerge as deviations associated with unstable or spurious correlations. To model the intricate distributions of normal and abnormal nodes, RagGAD integrates rationale-non-rationale Gaussian mixture modeling with a robust-fragile rationale mixture learning strategy. By mitigating spurious homophilic correlations and embracing the heterogeneity of normal patterns, RagGAD identifies anomalies as low-density regions within a structure-aware distribution space. Extensive experiments on multiple benchmark datasets demonstrate that RagGAD outperforms state-of-the-art methods.
摘要:圖形異常檢測旨在識別在圖中偏離正常行為模式的節點。
然而,現有的方法在很大程度上依賴於同質性假設,這使得區分虛假的親和力和捕捉正常節點的多樣行為變得困難,限制了它們在複雜現實場景中的穩健性。
為了解決這個問題,我們提出了 RagGAD,一個基於理性感知條件高斯混合正規化流的無監督圖形異常檢測框架。
RagGAD 引入了一個自適應的理性解耦器,以從節點之間的關係中解耦穩定的理性與虛假的相關性,並進一步將穩定的理性分解為穩健和脆弱的組件。
學習到的理性捕捉了在不同條件下表徵正常行為的潛在互動模式,而異常則作為與不穩定或虛假相關性相關的偏差出現。
為了建模正常和異常節點的複雜分佈,RagGAD 將理性-非理性高斯混合建模與穩健-脆弱理性混合學習策略相結合。
通過減少虛假的同質相關性並接受正常模式的異質性,RagGAD 將異常識別為結構感知分佈空間中的低密度區域。
在多個基準數據集上的廣泛實驗表明,RagGAD 超越了最先進的方法。
From Sequence to Structure: Relational Uncertainty Propagation for LLM Agents
2608.16002v1 by Zhengzhao Ma. Boxi Cao, Yaojie Lu, Hongyu Lin, Xianpei Han, Le Sun
Reliable uncertainty quantification (UQ) is essential for deploying large language model (LLM) agents in complex interactive environments. Existing UQ methods largely rely on local signals, such as token probabilities, predictive entropy, or per-step confidence, and therefore overlook the long-range dependencies through which errors accumulate across an execution trajectory. As a result, they may fail to identify agent failures whose causes originate several reasoning or interaction steps before the final answer. We propose RUPA (Relational Uncertainty Propagation for Agents), a trajectory-level UQ framework for LLM agents. RUPA represents an execution history as a directed trajectory graph in which reasoning states, tool interactions, and environment feedback are nodes connected by temporal and semantic dependency edges. It then propagates uncertainty over this graph to capture how execution risk accumulates and transfers across interaction steps. The propagated signal is combined with trajectory-level behavioral features and goal-alignment information to produce a confidence estimate for the full agent trajectory. We evaluate RUPA on representative agent benchmarks, including $τ$-2, Terminal-Bench-2, and GAIA, using 6 open-source LLMs spanning multiple model families. Experimental results show that RUPA consistently outperforms existing UQ methods by providing more accurate uncertainty estimates, enabling earlier failure detection, and improving uncertainty-guided agent execution across diverse agent tasks. These results demonstrate that explicitly modeling relational dependency is crucial to reliable UQ for long-horizon LLM agents, providing a practical foundation for trustworthy agent execution.
摘要:可靠的不確定性量化(UQ)對於在複雜互動環境中部署大型語言模型(LLM)代理至關重要。現有的UQ方法主要依賴於局部信號,例如標記概率、預測熵或每步信心,因此忽略了錯誤在執行軌跡中累積的長期依賴關係。結果,它們可能無法識別那些原因源於幾個推理或互動步驟之前的代理失敗。我們提出了RUPA(代理的關聯不確定性傳播),這是一個針對LLM代理的軌跡級UQ框架。RUPA將執行歷史表示為一個有向軌跡圖,其中推理狀態、工具互動和環境反饋是由時間和語義依賴邊連接的節點。然後,它在這個圖上傳播不確定性,以捕捉執行風險如何在互動步驟中累積和轉移。傳播的信號與軌跡級行為特徵和目標對齊信息相結合,以產生對整個代理軌跡的信心估計。我們在代表性的代理基準上評估RUPA,包括$τ$-2、Terminal-Bench-2和GAIA,使用6個跨多個模型系列的開源LLM。實驗結果顯示,RUPA通過提供更準確的不確定性估計、實現更早的失敗檢測以及改善不確定性引導的代理執行,始終優於現有的UQ方法。這些結果表明,明確建模關聯依賴性對於長期LLM代理的可靠UQ至關重要,為可信的代理執行提供了實用的基礎。
PLSQLBench: Benchmarking LLM Systems for Executable Procedural Database Programming
2608.15931v1 by Marianne Menglin Liu, Leonid Boytsov, Daniel W. Peterson, Pramuditha Perera, Rongguang Wang, Sai Ashish Somayajula, Syed Hamza Rafique, Rohit Saini, Shubham Pathak, Sujeeth Bharadwaj, Tao Sheng, Graham Horwood, Fahad Shah, Ankan Bansal, Sujith Ravi, Dan Roth
We present PLSQLBench, to our knowledge the first benchmark for evaluating whether LLMs can write executable PL/SQL programs, with correctness measured through execution-based tests. Existing LLM evaluations largely target general-purpose code generation or declarative text-to-SQL, leaving procedural database programming underexplored. PLSQLBench contains 2,865 instances: 2,594 single-turn tasks and 271 multi-turn conversations spanning 978 turns. The benchmark combines complex schema-grounded tasks over enterprise-style Spider 2 databases, simpler schema-grounded tasks derived from Spider, and MBPP-derived procedural problems, covering varying levels of database grounding and procedural complexity. Experiments with eight LLMs reveal recurring difficulties in schema grounding, PL/SQL dialect fidelity, procedural control flow, exception handling, and cross-turn consistency. Tool-augmented LLM agents improve performance on several schema-grounded evaluations, although substantial gaps remain. These results highlight procedural database programming capabilities not directly assessed by conventional code generation or text-to-SQL benchmarks. Our code is available at https://github.com/oracle-samples/plsqlbench.
摘要:我們介紹PLSQLBench,據我們所知,這是第一個用於評估LLMs是否能夠編寫可執行PL/SQL程序的基準,其正確性通過基於執行的測試來衡量。現有的LLM評估主要針對通用代碼生成或聲明式文本到SQL,導致程序性數據庫編程未得到充分探索。PLSQLBench包含2,865個實例:2,594個單回合任務和271個跨978回合的多回合對話。該基準結合了基於企業風格Spider 2數據庫的複雜架構任務、源自Spider的較簡單架構任務以及MBPP衍生的程序性問題,涵蓋了不同級別的數據庫基礎和程序複雜性。對八個LLM的實驗揭示了在架構基礎、PL/SQL方言忠實度、程序控制流程、異常處理和跨回合一致性方面的重複困難。工具增強的LLM代理在幾個架構基礎評估中提高了性能,儘管仍然存在重大差距。這些結果突顯了傳統代碼生成或文本到SQL基準未直接評估的程序性數據庫編程能力。我們的代碼可在https://github.com/oracle-samples/plsqlbench獲得。
Unified Pedestrian Path Prediction Using Inverse Reinforcement Learning
2608.15929v1 by Šimon Sukup, Ariyan Bighashdel, Pavol Jancura
Pedestrian path prediction is crucial for enhancing the safety of autonomous vehicles and advanced driver-assistance systems. Previous studies explored different learning-task formulations for pedestrian path prediction and compared these formulations using shallow neural networks, but did not extend this analysis to more complex deep-learning models. This paper adapts the Spatial-Temporal Graph Attention Network (STGAT) to a unified pedestrian path prediction framework and introduces state and action definitions specific to STGAT. The resulting formulations support deterministic and stochastic policies, one-time and sequential decision-making, and reinforcement-learning algorithms including REINFORCE and proximal policy optimization. The proposed learning-task formulations improve prediction performance across the selected benchmark datasets compared with the standard supervised-learning formulation. These results demonstrate that reformulating the decision process and training objective can improve an advanced pedestrian trajectory prediction architecture and may provide a path toward improving other graph-based prediction models.
摘要:行人路徑預測對於增強自動駕駛車輛和先進駕駛輔助系統的安全性至關重要。
以往的研究探討了行人路徑預測的不同學習任務公式,並使用淺層神經網絡比較了這些公式,但並未將此分析擴展到更複雜的深度學習模型。
本文將空間-時間圖注意力網絡(STGAT)適應於統一的行人路徑預測框架,並引入特定於STGAT的狀態和行動定義。
所得到的公式支持確定性和隨機策略、一次性和序列決策,以及包括REINFORCE和近端政策優化在內的強化學習算法。
與標準的監督學習公式相比,所提出的學習任務公式在所選基準數據集上提高了預測性能。
這些結果表明,重新構建決策過程和訓練目標可以改善先進的行人軌跡預測架構,並可能為改善其他基於圖的預測模型提供一條途徑。
Noesis: Bidirectional Graph-RAG with Adaptive Parallelism and Cross-Knowledge-Base Semantic Discovery
2608.15919v1 by Nicola Cogotti
Retrieval-Augmented Generation over knowledge graphs (Graph-RAG) has emerged as a powerful paradigm for grounding large language models in domain-specific corpora. However, existing systems face persistent limitations: (1) static chunking fragments long documents, losing cross-section semantic connections; (2) ingestion pipelines do not scale adaptively; and (3) multi-domain deployments require either a monolithic knowledge base that dilutes retrieval precision or manual user routing. We present Noesis, a decoupled Graph-RAG architecture addressing these limitations through four algorithms: (a) Bidirectional Graph Traversal with a Graph-Feedback Context Resolver simulating human reading with degrading memory; (b) an AIMD Concurrency Controller adapted from TCP congestion control, achieving 23x speedup with zero OOM events; (c) Moesis, domain-aware selective quantization for MoE models achieving 6.3x speedup on 12 GB consumer GPUs; and (d) Mesh, cross-KB semantic routing with runtime structural discovery enabling small on-premises models to perform multi-hop cross-domain reasoning. On HotpotQA (1,000 questions), Noesis achieves 59.5 EM / 74.7 F1, surpassing GraphRAG by +27.8 EM while using a 35B on-premises model for graph construction rather than GPT-4o. Source text verification on a 193-page document confirms 90% precision on long-range causal edges inaccessible to chunk-independent extraction.
摘要:檢索增強生成(Retrieval-Augmented Generation)在知識圖譜(Graph-RAG)上已成為一種強大的範式,用於將大型語言模型嵌入特定領域的語料庫中。
然而,現有系統面臨持續的限制:(1)靜態分塊使長文檔碎片化,失去了交叉部分的語義連接;(2)攝取管道無法自適應擴展;(3)多領域部署需要一個單體知識庫,這會稀釋檢索精度或需要手動用戶路由。
我們提出了 Noesis,一種解耦的 Graph-RAG 架構,通過四個算法解決這些限制:(a)雙向圖遍歷(Bidirectional Graph Traversal)與圖反饋上下文解析器(Graph-Feedback Context Resolver),模擬人類閱讀並隨著記憶衰退;(b)從 TCP 擁塞控制改編的 AIMD 並發控制器(AIMD Concurrency Controller),實現了 23 倍的加速,且無 OOM 事件;(c)Moesis,針對 MoE 模型的領域感知選擇性量化,實現了在 12 GB 消費者 GPU 上的 6.3 倍加速;以及(d)Mesh,跨知識庫的語義路由,通過運行時結構發現使小型本地模型能夠執行多跳跨域推理。
在 HotpotQA(1,000 個問題)上,Noesis 實現了 59.5 EM / 74.7 F1,超越了 GraphRAG,EM 提升了 27.8,同時使用 35B 的本地模型進行圖構建,而不是 GPT-4o。
對於一份 193 頁文檔的源文本驗證確認,對於長距離因果邊的精度達到 90%,這些邊是無法通過獨立於塊的提取來訪問的。
Large language model-assisted discovery of cohorts from scientific literature
2608.15909v1 by Moritz Sturm, Lisa M. Berg, Inken Berg, Harishny Sarma, Jasmin Hartmann, Denissa Girschik, Gemma Roig, Christine M. Freitag, Andreas G. Chiocchetti
Background: Planning multi-study analyses requires identifying cohorts with the relevant participants, phenotypes, and data modalities. This process commonly relies on prior knowledge, cohort catalogues, and manual literature searches. We developed a complementary question-driven framework that searches relevant scientific literature and extracts explicit cohort names. Methods: The framework first generates multiple PubMed queries from configurable vocabularies and templates and retrieves the resulting scientific literature automatically through the PubMed API. A large language model then screens the retrieved titles and abstracts and extracts explicit cohort names using a prompt tailored to the research question. The extracted names are deduplicated with human review. Configurable code, prompts, and example outputs are available at https://gitlab.rz.uni-frankfurt.de/cap_molgenlab/literature-cohort-discovery. Evaluation: As a use case, we applied the framework to youth aggression genetics. From 5,400 generated PubMed queries, the framework retrieved 5,254 unique records and identified 188 candidate cohorts. Manual screening using predefined criteria, including participant age and genetic-data availability, retained 44 eligible cohorts. Automated LLM-based name extraction was within the agreement range of human annotators. We also searched four established cohort catalogues using the same research question. Their combined results contained 27 of the 44 eligible cohorts, while 17 were not returned by any cohort catalogue search. Conclusion: The framework converts research-question-specific vocabulary into screenable cohort inventories via a large, automated literature search. It can be adapted across populations, phenotypes, data modalities, and study designs, and provides a literature-based complement to curated cohort catalogues.
摘要:背景:規劃多研究分析需要識別具有相關參與者、表型和數據模態的隊列。這一過程通常依賴於先前的知識、隊列目錄和手動文獻搜索。我們開發了一個補充的問題驅動框架,該框架搜索相關的科學文獻並提取明確的隊列名稱。方法:該框架首先從可配置的詞彙和模板生成多個PubMed查詢,並通過PubMed API自動檢索結果科學文獻。然後,大型語言模型篩選檢索到的標題和摘要,並使用針對研究問題量身定制的提示提取明確的隊列名稱。提取的名稱經過人工審核去重。可配置的代碼、提示和示例輸出可在 https://gitlab.rz.uni-frankfurt.de/cap_molgenlab/literature-cohort-discovery 獲得。評估:作為一個用例,我們將該框架應用於青少年攻擊性遺傳學。從5,400個生成的PubMed查詢中,該框架檢索到5,254個唯一記錄並識別了188個候選隊列。使用預定義標準進行的人工篩選,包括參與者年齡和基因數據可用性,保留了44個合格隊列。自動化的LLM基礎名稱提取與人類標註者的協議範圍內。 我們還使用相同的研究問題搜索了四個已建立的隊列目錄。它們的綜合結果包含44個合格隊列中的27個,而17個則未被任何隊列目錄搜索返回。結論:該框架將特定於研究問題的詞彙轉換為可篩選的隊列清單,通過大型自動化文獻搜索。它可以適應不同的人群、表型、數據模態和研究設計,並為策劃的隊列目錄提供文獻基礎的補充。
Scaling Manual-Grounded Appliance Manipulation with Data Synthesis and Unified Planning
2608.15863v1 by Yuxing Long, Lei Kang, Ziyan Yu, Yuzheng Gao, Bin Cheng, Jiyao Zhang, Xiaoqi Li, Haolin Yang, Dongjiang Li, Hui Shen, Hao Dong
Operating household appliances requires long-horizon planning that is state-dependent and robust to disturbances, yet existing large models fall short, as no sufficiently diverse, task-oriented dataset exists to support such planning. To bridge this gap, we propose MAGE, a scalable data synthesis pipeline that introduces a novel Hierarchical Appliance Graph (HAG) to automatically generate part grounding, long-horizon planning, and closed-loop recovery data from appliance manuals. With MAGE, we build UseAppliance, the first large-scale dataset for manual-grounded appliance manipulation planning, spanning 22 appliance categories with 89K+ part annotations, 53K+ manipulation tasks, and 33K+ closed-loop adjustment steps. Built on UseAppliance, we develop AppliancePlan, an end-to-end model for manual-grounded appliance manipulation planning. On RealAppliance-Bench, AppliancePlan with only 7B parameters achieves over 10x the best baseline on open-loop planning and consistently outperforms state-of-the-art models across all tasks. Real-robot experiments on six household appliances further confirm effective sim-to-real transfer, marking an important step toward general-purpose household robotics.
摘要:操作家用電器需要長期規劃,這種規劃依賴於狀態並且對擾動具有穩健性,然而現有的大型模型卻無法滿足需求,因為沒有足夠多樣化且以任務為導向的數據集來支持這種規劃。為了填補這一空白,我們提出了MAGE,一個可擴展的數據合成管道,該管道引入了一種新穎的層次家電圖(HAG),以自動從家電手冊生成部件定位、長期規劃和閉環恢復數據。通過MAGE,我們構建了UseAppliance,這是第一個基於手冊的家電操作規劃的大型數據集,涵蓋22個家電類別,擁有89K+的部件註釋、53K+的操作任務和33K+的閉環調整步驟。在UseAppliance的基礎上,我們開發了AppliancePlan,一個端到端的基於手冊的家電操作規劃模型。在RealAppliance-Bench上,僅有7B參數的AppliancePlan在開放式規劃上達到了超過10倍的最佳基準,並在所有任務中持續超越最先進的模型。對六種家用電器的真實機器人實驗進一步確認了有效的模擬到現實轉移,這標誌著朝向通用家用機器人邁出了重要一步。
RAGas: Retrieval-Augmented Gas Optimization for Smart Contracts with Continuous Knowledge Integration
2608.15857v1 by Yishun Wang, Wenjin Yi, Wenkai Li, Zongwei Li, Xiaoqi Li
Ethereum is now integral to mission-critical sectors, including finance, healthcare, and supply chain management. Execution fees, commonly referred to as Gas, scale with the computational complexity of their functions. Smart contracts on Ethereum incur execution fees, known as Gas, which increase with computational complexity. Thus, optimizing Gas-intensive code while preserving functional equivalence significantly lowers deployment costs. No existing system continuously exploits evolving Gas usage patterns. We systematically analyze syntactic and semantic constructs that drive excessive Gas use. This yields six high-level categories covering twelve fine-grained antipatterns underpinning a curated knowledge base. We operationalize these insights with RAGas, a three-stage retrieval-augmented generation framework that uses a large language model to pinpoint and automatically fix Gas inefficiencies. Experiments on deployed contracts demonstrate that RAGas reduces Gas usage by up to 11% and achieves high precision and recall in detecting code snippets exhibiting Gas wastage.
摘要:Ethereum 現在對於關鍵任務領域至關重要,包括金融、醫療保健和供應鏈管理。執行費用,通常稱為 Gas,隨著其功能的計算複雜性而增加。以太坊上的智能合約會產生執行費用,稱為 Gas,這些費用隨著計算複雜性的增加而上升。因此,在保持功能等價的同時優化高 Gas 消耗的代碼,可以顯著降低部署成本。現有系統無法持續利用不斷演變的 Gas 使用模式。我們系統地分析驅動過度 Gas 使用的語法和語義結構。這產生了六個高層次類別,涵蓋了十二個細緻的反模式,支撐著一個策劃的知識庫。我們利用這些洞見實現 RAGas,一個三階段的檢索增強生成框架,使用大型語言模型來精確定位並自動修復 Gas 效率低下的問題。對已部署合約的實驗表明,RAGas 將 Gas 使用量降低了多達 11%,並在檢測顯示 Gas 浪費的代碼片段時達到了高精度和高召回率。
Characterising cardiac tissue properties with graph neural networks
2608.15843v1 by Ching-En Chiu, Yoo Ri Kim, Magdi Saba, Danilo Mandic, Marta Varela
Characterising electrophysiological properties of cardiac tissue efficiently and accurately from spatially sparse intracardiac measurements is clinically important for localising ablation targets and improving arrhythmia treatment. We developed a graph neural network-based framework trained on synthetic electrogram signals on 2D flat surfaces to identify areas of interest in the context of cardiac ablation for premature ventricular complexes (PVCs). Our method achieved an average precision of 0.96, 0.97, and 0.95 for the detection of single-patch fibrosis, rapid depolarisation and high excitability, respectively. The trained model can then be applied to 2D curved surfaces with few-shot fine-tuning, demonstrating its generalisation capability. Future work will develop this framework further for clinical use in PVC ablation.
摘要:有效且準確地從空間稀疏的心內測量中描述心臟組織的電生理特性,對於定位消融目標和改善心律不整治療具有臨床重要性。
我們開發了一個基於圖神經網絡的框架,該框架在2D平面上對合成電圖信號進行訓練,以識別在心臟消融中與早期心室複雜(PVCs)相關的興趣區域。
我們的方法在檢測單一斑塊纖維化、快速去極化和高興奮性方面,分別達到了0.96、0.97和0.95的平均精度。
訓練好的模型可以通過少量調整應用於2D曲面,顯示出其泛化能力。
未來的工作將進一步開發這一框架,以便在PVC消融中用於臨床應用。
Schema-Agnostic Graph Reasoning Agent for Hybrid Knowledge Graphs
2608.15834v1 by Marius Dragic, Ruben Ifrah, Alexandre Rio
Tool-calling LLM agents navigate unfamiliar codebases with a handful of generic primitives for listing, reading and searching files (ls, cat, grep). A knowledge graph admits the same interface: listing neighbours, reading node content and searching descriptions are the same operations on a different substrate. Building on this correspondence, we present GRA, a Graph Reasoning Agent that explores hybrid knowledge graphs, whose nodes are either textual concepts or relational tables, with seven generic tools, discovering everything domain-specific at run time. On UFK-M (Unified Factory Knowledge Model), an industrial benchmark of 258 analytical questions whose gold answers are produced by executing validated SQL programs, GRA beats a full-context agent by 5.1 pp (88.4% vs. 83.3%), while reading under a third of its input tokens. A graph-free control shows the gain comes chiefly from selective agentic access rather than graph topology, and that the effect depends on a model able to drive tools reliably. Seeing less, the agent answers better: selective navigation over a structured substrate beats exhaustive context.
摘要:工具調用的 LLM 代理使用一小部分通用原語來導航不熟悉的代碼庫,以列出、閱讀和搜索文件(ls、cat、grep)。知識圖譜承認相同的接口:列出鄰居、閱讀節點內容和搜索描述在不同的基質上是相同的操作。基於這一對應關係,我們提出了 GRA,一個探索混合知識圖譜的圖推理代理,其節點可以是文本概念或關聯表,並使用七個通用工具,在運行時發現所有特定於領域的內容。在 UFK-M(統一工廠知識模型)上,這是一個包含 258 個分析問題的工業基準,其金標答案是通過執行經過驗證的 SQL 程序生成的,GRA 以 5.1 個百分點的優勢擊敗了全上下文代理(88.4% 對 83.3%),同時閱讀的輸入標記不到其三分之一。一個無圖控制顯示,這一增益主要來自於選擇性代理訪問,而非圖拓撲,並且這一效果依賴於能夠可靠驅動工具的模型。看到的越少,代理回答得越好:在結構化基質上進行選擇性導航優於全面上下文。
The Authority Resolution Framework: A Five-Domain Ontology for Governing Who and What Decides, at Scale
2608.15832v1 by Parviz Shariff
As AI systems become increasingly capable of autonomous action, determining whether an agent is technically capable of performing an action is insufficient: the system must also determine whether the action is authorised in its context. This paper introduces the Authority Resolution Framework (ARF), a five-domain ontology for representing and resolving authority across organisational roles and informal influence, business concepts, codified processes, machine-readable permissions and executable systems, and external real-world context. ARF defines the Authority Relation (AR) as a cross-domain primitive binding an actor, action, object, bounded context, justification chain, and a calibration measure termed the DNA-Coefficient, which captures divergence between documented authority structures and authority as practiced. The framework provides a machine-interpretable representation of authority provenance and scope, with JSON-LD representations and knowledge-graph query patterns for authority resolution. ARF is designed to support AI agents in determining the provenance, scope and contextual validity of authority before executing consequential actions. The framework positions authority resolution as a knowledge-representation and reasoning problem at the intersection of ontology engineering, semantic AI, agentic AI and AI governance.
摘要:隨著人工智慧系統越來越能夠自主行動,僅僅判斷一個代理是否在技術上能夠執行某個行動是不夠的:系統還必須確定該行動在其上下文中是否被授權。
本文介紹了權限解析框架(Authority Resolution Framework, ARF),這是一個五個領域的本體,用於表示和解決組織角色和非正式影響、商業概念、編碼過程、機器可讀的許可和可執行系統,以及外部現實世界上下文中的權限。
ARF 將權限關係(Authority Relation, AR)定義為一個跨領域的原始綁定,將行為者、行動、對象、有限上下文、理由鏈以及一個稱為 DNA-係數的校準度量結合在一起,該度量捕捉了記錄的權限結構與實際執行的權限之間的差異。
該框架提供了一種機器可解釋的權限來源和範圍的表示,並提供 JSON-LD 表示和知識圖譜查詢模式以進行權限解析。
ARF 設計旨在支持 AI 代理在執行有後果的行動之前,確定權限的來源、範圍和上下文有效性。
該框架將權限解析定位為一個知識表示和推理問題,位於本體工程、語義 AI、代理 AI 和 AI 治理的交集處。
QuantumPhaseNet: A Gauge-Covariant Geometric and Quantum-Spectral Theory of Semantic Concept Hierarchies with Prototype Validation of a Classical Quantum-Inspired Model
2608.15820v1 by Kiyotaka Kasubuchi, Kazuo Fukiya
We present QuantumPhaseNet, a gauge-covariant geometric and quantum-spectral extension of Transformer representations. Context-dependent semantic states are modeled as complex amplitudes; a covariant phase rate induces a semantic wavelength used as a proxy for conceptual scale; and low-frequency graph modes define a document-level discourse direction. The theoretical part establishes local gauge invariance, unitarity of the quantum block, boundedness and conditional stability of WavePhase Attention, and a calibratable hallucination-risk formulation. We also implemented a fully offline Validation Studio for the classical quantum-inspired pipeline in Section 14.1 and evaluated the five research questions in Section 16.1 on its built-in synthetic setting (n=240, observation noise 0.22, circuit noise 0.08, five seeds). RQ1 yielded a wavelength-hierarchy Spearman correlation of 0.852 versus 0.707 for the baseline, 87.3% direction accuracy, and AUC 0.953. RQ2 achieved discourse alignment 0.933 versus 0.589 and 41.2 versus 16.2 paragraphs before drift. RQ3 achieved AUROC 0.881 versus cosine 0.765 and phase-shuffle 0.536. RQ4 achieved error-detection AUROC 0.854 versus entropy 0.634, with Brier 0.150 and ECE 0.098. RQ5 did not show quantum advantage: target probability and end-to-end cost efficiency were 25.5% and 0.107, compared with 70.7% and 0.707 for the Chebyshev classical approximation. These results provide initial synthetic evidence for the classical quantum-inspired components, but not external validity or unconditional quantum speedup.
摘要:我們提出了QuantumPhaseNet,這是一種與規範協變的幾何和量子頻譜擴展的Transformer表示。上下文依賴的語義狀態被建模為複數振幅;協變相位速率引入了一種語義波長,作為概念尺度的代理;而低頻圖模式定義了文檔層級的話語方向。理論部分建立了局部規範不變性、量子區塊的單位性、WavePhase注意力的有界性和條件穩定性,以及可校準的幻覺風險公式。我們還為第14.1節中的經典量子啟發管道實施了一個完全離線的驗證工作室,並在第16.1節中對其內建的合成設置(n=240,觀察噪聲0.22,電路噪聲0.08,五個種子)評估了五個研究問題。RQ1產生了波長層次的Spearman相關性0.852,而基準為0.707,方向準確率87.3%,AUC 0.953。RQ2實現了話語對齊0.933,而基準為0.589,並且在漂移之前有41.2與16.2段落。RQ3實現了AUROC 0.881,而餘弦為0.765,隨機相位為0.536。RQ4實現了錯誤檢測AUROC 0.854,而熵為0.634,Brier為0.150,ECE為0.098。RQ5未顯示量子優勢:目標概率和端到端成本效率分別為25.5%和0.107,而Chebyshev經典近似為70.7%和0.707。這些結果為經典量子啟發組件提供了初步的合成證據,但不具備外部有效性或無條件的量子加速。
ALKEMIE Agent: an autonomous platform for computational materials design
2608.15776v1 by Hongfu Huang, Yuzhe Li, Ao Xu, Bo Liu, Changrui Wang, Kan Tang, Ning Yang, Shengxian Liu, Hanyu Liu, Pengpeng Zhang, Linggang Zhu, Fengkai Liu, Yichen Lu, Tong Zhao, Naihua Miao, Jian Zhou, Zhimei Sun
Despite the powerful multi-scale modeling methods and high-throughput infrastructures established in the materials community, real material computation workflows remain fragmented and heavily manual, requiring researchers to constantly bridge software tools, data analysis, and intermediate decisions. This growing gap between methodological capability and practical execution highlights the need for a new kind of autonomous computational framework, one that can coordinate tools, knowledge, and workflows in a more unified and adaptive way. Here, we introduce ALKEMIE Agent, an agentic platform in which retrieval-augmented generation, a materials-computation knowledge base, registered skills, database-supported provenance, AI-assisted structure modeling, bounded task execution, tool-calling iteration, and error-diagnostic assistance are integrated within a traceable control loop. The capabilities of ALKEMIE Agent are demonstrated through applications including materials recommendation, structure modeling, phonon calculations, machine-learned interatomic potential training, LAMMPS simulations, Ab Initio Monte Carlo (AIMC) sampling, and active-learning-based materials screening. Finally, we outline the future directions and challenges for the development of agentic platforms for computational materials design.
摘要:儘管材料社群中已建立強大的多尺度建模方法和高通量基礎設施,但真正的材料計算工作流程仍然是碎片化且高度手動的,這要求研究人員不斷地橋接軟體工具、數據分析和中間決策。這種方法能力與實際執行之間日益擴大的鴻溝突顯了對一種新型自主計算框架的需求,這種框架能以更統一和適應的方式協調工具、知識和工作流程。在這裡,我們介紹了ALKEMIE Agent,一個代理平台,其中檢索增強生成、材料計算知識庫、註冊技能、數據庫支持的來源、AI輔助結構建模、有限任務執行、工具調用迭代和錯誤診斷協助被整合在一個可追蹤的控制迴路中。ALKEMIE Agent的能力通過包括材料推薦、結構建模、聲子計算、機器學習的原子間勢訓練、LAMMPS模擬、Ab Initio Monte Carlo (AIMC) 取樣和基於主動學習的材料篩選等應用得以展示。最後,我們概述了計算材料設計的代理平台未來的方向和挑戰。
Large Models for Small Devices: Recent Advances and Empirical Analysis of Edge AI Deployment
2608.15693v1 by Subhransu Das, Jiaming Cheng, Arnav Kumar, Sadia Afrose, Mingzhe Han, Michael Silagy, Shreya Palande, Brijesh Soni, Rajiv Ramnath
Running large AI models on resource-constrained edge devices requires model compression to reduce model size and computation. What compresses well, however, need not deploy well. We survey dozens of recent works that report compression results on real hardware and extract practical deployment guidelines from them. Following these guidelines, we deploy compact language and image models on GPU, CPU, and Raspberry Pi platforms across question answering and image segmentation. No single technique wins across tasks. For question answering, Qwen3.5 0.8B reaches 93.85 SQuAD F1 and 92 EM under Q5_K_M GGUF quantization, while structured pruning at the same precision costs 16 F1 at a 1% ratio. For segmentation, the ranking reverses: default quantization leaves parameters and MACs unchanged, whereas pruning cuts model size by nearly 80% at near-constant mIoU. Pruning can even inflate the deployed artifact by 21-49% by breaking k-quant super-block alignment; combined with longer, less format-compliant outputs, this raises Raspberry Pi latency up to 3.4x. Compression can also manufacture the appearance of competence rather than destroy it visibly: one LoRA-recovered variant stays fully parseable and holds 71% strict BoolQ accuracy while sending 97 of 100 predictions to a single class, at 52.6% balanced accuracy. We explain these effects through neural-flow graph analysis and prefill-decode-level latency decomposition, and condense them into task-specific deployment research directions. The right technique depends on the task, the model, and the hardware. Our experiment code and artifacts are open-sourced at https://github.com/Arnavvvkumar/deployment
摘要:在資源受限的邊緣設備上運行大型 AI 模型需要模型壓縮,以減少模型大小和計算量。 然而,壓縮效果良好的模型不一定能夠良好部署。 我們調查了數十篇最近的研究,這些研究報告了在實際硬體上進行的壓縮結果,並從中提取實用的部署指南。 根據這些指南,我們在 GPU、CPU 和 Raspberry Pi 平台上部署了緊湊的語言和圖像模型,應用於問題回答和圖像分割。 沒有單一技術在所有任務中都能獲勝。 在問題回答中,Qwen3.5 0.8B 在 Q5_K_M GGUF 量化下達到 93.85 的 SQuAD F1 和 92 的 EM,而在相同精度下的結構化剪枝則以 1% 的比例損失 16 的 F1。 對於分割,排名則顛倒:默認量化保持參數和 MACs 不變,而剪枝則在幾乎不變的 mIoU 下將模型大小減少近 80%。 剪枝甚至可以通過打破 k-quant 超塊對齊來使部署的工件膨脹 21-49%;結合更長且格式不合規的輸出,這使得 Raspberry Pi 的延遲增加至 3.4 倍。 壓縮還可以製造出能力的外觀,而不是明顯地摧毀它:一個 LoRA 恢復的變體保持完全可解析,並在將 100 次預測中的 97 次發送至單一類別的同時,保持 71% 的嚴格 BoolQ 準確率,平衡準確率為 52.6%。 我們通過神經流圖分析和預填充解碼級延遲分解解釋這些效果,並將其濃縮為特定任務的部署研究方向。 正確的技術取決於任務、模型和硬體。 我們的實驗代碼和工件已在 https://github.com/Arnavvvkumar/deployment 開源。
BERTopic-Virality Prioritisation: A Scalable Framework for Thematic and Comparative Analysis of COVID-19 and Monkeypox Misinformation on Twitter
2608.15691v1 by Mkululi Sikosana, Sean Maudsley-Barton, Oluwaseun Ajao
Health misinformation circulating during pandemics can gain traction rapidly, creating harmful narratives that compete with public health guidance. Most topic-modelling pipelines treat engagement as an external outcome, limiting their ability to prioritise semantically coherent topics that are also rapidly diffusing. We introduce BERTopic-VP, a virality-prioritised topic-modelling framework that combines contextual embedding-based clustering (BERTopic) with a post hoc Virality Prioritisation (VP) layer. The pipeline is complemented by a two-stage hybrid misinformation detection module that fuses a supervised content-based classifier with an external verification signal derived from public-health knowledge bases. Applied to three benchmark datasets, COVID-19_FNIR, Monkeypox, and Constraint, the framework achieves strong classification performance, with F1 up to 0.950 and ROC-AUC up to 0.989, while identifying high-impact clusters under top 1%, 5%, and 10% VP thresholds. For datasets without native engagement metadata, prioritisation is based on a logistic propensity-to-spread score, used as an ordinal proxy for diffusion potential rather than a direct measure of engagement. The results show that integrating semantic structure, virality-aware ranking, and affective-linguistic profiling enables scalable and interpretable comparative analysis of misinformation across pandemics. The proposed framework supports monitoring-oriented early warning by surfacing low-volume but high-risk narratives for analyst review.
摘要:健康錯誤資訊在疫情期間迅速傳播,形成與公共衛生指導相競爭的有害敘事。大多數主題建模管道將參與度視為外部結果,限制了它們優先考慮語義一致且快速擴散主題的能力。我們介紹了BERTopic-VP,一種優先考慮傳播性的主題建模框架,將基於上下文嵌入的聚類(BERTopic)與事後傳播優先化(VP)層結合起來。該管道還配備了一個兩階段的混合錯誤資訊檢測模塊,該模塊將監督式內容分類器與來自公共衛生知識庫的外部驗證信號融合在一起。應用於三個基準數據集,COVID-19_FNIR、猴痘和Constraint,該框架實現了強大的分類性能,F1高達0.950,ROC-AUC高達0.989,同時在前1%、5%和10%的VP閾值下識別出高影響力的聚類。對於沒有原生參與度元數據的數據集,優先化基於邏輯傳播潛力分數,該分數用作擴散潛力的序數代理,而不是參與度的直接衡量。結果顯示,整合語義結構、考慮傳播性的排名和情感語言特徵分析,使得跨疫情的錯誤資訊進行可擴展且可解釋的比較分析成為可能。所提出的框架支持以監測為導向的早期預警,通過顯示低量但高風險的敘事供分析師審查。
THESIS-MoE: Trainable Hierarchical Extraction and SteerIng of Sycophancy in Mixture-of-Experts
2608.15687v1 by Kareem Hassani, Chaymaa Abbas, Lama Mawlawi, Mariette Awad
Sycophancy, the tendency of a language model to change its answer to match a user's stated belief, is a common alignment failure. Existing activation steering methods typically apply a single contrastive direction uniformly throughout the model, which is an unconditional intervention that alters activations even when no sycophantic behavior is present, trading knowledge retention for behavioral correction. In Mixture-of-Experts (MoE) models, prior work further suggests that behavior is encoded within expert computations rather than routing decisions alone, making precise behavioral steering particularly challenging. In this work, we introduce a shared contrastive signal, built from matched prompts with and without a stated belief, that identifies where sycophancy lives across the MoE hierarchy and drives interventions that act only where the behavior is present. We formulate localization as a causal search over a granularity ladder of MoE blocks, experts, attention blocks, and heads, and compare unconditional subtraction against two conditional alternatives: an analytic projection-based subtraction and a learned per-token gate that steers the model away from sycophancy while keeping its weights frozen. We evaluate on three MoE models measuring sycophancy alongside general knowledge and reasoning benchmarks. Our conditional interventions removed up to 90\% of the belief-induced sycophancy. Our results demonstrate that sycophancy resides in identifiable computational subcircuits and can be selectively steered while maintaining a favorable removal-retention trade-off.
摘要:拍馬屁是語言模型改變其答案以符合用戶所表達的信念的傾向,這是一種常見的對齊失敗。現有的激活引導方法通常在整個模型中均勻地應用單一的對比方向,這是一種無條件的干預,即使在沒有拍馬屁行為的情況下也會改變激活,從而以知識保留換取行為修正。在專家混合模型(MoE)中,先前的研究進一步表明,行為是編碼在專家計算中,而不僅僅是路由決策,這使得精確的行為引導特別具有挑戰性。在本研究中,我們引入了一個共享的對比信號,該信號由帶有和不帶有明確信念的匹配提示構建,能夠識別拍馬屁在MoE層級中的存在位置,並驅動僅在行為存在的地方進行干預。我們將定位公式化為對MoE區塊、專家、注意力區塊和頭部的粒度梯度進行因果搜索,並將無條件的減法與兩種條件替代方案進行比較:基於解析投影的減法和一個學習的每個標記門控,該門控在保持權重不變的情況下使模型遠離拍馬屁。我們在三個MoE模型上進行評估,測量拍馬屁以及一般知識和推理基準。我們的條件干預消除了高達90\%的信念引起的拍馬屁。我們的結果表明,拍馬屁存在於可識別的計算子電路中,並且可以在保持有利的去除-保留權衡的同時進行選擇性引導。
Agent Gym: A Framework for Continuous Evaluation and Evolution of LLM Agents Through Human-in-the-Loop Feedback
2608.15591v1 by Pouya Ghiasnezhad Omran, Michael Zimmermann, Duncan Cambridge, Ashmita Kapoor, Tanya Dixit
Large Language Model (LLM) agents deployed in production environments face a fundamental tension: the agent's behavior is frozen at deployment time, while the business rules and edge cases it must handle continue to evolve. Existing approaches address agent construction and one-time evaluation but provide no structured mechanism for continuous post-deployment behavioral correction without modifying the agent's source code. Most of the approaches offered in the market, require intense collection of logs and traces, and re-examining the agent design by the engineering team, a process which is heavy, long and negates the economical value of agentic transformation. We introduce Agent Gym, a modular, domain-agnostic framework that wraps any existing LLM-based agent in a continuous evaluation-and-evolution loop. The framework provides six composable capabilities --- Act, Evaluate, Investigate, Correct, Learn, and Observe --- organized across three architectural zones: a constitution layer that codifies domain knowledge in configuration artifacts, a runtime inference pipeline that chains acting, investigation, and adaptive correction, and a learning loop that enables subject matter experts to discover and validate new correction rules through natural language interaction. The key technical contributions include a hybrid deterministic-LLM correction engine with 21 condition operators and three-tier actions, a three-layer investigation architecture for ground-truth-free compliance validation, and a programmatic safety loop that guarantees rule correctness before human approval. We further introduce the Spec-to-Note Gap, an autoencoder-inspired view of agentic system transparency. An open-source reference implementation for invoice processing demonstrates that the framework is fully operational and ready for adoption.
摘要:大型語言模型(LLM)代理在生產環境中面臨著根本性的緊張關係:代理的行為在部署時被凍結,而必須處理的商業規則和邊緣案例則不斷演變。現有的方法解決了代理的構建和一次性評估,但未提供任何結構化機制以在不修改代理源代碼的情況下進行持續的部署後行為修正。市場上大多數提供的方法需要大量的日誌和追蹤數據收集,並由工程團隊重新檢查代理設計,這是一個繁重、漫長的過程,並削弱了代理轉型的經濟價值。我們介紹了Agent Gym,一個模組化的、與領域無關的框架,將任何現有的基於LLM的代理包裹在持續評估和演變的循環中。該框架提供六種可組合的能力——行動、評估、調查、修正、學習和觀察——這些能力組織在三個架構區域中:一個憲法層,將領域知識編碼為配置工件;一個運行時推理管道,鏈接行動、調查和自適應修正;以及一個學習循環,使主題專家能夠通過自然語言互動發現和驗證新的修正規則。關鍵的技術貢獻包括一個混合確定性-LLM修正引擎,具有21個條件運算符和三層行動;一個三層調查架構,用於無基準真相的合規驗證;以及一個程式化的安全循環,確保在人工批准之前規則的正確性。我們進一步介紹了Spec-to-Note Gap,一種受自編碼器啟發的代理系統透明度視角。一個開源的發票處理參考實現展示了該框架的完全運行狀態,並準備好被採用。
GraniKV: Asymmetric Granularity KV-Cache Paging for Multi-Agent Systems with Long Shared Prefix
2608.15584v1 by Jinhyun Jeon, Sungjoo Yoo
Production paged-serving engines apply uniform paging granularity to the KV cache, even though the two regions of a multi-agent workload have opposite storage requirements: a long shared prefix demands contiguity, while the per-request suffix demands fine-grained allocation. We present \textbf{GraniKV}, a KV-cache layer that allocates the shared prefix in a contiguous HOT pool and the suffix in a token-level COLD pool, combined with a per-step dispatcher which selects the appropriate backend among dual backends for each regime (compute-, memory-, or communication-bound). To the best of our knowledge, GraniKV is the first system to apply asymmetric paging granularity to the KV cache of a production paged-serving engine. At $L_p{=}16$\,K shared tokens GraniKV reaches $\mathbf{2.16\times}$, $\mathbf{1.98\times}$, and $\mathbf{1.57\times}$ output-token throughput over the production baseline on Llama-3.1-8B/TP=1, Qwen-2.5-14B/TP=2, and Qwen-2.5-32B/TP=4. The gain decomposes: cascade attention integration contributes the majority at saturation; the asymmetric storage layer adds $1.05$--$1.15\times$ end-to-end while being what makes the batched-GEMM prefix backend possible at all. Under heterogeneous multi-agent serving with \emph{distinct} prompts of different lengths, the attribution inverts: GraniKV sustains $\mathbf{1.95\times}$ while batch-global cascade collapses to parity --- the storage layer alone carries the win in the regime that motivates the paper.
摘要:生產頁面服務引擎對KV快取應用統一的分頁粒度,儘管多代理工作負載的兩個區域具有相反的存儲需求:長共享前綴要求連續性,而每個請求的後綴則要求細粒度分配。
我們提出了\textbf{GraniKV},這是一個KV快取層,將共享前綴分配在連續的HOT池中,後綴則分配在令牌級的COLD池中,並結合了一個每步調度器,該調度器在每個模式(計算、內存或通信限制)中選擇適當的後端。
據我們所知,GraniKV是第一個將非對稱分頁粒度應用於生產頁面服務引擎的KV快取系統。
在$L_p{=}16$\,K共享令牌下,GraniKV在Llama-3.1-8B/TP=1、Qwen-2.5-14B/TP=2和Qwen-2.5-32B/TP=4的生產基準上達到了$\mathbf{2.16\times}$、$\mathbf{1.98\times}$和$\mathbf{1.57\times}$的輸出令牌吞吐量。
增益分解如下:在飽和時,級聯注意整合貢獻了大部分;非對稱存儲層在端到端上增加了$1.05$--$1.15\times$,同時使得批次GEMM前綴後端成為可能。在異質多代理服務中,具有\emph{不同}長度的不同提示,歸因則反轉:GraniKV保持$\mathbf{1.95\times}$,而批次全局級聯則崩潰至平衡——存儲層單獨在促使本文的模式中獲得了勝利。
From Generalist to Specialist: A Context-Fusion Framework for Endoscopic Polyp Reporting with a Frozen VLM
2608.15580v1 by Ruijie Yang, Yan Zhu, Peiyao Fu, Siyuan Li, Te Luo, Zhihua Wang, Quanlin Li, Pinghong Zhou, Xian Yang, Shuo Wang
Reliable endoscopic polyp reporting requires integrating quantitative lesion sizing, standardized Paris classification, and clinically meaningful morphological description within a single record. General-purpose vision-language models (VLMs) offer a unified interface for image understanding and report generation. Existing specialization strategies, however, typically rely on task-specific models or model-weight adaptation, leaving unresolved how to introduce reliable specialist knowledge while preserving both this unified interface and the VLM's pretrained capabilities. We introduce a context-fusion framework that specializes a frozen general-purpose VLM through both implicit instruction context and explicit transduction context without modifying its pretrained weights. Specifically, a self-supervised polyp encoder retrieves related image-report pairs as explicit, query-specific evidence, while learned continuous specialist tokens provide implicit instruction context shared across cases. Experiments were conducted on 2,056 expert-annotated public endoscopic images. We compared the framework with general-purpose VLMs, task-specific predictors, and weight-adaptation methods to assess specialist performance, unified reporting, and adaptation efficiency. Across numerical, categorical, and report-generation metrics, the proposed framework substantially improved direct frozen-VLM inference and achieved the strongest overall performance among the evaluated methods. It added trainable parameters equal to only 0.006% of the frozen VLM's parameter count. When the top-1 retrieved case carried the correct target category, our framework corrected 70.5% of the errors made by a weight-adaptation baseline. These findings support the context-fusion framework as a lightweight and effective strategy for specialist adaptation of a frozen VLM.
摘要:可靠的內視鏡息肉報告需要將定量病變大小、標準化的巴黎分類和臨床上有意義的形態描述整合在單一記錄中。通用視覺-語言模型(VLMs)提供了一個統一的圖像理解和報告生成界面。然而,現有的專業化策略通常依賴於特定任務的模型或模型權重調整,尚未解決如何在保留這一統一界面和VLM的預訓練能力的同時引入可靠的專家知識。我們提出了一個上下文融合框架,通過隱式指令上下文和顯式轉導上下文專門化一個凍結的通用VLM,而不修改其預訓練權重。具體而言,自監督的息肉編碼器檢索相關的圖像-報告對作為顯式的查詢特定證據,而學習的連續專家標記提供了在案例之間共享的隱式指令上下文。實驗在2,056張專家標註的公共內視鏡圖像上進行。我們將該框架與通用VLMs、特定任務的預測器和權重調整方法進行比較,以評估專家性能、統一報告和適應效率。在數值、類別和報告生成指標上,所提出的框架顯著改善了直接凍結VLM推理,並在評估的方法中實現了最強的整體性能。它增加的可訓練參數僅佔凍結VLM參數總數的0.006%。當檢索到的頂級案例攜帶正確的目標類別時,我們的框架修正了70.5%的權重調整基線所犯的錯誤。這些發現支持上下文融合框架作為一種輕量且有效的策略,用於凍結VLM的專家適應。
Admission Without Answers: Label-Free Certification and Experience Learning for LLM-Based Optimization Modeling
2608.15565v2 by Junbo Jacob Lian, Huiling Chen, Hanzhang Qin, Chung-Piaw Teo
Experience-learning agents for optimization modeling improve by storing verified skills, but existing learners admit knowledge by checking against known answers, which real ticket streams do not provide. The natural label-free alternatives are unreliable: on a 300-problem label-blind stream, admitting every executable model poisons roughly one admission in four, while single-instance agreement accepts models that match at one value but differ elsewhere. We propose AdmitOR, an admission gate built on calibrated external behavioral evidence. Candidates from three model families, prompting strategies, and solver stacks are run on instances resampled from an extracted parameter domain; agreement across the resulting value-function traces is summarized by a cross-family clique, and a calibrated threshold returns accept, abstain, or escalate. The preregistered false-discovery criterion holds on calibration data but not on the wild stream. We report this negative result in full and trace most failures to benchmark texts that do not faithfully encode their labeled instances. Comparing four admission judges on one collection of logs inside a state-of-the-art skill learner, AdmitOR raises admission precision to 0.927, against 0.871 for majority vote and 0.726 for execution success, yielding 3.1x and 8.0x fewer poisoned admissions. Its library is the smallest and attains the highest macro accuracy across five public benchmarks, 58.4 against 54.8 for majority vote and 53.9 for the ground-truth-labeled library. The 3.5-point gain over majority vote is supported by a paired bootstrap and survives correction for a host-side anomaly. To our knowledge, AdmitOR is the first label-free admission mechanism designed around an explicitly calibrated false-discovery target. The transfer failure identifies a necessary condition for extending it to wild streams.
摘要:經驗學習代理在優化建模中通過儲存經過驗證的技能來提高性能,但現有的學習者通過檢查已知答案來承認知識,而這些答案在實際的票務流中並不存在。自然的無標籤替代方案不可靠:在一個300題的無標籤流中,承認每個可執行模型大約會使每四個承認中就有一個受到污染,而單實例一致性則接受在一個值上匹配但在其他地方不同的模型。我們提出了AdmitOR,一個基於經過校準的外部行為證據的承認閘。來自三個模型家族、提示策略和求解器堆棧的候選者在從提取的參數域重新抽樣的實例上運行;對於生成的值函數軌跡的一致性通過跨家族的團體進行總結,並且一個經過校準的閾值返回接受、放棄或升級。預註冊的假發現標準在校準數據上成立,但在野外流中不成立。我們全面報告這一負面結果,並將大多數失敗追溯到未忠實編碼其標記實例的基準文本。在一個最先進的技能學習者中的一組日誌上比較四個承認評審,AdmitOR將承認精度提高到0.927,而多數投票為0.871,執行成功為0.726,分別減少了3.1倍和8.0倍的污染承認。它的庫是最小的,並在五個公共基準中達到了最高的宏觀準確率,58.4對比多數投票的54.8和真實標記庫的53.9。相較於多數投票的3.5點增益得到了配對自助法的支持,並且在主機端異常的修正下仍然成立。據我們所知,AdmitOR是第一個圍繞明確校準的假發現目標設計的無標籤承認機制。轉移失敗確定了將其擴展到野外流的必要條件。
BengaliMCQ: Automatic Generation and Answer Prediction of Academic Multiple-Choice Questions in a Low-Resource Language
2608.15547v1 by Abu Tarabin Surzo, A. K. M. Nihalul Kabir, Sm Azmain Faysal, Ariana Haque Ami, Lawrence Amlan Gomes, Farig Sadeque
Traditional retrieval-augmented generation (RAG) frameworks process documents without attending to their hierarchical structure, leading to poor performance, especially in low-resource languages such as Bengali. To address this, we propose a structure-aware RAG framework that models Bengali textbooks as hierarchical graphs and uses a contrastively trained graph neural network to retrieve a small set of relevant passages. These passages provide focused context for a large language model, enabling topic-specific multiple-choice question (MCQ) generation and in-domain answer prediction. Experimental results demonstrate that our framework outperforms strong dense retrieval baselines across retrieval metrics, produces more relevant MCQs, and achieves superior answer prediction accuracy.
摘要:傳統的檢索增強生成(RAG)框架在處理文件時未考慮其層次結構,導致性能不佳,特別是在資源匱乏的語言如孟加拉語中。為了解決這個問題,我們提出了一種結構感知的 RAG 框架,將孟加拉語教科書建模為層次圖,並使用對比訓練的圖神經網絡來檢索一小組相關段落。這些段落為大型語言模型提供了集中上下文,使得能夠生成主題特定的多選題(MCQ)和在域內的答案預測。實驗結果顯示,我們的框架在檢索指標上超越了強大的密集檢索基準,產生了更相關的 MCQ,並實現了更高的答案預測準確性。
L3Cube-IndicQuest v2: A Large-Scale Multilingual Benchmark for Evaluating Factual Knowledge of Large Language Models Across Indic Languages
2608.15535v1 by Rinit Jain, Tirthraj Mahajan, Advait Joshi, Raviraj Joshi
We present L3Cube-IndicQuest v2, a large-scale gold-standard multilingual question-answering benchmark for evaluating the India-specific factual knowledge of Large Language Models (LLMs). The benchmark comprises 3,471 curriculum-grounded English question--answer pairs spanning nine domains, curated from educational curricula, competitive examination materials, and domain-specific reference books. We introduce a practical hybrid construction strategy that combines context-grounded LLM-based question generation and validation with semantic deduplication and human verification, enabling scalable creation of benchmark data while preserving annotation quality. The benchmark is translated into 19 Indic languages, yielding a publicly released multilingual dataset of 69,420 question--answer pairs across 20 languages. We evaluate six LLMs under three protocols: LLM-as-a-judge and two deterministic lexical criteria, exact-substring and word-overlap matching. All three produce almost the same model ranking, showing that the results do not depend on the choice of judge. The frontier commercial model leads by a wide margin, and among open-weight models Gemma4 31B outperforms the Indic-specialised Sarvam 30B in every evaluated Indic language.
摘要:我們推出 L3Cube-IndicQuest v2,這是一個大型的金標準多語言問答基準,用於評估大型語言模型(LLMs)在印度特定事實知識方面的表現。該基準包含 3,471 個基於課程的英語問答對,涵蓋九個領域,這些內容來自教育課程、競爭性考試材料和特定領域的參考書籍。我們引入了一種實用的混合建構策略,結合了基於上下文的 LLM 問題生成和驗證,以及語義去重和人工驗證,使得基準數據的可擴展創建成為可能,同時保持標註質量。該基準已翻譯成 19 種印度語言,產生了一個公開釋出的多語言數據集,包含 69,420 個問答對,涵蓋 20 種語言。我們在三個協議下評估了六個 LLM:LLM 作為評審以及兩個確定性詞彙標準,精確子字符串和詞重疊匹配。所有三種方法產生的模型排名幾乎相同,顯示結果不依賴於評審的選擇。最前沿的商業模型以較大優勢領先,而在開放權重模型中,Gemma4 31B 在每種評估的印度語言中均優於專注於印度的 Sarvam 30B。
Mental Model Management: An Operator-Based Framework for LLM Memory
2608.15451v1 by Oliver Kramer
Large language models process large amounts of information but usually lack an explicit mechanism for maintaining compact and evolving conceptual representations. We introduce Mental Model Management (3M), a framework in which knowledge is represented as mental models consisting of compact chunks. Rather than accumulating text passages, 3M continuously integrates new information into an existing conceptual representation. A set of operators extracts knowledge, retrieves relevant models, adds and updates chunks, reorganizes representations, detects inconsistencies, and derives new knowledge. We describe the main 3M operators and illustrate each operation using Evolution Strategies as a running example.
摘要:大型語言模型處理大量資訊,但通常缺乏明確的機制來維持緊湊且不斷演變的概念表徵。
我們介紹了心理模型管理(3M),這是一個將知識表示為由緊湊區塊組成的心理模型的框架。
3M並不是累積文本段落,而是持續將新資訊整合到現有的概念表徵中。
一組運算子提取知識、檢索相關模型、添加和更新區塊、重組表徵、檢測不一致性並推導新知識。
我們描述了主要的3M運算子,並以進化策略作為持續示例來說明每個操作。
Implementation of a Metacognition Framework for Self-Awareness and Self-Regulation in Ensembles of LLMs
2608.15400v1 by Charles Courchaine, Ricky J. Sethi, Hefei Qiu
Large Language Models (LLMs) are notorious for struggling with assessing their own uncertainty, detecting knowledge conflicts, or recognizing when problems exceed their expertise; such limitations inevitably undermine reliability and trust in LLMs. In this paper, we present the first implementation of a metacognitive framework for ensembles of LLMs that addresses these challenges through explicit monitoring and control mechanisms. Our system computes a Metacognitive State Vector (MSV) quantifying self-awareness for monitoring across five dimensions derived from cognitive psychology: Emotional Response, Correctness Evaluation, Experiential Match, Conflicting Information, and Problem Importance. MSV values also provide self-regulation for control, automatically switching between System 1 (fast, single- or multi-node) and System 2 (deliberative, multi-node) processing based on query complexity. For System 2 execution, graph-theoretic algorithms control the assignment of specialized roles (Domain Expert, Critic, Evaluator, Synthesizer, and Generalist) to ensemble nodes according to their MSV-quantified metacognitive states. Our implementation allows users to explore how different query types trigger distinct processing modes. The Proof-of-Concept (PoC) demo showcases the framework with illustrative examples showing appropriate System 1/System 2 routing and helps visualize the metacognitive process via real-time radar charts and decision indicators. This PoC implementation demonstrates the feasibility of creating a framework for metacognitive self-awareness and self-regulation in LLM systems.
摘要:大型語言模型(LLMs)因難以評估自身的不確定性、檢測知識衝突或識別問題超出其專業範疇而聞名;這些限制不可避免地削弱了對LLMs的可靠性和信任。在本文中,我們展示了首個針對LLMs集成體的元認知框架實現,通過明確的監控和控制機制來解決這些挑戰。
我們的系統計算一個元認知狀態向量(MSV),量化自我意識,以便在五個來自認知心理學的維度上進行監控:情感反應、正確性評估、經驗匹配、衝突信息和問題重要性。MSV值還提供自我調節以進行控制,根據查詢的複雜性自動在系統1(快速、單節點或多節點)和系統2(深思熟慮、多節點)處理之間切換。
在系統2執行中,圖論算法根據其MSV量化的元認知狀態控制專業角色(領域專家、批評者、評估者、綜合者和通才)在集成節點上的分配。
我們的實現允許用戶探索不同查詢類型如何觸發不同的處理模式。概念驗證(PoC)演示展示了該框架,並通過示例顯示適當的系統1/系統2路由,幫助通過實時雷達圖和決策指標可視化元認知過程。這個PoC實現展示了在LLM系統中創建元認知自我意識和自我調節框架的可行性。
Grounding Healthcare LLMs in a Causal Knowledge Graph: Framework, Metrics, and a Cardiovascular Pilot
2608.15382v1 by Ummara Mumtaz, Aimen Noor, Awais Ahmed
Large language models (LLMs) are increasingly proposed for healthcare decision support, but their evaluations still reward single-answer accuracy rather than reasoning about interventions, mechanisms, harms, evidence, and uncertainty. We propose a reproducible, graph-centered evaluation framework for intervention-oriented LLM behavior in healthcare and stress-test it in a cardiovascular pilot. The framework has four components: (i) a domain causal knowledge graph in which assertions are first-class, provenance-preserving nodes with stable identifiers; (ii) a scenario-conditioned subgraph extraction step that, given any clinical scenario, retrieves the relevant reified-assertion subgraph; (iii) four controlled grounding conditions that vary how the retrieved subgraph is composed into the model's context (ungrounded C1, knowledge-graph C2, causal-graph C3, integrated C4); and (iv) an automated scoring pipeline, anchored on assertion identifiers, that computes intervention accuracy, and other evaluation measures on a single pass. To test the framework, we built a category-balanced scenario generator across eight reasoning failure modes and instantiated it on a cardiovascular graph. The metric panel discriminates conditions along interpretable, non-redundant axes: C4 obtains the strongest causal edge F1 (0.838), adverse-effect F1 (0.833), evidence accuracy (0.738), and unsupported claim rate (0.114), while C1 obtains the highest raw intervention accuracy (0.948) with no measurable causal or evidential grounding.
摘要:大型語言模型(LLMs)越來越多地被提議用於醫療決策支持,但其評估仍然獎勵單一答案的準確性,而不是對干預、機制、危害、證據和不確定性進行推理。我們提出了一個可重複的、以圖為中心的評估框架,用於醫療保健中的干預導向LLM行為,並在心血管試點中進行壓力測試。該框架有四個組成部分:(i)一個領域因果知識圖,其中斷言是第一類的、保持來源的節點,具有穩定的標識符;(ii)一個情境條件的子圖提取步驟,根據任何臨床情境檢索相關的具體化斷言子圖;(iii)四個控制的基礎條件,變化檢索到的子圖如何組成模型的上下文(未基礎的C1、知識圖C2、因果圖C3、整合的C4);以及(iv)一個自動評分管道,以斷言標識符為基礎,計算干預準確性和其他評估指標,僅需一次通過。為了測試該框架,我們建立了一個跨越八種推理失敗模式的類別平衡情境生成器,並在心血管圖上實現了它。該指標面板沿著可解釋的、非冗餘的軸區分條件:C4獲得最強的因果邊緣F1(0.838)、不良影響F1(0.833)、證據準確性(0.738)和不支持的主張率(0.114),而C1獲得最高的原始干預準確性(0.948),卻沒有可測量的因果或證據基礎。