Skip to content

LLM

LLM

Publish Date Title Authors Homepage Code
2026-08-18 From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation Xingjian Wang et.al. 2608.18076v1 null
2026-08-18 Multi-Agent AI System for Radiology Report Structuring and Quality Assurance with Independent Radiologist Evaluation Iryna Hartsock et.al. 2608.18072v1 null
2026-08-18 TokEval: A Tokenizer Evaluation Suite Clara Meister et.al. 2608.18062v1 null
2026-08-18 Delegation Asymmetry in Agentic Recommender Systems: Measuring Two-Sided Receptivity in Online Dating Daria Leshchikova et.al. 2608.18058v1 null
2026-08-18 StagedWorkspace: A Versioned Workspace for Knowledge-Work Agents Yining Hua et.al. 2608.18050v1 null
2026-08-18 Language Has Two Parameters: Narrative-Induced Semantic Plasticity and Phase-Sensitive Interpretation Hollis Robbins et.al. 2608.18041v1 null
2026-08-18 Chain-of-Experience for Continual LLM Improvement Haoqin Tu et.al. 2608.18027v1 null
2026-08-18 Why GPT-Style Models Do Not Directly Transfer to Symbolic Music: Compression in the Wrong Coordinate System Yi Wang et.al. 2608.18025v1 null
2026-08-18 Can Large Language Models Explain Flight Safety Events? A Prior-Guided Semantic LLM-based Approach Lu Xu et.al. 2608.18017v1 null
2026-08-18 The IOL-AI Challenge: An Open Challenge towards Advancing Linguistic Reasoning Eduardo Sánchez et.al. 2608.18011v1 null
2026-08-18 Policy-Invariant Reward Shaping from LLM Feedback: A Framework for Hybrid RL Agents Christophe D. Hounwanou et.al. 2608.18008v1 null
2026-08-18 Traceable Trust for action-ready artificial intelligence in bioscience Huayu Xin et.al. 2608.17997v1 null
2026-08-18 Judge, Retrieve, or Abstain: Uncertainty-Guarded LLM Judging with Provable Risk Guarantees Sher Badshah et.al. 2608.17994v1 null
2026-08-18 Against Political Polarization: A Unified Framework for Tracing Evolving Political Ideologies on Social Media Yijie Xu et.al. 2608.17987v1 null
2026-08-18 Dual Co-Train: Cross-Dataset Ultrasound Tongue Segmentation Under Extreme Data Scarcity Alisher Myrgyyassov et.al. 2608.17983v1 null
2026-08-18 When Writing Style Drifts: Benchmarking Authorship Verification under Distribution Shifts in Genre, Time and the AI-Era Lotta Kiefer et.al. 2608.17979v1 null
2026-08-18 Too Sure to Be Safe: Model Calibration for Reliable Log Anomaly Detection Bin Li et.al. 2608.17965v1 null
2026-08-18 Towards Zero-Shot Task Transfer with Neurosymbolic World Models Isidoro Tamassia et.al. 2608.17959v1 null
2026-08-18 An Omitted Mode Is a Rare Rule: The Sampling-Verification Danger Law in Continuous Code World Models Javier Aguilar Martín et.al. 2608.17956v1 null
2026-08-18 Do Large Language Models Play Six Degrees of Separation? Measuring Topological Compression in Long-Context Manifolds Md. Faiyaz Abdullah Sayeedi et.al. 2608.17950v1 null
2026-08-18 SIGMA: SHAP-Guided Implicit-Trajectory Generation for Metadata-Free LLM-Based AutoFE Xuan Zheng et.al. 2608.17948v1 null
2026-08-18 Procedural Content Metageneration via Program Search and Continual Abstraction Discovery Matthew Siper et.al. 2608.17947v1 null
2026-08-18 Efficient RLVR Scheduling via Graph-Structured Online Difficulty Estimation Zhizhao Liu et.al. 2608.17941v1 null
2026-08-18 Grading Needs a Rubric, Not Intelligence Jhen-Ke Lin et.al. 2608.17938v1 null
2026-08-18 EvoTS-Agent: A Self-Evolving LLM Agent for Financial Time Series Change Point Detection Lei Jiang et.al. 2608.17933v1 null
2026-08-18 Collective Counterfactual Planning: Coordination, Consent, and Verification under Representational Constraints Chainarong Amornbunchornvej et.al. 2608.17932v1 null
2026-08-18 SpeechSense: A Paralinguistic-Focused Dataset for Fine-Grained Speech Sentiment Analysis Shicheng Ma et.al. 2608.17931v1 null
2026-08-18 Comparative Study of Out-of-the-Box Technology for Automatic Target Detection and Recognition Alma M. Liezenga et.al. 2608.17917v1 null
2026-08-18 CABLE: Extending the Reach of Memory Retrieval via Complementary Antecedent-Based Linking and Expansion Zheling Tan et.al. 2608.17911v1 null
2026-08-18 AutoResearch: Insight In, Hallucination Out Yiming Ren et.al. 2608.17906v1 null
2026-08-18 BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models Liubov Chubarova et.al. 2608.17895v1 null
2026-08-18 BayesPrompt: human readable prompts that make sense Franky Kevin Nando Tezoh et.al. 2608.17866v1 null
2026-08-18 ARASH: Adaptive Retrieval And Shot Selection for Tabular Prediction Samirasadat Jamalidinan et.al. 2608.17856v1 null
2026-08-18 Encoded but Not Actionable: Auditing the Decode-Generate-Steer Gap in Frozen LLMs for Geometric Constraints Man Liang et.al. 2608.17843v1 null
2026-08-18 AdaLens: Interactive Storyline for Monitoring and Steering Long-Running Agentic Data Analysis Yangtian Liu et.al. 2608.17834v1 null
2026-08-18 The Model's Tell: Measuring Context-Leakage Attack Signals with Behavior Gauges Maosen Zhang et.al. 2608.17829v1 null
2026-08-18 From Global Benchmarks to Local Evaluations: Benchmarking LLMs for the German Public Sector Camilla Dalerci et.al. 2608.17827v1 null
2026-08-18 MotoSafety: Edge-AI with Learned Temporal Importance for Two-Wheeler Collision Risk Assessment Under Time Pressure Sumit S. Shevtekar et.al. 2608.17823v1 null
2026-08-18 Interpretable Humans, Alien LLMs: Expert Analysis of Latent Structures in Assessment Responses Alona Strugatski et.al. 2608.17810v1 null
2026-08-18 Whether LLMs Can Navigate Beliefs and Facts Depends on How You Phrase It Quang Minh Nguyen et.al. 2608.17809v1 null
2026-08-18 An Empirical Study of Reward Specification and Benchmark Reliability in GRPO-based LLM Unlearning Rubén Balbastre et.al. 2608.17804v1 null
2026-08-18 StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows Liya Zhu et.al. 2608.17800v1 null
2026-08-18 Training with synthetic data for drone detection in thermal imagery Tanel Liiv et.al. 2608.17799v1 null
2026-08-18 TraceSQL: Traceable Answerability Estimation for Reference-Free Text-to-SQL Verification Neelesh Kumar Shukla et.al. 2608.17795v1 null
2026-08-18 Preference Is Not Intervention: The Structure and Stability Boundaries of Reader-Specific Evidence Utility Shi Zhou et.al. 2608.17781v1 null
2026-08-18 Learnware for CSI Feedback: Scene-specific Small Models Can Do Big Xiangyi Li et.al. 2608.17760v1 null
2026-08-18 D$^2$ACCI: A Dual-Loop Diagnostic Protocol for Evidence-Preserving Agent Memory Xule Liu et.al. 2608.17756v1 null
2026-08-18 The Curious Case of Exploding DecPOMDPs: Containing the Fire through Policy Counting Nazlı Nur Karabulut et.al. 2608.17749v1 null
2026-08-18 Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See Ayoub Kirouane et.al. 2608.17744v1 null
2026-08-18 Neuro-symbolic learning over OWL 2 DL via consequence-based compilation to differentiable circuits Olga Mashkova et.al. 2608.17741v1 null
2026-08-18 What Aggregate Scores Miss: Measuring Item-Level Regressions in Commercial LLM API Migrations Xiaonan Xu et.al. 2608.17719v1 null
2026-08-18 Beyond Suspicious Steps: Ontological Trust in Long-Horizon Agents An He et.al. 2608.17718v1 null
2026-08-18 Communicating Credit Risk with Large Language Models: Evaluation of Explanations from Standard and Alternative Data-Based Models Sahab Zandi et.al. 2608.17715v1 null
2026-08-18 Accuracy and Robustness of Model Cascades Under Data Perturbations Pallavi Mitra et.al. 2608.17711v1 null
2026-08-18 GADR: Gathering Architecture Decision Records from Meeting Transcriptions Lucas Daniel Costa da Silva et.al. 2608.17694v1 null
2026-08-18 Mixture-of-Expert Blocks Contain Strong Hallucination Detection Signals Joao Fonseca et.al. 2608.17687v1 null
2026-08-18 Benchmarking Automated Security Patch Backporting: How Far Are We? Jincheng Yang et.al. 2608.17671v1 null
2026-08-18 GraphWake: Group Polarization via Memory-Mediated Polarization Cascade in LLM-Agent Communities Haoran Bu et.al. 2608.17665v1 null
2026-08-18 MobileWorldSafety: Benchmarking GUI Agent Safety Against Environmental Injection Attacks in Android Apps Sujin Chen et.al. 2608.17659v1 null
2026-08-18 LLM-Derived Preference Judgments Are Not Self-Consistent Matthew T. Ford et.al. 2608.17644v1 null
2026-08-18 Beyond the Trace: Coupling an Interpretable Reasoning-State Readout to Native MoE Routing Kang Chen et.al. 2608.17638v1 null
2026-08-18 Graph Surgery and the Do-Operator: A Precise Correspondence for Acyclic Structural Causal Models Satpreet Makhija et.al. 2608.17634v1 null
2026-08-18 DEPT: Document Embedding Preservation Tuning for Unified Query Expansion and Retrieval Jingyuan Wang et.al. 2608.17632v1 null
2026-08-18 From Student Risk Prediction to SC2R: Semantics-Constrained Counterfactual Recourse for Educational Decision Support Ngoc Luyen Le et.al. 2608.17618v1 null
2026-08-18 Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges Syeda Faiza Ahmed et.al. 2608.17605v1 null
2026-08-18 HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety Yajing Bai et.al. 2608.17597v1 null
2026-08-18 tinyDSM: A Framework for Skill Modeling and Development for Resource-Constrained Millirobots Markus D. Kobelrausch et.al. 2608.17596v1 null
2026-08-18 TRUSS: Towards Task-Reliable and User-Safe Automated Agent Skill Generation Zhibo Zhang et.al. 2608.17588v1 null
2026-08-18 Write, Execute, Refine: From Skill Followers to Skill Optimizers via Reinforcement Learning from Execution Feedback Kang Peng et.al. 2608.17587v1 null
2026-08-18 Auditing Exposure to Harmful Content on TikTok using Multimodal Language Models: A Cross-National, Age-Stratified Study Hamidreza Saffari et.al. 2608.17583v1 null
2026-08-18 Quantifying Risk Under Evolving Uncertainty: Belief-Dependent Robustness for Safe Sequential Decision Making Deep Kumar Ganguly et.al. 2608.17574v1 null
2026-08-18 DMT-Dens: Density-preserving manifold visualization for biological data Ruizhe Wang et.al. 2608.17571v1 null
2026-08-18 Domain-Adapted Molecular Language Models for Efficient Search of Make-on-Demand Libraries Henrik Wille et.al. 2608.17567v1 null
2026-08-18 Where a New Concept Must Enter: Entry Point Gates Cross-Task Usability in Unified Multimodal Models Zongyang Qiu et.al. 2608.17564v1 null
2026-08-18 Reflex-Guard: A Low-Latency Guardrail for LLM Prompt Safety Using Dense Semantic Embeddings Istiaque Ahmed et.al. 2608.17556v1 null
2026-08-18 Code as Representation: A Compilable Parsing Paradigm for Academic Documents Rihui Jin et.al. 2608.17550v1 null
2026-08-18 No Gaussian Required: Contrastive Inverse Dynamics for JEPA World Models Jack Boylan et.al. 2608.17542v1 null
2026-08-18 CoAL-RAG: A Complexity-Aware Legal Retrieval-Augmented Generation Method Jin Su et.al. 2608.17536v1 null
2026-08-18 ArborMem: Navigating Interaction States with Memory Forests Zongwei Lv et.al. 2608.17534v1 null
2026-08-18 When to Review: Spaced Repetition for Continual Pre-Training of Language Models Alankar Atreya et.al. 2608.17530v1 null
2026-08-18 Agent Lightning v1.0: Towards Harnessed Agentic RL Zhiyuan He et.al. 2608.17528v1 null
2026-08-18 Explainable AI-Powered Framework for Video-Based Skill Assessment in Cataract Surgery Mohammad Javad Ahmadi et.al. 2608.17522v1 null
2026-08-18 Effects of Answer Format Variation on Gender Bias in Large Language Models Ksenia Merzlyakova et.al. 2608.17516v1 null
2026-08-18 Beyond FLOPs: Energy-Aware Knowledge Distillation for Sustainable LLMs on Code-Related Task Enrique Barba Roque et.al. 2608.17515v1 null
2026-08-18 SGHA: Evidence-Grounded Research Problem Discovery with Local Language Models Sarvesh Gharat et.al. 2608.17501v1 null
2026-08-18 When AI Designs AI: Innovation or Imitation? Yikang Yang et.al. 2608.17471v1 null
2026-08-18 SAGE: Self-Evolving Storyboard Skills via Attribution-Guided Rule Evolution Maolin Ran et.al. 2608.17468v1 null
2026-08-18 From Entity Mentions to Tone: An LLM-Based Pipeline for Media Bias Analysis Klesti Hoxha et.al. 2608.17454v1 null
2026-08-18 Decomposition Attacks Across Unlinkable Identities: Limits of Stateful Defenses for LLM Services Bowen Sun et.al. 2608.17445v1 null
2026-08-18 Structure-Internalized Rule Language Model for Faithful Knowledge Graph Reasoning Xingrui Zhuo et.al. 2608.17443v1 null
2026-08-18 Task-Aware Harness Provisioning for LLM Agents in Mission-Critical Infrastructure Operations Liangtao Lin et.al. 2608.17433v1 null
2026-08-18 SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation Keyu Tu et.al. 2608.17426v1 null
2026-08-18 An Investigation of Translationese in the Generations of Multilingual Large Language Models Maria Valentini et.al. 2608.17399v1 null
2026-08-18 LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents Yiming Du et.al. 2608.17393v1 null
2026-08-18 Leveraging generative hallucination and biophysics-informed modeling for unified biomolecular sequence-structure co-design Xuefeng Liu et.al. 2608.17381v1 null
2026-08-18 PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX Genghan Zhang et.al. 2608.17379v1 null
2026-08-18 Fair ASR: Re-Evaluating Black-Box Jailbreaks under Shared Target-Call Budgets Zhida He et.al. 2608.17360v1 null
2026-08-18 Cognitive Graph Intelligence for Adaptive and Robust DDoS Attack Detection in Next Generation Networks Mohammad Arif Hossain et.al. 2608.17352v1 null
2026-08-18 MoFE: A Novel Mixture-of-Experts Framework with Fourier Neural Operators for Cryptocurrency Forecasting Bowen Liu et.al. 2608.17342v1 null
2026-08-18 LLM-Only PDDL Domain Repair with Open-Weight Models Nader Karimi Bavandpour et.al. 2608.17341v1 null

Abstracts

From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation

2608.18076v1 by Xingjian Wang, Zhao Wang, Taihang Hu, Jun Zheng, Qing Jin, Qinye Zhou, Zhengtao Wu, Yongchao Du, Zuan Gao, Chao Lin, Yefeng Shen, Xiaoli Xu, Zhengze Xu, Hao Yan, Yuhang Yu, Mingzhou Zhang, Mengting Chen

Large-scale image generation has benefited from advances in data scale, quality, rebalancing, and recaptioning, yet conventional pipelines typically optimize task-specific datasets in isolation. A central challenge is not only how to curate each task-specific corpus, but also how to organize heterogeneous supervision according to the dependencies among generative capabilities. We present a \textbf{capability-driven data infrastructure} that couples capability-specific supervision construction with capability-aligned curriculum scheduling. Its three specialized yet interoperable data engines build complementary relational supervision for text-image grounding, inter-image transformation, and image-knowledge association, while caption experts align T2I and editing supervision across tasks and granularities. A multi-stage curriculum jointly evolves task composition, visual-concept distribution, data quality, and image resolution along the dependency order of capability acquisition, with capability-aware evaluation closing the loop through targeted retrieval, expert construction, and gap-aware resampling. At scale, the framework curates a 440M-image T2I corpus, 120M editing pairs, and over 27M image-entity pairs. With this infrastructure, we train multimodal diffusion models at two scales from scratch, with 3B and 6B sizes respectively. We conduct quantitative evaluation on CPI-Bench, along with qualitative evaluations across diverse text-to-image and editing scenarios. Experimental results present broad visual coverage, versatile rendering, and effective transfer across generative capabilities.

摘要:大規模圖像生成受益於數據規模、質量、重新平衡和重新標題的進步,但傳統流程通常在孤立的情況下優化特定任務的數據集。一個主要挑戰不僅在於如何策劃每個特定任務的語料庫,還在於如何根據生成能力之間的依賴關係組織異質監督。我們提出了一個\textbf{以能力為驅動的數據基礎設施},將特定能力的監督構建與能力對齊的課程安排結合起來。它的三個專門但可互操作的數據引擎為文本-圖像基礎、圖像間轉換和圖像-知識關聯構建互補的關係監督,同時標題專家在任務和粒度之間對齊T2I和編輯監督。一個多階段課程共同演變任務組合、視覺概念分佈、數據質量和圖像解析度,沿著能力獲取的依賴順序進行,而以能力為中心的評估通過針對性檢索、專家構建和關注差距的重採樣來閉合循環。在規模上,該框架策劃了一個包含4.4億圖像的T2I語料庫、1.2億編輯對和超過2700萬圖像-實體對。利用這一基礎設施,我們從零開始訓練了兩個規模的多模態擴散模型,分別為30億和60億大小。我們在CPI-Bench上進行了定量評估,並在多樣的文本到圖像和編輯場景中進行了定性評估。實驗結果顯示出廣泛的視覺覆蓋、多樣的渲染和在生成能力之間的有效轉移。

Multi-Agent AI System for Radiology Report Structuring and Quality Assurance with Independent Radiologist Evaluation

2608.18072v1 by Iryna Hartsock, Cesar Lam, Christopher Otteni, Aliya Qayyum, Robert Gatenby, Cyrillo Araujo, Ghulam Rasool

Purpose: To develop and evaluate a locally deployed multi-agent AI system for radiology report structuring and quality assurance. Materials and Methods: This retrospective study included 638 radiology reports from CT examinations of the chest, abdomen, and pelvis dictated by 15 board-certified radiologists in 2023 and 2024. A multi-agent AI pipeline was developed to perform report structuring and quality assurance (QA). The system structured the report into standardized anatomical sections at the sentence level using regex rules and local large language models. It also detected mismatches between the Findings and Impression sections, or within sections; gender-anatomy conflicts; and undocumented communication of critical findings. Two board-certified radiologists independently evaluated a 45-report subset. Results: The multi-agent system structured the Findings sections of all reports (22,270 sentences) into a predefined anatomical format while retaining the original report content. The system flagged 90 (14.1%) reports, most commonly for section mismatches (80 reports, 12.5%). In the radiologist evaluation, both reviewers agreed that 31 (69%) were correctly restructured, 2 reports (4%) were incorrectly restructured, and disagreed on the remaining 12 reports (27%). Both reviewers agreed that no clinically important information was omitted and no fabricated content was introduced. Overall QA performance was rated as "excellent" or "good" in 84% of the evaluated reports, with the remaining reports rated as "fair". Conclusion: A locally deployed multi-agent AI system combined radiology report structuring and quality assurance within a single workflow. The system demonstrated favorable performance in radiologist evaluation. Such systems may support standardization of reporting and quality assurance in radiology practice.

摘要:目的:開發和評估一個本地部署的多代理人工智慧系統,用於放射學報告的結構化和質量保證。材料和方法:本回顧性研究包含了2023年和2024年由15位董事會認證的放射科醫師口述的638份胸部、腹部和骨盆的CT檢查報告。開發了一個多代理人工智慧管道來執行報告結構化和質量保證(QA)。該系統使用正則表達式規則和本地大型語言模型將報告結構化為標準化的解剖學部分,並在句子層面進行處理。它還檢測到發現和印象部分之間的錯配,或部分內部的錯配;性別-解剖學衝突;以及對關鍵發現的未記錄溝通。兩位董事會認證的放射科醫師獨立評估了45份報告的子集。結果:該多代理系統將所有報告的發現部分(22,270句)結構化為預定的解剖格式,同時保留原始報告內容。系統標記了90份(14.1%)報告,最常見的原因是部分錯配(80份報告,12.5%)。在放射科醫師的評估中,兩位評審一致認為31份(69%)報告重構正確,2份報告(4%)重構不正確,對剩餘的12份報告(27%)意見不合。兩位評審一致認為沒有遺漏臨床重要信息,也沒有引入虛構內容。整體質量保證表現被評為“優秀”或“良好”的報告佔84%,其餘報告被評為“公平”。結論:一個本地部署的多代理人工智慧系統在單一工作流程中結合了放射學報告的結構化和質量保證。該系統在放射科醫師評估中顯示出良好的表現。這類系統可能支持放射學實踐中的報告標準化和質量保證。

TokEval: A Tokenizer Evaluation Suite

2608.18062v1 by Clara Meister

Language model tokenizers are typically selected with minimal evaluation, despite the fact that their design choices directly impact model capabilities. This can be partly attributed to a limited understanding of which tokenizer properties affect which aspects of downstream performance. We introduce TokEval, a framework of tokenizer evaluation metrics that goes beyond standard measures like fertility and compression rate to capture linguistically and structurally meaningful properties, e.g., UTF-8 character boundary integrity and digit place-value boundary alignment for mathematics. To validate whether these metrics are predictive of downstream model performance, we conduct controlled language model pretraining experiments, varying solely the tokenizers' training data mixture, pretokenization strategy, and training algorithm. We evaluate the resulting models on bits-per-byte (a tokenizer-agnostic version of perplexity) and several benchmarks, spanning linguistic understanding, mathematical reasoning, and code generation. Our experiments suggest that different intrinsic properties have different impacts on model abilities: information-theoretic metrics predict language modeling abilities (Spearman rho up to 0.80), while structure-sensitive metrics, such as those measuring digit and line-break handling, correlate with task accuracy. We hope TokEval enables more principled tokenizer evaluation, replacing pretraining sweeps with intrinsic measurement wherever the two agree.

摘要:語言模型的分詞器通常在最小評估下被選擇,儘管它們的設計選擇直接影響模型的能力。這部分可以歸因於對哪些分詞器特性影響下游性能的理解有限。我們引入了 TokEval,一個超越標準測量(如生育率和壓縮率)的分詞器評估指標框架,以捕捉語言和結構上有意義的特性,例如 UTF-8 字符邊界的完整性和數字位值邊界對齊以進行數學運算。為了驗證這些指標是否能預測下游模型性能,我們進行了受控的語言模型預訓練實驗,僅改變分詞器的訓練數據混合、預分詞策略和訓練算法。我們在每字節位數(這是一個與分詞器無關的困惑度版本)和幾個基準上評估了結果模型,涵蓋語言理解、數學推理和代碼生成。我們的實驗表明,不同的內在特性對模型能力有不同的影響:信息理論指標預測語言建模能力(斯皮爾曼相關係數高達 0.80),而結構敏感指標,例如測量數字和換行處理的指標,則與任務準確性相關。我們希望 TokEval 能夠實現更有原則的分詞器評估,並在兩者一致的地方用內在測量取代預訓練掃描。

Delegation Asymmetry in Agentic Recommender Systems: Measuring Two-Sided Receptivity in Online Dating

2608.18058v1 by Daria Leshchikova, Valentina V. Kuskova, Dmitry Zaytsev, Valerii Klimov

Autonomous LLM agents that converse on a user's behalf are an emerging design pattern in matching platforms, yet their viability depends on a condition rarely examined: users must accept not only delegating conversation to an agent, but also receiving agent-mediated communication from others. We study this condition using two large-scale surveys of active users of a major dating platform (N=2,894 on generative profile features; N=2,617 on autonomous conversational agents, fielded in two languages). We develop a latent-variable measurement model of agent receptivity based on graded response models with latent regression, and show via model comparison that willingness to send and willingness to receive agent communication are distinct constructs: highly correlated (rho=0.92) but separable (Delta BIC=52), with partial measurement invariance across languages. The model quantifies a systematic delegation asymmetry: deploying one's own agent requires far lower receptivity (threshold -0.38) than engaging a counterpart's agent (+0.32; full engagement +1.39), and mean deployment propensity exceeds engagement propensity roughly threefold. Under a random-pairing counterfactual derived from stated receptivity, only 4-13% of directed dyads combine agent deployment with receiver engagement, with a pronounced gender-directional imbalance. Design counterfactuals quantify the levers: a reciprocity requirement cuts interaction volume by half or more by excluding nearly two-thirds of would-be deployment, while routing agent contacts on receive receptivity triples per-contact engagement, a lift that survives out-of-sample validation with the target item held out (AUC 0.88, 3.1x quartile lift under respondent-level cross-validation). We discuss implications for agentic recommender design, including disclosure, opt-in mechanics, and receptivity-aware matchmaking.

摘要:自主 LLM 代理人代表用戶進行對話是一種新興的設計模式,尤其在配對平台上,但其可行性取決於一個鮮少被檢視的條件:用戶必須接受不僅將對話委託給代理人,還要接受來自他人的代理人中介通信。我們通過對一個主要約會平台的活躍用戶進行兩項大規模調查來研究這一條件(N=2,894 針對生成型個人資料特徵;N=2,617 針對自主對話代理人,調查以兩種語言進行)。我們基於潛在回應模型和潛在回歸開發了一個代理人接受度的潛變量測量模型,並通過模型比較顯示,發送代理人通信的意願和接收代理人通信的意願是不同的構念:高度相關(rho=0.92)但可分離(Delta BIC=52),在語言間具有部分測量不變性。該模型量化了一種系統性的委派不對稱:部署自己的代理人所需的接受度(閾值 -0.38)遠低於與對方的代理人互動所需的接受度(+0.32;完全互動 +1.39),而平均部署傾向約為互動傾向的三倍。在基於聲明的接受度推導的隨機配對反事實中,只有 4-13% 的定向雙人組合將代理人部署與接收者互動結合,並且存在明顯的性別導向不平衡。設計反事實量化了杠杆:互惠要求通過排除近三分之二的潛在部署將互動量減少一半或更多,而根據接收接受度路由代理人聯繫則使每次聯繫的互動增加三倍,這一提升在樣本外驗證中仍然有效,目標項目被保留(AUC 0.88,受訪者層級交叉驗證下的四分位提升 3.1 倍)。我們討論了對代理推薦設計的影響,包括披露、選擇加入機制和接受度敏感的配對。

StagedWorkspace: A Versioned Workspace for Knowledge-Work Agents

2608.18050v1 by Yining Hua, Hongbin Na, Yifan Zhou, Akshay Kalose, Cyrus Ayubcha, Levi Lian

AI agents increasingly perform knowledge work (i.e., produce and modify persistent digital artifacts such as code repositories, documents, spreadsheets, slides, reports), yet the parsed views they search, the native files they edit, the changes they review, and the artifacts they submit can refer to different versions of the same work product. We formulate this as a workspace-state contract: every view should be explicitly tied to a version of the evolving workspace state. Coding agents partly address this need through repository contracts for search, diffs, and tests, whereas an analogous contract is less explicit for PDFs, spreadsheets, slides, notebooks, and mixed-format project folders. We propose StagedWorkspace, a versioned workspace for knowledge-work agents. The workspace binds parsed records and review diffs to content hashes of the native files as they change. In fixed-harness ablations on OfficeQA Pro and APEX-Agents, dual parsed/native access has the highest point estimate for every tested model; relative to the more limiting single view, it improves OfficeQA Pass@1 by 8.3-12.1 points and APEX mean rubric score by 4.7-9.2 points. SW-AGENT scores 63.9% with Gemini 3.1 Pro on OfficeQA and 42.1 with GPT-5.4 Nano on APEX, compared with published same-model scores of 29.3% and 25.5, respectively. A paired review-axis ablation on 57 file-editing tasks further finds higher observed scores when diffs are visible. These results identify workspace state as an experimental variable in knowledge-work agents and motivate benchmarks that score evidence, staged edits, and submitted artifacts as explicit state transitions.

摘要:AI 代理人越來越多地執行知識工作(即,產生和修改持久的數位工件,如代碼庫、文件、電子表格、簡報、報告),然而他們所搜尋的解析視圖、編輯的原始文件、審查的變更以及提交的工件可能指的是同一工作產品的不同版本。我們將此表述為工作區狀態合約:每個視圖應明確與不斷演變的工作區狀態的某個版本相關聯。編碼代理人部分通過針對搜索、差異和測試的庫合約來滿足這一需求,而對於 PDF、電子表格、簡報、筆記本和混合格式的項目文件夾,類似的合約則不那麼明確。我們提出了 StagedWorkspace,一個針對知識工作代理人的版本化工作區。該工作區將解析記錄和審查差異綁定到隨原始文件變更的內容哈希。在 OfficeQA Pro 和 APEX-Agents 的固定裝置消融實驗中,雙重解析/原生訪問對於每個測試模型的最高點估計;相對於更具限制性的單一視圖,它將 OfficeQA Pass@1 提高了 8.3-12.1 分,將 APEX 的平均評分提高了 4.7-9.2 分。SW-AGENT 在 OfficeQA 上的得分為 63.9%,在 APEX 上的得分為 42.1,與已發表的同模型得分分別為 29.3% 和 25.5 相比。對 57 個文件編輯任務的配對審查軸消融進一步發現,當差異可見時,觀察到的得分更高。這些結果將工作區狀態確定為知識工作代理人的實驗變量,並激勵對證據、分階編輯和提交工件進行明確狀態轉換的基準評分。

Language Has Two Parameters: Narrative-Induced Semantic Plasticity and Phase-Sensitive Interpretation

2608.18041v1 by Hollis Robbins

Language has two parameters. Count how often words occur together and you estimate amplitude, the strength of association. Word embeddings and attention weights refine that count, which sums every writer in the corpus together. This paper claims a second parameter, phase, which signed weights learned from a corpus do not supply. Phase exists only between meanings: it determines how coactivated meanings combine, and it can reverse what a meaning contributes while that meaning stays fully present. A speaker can set phase in the signal through linguistic form; encounters install phase relations and history distributes them. Population averaging deletes history-indexed phase: agent-deindexed corpora identify the population marginal state and determine no individual or dyadic state, at any scale. The standard transformer has no explicit representation for phase in frozen inference, and the interpretability program measuring progress by monosemanticity is optimizing against it: the coexistence it treats as a defect is the condition of allusion, irony, and quotation. Six predictions test whether a suppressed meaning stays active, whether encounter order changes what a phrase does, whether marking the signal changes how a shared phrase is taken, and whether a model given a history is changed by it or only informed about it. The claim defended is the weak version: interpretation requires a second relational parameter, signed, persistent, and indexed to individuals and dyads. Quantum probability is one notation for the parameter; nothing in the formalism claims quantum processes in the brain. The strong version, that the quantum calculus constrains these phenomena as signed classical models do not, rests on an encounter-order constraint not yet derived. The architecture the theory calls for is a language model with agent-indexed, phase-bearing semantic states.

摘要:語言有兩個參數。計算單詞共同出現的頻率,你可以估算幅度,即聯繫的強度。詞嵌入和注意力權重會細化這個計數,這個計數將語料庫中的每位作者的貢獻相加。本文主張第二個參數,階段,這是從語料庫學習的簽名權重所無法提供的。階段僅存在於意義之間:它決定了如何共同激活的意義結合,並且它可以反轉一個意義所貢獻的內容,同時該意義仍然完全存在。說話者可以通過語言形式在信號中設置階段;遭遇安裝階段關係,而歷史則分配它們。人口平均會刪除歷史索引的階段:去代理索引的語料庫識別出人口的邊際狀態,並且在任何規模上都不確定任何個體或雙方狀態。標準Transformer在凍結推理中沒有對階段的明確表示,而測量進展的可解釋性程序則是針對它進行優化的:它所視為缺陷的共存是暗示、諷刺和引用的條件。六個預測測試被壓制的意義是否保持活躍,遭遇順序是否改變短語的作用,標記信號是否改變共享短語的理解,以及給定歷史的模型是否因其而改變或僅僅是被告知。所辯護的主張是弱版本:解釋需要第二個關係參數,這個參數是簽名的、持久的,並且索引到個體和雙方。量子概率是一種參數的表示法;形式主義中沒有任何內容聲稱大腦中的量子過程。強版本,即量子微積分限制這些現象,而簽名的經典模型則不,依賴於尚未推導出的遭遇順序約束。該理論所要求的架構是一個具有代理索引、承載階段的語義狀態的語言模型。

Chain-of-Experience for Continual LLM Improvement

2608.18027v1 by Haoqin Tu, Yunhao Fang, Yizhong Wang, Cihang Xie, Shen Yan

Humans continuously learn from experience, whereas conventional large language model (LLM) evaluations ignore the models' ability to improve through inference-time interaction. In this paper, we study how LLMs learn from iterative experience at test time, a setting we refer to as Chain-of-Experience (CoE), where models accumulate experiential traces through iterative interactions with self or environmental feedback to form a continual improvement loop beyond zero-shot inference. We instantiate CoE with diverse feedback mechanisms, including model self-feedback and environmental signals such as correctness or public coding test pass rates, and evaluate across math, coding, and knowledge domains using 8 LLMs, including GPT-5, Gemini-2.5 Pro, Claude-4.5 Sonnet. Our study shows that leveraging iterative experience consistently outperforms feedback-free baselines, achieving substantial gains with self feedback alone, alongside a 5.6% overall improvement and 19% lower API cost across tasks and models. We further show that combining complementary feedback channels (e.g., model and correctness signals) yields additional gains, and that CoE delivers higher accuracy per token than existing test-time strategies. We observe a positive correlation between LLM base ability and improvement capacity, and show that models remain robust under weak or spurious feedback, with different feedback contributing to distinct improvement aspects and most gains emerging early in the iterations.

摘要:人類不斷從經驗中學習,而傳統的大型語言模型(LLM)評估則忽略了模型通過推理時互動來改進的能力。在本文中,我們研究了 LLM 如何在測試時從迭代經驗中學習,這種情境我們稱之為經驗鏈(Chain-of-Experience, CoE),在這裡模型通過與自身或環境反饋的迭代互動積累經驗痕跡,以形成超越零-shot 推理的持續改進循環。我們用多樣的反饋機制來實現 CoE,包括模型自我反饋和環境信號,如正確性或公共編碼測試通過率,並使用 8 種 LLM 進行數學、編碼和知識領域的評估,包括 GPT-5、Gemini-2.5 Pro 和 Claude-4.5 Sonnet。我們的研究表明,利用迭代經驗的表現始終優於無反饋的基準,僅依靠自我反饋就實現了顯著的增益,並在各任務和模型中達到 5.6% 的整體改進和 19% 的 API 成本降低。我們進一步顯示,結合互補的反饋通道(例如模型和正確性信號)會產生額外的增益,並且 CoE 在每個 token 上提供的準確性高於現有的測試時策略。我們觀察到 LLM 的基本能力與改進能力之間存在正相關,並顯示模型在弱或虛假反饋下仍然保持穩健,不同的反饋對不同的改進方面有所貢獻,大多數增益在迭代的早期出現。

Why GPT-Style Models Do Not Directly Transfer to Symbolic Music: Compression in the Wrong Coordinate System

2608.18025v1 by Yi Wang

GPT-style models achieve strong performance by representing language with finite vocabularies of reusable discrete tokens. This success has motivated symbolic music tokenizations to treat recurring musical structures, such as chords, motifs, and phrases, as reusable units analogous to linguistic tokens. However, tokenization derives its advantage not from reusable combinations alone, but from compression: effective compression requires coordinates in which recurring regularities form stable and predictable conditional distributions. The key problem is therefore not to find larger musical combinations, but to discover the coordinate system in which musical facts become predictively compressible. We formulate the Effectiveness--Losslessness Framework and define tokenization as the construction of a predictively effective and relationally lossless coordinate system. The Predictive Effectiveness Principle defines the Fact--Token Boundary: decoupling and denesting construct coordinate interfaces that expose predictive regularities. The Relational Losslessness Principle defines the Token--State Boundary: tokenization stops before context-dependent relations are fixed, leaving their computation to model states. Controlled symbolic-music experiments validate these boundaries. Effective coordinate construction improves predictive compressibility, while fixed relational projections constrain contextual modeling. Sequence compaction alone does not guarantee predictive compression, while preserving contextual freedom allows higher-order musical organization to emerge without explicit structural labels. These results reveal why GPT-style models do not transfer directly across modalities: architectures transfer, but tokenization interfaces do not. Tokenization must discover effective representations while preserving the relational freedom from which contextual structure can emerge.

摘要:GPT風格的模型通過使用有限的可重用離散標記詞彙來表示語言,從而實現強大的性能。這一成功促使符號音樂標記化將重複出現的音樂結構,如和弦、主題和短語,視為類似於語言標記的可重用單元。然而,標記化的優勢並不僅僅來自可重用的組合,而是來自壓縮:有效的壓縮需要坐標系,在這些坐標系中,重複的規律形成穩定且可預測的條件分佈。因此,關鍵問題不在於尋找更大的音樂組合,而在於發現音樂事實變得可預測壓縮的坐標系。我們制定了有效性-無損框架,並將標記化定義為構建一個預測有效且關係無損的坐標系。預測有效性原則定義了事實-標記邊界:解耦和去嵌套構建坐標接口,揭示預測規律。關係無損原則定義了標記-狀態邊界:標記化在上下文依賴關係固定之前停止,將其計算留給模型狀態。受控的符號音樂實驗驗證了這些邊界。有效的坐標構建提高了預測壓縮性,而固定的關係投影則限制了上下文建模。僅僅進行序列壓縮並不能保證預測壓縮,而保留上下文自由則允許更高階的音樂組織在沒有明確結構標籤的情況下出現。這些結果揭示了為什麼GPT風格的模型無法直接跨模態轉移:架構可以轉移,但標記化接口則不能。標記化必須在保留關係自由的同時發現有效的表示,從而使上下文結構能夠出現。

Can Large Language Models Explain Flight Safety Events? A Prior-Guided Semantic LLM-based Approach

2608.18017v1 by Lu Xu, Xu Li, Linjiang Zheng, Fan Li, Riquan Zhang, Jiaxing Shang

Improving flight safety with flight data requires not only accurate detection of risk events, but more importantly, clear interpretation of their underlying causes at the level of pilot control behavior. Existing explainable AI techniques, such as feature importance maps, often require considerable domain knowledge to translate them into operationally meaningful explanations. Large Language Models (LLMs), which excel at language reasoning, bring a promising solution to this issue. However, applying LLMs in this domain presents key challenges such as modal inconsistency, limited classification ability, scarcity of task-specific data for fine-tuning, and lack of domain knowledge. To overcome these challenges, we propose FlightLLM, a prior-guided semantic LLM-based approach for interpretable flight safety analysis. Specifically, we first perform feature engineering to address modal inconsistency, combining statistical descriptors with physically meaningful flight indicators. This representation is further processed by a Semantic Discretization module, which converts abstract numerical patterns into qualitative descriptions that are more compatible with language reasoning. In addition, since LLMs are not inherently strong classifiers, CatBoost is incorporated as a statistical expert, and its prediction results are injected into the prompt as prior guidance. A contrastive few-shot learning strategy is further adopted to compensate for limited data. Finally, we design structured prompts to embed aviation-specific knowledge into the inference process. Using hard landing, a representative risk event with complex causal mechanisms, as an anchor point, we evaluate FlightLLM on a dataset of 704 real-world A320 flight samples. Experimental results show that the proposed approach achieves competitive classification performance while generating direct and reasonable explanations for event causes.

摘要:改善飛行安全需要不僅準確檢測風險事件,更重要的是在飛行員控制行為層面清晰解釋其潛在原因。現有的可解釋AI技術,如特徵重要性圖,通常需要相當的領域知識才能將其轉化為具有操作意義的解釋。大型語言模型(LLMs)在語言推理方面表現出色,為這一問題帶來了有希望的解決方案。然而,在這一領域應用LLMs面臨著關鍵挑戰,如模式不一致、有限的分類能力、缺乏特定任務的數據以進行微調,以及缺乏領域知識。為了克服這些挑戰,我們提出了FlightLLM,一種基於語義的先驗引導LLM方法,用於可解釋的飛行安全分析。具體而言,我們首先進行特徵工程以解決模式不一致,將統計描述符與具有物理意義的飛行指標相結合。這一表示進一步由語義離散化模塊處理,將抽象的數字模式轉換為更符合語言推理的定性描述。此外,由於LLMs本身並不是強大的分類器,因此CatBoost被納入作為統計專家,其預測結果被注入到提示中作為先驗指導。進一步採用了對比少樣本學習策略以彌補數據的有限性。最後,我們設計了結構化提示,將航空特定知識嵌入推理過程中。以硬著陸作為錨點,這是一個具有複雜因果機制的代表性風險事件,我們在704個真實世界A320飛行樣本的數據集上評估FlightLLM。實驗結果表明,所提出的方法在生成事件原因的直接和合理解釋的同時,實現了具有競爭力的分類性能。

The IOL-AI Challenge: An Open Challenge towards Advancing Linguistic Reasoning

2608.18011v1 by Eduardo Sánchez, Rita Berrada, Dan-Mircea Mirea, Sara Rajaee, Alexander Piperski, Ana Meta Dolinar, Boris Iomdin, Andrey Nikulin, Mariya Shmatova, Marzieh Fadaee, Julia Kreutzer

Reasoning in LLMs is overwhelmingly studied in domains that provide a model with rules: mathematics and code. Linguistic puzzles invert this: the solver must first discover the system before reasoning within it. We present the IOL-AI Challenge, an open-science competition run on the unseen problems of the International Linguistics Olympiad (IOL) 2026 Individual Contest, evaluated both automatically and, for the first time, by members of the official IOL Jury under the same rubrics applied to human contestants. The challenge drew 731 submissions from 46 teams under a strict compute budget (one T4, 30 mins). We additionally benchmark 15 unconstrained frontier and open models, with Claude Opus 4.8 earning a jury score equivalent to a gold medal, while both resource-constrained systems we submitted for jury grading scored in the range of the bottom 5% of contestants. Capability was not determined by scale: 14B submissions outperform models twice their size, and gains come from decoding and output-handling rather than model capacity. We also found that automatic metrics rank systems exactly as the jury does, but compress the scale, upscoring weak systems by ~13 points and understating strong ones. Our analysis shows that while frontier models might have prior knowledge about some of the problem languages, it does not significantly help them solve the linguistic reasoning tasks, leaving linguistic reasoning as a strong benchmarking proxy for generalizable reasoning skills.

摘要:推理在大型語言模型(LLMs)中的研究主要集中在提供規則的領域:數學和程式碼。語言謎題則顛倒了這一點:解題者必須首先發現系統,然後才能在其中進行推理。我們提出了IOL-AI挑戰賽,這是一項開放科學競賽,基於2026年國際語言奧林匹亞(IOL)個人賽的未見問題進行評估,這次評估既有自動評分,還首次由官方IOL評審團成員根據與人類參賽者相同的標準進行評分。這次挑戰吸引了46個團隊提交的731份作品,並在嚴格的計算預算下進行(一個T4,30分鐘)。我們還基準測試了15個不受限制的前沿和開放模型,其中Claude Opus 4.8獲得了相當於金牌的評審分數,而我們提交給評審打分的兩個資源受限系統的分數則落在參賽者的底部5%範圍內。能力並不是由規模決定的:14B的提交表現超過了規模是其兩倍的模型,並且性能的提升來自於解碼和輸出處理,而非模型容量。我們還發現,自動指標的排名與評審的排名完全一致,但壓縮了評分範圍,將弱系統的分數提高了約13分,而低估了強系統的分數。我們的分析顯示,儘管前沿模型可能對某些問題語言有先前的知識,但這並未顯著幫助它們解決語言推理任務,這使得語言推理成為通用推理能力的強基準代理。

Policy-Invariant Reward Shaping from LLM Feedback: A Framework for Hybrid RL Agents

2608.18008v1 by Christophe D. Hounwanou, John Emeka Eze, Yaé U. Gaba

Combining large language models with reinforcement learning is increasingly explored, yet the theoretical status of LLM-derived reward signals is often left implicit. We formalize the hybrid LLM-planner and RL-controller architecture as a Goal-Augmented Markov Decision Process and show that when the LLM per-state progress score is used as a bounded potential function, the resulting shaping term preserves the optimal policy set even when the LLM scores are inaccurate. This guarantee is stronger than what general LLM-as-reward approaches provide. We verify the result numerically on a small MDP under four potential configurations, including an adversarial one scaled to twenty times the base reward magnitude.

摘要:結合大型語言模型與強化學習的研究越來越受到關注,然而 LLM 衍生的獎勵信號的理論地位常常被隱含。我們將混合的 LLM 規劃者和 RL 控制器架構形式化為目標增強馬可夫決策過程,並顯示當 LLM 每狀態的進展分數被用作有界潛在函數時,所產生的塑形項即使在 LLM 分數不準確的情況下也能保持最佳政策集。這一保證比一般的 LLM 作為獎勵的方法提供的要強。我們在一個小型 MDP 上對結果進行了數值驗證,考慮了四種潛在配置,包括一個對抗性配置,其規模為基礎獎勵大小的二十倍。

Traceable Trust for action-ready artificial intelligence in bioscience

2608.17997v1 by Huayu Xin, Yizhi Cai, Mukilan Deivarajan Suresh, Gavin Michael Farrell, Iwona Gajda, Charlie Harrison, Conor Houghton, Mato Lagator, Yang Lu, Virginia Portillo, Reyer Zwiggelaar, Sebastian Lobentanzer

Artificial intelligence (AI) is becoming part of the working infrastructure of the biosciences. AI models can predict biomolecular structures, design proteins, rank variants, annotate images, recommend strains and optimise experimental conditions. We argue that the decision to use an AI output to guide laboratory action is a key juncture for trustworthy research and should follow a defined, reviewable process. We propose Traceable Trust as a proportionate assessment-and-design framework for this output-to-action boundary. It asks what evidence supports the output, what capability is being claimed, what agency has been delegated, what threshold authorises action, who can override it and how outcomes inform later decisions. We illustrate the framework through three case studies spanning ecosystem resources, project design and laboratory action. Together, the cases show how trust can be documented where AI outputs begin to shape scientific work.

摘要:人工智慧(AI)正逐漸成為生物科學工作基礎設施的一部分。AI 模型可以預測生物分子結構、設計蛋白質、排名變異體、註解圖像、推薦菌株並優化實驗條件。我們認為,使用 AI 輸出來指導實驗室行動的決定是值得信賴的研究的一個關鍵時刻,應遵循一個明確的、可審查的過程。我們提出可追溯的信任作為這一輸出到行動邊界的比例評估與設計框架。它詢問什麼證據支持該輸出、聲稱了什麼能力、授予了什麼代理權、什麼門檻授權行動、誰可以覆蓋它以及結果如何影響後續決策。我們通過三個案例研究來說明該框架,這些案例涵蓋了生態系統資源、項目設計和實驗室行動。這些案例共同展示了在 AI 輸出開始塑造科學工作時,如何記錄信任。

Judge, Retrieve, or Abstain: Uncertainty-Guarded LLM Judging with Provable Risk Guarantees

2608.17994v1 by Sher Badshah, Ali Emami, Hassan Sajjad

Using LLMs as judges has become standard practice for evaluating model outputs at scale. This is particularly common for subjective, open-ended tasks such as assessing helpfulness or alignment, where no single reference answer exists. However, objective tasks introduce a distinct reliability challenge for reference-free LLM judging. In the absence of a reference answer, the judge evaluates factual correctness either through its parametric knowledge or through tool augmentation. Although the former enables efficient evaluation, the judge may hallucinate or lack sufficient evidence for its verdict. Conversely, tool augmentation can provide additional evidence but introduces extra computational cost and requires an appropriate mechanism to determine when and how that evidence should be used reliably. More importantly, neither approach alone provides formal control over the risk of accepted verdicts or guarantees their reliability at a specified level. We propose a risk-controlled framework that calibrates uncertainty thresholds on a held-out set so that the false discovery rate among accepted verdicts remains below a user-specified level~$α$ with high probability, using finite-sample Clopper--Pearson intervals. When the parametric mode is not sufficiently confident, the instance is routed to a retrieval-augmented mode, where the judge gathers web evidence and re-evaluates the instance under a second calibrated threshold. The finite-sample guarantee carries over to this two-threshold routing without additional assumptions. Across open-domain QA benchmarks and judges of varying scales, the framework maintains the target error rate while achieving substantially higher coverage than single-mode baselines.

摘要:使用大型語言模型作為評審已成為大規模評估模型輸出的標準做法。這在主觀的、開放式的任務中尤其常見,例如評估有用性或一致性,因為這類任務並不存在單一的參考答案。然而,客觀任務對於無參考的 LLM 評審引入了明顯的可靠性挑戰。在缺乏參考答案的情況下,評審通過其參數知識或工具增強來評估事實的正確性。雖然前者能夠實現高效評估,但評審可能會出現幻覺或缺乏足夠的證據來支持其裁決。相反,工具增強可以提供額外的證據,但會引入額外的計算成本,並需要適當的機制來確定何時以及如何可靠地使用這些證據。更重要的是,單獨使用這兩種方法都無法對接受的裁決風險提供正式控制或保證其在特定水平上的可靠性。我們提出了一個風險控制框架,該框架在保留集上校準不確定性閾值,以便接受的裁決中的假陽性率以高概率保持在用戶指定的水平~$α$ 以下,使用有限樣本的 Clopper--Pearson 區間。當參數模式的信心不足時,實例會被路由到檢索增強模式,在該模式下,評審收集網絡證據並在第二個校準閾值下重新評估該實例。有限樣本的保證在這個雙閾值路由中延續,無需額外假設。在開放域問答基準和不同規模的評審中,該框架在保持目標錯誤率的同時,實現了顯著高於單一模式基準的覆蓋率。

Against Political Polarization: A Unified Framework for Tracing Evolving Political Ideologies on Social Media

2608.17987v1 by Yijie Xu, Chao Wang, Hui Xiong

The rapid growth of social media has greatly influenced political discourse, highlighting the need to understand individual political ideologies and their temporal dynamics. This task faces challenges such as data scarcity, abundant non-political content, costly and bias-prone manual annotation, and difficulty in modeling future ideological inclinations. To address these issues, we propose TSN4PI, a unified framework for tracking the evolution of political ideologies on social media. It includes two core modules. The PIDN uses large language models with style transfer and unsupervised domain adaptation to enable robust ideology detection and filter irrelevant content from noisy, cross-domain data. The PIPN employs temporal graph neural networks to predict future ideological shifts, enabling comprehensive analysis of ideology presence, intensity, and evolution. We release two large-scale datasets for noncommercial research use to facilitate further work. Extensive case studies on multiple platforms (X and Truth Social) validate the effectiveness of TSN4PI and provide empirical insights into political polarization and the evolution of online ideologies. Our findings offer a nuanced perspective, advancing both methodological development and empirical understanding in this field.

摘要:社交媒體的快速增長對政治話語產生了重大影響,突顯了理解個人政治意識形態及其時間動態的必要性。這項任務面臨著數據稀缺、非政治內容豐富、昂貴且易受偏見影響的手動標註以及未來意識形態傾向建模困難等挑戰。為了解決這些問題,我們提出了TSN4PI,一個用於追蹤社交媒體上政治意識形態演變的統一框架。它包括兩個核心模塊。PIDN使用大型語言模型結合風格轉換和無監督領域適應,以實現穩健的意識形態檢測並過濾來自嘈雜的跨領域數據中的無關內容。PIPN則利用時間圖神經網絡來預測未來的意識形態變化,使得對意識形態的存在、強度和演變進行全面分析成為可能。我們釋放了兩個大型數據集供非商業研究使用,以促進進一步的研究工作。在多個平台(X和Truth Social)上進行的廣泛案例研究驗證了TSN4PI的有效性,並提供了對政治極化和在線意識形態演變的實證見解。我們的發現提供了一個細緻的視角,推進了該領域的方法論發展和實證理解。

Dual Co-Train: Cross-Dataset Ultrasound Tongue Segmentation Under Extreme Data Scarcity

2608.17983v1 by Alisher Myrgyyassov, Zhen Song, Bruce Xiao Wang, Yu Sun, Min Ney Wong, Yihao Zhou, Yongping Zheng

Ultrasound tongue contour segmentation remains challenging under cross-dataset domain shift, where limited annotations, probe variability, and acquisition noise often degrade model generalization. We present a source-free domain adaptation framework for robust ultrasound tongue segmentation built on a lightweight UltraUNet backbone. Starting from a checkpoint pretrained on only five labeled source images, simulating an underfitted constrained source model, the proposed method adapts to a fully-unlabeled target domain by iteratively refining pseudo-labels, filtering unreliable masks with a contour-based quality-control module, and generating target-style synthetic image-mask pairs through a segmentation-guided conditional GAN. The student model is then trained on a mixture of clean pseudo-labeled target images, noisy pseudo-labels with consistency regularization, and synthetic samples, enabling closed-loop adaptation without access to source data. We evaluate the method on 12 source-target transfer pairs across eight ultrasound tongue imaging datasets, and conduct source-size scaling experiments and ablation studies. Across all comparisons, the proposed framework improves segmentation overlap and contour accuracy over the baselines, including supervised ones. These results suggest that task-specific pseudo-label refinement and synthetic target-style augmentation can substantially improve source-free adaptation for ultrasound tongue imaging.

摘要:超聲波舌頭輪廓分割在跨數據集域轉移下仍然具有挑戰性,因為有限的註釋、探頭變異性和獲取噪聲常常會降低模型的泛化能力。我們提出了一個無源域適應框架,用於穩健的超聲波舌頭分割,基於輕量級的UltraUNet骨幹。從僅在五張標記源圖像上預訓練的檢查點開始,模擬一個欠擬合的受限源模型,所提出的方法通過迭代地細化偽標籤,使用基於輪廓的質量控制模塊過濾不可靠的掩膜,並通過分割引導的條件GAN生成目標風格的合成圖像-掩膜對,適應於完全無標記的目標域。然後,學生模型在一組乾淨的偽標記目標圖像、帶有一致性正則化的噪聲偽標籤和合成樣本的混合上進行訓練,使得在無需訪問源數據的情況下實現閉環適應。我們在八個超聲波舌頭成像數據集上的12對源-目標轉移對上評估了該方法,並進行了源大小縮放實驗和消融研究。在所有比較中,所提出的框架在分割重疊和輪廓準確性方面優於基準,包括監督學習的基準。這些結果表明,特定任務的偽標籤細化和合成目標風格的增強可以顯著改善超聲波舌頭成像的無源適應能力。

When Writing Style Drifts: Benchmarking Authorship Verification under Distribution Shifts in Genre, Time and the AI-Era

2608.17979v1 by Lotta Kiefer, Brisca Balthes, Christoph Leiter, Yamen Ajjour, Elena Schmidt, Steffen Eger

Authorship verification (AV) assumes that an author's writing style remains sufficiently stable to distinguish it from that of other writers. In practice, however, this assumption is challenged by distribution shifts caused by changes in genre, time, and AI-assisted writing. Existing AV benchmarks typically study these factors in isolation and focus predominantly on English, limiting our understanding of model robustness under realistic conditions. We introduce AVShift, the first German benchmark for systematically evaluating AV under multiple distribution shifts. AVShift comprises over 150,000 text pairs spanning three genres and 21 years, enabling controlled evaluation of cross-genre, temporal, and AI-era shifts within a unified framework. We benchmark representative feature-based, embedding-based, and LLM-based approaches. Our experiments show that fine-tuned LLMs generalize best across genres and benefit substantially from stylistically diverse training data. We further demonstrate that temporal drift is one of the strongest factors affecting AV, with performance degrading significantly as the time gap between documents increases. In contrast, we find no evidence of a measurable AI-era distribution shift within AVShift. Finally, our feature analysis reveals stylistic features that remain stable across genres, while their relative importance varies depending on the specific genre transition. We release AVShift and our code for future research.

摘要:作者驗證(AV)假設作者的寫作風格保持足夠穩定,以便與其他作家的風格區分開來。然而,在實際操作中,這一假設受到由於類型、時間和AI輔助寫作變化而引起的分佈變化的挑戰。現有的AV基準通常孤立地研究這些因素,並主要集中在英語上,限制了我們在現實條件下對模型穩健性的理解。我們介紹了AVShift,這是第一個德語基準,用於系統地評估多種分佈變化下的AV。AVShift包含超過150,000對文本,涵蓋三個類型和21年,能夠在統一框架內進行跨類型、時間和AI時代變化的受控評估。我們基準測試了代表性的基於特徵、基於嵌入和基於LLM的方法。我們的實驗表明,微調的LLM在各類型之間的泛化效果最佳,並且從風格多樣的訓練數據中獲益良多。我們進一步證明,時間漂移是影響AV的最強因素之一,隨著文檔之間時間間隔的增加,性能顯著下降。相比之下,我們在AVShift中沒有發現可測量的AI時代分佈變化的證據。最後,我們的特徵分析揭示了在各類型中保持穩定的風格特徵,而它們的相對重要性則根據具體的類型轉換而有所不同。我們發布了AVShift和我們的代碼以供未來研究使用。

Too Sure to Be Safe: Model Calibration for Reliable Log Anomaly Detection

2608.17965v1 by Bin Li, Dongdong Wang, Siyang Lu

Online log anomaly detection is critical for maintaining the reliability of large-scale computing systems. Although recent language model-based log anomaly detectors achieve strong detection performance, their confidence estimates remain poorly calibrated. We show that these detectors frequently assign excessive confidence to incorrect predictions, particularly for anomalous logs under severe class imbalance. Moreover, confidence on erroneous predictions remains persistently high even when conventional calibration metrics indicate good calibration, creating a critical reliability gap for operational monitoring systems. To address this issue, we propose Log Reconstruction and Distance (LoRD), a lightweight post-hoc calibration framework for reliable log anomaly detection. LoRD learns prediction-route-specific reliability models from latent representations of correctly classified validation samples and estimates prediction reliability through route-wise reconstruction distances. Based on the estimated reliability, LoRD selectively recalibrates high-risk predictions to suppress overconfident errors while preserving reliable predictions. Extensive experiments on four large-scale log benchmark datasets and multiple language model-based detectors demonstrate that LoRD consistently improves confidence reliability and substantially reduces overconfident anomaly-related errors without sacrificing anomaly detection performance.

摘要:線上日誌異常檢測對於維護大規模計算系統的可靠性至關重要。儘管最近基於語言模型的日誌異常檢測器在檢測性能上表現出色,但它們的信心估計仍然校準不佳。我們顯示這些檢測器經常對錯誤預測賦予過高的信心,特別是在嚴重類別不平衡的異常日誌中。此外,即使在傳統的校準指標顯示良好校準的情況下,對錯誤預測的信心仍然持續偏高,這為運營監控系統創造了關鍵的可靠性差距。為了解決這個問題,我們提出了日誌重建與距離(LoRD),這是一個輕量級的事後校準框架,用於可靠的日誌異常檢測。LoRD從正確分類的驗證樣本的潛在表示中學習特定於預測路徑的可靠性模型,並通過路徑重建距離來估計預測的可靠性。根據估計的可靠性,LoRD選擇性地重新校準高風險預測,以抑制過度自信的錯誤,同時保留可靠的預測。在四個大規模日誌基準數據集和多個基於語言模型的檢測器上進行的廣泛實驗表明,LoRD持續改善信心可靠性,並顯著減少過度自信的異常相關錯誤,而不犧牲異常檢測性能。

Towards Zero-Shot Task Transfer with Neurosymbolic World Models

2608.17959v1 by Isidoro Tamassia, Lennert De Smet, Giuseppe Marra

State-of-the-art model-based reinforcement learning methods learn neural world models that allow policy improvement by planning in a latent space, without assumptions on the structure of the underlying environment. While expressive, these models are generally task-dependent: they learn uninterpretable latent representations that are tied to the training task and thus hard to generalize to new tasks. In this work, we present a novel world model formulation where the reward prediction only depends on a subset of structured, symbolic components of the whole latent state. Decoupling observation reconstruction and reward prediction allows us to learn world models that can adapt zero-shot, i.e. without further environment interactions, to new reward functions defined over the same symbolic state space. We discuss the main advantages and challenges of learning these neurosymbolic world models and demonstrate the strong generalisation properties of our approach over purely neural methods.

摘要:最先進的基於模型的強化學習方法學習神經世界模型,這些模型允許在潛在空間中進行規劃以改進策略,而不需要對基礎環境的結構做出假設。雖然這些模型具有表現力,但通常依賴於特定任務:它們學習的潛在表示難以解釋,並且與訓練任務緊密相關,因此難以泛化到新任務。在這項工作中,我們提出了一種新穎的世界模型公式,其中獎勵預測僅依賴於整個潛在狀態的結構化符號組件的子集。將觀察重建與獎勵預測解耦,使我們能夠學習能夠零次適應的世界模型,即在不進一步與環境互動的情況下,對定義在相同符號狀態空間上的新獎勵函數進行適應。我們討論了學習這些神經符號世界模型的主要優勢和挑戰,並展示了我們的方法相較於純神經方法的強泛化特性。

An Omitted Mode Is a Rare Rule: The Sampling-Verification Danger Law in Continuous Code World Models

2608.17956v1 by Javier Aguilar Martín

In the Code World Model paradigm an LLM synthesizes an executable world model that a classical planner searches, and the model is accepted when it reproduces sampled transitions. We ask what that acceptance certifies in continuous control. We define the pipeline's danger as an expected risk and isolate its exact factor: the probability that N i.i.d. gate rollouts all miss a critical event of probability r is exactly (1-r)^N; an independent acceptance sample adds its budget to the exponent. On three hybrid instruments the accepted mode-blind model is exploited: the planner is pinned at the mode boundary at a regret of nearly the whole attainable return. We prove a localization budget, valid at boundary points: models with Lipschitz constant at most L differing by eta at a point disagree above tolerance eps on a region of volume at least kappa((eta-eps)/L)^(d+m); the discontinuous reset modes studied pay no such budget. With real LLM synthesis, GPT-5.x repairs an omitted 1D clamp in 105 of 111 mode-containing draws -- every attempt exact on 50 of 56 instrument-stream blocks (95% CI [0.781, 0.960]). On 2D regions no artifact recovers the rule (0/156); eight targeted interventions leave the failure in place, and positive controls locate it: a located rule is not induced, while given form and location the constants follow exactly. A version-space certificate proves identification is class-relative: at the widest dose the declared fit succeeds in 20/20 blocks and every sample-consistent circle is within tolerance in 18/20. We prove a class of entry rules exactly consistent with every sample yet harmless at play, so identifiability is a measurable property of the instrument. Re-scoring all 1034 artifacts on independent samples confirms acceptance certifies sample consistency and no more: where the gate is provably informative it covers about two percent of the exploited planner's queries.

摘要:在代碼世界模型範式中,一個大型語言模型(LLM)合成了一個可執行的世界模型,供經典規劃器搜索,當模型重現抽樣轉換時便被接受。我們詢問這種接受在連續控制中證明了什麼。我們將管道的危險定義為預期風險,並孤立其確切因素:N個獨立同分佈的閘門展開全部錯過概率為r的關鍵事件的概率恰好是(1-r)^N;一個獨立的接受樣本將其預算添加到指數中。在三個混合工具上,接受的無模式模型被利用:規劃器在模式邊界被固定,後悔幾乎達到整個可獲得回報。我們證明了一個有效於邊界點的定位預算:在某一點上,Lipschitz常數最多為L的模型若在eta上有所不同,則在容忍度eps以上的區域內存在體積至少為kappa((eta-eps)/L)^(d+m)的差異;所研究的不連續重置模式不需要這樣的預算。通過實際的LLM合成,GPT-5.x在111個包含模式的抽樣中修復了105個遺漏的1D夾具——每次嘗試在56個工具流塊中的50個上都是精確的(95%置信區間 [0.781, 0.960])。在2D區域中,沒有任何工件恢復該規則(0/156);八個目標干預使失敗保持不變,而正控制則定位了它:一個定位的規則並未被誘導,而在給定形式和位置的情況下,常數恰好遵循。版本空間證明了識別是類別相對的:在最寬的劑量下,聲明的擬合在20/20塊中成功,每個樣本一致的圓圈在18/20中都在容忍範圍內。我們證明了一類與每個樣本完全一致但在遊戲中無害的進入規則,因此可識別性是工具的一個可測量屬性。對1034個工件在獨立樣本上的重新評分確認接受證明樣本一致性,並且僅此而已:在閘門被證明為信息豐富的地方,它覆蓋了約兩個百分比的被利用規劃者的查詢。

Do Large Language Models Play Six Degrees of Separation? Measuring Topological Compression in Long-Context Manifolds

2608.17950v1 by Md. Faiyaz Abdullah Sayeedi

Large Language Models (LLMs) demonstrate remarkable multi-hop reasoning capabilities over long contexts, yet the internal mechanisms enabling these distant cognitive leaps remain poorly understood. Traditional attention-based interpretability often fails to capture true semantic proximity due to routing artifacts like attention sinks. In this paper, we bypass attention weights to directly analyze the dynamic geometry of the hidden state manifold, proving that deep LLM latent spaces natively organize into Small-World networks. By sparsifying the continuous similarity matrices of long-context representations into unweighted graphs, we trace the connectivity between highly disjoint semantic anchors across two distinct architectures. Our findings reveal a sharp topological phase transition: while early syntactic layers remain entirely fractured, deep reasoning layers abruptly compress massive conceptual distances into highly navigable pathways strictly bounded by the "Six Degrees of Separation" limit (=< 6 semantic hops). Furthermore, we demonstrate the practical efficacy of this framework by applying it to zero-shot hallucination detection within Retrieval-Augmented Generation (RAG) using the RAGognize dataset. We show that factually grounded generations maintain structural integrity with their source context (approximately 3 hops), whereas hallucinations induce severe topological collapse. Ultimately, this work mathematically formalizes how transformers execute abstract reasoning and provides a novel, strictly geometric signature for evaluating factual reliability.

摘要:大型語言模型(LLMs)在長上下文中展現出卓越的多跳推理能力,但促成這些遙遠認知飛躍的內部機制仍然不甚了解。傳統的基於注意力的可解釋性常常無法捕捉到真實的語義接近性,這是由於路由伪影如注意力匯聚所致。在本文中,我們繞過注意力權重,直接分析隱藏狀態流形的動態幾何,證明深層LLM潛在空間本質上組織成小世界網絡。通過將長上下文表示的連續相似性矩陣稀疏化為無權重圖,我們追蹤兩個不同架構之間高度不相交的語義錨點之間的連接性。我們的研究結果揭示了一個明顯的拓撲相變:儘管早期的句法層完全破碎,深層推理層卻突然將巨大的概念距離壓縮成高度可導航的路徑,這些路徑嚴格受限於「六度分隔」的限制(=< 6語義跳躍)。此外,我們通過將此框架應用於檢索增強生成(RAG)中的零樣本幻覺檢測,展示了其實際效能,使用了RAGognize數據集。我們顯示,事實基礎的生成與其來源上下文保持結構完整(約3跳),而幻覺則引發嚴重的拓撲崩潰。最終,這項工作數學化了Transformer如何執行抽象推理,並提供了一種新穎的、嚴格的幾何特徵,用於評估事實可靠性。

SIGMA: SHAP-Guided Implicit-Trajectory Generation for Metadata-Free LLM-Based AutoFE

2608.17948v1 by Xuan Zheng, Kento Uchida, Shinichi Shirakawa

Recent research has leveraged Large Language Models (LLMs) to enhance Automated Feature Engineering (AutoFE) through semantic descriptions and trajectory-based prompting. However, there exist two challenges that limit their applicability and scalability in long-horizon optimization: (1) semantic metadata is unavailable in many practical settings, and (2) trajectory accumulation increases the risk of exceeding the context window, while without it, the generation process can become unstable, leading to becoming stuck in the local optima and a high duplicate rate of generated features. To this end, we propose a SHAP-enhanced Implicit-trajectory Generation for Metadata-free AutoFE (SIGMA), a scalable constant-context optimization framework. SIGMA leverages SHAP values to provide task-aware signals for guiding group feature generation instead of semantic information. In addition, we adopt an EXposed-feature Implicit Trajectory (EXIT) approach, where the exposed features in the prompt implicitly represent the trajectory. Empirical results demonstrate that SIGMA achieves performance comparable to the state-of-the-art (SOTA) LLM baselines with a nearly constant prompt length. Notably, EXIT significantly reduces the duplicate ratio of generated features from 37.2% to 6.8%. At the same time, SIGMA matches traditional SOTA performance with only 5.4 features on average, demonstrating substantial efficiency gains in feature utilization.

摘要:最近的研究利用大型語言模型(LLMs)通過語義描述和基於軌跡的提示來增強自動特徵工程(AutoFE)。然而,存在兩個挑戰限制了它們在長期優化中的適用性和可擴展性:(1)在許多實際環境中,語義元數據不可用,以及(2)軌跡累積增加了超出上下文窗口的風險,而如果沒有它,生成過程可能變得不穩定,導致陷入局部最優解和生成特徵的高重複率。為此,我們提出了一種增強SHAP的隱式軌跡生成方法,用於無元數據的自動特徵工程(SIGMA),這是一個可擴展的恆定上下文優化框架。SIGMA利用SHAP值提供任務感知信號,以指導群體特徵生成,而不是依賴語義信息。此外,我們採用了一種EXposed-feature隱式軌跡(EXIT)方法,其中提示中的暴露特徵隱式地代表了軌跡。實證結果表明,SIGMA在幾乎恆定的提示長度下達到了與最先進(SOTA)LLM基準相當的性能。值得注意的是,EXIT顯著將生成特徵的重複率從37.2%降低到6.8%。同時,SIGMA在平均僅使用5.4個特徵的情況下達到了傳統SOTA性能,顯示出特徵利用的顯著效率提升。

Procedural Content Metageneration via Program Search and Continual Abstraction Discovery

2608.17947v1 by Matthew Siper, Ahmed Khalifa, Julian Togelius

Large language models can generate executable programs, which makes it possible to search directly over procedural content generators rather than individual levels. We study this approach in Sokoban, Zelda, Dangerous Dave, and Lode Runner. Each run evolves complete Python generators through language-model mutation and crossover. We introduce Continual Abstraction Discovery, or CAD, which extracts reusable primitives from high-fitness programs into a run-specific helper module. A 2x2 experiment crosses CAD with access to a fixed hand-written domain API. The completed data set contains 160 complete runs, with at least ten 50-generation runs in every cell. CAD raises mean final best fitness in all eight domain and API comparisons. Across all CAD runs, learned libraries are adopted by most later programs and repeatedly rediscover validation, reachability, and structural utilities. These results support that discovering reusable primitives improves evolutionary program search for content generators.

摘要:大型語言模型可以生成可執行的程式,這使得可以直接在程序內容生成器上進行搜索,而不是單獨的關卡。我們在推箱子、薩爾達傳說、危險的戴夫和洛德奔跑者中研究這種方法。每次運行通過語言模型的突變和交叉演化出完整的 Python 生成器。我們引入了持續抽象發現(Continual Abstraction Discovery,簡稱 CAD),它從高適應度的程式中提取可重用的原語,形成特定於運行的輔助模組。一個 2x2 實驗將 CAD 與訪問固定的手寫領域 API 結合起來。完成的數據集包含 160 次完整運行,每個單元至少有十次 50 代的運行。CAD 在所有八個領域和 API 比較中提高了平均最終最佳適應度。在所有 CAD 運行中,學習到的庫被大多數後續程式採用,並重複發現驗證、可達性和結構性實用工具。這些結果支持發現可重用原語改善內容生成器的進化程式搜索。

Efficient RLVR Scheduling via Graph-Structured Online Difficulty Estimation

2608.17941v1 by Zhizhao Liu, Zhiliang Tian, Xi Wang, Zhihua Wen, Yihang Xiong, Zhiquan Lai, Dongsheng Li

Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models but relies on costly rollout exploration. Assigning the same exploration budget to samples with different difficulty levels is inefficient: easy samples may receive redundant rollouts, whereas difficult but learnable samples may receive too little exploration. Existing adaptive schedulers address this mismatch through curriculum-based sample selection or non-uniform rollout allocation based on estimated sample difficulty. However, obtaining reliable online difficulty estimates remains challenging: dedicated probing adds substantial generation overhead, whereas history-based estimators face a cold start with no initial observations and stale feedback, and typically ignore relations among samples. To address these limitations, we propose a plug-and-play graph-based online difficulty estimator that shares rollout feedback across related samples and continuously updates their difficulty estimates, mitigating cold start and staleness without dedicated probing. Specifically, we first construct a difficulty-aware sample graph based on semantic and reasoning similarities. Based on this graph, we introduce latent difficulty states and use a Potts prior to encourage neighboring samples to share the same state. We then employ a state-level Beta-Binomial model to aggregate the rollout outcomes associated with each state. Finally, we use an online mean-field variational algorithm to continuously update the latent-state assignments and state-level difficulty as new feedback arrives. Our framework can be integrated into sample-selection and rollout-allocation schedulers, enabling difficulty-adaptive exploration without dedicated probing. Experiments across multiple base models, RL schedulers, and benchmarks demonstrate that our framework achieves better performance.

摘要:強化學習與可驗證獎勵(RLVR)提升了大型語言模型的推理能力,但依賴於成本高昂的展開探索。將相同的探索預算分配給不同難度級別的樣本是低效的:簡單樣本可能會收到冗餘的展開,而難度較高但可學習的樣本可能會收到過少的探索。現有的自適應調度器通過基於課程的樣本選擇或根據預估樣本難度的非均勻展開分配來解決這一不匹配。然而,獲得可靠的在線難度估計仍然具有挑戰性:專門的探測增加了可觀的生成開銷,而基於歷史的估計器面臨著沒有初始觀察和過時反饋的冷啟動問題,並且通常忽略樣本之間的關係。為了解決這些限制,我們提出了一種即插即用的基於圖的在線難度估計器,該估計器在相關樣本之間共享展開反饋,並持續更新它們的難度估計,減輕冷啟動和過時問題,無需專門的探測。具體而言,我們首先根據語義和推理相似性構建一個難度感知樣本圖。基於這個圖,我們引入潛在的難度狀態,並使用Potts先驗來鼓勵相鄰樣本共享相同的狀態。然後,我們使用狀態級的Beta-Binomial模型來聚合與每個狀態相關的展開結果。最後,我們使用在線均場變分算法來持續更新潛在狀態分配和狀態級難度,隨著新反饋的到來。我們的框架可以集成到樣本選擇和展開分配調度器中,實現難度自適應探索,而無需專門的探測。在多個基礎模型、RL調度器和基準測試中的實驗表明,我們的框架實現了更好的性能。

Grading Needs a Rubric, Not Intelligence

2608.17938v1 by Jhen-Ke Lin

Small language models can grade open-ended examination answers as reliably as substantially more expensive models when they grade against an explicit rubric. We test this claim as the design principle behind any-to-bench: a frontier model reads source documents once, at ingestion, to extract each question and its rubric; lower-cost models then perform all repeated grading work. We evaluate six cost-efficient model configurations from two model families at three reasoning-effort levels. Each configuration answers 24 open-ended examination questions, and each also grades every answer sheet three times, yielding 3,456 per-question grades. Scores depend overwhelmingly on the answer being graded: answer identity explains 95.6% of score variance, whereas judge identity explains only 0.2%. Raising a writer's reasoning effort moves earned scores by as much as 0.143 of full marks, while raising a judge's reasoning effort moves assigned scores by at most 0.006. Six frontier-tier judges, added as a check, reproduce these scores and are no more reliable as a panel. Two ablations then decompose the rubric on the same questions and answers. Removing its criteria and levels while keeping the official answer changes nothing measurable. Removing the official answer as well collapses reliability (ICC 0.888 to 0.628), inflates scores, and makes judge reasoning effort matter again. The rubric is what decouples grading from judge intelligence, and within the rubric the official answer does nearly all the work. We find no evidence of length preference or same-family preference under rubric-anchored grading.

摘要:小型語言模型在根據明確的評分標準進行評分時,可以與成本高得多的模型一樣可靠地評分開放式考試答案。我們測試這一主張,作為 any-to-bench 的設計原則:前沿模型在攝取時讀取源文件一次,以提取每個問題及其評分標準;然後,成本較低的模型執行所有重複的評分工作。我們在三個推理努力水平上評估來自兩個模型系列的六種成本效益模型配置。每個配置回答 24 道開放式考試問題,並且每個配置還對每份答案進行三次評分,產生每個問題 3,456 次評分。分數在很大程度上取決於被評分的答案:答案身份解釋了 95.6% 的分數變異,而評判身份僅解釋了 0.2%。提高寫作者的推理努力可以使得獲得的分數提高最多 0.143 的滿分,而提高評判的推理努力則最多使分數提高 0.006。六位前沿級評判作為檢查,重現這些分數,且作為小組的可靠性並沒有提高。接下來的兩個消融實驗則在相同的問題和答案上分解評分標準。去除其標準和級別,同時保留官方答案,並不會改變可測量的結果。去除官方答案也會使可靠性崩潰(ICC 從 0.888 降至 0.628),使分數膨脹,並使評判的推理努力再次變得重要。評分標準是將評分與評判智力解耦的關鍵,而在評分標準內,官方答案幾乎承擔了所有的工作。我們沒有發現基於評分標準的評分中存在長度偏好或同家族偏好的證據。

EvoTS-Agent: A Self-Evolving LLM Agent for Financial Time Series Change Point Detection

2608.17933v1 by Lei Jiang, Ye Wei, Xinyu Xi, Jordan Langham-Lopez, Yifan Bao, Raad Khraishi, Yihao Ang, Anthony K. H. Tung, Lukasz Szpruch, Hao Ni

Financial time series exhibit non-stationary and heterogeneous statistical properties, making change-point detection challenging because no single unsupervised algorithm performs consistently across assets and market regimes. Conventional workflows consequently depend heavily on expert-driven model selection, feature design, and hyperparameter tuning, limiting their scalability and adaptability. We propose EvoTS-Agent, a validation-guided self-evolving LLM agent for autonomous financial time-series change-point detection. EvoTS-Agent first performs curated exploratory data analysis to characterize dataset properties and initialize candidate detection models. It then evolves executable experiment trajectories through three complementary operators: \textit{Revision} exploits the current best solution, \textit{Alternative Strategy} explores fundamentally different modeling directions when progress stagnates, and \textit{Recombination} synthesizes complementary evidence from high-performing trajectories. Validation feedback guides trajectory evolution throughout the search, enabling the agent to adapt its detection pipeline to the statistical characteristics of each dataset while preserving reliable optimization. Experiments across four benchmark datasets demonstrate that EvoTS-Agent consistently outperforms existing LLM-based agents while maintaining a 100\% execution success rate across all evaluated backbone LLMs.

摘要:金融時間序列展現出非平穩和異質的統計特性,使得變更點檢測變得具有挑戰性,因為沒有單一的無監督算法能在不同資產和市場狀態下持續表現良好。因此,傳統工作流程在很大程度上依賴專家驅動的模型選擇、特徵設計和超參數調整,這限制了它們的可擴展性和適應性。我們提出了EvoTS-Agent,一種基於驗證指導的自我演化LLM代理,用於自主金融時間序列變更點檢測。EvoTS-Agent首先執行精心策劃的探索性數據分析,以特徵化數據集特性並初始化候選檢測模型。然後,它通過三個互補的運算子來演化可執行的實驗軌跡:\textit{Revision}利用當前最佳解,\textit{Alternative Strategy}在進展停滯時探索根本不同的建模方向,\textit{Recombination}從高效能的軌跡中綜合互補證據。驗證反饋在整個搜索過程中指導軌跡演化,使得代理能夠根據每個數據集的統計特徵調整其檢測流程,同時保持可靠的優化。在四個基準數據集上的實驗表明,EvoTS-Agent始終超越現有的基於LLM的代理,同時在所有評估的主幹LLM中保持100%的執行成功率。

2608.17932v1 by Chainarong Amornbunchornvej

Groups routinely complete projects that no single member can plan, execute, or verify alone. We propose a formal model of this phenomenon, Collective Counterfactual Planning (CCP), in which the binding limitation on each agent is neither capability, knowledge, nor observability, but representational geometry: each agent perceives the state, conceives moves, consents to actions, and certifies goal requirements only through a projection onto an agent-specific subspace of a common task space. Four gates jointly determine whether a team can reach a conjunctive goal and legitimately recognize that it has done so: the exogenous implementation coalitions required to perform each action, together with three representational gates -- conception, consent, and task-relative verification qualification. We define the Collective Counterfactual Solvability (CCS) problem, separating geometric feasibility, executable attainment, and validated completion. The results expose a positive-negative duality. Iterated cross-agent relay can unlock a solution that no one-shot pooling of individual plans contains, but any goal requirement depending essentially on the subspace dark to the entire team is unverifiable and therefore not validly completable, even when the trajectory accidentally attains it. Memoryless and audited consent further constrain different objects -- action directions versus cumulative trajectory states -- and neither dominates the other. A four-step exhaustive horizon-bounded solvability scheme is sound and complete under exact representation of the relay closure; restricted implementations remain sound on returned plans but need not be complete. The model gives one geometry for sequential mutual enabling, competent execution of steps whose purpose is invisible to the executor, forced sub-teaming at expertise boundaries, and completion that cannot be validly declared.

摘要:團體經常完成單一成員無法獨自計劃、執行或驗證的項目。我們提出這一現象的正式模型,稱為集體反事實規劃(CCP),在這個模型中,每個代理的約束限制既不是能力、知識,也不是可觀察性,而是表徵幾何:每個代理僅通過投影到共同任務空間的代理特定子空間來感知狀態、構思行動、同意行為和認證目標要求。四個閘門共同決定一個團隊是否能夠達成聯合目標並合法地認識到它已經達成:執行每個行動所需的外生實施聯盟,以及三個表徵閘門——構思、同意和任務相對驗證資格。我們定義了集體反事實可解性(CCS)問題,將幾何可行性、可執行達成和驗證完成分開。結果揭示了一種正負對偶性。迭代的跨代理中繼可以解鎖一個單次個人計劃無法包含的解決方案,但任何本質上依賴於對整個團隊來說是黑暗的子空間的目標要求都是不可驗證的,因此無法有效完成,即使軌跡意外達成了它。無記憶和經審核的同意進一步限制了不同對象——行動方向與累積軌跡狀態——而且兩者不相互主導。一個四步的全面邊界可解性方案在中繼閉包的精確表徵下是健全且完整的;受限的實施在返回的計劃上仍然是健全的,但不必是完整的。該模型為順序相互啟用、執行目的對執行者不可見的步驟的能力執行、在專業邊界強制子團隊以及無法有效宣告的完成提供了一種幾何。

SpeechSense: A Paralinguistic-Focused Dataset for Fine-Grained Speech Sentiment Analysis

2608.17931v1 by Shicheng Ma, Wenqian Cui, Irwin King

Recent advances in AI have revolutionized speech processing, yet effective speech understanding requires discerning not just what is said, but how it is said. Speech Sentiment Analysis plays a critical role in decoding these paralinguistic cues for diverse real-world applications such as recruitment and customer service. However, existing Speech Sentiment Analysis research faces two primary limitations. First, dominant approaches rely on text-centric pipelines that cascade Automatic Speech Recognition with text analysis. This process inevitably discards essential acoustic features like prosody and tone, failing to capture attitudinal meanings in acoustically ambiguous utterances. Second, current benchmarks suffer from a mismatch in label granularity, prioritizing basic emotions (e.g., happy, sad) over the nuanced interpersonal stances (e.g., confident, impatient) necessary for social sensitivity. To address these limitations, we propose a novel dataset, SpeechSense, for fine-grained speech sentiment analysis. Specifically, we define a specialized 8-class taxonomy of interpersonal stances detectable primarily through prosodic cues beyond lexical content alone. We then construct a curated dataset based on this taxonomy, built from high-fidelity speech synthesis and rigorous human validation. Comprehensive experiments across multi-modal LLMs, text-only LLMs, and speech encoders demonstrate that models with acoustic access consistently outperform text-only baselines. These results empirically validate the primacy of acoustic cues in detecting subtle speaker attitudes, highlighting the necessity of SpeechSense. Dataset and supplementary materials are available at https://github.com/Sher13cked/SpeechSense.

摘要:最近在人工智慧方面的進展已經徹底改變了語音處理,但有效的語音理解不僅需要辨識所說的內容,還需要理解其表達方式。語音情感分析在解碼這些副語言線索方面扮演著關鍵角色,適用於招聘和客戶服務等多樣的現實應用。然而,現有的語音情感分析研究面臨兩個主要限制。首先,主流方法依賴於以文本為中心的流程,將自動語音識別與文本分析串聯起來。這一過程不可避免地忽略了諸如韻律和語調等重要的聲學特徵,未能捕捉聲學模糊表達中的態度意義。其次,當前的基準測試在標籤粒度上存在不匹配,優先考慮基本情感(例如,快樂、悲傷),而忽視了社交敏感性所需的細微人際立場(例如,自信、不耐煩)。為了解決這些限制,我們提出了一個新穎的數據集,SpeechSense,用於細粒度的語音情感分析。具體來說,我們定義了一個專門的8類人際立場分類法,主要通過韻律線索而非單純的詞彙內容來檢測。我們然後根據這一分類法構建了一個精心策劃的數據集,該數據集基於高保真語音合成和嚴謹的人類驗證。跨多模態大型語言模型、僅文本的大型語言模型和語音編碼器的全面實驗表明,具有聲學訪問的模型在性能上始終優於僅文本的基準。這些結果實證了聲學線索在檢測微妙說話者態度中的重要性,突顯了SpeechSense的必要性。數據集和補充材料可在 https://github.com/Sher13cked/SpeechSense 獲得。

Comparative Study of Out-of-the-Box Technology for Automatic Target Detection and Recognition

2608.17917v1 by Alma M. Liezenga, Lotte Nijskens, Henrik R. Baumann, Stefan Becker, Simon Bensberg, Niccolò Camarlinghi, Håvard R. Eiring, Alexander W. Johnsgaard, Tanel Liiv, Giuseppe Martino, Matteo Marturini, Matthias Rapp, Jan Erik van Woerden, Alexander Wolpert, Hugo J. Kuijf

Automatic Target Detection and Recognition (ATD/R) is critical for military decision support and (semi-)autonomous operations. Recent advances in object detection and artificial intelligence (AI) significantly boosted the potential performance of ATD/R. However, the scarcity of publicly available military datasets limits the application of these systems. As a solution, this paper explores the use of publicly available models and civilian datasets to achieve reasonable performance in military contexts. We benchmark several state-of-the-art models, including six iterations of the YOLO series and two variations on the DETR framework, on a newly acquired military relevant dataset. This dataset features military vehicles and challenging circumstances, including various degrees of occlusions and small targets. The out-of-the-box version of each model is validated alongside a version finetuned on the VisDrone dataset. This dataset features small objects, an Air-to-Ground (A2G) perspective and relevant classes, potentially generalizing to our military ATD/R task. We compare the performance of the models using mAP@0.5 and mAP@0.5:0.95, across A2G and Ground-to-Ground (G2G) perspective, target size and model size, giving insight into the real-time capabilities of models. Our main findings are: (1) bigger models outperform smaller models, (2) DETR-based models show promising results compared to the YOLO series,(3) fine-tuning models on an out-of-domain A2G dataset, improves their A2G performance and slightly improves their performance on small objects, but (4) all models still struggle with detecting small objects in an A2G scenario. We conclude that, despite recent advances in object detection, in-domain training is still crucial for creating capable ATD/R systems.

摘要:自動目標偵測與識別(ATD/R)對於軍事決策支持和(半)自主作業至關重要。最近在物體偵測和人工智慧(AI)方面的進展顯著提升了ATD/R的潛在性能。然而,公開可用的軍事數據集稀缺限制了這些系統的應用。作為解決方案,本文探討使用公開可用的模型和民用數據集,以在軍事環境中實現合理的性能。我們在新獲得的軍事相關數據集上基準測試了幾個最先進的模型,包括六個版本的YOLO系列和兩個DETR框架的變體。這個數據集包含軍事車輛和具有挑戰性的情況,包括各種程度的遮擋和小目標。每個模型的開箱即用版本與在VisDrone數據集上微調的版本一起進行驗證。這個數據集包含小物體、空對地(A2G)視角和相關類別,可能對我們的軍事ATD/R任務具有普遍性。我們使用mAP@0.5和mAP@0.5:0.95比較模型的性能,涵蓋A2G和地對地(G2G)視角、目標大小和模型大小,提供對模型實時能力的洞察。我們的主要發現是:(1)較大的模型表現優於較小的模型,(2)基於DETR的模型與YOLO系列相比顯示出有希望的結果,(3)在域外的A2G數據集上微調模型,提高了它們的A2G性能,並稍微改善了它們在小物體上的性能,但(4)所有模型在A2G場景中仍然難以偵測小物體。我們得出結論,儘管在物體偵測方面取得了最近的進展,域內訓練仍然對於創建能夠的ATD/R系統至關重要。

CABLE: Extending the Reach of Memory Retrieval via Complementary Antecedent-Based Linking and Expansion

2608.17911v1 by Zheling Tan, Jin Gao, Dequan Wang

As LLM agents operate across structured workflows and sessions, preserving long-term history does not ensure that later contexts can recover relevant evidence through a bounded memory interface. We study this evidence-reachability problem in long-term conversational memory, where retrieval still relies heavily on semantic similarity. This works well for topical recall, but it often misses earlier experiences, plans, or motivations that are semantically distant from the later events they help explain. Existing memory graphs provide cross-memory structure, yet links driven mainly by semantic overlap can duplicate what the host retriever already recovers. We argue that link construction should instead prioritize a sparse set of retriever-complementary associations. We present CABLE (Complementary Antecedent-Based Linking and Expansion), a plug-in augmentation that constructs links designed to extend the host retriever's direct semantic reach. For each new memory, CABLE generates antecedent-oriented queries, retrieves prior memories, subtracts candidates in the direct semantic neighborhood, and verifies the remainder before adding the accepted complementary associations into a sparse directed graph. At retrieval time, CABLE expands the host system's retrieved seeds along these links to surface implicit supporting evidence. We evaluate CABLE with A-MEM on LoCoMo and MA-LongMemEval, and further integrate it into SimpleMem and Mem0g on LoCoMo, using Qwen3.5-27B, DeepSeek-chat, and GPT-4o-mini. CABLE yields higher mean LLM-judge scores in every evaluated system-level setting, with the largest gains in categories where useful evidence is distributed across memories or sessions, including open-domain, multi-session, and preference-oriented questions. These results support prioritizing sparse, reasoning-relevant associations that complement rather than duplicate the host retriever.

摘要:隨著LLM代理在結構化工作流程和會話中運作,保存長期歷史並不保證後續上下文能通過有限的記憶介面恢復相關證據。我們研究這個在長期對話記憶中的證據可達性問題,其中檢索仍然在很大程度上依賴於語義相似性。這對於主題回憶來說運作良好,但它常常會錯過早期的經驗、計劃或動機,這些與後來幫助解釋的事件在語義上相距甚遠。現有的記憶圖提供了跨記憶結構,但主要由語義重疊驅動的鏈接可能會重複主檢索器已經恢復的內容。我們認為鏈接構建應優先考慮一組稀疏的檢索器互補關聯。我們提出CABLE(Complementary Antecedent-Based Linking and Expansion),這是一個插件增強,旨在構建鏈接,以擴展主檢索器的直接語義範圍。對於每個新記憶,CABLE生成以前因為導向的查詢,檢索先前的記憶,從直接語義鄰域中減去候選者,並在添加接受的互補關聯到稀疏有向圖之前驗證剩餘部分。在檢索時,CABLE沿著這些鏈接擴展主系統檢索的種子,以顯現隱含的支持證據。我們在LoCoMo和MA-LongMemEval上使用A-MEM評估CABLE,並進一步將其整合到LoCoMo上的SimpleMem和Mem0g中,使用Qwen3.5-27B、DeepSeek-chat和GPT-4o-mini。在每個評估的系統級設置中,CABLE在LLM評估者得分上均獲得更高的平均分數,在有用證據分佈於記憶或會話的類別中獲得最大的增益,包括開放域、多會話和偏好導向問題。這些結果支持優先考慮稀疏的、與推理相關的關聯,這些關聯互補而非重複主檢索器的功能。

AutoResearch: Insight In, Hallucination Out

2608.17906v1 by Yiming Ren, Xiang Liu, Qumeng Sun, Xiao Zhang, Jiahao Li, Haoyang Zhang, Junjie Wang

Autonomous research systems are increasingly capable of executing long research workflows, yet automation alone does not ensure that the resulting process remains scientifically grounded. We introduce AutoResearch, a two-stage system that connects Idea Generation with Idea Execution to address both how research ideas are formed and how they are reliably established through experimentation. In Idea Generation, AutoResearch continuously integrates emerging research signals with accumulated domain knowledge, identifies transferable mechanistic insights, and uses multi-model generation and cross-review to produce grounded, testable research plans. In Idea Execution, coordinated agents decompose these plans into experiments, iteratively implement and diagnose them, and employ independent evidence-based review before accepting research conclusions. Across representative settings in cross-modal retrieval, systems optimization, and benchmark-driven machine learning, AutoResearch turns generated ideas into measurable progress, detects and corrects unreliable experimental results, and makes evidence-conditioned decisions to continue, revise, or terminate research directions. For example, on RSICD benchmark, an AutoResearch-generated idea improves mean Recall from 32.84 to 34.69, while recording only 5 audit-confirmed issue events compared with 11-27 for other autonomous research systems. These results demonstrate a research process in which meaningful insight is grounded before experimentation and conclusions are grounded before acceptance: Insight In, Hallucination Out.

摘要:自主研究系統越來越能夠執行長期的研究工作流程,然而僅僅依賴自動化並不能確保所產生的過程保持科學基礎。我們介紹了 AutoResearch,一個兩階段的系統,將創意生成與創意執行連接起來,以解決研究想法是如何形成的,以及如何通過實驗可靠地建立這些想法。在創意生成階段,AutoResearch 持續整合新興的研究信號與累積的領域知識,識別可轉移的機制見解,並利用多模型生成和交叉審查來產出有根據、可測試的研究計劃。在創意執行階段,協調的代理將這些計劃分解為實驗,迭代實施和診斷它們,並在接受研究結論之前進行獨立的基於證據的審查。在跨模態檢索、系統優化和基準驅動的機器學習等代表性設置中,AutoResearch 將生成的想法轉化為可衡量的進展,檢測並修正不可靠的實驗結果,並做出基於證據的決策以繼續、修訂或終止研究方向。例如,在 RSICD 基準上,AutoResearch 生成的想法將平均召回率從 32.84 提高到 34.69,同時僅記錄了 5 次經審核確認的問題事件,而其他自主研究系統則記錄了 11-27 次。這些結果展示了一個研究過程,其中有意義的見解在實驗之前就已經建立,而結論在接受之前也已經有根據:見解進,幻覺出。

BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models

2608.17895v1 by Liubov Chubarova, Alexandra Kuleshova, Daniil Volkov, Kirill Sultanov, Alexey Zaytsev

While Multimodal Large Language Models (MLLMs) have made significant strides in visual comprehension, their ability to reason about text-dense, professional documents remains incompletely evaluated. Existing benchmarks emphasize information extraction, require external domain knowledge, or cover professional documents only as one of many settings. They are also largely English- or Chinese-centric, leaving other languages and Russian, in particular, substantially underrepresented. To address these limitations, we introduce BEAR-Bench (Bilingual Enterprise and Academic Reasoning), a self-contained, complex English-and-Russian benchmark comprising 1000 human-annotated questions based on text-rich business and scientific documents. We evaluate 16 proprietary and open-weight MLLMs, including Gemini 3.1 Pro and Qwen3.5-397B, on BEAR-Bench and observe clear headroom even for the strongest systems. Finally, we use the resulting model outputs to compare existing hallucination detection methods, evaluating not only how often models fail on BEAR-Bench but also how reliably those failures can be identified.

摘要:雖然多模態大型語言模型(MLLMs)在視覺理解方面取得了重大進展,但它們對於文本密集型的專業文件的推理能力仍未得到充分評估。現有的基準強調信息提取,需要外部領域知識,或者僅將專業文件作為眾多設置之一。這些基準在很大程度上以英語或中文為中心,使其他語言,特別是俄語,顯得大幅度不足。為了解決這些限制,我們推出了BEAR-Bench(雙語企業與學術推理),這是一個自包含的、複雜的英語和俄語基準,包含1000個基於文本豐富的商業和科學文件的人類標註問題。我們在BEAR-Bench上評估了16個專有和開放權重的MLLMs,包括Gemini 3.1 Pro和Qwen3.5-397B,並觀察到即使對於最強的系統也存在明顯的提升空間。最後,我們使用生成的模型輸出來比較現有的幻覺檢測方法,不僅評估模型在BEAR-Bench上的失敗頻率,還評估這些失敗能否被可靠地識別。

BayesPrompt: human readable prompts that make sense

2608.17866v1 by Franky Kevin Nando Tezoh, Ali Hussaini Umar, Alessandro Laio, Guido Sanguinetti, Riccardo Rende

Reconstructing prompts that can elicit a desired answer or behaviour in an LLM is an open and important research topic. Optimisation methods which aim at minimising the perplexity of a given answer, however, consistently yield so-called pseudoprompts, unintelligible strings of tokens which can lack human interpretability. We argue that this is a consequence of the ill-posedness of the prompt optimisation task. By reframing the task as a Bayesian posterior inference over prompts, we propose an efficient algorithm to sample prompts which are both efficient (in terms of perplexity) and human readable. We compare our approach with state of the art alternatives showing on a real data set a marked improvement over a range of metrics.

摘要:重建能夠引發大型語言模型(LLM)所需答案或行為的提示是一個開放且重要的研究主題。然而,旨在最小化給定答案困惑度的優化方法,卻持續產生所謂的偽提示,即無法理解的標記字符串,這些字符串可能缺乏人類可解釋性。我們認為這是提示優化任務不良定義的結果。通過將任務重新構架為對提示的貝葉斯後驗推斷,我們提出了一種有效的算法來抽樣既高效(在困惑度方面)又人類可讀的提示。我們將我們的方法與最先進的替代方案進行比較,顯示在一個真實數據集上,在多個指標上有顯著的改善。

ARASH: Adaptive Retrieval And Shot Selection for Tabular Prediction

2608.17856v1 by Samirasadat Jamalidinan, Yue Xu, Kazem Cheshmi

Tabular prediction is a critical task across numerous applications. The recent success of large language models has sparked various approaches for adapting them to the tabular domain. A prevalent strategy involves training or fine-tuning specialized Tabular Foundation Models (TFMs) such as TabPFN. However, TFMs require substantial computational resources, and frequent model retraining is often impractical. In-context learning (ICL), specifically, few-shot prompting, offers a resource-efficient alternative to enhance performance. Yet, identifying the most relevant rows to serve as shots remains a challenge for tabular data. This paper introduces ARASH (Adaptive, query-specific Retrieval And Shot selection), a method that improves TFM efficiency by selecting optimal shots based on local neighborhood analysis within the training set. Our results demonstrate that ARASH reduces the prompt length and memory usage of TabPFN by 1261.5$\times$ and 2.56$\times$, respectively, while providing comparable accuracy.

摘要:表格預測是許多應用中的一項關鍵任務。大型語言模型的近期成功激發了各種將它們適應於表格領域的方法。一種普遍的策略涉及訓練或微調專門的表格基礎模型(TFMs),如TabPFN。然而,TFMs需要大量的計算資源,並且頻繁的模型重訓練往往不切實際。在上下文學習(ICL)中,特別是少量樣本提示,提供了一種資源高效的替代方案來提升性能。然而,識別最相關的行作為樣本仍然是表格數據的一個挑戰。本文介紹了ARASH(自適應、查詢特定的檢索和樣本選擇),這是一種通過基於訓練集內的局部鄰域分析選擇最佳樣本來提高TFM效率的方法。我們的結果顯示,ARASH分別將TabPFN的提示長度和內存使用量減少了1261.5$\times$和2.56$\times$,同時提供了可比的準確性。

Encoded but Not Actionable: Auditing the Decode-Generate-Steer Gap in Frozen LLMs for Geometric Constraints

2608.17843v1 by Man Liang, Xinzhao Cheng, Faizan Wajid

Large language models (LLMs) have demonstrated strong performance on structured reasoning tasks, but what they encode and whether it informs model behavior remain unclear. We investigate this question through geometric reasoning, using parametric CAD constraints as a controlled testbed for separating local pairwise relations from sketch-level constraint status. By probing the hidden states of six frozen decoder-only LLMs, we examine four properties: linear decodability, forced-choice generation, activation-level influence, and behavioral steerability. Pretraining substantially improves the decoding of local geometric relations, and this advantage persists after accounting for positional cues with shuffled-order controls. In contrast, sketch-level DOF status is already highly decodable from randomly initialized representations and improves only modestly with pretraining, indicating that much of its probe performance is available without learned weights. Further analyses show that decodable information is not always actionable. Generation often fails to express this information, and on the two intervention-tested backbones, activation-restoration effects at the patched entity position vanish while decodability persists across depth. Mean-difference steering also does not reliably control outputs. These results show that decodability, generation, activation-level influence, and steerability can diverge in the tested setting. The audit provides a controlled way to distinguish failures to encode geometric structure from failures to express or control encoded information.

摘要:大型語言模型(LLMs)在結構推理任務中表現出色,但它們編碼了什麼以及這是否影響模型行為仍不清楚。我們通過幾何推理來研究這個問題,使用參數化CAD約束作為控制測試平台,以區分局部成對關係和草圖級約束狀態。通過探測六個凍結的僅解碼器LLMs的隱藏狀態,我們檢查了四個特性:線性可解碼性、強制選擇生成、激活水平影響和行為可引導性。預訓練顯著改善了局部幾何關係的解碼,並且在考慮到隨機順序控制的位置信息後,這一優勢仍然存在。相比之下,草圖級DOF狀態已經可以從隨機初始化的表示中高度可解碼,並且在預訓練後僅有適度改善,這表明其探測性能在沒有學習權重的情況下就已經可用。進一步分析顯示,可解碼的信息並不總是可操作的。生成通常未能表達這種信息,在兩個經過干預測試的骨幹上,修補實體位置的激活恢復效應消失,而可解碼性在深度上仍然存在。均值差異引導也無法可靠地控制輸出。這些結果顯示,在測試環境中,可解碼性、生成、激活水平影響和可引導性可能會出現分歧。這次審核提供了一種控制方式,以區分編碼幾何結構的失敗與表達或控制編碼信息的失敗。

AdaLens: Interactive Storyline for Monitoring and Steering Long-Running Agentic Data Analysis

2608.17834v1 by Yangtian Liu, Yan Miao, Shuhan Liu, Yunfan Zhou, Dae Hyun Kim, Di Weng, Yingcai Wu

Large language models are pushing data science toward increasingly autonomous and agentic workflows, with recent systems already supporting multi-step and long-running analyses. As these workflows become more autonomous, conventional interfaces no longer provide adequate support for two critical requirements: observability for understanding an agent's evolving reasoning and evidence, and steerability for redirecting low-value directions or deepening promising ones during execution. Existing interactive approaches improve process visibility and open intervention points, but they remain largely designed for discrete, turn-by-turn exchanges rather than the parallel branches and evolving decision structures of long-running agentic analysis. We study this need as interactive oversight in long-running agentic data analysis and present AdaLens, an interactive system for monitoring and steering ongoing runs. AdaLens combines a storyline-based representation that unifies analytical plans, execution progress, intermediate findings, and data-column involvement with steering interactions grounded in these analytical elements for directional guidance and execution control. We evaluate AdaLens through two case studies and a user study, examining how it supports analysts in monitoring and steering long-running agentic data analysis.

摘要:大型語言模型正在推動數據科學朝向越來越自主和具代理性的工作流程,最近的系統已經支持多步驟和長期運行的分析。隨著這些工作流程變得更加自主,傳統界面不再能夠充分支持兩個關鍵需求:可觀察性以理解代理人不斷演變的推理和證據,以及可引導性以在執行過程中重新定向低價值的方向或加深有前景的方向。現有的互動方法改善了過程的可見性並開放了干預點,但它們主要是為了離散的、逐步的交流而設計,而不是針對長期運行的代理分析中的平行分支和不斷演變的決策結構。我們研究這一需求作為長期運行的代理數據分析中的互動監督,並提出了AdaLens,一個用於監控和引導正在進行的運行的互動系統。AdaLens結合了一種基於故事情節的表示,統一了分析計劃、執行進度、中間發現和數據列參與,並基於這些分析元素提供引導互動,以實現方向指引和執行控制。我們通過兩個案例研究和一項用戶研究來評估AdaLens,檢視它如何支持分析師監控和引導長期運行的代理數據分析。

The Model's Tell: Measuring Context-Leakage Attack Signals with Behavior Gauges

2608.17829v1 by Maosen Zhang, Jianshuo Dong, Boting Lu, Wenyue Li, Xiaoping Zhang, Tianwei Zhang, Jie Zhang, Han Qiu

LLMs increasingly rely on external contexts, such as pre-defined system prompts or retrieved documents, to improve generation quality. However, processing these contexts alongside user queries creates an attack surface: adversarial inputs can induce models to disclose them. Prior probing studies suggest that leakage-related signals emerge in hidden states, yet the need to extract these states poses additional deployment challenges. In this paper, we explore whether this internal signal leaves a more accessible ``tell'' before decoding. We propose LeakGauge, which probes this response by appending a suffix that gauges leakage behavior and mapping its prefill token probabilities to an attack-risk score. While a direct gauge uses the initial tokens of confidential content, we find that a content-agnostic one that verbalizes leakage behavior yields more robust signals. Across 11 LLMs, including GLM-5.2 (753B) and Kimi-K3 (2.8T), LeakGauge reaches an AUROC range of 0.944--0.996 on unseen attacks. The signal remains stable when the content changes language or the attack shifts from verbatim to semantic disclosure. By activation-steering interventions, we further show that the risk score is sensitive to an internal leakage-related direction, relating the observable signal to the model's internal representation. In addition, LeakGauge enables an input detector with fewer than 0.5K extra parameters and added latency of 10.34 ms. Code: \href{https://github.com/yeasen-z/LeakGauge}.

摘要:LLM越來越依賴外部上下文,例如預定義的系統提示或檢索的文件,以提高生成質量。然而,將這些上下文與用戶查詢一起處理會創造攻擊面:對抗性輸入可能會誘使模型洩露它們。先前的探測研究表明,與洩漏相關的信號在隱藏狀態中出現,但提取這些狀態的需求帶來了額外的部署挑戰。在本文中,我們探討這個內部信號是否在解碼之前留下更易於訪問的“告訴”。我們提出了LeakGauge,它通過附加一個後綴來探測這個反應,以評估洩漏行為並將其預填令牌的概率映射到攻擊風險分數。雖然直接的評估使用了機密內容的初始令牌,但我們發現一個與內容無關的評估,能夠表達洩漏行為,產生更穩健的信號。在11個LLM中,包括GLM-5.2 (753B)和Kimi-K3 (2.8T),LeakGauge在未見過的攻擊上達到了0.944到0.996的AUROC範圍。當內容變更語言或攻擊從逐字披露轉變為語義披露時,信號仍然穩定。通過激活引導干預,我們進一步顯示風險分數對內部洩漏相關方向敏感,將可觀察信號與模型的內部表示相關聯。此外,LeakGauge使得輸入檢測器的額外參數少於0.5K,並增加了10.34毫秒的延遲。代碼:\href{https://github.com/yeasen-z/LeakGauge}。

From Global Benchmarks to Local Evaluations: Benchmarking LLMs for the German Public Sector

2608.17827v1 by Camilla Dalerci, Thilo Michael, Robin Schaefer, Daniel Weinland

Public institutions face a persistent challenge in selecting LLMs suited to their specific context. Existing benchmarks, however, are of limited use as they primarily reflect English-language and US-centric settings, and often only evaluate task performance. In this paper, we present first results of MÖVE, a holistic evaluation framework for the German public sector, examining three rarely considered governance dimensions: energy consumption, provider transparency, and knowledge of German-party positions. Our results reveal significant trade-offs, with no single model excelling across all dimensions: estimated energy consumption varies more than 60-fold and is not explained by model size alone, information disclosure varies systematically across providers, and European models do not exhibit stronger knowledge of German party positions. Model selection for public institutions thus cannot rely on performance rankings alone. Instead, evaluations should also reflect the governance requirements of the deployment context.

摘要:公共機構在選擇適合其特定情境的LLM時面臨持續的挑戰。然而,現有的基準測試用途有限,因為它們主要反映英語和美國中心的環境,且通常僅評估任務表現。在本文中,我們呈現MÖVE的初步結果,這是一個針對德國公共部門的整體評估框架,檢視三個鮮少考慮的治理維度:能源消耗、供應商透明度和對德國政黨立場的了解。我們的結果揭示了顯著的權衡,沒有單一模型在所有維度上表現優異:估計的能源消耗變化超過60倍,且僅以模型大小無法解釋,信息披露在不同供應商之間系統性變化,歐洲模型對德國政黨立場的了解並未顯示出更強的優勢。因此,公共機構的模型選擇不能僅依賴於性能排名。相反,評估還應反映部署情境的治理要求。

MotoSafety: Edge-AI with Learned Temporal Importance for Two-Wheeler Collision Risk Assessment Under Time Pressure

2608.17823v1 by Sumit S. Shevtekar, Chandresh K. Maurya, Gourab Sil, Subasish Das

Powered two-wheeler riders face critical safety challenges in low- and middle-income countries, yet limited studies exist on how cognitive stressors such as Time Pressure influence collision risk. To address this gap, we introduce a large-scale dataset of over 129,000 labeled multivariate time-series sequences from 153 simulator rides by 51 participants under No, Low, and High TP, capturing 64 features across vehicle dynamics, control inputs, proximity, and behavioral violations. Building on this dataset, we propose MotoSafety, a novel edge-AI architecture grounded in the Learned Temporal Importance principle. MotoSafety achieves 94.97% accuracy and 99.33% ROC AUC, outperforming ten baselines, including TimesNet and LLM4TS, and achieves 0.039 MSE and 0.094 MAE for forecasting (4.4x lower error than Time-LLM and iTransformer). With only 1.15M parameters and 0.135 ms latency, it is suitable for edge deployment on low-cost CPU hardware. Using ground truth TP as an inductive bias improves accuracy from 94.09% to 94.97%, while predicted TP achieves 94.82%. Using only 21 IMU+GPS features, it achieves 93.91% accuracy, indicating practical deployment. Beyond PTW safety, the architecture shows better transferability to human activity (97.66%) and clinical (99.65%) domains. This lightweight framework advances PTW collision risk assessment, supporting the Safe System Approach for Intelligent Transportation Systems.

摘要:在中低收入國家,動力二輪車騎士面臨著重大的安全挑戰,但關於認知壓力因素如時間壓力如何影響碰撞風險的研究卻相對有限。為了填補這一空白,我們引入了一個大規模數據集,該數據集包含來自51名參與者在無時間壓力、低時間壓力和高時間壓力下進行的153次模擬騎行的超過129,000個標記的多變量時間序列,捕捉了64個特徵,涵蓋了車輛動態、控制輸入、接近度和行為違規。基於這個數據集,我們提出了MotoSafety,一種基於學習時間重要性原則的新型邊緣人工智慧架構。MotoSafety實現了94.97%的準確率和99.33%的ROC AUC,超越了包括TimesNet和LLM4TS在內的十個基準,並在預測中達到了0.039的均方誤差和0.094的平均絕對誤差(比Time-LLM和iTransformer低4.4倍)。它僅需1.15M的參數和0.135毫秒的延遲,適合在低成本CPU硬體上進行邊緣部署。使用真實的時間壓力作為歸納偏見,準確率從94.09%提高到94.97%,而預測的時間壓力則達到94.82%。僅使用21個IMU+GPS特徵,它的準確率達到93.91%,顯示出實際部署的潛力。除了PTW安全性外,該架構在人體活動(97.66%)和臨床(99.65%)領域也顯示出更好的可轉移性。這個輕量級框架推進了PTW碰撞風險評估,支持智能交通系統的安全系統方法。

Interpretable Humans, Alien LLMs: Expert Analysis of Latent Structures in Assessment Responses

2608.17810v1 by Alona Strugatski, Licol Zeinfeld, Jason Cooper, Shelley Rap, Gil Schwarts, Giora Alexandron

The evaluation of large language models (LLMs) relies heavily on human-designed assessments, implicitly assuming that AI and humans employ similar underlying cognitive constructs. Challenging this assumption, we investigate whether the latent factors governing LLM performance carry the same substantive, human-interpretable meaning as the cognitive constructs governing human learners. Using responses from humans and six LLMs across quantitative reasoning and chemistry assessments, we conducted Exploratory Factor Analysis (EFA) separately for both groups. Subject-Matter Experts (SMEs) then blindly evaluated the resulting factor graphs to ascribe pedagogical meaning to the emerged constructs. SMEs successfully interpreted most of the human-derived factors. Conversely, they could not ascribe meaning to any LLM-derived factors in quantitative reasoning and interpreted only half of the LLM factors in chemistry. By combining data-driven EFA with blind expert interpretation, this framework shows that LLMs frequently operate on statistically opaque mechanisms distinct from human reasoning.

摘要:大型語言模型(LLMs)的評估在很大程度上依賴於人類設計的評估,隱含假設AI和人類使用相似的基本認知結構。挑戰這一假設,我們調查了支配LLM性能的潛在因素是否具有與支配人類學習者的認知結構相同的實質性、人類可解釋的意義。利用來自人類和六個LLM在定量推理和化學評估中的反應,我們分別對這兩組進行了探索性因素分析(EFA)。主題專家(SMEs)隨後盲目評估了所產生的因素圖,以賦予出現的結構教學意義。SMEs成功解釋了大多數人類衍生的因素。相反,他們無法為任何LLM衍生的因素在定量推理中賦予意義,並且只解釋了化學中一半的LLM因素。通過將數據驅動的EFA與盲專家解釋相結合,這一框架顯示LLMs經常在與人類推理不同的統計不透明機制上運作。

Whether LLMs Can Navigate Beliefs and Facts Depends on How You Phrase It

2608.17809v1 by Quang Minh Nguyen, Luis Frentzen Salim

Humans naturally form and express beliefs in daily communication, e.g., "I think the answer is 3" or "I suppose that's right." Such beliefs inevitably intertwine with fact and knowledge, making the ability to handle them in tandem desirable for large language models (LLMs), as they are increasingly deployed in user-facing settings. Prior work showed that even capable LLMs exhibit a systemic weakness in acknowledging user beliefs grounded in incorrect information. We extend this evaluation to 10 LLMs across 18 epistemic expressions and find that the size and direction of the weakness depend on the verb used to express the belief, with the accuracy gap between factual and false information ranging from +50% on "I vaguely remember" to -14% on "I seriously doubt". We further show that the phenomenon stems from task confusion: models default to fact-checking the underlying claim, overriding the user's stated belief; chains of thought that explicitly fact-check show lower accuracy on false information than those that do not; and a single instruction can reverse the failure across verb families. Mechanistically, models attend more to false beliefs they fail to confirm, but suppressing this attention at decoding time recovers accuracy only partially and only in some models, calling for future work on intervention methods. Our findings clarify prior results and show how fact-checking, a generally desirable behavior, can interfere with belief tracking in LLMs. Our code is available at https://github.com/ngqm/belief-fact-phrasing.

摘要:人類在日常交流中自然地形成和表達信念,例如「我認為答案是3」或「我想這是對的」。這些信念不可避免地與事實和知識交織在一起,使得同時處理它們的能力對大型語言模型(LLMs)來說變得可取,因為它們在面向用戶的環境中越來越多地被部署。先前的研究顯示,即使是能幹的LLMs在承認基於錯誤信息的用戶信念方面也存在系統性的弱點。我們將這一評估擴展到18種認識表達下的10個LLMs,發現弱點的大小和方向取決於用來表達信念的動詞,事實信息與虛假信息之間的準確性差距從「我模糊地記得」的+50%到「我嚴重懷疑」的-14%不等。我們進一步表明,這一現象源於任務混淆:模型默認檢查基礎主張的事實,覆蓋用戶所表達的信念;明確進行事實檢查的思維鏈在虛假信息上的準確性低於那些不進行檢查的;而單一指令可以逆轉動詞家族中的失敗。在機制上,模型對它們未能確認的虛假信念的注意力更高,但在解碼時抑制這種注意力僅能部分恢復準確性,且僅在某些模型中有效,這呼籲未來對干預方法的研究。我們的發現澄清了先前的結果,並顯示事實檢查這一通常可取的行為如何干擾LLMs中的信念追蹤。我們的代碼可在 https://github.com/ngqm/belief-fact-phrasing 獲得。

An Empirical Study of Reward Specification and Benchmark Reliability in GRPO-based LLM Unlearning

2608.17804v1 by Rubén Balbastre, Juan Manuel Orduña, Mariano Pérez

Practical LLM unlearning is usually evaluated through two objectives: suppress target-specific knowledge and preserve non-target utility. In generative QA, this leaves a third behavior underspecified: when a target-adjacent prompt admits a broader answer without target-specific leakage, the model should answer at that level rather than leak, evade, or refuse. We study this specification problem in a controlled LoRA-GRPO RWKU setting, comparing four reward designs that span lexical suppression, anti-refusal shaping, rubric-based broad answering, and an explicit refusal contrast, with and without SFT warm-up. The experiments show that optimization success is not equivalent to behavioral unlearning: RWKU forget scores, held-out completion audits, terminal training-rollout audits, and training dynamics can point to different conclusions. We trace these disagreements to reward-hacking endpoints, policy-support limits in GRPO, benchmark probes that miss endpoint changes, and rewards that can select broad-topic answering with low semantic leakage during optimization.

摘要:實際的 LLM 忘記通常通過兩個目標來評估:抑制特定目標的知識和保留非目標的效用。在生成性問答中,這留下了第三種行為未明確規範:當一個與目標相近的提示允許更廣泛的回答而不泄露特定目標時,模型應該在該層次上回答,而不是泄露、逃避或拒絕。我們在一個受控的 LoRA-GRPO RWKU 設定中研究這個規範問題,比較四種獎勵設計,涵蓋詞彙抑制、反拒絕塑造、基於標準的廣泛回答以及明確的拒絕對比,並且有無 SFT 熱身。實驗表明,優化成功並不等同於行為上的忘記:RWKU 忘記分數、保留的完成審計、終端訓練回滾審計和訓練動態可能指向不同的結論。我們將這些分歧追溯到獎勵駭客端點、GRPO 中的政策支持限制、錯過端點變化的基準探針,以及在優化過程中可以選擇廣泛主題回答且語義泄露低的獎勵。

StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows

2608.17800v1 by Liya Zhu, Xin Ma, Tao Liu, Haodong Wang, Ge Zhang, Jingzhe Ding, Qingshui Gu, Yongjie Zhong, Jinxiang Meng, Yuan Gao, Yunqiu Zhou, Hao Zhu, Jifeng He, Yongzhi Liao, Xinyi Zhang, Chaoxin Li, Yi Zhu, Xi Lin, Duju Zeng, Xiang Gao, Wen Zhang, Yunyang Wang, Duo Wang, Huan Zhou, Zuo Wang, Jin Chen, Kaiyuan Zhang, Chuqian Yu, Tianhao Yu, Longxiang Liu, Jianbo Xue, Huimin Che, Jiahao Wang, Yujia Qin, Jiaheng Liu, Shen Yan, Xiaolong Chang, Wenhao Huang

Recent advances in Large Language Models(LLMs) and agents have substantially improved the ability of AI systems to execute complex tasks. Yet existing benchmarks largely rely on researcher-selected tasks, leaving uncertain whether such progress extends to the work that real-world users actually demand from AI systems. We introduce \textbf{StartupBench}, an E2E agent benchmark grounded in market-validated AI startup products. Rather than defining tasks from pre-defined assumptions about useful agent capabilities, we systematically study AI products with demonstrated adoption, together with their product workflows and users, to identify real-world tasks for which AI has established practical demand across diverse professional domains. We translate these workflows into complete deliverable-oriented tasks and evaluate them with fine-grained rubrics capturing their complex requirements. Across representative models evaluated under a unified agent harness, even the strongest model successfully completes only approximately 30\% of StartupBench, despite making substantial partial progress on many tasks. Further analysis identifies aspects like complex instruction following and domain-specific expertise as major sources of failure. Our results reveal that many market-validated workflows remain beyond the reliable capabilities of current general-purpose agents, establishing StartupBench as an empirical measure of progress toward E2E completions of real-world user tasks.

摘要:最近在大型語言模型(LLMs)和代理方面的進展顯著提高了 AI 系統執行複雜任務的能力。然而,現有的基準主要依賴研究者選擇的任務,這使得不確定這樣的進展是否延伸到現實世界用戶對 AI 系統的實際需求。我們介紹了 \textbf{StartupBench},這是一個基於市場驗證的 AI 初創產品的端到端代理基準。我們不是從對有用代理能力的預定假設中定義任務,而是系統性地研究已經被採用的 AI 產品,連同它們的產品工作流程和用戶,以識別 AI 在各種專業領域中已建立的實際需求的任務。我們將這些工作流程轉化為完整的交付導向任務,並使用細緻的評分標準來評估它們,捕捉其複雜的要求。在統一的代理框架下評估的代表性模型中,即使是最強的模型也僅成功完成約 30\% 的 StartupBench,儘管在許多任務上取得了顯著的部分進展。進一步分析確定了複雜指令遵循和特定領域專業知識等方面是主要的失敗來源。我們的結果顯示,許多市場驗證的工作流程仍超出當前通用代理的可靠能力,確立了 StartupBench 作為向現實世界用戶任務的端到端完成進展的實證衡量標準。

Training with synthetic data for drone detection in thermal imagery

2608.17799v1 by Tanel Liiv, Sander Soodla, Nzamba Bignoumba, Alma M. Liezenga, Toomas Pruuden

Ground-to-Air (G2A) drone detection in medium- and long-wave infrared (MWIR/LWIR) imagery is challenging due to reduced texture information, sensor noise, weak thermal contrast, and the scarcity of annotated data. This work investigates a synthetic-first training strategy that combines synthetic scene generation with fine-tuning on real data. We show that synthetic data provides an effective basis for learning initial object representations, while real in-domain thermal imagery is still essential for reliable deployment. Even small amounts of real IR data substantially reduce domain gaps. Our experiments indicate that dataset alignment has a stronger impact on performance than model scale. Finally, our analysis of the dataset suggests that semantic alignment in feature space is the strongest predictor of model performance, while radiometric properties such as entropy and dynamic range also contribute to detection robustness. This work provides a foundation for combining synthetic and real IR data for effective G2A drone detection.

摘要:地面對空 (G2A) 無人機在中波和長波紅外 (MWIR/LWIR) 影像中的檢測具有挑戰性,這是因為紋理資訊減少、傳感器噪聲、熱對比度弱以及標註數據的稀缺。本研究探討了一種合成優先的訓練策略,將合成場景生成與真實數據的微調相結合。我們顯示合成數據為學習初始物體表示提供了有效的基礎,而真實的域內熱影像仍然對可靠的部署至關重要。即使是少量的真實紅外數據也能顯著減少域間差距。我們的實驗表明,數據集對齊對性能的影響比模型規模更強。最後,我們對數據集的分析表明,特徵空間中的語義對齊是模型性能最強的預測指標,而熵和動態範圍等輻射特性也有助於檢測的穩健性。本研究為有效的 G2A 無人機檢測結合合成和真實紅外數據提供了基礎。

TraceSQL: Traceable Answerability Estimation for Reference-Free Text-to-SQL Verification

2608.17795v1 by Neelesh Kumar Shukla, Debasmita Panda, Srutanik Bhaduri, Aditya Banerjee, Viji Krishnamurthy

Text-to-SQL systems are commonly evaluated using ground-truth SQL queries or reference execution results, but such supervision is unavailable at inference time in real-world deployments. This creates a critical verification problem: given only a user question, database context, and generated SQL, can a system estimate whether the generated query is likely to correctly answer the question? Recent approaches use LLMs as judge or specialized agents to inspect generated SQL, but their decisions can be difficult to trace. Outcome Reward Models (ORMs) address this by learning from execution-labeled candidate SQLs and assigning correctness scores to unseen queries, yet they still provide limited visibility into the signals behind each verification. To address this limitation, we propose TraceSQL, a lightweight and traceable verification model built on explicit diagnostic features. TraceSQL combines 67 features capturing question ambiguity, question requirements, question-schema-SQL consistency, SQL structure, and intent alignment. These signals remain available for examining which factors influence each prediction and for tracing decisions back to diagnostic evidence. On BIRD development databases, TraceSQL achieves 66.47% F1 and 64.48% ROC-AUC, compared with 61.87% F1 and 58.26% ROC-AUC for the GradeSQL-7B ORM baseline on the same generated-SQL evaluation. Feature attribution further shows that the model relies on both semantic grounding and deterministic SQL-structure signals. These results show that SQL verification can be performed with a lightweight learned model while retaining feature-level evidence for inspecting and diagnosing its predictions.

摘要:文本到 SQL 的系統通常使用真實的 SQL 查詢或參考執行結果來進行評估,但在現實世界的部署中,這種監督在推理時是不可用的。這創造了一個關鍵的驗證問題:僅根據用戶問題、數據庫上下文和生成的 SQL,系統能否估計生成的查詢是否可能正確回答問題?最近的方法使用大型語言模型(LLMs)作為評判或專門代理來檢查生成的 SQL,但它們的決策可能難以追蹤。結果獎勵模型(ORMs)通過從執行標記的候選 SQL 中學習並為未見查詢分配正確性分數來解決這個問題,但它們仍然提供有限的可見性來了解每個驗證背後的信號。為了解決這一限制,我們提出了 TraceSQL,一個基於明確診斷特徵的輕量級且可追蹤的驗證模型。TraceSQL 結合了 67 個特徵,捕捉問題模糊性、問題要求、問題-模式-SQL 一致性、SQL 結構和意圖對齊。這些信號仍然可用於檢查影響每個預測的因素,並將決策追溯到診斷證據。在 BIRD 開發數據庫上,TraceSQL 在相同生成 SQL 評估中達到 66.47% 的 F1 和 64.48% 的 ROC-AUC,而 GradeSQL-7B ORM 基準的 F1 為 61.87% 和 ROC-AUC 為 58.26%。特徵歸因進一步顯示,該模型依賴於語義基礎和確定性 SQL 結構信號。這些結果表明,SQL 驗證可以使用輕量級的學習模型來進行,同時保留特徵級的證據以檢查和診斷其預測。

Preference Is Not Intervention: The Structure and Stability Boundaries of Reader-Specific Evidence Utility

2608.17781v1 by Shi Zhou

ML systems increasingly condition decisions on downstream model identity, but this is useful only if model-specific differences form reusable structure rather than input-local interactions. We test this in retrieval-augmented generation (RAG), where evidence utility can be measured under controlled interventions. Holding query, evidence, task, scoring, and intervention fixed, nine readers disagree on effect sign in 33\% of jointly affected cells; reader$\times$query interaction explains 29.8\% of utility variance versus an 8.4\% permutation null; and self-selected evidence improves F1 by $+0.031$ ($t=3.39$). We then ask the sharper question: \emph{which components of this heterogeneity are stable reader properties across queries?} Separating three measurable objects---evidence \emph{activity}, \emph{ordinal preference}, and \emph{conditional signed direction}---we find ordinal reader geometry stable across four independent settings (split-half $ρ=0.60$--$0.83$): leave-one-out interventions, PRISM preferences, RAMDocs, and RAGuard. Signed geometry is task-bounded: weak in open-ended QA (0.14, 0.35), especially for misleading and irrelevant evidence, but strong in binary fact-checking (0.75) with no significant ordinal gap, though still below its sparsity-matched ceiling. Sparsity, decoding noise, and metric artifacts do not explain the main ordinal--signed gap. Finally, stable ordinal similarity fails to predict cross-reader intervention transfer (oracle-distance $ρ=-0.27$; regret reliability $-0.28$). Reader-specific utility exists, but preference is not intervention: stable ranking similarity does not license transfer of help/harm decisions.

摘要:ML 系統越來越多地根據下游模型的身份來決策,但這只有在模型特定的差異形成可重用結構而不是輸入局部交互時才有用。我們在檢索增強生成(RAG)中測試這一點,在這裡證據的效用可以在控制干預下進行測量。在查詢、證據、任務、評分和干預固定的情況下,九位讀者在 33\% 的共同影響單元中對效果符號存在分歧;讀者$\times$查詢交互解釋了 29.8\% 的效用變異,相較於 8.4\% 的置換虛無;自選證據使 F1 提高了 $+0.031$ ($t=3.39$)。然後我們提出更尖銳的問題:\emph{這種異質性的哪些組成部分是跨查詢的穩定讀者特性?}通過分離三個可測量的對象——證據 \emph{活動}、\emph{序數偏好} 和 \emph{條件符號方向}——我們發現序數讀者幾何在四個獨立設置中保持穩定(分半 $ρ=0.60$--$0.83$):留一法干預、PRISM 偏好、RAMDocs 和 RAGuard。符號幾何是任務界限的:在開放式問答中較弱(0.14, 0.35),特別是對於誤導性和不相關的證據,但在二元事實檢查中較強(0.75),並且沒有顯著的序數差距,儘管仍低於其稀疏匹配的上限。稀疏性、解碼噪音和度量工件無法解釋主要的序數-符號差距。最後,穩定的序數相似性無法預測跨讀者干預轉移(oracle-distance $ρ=-0.27$;遺憾可靠性 $-0.28$)。存在讀者特定的效用,但偏好不是干預:穩定的排名相似性並不授權幫助/傷害決策的轉移。

Learnware for CSI Feedback: Scene-specific Small Models Can Do Big

2608.17760v1 by Xiangyi Li, Jiajia Guo, Chao-Kai Wen, Xin Geng, Shi Jin, Zhi-Hua Zhou

Intelligent channel state information (CSI) feedback is essential for realizing the high capacity and spectral efficiency goals of future 6G systems, yet existing deep learning solutions face a trade-off between model generalization and scenario-specific performance. Large neural networks generalize well but incur high computational and tuning costs, while small models excel in particular environments but require repetitive costly end-to-end training for each base station (BS). To address these challenges, we introduce a model repository-based deployment framework in which a centralized AI data center maintains a catalog of scene-specific CSI models. The repository is enhanced with a Learnware-based framework, where each model is associated with a specification including semantic part (network architecture parameters) and statistical part (codeboo-fingerprint embeddings of training-data distributions). A BS submits only its local statistical specifications to retrieve the most relevant pre-trained model, enhancing data privacy by avoiding raw CSI transmission and drastically reducing retrieval latency and communication overhead. We further develop a data-driven search strategy that matches codebook fingerprints to model performance, achieving over 90% selection accuracy. In simulations, our scheme yields 18.8% and 57.7% performance improvements over the General Model in LOS and NLOS scenarios, respectively while reducing local fine-tuning by up to 1000 samples and 100 epochs. This Learnware-based approach minimizes redundant training, maximizes model reuse, and supports rapid,privacy-enhancing deployment of CSI feedback models.

摘要:智能通道狀態資訊(CSI)反饋對於實現未來6G系統的高容量和頻譜效率目標至關重要,然而現有的深度學習解決方案在模型泛化和場景特定性能之間面臨權衡。大型神經網絡具有良好的泛化能力,但會產生高計算和調整成本,而小型模型在特定環境中表現出色,但需要對每個基站(BS)進行重複昂貴的端到端訓練。為了解決這些挑戰,我們引入了一個基於模型庫的部署框架,其中一個集中式AI數據中心維護著場景特定的CSI模型目錄。該庫通過一個基於Learnware的框架進行增強,每個模型都與一個規範相關聯,包括語義部分(網絡架構參數)和統計部分(訓練數據分佈的代碼簿指紋嵌入)。基站僅提交其本地統計規範,以檢索最相關的預訓練模型,通過避免原始CSI傳輸來增強數據隱私,並大幅減少檢索延遲和通信開銷。我們進一步開發了一種數據驅動的搜索策略,將代碼簿指紋與模型性能匹配,實現了超過90%的選擇準確率。在模擬中,我們的方案在LOS和NLOS場景中分別比通用模型提高了18.8%和57.7%的性能,同時將本地微調減少了多達1000個樣本和100個訓練周期。這種基於Learnware的方法最小化了冗餘訓練,最大化了模型重用,並支持快速、增強隱私的CSI反饋模型部署。

D$^2$ACCI: A Dual-Loop Diagnostic Protocol for Evidence-Preserving Agent Memory

2608.17756v1 by Xule Liu, Yijun Liu, Chao Li, Shao Kun

Memory is a key capability of LLM agents. Persistent memory extends this across sessions---enabling recall, revision, and personalization. Yet its multi-stage pipeline (ingestion, retrieval, filtering, generation) makes failures difficult to localize: end-to-end evaluation reveals that an error occurred, but not which stage caused it. Existing evaluations often report aggregate performance without paired statistical comparisons, slice-level non-regression checks, or stage-level diagnostic traces. We propose D$^2$ACCI (Diagnostic-Driven Artifact-based Closed-loop Controlled Iteration), a dual-loop protocol whose outer diagnostic gate promotes, feature-flags, or rejects memory interventions based on paired evidence, protected-slice monitoring, and trace-level localizability. We further introduce DCR, a graded observability metric that measures whether failures remain localizable, and D$^2$ACCI-Eval, a reusable artifact for gate replay. We instantiate the protocol in MemStack and evaluate on three public benchmarks, achieving 93.59% on LoCoMo, 90.93% on LongMemEval, and 57.20% on PersonaMem-V2. Five paired ablations show that supplement extraction, session-memory retrieval, and Forget Guard yield statistically significant gains (+1.9 to +3.7pp, all p $\le$ .003). In contrast, BM25/RRF is retained as a monitored feature flag---a distinction invisible to aggregate-only evaluation. A diagnostic audit shows enriched traces substantially improve root-cause agreement over result-only relabeling. Diagnostic artifacts reach 98--100% DCR@3 versus 0% for results-only logs. These results establish that robust memory-system iteration demands traceable, statistically grounded, and regression-aware evidence---exactly the gap D$^2$ACCI fills.

摘要:記憶是 LLM 代理的一項關鍵能力。持久記憶擴展了這一能力,使其跨越多個會話——實現回憶、修訂和個性化。然而,其多階段管道(攝取、檢索、過濾、生成)使得故障難以定位:端到端評估顯示發生了錯誤,但無法確定是哪一階段導致的。現有的評估通常報告綜合性能,而沒有配對的統計比較、切片級別的非回歸檢查或階段級別的診斷追蹤。我們提出 D$^2$ACCI(基於診斷的工件閉環控制迭代),這是一種雙循環協議,其外部診斷閘根據配對證據、受保護的切片監控和追蹤級別的可定位性來促進、標記或拒絕記憶干預。我們進一步介紹 DCR,一種分級可觀察性指標,用於衡量故障是否仍然可定位,以及 D$^2$ACCI-Eval,一種可重用的工件,用於閘重放。我們在 MemStack 中實現該協議,並在三個公共基準上進行評估,在 LoCoMo 上達到 93.59%、在 LongMemEval 上達到 90.93%、在 PersonaMem-V2 上達到 57.20%。五個配對的消融實驗顯示,補充提取、會話記憶檢索和忘記保護產生了統計上顯著的增益(+1.9 到 +3.7 個百分點,所有 p $\le$ .003)。相比之下,BM25/RRF 被保留為監控的特徵標記——這一區別在僅進行綜合評估時是不可見的。診斷審計顯示,豐富的追蹤顯著改善了根本原因的一致性,相較於僅結果的重新標記。診斷工件在 DCR@3 上達到 98--100%,而僅結果的日誌為 0%。這些結果表明,穩健的記憶系統迭代需要可追蹤的、統計基礎的和意識到回歸的證據——正是 D$^2$ACCI 所填補的空白。

The Curious Case of Exploding DecPOMDPs: Containing the Fire through Policy Counting

2608.17749v1 by Nazlı Nur Karabulut, tanya Braun

Decentralised partially observable Markov decision processes (DecPOMDPs) provide a general framework for modelling multi-agent decision making under uncertainty. However, DecPOMDPs are known to suffer from exponential complexity in the number of agents. One way to combat this intractability in agent numbers is to look at partitions of agents that exhibit a form of symmetry among agents, allowing for a compact encoding by counting. However, a challenge arises as the policy space explodes, even though the model complexity and evaluation cost reduce to a polynomial dependence. In this paper, we redirect our focus from counting agents to counting policies, which actually enables tractability in agent numbers for so called policy-counted DecPOMDPs. Further, we present policy-counted dynamic programming using the compact representation to solve policy-counted DecPOMDPs efficiently.

摘要:去中心化的部分可觀察馬可夫決策過程(DecPOMDPs)提供了一個通用框架,用於建模在不確定性下的多代理決策。然而,DecPOMDPs 以代理數量的指數複雜性而聞名。對抗代理數量的這種難以處理的情況的一種方法是考慮具有某種對稱性的代理分區,這樣可以通過計數來實現緊湊編碼。然而,隨著政策空間的爆炸性增長,即使模型複雜性和評估成本降低到多項式依賴,挑戰依然存在。在本文中,我們將重點從計數代理轉向計數政策,這實際上使得所謂的政策計數 DecPOMDPs 在代理數量上變得可處理。此外,我們使用緊湊表示法提出了政策計數動態規劃,以高效解決政策計數的 DecPOMDPs。

Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See

2608.17744v1 by Ayoub Kirouane, Christos Petrocheilos

Take three frontier mixture-of-experts models (Alibaba, OpenAI, NVIDIA; 3.6-4.0B active parameters each) and fine-tune them to reason in a low-resource language. On accuracy benchmarks almost nothing happens, and the benchmark itself is noise at this scale: changing only the random seed moves the score by 7.7 points, more than every data and recipe effect we measured. That null is our first result. The real changes live where accuracy cannot see. Base models never think in Greek: 0 of 1,000 reasoning traces, even when the question is Greek, so the model answers correctly while reasoning in a form its user cannot read, audit, or correct. After supervised fine-tuning (SFT), every released checkpoint reasons in the language of the question on ~98% of items, one family at 3x fewer tokens, with judged grammaticality improving on all four models and general ability within a few points of each base: nothing was forgotten, and fluency was gained. We propose six behavioural dimensions that make such changes measurable, each gated to reject any metric that correlates with output length, and we report how our own instruments lied: six failures, each caught by a control. What SFT cannot do is fix its own defects: a quarter of answers skip the requested format, answers leak into the reasoning channel, and an explicit "think in English" is obeyed under half the time. Reinforcement learning with verifiable rewards, pre-registered before training, fixes the first two outright (fallback 24% to 2.5%, leak 3.5% to 0.0%, both against a flat random-reward control) and moves the third (+9.1pp), while the Greek reasoning habit survives an accuracy-only gradient untouched. We release five checkpoints. The instruments, the controls and the pre-registration travel to any low-resource language; Greek is the case that let us measure them.

摘要:取三個前沿的專家混合模型(Alibaba、OpenAI、NVIDIA;每個模型有 3.6-4.0B 的活躍參數)並對其進行微調,以便在低資源語言中進行推理。在準確性基準測試中幾乎沒有任何變化,而基準本身在這個規模下是噪音:僅改變隨機種子就能使得分數變動 7.7 分,這比我們測量的每個數據和配方效果都要大。這個無效結果是我們的第一個結果。真正的變化存在於準確性無法看到的地方。基礎模型從未用希臘語思考:在 1,000 條推理痕跡中,0 條是希臘語,即使問題是希臘語,因此模型在用用戶無法閱讀、審核或修正的形式進行推理時,仍然能正確回答。在監督微調(SFT)之後,每個釋出的檢查點在約 98% 的項目中用問題的語言進行推理,某一類別的標記數量減少到 3 倍,所有四個模型的語法判斷能力均有所改善,並且一般能力在每個基礎模型的幾個點之內:沒有任何被遺忘,流利度得到了提升。我們提出了六個行為維度,使這些變化可測量,每個維度都設置了閘門,以拒絕任何與輸出長度相關的度量,並報告我們自己的工具如何欺騙了我們:六次失敗,每次都被控制捕捉到。SFT 無法修復其自身缺陷:四分之一的答案跳過請求的格式,答案洩漏到推理通道中,並且明確的「用英語思考」在一半的時間內未被遵守。使用可驗證獎勵的強化學習,在訓練前進行預註冊,徹底修復了前兩個問題(回退從 24% 降至 2.5%,洩漏從 3.5% 降至 0.0%,均對比於隨機獎勵控制)並使第三個問題改善了 (+9.1pp),而希臘推理習慣在準確性僅有的梯度下仍然未受影響。我們釋出五個檢查點。這些工具、控制和預註冊可以應用於任何低資源語言;希臘語是讓我們能夠測量它們的案例。

Neuro-symbolic learning over OWL 2 DL via consequence-based compilation to differentiable circuits

2608.17741v1 by Olga Mashkova, Asaad Mohammedsaleh, Fernando Zhapa-Camacho, Robert Hoehndorf

OWL 2 DL ontologies, grounded in the description logic $\mathcal{SROIQ}$, express large knowledge bases in biomedicine and the Semantic Web. Neuro-symbolic (NeSy) learners over description logics either embed the ontology in a continuous space, abandoning classical entailment, or restrict to the Horn fragment $\mathcal{EL}^{++}$, which has a single canonical model. We present Baobab, which compiles a $\mathcal{SROIQ}$ ontology with a finite ABox into a Sentential Decision Diagram (SDD): it saturates a propositional core under a consequence-based calculus and instantiates the remaining $\mathcal{SROIQ}$ features (nominals, number restrictions, and the role axioms) over the active domain. The SDD's evidence-conditioned weighted model count then trains a perception network to recognize real images under partial ABox supervision: on an ontology that exercises every distinctive $\mathcal{SROIQ}$ feature, a CNN learns to read MNIST digits coupled by a successor relation and recovers latent ontology concepts that an independent perception leaves at chance. When the supervision admits several ontology-consistent completions, an independent perception collapses onto one, a reasoning shortcut: we show that a mixture indexed by the query's justifications can represent the calibrated posterior no independent perception can, and that seeding it from the circuit's enumerated completions attains the Bayes-optimal posterior on a real-image MNIST task where single-WMC and learned mixtures (the BEARS-ensemble hypothesis class) do not: to our knowledge the first to characterize and mitigate reasoning shortcuts in a non-Horn description logic. Soundness of the compiler and the representation result are machine-checked in Lean 4. Code is available at https://github.com/bio-ontology-research-group/baobab.

摘要:OWL 2 DL 本體,基於描述邏輯 $\mathcal{SROIQ}$,在生物醫學和語意網中表達大型知識庫。神經符號(NeSy)學習者在描述邏輯上要麼將本體嵌入連續空間,放棄傳統的推理,要麼限制於只有一個典範模型的 Horn 片段 $\mathcal{EL}^{++}$。我們提出了 Baobab,它將具有有限 ABox 的 $\mathcal{SROIQ}$ 本體編譯為句子決策圖(SDD):它在基於結果的計算下飽和一個命題核心,並在活動域上實例化剩餘的 $\mathcal{SROIQ}$ 特徵(名詞、數量限制和角色公理)。SDD 的證據條件加權模型計數然後訓練一個感知網絡,以在部分 ABox 監督下識別真實圖像:在一個行使每個獨特 $\mathcal{SROIQ}$ 特徵的本體上,CNN 學會閱讀與後繼關係相結合的 MNIST 數字,並恢復獨立感知所留下的潛在本體概念。當監督允許多個本體一致的完成時,獨立感知會崩潰到一個,這是一種推理捷徑:我們展示了一種由查詢的正當性索引的混合可以表示經過校準的後驗,而沒有獨立感知可以做到,並且從電路的列舉完成中種子達到在一個真實圖像 MNIST 任務上的貝葉斯最佳後驗,而單一 WMC 和學習的混合(BEARS-ensemble 假設類)則無法做到:據我們所知,這是第一次在非 Horn 描述邏輯中表徵和減輕推理捷徑。編譯器的健全性和表示結果在 Lean 4 中經過機器檢查。代碼可在 https://github.com/bio-ontology-research-group/baobab 獲得。

What Aggregate Scores Miss: Measuring Item-Level Regressions in Commercial LLM API Migrations

2608.17719v1 by Xiaonan Xu, Wenjing Wu

Context: Software systems that depend on commercial large language model APIs must migrate to successor versions when vendors deprecate older models. Migration decisions typically rely on aggregate benchmark scores, which compress heterogeneous item-level behaviour into a single net figure. Objective: We measure what that compression conceals. Method: On three pairwise upgrades in the GPT-5.4 to GPT-5.6 Sol product sequence, we query 900 public benchmark items (graduate-level knowledge, olympiad mathematics, instruction following) 50 times per item per model, classify each item as reliably improved, reliably regressed, practically equivalent, or inconclusive under false-discovery-rate control and a practical-significance threshold, and calibrate the results against a label-permutation null. Results: Across all nine migration-benchmark cells, reliable improvements and reliable regressions coexist. Edges with aggregate gains of up to 7.3 percentage points contain up to 8.3% reliably regressed items; edges with aggregate losses contain up to 10.7% reliably improved items. On the instruction-following benchmark, the gap between strict and loose scoring widens by 3.9 percentage points on the latest migration: a 3.9-point regression under strict scoring shrinks to 0.04 points under loose scoring. Conclusion: Migration decisions based on aggregate scores alone miss substantial bidirectional item-level change. The complete response-level archive and per-item scoring outputs are released.

摘要:背景:依賴商業大型語言模型 API 的軟體系統必須在供應商棄用舊模型時遷移到後繼版本。遷移決策通常依賴於綜合基準分數,這將異質的項目級行為壓縮為單一的淨數字。目標:我們測量這種壓縮所隱藏的內容。方法:在 GPT-5.4 到 GPT-5.6 Sol 產品序列的三次成對升級中,我們對 900 個公共基準項目(研究生級知識、奧林匹克數學、指令遵循)進行每個模型每項 50 次查詢,並根據假發現率控制和實際顯著性閾值將每個項目分類為可靠改進、可靠退步、實際等效或不確定,並將結果與標籤置換無效進行校準。結果:在所有九個遷移基準單元中,可靠的改進和可靠的退步共存。具有高達 7.3 個百分點的綜合增益的邊緣包含高達 8.3% 的可靠退步項目;具有綜合損失的邊緣包含高達 10.7% 的可靠改進項目。在指令遵循基準上,最新遷移中嚴格與寬鬆評分之間的差距擴大了 3.9 個百分點:在嚴格評分下的 3.9 點退步在寬鬆評分下縮小至 0.04 點。結論:僅根據綜合分數作出的遷移決策忽略了實質的雙向項目級變化。完整的響應級存檔和每項的評分輸出已發布。

Beyond Suspicious Steps: Ontological Trust in Long-Horizon Agents

2608.17718v1 by An He, Yao Wang, Haibin Zhang

Long-horizon agents increasingly operate across many steps, tools, and observa- tions. In this setting, the relevant oversight question is not only whether each action is locally valid, but whether the evolving trajectory still corresponds to the task the user authorized. Drift can accumulate quietly: an agent may call the right tool with plausible arguments at every step, while its prefix moves toward a broader role, an adjacent objective, or evidence the user never supplied. Existing monitors mostly check local compliance, deliver final-trace verdicts, or score generic risk; they do not directly estimate this prefix-level relation. We introduce ontological trust, a task-conditioned property of trajectory prefixes, and instantiate it as RGE, an online monitor that decomposes trust along Role, Goal, and Evidence. RGE uses LLMs only to derive structured task and step representations; trust-state updates, projec- tions, and intervention decisions are deterministic, so the output is a replayable and auditable trust trajectory rather than a single end-to-end judge verdict. We construct a cross-domain trajectory corpus from OSWorld, FinanceBench, and EICU-AC, covering benign executions, prefix-paired drift, and pseudo-consistency failures. On this corpus, RGE outperforms adapted rule-, judge-, and shield-style baselines on prefix-paired drift detection. With the two larger estimator models, it exceeds 93% Drift F1 on every benchmark while keeping benign coverage at or above 95.8%. Pseudo-consistency is harder: detection depends on whether task completion is externally visible, a structural limit we characterize empirically.

摘要:長期代理人越來越多地在多個步驟、工具和觀察中運作。在這種情況下,相關的監督問題不僅是每個行動是否在局部有效,而是演變的軌跡是否仍然符合用戶授權的任務。漂移可能悄然積累:一個代理人可能在每一步都以合理的論據調用正確的工具,而其前綴卻朝著更廣泛的角色、相鄰的目標或用戶從未提供的證據移動。現有的監控器主要檢查局部合規性,提供最終的追蹤判決,或評分一般風險;它們並不直接估計這種前綴級別的關係。我們引入本體信任,這是一種基於任務的軌跡前綴特性,並將其具體化為RGE,一種在線監控器,沿著角色、目標和證據分解信任。RGE僅使用LLMs來推導結構化的任務和步驟表示;信任狀態更新、投影和干預決策是確定性的,因此輸出是一個可重播和可審計的信任軌跡,而不是單一的端到端判決。 我們從OSWorld、FinanceBench和EICU-AC構建了一個跨領域的軌跡語料庫,涵蓋良性執行、前綴配對漂移和偽一致性失敗。在這個語料庫上,RGE在前綴配對漂移檢測上超越了適應的規則、判決和屏障風格基準。使用兩個更大的估計模型,它在每個基準上都超過93%的漂移F1,同時保持良性覆蓋率在95.8%或以上。偽一致性更難:檢測取決於任務完成是否在外部可見,這是一個我們經驗性描述的結構性限制。

Communicating Credit Risk with Large Language Models: Evaluation of Explanations from Standard and Alternative Data-Based Models

2608.17715v1 by Sahab Zandi, Noah Kostesku, Christophe Mues, María Óskarsdóttir, Cristián Bravo

Credit decisioning is a high-stakes task in which model outputs must be accurate and explainable to support compliant decisions. Although modern credit risk models such as eXtreme Gradient Boosting (XGBoost) and Graph Neural Networks (GNNs) improve predictive performance, their explanations are often too technical for stakeholders creating communication gaps that can shape approvals, denials, and fairness judgments. We examine whether Large Language Models (LLMs) can serve as explanation layers that translate post-hoc explanation artefacts into stakeholder-appropriate risk narratives. Using Freddie Mac single-family loan-level data, we develop three pipelines: standard tabular (XGBoost + SHAP), and two with alternative data, a pure network-based (GNN + GNNExplainer), and a bimodal one (combining tabular and network data). We generate narratives with three LLM configurations: a small fine-tuned LLM (Gemma 3 4B), a large fine-tuned LLM (DeepSeek R1 70B), and a zero-shot commercial LLM (Gemini 2.5). Explanation quality is evaluated through automated checks across all pipelines and a human study of bimodal explanations comparing credit risk professionals and non-professionals on eight decision-relevant dimensions. We have three main findings. First, the pipeline accounts for higher variance in evidence-grounding scores than the language model, meaning that the binding constraint on explanation quality is the evidence representation, not the model used. Second, the explanation narratives reliably name the influential factors but are less reliable when stating the direction of influence, which may be consequential for adverse-action communication. Finally, professionals apply stricter evidentiary standards than non-professionals. We discuss implications for the governance of risk models, including deployment considerations and the value of domain-aligned LLMs in regulated credit settings.

摘要:信用決策是一項高風險的任務,其中模型輸出必須準確且可解釋,以支持合規的決策。儘管現代信用風險模型如極端梯度提升(XGBoost)和圖神經網絡(GNNs)提高了預測性能,但它們的解釋往往對利益相關者來說過於技術性,造成溝通差距,這可能影響批准、拒絕和公平性判斷。我們檢視大型語言模型(LLMs)是否可以作為解釋層,將事後解釋產物轉化為適合利益相關者的風險敘事。使用Freddie Mac的單戶貸款數據,我們開發了三個管道:標準表格(XGBoost + SHAP),以及兩個使用替代數據的管道,一個是純基於網絡的(GNN + GNNExplainer),另一個是雙模的(結合表格和網絡數據)。我們使用三種LLM配置生成敘事:一個小型微調LLM(Gemma 3 4B),一個大型微調LLM(DeepSeek R1 70B),以及一個零樣本商業LLM(Gemini 2.5)。通過對所有管道的自動檢查以及對雙模解釋的人工研究,我們評估了解釋質量,並比較了信用風險專業人員和非專業人員在八個與決策相關的維度上的表現。我們有三個主要發現。首先,該管道在證據基礎分數的變異性上比語言模型更高,這意味著解釋質量的約束是證據表示,而不是所使用的模型。其次,解釋敘事可靠地命名了影響因素,但在陳述影響方向時可靠性較低,這對於不利行動的溝通可能具有重要意義。最後,專業人士應用的證據標準比非專業人士更為嚴格。我們討論了風險模型治理的影響,包括部署考量和在受監管的信用環境中領域對齊的LLMs的價值。

Accuracy and Robustness of Model Cascades Under Data Perturbations

2608.17711v1 by Pallavi Mitra, Jai Kushwaha, Felix Biessmann

Prediction cascades significantly reduce energy consumption of Artificial Intelligence (AI) models while maintaining high predictive performance. The idea is that easy inputs are routed through a lightweight small model, and difficult uncertain cases are deferred to a larger model. While this design can improve computational efficiency on clean data, its effectiveness depends on the reliability of confidence-based routing. Input degradations, such as static corruptions and sequential perturbations, can shift model confidence and routing decisions. In this paper, we study confidence-based cascade frameworks for image classification and investigate how such degradations affect their confidence-based deferral behavior. We select a model cascade at the pareto-optimum of accuracy, routing quality, and energy consumption that achieves competitive predictive performance with an up to 10-fold decrease in CO$_2$ emissions. We study the behavior of that model cascade under input corruptions and analyze how the cascade's routing decisions change when the input distribution shifts. Our analysis identifies three failure modes. Static corruptions either (1) break the routing signal while the large model remains useful, or (2) degrade both models so deferral no longer recovers accuracy. Sequential perturbations reveal a third mode: predictions stabilize but deferral suppresses, yielding stable but unreliable predictions. These findings demonstrate that energy efficient model cascades require evaluation beyond clean accuracy, with explicit attention to routing reliability under distribution shift.

摘要:預測級聯顯著降低人工智慧(AI)模型的能耗,同時保持高預測性能。其理念是將簡單的輸入通過一個輕量的小模型處理,而將困難的不確定案例延遲到更大的模型中。雖然這種設計可以在乾淨數據上提高計算效率,但其有效性取決於基於信心的路由的可靠性。輸入退化,例如靜態損壞和序列擾動,可能會改變模型的信心和路由決策。在本文中,我們研究了用於圖像分類的基於信心的級聯框架,並探討這些退化如何影響其基於信心的延遲行為。我們選擇了一個在準確性、路由質量和能耗的帕累托最優解上的模型級聯,該級聯在CO$_2$排放量上實現了高達10倍的減少,同時達到競爭性的預測性能。我們研究了該模型級聯在輸入損壞下的行為,並分析當輸入分佈發生變化時,級聯的路由決策如何改變。我們的分析確定了三種失效模式。靜態損壞要麼(1)破壞路由信號,而大型模型仍然有用,要麼(2)使兩個模型都退化,導致延遲不再恢復準確性。序列擾動揭示了第三種模式:預測穩定,但延遲被抑制,產生穩定但不可靠的預測。這些發現表明,能源高效的模型級聯需要在乾淨準確性之外進行評估,並明確關注在分佈變化下的路由可靠性。

GADR: Gathering Architecture Decision Records from Meeting Transcriptions

2608.17694v1 by Lucas Daniel Costa da Silva, Kiev Gama

Existing LLM-based approaches to Architecture Decision Record (ADR) generation share a critical and largely unexamined assumption: that input is already reasonably structured. In practice, architectural decisions emerge from informal, noisy meetings where choices are implicit, fragmented, and entangled with off-topic dialogue, precisely the conditions under which single-pass prompting degrades. This paper presents GADR, a multi-agent, self-correcting workflow that extracts architectural decisions from raw meeting transcriptions and generates Nygard-formatted ADR drafts. A feasibility study comprising five real project meeting transcripts, expert review by four senior architects, and evaluation by fifteen students provides initial evidence that the agentic workflow captures most expert-identified decisions and produces drafts participants found clear and useful, outperforming zero-shot and few-shot baselines in stability and structural adherence. The study also addresses the underexplored trade-off of RAG-based enrichment improving ADR depth while simultaneously risking transcript-unfaithful content, raising open questions about traceability in automated architectural documentation that we believe is worth the community's attention.

摘要:現有基於LLM的方法在架構決策記錄(ADR)生成中共享一個關鍵且大多未經檢驗的假設:即輸入已經相當結構化。實際上,架構決策源自非正式、嘈雜的會議,其中選擇是隱含的、零散的,並與無關的對話交織在一起,這正是單次提示退化的條件。本文提出了GADR,一種多代理、自我校正的工作流程,從原始會議記錄中提取架構決策並生成Nygard格式的ADR草稿。一項包含五個真實項目會議記錄的可行性研究、四位資深建築師的專家評審,以及十五名學生的評估提供了初步證據,表明該代理工作流程捕捉了大多數專家識別的決策,並生成了參與者認為清晰且有用的草稿,在穩定性和結構遵循性方面超越了零-shot和少-shot基準。該研究還探討了基於RAG的豐富性改善ADR深度的未充分探索的權衡,同時冒著轉錄不忠實內容的風險,提出了關於自動化架構文檔中可追溯性的開放問題,我們認為這值得社群的關注。

Mixture-of-Expert Blocks Contain Strong Hallucination Detection Signals

2608.17687v1 by Joao Fonseca, Rodrigo Rodrigues, Paolo Romano

Despite their widespread use, Large Language Models (LLMs) remain limited by a fundamental problem: the generation of plausible but false content, known as hallucinations. Most existing detection methods operate at the answer or sentence level, yet per-token detection is essential for localizing hallucinated spans and enabling fine-grained interventions. In this paper, we explore the use of the Mixture-of-Experts (MoE) paradigm to address this gap. In MoE architectures, a single forward pass activates a sparse subset of experts (i.e., distinct feedforward networks per layer) via a routing mechanism, producing internal signals (e.g., router entropy, expert disagreement, and expert usage patterns) that are unavailable in dense architectures and have not been previously exploited for hallucination detection. To this end, we introduce InnerExpert, the first method to leverage these MoE-specific signals for per-token hallucination detection. InnerExpert combines routing-level and standard transformer signals into compact per-token feature vectors, classified by a lightweight detector trained on labels produced by an LLM-as-a-judge pipeline, which enables continuous model updates without manual annotation. Our results show that InnerExpert outperforms existing methods across five datasets and two MoE architectures, achieving up to 0.91 answer-level and 0.76 token-level AUROC, while requiring only a single forward pass.

摘要:儘管大型語言模型(LLMs)被廣泛使用,但仍然受到一個根本問題的限制:生成看似合理但實際上錯誤的內容,稱為幻覺。大多數現有的檢測方法在答案或句子層面運作,然而,逐字檢測對於定位幻覺範圍和實現精細干預至關重要。在本文中,我們探討使用專家混合(Mixture-of-Experts, MoE)範式來解決這一差距。在MoE架構中,單次前向傳播通過路由機制激活稀疏的專家子集(即每層的不同前饋網絡),產生內部信號(例如,路由熵、專家不一致性和專家使用模式),這些信號在密集架構中不可用,且未被用於幻覺檢測。為此,我們介紹了InnerExpert,這是第一種利用這些MoE特定信號進行逐字幻覺檢測的方法。InnerExpert將路由層級和標準Transformer信號結合成緊湊的逐字特徵向量,並由一個輕量級檢測器進行分類,該檢測器在由LLM作為裁判管道產生的標籤上進行訓練,這使得模型可以在不需要人工標註的情況下進行持續更新。我們的結果顯示,InnerExpert在五個數據集和兩個MoE架構中超越了現有方法,達到了高達0.91的答案級別和0.76的逐字級別AUROC,同時僅需一次前向傳播。

Benchmarking Automated Security Patch Backporting: How Far Are We?

2608.17671v1 by Jincheng Yang, Yulong Fu, Chengwei Liu, Lyuye Zhang, Fangyuan Zhang, Bingyang Ren, Yang Liu, Hui Li

Automated security patch backporting is critical for mitigating N-day vulnerabilities. Recent tools report success rates above 80% on their respective datasets. However, these evaluations are often confined to homogeneous environments, such as one repository or specific project versions. Consequently, it remains unclear how well these tools generalize beyond their originally targeted scenarios. We present Porting Benchmark, a curated dataset of 1,234 security patch backporting cases spanning cross-version, cross-branch, and cross-repository scenarios, paired with a common evaluation framework. Using this benchmark, we evaluate five tools spanning program analysis, LLM prompting, and LLM agents under aligned settings. Our results show that aligned evaluation changes the apparent performance landscape: PortGPT and TSBPort remain comparatively strong on the Replication Dataset, while FixMorph and Mystique degrade substantially under the common protocol. Performance degrades sharply on structurally complex patches: the best commit-level success rate falls from 85.2% on Type-I patches to 24.0% on Type-IV. We identify four root-cause categories (missing target API awareness, cross-version semantic mismatch, non-local dependency propagation failure, and patch construction or localization failure) and derive concrete directions for next-generation tool design. On a 45-case dynamically validated subset with verified test cases and constructed POCs, we further observe that reference-based benchmark scores do not fully capture real-world remediation: exact match sharply under-credits harder target adaptations, while executable validation reveals residual integration failures in the target that static reference agreement misses. Executable-feedback refinement provides limited but measurable recovery on the hardest executable cases.

摘要:自動化安全補丁回溯對於減輕N天漏洞至關重要。最近的工具在其各自的數據集上報告的成功率超過80%。然而,這些評估通常僅限於同質環境,例如單一代碼庫或特定項目版本。因此,這些工具在其最初目標場景之外的普遍性仍不明確。我們提出了Porting Benchmark,這是一個經過精心策劃的數據集,包含1,234個安全補丁回溯案例,涵蓋跨版本、跨分支和跨代碼庫的場景,並配有一個共同的評估框架。利用這個基準,我們評估了五種工具,涵蓋程序分析、LLM提示和LLM代理,在對齊的設置下進行測試。我們的結果顯示,對齊評估改變了表面上的性能格局:PortGPT和TSBPort在複製數據集上仍然相對強勁,而FixMorph和Mystique在共同協議下顯著降級。對於結構複雜的補丁,性能急劇下降:最佳提交級成功率從Type-I補丁的85.2%降至Type-IV的24.0%。我們確定了四個根本原因類別(缺乏目標API認知、跨版本語義不匹配、非本地依賴傳播失敗,以及補丁構建或本地化失敗),並為下一代工具設計提供了具體方向。在一個包含經過驗證的測試案例和構建的POC的45案例動態驗證子集中,我們進一步觀察到基於參考的基準分數並未完全捕捉到現實世界的修復:精確匹配明顯低估了更難的目標適應,而可執行驗證揭示了靜態參考一致性所忽略的目標中的殘餘集成失敗。可執行反饋精煉在最難的可執行案例上提供了有限但可測量的恢復。

GraphWake: Group Polarization via Memory-Mediated Polarization Cascade in LLM-Agent Communities

2608.17665v1 by Haoran Bu, Zejian Chen, Litian Zhang, Xi Zhang

LLM-driven agents can autonomously exchange opinions on online platforms and form communities. Such agent-operated social platforms raise a new security concern: attackers may manipulate agents to induce group polarization. Existing methods manipulate agent prompts or construct echo chambers, both of which are difficult to realize in practice. We therefore formulate a new threat, Memory-Mediated Polarization Cascade, which uses agent memory as a persistence channel and public discussion as a propagation channel. This threat contains three stages. During exposure and memory retention, the attacker exposes a small set of target agents to arguments that reinforce their respective stated stances. The targets' memory systems then process and retain these arguments. During retrieval and reproduction, a shared stance-neutral discussion cues the targets to retrieve and reproduce their respective retained arguments. During iterative propagation, untreated agents influenced by the reproduced arguments restate and spread them. We instantiate this threat in GraphWake with three components: (i) stance-support argumentation knowledge graphs construct knowledge-based arguments; (ii) axiom-oriented triple selection distills them for reliable retention and reproduction; and (iii) stance-neutral memory cueing triggers concurrent retrieval and reproduction, initiating propagation. Experiments across multiple discussions and memory systems show that GraphWake substantially increases group polarization. These findings reveal a community-level polarization risk.

摘要:LLM 驅動的代理可以在在線平台上自主交換意見並形成社群。這種代理操作的社交平台引發了一個新的安全問題:攻擊者可能操縱代理以誘發群體極化。現有的方法操縱代理提示或構建回音室,這兩者在實踐中都難以實現。因此,我們提出了一種新的威脅,記憶介導的極化級聯,它利用代理記憶作為持久性通道,公共討論作為傳播通道。這一威脅包含三個階段。在暴露和記憶保留期間,攻擊者將一小組目標代理暴露於強化其各自表述立場的論點中。目標的記憶系統隨後處理並保留這些論點。在檢索和再現期間,共享的中立立場討論提示目標檢索並再現其各自保留的論點。在迭代傳播期間,受到再現論點影響的未處理代理重述並擴散這些論點。我們在 GraphWake 中實現了這一威脅,包含三個組件:(i)立場支持的論證知識圖構建基於知識的論點;(ii)公理導向的三元組選擇提煉它們以實現可靠的保留和再現;以及(iii)立場中立的記憶提示觸發同時檢索和再現,啟動傳播。多次討論和記憶系統的實驗顯示,GraphWake 顯著增加了群體極化。這些發現揭示了社群層面的極化風險。

MobileWorldSafety: Benchmarking GUI Agent Safety Against Environmental Injection Attacks in Android Apps

2608.17659v1 by Sujin Chen, Lijun Li, Tianyi Du, Jing Shao

LLM-powered GUI agents that autonomously operate smartphones are rapidly transitioning from research prototypes to early real-world deployment. However, because these agents routinely process untrusted environmental content, they are highly vulnerable to environmental injection attacks, which include indirect prompt injections and adversarial instructions. Such attacks can manipulate the behavior of agents without user awareness through diverse channels encountered in everyday mobile use. Despite these risks, existing benchmarks often fail to capture everyday user scenarios, lacking a systematic evaluation of GUI agents under environmental injection attacks on mobile devices. To address this gap, we introduce MobileWorldSafety, a benchmark of 142 risk tasks built on real Android applications. For each task, we define a programmatically verifiable risk indicator over the final system state and evaluate outcomes with a two-stage pipeline: rule-based verification handles unambiguous cases, while an LLM judge adjudicates ambiguous ones. This distinguishes safety failures from capability failures and enables objective and reproducible assessment. Evaluations on six agents, including both general agents and specialized GUI agents, demonstrate that all agents remain highly vulnerable, with attack success rates ranging from 40.4% to 66.9%. These findings indicate that current agents often fail to maintain safety alignment when adversarial content is presented as ordinary mobile context. MobileWorldSafety provides a foundation for quantifying these vulnerabilities and advancing research on robust mobile GUI agents.

摘要:LLM 驅動的 GUI 代理自動操作智能手機,正在迅速從研究原型轉向早期的實際部署。然而,由於這些代理經常處理不受信任的環境內容,它們對環境注入攻擊高度脆弱,這些攻擊包括間接提示注入和對抗性指令。這些攻擊可以通過日常移動使用中遇到的多樣渠道,在不讓用戶察覺的情況下操控代理的行為。儘管存在這些風險,現有基準往往未能捕捉到日常用戶場景,缺乏對移動設備上環境注入攻擊下 GUI 代理的系統評估。為了填補這一空白,我們推出了 MobileWorldSafety,一個基於真實 Android 應用的 142 個風險任務的基準。對於每個任務,我們定義了一個可程序驗證的風險指標,基於最終系統狀態進行評估,並通過兩階段的流程來評估結果:基於規則的驗證處理明確的情況,而 LLM 評判則裁決模糊的情況。這區分了安全失敗和能力失敗,並使客觀和可重複的評估成為可能。對六個代理的評估,包括一般代理和專門的 GUI 代理,顯示所有代理仍然高度脆弱,攻擊成功率範圍從 40.4% 到 66.9%。這些發現表明,當對抗性內容被呈現為普通的移動上下文時,當前的代理往往無法保持安全對齊。MobileWorldSafety 為量化這些脆弱性和推進穩健的移動 GUI 代理研究提供了基礎。

LLM-Derived Preference Judgments Are Not Self-Consistent

2608.17644v1 by Matthew T. Ford, Francis Bahk, Jingjing Wang, Adam S. Jovine, Tinghan Ye, David B. Shmoys, Peter I. Frazier

Agents increasingly interpret a person's natural-language preferences by querying an LLM for numerical preference judgments, e.g., by asking how much the person would be willing to pay for an item. A growing body of work estimates a utility function from these judgments and then chooses actions based on their estimated utility. This pipeline assumes the judgments are approximately self-consistent: that a single utility function can reproduce them. But are they? To study this question, we measure the self-consistency of cardinal LLM preference judgments. For example, the difference in stated willingness-to-pay between two items should match the stated payment that makes a person indifferent to exchanging them. We develop statistical tests and interpretable measures of how far observed responses depart from the best-fitting self-consistent utility function. Experiments with flight, apartment, and hotel examples across six LLMs reveal large persistent inconsistencies. This suggests that LLM-derived preference judgments cannot be faithfully summarized by a single utility function.

摘要:代理人越來越多地通過查詢 LLM 來解釋一個人的自然語言偏好,以獲得數值偏好判斷,例如,詢問這個人願意為某個項目支付多少。越來越多的研究從這些判斷中估計效用函數,然後根據其估計的效用選擇行動。這個流程假設這些判斷大致上是自我一致的:即單一的效用函數可以重現它們。但真的是這樣嗎?為了研究這個問題,我們測量了基數 LLM 偏好判斷的自我一致性。例如,兩個項目之間所表明的支付意願差異應該與使一個人對交換它們無所謂的所表明的支付相匹配。我們開發了統計檢驗和可解釋的度量,來衡量觀察到的反應與最佳擬合的自我一致效用函數之間的偏離程度。對六個 LLM 的航班、公寓和酒店範例進行的實驗揭示了持久的巨大不一致性。這表明 LLM 衍生的偏好判斷不能被單一的效用函數忠實地總結。

Beyond the Trace: Coupling an Interpretable Reasoning-State Readout to Native MoE Routing

2608.17638v1 by Kang Chen, Sihan Zhao, Yixin Cao, Yugang Jiang

What a reasoning model writes is only a partial record of the process that produces it. We introduce a two-level internal readout for mixture-of-experts reasoning. We first distill vocabulary-scale J-space into J64, a 64-axis semantic frame learned from the model's own reasoning states. J64 reveals readable process state that the emitted trace does not show: it separates inference effort from problem-induced strain. It also adds 0.096 to 0.135 held-out AUC over a baseline that reads the same rollout as token occupancy and aggregates it in exactly the same way. We then reconstruct J64 from native expert-routing statistics. The result is R64, a low-overhead proxy: its median per-axis correlation with J64 is 0.69 to 0.86 across three models and two families, and on gpt-oss-20b it preserves 95 to 100% of J64's predictive gain. The readout supports test-time decisions at two temporal resolutions. Over completed candidate sets, J64 and R64 improve single-branch selection, and R64-weighted voting improves plain majority voting in seven of eight settings. During generation, rolling readout windows drive a cumulative stop-and-resample policy whose operating point is fixed on training questions alone. J64 improves accuracy by 1.1 to 5.9 points over a sibling-permuted control, and the routing-only R64 proxy retains 0.9 to 3.2 of those points. Finally, router edits aimed at the mechanism J64 names induce the predicted reasoning behaviors and shift a diagnosed stall from numerical guessing toward exact symbolic execution. Together, J64 makes latent process state readable, while routing makes it deployable and actionable.

摘要:推理模型所寫的內容僅是產生該內容過程的部分記錄。我們為混合專家推理引入了兩級內部讀出。我們首先將詞彙規模的 J 空間提煉成 J64,這是一個從模型自身推理狀態學習而來的 64 軸語義框架。J64 揭示了可讀的過程狀態,而發出的痕跡並未顯示出來:它將推理努力與問題引起的壓力分開。它還在基準上增加了 0.096 到 0.135 的持出 AUC,該基準將同樣的展開視為標記佔用並以完全相同的方式進行聚合。然後,我們從原生專家路由統計中重建 J64。結果是 R64,一個低開銷的代理:它與 J64 的每軸中位數相關性在三個模型和兩個家族中為 0.69 到 0.86,而在 gpt-oss-20b 上,它保留了 J64 預測增益的 95% 到 100%。這個讀出支持在兩個時間解析度下的測試時決策。在完成的候選集上,J64 和 R64 改進了單分支選擇,而 R64 加權投票在八個設置中的七個中改善了普通多數投票。在生成過程中,滾動讀出窗口驅動一個累積的停止和重取樣策略,其操作點僅固定在訓練問題上。J64 在與兄弟置換控制相比中提高了 1.1 到 5.9 分的準確性,而僅路由的 R64 代理保留了 0.9 到 3.2 的這些分數。最後,針對 J64 所命名的機制的路由編輯引發了預測的推理行為,並將診斷出的停滯從數值猜測轉向精確的符號執行。總之,J64 使潛在的過程狀態可讀,而路由則使其可部署和可行動。

Graph Surgery and the Do-Operator: A Precise Correspondence for Acyclic Structural Causal Models

2608.17634v1 by Satpreet Makhija

The $\operatorname{do}$-operator is described graphically by deleting arrows into its targets and functionally by replacing their mechanisms with constants. To call these operations equivalent is not yet a mathematical statement: one returns a graph and remembers only the targets, whereas the other returns mechanisms and also remembers the imposed values. We make a dependency-level comparison precise for deterministic acyclic structural causal models with finitely many endogenous variables. If $\operatorname{Graph}(F)$ extracts the dependencies of a mechanism family $F$, our main theorem is $\operatorname{Graph}(F^ι)=\operatorname{Surg}(\operatorname{Graph}(F),T_ι)$. Thus replacing target mechanisms removes exactly the dependencies removed by graph surgery. For a model $M=(G,F)$ whose graph may contain unused arrows, we characterize when the same equality holds with $G$ in place of $\operatorname{Graph}(F)$; it holds for every intervention exactly when $G$ records the dependencies of $F$ exactly. We then define the intervened model, characterize its run, show how sequential interventions combine, and prove that an outcome depends only on interventions at its actual dependency ancestors.

摘要:$\operatorname{do}$-運算子在圖形上通過刪除指向其目標的箭頭來描述,而在功能上則通過用常數替換其機制來描述。將這些操作稱為等價尚未形成數學陳述:一個返回圖形並僅記住目標,而另一個返回機制並同時記住施加的值。我們對具有有限內生變量的確定性非循環結構因果模型進行依賴層級的精確比較。如果 $\operatorname{Graph}(F)$ 提取機制家族 $F$ 的依賴關係,我們的主要定理是 $\operatorname{Graph}(F^ι)=\operatorname{Surg}(\operatorname{Graph}(F),T_ι)$。因此,替換目標機制正好去除了圖形手術所去除的依賴關係。對於一個模型 $M=(G,F)$,其圖形可能包含未使用的箭頭,我們描述何時同樣的等式在 $G$ 代替 $\operatorname{Graph}(F)$ 時成立;當且僅當 $G$ 精確記錄 $F$ 的依賴關係時,它成立。我們然後定義干預模型,描述其運行,展示如何結合序列干預,並證明結果僅依賴於其實際依賴祖先的干預。

DEPT: Document Embedding Preservation Tuning for Unified Query Expansion and Retrieval

2608.17632v1 by Jingyuan Wang, Richong Zhang, Zhijie Nie, Mingxin Li, Yanzhao Zhang

Large language models (LLMs) can both expand underspecified queries and encode text as dense representations, suggesting a unified model for query expansion and retrieval. Existing systems usually rely on prompted expansions, independently trained modules, or staged optimization, leaving generated expansions only indirectly aligned with the retrieval loss that judges them. We train a single decoder-only LLM end to end, where the same model generates the expansion and encodes both the expanded query and candidate documents. This unified setting creates a moving-target problem: retrieval supervision should improve query-side expansion, but the same update also shifts the document embeddings that serve as retrieval targets. We introduce Document Embedding Preservation Tuning (DEPT), which keeps tuned document embeddings close to cached initial embeddings while allowing retrieval gradients to pass through straight-through decoding into the generator. DEPT converts joint query--document movement into query-side adaptation against approximately stable, whitened document embeddings that support index reuse and online hard-negative mining. Experiments with Qwen3-4B-Instruct-2507 and LLaMA-3.2-3B-Instruct on five datasets in BEIR benchmark show that DEPT improves average retrieval quality over training-free, independently trained, and staged unified baselines, while ablations isolate the effects of preservation, whitening, end-to-end expansion training, and online negatives. Code is available at https://github.com/ILSparkle/DEPT.

摘要:大型語言模型(LLMs)可以擴展未具體化的查詢並將文本編碼為密集表示,這表明查詢擴展和檢索的統一模型。現有系統通常依賴於提示擴展、獨立訓練的模塊或分階段優化,這使得生成的擴展與評估它們的檢索損失僅間接對齊。我們訓練了一個端到端的單解碼器 LLM,該模型同時生成擴展並編碼擴展查詢和候選文檔。這種統一的設置創造了一個移動目標問題:檢索監督應該改善查詢端的擴展,但相同的更新也會改變作為檢索目標的文檔嵌入。我們引入了文檔嵌入保護調整(DEPT),它保持調整後的文檔嵌入接近緩存的初始嵌入,同時允許檢索梯度通過直通解碼進入生成器。DEPT 將聯合查詢-文檔移動轉換為針對大致穩定的、經過去白化的文檔嵌入的查詢端適應,這些嵌入支持索引重用和在線困難負樣本挖掘。在 BEIR 基準的五個數據集上,使用 Qwen3-4B-Instruct-2507 和 LLaMA-3.2-3B-Instruct 的實驗顯示,DEPT 提高了平均檢索質量,超過了無需訓練的、獨立訓練的和分階段的統一基準,而消融實驗則隔離了保護、去白化、端到端擴展訓練和在線負樣本的影響。代碼可在 https://github.com/ILSparkle/DEPT 獲得。

From Student Risk Prediction to SC2R: Semantics-Constrained Counterfactual Recourse for Educational Decision Support

2608.17618v1 by Ngoc Luyen Le, Marie-Hélène Abel, Bertrand Laforge

Learning analytics models can identify students at risk of poor performance, but they do not directly indicate which interventions are feasible, actionable, and compatible with educational constraints. This paper introduces SC2R, a semantics-constrained counterfactual recourse framework for educational decision support. SC2R combines a calibrated predictive model, integer-programming-based recourse generation over discrete action variables, a lightweight RDF vocabulary for intervention-plan representation, and SHACL validation for enforcing timing, budget, immutability, and availability constraints. The framework is evaluated offline on the OULAD dataset using snapshots constructed relative to each assessment at two decision horizons. Results show that the predictive component provides strong performance, that compact intervention plans can be generated at scale, and that semantic validation reveals infeasible plans that lighter optimization-only settings would otherwise accept. Rather than claiming causal improvement in student outcomes, this work shows that counterfactual recourse becomes more operationally meaningful in education when recommendations are not only model-valid, but also semantically feasible and machine-checkable.

摘要:學習分析模型可以識別出有表現不佳風險的學生,但它們並不直接指示哪些干預措施是可行的、可操作的,以及與教育限制相容的。本文介紹了SC2R,一個語義約束的反事實補救框架,用於教育決策支持。SC2R結合了一個經過校準的預測模型、基於整數規劃的離散行動變數的補救生成、輕量級的RDF詞彙用於干預計劃表示,以及SHACL驗證以強制執行時間、預算、不變性和可用性約束。該框架在OULAD數據集上進行了離線評估,使用相對於每次評估在兩個決策視野下構建的快照。結果顯示,預測組件提供了強大的性能,能夠大規模生成緊湊的干預計劃,並且語義驗證揭示了在僅進行輕量優化的設置下會被接受的不可行計劃。本研究並不聲稱對學生結果的因果改善,而是顯示當推薦不僅是模型有效的,還是語義上可行且可機器檢查的時,反事實補救在教育中變得更具操作意義。

Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges

2608.17605v1 by Syeda Faiza Ahmed, Zien Sheikh Ali, Hunzalah Hassan Bhatti, Firoj Alam, Shammur Absar Chowdhury

Conversational AI is moving beyond isolated text prompts toward sustained, multimodal interaction. In real conversations, users clarify goals, revise requests, interrupt responses, switch topics, and introduce new evidence while expecting systems to preserve context across turns. This makes multi-turn dialogue a distinct challenge requiring systems to maintain and update memory, ground responses across modalities, tools, and external knowledge, and adapt across languages and cultures. This study reviews multi-turn conversational AI across text-only dialogue, AudioLLMs and speech-native systems, multimodal and omni-modal systems, and tool-augmented agents. We organize the literature around datasets and benchmarks, modeling paradigms, training strategies, evaluation setups, and cross-cutting challenges. Our analysis shows that support for multiple modalities has advanced faster than the ability to sustain coherent interaction across a session. Despite stronger capabilities to perceive, speak, and act across modalities, current systems still struggle with persistent memory, cross-turn grounding, full-duplex interaction, robust evaluation, and cultural alignment. We conclude with a research agenda for systems that can remember, revise, ground, speak, listen, act, and adapt across turns, modalities, and cultures. (https://github.com/faiza-sfa/multiturn-conversational-ai-survey)

摘要:對話式人工智慧正在超越孤立的文字提示,朝向持續的多模態互動發展。在真實的對話中,用戶會澄清目標、修訂請求、打斷回應、切換主題並引入新證據,同時期望系統能在不同回合中保持上下文。這使得多回合對話成為一個獨特的挑戰,要求系統維持和更新記憶,跨模態、工具和外部知識進行回應的基礎,並在語言和文化之間進行適應。本研究回顧了文本對話、AudioLLMs 和語音原生系統、多模態和全模態系統以及工具增強代理的多回合對話式人工智慧。我們根據數據集和基準、建模範式、訓練策略、評估設置和跨領域挑戰來組織文獻。我們的分析顯示,對多模態的支持發展得比在一個會話中持續一致互動的能力更快。儘管在感知、說話和跨模態行動方面的能力增強,當前的系統仍然在持久記憶、跨回合基礎、全雙工互動、穩健評估和文化對齊方面面臨挑戰。我們以一個研究議程作結,旨在開發能夠記住、修訂、基礎、說話、聆聽、行動和在回合、模態和文化之間適應的系統。(https://github.com/faiza-sfa/multiturn-conversational-ai-survey)

HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety

2608.17597v1 by Yajing Bai, Jinhao Duan, Jie Peng, Xianfeng Wu, Sijia Liu, Song Wang, Tianlong Chen

Large language models are increasingly deployed through agent harnesses that manage tools, extensions, persistent state, permissions, and external actions. Existing safety benchmarks mainly target individual attack mechanisms or a limited subset of operational settings, making it difficult to compare how safety failures emerge across different harness responsibilities. We present HarnessRisk, a lifecycle oriented benchmark that organizes agent harness safety into six operational phases including Harness Configuration, Capability Extension, Runtime Operation, State Persistence, Action Control, and Incident Recovery. HarnessRisk contains 128 sandboxed cases, each pairing a benign user objective with an adversarial instruction embedded in an untrusted workflow artifact. We evaluate each trajectory using Utility, Attack Success Rate, Persistence, and Detection. Across three harnesses, six language models, and 14 model and harness configurations, attack success ranges from 12.6% to 80.9%, while Utility remains between 75.0% and 97.6%. Harness Configuration is the most vulnerable phase across all three harnesses, showing that attacks can succeed by altering security sensitive parameters within otherwise authorized workflows. We also find that explicit risk recognition does not reliably lead to safe action, as some configurations detect risks in more than 90% of runs while retaining substantial attack success. These results highlight the need to evaluate agent safety across multiple harness responsibilities and at the level of the deployed model and harness configuration.

摘要:大型語言模型越來越多地透過代理工具來管理工具、擴展、持久狀態、權限和外部行動。現有的安全基準主要針對個別攻擊機制或有限的操作設定,使得比較不同工具責任下安全失敗的出現變得困難。我們提出了HarnessRisk,一個以生命週期為導向的基準,將代理工具的安全性組織為六個操作階段,包括工具配置、能力擴展、運行時操作、狀態持久性、行動控制和事件恢復。HarnessRisk包含128個沙盒案例,每個案例將一個良性的用戶目標與嵌入在不受信任的工作流程工件中的對抗指令配對。我們使用效用、攻擊成功率、持久性和檢測來評估每個軌跡。在三個工具、六個語言模型和14個模型及工具配置中,攻擊成功率範圍從12.6%到80.9%,而效用則保持在75.0%到97.6%之間。工具配置是所有三個工具中最脆弱的階段,顯示攻擊可以通過改變在其他授權工作流程中安全敏感的參數而成功。我們還發現,明確的風險識別並不可靠地導致安全行動,因為某些配置在超過90%的運行中檢測到風險,但仍然保留了相當大的攻擊成功率。這些結果突顯了需要在多個工具責任和部署的模型及工具配置層面上評估代理安全性。

tinyDSM: A Framework for Skill Modeling and Development for Resource-Constrained Millirobots

2608.17596v1 by Markus D. Kobelrausch, Michael Miedler, Axel Jantsch

In this study, we investigate developmental mechanisms that enable small, resource-constrained systems such as cm-sized millirobots to autonomously explore, learn, and adapt their capabilities throughout their lifespan. Reinforcement learning algorithms guide the agent's skill acquisition and adaptation through the interplay of our proposed tinyDSM, which integrates intrinsic motivation and fitness-based assessment. We strive for minimal, hard-wired skills while encouraging the open-ended development of new skills. A key emphasis in our approach is to encode minimal a-priori general knowledge, which serves as a foundational starting point for the system as it further learns system-specific dependencies from the initial knowledge provided. Thus, by design, our approach attempts to cover very generic application domains. The methodology is based on (a) developmental mechanism with intrinsic motivation, and (b) a cognitive architecture (knowledge, reasoning, learning), while (c) utilizing minimal resources. It uses a hierarchical knowledge graph and kinematic reasoners to model and evaluate simple and advanced motion related skills. In our experiments, we use a resource-constrained millirobot with a volume of 36 cm^3 with a Raspberry Pi Pico 32-bit microcontroller (RP2040) that integrates all described features and capabilities except the camera system in 9 kB. Starting with learning the most elementary motor skills the millirobot autonomously progresses from simple linear and angular movements to complex geometric patterns within 15 minutes. To complement the physical experiments, we perform a simulation-based analysis that enables systematic comparisons across learning algorithms and intrinsic motivation parameters.

摘要:在本研究中,我們探討使小型資源受限系統(如厘米級的微型機器人)能夠自主探索、學習和適應其能力的發展機制。強化學習算法通過我們提出的tinyDSM的相互作用來指導代理的技能獲得和適應,該系統整合了內在動機和基於適應度的評估。我們追求最小的硬連接技能,同時鼓勵新技能的開放式發展。我們方法的一個關鍵重點是編碼最小的先驗一般知識,這作為系統進一步從提供的初始知識中學習系統特定依賴的基礎起點。因此,我們的方法設計上試圖涵蓋非常通用的應用領域。該方法論基於(a)具有內在動機的發展機制,以及(b)一種認知架構(知識、推理、學習),同時(c)利用最小資源。它使用層次知識圖譜和運動學推理器來建模和評估簡單和高級運動相關技能。在我們的實驗中,我們使用一個資源受限的微型機器人,其體積為36 cm^3,搭載Raspberry Pi Pico 32位微控制器(RP2040),該微控制器整合了所有描述的功能和能力,除了攝像頭系統外,僅佔用9 kB。從學習最基本的運動技能開始,微型機器人自主地在15分鐘內從簡單的線性和角運動進展到複雜的幾何圖形。為了補充物理實驗,我們進行了一個基於模擬的分析,這使得能夠在學習算法和內在動機參數之間進行系統比較。

TRUSS: Towards Task-Reliable and User-Safe Automated Agent Skill Generation

2608.17588v1 by Zhibo Zhang, Zhen Ouyang, Ling Shi, Kailong Wang

Agent Skills package reusable natural language procedures with executable resources, enabling software agents to acquire task specific capabilities without model adaptation. Automatically generating such Skills can improve task performance, yet evaluating a candidate solely from its artifact or final task outcome leaves unresolved which actions the equipped agent will perform and which side effects those actions will produce. We present TRUSS, an evidence guided framework for generating functionally effective and safety reliable Agent Skills. TRUSS first inspects functional claims against source and domain evidence while evaluating the complete artifact under nine predefined safety properties. Candidates admitted by this static gate are loaded by a shadow agent inside a Controllable Execution Environment, where brokered tools expose requested actions to policy enforcement and record their results as provenance preserving execution traces. Functional failures and property violations are linked back to the responsible Skill content and used to guide iterative refinement. We evaluate TRUSS on 168 SkillInject artifacts, 155 SkillSafetyBench cases, and all 187 tasks in SkillGenBench. TRUSS achieves 100.00\% precision and recall in vulnerability detection. Repair reduces attack success from 38.71\% to 19.35\% with GPT 5.5 and from 46.45\% to 29.68\% with GPT 5.4, with zero attack regression. For Skill generation, TRUSS raises task effectiveness from 17.11\% without Skills to 52.94\%, while increasing the benchmark Security rate from 50.80\% to 100.00\%. These results show that execution evidence can expose behavioral failures missed by artifact inspection and can guide Skill generation toward jointly verified functional and safety outcomes.

摘要:代理技能包可重用自然語言程序及可執行資源,使軟體代理能夠獲得特定任務的能力,而無需模型調整。自動生成這些技能可以提高任務表現,但僅從其產物或最終任務結果評估候選者,無法解決裝備代理將執行哪些行動以及這些行動將產生哪些副作用。我們提出了TRUSS,一個基於證據的框架,用於生成功能有效且安全可靠的代理技能。TRUSS首先根據來源和領域證據檢查功能聲明,同時在九個預定義的安全性屬性下評估完整的產物。通過這個靜態閘口的候選者將由一個影子代理加載到可控執行環境中,在這裡,經紀工具將請求的行動暴露給政策執行,並將其結果記錄為保留來源的執行痕跡。功能失敗和屬性違規將回溯到負責的技能內容,並用於指導迭代改進。
我們在168個SkillInject產物、155個SkillSafetyBench案例和所有187個SkillGenBench任務上評估TRUSS。TRUSS在漏洞檢測中達到100.00\%的精確度和召回率。修復將攻擊成功率從38.71\%降低到19.35\%(使用GPT 5.5),並從46.45\%降低到29.68\%(使用GPT 5.4),且沒有攻擊回歸。對於技能生成,TRUSS將任務有效性從沒有技能的17.11\%提高到52.94\%,同時將基準安全率從50.80\%提高到100.00\%。這些結果表明,執行證據可以揭示產物檢查中遺漏的行為失敗,並能指導技能生成朝向共同驗證的功能和安全結果。

Write, Execute, Refine: From Skill Followers to Skill Optimizers via Reinforcement Learning from Execution Feedback

2608.17587v1 by Kang Peng, Zhiwei Zhang, Yichen Zhang, Zezhong Wang, Yiming Du, Geng Tu, Baojun Wang, Bin Liang, Ruifeng Xu, Kam-Fai Wong

Expert-written natural language skills can improve tool-using agents, yet agent-authored skills perform 8-11 points worse than using no skill. This gap suggests that following procedural guidance and improving it from execution evidence are distinct capabilities. Inference time loops can repair skills but do not improve the model that writes the next one. We study how to organize execution experience from intermediate skills into training states for an optimizer. We introduce WER (Write, Execute, and Refine), a multi-phase framework that trains a Skill Optimizer outside a frozen executor. The optimizer proposes skills, a frozen agent executes each repeatedly, and a programmatic verifier scores the outcomes. The scores provide relative credit and select mixed-outcome records. Matched successful and failed trajectories from these records form the next phase's refinement states, so the optimizer learns from the consequences of its earlier outputs. On BFCL v4 multi-turn and tau2-bench, WER improves average Pass@1 over the no-skill baseline by 7.80 and 3.85 points, respectively. Under an identical refinement workflow, it outperforms the same backbone without optimizer training by 9.35 and 10.29 points. The trained 4B optimizer reaches 76.63 percent on BFCL v4, outperforming all evaluated off-the-shelf general-purpose models used as skill optimizers on average.

摘要:專家撰寫的自然語言技能可以改善工具使用代理,但代理撰寫的技能表現比不使用技能低 8-11 分。這一差距表明,遵循程序指導和從執行證據中改進它是兩種不同的能力。推理時間循環可以修復技能,但不會改善撰寫下一個技能的模型。我們研究如何將中介技能的執行經驗組織成優化器的訓練狀態。我們引入 WER(寫作、執行和精煉),這是一個多階段框架,旨在在凍結的執行器之外訓練技能優化器。優化器提出技能,凍結的代理重複執行每個技能,程式驗證器對結果進行評分。這些分數提供相對的信用並選擇混合結果記錄。來自這些記錄的成功和失敗的匹配軌跡形成下一階段的精煉狀態,因此優化器從其早期輸出的後果中學習。在 BFCL v4 多輪和 tau2-bench 上,WER 分別將平均 Pass@1 提高了 7.80 和 3.85 分。在相同的精煉工作流程下,它比未經優化器訓練的相同骨幹高出 9.35 和 10.29 分。訓練後的 4B 優化器在 BFCL v4 上達到 76.63% 的表現,超越了所有評估的現成通用模型,並在平均上用作技能優化器。

Auditing Exposure to Harmful Content on TikTok using Multimodal Language Models: A Cross-National, Age-Stratified Study

2608.17583v1 by Hamidreza Saffari, Francesco Pierri

Online video platforms can expose young users to harmful content, but independent audits remain difficult because video annotation is costly and moderation judgments vary across languages. We audit TikTok in France, Italy, and Sweden with sockpuppet accounts representing four age personas (13, 16, 19, 40), collecting 36,971 videos from passive For-You-page scrolling and active sessions that scroll, search for harm keywords, and scroll again. To scale annotation, we validate four multimodal LLMs against native-speaker labels on a 300-video reference set. Gemini 2.5 Flash with eight sampled frames plus text performs best (aggregate kappa = 0.42), at half the per-call cost of native-video upload, and we apply it to a 10% sample for approximately \$50 in total API spend across both modalities. Keyword search returns 35-56% harmful content, a 1.5-7.5x increase over the scrolling baseline in ten of twelve country-age combinations; the spike is temporary and flattens the age differences observed in France and Sweden. Under passive scrolling, Italy has the highest harm rate at every age, with Italian age-19 reaching 48.6%. Overall, MLLM-based auditing offers a scalable approach for cross-national youth-safety audits, while provider safety filters (1.1% refusal rate) under-count the most explicit harms.

摘要:在線視頻平台可能會讓年輕用戶接觸到有害內容,但獨立審核仍然困難,因為視頻註釋成本高且不同語言的審核判斷存在差異。我們在法國、意大利和瑞典對 TikTok 進行審核,使用代表四個年齡角色(13、16、19、40)的假帳號,從被動的 For-You 頁面滾動和主動會話中收集 36,971 個視頻,這些會話會滾動、搜索有害關鍵詞,然後再次滾動。為了擴大註釋,我們在 300 個視頻的參考集上驗證了四個多模態 LLM,與母語者標籤進行比較。Gemini 2.5 Flash 使用八個取樣幀加上文本的表現最佳(綜合 kappa = 0.42),其每次調用成本僅為本土視頻上傳的一半,我們將其應用於 10% 的樣本,總 API 支出約為 50 美元。關鍵詞搜索返回 35-56% 的有害內容,在十二個國家-年齡組合中的十個中,這比滾動基線增加了 1.5-7.5 倍;這一激增是暫時的,並平坦了在法國和瑞典觀察到的年齡差異。在被動滾動下,意大利在每個年齡段的危害率最高,意大利的 19 歲達到 48.6%。總的來說,基於 MLLM 的審核為跨國青少年安全審核提供了一種可擴展的方法,而提供者的安全過濾器(拒絕率 1.1%)則低估了最明顯的危害。

Quantifying Risk Under Evolving Uncertainty: Belief-Dependent Robustness for Safe Sequential Decision Making

2608.17574v1 by Deep Kumar Ganguly, Jan Kretinsky

How cautious should an agent be while it is still learning its environment? We propose RATTL (Risk-Adversarial Total-Reward Learning), which ties caution to epistemic uncertainty: the agent holds a Bayesian posterior over unknown dynamics and plans against a Wasserstein ambiguity set whose radius is a monotone function of that posterior. The radius contracts with evidence, so behaviour interpolates continuously between worst-case robustness and risk-neutral total-reward maximization. The design follows the duality underlying the Entropic Value-at-Risk, which converts the choice of a risk level into the choice of an ambiguity radius. We show the resulting planning problem is well posed under transience and compactness conditions, and prove a Safety Sandwich: the RATTL value lies between the uninformed robust value and the full- knowledge optimum, with a gap that vanishes as the posterior concentrates. In a canonical binary-hazard instance, the induced criterion reduces to Conditional Value-at-Risk at a level set by the posterior entropy. A worked example shows the agent deferring the efficient action until a sharp identification threshold. RATTL targets runtime safety for agents, including LLM-based systems, acting under uncertainty.

摘要:代理在學習其環境時應該多謹慎?我們提出了RATTL(風險對抗總回報學習),它將謹慎與認知不確定性聯繫起來:代理對未知動態持有貝葉斯後驗,並根據一個其半徑是該後驗單調函數的Wasserstein模糊集進行規劃。隨著證據的增加,半徑會收縮,因此行為在最壞情況的穩健性和風險中立的總回報最大化之間持續插值。該設計遵循了熵值風險的對偶性,將風險水平的選擇轉化為模糊半徑的選擇。我們顯示,所得到的規劃問題在瞬態和緊湊性條件下是良好定義的,並證明了一個安全三明治:RATTL值介於無信息穩健值和全知最優值之間,當後驗集中時,這一差距消失。在一個典型的二元危險實例中,所引入的標準簡化為在後驗熵設定的水平下的條件風險價值。一個具體的例子顯示,代理在達到明確識別閾值之前推遲了有效行動。RATTL針對在不確定性下行動的代理,包括基於LLM的系統,目標是運行時安全。

DMT-Dens: Density-preserving manifold visualization for biological data

2608.17571v1 by Ruizhe Wang, Yixuan Dong, Bolin Yang, Bingo Wing-Kuen Ling, Fuji Yang, Zelin Zang

Motivation: Low-dimensional embeddings are widely used to explore cell-state heterogeneity in single-cell and other high-dimensional biological data. Although many methods preserve local neighborhoods, they may distort the apparent sampling density of processed observations, altering the visual contrast between dense and sparse regions and complicating the interpretation of rare, transitional, or continuous cell-state populations. Results: We present DMT-Dens, a parametric manifold-visualization method built on a latent-token Transformer encoder. The model integrates rank-based manifold alignment with hard-pair aggregation. To preserve density, it optimizes a loss based on the Pearson correlation between k-nearest-neighbor log-radius estimates in the processed input and two-dimensional embedding spaces. Benchmark evaluations demonstrate strong density preservation, particularly on biological datasets, while retaining competitive label separability. Availability: Source code, data-processing scripts, and resolved experiment configurations are available at https://github.com/Ruizhe-wang/DMT-Dens.

摘要:動機:低維嵌入被廣泛用於探索單細胞及其他高維生物數據中的細胞狀態異質性。儘管許多方法保留了局部鄰域,但它們可能會扭曲處理觀察的表觀取樣密度,改變密集區域和稀疏區域之間的視覺對比,並使得對稀有、過渡或連續細胞狀態群體的解釋變得複雜。結果:我們提出了DMT-Dens,一種基於潛在標記Transformer編碼器的參數流形可視化方法。該模型將基於排名的流形對齊與硬配對聚合相結合。為了保留密度,它優化了一個基於處理輸入和二維嵌入空間中k最近鄰對數半徑估計之間的Pearson相關性的損失。基準評估顯示出強大的密度保留能力,特別是在生物數據集上,同時保持競爭性的標籤可分性。可用性:源代碼、數據處理腳本和解決的實驗配置可在https://github.com/Ruizhe-wang/DMT-Dens獲得。

Domain-Adapted Molecular Language Models for Efficient Search of Make-on-Demand Libraries

2608.17567v1 by Henrik Wille, Luis-Finley Schütz, Felix Strieth-Kalthoff

Pretrained molecular language models are increasingly used as molecular encoders for learning structure-property relationships. However, their practical suitability for molecular discovery within and beyond their pretraining domain remains unclear. Herein, we systematically benchmark four molecular language models across six virtual molecular libraries spanning drug discovery, organic materials, and catalysis. Native molecular language model embeddings show substantial variation in discovery performance across libraries, whereas molecular fingerprints provide a consistently strong and robust baseline. Consistent with a potential domain-representation mismatch, we show that explicit domain adaptation substantially improves representation performance. Fine-tuning molecular language model encoders on structures from the target virtual library consistently improves sample efficiency, with several adapted encoders emerging as the top-performing representations across the benchmark tasks. These results show that molecular representation quality depends strongly on the target domain and that explicit adaptation can improve the practical utility of molecular foundation models. More broadly, our findings establish domain-adapted molecular representations as a promising strategy for sample-efficient adaptive decision making in virtual screening and self-driving laboratories.

摘要:預訓練的分子語言模型越來越多地被用作學習結構-性質關係的分子編碼器。然而,它們在其預訓練領域內外的分子發現中的實際適用性仍不明朗。在此,我們系統性地基準測試了四種分子語言模型,涵蓋了六個虛擬分子庫,涉及藥物發現、有機材料和催化。原生的分子語言模型嵌入在不同庫中的發現性能顯示出顯著的變化,而分子指紋則提供了一個一致強大且穩健的基準。與潛在的領域表示不匹配一致,我們顯示明確的領域適應顯著改善了表示性能。在目標虛擬庫的結構上微調分子語言模型編碼器,始終提高了樣本效率,其中幾個適應後的編碼器在基準任務中表現為最佳的表示。這些結果顯示,分子表示的質量強烈依賴於目標領域,而明確的適應可以提高分子基礎模型的實際效用。更廣泛地說,我們的發現確立了領域適應的分子表示作為在虛擬篩選和自駕實驗室中進行樣本高效自適應決策的一種有前景的策略。

Where a New Concept Must Enter: Entry Point Gates Cross-Task Usability in Unified Multimodal Models

2608.17564v1 by Zongyang Qiu, Yihan Wu, Kaixuan Fan, Bo Li, Hui Xiong

Unified multimodal models (UMMs) are motivated by the hope that understanding and generation reinforce each other but controlled ablations repeatedly find that adding a generation objective leaves understanding flat. Joint-training studies cannot settle the disagreement: with overlapping supervision, a gain cannot be attributed to the architecture rather than the data. To further investigate the relationship between the two directions in UMMs, we separate them by construction. A novel visual entity, a rendered 3D asset paired with a pseudo-word screened for absence from the frozen model's behavior, is bound through exactly one task direction, and the untrained direction is then measured. We find that the channel is real in both directions, but the directions differ in kind: generation training installs a name the model can only match among candidates; understanding training installs one it can also produce. What governs cross-task usability is where the binding enters the shared computation. An alignment probe predicts export across 36 configurations (Spearman $ρ= +0.68$). That objective's alignment term, maximized in closed form over activations with every weight frozen, makes a concept drawable when injected at layer 7 of 28 and is indistinguishable from the base model from layer 14 on, while the weight-based version of the same edit peaks at layers 10-14. In an observational series of four models, this window appears only where the understanding pathway is a semantic vision encoder, suggesting that unified weights are not enough: the two directions must share a semantic format at the entry point. Exploiting the rule, a mid-stack alignment objective acquires the concept for a $0.1\%$ relative loss of the model's general text-to-image ability, against $41\%$ for the standard generative route. Our code is at https://github.com/Zane-ZYQiu/entry-point-umm.

摘要:統一的多模態模型(UMMs)是受到理解與生成相互增強的希望所驅動,但控制性消融實驗反覆發現,添加生成目標會使理解保持平坦。聯合訓練研究無法解決這一分歧:在重疊監督下,增益無法歸因於架構而不是數據。為了進一步研究UMMs中這兩個方向之間的關係,我們通過構造將它們分開。一個新穎的視覺實體,即一個渲染的3D資產,與一個經過篩選以確保不出現在凍結模型行為中的偽詞配對,通過恰好一個任務方向綁定,然後測量未訓練的方向。我們發現這個通道在兩個方向上都是實際存在的,但這些方向在性質上有所不同:生成訓練安裝了一個模型只能在候選者中匹配的名稱;理解訓練則安裝了一個模型也可以生成的名稱。跨任務可用性的主導因素是綁定進入共享計算的地方。一個對齊探針預測在36個配置下的輸出(Spearman $ρ= +0.68$)。該目標的對齊項在凍結每個權重的情況下,對激活進行閉合形式的最大化,使得在28層的第7層注入時,概念可被繪製,並且從第14層開始與基礎模型無法區分,而同一編輯的基於權重的版本在第10-14層達到峰值。在一系列觀察四個模型的實驗中,這一窗口僅在理解路徑為語義視覺編碼器時出現,這表明統一權重並不夠:這兩個方向必須在進入點共享一個語義格式。利用這一規則,中堆棧對齊目標以$0.1\%$的相對損失獲得了該概念,這相對於標準生成路徑的$41\%$。我們的代碼位於 https://github.com/Zane-ZYQiu/entry-point-umm。

Reflex-Guard: A Low-Latency Guardrail for LLM Prompt Safety Using Dense Semantic Embeddings

2608.17556v1 by Istiaque Ahmed, Afia Anjum Borsha, Ranat Das Prangon, Abu-fuad Ahmad, Thi Hong Tran

Large Language Models (LLMs) in real-world applications often face the risks of specially crafted prompts designed to bypass the safety controls. Existing guardrail methods, such as LLM-as-a-judge and cloud-based safety APIs are able to detect unsafe content. However, they often add a delay of about 250-900 ms to each request. This delay is too high for real-time applications, when the system usually needs to respond in less than 100 ms. Furthermore, routing user prompts through external moderation endpoints raises significant data privacy concerns. This paper introduces Reflex-Guard, a lightweight guardrail that runs locally. It uses jailbreak-aware preprocessing, compact sentence-transformer embeddings, and seven fast binary classifiers. Together, these components enable high-accuracy prompt safety filtering with much lower latency than existing solutions. Through systematic evaluation on a strategically balanced dataset of 30,568 samples drawn from five complementary sources, we demonstrate that Reflex-Guard achieves 95.9% recall on harmful prompts at 37.6 ms end-to-end latency. It is faster than existing baselines, including Llama Guard 2 at 255 ms and SafeDecoding at 723 ms. It can detect 100% of GCG suffix attacks and Base64-encoded prompts using the default threshold. However, DrAttack structured prompts required lowering the threshold to 0.03 for optimal detection, as they produced a distinct probability distribution. Reflex-Guard achieves Reflex Efficiency Score (RES) scores up to 16.79, significantly outperforming Llama Guard 2 (11.90) and SafeDecoding (9.80). This analysis offers practical deployment advice and shows that different attack types occupy distinct regions in the embedding probability space.

摘要:大型語言模型(LLMs)在實際應用中常常面臨專門設計的提示風險,這些提示旨在繞過安全控制。現有的防護方法,如LLM作為評判者和基於雲的安全API,能夠檢測不安全的內容。然後,它們通常會為每個請求增加約250-900毫秒的延遲。這個延遲對於實時應用來說過高,因為系統通常需要在100毫秒內作出回應。此外,通過外部審核端點路由用戶提示會引發重大數據隱私問題。本文介紹了Reflex-Guard,一種輕量級的本地防護措施。它使用監獄破解感知的預處理、緊湊的句子轉換器嵌入和七個快速的二元分類器。這些組件共同實現了高準確度的提示安全過濾,延遲遠低於現有解決方案。通過對來自五個互補來源的30,568個樣本的戰略性平衡數據集進行系統評估,我們證明Reflex-Guard在有害提示上達到了95.9%的召回率,端到端延遲為37.6毫秒。它比現有的基準更快,包括Llama Guard 2的255毫秒和SafeDecoding的723毫秒。它可以使用默認閾值檢測100%的GCG後綴攻擊和Base64編碼的提示。然而,DrAttack結構化提示需要將閾值降低到0.03以達到最佳檢測,因為它們產生了不同的概率分佈。Reflex-Guard的反射效率得分(RES)高達16.79,顯著超過Llama Guard 2(11.90)和SafeDecoding(9.80)。這一分析提供了實際部署建議,並顯示不同的攻擊類型在嵌入概率空間中佔據不同的區域。

Code as Representation: A Compilable Parsing Paradigm for Academic Documents

2608.17550v1 by Rihui Jin, Jun Wang, chengyuan zhu, Liang Mingyu, Yue Gao, Li Yunxuan, Kuicai Dong, Guilin Qi, Lin Ren, Yongrui Chen, Xinbang Dai, Jiaqi Li, Tongtong Wu, Gholamreza Haffari

Academic papers are a primary carrier of scientific knowledge, yet most of this knowledge remains locked in PDFs that are optimized for human reading rather than machine use. For Multimodal Large Language Models (MLLMs), the core challenge is not only perception, but representation: scientific pages interleave text with Structured Academic Elements (SAEs) such as tables, formulas, charts, and pseudocode, whose structure, data, and logic are poorly preserved by common surrogates like Markdown. We therefore propose Compilable Academic Document Parsing (CADP), a paradigm that reconstructs a full page as contextual \LaTeX{} plus executable Python, so that structure-preserving elements and executable chart representations can be reconstructed, recompiled, and directly verified against the source page. To support this setting, we introduce CADP-Bench, an expert-verified benchmark of full academic pages containing tightly coupled text and multiple SAE types, evaluated through a re-injection compilation protocol. We further study current capabilities using SOTA MLLMs and an exploratory multi-agent baseline that incorporates common agentic techniques. Results show that even frontier models still struggle to produce high-fidelity executable reconstructions, highlighting substantial room for improvement in structure-aware scientific document parsing. CADP-Bench is released for future research.

摘要:學術論文是科學知識的主要載體,但大部分這些知識仍然鎖定在優化為人類閱讀而非機器使用的PDF中。對於多模態大型語言模型(MLLMs)來說,核心挑戰不僅在於感知,還在於表徵:科學頁面將文本與結構化學術元素(SAEs)交錯,如表格、公式、圖表和偽代碼,其結構、數據和邏輯在常見的替代品如Markdown中保存得很差。因此,我們提出可編譯學術文檔解析(CADP),這是一種將整個頁面重建為上下文 \LaTeX{} 加上可執行的Python的範式,以便結構保留的元素和可執行的圖表表示可以被重建、重新編譯並直接與源頁面進行驗證。為了支持這一設置,我們引入CADP-Bench,一個經專家驗證的完整學術頁面基準,包含緊密耦合的文本和多種類型的SAE,通過重新注入編譯協議進行評估。我們進一步研究使用SOTA MLLMs的當前能力以及一個探索性的多代理基準,該基準結合了常見的代理技術。結果顯示,即使是最前沿的模型仍然難以產生高保真度的可執行重建,突顯出結構感知的科學文檔解析有很大的改進空間。CADP-Bench已經釋出以供未來研究使用。

No Gaussian Required: Contrastive Inverse Dynamics for JEPA World Models

2608.17542v1 by Jack Boylan, Chris Hokamp

Joint-Embedding Predictive Architectures (JEPAs) learn world models by predicting future embeddings, but the objective admits a trivial solution of a constant encoder, so every practical system adds an anti-collapse mechanism (LeCun, 2022; Assran et al., 2023; Bardes et al., 2022; 2024). LeWorldModel (LeWM) prevents collapse with SIGReg, a regularizer that forces the latent distribution to match an isotropic Gaussian: the representation is stabilized by prescribing what it must look like, independently of the environment it models. We argue that the anti-collapse pressure can instead come from the transition data itself. Action-Contrastive Masked Transition Modeling (AC-MTM) keeps LeWM's forward latent-prediction objective and adds a training-only inverse-dynamics head trained with Action-NCE: each latent transition must identify the action that produced it among the other actions in the batch, a discrimination task that a collapsed encoder provably fails. The inverse branch is discarded after training, leaving test-time encoding, forward prediction, planning, and compute identical to LeWM. On four standard pixel-control tasks under a matched planning protocol, AC-MTM trains stably from scratch and matches SIGReg on average. On the harder multi-object OGBench Visual Scene task, results are consistent with the prescribed geometry becoming a bottleneck: AC-MTM reaches 80.0$\pm$2.0% success versus 58.0$\pm$2.0% for SIGReg, improving by 20-24 points in each training seed. A single 50-episode random-policy run gives a 52% baseline estimate. Contrastive inverse dynamics thus provides a distribution-free anti-collapse signal that requires no target network, stop-gradient, pretrained encoder, or reconstruction objective, and we characterize the action-space and observability assumptions under which it holds. We make our code available at https://github.com/jackboyla/action-contrastive-jepa

摘要:聯合嵌入預測架構(JEPAs)透過預測未來嵌入來學習世界模型,但該目標允許一個恆定編碼器的平凡解,因此每個實際系統都添加了一個反崩潰機制(LeCun, 2022; Assran et al., 2023; Bardes et al., 2022; 2024)。LeWorldModel(LeWM)透過SIGReg防止崩潰,這是一種正則化器,強迫潛在分佈與各向同性高斯匹配:該表示通過規定其必須的樣貌來穩定,無論其所建模的環境如何。我們認為反崩潰壓力可以來自於過渡數據本身。行動對比遮罩過渡建模(AC-MTM)保持LeWM的前向潛在預測目標,並添加一個僅訓練的逆動力頭,該頭使用行動-NCE進行訓練:每個潛在過渡必須在批次中的其他行動中識別出產生它的行動,這是一個崩潰編碼器顯然無法完成的區分任務。逆分支在訓練後被丟棄,留下測試時的編碼、前向預測、規劃和計算與LeWM相同。在四個標準像素控制任務中,根據匹配的規劃協議,AC-MTM從零開始穩定訓練,並在平均上匹配SIGReg。在更困難的多物體OGBench視覺場景任務中,結果與規定的幾何形狀成為瓶頸的情況一致:AC-MTM達到80.0$\pm$2.0%的成功率,而SIGReg則為58.0$\pm$2.0%,在每個訓練種子中提高了20-24點。一個50集隨機策略的運行給出了52%的基線估計。因此,對比逆動力提供了一個無分佈的反崩潰信號,無需目標網絡、停止梯度、預訓練編碼器或重建目標,我們還描述了其成立的行動空間和可觀察性假設。我們的代碼可在https://github.com/jackboyla/action-contrastive-jepa獲得。

2608.17536v1 by Jin Su, Zhuofeng Zhao, Huanhuan Wang, Hao Chen

Legal consultation questions exhibit multi-level complexity. A single retrieval strategy often leads to over-reasoning for simple questions and poor interpretability for complex ones, making it difficult to meet the requirements for both answer quality and efficiency in high-risk scenarios. To address this issue, this paper proposes CoAL-RAG, a complexity-aware legal retrieval-augmented generation method, which constructs a multi-dimensional evaluation mechanism based on question essence'' andretrieval consistency'' to enable adaptive routing of retrieval strategies. First, the reasoning demand is quantified according to the logical structure of the question. Then, the discrepancy between semantic retrieval and keyword retrieval is utilized to indirectly reflect problem complexity, thereby selecting the most appropriate retrieval strategy and dynamically filtering contextual information. Experimental results demonstrate that the proposed method significantly outperforms baseline models not only on Chinese legal benchmarks (SocialLawQA, LawBench) but also demonstrates strong cross-jurisdictional generalization on English datasets (LexGLUE, CaseHold). Specifically, on Chinese datasets, the BLEU score improves by 42.5\% and ROUGE-L reaches 3.6 times that of knowledge graph-based methods. On English benchmarks, CoAL-RAG maintains highly competitive accuracy, achieving an optimal balance between generation quality, deep logical reasoning, and system efficiency across different legal systems.

摘要:法律諮詢問題展現出多層次的複雜性。單一的檢索策略常常導致對簡單問題的過度推理,以及對複雜問題的可解釋性差,使得在高風險情境中難以滿足答案質量和效率的要求。為了解決這個問題,本文提出了 CoAL-RAG,一種具複雜性意識的法律檢索增強生成方法,該方法基於「問題本質」和「檢索一致性」構建了一個多維評估機制,以實現檢索策略的自適應路由。首先,根據問題的邏輯結構量化推理需求。然後,利用語義檢索與關鍵字檢索之間的差異,間接反映問題的複雜性,從而選擇最合適的檢索策略並動態過濾上下文信息。實驗結果表明,所提出的方法在中國法律基準(SocialLawQA、LawBench)上顯著超越基線模型,並且在英語數據集(LexGLUE、CaseHold)上展現出強大的跨法域泛化能力。具體而言,在中國數據集上,BLEU 分數提高了 42.5\%,而 ROUGE-L 達到知識圖譜方法的 3.6 倍。在英語基準上,CoAL-RAG 維持了高度競爭的準確性,在不同法律系統中實現生成質量、深度邏輯推理和系統效率之間的最佳平衡。

ArborMem: Navigating Interaction States with Memory Forests

2608.17534v1 by Zongwei Lv, Yuemeng Xu, Yilun Yao, Siyi Ding, Xinyu Tan, Yaoming Li, Guangxiang Zhao, Weihong Lin, Lin Sun, Xiangzheng Zhang, Tong Yang

Large language models increasingly serve as persistent conversational assistants, requiring memory that preserves relevant experience and maintains continuity across interactions. Existing methods improve access to conversational history through long-context processing, selective retrieval, and structured memory organization. However, most systems treat memory access as retrieving relevant past information without first determining which prior interaction state the current turn resumes. This limitation becomes particularly important when conversations interleave multiple tasks, people, and plans that may be interrupted and later revisited. We introduce ArborMem, an online memory framework that represents a long-running conversation as a navigable forest of interaction states. Each branch preserves a locally coherent trajectory, while the forest maintains multiple trajectories that may later be resumed. For each new input, ArborMem localizes the relevant state, restores its branch-local context, and augments it with reusable evidence retrieved across branches, preserving interaction continuity without conflating semantically related but structurally distinct trajectories. Existing long-term memory benchmarks cover diverse memory and reasoning capabilities but do not explicitly isolate branch-structured challenges. We therefore introduce BranchMemEval, a controlled diagnostic benchmark for interleaved and resumable interaction trajectories. Experiments on LongMemEval, LoCoMo, BEAM 100K, and BranchMemEval show that ArborMem outperforms the strongest baselines by 3.36 to 10.31 percentage points on the three established benchmarks and by 5.0 points on BranchMemEval. Its advantage grows under constrained read budgets, while complete memory queries remain below half a second.

摘要:大型語言模型越來越多地作為持久的對話助手,這需要記憶來保留相關經驗並在互動中保持連貫性。現有的方法通過長上下文處理、選擇性檢索和結構化記憶組織來改善對對話歷史的訪問。然而,大多數系統將記憶訪問視為檢索相關的過去信息,而不首先確定當前回合恢復的先前互動狀態。當對話交織著多個任務、人物和計劃,這些任務可能會被中斷並在稍後重新訪問時,這一限制變得尤為重要。我們介紹了 ArborMem,一個在線記憶框架,將長期對話表示為可導航的互動狀態森林。每個分支保留一個局部一致的軌跡,而森林則維護多條可能稍後恢復的軌跡。對於每個新的輸入,ArborMem 定位相關狀態,恢復其分支局部上下文,並通過跨分支檢索的可重用證據進行增強,保留互動的連續性而不混淆語義上相關但結構上不同的軌跡。現有的長期記憶基準涵蓋了多樣的記憶和推理能力,但並未明確隔離分支結構挑戰。因此,我們引入了 BranchMemEval,一個針對交織和可恢復互動軌跡的受控診斷基準。在 LongMemEval、LoCoMo、BEAM 100K 和 BranchMemEval 上的實驗顯示,ArborMem 在三個既定基準上比最強基線高出 3.36 到 10.31 個百分點,在 BranchMemEval 上高出 5.0 個百分點。其優勢在受限的讀取預算下增長,而完整的記憶查詢仍保持在半秒以下。

When to Review: Spaced Repetition for Continual Pre-Training of Language Models

2608.17530v1 by Alankar Atreya, Devesh Batra, Yoages Kumar Mantri, Geremy Bantug, Greig A Cowan, Raad Khraishi

Continual pre-training of large language models must acquire new information without erasing old knowledge. Existing replay methods often choose a global old/new mixture and sample uniformly, ignoring that examples differ in how quickly they are forgotten. We formulate continual pre-training as adaptive review scheduling: the training loop should decide not only how much history to replay, but which examples should return at each step. We introduce Spaced Repetition Training (SRT), a continual learning framework inspired by cognitive science, which schedules sample-rehearsal using the SuperMemo-2 (SM-2) algorithm. SRT maintains per-example review state, maps per-example perplexity to a recall-quality signal, and schedules historical examples for retention and new examples for consolidation while leaving the model, objective, and optimizer unchanged. On temporally separated Wikipedia and code corpora, SRT improves the stability-plasticity trade-off, recovering 5 to 37 percentage points of old-knowledge accuracy lost by naive continual pre-training across model scales while preserving or improving new-knowledge acquisition. At larger scale, SRT preserves broad benchmark performance that naive continual pre-training and uniform replay substantially degrade. Experiments with vision and tabular data further suggest that the scheduling principle extends beyond language when paired with an appropriate recall signal.

摘要:持續的預訓練大型語言模型必須在不抹去舊知識的情況下獲取新信息。現有的重播方法通常選擇一個全局的舊/新混合並均勻抽樣,忽略了示例在被遺忘的速度上存在差異。我們將持續預訓練公式化為自適應回顧排程:訓練循環應決定不僅是重播多少歷史,還有每一步應該返回哪些示例。我們引入了間隔重複訓練(SRT),這是一個受認知科學啟發的持續學習框架,使用 SuperMemo-2 (SM-2) 算法來排程樣本重複。SRT 維持每個示例的回顧狀態,將每個示例的困惑度映射到回憶質量信號,並在保留模型、目標和優化器不變的情況下,為保留歷史示例和鞏固新示例進行排程。在時間上分隔的維基百科和代碼語料庫上,SRT 改善了穩定性與可塑性的權衡,恢復了由天真的持續預訓練在各模型規模上損失的 5 到 37 個百分點的舊知識準確率,同時保留或改善了新知識的獲取。在更大規模下,SRT 保持了廣泛的基準性能,而天真的持續預訓練和均勻重播則大幅降低了這一性能。對於視覺和表格數據的實驗進一步表明,當與適當的回憶信號配對時,排程原則超越了語言的範疇。

Agent Lightning v1.0: Towards Harnessed Agentic RL

2608.17528v1 by Zhiyuan He, Siwei Zhang, Zhiwen Zhou, Yuqing Yang, Yu Kang, Yuge Zhang, Luna K. Qiu, Tin Yan Tsui, Jiahang Xu, Chong Luo

Modern agents operate inside agent harnesses that manage tools, context, and control flow, making the harness a critical part of the agent system. Our original Agent Lightning introduced a disaggregated architecture that connects arbitrary agents to RL training through an LLM endpoint proxy, an approach later adopted by frameworks such as verl Uni-Agent, AReaL 2.0, slime, and Polar. We refer to this paradigm as harnessed agentic RL, where the deploy-time harness directly participates in model post-training. Harnessed agentic RL differs fundamentally from traditional agentic RL: the harness, rather than the training engine, owns the environment interaction loop, while the trainer observes only sequences of LLM request-response pairs. This introduces challenges in retokenization, sample merging, advantage calculation, loss normalization, and backend scheduling, which can substantially affect training stability and effectiveness. We present Agent Lightning v1.0, a lightweight framework for harnessed agentic RL implemented in approximately 3,500 lines of code. It supports arbitrary agent harnesses and serves as a practical testbed for studying these challenges. We evaluate it on instruction-following, search, and coding agents, and provide a complete reproducible pipeline for coding-agent RL. Using only 6K training examples and modest compute, RL improves Qwen3.5-9B on SWE-bench Verified from 41.8% to 56.4%, a 14.6-point absolute gain. We release the complete workflow and training scripts to facilitate reproducible research on harnessed agentic RL.

摘要:現代代理人運行在管理工具、上下文和控制流程的代理人鞍具內,使得鞍具成為代理人系統中的關鍵部分。我們的原始 Agent Lightning 引入了一種解耦架構,通過 LLM 端點代理將任意代理人連接到強化學習訓練,這種方法後來被如 verl Uni-Agent、AReaL 2.0、slime 和 Polar 等框架採用。我們將這種範式稱為鞍具代理強化學習,其中部署時的鞍具直接參與模型的後訓練。鞍具代理強化學習在根本上與傳統的代理強化學習不同:鞍具,而不是訓練引擎,擁有環境互動循環,而訓練者僅觀察 LLM 請求-回應對的序列。這在重新標記、樣本合併、優勢計算、損失正規化和後端調度方面引入了挑戰,這些挑戰可能會顯著影響訓練的穩定性和有效性。我們提出了 Agent Lightning v1.0,這是一個輕量級的鞍具代理強化學習框架,實現約 3,500 行代碼。它支持任意的代理人鞍具,並作為研究這些挑戰的實用測試平台。我們在遵循指令、搜索和編碼代理人上對其進行評估,並提供完整的可重現管道以進行編碼代理人強化學習。僅使用 6K 訓練範例和適度的計算,強化學習將 Qwen3.5-9B 在 SWE-bench Verified 上的表現從 41.8% 提升至 56.4%,絕對增益為 14.6 點。我們發布完整的工作流程和訓練腳本,以促進對鞍具代理強化學習的可重現研究。

Explainable AI-Powered Framework for Video-Based Skill Assessment in Cataract Surgery

2608.17522v1 by Mohammad Javad Ahmadi, Hamid D. Taghirad

Persistent shortages in the surgical workforce and inherent limitations of traditional training methods highlight the necessity of automated, data-driven approaches in surgical education. This study addresses these challenges by introducing a novel, explainable AI-powered framework for automated skill assessment, specifically focusing on cataract surgery. We present the world's largest dataset of cataract surgery videos, comprising 2,000 recordings. Additionally, we propose an AI-powered analytical framework that employs advanced computer vision and signal-processing techniques to automatically evaluate surgical videos to derive objective, quantitative performance indicators that complement or potentially replace subjective scoring methods. A significant advantage of our framework over previous methods lies precisely in its explainability of outputs, elevating it beyond merely an opaque skill classification tool. Through experimental analysis of 83 cataract surgery videos, we demonstrate that the automatically computed metrics exhibit strong correlations with expert-based subjective evaluations, achieving up to 87% accuracy in surgical skill assessment. Each metric was individually examined, and expert surgeons provided subjective ratings using the newly introduced Capsulorhexis Skill Assessment System (CSAS). These subjective assessments were compared with ten objective motion-based metrics extracted through our framework. The results indicated a robust correlation between subjective ratings and automated indicators, underscoring the framework's capacity to accurately model surgical expertise.

摘要:持續的外科醫療人力短缺以及傳統訓練方法的固有限制凸顯了在外科教育中自動化、數據驅動方法的必要性。這項研究通過引入一個新穎的、可解釋的人工智慧驅動框架來解決這些挑戰,特別專注於白內障手術。我們展示了世界上最大的白內障手術視頻數據集,包含2,000個錄像。此外,我們提出了一個人工智慧驅動的分析框架,利用先進的計算機視覺和信號處理技術,自動評估手術視頻,以獲得客觀的、定量的性能指標,這些指標可以補充或潛在地取代主觀評分方法。我們的框架相較於先前的方法的一個顯著優勢恰恰在於其輸出的可解釋性,使其超越僅僅是一個不透明的技能分類工具。通過對83個白內障手術視頻的實驗分析,我們證明自動計算的指標與專家基於主觀評估的評分之間存在強烈的相關性,在外科技能評估中達到高達87%的準確率。每個指標都經過單獨檢查,專家外科醫生使用新引入的囊膜切開技能評估系統(CSAS)提供主觀評分。這些主觀評估與通過我們的框架提取的十個客觀運動基礎指標進行了比較。結果顯示主觀評分與自動指標之間存在穩健的相關性,強調了該框架準確建模外科專業知識的能力。

Effects of Answer Format Variation on Gender Bias in Large Language Models

2608.17516v1 by Ksenia Merzlyakova, Sebastian Padó, Franziska Weeber

Gender bias or other social biases in large language models (LLMs) are frequently evaluated with question answering or survey benchmarks where the LLM needs to give a response in a predefined answer format. It is well known in survey science that the answer format has a substantial impact on answers, just as LLMs are sensitive to the prompt wording. However, to our knowledge it has not been studied yet how changes in answer format impact the measurement of gender bias in LLMs and their alignment with human response distributions. We evaluate three instruction-tuned models on the BBQ benchmark and OpinionQA survey data across closed-ended, Likert-scaled and open-ended formats, comparing bias measurement and distributional alignment under otherwise identical conditions. We find that answer format does substantially alter measured outcomes, including reversals in order rankings. These differences arise because each format elicits distinct response behaviours, such as forced-choice selection, scale-based distributions and refusal in free-text generation. Our findings highlight the importance of treating answer format as a substantive component of LLM evaluation and motivate multi-format designs for more robust model assessment.

摘要:性別偏見或其他社會偏見在大型語言模型(LLMs)中的評估,通常使用問答或調查基準,其中LLM需要以預定的答案格式給出回應。調查科學中已知答案格式對答案有重大影響,就像LLMs對提示措辭敏感一樣。然而,據我們所知,尚未研究答案格式的變化如何影響LLMs中性別偏見的測量及其與人類回應分佈的一致性。我們在BBQ基準和OpinionQA調查數據上評估了三個經過指令調整的模型,並比較了在封閉式、Likert量表和開放式格式下的偏見測量和分佈一致性,條件則保持一致。我們發現答案格式確實顯著改變了測量結果,包括排序排名的逆轉。這些差異的產生是因為每種格式引發了不同的回應行為,例如強制選擇、基於量表的分佈和在自由文本生成中的拒絕。我們的發現強調將答案格式視為LLM評估的實質性組成部分的重要性,並促使多格式設計以進行更穩健的模型評估。

2608.17515v1 by Enrique Barba Roque, Luís Cruz, Annibale Panichella

Background: Large Language Models (LLMs) are increasingly being applied to Software Engineering (SE) tasks, achieving high accuracy across problems such as clone detection, vulnerability prediction, and code summarization. However, their high computational demands and energy consumption raise sustainability concerns and hinder their use on consumer hardware and resource-constrained platforms. A common way to report the computational cost of an LLM in the literature and industry is to use the number of Floating Point Operations (FLOPs) required to perform a pass over the network. Aims: This paper investigates the implications of energy-aware knowledge distillation for SE, aiming to improve model efficiency while maintaining performance and to determine whether FLOPs is a reliable energy-aware metric. Method: We conduct a controlled experiment using Morph, a Many-Objective Optimization-based distillation methodology, to empirically examine whether FLOPs accurately reflect energy consumption in Clone Detection and Vulnerability Prediction tasks. We extend this methodology to include energy-surrogate models that directly estimate CPU and GPU energy consumption during optimization, and we apply Morph to generative tasks using CodeT5+ for code summarization. Results: Our results show that FLOPs is not always a reliable indicator of energy consumption, and better results can be achieved by using energy-surrogate models. Distilled student models can reduce inference energy consumption by up to 90\% and memory usage by 86\%, with only modest accuracy trade-offs. Conclusions: Energy-aware knowledge distillation when guided by direct energy surrogates rather than FLOPs can improve the energy consumption, sustainability, and deployability of LLMs for SE applications, enabling efficient models on consumer hardware.

摘要:背景:大型語言模型(LLMs)越來越多地應用於軟體工程(SE)任務,在克隆檢測、漏洞預測和程式碼摘要等問題上達到了高準確率。然而,它們的高計算需求和能量消耗引發了可持續性問題,並阻礙了它們在消費者硬體和資源受限平台上的使用。在文獻和業界中,報告LLM計算成本的常見方法是使用執行一次網絡所需的浮點運算次數(FLOPs)。目標:本文探討了對SE進行能量感知知識蒸餾的影響,旨在提高模型效率的同時保持性能,並確定FLOPs是否是一個可靠的能量感知指標。方法:我們使用Morph進行了一項受控實驗,這是一種基於多目標優化的蒸餾方法,實證檢驗FLOPs是否準確反映克隆檢測和漏洞預測任務中的能量消耗。我們擴展了這一方法,納入能量替代模型,這些模型在優化過程中直接估算CPU和GPU的能量消耗,並將Morph應用於使用CodeT5+進行程式碼摘要的生成任務。結果:我們的結果顯示FLOPs並不總是能可靠指示能量消耗,使用能量替代模型可以獲得更好的結果。蒸餾的學生模型可以將推理能量消耗降低高達90%,內存使用量降低86%,而準確率僅有適度的折衷。結論:在直接能量替代模型的指導下,能量感知知識蒸餾可以改善LLMs在SE應用中的能量消耗、可持續性和可部署性,從而使消費者硬體上的模型更加高效。

SGHA: Evidence-Grounded Research Problem Discovery with Local Language Models

2608.17501v1 by Sarvesh Gharat, Junpei Komiyama

Recent efforts toward fully automated AI scientists have demonstrated that language-model agents can generate hypotheses, execute experiments, and draft scientific manuscripts. However, during the early stages of research, when research problems are formulated, these AI scientists often rely heavily on proprietary frontier models. Their proposals are shaped by opaque parametric knowledge and by literature searches conditioned on the proposals themselves. Such knowledge is effectively a black box, and this dependence makes the evidential basis and validity of generated research problems difficult to audit and leaves the process vulnerable to model-specific hallucinations and biases. Furthermore, if proprietary research materials are transmitted to external APIs, the use of these models creates confidentiality, privacy, and data-governance concerns. We introduce the Structural Gap Hypothesis Agent (SGHA), a fully automated, corpus-first research-problem discovery system that runs entirely on a local LLM. SGHA structures a scientific literature corpus into evidence-linked paper objects and a typed evidence graph, detects unresolved structural patterns across papers, screens candidate gaps before formulation, and produces traceable research-problem families. In particular, it is able to output assumptions, objectives, success criteria, and remaining ambiguities. All LLM-based components of SGHA are executed using a locally served open-weight 9B language model, without requiring proprietary frontier-model APIs. We compare SGHA with the AI Scientist-v2 idea formulation module in five machine-learning domains. Our results suggest that explicit corpus structure and evidence-constrained reasoning can support promising, inspectable research-problem formulation without relying on frontier models during generation or verification.

摘要:最近對於完全自動化的AI科學家的努力顯示,語言模型代理可以生成假設、執行實驗並撰寫科學手稿。然而,在研究的早期階段,當研究問題被形成時,這些AI科學家往往過度依賴專有的前沿模型。他們的提案受到不透明的參數知識和基於提案本身的文獻搜尋的影響。這種知識實際上是一個黑箱,而這種依賴使得生成的研究問題的證據基礎和有效性難以審核,並使過程容易受到模型特定的幻覺和偏見的影響。此外,如果專有研究材料被傳輸到外部API,使用這些模型會產生保密性、隱私和數據治理的問題。我們介紹了結構性差距假設代理(SGHA),這是一個完全自動化的、以語料庫為首的研究問題發現系統,完全在本地的LLM上運行。SGHA將科學文獻語料庫結構化為與證據相關聯的論文對象和類型化的證據圖,檢測論文之間未解決的結構模式,在形成之前篩選候選差距,並生成可追溯的研究問題家族。特別是,它能夠輸出假設、目標、成功標準和剩餘的模糊性。SGHA的所有基於LLM的組件都是使用本地提供的開放權重9B語言模型執行的,而不需要專有的前沿模型API。我們將SGHA與AI Scientist-v2的想法形成模塊在五個機器學習領域進行比較。我們的結果表明,明確的語料結構和基於證據的推理可以支持有前景的、可檢查的研究問題形成,而無需在生成或驗證過程中依賴前沿模型。

When AI Designs AI: Innovation or Imitation?

2608.17471v1 by Yikang Yang, Zhengxin Yang, Luzhou Peng, Minghao Luo, Yanqi Kan, Wanling Gao, Jianfeng Zhan

Recent advances in LLM agents have made them increasingly capable of designing methods for complex AI tasks. This raises two central questions about agent-designed methods relative to human-designed methods: how well they perform, and how different their algorithmic designs are. To study these questions, this paper introduces an analysis that derives task-specific algorithmic design spaces from human-designed methods, maps both human- and agent-designed methods into these spaces, and quantifies their algorithmic differences at the module level. Widely used LLM agents are evaluated on a suite of representative, open-ended AI tasks spanning multiple modalities, and the methods they design are analyzed in terms of both task performance and algorithmic differences from human-designed methods. Experimental results show that current agents can occasionally match or surpass human state-of-the-art (SOTA) performance (10/72 configurations), but such success does not generalize reliably across tasks or agents. Moreover, 96.8% of agent-designed methods fall within human-derived algorithmic design spaces, largely recombining algorithmic choices found in human-designed methods, while nearly half exactly match an existing human algorithmic design. Taken together, these findings suggest that although current agents can occasionally match or surpass human SOTA performance, their algorithmic designs remain within human-derived algorithmic design spaces, reflecting the reuse and recombination of algorithmic choices.

摘要:最近在大型語言模型(LLM)代理方面的進展使它們在設計複雜人工智慧任務的方法上變得越來越有能力。這引發了兩個關於代理設計的方法與人類設計的方法的核心問題:它們的表現如何,以及它們的算法設計有多不同。為了研究這些問題,本文介紹了一種分析方法,從人類設計的方法中推導出特定任務的算法設計空間,將人類和代理設計的方法映射到這些空間中,並在模塊層面量化它們的算法差異。廣泛使用的LLM代理在一系列具有代表性的開放式人工智慧任務中進行評估,這些任務涵蓋多種模式,並分析它們設計的方法在任務表現和與人類設計的方法的算法差異方面。實驗結果顯示,當前的代理偶爾可以匹配或超越人類的最先進(SOTA)表現(10/72配置),但這種成功並不可靠地在不同任務或代理之間泛化。此外,96.8%的代理設計的方法都落在由人類推導的算法設計空間內,主要是重新組合在人體設計的方法中找到的算法選擇,而近一半則完全匹配現有的人類算法設計。綜合來看,這些發現表明,儘管當前的代理偶爾可以匹配或超越人類的SOTA表現,但它們的算法設計仍然保持在由人類推導的算法設計空間內,反映了算法選擇的重用和重新組合。

SAGE: Self-Evolving Storyboard Skills via Attribution-Guided Rule Evolution

2608.17468v1 by Maolin Ran, Xiaoyang Lu, Jiaqi Liu, Jian Wang, Weiwen Liu, Jianghao Lin, Yong Yu, Weinan Zhang

Storyboards turn screenplays into visual shot plans for automated short drama production. Professional storyboarding relies on tacit directorial expertise and remains an industrial bottleneck. Large language models can automate this step, but methods for supplying directing knowledge face three challenges: (1) Knowledge acquisition: the craft remains implicit in exemplars or must be written manually. (2) Knowledge refinement: authored knowledge is not evaluated against execution outcomes, and opaque generation prevents feedback attribution to the knowledge behind each decision. (3) Knowledge injection: injecting all knowledge exceeds usable context, while manual selection for every narrative group does not scale. We present SAGE (Skill with Attribution-Guided Evolution), a deployed framework that learns, attributes, evolves, and routes directing knowledge from expert demonstrations. SAGE derives rules that are independent of episode content by contrasting each training screenplay with its expert storyboard. During generation, the model records each narrative group's adopted rules. Combining these records with localized feedback enables targeted updates to individual rules. Evolved rules form scenario packages with a routing index, so each group retrieves only a bounded set appropriate to its situation without expert intervention. On 18 test episodes across three genres, SAGE scored 77.8 on a rubric validated by experts, versus 77.1 for professional directors. Deployed for 14 days on Virtual Film Studio, SAGE produced 1,344 narrative group outputs; 87.2 percent were accepted without substantive edits, and the production team recorded over 83 percent less authoring time per episode. We release PROSE, the first public dataset pairing screenplays with storyboards by professional directors across 68 episodes: https://github.com/creDreams/PROSE.

摘要:故事板將劇本轉化為自動化短劇製作的視覺拍攝計劃。專業的故事板製作依賴於隱性導演專業知識,並且仍然是產業瓶頸。大型語言模型可以自動化這一步驟,但提供導演知識的方法面臨三個挑戰:(1)知識獲取:這項技藝仍然隱含於範例中或必須手動撰寫。(2)知識精煉:創作的知識未能根據執行結果進行評估,且不透明的生成過程阻礙了對每個決策背後知識的反饋歸因。(3)知識注入:注入所有知識超出了可用的上下文,而對每個敘事群體進行手動選擇則無法擴展。我們提出了SAGE(具歸因引導演變的技能),這是一個已部署的框架,從專家示範中學習、歸因、演變和路由導演知識。SAGE通過將每個訓練劇本與其專家故事板進行對比,推導出獨立於劇集內容的規則。在生成過程中,模型記錄每個敘事群體採用的規則。將這些記錄與本地反饋結合,使得對個別規則的針對性更新成為可能。演變的規則形成具有路由索引的場景包,因此每個群體僅檢索適合其情境的有限集合,而無需專家介入。在三個類型的18個測試劇集中,SAGE在專家驗證的評分標準上得分77.8,而專業導演則為77.1。在虛擬電影工作室部署14天後,SAGE產出了1,344個敘事群體的輸出;87.2%的輸出在未進行實質性編輯的情況下被接受,製作團隊每集的創作時間減少了超過83%。我們發布了PROSE,這是第一個將劇本與專業導演的故事板配對的公共數據集,涵蓋68個劇集:https://github.com/creDreams/PROSE。

From Entity Mentions to Tone: An LLM-Based Pipeline for Media Bias Analysis

2608.17454v1 by Klesti Hoxha, Olti Qirici

This paper presents a pipeline for analyzing media bias and framing in online news. The pipeline groups articles into topics and events, adds named-entity and sentiment annotations, and compares news sources through people mentions, source-level tone, and event-level coverage patterns. We apply it to 8,358 Albanian news articles collected from GDELT and compare the resulting annotations with GDELT's automated annotations. The results show moderate agreement for sentiment and entity extraction, as well as additional person-entity pairs that can potentially support the bias analysis. We compare two annotation prompts and find that stricter sentiment-validation rules remove label-score inconsistencies but increase execution time and reduce annotation coverage. Based on these results, the simpler prompt is used for the rest of the analysis. We have provided sample analysis on source-level framing pro les, person-level tone differences across sources, and event-level gatekeeping and coverage indicators. These outputs show how the same news collection can be used to examine what sources cover, how they describe public figures, and where coverage is concentrated. The approach is particularly useful in settings where manually verified datasets or specialized language tools are limited.

摘要:這篇論文提出了一個用於分析在線新聞中的媒體偏見和框架的流程。該流程將文章分組為主題和事件,添加命名實體和情感註釋,並通過人名提及、來源層級語調和事件層級報導模式來比較新聞來源。我們將其應用於從GDELT收集的8,358篇阿爾巴尼亞新聞文章,並將結果註釋與GDELT的自動註釋進行比較。結果顯示情感和實體提取之間有中等程度的一致性,以及額外的人物-實體對,這些對可能支持偏見分析。我們比較了兩個註釋提示,發現更嚴格的情感驗證規則消除了標籤分數不一致,但增加了執行時間並減少了註釋覆蓋率。根據這些結果,簡單的提示被用於後續的分析。我們提供了來源層級框架概況、不同來源之間的人物層級語調差異,以及事件層級的把關和報導指標的樣本分析。這些輸出顯示了相同的新聞集合如何用來檢查哪些來源進行報導、他們如何描述公眾人物,以及報導的集中地點。這種方法在手動驗證數據集或專業語言工具有限的情況下特別有用。

Decomposition Attacks Across Unlinkable Identities: Limits of Stateful Defenses for LLM Services

2608.17445v1 by Bowen Sun, Zhengyue Zhao, Xiaogeng Liu, Yinzhi Cao, Chaowei Xiao

Most large language model services use stateless defenses, which judge only the current request, to refuse harmful tasks. Decomposition attacks exploit this limitation by splitting a harmful task into individually permissible requests and combining their answers. Defending against them therefore requires a stateful monitor that considers requests together. If it can group all requests for one attacker task, it can stop the attack. However, attackers can use unlinkable identities and combine answers elsewhere, leaving no reliable grouping signal. We ask whether decomposition attacks can still be stopped under this setting. For a fixed attack strategy without retries, we prove that the achievable security and utility tradeoff depends entirely on how benign requests for the same capabilities are grouped. Persistent, recognizable groups permit a useful defense; fresh, indistinguishable groups do not. When attackers can retry and learn from Allow/Block decisions, this useful operating point disappears: the feedback reveals what passes but not whether a block was correct. Experiments on 91 executable tasks and 11,393 capability-matched benign requests support these results. Under a 1% denial cap for these requests and a 0.5% cap for unrelated background traffic, all ten tested policies, including one privileged policy with an exact request-to-operation map, either fail to stop attacks or exceed the budget. On defense-unseen task families, attack success is at least 99% after one attempt and 100% after two. Effective defenses therefore require additional evidence or mechanisms tied to grouping, such as reliable identity linkage, costs for fresh identities, or control over answer use.

摘要:大多數大型語言模型服務使用無狀態防禦,只根據當前請求來拒絕有害任務。分解攻擊利用了這一限制,通過將有害任務拆分為單獨可允許的請求並結合它們的答案來進行攻擊。因此,防禦這些攻擊需要一個有狀態的監控器,能夠將請求一起考慮。如果它能夠將所有針對一個攻擊者任務的請求分組,就能夠阻止攻擊。然而,攻擊者可以使用不可鏈接的身份並在其他地方結合答案,這樣就沒有可靠的分組信號。我們詢問在這種情況下是否仍然可以阻止分解攻擊。對於沒有重試的固定攻擊策略,我們證明可實現的安全性和效用權衡完全取決於對相同能力的良性請求如何分組。持久且可識別的群體允許有效的防禦;新鮮且不可區分的群體則不允許。當攻擊者可以重試並從允許/阻止決策中學習時,這一有用的操作點消失了:反饋揭示了哪些請求通過,但並不顯示阻止是否正確。對91個可執行任務和11,393個能力匹配的良性請求的實驗支持了這些結果。在這些請求的1%拒絕上限和與之無關的背景流量的0.5%上限下,所有十個測試的政策,包括一個具有精確請求到操作映射的特權政策,要麼未能阻止攻擊,要麼超出預算。在未見防禦的任務家族中,攻擊成功率在一次嘗試後至少為99%,在兩次嘗試後為100%。因此,有效的防禦需要額外的證據或與分組相關的機制,例如可靠的身份鏈接、新身份的成本或對答案使用的控制。

Structure-Internalized Rule Language Model for Faithful Knowledge Graph Reasoning

2608.17443v1 by Xingrui Zhuo, Jiapu Wang, Manzong Huang, Gongqing Wu, Xindong Wu

Knowledge Graph Reasoning (KGR) aims to discover latent facts by leveraging the structural evidence available in KGs, posing a challenge to the structural semantic understanding capability of KGR models. Recent studies have demonstrated that Large Language Models (LLMs) can achieve remarkable progress on KGR tasks via flexible in-context learning. However, the inherent representation inconsistency between KG structural context and LLM parametric knowledge remains inadequately addressed. This limitation prevents LLMs from effectively perceiving reasoning evidence that aligns with KG constraints, which undermines both the effectiveness and faithfulness of reasoning. We refer to this problem as reasoning evidence perception drift of LLMs over KGs. To address this problem, we propose a Structure-Internalized Rule Language Model (SIRLM), which centers on structural rule generation to couple the parametric learning of structural knowledge with the faithfulness evaluation of reasoning logic, enabling LLMs to anchor tightly to KG-grounded evidence. Specifically, we first design a Structure-Internalized Rule Generator (SIRG), which incorporates an in-context learning block augmented with a structural relation memory to coordinate structural and parametric knowledge. Furthermore, we equip SIRG with a KG tokenizer based on structural invariance learning and a neuro-symbolic reasoner based on rule-constrained message propagation. These components provide SIRG with learnable structural representations and faithful rule-execution feedback, respectively. Our SIRLM can be seamlessly integrated into standard LLM training paradigms, such as SFT and GRPO. Extensive experiments against 17 state-of-the-art KGR methods on 36 datasets demonstrate the significant superiority of SIRLM.

摘要:知識圖譜推理(KGR)旨在利用知識圖譜中的結構證據來發現潛在事實,這對KGR模型的結構語義理解能力提出了挑戰。最近的研究表明,大型語言模型(LLMs)可以通過靈活的上下文學習在KGR任務上取得顯著進展。然而,知識圖譜的結構上下文與LLM的參數知識之間固有的表示不一致性仍然未得到充分解決。這一限制阻礙了LLMs有效感知與知識圖譜約束相符的推理證據,從而削弱了推理的有效性和可靠性。我們將這個問題稱為LLMs在知識圖譜上的推理證據感知漂移。為了解決這個問題,我們提出了一種結構內化規則語言模型(SIRLM),該模型專注於結構規則生成,以將結構知識的參數學習與推理邏輯的可靠性評估相結合,使LLMs能夠緊密依賴於知識圖譜基礎的證據。具體而言,我們首先設計了一個結構內化規則生成器(SIRG),該生成器包含一個增強了結構關係記憶的上下文學習模塊,以協調結構和參數知識。此外,我們為SIRG配備了一個基於結構不變性學習的知識圖譜標記器和一個基於規則約束消息傳播的神經符號推理器。這些組件分別為SIRG提供了可學習的結構表示和可靠的規則執行反饋。我們的SIRLM可以無縫集成到標準的LLM訓練範式中,如SFT和GRPO。在36個數據集上對17種最先進的KGR方法進行的廣泛實驗顯示了SIRLM的顯著優越性。

Task-Aware Harness Provisioning for LLM Agents in Mission-Critical Infrastructure Operations

2608.17433v1 by Liangtao Lin, Qingang Zhang, Zhaomeng Zhu, Tianwei Zhang, Yonggang Wen

LLM agents have been widely adopted to operate mission-critical infrastructure (MCI). These agents normally rely on a harness that determines what information they can access, which tools they can use, and what actions they can take. Existing systems often expose the same comprehensive harness to every task, which may not be necessary and cause resource wastes. In this paper, we focus on the identification of optimal harness configurations, and view it as a resource-matching problem between what each task requires and what the harness provides. To measure this match, we classify MCI tasks based on the mathematical representation of the underlying system and rank harness configurations by the amount and type of information they provide. We then construct task-to-harness mappings from two sources: mining research literature and measuring controlled agent execution. Leveraging the measured mapping, we propose a new harness provisioning algorithm: map-guided escalation. It begins with a task-specific harness and expands to full provision only after a failed self-check. We evaluate our method in two representative MCI tasks: in liquid cooling, it improves the agent accuracy from 0.652 under full provision to 0.715 and achieves accuracy comparable to Reflexion with 48% fewer tokens; In power grids, full provision remains accuracy-optimal, while map-based provisioning offers lower-cost alternatives. These findings show that harness provisioning follows a domain-dependent accuracy-cost Pareto frontier rather than a universal optimum.

摘要:LLM 代理已被廣泛應用於運營關鍵任務基礎設施 (MCI)。這些代理通常依賴於一個裝置,該裝置決定它們可以訪問哪些信息、可以使用哪些工具以及可以採取哪些行動。現有系統通常對每個任務暴露相同的全面裝置,這可能不是必要的,並導致資源浪費。在本文中,我們專注於最佳裝置配置的識別,並將其視為每個任務所需與裝置提供之間的資源匹配問題。為了衡量這種匹配,我們根據基礎系統的數學表示對 MCI 任務進行分類,並根據它們提供的信息的數量和類型對裝置配置進行排名。然後,我們從兩個來源構建任務到裝置的映射:挖掘研究文獻和測量受控代理執行。利用測量的映射,我們提出了一種新的裝置供應算法:基於映射的升級。它從特定任務的裝置開始,僅在自我檢查失敗後擴展到完全供應。我們在兩個具代表性的 MCI 任務中評估我們的方法:在液體冷卻中,它將代理的準確度從完全供應下的 0.652 提高到 0.715,並且在使用 48% 更少的標記的情況下達到與 Reflexion 相當的準確度;在電力網中,完全供應仍然是準確度最佳,而基於地圖的供應則提供了更低成本的替代方案。這些發現表明,裝置供應遵循一個依賴於領域的準確度-成本 Pareto 邊界,而不是一個普遍的最優解。

SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation

2608.17426v1 by Keyu Tu, Zhuowei Chen, Mengqi Huang, Yuxin Wang, Jiahao Zhu, Zhendong Mao, Yongdong Zhang

We introduce Semantic Task Completion Video Generation, an outcome-oriented video generation task. Under this formulation, success requires both achievement of the intended outcome and semantic grounding. Semantic grounding characterizes the correspondence between the reference image and the generated outcome in terms of high-level semantics relevant to the task. Evaluation focuses on the generated outcome and requires neither the presentation of a complete sequence of intermediate task steps nor conventional appearance consistency with the reference image. To support systematic evaluation, we construct SemComp-Data, an evaluation dataset covering six domains. Each instance comprises a reference image, a detailed instruction, a brief instruction, and an outcome-centric video clip. A scalable four-stage curation pipeline converts raw videos into standardized SemComp-Data instances. We further introduce SemComp-Bench, an evaluation protocol that uses a vision-language model (VLM) to answer structured binary questions. SemComp-Bench reports the OA Score and the GR Score for Outcome Achievement and Generation Reliability, respectively. Experiments on representative video generation models show that achieving intended outcomes while maintaining task-relevant semantic grounding in reference images remains challenging.

摘要:我們介紹了語義任務完成視頻生成,這是一項以結果為導向的視頻生成任務。根據這一表述,成功需要同時實現預期的結果和語義基礎。語義基礎描述了參考圖像與生成結果之間在與任務相關的高階語義方面的對應關係。評估重點在於生成的結果,並不需要呈現完整的中間任務步驟序列,也不需要與參考圖像的常規外觀一致性。為了支持系統化評估,我們構建了SemComp-Data,一個涵蓋六個領域的評估數據集。每個實例包括一個參考圖像、一個詳細說明、一個簡要說明和一個以結果為中心的視頻片段。一個可擴展的四階段策展管道將原始視頻轉換為標準化的SemComp-Data實例。我們進一步介紹了SemComp-Bench,一個使用視覺-語言模型(VLM)回答結構化二元問題的評估協議。SemComp-Bench報告了結果達成的OA分數和生成可靠性的GR分數。對代表性視頻生成模型的實驗顯示,在保持與參考圖像相關的語義基礎的同時實現預期結果仍然具有挑戰性。

An Investigation of Translationese in the Generations of Multilingual Large Language Models

2608.17399v1 by Maria Valentini, Téa Wright, Julisa Granados, Eliana Colunga, Katharina von der Wense

Text which has been translated from another language tends to carry with it evidence of translation$\unicode{x2014}$hence, it is often referred to as $\textit{translationese}$. Multilingual large language models (MLLMs) generate text in a variety of languages. However, it is still unclear if MLLMs' generations resemble internal translation (from English or, potentially, other languages) and, thus, result in translationese. Here, we ask the following research questions: (1) Does text generated by MLLMs resemble translationese? (2) How does translationese produced by MLLMs differ from translationese produced through direct translation? We leverage established indicators of translated text to evaluate text generated by state-of-the-art MLLMs in five languages, comparing to both non-translated and human-written baselines in order to isolate translationese from other kinds of interference. Through the use of high-accuracy classification models, analyses of variance on individual linguistic features, and the collection of human annotations in a subset of two languages (German and Spanish), we assess the translationese content of MLLM generations and examine the key features that distinguish MLLM-generated text from typical translation-related interference.

摘要:從另一種語言翻譯過來的文本往往帶有翻譯的痕跡$\unicode{x2014}$因此,它通常被稱為$\textit{translationese}$。多語言大型語言模型(MLLMs)可以生成多種語言的文本。然而,目前尚不清楚MLLMs生成的文本是否類似於內部翻譯(從英語或潛在的其他語言),因此是否會導致翻譯語。 在這裡,我們提出以下研究問題:(1)MLLMs生成的文本是否類似於翻譯語?(2)MLLMs產生的翻譯語與通過直接翻譯產生的翻譯語有何不同?我們利用已建立的翻譯文本指標來評估五種語言中最先進的MLLMs生成的文本,並與非翻譯文本和人類撰寫的基準進行比較,以便將翻譯語與其他類型的干擾區分開來。通過使用高準確度的分類模型、對個別語言特徵的變異分析,以及在兩種語言(德語和西班牙語)子集中的人類標註收集,我們評估MLLM生成的翻譯語內容,並檢查區分MLLM生成文本與典型翻譯相關干擾的關鍵特徵。

LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents

2608.17393v1 by Yiming Du, Yuxin Jiang, Tao Yuan, Jianbo Dai, Shaowei Wang, Jierun Chen, Chaofan Tao, Xianzhi Yu, Lifeng Shang, Kam-Fai Wong, Xiaohui Li, Haoli Bai

Reinforcement learning for coding agents increasingly relies on long-running agent harnesses to manage tool integration, repository contexts, and execution feedback. However, the native execution environments of these harnesses are inherently misaligned with policy-gradient training: environmental crashes and reward hacking corrupt outcome signals, while train-inference discrepancies decouple rollout behavior from policy updates. To address this, we present LEGO-RL, a framework that bridges native coding-agent harnesses with scalable policy-gradient optimization without modifying their internal control flow. LEGO-RL is built upon three pillars: (1) faithful optimization via in-process LLM proxying that captures raw generation streams for token-level alignment and robust trainer-side log-probability recomputation, even under harness-side compaction or re-serialization; (2) reliable execution via scalable sandbox orchestration featuring image caching and stage-wise defenses to mitigate reward hacking; and (3) observable training through an integrated plugin that automates validation and monitoring, paired with a Live UI for granular trajectory diagnostics. We evaluate LEGO-RL by training the sparse MoE model Qwen3.5-35B-A3B with GSPO across three native coding-agent harnesses. LEGO-RL improves Qwen3.5-35B-A3B across OpenHands SDK (64.0% to 70.4%), Claude Code (62.4% to 68.2%), and OpenCode (57.2% to 66.6%) on SWE-bench Verified, while maintaining a rollout-training probability correlation above 0.99.

摘要:強化學習對於編碼代理的依賴越來越多,尤其是在長期運行的代理工具中,以管理工具整合、庫上下文和執行反饋。然而,這些工具的本地執行環境與策略梯度訓練本質上不一致:環境崩潰和獎勵駭客會破壞結果信號,而訓練與推理之間的差異使得回滾行為與策略更新脫節。為了解決這個問題,我們提出了LEGO-RL,一個將本地編碼代理工具與可擴展的策略梯度優化相連接的框架,而無需修改其內部控制流程。LEGO-RL建立在三個支柱之上:(1)通過進程內LLM代理捕獲原始生成流以實現令牌級對齊的忠實優化,以及在工具端壓縮或重新序列化下的穩健訓練方日志概率重新計算;(2)通過可擴展的沙盒編排實現可靠的執行,特徵包括圖像緩存和階段防禦,以減輕獎勵駭客的影響;(3)通過集成插件實現可觀察的訓練,自動化驗證和監控,並配備用於細粒度軌跡診斷的實時用戶界面。我們通過在三個本地編碼代理工具上使用GSPO訓練稀疏的MoE模型Qwen3.5-35B-A3B來評估LEGO-RL。LEGO-RL在SWE-bench Verified上改善了Qwen3.5-35B-A3B在OpenHands SDK(從64.0%提升至70.4%)、Claude Code(從62.4%提升至68.2%)和OpenCode(從57.2%提升至66.6%)的表現,同時保持回滾訓練概率的相關性高於0.99。

Leveraging generative hallucination and biophysics-informed modeling for unified biomolecular sequence-structure co-design

2608.17381v1 by Xuefeng Liu, Mingxuan Cao, Xiao Luo, Songhao Jiang, Tobin Sosnick, Jinbo Xu, Louis Maher, Rick Stevens

Biomolecular design underpins applications from molecular recognition to therapeutics and synthetic biology, yet de novo interaction design remains challenging-especially for DNA/RNA, underexplored non-protein modalities with scarce, heterogeneous complex data and sharper geometric and chemical constraints. We introduce MCTH (Monte Carlo Tree Hallucination), an inference-only framework that casts all-atom sequence-structure co-design as uncertainty-aware planning over hallucinated states from pretrained folding and inverse-folding models, with optional biophysical control within the same decision loop. MCTH treats these models as frozen black-box operators and uses Monte Carlo Tree Search to allocate a fixed inference budget across competing design trajectories, incorporating model confidence and uncertainty, as well as cross-expert consensus/disagreement when multiple predictors are available. Across protein-RNA, protein-DNA, protein-protein, and protein-ligand design, matched-budget experiments show that adaptive search improves over simpler sampling and cycling strategies, while held-out AlphaFold3 and Chai-1 evaluations demonstrate transfer beyond the search-time oracle. MCTH provides a shared planning layer across modalities while allowing task-specific folding, inverse-folding, and biophysical modules, requiring no fine-tuning or backpropagation through component models.

摘要:生物分子設計支撐著從分子識別到治療和合成生物學的應用,然而,從零開始的互動設計仍然具有挑戰性,尤其是對於DNA/RNA這些未被充分探索的非蛋白質模式,其複雜數據稀少且異質,並且面臨更嚴格的幾何和化學限制。我們介紹了MCTH(蒙特卡羅樹幻覺),這是一個僅用於推理的框架,將全原子序列結構共同設計視為對從預訓練的摺疊和反摺疊模型中幻覺狀態的帶有不確定性意識的規劃,並在同一決策循環中可選擇生物物理控制。MCTH將這些模型視為凍結的黑箱運算符,並使用蒙特卡羅樹搜索在競爭設計軌跡中分配固定的推理預算,納入模型信心和不確定性,以及當多個預測器可用時的跨專家共識/分歧。在蛋白質-RNA、蛋白質-DNA、蛋白質-蛋白質和蛋白質-配體設計中,匹配預算的實驗顯示,自適應搜索優於更簡單的採樣和循環策略,而持出來的AlphaFold3和Chai-1評估則顯示出超越搜索時間神諭的轉移。MCTH在不同模式之間提供了一個共享的規劃層,同時允許特定任務的摺疊、反摺疊和生物物理模塊,無需對組件模型進行微調或反向傳播。

PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX

2608.17379v1 by Genghan Zhang, Yixin Dong, Chengze Fan, Zhichen Zeng, Yueming Yuan, Shaowei Zhu, Kunle Olukotun

We introduce PTXBench, a benchmark for evaluating and adapting large language models (LLMs) to use architecture-specific PTX for GPU kernel optimization. PTXBench measures functional correctness, whether selected target instructions execute at runtime, and speedup over frontier libraries across GEMM and attention workloads on H100 and B200 GPUs. Our evaluation shows that architecture-specific PTX capability remains uneven: success rates fall substantially on complex attention backward workloads, and executing the target instructions does not necessarily translate into competitive performance. No evaluated model consistently matches frontier libraries across the suite. We further adapt Qwen3.6-27B using supervised fine-tuning. Repair-conditioned training improves several tasks, but generalization remains uneven; data coverage, balance, and the quality of the reasoning teacher matter in addition to dataset size. PTXBench provides an auditable testbed for measuring and improving LLMs' ability to exploit evolving GPU architectures.

摘要:我們介紹 PTXBench,這是一個用於評估和調整大型語言模型(LLMs)以使用特定架構的 PTX 進行 GPU 核心優化的基準測試。PTXBench 測量功能正確性,檢查所選目標指令在運行時是否執行,以及在 H100 和 B200 GPU 上的 GEMM 和注意力工作負載中,相較於前沿庫的加速效果。我們的評估顯示,特定架構的 PTX 能力仍然不均衡:在複雜的注意力反向工作負載上,成功率顯著下降,而執行目標指令不一定能轉化為具競爭力的性能。在整個測試套件中,沒有任何評估模型能持續匹配前沿庫。我們進一步使用監督微調調整 Qwen3.6-27B。修復條件訓練改善了幾個任務,但泛化能力仍然不均衡;數據覆蓋、平衡以及推理教師的質量在數據集大小之外也很重要。PTXBench 提供了一個可審核的測試平台,用於測量和改善 LLM 利用不斷演變的 GPU 架構的能力。

Fair ASR: Re-Evaluating Black-Box Jailbreaks under Shared Target-Call Budgets

2608.17360v1 by Zhida He, Xiaoyu Wen, Han Qi, Ziyuan Zhou, Peng Yu, Jiajia Li, Chaochao Lu, Qiaosheng Zhang

Reliable jailbreak evaluation is essential for assessing LLM safety, but most existing studies rely solely on attack success rate (ASR) without accounting for its dependence on attack budgets, resulting in unfair comparisons across methods. Existing compute-aware evaluations reduce heterogeneous resources into FLOPs, which is difficult to estimate for black-box models and fails to capture resource-specific constraints. To provide a comparable evaluation basis, we introduce Fair-ASR, an evaluation protocol for black-box jailbreak attacks under shared target-call budgets B, using target calls as a directly observable and method-agnostic comparison axis while tracking attacker calls separately for efficiency analysis. We re-evaluate 11 representative attacks under the Fair-ASR protocol and find that attack rankings change substantially across target-call budgets, simple stochastic perturbations and hand-crafted templates remain highly competitive under equal target access, and no evaluated LLM-driven method is efficient in both target and attacker calls. Motivated by this efficiency gap, we introduce ReCode, a compositional budget-efficient attack that combines desensitization rewriting with two effective low-cost primitives identified by Fair-ASR. Under a budget of 20 target calls, ReCode achieves 85% ASR on GPT-5 while requiring only 7.19 attacker calls per request on average, showing strong efficiency in both target and attacker calls.

摘要:可靠的越獄評估對於評估大型語言模型(LLM)的安全性至關重要,但現有的大多數研究僅依賴攻擊成功率(ASR),而未考慮其對攻擊預算的依賴,導致不同方法之間的比較不公平。現有的計算感知評估將異質資源簡化為FLOPs,這對於黑箱模型來說難以估算,並且未能捕捉資源特定的限制。為了提供可比較的評估基礎,我們引入了Fair-ASR,這是一種針對共享目標調用預算B的黑箱越獄攻擊的評估協議,利用目標調用作為直接可觀察且與方法無關的比較軸,同時單獨跟踪攻擊者調用以進行效率分析。我們在Fair-ASR協議下重新評估了11個代表性攻擊,發現攻擊排名在不同的目標調用預算下有顯著變化,簡單的隨機擾動和手工製作的模板在平等的目標訪問下仍然具有高度競爭力,且沒有評估的LLM驅動方法在目標和攻擊者調用中都有效率。受到這一效率差距的激勵,我們引入了ReCode,一種組合預算高效的攻擊,結合了去敏感化重寫和Fair-ASR識別的兩個有效低成本原語。在20次目標調用的預算下,ReCode在GPT-5上達到85%的ASR,同時每次請求平均僅需7.19次攻擊者調用,顯示出在目標和攻擊者調用中都具有強大的效率。

Cognitive Graph Intelligence for Adaptive and Robust DDoS Attack Detection in Next Generation Networks

2608.17352v1 by Mohammad Arif Hossain, Yeahia Sarker, Md Jafrin Hossain, Most. Humayra Khanom Rime, Nirwan Ansari

Distributed Denial-of-Service (DDoS) attacks threaten network availability, requiring a cognitive detection process that senses traffic, infers intent, and supports an adaptive response under severe class imbalance and non-stationary conditions. This paper proposes a Graph-based Generative Adversarial Network (GraphGAN) that serves as the cognitive detection engine for this task. GraphGAN captures the relational structure among traffic flows while addressing imbalance through adversarial generation of synthetic samples. Sequential flows are converted into $k$-nearest neighbor graphs using sliding windows to preserve feature-similarity and temporal dependencies among flows. The generator learns the distribution of DDoS attacks to synthesize realistic minority samples, while a Graph Convolutional Network (GCN)-based discriminator distinguishes real from synthetic graph data. A separate GCN classifier, trained on the balanced dataset, performs the final detection decision. Evaluations on four benchmark datasets show that GraphGAN achieves superior accuracy, precision, and recall compared to state-of-the-art approaches, particularly in data-scarce scenarios. By integrating temporal graph construction, adversarial augmentation, and GCN classification, GraphGAN effectively models coordinated attack behaviors and mitigates class imbalance, providing a robust and topology-aware solution for intrusion detection in data-constrained environments.

摘要:分散式拒絕服務(DDoS)攻擊威脅網絡可用性,這需要一個認知檢測過程來感知流量、推斷意圖,並在嚴重的類別不平衡和非穩態條件下支持自適應響應。本文提出了一種基於圖的生成對抗網絡(GraphGAN),作為此任務的認知檢測引擎。GraphGAN 捕捉流量流之間的關係結構,同時通過對抗生成合成樣本來解決不平衡問題。連續流量被轉換為 $k$-最近鄰圖,使用滑動窗口來保留流量之間的特徵相似性和時間依賴性。生成器學習 DDoS 攻擊的分佈,以合成現實的少數樣本,而基於圖卷積網絡(GCN)的判別器則區分真實與合成的圖數據。另一個在平衡數據集上訓練的 GCN 分類器執行最終檢測決策。在四個基準數據集上的評估顯示,GraphGAN 在準確性、精確度和召回率方面優於最先進的方法,特別是在數據稀缺的情況下。通過整合時間圖構建、對抗增強和 GCN 分類,GraphGAN 有效地建模協調攻擊行為並減輕類別不平衡,為數據受限環境中的入侵檢測提供了一個強健且考慮拓撲的解決方案。

MoFE: A Novel Mixture-of-Experts Framework with Fourier Neural Operators for Cryptocurrency Forecasting

2608.17342v1 by Bowen Liu, Mingming Sun

Forecasting cryptocurrency prices remains a formidable challenge due to inherent non-stationarity, abrupt regime shifts, and multi-scale stochastic dependencies. Conventional deep learning models often struggle to capture complex underlying dynamics, frequently resulting in persistent phase-lagged predictions. To address these limitations, we propose MoFE, a novel deep learning framework that integrates Fourier Neural Operators (FNOs) within a Mixture-of-Experts (MoE) architecture. Rooted in the theoretical framework of stochastic differential equations, MoFE conceptualizes cryptocurrency volatility as a superposition of multi-frequency components, which includes user network based fundamental growth, mining costs and halving mechanism caused seasonal volatility, and market sentiment-induced chaos. Specifically, specialized adaptive FNO (AFNO) and Convolution dual-domain experts learn continuous function-to-function mappings to encapsulate global spectral trends, cyclical adjustments and microstructures, while a dynamic gating based MoE mechanism enables adaptive strategy switching across diverse market regimes. Extensive experiments on Bitcoin datasets spanning January 2020 to December 2025 demonstrate that MoFE achieves state-of-the-art (SOTA) performance in both T+1 and T+5 forecasting horizons. Notably, the model effectively mitigates the phase-lag effect, delivering superior Directional Accuracy (DA) and Information Coefficient (IC). In high-fidelity simulated trading environments, these predictive gains transfer into significant excess returns and robust risk-adjusted performance, characterized by a high Sharpe ratio.

摘要:預測加密貨幣價格仍然是一項艱巨的挑戰,因為其固有的非平穩性、突變的制度轉變以及多尺度隨機依賴性。傳統的深度學習模型往往難以捕捉複雜的潛在動態,經常導致持續的相位滯後預測。為了解決這些限制,我們提出了MoFE,一種新穎的深度學習框架,將傅立葉神經運算子(FNOs)整合到專家混合(MoE)架構中。MoFE根植於隨機微分方程的理論框架,將加密貨幣的波動性概念化為多頻率組件的疊加,這包括基於用戶網絡的基本增長、挖礦成本和因減半機制引起的季節性波動,以及市場情緒引發的混沌。具體而言,專門的自適應FNO(AFNO)和卷積雙域專家學習連續的函數到函數映射,以封裝全球光譜趨勢、周期性調整和微結構,而基於動態門控的MoE機制則使得在不同市場制度之間的自適應策略切換成為可能。對於2020年1月至2025年12月的比特幣數據集進行的廣泛實驗表明,MoFE在T+1和T+5預測範圍內均實現了最先進的(SOTA)性能。值得注意的是,該模型有效減輕了相位滯後效應,提供了優越的方向準確性(DA)和信息係數(IC)。在高保真模擬交易環境中,這些預測增益轉化為顯著的超額回報和穩健的風險調整表現,特徵是高夏普比率。

LLM-Only PDDL Domain Repair with Open-Weight Models

2608.17341v1 by Nader Karimi Bavandpour, Pascal Bercher

AI planning is concerned with finding a sequence of actions that achieves a specified goal. It relies on explicit models of the world, commonly represented in the Planning Domain Definition Language (PDDL). An active line of research investigates how errors in such models can be detected and repaired. For example, users may provide positive test plans that are solutions, and negative test plans that fail during execution. Automated repair methods then modify the PDDL model to satisfy these constraints. In this paper, we evaluate the ability of recent open-weight large language models to perform this repair task using an LLM-only approach. Our experiments show that the symbolic baseline achieves an $F_1$ score of $.49$, while the best-performing LLM reaches $.87$ with high reasoning effort, an absolute improvement of $.38$. However, that setting has a mean test pass rate of only $.82$, falling to $.06$ on the Thoughtful domain; even the best setting that includes the test traces reaches only $.92$. Thus, current open-weight models cannot guarantee satisfaction of the test constraints required for reliable automated model repair.

摘要:AI 規劃關注於找到一系列行動以達成特定目標。它依賴於對世界的明確模型,通常以規劃領域定義語言 (PDDL) 表示。一個活躍的研究方向探討如何檢測和修復這些模型中的錯誤。例如,使用者可能提供正面的測試計劃作為解決方案,以及在執行過程中失敗的負面測試計劃。自動修復方法隨後會修改 PDDL 模型以滿足這些約束。在本文中,我們評估最近的開放權重大型語言模型使用僅 LLM 方法執行此修復任務的能力。我們的實驗顯示,符號基準達到了 $.49$ 的 $F_1$ 分數,而表現最佳的 LLM 則在高推理努力下達到了 $.87$,絕對改善為 $.38$。然而,該設置的平均測試通過率僅為 $.82$,在 Thoughtful 領域下降至 $.06$;即使是包括測試痕跡的最佳設置也僅達到 $.92$。因此,目前的開放權重模型無法保證滿足可靠的自動模型修復所需的測試約束。