40 papers across AI, ML, NLP, and CV from the last 24 hours.
Three interconnected threads dominate today's arXiv submissions. First, the community is pivoting from asking what LLMs can do to how reliably they execute -- four separate papers examine procedural fidelity, tool-calling decisions, plan-constrained agentic execution, and the case for Bayesian control layers in multi-agent systems. Second, vision-language models are hitting practical bottlenecks: KV cache bloat during autoregressive decoding, visual signal dilution as textual history expands, and the persistent difficulty of aligning vision encoders with language objectives without contrastive overhead. Third, applied AI is confronting real-world constraints head-on -- security audits of patient-facing RAG systems, federated unlearning with cross-modal entanglement, and hospital readmission prediction grounded in EHR data windows.
A position paper arguing that agentic AI orchestration should be Bayes-consistent stands out. Rather than proposing another method to make LLMs "uncertain," it makes a structural claim: the control layer that routes between tools, experts, and resources is where Bayesian decision theory applies cleanly, separate from the opaque inference happening inside the model. This reframes the reliability problem as an engineering question about system architecture rather than model alignment.
Today's batch collectively suggests the field is maturing past capability demonstrations and into the mechanics of dependable deployment -- efficiency, verification, and governance are no longer afterthoughts but primary research questions.
George Stoica, Sayak Paul, Matthew Wallingford · 2026-05-01
Flow matching (FM) trains a time-dependent vector field that transports samples from a simple prior to a complex data distribution. However, for high-dimensional images, each training sample supervises only a single trajectory and intermediate point, yielding an extremely sparse and high-variance training signal. This under-constrained supervision can cause flow collapse, where the learned dynamic…
Siyuan Huang, Xiaoye Qu, Yafu Li · 2026-05-01
While autoregressive Large Vision-Language Models (LVLMs) demonstrate remarkable proficiency in multimodal tasks, they face a "Visual Signal Dilution" phenomenon, where the accumulation of textual history expands the attention partition function, causing visual attention to decay inversely with generated sequence length. To counteract this, we propose Persistent Visual Memory (PVM), a lightweight …
Yan Fang, Mengcheng Lan, Zilong Huang · 2026-05-01
In this paper, we present erative anguage-mage re-training (GenLIP), a minimalist generative pretraining framework for Vision Transformers (ViTs) designed for multimodal large language models (MLLMs). To better align vision encoders with the autoregressive nature of LLMs, GenLIP trains a ViT to predict language tokens directly from visual tokens using a st…
Xinyuan Zhao, Yihang Wu, Ahmad Chaddad · 2026-05-01
Gaze estimation methods commonly use facial appearances to predict the direction of a person gaze. However, previous studies show three major challenges with convolutional neural network (CNN)-based, transformer-based, and contrastive language-image pre-training (CLIP)-based methods, including late fusion of image features, lack of factor-aware conditioning, and impractical capacity scaling. To ad…
Xihao Chen, Yangyang Guo, Roger Zimmermann · 2026-05-01
Key-Value (KV) cache has become a de facto component of modern Large Vision-Language Models (LVLMs) for inference. While it enhances decoding efficiency in Large Language Models (LLMs), its direct adoption in LVLMs introduces substantial GPU memory overhead due to the large number of vision tokens processed during the prefill stage. To tackle this problem, we propose LightKV, a novel approach that…
Jaeyoung Chung, Suyoung Lee, Jianfeng Xiang · 2026-05-01
3D world generation is essential for applications such as immersive content creation or autonomous driving simulation. Recent advances in 3D world generation have shown promising results; however, these methods are constrained by grid layouts and suffer from inconsistencies in object scale throughout the entire world. In this work, we introduce a novel framework, Map2World, that first enables 3D w…
Lin Che, Xi Wang, Marc Pollefeys · 2026-05-01
Urban perception describes how people subjectively evaluate urban environments, shaping how cities are experienced and understood. Existing computational approaches primarily model urban perception directly from street view images, but largely ignore the human perceptual process through which such judgments are formed. In this paper, we introduce Place Pulse-Gaze, an urban perception dataset that …
Mohammad Aamir Sohail, Gabriela Pinheiro, Yasemin Poyraz Kocak · 2026-05-01
Edge detection refers to identifying points in a digital image where intensity changes sharply, indicating object boundaries or structural features. Corners are locations where gray-level intensity changes abruptly in multiple directions and are widely used in feature extraction, object tracking, and 3D modeling. In this study, we present a quantum implementation of Sobel-based edge detection and …
Qiancheng Zhou, Wenhua Zhang · 2026-05-01
Single-point supervised infrared small target detection (IRSTD) drastically reduces dense annotation costs. Current state-of-the-art (SOTA) methods achieve high precision by recovering mask supervision through explicit, offline pseudo-label construction, such as multi-stage active learning and physics-driven mask generation. In this paper, we study a minimalist alternative: generating point-to-mas…
Yinghao Chen, Yeying Jin, Xiang Chen · 2026-05-01
Unsupervised deraining has attracted attention for its ability to learn the real-world distribution of rain without paired supervision. However, the lack of strong constraints makes it difficult for the network to converge, especially with the complex diversity of rain degradation. A key motivation is that high-quality deraining results occasionally emerge during training, which can be leveraged t…
Tongxu Zhang · 2026-05-01
Knee osteoarthritis (OA) assessment involves a natural but often underused label hierarchy: a coarse binary OA decision and a fine-grained Kellgren--Lawrence (KL) severity grade. Existing deep learning studies commonly treat these targets as separate classification problems, either reducing OA assessment to disease presence or directly optimizing noisy ordinal KL labels. In this work, we ask wheth…
Pavlin G. Poličar, Andraž Pevcin, Blaž Zupan · 2026-05-01
Generating diverse, readable statistical charts from tabular data remains challenging for LLMs, as many failures become apparent after rendering and are not detectable from data or code alone. Existing chart datasets also rarely provide fully aligned artifacts, such as executable code, dataset context, and question-answer pairs. We present a structured LLM-based workflow that decomposes chart gene…
Arunabh Srivastava, Mohammad A., Khojastepour · 2026-05-01
Humans solve problems by executing targeted plans, yet large language models (LLMs) remain unreliable for structured workflow execution. We propose RunAgent, a multi-agent plan execution platform that interprets natural-language plans while enforcing stepwise execution through constraints and rubrics. RunAgent bridges the expressiveness of natural language with the determinism of programming via a…
Stavros Orfanoudakis, Pedro P. Vergara · 2026-05-01
While representation and similarity learning have improved the sample efficiency of Reinforcement Learning (RL), they are rarely used to shape policy updates directly in the action space. To bridge this gap, a geometry-aware RL algorithm that explicitly incorporates value-based similarity into the policy update, State-Action Value Geometry Optimization (SAVGO), is proposed. In detail, SAVGO learns…
Jacques Raynal, Pierre Slangen, Jacques Margerit · 2026-05-01
In biomechanical systems, observable performance is often used as a proxy for underlying system organization. However, this assumption implicitly presumes a correspondence between output metrics and internal system states that may not hold in adaptive systems. In this study, the vertical dimension of occlusion (VDO) is considered as a constraint applied to an adaptive neuromechanical system, enabl…
Shradha Sharma, Swapnil Dhamal, Shweta Jain · 2026-05-01
We propose a new framework for meritocratic fairness in budgeted combinatorial multi-armed bandits with full-bandit feedback (BCMAB-FBF). Unlike semi-bandit feedback, the contribution of individual arms is not received in full-bandit feedback, making the setting significantly more challenging. To compute arm contributions in BCMAB-FBF, we first extend the Shapley value, a classical solution concep…
Rodolphe Barlogis, Ferhat Tamssaouet, Quentin Falcoz · 2026-05-01
This paper deals with solving the 2D Helmholtz equation on non-parametric domains, leveraging a physics-informed neural operator network based on the DeepONet framework. We consider a 2D square domain with an inclusion of arbitrary boundary geometry at its center. This inclusion acts as a scatterer for an incoming harmonic wave. The aim is to learn the operator linking the geometry of the scattere…
Sizhe Tang, Zuyuan Zhang, Mahdi Imani · 2026-05-01
Monte Carlo Tree Search (MCTS) scales poorly in cooperative multi-agent domains because expansion must consider an exponentially large set of joint actions, severely limiting exploration under realistic search budgets. We propose NonZero, which keeps multi-agent MCTS tractable by running surrogate-guided selection over a low-dimensional nonlinear representation using an interaction-guided proposal…
Ramin Mohammadi, Vahab vahdat, Sarthak Jain · 2026-05-01
With the proliferation of Electronic Health Records (EHRs), a critical challenge in building predictive models is determining the optimal historical data time window to maximize accuracy. This study investigates the impact of various observation windows ranging from the day of surgery to three years prior on predicting 30-day readmission following hip and knee arthroplasties. The dataset encompass…
Jiawen Chen, Qi Shao, Duxin Chen · 2026-05-01
Combinatorial complexes have unified set-based (e.g., graphs, hypergraphs) and part-whole (e.g., simplicial, cellular complexes) structures into a common topological framework. Existing topological neural networks and Weisfeiler-Lehman variants remain fragmented, lacking a unified theoretical foundation for topological deep learning. In this work, we introduce the Combinatorial Complex Weisfeiler-…
Nikolaos Nakis, Chrysoula Kosma, Panagiotis Promponas · 2026-05-01
Representation learning is central to graph machine learning, powering tasks such as link prediction and node classification. However, most graph embeddings are hard to interpret, offering limited insight into how learned features relate to graph structure. Many networks naturally admit a role-mixture view, where nodes are best described as mixtures over latent archetypal factors. Motivated by thi…
Sailesh Panda, Pritam Kadasi, Abhishek Upperwal · 2026-05-01
Large language models (LLMs) often achieve strong performance on reasoning benchmarks, but final-answer accuracy alone does not show whether they faithfully execute the procedure specified in a prompt. We study this question through a controlled diagnostic benchmark for procedural execution, where models are given a step-wise arithmetic algorithm and two numeric inputs, and must return the final c…
Scott Friedman, Ruta Wheelock, Sonja Schmer-Galunder · 2026-05-01
The language in online platforms, influence operations, and political rhetoric frequently directs a mix of pro-social sentiment (e.g., advocacy, helpfulness, compassion) and anti-social sentiment (e.g., threats, opposition, blame) at different topics, all in the same message. While many natural language processing (NLP) tools classify or score a text's overall sentiment as positive, neutral, or ne…
Jiaoda Li, Ryan Cotterell · 2026-05-01
The transformer is the most popular neural architecture for language modeling. The cornerstone of the transformer is its global attention mechanism, which lets the model aggregate information from all preceding tokens before generating the next token. One common variant of attention is called local attention, which restricts each token to aggregating information from a bounded window of predecesso…
Ziyang Huang, Yi Cao, Ali K. Shargh · 2026-05-01
Large language models are increasingly deployed as autonomous coding agents and have achieved remarkably strong performance on software engineering benchmarks. However, it is unclear whether such success transfers to computational scientific workflows, where tasks require not only strong coding ability, but also the ability to navigate complex, domain-specific procedures and to interpret results i…
This digest is generated automatically from arXiv submissions. Not affiliated with arXiv or Cornell University.