Three themes dominate today's batch of 40 papers. First, video generation and 3D spatial reasoning are merging: models are no longer content with flat image outputs. VLM-IE3D equips vision-language models with implicit and explicit 3D geometry from video, while SANA-Video 2.0 hybridizes linear and softmax attention to generate 720p video at 5B and 14B scales on a single GPU. WorldWeaver introduces cross-agent state registers to maintain persistent world states across multi-view video rollouts, and GraphVid brings graph-based control to multi-object video generation. Second, there is a push to make agent training end-to-end tractable. OpenForgeRL exposes harness-level inference to RL stacks that cannot natively reason about stateful, multi-process tool use, while Agentic Context Management reframes the memory-and-cost problem as an architecture and lifecycle concern rather than a retrieval problem. Third, several papers probe the theoretical boundaries of what machine learning models can and should do, from proving that surprisal theory of human language processing is tautological without rational grounding, to showing that LLMs can appear safer under direct exposure to a dangerous objective than when that objective is relayed through mediating agents.
Ryan Cotterell's single-author paper on surprisal theory stands out. It argues that any non-negative difficulty measure over linguistic units can be fit affirmatively by some language model's surprisal under mild conditions, rendering the entire theory vacuous without independent constraints on which language model counts as rational. This is a clean, potentially field-altering challenge to a framework widely used in computational psycholinguistics.
Notably absent from today's batch are papers on new large-scale model architectures or benchmark-setting pretraining runs. The focus has shifted toward making existing capabilities more controllable, more efficient, and more grounded. Taken together, this batch suggests the field is maturing from capability expansion into a phase of structural refinement: better spatial reasoning, better inference economics, and sharper theoretical self-examination.
Wenhao Li, Xueying Jiang, Quanhao Qian · 2026-07-23
Despite rapid progress, most existing vision-language models (VLMs) built from 2D visual inputs often struggle when handling various 3D tasks that require fine-grained spatial understanding and reasoning. To bridge this gap, we present VLM-IE3D, a unified framework that enhances the 3D spatial awareness of VLMs by equipping them with both implicit and explicit 3D geometries learned from RGB videos...
Sicheng Mo, Yuheng Li, Ziyang Leng · 2026-07-23
Multi-agent interactive world models should not only generate consistent observations, but also maintain world states that persist across agents and evolve across views. Existing autoregressive video diffusion pipelines carry forward observation history as conditioning context, which makes shared state difficult to maintain in multi-agent and multi-view settings. We present WorldWeaver (W^2), a st...
Yihong Sun, Seoung Wug Oh, Jiahui Huang · 2026-07-23
Scene understanding requires simultaneous prediction about geometry, appearance, and semantics. However, existing task-specific annotations are fragmented across incompatible, domain-specific datasets. Current unified systems circumvent this by restricting training to fully co-annotated data, or by incurring the large computational cost of pseudo-labeling. To mitigate this, we introduce UniD, a un...
Rogerio Guimaraes, Pietro Perona · 2026-07-23
Diffusion and flow-matching models dominate conditional image generation, yet inference-time scaling for these models is far less developed than for autoregressive language models. Because final quality is highly sensitive to the initial noise seed, many approaches spend extra compute on seed search or resampling under a black-box reward, but typically maintaining a constant memory footprint throu...
Vedant Shah, Onkar Susladkar, Tushar Prakash · 2026-07-23
Controllable video generation remains challenging due to the difficulty of specifying precise multi-object interactions using text prompts or motion-control inputs that primarily constrain pixel movement. In practice, trajectory-based control often requires users to draw accurate tracks for multiple objects, which scales poorly with scene complexity and becomes ambiguous under occlusion or overlap...
Korota Arsène Coulibaly, Mohamed Hamlich, Khalid Hmali · 2026-07-23
Quality control in printing, particularly in rotogravure printing, still depends on slow, costly, and subjective manual inspection. Automated surface defect detection is critical for maintaining high-quality standards in rotogravure printing. Deep learning models give prospects for automation. However, training robust deep learning models, such as YOLO or Vision Transformers, is heavily hindered b...
Dawei Li, Xiaotian Jiang, Mingyi Hong · 2026-07-23
Barzilai--Borwein (BB) method has shown strong practical performance in continuous optimization, yet its convergence dynamics remains poorly understood. In particular, a central unresolved question is whether BB converges superlinearly for almost every strictly convex quadratic problem and initialization. We provide a negative answer to this question. Specifically, for every finite dimension $n\ge...
T. Ansah-Narh, Y. Asare Afrane · 2026-07-23
A consensus anomaly detection framework was applied to monthly malaria surveillance data from Ghana (2014-2023) to identify atypical transmission patterns. Anomalies were highly structured in space and time. Ashanti and Northern Regions accounted for most recurrent anomalies, with persistent hotspots at Tamale, Kumasi, and Accra. A key finding was the spatial distinction between anomaly burden (cu...
Baihui Wang, Bernard Koch · 2026-07-23
Building socially calibrated large language models, which can learn from others without simply yielding to them, requires more than reducing sycophancy as a one-dimensional failure mode. Models must distinguish when to incorporate others' perspectives from when to maintain a well-grounded moral judgment. We study the broader resistance-compliance process governing this distinction. Across three st...
Xiao Yu, Baolin Peng, Ruize Xu · 2026-07-23
Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn reasoning, tool use, and access to external systems. While powerful, these complex harnesses also make agents hard to train end-to-end with open infrastructure, whose SFT/RL stacks cannot natively express stateful, multi-process harness inference. To address this, we present OpenForg...
Wen Ye, Yuxiao Qu, Aviral Kumar · 2026-07-23
Unlike large language models (LLMs) that exhibit strong reasoning capabilities, vision-language models (VLMs) struggle with visual reasoning, even on geometry problems that admit equivalent text, diagram, and combined diagram+text views. We show that these views often elicit different behaviors: a model may solve a problem from text but fail on the corresponding diagram, or succeed visually while...
Sophia Tang, Pranam Chatterjee · 2026-07-23
Flow-based generative models have enabled remarkable progress in fast and controllable generation across continuous and discrete state spaces, yet existing parameterizations are constrained to fixed dimensions or fixed sequence lengths. Here, we introduce Expanding Generative Flows (EFlows), which define flows between distributions of increasing dimensionality along an expanding interpolant that g...
Hongnan Ma, Yiwei Shi, Mengyue Yang · 2026-07-23
Faithful explanations of time-series classifiers should identify subsequences that are not only sufficient to preserve a black-box model's prediction, but also necessary for maintaining it. However, existing sufficiency-oriented methods can assign high importance to spurious subsequences that support the prediction without being essential to the model's decision. We introduce TimePNS, a necessity-...
Aaron Feller, Kris Deibler, Maxim Secor · 2026-07-23
Molecular property prediction from structure often uses a single representative conformation, even though many molecules exist as conformational ensembles in solution. We introduce EnsembleEGNN, a molecular ensemble foundation model that encodes an ensemble by first encoding each conformer with shared Equivariant Graph Neural Network (EGNN) layers, then pooling the resulting conformer representati...
Dongjie Fu, Di Cao, Xize Cheng · 2026-07-23
While large audio-language models have achieved remarkable progress in auditory perception, they still lag behind text-based large language models in deep logical reasoning, primarily due to the scarcity of high-quality audio reasoning data. To bridge this gap, we propose X$^3$-OPD, a cross-modal on-policy distillation framework that transfers reasoning capabilities from a powerful text teacher to...
Rodrigo Carmo Terin · 2026-07-23
The coupled ghost and gluon Dyson--Schwinger equations (DSEs) of four-dimensional Landau-gauge Yang--Mills (YM) theory are solved with a neural representation trained only from renormalized equation residuals. The neural and fixed-point solutions agree at the percent level and remain stable under changes of initialization, network size, integration grid, and infrared boundary condition. Variations...
Ryan Cotterell · 2026-07-23
Surprisal theory holds that the human processing difficulty of a linguistic unit in context is an affine function of its surprisal under some language model. I argue this claim is a tautology without further constraint: for any non-negative difficulty measure over units in context, there exists a language model whose surprisal is an affine function of it under mild technical conditions. Therefore,...
Qian Wu, Xinrong Zhou, Zizhan Ma · 2026-07-23
Large Language Models (LLMs) show promise for medical education, but most existing systems focus on localized interactions such as question answering or single-turn feedback, rather than organizing an entire clinical case into a decision-centered learning trajectory. We introduce MedGame, a framework that transforms static clinical cases into structured, executable storytelling games. MedGame uses...
Paul Azunre · 2026-07-23
We present DONDO, a family of open, permissively licensed automatic speech recognition (ASR) base models for African languages, built on the w2v-BERT 2.0 self-supervised speech encoder. DONDO comprises twenty-one monolingual models and five multilingual models spanning twenty-seven language varieties across Ghana, Sierra Leone, Nigeria, Senegal, Kenya and Zimbabwe. Models are fine-tuned primarily...
Federico Boggia · 2026-07-23
A rhetorical figure that Cicero and Quintilian catalogued two thousand years ago reappears, systematically, in the text of large language models: epanorthosis, the self-correction of the specimen «This is not a course. It is a journey of transformation». This essay argues that the overuse is a trained disposition, driven mainly by a training distribution rich in promotional prose and by preference...
Yu Qi, Zhang Ye, Xinyi Xu · 2026-07-23
Compositional generalization is essential for robot to follow diverse instructions. However, pretrained policies are known to take shortcuts, deferring to salient cues rather than grounding language. We introduce a diagnostic framework that localizes this failure to individual instruction factors, e.g., reusable semantic components such as color, verb, object, size, and spatial attribute. Our fram...
Hongxin Zhang, Chunru Lin, Junyan Li · 2026-07-23
Creating dynamic and physically realistic 4D worlds from natural language descriptions is both fascinating and challenging. Traditional computer graphics methods rely on manual creation, requiring extensive human effort to fine-tune materials, motions, and visual fidelity. Recent advances in generative foundation models have sparked interest in learning to generate such 4D worlds from large-scale...
Yihong Gu · 2026-07-23
Consider the partial linear model $Y = μ_0(X) + β_0 \cdot T + \varepsilon$ and $T = π_0(X) + u$ in the structure-agnostic setting, where we are blind to the structure $μ_0$ and $π_0$ and estimate the nuisances by a black-box hypothesis class. The learnability of the class is characterized by the estimation error $δ_s$ in the absence of model misspecification and its $L_2$ mis-specification error $...
This digest is generated automatically from arXiv submissions. Not affiliated with arXiv or Cornell University.