40 papers across AI, ML, NLP, and CV from the last 24 hours.
Today's batch reflects a field pivoting from capability demonstration to structural engineering of AI systems. Three themes dominate: agentic infrastructure, inference-time reasoning, and the quiet maturation of evaluation methodology.
The agentic theme is unavoidable. Sixteen papers orbit it directly — from coding agents that construct executable world models for unknown games (Twin), to multi-agent frameworks for grounded technical reports (Wyvern), to spreadsheet reasoning via hierarchical relation graphs (SheetCompass). The research is no longer asking whether agents can be built; it's solving the plumbing of session handovers, evidence aggregation, and cross-environment adaptation. The PACE-Bench benchmark tests whether agents can recover when the physics of their world changes — a stress test for systems that previously only saw static evaluations.
Inference-time computation draws equal attention. "You Only Pass Once" extracts both answers and abstention signals from a single frozen forward pass, while Power Sampling exposes a paradox where sharpening probability mass toward correct trajectories actually degrades downstream performance. These papers treat the model as fixed and the computation around it as the design surface.
The standout is Toby Ord's mathematics of intelligence explosions. Most discussions of AI self-improvement rely on economic growth analogies; Ord shows that singular growth toward a vertical asymptote is harder to achieve than the models assume, and identifies a neglected class of feedback dynamics. It's one of the few papers today that treats AI capability scaling as a rigorous mathematical problem rather than an empirical bet.
What's absent is telling: no major papers on new foundation model architectures, no scaling law results, no new multimodal pretraining runs. The field appears to be consolidating around existing models and investing in everything around them — evaluation, optimization, safety, and deployment infrastructure. The collective signal suggests a discipline entering its engineering phase: less about what models can do, more about how to build systems that use them reliably.
Masahiro Kato, Taka Kato · 2026-08-14
This study investigates the methodological and theoretical properties of session handover in applications that use large language models. A task may continue in a new session when the context reaches the model's input limit, when the application restarts, or when another agent is asked to finish the task. The application must then decide which information from the earlier session to pass on. We fo
Taenyun Kim, Edyta Bogucka, Daniele Quercia · 2026-08-14
As AI systems make more morally loaded decisions across society, one response has been moral preference elicitation. In this approach, researchers poll participants on hypothetical dilemmas and use the aggregated votes to train a policy that an AI model then applies at scale. Before any vote is cast, developers make three key choices in the moral AI elicitation pipeline: feature scoping, voter sam
Yubo Zhang, Yiyao Liu, Xiaodong Wang · 2026-08-14
High-order multiple-input multiple-output (MIMO) detection requires efficient search over a large discrete symbol space while producing reliable soft information for channel decoding. This paper develops a learning-to-transition (L2T) framework that formulates MIMO detection as a stochastic sequence of complete-vector transitions. At each transition, a channel-coupled Transformer updates both the
Zhelun Wu · 2026-08-14
Systems that ask a language model to reach a conclusion from many sources usually concatenate them into one prompt. This conflates two operations with different requirements. Interpreting a source rewards capacity and context. Combining interpretations rewards fixed arithmetic, comparability across instances, and the option to return nothing. Once separated, the design problem becomes the interfac
Juno Nam, Bowen Deng, Xiaochen Du · 2026-08-14
Free energies govern solid-state phase stability, yet computational materials discovery still relies largely on ground-state energies because free energy calculations require ensemble averages. We introduce the thermodynamic interatomic potential (TIP), which extends an interatomic potential from its static energy to a thermodynamically consistent Gibbs free energy model, with thermodynamic respon
Charitha Nandepu, Lohitha Kalepu, Gabriele Ciavarella · 2026-08-14
Network-level maintenance planning requires repeated evaluations of equilibrium traffic flows under road capacity reductions. While equilibrium traffic assignment models are well established, their repeated solution quickly becomes computationally prohibitive and challenging to embed within maintenance scheduling problems. This paper investigates data-driven surrogate models that approximate equil
Alexy Skoutnev, Kirill Acharya, Gaston Longhitano · 2026-08-14
We present a Test-time World-model Inference (Twin) system, in which a frontier coding agent writes an executable world model for completing continual learning tasks, such as ARC-AGI-3 games. Traditional approaches hand-engineer such models, one custom design per task. Each game hides its rules and goal, and our system constructs them from simulation and interaction alone. Its inductive prior over
Shahab Band, Hamed Mohammadi · 2026-08-14
Short-horizon forecasting of fine particulate matter (PM2.5) remains difficult when observations from the target domain are limited and the statistical properties of the source and target domains differ. In these settings, models trained only on local data may not capture complex temporal dynamics, while direct transfer learning can result in negative transfer. This study develops a shift-aware du
Panjing He, Mingyue Cheng, Yucong Luo · 2026-08-14
Spreadsheets are widely used to organize, analyze, and manipulate semi-structured data, yet automated spreadsheet reasoning remains challenging for large language models (LLMs). Real-world workbooks often contain implicit cross-table associations, fine-grained column dependencies, and complex spatial layouts. Existing methods typically flatten these multidimensional structures into sequential stri
Beatrice Alessandra Motetti, Emilien Guandalino, Daniele Jahier Pagliari · 2026-08-14
In the current artificial intelligence-driven innovation era, the pace of knowledge growth is accelerating, and is hard to keep up with. While generative models are increasingly used to synthesize content, they often lack in information grounding. To address these peculiarities of our time, we propose Wyvern, a multi-agent framework for the automated generation of grounded, multimodal technical re
Yuhao Zhan, Bingxiang He, Zecong Tang · 2026-08-14
Self-evolving agents improve future behavior from interaction experience, yet existing evaluations typically optimize under fixed execution conditions and do not test recovery after those conditions change. To address this gap, we introduce PACE-Bench (Physics Adaptation via Code Evolution), a simulator-grounded benchmark of 144 source-to-target adaptation pairs across six physics domains. Each pa
Toby Ord · 2026-08-14
AI is increasingly being used to help with AI R&D. Under certain conditions this feedback loop might be able to produce an intelligence explosion, with rapidly escalating AI capabilities. I explore the mathematics of the most explosive possibilities, with an eye to understanding what drives the dynamics. I show that singular growth (towards a vertical asymptote) is harder to achieve than would be
Toby D. Pilditch · 2026-08-14
LLM evaluations often use fixed sampling budgets, testing every item the same number of times even after estimates are precise. We introduce optstop, a precision-based adaptive stopping framework that treats evaluation as a sequential measurement problem: keep sampling where uncertainty remains high, and stop where estimates are precise or stable enough. The framework builds on hierarchical Bayesi
Ross D. King · 2026-08-14
We present a survey of the past and future of AI Scientists: machines capable of automating science. AI Scientists can originate hypotheses, deduce their consequences, design and execute experiments, interpret their results, and revise their beliefs. Such systems are integrated scientific agents, connected to the literature, formal knowledge, mathematical models, simulations, data-analysis systems
Wen-Fan Wang, TsaiHsuan Lin, Chi-Lan Yang · 2026-08-14
Art style is a signature of professional digital artists that develops through repeated experimentation, reflection, and adaptation. While generative AI (GenAI) can reproduce styles with high fidelity, current tools provide limited support for exploring new stylistic directions and may encourage style replication over exploration. To address this gap, we propose Analyze-Experiment-Resituate (AER),
Syeda Anshrah Gillani, Mirza Samad Ahmed Baig · 2026-08-14
Patients increasingly ask large language model (LLM) assistants which doctor to see, making these systems AI infomediaries: algorithms that intermediate one person's choice among other people and thereby decide, silently and at scale, which physicians become visible. We report a prespecified randomized algorithm audit of what causally moves those recommendations. Seven models (six open-weight; gpt
Pin-Yen Huang, Sachin Chhabra, Prasanth Sai Gouripeddi · 2026-08-14
Recipe data arises in domains such as materials synthesis, pharmaceutical formulation, and industrial manufacturing, where procedures are represented as ordered sequences of steps containing heterogeneous structured fields. Existing tabular learning methods typically flatten this structure into fixed-schema representations, limiting their ability to capture hierarchical field interactions and proc
Hanfeng Lu, Tianyu Feng, Suyi Li · 2026-08-14
Vision-language models (VLMs) enable embodied agents to reason and act from visual observations and language instructions. Reinforcement learning (RL) post-training enhances these capabilities using task feedback, but current on-policy RL runtimes execute rollout, reference scoring, and actor training in strict serial phases. While effective for text-only RL, this phase-granular execution is waste
Hao Yan, Lisa Pilgram, Dan Liu · 2026-08-14
Cross-Tabular Data Generation (CTDG) seeks to learn a generative model from multiple heterogeneous tables and produce new synthetic tabular datasets. However, existing synthetic tabular data generation methods are largely restricted to single-input-table scenarios and struggle to effectively handle multiple heterogeneous tables with diverse feature sets. To address this limitation, we propose a tw
Ben Anson, Conor Houghton, Edward Milsom · 2026-08-14
The Muon optimizer shows clear benefits versus alternatives when pretraining neural networks. However, it is used less frequently for parameter-efficient fine-tuning (PEFT). One potential reason is that the most common PEFT method, LoRA, does not naturally combine with Muon since it is not mathematically possible to orthogonalize the weight update given by a low-rank parameterization. In this pape
Abhishek Shukla, Ankur Sinha, Faiz Hamid · 2026-08-14
Neural Architecture Search (NAS) aims to automate neural network architecture design, reducing reliance on human expertise. Among the various NAS methods, differentiable NAS has gained prominence due to its efficiency and accuracy compared to conventional NAS approaches. Since differentiable NAS relaxes the architecture search space into a continuous domain, it is possible to apply principles from
Abhishek Shukla, Ankur Sinha, Faiz Hamid · 2026-08-14
Neural Architecture Search (NAS) is naturally formulated as a bilevel optimization problem, where the upper-level optimizes the architecture using validation performance and the lower-level trains network parameters using training loss. However, NAS is computationally expensive due to discrete architectural decisions, exponentially growing search spaces, and the high cost of training candidate arc
Yixian Xu, Yuanrui Zhang, Shengjie Luo · 2026-08-14
Reinforcement learning (RL) post-training provides a direct way to align diffusion models with human preferences and task-specific rewards. However, current RL algorithms for diffusion models remain fragmented: reverse-trajectory methods rely on discretized likelihood ratios, whereas forward-matching methods train on reward-labeled noising versions of the rollout samples. This paper shows that the
Haohui Yang, Jiaxing Sun, Xiujun Ma · 2026-08-14
Power Sampling sharpens a language model's distribution over complete generation trajectories, offering a verifier-free way to improve reasoning at inference time. It also has the potential to serve as a general-purpose front end for a broad range of downstream sampling methods. However, we uncover a striking paradox: Power Sampling can drive more probability mass toward correct trajectories while
Syed Abdul Haseeb Qadri, Bjarne C. Hiller, Felix Blanke · 2026-08-14
Cytometry measures the complex characteristics of single cells (e.g., counts and protein expression of immune cells) and is widely used across immunological research and clinical settings. However, cytometry data is highly heterogeneous and unstandardized due to experimental protocols and the choice of measured features. While machine learning methods hold the potential to gain deeper insights int
Qinye Zhou, Jun Zheng, Yongchao Du · 2026-08-14
With the rapid advancement of image editing models and their widespread application across various domains, there is an increasingly urgent need to deploy these model capabilities directly into real-world scenarios. However, existing benchmarks remain confined to simple single-image tasks, suffering from limited coverage dimensions and an inability to effectively differentiate performance among di
Mahesh Reddy, Yashesh Savani, Antoine Mercier · 2026-08-14
High-resolution image restoration from degraded inputs is challenging because it must preserve global structural consistency while recovering fine-grained local details, especially at 4K resolution where direct diffusion-based restoration is computationally expensive and prone to repeated or inconsistent textures. In this work, we introduce MagnifiQ, an image restoration framework that progressive
Karel Becerra, Boris Mederos, Dean Snow · 2026-08-14
Determining the biological sex of the individuals who created Upper Paleolithic hand stencils remains a challenging problem due to the absence of ground truth, population differences between contemporary and prehistoric groups, and the uncertainty introduced by image degradation. Traditional morphometric methods suffer from high structural overlap across sexes, poor cross-population generalizabili
Zian Meng, Zhen Li, Chuanhao Li · 2026-08-14
Interactive game world models typically autoregress visual observations directly in pixel or latent space, forcing structured properties such as pose, geometry, and occlusion to be implicitly maintained by the same generative sequence. Over long horizons, errors in these latent world properties accumulate, making consistency and controllability fragile. We explicitly model the evolving world state
Rory Ashton · 2026-08-14
Frozen image embeddings from models such as CLIP are increasingly used to classify paintings by art-historical style, with high reported accuracy. We ask whether this accuracy reflects an understanding of style or the recognition of individual artists. Standard evaluation uses random splits in which works by the same artist appear on both sides, so a classifier can succeed by recognising the paint
Mohamed Abdelsamad, Bin Yang, Michael Ulrich · 2026-08-14
3D object detection from LiDAR point clouds is a core problem in autonomous driving. Recent advances in self-supervised learning (SSL) enable scalable pretraining and transfers well to per-point tasks such as semantic and panoptic segmentation, but transfer to 3D detection remains weaker. We analyze recent SSL methods and find that most objectives are defined only on measured LiDAR returns from vi
Mahdi Saberi, Toygan Kiliç, Mehmet Akçakaya · 2026-08-14
MRI reconstruction from undersampled k-space measurements is an ill-posed inverse problem. Physics-driven deep learning (PD-DL) methods have shown strong performance for this task by combining the MRI forward model with learned image regularization within algorithm-unrolling frameworks. However, most existing PD-DL methods reconstruct complex-valued images directly, thereby implicitly coupling mag
Jihun Park, Kyoungmin Lee, Jongmin Gim · 2026-08-14
Subject-driven image personalization---generating new images that preserve the identity of one or several reference subjects in novel scenes---is a foundational capability for modern visual content creation. It is currently dominated by generalized methods that fine-tune a pretrained multimodal diffusion transformer (MMDiT) on hundreds of thousands to millions of paired \emph{(reference, composed-
Yijiao Zhang, Hongzhe Li · 2026-08-14
Modern generative models increasingly produce distribution-valued outputs, such as predicted cellular responses to genetic perturbations in single-cell genomics. While these models provide valuable auxiliary information, they are inherently imperfect, creating a need for statistical methods that leverage their predictions without relying on their correctness. We propose generation-powered inferenc
Yang Peng, Liangyu Zhang · 2026-08-14
We study online statistical inference for functionals of the return distribution under a fixed policy. The return distribution is estimated by nonparametric distributional temporal-difference learning from a single Markov trajectory. For the Polyak--Ruppert averaged estimator, we prove that its root-$T$ error converges weakly to a centered Gaussian random element in Cramér space. We also prove tha
Xiaohong Chen, Yuling Jiao, Lican Kang · 2026-08-14
In offline RL, estimating the optimal action-value function $Q^*$ can be formulated as solving the optimal Bellman equation based solely on offline observations. A fundamental challenge is that the reward function and transition kernel are unknown, so the optimal Bellman operator is not directly observable from data. To address this issue, we propose a novel framework that decouples operator estim
Alexei Odinokov, Rostislav Yavorskiy · 2026-08-14
As heterogeneous robotic systems deploy across diverse urban zones, maintaining safety amid complex human-robot interactions remains a critical challenge. We present a unified framework that bridges systematic hazard analysis and runtime enforcement using hazard-informed safety envelopes. Rather than treating safety as a static constraint isolated within individual software modules, we introduce a
Ajith Anil Meera, Pablo Lanillos, Wouter Kouw · 2026-08-14
An autonomous robot efficiently exploring an unknown environment, such as looking for water sources on Mars, faces two simultaneous demands: building an accurate information map while quickly finding the regions of greatest value, and paying for every meter of travel and the cost of every measurement it takes. Classical information-seeking and reward-seeking criteria address only one of these obje
Ziyang Luo, Zhongyao Chu, Xinjie He · 2026-08-14
A frozen language model on reasoning tasks has two coupled weaknesses: it under-uses evidence its own residual stream already encodes, and it fails to detect when the input is insufficient to answer, so it confabulates. This paper consolidates two research lines that address these on the same residual stream: a conditional steering probe writes the stream at mid-stack layers and recovers reasoning
Isabel Cachola, William Walden, Reno Kriz · 2026-08-14
The majority of work on summarization evaluation focuses on general summary quality (e.g., ROUGE, BERTScore) or specific desired properties (e.g., readability, factuality). However, these metrics fail to measure the utility of a summary to an individual user. For example, a biomedical researcher learning about the latest vaccine research will have different informational needs from a family doctor
This digest is generated automatically from arXiv submissions. Not affiliated with arXiv or Cornell University.