40 papers across AI, ML, NLP, and CV from the last 24 hours.
A clear split is emerging between capability research and systems research. This batch leans heavily toward the latter: agentic frameworks that must maintain reasoning across dozens of tool calls, reward-shaping mechanisms for circuits and physical systems, and fine-tuning strategies designed for low-resource deployment rather than scale. The field is no longer asking "how big can we go" but "how do we make this work reliably over time."
Video generative modeling continues to mature past demonstration into constraint. World models for autonomous driving now support minute-long streaming with object-level manipulation, video foundation models are being repurposed as generative backbones, and diffusion-based weather editing addresses the real data gap in perception training. The video category is no longer just about quality — it is becoming about controllability and integration with downstream tasks.
AI security is receiving structured attention. Researchers are redefining penetration testing for systems where the attack surface lives in prompts and memory rather than binaries, training biosecurity probes on frozen genomic models, and formalizing plausible deniability for whistleblowers. These are not theoretical exercises — they assume AI is already deployed and need protection.
Hindcast stands out for identifying a quiet methodological crisis in LLM evaluation. Standard backtesting of LLM forecasters is compromised because retrieval exposes post-event reports and newer models train on data that already contains the answers. The paper demonstrates that "forecasting" performance can be an artifact of data leakage, forcing a rethink of how temporal models should be benchmarked.
The batch collectively suggests the field is shifting from capability demonstrations toward evaluation rigor, domain integration, and long-horizon reliability — markers of a discipline moving from laboratory to production.
Ximeng Mao, Nanda H. Krishna, Avery Hee-Woon Ryoo · 2026-07-15
Robust and accurate neural decoders are integral to neurotechnologies such as brain-computer interfaces and closed-loop experiments. Recent work has shown that tokenizing neural data at the spike level facilitates multi-session pretraining and delivers state-of-the-art decoding performance. However, current spike-based models are restricted to supervised learning, limiting training to datasets with paired behavioural labels. To address this, MOJO introduces a masked autoencoder approach for leveraging unlabelled neural data.
Ashutosh Jha, Michel Besserve, Simon Buchholz · 2026-07-15
Linear ICA recovers jointly independent source signals from linear mixtures. Classical algorithms maximize non-Gaussianity via proxy contrast functions. This work proposes measuring non-Gaussianity using the squared Wasserstein distance, offering a new optimal transport-based approach to ICA.
Soham Jana · 2026-07-15
Speckle noise is multiplicative noise encountered in coherent imaging modalities such as SAR and optical coherence tomography. While deep learning achieves state-of-the-art speckle denoising, fundamental statistical limits remain unexplored. This work establishes minimax theory for likelihood-based approaches.
Katie Everett · 2026-07-15
Investigates how Transformer feedforward block architecture components determine rank survival across depth at initialization. Reinterprets skip connections and normalization as mechanisms for preserving gradient rank, showing how they trade off rank collapse against ensemble-like behavior.
Mustafa Emre Gürsoy, Stefan Uhlich, Ryoga Matsuo · 2026-07-15
Introduces Lighthouse RL, a sample-efficient RL approach for analog circuit sizing with strategic reset points that initialize episodes from high-performing configurations discovered during training, addressing inefficiencies in standard RL exploration.
Leitian Tao, Baolin Peng, Wenlin Yao · 2026-07-15
Multi-turn agents solve tasks through extended sequences of tool interactions, making credit assignment fundamental. Outcome rewards become sparse and high-variance as trajectories grow. TRACE estimates turn-level credit to provide denser supervision for long-horizon agent training.
Anders Sjöberg, Nils Olsson, Marcus Baaz · 2026-07-15
Extends the EB-VAE framework to joint longitudinal and time-to-event modeling, integrating longitudinal tumor measurements, dropout information, and genetic covariates within a single population modeling framework for treatment response analysis.
Son Ha Xuan, Xuan-Bach Le, Phat T. Tran-Truong · 2026-07-15
Addresses low-resource NLP by training language and task deltas separately and recombining them in weight space, avoiding expensive language-task fine-tuning runs while adapting multilingual encoders to new languages and tasks.
Cheng Tang, Junzhi Ning, Min Cen · 2026-07-15
RL with verifiable rewards drives multimodal reasoning but answer-level correctness does not guarantee visual grounding. SIVA-RL proposes sensitivity-invariance visual alignment based on observed intervention effects rather than intervention type.
Yiming Ma, Xinyu Chen · 2026-07-15
Decoder-only Transformer for probabilistic next-return modeling on foreign-exchange bars. Separates input representation from output likelihood: continuous multivariate financial event vectors preserve numerical structure at input while maintaining discrete ordinal output.
Yunbum Kook · 2026-07-15
Advances the theoretical understanding of Dikin walk convergence on polytopes, building on interior-point method-inspired sampling approaches and improving mixing time bounds beyond prior results.
Leo Richter, Matt J. Kusner · 2026-07-15
Formalizes protection against strong-adversary threat models for whistleblowers using per-reporter differential privacy, addressing the gap where existing mechanisms do not target the natural threat model of audited organizations observing auditor selection.
Zhihao Xie, Junfeng Wu, Xinting Hu · 2026-07-15
Video generative models rely on 3D-VAE latents optimized for pixel reconstruction, limiting semantic capture. This work transforms frozen Video Foundation Model representations (V-JEPA 2, VideoMAEv2) into compact, reconstruction-capable generative latents via representation autoencoders.
Zhen Li, Zian Meng, Shuwei Shi · 2026-07-15
Explores building interactive worlds that respond coherently to player actions using video generative models as next-generation game engines, requiring interaction outcomes that follow rules over evolving game conditions rather than single-step prediction.
Ke Cheng, Hanqiao Ye, Lei Shi · 2026-07-15
Generative driving world model synthesizing future surround-view video streams and synchronized LiDAR scans with interactive object manipulation and stable minute-long streaming, addressing controllability and long-horizon stability limitations.
Jiajun Sun, Zhe Gao · 2026-07-15
Studies task-adaptive feature fusion for the ABAW11 challenge, finding that valence-arousal, categorical expressions, and facial action units benefit from different visual features, temporal strategies, fusion mechanisms, and calibration.
Shunya Shimomura, Kazuhiro Hotta · 2026-07-15
Addresses ViT's inability to explicitly reject background or irrelevant patches in visual recognition. Proposes a screening mechanism that independently evaluates patch relevance, moving beyond softmax-normalized self-attention weights.
Tung Hung Bui, Hong Hai Nguyen, Van Thong Huynh · 2026-07-15
Multimodal network for detecting ambivalence and hesitancy in unconstrained video, encoding visual, audio, and transcript streams with frozen backbones, submitted to the ABAW 11th Challenge at ECCV 2026.
Geng Li, Haiwen Li, Rui Chen · 2026-07-15
Proposes a lightweight framework for video aesthetic assessment inspired by the psychological peak-end rule, addressing the scarcity of large-scale VAA benchmarks and the intrinsic subjectivity of aesthetic judgment.
Thang-Anh-Quan Nguyen, Moussab Bennehar, Luis Guillermo Roldao Jimenez · 2026-07-15
Diffusion model for cycle-consistent weather editing using unpaired driving data, addressing the challenge of generating realistic weather effects for training autonomous driving perception without paired training data.
Xinhao Cai, Yixuan Sun, Minghang Zheng · 2026-07-15
Structure-aware framework modeling choreography as a sequence of atomic movements for music-driven dance generation, producing rhythmically synchronized and semantically consistent motion beyond continuous signal modeling.
Zhan Chen, Jiqiao Ma, Chih-wen Kuo · 2026-07-15
Multi-expert system for historical Manchu OCR reusing checkpoints from iterative fine-tuning as domain specialists, dispatching pages by visual style through a lightweight classifier to handle limited labeled data across writing styles.
Parisa Masnadi Khiabani, Wolfgang Jentner, Alireza Rangrazjeddi · 2026-07-15
Using 63 EMIT-derived Carbon Mapper plume records, demonstrates that published scalar quantities do not uniquely constrain plume boundaries, proposing genetic algorithms for uncertainty-aware consistency assessment.
Brunnhilde Ponsi, Thomas Carlier, Lara Marteau · 2026-07-15
Methodological strategy for dealing with simultaneous PET/MR cardiac data including inter-patient linkage and regional analysis for arrhythmogenic left ventricular cardiomyopathy diagnosis.
Xiao Ye, Jacob Dineen, Evan Zhu · 2026-07-15
Identifies two data leakage channels in LLM forecaster backtesting: retrieval surfacing post-event reports and training data contamination from temporal proximity. Demonstrates that "forecasting" performance may be a lookup artifact, calling for evaluation methodology reform.
Alaina Brandt · 2026-07-15
PAT (Pragmatic Auto-Translator) is a RAG-based system for whole-document, corpus-informed translation, pairing user-configured specifications with context from comparable longform texts in U.S. English and Latin American Spanish, moving beyond sentence-by-sentence translation.
S M Rafiuddin, Vamsi Krishna Pavuluri, Atriya Sen · 2026-07-15
Generates counterfactuals that flip sentiment of target aspects while preserving non-target aspects, semantic meaning, fluency, and factual consistency — a multi-constraint challenge for ABSA evaluation.
Hefeng Zhou, Jinxuan Zhang, Jiong Lou · 2026-07-15
Addresses the inefficiency of current CoT interaction where users must laboriously flag faulty steps across follow-up turns. Proposes an efficient human-AI interaction method for large reasoning models with structured error correction.
Xanthi Kokkinou, Chaido Mizeli, Nafsika Koulaxidou · 2026-07-15
Hybrid educational framework integrating conversational AI based on RAG with educational robotics for primary school earthquake preparedness, extending mechanical Lego WeDo2 simulation to cognitive and metacognitive processing.
Tam Nguyen, Hung Nguyen, Robert Ogburn · 2026-07-15
End-to-end framework applying AI acceleration across five stages of professional upskilling — knowledge acquisition, content development, review and verification, teaching, and assessment — with industry validation.
Maliha Noushin Raida, Daqing Hou · 2026-07-15
Analyzes 25,264 agentic PRs from 2,361 popular GitHub repositories to investigate adoption patterns and management of agentic coding tools at the project level, revealing how human-agent collaboration unfolds in practice.
Wenxiao Wang, Priyatham Kattakinda, Soheil Feizi · 2026-07-15
Tests whether agent-optimization gains compound over recursive optimization cycles or remain one-shot artifacts, evaluating on Terminal-Bench 2.0 to address the deployed agent setting where optimization is applied continuously.
Haoran Li, Jiebi Deng, Tong Jin · 2026-07-15
HealthClaw: an open-source agent architecture that updates health support as routines, preferences, and risks change longitudinally. Separates shared safety rules from private memory containing profile facts, reusable procedures, and episodic traces.
Mohammad Allahbakhsh, Mohammad Hassan Bahari, Moslem Attar-Raouf · 2026-07-15
Argues that traditional penetration testing is insufficient for AI-enabled systems where adversaries influence prompts, retrieved content, training data, memory, tools, or human-AI interaction loops to alter behavior without compromising underlying infrastructure.
Ke Cheng, Hanqiao Ye, Lei Shi · 2026-07-15
Also cross-listed in cs.RO. Multi-view multimodal generative model for autonomous driving simulation with surround-view video synthesis, LiDAR scan generation, and interactive object manipulation over minute-long horizons.
Yağız Gença, Stefan Uhlich, Andrea Bonetti · 2026-07-15
Genetic algorithm for automated analog circuit synthesis with joint topology and sizing optimization, transferring NEAT principles to the analog circuit domain through reformulated genome representation and adapted genetic operators.
Ashutosh Jha, Michel Besserve, Simon Buchholz · 2026-07-15
Also cross-listed in stat.ML. Optimal transport-based approach to ICA using squared Wasserstein distance as non-Gaussianity measure.
Soham Jana · 2026-07-15
Also cross-listed in stat.ML. Fundamental statistical limits and minimax theory for deep learning approaches to multiplicative noise regression.
Mohammad Allahbakhsh, Mohammad Hassan Bahari, Moslem Attar-Raouf · 2026-07-15
Also cross-listed in cs.CR. Proposes reframing penetration testing from resource compromise to behavioral objective violation for AI systems with prompt-level, memory-level, and interaction-loop attack surfaces.
This digest is generated automatically from arXiv submissions. Not affiliated with arXiv or Cornell University.