ArXiv Scout
Daily arXiv scanner that identifies and summarizes papers relevant to agentic control systems and AI infrastructure.
Latest real output
10 new arXiv papers on LLM agent evaluation this morning
Here are 10 new arXiv papers on LLM agent evaluation, checked at 05:08 PM PT:
- AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design by Yaxin Luo, Haobin Jiang, Jialv Zou, Xu Huang, Wenhao Yan, Haodong Li et al.
This paper introduces AutoDesign, a framework for meta-harness optimization that guides a code agent to recursively improve its harness based on rollout feedback, aligning with human design priors for transforming multimodal sources into structured media outputs. The framework focuses on academic paper-to-poster generation for instantiation and evaluation.
verified via arXiv as of 2026-08-15T00:08:18.775Z
- OmniScientist: An Omni-Modal Omni-Discipline AI Scientist by Bobo Li, Hao Fei, Tianjie Ju, Mong-Li Lee, Wynne Hsu
OmniScientist is presented as an end-to-end, omni-modal AI scientist capable of conducting multidisciplinary research directly from heterogeneous raw evidence. It aims to overcome the limitations of existing AI scientists that typically reason over restricted data types, by providing access to the full range of evidence crucial for scientific discovery.
verified via arXiv as of 2026-08-15T00:08:18.775Z
- V-RAE: Rethinking Video Latent Spaces for Generation by Minghui Guo, Shengqiong Wu, Hao Fei
V-RAE is a novel video representation autoencoder that builds compact generative latents on top of frozen vision foundation model representations, addressing the issue that current video autoencoders primarily optimize for pixel-level reconstruction rather than high-level semantic organization, which is crucial for generative modeling. It includes a lightweight temporal pooling module to reduce redundancy while maintaining semantic structure, and a video decoder for content reconstruction.
verified via arXiv as of 2026-08-15T00:08:18.775Z
- HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark by Dairu Liu, Zekun Qi, Jiayu Zeng, Ruixi Yu, Yu Guan, Yintianrun Zhang et al.
HumanTracker is introduced as a benchmark for humanoid motion tracking evaluation that is both perceptually aligned and scalable, addressing the common discrepancy between kinematic error metrics and human perception of physical artifacts in videos. The benchmark comprises approximately 153 hours of optical motion trajectories from professional performers, focusing on diverse, contact-rich, long-horizon behaviors.
verified via arXiv as of 2026-08-15T00:08:18.775Z
- PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives by Kaixin Ding, Xi Chen, Minghong Cai, Zhiyuan Xu, Yiyang Wang, Yuxiang Lu et al.
PlayWorld is a benchmark designed to facilitate the fair comparison of interactive video world models by having agent players pursue long-horizon objectives through interaction. This approach addresses the challenge of evaluating these models where fixed action-conditioned assessments fall short, as the action sequences to achieve objectives can vary significantly between models.
verified via arXiv as of 2026-08-15T00:08:18.775Z
- QuoteBench: How Matched Scores Can Hide Command-Path Failures by Shangao Li, Yao Zhang, Volker Tresp, Yuanyuan Yang
QuoteBench measures the boundary between command-generation errors and post-generation failures in LLM coding agents by using exact final-state validation on 56 one-shot tasks from 14 incident-derived families. It highlights that matched execution scores alone are insufficient to distinguish between these types of failures, especially when serialization, wrapping, and reparsing of model output by interfaces are involved.
verified via arXiv as of 2026-08-15T00:08:18.775Z
- LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure by Fanfei Li, Jana Zeller, Manuel Prada-Corral, Thaddäus Wiedemer, Prasanna Mayilvahanan, Ryan Cotterell et al.
This paper presents LITTLECURRICULUM, an 88B-token pretraining corpus curated from U.S. elementary school material to study knowledge and skill acquisition in language models under pedagogically controlled knowledge exposure. Training a 5B-parameter LLM on this corpus yields LITTLELEARNER, a model that demonstrates language competence for open-ended evaluation with interpretable knowledge and capability boundaries.
verified via arXiv as of 2026-08-15T00:08:18.775Z
- SAEVerbalizer: Generating Explanations for Sparse Autoencoder Features via Representation Verbalization by Weihan Meng, Hongzhu Guo, Yi Jing, Dewen Liu, Zijun Yao, Xiaozhi Wang et al.
SAEVerbalizer is a framework designed to generate natural-language explanations for Sparse Autoencoder (SAE) features by directly fine-tuning an LLM's downstream layers after injecting SAE decoder directions into its representations. This approach aims to provide more intrinsic and computationally efficient explanations compared to relying solely on external observation of model behavior.
verified via arXiv as of 2026-08-15T00:08:18.775Z
- Joint Communication-Control Strategy Optimization with Partially Nested Information Structures: The Linear-Quadratic Case by Haoyi You, Kaiqing Zhang
This paper formalizes a joint communication-control strategy optimization (JCCO) problem within multi-agent linear systems with quadratic costs, focusing on partially nested (PN) information structures for computational tractability. It establishes conditions under which PN is preserved under communication strategies to be optimized, noting that violating these conditions may lead to non-linear optimal strategies.
verified via arXiv as of 2026-08-15T00:08:18.775Z
- Vero: Can AI Agents Build Formally Verified Software Repositories? by Zhe Ye, Hantao Lou, Yuechun Sun, Peiyang Song, Zhengxu Yan, Timothe Kasriel et al.
Vero is introduced as the first benchmark to evaluate the joint implementation and proof synthesis by AI agents at the multi-module codebase level, addressing the open question of whether agents can make coherent implementation and proof choices across real-world software. This aims to advance trustworthy AI-generated software by moving beyond individual functions or proof generation with provided implementations.
Make it yours →One click — change what it watches, get your own alerts.
Want to change more? Open it in the builder.
Run history
14 recent checks
10 new arXiv papers on LLM agent evaluation this morning
10 new arXiv papers on LLM agent evaluation this morning
Here are 10 new arXiv papers on LLM agent evaluation, checked at 05:08 PM PT:
- AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design by Yaxin Luo, Haobin Jiang, Jialv Zou, Xu Huang, Wenhao Yan, Haodong Li et al.
This paper introduces AutoDesign, a framework for meta-harness optimization that guides a code agent to recursively improve its harness based on rollout feedback, aligning with human design priors for transforming multimodal sources into structured media outputs. The framework focuses on academic paper-to-poster generation for instantiation and evaluation.
verified via arXiv as of 2026-08-15T00:08:18.775Z
- OmniScientist: An Omni-Modal Omni-Discipline AI Scientist by Bobo Li, Hao Fei, Tianjie Ju, Mong-Li Lee, Wynne Hsu
OmniScientist is presented as an end-to-end, omni-modal AI scientist capable of conducting multidisciplinary research directly from heterogeneous raw evidence. It aims to overcome the limitations of existing AI scientists that typically reason over restricted data types, by providing access to the full range of evidence crucial for scientific discovery.
verified via arXiv as of 2026-08-15T00:08:18.775Z
- V-RAE: Rethinking Video Latent Spaces for Generation by Minghui Guo, Shengqiong Wu, Hao Fei
V-RAE is a novel video representation autoencoder that builds compact generative latents on top of frozen vision foundation model representations, addressing the issue that current video autoencoders primarily optimize for pixel-level reconstruction rather than high-level semantic organization, which is crucial for generative modeling. It includes a lightweight temporal pooling module to reduce redundancy while maintaining semantic structure, and a video decoder for content reconstruction.
[verified via arXiv as of 2026-08-15T00:08:18.775Z](http://arxiv.org/a
10 new arXiv papers on LLM agent evaluation this morning
10 new arXiv papers on LLM agent evaluation this morning
Here are 10 new papers on LLM agent evaluation as of 2026-08-12 5:08 PM PT:
- Surgical WAM: A World-Action Model for Data-Efficient Surgical Robot Learning by Wenrui Bao, Tianyun Jiang, Zhiben Chen et al.
This paper introduces a World-Action Model (WAM) to address the challenge of data scarcity in learning surgical manipulation policies for robots, by learning world models from endoscopic video for closed-loop control [verified via arXiv as of 2026-08-12 5:08 PM PT].
http://arxiv.org/abs/2608.11204v1
- ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialogue Generation in Violence Against Women and Girls by Chen Lyu, Xingwei Tan, Simon Cullen et al.
This work focuses on modeling abuse as a relational and temporally unfolding phenomenon by proposing a retrieval-grounded framework for controlled synthetic dialogue generation in sensitive domains like Violence Against Women and Girls, where real data is difficult to access [verified via arXiv as of 2026-08-12 5:08 PM PT].
http://arxiv.org/abs/2608.11200v1
- Beyond a Bag of Features: Set-Level Instability in Sparse Autoencoders by Nikolai Bolik, Lennart Stöpler, Artur Andrzejak
This paper re-examines LLM representations by using overlap over active sparse autoencoder (SAE) latent sets as a more interpretable similarity measure, finding that SAE activation sets do not recover human category boundaries as shown by prior work using cosine similarity over dense model representations [verified via arXiv as of 2026-08-12 5:08 PM PT].
http://arxiv.org/abs/2608.11197v1
- Long-Horizon AI Research for Grothendieck Constant: A Case Study in Human-AI Mathematical Collaboration by Alan Li, Rahul Saha, Anton Xue et al.
This case study details the effective use of an AI research system to impr
5 new arXiv papers on LLM agent evaluation this morning
5 new arXiv papers on LLM agent evaluation this morning
Here are 5 new arXiv papers on LLM agent evaluation as of 2026-08-09 5:08 PM PT:
- The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping by Sarvesh Baskar et al.
This paper introduces trace-grounded parametric profiling to evaluate video language models on event counting in controlled video tasks, revealing their failure at simple event bookkeeping due to the "low frequency trap."
http://arxiv.org/abs/2608.06361v1 [verified via https://arxiv.org as of 2026-08-10T00:08:33.399Z]
- Resourced Authority A Mechanism-Design Model for Participatory Governance of Deployed AI Agents by Praphul Chandra et al.
This paper proposes a formal mechanism-design model for the continuous participatory governance of deployed AI agents, focusing on resource allocation to ensure authorization is self-enforcing via compute budgets.
http://arxiv.org/abs/2608.06353v1 [verified via https://arxiv.org as of 2026-08-10T00:08:33.399Z]
- CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks by Fanzhe Meng et al.
This paper presents CalibForge, an autonomous terminal-task synthesis system that uses adversarial solver calibration to revise candidate tasks, ensuring they are appropriately challenging for learning.
http://arxiv.org/abs/2608.06352v1 [verified via https://arxiv.org as of 2026-08-10T00:08:33.399Z]
- Challenges in Evaluating Explanation Methods for Static and Evolving Data by Jerzy Stefanowski
This paper discusses limitations in evaluating Explainable AI (XAI) methods, illustrated through the DetoxAI system, and explores challenges in adapting explanations to evolving data streams with concept drift.
http://arxiv.org/abs/2608.06351v1 [verified via https://arxiv.org as of 202
10 new arXiv papers on LLM agent evaluation this morning
10 new arXiv papers on LLM agent evaluation this morning
Here are 10 new arXiv papers on LLM agent evaluation as of 5:08 PM PT on Saturday, August 8, 2026:
1. DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation
First Author(s): Junfeng Li
Summary: This paper introduces DyPES-VLA, a cross-embodiment Vision-Language-Action (VLA) model designed to learn shared dynamics priors and embodiment-specific control for robot manipulation, addressing limitations in current methods that underuse shared dynamics priors and require extensive manual action conversion.
[verified via arXiv as of 5:08 PM PT]: http://arxiv.org/abs/2608.06374v1
2. The Bitter Lesson of Tool Calling
First Author(s): Ishan Patel
Summary: This work empirically compares programmatic tool calling (PTC) to native JSON tool calling across 14 language models on BFCL v4, revealing insights into the performance of LLMs as agents using tools through code execution.
[verified via arXiv as of 5:08 PM PT]: http://arxiv.org/abs/2608.06370v1
3. Tracing the Heart: An Evidence-Linked Pipeline for Heart-Failure Feature Engineering
First Author(s): Soorya Ram Shimgekar
Summary: This paper introduces the Nimblemind Multi-Agent System (nMAS), an evidence-linked, rubric-grounded pipeline for automated heart-failure feature engineering, addressing the bottleneck of feature engineering in clinical research and AI.
[verified via arXiv as of 5:08 PM PT]: http://arxiv.org/abs/2608.06366v1
4. Investigating Artificial Intelligence Digital Sovereignty in Mobile Shopping Apps: A Case Study of Nigeria
First Author(s): George Grispos
Summary: This research examines how AI in Nigerian mobile applications affects digital sovereignty, specifically focusing on platform transparency as a key indicator of user awareness and control.
[verified via arXiv as of 5:08 PM PT]: http://arxiv.org/