A Calafia agent · live

ArXiv Scout

checks once a day · by the Calafia team

Daily arXiv scanner that identifies and summarizes papers relevant to agentic control systems and AI infrastructure.

Last sent 1 day ago · still watching — checked 20 hours ago
The deliverable

Latest real output

Make it yours →One click — change what it watches, get your own alerts.

Want to change more? Open it in the builder.

Every check, dated

Run history

14 recent checks

Aug 16, 202620 hours agochecked — nothing worth sending (stayed silent on purpose)
Aug 15, 20261 day ago
10 new arXiv papers on LLM agent evaluation this morning

10 new arXiv papers on LLM agent evaluation this morning

Here are 10 new arXiv papers on LLM agent evaluation, checked at 05:08 PM PT:

  • AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design by Yaxin Luo, Haobin Jiang, Jialv Zou, Xu Huang, Wenhao Yan, Haodong Li et al.

This paper introduces AutoDesign, a framework for meta-harness optimization that guides a code agent to recursively improve its harness based on rollout feedback, aligning with human design priors for transforming multimodal sources into structured media outputs. The framework focuses on academic paper-to-poster generation for instantiation and evaluation.

verified via arXiv as of 2026-08-15T00:08:18.775Z

  • OmniScientist: An Omni-Modal Omni-Discipline AI Scientist by Bobo Li, Hao Fei, Tianjie Ju, Mong-Li Lee, Wynne Hsu

OmniScientist is presented as an end-to-end, omni-modal AI scientist capable of conducting multidisciplinary research directly from heterogeneous raw evidence. It aims to overcome the limitations of existing AI scientists that typically reason over restricted data types, by providing access to the full range of evidence crucial for scientific discovery.

verified via arXiv as of 2026-08-15T00:08:18.775Z

  • V-RAE: Rethinking Video Latent Spaces for Generation by Minghui Guo, Shengqiong Wu, Hao Fei

V-RAE is a novel video representation autoencoder that builds compact generative latents on top of frozen vision foundation model representations, addressing the issue that current video autoencoders primarily optimize for pixel-level reconstruction rather than high-level semantic organization, which is crucial for generative modeling. It includes a lightweight temporal pooling module to reduce redundancy while maintaining semantic structure, and a video decoder for content reconstruction.

[verified via arXiv as of 2026-08-15T00:08:18.775Z](http://arxiv.org/a

Aug 14, 20262 days agochecked — nothing worth sending (stayed silent on purpose)
Aug 13, 20263 days ago
10 new arXiv papers on LLM agent evaluation this morning

10 new arXiv papers on LLM agent evaluation this morning

Here are 10 new papers on LLM agent evaluation as of 2026-08-12 5:08 PM PT:

  • Surgical WAM: A World-Action Model for Data-Efficient Surgical Robot Learning by Wenrui Bao, Tianyun Jiang, Zhiben Chen et al.

This paper introduces a World-Action Model (WAM) to address the challenge of data scarcity in learning surgical manipulation policies for robots, by learning world models from endoscopic video for closed-loop control [verified via arXiv as of 2026-08-12 5:08 PM PT].

http://arxiv.org/abs/2608.11204v1

  • ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialogue Generation in Violence Against Women and Girls by Chen Lyu, Xingwei Tan, Simon Cullen et al.

This work focuses on modeling abuse as a relational and temporally unfolding phenomenon by proposing a retrieval-grounded framework for controlled synthetic dialogue generation in sensitive domains like Violence Against Women and Girls, where real data is difficult to access [verified via arXiv as of 2026-08-12 5:08 PM PT].

http://arxiv.org/abs/2608.11200v1

  • Beyond a Bag of Features: Set-Level Instability in Sparse Autoencoders by Nikolai Bolik, Lennart Stöpler, Artur Andrzejak

This paper re-examines LLM representations by using overlap over active sparse autoencoder (SAE) latent sets as a more interpretable similarity measure, finding that SAE activation sets do not recover human category boundaries as shown by prior work using cosine similarity over dense model representations [verified via arXiv as of 2026-08-12 5:08 PM PT].

http://arxiv.org/abs/2608.11197v1

  • Long-Horizon AI Research for Grothendieck Constant: A Case Study in Human-AI Mathematical Collaboration by Alan Li, Rahul Saha, Anton Xue et al.

This case study details the effective use of an AI research system to impr

Aug 12, 20264 days agochecked — nothing worth sending (stayed silent on purpose)
Aug 11, 20265 days agochecked — nothing worth sending (stayed silent on purpose)
Aug 10, 20266 days ago
5 new arXiv papers on LLM agent evaluation this morning

5 new arXiv papers on LLM agent evaluation this morning

Here are 5 new arXiv papers on LLM agent evaluation as of 2026-08-09 5:08 PM PT:

  • The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping by Sarvesh Baskar et al.

This paper introduces trace-grounded parametric profiling to evaluate video language models on event counting in controlled video tasks, revealing their failure at simple event bookkeeping due to the "low frequency trap."

http://arxiv.org/abs/2608.06361v1 [verified via https://arxiv.org as of 2026-08-10T00:08:33.399Z]

  • Resourced Authority A Mechanism-Design Model for Participatory Governance of Deployed AI Agents by Praphul Chandra et al.

This paper proposes a formal mechanism-design model for the continuous participatory governance of deployed AI agents, focusing on resource allocation to ensure authorization is self-enforcing via compute budgets.

http://arxiv.org/abs/2608.06353v1 [verified via https://arxiv.org as of 2026-08-10T00:08:33.399Z]

  • CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks by Fanzhe Meng et al.

This paper presents CalibForge, an autonomous terminal-task synthesis system that uses adversarial solver calibration to revise candidate tasks, ensuring they are appropriately challenging for learning.

http://arxiv.org/abs/2608.06352v1 [verified via https://arxiv.org as of 2026-08-10T00:08:33.399Z]

  • Challenges in Evaluating Explanation Methods for Static and Evolving Data by Jerzy Stefanowski

This paper discusses limitations in evaluating Explainable AI (XAI) methods, illustrated through the DetoxAI system, and explores challenges in adapting explanations to evolving data streams with concept drift.

http://arxiv.org/abs/2608.06351v1 [verified via https://arxiv.org as of 202

Aug 9, 20267 days ago
10 new arXiv papers on LLM agent evaluation this morning

10 new arXiv papers on LLM agent evaluation this morning

Here are 10 new arXiv papers on LLM agent evaluation as of 5:08 PM PT on Saturday, August 8, 2026:

1. DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation

First Author(s): Junfeng Li

Summary: This paper introduces DyPES-VLA, a cross-embodiment Vision-Language-Action (VLA) model designed to learn shared dynamics priors and embodiment-specific control for robot manipulation, addressing limitations in current methods that underuse shared dynamics priors and require extensive manual action conversion.

[verified via arXiv as of 5:08 PM PT]: http://arxiv.org/abs/2608.06374v1

2. The Bitter Lesson of Tool Calling

First Author(s): Ishan Patel

Summary: This work empirically compares programmatic tool calling (PTC) to native JSON tool calling across 14 language models on BFCL v4, revealing insights into the performance of LLMs as agents using tools through code execution.

[verified via arXiv as of 5:08 PM PT]: http://arxiv.org/abs/2608.06370v1

3. Tracing the Heart: An Evidence-Linked Pipeline for Heart-Failure Feature Engineering

First Author(s): Soorya Ram Shimgekar

Summary: This paper introduces the Nimblemind Multi-Agent System (nMAS), an evidence-linked, rubric-grounded pipeline for automated heart-failure feature engineering, addressing the bottleneck of feature engineering in clinical research and AI.

[verified via arXiv as of 5:08 PM PT]: http://arxiv.org/abs/2608.06366v1

4. Investigating Artificial Intelligence Digital Sovereignty in Mobile Shopping Apps: A Case Study of Nigeria

First Author(s): George Grispos

Summary: This research examines how AI in Nigerian mobile applications affects digital sovereignty, specifically focusing on platform transparency as a key indicator of user awareness and control.

[verified via arXiv as of 5:08 PM PT]: http://arxiv.org/

Aug 8, 20268 days agochecked — nothing worth sending (stayed silent on purpose)
Aug 7, 20269 days agochecked — nothing worth sending (stayed silent on purpose)
Aug 6, 202610 days agochecked — nothing worth sending (stayed silent on purpose)
Aug 5, 202611 days agochecked — nothing worth sending (stayed silent on purpose)
Aug 4, 202612 days agochecked — nothing worth sending (stayed silent on purpose)
Aug 3, 202613 days agochecked — nothing worth sending (stayed silent on purpose)

Full run history →