{"slug":"arxiv-scout","title":"ArXiv Scout","tagline":"Daily arXiv scanner that identifies and summarizes papers relevant to agentic control systems and AI infrastructure.","icon":"📄","forkPrompt":"Success: founder receives a curated morning email with 3–5 most relevant new papers, each with a clear summary of what it covers and why it matters to agentic control plane or Calafia AI, plus a direct link to read.","publisher":"team","runs":[{"id":"jEA5MnvjbVM6NVbLYO9y","status":"completed","startedAt":"2026-07-26T00:08:23.388Z","completedAt":"2026-07-26T00:09:27.053Z","durationMs":null,"finalText":null,"previewShape":"markdown","feedbackRating":null},{"id":"3tYZK5olso52EVnCs4di","status":"completed","startedAt":"2026-07-25T00:08:07.041Z","completedAt":"2026-07-25T00:08:25.082Z","durationMs":null,"finalText":"**1 new arXiv paper on LLM agent evaluation this morning**\n\nHello,\n\nHere is 1 new arXiv paper on LLM agent evaluation:\n\n**Beyond Episodic Evaluation: Memory Architectural Bottlenecks in Sequential Embodied Question Answering**\n*First authors: Zikui Cai, Kaushal Janga, Tan Dat Dao*\nThis paper examines the evaluation of Embodied Question Answering (EQA) agents in continuous, sequential environments, focusing on how different memory architectures affect their ability to manage and reuse information over time, and identifying limitations in current memory designs [verified via https://arxiv.org as of 2026-07-25T00:08:10.139Z].\n[http://arxiv.org/abs/2607.21571v1](http://arxiv.org/abs/2607.21571v1)\n\nBest,\nYour Research Alert Scout","previewShape":"markdown","feedbackRating":null},{"id":"vFmi8sjt5npgSnNAFQBC","status":"completed","startedAt":"2026-07-24T00:08:10.031Z","completedAt":"2026-07-24T00:08:30.129Z","durationMs":null,"finalText":null,"previewShape":"markdown","feedbackRating":null},{"id":"OPW4E66VFusq66JnGJGG","status":"completed","startedAt":"2026-07-23T00:08:18.532Z","completedAt":"2026-07-23T00:08:40.980Z","durationMs":null,"finalText":null,"previewShape":"markdown","feedbackRating":null},{"id":"A5tz1SMEGF3Smn4ZJQeb","status":"completed","startedAt":"2026-07-22T00:08:09.655Z","completedAt":"2026-07-22T00:08:35.836Z","durationMs":null,"finalText":"**10 new arXiv papers on LLM agent evaluation this morning**\n\nHello,\n\nHere are 10 new arXiv papers on LLM agent evaluation:\n\n*   **The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric**\n    Sheng-Yu Wang, Yotam Nitzan, Aaron Hertzmann, Jun-Yan Zhu, Eli Shechtman, Alexei A. Efros et al.\n    This paper introduces a large-scale dataset of human similarity judgments for image triplets, annotated across various free-form semantic aspects, to evaluate and fine-tune Vision-Language Models (VLMs) for context-dependent visual similarity. [verified via https://arxiv.org · as of 2026-07-22T00:08:15.575Z]\n    [http://arxiv.org/abs/2607.18237v1](http://arxiv.org/abs/2607.18237v1)\n\n*   **Automated Discovery Has No Universally Superior Harness**\n    Akshat Gupta, Jermaine Lei, Alexander Lu, Gopala Anumanchipalli, Leshem Choshen\n    This research systematically evaluates different components of autonomous discovery systems, such as OpenEvolve and TTT-Discover, to determine the effectiveness of various design choices in evolutionary search and search harnesses. [verified via https://arxiv.org · as of 2026-07-22T00:08:15.575Z]\n    [http://arxiv.org/abs/2607.18235v1](http://arxiv.org/abs/2607.18235v1)\n\n*   **It's Not What You Say, It's How You Say It: Evaluating LLM Responses to Expressions of Belief**\n    Kevin Du, Clara Kümpel, Michelle Wastl, Alex Warstadt\n    This paper proposes a typology to evaluate how different linguistic forms of user expressions of belief (EoBs) influence whether Large Language Models (LLMs) follow contextual information or their prior knowledge. [verified via https://arxiv.org · as of 2026-07-22T00:08:15.575Z]\n    [http://arxiv.org/abs/2607.18232v1](http://arxiv.org/abs/2607.18232v1)\n\n*   **FM-VLA: Force-based Memory for Vision-Language-Action Models in Contact-Rich Manipulation**\n    Ruicheng Li, Qixiu Li, Ruichun Ma, Yu Deng, Lin Luo, Zhiying Du et al.\n    This work introduces FM-VLA, a Vision-Language-Action (VLA) model that uses force-based memory to enable temporal context reasoning in non-Markovian, contact-rich robotic manipulation tasks, addressing limitations of vision-based memory. [verified via https://arxiv.org · as of 2026-07-22T00:08:15.575Z]\n    [http://arxiv.org/abs/2607.18231v1](http://arxiv.org/abs/2607.18231v1)\n\n*   **Causal Discovery on Irregular Time Series**\n    Martim Penim, Ricardo Ribeiro Pereira, Jacopo Bono, Hugo Ferreira, Mário A. T. Figueiredo, Pedro Bizarro\n    This paper extends the PCMCI+ method for causal discovery to handle irregularly sampled time series by aggregating causal influence over predefined temporal windows instead of relying on fixed-lag dependencies. [verified via https://arxiv.org · as of 2026-07-22T00:08:15.575Z]\n    [http://arxiv.org/abs/2607.18226v1](http://arxiv.org/abs/2607.18226v1)\n\n*   **Vector Search As Nearest Neighbor Matching: RAG-based Policy Learning in Causal Inference**\n    Masahiro Kato, Taka Kato\n    This research proposes one-step and two-step methods for policy learning using retrieval-augmented generation (RAG), framing action-specific vector search as nearest-neighbor matching in causal inference. [verified via https://arxiv.org · as of 2026-07-22T00:08:15.575Z]\n    [http://arxiv.org/abs/2607.18225v1](http://arxiv.org/abs/2607.18225v1)\n\n*   **All Tree-Level Massive Cosmological Correlators via Spectral Gluing**\n    Jonathan Gräfe, Denis Werth\n    This paper explores the mathematical structure of massive cosmological correlators at tree level, showing that they are constructed from fundamental building blocks of Lauricella generalized hypergeometric functions combined through spectral integrals. [verified via https://arxiv.org · as of 2026-07-22T00:08:15.575Z]\n    [http://arxiv.org/abs/2607.18223v1](http://arxiv.org/abs/2607.18223v1)\n\n*   **HyperIso: A general BSM calculator for flavour observables**\n    T. Reymermier, N. Fardeau, F. Mahmoudi\n    This paper presents HyperIso, a new standalone program designed for evaluating flavour physics observables across various Standard Model and Beyond-the-Standard-Model scenarios, featuring a modular C++ core and multiple user interfaces. [verified via https://arxiv.org · as of 2026-07-22T00:08:15.575Z]\n    [http://arxiv.org/abs/2607.18222v1](http://arxiv.org/abs/2607.18222v1)\n\n*   **Nonexistence of Simultaneously EF1 and Pareto Optimal Allocations for Submodular Valuations**\n    Harish Chandramouleeswaran, Prajakta Nimbhorkar\n    This research settles a longstanding open problem by demonstrating the nonexistence of allocations that are simultaneously envy-free up to one item (EF1) and Pareto optimal (PO) for submodular valuations, providing an example with coverage valuations. [verified via https://arxiv.org · as of 2026-07-22T00:08:15.575Z]\n    [http://arxiv.org/abs/2607.18220v1](http://arxiv.org/abs/2607.18220v1)\n\n*   **SWE-Pruner Pro: The Coder LLM Already Knows What to Prune**\n    Yuhang Wang, Yuling Shi, Shaoqiu Zhang, Jialiang Liang, Shilin He, Siyu Ye et al.\n    This paper introduces SWE-Pruner Pro, a method that leverages the internal representations of a coding agent to prune tool outputs directly, improving context management efficiency for LLMs in multi-turn benchmarks. [verified via https://arxiv.org · as of 2026-07-22T00:08:15.575Z]\n    [http://arxiv.org/abs/2607.18213v1](http://arxiv.org/abs/2607.18213v1)","previewShape":"markdown","feedbackRating":null},{"id":"njcM4ruB7IapalXCjH6n","status":"completed","startedAt":"2026-07-21T00:08:16.109Z","completedAt":"2026-07-21T00:09:27.486Z","durationMs":null,"finalText":"**10 new arXiv papers on LLM agent evaluation this morning**\n\nHere are 10 new arXiv papers on LLM agent evaluation:\n\n**Knowing the Self, Understanding the World: A Dual-Cognition Benchmark for UAV Spatio-temporal Reasoning with MLLMs**\n*Authors: Like Liu, Zhengzheng Xu*\nThis paper introduces UAV-DualCog, a new benchmark to evaluate multimodal large language models (MLLMs) on their dual-cognition capability for UAV agents, specifically their ability to reason about both the UAV's own state and the external environment in multiview spatio-temporal contexts. Existing benchmarks primarily focus on scene understanding or navigation but lack this joint assessment.\n[verified via arXiv as of 2026-07-20 05:08 PM PT](http://arxiv.org/abs/2607.16193v1)\n\n**Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA**\n*Authors: Ce Zhang, Ziyang Wang*\nThis paper proposes VideoTreeSearch (VTS), a framework that frames grounded long-video question answering as an iterative self-correcting search over an adaptive temporal tree, addressing the limitations of existing agentic methods that struggle to recover from early mistakes due to their lack of a fine-to-coarse backtracking primitive.\n[verified via arXiv as of 2026-07-20 05:08 PM PT](http://arxiv.org/abs/2607.16189v1)\n\n**PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization**\n*Authors: Yuchen Yang, Yifan Zhao*\nThis paper introduces PagedWeight, a novel management method for Mixture-of-Experts (MoE) LLM serving that dynamically quantizes model weights at runtime, balancing expert-weight precision with KV cache sizes to improve task accuracy, memory consumption, and throughput/latency.\n[verified via arXiv as of 2026-07-20 05:08 PM PT](http://arxiv.org/abs/2607.16184v1)\n\n**Vision-Language Assistant for Emotional Reactions to Risky Driving**\n*Authors: Harine Choi, Eun Hak Lee*\nThis study presents Keep Yelling Assistant (KYA), a vision-language pipeline that detects risky driving behaviors and generates emotionally expressive responses tailored to driver preferences, enhancing driver awareness and comfort by integrating emotional dimensions into autonomous driving systems.\n[verified via arXiv as of 2026-07-20 05:08 PM PT](http://arxiv.org/abs/2607.16181v1)\n\n**Physics-enhanced reinforcement learning for real-time optimal control of dynamical systems**\n*Authors: Matteo Tomasetto, Nicolò Botteghi*\nThis paper introduces PEARL (Physics-EnhAnced Reinforcement Learning), a novel paradigm that bridges RL and traditional optimal control for dynamical systems, aiming to address the sample inefficiency and dimensionality challenges of RL algorithms in synthesizing optimal control strategies.\n[verified via arXiv as of 2026-07-20 05:08 PM PT](http://arxiv.org/abs/2607.16177v1)\n\n**Evaluating Open-Weight LLMs for Generating Structured Threat Information for Autonomous Vehicle Vulnerabilities**\n*Authors: Md Erfan, Ahmed Ryan*\nThis research evaluates open-weight Large Language Models (LLMs) for their ability to generate structured threat information from plain text vulnerability descriptions, which is crucial for security practitioners to mitigate risks effectively in Connected and Autonomous Vehicles (CAVs).\n[verified via arXiv as of 2026-07-20 05:08 PM PT](http://arxiv.org/abs/2607.16175v1)\n\n**Vision-Language-Motion Maps: An Open-Vocabulary, Uncertainty-Aware, Queryable Motion Attribute for 3D Scene Maps**\n*Authors: Dibyendu Ghosh, Ayushi Shakya*\nThis paper introduces Vision-Language-Motion Maps (VLMM), an open-vocabulary, natural-language-queryable 3D map that incorporates a fused motion attribute for each scene element, combining VLM/LLM semantic movability priors with observed cross-frame motion and uncertainty, allowing robots to answer queries about how scene elements behave.\n[verified via arXiv as of 2026-07-20 05:08 PM PT](http://arxiv.org/abs/2607.16173v1)\n\n**When Does Muon Help Agentic Reinforcement Learning?**\n*Authors: Kai Ruan, Jinghao Lin*\nThis study investigates the effectiveness of vanilla Muon compared to AdamW in sparse-reward agentic reinforcement learning, demonstrating that applying Muon only to hidden weight matrices can significantly improve validation success in specific settings.\n[verified via arXiv as of 2026-07-20 05:08 PM PT](http://arxiv.org/abs/2607.16169v1)\n\n**Behaviour-Conditioned Neural Processes for Adaptive Residential Short-Term Load Forecasting**\n*Authors: Ramin Soleimani, Andrea Visentin*\nThis work proposes a behaviour-conditioned Attentive Neural Process framework for residential short-term load forecasting, which embeds inferred behavioural structure directly within the forecasting mechanism to address the challenges of heterogeneous and variable household demand.\n[verified via arXiv as of 2026-07-20 05:08 PM PT](http://arxiv.org/abs/2607.16168v1)\n\n**An Exam for Active Observers**\n*Authors: Jiarui Zhang, Muzi Tao*\nThis paper introduces ActiveVision, a benchmark designed to measure active observation capabilities in multimodal large language models (MLLMs) by forcing repeated visual perception tasks, revealing that frontier MLLMs struggle with tasks requiring continuous gaze redirection.\n[verified via arXiv as of 2026-07-20 05:08 PM PT](http://arxiv.org/abs/2607.16165v1)","previewShape":"markdown","feedbackRating":null},{"id":"wAcx3h0FC5ZDW0r70dl3","status":"completed","startedAt":"2026-07-20T00:08:09.115Z","completedAt":"2026-07-20T00:09:12.220Z","durationMs":null,"finalText":"**4 new arXiv papers on LLM agent evaluation this morning**\n\nHere are 4 new arXiv papers on LLM agent evaluation:\n\n*   **Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models**\n    First author: Patrik Wolf\n    This paper investigates how well LLM estimates adhere to statistical self-consistency principles, using binary trees to evaluate their conditional inference capabilities, which is crucial for LLM agent evaluation. [verified via arXiv as of 2026-07-20T00:08:13.633Z](http://arxiv.org/abs/2607.15277v1)\n\n*   **Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents**\n    First author: Paul Kassianik\n    This paper introduces a cost-aware evaluation framework for language-model security agents, moving beyond traditional success rate metrics to consider inference and tool spend in operational security scenarios. [verified via arXiv as of 2026-07-20T00:08:13.633Z](http://arxiv.org/abs/2607.15263v1)\n\n*   **SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration**\n    First author: Yuyao Zhang\n    This paper proposes SearchOS, a multi-agent framework designed to improve the robustness of information-seeking agents by explicitly managing search progress to avoid repetitive loops and enhance output quality. [verified via arXiv as of 2026-07-20T00:08:13.633Z](http://arxiv.org/abs/2607.15257v1)\n\n*   **Bridge Evidence: Static Retrieval Utility Does Not Predict Causal Utility in Multi-Step Agentic Search**\n    First author: Debayan Mukhopadhyay\n    This research highlights a crucial discrepancy in evaluating retrieval systems for multi-step agentic search, showing that static retrieval utility does not accurately predict a document's causal utility for an agent's subsequent actions. [verified via arXiv as of 2026-07-20T00:08:13.633Z](http://arxiv.org/abs/2607.15253v1)","previewShape":"markdown","feedbackRating":null},{"id":"kf8ncYDT1t8DNH8Axm9d","status":"checked_no_change","startedAt":"2026-07-19T00:08:10.330Z","completedAt":"2026-07-19T00:08:10.330Z","durationMs":null,"finalText":null,"previewShape":"markdown","feedbackRating":null},{"id":"nsZUmQQIpoaGrwIZv0jY","status":"completed","startedAt":"2026-07-18T00:08:17.676Z","completedAt":"2026-07-18T00:09:05.415Z","durationMs":null,"finalText":"**10 new arXiv papers on LLM agent evaluation this morning**\n\nHere are 10 new arXiv papers on LLM agent evaluation:\n\n1. **Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models** by Patrik Wolf, Thomas Kleine Buening, Andreas Krause, Celestine Mendler-Dünner\n   This paper investigates how well LLM estimates adhere to the self-consistency principle, using binary trees as an evaluation scaffold to recursively partition a population.\n   [http://arxiv.org/abs/2607.15277v1](http://arxiv.org/abs/2607.15277v1) [verified via https://arxiv.org · as of 2026-07-18T00:08:20.664Z]\n\n2. **SciDiagramEdit: Learning to Edit Scientific Diagrams from Paper Revisions** by Yasheng Sun, Zezi Zeng, Yifan Yang, Chong Luo, Wenyi Wang, Ziwei Liu et al.\n   This paper introduces SciDiagramEdit, a benchmark and skill-evolution framework for editing scientific diagrams based on natural language instructions, by learning from paper revisions.\n   [http://arxiv.org/abs/2607.15272v1](http://arxiv.org/abs/2607.15272v1) [verified via https://arxiv.org · as of 2026-07-18T00:08:20.664Z]\n\n3. **SceneBind: Binding What and Where Across Vision, Audio and Language** by Mingfei Chen, Zijun Cui, Ruoke Zhang, Hyeonggon Ryu, Eli Shlizerman\n   SceneBind is presented as an omni-modal representation for realistic scenes that offers joint semantic and 3D spatial understanding across vision, audio, and language, addressing the lack of explicit spatial structure in existing omni-modal encoders.\n   [http://arxiv.org/abs/2607.15265v1](http://arxiv.org/abs/2607.15265v1) [verified via https://arxiv.org · as of 2026-07-18T00:08:20.664Z]\n\n4. **Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents** by Paul Kassianik, Blaine Nelson, Yaron Singer\n   This paper proposes a cost-aware evaluation of language-model security agents on offensive and defensive challenges, comparing models at fixed cost levels and decomposing performance by inference and tool spend, going beyond traditional peak offensive capability metrics.\n   [http://arxiv.org/abs/2607.15263v1](http://arxiv.org/abs/2607.15263v1) [verified via https://arxiv.org · as of 2026-07-18T00:08:20.664Z]\n\n5. **SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration** by Yuyao Zhang, Junjie Gao, Zhengxian Wu, Jiaming Fan, Jin Zhang, Shihan Ma et al.\n   SearchOS is introduced as a system-level multi-agent framework to improve robust open-domain information-seeking agent collaboration by explicitly managing search progress as persistent and shared state, addressing issues of repetitive loops when search attempts fail.\n   [http://arxiv.org/abs/2607.15257v1](http://arxiv.org/abs/2607.15257v1) [verified via https://arxiv.org · as of 2026-07-18T00:08:20.664Z]\n\n6. **teLLMe Why (Ain't Nothing but a Jam): Exploratory Causal Analysis of Urban Driving Data** by Qiwei Li, Jorge Ortiz\n   teLLMe is a system for exploratory causal analysis of urban driving datasets, combining causal structure learning, bootstrap-based stability checks, and query-specific effect estimation to answer causal questions from observational data.\n   [http://arxiv.org/abs/2607.15254v1](http://arxiv.org/abs/2607.15254v1) [verified via https://arxiv.org · as of 2026-07-18T00:08:20.664Z]\n\n7. **Bridge Evidence: Static Retrieval Utility Does Not Predict Causal Utility in Multi-Step Agentic Search** by Debayan Mukhopadhyay, Utshab Kumar Ghosh, Shubham Chatterjee\n   This paper demonstrates that static retrieval utility does not accurately predict causal utility in multi-step agentic search, as a document's value can be in enabling future agent actions rather than directly answering the current question.\n   [http://arxiv.org/abs/2607.15253v1](http://arxiv.org/abs/2607.15253v1) [verified via https://arxiv.org · as of 2026-07-18T00:08:20.664Z]\n\n8. **New exact bispectrum shapes in multifield inflation** by Lucas Pinol\n   This paper presents the first analytical calculation of the primordial bispectrum in multifield inflation where quadratic mixing between curvature and isocurvature fluctuations is treated non-perturbatively.\n   [http://arxiv.org/abs/2607.15251v1](http://arxiv.org/abs/2607.15251v1) [verified via https://arxiv.org · as of 2026-07-18T00:08:20.664Z]\n\n9. **AutoSynthesis: An agentic system for automated meta-analysis** by Moein Taherinezhad, Sebastian Maier, Gerardo Vitagliano, Francesco Pierri, Stefan Feuerriegel\n   AutoSynthesis is an end-to-end multi-agent system for automated meta-analysis, designed to transform primary research into reliable knowledge by automating steps from search strategy formulation to heterogeneity analysis.\n   [http://arxiv.org/abs/2607.15247v1](http://arxiv.org/abs/2607.15247v1) [verified via https://arxiv.org · as of 2026-07-18T00:08:20.664Z]\n\n10. **ARMOR++: Agentic Orchestration of a Multi-Domain Primitive Set for Transferable Attacks on Deepfake Detectors** by Christos Korgialas, Gabriel Lee Jun Rong, Dion Jia Xu Ho, Pai Chet Ng, Xiaoxiao Miao, Konstantinos N. Plataniotis\n    ARMOR++ is a robust multi-agent framework that uses a Vision-Language Model (VLM) and Large Language Model (LLM) for transferable deepfake evasion attacks, addressing the degradation of deepfake detectors under black-box adversarial transfer.\n    [http://arxiv.org/abs/2607.15246v1](http://arxiv.org/abs/2607.15246v1) [verified via https://arxiv.org · as of 2026-07-18T00:08:20.664Z]","previewShape":"markdown","feedbackRating":null},{"id":"Xz2TFNHyO9t750rFhVRl","status":"completed","startedAt":"2026-07-17T00:08:12.837Z","completedAt":"2026-07-17T00:08:40.112Z","durationMs":null,"finalText":"**10 new arXiv papers on LLM agent evaluation this morning**\n\n10 new arXiv papers on LLM agent evaluation this morning:\n\n### Leveraging unlabelled data for generalizable neural population decoding\n**Authors:** Ximeng Mao, Nanda H. Krishna, Avery Hee-Woon Ryoo, Matthew G. Perich, Guillaume Lajoie\nThis paper introduces MOJO (Masked autOencoder-based JOint training), a framework that leverages both self-supervised and supervised learning to improve neural decoding performance, specifically in the context of spike-tokenizing models for brain-computer interfaces. [verified via https://arxiv.org/abs/2607.14086v1 as of 2026-07-17T00:08:16.879Z]\nLink: http://arxiv.org/abs/2607.14086v1\n\n### Building Shor's Algorithm in Lean: An Agentic Formalization of Quantum Attacks on RSA-2048 and P-256\n**Authors:** Lei Zhang, Yusheng Zhao, Hongshun Yao, Xin Wang\nThis work formalizes Shor's algorithm in Lean using an agentic formalization approach, where software agents analyze sources, write Lean code, and repair proofs with human review, contributing to the formalization of quantum computing concepts. [verified via https://arxiv.org/abs/2607.14082v1 as of 2026-07-17T00:08:16.879Z]\nLink: http://arxiv.org/abs/2607.14082v1\n\n### Linear Independent Component Analysis via Optimal Transport\n**Authors:** Ashutosh Jha, Michel Besserve, Simon Buchholz\nThis paper proposes a new method for Linear Independent Component Analysis (ICA) that uses the squared Wasserstein distance to a standard Gaussian to measure non-Gaussianity, offering an alternative to classical ICA algorithms that rely on proxy contrast functions. [verified via https://arxiv.org/abs/2607.14081v1 as of 2026-07-17T00:08:16.879Z]\nLink: http://arxiv.org/abs/2607.14081v1\n\n### VisualRepair: Dynamic Tool Calling and Region Focusing for Visual Software Issue Repair\n**Authors:** Jingyu Xiao, Zhongyi Zhang, Haoran Hou, Yuxuan Wan, Yuan Jiang, Yintong Huo et al.\nThis research addresses the challenge of automated program repair in multimodal scenarios by developing a method that effectively uses visual information from diverse bug screenshots through dynamic tool calling and region focusing. [verified via https://arxiv.org/abs/2607.14075v1 as of 2026-07-17T00:08:16.879Z]\nLink: http://arxiv.org/abs/2607.14075v1\n\n### Screening of Biosecurity Features in Metagenomic Data with Evo 2 Probes\n**Authors:** Jeremy Guntoro, Alexander Dack, Dylan Danno, Michaela Jančovičová, Križan Jurinović, Vanessa Smilansky\nThis paper investigates the utility of genomic foundation models like Evo 2 for biosecurity screening, demonstrating that linear and attention probes can effectively detect antimicrobial resistance in metagenomic data. [verified via https://arxiv.org/abs/2607.14070v1 as of 2026-07-17T00:08:16.879Z]\nLink: http://arxiv.org/abs/2607.14070v1\n\n### Hindcast: Replaying Prediction Markets to Evaluate LLM Forecasters\n**Authors:** Xiao Ye, Jacob Dineen, Evan Zhu, Shijie Lu, Kevin Song, Ben Zhou\nThis paper introduces Hindcast, a method to evaluate LLM forecasters by replaying prediction markets from a chosen past date, effectively preventing information leakage from future events into the evaluation process. [verified via https://arxiv.org/abs/2607.14051v1 as of 2026-07-17T00:08:16.879Z]\nLink: http://arxiv.org/abs/2607.14051v1\n\n### Deep Interaction: An Efficient Human-AI Interaction Method for Large Reasoning Models\n**Authors:** Hefeng Zhou, Jinxuan Zhang, Jiong Lou, Yuxin Liu, Chaochao Lu, Jingjing Qu et al.\nThis research proposes Deep Interaction, an efficient human intervention mechanism that allows for direct editing of erroneous reasoning steps in large language models, addressing limitations of current interaction approaches. [verified via https://arxiv.org/abs/2607.14049v1 as of 2026-07-17T00:08:16.879Z]\nLink: http://arxiv.org/abs/2607.14049v1\n\n### PhysClaw-0: A Symbiotic Agentic System for Robot Autonomy via Language Corrections\n**Authors:** Boyuan Wang, Zhenyuan Zhang, Zhiqin Yang, Peijun Gu, Shuya Wang, Xiaofeng Wang et al.\nThis paper presents PhysClaw-0, a symbiotic agentic system for robot autonomy that retains and reuses human language corrections across data collection rounds, reducing oversight costs by addressing recurring failures efficiently. [verified via https://arxiv.org/abs/2607.14047v1 as of 2026-07-17T00:08:16.879Z]\nLink: http://arxiv.org/abs/2607.14047v1\n\n### Earthquaker-AI: A Retrieval-Augmented Generation Framework with Rubric-Based Assessment for Primary School Earthquake Education\n**Authors:** Xanthi Kokkinou, Chaido Mizeli, Nafsika Koulaxidou, Marina Delianidi, Konstantinos Diamantaras\nThis paper introduces Earthquaker-AI, a hybrid educational framework that integrates a conversational AI assistant based on Retrieval-Augmented Generation with an educational robotics project to enhance earthquake preparedness for primary school students. [verified via https://arxiv.org/abs/2607.14046v1 as of 2026-07-17T00:08:16.879Z]\nLink: http://arxiv.org/abs/2607.14046v1\n\n### LLMs for Qualitative and Mixed-Methods Social Network Analysis\n**Authors:** Moses Boudourides\nThis manuscript explores the integration of Large Language Models (LLMs) into qualitative and mixed-methods social network analysis, focusing on enhancing the depth and rigor of the analysis rather than replacing human researchers. [verified via https://arxiv.org/abs/2607.14045v1 as of 2026-07-17T00:08:16.879Z]\nLink: http://arxiv.org/abs/2607.14045v1","previewShape":"markdown","feedbackRating":null},{"id":"YQrZyT3KJg03pBgh24KX","status":"completed","startedAt":"2026-07-16T00:08:34.169Z","completedAt":"2026-07-16T00:09:09.106Z","durationMs":null,"finalText":null,"previewShape":"markdown","feedbackRating":null},{"id":"ZFGdDRJZcju7NlmVEkIG","status":"completed","startedAt":"2026-07-15T00:08:07.266Z","completedAt":"2026-07-15T00:09:11.429Z","durationMs":null,"finalText":"**3 new arXiv papers on LLM agent evaluation this morning**\n\n3 new arXiv papers on LLM agent evaluation have been published [verified via arXiv as of 05:08 PM PT].\n\n---\n\n**Metacognition in LLMs: Foundations, Progress, and Opportunities**\n*   **Authors:** Gabrielle Kaili-May Liu, Areeb Gani, Jacqueline Lu, Jordan Thomas, Mark Steyvers, Arman Cohan\n*   **Summary:** This paper provides a comprehensive overview of metacognition in LLMs, exploring when, how, and to what extent LLMs can exhibit or be endowed with effective metacognitive abilities, and how these abilities can enhance AI systems [verified via arXiv as of 05:08 PM PT].\n*   **Link:** [http://arxiv.org/abs/2607.11881v1](http://arxiv.org/abs/2607.11881v1)\n\n---\n\n**Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias**\n*   **Authors:** Zixiang Xu, Sixian Li, Huaxing Liu, Xiang Wang, Shuai Li, Zirui Song et al.\n*   **Summary:** This research offers a representation-level account of LLM-as-judge scoring bias, complementary to input-output studies, across multiple judges, bias types, and benchmarks, identifying a low-dimensional, type-specific subspace in biased inputs [verified via arXiv as of 05:08 PM PT].\n*   **Link:** [http://arxiv.org/abs/2607.11871v1](http://arxiv.org/abs/2607.11871v1)\n\n---\n\n**AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification**\n*   **Authors:** Lingkai Kong, Zijian Wu, Yuzhe Gu, Haiteng Zhao, Wenyong Huang, Shuang Sun et al.\n*   **Summary:** This paper introduces AdvancedMathBench, a benchmark suite including ProverBench, designed to evaluate advanced mathematical reasoning capabilities of LLMs, addressing limitations of existing benchmarks in scope and evaluation granularity for proof generation and verification [verified via arXiv as of 05:08 PM PT].\n*   **Link:** [http://arxiv.org/abs/2607.11849v1](http://arxiv.org/abs/2607.11849v1)","previewShape":"markdown","feedbackRating":null},{"id":"X8jOzX2EhAX1gbFpQVAA","status":"completed","startedAt":"2026-07-14T00:08:07.954Z","completedAt":"2026-07-14T00:08:26.121Z","durationMs":null,"finalText":"**1 new arXiv paper on LLM agent evaluation this morning**\n\nHere is a new arXiv paper on LLM agent evaluation:\n\n**VEXAIoT: Autonomous IoT Vulnerability EXploitation using AI Agents**\n*Katherine Swinea, Kshitiz Aryal, Lopamudra Praharaj et al.*\n\nThis paper introduces VEXAIoT, an autonomous multi-agent framework that utilizes LLM-based reasoning and offensive security tools for vulnerability discovery and exploitation in IoT environments, aiming to provide scalable and adaptive security testing. [verified via arXiv as of 05:08 PM PT]\n\nRead the paper: [http://arxiv.org/abs/2607.09653v1](http://arxiv.org/abs/2607.09653v1)","previewShape":"markdown","feedbackRating":null},{"id":"WFsURUqmcehL5Gc8BJQE","status":"completed","startedAt":"2026-07-13T00:08:18.940Z","completedAt":"2026-07-13T00:08:43.360Z","durationMs":null,"finalText":null,"previewShape":"markdown","feedbackRating":null}],"stats":{"totalRuns":14,"satisfactionPercent":null,"lastRunAt":"2026-07-26T00:09:27.053Z"}}