ToolClaude Code8 recent entries27 May 2026ChartAct: A Benchmark for Dynamic Chart UnderstandingarXiv:2605.26994v1 Announce Type: new Abstract: Charts are widely used to present complex data for analysis and decision making. Existing chart understanding benchmarks mainly focus on static charts, →9 Jun 2026VideoWeaver: Evaluating and Evolving Skills for Agentic Long Video GenerationarXiv:2606.08091v1 Announce Type: new Abstract: Recent agent frameworks such as Claude Code, Codex, and OpenClaw are strong at tool use and orchestration, but whether they can handle long video genera
ToolCursor2 recent entries15 Apr 2026See, Point, Refine: Multi-Turn Approach to GUI Grounding with Visual FeedbackarXiv:2604.13019v1 Announce Type: new Abstract: Computer Use Agents (CUAs) fundamentally rely on graphical user interface (GUI) grounding to translate language instructions into executable screen acti→23 Jul 2026Pathologist Attention-Aligned Report Generation for Prostate HistopathologyarXiv:2607.19624v1 Announce Type: new Abstract: The allocation of visual attention by pathologists during cancer diagnosis is a highly selective process that critically shapes the information extracte
ToolLangChain2 recent entries5 Jun 2026Entropy-Based Evaluation of AI Agents: A Lightweight Framework for Measuring Behavioral PatternsarXiv:2606.05872v1 Announce Type: cross Abstract: AI agents are commonly evaluated using task success, reward, latency, and cost. These metrics are useful, but they often miss important aspects of age→28 Jul 2026Energy Constrained Hierarchical Underwater Monitoring via Local Multi-Agent RAGarXiv:2607.24313v1 Announce Type: cross Abstract: Marine life monitoring is limited by strict energy constraints, poor underwater connectivity, and the high cost of transmitting raw multimodal data fr
ToolOllama7 recent entries16 Apr 2026Multi-Agent Object Detection Framework Based on Raspberry Pi YOLO Detector and Slack-Ollama Natural Language InterfacearXiv:2604.13345v1 Announce Type: new Abstract: The paper presents design and prototype implementation of an edge based object detection system within the new paradigm of AI agents orchestration. It g→20 Apr 2026Two-Stage Framework for Efficient UAV-Based Wildfire Video Analysis with Adaptive Compression and Fire Source DetectionarXiv:2508.16739v2 Announce Type: replace Abstract: Unmanned Aerial Vehicles (UAVs) have become increasingly important in disaster emergency response by facilitating aerial video analysis. Due to the →20 Apr 2026Fed3D: Federated 3D Object DetectionarXiv:2604.15795v1 Announce Type: new Abstract: 3D object detection models trained in one server plays an important role in autonomous driving, robotics manipulation, and augmented reality scenarios. →21 Apr 2026LOD-Net: Locality-Aware 3D Object Detection Using Multi-Scale Transformer NetworkarXiv:2604.16696v1 Announce Type: new Abstract: 3D object detection in point cloud data remains a challenging task due to the sparsity and lack of global structure inherent in the input. In this work,→21 Apr 2026Fringe Projection Based Vision Pipeline for Autonomous Hard Drive DisassemblyarXiv:2604.17231v1 Announce Type: new Abstract: Unrecovered e-waste represents a significant economic loss. Hard disk drives (HDDs) comprise a valuable e-waste stream necessitating robotic disassembly→22 Apr 2026HMR-Net: Hierarchical Modular Routing for Cross-Domain Object Detection in Aerial ImagesarXiv:2604.18866v1 Announce Type: new Abstract: Despite advances in object detection, aerial imagery remains a challenging domain, as models often fail to generalize across variations in spatial resol→27 Apr 2026Depth-Aware Rover: A Study of Edge AI and Monocular Vision for Real-World ImplementationarXiv:2604.22331v1 Announce Type: new Abstract: This study analyses simulated and real-world implementations of depth-aware rover navigation, highlighting the transition from stereo vision to monocula
ToolHugging Face8 recent entries23 Jun 2026Open Annotations and Synthetic Data for Field Localisation in Indian Bank ChequesarXiv:2606.20682v1 Announce Type: new Abstract: Automated cheque processing requires localising key fields (date, legal amount, IFSC code, account number, signature, and payee name) before any recogni→23 Jun 2026ConnectomeBench2: A Unified Benchmark for Automated Connectomic ProofreadingarXiv:2606.21116v1 Announce Type: new Abstract: Proofreading--correcting segmentation errors in 3D brain reconstructions--is the rate-limiting step in synapse-resolution connectomics. We release Conne→26 Jun 2026TMP: Tree-structured Mixed-policy Pruning for Large-scale Image Generation and EditingarXiv:2606.27089v1 Announce Type: new Abstract: Modern image generation model rapidly grows their sizes to meet high-fidelity image synthesis. However, they gradually become unaffordable for their eno→7 Jul 2026SteelBench: Evaluating Vision-Language Models in Real-World Industrial EnvironmentsarXiv:2607.05264v1 Announce Type: new Abstract: Existing video benchmarks evaluate action recognition on consumer videos, egocentric recordings, or simulated industrial environments. They do not test →16 Jul 2026Quantum Circuit Vision: Cost-Aware Evaluation of Visual AI Agents for Quantum Code GenerationarXiv:2607.10057v1 Announce Type: cross Abstract: Can AI agents visually comprehend quantum circuit diagrams and generate verified executable code--and at what cost? We present Quantum Circuit Vision,→24 Jul 2026Future Rendering neq Future Surface: A Benchmark and Dataset for Dynamic Surface Reconstruction Beyond the Observed WindowarXiv:2607.21471v1 Announce Type: new Abstract: Dynamic-scene reconstruction is almost always evaluated inside the observed time window, yet deployment settings such as AR overlays, robot interaction,→4 Aug 2026OSSDD - a New Open Dataset for Sentinel-1 Ship DetectionarXiv:2608.01963v1 Announce Type: new Abstract: Ship detection in Synthetic Aperture Radar (SAR) images plays an important role for maritime situational awareness, especially with respect to different→6 Aug 2026LiveXiv -- A Multi-Modal Live Benchmark Based on Arxiv Papers ContentarXiv:2410.10783v4 Announce Type: replace Abstract: The large-scale training of multi-modal models on data scraped from the web has shown outstanding utility in infusing these models with the required