Conformal Agent Error Attribution
arXiv:2605.06788v1 Announce Type: new Abstract: When multi-agent systems (MAS) fail, identifying where the decisive error occurred is the first step for automated recovery to an earlier state. Error a
Knowledge catalogue
arXiv:2605.06788v1 Announce Type: new Abstract: When multi-agent systems (MAS) fail, identifying where the decisive error occurred is the first step for automated recovery to an earlier state. Error a
arXiv:2605.07214v1 Announce Type: new Abstract: Large Language Models have recently emerged as a promising paradigm for automated heuristic design for NP-hard combinatorial optimization problems. Desp
Cofounder.co is a platform by Intelligence Co that provides AI agent management, allowing users to delegate agent oversight and coordination to an automated system rather than managing multiple agents
arXiv:2605.07433v1 Announce Type: cross Abstract: Qualitative models provide crucial instruments for modelling complex biological systems. While advances in automated reasoning and symbolic encodings
arXiv:2510.22944v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have become indispensable for automated code generation, yet the quality and security of their outputs remain a c
arXiv:2605.07835v1 Announce Type: new Abstract: Multi-robot systems in automated warehouses must manage continuous streams of pickup-and-delivery tasks while ensuring efficiency and safety. Prior work
arXiv:2602.00513v3 Announce Type: replace Abstract: Cyber threat intelligence (CTI) analysts routinely convert noisy, unstructured security artifacts into standardized, automation-ready representation
arXiv:2605.07635v1 Announce Type: new Abstract: Automated assistants for Grammatical Error Correction are now embedded in educational platforms serving millions of learners, yet three critical gaps re
arXiv:2605.06940v1 Announce Type: new Abstract: Annotation automation via Large Language Models (LLMs) is the core approach for scaling NLP datasets; however, LLM behavior with respect to closed-set i
arXiv:2605.06173v2 Announce Type: replace-cross Abstract: Diabetic Retinopathy (DR) is a leading cause of preventable blindness among working-age adults worldwide, yet most automated screening systems
arXiv:2605.07507v1 Announce Type: new Abstract: The exponential growth of academic publications has created an urgent need for automated tools capable of extracting structured knowledge from unstructu
arXiv:2605.06913v1 Announce Type: cross Abstract: We present You Only Stack Once (YOSO), an automated pipeline designed to detect faint, slow-moving Solar System objects in wide-field astronomical sur
Sam Altman posted about starting multiple Codex tasks, spending time outdoors with his child, and returning after naptime to find all the tasks had been completed, suggesting automated or background t
This article discusses how AI technologies can help Human Resources departments address growing capacity challenges and workload demands. It likely explores how AI tools can automate routine HR tasks,
This OpenAI article outlines safety practices and guidelines for responsibly deploying and using Codex, their code generation model. It likely covers potential risks associated with automated code gen
With AI coding tools like Antigravity and Claude Code, I can build a working web app in record time. But deploying it? That's where I'd historically lose the rest of the afternoon to Dockerfiles, IAM
Financial services are becoming more personal, more automated and increasingly AI-driven. However, a big hurdle remains: AI trust. In an industry that handles sensitive financial data, there is little
arXiv:2605.04888v1 Announce Type: new Abstract: The exponential growth of social media has created an urgent need for automated systems to analyze unstructured public sentiment in real time. This stud
arXiv:2605.03537v1 Announce Type: cross Abstract: This paper presents a modular AI agentic skill pipeline for automating subject indexing with Library of Congress Subject Headings (LCSH). Subject inde
arXiv:2605.03117v1 Announce Type: cross Abstract: Repository-level fault localization (FL) and automated program repair (APR) require an agent to identify the relevant code units across files, follow
arXiv:2605.05090v1 Announce Type: new Abstract: We present an automated, contrastive evaluation pipeline for auditing the behavioral impact of interventions on large language models. Given a base mode
arXiv:2605.05121v1 Announce Type: new Abstract: Automated mental health prediction using textual data has shown promising results with deep learning and large language models. However, deploying these
arXiv:2605.05031v1 Announce Type: new Abstract: Recent deep learning approaches seek to automate CAD creation by representing a model as a sequence of discrete commands and parameters, and then genera
arXiv:2605.05092v1 Announce Type: cross Abstract: Safe L2/L3 driving automation requires anticipating human-in-the-loop reactions during shared-control transitions. While most driving world models for
arXiv:2605.04882v1 Announce Type: new Abstract: Automated glaucoma detection is critical for preventing irreversible vision loss and reducing the burden on healthcare systems. However, ensuring fairne
Hacker News → LLM Artifact I built the most personalized HN feed. It only tracks topics I do research around based on memory and LLM wiki. No point in storing bookmarks. With a few automations, rules,
arXiv:2602.12095v3 Announce Type: replace Abstract: The automation of warehouse operations is crucial for improving productivity and reducing human exposure to hazardous environments. One operation fr
arXiv:2605.04564v1 Announce Type: new Abstract: The representativeness of synthetic pre-crash scenarios is crucial for assessing the safety impact of Driving Automation Systems through virtual simulat
Thoth v3.21.0 is a local-first AI assistant with integrated tools including a personal knowledge graph, voice, vision, shell, and browser automation capabilities. This release introduces Buddy Compani
arXiv:2605.04298v1 Announce Type: new Abstract: Automated essay scoring (AES) research often relies on rank-based correlation metrics to validate analytic assessment. However, such metrics obscure bot
arXiv:2605.02741v1 Announce Type: cross Abstract: The promise of Large Language Models in automated software engineering is often measured by functional correctness, overlooking the critical issue of
arXiv:2605.01727v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for automated news credibility assessment, yet it remains unclear whether they apply even-handed stan
arXiv:2605.02690v1 Announce Type: cross Abstract: Hyperledger Fabric performance depends on many interacting configuration parameters, making manual tuning difficult. We study automated throughput tun
arXiv:2605.02921v1 Announce Type: cross Abstract: As LLMs continue to shape real-world applications, automated jailbreak generation becomes essential to reveal safety weaknesses and guide model improv
arXiv:2605.03686v1 Announce Type: cross Abstract: Automated Machine Learning (AutoML) frameworks increasingly leverage Large Language Models (LLMs) for tasks such as hyperparameter optimization and ne
arXiv:2605.01423v1 Announce Type: cross Abstract: The escalating data scale in High-Energy Physics (HEP) fuels a growing aspiration for higher analytical efficiency. While Large Language Models (LLMs)
arXiv:2605.02965v1 Announce Type: new Abstract: Artificial intelligence-generated content (AIGC) has emerged as a transformative paradigm for automating the creation of diverse and customized content,
arXiv:2605.02092v1 Announce Type: new Abstract: The automation of scientific research workflows has emerged as a transformative frontier in artificial intelligence, yet existing autonomous research ag
arXiv:2605.02917v1 Announce Type: new Abstract: Supervised deep learning models for automated CTG analysis are typically constrained by narrowly curated labelled datasets and limited patient cohorts,
arXiv:2605.03784v1 Announce Type: new Abstract: Rising global food demand and growing climate pressure increase the need for sustainable, precise agricultural practices. Automated, individualized plan
Singular Bank leverages OpenAI's ChatGPT and Codex to accelerate banking operations and developer productivity. The integration enables bankers and software engineers to streamline workflows, automate
arXiv:2605.03842v1 Announce Type: cross Abstract: Robotic Mobile Fulfillment Systems (RMFS) rely on mobile robots for automated inventory transportation, coordinating order allocation and robot schedu
Robobun is a GitHub bot developed as part of the Bun project that automates contributions to the repository. Developers Jared Sumner and Bun Cherny discussed the bot's functionality and capabilities d
arXiv:2605.01336v1 Announce Type: new Abstract: News outlets shape public opinion at a scale that makes automated detection of political bias and factuality essential. However, the field still lacks u
arXiv:2602.00074v2 Announce Type: replace-cross Abstract: While large language models (LLMs) can support clinical documentation needs, standalone tools struggle with 'workflow friction' from manual da
arXiv:2605.01355v1 Announce Type: new Abstract: Automated leaf disease classification is critical for early disease detection in resource-constrained field environments. Vision Transformers (ViTs) pro
Cursor has introduced an automated CI failure resolution feature that uses always-on agents to monitor GitHub repositories, analyze build failures, and automatically generate pull requests with fixes.
arXiv:2605.02245v1 Announce Type: new Abstract: Automated sleep stage classification typically employs a single population-agnostic model, disregarding established demographic variations in sleep arch
arXiv:2605.02011v1 Announce Type: new Abstract: Automating the drafting of judgment documents is pivotal to judicial efficiency, yet it remains challenging due to the dual requirements of comprehensiv
arXiv:2605.02376v1 Announce Type: new Abstract: Automated medical report generation, MRG, holds substantial value for alleviating radiologist workload and enhancing diagnostic efficiency. However, mai
arXiv:2605.02288v1 Announce Type: new Abstract: Automated laboratories hold the promise of accelerating scientific discovery, yet their deployment is bottlenecked by the difficulty of designing safe a
arXiv:2605.01731v1 Announce Type: new Abstract: Platooning of connected and automated vehicles provides significant benefits in terms of energy efficiency, traffic throughput, and, most critically, sa
arXiv:2605.01779v1 Announce Type: new Abstract: Vision-language models (VLMs) have shown potential for automated radiology report generation, yet existing approaches rely on global embedding compressi
arXiv:2512.05525v2 Announce Type: replace-cross Abstract: Businesses increasingly rely on large language models (LLMs) to automate simple repetitive tasks instead of developing custom machine learning
arXiv:2605.01450v1 Announce Type: new Abstract: Recent frameworks like ToFu and TEMPEH provide an automated alternative to classical registration pipelines by predicting 3D meshes in dense semantic co
arXiv:2602.12361v2 Announce Type: replace Abstract: Human-machine interfaces in industrial automation need sensing modules that monitor operator actions and physiological state. This is important in f
arXiv:2605.00238v1 Announce Type: new Abstract: Automated short answer grading (ASAG) with large language models (LLMs) is commonly evaluated with aggregate metrics such as macro-F1 and Cohen's kappa.
Deepsec is a security tool introduced by Vercel designed to identify and help remediate vulnerabilities within codebases. The tool functions as a security harness that automates vulnerability detectio
arXiv:2605.00647v1 Announce Type: new Abstract: Automated pediatric electrocardiogram (ECG) diagnosis remains challenging because models trained predominantly on adult data suffer from substantial cro
arXiv:2605.00595v1 Announce Type: new Abstract: Perception for automated driving is largely based on onboard environmental sensors, such as cameras and radar, which are cost-effective but limited by l