CompanyAnthropic8 recent entries11 Jun 2026NightFeats @ MMU-RAGent NeurIPS 2025: A Context-Optimized Multi-Agent RAG System for the Text-to-Text TrackarXiv:2606.11199v1 Announce Type: cross Abstract: We present NightFeats, a structured multi-agent retrieval-augmented generation (RAG) system submitted to the MMU-RAGent competition at NeurIPS 2025, w→30 Jun 2026meta-pipe: An LLM-agent pipeline for end-to-end automated systematic review and meta-analysisarXiv:2606.28363v1 Announce Type: cross Abstract: Objective: To describe the architecture and design rationale of meta-pipe, an open-source large language model (LLM)-agent pipeline that integrates th
CompanyOpenAI7 recent entries21 Apr 2026RefineStat: Efficient Exploration for Probabilistic Program SynthesisarXiv:2509.01082v3 Announce Type: replace Abstract: Probabilistic programming offers a powerful framework for modeling uncertainty, yet statistical model discovery in this domain entails navigating an→19 May 2026A-ProS: Towards Reliable Autonomous Programming Through Multi-Model FeedbackarXiv:2605.18073v1 Announce Type: cross Abstract: Large Language Models (LLMs) demonstrate strong potential for automated code generation, yet their ability to iteratively refine solutions using execu→26 May 2026MDIA: A Multi-Agent Diagnostic Intelligence Pipeline on HealthBench ProfessionalarXiv:2605.24699v1 Announce Type: new Abstract: Most reported gains on agentic-LLM clinical benchmarks are often attributed to prompt engineering, yet our results suggest that larger improvements can →2 Jun 2026Dynamic Coordination Strategy Selection for Enterprise Multi-Agent SystemsarXiv:2606.00804v1 Announce Type: cross Abstract: Enterprise multi-agent systems increasingly expose multiple coordination patterns, but deployments often lack evidence for when to use consensus, deba→2 Jun 2026Characterizing Web Search in The Age of Generative AIarXiv:2510.11560v2 Announce Type: replace-cross Abstract: The advent of LLMs has given rise to generative search, a new search paradigm in which LLMs retrieve information from the web related to a que→10 Jun 2026ABC-Bench: An Agentic Bio-Capabilities Benchmark for BiosecurityarXiv:2606.11150v1 Announce Type: new Abstract: Large language models (LLMs) are rapidly acquiring capabilities relevant to biological research, from literature synthesis to interpretation of experime→30 Jul 2026The Reliability of LLMs for Medical Diagnosis: An Examination of Consistency, Manipulation, and Contextual AwarenessarXiv:2503.10647v2 Announce Type: replace Abstract: This study evaluated the diagnostic reliability of two Large Language Models (LLMs), Google Gemini 2.0 Flash and OpenAI ChatGPT-4o, across three dim
CompanyGoogle8 recent entries3 Aug 2026Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model ReviewarXiv:2607.28631v1 Announce Type: new Abstract: AI Scientist systems capable of autonomous research have the potential to significantly accelerate scientific discovery. However, evaluating and compari→3 Aug 2026Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon StatementsarXiv:2607.28661v1 Announce Type: new Abstract: Do Large Language Models (LLMs) possess genuine structural reasoning, or merely rely on surface-level pattern matching? The financial domain, demanding →4 Aug 2026Domain-Specific Evaluation of Text-to-Speech Systems: A Multi-Metric Benchmarking StudyarXiv:2608.02235v1 Announce Type: new Abstract: Recent advances in neural text-to-speech (TTS) systems have substantially improved speech naturalness and intelligibility across many languages. However→6 Aug 2026NuclearDiffusion: Text-to-Image Foundation Models for Learning Nuclear Energy ConceptsarXiv:2608.04030v1 Announce Type: cross Abstract: Generative artificial intelligence (AI) has transformed text-to-image synthesis, yet its ability to represent specialized engineering domains remains →7 Aug 2026Vibe Compiler: A Research-Logic Synthesis Tool That Runs without Prompt Engineering -Toward Enhancing Metacognition for Sustaining Agency in the Age of Generative AI-arXiv:2608.05545v1 Announce Type: cross Abstract: Generative AI used as a capable servant has greatly accelerated intellectual work, but it also risks eroding human epistemic agency by encouraging unc→10 Aug 2026Multi-Agent Forensic Reasoning for Generalizable Deepfake Video DetectionarXiv:2608.06865v1 Announce Type: cross Abstract: The malicious use of generative artificial intelligence to create highly realistic deepfake videos raises serious ethical concerns and poses substanti→10 Aug 2026Artificial Intelligence Can Match Domain Experts in Evidence Extraction and Critical Appraisal of Microbial Oncogenesis Research PublicationsarXiv:2608.07250v1 Announce Type: cross Abstract: Confirmed oncogenic microbes contribute significantly to cancer burden. Identifying novel microbial oncogenicity could yield strategies that will redu→12 Aug 2026Situation Graph Prediction for User Perspective ModelingarXiv:2602.13319v2 Announce Type: replace Abstract: Perspective-aware AI requires modeling evolving internal states---goals, emotions, contexts---not merely preferences. Progress is limited by a data
CompanyMeta8 recent entries27 Jul 2026Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement LearningarXiv:2607.21971v1 Announce Type: cross Abstract: Test-time scaling through iterative self-evolution with environment feedback, as demonstrated by AlphaEvolve, shows remarkable performance gains. We h→28 Jul 2026Pointer-Augmented Autoregressive Generation of Patent Claims with Joint Topology and Content DecodingarXiv:2607.24040v1 Announce Type: new Abstract: Autoregressive decoders emit flat token sequences and cannot enforce hierarchical constraints across output segments, a limitation that becomes acute in→28 Jul 2026Formalizing Flag Algebras in LeanarXiv:2607.23500v1 Announce Type: cross Abstract: Razborov's flag algebra method is a powerful tool for proving asymptotic inequalities in extremal graph theory, often reducing the task to finding a f→28 Jul 2026Do Language Models Converge to Themselves? Recursive Self-Refinement as Textual RelaxationarXiv:2607.22653v1 Announce Type: new Abstract: Large language models are increasingly used in recursive refinement workflows, where an initial draft is repeatedly revised by the same model. Despite t→28 Jul 2026Delegation Intelligence in Deep Search: A Controllable Framework for Disentangled Capability DiagnosisarXiv:2607.23524v1 Announce Type: new Abstract: Deep search is becoming a core capability of modern agent systems, yet it is typically evaluated solely based on end-to-end answer accuracy. This couple→28 Jul 2026AutoCluster, AutoTopicModeling, AutoTrendAnalysis: A Complete AutoML Pipeline for Predicting Emerging TrendsarXiv:2607.22641v1 Announce Type: cross Abstract: Predicting emerging trends is vital for businesses, researchers, and policymakers; yet traditional approaches often lack scalability and adaptability.→3 Aug 2026Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model ReviewarXiv:2607.28631v1 Announce Type: new Abstract: AI Scientist systems capable of autonomous research have the potential to significantly accelerate scientific discovery. However, evaluating and compari→12 Aug 2026Physics-informed Diffusion Generative Model for Time-Series Data Synthesis in Dynamic SystemsarXiv:2608.10941v1 Announce Type: new Abstract: Industrial time-series signals, such as turbine temperature and rotational speed in aero-engines, are essential for monitoring the health and operationa
CompanyMistral7 recent entries14 Apr 2026MathAgent: Adversarial Evolution of Constraint Graphs for Mathematical Reasoning Data SynthesisarXiv:2604.11188v1 Announce Type: cross Abstract: Synthesizing high-quality mathematical reasoning data without human priors remains a significant challenge. Current approaches typically rely on seed →7 May 2026Delta-Based Neural Architecture Search: LLM Fine-Tuning via Code DiffsarXiv:2605.04903v1 Announce Type: cross Abstract: Large language models (LLMs) show strong potential for neural architecture generation, yet existing approaches produce complete model implementations →26 May 2026BODHI: Precise OS Kernel Specification InferencearXiv:2605.23931v1 Announce Type: new Abstract: The formal verification of operating system kernels requires precise specifications that capture the intended behavior of system calls. Writing these sp→29 May 2026Less is Enough: Synthesizing Diverse Data in LLM Feature Space with Sparse AutoencodersarXiv:2602.10388v3 Announce Type: replace-cross Abstract: The diversity of post-training data is critical for effective downstream performance in large language models (LLMs). Many existing approaches→11 Jun 2026Multi-Agent Reasoning with Adaptive Worker Allocation for Stance DetectionarXiv:2606.11609v1 Announce Type: new Abstract: Stance detection requires identifying an author's position toward a target, often from short-form texts where stance is implicit, indirect, or rhetorica→24 Jun 2026Not All Invariants Are Equal: Curating Training Data to Accelerate Program Verification with SLMsarXiv:2603.15510v2 Announce Type: replace Abstract: The synthesis of inductive loop invariants remains a critical bottleneck in automated program verification. While Large Language Models (LLMs) show →2 Jul 2026GRACE-RAG: Governed Retrieval Architecture for Canonical Evidence Synthesis, Enabling Lightweight Deployment in Closed-Domain Institutional SettingsarXiv:2607.00013v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) systems are widely used in institutional question answering settings where responses must be grounded in authorit
CompanyxAI6 recent entries17 Apr 2026Interpretable and Explainable Surrogate Modeling for Simulations: A State-of-the-Art Survey and Perspectives on Explainable AI for Decision-MakingarXiv:2604.14240v1 Announce Type: cross Abstract: The simulation of complex systems increasingly relies on sophisticated but fundamentally opaque computational black-box simulators. Surrogate models p→24 Apr 2026C-SHAP for time series: An approach to high-level temporal explanationsarXiv:2504.11159v2 Announce Type: replace Abstract: In high-stakes domains, such as healthcare and industry, the explainability of AI-based decision-making has become crucial. Without insight into mod→28 Apr 2026Seeing Is No Longer Believing: Frontier Image Generation Models, Synthetic Visual Evidence, and Real-World RiskarXiv:2604.24197v1 Announce Type: cross Abstract: Frontier image generation has moved from artistic synthesis toward synthetic visual evidence. Systems such as GPT Image 2, Nano Banana Pro, Nano Banan→22 May 2026Evaluating Commercial AI Chatbots as News IntermediariesarXiv:2605.22785v1 Announce Type: new Abstract: AI chatbots are rapidly shaping how people encounter the news, yet no prior study has systematically measured how accurately these systems, with their p→16 Jul 2026STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair OraclearXiv:2607.13618v1 Announce Type: new Abstract: LLM agents are increasingly evaluated on multi-week decision tasks in which the state that drives cost is never directly observed. On such tasks the fin→5 Aug 2026Explainable AI for the EU Right to Explanation: A Systematic Review of the Law-XAI Translation GaparXiv:2608.02699v1 Announce Type: new Abstract: When algorithms make or influence consequential decisions---about loan eligibility, hiring, or healthcare---EU law grants affected individuals a Right t
CompanyDeepSeek8 recent entries8 Jul 2026Foundation Models for Automatic CAD GenerationarXiv:2607.05573v1 Announce Type: new Abstract: Recent advances in Large Language Models (LLMs) and Vision-Language Models (VLMs) enable the automatic generation of parametric 3D designs from natural-→8 Jul 2026FirstResearch: Auditable Question Formation for LLM Scientific Discovery AgentsarXiv:2607.05682v1 Announce Type: new Abstract: LLM systems for scientific discovery increasingly assist with ideation, literature synthesis, experiment planning, and report generation, but the first →9 Jul 2026Cost-Effective Agent Harnesses for Abstract Reasoning and Generalization on ARC-AGI-1arXiv:2607.06764v1 Announce Type: new Abstract: Recent progress on ARC-AGI-1 from disclosed architectures has come broadly from two regimes: heavy test-time compute over frontier models (evolutionary →10 Jul 2026Infinity-Parser2 Technical ReportarXiv:2607.07836v1 Announce Type: new Abstract: We present Infinity-Parser2, a large multimodal model that couples a controllable data-synthesis pipeline with multi-task reinforcement learning for end→16 Jul 2026STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair OraclearXiv:2607.13618v1 Announce Type: new Abstract: LLM agents are increasingly evaluated on multi-week decision tasks in which the state that drives cost is never directly observed. On such tasks the fin→31 Jul 2026AfriEconQA: A Benchmark for Quantitative and Temporal Reasoning over World Bank Economic ReportsarXiv:2601.15297v3 Announce Type: replace Abstract: Reliable question answering over long institutional documents requires more than topical retrieval: a system must localize the exact passage that su→7 Aug 2026Recursive Synthesis for Long-Horizon Terminal TasksarXiv:2608.05466v1 Announce Type: new Abstract: High-quality long-horizon training data for terminal agents is expensive to produce, often costing hundreds to thousands of dollars per task, because ea→10 Aug 2026Policy-Masked Private Experts: Auditable and Reversible Capability Access Control in Sparse MoE ModelsarXiv:2608.06690v1 Announce Type: cross Abstract: Most language-model access controls regulate behavior while leaving the same computation available to every request. We study a different systems ques
CompanyNVIDIA8 recent entries7 Jul 2026ELiTeFormer: An Efficient Transformer for FPGAsarXiv:2607.03652v1 Announce Type: cross Abstract: Transformer blocks are prevalent in large language model (LLM) but present deployment challenges due to their challenging computational and memory dem→8 Jul 2026KernelEvolve: Scaling Agentic Kernel Coding for Heterogeneous AI Accelerators at MetaarXiv:2512.23236v4 Announce Type: replace-cross Abstract: Making deep learning recommendation model (DLRM) training and inference fast and efficient is important. However, this presents three key syst→15 Jul 2026WanToFight: Real-Time Generative Game Engine for Multi-Player Combat InteractionarXiv:2607.12592v1 Announce Type: new Abstract: We present WanToFight, a generative game engine that simulates real-time, two-player The King of Fighters '97 (KOF~'97) gameplay from keyboard input. Pr→24 Jul 2026Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUsarXiv:2607.21042v1 Announce Type: new Abstract: Autoregressive text-to-speech models achieve strong naturalness but suffer from slow inference due to sequential token generation, limiting their deploy→3 Aug 2026TRACE: High-Fidelity 3D Scene Editing via Tangible Reconstruction and Geometry-Aligned Contextual Video MaskingarXiv:2604.01207v2 Announce Type: replace Abstract: Existing 3D Gaussian Splatting (3DGS) editing methods primarily focus on appearance modification and often struggle to support flexible geometry edi→3 Aug 2026Inference-time Trajectory Optimization for Structure-Preserving Manga Image EditingarXiv:2603.27790v2 Announce Type: replace Abstract: We present a lightweight, training-free trajectory correction method that adapts a pretrained image editing model to each input manga image using on→4 Aug 2026DiffusionGemma Technical ReportarXiv:2608.00146v1 Announce Type: new Abstract: We introduce DiffusionGemma, an experimental open-weight language model that uses discrete diffusion to generate text at exceptionally high speed. Rathe→11 Aug 2026EsaacSim: A Multimodal Event Camera Add-on for NVIDIA Isaac SimarXiv:2608.08522v1 Announce Type: new Abstract: Event-based vision is becoming an increasingly important sensing paradigm for robotics, yet its adoption remains limited by sensor availability and the