CompanyAnthropic8 recent entries5 Aug 2026Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent EvaluationarXiv:2608.03166v1 Announce Type: new Abstract: Role-Playing Language Agents (RPLAs) are increasingly deployed in high-stakes applications such as healthcare assistance, customer support, and educatio→6 Aug 2026EDATracer: An Agentic Framework for Large-Scale EDA Artifact AnalysisarXiv:2608.04032v1 Announce Type: cross Abstract: Modern chip design relies on electronic design automation (EDA) tools that generate large, heterogeneous artifacts, including source files, scripts, l
CompanyOpenAI8 recent entries15 Jul 2026Evaluating Nonuniform Dependability Across Response Conditions: A Conditional Generalizability Framework Illustrated in Automated Essay ScoringarXiv:2607.11981v1 Announce Type: cross Abstract: Aggregate reliability estimates can obscure heterogeneity in measurement-design burden across response conditions, so a single G- or D-study may misch→24 Jul 2026Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent MediationarXiv:2607.21518v1 Announce Type: new Abstract: Even a current high-capability LLM can appear safer when shown a dangerous objective directly than when other agents transform and relay its direction. →24 Jul 2026InferenceBench: A Benchmark for Open-Ended LLM Inference Optimization by AI AgentsarXiv:2607.20468v1 Announce Type: new Abstract: AI agents are increasingly used to automate research and development tasks, yet existing benchmarks typically evaluate them on prescribed workflows or n→3 Aug 2026A robust association between LLM use and scientific productivity: Assessing stopping-time selectionarXiv:2607.28968v1 Announce Type: cross Abstract: Renault, Bergeaud, and Bosquet (hereafter RBB) argue that dating LLM adoption as the first month in which an author's abstract is flagged induces a st→7 Aug 2026AI Playing Business Games: Benchmarking Large Language Models on Managerial Decision-Making in Dynamic SimulationsarXiv:2509.26331v2 Announce Type: replace Abstract: The rapid advancement of LLMs sparked significant interest in their potential to augment or automate managerial functions. One of the most recent tr→10 Aug 2026LMM Modality Transfer: A Pre-requisite for Autonomous GIS AgentsarXiv:2608.06948v1 Announce Type: new Abstract: AI models are becoming increasingly adept at understanding and processing spatial information, thereby facilitating agentic problem-solving in spatial t→11 Aug 2026When LLM Agents Negotiate: Private Information and Dynamic Bargaining in Supply ChainsarXiv:2608.07538v1 Announce Type: new Abstract: As LLM agents move from decision support to autonomous procurement, firms need to know whether delegated negotiators create value, divide it predictably→11 Aug 2026Curriculum Generation under Structured Parametric Environments for Robust Navigation PoliciesarXiv:2608.08545v1 Announce Type: cross Abstract: Robust navigation policies for autonomous agents must generalize across continuously varying environmental conditions such as turn rates, obstacles, f
CompanyGoogle8 recent entries5 Aug 2026MDArena: Evaluating Coding Agents on Realistic Molecular Dynamics WorkflowsarXiv:2608.02642v1 Announce Type: cross Abstract: Accelerating scientific discovery is among the most consequential applications of AI, and computational biomolecular simulation stands out as a partic→7 Aug 2026AI Playing Business Games: Benchmarking Large Language Models on Managerial Decision-Making in Dynamic SimulationsarXiv:2509.26331v2 Announce Type: replace Abstract: The rapid advancement of LLMs sparked significant interest in their potential to augment or automate managerial functions. One of the most recent tr→10 Aug 2026Artificial Intelligence Can Match Domain Experts in Evidence Extraction and Critical Appraisal of Microbial Oncogenesis Research PublicationsarXiv:2608.07250v1 Announce Type: cross Abstract: Confirmed oncogenic microbes contribute significantly to cancer burden. Identifying novel microbial oncogenicity could yield strategies that will redu→11 Aug 2026Towards Expert-level Medical AI for Real-time Video ConsultationsarXiv:2608.09861v1 Announce Type: new Abstract: Audio-visual interaction is the standard for patient-physician consultations, enabling natural communication and effective assessment of illness through→11 Aug 2026LLM-Guided Heuristic Design from Simulation Traces: A Case Study in Dynamic Production and AGV SchedulingarXiv:2608.09343v1 Announce Type: new Abstract: Simulation-based optimization (SBO) evaluates executable policies under stochastic dynamics, but most methods treat the simulator as a black box: aggreg→11 Aug 2026Coupled Graph--Policy Distillation for Personalized Medication Safety in Older Adults with MultimorbidityarXiv:2608.09443v1 Announce Type: new Abstract: Large language model (LLM) agents can support medication review between clinical visits, but safe choices for older adults with multimorbidity depend on→11 Aug 2026BibTeX Citation Errors in Scientific Publishing Agents: Evaluation and MitigationarXiv:2604.03159v2 Announce Type: replace-cross Abstract: Large language models with web search are increasingly used in scientific publishing agents, yet they produce BibTeX entries with pervasive fi→11 Aug 2026Automating Deception: Scalable Multi-Turn LLM JailbreaksarXiv:2511.19517v3 Announce Type: replace-cross Abstract: Multi-turn conversational attacks, which leverage psychological principles like Foot-in-the-Door (FITD), where a small initial request paves t
CompanyMeta8 recent entries5 Aug 2026Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent EvaluationarXiv:2608.03166v1 Announce Type: new Abstract: Role-Playing Language Agents (RPLAs) are increasingly deployed in high-stakes applications such as healthcare assistance, customer support, and educatio→7 Aug 2026AI Playing Business Games: Benchmarking Large Language Models on Managerial Decision-Making in Dynamic SimulationsarXiv:2509.26331v2 Announce Type: replace Abstract: The rapid advancement of LLMs sparked significant interest in their potential to augment or automate managerial functions. One of the most recent tr→10 Aug 2026Latent Fact-Checking: Detecting Misinformation through Activation EngineeringarXiv:2608.06417v1 Announce Type: cross Abstract: The proliferation of misinformation online has driven demand for scalable detection systems. While most existing approaches rely on surface-level ling→10 Aug 2026Beyond Fluency: A Clinical Benchmark and Anomaly-Enhanced Baseline for Spine MRI Report GenerationarXiv:2608.07117v1 Announce Type: new Abstract: Radiology reporting is time-consuming and subject to inter-rater variability, making automated report generation an attractive clinical application for →11 Aug 2026Contamination Means Overestimation? A Fine-Grained Empirical Study in Code IntelligencearXiv:2506.02791v4 Announce Type: replace-cross Abstract: In recent years, code intelligence has gained increasing importance in the field of automated software engineering. Meanwhile, the widespread →11 Aug 2026Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded DimensionsarXiv:2608.09930v1 Announce Type: cross Abstract: Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expecte→11 Aug 2026Automated Generation of Complexity-Validated Decision Scenarios Using Large Language ModelsarXiv:2608.08822v1 Announce Type: new Abstract: Cognitive decision-making research depends on diverse scenarios with carefully controlled complexity, yet manual production is slow, inconsistent, and b→12 Aug 2026VidForensics-M1: Meta-Detection Reinforcement Learning with Verifiable Temporal Grounding for AI-Generated Video ForensicsarXiv:2608.11201v1 Announce Type: new Abstract: Recent advances in video generation models have significantly improved the realism of synthetic videos, blurring the boundary between generated and auth
CompanyMistral8 recent entries2 Jun 2026Truth, Trust, and Trouble: Medical AI on the EdgearXiv:2507.02983v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) hold significant promise for transforming digital health by enabling automated medical question answering. Howeve→2 Jun 2026Citation Grounding: Detecting and Reducing LLM Citation Hallucinations via Legal Citation GraphsarXiv:2606.00898v1 Announce Type: new Abstract: Large language models systematically hallucinate legal citations -- fabricating statute references, citing repealed provisions, and confusing jurisdicti→3 Jun 2026GTBench: A Curriculum-Grounded Benchmark for Evaluating LLMs as Mathematical Research Assistants in Graph TheoryarXiv:2606.03144v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as self-study assistants in technical disciplines, yet their reliability as mathematical reasoning as→3 Jun 2026Fully Automated Identification of Lexical Alignment and Preference-Stage Shifts in Large Language ModelsarXiv:2606.03165v1 Announce Type: cross Abstract: The language used by digital chat assistants such as ChatGPT can diverge from human expectations (misalignment). Research, mostly on Scientific Englis→6 Jun 2026Synthetic Contrastive Reasoning for Multi-Table Q&AarXiv:2606.05382v1 Announce Type: new Abstract: Multi-table question answering requires models to retrieve relevant evidence, link schemas, and perform compositional reasoning across relational tables→24 Jun 2026Not All Invariants Are Equal: Curating Training Data to Accelerate Program Verification with SLMsarXiv:2603.15510v2 Announce Type: replace Abstract: The synthesis of inductive loop invariants remains a critical bottleneck in automated program verification. While Large Language Models (LLMs) show →5 Aug 2026Evaluating LLM Trade-offs for Enterprise Automation: Lessons from Workflow Generation in a Production Enterprise PlatformarXiv:2608.03311v1 Announce Type: cross Abstract: Enterprise compliance management requires rapid adaptation to evolving regulatory frameworks (e.g., DORA, AI RMF, FedRAMP) and tight remediation SLAs.→7 Aug 2026AI Playing Business Games: Benchmarking Large Language Models on Managerial Decision-Making in Dynamic SimulationsarXiv:2509.26331v2 Announce Type: replace Abstract: The rapid advancement of LLMs sparked significant interest in their potential to augment or automate managerial functions. One of the most recent tr
CompanyxAI8 recent entries16 Jul 2026Explainable Artificial Intelligence for Anomaly Detection in Banking Transactions: An Internal Audit PerspectivearXiv:2607.13469v1 Announce Type: cross Abstract: The banking sector increasingly relies on automated systems to monitor electronic transactions for signs of fraud, yet conventional rule-based approac→23 Jul 2026Scaling Time Series Classification via XAI-Driven Data ReductionarXiv:2607.15774v2 Announce Type: replace-cross Abstract: Explainable AI (XAI) for time series has seen significant algorithmic growth, but its utility in providing measurable performance gains for do→24 Jul 2026Explanation-Based Runtime Verification for Trustworthy ML-driven Optical NetworksarXiv:2607.20675v1 Announce Type: new Abstract: Machine learning (ML) models are increasingly integrated into optical network automation frameworks to support tasks such as failure management, perform→24 Jul 2026Counterfactual Explainability Framework With CycleGAN And Counterfactual-Classifier Alignnment Score for Retinal Disease ClassificationarXiv:2607.21068v1 Announce Type: cross Abstract: Automated detection of vision impairing retina-based ocular conditions from fundus images is important for early screening, timely referral and reduci→28 Jul 2026Guiding Language Models to Be More Empathetic: Culturally Sensitive Mental Health Advice Generation Through Human-LLM CollaborationarXiv:2607.23538v1 Announce Type: new Abstract: Despite recent advances in large language models (LLMs), their ability to generate empathetic mental health counseling responses in low-resource languag→28 Jul 2026Evaluating the Impact of Explainable AI on Trust in AI-Assisted Code ReviewarXiv:2607.24601v1 Announce Type: cross Abstract: Background: Large language models (LLMs) are increasingly used to automate code review, but the reasoning behind their decisions remains hard to under→5 Aug 2026Explainable AI for the EU Right to Explanation: A Systematic Review of the Law-XAI Translation GaparXiv:2608.02699v1 Announce Type: new Abstract: When algorithms make or influence consequential decisions---about loan eligibility, hiring, or healthcare---EU law grants affected individuals a Right t→7 Aug 2026AI Playing Business Games: Benchmarking Large Language Models on Managerial Decision-Making in Dynamic SimulationsarXiv:2509.26331v2 Announce Type: replace Abstract: The rapid advancement of LLMs sparked significant interest in their potential to augment or automate managerial functions. One of the most recent tr
CompanyDeepSeek8 recent entries24 Jul 2026What is Good? Extracting and Testing Implicit Theories of Literary Quality from LLM Reasoning TracesarXiv:2607.20425v1 Announce Type: new Abstract: What makes writing 'good' remains a persistent question in literary studies and computational linguistics. We present a two-study investigation of how r→28 Jul 2026DualityCert: Verifier-Gated Language-Model Repair of Broken Duality Claims in Quantum Field TheoryarXiv:2607.23614v1 Announce Type: cross Abstract: We present DualityCert, a symbolic verifier for candidate Seiberg-duality claims in four-dimensional N=1 quiver gauge theories. The verifier evaluates→4 Aug 2026Cost-Effective Automated Judging of Natural-Language Mathematical ProofsarXiv:2608.00004v1 Announce Type: new Abstract: Grading natural-language mathematical proofs is a recurring cost in evaluating math-reasoning systems, and frontier LLM judges are expensive. We ask whe→6 Aug 2026Same Formulas, Different Semantics: Do Language Models Follow Modal Logic Specifications?arXiv:2608.05097v1 Announce Type: new Abstract: Reasoning about necessity and possibility depends on assumptions about accessibility between worlds and about which objects exist at each one. The same →10 Aug 2026Fisher-R1: Training LLM Agents for Reliable Hypothesis TestingarXiv:2608.07437v1 Announce Type: new Abstract: Reliable hypothesis testing is the foundation of many empirical scientific claims. Large language model (LLM) agents are increasingly used to automate t→11 Aug 2026Automated Generation of Complexity-Validated Decision Scenarios Using Large Language ModelsarXiv:2608.08822v1 Announce Type: new Abstract: Cognitive decision-making research depends on diverse scenarios with carefully controlled complexity, yet manual production is slow, inconsistent, and b→12 Aug 2026Persistent Recursive Worlds Enable Autonomous Software EvolutionarXiv:2608.10450v1 Announce Type: cross Abstract: Complex software systems develop over timescales that exceed the lifespan of any individual coding agent. Most agentic software systems preserve conti→12 Aug 2026CHORUS: Complementary Experts for High-Coverage Testbench Stimulus GenerationarXiv:2608.10090v1 Announce Type: new Abstract: Large language models (LLMs) have advanced code generation, where executable feedback provides a more reliable learning signal than textual imitation al
CompanyNVIDIA8 recent entries24 Jun 2026Systematic Exploration of 4-Expert Heterogeneous Mixture-of-Experts via Automated Pipeline SearcharXiv:2606.23739v1 Announce Type: cross Abstract: We present an automated large-scale search pipeline for heterogeneous 4-Expert Mixture-of-Experts (MoE4) architectures within the LEMUR neural network→30 Jun 2026Multi-Agentic System Leveraging Open-Source LLMs to Mitigate Disinformation ThreatsarXiv:2606.30259v1 Announce Type: new Abstract: In contemporary societies, the threat of disinformation has reached alarming levels, exacerbated by the proliferation of electronic communication, socia→8 Jul 2026KernelEvolve: Scaling Agentic Kernel Coding for Heterogeneous AI Accelerators at MetaarXiv:2512.23236v4 Announce Type: replace-cross Abstract: Making deep learning recommendation model (DLRM) training and inference fast and efficient is important. However, this presents three key syst→9 Jul 2026DYNA-PRUNER: Input-Adaptive Data-Model Co-Pruning for Efficient and Scalable Spatio-Temporal Media PredictionarXiv:2606.15346v2 Announce Type: replace Abstract: Spatio-temporal prediction supports radar/satellite nowcasting and city-scale traffic monitoring, but modern models are often too expensive for real→15 Jul 2026WanToFight: Real-Time Generative Game Engine for Multi-Player Combat InteractionarXiv:2607.12592v1 Announce Type: new Abstract: We present WanToFight, a generative game engine that simulates real-time, two-player The King of Fighters '97 (KOF~'97) gameplay from keyboard input. Pr→28 Jul 2026Energy Constrained Hierarchical Underwater Monitoring via Local Multi-Agent RAGarXiv:2607.24313v1 Announce Type: cross Abstract: Marine life monitoring is limited by strict energy constraints, poor underwater connectivity, and the high cost of transmitting raw multimodal data fr→7 Aug 2026CUDA-L2: Surpassing cuBLAS Performance for Matrix Multiplication through Reinforcement LearningarXiv:2512.02551v4 Announce Type: replace-cross Abstract: In this paper, we propose CUDA-L2, a system that combines large language models (LLMs) and reinforcement learning (RL) to automatically optimi→12 Aug 2026TransitReID: Transit OD Data Collection with Occlusion-Resistant Dynamic Passenger Re-IdentificationarXiv:2504.11500v3 Announce Type: replace-cross Abstract: Transit Origin-Destination (OD) data are fundamental for optimizing public transit services, yet current collection methods, such as manual su