CompanyAnthropic8 recent entries5 Aug 2026How Closely Do LLM Reviews Align with Human Peer Review?arXiv:2608.03659v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to generate scientific reviews, yet existing evaluations rarely examine whether different providers→5 Aug 2026AI World Cup 2026: Benchmarking Large Language Models for End-to-End Football Tournament PredictionarXiv:2608.03416v1 Announce Type: new Abstract: Large language models (LLMs) are now regularly asked to forecast real-world events, but comparisons are often difficult because models receive different
CompanyOpenAI8 recent entries5 Aug 2026How Closely Do LLM Reviews Align with Human Peer Review?arXiv:2608.03659v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to generate scientific reviews, yet existing evaluations rarely examine whether different providers→6 Aug 2026Can Post-Training Transform LLMs into Causal Reasoners?arXiv:2602.06337v2 Announce Type: replace-cross Abstract: Causal inference is essential for decision-making but remains challenging for non-experts. While large language models (LLMs) show promise in →7 Aug 2026Stochastic Parrots or Singing in Harmony? Testing Five Leading LLMs for their Ability to Replicate a Human Survey with Synthetic DataarXiv:2603.00059v3 Announce Type: replace-cross Abstract: How well can AI-derived synthetic research data replicate the responses of human participants? An emerging literature has begun to engage with→7 Aug 2026ProDVI: Programmatic Dynamics Priors for Value Network InitializationarXiv:2608.06015v1 Announce Type: cross Abstract: Deep Reinforcement Learning (RL) is notoriously sample inefficient. One contributing factor is that RL agents are typically initialized from scratch, →10 Aug 2026Social World ModelsarXiv:2509.00559v3 Announce Type: replace Abstract: Humans intuitively navigate social interactions by simulating unspoken dynamics and reasoning about others' perspectives, even with limited informat→10 Aug 2026Grammar Engineering Meets LLMs: Development of Cantonese and Irish ParGram TreebanksarXiv:2608.07283v1 Announce Type: new Abstract: Grammar engineering requires expertise in linguistic formalism and computational implementation, especially in parallel grammar projects that balance cr→11 Aug 2026How to Ask the AI: A User Perspective Survey for Large Language Model PromptingarXiv:2608.07494v1 Announce Type: cross Abstract: AI tools like ChatGPT and DeepSeek, powered by Large Language Models (LLMs), allow users to obtain instant and effective content responses simply by t→11 Aug 2026Curriculum Generation under Structured Parametric Environments for Robust Navigation PoliciesarXiv:2608.08545v1 Announce Type: cross Abstract: Robust navigation policies for autonomous agents must generalize across continuously varying environmental conditions such as turn rates, obstacles, f
CompanyGoogle8 recent entries7 Aug 2026Stochastic Parrots or Singing in Harmony? Testing Five Leading LLMs for their Ability to Replicate a Human Survey with Synthetic DataarXiv:2603.00059v3 Announce Type: replace-cross Abstract: How well can AI-derived synthetic research data replicate the responses of human participants? An emerging literature has begun to engage with→7 Aug 2026ProDVI: Programmatic Dynamics Priors for Value Network InitializationarXiv:2608.06015v1 Announce Type: cross Abstract: Deep Reinforcement Learning (RL) is notoriously sample inefficient. One contributing factor is that RL agents are typically initialized from scratch, →7 Aug 2026Observation-Grounded Self-Predictive Reinforcement Learning for Visual Continuous ControlarXiv:2608.05989v1 Announce Type: new Abstract: Sample-efficient policy learning from pixels is a long-standing challenge in reinforcement learning (RL). Recent dynamics-based representation learning →7 Aug 2026Hijacking Robots with a Piece of Paper: A Systematic Study of Physical Prompt Injection in VLM-Controlled RobotsarXiv:2608.05715v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) are increasingly deployed as planners in robotic systems, where they translate natural-language commands into executable→10 Aug 2026Artificial Intelligence Can Match Domain Experts in Evidence Extraction and Critical Appraisal of Microbial Oncogenesis Research PublicationsarXiv:2608.07250v1 Announce Type: cross Abstract: Confirmed oncogenic microbes contribute significantly to cancer burden. Identifying novel microbial oncogenicity could yield strategies that will redu→11 Aug 2026When Grammar Guides the Attack: Uncovering Control-Plane Vulnerabilities in LLMs with Structured OutputarXiv:2503.24191v4 Announce Type: replace-cross Abstract: Content Warning: This paper may contain unsafe or harmful content generated by LLMs that may be offensive to readers. Large Language Models (L→11 Aug 2026BibTeX Citation Errors in Scientific Publishing Agents: Evaluation and MitigationarXiv:2604.03159v2 Announce Type: replace-cross Abstract: Large language models with web search are increasingly used in scientific publishing agents, yet they produce BibTeX entries with pervasive fi→11 Aug 2026Automating Deception: Scalable Multi-Turn LLM JailbreaksarXiv:2511.19517v3 Announce Type: replace-cross Abstract: Multi-turn conversational attacks, which leverage psychological principles like Foot-in-the-Door (FITD), where a small initial request paves t
CompanyMeta8 recent entries10 Aug 2026Counterfactual Simulation Training for Chain-of-Thought FaithfulnessarXiv:2602.20710v2 Announce Type: replace Abstract: Inspecting Chain-of-Thought reasoning is among the most common means of understanding why an LLM produced its output. But well-known problems with C→11 Aug 2026SAGE: SLO-Aware Adaptive Retrieval for Production RAG SystemsarXiv:2608.08237v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) systems in production operate under strict service level objectives (SLOs) on tail latency and infrastructure cos→11 Aug 2026Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful JailbreaksarXiv:2608.09624v1 Announce Type: cross Abstract: Internal safety scores judge a prompt before any text is generated, and they are validated by how well they separate harmful prompts from benign ones.→11 Aug 2026Contamination Means Overestimation? A Fine-Grained Empirical Study in Code IntelligencearXiv:2506.02791v4 Announce Type: replace-cross Abstract: In recent years, code intelligence has gained increasing importance in the field of automated software engineering. Meanwhile, the widespread →11 Aug 2026Agentic AI-driven Immersive Simulation: A Knowledge-Aware Virtual Training Platform forHigh Dose Rate (HDR) BrachytherapyarXiv:2608.08163v1 Announce Type: new Abstract: The convergence of the Metaverse and Large Language Model (LLM)-based AI agent is catalyzing a shift toward autonomous, immersive, and personalized peda→11 Aug 2026A New Approach to Characterising Optimisation Problems Using Programmatic Representation and Complexity MeasuresarXiv:2608.08898v1 Announce Type: cross Abstract: Characterising optimisation problem instances is a fundamental part of understanding the behaviour and performance of different algorithms as well as →12 Aug 2026Eleven Years of BRACIS: A Meta-Scientific Study of the Brazilian Conference on Intelligent SystemsarXiv:2608.09964v1 Announce Type: cross Abstract: The Brazilian Conference on Intelligent Systems (BRACIS) is the main national venue for Artificial Intelligence research in Brazil, hosted by the Braz→12 Aug 2026Astrolabe: Balancing Load in LLM Serving with Randomized Prediction-Guided SchedulingarXiv:2508.03611v3 Announce Type: replace-cross Abstract: This paper presents Astrolabe, a randomized prediction-guided scheduler for one-shot request dispatch in multi-instance large language model (
CompanyMistral8 recent entries9 Jul 2026Measuring the metacognition of AIarXiv:2603.29693v3 Announce Type: replace Abstract: A robust decision-making process must take into account uncertainty, especially when the choice involves inherent risks. Because artificial intellig→23 Jul 2026Benchmarking Confidential GPU Inference on NVIDIA H100 under Intel TDXarXiv:2607.19353v1 Announce Type: new Abstract: Confidential computing is becoming a practical deployment requirement for AI inference workloads that process sensitive inputs or protect proprietary mo→28 Jul 2026LLM-Assisted Ontology Engineering and Construction of a French Legal Knowledge GrapharXiv:2607.24551v1 Announce Type: new Abstract: Maintenance regulations are complex legal texts that are difficult to exploit when addressing a specific case and challenging to integrate into operatio→29 Jul 2026Decision-Level Hijacking: Injecting Cognitive Bias into Large Language Models via Bit-Flip AttacksarXiv:2607.25227v1 Announce Type: cross Abstract: Large Language Models (LLMs) have been widely applied in high-stakes decision-making scenarios such as corporate strategy, and users are increasingly →30 Jul 2026Evaluating Prompt Scope and Demonstration Similarity in Local LLM Machine TranslationarXiv:2607.26286v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as general-purpose translation systems, but their behavior is usually evaluated under a single prompt→5 Aug 2026Evaluating LLM Trade-offs for Enterprise Automation: Lessons from Workflow Generation in a Production Enterprise PlatformarXiv:2608.03311v1 Announce Type: cross Abstract: Enterprise compliance management requires rapid adaptation to evolving regulatory frameworks (e.g., DORA, AI RMF, FedRAMP) and tight remediation SLAs.→5 Aug 2026DenialRAG: Single-Document RAG Poisoning via Embedded Parametric DenialarXiv:2608.02678v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) systems are vulnerable to corpus poisoning: an attacker who inserts a crafted document into the retrieval corpus →11 Aug 2026SAGE: SLO-Aware Adaptive Retrieval for Production RAG SystemsarXiv:2608.08237v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) systems in production operate under strict service level objectives (SLOs) on tail latency and infrastructure cos
CompanyxAI8 recent entries23 Jul 2026Scaling Time Series Classification via XAI-Driven Data ReductionarXiv:2607.15774v2 Announce Type: replace-cross Abstract: Explainable AI (XAI) for time series has seen significant algorithmic growth, but its utility in providing measurable performance gains for do→29 Jul 2026On the Design and Evaluation of Human-centered Explainable AI Systems: A Systematic Review and TaxonomyarXiv:2510.12201v2 Announce Type: replace Abstract: As AI becomes more common in everyday living, there is an increasing demand for intelligent systems that are both performant and understandable. Exp→4 Aug 2026Explainable Multimodal AI for Adaptive Calibration of Archaeological Sensing WorkflowsarXiv:2608.00074v1 Announce Type: new Abstract: This paper presents a multimodal machine-learning framework for calibration monitoring, quality assessment, and adaptive acquisition support in archaeol→4 Aug 2026Does Explainability Transfer? A Controlled Benchmark of Attribution Methods on Vision Transformers and CNNsarXiv:2608.02396v1 Announce Type: new Abstract: Most evidence on the effectiveness of explainable artificial intelligence (XAI) attribution methods has been established on convolutional neural network→5 Aug 2026Explainable AI for the EU Right to Explanation: A Systematic Review of the Law-XAI Translation GaparXiv:2608.02699v1 Announce Type: new Abstract: When algorithms make or influence consequential decisions---about loan eligibility, hiring, or healthcare---EU law grants affected individuals a Right t→7 Aug 2026Challenges in Evaluating Explanation Methods for Static and Evolving DataarXiv:2608.06351v1 Announce Type: new Abstract: This paper addresses the limitations of Explainable Artificial Intelligence (XAI) with respect to insufficient evaluation. They are illustrated through →12 Aug 2026The Epistemic Politics of AI AnthropomorphismarXiv:2608.00961v2 Announce Type: replace-cross Abstract: AI anthropomorphism is typically treated as a problem of user misperception requiring institutional correction. Users who engage in sustained →12 Aug 2026Entropy-Centric Explainable AI for Remote Sensing Image SegmentationarXiv:2608.11064v1 Announce Type: cross Abstract: Artificial intelligence (AI) has become a powerful approach to solving complex problems in critical domains. Many concerns arise regarding the decisio
CompanyDeepSeek8 recent entries31 Jul 2026Safety-Gated Agentic Supervisory Control on a Coupled Distillation Benchmark: Regime Map, Auditable Gate, and Co-Design FindingsarXiv:2607.27849v1 Announce Type: cross Abstract: An open-weight LLM can write composition setpoints every five minutes. What a plant still needs is a hard check: named constraints, logged margins, an→5 Aug 2026Approximate Speculative DecodingarXiv:2608.03447v1 Announce Type: cross Abstract: Speculative decoding accelerates autoregressive generation by verifying a draft block with a target model in parallel. Under standard greedy verificat→7 Aug 2026Stochastic Parrots or Singing in Harmony? Testing Five Leading LLMs for their Ability to Replicate a Human Survey with Synthetic DataarXiv:2603.00059v3 Announce Type: replace-cross Abstract: How well can AI-derived synthetic research data replicate the responses of human participants? An emerging literature has begun to engage with→7 Aug 2026Runtime Observability for Heterogeneous Attention MemoryarXiv:2608.05863v1 Announce Type: new Abstract: Modern models no longer keep a plain KV cache: latent caches, learned sparse selectors and recurrent states each carry the model's memory in a different→10 Aug 2026Transformers Struggle to Use Their Emergent World Models: Revisiting the Tower of Hanoi, and the Illusion of ThinkingarXiv:2608.07077v1 Announce Type: new Abstract: The Tower of Hanoi is a simple planning puzzle that in prior work has proven challenging for large reasoning models (LRMs). Current models solve the sta→11 Aug 2026When Grammar Guides the Attack: Uncovering Control-Plane Vulnerabilities in LLMs with Structured OutputarXiv:2503.24191v4 Announce Type: replace-cross Abstract: Content Warning: This paper may contain unsafe or harmful content generated by LLMs that may be offensive to readers. Large Language Models (L→11 Aug 2026How to Ask the AI: A User Perspective Survey for Large Language Model PromptingarXiv:2608.07494v1 Announce Type: cross Abstract: AI tools like ChatGPT and DeepSeek, powered by Large Language Models (LLMs), allow users to obtain instant and effective content responses simply by t→11 Aug 2026Adversarial Attacks on Deep OCR SystemsarXiv:2608.07636v1 Announce Type: cross Abstract: Deep-OCR (DeepSeek-OCR) advances document recognition by treating the visual modality as an optical compression medium, enabling long-context OCR at l
CompanyNVIDIA8 recent entries6 Aug 2026An Explainable LLM Agent Layer for Open-World Anomaly Detection in Oil WellsarXiv:2608.04041v1 Announce Type: new Abstract: Open-World Learning (OWL) pipelines for oil well anomaly detection have recently been shown to combine autoencoder-based detection, multiclass classific→7 Aug 2026Unified Planning-Learning Framework for Robust UUV Navigation Under Partial ObservabilityarXiv:2608.05365v1 Announce Type: new Abstract: This paper presents an observation-only autonomy framework for Unmanned Underwater Vehicles (UUVs) navigation in dynamic underwater environments that in→7 Aug 2026Parallel-in-Time Nonlinear Optimal Control via GPU-native Sequential Convex ProgrammingarXiv:2603.10711v3 Announce Type: replace Abstract: Real-time solution of nonlinear optimal control problems remains challenging on embedded robotic hardware, where conventional solvers often rely on →7 Aug 2026IcFuzz: Fuzzing Isaac Sim with Semantic Stage Guidance and Multi-level MutationarXiv:2608.06088v1 Announce Type: new Abstract: Robotics simulators serve as a foundational infrastructure for embodied AI, facilitating safe and scalable robotic system development. NVIDIA Isaac Sim →7 Aug 2026CUDA-L2: Surpassing cuBLAS Performance for Matrix Multiplication through Reinforcement LearningarXiv:2512.02551v4 Announce Type: replace-cross Abstract: In this paper, we propose CUDA-L2, a system that combines large language models (LLMs) and reinforcement learning (RL) to automatically optimi→11 Aug 2026EsaacSim: A Multimodal Event Camera Add-on for NVIDIA Isaac SimarXiv:2608.08522v1 Announce Type: new Abstract: Event-based vision is becoming an increasingly important sensing paradigm for robotics, yet its adoption remains limited by sensor availability and the →11 Aug 2026City Sentinel: A Unified AI-Based Smart Surveillance Framework for Real-Time Multi-Threat Detection Using Deep LearningarXiv:2608.08887v1 Announce Type: new Abstract: Rapid urbanization has increased the need for surveillance systems that can monitor multiple public safety risks at the same time. Traditional systems o→12 Aug 2026Hand-Written PTX Tensor-Core GEMM Kernels: A Multi-Precision Study on NVIDIA L4arXiv:2608.10103v1 Announce Type: cross Abstract: High-performance Tensor Core kernels rely on a low-level PTX pipeline built from asynchronous data movement with cp.async, warp-level matrix loads wit