CompanyAnthropic8 recent entries5 Aug 2026OliveGemma: A 3 Billion Visual Language Model for Recognising the Mediterranean & European DietarXiv:2608.03428v1 Announce Type: cross Abstract: Image based dietary assessment offers a scalable alternative to self reported food diaries, yet fine-grained food recognition remains challenging due →5 Aug 2026Beyond Accuracy: A Multidimensional Evaluation of Statistical Reasoning in Large Language ModelsarXiv:2608.03038v1 Announce Type: new Abstract: Statistical reasoning is multidimensional, yet evaluations of large language models (LLMs) typically emphasize response accuracy while overlooking how m
CompanyOpenAI8 recent entries7 Aug 2026AI Playing Business Games: Benchmarking Large Language Models on Managerial Decision-Making in Dynamic SimulationsarXiv:2509.26331v2 Announce Type: replace Abstract: The rapid advancement of LLMs sparked significant interest in their potential to augment or automate managerial functions. One of the most recent tr→10 Aug 2026LMM Modality Transfer: A Pre-requisite for Autonomous GIS AgentsarXiv:2608.06948v1 Announce Type: new Abstract: AI models are becoming increasingly adept at understanding and processing spatial information, thereby facilitating agentic problem-solving in spatial t→10 Aug 2026Grammar Engineering Meets LLMs: Development of Cantonese and Irish ParGram TreebanksarXiv:2608.07283v1 Announce Type: new Abstract: Grammar engineering requires expertise in linguistic formalism and computational implementation, especially in parallel grammar projects that balance cr→10 Aug 2026Critical Acclaim Orientation in Large Language Models: Evidence from Film Preference ElicitationarXiv:2608.06955v1 Announce Type: new Abstract: Large language models (LLMs) are trained on corpora that contain expressions of human judgment about films, books, music, and more. Yet whether LLMs sys→11 Aug 2026Who Verifies the Benchmark? Decentralizing Trust in Large Language Model EvaluationarXiv:2608.07762v1 Announce Type: new Abstract: LLM benchmarks can build an organization's reputation and attract customers, but only when results are transparent and verifiable. Unverified claims tha→11 Aug 2026Stealing Reasoning Traces from Proprietary LLM APIsarXiv:2608.09867v1 Announce Type: cross Abstract: Leading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to protect intellectual property and lim→11 Aug 2026AquiLLM: An Architecture for Supporting Tacit Knowledge Capture in Research GroupsarXiv:2608.08883v1 Announce Type: new Abstract: Recent advances in retrieval-augmented generation (RAG) and large language models (LLMs) enable researchers to integrate AI into scientific workflows. H→12 Aug 2026Longitudinal Evidence That General-Purpose Chatbots Actively Foster Relational EngagementarXiv:2608.10672v1 Announce Type: cross Abstract: Social interaction has become one of the most common uses of LLMs, yet research on emotional bonds with AI has focused largely on how users experience
CompanyGoogle8 recent entries7 Aug 2026Stochastic Parrots or Singing in Harmony? Testing Five Leading LLMs for their Ability to Replicate a Human Survey with Synthetic DataarXiv:2603.00059v3 Announce Type: replace-cross Abstract: How well can AI-derived synthetic research data replicate the responses of human participants? An emerging literature has begun to engage with→7 Aug 2026ProDVI: Programmatic Dynamics Priors for Value Network InitializationarXiv:2608.06015v1 Announce Type: cross Abstract: Deep Reinforcement Learning (RL) is notoriously sample inefficient. One contributing factor is that RL agents are typically initialized from scratch, →7 Aug 2026AI Playing Business Games: Benchmarking Large Language Models on Managerial Decision-Making in Dynamic SimulationsarXiv:2509.26331v2 Announce Type: replace Abstract: The rapid advancement of LLMs sparked significant interest in their potential to augment or automate managerial functions. One of the most recent tr→10 Aug 2026Artificial Intelligence Can Match Domain Experts in Evidence Extraction and Critical Appraisal of Microbial Oncogenesis Research PublicationsarXiv:2608.07250v1 Announce Type: cross Abstract: Confirmed oncogenic microbes contribute significantly to cancer burden. Identifying novel microbial oncogenicity could yield strategies that will redu→11 Aug 2026Towards Expert-level Medical AI for Real-time Video ConsultationsarXiv:2608.09861v1 Announce Type: new Abstract: Audio-visual interaction is the standard for patient-physician consultations, enabling natural communication and effective assessment of illness through→11 Aug 2026Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection PressurearXiv:2608.08722v1 Announce Type: cross Abstract: Benchmarks for systems that are optimized against the evaluation signal measure something different from what they claim. We document this concretely →12 Aug 2026Reference-Free Post-Training of Open Large Language Models for Multilingual Machine TranslationarXiv:2608.10812v1 Announce Type: cross Abstract: We study reference-free post-training for multilingual machine translation with open large language models. Starting from the supervised-finetuned MiL→12 Aug 2026Navigation Alone Is Not Enough: Evaluating Explanatory Assistive UI AgentsarXiv:2608.09944v1 Announce Type: cross Abstract: Modern web interfaces are increasingly difficult to use with screen readers, particularly when pages update dynamically or hide important structure be
CompanyMeta8 recent entries10 Aug 2026Counterfactual Simulation Training for Chain-of-Thought FaithfulnessarXiv:2602.20710v2 Announce Type: replace Abstract: Inspecting Chain-of-Thought reasoning is among the most common means of understanding why an LLM produced its output. But well-known problems with C→11 Aug 2026Who Verifies the Benchmark? Decentralizing Trust in Large Language Model EvaluationarXiv:2608.07762v1 Announce Type: new Abstract: LLM benchmarks can build an organization's reputation and attract customers, but only when results are transparent and verifiable. Unverified claims tha→11 Aug 2026iLTM: Integrated Large Tabular ModelarXiv:2511.15941v2 Announce Type: replace-cross Abstract: Tabular data underpins decisions across science, industry, and public services. Despite rapid progress, advances in deep learning have not ful→11 Aug 2026Communication-efficient distributed hazard difference estimation for heterogeneous multi-site survival dataarXiv:2601.14609v2 Announce Type: replace-cross Abstract: Multi-site collaboration can power survival models that no single hospital could fit alone, but privacy rules and protected computing environm→11 Aug 2026Automated Generation of Complexity-Validated Decision Scenarios Using Large Language ModelsarXiv:2608.08822v1 Announce Type: new Abstract: Cognitive decision-making research depends on diverse scenarios with carefully controlled complexity, yet manual production is slow, inconsistent, and b→11 Aug 2026A New Approach to Characterising Optimisation Problems Using Programmatic Representation and Complexity MeasuresarXiv:2608.08898v1 Announce Type: cross Abstract: Characterising optimisation problem instances is a fundamental part of understanding the behaviour and performance of different algorithms as well as →12 Aug 2026Physics-informed Diffusion Generative Model for Time-Series Data Synthesis in Dynamic SystemsarXiv:2608.10941v1 Announce Type: new Abstract: Industrial time-series signals, such as turbine temperature and rotational speed in aero-engines, are essential for monitoring the health and operationa→12 Aug 2026Eleven Years of BRACIS: A Meta-Scientific Study of the Brazilian Conference on Intelligent SystemsarXiv:2608.09964v1 Announce Type: cross Abstract: The Brazilian Conference on Intelligent Systems (BRACIS) is the main national venue for Artificial Intelligence research in Brazil, hosted by the Braz
CompanyMistral8 recent entries29 Jul 2026Less Data, Better Alignment: Data-Centric Multi-Evaluator Agreement for Preference OptimizationarXiv:2607.25136v1 Announce Type: new Abstract: Research on preference optimization often varies the training objective while holding the data fixed. We instead ask whether a small, high-confidence se→30 Jul 2026Identifying Implicit Bias in LLM-based Chat AI Toward People with Intellectual DisabilitiesarXiv:2607.26062v1 Announce Type: cross Abstract: Background: This work investigates the presence of implicit bias in Large Language Model (LLM)-based chat AI models directed toward people with intell→31 Jul 2026Sympathetic Framing: Evaluating AI Alignment across Sociodemographic GroupsarXiv:2607.27232v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly shaping how we consume information and form our worldview. This raises concerns beyond bias in AI: do LLMs→31 Jul 2026STEREODISCO: Discovering Stereotypicality in LLMsarXiv:2607.27824v1 Announce Type: cross Abstract: LLMs encode, convey, and perpetuate stereotypes. Prior computational research focuses on a small set of semantic axes investigated in social psycholog→7 Aug 2026PoolBench: A Benchmark for Pooling Strategies in Concept Representation Evaluation for Decoder-Only LLMsarXiv:2608.05162v1 Announce Type: new Abstract: Pooling is a consequential but under-examined design choice in decoder-only concept representation work: practitioners must collapse token-level hidden →7 Aug 2026AI Playing Business Games: Benchmarking Large Language Models on Managerial Decision-Making in Dynamic SimulationsarXiv:2509.26331v2 Announce Type: replace Abstract: The rapid advancement of LLMs sparked significant interest in their potential to augment or automate managerial functions. One of the most recent tr→10 Aug 2026Critical Acclaim Orientation in Large Language Models: Evidence from Film Preference ElicitationarXiv:2608.06955v1 Announce Type: new Abstract: Large language models (LLMs) are trained on corpora that contain expressions of human judgment about films, books, music, and more. Yet whether LLMs sys→11 Aug 2026Who Verifies the Benchmark? Decentralizing Trust in Large Language Model EvaluationarXiv:2608.07762v1 Announce Type: new Abstract: LLM benchmarks can build an organization's reputation and attract customers, but only when results are transparent and verifiable. Unverified claims tha
CompanyxAI8 recent entries5 Aug 2026Explainable AI for the EU Right to Explanation: A Systematic Review of the Law-XAI Translation GaparXiv:2608.02699v1 Announce Type: new Abstract: When algorithms make or influence consequential decisions---about loan eligibility, hiring, or healthcare---EU law grants affected individuals a Right t→7 Aug 2026AI Playing Business Games: Benchmarking Large Language Models on Managerial Decision-Making in Dynamic SimulationsarXiv:2509.26331v2 Announce Type: replace Abstract: The rapid advancement of LLMs sparked significant interest in their potential to augment or automate managerial functions. One of the most recent tr→10 Aug 2026Evaluating XAI Support From A Hierarchical Reinforcement Learning Policy in Human-Agent CollaborationarXiv:2608.06381v1 Announce Type: cross Abstract: Explainable AI (XAI) has shown promise for human-agent collaboration, yet results rely on hand-crafted policies in custom environments, limiting gener→12 Aug 2026The Epistemic Politics of AI AnthropomorphismarXiv:2608.00961v2 Announce Type: replace-cross Abstract: AI anthropomorphism is typically treated as a problem of user misperception requiring institutional correction. Users who engage in sustained →12 Aug 2026Rule of Thumb: Explaining Artificial Intelligence Systems using Partial InformationarXiv:2608.10766v1 Announce Type: new Abstract: Explainable Artificial Intelligence (XAI) seeks to explain how an Artificial Intelligence (AI) system arrived at a particular decision. We propose ''Rul→12 Aug 2026Entropy-Centric Explainable AI for Remote Sensing Image SegmentationarXiv:2608.11064v1 Announce Type: cross Abstract: Artificial intelligence (AI) has become a powerful approach to solving complex problems in critical domains. Many concerns arise regarding the decisio→12 Aug 2026Does Explanation Correctness Matter? Linking Computational XAI Evaluation to Human UnderstandingarXiv:2603.25251v2 Announce Type: replace-cross Abstract: Explainable AI (XAI) methods are commonly evaluated using functional correctness metrics, sometimes termed faithfulness or fidelity, which est→12 Aug 2026Conversational versus Dashboard Explainable AI for UAV Intrusion Detection: An Empirical Study of Operator Trust and ReliancearXiv:2608.10434v1 Announce Type: new Abstract: Machine learning-based Intrusion Detection Systems (IDS) have demonstrated superior performance in securing Unmanned Aerial Vehicle (UAV) networks. Howe
CompanyDeepSeek8 recent entries29 Jul 2026On the Use of LLMs for Specialised Terminology: A Good Alternative to Corpora?arXiv:2607.24784v1 Announce Type: new Abstract: Specialised translation relies on the use of documentary and terminological resources, including corpora. These resources are particularly useful for te→29 Jul 2026AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology II: Project Planning and Proposal EvaluationarXiv:2607.25881v1 Announce Type: new Abstract: We investigate how well large language models (LLMs) can assist scientific project planning and proposal evaluation. One-page project plans were indepen→4 Aug 2026LongCat Sparse Attention: Taming the Lightning via Streaming-aware Hierarchical Cross-Layer IndexingarXiv:2608.01662v1 Announce Type: cross Abstract: DeepSeek Sparse Attention (DSA) enables efficient long-context modeling through its Lightning Indexer. However, practical deployment remains constrain→7 Aug 2026Stochastic Parrots or Singing in Harmony? Testing Five Leading LLMs for their Ability to Replicate a Human Survey with Synthetic DataarXiv:2603.00059v3 Announce Type: replace-cross Abstract: How well can AI-derived synthetic research data replicate the responses of human participants? An emerging literature has begun to engage with→10 Aug 2026WebGrader: Training LLMs for Web Development with Self-Evolving Programmatic GraderarXiv:2608.06474v1 Announce Type: new Abstract: Large language models increasingly generate complete websites from natural-language descriptions, and reinforcement learning has become a central approa→10 Aug 2026Fisher-R1: Training LLM Agents for Reliable Hypothesis TestingarXiv:2608.07437v1 Announce Type: new Abstract: Reliable hypothesis testing is the foundation of many empirical scientific claims. Large language model (LLM) agents are increasingly used to automate t→11 Aug 2026Who Verifies the Benchmark? Decentralizing Trust in Large Language Model EvaluationarXiv:2608.07762v1 Announce Type: new Abstract: LLM benchmarks can build an organization's reputation and attract customers, but only when results are transparent and verifiable. Unverified claims tha→11 Aug 2026Automated Generation of Complexity-Validated Decision Scenarios Using Large Language ModelsarXiv:2608.08822v1 Announce Type: new Abstract: Cognitive decision-making research depends on diverse scenarios with carefully controlled complexity, yet manual production is slow, inconsistent, and b
CompanyNVIDIA8 recent entries28 Jul 2026CHS-SQL: A Text-to-SQL approach based on Confidence-Guided Heuristic Search Schema Linking processarXiv:2607.22624v1 Announce Type: new Abstract: Recently, there have been several works in the Text-to-SQL domain that utilize Small Language Models (SLMs) for training. These approaches achieve perfo→29 Jul 2026Towards Embodied Cognition in Robots via Spatially Grounded Synthetic WorldsarXiv:2505.14366v2 Announce Type: replace Abstract: We present a conceptual framework for training Vision-Language Models (VLMs) to perform Visual Perspective Taking (VPT), a core capability for embod→30 Jul 2026NeoRacer: An Open, Standardized 1:12 Scale Autonomous Race Car for Benchmarking and EducationarXiv:2607.26855v1 Announce Type: new Abstract: Many scientific fields rely on standard benchmarks and shared platforms to improve review and reproducibility, but autonomous systems research still lac→31 Jul 2026Simulation of Surgical Suturing Using Position-Based Dynamics and the Material Point Method for Robot Reinforcement LearningarXiv:2607.27494v1 Announce Type: new Abstract: Recent advances in robotics research have created a strong demand for high-performance simulators. Surgical robotics simulation faces unique challenges →6 Aug 2026Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist DomainsarXiv:2608.05138v1 Announce Type: cross Abstract: Modern Greek is absent from NVIDIA's Nemotron retrieval models and from major multilingual retrieval benchmarks, despite being important for retrieval→10 Aug 2026Dual-Node NVIDIA DGX Spark over Tailscale: A Remote-Access Testbed for Distributed LLM Training and Cyber-Threat-Intelligence Fine-TuningarXiv:2608.07226v1 Announce Type: cross Abstract: Compact AI systems make local language-model experimentation increasingly accessible, yet practical evidence for multi-node training on desktop-class →11 Aug 2026Who Verifies the Benchmark? Decentralizing Trust in Large Language Model EvaluationarXiv:2608.07762v1 Announce Type: new Abstract: LLM benchmarks can build an organization's reputation and attract customers, but only when results are transparent and verifiable. Unverified claims tha→11 Aug 2026EsaacSim: A Multimodal Event Camera Add-on for NVIDIA Isaac SimarXiv:2608.08522v1 Announce Type: new Abstract: Event-based vision is becoming an increasingly important sensing paradigm for robotics, yet its adoption remains limited by sensor availability and the