CompanyAnthropic8 recent entries5 Aug 2026Single Canonical Prompts Underestimate LLM Safety's Surface-Form SensitivityarXiv:2608.02665v1 Announce Type: cross Abstract: A benchmark score is a measurement instrument, yet most benchmarks read each item at a single canonical surface form. We ask whether that reading is f→5 Aug 2026Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent EvaluationarXiv:2608.03166v1 Announce Type: new Abstract: Role-Playing Language Agents (RPLAs) are increasingly deployed in high-stakes applications such as healthcare assistance, customer support, and educatio
CompanyOpenAI8 recent entries28 Jul 2026OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical QueriesarXiv:2509.02594v3 Announce Type: replace-cross Abstract: Evaluating large language models (LLMs) on their ability to generate high-quality, accurate, situationally aware answers to clinical questions→30 Jul 2026The Reliability of LLMs for Medical Diagnosis: An Examination of Consistency, Manipulation, and Contextual AwarenessarXiv:2503.10647v2 Announce Type: replace Abstract: This study evaluated the diagnostic reliability of two Large Language Models (LLMs), Google Gemini 2.0 Flash and OpenAI ChatGPT-4o, across three dim→5 Aug 2026Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI SafetyarXiv:2601.17003v2 Announce Type: replace-cross Abstract: Mental-health AI safety is typically evaluated with small, simulation-based benchmarks that may not reflect the linguistic and contextual dive→7 Aug 2026What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations)arXiv:2608.06202v1 Announce Type: cross Abstract: Large language model (LLM) benchmark evaluations are routinely used to support claims about model safety, reliability, and deployment readiness. Yet m→10 Aug 2026NiyamAI - An Intent-Bound AI Agent with Cryptographically Verifiable Guardrails using Zero-Knowledge ProofsarXiv:2608.07167v1 Announce Type: new Abstract: Giving an AI agent the ability to send emails, query databases, or execute commands is useful--until the agent is tricked into doing something it should→10 Aug 2026LMM Modality Transfer: A Pre-requisite for Autonomous GIS AgentsarXiv:2608.06948v1 Announce Type: new Abstract: AI models are becoming increasingly adept at understanding and processing spatial information, thereby facilitating agentic problem-solving in spatial t→11 Aug 2026Stealing Reasoning Traces from Proprietary LLM APIsarXiv:2608.09867v1 Announce Type: cross Abstract: Leading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to protect intellectual property and lim→12 Aug 2026Rethinking LLM Verification: Evidence Structure, Uncertainty, and Selective RefinementarXiv:2608.10725v1 Announce Type: new Abstract: Large language models (LLMs) often rely on shortcuts rather than systematic reasoning, raising safety concerns in medical applications. Allowing models
CompanyGoogle8 recent entries7 Aug 2026Innocent Panels, Hateful Stories: Evaluating and Detecting Hateful Intent in Multi-Turn Visual Story GenerationarXiv:2608.05210v1 Announce Type: cross Abstract: Picture books and comics have long been used to disseminate hateful narratives because they are easily understood even by children, as exemplified by →7 Aug 2026Clinician input steers AI toward accurate and harmful recommendationsarXiv:2603.14158v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are entering clinical workflows, yet evaluations rarely assess how clinician reasoning shapes model behavior duri→10 Aug 2026Multi-Agent Forensic Reasoning for Generalizable Deepfake Video DetectionarXiv:2608.06865v1 Announce Type: cross Abstract: The malicious use of generative artificial intelligence to create highly realistic deepfake videos raises serious ethical concerns and poses substanti→10 Aug 2026Diffusion LLMs as Targets and Adversaries: Mechanistic Safety ExploitsarXiv:2608.07430v1 Announce Type: cross Abstract: Diffusion Large Language Models (DLLMs) replace autoregressive next-token prediction with iterative parallel denoising, yet their internal safety mech→11 Aug 2026When Grammar Guides the Attack: Uncovering Control-Plane Vulnerabilities in LLMs with Structured OutputarXiv:2503.24191v4 Announce Type: replace-cross Abstract: Content Warning: This paper may contain unsafe or harmful content generated by LLMs that may be offensive to readers. Large Language Models (L→11 Aug 2026Coupled Graph--Policy Distillation for Personalized Medication Safety in Older Adults with MultimorbidityarXiv:2608.09443v1 Announce Type: new Abstract: Large language model (LLM) agents can support medication review between clinical visits, but safe choices for older adults with multimorbidity depend on→11 Aug 2026Automating Deception: Scalable Multi-Turn LLM JailbreaksarXiv:2511.19517v3 Announce Type: replace-cross Abstract: Multi-turn conversational attacks, which leverage psychological principles like Foot-in-the-Door (FITD), where a small initial request paves t→12 Aug 2026Navigation Alone Is Not Enough: Evaluating Explanatory Assistive UI AgentsarXiv:2608.09944v1 Announce Type: cross Abstract: Modern web interfaces are increasingly difficult to use with screen readers, particularly when pages update dynamically or hide important structure be
CompanyMeta8 recent entries10 Aug 2026Diffusion LLMs as Targets and Adversaries: Mechanistic Safety ExploitsarXiv:2608.07430v1 Announce Type: cross Abstract: Diffusion Large Language Models (DLLMs) replace autoregressive next-token prediction with iterative parallel denoising, yet their internal safety mech→11 Aug 2026When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMsarXiv:2608.08542v1 Announce Type: new Abstract: Model merging has become the default way to give an aligned language model new skills without retraining: a practitioner folds task vectors from math, c→11 Aug 2026When Do Task Vectors Interfere? Mapping the Validity Boundaries of Weight-Space CompositionarXiv:2608.09490v1 Announce Type: new Abstract: Task arithmetic treats fine-tuning displacements as composable directions in weight space, yet it remains unclear when parameter addition reflects predi→11 Aug 2026Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful JailbreaksarXiv:2608.09624v1 Announce Type: cross Abstract: Internal safety scores judge a prompt before any text is generated, and they are validated by how well they separate harmful prompts from benign ones.→11 Aug 2026Learning Preference Adaptation for Large Language Model Personalization via Verbal Reinforcement LearningarXiv:2608.09507v1 Announce Type: cross Abstract: Natural language user preferences provide an interpretable interface for LLM personalization. However, universal preference summaries often contain in→11 Aug 2026Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model FamiliesarXiv:2608.08029v1 Announce Type: cross Abstract: Khatri et al. (2026) [DOI: 10.1109/DSN-W70714.2026.00027] show that lightweight MLP probes on final-layer activations of a single 8B model (LLaMA-3.1-→11 Aug 2026Agentic AI-driven Immersive Simulation: A Knowledge-Aware Virtual Training Platform forHigh Dose Rate (HDR) BrachytherapyarXiv:2608.08163v1 Announce Type: new Abstract: The convergence of the Metaverse and Large Language Model (LLM)-based AI agent is catalyzing a shift toward autonomous, immersive, and personalized peda→12 Aug 2026SBCO: Self-Supervised, Verifier-Grounded Harness Optimization For Planning AgentsarXiv:2608.10157v1 Announce Type: new Abstract: Self-improving agents seek to reduce the human engineering effort behind AI systems by enabling them to evolve and self-improve their performance over t
CompanyMistral8 recent entries3 Jul 2026Robust for the Wrong Reasons: The Representational Geometry of LLM Robustness to Science SkepticismarXiv:2607.01951v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly consulted on contested scientific questions, raising the concern that they will sycophantically retreat →3 Jul 2026Breaking Safety at the Token Boundary: How BPE Tokenization Creates Exploitable Gaps in LLM AlignmentarXiv:2607.01239v1 Announce Type: cross Abstract: Character-level perturbations bypass safety alignment in modern LLMs despite leaving prompts human-readable. We identify and test a central structural→7 Jul 2026Governed MCP: Kernel-Level Tool Governance for AI Agents via Logit-Based Safety PrimitivesarXiv:2604.16870v2 Announce Type: replace-cross Abstract: AI agents increasingly call external tools (file system, network, APIs) through the Model Context Protocol (MCP). These tool calls are the age→31 Jul 2026STEREODISCO: Discovering Stereotypicality in LLMsarXiv:2607.27824v1 Announce Type: cross Abstract: LLMs encode, convey, and perpetuate stereotypes. Prior computational research focuses on a small set of semantic axes investigated in social psycholog→4 Aug 2026Efficiency vs. Alignment: Investigating Safety and Fairness Risks in Parameter-Efficient Fine-Tuning of LLMsarXiv:2511.00382v2 Announce Type: replace-cross Abstract: Organizations increasingly adapt Large Language Models (LLMs) from public repositories such as HuggingFace to downstream tasks. Prior work sho→11 Aug 2026Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model FamiliesarXiv:2608.08029v1 Announce Type: cross Abstract: Khatri et al. (2026) [DOI: 10.1109/DSN-W70714.2026.00027] show that lightweight MLP probes on final-layer activations of a single 8B model (LLaMA-3.1-→12 Aug 2026The Truth Stays in the Family: Enhancing Contextual Grounding via Inherited Truthful Heads in Model LineagesarXiv:2606.15821v2 Announce Type: replace-cross Abstract: Recent advances in large language models (LLMs) have produced many specialized multimodal LLMs (MLLMs) that share common foundational LLMs, fo→12 Aug 2026REAP: Relation-Aware Elicitation and Parsing for Closed-Book Knowledge Base Construction from LLMsarXiv:2608.10963v1 Announce Type: new Abstract: We present the REAP system for the AKBC Shared Task 2026 on constructing knowledge bases from language models in a closed-book setting, subject to a bud
CompanyxAI8 recent entries28 Jul 2026Disentangling Acoustic Cues in Alzheimer's Pathology and Perception: The Roles of Language and GenderarXiv:2607.23977v1 Announce Type: cross Abstract: Acoustic biomarkers show promise for detecting Alzheimer's Disease (AD), yet whether the cues driving diagnostic AI align with those salient to human →28 Jul 2026Context-Aware Concept Distillation for Trustworthy Flood PredictionarXiv:2607.23237v1 Announce Type: cross Abstract: Effective flood risk management relies on accurate forecasting, yet the 'black box' nature of stateof-the-art Deep Learning models creates a barrier t→28 Jul 2026Auditing Alignment Controllability in LLMs via Political AxesarXiv:2607.23519v1 Announce Type: cross Abstract: Political audits of large language models (LLMs) usually reduce each to one point on a political compass. But that resting point barely matters in dep→29 Jul 2026From Dyad to Triad: Eliciting XAI Requirements in Stroke RehabilitationarXiv:2607.25423v1 Announce Type: cross Abstract: Eliciting explainable AI (XAI) requirements from stroke survivors presents a methodological challenge with direct implications for the design of trust→31 Jul 2026Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction GamearXiv:2607.28146v1 Announce Type: new Abstract: As large language models (LLMs) are deployed as agents in high-stakes settings, such as medical and legal systems, understanding their deceptive capabil→31 Jul 2026AI LEGO: Scaffolding Cross-Functional Collaboration in Industrial Responsible AI Practices during Early Design StagesarXiv:2505.10300v2 Announce Type: replace-cross Abstract: Responsible AI (RAI) efforts increasingly emphasize the importance of addressing potential harms early in the AI development lifecycle through→7 Aug 2026Explanations of Large Language Models Explain Language Representations in the BrainarXiv:2502.14671v4 Announce Type: replace-cross Abstract: Large Language Model (LLM) representations are known to align with brain activity during language processing, but it remains unclear what driv→7 Aug 2026Challenges in Evaluating Explanation Methods for Static and Evolving DataarXiv:2608.06351v1 Announce Type: new Abstract: This paper addresses the limitations of Explainable Artificial Intelligence (XAI) with respect to insufficient evaluation. They are illustrated through
CompanyDeepSeek8 recent entries31 Jul 2026Safety-Gated Agentic Supervisory Control on a Coupled Distillation Benchmark: Regime Map, Auditable Gate, and Co-Design FindingsarXiv:2607.27849v1 Announce Type: cross Abstract: An open-weight LLM can write composition setpoints every five minutes. What a plant still needs is a hard check: named constraints, logged margins, an→31 Jul 2026Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction GamearXiv:2607.28146v1 Announce Type: new Abstract: As large language models (LLMs) are deployed as agents in high-stakes settings, such as medical and legal systems, understanding their deceptive capabil→4 Aug 2026Why Formal Monitors Fail: Attack Distribution Entropy as a Coverage Bound for LTL-Based LLM Agent SafetyarXiv:2608.01388v1 Announce Type: cross Abstract: Runtime safety monitors based on Linear Temporal Logic (LTL) and finite automata (FSA) are increasingly deployed to intercept unsafe tool-call sequenc→4 Aug 2026From Direction to Magnitude: How Multimodal Instruction-Tuning Reorganizes the Geometric Encoding of Identity-Specifying Prompts in Transformer Hidden StatesarXiv:2607.09842v2 Announce Type: replace-cross Abstract: We investigate whether identity-specifying system prompts produce statistically distinguishable geometric fingerprints in the hidden-state tra→5 Aug 2026Risky Business: Measuring The Faithfulness-Safety TensionarXiv:2608.03745v1 Announce Type: new Abstract: Chain-of-Thought (CoT) reasoning offers a promising window into model monitoring. However, monitoring relies on faithfulness, i.e., the model output str→5 Aug 2026Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI SafetyarXiv:2601.17003v2 Announce Type: replace-cross Abstract: Mental-health AI safety is typically evaluated with small, simulation-based benchmarks that may not reflect the linguistic and contextual dive→11 Aug 2026When Grammar Guides the Attack: Uncovering Control-Plane Vulnerabilities in LLMs with Structured OutputarXiv:2503.24191v4 Announce Type: replace-cross Abstract: Content Warning: This paper may contain unsafe or harmful content generated by LLMs that may be offensive to readers. Large Language Models (L→12 Aug 2026Rethinking LLM Verification: Evidence Structure, Uncertainty, and Selective RefinementarXiv:2608.10725v1 Announce Type: new Abstract: Large language models (LLMs) often rely on shortcuts rather than systematic reasoning, raising safety concerns in medical applications. Allowing models
CompanyNVIDIA8 recent entries31 Jul 2026Safety-Gated Agentic Supervisory Control on a Coupled Distillation Benchmark: Regime Map, Auditable Gate, and Co-Design FindingsarXiv:2607.27849v1 Announce Type: cross Abstract: An open-weight LLM can write composition setpoints every five minutes. What a plant still needs is a hard check: named constraints, logged margins, an→3 Aug 2026TRACE: High-Fidelity 3D Scene Editing via Tangible Reconstruction and Geometry-Aligned Contextual Video MaskingarXiv:2604.01207v2 Announce Type: replace Abstract: Existing 3D Gaussian Splatting (3DGS) editing methods primarily focus on appearance modification and often struggle to support flexible geometry edi→3 Aug 2026Temporal Policy: History-Initialized Action Generation for Robotic Learning from DemonstrationarXiv:2607.29482v1 Announce Type: new Abstract: By relying on independent couplings from uninformative Gaussian priors, standard diffusion and flow matching models are forced to learn complex, high-co→4 Aug 2026Meganeura: Portable GPU Training and Inference through Vulkan and MetalarXiv:2608.01563v1 Announce Type: new Abstract: Training and deployed inference often cross export, conversion, and platform-specific runtime boundaries. Meganeura asks whether one compact native comp→5 Aug 2026Bimanual Manipulation Within an 8 GB Budget: Zero-Copy Sensing and Quantized ACT on an Entry-Level JetsonarXiv:2608.03938v1 Announce Type: new Abstract: Bimanual manipulation policies trained with imitation learning are typically evaluated on workstation or datacenter-class GPUs, leaving the cost of depl→7 Aug 2026Unified Planning-Learning Framework for Robust UUV Navigation Under Partial ObservabilityarXiv:2608.05365v1 Announce Type: new Abstract: This paper presents an observation-only autonomy framework for Unmanned Underwater Vehicles (UUVs) navigation in dynamic underwater environments that in→10 Aug 2026SNI-GNN: SmartNIC-Assisted Full-Graph GNN Training with In-Network Embedding PredictionarXiv:2608.06441v1 Announce Type: new Abstract: Full-graph GNN training delivers high accuracy but scales poorly on multi-server clusters due to heavy, irregular inter-node embedding exchanges. We pre→11 Aug 2026City Sentinel: A Unified AI-Based Smart Surveillance Framework for Real-Time Multi-Threat Detection Using Deep LearningarXiv:2608.08887v1 Announce Type: new Abstract: Rapid urbanization has increased the need for surveillance systems that can monitor multiple public safety risks at the same time. Traditional systems o