AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,202 results
CompaniesToolsTechniques

Each lane shows up to 8 recent matching entries, ordered from earlier to later. Tracks load separately to keep the 75,000+ entry wiki fast.

Companies

CompanyAnthropic8 recent entries
5 Aug 2026Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity

arXiv:2608.02665v1 Announce Type: cross Abstract: A benchmark score is a measurement instrument, yet most benchmarks read each item at a single canonical surface form. We ask whether that reading is f

→5 Aug 2026Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation

arXiv:2608.03166v1 Announce Type: new Abstract: Role-Playing Language Agents (RPLAs) are increasingly deployed in high-stakes applications such as healthcare assistance, customer support, and educatio


Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
→6 Aug 2026Diagnosing Tool-Selection Reasoning in LLM Agents with Canary Tools

arXiv:2608.04719v1 Announce Type: new Abstract: Agent evaluations tell us that a model picked the wrong tool, but rarely why. We introduce canary tools: diagnostic probe tools planted in an agent's Mo

→6 Aug 2026DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots

arXiv:2608.05004v1 Announce Type: new Abstract: Mental health professionals have raised concerns about risks of psychological harm from interaction with large language models (LLMs), including 'delusi

→7 Aug 2026Clinician input steers AI toward accurate and harmful recommendations

arXiv:2603.14158v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are entering clinical workflows, yet evaluations rarely assess how clinician reasoning shapes model behavior duri

→10 Aug 2026Divergent Response Modes in Frontier Language Models Under Steering Pressure

arXiv:2608.06578v1 Announce Type: new Abstract: Frontier language models are trained using distinct data, objectives, and safety pipelines. Whether these differences produce measurably different behav

→11 Aug 2026Stealing Reasoning Traces from Proprietary LLM APIs

arXiv:2608.09867v1 Announce Type: cross Abstract: Leading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to protect intellectual property and lim

→11 Aug 2026Automating Deception: Scalable Multi-Turn LLM Jailbreaks

arXiv:2511.19517v3 Announce Type: replace-cross Abstract: Multi-turn conversational attacks, which leverage psychological principles like Foot-in-the-Door (FITD), where a small initial request paves t

CompanyOpenAI8 recent entries
28 Jul 2026OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries

arXiv:2509.02594v3 Announce Type: replace-cross Abstract: Evaluating large language models (LLMs) on their ability to generate high-quality, accurate, situationally aware answers to clinical questions

→30 Jul 2026The Reliability of LLMs for Medical Diagnosis: An Examination of Consistency, Manipulation, and Contextual Awareness

arXiv:2503.10647v2 Announce Type: replace Abstract: This study evaluated the diagnostic reliability of two Large Language Models (LLMs), Google Gemini 2.0 Flash and OpenAI ChatGPT-4o, across three dim

→5 Aug 2026Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI Safety

arXiv:2601.17003v2 Announce Type: replace-cross Abstract: Mental-health AI safety is typically evaluated with small, simulation-based benchmarks that may not reflect the linguistic and contextual dive

→7 Aug 2026What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations)

arXiv:2608.06202v1 Announce Type: cross Abstract: Large language model (LLM) benchmark evaluations are routinely used to support claims about model safety, reliability, and deployment readiness. Yet m

→10 Aug 2026NiyamAI - An Intent-Bound AI Agent with Cryptographically Verifiable Guardrails using Zero-Knowledge Proofs

arXiv:2608.07167v1 Announce Type: new Abstract: Giving an AI agent the ability to send emails, query databases, or execute commands is useful--until the agent is tricked into doing something it should

→10 Aug 2026LMM Modality Transfer: A Pre-requisite for Autonomous GIS Agents

arXiv:2608.06948v1 Announce Type: new Abstract: AI models are becoming increasingly adept at understanding and processing spatial information, thereby facilitating agentic problem-solving in spatial t

→11 Aug 2026Stealing Reasoning Traces from Proprietary LLM APIs

arXiv:2608.09867v1 Announce Type: cross Abstract: Leading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to protect intellectual property and lim

→12 Aug 2026Rethinking LLM Verification: Evidence Structure, Uncertainty, and Selective Refinement

arXiv:2608.10725v1 Announce Type: new Abstract: Large language models (LLMs) often rely on shortcuts rather than systematic reasoning, raising safety concerns in medical applications. Allowing models

CompanyGoogle8 recent entries
7 Aug 2026Innocent Panels, Hateful Stories: Evaluating and Detecting Hateful Intent in Multi-Turn Visual Story Generation

arXiv:2608.05210v1 Announce Type: cross Abstract: Picture books and comics have long been used to disseminate hateful narratives because they are easily understood even by children, as exemplified by

→7 Aug 2026Clinician input steers AI toward accurate and harmful recommendations

arXiv:2603.14158v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are entering clinical workflows, yet evaluations rarely assess how clinician reasoning shapes model behavior duri

→10 Aug 2026Multi-Agent Forensic Reasoning for Generalizable Deepfake Video Detection

arXiv:2608.06865v1 Announce Type: cross Abstract: The malicious use of generative artificial intelligence to create highly realistic deepfake videos raises serious ethical concerns and poses substanti

→10 Aug 2026Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits

arXiv:2608.07430v1 Announce Type: cross Abstract: Diffusion Large Language Models (DLLMs) replace autoregressive next-token prediction with iterative parallel denoising, yet their internal safety mech

→11 Aug 2026When Grammar Guides the Attack: Uncovering Control-Plane Vulnerabilities in LLMs with Structured Output

arXiv:2503.24191v4 Announce Type: replace-cross Abstract: Content Warning: This paper may contain unsafe or harmful content generated by LLMs that may be offensive to readers. Large Language Models (L

→11 Aug 2026Coupled Graph--Policy Distillation for Personalized Medication Safety in Older Adults with Multimorbidity

arXiv:2608.09443v1 Announce Type: new Abstract: Large language model (LLM) agents can support medication review between clinical visits, but safe choices for older adults with multimorbidity depend on

→11 Aug 2026Automating Deception: Scalable Multi-Turn LLM Jailbreaks

arXiv:2511.19517v3 Announce Type: replace-cross Abstract: Multi-turn conversational attacks, which leverage psychological principles like Foot-in-the-Door (FITD), where a small initial request paves t

→12 Aug 2026Navigation Alone Is Not Enough: Evaluating Explanatory Assistive UI Agents

arXiv:2608.09944v1 Announce Type: cross Abstract: Modern web interfaces are increasingly difficult to use with screen readers, particularly when pages update dynamically or hide important structure be

CompanyMeta8 recent entries
10 Aug 2026Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits

arXiv:2608.07430v1 Announce Type: cross Abstract: Diffusion Large Language Models (DLLMs) replace autoregressive next-token prediction with iterative parallel denoising, yet their internal safety mech

→11 Aug 2026When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs

arXiv:2608.08542v1 Announce Type: new Abstract: Model merging has become the default way to give an aligned language model new skills without retraining: a practitioner folds task vectors from math, c

→11 Aug 2026When Do Task Vectors Interfere? Mapping the Validity Boundaries of Weight-Space Composition

arXiv:2608.09490v1 Announce Type: new Abstract: Task arithmetic treats fine-tuning displacements as composable directions in weight space, yet it remains unclear when parameter addition reflects predi

→11 Aug 2026Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks

arXiv:2608.09624v1 Announce Type: cross Abstract: Internal safety scores judge a prompt before any text is generated, and they are validated by how well they separate harmful prompts from benign ones.

→11 Aug 2026Learning Preference Adaptation for Large Language Model Personalization via Verbal Reinforcement Learning

arXiv:2608.09507v1 Announce Type: cross Abstract: Natural language user preferences provide an interpretable interface for LLM personalization. However, universal preference summaries often contain in

→11 Aug 2026Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families

arXiv:2608.08029v1 Announce Type: cross Abstract: Khatri et al. (2026) [DOI: 10.1109/DSN-W70714.2026.00027] show that lightweight MLP probes on final-layer activations of a single 8B model (LLaMA-3.1-

→11 Aug 2026Agentic AI-driven Immersive Simulation: A Knowledge-Aware Virtual Training Platform forHigh Dose Rate (HDR) Brachytherapy

arXiv:2608.08163v1 Announce Type: new Abstract: The convergence of the Metaverse and Large Language Model (LLM)-based AI agent is catalyzing a shift toward autonomous, immersive, and personalized peda

→12 Aug 2026SBCO: Self-Supervised, Verifier-Grounded Harness Optimization For Planning Agents

arXiv:2608.10157v1 Announce Type: new Abstract: Self-improving agents seek to reduce the human engineering effort behind AI systems by enabling them to evolve and self-improve their performance over t

CompanyMistral8 recent entries
3 Jul 2026Robust for the Wrong Reasons: The Representational Geometry of LLM Robustness to Science Skepticism

arXiv:2607.01951v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly consulted on contested scientific questions, raising the concern that they will sycophantically retreat

→3 Jul 2026Breaking Safety at the Token Boundary: How BPE Tokenization Creates Exploitable Gaps in LLM Alignment

arXiv:2607.01239v1 Announce Type: cross Abstract: Character-level perturbations bypass safety alignment in modern LLMs despite leaving prompts human-readable. We identify and test a central structural

→7 Jul 2026Governed MCP: Kernel-Level Tool Governance for AI Agents via Logit-Based Safety Primitives

arXiv:2604.16870v2 Announce Type: replace-cross Abstract: AI agents increasingly call external tools (file system, network, APIs) through the Model Context Protocol (MCP). These tool calls are the age

→31 Jul 2026STEREODISCO: Discovering Stereotypicality in LLMs

arXiv:2607.27824v1 Announce Type: cross Abstract: LLMs encode, convey, and perpetuate stereotypes. Prior computational research focuses on a small set of semantic axes investigated in social psycholog

→4 Aug 2026Efficiency vs. Alignment: Investigating Safety and Fairness Risks in Parameter-Efficient Fine-Tuning of LLMs

arXiv:2511.00382v2 Announce Type: replace-cross Abstract: Organizations increasingly adapt Large Language Models (LLMs) from public repositories such as HuggingFace to downstream tasks. Prior work sho

→11 Aug 2026Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families

arXiv:2608.08029v1 Announce Type: cross Abstract: Khatri et al. (2026) [DOI: 10.1109/DSN-W70714.2026.00027] show that lightweight MLP probes on final-layer activations of a single 8B model (LLaMA-3.1-

→12 Aug 2026The Truth Stays in the Family: Enhancing Contextual Grounding via Inherited Truthful Heads in Model Lineages

arXiv:2606.15821v2 Announce Type: replace-cross Abstract: Recent advances in large language models (LLMs) have produced many specialized multimodal LLMs (MLLMs) that share common foundational LLMs, fo

→12 Aug 2026REAP: Relation-Aware Elicitation and Parsing for Closed-Book Knowledge Base Construction from LLMs

arXiv:2608.10963v1 Announce Type: new Abstract: We present the REAP system for the AKBC Shared Task 2026 on constructing knowledge bases from language models in a closed-book setting, subject to a bud

CompanyxAI8 recent entries
28 Jul 2026Disentangling Acoustic Cues in Alzheimer's Pathology and Perception: The Roles of Language and Gender

arXiv:2607.23977v1 Announce Type: cross Abstract: Acoustic biomarkers show promise for detecting Alzheimer's Disease (AD), yet whether the cues driving diagnostic AI align with those salient to human

→28 Jul 2026Context-Aware Concept Distillation for Trustworthy Flood Prediction

arXiv:2607.23237v1 Announce Type: cross Abstract: Effective flood risk management relies on accurate forecasting, yet the 'black box' nature of stateof-the-art Deep Learning models creates a barrier t

→28 Jul 2026Auditing Alignment Controllability in LLMs via Political Axes

arXiv:2607.23519v1 Announce Type: cross Abstract: Political audits of large language models (LLMs) usually reduce each to one point on a political compass. But that resting point barely matters in dep

→29 Jul 2026From Dyad to Triad: Eliciting XAI Requirements in Stroke Rehabilitation

arXiv:2607.25423v1 Announce Type: cross Abstract: Eliciting explainable AI (XAI) requirements from stroke survivors presents a methodological challenge with direct implications for the design of trust

→31 Jul 2026Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game

arXiv:2607.28146v1 Announce Type: new Abstract: As large language models (LLMs) are deployed as agents in high-stakes settings, such as medical and legal systems, understanding their deceptive capabil

→31 Jul 2026AI LEGO: Scaffolding Cross-Functional Collaboration in Industrial Responsible AI Practices during Early Design Stages

arXiv:2505.10300v2 Announce Type: replace-cross Abstract: Responsible AI (RAI) efforts increasingly emphasize the importance of addressing potential harms early in the AI development lifecycle through

→7 Aug 2026Explanations of Large Language Models Explain Language Representations in the Brain

arXiv:2502.14671v4 Announce Type: replace-cross Abstract: Large Language Model (LLM) representations are known to align with brain activity during language processing, but it remains unclear what driv

→7 Aug 2026Challenges in Evaluating Explanation Methods for Static and Evolving Data

arXiv:2608.06351v1 Announce Type: new Abstract: This paper addresses the limitations of Explainable Artificial Intelligence (XAI) with respect to insufficient evaluation. They are illustrated through

CompanyDeepSeek8 recent entries
31 Jul 2026Safety-Gated Agentic Supervisory Control on a Coupled Distillation Benchmark: Regime Map, Auditable Gate, and Co-Design Findings

arXiv:2607.27849v1 Announce Type: cross Abstract: An open-weight LLM can write composition setpoints every five minutes. What a plant still needs is a hard check: named constraints, logged margins, an

→31 Jul 2026Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game

arXiv:2607.28146v1 Announce Type: new Abstract: As large language models (LLMs) are deployed as agents in high-stakes settings, such as medical and legal systems, understanding their deceptive capabil

→4 Aug 2026Why Formal Monitors Fail: Attack Distribution Entropy as a Coverage Bound for LTL-Based LLM Agent Safety

arXiv:2608.01388v1 Announce Type: cross Abstract: Runtime safety monitors based on Linear Temporal Logic (LTL) and finite automata (FSA) are increasingly deployed to intercept unsafe tool-call sequenc

→4 Aug 2026From Direction to Magnitude: How Multimodal Instruction-Tuning Reorganizes the Geometric Encoding of Identity-Specifying Prompts in Transformer Hidden States

arXiv:2607.09842v2 Announce Type: replace-cross Abstract: We investigate whether identity-specifying system prompts produce statistically distinguishable geometric fingerprints in the hidden-state tra

→5 Aug 2026Risky Business: Measuring The Faithfulness-Safety Tension

arXiv:2608.03745v1 Announce Type: new Abstract: Chain-of-Thought (CoT) reasoning offers a promising window into model monitoring. However, monitoring relies on faithfulness, i.e., the model output str

→5 Aug 2026Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI Safety

arXiv:2601.17003v2 Announce Type: replace-cross Abstract: Mental-health AI safety is typically evaluated with small, simulation-based benchmarks that may not reflect the linguistic and contextual dive

→11 Aug 2026When Grammar Guides the Attack: Uncovering Control-Plane Vulnerabilities in LLMs with Structured Output

arXiv:2503.24191v4 Announce Type: replace-cross Abstract: Content Warning: This paper may contain unsafe or harmful content generated by LLMs that may be offensive to readers. Large Language Models (L

→12 Aug 2026Rethinking LLM Verification: Evidence Structure, Uncertainty, and Selective Refinement

arXiv:2608.10725v1 Announce Type: new Abstract: Large language models (LLMs) often rely on shortcuts rather than systematic reasoning, raising safety concerns in medical applications. Allowing models

CompanyNVIDIA8 recent entries
31 Jul 2026Safety-Gated Agentic Supervisory Control on a Coupled Distillation Benchmark: Regime Map, Auditable Gate, and Co-Design Findings

arXiv:2607.27849v1 Announce Type: cross Abstract: An open-weight LLM can write composition setpoints every five minutes. What a plant still needs is a hard check: named constraints, logged margins, an

→3 Aug 2026TRACE: High-Fidelity 3D Scene Editing via Tangible Reconstruction and Geometry-Aligned Contextual Video Masking

arXiv:2604.01207v2 Announce Type: replace Abstract: Existing 3D Gaussian Splatting (3DGS) editing methods primarily focus on appearance modification and often struggle to support flexible geometry edi

→3 Aug 2026Temporal Policy: History-Initialized Action Generation for Robotic Learning from Demonstration

arXiv:2607.29482v1 Announce Type: new Abstract: By relying on independent couplings from uninformative Gaussian priors, standard diffusion and flow matching models are forced to learn complex, high-co

→4 Aug 2026Meganeura: Portable GPU Training and Inference through Vulkan and Metal

arXiv:2608.01563v1 Announce Type: new Abstract: Training and deployed inference often cross export, conversion, and platform-specific runtime boundaries. Meganeura asks whether one compact native comp

→5 Aug 2026Bimanual Manipulation Within an 8 GB Budget: Zero-Copy Sensing and Quantized ACT on an Entry-Level Jetson

arXiv:2608.03938v1 Announce Type: new Abstract: Bimanual manipulation policies trained with imitation learning are typically evaluated on workstation or datacenter-class GPUs, leaving the cost of depl

→7 Aug 2026Unified Planning-Learning Framework for Robust UUV Navigation Under Partial Observability

arXiv:2608.05365v1 Announce Type: new Abstract: This paper presents an observation-only autonomy framework for Unmanned Underwater Vehicles (UUVs) navigation in dynamic underwater environments that in

→10 Aug 2026SNI-GNN: SmartNIC-Assisted Full-Graph GNN Training with In-Network Embedding Prediction

arXiv:2608.06441v1 Announce Type: new Abstract: Full-graph GNN training delivers high accuracy but scales poorly on multi-server clusters due to heavy, irregular inter-node embedding exchanges. We pre

→11 Aug 2026City Sentinel: A Unified AI-Based Smart Surveillance Framework for Real-Time Multi-Threat Detection Using Deep Learning

arXiv:2608.08887v1 Announce Type: new Abstract: Rapid urbanization has increased the need for surveillance systems that can monitor multiple public safety risks at the same time. Traditional systems o