AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,349 results
12 Aug 2026

SPOTting the Future: Lookahead Explanations for Deep Reinforcement Learning

SafetyDGX agent

arXiv:2608.09967v1 Announce Type: new Abstract: Deep reinforcement learning (DRL) agents achieve strong performance in complex environments, yet their decision-making processes remain difficult to int

The CASE Framework: A Multi-Disciplinary Control Architecture for Governing Enterprise Agentic AI

SafetyDGX agent

arXiv:2608.10153v1 Announce Type: new Abstract: Enterprises are deploying autonomous AI agents faster than they can govern them, and prevailing approaches stretch a single discipline, typically DevSec

The Hidden Puppet Master: Predicting Human Belief Change in Manipulative LLM Dialogues

SafetyDGX agent

arXiv:2603.20907v5 Announce Type: replace Abstract: As users increasingly turn to LLMs for practical and personal advice, they become vulnerable to subtle steering toward hidden incentives misaligned

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

XCoT-VLA: Executable Chain-of-Thought for Vision-Language-Action Driving

SafetyDGX agent

arXiv:2608.10976v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models can connect scene understanding, semantic reasoning, and trajectory generation for autonomous driving. However, verb

11 Aug 2026

An Agentic Generative Large Language Model for Treatment Planning of Colorectal Cancer

SafetyDGX agent

arXiv:2608.09142v1 Announce Type: new Abstract: Treatment planning in precision oncology requires synthesizing heterogeneous patient information with rapidly evolving clinical guidelines to ensure gui

Decoy Images Amplify Caption-Mediated Defenses Against Encoded Jailbreaks

SafetyDGX agent

arXiv:2608.01043v2 Announce Type: replace-cross Abstract: We report a counter-intuitive interaction between image inputs and existing black-box defenses on Vision--Language Models (VLMs): pairing an e

From Operational Design Domain to Action: A Systematic Behavioral Taxonomy for Autonomous Driving

SafetyDGX agent

arXiv:2608.08941v1 Announce Type: cross Abstract: Operational Design Domain (ODD) specifications describe where an automated driving system (ADS) is permitted to operate, but they do not prescribe wha

MedCalc-R1: Knowledge-Guided Reward Framework for Medical Mathematical Reasoning

SafetyDGX agent

arXiv:2608.08623v1 Announce Type: new Abstract: In Reinforcement Learning with Verifiable Rewards (RLVR) frameworks for mathematical reasoning tasks, floating-point results are typically evaluated usi

Multimodal Model Diffing for Feature Discovery and Control

SafetyDGX agent

arXiv:2608.09928v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) exhibit strong visual understanding, yet the internal features that cause these behaviors remain difficult to

PQC in Plaintext: Google Cloud’s post-quantum cryptography roadmap

SafetyDGX agent

Securing infrastructure and services against a future cryptographically-relevant quantum computer has been a goal for Google for a decade, and we’ve dedicated ourselves to help developers by advancing

What Keeps Agent Skills from Being Reusable? Evidence from 138K SKILL.md Files

SafetyDGX agent

arXiv:2608.08453v1 Announce Type: new Abstract: Under the current standard, Agent Skills are SKILL.md files that combine instructions with supporting resources, enabling Large Language Model (LLM) age

10 Aug 2026

bioMoR: Biology-Guided Mixture-of-Recursions for Effective Genomic Learning

SafetyDGX agent

arXiv:2608.06727v1 Announce Type: new Abstract: Transformer models for high-dimensional omics analysis process thousands of genes or pathways, although only a subset requires deep computation. Mixture

From Cheap Fakes to Pure Synthesis: Addressing the New Era of T2V Fake News Videos

SafetyDGX agent

arXiv:2608.06732v1 Announce Type: new Abstract: Recent text-to-video (T2V) generation models enable fake news videos to be synthesized from scratch, shifting the threat beyond cheap fakes assembled fr

SoRoMoX: Fast, Differentiable, and Parallelizable Soft Robot Models

SafetyDGX agent

arXiv:2608.06650v1 Announce Type: cross Abstract: Reduced-order models based on Cosserat-rod theory are now well established, and modeling theory is no longer the primary bottleneck in soft-robot cont

7 Aug 2026

A System for Train Condition Monitoring and Structural Health Assessment of Rail Vehicles

SafetyDGX agent

arXiv:2608.05221v1 Announce Type: new Abstract: The ongoing digitalization of rail systems and the increasing use of artificial intelligence (AI) are fundamentally transforming the design, operation,

DreamGuard: Efficient Runtime Guardrail for LLM Agents via Risk-Aware World Model

SafetyDGX agent

arXiv:2608.05695v1 Announce Type: new Abstract: As large language model (LLM) agents increasingly invoke external tools and interact with real-world systems, unsafe actions may cause irreversible cons

6 Aug 2026

Looking in the Mirror: Introspecting Side-Effect Misalignments Induced by Fine-Tuning

SafetyDGX agent

arXiv:2608.04347v1 Announce Type: new Abstract: Fine-tuning enables a source model to acquire desired capabilities and behaviors in a target domain while retaining much of its general-purpose competen

5 Aug 2026

A Blind Spot in Alignment: Quantifying Biosecurity Risks in Large Language Models

SafetyDGX agent

arXiv:2608.02684v1 Announce Type: cross Abstract: Large Language Models (LLMs) are accelerating biological research, yet this same capability poses a critical biosecurity threat: models that assist in

LatentGuard: Efficient and Inspectable Latent Reasoning for LLM Safeguards

SafetyDGX agent

arXiv:2608.03838v1 Announce Type: new Abstract: Reasoning-based guard models improve LLM safeguards, but decoding explicit rationales for every interaction makes them costly to deploy. Although latent

Poisoning Prompt-Guided Sampling in Video Large Language Models

SafetyDGX agent

arXiv:2509.20851v2 Announce Type: replace Abstract: Video Large Language Models (VideoLLMs) are increasingly deployed as automated moderators on user-generated video platforms, where a few unwatched s

PULSE: An Executable Contract Language for Spatiotemporal Knowledge Graph Engineering

SafetyDGX agent

arXiv:2608.02630v1 Announce Type: new Abstract: Knowledge graph engineering often distributes accepted state, observations, constraints, processes, and hypothetical scenarios across artifacts whose co

4 Aug 2026

HapticVLA: Contact-Rich Manipulation via Vision-Language-Action Model without Inference-Time Tactile Sensing

SafetyDGX agent

arXiv:2603.15257v2 Announce Type: replace Abstract: Tactile sensing is a crucial capability for Vision-Language-Action (VLA) architectures, as it enables dexterous and safe manipulation in contact-ric

MedPRESS: A Multi-turn Benchmark for Patient-Pressure-Induced Medical Sycophancy in LLMs

Model ReleasesDGX agent

arXiv:2608.02520v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for health-related advice. Existing research measures their safety with static questions rather than

Quality-Diversity Red-Teaming: Automated Generation of High-Quality and Diverse Attackers for Large Language Models

Model ReleasesDGX agent

arXiv:2506.07121v2 Announce Type: replace Abstract: Ensuring the safety and robustness of large language models (LLMs) is a fundamental challenge and a critical prerequisite for the responsible deploy

Self-Improving Large Language Models via Progressive Experience Evolution

SafetyDGX agent

arXiv:2608.02139v1 Announce Type: new Abstract: Large language models (LLMs) capable of self-improvement require not only effective policy optimization, but also a principled mechanism for transformin

When May a Model Replace the Experiment? Audits, Licenses, and the Price of Trust in Surrogate-Driven Design

SafetyDGX agent

arXiv:2608.01378v1 Announce Type: new Abstract: Design campaigns in chemistry, materials science, and machine learning share a bottleneck: determining how good a candidate truly is requires an expensi

3 Aug 2026

CAGE: Certified Authorization under Typed-Return Uncertainty for Tool-Using Agents

SafetyDGX agent

arXiv:2607.29190v1 Announce Type: new Abstract: Tool-using LLM agents act on typed tool returns, records pairing provenance and categorical fields with numerical values. Runtime permission gates gener

31 Jul 2026

Active Lubrication of Transluminal Medical Instruments

SafetyDGX agent

arXiv:2506.07225v2 Announce Type: replace-cross Abstract: Transluminal minimally invasive surgery uses natural orifices and small incisions to access internal anatomical structures, promoting quicker

Ask don't tell: Reducing sycophancy in large language models

SafetyDGX agent

arXiv:2602.23971v4 Announce Type: replace-cross Abstract: Sycophancy, the tendency of large language models to favour user-affirming responses over critical engagement, has been identified as an align

Corrigible Assistance in One Round: Pragmatic-Pedagogic Best Response

SafetyDGX agent

arXiv:2607.27508v1 Announce Type: new Abstract: Assistance games formalize human-robot collaboration under asymmetric information: the human knows the goal, while the robot must infer it from observat

EmoFeedback^2: Reinforcement of Continuous Emotional Image Generation via LVLM-based Reward and Textual Feedback

SafetyDGX agent

arXiv:2511.19982v3 Announce Type: replace Abstract: Continuous emotional image content generation (C-EICG) is emerging rapidly due to its ability to produce images aligned with both user descriptions

30 Jul 2026

FleetScape: A Mixed Reality Sandtable for Spatial Supervision and Control of Scalable Drone Fleets

SafetyDGX agent

arXiv:2607.26423v1 Announce Type: cross Abstract: As autonomous drone deployments scale from individual units to coordinated swarms, the human operator's role shifts from direct piloting to high-level

Prosody-driven Jailbreaks in Audio LLMs: A Controlled Study and Mechanistic Analysis

Model ReleasesDGX agent

arXiv:2607.26541v1 Announce Type: cross Abstract: Audio-capable foundation models enable end-to-end spoken interaction, but they also introduce safety risks beyond transcript content. It remains uncle

29 Jul 2026

Agentic AI in medicine: architectures, applications, evaluation, and challenges for clinical translation

SafetyDGX agent

arXiv:2607.25489v1 Announce Type: new Abstract: Large language models and multimodal foundation models are enabling medical artificial intelligence (AI) systems to move beyond isolated prediction and

Inspect India Evals: An Open Benchmarking Framework for Evaluating Large Language Models in the Indian Linguistic and Cultural Context

Model ReleasesDGX agent

arXiv:2607.25375v1 Announce Type: new Abstract: India is a vast nation of over 1.4 billion people, varied by hundreds of diverse and locally specific traditions and cultures and 22 officially recogniz

Shieldstral

Model ReleasesDGX agent

arXiv:2607.25857v1 Announce Type: new Abstract: We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7imes its size on text s

SONG: A Photorealistic 3D Gaussian Simulation Platform for Benchmarking Social Navigation

SafetyDGX agent

arXiv:2607.25219v1 Announce Type: new Abstract: Social navigation has progressed from simplified 2D environments toward a more general vision-based setting, in which a robot needs to achieve socially

28 Jul 2026

A Computational Ethical Framework for Financial Digital Phenotyping for Mental Health

SafetyDGX agent

arXiv:2607.24275v1 Announce Type: cross Abstract: Ethical governance of AI-driven systems is often expressed through high-level principles and static documentation, creating a gap between regulatory r

Adapter Merging Reactivates Latent Reasoning Traces: A Mechanism Analysis

SafetyDGX agent

arXiv:2601.18350v5 Announce Type: replace-cross Abstract: Large language models fine-tuned via a two-stage pipeline (domain adaptation followed by instruction alignment) can exhibit non-trivial interf

ARdena: Scenario-driven control of real-time LLM agents

SafetyDGX agent

arXiv:2607.22651v1 Announce Type: new Abstract: Large language models (LLMs) have enabled increasingly capable conversational agents, but reliably controlling their behavior in real-time interactive e

Frustratingly Simple Black-Box Adaptation of Language Models via Logit Bias

SafetyDGX agent

arXiv:2607.22837v1 Announce Type: cross Abstract: Many organizations aim to adapt language models for internal use, both to improve performance on domain-specific tasks and to address privacy concerns

How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift

SafetyDGX agent

arXiv:2607.22676v1 Announce Type: new Abstract: Post-training is a key mechanism for adapting large language models to downstream tasks. While prior work suggests that task adaptation can alter a mode

IJCB-AFMFR 2026: Competition on Adapting Foundation Models for Face Recognition Using Synthetic Training Data

SafetyDGX agent

arXiv:2607.24422v1 Announce Type: new Abstract: This paper presents a summary of the Competition on Adapting Foundation Models for Face Recognition Using Synthetic Training Data (AFMFR), held at the 2

Position: AI/ML Deepfake Research is Misaligned with AI-Generated Non-Consensual Intimate Imagery (AIG-NCII)

SafetyDGX agent

arXiv:2607.18263v2 Announce Type: replace Abstract: AI-generated non-consensual intimate imagery (AIG-NCII) is not adequately addressed in AI/ML literature regarding AI-generated media, commonly refer

RM-Distiller: Exploiting Generative LLM for Reward Model Distillation

SafetyDGX agent

arXiv:2601.14032v2 Announce Type: replace Abstract: Reward models (RMs) play a pivotal role in aligning large language models (LLMs) with human preferences. Due to the difficulty of obtaining high-qua

Self-Authored Verification Is Unreliable in Heuristic Self-Improving Agents

SafetyDGX agent

arXiv:2607.24300v1 Announce Type: new Abstract: Self-improving agents accumulate capability by repeatedly rewriting procedural policies, controllers, or heuristic rules. They typically rely on self-au

STAIF: A Stage-wise Optimization for Complex Instruction Following

SafetyDGX agent

arXiv:2607.22649v1 Announce Type: new Abstract: Following complex instructions with multiple explicit constraints remains a fundamental challenge for large language models (LLMs). Existing alignment m

Tailored untruths: How personalisation challenges LLM safeguards

SafetyDGX agent

arXiv:2510.12993v3 Announce Type: replace Abstract: Large Language Models (LLMs) can generate highly persuasive disinformation, yet little is known about how effectively they personalise it across lan

Unequal Trips, Unequal Places: Diagnosing and Mitigating Delay Inequity in Autonomous Vehicle Fleet Coordination

SafetyDGX agent

arXiv:2607.24336v1 Announce Type: new Abstract: City-scale autonomous vehicle fleet coordinators are typically optimized for aggregate travel time, yet fleet averages conceal how delay is distributed

White-hat hacking IS the defense to black-hat hacking. The techniques are the same. How does Dario expect companies to do it if their models refuse?

SafetyDGX agent

You patch security holes by intentionally finding them. If the models refuse to do it, how can companies protect themselves against rogue AIs, whether they are Chinese or OpenAI/Anthropic themselves?

Why Anthropic's battle is meant to poison the wells of open weight models, in 3 steps.

SafetyDGX agent

It doesn't solve any problems. Just a few paragraphs above, he says he fears that authoritarian states (he names China, and possibly others) can use their models to do evil stuff. And surely enough, m

25 Jul 2026

We recognize there are a lot of questions and speculative details circulating related to the Hugging Face incident. This is an unprecedented…

SafetyDGX agent

We recognize there are a lot of questions and speculative details circulating related to the Hugging Face incident. This is an unprecedented incident, and we think it marks an important moment for AI

24 Jul 2026

Interaction Dynamics Modeling and Predictive Control for Safe Steerable Catheter--Tissue Interaction

SafetyDGX agent

arXiv:2607.20939v1 Announce Type: cross Abstract: Safe steerable catheter control is fundamentally a problem of interaction dynamics: the tip must follow a planned motion, remain compliant against mov

OpenAI and Anthropic are now competing to see who can make the scariest AI for marketing purposes. What could possibly go wrong?

SafetyDGX agent

Pedro Domingos noted that OpenAI and Anthropic are competing to develop the scariest AI for marketing, raising concerns about the potential misuse of persuasive machine‑learning models. His comment hi

Robust Critics: Defending LLMs Against Multi-Turn Attacks

SafetyDGX agent

arXiv:2607.20472v1 Announce Type: new Abstract: When a user asks a language model something harmful, is it a genuine attack or a misunderstood but well-meaning question? This ambiguity is one of the c

Safe and Scalable Multi-Drone Payload Transport via CBF-based Reinforcement Learning with Zero-Shot Sim-to-Real Transfer

SafetyDGX agent

arXiv:2607.20665v1 Announce Type: new Abstract: Multi-drone payload transportation has emerged as a promising research paradigm with potential applications in construction, logistics, and disaster res

23 Jul 2026

Formal Foundations for Known Good Reliable Die Screening in Chiplet-Based AI Systems-on-Chip

SafetyDGX agent

arXiv:2607.20141v1 Announce Type: cross Abstract: The rapid growth of chiplet-based artificial intelligence systems-on-chip (SoCs) has exposed a fundamental gap in semiconductor test methodology. Exis

No Training, Better Flights: Test-Time Scaled VLMs for UAV Navigation

SafetyDGX agent

arXiv:2607.19288v1 Announce Type: new Abstract: Test-time scaling offers a promising method to improve the inference performance of Vision-Language Models (VLMs) without additional training. Existing

OPIUM: Mitigating Steering Externalities and Over-Refusal via Dual Objective Latent Optimization

SafetyDGX agent

arXiv:2607.19806v1 Announce Type: cross Abstract: Activation steering provides a lightweight mechanism for controlling large language models at inference time, but steering vectors can have unintended

20 Jul 2026

SSI should open-source an opus-tier model. 1. it nearly closes the us/china gap on the os frontier 2. they don't want to waste compute alloc…

SafetyDGX agent

SSI should open-source an opus-tier model. 1. it nearly closes the us/china gap on the os frontier 2. they don't want to waste compute allocation on inference, releasing os will enable continued full

← Previous
1…2627282930…240
Next →