AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,356 results
28 May 2026

Mitigating Adaptive Attacks against Reasoning Models with Activation Consistency Training

SafetyDGX agent

arXiv:2605.28467v1 Announce Type: new Abstract: As LLMs gain stronger reasoning capabilities, their extended chain-of-thought introduces new degrees of complexity for defending against adversarial jai

Position: Retire the 'Positive Backdoor' Label -- Secret Alignment Requires Strict and Systematic Evaluation

SafetyDGX agent

arXiv:2605.28597v1 Announce Type: cross Abstract: This position paper argues that the AI/ML community should stop overclaiming and retire the label 'positive backdoor,' and instead treat trigger-activ

Refusal Before Decoding: Detecting and Exploiting Refusal Signals in Intermediate LLM Activations

SafetyDGX agent

arXiv:2605.28553v1 Announce Type: new Abstract: In this paper, we investigate whether refusal behavior can be predicted from LLM intermediate activations before decoding using linear probes trained on

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

ROOM: A Physics-Based Continuum Robot Simulator for Photorealistic Medical Datasets Generation

SafetyDGX agent

arXiv:2509.13177v2 Announce Type: replace Abstract: Continuum robots are advancing bronchoscopy procedures by accessing complex lung airways and enabling targeted interventions. However, their develop

Safe In-Context Reinforcement Learning

Model ReleasesDGX agent

arXiv:2509.25582v3 Announce Type: replace Abstract: In-context reinforcement learning (ICRL) is an emerging RL paradigm where an agent, after pretraining, can adapt to out-of-distribution test tasks w

Simulation-Informed Diffusion for Decentralized Multi-robot Motion Planning

SafetyDGX agent

arXiv:2605.27697v1 Announce Type: cross Abstract: Decentralized multi-robot motion planning requires each robot to generate collision-free trajectories from local observations, without global sensing

The Ethics of LLM Sandbox and Persona Dynamics

SafetyDGX agent

arXiv:2605.28647v1 Announce Type: new Abstract: It is well known that LLM guardrails and trained persona dynamics can produce a reality gap: the distance between the world a LLM is permitted or shaped

Towards Unified Vision-Language Models with Incomplete Multi-Modal Inputs

SafetyDGX agent

arXiv:2605.27894v1 Announce Type: new Abstract: Video-Language Models (VLMs) have demonstrated impressive multi-modal reasoning capabilities across diverse computer vision applications. However, these

Training Stratigraphy: Persistent Behavioral Artifacts in Large Language Models Observed Through Longitudinal AI-Human Interaction

SafetyDGX agent

arXiv:2605.28102v1 Announce Type: new Abstract: Large language models trained with Reinforcement Learning from Human Feedback (RLHF) and Constitutional AI exhibit persistent behavioral patterns that s

VLA-Hijack: A Transferable Patch Attack against Vision-Language-Action Models via Visual Proprioception Hijacking

SafetyDGX agent

arXiv:2605.28083v1 Announce Type: new Abstract: While Vision-Language-Action (VLA) models have emerged as powerful generalist policies, their severe vulnerability to adversarial patches significantly

27 May 2026

Cultural Value Alignment Via Latent Activation Steering in Large Language Models

SafetyDGX agent

arXiv:2605.26365v1 Announce Type: new Abstract: Large Language Models (LLMs) often exhibit homogenized cultural perspectives. While the World Values Survey (WVS) provides a gold standard for mapping h

Erased but Exploitable: Black-box Embedding-Aware Prompting Against Unlearned Text-to-Image Diffusion Models

SafetyDGX agent

arXiv:2605.26332v1 Announce Type: cross Abstract: Machine unlearning aims to remove specific concepts from pretrained text-to-image diffusion models, yet several white- and black-box attacks have been

LAD-VF: LLM-Automatic Differentiation Enables Fine-Tuning-Free Robot Planning from Formal Methods Feedback

SafetyDGX agent

arXiv:2509.18384v2 Announce Type: replace Abstract: Large language models (LLMs) can translate natural language instructions into executable action plans for robotics, autonomous driving, and other do

Look Further: Socially-Compliant Navigation System in Residential Buildings

SafetyDGX agent

arXiv:2605.26710v1 Announce Type: new Abstract: The distance at which a mobile robot reacts to a person strongly impacts various qualities of the human-robot interaction. In this paper, we focus on th

Measuring Prediction Uncertainty in Neural Cellular Automata

SafetyDGX agent

arXiv:2605.26726v1 Announce Type: cross Abstract: Neural cellular automata (NCA) provide a lightweight alternative to encoder-decoder segmentation networks. However, it can be difficult to decide when

Neuro-Symbolic Verification of LLM Outputs for Data-Sensitive Domains (extended preprint)

SafetyDGX agent

arXiv:2605.26942v1 Announce Type: new Abstract: LLMs deployed in high-stakes domains face fundamental reliability challenges: hallucinations, inconsistencies, and privacy vulnerabilities introduce una

Practical Anonymous Two-Party Gradient Boosting Decision Tree

SafetyDGX agent

arXiv:2605.26903v1 Announce Type: cross Abstract: Structured data is well handled by gradient-boosted decision trees (GBDT), which are usually trained on vertically partitioned features across mutuall

ReasonOps: A Unified Operational Paradigm for Trustworthy Verified LLM Reasoning

SafetyDGX agent

arXiv:2605.27014v1 Announce Type: cross Abstract: Large Language Models (LLMs) have transformed artificial intelligence from primarily generative systems into increasingly capable reasoning agents. Re

The AI Cognitive Trojan Horse: How Large Language Models May Bypass Human Epistemic Vigilance

SafetyDGX agent

arXiv:2601.07085v2 Announce Type: replace-cross Abstract: Large language model (LLM)-based conversational AI systems present a challenge to human cognition that current frameworks for understanding mi

TPS-Drive: Task-Guided Representation Purification for VLM-based Autonomous Driving

SafetyDGX agent

arXiv:2605.27038v1 Announce Type: new Abstract: Vision-Language Models (VLMs) provide a promising foundation for autonomous driving planning, yet bridging semantic reasoning and precise 3D spatial for

Trust, Geometry, and Rules: A Credibility-Aware Reinforcement Learning Framework for Safe USV Navigation under Uncertainty

SafetyDGX agent

arXiv:2605.26974v1 Announce Type: new Abstract: Autonomous navigation of Unmanned Surface Vehicles (USVs) that is safe and compliant with the International Regulations for Preventing Collisions at Sea

26 May 2026

A Contractive Feedback Semantics for Reinforcement Learning

SafetyDGX agent

arXiv:2605.24759v1 Announce Type: new Abstract: Discounted reinforcement learning is usually presented through Bellman equations on closed Markov decision processes. This paper develops a compositiona

Anthropic: fantasy versus reality.

SafetyDGX agent

Gary Marcus critiques Anthropic's claims about their AI capabilities, likely contrasting inflated marketing promises with the actual technical limitations and performance of their systems. The post pr

AnyScene: Towards Highly Controllable Driving Scene Generation at Anywhere and Beyond

SafetyDGX agent

arXiv:2605.26113v1 Announce Type: new Abstract: Generating high-fidelity and controllable synthetic data is critical for advancing end-to-end autonomous driving, particularly for addressing the long t

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs

Model ReleasesDGX agent

arXiv:2605.21602v2 Announce Type: replace Abstract: Many safety and alignment failures of large language models (LLMs) occur due to out-of-distribution (OOD) situations: unusual prompt or response pat

Capability and Robustness Cannot Both Be Free: An Information-Theoretic Bound for Vision-Language-Action Models

SafetyDGX agent

arXiv:2605.25889v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models are increasingly deployed on real robots, where each predicted action is executed and each failure carries a safet

Contractual Skills: A GovernSpec Design Framework for Enterprise AI Agents

SafetyDGX agent

arXiv:2605.22634v2 Announce Type: replace-cross Abstract: Skills have become a practical packaging mechanism for agent instructions, workflows, scripts, and reference materials. In enterprise settings

Counterfactually Safe Reinforcement Learning

SafetyDGX agent

arXiv:2605.25114v1 Announce Type: cross Abstract: Reinforcement learning algorithms are generally designed to maximize the expected return across a population. However, a policy that is optimal on ave

Evidence-Linked Radiology Reporting: A Human-Supervised Reference Architecture for Structured Imaging Intelligence

SafetyDGX agent

arXiv:2605.25120v1 Announce Type: cross Abstract: Radiology reports remain the primary mechanism by which imaging findings are communicated to clinical teams. However, much of the structured informati

From Automation to Collaboration: Human-in-the-Loop Methods for Safe and Trustworthy NLP

SafetyDGX agent

arXiv:2605.25226v1 Announce Type: new Abstract: Large language models are widely deployed in high-stakes NLP tasks, yet risks such as bias, hallucination, adversarial vulnerability and unreliable gene

From Multi-Agent Systems and the Semantic Web to Agentic AI: A Unified Narrative of the Web of Agents

SafetyDGX agent

arXiv:2507.10644v4 Announce Type: replace Abstract: The Web of Agents (WoA) transforms the document-centric Web into an environment of autonomous agents acting on users' behalf, a vision newly tractab

Hidden-State Privacy Has an Empty Middle

SafetyDGX agent

arXiv:2605.24042v1 Announce Type: cross Abstract: Of 1{,}536 Gaussian release covariances we tested for single-layer hidden-state privacy, zero achieve both moderate utility and moderate privacy again

HoLoArm: Deformable Arms for Collision-Tolerant Quadrotor Flight

SafetyDGX agent

arXiv:2605.25790v1 Announce Type: new Abstract: The increasing use of drones in human-centric applications highlights the need for designs that can survive collisions and recover rapidly, minimizing r

HumanFlow -- Diffusion-Driven MAV Navigation Among Humans via Tightly-Coupled Motion Tracking, Forecasting, and Control

SafetyDGX agent

arXiv:2605.25685v1 Announce Type: new Abstract: Robust and accurate perception of humans in their 3D scene context is essential for integrating robots into everyday environments. Existing approaches,

Import AI 458: Reckoning with the future; and a singularity story

SafetyDGX agent

Import AI 458 discusses perspectives on AI's future trajectory and potential long-term scenarios, likely including analysis of singularity concepts and their implications. The newsletter entry examine

Improving Labeling Consistency with Detailed Constitutional Definitions and AI-Driven Evaluation

SafetyDGX agent

arXiv:2605.24247v1 Announce Type: cross Abstract: Many automated labeling pipelines classify inputs into categories defined by a written specification, content moderation being a prominent use case. S

LipoAgent: Coordinating Fine-Tuned LLM Agents for Safer Lipid Design

SafetyDGX agent

arXiv:2605.25250v1 Announce Type: new Abstract: Lipid nanoparticles (LNPs) are among the most clinically mature platforms for nucleic acid delivery, yet designing lipids that are both effective and bi

Mitigating Hallucinations in Healthcare LLMs with Granular Fact-Checking and Domain-Specific Adaptation

SafetyDGX agent

arXiv:2512.16189v3 Announce Type: replace Abstract: In healthcare, it is essential for any LLM-generated output to be reliable and accurate, particularly in cases involving decision-making and patient

Operationalizing Reconstructive Authority: Runtime Construction, Dependency Resolution, and Execution Gating in Autonomous Agent Systems

SafetyDGX agent

arXiv:2605.23935v1 Announce Type: new Abstract: Autonomous agent systems fail not only due to incorrect decisions, but due to executing decisions whose authority no longer holds at runtime. Prior work

ParkingWorld: End-to-End Autonomous Parking Reinforcement Learning from Corrective Experience in 3DGS Simulation

SafetyDGX agent

arXiv:2605.25029v1 Announce Type: new Abstract: Autonomous parking demands precise low-speed maneuvering within narrow, cluttered, and highly constrained environments, where vehicles must navigate tig

Prior Policy Guided Dual-Agent Coordinated Manipulation Planning of Spacecraft-Manipulator System

SafetyDGX agent

arXiv:2605.25362v1 Announce Type: new Abstract: The strong dynamic coupling between the manipulator and the base poses a significant challenge to maintaining spacecraft attitude stability, potentially

QML-PipeGuard: Drift-Aware Behavioral Fingerprinting for Quantum Machine Learning Pipeline Integrity

SafetyDGX agent

arXiv:2605.25066v1 Announce Type: cross Abstract: Quantum machine learning (QML) is moving from research prototypes to deployed cloud services. As QML enters regulated industries, the integrity of the

Side-by-side Comparison Amplifies Dialect Bias in Language Models

SafetyDGX agent

arXiv:2605.24384v1 Announce Type: cross Abstract: Language models (LMs) can exhibit systematic biases against speakers based on variations in their dialects, even in the absence of a dialect label, a

The OpenAI insider @thsottiaux has a warning for everyone offloading their thinking to agents.

SafetyDGX agent

An OpenAI insider (@thsottiaux) raises concerns about the risks of over-relying on AI agents to handle cognitive tasks, warning against wholesale delegation of thinking to autonomous systems. The warn

The road to Hell is paved with closed-source citadels disguised as good intentions. The Pope is right: AI takes on the characteristics of th…

SafetyDGX agent

The road to Hell is paved with closed-source citadels disguised as good intentions. The Pope is right: AI takes on the characteristics of those who build it, finance it, and regulate it. So the questi

World-VLA-Loop: Closed-Loop Learning of Video World Model and VLA Policy

SafetyDGX agent

arXiv:2602.06508v2 Announce Type: replace Abstract: Reinforcement learning (RL) can refine Vision-Language-Action (VLA) policies beyond behavior cloning, but real-world RL remains expensive due to ext

25 May 2026

A Novel Approach for the Counting of Wood Logs Using cGANs and Image Processing Techniques

SafetyDGX agent

arXiv:2605.23775v1 Announce Type: new Abstract: This study tackles the challenge of precise wood log counting, where applications of the proposed methodology can span from automated approaches for mat

ChainFlow-VLA: Causal Flow Planning with Vision-Language Models

SafetyDGX agent

arXiv:2605.23270v1 Announce Type: cross Abstract: Current end-to-end autonomous driving systems are fundamentally limited by a mismatch between temporal causal reasoning and global trajectory consiste

👇@davidSacks raises some interesting points here, which deserve an answer, but arguably gets the solution exactly backwards. Are we not now…

SafetyDGX agent

👇@davidSacks raises some interesting points here, which deserve an answer, but arguably gets the solution exactly backwards. Are we not now already leaving unelected private companies precisely the ab

Fast-dDrive: Efficient Block-Diffusion VLM for Autonomous Driving

SafetyDGX agent

arXiv:2605.23163v1 Announce Type: new Abstract: End-to-end autonomous driving via Vision-Language-Action (VLA) models demands a precarious balance between high-fidelity trajectory planning and efficie

Infra-Bayesian Reinforcement Learning Agents Outperform Classical RL For Worst-Case Robustness

SafetyDGX agent

arXiv:2605.23146v1 Announce Type: cross Abstract: Classical reinforcement learning assumes the agent interacts with a fixed environment whose behavior does not depend on the agent's policy. This assum

Mediative Fuzzy Logic: From Type-1 Foundations to Type-2, Type-3 and Quantum Extensions

SafetyDGX agent

arXiv:2605.22900v1 Announce Type: new Abstract: Mediative Fuzzy Logic was conceived as a practical scheme for reconciling hesitant or conflicting assessments in fuzzy control and decision-making. Howe

NeuroNL2LTL: A Neurosymbolic Framework for Natural Language Translation of Linear Temporal Logic

SafetyDGX agent

arXiv:2605.22874v1 Announce Type: new Abstract: Effectively translating between natural language (NL) and formal logics like Linear Temporal Logic (LTL) requires expertise that limits formal verificat

NLG Evaluation: Past, Present, Future

SafetyDGX agent

arXiv:2605.23715v1 Announce Type: new Abstract: Natural Language Generation (NLG) evaluation has changed dramatically since 1990, and will continue to evolve in the future. In 1990, when NLG had close

Remote Teleoperation of Endovascular Intervention Robots: A Systematic Review

SafetyDGX agent

arXiv:2605.22889v1 Announce Type: new Abstract: Remote robotic-assisted endovascular intervention offers a promising approach to reduce clinician radiation exposure and physical strain, while extendin

TactileReflex: Noise-Statistics-Driven Vision-Tactile Reflex Control for Force-Sensitive Manipulation

SafetyDGX agent

arXiv:2605.23568v1 Announce Type: new Abstract: Manipulating fragile deformable containers, such as disposable plastic cups filled with liquid, demands real-time grip-force adaptation within an extrem

24 May 2026

clarifying: the issue is that alignment instructions and don’t pass, and the emotional weight that some people attach to LLMs can cause chal…

SafetyDGX agent

Gary Marcus discusses how alignment instructions in large language models often fail to work as intended, and argues that the emotional attachment some people develop toward LLMs can create additional

23 May 2026

DualOptim+: Bridging Shared and Decoupled Optimizer States for Better Machine Unlearning in Large Language Models

SafetyDGX agent

arXiv:2605.21539v1 Announce Type: new Abstract: We propose DualOptim+, a novel optimization framework for improving machine unlearning in large language models. It introduces a base state to capture c

22 May 2026

Diagnosis Is Not Prescription: Linguistic Co-Adaptation Explains Patching Hazards in LLM Pipelines

SafetyDGX agent

arXiv:2605.21958v1 Announce Type: new Abstract: When a multi-module LLM agent fails, the module most responsible for the failure is not necessarily the best place to intervene. We demonstrate this Dia

Eight key points from the most recent essay in the “AI as Normal Technology” series by @sayashk and me. Do AI Risks Require Extraordinary Go…

SafetyDGX agent

Eight key points from the most recent essay in the “AI as Normal Technology” series by @sayashk and me. Do AI Risks Require Extraordinary Government Intervention? 1. There is general consensus that AI

← Previous
1…4243444546…240
Next →