AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,351 results
19 May 2026

State Contamination in Memory-Augmented LLM Agents

SafetyDGX agent

arXiv:2605.16746v1 Announce Type: new Abstract: LLM agents increasingly rely on persistent state, including transcripts, summaries, retrieved context, and memory buffers, to support long-horizon inter

15 May 2026

EVA: Editing for Versatile Alignment against Jailbreaks

Local AiDGX agent

arXiv:2605.14750v1 Announce Type: cross Abstract: Large Language Models (LLMs) and Vision Language Models (VLMs) have demonstrated impressive capabilities but remain vulnerable to jailbreaking attacks

Synthesizing POMDP Policies: Sampling Meets Model-checking via Learning

SafetyDGX agent

arXiv:2605.14440v1 Announce Type: new Abstract: Partially Observable Markov Decision Processes (POMDPs) are the standard framework for decision-making under uncertainty. While sampling-based methods s

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
14 May 2026

AgenticAITA: A Proof-Of-Concept About Deliberative Multi-Agent Reasoning for Autonomous Trading Systems

SafetyDGX agent

arXiv:2605.12532v1 Announce Type: cross Abstract: Conventional algorithmic trading systems are grounded in deterministic heuristics or offline-trained statistical models that cannot adapt to the seman

Automated alignment is harder than you think

SafetyDGX agent

arXiv:2605.06390v2 Announce Type: replace Abstract: A leading proposal for aligning artificial superintelligence (ASI) is to use AI agents to automate an increasing fraction of alignment research as c

Beyond VMAF: Towards Application-Specific Metrics for Teleoperation Video

SafetyDGX agent

arXiv:2605.13525v1 Announce Type: cross Abstract: Automated driving has made remarkable progress, yet situations still arise where human intervention is necessary. Teleoperation provides a scalable so

Certified Robustness under Heterogeneous Perturbations via Hybrid Randomized Smoothing

SafetyDGX agent

arXiv:2605.12876v1 Announce Type: new Abstract: Randomized smoothing provides strong, model-agnostic robustness certificates, but existing guarantees are limited to single modalities, treating continu

History Anchors: How Prior Behavior Steers LLM Decisions Toward Unsafe Actions

SafetyDGX agent

arXiv:2605.13825v1 Announce Type: new Abstract: Frontier LLMs are increasingly deployed as agents that pick the next action after a long log of prior tool calls produced by the same or a different mod

Learning When to Act: Communication-Efficient Reinforcement Learning via Run-Time Assurance

SafetyDGX agent

arXiv:2605.12561v1 Announce Type: new Abstract: Safe reinforcement learning (RL) typically asks extit{what} an agent should do. We ask extit{when} it needs to act, and show that a single policy can jo

13 May 2026

Certified Gradient-Based Contact-Rich Manipulation via Smoothing-Error Reachable Tubes

SafetyDGX agent

arXiv:2602.09368v2 Announce Type: replace Abstract: Gradient-based methods can efficiently optimize controllers by leveraging differentiable simulation and physical priors. However, contact-rich manip

Generative AI for Visualizing Highway Construction Hazards Through Synthetic Images and Temporal Sequences

SafetyDGX agent

arXiv:2605.11276v1 Announce Type: new Abstract: Highway construction workers face a high risk of serious injury or death. Image-based training materials depicting hazardous scenarios are essential for

Robust Policy Optimization to Prevent Catastrophic Forgetting

SafetyDGX agent

arXiv:2602.08813v2 Announce Type: replace Abstract: Large language models are commonly trained through multi-stage post-training: first via RLHF, then fine-tuned for other downstream objectives. Yet e

Transformer-Based Autonomous Driving Models and Deployment-Oriented Compression: A Survey

SafetyDGX agent

arXiv:2304.10891v2 Announce Type: replace-cross Abstract: Transformer-based models are becoming a central paradigm in autonomous driving because they can capture long-range spatial dependencies, multi

12 May 2026

Causal Explanations from the Geometric Properties of ReLU Neural Networks

SafetyDGX agent

arXiv:2605.10396v1 Announce Type: new Abstract: Neural networks have proved an effective means of learning control policies for autonomous systems, but these learned policies are difficult to understa

Code Mixologist : A Practitioner's Guide to Building Code-Mixed LLMs

SafetyDGX agent

arXiv:2602.11181v2 Announce Type: replace Abstract: Code-mixing and code-switching (CSW) remain challenging phenomena for large language models (LLMs). Despite recent advances in multilingual modeling

Constraint-Aware Diffusion Priors for High-Fidelity and Versatile Quadruped Locomotion

SafetyDGX agent

arXiv:2605.08804v1 Announce Type: new Abstract: Reinforcement learning combined with imitation learning has significantly advanced biomimetic quadrupedal locomotion. However, scaling these frameworks

EROAS: 3D Efficient Reactive Obstacle Avoidance System for Autonomous Underwater Vehicles using 2.5D Forward-Looking Sonar

SafetyDGX agent

arXiv:2411.05516v3 Announce Type: replace Abstract: Autonomous Underwater Vehicles (AUVs) have advanced significantly in obstacle detection and path planning through sonar, cameras, and learning-based

Hierarchical End-to-End Taylor Bounds for Complete Neural Network Verification

SafetyDGX agent

arXiv:2605.10621v1 Announce Type: new Abstract: Reachability analysis of neural networks, which seeks to compute or bound the set of outputs attainable over a given input domain, is central to certify

Mechanism Design Is Not Enough: Prosocial Agents for Cooperative AI

SafetyDGX agent

arXiv:2605.08426v1 Announce Type: cross Abstract: Ensuring that AI agents behave safely and beneficially when interacting with other parties has emerged as one of the central challenges of modern AI s

NEXUS: Continual Learning of Symbolic Constraints for Safe and Robust Embodied Planning

SafetyDGX agent

arXiv:2605.09387v1 Announce Type: new Abstract: While Large Language Models (LLMs) have catalyzed progress in embodied intelligence, a fundamental gap between their inherent probabilistic uncertainty

Reinforcement Learning for Scalable and Trustworthy Intelligent Systems

SafetyDGX agent

arXiv:2605.08378v1 Announce Type: cross Abstract: Reinforcement learning has become a powerful paradigm for improving the capability of intelligent systems, but its practical deployment faces two cent

RUBEN: Rule-Based Explanations for Retrieval-Augmented LLM Systems

SafetyDGX agent

arXiv:2605.10862v1 Announce Type: new Abstract: This paper demonstrates RUBEN, an interactive tool for discovering minimal rules to explain the outputs of retrieval-augmented large language models (LL

Verification Mirage: Mapping the Reliability Boundary of Self-Verification in Medical VQA

SafetyDGX agent

arXiv:2605.10850v1 Announce Type: new Abstract: Self-verification, re-invoking the same vision language model (VLM) in a fresh context to check its own generated answer, is increasingly used as a defa

11 May 2026

GLiGuard: Schema-Conditioned Classification for LLM Safeguard

Model ReleasesDGX agent

arXiv:2605.07982v1 Announce Type: new Abstract: Ensuring safe, policy-compliant outputs from large language models requires real-time content moderation that can scale across multiple safety dimension

GPO-V: Jailbreak Diffusion Vision Language Model by Global Probability Optimization

SafetyDGX agent

arXiv:2605.07399v1 Announce Type: new Abstract: Diffusion Vision-Language Models (dVLMs), built upon the non-causal foundations of Diffusion Large Language Models (dLLMs), have demonstrated remarkable

MORPH-U: Multi-Objective Resilient Motion Planning for V2X-Enabled Autonomous Driving in High-Uncertainty Environments via Simulation

SafetyDGX agent

arXiv:2605.07370v1 Announce Type: cross Abstract: V2X can warn an autonomous vehicle about hazards beyond line-of-sight, but it also brings uncertainty: messages may be delayed, dropped, or even forge

Operating Within the Operational Design Domain: Zero-Shot Perception with Vision-Language Models

SafetyDGX agent

arXiv:2605.07649v1 Announce Type: cross Abstract: Over the last few years, research on autonomous systems has matured to such a degree that the field is increasingly well-positioned to translate resea

OrchJail: Jailbreaking Tool-Calling Text-to-Image Agents by Orchestration-Guided Fuzzing

SafetyDGX agent

arXiv:2605.07414v1 Announce Type: cross Abstract: Tool-calling text-to-image (T2I) agents can plan and execute multi-step tool chains to accomplish complex generation and editing queries. However, thi

Sparse Autoencoders as Plug-and-Play Firewalls for Adversarial Attack Detection in VLMs

SafetyDGX agent

arXiv:2605.07447v1 Announce Type: cross Abstract: Vision-language models (VLMs) have advanced rapidly and are increasingly deployed in real-world applications, especially with the rise of agent-based

7 May 2026

I really enjoyed chatting with @mattturck, was a great discussion.

SafetyDGX agent

I really enjoyed chatting with @mattturck, was a great discussion. Deeply thoughtful conversation with @zicokolter, board member at @OpenAI and head of the machine learning department at @CarnegieMell

LLM-Based Human-Agent Collaboration and Interaction Systems: A Survey

SafetyDGX agent

arXiv:2505.00753v5 Announce Type: replace Abstract: Recent advances in large language models (LLMs) have sparked growing interest in building fully autonomous agents. However, fully autonomous LLM-bas

SafeRedir: Prompt Embedding Redirection for Robust Unlearning in Image Generation Models

SafetyDGX agent

arXiv:2601.08623v2 Announce Type: replace Abstract: Image generation models (IGMs), while capable of producing impressive and creative content, often memorize a wide range of undesirable concepts from

Software Engineering for Self-Adaptive Robotics: A Research Agenda

SafetyDGX agent

arXiv:2505.19629v3 Announce Type: replace-cross Abstract: Self-adaptive robotic systems operate autonomously in dynamic and uncertain environments, requiring robust real-time monitoring and adaptive b

6 May 2026

Learning Reactive Dexterous Grasping via Hierarchical Task-Space RL Planning and Joint-Space QP Control

SafetyDGX agent

arXiv:2605.03363v1 Announce Type: new Abstract: In this work, we propose a hybrid hierarchical control framework for reactive dexterous grasping that explicitly decouples high-level spatial intent fro

MAGE: Safeguarding LLM Agents against Long-Horizon Threats via Shadow Memory

SafetyDGX agent

arXiv:2605.03228v1 Announce Type: cross Abstract: As large language model (LLM)-powered agents are increasingly deployed to perform complex, real-world tasks, they face a growing class of attacks that

5 May 2026

Analyzing Adversarial Inputs in Deep Reinforcement Learning

SafetyDGX agent

arXiv:2402.05284v2 Announce Type: replace Abstract: In recent years, Deep Reinforcement Learning (DRL) has become a popular paradigm in machine learning due to its successful applications to real-worl

Cut-In Gap Acceptance Toward Autonomous vs. Human-Driven Vehicles: Evidence from the Waymo Open Motion Dataset

SafetyDGX agent

arXiv:2605.01485v1 Announce Type: new Abstract: Autonomous vehicles (AVs) are widely known to follow conservative, rule-based motion policies that surrounding drivers can learn to anticipate. A direct

Reliability-Oriented Multilingual Orthopedic Diagnosis: A Domain-Adaptive Modeling and a Conceptual Validation Framework

SafetyDGX agent

arXiv:2605.02266v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly proposed for clinical decision support including multilingual diagnosis in low-resource settings. However,

SAGA: A Robust Self-Attention and Goal-Aware Anchor-based Planner for Safe UAV Autonomous Navigation

SafetyDGX agent

arXiv:2605.02301v1 Announce Type: new Abstract: Agile unmanned aerial vehicle (UAV) navigation in cluttered environments demands a planning architecture that is both computationally efficient and stru

Visibility-Aware Mobile Grasping in Dynamic Environments

SafetyDGX agent

arXiv:2605.02487v1 Announce Type: new Abstract: This paper addresses the problem of mobile grasping in dynamic, unknown environments where a robot must operate under a limited field-of-view. The funda

Zero-Shot, Safe and Time-Efficient UAV Navigation via Potential-Based Reward Shaping, Control Lyapunov and Barrier Functions

SafetyDGX agent

arXiv:2605.01787v1 Announce Type: cross Abstract: Autonomous navigation and obstacle avoidance remain a core challenge of modern Unmanned Aerial Vehicles (UAVs). While traditional control methods stru

4 May 2026

Jailbreaking Vision-Language Models Through the Visual Modality

Model ReleasesDGX agent

arXiv:2605.00583v1 Announce Type: new Abstract: The visual modality of vision-language models (VLMs) is an underexplored attack surface for bypassing safety alignment. We introduce four jailbreak atta

ReLay: Personalized LLM-Generated Plain-Language Summaries for Better Understanding, but at What Cost?

SafetyDGX agent

arXiv:2605.00468v1 Announce Type: new Abstract: Plain Language Summaries (PLS) aim to make research accessible to lay readers, but they are typically written in a one-size-fits-all style that ignores

3 May 2026

AI will create more jobs than any other technology in history. The doomers' fundamental error isn't just the lump of labor fallacy. It's dee…

SafetyDGX agent

AI will create more jobs than any other technology in history. The doomers' fundamental error isn't just the lump of labor fallacy. It's deeper than that. They assume a finite problem space. This is t

1 May 2026

Detecting Clinical Discrepancies in Health Coaching Agents: A Dual-Stream Memory and Reconciliation Architecture

SafetyDGX agent

arXiv:2604.27045v1 Announce Type: cross Abstract: As Large Language Model (LLM) agents transition from single-session tools to persistent systems managing longitudinal healthcare journeys, their memor

Mechanized Foundations of Structural Governance: Machine-Checked Proofs for Governed Intelligence

SafetyDGX agent

arXiv:2604.27289v1 Announce Type: new Abstract: We present five results in the theory of structural governance for cognitive workflow systems. Three are mechanized in Coq 8.19 using the Interaction Tr

Test Before You Deploy: Governing Updates in the LLM Supply Chain

SafetyDGX agent

arXiv:2604.27789v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used as core dependencies in software systems. However, the hosted LLM services evolve continuously thro

30 Apr 2026

A Scoping Review of LLM-as-a-Judge in Healthcare and the MedJUDGE Framework

SafetyDGX agent

arXiv:2604.25933v1 Announce Type: cross Abstract: As large language models (LLMs) increasingly generate and process clinical text, scalable evaluation has become critical. LLM-as-a-Judge (LaaJ), which

Rule-based High-Level Coaching for Goal-Conditioned Reinforcement Learning in Search-and-Rescue UAV Missions Under Limited-Simulation Training

SafetyDGX agent

arXiv:2604.26833v1 Announce Type: cross Abstract: This paper presents a hierarchical decision-making framework for unmanned aerial vehicle (UAV) missions motivated by search-and-rescue (SAR) scenarios

Walk With Me: Long-Horizon Social Navigation for Human-Centric Outdoor Assistance

SafetyDGX agent

arXiv:2604.26839v1 Announce Type: new Abstract: Assisting humans in open-world outdoor environments requires robots to translate high-level natural-language intentions into safe, long-horizon, and soc

29 Apr 2026

Generative AI Carries Non-Democratic Biases and Stereotypes: Representation of Women, Black Individuals, Age Groups, and People with Disability in AI-Generated Images across Occupations

SafetyDGX agent

arXiv:2409.13869v2 Announce Type: replace-cross Abstract: In this study, I investigate how generative artificial intelligence (AI) systems reproduce and reinforce societal biases, with a specific focu

Learning from Medical Entity Trees: An Entity-Centric Medical Data Engineering Framework for MLLMs

SafetyDGX agent

arXiv:2604.25296v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have shown transformative potential in medical applications, yet their performance is hindered by conventional

Sustained Gradient Alignment Mediates Subliminal Learning in a Multi-Step Setting: Evidence from MNIST Auxiliary Logit Distillation Experiment

SafetyDGX agent

arXiv:2604.25779v1 Announce Type: new Abstract: In the MNIST auxiliary logit distillation experiment, a student can acquire an unintended teacher trait despite distilling only on no-class logits throu

Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution

SafetyDGX agent

arXiv:2604.03472v2 Announce Type: replace Abstract: Co-evolutionary self-play, where one language model generates problems and another solves them, promises autonomous curriculum learning without huma

28 Apr 2026

Adaptive Multi-Subspace Representation Steering for Attribute Alignment in Large Language Models

SafetyDGX agent

arXiv:2508.10599v4 Announce Type: replace Abstract: Activation steering offers a promising approach to controlling the behavior of Large Language Models by directly manipulating their internal activat

AMAVA: Adaptive Motion-Aware Video-to-Audio Framework for Visually-Impaired Assistance

SafetyDGX agent

arXiv:2604.23909v1 Announce Type: new Abstract: Navigational aids for blind and low vision individuals struggle conveying dynamic real-world environments, leading to cognitive overload from continuous

An empirical evaluation of the risks of AI model updates using clinical data: stability, arbitrariness, and fairness

SafetyDGX agent

arXiv:2604.23954v1 Announce Type: new Abstract: Artificial Intelligence and Machine Learning (AI/ML) models used in clinical settings are increasingly deployed to support clinical decision-making. How

Analytica: Soft Propositional Reasoning for Robust and Scalable LLM-Driven Analysis

SafetyDGX agent

arXiv:2604.23072v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly tasked with complex real-world analysis (e.g., in financial forecasting, scientific discovery), yet t

ArgRE: Formal Argumentation for Conflict Resolution in Multi-Agent Requirements Negotiation

SafetyDGX agent

arXiv:2604.23124v1 Announce Type: cross Abstract: As software systems grow in complexity, they must satisfy an increasing number of competing quality attributes, making it essential to balance them in

Autocorrelation Reintroduces Spectral Bias in KANs for Time Series Forecasting

SafetyDGX agent

arXiv:2604.23518v1 Announce Type: cross Abstract: Existing theory suggests that Kolmogorov-Arnold Networks (KANs) can overcome the spectral bias commonly observed in neural networks under the assumpti

← Previous
1…2930313233…240
Next →