AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlog
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,351 results
Safety

GPO-V: Jailbreak Diffusion Vision Language Model by Global Probability Optimization

DGX agent

arXiv:2605.07399v1 Announce Type: new Abstract: Diffusion Vision-Language Models (dVLMs), built upon the non-causal foundations of Diffusion Large Language Models (dLLMs), have demonstrated remarkable

safetyarxiv-cs-cv
11 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Safety

MORPH-U: Multi-Objective Resilient Motion Planning for V2X-Enabled Autonomous Driving in High-Uncertainty Environments via Simulation

DGX agent

arXiv:2605.07370v1 Announce Type: cross Abstract: V2X can warn an autonomous vehicle about hazards beyond line-of-sight, but it also brings uncertainty: messages may be delayed, dropped, or even forge

safetyarxiv-cs-ai
11 May 2026
Safety

Operating Within the Operational Design Domain: Zero-Shot Perception with Vision-Language Models

DGX agent

arXiv:2605.07649v1 Announce Type: cross Abstract: Over the last few years, research on autonomous systems has matured to such a degree that the field is increasingly well-positioned to translate resea

safetyarxiv-cs-ai
11 May 2026
Safety

OrchJail: Jailbreaking Tool-Calling Text-to-Image Agents by Orchestration-Guided Fuzzing

DGX agent

arXiv:2605.07414v1 Announce Type: cross Abstract: Tool-calling text-to-image (T2I) agents can plan and execute multi-step tool chains to accomplish complex generation and editing queries. However, thi

safetyarxiv-cs-ai
11 May 2026
Safety

Sparse Autoencoders as Plug-and-Play Firewalls for Adversarial Attack Detection in VLMs

DGX agent

arXiv:2605.07447v1 Announce Type: cross Abstract: Vision-language models (VLMs) have advanced rapidly and are increasingly deployed in real-world applications, especially with the rise of agent-based

safetyarxiv-cs-ai
11 May 2026
Safety

I really enjoyed chatting with @mattturck, was a great discussion.

DGX agent

I really enjoyed chatting with @mattturck, was a great discussion. Deeply thoughtful conversation with @zicokolter, board member at @OpenAI and head of the machine learning department at @CarnegieMell

safetyjeremy-howard--x
7 May 2026
Safety

LLM-Based Human-Agent Collaboration and Interaction Systems: A Survey

DGX agent

arXiv:2505.00753v5 Announce Type: replace Abstract: Recent advances in large language models (LLMs) have sparked growing interest in building fully autonomous agents. However, fully autonomous LLM-bas

safetyarxiv-cs-cl
7 May 2026
Safety

SafeRedir: Prompt Embedding Redirection for Robust Unlearning in Image Generation Models

DGX agent

arXiv:2601.08623v2 Announce Type: replace Abstract: Image generation models (IGMs), while capable of producing impressive and creative content, often memorize a wide range of undesirable concepts from

safetyarxiv-cs-cv
7 May 2026
Safety

Software Engineering for Self-Adaptive Robotics: A Research Agenda

DGX agent

arXiv:2505.19629v3 Announce Type: replace-cross Abstract: Self-adaptive robotic systems operate autonomously in dynamic and uncertain environments, requiring robust real-time monitoring and adaptive b

safetyarxiv-cs-ro
7 May 2026
Safety

Learning Reactive Dexterous Grasping via Hierarchical Task-Space RL Planning and Joint-Space QP Control

DGX agent

arXiv:2605.03363v1 Announce Type: new Abstract: In this work, we propose a hybrid hierarchical control framework for reactive dexterous grasping that explicitly decouples high-level spatial intent fro

safetyarxiv-cs-ro
6 May 2026
Safety

MAGE: Safeguarding LLM Agents against Long-Horizon Threats via Shadow Memory

DGX agent

arXiv:2605.03228v1 Announce Type: cross Abstract: As large language model (LLM)-powered agents are increasingly deployed to perform complex, real-world tasks, they face a growing class of attacks that

safetyarxiv-cs-cl
6 May 2026
Safety

Analyzing Adversarial Inputs in Deep Reinforcement Learning

DGX agent

arXiv:2402.05284v2 Announce Type: replace Abstract: In recent years, Deep Reinforcement Learning (DRL) has become a popular paradigm in machine learning due to its successful applications to real-worl

safetyarxiv-cs-lg
5 May 2026
Safety

Cut-In Gap Acceptance Toward Autonomous vs. Human-Driven Vehicles: Evidence from the Waymo Open Motion Dataset

DGX agent

arXiv:2605.01485v1 Announce Type: new Abstract: Autonomous vehicles (AVs) are widely known to follow conservative, rule-based motion policies that surrounding drivers can learn to anticipate. A direct

safetyarxiv-cs-ro
5 May 2026
Safety

Reliability-Oriented Multilingual Orthopedic Diagnosis: A Domain-Adaptive Modeling and a Conceptual Validation Framework

DGX agent

arXiv:2605.02266v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly proposed for clinical decision support including multilingual diagnosis in low-resource settings. However,

safetyarxiv-cs-cl
5 May 2026
Safety

SAGA: A Robust Self-Attention and Goal-Aware Anchor-based Planner for Safe UAV Autonomous Navigation

DGX agent

arXiv:2605.02301v1 Announce Type: new Abstract: Agile unmanned aerial vehicle (UAV) navigation in cluttered environments demands a planning architecture that is both computationally efficient and stru

safetyarxiv-cs-ro
5 May 2026
Safety

Visibility-Aware Mobile Grasping in Dynamic Environments

DGX agent

arXiv:2605.02487v1 Announce Type: new Abstract: This paper addresses the problem of mobile grasping in dynamic, unknown environments where a robot must operate under a limited field-of-view. The funda

safetyarxiv-cs-ro
5 May 2026
Safety

Zero-Shot, Safe and Time-Efficient UAV Navigation via Potential-Based Reward Shaping, Control Lyapunov and Barrier Functions

DGX agent

arXiv:2605.01787v1 Announce Type: cross Abstract: Autonomous navigation and obstacle avoidance remain a core challenge of modern Unmanned Aerial Vehicles (UAVs). While traditional control methods stru

safetyarxiv-cs-lg
5 May 2026
Model Releases

Jailbreaking Vision-Language Models Through the Visual Modality

DGX agent

arXiv:2605.00583v1 Announce Type: new Abstract: The visual modality of vision-language models (VLMs) is an underexplored attack surface for bypassing safety alignment. We introduce four jailbreak atta

model-releasesarxiv-cs-cv
4 May 2026
Safety

ReLay: Personalized LLM-Generated Plain-Language Summaries for Better Understanding, but at What Cost?

DGX agent

arXiv:2605.00468v1 Announce Type: new Abstract: Plain Language Summaries (PLS) aim to make research accessible to lay readers, but they are typically written in a one-size-fits-all style that ignores

safetyarxiv-cs-cl
4 May 2026
Safety

AI will create more jobs than any other technology in history. The doomers' fundamental error isn't just the lump of labor fallacy. It's dee…

DGX agent

AI will create more jobs than any other technology in history. The doomers' fundamental error isn't just the lump of labor fallacy. It's deeper than that. They assume a finite problem space. This is t

safetyyann-lecun--x
3 May 2026
Safety

Detecting Clinical Discrepancies in Health Coaching Agents: A Dual-Stream Memory and Reconciliation Architecture

DGX agent

arXiv:2604.27045v1 Announce Type: cross Abstract: As Large Language Model (LLM) agents transition from single-session tools to persistent systems managing longitudinal healthcare journeys, their memor

safetyarxiv-cs-ai
1 May 2026
Safety

Mechanized Foundations of Structural Governance: Machine-Checked Proofs for Governed Intelligence

DGX agent

arXiv:2604.27289v1 Announce Type: new Abstract: We present five results in the theory of structural governance for cognitive workflow systems. Three are mechanized in Coq 8.19 using the Interaction Tr

safetyarxiv-cs-ai
1 May 2026
Safety

Test Before You Deploy: Governing Updates in the LLM Supply Chain

DGX agent

arXiv:2604.27789v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used as core dependencies in software systems. However, the hosted LLM services evolve continuously thro

safetyarxiv-cs-ai
1 May 2026
Safety

A Scoping Review of LLM-as-a-Judge in Healthcare and the MedJUDGE Framework

DGX agent

arXiv:2604.25933v1 Announce Type: cross Abstract: As large language models (LLMs) increasingly generate and process clinical text, scalable evaluation has become critical. LLM-as-a-Judge (LaaJ), which

safetyarxiv-cs-ai
30 Apr 2026
Safety

Rule-based High-Level Coaching for Goal-Conditioned Reinforcement Learning in Search-and-Rescue UAV Missions Under Limited-Simulation Training

DGX agent

arXiv:2604.26833v1 Announce Type: cross Abstract: This paper presents a hierarchical decision-making framework for unmanned aerial vehicle (UAV) missions motivated by search-and-rescue (SAR) scenarios

safetyarxiv-cs-ai
30 Apr 2026
Safety

Walk With Me: Long-Horizon Social Navigation for Human-Centric Outdoor Assistance

DGX agent

arXiv:2604.26839v1 Announce Type: new Abstract: Assisting humans in open-world outdoor environments requires robots to translate high-level natural-language intentions into safe, long-horizon, and soc

safetyarxiv-cs-ro
30 Apr 2026
Safety

Generative AI Carries Non-Democratic Biases and Stereotypes: Representation of Women, Black Individuals, Age Groups, and People with Disability in AI-Generated Images across Occupations

DGX agent

arXiv:2409.13869v2 Announce Type: replace-cross Abstract: In this study, I investigate how generative artificial intelligence (AI) systems reproduce and reinforce societal biases, with a specific focu

safetyarxiv-cs-cl
29 Apr 2026
Safety

Learning from Medical Entity Trees: An Entity-Centric Medical Data Engineering Framework for MLLMs

DGX agent

arXiv:2604.25296v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have shown transformative potential in medical applications, yet their performance is hindered by conventional

safetyarxiv-cs-cl
29 Apr 2026
Safety

Sustained Gradient Alignment Mediates Subliminal Learning in a Multi-Step Setting: Evidence from MNIST Auxiliary Logit Distillation Experiment

DGX agent

arXiv:2604.25779v1 Announce Type: new Abstract: In the MNIST auxiliary logit distillation experiment, a student can acquire an unintended teacher trait despite distilling only on no-class logits throu

safetyarxiv-cs-lg
29 Apr 2026
Safety

Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution

DGX agent

arXiv:2604.03472v2 Announce Type: replace Abstract: Co-evolutionary self-play, where one language model generates problems and another solves them, promises autonomous curriculum learning without huma

safetyarxiv-cs-cl
29 Apr 2026
Safety

Adaptive Multi-Subspace Representation Steering for Attribute Alignment in Large Language Models

DGX agent

arXiv:2508.10599v4 Announce Type: replace Abstract: Activation steering offers a promising approach to controlling the behavior of Large Language Models by directly manipulating their internal activat

safetyarxiv-cs-ai
28 Apr 2026
Safety

AMAVA: Adaptive Motion-Aware Video-to-Audio Framework for Visually-Impaired Assistance

DGX agent

arXiv:2604.23909v1 Announce Type: new Abstract: Navigational aids for blind and low vision individuals struggle conveying dynamic real-world environments, leading to cognitive overload from continuous

safetyarxiv-cs-cv
28 Apr 2026
Safety

An empirical evaluation of the risks of AI model updates using clinical data: stability, arbitrariness, and fairness

DGX agent

arXiv:2604.23954v1 Announce Type: new Abstract: Artificial Intelligence and Machine Learning (AI/ML) models used in clinical settings are increasingly deployed to support clinical decision-making. How

safetyarxiv-cs-ai
28 Apr 2026
Safety

Analytica: Soft Propositional Reasoning for Robust and Scalable LLM-Driven Analysis

DGX agent

arXiv:2604.23072v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly tasked with complex real-world analysis (e.g., in financial forecasting, scientific discovery), yet t

safetyarxiv-cs-ai
28 Apr 2026
Safety

ArgRE: Formal Argumentation for Conflict Resolution in Multi-Agent Requirements Negotiation

DGX agent

arXiv:2604.23124v1 Announce Type: cross Abstract: As software systems grow in complexity, they must satisfy an increasing number of competing quality attributes, making it essential to balance them in

safetyarxiv-cs-ai
28 Apr 2026
Safety

Autocorrelation Reintroduces Spectral Bias in KANs for Time Series Forecasting

DGX agent

arXiv:2604.23518v1 Announce Type: cross Abstract: Existing theory suggests that Kolmogorov-Arnold Networks (KANs) can overcome the spectral bias commonly observed in neural networks under the assumpti

safetyarxiv-cs-ai
28 Apr 2026
Model Releases

Beyond Context: Large Language Models' Failure to Grasp Users' Intent

DGX agent

arXiv:2512.21110v3 Announce Type: replace Abstract: Current Large Language Models (LLMs) safety approaches focus on explicitly harmful content while overlooking a critical vulnerability: the inability

model-releasesarxiv-cs-ai
28 Apr 2026
Safety

Context-Aware Hospitalization Forecasting Evaluations for Decision Support using LLMs

DGX agent

arXiv:2604.23949v1 Announce Type: new Abstract: Medical and public health experts must make real-time resource decisions, such as expanding hospital bed capacity, based on projected hospitalization tr

safetyarxiv-cs-ai
28 Apr 2026
Safety

Designing escalation criteria for international AI incident response: criteria, triggers, and thresholds

DGX agent

arXiv:2604.23183v1 Announce Type: cross Abstract: AI incident reporting requirements are emerging in regulation and policy, yet no operational criteria exist for determining when a detected AI inciden

safetyarxiv-cs-ai
28 Apr 2026
Safety

Designing Instance-Level Sampling Schedules via REINFORCE with James-Stein Shrinkage

DGX agent

arXiv:2511.22177v2 Announce Type: replace-cross Abstract: Most post-training methods for text-to-image samplers focus on model weights: either fine-tuning the backbone for alignment or distilling it f

safetyarxiv-cs-cv
28 Apr 2026
Safety

From Stateless Queries to Autonomous Actions: A Layered Security Framework for Agentic AI Systems

DGX agent

arXiv:2604.23338v1 Announce Type: cross Abstract: Agentic AI systems face security challenges that stateless large language models do not. They plan across extended horizons, maintain persistent memor

safetyarxiv-cs-lg
28 Apr 2026
Safety

Governing What You Cannot Observe: Adaptive Runtime Governance for Autonomous AI Agents

DGX agent

arXiv:2604.24686v1 Announce Type: new Abstract: Autonomous AI agents can remain fully authorized and still become unsafe as behavior drifts, adversaries adapt, and decision patterns shift without any

safetyarxiv-cs-ai
28 Apr 2026
Safety

LLM-Auction: Generative Auction towards LLM-Native Advertising

DGX agent

arXiv:2512.10551v2 Announce Type: replace-cross Abstract: The commercialization of LLM applications is the next frontier in online advertising, with LLM-native advertising emerging as a promising para

safetyarxiv-cs-ai
28 Apr 2026
Safety

Polychromic Objectives for Reinforcement Learning

DGX agent

arXiv:2509.25424v5 Announce Type: replace-cross Abstract: Reinforcement learning fine-tuning (RLFT) is a dominant paradigm for improving pretrained policies for downstream tasks. These pretrained poli

safetyarxiv-cs-ai
28 Apr 2026
Safety

Position: Logical Soundness is not a Reliable Criterion for Neurosymbolic Fact-Checking with LLMs

DGX agent

arXiv:2604.04177v2 Announce Type: replace Abstract: As large language models (LLMs) are increasing integrated into fact-checking pipelines, formal logic is often proposed as a rigorous means by which

safetyarxiv-cs-cl
28 Apr 2026
Safety

ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI Evaluation

DGX agent

arXiv:2604.23099v1 Announce Type: cross Abstract: Evaluating generative AI models is increasingly resource-intensive due to slow inference, expensive raters, and a rapidly growing landscape of models

safetyarxiv-cs-ai
28 Apr 2026
Safety

Protecting the Trace: A Principled Black-Box Approach Against Distillation Attacks

DGX agent

arXiv:2604.23238v1 Announce Type: cross Abstract: Frontier models push the boundaries of what is learnable at extreme computational costs, yet distillation via sampling reasoning traces exposes closed

safetyarxiv-cs-ai
28 Apr 2026
Safety

Safe Navigation in Unknown and Cluttered Environments via Direction-Aware Convex Free-Region Generation

DGX agent

arXiv:2604.23648v1 Announce Type: new Abstract: Convex free regions provide a structured and optimization-friendly representation of collision-free space for robot navigation in unknown and cluttered

safetyarxiv-cs-ro
28 Apr 2026
← Previous
1…3738394041…299
Next →