AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,314 results
Safety

Viverra: Text-to-Code with Guarantees

DGX agent

arXiv:2605.14972v1 Announce Type: cross Abstract: A fundamental limitation of Text-to-Code is that no guarantee can be obtained about the correctness of the generated code. Therefore, to ensure its co

safetyarxiv-cs-ai
15 May 2026
Safety

Asynchronous Reasoning: Training-Free Interactive Thinking LLMs

DGX agent
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2512.10931v3 Announce Type: replace Abstract: Many state-of-the-art LLMs are trained to think before giving their answer. Reasoning can greatly improve language model capabilities, but it also m

safetyarxiv-cs-lg
14 May 2026
Safety

Automated Rubrics for Reliable Evaluation of Medical Dialogue Systems

DGX agent

arXiv:2601.15161v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly used for clinical decision support, where hallucinations and unsafe suggestions may pose direct

safetyarxiv-cs-ai
14 May 2026
Safety

BEHAVE: A Hybrid AI Framework for Real-Time Modeling of Collective Human Dynamics

DGX agent

arXiv:2605.12730v1 Announce Type: new Abstract: Existing AI systems for modeling human behavior operate at the level of individuals or detect events after they occur. As a result, they systematically

safetyarxiv-cs-ai
14 May 2026
Safety

DiffusionHijack: Supply-Chain PRNG Backdoor Attack on Diffusion Models and Quantum Random Number Defense

DGX agent

arXiv:2605.13115v1 Announce Type: cross Abstract: Diffusion models depend on pseudo-random number generators (PRNGs) for latent noise sampling. We present DiffusionHijack, a supply-chain backdoor atta

safetyarxiv-cs-lg
14 May 2026
Safety

Loiter UAV Reinsertion Guidance for Fixed-wing UAV Corridors

DGX agent

arXiv:2605.13822v1 Announce Type: new Abstract: This paper considers fixed-wing unmanned aerial vehicle (UAV) corridors comprising a main lane, a circular loiter lane for managing traffic congestion,

safetyarxiv-cs-ro
14 May 2026
Model Releases

Neurosymbolic Auditing of Natural-Language Software Requirements

DGX agent

arXiv:2605.13817v1 Announce Type: cross Abstract: Natural-language software requirements are often ambiguous, inconsistent, and underspecified; in safety-critical domains, these defects propagate into

model-releasesarxiv-cs-ai
14 May 2026
Safety

No Attack Required: Semantic Fuzzing for Specification Violations in Agent Skills

DGX agent

arXiv:2605.13044v1 Announce Type: cross Abstract: LLM-powered agents can silently delete documents, leak credentials, or transfer funds on a routine user request, not because the agent was attacked, b

safetyarxiv-cs-ai
14 May 2026
Safety

Runtime Monitoring of Perception-Based Autonomous Systems via Embedding Temporal Logic

DGX agent

arXiv:2605.12651v1 Announce Type: new Abstract: Runtime monitoring of autonomous systems traditionally relies on mapping continuous sensor observations to discrete logical propositions defined over lo

safetyarxiv-cs-lg
14 May 2026
Safety

The Unified Autonomy Stack: Toward a Blueprint for Generalizable Robot Autonomy

DGX agent

arXiv:2605.12735v1 Announce Type: new Abstract: We introduce and open-source the Unified Autonomy Stack, a system-level solution that enables resilient autonomy across diverse aerial and ground robot

safetyarxiv-cs-ro
14 May 2026
Safety

VERA-MH: Validation of Ethical and Responsible AI in Mental Health

DGX agent

arXiv:2605.13318v1 Announce Type: new Abstract: Chatbot usage has increased, including in fields for which they were never developed for--notably mental health support. To that end, we introduce Valid

safetyarxiv-cs-ai
14 May 2026
Safety

Can a Single Message Paralyze the AI Infrastructure? The Rise of AbO-DDoS Attacks through Targeted Mobius Injection

DGX agent

arXiv:2605.11442v1 Announce Type: cross Abstract: Large Language Model (LLM) agents have emerged as key intermediaries, orchestrating complex interactions between human users and a wide range of digit

safetyarxiv-cs-cl
13 May 2026
Safety

Characterizing the Robustness of Black-Box LLM Planners Under Perturbed Observations with Adaptive Stress Testing

DGX agent

arXiv:2505.05665v4 Announce Type: replace-cross Abstract: Large language models (LLMs) have recently demonstrated success in decision-making tasks including planning, control, and prediction, but thei

safetyarxiv-cs-cl
13 May 2026
Safety

Continuous Discovery of Vulnerabilities in LLM Serving Systems with Fuzzing

DGX agent

arXiv:2605.11202v1 Announce Type: cross Abstract: LLM inference and serving systems have become security-critical infrastructure; however, many of their most concerning failures arise from the serving

safetyarxiv-cs-lg
13 May 2026
Safety

Leveraging RAG for Training-Free Alignment of LLMs

DGX agent

arXiv:2605.11217v1 Announce Type: new Abstract: Large language model (LLM) alignment algorithms typically consist of post-training over preference pairs. While such algorithms are widely used to enabl

safetyarxiv-cs-lg
13 May 2026
Safety

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue

DGX agent

arXiv:2605.05630v2 Announce Type: replace Abstract: Hidden malicious intent in multi-turn dialogue poses a growing threat to deployed large language models (LLMs). Rather than exposing a harmful objec

safetyarxiv-cs-cl
13 May 2026
Safety

Robust LLM Unlearning Against Relearning Attacks: The Minor Components in Representations Matter

DGX agent

arXiv:2605.11685v1 Announce Type: new Abstract: Large language model (LLM) unlearning aims to remove specific data influences from pre-trained model without costly retraining, addressing privacy, copy

safetyarxiv-cs-cl
13 May 2026
Safety

The DAWN of World-Action Interactive Models

DGX agent

arXiv:2605.11550v1 Announce Type: new Abstract: A plausible scene evolution depends on the maneuver being considered, while a good maneuver depends on how the scene may evolve. Existing World Action M

safetyarxiv-cs-cv
13 May 2026
Safety

ActivationReasoning: Logical Reasoning in Latent Activation Spaces

DGX agent

arXiv:2510.18184v3 Announce Type: replace-cross Abstract: Large language models (LLMs) excel at generating fluent text, but their internal reasoning remains opaque and difficult to control. Sparse aut

safetyarxiv-cs-ai
12 May 2026
Safety

Adversarial Attacks Against MLLMs via Progressive Resolution Processing and Adaptive Feature Alignment

DGX agent

arXiv:2605.09902v1 Announce Type: new Abstract: Adversarial perturbations can mislead Multimodal Large Language Models (MLLMs) recognize a benign image as a specific target object, posing serious risk

safetyarxiv-cs-cv
12 May 2026
Safety

Beyond Autonomy: A Dynamic Tiered AgentRunner Framework for Governable and Resilient Enterprise AI Execution

DGX agent

arXiv:2605.10223v1 Announce Type: new Abstract: Current large language model agent frameworks prioritize autonomy but lack the governability mechanisms required for enterprise deployment. High-risk wr

safetyarxiv-cs-ai
12 May 2026
Safety

Beyond Self-Play: Hierarchical Reasoning for Continuous Motion in Closed-Loop Traffic Simulation

DGX agent

arXiv:2605.09153v1 Announce Type: cross Abstract: Closed-loop traffic simulation requires agents that are both scalable and behaviorally realistic. Recent self-play reinforcement learning approaches d

safetyarxiv-cs-ai
12 May 2026
Safety

Break the Brake, Not the Wheel: Untargeted Jailbreak via Entropy Maximization

DGX agent

arXiv:2605.10764v1 Announce Type: cross Abstract: Recent studies show that gradient-based universal image jailbreaks on vision-language models (VLMs) exhibit little or no cross-model transferability,

safetyarxiv-cs-ai
12 May 2026
Model Releases

C-CoT: Counterfactual Chain-of-Thought with Vision-Language Models for Safe Autonomous Driving

DGX agent

arXiv:2605.10744v1 Announce Type: new Abstract: Safety-critical planning in complex environments, particularly at urban intersections, remains a fundamental challenge for autonomous driving. Existing

model-releasesarxiv-cs-cv
12 May 2026
Safety

DynaMiCS: Fine-tuning LLMs with Performance Constraints using Dynamic Mixtures

DGX agent

arXiv:2605.10770v1 Announce Type: new Abstract: Multi-domain fine-tuning of large language models requires improving performance on target domains while preserving performance on constrained domains,

safetyarxiv-cs-lg
12 May 2026
Safety

f-GRPO and Beyond: Divergence-Based Reinforcement Learning Algorithms for General LLM Alignment

DGX agent

arXiv:2602.05946v3 Announce Type: replace Abstract: Recent work shows that preference alignment objectives can be interpreted as divergence estimators between aligned (preferred) & unaligned (less-pre

safetyarxiv-cs-lg
12 May 2026
Safety

Intelligent Autonomous Orchestration for Distributed Cloud Resources using Complex-Stability Analysis

DGX agent

arXiv:2605.08139v1 Announce Type: cross Abstract: In modern distributed cloud environments, efficient resource allocation is required as traditional scaling mechanisms are often subject to cloud thras

safetyarxiv-cs-ai
12 May 2026
Safety

MedFL-Stress: A Systematic Robustness Evaluation of Federated Brain Tumor Segmentation under Cross-Hospital MRI Appearance Shift

DGX agent

arXiv:2605.09025v1 Announce Type: new Abstract: Federated learning enables hospitals to collaboratively train segmentation models without sharing patient data. However, current evaluation protocols re

safetyarxiv-cs-cv
12 May 2026
Safety

MoMo: Conditioned Contrastive Representation Learning for Preference-Modulated Planning

DGX agent

arXiv:2605.08512v1 Announce Type: new Abstract: Temporally contrastive representation learning induces a latent structure capable of reducing long-horizon planning to inference in a low-dimensional li

safetyarxiv-cs-lg
12 May 2026
Safety

Muninn: Your Trajectory Diffusion Model But Faster

DGX agent

arXiv:2605.09999v1 Announce Type: new Abstract: Diffusion-based trajectory planners can synthesize rich, multimodal robot motions, but their iterative denoising makes online planning and control prohi

safetyarxiv-cs-ro
12 May 2026
Safety

Mutual Information Optimal Density Control of Linear Systems and Generalized Schrodinger Bridges with Reference Refinement

DGX agent

arXiv:2605.09349v1 Announce Type: cross Abstract: We consider a mutual information (MI) regularized version of optimal density control of a discrete-time linear system. MI optimal control has been pro

safetyarxiv-cs-lg
12 May 2026
Safety

Not All Turns Matter: Credit Assignment for Multi-Turn Jailbreaking

DGX agent

arXiv:2605.08778v1 Announce Type: new Abstract: Deploying LLMs in multi-turn dialogues facilitates jailbreak attacks that distribute harmful intent across seemingly benign turns. Recent training-based

safetyarxiv-cs-ai
12 May 2026
Safety

On Uniform Error Bounds for Kernel Regression under Non-Gaussian Noise

DGX agent

arXiv:2605.09757v1 Announce Type: new Abstract: Providing non-conservative uncertainty quantification for function estimates derived from noisy observations remains a fundamental challenge in statisti

safetyarxiv-cs-lg
12 May 2026
Safety

PHMForge: Evaluating LLM Agents on Industrial Prognostics through MCP-Native, Algorithm-Grounded Tools

DGX agent

arXiv:2604.01532v2 Announce Type: replace Abstract: LLM agents are beginning to invoke industrial asset-management tools through the Model Context Protocol (MCP), yet whether they can act reliably on

safetyarxiv-cs-ai
12 May 2026
Safety

Research on Security Enhancement Methods for Adversarial Robust Large Language Model Intelligent Agents for Medical Decision-Making Tasks

DGX agent

arXiv:2605.08257v1 Announce Type: cross Abstract: Motivated by the challenge to improve the adversarial robustness, security, and trust of medical decision making intelligent agents, this study develo

safetyarxiv-cs-ai
12 May 2026
Safety

Self-ReSET: Learning to Self-Recover from Unsafe Reasoning Trajectories

DGX agent

arXiv:2605.08936v1 Announce Type: new Abstract: Large Reasoning Models possess remarkable capabilities for self-correction in general domain; however, they frequently struggle to recover from unsafe r

safetyarxiv-cs-ai
12 May 2026
Safety

SHIELD: Scalable Optimal Control with Certification using Duality and Convexity

DGX agent

arXiv:2605.09171v1 Announce Type: new Abstract: We present SHIELD, a hierarchical algorithm that reduces both the decision-variable dimension and the constraint set in ell_1-regularized convex program

safetyarxiv-cs-ro
12 May 2026
Safety

SnareNet: Flexible Repair Layers for Neural Networks with Hard Constraints

DGX agent

arXiv:2602.09317v2 Announce Type: replace-cross Abstract: Neural networks are increasingly used as fast surrogate models across various domains, but unconstrained predictions can violate physical, ope

safetyarxiv-cs-ai
12 May 2026
Safety

Supervised Mixture-of-Experts for Surgical Grasping and Retraction

DGX agent

arXiv:2601.21971v2 Announce Type: replace-cross Abstract: Imitation learning has achieved remarkable success in robotic manipulation, yet its application to surgical robotics remains challenging due t

safetyarxiv-cs-ai
12 May 2026
Safety

The Art of the Jailbreak: Formulating Jailbreak Attacks for LLM Security Beyond Binary Scoring

DGX agent

arXiv:2605.09225v1 Announce Type: cross Abstract: Jailbreak attacks -- adversarial prompts that bypass LLM alignment through purely linguistic manipulation -- pose a growing operational security threa

safetyarxiv-cs-ai
12 May 2026
Safety

The Value of Mechanistic Priors in Sequential Decision Making

DGX agent

arXiv:2605.10018v1 Announce Type: new Abstract: Hybrid mechanistic models, physical priors with learned residuals, promise to reduce the data required for good decisions, but have no computable criter

safetyarxiv-cs-lg
12 May 2026
Safety

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning

DGX agent

arXiv:2508.20697v3 Announce Type: replace-cross Abstract: As large language models (LLMs) continue to grow in capability, so do the risks of harmful misuse through fine-tuning. While most prior studie

safetyarxiv-cs-cl
12 May 2026
Safety

Unlearners Can Lie: Evaluating and Improving Honesty in LLM Unlearning

DGX agent

arXiv:2605.08765v1 Announce Type: cross Abstract: Unlearning in large language models (LLMs) aims to remove harmful training data while preserving overall utility. However, we find that existing metho

safetyarxiv-cs-ai
12 May 2026
Safety

Variational Inference for Levy Process-Driven SDEs via Neural Tilting

DGX agent

arXiv:2605.10934v1 Announce Type: cross Abstract: Modelling extreme events and heavy-tailed phenomena is central to building reliable predictive systems in domains such as finance, climate science, an

safetyarxiv-cs-ai
12 May 2026
Safety

When a Robot is More Capable than a Human: Learning from Constrained Demonstrators

DGX agent

arXiv:2510.09096v3 Announce Type: replace-cross Abstract: Learning from demonstrations enables experts to teach robots complex tasks using interfaces such as kinesthetic teaching, joystick control, an

safetyarxiv-cs-ai
12 May 2026
Safety

Beyond Confidence: Rethinking Self-Assessments for Performance Prediction in LLMs

DGX agent

arXiv:2605.07806v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used in settings where reliable self-assessment is critical. Assessing model reliability has evolved fro

safetyarxiv-cs-ai
11 May 2026
Safety

Intention assimilation control for accurate tracking with variable impedance in teleoperation

DGX agent

arXiv:2605.07037v1 Announce Type: new Abstract: Robot systems for teleoperation commonly use a spring-like force pulling the follower robot towards the leader's position to track their movements. With

safetyarxiv-cs-ro
11 May 2026
Safety

Learned Lyapunov Shielding for Adaptive Control

DGX agent

arXiv:2605.06934v1 Announce Type: new Abstract: We augment the Slotine--Li adaptive controller for Euler--Lagrange systems with three learned components: a structured-quadratic Lyapunov function (V_ps

safetyarxiv-cs-lg
11 May 2026
← Previous
1…5051525354…257
Next →