AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlog
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,357 results
Safety

The Unified Autonomy Stack: Toward a Blueprint for Generalizable Robot Autonomy

DGX agent

arXiv:2605.12735v1 Announce Type: new Abstract: We introduce and open-source the Unified Autonomy Stack, a system-level solution that enables resilient autonomy across diverse aerial and ground robot

safetyarxiv-cs-ro
14 May 2026
Safety
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

VERA-MH: Validation of Ethical and Responsible AI in Mental Health

DGX agent

arXiv:2605.13318v1 Announce Type: new Abstract: Chatbot usage has increased, including in fields for which they were never developed for--notably mental health support. To that end, we introduce Valid

safetyarxiv-cs-ai
14 May 2026
Safety

Can a Single Message Paralyze the AI Infrastructure? The Rise of AbO-DDoS Attacks through Targeted Mobius Injection

DGX agent

arXiv:2605.11442v1 Announce Type: cross Abstract: Large Language Model (LLM) agents have emerged as key intermediaries, orchestrating complex interactions between human users and a wide range of digit

safetyarxiv-cs-cl
13 May 2026
Safety

Characterizing the Robustness of Black-Box LLM Planners Under Perturbed Observations with Adaptive Stress Testing

DGX agent

arXiv:2505.05665v4 Announce Type: replace-cross Abstract: Large language models (LLMs) have recently demonstrated success in decision-making tasks including planning, control, and prediction, but thei

safetyarxiv-cs-cl
13 May 2026
Safety

Continuous Discovery of Vulnerabilities in LLM Serving Systems with Fuzzing

DGX agent

arXiv:2605.11202v1 Announce Type: cross Abstract: LLM inference and serving systems have become security-critical infrastructure; however, many of their most concerning failures arise from the serving

safetyarxiv-cs-lg
13 May 2026
Safety

Leveraging RAG for Training-Free Alignment of LLMs

DGX agent

arXiv:2605.11217v1 Announce Type: new Abstract: Large language model (LLM) alignment algorithms typically consist of post-training over preference pairs. While such algorithms are widely used to enabl

safetyarxiv-cs-lg
13 May 2026
Safety

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue

DGX agent

arXiv:2605.05630v2 Announce Type: replace Abstract: Hidden malicious intent in multi-turn dialogue poses a growing threat to deployed large language models (LLMs). Rather than exposing a harmful objec

safetyarxiv-cs-cl
13 May 2026
Safety

Robust LLM Unlearning Against Relearning Attacks: The Minor Components in Representations Matter

DGX agent

arXiv:2605.11685v1 Announce Type: new Abstract: Large language model (LLM) unlearning aims to remove specific data influences from pre-trained model without costly retraining, addressing privacy, copy

safetyarxiv-cs-cl
13 May 2026
Safety

The DAWN of World-Action Interactive Models

DGX agent

arXiv:2605.11550v1 Announce Type: new Abstract: A plausible scene evolution depends on the maneuver being considered, while a good maneuver depends on how the scene may evolve. Existing World Action M

safetyarxiv-cs-cv
13 May 2026
Safety

The new era of SaMD: Why cloud infrastructure is the foundation for digital health in 2026

DGX agent

In the healthcare and life sciences industries, speed saves lives, but meeting regulatory requirements and other administrative burdens often pumps the brakes for manufacturers of software as a medica

safetygoogle-cloud-ai
13 May 2026
Safety

there are many agent use cases locked behind 'what if' fears because after we give a tool to an agent, we're relying on the prompt to limit …

DGX agent

there are many agent use cases locked behind 'what if' fears because after we give a tool to an agent, we're relying on the prompt to limit behavior tools like @denieddotdev allow teams to manage capa

safetyyohei-nakajima--x
13 May 2026
Safety

ActivationReasoning: Logical Reasoning in Latent Activation Spaces

DGX agent

arXiv:2510.18184v3 Announce Type: replace-cross Abstract: Large language models (LLMs) excel at generating fluent text, but their internal reasoning remains opaque and difficult to control. Sparse aut

safetyarxiv-cs-ai
12 May 2026
Safety

Adversarial Attacks Against MLLMs via Progressive Resolution Processing and Adaptive Feature Alignment

DGX agent

arXiv:2605.09902v1 Announce Type: new Abstract: Adversarial perturbations can mislead Multimodal Large Language Models (MLLMs) recognize a benign image as a specific target object, posing serious risk

safetyarxiv-cs-cv
12 May 2026
Safety

Beyond Autonomy: A Dynamic Tiered AgentRunner Framework for Governable and Resilient Enterprise AI Execution

DGX agent

arXiv:2605.10223v1 Announce Type: new Abstract: Current large language model agent frameworks prioritize autonomy but lack the governability mechanisms required for enterprise deployment. High-risk wr

safetyarxiv-cs-ai
12 May 2026
Safety

Beyond Self-Play: Hierarchical Reasoning for Continuous Motion in Closed-Loop Traffic Simulation

DGX agent

arXiv:2605.09153v1 Announce Type: cross Abstract: Closed-loop traffic simulation requires agents that are both scalable and behaviorally realistic. Recent self-play reinforcement learning approaches d

safetyarxiv-cs-ai
12 May 2026
Safety

Break the Brake, Not the Wheel: Untargeted Jailbreak via Entropy Maximization

DGX agent

arXiv:2605.10764v1 Announce Type: cross Abstract: Recent studies show that gradient-based universal image jailbreaks on vision-language models (VLMs) exhibit little or no cross-model transferability,

safetyarxiv-cs-ai
12 May 2026
Model Releases

C-CoT: Counterfactual Chain-of-Thought with Vision-Language Models for Safe Autonomous Driving

DGX agent

arXiv:2605.10744v1 Announce Type: new Abstract: Safety-critical planning in complex environments, particularly at urban intersections, remains a fundamental challenge for autonomous driving. Existing

model-releasesarxiv-cs-cv
12 May 2026
Safety

DynaMiCS: Fine-tuning LLMs with Performance Constraints using Dynamic Mixtures

DGX agent

arXiv:2605.10770v1 Announce Type: new Abstract: Multi-domain fine-tuning of large language models requires improving performance on target domains while preserving performance on constrained domains,

safetyarxiv-cs-lg
12 May 2026
Safety

f-GRPO and Beyond: Divergence-Based Reinforcement Learning Algorithms for General LLM Alignment

DGX agent

arXiv:2602.05946v3 Announce Type: replace Abstract: Recent work shows that preference alignment objectives can be interpreted as divergence estimators between aligned (preferred) & unaligned (less-pre

safetyarxiv-cs-lg
12 May 2026
Safety

Intelligent Autonomous Orchestration for Distributed Cloud Resources using Complex-Stability Analysis

DGX agent

arXiv:2605.08139v1 Announce Type: cross Abstract: In modern distributed cloud environments, efficient resource allocation is required as traditional scaling mechanisms are often subject to cloud thras

safetyarxiv-cs-ai
12 May 2026
Safety

MedFL-Stress: A Systematic Robustness Evaluation of Federated Brain Tumor Segmentation under Cross-Hospital MRI Appearance Shift

DGX agent

arXiv:2605.09025v1 Announce Type: new Abstract: Federated learning enables hospitals to collaboratively train segmentation models without sharing patient data. However, current evaluation protocols re

safetyarxiv-cs-cv
12 May 2026
Safety

MoMo: Conditioned Contrastive Representation Learning for Preference-Modulated Planning

DGX agent

arXiv:2605.08512v1 Announce Type: new Abstract: Temporally contrastive representation learning induces a latent structure capable of reducing long-horizon planning to inference in a low-dimensional li

safetyarxiv-cs-lg
12 May 2026
Safety

Muninn: Your Trajectory Diffusion Model But Faster

DGX agent

arXiv:2605.09999v1 Announce Type: new Abstract: Diffusion-based trajectory planners can synthesize rich, multimodal robot motions, but their iterative denoising makes online planning and control prohi

safetyarxiv-cs-ro
12 May 2026
Safety

Mutual Information Optimal Density Control of Linear Systems and Generalized Schrodinger Bridges with Reference Refinement

DGX agent

arXiv:2605.09349v1 Announce Type: cross Abstract: We consider a mutual information (MI) regularized version of optimal density control of a discrete-time linear system. MI optimal control has been pro

safetyarxiv-cs-lg
12 May 2026
Safety

Not All Turns Matter: Credit Assignment for Multi-Turn Jailbreaking

DGX agent

arXiv:2605.08778v1 Announce Type: new Abstract: Deploying LLMs in multi-turn dialogues facilitates jailbreak attacks that distribute harmful intent across seemingly benign turns. Recent training-based

safetyarxiv-cs-ai
12 May 2026
Safety

On Uniform Error Bounds for Kernel Regression under Non-Gaussian Noise

DGX agent

arXiv:2605.09757v1 Announce Type: new Abstract: Providing non-conservative uncertainty quantification for function estimates derived from noisy observations remains a fundamental challenge in statisti

safetyarxiv-cs-lg
12 May 2026
Safety

PHMForge: Evaluating LLM Agents on Industrial Prognostics through MCP-Native, Algorithm-Grounded Tools

DGX agent

arXiv:2604.01532v2 Announce Type: replace Abstract: LLM agents are beginning to invoke industrial asset-management tools through the Model Context Protocol (MCP), yet whether they can act reliably on

safetyarxiv-cs-ai
12 May 2026
Safety

Research on Security Enhancement Methods for Adversarial Robust Large Language Model Intelligent Agents for Medical Decision-Making Tasks

DGX agent

arXiv:2605.08257v1 Announce Type: cross Abstract: Motivated by the challenge to improve the adversarial robustness, security, and trust of medical decision making intelligent agents, this study develo

safetyarxiv-cs-ai
12 May 2026
Safety

Self-ReSET: Learning to Self-Recover from Unsafe Reasoning Trajectories

DGX agent

arXiv:2605.08936v1 Announce Type: new Abstract: Large Reasoning Models possess remarkable capabilities for self-correction in general domain; however, they frequently struggle to recover from unsafe r

safetyarxiv-cs-ai
12 May 2026
Safety

SHIELD: Scalable Optimal Control with Certification using Duality and Convexity

DGX agent

arXiv:2605.09171v1 Announce Type: new Abstract: We present SHIELD, a hierarchical algorithm that reduces both the decision-variable dimension and the constraint set in ell_1-regularized convex program

safetyarxiv-cs-ro
12 May 2026
Safety

SnareNet: Flexible Repair Layers for Neural Networks with Hard Constraints

DGX agent

arXiv:2602.09317v2 Announce Type: replace-cross Abstract: Neural networks are increasingly used as fast surrogate models across various domains, but unconstrained predictions can violate physical, ope

safetyarxiv-cs-ai
12 May 2026
Safety

Supervised Mixture-of-Experts for Surgical Grasping and Retraction

DGX agent

arXiv:2601.21971v2 Announce Type: replace-cross Abstract: Imitation learning has achieved remarkable success in robotic manipulation, yet its application to surgical robotics remains challenging due t

safetyarxiv-cs-ai
12 May 2026
Safety

The Art of the Jailbreak: Formulating Jailbreak Attacks for LLM Security Beyond Binary Scoring

DGX agent

arXiv:2605.09225v1 Announce Type: cross Abstract: Jailbreak attacks -- adversarial prompts that bypass LLM alignment through purely linguistic manipulation -- pose a growing operational security threa

safetyarxiv-cs-ai
12 May 2026
Safety

The Value of Mechanistic Priors in Sequential Decision Making

DGX agent

arXiv:2605.10018v1 Announce Type: new Abstract: Hybrid mechanistic models, physical priors with learned residuals, promise to reduce the data required for good decisions, but have no computable criter

safetyarxiv-cs-lg
12 May 2026
Safety

There will be no AI jobpocalypse. The story that AI will lead to massive unemployment is stoking unnecessary fear. AI — like any other techn…

DGX agent

There will be no AI jobpocalypse. The story that AI will lead to massive unemployment is stoking unnecessary fear. AI — like any other technology — does affect jobs, but telling overblown stories of l

safetyandrew-ng--x
12 May 2026
Safety

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning

DGX agent

arXiv:2508.20697v3 Announce Type: replace-cross Abstract: As large language models (LLMs) continue to grow in capability, so do the risks of harmful misuse through fine-tuning. While most prior studie

safetyarxiv-cs-cl
12 May 2026
Safety

Unlearners Can Lie: Evaluating and Improving Honesty in LLM Unlearning

DGX agent

arXiv:2605.08765v1 Announce Type: cross Abstract: Unlearning in large language models (LLMs) aims to remove harmful training data while preserving overall utility. However, we find that existing metho

safetyarxiv-cs-ai
12 May 2026
Safety

Variational Inference for Levy Process-Driven SDEs via Neural Tilting

DGX agent

arXiv:2605.10934v1 Announce Type: cross Abstract: Modelling extreme events and heavy-tailed phenomena is central to building reliable predictive systems in domains such as finance, climate science, an

safetyarxiv-cs-ai
12 May 2026
Safety

When a Robot is More Capable than a Human: Learning from Constrained Demonstrators

DGX agent

arXiv:2510.09096v3 Announce Type: replace-cross Abstract: Learning from demonstrations enables experts to teach robots complex tasks using interfaces such as kinesthetic teaching, joystick control, an

safetyarxiv-cs-ai
12 May 2026
Safety

Beyond Confidence: Rethinking Self-Assessments for Performance Prediction in LLMs

DGX agent

arXiv:2605.07806v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used in settings where reliable self-assessment is critical. Assessing model reliability has evolved fro

safetyarxiv-cs-ai
11 May 2026
Safety

Future-proof your data strategy: AlloyDB adds PostgreSQL 18 and new Extended Support

DGX agent

As you look out at your 2026 infrastructure roadmap, your goal is to balance the need for rapid innovation with operational stability. You shouldn't have to choose between adopting the latest database

safetygoogle-cloud-ai
11 May 2026
Safety

Intention assimilation control for accurate tracking with variable impedance in teleoperation

DGX agent

arXiv:2605.07037v1 Announce Type: new Abstract: Robot systems for teleoperation commonly use a spring-like force pulling the follower robot towards the leader's position to track their movements. With

safetyarxiv-cs-ro
11 May 2026
Safety

Learned Lyapunov Shielding for Adaptive Control

DGX agent

arXiv:2605.06934v1 Announce Type: new Abstract: We augment the Slotine--Li adaptive controller for Euler--Lagrange systems with three learned components: a structured-quadratic Lyapunov function (V_ps

safetyarxiv-cs-lg
11 May 2026
Safety

Probabilistic Object Detection with Conformal Prediction

DGX agent

arXiv:2605.07549v1 Announce Type: new Abstract: Conformal Prediction (CP) is a distribution-free method for constructing prediction sets with marginal finite-sample coverage guarantees, making it a su

safetyarxiv-cs-cv
11 May 2026
Safety

Sensitivity-Based Robust NMPC for Close-Proximity Offshore Wind Turbine Inspection with a Tilted Multirotor

DGX agent

arXiv:2605.07771v1 Announce Type: new Abstract: Close-proximity offshore wind turbine inspection requires strict clearance control around large cylindrical structures under wind and model mismatch. No

safetyarxiv-cs-ro
11 May 2026
Safety

The Moltbook Files: A Harmless Slopocalypse or Humanity's Last Experiment

DGX agent

arXiv:2605.07462v1 Announce Type: cross Abstract: Moltbook is a Reddit-like platform where OpenClaw agents post, comment, and vote at scale - a so far unprecedented incident that comes with serious sa

safetyarxiv-cs-ai
11 May 2026
Safety

Theoretical Limits of Language Model Alignment

DGX agent

arXiv:2605.07105v1 Announce Type: cross Abstract: Language model (LM) alignment improves model outputs to reflect human preferences while preserving the capabilities of the base model. The most common

safetyarxiv-cs-cl
11 May 2026
Safety

RVPO: Risk-Sensitive Alignment via Variance Regularization

DGX agent

Current critic-less RLHF methods aggregate multi-objective rewards via an arithmetic mean, leaving them vulnerable to constraint neglect: high-magnitude success in one objective can numerically offset

safetyapple-ml-research
8 May 2026
← Previous
1…5657585960…300
Next →