AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,356 results
12 May 2026

On Uniform Error Bounds for Kernel Regression under Non-Gaussian Noise

SafetyDGX agent

arXiv:2605.09757v1 Announce Type: new Abstract: Providing non-conservative uncertainty quantification for function estimates derived from noisy observations remains a fundamental challenge in statisti

PHMForge: Evaluating LLM Agents on Industrial Prognostics through MCP-Native, Algorithm-Grounded Tools

SafetyDGX agent

arXiv:2604.01532v2 Announce Type: replace Abstract: LLM agents are beginning to invoke industrial asset-management tools through the Model Context Protocol (MCP), yet whether they can act reliably on

Research on Security Enhancement Methods for Adversarial Robust Large Language Model Intelligent Agents for Medical Decision-Making Tasks

SafetyDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.08257v1 Announce Type: cross Abstract: Motivated by the challenge to improve the adversarial robustness, security, and trust of medical decision making intelligent agents, this study develo

Self-ReSET: Learning to Self-Recover from Unsafe Reasoning Trajectories

SafetyDGX agent

arXiv:2605.08936v1 Announce Type: new Abstract: Large Reasoning Models possess remarkable capabilities for self-correction in general domain; however, they frequently struggle to recover from unsafe r

SHIELD: Scalable Optimal Control with Certification using Duality and Convexity

SafetyDGX agent

arXiv:2605.09171v1 Announce Type: new Abstract: We present SHIELD, a hierarchical algorithm that reduces both the decision-variable dimension and the constraint set in ell_1-regularized convex program

SnareNet: Flexible Repair Layers for Neural Networks with Hard Constraints

SafetyDGX agent

arXiv:2602.09317v2 Announce Type: replace-cross Abstract: Neural networks are increasingly used as fast surrogate models across various domains, but unconstrained predictions can violate physical, ope

Supervised Mixture-of-Experts for Surgical Grasping and Retraction

SafetyDGX agent

arXiv:2601.21971v2 Announce Type: replace-cross Abstract: Imitation learning has achieved remarkable success in robotic manipulation, yet its application to surgical robotics remains challenging due t

The Art of the Jailbreak: Formulating Jailbreak Attacks for LLM Security Beyond Binary Scoring

SafetyDGX agent

arXiv:2605.09225v1 Announce Type: cross Abstract: Jailbreak attacks -- adversarial prompts that bypass LLM alignment through purely linguistic manipulation -- pose a growing operational security threa

The Value of Mechanistic Priors in Sequential Decision Making

SafetyDGX agent

arXiv:2605.10018v1 Announce Type: new Abstract: Hybrid mechanistic models, physical priors with learned residuals, promise to reduce the data required for good decisions, but have no computable criter

There will be no AI jobpocalypse. The story that AI will lead to massive unemployment is stoking unnecessary fear. AI — like any other techn…

SafetyDGX agent

There will be no AI jobpocalypse. The story that AI will lead to massive unemployment is stoking unnecessary fear. AI — like any other technology — does affect jobs, but telling overblown stories of l

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning

SafetyDGX agent

arXiv:2508.20697v3 Announce Type: replace-cross Abstract: As large language models (LLMs) continue to grow in capability, so do the risks of harmful misuse through fine-tuning. While most prior studie

Unlearners Can Lie: Evaluating and Improving Honesty in LLM Unlearning

SafetyDGX agent

arXiv:2605.08765v1 Announce Type: cross Abstract: Unlearning in large language models (LLMs) aims to remove harmful training data while preserving overall utility. However, we find that existing metho

Variational Inference for Levy Process-Driven SDEs via Neural Tilting

SafetyDGX agent

arXiv:2605.10934v1 Announce Type: cross Abstract: Modelling extreme events and heavy-tailed phenomena is central to building reliable predictive systems in domains such as finance, climate science, an

When a Robot is More Capable than a Human: Learning from Constrained Demonstrators

SafetyDGX agent

arXiv:2510.09096v3 Announce Type: replace-cross Abstract: Learning from demonstrations enables experts to teach robots complex tasks using interfaces such as kinesthetic teaching, joystick control, an

11 May 2026

Beyond Confidence: Rethinking Self-Assessments for Performance Prediction in LLMs

SafetyDGX agent

arXiv:2605.07806v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used in settings where reliable self-assessment is critical. Assessing model reliability has evolved fro

Future-proof your data strategy: AlloyDB adds PostgreSQL 18 and new Extended Support

SafetyDGX agent

As you look out at your 2026 infrastructure roadmap, your goal is to balance the need for rapid innovation with operational stability. You shouldn't have to choose between adopting the latest database

Intention assimilation control for accurate tracking with variable impedance in teleoperation

SafetyDGX agent

arXiv:2605.07037v1 Announce Type: new Abstract: Robot systems for teleoperation commonly use a spring-like force pulling the follower robot towards the leader's position to track their movements. With

Learned Lyapunov Shielding for Adaptive Control

SafetyDGX agent

arXiv:2605.06934v1 Announce Type: new Abstract: We augment the Slotine--Li adaptive controller for Euler--Lagrange systems with three learned components: a structured-quadratic Lyapunov function (V_ps

Probabilistic Object Detection with Conformal Prediction

SafetyDGX agent

arXiv:2605.07549v1 Announce Type: new Abstract: Conformal Prediction (CP) is a distribution-free method for constructing prediction sets with marginal finite-sample coverage guarantees, making it a su

Sensitivity-Based Robust NMPC for Close-Proximity Offshore Wind Turbine Inspection with a Tilted Multirotor

SafetyDGX agent

arXiv:2605.07771v1 Announce Type: new Abstract: Close-proximity offshore wind turbine inspection requires strict clearance control around large cylindrical structures under wind and model mismatch. No

The Moltbook Files: A Harmless Slopocalypse or Humanity's Last Experiment

SafetyDGX agent

arXiv:2605.07462v1 Announce Type: cross Abstract: Moltbook is a Reddit-like platform where OpenClaw agents post, comment, and vote at scale - a so far unprecedented incident that comes with serious sa

Theoretical Limits of Language Model Alignment

SafetyDGX agent

arXiv:2605.07105v1 Announce Type: cross Abstract: Language model (LM) alignment improves model outputs to reflect human preferences while preserving the capabilities of the base model. The most common

8 May 2026

RVPO: Risk-Sensitive Alignment via Variance Regularization

SafetyDGX agent

Current critic-less RLHF methods aggregate multi-objective rewards via an arithmetic mean, leaving them vulnerable to constraint neglect: high-magnitude success in one objective can numerically offset

The human brain🧠 is incredibly efficient because it only activates the specific neurons needed for a thought. Modern LLMs naturally try to …

SafetyDGX agent

The human brain🧠 is incredibly efficient because it only activates the specific neurons needed for a thought. Modern LLMs naturally try to do this too (> 95% of neurons in feedforward layers stay sile

7 May 2026

Anatomy of a failure: When, how, and why deep vision fails in scientific domains

SafetyDGX agent

arXiv:2605.04231v1 Announce Type: new Abstract: Mirroring its ubiquity in popular media and all human activities, the use of deep learning (DL) is rapidly growing in scientific imaging modalities. How

Beyond Fixed Thresholds and Domain-Specific Benchmarks for Explainable Multi-Task Classification in Autonomous Vehicles

SafetyDGX agent

arXiv:2605.04299v1 Announce Type: new Abstract: Scene understanding is a vital part of autonomous driving systems, which requires the use of deep learning models. Deep learning methods are intrinsical

Encoding Predictability and Legibility for Style-Conditioned Diffusion Policy

SafetyDGX agent

arXiv:2603.16368v2 Announce Type: replace-cross Abstract: Striking a balance between efficiency and transparent motion is a core challenge in human-robot collaboration, as highly expressive movements

From Reach to Insert: Tactile-Augmented Precision Assembly under Sub-Millimeter Tolerances

SafetyDGX agent

arXiv:2605.04649v1 Announce Type: new Abstract: High-precision assembly frequently involves tight-tolerance insertions, where even slight pose errors can cause jamming or excessive interaction forces,

@GaryMarcus is on fire lately... follow him for #AI

SafetyDGX agent

Gary Marcus is an AI researcher and public intellectual who frequently shares commentary and insights about artificial intelligence developments on social media. His posts on X (formerly Twitter) cove

How Does the Lagrangian Guide Safe Reinforcement Learning through Diffusion Models?

SafetyDGX agent

arXiv:2602.02924v2 Announce Type: replace Abstract: Diffusion policy sampling enables reinforcement learning (RL) to represent multimodal action distributions beyond suboptimal unimodal Gaussian polic

I'm really excited about this as a new tool in our interpretability tool kit

SafetyDGX agent

I'm really excited about this as a new tool in our interpretability tool kit In a new paper, we present NLAs, an unsupervised method for converting an LLM's internal state into human-readable text. I'

InterFuserDVS: Event-Enhanced Sensor Fusion for Safe RL-Based Decision Making

SafetyDGX agent

arXiv:2605.04355v1 Announce Type: new Abstract: Autonomous driving systems rely heavily on robust sensor fusion to perceive complex envi- ronments. Traditional setups using RGB cameras and LiDAR often

It'll be 'impossible to slow down the ASI race' until it (very) suddenly isn't.

SafetyDGX agent

This post discusses the dynamics of artificial superintelligence (ASI) development as a competitive race, arguing that competitive pressures make it difficult to slow progress until a critical inflect

Manifold of Failure: Behavioral Attraction Basins in Language Models

Model ReleasesDGX agent

arXiv:2602.22291v3 Announce Type: replace Abstract: While prior work has focused on projecting adversarial examples back onto the manifold of natural data to restore safety, we argue that a comprehens

Misaligned by Reward: Socially Undesirable Preferences in LLMs

SafetyDGX agent

arXiv:2605.05003v1 Announce Type: new Abstract: Reward models are a key component of large language model alignment, serving as proxies for human preferences during training. However, existing evaluat

Predict-then-Diffuse: Adaptive Response Length for Compute-Budgeted Inference in Diffusion LLMs

SafetyDGX agent

arXiv:2605.04215v1 Announce Type: new Abstract: Diffusion-based Large Language Models (D-LLMs) represent a promising frontier in generative AI, offering fully parallel token generation that can lead t

Road Risk Monitor: A Deployable U.S. Road Incident Forecasting System with Live Weather and Road-Level Tiles

SafetyDGX agent

arXiv:2605.04242v1 Announce Type: new Abstract: Nationwide road-incident forecasting is a systems problem before it is a modeling problem. A usable service must connect historical incident archives, h

Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization

SafetyDGX agent

arXiv:2605.04700v1 Announce Type: cross Abstract: Jailbreak attacks on audio language models (ALMs) optimize audio perturbations to elicit unsafe generations, and they typically update the entire wave

Terence Tao recognized that plausibility and veracity are not the same, and that current tools are better at the former than the latter. [ed…

SafetyDGX agent

Terence Tao recognized that plausibility and veracity are not the same, and that current tools are better at the former than the latter. [edit: the video is from 2024, and affirms what i said in 2019

6 May 2026

A Knowledge-Driven LLM-Based Decision-Support System for Explainable Defect Analysis and Mitigation Guidance in Laser Powder Bed Fusion

SafetyDGX agent

arXiv:2605.01100v1 Announce Type: new Abstract: This work presents a knowledge-driven decision-support system that integrates structured defect knowledge with LLM-based reasoning to provide explainabl

Algebraic Semantics of Governed Execution: Monoidal Categories, Effect Algebras, and Coterminous Boundaries

SafetyDGX agent

arXiv:2605.01032v2 Announce Type: new Abstract: We present an algebraic semantics for governed execution in which governance is axiomatized, compositional, and coterminous with expressibility. The fra

Architectural Obsolescence of Unhardened Agentic-AI Runtimes

SafetyDGX agent

arXiv:2605.01740v1 Announce Type: cross Abstract: An agentic-AI runtime issues tool calls, sends messages, and actuates devices on behalf of an LLM. Catching the four ways an action can diverge from i

Audio-Visual Intelligence in Large Foundation Models

SafetyDGX agent

arXiv:2605.04045v1 Announce Type: new Abstract: Audio-Visual Intelligence (AVI) has emerged as a central frontier in artificial intelligence, bridging auditory and visual modalities to enable machines

Governing What the EU AI Act Excludes: Accountability for Autonomous AI Agents in Smart City Critical Infrastructure

SafetyDGX agent

arXiv:2605.01091v1 Announce Type: cross Abstract: When a traffic signal controller adjusts green phases and a grid manager curtails power on the same corridor, each system may comply with its own obli

Height Control and Optimal Torque Planning for Jumping With Wheeled-Bipedal Robots

SafetyDGX agent

arXiv:2605.03302v1 Announce Type: new Abstract: This paper mainly studies the accurate height jumping control of wheeled-bipedal robots based on torque planning and energy consumption optimization. Du

Logic-Constrained Shortest Paths for Flight Planning

SafetyDGX agent

arXiv:2412.13235v4 Announce Type: replace Abstract: The logic-constrained shortest path problem (LCSPP) combines a one-to-one shortest path problem with satisfiability constraints imposed on the routi

MILD: Mediator Agent System with Bidirectional Perception and Multi-Layered Alignment for Human-Vehicle Collaboration

SafetyDGX agent

arXiv:2605.01507v1 Announce Type: new Abstract: Prior studies report that partial driving automation can increase the cognitive demands on human drivers. This effect largely arises from human drivers'

Model Spec Midtraining: Improving How Alignment Training Generalizes

SafetyDGX agent

arXiv:2605.02087v1 Announce Type: new Abstract: Some frontier AI developers aim to align language models to a Model Spec or Constitution that describes the intended model behavior. However, standard a

NORA: A Harness-Engineered Autonomous Research Agent for End-to-End Spatial Data Science

SafetyDGX agent

arXiv:2605.02092v1 Announce Type: new Abstract: The automation of scientific research workflows has emerged as a transformative frontier in artificial intelligence, yet existing autonomous research ag

The AI risk repository: A meta-review, database, and taxonomy of risks from artificial intelligence

SafetyDGX agent

arXiv:2408.12622v3 Announce Type: replace-cross Abstract: Artificial intelligence (AI) is reshaping society, from video generation to medical diagnosis, coding agents to autonomous vehicles. Yet resea

Viewpoint-Agnostic Grasp Pipeline using VLM and Partial Observations

SafetyDGX agent

arXiv:2603.07866v2 Announce Type: replace-cross Abstract: Robust grasping in cluttered, unstructured environments remains challenging for mobile legged manipulators due to occlusions that lead to part

5 May 2026

A Deep Learning Model for Battery State Prediction towards Intelligent Energy Management

SafetyDGX agent

arXiv:2605.00898v1 Announce Type: cross Abstract: Accurate forecasting of battery health indicators, including remaining capacity and lifetime, is of paramount importance for ensuring the reliability,

AgentReputation: A Decentralized Agentic AI Reputation Framework

SafetyDGX agent

arXiv:2605.00073v1 Announce Type: new Abstract: Decentralized, agentic AI marketplaces are rapidly emerging to support software engineering tasks such as debugging, patch generation, and security audi

Artificial intelligence language technologies in multilingual healthcare: Grand challenges ahead

SafetyDGX agent

arXiv:2605.01441v1 Announce Type: new Abstract: AI language technologies (AILTs), increasingly enabled by large language models (LLMs), are becoming embedded in multilingual healthcare workflows for t

Beyond Semantic Relevance: Counterfactual Risk Minimization for Robust Retrieval-Augmented Generation

SafetyDGX agent

arXiv:2605.01302v1 Announce Type: new Abstract: Standard Retrieval-Augmented Generation (RAG) systems predominantly rely on semantic relevance as a proxy for utility. However, this assumption collapse

Experience Constrained Hierarchical Federated Reinforcement Learning for Large-scale UAV Teams in Hazardous Environments

SafetyDGX agent

arXiv:2605.02165v1 Announce Type: new Abstract: Conventional federated learning assumes that greater learner participation improves training performance, by leveraging abundant, independently generate

Hazard-Aware Traffic Scene Graph Generation

SafetyDGX agent

arXiv:2603.03584v2 Announce Type: replace Abstract: Maintaining situational awareness in complex driving scenarios is challenging. It requires continuously prioritizing attention among extensive scene

Logit-Gap Steering: A Forward-Pass Diagnostic for Alignment Robustness

SafetyDGX agent

arXiv:2506.24056v2 Announce Type: replace-cross Abstract: RLHF-style alignment trains language models to refuse unsafe requests, but how much operational margin does this refusal rest on? We introduce

Machine Learning Enhanced Laser Spectroscopy for Multi-Species Gas Detection in Complex and Harsh Environments

SafetyDGX agent

arXiv:2605.01306v1 Announce Type: cross Abstract: Laser absorption spectroscopy (LAS) is a well-established technique for non-intrusive measurement of gas species in combustion and atmospheric environ

Patient-Specific Optimization for Mandibular Reconstruction Planning with Enhanced Bone Union

SafetyDGX agent

arXiv:2605.01084v1 Announce Type: new Abstract: Mandibular reconstruction with vascularized bone grafts is complicated by donor-host nonunion, and current virtual surgical planning produces a geometri

← Previous
1…4546474849…240
Next →