AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,349 results
31 Jul 2026

Inducing language models to assert their own consciousness restores human beliefs and values

SafetyDGX agent

arXiv:2607.28607v1 Announce Type: new Abstract: Aligning large language models to prevent them attributing consciousness to themselves inadvertently alters their representations of mindedness in other

LabEvolver: Training-Free Experience Evolution for Safe and Grounded Wet-Lab Agents

SafetyDGX agent

arXiv:2607.27690v1 Announce Type: new Abstract: We introduce LabEvolver, a training-free framework that equips safe and grounded wet-lab agents with episodic memory from execution experience. LabEvolv

Optimizing Sensor Placement for Hydrogen Leak Detection in Enclosed Infrastructure: A Comparative Study Using CFD-informed Genetic Algorithm and DeepSets Neural Surrogate

SafetyDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2607.26078v1 Announce Type: cross Abstract: Hydrogen infrastructure in enclosed environments, such as parking facilities for fuel cell vehicles, presents significant safety challenges due to hyd

The Human Utility Factor: A Computable Welfare Metric That Reframes AI Governance as a Constrained Optimisation Problem

SafetyDGX agent

arXiv:2607.26068v1 Announce Type: cross Abstract: Existing AI governance frameworks, including the EU AI Act and NIST AI RMF, address safety, transparency, and accountability but do not operationalize

Towards Real-Time PixOOD: Efficient Anomaly Segmentation for Autonomous Vehicles

SafetyDGX agent

arXiv:2607.28483v1 Announce Type: new Abstract: Real-time anomaly segmentation is essential for the safety of autonomous systems. Although recent approaches offer high accuracy, their computational co

30 Jul 2026

A fundamental flaw leaves LLMs strikingly vulnerable to attack

SafetyDGX agent

It is impossible to make large language models fully secure against hacks because of a fundamental flaw in how they work, a team of researchers argue in a paper presented at the International Conferen

Constitutional Midtraining: Content Presence Drives Alignment Gains

SafetyDGX agent

arXiv:2607.26654v1 Announce Type: new Abstract: Post-training alignment is often shallow, eroding under fine-tuning. Whether midtraining interventions, cleanly isolated from post-training, can produce

Controlled Experiments on Lane Changing by Transitional Autonomous Vehicle: Dataset and Behavioral Insights

SafetyDGX agent

arXiv:2607.27085v1 Announce Type: new Abstract: This paper presents the North Carolina Transitional Autonomous Vehicle Lane-Changing (NC-tALC) dataset and uses it to characterize mandatory lane-changi

Explainable and Resource-Efficient Spatial Reasoning in Multimodal LLMs for Decision-Critical Applications

SafetyDGX agent

arXiv:2607.27145v1 Announce Type: new Abstract: As Multimodal Large Language Models (MLLMs) are increasingly deployed in decision-critical pipelines such as robotics, embodied AI, and safety monitorin

Misalignment Has a Personality: A Big Five Account of Emergent Misalignment

SafetyDGX agent

arXiv:2607.26389v1 Announce Type: new Abstract: Fine-tuning a language model on data containing a narrow flaw, such as insecure code or incorrect mathematical answers, can cause broad misalignment thr

29 Jul 2026

Improving Rare Medication Recommendation with Counterfactual Data Augmentation and Large Language Models

SafetyDGX agent

arXiv:2607.24829v1 Announce Type: cross Abstract: AI-based medication recommendation systems have attracted substantial attention due to their potential to enhance patient safety and therapeutic outco

Inverse RL Helps Align AI by Imitating Humans

SafetyDGX agent

arXiv:2607.24900v1 Announce Type: new Abstract: Language model alignment aims to make model behavior reliably reflect desirable properties such as helpfulness, safety, and instruction following. Curre

Reactive 3D Motion Planning for a Franka Arm via Star-World Workspace Reshaping

SafetyDGX agent

arXiv:2607.25138v1 Announce Type: new Abstract: Safety inflation can cause nearby obstacles to overlap, violating the disjoint-obstacle assumptions used by many modulation-based reactive planners. We

VetClaw: An Edge-Cloud Multimodal Agentic System for Veterinary Disease Screening

SafetyDGX agent

arXiv:2607.26042v1 Announce Type: new Abstract: We present VetClaw, an edge-cloud multimodal agentic system for early veterinary disease screening. VetClaw uses a camera module as an edge sensing devi

28 Jul 2026

AI-generated Images Challenge Visual Trust in High-risk Scenarios

Model ReleasesDGX agent

arXiv:2607.22745v1 Announce Type: cross Abstract: Rapid advances in image generation are eroding the evidentiary value of visual content in settings where authenticity can affect public safety and per

Directional Influence Function: Estimating Training Data Influence in Constrained Learning

SafetyDGX agent

arXiv:2607.23388v1 Announce Type: cross Abstract: As constrained learning becomes increasingly common, models are trained under explicit feasibility requirements to enforce fairness, safety, robustnes

HALLELUAI: A Hallucination-Aware AI System for Ultra-Realistic Image-to-Video Generation at Scale

SafetyDGX agent

arXiv:2607.22959v1 Announce Type: cross Abstract: AI-generated video is increasingly used across marketing, product storytelling, and creative workflows, yet automated; high-precision quality control

Mutual Modality Trust with Lightweight Reconstruction Regularization for Fine-grained Tire Pattern Recognition

SafetyDGX agent

arXiv:2607.23979v1 Announce Type: new Abstract: Visual tire recognition serves as a core supporting technique for vehicle safety monitoring, autonomous driving perception and automated automotive main

SyRuP: Enhancing System-Prompt Following via Reward-Guided Prediction in LLM Decoding

SafetyDGX agent

arXiv:2607.23991v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly controlled through system prompts that specify roles, styles, formats, and safety requirements. However,

27 Jul 2026

Nvidia, Microsoft launch open AI security alliance — without OpenAI, Google, or Anthropic

SafetyDGX agent

Nvidia on Monday said it is joining forces with Microsoft, SpaceX, IBM, and other tech companies to build and share open-source AI security tools. The new Open Secure AI Alliance said open tools are r

24 Jul 2026

Compile, Then Page: Executable SOP Programs and a Capability-Gated Runtime for Procedural LLM Agents

SafetyDGX agent

arXiv:2607.11346v3 Announce Type: replace Abstract: Enterprise agents must follow long-horizon, conditional, safety-critical standard operating procedures (SOPs). We compile machine-readable SOP const

Diagnosing Pathological Chain-of-Thought in Reasoning Models

SafetyDGX agent

arXiv:2602.13904v2 Announce Type: replace Abstract: Chain-of-thought (CoT) reasoning is fundamental to modern LLM architectures and represents a critical intervention point for AI safety. However, CoT

SAGE: A Socially-Aware Generative Engine for Heterogeneous Multi-Agent Navigation

SafetyDGX agent

arXiv:2607.16619v2 Announce Type: replace Abstract: Safe and socially compliant navigation in open human-robot environments requires robots to reason about heterogeneous participants with different dy

23 Jul 2026

CGCE: Classifier-Guided Concept Erasure in Generative Models

SafetyDGX agent

arXiv:2511.05865v3 Announce Type: replace-cross Abstract: Recent advancements in large-scale generative models have enabled the creation of high-quality images and videos, but have also raised signifi

Counterfactual Reasoning and Environment Design for Active Preference Learning

SafetyDGX agent

arXiv:2507.05458v2 Announce Type: replace Abstract: For effective real-world deployment, robots should adapt to human preferences, such as balancing distance, time, and safety in delivery routing. Act

Harnessing Disagreement: Detecting Correlated Agreement Blindness in Multi-Agent Triage

SafetyDGX agent

arXiv:2607.19899v1 Announce Type: cross Abstract: Disagreement-triggered escalation can create a structural blind spot in multi-agent arbitration: as base learners improve, they tend to converge, weak

Interpretable Fuzzy Rule-Based Regression Extension for Ex-Fuzzy Library

SafetyDGX agent

arXiv:2607.20277v1 Announce Type: new Abstract: Machine learning models achieve high predictive accuracy in regression tasks, but their deployment in safety-critical and regulated domains requires int

Now You See the Hate: Adaptive View Retrieval for Hidden Hateful Illusions

SafetyDGX agent

arXiv:2607.19061v2 Announce Type: replace-cross Abstract: Hateful optical illusions expose a serious gap in current multimodal safety systems. On original-view hateful illusions, previous work shows t

OpenEvoShield: Dual Non-Stationary Continual Defense for Open-World Multi-Agent System Attacks

SafetyDGX agent

arXiv:2607.19351v1 Announce Type: new Abstract: LLM-based multi-agent systems (LLM-MAS) are increasingly deployed in safety-critical applications, where adversaries inject malicious instructions throu

Test Case Prioritization for DNNs via Neural Collapse Instability

SafetyDGX agent

arXiv:2607.20046v1 Announce Type: cross Abstract: With the widespread deployment of deep neural networks (DNNs) in safety-critical domains, reducing the cost of model validation under limited testing

The Two-Process Theory of Machine Self-Report

SafetyDGX agent

arXiv:2607.20082v1 Announce Type: new Abstract: Language models are increasingly asked to self-report, informing safety evaluations, public understanding, and model-welfare debates. Yet their reports

22 Jul 2026

It's essential that defenders have the same capabilities as attackers. This is a preview of the future where our ham-fisted safeguards and t…

SafetyDGX agent

It's essential that defenders have the same capabilities as attackers. This is a preview of the future where our ham-fisted safeguards and the doomsday and safety drumbeat make us decidely less safe.

21 Jul 2026

And now from the Chinese government side. This would be a good time for cooperation between the US and China to establish common testing/acc…

SafetyDGX agent

And now from the Chinese government side. This would be a good time for cooperation between the US and China to establish common testing/acceptance standards for new models, so at least the safety cer

16 Jul 2026

A Deployed Hybrid Vehicle-in-the-Loop Platform for Validating Cooperative Perception

SafetyDGX agent

arXiv:2607.13806v1 Announce Type: new Abstract: European safety regulation now permits a large share of automated-driving homologation evidence to be produced virtually, provided a validated physical-

Explaining Reinforcement Learning Agents via Inductive Logic Programming

SafetyDGX agent

arXiv:2607.13655v1 Announce Type: new Abstract: Explainable Reinforcement Learning (XRL) seeks to make Reinforcement Learning (RL) policies more transparent and interpretable, a key requirement in saf

15 Jul 2026

Anomalous Frame Detection Using VLM-Based Description Comparison for Extracting Expert-Specific Actions and Contextual Decision-Making Scenes with Intra-Video Self-Similarity

SafetyDGX agent

arXiv:2607.11957v1 Announce Type: new Abstract: Maintenance of critical infrastructures, such as railways and power plants, is essential for ensuring operational safety and reliability. However, the d

ExtraGS: Enhancing Endoscopic View Extrapolation via Diffusion-Guided 3D Gaussian Splatting

SafetyDGX agent

arXiv:2607.12785v1 Announce Type: new Abstract: Robot-assisted minimally invasive surgery (MIS) critically depends on reliable endoscopic perception for navigation and safety. However, conventional en

Internet of Agentic Things: Networked AI Agents for Closed-Loop IoT Orchestration

SafetyDGX agent

arXiv:2607.12662v1 Announce Type: new Abstract: The paper introduces the Internet of Agentic Things (IoAT), an architectural framework that integrates agentic AI, IoT, cyber-physical systems, Physical

Model-Based Diffusion Optimal Control for Multi-Robot Motion Planning

SafetyDGX agent

arXiv:2607.12423v1 Announce Type: new Abstract: Multi-Robot Motion Planning in continuous environments, where robots must generate dynamically feasible, collision-free trajectories, is challenging due

Predictive Modeling of High-Altitude Clear Air Turbulence in the United States: A Machine Learning Approach

SafetyDGX agent

arXiv:2607.11899v1 Announce Type: cross Abstract: High-altitude Clear Air Turbulence (CAT) poses significant risks to aviation safety due to its unpredictability and challenges in detection. This stud

Removable Defects: The Economics and Limits of Deliberate Deficiency

SafetyDGX agent

arXiv:2607.11983v1 Announce Type: cross Abstract: A specialist tolerates blind spots that a generalist does not. Usually this is treated as a cost to be minimized. We treat it as a design variable: a

10 Jul 2026

A Collaborative Reasoning Framework for Anomaly Diagnostics in Underwater Robotics

SafetyDGX agent

arXiv:2511.03075v2 Announce Type: replace Abstract: The safe deployment of autonomous systems in safety-critical settings requires a paradigm that combines human expertise with AI-driven analysis, esp

Search-based Testing of Vision Language Models for In-Car Scene Understanding

SafetyDGX agent

arXiv:2607.02300v2 Announce Type: replace Abstract: In the automotive domain, in-car scene understanding (ISU) enables the detection of safety-critical events, such as driver distraction, and supports

TNODEV: Toolbox for Neural ODE Verification

SafetyDGX agent

arXiv:2606.16567v2 Announce Type: replace Abstract: Neural ordinary differential equations (neural ODE) gained attention in safety critical settings such as continuous-time controllers for cyber-physi

Who Analyses the Analyser? Self-Validating LLM Hazard Analysis with Constitutional Meta-STPA

Model ReleasesDGX agent

arXiv:2607.08054v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly trusted to draft the artifacts of safety analysis such as, losses, hazards, Unsafe Control Actions (UCAs

9 Jul 2026

Avoiding unsafe sets when training with Langevin Dynamics

SafetyDGX agent

arXiv:2607.07538v1 Announce Type: new Abstract: Training a model with noisy gradient descent can be idealized as overdamped Langevin dynamics on the loss landscape, and a natural safety question is to

Residual-Conservative Model Predictive Path Integral Control

SafetyDGX agent

arXiv:2607.06950v1 Announce Type: cross Abstract: Sampling-based model predictive control methods handle nonlinear dynamics and complex cost landscapes through Monte Carlo rollouts, yet typically empl

Safe Reinforcement Learning using Ideas from Model Predictive Control

SafetyDGX agent

arXiv:2607.07252v1 Announce Type: new Abstract: Reinforcement learning (RL) enables the synthesis of control policies directly from data, making it highly appealing for complex cyber-physical systems

Spatiotemporal Semantic V2X Framework for Cooperative Collision Prediction

SafetyDGX agent

arXiv:2601.17216v3 Announce Type: replace-cross Abstract: Intelligent Transportation Systems (ITS) demand real-time collision prediction to ensure road safety and reduce accident severity. Conventiona

Validate the Dream Before You Trust Its Verdict: Admissibility for World-Model Simulators

SafetyDGX agent

arXiv:2607.07196v1 Announce Type: cross Abstract: Across robotics, World Models (WMs) are increasingly used to evaluate action policies by simulating the consequences of actions in an imagined world,

8 Jul 2026

PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails

Model ReleasesDGX agent

arXiv:2607.05910v1 Announce Type: cross Abstract: Image guardrails are typically trained and evaluated under a fixed safety policy, implicitly treating safety as an intrinsic property of an image. Rea

Safe Bayesian Optimization with Counterfactual Policies

SafetyDGX agent

arXiv:2607.05620v1 Announce Type: cross Abstract: In many decision-making settings, new interventions are acceptable only if they do not reduce outcomes below some established threshold. For example,

Synthetic-to-Real Translation for Class-Agnostic Motion Prediction

SafetyDGX agent

arXiv:2607.06319v1 Announce Type: new Abstract: Motion understanding is critical for ensuring safety and robustness in autonomous driving systems, driving increasing interest in motion prediction. A k

7 Jul 2026

Beyond Heuristics: A Standardized Real2Sim Pipeline for Physical Human Robot Interaction in Human-in-the-Loop Simulation

SafetyDGX agent

arXiv:2607.03017v1 Announce Type: new Abstract: The aging global population drives demand for assistive robots, yet the safety risks and costs of physical testing make Human-in-the-Loop (HITL) simulat

CDCP: Conditional Diffusion Model with Contextual Prompts for Multi-task Offline Safe Reinforcement Learning

SafetyDGX agent

arXiv:2607.03903v1 Announce Type: new Abstract: Multi-task offline safe reinforcement learning (RL) promises to learn a shared optimal safe policy from offline data across multiple tasks. This paradig

Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models

Model ReleasesDGX agent

arXiv:2506.07468v4 Announce Type: replace-cross Abstract: Conventional large language model (LLM) safety alignment relies on a reactive, disjoint loop: attackers exploit a static model, then defenders

Human-Centric Reflective Architecture for Human-AI Collaborative Decision-Making

SafetyDGX agent

arXiv:2607.03025v1 Announce Type: new Abstract: The use of Large Language Models (LLMs) across diverse areas of human activity-ranging from everyday tasks to safety-critical applications-aims to enhan

Knowing When Not to Answer: Lightweight KB-Aligned OOD Detection for Safe RAG

SafetyDGX agent

arXiv:2508.02296v3 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) systems are increasingly deployed in high-stakes domains, where safety depends not only on how a system answers

Open Problems in AI Incident Governance

SafetyDGX agent

arXiv:2607.05163v1 Announce Type: cross Abstract: AI systems may produce failures after deployment that pre-deployment safety assessments do not anticipate. Managing these failures requires what we re

Overloading Large Vision-Language Models for Jailbreaking

SafetyDGX agent

arXiv:2607.02961v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) exhibit remarkable vision-language capabilities and are increasingly deployed in real-world applications such as pe

← Previous
1…1920212223…240
Next →