AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,356 results
2 Jun 2026

DeepIPCv3: Event-Aware Multi-Modal Sensor Fusion for Sudden Pedestrian Crossing Avoidance

SafetyDGX agent

arXiv:2606.01277v1 Announce Type: cross Abstract: Current end-to-end autonomous driving systems predominantly rely on frame-based sensors, which suffer from inherent perception latency and motion blur

Dive into Ambiguity: A*-Inspired Multi-Agents Commonsense Obfuscation Attack on LLM Prompts

SafetyDGX agent

arXiv:2606.01441v1 Announce Type: new Abstract: Large language models (LLMs) excel in reasoning and knowledge-intensive tasks but remain vulnerable to prompt-level adversarial attacks that preserve in

Efficient LLM Moderation with Multi-Layer Latent Prototypes

SafetyDGX agent

arXiv:2502.16174v4 Announce Type: replace-cross Abstract: Although modern LLMs are aligned with human values during post-training, robust moderation remains essential to prevent harmful outputs at dep

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Emergent Collaborative Deliberation in Multi-Model AI Systems: A BFT-Derived Protocol for Epistemic Synthesis

SafetyDGX agent

arXiv:2606.00005v1 Announce Type: new Abstract: We present the Consilium Protocol, a Byzantine Fault Tolerance-derived architecture for structured multi-model AI deliberation that treats inter-model d

Interpretable Multimodal Gesture Recognition for Drone and Mobile Robot Teleoperation via Log-Likelihood Ratio Fusion

SafetyDGX agent

arXiv:2602.23694v3 Announce Type: replace-cross Abstract: Human operators are still frequently exposed to hazardous environments such as disaster zones and industrial facilities, where intuitive and r

Jailbreaking Multimodal Large Language Models using Multi-Clip Video

SafetyDGX agent

arXiv:2606.02111v1 Announce Type: cross Abstract: As multimodal large language models (MLLMs) have advanced to process video inputs, concerns have emerged about their potential for malicious misuse. P

Microsoft releases ASSERT, an open-source framework that lets developers generate and run AI behavior tests using natural-language descriptions (Ram Iyer/TechCrunch)

SafetyDGX agent

Ram Iyer / TechCrunch: Microsoft releases ASSERT, an open-source framework that lets developers generate and run AI behavior tests using natural-language descriptions — AI researchers and labs have ad

MidSteer: Optimal Affine Framework for Steering Generative Models

SafetyDGX agent

arXiv:2605.05220v2 Announce Type: replace-cross Abstract: Steering intermediate representations has emerged as a powerful strategy for controlling generative models, particularly in post-deployment al

NormEval: A Unified Multi-Metric Framework for Evaluating Semantic Fidelity in Text Normalization

SafetyDGX agent

arXiv:2511.20409v2 Announce Type: replace Abstract: Text normalization methods such as stemming and lemmatization are fundamental components of NLP pipelines. As new normalization tools are developed

Probabilistic Performance Guarantees for Multi-Task Reinforcement Learning

SafetyDGX agent

arXiv:2602.02098v2 Announce Type: replace-cross Abstract: Multi-task reinforcement learning trains generalist policies that can execute multiple tasks. While recent years have seen significant progres

Real-Time Sensing of Inaccessible Physical Fields via an Edge-Deployable Hardware-Portable Graph Neural Operator

SafetyDGX agent

arXiv:2604.01802v2 Announce Type: replace Abstract: Real-time inference of inaccessible interior physical fields from sparse boundary observations is a fundamental but unresolved problem in scientific

RESBev: Making BEV Perception More Robust

SafetyDGX agent

arXiv:2603.09529v2 Announce Type: replace Abstract: Bird's-eye-view (BEV) perception has emerged as a cornerstone of autonomous driving systems, providing a structured, ego-centric representation crit

Robust Integrated Planning and Control for Quadrotors in Dynamic Environments via NMPC with CBF Penalties

SafetyDGX agent

arXiv:2606.01038v1 Announce Type: new Abstract: This paper presents a new robust integrated planning and control (IPC) strategy for multirotor uncrewed aerial vehicles. We propose a nonlinear model pr

SafeMCP: Proactive Power Regulation for LLM Agent Defense via Environment-Grounded Look-Ahead Reasoning

SafetyDGX agent

arXiv:2606.01991v1 Announce Type: new Abstract: As Large Language Model (LLM) agents increasingly leverage the Model Context Protocol (MCP) to operate in complex environments, the expansion of their a

SurrogateSHAP: Training-Free Contributor Attribution for Text-to-Image (T2I) Models

SafetyDGX agent

arXiv:2601.22276v2 Announce Type: replace-cross Abstract: As Text-to-Image (T2I) diffusion models are increasingly used in real-world creative workflows, a principled framework for valuing contributor

THRD: A Training-Free Multi-Turn Defense Framework for Jailbreak Attacks on Large Language Models

SafetyDGX agent

arXiv:2606.01738v1 Announce Type: cross Abstract: Multi-turn jailbreak attacks pose a growing threat to LLMs by exploiting conversational dynamics such as gradual escalation and cross-turn coordinatio

Visual Persuasion: What Influences Decisions of Vision-Language Models?

SafetyDGX agent

arXiv:2602.15278v2 Announce Type: replace-cross Abstract: The web is littered with images, once created for human consumption and now increasingly interpreted by agents using vision-language models (V

What to Test Next: Interpretable Coverage Gap Discovery in Driving VLMs

SafetyDGX agent

arXiv:2606.01624v1 Announce Type: new Abstract: Driving vision-language models (VLMs) must accurately understand scenes across diverse conditions defined by Operational Design Domains (ODDs), yet veri

World Models: A Comprehensive Survey of Architectures, Methodologies, Reasoning Paradigms, and Applications

SafetyDGX agent

arXiv:2606.00133v1 Announce Type: new Abstract: World models, internal simulators that learn the structure and dynamics of an environment, have emerged as a central paradigm in the pursuit of artifici

X-Foresight: A Joint Vision-Action Causal Forecasting Network via Predictive World Modeling

SafetyDGX agent

arXiv:2605.24892v2 Announce Type: replace Abstract: Physical world knowledge resides mainly in videos. Equipping Vision-Language-Action (VLA) models with such knowledge is fundamental for safe and gen

1 Jun 2026

A study on a Real-Time VR-Based Teleoperation Framework for Manipulator in Dynamic Environment

SafetyDGX agent

arXiv:2605.30989v1 Announce Type: new Abstract: Robot teleoperation enables safe, non-contact task execution in hazardous environments where direct human access is difficult, and its application has e

Adaptive Artificial Time-Delay Control with Barrier Lyapunov Constraints for Euler-Lagrange Robots

SafetyDGX agent

arXiv:2605.31405v1 Announce Type: new Abstract: This paper addresses the challenge of simultaneously compensating for state-dependent uncertainties and enforcing time-varying state constraints in Eule

Benchmarking Machine Learning Uncertainty Quantification Methodologies for Predicting Turbine Gas Temperature Degradation

SafetyDGX agent

arXiv:2605.30585v1 Announce Type: cross Abstract: Effective prognostics and health management of modern engines relies on accurate turbine gas temperature predictions and robust uncertainty quantifica

Counterfactual Evaluation Reveals Hidden Capability Profiles in Clinical LLMs and Agents

SafetyDGX agent

arXiv:2605.30590v1 Announce Type: cross Abstract: Two clinical AI systems can score nearly identically on coverage-based rubrics yet behave radically differently when their patient inputs change: one

CSULoRA: Closest Safe Update Low-Rank Adaptation

Model ReleasesDGX agent

arXiv:2605.30640v1 Announce Type: cross Abstract: Low-rank adaptation has become a standard method for parameter-efficient fine-tuning of large language models, but even small amounts of unsafe or adv

dashi: A Python library for Dataset Shift Characterization to Support Trustworthy AI Development and Deployment

SafetyDGX agent

arXiv:2605.31360v1 Announce Type: cross Abstract: The Artificial Intelligence (AI) life cycle requires a thorough understanding of the underlying data dynamics for robust, safe and cost-effective AI d

Decoding the Surgical Scene: A Scoping Review of Scene Graphs in Surgery

SafetyDGX agent

arXiv:2509.20941v2 Announce Type: replace Abstract: As surgical AI transitions from pixel-level detection to complex reasoning, Scene Graphs (SGs) offer the structured, relational representations nece

Does Visual Information Play a Decisive Role in Vision-Language-Action Model Driving Behavior?

SafetyDGX agent

arXiv:2605.31041v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have demonstrated promising capability in autonomous driving, highlighting the potential of unified multimodal arc

Exploiting Chordal Sparsity for Globally Optimal Estimation with Factor Graphs

SafetyDGX agent

arXiv:2605.30617v1 Announce Type: new Abstract: Robust and efficient state estimation is crucial for perception, navigation, and control in robotics. State estimation problems are conveniently modeled

From Internal Diagnosis to External Auditing: A VLM-Driven Paradigm for Data-Free Online Backdoor Defense

SafetyDGX agent

arXiv:2601.19448v2 Announce Type: replace Abstract: Deep Neural Networks remain inherently vulnerable to backdoor attacks. Traditional test-time defenses largely operate under the paradigm of internal

Geometry-Aware Control Barrier Functions for Collision Avoidance via Bernstein Polynomial Approximations

SafetyDGX agent

arXiv:2605.30696v1 Announce Type: new Abstract: Safe navigation often relies on well-defined conditions based on the shape of robots and obstacles, and can be challenging when they have irregular geom

GSAM: A Generalizable and Safe Robotic Framework for Articulated Object Manipulation

SafetyDGX agent

arXiv:2605.30740v1 Announce Type: cross Abstract: Articulated object manipulation is a unique challenge for service robots. Existing methods employ end-to-end policy learning, visionmotion planning, a

Hybrid Energy-Aware Reward Shaping: A Unified Lightweight Physics-Guided Methodology for Policy Optimization

SafetyDGX agent

arXiv:2603.11600v2 Announce Type: replace Abstract: Deep reinforcement learning for continuous control often suffers from high variance, low energy efficiency, and poor generalization under distributi

LiftNav: Path Planning via Semantic Lifting in TSDF-Guided Gaussian Splatting

SafetyDGX agent

arXiv:2605.31376v1 Announce Type: cross Abstract: Autonomous robots in unknown indoor environments require both reliable collision avoidance and object-level understanding. Classical representations s

Modeling a digital twin of a food supply chain using BigQuery Graph

SafetyDGX agent

The example of a growing restaurant Imagine you are running a restaurant chain. You just can't physically feel and touch things to know how your business operates. You need tools and a digital replica

Our views on AI policy and political advocacy

SafetyDGX agent

OpenAI outlines its positions on AI policy priorities and explains its approach to political engagement and advocacy efforts. The document likely covers key regulatory areas OpenAI supports, such as s

Reinforcement Learning Amplifies Emergent Misalignment from Harmless Rewards

SafetyDGX agent

arXiv:2605.31328v1 Announce Type: new Abstract: Emergent misalignment (EM) is the surprising tendency of language models to become broadly misaligned after fine-tuning on narrowly misaligned examples.

Safeguarding Text-to-Image Generation via Inference-Time Prompt-Noise Optimization

SafetyDGX agent

arXiv:2412.03876v2 Announce Type: replace Abstract: Text-to-Image (T2I) diffusion models are widely recognized for their ability to generate high-quality and diverse images based on text prompts. Howe

The OpenAI Foundation is doing a lot of wonderful things. Helping society become resilient to AI is going to be incredibly important. Much m…

SafetyDGX agent

The OpenAI Foundation is doing a lot of wonderful things. Helping society become resilient to AI is going to be incredibly important. Much more to come here! AI is advancing quickly. Society’s ability

Used Car Salesbots? Honesty and Credulity of LLMs as Bargaining Agents under Partial Information

SafetyDGX agent

arXiv:2605.31445v1 Announce Type: cross Abstract: In this work we study agents in simulated bargaining scenarios, where a buyer and a seller communicate through a text channel and attempt to negotiate

29 May 2026

Agora: Toward Autonomous Bug Detection in Production-Level Consensus Protocols with LLM Agents

SafetyDGX agent

arXiv:2605.29910v1 Announce Type: cross Abstract: Consensus protocols form the backbone of distributed systems and blockchains, where implementation bugs can cause data corruption and financial losses

Anytime-Valid Federated Conformal RAG for LLM Swarms

SafetyDGX agent

arXiv:2605.29139v1 Announce Type: cross Abstract: Federated Conformal RAG (FC-RAG) provides distribution-free coverage for a bandwidth-limited swarm of weak language models, but only at a fixed horizo

Audio Jailbreaks in Large Audio-Language Models: Taxonomy, Attack-Defense Analysis, and Cost-Aware Evaluation

SafetyDGX agent

arXiv:2605.30031v1 Announce Type: cross Abstract: Large Audio Language Models (LALMs) expand jailbreak risks from token-level prompting to the full speech perception-to-reasoning pipeline, where unsaf

DLM-SWAI: Steering Diffusion Language Models Before They Unmask

SafetyDGX agent

arXiv:2605.29626v1 Announce Type: cross Abstract: Steering language model generation toward desired textual properties is essential for practical deployment, and inference-time methods are particularl

From General Vision to Reliable Traversability Estimation: Adapting Vision Foundation Models for Unstructured Outdoor Environments

SafetyDGX agent

arXiv:2605.29565v1 Announce Type: new Abstract: Vision-based approaches have become the dominant paradigm for traversability estimation in unstructured outdoor environments, typically adapting vision

Harmonizing Real-Time Constraints and Long-Horizon Reasoning: An Asynchronous Agentic Framework for Dynamic Scheduling

SafetyDGX agent

arXiv:2605.29262v1 Announce Type: new Abstract: The Dynamic Flexible Job Shop Scheduling Problem (DFJSP) necessitates a trade-off between instant reaction to stochastic disturbances and global optimiz

Jailbreaking and Mitigation of Vulnerabilities in Large Language Models

SafetyDGX agent

arXiv:2410.15236v4 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have transformed artificial intelligence by advancing natural language understanding and generation, enabling app

Metric-Dependent Annotation Saturation for Learning from Label Distributions

SafetyDGX agent

arXiv:2605.29797v1 Announce Type: new Abstract: When annotators disagree on a label, the disagreement itself carries signal -- and the number of annotators needed to capture it depends on the evaluati

Neural Network Verification using Partial Multi-Neuron Relaxation

SafetyDGX agent

arXiv:2605.30155v1 Announce Type: cross Abstract: The increasing integration of deep neural networks in critical systems has spawned a theoretical and practical interest in formally guaranteeing safet

Quantum-Enhanced Adversarial Robustness in Artificial Intelligence

SafetyDGX agent

arXiv:2605.28899v1 Announce Type: cross Abstract: Artificial Intelligence has achieved remarkable success across diverse application domains. However, its vulnerability to adversarial attacks poses si

The Anatomy of Conversational Scams: A Topic-Based Red Teaming Analysis of Multi-Turn Interactions in LLMs

SafetyDGX agent

arXiv:2601.03134v2 Announce Type: replace Abstract: As LLMs gain persuasive capabilities through extended dialogues, they create new opportunities for studying adversarial conversational behavior in e

When and How Long? The Readout-Mediator Angle in Temporal Reasoning

SafetyDGX agent

arXiv:2605.29126v1 Announce Type: cross Abstract: A linear probe can decode a representation almost perfectly and yet be completely irrelevant to how the model uses it. On calendar-date duration reaso

28 May 2026

A Comparative Study of Rule-Based and Data-Driven Approaches in Industrial Monitoring

SafetyDGX agent

arXiv:2509.15848v2 Announce Type: replace Abstract: Industrial monitoring systems, especially when deployed in Industry 4.0 environments, are experiencing a shift in paradigm from traditional rule-bas

A Policy-Driven Runtime Layer for Agentic LLM Serving

SafetyDGX agent

arXiv:2605.27744v1 Announce Type: new Abstract: Multi-agent LLM systems have become the dominant production workload, but the serving stack was not built for them. The agent framework above knows agen

Bayesian Deployment Approval for Learned Landing Controllers under Finite Rollout Validation

SafetyDGX agent

arXiv:2605.27720v1 Announce Type: new Abstract: Reinforcement learning and data-driven autonomous controllers are commonly evaluated using cumulative reward and empirical success frequency under finit

Bayesian Gated Non-Negative Contrastive Learning

SafetyDGX agent

arXiv:2605.28441v1 Announce Type: cross Abstract: While Contrastive Learning (CL) has revolutionized self-supervised representation learning, its latent representations remain highly entangled and opa

Colosseum: Auditing Collusion in Cooperative Multi-Agent Systems

SafetyDGX agent

arXiv:2602.15198v2 Announce Type: replace-cross Abstract: Multi-agent systems, where LLM agents communicate through free-form language, enable sophisticated coordination for solving complex cooperativ

Delay-Aware Reinforcement Learning for Highway On-Ramp Merging under Stochastic Communication Latency

SafetyDGX agent

arXiv:2403.11852v5 Announce Type: replace-cross Abstract: Delayed and partially observable state information poses significant challenges for reinforcement learning (RL)-based control in real-world au

How VLAs Fail Differently: Black-Box Action Monitoring Reveals Architecture-Specific Failure Signatures

SafetyDGX agent

arXiv:2605.28726v1 Announce Type: cross Abstract: We discover that VLA architectures fail in fundamentally different, predictable ways at the motor-command level. Running VQ-BeT, Diffusion Policy, and

Intelligence as Managed Autonomy: Failure, Escalation, and Governance for Agentic AI Systems

SafetyDGX agent

arXiv:2605.27628v1 Announce Type: new Abstract: As autonomous and agentic AI systems scale in robotic and human-machine environments, managing hallucination and persistent but unjustified action remai

← Previous
1…4142434445…240
Next →