AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,357 results
Safety

Counterfactual Evaluation Reveals Hidden Capability Profiles in Clinical LLMs and Agents

DGX agent

arXiv:2605.30590v1 Announce Type: cross Abstract: Two clinical AI systems can score nearly identically on coverage-based rubrics yet behave radically differently when their patient inputs change: one

safetyarxiv-cs-ai
1 Jun 2026
Model Releases
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

CSULoRA: Closest Safe Update Low-Rank Adaptation

DGX agent

arXiv:2605.30640v1 Announce Type: cross Abstract: Low-rank adaptation has become a standard method for parameter-efficient fine-tuning of large language models, but even small amounts of unsafe or adv

model-releasesarxiv-cs-cl
1 Jun 2026
Safety

dashi: A Python library for Dataset Shift Characterization to Support Trustworthy AI Development and Deployment

DGX agent

arXiv:2605.31360v1 Announce Type: cross Abstract: The Artificial Intelligence (AI) life cycle requires a thorough understanding of the underlying data dynamics for robust, safe and cost-effective AI d

safetyarxiv-cs-ai
1 Jun 2026
Safety

Decoding the Surgical Scene: A Scoping Review of Scene Graphs in Surgery

DGX agent

arXiv:2509.20941v2 Announce Type: replace Abstract: As surgical AI transitions from pixel-level detection to complex reasoning, Scene Graphs (SGs) offer the structured, relational representations nece

safetyarxiv-cs-cv
1 Jun 2026
Safety

Does Visual Information Play a Decisive Role in Vision-Language-Action Model Driving Behavior?

DGX agent

arXiv:2605.31041v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have demonstrated promising capability in autonomous driving, highlighting the potential of unified multimodal arc

safetyarxiv-cs-ai
1 Jun 2026
Safety

Exploiting Chordal Sparsity for Globally Optimal Estimation with Factor Graphs

DGX agent

arXiv:2605.30617v1 Announce Type: new Abstract: Robust and efficient state estimation is crucial for perception, navigation, and control in robotics. State estimation problems are conveniently modeled

safetyarxiv-cs-ro
1 Jun 2026
Safety

From Internal Diagnosis to External Auditing: A VLM-Driven Paradigm for Data-Free Online Backdoor Defense

DGX agent

arXiv:2601.19448v2 Announce Type: replace Abstract: Deep Neural Networks remain inherently vulnerable to backdoor attacks. Traditional test-time defenses largely operate under the paradigm of internal

safetyarxiv-cs-lg
1 Jun 2026
Safety

Geometry-Aware Control Barrier Functions for Collision Avoidance via Bernstein Polynomial Approximations

DGX agent

arXiv:2605.30696v1 Announce Type: new Abstract: Safe navigation often relies on well-defined conditions based on the shape of robots and obstacles, and can be challenging when they have irregular geom

safetyarxiv-cs-ro
1 Jun 2026
Safety

GSAM: A Generalizable and Safe Robotic Framework for Articulated Object Manipulation

DGX agent

arXiv:2605.30740v1 Announce Type: cross Abstract: Articulated object manipulation is a unique challenge for service robots. Existing methods employ end-to-end policy learning, visionmotion planning, a

safetyarxiv-cs-ai
1 Jun 2026
Safety

Hybrid Energy-Aware Reward Shaping: A Unified Lightweight Physics-Guided Methodology for Policy Optimization

DGX agent

arXiv:2603.11600v2 Announce Type: replace Abstract: Deep reinforcement learning for continuous control often suffers from high variance, low energy efficiency, and poor generalization under distributi

safetyarxiv-cs-lg
1 Jun 2026
Safety

LiftNav: Path Planning via Semantic Lifting in TSDF-Guided Gaussian Splatting

DGX agent

arXiv:2605.31376v1 Announce Type: cross Abstract: Autonomous robots in unknown indoor environments require both reliable collision avoidance and object-level understanding. Classical representations s

safetyarxiv-cs-cv
1 Jun 2026
Safety

Modeling a digital twin of a food supply chain using BigQuery Graph

DGX agent

The example of a growing restaurant Imagine you are running a restaurant chain. You just can't physically feel and touch things to know how your business operates. You need tools and a digital replica

safetygoogle-cloud-ai
1 Jun 2026
Safety

Our views on AI policy and political advocacy

DGX agent

OpenAI outlines its positions on AI policy priorities and explains its approach to political engagement and advocacy efforts. The document likely covers key regulatory areas OpenAI supports, such as s

safetyopenai
1 Jun 2026
Safety

Reinforcement Learning Amplifies Emergent Misalignment from Harmless Rewards

DGX agent

arXiv:2605.31328v1 Announce Type: new Abstract: Emergent misalignment (EM) is the surprising tendency of language models to become broadly misaligned after fine-tuning on narrowly misaligned examples.

safetyarxiv-cs-cl
1 Jun 2026
Safety

Safeguarding Text-to-Image Generation via Inference-Time Prompt-Noise Optimization

DGX agent

arXiv:2412.03876v2 Announce Type: replace Abstract: Text-to-Image (T2I) diffusion models are widely recognized for their ability to generate high-quality and diverse images based on text prompts. Howe

safetyarxiv-cs-cv
1 Jun 2026
Safety

The OpenAI Foundation is doing a lot of wonderful things. Helping society become resilient to AI is going to be incredibly important. Much m…

DGX agent

The OpenAI Foundation is doing a lot of wonderful things. Helping society become resilient to AI is going to be incredibly important. Much more to come here! AI is advancing quickly. Society’s ability

safetysam-altman--x
1 Jun 2026
Safety

Used Car Salesbots? Honesty and Credulity of LLMs as Bargaining Agents under Partial Information

DGX agent

arXiv:2605.31445v1 Announce Type: cross Abstract: In this work we study agents in simulated bargaining scenarios, where a buyer and a seller communicate through a text channel and attempt to negotiate

safetyarxiv-cs-ai
1 Jun 2026
Safety

Agora: Toward Autonomous Bug Detection in Production-Level Consensus Protocols with LLM Agents

DGX agent

arXiv:2605.29910v1 Announce Type: cross Abstract: Consensus protocols form the backbone of distributed systems and blockchains, where implementation bugs can cause data corruption and financial losses

safetyarxiv-cs-ai
29 May 2026
Safety

Anytime-Valid Federated Conformal RAG for LLM Swarms

DGX agent

arXiv:2605.29139v1 Announce Type: cross Abstract: Federated Conformal RAG (FC-RAG) provides distribution-free coverage for a bandwidth-limited swarm of weak language models, but only at a fixed horizo

safetyarxiv-cs-lg
29 May 2026
Safety

Audio Jailbreaks in Large Audio-Language Models: Taxonomy, Attack-Defense Analysis, and Cost-Aware Evaluation

DGX agent

arXiv:2605.30031v1 Announce Type: cross Abstract: Large Audio Language Models (LALMs) expand jailbreak risks from token-level prompting to the full speech perception-to-reasoning pipeline, where unsaf

safetyarxiv-cs-ai
29 May 2026
Safety

DLM-SWAI: Steering Diffusion Language Models Before They Unmask

DGX agent

arXiv:2605.29626v1 Announce Type: cross Abstract: Steering language model generation toward desired textual properties is essential for practical deployment, and inference-time methods are particularl

safetyarxiv-cs-ai
29 May 2026
Safety

From General Vision to Reliable Traversability Estimation: Adapting Vision Foundation Models for Unstructured Outdoor Environments

DGX agent

arXiv:2605.29565v1 Announce Type: new Abstract: Vision-based approaches have become the dominant paradigm for traversability estimation in unstructured outdoor environments, typically adapting vision

safetyarxiv-cs-cv
29 May 2026
Safety

Harmonizing Real-Time Constraints and Long-Horizon Reasoning: An Asynchronous Agentic Framework for Dynamic Scheduling

DGX agent

arXiv:2605.29262v1 Announce Type: new Abstract: The Dynamic Flexible Job Shop Scheduling Problem (DFJSP) necessitates a trade-off between instant reaction to stochastic disturbances and global optimiz

safetyarxiv-cs-ai
29 May 2026
Safety

Jailbreaking and Mitigation of Vulnerabilities in Large Language Models

DGX agent

arXiv:2410.15236v4 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have transformed artificial intelligence by advancing natural language understanding and generation, enabling app

safetyarxiv-cs-ai
29 May 2026
Safety

Metric-Dependent Annotation Saturation for Learning from Label Distributions

DGX agent

arXiv:2605.29797v1 Announce Type: new Abstract: When annotators disagree on a label, the disagreement itself carries signal -- and the number of annotators needed to capture it depends on the evaluati

safetyarxiv-cs-cl
29 May 2026
Safety

Neural Network Verification using Partial Multi-Neuron Relaxation

DGX agent

arXiv:2605.30155v1 Announce Type: cross Abstract: The increasing integration of deep neural networks in critical systems has spawned a theoretical and practical interest in formally guaranteeing safet

safetyarxiv-cs-ai
29 May 2026
Safety

Quantum-Enhanced Adversarial Robustness in Artificial Intelligence

DGX agent

arXiv:2605.28899v1 Announce Type: cross Abstract: Artificial Intelligence has achieved remarkable success across diverse application domains. However, its vulnerability to adversarial attacks poses si

safetyarxiv-cs-ai
29 May 2026
Safety

The Anatomy of Conversational Scams: A Topic-Based Red Teaming Analysis of Multi-Turn Interactions in LLMs

DGX agent

arXiv:2601.03134v2 Announce Type: replace Abstract: As LLMs gain persuasive capabilities through extended dialogues, they create new opportunities for studying adversarial conversational behavior in e

safetyarxiv-cs-cl
29 May 2026
Safety

When and How Long? The Readout-Mediator Angle in Temporal Reasoning

DGX agent

arXiv:2605.29126v1 Announce Type: cross Abstract: A linear probe can decode a representation almost perfectly and yet be completely irrelevant to how the model uses it. On calendar-date duration reaso

safetyarxiv-cs-ai
29 May 2026
Safety

A Comparative Study of Rule-Based and Data-Driven Approaches in Industrial Monitoring

DGX agent

arXiv:2509.15848v2 Announce Type: replace Abstract: Industrial monitoring systems, especially when deployed in Industry 4.0 environments, are experiencing a shift in paradigm from traditional rule-bas

safetyarxiv-cs-ai
28 May 2026
Safety

A Policy-Driven Runtime Layer for Agentic LLM Serving

DGX agent

arXiv:2605.27744v1 Announce Type: new Abstract: Multi-agent LLM systems have become the dominant production workload, but the serving stack was not built for them. The agent framework above knows agen

safetyarxiv-cs-ai
28 May 2026
Safety

Bayesian Deployment Approval for Learned Landing Controllers under Finite Rollout Validation

DGX agent

arXiv:2605.27720v1 Announce Type: new Abstract: Reinforcement learning and data-driven autonomous controllers are commonly evaluated using cumulative reward and empirical success frequency under finit

safetyarxiv-cs-lg
28 May 2026
Safety

Bayesian Gated Non-Negative Contrastive Learning

DGX agent

arXiv:2605.28441v1 Announce Type: cross Abstract: While Contrastive Learning (CL) has revolutionized self-supervised representation learning, its latent representations remain highly entangled and opa

safetyarxiv-cs-ai
28 May 2026
Safety

Colosseum: Auditing Collusion in Cooperative Multi-Agent Systems

DGX agent

arXiv:2602.15198v2 Announce Type: replace-cross Abstract: Multi-agent systems, where LLM agents communicate through free-form language, enable sophisticated coordination for solving complex cooperativ

safetyarxiv-cs-ai
28 May 2026
Safety

Delay-Aware Reinforcement Learning for Highway On-Ramp Merging under Stochastic Communication Latency

DGX agent

arXiv:2403.11852v5 Announce Type: replace-cross Abstract: Delayed and partially observable state information poses significant challenges for reinforcement learning (RL)-based control in real-world au

safetyarxiv-cs-ai
28 May 2026
Safety

How VLAs Fail Differently: Black-Box Action Monitoring Reveals Architecture-Specific Failure Signatures

DGX agent

arXiv:2605.28726v1 Announce Type: cross Abstract: We discover that VLA architectures fail in fundamentally different, predictable ways at the motor-command level. Running VQ-BeT, Diffusion Policy, and

safetyarxiv-cs-lg
28 May 2026
Safety

Intelligence as Managed Autonomy: Failure, Escalation, and Governance for Agentic AI Systems

DGX agent

arXiv:2605.27628v1 Announce Type: new Abstract: As autonomous and agentic AI systems scale in robotic and human-machine environments, managing hallucination and persistent but unjustified action remai

safetyarxiv-cs-ai
28 May 2026
Safety

Mitigating Adaptive Attacks against Reasoning Models with Activation Consistency Training

DGX agent

arXiv:2605.28467v1 Announce Type: new Abstract: As LLMs gain stronger reasoning capabilities, their extended chain-of-thought introduces new degrees of complexity for defending against adversarial jai

safetyarxiv-cs-lg
28 May 2026
Safety

Position: Retire the 'Positive Backdoor' Label -- Secret Alignment Requires Strict and Systematic Evaluation

DGX agent

arXiv:2605.28597v1 Announce Type: cross Abstract: This position paper argues that the AI/ML community should stop overclaiming and retire the label 'positive backdoor,' and instead treat trigger-activ

safetyarxiv-cs-ai
28 May 2026
Safety

Refusal Before Decoding: Detecting and Exploiting Refusal Signals in Intermediate LLM Activations

DGX agent

arXiv:2605.28553v1 Announce Type: new Abstract: In this paper, we investigate whether refusal behavior can be predicted from LLM intermediate activations before decoding using linear probes trained on

safetyarxiv-cs-ai
28 May 2026
Safety

ROOM: A Physics-Based Continuum Robot Simulator for Photorealistic Medical Datasets Generation

DGX agent

arXiv:2509.13177v2 Announce Type: replace Abstract: Continuum robots are advancing bronchoscopy procedures by accessing complex lung airways and enabling targeted interventions. However, their develop

safetyarxiv-cs-ro
28 May 2026
Model Releases

Safe In-Context Reinforcement Learning

DGX agent

arXiv:2509.25582v3 Announce Type: replace Abstract: In-context reinforcement learning (ICRL) is an emerging RL paradigm where an agent, after pretraining, can adapt to out-of-distribution test tasks w

model-releasesarxiv-cs-lg
28 May 2026
Safety

Simulation-Informed Diffusion for Decentralized Multi-robot Motion Planning

DGX agent

arXiv:2605.27697v1 Announce Type: cross Abstract: Decentralized multi-robot motion planning requires each robot to generate collision-free trajectories from local observations, without global sensing

safetyarxiv-cs-ai
28 May 2026
Safety

The Ethics of LLM Sandbox and Persona Dynamics

DGX agent

arXiv:2605.28647v1 Announce Type: new Abstract: It is well known that LLM guardrails and trained persona dynamics can produce a reality gap: the distance between the world a LLM is permitted or shaped

safetyarxiv-cs-ai
28 May 2026
Safety

Towards Unified Vision-Language Models with Incomplete Multi-Modal Inputs

DGX agent

arXiv:2605.27894v1 Announce Type: new Abstract: Video-Language Models (VLMs) have demonstrated impressive multi-modal reasoning capabilities across diverse computer vision applications. However, these

safetyarxiv-cs-cv
28 May 2026
Safety

Training Stratigraphy: Persistent Behavioral Artifacts in Large Language Models Observed Through Longitudinal AI-Human Interaction

DGX agent

arXiv:2605.28102v1 Announce Type: new Abstract: Large language models trained with Reinforcement Learning from Human Feedback (RLHF) and Constitutional AI exhibit persistent behavioral patterns that s

safetyarxiv-cs-ai
28 May 2026
Safety

VLA-Hijack: A Transferable Patch Attack against Vision-Language-Action Models via Visual Proprioception Hijacking

DGX agent

arXiv:2605.28083v1 Announce Type: new Abstract: While Vision-Language-Action (VLA) models have emerged as powerful generalist policies, their severe vulnerability to adversarial patches significantly

safetyarxiv-cs-cv
28 May 2026
Safety

Cultural Value Alignment Via Latent Activation Steering in Large Language Models

DGX agent

arXiv:2605.26365v1 Announce Type: new Abstract: Large Language Models (LLMs) often exhibit homogenized cultural perspectives. While the World Values Survey (WVS) provides a gold standard for mapping h

safetyarxiv-cs-cl
27 May 2026
← Previous
1…5253545556…300
Next →