AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,314 results
Safety

Metric-Dependent Annotation Saturation for Learning from Label Distributions

DGX agent

arXiv:2605.29797v1 Announce Type: new Abstract: When annotators disagree on a label, the disagreement itself carries signal -- and the number of annotators needed to capture it depends on the evaluati

safetyarxiv-cs-cl
29 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

Neural Network Verification using Partial Multi-Neuron Relaxation

DGX agent

arXiv:2605.30155v1 Announce Type: cross Abstract: The increasing integration of deep neural networks in critical systems has spawned a theoretical and practical interest in formally guaranteeing safet

safetyarxiv-cs-ai
29 May 2026
Safety

Quantum-Enhanced Adversarial Robustness in Artificial Intelligence

DGX agent

arXiv:2605.28899v1 Announce Type: cross Abstract: Artificial Intelligence has achieved remarkable success across diverse application domains. However, its vulnerability to adversarial attacks poses si

safetyarxiv-cs-ai
29 May 2026
Safety

The Anatomy of Conversational Scams: A Topic-Based Red Teaming Analysis of Multi-Turn Interactions in LLMs

DGX agent

arXiv:2601.03134v2 Announce Type: replace Abstract: As LLMs gain persuasive capabilities through extended dialogues, they create new opportunities for studying adversarial conversational behavior in e

safetyarxiv-cs-cl
29 May 2026
Safety

When and How Long? The Readout-Mediator Angle in Temporal Reasoning

DGX agent

arXiv:2605.29126v1 Announce Type: cross Abstract: A linear probe can decode a representation almost perfectly and yet be completely irrelevant to how the model uses it. On calendar-date duration reaso

safetyarxiv-cs-ai
29 May 2026
Safety

A Comparative Study of Rule-Based and Data-Driven Approaches in Industrial Monitoring

DGX agent

arXiv:2509.15848v2 Announce Type: replace Abstract: Industrial monitoring systems, especially when deployed in Industry 4.0 environments, are experiencing a shift in paradigm from traditional rule-bas

safetyarxiv-cs-ai
28 May 2026
Safety

A Policy-Driven Runtime Layer for Agentic LLM Serving

DGX agent

arXiv:2605.27744v1 Announce Type: new Abstract: Multi-agent LLM systems have become the dominant production workload, but the serving stack was not built for them. The agent framework above knows agen

safetyarxiv-cs-ai
28 May 2026
Safety

Bayesian Deployment Approval for Learned Landing Controllers under Finite Rollout Validation

DGX agent

arXiv:2605.27720v1 Announce Type: new Abstract: Reinforcement learning and data-driven autonomous controllers are commonly evaluated using cumulative reward and empirical success frequency under finit

safetyarxiv-cs-lg
28 May 2026
Safety

Bayesian Gated Non-Negative Contrastive Learning

DGX agent

arXiv:2605.28441v1 Announce Type: cross Abstract: While Contrastive Learning (CL) has revolutionized self-supervised representation learning, its latent representations remain highly entangled and opa

safetyarxiv-cs-ai
28 May 2026
Safety

Colosseum: Auditing Collusion in Cooperative Multi-Agent Systems

DGX agent

arXiv:2602.15198v2 Announce Type: replace-cross Abstract: Multi-agent systems, where LLM agents communicate through free-form language, enable sophisticated coordination for solving complex cooperativ

safetyarxiv-cs-ai
28 May 2026
Safety

Delay-Aware Reinforcement Learning for Highway On-Ramp Merging under Stochastic Communication Latency

DGX agent

arXiv:2403.11852v5 Announce Type: replace-cross Abstract: Delayed and partially observable state information poses significant challenges for reinforcement learning (RL)-based control in real-world au

safetyarxiv-cs-ai
28 May 2026
Safety

How VLAs Fail Differently: Black-Box Action Monitoring Reveals Architecture-Specific Failure Signatures

DGX agent

arXiv:2605.28726v1 Announce Type: cross Abstract: We discover that VLA architectures fail in fundamentally different, predictable ways at the motor-command level. Running VQ-BeT, Diffusion Policy, and

safetyarxiv-cs-lg
28 May 2026
Safety

Intelligence as Managed Autonomy: Failure, Escalation, and Governance for Agentic AI Systems

DGX agent

arXiv:2605.27628v1 Announce Type: new Abstract: As autonomous and agentic AI systems scale in robotic and human-machine environments, managing hallucination and persistent but unjustified action remai

safetyarxiv-cs-ai
28 May 2026
Safety

Mitigating Adaptive Attacks against Reasoning Models with Activation Consistency Training

DGX agent

arXiv:2605.28467v1 Announce Type: new Abstract: As LLMs gain stronger reasoning capabilities, their extended chain-of-thought introduces new degrees of complexity for defending against adversarial jai

safetyarxiv-cs-lg
28 May 2026
Safety

Position: Retire the 'Positive Backdoor' Label -- Secret Alignment Requires Strict and Systematic Evaluation

DGX agent

arXiv:2605.28597v1 Announce Type: cross Abstract: This position paper argues that the AI/ML community should stop overclaiming and retire the label 'positive backdoor,' and instead treat trigger-activ

safetyarxiv-cs-ai
28 May 2026
Safety

Refusal Before Decoding: Detecting and Exploiting Refusal Signals in Intermediate LLM Activations

DGX agent

arXiv:2605.28553v1 Announce Type: new Abstract: In this paper, we investigate whether refusal behavior can be predicted from LLM intermediate activations before decoding using linear probes trained on

safetyarxiv-cs-ai
28 May 2026
Safety

ROOM: A Physics-Based Continuum Robot Simulator for Photorealistic Medical Datasets Generation

DGX agent

arXiv:2509.13177v2 Announce Type: replace Abstract: Continuum robots are advancing bronchoscopy procedures by accessing complex lung airways and enabling targeted interventions. However, their develop

safetyarxiv-cs-ro
28 May 2026
Model Releases

Safe In-Context Reinforcement Learning

DGX agent

arXiv:2509.25582v3 Announce Type: replace Abstract: In-context reinforcement learning (ICRL) is an emerging RL paradigm where an agent, after pretraining, can adapt to out-of-distribution test tasks w

model-releasesarxiv-cs-lg
28 May 2026
Safety

Simulation-Informed Diffusion for Decentralized Multi-robot Motion Planning

DGX agent

arXiv:2605.27697v1 Announce Type: cross Abstract: Decentralized multi-robot motion planning requires each robot to generate collision-free trajectories from local observations, without global sensing

safetyarxiv-cs-ai
28 May 2026
Safety

The Ethics of LLM Sandbox and Persona Dynamics

DGX agent

arXiv:2605.28647v1 Announce Type: new Abstract: It is well known that LLM guardrails and trained persona dynamics can produce a reality gap: the distance between the world a LLM is permitted or shaped

safetyarxiv-cs-ai
28 May 2026
Safety

Towards Unified Vision-Language Models with Incomplete Multi-Modal Inputs

DGX agent

arXiv:2605.27894v1 Announce Type: new Abstract: Video-Language Models (VLMs) have demonstrated impressive multi-modal reasoning capabilities across diverse computer vision applications. However, these

safetyarxiv-cs-cv
28 May 2026
Safety

Training Stratigraphy: Persistent Behavioral Artifacts in Large Language Models Observed Through Longitudinal AI-Human Interaction

DGX agent

arXiv:2605.28102v1 Announce Type: new Abstract: Large language models trained with Reinforcement Learning from Human Feedback (RLHF) and Constitutional AI exhibit persistent behavioral patterns that s

safetyarxiv-cs-ai
28 May 2026
Safety

VLA-Hijack: A Transferable Patch Attack against Vision-Language-Action Models via Visual Proprioception Hijacking

DGX agent

arXiv:2605.28083v1 Announce Type: new Abstract: While Vision-Language-Action (VLA) models have emerged as powerful generalist policies, their severe vulnerability to adversarial patches significantly

safetyarxiv-cs-cv
28 May 2026
Safety

Cultural Value Alignment Via Latent Activation Steering in Large Language Models

DGX agent

arXiv:2605.26365v1 Announce Type: new Abstract: Large Language Models (LLMs) often exhibit homogenized cultural perspectives. While the World Values Survey (WVS) provides a gold standard for mapping h

safetyarxiv-cs-cl
27 May 2026
Safety

Erased but Exploitable: Black-box Embedding-Aware Prompting Against Unlearned Text-to-Image Diffusion Models

DGX agent

arXiv:2605.26332v1 Announce Type: cross Abstract: Machine unlearning aims to remove specific concepts from pretrained text-to-image diffusion models, yet several white- and black-box attacks have been

safetyarxiv-cs-ai
27 May 2026
Safety

LAD-VF: LLM-Automatic Differentiation Enables Fine-Tuning-Free Robot Planning from Formal Methods Feedback

DGX agent

arXiv:2509.18384v2 Announce Type: replace Abstract: Large language models (LLMs) can translate natural language instructions into executable action plans for robotics, autonomous driving, and other do

safetyarxiv-cs-ro
27 May 2026
Safety

Look Further: Socially-Compliant Navigation System in Residential Buildings

DGX agent

arXiv:2605.26710v1 Announce Type: new Abstract: The distance at which a mobile robot reacts to a person strongly impacts various qualities of the human-robot interaction. In this paper, we focus on th

safetyarxiv-cs-ro
27 May 2026
Safety

Measuring Prediction Uncertainty in Neural Cellular Automata

DGX agent

arXiv:2605.26726v1 Announce Type: cross Abstract: Neural cellular automata (NCA) provide a lightweight alternative to encoder-decoder segmentation networks. However, it can be difficult to decide when

safetyarxiv-cs-ai
27 May 2026
Safety

Neuro-Symbolic Verification of LLM Outputs for Data-Sensitive Domains (extended preprint)

DGX agent

arXiv:2605.26942v1 Announce Type: new Abstract: LLMs deployed in high-stakes domains face fundamental reliability challenges: hallucinations, inconsistencies, and privacy vulnerabilities introduce una

safetyarxiv-cs-ai
27 May 2026
Safety

Practical Anonymous Two-Party Gradient Boosting Decision Tree

DGX agent

arXiv:2605.26903v1 Announce Type: cross Abstract: Structured data is well handled by gradient-boosted decision trees (GBDT), which are usually trained on vertically partitioned features across mutuall

safetyarxiv-cs-ai
27 May 2026
Safety

ReasonOps: A Unified Operational Paradigm for Trustworthy Verified LLM Reasoning

DGX agent

arXiv:2605.27014v1 Announce Type: cross Abstract: Large Language Models (LLMs) have transformed artificial intelligence from primarily generative systems into increasingly capable reasoning agents. Re

safetyarxiv-cs-ai
27 May 2026
Safety

The AI Cognitive Trojan Horse: How Large Language Models May Bypass Human Epistemic Vigilance

DGX agent

arXiv:2601.07085v2 Announce Type: replace-cross Abstract: Large language model (LLM)-based conversational AI systems present a challenge to human cognition that current frameworks for understanding mi

safetyarxiv-cs-ai
27 May 2026
Safety

TPS-Drive: Task-Guided Representation Purification for VLM-based Autonomous Driving

DGX agent

arXiv:2605.27038v1 Announce Type: new Abstract: Vision-Language Models (VLMs) provide a promising foundation for autonomous driving planning, yet bridging semantic reasoning and precise 3D spatial for

safetyarxiv-cs-ro
27 May 2026
Safety

Trust, Geometry, and Rules: A Credibility-Aware Reinforcement Learning Framework for Safe USV Navigation under Uncertainty

DGX agent

arXiv:2605.26974v1 Announce Type: new Abstract: Autonomous navigation of Unmanned Surface Vehicles (USVs) that is safe and compliant with the International Regulations for Preventing Collisions at Sea

safetyarxiv-cs-ro
27 May 2026
Safety

A Contractive Feedback Semantics for Reinforcement Learning

DGX agent

arXiv:2605.24759v1 Announce Type: new Abstract: Discounted reinforcement learning is usually presented through Bellman equations on closed Markov decision processes. This paper develops a compositiona

safetyarxiv-cs-lg
26 May 2026
Safety

AnyScene: Towards Highly Controllable Driving Scene Generation at Anywhere and Beyond

DGX agent

arXiv:2605.26113v1 Announce Type: new Abstract: Generating high-fidelity and controllable synthetic data is critical for advancing end-to-end autonomous driving, particularly for addressing the long t

safetyarxiv-cs-ro
26 May 2026
Model Releases

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs

DGX agent

arXiv:2605.21602v2 Announce Type: replace Abstract: Many safety and alignment failures of large language models (LLMs) occur due to out-of-distribution (OOD) situations: unusual prompt or response pat

model-releasesarxiv-cs-ai
26 May 2026
Safety

Capability and Robustness Cannot Both Be Free: An Information-Theoretic Bound for Vision-Language-Action Models

DGX agent

arXiv:2605.25889v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models are increasingly deployed on real robots, where each predicted action is executed and each failure carries a safet

safetyarxiv-cs-lg
26 May 2026
Safety

Contractual Skills: A GovernSpec Design Framework for Enterprise AI Agents

DGX agent

arXiv:2605.22634v2 Announce Type: replace-cross Abstract: Skills have become a practical packaging mechanism for agent instructions, workflows, scripts, and reference materials. In enterprise settings

safetyarxiv-cs-ai
26 May 2026
Safety

Counterfactually Safe Reinforcement Learning

DGX agent

arXiv:2605.25114v1 Announce Type: cross Abstract: Reinforcement learning algorithms are generally designed to maximize the expected return across a population. However, a policy that is optimal on ave

safetyarxiv-cs-lg
26 May 2026
Safety

Evidence-Linked Radiology Reporting: A Human-Supervised Reference Architecture for Structured Imaging Intelligence

DGX agent

arXiv:2605.25120v1 Announce Type: cross Abstract: Radiology reports remain the primary mechanism by which imaging findings are communicated to clinical teams. However, much of the structured informati

safetyarxiv-cs-ai
26 May 2026
Safety

From Automation to Collaboration: Human-in-the-Loop Methods for Safe and Trustworthy NLP

DGX agent

arXiv:2605.25226v1 Announce Type: new Abstract: Large language models are widely deployed in high-stakes NLP tasks, yet risks such as bias, hallucination, adversarial vulnerability and unreliable gene

safetyarxiv-cs-cl
26 May 2026
Safety

From Multi-Agent Systems and the Semantic Web to Agentic AI: A Unified Narrative of the Web of Agents

DGX agent

arXiv:2507.10644v4 Announce Type: replace Abstract: The Web of Agents (WoA) transforms the document-centric Web into an environment of autonomous agents acting on users' behalf, a vision newly tractab

safetyarxiv-cs-ai
26 May 2026
Safety

Hidden-State Privacy Has an Empty Middle

DGX agent

arXiv:2605.24042v1 Announce Type: cross Abstract: Of 1{,}536 Gaussian release covariances we tested for single-layer hidden-state privacy, zero achieve both moderate utility and moderate privacy again

safetyarxiv-cs-ai
26 May 2026
Safety

HoLoArm: Deformable Arms for Collision-Tolerant Quadrotor Flight

DGX agent

arXiv:2605.25790v1 Announce Type: new Abstract: The increasing use of drones in human-centric applications highlights the need for designs that can survive collisions and recover rapidly, minimizing r

safetyarxiv-cs-ro
26 May 2026
Safety

HumanFlow -- Diffusion-Driven MAV Navigation Among Humans via Tightly-Coupled Motion Tracking, Forecasting, and Control

DGX agent

arXiv:2605.25685v1 Announce Type: new Abstract: Robust and accurate perception of humans in their 3D scene context is essential for integrating robots into everyday environments. Existing approaches,

safetyarxiv-cs-ro
26 May 2026
Safety

Improving Labeling Consistency with Detailed Constitutional Definitions and AI-Driven Evaluation

DGX agent

arXiv:2605.24247v1 Announce Type: cross Abstract: Many automated labeling pipelines classify inputs into categories defined by a written specification, content moderation being a prominent use case. S

safetyarxiv-cs-ai
26 May 2026
Safety

LipoAgent: Coordinating Fine-Tuned LLM Agents for Safer Lipid Design

DGX agent

arXiv:2605.25250v1 Announce Type: new Abstract: Lipid nanoparticles (LNPs) are among the most clinically mature platforms for nucleic acid delivery, yet designing lipids that are both effective and bi

safetyarxiv-cs-ai
26 May 2026
← Previous
1…4748495051…257
Next →