AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,314 results
Safety

Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook

DGX agent

arXiv:2604.06210v2 Announce Type: cross Abstract: As LLMs are globally deployed, aligning their cultural value orientations is critical for safety and user engagement. However, existing benchmarks fac

safetyarxiv-cs-ai
10 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

EMMa: End-Effector Stability-Oriented Mobile Manipulation for Tracked Rescue Robots

DGX agent

arXiv:2604.08292v1 Announce Type: new Abstract: The autonomous operation of tracked mobile manipulators in rescue missions requires not only ensuring the reachability and safety of robot motion but al

safetyarxiv-cs-ro
10 Apr 2026
Safety

Evaluation as Evolution: Transforming Adversarial Diffusion into Closed-Loop Curricula for Autonomous Vehicles

DGX agent

arXiv:2604.07378v1 Announce Type: new Abstract: Autonomous vehicles in interactive traffic environments are often limited by the scarcity of safety-critical tail events in static datasets, which biase

safetyarxiv-cs-ro
10 Apr 2026
Safety

Event-Centric World Modeling with Memory-Augmented Retrieval for Embodied Decision-Making

DGX agent

arXiv:2604.07392v1 Announce Type: cross Abstract: Autonomous agents operating in dynamic and safety-critical environments require decision-making frameworks that are both computationally efficient and

safetyarxiv-cs-ro
10 Apr 2026
Safety

Event-Level Detection of Surgical Instrument Handovers in Videos with Interpretable Vision Models

DGX agent

arXiv:2604.07577v1 Announce Type: new Abstract: Reliable monitoring of surgical instrument exchanges is essential for maintaining procedural efficiency and patient safety in the operating room. Automa

safetyarxiv-cs-cv
10 Apr 2026
Safety

Forward Trajectory Steering for Hamilton-Jacobi Reachability Analysis

DGX agent

arXiv:2608.11480v1 Announce Type: cross Abstract: Hamilton-Jacobi (HJ) reachability provides a mathematically rigorous framework for safe control of dynamical systems, but its practical application is

safetyarxiv-cs-lg
13 Aug 2026
Safety

Learning-Based Behavior Planning for Automated Driving: Real-World Integration and Deployment

DGX agent

arXiv:2608.12198v1 Announce Type: cross Abstract: Recent research in machine and deep learning has shown the potential of learningbased motion planning approaches to improve the driving behavior of au

safetyarxiv-cs-ai
13 Aug 2026
Safety

ELMER: Evolutionary Language Model that Explores and Refines

DGX agent

arXiv:2608.10196v1 Announce Type: cross Abstract: Program evolution can measure whether a mutation helped, but it rarely controls how far the mutation moves in behavior space. Syntactic edit size is a

safetyarxiv-cs-ai
12 Aug 2026
Safety

Hidden in Plain Sight: Diffusion-Based Unrestricted Robotic Attacks on Vision-Language-Action Models

DGX agent

arXiv:2608.10393v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have shown strong capabilities in controlling robots across diverse manipulation tasks. However, their adversarial r

safetyarxiv-cs-ai
12 Aug 2026
Safety

How to Verify Consistency of Probabilistic Claims

DGX agent

arXiv:2608.11181v1 Announce Type: cross Abstract: When a probabilistic predictor answers many conditional-probability queries, are its answers self-consistent, and can this be verified in polynomial t

safetyarxiv-cs-ai
12 Aug 2026
Safety

MedUP: Awakening Unified Understanding and Perception in Medical Vision-Language Models

DGX agent

arXiv:2608.10635v1 Announce Type: cross Abstract: Medical Vision-Language Models (Med-VLMs) excel at verbalizing visual content, yet precise visual perception, segmentation, and grounding remain chall

safetyarxiv-cs-ai
12 Aug 2026
Safety

Never Stop Speaking: a Denial-of-Service Attack on End-to-End Speech Language Models

DGX agent

arXiv:2608.10405v1 Announce Type: cross Abstract: Many studies have shown that specially crafted inputs can induce large language models (LLMs) to generate excessively long outputs, resulting in signi

safetyarxiv-cs-ai
12 Aug 2026
Safety

Rethinking Data Efficiency in Industrial Dense Prediction: Pretraining Coherence, Not Inductive Bias, Determines ViTs Low-Data Advantage

DGX agent

arXiv:2608.10590v1 Announce Type: new Abstract: Vision Transformers (ViTs) are widely believed to require more labeled data than CNNs for industrial dense prediction. Through controlled experiments on

safetyarxiv-cs-cv
12 Aug 2026
Safety

SPOTting the Future: Lookahead Explanations for Deep Reinforcement Learning

DGX agent

arXiv:2608.09967v1 Announce Type: new Abstract: Deep reinforcement learning (DRL) agents achieve strong performance in complex environments, yet their decision-making processes remain difficult to int

safetyarxiv-cs-ai
12 Aug 2026
Safety

The CASE Framework: A Multi-Disciplinary Control Architecture for Governing Enterprise Agentic AI

DGX agent

arXiv:2608.10153v1 Announce Type: new Abstract: Enterprises are deploying autonomous AI agents faster than they can govern them, and prevailing approaches stretch a single discipline, typically DevSec

safetyarxiv-cs-ai
12 Aug 2026
Safety

The Hidden Puppet Master: Predicting Human Belief Change in Manipulative LLM Dialogues

DGX agent

arXiv:2603.20907v5 Announce Type: replace Abstract: As users increasingly turn to LLMs for practical and personal advice, they become vulnerable to subtle steering toward hidden incentives misaligned

safetyarxiv-cs-cl
12 Aug 2026
Safety

XCoT-VLA: Executable Chain-of-Thought for Vision-Language-Action Driving

DGX agent

arXiv:2608.10976v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models can connect scene understanding, semantic reasoning, and trajectory generation for autonomous driving. However, verb

safetyarxiv-cs-ai
12 Aug 2026
Safety

An Agentic Generative Large Language Model for Treatment Planning of Colorectal Cancer

DGX agent

arXiv:2608.09142v1 Announce Type: new Abstract: Treatment planning in precision oncology requires synthesizing heterogeneous patient information with rapidly evolving clinical guidelines to ensure gui

safetyarxiv-cs-cl
11 Aug 2026
Safety

Decoy Images Amplify Caption-Mediated Defenses Against Encoded Jailbreaks

DGX agent

arXiv:2608.01043v2 Announce Type: replace-cross Abstract: We report a counter-intuitive interaction between image inputs and existing black-box defenses on Vision--Language Models (VLMs): pairing an e

safetyarxiv-cs-ai
11 Aug 2026
Safety

From Operational Design Domain to Action: A Systematic Behavioral Taxonomy for Autonomous Driving

DGX agent

arXiv:2608.08941v1 Announce Type: cross Abstract: Operational Design Domain (ODD) specifications describe where an automated driving system (ADS) is permitted to operate, but they do not prescribe wha

safetyarxiv-cs-ai
11 Aug 2026
Safety

MedCalc-R1: Knowledge-Guided Reward Framework for Medical Mathematical Reasoning

DGX agent

arXiv:2608.08623v1 Announce Type: new Abstract: In Reinforcement Learning with Verifiable Rewards (RLVR) frameworks for mathematical reasoning tasks, floating-point results are typically evaluated usi

safetyarxiv-cs-ai
11 Aug 2026
Safety

Multimodal Model Diffing for Feature Discovery and Control

DGX agent

arXiv:2608.09928v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) exhibit strong visual understanding, yet the internal features that cause these behaviors remain difficult to

safetyarxiv-cs-ai
11 Aug 2026
Safety

What Keeps Agent Skills from Being Reusable? Evidence from 138K SKILL.md Files

DGX agent

arXiv:2608.08453v1 Announce Type: new Abstract: Under the current standard, Agent Skills are SKILL.md files that combine instructions with supporting resources, enabling Large Language Model (LLM) age

safetyarxiv-cs-ai
11 Aug 2026
Safety

bioMoR: Biology-Guided Mixture-of-Recursions for Effective Genomic Learning

DGX agent

arXiv:2608.06727v1 Announce Type: new Abstract: Transformer models for high-dimensional omics analysis process thousands of genes or pathways, although only a subset requires deep computation. Mixture

safetyarxiv-cs-ai
10 Aug 2026
Safety

From Cheap Fakes to Pure Synthesis: Addressing the New Era of T2V Fake News Videos

DGX agent

arXiv:2608.06732v1 Announce Type: new Abstract: Recent text-to-video (T2V) generation models enable fake news videos to be synthesized from scratch, shifting the threat beyond cheap fakes assembled fr

safetyarxiv-cs-ai
10 Aug 2026
Safety

SoRoMoX: Fast, Differentiable, and Parallelizable Soft Robot Models

DGX agent

arXiv:2608.06650v1 Announce Type: cross Abstract: Reduced-order models based on Cosserat-rod theory are now well established, and modeling theory is no longer the primary bottleneck in soft-robot cont

safetyarxiv-cs-ai
10 Aug 2026
Safety

A System for Train Condition Monitoring and Structural Health Assessment of Rail Vehicles

DGX agent

arXiv:2608.05221v1 Announce Type: new Abstract: The ongoing digitalization of rail systems and the increasing use of artificial intelligence (AI) are fundamentally transforming the design, operation,

safetyarxiv-cs-ro
7 Aug 2026
Safety

DreamGuard: Efficient Runtime Guardrail for LLM Agents via Risk-Aware World Model

DGX agent

arXiv:2608.05695v1 Announce Type: new Abstract: As large language model (LLM) agents increasingly invoke external tools and interact with real-world systems, unsafe actions may cause irreversible cons

safetyarxiv-cs-ai
7 Aug 2026
Safety

Looking in the Mirror: Introspecting Side-Effect Misalignments Induced by Fine-Tuning

DGX agent

arXiv:2608.04347v1 Announce Type: new Abstract: Fine-tuning enables a source model to acquire desired capabilities and behaviors in a target domain while retaining much of its general-purpose competen

safetyarxiv-cs-lg
6 Aug 2026
Safety

A Blind Spot in Alignment: Quantifying Biosecurity Risks in Large Language Models

DGX agent

arXiv:2608.02684v1 Announce Type: cross Abstract: Large Language Models (LLMs) are accelerating biological research, yet this same capability poses a critical biosecurity threat: models that assist in

safetyarxiv-cs-ai
5 Aug 2026
Safety

LatentGuard: Efficient and Inspectable Latent Reasoning for LLM Safeguards

DGX agent

arXiv:2608.03838v1 Announce Type: new Abstract: Reasoning-based guard models improve LLM safeguards, but decoding explicit rationales for every interaction makes them costly to deploy. Although latent

safetyarxiv-cs-ai
5 Aug 2026
Safety

Poisoning Prompt-Guided Sampling in Video Large Language Models

DGX agent

arXiv:2509.20851v2 Announce Type: replace Abstract: Video Large Language Models (VideoLLMs) are increasingly deployed as automated moderators on user-generated video platforms, where a few unwatched s

safetyarxiv-cs-cv
5 Aug 2026
Safety

PULSE: An Executable Contract Language for Spatiotemporal Knowledge Graph Engineering

DGX agent

arXiv:2608.02630v1 Announce Type: new Abstract: Knowledge graph engineering often distributes accepted state, observations, constraints, processes, and hypothetical scenarios across artifacts whose co

safetyarxiv-cs-ai
5 Aug 2026
Safety

HapticVLA: Contact-Rich Manipulation via Vision-Language-Action Model without Inference-Time Tactile Sensing

DGX agent

arXiv:2603.15257v2 Announce Type: replace Abstract: Tactile sensing is a crucial capability for Vision-Language-Action (VLA) architectures, as it enables dexterous and safe manipulation in contact-ric

safetyarxiv-cs-ro
4 Aug 2026
Model Releases

MedPRESS: A Multi-turn Benchmark for Patient-Pressure-Induced Medical Sycophancy in LLMs

DGX agent

arXiv:2608.02520v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for health-related advice. Existing research measures their safety with static questions rather than

model-releasesarxiv-cs-cl
4 Aug 2026
Model Releases

Quality-Diversity Red-Teaming: Automated Generation of High-Quality and Diverse Attackers for Large Language Models

DGX agent

arXiv:2506.07121v2 Announce Type: replace Abstract: Ensuring the safety and robustness of large language models (LLMs) is a fundamental challenge and a critical prerequisite for the responsible deploy

model-releasesarxiv-cs-lg
4 Aug 2026
Safety

Self-Improving Large Language Models via Progressive Experience Evolution

DGX agent

arXiv:2608.02139v1 Announce Type: new Abstract: Large language models (LLMs) capable of self-improvement require not only effective policy optimization, but also a principled mechanism for transformin

safetyarxiv-cs-cl
4 Aug 2026
Safety

When May a Model Replace the Experiment? Audits, Licenses, and the Price of Trust in Surrogate-Driven Design

DGX agent

arXiv:2608.01378v1 Announce Type: new Abstract: Design campaigns in chemistry, materials science, and machine learning share a bottleneck: determining how good a candidate truly is requires an expensi

safetyarxiv-cs-lg
4 Aug 2026
Safety

CAGE: Certified Authorization under Typed-Return Uncertainty for Tool-Using Agents

DGX agent

arXiv:2607.29190v1 Announce Type: new Abstract: Tool-using LLM agents act on typed tool returns, records pairing provenance and categorical fields with numerical values. Runtime permission gates gener

safetyarxiv-cs-ai
3 Aug 2026
Safety

Active Lubrication of Transluminal Medical Instruments

DGX agent

arXiv:2506.07225v2 Announce Type: replace-cross Abstract: Transluminal minimally invasive surgery uses natural orifices and small incisions to access internal anatomical structures, promoting quicker

safetyarxiv-cs-ro
31 Jul 2026
Safety

Ask don't tell: Reducing sycophancy in large language models

DGX agent

arXiv:2602.23971v4 Announce Type: replace-cross Abstract: Sycophancy, the tendency of large language models to favour user-affirming responses over critical engagement, has been identified as an align

safetyarxiv-cs-ai
31 Jul 2026
Safety

Corrigible Assistance in One Round: Pragmatic-Pedagogic Best Response

DGX agent

arXiv:2607.27508v1 Announce Type: new Abstract: Assistance games formalize human-robot collaboration under asymmetric information: the human knows the goal, while the robot must infer it from observat

safetyarxiv-cs-ro
31 Jul 2026
Safety

EmoFeedback^2: Reinforcement of Continuous Emotional Image Generation via LVLM-based Reward and Textual Feedback

DGX agent

arXiv:2511.19982v3 Announce Type: replace Abstract: Continuous emotional image content generation (C-EICG) is emerging rapidly due to its ability to produce images aligned with both user descriptions

safetyarxiv-cs-cv
31 Jul 2026
Safety

FleetScape: A Mixed Reality Sandtable for Spatial Supervision and Control of Scalable Drone Fleets

DGX agent

arXiv:2607.26423v1 Announce Type: cross Abstract: As autonomous drone deployments scale from individual units to coordinated swarms, the human operator's role shifts from direct piloting to high-level

safetyarxiv-cs-ro
30 Jul 2026
Model Releases

Prosody-driven Jailbreaks in Audio LLMs: A Controlled Study and Mechanistic Analysis

DGX agent

arXiv:2607.26541v1 Announce Type: cross Abstract: Audio-capable foundation models enable end-to-end spoken interaction, but they also introduce safety risks beyond transcript content. It remains uncle

model-releasesarxiv-cs-cl
30 Jul 2026
Safety

Agentic AI in medicine: architectures, applications, evaluation, and challenges for clinical translation

DGX agent

arXiv:2607.25489v1 Announce Type: new Abstract: Large language models and multimodal foundation models are enabling medical artificial intelligence (AI) systems to move beyond isolated prediction and

safetyarxiv-cs-cv
29 Jul 2026
Model Releases

Inspect India Evals: An Open Benchmarking Framework for Evaluating Large Language Models in the Indian Linguistic and Cultural Context

DGX agent

arXiv:2607.25375v1 Announce Type: new Abstract: India is a vast nation of over 1.4 billion people, varied by hundreds of diverse and locally specific traditions and cultures and 22 officially recogniz

model-releasesarxiv-cs-cl
29 Jul 2026
Model Releases

Shieldstral

DGX agent

arXiv:2607.25857v1 Announce Type: new Abstract: We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7imes its size on text s

model-releasesarxiv-cs-cl
29 Jul 2026
← Previous
1…2930313233…257
Next →