AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,349 results
2 Jul 2026

Seahorse: A Unified Benchmarking Framework for Spatiotemporal Event Modeling

Model ReleasesDGX agent

arXiv:2607.01022v1 Announce Type: new Abstract: Spatiotemporal point processes (STPPs) model event data in continuous time and space, with applications in mobility, epidemiology, and public safety. Re

30 Jun 2026

A Physics-Grounded Benchmark for Multi-Agent Dynamics in World Models

Model ReleasesDGX agent

arXiv:2606.28757v1 Announce Type: new Abstract: Generative world models hold immense promise as scalable simulators for autonomous systems, particularly for synthesizing rare but safety-critical multi

Agree

ToolsDGX agent

Agree is a TypeScript library by Boris Cherny that provides a schema validation and serialization system, enabling developers to define data schemas with type safety and validate data at runtime. The

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Closing the Activation-Cone Blind Spot: Response-Time Probing and Unified Defense

Model ReleasesDGX agent

arXiv:2606.29441v1 Announce Type: cross Abstract: Inference-time safety methods for large language models have proliferated, yet no systematic comparison exists. We evaluate five defense paradigms (no

Fine-Tuning General-Purpose Large Language Models for Agricultural Applications:A Reproducible Framework and Evaluation Protocol Based on Qwen3-8B

Model ReleasesDGX agent

arXiv:2606.28992v1 Announce Type: cross Abstract: General-purpose large language models (LLMs) have demonstrated strong abilities in opendomain question answering, information extraction, and text gen

Hard-constraint physics-residual networks for hydrogen crossover prediction and high-pressure extrapolation in PEM water electrolysis

Model ReleasesDGX agent

arXiv:2511.05879v5 Announce Type: replace-cross Abstract: Hydrogen crossover is a critical safety and efficiency constraint in high-pressure polymer electrolyte membrane water electrolysis (PEMWE), bu

IMCBench: A benchmark for multimodal LLMs in Image-grounded Medical Conversations

Model ReleasesDGX agent

arXiv:2606.28556v1 Announce Type: new Abstract: Recent advances in large language models and vision-language models have enabled reasoning over multimodal data, offering opportunities for clinical app

Redefining Maritime Anomaly Detection via Equation-Grounded Synthetic Anomalies

Model ReleasesDGX agent

arXiv:2606.29721v1 Announce Type: cross Abstract: Maritime anomaly detection is essential for ensuring maritime safety, security, and efficient traffic management at sea, with Automatic Identification

Representational Depth of Evaluation Awareness Shifts With Scale in Open-Weight Language Models

Model ReleasesDGX agent

arXiv:2606.29196v1 Announce Type: cross Abstract: Do language models know when they are being tested? This question matters for AI safety: a model that recognises an evaluation context could alter its

SCARCE: Scalable Cascade Analysis for Rare-event Characterisation via Embeddings

Model ReleasesDGX agent

arXiv:2606.29623v1 Announce Type: new Abstract: Rare events govern the safety profile of modern AI systems, yet their probabilities are extremely difficult to estimate: direct Monte Carlo requires pro

ViPSim: Collaborating Visual and Parameter Spaces for Consistent Long-Horizon Embodied World Models

Model ReleasesDGX agent

arXiv:2606.28804v1 Announce Type: new Abstract: Embodied World Models (EWMs) have emerged as a scalable and risk-free paradigm for advancing embodied intelligence, enabling the safety-critical evaluat

29 Jun 2026

Physics-Guided Robotic Radiation Source Localization along Arbitrary Measurement Paths in Unstructured Environments

Local AiDGX agent

arXiv:2606.27624v1 Announce Type: cross Abstract: Using robots to estimate the location of the radiation source is an effective way to improve efficiency and safety. Existing methods focus on planning

When the Prompt Becomes Visual: Vision-Centric Jailbreak Attacks for Large Image Editing Models

Model ReleasesDGX agent

arXiv:2602.10179v2 Announce Type: replace-cross Abstract: Recent advances in large image editing models have shifted the paradigm from text-driven instructions to vision-prompt editing, where user int

26 Jun 2026

Autoformalization of Agent Instructions into Policy-as-Code

Model ReleasesDGX agent

arXiv:2606.26649v1 Announce Type: new Abstract: Agent safety in high-stakes domains requires formal policy enforcement, but most existing approaches either rely on probabilistic guardrails (fine-tuned

Digital Twin-Driven Communication-Efficient Federated Anomaly Detection for Industrial IoT

Model ReleasesDGX agent

arXiv:2601.01701v2 Announce Type: replace-cross Abstract: Anomaly detection is increasingly becoming crucial for maintaining the safety, reliability, and efficiency of industrial systems. Recently, wi

NavIsaacLab: Generating Realistic Crowd via Parallel Robot Learning for Benchmarking Human-aware Navigation

Model ReleasesDGX agent

arXiv:2606.26265v1 Announce Type: new Abstract: Robot autonomous navigation that accounts for surrounding human activities is crucial for ensuring both safety and natural human-robot interaction in re

Parametric Generalized Adaptive Moment Features (PG-AMF) for Bearing Fault Diagnosis and Machine Health Monitoring

Model ReleasesDGX agent

arXiv:2606.26317v1 Announce Type: cross Abstract: Accurate fault diagnosis of rolling element bearings in rotating machinery is considered essential for ensuring industrial safety and enabling predict

25 Jun 2026

A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation

Model ReleasesDGX agent

arXiv:2606.25476v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated remarkable performance across natural language processing tasks, yet their deployment in high-stakes appl

Auto-Labelling-Based Domain Transfer for 3D Object Detection on a Bicycle-Mounted LiDAR Platform

Model ReleasesDGX agent

arXiv:2606.25652v1 Announce Type: new Abstract: Reliable 3D perception of vulnerable road users (VRUs) such as cyclists and pedestrians is essential for their safety in urban traffic and a core requir

Distribution Preference Optimization: A Fine-grained Perspective for LLM Unlearning

Model ReleasesDGX agent

arXiv:2510.04773v2 Announce Type: replace Abstract: As Large Language Models (LLMs) demonstrate remarkable capabilities learned from vast corpora, concerns regarding data privacy and safety are receiv

Model Forensics: Investigating Whether Concerning Behavior Reflects Misalignment

Model ReleasesDGX agent

arXiv:2606.26071v1 Announce Type: new Abstract: A central goal of safety research is determining whether a model is misaligned. Prior work has largely focused on detecting concerning behavior. But beh

24 Jun 2026

PDS Joint: A Parametric Double-Spiral Joint Tailored for Dexterous Hands

Model ReleasesDGX agent

arXiv:2606.24377v1 Announce Type: new Abstract: Compliant joints can embed safety and adaptability into dexterous hands, but achieving large-stroke anthropomorphic motion while maintaining joint-speci

23 Jun 2026

Local Causal Attribution of Chain-of-Thought Reasoning

Local AiDGX agent

arXiv:2606.21821v1 Announce Type: new Abstract: Understanding the causal structure of a language model's thought process is a problem of significant importance for both transparency and safety. In thi

LOGOS: LiDAR-Only Gaussian Elevation Splatting for Unified Tiny Obstacle Segmentation

Model ReleasesDGX agent

arXiv:2606.21527v1 Announce Type: cross Abstract: Robust obstacle segmentation is essential for the safety of intelligent robots, where LiDAR-based perception systems play a fundamental role in the ro

Mirage: a Clean-Label Backdoor against LiDAR 3D Object Detection

Model ReleasesDGX agent

arXiv:2606.20752v1 Announce Type: new Abstract: Deep neural network-based LiDAR 3D object detection serves as a critical perception component in safety-critical autonomous systems. However, recent stu

11 Jun 2026

A Lightweight Multi-Agent Framework for Automated Concrete Barrier Design

Model ReleasesDGX agent

arXiv:2606.12040v1 Announce Type: new Abstract: The design of reinforced concrete highway barriers is a safety-critical process that requires strict compliance with regulatory provisions such as the A

PCS-UQ: Uncertainty Quantification via the Predictability-Computability-Stability Framework

Model ReleasesDGX agent

arXiv:2505.08784v2 Announce Type: replace-cross Abstract: As machine learning (ML) enters high-stakes domains, trustworthy uncertainty quantification (UQ) is essential for safety. In this paper we int

SceneMiner: Identity-Preserving Multi-Task Fine-Tuning for Unified BEV Scene Mining

Model ReleasesDGX agent

arXiv:2606.11507v1 Announce Type: new Abstract: Mining hard, safety-critical scenes from driving logs is bottlenecked by the absence of difficulty labels, and no single proxy, collision risk, trajecto

10 Jun 2026

[AINews] Anthropic Claude Fable 5 — Mythos but Safe, with Controversial Terms

Model ReleasesDGX agent

This article from Latent Space discusses Anthropic's Claude Fable 5 model, examining its capabilities in handling creative and mythological content while maintaining safety guardrails, along with cove

9 Jun 2026

A message to Anthropic leadership: You're not special. Making sure AI goes well is a team effort not a 'you effort.'

TutorialsDGX agent

Jeremy Howard argues that Anthropic's leadership should recognize that ensuring AI safety and positive outcomes requires collaborative effort across the industry rather than positioning any single org

Claude Mythos went from “too dangerous to release” to publicly available (with some extra guard rails) in two months. And y’all fell for Ant…

Model ReleasesDGX agent

Gary Marcus critiques Anthropic's rapid shift in positioning Claude from a model deemed too dangerous for public release to one made widely available with safety measures, suggesting this represents i

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models

Model ReleasesDGX agent

arXiv:2606.09142v1 Announce Type: cross Abstract: Egocentric vision offers a first-person view of human perception and decision making, yet its potential for traffic-safety prediction remains underexp

Defending Against Malicious Finetuning by Scaling Train-time Adversarial Attacks

Model ReleasesDGX agent

arXiv:2606.07970v1 Announce Type: cross Abstract: Current open-weight large language models (LLMs) are prone to malicious finetuning attacks, which could compromise the safety alignment of LLMs with o

Driving Video Retrieval for Complex Queries with Structured Grounding

Model ReleasesDGX agent

arXiv:2606.09109v1 Announce Type: new Abstract: Video retrieval at scale is central to data curation and safety validation in autonomous driving, where users want to find not only scenes but also dyna

Hybrid Robustness Verification for Spatio-Temporal Neural Networks

Model ReleasesDGX agent

arXiv:2606.09746v1 Announce Type: cross Abstract: With AI increasingly deployed in safety-critical systems, providing formal robustness guarantees for the underlying models is essential. Existing veri

Learning Predictive Control with Deep Koopman Operators for Autonomous Vehicle Motion Planning

Model ReleasesDGX agent

arXiv:2606.08136v1 Announce Type: new Abstract: Model Predictive Control (MPC) is widely used for autonomous-vehicle (AV) motion planning, but its real-time applicability is often limited by the need

SENTRY: Statistical Reliability Analysis of Vision Transformers Under Soft Errors

Model ReleasesDGX agent

arXiv:2606.07620v1 Announce Type: cross Abstract: With the growth of Vision Transformers in safety-critical domains like autonomous systems and medical imaging, ensuring their reliability against soft

6 Jun 2026

TLA-Prover: Verifiable TLA+ Specification Synthesis via Preference-Optimized Low-Rank Adaptation

Model ReleasesDGX agent

arXiv:2606.06133v1 Announce Type: cross Abstract: TLA+ is a formal specification language for verifying distributed systems and safety-critical protocols. Large language models (LLMs) frequently produ

5 Jun 2026

Coding with 'Enemy': Can Human Developers Detect AI Agent Sabotage?

Model ReleasesDGX agent

arXiv:2606.05647v1 Announce Type: cross Abstract: AI coding agents are increasingly embedded in real-world software development, collaborating with human developers while gaining broader access to cod

Drive-KD: Multi-Teacher Distillation for VLMs in Autonomous Driving

Model ReleasesDGX agent

arXiv:2601.21288v2 Announce Type: replace-cross Abstract: Autonomous driving is an important and safety-critical task, and recent advances in LLMs/VLMs have opened new possibilities for reasoning and

4 Jun 2026

CADET: A Modular Platform for Evaluating Distributed Cooperative Autonomy in Connected Autonomous Vehicles

Model ReleasesDGX agent

arXiv:2606.04072v1 Announce Type: cross Abstract: Deep learning models are increasingly central to autonomous vehicle (AV) pipelines, yet their integration has traditionally followed a monolithic desi

Latent Anchor-Driven Test Generation for Deep Neural Networks

Local AiDGX agent

arXiv:2606.04310v1 Announce Type: new Abstract: Deep Neural Networks (DNNs) are increasingly being deployed in security-critical and safety-sensitive applications, which makes rigorous testing essenti

Reality: The Final Eval — Lukas Petersson and Axel Backlund of Andon Labs

ToolsDGX agent

Lukas Petersson and Axel Backlund of Andon Labs discuss their work on evaluation methods and benchmarking for AI systems, likely covering approaches to assessing AI capabilities and safety in real-wor

The Saturation Trap and the Subjectivity of Intervention Timing: Why Affect-Based Triggers and LLM Judges Fail to Time Interventions on Autonomous Agents

Model ReleasesDGX agent

arXiv:2606.04296v1 Announce Type: new Abstract: As autonomous AI agents move from conversational systems to long-horizon software execution, runtime safety layers that decide when to interrupt an agen

3 Jun 2026

Don't Forget Your Embeddings: Robust Knowledge Erasure via Precise Editing of Embeddings

Model ReleasesDGX agent

arXiv:2606.03695v1 Announce Type: new Abstract: As language models are increasingly deployed in real-world applications, the ability to erase specific knowledge from them becomes critical for safety a

theUSshould lead on AI by continuing to develop the very best models, making sure they're safe, and getting cyber tools into the hands of tr…

IndustryDGX agent

Sam Altman argues that US leadership in artificial intelligence requires three concurrent priorities: advancing cutting-edge AI model development, ensuring these models incorporate robust safety measu

Unsupervised Robotaxi now in the entire Austin Metro area

IndustryDGX agent

Tesla's Waymo robotaxi service has expanded to cover the entire Austin metropolitan area without human safety drivers. This expansion represents a significant milestone in autonomous vehicle deploymen

2 Jun 2026

Beyond the Simplex: Balanced Prototype Geometry for Scorer-Agnostic Open-Set Recognition

Model ReleasesDGX agent

arXiv:2606.01883v1 Announce Type: cross Abstract: Open-set recognition (OSR) requires a classifier to reject inputs from unseen classes which is essential in safety-critical settings such as medical i

POIROT: Interrogating Agents for Failure Detection in Multi-Agent Systems

Model ReleasesDGX agent

arXiv:2606.02282v1 Announce Type: new Abstract: Orchestrating Large Language Models into Multi-Agent Systems (LLM-MAS) has unlocked remarkable reasoning capabilities, yet emergent failures and halluci

Reinforcement Learning for Optimal Experiment Design in Parameter Identification of Mechatronic Systems

Model ReleasesDGX agent

arXiv:2606.00059v1 Announce Type: cross Abstract: Informative excitation signals are critical for accurate system identification of mechatronic systems, yet classical system identification (SI) approa

ROGUE: Misaligned Agent Behavior Arising from Ordinary Computer Use

Model ReleasesDGX agent

arXiv:2606.00341v1 Announce Type: cross Abstract: As AI agents are increasingly deployed in real personal and corporate settings (email accounts, development workflows, company databases, etc.), safet

Truth, Trust, and Trouble: Medical AI on the Edge

Model ReleasesDGX agent

arXiv:2507.02983v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) hold significant promise for transforming digital health by enabling automated medical question answering. Howeve

When Safe Skills Collide: Measuring Compositional Risk in Agent Skill Ecosystems

Model ReleasesDGX agent

arXiv:2606.00448v1 Announce Type: cross Abstract: LLM agents increasingly rely on community-contributed skills that expand an agent's operational capability set. We study a core safety problem in agen

1 Jun 2026

McDonald’s advertising on 𝕏

IndustryDGX agent

McDonald's paused advertising on 𝕏 (formerly Twitter) in late 2024, likely due to concerns about content moderation and brand safety on the platform following Elon Musk's ownership changes. This move

MedFact: Benchmarking the Fact-Checking Capabilities of Large Language Models on Chinese Medical Texts

Model ReleasesDGX agent

arXiv:2509.12440v3 Announce Type: replace-cross Abstract: Deploying Large Language Models (LLMs) in medical applications requires fact-checking capabilities to ensure patient safety and regulatory com

Probabilistic Precipitation Nowcasting with Rectified Flow Transformers

Model ReleasesDGX agent

arXiv:2605.31204v1 Announce Type: new Abstract: Accurate weather forecasts are essential across various domains and are safety-critical in extreme weather conditions. Compared to simulation-based fore

Safe Equilibrium Policy Optimization for Strategic Agent Policies

Model ReleasesDGX agent

arXiv:2605.30854v1 Announce Type: cross Abstract: Language models fine-tuned with reinforcement learning typically optimize for task reward, ignoring multi-agent strategic structure. Because these age

Send a SCOUT First: Pre-hoc Reasoning for Adaptive Detector Allocation in Prompt-Injection Defense

Model ReleasesDGX agent

arXiv:2605.30837v1 Announce Type: cross Abstract: Prompt-injection detectors are heterogeneous: each is strong on a different slice of attacks, and none is always reliable. Yet existing systems still

29 May 2026

Dissecting the Black Box: Circuit-Level Analysis of LLM Vulnerability Detection

Model ReleasesDGX agent

arXiv:2605.29901v1 Announce Type: cross Abstract: Large language models (LLMs) can detect software vulnerabilities, but how do they actually identify vulnerable code? We address this question using me

Fairness Beyond Demographics: Optimizing Performance Across Appearance-Based Hidden Cohorts in Medical Imaging

Model ReleasesDGX agent

arXiv:2605.29827v1 Announce Type: new Abstract: Medical image analysis models can exhibit performance disparities across patient subgroups, threatening clinical safety and fairness. Existing methods t

← Previous
1…220221222223224…240
Next →