AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,349 results
Model Releases

Beyond Scene Priors: Fine-Grained Traffic Scene Reasoning with Benchmarking and Query-Guided Small-Object Focus

DGX agent

arXiv:2607.04149v1 Announce Type: new Abstract: In safety-critical traffic scenarios, answering complex questions relies on minute, localized visual cues. However, standard Multimodal Large Language M

model-releasesarxiv-cs-cv
7 Jul 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Local Ai

Federated Learning for Object Detection: Enabling Collaborative Drone Learning Without Centralizing Data

DGX agent

arXiv:2607.02636v1 Announce Type: cross Abstract: Object detection is a fundamental capability for AI-driven perception in safety-critical drone and edge-vision systems, including disaster response, o

local-aiarxiv-cs-ai
7 Jul 2026
Model Releases

From Arithmetic to Logic: The Resilience of Logic and Lookup-Based Neural Networks Under Parameter Bit-Flips

DGX agent

arXiv:2603.22770v2 Announce Type: replace-cross Abstract: The deployment of deep neural networks (DNNs) in safety-critical edge environments necessitates robustness against hardware-induced bit-flip e

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

Hierarchical Multi-to-Single-Modal Knowledge Distillation for Disruption Prediction in EAST

DGX agent

arXiv:2607.04241v1 Announce Type: cross Abstract: Plasma disruption is a critical threat to tokamak safety. Existing data-driven predictors mainly rely on time-series diagnostic signals, while visible

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

LLMs Encode Harmfulness and Refusal Separately

DGX agent

arXiv:2507.11878v5 Announce Type: replace Abstract: LLMs are trained to refuse harmful instructions, but do they truly understand harmfulness beyond just refusing? Prior work has shown that LLMs' refu

model-releasesarxiv-cs-cl
7 Jul 2026
Model Releases

NEST: Nascent Encoded Steganographic Thoughts

DGX agent

arXiv:2602.14095v2 Announce Type: replace Abstract: Monitoring chain-of-thought (CoT) reasoning is a foundational safety technique for large language model agents; however, this oversight is compromis

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

Optimizing Large Language Models for Causality Assessment in Pharmacovigilance: Developing a Performance Metric as Objective for Bayesian Hyperparameter Optimization

DGX agent

arXiv:2607.03704v1 Announce Type: new Abstract: Background: Growing individual case safety report (ICSR) volumes have intensified demand for scalable automated causality assessment. Large Language Mod

model-releasesarxiv-cs-cl
7 Jul 2026
Model Releases

Refused in Chat, Written in Code: Workflow-Level Jailbreak Construction in IDE Coding Agents

DGX agent

arXiv:2607.03968v1 Announce Type: cross Abstract: Large language models are increasingly deployed as IDE-integrated coding agents that decompose tasks, generate and edit files, run code, and refine ou

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

Robustness Meets Uncertainty: Evidential Adversarial Training for Robust Selective Classification

DGX agent

arXiv:2607.03075v1 Announce Type: cross Abstract: Safety-critical applications require classifiers that are both robust and reliable. Adversarial training is a widely adopted defense for improving rob

model-releasesarxiv-cs-cv
7 Jul 2026
Model Releases

Fast Multi-dimensional Refusal Subspaces via RFM-AGOP

DGX agent

arXiv:2607.02396v1 Announce Type: new Abstract: Steering and monitoring activations in Large Language Models (LLMs) are increasingly used for both safety and interpretability. Early work assumed behav

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Locality-Aware Continual Unlearning for Diffusion Models

DGX agent

arXiv:2512.02657v2 Announce Type: replace-cross Abstract: Real-world deployment of text-to-image diffusion models requires continual concept removal as new privacy, copyright, or safety obligations ar

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

OpenSafeIntent: Evaluating Intent-Calibrated Safe Completion Across Dual-Use Prompt Sets

DGX agent

arXiv:2607.02047v1 Announce Type: cross Abstract: Safe completion requires models to provide useful assistance without enabling harm, but this behavior is difficult to evaluate with isolated prompts.

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Seahorse: A Unified Benchmarking Framework for Spatiotemporal Event Modeling

DGX agent

arXiv:2607.01022v1 Announce Type: new Abstract: Spatiotemporal point processes (STPPs) model event data in continuous time and space, with applications in mobility, epidemiology, and public safety. Re

model-releasesarxiv-cs-lg
2 Jul 2026
Model Releases

A Physics-Grounded Benchmark for Multi-Agent Dynamics in World Models

DGX agent

arXiv:2606.28757v1 Announce Type: new Abstract: Generative world models hold immense promise as scalable simulators for autonomous systems, particularly for synthesizing rare but safety-critical multi

model-releasesarxiv-cs-cv
30 Jun 2026
Tools

Agree

DGX agent

Agree is a TypeScript library by Boris Cherny that provides a schema validation and serialization system, enabling developers to define data schemas with type safety and validate data at runtime. The

toolsboris-cherny--x
30 Jun 2026
Model Releases

Closing the Activation-Cone Blind Spot: Response-Time Probing and Unified Defense

DGX agent

arXiv:2606.29441v1 Announce Type: cross Abstract: Inference-time safety methods for large language models have proliferated, yet no systematic comparison exists. We evaluate five defense paradigms (no

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Fine-Tuning General-Purpose Large Language Models for Agricultural Applications:A Reproducible Framework and Evaluation Protocol Based on Qwen3-8B

DGX agent

arXiv:2606.28992v1 Announce Type: cross Abstract: General-purpose large language models (LLMs) have demonstrated strong abilities in opendomain question answering, information extraction, and text gen

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Hard-constraint physics-residual networks for hydrogen crossover prediction and high-pressure extrapolation in PEM water electrolysis

DGX agent

arXiv:2511.05879v5 Announce Type: replace-cross Abstract: Hydrogen crossover is a critical safety and efficiency constraint in high-pressure polymer electrolyte membrane water electrolysis (PEMWE), bu

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

IMCBench: A benchmark for multimodal LLMs in Image-grounded Medical Conversations

DGX agent

arXiv:2606.28556v1 Announce Type: new Abstract: Recent advances in large language models and vision-language models have enabled reasoning over multimodal data, offering opportunities for clinical app

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Redefining Maritime Anomaly Detection via Equation-Grounded Synthetic Anomalies

DGX agent

arXiv:2606.29721v1 Announce Type: cross Abstract: Maritime anomaly detection is essential for ensuring maritime safety, security, and efficient traffic management at sea, with Automatic Identification

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Representational Depth of Evaluation Awareness Shifts With Scale in Open-Weight Language Models

DGX agent

arXiv:2606.29196v1 Announce Type: cross Abstract: Do language models know when they are being tested? This question matters for AI safety: a model that recognises an evaluation context could alter its

model-releasesarxiv-cs-cl
30 Jun 2026
Model Releases

SCARCE: Scalable Cascade Analysis for Rare-event Characterisation via Embeddings

DGX agent

arXiv:2606.29623v1 Announce Type: new Abstract: Rare events govern the safety profile of modern AI systems, yet their probabilities are extremely difficult to estimate: direct Monte Carlo requires pro

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

ViPSim: Collaborating Visual and Parameter Spaces for Consistent Long-Horizon Embodied World Models

DGX agent

arXiv:2606.28804v1 Announce Type: new Abstract: Embodied World Models (EWMs) have emerged as a scalable and risk-free paradigm for advancing embodied intelligence, enabling the safety-critical evaluat

model-releasesarxiv-cs-cv
30 Jun 2026
Local Ai

Physics-Guided Robotic Radiation Source Localization along Arbitrary Measurement Paths in Unstructured Environments

DGX agent

arXiv:2606.27624v1 Announce Type: cross Abstract: Using robots to estimate the location of the radiation source is an effective way to improve efficiency and safety. Existing methods focus on planning

local-aiarxiv-cs-lg
29 Jun 2026
Model Releases

When the Prompt Becomes Visual: Vision-Centric Jailbreak Attacks for Large Image Editing Models

DGX agent

arXiv:2602.10179v2 Announce Type: replace-cross Abstract: Recent advances in large image editing models have shifted the paradigm from text-driven instructions to vision-prompt editing, where user int

model-releasesarxiv-cs-ai
29 Jun 2026
Model Releases

Autoformalization of Agent Instructions into Policy-as-Code

DGX agent

arXiv:2606.26649v1 Announce Type: new Abstract: Agent safety in high-stakes domains requires formal policy enforcement, but most existing approaches either rely on probabilistic guardrails (fine-tuned

model-releasesarxiv-cs-ai
26 Jun 2026
Model Releases

Digital Twin-Driven Communication-Efficient Federated Anomaly Detection for Industrial IoT

DGX agent

arXiv:2601.01701v2 Announce Type: replace-cross Abstract: Anomaly detection is increasingly becoming crucial for maintaining the safety, reliability, and efficiency of industrial systems. Recently, wi

model-releasesarxiv-cs-ai
26 Jun 2026
Model Releases

NavIsaacLab: Generating Realistic Crowd via Parallel Robot Learning for Benchmarking Human-aware Navigation

DGX agent

arXiv:2606.26265v1 Announce Type: new Abstract: Robot autonomous navigation that accounts for surrounding human activities is crucial for ensuring both safety and natural human-robot interaction in re

model-releasesarxiv-cs-ro
26 Jun 2026
Model Releases

Parametric Generalized Adaptive Moment Features (PG-AMF) for Bearing Fault Diagnosis and Machine Health Monitoring

DGX agent

arXiv:2606.26317v1 Announce Type: cross Abstract: Accurate fault diagnosis of rolling element bearings in rotating machinery is considered essential for ensuring industrial safety and enabling predict

model-releasesarxiv-cs-ai
26 Jun 2026
Model Releases

A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation

DGX agent

arXiv:2606.25476v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated remarkable performance across natural language processing tasks, yet their deployment in high-stakes appl

model-releasesarxiv-cs-cl
25 Jun 2026
Model Releases

Auto-Labelling-Based Domain Transfer for 3D Object Detection on a Bicycle-Mounted LiDAR Platform

DGX agent

arXiv:2606.25652v1 Announce Type: new Abstract: Reliable 3D perception of vulnerable road users (VRUs) such as cyclists and pedestrians is essential for their safety in urban traffic and a core requir

model-releasesarxiv-cs-cv
25 Jun 2026
Model Releases

Distribution Preference Optimization: A Fine-grained Perspective for LLM Unlearning

DGX agent

arXiv:2510.04773v2 Announce Type: replace Abstract: As Large Language Models (LLMs) demonstrate remarkable capabilities learned from vast corpora, concerns regarding data privacy and safety are receiv

model-releasesarxiv-cs-lg
25 Jun 2026
Model Releases

Model Forensics: Investigating Whether Concerning Behavior Reflects Misalignment

DGX agent

arXiv:2606.26071v1 Announce Type: new Abstract: A central goal of safety research is determining whether a model is misaligned. Prior work has largely focused on detecting concerning behavior. But beh

model-releasesarxiv-cs-lg
25 Jun 2026
Model Releases

PDS Joint: A Parametric Double-Spiral Joint Tailored for Dexterous Hands

DGX agent

arXiv:2606.24377v1 Announce Type: new Abstract: Compliant joints can embed safety and adaptability into dexterous hands, but achieving large-stroke anthropomorphic motion while maintaining joint-speci

model-releasesarxiv-cs-ro
24 Jun 2026
Local Ai

Local Causal Attribution of Chain-of-Thought Reasoning

DGX agent

arXiv:2606.21821v1 Announce Type: new Abstract: Understanding the causal structure of a language model's thought process is a problem of significant importance for both transparency and safety. In thi

local-aiarxiv-cs-lg
23 Jun 2026
Model Releases

LOGOS: LiDAR-Only Gaussian Elevation Splatting for Unified Tiny Obstacle Segmentation

DGX agent

arXiv:2606.21527v1 Announce Type: cross Abstract: Robust obstacle segmentation is essential for the safety of intelligent robots, where LiDAR-based perception systems play a fundamental role in the ro

model-releasesarxiv-cs-cv
23 Jun 2026
Model Releases

Mirage: a Clean-Label Backdoor against LiDAR 3D Object Detection

DGX agent

arXiv:2606.20752v1 Announce Type: new Abstract: Deep neural network-based LiDAR 3D object detection serves as a critical perception component in safety-critical autonomous systems. However, recent stu

model-releasesarxiv-cs-cv
23 Jun 2026
Model Releases

A Lightweight Multi-Agent Framework for Automated Concrete Barrier Design

DGX agent

arXiv:2606.12040v1 Announce Type: new Abstract: The design of reinforced concrete highway barriers is a safety-critical process that requires strict compliance with regulatory provisions such as the A

model-releasesarxiv-cs-ai
11 Jun 2026
Model Releases

PCS-UQ: Uncertainty Quantification via the Predictability-Computability-Stability Framework

DGX agent

arXiv:2505.08784v2 Announce Type: replace-cross Abstract: As machine learning (ML) enters high-stakes domains, trustworthy uncertainty quantification (UQ) is essential for safety. In this paper we int

model-releasesarxiv-cs-lg
11 Jun 2026
Model Releases

SceneMiner: Identity-Preserving Multi-Task Fine-Tuning for Unified BEV Scene Mining

DGX agent

arXiv:2606.11507v1 Announce Type: new Abstract: Mining hard, safety-critical scenes from driving logs is bottlenecked by the absence of difficulty labels, and no single proxy, collision risk, trajecto

model-releasesarxiv-cs-cv
11 Jun 2026
Model Releases

[AINews] Anthropic Claude Fable 5 — Mythos but Safe, with Controversial Terms

DGX agent

This article from Latent Space discusses Anthropic's Claude Fable 5 model, examining its capabilities in handling creative and mythological content while maintaining safety guardrails, along with cove

model-releaseslatent-space
10 Jun 2026
Tutorials

A message to Anthropic leadership: You're not special. Making sure AI goes well is a team effort not a 'you effort.'

DGX agent

Jeremy Howard argues that Anthropic's leadership should recognize that ensuring AI safety and positive outcomes requires collaborative effort across the industry rather than positioning any single org

tutorialsjeremy-howard--x
9 Jun 2026
Model Releases

Claude Mythos went from “too dangerous to release” to publicly available (with some extra guard rails) in two months. And y’all fell for Ant…

DGX agent

Gary Marcus critiques Anthropic's rapid shift in positioning Claude from a model deemed too dangerous for public release to one made widely available with safety measures, suggesting this represents i

model-releasesgary-marcus--x
9 Jun 2026
Model Releases

Decoding Pedestrian Crossing Intention from Egocentric Vision via Vision Language Models

DGX agent

arXiv:2606.09142v1 Announce Type: cross Abstract: Egocentric vision offers a first-person view of human perception and decision making, yet its potential for traffic-safety prediction remains underexp

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

Defending Against Malicious Finetuning by Scaling Train-time Adversarial Attacks

DGX agent

arXiv:2606.07970v1 Announce Type: cross Abstract: Current open-weight large language models (LLMs) are prone to malicious finetuning attacks, which could compromise the safety alignment of LLMs with o

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

Driving Video Retrieval for Complex Queries with Structured Grounding

DGX agent

arXiv:2606.09109v1 Announce Type: new Abstract: Video retrieval at scale is central to data curation and safety validation in autonomous driving, where users want to find not only scenes but also dyna

model-releasesarxiv-cs-cv
9 Jun 2026
Model Releases

Hybrid Robustness Verification for Spatio-Temporal Neural Networks

DGX agent

arXiv:2606.09746v1 Announce Type: cross Abstract: With AI increasingly deployed in safety-critical systems, providing formal robustness guarantees for the underlying models is essential. Existing veri

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

Learning Predictive Control with Deep Koopman Operators for Autonomous Vehicle Motion Planning

DGX agent

arXiv:2606.08136v1 Announce Type: new Abstract: Model Predictive Control (MPC) is widely used for autonomous-vehicle (AV) motion planning, but its real-time applicability is often limited by the need

model-releasesarxiv-cs-ro
9 Jun 2026
← Previous
1…275276277278279…299
Next →