AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,349 results
Model Releases

Mechanistic Steering of LLMs Reveals Layer-wise Feature Vulnerabilities in Adversarial Settings

DGX agent

arXiv:2604.23130v1 Announce Type: cross Abstract: Large language models (LLMs) can still be jailbroken into producing harmful outputs despite safety alignment. Existing attacks show this vulnerability

model-releasesarxiv-cs-ai
28 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

The Price of Agreement: Measuring LLM Sycophancy in Agentic Financial Applications

DGX agent

arXiv:2604.24668v1 Announce Type: new Abstract: Given the increased use of LLMs in financial systems today, it becomes important to evaluate the safety and robustness of such systems. One failure mode

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

TRACE: Topology-aware Reconstruction of Accidents in CARLA for AV Evaluation

DGX agent

arXiv:2604.22068v1 Announce Type: cross Abstract: Validating Autonomous Vehicles (AVs) requires exposure to rare, safety-critical scenarios, infrequent in routine driving data. Existing benchmarks add

model-releasesarxiv-cs-ro
27 Apr 2026
Model Releases

Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers

DGX agent

arXiv:2604.21700v1 Announce Type: cross Abstract: The growing application of large language models (LLMs) in safety-critical domains has raised urgent concerns about their security. Many recent studie

model-releasesarxiv-cs-ai
24 Apr 2026
Model Releases

Transient Turn Injection: Exposing Stateless Multi-Turn Vulnerabilities in Large Language Models

DGX agent

arXiv:2604.21860v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly integrated into sensitive workflows, raising the stakes for adversarial robustness and safety. This pape

model-releasesarxiv-cs-ai
24 Apr 2026
Local Ai

Ufil: A Unified Framework for Infrastructure-based Localization

DGX agent

arXiv:2604.21471v1 Announce Type: new Abstract: Infrastructure-based localization enhances road safety and traffic management by providing state estimates of road users. Development is hindered by fra

local-aiarxiv-cs-ro
24 Apr 2026
Model Releases

When to Trust the Answer: Question-Aligned Semantic Nearest Neighbor Entropy for Safer Surgical VQA

DGX agent

arXiv:2511.01458v2 Announce Type: replace-cross Abstract: Safety and reliability are critical for deploying visual question answering (VQA) systems in surgery, where incorrect or ambiguous responses c

model-releasesarxiv-cs-ai
24 Apr 2026
Model Releases

1. We believe in iterative deployment; although GPT-5.5 is already a smart model, we expect rapid improvements. Iterative deployment is a bi…

DGX agent

1. We believe in iterative deployment; although GPT-5.5 is already a smart model, we expect rapid improvements. Iterative deployment is a big part of our safety strategy; we believe the world will be

model-releasessam-altman--x
23 Apr 2026
Model Releases

Automated Detection of Dosing Errors in Clinical Trial Narratives: A Multi-Modal Feature Engineering Approach with LightGBM

DGX agent

arXiv:2604.19759v1 Announce Type: new Abstract: Clinical trials require strict adherence to medication protocols, yet dosing errors remain a persistent challenge affecting patient safety and trial int

model-releasesarxiv-cs-ai
23 Apr 2026
Model Releases

Bridging Mechanistic Interpretability and Prompt Engineering with Gradient Ascent for Interpretable Persona Control

DGX agent

arXiv:2601.02896v2 Announce Type: replace Abstract: Controlling emergent behavioral personas (e.g., sycophancy, hallucination) in Large Language Models (LLMs) is critical for AI safety, yet remains a

model-releasesarxiv-cs-lg
23 Apr 2026
Model Releases

CCTVBench: Contrastive Consistency Traffic VideoQA Benchmark for Multimodal LLMs

DGX agent

arXiv:2604.20460v1 Announce Type: new Abstract: Safety-critical traffic reasoning requires contrastive consistency: models must detect true hazards when an accident occurs, and reliably reject plausib

model-releasesarxiv-cs-cv
23 Apr 2026
Model Releases

Evaluating the Quality of the Quantified Uncertainty for (Re)Calibration of Data-Driven Regression Models

DGX agent

arXiv:2508.17761v3 Announce Type: replace Abstract: In safety-critical applications data-driven models must not only be accurate but also provide reliable uncertainty estimates. This property, commonl

model-releasesarxiv-cs-lg
23 Apr 2026
Model Releases

Model Capability Assessment and Safeguards for Biological Weaponization

DGX agent

arXiv:2604.19811v1 Announce Type: cross Abstract: AI leaders and safety reports increasingly warn that advances in model reasoning may enable biological misuse, including by low-expertise users, while

model-releasesarxiv-cs-ai
23 Apr 2026
Model Releases

PipeMFL-240K: A Large-scale Dataset and Benchmark for Object Detection in Pipeline Magnetic Flux Leakage Imaging

DGX agent

arXiv:2602.07044v2 Announce Type: replace-cross Abstract: Pipeline integrity is critical to industrial safety and environmental protection, with Magnetic Flux Leakage (MFL) detection being a primary n

model-releasesarxiv-cs-ai
23 Apr 2026
Local Ai

A Heterogeneous Long-Micro Scale Cascading Architecture for General Aviation Health Management

DGX agent

arXiv:2603.22885v4 Announce Type: replace Abstract: BACKGROUND: General aviation fleet expansion demands intelligent health monitoring under computational constraints. Real-world aircraft health diagn

local-aiarxiv-cs-lg
22 Apr 2026
Model Releases

... and Anthropic reverted this change. Claude Code is now part of Pro, as per the Pricing page. Important note on the growth hack: Anthropi…

DGX agent

... and Anthropic reverted this change. Claude Code is now part of Pro, as per the Pricing page. Important note on the growth hack: Anthropic advertises safety and integrity as their values. A 'fake d

model-releasesjeremy-howard--x
22 Apr 2026
Local Ai

HardNet++: Nonlinear Constraint Enforcement in Neural Networks

DGX agent

arXiv:2604.19669v1 Announce Type: new Abstract: Enforcing constraint satisfaction in neural network outputs is critical for safety, reliability, and physical fidelity in many control and decision-maki

local-aiarxiv-cs-lg
22 Apr 2026
Model Releases

Harness as an Asset: Enforcing Determinism via the Convergent AI Agent Framework (CAAF)

DGX agent

arXiv:2604.17025v1 Announce Type: cross Abstract: Large Language Models (LLMs) produce a controllability gap in safety-critical engineering: even low rates of undetected constraint violations render a

model-releasesarxiv-cs-lg
21 Apr 2026
Model Releases

Learning to Retrieve User History and Generate User Profiles for Personalized Persuasiveness Prediction

DGX agent

arXiv:2601.05654v3 Announce Type: replace Abstract: Estimating the persuasiveness of messages is critical in various applications, from recommender systems to safety assessment of LLMs. While it is im

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Screen Before You Interpret: A Portable Validity Protocol for Benchmark-Based LLM Confidence Signals

DGX agent

arXiv:2604.17714v1 Announce Type: new Abstract: LLM confidence signals are used for abstention, routing, and safety-critical decisions. No standard practice exists for checking whether a confidence si

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Automating Crash Diagram Generation Using Vision-Language Models: A Case Study on Multi-Lane Roundabouts

DGX agent

arXiv:2604.15332v1 Announce Type: cross Abstract: Crash diagrams are essential tools in transportation safety analysis, yet their manual preparation remains time-consuming and prone to human variabili

model-releasesarxiv-cs-ai
20 Apr 2026
Model Releases

LinuxArena: A Control Setting for AI Agents in Live Production Software Environments

DGX agent

arXiv:2604.15384v1 Announce Type: cross Abstract: We introduce LinuxArena, a control setting in which agents operate directly on live, multi-service production environments. LinuxArena contains 20 env

model-releasesarxiv-cs-ai
20 Apr 2026
Model Releases

Pruning Unsafe Tickets: A Resource-Efficient Framework for Safer and More Robust LLMs

DGX agent

arXiv:2604.15780v1 Announce Type: cross Abstract: Machine learning models are increasingly deployed in real-world applications, but even aligned models such as Mistral and LLaVA still exhibit unsafe b

model-releasesarxiv-cs-cl
20 Apr 2026
Model Releases

Reasoning-targeted Jailbreak Attacks on Large Reasoning Models via Semantic Triggers and Psychological Framing

DGX agent

arXiv:2604.15725v1 Announce Type: cross Abstract: Large Reasoning Models (LRMs) have demonstrated strong capabilities in generating step-by-step reasoning chains alongside final answers, enabling thei

model-releasesarxiv-cs-ai
20 Apr 2026
Model Releases

RedBench: A Universal Dataset for Comprehensive Red Teaming of Large Language Models

DGX agent

arXiv:2601.03699v2 Announce Type: replace Abstract: As large language models (LLMs) become integral to safety-critical applications, ensuring their robustness against adversarial prompts is paramount.

model-releasesarxiv-cs-cl
20 Apr 2026
Tutorials

@AmandaAskell are you the person to thank for this?

DGX agent

Amanda Askell is likely a researcher or professional involved in AI safety or alignment work, and Jeremy Howard is publicly crediting or thanking her for a contribution or achievement on social media.

tutorialsjeremy-howard--x
17 Apr 2026
Local Ai

MS-SSE-Net: A Multi-Scale Spatial Squeeze-and-Excitation Network for Structural Damage Detection in Civil and Geotechnical Engineering

DGX agent

arXiv:2604.14711v1 Announce Type: new Abstract: Structural damage detection is essential for maintaining the safety and reliability of civil infrastructure. However, accurately identifying different t

local-aiarxiv-cs-cv
17 Apr 2026
Model Releases

DeEscalWild: A Real-World Benchmark for Automated De-Escalation Training with SLMs

DGX agent

arXiv:2604.13075v1 Announce Type: new Abstract: Effective de-escalation is critical for law enforcement safety and community trust, yet traditional training methods lack scalability and realism. While

model-releasesarxiv-cs-cl
16 Apr 2026
Model Releases

Dental-TriageBench: Benchmarking Multimodal Reasoning for Hierarchical Dental Triage

DGX agent

arXiv:2604.13060v1 Announce Type: new Abstract: Dental triage is a safety-critical clinical routing task that requires integrating multimodal clinical information (e.g., patient complaints and radiogr

model-releasesarxiv-cs-cl
16 Apr 2026
Model Releases

FlexGuard: Continuous Risk Scoring for Strictness-Adaptive LLM Content Moderation

DGX agent

arXiv:2602.23636v3 Announce Type: replace Abstract: Ensuring the safety of LLM-generated content is essential for real-world deployment. Most existing guardrail models formulate moderation as a fixed

model-releasesarxiv-cs-lg
16 Apr 2026
Tools

[AINews] Humanity's Last Gasp

DGX agent

'Humanity's Last Gasp' is an AI news roundup from the Latent Space newsletter, likely covering significant developments in AI safety, existential risk discussions, or major industry milestones that pr

toolslatent-space
15 Apr 2026
Model Releases

GF-Score: Certified Class-Conditional Robustness Evaluation with Fairness Guarantees

DGX agent

arXiv:2604.12757v1 Announce Type: cross Abstract: Adversarial robustness is essential for deploying neural networks in safety-critical applications, yet standard evaluation methods either require expe

model-releasesarxiv-cs-ai
15 Apr 2026
Model Releases

Monte Carlo Stochastic Depth for Uncertainty Estimation in Deep Learning

DGX agent

arXiv:2604.12719v1 Announce Type: new Abstract: The deployment of deep neural networks in safety-critical systems necessitates reliable and efficient uncertainty quantification (UQ). A practical and w

model-releasesarxiv-cs-lg
15 Apr 2026
Model Releases

Reading Between the Pixels: Linking Text-Image Embedding Alignment to Typographic Attack Success on Vision-Language Models

DGX agent

arXiv:2604.12371v1 Announce Type: new Abstract: We study typographic prompt injection attacks on vision-language models (VLMs), where adversarial text is rendered as images to bypass safety mechanisms

model-releasesarxiv-cs-cv
15 Apr 2026
Local Ai

Abliterated (uncensored) models

DGX agent

This r/ollama discussion covers 'abliterated' models — LLMs that have had their built-in refusal mechanisms removed through a technique called abliteration, allowing them to respond to prompts without

local-air-ollama
14 Apr 2026
Model Releases

Cybersecurity Looks Like Proof of Work Now

DGX agent

Cybersecurity Looks Like Proof of Work Now The UK's AI Safety Institute recently published Our evaluation of Claude Mythos Preview’s cyber capabilities, their own independent analysis of Claude Mythos

model-releasessimon-willison
14 Apr 2026
Model Releases

Doc-PP: Document Policy Preservation Benchmark for Large Vision-Language Models

DGX agent

arXiv:2601.03926v2 Announce Type: replace Abstract: The deployment of Large Vision-Language Models (LVLMs) for real-world document question answering is often constrained by dynamic, user-defined poli

model-releasesarxiv-cs-cl
14 Apr 2026
Model Releases

FedKLPR: KL-Guided Pruning-Aware Federated Learning for Person Re-Identification

DGX agent

arXiv:2508.17431v3 Announce Type: replace-cross Abstract: Person re-identification (re-ID) is a fundamental task in intelligent surveillance and public safety. Federated learning (FL) provides a priva

model-releasesarxiv-cs-ai
14 Apr 2026
Local Ai

Global monitoring of methane point sources using deep learning on hyperspectral radiance measurements from EMIT

DGX agent

arXiv:2604.10094v1 Announce Type: new Abstract: Anthropogenic methane (CH4) point sources drive near-term climate forcing, safety hazards, and system inefficiencies. Space-based imaging spectroscopy i

local-aiarxiv-cs-cv
14 Apr 2026
Model Releases

OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language World Models

DGX agent

arXiv:2604.10866v1 Announce Type: new Abstract: AI agents are expected to perform professional work across hundreds of occupational domains (from emergency department triage to nuclear reactor safety

model-releasesarxiv-cs-cl
14 Apr 2026
Model Releases

Quantitative Introspection in Language Models: Tracking Emotive States Across Conversation

DGX agent

arXiv:2603.18893v2 Announce Type: replace Abstract: Tracking the internal states of large language models across conversations is important for safety, interpretability, and model welfare, yet current

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

SimScale: Learning to Drive via Real-World Simulation at Scale

DGX agent

arXiv:2511.23369v3 Announce Type: replace Abstract: Achieving fully autonomous driving systems requires learning rational decisions in a wide span of scenarios, including safety-critical and out-of-di

model-releasesarxiv-cs-cv
13 Apr 2026
Model Releases

VOLTA: The Surprising Ineffectiveness of Auxiliary Losses for Calibrated Deep Learning

DGX agent

arXiv:2604.08639v1 Announce Type: cross Abstract: Uncertainty quantification (UQ) is essential for deploying deep learning models in safety critical applications, yet no consensus exists on which UQ m

model-releasesarxiv-cs-ai
13 Apr 2026
Model Releases

AgentOpt v0.1 Technical Report: Client-Side Optimization for LLM-Based Agent

DGX agent

arXiv:2604.06296v1 Announce Type: cross Abstract: AI agents are increasingly deployed in real-world applications, including systems such as Manus, OpenClaw, and coding agents. Existing research has pr

model-releasesarxiv-cs-ai
10 Apr 2026
Local Ai

CHiQPM: Calibrated Hierarchical Interpretable Image Classification

DGX agent

arXiv:2511.20779v2 Announce Type: replace Abstract: Globally interpretable models are a promising approach for trustworthy AI in safety-critical domains. Alongside global explanations, detailed local

local-aiarxiv-cs-lg
10 Apr 2026
Model Releases

SearchAD: Large-Scale Rare Image Retrieval Dataset for Autonomous Driving

DGX agent

arXiv:2604.08008v1 Announce Type: new Abstract: Retrieving rare and safety-critical driving scenarios from large-scale datasets is essential for building robust autonomous driving (AD) systems. As dat

model-releasesarxiv-cs-cv
10 Apr 2026
Model Releases

Florida launches investigation into OpenAI

DGX agent

Florida Attorney General James Uthmeier is launching an investigation into OpenAI over public safety and national security risks, as reported earlier by Reuters. In a statement on Thursday, Uthmeier s

model-releasesthe-verge-ai
9 Apr 2026
Model Releases

Benchmarking LLM-Guided Control-Plane Policies for Backend Fault Isolation in HAProxy

DGX agent

arXiv:2608.10532v1 Announce Type: cross Abstract: Static load balancers cannot mitigate a backend that is degraded rather than down: round-robin and least-connections keep routing traffic to a server

model-releasesarxiv-cs-lg
12 Aug 2026
← Previous
1…278279280281282…299
Next →