AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlog
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,237 results
Model Releases

Evaluating the Quality of the Quantified Uncertainty for (Re)Calibration of Data-Driven Regression Models

DGX agent

arXiv:2508.17761v3 Announce Type: replace Abstract: In safety-critical applications data-driven models must not only be accurate but also provide reliable uncertainty estimates. This property, commonl

model-releasesarxiv-cs-lg
23 Apr 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Model Capability Assessment and Safeguards for Biological Weaponization

DGX agent

arXiv:2604.19811v1 Announce Type: cross Abstract: AI leaders and safety reports increasingly warn that advances in model reasoning may enable biological misuse, including by low-expertise users, while

model-releasesarxiv-cs-ai
23 Apr 2026
Model Releases

PipeMFL-240K: A Large-scale Dataset and Benchmark for Object Detection in Pipeline Magnetic Flux Leakage Imaging

DGX agent

arXiv:2602.07044v2 Announce Type: replace-cross Abstract: Pipeline integrity is critical to industrial safety and environmental protection, with Magnetic Flux Leakage (MFL) detection being a primary n

model-releasesarxiv-cs-ai
23 Apr 2026
Local Ai

A Heterogeneous Long-Micro Scale Cascading Architecture for General Aviation Health Management

DGX agent

arXiv:2603.22885v4 Announce Type: replace Abstract: BACKGROUND: General aviation fleet expansion demands intelligent health monitoring under computational constraints. Real-world aircraft health diagn

local-aiarxiv-cs-lg
22 Apr 2026
Model Releases

... and Anthropic reverted this change. Claude Code is now part of Pro, as per the Pricing page. Important note on the growth hack: Anthropi…

DGX agent

... and Anthropic reverted this change. Claude Code is now part of Pro, as per the Pricing page. Important note on the growth hack: Anthropic advertises safety and integrity as their values. A 'fake d

model-releasesjeremy-howard--x
22 Apr 2026
Local Ai

HardNet++: Nonlinear Constraint Enforcement in Neural Networks

DGX agent

arXiv:2604.19669v1 Announce Type: new Abstract: Enforcing constraint satisfaction in neural network outputs is critical for safety, reliability, and physical fidelity in many control and decision-maki

local-aiarxiv-cs-lg
22 Apr 2026
Model Releases

Harness as an Asset: Enforcing Determinism via the Convergent AI Agent Framework (CAAF)

DGX agent

arXiv:2604.17025v1 Announce Type: cross Abstract: Large Language Models (LLMs) produce a controllability gap in safety-critical engineering: even low rates of undetected constraint violations render a

model-releasesarxiv-cs-lg
21 Apr 2026
Model Releases

Learning to Retrieve User History and Generate User Profiles for Personalized Persuasiveness Prediction

DGX agent

arXiv:2601.05654v3 Announce Type: replace Abstract: Estimating the persuasiveness of messages is critical in various applications, from recommender systems to safety assessment of LLMs. While it is im

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Screen Before You Interpret: A Portable Validity Protocol for Benchmark-Based LLM Confidence Signals

DGX agent

arXiv:2604.17714v1 Announce Type: new Abstract: LLM confidence signals are used for abstention, routing, and safety-critical decisions. No standard practice exists for checking whether a confidence si

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Automating Crash Diagram Generation Using Vision-Language Models: A Case Study on Multi-Lane Roundabouts

DGX agent

arXiv:2604.15332v1 Announce Type: cross Abstract: Crash diagrams are essential tools in transportation safety analysis, yet their manual preparation remains time-consuming and prone to human variabili

model-releasesarxiv-cs-ai
20 Apr 2026
Model Releases

LinuxArena: A Control Setting for AI Agents in Live Production Software Environments

DGX agent

arXiv:2604.15384v1 Announce Type: cross Abstract: We introduce LinuxArena, a control setting in which agents operate directly on live, multi-service production environments. LinuxArena contains 20 env

model-releasesarxiv-cs-ai
20 Apr 2026
Model Releases

Pruning Unsafe Tickets: A Resource-Efficient Framework for Safer and More Robust LLMs

DGX agent

arXiv:2604.15780v1 Announce Type: cross Abstract: Machine learning models are increasingly deployed in real-world applications, but even aligned models such as Mistral and LLaVA still exhibit unsafe b

model-releasesarxiv-cs-cl
20 Apr 2026
Model Releases

Reasoning-targeted Jailbreak Attacks on Large Reasoning Models via Semantic Triggers and Psychological Framing

DGX agent

arXiv:2604.15725v1 Announce Type: cross Abstract: Large Reasoning Models (LRMs) have demonstrated strong capabilities in generating step-by-step reasoning chains alongside final answers, enabling thei

model-releasesarxiv-cs-ai
20 Apr 2026
Model Releases

RedBench: A Universal Dataset for Comprehensive Red Teaming of Large Language Models

DGX agent

arXiv:2601.03699v2 Announce Type: replace Abstract: As large language models (LLMs) become integral to safety-critical applications, ensuring their robustness against adversarial prompts is paramount.

model-releasesarxiv-cs-cl
20 Apr 2026
Tutorials

@AmandaAskell are you the person to thank for this?

DGX agent

Amanda Askell is likely a researcher or professional involved in AI safety or alignment work, and Jeremy Howard is publicly crediting or thanking her for a contribution or achievement on social media.

tutorialsjeremy-howard--x
17 Apr 2026
Local Ai

MS-SSE-Net: A Multi-Scale Spatial Squeeze-and-Excitation Network for Structural Damage Detection in Civil and Geotechnical Engineering

DGX agent

arXiv:2604.14711v1 Announce Type: new Abstract: Structural damage detection is essential for maintaining the safety and reliability of civil infrastructure. However, accurately identifying different t

local-aiarxiv-cs-cv
17 Apr 2026
Model Releases

DeEscalWild: A Real-World Benchmark for Automated De-Escalation Training with SLMs

DGX agent

arXiv:2604.13075v1 Announce Type: new Abstract: Effective de-escalation is critical for law enforcement safety and community trust, yet traditional training methods lack scalability and realism. While

model-releasesarxiv-cs-cl
16 Apr 2026
Model Releases

Dental-TriageBench: Benchmarking Multimodal Reasoning for Hierarchical Dental Triage

DGX agent

arXiv:2604.13060v1 Announce Type: new Abstract: Dental triage is a safety-critical clinical routing task that requires integrating multimodal clinical information (e.g., patient complaints and radiogr

model-releasesarxiv-cs-cl
16 Apr 2026
Model Releases

FlexGuard: Continuous Risk Scoring for Strictness-Adaptive LLM Content Moderation

DGX agent

arXiv:2602.23636v3 Announce Type: replace Abstract: Ensuring the safety of LLM-generated content is essential for real-world deployment. Most existing guardrail models formulate moderation as a fixed

model-releasesarxiv-cs-lg
16 Apr 2026
Tools

[AINews] Humanity's Last Gasp

DGX agent

'Humanity's Last Gasp' is an AI news roundup from the Latent Space newsletter, likely covering significant developments in AI safety, existential risk discussions, or major industry milestones that pr

toolslatent-space
15 Apr 2026
Model Releases

GF-Score: Certified Class-Conditional Robustness Evaluation with Fairness Guarantees

DGX agent

arXiv:2604.12757v1 Announce Type: cross Abstract: Adversarial robustness is essential for deploying neural networks in safety-critical applications, yet standard evaluation methods either require expe

model-releasesarxiv-cs-ai
15 Apr 2026
Model Releases

Monte Carlo Stochastic Depth for Uncertainty Estimation in Deep Learning

DGX agent

arXiv:2604.12719v1 Announce Type: new Abstract: The deployment of deep neural networks in safety-critical systems necessitates reliable and efficient uncertainty quantification (UQ). A practical and w

model-releasesarxiv-cs-lg
15 Apr 2026
Model Releases

Reading Between the Pixels: Linking Text-Image Embedding Alignment to Typographic Attack Success on Vision-Language Models

DGX agent

arXiv:2604.12371v1 Announce Type: new Abstract: We study typographic prompt injection attacks on vision-language models (VLMs), where adversarial text is rendered as images to bypass safety mechanisms

model-releasesarxiv-cs-cv
15 Apr 2026
Local Ai

Abliterated (uncensored) models

DGX agent

This r/ollama discussion covers 'abliterated' models — LLMs that have had their built-in refusal mechanisms removed through a technique called abliteration, allowing them to respond to prompts without

local-air-ollama
14 Apr 2026
Model Releases

Cybersecurity Looks Like Proof of Work Now

DGX agent

Cybersecurity Looks Like Proof of Work Now The UK's AI Safety Institute recently published Our evaluation of Claude Mythos Preview’s cyber capabilities, their own independent analysis of Claude Mythos

model-releasessimon-willison
14 Apr 2026
Model Releases

Doc-PP: Document Policy Preservation Benchmark for Large Vision-Language Models

DGX agent

arXiv:2601.03926v2 Announce Type: replace Abstract: The deployment of Large Vision-Language Models (LVLMs) for real-world document question answering is often constrained by dynamic, user-defined poli

model-releasesarxiv-cs-cl
14 Apr 2026
Model Releases

FedKLPR: KL-Guided Pruning-Aware Federated Learning for Person Re-Identification

DGX agent

arXiv:2508.17431v3 Announce Type: replace-cross Abstract: Person re-identification (re-ID) is a fundamental task in intelligent surveillance and public safety. Federated learning (FL) provides a priva

model-releasesarxiv-cs-ai
14 Apr 2026
Local Ai

Global monitoring of methane point sources using deep learning on hyperspectral radiance measurements from EMIT

DGX agent

arXiv:2604.10094v1 Announce Type: new Abstract: Anthropogenic methane (CH4) point sources drive near-term climate forcing, safety hazards, and system inefficiencies. Space-based imaging spectroscopy i

local-aiarxiv-cs-cv
14 Apr 2026
Model Releases

OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language World Models

DGX agent

arXiv:2604.10866v1 Announce Type: new Abstract: AI agents are expected to perform professional work across hundreds of occupational domains (from emergency department triage to nuclear reactor safety

model-releasesarxiv-cs-cl
14 Apr 2026
Model Releases

Quantitative Introspection in Language Models: Tracking Emotive States Across Conversation

DGX agent

arXiv:2603.18893v2 Announce Type: replace Abstract: Tracking the internal states of large language models across conversations is important for safety, interpretability, and model welfare, yet current

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

SimScale: Learning to Drive via Real-World Simulation at Scale

DGX agent

arXiv:2511.23369v3 Announce Type: replace Abstract: Achieving fully autonomous driving systems requires learning rational decisions in a wide span of scenarios, including safety-critical and out-of-di

model-releasesarxiv-cs-cv
13 Apr 2026
Model Releases

VOLTA: The Surprising Ineffectiveness of Auxiliary Losses for Calibrated Deep Learning

DGX agent

arXiv:2604.08639v1 Announce Type: cross Abstract: Uncertainty quantification (UQ) is essential for deploying deep learning models in safety critical applications, yet no consensus exists on which UQ m

model-releasesarxiv-cs-ai
13 Apr 2026
Model Releases

AgentOpt v0.1 Technical Report: Client-Side Optimization for LLM-Based Agent

DGX agent

arXiv:2604.06296v1 Announce Type: cross Abstract: AI agents are increasingly deployed in real-world applications, including systems such as Manus, OpenClaw, and coding agents. Existing research has pr

model-releasesarxiv-cs-ai
10 Apr 2026
Local Ai

CHiQPM: Calibrated Hierarchical Interpretable Image Classification

DGX agent

arXiv:2511.20779v2 Announce Type: replace Abstract: Globally interpretable models are a promising approach for trustworthy AI in safety-critical domains. Alongside global explanations, detailed local

local-aiarxiv-cs-lg
10 Apr 2026
Model Releases

SearchAD: Large-Scale Rare Image Retrieval Dataset for Autonomous Driving

DGX agent

arXiv:2604.08008v1 Announce Type: new Abstract: Retrieving rare and safety-critical driving scenarios from large-scale datasets is essential for building robust autonomous driving (AD) systems. As dat

model-releasesarxiv-cs-cv
10 Apr 2026
Model Releases

Florida launches investigation into OpenAI

DGX agent

Florida Attorney General James Uthmeier is launching an investigation into OpenAI over public safety and national security risks, as reported earlier by Reuters. In a statement on Thursday, Uthmeier s

model-releasesthe-verge-ai
9 Apr 2026
Model Releases

Benchmarking LLM-Guided Control-Plane Policies for Backend Fault Isolation in HAProxy

DGX agent

arXiv:2608.10532v1 Announce Type: cross Abstract: Static load balancers cannot mitigate a backend that is degraded rather than down: round-robin and least-connections keep routing traffic to a server

model-releasesarxiv-cs-lg
12 Aug 2026
Local Ai

Certify or Refuse: A Cross-Model Map for Selective Risk Control with Coverage Floors under Covariate Shift

DGX agent

arXiv:2608.10893v1 Announce Type: new Abstract: Certified selective predictors attain whatever coverage they attain; operators impose an automation floor: answer at least a eta-fraction of shifted tar

local-aiarxiv-cs-cl
12 Aug 2026
Model Releases

Diffract: Spectral View of LLM Domain Adaptation

DGX agent

arXiv:2608.10850v1 Announce Type: new Abstract: We study continual pre-training (CPT) as a mechanism for adapting general-purpose large language models to specialized domains: mathematics, instruction

model-releasesarxiv-cs-lg
12 Aug 2026
Model Releases

Logit-Boundary Geometric Belief Interfaces and Sparse Sheaf-Enclave Protocols: A Self-Contained Substrate for Secure Network Electronic Health Record (EHR) Interoperability

DGX agent

arXiv:2608.10300v1 Announce Type: new Abstract: Electronic health-record interoperability is a boundary problem: legacy systems, generative models, terminology services, identity systems, and human re

model-releasesarxiv-cs-ai
12 Aug 2026
Model Releases

The Truth Stays in the Family: Enhancing Contextual Grounding via Inherited Truthful Heads in Model Lineages

DGX agent

arXiv:2606.15821v2 Announce Type: replace-cross Abstract: Recent advances in large language models (LLMs) have produced many specialized multimodal LLMs (MLLMs) that share common foundational LLMs, fo

model-releasesarxiv-cs-ai
12 Aug 2026
Local Ai

Toward the Cognitive--Physical Limits of Embodied Intelligence through a World-Model-Centric Autonomous Racing Agent

DGX agent

arXiv:2608.10618v1 Announce Type: new Abstract: Embodied artificial intelligence aims to develop agents that perceive, reason, and act through continuous interaction with the physical world. However,

local-aiarxiv-cs-ro
12 Aug 2026
Model Releases

We are in an insane run of open-weight drops. Every modality, open source is winning. This is what an open source AI summer ☀️ looks like: …

DGX agent

We are in an insane run of open-weight drops. Every modality, open source is winning. This is what an open source AI summer ☀️ looks like: 🧠 LLMs & Reasoning → DeepSeek-V4-Flash-0731 (my king 👑): 304B

model-releasesclem-delangue--x
12 Aug 2026
Model Releases

Automating Deception: Scalable Multi-Turn LLM Jailbreaks

DGX agent

arXiv:2511.19517v3 Announce Type: replace-cross Abstract: Multi-turn conversational attacks, which leverage psychological principles like Foot-in-the-Door (FITD), where a small initial request paves t

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness

DGX agent

arXiv:2608.09900v1 Announce Type: new Abstract: Large language model evaluations typically focus on performance under nominal conditions, creating an illusion of capability where models comfortably wa

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

MetaSpace: Metamorphic Testing for Spatial Cognition in Embodied Agents

DGX agent

arXiv:2608.07533v1 Announce Type: new Abstract: An embodied agent is an intelligent entity that interacts with its environment through a physical body. Currently, the evaluation of embodied agents pri

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Verication-driven closed-loop multi-agent large language modelframework for code-compliant structural design

DGX agent

arXiv:2608.07978v1 Announce Type: cross Abstract: Multi-agent large language model(LLM)systems are applied to structural design,yet most use one-shot generation and cannot verify their output,leaving

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

When Do Task Vectors Interfere? Mapping the Validity Boundaries of Weight-Space Composition

DGX agent

arXiv:2608.09490v1 Announce Type: new Abstract: Task arithmetic treats fine-tuning displacements as composable directions in weight space, yet it remains unclear when parameter addition reflects predi

model-releasesarxiv-cs-lg
11 Aug 2026
← Previous
1…276277278279280…297
Next →