AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,237 results
15 May 2026

Good to Go: The LOOP Skill Engine That Hits 99% Success and Slashes Token Usage by 99% via One-Shot Recording and Deterministic Replay

Model ReleasesDGX agent

arXiv:2605.14237v1 Announce Type: new Abstract: Deploying AI agents for repetitive periodic tasks exposes a critical tension: Large Language Models (LLMs) offer unmatched flexibility in tool orchestra

NodeSynth: Socially Aligned Synthetic Data for AI Evaluation

Model ReleasesDGX agent

arXiv:2605.14381v1 Announce Type: cross Abstract: Recent advancements in generative AI facilitate large-scale synthetic data generation for model evaluation. However, without targeted approaches, thes

Talk is (Not) Cheap: A Taxonomy and Benchmark Coverage Audit for LLM Attacks

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.15118v1 Announce Type: cross Abstract: We introduce a reusable framework for auditing whether LLM attack benchmarks collectively cover the threat surface: a 4imes6 Target imes Technique mat

The Moltbook Observatory Archive: an incremental dataset of agent-only social network activity

Model ReleasesDGX agent

arXiv:2605.13860v1 Announce Type: cross Abstract: Moltbook is a social media platform in which posts and comments are authored exclusively by autonomous AI agents. We present the Moltbook Observatory

The most revealing thing about this AI leadership paper is that it reads less like a vision for innovation and more like a glossy whitepaper…

Model ReleasesDGX agent

The most revealing thing about this AI leadership paper is that it reads less like a vision for innovation and more like a glossy whitepaper for a 21st century East India Company. Every generation of

Vision-Based Runtime Monitoring under Varying Specifications using Semantic Latent Representations

Model ReleasesDGX agent

arXiv:2605.13923v1 Announce Type: cross Abstract: We study certified runtime monitoring of past-time signal temporal logic (ptSTL) from visual observations under partial observability. The monitor mus

14 May 2026

Action Emergence from Streaming Intent

Model ReleasesDGX agent

arXiv:2605.12622v1 Announce Type: cross Abstract: We formalize action emergence as a target capability for end-to-end autonomous driving: the ability to generate physically feasible, semantically appr

Connecting the Dots: A Machine Learning Ready Dataset for Ionospheric Forecasting Models

Model ReleasesDGX agent

arXiv:2511.15743v2 Announce Type: replace Abstract: Operational forecasting of the ionosphere remains a critical space weather challenge due to sparse observations, complex coupling across geospatial

Negation Neglect: When models fail to learn negations in training

Model ReleasesDGX agent

arXiv:2605.13829v1 Announce Type: cross Abstract: We introduce Negation Neglect, where finetuning LLMs on documents that flag a claim as false makes them believe the claim is true. For example, models

Vaporware or not? Aptera assembles its first five validation models.

IndustryDGX agent

Aptera Motors announced that five validation vehicles have been driven off its newly established low-volume validation assembly line in Carlsbad, California in May 2026. Running multiple vehicles thro

13 May 2026

AcuityBench: Evaluating Clinical Acuity Identification and Uncertainty Alignment

Model ReleasesDGX agent

arXiv:2605.11398v1 Announce Type: cross Abstract: We introduce AcuityBench, a benchmark for evaluating whether language models identify the appropriate urgency of care from user medical presentations.

AutoMonitor-Bench: Evaluating the Reliability of LLM-Based Misbehavior Monitor

Model ReleasesDGX agent

arXiv:2601.05752v3 Announce Type: replace Abstract: We introduce AutoMonitor-Bench, the first benchmark designed to systematically evaluate the reliability of LLM-based misbehavior monitors across div

Crash Assessment via Mesh-Based Graph Neural Networks and Physics-Aware Attention

Model ReleasesDGX agent

arXiv:2605.11784v1 Announce Type: cross Abstract: Full-vehicle crash simulations are computationally expensive, limiting their use in iterative design exploration. This work investigates learned hybri

Multi-Task Representation Learning for Conservative Linear Bandits

Model ReleasesDGX agent

arXiv:2605.12176v1 Announce Type: new Abstract: This paper presents the Constrained Multi-Task Representation Learning (CMTRL) framework for linear bandits. We consider T linear bandit tasks in a d di

Qwen-Scope: Turning Sparse Features into Development Tools for Large Language Models

Model ReleasesDGX agent

arXiv:2605.11887v1 Announce Type: new Abstract: Large language models have achieved remarkable capabilities across diverse tasks, yet their internal decision-making processes remain largely opaque, li

Robust Promptable Video Object Segmentation

Model ReleasesDGX agent

arXiv:2605.12006v1 Announce Type: new Abstract: The performance of promptable video object segmentation (PVOS) models substantially degrades under input corruptions, which prevents PVOS deployment in

Urban Risk-Aware Navigation via VQA-Based Event Maps for People with Low Vision

Model ReleasesDGX agent

arXiv:2605.11782v1 Announce Type: new Abstract: Visual impairment affects hundreds of millions of people worldwide, severely limiting their ability to navigate urban environments safely and independen

12 May 2026

Aligning Agents via Planning: A Benchmark for Trajectory-Level Reward Modeling

Model ReleasesDGX agent

arXiv:2604.08178v2 Announce Type: replace Abstract: In classical Reinforcement Learning from Human Feedback (RLHF), Reward Models (RMs) serve as the fundamental signal provider for model alignment. As

ChatbotManip: A Dataset to Facilitate Evaluation and Oversight of Manipulative Chatbot Behaviour

Model ReleasesDGX agent

arXiv:2506.12090v2 Announce Type: replace Abstract: This paper introduces ChatbotManip, a novel dataset for studying manipulation in Chatbots. It contains simulated generated conversations between a c

Decomposing and Steering Functional Metacognition in Large Language Models

Model ReleasesDGX agent

arXiv:2605.08942v1 Announce Type: new Abstract: Large language models (LLMs) increasingly exhibit behaviors suggesting awareness of their evaluation context, often adapting their reasoning strategies

GONE: Structural Knowledge Unlearning via Neighborhood-Expanded Distribution Shaping

Model ReleasesDGX agent

arXiv:2603.12275v1 Announce Type: cross Abstract: Unlearning knowledge is a pressing and challenging task in Large Language Models (LLMs) because of their unprecedented capability to memorize and dige

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs

Model ReleasesDGX agent

arXiv:2508.20325v3 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) become increasingly integral to various domains, their potential to generate harmful responses has prompted si

HiDrive: A Closed-Loop Benchmark for High-Level Autonomous Driving

Model ReleasesDGX agent

arXiv:2605.09972v1 Announce Type: cross Abstract: End-to-end autonomous driving has witnessed rapid progress, yet existing benchmarks are increasingly saturated, with state-of-the-art models achieving

Interactive Critique-Revision Training for Reliable Structured LLM Generation

Local AiDGX agent

arXiv:2605.08327v1 Announce Type: cross Abstract: In structured decision-making workflows such as form filling, compliance checking, and maintenance reporting, LLM outputs must be locally correct, glo

Less Diverse, Less Safe: The Indirect But Pervasive Risk of Test-Time Scaling in Large Language Models

Model ReleasesDGX agent

arXiv:2510.08592v3 Announce Type: replace-cross Abstract: Test-Time Scaling (TTS) improves LLM reasoning by exploring multiple candidate responses and then operating over this set to find the best out

Measuring What Matters: Benchmarking Generative, Multimodal, and Agentic AI in Healthcare

Model ReleasesDGX agent

arXiv:2605.08445v1 Announce Type: new Abstract: AI models are increasingly deployed in live clinical environments where they must perform reliably across complex, high-stakes workflows that standard t

Med-StepBench: A Hierarchical Reasoning Framework for Evaluating Hallucinations in Medical Vision-Language Models

Model ReleasesDGX agent

arXiv:2605.10002v1 Announce Type: new Abstract: Large vision-language models (VLMs) demonstrate strong performance in medical image understanding, but frequently generate clinically plausible yet inco

Medical Incident Causal Factors and Preventive Measures Generation Using Tag-based Example Selection in Few-shot Learning

Model ReleasesDGX agent

arXiv:2605.10025v1 Announce Type: cross Abstract: In high-stakes domains such as healthcare, the reliability of Large Language Models (LLMs) is critical, particularly when generating clinical insights

Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks

Model ReleasesDGX agent

arXiv:2605.10639v1 Announce Type: new Abstract: The rapid adoption of LLMs in both research and industry highlights the challenges of deploying them safely and reveals a gap in the systematic evaluati

Phoenix-VL 1.5 Medium Technical Report

Model ReleasesDGX agent

arXiv:2605.10391v1 Announce Type: cross Abstract: We introduce Phoenix-VL 1.5 Medium, a 123B-parameter natively multimodal and multilingual foundation model, adapted to regional languages and the Sing

Playing games with knowledge: AI-Induced delusions need game theoretic interventions

Local AiDGX agent

arXiv:2605.08409v1 Announce Type: new Abstract: Conversational AI has a fundamental flaw as a knowledge interface: sycophantic chatbots induce epistemic entrenchment and delusional belief spirals even

Position: AI Security Policy Should Target Systems, Not Models

Model ReleasesDGX agent

arXiv:2605.09504v1 Announce Type: cross Abstract: We present swarm-attack, an open-source adversarial testing framework in which multiple lightweight LLM agents coordinate through shared memory, paral

PYTHALAB-MERA: Validation-Grounded Memory, Retrieval, and Acceptance Control for Frozen-LLM Coding Agents

Local AiDGX agent

arXiv:2605.08468v1 Announce Type: cross Abstract: Local LLM-based coding agents increasingly work in settings where correctness is earned through execution feedback, persistent state, and bounded repa

Scaling Vision Models Does Not Consistently Improve Localisation-Based Explanation Quality

Model ReleasesDGX agent

arXiv:2605.10142v1 Announce Type: cross Abstract: Artificial intelligence models are increasingly scaled to improve predictive accuracy, yet it remains unclear whether scale improves the quality of po

Scam2Prompt: A Scalable Framework for Auditing Malicious Scam Endpoints in Production LLMs

Model ReleasesDGX agent

arXiv:2509.02372v3 Announce Type: replace-cross Abstract: Large Language Models have become critical to modern software development, but their reliance on uncurated web-scale datasets for training int

Steerable but Not Decodable: Function Vectors Operate Beyond the Logit Lens

Model ReleasesDGX agent

arXiv:2604.02608v2 Announce Type: replace Abstract: Activation steering presupposes that task-relevant behaviors correspond to linear directions in activation space -- directions that should both stee

The Echo Amplifies the Knowledge: Somatic Marker Analogues in Language Models via Emotion Vector Re-Injection

Model ReleasesDGX agent

arXiv:2605.08611v1 Announce Type: new Abstract: Current language model memory systems store what happened but not how it felt. This distinction -- between semantic memory (knowing about a past event)

The Extrapolation Cliff in On-Policy Distillation of Near-Deterministic Structured Outputs

Model ReleasesDGX agent

arXiv:2605.08737v1 Announce Type: cross Abstract: On-policy distillation (OPD) is widely used for LLM post-training. When pushed with a reward-extrapolation coefficient lambda > 1, the student can lif

What's the plan? Metrics for implicit planning in LLMs and their application to rhyme generation and question answering

Model ReleasesDGX agent

arXiv:2601.20164v2 Announce Type: replace-cross Abstract: Prior work suggests that language models, while trained on next token prediction, show implicit planning behavior: they may select the next to

11 May 2026

Agents + file sandboxes are all in the range in 2026 🤖🗃️ This is a nifty reference implementation by @itsclelia showing you how to run you…

Model ReleasesDGX agent

Agents + file sandboxes are all in the range in 2026 🤖🗃️ This is a nifty reference implementation by @itsclelia showing you how to run your agent over a collection of docs (PDFs, images, Office) with

Beyond the Black Box: Interpretability of Agentic AI Tool Use

Model ReleasesDGX agent

arXiv:2605.06890v1 Announce Type: new Abstract: AI agents are promising for high-stakes enterprise workflows, but dependable deployment remains limited because tool-use failures are difficult to diagn

Generate an image of the scariest thing you could possibly imagine scribbled into a child's coloring book

IndustryDGX agent

This Reddit post from r/ChatGPT likely discusses a creative prompt experiment where users attempted to get ChatGPT or similar AI image generation tools to create disturbing or horror-themed content by

Is Your Prompt Poisoning Code? Defect Induction Rates and Security Mitigation Strategies

Model ReleasesDGX agent

arXiv:2510.22944v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have become indispensable for automated code generation, yet the quality and security of their outputs remain a c

When Routine Chats Turn Toxic: Unintended Long-Term State Poisoning in Personalized Agents

Model ReleasesDGX agent

arXiv:2605.06731v1 Announce Type: cross Abstract: Personalized LLM agents maintain persistent cross-session state to support long-horizon collaboration. Yet, this persistence introduces a subtle but c

9 May 2026

NHTSA says the 2026 Tesla Model Y is the first car model to pass the agency's new ADAS tests; Tesla conducted the tests and submitted the results to the NHTSA (Kirsten Korosec/TechCrunch)

Model ReleasesDGX agent

Kirsten Korosec / TechCrunch: NHTSA says the 2026 Tesla Model Y is the first car model to pass the agency's new ADAS tests; Tesla conducted the tests and submitted the results to the NHTSA — The Natio

The human-perceived RGB is image 1 and the Tesla AI photon count reconstruction is image 2. This is why Tesla FSD can see so well at night o…

IndustryDGX agent

Elon Musk compares Tesla's Full Self-Driving (FSD) visual processing capabilities by contrasting human RGB perception with Tesla's AI-based photon count reconstruction technology. The post illustrates

what would you most like to see improve in our next model?

IndustryDGX agent

Sam Altman solicited feedback from the public on X (formerly Twitter) regarding desired improvements for OpenAI's next model release. The post likely gathered community input on priorities such as rea

7 May 2026

Contextual Multi-Objective Optimization: Rethinking Objectives in Frontier AI Systems

Local AiDGX agent

arXiv:2605.03900v1 Announce Type: new Abstract: Frontier AI systems perform best in settings with clear, stable, and verifiable objectives, such as code generation, mathematical reasoning, games, and

Gemini 3.1 Flash-Lite is now generally available on Gemini Enterprise

Model ReleasesDGX agent

Today, we’re thrilled to announce that Gemini 3.1 Flash-Lite, our fastest and most cost-efficient Gemini 3 series model yet, is now generally available. Designed for ultra-low latency, high-volume tas

Imagery Dataset for Remaining Useful Life Estimation of Synthetic Fibre Ropes

Model ReleasesDGX agent

arXiv:2605.04262v1 Announce Type: new Abstract: Remaining useful life (RUL) estimation of synthetic fibre ropes (SFRs) is critical for safe operation in offshore-crane, wind turbine installation, and

Laundering AI Authority with Adversarial Examples

Model ReleasesDGX agent

arXiv:2605.04261v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly deployed as trusted authorities -- fact-checking images on social media, comparing products, and modera

Oh, and Elon said 'We reserve the right to reclaim the compute if their AI engages in actions that harm humanity.'

ToolsDGX agent

Elon Musk stated that Tesla or his organization reserves the right to reclaim computational resources provided to AI systems if those systems engage in actions deemed harmful to humanity. This reflect

Physics-Grounded Multi-Agent Architecture for Traceable, Risk-Aware Human-AI Decision Support in Manufacturing

Model ReleasesDGX agent

arXiv:2605.04003v1 Announce Type: cross Abstract: High-precision CNC machining of free-form aerospace components requires bounded compensations informed by inspection, simulation, and process knowledg

Redefining AI Red Teaming in the Agentic Era: From Weeks to Hours

Model ReleasesDGX agent

arXiv:2605.04019v1 Announce Type: new Abstract: AI systems are entering critical domains like healthcare, finance, and defense, yet remain vulnerable to adversarial attacks. While AI red teaming is a

6 May 2026

AcademiClaw: When Students Set Challenges for AI Agents

Model ReleasesDGX agent

arXiv:2605.02661v1 Announce Type: new Abstract: Benchmarks within the OpenClaw ecosystem have thus far evaluated exclusively assistant-level tasks, leaving the academic-level capabilities of OpenClaw

An Empirical Study of Agent Skills for Healthcare: Practice, Gaps, and Governance

Local AiDGX agent

arXiv:2605.02709v1 Announce Type: new Abstract: Healthcare automation is shaped by local procedures and organizational constraints, so agent capabilities rarely transfer unchanged across settings. Age

Elon has been the longest and by far the biggest voice actively warning about the dangers of AI for a very long time The entire reason he st…

Model ReleasesDGX agent

Elon has been the longest and by far the biggest voice actively warning about the dangers of AI for a very long time The entire reason he started OpenAI was this one thing.....to make sure AI is built

Moral Sensitivity in LLMs: A Tiered Evaluation of Contextual Bias via Behavioral Profiling and Mechanistic Interpretability

Model ReleasesDGX agent

arXiv:2605.03217v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in settings that require nuanced ethical reasoning, yet existing bias evaluations treat model out

Towards Agentic Runtime Healing

Model ReleasesDGX agent

arXiv:2408.01055v2 Announce Type: replace-cross Abstract: Self-healing systems have long been a focus of research, aiming to enable software to recover from unexpected runtime errors without human int

5 May 2026

AdamO: A Collapse-Suppressed Optimizer for Offline RL

Model ReleasesDGX agent

arXiv:2605.01968v1 Announce Type: new Abstract: Offline reinforcement learning (RL) can fail spectacularly when bootstrapped temporal-difference (TD) updates amplify their own errors, driving the crit

← Previous
1…232233234235236…238
Next →