AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,234 results
24 Apr 2026

Serialisation Strategy Matters: How FHIR Data Format Affects LLM Medication Reconciliation

Model ReleasesDGX agent

arXiv:2604.21076v1 Announce Type: cross Abstract: Medication reconciliation at clinical handoffs is a high-stakes, error-prone process. Large language models are increasingly proposed to assist with t

Tesla FSD is the first AI saving lives at scale....on real roads every single day Road accidents kill 1.19 million people every year - the #…

Model ReleasesDGX agent

Tesla FSD is the first AI saving lives at scale....on real roads every single day Road accidents kill 1.19 million people every year - the #1 cause of death for ages 5–29 94%+ of crashes are caused by

We’ve invested deeply in security at Replit, including our recent launches with Security Agent + Auto-Protect. If you want to move your app …

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases
DGX agent

We’ve invested deeply in security at Replit, including our recent launches with Security Agent + Auto-Protect. If you want to move your app to Replit, we’re offering free app imports for a limited tim

23 Apr 2026

A pelican for GPT-5.5 via the semi-official Codex backdoor API

Model ReleasesDGX agent

GPT-5.5 is out. It's available in OpenAI Codex and is rolling out to paid ChatGPT subscribers. I've had some preview access and found it to be a fast, effective and highly capable model. As is usually

A Vision-Language-Action Model for Adaptive Ultrasound-Guided Needle Insertion and Needle Tracking

Model ReleasesDGX agent

arXiv:2604.20347v1 Announce Type: cross Abstract: Ultrasound (US)-guided needle insertion is a critical yet challenging procedure due to dynamic imaging conditions and difficulties in needle visualiza

Foundation Models in Biomedical Imaging: Turning Hype into Reality

Model ReleasesDGX agent

arXiv:2512.15808v2 Announce Type: replace-cross Abstract: Foundation models (FMs) are driving a prominent shift in biomedical imaging from task-specific models to unified backbone models for diverse t

From Scene to Object: Text-Guided Dual-Gaze Prediction

Model ReleasesDGX agent

arXiv:2604.20191v1 Announce Type: cross Abstract: Interpretable driver attention prediction is crucial for human-like autonomous driving. However, existing datasets provide only scene-level global gaz

Large Language Models Outperform Humans in Fraud Detection and Resistance to Motivated Investor Pressure

Model ReleasesDGX agent

arXiv:2604.20652v1 Announce Type: new Abstract: Large language models trained on human feedback may suppress fraud warnings when investors arrive already persuaded of a fraudulent opportunity. We test

Last night was the biggest disaster in the history of Tesla. Let me walk you through what actually happened on that earnings call, because t…

Model ReleasesDGX agent

Last night was the biggest disaster in the history of Tesla. Let me walk you through what actually happened on that earnings call, because the headlines are doing you a disservice: Elon Musk got on th

Open-Architecture End-to-End System for Real-World Autonomous Robot Navigation

Local AiDGX agent

arXiv:2410.06239v3 Announce Type: replace Abstract: Enabling robots to autonomously navigate unknown, complex, and dynamic real-world environments presents several challenges, including imperfect perc

OVPD: A Virtual-Physical Fusion Testing Dataset of OnSite Auton-omous Driving Challenge

Model ReleasesDGX agent

arXiv:2604.20423v1 Announce Type: new Abstract: The rapid iteration of autonomous driving algorithms has created a growing demand for high-fidelity, replayable, and diagnosable testing data. However,

QuadPiPS: A Perception-informed Footstep Planner for Quadrupeds With Semantic Affordance Prediction

Local AiDGX agent

arXiv:2501.00112v2 Announce Type: replace Abstract: This work proposes QuadPiPS, a perception-informed framework for quadrupedal foothold planning in the perception space. QuadPiPS employs a novel ego

SkillLearnBench: Benchmarking Continual Learning Methods for Agent Skill Generation on Real-World Tasks

Model ReleasesDGX agent

arXiv:2604.20087v1 Announce Type: new Abstract: Skills have become the de facto way to enable LLM agents to perform complex real-world tasks with customized instructions, workflows, and tools, but how

WildFireVQA: A Large-Scale Radiometric Thermal VQA Benchmark for Aerial Wildfire Monitoring

Model ReleasesDGX agent

arXiv:2604.20190v1 Announce Type: new Abstract: Wildfire monitoring requires timely, actionable situational awareness from airborne platforms, yet existing aerial visual question answering (VQA) bench

22 Apr 2026

3D Foundation Model for Generalizable Disease Detection in Head Computed Tomography

Model ReleasesDGX agent

arXiv:2502.02779v3 Announce Type: replace-cross Abstract: Head computed tomography (CT) imaging is a widely-used imaging modality with multitudes of medical indications, particularly in assessing path

Announcing Spanner Omni: Your infrastructure, Google’s innovation

Model ReleasesDGX agent

Today, we announced the preview of Spanner Omni, a downloadable version of Spanner, that expands its industry-leading distributed database capabilities beyond Google Cloud. This enables enterprises to

Are Large Language Models Economically Viable for Industry Deployment?

Model ReleasesDGX agent

arXiv:2604.19342v1 Announce Type: new Abstract: Generative AI-powered by Large Language Models (LLMs)-is increasingly deployed in industry across healthcare decision support, financial analytics, ente

Beyond Explicit Refusals: Soft-Failure Attacks on Retrieval-Augmented Generation

Model ReleasesDGX agent

arXiv:2604.18663v1 Announce Type: cross Abstract: Existing jamming attacks on Retrieval-Augmented Generation (RAG) systems typically induce explicit refusals or denial-of-service behaviors, which are

Cross-cloud infrastructure innovation for the agentic enterprise

Model ReleasesDGX agent

The era of agentic AI is accelerating from human- to machine-speed operations, while also creating profound stress on legacy technology infrastructure. This new reality pushes foundational systems to

Distillation Traps and Guards: A Calibration Knob for LLM Distillability

Local AiDGX agent

arXiv:2604.18963v1 Announce Type: cross Abstract: Knowledge distillation (KD) transfers capabilities from large language models (LLMs) to smaller students, yet it can fail unpredictably and also under

GenerativeMPC: VLM-RAG-guided Whole-Body MPC with Virtual Impedance for Bimanual Mobile Manipulation

Model ReleasesDGX agent

arXiv:2604.19522v1 Announce Type: new Abstract: Bimanual mobile manipulation requires a seamless integration between high-level semantic reasoning and safe, compliant physical interaction - a challeng

Harmful Intent as a Geometrically Recoverable Feature of LLM Residual Streams

Model ReleasesDGX agent

arXiv:2604.18901v1 Announce Type: cross Abstract: Harmful intent is geometrically recoverable from large language model residual streams: as a linear direction in most layers, and as angular deviation

Human-Guided Harm Recovery for Computer Use Agents

Model ReleasesDGX agent

arXiv:2604.18847v1 Announce Type: new Abstract: As LM agents gain the ability to execute actions on real computer systems, we need ways to not only prevent harmful actions at scale but also effectivel

Introducing Gemini Enterprise Agent Platform, powering the next wave of agents

Model ReleasesDGX agent

In the early days of generative AI, building safe and reliable business tools took massive engineering effort and a high tolerance for trial and error. We helped solve that with Vertex AI, our trusted

SpikeMLLM: Spike-based Multimodal Large Language Models via Modality-Specific Temporal Scales and Temporal Compression

HardwareDGX agent

arXiv:2604.18610v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable progress but incur substantial computational overhead and energy consumption during

The new Gemini Enterprise: one platform for agent development, orchestration, and governance

Model ReleasesDGX agent

The first wave of AI changed how we find information; the next wave is changing how we get work done. Today, we’re enhancing our most powerful AI tools and bringing them together under one roof. Gemin

Towards Optimal Agentic Architectures for Offensive Security Tasks

Model ReleasesDGX agent

arXiv:2604.18718v1 Announce Type: cross Abstract: Agentic security systems increasingly audit live targets with tool-using LLMs, but prior systems fix a single coordination topology, leaving unclear w

Uncertainty Quantification in Detection Transformers: Object-Level Calibration and Image-Level Reliability

Local AiDGX agent

arXiv:2412.01782v4 Announce Type: replace-cross Abstract: DETR and its variants have emerged as promising architectures for object detection, offering an end-to-end prediction pipeline. In practice, h

Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMs

Model ReleasesDGX agent

arXiv:2508.00161v3 Announce Type: replace-cross Abstract: The releases of powerful open-weight large language models (LLMs) are often not accompanied by access to their full training data. Existing in

21 Apr 2026

AIM 2025 Rip Current Segmentation (RipSeg) Challenge Report

Model ReleasesDGX agent

arXiv:2508.13401v3 Announce Type: replace Abstract: This report presents an overview of the AIM 2025 RipSeg Challenge, a competition designed to advance techniques for automatic rip current segmentati

Causally-Constrained Probabilistic Forecasting for Time-Series Anomaly Detection

Local AiDGX agent

arXiv:2604.17998v1 Announce Type: new Abstract: Anomaly detection in multivariate time series is a central challenge in industrial monitoring, as failures frequently arise from complex temporal dynami

Cognitive Chain-of-Thought (CoCoT): Structured Multimodal Reasoning about Social Situations

Model ReleasesDGX agent

arXiv:2507.20409v2 Announce Type: replace Abstract: Chain-of-Thought (CoT) prompting helps models think step by step. But naive CoT breaks down in visually grounded social tasks, where models must per

From Heads to Neurons: Causal Attribution and Steering in Multi-Task Vision-Language Models

Model ReleasesDGX agent

arXiv:2604.17941v1 Announce Type: cross Abstract: Recent work has increasingly explored neuron-level interpretation in vision-language models (VLMs) to identify neurons critical to final predictions.

Fuzzy Encoding-Decoding to Improve Spiking Q-Learning Performance in Autonomous Driving

Model ReleasesDGX agent

arXiv:2604.16436v1 Announce Type: cross Abstract: This paper develops an end-to-end fuzzy encoder-decoder architecture for enhancing vision-based multi-modal deep spiking Q-networks in autonomous driv

Injecting Structured Biomedical Knowledge into Language Models: Continual Pretraining vs. GraphRAG

Model ReleasesDGX agent

arXiv:2604.16422v1 Announce Type: new Abstract: The injection of domain-specific knowledge is crucial for adapting language models (LMs) to specialized fields such as biomedicine. While most current a

MedRedFlag: Investigating how LLMs Redirect Misconceptions in Real-World Health Communication

Model ReleasesDGX agent

arXiv:2601.09853v2 Announce Type: replace Abstract: Real-world health questions from patients often unintentionally embed false assumptions or premises. In such cases, safe medical communication typic

NTIRE 2026 Rip Current Detection and Segmentation (RipDetSeg) Challenge Report

Model ReleasesDGX agent

arXiv:2604.17070v1 Announce Type: new Abstract: This report presents the NTIRE 2026 Rip Current Detection and Segmentation (RipDetSeg) Challenge, which targets automatic rip current understanding in i

On Inverse Problems, Parameter Estimation, and Domain Generalization

Model ReleasesDGX agent

arXiv:2506.06024v2 Announce Type: replace-cross Abstract: Signal restoration and inverse problems are key elements in most real-world data science applications. In the past decades, with the emergence

RoTRAG: Rule of Thumb Reasoning for Conversation Harm Detection with Retrieval-Augmented Generation

Model ReleasesDGX agent

arXiv:2604.17301v1 Announce Type: new Abstract: Detecting harmful content in multi turn dialogue requires reasoning over the full conversational context rather than isolated utterances. However, most

Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification

Model ReleasesDGX agent

arXiv:2604.16993v1 Announce Type: cross Abstract: As embodied AI transitions to real-world deployment, the success of the Vision-and-Language Navigation (VLN) task tends to evolve from mere reachabili

TowerDataset: A Heterogeneous Benchmark for Transmission Corridor Segmentation with a Global-Local Fusion Framework

Model ReleasesDGX agent

arXiv:2604.16848v1 Announce Type: new Abstract: Fine-grained semantic segmentation of transmission-corridor point clouds is fundamental for intelligent power-line inspection. However, current progress

20 Apr 2026

HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents?

Model ReleasesDGX agent

arXiv:2604.15415v1 Announce Type: cross Abstract: Large language models (LLMs) have evolved into autonomous agents that rely on open skill ecosystems (e.g., ClawHub and Skills.Rest), hosting numerous

Here's how F1 is tweaking its hybrid systems to try to save the show

IndustryDGX agent

F1's 2026 regulations feature new V6 hybrid engines with a near 50/50 split between combustion and electrical power, creating challenges with energy management and an unpopular driving style among dri

HiPreNets: High-Precision Neural Networks through Progressive Training

Model ReleasesDGX agent

arXiv:2506.15064v3 Announce Type: replace Abstract: Deep neural networks are powerful tools for solving nonlinear problems in science and engineering, but training highly accurate models becomes chall

LLMs Corrupt Your Documents When You Delegate

Model ReleasesDGX agent

arXiv:2604.15597v1 Announce Type: new Abstract: Large Language Models (LLMs) are poised to disrupt knowledge work, with the emergence of delegated work as a new interaction paradigm (e.g., vibe coding

MTR-DuplexBench: Towards a Comprehensive Evaluation of Multi-Round Conversations for Full-Duplex Speech Language Models

Model ReleasesDGX agent

arXiv:2511.10262v3 Announce Type: replace-cross Abstract: Full-Duplex Speech Language Models (FD-SLMs) enable real-time, overlapping conversational interactions, offering a more dynamic user experienc

TwoHamsters: Benchmarking Multi-Concept Compositional Unsafety in Text-to-Image Models

Model ReleasesDGX agent

arXiv:2604.15967v1 Announce Type: cross Abstract: Despite the remarkable synthesis capabilities of text-to-image (T2I) models, safeguarding them against content violations remains a persistent challen

19 Apr 2026

Only Grok 4.3 lets me drive my car to get gas. ChatGPT, Claude, and Gemini want me to walk.

Model ReleasesDGX agent

This post compares AI assistants' responses to a request about driving to get gas, claiming that Grok 4.3 provides the requested information while ChatGPT, Claude, and Gemini decline or suggest altern

Since Anthropic publish their system prompts we can generate a diff between Claude Opus 4.6 and 4.7 - here are my notes on what's changed ht…

Model ReleasesDGX agent

Simon Willison documents the differences between Anthropic's Claude Opus 4.6 and 4.7 system prompts, analyzing changes that Anthropic made public. The notes likely highlight modifications to model beh

17 Apr 2026

ECM Contracts: Contract-Aware, Versioned, and Governable Capability Interfaces for Embodied Agents

Model ReleasesDGX agent

arXiv:2604.13097v1 Announce Type: cross Abstract: Embodied agents increasingly rely on modular capabilities that can be installed, upgraded, composed, and governed at runtime. Prior work has introduce

Mechanistic Decoding of Cognitive Constructs in LLMs

Model ReleasesDGX agent

arXiv:2604.14593v1 Announce Type: new Abstract: While Large Language Models (LLMs) demonstrate increasingly sophisticated affective capabilities, the internal mechanisms by which they process complex

Rethinking AI Hardware: A Three-Layer Cognitive Architecture for Autonomous Agents

Local AiDGX agent

arXiv:2604.13757v1 Announce Type: new Abstract: The next generation of autonomous AI systems will be constrained not only by model capability, but by how intelligence is structured across heterogeneou

Still refuses to write sestinas for some reason, so I don't think all the rough edges are gone.

ApplicationsDGX agent

Ethan Mollick observes that an AI system (likely Claude or another large language model) still declines to write sestinas, suggesting that certain behavioral constraints or limitations persist despite

16 Apr 2026

A Proactive EMR Assistant for Doctor-Patient Dialogue: Streaming ASR, Belief Stabilization, and Preliminary Controlled Evaluation

Model ReleasesDGX agent

arXiv:2604.13059v1 Announce Type: new Abstract: Most dialogue-based electronic medical record (EMR) systems still behave as passive pipelines: transcribe speech, extract information, and generate the

Built an political benchmark for LLMs. KIMI K2 can't answer about Taiwan (Obviously). GPT-5.3 refuses 100% of questions when given an opt-out. [P]

Model ReleasesDGX agent

A researcher on r/MachineLearning built a political benchmark to evaluate how various LLMs handle sensitive geopolitical and politically contentious questions. Key findings include that Kimi K2 (Moons

Document-tuning for robust alignment to animals

Model ReleasesDGX agent

arXiv:2604.13076v1 Announce Type: new Abstract: We investigate the robustness of value alignment via finetuning with synthetic documents, using animal compassion as a value that is both important in i

FieldWorkArena: Agentic AI Benchmark for Real Field Work Tasks

Model ReleasesDGX agent

arXiv:2505.19662v3 Announce Type: replace-cross Abstract: This paper introduces FieldWorkArena, a benchmark for agentic AI targeting real-world field work. With the recent increase in demand for agent

Gaslight, Gatekeep, V1-V3: Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation

Model ReleasesDGX agent

arXiv:2604.13803v1 Announce Type: new Abstract: Vision-language models are increasingly deployed in high-stakes settings, yet their susceptibility to sycophantic manipulation remains poorly understood

How WPP accelerates humanoid robot training 10x with G4 VMs

Model ReleasesDGX agent

Editor’s note: Today we hear from Perry Nightingale, SVP of Creative AI at WPP about the workflow that cuts training time for humanoid robots from days to minutes — plus access to the open-source code

Mosaic: An Extensible Framework for Composing Rule-Based and Learned Motion Planners

Model ReleasesDGX agent

arXiv:2604.13853v1 Announce Type: new Abstract: Safe and explainable motion planning remains a central challenge in autonomous driving. While rule-based planners offer predictable and explainable beha

← Previous
1…234235236237238
Next →