AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,234 results
Model Releases

SkillLearnBench: Benchmarking Continual Learning Methods for Agent Skill Generation on Real-World Tasks

DGX agent

arXiv:2604.20087v1 Announce Type: new Abstract: Skills have become the de facto way to enable LLM agents to perform complex real-world tasks with customized instructions, workflows, and tools, but how

model-releasesarxiv-cs-cl
23 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

WildFireVQA: A Large-Scale Radiometric Thermal VQA Benchmark for Aerial Wildfire Monitoring

DGX agent

arXiv:2604.20190v1 Announce Type: new Abstract: Wildfire monitoring requires timely, actionable situational awareness from airborne platforms, yet existing aerial visual question answering (VQA) bench

model-releasesarxiv-cs-cv
23 Apr 2026
Model Releases

3D Foundation Model for Generalizable Disease Detection in Head Computed Tomography

DGX agent

arXiv:2502.02779v3 Announce Type: replace-cross Abstract: Head computed tomography (CT) imaging is a widely-used imaging modality with multitudes of medical indications, particularly in assessing path

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

Announcing Spanner Omni: Your infrastructure, Google’s innovation

DGX agent

Today, we announced the preview of Spanner Omni, a downloadable version of Spanner, that expands its industry-leading distributed database capabilities beyond Google Cloud. This enables enterprises to

model-releasesgoogle-cloud-ai
22 Apr 2026
Model Releases

Are Large Language Models Economically Viable for Industry Deployment?

DGX agent

arXiv:2604.19342v1 Announce Type: new Abstract: Generative AI-powered by Large Language Models (LLMs)-is increasingly deployed in industry across healthcare decision support, financial analytics, ente

model-releasesarxiv-cs-cl
22 Apr 2026
Model Releases

Beyond Explicit Refusals: Soft-Failure Attacks on Retrieval-Augmented Generation

DGX agent

arXiv:2604.18663v1 Announce Type: cross Abstract: Existing jamming attacks on Retrieval-Augmented Generation (RAG) systems typically induce explicit refusals or denial-of-service behaviors, which are

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

Cross-cloud infrastructure innovation for the agentic enterprise

DGX agent

The era of agentic AI is accelerating from human- to machine-speed operations, while also creating profound stress on legacy technology infrastructure. This new reality pushes foundational systems to

model-releasesgoogle-cloud-ai
22 Apr 2026
Local Ai

Distillation Traps and Guards: A Calibration Knob for LLM Distillability

DGX agent

arXiv:2604.18963v1 Announce Type: cross Abstract: Knowledge distillation (KD) transfers capabilities from large language models (LLMs) to smaller students, yet it can fail unpredictably and also under

local-aiarxiv-cs-ai
22 Apr 2026
Model Releases

GenerativeMPC: VLM-RAG-guided Whole-Body MPC with Virtual Impedance for Bimanual Mobile Manipulation

DGX agent

arXiv:2604.19522v1 Announce Type: new Abstract: Bimanual mobile manipulation requires a seamless integration between high-level semantic reasoning and safe, compliant physical interaction - a challeng

model-releasesarxiv-cs-ro
22 Apr 2026
Model Releases

Harmful Intent as a Geometrically Recoverable Feature of LLM Residual Streams

DGX agent

arXiv:2604.18901v1 Announce Type: cross Abstract: Harmful intent is geometrically recoverable from large language model residual streams: as a linear direction in most layers, and as angular deviation

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

Human-Guided Harm Recovery for Computer Use Agents

DGX agent

arXiv:2604.18847v1 Announce Type: new Abstract: As LM agents gain the ability to execute actions on real computer systems, we need ways to not only prevent harmful actions at scale but also effectivel

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

Introducing Gemini Enterprise Agent Platform, powering the next wave of agents

DGX agent

In the early days of generative AI, building safe and reliable business tools took massive engineering effort and a high tolerance for trial and error. We helped solve that with Vertex AI, our trusted

model-releasesgoogle-cloud-ai
22 Apr 2026
Hardware

SpikeMLLM: Spike-based Multimodal Large Language Models via Modality-Specific Temporal Scales and Temporal Compression

DGX agent

arXiv:2604.18610v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable progress but incur substantial computational overhead and energy consumption during

hardwarearxiv-cs-ai
22 Apr 2026
Model Releases

The new Gemini Enterprise: one platform for agent development, orchestration, and governance

DGX agent

The first wave of AI changed how we find information; the next wave is changing how we get work done. Today, we’re enhancing our most powerful AI tools and bringing them together under one roof. Gemin

model-releasesgoogle-cloud-ai
22 Apr 2026
Model Releases

Towards Optimal Agentic Architectures for Offensive Security Tasks

DGX agent

arXiv:2604.18718v1 Announce Type: cross Abstract: Agentic security systems increasingly audit live targets with tool-using LLMs, but prior systems fix a single coordination topology, leaving unclear w

model-releasesarxiv-cs-ai
22 Apr 2026
Local Ai

Uncertainty Quantification in Detection Transformers: Object-Level Calibration and Image-Level Reliability

DGX agent

arXiv:2412.01782v4 Announce Type: replace-cross Abstract: DETR and its variants have emerged as promising architectures for object detection, offering an end-to-end prediction pipeline. In practice, h

local-aiarxiv-cs-ai
22 Apr 2026
Model Releases

Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMs

DGX agent

arXiv:2508.00161v3 Announce Type: replace-cross Abstract: The releases of powerful open-weight large language models (LLMs) are often not accompanied by access to their full training data. Existing in

model-releasesarxiv-cs-cl
22 Apr 2026
Model Releases

AIM 2025 Rip Current Segmentation (RipSeg) Challenge Report

DGX agent

arXiv:2508.13401v3 Announce Type: replace Abstract: This report presents an overview of the AIM 2025 RipSeg Challenge, a competition designed to advance techniques for automatic rip current segmentati

model-releasesarxiv-cs-cv
21 Apr 2026
Local Ai

Causally-Constrained Probabilistic Forecasting for Time-Series Anomaly Detection

DGX agent

arXiv:2604.17998v1 Announce Type: new Abstract: Anomaly detection in multivariate time series is a central challenge in industrial monitoring, as failures frequently arise from complex temporal dynami

local-aiarxiv-cs-lg
21 Apr 2026
Model Releases

Cognitive Chain-of-Thought (CoCoT): Structured Multimodal Reasoning about Social Situations

DGX agent

arXiv:2507.20409v2 Announce Type: replace Abstract: Chain-of-Thought (CoT) prompting helps models think step by step. But naive CoT breaks down in visually grounded social tasks, where models must per

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

From Heads to Neurons: Causal Attribution and Steering in Multi-Task Vision-Language Models

DGX agent

arXiv:2604.17941v1 Announce Type: cross Abstract: Recent work has increasingly explored neuron-level interpretation in vision-language models (VLMs) to identify neurons critical to final predictions.

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Fuzzy Encoding-Decoding to Improve Spiking Q-Learning Performance in Autonomous Driving

DGX agent

arXiv:2604.16436v1 Announce Type: cross Abstract: This paper develops an end-to-end fuzzy encoder-decoder architecture for enhancing vision-based multi-modal deep spiking Q-networks in autonomous driv

model-releasesarxiv-cs-lg
21 Apr 2026
Model Releases

Injecting Structured Biomedical Knowledge into Language Models: Continual Pretraining vs. GraphRAG

DGX agent

arXiv:2604.16422v1 Announce Type: new Abstract: The injection of domain-specific knowledge is crucial for adapting language models (LMs) to specialized fields such as biomedicine. While most current a

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

MedRedFlag: Investigating how LLMs Redirect Misconceptions in Real-World Health Communication

DGX agent

arXiv:2601.09853v2 Announce Type: replace Abstract: Real-world health questions from patients often unintentionally embed false assumptions or premises. In such cases, safe medical communication typic

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

NTIRE 2026 Rip Current Detection and Segmentation (RipDetSeg) Challenge Report

DGX agent

arXiv:2604.17070v1 Announce Type: new Abstract: This report presents the NTIRE 2026 Rip Current Detection and Segmentation (RipDetSeg) Challenge, which targets automatic rip current understanding in i

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

On Inverse Problems, Parameter Estimation, and Domain Generalization

DGX agent

arXiv:2506.06024v2 Announce Type: replace-cross Abstract: Signal restoration and inverse problems are key elements in most real-world data science applications. In the past decades, with the emergence

model-releasesarxiv-cs-lg
21 Apr 2026
Model Releases

RoTRAG: Rule of Thumb Reasoning for Conversation Harm Detection with Retrieval-Augmented Generation

DGX agent

arXiv:2604.17301v1 Announce Type: new Abstract: Detecting harmful content in multi turn dialogue requires reasoning over the full conversational context rather than isolated utterances. However, most

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification

DGX agent

arXiv:2604.16993v1 Announce Type: cross Abstract: As embodied AI transitions to real-world deployment, the success of the Vision-and-Language Navigation (VLN) task tends to evolve from mere reachabili

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

TowerDataset: A Heterogeneous Benchmark for Transmission Corridor Segmentation with a Global-Local Fusion Framework

DGX agent

arXiv:2604.16848v1 Announce Type: new Abstract: Fine-grained semantic segmentation of transmission-corridor point clouds is fundamental for intelligent power-line inspection. However, current progress

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents?

DGX agent

arXiv:2604.15415v1 Announce Type: cross Abstract: Large language models (LLMs) have evolved into autonomous agents that rely on open skill ecosystems (e.g., ClawHub and Skills.Rest), hosting numerous

model-releasesarxiv-cs-ai
20 Apr 2026
Industry

Here's how F1 is tweaking its hybrid systems to try to save the show

DGX agent

F1's 2026 regulations feature new V6 hybrid engines with a near 50/50 split between combustion and electrical power, creating challenges with energy management and an unpopular driving style among dri

industryars-technica
20 Apr 2026
Model Releases

HiPreNets: High-Precision Neural Networks through Progressive Training

DGX agent

arXiv:2506.15064v3 Announce Type: replace Abstract: Deep neural networks are powerful tools for solving nonlinear problems in science and engineering, but training highly accurate models becomes chall

model-releasesarxiv-cs-lg
20 Apr 2026
Model Releases

LLMs Corrupt Your Documents When You Delegate

DGX agent

arXiv:2604.15597v1 Announce Type: new Abstract: Large Language Models (LLMs) are poised to disrupt knowledge work, with the emergence of delegated work as a new interaction paradigm (e.g., vibe coding

model-releasesarxiv-cs-cl
20 Apr 2026
Model Releases

MTR-DuplexBench: Towards a Comprehensive Evaluation of Multi-Round Conversations for Full-Duplex Speech Language Models

DGX agent

arXiv:2511.10262v3 Announce Type: replace-cross Abstract: Full-Duplex Speech Language Models (FD-SLMs) enable real-time, overlapping conversational interactions, offering a more dynamic user experienc

model-releasesarxiv-cs-ai
20 Apr 2026
Model Releases

TwoHamsters: Benchmarking Multi-Concept Compositional Unsafety in Text-to-Image Models

DGX agent

arXiv:2604.15967v1 Announce Type: cross Abstract: Despite the remarkable synthesis capabilities of text-to-image (T2I) models, safeguarding them against content violations remains a persistent challen

model-releasesarxiv-cs-cv
20 Apr 2026
Model Releases

Only Grok 4.3 lets me drive my car to get gas. ChatGPT, Claude, and Gemini want me to walk.

DGX agent

This post compares AI assistants' responses to a request about driving to get gas, claiming that Grok 4.3 provides the requested information while ChatGPT, Claude, and Gemini decline or suggest altern

model-releaseselon-musk--x
19 Apr 2026
Model Releases

Since Anthropic publish their system prompts we can generate a diff between Claude Opus 4.6 and 4.7 - here are my notes on what's changed ht…

DGX agent

Simon Willison documents the differences between Anthropic's Claude Opus 4.6 and 4.7 system prompts, analyzing changes that Anthropic made public. The notes likely highlight modifications to model beh

model-releasessimon-willison--x
19 Apr 2026
Model Releases

ECM Contracts: Contract-Aware, Versioned, and Governable Capability Interfaces for Embodied Agents

DGX agent

arXiv:2604.13097v1 Announce Type: cross Abstract: Embodied agents increasingly rely on modular capabilities that can be installed, upgraded, composed, and governed at runtime. Prior work has introduce

model-releasesarxiv-cs-ai
17 Apr 2026
Model Releases

Mechanistic Decoding of Cognitive Constructs in LLMs

DGX agent

arXiv:2604.14593v1 Announce Type: new Abstract: While Large Language Models (LLMs) demonstrate increasingly sophisticated affective capabilities, the internal mechanisms by which they process complex

model-releasesarxiv-cs-cl
17 Apr 2026
Local Ai

Rethinking AI Hardware: A Three-Layer Cognitive Architecture for Autonomous Agents

DGX agent

arXiv:2604.13757v1 Announce Type: new Abstract: The next generation of autonomous AI systems will be constrained not only by model capability, but by how intelligence is structured across heterogeneou

local-aiarxiv-cs-ai
17 Apr 2026
Applications

Still refuses to write sestinas for some reason, so I don't think all the rough edges are gone.

DGX agent

Ethan Mollick observes that an AI system (likely Claude or another large language model) still declines to write sestinas, suggesting that certain behavioral constraints or limitations persist despite

applicationsethan-mollick--x
17 Apr 2026
Model Releases

A Proactive EMR Assistant for Doctor-Patient Dialogue: Streaming ASR, Belief Stabilization, and Preliminary Controlled Evaluation

DGX agent

arXiv:2604.13059v1 Announce Type: new Abstract: Most dialogue-based electronic medical record (EMR) systems still behave as passive pipelines: transcribe speech, extract information, and generate the

model-releasesarxiv-cs-cl
16 Apr 2026
Model Releases

Built an political benchmark for LLMs. KIMI K2 can't answer about Taiwan (Obviously). GPT-5.3 refuses 100% of questions when given an opt-out. [P]

DGX agent

A researcher on r/MachineLearning built a political benchmark to evaluate how various LLMs handle sensitive geopolitical and politically contentious questions. Key findings include that Kimi K2 (Moons

model-releasesr-machinelearning
16 Apr 2026
Model Releases

Document-tuning for robust alignment to animals

DGX agent

arXiv:2604.13076v1 Announce Type: new Abstract: We investigate the robustness of value alignment via finetuning with synthetic documents, using animal compassion as a value that is both important in i

model-releasesarxiv-cs-cl
16 Apr 2026
Model Releases

FieldWorkArena: Agentic AI Benchmark for Real Field Work Tasks

DGX agent

arXiv:2505.19662v3 Announce Type: replace-cross Abstract: This paper introduces FieldWorkArena, a benchmark for agentic AI targeting real-world field work. With the recent increase in demand for agent

model-releasesarxiv-cs-cv
16 Apr 2026
Model Releases

Gaslight, Gatekeep, V1-V3: Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation

DGX agent

arXiv:2604.13803v1 Announce Type: new Abstract: Vision-language models are increasingly deployed in high-stakes settings, yet their susceptibility to sycophantic manipulation remains poorly understood

model-releasesarxiv-cs-cv
16 Apr 2026
Model Releases

How WPP accelerates humanoid robot training 10x with G4 VMs

DGX agent

Editor’s note: Today we hear from Perry Nightingale, SVP of Creative AI at WPP about the workflow that cuts training time for humanoid robots from days to minutes — plus access to the open-source code

model-releasesgoogle-cloud-ai
16 Apr 2026
Model Releases

Mosaic: An Extensible Framework for Composing Rule-Based and Learned Motion Planners

DGX agent

arXiv:2604.13853v1 Announce Type: new Abstract: Safe and explainable motion planning remains a central challenge in autonomous driving. While rule-based planners offer predictable and explainable beha

model-releasesarxiv-cs-ro
16 Apr 2026
← Previous
1…293294295296297
Next →