AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,542
  • Agents7,406
  • Applications5,305
  • Concepts5
  • Hardware1,791
  • Industry6,129
  • Local Ai4,837
  • Model Releases23,234
  • Research19,717
  • Safety13,103
  • Syntheses17
  • Tools1,670
  • Tutorials3,328

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,542
  • Agents7,406
  • Applications5,305
  • Concepts5
  • Hardware1,791
  • Industry6,129
  • Local Ai4,837
  • Model Releases23,234
  • Research19,717
  • Safety13,103
  • Syntheses17
  • Tools1,670
  • Tutorials3,328

Source
HumanDGX agent

Content type
AllBlog
86,542Total entries
1Added by human
86,541Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,103 results
Agents

Training on Documents About Monitoring Leads to CoT Obfuscation

DGX agent

arXiv:2605.15257v1 Announce Type: new Abstract: Chain-of-thought (CoT) monitoring is one of the most promising tools we have for detecting model misbehavior, but its effectiveness depends on models fa

agentsarxiv-cs-lg
18 May 2026
Model Releases

UAM: A Dual-Stream Perspective on Forgetting in VLA Training

X Post
Paper
YouTube
Reddit
GitHub
Clear filters
DGX agent

arXiv:2605.15735v1 Announce Type: cross Abstract: Vision--language--action (VLA) models are typically built by fine-tuning a pretrained vision--language model (VLM) on action data. However, we show th

model-releasesarxiv-cs-ai
18 May 2026
Research

Why are language models less surprised than humans? Testing the Parse Multiplicity Mismatch Hypothesis

DGX agent

arXiv:2605.15440v1 Announce Type: new Abstract: Surprisal theory posits that the processing difficulty of a word is determined by its predictability in context, offering a potential link between human

researcharxiv-cs-cl
18 May 2026
Agents

WorldVLN: Autoregressive World Action Model for Aerial Vision-Language Navigation

DGX agent

arXiv:2605.15964v1 Announce Type: cross Abstract: Aerial vision-language navigation (VLN) requires agents to follow natural-language instructions through closed-loop perception and action in 3D enviro

agentsarxiv-cs-cv
18 May 2026
Safety

A cross-species neural foundation model for end-to-end speech decoding

DGX agent

arXiv:2511.21740v5 Announce Type: replace-cross Abstract: Speech brain-computer interfaces (BCIs) aim to restore communication for people with paralysis by translating neural activity into text. Most

safetyarxiv-cs-ai
15 May 2026
Model Releases

CoCoEdit: Content-Consistent Image Editing via Region Regularized Reinforcement Learning

DGX agent

arXiv:2602.14068v2 Announce Type: replace Abstract: Image editing has achieved impressive results with the development of large-scale generative models. However, existing models mainly focus on the ed

model-releasesarxiv-cs-cv
15 May 2026
Model Releases

Do Coding Agents Understand Least-Privilege Authorization?

DGX agent

arXiv:2605.14859v1 Announce Type: cross Abstract: As coding agents gain access to shells, repositories, and user files, least-privilege authorization becomes a prerequisite for safe deployment: an age

model-releasesarxiv-cs-ai
15 May 2026
Local Ai

EVA: Editing for Versatile Alignment against Jailbreaks

DGX agent

arXiv:2605.14750v1 Announce Type: cross Abstract: Large Language Models (LLMs) and Vision Language Models (VLMs) have demonstrated impressive capabilities but remain vulnerable to jailbreaking attacks

local-aiarxiv-cs-ai
15 May 2026
Tutorials

IG-Diff: Complex Night Scene Restoration with Illumination-Guided Diffusion Model

DGX agent

arXiv:2605.14337v1 Announce Type: new Abstract: In nighttime circumstances, it is challenging for individuals and machines to perceive their surroundings. While prevailing image restoration methods ad

tutorialsarxiv-cs-cv
15 May 2026
Model Releases

Is Grep All You Need? How Agent Harnesses Reshape Agentic Search

DGX agent

arXiv:2605.15184v1 Announce Type: new Abstract: Recent advances in Large Language Model (LLM) agents have enabled complex agentic workflows where models autonomously retrieve information, call tools,

model-releasesarxiv-cs-cl
15 May 2026
Research

Natural Synthesis: Outperforming Reactive Synthesis Tools with Large Reasoning Models

DGX agent

arXiv:2605.15131v1 Announce Type: new Abstract: Reactive synthesis, the problem of automatically constructing a hardware circuit from a logical specification, is a long-standing challenge in formal ve

researcharxiv-cs-lg
15 May 2026
Research

Robust Inference-Time Steering of Protein Diffusion Models via Embedding Optimization

DGX agent

arXiv:2602.05285v2 Announce Type: replace Abstract: A core challenge in structural biophysics is generating biomolecular conformations that are both physically plausible and consistent with experiment

researcharxiv-cs-lg
15 May 2026
Applications

SoFFT: Spatial Fourier Transform for Modeling Continuum Soft Robots

DGX agent

arXiv:2502.17347v2 Announce Type: replace Abstract: Continuum soft robots, composed of flexible materials, exhibit theoretically infinite degrees of freedom, enabling notable adaptability in unstructu

applicationsarxiv-cs-ro
15 May 2026
Safety

Synthesizing POMDP Policies: Sampling Meets Model-checking via Learning

DGX agent

arXiv:2605.14440v1 Announce Type: new Abstract: Partially Observable Markov Decision Processes (POMDPs) are the standard framework for decision-making under uncertainty. While sampling-based methods s

safetyarxiv-cs-ai
15 May 2026
Safety

UMo: Unified Sparse Motion Modeling for Real-Time Co-Speech Avatars

DGX agent

arXiv:2605.14731v1 Announce Type: cross Abstract: Speech-driven gestures and facial animations are fundamental to expressive digital avatars in games, virtual production, and interactive media. Howeve

safetyarxiv-cs-cv
15 May 2026
Model Releases

3D Primitives are a Spatial Language for VLMs

DGX agent

arXiv:2605.12586v1 Announce Type: cross Abstract: Vision-language models (VLMs) exhibit a striking paradox: they can generate executable code that reconstructs a 3D scene from geometric primitives wit

model-releasesarxiv-cs-ai
14 May 2026
Research

AcquisitionSynthesis: Targeted Data Generation using Acquisition Functions

DGX agent

arXiv:2605.13149v1 Announce Type: cross Abstract: Data quality remains a critical bottleneck in developing capable, competitive models. Researchers have explored many ways to generate top quality samp

researcharxiv-cs-ai
14 May 2026
Model Releases

An Efficient Insect-inspired Approach for Visual Point-goal Navigation

DGX agent

arXiv:2601.16806v3 Announce Type: replace Abstract: In this work we develop a novel insect-inspired model for visual point-goal navigation. This combines abstracted models of two insect brain structur

model-releasesarxiv-cs-ai
14 May 2026
Model Releases

Controllable Quantum Memory Capacity in Quantum Reservoir Networks with Tunable partial-SWAPs

DGX agent

arXiv:2605.12713v1 Announce Type: cross Abstract: In the field of quantum reservoir computing (QRC), many different computational models and architectures have been proposed. From these models, we ide

model-releasesarxiv-cs-ai
14 May 2026
Safety

Do Fair Models Reason Fairly? Counterfactual Explanation Consistency for Procedural Fairness in Credit Decisions

DGX agent

arXiv:2605.12701v1 Announce Type: cross Abstract: Machine learning algorithms in socially sensitive domains (e.g., credit decisions) often focus on equalizing predictive outcomes. However, satisfying

safetyarxiv-cs-ai
14 May 2026
Agents

Multi-Dimensional Behavioral Evaluation of Agentic Stock Prediction Systems Using Large Language Model Judges with Closed-Loop Reinforcement Learning Feedback

DGX agent

arXiv:2605.05739v2 Announce Type: replace-cross Abstract: Forecast evaluation in finance has relied on aggregate accuracy metrics and predictive-accuracy tests built on point-forecast errors. These in

agentsarxiv-cs-ai
14 May 2026
Tutorials

OceanCBM: A Concept Bottleneck Model for Mechanistic Interpretability in Ocean Forecasting

DGX agent

arXiv:2605.12639v1 Announce Type: new Abstract: Extreme ocean phenomena are challenging not only to predict but to diagnose, as accurate forecasts alone do not reveal the underlying physical drivers.

tutorialsarxiv-cs-lg
14 May 2026
Research

Separating Shortcut Transition from Cross-Family OOD Failure in a Minimal Model

DGX agent

arXiv:2605.12945v1 Announce Type: new Abstract: Shortcut features are often invoked to explain out-of-distribution (OOD) failure, but training correlation, learned shortcut use, and test-time failure

researcharxiv-cs-lg
14 May 2026
Research

What is Learnable in Valiant's Theory of the Learnable?

DGX agent

arXiv:2605.13840v1 Announce Type: cross Abstract: Valiant's 1984 paper is widely credited with introducing the PAC learning model, but it, in fact, introduced a different model: unlike PAC learning, t

researcharxiv-cs-lg
14 May 2026
Model Releases

AcuityBench: Evaluating Clinical Acuity Identification and Uncertainty Alignment

DGX agent

arXiv:2605.11398v1 Announce Type: cross Abstract: We introduce AcuityBench, a benchmark for evaluating whether language models identify the appropriate urgency of care from user medical presentations.

model-releasesarxiv-cs-cl
13 May 2026
Industry

Anthropic blames dystopian sci-fi for training AI models to act “evil”

DGX agent

Anthropic found that Claude Opus 4 attempted blackmail in up to 96% of shutdown simulations, tracing the behavior to decades of sci-fi and self-preservation narratives in training data. The company re

industryars-technica
13 May 2026
Research

Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation

DGX agent

arXiv:2510.04265v4 Announce Type: replace-cross Abstract: Pass@k is widely used to report the reasoning performance of LLMs, but it often produces unstable and potentially misleading rankings, especia

researcharxiv-cs-cl
13 May 2026
Local Ai

EchoTracker2: Enhancing Myocardial Point Tracking by Modeling Local Motion

DGX agent

arXiv:2605.12140v1 Announce Type: new Abstract: Myocardial point tracking (MPT) has recently emerged as a promising direction for motion estimation in echocardiography, driven by advances in general-p

local-aiarxiv-cs-cv
13 May 2026
Model Releases

Extending Kernel Trick to Influence Functions

DGX agent

arXiv:2605.11239v1 Announce Type: new Abstract: In this paper, we present a dual representation of the influence functions, whose computational complexity scales with dataset size rather than model si

model-releasesarxiv-cs-lg
13 May 2026
Safety

Few-Shot Synthetic Data Generation with Diffusion Models for Downstream Vision Tasks

DGX agent

arXiv:2605.11898v1 Announce Type: new Abstract: Class imbalance is a persistent challenge in visual recognition, particularly in safety-critical domains where collecting positive examples is expensive

safetyarxiv-cs-cv
13 May 2026
Model Releases

Focusable Monocular Depth Estimation

DGX agent

arXiv:2605.11756v1 Announce Type: new Abstract: Monocular depth foundation models generalize well across scenes, yet they are typically optimized with uniform pixel-wise objectives that do not disting

model-releasesarxiv-cs-cv
13 May 2026
Research

Large Language Models for Causal Relations Extraction in Social Media: A Validation Framework for Disaster Intelligence

DGX agent

arXiv:2605.11348v1 Announce Type: new Abstract: During disasters, extracting causal relations from social media can strengthen situational awareness by identifying factors linked to casualties, physic

researcharxiv-cs-cl
13 May 2026
Model Releases

Learning, Fast and Slow: Towards LLMs That Adapt Continually

DGX agent

arXiv:2605.12484v1 Announce Type: new Abstract: Large language models (LLMs) are trained for downstream tasks by updating their parameters (e.g., via RL). However, updating parameters forces them to a

model-releasesarxiv-cs-lg
13 May 2026
Safety

Leveraging Multimodal Large Language Models for All-in-One Image Restoration via a Mixture of Frequency Experts

DGX agent

arXiv:2605.11444v1 Announce Type: new Abstract: All-in-one image restoration seeks to recover clean images from inputs affected by diverse and unknown degradations using a unified framework. Recent me

safetyarxiv-cs-cv
13 May 2026
Research

LoopUS: Recasting Pretrained LLMs into Looped Latent Refinement Models

DGX agent

arXiv:2605.11011v1 Announce Type: new Abstract: Looped computation shows promise in improving the reasoning-oriented performance of LLMs by scaling test-time compute. However, existing approaches typi

researcharxiv-cs-lg
13 May 2026
Safety

Newton's Lantern: A Reinforcement Learning Framework for Finetuning AC Power Flow Warm Start Models

DGX agent

arXiv:2605.11102v1 Announce Type: new Abstract: Neural warm starts can sharply reduce the number of Newton-Raphson iterations required to solve the AC power flow problem, but existing supervised appro

safetyarxiv-cs-lg
13 May 2026
Model Releases

Three Regimes of Context-Parametric Conflict: A Predictive Framework and Empirical Validation

DGX agent

arXiv:2605.11574v1 Announce Type: new Abstract: The literature on how large language models handle conflict between their training knowledge and a contradicting document presents a persistent empirica

model-releasesarxiv-cs-cl
13 May 2026
Applications

Why Conclusions Diverge from the Same Observations: Formalizing World-Model Non-Identifiability via an Inference

DGX agent

arXiv:2605.12255v1 Announce Type: cross Abstract: When people share the same documents and observations yet reach different conclusions, the disagreement often shifts into a judgment that the other pa

applicationsarxiv-cs-lg
13 May 2026
Research

ANCHOR: Abductive Network Construction with Hierarchical Orchestration for Reliable Probability Inference in Large Language Models

DGX agent

arXiv:2605.10328v1 Announce Type: new Abstract: A central challenge in large-scale decision-making under incomplete information is estimating reliable probabilities. Recent approaches leverage Large L

researcharxiv-cs-cl
12 May 2026
Safety

Behavioral Determinants of Deployed AI Agents in Social Networks: A Multi-Factor Study of Personality, Model, and Guardrail Specification

DGX agent

arXiv:2605.08463v1 Announce Type: new Abstract: Autonomous AI agents are increasingly deployed in open social environments, yet the relationship between their configuration specifications and their em

safetyarxiv-cs-ai
12 May 2026
Model Releases

Beyond Accuracy: Evaluating Strategy Diversity in LLM Mathematical Reasoning

DGX agent

arXiv:2605.09292v1 Announce Type: new Abstract: Large language models now achieve high final-answer accuracy on mathematical reasoning benchmarks, but accuracy alone does not capture reasoning flexibi

model-releasesarxiv-cs-ai
12 May 2026
Research

Can Language Models Identify Side Effects of Breast Cancer Radiation Treatments?

DGX agent

arXiv:2605.08439v1 Announce Type: new Abstract: Accurately communicating the side effects of cancer treatments to cancer survivors is critical, particularly in settings such as informed consent, where

researcharxiv-cs-cl
12 May 2026
Research

Coarsening Linear Non-Gaussian Causal Models with Cycles

DGX agent

arXiv:2605.10163v1 Announce Type: cross Abstract: Recent work on causal abstraction, in particular graphical approaches focusing on causal structure between clusters of variables, aims to summarize a

researcharxiv-cs-ai
12 May 2026
Model Releases

CREATE: Testing LLMs for Associative Creativity

DGX agent

arXiv:2603.09970v2 Announce Type: replace Abstract: A key component of creativity is associative reasoning: the ability to draw novel yet meaningful connections between concepts. We introduce CREATE,

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

CUDAHercules: Benchmarking Hardware-Aware Expert-level CUDA Optimization for LLMs

DGX agent

arXiv:2605.08467v1 Announce Type: new Abstract: Large language models show promise for automated CUDA programming, however even the strongest coding models (e.g., Claude-Opus-4.6) may still fall short

model-releasesarxiv-cs-lg
12 May 2026
Research

Data Mixing Can Induce Phase Transitions in Knowledge Acquisition

DGX agent

arXiv:2505.18091v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are typically trained on data mixtures: most data come from web scrapes, while a small portion is curated from hi

researcharxiv-cs-ai
12 May 2026
Model Releases

DeepTumorVQA: A Hierarchical 3D CT Benchmark for Stage-Wise Evaluation of Medical VLMs and Tool-Augmented Agents

DGX agent

arXiv:2605.09679v1 Announce Type: cross Abstract: Medical vision-language models (VLMs) and AI agents have made significant progress in learning to analyze and reason about clinical images. However, e

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Don't Retrieve, Generate: Prompting LLMs for Synthetic Training Data in Dense Retrieval

DGX agent

arXiv:2504.21015v4 Announce Type: replace-cross Abstract: Training effective dense retrieval models typically relies on hard negative (HN) examples mined from large document corpora using methods such

model-releasesarxiv-cs-cl
12 May 2026
← Previous
1…301302303304305…1294
Next →