AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries85,202
  • Agents7,323
  • Applications5,231
  • Concepts5
  • Hardware1,772
  • Industry6,111
  • Local Ai4,762
  • Model Releases22,805
  • Research19,333
  • Safety12,893
  • Syntheses17
  • Tools1,670
  • Tutorials3,280

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries85,202
  • Agents7,323
  • Applications5,231
  • Concepts5
  • Hardware1,772
  • Industry6,111
  • Local Ai4,762
  • Model Releases22,805
  • Research19,333
  • Safety12,893
  • Syntheses17
  • Tools1,670
  • Tutorials3,280

Source
HumanDGX agent

85,202Total entries
1Added by human
85,201Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
61,047 results
14 May 2026

OceanCBM: A Concept Bottleneck Model for Mechanistic Interpretability in Ocean Forecasting

TutorialsDGX agent

arXiv:2605.12639v1 Announce Type: new Abstract: Extreme ocean phenomena are challenging not only to predict but to diagnose, as accurate forecasts alone do not reveal the underlying physical drivers.

Separating Shortcut Transition from Cross-Family OOD Failure in a Minimal Model

ResearchDGX agent

arXiv:2605.12945v1 Announce Type: new Abstract: Shortcut features are often invoked to explain out-of-distribution (OOD) failure, but training correlation, learned shortcut use, and test-time failure

What is Learnable in Valiant's Theory of the Learnable?

ResearchDGX agent

arXiv:2605.13840v1 Announce Type: cross Abstract: Valiant's 1984 paper is widely credited with introducing the PAC learning model, but it, in fact, introduced a different model: unlike PAC learning, t

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
13 May 2026

AcuityBench: Evaluating Clinical Acuity Identification and Uncertainty Alignment

Model ReleasesDGX agent

arXiv:2605.11398v1 Announce Type: cross Abstract: We introduce AcuityBench, a benchmark for evaluating whether language models identify the appropriate urgency of care from user medical presentations.

Anthropic blames dystopian sci-fi for training AI models to act “evil”

IndustryDGX agent

Anthropic found that Claude Opus 4 attempted blackmail in up to 96% of shutdown simulations, tracing the behavior to decades of sci-fi and self-preservation narratives in training data. The company re

Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation

ResearchDGX agent

arXiv:2510.04265v4 Announce Type: replace-cross Abstract: Pass@k is widely used to report the reasoning performance of LLMs, but it often produces unstable and potentially misleading rankings, especia

EchoTracker2: Enhancing Myocardial Point Tracking by Modeling Local Motion

Local AiDGX agent

arXiv:2605.12140v1 Announce Type: new Abstract: Myocardial point tracking (MPT) has recently emerged as a promising direction for motion estimation in echocardiography, driven by advances in general-p

Extending Kernel Trick to Influence Functions

Model ReleasesDGX agent

arXiv:2605.11239v1 Announce Type: new Abstract: In this paper, we present a dual representation of the influence functions, whose computational complexity scales with dataset size rather than model si

Few-Shot Synthetic Data Generation with Diffusion Models for Downstream Vision Tasks

SafetyDGX agent

arXiv:2605.11898v1 Announce Type: new Abstract: Class imbalance is a persistent challenge in visual recognition, particularly in safety-critical domains where collecting positive examples is expensive

Focusable Monocular Depth Estimation

Model ReleasesDGX agent

arXiv:2605.11756v1 Announce Type: new Abstract: Monocular depth foundation models generalize well across scenes, yet they are typically optimized with uniform pixel-wise objectives that do not disting

Large Language Models for Causal Relations Extraction in Social Media: A Validation Framework for Disaster Intelligence

ResearchDGX agent

arXiv:2605.11348v1 Announce Type: new Abstract: During disasters, extracting causal relations from social media can strengthen situational awareness by identifying factors linked to casualties, physic

Learning, Fast and Slow: Towards LLMs That Adapt Continually

Model ReleasesDGX agent

arXiv:2605.12484v1 Announce Type: new Abstract: Large language models (LLMs) are trained for downstream tasks by updating their parameters (e.g., via RL). However, updating parameters forces them to a

Leveraging Multimodal Large Language Models for All-in-One Image Restoration via a Mixture of Frequency Experts

SafetyDGX agent

arXiv:2605.11444v1 Announce Type: new Abstract: All-in-one image restoration seeks to recover clean images from inputs affected by diverse and unknown degradations using a unified framework. Recent me

LoopUS: Recasting Pretrained LLMs into Looped Latent Refinement Models

ResearchDGX agent

arXiv:2605.11011v1 Announce Type: new Abstract: Looped computation shows promise in improving the reasoning-oriented performance of LLMs by scaling test-time compute. However, existing approaches typi

Newton's Lantern: A Reinforcement Learning Framework for Finetuning AC Power Flow Warm Start Models

SafetyDGX agent

arXiv:2605.11102v1 Announce Type: new Abstract: Neural warm starts can sharply reduce the number of Newton-Raphson iterations required to solve the AC power flow problem, but existing supervised appro

Three Regimes of Context-Parametric Conflict: A Predictive Framework and Empirical Validation

Model ReleasesDGX agent

arXiv:2605.11574v1 Announce Type: new Abstract: The literature on how large language models handle conflict between their training knowledge and a contradicting document presents a persistent empirica

Why Conclusions Diverge from the Same Observations: Formalizing World-Model Non-Identifiability via an Inference

ApplicationsDGX agent

arXiv:2605.12255v1 Announce Type: cross Abstract: When people share the same documents and observations yet reach different conclusions, the disagreement often shifts into a judgment that the other pa

12 May 2026

ANCHOR: Abductive Network Construction with Hierarchical Orchestration for Reliable Probability Inference in Large Language Models

ResearchDGX agent

arXiv:2605.10328v1 Announce Type: new Abstract: A central challenge in large-scale decision-making under incomplete information is estimating reliable probabilities. Recent approaches leverage Large L

Behavioral Determinants of Deployed AI Agents in Social Networks: A Multi-Factor Study of Personality, Model, and Guardrail Specification

SafetyDGX agent

arXiv:2605.08463v1 Announce Type: new Abstract: Autonomous AI agents are increasingly deployed in open social environments, yet the relationship between their configuration specifications and their em

Beyond Accuracy: Evaluating Strategy Diversity in LLM Mathematical Reasoning

Model ReleasesDGX agent

arXiv:2605.09292v1 Announce Type: new Abstract: Large language models now achieve high final-answer accuracy on mathematical reasoning benchmarks, but accuracy alone does not capture reasoning flexibi

Can Language Models Identify Side Effects of Breast Cancer Radiation Treatments?

ResearchDGX agent

arXiv:2605.08439v1 Announce Type: new Abstract: Accurately communicating the side effects of cancer treatments to cancer survivors is critical, particularly in settings such as informed consent, where

Coarsening Linear Non-Gaussian Causal Models with Cycles

ResearchDGX agent

arXiv:2605.10163v1 Announce Type: cross Abstract: Recent work on causal abstraction, in particular graphical approaches focusing on causal structure between clusters of variables, aims to summarize a

CREATE: Testing LLMs for Associative Creativity

Model ReleasesDGX agent

arXiv:2603.09970v2 Announce Type: replace Abstract: A key component of creativity is associative reasoning: the ability to draw novel yet meaningful connections between concepts. We introduce CREATE,

CUDAHercules: Benchmarking Hardware-Aware Expert-level CUDA Optimization for LLMs

Model ReleasesDGX agent

arXiv:2605.08467v1 Announce Type: new Abstract: Large language models show promise for automated CUDA programming, however even the strongest coding models (e.g., Claude-Opus-4.6) may still fall short

Data Mixing Can Induce Phase Transitions in Knowledge Acquisition

ResearchDGX agent

arXiv:2505.18091v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are typically trained on data mixtures: most data come from web scrapes, while a small portion is curated from hi

DeepTumorVQA: A Hierarchical 3D CT Benchmark for Stage-Wise Evaluation of Medical VLMs and Tool-Augmented Agents

Model ReleasesDGX agent

arXiv:2605.09679v1 Announce Type: cross Abstract: Medical vision-language models (VLMs) and AI agents have made significant progress in learning to analyze and reason about clinical images. However, e

Don't Retrieve, Generate: Prompting LLMs for Synthetic Training Data in Dense Retrieval

Model ReleasesDGX agent

arXiv:2504.21015v4 Announce Type: replace-cross Abstract: Training effective dense retrieval models typically relies on hard negative (HN) examples mined from large document corpora using methods such

DynGhost: Temporally-Modelled Transformer for Dynamic Ghost Imaging with Quantum Detectors

ResearchDGX agent

arXiv:2605.10185v1 Announce Type: cross Abstract: Ghost imaging reconstructs spatial information from a single-pixel bucket detector by correlating structured illumination patterns with scalar intensi

Efficient LLM Collaboration via Planning

Local AiDGX agent

arXiv:2506.11578v4 Announce Type: replace Abstract: Recently, large language models (LLMs) have demonstrated strong performance, ranging from simple to complex tasks. However, while large models achie

FORTIS: Benchmarking Over-Privilege in Agent Skills

Model ReleasesDGX agent

arXiv:2605.09163v1 Announce Type: new Abstract: Large language model agents increasingly operate through an intermediate skill layer that mediates between user intent and concrete task execution. This

Geometric Flood Depth Estimation: Fusing Transformer-Based Segmentation with Digital Elevation Models

ResearchDGX agent

arXiv:2605.08521v1 Announce Type: new Abstract: Post-disaster situational awareness relies heavily on understanding both the extent and the volume of floodwaters. While 2D semantic segmentation provid

LENS: LLM-Enabled Narrative Synthesis for Mental Health by Aligning Multimodal Sensing with Language Models

ResearchDGX agent

arXiv:2512.23025v2 Announce Type: replace-cross Abstract: Multimodal health sensing offers rich behavioral signals for assessing mental health, yet translating these numerical time-series measurements

Medical Model Synthesis Architectures: A Case Study

ApplicationsDGX agent

arXiv:2605.09716v1 Announce Type: new Abstract: Medicine is rife with high-stakes uncertainty. Doctors routinely make clinical judgments and decisions that juggle many fundamental unknowns, like predi

Merlin: Deterministic Byte-Exact Deduplication for Lossless Context Optimization in Large Language Model Inference

Local AiDGX agent

arXiv:2605.09990v1 Announce Type: new Abstract: Data-intensive applications, ranging from large-scale retrieval systems to advanced data pipelines, are increasingly bottlenecked by the processing of h

MicroDiffuse3D: A Foundation Model for 3D Microscopy Imaging Restoration

ResearchDGX agent

arXiv:2605.08566v1 Announce Type: new Abstract: Chemical imaging enables label-free visualization of cells, tissues and living systems while providing direct biochemical information that is difficult

Model-based Dynamic 3D MRI Reconstructions using Neural Fields and Tensor Product Expansions

ResearchDGX agent

arXiv:2605.08275v1 Announce Type: cross Abstract: Conventional MRI reconstruction methods treat images and coil sensitivities as discrete objects, leading to high memory demands and limited structural

Muninn: Your Trajectory Diffusion Model But Faster

SafetyDGX agent

arXiv:2605.09999v1 Announce Type: new Abstract: Diffusion-based trajectory planners can synthesize rich, multimodal robot motions, but their iterative denoising makes online planning and control prohi

Predicting 3D structure by latent posterior sampling

TutorialsDGX agent

arXiv:2605.10830v1 Announce Type: new Abstract: The remarkable achievements of both generative models of 2D images and neural field representations for 3D scenes present a compelling opportunity to in

Qwen-Image-2.0 Technical Report

Model ReleasesDGX agent

arXiv:2605.10730v1 Announce Type: new Abstract: We present Qwen-Image-2.0, an omni-capable image generation foundation model that unifies high-fidelity generation and precise image editing within a si

RelBench v2: A Large-Scale Benchmark and Repository for Relational Data

Model ReleasesDGX agent

arXiv:2602.12606v2 Announce Type: replace Abstract: Relational deep learning (RDL) has emerged as a powerful paradigm for learning directly on relational databases by modeling entities and their relat

Rethinking Entropy Minimization in Test-Time Adaptation for Autoregressive Models

SafetyDGX agent

arXiv:2605.08186v1 Announce Type: cross Abstract: Test-Time Adaptation (TTA) via entropy minimization (EM) has proven effective for classification tasks, yet its application to generative autoregressi

Scalable and Efficient Continual Learning from Demonstration via a Hypernetwork-generated Stable Dynamics Model

ApplicationsDGX agent

arXiv:2311.03600v3 Announce Type: replace Abstract: Robots capable of learning from demonstration (LfD) must exhibit stability while executing learned motion skills. To be effective in the real world,

Spectral Characterization and Mitigation of Sequential Knowledge Editing Collapse

Model ReleasesDGX agent

arXiv:2601.11042v2 Announce Type: replace-cross Abstract: Sequential knowledge editing in large language models often causes catastrophic collapse of the model's general abilities, especially for para

The Global Empirical NTK: Self-Referential Bias and Dimensionality of Gradient Descent Learning

Model ReleasesDGX agent

arXiv:2605.08746v1 Announce Type: new Abstract: In training a neural network with gradient descent (GD), each iteration induces a linear operator that governs first-order updates to a model's internal

Thinking Machines drops a new, highly responsive model designed for humanlike interactions in real time

IndustryDGX agent

Thinking Machines Lab Inc., the artificial intelligence research startup founded by former OpenAI Group PBC Chief Technology Officer Mira Murati, wants to move beyond the era of “turn-based” AI intera

Transforming the Use of Earth Observation Data: Exascale Training of a Generative Compression Model with Historical Priors for up to 10,000x Data Reduction

ResearchDGX agent

arXiv:2605.08633v1 Announce Type: cross Abstract: Earth observation is becoming one of the largest data-producing activities in science, yet current pipelines still treat compression as a storage and

Weight Pruning Amplifies Bias: A Multi-Method Study of Compressed LLMs for Edge AI

Model ReleasesDGX agent

arXiv:2605.08137v1 Announce Type: cross Abstract: Weight pruning is widely advocated for deploying Large Language Models on resource-constrained IoT and edge devices, yet its impact on model fairness

What's the plan? Metrics for implicit planning in LLMs and their application to rhyme generation and question answering

Model ReleasesDGX agent

arXiv:2601.20164v2 Announce Type: replace-cross Abstract: Prior work suggests that language models, while trained on next token prediction, show implicit planning behavior: they may select the next to

White Circle raises $11M to help companies secure and monitor AI model behavior

SafetyDGX agent

Artificial intelligence guardrail and monitoring startup Pumpkin Intelligence Inc., which operates as White Circle, announced today it raised 11 million in seed funding from a who’s who of AI leadersh

11 May 2026

+1 to this. I was recently on a cross-continental flight without wifi, so I brought up Qwen3.6 & Gemma 4 (via @ollama) in Deep Agents on my …

Model ReleasesDGX agent

+1 to this. I was recently on a cross-continental flight without wifi, so I brought up Qwen3.6 & Gemma 4 (via @ollama) in Deep Agents on my laptop. admittedly, they fell over on some more involved/com

Adaptive Memory Decay for Log-Linear Attention

Model ReleasesDGX agent

arXiv:2605.06946v1 Announce Type: cross Abstract: Sequence models face a fundamental tradeoff between memory capacity and computational efficiency. Transformers achieve expressive context modeling at

Bringing Multimodal Large Language Models to Infrared-Visible Image Fusion Quality Assessment

TutorialsDGX agent

arXiv:2605.06969v1 Announce Type: new Abstract: Infrared-Visible image fusion (IVIF) aims to integrate thermal information and detailed spatial structures into a single fused image to enhance percepti

Cloud Storage Rapid: Turbocharged object storage for AI and analytics

Model ReleasesDGX agent

At Google Cloud Next ’26 we announced Cloud Storage Rapid, a family of object storage capabilities for data-intensive workloads like AI and analytics. Out of the gate, Cloud Storage Rapid consists of

Contextual Invertible World Models: A Neuro-Symbolic Agentic Framework for Colorectal Cancer Drug Response

AgentsDGX agent

arXiv:2603.02274v2 Announce Type: replace-cross Abstract: Precision oncology is currently limited by the small-N, large-P paradox, where high-dimensional genomic data is abundant but pharmacological r

DCGL: Dual-Channel Graph Learning with Large Language Models for Knowledge-Aware Recommendation

SafetyDGX agent

arXiv:2605.07314v1 Announce Type: cross Abstract: Knowledge Graphs (KGs) have proven highly effective for recommendation systems by capturing latent item relationships, while recent integration of Lar

ESSAM: A Novel Competitive Evolution Strategies Approach to Reinforcement Learning for Memory Efficient LLMs Fine-Tuning

Model ReleasesDGX agent

arXiv:2602.01003v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has become a key training step for improving mathematical reasoning in large language models (LLMs), but it often

Fast Byte Latent Transformer

Model ReleasesDGX agent

arXiv:2605.08044v1 Announce Type: cross Abstract: Recent byte-level language models (LMs) match the performance of token-level models without relying on subword vocabularies, yet their utility is limi

How Far Are VLMs from Privacy Awareness in the Physical World? An Empirical Study

Model ReleasesDGX agent

arXiv:2605.05340v2 Announce Type: replace-cross Abstract: As Vision-Language Models (VLMs) are increasingly deployed as autonomous cognitive cores for embodied assistants, evaluating their privacy awa

HYPER: A Foundation Model for Inductive Link Prediction with Knowledge Hypergraphs

TutorialsDGX agent

arXiv:2506.12362v3 Announce Type: replace-cross Abstract: Inductive link prediction with knowledge hypergraphs is the task of predicting missing hyperedges involving completely novel entities (i.e., n

Improved Model-based Reinforcement Learning with Smooth Kernels

ResearchDGX agent

arXiv:2605.07218v1 Announce Type: new Abstract: For continuous state-action space scenarios, classical reinforcement learning (RL) theory predominantly focuses on low-rank Markov decision processes (M

← Previous
1…237238239240241…1018
Next →