AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlog
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
59,904 results
Model Releases

World-Ego Modeling for Long-Horizon Evolution in Hybrid Embodied Tasks

DGX agent

arXiv:2605.19957v1 Announce Type: cross Abstract: World models are widely explored in embodied intelligence, yet they typically predict distinct evolutions of the world and the ego within a single str

model-releasesarxiv-cs-ai
20 May 2026
Agents
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

A Mechanistic Model for Collective Motion from Sensorimotor Regularities

DGX agent

arXiv:2605.16522v1 Announce Type: new Abstract: Collective behavior in animals has long been modeled through self-propelled particle models, which reproduce striking group-level phenomena through abst

agentsarxiv-cs-ro
19 May 2026
Model Releases

DevBench: A Realistic, Developer-Informed Benchmark for Code Generation Models

DGX agent

arXiv:2601.11895v3 Announce Type: replace-cross Abstract: DevBench is a telemetry-driven benchmark designed to evaluate Large Language Models (LLMs) on realistic code completion tasks. It includes 1,8

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

EPIC-Bench: A Perception-Centric Benchmark for Fine-Grained Embodied Visual Grounding in Vision-Language Models

DGX agent

arXiv:2605.17070v1 Announce Type: new Abstract: While large vision-language models (VLMs) are increasingly adopted as the perceptual backbone for embodied agents, existing benchmarks often rely on que

model-releasesarxiv-cs-cv
19 May 2026
Model Releases

GIM: Evaluating models via tasks that integrate multiple cognitive domains

DGX agent

arXiv:2605.18663v1 Announce Type: new Abstract: As LLM benchmarks saturate, the evaluation community has pursued two strategies to increase difficulty: escalating knowledge demands (GPQA, HLE) or remo

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

I remember when people were saying 'It's useless to open-source big models because nobody will be able to run them fast'....

DGX agent

I remember when people were saying 'It's useless to open-source big models because nobody will be able to run them fast'.... Cerebras is now running Kimi K2.6 – a trillion parameter model – in enterpr

model-releasesclem-delangue--x
19 May 2026
Model Releases

Machine Unlearning for Masked Diffusion Language Models

DGX agent

arXiv:2605.18253v1 Announce Type: cross Abstract: Recent masked diffusion language models (MDLMs), such as LLaDA and Dream, have achieved performance comparable to autoregressive large language models

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Prompt Compression in Diffusion Large Language Models: Evaluating LLMLingua-2 on LLaDA

DGX agent

arXiv:2605.17932v1 Announce Type: cross Abstract: Prompt compression reduces inference cost and context length in large language models, but prior evaluations focus primarily on autoregressive archite

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

ProxyKV: Cross-Model Proxy Pruning for Efficient Long-Context LLM Inference

DGX agent

arXiv:2605.16360v1 Announce Type: cross Abstract: Efficient long-context inference in Large Language Models (LLMs) is severely constrained by the Key-Value (KV) cache memory wall, yet existing pruning

model-releasesarxiv-cs-ai
19 May 2026
Safety

Real-Time Aligned Reward Model beyond Semantics

DGX agent

arXiv:2601.22664v4 Announce Type: replace Abstract: Reinforcement Learning from Human Feedback (RLHF) is a pivotal technique for aligning large language models (LLMs) with human preferences, yet it is

safetyarxiv-cs-ai
19 May 2026
Safety

Right Predictions, Misleading Explanations: On the Vulnerability of Vision-Language Model Explanations

DGX agent

arXiv:2605.16651v1 Announce Type: new Abstract: Explanation mechanisms are increasingly used to support transparency and trust in vision-language models (VLMs), particularly in settings where model de

safetyarxiv-cs-cv
19 May 2026
Research

SAFE-SVD: Sensitivity-Aware Fidelity-Enforcing SVD for Physics Foundation Models

DGX agent

arXiv:2605.17985v1 Announce Type: cross Abstract: We propose a new method for compressing physics foundation models (PFMs) which is a new trend in AI for Science. While model compression is essential

researcharxiv-cs-ai
19 May 2026
Model Releases

Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models

DGX agent

arXiv:2605.18160v1 Announce Type: cross Abstract: In recent years, multimodal large language models (MLLMs) have achieved remarkable progress, primarily attributed to effective paradigms for integrati

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Adapting Foundation Vision-Language Models to Medical Diagnosis via Query-Driven Expert Bridging

DGX agent

arXiv:2505.21698v3 Announce Type: replace Abstract: Vision-language foundation models achieve promising performance in natural image classification, yet their direct application to medical imaging is

model-releasesarxiv-cs-cv
18 May 2026
Model Releases

Golden Layers and Where to Find Them: Improved Knowledge Editing for Large Language Models Via Layer Gradient Analysis

DGX agent

arXiv:2602.20207v3 Announce Type: replace-cross Abstract: Knowledge editing in Large Language Models (LLMs) aims to update the model's prediction for a specific query to a desired target while preserv

model-releasesarxiv-cs-ai
18 May 2026
Research

Latent Video Prediction Learns Better World Models

DGX agent

arXiv:2605.15618v1 Announce Type: cross Abstract: Self-supervised video models are increasingly framed as world models, yet their evaluation remains largely confined to a single top-1 accuracy score o

researcharxiv-cs-ai
18 May 2026
Model Releases

VideoGameBench: Can Vision-Language Models complete popular video games?

DGX agent

arXiv:2505.18134v3 Announce Type: replace Abstract: Vision-language models (VLMs) have achieved strong results on coding and math benchmarks that are challenging for humans, yet their ability to perfo

model-releasesarxiv-cs-ai
18 May 2026
Model Releases

Latest open artifacts (#21): Open model bonanza! Gemma 4, DeepSeek V4, Kimi K2.6, MiMo 2.5, GLM-5.1 & others. On CAISI's V4 assessment.

DGX agent

This article covers recent releases of open-source AI models including Gemma 4, DeepSeek V4, Kimi K2.6, MiMo 2.5, and GLM-5.1, along with discussion of CAISI's V4 model assessment framework. The piece

model-releasesinterconnects
16 May 2026
Model Releases

Mini-JEPA Foundation Model Fleet Enables Agentic Hydrologic Intelligence

DGX agent

arXiv:2605.14120v1 Announce Type: cross Abstract: Geospatial foundation models compress multispectral observations into dense embeddings increasingly used in natural-language environmental reasoning s

model-releasesarxiv-cs-cl
15 May 2026
Agents

Model-Adaptive Tool Necessity Reveals the Knowing-Doing Gap in LLM Tool Use

DGX agent

arXiv:2605.14038v1 Announce Type: new Abstract: Large language models (LLMs) increasingly act as autonomous agents that must decide when to answer directly vs. when to invoke external tools. Prior wor

agentsarxiv-cs-ai
15 May 2026
Model Releases

NeuroAtlas: Benchmarking Foundation Models for Clinical EEG and Brain-Computer Interfaces

DGX agent

arXiv:2605.14698v1 Announce Type: cross Abstract: Foundation models (FMs) promise to extract unified representations that generalize across downstream tasks. They have emerged across fields, including

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

OPT-Engine: Benchmarking the Limits of LLMs in Optimization Modeling via Complexity Scaling

DGX agent

arXiv:2601.19924v2 Announce Type: replace-cross Abstract: We investigate the capabilities and scalability of Large Language Models (LLMs) in optimization modeling, a domain requiring structured reason

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

Small Language Models (SLMs) Can Still Pack a Punch: A survey (updated 2026)

DGX agent

arXiv:2501.05465v2 Announce Type: replace Abstract: As foundation AI models continue to increase in size, an important question arises - is massive scale the only path forward? This survey of about 16

model-releasesarxiv-cs-cl
15 May 2026
Model Releases

Unsteady Metrics and Benchmarking Cultures of AI Model Builders

DGX agent

arXiv:2605.14164v1 Announce Type: new Abstract: The primary way to establish and compare competencies in foundation and generative AI models has shifted from peer-reviewed literature to press releases

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

Another banger of a model, free for Hermes agent users via Nous Portal!

DGX agent

Nous Research announced the release of a new model available for free to Hermes agent users through the Nous Portal. The post suggests this is a significant model release from Nous Research, their org

model-releasesnous-research--x
14 May 2026
Model Releases

CoT-Guard: Small Models for Strong Monitoring

DGX agent

arXiv:2605.12746v1 Announce Type: cross Abstract: Monitoring the chain-of-thought (CoT) of reasoning models is a promising approach for detecting covert misbehavior (i.e., hidden objectives) in code g

model-releasesarxiv-cs-ai
14 May 2026
Model Releases

D-VLA: A High-Concurrency Distributed Asynchronous Reinforcement Learning Framework for Vision-Language-Action Models

DGX agent

arXiv:2605.13276v1 Announce Type: new Abstract: The rapid evolution of Embodied AI has enabled Vision-Language-Action (VLA) models to excel in multimodal perception and task execution. However, applyi

model-releasesarxiv-cs-ai
14 May 2026
Model Releases

Embodied Multi-Agent Coordination by Aligning World Models Through Dialogue

DGX agent

arXiv:2605.12920v1 Announce Type: cross Abstract: Effective collaboration between embodied agents requires more than acting in a shared environment; it demands communication grounded in each agent's e

model-releasesarxiv-cs-ai
14 May 2026
Local Ai

PROMETHEUS: Automating Deep Causal Research Integrating Text, Data and Models

DGX agent

arXiv:2605.12835v1 Announce Type: new Abstract: Large language models can extract local causal claims from text, but those claims become more useful when organized as persistent, navigable world model

local-aiarxiv-cs-ai
14 May 2026
Model Releases

Query-Conditioned Test-Time Self-Training for Large Language Models

DGX agent

arXiv:2605.13369v1 Announce Type: cross Abstract: Large language models (LLMs) are typically deployed with fixed parameters, and their performance is often improved by allocating more computation at i

model-releasesarxiv-cs-ai
14 May 2026
Model Releases

Safe Bayesian Optimization for Uncertain Correlations Matrices in Linear Models of Co-Regionalization

DGX agent

arXiv:2605.13302v1 Announce Type: new Abstract: This paper extends safety guarantees for multi-task Bayesian optimization with uncertain correlation matrices from intrinsic co-reginalization models to

model-releasesarxiv-cs-lg
14 May 2026
Model Releases

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs

DGX agent

arXiv:2510.18245v3 Announce Type: replace-cross Abstract: Scaling the number of parameters and the size of training data has proven to be an effective strategy for improving large language model (LLM)

model-releasesarxiv-cs-ai
14 May 2026
Model Releases

Visual Aesthetic Benchmark: Can Frontier Models Judge Beauty?

DGX agent

arXiv:2605.12684v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) are now routinely deployed for visual understanding, generation, and curation. A substantial fraction of thes

model-releasesarxiv-cs-ai
14 May 2026
Applications

Coevolutionary Continuous Discrete Diffusion: Make Your Diffusion Language Model a Latent Reasoner

DGX agent

arXiv:2510.03206v2 Announce Type: replace-cross Abstract: Diffusion language models, especially masked discrete diffusion models, have achieved great success recently. While there are some theoretical

applicationsarxiv-cs-cl
13 May 2026
Safety

Enabling clinical use of foundation models for computational pathology

DGX agent

arXiv:2602.22347v2 Announce Type: replace Abstract: Foundation models for computational pathology are expected to facilitate the development of high-performing, generalisable deep learning systems. Ho

safetyarxiv-cs-cv
13 May 2026
Model Releases

More Edits, More Stable: Understanding the Lifelong Normalization in Sequential Model Editing

DGX agent

arXiv:2605.11836v1 Announce Type: cross Abstract: Lifelong Model Editing aims to continuously update evolving facts in Large Language Models while preserving unrelated knowledge and general capabiliti

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

Can We Go Beyond Visual Features? Neural Tissue Relation Modeling for Relational Graph Analysis in Non-Melanoma Skin Histology

DGX agent

arXiv:2512.06949v3 Announce Type: replace Abstract: Histopathology image segmentation is essential for delineating tissue structures in skin cancer diagnostics, but modeling spatial context and inter-

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

ColorConceptBench: A Benchmark for Probabilistic Color-Concept Understanding in Text-to-Image Models

DGX agent

arXiv:2601.16836v3 Announce Type: replace-cross Abstract: Text-to-image (T2I) models have advanced considerably in generating high-quality images from textual descriptions. However, their ability to a

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Decomposing and Steering Functional Metacognition in Large Language Models

DGX agent

arXiv:2605.08942v1 Announce Type: new Abstract: Large language models (LLMs) increasingly exhibit behaviors suggesting awareness of their evaluation context, often adapting their reasoning strategies

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

DeformMaster: An Interactive Physics-Neural World Model for Deformable Objects from Videos

DGX agent

arXiv:2605.09586v1 Announce Type: new Abstract: World models for deformable objects should recover not only geometry and appearance, but also underlying physical dynamics, interaction grounding, and m

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Diffusion Models are Evolutionary Algorithms

DGX agent

arXiv:2410.02543v3 Announce Type: replace-cross Abstract: In a convergence of machine learning and biology, we reveal that diffusion models are evolutionary algorithms. By considering evolution as a d

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

jNO: A JAX Library for Neural Operator and Foundation Model Training

DGX agent

arXiv:2605.10159v1 Announce Type: new Abstract: jNO (jax Neural Operators) is a JAX-native library for neural operators and foundation models with unified support for both data-driven and physics-info

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Large Language Models as Students Who Think Aloud: Overly Coherent, Verbose, and Confident

DGX agent

arXiv:2602.01015v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly embedded in AI-based tutoring systems. Can they faithfully model novice reasoning and metacognitive ju

model-releasesarxiv-cs-cl
12 May 2026
Research

PARD-2: Target-Aligned Parallel Draft Model for Dual-Mode Speculative Decoding

DGX agent

arXiv:2605.08632v1 Announce Type: cross Abstract: Speculative decoding accelerates Large Language Models (LLMs) inference by using a lightweight draft model to propose candidate tokens that are verifi

researcharxiv-cs-ai
12 May 2026
Model Releases

Reinforcement Learning Measurement Model

DGX agent

arXiv:2605.09305v1 Announce Type: cross Abstract: Interactive assessments generate sequential process data that are not well handled by conventional item response models. Existing MDP-based measuremen

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Selective Deficits in LLM Mental Self-Modeling in a Behavior-Based Test of Theory of Mind

DGX agent

arXiv:2603.26089v2 Announce Type: replace-cross Abstract: The ability to represent oneself and others as agents with knowledge, intentions, and belief states that guide their behavior - Theory of Mind

model-releasesarxiv-cs-ai
12 May 2026
Research

Sparse Layers are Critical to Scaling Looped Language Models

DGX agent

arXiv:2605.09165v1 Announce Type: cross Abstract: Looped language models repeat a set of transformer layers through depth, reducing memory costs and providing natural early-exit points at loop boundar

researcharxiv-cs-cl
12 May 2026
Model Releases

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data

DGX agent

arXiv:2605.10129v1 Announce Type: new Abstract: Large language models (LLMs) rely on web-scale corpora for pre-training. The noise inherent in these datasets tends to obscure meaningful patterns and u

model-releasesarxiv-cs-cl
12 May 2026
← Previous
1…5051525354…1248
Next →