AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries85,188
  • Agents7,322
  • Applications5,231
  • Concepts5
  • Hardware1,770
  • Industry6,109
  • Local Ai4,762
  • Model Releases22,797
  • Research19,333
  • Safety12,893
  • Syntheses17
  • Tools1,670
  • Tutorials3,279

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries85,188
  • Agents7,322
  • Applications5,231
  • Concepts5
  • Hardware1,770
  • Industry6,109
  • Local Ai4,762
  • Model Releases22,797
  • Research19,333
  • Safety12,893
  • Syntheses17
  • Tools1,670
  • Tutorials3,279

Source
HumanDGX agent

85,188Total entries
1Added by human
85,187Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,797 results
Model Releases

Hylos: Operability Contracts for Model-Native Spatial Intelligence

DGX agent

arXiv:2605.24728v1 Announce Type: new Abstract: Foundation models can increasingly describe, reconstruct, and generate 3D objects, assemblies, scenes, and environments, but visually plausible spatial

model-releasesarxiv-cs-ai
26 May 2026
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

I don't comment on every article that uses out-of-date measures on AI ability, but I felt (probably wrongly) that the article was a response…

DGX agent

I don't comment on every article that uses out-of-date measures on AI ability, but I felt (probably wrongly) that the article was a response to my viral tweet, so I felt I needed to say something! htt

model-releasesethan-mollick--x
26 May 2026
Model Releases

I found this Wired article on AI fact-checking frustrating. It could have been about why we continue to need human fact checkers (talk to pe…

DGX agent

I found this Wired article on AI fact-checking frustrating. It could have been about why we continue to need human fact checkers (talk to people, use judgement, resolve conflict). Instead it is full o

model-releasesethan-mollick--x
26 May 2026
Model Releases

IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference

DGX agent

arXiv:2605.25475v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly expected to operate over long contexts, yet standard softmax attention incurs a KV cache that grows line

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

INDUCTION: Finite-Structure Concept Synthesis in First-Order Logic

DGX agent

arXiv:2602.18956v3 Announce Type: replace Abstract: We introduce INDUCTION, a benchmark for finite structure concept synthesis in first order logic. Given small finite relational worlds with extension

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Inference Time Optimization with Confidence Dynamics

DGX agent

arXiv:2605.25244v1 Announce Type: new Abstract: Inference time optimization techniques, such as repeated sampling, have significantly advanced the reasoning capabilities of Large Language Models (LLMs

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

InfiFPO: Implicit Model Fusion via Preference Optimization in Large Language Models

DGX agent

arXiv:2505.13878v3 Announce Type: replace-cross Abstract: Model fusion combines multiple Large Language Models (LLMs) with different strengths into a more powerful, integrated model through lightweigh

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

Insuring Every Action: An Authority Frontier Framework for Runtime Actuarial Control of Autonomous AI Agents

DGX agent

arXiv:2605.25632v1 Announce Type: new Abstract: Autonomous AI agents increasingly issue side-effect-bearing actions: database mutations, refunds, payments, external commitments. We propose the Actuari

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Introducing CHI-Bench on @huggingface: the world’s first long-horizon healthcare benchmark for AI agents. 75 real healthcare workflows + 20 …

DGX agent

Introducing CHI-Bench on @huggingface: the world’s first long-horizon healthcare benchmark for AI agents. 75 real healthcare workflows + 20 apps + 200+ MCP tools + 1,290 skills + process / outcome rew

model-releasesclem-delangue--x
26 May 2026
Model Releases

IterInject: Indirect Prompt Injection Against LLM Agents via Feedback-Guided Iterative Optimization

DGX agent

arXiv:2605.24659v1 Announce Type: new Abstract: LLM-based agents are increasingly deployed for complex tasks requiring planning, tool use, and interaction with external services. Their reliance on unt

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

JacQuant: STE-Free Quantization-Aware Training via Learned Jacobian Surrogates

DGX agent

arXiv:2605.25469v1 Announce Type: new Abstract: Quantization-aware training (QAT) is widely deployed but typically relies on the Straight-Through Estimator (STE), which passes gradients through non-di

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments

DGX agent

arXiv:2602.18527v2 Announce Type: replace-cross Abstract: Current audio-visual large language models (AV-LLMs) are predominantly restricted to 2D perception, relying on RGB video and monaural audio. T

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

JEPA-DNA: Grounding Genomic Foundation Models through Joint-Embedding Predictive Architectures

DGX agent

arXiv:2602.17162v2 Announce Type: replace Abstract: Genomic Foundation Models (GFMs) typically rely on Masked Language Modeling (MLM) or Next-Token Prediction (NTP) to learn the 'Laws of Nature'. Whil

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

JudgmentBench: Comparing Rubric and Preference Evaluation for Quality Assessment

DGX agent

arXiv:2605.25240v1 Announce Type: cross Abstract: Two methodologies dominate current practices of benchmarking: rubric-based scoring evaluates items against predefined criteria, whereas comparative ju

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

KAME: Tandem Architecture for Enhancing Knowledge in Real-Time Speech-to-Speech Conversational AI

DGX agent

arXiv:2510.02327v2 Announce Type: replace-cross Abstract: Real-time speech-to-speech (S2S) models excel at generating natural, low-latency conversational responses but often lack deep knowledge and se

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Keep the Proof State Live: Snapshotting for Efficient Tactic Search in Lean 4

DGX agent

arXiv:2605.25556v1 Announce Type: cross Abstract: Automated theorem proving systems built on Lean 4 increasingly rely on parallel tactic search over partially specified proofs, such as those generated

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Kolmogorov-Arnold Fourier Networks

DGX agent

arXiv:2502.06018v3 Announce Type: replace-cross Abstract: Although Kolmogorov-Arnold-based interpretable networks (KANs) possess strong theoretical expressiveness, they suffer from severe parameter ex

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Latent Q-Barrier Shielding for Safe In-Context Reinforcement Learning

DGX agent

arXiv:2605.25267v1 Announce Type: cross Abstract: Safe in-context reinforcement learning (ICRL) adapts online from interaction history without test-time parameter updates while controlling episode cos

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Learning Fine-grained Parameter Sharing via Sparse Tensor Decomposition

DGX agent

arXiv:2411.09816v4 Announce Type: replace Abstract: Large neural networks achieve state-of-the-art performance on many tasks, yet their sheer size hinders deployment on resource-constrained devices. A

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

Learning Sparse Compositional Functions with Norm-Constrained Neural Networks

DGX agent

arXiv:2605.25608v1 Announce Type: cross Abstract: The ability of deep neural networks to learn hierarchical features is widely regarded as a key mechanism underlying their success in high-dimensional

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

Learning to Reason Efficiently with A* Post-Training

DGX agent

arXiv:2605.24597v1 Announce Type: new Abstract: Many applications of large language models (LLMs) require deductive reasoning, yet models frequently produce incorrect or redundant inference steps. We

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization

DGX agent

arXiv:2510.03827v2 Announce Type: replace-cross Abstract: LIBERO has emerged as a widely adopted benchmark for evaluating Vision-Language-Action (VLA) models; however, its current training and evaluat

model-releasesarxiv-cs-ro
26 May 2026
Model Releases

LiveMCP-101: Stress Testing and Diagnosing MCP-enabled Agents on Challenging Queries

DGX agent

arXiv:2508.15760v2 Announce Type: replace-cross Abstract: Tool calling has emerged as a critical capability for AI agents. In contrast to conventional tool calling frameworks that rely on static, prov

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Llamion Technical Report

DGX agent

arXiv:2605.25676v1 Announce Type: new Abstract: We release Llamion, a family of 14B-parameter open-weight language models obtained by transforming Orion-14B into the standardized Llama-family architec

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

LLM-as-a-Reviewer: Benchmarking Their Ability, Divergence, and Prompt Injection Resistance as Paper Reviewers

DGX agent

arXiv:2605.25415v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in academic peer review, yet their reliability, alignment with human judgment, and robustness to adve

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

LLMs on prem. but the prem is your wardrobe. http://shop.cohere.com The design team cooked on this one.

DGX agent

Cohere has announced a creative marketing initiative featuring LLM-themed merchandise available through their shop, with the tagline playing on the phrase 'on premises' to humorously suggest their lan

model-releasescohere--x
26 May 2026
Model Releases

LLMTabBench: Evaluating LLMs on Binary Tabular Classification From Zero to Few Shots

DGX agent

arXiv:2605.24417v1 Announce Type: new Abstract: Supervised classification for tabular data remains a core machine learning task, yet its reliance on large labeled datasets limits applicability in data

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

Looped Diffusion Language Models

DGX agent

arXiv:2605.26106v1 Announce Type: new Abstract: Masked diffusion models (MDMs) have emerged as a promising alternative to autoregressive models for language modeling, yet the effective design of trans

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

LWiAI Podcast #246 - Gemini 3.5 + Omni, Musk Loses, OpenAI vs Erdős

DGX agent

This podcast episode discusses Google's release of Gemini 3.5 and its Omni multimodal capabilities, covers recent developments in Elon Musk's AI ventures, and examines tensions or competition between

model-releaseslast-week-in-ai
26 May 2026
Model Releases

M^3-Verse: A 'Spot the Difference' Challenge for Large Multimodal Models

DGX agent

arXiv:2512.18735v2 Announce Type: replace-cross Abstract: Modern Large Multimodal Models (LMMs) have demonstrated extraordinary ability in static image and single-state spatial-temporal understanding.

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

MDIA: A Multi-Agent Diagnostic Intelligence Pipeline on HealthBench Professional

DGX agent

arXiv:2605.24699v1 Announce Type: new Abstract: Most reported gains on agentic-LLM clinical benchmarks are often attributed to prompt engineering, yet our results suggest that larger improvements can

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework

DGX agent

arXiv:2605.24661v1 Announce Type: new Abstract: LLMs have achieved remarkable success in complex reasoning tasks, yet current evaluation approaches predominantly rely on final-answer correctness, offe

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

MEDAL: Manifold Embedding Distillation via Autoencoder Learning

DGX agent

arXiv:2605.24244v1 Announce Type: cross Abstract: Low-dimensional embeddings are widely used as visual summaries of high-dimensional data and to enable downstream scientific discoveries. Yet, popular

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

Memory-Induced Tool-Drift in LLM Agents

DGX agent

arXiv:2605.24941v1 Announce Type: cross Abstract: Modern LLM agents combine long-term memory for personalization with tool-calling interfaces for taking actions in the world -- a combination underpinn

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

Merch, on-premises. http://shop.cohere.com

DGX agent

Cohere announced merchandise available for purchase at their on-premises shop (shop.cohere.com), likely offering branded items such as apparel, stickers, or other company merchandise for employees, cu

model-releasescohere--x
26 May 2026
Model Releases

Merge-Bench: Resolve Merge Conflicts with Large Language Models

DGX agent

arXiv:2605.25890v1 Announce Type: new Abstract: This paper applies machine learning to the difficult and important task of version control merging. (1) We constructed a dataset, Merge-Bench, of 7938 r

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

Meta-Soft: Leveraging Composable Meta-Tokens for Context-Preserving KV Cache Compression

DGX agent

arXiv:2605.22337v2 Announce Type: replace Abstract: The KV cache used in large language models has linearly growing time complexity, so LLMs face memory blow-up and reduced decoding efficiency when th

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

MimirRAG: A Multi-Agent RAG Framework for Financial Data Retrieval with Metadata Integration

DGX agent

arXiv:2605.25030v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) systems offer a promising approach to reduce hallucinations and improve answer accuracy in large language models (L

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

MIND: Multi-Scale Intent Diffusion for Text-Driven Physics-Based Humanoid Control

DGX agent

arXiv:2605.26006v1 Announce Type: cross Abstract: Enabling physics-based humanoids to execute diverse behaviors from high-level textual commands remains a significant challenge. Existing methods typic

model-releasesarxiv-cs-ro
26 May 2026
Model Releases

MindAlign: Bridging EEG, Vision, and Language for Zero-Shot Visual Decoding

DGX agent

arXiv:2605.24523v1 Announce Type: cross Abstract: Visual decoding from brain signals is a key challenge at the intersection of computer vision and neuroscience, requiring methods that bridge neural re

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

Minimax Limits of k-Fold Cross-Validation via Majority

DGX agent

arXiv:2605.25859v1 Announce Type: cross Abstract: We study the mean-squared error of k-fold cross-validation as a risk estimator, with particular emphasis on how its accuracy depends on the number of

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

Mistral-7B v0.3 at 128K in llama.cpp: 22,657 → 13,235 MiB live VRAM with ≤0.004 PPL drift

DGX agent

Mistral-7B v0.3 model achieves significant memory optimization when running at 128K context length in llama.cpp, reducing live VRAM usage from 22,657 MiB to 13,235 MiB while maintaining minimal perfor

model-releasesr-ollama
26 May 2026
Model Releases

Mix-MoE: Improving Multilingual Machine Translation of Large Language Models through Mixed MoEs

DGX agent

arXiv:2605.24681v1 Announce Type: cross Abstract: Large Language Models (LLMs) have shown great promise in multilingual machine translation (MT), even with limited bilingual supervision. However, fine

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence

DGX agent

arXiv:2505.23764v3 Announce Type: replace-cross Abstract: Spatial intelligence is essential for multimodal large language models (MLLMs) operating in the complex physical world. Existing benchmarks, h

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

Mosaic: Compositional Multi-Concept Erasure via Vector Field Blending

DGX agent

arXiv:2605.25574v1 Announce Type: cross Abstract: Concept erasure has emerged as a key research direction for ensuring safe and ethical image synthesis in Text-to-Image (T2I) models. While existing st

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

MR-LiDAR: A Multi-Resolution Roadside LiDAR Benchmark for Perception Diagnostics and Deployment Guidance

DGX agent

arXiv:2605.24777v1 Announce Type: new Abstract: LiDAR model selection is a critical issue in roadside sensing systems, as it directly determines both perception capability and deployment cost. However

model-releasesarxiv-cs-ro
26 May 2026
Model Releases

@mteamisloading the models from 6 months ago kinda feel the same like the recently released models. currently not holding my breath for more…

DGX agent

Jeremy Howard comments that large language models released 6 months ago feel comparable in capability to recently released models, suggesting that the pace of improvement in model development may be s

model-releasesjeremy-howard--x
26 May 2026
Model Releases

MuCRASP: Multimodal Chain-of-thought Reasoning aware Structured Pruning

DGX agent

arXiv:2605.25842v1 Announce Type: new Abstract: Vision-language models (VLMs) increasingly rely on chain-of-thought (CoT) reasoning to solve complex multimodal tasks, but their large parameter sizes m

model-releasesarxiv-cs-ai
26 May 2026
← Previous
1…268269270271272…475
Next →