AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,630Total entries
1Added by human
84,629Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,611 results
26 May 2026

InfiFPO: Implicit Model Fusion via Preference Optimization in Large Language Models

Model ReleasesDGX agent

arXiv:2505.13878v3 Announce Type: replace-cross Abstract: Model fusion combines multiple Large Language Models (LLMs) with different strengths into a more powerful, integrated model through lightweigh

Insuring Every Action: An Authority Frontier Framework for Runtime Actuarial Control of Autonomous AI Agents

Model ReleasesDGX agent

arXiv:2605.25632v1 Announce Type: new Abstract: Autonomous AI agents increasingly issue side-effect-bearing actions: database mutations, refunds, payments, external commitments. We propose the Actuari

Introducing CHI-Bench on @huggingface: the world’s first long-horizon healthcare benchmark for AI agents. 75 real healthcare workflows + 20 …


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases
DGX agent

Introducing CHI-Bench on @huggingface: the world’s first long-horizon healthcare benchmark for AI agents. 75 real healthcare workflows + 20 apps + 200+ MCP tools + 1,290 skills + process / outcome rew

IterInject: Indirect Prompt Injection Against LLM Agents via Feedback-Guided Iterative Optimization

Model ReleasesDGX agent

arXiv:2605.24659v1 Announce Type: new Abstract: LLM-based agents are increasingly deployed for complex tasks requiring planning, tool use, and interaction with external services. Their reliance on unt

JacQuant: STE-Free Quantization-Aware Training via Learned Jacobian Surrogates

Model ReleasesDGX agent

arXiv:2605.25469v1 Announce Type: new Abstract: Quantization-aware training (QAT) is widely deployed but typically relies on the Straight-Through Estimator (STE), which passes gradients through non-di

JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments

Model ReleasesDGX agent

arXiv:2602.18527v2 Announce Type: replace-cross Abstract: Current audio-visual large language models (AV-LLMs) are predominantly restricted to 2D perception, relying on RGB video and monaural audio. T

JEPA-DNA: Grounding Genomic Foundation Models through Joint-Embedding Predictive Architectures

Model ReleasesDGX agent

arXiv:2602.17162v2 Announce Type: replace Abstract: Genomic Foundation Models (GFMs) typically rely on Masked Language Modeling (MLM) or Next-Token Prediction (NTP) to learn the 'Laws of Nature'. Whil

JudgmentBench: Comparing Rubric and Preference Evaluation for Quality Assessment

Model ReleasesDGX agent

arXiv:2605.25240v1 Announce Type: cross Abstract: Two methodologies dominate current practices of benchmarking: rubric-based scoring evaluates items against predefined criteria, whereas comparative ju

KAME: Tandem Architecture for Enhancing Knowledge in Real-Time Speech-to-Speech Conversational AI

Model ReleasesDGX agent

arXiv:2510.02327v2 Announce Type: replace-cross Abstract: Real-time speech-to-speech (S2S) models excel at generating natural, low-latency conversational responses but often lack deep knowledge and se

Keep the Proof State Live: Snapshotting for Efficient Tactic Search in Lean 4

Model ReleasesDGX agent

arXiv:2605.25556v1 Announce Type: cross Abstract: Automated theorem proving systems built on Lean 4 increasingly rely on parallel tactic search over partially specified proofs, such as those generated

Kolmogorov-Arnold Fourier Networks

Model ReleasesDGX agent

arXiv:2502.06018v3 Announce Type: replace-cross Abstract: Although Kolmogorov-Arnold-based interpretable networks (KANs) possess strong theoretical expressiveness, they suffer from severe parameter ex

Latent Q-Barrier Shielding for Safe In-Context Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.25267v1 Announce Type: cross Abstract: Safe in-context reinforcement learning (ICRL) adapts online from interaction history without test-time parameter updates while controlling episode cos

Learning Fine-grained Parameter Sharing via Sparse Tensor Decomposition

Model ReleasesDGX agent

arXiv:2411.09816v4 Announce Type: replace Abstract: Large neural networks achieve state-of-the-art performance on many tasks, yet their sheer size hinders deployment on resource-constrained devices. A

Learning Sparse Compositional Functions with Norm-Constrained Neural Networks

Model ReleasesDGX agent

arXiv:2605.25608v1 Announce Type: cross Abstract: The ability of deep neural networks to learn hierarchical features is widely regarded as a key mechanism underlying their success in high-dimensional

Learning to Reason Efficiently with A* Post-Training

Model ReleasesDGX agent

arXiv:2605.24597v1 Announce Type: new Abstract: Many applications of large language models (LLMs) require deductive reasoning, yet models frequently produce incorrect or redundant inference steps. We

LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization

Model ReleasesDGX agent

arXiv:2510.03827v2 Announce Type: replace-cross Abstract: LIBERO has emerged as a widely adopted benchmark for evaluating Vision-Language-Action (VLA) models; however, its current training and evaluat

LiveMCP-101: Stress Testing and Diagnosing MCP-enabled Agents on Challenging Queries

Model ReleasesDGX agent

arXiv:2508.15760v2 Announce Type: replace-cross Abstract: Tool calling has emerged as a critical capability for AI agents. In contrast to conventional tool calling frameworks that rely on static, prov

Llamion Technical Report

Model ReleasesDGX agent

arXiv:2605.25676v1 Announce Type: new Abstract: We release Llamion, a family of 14B-parameter open-weight language models obtained by transforming Orion-14B into the standardized Llama-family architec

LLM-as-a-Reviewer: Benchmarking Their Ability, Divergence, and Prompt Injection Resistance as Paper Reviewers

Model ReleasesDGX agent

arXiv:2605.25415v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in academic peer review, yet their reliability, alignment with human judgment, and robustness to adve

LLMs on prem. but the prem is your wardrobe. http://shop.cohere.com The design team cooked on this one.

Model ReleasesDGX agent

Cohere has announced a creative marketing initiative featuring LLM-themed merchandise available through their shop, with the tagline playing on the phrase 'on premises' to humorously suggest their lan

LLMTabBench: Evaluating LLMs on Binary Tabular Classification From Zero to Few Shots

Model ReleasesDGX agent

arXiv:2605.24417v1 Announce Type: new Abstract: Supervised classification for tabular data remains a core machine learning task, yet its reliance on large labeled datasets limits applicability in data

Looped Diffusion Language Models

Model ReleasesDGX agent

arXiv:2605.26106v1 Announce Type: new Abstract: Masked diffusion models (MDMs) have emerged as a promising alternative to autoregressive models for language modeling, yet the effective design of trans

LWiAI Podcast #246 - Gemini 3.5 + Omni, Musk Loses, OpenAI vs Erdős

Model ReleasesDGX agent

This podcast episode discusses Google's release of Gemini 3.5 and its Omni multimodal capabilities, covers recent developments in Elon Musk's AI ventures, and examines tensions or competition between

M^3-Verse: A 'Spot the Difference' Challenge for Large Multimodal Models

Model ReleasesDGX agent

arXiv:2512.18735v2 Announce Type: replace-cross Abstract: Modern Large Multimodal Models (LMMs) have demonstrated extraordinary ability in static image and single-state spatial-temporal understanding.

MDIA: A Multi-Agent Diagnostic Intelligence Pipeline on HealthBench Professional

Model ReleasesDGX agent

arXiv:2605.24699v1 Announce Type: new Abstract: Most reported gains on agentic-LLM clinical benchmarks are often attributed to prompt engineering, yet our results suggest that larger improvements can

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework

Model ReleasesDGX agent

arXiv:2605.24661v1 Announce Type: new Abstract: LLMs have achieved remarkable success in complex reasoning tasks, yet current evaluation approaches predominantly rely on final-answer correctness, offe

MEDAL: Manifold Embedding Distillation via Autoencoder Learning

Model ReleasesDGX agent

arXiv:2605.24244v1 Announce Type: cross Abstract: Low-dimensional embeddings are widely used as visual summaries of high-dimensional data and to enable downstream scientific discoveries. Yet, popular

Memory-Induced Tool-Drift in LLM Agents

Model ReleasesDGX agent

arXiv:2605.24941v1 Announce Type: cross Abstract: Modern LLM agents combine long-term memory for personalization with tool-calling interfaces for taking actions in the world -- a combination underpinn

Merch, on-premises. http://shop.cohere.com

Model ReleasesDGX agent

Cohere announced merchandise available for purchase at their on-premises shop (shop.cohere.com), likely offering branded items such as apparel, stickers, or other company merchandise for employees, cu

Merge-Bench: Resolve Merge Conflicts with Large Language Models

Model ReleasesDGX agent

arXiv:2605.25890v1 Announce Type: new Abstract: This paper applies machine learning to the difficult and important task of version control merging. (1) We constructed a dataset, Merge-Bench, of 7938 r

Meta-Soft: Leveraging Composable Meta-Tokens for Context-Preserving KV Cache Compression

Model ReleasesDGX agent

arXiv:2605.22337v2 Announce Type: replace Abstract: The KV cache used in large language models has linearly growing time complexity, so LLMs face memory blow-up and reduced decoding efficiency when th

MimirRAG: A Multi-Agent RAG Framework for Financial Data Retrieval with Metadata Integration

Model ReleasesDGX agent

arXiv:2605.25030v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) systems offer a promising approach to reduce hallucinations and improve answer accuracy in large language models (L

MIND: Multi-Scale Intent Diffusion for Text-Driven Physics-Based Humanoid Control

Model ReleasesDGX agent

arXiv:2605.26006v1 Announce Type: cross Abstract: Enabling physics-based humanoids to execute diverse behaviors from high-level textual commands remains a significant challenge. Existing methods typic

MindAlign: Bridging EEG, Vision, and Language for Zero-Shot Visual Decoding

Model ReleasesDGX agent

arXiv:2605.24523v1 Announce Type: cross Abstract: Visual decoding from brain signals is a key challenge at the intersection of computer vision and neuroscience, requiring methods that bridge neural re

Minimax Limits of k-Fold Cross-Validation via Majority

Model ReleasesDGX agent

arXiv:2605.25859v1 Announce Type: cross Abstract: We study the mean-squared error of k-fold cross-validation as a risk estimator, with particular emphasis on how its accuracy depends on the number of

Mistral-7B v0.3 at 128K in llama.cpp: 22,657 → 13,235 MiB live VRAM with ≤0.004 PPL drift

Model ReleasesDGX agent

Mistral-7B v0.3 model achieves significant memory optimization when running at 128K context length in llama.cpp, reducing live VRAM usage from 22,657 MiB to 13,235 MiB while maintaining minimal perfor

Mix-MoE: Improving Multilingual Machine Translation of Large Language Models through Mixed MoEs

Model ReleasesDGX agent

arXiv:2605.24681v1 Announce Type: cross Abstract: Large Language Models (LLMs) have shown great promise in multilingual machine translation (MT), even with limited bilingual supervision. However, fine

MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence

Model ReleasesDGX agent

arXiv:2505.23764v3 Announce Type: replace-cross Abstract: Spatial intelligence is essential for multimodal large language models (MLLMs) operating in the complex physical world. Existing benchmarks, h

Mosaic: Compositional Multi-Concept Erasure via Vector Field Blending

Model ReleasesDGX agent

arXiv:2605.25574v1 Announce Type: cross Abstract: Concept erasure has emerged as a key research direction for ensuring safe and ethical image synthesis in Text-to-Image (T2I) models. While existing st

MR-LiDAR: A Multi-Resolution Roadside LiDAR Benchmark for Perception Diagnostics and Deployment Guidance

Model ReleasesDGX agent

arXiv:2605.24777v1 Announce Type: new Abstract: LiDAR model selection is a critical issue in roadside sensing systems, as it directly determines both perception capability and deployment cost. However

@mteamisloading the models from 6 months ago kinda feel the same like the recently released models. currently not holding my breath for more…

Model ReleasesDGX agent

Jeremy Howard comments that large language models released 6 months ago feel comparable in capability to recently released models, suggesting that the pace of improvement in model development may be s

MuCRASP: Multimodal Chain-of-thought Reasoning aware Structured Pruning

Model ReleasesDGX agent

arXiv:2605.25842v1 Announce Type: new Abstract: Vision-language models (VLMs) increasingly rely on chain-of-thought (CoT) reasoning to solve complex multimodal tasks, but their large parameter sizes m

MuJoCoUni:Persistent Batched Runtime Primitives for MuJoCo

Model ReleasesDGX agent

arXiv:2605.24922v1 Announce Type: new Abstract: We present MuJoCoUni, a downstream MuJoCo distribution for online robot learning and batched physics evaluation. Alongside the open-loop batched traject

MultiHaluDet: Multilingual Hallucination Detection via LLM Hidden State Probing

Model ReleasesDGX agent

arXiv:2605.24919v1 Announce Type: new Abstract: Hallucinations in Large Language Models (LLMs) represent a critical barrier to their reliable deployment, a vulnerability heavily exacerbated in non-Eng

MuNet: A Mutualistic Network for Joint 3D Human Mesh Recovery and 3D Clothed Human Reconstruction from Single Images

Model ReleasesDGX agent

arXiv:2605.25861v1 Announce Type: cross Abstract: 3D human mesh recovery and 3D clothed human reconstruction are inherently related, yet they have long been studied in isolation, thereby overlooking t

Muon in Associative Memory Learning: Training Dynamics and Scaling Laws

Model ReleasesDGX agent

arXiv:2602.05725v2 Announce Type: replace Abstract: Muon updates matrix parameters via the matrix sign of the gradient and has shown strong empirical gains, yet its dynamics and scaling behavior remai

Neural Integral Operators for Inverse Problems: An Operator-Learning Framework for Small-Sample Spectroscopic Classification

Model ReleasesDGX agent

arXiv:2505.03677v3 Announce Type: replace Abstract: Learning maps between function spaces with a strong inductive bias is a central challenge in soft computing, especially when training data are scarc

Neural Router: Semantic Content Matching for Agentic AI

Model ReleasesDGX agent

arXiv:2605.25701v1 Announce Type: cross Abstract: Large language models (LLMs) can serve as the semantic-matching engine of a content-based publish/subscribe broker for agentic AI across the edge-clou

Neural Scalable Symbolic Search Framework for Complex Logical Queries with Multiple Free Variables

Model ReleasesDGX agent

arXiv:2605.25985v1 Announce Type: new Abstract: Complex Query Answering (CQA) is a fundamental knowledge representation and reasoning task over incomplete knowledge graphs (KGs). Answering existential

Neuromorphic LiDAR-based Bird's Eye View Object Detection using Energy-efficient Spiking Neural Networks

Model ReleasesDGX agent

arXiv:2605.25293v1 Announce Type: cross Abstract: Autonomous driving perception demands accurate and efficient processing of three-dimensional sensor data under strict power constraints. Traditional c

Neurosymbolic AI is saving deep learning from hitting the wall (*exactly* as I said my 2022 paper “deep learning is hitting a wall”) Really …

Model ReleasesDGX agent

Neurosymbolic AI is saving deep learning from hitting the wall (*exactly* as I said my 2022 paper “deep learning is hitting a wall”) Really sad to see someone as smart @peterwildeford confusing the or

New on the Engineering Blog: The access and permissions we grant agents should evolve with their capabilities. In our own products, we set t…

Model ReleasesDGX agent

New on the Engineering Blog: The access and permissions we grant agents should evolve with their capabilities. In our own products, we set these parameters through sandboxing, which limits the scope o

Nine reasons why I warned OpenAI might fail and turn out to be the WeWork of AI, from two years ago.

Model ReleasesDGX agent

Nine reasons why I warned OpenAI might fail and turn out to be the WeWork of AI, from two years ago. 9 reasons that OpenAI could someday be seen as the WeWork of AI: 👉 Lots of competitors are catching

Noise-Robust Financial Numerical Entity Attribute Tagging

Model ReleasesDGX agent

arXiv:2605.24910v1 Announce Type: new Abstract: Financial Numerical Entity (FNE) understanding aims to recover the meaning of numerical mentions in financial reports. Existing studies primarily focus

Nonstationary Generalized Linear Bandits with Discounted Online Mirror Descent

Model ReleasesDGX agent

arXiv:2605.25590v1 Announce Type: cross Abstract: We study nonstationary generalized linear bandits (GLBs), where the expected reward is modeled through a nonlinear link function with an unknown time-

NormimesDirection: Restoring the Missing Query Norm in Vision Linear Attention

Model ReleasesDGX agent

arXiv:2506.21137v3 Announce Type: replace Abstract: Linear attention mitigates the quadratic complexity of softmax attention but suffers from a critical loss of expressiveness. We identify two primary

Novee debuts Agentic Fix, pushing pentest findings into Claude, Copilot and Cursor

Model ReleasesDGX agent

Artificial intelligence penetration testing startup Novee Cyber Security Ltd. today launched Agentic Fix, a new capability that pushes validated exploit findings directly into the AI coding agents dev

NVIDIA Vera CPU Is ‘Packing a Heavy-Hitting Punch’ Against Competition

Model ReleasesDGX agent

The shift to agentic AI creates a new CPU requirement for the AI factory: fast cores, massive memory bandwidth and the ability to sustain high performance when all cores are active. Initial benchmark

Omissive Bias in Religious Representation: Benchmarking LLM Answers to Everyday Ethical Decision-making

Model ReleasesDGX agent

arXiv:2605.24319v1 Announce Type: new Abstract: As large language models become a default source of guidance on personal, moral, and existential questions, it matters whether they draw on the religiou

On the Epistemic Uncertainty of Overparametrized Neural Networks

Model ReleasesDGX agent

arXiv:2605.25234v1 Announce Type: cross Abstract: Epistemic uncertainty is often viewed as a reducible uncertainty that vanishes with increasing data. This perspective implicitly assumes parameter ide

← Previous
1…212213214215216…377
Next →