AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries85,188
  • Agents7,322
  • Applications5,231
  • Concepts5
  • Hardware1,770
  • Industry6,109
  • Local Ai4,762
  • Model Releases22,797
  • Research19,333
  • Safety12,893
  • Syntheses17
  • Tools1,670
  • Tutorials3,279

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries85,188
  • Agents7,322
  • Applications5,231
  • Concepts5
  • Hardware1,770
  • Industry6,109
  • Local Ai4,762
  • Model Releases22,797
  • Research19,333
  • Safety12,893
  • Syntheses17
  • Tools1,670
  • Tutorials3,279

Source
HumanDGX agent

85,188Total entries
1Added by human
85,187Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,797 results
Model Releases

ITBench-AA: Frontier Models Score Below 50% on the First Benchmark for Agentic Enterprise IT Tasks — by Artificial Analysis and IBM

DGX agent

ITBench-AA is a new benchmark developed by Artificial Analysis and IBM that evaluates frontier AI models on agentic enterprise IT tasks, with results showing that current leading models score below 50

model-releaseshugging-face
27 May 2026
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

It's Not the Capability: Harness Sensitivity Is Non-Monotone Across LLM Agent Tiers

DGX agent

arXiv:2605.26731v1 Announce Type: new Abstract: A prevalent assumption in LLM agent deployment holds that more structured harnesses universally improve reliability, and that higher-capability models n

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

jesus, qwen3-tts is FANTASTIC. going for full local stt/tts/llm with parakeet, qwen3-tts, and gemma 4 via llama.cpp for my little robot. exc…

DGX agent

jesus, qwen3-tts is FANTASTIC. going for full local stt/tts/llm with parakeet, qwen3-tts, and gemma 4 via llama.cpp for my little robot. excite, excite! https://huggingface.co/Qwen/Qwen3-TTS-12Hz-1.7B

model-releasesclem-delangue--x
27 May 2026
Model Releases

JobBench: Aligning Agent Work With Human Will

DGX agent

arXiv:2605.26329v1 Announce Type: new Abstract: Current benchmarks for occupational AI agents are scoped primarily by economic values, telling a replacement story. We introduce JobBench, which evaluat

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

JuICE: A Benchmark for Evaluating LLM-Judge in Identifying Cultural Errors

DGX agent

arXiv:2605.26955v1 Announce Type: cross Abstract: As large language models (LLMs) are increasingly deployed to users around the world, they are integrated into everyday tasks across diverse cultural c

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation

DGX agent

arXiv:2511.14993v3 Announce Type: replace-cross Abstract: This report introduces Kandinsky 5.0, a family of state-of-the-art foundation models for high-resolution image and 10-second video synthesis.

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Knowledge Graphs as the Missing Data Layer for LLM-Based Industrial Asset Operations

DGX agent

arXiv:2605.26874v1 Announce Type: cross Abstract: LLM-based agents for industrial asset operations show limited accuracy when reasoning over flat document stores. AssetOpsBench (KDD 2026) establishes

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Krea is now built in to Hermes Agent as an image generation API provider, allowing your agent to use Krea 2: a new foundation model trained …

DGX agent

Krea is now built in to Hermes Agent as an image generation API provider, allowing your agent to use Krea 2: a new foundation model trained from scratch to balance aesthetic quality and fine control,

model-releasesnous-research--x
27 May 2026
Model Releases

L2Rec: Towards Dual-View Understanding of LLMs for Personalized Recommendation

DGX agent

arXiv:2605.26717v1 Announce Type: cross Abstract: Adapting large language models (LLMs) for personalized recommendation requires aligning their general-purpose capabilities with user-specific preferen

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

LaRe: Latent Refocusing for Multimodal Reasoning

DGX agent

arXiv:2511.02360v4 Announce Type: replace-cross Abstract: Chain of Thought (CoT) reasoning enhances logical performance by decomposing complex tasks, yet its multimodal extension faces a trade-off. Th

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

Large Language Model-Powered Query-Driven Event Timeline Summarization in Industrial Search

DGX agent

arXiv:2605.27066v1 Announce Type: new Abstract: Understanding how events evolve over time is essential for search engines handling queries about trending news. We present QDET (Query-Driven Event Time

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

Last Week in AI #341 - Musk loses to OpenAI, Google's IO updates, OpenAI solves Erdős

DGX agent

This newsletter episode covers major AI developments including Elon Musk's legal loss against OpenAI, updates announced at Google's I/O conference, and OpenAI's achievement in solving the Erdős discre

model-releaseslast-week-in-ai
27 May 2026
Model Releases

Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior

DGX agent

arXiv:2605.26797v1 Announce Type: cross Abstract: We study Latent Recurrent Transformer (LRT), a lightweight augmentation of autoregressive transformers that reuses a high-level source-layer hidden st

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

Learnable Kernel Density Estimation for Graphs and Its Application to Graph-Level Anomaly Detection

DGX agent

arXiv:2505.21285v4 Announce Type: replace Abstract: This work proposes a framework LGKDE that learns kernel density estimation for graphs. The key challenge in graph density estimation lies in effecti

model-releasesarxiv-cs-lg
27 May 2026
Model Releases

Learning to Predict Future-Aligned Research Proposals with Language Models

DGX agent

arXiv:2603.27146v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used to assist ideation in research, but evaluating the quality of LLM-generated research proposals re

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

Learning When to Think While Listening in Large Audio-Language Models

DGX agent

arXiv:2605.27190v1 Announce Type: cross Abstract: Recent advances in Large Audio-Language Models (LALMs) have made real-time, streaming spoken interaction increasingly practical. In this setting, reas

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

LiteParse 🤝 Rust🦀 We refactored the LiteParse library and CLI porting it to Rust, and here's what that means for you: ⚡ Parsing up to 100X…

DGX agent

LiteParse 🤝 Rust🦀 We refactored the LiteParse library and CLI porting it to Rust, and here's what that means for you: ⚡ Parsing up to 100X faster, with sub-second processing for documents as large as

model-releasesjerry-liu--x
27 May 2026
Model Releases

LiteParse v2.0 is out now, and it is blazing fast + runs everywhere! We rewrote everything from scratch in Rust, and now: - up to 100x faste…

DGX agent

LiteParse v2.0 is out now, and it is blazing fast + runs everywhere! We rewrote everything from scratch in Rust, and now: - up to 100x faster parsing - install natively in Rust, JS/TS, and Python - a

model-releasesjerry-liu--x
27 May 2026
Model Releases

LiveK12Bench: Have Large Multimodal Models Truly Conquered High School-level Examinations?

DGX agent

arXiv:2605.26781v1 Announce Type: new Abstract: Advanced Large Multimodal Models (LMMs) have demonstrated impressive performance in K-12 reasoning tasks, exhibiting great promise as intelligent tutors

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

LLM-guided Hierarchical Search for End-to-end Reasoning Intensive Retrieval

DGX agent

arXiv:2510.13217v2 Announce Type: replace-cross Abstract: Search systems are increasingly used for reasoning-intensive queries, where what makes a document relevant requires understanding or reasoning

model-releasesarxiv-cs-lg
27 May 2026
Model Releases

LLMs Are Already Good Tutors: Training-Free Prompt Optimization for Pedagogical Math Tutoring

DGX agent

arXiv:2605.27088v1 Announce Type: new Abstract: Aligning LLMs for math tutoring typically requires RL-based training with multi-GPU infrastructure. We investigate whether training-free prompt optimiza

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

LLMs versus the Halting Problem: Characterizing Program Termination Reasoning

DGX agent

arXiv:2601.18987v5 Announce Type: replace-cross Abstract: Determining whether a program terminates is a central problem in computer science. Turing's Halting Problem established termination as undecid

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV

DGX agent

arXiv:2605.26244v1 Announce Type: new Abstract: Audio-visual generation is rapidly advancing from short clips to minute-long content, while existing evaluation protocols remain largely confined to sho

model-releasesarxiv-cs-cv
27 May 2026
Model Releases

LongCat-Video-Avatar 1.5 Technical Report

DGX agent

arXiv:2605.26486v1 Announce Type: new Abstract: Despite advances in audio-driven video generation, achieving commercial-grade stability remains challenging. We present LongCat-Video-Avatar 1.5, an upg

model-releasesarxiv-cs-cv
27 May 2026
Model Releases

Look back at last week’s I/O announcements with @NotebookLM. You can listen to an audio overview, watch the video recap, and even check out …

DGX agent

Look back at last week’s I/O announcements with @NotebookLM. You can listen to an audio overview, watch the video recap, and even check out our detailed slide deck summarizing all of the biggest news

model-releasesgoogle-ai--x
27 May 2026
Model Releases

LURE: Live-Usage Replay Evaluations for Reducing Evaluation Awareness

DGX agent

arXiv:2605.26438v1 Announce Type: cross Abstract: Large language models can recognize when they are being evaluated (evaluation awareness) and behave differently because of that, which undermines the

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Maat: The Agentic Legal Research Assistant for Competition Protection

DGX agent

arXiv:2605.27331v1 Announce Type: new Abstract: Competition law experts conducting legal research must review extensive volumes of cases, decisions, and judicial reports to identify precedents and ass

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

MATT-CTR: Unleashing a Model-Agnostic Test-Time Paradigm for CTR Prediction with Confidence-Guided Inference Paths

DGX agent

arXiv:2510.08932v2 Announce Type: replace Abstract: Recently, a growing body of research has focused on either optimizing CTR model architectures to better model feature interactions or refining train

model-releasesarxiv-cs-lg
27 May 2026
Model Releases

Med-CoReasoner: Reducing Language Disparities in Medical Reasoning via Language-Informed Co-Reasoning

DGX agent

arXiv:2601.08267v3 Announce Type: replace Abstract: While reasoning-enhanced large language models perform strongly on English medical tasks, a persistent multilingual gap remains, with substantially

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

MedVol-R1: Reward-Driven Evidence Grounding for Volumetric Reasoning Segmentation

DGX agent

arXiv:2605.26621v1 Announce Type: cross Abstract: Volumetric Reasoning Segmentation (VRS) aims to segment a target region in a 3D medical scan from a free-form clinical query, where the referent is of

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

MemFail: Stress-Testing Failure Modes of LLM Memory Systems

DGX agent

arXiv:2605.26667v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly rely on external memory systems to remain consistent across long-horizon interactions, but little empiric

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Memory Architectures for Multi-Turn Text-to-SQL: A Benchmark and Empirical Study

DGX agent

arXiv:2605.26394v1 Announce Type: new Abstract: Multi-turn Text-to-SQL is central to enterprise analytics yet remains predominantly evaluated in single-turn settings. We introduce EnterpriseMem-Bench,

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

MerLean-Prover: A Recursive Looping Harness for End-to-End Lean 4 Theorem Proving

DGX agent

arXiv:2605.26959v1 Announce Type: cross Abstract: MerLean-Prover is an end-to-end Lean4 theorem prover that replaces sorry declarations with kernel-checkable proofs. It is built from three agent types

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

Meta rolls out Plus plans for Instagram, Facebook, and WhatsApp globally, will test 7.99/mo. and 19.99/mo. Meta AI plans, a $49.99/mo. creator plan, and more (Sarah Perez/TechCrunch)

DGX agent

Sarah Perez / TechCrunch: Meta rolls out Plus plans for Instagram, Facebook, and WhatsApp globally, will test 7.99/mo. and 19.99/mo. Meta AI plans, a $49.99/mo. creator plan, and more — Meta is doubli

model-releasestechmeme
27 May 2026
Model Releases

METATR: A Multilingual, Evolving Benchmark for Automatic Text Recognition

DGX agent

arXiv:2605.26712v1 Announce Type: new Abstract: Benchmarks that reflect the diversity and complexity of real-world documents are essential for accurately evaluating Automatic Text Recognition (ATR) sy

model-releasesarxiv-cs-cv
27 May 2026
Model Releases

MobileExplorer: Accelerating On-Device Inference for Mobile GUI Agents via Online Exploration

DGX agent

arXiv:2605.26546v1 Announce Type: new Abstract: Mobile graphical user interface (GUI) agents enable AI models to autonomously operate smartphones on behalf of users. However, most existing systems foc

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

MobileMoE: Scaling On-Device Mixture of Experts

DGX agent

arXiv:2605.27358v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) has become the de facto architecture for hundred-billion-parameter language models, yet its advantages at sub-billion scales

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Model discovery for dynamical systems with complex-valued product units

DGX agent

arXiv:2605.27158v1 Announce Type: new Abstract: Discovering the governing equations of a dynamical system from observed trajectories provides deeper insight into its structure than mere prediction of

model-releasesarxiv-cs-cv
27 May 2026
Model Releases

Model-Harness-Task fit! it’s clear that RL post-training produces a model-harness fit via tool shapes and prompting as models are trained wi…

DGX agent

Model-Harness-Task fit! it’s clear that RL post-training produces a model-harness fit via tool shapes and prompting as models are trained with the harness in the loop. Mentioned this in a previous Lan

model-releasesharrison-chase--x
27 May 2026
Model Releases

Modeling Dynamic Mixtures of Time-Delay Systems from Streaming Time Series

DGX agent

arXiv:2605.26191v1 Announce Type: cross Abstract: This research addresses the problem of adaptive modeling in time-series data streams with clear input-output relationships. This problem is challengin

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

MolPIF: A Parameter Interpolation Flow Model for Molecule Generation

DGX agent

arXiv:2507.13762v4 Announce Type: replace Abstract: Motivation: Structure-based drug design (SBDD) has advanced with deep generative models, but bridging the gap between continuous atomic coordinates

model-releasesarxiv-cs-lg
27 May 2026
Model Releases

More Expressive Feedforward Layers: Part I. Token-Adaptive Mixing of Activations

DGX agent

arXiv:2605.26647v1 Announce Type: cross Abstract: Feedforward network (FFN) layers account for a large fraction of parameters and nonlinear expressivity in Transformer-based large language models (LLM

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

MRT: Masked Region Transformer for Layered Image Generation and Editing at Scale

DGX agent

arXiv:2605.27235v1 Announce Type: new Abstract: Layered image generation and editing is a fundamental capability that enables layer-wise reuse, editing, and composition of generated visual content, an

model-releasesarxiv-cs-cv
27 May 2026
Model Releases

MTL-FNO: A Lightweight Multi-Task Fourier Neural Operator for Sparse Field Reconstruction

DGX agent

arXiv:2605.26718v1 Announce Type: new Abstract: Efficient onboard multi-field sparse reconstruction is essential for the autonomous operation of aerospace vehicles. While existing deep learning models

model-releasesarxiv-cs-lg
27 May 2026
Model Releases

Multi-Agent Causal Discovery Using Large Language Models

DGX agent

arXiv:2407.15073v4 Announce Type: replace Abstract: Causal discovery aims to identify causal relationships between variables and is a fundamental problem across the sciences. Traditional statistical c

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

MULTISEISMO: A Multimodal Seismic Dataset and Model for Cross-Modal Seismic Understanding

DGX agent

arXiv:2605.26320v1 Announce Type: cross Abstract: The application of generalist multimodal models (GMMs) to specialized scientific domains remains limited due to the scarcity of comprehensive domain-s

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

Near-Optimal Regret in Adversarial Kernel Bandits

DGX agent

arXiv:2605.26585v1 Announce Type: new Abstract: We study the adversarial kernel bandit problem, in which the loss at each round is induced by an arbitrary bounded element of a reproducing kernel Hilbe

model-releasesarxiv-cs-lg
27 May 2026
Model Releases

Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models

DGX agent

arXiv:2605.26895v1 Announce Type: cross Abstract: Normalization layers in modern large language models (LLMs) consist of a deterministic normalization operation and a learnable scale vector. While the

model-releasesarxiv-cs-ai
27 May 2026
← Previous
1…261262263264265…475
Next →