AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
All
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,332 results
Model Releases

Inflated Excellence or True Performance? Rethinking Medical Diagnostic Benchmarks with Dynamic Evaluation

DGX agent

arXiv:2510.09275v2 Announce Type: replace Abstract: Medical diagnostics is a high-stakes and complex domain that is critical to patient care. However, current evaluations of large language models (LLM

model-releasesarxiv-cs-cl
21 Apr 2026
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Injecting Structured Biomedical Knowledge into Language Models: Continual Pretraining vs. GraphRAG

DGX agent

arXiv:2604.16422v1 Announce Type: new Abstract: The injection of domain-specific knowledge is crucial for adapting language models (LMs) to specialized fields such as biomedicine. While most current a

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

INTENT: Invariance and Discrimination-aware Noise Mitigation for Robust Composed Image Retrieval

DGX agent

arXiv:2604.18051v1 Announce Type: new Abstract: Composed Image Retrieval (CIR) is a challenging image retrieval paradigm that enables to retrieve target images based on multimodal queries consisting o

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

InternScenes: A Large-scale Simulatable Indoor Scene Dataset with Realistic Layouts

DGX agent

arXiv:2509.10813v3 Announce Type: replace Abstract: The advancement of Embodied AI heavily relies on large-scale, simulatable 3D scene datasets characterized by scene diversity and realistic layouts.

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

Interpolating Discrete Diffusion Models with Controllable Resampling

DGX agent

arXiv:2604.17310v1 Announce Type: new Abstract: Discrete diffusion models form a powerful class of generative models across diverse domains, including text and graphs. However, existing approaches fac

model-releasesarxiv-cs-lg
21 Apr 2026
Model Releases

Introducing ChatGPT Images 2.0

DGX agent

ChatGPT Images 2.0 is an updated version of OpenAI's image generation feature integrated into ChatGPT, offering improvements to image creation capabilities for users. The update likely includes enhanc

model-releasesopenai
21 Apr 2026
Model Releases

Introducing ChatGPT Images 2.0 A state-of-the-art image model that can take on complex visual tasks and produce precise, immediately usable …

DGX agent

Introducing ChatGPT Images 2.0 A state-of-the-art image model that can take on complex visual tasks and produce precise, immediately usable visuals, with sharper editing, richer layouts, and thinking-

model-releasesopenai--x
21 Apr 2026
Model Releases

Introducing ml-intern, the agent that just automated the post-training team @huggingface It's an open-source implementation of the real rese…

DGX agent

Introducing ml-intern, the agent that just automated the post-training team @huggingface It's an open-source implementation of the real research loop that our ML researchers do every day. You give it

model-releasesclem-delangue--x
21 Apr 2026
Model Releases

Judge a Book by its Cover: Investigating Multi-Modal LLMs for Multi-Page Handwritten Document Transcription

DGX agent

arXiv:2502.20295v2 Announce Type: replace-cross Abstract: Handwriting text recognition (HTR) remains a challenging task. Existing approaches require fine-tuning on labeled data, which is impractical t

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

JudgeMeNot: Personalizing Large Language Models to Emulate Judicial Reasoning in Hebrew

DGX agent

arXiv:2604.18041v1 Announce Type: new Abstract: Despite significant advances in large language models, personalizing them for individual decision-makers remains an open problem. Here, we introduce a s

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Jupiter-N Technical Report

DGX agent

arXiv:2604.17429v1 Announce Type: new Abstract: We present Jupiter-N, a hybrid reasoning model post-trained from Nemotron 3 Super, a fully open-source 120 billion parameter LLM. We target three object

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

K2.6 + hermes = 4 hr session setting up qwen 3.6 training regime on dgx spark with the current autnomous session currently lasting 70+ min w…

DGX agent

K2.6 + hermes = 4 hr session setting up qwen 3.6 training regime on dgx spark with the current autnomous session currently lasting 70+ min without any prompting. Kimi with hermes is next level. @NousR

model-releasesnous-research--x
21 Apr 2026
Model Releases

Keep improving 🚀🚀 @arena

DGX agent

Keep improving 🚀🚀 @arena Qwen3.6 Plus lands at #7 in Code Arena with a score of 1476 - up +16 points since the Preview. The new score also moves @AlibabaGroup to #3 lab in Code Arena. In the Text Aren

model-releasesqwen--x
21 Apr 2026
Model Releases

📢 Kimi K2.6 API is live • Input Price (Cache Hit): 0.16 / M tokens • Input Price (Cache Miss): 0.95 / M tokens • Output: $4.00 / M tokens…

DGX agent

📢 Kimi K2.6 API is live • Input Price (Cache Hit): 0.16 / M tokens • Input Price (Cache Miss): 0.95 / M tokens • Output: $4.00 / M tokens Kimi K2.6 is our latest + most intelligent model - stronger lo

model-releaseskimi-moonshot--x
21 Apr 2026
Model Releases

Kimi K2.6 autonomously overhauled exchange-core, an 8-year-old open-source financial matching engine. Over a 13-hour execution, the model it…

DGX agent

Kimi K2.6 autonomously overhauled exchange-core, an 8-year-old open-source financial matching engine. Over a 13-hour execution, the model iterated through 12 optimization strategies, initiating over 1

model-releaseskimi-moonshot--x
21 Apr 2026
Model Releases

Kimi K2.6 has captured #1 on the open-weight Vals Index, and is #7 overall.

DGX agent

Kimi K2.6, a language model developed by Moonshot AI, has achieved the top ranking on the open-weight category of the Vals Index benchmark, while placing 7th overall across all model categories. The V

model-releaseskimi-moonshot--x
21 Apr 2026
Model Releases

Kimi K2.6 is now live inside Anything!

DGX agent

Kimi K2.6, an AI model from Moonshot, has been integrated into the Anything platform. This update likely enables users to access Kimi's capabilities directly within the Anything application interface.

model-releaseskimi-moonshot--x
21 Apr 2026
Model Releases

Kimi K2.6 @Kimi_Moonshot is the new leading open-weights agent model, landing at #4 on Claw-Eval (Pass^3: 62.3%). Key takeaways: - 👑 Best o…

DGX agent

Kimi K2.6 @Kimi_Moonshot is the new leading open-weights agent model, landing at #4 on Claw-Eval (Pass^3: 62.3%). Key takeaways: - 👑 Best open-source agent, period: Pass^3 of 62.3% is the highest of a

model-releaseskimi-moonshot--x
21 Apr 2026
Model Releases

KIRA: Knowledge-Intensive Image Retrieval and Reasoning Architecture for Specialized Visual Domains

DGX agent

arXiv:2604.16915v1 Announce Type: new Abstract: Retrieval augmented generation (RAG) has transformed text based question answering, yet its extension to visual domains remains hindered by fundamental

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

Knowing When to Quit: A Principled Framework for Dynamic Abstention in LLM Reasoning

DGX agent

arXiv:2604.18419v1 Announce Type: cross Abstract: Large language models (LLMs) using chain-of-thought reasoning often waste substantial compute by producing long, incorrect responses. Abstention can m

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Knowledge without Wisdom: Measuring Misalignment between LLMs and Intended Impact

DGX agent

arXiv:2603.00883v2 Announce Type: replace Abstract: LLMs increasingly excel on AI benchmarks, but doing so does not guarantee validity for downstream tasks. This study contrasts LLM alignment on bench

model-releasesarxiv-cs-lg
21 Apr 2026
Model Releases

Language Models Don't Know What You Want: Evaluating Personalization in Deep Research Needs Real Users

DGX agent

arXiv:2603.16120v2 Announce Type: replace Abstract: Deep Research (DR) systems help researchers cope with ballooning publishing counts. Such tools synthesize scientific papers to answer research queri

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Large Language Models Are Still Misled by Simple Bias Ensembles

DGX agent

arXiv:2505.16522v3 Announce Type: replace Abstract: With the evolution of large language models (LLMs), their robustness against individual simple biases has been enhanced. However, we observe that th

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Late Fusion Neural Operators for Extrapolation Across Parameter Space in Partial Differential Equations

DGX agent

arXiv:2604.16721v1 Announce Type: new Abstract: Developing neural operators that accurately predict the behavior of systems governed by partial differential equations (PDEs) across unseen parameter re

model-releasesarxiv-cs-lg
21 Apr 2026
Model Releases

Latent Preference Modeling for Cross-Session Personalized Tool Calling

DGX agent

arXiv:2604.17886v1 Announce Type: new Abstract: Users often omit essential details in their requests to LLM-based agents, resulting in under-specified inputs for tool use. This poses a fundamental cha

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

LayerCache: Exploiting Layer-wise Velocity Heterogeneity for Efficient Flow Matching Inference

DGX agent

arXiv:2604.16492v1 Announce Type: new Abstract: Flow Matching models achieve state-of-the-art image generation quality but incur substantial inference cost due to iterative denoising through large Tra

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

LEAF: Knowledge Distillation of Text Embedding Models with Teacher-Aligned Representations

DGX agent

arXiv:2509.12539v2 Announce Type: replace-cross Abstract: We present LEAF ('Lightweight Embedding Alignment Framework'), a knowledge distillation framework for text embedding models. A key distinguish

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Learned Nonlocal Feature Matching and Filtering for RAW Image Denoising

DGX agent

arXiv:2604.17453v1 Announce Type: cross Abstract: Being one of the oldest and most basic problems in image processing, image denoising has seen a resurgence spurred by rapid advances in deep learning.

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

Learning Stable Predictors from Weak Supervision under Distribution Shift

DGX agent

arXiv:2604.05002v2 Announce Type: replace Abstract: Learning from weak, proxy, or relative supervision is common when ground-truth labels are unavailable, but robustness under distribution shift remai

model-releasesarxiv-cs-lg
21 Apr 2026
Model Releases

Learning to Control Summaries with Score Ranking

DGX agent

arXiv:2604.17197v1 Announce Type: new Abstract: Recent advances in summarization research focus on improving summary quality across multiple criteria, such as completeness, conciseness, and faithfulne

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Learning to Retrieve User History and Generate User Profiles for Personalized Persuasiveness Prediction

DGX agent

arXiv:2601.05654v3 Announce Type: replace Abstract: Estimating the persuasiveness of messages is critical in various applications, from recommender systems to safety assessment of LLMs. While it is im

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Let's talk parsing charts 📊📈. Last week we released ParseBench, the first document OCR benchmark for AI agents. New in ParseBench: ChartDa…

DGX agent

Let's talk parsing charts 📊📈. Last week we released ParseBench, the first document OCR benchmark for AI agents. New in ParseBench: ChartDataPointMatch. Most document look at a chart and OCR the captio

model-releasesjerry-liu--x
21 Apr 2026
Model Releases

Leveraging Large Language Models for Sarcastic Speech Annotation in Sarcasm Detection

DGX agent

arXiv:2506.00955v2 Announce Type: replace Abstract: Sarcasm fundamentally alters meaning through tone and context, yet detecting it in speech remains a challenge due to data scarcity. In addition, exi

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

LexRel: Benchmarking Legal Relation Extraction for Chinese Civil Cases

DGX agent

arXiv:2512.12643v2 Announce Type: replace Abstract: Legal relations serve as an important analytical framework for dispute resolution in civil cases. However, legal relations in Chinese civil cases re

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

LiFT: Does Instruction Fine-Tuning Improve In-Context Learning for Longitudinal Modelling by Large Language Models?

DGX agent

arXiv:2604.16382v1 Announce Type: new Abstract: Longitudinal NLP tasks require reasoning over temporally ordered text to detect persistence and change in human behavior and opinions. However, in-conte

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

LIFT the Veil for the Truth: Principal Weights Emerge after Rank Reduction for Reasoning-Focused Supervised Fine-Tuning

DGX agent

arXiv:2506.00772v2 Announce Type: replace-cross Abstract: Recent studies have shown that supervised fine-tuning of LLMs on a small number of high-quality datasets can yield strong reasoning capabiliti

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

LiquidTAD: An Efficient Method for Temporal Action Detection via Liquid Neural Dynamics

DGX agent

arXiv:2604.18274v1 Announce Type: new Abstract: Temporal Action Detection (TAD) in untrimmed videos is currently dominated by Transformer-based architectures. While high-performing, their quadratic co

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length

DGX agent

arXiv:2512.04677v5 Announce Type: replace Abstract: Audio-driven avatar interaction demands real-time, streaming, and infinite-length generation -- capabilities fundamentally at odds with the sequenti

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

LIVE: Leveraging Image Manipulation Priors for Instruction-based Video Editing

DGX agent

arXiv:2604.17021v1 Announce Type: new Abstract: Video editing aims to modify input videos according to user intent. Recently, end-to-end training methods have garnered widespread attention, constructi

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

LiveFact: A Dynamic, Time-Aware Benchmark for LLM-Driven Fake News Detection

DGX agent

arXiv:2604.04815v2 Announce Type: replace Abstract: The rapid development of Large Language Models (LLMs) has transformed fake news detection and fact-checking tasks from simple classification to comp

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Lizard: An Efficient Linearization Framework for Large Language Models

DGX agent

arXiv:2507.09025v4 Announce Type: replace Abstract: We propose Lizard, a linearization framework that transforms pretrained Transformer-based Large Language Models (LLMs) into subquadratic architectur

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

LLaMA-XR: A Novel Framework for Radiology Report Generation using LLaMA and QLoRA Fine Tuning

DGX agent

arXiv:2506.03178v2 Announce Type: replace-cross Abstract: Automated radiology report generation holds significant potential to reduce radiologists' workload and enhance diagnostic accuracy. However, g

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

LLMs are still not consistent judges of qualitative work, and small changes to how that work is presented affect outcomes. Better harnessing…

DGX agent

LLMs are still not consistent judges of qualitative work, and small changes to how that work is presented affect outcomes. Better harnessing and methods (multiple judging runs with randomized orders,

model-releasesethan-mollick--x
21 Apr 2026
Model Releases

LOGICAL-COMMONSENSEQA: A Benchmark for Logical Commonsense Reasoning

DGX agent

arXiv:2601.16504v3 Announce Type: replace Abstract: Commonsense reasoning often involves evaluating multiple plausible interpretations rather than selecting a single atomic answer, yet most benchmarks

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Logit Arithmetic Elicits Long Reasoning Capabilities Without Training

DGX agent

arXiv:2510.09354v2 Announce Type: replace Abstract: Large reasoning models exhibit long chain-of-thought reasoning with complex strategies such as backtracking and self-verification. Yet, these capabi

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Long-CODE: Isolating Pure Long-Context as an Orthogonal Dimension in Video Evaluation

DGX agent

arXiv:2604.17428v1 Announce Type: new Abstract: As video generation models achieve unprecedented capabilities, the demand for robust video evaluation metrics becomes increasingly critical. Traditional

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

Long-Text-to-Image Generation via Compositional Prompt Decomposition

DGX agent

arXiv:2604.18258v1 Announce Type: new Abstract: While modern text-to-image (T2I) models excel at generating images from intricate prompts, they struggle to capture the key details when the inputs are

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

LongBench: Evaluating Robotic Manipulation Policies on Real-World Long-Horizon Tasks

DGX agent

arXiv:2604.16788v1 Announce Type: new Abstract: Robotic manipulation policies often degrade over extended horizons, yet existing benchmarks provide limited insight into why such failures occur. Most p

model-releasesarxiv-cs-ro
21 Apr 2026
← Previous
1…409410411412413…466
Next →