AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,323 results
21 Apr 2026

Language Models Don't Know What You Want: Evaluating Personalization in Deep Research Needs Real Users

Model ReleasesDGX agent

arXiv:2603.16120v2 Announce Type: replace Abstract: Deep Research (DR) systems help researchers cope with ballooning publishing counts. Such tools synthesize scientific papers to answer research queri

Large Language Models Are Still Misled by Simple Bias Ensembles

Model ReleasesDGX agent

arXiv:2505.16522v3 Announce Type: replace Abstract: With the evolution of large language models (LLMs), their robustness against individual simple biases has been enhanced. However, we observe that th

Late Fusion Neural Operators for Extrapolation Across Parameter Space in Partial Differential Equations

Model Releases

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2604.16721v1 Announce Type: new Abstract: Developing neural operators that accurately predict the behavior of systems governed by partial differential equations (PDEs) across unseen parameter re

Latent Preference Modeling for Cross-Session Personalized Tool Calling

Model ReleasesDGX agent

arXiv:2604.17886v1 Announce Type: new Abstract: Users often omit essential details in their requests to LLM-based agents, resulting in under-specified inputs for tool use. This poses a fundamental cha

LayerCache: Exploiting Layer-wise Velocity Heterogeneity for Efficient Flow Matching Inference

Model ReleasesDGX agent

arXiv:2604.16492v1 Announce Type: new Abstract: Flow Matching models achieve state-of-the-art image generation quality but incur substantial inference cost due to iterative denoising through large Tra

LEAF: Knowledge Distillation of Text Embedding Models with Teacher-Aligned Representations

Model ReleasesDGX agent

arXiv:2509.12539v2 Announce Type: replace-cross Abstract: We present LEAF ('Lightweight Embedding Alignment Framework'), a knowledge distillation framework for text embedding models. A key distinguish

Learned Nonlocal Feature Matching and Filtering for RAW Image Denoising

Model ReleasesDGX agent

arXiv:2604.17453v1 Announce Type: cross Abstract: Being one of the oldest and most basic problems in image processing, image denoising has seen a resurgence spurred by rapid advances in deep learning.

Learning Stable Predictors from Weak Supervision under Distribution Shift

Model ReleasesDGX agent

arXiv:2604.05002v2 Announce Type: replace Abstract: Learning from weak, proxy, or relative supervision is common when ground-truth labels are unavailable, but robustness under distribution shift remai

Learning to Control Summaries with Score Ranking

Model ReleasesDGX agent

arXiv:2604.17197v1 Announce Type: new Abstract: Recent advances in summarization research focus on improving summary quality across multiple criteria, such as completeness, conciseness, and faithfulne

Learning to Retrieve User History and Generate User Profiles for Personalized Persuasiveness Prediction

Model ReleasesDGX agent

arXiv:2601.05654v3 Announce Type: replace Abstract: Estimating the persuasiveness of messages is critical in various applications, from recommender systems to safety assessment of LLMs. While it is im

Let's talk parsing charts 📊📈. Last week we released ParseBench, the first document OCR benchmark for AI agents. New in ParseBench: ChartDa…

Model ReleasesDGX agent

Let's talk parsing charts 📊📈. Last week we released ParseBench, the first document OCR benchmark for AI agents. New in ParseBench: ChartDataPointMatch. Most document look at a chart and OCR the captio

Leveraging Large Language Models for Sarcastic Speech Annotation in Sarcasm Detection

Model ReleasesDGX agent

arXiv:2506.00955v2 Announce Type: replace Abstract: Sarcasm fundamentally alters meaning through tone and context, yet detecting it in speech remains a challenge due to data scarcity. In addition, exi

LexRel: Benchmarking Legal Relation Extraction for Chinese Civil Cases

Model ReleasesDGX agent

arXiv:2512.12643v2 Announce Type: replace Abstract: Legal relations serve as an important analytical framework for dispute resolution in civil cases. However, legal relations in Chinese civil cases re

LiFT: Does Instruction Fine-Tuning Improve In-Context Learning for Longitudinal Modelling by Large Language Models?

Model ReleasesDGX agent

arXiv:2604.16382v1 Announce Type: new Abstract: Longitudinal NLP tasks require reasoning over temporally ordered text to detect persistence and change in human behavior and opinions. However, in-conte

LIFT the Veil for the Truth: Principal Weights Emerge after Rank Reduction for Reasoning-Focused Supervised Fine-Tuning

Model ReleasesDGX agent

arXiv:2506.00772v2 Announce Type: replace-cross Abstract: Recent studies have shown that supervised fine-tuning of LLMs on a small number of high-quality datasets can yield strong reasoning capabiliti

LiquidTAD: An Efficient Method for Temporal Action Detection via Liquid Neural Dynamics

Model ReleasesDGX agent

arXiv:2604.18274v1 Announce Type: new Abstract: Temporal Action Detection (TAD) in untrimmed videos is currently dominated by Transformer-based architectures. While high-performing, their quadratic co

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length

Model ReleasesDGX agent

arXiv:2512.04677v5 Announce Type: replace Abstract: Audio-driven avatar interaction demands real-time, streaming, and infinite-length generation -- capabilities fundamentally at odds with the sequenti

LIVE: Leveraging Image Manipulation Priors for Instruction-based Video Editing

Model ReleasesDGX agent

arXiv:2604.17021v1 Announce Type: new Abstract: Video editing aims to modify input videos according to user intent. Recently, end-to-end training methods have garnered widespread attention, constructi

LiveFact: A Dynamic, Time-Aware Benchmark for LLM-Driven Fake News Detection

Model ReleasesDGX agent

arXiv:2604.04815v2 Announce Type: replace Abstract: The rapid development of Large Language Models (LLMs) has transformed fake news detection and fact-checking tasks from simple classification to comp

Lizard: An Efficient Linearization Framework for Large Language Models

Model ReleasesDGX agent

arXiv:2507.09025v4 Announce Type: replace Abstract: We propose Lizard, a linearization framework that transforms pretrained Transformer-based Large Language Models (LLMs) into subquadratic architectur

LLaMA-XR: A Novel Framework for Radiology Report Generation using LLaMA and QLoRA Fine Tuning

Model ReleasesDGX agent

arXiv:2506.03178v2 Announce Type: replace-cross Abstract: Automated radiology report generation holds significant potential to reduce radiologists' workload and enhance diagnostic accuracy. However, g

LLMs are still not consistent judges of qualitative work, and small changes to how that work is presented affect outcomes. Better harnessing…

Model ReleasesDGX agent

LLMs are still not consistent judges of qualitative work, and small changes to how that work is presented affect outcomes. Better harnessing and methods (multiple judging runs with randomized orders,

LOGICAL-COMMONSENSEQA: A Benchmark for Logical Commonsense Reasoning

Model ReleasesDGX agent

arXiv:2601.16504v3 Announce Type: replace Abstract: Commonsense reasoning often involves evaluating multiple plausible interpretations rather than selecting a single atomic answer, yet most benchmarks

Logit Arithmetic Elicits Long Reasoning Capabilities Without Training

Model ReleasesDGX agent

arXiv:2510.09354v2 Announce Type: replace Abstract: Large reasoning models exhibit long chain-of-thought reasoning with complex strategies such as backtracking and self-verification. Yet, these capabi

Long-CODE: Isolating Pure Long-Context as an Orthogonal Dimension in Video Evaluation

Model ReleasesDGX agent

arXiv:2604.17428v1 Announce Type: new Abstract: As video generation models achieve unprecedented capabilities, the demand for robust video evaluation metrics becomes increasingly critical. Traditional

Long-Text-to-Image Generation via Compositional Prompt Decomposition

Model ReleasesDGX agent

arXiv:2604.18258v1 Announce Type: new Abstract: While modern text-to-image (T2I) models excel at generating images from intricate prompts, they struggle to capture the key details when the inputs are

LongBench: Evaluating Robotic Manipulation Policies on Real-World Long-Horizon Tasks

Model ReleasesDGX agent

arXiv:2604.16788v1 Announce Type: new Abstract: Robotic manipulation policies often degrade over extended horizons, yet existing benchmarks provide limited insight into why such failures occur. Most p

LoRA on the Go: Instance-level Dynamic LoRA Selection and Merging

Model ReleasesDGX agent

arXiv:2511.07129v3 Announce Type: replace Abstract: Low-Rank Adaptation (LoRA) has emerged as a parameter-efficient approach for fine-tuning large language models. However, conventional LoRA adapters

Love this work from Aksel and the post-training team at Hugging Face! Turns out the HF ecosystem (papers, datasets, models all accessible th…

Model ReleasesDGX agent

Love this work from Aksel and the post-training team at Hugging Face! Turns out the HF ecosystem (papers, datasets, models all accessible through CLI, skills and md files) is perfect for running SOTA

Low-rank Orthogonalization for Large-scale Matrix Optimization with Applications to Foundation Model Training

Model ReleasesDGX agent

arXiv:2509.11983v2 Announce Type: replace Abstract: Neural network (NN) training is inherently a large-scale matrix optimization problem, yet the matrix structure of NN parameters has long been overlo

ltzGLUE: Luxembourgish General Language Understanding Evaluation

Model ReleasesDGX agent

arXiv:2604.17976v1 Announce Type: new Abstract: This paper presents ltzGLUE, the first Natural Language Understanding (NLU) benchmark for Luxembourgish (LTZ) based on the popular GLUE benchmark for En

Lumos3D: A Single-Forward Framework for Low-Light 3D Scene Restoration

Model ReleasesDGX agent

arXiv:2511.09818v2 Announce Type: replace Abstract: Restoring 3D scenes with low-light conditions is challenging, and most existing methods depend on precomputed camera poses and scene-specific optimi

M100: An Orchestrated Dataflow Architecture Powering General AI Computing

Model ReleasesDGX agent

arXiv:2604.17862v1 Announce Type: new Abstract: As deep learning-based AI technologies gain momentum, the demand for general-purpose AI computing architectures continues to grow. While GPGPU-based arc

Macaron: Controlled, Human-Written Benchmark for Multilingual and Multicultural Reasoning via Template-Filling

Model ReleasesDGX agent

arXiv:2602.10732v2 Announce Type: replace Abstract: Multilingual benchmarks rarely test reasoning over culturally grounded premises: translated datasets keep English-centric scenarios, while culture-f

Machine Learning Hamiltonian Dynamical Systems with Sparse and Noisy Data

Model ReleasesDGX agent

arXiv:2604.17470v1 Announce Type: new Abstract: Machine learning has become a powerful tool for discovering governing laws of dynamical systems from data. However, most existing approaches degrade sev

MARCO: Navigating the Unseen Space of Semantic Correspondence

Model ReleasesDGX agent

arXiv:2604.18267v1 Announce Type: new Abstract: Recent advances in semantic correspondence rely on dual-encoder architectures, combining DINOv2 with diffusion backbones. While accurate, these billion-

Marrying Text-to-Motion Generation with Skeleton-Based Action Recognition

Model ReleasesDGX agent

arXiv:2604.17090v1 Announce Type: new Abstract: Human action recognition and motion generation are two active research problems in human-centric computer vision, both aiming to align motion with textu

MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems

Model ReleasesDGX agent

arXiv:2503.16549v2 Announce Type: replace Abstract: Despite strong results on many tasks, multimodal large language models (MLLMs) still underperform on visual mathematical problem solving, especially

MathNet: a Global Multimodal Benchmark for Mathematical Reasoning and Retrieval

Model ReleasesDGX agent

arXiv:2604.18584v1 Announce Type: cross Abstract: Mathematical problem solving remains a challenging test of reasoning for large language and multimodal models, yet existing benchmarks are limited in

Maximizing Local Entropy Where It Matters: Prefix-Aware Localized LLM Unlearning

Model ReleasesDGX agent

arXiv:2601.03190v3 Announce Type: replace Abstract: Machine unlearning aims to forget sensitive knowledge from Large Language Models (LLMs) while maintaining general utility. However, existing approac

MeasHalu: Mitigation of Scientific Measurement Hallucinations for Large Language Models with Enhanced Reasoning

Model ReleasesDGX agent

arXiv:2604.16929v1 Announce Type: new Abstract: The accurate extraction of scientific measurements from literature is a critical yet challenging task in AI4Science, enabling large-scale analysis and i

Measuring Representation Robustness in Large Language Models for Geometry

Model ReleasesDGX agent

arXiv:2604.16421v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly evaluated on mathematical reasoning, yet their robustness to equivalent problem representations remains po

Measuring Social Bias in Vision-Language Models with Face-Only Counterfactuals from Real Photos

Model ReleasesDGX agent

arXiv:2601.06931v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) are increasingly deployed in socially consequential settings, raising concerns about social bias driven by demog

Medical Image Understanding Improves Survival Prediction via Visual Instruction Tuning

Model ReleasesDGX agent

arXiv:2604.18250v1 Announce Type: new Abstract: Accurate prognostication and risk estimation are essential for guiding clinical decision-making and optimizing patient management. While radiologist-ass

Medical thinking with multiple images

Model ReleasesDGX agent

arXiv:2604.16506v1 Announce Type: cross Abstract: Large language models perform well on many medical QA benchmarks, but real clinical reasoning often requires integrating evidence across multiple imag

MEDN: Motion-Emotion Feature Decoupling Network for Micro-Expression Recognition

Model ReleasesDGX agent

arXiv:2604.17899v1 Announce Type: new Abstract: Unlike macro-expression, micro-expression does not follow a strictly consistent mapping rule between emotions and Action Units (AUs). As a result, some

MedPRMBench: A Fine-grained Benchmark for Process Reward Models in Medical Reasoning

Model ReleasesDGX agent

arXiv:2604.17282v1 Announce Type: new Abstract: Process-Level Reward Models (PRMs) are essential for guiding complex reasoning in large language models, yet existing PRM benchmarks cover only general

MedProbeBench: Systematic Benchmarking at Deep Evidence Integration for Expert-level Medical Guideline

Model ReleasesDGX agent

arXiv:2604.18418v1 Announce Type: new Abstract: Recent advances in deep research systems enable large language models to retrieve, synthesize, and reason over large-scale external knowledge. In medici

MedRedFlag: Investigating how LLMs Redirect Misconceptions in Real-World Health Communication

Model ReleasesDGX agent

arXiv:2601.09853v2 Announce Type: replace Abstract: Real-world health questions from patients often unintentionally embed false assumptions or premises. In such cases, safe medical communication typic

MemBuilder: Reinforcing LLMs for Long-Term Memory Construction via Attributed Dense Rewards

Model ReleasesDGX agent

arXiv:2601.05488v3 Announce Type: replace Abstract: Maintaining consistency in long-term dialogues remains a fundamental challenge for LLMs, as standard retrieval mechanisms often fail to capture the

mEOL: Training-Free Instruction-Guided Multimodal Embedder for Vector Graphics and Image Retrieval

Model ReleasesDGX agent

arXiv:2604.17054v1 Announce Type: new Abstract: Scalable Vector Graphics (SVGs) function both as visual images and as structured code that encode rich geometric and layout information, yet most method

MerLin: A Discovery Engine for Photonic and Hybrid Quantum Machine Learning

Model ReleasesDGX agent

arXiv:2602.11092v2 Announce Type: replace Abstract: Identifying where quantum models may offer practical benefits in near term quantum machine learning (QML) requires moving beyond isolated algorithmi

MeSH: Memory-as-State-Highways for Recursive Transformers

Model ReleasesDGX agent

arXiv:2510.07739v2 Announce Type: replace Abstract: Recursive transformers reuse parameters and iterate over hidden states multiple times, decoupling compute depth from parameter depth. However, under

MetaLint: Easy-to-Hard Generalization for Code Linting

Model ReleasesDGX agent

arXiv:2507.11687v4 Announce Type: replace-cross Abstract: Large language models excel at code generation but struggle with code linting, particularly in generalizing to unseen or evolving best practic

Method for Aggregating Unstructured Data Using Large Language Models

Model ReleasesDGX agent

arXiv:2604.16425v1 Announce Type: cross Abstract: This paper presents a method for the automated collection and aggregation of unstructured data from diverse web sources, utilizing Large Language Mode

Mind the Way You Select Negative Texts: Pursuing the Distance Consistency in OOD Detection with VLMs

Model ReleasesDGX agent

arXiv:2603.02618v3 Announce Type: replace Abstract: Out-of-distribution (OOD) detection seeks to identify samples from unknown classes, a critical capability for deploying machine learning models in o

Missing-by-Design: Certifiable Modality Deletion for Revocable Multimodal Sentiment Analysis

Model ReleasesDGX agent

arXiv:2602.16144v3 Announce Type: replace Abstract: As multimodal systems increasingly process sensitive personal data, the ability to selectively revoke specific data modalities has become a critical

Missing Pattern Tree based Decision Grouping and Ensemble for Enhancing Pair Utilization in Deep Incomplete Multi-View Clustering

Model ReleasesDGX agent

arXiv:2512.21510v2 Announce Type: replace-cross Abstract: Real-world multi-view data often exhibit highly inconsistent missing patterns, posing significant challenges for incomplete multi-view cluster

MM-JudgeBias: A Benchmark for Evaluating Compositional Biases in MLLM-as-a-Judge

Model ReleasesDGX agent

arXiv:2604.18164v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have been increasingly used as automatic evaluators-a paradigm known as MLLM-as-a-Judge. However, their reliabi

MMErroR: A Benchmark for Erroneous Reasoning in Vision-Language Models

Model ReleasesDGX agent

arXiv:2601.03331v2 Announce Type: replace Abstract: Recent advances in Vision-Language Models (VLMs) have improved performance in multi-modal learning, raising the question of whether these models tru

← Previous
1…328329330331332…373
Next →