AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
59,904 results
24 Apr 2026

China’s DeepSeek previews new AI model a year after jolting US rivals

Model ReleasesDGX agent

Chinese AI company DeepSeek released a preview of its hotly anticipated next-generation AI model V4 on Friday, saying that the open-source model can compete with leading closed-source systems from US

GPT-5.5 is now available in Cursor! It's currently the top model on CursorBench at 72.8%. We've partnered with OpenAI to offer it for 50% of…

Model ReleasesDGX agent

Cursor has integrated OpenAI's GPT-5.5 model into its IDE, positioning it as the top performer on CursorBench with a 72.8% score. Through a partnership with OpenAI, the model is being offered at a 50%

On the Relationship between Bayesian Networks and Probabilistic Structural Causal Models

ResearchDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2603.27406v2 Announce Type: replace Abstract: In this paper, the relationship between probabilistic graphical models, in particular Bayesian networks, and causal diagrams, also called structural

We're the #1 open model in Vision and Document Arena!

Model ReleasesDGX agent

We're the #1 open model in Vision and Document Arena! Kimi K2.6 is the new SOTA open model in Vision and Document Arena, with solid gains since Kimi K2.5: - #1 open on Vision Arena (#15 overall), +14

We’ve been using Sakana Fugu internally for our own research and coding. Instead of relying on a single model, it dynamically orchestrates t…

AgentsDGX agent

We’ve been using Sakana Fugu internally for our own research and coding. Instead of relying on a single model, it dynamically orchestrates the best combination of open and closed models for any task.

Why Do Language Model Agents Whistleblow?

SafetyDGX agent

arXiv:2511.17085v3 Announce Type: replace-cross Abstract: The deployment of Large Language Models (LLMs) as tool-using agents causes their alignment training to manifest in new ways. Recent work finds

23 Apr 2026

AI models of unstable flow exhibit hallucination

SafetyDGX agent

arXiv:2604.20372v1 Announce Type: cross Abstract: We report the first systematic evidence of hallucination in AI models of fluid dynamics, demonstrated in the canonical problem of hydrodynamically uns

Cross-Modal Taxonomic Generalization in (Vision-) Language Models

ApplicationsDGX agent

arXiv:2603.07474v2 Announce Type: replace-cross Abstract: What is the interplay between semantic representations learned by language models (LM) from surface form alone to those learned from more grou

Explainability in Generative Medical Diffusion Models: A Faithfulness-Based Analysis on MRI Synthesis

ApplicationsDGX agent

arXiv:2602.09781v2 Announce Type: replace-cross Abstract: This study investigates the explainability of generative diffusion models in the context of medical imaging, focusing on Magnetic resonance im

Local Diffusion Models and Phases of Data Distributions

TutorialsDGX agent

arXiv:2508.06614v2 Announce Type: replace Abstract: As a class of generative artificial intelligence frameworks inspired by statistical physics, diffusion models have shown extraordinary performance i

Membership Inference for Contrastive Pre-training Models with Text-only PII Queries

SafetyDGX agent

arXiv:2603.14222v2 Announce Type: replace-cross Abstract: Contrastive pretraining models such as CLIP and CLAP, serve as the ubiquitous perceptual backbones for modern multimodal large models, yet the

WebGen-R1: Incentivizing Large Language Models to Generate Functional and Aesthetic Websites with Reinforcement Learning

Model ReleasesDGX agent

arXiv:2604.20398v1 Announce Type: new Abstract: While Large Language Models (LLMs) excel at function-level code generation, project-level tasks such as generating functional and visually aesthetic mul

22 Apr 2026

GRASPrune: Global Gating for Budgeted Structured Pruning of Large Language Models

Model ReleasesDGX agent

arXiv:2604.19398v1 Announce Type: new Abstract: Large language models (LLMs) are expensive to serve because model parameters, attention computation, and KV caches impose substantial memory and latency

HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language Models

Model ReleasesDGX agent

arXiv:2604.19300v1 Announce Type: cross Abstract: Large Audio-Language Models (LALMs) have recently achieved strong performance across various audio-centric tasks. However, hallucination, where models

Hybrid Architectures for Language Models: Systematic Analysis and Design Insights

ResearchDGX agent

arXiv:2510.04800v3 Announce Type: replace Abstract: Recent progress in large language models demonstrates that hybrid architectures--combining self-attention mechanisms with structured state space mod

Learning Lifted Action Models from Unsupervised Visual Traces

TutorialsDGX agent

arXiv:2604.19043v1 Announce Type: new Abstract: Efficient construction of models capturing the preconditions and effects of actions is essential for applying AI planning in real-world domains. Extensi

LegalBench-BR: A Benchmark for Evaluating Large Language Models on Brazilian Legal Decision Classification

Model ReleasesDGX agent

arXiv:2604.18878v1 Announce Type: new Abstract: We introduce LegalBench-BR, the first public benchmark for evaluating language models on Brazilian legal text classification. The dataset comprises 3,10

Props to @augmentcode for being one of the first to ship Kimi K2.6 to their users. Open models are now the frontier.

ToolsDGX agent

Props to @augmentcode for being one of the first to ship Kimi K2.6 to their users. Open models are now the frontier. Support for Kimi K2.6 is now here in Augment Code! We teamed up with @Kimi_Moonshot

21 Apr 2026

CreditDecoding: Accelerating Parallel Decoding in Diffusion Large Language Models with Trace Credit

Model ReleasesDGX agent

arXiv:2510.06133v2 Announce Type: replace Abstract: Diffusion large language models (dLLMs) generate text through iterative denoising. In commonly adopted parallel decoding schemes, each step confirms

CSF: Black-box Fingerprinting via Compositional Semantics for Text-to-Image Models

ResearchDGX agent

arXiv:2604.16363v1 Announce Type: cross Abstract: Text-to-image models are commercially valuable assets often distributed under restrictive licenses, but such licenses are enforceable only when violat

D-QRELO: Training- and Data-Free Delta Compression for Large Language Models via Quantization and Residual Low-Rank Approximation

Model ReleasesDGX agent

arXiv:2604.16940v1 Announce Type: new Abstract: Supervised Fine-Tuning (SFT) accelerates taskspecific large language models (LLMs) development, but the resulting proliferation of finetuned models incu

DGSSM: Diffusion guided state-space models for multimodal salient object detection

ResearchDGX agent

arXiv:2604.17585v1 Announce Type: new Abstract: Salient object detection (SOD) requires modeling both long-range contextual dependencies and fine-grained structural details, which remains challenging

Evaluating Multimodal LLMs for Inpatient Diagnosis: Real-World Performance, Safety, and Cost Across Ten Frontier Models

Model ReleasesDGX agent

arXiv:2604.16980v1 Announce Type: new Abstract: Background: Large language models (LLMs) are increasingly proposed for diagnostic support, but few evaluations use real-world multimodal inpatient data,

Generalization Boundaries of Fine-Tuned Small Language Models for Graph Structural Inference

Model ReleasesDGX agent

arXiv:2604.18092v1 Announce Type: new Abstract: Small language models fine-tuned for graph property estimation have demonstrated strong in-distribution performance, yet their generalization capabiliti

MultiWorld: Scalable Multi-Agent Multi-View Video World Models

AgentsDGX agent

arXiv:2604.18564v1 Announce Type: new Abstract: Video world models have achieved remarkable success in simulating environmental dynamics in response to actions by users or agents. They are modeled as

New Fourth-Order Grayscale Indicator-Based Telegraph Diffusion Model for Image Despeckling

ResearchDGX agent

arXiv:2509.26010v2 Announce Type: replace Abstract: Second-order PDE models have been widely used for suppressing multiplicative noise, but they often introduce blocky artifacts in the early stages of

OmniVLA-RL: A Vision-Language-Action Model with Spatial Understanding and Online RL

SafetyDGX agent

arXiv:2604.17706v1 Announce Type: new Abstract: Visual-Language-Action (VLA) models represent a paradigm shift in embodied AI, yet existing frameworks often struggle with imprecise spatial perception,

PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling

SafetyDGX agent

arXiv:2510.24235v3 Announce Type: replace Abstract: Reward models (RMs) are central to reinforcement learning from human feedback (RLHF), providing the critical supervision signals that align large la

Please don’t trust your chatbot for medical advice. 🙏 Remember how I used to say that large language models are “frequently wrong, never in…

Model ReleasesDGX agent

Please don’t trust your chatbot for medical advice. 🙏 Remember how I used to say that large language models are “frequently wrong, never in doubt”, and how I warned three years ago on 60 Minutes that

ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction

SafetyDGX agent

arXiv:2506.01770v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have achieved tremendous success in various tasks, yet concerns about their safety and security have emerged. In

Rethinking Post-Unlearning Behavior of Large Vision-Language Models

ResearchDGX agent

arXiv:2506.02541v2 Announce Type: replace-cross Abstract: Large Vision-Language Models (LVLMs) can recognize individuals in images and disclose sensitive personal information about them, raising criti

Robust Bias Evaluation with FilBBQ: A Filipino Bias Benchmark for Question-Answering Language Models

Model ReleasesDGX agent

arXiv:2602.14466v2 Announce Type: replace Abstract: With natural language generation becoming a popular use case for language models, the Bias Benchmark for Question-Answering (BBQ) has grown to be an

Sonata: A Hybrid World Model for Inertial Kinematics under Clinical Data Scarcity

Model ReleasesDGX agent

arXiv:2604.18058v1 Announce Type: new Abstract: We introduce Sonata, a compact latent world model for six-axis trunk IMU representation learning under clinical data scarcity. Clinical cohorts typicall

(Sparse) Attention to the Details: Preserving Spectral Fidelity in ML-based Weather Forecasting Models

SafetyDGX agent

arXiv:2604.16429v1 Announce Type: cross Abstract: We introduce Mosaic, a probabilistic weather forecasting model that addresses two principal sources of spectral degradation in ML-based weather predic

UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models

Model ReleasesDGX agent

arXiv:2604.18518v1 Announce Type: new Abstract: Uniform Discrete Diffusion Model (UDM) has recently emerged as a promising paradigm for discrete generative modeling; however, its integration with rein

When Pretty Isn't Useful: Investigating Why Modern Text-to-Image Models Fail as Reliable Training Data Generators

Model ReleasesDGX agent

arXiv:2602.19946v4 Announce Type: replace Abstract: Recent text-to-image (T2I) diffusion models produce visually stunning images and demonstrate excellent prompt following. But do they perform well as

20 Apr 2026

ECG-Lens: Benchmarking ML & DL Models on PTB-XL Dataset

Model ReleasesDGX agent

arXiv:2604.15822v1 Announce Type: cross Abstract: Automated classification of electrocardiogram (ECG) signals is a useful tool for diagnosing and monitoring cardiovascular diseases. This study compare

Efficient Video Diffusion Models: Advancements and Challenges

ApplicationsDGX agent

arXiv:2604.15911v1 Announce Type: new Abstract: Video diffusion models have rapidly become the dominant paradigm for high-fidelity generative video synthesis, but their practical deployment remains co

Exascale Multi-Task Graph Foundation Models for Imbalanced, Multi-Fidelity Atomistic Data

ResearchDGX agent

arXiv:2604.15380v1 Announce Type: cross Abstract: We present an exascale workflow for materials discovery using atomistic graph foundation models built on HydraGNN. We jointly train on 16 open first-p

Kimi K2.6 is now available in Hermes Agent. Simply run `hermes update` and use `hermes model` to select a compatible provider hosting the mo…

AgentsDGX agent

Kimi K2.6 is now available in Hermes Agent. Simply run `hermes update` and use `hermes model` to select a compatible provider hosting the model! Meet Kimi K2.6: Advancing Open-Source Coding 🔹Open-sour

neuralCAD-Edit: An Expert Benchmark for Multimodal-Instructed 3D CAD Model Editing

Model ReleasesDGX agent

arXiv:2604.16170v1 Announce Type: new Abstract: We introduce neuralCAD-Edit, the first benchmark for editing 3D CAD models collected from expert CAD engineers. Instead of text conditioning as in prior

TabularMath: Understanding Math Reasoning over Tables with Large Language Models

Model ReleasesDGX agent

arXiv:2505.19563v4 Announce Type: replace Abstract: Mathematical reasoning has long been a key benchmark for evaluating large language models. Although substantial progress has been made on math word

TTL: Test-time Textual Learning for OOD Detection with Pretrained Vision-Language Models

ResearchDGX agent

arXiv:2604.15756v1 Announce Type: new Abstract: Vision-language models (VLMs) such as CLIP exhibit strong Out-of-distribution (OOD) detection capabilities by aligning visual and textual representation

vla-eval: A Unified Evaluation Harness for Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2603.13966v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models are increasingly evaluated across multiple simulation benchmarks, yet adding each benchmark to an evaluation pip

19 Apr 2026

The continuing gap between the capabilities of Gemini Pro 3.1 (very good model) and the capabilities of the Gemini app/website is odd. The m…

Model ReleasesDGX agent

The continuing gap between the capabilities of Gemini Pro 3.1 (very good model) and the capabilities of the Gemini app/website is odd. The model can do what Claude/GPT can do, but there is a minimal h

17 Apr 2026

AutoRAN: Automated Hijacking of Safety Reasoning in Large Reasoning Models

Model ReleasesDGX agent

arXiv:2505.10846v3 Announce Type: replace Abstract: This paper presents AutoRAN, the first framework to automate the hijacking of internal safety reasoning in large reasoning models (LRMs). At its cor

Compressed-Sensing-Guided, Inference-Aware Structured Reduction for Large Language Models

Model ReleasesDGX agent

arXiv:2604.14156v1 Announce Type: new Abstract: Large language models deliver strong generative performance but at the cost of massive parameter counts, memory use, and decoding latency. Prior work ha

Finetuning-Free Diffusion Model with Adaptive Constraint Guidance for Inorganic Crystal Structure Generation

ResearchDGX agent

arXiv:2604.13354v1 Announce Type: cross Abstract: The discovery of inorganic crystal structures with targeted properties is a significant challenge in materials science. Generative models, especially

From Black Box to Glass Box: Cross-Model ASR Disagreement to Prioto Review in Ambient AI Scribe Documentation

ApplicationsDGX agent

arXiv:2604.14152v1 Announce Type: cross Abstract: Ambient AI 'scribe' systems promise to reduce clinical documentation burden, but automatic speech recognition (ASR) errors can remain unnoticed withou

POP: Prefill-Only Pruning for Efficient Large Model Inference

Model ReleasesDGX agent

arXiv:2602.03295v2 Announce Type: replace Abstract: Large Language Models (LLMs) and Vision-Language Models (VLMs) have demonstrated remarkable capabilities. However, their deployment is hindered by s

ReSS: Learning Reasoning Models for Tabular Data Prediction via Symbolic Scaffold

TutorialsDGX agent

arXiv:2604.13392v1 Announce Type: new Abstract: Tabular data remains prevalent in high-stakes domains such as healthcare and finance, where predictive models are expected to provide both high accuracy

Robustness of Vision Foundation Models to Common Perturbations

ResearchDGX agent

arXiv:2604.14973v1 Announce Type: cross Abstract: A vision foundation model outputs an embedding vector for an image, which can be affected by common editing operations (e.g., JPEG compression, bright

16 Apr 2026

Data-driven Learning of Probabilistic Model of Binary Droplet Collision for Spray Simulation

Model ReleasesDGX agent

arXiv:2604.13594v1 Announce Type: cross Abstract: Binary droplet collisions are ubiquitous in dense sprays. Traditional deterministic models cannot adequately represent transitional and stochastic beh

Gaslight, Gatekeep, V1-V3: Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation

Model ReleasesDGX agent

arXiv:2604.13803v1 Announce Type: new Abstract: Vision-language models are increasingly deployed in high-stakes settings, yet their susceptibility to sycophantic manipulation remains poorly understood

Guided Transfer Learning for Discrete Diffusion Models

ApplicationsDGX agent

arXiv:2512.10877v4 Announce Type: replace Abstract: Discrete diffusion models (DMs) have achieved strong performance in language and other discrete domains, offering a compelling alternative to autore

Hydra: Unifying Document Retrieval and Generation in a Single Vision-Language Model

HardwareDGX agent

arXiv:2603.28554v2 Announce Type: replace Abstract: Visual document understanding typically requires separate retrieval and generation models, doubling memory and system complexity. We present Hydra,

Scaling Test-Time Compute to Achieve IOI Gold Medal with Open-Weight Models

Model ReleasesDGX agent

arXiv:2510.14232v2 Announce Type: replace-cross Abstract: Competitive programming has become a rigorous benchmark for evaluating the reasoning and problem-solving capabilities of large language models

15 Apr 2026

Are Video Reasoning Models Ready to Go Outside?

Model ReleasesDGX agent

arXiv:2603.10652v2 Announce Type: replace-cross Abstract: In real-world deployment, vision-language models often encounter disturbances such as weather, occlusion, and camera motion. Under such condit

GlotOCR Bench: OCR Models Still Struggle Beyond a Handful of Unicode Scripts

Model ReleasesDGX agent

arXiv:2604.12978v1 Announce Type: new Abstract: Optical character recognition (OCR) has advanced rapidly with the rise of vision-language models, yet evaluation has remained concentrated on a small cl

MAML-KT: Addressing Cold Start Problem in Knowledge Tracing for New Students via Few-Shot Model-Agnostic Meta Learning

ResearchDGX agent

arXiv:2603.00137v2 Announce Type: replace-cross Abstract: Knowledge tracing (KT) models are commonly evaluated by training on early interactions from all students and testing on later responses. While

← Previous
1…4243444546…999
Next →