AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,323 results
21 Apr 2026

From Handwriting to Structured Data: Benchmarking AI Digitisation of Handwritten Forms

Model ReleasesDGX agent

arXiv:2604.16504v1 Announce Type: new Abstract: Manual digitisation of structured handwritten documents is slow and costly. We benchmark 17 leading frontier multi-modal large language models and open-

From Heads to Neurons: Causal Attribution and Steering in Multi-Task Vision-Language Models

Model ReleasesDGX agent

arXiv:2604.17941v1 Announce Type: cross Abstract: Recent work has increasingly explored neuron-level interpretation in vision-language models (VLMs) to identify neurons critical to final predictions.

From Inheritance to Saturation: Disentangling the Evolution of Visual Redundancy for Architecture-Aware MLLM Inference Acceleration


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases
DGX agent

arXiv:2604.16462v1 Announce Type: new Abstract: High-resolution Multimodal Large Language Models (MLLMs) face prohibitive computational costs during inference due to the explosion of visual tokens. Ex

From keynote to the terminal: Join our Next ‘26 developer livestreams

Model ReleasesDGX agent

The main stage at Google Cloud Next is where the vision is set. This year, we’re bridging the gap between those massive 'Cloud-scale' announcements and your local terminal. We are thrilled to announce

From log pi to pi: Taming Divergence in Soft Clipping via Bilateral Decoupled Decay of Probability Gradient Weight

Model ReleasesDGX agent

arXiv:2603.14389v2 Announce Type: replace Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has catalyzed a leap in Large Language Model (LLM) reasoning, yet its optimization dynamics re

Functional Similarity Metric for Neural Networks: Overcoming Parametric Ambiguity via Activation Region Analysis

Model ReleasesDGX agent

arXiv:2604.16426v1 Announce Type: new Abstract: As modern deep learning architectures grow in complexity, representational ambiguity emerges as a critical barrier to their interpretability and reliabl

Fuzzy Encoding-Decoding to Improve Spiking Q-Learning Performance in Autonomous Driving

Model ReleasesDGX agent

arXiv:2604.16436v1 Announce Type: cross Abstract: This paper develops an end-to-end fuzzy encoder-decoder architecture for enhancing vision-based multi-modal deep spiking Q-networks in autonomous driv

Generalizable Face Forgery Detection via Separable Prompt Learning

Model ReleasesDGX agent

arXiv:2604.17307v1 Announce Type: new Abstract: Detecting face forgeries using CLIP has recently emerged as a promising and increasingly popular research direction. Owing to its rich visual knowledge

Generalization Boundaries of Fine-Tuned Small Language Models for Graph Structural Inference

Model ReleasesDGX agent

arXiv:2604.18092v1 Announce Type: new Abstract: Small language models fine-tuned for graph property estimation have demonstrated strong in-distribution performance, yet their generalization capabiliti

GeoRC: A Benchmark for Geolocation Reasoning Chains

Model ReleasesDGX agent

arXiv:2601.21278v2 Announce Type: replace-cross Abstract: Vision Language Models (VLMs) are good at recognizing the global location of a photograph -- their geolocation prediction accuracy rivals the

Google Cloud Next 2026 preview: The real story isn’t AI — it’s the control plane

Model ReleasesDGX agent

Everyone heading into Google Cloud Next this week is bracing for another wave of artificial intelligence announcements. More Gemini. More agents. More benchmarks. More onstage demos that look great in

Google's Chief AI Architect Koray Kavukcuoglu is working to unite its internal AI coding tools under the Antigravity platform, to counter Claude Code and Codex (Julia Love/Bloomberg)

Model ReleasesDGX agent

Julia Love / Bloomberg: Google's Chief AI Architect Koray Kavukcuoglu is working to unite its internal AI coding tools under the Antigravity platform, to counter Claude Code and Codex — At Google, lea

GR4CIL: Gap-compensated Routing for CLIP-based Class Incremental Learning

Model ReleasesDGX agent

arXiv:2604.17822v1 Announce Type: new Abstract: Class-Incremental Learning (CIL) aims to continuously acquire new categories while preserving previously learned knowledge. Recently, Contrastive Langua

GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling

Model ReleasesDGX agent

arXiv:2604.18556v1 Announce Type: new Abstract: Weight quantization has become a standard tool for efficient LLM deployment, especially for local inference, where models are now routinely served at 2-

Guardrails in Logit Space: Safety Token Regularization for LLM Alignment

Model ReleasesDGX agent

arXiv:2604.17210v1 Announce Type: new Abstract: Fine-tuning well-aligned large language models (LLMs) on new domains often degrades their safety alignment, even when using benign datasets. Existing sa

HalluSAE: Detecting Hallucinations in Large Language Models via Sparse Auto-Encoders

Model ReleasesDGX agent

arXiv:2604.16430v1 Announce Type: new Abstract: Large Language Models (LLMs) are powerful and widely adopted, but their practical impact is limited by the well-known hallucination phenomenon. While re

Harness as an Asset: Enforcing Determinism via the Convergent AI Agent Framework (CAAF)

Model ReleasesDGX agent

arXiv:2604.17025v1 Announce Type: cross Abstract: Large Language Models (LLMs) produce a controllability gap in safety-critical engineering: even low rates of undetected constraint violations render a

Having a 20-agent system is a million times more powerful than 100 agents working in silos. Last night, I stayed up way too late fixing and …

Model ReleasesDGX agent

Having a 20-agent system is a million times more powerful than 100 agents working in silos. Last night, I stayed up way too late fixing and improving my digital workforce for my AI Agent Mastermind pr

HiGMem: A Hierarchical and LLM-Guided Memory System for Long-Term Conversational Agents

Model ReleasesDGX agent

arXiv:2604.18349v1 Announce Type: new Abstract: Long-term conversational large language model (LLM) agents require memory systems that can recover relevant evidence from historical interactions withou

HiP-LoRA: Budgeted Spectral Plasticity for Robust Low-Rank Adaptation

Model ReleasesDGX agent

arXiv:2604.17751v1 Announce Type: cross Abstract: Adapting foundation models under resource budgets relies heavily on Parameter-Efficient Fine-Tuning (PEFT), with LoRA being a standard modular solutio

HiRAS: A Hierarchical Multi-Agent Framework for Paper-to-Code Generation and Execution

Model ReleasesDGX agent

arXiv:2604.17745v1 Announce Type: new Abstract: Recent advances in large language models have highlighted their potential to automate computational research, particularly reproducing experimental resu

HORIZON: A Benchmark for In-the-wild User Behaviour Modeling

Model ReleasesDGX agent

arXiv:2604.17259v1 Announce Type: cross Abstract: User behavior in the real world is diverse, cross-domain, and spans long time horizons. Existing user modeling benchmarks however remain narrow, focus

HorizonBench: Long-Horizon Personalization with Evolving Preferences

Model ReleasesDGX agent

arXiv:2604.17283v1 Announce Type: new Abstract: User preferences evolve across months of interaction, and tracking them requires inferring when a stated preference has been changed by a subsequent lif

How Robustly do LLMs Understand Execution Semantics?

Model ReleasesDGX agent

arXiv:2604.16320v1 Announce Type: cross Abstract: LLMs demonstrate remarkable reasoning capabilities, yet whether they utilize internal world models or rely on sophisticated pattern matching remains o

How Should We Enhance the Safety of Large Reasoning Models: An Empirical Study

Model ReleasesDGX agent

arXiv:2505.15404v2 Announce Type: replace Abstract: Large Reasoning Models (LRMs) have achieved remarkable success on reasoning-intensive tasks such as mathematics and programming. However, their enha

HQA-VLAttack: Towards High Quality Adversarial Attack on Vision-Language Pre-Trained Models

Model ReleasesDGX agent

arXiv:2604.16499v1 Announce Type: new Abstract: Black-box adversarial attack on vision-language pre-trained models is a practical and challenging task, as text and image perturbations need to be consi

HUGGING FACE JUST AUTOMATED THEIR ENTIRE POST-TRAINING TEAM WITH AN AGENT. It reads papers, runs GPU experiments, iterates, and builds resea…

Model ReleasesDGX agent

HUGGING FACE JUST AUTOMATED THEIR ENTIRE POST-TRAINING TEAM WITH AN AGENT. It reads papers, runs GPU experiments, iterates, and builds research-backed models autonomously. Pushed a benchmark from 10%

Hybrid-Vector Retrieval for Visually Rich Documents: Combining Single-Vector Efficiency and Multi-Vector Accuracy

Model ReleasesDGX agent

arXiv:2510.22215v2 Announce Type: replace-cross Abstract: Retrieval over visually rich documents is essential for tasks such as legal discovery, scientific search, and enterprise knowledge management.

I agree, and I will add to this. Companies should own their memories, skills and other resources agents need. Use Claude code, codex, copilo…

Model ReleasesDGX agent

I agree, and I will add to this. Companies should own their memories, skills and other resources agents need. Use Claude code, codex, copilot, etc for intelligence but your execution layer should plug

I came up with a somewhat foolish new benchmark for testing image generation models, to exercise the new ChatGPT Images 2.0: 'Do a where's W…

Model ReleasesDGX agent

I came up with a somewhat foolish new benchmark for testing image generation models, to exercise the new ChatGPT Images 2.0: 'Do a where's Waldo style image but it's where is the raccoon holding a ham

I find that open weights models over-perform on benchmarks compared to actual real-world usage, and Kimi feels like no exception. For exampl…

Model ReleasesDGX agent

I find that open weights models over-perform on benchmarks compared to actual real-world usage, and Kimi feels like no exception. For example, a small amount of use will show that Kimi is not as good

I hope this helps. If you need more help let me know: Connecting OpenClaw to the X API is straightforward now thanks to X’s official native …

Model ReleasesDGX agent

I hope this helps. If you need more help let me know: Connecting OpenClaw to the X API is straightforward now thanks to X’s official native support... The best and most direct method uses the official

ICAT: Incident-Case-Grounded Adaptive Testing for Physical-Risk Prediction in Embodied World Models

Model ReleasesDGX agent

arXiv:2604.16405v1 Announce Type: cross Abstract: Video-generative world models are increasingly used as neural simulators for embodied planning and policy learning, yet their ability to predict physi

IDOBE: Infectious Disease Outbreak forecasting Benchmark Ecosystem

Model ReleasesDGX agent

arXiv:2604.18521v1 Announce Type: new Abstract: Epidemic forecasting has become an integral part of real-time infectious disease outbreak response. While collaborative ensembles composed of statistica

iDocV2: Leveraging Self-Supervision and Open-Set Detection for Improving Pattern Spotting in Historical Documents

Model ReleasesDGX agent

arXiv:2604.16726v1 Announce Type: new Abstract: Considering the imminent massification of digital books, it has become critical to facilitate searching collections through graphical patterns. Current

Improving LLM Code Reasoning via Semantic Equivalence Self-Play with Formal Verification

Model ReleasesDGX agent

arXiv:2604.17010v1 Announce Type: new Abstract: We introduce a self-play framework for semantic equivalence in Haskell, utilizing formal verification to guide adversarial training between a generator

In-Context Symbolic Regression for Robustness-Improved Kolmogorov-Arnold Networks

Model ReleasesDGX agent

arXiv:2603.15250v2 Announce Type: replace Abstract: Symbolic regression aims to replace black-box predictors with concise analytical expressions that can be inspected and validated in scientific machi

Incoherent Deformation, Not Capacity: Diagnosing and Mitigating Overfitting in Dynamic Gaussian Splatting

Model ReleasesDGX agent

arXiv:2604.16747v1 Announce Type: new Abstract: Dynamic 3D Gaussian Splatting methods achieve strong training-view PSNR on monocular video but generalize poorly: on the D-NeRF benchmark we measure an

IncreFA: Breaking the Static Wall of Generative Model Attribution

Model ReleasesDGX agent

arXiv:2604.17736v1 Announce Type: new Abstract: As AI generative models evolve at unprecedented speed, image attribution has become a moving target. New diffusion, adversarial and autoregressive gener

Inflated Excellence or True Performance? Rethinking Medical Diagnostic Benchmarks with Dynamic Evaluation

Model ReleasesDGX agent

arXiv:2510.09275v2 Announce Type: replace Abstract: Medical diagnostics is a high-stakes and complex domain that is critical to patient care. However, current evaluations of large language models (LLM

Injecting Structured Biomedical Knowledge into Language Models: Continual Pretraining vs. GraphRAG

Model ReleasesDGX agent

arXiv:2604.16422v1 Announce Type: new Abstract: The injection of domain-specific knowledge is crucial for adapting language models (LMs) to specialized fields such as biomedicine. While most current a

INTENT: Invariance and Discrimination-aware Noise Mitigation for Robust Composed Image Retrieval

Model ReleasesDGX agent

arXiv:2604.18051v1 Announce Type: new Abstract: Composed Image Retrieval (CIR) is a challenging image retrieval paradigm that enables to retrieve target images based on multimodal queries consisting o

InternScenes: A Large-scale Simulatable Indoor Scene Dataset with Realistic Layouts

Model ReleasesDGX agent

arXiv:2509.10813v3 Announce Type: replace Abstract: The advancement of Embodied AI heavily relies on large-scale, simulatable 3D scene datasets characterized by scene diversity and realistic layouts.

Interpolating Discrete Diffusion Models with Controllable Resampling

Model ReleasesDGX agent

arXiv:2604.17310v1 Announce Type: new Abstract: Discrete diffusion models form a powerful class of generative models across diverse domains, including text and graphs. However, existing approaches fac

Introducing ChatGPT Images 2.0

Model ReleasesDGX agent

ChatGPT Images 2.0 is an updated version of OpenAI's image generation feature integrated into ChatGPT, offering improvements to image creation capabilities for users. The update likely includes enhanc

Introducing ChatGPT Images 2.0 A state-of-the-art image model that can take on complex visual tasks and produce precise, immediately usable …

Model ReleasesDGX agent

Introducing ChatGPT Images 2.0 A state-of-the-art image model that can take on complex visual tasks and produce precise, immediately usable visuals, with sharper editing, richer layouts, and thinking-

Introducing ml-intern, the agent that just automated the post-training team @huggingface It's an open-source implementation of the real rese…

Model ReleasesDGX agent

Introducing ml-intern, the agent that just automated the post-training team @huggingface It's an open-source implementation of the real research loop that our ML researchers do every day. You give it

Judge a Book by its Cover: Investigating Multi-Modal LLMs for Multi-Page Handwritten Document Transcription

Model ReleasesDGX agent

arXiv:2502.20295v2 Announce Type: replace-cross Abstract: Handwriting text recognition (HTR) remains a challenging task. Existing approaches require fine-tuning on labeled data, which is impractical t

JudgeMeNot: Personalizing Large Language Models to Emulate Judicial Reasoning in Hebrew

Model ReleasesDGX agent

arXiv:2604.18041v1 Announce Type: new Abstract: Despite significant advances in large language models, personalizing them for individual decision-makers remains an open problem. Here, we introduce a s

Jupiter-N Technical Report

Model ReleasesDGX agent

arXiv:2604.17429v1 Announce Type: new Abstract: We present Jupiter-N, a hybrid reasoning model post-trained from Nemotron 3 Super, a fully open-source 120 billion parameter LLM. We target three object

K2.6 + hermes = 4 hr session setting up qwen 3.6 training regime on dgx spark with the current autnomous session currently lasting 70+ min w…

Model ReleasesDGX agent

K2.6 + hermes = 4 hr session setting up qwen 3.6 training regime on dgx spark with the current autnomous session currently lasting 70+ min without any prompting. Kimi with hermes is next level. @NousR

Keep improving 🚀🚀 @arena

Model ReleasesDGX agent

Keep improving 🚀🚀 @arena Qwen3.6 Plus lands at #7 in Code Arena with a score of 1476 - up +16 points since the Preview. The new score also moves @AlibabaGroup to #3 lab in Code Arena. In the Text Aren

📢 Kimi K2.6 API is live • Input Price (Cache Hit): 0.16 / M tokens • Input Price (Cache Miss): 0.95 / M tokens • Output: $4.00 / M tokens…

Model ReleasesDGX agent

📢 Kimi K2.6 API is live • Input Price (Cache Hit): 0.16 / M tokens • Input Price (Cache Miss): 0.95 / M tokens • Output: $4.00 / M tokens Kimi K2.6 is our latest + most intelligent model - stronger lo

Kimi K2.6 autonomously overhauled exchange-core, an 8-year-old open-source financial matching engine. Over a 13-hour execution, the model it…

Model ReleasesDGX agent

Kimi K2.6 autonomously overhauled exchange-core, an 8-year-old open-source financial matching engine. Over a 13-hour execution, the model iterated through 12 optimization strategies, initiating over 1

Kimi K2.6 has captured #1 on the open-weight Vals Index, and is #7 overall.

Model ReleasesDGX agent

Kimi K2.6, a language model developed by Moonshot AI, has achieved the top ranking on the open-weight category of the Vals Index benchmark, while placing 7th overall across all model categories. The V

Kimi K2.6 is now live inside Anything!

Model ReleasesDGX agent

Kimi K2.6, an AI model from Moonshot, has been integrated into the Anything platform. This update likely enables users to access Kimi's capabilities directly within the Anything application interface.

Kimi K2.6 @Kimi_Moonshot is the new leading open-weights agent model, landing at #4 on Claw-Eval (Pass^3: 62.3%). Key takeaways: - 👑 Best o…

Model ReleasesDGX agent

Kimi K2.6 @Kimi_Moonshot is the new leading open-weights agent model, landing at #4 on Claw-Eval (Pass^3: 62.3%). Key takeaways: - 👑 Best open-source agent, period: Pass^3 of 62.3% is the highest of a

KIRA: Knowledge-Intensive Image Retrieval and Reasoning Architecture for Specialized Visual Domains

Model ReleasesDGX agent

arXiv:2604.16915v1 Announce Type: new Abstract: Retrieval augmented generation (RAG) has transformed text based question answering, yet its extension to visual domains remains hindered by fundamental

Knowing When to Quit: A Principled Framework for Dynamic Abstention in LLM Reasoning

Model ReleasesDGX agent

arXiv:2604.18419v1 Announce Type: cross Abstract: Large language models (LLMs) using chain-of-thought reasoning often waste substantial compute by producing long, incorrect responses. Abstention can m

Knowledge without Wisdom: Measuring Misalignment between LLMs and Intended Impact

Model ReleasesDGX agent

arXiv:2603.00883v2 Announce Type: replace Abstract: LLMs increasingly excel on AI benchmarks, but doing so does not guarantee validity for downstream tasks. This study contrasts LLM alignment on bench

← Previous
1…327328329330331…373
Next →