AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,340 results
22 Apr 2026

LegalBench-BR: A Benchmark for Evaluating Large Language Models on Brazilian Legal Decision Classification

Model ReleasesDGX agent

arXiv:2604.18878v1 Announce Type: new Abstract: We introduce LegalBench-BR, the first public benchmark for evaluating language models on Brazilian legal text classification. The dataset comprises 3,10

LePREC: Reasoning as Classification over Structured Factors for Assessing Relevance of Legal Issues

Model ReleasesDGX agent

arXiv:2604.19464v1 Announce Type: cross Abstract: More than half of the global population struggles to meet their civil justice needs due to limited legal resources. While Large Language Models (LLMs)

Less Is More: Cognitive Load and the Single-Prompt Ceiling in LLM Mathematical Reasoning


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases
DGX agent

arXiv:2604.18897v1 Announce Type: new Abstract: We present a systematic empirical study of prompt engineering for formal mathematical reasoning in the context of the SAIR Equational Theories Stage 1 c

Level Up Your Agents: Announcing Google's Official Skills Repository

Model ReleasesDGX agent

As AI models improve, technical practitioners are increasingly turning to agentic AI tools to build with Google Cloud products, from Firebase and the Gemini API, to BigQuery and GKE. But how can you e

Lingua-SafetyBench: A Benchmark for Safety Evaluation of Multilingual Vision-Language Models

Model ReleasesDGX agent

arXiv:2601.22737v2 Announce Type: replace Abstract: The robust safety of Vision-Language Large Models (VLLMs) against joint multilingual and multimodal threats remains severely underexplored. Current

LiteParse, our OSS document parser, is really good at parsing complex PDF layouts, text, and tables into a clean spatial grid. The best part…

Model ReleasesDGX agent

LiteParse, our OSS document parser, is really good at parsing complex PDF layouts, text, and tables into a clean spatial grid. The best part is it doesn't use VLMs or any ML models at all. It's entire

LiveVLN: Breaking the Stop-and-Go Loop in Vision-Language Navigation

Model ReleasesDGX agent

arXiv:2604.19536v1 Announce Type: new Abstract: Recent navigation systems achieve strong benchmark results, yet real-world deployment often remains visibly stop-and-go. This bottleneck arises because

llama-server -hf ggml-org/Qwen3.6-27B-GGUF --spec-default

Model ReleasesDGX agent

This post likely demonstrates running Qwen2 3.6B or 27B model in GGUF format using llama-server with default specifications, showcasing inference capabilities of quantized open-source models. The comm

LLM-as-Judge Framework for Evaluating Tone-Induced Hallucination in Vision-Language Models

Model ReleasesDGX agent

arXiv:2604.18803v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) are increasingly deployed in settings where reliable visual grounding carries operational consequences, yet their behavi

LM Performance:With only 27B parameters, Qwen3.6-27B outperforms the Qwen3.5-397B-A17B (397B total / 17B active, ~15x larger!) on every majo…

Model ReleasesDGX agent

LM Performance:With only 27B parameters, Qwen3.6-27B outperforms the Qwen3.5-397B-A17B (397B total / 17B active, ~15x larger!) on every major coding benchmark — including SWE-bench Verified (77.2 vs.

Lost in Translation: Do LVLM Judges Generalize Across Languages?

Model ReleasesDGX agent

arXiv:2604.19405v1 Announce Type: new Abstract: Automatic evaluators such as reward models play a central role in the alignment and evaluation of large vision-language models (LVLMs). Despite their gr

LoViF 2026 Challenge on Real-World All-in-One Image Restoration: Methods and Results

Model ReleasesDGX agent

arXiv:2604.19445v1 Announce Type: new Abstract: This paper presents a review for the LoViF Challenge on Real-World All-in-One Image Restoration. The challenge aimed to advance research on real-world a

LSTM-MAS: A Long Short-Term Memory Inspired Multi-Agent System for Long-Context Understanding

Model ReleasesDGX agent

arXiv:2601.11913v2 Announce Type: replace-cross Abstract: Effectively processing long contexts remains a fundamental yet unsolved challenge for large language models (LLMs). Existing single-LLM-based

Making ChatGPT better for clinicians

Model ReleasesDGX agent

OpenAI describes improvements and adaptations made to ChatGPT to better serve clinical and healthcare professionals. The work likely addresses how the model can be made more reliable, accurate, and us

Mango: Multi-Agent Web Navigation via Global-View Optimization

Model ReleasesDGX agent

arXiv:2604.18779v1 Announce Type: new Abstract: Existing web agents typically initiate exploration from the root URL, which is inefficient for complex websites with deep hierarchical structures. Witho

Maximizing Gemini: Google Cloud makes its bid to build the operating system for enterprise AI

Model ReleasesDGX agent

Google LLC has emerged as the only cloud “hyperscaler” with a leading frontier artificial intelligence large language model – Gemini – and today it issued a raft of announcements designed to capitaliz

Mechanistic Anomaly Detection via Functional Attribution

Model ReleasesDGX agent

arXiv:2604.18970v1 Announce Type: new Abstract: We can often verify the correctness of neural network outputs using ground truth labels, but we cannot reliably determine whether the output was produce

🚀 Meet Qwen3.6-27B, our latest dense, open-source model, packing flagship-level coding power! Yes, 27B, and Qwen3.6-27B punches way above i…

Model ReleasesDGX agent

🚀 Meet Qwen3.6-27B, our latest dense, open-source model, packing flagship-level coding power! Yes, 27B, and Qwen3.6-27B punches way above its weight. 👇 What's new: 🧠 Outstanding agentic coding — surpa

Michele Catasta (@pirroh) at Google Cloud Next talking about what's new with Gemini, agentic applications, and what's next for builders. Fre…

Model ReleasesDGX agent

Michele Catasta (@pirroh) at Google Cloud Next talking about what's new with Gemini, agentic applications, and what's next for builders. Fresh off Replit being named Google Cloud Partner of the Year.

Micro Language Models Enable Instant Responses

Model ReleasesDGX agent

arXiv:2604.19642v1 Announce Type: new Abstract: Edge devices such as smartwatches and smart glasses cannot continuously run even the smallest 100M-1B parameter language models due to power and compute

MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling

Model ReleasesDGX agent

arXiv:2511.11793v3 Announce Type: replace Abstract: We present MiroThinker v1.0, an open-source research agent designed to advance tool-augmented reasoning and information-seeking capabilities. Unlike

ml-intern by @huggingface is wild 🔥 You drop a high-level prompt (“build the best scientific reasoning model” or “crush healthcare benchmar…

Model ReleasesDGX agent

ml-intern by @huggingface is wild 🔥 You drop a high-level prompt (“build the best scientific reasoning model” or “crush healthcare benchmarks”) and this open-source agent does the entire post-training

MORPHOGEN: A Multilingual Benchmark for Evaluating Gender-Aware Morphological Generation

Model ReleasesDGX agent

arXiv:2604.18914v1 Announce Type: cross Abstract: While multilingual large language models (LLMs) perform well on high-level tasks like translation and question answering, their ability to handle gram

MSDS: Deep Structural Similarity with Multiscale Representation

Model ReleasesDGX agent

arXiv:2604.19159v1 Announce Type: new Abstract: Deep-feature-based perceptual similarity models have demonstrated strong alignment with human visual perception in Image Quality Assessment (IQA). Howev

Multi-Domain Learning with Global Expert Mapping

Model ReleasesDGX agent

arXiv:2604.18842v1 Announce Type: new Abstract: Human perception generalizes well across different domains, but most vision models struggle beyond their training data. This gap motivates multi-dataset

Multi-modal Test-time Adaptation via Adaptive Probabilistic Gaussian Calibration

Model ReleasesDGX agent

arXiv:2604.19093v1 Announce Type: cross Abstract: Multi-modal test-time adaptation (TTA) enhances the resilience of benchmark multi-modal models against distribution shifts by leveraging the unlabeled

Multilingual Language Models Encode Script Over Linguistic Structure

Model ReleasesDGX agent

arXiv:2604.05090v2 Announce Type: replace Abstract: Multilingual language models (LMs) organize representations for typologically and orthographically diverse languages into a shared parameter space,

NemeSys: Toward Online Underwater Exploration with Remote Operator-in-the-loop Adaptive Autonomy

Model ReleasesDGX agent

arXiv:2507.11889v2 Announce Type: replace Abstract: Adaptive mission control and dynamic parameter reconfiguration are essential for autonomous underwater vehicles (AUVs) operating in GPS-denied, comm

New in Claude Code: /ultrareview (research preview) runs a fleet of bug-hunting agents in the cloud. Findings land in the CLI or Desktop aut…

Model ReleasesDGX agent

New in Claude Code: /ultrareview (research preview) runs a fleet of bug-hunting agents in the cloud. Findings land in the CLI or Desktop automatically. Run it before merging critical changes—auth, dat

New innovations in Google Distributed Cloud

Model ReleasesDGX agent

Today at Google Cloud Next, we’re announcing new capabilities in Google Distributed Cloud (GDC) that bring Gemini and our advanced AI stack to wherever your data is, so you don’t need to compromise be

Not the first time either - they shut down a bunch of of their original proprietary hosted embedding models in this announcement back in Apr…

Model ReleasesDGX agent

Not the first time either - they shut down a bunch of of their original proprietary hosted embedding models in this announcement back in April 2024 https://openai.com/index/gpt-4-api-general-availabil

ODMA: On-Demand Memory Allocation Strategy for LLM Serving on LPDDR-Class Accelerators

Model ReleasesDGX agent

arXiv:2512.09427v5 Announce Type: replace-cross Abstract: Existing memory management techniques severely hinder efficient Large Language Model serving on accelerators constrained by poor random-access

OLLM: Options-based Large Language Models

Model ReleasesDGX agent

arXiv:2604.19087v1 Announce Type: new Abstract: We introduce Options LLM (OLLM), a simple, general method that replaces the single next-token prediction of standard LLMs with a extit{set of learned op

OmniGen2: Towards Instruction-Aligned Multimodal Generation

Model ReleasesDGX agent

arXiv:2506.18871v4 Announce Type: replace-cross Abstract: In this work, we introduce OmniGen2, a versatile and open-source generative model designed to provide a unified solution for diverse generatio

OmniMouse: Scaling properties of multi-modal, multi-task Brain Models on 150B Neural Tokens

Model ReleasesDGX agent

arXiv:2604.18827v1 Announce Type: cross Abstract: Scaling data and artificial neural networks has transformed AI, driving breakthroughs in language and vision. Whether similar principles apply to mode

Once Correct, Still Wrong: Counterfactual Hallucination in Multilingual Vision-Language Models

Model ReleasesDGX agent

arXiv:2602.05437v2 Announce Type: replace Abstract: Vision-language models (VLMs) can achieve high accuracy while still accepting culturally plausible but visually incorrect interpretations. Existing

One Step Forward and K Steps Back: Better Reasoning with Denoising Recursion Models

Model ReleasesDGX agent

arXiv:2604.18839v1 Announce Type: cross Abstract: Looped transformers scale computational depth without increasing parameter count by repeatedly applying a shared transformer block and can be used for

OpenAI dropped a new model on HF today!

Model ReleasesDGX agent

OpenAI released a new model that was made available on Hugging Face, as announced by Hugging Face CEO Clem Delangue on X (formerly Twitter). The specific details about which model was released are not

OpenAI just released a new open-source model it's 'a bidirectional token-classification model for personally identifiable information (PII) …

Model ReleasesDGX agent

OpenAI just released a new open-source model it's 'a bidirectional token-classification model for personally identifiable information (PII) detection and masking in text' https://github.com/openai/pri

Oracle expands Google Cloud partnership with natural language database agent

Model ReleasesDGX agent

Oracle Corp. is extending its partnership with Google LLC’s Cloud to simplify how enterprise users interact with data, introducing a new natural-language interface for queries directly against Oracle

ORCA: An Agentic Reasoning Framework for Hallucination and Adversarial Robustness in Vision-Language Models

Model ReleasesDGX agent

arXiv:2509.15435v2 Announce Type: replace-cross Abstract: Large Vision-Language Models (LVLMs) exhibit strong multimodal capabilities but remain vulnerable to hallucinations from intrinsic errors and

Owner-Harm: A Missing Threat Model for AI Agent Safety

Model ReleasesDGX agent

arXiv:2604.18658v1 Announce Type: cross Abstract: Existing AI agent safety benchmarks focus on generic criminal harm (cybercrime, harassment, weapon synthesis), leaving a systematic blind spot for a d

Paparazzo: Active Mapping of Moving 3D Objects

Model ReleasesDGX agent

arXiv:2604.19556v1 Announce Type: new Abstract: Current 3D mapping pipelines generally assume static environments, which limits their ability to accurately capture and reconstruct moving objects. To a

PC2Model: ISPRS benchmark on 3D point cloud to model registration

Model ReleasesDGX agent

arXiv:2604.19596v1 Announce Type: new Abstract: Point cloud registration involves aligning one point cloud with another or with a three-dimensional (3D) model, enabling the integration of multimodal d

Pixels or Positions? Benchmarking Modalities in Group Activity Recognition

Model ReleasesDGX agent

arXiv:2511.12606v3 Announce Type: replace Abstract: Group Activity Recognition (GAR) is well studied on the video modality for surveillance and indoor team sports (e.g., volleyball, basketball). Yet,

Polymarket says perpetuals are coming to the platform and lets users sign up for early access; it hasn't specified whether crypto perpetual futures are included (Tanaya Macheel/CNBC)

Model ReleasesDGX agent

Tanaya Macheel / CNBC: Polymarket says perpetuals are coming to the platform and lets users sign up for early access; it hasn't specified whether crypto perpetual futures are included — Prediction mar

Presumably GPT-imagegen-2 (aka ChatGPT Images 2.0 aka gpt-image-2) works as a tool which the models generate prompts for? I wish we could se…

Model ReleasesDGX agent

Presumably GPT-imagegen-2 (aka ChatGPT Images 2.0 aka gpt-image-2) works as a tool which the models generate prompts for? I wish we could see those prompts, like back in the DALL-E 3 days https://simo

PriorGuide: Test-Time Prior Adaptation for Simulation-Based Inference

Model ReleasesDGX agent

arXiv:2510.13763v2 Announce Type: replace-cross Abstract: Amortized simulator-based inference offers a powerful framework for tackling Bayesian inference in computational fields such as engineering or

Protecting Bystander Privacy via Selective Hearing in Audio LLMs

Model ReleasesDGX agent

arXiv:2512.06380v3 Announce Type: replace-cross Abstract: Audio Large language models (LLMs) are increasingly deployed in the real world, where they inevitably capture speech from unintended nearby by

PuzzleWorld: A Benchmark for Multimodal, Open-Ended Reasoning in Puzzlehunts

Model ReleasesDGX agent

arXiv:2506.06211v2 Announce Type: replace-cross Abstract: Puzzlehunts are a genre of complex, multi-step puzzles lacking well-defined problem definitions. In contrast to conventional reasoning benchma

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models

Model ReleasesDGX agent

arXiv:2604.00161v2 Announce Type: replace Abstract: Optical Character Recognition (OCR) is increasingly regarded as a foundational capability for modern vision-language models (VLMs), enabling them no

Quadruped Parkour Learning: Sparsely Gated Mixture of Experts with Visual Input

Model ReleasesDGX agent

arXiv:2604.19344v1 Announce Type: new Abstract: Robotic parkour provides a compelling benchmark for advancing locomotion over highly challenging terrain, including large discontinuities such as elevat

Quoting Bobby Holley

Model ReleasesDGX agent

As part of our continued collaboration with Anthropic, we had the opportunity to apply an early version of Claude Mythos Preview to Firefox. This week’s release of Firefox 150 includes fixes for 271 v

Qwen 3.6 27B model is available on Ollama! Use it with all the integrations in Ollama or chat with the model. Chat with the model: ollama ru…

Model ReleasesDGX agent

Qwen 3.6 27B model is available on Ollama! Use it with all the integrations in Ollama or chat with the model. Chat with the model: ollama run qwen3.6:27b OpenClaw: ollama launch openclaw --model qwen3

Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model

Model ReleasesDGX agent

Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model Big claims from Qwen about their latest open weight model: Qwen3.6-27B delivers flagship-level agentic coding performance, surpassing the previo

Qwen3.6 27B is now in LM Studio! Vision, reasoning, and agentic tool calling - running locally on your computer. Surpasses previous Qwen mod…

Model ReleasesDGX agent

Qwen3.6 27B is now in LM Studio! Vision, reasoning, and agentic tool calling - running locally on your computer. Surpasses previous Qwen models many times its size 🚀👾🔥 https://lmstudio.ai/models/qwen/

Qwen3.6 and gemma 4 just shown that we were still far from capability limit on small models, we still are. Huge models are good, but is ther…

Model ReleasesDGX agent

Qwen3.6 and gemma 4 just shown that we were still far from capability limit on small models, we still are. Huge models are good, but is there really a point in going bigger. Just realize something. On

RAFT-MSF++: Temporal Geometry-Motion Feature Fusion for Self-Supervised Monocular Scene Flow

Model ReleasesDGX agent

arXiv:2604.19349v1 Announce Type: new Abstract: Monocular scene flow estimation aims to recover dense 3D motion from image sequences, yet most existing methods are limited to two-frame inputs, restric

Rank-Turbulence Delta and Interpretable Approaches to Stylometric Delta Metrics

Model ReleasesDGX agent

arXiv:2604.19499v1 Announce Type: new Abstract: This article introduces two new measures for authorship attribution - Rank-Turbulence Delta and Jensen-Shannon Delta - which generalise Burrows's classi

RARE: Redundancy-Aware Retrieval Evaluation Framework for High-Similarity Corpora

Model ReleasesDGX agent

arXiv:2604.19047v1 Announce Type: cross Abstract: Existing QA benchmarks typically assume distinct documents with minimal overlap, yet real-world retrieval-augmented generation (RAG) systems operate o

← Previous
1…322323324325326…373
Next →