AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,284 results
16 Apr 2026

LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks

Model ReleasesDGX agent

arXiv:2604.13072v1 Announce Type: new Abstract: LLM-based agents are increasingly expected to handle real-world assistant tasks, yet existing benchmarks typically evaluate them under isolated sources

llm-anthropic 0.25

Model ReleasesDGX agent

Release: llm-anthropic 0.25 New model: claude-opus-4.7, which supports thinking_effort: xhigh. #66 New thinking_display and thinking_adaptive boolean options. thinking_display summarized output is cur

LongCoT: Benchmarking Long-Horizon Chain-of-Thought Reasoning

Model ReleasesDGX agent

arXiv:2604.14140v1 Announce Type: new Abstract: As language models are increasingly deployed for complex autonomous tasks, their ability to reason accurately over longer horizons becomes critical. An


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding

Model ReleasesDGX agent

arXiv:2602.20913v2 Announce Type: replace Abstract: This paper addresses the critical and underexplored challenge of long video understanding with low computational budgets. We propose LongVideo-R1, a

LoRA-MME: Multi-Model Ensemble of LoRA-Tuned Encoders for Code Comment Classification

Model ReleasesDGX agent

arXiv:2603.03959v4 Announce Type: replace-cross Abstract: Code comment classification is a critical task for automated software documentation and analysis. In the context of the NLBSE'26 Tool Competit

Lossless Prompt Compression via Dictionary-Encoding and In-Context Learning: Enabling Cost-Effective LLM Analysis of Repetitive Data

Model ReleasesDGX agent

arXiv:2604.13066v1 Announce Type: new Abstract: In-context learning has established itself as an important learning paradigm for Large Language Models (LLMs). In this paper, we demonstrate that LLMs c

MAny: Merge Anything for Multimodal Continual Instruction Tuning

Model ReleasesDGX agent

arXiv:2604.14016v1 Announce Type: new Abstract: Multimodal Continual Instruction Tuning (MCIT) is essential for sequential task adaptation of Multimodal Large Language Models (MLLMs) but is severely r

MDPs with a State Sensing Cost

Model ReleasesDGX agent

arXiv:2505.03280v3 Announce Type: replace Abstract: In many practical sequential decision-making problems, tracking the state of the environment incurs a sensing/communication/computation cost. In the

MedRCube: A Multidimensional Framework for Fine-Grained and In-Depth Evaluation of MLLMs in Medical Imaging

Model ReleasesDGX agent

arXiv:2604.13756v1 Announce Type: new Abstract: The potential of Multimodal Large Language Models (MLLMs) in domain of medical imaging raise the demands of systematic and rigorous evaluation framework

⚡ Meet Qwen3.6-35B-A3B:Now Open-Source!🚀🚀 A sparse MoE model, 35B total params, 3B active. Apache 2.0 license. 🔥 Agentic coding on par wi…

Model ReleasesDGX agent

⚡ Meet Qwen3.6-35B-A3B:Now Open-Source!🚀🚀 A sparse MoE model, 35B total params, 3B active. Apache 2.0 license. 🔥 Agentic coding on par with models 10x its active size 📷 Strong multimodal perception an

MERRIN: A Benchmark for Multimodal Evidence Retrieval and Reasoning in Noisy Web Environments

Model ReleasesDGX agent

arXiv:2604.13418v1 Announce Type: new Abstract: Motivated by the underspecified, multi-hop nature of search queries and the multimodal, heterogeneous, and often conflicting nature of real-world web re

Merry Claude-mas! Opus 4.7 in Claude Code is a monster. Very happy camper! https://www.anthropic.com/news/claude-opus-4-7

Model ReleasesDGX agent

This post celebrates Claude Opus 4.7's capabilities within Claude Code, expressing enthusiasm about its performance and describing it as exceptionally powerful. The entry references an Anthropic annou

Mitigating Catastrophic Forgetting in Target Language Adaptation of LLMs via Source-Shielded Updates

Model ReleasesDGX agent

arXiv:2512.04844v2 Announce Type: replace Abstract: Expanding the linguistic diversity of instruct large language models (LLMs) is crucial for global accessibility but is often hindered by the relianc

MM-Doc-R1: Training Agents for Long Document Visual Question Answering through Multi-turn Reinforcement Learning

Model ReleasesDGX agent

arXiv:2604.13579v1 Announce Type: new Abstract: Conventional Retrieval-Augmented Generation (RAG) systems often struggle with complex multi-hop queries over long documents due to their single-pass ret

Models are getting smaller, smarter and Apache licensed. Love to see Gemma and Qwen doing it.

Model ReleasesDGX agent

Models are getting smaller, smarter and Apache licensed. Love to see Gemma and Qwen doing it. Qwen 3.6 is here, and open-source! Run it locally with improved agentic coding capabilities. Try it with C

MolCryst-MLIPs: A Machine-Learned Interatomic Potentials Database for Molecular Crystals

Model ReleasesDGX agent

arXiv:2604.13897v1 Announce Type: new Abstract: We present an open Molecular Crystal (MC) database of Machine-Learned Interatomic Potentials (MLIP) called MolCryst-MLIPs. The first release comprises f

MOONSHOT : A Framework for Multi-Objective Pruning of Vision and Large Language Models

Model ReleasesDGX agent

arXiv:2604.13287v1 Announce Type: new Abstract: Weight pruning is a common technique for compressing large neural networks. We focus on the challenging post-training one-shot setting, where a pre-trai

More on my blog, including results from the previously secret 'flamingo on a unicycle' test https://simonwillison.net/2026/Apr/16/qwen-beats…

Model ReleasesDGX agent

Simon Willison discusses results from a 'flamingo on a unicycle' test on his blog, likely comparing AI model performance including Qwen. The post appears to reference previously undisclosed or unconve

Mosaic: An Extensible Framework for Composing Rule-Based and Learned Motion Planners

Model ReleasesDGX agent

arXiv:2604.13853v1 Announce Type: new Abstract: Safe and explainable motion planning remains a central challenge in autonomous driving. While rule-based planners offer predictable and explainable beha

Mozilla launches Thunderbolt AI client with focus on self-hosted infrastructure

Model ReleasesDGX agent

Thunderbolt is a new open-source AI client from Mozilla-owned MZLA Technologies aimed at enterprises who want to run self-hosted chatbots on their own infrastructure. The platform allows organizations

Mozilla launches Thunderbolt, an open-source AI client for users and businesses who want to run their own self-hosted AI infrastructure, available on GitHub (Kyle Orland/Ars Technica)

Model ReleasesDGX agent

Kyle Orland / Ars Technica: Mozilla launches Thunderbolt, an open-source AI client for users and businesses who want to run their own self-hosted AI infrastructure, available on GitHub — Mozilla is th

MulDimIF: A Multi-Dimensional Constraint Framework for Evaluating and Improving Instruction Following in Large Language Models

Model ReleasesDGX agent

arXiv:2505.07591v2 Announce Type: replace Abstract: Instruction following refers to the ability of large language models (LLMs) to generate outputs that satisfy all specified constraints. Existing res

Multi-Task LLM with LoRA Fine-Tuning for Automated Cancer Staging and Biomarker Extraction

Model ReleasesDGX agent

arXiv:2604.13328v1 Announce Type: new Abstract: Pathology reports serve as the definitive record for breast cancer staging, yet their unstructured format impedes large-scale data curation. While Large

New insane model from Jackrong on @huggingface 🤯 Qwen3.5-9B-GLM5.1-Distill-v1 🧠 Distilled on GLM-5.1 reasoning ⚙️ Deeper thinking than bas…

Model ReleasesDGX agent

New insane model from Jackrong on @huggingface 🤯 Qwen3.5-9B-GLM5.1-Distill-v1 🧠 Distilled on GLM-5.1 reasoning ⚙️ Deeper thinking than base model 🧪 Benchmarks coming soon ✅ Fits on 8GB VRAM ✍️ New mod

New ways to create personalized images in the Gemini app

Model ReleasesDGX agent

Gemini has been updated to generate highly personalized images by accessing a user's Google Photos and personal preferences. This new feature allows users to simply request scenes featuring themselves

No Need for Space Gear — Capcom’s ‘PRAGMATA’ Joins GeForce NOW on Launch Day

Model ReleasesDGX agent

Head straight for orbit with GeForce NOW — no space helmet required. PRAGMATA, Capcom’s long-awaited sci-fi action adventure, touches down on GeForce NOW the same day it launches worldwide. The futuri

Olfactory pursuit: catching a moving odor source in complex flows

Model ReleasesDGX agent

arXiv:2604.13121v1 Announce Type: new Abstract: Locating and intercepting a moving target from possibly delayed, intermittent sensory signals is a paradigmatic problem in decision-making under uncerta

🎬 Ollama Gemma Day Recap: SGLang at the Ollama Gemma 4 Party in Palo Alto 🍾 Last night, @ollama hosted a packed Gemma Day at the Palo Alto…

Model ReleasesDGX agent

🎬 Ollama Gemma Day Recap: SGLang at the Ollama Gemma 4 Party in Palo Alto 🍾 Last night, @ollama hosted a packed Gemma Day at the Palo Alto office alongside the @GoogleDeepMind Gemma team. SGLang was i

Online learning with noisy side observations

Model ReleasesDGX agent

arXiv:2604.13740v1 Announce Type: new Abstract: We propose a new partial-observability model for online learning problems where the learner, besides its own loss, also observes some noisy feedback abo

OpenAI launches GPT-Rosalind, an AI model for life sciences research, including drug discovery, as a research preview for customers such as Moderna and Amgen (Megan Morrone/Axios)

Model ReleasesDGX agent

Megan Morrone / Axios: OpenAI launches GPT-Rosalind, an AI model for life sciences research, including drug discovery, as a research preview for customers such as Moderna and Amgen — OpenAI announced

OpenAI’s big Codex update is a direct shot at Claude Code

Model ReleasesDGX agent

OpenAI is beefing up its agentic coding and development system, Codex, with a suite of updates that let it use your computer, generate images, and remember from past experiences. The package of update

OPTED: Open Preprocessed Trachoma Eye Dataset Using Zero-Shot SAM 3 Segmentation

Model ReleasesDGX agent

arXiv:2603.06885v2 Announce Type: replace Abstract: Trachoma remains the leading infectious cause of blindness worldwide, with Sub-Saharan Africa bearing over 85% of the global burden and Ethiopia alo

Optimization with SpotOptim

Model ReleasesDGX agent

arXiv:2604.13672v1 Announce Type: new Abstract: The `spotoptim` package implements surrogate-model-based optimization of expensive black-box functions in Python. Building on two decades of Sequential

Opus 4.7 feels more intelligent, agentic, and precise than 4.6. It took a few days for me to learn how to work with it effectively, to fully…

Model ReleasesDGX agent

Opus 4.7 feels more intelligent, agentic, and precise than 4.6. It took a few days for me to learn how to work with it effectively, to fully take advantage of its new capabilities. Will post a few mor

Opus 4.7 is a model I’ve loved working with in Claude Code. It’s more agentic and instruction following but also incredibly smart and creati…

Model ReleasesDGX agent

Opus 4.7 is a model I’ve loved working with in Claude Code. It’s more agentic and instruction following but also incredibly smart and creative. I think it takes a slight adjustment to get used to, but

Opus 4.7 is in Claude Code today. It's more agentic, more precise, and a lot better at long-running work. It carries context across sessions…

Model ReleasesDGX agent

Opus 4.7 is in Claude Code today. It's more agentic, more precise, and a lot better at long-running work. It carries context across sessions and handles ambiguity much better. Introducing Claude Opus

Opus 4.7 is now supported in Hermes Agent 🚀🚀

Model ReleasesDGX agent

Opus 4.7 is now supported in Hermes Agent 🚀🚀 Introducing Claude Opus 4.7, our most capable Opus model yet. It handles long-running tasks with more rigor, follows instructions more precisely, and verif

Ordinary Least Squares is a Special Case of Transformer

Model ReleasesDGX agent

arXiv:2604.13656v1 Announce Type: new Abstract: The statistical essence of the Transformer architecture has long remained elusive: Is it a universal approximator, or a neural network version of known

Our Life Sciences model series is available as a research preview starting today for qualified customers including @Amgen, @moderna_tx, the …

Model ReleasesDGX agent

Our Life Sciences model series is available as a research preview starting today for qualified customers including @Amgen, @moderna_tx, the @AllenInstitute, and @thermofisher Scientific through ChatGP

Out of Context: Reliability in Multimodal Anomaly Detection Requires Contextual Inference

Model ReleasesDGX agent

arXiv:2604.13252v1 Announce Type: new Abstract: Anomaly detection aims to identify observations that deviate from expected behavior. Because anomalous events are inherently sparse, most frameworks are

Parameter-efficient Quantum Multi-task Learning

Model ReleasesDGX agent

arXiv:2604.13560v1 Announce Type: new Abstract: Multi-task learning (MTL) improves generalization and data efficiency by jointly learning related tasks through shared representations. In the widely us

Parameter-Free Non-Ergodic Extragradient Algorithms for Solving Monotone Variational Inequalities

Model ReleasesDGX agent

arXiv:2604.07662v2 Announce Type: replace-cross Abstract: Monotone variational inequalities (VIs) provide a unifying framework for convex minimization, equilibrium computation, and convex-concave sadd

Parameter Importance is Not Static: Evolving Parameter Isolation for Supervised Fine-Tuning

Model ReleasesDGX agent

arXiv:2604.14010v1 Announce Type: cross Abstract: Supervised Fine-Tuning (SFT) of large language models often suffers from task interference and catastrophic forgetting. Recent approaches alleviate th

PatchPoison: Poisoning Multi-View Datasets to Degrade 3D Reconstruction

Model ReleasesDGX agent

arXiv:2604.13153v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has recently enabled highly photorealistic 3D reconstruction from casually captured multi-view images. However, this access

PBE-UNet: A light weight Progressive Boundary-Enhanced U-Net with Scale-Aware Aggregation for Ultrasound Image Segmentation

Model ReleasesDGX agent

arXiv:2604.13791v1 Announce Type: new Abstract: Accurate lesion segmentation in ultrasound images is essential for preventive screening and clinical diagnosis, yet remains challenging due to low contr

Peer-Predictive Self-Training for Language Model Reasoning

Model ReleasesDGX agent

arXiv:2604.13356v1 Announce Type: new Abstract: Mechanisms for continued self-improvement of language models without external supervision remain an open challenge. We propose Peer-Predictive Self-Trai

PersonaVLM: Long-Term Personalized Multimodal LLMs

Model ReleasesDGX agent

arXiv:2604.13074v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) serve as daily assistants for millions. However, their ability to generate responses aligned with individual pr

Predicting Time Pressure of Powered Two-Wheeler Riders for Proactive Safety Interventions

Model ReleasesDGX agent

arXiv:2601.03173v2 Announce Type: replace Abstract: Time pressure critically influences risky maneuvers and crash proneness among powered two-wheeler riders, yet its prediction remains underexplored i

Qwen 3.6 is here, and open-source! Run it locally with improved agentic coding capabilities. Try it with Claude Code: ollama launch claude -…

Model ReleasesDGX agent

Qwen 3.6 is here, and open-source! Run it locally with improved agentic coding capabilities. Try it with Claude Code: ollama launch claude --model qwen3.6 Try it with OpenClaw: ollama launch openclaw

Qwen3.6-35B-A3B on my laptop drew me a better pelican than Claude Opus 4.7

Model ReleasesDGX agent

For anyone who has been (inadvisably) taking my pelican riding a bicycle benchmark seriously as a robust way to test models, here are pelicans from this morning's two big model releases - Qwen3.6-35B-

RAG or Learning? Understanding the Limits of LLM Adaptation under Continuous Knowledge Drift in the Real World

Model ReleasesDGX agent

arXiv:2604.05096v2 Announce Type: replace Abstract: Large language models (LLMs) acquire most of their knowledge during pretraining, which ties them to a fixed snapshot of the world and makes adaptati

RANDPOL: Parameter-Efficient End-to-End Quadruped Locomotion via Randomized Policy Learning

Model ReleasesDGX agent

arXiv:2505.19054v2 Announce Type: replace Abstract: Modern learning-based locomotion controllers typically rely on fully trainable deep neural networks with a large number of parameters. This paper st

Rare Event Analysis via Stochastic Optimal Control

Model ReleasesDGX agent

arXiv:2604.13213v1 Announce Type: cross Abstract: Rare events such as conformational changes in biomolecules, phase transitions, and chemical reactions are central to the behavior of many physical sys

ReConText3D: Replay-based Continual Text-to-3D Generation

Model ReleasesDGX agent

arXiv:2604.13730v1 Announce Type: new Abstract: Continual learning enables models to acquire new knowledge over time while retaining previously learned capabilities. However, its application to text-t

Red Skills or Blue Skills? A Dive Into Skills Published on ClawHub

Model ReleasesDGX agent

arXiv:2604.13064v1 Announce Type: new Abstract: Skill ecosystems have emerged as an increasingly important layer in Large Language Model (LLM) agent systems, enabling reusable task packaging, public d

Replit Agent 4 is even smarter now with Claude Opus 4.7! 50% off for a limited time. Go try it now ↓

Model ReleasesDGX agent

Replit Agent 4 has been upgraded to use Claude Opus 4.7, an advanced AI model, enhancing its code generation and problem-solving capabilities. The company is offering a 50% discount for a limited time

Response.

Model ReleasesDGX agent

Response. Hey Ethan! Sean here, PM on http://Claude.ai - thanks for the feedback. This isn't a router, this is the model being trained to decide when to think based on the context -- we've been runnin

Reward Design for Physical Reasoning in Vision-Language Models

Model ReleasesDGX agent

arXiv:2604.13993v1 Announce Type: cross Abstract: Physical reasoning over visual inputs demands tight integration of visual perception, domain knowledge, and multi-step symbolic inference. Yet even st

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges

Model ReleasesDGX agent

arXiv:2604.13602v1 Announce Type: new Abstract: Reinforcement Learning from Human Feedback (RLHF) and related alignment paradigms have become central to steering large language models (LLMs) and multi

RiskWebWorld: A Realistic Interactive Benchmark for GUI Agents in E-commerce Risk Management

Model ReleasesDGX agent

arXiv:2604.13531v1 Announce Type: cross Abstract: Graphical User Interface (GUI) agents show strong capabilities for automating web tasks, but existing interactive benchmarks primarily target benign,

← Previous
1…341342343344345…372
Next →