AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,292 results
19 Apr 2026

Thanks for the shoutout re: ParseBench! 📑 🙂 Opus 4.7 is definitely a step up from Opus 4.6 on document understanding capabilities. Great t…

Model ReleasesDGX agent

Thanks for the shoutout re: ParseBench! 📑 🙂 Opus 4.7 is definitely a step up from Opus 4.6 on document understanding capabilities. Great to see labs start to push a bit more on this front. [16 Apr 202

The biggest challenge of using chat-based AI systems is that the details of what they can do are invisible - those tool descriptions are the…

Model ReleasesDGX agent

The biggest challenge of using chat-based AI systems is that the details of what they can do are invisible - those tool descriptions are the missing manual, publishing them would be a huge benefit to

The continuing gap between the capabilities of Gemini Pro 3.1 (very good model) and the capabilities of the Gemini app/website is odd. The m…


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

The continuing gap between the capabilities of Gemini Pro 3.1 (very good model) and the capabilities of the Gemini app/website is odd. The model can do what Claude/GPT can do, but there is a minimal h

The Top AI Papers of the Week (April 13 - 19) - AlphaEval - AiScientist - Auto-Diagnose - Nemotron 3 Super - Subliminal Learning - Automated…

Model ReleasesDGX agent

The Top AI Papers of the Week (April 13 - 19) - AlphaEval - AiScientist - Auto-Diagnose - Nemotron 3 Super - Subliminal Learning - Automated W2S Researcher - Memory Transfer Learning Read on for more:

Who is running local models on GPUs on OpenClaw? I have started benchmarking different models this week. I am working on improving model sel…

Model ReleasesDGX agent

Who is running local models on GPUs on OpenClaw? I have started benchmarking different models this week. I am working on improving model selection and switching UX on OpenClaw, i.e. I run /model vllm/

Wordle 1,764 3/6 ⬛🟨🟩⬛⬛ ⬛⬛⬛⬛🟨 🟩🟩🟩🟩🟩

Model ReleasesDGX agent

This post documents a Wordle game result where the player solved puzzle #1,764 in three attempts, using the standard color-coded emoji system (black for incorrect letters, yellow for correct letters i

You ever go on Huggingface and see: - GGUF - Unsloth - Llama.cpp - Dynamic GGUF - Q_4_M / IQ_4XL etc. Here's what's going on under the hood.…

Model ReleasesDGX agent

This post explains the technical details behind common terms and tools encountered on Hugging Face for running large language models locally, including quantization formats (GGUF, Q_4_M, IQ_4XL), opti

You should watch this. It just shows how disconnected we are from the small group of people making decisions that will impact our future hea…

Model ReleasesDGX agent

You should watch this. It just shows how disconnected we are from the small group of people making decisions that will impact our future heavily. These people have so much ai psychosis. If you listen

18 Apr 2026

A downside with using VLMs to parse PDFs is guaranteeing that the output text is *correct* and output in the correct reading order. 1️⃣ Text…

Model ReleasesDGX agent

A downside with using VLMs to parse PDFs is guaranteeing that the output text is *correct* and output in the correct reading order. 1️⃣ Text correctness: making sure that digits, words, sentences are

Adding a new content type to my blog-to-newsletter tool

Model ReleasesDGX agent

Agentic Engineering Patterns > Here's an example of a deceptively short prompt that got a lot of work done in a single shot. First, some background. I send out a free Substack newsletter around once a

Airbnb launches a pilot in NYC, LA, and other cities that lets users to select from a range of boutique hotels alongside private homes in a bid to boost growth (Stephanie Stacey/Financial Times)

Model ReleasesDGX agent

Stephanie Stacey / Financial Times: Airbnb launches a pilot in NYC, LA, and other cities that lets users to select from a range of boutique hotels alongside private homes in a bid to boost growth — Co

Anthropic shut down an entire company's Claude access overnight 60+ employees. No explanation. Just an email. Want to appeal? Fill out a Goo…

Model ReleasesDGX agent

Anthropic shut down an entire company's Claude access overnight 60+ employees. No explanation. Just an email. Want to appeal? Fill out a Google Form. Integrations gone. Histories gone. Everything buil

Anthropic’s CEO keeps talking about AI wiping out jobs because he’s trying to IPO this year. If he positions Claude as armageddon for jobs, …

Model ReleasesDGX agent

Anthropic’s CEO keeps talking about AI wiping out jobs because he’s trying to IPO this year. If he positions Claude as armageddon for jobs, his TAM becomes “all white-collar human labor,” not just AI

Both Gemini and Grok are underrated. I used to champion Gemini for a while, but lately I've been very happy with Grok, especially since the …

Model ReleasesDGX agent

Both Gemini and Grok are underrated. I used to champion Gemini for a while, but lately I've been very happy with Grok, especially since the Grok 4.20 release with multi-agent. And now we have Grok 4.3

Claude system prompts as a git timeline

Model ReleasesDGX agent

Research: Claude system prompts as a git timeline Anthropic publish the system prompts for Claude chat and make that page available as Markdown. I had Claude Code turn that page into separate files fo

Fantastic to see GLM being applied to such fresh, dynamic scenarios.

Model ReleasesDGX agent

Fantastic to see GLM being applied to such fresh, dynamic scenarios. Doing some stress tests on OpenMAIC’s Interactive Simulation with a DNA Replication case. 💻 Both powered by @Zai_org — with GLM-5.1

GDPval is one of the most important benchmarks of AI ability because it is based on human expertise. It compares expert human performance to…

Model ReleasesDGX agent

GDPval is one of the most important benchmarks of AI ability because it is based on human expertise. It compares expert human performance to AI performance using expert human judges who spend an avera

Hermes Agent is model & tool backend agnostic for a reason, everyone should have access to AI. We don't dictate the rules of use for your ag…

Model ReleasesDGX agent

Hermes Agent is model & tool backend agnostic for a reason, everyone should have access to AI. We don't dictate the rules of use for your agent, YOU do Anthropic shut down an entire company's Claude a

I prefer my design tool to be more closely integrated with where my agents work. I spent a few hours building my own design tool (inspired b…

Model ReleasesDGX agent

I prefer my design tool to be more closely integrated with where my agents work. I spent a few hours building my own design tool (inspired by Claude Design) inside my orchestrator. I can use this with

I think Artificial Analysis does a good job overall and provides transparency in benchmarking, but GDPval-AA is not a good benchmark and nee…

Model ReleasesDGX agent

I think Artificial Analysis does a good job overall and provides transparency in benchmarking, but GDPval-AA is not a good benchmark and needs to stop being reported. It is Gemini 3.1 judging other mo

Mistral, which once aimed for top open models, now leans on being an alternative to Chinese and US labs, says it's on track for $80M in monthly revenue by Dec. (Iain Martin/Forbes)

Model ReleasesDGX agent

Iain Martin / Forbes: Mistral, which once aimed for top open models, now leans on being an alternative to Chinese and US labs, says it's on track for $80M in monthly revenue by Dec. — Paris-based Mist

Okay this one is insane. A new 18B frankenstein model was just released on @huggingface — Beats the new Qwen3.6-35B-A3B on a 44-test suite d…

Model ReleasesDGX agent

Okay this one is insane. A new 18B frankenstein model was just released on @huggingface — Beats the new Qwen3.6-35B-A3B on a 44-test suite despite requiring 12GB VRAM instead of 24GB 🤯 Runs on a SINGL

RIP Anthropic API fees. Someone just made Claude Code run 100% locally on a MacBook for $0/month. 122B parameter model. 65 tokens per second…

Model ReleasesDGX agent

RIP Anthropic API fees. Someone just made Claude Code run 100% locally on a MacBook for $0/month. 122B parameter model. 65 tokens per second. Nothing touches the cloud. The trick everyone else missed:

Salesforce announces Headless 360, an initiative that will give AI agents access to Salesforce's platform capabilities through APIs, MCP tools or CLI commands (Michael Nuñez/VentureBeat)

Model ReleasesDGX agent

Michael Nuñez / VentureBeat: Salesforce announces Headless 360, an initiative that will give AI agents access to Salesforce's platform capabilities through APIs, MCP tools or CLI commands — Salesforce

Stunning views and just as jaw-dropping of presentations at the @cerebral_valley Gemma 4 Launch Party tonight! @vllm_project @UnslothAI @Oll…

Model ReleasesDGX agent

Stunning views and just as jaw-dropping of presentations at the @cerebral_valley Gemma 4 Launch Party tonight! @vllm_project @UnslothAI @Ollama @huggingface @cactuscompute @apple @nvidia @pytorch and

We push Prefill/Decode disaggregation beyond a single cluster: cross-datacenter + heterogeneous hardware, unlocking the potential for signif…

Model ReleasesDGX agent

We push Prefill/Decode disaggregation beyond a single cluster: cross-datacenter + heterogeneous hardware, unlocking the potential for significantly lower cost per token. This was previously blocked by

Wordle 1,763 5/6 ⬛⬛⬛⬛🟩 🟨⬛⬛⬛⬛ ⬛🟨⬛⬛🟩 ⬛⬛⬛🟩🟩 🟩🟩🟩🟩🟩

Model ReleasesDGX agent

This post documents a Wordle game result where the player solved puzzle #1,763 in 5 attempts, with the final answer being a five-letter word. The emoji grid shows the progression of guesses, indicatin

17 Apr 2026

30K+ likes in the first hour. 🤯 That is crazy! Design is unsolved with agents. But lots of impactful work generated by agents is around des…

Model ReleasesDGX agent

30K+ likes in the first hour. 🤯 That is crazy! Design is unsolved with agents. But lots of impactful work generated by agents is around design. Claude Design is Anthropic's way of saying that they are

3D Instruction Ambiguity Detection

Model ReleasesDGX agent

arXiv:2601.05991v2 Announce Type: replace Abstract: In safety-critical domains, linguistic ambiguity can have severe consequences; a vague command like 'Pass me the vial' in a surgical setting could l

A multi-platform LiDAR dataset for standardized forest inventory measurement at long term ecological monitoring sites

Model ReleasesDGX agent

arXiv:2604.14635v1 Announce Type: new Abstract: We present a curated multi-platform LiDAR reference dataset from an instrumented ICOS forest plot, explicitly designed to support calibration, benchmark

A Nonlinear Separation Principle: Applications to Neural Networks, Control and Learning

Model ReleasesDGX agent

arXiv:2604.15238v1 Announce Type: cross Abstract: This paper investigates continuous-time and discrete-time firing-rate and Hopfield recurrent neural networks (RNNs), with applications in nonlinear co

AccelOpt: A Self-Improving LLM Agentic System for AI Accelerator Kernel Optimization

Model ReleasesDGX agent

arXiv:2511.15915v2 Announce Type: replace-cross Abstract: We present AccelOpt, a self-improving large language model (LLM) agentic system that autonomously optimizes kernels for emerging AI acclerator

Acceptance Dynamics Across Cognitive Domains in Speculative Decoding

Model ReleasesDGX agent

arXiv:2604.14682v1 Announce Type: cross Abstract: Speculative decoding accelerates large language model (LLM) inference. It uses a small draft model to propose a tree of future tokens. A larger target

Active Learning with Selective Time-Step Acquisition for PDEs

Model ReleasesDGX agent

arXiv:2511.18107v2 Announce Type: replace Abstract: Accurately solving partial differential equations (PDEs) is critical to understanding complex scientific and engineering phenomena, yet traditional

AD4AD: Benchmarking Visual Anomaly Detection Models for Safer Autonomous Driving

Model ReleasesDGX agent

arXiv:2604.15291v1 Announce Type: new Abstract: The reliability of a machine vision system for autonomous driving depends heavily on its training data distribution. When a vehicle encounters significa

ADAPT: Benchmarking Commonsense Planning under Unspecified Affordance Constraints

Model ReleasesDGX agent

arXiv:2604.14902v1 Announce Type: cross Abstract: Intelligent embodied agents should not simply follow instructions, as real-world environments often involve unexpected conditions and exceptions. Howe

Adaptive Test-Time Compute Allocation for Reasoning LLMs via Constrained Policy Optimization

Model ReleasesDGX agent

arXiv:2604.14853v1 Announce Type: new Abstract: Test-time compute scaling, the practice of spending extra computation during inference via repeated sampling, search, or extended reasoning, has become

AgentGA: Evolving Code Solutions in Agent-Seed Space

Model ReleasesDGX agent

arXiv:2604.14655v1 Announce Type: cross Abstract: We present AgentGA, a framework that evolves autonomous code-generation runs by optimizing the agent seed: the task prompt plus optional parent archiv

AgentIAD: Agentic Industrial Anomaly Detection via Adaptive Memory Augmentation

Model ReleasesDGX agent

arXiv:2512.13671v2 Announce Type: replace Abstract: Industrial anomaly detection (IAD) is challenging due to the subtle and highly localized nature of many defects, which single-pass vision--language

AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot

Model ReleasesDGX agent

arXiv:2604.13940v1 Announce Type: new Abstract: Scientific peer review faces mounting strain as submission volumes surge, making it increasingly difficult to sustain review quality, consistency, and t

[AINews] Anthropic Claude Opus 4.7 - literally one step better than 4.6 in every dimension

Model ReleasesDGX agent

This article from Latent Space discusses Anthropic's Claude Opus 4.7 release, highlighting incremental improvements across multiple performance dimensions compared to the previous 4.6 version. The pie

AnimationBench: Are Video Models Good at Character-Centric Animation?

Model ReleasesDGX agent

arXiv:2604.15299v1 Announce Type: new Abstract: Video generation has advanced rapidly, with recent methods producing increasingly convincing animated results. However, existing benchmarks-largely desi

Anonpsy: A Graph-Based Framework for Structure-Preserving De-identification of Psychiatric Narratives

Model ReleasesDGX agent

arXiv:2601.13503v2 Announce Type: replace Abstract: Psychiatric narratives encode patient identity not only through explicit identifiers but also through idiosyncratic life events embedded in their cl

Anthropic launches Claude Design to speed up graphic design projects

Model ReleasesDGX agent

The latest addition to Anthropic PBC’s product portfolio is Claude Design, a tool that enables users to generate visual assets with prompts. The company launched the offering into public preview today

Anthropic’s new cybersecurity model could get it back in the government’s good graces

Model ReleasesDGX agent

The Trump administration has spent nearly two months fighting with AI company Anthropic. It's dubbed the company a 'RADICAL LEFT, WOKE COMPANY' full of 'Leftwing nut jobs' and a menace to national sec

Applying an Agentic Coding Tool for Improving Published Algorithm Implementations

Model ReleasesDGX agent

arXiv:2604.13109v1 Announce Type: cross Abstract: We present a two-stage pipeline for AI-assisted improvement of published algorithm implementations. In the first stage, a large language model with re

Assessment Design in the AI Era: A Method for Identifying Items Functioning Differentially for Humans and Chatbots

Model ReleasesDGX agent

arXiv:2603.23682v2 Announce Type: replace-cross Abstract: The rapid adoption of large language models (LLMs) in education raises profound challenges for assessment design. To adapt assessments to the

AutoRAN: Automated Hijacking of Safety Reasoning in Large Reasoning Models

Model ReleasesDGX agent

arXiv:2505.10846v3 Announce Type: replace Abstract: This paper presents AutoRAN, the first framework to automate the hijacking of internal safety reasoning in large reasoning models (LRMs). At its cor

Auxiliary Finite-Difference Residual-Gradient Regularization for PINNs

Model ReleasesDGX agent

arXiv:2604.14472v1 Announce Type: new Abstract: Physics-informed neural networks (PINNs) are often selected by a single scalar loss even when the quantity of interest is more specific. We study a hybr

Benchmarking Classical Coverage Path Planning Heuristics on Irregular Hexagonal Grids for Maritime Coverage Scenarios

Model ReleasesDGX agent

arXiv:2604.15202v1 Announce Type: new Abstract: Coverage path planning on irregular hexagonal grids is relevant to maritime surveillance, search and rescue and environmental monitoring, yet classical

Benchmarking Linguistic Adaptation in Comparable-Sized LLMs: A Study of Llama-3.1-8B, Mistral-7B-v0.1, and Qwen3-8B on Romanized Nepali

Model ReleasesDGX agent

arXiv:2604.14171v1 Announce Type: new Abstract: Romanized Nepali, the Nepali language written in the Latin alphabet, is the dominant medium for informal digital communication in Nepal, yet it remains

Benchmarking Optimizers for MLPs in Tabular Deep Learning

Model ReleasesDGX agent

arXiv:2604.15297v1 Announce Type: new Abstract: MLP is a heavily used backbone in modern deep learning (DL) architectures for supervised learning on tabular data, and AdamW is the go-to optimizer used

Best of both worlds: Stochastic & adversarial best-arm identification

Model ReleasesDGX agent

arXiv:2604.14860v1 Announce Type: cross Abstract: We study bandit best-arm identification with arbitrary and potentially adversarial rewards. A simple random uniform learner obtains the optimal rate o

BiCon-Gate: Consistency-Gated De-colloquialisation for Dialogue Fact-Checking

Model ReleasesDGX agent

arXiv:2604.14389v1 Announce Type: new Abstract: Automated fact-checking in dialogue involves multi-turn conversations where colloquial language is frequent yet understudied. To address this gap, we pr

Bridging the Gap between Learning and Inference for Diffusion-Based Molecule Generation

Model ReleasesDGX agent

arXiv:2411.05472v2 Announce Type: replace Abstract: The paradigm shift toward structure-driven molecule generation has been propelled by advances in deep generative models, such as variational auto-en

Building Extraction from Remote Sensing Imagery under Hazy and Low-light Conditions: Benchmark and Baseline

Model ReleasesDGX agent

arXiv:2604.15088v1 Announce Type: new Abstract: Building extraction from optical Remote Sensing (RS) imagery suffers from performance degradation under real-world hazy and low-light conditions. Howeve

CaptionQA: Is Your Caption as Useful as the Image Itself?

Model ReleasesDGX agent

arXiv:2511.21025v2 Announce Type: replace Abstract: Image captions serve as efficient surrogates for visual content in multimodal systems such as retrieval, recommendation, and multi-step agentic infe

Catching Every Ripple: Enhanced Anomaly Awareness via Dynamic Concept Adaptation

Model ReleasesDGX agent

arXiv:2604.14726v1 Announce Type: new Abstract: Online anomaly detection (OAD) plays a pivotal role in real-time analytics and decision-making for evolving data streams. However, existing methods ofte

CausalDetox: Causal Head Selection and Intervention for Language Model Detoxification

Model ReleasesDGX agent

arXiv:2604.14602v1 Announce Type: new Abstract: Large language models (LLMs) frequently generate toxic content, posing significant risks for safe deployment. Current mitigation strategies often degrad

CAVERS: Multimodal SLAM Data from a Natural Karstic Cave with Ground Truth Motion Capture

Model ReleasesDGX agent

arXiv:2604.15052v1 Announce Type: new Abstract: Autonomous robots operating in natural karstic caves face perception and navigation challenges that are qualitatively distinct from those encountered in

← Previous
1…336337338339340…372
Next →