AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,292 results
20 Apr 2026

SocialGrid: A Benchmark for Planning and Social Reasoning in Embodied Multi-Agent Systems

Model ReleasesDGX agent

arXiv:2604.16022v1 Announce Type: new Abstract: As Large Language Models (LLMs) transition from text processors to autonomous agents, evaluating their social reasoning in embodied multi-agent settings

Softpick: No Attention Sink, No Massive Activations with Rectified Softmax

Model ReleasesDGX agent

arXiv:2504.20966v4 Announce Type: replace Abstract: We introduce softpick, a rectified, not sum-to-one, drop-in replacement for softmax in transformer attention mechanisms that eliminates attention si

Solving Inverse Parametrized Problems via Finite Elements and Extreme Learning Networks

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2602.14757v2 Announce Type: replace-cross Abstract: We develop an interpolation-based modeling framework for parameter-dependent partial differential equations arising in control, inverse proble

SSFT: A Lightweight Spectral-Spatial Fusion Transformer for Generic Hyperspectral Classification

Model ReleasesDGX agent

arXiv:2604.15828v1 Announce Type: new Abstract: Hyperspectral imaging enables fine-grained recognition of materials by capturing rich spectral signatures, but learning robust classifiers is challengin

Stargazer: A Scalable Model-Fitting Benchmark Environment for AI Agents under Astrophysical Constraints

Model ReleasesDGX agent

arXiv:2604.15664v1 Announce Type: new Abstract: The rise of autonomous AI agents suggests that dynamic benchmark environments with built-in feedback on scientifically grounded tasks are needed to eval

Stein Variational Black-Box Combinatorial Optimization

Model ReleasesDGX agent

arXiv:2604.15837v1 Announce Type: new Abstract: Combinatorial black-box optimization in high-dimensional settings demands a careful trade-off between exploiting promising regions of the search space a

Stochasticity in Tokenisation Improves Robustness

Model ReleasesDGX agent

arXiv:2604.16037v1 Announce Type: new Abstract: The widespread adoption of large language models (LLMs) has increased concerns about their robustness. Vulnerabilities in perturbations of tokenisation

Subjective and Objective Quality-of-Experience Evaluation Study for Live Video Streaming

Model ReleasesDGX agent

arXiv:2409.17596v2 Announce Type: replace-cross Abstract: In recent years, live video streaming has gained widespread popularity across various social media platforms. Quality of experience (QoE), whi

SwanNLP at SemEval-2026 Task 5: An LLM-based Framework for Plausibility Scoring in Narrative Word Sense Disambiguation

Model ReleasesDGX agent

arXiv:2604.16262v1 Announce Type: new Abstract: Recent advances in language models have substantially improved Natural Language Understanding (NLU). Although widely used benchmarks suggest that Large

TabularMath: Understanding Math Reasoning over Tables with Large Language Models

Model ReleasesDGX agent

arXiv:2505.19563v4 Announce Type: replace Abstract: Mathematical reasoning has long been a key benchmark for evaluating large language models. Although substantial progress has been made on math word

the Codex x @skybysoftware acquisition may have been one of the best @openai deals made in the last year. I've been waiting for 'real' compu…

Model ReleasesDGX agent

the Codex x @skybysoftware acquisition may have been one of the best @openai deals made in the last year. I've been waiting for 'real' computer use since @romainhuet demoed the ChatGPT App with 4o Vis

The Illusion of Equivalence: Systematic FP16 Divergence in KV-Cached Autoregressive Inference

Model ReleasesDGX agent

arXiv:2604.15409v1 Announce Type: cross Abstract: KV caching is a ubiquitous optimization in autoregressive transformer inference, long presumed to be numerically equivalent to cache-free computation.

The Jensen + @dwarkesh_sp podcast was fantastic. Jensen is someone who understood how ecosystems work and someone who understands real-world…

Model ReleasesDGX agent

The Jensen + @dwarkesh_sp podcast was fantastic. Jensen is someone who understood how ecosystems work and someone who understands real-world trade, policy and controls work. And in some deeper sense h

The Metacognitive Monitoring Battery: A Cross-Domain Benchmark for LLM Self-Monitoring

Model ReleasesDGX agent

arXiv:2604.15702v1 Announce Type: new Abstract: We introduce a cross-domain behavioural assay of monitoring-control coupling in LLMs, grounded in the Nelson and Narens (1990) metacognitive framework a

The Reasoning Trap: How Enhancing LLM Reasoning Amplifies Tool Hallucination

Model ReleasesDGX agent

arXiv:2510.22977v2 Announce Type: replace-cross Abstract: Enhancing the reasoning capabilities of Large Language Models (LLMs) is a key strategy for building Agents that 'think then act.' However, rec

The Relic Condition: When Published Scholarship Becomes Material for Its Own Replacement

Model ReleasesDGX agent

arXiv:2604.16116v1 Announce Type: cross Abstract: We extracted the scholarly reasoning systems of two internationally prominent humanities and social science scholars from their published corpora alon

The Spectral Geometry of Thought: Phase Transitions, Instruction Reversal, Token-Level Dynamics, and Perfect Correctness Prediction in How Transformers Reason

Model ReleasesDGX agent

arXiv:2604.15350v1 Announce Type: new Abstract: We discover that large language models exhibit spectral phase transitions in their hidden activation spaces when engaging in reasoning versus factual re

.@TheOnion is replacing Alex Jones' Infowars with the comedians behind 'Tim & Eric' and 'Nathan for You' — and finally paying Sandy Hook fam…

Model ReleasesDGX agent

.@TheOnion is replacing Alex Jones' Infowars with the comedians behind 'Tim & Eric' and 'Nathan for You' — and finally paying Sandy Hook families after years of lies. 'What happened to them won't happ

Time to really give Kimi K2.6 a go. Thank you @ollama! Love the ollama cloud setup!

Model ReleasesDGX agent

Time to really give Kimi K2.6 a go. Thank you @ollama! Love the ollama cloud setup! Kimi K2.6 raises the bar for open-source models. 🦙 available on Ollama's cloud! Try it with OpenClaw: ollama launch

Towards Realistic Open-Vocabulary Remote Sensing Segmentation: Benchmark and Baseline

Model ReleasesDGX agent

arXiv:2604.15652v1 Announce Type: new Abstract: Open-vocabulary remote sensing image segmentation (OVRSIS) remains underexplored due to fragmented datasets, limited training diversity, and the lack of

Towards Universal Convergence of Backward Error in Linear System Solvers

Model ReleasesDGX agent

arXiv:2604.16075v1 Announce Type: cross Abstract: The quest for an algorithm that solves an nimes n linear system in O(n^2) time complexity, or O(n^2 ext{poly}(1/epsilon)) when solving up to epsilon r

Transformer Neural Processes - Kernel Regression

Model ReleasesDGX agent

arXiv:2411.12502v4 Announce Type: replace-cross Abstract: Neural Processes (NPs) are a rapidly evolving class of models designed to directly model the posterior predictive distribution of stochastic p

TRIDENT: Enhancing Large Language Model Safety with Tri-Dimensional Diversified Red-Teaming Data Synthesis

Model ReleasesDGX agent

arXiv:2505.24672v2 Announce Type: replace Abstract: Large Language Models (LLMs) excel in various natural language processing tasks but remain vulnerable to generating harmful content or being exploit

Try it out

Model ReleasesDGX agent

Try it out Anyone telling you that you need to use Claude to run just about any OpenClaw spending a wacky 1000 or more per month is no expert you should trust. Use @Grok. Great pricing and 4.3 is in a

TwinTrack: Post-hoc Multi-Rater Calibration for Medical Image Segmentation

Model ReleasesDGX agent

arXiv:2604.15950v1 Announce Type: new Abstract: Pancreatic ductal adenocarcinoma (PDAC) segmentation on contrast-enhanced CT is inherently ambiguous: inter-rater disagreement among experts reflects ge

Two-Dimensional Deep ReLU CNN Approximation for Korobov Functions: A Constructive Approach

Model ReleasesDGX agent

arXiv:2503.07976v2 Announce Type: replace-cross Abstract: This paper investigates approximation capabilities of two-dimensional (2D) deep convolutional neural networks (CNNs), with Korobov functions s

TwoHamsters: Benchmarking Multi-Concept Compositional Unsafety in Text-to-Image Models

Model ReleasesDGX agent

arXiv:2604.15967v1 Announce Type: cross Abstract: Despite the remarkable synthesis capabilities of text-to-image (T2I) models, safeguarding them against content violations remains a persistent challen

UniEditBench: A Unified and Cost-Effective Benchmark for Image and Video Editing via Distilled MLLMs

Model ReleasesDGX agent

arXiv:2604.15871v1 Announce Type: cross Abstract: The evaluation of visual editing models remains fragmented across methods and modalities. Existing benchmarks are often tailored to specific paradigms

VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effects

Model ReleasesDGX agent

arXiv:2604.16272v1 Announce Type: cross Abstract: As AI-assisted video creation becomes increasingly practical, instruction-guided video editing has become essential for refining generated or captured

vla-eval: A Unified Evaluation Harness for Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2603.13966v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models are increasingly evaluated across multiple simulation benchmarks, yet adding each benchmark to an evaluation pip

VLegal-Bench: Cognitively Grounded Benchmark for Vietnamese Legal Reasoning of Large Language Models

Model ReleasesDGX agent

arXiv:2512.14554v5 Announce Type: replace-cross Abstract: The rapid advancement of large language models (LLMs) has enabled new possibilities for applying artificial intelligence within the legal doma

Was happy with Gemma 4 Cloud, but had to change due to API Errors, GLM 5.1 spends a lot more ressources

Model ReleasesDGX agent

A user reported satisfaction with Gemma 4 Cloud but switched to GLM 5.1 due to API errors, noting that the alternative model consumes significantly more resources. The post likely discusses the perfor

Was super interesting to chat with one of @cohere's senior PMs about its AI speech-to-text model, which it hopes to incorporate into its Nor…

Model ReleasesDGX agent

Was super interesting to chat with one of @cohere's senior PMs about its AI speech-to-text model, which it hopes to incorporate into its North platform soon. .@Cohere's Cassie Cao takes us behind the

Watching Movies Like a Human: Egocentric Emotion Understanding for Embodied Companions

Model ReleasesDGX agent

arXiv:2604.15823v1 Announce Type: new Abstract: Embodied robotic agents often perceive movies through an egocentric screen-view interface rather than native cinematic footage, introducing domain shift

We have been named one of the 40 Most Innovative AI-Native Prosumer Companies by @notablecap. The Prosumer AI 40 recognizes companies buildi…

Model ReleasesDGX agent

We have been named one of the 40 Most Innovative AI-Native Prosumer Companies by @notablecap. The Prosumer AI 40 recognizes companies building the tools that blur the line between professional and con

What a week in SF for Gemma! 💎 - Gemma 4 SF Meetup, with Unsloth, MLX, VLLM, NousResearch, and Cactus Compute - Hackathon at Y Combinator w…

Model ReleasesDGX agent

What a week in SF for Gemma! 💎 - Gemma 4 SF Meetup, with Unsloth, MLX, VLLM, NousResearch, and Cactus Compute - Hackathon at Y Combinator with Cactus Compute - Ollama Gemma 4 Meetup, with SGLang - NVI

What to expect during the Phi Moments @ Next event: Join theCUBE April 22-23

Model ReleasesDGX agent

Advances in the use of artificial intelligence by major providers such as Google Cloud are not confined solely to the hyperscaler’s own customers. Partners, such as the AI-first digital engineering co

When Cultures Meet: Multicultural Text-to-Image Generation

Model ReleasesDGX agent

arXiv:2502.15972v2 Announce Type: replace-cross Abstract: Text-to-image generation models have achieved strong performance in culturally homogeneous settings, yet their ability to generate multicultur

When Surfaces Lie: Exploiting Wrinkle-Induced Attention Shift to Attack Vision-Language Models

Model ReleasesDGX agent

arXiv:2603.27759v3 Announce Type: replace Abstract: Visual-Language Models (VLMs) have demonstrated exceptional cross-modal understanding across various tasks, including zero-shot classification, imag

Why Fine-Tuning Encourages Hallucinations and How to Fix It

Model ReleasesDGX agent

arXiv:2604.15574v1 Announce Type: cross Abstract: Large language models are prone to hallucinating factually incorrect statements. A key source of these errors is exposure to new factual information t

Wisdom is Knowing What not to Say: Hallucination-Free LLMs Unlearning via Attention Shifting

Model ReleasesDGX agent

arXiv:2510.17210v3 Announce Type: replace Abstract: The increase in computing power and the necessity of AI-assisted decision-making boost the growing application of large language models (LLMs). Alon

Wordle 1,765 3/6 🟩🟩🟩⬛⬛ ⬛⬛🟨⬛⬛ 🟩🟩🟩🟩🟩

Model ReleasesDGX agent

This post documents a Wordle game result where the player solved puzzle #1,765 in 3 attempts, with the emoji grid showing the letter placement feedback for each guess. The final guess revealed all fiv

Wordle 1,766 4/6 ⬛⬛🟩⬛🟩 ⬛⬛⬛⬛⬛ ⬛🟨🟩⬛🟩 🟩🟩🟩🟩🟩

Model ReleasesDGX agent

This post shows Anthropic's Wordle game result for puzzle #1,766, solved in 4 attempts using a specific sequence of guesses indicated by the emoji color codes (gray for incorrect letters, yellow for c

19 Apr 2026

An obvious way to release Mythos class models with uncertain autonomous ability is to make them only available on the website, like Gemini D…

Model ReleasesDGX agent

An obvious way to release Mythos class models with uncertain autonomous ability is to make them only available on the website, like Gemini Deep Think or ChatGPT Pro. Minimal risk of being used for aut

Anthropic locked Claude Code to native apps in Jan 2026. Are we still comparing models or just ecosystems

Model ReleasesDGX agent

Anthropic's Claude Code desktop app is strictly optimized for Anthropic's models , creating a 'walled garden' effect that restricts users to Claude exclusively. The redesigned Claude Code desktop app

Appfigures: app releases across the App Store and Google Play grew 60% YoY in Q1, with App Store releases alone up 80%, possibly driven by AI coding tools (Sarah Perez/TechCrunch)

Model ReleasesDGX agent

Sarah Perez / TechCrunch: Appfigures: app releases across the App Store and Google Play grew 60% YoY in Q1, with App Store releases alone up 80%, possibly driven by AI coding tools — Everyone said AI

🔴 BREAKING Speculative decoding just got real for local LLMs. llama.cpp merged speculative checkpointing (PR #19493). Expect 0-50% speedups…

Model ReleasesDGX agent

🔴 BREAKING Speculative decoding just got real for local LLMs. llama.cpp merged speculative checkpointing (PR #19493). Expect 0-50% speedups on coding tasks with `--spec-type ngram-mod`. This means: •

Falcon 9 launches 25 @Starlink satellites from California ahead of completing the 600th overall landing of an orbital class rocket

Model ReleasesDGX agent

SpaceX launched 25 Starlink satellites aboard a Falcon 9 rocket from California, marking a significant milestone as the mission represented the 600th successful landing of an orbital-class rocket by t

I gave two MoE models the same vibe coding challenge Qwen3.6 35B A3B (31.8GB) vs Gemma4 26B A4B (23.3GB) Stack: > Unsloth Q6_K_XL > llama.cp…

Model ReleasesDGX agent

I gave two MoE models the same vibe coding challenge Qwen3.6 35B A3B (31.8GB) vs Gemma4 26B A4B (23.3GB) Stack: > Unsloth Q6_K_XL > llama.cpp > Model-card recommended sampling for each 4 prompts, side

Introducing OpenMythos An open-source, first-principles theoretical reconstruction of Claude Mythos, implemented in PyTorch. The architectur…

Model ReleasesDGX agent

Introducing OpenMythos An open-source, first-principles theoretical reconstruction of Claude Mythos, implemented in PyTorch. The architecture instantiates a looped transformer with a Mixture-of-Expert

Just got access to Grok 4.3 and the first thing I did was throw the hardest task I usually run on claude to test it gave it a complex resear…

Model ReleasesDGX agent

Just got access to Grok 4.3 and the first thing I did was throw the hardest task I usually run on claude to test it gave it a complex research prompt on deepseek: >thought for 13 mins and 17 seconds >

LiteParse is the best model-free, open-source document parser for AI agents. It now gets a first-class landing page on our website 💫 Our co…

Model ReleasesDGX agent

LiteParse is the best model-free, open-source document parser for AI agents. It now gets a first-class landing page on our website 💫 Our company mission is building the world's best agentic document p

Make neural network cells inside a “Digital Petri Dish” fight for control and dominance in a web browser tab. https://pub.sakana.ai/digital-…

Model ReleasesDGX agent

Make neural network cells inside a “Digital Petri Dish” fight for control and dominance in a web browser tab. https://pub.sakana.ai/digital-ecosystem/ Media What happens when you put competing neural

Note to @AnthropicAI - much as I appreciate the public system prompts this would be so much more valuable to me as a Claude power user if yo…

Model ReleasesDGX agent

Note to @AnthropicAI - much as I appreciate the public system prompts this would be so much more valuable to me as a Claude power user if you published the tool descriptions as well Since Anthropic pu

OMG. Let’s get one thing straight. Claude doesn’t get anxious. It mimics people who get anxious. Those two things are NOT the same. My head …

Model ReleasesDGX agent

OMG. Let’s get one thing straight. Claude doesn’t get anxious. It mimics people who get anxious. Those two things are NOT the same. My head is shaking so much I need medical attention. anthropic's in-

Only Grok 4.3 lets me drive my car to get gas. ChatGPT, Claude, and Gemini want me to walk.

Model ReleasesDGX agent

This post compares AI assistants' responses to a request about driving to get gas, claiming that Grok 4.3 provides the requested information while ChatGPT, Claude, and Gemini decline or suggest altern

releasing qwen3.6 35b a3b carnice edition (inspired by @kaiostephens) to enjoy an hermes agent tailored finetuned on unified memory hardware…

Model ReleasesDGX agent

releasing qwen3.6 35b a3b carnice edition (inspired by @kaiostephens) to enjoy an hermes agent tailored finetuned on unified memory hardware! safetensors and gguf versions (q8, 6,5 and 4) available! f

Since Anthropic publish their system prompts we can generate a diff between Claude Opus 4.6 and 4.7 - here are my notes on what's changed ht…

Model ReleasesDGX agent

Simon Willison documents the differences between Anthropic's Claude Opus 4.6 and 4.7 system prompts, analyzing changes that Anthropic made public. The notes likely highlight modifications to model beh

Source tweet is here: https://x.com/jerryjliu0/status/2044902620746363016?s=20

Model ReleasesDGX agent

Source tweet is here: https://x.com/jerryjliu0/status/2044902620746363016?s=20 We comprehensively benchmarked Opus 4.7 on document understanding. We evaluated it through ParseBench - our comprehensive

Sources: the glowing '26' in Apple's WWDC invite is teasing a revamped Siri, memory shortages may push Mac Studio and touch MacBook Pro launches by a few months (Mark Gurman/Bloomberg)

Model ReleasesDGX agent

Mark Gurman / Bloomberg: Sources: the glowing “26” in Apple's WWDC invite is teasing a revamped Siri, memory shortages may push Mac Studio and touch MacBook Pro launches by a few months — Also: Memory

← Previous
1…335336337338339…372
Next →