AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,092 results
CompaniesToolsTechniques

Each lane shows up to 8 recent matching entries, ordered from earlier to later. Tracks load separately to keep the 75,000+ entry wiki fast.

Companies

CompanyAnthropic8 recent entries
12 Aug 2026Grok 4.6 is now one of the top models in the world for agentic workflows It ranks #1 on the Artificial Analysis Agentic Index, tied with Cla…

Grok 4.6 is now one of the top models in the world for agentic workflows It ranks #1 on the Artificial Analysis Agentic Index, tied with Claude Opus 5 Max Agentic AI is about more than answering quest

→12 Aug 2026Grok 4.6 is an excellent model. I’ve been using it heavily for the past couple of weeks and it handles everything from simple coding & code …

Grok 4.6 is an excellent model. I’ve been using it heavily for the past couple of weeks and it handles everything from simple coding & code review all the way to designing and debugging complex system

3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
→12 Aug 2026Everyone else is talking about building ASI to like monopolize b2b saas and Elon is talking about building a kardashev II sentient sun

Everyone else is talking about building ASI to like monopolize b2b saas and Elon is talking about building a kardashev II sentient sun Media It’s always funny how people in SF twitter bubble will say

→12 Aug 2026DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?

arXiv:2608.10366v1 Announce Type: new Abstract: Real-world data science involves long-horizon workflows that span data wrangling, exploration, modeling, visualization, and validation, and require coor

→12 Aug 2026Do LLM Recommenders Know When They're Hallucinating? Auditing Confidence Calibration in Catalog Faithfulness

arXiv:2608.10008v1 Announce Type: cross Abstract: LLM recommenders for top-K item suggestion regularly emit titles outside the target catalog. Prior audits measure this as a binary out-of-domain rate;

→12 Aug 2026Bayesian-Agent: Posterior-Guided Skill Evolution Across LLM Agent Harnesses

arXiv:2606.08348v2 Announce Type: replace Abstract: LLM agents increasingly rely on prompts, tools, memory, SOPs, skills, and harness feedback, yet current self-evolution pipelines often update these

→12 Aug 2026AIFS-TC: A simple correction competitive with the operational frontier for tropical cyclone intensity forecasting

arXiv:2608.09959v1 Announce Type: cross Abstract: AI weather models are in the process of revolutionising weather forecasting. While these models have been shown to achieve superior performance to phy

→12 Aug 202618 two-word AI prompts I'm kind of obsessed with: 1) now what - great for when you've wrapped up a project or big push and you still have en…

18 two-word AI prompts I'm kind of obsessed with: 1) now what - great for when you've wrapped up a project or big push and you still have energy and want AI to give you more 2) plz fix - usually accom

CompanyOpenAI8 recent entries
11 Aug 2026Curriculum Generation under Structured Parametric Environments for Robust Navigation Policies

arXiv:2608.08545v1 Announce Type: cross Abstract: Robust navigation policies for autonomous agents must generalize across continuously varying environmental conditions such as turn rates, obstacles, f

→11 Aug 2026Claude's watermark probably doesn't work how you think. As the CTO of GPTZero, I'll explain how Anthropic, Google and OpenAI are building te…

Claude's watermark probably doesn't work how you think. As the CTO of GPTZero, I'll explain how Anthropic, Google and OpenAI are building text watermarking in this brief explainer and whether it can b

→11 Aug 2026ChatGPT and Gemini both just passed 1 billion users

For the 14th time, a Google product has hit 1 billion users. Google CEO Sundar Pichai posted on X that a billion people are using Gemini every month, and that Gemini is Google's fastest-growing produc

→11 Aug 2026Can Gemma and Qwen models catch hallucinations by looking at their own logprobs?

Hi! I'm really obsessed with LLM hallucinations for the last 6 days 😭 I started by designing system prompts to attack hallucinations but failed, obviously. Now I tried reading logprobs and... I think

→11 Aug 2026An Expectation-Maximization Perspective on Reinforcement Learning for LLM Reasoning

arXiv:2504.18587v2 Announce Type: replace-cross Abstract: Reinforcement learning has emerged as a powerful approach for improving the reasoning capabilities of large language models, as demonstrated b

→11 Aug 2026⚡️A coalition that secures long-term AI capacity: We’re aggregating long-term compute demand in Europe to determine what capacity is built, …

⚡️A coalition that secures long-term AI capacity: We’re aggregating long-term compute demand in Europe to determine what capacity is built, where it’s located, and whom it serves. Through these multi-

→12 Aug 2026Rethinking LLM Verification: Evidence Structure, Uncertainty, and Selective Refinement

arXiv:2608.10725v1 Announce Type: new Abstract: Large language models (LLMs) often rely on shortcuts rather than systematic reasoning, raising safety concerns in medical applications. Allowing models

→12 Aug 2026Interesting research suggests caution in determining which AI company is winning by looking at any one source.. OpenRouter seems to show ope…

Interesting research suggests caution in determining which AI company is winning by looking at any one source.. OpenRouter seems to show open weights winning over time, but work submitted to Pangram i

CompanyGoogle8 recent entries
12 Aug 2026Sources detail moves behind Google's AI reshuffle; Sergey Brin urged key staff to go all in on Gemini, and some teams shifted from DeepMind to corporate Google (Kenrick Cai/Reuters)

Kenrick Cai / Reuters: Sources detail moves behind Google's AI reshuffle; Sergey Brin urged key staff to go all in on Gemini, and some teams shifted from DeepMind to corporate Google — Google co-found

→12 Aug 2026Situation Graph Prediction for User Perspective Modeling

arXiv:2602.13319v2 Announce Type: replace Abstract: Perspective-aware AI requires modeling evolving internal states---goals, emotions, contexts---not merely preferences. Progress is limited by a data

→12 Aug 2026Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation

arXiv:2608.10812v1 Announce Type: cross Abstract: We study reference-free post-training for multilingual machine translation with open large language models. Starting from the supervised-finetuned MiL

→12 Aug 2026Putting sign language AI into users’ hands

Google DeepMind’s new sign‑language‑to‑text (SL2T) model powers real‑time sign-to-text dictation in Gboard and Live Transcribe on Pixel 11, currently supporting American Sign Language to English. The

→12 Aug 2026Navigation Alone Is Not Enough: Evaluating Explanatory Assistive UI Agents

arXiv:2608.09944v1 Announce Type: cross Abstract: Modern web interfaces are increasingly difficult to use with screen readers, particularly when pages update dynamically or hide important structure be

→12 Aug 2026Google unveils the 899+ Pixel 11, 1,099+ 11 Pro, and $1,299+ 11 Pro XL, with a Tensor G6, new Gemini features, Magic Capture to pick the best frames, and more (Ivan Mehta/TechCrunch)

Ivan Mehta / TechCrunch: Google unveils the 899+ Pixel 11, 1,099+ 11 Pro, and $1,299+ 11 Pro XL, with a Tensor G6, new Gemini features, Magic Capture to pick the best frames, and more — For the last f

→12 Aug 2026Google unveils the $399 Pixel Watch 5 with a satin pyrite case finish, offline Gemini, proactive AI suggestions, better GPS maps, and insulin resistance trends (Victoria Song/The Verge)

Victoria Song / The Verge: Google unveils the 399 Pixel Watch 5 with a satin pyrite case finish, offline Gemini, proactive AI suggestions, better GPS maps, and insulin resistance trends — The 399 Goog

→12 Aug 2026Google DeepMind launches SL2T, a multilingual sign-language-to-text model debuting on the Pixel 11 in Gboard and Live Transcribe, first with ASL and English (Mike Wheatley/SiliconANGLE)

Mike Wheatley / SiliconANGLE: Google DeepMind launches SL2T, a multilingual sign-language-to-text model debuting on the Pixel 11 in Gboard and Live Transcribe, first with ASL and English — Google Deep

CompanyMeta8 recent entries
12 Aug 2026Do LLM Recommenders Know When They're Hallucinating? Auditing Confidence Calibration in Catalog Faithfulness

arXiv:2608.10008v1 Announce Type: cross Abstract: LLM recommenders for top-K item suggestion regularly emit titles outside the target catalog. Prior audits measure this as a binary out-of-domain rate;

→12 Aug 2026Cracks in the Foundation: Seemingly Minor Architectural Choices Impact Long Context Extension

arXiv:2608.10296v1 Announce Type: new Abstract: One might imagine that architectural variations within the dense transformer paradigm have a limited effect on accuracy. However, we demonstrate that th

→12 Aug 2026Behavioral Inference at Scale: The Fundamental Asymmetry Between Motivations and Belief Systems

arXiv:2509.05624v3 Announce Type: replace-cross Abstract: How much information about an agent's underlying values can be recovered from its observable behavior? This question matters for any approach

→12 Aug 2026b10375

chat : tighten bare function parsing for Qwen models (#26793) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64)

→12 Aug 2026b10373

imatrix.cpp: Move finite check and only check touched experts (#26861) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS In

→12 Aug 2026b10369

mtmd: support pocket-tts (#26871) adapt the api text model ok working impl, need verify and clean up mtmd: build the pocket-tts transposed convolutions as GEMM + col2im ggml_conv_transpose_1d has no g

→12 Aug 2026Astrolabe: Balancing Load in LLM Serving with Randomized Prediction-Guided Scheduling

arXiv:2508.03611v3 Announce Type: replace-cross Abstract: This paper presents Astrolabe, a randomized prediction-guided scheduler for one-shot request dispatch in multi-instance large language model (

→12 Aug 2026Actionable Hallucination Detection: Translating Latent Uncertainty into Agentic Critique

arXiv:2608.10430v1 Announce Type: cross Abstract: Large Language Models (LLMs) deployed as AI agents frequently exhibit user specification-grounding failures, executing hallucinated, undesired actions

CompanyMistral8 recent entries
11 Aug 2026DistillCache: KL-Guided Adaptive KV-Cache Eviction for Memory-Efficient LLM Inference

arXiv:2608.08878v1 Announce Type: cross Abstract: Transformer-based large language models (LLMs) achieve strong performance across many tasks, but their Key-Value (KV) cache grows linearly with sequen

→11 Aug 2026Can Open-Weight Models Compete on Financial Text Comprehension?

arXiv:2608.08634v1 Announce Type: new Abstract: Open-weight language models from Chinese AI labs caught up on benchmarks relative to proprietary frontier models in recent months. Yet their reliability

→11 Aug 2026⚡️A coalition that secures long-term AI capacity: We’re aggregating long-term compute demand in Europe to determine what capacity is built, …

⚡️A coalition that secures long-term AI capacity: We’re aggregating long-term compute demand in Europe to determine what capacity is built, where it’s located, and whom it serves. Through these multi-

→12 Aug 2026We are in an insane run of open-weight drops. Every modality, open source is winning. This is what an open source AI summer ☀️ looks like: …

We are in an insane run of open-weight drops. Every modality, open source is winning. This is what an open source AI summer ☀️ looks like: 🧠 LLMs & Reasoning → DeepSeek-V4-Flash-0731 (my king 👑): 304B

→12 Aug 2026The Truth Stays in the Family: Enhancing Contextual Grounding via Inherited Truthful Heads in Model Lineages

arXiv:2606.15821v2 Announce Type: replace-cross Abstract: Recent advances in large language models (LLMs) have produced many specialized multimodal LLMs (MLLMs) that share common foundational LLMs, fo

→12 Aug 2026REAP: Relation-Aware Elicitation and Parsing for Closed-Book Knowledge Base Construction from LLMs

arXiv:2608.10963v1 Announce Type: new Abstract: We present the REAP system for the AKBC Shared Task 2026 on constructing knowledge bases from language models in a closed-book setting, subject to a bud

→12 Aug 2026Mistral says its platform will support third-party open models, starting with Z.ai's GLM-5.2, and run them on the same infrastructure as its own models (Mistral AI Blog)

Mistral AI Blog: Mistral says its platform will support third-party open models, starting with Z.ai's GLM-5.2, and run them on the same infrastructure as its own models — At Mistral, we believe every

→12 Aug 2026Do LLM Recommenders Know When They're Hallucinating? Auditing Confidence Calibration in Catalog Faithfulness

arXiv:2608.10008v1 Announce Type: cross Abstract: LLM recommenders for top-K item suggestion regularly emit titles outside the target catalog. Prior audits measure this as a binary out-of-domain rate;

CompanyxAI8 recent entries
12 Aug 2026SPACEXAI: Grok 4.6 leads on the two strongest knowledge-work / real-world productivity benchmarks (GDPVal-AA and AA-Briefcase) and on the le…

SPACEXAI: Grok 4.6 leads on the two strongest knowledge-work / real-world productivity benchmarks (GDPVal-AA and AA-Briefcase) and on the legal benchmark, while remaining highly competitive on coding-

→12 Aug 2026imo GDPVal is probably the most important benchmark, it measures the performance of models on real world tasks Big leap in performance here …

imo GDPVal is probably the most important benchmark, it measures the performance of models on real world tasks Big leap in performance here to top it at a great price, congrats to @SpaceXAI team & loo

→12 Aug 2026Grok Bot

Grok Bot Here's my Grok Bot team: - Webby: Web designer - Shotry: Short-form content creator - Writey: Article/Newsletter writer - Claude Code: Grok agent that specializes in CC - Codex: Same as the a

→12 Aug 2026Grok 4.6 reaches 1753 ELO

Grok 4.6 topped the GDPVal-AA benchmark with an Elo score of 1,753. It surpassed competitors Fable 5 Max (1,741 Elo), GPT‑5.6 Sol Max (1,728 Elo) and Grok 4.5 High (1,526 Elo). Elon Musk publicly ackn

→12 Aug 2026Grok 4.6 is objectively #1 when considering intelligence, speed & cost

Grok 4.6 is objectively #1 when considering intelligence, speed & cost SpaceXAI's Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, joining the frontier in line with GPT-5.6 Sol, with

→12 Aug 2026Grok 4.6 is now one of the top models in the world for agentic workflows It ranks #1 on the Artificial Analysis Agentic Index, tied with Cla…

Grok 4.6 is now one of the top models in the world for agentic workflows It ranks #1 on the Artificial Analysis Agentic Index, tied with Claude Opus 5 Max Agentic AI is about more than answering quest

→12 Aug 2026Grok 4.6 is an excellent model. I’ve been using it heavily for the past couple of weeks and it handles everything from simple coding & code …

Grok 4.6 is an excellent model. I’ve been using it heavily for the past couple of weeks and it handles everything from simple coding & code review all the way to designing and debugging complex system

→12 Aug 2026give 4.6 a try and let us know how it goes. your feedback is a big part of why the model gets better with each iteration.

give 4.6 a try and let us know how it goes. your feedback is a big part of why the model gets better with each iteration. SpaceXAI's Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, j

CompanyDeepSeek8 recent entries
12 Aug 2026Qwen3.8-2.4T-A95B is now live on Fireworks with Day-0 support! This 2.4T parameter MoE model is built for autonomous agents, heavy coding, l…

Qwen3.8-2.4T-A95B is now live on Fireworks with Day-0 support! This 2.4T parameter MoE model is built for autonomous agents, heavy coding, large context windows, and is ideal for coding and agentic pe

→12 Aug 2026Persistent Recursive Worlds Enable Autonomous Software Evolution

arXiv:2608.10450v1 Announce Type: cross Abstract: Complex software systems develop over timescales that exceed the lifespan of any individual coding agent. Most agentic software systems preserve conti

→12 Aug 2026Measuring Semantic Abstractness of SAE Features via Nonlocality

arXiv:2608.10537v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) have helped uncover mechanistic explanations for LLM behaviours such as reasoning, jailbreaking etc., via understanding the c

→12 Aug 2026Interesting research suggests caution in determining which AI company is winning by looking at any one source.. OpenRouter seems to show ope…

Interesting research suggests caution in determining which AI company is winning by looking at any one source.. OpenRouter seems to show open weights winning over time, but work submitted to Pangram i

→12 Aug 2026Idea for a deepseek-v4-flash-0731 backed automated research workflow to be leveraged via qwen3.6/3.8 27b for difficult tasks that require highly technical, not easy to find information.

Sometimes you have tasks that are outside of your expertise and the idea is this workflow automation could be leveraged to manage to have local AI figure it out using research from his workflow gather

→12 Aug 2026DeepSeek V4 Flash 0731 uncensored (jailbreak pt2)

Since lot's of people were sceptical or whatever, heres how to uncensor / jailbreak V4 flash and proof. No it is not lead on whatever, first prompt, first try, every time. Put this in System message:

→12 Aug 2026CohereLabs/North-Micro-Vision-Instruct · Hugging Face

North Micro Vision Instruct is a 2.4B-parameter open-weight vision-language model with native-resolution image support, released under the Apache 2.0 license. It is designed as a compact foundation fo

→12 Aug 2026CHORUS: Complementary Experts for High-Coverage Testbench Stimulus Generation

arXiv:2608.10090v1 Announce Type: new Abstract: Large language models (LLMs) have advanced code generation, where executable feedback provides a more reliable learning signal than textual imitation al

CompanyNVIDIA8 recent entries
11 Aug 202610 year garbage card for local llms

Hello everyone! ​I like dumb things. I like working with weak computers and microcontrollers. I like the simplicity and low electricity usage. Simply put, the efficiency of a 'dumb' PC. ​The first tim

→12 Aug 2026We are in an insane run of open-weight drops. Every modality, open source is winning. This is what an open source AI summer ☀️ looks like: …

We are in an insane run of open-weight drops. Every modality, open source is winning. This is what an open source AI summer ☀️ looks like: 🧠 LLMs & Reasoning → DeepSeek-V4-Flash-0731 (my king 👑): 304B

→12 Aug 2026Tested Nemotron 3.5 Lightning locally on coding, Hermes Agent and agentic work

Ran the model with quants (Q5) and MTP by bartowski with llama.cpp server. It takes ~24GB ram running on M5 Pro with 48GB at about 65t/s. On some tasks it was quite the overthinker. Overall, the quali

→12 Aug 2026Serve Qwen3.8-2.4T-A95B, a 2.4T-Parameter Model, with Configurable Reasoning on NVIDIA GB300 NVL72

Alibaba released the open‑weights Qwen3.8‑2.4T‑A95B (Qwen3.8‑Max), a fine‑grained mixture‑of‑experts model with 2.4 trillion parameters, hybrid full‑ and linear‑attention, a one‑million‑token context

→12 Aug 2026Qwen3.8-2.4T-A95B is now live on Fireworks with Day-0 support! This 2.4T parameter MoE model is built for autonomous agents, heavy coding, l…

Qwen3.8-2.4T-A95B is now live on Fireworks with Day-0 support! This 2.4T parameter MoE model is built for autonomous agents, heavy coding, large context windows, and is ideal for coding and agentic pe

→12 Aug 2026Persona Conditioning as an Assessor-Sensitivity Probe for LLM-Based IR Evaluation

arXiv:2608.10385v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as relevance assessors in information retrieval (IR) evaluation, raising questions about how assess

→12 Aug 2026NVIDIA CEO Tops Glassdoor’s 2026 List of Best CEOs

NVIDIA founder and CEO Jensen Huang is ranked No. 1 on Glassdoor’s Best CEOs list for 2026. In the just-released ranking, recognition is earned directly from the people who know their leadership the b

→12 Aug 2026Four small architecture decisions can cost up to 47% of a model's long-context performance. New research from Ai2, Carnegie Mellon, and the …

Four small architecture decisions can cost up to 47% of a model's long-context performance. New research from Ai2, Carnegie Mellon, and the University of Washington isolates them. Normalization, GQA,