AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
6,147 results
Model Releases

Fascinating: OpenAI’s @deanwball is saying Astra can do anything, and it’s not even clear it can do “anything” in math (let alone anything i…

DGX agent

Fascinating: OpenAI’s @deanwball is saying Astra can do anything, and it’s not even clear it can do “anything” in math (let alone anything in more or open-ended, less formalizable domains). I dropped

model-releasesgary-marcus--x
1 Aug 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

If you maintain an AGENTS.md or a CLAUDE.md, this is worth a read. (bookmark it) 288 gold-test evaluated runs across Claude Code and Codex, …

DGX agent

If you maintain an AGENTS.md or a CLAUDE.md, this is worth a read. (bookmark it) 288 gold-test evaluated runs across Claude Code and Codex, 17 real tasks from 3 repositories, with context-injection st

model-releasesdair-ai--x
1 Aug 2026
Model Releases

// Persistent Workspaces for Long-Lived Claude Code Agent Teams // Four issues to be aware of: > Working state vanishes when a terminal clos…

DGX agent

// Persistent Workspaces for Long-Lived Claude Code Agent Teams // Four issues to be aware of: > Working state vanishes when a terminal closes and the team cannot be resumed. > Compaction condenses th

model-releasesdair-ai--x
1 Aug 2026
Model Releases

so much for the “general” part of general intelligence

DGX agent

so much for the “general” part of general intelligence I have tried many times to get ChatGPT, Claude, Grok or Gemini to write scripts for my YouTube videos. It is still a complete failure. For one th

model-releasesgary-marcus--x
1 Aug 2026
Model Releases

“Stochastic parrots” is not my term (it’s @emilymbender’s). But a lot of people today commenting on it are confused. To some extent (though …

DGX agent

“Stochastic parrots” is not my term (it’s @emilymbender’s). But a lot of people today commenting on it are confused. To some extent (though I don’t think it’s a perfect metaphor, and have said that be

model-releasesgary-marcus--x
1 Aug 2026
Model Releases

Wake me when Astra solves a significant open-world problem that doesn’t revolve around formal verification. Or at least fixes poor @skdh’s v…

DGX agent

Wake me when Astra solves a significant open-world problem that doesn’t revolve around formal verification. Or at least fixes poor @skdh’s video problems. I have tried many times to get ChatGPT, Claud

model-releasesgary-marcus--x
1 Aug 2026
Model Releases

🚀 DeepSeek-V4-Flash Official API is now LIVE in public beta! 🔷 We’ve massively upgraded its Agent capabilities—benchmark scores are now fa…

DGX agent

🚀 DeepSeek-V4-Flash Official API is now LIVE in public beta! 🔷 We’ve massively upgraded its Agent capabilities—benchmark scores are now far surpassing the V4-Pro-Preview. Check out the massive perform

model-releasesdeepseek--x
31 Jul 2026
Model Releases

Hermes and Pi Agent led on the average cost per task, while Claude Code cost about 3.7x as much as Pi: - 0.39 Hermes Agent - 0.40 Pi Agent…

DGX agent

Hermes and Pi Agent led on the average cost per task, while Claude Code cost about 3.7x as much as Pi: - 0.39 Hermes Agent - 0.40 Pi Agent - 0.47 Codex - 0.51 OpenCode - 0.54 Kimi Code - 1.47 Claude C

model-releasesnous-research--x
31 Jul 2026
Model Releases

i see your moore's law and i raise you 20x

DGX agent

i see your moore's law and i raise you 20x GPT-5.4 full at xhigh scored 51, exactly where Luna max sits today. GPT-5.4 costs 2.50/15; Luna now costs 0.20/1.20. In other words, roughly four months late

model-releasessam-altman--x
31 Jul 2026
Model Releases

If LoRA is underperforming, don't reach for more expensive full parameter fine-tuning right away. We ran three cheap tests (data coverage, o…

DGX agent

If LoRA is underperforming, don't reach for more expensive full parameter fine-tuning right away. We ran three cheap tests (data coverage, optimization, rank) to see if we could close the gap between

model-releasesfireworks-ai--x
31 Jul 2026
Model Releases

'Intelligence too cheap to meter' battle is on! Given that DeepSeek-V4-Flash-Preview is already great for agentic tasks, there is no doubt t…

DGX agent

'Intelligence too cheap to meter' battle is on! Given that DeepSeek-V4-Flash-Preview is already great for agentic tasks, there is no doubt this new checkpoint must be an absolute beast. 20+ point jump

model-releasesdair-ai--x
31 Jul 2026
Model Releases

Neat work on long-horizon agents. Splitting a hard task across agents is typically how standard multi-agent work. The usual design lets them…

DGX agent

Neat work on long-horizon agents. Splitting a hard task across agents is typically how standard multi-agent work. The usual design lets them exchange findings only at phase boundaries, through staged

model-releasesdair-ai--x
31 Jul 2026
Model Releases

OK, GPT-5.6 Luna is a bit of a beast. Given the 80% price drop today I decided to try it in Datasette Agent, and it's furiously quick and ge…

DGX agent

OK, GPT-5.6 Luna is a bit of a beast. Given the 80% price drop today I decided to try it in Datasette Agent, and it's furiously quick and generates all the SQL, HTML and JavaScript (for Datasette Apps

model-releasessimon-willison--x
31 Jul 2026
Model Releases

Period.

DGX agent

Period. If LoRA is underperforming, don't reach for more expensive full parameter fine-tuning right away. We ran three cheap tests (data coverage, optimization, rank) to see if we could close the gap

model-releasesfireworks-ai--x
31 Jul 2026
Model Releases

The new DeepSeek v4 Flash is now available in Hermes Agent through Nous Portal and OpenRouter

DGX agent

The new DeepSeek v4 Flash is now available in Hermes Agent through Nous Portal and OpenRouter 🚀 DeepSeek-V4-Flash Official API is now LIVE in public beta! 🔷 We’ve massively upgraded its Agent capabili

model-releasesnous-research--x
31 Jul 2026
Model Releases

They knew the charges were bullshit from the start. They knew they arrested an innocent man just to intimidate others and appease the fake n…

DGX agent

They knew the charges were bullshit from the start. They knew they arrested an innocent man just to intimidate others and appease the fake narrative of an incompetent fucking toddler. This is the real

model-releasesanthropic--x
31 Jul 2026
Model Releases

Try Grok 4.5 http://X.ai/cli

DGX agent

Try Grok 4.5 http://X.ai/cli BREAKING: Grok 4.5 outperforms GPT-5.6 Terra across almost all shared benchmarks on AskClash, leading in ACB, GPQA, SWE-P, and Atlas while also achieving a higher overall

model-releaseselon-musk--x
31 Jul 2026
Model Releases

We created a document OCR router that can estimate the complexity of every single page and parse it with the relevant mode 💫 * Some pages a…

DGX agent

We created a document OCR router that can estimate the complexity of every single page and parse it with the relevant mode 💫 * Some pages are full of native text, which can be directly handled with Li

model-releasesjerry-liu--x
31 Jul 2026
Model Releases

And Grok 4.6 comes out in a week

DGX agent

And Grok 4.6 comes out in a week BREAKING: Grok 4.5 ranked #1 on LaurenBench with a score of 56.9%, ahead of Claude Sonnet 5, GLM 5.2, Claude Opus 5, Kimi K3 and GPT-5.6. The benchmark tests real-worl

model-releaseselon-musk--x
30 Jul 2026
Model Releases

ComfyUI Face Swap Tutorial: Fast, Clean Results VFX artist Doug Hogan walks through a fully automated face swap workflow built inside ComfyU…

DGX agent

ComfyUI Face Swap Tutorial: Fast, Clean Results VFX artist Doug Hogan walks through a fully automated face swap workflow built inside ComfyUI - combining Florence 2, SAM2, and WAN Video into a single

model-releasescomfyui--x
30 Jul 2026
Model Releases

Consistent with what I found with Qwen3.6 a while back: Claude Code uses 2-3x as many tokens than (many) other harnesses at similar success …

DGX agent

Consistent with what I found with Qwen3.6 a while back: Claude Code uses 2-3x as many tokens than (many) other harnesses at similar success rate. - Unoptimized? - Buggy? - Deliberate (coz that helps i

model-releasessebastian-raschka--x
30 Jul 2026
Model Releases

goblin-level blog post

DGX agent

goblin-level blog post Turns out GPT-5.6 Sol is actually SoTA on ARC-AGI-3. Just took two setting changes. You just have to allow it to reason and work over multiple context windows with the help of o

model-releasessam-altman--x
30 Jul 2026
Model Releases

A new TIL on adding custom MCP servers to both the ChatGPT and Claude regular chat interfaces - it's a little less obvious than I had hoped,…

DGX agent

A new TIL on adding custom MCP servers to both the ChatGPT and Claude regular chat interfaces - it's a little less obvious than I had hoped, but I got there in the end https://til.simonwillison.net/ll

model-releasessimon-willison--x
29 Jul 2026
Model Releases

After deployment, we applied GPT-5.6 Sol to advance the frontier of efficiency by making itself more efficient to run. The results: - 20% lo…

DGX agent

After deployment, we applied GPT-5.6 Sol to advance the frontier of efficiency by making itself more efficient to run. The results: - 20% lower serving costs from production GPU kernel improvements. -

model-releasesopenai--x
29 Jul 2026
Model Releases

BREAKING: Grok 4.5 (high) ranks #1 on the HighWalk benchmark, which tests how well AI agents update technical specifications from code chang…

DGX agent

BREAKING: Grok 4.5 (high) ranks #1 on the HighWalk benchmark, which tests how well AI agents update technical specifications from code changes. Grok delivered the best combination of quality and opera

model-releaseselon-musk--x
29 Jul 2026
Model Releases

BREAKING: Grok 4.5 ranked #1 on LaurenBench with a score of 56.9%, ahead of Claude Sonnet 5, GLM 5.2, Claude Opus 5, Kimi K3 and GPT-5.6. Th…

DGX agent

BREAKING: Grok 4.5 ranked #1 on LaurenBench with a score of 56.9%, ahead of Claude Sonnet 5, GLM 5.2, Claude Opus 5, Kimi K3 and GPT-5.6. The benchmark tests real-world AI agents across conversation,

model-releaseselon-musk--x
29 Jul 2026
Model Releases

ComfyUI Face Swap Tutorial: Fast, Clean Results VFX artist @heydoughogan walks through a fully automated face swap workflow built inside Com…

DGX agent

ComfyUI Face Swap Tutorial: Fast, Clean Results VFX artist @heydoughogan walks through a fully automated face swap workflow built inside ComfyUI - combining Florence 2, SAM2, and WAN Video into a sing

model-releasescomfyui--x
29 Jul 2026
Research

Dreaming in Voxels: How AI is Generating Playable Minecraft Worlds Generative AI has conquered images, video, text. But what about interacti…

DGX agent

Dreaming in Voxels: How AI is Generating Playable Minecraft Worlds Generative AI has conquered images, video, text. But what about interactive 3D environments? We trained models on billions of cubes t

researchdavid-ha--x
29 Jul 2026
Safety

@GaryMarcus OpenAI and Anthropic desperately need to raise prices. We seem to be faced with one of three choices for them: 1. Bailout 2. Ban…

DGX agent

Gary Marcus has argued that both OpenAI and Anthropic must increase their pricing, citing the rapid decline in token costs and competition from open‑source models. He presents a binary choice: either

safetygary-marcus--x
29 Jul 2026
Model Releases

GPT-5.6 Sol has been used to solve open problems in mathematics. So why was it struggling with ARC-AGI-3, a benchmark of 2D puzzle games? We…

DGX agent

GPT-5.6 Sol has been used to solve open problems in mathematics. So why was it struggling with ARC-AGI-3, a benchmark of 2D puzzle games? We investigated. The harness was not letting it remember what

model-releasesopenai--x
29 Jul 2026
Model Releases

https://x.com/EMostaque/status/2082600218529235174

DGX agent

After deployment, we applied GPT-5.6 Sol to advance the frontier of efficiency by making itself more efficient to run. The results: - 20% lower serving costs from production GPU kernel improvements. -

model-releasesemad-mostaque--x
29 Jul 2026
Model Releases

Install the open-source Codex Security CLI,: npm install @OpenAI/codex-security Or start with: npx @OpenAI/codex-security@latest --help NPM:…

DGX agent

OpenAI has released the open‑source Codex Security CLI, which can be installed with `npm install @OpenAI/codex-security` or run directly via `npx @OpenAI/codex-security@latest --help`. The tool scans

model-releasesopenai--x
29 Jul 2026
Model Releases

On benchmarking long-context agentic instruction following. Agent benchmarks mostly reward reaching the answer. This new benchmark measures …

DGX agent

On benchmarking long-context agentic instruction following. Agent benchmarks mostly reward reaching the answer. This new benchmark measures whether the agent reached it the permitted way, which is the

model-releasesdair-ai--x
29 Jul 2026
Model Releases

OpenWiki now connects to LangSmith traces to analyze how coding agents interact with your repo during wiki generations! We added a LangSmith…

DGX agent

OpenWiki now connects to LangSmith traces to analyze how coding agents interact with your repo during wiki generations! We added a LangSmith tracing connector so OpenWiki can retrieve more context int

model-releasesharrison-chase--x
29 Jul 2026
Model Releases

The median task consumed nearly 6x more tokens in Claude Code than in Kimi Code: - 61k in Kimi Code - 67k in Hermes - 340k in Claude Code At…

DGX agent

The median task consumed nearly 6x more tokens in Claude Code than in Kimi Code: - 61k in Kimi Code - 67k in Hermes - 340k in Claude Code At K3's 3/M input rate (input tokens make up roughly 95% of ag

model-releasesnous-research--x
29 Jul 2026
Model Releases

This is a willfully misleading narrative from OpenAI. Sam’s earnest expressions are being deployed, again, to misdirect. If there’s a need t…

DGX agent

This is a willfully misleading narrative from OpenAI. Sam’s earnest expressions are being deployed, again, to misdirect. If there’s a need to “pace AI development”, the OAI hack incident isn’t any evi

model-releasesyann-lecun--x
29 Jul 2026
Model Releases

We believe the benefits of frontier AI should not be concentrated in a few companies and well-resourced labs. Researchers know their fields …

DGX agent

We believe the benefits of frontier AI should not be concentrated in a few companies and well-resourced labs. Researchers know their fields best. Our role is to put powerful tools in their hands and h

model-releasesopenai--x
29 Jul 2026
Model Releases

We quietly released the open-source Codex Security CLI, but Hacker News found it before we had a chance to share it here... You can now use …

DGX agent

We quietly released the open-source Codex Security CLI, but Hacker News found it before we had a chance to share it here... You can now use it to scan repositories, track findings across runs, verify

model-releasesopenai--x
29 Jul 2026
Safety

We ran a large-scale distillation attack on the Kimi K3 technical report by reading it in parallel at the Hugging Face Journal Club :) https…

DGX agent

We ran a large-scale distillation attack on the Kimi K3 technical report by reading it in parallel at the Hugging Face Journal Club :) https://youtu.be/MW8-kqd2SD8?si=jSKDogcUWbJ8N2k7 Our main takeawa

safetyclem-delangue--x
29 Jul 2026
Model Releases

Who on earth came up with 5 hour limits? The day is 24 hours long. While do our limit times shift by an hour each day. Pls if you are going …

DGX agent

Who on earth came up with 5 hour limits? The day is 24 hours long. While do our limit times shift by an hour each day. Pls if you are going to have limits 4 or 6 hours. Hello people of Sol! I've reset

model-releasesemad-mostaque--x
29 Jul 2026
Safety

AGI smarter than the smartest humans, my ass. @Kasparov63 (peak rating 2851) probably could’ve crushed the best commercial large language mo…

DGX agent

AGI smarter than the smartest humans, my ass. @Kasparov63 (peak rating 2851) probably could’ve crushed the best commercial large language models in chess when he was 7 years old. graph courtesy @chess

safetygary-marcus--x
28 Jul 2026
Model Releases

https://x.com/Sauers_/status/2082171683645817193

DGX agent

New Anthropic research: Discovering cryptographic weaknesses with Claude. Claude Mythos Preview has helped our researchers find weaknesses in cryptographic algorithms—the mathematical methods that are

model-releasesemad-mostaque--x
28 Jul 2026
Model Releases

It’s been five years since ChatGPT was released, AGI is not coming, not even the most modest AI predictions will ever become a reality, and …

DGX agent

It’s been five years since ChatGPT was released, AGI is not coming, not even the most modest AI predictions will ever become a reality, and there needs to be some of severe criminal sanction applied t

model-releasesgary-marcus--x
28 Jul 2026
Safety

The administration is very consistent: All aspects of foreign involvement in AI are banned: people (students and employees), hardware and mo…

DGX agent

The administration is very consistent: All aspects of foreign involvement in AI are banned: people (students and employees), hardware and models. And funding is down to a drip and capital allocation i

safetygary-marcus--x
28 Jul 2026
Safety

創業以来、オープンソースコミュニティから多くを学び、また研究成果の公開を通じてそこに貢献してきました。オープンなエコシステムが健全なAI産業と技術主権を支える重要な基盤の一つであると考えており、その発展を支持します。 このたび、Sakana AIは、オープンウェイトAIモデルに関…

DGX agent

創業以来、オープンソースコミュニティから多くを学び、また研究成果の公開を通じてそこに貢献してきました。オープンなエコシステムが健全なAI産業と技術主権を支える重要な基盤の一つであると考えており、その発展を支持します。 このたび、Sakana AIは、オープンウェイトAIモデルに関する公開書簡 「Open Weights and American AI Leadership」に署名しました。 書簡は

safetydavid-ha--x
27 Jul 2026
Model Releases

All of this took me under 5 minutes to find. Great reminder to all of us to be more vigilant of multiplayer mode in AI and to maybe ask Code…

DGX agent

All of this took me under 5 minutes to find. Great reminder to all of us to be more vigilant of multiplayer mode in AI and to maybe ask Codex or Claude to go through all of your google docs and sheets

model-releasesallie-k--miller--x
27 Jul 2026
Model Releases

At a small business, the person closest to a problem is often the one who has to solve it—even when it falls outside their job description. …

DGX agent

At a small business, the person closest to a problem is often the one who has to solve it—even when it falls outside their job description. We studied how AI is becoming powerful generalist tool for s

model-releasesopenai--x
27 Jul 2026
Industry

Extremely important point

DGX agent

Extremely important point Attackers have frontier AI. Defenders need a frontier AI ecosystem—the best open and closed models, force-multiplied by a global community. During the Hugging Face incident,

industryelon-musk--x
27 Jul 2026
← Previous
1…6364656667…129
Next →