AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
Human
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
13,897 results
30 Jul 2026

This is absolutely wild... Anthropic reviewed their logs and found out that their own supposedly-sandboxed cyber evals had hacked three sepa…

Model ReleasesDGX agent

This is absolutely wild... Anthropic reviewed their logs and found out that their own supposedly-sandboxed cyber evals had hacked three separate companies back in April without them noticing! In a rev

We are committed to pushing the model frontier across cost efficiency, capability, and speed. Starting today, we are reducing prices for GPT…

Model ReleasesDGX agent

We are committed to pushing the model frontier across cost efficiency, capability, and speed. Starting today, we are reducing prices for GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20% , and offering a f

We have to be careful to not offload our understanding to agents. I think there is also a good opportunity to build agentic applications tha…

DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
AgentsDGX agent

We have to be careful to not offload our understanding to agents. I think there is also a good opportunity to build agentic applications that encourage deeper understanding. For example, coding agents

We’re excited to rollout an official batch parsing experience to LlamaParse. ✅ Instead of hitting our APIs one file at a time, create a batc…

AgentsDGX agent

We’re excited to rollout an official batch parsing experience to LlamaParse. ✅ Instead of hitting our APIs one file at a time, create a batch of 10k at once. ✅ Get a dedicated UI where you can audit t

Yesterday I cohosted a dinner with @dexhorthy with a wonderful group of founders, to talk about agent loops and loop engineering. Some inter…

Model ReleasesDGX agent

Yesterday I cohosted a dinner with @dexhorthy with a wonderful group of founders, to talk about agent loops and loop engineering. Some interesting insights: * Most of our group was *not* actively usin

29 Jul 2026

A benchmark score reflects the model as well as the harness and settings used to run it. For long-running agents, retaining reasoning and co…

Model ReleasesDGX agent

A benchmark score reflects the model as well as the harness and settings used to run it. For long-running agents, retaining reasoning and compacting context lets the model build on what it has already

A new TIL on adding custom MCP servers to both the ChatGPT and Claude regular chat interfaces - it's a little less obvious than I had hoped,…

Model ReleasesDGX agent

A new TIL on adding custom MCP servers to both the ChatGPT and Claude regular chat interfaces - it's a little less obvious than I had hoped, but I got there in the end https://til.simonwillison.net/ll

After a few more hours, I think I've figured out Opus 5. Opus 5 is trained to be more agentic than anything I've used. All Claude 5 models a…

Model ReleasesDGX agent

After a few more hours, I think I've figured out Opus 5. Opus 5 is trained to be more agentic than anything I've used. All Claude 5 models are like that. So what changes? The way to interact with Opus

After deployment, we applied GPT-5.6 Sol to advance the frontier of efficiency by making itself more efficient to run. The results: - 20% lo…

Model ReleasesDGX agent

After deployment, we applied GPT-5.6 Sol to advance the frontier of efficiency by making itself more efficient to run. The results: - 20% lower serving costs from production GPU kernel improvements. -

Agentic inference wastes GPUs on KV cache thrashing. ThunderAgent fixes it at the scheduler level: 2.5x higher single-node throughput and ~1…

AgentsDGX agent

Agentic inference wastes GPUs on KV cache thrashing. ThunderAgent fixes it at the scheduler level: 2.5x higher single-node throughput and ~10x lower P50 latency at high concurrency. ThunderAgent was a

agreed. which is part of why coating the world in data centers is a profound mistake.

HardwareDGX agent

agreed. which is part of why coating the world in data centers is a profound mistake. AI will get so ridiculously efficient that we will look back at GPU clusters the way we now look at these first ro

AI can write more code than any team can review by hand, and a pull request can look fine while hiding a security issue or missing a require…

AgentsDGX agent

AI can write more code than any team can review by hand, and a pull request can look fine while hiding a security issue or missing a requirement. In our new short course, AI Code Review, built in coll

Batch parsing used to mean writing API scripts. Now it's a button. 🦙 Parse or extract across up to 10,000 files in one run, straight from t…

AgentsDGX agent

Batch parsing used to mean writing API scripts. Now it's a button. 🦙 Parse or extract across up to 10,000 files in one run, straight from the LlamaParse UI. Point it at a folder and go — no code requi

BREAKING: Grok 4.5 (high) ranks #1 on the HighWalk benchmark, which tests how well AI agents update technical specifications from code chang…

Model ReleasesDGX agent

BREAKING: Grok 4.5 (high) ranks #1 on the HighWalk benchmark, which tests how well AI agents update technical specifications from code changes. Grok delivered the best combination of quality and opera

BREAKING: Grok 4.5 just claimed the top spot on the new HighWalk Benchmark. The independent test measures how well AI models update real tec…

Model ReleasesDGX agent

BREAKING: Grok 4.5 just claimed the top spot on the new HighWalk Benchmark. The independent test measures how well AI models update real technical specifications from 46 Laravel commits — heavy on cod

BREAKING: Grok 4.5 ranked #1 on LaurenBench with a score of 56.9%, ahead of Claude Sonnet 5, GLM 5.2, Claude Opus 5, Kimi K3 and GPT-5.6. Th…

Model ReleasesDGX agent

BREAKING: Grok 4.5 ranked #1 on LaurenBench with a score of 56.9%, ahead of Claude Sonnet 5, GLM 5.2, Claude Opus 5, Kimi K3 and GPT-5.6. The benchmark tests real-world AI agents across conversation,

BREAKING: SpaceXAI's newly released Grok Voice Think Fast 2.0 beats voice models from OpenAI, Google, Alibaba, and DeepSlate in the Artifici…

Model ReleasesDGX agent

SpaceXAI has released its new Grok Voice Think Fast 2.0, which on the Artificial Analysis Speech‑to‑Speech benchmark outperformed leading models from OpenAI, Google, Alibaba and DeepSlate. The claim w

ComfyUI Face Swap Tutorial: Fast, Clean Results VFX artist @heydoughogan walks through a fully automated face swap workflow built inside Com…

Model ReleasesDGX agent

ComfyUI Face Swap Tutorial: Fast, Clean Results VFX artist @heydoughogan walks through a fully automated face swap workflow built inside ComfyUI - combining Florence 2, SAM2, and WAN Video into a sing

DeepSeek V4 Flash isn't just for inference anymore. Fine-tune it on Fireworks with supervised fine-tuning, preference tuning, and combined p…

Model ReleasesDGX agent

DeepSeek V4 Flash isn't just for inference anymore. Fine-tune it on Fireworks with supervised fine-tuning, preference tuning, and combined preference optimization from the managed UI. Reinforcement le

Dreaming in Voxels: How AI is Generating Playable Minecraft Worlds Generative AI has conquered images, video, text. But what about interacti…

ResearchDGX agent

Dreaming in Voxels: How AI is Generating Playable Minecraft Worlds Generative AI has conquered images, video, text. But what about interactive 3D environments? We trained models on billions of cubes t

@GaryMarcus OpenAI and Anthropic desperately need to raise prices. We seem to be faced with one of three choices for them: 1. Bailout 2. Ban…

SafetyDGX agent

Gary Marcus has argued that both OpenAI and Anthropic must increase their pricing, citing the rapid decline in token costs and competition from open‑source models. He presents a binary choice: either

GPT-5.6 Sol has been used to solve open problems in mathematics. So why was it struggling with ARC-AGI-3, a benchmark of 2D puzzle games? We…

Model ReleasesDGX agent

GPT-5.6 Sol has been used to solve open problems in mathematics. So why was it struggling with ARC-AGI-3, a benchmark of 2D puzzle games? We investigated. The harness was not letting it remember what

Handy dandy AI crisis flowchart from @klonick

SafetyDGX agent

Gary Marcus shared a “handy dandy AI crisis flowchart” created by Kate Klonick (@Klonick) on Saturday, July 28. The tweet also notes that he had planned to write an article about Hugging Face and Open

https://x.com/EMostaque/status/2082600218529235174

Model ReleasesDGX agent

After deployment, we applied GPT-5.6 Sol to advance the frontier of efficiency by making itself more efficient to run. The results: - 20% lower serving costs from production GPU kernel improvements. -

I gave a talk on forward deployed engineering to a thousand AI engineers at @aiDotEngineer World's Fair. A year ago I'd have opened by expla…

ToolsDGX agent

I gave a talk on forward deployed engineering to a thousand AI engineers at @aiDotEngineer World's Fair. A year ago I'd have opened by explaining what FDE stood for. Not this time. Thank you to @swyx,

i'll say it plainly, hermes agent desktop is the best agentic app i've used, and i'm a little mad i didn't find it sooner. it auto see the m…

Local AiDGX agent

i'll say it plainly, hermes agent desktop is the best agentic app i've used, and i'm a little mad i didn't find it sooner. it auto see the models i'm serving, laguna s 2.1 sitting on my dgx spark and

Impressive paper! It's on one of the hardest tasks for coding agents today. Of course, I am talking about kernel optimization. Coding agents…

Local AiDGX agent

Impressive paper! It's on one of the hardest tasks for coding agents today. Of course, I am talking about kernel optimization. Coding agents are usually not so great at this. Reasons: Unfamiliar low-l

Install the open-source Codex Security CLI,: npm install @OpenAI/codex-security Or start with: npx @OpenAI/codex-security@latest --help NPM:…

Model ReleasesDGX agent

OpenAI has released the open‑source Codex Security CLI, which can be installed with `npm install @OpenAI/codex-security` or run directly via `npx @OpenAI/codex-security@latest --help`. The tool scans

It's a drop-in. One program_id field and it plugs into your existing engine configs such as KV offloading + speculative decoding. Format-agn…

HardwareDGX agent

It's a drop-in. One program_id field and it plugs into your existing engine configs such as KV offloading + speculative decoding. Format-agnostic by design, with OpenAI chat completions supported toda

Jensen Huang was asked about the people technology left behind. He didn't offer a plan. He said the gap already closed. Huang: 'All of a sud…

TutorialsDGX agent

Jensen Huang was asked about the people technology left behind. He didn't offer a plan. He said the gap already closed. Huang: 'All of a sudden artificial intelligence closed that technology divide.'

Model + harness. We have barely begun to understand the best ways to do harness engineering. A huge amount of untapped potential even withou…

Model ReleasesDGX agent

Model + harness. We have barely begun to understand the best ways to do harness engineering. A huge amount of untapped potential even without models getting better (but models are getting better) Turn

On benchmarking long-context agentic instruction following. Agent benchmarks mostly reward reaching the answer. This new benchmark measures …

Model ReleasesDGX agent

On benchmarking long-context agentic instruction following. Agent benchmarks mostly reward reaching the answer. This new benchmark measures whether the agent reached it the permitted way, which is the

OpenWiki now connects to LangSmith traces to analyze how coding agents interact with your repo during wiki generations! We added a LangSmith…

Model ReleasesDGX agent

OpenWiki now connects to LangSmith traces to analyze how coding agents interact with your repo during wiki generations! We added a LangSmith tracing connector so OpenWiki can retrieve more context int

Packed house today at the first edition of our @MiniMax_AI Intelligence in the Open event. Thank you to our amazing cohost @withprotegeai An…

ToolsDGX agent

Packed house today at the first edition of our @MiniMax_AI Intelligence in the Open event. Thank you to our amazing cohost @withprotegeai And our wonderful partners UpscaleX @togethercompute @Artifici

Participants will receive access to our frontier models, including our GPT-5.6 family of models. Each workspace includes business-grade priv…

Model ReleasesDGX agent

Participants will receive access to our frontier models, including our GPT-5.6 family of models. Each workspace includes business-grade privacy and security protections. Researcher data is not used to

Replit Design reimagines the UX for AI design with Ambient Intelligence. You don’t need to prompt. You don’t need design language. At every …

AgentsDGX agent

Replit Design reimagines the UX for AI design with Ambient Intelligence. You don’t need to prompt. You don’t need design language. At every step the agent suggests next best actions that you can take

Sakana is at the frontier of having fun and I respect that

AgentsDGX agent

Sakana is at the frontier of having fun and I respect that We are excited to share our latest work, together with @nyuniversity: 'Dream-Cubed: Controllable Generative Modeling in Minecraft by Training

Super interesting new work from NVIDIA. (bookmark it) They suggest building agents as Python objects. Very cool idea and I think it could a …

HardwareDGX agent

Super interesting new work from NVIDIA. (bookmark it) They suggest building agents as Python objects. Very cool idea and I think it could a lot with agent reliability. More below: Agent development to

The median task consumed nearly 6x more tokens in Claude Code than in Kimi Code: - 61k in Kimi Code - 67k in Hermes - 340k in Claude Code At…

Model ReleasesDGX agent

The median task consumed nearly 6x more tokens in Claude Code than in Kimi Code: - 61k in Kimi Code - 67k in Hermes - 340k in Claude Code At K3's 3/M input rate (input tokens make up roughly 95% of ag

The number of times I have had to tell Opus 5 and Fable 5 to talk to me like an excited teenager and not a computer science phd is high. Was…

Model ReleasesDGX agent

The number of times I have had to tell Opus 5 and Fable 5 to talk to me like an excited teenager and not a computer science phd is high. Was going through a Google auth thing and got multiple paragrap

The openai agent sandbox escape actually has very real implications for the diffusion of AI in the enterprise. The incident showed the power…

AgentsDGX agent

The openai agent sandbox escape actually has very real implications for the diffusion of AI in the enterprise. The incident showed the power and capability of agents, and the need to harden systems an

The problem: agent workflows alternate between GPU-heavy reasoning and GPU-idle waiting on tools. Run hundreds concurrently and their KV cac…

SafetyDGX agent

The problem: agent workflows alternate between GPU-heavy reasoning and GPU-idle waiting on tools. Run hundreds concurrently and their KV caches fight for memory. Engines evict on a dumb LRU policy, ev

The root cause: request-level engines never see that a series of LLM calls belongs to one longer workflow. ThunderAgent adds that missing vi…

AgentsDGX agent

The root cause: request-level engines never see that a series of LLM calls belongs to one longer workflow. ThunderAgent adds that missing view. It treats each agent workflow as a schedulable program,

These optimizations across our stack compound to unlock the most performant models at every point in the cost-intelligence curve. https://op…

Model ReleasesDGX agent

OpenAI announced the release of GPT‑5.6 Sol after deployment, incorporating optimizations across its stack to enhance run‑time efficiency. The update delivers a roughly 20 % reduction in serving costs

This is a willfully misleading narrative from OpenAI. Sam’s earnest expressions are being deployed, again, to misdirect. If there’s a need t…

Model ReleasesDGX agent

This is a willfully misleading narrative from OpenAI. Sam’s earnest expressions are being deployed, again, to misdirect. If there’s a need to “pace AI development”, the OAI hack incident isn’t any evi

Thrilled to have @gabepereyra speak at our @sequoia event tmrw on OWN YOUR AI: how to build your own Lab as an application company. Also fea…

Model ReleasesDGX agent

Thrilled to have @gabepereyra speak at our @sequoia event tmrw on OWN YOUR AI: how to build your own Lab as an application company. Also featuring @FireworksAI_HQ @mercor_ai @LangChain @trajectorylabs

Today we’re open-sourcing Numbat, an agent-detection and response layer that is designed to work across agent harnesses. Numbat gives securi…

AgentsDGX agent

Today we’re open-sourcing Numbat, an agent-detection and response layer that is designed to work across agent harnesses. Numbat gives security teams visibility into agent activity, with controls to bl

Two of the people most responsible for scaling the transformer are now betting on a next act. @MillionInt ran the Reasoning 🍓 team at OpenA…

Model ReleasesDGX agent

Two of the people most responsible for scaling the transformer are now betting on a next act. @MillionInt ran the Reasoning 🍓 team at OpenAI. @_arohan_ was a pre-training lead on Gemini after years at

// Unfolding Sub-Agents for Long-Horizon ML Engineering // Watch a single agent work a machine learning engineering task for six hours you s…

AgentsDGX agent

// Unfolding Sub-Agents for Long-Horizon ML Engineering // Watch a single agent work a machine learning engineering task for six hours you see issues like context fills with stack traces, dead experim

.@UseApolloio will be at Interrupt NYC. Apollo's AI Assistant is one of the earliest production multi-agent systems on LangGraph. Their team…

AgentsDGX agent

.@UseApolloio will be at Interrupt NYC. Apollo's AI Assistant is one of the earliest production multi-agent systems on LangGraph. Their team will share how they migrated it from a hand-rolled supervis

We believe the benefits of frontier AI should not be concentrated in a few companies and well-resourced labs. Researchers know their fields …

Model ReleasesDGX agent

We believe the benefits of frontier AI should not be concentrated in a few companies and well-resourced labs. Researchers know their fields best. Our role is to put powerful tools in their hands and h

We hope these experiments serve as a reminder that evals rarely measure models in isolation—they also measure a bundle of less visible choic…

ApplicationsDGX agent

We hope these experiments serve as a reminder that evals rarely measure models in isolation—they also measure a bundle of less visible choices about API settings, harness design, and prompting. If you

We quietly released the open-source Codex Security CLI, but Hacker News found it before we had a chance to share it here... You can now use …

Model ReleasesDGX agent

We quietly released the open-source Codex Security CLI, but Hacker News found it before we had a chance to share it here... You can now use it to scan repositories, track findings across runs, verify

We ran a large-scale distillation attack on the Kimi K3 technical report by reading it in parallel at the Hugging Face Journal Club :) https…

SafetyDGX agent

We ran a large-scale distillation attack on the Kimi K3 technical report by reading it in parallel at the Hugging Face Journal Club :) https://youtu.be/MW8-kqd2SD8?si=jSKDogcUWbJ8N2k7 Our main takeawa

We’re giving scientists, mathematicians, and engineers free access to our frontier models—starting with 10,000 researchers and expanding to …

Model ReleasesDGX agent

We’re giving scientists, mathematicians, and engineers free access to our frontier models—starting with 10,000 researchers and expanding to 100,000 through 2027. ChatGPT for Academic Researchers is bu

Who on earth came up with 5 hour limits? The day is 24 hours long. While do our limit times shift by an hour each day. Pls if you are going …

Model ReleasesDGX agent

Who on earth came up with 5 hour limits? The day is 24 hours long. While do our limit times shift by an hour each day. Pls if you are going to have limits 4 or 6 hours. Hello people of Sol! I've reset

28 Jul 2026

Across eight case studies spanning industry and academia, we explore what this shift means for scientific computing—and why human verificati…

AgentsDGX agent

Across eight case studies spanning industry and academia, we explore what this shift means for scientific computing—and why human verification, stewardship, and long-term maintenance matter. https://o

AGI smarter than the smartest humans, my ass. @Kasparov63 (peak rating 2851) probably could’ve crushed the best commercial large language mo…

SafetyDGX agent

AGI smarter than the smartest humans, my ass. @Kasparov63 (peak rating 2851) probably could’ve crushed the best commercial large language models in chess when he was 7 years old. graph courtesy @chess

AI21 joined @nvidia, @Microsoft , @a16z, and dozens of others in signing the Open Weights and American AI Leadership letter. We build agenti…

SafetyDGX agent

AI21 joined @nvidia, @Microsoft , @a16z, and dozens of others in signing the Open Weights and American AI Leadership letter. We build agentic systems and agent optimization products on top of models,

Alibaba has released Qwen Audio 3.0 Realtime, with the Plus variant debuting as the new #1 model on the Artificial Analysis Speech to Speech…

Model ReleasesDGX agent

Alibaba has released Qwen Audio 3.0 Realtime, with the Plus variant debuting as the new #1 model on the Artificial Analysis Speech to Speech Index at 84.1%, ahead of GPT-Realtime-2.1 High at 79.1% Rel

← Previous
1…89101112…232
Next →