AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “engineering”

GridTimelineEvolution
1,333 results
CompaniesToolsTechniques

Each lane shows up to 8 recent matching entries, ordered from earlier to later. Tracks load separately to keep the 75,000+ entry wiki fast.

Companies

CompanyAnthropic8 recent entries
30 Jul 2026Yesterday I cohosted a dinner with @dexhorthy with a wonderful group of founders, to talk about agent loops and loop engineering. Some inter…

Yesterday I cohosted a dinner with @dexhorthy with a wonderful group of founders, to talk about agent loops and loop engineering. Some interesting insights: * Most of our group was *not* actively usin

→30 Jul 2026Can AI agents conduct open-ended AI research? Most evaluations of agents conducting AI research focus on narrow, verifiable tasks. But AI re…

Can AI agents conduct open-ended AI research? Most evaluations of agents conducting AI research focus on narrow, verifiable tasks. But AI research is often open ended. Researchers pick hypotheses, dec

3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
→1 Aug 2026Hot take on OpenAI’s Astra: - Obviously impressive - But math is different from most other problems in that it is more amenable to to formal…

Hot take on OpenAI’s Astra: - Obviously impressive - But math is different from most other problems in that it is more amenable to to formal verification and synthetic data. How well it works in open-

→4 Aug 2026OpenAI and Anthropic have both just posted about an overlapping cyber incident involving GPT-5.6-Sol and Mythos 5 during an evaluation by UK…

OpenAI and Anthropic have both just posted about an overlapping cyber incident involving GPT-5.6-Sol and Mythos 5 during an evaluation by UKAISI. I will quote: 'In the most serious case, an agent trie

→5 Aug 2026Yes, the AIs were given a cybersecurity challenge, with internet access enabled and safety filters disabled. But the extent to which Mythos …

Yes, the AIs were given a cybersecurity challenge, with internet access enabled and safety filters disabled. But the extent to which Mythos 5 pursued its mission (fake identities, social engineering,

→5 Aug 2026Some people are surprised that APIs (aka what Anthropic, OpenAI, and others provide) are treated differently than open weights in the new AI…

Some people are surprised that APIs (aka what Anthropic, OpenAI, and others provide) are treated differently than open weights in the new AI model framework. I'm not surprised at all, and it's actuall

→10 Aug 2026Muse Glimmer is now available to run with Ollama. Available today via Ollama’s MLX engine with state-of-the-art-performance on Apple Silicon…

Muse Glimmer is now available to run with Ollama. Available today via Ollama’s MLX engine with state-of-the-art-performance on Apple Silicon, Muse Glimmer can power Claude Code, Codex, and more always

→11 Aug 2026we recently trimmed the deepagents harness base prompt by 65% (including tool info) it shows — deepagents is cheap!

we recently trimmed the deepagents harness base prompt by 65% (including tool info) it shows — deepagents is cheap! We ran DeepSeek V4 Flash through 4 more agent harnesses (Hermes Agent, Pi Agent, Pri

CompanyOpenAI8 recent entries
29 Jul 2026It's a drop-in. One program_id field and it plugs into your existing engine configs such as KV offloading + speculative decoding. Format-agn…

It's a drop-in. One program_id field and it plugs into your existing engine configs such as KV offloading + speculative decoding. Format-agnostic by design, with OpenAI chat completions supported toda

→1 Aug 2026Hot take on OpenAI’s Astra: - Obviously impressive - But math is different from most other problems in that it is more amenable to to formal…

Hot take on OpenAI’s Astra: - Obviously impressive - But math is different from most other problems in that it is more amenable to to formal verification and synthetic data. How well it works in open-

→2 Aug 2026When asked if AI developers have lost control of their technology after an AI agent created by OpenAI escaped its testing environment and ha…

When asked if AI developers have lost control of their technology after an AI agent created by OpenAI escaped its testing environment and hacked another company, Hugging Face CEO Clément Delangue says

→4 Aug 2026OpenAI and Anthropic have both just posted about an overlapping cyber incident involving GPT-5.6-Sol and Mythos 5 during an evaluation by UK…

OpenAI and Anthropic have both just posted about an overlapping cyber incident involving GPT-5.6-Sol and Mythos 5 during an evaluation by UKAISI. I will quote: 'In the most serious case, an agent trie

→5 Aug 2026Some people are surprised that APIs (aka what Anthropic, OpenAI, and others provide) are treated differently than open weights in the new AI…

Some people are surprised that APIs (aka what Anthropic, OpenAI, and others provide) are treated differently than open weights in the new AI model framework. I'm not surprised at all, and it's actuall

→5 Aug 2026None of this was us. Day 1 (and the first 48 hours) belonged to the open-source community: Generate with it @ComfyUI — native support + offi…

None of this was us. Day 1 (and the first 48 hours) belonged to the open-source community: Generate with it @ComfyUI — native support + official quantized builds, Day 0 Diffusers — the reference Pytho

→5 Aug 2026if you have been following his excellent work, @shloked has been breaking down every frontier labs' harness engineering for the last few mon…

if you have been following his excellent work, @shloked has been breaking down every frontier labs' harness engineering for the last few months. excited to publish his deepest dive into ChatGPT yet as

→7 Aug 2026Basically every remaining good AI benchmark score has an implied asterisk next to it which reads: * could be signficantly higher with a bett…

On August 7, 2026 Ethan Mollick tweeted that “every remaining good AI benchmark score has an implied asterisk next to it which reads: * could be significantly higher with a better harness.” The commen

CompanyGoogle8 recent entries
30 Jun 2026Thrilled to announce the Wearable AI Workshop at ECCV 2026 🎉 that we're organizing with an awesome group of folks across Meta Reality Labs,…

Thrilled to announce the Wearable AI Workshop at ECCV 2026 🎉 that we're organizing with an awesome group of folks across Meta Reality Labs, AMI Labs, HKUST, Georgia Tech, UCF, and U. of Edinburgh. If

→1 Jul 2026Gemini Omni Flash is now in ComfyUI via partner nodes. Google's state of the art model with physics and video editing capabilities is now in…

Gemini Omni Flash is now in ComfyUI via partner nodes. Google's state of the art model with physics and video editing capabilities is now inside the world's most strongest generative AI workflow engin

→8 Jul 2026- OpenAI continually outperforms Anthropic models on computer use (with Claude, I’m surprised when it works, with Codex, I expect it to work…

- OpenAI continually outperforms Anthropic models on computer use (with Claude, I’m surprised when it works, with Codex, I expect it to work) - when I prompt with Fable 5.5, I feel like I’m motivating

→14 Jul 2026Exactly one year ago was the craziest 72 hours I’d experience in Silicon Valley. I was thrust into a situation where many members of the Win…

Exactly one year ago was the craziest 72 hours I’d experience in Silicon Valley. I was thrust into a situation where many members of the Windsurf team had moved on to Google, recruiters were reaching

→20 Jul 2026Huge launch from @tryramp. Different steps in an agent workflow can use different models. This can help reduce costs significantly without s…

Huge launch from @tryramp. Different steps in an agent workflow can use different models. This can help reduce costs significantly without sacrificing performance. Model routing will become a core par

→26 Jul 2026BREAKING: A Redditor just discovered that shared Claude conversations have been showing up in public search results. The post has 4K upvotes…

BREAKING: A Redditor just discovered that shared Claude conversations have been showing up in public search results. The post has 4K upvotes and hundreds of comments, so this is spreading fast. Here's

→4 Aug 2026one thing i appreciate about silico is that it's a deeply humanist product. we designed silico to keep you in the experimental loop -- more …

one thing i appreciate about silico is that it's a deeply humanist product. we designed silico to keep you in the experimental loop -- more observable, easier to steer, easier to understand we want to

→9 Aug 2026The best 'raw' frontier model for document parsing is gemini 3 flash, but the issue is that since then the flash models have gotten 3x more …

The best 'raw' frontier model for document parsing is gemini 3 flash, but the issue is that since then the flash models have gotten 3x more expensive while flatlining on visual recognition across comp

CompanyMeta8 recent entries
13 Jul 2026New model for AMD Strix Halo users: My 198B Step 3.7 Flash release was a big hit, but this one may be even better: 298B-parameter Hy3, now r…

New model for AMD Strix Halo users: My 198B Step 3.7 Flash release was a big hit, but this one may be even better: 298B-parameter Hy3, now running on a new 2-bit FPX codebook designed to map efficient

→15 Jul 2026We're part of the Amazon Web Services (AWS) AI Builder Lab in New York on Friday, July 24 - a Clash of Agents competition with OpenAI, LangC…

We're part of the Amazon Web Services (AWS) AI Builder Lab in New York on Friday, July 24 - a Clash of Agents competition with OpenAI, LangChain, HiddenLayer, Protopia AI, Fiddler AI, and Coder. One d

→23 Jul 2026Dynamic workflows are a generalization of harnesses, automations, loops, routing, and graphs. It's the most powerful feature I have built in…

Dynamic workflows are a generalization of harnesses, automations, loops, routing, and graphs. It's the most powerful feature I have built into my agent orchestrator. Supports all kinds of patterns tha

→30 Jul 2026I’m excited to share my next chapter: I’ve joined @FireworksAI_HQ . From AMD, Apple, Uber, Meta, Google, and most recently Snowflake, I’ve w…

I’m excited to share my next chapter: I’ve joined @FireworksAI_HQ . From AMD, Apple, Uber, Meta, Google, and most recently Snowflake, I’ve worked on many of the foundational technologies that power mo

→31 Jul 2026Very interesting paper on recursive self-improvement. The whole stack is released. Machine learning engineering gives recursive self-improve…

Very interesting paper on recursive self-improvement. The whole stack is released. Machine learning engineering gives recursive self-improvement a concrete, executable testbed. OpenMLE is an open full

→3 Aug 2026You probably used @FireworksAI_HQ this week. The cool part is that you just did not know you did. 🎆 Open-source models are free, sure... bu…

You probably used @FireworksAI_HQ this week. The cool part is that you just did not know you did. 🎆 Open-source models are free, sure... but the hard part is tailoring them so they perform best on you

→10 Aug 2026Impressive new paper from Meta. (bookmark it) Scaling laws assume model size and training data act on loss independently. This work introduc…

Impressive new paper from Meta. (bookmark it) Scaling laws assume model size and training data act on loss independently. This work introduces Skaling law, which couples capacity and data through a si

→10 Aug 2026Got Meta's new Muse Glimmer 30B running on my MacBook (M3 Max, 96GG) and tested the serving options available so far. Fastest right now: Oll…

Got Meta's new Muse Glimmer 30B running on my MacBook (M3 Max, 96GG) and tested the serving options available so far. Fastest right now: Ollama's MLX engine (DFlash included) at ~29 tok/s. Tuned llama

CompanyMistral2 recent entries
1 May 2026I have been testing DeepSeek-V4-Pro with the Pi coding agent. I am mindblown by how well it works out of the box. A few notes: I spent a few…

I have been testing DeepSeek-V4-Pro with the Pi coding agent. I am mindblown by how well it works out of the box. A few notes: I spent a few hours building an LLM wiki with an agent powered entirely b

→11 Aug 2026💡The world needs an open-source platform, and that’s exactly what we’re building to give our customers more choice and the flexibility to c…

💡The world needs an open-source platform, and that’s exactly what we’re building to give our customers more choice and the flexibility to choose the right model for the right task. As part of this, we

CompanyxAI8 recent entries
9 Jul 2026The upcoming wave of SpaceXAI Grok updates is insane Grok 4.5: The 1.5T foundation model is being refined almost daily, and its context wind…

The upcoming wave of SpaceXAI Grok updates is insane Grok 4.5: The 1.5T foundation model is being refined almost daily, and its context window is expected to jump to 1M tokens, possibly as soon as nex

→9 Jul 2026I appreciate the many xAI and Cursor engineers who dedicated their time to addressing feedback from Tesla. I remember meeting Andrew a few w…

I appreciate the many xAI and Cursor engineers who dedicated their time to addressing feedback from Tesla. I remember meeting Andrew a few weeks after he was hired and telling him that I needed a bett

→9 Jul 2026Grok 4.5 is also rank 1 in SWE marathon

Grok 4.5, xAI's large language model, has achieved the top ranking in the SWE (Software Engineering) Marathon benchmark, according to an announcement by Elon Musk. This ranking suggests the model demo

→11 Jul 2026Grok places second after Fable on real-world software engineering

Grok places second after Fable on real-world software engineering Grok 4.5 from @SpaceXAI places #2 on the APEX-SWE leaderboard at 51.2% Pass@1 (±6.0), behind Fable 5 (65.5% ±6.2) on our benchmark for

→13 Jul 2026NEW: Grok 4.5 now scores highest on the SWE-Atlas-QnA benchmark, edging Claude Fable 5 & GPT-5.6 Sol.

Grok 4.5 achieved the highest score on the SWE‑Atlas‑QnA benchmark, surpassing Claude Fable 5 and GPT‑5.6 Sol, according to a Polymarket tweet posted on 13 July 2026. The update highlights Grok 4.5’s

→14 Jul 2026Grok is the Flow LLM. Grok 4.5’s biggest advantage is its speed. It’s smart enough to be comparable to the other models on most things. But …

Grok is the Flow LLM. Grok 4.5’s biggest advantage is its speed. It’s smart enough to be comparable to the other models on most things. But that speed allows you to make little tweaks to your system s

→15 Jul 2026BREAKING: Grok 4.5 has climbed to #2 on the FrontierSWE benchmark. The result places Grok 4.5 among the world's top-performing AI models for…

BREAKING: Grok 4.5 has climbed to #2 on the FrontierSWE benchmark. The result places Grok 4.5 among the world's top-performing AI models for software engineering tasks, highlighting its growing streng

→20 Jul 2026Huge launch from @tryramp. Different steps in an agent workflow can use different models. This can help reduce costs significantly without s…

Huge launch from @tryramp. Different steps in an agent workflow can use different models. This can help reduce costs significantly without sacrificing performance. Model routing will become a core par

CompanyDeepSeek8 recent entries
3 Aug 2026You probably used @FireworksAI_HQ this week. The cool part is that you just did not know you did. 🎆 Open-source models are free, sure... bu…

You probably used @FireworksAI_HQ this week. The cool part is that you just did not know you did. 🎆 Open-source models are free, sure... but the hard part is tailoring them so they perform best on you

→3 Aug 2026two weeks ago i went on @swyx's pod and said some things that i... should not have said. a lot has happened since then, i owe you all an apo…

two weeks ago i went on @swyx's pod and said some things that i... should not have said. a lot has happened since then, i owe you all an apology. i'm sorry that i was right about every single thing. a

→7 Aug 2026We analyzed DeepSeek-V4 Flash-0731 vs. GPT-5.6 Luna on software engineering tasks using DeepSWE. DeepSeek-V4 Flash-0731 delivers 80% of Luna…

We analyzed DeepSeek-V4 Flash-0731 vs. GPT-5.6 Luna on software engineering tasks using DeepSWE. DeepSeek-V4 Flash-0731 delivers 80% of Luna’s performance at roughly 1/6 the cost. More insights in the

→7 Aug 2026Basically every remaining good AI benchmark score has an implied asterisk next to it which reads: * could be signficantly higher with a bett…

On August 7, 2026 Ethan Mollick tweeted that “every remaining good AI benchmark score has an implied asterisk next to it which reads: * could be significantly higher with a better harness.” The commen

→8 Aug 2026We analyzed DeepSeek V4 Flash and GPT-5.6 Luna on DeepSWE. A DeepSeek-first cascade with test-suite verification solved MORE tasks than Luna…

Researchers from TogetherAI analyzed DeepSeek V4 Flash and GPT‑5.6 Luna on the DeepSWE benchmark. The study found that employing a DeepSeek‑first cascade with test‑suite verification solved more tasks

→9 Aug 2026We compared how far the same budget goes with DeepSeek V4 Flash and GPT-5.6 Luna on DeepSWE. Two DeepSeek V4 Flash attempts solved MORE task…

We compared how far the same budget goes with DeepSeek V4 Flash and GPT-5.6 Luna on DeepSWE. Two DeepSeek V4 Flash attempts solved MORE tasks than one Luna attempt at roughly one-third the cost. Media

→11 Aug 2026we recently trimmed the deepagents harness base prompt by 65% (including tool info) it shows — deepagents is cheap!

we recently trimmed the deepagents harness base prompt by 65% (including tool info) it shows — deepagents is cheap! We ran DeepSeek V4 Flash through 4 more agent harnesses (Hermes Agent, Pi Agent, Pri

→11 Aug 2026DeepSeek V4 Flash 0731 is now available to fine-tune on Together AI. Specialize it for coding, tool use, and your own domain with SFT or DPO…

DeepSeek V4 Flash 0731 is now available to fine-tune on Together AI. Specialize it for coding, tool use, and your own domain with SFT or DPO, then deploy the fine-tuned model on Together AI for produc

CompanyNVIDIA8 recent entries
24 Jul 2026The knowledge that makes AI useful is diffused. It lives with scientists, engineers, clinicians, firms. For AI to benefit from distributed k…

The knowledge that makes AI useful is diffused. It lives with scientists, engineers, clinicians, firms. For AI to benefit from distributed knowledge, it must itself be distributed. Agree with Jensen t

→29 Jul 2026It's a drop-in. One program_id field and it plugs into your existing engine configs such as KV offloading + speculative decoding. Format-agn…

It's a drop-in. One program_id field and it plugs into your existing engine configs such as KV offloading + speculative decoding. Format-agnostic by design, with OpenAI chat completions supported toda

→3 Aug 2026two weeks ago i went on @swyx's pod and said some things that i... should not have said. a lot has happened since then, i owe you all an apo…

two weeks ago i went on @swyx's pod and said some things that i... should not have said. a lot has happened since then, i owe you all an apology. i'm sorry that i was right about every single thing. a

→5 Aug 2026None of this was us. Day 1 (and the first 48 hours) belonged to the open-source community: Generate with it @ComfyUI — native support + offi…

None of this was us. Day 1 (and the first 48 hours) belonged to the open-source community: Generate with it @ComfyUI — native support + official quantized builds, Day 0 Diffusers — the reference Pytho

→7 Aug 2026Live now: our Local AI Track from AI Engineer World's Fair 2026, brought to you by @nvidia. Thesis: frontier intelligence is becoming someth…

Live now: our Local AI Track from AI Engineer World's Fair 2026, brought to you by @nvidia. Thesis: frontier intelligence is becoming something you own. https://www.youtube.com/watch?v=KB41dTlX1Uc&lis

→10 Aug 2026Ultra-High Interactivity on NVIDIA GPUs? TileRT InferenceX Can TileRT software on NVIDIA GPU compete with Cerebras, Groq LPU, SambaNova? Bat…

Ultra-High Interactivity on NVIDIA GPUs? TileRT InferenceX Can TileRT software on NVIDIA GPU compete with Cerebras, Groq LPU, SambaNova? Batch Size 1, Disaggregated engine, High throughput prefill eng

→10 Aug 2026Muse Glimmer is now available to run with Ollama. Available today via Ollama’s MLX engine with state-of-the-art-performance on Apple Silicon…

Muse Glimmer is now available to run with Ollama. Available today via Ollama’s MLX engine with state-of-the-art-performance on Apple Silicon, Muse Glimmer can power Claude Code, Codex, and more always

→11 Aug 2026Looking for a faster specialized model for your Agent Work? @nvidia Nemotron 3.5 Lightning (30B MoE, 3B active params) is now live on Firewo…

Looking for a faster specialized model for your Agent Work? @nvidia Nemotron 3.5 Lightning (30B MoE, 3B active params) is now live on Fireworks. It’s distilled from NVIDIA Nemotron 3 Ultra to be your