AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
4,322 results
Model Releases

so proud to be working with a bunch of people who are absolutely crushing it! check out all these new things if you haven’t had a chance to …

DGX agent

so proud to be working with a bunch of people who are absolutely crushing it! check out all these new things if you haven’t had a chance to yet. huge week for us here @LangChain!! big week at langchai

model-releasesharrison-chase--x
2 Jul 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Hardware

🌍 World’s Largest Hermes Buildathon 10 cities. 5 countries. $750K credits Backed by OpenAI, Convex, Cloudflare, ElevenLabs, Linkup, Dodo Pa…

DGX agent

🌍 World’s Largest Hermes Buildathon 10 cities. 5 countries. $750K credits Backed by OpenAI, Convex, Cloudflare, ElevenLabs, Linkup, Dodo Payments, Hissa Fund & Wispr Flow. Proudly presented by GrowthX

hardwarenous-research--x
2 Jul 2026
Research

ENPIRE -> ASPIRE, our 2nd work in the series for Physical AutoResearch. We are building the components for robot self-improvement, one /skil…

DGX agent

ENPIRE -> ASPIRE, our 2nd work in the series for Physical AutoResearch. We are building the components for robot self-improvement, one /skill at a time. Today, we give robots a /skills library that se

researchjim-fan--x
1 Jul 2026
Model Releases

You really need to benchmark models for your use case. As soon as judgements & decisions stack on top of each other, the differences between…

DGX agent

You really need to benchmark models for your use case. As soon as judgements & decisions stack on top of each other, the differences between models amplifies, and no standard benchmark will tell you t

model-releasesethan-mollick--x
1 Jul 2026
Tutorials

Showing how to display dynamic subagents was tricky - this is what we aligned on

DGX agent

This post likely discusses technical decisions and design patterns for dynamically displaying subagents within an AI system, documenting the alignment the team reached after encountering implementatio

tutorialsharrison-chase--x
30 Jun 2026
Model Releases

𝗚𝗟𝗠-𝟱.𝟮 (the latest open weights model) is having an Enterprise moment, and it is not an exaggeration.🚀 🔥 We have been impressed by h…

DGX agent

𝗚𝗟𝗠-𝟱.𝟮 (the latest open weights model) is having an Enterprise moment, and it is not an exaggeration.🚀 🔥 We have been impressed by how strongly GLM-5.2 is pushing long-horizon performance .. not just

model-releasesclem-delangue--x
30 Jun 2026
Model Releases

We’re shipping two major updates to streamline your creative workflow, allowing you to generate high-speed images with one model and then in…

DGX agent

We’re shipping two major updates to streamline your creative workflow, allowing you to generate high-speed images with one model and then instantly animate them with the other—all at a fraction of the

model-releasesgoogle-ai--x
30 Jun 2026
Tools

because i'm not a design engineer myself, this track is one of the harder ones I struggle to curate. very fortunate to befriend Geoff who ha…

DGX agent

because i'm not a design engineer myself, this track is one of the harder ones I struggle to curate. very fortunate to befriend Geoff who has lent a hand to the past 2 years of AI UX meetups, and now

toolsswyx--x
29 Jun 2026
Model Releases

Chess engines tell you the best move. But grandmasters are human, they don’t always play it. So I built 'Kibitz': a human move predictor for…

DGX agent

Chess engines tell you the best move. But grandmasters are human, they don’t always play it. So I built 'Kibitz': a human move predictor for chess broadcasts. I trained this model on my Nvidia RTX 508

model-releasesnous-research--x
29 Jun 2026
Model Releases

Most popular model on @OpenRouter (10tr tokens) turns out to be a 1.6tr MoE by @Meituan_LongCat (superapp/DoorDash of China) Basically Gemin…

DGX agent

Most popular model on @OpenRouter (10tr tokens) turns out to be a 1.6tr MoE by @Meituan_LongCat (superapp/DoorDash of China) Basically Gemini / Opus 4.6 level 35tr tokens trained entirely on 50k Chine

model-releasesemad-mostaque--x
29 Jun 2026
Industry

Après avoir vu Elon répondre au Programme alimentaire mondial de l'ONU qui lui réclamait 6 milliards pour 'résoudre la faim dans le monde', …

DGX agent

Après avoir vu Elon répondre au Programme alimentaire mondial de l'ONU qui lui réclamait 6 milliards pour 'résoudre la faim dans le monde', j'ai compris quelque chose que je vais essayer de prouver ic

industryelon-musk--x
28 Jun 2026
Applications

Harness: ✅ (DeepAgents) Sandboxes: ✅ (LangSmith Sandboxes) Eval: ✅ (LangSmith Sandboxes Model: integrate with all the popular models and pro…

DGX agent

Harness: ✅ (DeepAgents) Sandboxes: ✅ (LangSmith Sandboxes) Eval: ✅ (LangSmith Sandboxes Model: integrate with all the popular models and providers Plus we have the engine that helps you turn this flyw

applicationsharrison-chase--x
28 Jun 2026
Applications

My conversation with @ScottWu46, founder and CEO of @Cognition, the company behind Devin, the first AI software engineer. 0:00 Scott Wu's Ob…

DGX agent

My conversation with @ScottWu46, founder and CEO of @Cognition, the company behind Devin, the first AI software engineer. 0:00 Scott Wu's Obsession With Winning 2:06 Competitive Programming, Games And

applicationscognition-ai--x
28 Jun 2026
Tutorials

NEW paper worth reading. Reasoning-data curation is expensive because scoring a trace usually means reading it to the end. This new work fro…

DGX agent

NEW paper worth reading. Reasoning-data curation is expensive because scoring a trace usually means reading it to the end. This new work from UCLA shows you may not have to. The quality of a reasoning

tutorialsdair-ai--x
28 Jun 2026
Model Releases

Why do RL runs on LLMs blow up even when the recipe looks right? GEOALIGN, from the Alibaba team behind Qwen, points at the rollouts. A hand…

DGX agent

Why do RL runs on LLMs blow up even when the recipe looks right? GEOALIGN, from the Alibaba team behind Qwen, points at the rollouts. A handful of bad batches push the policy in incoherent directions,

model-releasesdair-ai--x
28 Jun 2026
Tutorials

If you use LLM-as-judge, this one is worth reading. (bookmark it) It's actually one of the most effective ways to use LLM-as-a-Judge for eva…

DGX agent

If you use LLM-as-judge, this one is worth reading. (bookmark it) It's actually one of the most effective ways to use LLM-as-a-Judge for evals. Holistic judge scores hide both their reasoning and thei

tutorialsdair-ai--x
27 Jun 2026
Model Releases

LiteParse is unreasonably good for document parsing ✅ It is the fastest document parsing tool out there - average parse time per page is 3ms…

DGX agent

LiteParse is unreasonably good for document parsing ✅ It is the fastest document parsing tool out there - average parse time per page is 3ms ⚡️⚡️ ✅ Now that we support markdown, it tops opendataloader

model-releasesjerry-liu--x
27 Jun 2026
Hardware

NEW paper from NVIDIA. (bookmark it) Speed-of-light performance analysis tells you the theoretical floor of a workload, but teams still deri…

DGX agent

NEW paper from NVIDIA. (bookmark it) Speed-of-light performance analysis tells you the theoretical floor of a workload, but teams still derive it by hand and freeze it. SOLAR automates the whole thing

hardwaredair-ai--x
27 Jun 2026
Model Releases

Run Ornith with Ollama: ollama run ornith For coding, use it with Claude or Pi: ollama launch claude --model ornith ollama launch pi --model…

DGX agent

Run Ornith with Ollama: ollama run ornith For coding, use it with Claude or Pi: ollama launch claude --model ornith ollama launch pi --model ornith For the more capable 35B model, use: ollama launch c

model-releasesollama--x
27 Jun 2026
Tutorials

Forward Deployed Engineering is the most critical ingredient for Enterprise AI adoption. It’s why the most important companies on the planet…

DGX agent

Forward Deployed Engineering is the most critical ingredient for Enterprise AI adoption. It’s why the most important companies on the planet are investing so heavily into it. That’s why we put togethe

tutorialsswyx--x
26 Jun 2026
Tools

I was at the first AI Engineer Summit @aiDotEngineer ~500 people, limited admission, felt like a secret. Next week it takes over Moscone Wes…

DGX agent

I was at the first AI Engineer Summit @aiDotEngineer ~500 people, limited admission, felt like a secret. Next week it takes over Moscone West: thousands of engineers, 400+ sessions. Huge props to @swy

toolsswyx--x
26 Jun 2026
Model Releases

Ok, I am interviewing the legend @trq212 on Friday. I'm planning to ask him to show me: → His Claude Code setup and how he uses /goal and /l…

DGX agent

Ok, I am interviewing the legend @trq212 on Friday. I'm planning to ask him to show me: → His Claude Code setup and how he uses /goal and /loop and dynamic workflows → How he does planning with HTML +

model-releasesthariq--x
25 Jun 2026
Model Releases

We open-source Qwen-AgentWorld-35B-A3B (MoE, 35B/3B active, 256K context) and AgentWorldBench. Two routes, one roadmap: 🔬 Build the simulat…

DGX agent

We open-source Qwen-AgentWorld-35B-A3B (MoE, 35B/3B active, 256K context) and AgentWorldBench. Two routes, one roadmap: 🔬 Build the simulator — scalable, controllable, surpassing real environments 🧠 I

model-releasesqwen--x
24 Jun 2026
Applications

Fugu stands shoulder-to-shoulder with leading models like Fable and Mythos across the industry's most rigorous engineering, scientific, and …

DGX agent

Fugu stands shoulder-to-shoulder with leading models like Fable and Mythos across the industry's most rigorous engineering, scientific, and reasoning benchmarks. Read the full blog: https://sakana.ai/

applicationsdavid-ha--x
22 Jun 2026
Model Releases

GLM-5.2 has been the most popular new model on Fireworks this past week. @ArtificialAnlys confirms why: #3 overall on GDPval-AA (1524 Elo), …

DGX agent

GLM-5.2 has been the most popular new model on Fireworks this past week. @ArtificialAnlys confirms why: #3 overall on GDPval-AA (1524 Elo), #1 open weights by 116 points. Interest is showing no signs

model-releasesfireworks-ai--x
22 Jun 2026
Model Releases

Gray Swan: Red-Teaming after Mythos & the coming AI security crisis https://www.latent.space/p/gray-swan @GraySwanAI cofounders @zicokolter …

DGX agent

Gray Swan: Red-Teaming after Mythos & the coming AI security crisis https://www.latent.space/p/gray-swan @GraySwanAI cofounders @zicokolter and Matt Fredrikson explain why AI security is fundamentally

model-releasesswyx--x
22 Jun 2026
Safety

have been thinking a bunch about model routing and related things current thoughts here, would love feedback: 1/ there is a difference betwe…

DGX agent

have been thinking a bunch about model routing and related things current thoughts here, would love feedback: 1/ there is a difference between 'model routing' and 'model council' 'model routing' = rou

safetyharrison-chase--x
22 Jun 2026
Safety

The largest LLM-as-a-Judge reliability audit yet. Researchers ran 21 judges from nine providers over roughly 541,000 judgments on MT-Bench, …

DGX agent

The largest LLM-as-a-Judge reliability audit yet. Researchers ran 21 judges from nine providers over roughly 541,000 judgments on MT-Bench, JudgeBench, and RewardBench. Findings: Validating a judge wi

safetydair-ai--x
22 Jun 2026
Hardware

MaineCoon is the first video model that focuses on social interactions: facial expressions, emotions, fluid conversation, audio-lip sync, et…

DGX agent

MaineCoon is the first video model that focuses on social interactions: facial expressions, emotions, fluid conversation, audio-lip sync, etc. Really impressive inference specs: 22B params, 47.5 FPS o

hardwarefrancois-chollet--x
21 Jun 2026
Model Releases

The more you embrace AI, the more you need SaaS. This is not obvious to armchair market analysts who love disruption narratives, but it is o…

DGX agent

The more you embrace AI, the more you need SaaS. This is not obvious to armchair market analysts who love disruption narratives, but it is obvious to people actually running companies. Levie now uses

model-releasesfrancois-chollet--x
21 Jun 2026
Tutorials

GLM-5.2 is great at design (Opus level IMO). I am also starting to see great results with long-running tasks, too. How is this possible? I t…

DGX agent

GLM-5.2 is great at design (Opus level IMO). I am also starting to see great results with long-running tasks, too. How is this possible? I think there are a few clever hacks. But I just came across th

tutorialsdair-ai--x
20 Jun 2026
Model Releases

I have some very big news... KernelBench-Hard with H100 and B200 (single gpu results) AND KernelBench-Mega tested on RTX PRO 6000, H100, B20…

DGX agent

I have some very big news... KernelBench-Hard with H100 and B200 (single gpu results) AND KernelBench-Mega tested on RTX PRO 6000, H100, B200 is finally out! Starting with Mega, each of models wrote a

model-releasesclem-delangue--x
20 Jun 2026
Model Releases

it looks like AIE 2026 is mostly sold out, for anyone who couldn't get a ticket for this year, I love that AIE would livestream and have thi…

DGX agent

it looks like AIE 2026 is mostly sold out, for anyone who couldn't get a ticket for this year, I love that AIE would livestream and have this archive of the conference and workshop. A few of my favori

model-releasesswyx--x
20 Jun 2026
Safety

// Self-play with a pinch of human data // Really cool paper combining human demonstrations and self-play RL. 30 minutes of human data, 2500…

DGX agent

// Self-play with a pinch of human data // Really cool paper combining human demonstrations and self-play RL. 30 minutes of human data, 2500x less than imitation learning, is enough to make self-play

safetydair-ai--x
20 Jun 2026
Tools

Markie is one of my favorite people. I have learned so much from her perspective that is equal parts soulful clarity about what computing sh…

DGX agent

Markie is one of my favorite people. I have learned so much from her perspective that is equal parts soulful clarity about what computing should aspire to be for humans, and the demanding complexity t

toolslinus-lee--x
10 Jun 2026
Model Releases

This is a great article on how startups/frontier labs can coexist. Another way to look at this is task complexity - the number of bits of in…

DGX agent

This is a great article on how startups/frontier labs can coexist. Another way to look at this is task complexity - the number of bits of information needed to specify a task such that AI can solve th

model-releasesjerry-liu--x
10 Jun 2026
Model Releases

BREAKING: Anthropic just dropped Claude Fable 5—this is Mythos, made safe for public release. It is the best coding model in the world. We'v…

DGX agent

BREAKING: Anthropic just dropped Claude Fable 5—this is Mythos, made safe for public release. It is the best coding model in the world. We've been testing it internally @every for the last week or so

model-releasessimon-willison--x
9 Jun 2026
Model Releases

🚨 Fable 5 is something to pay attention to. This is another 'I had early access to the new Mythos-class Anthropic model and I want to tell …

DGX agent

🚨 Fable 5 is something to pay attention to. This is another 'I had early access to the new Mythos-class Anthropic model and I want to tell you what I thought of it' post. I know, I know, it's annoying

model-releasesallie-k--miller--x
9 Jun 2026
Model Releases

Fable 5 is the biggest step up I’ve felt in our models since Opus 4.5 back in November. After 4.5 came out I uninstalled my IDE when I reali…

DGX agent

Fable 5 is the biggest step up I’ve felt in our models since Opus 4.5 back in November. After 4.5 came out I uninstalled my IDE when I realized that I’d been doing 100% of my coding in a terminal for

model-releasesboris-cherny--x
9 Jun 2026
Applications

I want to share some thoughts about the impact of AI on org structure. This is based on early conversations with some of the top CHROs in Fo…

DGX agent

I want to share some thoughts about the impact of AI on org structure. This is based on early conversations with some of the top CHROs in Fortune 500 companies as well as other C-Suite leaders and sta

applicationsallie-k--miller--x
9 Jun 2026
Tutorials

It's confusing that Trump constantly warns about the security of our elections & foreign interference, yet his administration has taken seve…

DGX agent

It's confusing that Trump constantly warns about the security of our elections & foreign interference, yet his administration has taken several steps to dismantle or divert resources away from key saf

tutorialsyann-lecun--x
9 Jun 2026
Local Ai

this model is the opposite of mythos. Its small, cost effective, apache 2.0, and locally deployable. This is the way LLMs should go. small, …

DGX agent

this model is the opposite of mythos. Its small, cost effective, apache 2.0, and locally deployable. This is the way LLMs should go. small, open source, transparent and sovereign vs large, expensive,

local-aiclem-delangue--x
9 Jun 2026
Hardware

Good take My guess is - demand for intelligence is near infinite - but 80% of workloads will be running on 99% cheaper models within 12-18 m…

DGX agent

Good take My guess is - demand for intelligence is near infinite - but 80% of workloads will be running on 99% cheaper models within 12-18 months - 20% of workloads will still run on latest gen models

hardwareclem-delangue--x
8 Jun 2026
Applications

For some industries and dev orgs, this will happen already in 2026. For others, it may take until 2030. A short list of key differences for …

DGX agent

For some industries and dev orgs, this will happen already in 2026. For others, it may take until 2030. A short list of key differences for whether orgs will reach that level soon or later: • How much

applicationsitamar-friedman--x
7 Jun 2026
Tutorials

Super-powerful AI models will launch in the coming weeks. We are looking at a potential step change in model capabilities. The biggest mista…

DGX agent

Super-powerful AI models will launch in the coming weeks. We are looking at a potential step change in model capabilities. The biggest mistake right now is to lock into one vendor. I say this not only

tutorialsdair-ai--x
7 Jun 2026
Model Releases

Your margin is my opportunity: AI version… The biggest surprise of 2026 is that the capability gap between the best open-weight/source model…

DGX agent

Your margin is my opportunity: AI version… The biggest surprise of 2026 is that the capability gap between the best open-weight/source models and the best closed models has narrowed much faster than t

model-releasesclem-delangue--x
6 Jun 2026
Tools

'Reality: The Final Eval' — 現実タスク完了率こそが最終評価指標(@swyx / Andon Labs)。 複数のAI実装を並列で回していると、実感として正確だと思う。SWE-Benchの数字より「本番で動くか」が判断軸。エージェント設計で最初に決めるの…

DGX agent

# Reality: The Final Eval This post argues that real-world task completion rate is the ultimate metric for evaluating AI systems, prioritizing actual production performance over benchmark scores like

toolsswyx--x
5 Jun 2026
Model Releases

The NVIDIA Nemotron Coalition continues to grow. We're excited to welcome new members: @hcompany_ai, @NousResearch, and @PrimeIntellect. And…

DGX agent

The NVIDIA Nemotron Coalition continues to grow. We're excited to welcome new members: @hcompany_ai, @NousResearch, and @PrimeIntellect. And a big thank you to our existing members: @bfl_ai, @cursor_a

model-releasesnous-research--x
5 Jun 2026
← Previous
1…8384858687…91
Next →