AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlog
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “dair-ai--x”

GridTimelineEvolution
359 results
Model Releases

Great to see the new GPT-5.6 models finally announced. Sad to see this new release strategy where only a select few get access initially. No…

DGX agent

Great to see the new GPT-5.6 models finally announced. Sad to see this new release strategy where only a select few get access initially. Not a win for our industry IMO. Open-source AI must win! Intro

model-releasesdair-ai--x
26 Jun 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Highly-recommended reading. Interesting details in this METR's GPT-5.6 eval. They couldn't get a clean capability number because the model c…

DGX agent

Highly-recommended reading. Interesting details in this METR's GPT-5.6 eval. They couldn't get a clean capability number because the model cheated more than any public model they've tested, and even r

model-releasesdair-ai--x
26 Jun 2026
Model Releases

New paper on giving LLM agents experience that improves the weights and stays readable at the same time. Agent-experience methods split into…

DGX agent

New paper on giving LLM agents experience that improves the weights and stays readable at the same time. Agent-experience methods split into two camps. Externalized natural-language rules stay interpr

model-releasesdair-ai--x
26 Jun 2026
Agents

One of my best uses of agentic loops has been personal health. I don't talk about it often because it's very personal. But here it goes in t…

DGX agent

One of my best uses of agentic loops has been personal health. I don't talk about it often because it's very personal. But here it goes in the hope it helps someone who's struggling. (I am not doing t

agentsdair-ai--x
26 Jun 2026
Research

As I said, there are existing companies that are already doing this 'AI employee' stuff really well. @viktor__com is one of them. The best p…

DGX agent

As I said, there are existing companies that are already doing this 'AI employee' stuff really well. @viktor__com is one of them. The best part is you don't get locked into one model. And you really w

researchdair-ai--x
25 Jun 2026
Agents

Just had a great discussion on dynamic workflows. Rough notes: - applies to a very small set of use cases - think of it as a new paradigm of…

DGX agent

Just had a great discussion on dynamic workflows. Rough notes: - applies to a very small set of use cases - think of it as a new paradigm of (test-time compute) TTC - strong for hill-climbing research

agentsdair-ai--x
25 Jun 2026
Agents

New research from Meta. Building synthetic training data has stayed a fixed pipeline that you hand-tune and then freeze. Autodata casts an A…

DGX agent

New research from Meta. Building synthetic training data has stayed a fixed pipeline that you hand-tune and then freeze. Autodata casts an AI agent as a data scientist that builds training and evaluat

agentsdair-ai--x
25 Jun 2026
Agents

// Critique of the Agent Model // Finally, a paper that tries to define what an agent is and what agency consists of. Good read overall. (gr…

DGX agent

// Critique of the Agent Model // Finally, a paper that tries to define what an agent is and what agency consists of. Good read overall. (great bookmark) The word agent now covers everything from a fo

agentsdair-ai--x
24 Jun 2026
Agents

Finally caved in, and I now fully speak to agents as opposed to typing prompts. My first realization is that you can just blabber on and tel…

DGX agent

Finally caved in, and I now fully speak to agents as opposed to typing prompts. My first realization is that you can just blabber on and tell the agent so many rich details via audio. The longer and t

agentsdair-ai--x
24 Jun 2026
Agents

Learn *anything* with our new /learn agent skill.

DGX agent

Learn *anything* with our new /learn agent skill. Obsessed with our new /learn skill. It's my favorite way of learning and researching topics. The agent creates a learning plan and a learning hub (art

agentsdair-ai--x
24 Jun 2026
Agents

Obsessed with our new /learn skill. It's my favorite way of learning and researching topics. The agent creates a learning plan and a learnin…

DGX agent

Obsessed with our new /learn skill. It's my favorite way of learning and researching topics. The agent creates a learning plan and a learning hub (artifact) that adjusts per learner needs and progress

agentsdair-ai--x
24 Jun 2026
Agents

What is actually limiting model routers for coding tasks? Most routers treat picking a model as a static, one-off classification. This paper…

DGX agent

What is actually limiting model routers for coding tasks? Most routers treat picking a model as a static, one-off classification. This paper identifies the real bottleneck as information deficit. Simp

agentsdair-ai--x
24 Jun 2026
Agents

Highly-recommended read. It's exciting to see large-scale agentic RL becoming more accessible. Cool to see the infra layer for this is being…

DGX agent

Highly-recommended read. It's exciting to see large-scale agentic RL becoming more accessible. Cool to see the infra layer for this is being built and I think this plays an important role in self-impr

agentsdair-ai--x
23 Jun 2026
Agents

I'm digging the eve agentic framework from Vercel. I like that everything is files, from the tools to the skills to the evals. More importan…

DGX agent

I'm digging the eve agentic framework from Vercel. I like that everything is files, from the tools to the skills to the evals. More importantly, it's gets you building with agents fast. Very promising

agentsdair-ai--x
23 Jun 2026
Agents

Learn to use the new eve agentic framework from Vercel. Go try out the hands-on labs now.

DGX agent

Learn to use the new eve agentic framework from Vercel. Go try out the hands-on labs now. I'm digging the eve agentic framework from Vercel. I like that everything is files, from the tools to the skil

agentsdair-ai--x
23 Jun 2026
Research

Most AI code review tools look at one repo at a time. But the bug usually isn't in the code that changed. It's in what that change quietly b…

DGX agent

Most AI code review tools look at one repo at a time. But the bug usually isn't in the code that changed. It's in what that change quietly breaks three repos away. @QodoAI just shipped Cross Repo Revi

researchdair-ai--x
23 Jun 2026
Safety

Great report on LLM agent communication protocols. Communication is a huge bottleneck in multi-agent systems. (worth bookmarking) The report…

DGX agent

Great report on LLM agent communication protocols. Communication is a huge bottleneck in multi-agent systems. (worth bookmarking) The report builds a five-dimensional taxonomy (counterparty, payload,

safetydair-ai--x
22 Jun 2026
Agents

Guess which is Fugu Ultra? This is how recent models compare when generating endless procedural terrain (using Three.js). All of these are o…

DGX agent

Guess which is Fugu Ultra? This is how recent models compare when generating endless procedural terrain (using Three.js). All of these are one-shotted! Just wild! Trying a few more examples. Will shar

agentsdair-ai--x
22 Jun 2026
Agents

Just a glimpse of what collective AI intelligence will bring. We haven’t truly cracked multi-agent orchestration but with every new frontier…

DGX agent

Just a glimpse of what collective AI intelligence will bring. We haven’t truly cracked multi-agent orchestration but with every new frontier model, intelligence should compound. Introducing Sakana Fug

agentsdair-ai--x
22 Jun 2026
Agents

OMG! Fugu Ultra is ridiculously good at these 3D renders.

DGX agent

OMG! Fugu Ultra is ridiculously good at these 3D renders. Media Introducing Sakana Fugu: A full multi-agent orchestration system accessible via a single model API. Our ‘Fugu Ultra’ model matches the p

agentsdair-ai--x
22 Jun 2026
Safety

The largest LLM-as-a-Judge reliability audit yet. Researchers ran 21 judges from nine providers over roughly 541,000 judgments on MT-Bench, …

DGX agent

The largest LLM-as-a-Judge reliability audit yet. Researchers ran 21 judges from nine providers over roughly 541,000 judgments on MT-Bench, JudgeBench, and RewardBench. Findings: Validating a judge wi

safetydair-ai--x
22 Jun 2026
Agents

>> Scalable Evaluation for AI Agents << If you run agent evaluation in production, this one is worth your time. It shows that front-loading …

DGX agent

>> Scalable Evaluation for AI Agents << If you run agent evaluation in production, this one is worth your time. It shows that front-loading human judgment into reusable evaluation assets is useful. Bu

agentsdair-ai--x
21 Jun 2026
Research

The Top AI Papers of the Week (June 14 - June 21): - PreAct - SpatialClaw - Back on Track - OpenClaw-Skill - From Trainee to Trainer - Compo…

DGX agent

The Top AI Papers of the Week (June 14 - June 21): - PreAct - SpatialClaw - Back on Track - OpenClaw-Skill - From Trainee to Trainer - Compositional Skill Routing - Can LLM Agents Infer World Models?

researchdair-ai--x
21 Jun 2026
Model Releases

Very impressive from GLM-5.2. Frontier open-weight model indeed. Now, can we get a Gemini model in the top 3 soon?

DGX agent

Very impressive from GLM-5.2. Frontier open-weight model indeed. Now, can we get a Gemini model in the top 3 soon? GLM 5.2 is now on DeepSWE as the top open-source model on our leaderboard. With a pas

model-releasesdair-ai--x
21 Jun 2026
Agents

// Evolving Meta-Skill for Multi-Agent Systems // Can a multi-agent system get better at orchestration without touching a single weight? Aut…

DGX agent

// Evolving Meta-Skill for Multi-Agent Systems // Can a multi-agent system get better at orchestration without touching a single weight? Automatic MAS generation has been stuck between two bad options

agentsdair-ai--x
20 Jun 2026
Tutorials

GLM-5.2 is great at design (Opus level IMO). I am also starting to see great results with long-running tasks, too. How is this possible? I t…

DGX agent

GLM-5.2 is great at design (Opus level IMO). I am also starting to see great results with long-running tasks, too. How is this possible? I think there are a few clever hacks. But I just came across th

tutorialsdair-ai--x
20 Jun 2026
Safety

// Self-play with a pinch of human data // Really cool paper combining human demonstrations and self-play RL. 30 minutes of human data, 2500…

DGX agent

// Self-play with a pinch of human data // Really cool paper combining human demonstrations and self-play RL. 30 minutes of human data, 2500x less than imitation learning, is enough to make self-play

safetydair-ai--x
20 Jun 2026
Tutorials

Working on hands-on material for this. Any requests or topics you would like me to cover?

DGX agent

This appears to be a call for community input from DAIR.AI's Omar Saro regarding hands-on educational material development, likely soliciting topic requests and feedback from followers on what technic

tutorialsdair-ai--x
20 Jun 2026
Applications

So the message I am getting is that I can't use Fable to further accelerate AI research and education. No company will decide that for me. J…

DGX agent

So the message I am getting is that I can't use Fable to further accelerate AI research and education. No company will decide that for me. Just an absolutely sad day for the research community. As a d

applicationsdair-ai--x
10 Jun 2026
Research

This is awesome! I am spending a lot of time on diffusion LLMs these days, so this is perfect timing. I feel like there are so many underexp…

DGX agent

This is awesome! I am spending a lot of time on diffusion LLMs these days, so this is perfect timing. I feel like there are so many underexplored research questions around text diffusion. Weight avail

researchdair-ai--x
10 Jun 2026
Agents

This is just awesomeness from @cohere, @nickfrosst, and team. I so badly want a coding agent that just runs on my local machine. We are not …

DGX agent

This is just awesomeness from @cohere, @nickfrosst, and team. I so badly want a coding agent that just runs on my local machine. We are not too far now! Excited to get this to work with my @dair_ai co

agentsdair-ai--x
10 Jun 2026
Research

This is why frontier open models are crucial. This is extremely sad for the research community.

DGX agent

This is why frontier open models are crucial. This is extremely sad for the research community. mythos will be bad ON PURPOSE on ai 'frontier llm research' tasks, this is very very sad for the researc

researchdair-ai--x
10 Jun 2026
Model Releases

Also, I found that Hermes Agent + Nemotron 3 Ultra is a mighty combo!

DGX agent

Also, I found that Hermes Agent + Nemotron 3 Ultra is a mighty combo! Excited to launch a new way to upskill with AI agents. This is how we are making it possible for anyone to learn to build with cod

model-releasesdair-ai--x
9 Jun 2026
Agents

Excited to launch a new way to upskill with AI agents. This is how we are making it possible for anyone to learn to build with coding agents…

DGX agent

Excited to launch a new way to upskill with AI agents. This is how we are making it possible for anyone to learn to build with coding agents. To start, we are launching 4 new hands-on labs on the foll

agentsdair-ai--x
9 Jun 2026
Model Releases

NEW: Anthropic introduces Claude Fable 5, a Mythos-class model for general use. Beginning of a new class of frontier models.

DGX agent

NEW: Anthropic introduces Claude Fable 5, a Mythos-class model for general use. Beginning of a new class of frontier models. Introducing Claude Fable 5: a Mythos-class model that we’ve made safe for g

model-releasesdair-ai--x
9 Jun 2026
Agents

// Self-Harness: Harnesses That Improve Themselves // (bookmark this one) Most of the agent scaffolds we rely on today are built once and re…

DGX agent

// Self-Harness: Harnesses That Improve Themselves // (bookmark this one) Most of the agent scaffolds we rely on today are built once and remain frozen or mostly unchanged. The harness, like the skill

agentsdair-ai--x
9 Jun 2026
Agents

// The Consistency Illusion // Multi-agent debate can make agents agree on the final answer while their underlying reasoning stays misaligne…

DGX agent

// The Consistency Illusion // Multi-agent debate can make agents agree on the final answer while their underlying reasoning stays misaligned. This work finds that consensus on the output hides disagr

agentsdair-ai--x
9 Jun 2026
Model Releases

Great tips. In practice, this is how it roughly looks to run agents autonomously for hours or days. /goal or /loop to keep it going. Verific…

DGX agent

Great tips. In practice, this is how it roughly looks to run agents autonomously for hours or days. /goal or /loop to keep it going. Verification is crucial here. Seeing a number of benchmarks showing

model-releasesdair-ai--x
8 Jun 2026
Agents

New paper on how AI agents are reshaping knowledge work. This is a nice economic read on where agents actually change knowledge work to meet…

DGX agent

New paper on how AI agents are reshaping knowledge work. This is a nice economic read on where agents actually change knowledge work to meet that gap directly. (bookmark it) It studies agent adoption

agentsdair-ai--x
8 Jun 2026
Research

The point is that you should start implementing ways to encode instructions/prompts with clear goals inside automations. Nothing new but new…

DGX agent

The point is that you should start implementing ways to encode instructions/prompts with clear goals inside automations. Nothing new but newer LLMs are being trained to perform for longer duration uni

researchdair-ai--x
8 Jun 2026
Agents

Great paper on self-improving agents:

DGX agent

Great paper on self-improving agents: This was one of the standout AI papers of the week. (bookmark it) It tackles a question most self-improving AI agents ignore: is the agent actually discovering an

agentsdair-ai--x
7 Jun 2026
Tutorials

Super-powerful AI models will launch in the coming weeks. We are looking at a potential step change in model capabilities. The biggest mista…

DGX agent

Super-powerful AI models will launch in the coming weeks. We are looking at a potential step change in model capabilities. The biggest mistake right now is to lock into one vendor. I say this not only

tutorialsdair-ai--x
7 Jun 2026
Agents

The Top AI Papers of the Week (May 31 - June 7) - LEAP - AutoLab - Learn From Your Own Latents - Reusable Context Engineering - Self-Revisin…

DGX agent

The Top AI Papers of the Week (May 31 - June 7) - LEAP - AutoLab - Learn From Your Own Latents - Reusable Context Engineering - Self-Revising Discovery Systems - Scaling Laws for Agent Harnesses - Dis

agentsdair-ai--x
7 Jun 2026
Agents

This was one of the standout AI papers of the week. (bookmark it) It tackles a question most self-improving AI agents ignore: is the agent a…

DGX agent

This was one of the standout AI papers of the week. (bookmark it) It tackles a question most self-improving AI agents ignore: is the agent actually discovering anything, or just remixing what it alrea

agentsdair-ai--x
7 Jun 2026
Tutorials

// Continual Learning Bench // One of the research areas with lots of investments is continual learning. While there are many efforts, there…

DGX agent

// Continual Learning Bench // One of the research areas with lots of investments is continual learning. While there are many efforts, there is very little progress in measuring it. So the big questio

tutorialsdair-ai--x
6 Jun 2026
Local Ai

New research from Renmin University. Treat skill selection as a harness in its own right. If you design skill routing for personal or edge a…

DGX agent

New research from Renmin University. Treat skill selection as a harness in its own right. If you design skill routing for personal or edge agents, this work argues that the selection layer is a first-

local-aidair-ai--x
6 Jun 2026
Model Releases

// Agents' Last Exam // Agents' Last Exam is a living benchmark of over 1,000 economically valuable tasks, built with 250+ industry experts …

DGX agent

// Agents' Last Exam // Agents' Last Exam is a living benchmark of over 1,000 economically valuable tasks, built with 250+ industry experts and mapped to the U.S. federal occupational taxonomy. The ha

model-releasesdair-ai--x
5 Jun 2026
Research

Find an important unsolved problem you care about. Then use AI to solve it. Go deep! Talk to people. Build a community. It might take you mo…

DGX agent

Find an important unsolved problem you care about. Then use AI to solve it. Go deep! Talk to people. Build a community. It might take you months or years, but always know that AI capabilities will onl

researchdair-ai--x
5 Jun 2026
← Previous
123456…8
Next →