AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlog
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “dair-ai--x”

GridTimelineEvolution
359 results
Tutorials

// In-context Retrieval at Million-token Scale // Great study providing better understanding of retrieval at million-token scale. They run t…

DGX agent

// In-context Retrieval at Million-token Scale // Great study providing better understanding of retrieval at million-token scale. They run the first systematic study of in-context retrieval at the sca

tutorialsdair-ai--x
6 Jul 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Must-read research by Anthropic. Here is the simple explanation and why this is a big deal. We suspect LLMs perform 'internal reasoning'. Bu…

DGX agent

Must-read research by Anthropic. Here is the simple explanation and why this is a big deal. We suspect LLMs perform 'internal reasoning'. But little is known or do good methods exist to understand it.

model-releasesdair-ai--x
6 Jul 2026
Tutorials

// ReContext // Models now support 128K context windows and still fail to use evidence that is already in the prompt. Where is the gap? New …

DGX agent

// ReContext // Models now support 128K context windows and still fail to use evidence that is already in the prompt. Where is the gap? New paper introduces ReContext, a training-free inference harnes

tutorialsdair-ai--x
6 Jul 2026
Agents

// What MCP, A2A, and ACP cannot express // MCP and A2A solve capability discovery and message passing, then stop right where enterprise dep…

DGX agent

// What MCP, A2A, and ACP cannot express // MCP and A2A solve capability discovery and message passing, then stop right where enterprise deployment begins. New research runs a systematic gap analysis

agentsdair-ai--x
6 Jul 2026
Research

Fable 5 is an absolute beast at threejs. It's the best LLM at generating 3D simulations/worlds. Watch how I combine it with gpt-realtime-2 t…

DGX agent

Fable 5 is an absolute beast at threejs. It's the best LLM at generating 3D simulations/worlds. Watch how I combine it with gpt-realtime-2 to generate interactive educational 3D worlds @dair_ai. Comme

researchdair-ai--x
5 Jul 2026
Tutorials

Great overview of always-on agents. (bookmark it) It's a new 130+ pages survey on always-on agents. Simply put it, always-on agents are syst…

DGX agent

Great overview of always-on agents. (bookmark it) It's a new 130+ pages survey on always-on agents. Simply put it, always-on agents are systems whose future behavior depends on durable state built up

tutorialsdair-ai--x
5 Jul 2026
Agents

// HASTE: tiered skills for ML engineering agents // Why do ML engineering agents keep rediscovering the same techniques on every new task? …

DGX agent

// HASTE: tiered skills for ML engineering agents // Why do ML engineering agents keep rediscovering the same techniques on every new task? Because each competition is a cold start. New research intro

agentsdair-ai--x
5 Jul 2026
Agents

The Top AI Papers of the Week (June 28 - July 5): - RLMF - AutoMem - Paper Assistant Tool - MCP Server Patterns - The Verification Horizon -…

DGX agent

The Top AI Papers of the Week (June 28 - July 5): - RLMF - AutoMem - Paper Assistant Tool - MCP Server Patterns - The Verification Horizon - Red Queen Gödel Machine - Generative Skill Composition Read

agentsdair-ai--x
5 Jul 2026
Tutorials

Learn why multimodal prompting is a big deal when working with coding agents.

DGX agent

Multimodal prompting enhances coding agents by enabling them to process and integrate multiple types of input data—such as text, images, and code snippets—simultaneously, improving their ability to un

tutorialsdair-ai--x
4 Jul 2026
Tutorials

Multimodal prompting is clearly the future. How we interact with agents is evolving. I share a bit (including a video walkthrough) of how I …

DGX agent

This post discusses the evolution of multimodal prompting in AI agent interactions, highlighting how users can now communicate with AI systems using multiple input types beyond text. The author provid

tutorialsdair-ai--x
4 Jul 2026
Agents

Can coding-agents replicate scientific ML papers? We know this is possible because we can already do this @dair_ai. Still a great read. So t…

DGX agent

Can coding-agents replicate scientific ML papers? We know this is possible because we can already do this @dair_ai. Still a great read. So they try to replicate an ML paper from its materials alone. T

agentsdair-ai--x
3 Jul 2026
Model Releases

Highly-recommended read from MIT on the part of RL with verifiable rewards that everyone keeps hitting. RLVR only optimizes what you can obj…

DGX agent

Highly-recommended read from MIT on the part of RL with verifiable rewards that everyone keeps hitting. RLVR only optimizes what you can objectively score, so style, structure, and diversity quietly c

model-releasesdair-ai--x
3 Jul 2026
Agents

Multimodal prompting is clearly the future. I love experimenting with new ways to interact with agents. As a researcher and engineer, I've f…

DGX agent

Multimodal prompting is clearly the future. I love experimenting with new ways to interact with agents. As a researcher and engineer, I've found that the richer the inputs to the agent and the richer

agentsdair-ai--x
3 Jul 2026
Tutorials

NEW paper worth reading. (bookmark it) The basic idea is to pair a compressive recurrent state with a small exact memory, which helps to rec…

DGX agent

NEW paper worth reading. (bookmark it) The basic idea is to pair a compressive recurrent state with a small exact memory, which helps to recover long-range recall without giving up the efficiency of l

tutorialsdair-ai--x
3 Jul 2026
Research

AI sovereignty isn’t optional. Don’t give away your alpha so easily. Protect it as much as you can. Open source models are critical and shou…

DGX agent

AI sovereignty isn’t optional. Don’t give away your alpha so easily. Protect it as much as you can. Open source models are critical and should be an important part of any individual’s, organisation’s,

researchdair-ai--x
2 Jul 2026
Model Releases

Another fascinating paper on LLM Judges. (bookmark it) It's from Amazon, and they show that if you run panels of LLM judges, averaging their…

DGX agent

Another fascinating paper on LLM Judges. (bookmark it) It's from Amazon, and they show that if you run panels of LLM judges, averaging their scores is a trap. 'Overall, we establish that robust aggreg

model-releasesdair-ai--x
2 Jul 2026
Model Releases

// AutoMem // I quite like this idea of metamemory. (bookmark it) This new research from Stanford treats agent's memory management as a trai…

DGX agent

// AutoMem // I quite like this idea of metamemory. (bookmark it) This new research from Stanford treats agent's memory management as a trainable skill instead of a fixed module. The model decides wha

model-releasesdair-ai--x
2 Jul 2026
Model Releases

LLM Wikis are being slept on. I argue that creating knowledge bases with LLMs or coding agents is one of the most valuable applications of A…

DGX agent

LLM Wikis are being slept on. I argue that creating knowledge bases with LLMs or coding agents is one of the most valuable applications of AI today. It's about being intentional in building and scalin

model-releasesdair-ai--x
2 Jul 2026
Safety

NEW paper from NVIDIA. They discuss robot programming that compounds experience instead of throwing it away. Traditional robot programming f…

DGX agent

NEW paper from NVIDIA. They discuss robot programming that compounds experience instead of throwing it away. Traditional robot programming forces you to orchestrate perception, contact dynamics, diver

safetydair-ai--x
2 Jul 2026
Tutorials

New research from Google. LLMs hallucinate with high confidence, miss their own knowledge boundaries, and misreport uncertainty. Most fixes …

DGX agent

New research from Google. LLMs hallucinate with high confidence, miss their own knowledge boundaries, and misreport uncertainty. Most fixes bolt calibration on from the outside. RLMF turns the model o

tutorialsdair-ai--x
2 Jul 2026
Model Releases

Great paper on managing agent skills. Skill libraries keep growing, and picking the right skills has become a bottleneck for coding agents. …

DGX agent

Great paper on managing agent skills. Skill libraries keep growing, and picking the right skills has become a bottleneck for coding agents. The defaults are to expose the agent to the whole skill coll

model-releasesdair-ai--x
1 Jul 2026
Model Releases

My prediction: the excitement for Fable 5 will wear off really fast. Reposting this to help those who will be extremely disappointed after t…

DGX agent

My prediction: the excitement for Fable 5 will wear off really fast. Reposting this to help those who will be extremely disappointed after they play with Fable 5 and run out of tokens or can't do much

model-releasesdair-ai--x
1 Jul 2026
Agents

NEW paper worth reading. (bookmark it) Autonomous research systems usually prove themselves on cherry-picked wins, human-framed topics, or a…

DGX agent

NEW paper worth reading. (bookmark it) Autonomous research systems usually prove themselves on cherry-picked wins, human-framed topics, or a handful of preset tasks. FARS runs the full loop at scale i

agentsdair-ai--x
1 Jul 2026
Model Releases

Really confused by all the excitement I see in my timeline for a nerfed model. Never seen anything like it. So many will end up very disappo…

DGX agent

Really confused by all the excitement I see in my timeline for a nerfed model. Never seen anything like it. So many will end up very disappointed. Time to rethink how to build around frontier and open

model-releasesdair-ai--x
1 Jul 2026
Research

Who did it best? GLM-5.2 (left) | Fugu Ultra (middle) | Fable 5 (right) Same one-shot prompt. The last one is my favorite!

DGX agent

This post compares the outputs of three AI models—GLM-5.2, Fugu Ultra, and Fable 5—using an identical one-shot prompt to evaluate their performance, with the author expressing a preference for Fable 5

researchdair-ai--x
1 Jul 2026
Model Releases

Cool new paper from NVIDIA. Looks like agentic coding is moving into hardware design. HORIZON treats hardware design as repository-level cod…

DGX agent

Cool new paper from NVIDIA. Looks like agentic coding is moving into hardware design. HORIZON treats hardware design as repository-level code evolution. A Markdown harness becomes a project pack with

model-releasesdair-ai--x
30 Jun 2026
Agents

If you build with MCPs, this one is worth reading. (bookmark it) The paper covers five recurring MCP server patterns across fifteen independ…

DGX agent

If you build with MCPs, this one is worth reading. (bookmark it) The paper covers five recurring MCP server patterns across fifteen independently developed servers. That taxonomy is useful because I s

agentsdair-ai--x
30 Jun 2026
Model Releases

Love how Google continues to drive down the cost of building with their models. <4s image and $0.034 / 1K image. Wow! We have a bunch of stu…

DGX agent

Love how Google continues to drive down the cost of building with their models. <4s image and 0.034 / 1K image. Wow! We have a bunch of stuff (education & research) we're building @dair_ai using Nano

model-releasesdair-ai--x
30 Jun 2026
Model Releases

// Neural procedural memory // Good paper on agent memory beyond prompt retrieval. NPM stores procedural skills as activation steering vecto…

DGX agent

// Neural procedural memory // Good paper on agent memory beyond prompt retrieval. NPM stores procedural skills as activation steering vectors distilled from contrastive historical experience. Textual

model-releasesdair-ai--x
30 Jun 2026
Model Releases

Qwen publishes new work on RL coding agents. (bookmark it) The idea is to continually build a verification system that co-evolves with AI ag…

DGX agent

Qwen publishes new work on RL coding agents. (bookmark it) The idea is to continually build a verification system that co-evolves with AI agents. LLMs suffer from all sorts of reward hacking issues. T

model-releasesdair-ai--x
30 Jun 2026
Applications

Recommended reading if you are looking to scale with open models in production.

DGX agent

This post likely recommends resources or reading materials for practitioners looking to deploy and scale open-source language models in production environments. DAIR.AI, a research organization focuse

applicationsdair-ai--x
30 Jun 2026
Tutorials

Recommended reading if you are scaling with open models. BTW, you should be thinking about how to scale with open-weight models.

DGX agent

This post from DAIR.AI recommends resources for developers and organizations scaling applications using open-weight language models, emphasizing the importance of planning infrastructure and deploymen

tutorialsdair-ai--x
30 Jun 2026
Model Releases

Sonnet 5 is here! This is going to support better long-running agents. Previous Sonnet models were unreliable, so it's great to see the impr…

DGX agent

Sonnet 5 is here! This is going to support better long-running agents. Previous Sonnet models were unreliable, so it's great to see the improved version that can complete agentic tasks more reliably.

model-releasesdair-ai--x
30 Jun 2026
Agents

The gap in autonomous agentic loops that gets ignored: agents can plan and call APIs but can't acquire tools they don't have access to. x402…

DGX agent

The gap in autonomous agentic loops that gets ignored: agents can plan and call APIs but can't acquire tools they don't have access to. x402 + Apify's 20,000+ Actors is a concrete fix for that. Worth

agentsdair-ai--x
30 Jun 2026
Research

Finally, a people-search tool that actually works. Most people-search tools sell you a frozen list. @CLODOAI searches the live web, reads th…

DGX agent

Finally, a people-search tool that actually works. Most people-search tools sell you a frozen list. @CLODOAI searches the live web, reads the signals, and tells you why this specific person, right now

researchdair-ai--x
29 Jun 2026
Research

Highly recommended reading. What an impressive use of LLMs and deep learning. Achieves 'real-time sentence decoding from non-invasive brain …

DGX agent

Highly recommended reading. What an impressive use of LLMs and deep learning. Achieves 'real-time sentence decoding from non-invasive brain recordings, approaching levels of accuracy previously exclus

researchdair-ai--x
29 Jun 2026
Tutorials

LLM-as-a-Judge explained in ~10 mins. Knowing how to build AI verifiers and judges is one of the most important emerging AI skills today. He…

DGX agent

LLM-as-a-Judge explained in ~10 mins. Knowing how to build AI verifiers and judges is one of the most important emerging AI skills today. Here is a quick intro on the topic and where to learn how to a

tutorialsdair-ai--x
29 Jun 2026
Agents

NEW paper from Google (bookmark it) It's on advancing automated scientific review. Just pay attention to the focus on agentic verification w…

DGX agent

NEW paper from Google (bookmark it) It's on advancing automated scientific review. Just pay attention to the focus on agentic verification which is something I've been writing about recently. AI is ac

agentsdair-ai--x
29 Jun 2026
Model Releases

This is smart from Cline. They just launched ClinePass, which makes it easy to access the latest open-weight models like GLM 5.2, Kimi k2.7-…

DGX agent

This is smart from Cline. They just launched ClinePass, which makes it easy to access the latest open-weight models like GLM 5.2, Kimi k2.7-code, Mimo 2.5, Deepseek v4 pro, Minimax M3, and more. Alway

model-releasesdair-ai--x
29 Jun 2026
Agents

Fascinating paper on self-improving agents. (bookmark it) If you are working on agentic loops, you will quickly realize that they are only a…

DGX agent

Fascinating paper on self-improving agents. (bookmark it) If you are working on agentic loops, you will quickly realize that they are only as good as the effectiveness of the evaluator. Self-improveme

agentsdair-ai--x
28 Jun 2026
Tutorials

NEW paper worth reading. Reasoning-data curation is expensive because scoring a trace usually means reading it to the end. This new work fro…

DGX agent

NEW paper worth reading. Reasoning-data curation is expensive because scoring a trace usually means reading it to the end. This new work from UCLA shows you may not have to. The quality of a reasoning

tutorialsdair-ai--x
28 Jun 2026
Agents

The Top AI Papers of the Week (June 21 - June 28) - Autodata - Sakana Fugu - Agent-as-a-Router - Agent-Native Memory - A Pinch of Human Data…

DGX agent

The Top AI Papers of the Week (June 21 - June 28) - Autodata - Sakana Fugu - Agent-as-a-Router - Agent-Native Memory - A Pinch of Human Data - Critique of the Agent Model - Agent Communication Protoco

agentsdair-ai--x
28 Jun 2026
Model Releases

Why do RL runs on LLMs blow up even when the recipe looks right? GEOALIGN, from the Alibaba team behind Qwen, points at the rollouts. A hand…

DGX agent

Why do RL runs on LLMs blow up even when the recipe looks right? GEOALIGN, from the Alibaba team behind Qwen, points at the rollouts. A handful of bad batches push the policy in incoherent directions,

model-releasesdair-ai--x
28 Jun 2026
Tutorials

If you use LLM-as-judge, this one is worth reading. (bookmark it) It's actually one of the most effective ways to use LLM-as-a-Judge for eva…

DGX agent

If you use LLM-as-judge, this one is worth reading. (bookmark it) It's actually one of the most effective ways to use LLM-as-a-Judge for evals. Holistic judge scores hide both their reasoning and thei

tutorialsdair-ai--x
27 Jun 2026
Research

Loop engineering is just prompt engineering with great system design.

DGX agent

Loop engineering refers to an iterative approach to AI system design that combines prompt engineering with robust system architecture, emphasizing that effective AI applications depend equally on how

researchdair-ai--x
27 Jun 2026
Hardware

NEW paper from NVIDIA. (bookmark it) Speed-of-light performance analysis tells you the theoretical floor of a workload, but teams still deri…

DGX agent

NEW paper from NVIDIA. (bookmark it) Speed-of-light performance analysis tells you the theoretical floor of a workload, but teams still derive it by hand and freeze it. SOLAR automates the whole thing

hardwaredair-ai--x
27 Jun 2026
Safety

When does combining LLMs help? Great analysis on combining language models, measured across 67 models from 21 providers. Any policy that rou…

DGX agent

When does combining LLMs help? Great analysis on combining language models, measured across 67 models from 21 providers. Any policy that routes, votes, cascades, or runs a mixture of agents and then r

safetydair-ai--x
27 Jun 2026
Model Releases

Dynamic workflows (generating harnesses on the fly) are a new form of test-time compute. But LLMs aren't great at building them. I often hav…

DGX agent

Dynamic workflows (generating harnesses on the fly) are a new form of test-time compute. But LLMs aren't great at building them. I often have to steer agents to generate complex patterns. Curious how

model-releasesdair-ai--x
26 Jun 2026
← Previous
12345…8
Next →