AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “papers”

GridTimelineEvolution
693 results
Model Releases

// Persistent Workspaces for Long-Lived Claude Code Agent Teams // Four issues to be aware of: > Working state vanishes when a terminal clos…

DGX agent

// Persistent Workspaces for Long-Lived Claude Code Agent Teams // Four issues to be aware of: > Working state vanishes when a terminal closes and the team cannot be resumed. > Compaction condenses th

model-releasesdair-ai--x
1 Aug 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

The “AGI-is-near” community keeps committing the same logical fallacy over and over; I have seen it at least half a dozen times today alone.…

DGX agent

The “AGI-is-near” community keeps committing the same logical fallacy over and over; I have seen it at least half a dozen times today alone. Every time there’s an advance, I see the same error. Here’s

safetygary-marcus--x
1 Aug 2026
Model Releases

Neat work on long-horizon agents. Splitting a hard task across agents is typically how standard multi-agent work. The usual design lets them…

DGX agent

Neat work on long-horizon agents. Splitting a hard task across agents is typically how standard multi-agent work. The usual design lets them exchange findings only at phase boundaries, through staged

model-releasesdair-ai--x
31 Jul 2026
Model Releases

New research from Microsoft. This one is on training computer-use agents at scale. Recent pipelines generate synthetic environments in bulk,…

DGX agent

New research from Microsoft. This one is on training computer-use agents at scale. Recent pipelines generate synthetic environments in bulk, which moved the bottleneck from how many exist to what is i

model-releasesdair-ai--x
31 Jul 2026
Applications

As the benchmarks that test frontier AI on get more complex, we are losing one of the most important aspects of benchmarking: comparisons to…

DGX agent

As the benchmarks that test frontier AI on get more complex, we are losing one of the most important aspects of benchmarking: comparisons to humans Validated benchmarks need to have human (ideally mul

applicationsethan-mollick--x
30 Jul 2026
Agents

We have to be careful to not offload our understanding to agents. I think there is also a good opportunity to build agentic applications tha…

DGX agent

We have to be careful to not offload our understanding to agents. I think there is also a good opportunity to build agentic applications that encourage deeper understanding. For example, coding agents

agentsdair-ai--x
30 Jul 2026
Agents

// Unfolding Sub-Agents for Long-Horizon ML Engineering // Watch a single agent work a machine learning engineering task for six hours you s…

DGX agent

// Unfolding Sub-Agents for Long-Horizon ML Engineering // Watch a single agent work a machine learning engineering task for six hours you see issues like context fills with stack traces, dead experim

agentsdair-ai--x
29 Jul 2026
Agents

Across eight case studies spanning industry and academia, we explore what this shift means for scientific computing—and why human verificati…

DGX agent

Across eight case studies spanning industry and academia, we explore what this shift means for scientific computing—and why human verification, stewardship, and long-term maintenance matter. https://o

agentsopenai--x
28 Jul 2026
Applications

Everything you need to know about Kimi K3 with @Kimi_Moonshot and @togethercompute team. If you're looking into evaling and building with Ki…

DGX agent

Everything you need to know about Kimi K3 with @Kimi_Moonshot and @togethercompute team. If you're looking into evaling and building with Kimi K3 I would not miss this one! Kimi K3 has everyone’s atte

applicationstogether-ai--x
28 Jul 2026
Research

Recent discussions about open-sourcing make me feel that I should go back and revisit these important open-source works in representation le…

DGX agent

Recent discussions about open-sourcing make me feel that I should go back and revisit these important open-source works in representation learning that pushed the field forward and eventually made vis

researchyann-lecun--x
26 Jul 2026
Agents

Diffusion LLMs can now handle real agentic work. LLaDA 2.2 is the first large-scale diffusion LLM built to operate as a real agent, planning…

DGX agent

Diffusion LLMs can now handle real agentic work. LLaDA 2.2 is the first large-scale diffusion LLM built to operate as a real agent, planning, calling tools, and self-correcting across long multi-turn

agentsdair-ai--x
25 Jul 2026
Model Releases

Ha! It did it: 'We introduce BenchBenchBenchBenchBench (BBBBB), an executable benchmark of AI-authored conformance suites for benchmark-eval…

DGX agent

Ha! It did it: 'We introduce BenchBenchBenchBenchBench (BBBBB), an executable benchmark of AI-authored conformance suites for benchmark-evaluation metrics' I really thought it would treat 'now do benc

model-releasesethan-mollick--x
25 Jul 2026
Model Releases

As a joke I prompted Codex 'Build and run BenchBench, a benchmark of now good ai is at creating benchmarks. then figure out what benchbenchb…

DGX agent

As a joke I prompted Codex 'Build and run BenchBench, a benchmark of now good ai is at creating benchmarks. then figure out what benchbenchbench is and run that. and then write benchbenchbench up as a

model-releasesethan-mollick--x
24 Jul 2026
Safety

Wish Dario had taken this bet.

DGX agent

Wish Dario had taken this bet. 🎺 I am hereby publicly offering to bet @darioamodei $1,000,000 that AI in 2027 will NOT be “smarter than Nobel Prize winners across most fields in science and engineerin

safetygary-marcus--x
23 Jul 2026
Hardware

Decoupling embedding from ingestion means whatever hardware is on hand can do the work. One script, runtime device check: MPS on Apple Silic…

DGX agent

Decoupling embedding from ingestion means whatever hardware is on hand can do the work. One script, runtime device check: MPS on Apple Silicon, CUDA on NVIDIA, CPU fallback otherwise. 10K-20K records

hardwarepinecone--x
22 Jul 2026
Model Releases

🎨 Meet Qwen-Image-3.0 — the third generation of our foundational image generation model. If 1.0 was about 'Precision,' and 2.0 added 'Varie…

DGX agent

🎨 Meet Qwen-Image-3.0 — the third generation of our foundational image generation model. If 1.0 was about 'Precision,' and 2.0 added 'Variety, Completeness, Beauty & Authenticity,' then 3.0 comes down

model-releasesqwen--x
22 Jul 2026
Agents

Highly recommended. I've often claimed there's huge alpha in building agent harnesses. Turns out harnesses are compositional generalizers. T…

DGX agent

Highly recommended. I've often claimed there's huge alpha in building agent harnesses. Turns out harnesses are compositional generalizers. The RLM harness is an instance of this. This could lead to in

agentsdair-ai--x
20 Jul 2026
Industry

Today, we are introducing Inkling. Inkling reasons efficiently across text, image, and audio modalities. We are making the full weights avai…

DGX agent

Today, we are introducing Inkling. Inkling reasons efficiently across text, image, and audio modalities. We are making the full weights available. https://thinkingmachines.ai/news/introducing-inkling/

industryclem-delangue--x
15 Jul 2026
Model Releases

Fable, turn my tweet into a thinkpiece (this was pretty funny): There has never been a better time to have opinions about artificial intelli…

DGX agent

Fable, turn my tweet into a thinkpiece (this was pretty funny): There has never been a better time to have opinions about artificial intelligence. I say this with some authority, because I am currentl

model-releasesethan-mollick--x
14 Jul 2026
Tutorials

Everyone keeps asking me how to build a second brain or an LLM wiki. Here is the easiest setup I have found. I took my Wiki Builder skill, i…

DGX agent

Everyone keeps asking me how to build a second brain or an LLM wiki. Here is the easiest setup I have found. I took my Wiki Builder skill, installed it into HyperAgent (@hyperagentapp) as a reusable s

tutorialsdair-ai--x
13 Jul 2026
Model Releases

Great writeup from the University of Oxford. It's a taxonomy of LLM-based agent limitations. Good read for anyone shipping with agents. Benc…

DGX agent

Great writeup from the University of Oxford. It's a taxonomy of LLM-based agent limitations. Good read for anyone shipping with agents. Benchmark scores keep climbing, yet the same agent failures resu

model-releasesdair-ai--x
8 Jul 2026
Model Releases

Qwen publishes new work on RL coding agents. (bookmark it) The idea is to continually build a verification system that co-evolves with AI ag…

DGX agent

Qwen publishes new work on RL coding agents. (bookmark it) The idea is to continually build a verification system that co-evolves with AI agents. LLMs suffer from all sorts of reward hacking issues. T

model-releasesdair-ai--x
30 Jun 2026
Safety

When does combining LLMs help? Great analysis on combining language models, measured across 67 models from 21 providers. Any policy that rou…

DGX agent

When does combining LLMs help? Great analysis on combining language models, measured across 67 models from 21 providers. Any policy that routes, votes, cascades, or runs a mixture of agents and then r

safetydair-ai--x
27 Jun 2026
Agents

// The Consistency Illusion // Multi-agent debate can make agents agree on the final answer while their underlying reasoning stays misaligne…

DGX agent

// The Consistency Illusion // Multi-agent debate can make agents agree on the final answer while their underlying reasoning stays misaligned. This work finds that consensus on the output hides disagr

agentsdair-ai--x
9 Jun 2026
Hardware

I'll never forget @GaryMarcus' comment at last autumn's @MediaTechDem conference in MTL - something to the effect of 'what if Nvidia are jus…

DGX agent

I'll never forget @GaryMarcus' comment at last autumn's @MediaTechDem conference in MTL - something to the effect of 'what if Nvidia are just the people who sold shovels in the gold rush'? Striking pa

hardwaregary-marcus--x
8 Jun 2026
Agents

New MIT study. Code volume surges by 300%, but output increases by only 30%: The AI dividend meets an awkward reality Autonomous AI coding a…

DGX agent

New MIT study. Code volume surges by 300%, but output increases by only 30%: The AI dividend meets an awkward reality Autonomous AI coding agents raised commits by 180%, but releases rose only 30%. Th

agentsgary-marcus--x
7 Jun 2026
Tutorials

Nice primer on post-training reasoning data. (bookmark it) This is one of the first primers to pull the scattered post-training reasoning-da…

DGX agent

Nice primer on post-training reasoning data. (bookmark it) This is one of the first primers to pull the scattered post-training reasoning-data literature into one place, synthesizing over 150 public s

tutorialsdair-ai--x
3 Jun 2026
Model Releases

Highlighting the new WebGPU backend in llama.cpp/ggml The work to bring full-fledged WebGPU support in llama.cpp started about an year and a…

DGX agent

Highlighting the new WebGPU backend in llama.cpp/ggml The work to bring full-fledged WebGPU support in llama.cpp started about an year and a half ago. It has been lead by @reeselevine and team at USCS

model-releasesgeorgi-gerganov--x
22 May 2026
Industry

Can we transform the Hugging Face Hub🤗—with its enormous sea of artifacts—into a self-evolving discovery machine? WE CAN. Introducing Artif…

DGX agent

Can we transform the Hugging Face Hub🤗—with its enormous sea of artifacts—into a self-evolving discovery machine? WE CAN. Introducing ArtifactLinker! 🚀🔥 We are fundamentally rethinking how we interact

industryclem-delangue--x
20 May 2026
Safety

The pure LLM debate - which I had for many years, here and elsewhere - is indeed no longer relevant. Why? Because I won; nobody uses pure LL…

DGX agent

The pure LLM debate - which I had for many years, here and elsewhere - is indeed no longer relevant. Why? Because I won; nobody uses pure LLMs anymore. Nowadays all deployed objects are neurosymbolic,

safetygary-marcus--x
18 May 2026
Safety

Big move from arXiv: a one-year ban for authors who submit AI-generated content without proper checking. This is not about banning AI from a…

DGX agent

Big move from arXiv: a one-year ban for authors who submit AI-generated content without proper checking. This is not about banning AI from academia. AI can be extremely useful. It can help us write be

safetygary-marcus--x
15 May 2026
Safety

Pay attention to this one if you build research or knowledge-work agents. Most research-agent systems produce uniform outputs regardless of …

DGX agent

Pay attention to this one if you build research or knowledge-work agents. Most research-agent systems produce uniform outputs regardless of who is driving them. This new work, NanoResearch, argues tha

safetydair-ai--x
12 May 2026
Safety

Am old enough to remember when @GeoffreyHinton told me I was stupid for saying that LLMs regurgitate training data. He was wrong. LLM regurg…

DGX agent

Am old enough to remember when @GeoffreyHinton told me I was stupid for saying that LLMs regurgitate training data. He was wrong. LLM regurgitation is now one of the best-established findings in the f

safetygary-marcus--x
11 May 2026
Tutorials

// LLMs Improving LLMs // Interesting progress the past of couple of weeks around self-improving AI agents. If autoresearch was interesting,…

DGX agent

// LLMs Improving LLMs // Interesting progress the past of couple of weeks around self-improving AI agents. If autoresearch was interesting, you will like this read. (bookmark it) We've been hand-tuni

tutorialsdair-ai--x
11 May 2026
Agents

This seems like a critical reason to open up about AI use in academia. Scholars are using old AI models, badly, and not talking about it. Ne…

DGX agent

This seems like a critical reason to open up about AI use in academia. Scholars are using old AI models, badly, and not talking about it. New models hallucinate very few citations, and good agentic ha

agentsethan-mollick--x
11 May 2026
Safety

The human brain🧠 is incredibly efficient because it only activates the specific neurons needed for a thought. Modern LLMs naturally try to …

DGX agent

The human brain🧠 is incredibly efficient because it only activates the specific neurons needed for a thought. Modern LLMs naturally try to do this too (> 95% of neurons in feedforward layers stay sile

safetydavid-ha--x
8 May 2026
Model Releases

// HeavySkill // One of the cleaner takes on agentic harness design I've read. They argue that what actually drives agent harness performanc…

DGX agent

// HeavySkill // One of the cleaner takes on agentic harness design I've read. They argue that what actually drives agent harness performance is not the orchestration code. It's a single inner skill:

model-releasesdair-ai--x
5 May 2026
Safety

Are AIs about to subjugate humanity? In the debate about catastrophic AI risk, evolutionary scenarios receive far too little attention compa…

DGX agent

Are AIs about to subjugate humanity? In the debate about catastrophic AI risk, evolutionary scenarios receive far too little attention compared to largely speculative arguments about “instrumental con

safetydan-hendrycks--x
2 May 2026
Model Releases

Claude Opus 4.7 just implemented an AlphaZero-style self-play pipeline from scratch. It did this on consumer hardware in three hours, then b…

DGX agent

Claude Opus 4.7 just implemented an AlphaZero-style self-play pipeline from scratch. It did this on consumer hardware in three hours, then beat the Pascal Pons solver 7 of 8 as first-mover on Connect

model-releasesdair-ai--x
2 May 2026
Agents

'AI should elevate your thinking, not replace it.' I don't disagree, but the issue is that current LLMs are not really trained to support th…

DGX agent

'AI should elevate your thinking, not replace it.' I don't disagree, but the issue is that current LLMs are not really trained to support that out of the box. I've solved this by building my own agent

agentsdair-ai--x
27 Apr 2026
Agents

For the past few years, humans have been doing “prompt engineering” to coax the best performance out of different LLMs. In this work, we exp…

DGX agent

For the past few years, humans have been doing “prompt engineering” to coax the best performance out of different LLMs. In this work, we explored what happens if we train an AI to do that job instead.

agentsdavid-ha--x
27 Apr 2026
Industry

Looking forward to the work coming out of @IneffableLabs 🇬🇧 Largest EU/UK raise ever 😮 Beat the previous UK largest seed raise ($101m Sta…

DGX agent

Looking forward to the work coming out of @IneffableLabs 🇬🇧 Largest EU/UK raise ever 😮 Beat the previous UK largest seed raise (101m Stability AI!) by 999m 🚀 We have seen amazing things on our self-le

industryemad-mostaque--x
27 Apr 2026
Applications

Very cool analysis of the submissions to a major management journal that shows how much the system of science, built for humans, is under st…

DGX agent

Very cool analysis of the submissions to a major management journal that shows how much the system of science, built for humans, is under strain as a result of AI. AI can be used to do better science

applicationsethan-mollick--x
27 Apr 2026
Safety

Here is a very common problem when building complex agents. Long-horizon agents (in particular) fail in two ways: the decision-maker can't d…

DGX agent

Here is a very common problem when building complex agents. Long-horizon agents (in particular) fail in two ways: the decision-maker can't decompose well, or the skill library goes stale. This new res

safetydair-ai--x
26 Apr 2026
Industry

👀 The most interesting people in the most exotic places can occasionally be persuaded to talk about NFTs, digital collectibles, and even @C…

DGX agent

👀 The most interesting people in the most exotic places can occasionally be persuaded to talk about NFTs, digital collectibles, and even @CandyDigital! Try it sometime just in case it catches on… 🔥🚀📈

industryemad-mostaque--x
26 Apr 2026
Model Releases

Alpha Eval: Agents Making Evals as a Multi-Player Game This diagram and blurb is largely a research riff with Claude on building data genera…

DGX agent

Alpha Eval: Agents Making Evals as a Multi-Player Game This diagram and blurb is largely a research riff with Claude on building data generation systems to get closer to the holy grail of self-improvi

model-releasesharrison-chase--x
25 Apr 2026
Industry

Bet this happens with Navier Stokes and it’s going to be something not even related to PDEs that solves it

DGX agent

Bet this happens with Navier Stokes and it’s going to be something not even related to PDEs that solves it 23 years old with no advanced mathematics training solves Erdős problem with ChatGPT Pro. 'Wh

industryemad-mostaque--x
25 Apr 2026
Industry

5B tokens in ml-intern in 48h 😅😅😅

DGX agent

5B tokens in ml-intern in 48h 😅😅😅 we burned through 5 billion tokens in 48h. turns out giving everyone unlimited access to the most expensive model on the planet is not a great business strategy 🙈 So

industryclem-delangue--x
23 Apr 2026
← Previous
1…7891011…15
Next →