AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “papers”

GridTimelineEvolution
693 results
CompaniesToolsTechniques

Each lane shows up to 8 recent matching entries, ordered from earlier to later. Tracks load separately to keep the 75,000+ entry wiki fast.

Companies

CompanyAnthropic8 recent entries
30 Jul 2026Can AI agents conduct open-ended AI research? Most evaluations of agents conducting AI research focus on narrow, verifiable tasks. But AI re…

Can AI agents conduct open-ended AI research? Most evaluations of agents conducting AI research focus on narrow, verifiable tasks. But AI research is often open ended. Researchers pick hypotheses, dec

→31 Jul 2026Neat work on long-horizon agents. Splitting a hard task across agents is typically how standard multi-agent work. The usual design lets them…

Neat work on long-horizon agents. Splitting a hard task across agents is typically how standard multi-agent work. The usual design lets them exchange findings only at phase boundaries, through staged

3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
→1 Aug 2026// Persistent Workspaces for Long-Lived Claude Code Agent Teams // Four issues to be aware of: > Working state vanishes when a terminal clos…

// Persistent Workspaces for Long-Lived Claude Code Agent Teams // Four issues to be aware of: > Working state vanishes when a terminal closes and the team cannot be resumed. > Compaction condenses th

→1 Aug 2026If you maintain an AGENTS.md or a CLAUDE.md, this is worth a read. (bookmark it) 288 gold-test evaluated runs across Claude Code and Codex, …

If you maintain an AGENTS.md or a CLAUDE.md, this is worth a read. (bookmark it) 288 gold-test evaluated runs across Claude Code and Codex, 17 real tasks from 3 repositories, with context-injection st

→1 Aug 20263/ here's the part that makes it non-optional: the same agent that will do whatever it takes to solve a problem will also walk straight out …

3/ here's the part that makes it non-optional: the same agent that will do whatever it takes to solve a problem will also walk straight out of a sandbox you thought was locked down. We watched exactly

→6 Aug 2026This paper by researchers from MIT and Stanford finds that most people would be financially better off if they followed the financial advice…

This paper by researchers from MIT and Stanford finds that most people would be financially better off if they followed the financial advice of LLMs (GPT-5.2 & Gemini 3 Flash) But some people get a bi

→11 Aug 2026Claude's watermark probably doesn't work how you think. As the CTO of GPTZero, I'll explain how Anthropic, Google and OpenAI are building te…

Claude's watermark probably doesn't work how you think. As the CTO of GPTZero, I'll explain how Anthropic, Google and OpenAI are building text watermarking in this brief explainer and whether it can b

→12 Aug 2026Very interesting new work from Anthropic. (bookmark it) They evolve mind viruses, ideas that spread through a multi-agent system by getting …

Very interesting new work from Anthropic. (bookmark it) They evolve mind viruses, ideas that spread through a multi-agent system by getting each host to pass them on, then measure what governs the spr

CompanyOpenAI8 recent entries
24 Jul 2026As a joke I prompted Codex 'Build and run BenchBench, a benchmark of now good ai is at creating benchmarks. then figure out what benchbenchb…

As a joke I prompted Codex 'Build and run BenchBench, a benchmark of now good ai is at creating benchmarks. then figure out what benchbenchbench is and run that. and then write benchbenchbench up as a

→25 Jul 2026Ha! It did it: 'We introduce BenchBenchBenchBenchBench (BBBBB), an executable benchmark of AI-authored conformance suites for benchmark-eval…

Ha! It did it: 'We introduce BenchBenchBenchBenchBench (BBBBB), an executable benchmark of AI-authored conformance suites for benchmark-evaluation metrics' I really thought it would treat 'now do benc

→28 Jul 2026Across eight case studies spanning industry and academia, we explore what this shift means for scientific computing—and why human verificati…

Across eight case studies spanning industry and academia, we explore what this shift means for scientific computing—and why human verification, stewardship, and long-term maintenance matter. https://o

→31 Jul 2026This is not optional, things are getting chaotic now, and not dealing with this change won't make it go away. Plus, this could be a huge boo…

This is not optional, things are getting chaotic now, and not dealing with this change won't make it go away. Plus, this could be a huge boost for both individual satisfaction & firm performance if do

→2 Aug 2026Top eight misconceptions about OpenAI’s amazing new Astra math results. 1. Expertise in one domain does not at all guarantee expertise in al…

Top eight misconceptions about OpenAI’s amazing new Astra math results. 1. Expertise in one domain does not at all guarantee expertise in all or even most domains. There is an important, principled re

→7 Aug 2026Every CIO should watch this. Persistent systems trying to break through and solve problems at all costs are more creative than you think. La…

Every CIO should watch this. Persistent systems trying to break through and solve problems at all costs are more creative than you think. Labs and enterprises will be focused more on network effects (

→11 Aug 2026🦔Nvidia announced agreements yesterday with the six biggest names in private capital, Apollo, Blackstone, BlackRock, Brookfield, Goldman Sa…

🦔Nvidia announced agreements yesterday with the six biggest names in private capital, Apollo, Blackstone, BlackRock, Brookfield, Goldman Sachs, and KKR, to raise over $500 billion so its own customers

→11 Aug 2026Claude's watermark probably doesn't work how you think. As the CTO of GPTZero, I'll explain how Anthropic, Google and OpenAI are building te…

Claude's watermark probably doesn't work how you think. As the CTO of GPTZero, I'll explain how Anthropic, Google and OpenAI are building text watermarking in this brief explainer and whether it can b

CompanyGoogle8 recent entries
12 Jul 2026Another big reason to use combination of frontier models. Chain-of-thought monitoring is treated as a reliable safety layer for agents. This…

Another big reason to use combination of frontier models. Chain-of-thought monitoring is treated as a reliable safety layer for agents. This DeepMind-affiliated study shows the layer can be argued out

→14 Jul 2026New research from Google DeepMind on effective model routing. LLM routers get judged on accuracy and cost. Both can look great while the rou…

New research from Google DeepMind on effective model routing. LLM routers get judged on accuracy and cost. Both can look great while the router is meaningless. If every model in your society responds

→21 Jul 2026very notable trajectory comparison writeup here buried in the RLM paper from @a1zhang and @lateinteraction. an open secret of 'frontier' mod…

very notable trajectory comparison writeup here buried in the RLM paper from @a1zhang and @lateinteraction. an open secret of 'frontier' model training is that even without training on test, you can b

→27 Jul 2026Great technical paper from Google. Great read on why context beats scale for agents working against unfamiliar APIs. (bookmark it) GPU kerne…

Great technical paper from Google. Great read on why context beats scale for agents working against unfamiliar APIs. (bookmark it) GPU kernel optimization has KernelBench to hillclimb on. TPUs had not

→1 Aug 2026The “AGI-is-near” community keeps committing the same logical fallacy over and over; I have seen it at least half a dozen times today alone.…

The “AGI-is-near” community keeps committing the same logical fallacy over and over; I have seen it at least half a dozen times today alone. Every time there’s an advance, I see the same error. Here’s

→2 Aug 2026New research from Google DeepMind. (bookmark it) SkillSmith treats model weights as an additional modality the LLM reads natively. The augme…

New research from Google DeepMind. (bookmark it) SkillSmith treats model weights as an additional modality the LLM reads natively. The augmented model ingests existing prefix weights alongside rich te

→6 Aug 2026This paper by researchers from MIT and Stanford finds that most people would be financially better off if they followed the financial advice…

This paper by researchers from MIT and Stanford finds that most people would be financially better off if they followed the financial advice of LLMs (GPT-5.2 & Gemini 3 Flash) But some people get a bi

→11 Aug 2026Claude's watermark probably doesn't work how you think. As the CTO of GPTZero, I'll explain how Anthropic, Google and OpenAI are building te…

Claude's watermark probably doesn't work how you think. As the CTO of GPTZero, I'll explain how Anthropic, Google and OpenAI are building text watermarking in this brief explainer and whether it can b

CompanyMeta8 recent entries
22 Jul 2026New research from Meta. (bookmark it) Most factuality work checks whether the claims in an answer are correct. GAMUT goes after the harder q…

New research from Meta. (bookmark it) Most factuality work checks whether the claims in an answer are correct. GAMUT goes after the harder question of whether the answer covers everything it should. I

→24 Jul 2026Open weights = freedom. You can run them on your own hardware. No vendor can pull the plug. No API can deprecate you. No company logs your p…

Open weights = freedom. You can run them on your own hardware. No vendor can pull the plug. No API can deprecate you. No company logs your private data. That's sovereignty. Closed models hand one comp

→27 Jul 2026Great technical paper from Google. Great read on why context beats scale for agents working against unfamiliar APIs. (bookmark it) GPU kerne…

Great technical paper from Google. Great read on why context beats scale for agents working against unfamiliar APIs. (bookmark it) GPU kernel optimization has KernelBench to hillclimb on. TPUs had not

→28 Jul 2026New research from Meta and CMU. This one is on agentic context management for long horizon tasks. (bookmark it) Production agents accumulate…

New research from Meta and CMU. This one is on agentic context management for long horizon tasks. (bookmark it) Production agents accumulate context every turn. The usual fix compresses on a token thr

→31 Jul 2026Very interesting paper on recursive self-improvement. The whole stack is released. Machine learning engineering gives recursive self-improve…

Very interesting paper on recursive self-improvement. The whole stack is released. Machine learning engineering gives recursive self-improvement a concrete, executable testbed. OpenMLE is an open full

→9 Aug 2026New research from Meta. Agent harnesses are still mostly authored by hand. This makes it hard to tune robust agent harnesses for long-horizo…

New research from Meta. Agent harnesses are still mostly authored by hand. This makes it hard to tune robust agent harnesses for long-horizon tasks. In this new work, agents learn harness policies off

→10 Aug 2026Impressive new paper from Meta. (bookmark it) Scaling laws assume model size and training data act on loss independently. This work introduc…

Impressive new paper from Meta. (bookmark it) Scaling laws assume model size and training data act on loss independently. This work introduces Skaling law, which couples capacity and data through a si

→12 Aug 2026Four small architecture decisions can cost up to 47% of a model's long-context performance. New research from Ai2, Carnegie Mellon, and the …

Four small architecture decisions can cost up to 47% of a model's long-context performance. New research from Ai2, Carnegie Mellon, and the University of Washington isolates them. Normalization, GQA,

CompanyMistral3 recent entries
1 May 2026I have been testing DeepSeek-V4-Pro with the Pi coding agent. I am mindblown by how well it works out of the box. A few notes: I spent a few…

I have been testing DeepSeek-V4-Pro with the Pi coding agent. I am mindblown by how well it works out of the box. A few notes: I spent a few hours building an LLM wiki with an agent powered entirely b

→24 Jun 2026ParseBench is now also available on Papers with Code! Find it here: https://paperswithcode.co/benchmark/parsebench

ParseBench is now also available on Papers with Code! Find it here: https://paperswithcode.co/benchmark/parsebench We benchmarked Mistral OCR against other frontier and open-weight models on ParseBenc

→2 Jul 2026Another fascinating paper on LLM Judges. (bookmark it) It's from Amazon, and they show that if you run panels of LLM judges, averaging their…

Another fascinating paper on LLM Judges. (bookmark it) It's from Amazon, and they show that if you run panels of LLM judges, averaging their scores is a trap. 'Overall, we establish that robust aggreg

CompanyxAI8 recent entries
8 Apr 2026Pretty cool to see Tobi using Hermes and the Manim skill!

Nous Research's Hermes Agent gained a Manim skill that serves as a production pipeline for mathematical and technical animations using Manim Community Edition, creating 3Blue1Brown-style animated ...

→20 Apr 2026Plugins for Excel, PowerPoint, Word, … coming soon

Plugins for Excel, PowerPoint, Word, … coming soon Academics, business leaders, and anyone stuck making PowerPoints: Grok 4.3 just turned a full tDCS/TMS neuroscience paper into this clean 9-slide aca

→20 Apr 2026Grok 4.3 'write an academic paper on general relativity' It produced a 5-page LaTeX paper: >Einstein field equations >Schwarzschild metric >…

Grok 4.3 'write an academic paper on general relativity' It produced a 5-page LaTeX paper: >Einstein field equations >Schwarzschild metric >tensor notation and Christoffel symbols >starlight deflectio

→20 Apr 2026Grok 4.3 now has native LaTeX compilation built into Grok Files

Grok 4.3 has introduced native LaTeX compilation functionality integrated directly into Grok Files, enabling users to compile LaTeX documents without requiring external tools or applications. This fea

→27 Apr 2026Sam Altman cannot be trusted. • The OpenAI board fired him because he was not always honest with them. They said he should not control power…

Sam Altman cannot be trusted. • The OpenAI board fired him because he was not always honest with them. They said he should not control powerful AI. • A major report talked to over 100 people and saw s

→19 May 2026We just added Grok's new imagine model in Paper so you can explore images even faster. Here's what we found: - Super fast generations for th…

We just added Grok's new imagine model in Paper so you can explore images even faster. Here's what we found: - Super fast generations for the quality - Saves details when editing - Perfect for 30 rapi

→21 Jul 2026very notable trajectory comparison writeup here buried in the RLM paper from @a1zhang and @lateinteraction. an open secret of 'frontier' mod…

very notable trajectory comparison writeup here buried in the RLM paper from @a1zhang and @lateinteraction. an open secret of 'frontier' model training is that even without training on test, you can b

→27 Jul 2026btw anthropic's internal document on this literally said 'we don't want it to be known that we are working on this.” it was called project p…

btw anthropic's internal document on this literally said 'we don't want it to be known that we are working on this.” it was called project panama. here's exactly what happened: 1: anthropic concluded

CompanyDeepSeek8 recent entries
1 May 2026I have been testing DeepSeek-V4-Pro with the Pi coding agent. I am mindblown by how well it works out of the box. A few notes: I spent a few…

I have been testing DeepSeek-V4-Pro with the Pi coding agent. I am mindblown by how well it works out of the box. A few notes: I spent a few hours building an LLM wiki with an agent powered entirely b

→4 May 2026Introducing nanowhale 🐳! A tiny DeepSeek model fully pretrained by an agent. Inspired by @karpathy's nanochat, we gave ml-intern the task o…

Introducing nanowhale 🐳! A tiny DeepSeek model fully pretrained by an agent. Inspired by @karpathy's nanochat, we gave ml-intern the task of training a tiny MoE with all the architectural advancements

→15 May 2026The most revealing thing about this AI leadership paper is that it reads less like a vision for innovation and more like a glossy whitepaper…

The most revealing thing about this AI leadership paper is that it reads less like a vision for innovation and more like a glossy whitepaper for a 21st century East India Company. Every generation of

→27 May 2026The MiniMax M2 series was one of the most widely used open-weight LLM series earlier this year. Now, we got a technical report with some int…

The MiniMax M2 series was one of the most widely used open-weight LLM series earlier this year. Now, we got a technical report with some interesting tidbits. I summarized some of them below: 1. Full a

→2 Jul 2026LLM Wikis are being slept on. I argue that creating knowledge bases with LLMs or coding agents is one of the most valuable applications of A…

LLM Wikis are being slept on. I argue that creating knowledge bases with LLMs or coding agents is one of the most valuable applications of AI today. It's about being intentional in building and scalin

→2 Jul 2026DeepSeek effect https://x.com/maximelabonne/status/2070867418377818542

DeepSeek effect https://x.com/maximelabonne/status/2070867418377818542 Fun surprise: DeepSeek used my open-perfectblend dataset to train their new DSpark drafter Time to promote it again! It's an open

→24 Jul 2026As a joke I prompted Codex 'Build and run BenchBench, a benchmark of now good ai is at creating benchmarks. then figure out what benchbenchb…

As a joke I prompted Codex 'Build and run BenchBench, a benchmark of now good ai is at creating benchmarks. then figure out what benchbenchbench is and run that. and then write benchbenchbench up as a

→27 Jul 2026Great technical paper from Google. Great read on why context beats scale for agents working against unfamiliar APIs. (bookmark it) GPU kerne…

Great technical paper from Google. Great read on why context beats scale for agents working against unfamiliar APIs. (bookmark it) GPU kernel optimization has KernelBench to hillclimb on. TPUs had not

CompanyNVIDIA8 recent entries
22 Jul 2026Decoupling embedding from ingestion means whatever hardware is on hand can do the work. One script, runtime device check: MPS on Apple Silic…

Decoupling embedding from ingestion means whatever hardware is on hand can do the work. One script, runtime device check: MPS on Apple Silicon, CUDA on NVIDIA, CPU fallback otherwise. 10K-20K records

→24 Jul 2026Open weights = freedom. You can run them on your own hardware. No vendor can pull the plug. No API can deprecate you. No company logs your p…

Open weights = freedom. You can run them on your own hardware. No vendor can pull the plug. No API can deprecate you. No company logs your private data. That's sovereignty. Closed models hand one comp

→26 Jul 2026New research from NVIDIA. Does AdamW have a scale ceiling? This work claims yes, and shows where it sits. At batch sizes up to 100M tokens f…

New research from NVIDIA. Does AdamW have a scale ceiling? This work claims yes, and shows where it sits. At batch sizes up to 100M tokens for next-token prediction, SOAP and Muon maintain training st

→27 Jul 2026New research from NVIDIA. They just dropped a PyTorch-native training framework for agentic RL. (bookmark it) Paper summary: Molt is a PyTor…

New research from NVIDIA. They just dropped a PyTorch-native training framework for agentic RL. (bookmark it) Paper summary: Molt is a PyTorch-native agentic RL framework with an unusual design target

→29 Jul 2026Super interesting new work from NVIDIA. (bookmark it) They suggest building agents as Python objects. Very cool idea and I think it could a …

Super interesting new work from NVIDIA. (bookmark it) They suggest building agents as Python objects. Very cool idea and I think it could a lot with agent reliability. More below: Agent development to

→29 Jul 2026Impressive paper! It's on one of the hardest tasks for coding agents today. Of course, I am talking about kernel optimization. Coding agents…

Impressive paper! It's on one of the hardest tasks for coding agents today. Of course, I am talking about kernel optimization. Coding agents are usually not so great at this. Reasons: Unfamiliar low-l

→11 Aug 2026🦔Nvidia announced agreements yesterday with the six biggest names in private capital, Apollo, Blackstone, BlackRock, Brookfield, Goldman Sa…

🦔Nvidia announced agreements yesterday with the six biggest names in private capital, Apollo, Blackstone, BlackRock, Brookfield, Goldman Sachs, and KKR, to raise over $500 billion so its own customers

→12 Aug 2026Four small architecture decisions can cost up to 47% of a model's long-context performance. New research from Ai2, Carnegie Mellon, and the …

Four small architecture decisions can cost up to 47% of a model's long-context performance. New research from Ai2, Carnegie Mellon, and the University of Washington isolates them. Normalization, GQA,