AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “papers”

GridTimelineEvolution
693 results
Model Releases

Highly-recommended read from MIT on the part of RL with verifiable rewards that everyone keeps hitting. RLVR only optimizes what you can obj…

DGX agent

Highly-recommended read from MIT on the part of RL with verifiable rewards that everyone keeps hitting. RLVR only optimizes what you can objectively score, so style, structure, and diversity quietly c

model-releasesdair-ai--x
3 Jul 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Research

We are pleased to present our latest research at #ICML2026, “Bridging Spherical Black-Box Optimizers” https://arxiv.org/abs/2606.25761 When …

DGX agent

We are pleased to present our latest research at #ICML2026, “Bridging Spherical Black-Box Optimizers” https://arxiv.org/abs/2606.25761 When optimizing through simulators, external APIs, or in reinforc

researchdavid-ha--x
3 Jul 2026
Agents

When brainstorming new AI agent uses, I always remind myself what used to be expensive and is now nearly zero… 1) cost of reading everything…

DGX agent

When brainstorming new AI agent uses, I always remind myself what used to be expensive and is now nearly zero… 1) cost of reading everything fell → you can now watch 100% instead of a sample (every ex

agentsallie-k--miller--x
3 Jul 2026
Model Releases

// AutoMem // I quite like this idea of metamemory. (bookmark it) This new research from Stanford treats agent's memory management as a trai…

DGX agent

// AutoMem // I quite like this idea of metamemory. (bookmark it) This new research from Stanford treats agent's memory management as a trainable skill instead of a fixed module. The model decides wha

model-releasesdair-ai--x
2 Jul 2026
Model Releases

Bridgewater just published numbers that should make every frontier lab nervous. The world's largest hedge fund tested Gemini, Claude, and GP…

DGX agent

Bridgewater just published numbers that should make every frontier lab nervous. The world's largest hedge fund tested Gemini, Claude, and GPT on six document filtering tasks its investors do every day

model-releasesclem-delangue--x
2 Jul 2026
Model Releases

DeepSeek effect https://x.com/maximelabonne/status/2070867418377818542

DGX agent

DeepSeek effect https://x.com/maximelabonne/status/2070867418377818542 Fun surprise: DeepSeek used my open-perfectblend dataset to train their new DSpark drafter Time to promote it again! It's an open

model-releasesclem-delangue--x
2 Jul 2026
Tutorials

New research from Google. LLMs hallucinate with high confidence, miss their own knowledge boundaries, and misreport uncertainty. Most fixes …

DGX agent

New research from Google. LLMs hallucinate with high confidence, miss their own knowledge boundaries, and misreport uncertainty. Most fixes bolt calibration on from the outside. RLMF turns the model o

tutorialsdair-ai--x
2 Jul 2026
Research

1/ On Training in Imagination - Dwarkesh's episode has a segment on dreaming as one of the next training paradigms. The idea is that a model…

DGX agent

1/ On Training in Imagination - Dwarkesh's episode has a segment on dreaming as one of the next training paradigms. The idea is that a model learns mostly inside its own, by imagining what would happe

researchyann-lecun--x
1 Jul 2026
Research

LeWorld model becomes ADAPTIVE and meets MODEL-PREDICTIVE CONTROL AdaJEPA by Yann LeCun and colleagues performs actions, then checks the pre…

DGX agent

LeWorld model becomes ADAPTIVE and meets MODEL-PREDICTIVE CONTROL AdaJEPA by Yann LeCun and colleagues performs actions, then checks the predicted latent state versus the observed and adapts at TEST T

researchyann-lecun--x
1 Jul 2026
Model Releases

Introducing SWE-Together: a multi-turn benchmark built from real user–agent coding sessions. Coding agents are often benchmarked like exam-t…

DGX agent

Introducing SWE-Together: a multi-turn benchmark built from real user–agent coding sessions. Coding agents are often benchmarked like exam-takers: given the full spec up front, then graded on the fina

model-releasesjeremy-howard--x
30 Jun 2026
Industry

The interesting thing about viral narratives is that they don't have to be true to spread. There's now enough evidence pointing to the fact …

DGX agent

The interesting thing about viral narratives is that they don't have to be true to spread. There's now enough evidence pointing to the fact that AI usage does not kill jobs, but that might not matter

industrycristobal-valenzuela--x
30 Jun 2026
Tools

very successful first poster sessions day at AIEWF, thanks to @heathercmiller and @ACM_President for the support!! tomorrow: Poaster session…

DGX agent

very successful first poster sessions day at AIEWF, thanks to @heathercmiller and @ACM_President for the support!! tomorrow: Poaster sessions! submit your hottest tweets to @vibhuuuus for printing! sp

toolsswyx--x
30 Jun 2026
Model Releases

Why do RL runs on LLMs blow up even when the recipe looks right? GEOALIGN, from the Alibaba team behind Qwen, points at the rollouts. A hand…

DGX agent

Why do RL runs on LLMs blow up even when the recipe looks right? GEOALIGN, from the Alibaba team behind Qwen, points at the rollouts. A handful of bad batches push the policy in incoherent directions,

model-releasesdair-ai--x
28 Jun 2026
Tutorials

If you use LLM-as-judge, this one is worth reading. (bookmark it) It's actually one of the most effective ways to use LLM-as-a-Judge for eva…

DGX agent

If you use LLM-as-judge, this one is worth reading. (bookmark it) It's actually one of the most effective ways to use LLM-as-a-Judge for evals. Holistic judge scores hide both their reasoning and thei

tutorialsdair-ai--x
27 Jun 2026
Model Releases

Thanks for running our open-source work on current frontier models “The results are: the most capable models today (GPT-5.5 Pro) did outperf…

DGX agent

Thanks for running our open-source work on current frontier models “The results are: the most capable models today (GPT-5.5 Pro) did outperform the best models from before (79/100 vs 69/100), but did

model-releasesgary-marcus--x
27 Jun 2026
Model Releases

The take that frontier models aren't ready for use in medicine is dead wrong. I've been using Opus 4.x in my clinical workflow every day for…

DGX agent

The take that frontier models aren't ready for use in medicine is dead wrong. I've been using Opus 4.x in my clinical workflow every day for 6 months. I'm a deep sub-sub-specialist in dermatology - th

model-releasesjeremy-howard--x
27 Jun 2026
Agents

New research from Meta. Building synthetic training data has stayed a fixed pipeline that you hand-tune and then freeze. Autodata casts an A…

DGX agent

New research from Meta. Building synthetic training data has stayed a fixed pipeline that you hand-tune and then freeze. Autodata casts an AI agent as a data scientist that builds training and evaluat

agentsdair-ai--x
25 Jun 2026
Agents

This has a lot in common with how controlled self-evolution works in Row-Bot. And that built on top on LangGraph too. Here's an article I wr…

DGX agent

This has a lot in common with how controlled self-evolution works in Row-Bot. And that built on top on LangGraph too. Here's an article I wrote about how it works a few days ago: https://x.com/i/statu

agentsharrison-chase--x
25 Jun 2026
Model Releases

Video models are hard. Real-time video models are even harder. Loved this breakdown by @itunpredictable on how we productionized Runway Char…

DGX agent

Video models are hard. Real-time video models are even harder. Loved this breakdown by @itunpredictable on how we productionized Runway Characters, our real-time interactive avatar model. We had to fi

model-releasescristobal-valenzuela--x
25 Jun 2026
Model Releases

📣📣 Meet Qwen-AgentWorld — a native language world model that simulates 7 agent environments (MCP, Search, Terminal, SWE, Web, OS, Android)…

DGX agent

📣📣 Meet Qwen-AgentWorld — a native language world model that simulates 7 agent environments (MCP, Search, Terminal, SWE, Web, OS, Android) within a single model. Environment modeling is the training o

model-releasesqwen--x
24 Jun 2026
Model Releases

We open-source Qwen-AgentWorld-35B-A3B (MoE, 35B/3B active, 256K context) and AgentWorldBench. Two routes, one roadmap: 🔬 Build the simulat…

DGX agent

We open-source Qwen-AgentWorld-35B-A3B (MoE, 35B/3B active, 256K context) and AgentWorldBench. Two routes, one roadmap: 🔬 Build the simulator — scalable, controllable, surpassing real environments 🧠 I

model-releasesqwen--x
24 Jun 2026
Agents

auto-research style proposal loops should be data driven! they largely work best only when Data/Evals/Feedback give a useful gradient to hil…

DGX agent

auto-research style proposal loops should be data driven! they largely work best only when Data/Evals/Feedback give a useful gradient to hill-climb against increasingly auto-research is a very good to

agentsharrison-chase--x
23 Jun 2026
Agents

self-harness: harnesses that improve themselves spoiler alert, it's all loop engineering! 1. weakness mining - run agent and observe failure…

DGX agent

self-harness: harnesses that improve themselves spoiler alert, it's all loop engineering! 1. weakness mining - run agent and observe failures 2. propose improvements to the harness 3. confirm improvem

agentsharrison-chase--x
23 Jun 2026
Tools

Single-shot generation still surfaces net-new kernels with no public reference: NeMo vocab-parallel log-probs, Hyena context parallelism, SA…

DGX agent

Single-shot generation still surfaces net-new kernels with no public reference: NeMo vocab-parallel log-probs, Hyena context parallelism, SAM 3 mask suppression. One GEMM + All-Gather kernel hit 87.9µ

toolstogether-ai--x
23 Jun 2026
Safety

Great report on LLM agent communication protocols. Communication is a huge bottleneck in multi-agent systems. (worth bookmarking) The report…

DGX agent

Great report on LLM agent communication protocols. Communication is a huge bottleneck in multi-agent systems. (worth bookmarking) The report builds a five-dimensional taxonomy (counterparty, payload,

safetydair-ai--x
22 Jun 2026
Safety

The largest LLM-as-a-Judge reliability audit yet. Researchers ran 21 judges from nine providers over roughly 541,000 judgments on MT-Bench, …

DGX agent

The largest LLM-as-a-Judge reliability audit yet. Researchers ran 21 judges from nine providers over roughly 541,000 judgments on MT-Bench, JudgeBench, and RewardBench. Findings: Validating a judge wi

safetydair-ai--x
22 Jun 2026
Agents

>> Scalable Evaluation for AI Agents << If you run agent evaluation in production, this one is worth your time. It shows that front-loading …

DGX agent

>> Scalable Evaluation for AI Agents << If you run agent evaluation in production, this one is worth your time. It shows that front-loading human judgment into reusable evaluation assets is useful. Bu

agentsdair-ai--x
21 Jun 2026
Agents

🚨🚨🚨A research project idea! How to measure world models? Everyone's talking about world models these days. World model here, world model …

DGX agent

🚨🚨🚨A research project idea! How to measure world models? Everyone's talking about world models these days. World model here, world model there. We can argue about what 'world model' actually means, an

agentsyann-lecun--x
20 Jun 2026
Agents

// Evolving Meta-Skill for Multi-Agent Systems // Can a multi-agent system get better at orchestration without touching a single weight? Aut…

DGX agent

// Evolving Meta-Skill for Multi-Agent Systems // Can a multi-agent system get better at orchestration without touching a single weight? Automatic MAS generation has been stuck between two bad options

agentsdair-ai--x
20 Jun 2026
Tutorials

GLM-5.2 is great at design (Opus level IMO). I am also starting to see great results with long-running tasks, too. How is this possible? I t…

DGX agent

GLM-5.2 is great at design (Opus level IMO). I am also starting to see great results with long-running tasks, too. How is this possible? I think there are a few clever hacks. But I just came across th

tutorialsdair-ai--x
20 Jun 2026
Local Ai

GLM 5.2 quantized GGUFs: https://huggingface.co/unsloth/GLM-5.2-GGUF GLM 5.2 on Open Router: https://openrouter.ai/z-ai/glm-5.2 Minion codin…

DGX agent

GLM 5.2 quantized GGUFs: https://huggingface.co/unsloth/GLM-5.2-GGUF GLM 5.2 on Open Router: https://openrouter.ai/z-ai/glm-5.2 Minion coding agent: https://github.com/Sentdex/minion A decent consolid

local-aiclem-delangue--x
19 Jun 2026
Model Releases

Over the last two weeks, both the U.S. Government and Anthropic took significant actions that demonstrated their power to control access to …

DGX agent

Over the last two weeks, both the U.S. Government and Anthropic took significant actions that demonstrated their power to control access to AI by restricting what others can do with frontier models. T

model-releasesandrew-ng--x
19 Jun 2026
Tutorials

Causal Clothes-Invariant Feature Learning for Cloth-Changing Person Re-ID

DGX agent

arXiv:2305.06145v2 Announce Type: replace Abstract: In cloth-changing person re-identification (CCReID), it is critical to learn clothes-invariant feature, which can provide discriminative ID features

tutorialsarxiv-cs-cv
11 Jun 2026
Hardware

GPU MODE has powered much of the public GPU kernel work online, with a permissive license from day one and generous credit from researchers,…

DGX agent

GPU MODE has powered much of the public GPU kernel work online, with a permissive license from day one and generous credit from researchers, NVIDIA, AMD, and others. Today we’re moving our datasets to

hardwarejeremy-howard--x
9 Jun 2026
Agents

// Self-Harness: Harnesses That Improve Themselves // (bookmark this one) Most of the agent scaffolds we rely on today are built once and re…

DGX agent

// Self-Harness: Harnesses That Improve Themselves // (bookmark this one) Most of the agent scaffolds we rely on today are built once and remain frozen or mostly unchanged. The harness, like the skill

agentsdair-ai--x
9 Jun 2026
Agents

Kimi Code, our open-source coding agent, just got a major upgrade! 🔹One-line CLI install, zero setup, fast startup​ 🔹Drag in videos as cod…

DGX agent

Kimi Code, our open-source coding agent, just got a major upgrade! 🔹One-line CLI install, zero setup, fast startup​ 🔹Drag in videos as coding context: reference-to-LUT, long-video-to-short, screen-rec

agentskimi-moonshot--x
8 Jun 2026
Safety

No. Not by itself. Sergey Brin is absolutely wrong. Transformers by themselves are not “sufficient” for AGI. Nobody uses transformers on the…

DGX agent

No. Not by itself. Sergey Brin is absolutely wrong. Transformers by themselves are not “sufficient” for AGI. Nobody uses transformers on their own anymore. Everybody is supplementing them tools and ha

safetygary-marcus--x
8 Jun 2026
Applications

The Matrix idea of keeping humans as batteries is obviously weird... we would be more useful as dice. LLMs default to very similar kinds of …

DGX agent

The Matrix idea of keeping humans as batteries is obviously weird... we would be more useful as dice. LLMs default to very similar kinds of arguments & structure, and even different LLMs seem to colla

applicationsethan-mollick--x
8 Jun 2026
Safety

One of the more interesting takes on positive alignment that have recently come out-it’s long and interesting, combining philosophy and trai…

DGX agent

One of the more interesting takes on positive alignment that have recently come out-it’s long and interesting, combining philosophy and training setups (eg reward proposals), and worth a read. What ha

safetydan-hendrycks--x
7 Jun 2026
Tutorials

// Continual Learning Bench // One of the research areas with lots of investments is continual learning. While there are many efforts, there…

DGX agent

// Continual Learning Bench // One of the research areas with lots of investments is continual learning. While there are many efforts, there is very little progress in measuring it. So the big questio

tutorialsdair-ai--x
6 Jun 2026
Safety

https://x.com/CAIS/status/2060031683420999844?s=20

DGX agent

https://x.com/CAIS/status/2060031683420999844?s=20 AI systems may soon help run economies, infrastructure, and military operations. But these systems are not reliably loyal or secure. An adversary can

safetydan-hendrycks--x
6 Jun 2026
Local Ai

New research from Renmin University. Treat skill selection as a harness in its own right. If you design skill routing for personal or edge a…

DGX agent

New research from Renmin University. Treat skill selection as a harness in its own right. If you design skill routing for personal or edge agents, this work argues that the selection layer is a first-

local-aidair-ai--x
6 Jun 2026
Model Releases

// Agents' Last Exam // Agents' Last Exam is a living benchmark of over 1,000 economically valuable tasks, built with 250+ industry experts …

DGX agent

// Agents' Last Exam // Agents' Last Exam is a living benchmark of over 1,000 economically valuable tasks, built with 250+ industry experts and mapped to the U.S. federal occupational taxonomy. The ha

model-releasesdair-ai--x
5 Jun 2026
Model Releases

Critical context on the new Anthropic blog: 1, AGI is *harder* than RSI (as used below). AGI: machine can do anything human can do, autonomo…

DGX agent

Critical context on the new Anthropic blog: 1, AGI is *harder* than RSI (as used below). AGI: machine can do anything human can do, autonomously [not achieved] RSI (as used below): AI is a useful codi

model-releasesgary-marcus--x
5 Jun 2026
Model Releases

I always appreciate the opportunity to discuss @LawZero_ and our approach to honest, reliable AI. Working on the Scientist AI with my brilli…

DGX agent

I always appreciate the opportunity to discuss @LawZero_ and our approach to honest, reliable AI. Working on the Scientist AI with my brilliant colleagues at LawZero has made me very confident that we

model-releasesyoshua-bengio--x
5 Jun 2026
Model Releases

Leaving aside the question of consciousness, the Ted Chiang piece has a reasonable point about moral atrophy if you let AI make choices. But…

DGX agent

Leaving aside the question of consciousness, the Ted Chiang piece has a reasonable point about moral atrophy if you let AI make choices. But it is also interesting in light of the fact that repeated r

model-releasesethan-mollick--x
4 Jun 2026
Model Releases

We're presenting ParseBench at CVPR 2026! ParseBench is the most comprehensive document understanding benchmark for VLMs. ✅ It contains 2k p…

DGX agent

We're presenting ParseBench at CVPR 2026! ParseBench is the most comprehensive document understanding benchmark for VLMs. ✅ It contains 2k pages of real-world enterprise documents ✅ It has comprehensi

model-releasesjerry-liu--x
4 Jun 2026
Agents

New research from Google. Just shows the impressive results you can get from custom agent harnesses. LEAP wraps a general-purpose LLM in an …

DGX agent

New research from Google. Just shows the impressive results you can get from custom agent harnesses. LEAP wraps a general-purpose LLM in an agentic scaffold that grounds every step in the Lean compile

agentsdair-ai--x
3 Jun 2026
← Previous
1…101112131415
Next →