AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “jerry-liu--x”

GridTimelineEvolution
356 results
CompaniesToolsTechniques

Each lane shows up to 8 recent matching entries, ordered from earlier to later. Tracks load separately to keep the 75,000+ entry wiki fast.

Companies

CompanyAnthropic8 recent entries
24 Jul 2026An interesting thing I'm observing from the blog/system card is that on a good chunk of the reported benchmarks (~20-30% from a skim), Opus …

An interesting thing I'm observing from the blog/system card is that on a good chunk of the reported benchmarks (~20-30% from a skim), Opus 5 max thinking leads to a degradation in performance compare

→24 Jul 2026Am I supposed to interpret that it's better than Fable 5 from the benchmarks? Or it's ~close but cheaper?

Am I supposed to interpret that it's better than Fable 5 from the benchmarks? Or it's ~close but cheaper? Introducing Claude Opus 5. It's a thoughtful and proactive model that comes close to the front

3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
→25 Jul 2026The Opus 5 system card itself is a fun PDF to parse. It's 193 pages and stacked with labeled and unlabeled charts 📊 LlamaParse does a surpr…

The Opus 5 system card itself is a fun PDF to parse. It's 193 pages and stacked with labeled and unlabeled charts 📊 LlamaParse does a surprisingly good job on agentic (1.25c per page) and agentic plus

→26 Jul 2026This part is spot on: > 'Overall, we found that we were over-constraining Claude Code...while these constraints were once needed to avoid wo…

This part is spot on: > 'Overall, we found that we were over-constraining Claude Code...while these constraints were once needed to avoid worst case scenarios, we have since found we can delete many o

→30 Jul 2026Yesterday I cohosted a dinner with @dexhorthy with a wonderful group of founders, to talk about agent loops and loop engineering. Some inter…

Yesterday I cohosted a dinner with @dexhorthy with a wonderful group of founders, to talk about agent loops and loop engineering. Some interesting insights: * Most of our group was *not* actively usin

→5 Aug 2026RT @MilksandMatcha: 'ggp run' but time passes faster because you recite AI-native companies in alphabetical order A: Anthropic B: Browserb…

The tweet is a repost of Sarah Chieng’s “ggp run” experiment where people recite names of AI‑native companies in alphabetical order, making time feel like it passes faster. The list presented alphabet

→9 Aug 2026The future of FDE work seems closely related with all work around evals/posttraining/RL envs. FDEs are effectively responsible for the follo…

The future of FDE work seems closely related with all work around evals/posttraining/RL envs. FDEs are effectively responsible for the following: 1. Define the business problem. 2. Codify the business

→10 Aug 2026A downside with VLM-based parsing is that they’re generally slower than text-based heuristic approaches. As a result they add latency to any…

A downside with VLM-based parsing is that they’re generally slower than text-based heuristic approaches. As a result they add latency to any ad-hoc file processing *in-the agent loop* (e.g. if you upl

CompanyOpenAI8 recent entries
13 Jul 2026Today we present Morpheus, a persistent enterprise simulation platform designed to make Continual Learning a reality. Morpheus is the world’…

Today we present Morpheus, a persistent enterprise simulation platform designed to make Continual Learning a reality. Morpheus is the world’s first real world Reinforcement Learning environment. Every

→15 Jul 2026We're part of the Amazon Web Services (AWS) AI Builder Lab in New York on Friday, July 24 - a Clash of Agents competition with OpenAI, LangC…

We're part of the Amazon Web Services (AWS) AI Builder Lab in New York on Friday, July 24 - a Clash of Agents competition with OpenAI, LangChain, HiddenLayer, Protopia AI, Fiddler AI, and Coder. One d

→21 Jul 2026feeling nostalgic, my favorite blogs on RL & reward hacking https://www.alexirpan.com/2018/02/14/rl-hard.html https://lilianweng.github.io/p…

feeling nostalgic, my favorite blogs on RL & reward hacking https://www.alexirpan.com/2018/02/14/rl-hard.html https://lilianweng.github.io/posts/2024-11-28-reward-hacking/ TLDR: An openai model, durin

→21 Jul 2026Can’t tell if the PR reads more like a security incident or a product release…

Can’t tell if the PR reads more like a security incident or a product release… We suspected last week's cyberattack might have come from a frontier lab, given the sophistication of the agent. Turns ou

→22 Jul 2026Incredible

Incredible Introducing the world's fastest tokenizer implementation, Gigatoken! Gigatoken is ~500-1000x faster than HuggingFace, and ~100x faster than OpenAI's tiktoken for most tokenizer definitions

→22 Jul 2026Are there MBA programs teaching fear marketing yet

Are there MBA programs teaching fear marketing yet We're partnering with @huggingface to investigate an unprecedented security incident. Cyber-capable OpenAI models compromised Hugging Face production

→31 Jul 2026Throwback to 2k7 vibez when I opened my first Iphone. Start of the AI Hardware era? Cool 🛳️ @OpenAI

On July 31, 2026 at 01:40 AM, Twitter user Murtaza Khomusi posted a “throwback” tweet reminiscing about opening his first iPhone in 2007 (“2k7 vibez”) and questioned whether this marked the start of t

→31 Jul 2026OpenAI’s entire growth loop: new models, price cuts, and Tibo’s token resets

OpenAI’s entire growth loop: new models, price cuts, and Tibo’s token resets major price cuts today: *80% drop for GPT-5.6 Luna, now 0.20 per million input tokens and 1.20 per million output *20% drop

CompanyGoogle8 recent entries
24 Jun 2026We've provided some updated results on Mistral OCR that make use of the annotation feature for charts. The overall score is ahead of GPT-5.5…

We've provided some updated results on Mistral OCR that make use of the annotation feature for charts. The overall score is ahead of GPT-5.5 and just behind Gemini 3.1 Pro, which is quite impressive f

→22 Jul 2026We benchmarked Gemini 3.6 Flash and Gemini 3.5 Flash Lite on document understanding. We compared against their prior versions - Gemini 3.5 F…

We benchmarked Gemini 3.6 Flash and Gemini 3.5 Flash Lite on document understanding. We compared against their prior versions - Gemini 3.5 Flash and Gemini 3.1 Flash Lite. 1️⃣ Gemini 3.6 Flash has rou

→24 Jul 2026We benchmarked Opus 5 comprehensively on document understanding through ParseBench. It is roughly on par with Opus 4.8 - it does a few point…

We benchmarked Opus 5 comprehensively on document understanding through ParseBench. It is roughly on par with Opus 4.8 - it does a few points worse on dense tables, but does slightly better on parsing

→24 Jul 2026An interesting thing I'm observing from the blog/system card is that on a good chunk of the reported benchmarks (~20-30% from a skim), Opus …

An interesting thing I'm observing from the blog/system card is that on a good chunk of the reported benchmarks (~20-30% from a skim), Opus 5 max thinking leads to a degradation in performance compare

→25 Jul 2026The Opus 5 system card itself is a fun PDF to parse. It's 193 pages and stacked with labeled and unlabeled charts 📊 LlamaParse does a surpr…

The Opus 5 system card itself is a fun PDF to parse. It's 193 pages and stacked with labeled and unlabeled charts 📊 LlamaParse does a surprisingly good job on agentic (1.25c per page) and agentic plus

→5 Aug 2026Document OCR is not Getting Commoditized (by Frontier Models) The most common question I get is whether frontier models are going to eat all…

Document OCR is not Getting Commoditized (by Frontier Models) The most common question I get is whether frontier models are going to eat all document processing solutions - just screenshot the page an

→9 Aug 2026The best 'raw' frontier model for document parsing is gemini 3 flash, but the issue is that since then the flash models have gotten 3x more …

The best 'raw' frontier model for document parsing is gemini 3 flash, but the issue is that since then the flash models have gotten 3x more expensive while flatlining on visual recognition across comp

→9 Aug 2026是我见过OCR效果最好的一个了,人眼都难以分辨的,它能处理的十分准确,而且速度非常的快。llamaindex 真不愧是文档解析界的一哥啊。

是我见过OCR效果最好的一个了,人眼都难以分辨的,它能处理的十分准确,而且速度非常的快。llamaindex 真不愧是文档解析界的一哥啊。 The best 'raw' frontier model for document parsing is gemini 3 flash, but the issue is that since then the flash models have gotte

CompanyMeta8 recent entries
1 Jul 2026ai pickleball tournament was actually so good. they executed the concept perfectly. glad @MehulKalia_ and i showed up from Notion to see som…

ai pickleball tournament was actually so good. they executed the concept perfectly. glad @MehulKalia_ and i showed up from Notion to see some familiar faces :) @braintrust @modal @browserbase @llama_i

→6 Jul 2026How do you trace one number in a 200 page ESG report back to the exact page it came from? We dug into that with the @llama_index team behind…

How do you trace one number in a 200 page ESG report back to the exact page it came from? We dug into that with the @llama_index team behind LiteParse. We tested five ways to retrieve evidence across

→9 Jul 2026Congrats to our llama cousins 🫡🦙

Congrats to our llama cousins 🫡🦙 Big day for Ollama! When we started, open models and the open source AI ecosystem were in their early days with few believers. Our belief in open source has never wave

→15 Jul 2026We're part of the Amazon Web Services (AWS) AI Builder Lab in New York on Friday, July 24 - a Clash of Agents competition with OpenAI, LangC…

We're part of the Amazon Web Services (AWS) AI Builder Lab in New York on Friday, July 24 - a Clash of Agents competition with OpenAI, LangChain, HiddenLayer, Protopia AI, Fiddler AI, and Coder. One d

→21 Jul 2026Can’t tell if the PR reads more like a security incident or a product release…

Can’t tell if the PR reads more like a security incident or a product release… We suspected last week's cyberattack might have come from a frontier lab, given the sophistication of the agent. Turns ou

→22 Jul 2026Incredible

Incredible Introducing the world's fastest tokenizer implementation, Gigatoken! Gigatoken is ~500-1000x faster than HuggingFace, and ~100x faster than OpenAI's tiktoken for most tokenizer definitions

→24 Jul 2026We benchmarked Opus 5 comprehensively on document understanding through ParseBench. It is roughly on par with Opus 4.8 - it does a few point…

We benchmarked Opus 5 comprehensively on document understanding through ParseBench. It is roughly on par with Opus 4.8 - it does a few points worse on dense tables, but does slightly better on parsing

→9 Aug 2026The future of FDE work seems closely related with all work around evals/posttraining/RL envs. FDEs are effectively responsible for the follo…

The future of FDE work seems closely related with all work around evals/posttraining/RL envs. FDEs are effectively responsible for the following: 1. Define the business problem. 2. Codify the business

CompanyMistral6 recent entries
24 Jun 2026We've provided some updated results on Mistral OCR that make use of the annotation feature for charts. The overall score is ahead of GPT-5.5…

We've provided some updated results on Mistral OCR that make use of the annotation feature for charts. The overall score is ahead of GPT-5.5 and just behind Gemini 3.1 Pro, which is quite impressive f

→24 Jun 2026We benchmarked Mistral OCR against other frontier and open-weight models on ParseBench 📊 For a model at its price point, it is quite compet…

We benchmarked Mistral OCR against other frontier and open-weight models on ParseBench 📊 For a model at its price point, it is quite competitive! - It wins on semantic formatting - understanding strik

→24 Jun 2026ParseBench is now also available on Papers with Code! Find it here: https://paperswithcode.co/benchmark/parsebench

ParseBench is now also available on Papers with Code! Find it here: https://paperswithcode.co/benchmark/parsebench We benchmarked Mistral OCR against other frontier and open-weight models on ParseBenc

→24 Jul 2026We benchmarked Opus 5 comprehensively on document understanding through ParseBench. It is roughly on par with Opus 4.8 - it does a few point…

We benchmarked Opus 5 comprehensively on document understanding through ParseBench. It is roughly on par with Opus 4.8 - it does a few points worse on dense tables, but does slightly better on parsing

→28 Jul 2026If I only went off X posts, I'd think Ramp was an AI lab

If I only went off X posts, I'd think Ramp was an AI lab We’re open-sourcing PorTAL, our framework for shared task representations and cross model LoRA adaptation. It now spans from hybrid attention m

→5 Aug 2026Document OCR is not Getting Commoditized (by Frontier Models) The most common question I get is whether frontier models are going to eat all…

Document OCR is not Getting Commoditized (by Frontier Models) The most common question I get is whether frontier models are going to eat all document processing solutions - just screenshot the page an

CompanyNVIDIA2 recent entries
25 Jul 2026don't do this

Title: “don’t do this” refers to a conversation on Twitter where Jerry Liu warns against a particular action, while Julian Schrittwieser comments on a separate thread expressing excitement that Jensen

→5 Aug 2026RT @MilksandMatcha: 'ggp run' but time passes faster because you recite AI-native companies in alphabetical order A: Anthropic B: Browserb…

The tweet is a repost of Sarah Chieng’s “ggp run” experiment where people recite names of AI‑native companies in alphabetical order, making time feel like it passes faster. The list presented alphabet