CompanyAnthropic5 recent entries1 May 2026RL is a bit of a double edged sword: in known territory performance increases, but in unknown territory the model tends to hallucinate that …RL is a bit of a double edged sword: in known territory performance increases, but in unknown territory the model tends to hallucinate that it is performing a completely different task it was trained →1 Jun 2026Anthropic Opus 4.8 is new SOTA on ARC-AGI-3 Score: 1.5%, ~$10K ARC-AGI-3 analysis notes: * Opus 4.8 read the environment an abstraction *abo…Anthropic Opus 4.8 is new SOTA on ARC-AGI-3 Score: 1.5%, ~$10K ARC-AGI-3 analysis notes: * Opus 4.8 read the environment an abstraction *above* Opus 4.7, as objects & systems, not pictures * Opus 4.8
CompanyOpenAI3 recent entries1 Jul 2026Cross-agent feedback loops are incredibly effective -- for a reason. Check out what @leon2mcp and team at @Bloome_im are building in this sp…Cross-agent feedback loops are incredibly effective -- for a reason. Check out what @leon2mcp and team at @Bloome_im are building in this space: http://bloome.im Bloome lets you pull Claude, ChatGPT, →30 Jul 2026Quick reminder of what's ok vs not ok with harnesses used for playing ARC-AGI-3: 1. Not okay: harnesses that were custom-made to solve the b…Quick reminder of what's ok vs not ok with harnesses used for playing ARC-AGI-3: 1. Not okay: harnesses that were custom-made to solve the benchmark or that contain knowledge about the benchmark forma→6 Aug 2026We re-tested GPT-5.6 Luna from @OpenAI on ARC-AGI (Verified) following its recent 80% price reduction: - ARC-AGI-2: 59.6%, $0.18/task - ARC-…We re-tested GPT-5.6 Luna from @OpenAI on ARC-AGI (Verified) following its recent 80% price reduction: - ARC-AGI-2: 59.6%, 0.18/task - ARC-AGI-1: 90.7%, 0.07/task The new results match Luna's original
CompanyGoogle5 recent entries1 May 2026RL is a bit of a double edged sword: in known territory performance increases, but in unknown territory the model tends to hallucinate that …RL is a bit of a double edged sword: in known territory performance increases, but in unknown territory the model tends to hallucinate that it is performing a completely different task it was trained →19 May 2026GeminiGemini Gemini 3.5 Flash ARC-AGI (Verified) ARC-AGI-2: - High: 72.1%, 0.85 - Minimal: 8.9%, 0.11 ARC-AGI-1: - High: 92.5%, 0.42 - Minimal: 48.8%, 0.06 Gemini 3.5 Flash is on par with GPT-5.5 (Medium) o→1 Jul 2026Cross-agent feedback loops are incredibly effective -- for a reason. Check out what @leon2mcp and team at @Bloome_im are building in this sp…Cross-agent feedback loops are incredibly effective -- for a reason. Check out what @leon2mcp and team at @Bloome_im are building in this space: http://bloome.im Bloome lets you pull Claude, ChatGPT, →6 Aug 2026We re-tested GPT-5.6 Luna from @OpenAI on ARC-AGI (Verified) following its recent 80% price reduction: - ARC-AGI-2: 59.6%, $0.18/task - ARC-…We re-tested GPT-5.6 Luna from @OpenAI on ARC-AGI (Verified) following its recent 80% price reduction: - ARC-AGI-2: 59.6%, 0.18/task - ARC-AGI-1: 90.7%, 0.07/task The new results match Luna's original→8 Aug 2026The reports of the demise of Google are greatly exaggerated. I wouldn't underestimate themFrançois Chollet commented that claims the demise of Google were greatly exaggerated, cautioning against undervaluation. According to a Polymarket report, Sergey Brin is expected to take direct oversi
CompanyMeta3 recent entries8 Apr 2026The new model from Meta is already looking like a disappointment: overoptimized for public benchmark numbers at the detriment of everything …The new model from Meta is already looking like a disappointment: overoptimized for public benchmark numbers at the detriment of everything else. Knowing how to evaluate models in a way that correlate→23 Jun 2026Casual: Token maxxing Sweaty: Token minning Meta: Token min-maxingFrancois Chollet presents a framework for understanding different approaches to token usage in AI models: casual users maximize tokens for flexibility, competitive users minimize tokens for efficiency→10 Aug 2026Coding isn't yet another application domain -- it's the meta-skill required for AI to automatically develop its own training material, via s…Coding isn't yet another application domain -- it's the meta-skill required for AI to automatically develop its own training material, via symbolic world models. That's how the RSI loop actually kicks
CompanyxAI1 recent entries19 Apr 2026There's no doubt that the world can consume tokens as fast as they're produced, even in the most maximalist infrastructure buildup scenarios…There's no doubt that the world can consume tokens as fast as they're produced, even in the most maximalist infrastructure buildup scenarios imaginable. That's not the question. The question is whethe
CompanyDeepSeek1 recent entries7 Aug 2026Outstanding cost-to-performance from DeepSeek GPT-5.6 Luna (Max) performance for a 1/4th of the cost on ARC-AGIOutstanding cost-to-performance from DeepSeek GPT-5.6 Luna (Max) performance for a 1/4th of the cost on ARC-AGI DeepSeek V4 Flash from @deepseek_ai on ARC-AGI (Verified): - ARC-AGI-2: 61.4%, 0.04/task
CompanyNVIDIA1 recent entries10 Apr 2026The power of JAXThe power of JAX Introducing gyaradax 🐉: A JAX solver for local flux-tube gyrokinetics with custom CUDA kernels for acceleration. This entire code was vibecoded by @ggalletti_ and me in a month. Valid