AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “francois-chollet--x”

GridTimelineEvolution
169 results
1 May 2026

Make sure to read the blog post for a detailed analysis of frontier model failure modes: https://arcprize.org/blog/arc-agi-3-gpt-5-5-opus-4-…

Model ReleasesDGX agent

Francois Chollet shared a blog post analyzing failure modes of frontier AI models, specifically examining performance on the ARC (Abstraction and Reasoning Corpus) AGI benchmark with models including

RL is a bit of a double edged sword: in known territory performance increases, but in unknown territory the model tends to hallucinate that …

Model ReleasesDGX agent

RL is a bit of a double edged sword: in known territory performance increases, but in unknown territory the model tends to hallucinate that it is performing a completely different task it was trained

The latest crop of models remains below 1% on ARC-AGI-3 -- for now. Where will the scores be by the end of the year?


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

The latest crop of models remains below 1% on ARC-AGI-3 -- for now. Where will the scores be by the end of the year? GPT-5.5 & Opus 4.7 on ARC-AGI-3 - GPT-5.5: 0.43% - Opus 4.7: 0.18% We found 3 failu

30 Apr 2026

AI automates tasks, not jobs, and when a task gets cheaper, demand for the job grows. AI cannot automate jobs end-to-end because it lacks au…

ResearchDGX agent

AI automates tasks, not jobs, and when a task gets cheaper, demand for the job grows. AI cannot automate jobs end-to-end because it lacks autonomy and cannot operate without supervision. There is stil

29 Apr 2026

I built a TPU-native medical Q&A fine-tuning pipeline using Gemma 3 with Keras and JAX The project fine-tunes Gemma-3 on medical dialogue da…

Model ReleasesDGX agent

I built a TPU-native medical Q&A fine-tuning pipeline using Gemma 3 with Keras and JAX The project fine-tunes Gemma-3 on medical dialogue data from ChatDoctor and evaluates it on MedMCQA, with a large

28 Apr 2026

ARC Prize 2026: ARC-AGI-3 now has H100s Thank you to our partners at @kaggle for powering the ARC-AGI-3 competition with H100s Participants …

ResearchDGX agent

ARC Prize 2026: ARC-AGI-3 now has H100s Thank you to our partners at @kaggle for powering the ARC-AGI-3 competition with H100s Participants now have a pool of accelerators (subject to submission limit

27 Apr 2026

Keras Kinetic has a new alpha release: v0.0.2! Including a new docs website: http://kinetic.readthedocs.io Kinetic is my favorite new releas…

HardwareDGX agent

Keras Kinetic has a new alpha release: v0.0.2! Including a new docs website: http://kinetic.readthedocs.io Kinetic is my favorite new release from the Keras team: a super simple Modal-like API to run

Read the release notes: https://github.com/keras-team/kinetic/releases/tag/0.0.2

ResearchDGX agent

Kinetic version 0.0.2 release notes are available on GitHub, detailing updates and features for this Keras-related project. The announcement was shared by François Chollet, Keras creator, on X (former

26 Apr 2026

Describe Seattle:

ResearchDGX agent

Describe Seattle: I want to live near the sea and I want to live near the forest and I want to live near the mountains and I want to live near a café and I want to live near a library and I want to li

If you want to get dumber, confused, and lacking any situational awareness, listen to anonymous 'AI influencer' drivel accounts...

ResearchDGX agent

François Chollet criticizes anonymous AI influencer accounts on social media, warning that following their content can lead to diminished critical thinking and poor judgment. The post cautions against

No, the top score if you didn't account for action efficiency would be 100%, achievable with 20 lines of Python. All you need is to brute-fo…

ResearchDGX agent

No, the top score if you didn't account for action efficiency would be 100%, achievable with 20 lines of Python. All you need is to brute-force the state space. Please stop spreading complete disinfor

Though to be fair it's great engagement farming, the bi-weekly Twitter payouts must be juicy.

ResearchDGX agent

Francois Chollet comments on Twitter's bi-weekly payout system, noting that while it may be effective for engagement farming, the financial incentives are likely substantial. The post appears to criti

To all the clowns saying that Seattle is always rainy and never sees the sun: Seattle has *less* precipitation than NYC, Boston, Atlanta, Mi…

ResearchDGX agent

To all the clowns saying that Seattle is always rainy and never sees the sun: Seattle has *less* precipitation than NYC, Boston, Atlanta, Miami, New Orleans, etc... in fact it rains less in Seattle th

25 Apr 2026

#ICLR2026 Keynote talk from Percy Liang @percyliang happening now 🔥

ResearchDGX agent

Percy Liang delivered a keynote address at ICLR 2026, a major machine learning conference. The talk was announced and promoted by Francois Chollet on social media. The specific content of Liang's keyn

23 Apr 2026

GPT-5.5 on ARC-AGI (Verified) ARC-AGI-2: - Max: 85.0%, 1.87 - High: 83.3%, 1.45 - Med: 70.4%, 0.86 - Low: 33%, 0.35 GPT-5.5 is now state…

Model ReleasesDGX agent

GPT-5.5 achieved state-of-the-art performance on the ARC-AGI-2 benchmark, with scores ranging from 85.0% on maximum difficulty tasks to 33% on low difficulty tasks. The model demonstrated consistent i

22 Apr 2026

Judging AGI by how well it can mimic us is a category error, because mimicry isn't intelligence and isn't general. We should judge AGI by ho…

TutorialsDGX agent

Judging AGI by how well it can mimic us is a category error, because mimicry isn't intelligence and isn't general. We should judge AGI by how well it learns to do things we didn't teach it (including

21 Apr 2026

My favorite military drone company is Paris-based Harmattan (http://harmattan.ai). They make more than just quadcopters.

ResearchDGX agent

Harmattan is a Paris-based military drone company that produces more than quadcopters, according to a post by AI researcher François Chollet. The company's product line extends beyond standard quadrot

One of the most jarring things about current AI is its lack of introspection ability and metacognition. It doesn't know what it doesn't know…

ResearchDGX agent

One of the most jarring things about current AI is its lack of introspection ability and metacognition. It doesn't know what it doesn't know, how it knows, or how it could find out. It's a one-way sys

The very first quadcopter prototype was the Breguet-Richet Gyroplane No. 1, created in 1907 by Breguet Aviation in France. The first consume…

Model ReleasesDGX agent

The very first quadcopter prototype was the Breguet-Richet Gyroplane No. 1, created in 1907 by Breguet Aviation in France. The first consumer quadcopter drone was the Parrot AR.Drone, released in 2010

20 Apr 2026

Human biological limits, like our tiny working memory and shallow calculation depth, are actually a feature. They force us to abstract, comp…

ResearchDGX agent

Human biological limits, like our tiny working memory and shallow calculation depth, are actually a feature. They force us to abstract, compress, intuit. If we had infinite resources, we would never h

In a capitalist economy, absolute levels of value creation and productivity take a backseat to relative competitive advantage. That's why jo…

ResearchDGX agent

In a capitalist economy, absolute levels of value creation and productivity take a backseat to relative competitive advantage. That's why jobs still exist despite 150 years of intense automation. As l

In a hypothetical world where literally everything could be automated at no cost (we're very far from that), everybody would work in interpe…

ResearchDGX agent

In a hypothetical scenario where all tasks could be automated at zero cost, François Chollet argues that everyone would work in interpersonal roles, suggesting that human value would shift entirely to

19 Apr 2026

Exactly. The number of iterations you do is far more important than the amount of 'thinking' you do. It doesn't matter how intelligent or kn…

ResearchDGX agent

Exactly. The number of iterations you do is far more important than the amount of 'thinking' you do. It doesn't matter how intelligent or knowledgeable you are. Reality will always push back. The worl

Human cognitive friction has long been acting as a regularizer for a lot of digital infrastructure. It made software APIs less terrible and …

ResearchDGX agent

Human cognitive friction has long been acting as a regularizer for a lot of digital infrastructure. It made software APIs less terrible and codebases less complex. Now LLM disintermediation is causing

I'd say cognitive friction is even more than just a regularizer: it's an incentive to find the right interface abstractions, and in turn goo…

ResearchDGX agent

I'd say cognitive friction is even more than just a regularizer: it's an incentive to find the right interface abstractions, and in turn good abstractions are what enables compounding over time. Piles

There's no doubt that the world can consume tokens as fast as they're produced, even in the most maximalist infrastructure buildup scenarios…

ApplicationsDGX agent

There's no doubt that the world can consume tokens as fast as they're produced, even in the most maximalist infrastructure buildup scenarios imaginable. That's not the question. The question is whethe

Unlimited demand for something doesn't necessarily mean it is viable as a business, much less as an economy-wide all-in bet.

ResearchDGX agent

Unlimited demand for a product or service does not guarantee business viability or economic sustainability, as other critical factors such as cost structure, scalability, profitability, and resource c

18 Apr 2026

Constraints are the catalyst of invention. An infinite search space leads to paralysis. The most creative inventions happen when you are for…

ResearchDGX agent

Constraints are the catalyst of invention. An infinite search space leads to paralysis. The most creative inventions happen when you are forced to solve a problem within appropriately narrow constrain

The quality of your thinking is a multiplier for the amount of progress you make at each iteration. But the dominant factor behind success i…

ResearchDGX agent

The quality of thinking acts as a multiplier on progress made during iterative work, but iteration frequency itself is the primary driver of success. This insight from AI researcher François Chollet e

There's a bit of a 2010s PHP vs Go flavor in there.

ResearchDGX agent

Francois Chollet compares programming language philosophies between PHP and Go, likely discussing design trade-offs, developer experience, or use cases that differentiate the two languages. The compar

When looking at deep learning profiles, one of the most obvious tells between a mediocre and great candidate is whether they list PyTorch or…

ResearchDGX agent

Francois Chollet discusses a hiring signal for identifying high-quality deep learning candidates based on their choice of machine learning frameworks, suggesting PyTorch proficiency (or lack thereof)

You cannot think your way to a perfect design. Only building and testing, over many iterations, can reveal the flaws in your mental model an…

ResearchDGX agent

You cannot think your way to a perfect design. Only building and testing, over many iterations, can reveal the flaws in your mental model and provide the feedback you need to create the best design po

16 Apr 2026

The same is of course true of software -- to solve a problem in a scalable manner requires a lot more work and a lot more code compared to a…

TutorialsDGX agent

This post discusses how software development, like other engineering disciplines, requires substantially more effort and code when building scalable solutions compared to quick prototypes or one-off i

There's a broadly held misconception in AI that methods that scale well are simple methods -- even, that simple methods usually scale. This …

ResearchDGX agent

There's a broadly held misconception in AI that methods that scale well are simple methods -- even, that simple methods usually scale. This is completely wrong. Pretty much none of the truly simple me

15 Apr 2026

Any smart human giving it real effort should score >90% on ARC-AGI-3

ResearchDGX agent

ARC-AGI-3 is a new benchmark iteration designed so that any intelligent human applying genuine effort should achieve greater than 90% accuracy, maintaining the benchmark's core principle that tasks mu

ARC-AGI-3 has the lowest human bar of any AI benchmark out there. Almost all benchmarks require specialized knowledge that make them inacces…

Model ReleasesDGX agent

ARC-AGI-3 has the lowest human bar of any AI benchmark out there. Almost all benchmarks require specialized knowledge that make them inaccessible to 99%+ of humans (like, say SWE-Bench). ARC-AGI-3 is

To score 100% on a game, you just need to beat the median action efficiency of an unfiltered pool of random people. Easy if you're a bit sma…

ResearchDGX agent

Francois Chollet makes the observation that achieving a perfect score on certain AI benchmark games or tasks only requires surpassing the median performance of an unfiltered random population sample,

14 Apr 2026

ARC-AGI-3 Human Baseline Dataset Today we're open-sourcing the ARC-AGI-3 Human Baseline. This is the most exhaustive human testing study in …

ResearchDGX agent

ARC-AGI-3 Human Baseline Dataset Today we're open-sourcing the ARC-AGI-3 Human Baseline. This is the most exhaustive human testing study in the ARC-AGI series Every environment was solved by at least

In order to make sure ARC-AGI-3 games were solvable by humans we tested over 450 people If a game was too hard we either revised it or tosse…

ResearchDGX agent

In order to make sure ARC-AGI-3 games were solvable by humans we tested over 450 people If a game was too hard we either revised it or tossed it completely For maximum transparency we just open source

Simply retrieving a reasoning trace looks a lot like human reasoning, until it's time to navigate uncharted territory. If you memorized all …

ResearchDGX agent

Simply retrieving a reasoning trace looks a lot like human reasoning, until it's time to navigate uncharted territory. If you memorized all reasoning traces of humans from 10,000 BC, you could automat

The role of memorization and knowledge is to cache & reuse past cognitive work. It should be leveraged as a way to speed up cognition, not a…

ResearchDGX agent

Francois Chollet argues that memorization and knowledge function as a cache for previously completed cognitive work, allowing it to be reused efficiently rather than recomputed. The purpose of storing

11 Apr 2026

Good design is the art of packing 1,000 'hows' into a single 'what'. Good design is compression: making the numerator trend towards infinity…

ResearchDGX agent

Francois Chollet, AI researcher and creator of Keras, argues that good design is fundamentally an act of compression — distilling an enormous number of implementation decisions, constraints, and consi

10 Apr 2026

The power of JAX

HardwareDGX agent

The power of JAX Introducing gyaradax 🐉: A JAX solver for local flux-tube gyrokinetics with custom CUDA kernels for acceleration. This entire code was vibecoded by @ggalletti_ and me in a month. Valid

The reason symmetry is so important in physics is because symmetry is a highly effective compression operator. If a system is invariant unde…

ResearchDGX agent

The reason symmetry is so important in physics is because symmetry is a highly effective compression operator. If a system is invariant under some symmetry, you only need to explain one axis of it. Sc

9 Apr 2026

Science needs a way to process models that are only 'mostly correct' in terms of their predictions, but are very compressive (high ratio bet…

ResearchDGX agent

Science needs a way to process models that are only 'mostly correct' in terms of their predictions, but are very compressive (high ratio between predictive power and model complexity). They are likely

Simplicity is a very strong signal of model quality. Galileo's heliocentric model was directionally correct, but its predictive power was qu…

ResearchDGX agent

Simplicity is a very strong signal of model quality. Galileo's heliocentric model was directionally correct, but its predictive power was quite bad compared to the much older Ptolemaic model, because

We should view the history of physics as a long-running program synthesis task. Kepler and Newton were searching the space of possible symbo…

ResearchDGX agent

We should view the history of physics as a long-running program synthesis task. Kepler and Newton were searching the space of possible symbolic models to find the simplest one that would best satisfy

8 Apr 2026

The new model from Meta is already looking like a disappointment: overoptimized for public benchmark numbers at the detriment of everything …

Model ReleasesDGX agent

The new model from Meta is already looking like a disappointment: overoptimized for public benchmark numbers at the detriment of everything else. Knowing how to evaluate models in a way that correlate

7 Apr 2026

Join the ARC Prize team -- help us build ARC-AGI-4 and ARC-AGI-5

Model ReleasesDGX agent

Join the ARC Prize team -- help us build ARC-AGI-4 and ARC-AGI-5 Platform Engineer - Benchmark Lead ARC Prize Foundation is hiring a senior engineer to build our benchmark platform * Expand ARC-AGI-3

← Previous
123
Next →