Three months ago
Three months ago likely refers to a post by Ethan Mollick, a Wharton professor and AI researcher known for sharing observations about AI progress and capabilities, reflecting on developments or milest
Knowledge catalogue
Three months ago likely refers to a post by Ethan Mollick, a Wharton professor and AI researcher known for sharing observations about AI progress and capabilities, reflecting on developments or milest
Wish there was information about where this data came from, but this is a very significant change. Since AI use comes from experience, the persistent gender gap in AI use across every study of AI was
AI keeps getting better but the last time the shape of the jagged frontier changed radically was o1 & the Reasoner A good mental model of the coming months is that models get very good at the things t
Ethan Mollick's post addresses the trajectory of AI development, arguing that progress often described as 'gradual' is actually following an exponential curve rather than a linear one. He likely empha
And this is a very generous definition of notable. If we are talking frontier models, only the US and China that are even in the race. And that obscures the fact that the Big Three US labs really do s
Given the messy naming scheme used by all the AI companies, I caused a chart to be made showing the gain in GPQA per 0.1 version in model names (estimated, since model names skip version numbers). The
Interesting: 'Currently, 38% of Americans live within 5 miles of at least one operational data center... Living near a data center doesn’t have much of an effect on public opinion about the facilities
It is going to be like what happened in coding: as soon as models crossed a certain threshold (Opus 4.5, GPT-5.2, Gemini 3), suddenly Claude Code & Codex were viable. Before that, it was all about cod
Soon, at each gradual improvement level of AI, you will start to see large discrete jumps in ability in economically important areas, because the previous AI ability level in some aspect of the job bo
The idea of sovereign frontier models is only viable as long as the main Chinese model makers keep shipping open weights models that anyone can build on. I am not sure how long that will continues but
Ethan Mollick discusses the concept of 'emergence' in large language models (LLMs), referring to the phenomenon where AI systems appear to suddenly develop unexpected capabilities as they scale. The p
Version numbers are not a very useful way to understand model ability gains at this stage. Unfortunately that means that if you aren’t following closely, you would expect that 5.4 is a small gain over
This appears to be a whimsical, imaginative post from Ethan Mollick on X (formerly Twitter) framing an evolutionary history scenario as a fictional mech battle, pitting a machine piloted by Tiktaalik
Ethan Mollick, a prominent researcher and commentator on AI, expresses cautious optimism about the trajectory of AI-powered coding tools, acknowledging a scenario where increasingly capable coding ass
As I look at some of the replies (which are startlingly human, whats up with X?), it appears a lot of the concern is that Mythos was presented as some magic leap. I didn't read Anthropic's announcemen
At this point, I assume that every internal memo at an AI lab is just written for public release. The labs are certainly capable of keeping secrets that don’t leak, so they must realize that all-hands
This post by Ethan Mollick likely explores AI-generated or conceptual visualizations of battles between single-celled and multicellular organisms framed as mech combat, drawing on biological history f
Ethan Mollick shared a post on X expressing frustration that a piece of content—likely a documentary or documentary-style video—was automatically or incorrectly labeled as 'made with AI.' The post hig
Folks, this is not Jevon's Paradox. this is just normal supply and demand. It turns out the utility of AI is high enough that people have high demand, which is outstripping supply (so prices will go u
General Purpose Technologies have downstream effects everywhere, good and bad. Those outcomes can be mitigated, or encouraged, with the right sorts of policy choices If the only options are being for
I am catching glimpses in my feed that there is a backlash against Mythos as 'marketing hype,' and it is a little confusing. I don't think anyone who has used the latest agentic coding tools, would th
Ethan Mollick shares an example of Seedance 2.0, a video generation AI model, demonstrating impressive creative and historical scene generation capabilities by producing a depiction of a mech battle b
Ethan Mollick, a Wharton professor and prominent AI researcher, posted commentary on what appears to be a notable AI capability advancement, framing it as a predictable progression along established p
Six months ago, there was a lot of focus on the idea that the there would be a massive glut of unused computing power which would could a recession as AI use plateaued. The 'compute bubble' belief was
So the concern over Mythos and cybersecurity seems warranted. We conducted cyber evaluations of Claude Mythos Preview and found that it is the first model to complete an AISI cyber range end-to-end. 🧵
Ethan Mollick discusses the economic and financial risks surrounding data center construction, distinguishing between the risks inherent in financing methods and the broader arguments driving investme
The trend of treating all of AI as One Big Thing that always includes data centers & job changes & education changes & power & accelerating science & misinformation & national security & corporate con
According to Waymo's published data, their technology is preventing injuries & deaths. My view is that if this is true, and I have yet to see a debunking of their data, then we safety advocates should
And if you are interested in the answer, Claude and ChatGPT 5.4 Pro think running OpenClaw with local inference on your Mac uses more total power. Gemini doesn't think so (but appears to have not thou
Currently, ChatGPT has the best way of viewing thinking traces, a short summary of steps in the main window, and a detailed audit in the sidebar if you want it Claude does almost as well, but more sum
From this perspective, Gemini is also worse than the original Bard. Sydney was the original sin of LLMs anthropormism, but also got the idea that AIs can sometimes be better with personalities right.
I think Muse Spark came in far better than most were expecting as the first new model attempt from Meta, especially given the fact that it has been a year since Llama 4 with no models at all (and that
It is notable that we are all debating exactly which markdown files are most important to feed AI (skills, memory, tool instructions) and in which order to feed them to get the best output. Feels that
Ethan Mollick shares an assessment of an AI model (likely a newer or smaller model) that, while not matching the performance of the leading frontier models (likely referring to top offerings from Open
OpenAI should probably bite the bullet and just name their next set of models something more human sounding. Everyone anthropomorphizes their AIs anyway, and 'Claude' is an easier name to refer to tha
Really interesting ideas are going to be increasingly at a premium as the cost of executing those ideas drops. (Our research and others shows AI is quite good at generating interesting ideas, but not
(thinking traces are always summaries of the internal reasoning, since the actual internal chain of thought is a trade secret, but that doesn't mean they aren't useful when summarized, especially when
A lot of our education on writing well focuses on logic, clarity, and argument. AI will force us to think more about style. The boredom that comes from everything on the internet reading Claude-y now,
Chiasmus (reversing grammatical structures in two sentences for drama). Asyndetic tricolon (three items listed without a conjunction). Parataxis (short and somewhat disconnected dramatic sentences). S
Ethan Mollick shared a humorous AI-generated image that 'fixes' Goya's famous dark painting *Saturn Devouring His Son* by replacing the disturbing scene with Saturn happily eating a hot dog instead. T
I am not just talking about the em-dashes and phrasing ('doing the heavy lifting'). Those are more fixable than style, which will always default to an LLM-normal even if that normal changes with new L
Neat experiment finds AI fact checks are rated as more helpful & less ideological than human ones 'LLM-generated Community Notes can achieve broader cross-ideological acceptance than human-written not
I was unable to retrieve the specific content of the arXiv paper `2604.02592` or the specific tweet referenced. The search did not return the correct paper (the arXiv ID `2604.02592` as a 2026 pape...
Tired of the old version of Michelangelo's 'Creation of Man' just sitting there on the ceiling of the Sistine Chapel and not even moving or doing anything cool? Don't worry, its been updated while sta
A post by Ethan Mollick (@emollick) on X highlights an AI-generated video reimagining of William Blake's *The Ancient of Days* (1794) — a design originally published as the frontispiece to *Europe...
AI finally lets us see Raphael's The School of Athens the way Raphael obviously intended it, illustrating the delicate dance and subtle conflicts between Plato and Artistotle. (Seedance 2.0 is very fu
All is not lost. Duckerton is still possible. Here is Seedance 2.0 with the same prompt. Media My most popular Sora video was “an Elaborate regency romance where everyone is wearing a live duck for a
Our Lab just posted a new research report from Zimran Ahmed about how the game industry is adapting to AI. He spoke to people at 20 different studios and found a wide range of approaches to adapt (or
The pace at which useful things are shipping also seems to be accelerating. Model releases are coming faster, of course, but so are significant application and enterprise products (especially from Ant
Things that make the jagged intelligence of AI harder to deal with than the jaggedness of humans: 1) Weaknesses are not always intuitive or identifiable in advanced 2) All LLMs have similar weaknesses
After playing with it a bit, Meta's Muse Spark Thinking is fine so far, but really doesn't match the current Big Three models. It also is a bit... weird. Like some strange language & tone, a little lo
AI is jagged, but I think sometimes it is easy to overly focus on that. The generalness is a surprise too! LLMs may be optimized for verifiable fields like coding, but AI is also not bad at corporate
And all the evidence is is that is that models are getting better all this other stuff at the same time as they are improving in coding. More recent models are more creative, for example. Still plenty
Anyhow, its not bad. Just not the vibe level that the benchmarks might indicate. And, for a first re-entry into the frontier model space, given the engineering efficiencies they achieved, it feels lik
I was unable to retrieve the specific content of the referenced X (Twitter) post (status ID 2042355048311575010), as it is not accessible through web search results and X/Twitter requires a login t...
Meh on the Lem Test as well https://x.com/emollick/status/2019478314767835634?s=20 Opus 4.6 saturates my Lem test, which I've done since GPT-3.5 SciFi author Stanislaw Lem wrote of two rival construct
One fun thing about AI is that it lets you play with interfaces and approaches to displaying information in new ways without a lot of effort. I got a an internet connected e-ink display and set it up
So we now have a pretty good picture of the state of the frontier AI model makers. US closed source models continue to lead. Google, OpenAI, and Anthropic stand well ahead of the pack, and may have si
Amazon released the Nova 2 model family on December 2, 2025 at AWS re:Invent, comprising Nova 2 Lite, Nova 2 Pro (Preview), Nova 2 Omni, and Nova 2 Sonic, all featuring a 1-million-token context wi...
Sparks unicorn https://x.com/emollick/status/2024756029121020236?s=20 Here is the Gemini 3.1 'Sparks unicorn' (This is created using TikZ, which is a language built for scientific diagrams & very much