Can't pull off a sestina.
This post likely discusses the challenges or difficulties of writing a sestina, the complex poetic form requiring six stanzas with a specific pattern of end-word repetition. Ethan Mollick, known for d
Knowledge catalogue
This post likely discusses the challenges or difficulties of writing a sestina, the complex poetic form requiring six stanzas with a specific pattern of end-word repetition. Ethan Mollick, known for d
I find that open weights models over-perform on benchmarks compared to actual real-world usage, and Kimi feels like no exception. For example, a small amount of use will show that Kimi is not as good
I have been using GPT ImageGen-2 for the past weeks I didn't think that better image-generators would be a big deal but it turns out that there is a quality threshold I didn't expect, where you can no
Kimi 2.6 Thinking seems very good for an open weights model, but many rough edges compared to closed SoTA. The Lem Test resulted in a 74 page thinking trace... and an okay-ish answer. It did an okay T
LLMs are still not consistent judges of qualitative work, and small changes to how that work is presented affect outcomes. Better harnessing and methods (multiple judging runs with randomized orders,
My most popular AI post was a bunch of made-up 'graphs' four years ago. Now, the new GPT-2 image generator does it for real (though not perfect) Here's the famous AI task horizons graph with a touch o
Same prompts as before, but now in GPT image-generator 2, page excerpts from: 'Eldritch Horrors as Pets: A Guide' 'How Womblenauts Work' 'Photographs of the People of New York Who Look Like Birds' 'Ca
Though the images are very good, ChatGPT Image 2.0 does have the typical imagegen problem, which is that editing can be 'stubborn', and attempts to get the AI to change details work well for the first
A useful ward against slop story/science posts on X is noting which is in the character limit. All of the models struggle to do 280 character summaries on their first pass, and most of the people crea
AI reviewers then ranked the submissions, and gave the same ordering every time, regardless of model doing the ranking: Codex GPT-5.4 > GPT-5.3-Codex > Opus 4.6 > humans. Paper: http://claude-code-eco
Classic study gave 146 economist teams the same dataset & got wildly different answers New paper reruns it with agentic AI. Claude Code & Codex land near the human median, but with far tighter dispers
Ethan Mollick shared context about OpenAI's o1-preview model launch, likely discussing the capabilities and implications of this new reasoning-focused AI system that represents a shift toward models d
The second most important release of the LLM era (after GPT-3.5), featuring what was likely the most important chart. Still seems surprising to me that OpenAI told everyone about the biggest advance i
This is also consistent with what OpenAI said at the time! https://x.com/polynoamial/status/2046064264189026587?s=20 @emollick We felt it was the right thing to do. People deserved to know what was co
An obvious way to release Mythos class models with uncertain autonomous ability is to make them only available on the website, like Gemini Deep Think or ChatGPT Pro. Minimal risk of being used for aut
And its not just economists: I am an economic sociologist, there are plenty of us, and management scholars and organizational psychologists, and many other scholars looking at topics related AI & work
And there are many really good AI components at Google: they have top-flight image, music, & video generation, and good UIs for each of them. Google AI studio is probably the best playground for AI ex
I am not convinced that we should be comfortable calling 'problem solving' or 'judgement' or whatever as skills that are impossible for AI to do well. Like any other skill, there are humans who are re
The continuing gap between the capabilities of Gemini Pro 3.1 (very good model) and the capabilities of the Gemini app/website is odd. The model can do what Claude/GPT can do, but there is a minimal h
This paper shows people are asking a lot of medical questions of AI already, but we have little evidence of how good or bad this is. Most of the published research uses old models & compares to doctor
We only have spotty information about this very important topic. It suggests AI can be good at diagnosis, but the real world doesn't always match the experiments. https://x.com/emollick/status/1980474
A major lesson to take away from Opus 4.7 is that, while there is a lot of arguments about implementation choices and personality, models keep improving measurably on economically important tasks with
Anyhow, I realize nobody cares and all the AI labs have started presenting their GDPval-AA score, but it is an incredibly gameable output with low face validity and we really need trustworthy measures
This XKCD comic likely explores themes related to artificial intelligence, technology, or machine learning, presented through the comic's characteristic minimalist style and witty commentary. Ethan Mo
GDPval is one of the most important benchmarks of AI ability because it is based on human expertise. It compares expert human performance to AI performance using expert human judges who spend an avera
I think Artificial Analysis does a good job overall and provides transparency in benchmarking, but GDPval-AA is not a good benchmark and needs to stop being reported. It is Gemini 3.1 judging other mo
Key to note that AI scientists are not experts on labor. Some other economists active on X doing work on AI & labor: @alexolegimas, @danielrock, @joshgans & @robseamans (among many others) But worth n
One of the premier journals in my field... I think there are very valid reasons to set rules on AI in peer review (including disclosure), but the idea that all AI models steal your data is very 2023.
Ethan Mollick observes that AI has fundamentally changed the landscape of human expertise and effort, as AI-generated content now permeates many domains that were previously built exclusively through
This post references a non-exhaustive collection of academics working on a particular topic, with Ethan Mollick directing readers to his previous tweets for additional examples and research. The speci
Tower of Babel is a GitHub repository by researcher Ethan Mollick that likely explores how AI language models handle multilingual tasks and cross-language communication, potentially investigating chal
I was told by Anthropic that they are looking at ways of fixing this, which is good (you can also see a reply from a Claude PM in the thread). I think the adaptive thinking requirement in Claude Opus
I'll give Anthropic credit for moving quickly. Opus 4.7 Adaptive Thinking now triggers thinking much more often, including for the tasks it failed at yesterday. That also means it is doing a lot more
This post likely discusses an AI system that procedurally generates artwork in the style of Pieter Bruegel the Elder, the 16th-century Flemish painter famous for detailed paintings depicting peasant l
Ethan Mollick observes that an AI system (likely Claude or another large language model) still declines to write sestinas, suggesting that certain behavioral constraints or limitations persist despite
We need a new document that AI labs should release with each new model, besides the model card: a sort of changelog I want to see how & in what way the new model changes, breaks, or improves at a rang
With max thinking Opus 4.7 is quite impressive, with a real sense of style In two prompts: 'implement the Tower of Babel, in 3D, in as sophisticated and visually interesting a way as possible. It shou
You can watch the accelerated shipping from the AI labs to get a feeling of what AI-driven product development makes possible. A tremendous number of products are coming out, many of them are really g
A lot of papers coming out are still focused on GPT-4, but you could extrapolate their effects to GPT-5, etc. Much harder to know what the impacts of Claude Code/Codex etc. are because they are so new
A real issue with the current state of our knowledge on the work implications of AI is that there was a genuine discontinuity in AI ability with the rise of practical agentic systems in 2026. We were
Claude remains irreducibly Claude. If you know, you know. (The fact that models have distinct personalities that are consistent across generations is technically interesting, it also makes it very eas
Ethan Mollick reported that requesting Claude Opus 4.7 to write sestinas—a complex poetic form with strict structural requirements—frequently triggers the model's safety guardrails, suggesting the AI
I think the adaptive thinking requirement in Claude Opus 4.7 is bad in the ways that all AI effort routers are bad, but magnified by the fact that there is no manual override like in ChatGPT. It regul
It is not well-explained, but with the adaptive switch off, I get no thinking. I can set thinking levels in Claude Code, but not in Claude Cowork. AI companies keep seeming to assume that coding/techn
Its noticeable how much of the whole practice of working with AI - the prompts, the skill files, the connectors, retrieval work, the markdown files, etc. - is a substitute for the real problem of cont
Maybe because of this paper? https://x.com/emollick/status/1991624198855561508?s=20 Tell all the truth but tell it slant— Success in Circuit lies Too bright for our infirm Delight The Truth's superb s
On the plus side with Opus 4.7, if it does decide to think it produces BY FAR the best Sparks unicorn* ever, even non-thinking is pretty good, if not great. * This is created using TikZ, which is a la
Response. Hey Ethan! Sean here, PM on http://Claude.ai - thanks for the feedback. This isn't a router, this is the model being trained to decide when to think based on the context -- we've been runnin
As someone who teaches at a business school, I can tell you that the pre-professional student population (as opposed to liberal arts or those pursuing academia or a topic of personal excitement) is ex
Compute constraints are a double bind: On the inference side you need to either (a) raise prices, (b) ration use, and/or (c) serve worse models. This hurts current growth On the training side, you can
Researchers are making progress in decoding the communication systems of whales, with recent findings suggesting their vocalizations contain structural elements that function similarly to vowels in hu
Instead of the gold standard, we can imagine an inference standard of exchange, the FLOP. (As opposed to tokens, this accounts for AI ability) With some AI help, I figure 1 buys roughly 10^17 managed-
Markets seem to love this stuff In 2018, Kodak said out of nowhere that it was a crypto company and its stock shot up In 1998, Zapata (an oil firm founded by the Bushes that became a fish meal company
On the quality of the current round of proofs. Paul Erdos had a concept of 'Proofs from The Book', meaning that the argument is so compact and elegant that this is the proof God would've written down
This is becoming a pattern in AI that makes talking about capabilities challenging. First, there are overstated claims (like the flubbed Erdos problems last year), then minor wins (AI helps with disco
Three months ago likely refers to a post by Ethan Mollick, a Wharton professor and AI researcher known for sharing observations about AI progress and capabilities, reflecting on developments or milest
Wish there was information about where this data came from, but this is a very significant change. Since AI use comes from experience, the persistent gender gap in AI use across every study of AI was
AI keeps getting better but the last time the shape of the jagged frontier changed radically was o1 & the Reasoner A good mental model of the coming months is that models get very good at the things t
Ethan Mollick's post addresses the trajectory of AI development, arguing that progress often described as 'gradual' is actually following an exponential curve rather than a linear one. He likely empha
And this is a very generous definition of notable. If we are talking frontier models, only the US and China that are even in the race. And that obscures the fact that the Big Three US labs really do s