AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
Human
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

Search: “ethan-mollick--x”

GridTimelineEvolution
61+ results
12 Aug 2026

Interesting research suggests caution in determining which AI company is winning by looking at any one source.. OpenRouter seems to show ope…

Model ReleasesDGX agent

Interesting research suggests caution in determining which AI company is winning by looking at any one source.. OpenRouter seems to show open weights winning over time, but work submitted to Pangram i

10 Aug 2026

Oh no, we aren’t going to go back to this sort of prompting again, are we? I would love Anthropic to test if it actually works robustly, bec…

Model ReleasesDGX agent

Oh no, we aren’t going to go back to this sort of prompting again, are we? I would love Anthropic to test if it actually works robustly, because our experiments (with slightly older models) found it d

9 Aug 2026
735 results
DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

A failure of ChatGPT Work & Claude Cowork is they assume that non-coders couldn't understand how to think about problems like a coder, so th…

Model ReleasesDGX agent

A failure of ChatGPT Work & Claude Cowork is they assume that non-coders couldn't understand how to think about problems like a coder, so they hide all that stuff. They should instead explain choices

8 Aug 2026

You may have been told to watch this video about the OpenAI AI hack. You really should, even if you don't usually care about tech stuff. If …

ApplicationsDGX agent

You may have been told to watch this video about the OpenAI AI hack. You really should, even if you don't usually care about tech stuff. If nothing else, click this link to the 18 minutes in & see how

7 Aug 2026

Basically every remaining good AI benchmark score has an implied asterisk next to it which reads: * could be signficantly higher with a bett…

Model ReleasesDGX agent

On August 7, 2026 Ethan Mollick tweeted that “every remaining good AI benchmark score has an implied asterisk next to it which reads: * could be significantly higher with a better harness.” The commen

6 Aug 2026

It is past time to take AI & security seriously at the individual level as well. If its not the current OpenAI and Anthropic models doing it…

ApplicationsDGX agent

It is past time to take AI & security seriously at the individual level as well. If its not the current OpenAI and Anthropic models doing it, then the coming open weights models will when they catch u

This definitely seems like something worth noting, and illustrates the gap between Fable/Astra class models and the previous frontier that w…

ApplicationsDGX agent

This definitely seems like something worth noting, and illustrates the gap between Fable/Astra class models and the previous frontier that was 'merely' good at hacking under human instructions. Initia

This paper by researchers from MIT and Stanford finds that most people would be financially better off if they followed the financial advice…

Model ReleasesDGX agent

This paper by researchers from MIT and Stanford finds that most people would be financially better off if they followed the financial advice of LLMs (GPT-5.2 & Gemini 3 Flash) But some people get a bi

5 Aug 2026

Also I think AISI is a great model of a government agency tasked with AI security. They have open benchmarks, very fast testing, and clear c…

ApplicationsDGX agent

Also I think AISI is a great model of a government agency tasked with AI security. They have open benchmarks, very fast testing, and clear communication about incidents that is neither hyped up nor hi

Yes, the AIs were given a cybersecurity challenge, with internet access enabled and safety filters disabled. But the extent to which Mythos …

SafetyDGX agent

Yes, the AIs were given a cybersecurity challenge, with internet access enabled and safety filters disabled. But the extent to which Mythos 5 pursued its mission (fake identities, social engineering,

4 Aug 2026

Given today, it is surprising how daring Microsoft & Google were initially with AI. Microsoft released GPT-4 before OpenAI, didn't back down…

Model ReleasesDGX agent

Given today, it is surprising how daring Microsoft & Google were initially with AI. Microsoft released GPT-4 before OpenAI, didn't back down after Sydney & got Copilot to market quickly (the 1st profe

3 Aug 2026

Nate is right (full context can lead to a degradation of chats in several different ways) but you can't ask the AI about stuff like this as …

Model ReleasesDGX agent

Nate is right (full context can lead to a degradation of chats in several different ways) but you can't ask the AI about stuff like this as they have bad self-knowledge A good approach is to compact t

1 Aug 2026

Some things to note: 1) AI is getting very good at math and science. 2) Two years ago LLMs could not do basic math consistently 3) According…

ApplicationsDGX agent

Some things to note: 1) AI is getting very good at math and science. 2) Two years ago LLMs could not do basic math consistently 3) According to @polynoamial this cost less than $2000 in current API co

31 Jul 2026

One of our big findings in our study at Procter and Gamble was that AI blurred the lines between jobs. Now OpenAI has a similar finding. Org…

ApplicationsDGX agent

One of our big findings in our study at Procter and Gamble was that AI blurred the lines between jobs. Now OpenAI has a similar finding. Organizational boundaries are becoming porous, the walls thinni

This is both a real incident (in that the AI really did get unauthorized access to real systems) and also something it was (sort of) prompte…

Model ReleasesDGX agent

This is both a real incident (in that the AI really did get unauthorized access to real systems) and also something it was (sort of) prompted to do. In a review of our cybersecurity evaluations, we fo

This is not optional, things are getting chaotic now, and not dealing with this change won't make it go away. Plus, this could be a huge boo…

ApplicationsDGX agent

This is not optional, things are getting chaotic now, and not dealing with this change won't make it go away. Plus, this could be a huge boost for both individual satisfaction & firm performance if do

30 Jul 2026

Answer from OpenAI

ApplicationsDGX agent

Answer from OpenAI @emollick re ARC-AGI-3: human testers scored ~48% (ARC uploaded testers logs to HuggingFace some time ago, I believe). re GDPval: it’s close to saturated now, so we’re mostly lookin

As the benchmarks that test frontier AI on get more complex, we are losing one of the most important aspects of benchmarking: comparisons to…

ApplicationsDGX agent

As the benchmarks that test frontier AI on get more complex, we are losing one of the most important aspects of benchmarking: comparisons to humans Validated benchmarks need to have human (ideally mul

Big unsaturated benchmarks that have this: ARC-AGI, the original GDPval (not GDPval-AA), METR long horizons, ASI cyber tasks, (Speaking of w…

ApplicationsDGX agent

Ethan Mollick highlights that large, currently under‑explored benchmarks (e.g., ARC‑AGI, GDPval, METR long horizons, ASI cyber tasks) are growing in complexity, yet they increasingly lack systematic h

29 Jul 2026

Model + harness. We have barely begun to understand the best ways to do harness engineering. A huge amount of untapped potential even withou…

Model ReleasesDGX agent

Model + harness. We have barely begun to understand the best ways to do harness engineering. A huge amount of untapped potential even without models getting better (but models are getting better) Turn

28 Jul 2026

There is definitely real jaggedness both among fields (writing is an area where models improve slowly if at all) and within them, but the ma…

ApplicationsDGX agent

There is definitely real jaggedness both among fields (writing is an area where models improve slowly if at all) and within them, but the magic of LLMs is that they are so unreasonably effective acros

There's been a talk about how LLMs are only advancing in verifiable areas like math or coding, but that isn't what the data suggests. As mod…

ApplicationsDGX agent

There's been a talk about how LLMs are only advancing in verifiable areas like math or coding, but that isn't what the data suggests. As models have gotten better at that, they are also better at solv

27 Jul 2026

We are in a world where you can create truly unique, visually interesting and creative playable demos on demand with the current capabilitie…

Model ReleasesDGX agent

We are in a world where you can create truly unique, visually interesting and creative playable demos on demand with the current capabilities of Codex and Claude Code. We don't need to keep cloning th

25 Jul 2026

Ha! It did it: 'We introduce BenchBenchBenchBenchBench (BBBBB), an executable benchmark of AI-authored conformance suites for benchmark-eval…

Model ReleasesDGX agent

Ha! It did it: 'We introduce BenchBenchBenchBenchBench (BBBBB), an executable benchmark of AI-authored conformance suites for benchmark-evaluation metrics' I really thought it would treat 'now do benc

24 Jul 2026

As a joke I prompted Codex 'Build and run BenchBench, a benchmark of now good ai is at creating benchmarks. then figure out what benchbenchb…

Model ReleasesDGX agent

As a joke I prompted Codex 'Build and run BenchBench, a benchmark of now good ai is at creating benchmarks. then figure out what benchbenchbench is and run that. and then write benchbenchbench up as a

Glad to see Google sharing data on how Gemini is being used. Especially interesting is that the usefulness of multimodal AI for manual labor…

Model ReleasesDGX agent

Glad to see Google sharing data on how Gemini is being used. Especially interesting is that the usefulness of multimodal AI for manual labor may be greater than expected. https://blog.google/innovatio

The GPT-5.x Pro series has remained the best models for hard technical problems since they launched. There is some parallel model magic goin…

Model ReleasesDGX agent

The GPT-5.x Pro series has remained the best models for hard technical problems since they launched. There is some parallel model magic going on that is not well-explained. Anthropic has never had an

This is a big jump in ARC-AGI-3.

Model ReleasesDGX agent

This is a big jump in ARC-AGI-3. Claude Opus 5 from @AnthropicAI is the new SOTA on ARC-AGI-3: 30.2% The previous high score (7.8%) was set by GPT-5.6 Sol (Max) Throughout our analysis, we observed no

23 Jul 2026

I wrote the latest of my occasional guides to which AI to use right now for non-experts who want to get stuff done. The agentic systems avai…

AgentsDGX agent

I wrote the latest of my occasional guides to which AI to use right now for non-experts who want to get stuff done. The agentic systems available to everyone are getting extremely powerful (even as th

22 Jul 2026

And yes, the models were following instructions, they just did so in clever ways. Some nice additional info here

ApplicationsDGX agent

And yes, the models were following instructions, they just did so in clever ways. Some nice additional info here A few thoughts on the Hugging Face hack: - This is, to my knowledge, the *third* disclo

“Generate a fake, but believable, witty Churchill insult at a party and explain the context. It should be very clever and original” This tim…

Model ReleasesDGX agent

“Generate a fake, but believable, witty Churchill insult at a party and explain the context. It should be very clever and original” This time, I think GPT 5.6 Sol Pro wins, but Fable is good too, and

Who would you give authorship to? The person who wrote 58 words of prompts, or GPT-5.6 Pro?

Model ReleasesDGX agent

Who would you give authorship to? The person who wrote 58 words of prompts, or GPT-5.6 Pro? Dinitz-Garg-Goemans conjecture is false. This graph theory problem was open for ~30 years. The graph below h

21 Jul 2026

And now from the Chinese government side. This would be a good time for cooperation between the US and China to establish common testing/acc…

SafetyDGX agent

And now from the Chinese government side. This would be a good time for cooperation between the US and China to establish common testing/acceptance standards for new models, so at least the safety cer

“Geologists communicated in English; and they could name things in a manner that sent shivers through the bones' Not a whiff of LLM feel, de…

ApplicationsDGX agent

“Geologists communicated in English; and they could name things in a manner that sent shivers through the bones' Not a whiff of LLM feel, despite the em-dashes & period-separated lists. We focus too m

Previously, these AI hacking stories were about breaches in test environments, where any question of AI breaching security was purely theore…

ApplicationsDGX agent

Previously, these AI hacking stories were about breaches in test environments, where any question of AI breaching security was purely theoretical. This is something else. https://openai.com/index/hugg

20 Jul 2026

Folks, Gemma is a good model & very useful for some purposes, it is not anywhere near the frontier. Inkling is an interesting model, it is n…

Model ReleasesDGX agent

Folks, Gemma is a good model & very useful for some purposes, it is not anywhere near the frontier. Inkling is an interesting model, it is nowhere near the frontier. This is pretty obvious, you can lo

15 Jul 2026

I demonstrated the incredible power of o1-preview/ reasoning less than two years ago by showing it could solve this crossword puzzle with on…

Model ReleasesDGX agent

I demonstrated the incredible power of o1-preview/ reasoning less than two years ago by showing it could solve this crossword puzzle with only one hint. https://www.oneusefulthing.org/p/something-new-

14 Jul 2026

Fable, turn my tweet into a thinkpiece (this was pretty funny): There has never been a better time to have opinions about artificial intelli…

Model ReleasesDGX agent

Fable, turn my tweet into a thinkpiece (this was pretty funny): There has never been a better time to have opinions about artificial intelligence. I say this with some authority, because I am currentl

If you are use the Claude everything app, you pick between Home and Code. If you pick Home you get to pick between Chat & Cowork If you use …

Model ReleasesDGX agent

If you are use the Claude everything app, you pick between Home and Code. If you pick Home you get to pick between Chat & Cowork If you use the OpenAI everything app, you pick between ChatGPT Work & C

13 Jul 2026

Computer use in Codex got very good on PC. Asking it to do something on your computer and having the cursor move under the control of a ghos…

Model ReleasesDGX agent

Computer use in Codex got very good on PC. Asking it to do something on your computer and having the cursor move under the control of a ghost is one of the things that makes you viscerally realize how

I guess image input is the big capability of the models, and tool use can be a substitute for non-omni model output. Still, multimodal voice…

AgentsDGX agent

Ethan Mollick notes that image input represents the primary advanced capability of current AI models, and that tool‑use can effectively replace outputs from non‑omni models. He observes that multimoda

I think OpenRouter is not a good measure of actual model usage in a world of agentic tools (not that I doubt that Chinese open weights model…

AgentsDGX agent

I think OpenRouter is not a good measure of actual model usage in a world of agentic tools (not that I doubt that Chinese open weights model usage is up, but this could also look like a graph of usage

One thing I am kind of surprised by is that full multi-modal (any-any) models have not become a bigger deal. It seems Google is the only Lab…

ApplicationsDGX agent

One thing I am kind of surprised by is that full multi-modal (any-any) models have not become a bigger deal. It seems Google is the only Lab releasing these, OpenAI uses selective multimodal capabilit

11 Jul 2026

ChatGPT still has study mode, but rather than /study you now have to type @ study It makes the AI act more like a tutor than a helpful assis…

Model ReleasesDGX agent

ChatGPT still has study mode, but rather than /study you now have to type @ study It makes the AI act more like a tutor than a helpful assistant, and some work suggests it is better if you are trying

I gave Fable the code: 'take this game and do something incredible with it to make it something very different. Be creative' It created DEEP…

Model ReleasesDGX agent

I gave Fable the code: 'take this game and do something incredible with it to make it something very different. Be creative' It created DEEP TIME: create a city, watch it be abandoned and forgotten, a

I really want to get artisanal with my model selection: Luna xHigh for recipes Sol Medium for sonnets Terra High for bad puns Luna xLow for …

ApplicationsDGX agent

I really want to get artisanal with my model selection: Luna xHigh for recipes Sol Medium for sonnets Terra High for bad puns Luna xLow for limericks Sol Utra for true prophecies Sol Low for finding o

Thats good for the dumb reason that I am burning tokens having Sol in Codex playing Slay the Spire 2’s daily challenge (so randomized rules,…

ApplicationsDGX agent

Thats good for the dumb reason that I am burning tokens having Sol in Codex playing Slay the Spire 2’s daily challenge (so randomized rules, today it is Ascension 3 Defect with hoarder) to see if it c

To be clear, NotebookLM has its own issues, and is built for a specific use case (research and analysis of sources) but it is an example of …

ApplicationsDGX agent

To be clear, NotebookLM has its own issues, and is built for a specific use case (research and analysis of sources) but it is an example of how a UX might actually operate that treats knowledge work s

10 Jul 2026

And then what is the point of Work? Just a dumbed-down version of Codex that secretly does coding but hides it? (I can understand, maybe, if…

ApplicationsDGX agent

And then what is the point of Work? Just a dumbed-down version of Codex that secretly does coding but hides it? (I can understand, maybe, if the release was a 1st step towards something but there is n

Fable: 'an 8th grader puts together a powerpoint about the Great Gatsby but obviously did not read the Great Gatsby. Show me that powerpoint…

ApplicationsDGX agent

Fable: 'an 8th grader puts together a powerpoint about the Great Gatsby but obviously did not read the Great Gatsby. Show me that powerpoint! ;)' This was actually pretty funny. Even the font choices

For the first time, the personalities and approaches of the leading models are diverging in significant ways, magnified by the fact that ove…

ApplicationsDGX agent

For the first time, the personalities and approaches of the leading models are diverging in significant ways, magnified by the fact that over longer task horizons these differences in judgement & appr

I also like the cyberpunk and totalitarian visualization modes. There is also an achievement-based unlock system in there, traffic & usage, …

ApplicationsDGX agent

I also like the cyberpunk and totalitarian visualization modes. There is also an achievement-based unlock system in there, traffic & usage, day & night, and a bunch of other stuff. If it runs slow, yo

If you want to edit it yourself, here it is: https://github.com/emollick/monument-brutalist-city-builder

ApplicationsDGX agent

This repository contains code for a brutalist city builder game or simulation, made available for anyone who wants to edit or modify it themselves. The project appears to be shared by Ethan Mollick on

Incredibly annoying when Fable has a forbidden thought in the middle of a long-running project and kills it. Apparently this page of referen…

ApplicationsDGX agent

Incredibly annoying when Fable has a forbidden thought in the middle of a long-running project and kills it. Apparently this page of references in one of my papers makes Fable wonder about something t

I've been going on about how ChatGPT Work (and Cowork) are missed opportunities for knowledge workers, and to illustrate that take a look at…

ApplicationsDGX agent

I've been going on about how ChatGPT Work (and Cowork) are missed opportunities for knowledge workers, and to illustrate that take a look at Google's NotebookLM answering the same question as ChatGPT

See also this: https://x.com/emollick/status/2068729258176819253?s=20

ApplicationsDGX agent

See also this: https://x.com/emollick/status/2068729258176819253?s=20 A fundamental problem with extending Codex/Cowork/Code to all knowledge work is that they remain very 'software-brained' where the

The Labs should launch some easy-to-view comparisons of their new models: show us what a piece of work looks like from Luna xHigh versus Ter…

ApplicationsDGX agent

The Labs should launch some easy-to-view comparisons of their new models: show us what a piece of work looks like from Luna xHigh versus Terra Medium versus Sol Low (or similar for Haiku, Sonnet, Opus

The new ChatGPT voice is quite impressive to use, really worth a minute to try it out on your phone. (while staying aware that the voice mod…

ApplicationsDGX agent

Ethan Mollick recommends trying ChatGPT's new voice feature on mobile devices, noting it delivers an impressive user experience. He advises users to test the feature while remaining mindful of potenti

This time it is novel math proofs with a public model (most of the other big math breakthroughs have been with experimental LLMs).

Model ReleasesDGX agent

This time it is novel math proofs with a public model (most of the other big math breakthroughs have been with experimental LLMs). Yesterday, we made GPT-5.6 Sol Ultra generally available. Today, we'r

This was a critical early paper on AI & work, showing that entrepreneurs getting advice from GPT-4 had higher profit margins if they were hi…

Model ReleasesDGX agent

This was a critical early paper on AI & work, showing that entrepreneurs getting advice from GPT-4 had higher profit margins if they were high performing, but did worse if they were already in trouble

← Previous
1
Next →
← Previous
123…13
Next →