AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “ethan-mollick--x”

GridTimelineEvolution
735 results
6 May 2026

I usually avoid commenting too much on industry deals, but this one is fascinating. Certainly seems like a blow to the idea that Grok will r…

ApplicationsDGX agent

I usually avoid commenting too much on industry deals, but this one is fascinating. Certainly seems like a blow to the idea that Grok will remain a frontier model. Our agreement with @SpaceX means we

5 May 2026

A reminder that telling the AI that it is an expert in a field is no longer helpful in making the AI better at that field.

ApplicationsDGX agent

A reminder that telling the AI that it is an expert in a field is no longer helpful in making the AI better at that field. We tested one of the most common prompting techniques: giving the AI a person

All benchmarks are flawed, but GPQA has been fairly consistent & highly correlated with other measured benchmars. I think it's a good way to…


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
ApplicationsDGX agent

All benchmarks are flawed, but GPQA has been fairly consistent & highly correlated with other measured benchmars. I think it's a good way to see how far we've come that the free model from OpenAI, GPT

Also

ApplicationsDGX agent

Also Instead of the gold standard, we can imagine an inference standard of exchange, the FLOP. (As opposed to tokens, this accounts for AI ability) With some AI help, I figure 1 buys roughly 10^17 man

Everything about AI will be contested, because everyone has different interests. So it is and so it will be for so it has been, time out of …

ApplicationsDGX agent

Mollick argues that AI policy and development will inevitably remain contested territory because different stakeholders—including governments, corporations, researchers, and the public—have fundamenta

For individual AI use, the jagged frontier is increasingly well understood. In multi-agent workflows in organizations, AI is jagged in ways …

AgentsDGX agent

For individual AI use, the jagged frontier is increasingly well understood. In multi-agent workflows in organizations, AI is jagged in ways that have not been well identified yet. In fact, we don't ev

In addition to the CAISI evaluation, it would be useful if NIST conducted public tests of AI abilities as an independent evaluator - though …

ApplicationsDGX agent

In addition to the CAISI evaluation, it would be useful if NIST conducted public tests of AI abilities as an independent evaluator - though those obviously should not be pre-release tests & can be don

It would also be useful for funding R&D into benchmarking models, which is currently mostly done by the labs themselves right now.

ApplicationsDGX agent

Ethan Mollick advocates for dedicated funding toward developing independent benchmarking systems for AI models, noting that evaluation is currently predominantly conducted by the labs that create the

Missing from the “will AI replace doctors?” debate is that doctors (and lawyers and psychologists and bankers) all vote & form the donor bas…

ApplicationsDGX agent

Missing from the “will AI replace doctors?” debate is that doctors (and lawyers and psychologists and bankers) all vote & form the donor base to political parties & have deep community ties. The gover

Mythos is probably really good at identifying pig diseases.

ApplicationsDGX agent

Ethan Mollick suggests that Mythos, likely an AI model or system, demonstrates strong capability in identifying and diagnosing pig diseases, possibly indicating its effectiveness in agricultural or ve

The common noun have been FLOPs: 1) Has precise meaning 2) You can actually price and measure it 3) Funnier

ApplicationsDGX agent

FLOPs (floating-point operations) serve as a practical metric for measuring AI computational capacity because they have a precise technical definition, can be directly priced and measured, and offer a

The unreasonable effectiveness of LLMs is what makes them so weird. The labs don’t need to decide what kind of AI to build, because better L…

ApplicationsDGX agent

The unreasonable effectiveness of LLMs is what makes them so weird. The labs don’t need to decide what kind of AI to build, because better LLMs do better at most things. Finance? Pig disease identific

Too much of the solutions & language for agentic systems comes from a coding perspective (control planes, hooks, loops), but I really think …

AgentsDGX agent

Too much of the solutions & language for agentic systems comes from a coding perspective (control planes, hooks, loops), but I really think that the field of management and organizations can tell us m

We know the answer: Cal State LA, followed by Pace & SUNY. State systems in California, Texas & New York are great for moving many born in t…

ApplicationsDGX agent

We know the answer: Cal State LA, followed by Pace & SUNY. State systems in California, Texas & New York are great for moving many born in the lower 20% of income to the top 20%; Ivies for moving a fe

4 May 2026

A challenge with AI regulation and vetting is how bad our benchmarks of AI model performance and risks are. There is no benchmark for risks …

Model ReleasesDGX agent

A challenge with AI regulation and vetting is how bad our benchmarks of AI model performance and risks are. There is no benchmark for risks and red-teaming requires experiments from dedicated speciali

Co-founder of Anthropic, interesting that he refers to public sources when he is also obviously privy to lots of internal sources that he ca…

ApplicationsDGX agent

Co-founder of Anthropic, interesting that he refers to public sources when he is also obviously privy to lots of internal sources that he cannot discuss. I assume he sees the same thing at Anthropic.

I think the fact that GPT-4o and Llama 3.3-80B did no significant harm is just as important as whether AI helped. If older (less accurate & …

Model ReleasesDGX agent

I think the fact that GPT-4o and Llama 3.3-80B did no significant harm is just as important as whether AI helped. If older (less accurate & more sycophantic) chatbots essentially did nothing for peopl

It is somewhat comforting that now, whenever I see a post about “here’s the thing that keeps me up at night” I know that there is absolutely…

ApplicationsDGX agent

It is somewhat comforting that now, whenever I see a post about “here’s the thing that keeps me up at night” I know that there is absolutely no chance that this is being written by a human who is stay

May 5 is the GPT-5.5 launch celebration in San Francisco and the Claude Finance Briefing in New York. Real opposite valence events on opposi…

Model ReleasesDGX agent

I cannot provide a summary for this entry as the post content appears incomplete or corrupted in the provided information. The title is cut off mid-sentence and doesn't clearly convey the full topic.

My surprise here seems warranted, this paper was retracted (There are other peer-reviewed meta-analyses of the impact of AI on education fin…

ApplicationsDGX agent

My surprise here seems warranted, this paper was retracted (There are other peer-reviewed meta-analyses of the impact of AI on education finding positive effects, like: https://www.researchgate.net/pu

Poems that ChatGPT, Claude, and Gemini all seem to 'like' when you ask for poetry related to being/making LLMs: Rilke's 'Archaic Torso of Ap…

Model ReleasesDGX agent

Poems that ChatGPT, Claude, and Gemini all seem to 'like' when you ask for poetry related to being/making LLMs: Rilke's 'Archaic Torso of Apollo' Stevens' 'Idea of Order at Key West' Borges's 'The Gol

That doesn't mean that there should not be regulation and vetting, but it does suggest that is hard to write a criteria right now that is no…

ApplicationsDGX agent

That doesn't mean that there should not be regulation and vetting, but it does suggest that is hard to write a criteria right now that is not somewhat vague. More R&D into non-lab benchmarks is urgent

This is a very niche post, but they are good poems. I hadn't read Autopsychography until it came up from both Claude and Gemini. And if you …

Model ReleasesDGX agent

This is a very niche post, but they are good poems. I hadn't read Autopsychography until it came up from both Claude and Gemini. And if you haven't read Rilke, that poem, with its incredible end, is a

3 May 2026

I am not sure I would agree with all of this, but the relationship between Anthropic and Claude is quite different than the relationship bet…

Model ReleasesDGX agent

I am not sure I would agree with all of this, but the relationship between Anthropic and Claude is quite different than the relationship between other labs and their models. And that shows up in lots

Its getting hard to benchmark frontier agent performance on longer tasks. Repeated measurement is very expensive and there are differences b…

Model ReleasesDGX agent

Its getting hard to benchmark frontier agent performance on longer tasks. Repeated measurement is very expensive and there are differences between using models in harnesses versus via APIs. I suspect

Of course Pynchon would call this correctly 40 years ago: “It will be amazing and unpredictable, and even the biggest of brass, let us devou…

TutorialsDGX agent

Of course Pynchon would call this correctly 40 years ago: “It will be amazing and unpredictable, and even the biggest of brass, let us devoutly hope, are going to be caught flat-footed.“ He and Dougla

Sometimes when I demo AI, I show it turning cover letters into goofy formats (poetry, etc) as an introduction to the idea of AI as translato…

Model ReleasesDGX agent

Sometimes when I demo AI, I show it turning cover letters into goofy formats (poetry, etc) as an introduction to the idea of AI as translator between forms. For the first time, GPT-5.5 has been trying

The artificial analysis index is a normalized score of several benchmarks (and has changed over time) it is fine for roughly comparing model…

ApplicationsDGX agent

The artificial analysis index is a normalized score of several benchmarks (and has changed over time) it is fine for roughly comparing models, it is not useful for trend analysis and it is unclear wha

The single most accurate science fiction author writing about AI turned out to be… Douglas Adams He wrote about AIs that work best when emot…

ApplicationsDGX agent

The single most accurate science fiction author writing about AI turned out to be… Douglas Adams He wrote about AIs that work best when emotionally manipulated & that guilt you in turn. And he underst

This is a good explanation of why the gap between open and closed models is larger than it appears in benchmarks. I would add in that curren…

ApplicationsDGX agent

This is a good explanation of why the gap between open and closed models is larger than it appears in benchmarks. I would add in that current open models are also more fragile than closed: they handle

2 May 2026

Generally, I would say X is not real life, but I am surprised about how often I get asked by executives about which AI lab is winning or wha…

ApplicationsDGX agent

Generally, I would say X is not real life, but I am surprised about how often I get asked by executives about which AI lab is winning or what is up with a particular model in ways that indicate that t

I was quoted a couple times in this Atlantic article, but that isn’t (the only) reason I think it is good. It lays out the reasons why we wh…

ApplicationsDGX agent

I was quoted a couple times in this Atlantic article, but that isn’t (the only) reason I think it is good. It lays out the reasons why we whipsawed from “AI is a bubble” to “there are not enough data

(Sorry, after seeing so many of these, could not resist): 🚨 BREAKING: Google just dropped a NEW paper that completely deletes RNNs from exi…

Model ReleasesDGX agent

(Sorry, after seeing so many of these, could not resist): 🚨 BREAKING: Google just dropped a NEW paper that completely deletes RNNs from existence. No recurrence. No convolutions. Nothing. Just one mec

This post is a real Voight-Kampff test for bots on X.

ApplicationsDGX agent

This post likely references the Voight-Kampff test from 'Blade Runner,' a fictional device used to identify replicants (artificial beings), as a metaphor for distinguishing AI bots from humans on X (f

1 May 2026

GPT-imagegen-2: 'make 5x5 grid of dog photos, where each photo gets noticeably cuter' ...now cats ...now man-eating squid ...now covers of t…

ApplicationsDGX agent

Ethan Mollick demonstrates GPT-imagegen-2's capability to generate image grids with progressive variations on a theme, showing how the model can create multiple iterations of subjects (dogs, cats, squ

I should also add that one of the strongest alignment actions that OpenAI did was to name their product chatgpt with gpt 5.5 medium, names s…

SafetyDGX agent

I should also add that one of the strongest alignment actions that OpenAI did was to name their product chatgpt with gpt 5.5 medium, names so uninspired that nobody could see it as a friend. Unlike Cl

I think everyone would be okay with this, though.

ApplicationsDGX agent

I think everyone would be okay with this, though. Many people do not seem to want data centres built near them, despite the fact that they don't cause that much traffic and often generate a lot of loc

New paper (on an old AI) tests o1 against doctors on medical benchmarks & real ER cases: “across a variety of scenarios and applications, th…

ApplicationsDGX agent

New paper (on an old AI) tests o1 against doctors on medical benchmarks & real ER cases: “across a variety of scenarios and applications, the large language model outperformed both human physicians an

Nice thread by an author

ApplicationsDGX agent

Nice thread by an author 🧵1/ Our new study on AI and physician reasoning just came out in @ScienceMagazine. As co-senior author, I'm excited about our findings, and I do think AI will reshape medicine

Organizations are already superhuman intelligences. The University of Pennsylvania or Walmart or whatever is far more capable than any human…

ApplicationsDGX agent

Organizations are already superhuman intelligences. The University of Pennsylvania or Walmart or whatever is far more capable than any human. That is why the focus on AIs as individual productivity to

Randomized trial of an AI therapy chatbot on Mexican women found “improved mental health by 0.3 SD over 6 months with no evidence of an incr…

AgentsDGX agent

Randomized trial of an AI therapy chatbot on Mexican women found “improved mental health by 0.3 SD over 6 months with no evidence of an increase of severe cases; improved sleep, healthful behaviors, d

The goblin thing was fun as it was a real quirk that was emblematic of what makes AI interesting, and it organically came out of an AI user …

ApplicationsDGX agent

The goblin thing was fun as it was a real quirk that was emblematic of what makes AI interesting, and it organically came out of an AI user discovery. So was, for what it was worth, Ghiblitization Whe

The new Grok comes in below the latest Chinese open weights models, Grok 4 was at the frontier when released. (& Artificial Analysis: please…

Model ReleasesDGX agent

The new Grok comes in below the latest Chinese open weights models, Grok 4 was at the frontier when released. (& Artificial Analysis: please stop using GDPval-AA which is not a useful test of anything

We need more work on AI inequality, but this study is not about GenAI, the survey was fielded in 2022. “In this study, we selected items fro…

ApplicationsDGX agent

We need more work on AI inequality, but this study is not about GenAI, the survey was fielded in 2022. “In this study, we selected items from Wave 119 (N = 10,087), which were collected from December

Yes, that includes the road ending in the river for some reason.

ApplicationsDGX agent

This post likely discusses an unexpected or unusual infrastructure discovery, possibly referencing a road that literally terminates at a river without proper connection or explanation. Based on Ethan

30 Apr 2026

AGI is GPT-X doing all of the work after you say 'We would love you to throw a party for yourself as a marketing event for OpenAI, so do tha…

ApplicationsDGX agent

This post by Ethan Mollick likely discusses how advanced AI systems like hypothetical future GPT versions could autonomously execute complex, real-world tasks with minimal human direction, using a hum

Ancient catastrophes, silence, unsaid things that everyone knows (even when this doesn't make a lot of sense in the story). Also terrible an…

ApplicationsDGX agent

This post appears to discuss narrative and storytelling techniques that rely on implied catastrophes, unspoken tensions, and shared understanding between author and audience—elements that create drama

ASI* is GPT-X deciding I want to throw a party for myself and every human on earth will get a personalized email them that will help them in…

ApplicationsDGX agent

ASI* is GPT-X deciding I want to throw a party for myself and every human on earth will get a personalized email them that will help them in some way, and also here is a cure for a bunch of diseases *

For better or worse, regulation for closed-source models served by a few (quite large) companies is easy. It is not as easy to imagine how y…

SafetyDGX agent

For better or worse, regulation for closed-source models served by a few (quite large) companies is easy. It is not as easy to imagine how you regulate open-source models that can be served by a range

Forget goblins, things that GPT-5.5 really likes in its fiction: lighthouses, the ocean, maps, bells, clock towers with bells that ring impo…

Model ReleasesDGX agent

Forget goblins, things that GPT-5.5 really likes in its fiction: lighthouses, the ocean, maps, bells, clock towers with bells that ring impossible times, Mira Vale, resonances and echoes (Claude and G

I think the Gemini chatbot has all the pieces to be a useful tool, but struggles to put it all together. It still doesn't seem to know what …

Model ReleasesDGX agent

I think the Gemini chatbot has all the pieces to be a useful tool, but struggles to put it all together. It still doesn't seem to know what files it can create or how its tools work together. It also

Illustration of the jagged frontier as a PR thing: 1) People had to ask the AI for a party date 2) People wrote the social media posts about…

Model ReleasesDGX agent

Illustration of the jagged frontier as a PR thing: 1) People had to ask the AI for a party date 2) People wrote the social media posts about the party, set up the invite list 3) People had to solicit

Increasingly, I think, we will see a gap between what you can do with frontier model APIs & what you can do with the native apps from the fr…

Model ReleasesDGX agent

Increasingly, I think, we will see a gap between what you can do with frontier model APIs & what you can do with the native apps from the frontier labs (Codex, Claude Code). Models developed and train

It is really interesting that Microsoft and OpenAI have access to the exact same models at the exact same time, and they have done such diff…

ApplicationsDGX agent

It is really interesting that Microsoft and OpenAI have access to the exact same models at the exact same time, and they have done such different things with them. A rare pure experiment with a no-nam

'Load bearing,' 'I keep coming back to,' 'Not X, but Y' A curse of using AI a lot is that you realize how much of the writing around you is …

ApplicationsDGX agent

'Load bearing,' 'I keep coming back to,' 'Not X, but Y' A curse of using AI a lot is that you realize how much of the writing around you is just AI, now People who don't use AI have been unable to ide

Mythos seems to be a very capable model based on available information, but it is not a cybersecurity model - it is an advanced general purp…

ApplicationsDGX agent

Mythos seems to be a very capable model based on available information, but it is not a cybersecurity model - it is an advanced general purpose model that happens to be good at cyber because it is goo

29 Apr 2026

Gemini now can create documents, and it is a nice start, but not up to the frontier yet, as you can see from my 'LBO of Hogwarts' test. Powe…

Model ReleasesDGX agent

Gemini now can create documents, and it is a nice start, but not up to the frontier yet, as you can see from my 'LBO of Hogwarts' test. PowerPoints are substantially worse than NotebookLM, spreadsheet

Its a bit frustrating, because Gemini 3.1 Pro is an excellent model and can deliver really good results. But here is GPT-5.5 Pro for compari…

Model ReleasesDGX agent

Its a bit frustrating, because Gemini 3.1 Pro is an excellent model and can deliver really good results. But here is GPT-5.5 Pro for comparison. Sadly, it took this assignment very seriously and ethic

One reason I don’t think “judgment” is going to be a distinctly human role in working with AI is that the most recent agentic models have go…

AgentsDGX agent

One reason I don’t think “judgment” is going to be a distinctly human role in working with AI is that the most recent agentic models have gotten quite good at some types of judgment. You can’t do the

Yes, just having students “use AI to study” hurts learning (a helpful assistant is not a tutor), but using AI prompted to act like a tutor, …

ApplicationsDGX agent

Yes, just having students “use AI to study” hurts learning (a helpful assistant is not a tutor), but using AI prompted to act like a tutor, especially with teacher support, seems to have large positiv

← Previous
1…7891011…13
Next →