AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

Search: “ethan-mollick--x”

GridTimelineEvolution
735 results
CompaniesToolsTechniques

Each lane shows up to 8 recent matching entries, ordered from earlier to later. Tracks load separately to keep the 75,000+ entry wiki fast.

Companies

CompanyAnthropic8 recent entries
3 Aug 2026Nate is right (full context can lead to a degradation of chats in several different ways) but you can't ask the AI about stuff like this as …

Nate is right (full context can lead to a degradation of chats in several different ways) but you can't ask the AI about stuff like this as they have bad self-knowledge A good approach is to compact t

→5 Aug 2026Yes, the AIs were given a cybersecurity challenge, with internet access enabled and safety filters disabled. But the extent to which Mythos …

Yes, the AIs were given a cybersecurity challenge, with internet access enabled and safety filters disabled. But the extent to which Mythos 5 pursued its mission (fake identities, social engineering,

3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
→5 Aug 2026Also I think AISI is a great model of a government agency tasked with AI security. They have open benchmarks, very fast testing, and clear c…

Also I think AISI is a great model of a government agency tasked with AI security. They have open benchmarks, very fast testing, and clear communication about incidents that is neither hyped up nor hi

→6 Aug 2026This paper by researchers from MIT and Stanford finds that most people would be financially better off if they followed the financial advice…

This paper by researchers from MIT and Stanford finds that most people would be financially better off if they followed the financial advice of LLMs (GPT-5.2 & Gemini 3 Flash) But some people get a bi

→6 Aug 2026It is past time to take AI & security seriously at the individual level as well. If its not the current OpenAI and Anthropic models doing it…

It is past time to take AI & security seriously at the individual level as well. If its not the current OpenAI and Anthropic models doing it, then the coming open weights models will when they catch u

→9 Aug 2026A failure of ChatGPT Work & Claude Cowork is they assume that non-coders couldn't understand how to think about problems like a coder, so th…

A failure of ChatGPT Work & Claude Cowork is they assume that non-coders couldn't understand how to think about problems like a coder, so they hide all that stuff. They should instead explain choices

→10 Aug 2026Oh no, we aren’t going to go back to this sort of prompting again, are we? I would love Anthropic to test if it actually works robustly, bec…

Oh no, we aren’t going to go back to this sort of prompting again, are we? I would love Anthropic to test if it actually works robustly, because our experiments (with slightly older models) found it d

→12 Aug 2026Interesting research suggests caution in determining which AI company is winning by looking at any one source.. OpenRouter seems to show ope…

Interesting research suggests caution in determining which AI company is winning by looking at any one source.. OpenRouter seems to show open weights winning over time, but work submitted to Pangram i

CompanyOpenAI8 recent entries
1 Aug 2026Some things to note: 1) AI is getting very good at math and science. 2) Two years ago LLMs could not do basic math consistently 3) According…

Some things to note: 1) AI is getting very good at math and science. 2) Two years ago LLMs could not do basic math consistently 3) According to @polynoamial this cost less than $2000 in current API co

→4 Aug 2026Given today, it is surprising how daring Microsoft & Google were initially with AI. Microsoft released GPT-4 before OpenAI, didn't back down…

Given today, it is surprising how daring Microsoft & Google were initially with AI. Microsoft released GPT-4 before OpenAI, didn't back down after Sydney & got Copilot to market quickly (the 1st profe

→6 Aug 2026This definitely seems like something worth noting, and illustrates the gap between Fable/Astra class models and the previous frontier that w…

This definitely seems like something worth noting, and illustrates the gap between Fable/Astra class models and the previous frontier that was 'merely' good at hacking under human instructions. Initia

→6 Aug 2026It is past time to take AI & security seriously at the individual level as well. If its not the current OpenAI and Anthropic models doing it…

It is past time to take AI & security seriously at the individual level as well. If its not the current OpenAI and Anthropic models doing it, then the coming open weights models will when they catch u

→7 Aug 2026Basically every remaining good AI benchmark score has an implied asterisk next to it which reads: * could be signficantly higher with a bett…

On August 7, 2026 Ethan Mollick tweeted that “every remaining good AI benchmark score has an implied asterisk next to it which reads: * could be significantly higher with a better harness.” The commen

→8 Aug 2026You may have been told to watch this video about the OpenAI AI hack. You really should, even if you don't usually care about tech stuff. If …

You may have been told to watch this video about the OpenAI AI hack. You really should, even if you don't usually care about tech stuff. If nothing else, click this link to the 18 minutes in & see how

→9 Aug 2026A failure of ChatGPT Work & Claude Cowork is they assume that non-coders couldn't understand how to think about problems like a coder, so th…

A failure of ChatGPT Work & Claude Cowork is they assume that non-coders couldn't understand how to think about problems like a coder, so they hide all that stuff. They should instead explain choices

→12 Aug 2026Interesting research suggests caution in determining which AI company is winning by looking at any one source.. OpenRouter seems to show ope…

Interesting research suggests caution in determining which AI company is winning by looking at any one source.. OpenRouter seems to show open weights winning over time, but work submitted to Pangram i

CompanyGoogle8 recent entries
1 Jul 2026You really need to benchmark models for your use case. As soon as judgements & decisions stack on top of each other, the differences between…

You really need to benchmark models for your use case. As soon as judgements & decisions stack on top of each other, the differences between models amplifies, and no standard benchmark will tell you t

→2 Jul 2026You really need your own benchmarks. If you are translating hieroglyphics, use Gemini 3.5 Flash. If you are running a vending machine use Op…

You really need your own benchmarks. If you are translating hieroglyphics, use Gemini 3.5 Flash. If you are running a vending machine use Opus 4.8. (This is one reason why I am skeptical of just swapp

→6 Jul 2026Less than 2 years later, with Fable: 'simulate an encounter between a mind flayer and a drow warrior of equal CR. Set up initial stats then …

Less than 2 years later, with Fable: 'simulate an encounter between a mind flayer and a drow warrior of equal CR. Set up initial stats then simulate each move, including initiative, rolling dice as ne

→11 Jul 2026ChatGPT still has study mode, but rather than /study you now have to type @ study It makes the AI act more like a tutor than a helpful assis…

ChatGPT still has study mode, but rather than /study you now have to type @ study It makes the AI act more like a tutor than a helpful assistant, and some work suggests it is better if you are trying

→13 Jul 2026Computer use in Codex got very good on PC. Asking it to do something on your computer and having the cursor move under the control of a ghos…

Computer use in Codex got very good on PC. Asking it to do something on your computer and having the cursor move under the control of a ghost is one of the things that makes you viscerally realize how

→22 Jul 2026“Generate a fake, but believable, witty Churchill insult at a party and explain the context. It should be very clever and original” This tim…

“Generate a fake, but believable, witty Churchill insult at a party and explain the context. It should be very clever and original” This time, I think GPT 5.6 Sol Pro wins, but Fable is good too, and

→24 Jul 2026Glad to see Google sharing data on how Gemini is being used. Especially interesting is that the usefulness of multimodal AI for manual labor…

Glad to see Google sharing data on how Gemini is being used. Especially interesting is that the usefulness of multimodal AI for manual labor may be greater than expected. https://blog.google/innovatio

→6 Aug 2026This paper by researchers from MIT and Stanford finds that most people would be financially better off if they followed the financial advice…

This paper by researchers from MIT and Stanford finds that most people would be financially better off if they followed the financial advice of LLMs (GPT-5.2 & Gemini 3 Flash) But some people get a bi

CompanyMeta8 recent entries
4 May 2026My surprise here seems warranted, this paper was retracted (There are other peer-reviewed meta-analyses of the impact of AI on education fin…

My surprise here seems warranted, this paper was retracted (There are other peer-reviewed meta-analyses of the impact of AI on education finding positive effects, like: https://www.researchgate.net/pu

→4 May 2026I think the fact that GPT-4o and Llama 3.3-80B did no significant harm is just as important as whether AI helped. If older (less accurate & …

I think the fact that GPT-4o and Llama 3.3-80B did no significant harm is just as important as whether AI helped. If older (less accurate & more sycophantic) chatbots essentially did nothing for peopl

→7 May 2026If Llama 4 didn’t fail, if Microsoft had pulled Sydney after the Roose article, if New Sonnet hadn’t been so good, if Orion hadn’t been so m…

If Llama 4 didn’t fail, if Microsoft had pulled Sydney after the Roose article, if New Sonnet hadn’t been so good, if Orion hadn’t been so meh, if the leadership change at OpenAI had happened, if a re

→2 Jun 2026It is difficult to know how good MAI-Thinking-1 is from the scores alone (like weirdly low GPQA & Terminal Bench 2.0) But Microsoft makes it…

It is difficult to know how good MAI-Thinking-1 is from the scores alone (like weirdly low GPQA & Terminal Bench 2.0) But Microsoft makes it really hard to try its models upon release (a general issue

→5 Jun 2026At least until (if?) rapid improvement stops, it seems less likely someone is going to catch the Big Three AI Labs. Microsoft and Meta relea…

At least until (if?) rapid improvement stops, it seems less likely someone is going to catch the Big Three AI Labs. Microsoft and Meta released their models, which were fine, but not frontier. SpaceX

→3 Jul 2026This is true… but maybe less important than the fact that people don’t try ambitious things with these systems. Many models are excellent as…

This is true… but maybe less important than the fact that people don’t try ambitious things with these systems. Many models are excellent as a Google replacement, for homework “help,” etc. It is someo

→13 Jul 2026I guess image input is the big capability of the models, and tool use can be a substitute for non-omni model output. Still, multimodal voice…

Ethan Mollick notes that image input represents the primary advanced capability of current AI models, and that tool‑use can effectively replace outputs from non‑omni models. He observes that multimoda

→14 Jul 2026Fable, turn my tweet into a thinkpiece (this was pretty funny): There has never been a better time to have opinions about artificial intelli…

Fable, turn my tweet into a thinkpiece (this was pretty funny): There has never been a better time to have opinions about artificial intelligence. I say this with some authority, because I am currentl

CompanyMistral1 recent entries
9 Apr 2026So we now have a pretty good picture of the state of the frontier AI model makers. US closed source models continue to lead. Google, OpenAI,…

So we now have a pretty good picture of the state of the frontier AI model makers. US closed source models continue to lead. Google, OpenAI, and Anthropic stand well ahead of the pack, and may have si

CompanyxAI8 recent entries
9 Apr 2026So we now have a pretty good picture of the state of the frontier AI model makers. US closed source models continue to lead. Google, OpenAI,…

So we now have a pretty good picture of the state of the frontier AI model makers. US closed source models continue to lead. Google, OpenAI, and Anthropic stand well ahead of the pack, and may have si

→1 May 2026The new Grok comes in below the latest Chinese open weights models, Grok 4 was at the frontier when released. (& Artificial Analysis: please…

The new Grok comes in below the latest Chinese open weights models, Grok 4 was at the frontier when released. (& Artificial Analysis: please stop using GDPval-AA which is not a useful test of anything

→6 May 2026I usually avoid commenting too much on industry deals, but this one is fascinating. Certainly seems like a blow to the idea that Grok will r…

I usually avoid commenting too much on industry deals, but this one is fascinating. Certainly seems like a blow to the idea that Grok will remain a frontier model. Our agreement with @SpaceX means we

→7 May 2026I’m curious- @grok explain all those references

Ethan Mollick asked Grok (xAI's AI assistant) to explain various cultural, historical, or technical references, likely exploring Grok's knowledge comprehension and ability to contextualize diverse top

→7 May 2026@grok, explain for those who don't want to look up the vocab words

This post likely features Grok (Elon Musk's AI chatbot) being used to explain complex concepts or vocabulary terms in simplified language for readers who prefer not to independently research definitio

→8 Jul 2026Here is the same series of prompts in Grok 4.5 for a twigl shader of a neo-gothic city drowning in a stormy sea. (Had an interesting error i…

Here is the same series of prompts in Grok 4.5 for a twigl shader of a neo-gothic city drowning in a stormy sea. (Had an interesting error initially where it made technically correct code that complet

→8 Jul 2026Added Grok 4.5 to the Harbor Town Playable gallery. Here is it's 'build me a procedurally generated 3D simulation showing the evolution of a…

Added Grok 4.5 to the Harbor Town Playable gallery. Here is it's 'build me a procedurally generated 3D simulation showing the evolution of a harbor town from 3000 BCE to 3000 AD, it should look beauti

→9 Jul 2026Since I am needling every model maker tonight about minor but important issues, one more: Grok 4.5 has no model card. Companies that are try…

Since I am needling every model maker tonight about minor but important issues, one more: Grok 4.5 has no model card. Companies that are trying to compete near the frontier should be releasing model c

CompanyDeepSeek8 recent entries
9 Apr 2026So we now have a pretty good picture of the state of the frontier AI model makers. US closed source models continue to lead. Google, OpenAI,…

So we now have a pretty good picture of the state of the frontier AI model makers. US closed source models continue to lead. Google, OpenAI, and Anthropic stand well ahead of the pack, and may have si

→24 Apr 2026My first two TiKZ Sparks unicorns from DeepSeek v4. (Expert mode, from the DeepSeek site, which is supposed to be v4 Pro according to the re…

Ethan Mollick documents his first attempts at using DeepSeek v4's expert mode to generate TiKZ code for creating unicorn graphics, sharing results from the DeepSeek website's v4 Pro interface. The pos

→24 Apr 2026I hope the upgrade to DeepSeek v4 will make the bot comments on here more bearable.

This post expresses hope that upgrading to DeepSeek v4 (an AI model) will improve the quality of bot-generated comments on a platform or service. The statement implies current bot comments are conside

→24 Apr 2026Here's DeepSeek v4 Pro. Added to the playable gallery as well.

Here's DeepSeek v4 Pro. Added to the playable gallery as well. Media I had a range of models 'build me a procedurally generated 3D simulation showing the evolution of a harbor town from 3000 BCE to 30

→24 Apr 2026And now a new DeepSeek model, and appears to be fully open weights. Good benchmarks, but with open models, that isn't always as meaningful. …

DeepSeek released a new open-weights model with strong benchmark performance, though Mollick notes that benchmark results may not fully capture the capabilities of open models compared to closed syste

→24 Jul 2026As a joke I prompted Codex 'Build and run BenchBench, a benchmark of now good ai is at creating benchmarks. then figure out what benchbenchb…

As a joke I prompted Codex 'Build and run BenchBench, a benchmark of now good ai is at creating benchmarks. then figure out what benchbenchbench is and run that. and then write benchbenchbench up as a

→7 Aug 2026Basically every remaining good AI benchmark score has an implied asterisk next to it which reads: * could be signficantly higher with a bett…

On August 7, 2026 Ethan Mollick tweeted that “every remaining good AI benchmark score has an implied asterisk next to it which reads: * could be significantly higher with a better harness.” The commen

→12 Aug 2026Interesting research suggests caution in determining which AI company is winning by looking at any one source.. OpenRouter seems to show ope…

Interesting research suggests caution in determining which AI company is winning by looking at any one source.. OpenRouter seems to show open weights winning over time, but work submitted to Pangram i