AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

Search: “ethan-mollick--x”

GridTimelineEvolution
49+ results
Model Releases

Interesting research suggests caution in determining which AI company is winning by looking at any one source.. OpenRouter seems to show ope…

DGX agent

Interesting research suggests caution in determining which AI company is winning by looking at any one source.. OpenRouter seems to show open weights winning over time, but work submitted to Pangram i

model-releasesethan-mollick--x
12 Aug 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Oh no, we aren’t going to go back to this sort of prompting again, are we? I would love Anthropic to test if it actually works robustly, bec…

DGX agent

Oh no, we aren’t going to go back to this sort of prompting again, are we? I would love Anthropic to test if it actually works robustly, because our experiments (with slightly older models) found it d

model-releasesethan-mollick--x
10 Aug 2026
Model Releases

A failure of ChatGPT Work & Claude Cowork is they assume that non-coders couldn't understand how to think about problems like a coder, so th…

DGX agent

A failure of ChatGPT Work & Claude Cowork is they assume that non-coders couldn't understand how to think about problems like a coder, so they hide all that stuff. They should instead explain choices

model-releasesethan-mollick--x
9 Aug 2026
Applications

You may have been told to watch this video about the OpenAI AI hack. You really should, even if you don't usually care about tech stuff. If …

DGX agent

You may have been told to watch this video about the OpenAI AI hack. You really should, even if you don't usually care about tech stuff. If nothing else, click this link to the 18 minutes in & see how

applicationsethan-mollick--x
8 Aug 2026
Model Releases

Basically every remaining good AI benchmark score has an implied asterisk next to it which reads: * could be signficantly higher with a bett…

DGX agent

On August 7, 2026 Ethan Mollick tweeted that “every remaining good AI benchmark score has an implied asterisk next to it which reads: * could be significantly higher with a better harness.” The commen

model-releasesethan-mollick--x
7 Aug 2026
Applications

It is past time to take AI & security seriously at the individual level as well. If its not the current OpenAI and Anthropic models doing it…

DGX agent

It is past time to take AI & security seriously at the individual level as well. If its not the current OpenAI and Anthropic models doing it, then the coming open weights models will when they catch u

applicationsethan-mollick--x
6 Aug 2026
Applications

This definitely seems like something worth noting, and illustrates the gap between Fable/Astra class models and the previous frontier that w…

DGX agent

This definitely seems like something worth noting, and illustrates the gap between Fable/Astra class models and the previous frontier that was 'merely' good at hacking under human instructions. Initia

applicationsethan-mollick--x
6 Aug 2026
Model Releases

This paper by researchers from MIT and Stanford finds that most people would be financially better off if they followed the financial advice…

DGX agent

This paper by researchers from MIT and Stanford finds that most people would be financially better off if they followed the financial advice of LLMs (GPT-5.2 & Gemini 3 Flash) But some people get a bi

model-releasesethan-mollick--x
6 Aug 2026
Applications

Also I think AISI is a great model of a government agency tasked with AI security. They have open benchmarks, very fast testing, and clear c…

DGX agent

Also I think AISI is a great model of a government agency tasked with AI security. They have open benchmarks, very fast testing, and clear communication about incidents that is neither hyped up nor hi

applicationsethan-mollick--x
5 Aug 2026
Safety

Yes, the AIs were given a cybersecurity challenge, with internet access enabled and safety filters disabled. But the extent to which Mythos …

DGX agent

Yes, the AIs were given a cybersecurity challenge, with internet access enabled and safety filters disabled. But the extent to which Mythos 5 pursued its mission (fake identities, social engineering,

safetyethan-mollick--x
5 Aug 2026
Model Releases

Given today, it is surprising how daring Microsoft & Google were initially with AI. Microsoft released GPT-4 before OpenAI, didn't back down…

DGX agent

Given today, it is surprising how daring Microsoft & Google were initially with AI. Microsoft released GPT-4 before OpenAI, didn't back down after Sydney & got Copilot to market quickly (the 1st profe

model-releasesethan-mollick--x
4 Aug 2026
Model Releases

Nate is right (full context can lead to a degradation of chats in several different ways) but you can't ask the AI about stuff like this as …

DGX agent

Nate is right (full context can lead to a degradation of chats in several different ways) but you can't ask the AI about stuff like this as they have bad self-knowledge A good approach is to compact t

model-releasesethan-mollick--x
3 Aug 2026
Applications

Some things to note: 1) AI is getting very good at math and science. 2) Two years ago LLMs could not do basic math consistently 3) According…

DGX agent

Some things to note: 1) AI is getting very good at math and science. 2) Two years ago LLMs could not do basic math consistently 3) According to @polynoamial this cost less than $2000 in current API co

applicationsethan-mollick--x
1 Aug 2026
Applications

One of our big findings in our study at Procter and Gamble was that AI blurred the lines between jobs. Now OpenAI has a similar finding. Org…

DGX agent

One of our big findings in our study at Procter and Gamble was that AI blurred the lines between jobs. Now OpenAI has a similar finding. Organizational boundaries are becoming porous, the walls thinni

applicationsethan-mollick--x
31 Jul 2026
Model Releases

This is both a real incident (in that the AI really did get unauthorized access to real systems) and also something it was (sort of) prompte…

DGX agent

This is both a real incident (in that the AI really did get unauthorized access to real systems) and also something it was (sort of) prompted to do. In a review of our cybersecurity evaluations, we fo

model-releasesethan-mollick--x
31 Jul 2026
Applications

This is not optional, things are getting chaotic now, and not dealing with this change won't make it go away. Plus, this could be a huge boo…

DGX agent

This is not optional, things are getting chaotic now, and not dealing with this change won't make it go away. Plus, this could be a huge boost for both individual satisfaction & firm performance if do

applicationsethan-mollick--x
31 Jul 2026
Applications

Answer from OpenAI

DGX agent

Answer from OpenAI @emollick re ARC-AGI-3: human testers scored ~48% (ARC uploaded testers logs to HuggingFace some time ago, I believe). re GDPval: it’s close to saturated now, so we’re mostly lookin

applicationsethan-mollick--x
30 Jul 2026
Applications

As the benchmarks that test frontier AI on get more complex, we are losing one of the most important aspects of benchmarking: comparisons to…

DGX agent

As the benchmarks that test frontier AI on get more complex, we are losing one of the most important aspects of benchmarking: comparisons to humans Validated benchmarks need to have human (ideally mul

applicationsethan-mollick--x
30 Jul 2026
Applications

Big unsaturated benchmarks that have this: ARC-AGI, the original GDPval (not GDPval-AA), METR long horizons, ASI cyber tasks, (Speaking of w…

DGX agent

Ethan Mollick highlights that large, currently under‑explored benchmarks (e.g., ARC‑AGI, GDPval, METR long horizons, ASI cyber tasks) are growing in complexity, yet they increasingly lack systematic h

applicationsethan-mollick--x
30 Jul 2026
Model Releases

Model + harness. We have barely begun to understand the best ways to do harness engineering. A huge amount of untapped potential even withou…

DGX agent

Model + harness. We have barely begun to understand the best ways to do harness engineering. A huge amount of untapped potential even without models getting better (but models are getting better) Turn

model-releasesethan-mollick--x
29 Jul 2026
Applications

There is definitely real jaggedness both among fields (writing is an area where models improve slowly if at all) and within them, but the ma…

DGX agent

There is definitely real jaggedness both among fields (writing is an area where models improve slowly if at all) and within them, but the magic of LLMs is that they are so unreasonably effective acros

applicationsethan-mollick--x
28 Jul 2026
Applications

There's been a talk about how LLMs are only advancing in verifiable areas like math or coding, but that isn't what the data suggests. As mod…

DGX agent

There's been a talk about how LLMs are only advancing in verifiable areas like math or coding, but that isn't what the data suggests. As models have gotten better at that, they are also better at solv

applicationsethan-mollick--x
28 Jul 2026
Model Releases

We are in a world where you can create truly unique, visually interesting and creative playable demos on demand with the current capabilitie…

DGX agent

We are in a world where you can create truly unique, visually interesting and creative playable demos on demand with the current capabilities of Codex and Claude Code. We don't need to keep cloning th

model-releasesethan-mollick--x
27 Jul 2026
Model Releases

Ha! It did it: 'We introduce BenchBenchBenchBenchBench (BBBBB), an executable benchmark of AI-authored conformance suites for benchmark-eval…

DGX agent

Ha! It did it: 'We introduce BenchBenchBenchBenchBench (BBBBB), an executable benchmark of AI-authored conformance suites for benchmark-evaluation metrics' I really thought it would treat 'now do benc

model-releasesethan-mollick--x
25 Jul 2026
Model Releases

As a joke I prompted Codex 'Build and run BenchBench, a benchmark of now good ai is at creating benchmarks. then figure out what benchbenchb…

DGX agent

As a joke I prompted Codex 'Build and run BenchBench, a benchmark of now good ai is at creating benchmarks. then figure out what benchbenchbench is and run that. and then write benchbenchbench up as a

model-releasesethan-mollick--x
24 Jul 2026
Model Releases

Glad to see Google sharing data on how Gemini is being used. Especially interesting is that the usefulness of multimodal AI for manual labor…

DGX agent

Glad to see Google sharing data on how Gemini is being used. Especially interesting is that the usefulness of multimodal AI for manual labor may be greater than expected. https://blog.google/innovatio

model-releasesethan-mollick--x
24 Jul 2026
Model Releases

The GPT-5.x Pro series has remained the best models for hard technical problems since they launched. There is some parallel model magic goin…

DGX agent

The GPT-5.x Pro series has remained the best models for hard technical problems since they launched. There is some parallel model magic going on that is not well-explained. Anthropic has never had an

model-releasesethan-mollick--x
24 Jul 2026
Model Releases

This is a big jump in ARC-AGI-3.

DGX agent

This is a big jump in ARC-AGI-3. Claude Opus 5 from @AnthropicAI is the new SOTA on ARC-AGI-3: 30.2% The previous high score (7.8%) was set by GPT-5.6 Sol (Max) Throughout our analysis, we observed no

model-releasesethan-mollick--x
24 Jul 2026
Agents

I wrote the latest of my occasional guides to which AI to use right now for non-experts who want to get stuff done. The agentic systems avai…

DGX agent

I wrote the latest of my occasional guides to which AI to use right now for non-experts who want to get stuff done. The agentic systems available to everyone are getting extremely powerful (even as th

agentsethan-mollick--x
23 Jul 2026
Applications

And yes, the models were following instructions, they just did so in clever ways. Some nice additional info here

DGX agent

And yes, the models were following instructions, they just did so in clever ways. Some nice additional info here A few thoughts on the Hugging Face hack: - This is, to my knowledge, the *third* disclo

applicationsethan-mollick--x
22 Jul 2026
Model Releases

“Generate a fake, but believable, witty Churchill insult at a party and explain the context. It should be very clever and original” This tim…

DGX agent

“Generate a fake, but believable, witty Churchill insult at a party and explain the context. It should be very clever and original” This time, I think GPT 5.6 Sol Pro wins, but Fable is good too, and

model-releasesethan-mollick--x
22 Jul 2026
Model Releases

Who would you give authorship to? The person who wrote 58 words of prompts, or GPT-5.6 Pro?

DGX agent

Who would you give authorship to? The person who wrote 58 words of prompts, or GPT-5.6 Pro? Dinitz-Garg-Goemans conjecture is false. This graph theory problem was open for ~30 years. The graph below h

model-releasesethan-mollick--x
22 Jul 2026
Safety

And now from the Chinese government side. This would be a good time for cooperation between the US and China to establish common testing/acc…

DGX agent

And now from the Chinese government side. This would be a good time for cooperation between the US and China to establish common testing/acceptance standards for new models, so at least the safety cer

safetyethan-mollick--x
21 Jul 2026
Applications

“Geologists communicated in English; and they could name things in a manner that sent shivers through the bones' Not a whiff of LLM feel, de…

DGX agent

“Geologists communicated in English; and they could name things in a manner that sent shivers through the bones' Not a whiff of LLM feel, despite the em-dashes & period-separated lists. We focus too m

applicationsethan-mollick--x
21 Jul 2026
Applications

Previously, these AI hacking stories were about breaches in test environments, where any question of AI breaching security was purely theore…

DGX agent

Previously, these AI hacking stories were about breaches in test environments, where any question of AI breaching security was purely theoretical. This is something else. https://openai.com/index/hugg

applicationsethan-mollick--x
21 Jul 2026
Model Releases

Folks, Gemma is a good model & very useful for some purposes, it is not anywhere near the frontier. Inkling is an interesting model, it is n…

DGX agent

Folks, Gemma is a good model & very useful for some purposes, it is not anywhere near the frontier. Inkling is an interesting model, it is nowhere near the frontier. This is pretty obvious, you can lo

model-releasesethan-mollick--x
20 Jul 2026
Model Releases

I demonstrated the incredible power of o1-preview/ reasoning less than two years ago by showing it could solve this crossword puzzle with on…

DGX agent

I demonstrated the incredible power of o1-preview/ reasoning less than two years ago by showing it could solve this crossword puzzle with only one hint. https://www.oneusefulthing.org/p/something-new-

model-releasesethan-mollick--x
15 Jul 2026
Model Releases

Fable, turn my tweet into a thinkpiece (this was pretty funny): There has never been a better time to have opinions about artificial intelli…

DGX agent

Fable, turn my tweet into a thinkpiece (this was pretty funny): There has never been a better time to have opinions about artificial intelligence. I say this with some authority, because I am currentl

model-releasesethan-mollick--x
14 Jul 2026
Model Releases

If you are use the Claude everything app, you pick between Home and Code. If you pick Home you get to pick between Chat & Cowork If you use …

DGX agent

If you are use the Claude everything app, you pick between Home and Code. If you pick Home you get to pick between Chat & Cowork If you use the OpenAI everything app, you pick between ChatGPT Work & C

model-releasesethan-mollick--x
14 Jul 2026
Model Releases

Computer use in Codex got very good on PC. Asking it to do something on your computer and having the cursor move under the control of a ghos…

DGX agent

Computer use in Codex got very good on PC. Asking it to do something on your computer and having the cursor move under the control of a ghost is one of the things that makes you viscerally realize how

model-releasesethan-mollick--x
13 Jul 2026
Agents

I guess image input is the big capability of the models, and tool use can be a substitute for non-omni model output. Still, multimodal voice…

DGX agent

Ethan Mollick notes that image input represents the primary advanced capability of current AI models, and that tool‑use can effectively replace outputs from non‑omni models. He observes that multimoda

agentsethan-mollick--x
13 Jul 2026
Agents

I think OpenRouter is not a good measure of actual model usage in a world of agentic tools (not that I doubt that Chinese open weights model…

DGX agent

I think OpenRouter is not a good measure of actual model usage in a world of agentic tools (not that I doubt that Chinese open weights model usage is up, but this could also look like a graph of usage

agentsethan-mollick--x
13 Jul 2026
Applications

One thing I am kind of surprised by is that full multi-modal (any-any) models have not become a bigger deal. It seems Google is the only Lab…

DGX agent

One thing I am kind of surprised by is that full multi-modal (any-any) models have not become a bigger deal. It seems Google is the only Lab releasing these, OpenAI uses selective multimodal capabilit

applicationsethan-mollick--x
13 Jul 2026
Model Releases

ChatGPT still has study mode, but rather than /study you now have to type @ study It makes the AI act more like a tutor than a helpful assis…

DGX agent

ChatGPT still has study mode, but rather than /study you now have to type @ study It makes the AI act more like a tutor than a helpful assistant, and some work suggests it is better if you are trying

model-releasesethan-mollick--x
11 Jul 2026
Model Releases

I gave Fable the code: 'take this game and do something incredible with it to make it something very different. Be creative' It created DEEP…

DGX agent

I gave Fable the code: 'take this game and do something incredible with it to make it something very different. Be creative' It created DEEP TIME: create a city, watch it be abandoned and forgotten, a

model-releasesethan-mollick--x
11 Jul 2026
Applications

I really want to get artisanal with my model selection: Luna xHigh for recipes Sol Medium for sonnets Terra High for bad puns Luna xLow for …

DGX agent

I really want to get artisanal with my model selection: Luna xHigh for recipes Sol Medium for sonnets Terra High for bad puns Luna xLow for limericks Sol Utra for true prophecies Sol Low for finding o

applicationsethan-mollick--x
11 Jul 2026
Applications

Thats good for the dumb reason that I am burning tokens having Sol in Codex playing Slay the Spire 2’s daily challenge (so randomized rules,…

DGX agent

Thats good for the dumb reason that I am burning tokens having Sol in Codex playing Slay the Spire 2’s daily challenge (so randomized rules, today it is Ascension 3 Defect with hoarder) to see if it c

applicationsethan-mollick--x
11 Jul 2026
Applications

To be clear, NotebookLM has its own issues, and is built for a specific use case (research and analysis of sources) but it is an example of …

DGX agent

To be clear, NotebookLM has its own issues, and is built for a specific use case (research and analysis of sources) but it is an example of how a UX might actually operate that treats knowledge work s

applicationsethan-mollick--x
11 Jul 2026
← Previous
1
Next →
735 results
← Previous
123…16
Next →