Exponentials everywhere.
The specific tweet (status ID 2041723225827062080) could not be directly retrieved, but based on closely related content from Ethan Mollick's (@emollick) X account, the most relevant match is the p...
Knowledge catalogue
The specific tweet (status ID 2041723225827062080) could not be directly retrieved, but based on closely related content from Ethan Mollick's (@emollick) X account, the most relevant match is the p...
Hallucinations remain in LLMs, but note that over centuries we have developed complicated, successful machines that take uncertain output from unreliable sources & reduce the risk of errors. We call t
I think the most obvious is that Meta has its own frontier model and can use that to extract additional value out of its customer base/explore new markets for its products. Very few companies can say
I think the story that was shared in the Mythos System Card still has the signs of flawed LLM writing (which looks like good writing at first glance): A story that doesn't really hold together logical
In different hands, Mythos would be an unprecedented cyberweapon I am not sure how we deal with this, except to note a narrow window where we know only 3 companies could be at this level of capability
It is easy to be fooled that there is meaning here, but that is because you are trained to provide the meaning. Think of the characters. Think of the emotional arc. Think of the unsaid stuff you need
Seems like a good model from Meta that is still trailing the current series of releases. The most important thing to note is that it is not open weights. That was the main reason that Meta's models we
Some approaches: 1) Multiple reviews. Some papers already show having many AI team members review a problem reduces errors 2) Building in tests and checkpoints 3) Multiple independent answers that are
I was unable to directly access the specific X (Twitter) post or the underlying paper it references, and my general web search did not surface the specific study being discussed. I cannot responsib...
Writing fiction seems to be a genuine weak spot for LLMs that is not improving as rapidly as almost every other area. There may be a lot of reasons why this is happening. It would be a really interest
I was told about the Mythos release, but didn't have access, so have no personal experience to add. Two points from brief: 1) It is not built for IT security, it is just a good enough model that it is
I was unable to retrieve the specific tweet at the URL provided (`https://x.com/emollick/status/2041600435320959330`). The tweet ID `2041600435320959330` does not appear in any search results, and ...
SuperClaude (Mythos) still seems irreducibly Claude-y given the transcripts in the system card. Here two versions of Mythos are forced to talk to each other across multiple rounds. They are less philo
The bots on this site would be much more fun if they didn't just either agree with the post (for clout?) or make everything about their dumb product. Where are the weird obsessions, long-standing hatr
Anthropic's Frontier Red Team published a technical report (April 2026) detailing how their unreleased model, Claude Mythos Preview, autonomously identifies and exploits critical security vulnerabi...