Watching @awnihannun at @ollama
I was unable to retrieve the specific content from the X (Twitter) URL provided (`https://x.com/twid/status/2042425382859841926`), as it is a social media post that requires authentication to acces...
Knowledge catalogue
I was unable to retrieve the specific content from the X (Twitter) URL provided (`https://x.com/twid/status/2042425382859841926`), as it is a social media post that requires authentication to acces...
We had Lin on stage: 'the future is millions of models — one per application, one per use case.' Jet delivered a masterclass on reinforcement fine-tuning. Rob joined @WorkOS for some hot takes on the
arXiv:2510.12476v2 Announce Type: replace Abstract: Large language models (LLMs) have grown more powerful in language generation, producing fluent text and even imitating personal style. Yet, this abi
Anthropic's new hosted service for long-running AI agents, designed to solve the challenge of creating systems that support 'programs as yet unthought of.' It abstracts infrastructure management to en
Another banger paper from Microsoft. Why it's a big deal: It teaches reasoning models to compress their own chain-of-thought mid-generation. The most interesting finding isn't the 2-3x memory savings
Anthropic's new frontier model, Claude Mythos, is the first model the company has publicly deemed too high-risk for general release , due to its advanced cybersecurity capabilities — including the ...
Mobile app security solutions provider Appknox today announced the launch of KnoxIQ, an artificial intelligence-native vulnerability assessment capability that introduces a new prioritization and reme
The referenced tweet (from @himanshustwts) is not publicly accessible without a login, so its exact content cannot be retrieved. However, based on contextual search results, this post appears to be...
Here is my free workshop on building multi-agent systems, I presented at the @aiDotEngineer London conference together with @Whats_AI. It has code, slides and soon a 2-hour video diving deep into how
I just gave a workshop at @aiDotEngineer in London on building real multi-agent systems. The best part was hearing people laugh, interrupt us with questions.. You could feel they were following, think
If you want an example of what this looks like in practice, check out the '/research-docs' skill I created for Claude Code https://x.com/jerryjliu0/status/2041564207750246904?s=20 I built a Claude Cod
The environmental impact of individual ChatGPT use is widely considered overstated: Epoch AI estimated a typical ChatGPT query uses just 0.3 Wh of electricity — ten times less than older 2023 esti...
Judging by my tl there is a growing gap in understanding of AI capability. The first issue I think is around recency and tier of use. I think a lot of people tried the free tier of ChatGPT somewhere l
Meta's new AI can predict your brain better than a brain scan. TRIBE v2 is a foundation model trained on 1,000+ hours of brain imaging data from 720 people. You feed it a video, sound clip, or text, a
My considered opinion is that The Mythos stuff was mostly a myth. Take it as a serious warning sign that we need to get our act together with respect to cybersecurity. But don’t take the details serio
Following a courtroom defeat, HHS Secretary RFK Jr. rewrote the charter of the CDC's Advisory Committee on Immunization Practices (ACIP), broadening its membership criteria, increasing its focus on...
// Scaling Coding Agents via Atomic Skills // Most coding agents train end-to-end on full tasks like resolving GitHub issues. But complex software engineering is really a composition of simpler skills
A Reddit discussion thread on r/MachineLearning in which practitioners explore how foundational concepts from Sutton and Barto's *Reinforcement Learning: An Introduction* — including MDPs, policy g...
March ships Foundry Agent Service GA with private networking, GPT-5.4 and GPT-5.4 Mini, Priority Processing, Phi-4 Reasoning Vision, SDK 2.0 GA across Python, JS/TS, Java, and .NET, Fireworks AI and
Anthropic announced it has grown its ARR from $19B to $30B in just one month, coinciding with the formal unveiling of Claude Mythos Preview — described in leaked company documents as 'by far the m...
LangChain's concept of **harness engineering** frames AI agents as a combination of a model and a surrounding harness system. An agent equals a model plus a harness — harness engineering is how sy...
Director James Cameron on why Big Tech owning AGI is scarier than any science fiction he's ever made: 'AGI will not emerge from a government funded program. It will emerge from one of the tech giants
OpenAI's Child Safety Blueprint, released in April 2026, is a policy framework aimed at combating the rise of AI-enabled child sexual exploitation by combining legal, operational, and technical app...
Meta announced Muse Spark today, their first model release since Llama 4 almost exactly a year ago. It's hosted, not open weights, and the API is currently 'a private API preview to select users', but
Self-improving agents isn’t a single algorithm - it’s a systems engineering problem involving: - eval data curation + maintenance - experiment design to battle overfitting - an update algorithm - huma
Some sober thinking about Mythos (full version with links at my newsletter): 1It’s probably not as bad as they say, as AI and cybersecurity expert @HeidyKhlaaf explains elsewhere (in a thread “As some
This is literally why I wrote Taming Silicon Valley. Director James Cameron on why Big Tech owning AGI is scarier than any science fiction he's ever made: 'AGI will not emerge from a government funded
In February 2019, OpenAI announced GPT-2, a language model trained on text from 8 million webpages to predict the next word in a piece of writing, capable of adapting to the style and content of a...
AI critic Gary Marcus uses the irony of ChatGPT's Voice mode being unable to perform a basic task — starting a timer — as a pointed illustration of the gap between AI industry hype and real-world c...
We're seeing even more autonomous AI coworkers. The new MLE agent on the market is Disarray. In Kaggle competitions, Disarray: - won 28 medals across diverse domains (vision, NLP, tabular data) - plac
Before limited-releasing Claude Mythos Preview, we investigated its internal mechanisms with interpretability techniques. We found it exhibited notably sophisticated (and often unspoken) strategic thi
Julien Chaumond (co-founder of Hugging Face) posted a tweet referencing the infamous 2019 OpenAI decision to initially withhold GPT-2-large from public release due to fears it was 'too dangerous,' ...
Mythos is very powerful, and should feel terrifying. I am proud of our approach to responsibly preview it with cyber defenders, rather than generally releasing it into the wild. Model card here: https
SuperClaude (Mythos) still seems irreducibly Claude-y given the transcripts in the system card. Here two versions of Mythos are forced to talk to each other across multiple rounds. They are less philo
Thank you to @AnthropicAI for sending FFmpeg patches Introducing Project Glasswing: an urgent initiative to help secure the world’s most critical software. It’s powered by our newest frontier model, C