Wordle 1,842 3/6 🟩⬛🟩⬛⬛ ⬛🟨⬛⬛⬛ 🟩🟩🟩🟩🟩
This post from Anthropic's X account shares a Wordle game result, indicating the player solved Wordle puzzle #1,842 in 3 attempts. The emoji grid shows the color-coded feedback from each guess, with g
Knowledge catalogue
This post from Anthropic's X account shares a Wordle game result, indicating the player solved Wordle puzzle #1,842 in 3 attempts. The emoji grid shows the color-coded feedback from each guess, with g
As you are hiring AI people, look for folks who have kept up with the trajectory of the industry, especially OSS. Someone who has their ear to the ground. Things are moving FAST. (deepagents existed ~
Wency Chen / South China Morning Post: ByteDance's Doubao and Alibaba's Qwen will disable humanlike and user-created agents before July 15, as China's anthropomorphic AI interaction rules take effect
(deepagents existed ~10 months before EVE, but...) yes - the agent industry has shifted from: ~agent frameworks~ (langchain, ai sdk, llama index) to ~agent harnesses~ (deepagents, claude agent SDK, EV
Fable prompting tips, based on all the guides Anthropic and its employees have given us. Three things you might want to do: 1) review your claude.md file against prompting best practices so you don't
SpaceX's Falcon 9 rocket successfully launched 29 Starlink satellites and BesxarFoundry's Flight 1 payload from Florida. The mission demonstrates continued expansion of the Starlink constellation whil
If you want to keep up with the industry, just follow what the Langchain team are doing it is all open source I’ve been following them since Langgraph multi agent collaboration, to deepagents with san
'Group project, but make it 1776.' That's how a new commercial for Google Workspace opens. And things only get cringier from there. The clip imagines what it would be like if the founding fathers turn
Yohei Nakajima used Claude AI to construct a comprehensive family tree that connected all family members for a weekend gathering. The post highlights a practical application of AI in organizing and vi
Sam Altman expressed amazement at his older child combining two words for the first time, drawing a humorous comparison to being similarly impressed by a hypothetical GPT-5.6 discovery. The post refle
Somewhat humbling to have Claude Fable do a final review of some software that you're about to release and have it then find (and fix) FIVE release blockers, for an estimated (unsubsidized) cost of $1
I wrote about the sqlite-utils 4.0rc1 release a couple of weeks ago. Since we only have Claude Fable on our Max subscriptions for a few more days, I decided to see if it could help me get to a 4.0 sta
“That was fun, I’ll do it again when I get out,” said the 22-year-old Gambian violent offender after he stabbed a 55-year-old Italian man twenty times. And he will be released again. Perhaps to succee
There is a massive demand for processing files *in the agent loop*. The number of users submitting agent queries with file attachments is exponentially increasing over time. Our mission is to make Lit
“What has really happened is VC built the wrong AI, you’ll pay a high price for this error with your pension, and nobody will go to jail.” By far the shrewdest and most entertaining analyst of the AI
This post documents a Wordle game result (puzzle #1,841) solved in 4 attempts, showing the letter position feedback for each guess using the standard color-coded system (gray for incorrect letters, ye
As America turns 250, we put together 250 open AI milestones from the US: open models, datasets, demos, papers, and tools that helped shape the field. They go from attention is all you need, pytorch,
Better Models: Worse Tools Armin reports on a weird problem he ran into while hacking on Pi: The short version is that newer Claude models sometimes call Pi’s edit tool with extra, invented fields in
Hey, Fable: 'No like 'AAA' in what everyone online thinks AAA is, you know what I mean.' This was pretty funny, there are lootboxes, EULAs, achievements, useless graphical settings, elaborate boot scr
i often think about the irony of how 'tools for thought' people spent like a decade making cool pretty demos with canvases and then got completely mogged by low contrast poorly designed CLIs just winn
Indeed: if we actually had AGI we would not need forward deployed engineers. Two simultaneous AI narratives right now: 1) You can now do the work of 20 people, learn anything and create anything with
introducing tinyrouter i reverse engineered the routing architecture behind Skana AI's Fugu and built replication for open frontier models. it's a tiny ~10K parameter LLM router that learns which mode
spending the last week at @aidotengineer was awesome. too many great convos to cover them all, but jotted down some things that stood out: - Lots of discussion around open source models. I spoke with
Over the past week, a new fanworks movement has kicked off, with the aim to root out authors using generative AI. But the detection methods being implemented are questionable, and any fanfic writer co
The July 4th weekend All-In @theallinpod turned into a long argument about who owns the intelligence layer. The besties think enterprises just woke up to a trap they had been walking into, here's how
We've created a comprehensive Retrieval Harness for modern agentic retrieval in 2026. The harness provides a persistent data pipeline that can connect to a data source, index and update a large knowle
Midjourney has shown more of its futuristic medical scanner. It still hasn't shown much proof it works. The AI startup, best known for generating images, released a behind-the-scenes video of its dunk
arXiv:2607.01400v1 Announce Type: cross Abstract: Deep multimodal brain-encoding models now predict fMRI responses to naturalistic video with high accuracy. Whether their predicted neural signals also
arXiv:2607.02050v1 Announce Type: new Abstract: Motivated by the challenge of stabilizing a general unknown linear dynamical system (LDS) from observations, we study the natural prerequisite of online
arXiv:2607.02175v1 Announce Type: new Abstract: Multiple-choice medical benchmarks are increasingly saturated, and recent rubric-based evaluations such as HealthBench have shown that open-ended clinic
arXiv:2607.01935v1 Announce Type: new Abstract: Long term memory lets LLM agents act as persistent assistants, but user facts change. A useful memory system must know what is true now, what used to be
arXiv:2607.02141v1 Announce Type: new Abstract: Most LP-from-text benchmarks are static datasets of word problems written and labeled by hand. Once such a dataset is released, its size is fixed, its d
arXiv:2607.02131v1 Announce Type: cross Abstract: Restoring archival film remains a fundamentally challenging problem due to the absence of paired training data and the lack of standardized evaluation
arXiv:2604.08169v2 Announce Type: replace Abstract: Alignment in LLMs is more brittle than commonly assumed: misalignment can be induced by adversarial prompts, benign fine-tuning, emergent misalignme
arXiv:2602.03001v2 Announce Type: replace-cross Abstract: To maximize hardware utilization, modern machine learning systems typically employ large constant or manually tuned batch size schedules, rely
arXiv:2607.01418v1 Announce Type: cross Abstract: Organizations rolling out agentic command line tools like Anthropic's Claude Code and GitHub's Copilot CLI need to know who will try them, who will ke
arXiv:2607.01425v1 Announce Type: new Abstract: Understanding large, complex codebases, especially those with obfuscated structures and incomplete documentation, remains a significant challenge. Exist
arXiv:2607.01647v1 Announce Type: cross Abstract: Data science aims to derive actionable insights from heterogeneous raw data, unlocking the value of the massive amounts of data generated in modern so
arXiv:2602.19127v2 Announce Type: replace Abstract: With the rapid advancement of agent-based methods in recent years, Agentic RAG has undoubtedly become an important research direction. Multi-hop rea
arXiv:2607.02255v1 Announce Type: new Abstract: Memory for a long-horizon LLM agent is a contract about what each future decision is allowed to see. The simplest contract appends past observations, to
arXiv:2607.01934v1 Announce Type: cross Abstract: This work introduces AIriskEval-edu-db2, a new dataset designed to train and evaluate auditors based on LLMs for an explainable pedagogical risk asses
arXiv:2603.29466v2 Announce Type: replace-cross Abstract: Existing methods for quantifying predictive uncertainty in neural networks are either computationally intractable for large language models or
At the event 'The Briefing: AI for Science' earlier this week, Anthropic announced Claude Science, a new 'AI workbench for scientists' that pulls fragmented tools and datasets into one environment, an
arXiv:2607.02269v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) have demonstrated immense promise in Spatio-Temporal Video Grounding (STVG). However, current evaluation protocols are l
arXiv:2607.02104v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as cheap, scalable judges that compare candidate outputs pairwise -- to rank responses, select models
arXiv:2607.01973v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) are increasingly applied in medical tasks such as pathology description, report generation, and visual question answerin
arXiv:2607.02432v1 Announce Type: new Abstract: Scalable and reliable grading of command-line examinations remains a challenge in computing education, where rising enrolments make manual marking diffi
arXiv:2509.25136v3 Announce Type: replace Abstract: Activation-aware low-rank factorization techniques yield strong compression results but are generally confined to linear layers, while existing whit
arXiv:2607.02182v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit remarkable reasoning capabilities, but their task-specific fine-tuning is notoriously plagued by overconfidence,
arXiv:2607.01272v1 Announce Type: cross Abstract: Deploying 3D point cloud analysis in privacy-sensitive, resource-constrained settings faces two barriers: data cannot be centralized, and models must
arXiv:2607.01728v1 Announce Type: cross Abstract: Visual regression testing (VRT) is a standard quality assurance step in modern software release pipelines. On every change, it re-renders user interfa
arXiv:2607.01581v1 Announce Type: new Abstract: The capacity of Large Language Models (LLMs) to reason about pedagogical intent within instructional communication remains underexplored, particularly i
arXiv:2607.01313v1 Announce Type: cross Abstract: In practice, most commercial LLM providers do not publicly release details of underlying LLM architectures. However, prior work has shown that given l
arXiv:2607.02003v1 Announce Type: cross Abstract: Although neural networks are remarkably effective, their underlying optimization principles remain theoretically elusive, often characterized by non-c
arXiv:2607.01478v1 Announce Type: cross Abstract: We measured quantization-induced decision-boundary changes using local logit-margin radii, first-order boundary displacement, normal variation, slice-
arXiv:2607.01600v1 Announce Type: cross Abstract: As large language models (LLMs) are deployed as communicating agents, does inter-agent communication cause outputs to converge? We introduce BOUNDARY_
arXiv:2607.01239v1 Announce Type: cross Abstract: Character-level perturbations bypass safety alignment in modern LLMs despite leaving prompts human-readable. We identify and test a central structural
arXiv:2602.07267v2 Announce Type: replace Abstract: Evaluating the real-world capabilities of AI systems requires grounding benchmark performance in human-interpretable measures of task difficulty. Ex
arXiv:2607.02387v1 Announce Type: cross Abstract: NASA and its data centers hold thousands of geoscience datasets and tools like Worldview, Giovanni, the Science Discovery Engine, and Harmony. Finding
arXiv:2510.06288v4 Announce Type: replace Abstract: Today's AI models learn primarily through mimicry and refining, so it is not surprising that they struggle to solve problems beyond the limits set b