Quoting Boris Cherny
More than any of these eval scores, what is most exciting to me is something else: Opus 5 is our least prompt injectable model yet. It is a bit buried in the system card, but across PI evals and red t
Knowledge catalogue
More than any of these eval scores, what is most exciting to me is something else: Opus 5 is our least prompt injectable model yet. It is a bit buried in the system card, but across PI evals and red t
Ruff v0.16.0 Astral shipped a significant new version of their Ruff Python linting tool a few days ago on July 23rd. I noticed today because my various CI jobs all started failing thanks to new defaul
Pei Li / Bloomberg: Sources: DeepSeek told investors it is suspending its second funding round after remarks attributed to Liang Wenfeng on US-China AI competition went viral — DeepSeek has told prosp
The Gemma series is an amazing set of highly performant open-weight models. They have proven extremely effective in industrial settings where site-deployed agents need exactly this as a base for domai
The Opus 5 system card itself is a fun PDF to parse. It's 193 pages and stacked with labeled and unlabeled charts 📊 LlamaParse does a surprisingly good job on agentic (1.25c per page) and agentic plus
Very happy to support this on behalf of Google. We have long benefited from open source, are big contributors to open source and in fact have consistently made open weights models with Gemma available
arXiv:2607.21274v1 Announce Type: cross Abstract: We present CUP, a Greek book retrieval benchmark consisting of 868 catalog records and 104 expert-annotated queries with graded relevance judgments. W
arXiv:2607.20453v1 Announce Type: cross Abstract: Large language models show promise for clinical prediction, but zero-shot performance on specialized tasks is limited by incomplete domain knowledge,
A new model launch is not a product update ‼ It's one of three things, and you don't know which until you test it. Sometimes it's nothing: the model improved inside the same distribution, your harness
mr‑r0b0t announced that Claude Opus 5 is available at a 20 % discount through the Nous Portal. The model can be accessed via the Hermes Agent on the Nous Portal, as well as through OpenRouter and Anth
arXiv:2607.09424v3 Announce Type: replace-cross Abstract: We present Soofi S 30B-A3B, a sovereign, open-source Mixture-of-Experts (MoE) hybrid Mamba Transformer foundation model for German and English
arXiv:2607.21057v1 Announce Type: new Abstract: Text-based person retrieval faces a critical but under-explored challenge: the inherent uncertainty of query granularity in real-world scenarios. This p
arXiv:2607.21003v1 Announce Type: new Abstract: Ordinal Classification (OC) deals with classification tasks where the classes follow a natural order. Despite the progress in OC, many existing approach
arXiv:2607.20656v1 Announce Type: cross Abstract: Effective decision-making in complex and changing environments requires balancing short-term and long-term consequences. In reinforcement learning (RL
arXiv:2607.21482v1 Announce Type: new Abstract: Large language models (LLMs) and agents are now widely used tools in code development, with data typically sent to third-party cloud-based models. Their
arXiv:2607.20866v1 Announce Type: new Abstract: Generating realistic interior furniture layouts that strictly adhere to architectural constraints (e.g., walls, doors, and windows) remains a fundamenta
arXiv:2607.21306v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as tutors and thought partners, helping users reason through problems. While guidance from AI assis
arXiv:2607.20498v1 Announce Type: new Abstract: Large language models (LLMs) augmented with tools are emerging as autonomous agents capable of using Web engine, APIs, and code to solve complex, long-h
Am I supposed to interpret that it's better than Fable 5 from the benchmarks? Or it's ~close but cheaper? Introducing Claude Opus 5. It's a thoughtful and proactive model that comes close to the front
arXiv:2607.20943v1 Announce Type: cross Abstract: Quantum Phase Estimation (QPE) is a foundational algorithm for molecular ground-state energy estimation, but its deep circuit requirements make direct
An interesting thing I'm observing from the blog/system card is that on a good chunk of the reported benchmarks (~20-30% from a skim), Opus 5 max thinking leads to a degradation in performance compare
arXiv:2607.21292v1 Announce Type: new Abstract: We present a structured large-language-model-driven workflow for automated multi-variable control design from dynamic process models. The workflow decom
Announcing Fugu-Ultra v1.1 🐡 We’ve been thrilled by the reception to the Fugu model family. Thanks to everyone who tried it, shared feedback, and trusted Fugu with real work. Today, we’re releasing Fu
Anthropic: Anthropic launches Claude Opus 5, which it says comes close to Fable 5 performance at half the price; it is the new default model on Claude Max — Claude Opus 5 is available today. It's a th
Weeks after Anthropic's latest toe-to-toe with the US government, and days after an OpenAI security incident that dominated tech industry discussions, Anthropic on Thursday released its newest model,
Madison Mills / Axios: Anthropic says Opus 5 is the company's “most aligned model to date”; it is Anthropic's fourth model release in less than two months — Anthropic on Thursday is releasing Claude O
arXiv:2607.20536v1 Announce Type: new Abstract: Tool-use agents that address day-to-day digital tasks such as ordering groceries must not only operate applications, but also interact with the user, e.
arXiv:2607.20764v1 Announce Type: new Abstract: We introduce ARBIGRAPH, a benchmark generator for evaluating whether tool-assisted language agents can retain, update, compose, and discard task-relevan
arXiv:2607.20596v1 Announce Type: cross Abstract: Sparse autoencoder (SAE) features are used to interpret and steer large language models, yet whether a feature's causal role is stable across SAE fami
arXiv:2607.21461v1 Announce Type: new Abstract: Deep research requires agents to find answers that jointly satisfy multiple constraints. Discovering such answers is costly, whereas verifying a candida
As a joke I prompted Codex 'Build and run BenchBench, a benchmark of now good ai is at creating benchmarks. then figure out what benchbenchbench is and run that. and then write benchbenchbench up as a
arXiv:2607.20524v1 Announce Type: new Abstract: Mean cross-positional attention degradation is widely reported in transformer interpretability, yet whether it causally limits contextual retrieval rema
arXiv:2607.20694v1 Announce Type: new Abstract: Personal and organizational planning systems maintain two records that drift apart: what was planned (a task's effort budget) and what was done (a logge
audio.cpp again :) Release 0.4 is out. The headline this time is new high-quality TTS coverage plus GGUF becoming a first-class across the project. What’s new: Added Higgs Audio v3 TTS 4B, Fish Audio
arXiv:2607.21173v1 Announce Type: new Abstract: While automated research systems promise to accelerate empirical analysis, they are prone to silent failures: instances in which analysis code executes
arXiv:2607.20525v1 Announce Type: new Abstract: OpenAI's recent disproof of the Erdos unit distance conjecture marked a milestone for AI in mathematics. It also inspired another breakthrough: a human
arXiv:2607.20488v1 Announce Type: new Abstract: Multi-agent LLM frameworks typically fix their team topology at boot time. When an individual agent becomes overloaded at runtime, for example by mixing
arXiv:2607.21588v1 Announce Type: new Abstract: Learning effective robot manipulation policies requires diverse, high-quality demonstrations, yet existing data pipelines are often difficult to scale b
metal : add f16 type support to leaky relu (#25981) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFra
args: refactor mlock/mmap/directio into load-mode (#20834) args: overhaul mmap/mlock/dio into single arg Signed-off-by: Aaron Teo aaron.teo1@ibm.com docs: update docs with llama-gen-docs Signed-off-by
CUDA: fix external compilation of q1_0 MMQ (#25778) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFra
hexagon: fix Windows crash when op_poll is enabled (#26029) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) i
arXiv:2607.20476v1 Announce Type: new Abstract: We present an empirical benchmark evaluating how five large language models assess multisensor physical hazard data. Testing 60 scenarios across three c
arXiv:2607.20471v1 Announce Type: new Abstract: Personalization, the act of varying a message to induce action from a specific receiver while keeping sender, channel, and time fixed, has a long tradit
arXiv:2602.20114v2 Announce Type: replace-cross Abstract: Machine unlearning (MU) refers to the post-training capability to remove (the influence of) training examples that are incorrect, biased, or l
arXiv:2607.20479v1 Announce Type: new Abstract: Training probes to detect deceptive outputs from large language models is still an open problem. Recent work has demonstrated that detection probes fail
arXiv:2607.20434v1 Announce Type: cross Abstract: As the parameter size of language models continues to grow, effective model compression is required to reduce their computational and memory overhead.
I’m not affiliated with this project, but I’ve been running it recently and I’m surprised it hasn’t received more attention here: https://github.com/fewtarius/CachyLLama CachyLLama is a fork of llama.
arXiv:2607.20458v1 Announce Type: cross Abstract: Large language model (LLM) agents operating over extended dialogues accumulate vast amounts of information, yet existing memory systems either retain
https://reddit.com/link/1v5rvuq/video/bgmwc754i9fh1/player My goal was to create a benchmark to measure the spatial awareness and memory of models. Eventually, I came up with the simple idea of a maze
arXiv:2607.20518v1 Announce Type: new Abstract: AI agents are now capable of writing, compiling, and iteratively optimizing low-level operator kernels on different hardware platforms. Existing benchma
arXiv:2607.21340v1 Announce Type: new Abstract: In capital-markets workflows the question is rarely whether a large language model can produce a fluent draft, but whether the draft is bankable: defens
arXiv:2607.20737v1 Announce Type: new Abstract: Graph Neural Networks trained on heterogenous bipartite graphs form a common basis in recommendation systems. These graphs often express relations that
arXiv:2607.21196v1 Announce Type: cross Abstract: Ninety-Nine Prolog Problems (P-99) is a famous set of Prolog exercises. We solved the first thirty three just by prompting an LLM (Large Language Mode
arXiv:2607.20560v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) systems retrieve and integrate external knowledge to ground large language model (LLM) outputs. However, current RA
arXiv:2412.01748v2 Announce Type: replace Abstract: Complex dynamical systems, such as particle accelerators, often require intricate and time-consuming tuning procedures to achieve optimal performanc
arXiv:2607.20526v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in settings where fluent but incorrect answers can be costly. In these settings, accuracy alone i
arXiv:2607.21155v1 Announce Type: cross Abstract: Knowledge-Intensive Visual Question Answering (KI-VQA) benchmarks evaluate Vision-Language Models (VLMs) as multimodal knowledge assistants by requiri
arXiv:2607.20561v1 Announce Type: cross Abstract: LoRA adapters provide an efficient way to specialize a pretrained model for many downstream tasks, but deploying one adapter per task requires adapter
arXiv:2607.21016v1 Announce Type: new Abstract: Culture is lived through conversation, yet existing Indonesian cultural commonsense benchmarks evaluate LLMs on short and isolated prompts, stripping aw