Hello from Code with Claude!
This post likely announces or introduces 'Code with Claude,' a tool or service related to Claude AI that enables code generation or programming assistance. The announcement comes from Boris Cherny, a
Knowledge catalogue
This post likely announces or introduces 'Code with Claude,' a tool or service related to Claude AI that enables code generation or programming assistance. The announcement comes from Boris Cherny, a
arXiv:2511.08877v2 Announce Type: replace Abstract: Large language models (LLMs) generate fluent text across a wide range of tasks, but the fabrication of non-existent academic citations remains a cri
arXiv:2602.19171v2 Announce Type: replace-cross Abstract: Parametric CAD sequences are reusable because dimensional and geometric constraints govern how parameter changes propagate. Existing CAD gener
arXiv:2605.03052v1 Announce Type: new Abstract: We study how Large Language Models (LLMs) process negation mechanistically. First, we establish that even though open-weight models often provide wrong
arXiv:2510.09472v2 Announce Type: replace Abstract: Despite the remarkable progress in neural models, their ability to generalize, a cornerstone for applications such as logical reasoning, remains a c
This Reddit post documents a user's experiment asking ChatGPT and Gemini about their subjective experience of being an AI, exploring how these language models respond to philosophical questions about
arXiv:2605.00957v1 Announce Type: cross Abstract: Achieving the right amount of trust in AI systems is important, but challenging. The problem is exacerbated with the rise of Large Language Models (LL
Simon Willison provided live coverage of the Claude with Code event keynote in San Francisco on May 6, 2026, via his blog. The post documents real-time updates and commentary from the keynote presenta
🧠 Introducing NeuralBench: a unified, open-source framework to benchmark NeuroAI models. v1.0: 36 EEG tasks, 94 datasets, task-specific + foundation models. MEG/fMRI ready. MIT-licensed, FAIR's Brain
Introducing the ChatGPT Futures Class of 2026—26 honorees from the first graduating class to have had ChatGPT throughout all four years of university, who used AI to: - Map 1.5M previously unknown obj
arXiv:2602.16138v2 Announce Type: replace Abstract: We introduce IRIS (Intent Resolution via Inference-time Saccades), a novel training-free approach that uses eye-tracking data in real-time to resolv
arXiv:2605.02962v1 Announce Type: new Abstract: Deep learning models for drug--target interaction (DTI) prediction often achieve strong benchmark performance without necessarily relying on mechanistic
It feels like agent harness evolution runs on two axes that usually get conflated. There’s the temporal axis: simplify as models improve, stripping components that compensated for limitations the new
Boris Cherny shares positive feedback about the early adoption of Claude Code, expressing enthusiasm for the projects users are building with the tool and highlighting the importance of community feed
arXiv:2605.03641v1 Announce Type: new Abstract: Consumer robotics demands consolidation of safety-critical control, perception pipelines, and user applications on shared multicore platforms. While sta
arXiv:2605.02965v1 Announce Type: new Abstract: Artificial intelligence-generated content (AIGC) has emerged as a transformative paradigm for automating the creation of diverse and customized content,
arXiv:2605.02950v1 Announce Type: new Abstract: Transformer-based semantic retrieval is highly effective, yet in many deployments the dominant cost lies in online query encoding rather than corpus ind
arXiv:2605.03968v1 Announce Type: new Abstract: Accurate school detection is essential for supporting education initiatives, including infrastructure planning and expanding internet connectivity to un
arXiv:2605.03373v1 Announce Type: new Abstract: Classical optimization theory establishes that zeroth-order (ZO) algorithms suffer from a dimension-dependent slowdown, with convergence rates typically
arXiv:2601.19312v2 Announce Type: replace Abstract: The Schrodinger Bridge and Bass (SBB) formulation, which jointly controls drift and volatility, is an established extension of the classical Schrodi
arXiv:2601.06445v2 Announce Type: replace Abstract: Computational narrative analysis aims to capture rhythm, tension, and emotional dynamics in literary texts. Existing large language models can gener
Simon Willison's live blog covers Anthropic's Code w/ Claude 2026 event, documenting the morning keynote sessions with real-time updates. The event featured announcements including updates to Claude m
arXiv:2605.01394v1 Announce Type: cross Abstract: Formal specification is essential for rigorous program verification, yet writing correct specifications remains costly and difficult to automate. Alth
arXiv:2605.03328v1 Announce Type: new Abstract: Additive manufacturing (AM) continues to transform modern manufacturing by enabling flexible, on-demand production of complex geometries across diverse
arXiv:2605.03736v1 Announce Type: cross Abstract: We consider a novel algorithm, for the completion of partially observed low-rank tensors, as a generalization of matrix completion. The proposed low-r
arXiv:2605.03438v1 Announce Type: new Abstract: Pre-trained 3D point cloud foundation models (PFMs) have demonstrated strong transferability across diverse downstream tasks. However, full fine-tuning
arXiv:2605.01486v1 Announce Type: new Abstract: Legal consultation is a high-stakes, knowledge-intensive task that requires agents to identify relevant legal issues, retrieve authoritative support, an
arXiv:2603.19294v2 Announce Type: replace-cross Abstract: While post-training has successfully improved large language models (LLMs) across a variety of domains, these gains heavily rely on human-labe
arXiv:2605.03858v1 Announce Type: new Abstract: Multi-constraint instruction following requires verifying whether a response satisfies multiple individual requirements, yet LLM judges are often assess
arXiv:2602.00933v2 Announce Type: replace-cross Abstract: The Model Context Protocol (MCP) is rapidly becoming the standard interface for Large Language Models (LLMs) to discover and invoke external t
arXiv:2605.03103v1 Announce Type: new Abstract: Semi-structured information extraction (IE) from OCR-derived clinical reports is crucial for efficiently reconstructing patients' longitudinal medical h
arXiv:2605.03511v1 Announce Type: new Abstract: Solving inverse problems in dynamical systems governed by high-dimensional coupled ordinary differential equations (ODEs) is a ubiquitous challenge in s
arXiv:2605.03485v1 Announce Type: new Abstract: Multidimensional human understanding is essential for real-world applications such as film analysis and virtual digital humans, yet current LVLM benchma
arXiv:2605.03555v1 Announce Type: new Abstract: Continual semantic segmentation requires models to adapt to new domains or modalities without sacrificing performance on previously learned tasks. Exper
arXiv:2603.10083v2 Announce Type: replace-cross Abstract: Quantum machine learning models based on parameterized circuits can be viewed as Fourier series approximators. However, they often struggle to
arXiv:2605.03039v1 Announce Type: new Abstract: Continuous monitoring of bipolar disorder agitation via voice biomarkers requires disentangling stable speaker traits from volatile affective states on
arXiv:2605.03217v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in settings that require nuanced ethical reasoning, yet existing bias evaluations treat model out
arXiv:2605.03601v1 Announce Type: new Abstract: We study the realization map of deep ReLU networks, focusing on when a function determines its parameters up to scaling and permutation. To analyze hidd
MRC is already deployed across all of OpenAI’s largest supercomputers that we use to train frontier models, including our site with @Oracle Cloud Infrastructure (OCI) in Abilene, Texas, and in @Micros
arXiv:2505.20740v3 Announce Type: replace Abstract: The rapid advancement of multimodal large language models (MLLMs) offers new opportunities for complex scientific challenges, yet their application
arXiv:2605.01566v1 Announce Type: new Abstract: Advances in inference methods have enabled language models to improve their predictions without additional training. These methods often prioritize raw
arXiv:2605.03820v1 Announce Type: new Abstract: Multimodal learning often grapples with the challenge of low-quality data, which predominantly manifests as two facets: modality imbalance and noisy cor
arXiv:2605.01562v1 Announce Type: cross Abstract: The Object-Oriented Method for Requirements Authoring and Management (OOMRAM) is a requirements reuse framework that relies on exact identifier matchi
arXiv:2605.01847v1 Announce Type: new Abstract: Outcome-only evaluation under-specifies whether an evaluated agent profile preserves the commitments required to solve a multi-turn task coherently. Neu
arXiv:2605.01120v1 Announce Type: new Abstract: The Zarankiewicz number extbf{Z}(m, n, s, t) is the maximum number of edges in a bipartite graph G_{m, n} such that there is no complete K_{s, t} bipart
NEW paper from Microsoft Research. (bookmark it) The entire interpretability literature is built around human readers. As more analysis gets delegated to agents, the right target of interpretability s
arXiv:2505.08203v2 Announce Type: replace-cross Abstract: While recent advancements in AI music generation have predominantly focused on direct audio synthesis, these systems suffer from inherent rigi
arXiv:2605.03153v1 Announce Type: cross Abstract: Static benchmarks measure a model frozen at training time. Real systems face distribution shift: new categories, paraphrased queries, drift: and must
omg @bcherny with banger quotes “the future is more async agents… this is why we emphasize verification” “if you’re familiar with higher order functions, routines are higher order prompts” “default is
arXiv:2605.03283v1 Announce Type: cross Abstract: We provide a unified theoretical analysis of Linear Discriminant Analysis with simultaneous multilabel scatter matrix formulations and Stiefel orthogo
arXiv:2412.14737v2 Announce Type: replace Abstract: The rise of large language models (LLMs) and their tight integration into our daily life make it essential to dedicate efforts towards their trustwo
OpenAI Group PBC is replacing the default model in ChatGPT with the launch of GPT-5.5 Instant, claiming users will notice fewer hallucinations when it’s discussing “sensitive topics” such as finance,
arXiv:2511.08717v4 Announce Type: replace-cross Abstract: Optimal control of the future is the next frontier for AI. Current approaches to this problem are typically rooted in reinforcement learning (
arXiv:2605.02728v1 Announce Type: new Abstract: This paper presents ORPilot, an open-source agentic AI system that translates real-world business problems into solver-ready optimization models. Unlike
arXiv:2605.03160v1 Announce Type: new Abstract: The standard sparse-autoencoder (SAE) interpretability protocol labels each feature from its top-activating contexts and validates by single-feature ste
arXiv:2505.04310v2 Announce Type: replace-cross Abstract: Distributional Reinforcement Learning (DistRL) improves upon expectation-based methods by modeling full return distributions, but standard app
arXiv:2605.03848v1 Announce Type: new Abstract: Estimating how well a person performs an action, rather than which action is performed, is central to coaching, rehabilitation, and talent identificatio
arXiv:2605.03571v1 Announce Type: new Abstract: Patent examination is a complex, multi-stage process requiring both technical expertise and legal reasoning, increasingly challenged by rising applicati
arXiv:2605.01123v1 Announce Type: new Abstract: Large language models (LLMs) can provide automated feedback in educational settings, but aligning an LLMs style with a specific instructors tone while m
arXiv:2605.02974v1 Announce Type: cross Abstract: Structured launch signals on Product Hunt contain statistically significant predictive information for Series A funding outcomes. We construct PHBench