Grok 4.3 - excellent intelligence per unit cost
Grok 4.3 - excellent intelligence per unit cost xAI has launched Grok 4.3, achieving 53 on the Artificial Analysis Intelligence Index with improved agentic performance, ~40% lower input price, and ~60
Knowledge catalogue
Grok 4.3 - excellent intelligence per unit cost xAI has launched Grok 4.3, achieving 53 on the Artificial Analysis Intelligence Index with improved agentic performance, ~40% lower input price, and ~60
Grok Voice is used by Starlink right now Grok Voice brutally dominates the top of the τ-voice Bench Grok scores 67.3%, while Gemini sits at 43.8% and GPT Realtime at 35.3% This is a massive lead over
here’s the essay: “Consciousness is not about what a creature says, but how it *feels*. And there is no reason to think that Claude feels anything at all. I am sure Claude can draw on its training dat
I added a new feature to my blog (built entirely on my phone with Claude code for web) that imports my iNaturalist photos and adds them to my site's overall timeline https://simonwillison.net/2026/May
Neural networks were declared scientifically dead in 1987. A French PhD student bet his entire career on them anyway ~ and won. 🤯 >Meet Yann LeCun 🇫🇷 >Paris-born. PhD from Sorbonne in 1987. >Joined Be
This article by cognitive scientist Gary Marcus likely critiques Richard Dawkins' views or statements regarding Claude (an AI system), examining whether Dawkins may be mistaken or overly optimistic ab
see also this essay: “Consciousness is not about what a creature says, but how it *feels*. And there is no reason to think that Claude feels anything at all. I am sure Claude can draw on its training
OpenAI announced a community contest inviting users to create custom Codex pets using a /hatch command, with 10 winning entries receiving 30 days of ChatGPT Pro as prizes. The promotion encouraged use
/elsewhere/sightings/ I have a new camera (a Canon R6 Mark II) so I'm taking a lot more photos of birds. I share my best wildlife photos on iNaturalist, and based on yesterday's successful prototype I
small milestone: uninstalled the chatgpt app. codex is strict superset now! found something cool - among frontier models, @xai @grok 4.30 is the most intelligence per dollar you can get, beating even
(Sorry, after seeing so many of these, could not resist): 🚨 BREAKING: Google just dropped a NEW paper that completely deletes RNNs from existence. No recurrence. No convolutions. Nothing. Just one mec
The Harness is a Context Manager on Behalf of the Model What happens when the context window fills up and who decides? This decision is external to the model - The Harness designer must have some opin
This is actually a version of an alignment problem. Humans have background beliefs (don’t waste large sums of money without telling me) and Claude doesn’t respect those. Caveat emptor. THIS GUY ACCIDE
We are honored to be featured in the latest @TwoMinutePapers video! You all can watch the full video here: https://youtu.be/QzZ4VwDHAT4 Here’s a short clip from it: Media What happens when you put com
arXiv:2604.27421v1 Announce Type: cross Abstract: Large Language Models (LLMs) are now widely used for query reformulation and expansion in Information Retrieval, with many studies reporting substanti
arXiv:2604.28126v1 Announce Type: cross Abstract: Diffusion models offer superior generation quality at the expense of extensive sampling steps. Distillation methods, with Distribution Matching Distil
arXiv:2604.28177v1 Announce Type: new Abstract: We introduce AEGIS, A holistic benchmark for Evaluating forensic analysis of AI-Generated academic ImageS. Compared to existing benchmarks, AEGIS featur
arXiv:2604.28078v1 Announce Type: new Abstract: Despite rapid advances in photorealistic video generation, real-world applications such as filmmaking require video aesthetics, e.g., harmonious colors
arXiv:2602.11897v3 Announce Type: replace-cross Abstract: Cybersecurity decision-making increasingly occurs in environments characterized by uncertainty, partial observability, and adversarial manipul
arXiv:2604.17460v2 Announce Type: replace-cross Abstract: AI coding assistants have proliferated rapidly, yet structured pedagogical frameworks for learning these tools remain scarce. Developers face
arXiv:2604.26969v1 Announce Type: cross Abstract: Modern large-scale recommendation systems are typically constructed as multi-stage pipelines, encompassing pre-ranking, ranking, and re-ranking phases
This article discusses the emerging landscape of AI agents specialized for different domains, contrasting Codex's capabilities for knowledge work and structured tasks with Claude's strengths in creati
Lauren Forristal / TechCrunch: Amazon debuts “Join the chat”, an AI-powered feature that lets users ask questions about products and get conversational audio responses generated in real time — Amazon
arXiv:2604.27550v1 Announce Type: cross Abstract: Privacy policies are essential for users to understand how service providers handle their personal data. However, these documents are often long and c
arXiv:2604.27543v1 Announce Type: new Abstract: Evaluating English ASR systems for conversational AI applications remains difficult, as many publicly available corpora are either pre-segmented into sh
arXiv:2604.27582v1 Announce Type: new Abstract: Surgical resection remains the only potentially curative treatment for pancreatic ductal adenocarcinoma (PDAC), and eligibility depends on accurate asse
arXiv:2604.27720v1 Announce Type: new Abstract: Deploying vision-language models (VLMs) in clinical settings demands auditable behavior under realistic failure conditions, yet the failure landscape of
arXiv:2604.28109v1 Announce Type: new Abstract: Model merging has attracted attention as an effective path toward multi-task adaptation by integrating knowledge from multiple task-specific models. Amo
arXiv:2603.01444v2 Announce Type: replace Abstract: Synthetic data generation is an important capability for privacy-preserving data sharing, system benchmarking and test data provisioning. For mixed-
arXiv:2604.27089v1 Announce Type: new Abstract: Large-language-models (LLMs) demonstrate enormous utility in long-context tasks which require processing prompts that consist of tens to hundreds of tho
arXiv:2604.27253v1 Announce Type: new Abstract: Recent advances in multimodal large language models (LLMs) have revolutionized web agents that can automate complex tasks on websites. However, their ac
arXiv:2604.26986v1 Announce Type: new Abstract: We introduce a novel task of digital battery passport (DBP) conformance classification and introduce the first public benchmark for the task: BatteryPas
arXiv:2603.10252v2 Announce Type: replace-cross Abstract: Bayesian hierarchical models are frequently used in practical data analysis contexts. One interpretation of these models is that they provide
arXiv:2604.27394v1 Announce Type: cross Abstract: Conditional Average Treatment Effect (CATE) estimation in practice demands three properties simultaneously: heterogeneous effects au(x), calibrated un
arXiv:2604.27124v1 Announce Type: new Abstract: Training stable biological foundation models requires rethinking attention mechanisms: we find that using sigmoid attention as a drop in replacement for
arXiv:2604.27006v1 Announce Type: cross Abstract: Context: Study screening in systematic literature reviews is costly, inconsistency-prone, and risk-asymmetric, since false negatives can compromise va
arXiv:2604.27920v1 Announce Type: cross Abstract: Preserving affective nuance remains a challenge in Machine Translation (MT), where semantic equivalence often takes precedence over emotional fidelity
arXiv:2604.27405v1 Announce Type: cross Abstract: We adapted the Reliable Change Index (RCI; Jacobson and Truax, 1991) from clinical psychology to item-level LLM version comparison on 2,000 MMLU-Pro i
This post appears to be a sports-related social media update about an upcoming Hawks game, though the message is incomplete and cuts off mid-sentence before providing specific details about the score
arXiv:2604.27308v1 Announce Type: cross Abstract: Parameter-efficient fine-tuning (PEFT) methods face a tradeoff between adapter size and expressivity: ultra-low-parameter adapters are confined to fix
OpenAI announced a feature that allows users to quickly import their existing workflow configurations, including settings, plugins, agents, and project configurations, into Codex with minimal effort.
arXiv:2602.10140v2 Announce Type: replace-cross Abstract: Large language models (LLMs) can now synthesize non-trivial executable code from textual descriptions, raising an important question: can LLMs
arXiv:2604.26959v1 Announce Type: cross Abstract: Integrating large language models (LLMs) into patient-facing healthcare systems offers significant potential to improve access to medical information.
arXiv:2602.07915v2 Announce Type: replace-cross Abstract: Causal discovery from time series is a fundamental task in machine learning. However, its widespread adoption is hindered by a reliance on unt
arXiv:2604.28082v1 Announce Type: new Abstract: Fine-tuning large language models (LLMs) on narrowly misaligned data generalizes to broadly misaligned behavior, a phenomenon termed emergent misalignme
arXiv:2604.27415v1 Announce Type: new Abstract: With the rapid advancement of semiconductor technology, Electronic Design Automation (EDA) has become an increasingly knowledge-intensive and document-d
arXiv:2604.27043v1 Announce Type: new Abstract: Today's AI assistants such as OpenClaw are designed to handle context effectively, making context learning an increasingly important capability for mode
arXiv:2604.28139v1 Announce Type: cross Abstract: LLM agents are expected to complete end-to-end units of work across software tools, business services, and local workspaces. Yet many agent benchmarks
Code with Claude, our developer conference, returns next week. Whether you're just getting started with Claude Code or you've been building for a while, there's a session for you. Register for the liv
Yohei Nakajima announced the launch of Cofounder 2 on May 4th via X. Cofounder is an AI agent tool designed to assist with business and startup tasks. The announcement was shared on social media to in
arXiv:2604.27389v1 Announce Type: cross Abstract: In recent years, Multimodal Large Language Models (MLLMs) have achieved remarkable progress on a wide range of multimodal benchmarks. Despite these ad
arXiv:2604.27707v1 Announce Type: new Abstract: Current agentic memory systems (vector stores, retrieval-augmented generation, scratchpads, and context-window management) do not implement memory: they
arXiv:2604.27137v1 Announce Type: new Abstract: This paper introduces a systematic evaluation framework grounded in the Interagency Language Roundtable (ILR) Skill Level Descriptions and applies it to
OpenAI announced a migration feature allowing users to switch to Codex directly through both the Codex app and command-line interface (CLI). The announcement indicates that Codex migration tools are n
arXiv:2604.27604v1 Announce Type: new Abstract: We introduce SPUR, a comprehensive benchmark for scientific experimental image perception, understanding, and reasoning, comprising 4,264 question-answe
arXiv:2604.26962v1 Announce Type: cross Abstract: Education represents one of the most promising real-world applications for Large Language Models (LLMs). However, conventional tutoring systems rely o
arXiv:2604.28118v1 Announce Type: cross Abstract: Transformer models are widely deployed in critical AI applications, yet faults in their attention mechanisms, projections, and other internal componen
arXiv:2604.27547v1 Announce Type: new Abstract: Fine-tuning large language models (LLMs) for domain-specific tasks requires training datasets that comprehensively cover the target capabilities a pract
arXiv:2603.22018v2 Announce Type: replace Abstract: Ensuring consistency between research papers and their corresponding software code implementations is a fundamental prerequisite for guaranteeing th
arXiv:2603.09881v2 Announce Type: replace Abstract: Speech Large Language Models (SLLMs) have rapidly expanded, supporting a wide range of tasks. These models are typically evaluated using text prompt