Towards Tight Bounds for Streaming Attention
arXiv:2606.07205v1 Announce Type: cross Abstract: The attention mechanism is a cornerstone of modern transformer architectures. However, its expressive power comes at the cost of quadratic runtime and
Knowledge catalogue
arXiv:2606.07205v1 Announce Type: cross Abstract: The attention mechanism is a cornerstone of modern transformer architectures. However, its expressive power comes at the cost of quadratic runtime and
arXiv:2606.04550v1 Announce Type: cross Abstract: E-commerce recommender systems strongly influence which products users consider and purchase, yet sustainability signals such as Product Carbon Footpr
arXiv:2606.06960v1 Announce Type: new Abstract: Experience-based self-evolution is crucial for LLM agents, but existing benchmarks often assume explicit goals, stable task patterns, and clear feedback
arXiv:2601.23204v2 Announce Type: replace Abstract: Time series data are integral to critical applications across domains such as finance, healthcare, transportation, and environmental science. While
arXiv:2403.05532v2 Announce Type: replace-cross Abstract: We introduce Tune without Validation (Twin), a simple and effective pipeline for tuning learning rate and weight decay of homogeneous classifi
arXiv:2606.06673v1 Announce Type: new Abstract: Sparse rewards and heterogeneous task sequences remain persistent challenges in Reinforcement Learning (RL), often resulting in slow convergence, weak g
arXiv:2603.25670v3 Announce Type: replace Abstract: Safety monitoring is essential for Cyber-Physical Systems (CPSs). However, unsafe events are rare in real-world CPS operations, creating an extreme
arXiv:2606.06934v1 Announce Type: new Abstract: We analyze generalization error, uniform stability, and uniform argument stability of gradient descent (GD) and stochastic gradient descent (SGD) over d
arXiv:2606.07514v1 Announce Type: new Abstract: In this work, we focus on extending SHARP, the popular photorealistic view synthesis method, for universal monocular rendering across a continuum of cam
arXiv:2606.06622v1 Announce Type: new Abstract: We introduce UnpredictaBench, an evaluation that tests the ability of large language models (LLMs) to capture true underlying distributions. As LLMs are
arXiv:2606.07474v1 Announce Type: new Abstract: Unsupervised Continual Learning (UCL) aims to enable neural networks to learn sequential tasks without labels or access to past data. A major challenge
arXiv:2606.07167v1 Announce Type: cross Abstract: Meaningful multilingual evaluation must test models in the target language and educational context. Urdu, spoken by more than 230 million people, lack
arXiv:2606.07338v1 Announce Type: new Abstract: Vision-language driving models increasingly use reasoning supervision to bridge perception, prediction, and planning, but existing driving rationales ar
arXiv:2606.06819v1 Announce Type: new Abstract: Reasoning Video Object Segmentation (RVOS) demands a sophisticated integration of temporal dynamics, spatial details, and linguistic reasoning to achiev
we just shipped support for rubrics in deepagents ✅ give your agent a clear definition of what 'done' looks like, and force it to run in a loop until said goal is complete this is similar to /goal in
arXiv:2606.06784v1 Announce Type: cross Abstract: Public social media posts can reveal private information through weak cues scattered across text, images, or metadata. Such leakage is often cumulativ
arXiv:2606.07237v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used in healthcare for tasks such as clinical question answering, diagnosis support, and report summariz
arXiv:2606.07171v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) enable flexible instruction-driven image editing, but privacy risks arise when user images expose diverse and u
arXiv:2606.06745v1 Announce Type: new Abstract: Reasoning Large Language Models can improve problem-solving performance through deliberative inference, but invoking slow reasoning for every input is c
When we first demoed Claude Code internally, it got two reactions on Slack. A year after GA, @_catwu and I sat down to talk about what's changed: why I use auto mode instead of plan mode, how routines
arXiv:2606.06509v1 Announce Type: cross Abstract: Numerous medical imaging problems must be solved under limited labels and constrained compute, yet it remains unclear whether performance gains are dr
This post documents a Wordle game result where the player solved puzzle #1,814 in four attempts, using color-coded emoji tiles (⬛ for incorrect letters, 🟨 for correct letters in wrong positions, 🟩 for
arXiv:2606.06538v1 Announce Type: new Abstract: In real-world applications, models are expected to perform reliably across diverse settings. Yet, many existing multimodal benchmarks expand task types
Jose Antonio Lanz / Decrypt: Xiaomi claims MiMo-V2.5-Pro-UltraSpeed tops 1K tokens/second, a first at the 1T-parameter scale, using a standard 8-GPU commodity node; API trial starts June 9 — Most peop
arXiv:2601.12359v1 Announce Type: cross Abstract: Prompt injection attacks have become an increasing vulnerability for LLM applications, where adversarial prompts exploit indirect input channels such
Aston Martin F1 team secured a point at the 2026 Monaco Grand Prix, acknowledging contributions from their team members, partners, and supporters. The post expresses gratitude for the collaborative ef
Release: datasette-agent-edit 0.1a0 I'm planning several plugins for Datasette Agent which can make edits to existing pieces of text - things like collaborative Markdown editing, updating large SQL qu
SpaceX's Falcon 9 rocket successfully launched 21 Starlink internet satellites and two Starshield military satellites from a California launch facility. Starlink satellites are part of SpaceX's global
Grep timeout issue fixed in latest Grok Build Grok Build update just released v0.2.31 Release Notes: Bug Fixes: • Marketplace skills without proper descriptions are now hidden from listings instead of
Have been extensively testing Claude Workflows this weekend, with the best model possible. Threw it at my whole code base, combing for bugs. 144 found and fixed! Geez... It is a large code base, for s
LOL. Even Claude sees through Hinton’s nonsense. (though see the articles I posted earlier, for converging sources I put more weight on) This is my prompt and this is Claude's response: ME: What do yo
see eg https://www.scmp.com/tech/tech-trends/article/3271858/ai-race-alibaba-tencent-quickly-adopt-metas-new-llama-31-model-amid-excitement and https://medium.com/the-endless-forge/zuckerbergs-llama-f
The first wave of AI-native applications is wrapping tokens and providing in-app agents. As agent usage centralizes around core apps (e.g. Claude Code, Codex), there's this emerging wave of building s
There was an inflection point recently where the tide shifted to model pickers and OSS Mix of tokenmaxxing/cost fatigue, nemotron coalition, harness step functions, brains/claws, etc Long live those w
Krea 2, Krea's first foundation image model built from scratch, was announced on May 12, 2026 , with a focus on aesthetics, style transfer, and creative control . Krea 2 became available to everyone s
This post documents a Wordle game result shared by Anthropic on X (Twitter), showing puzzle #1,813 solved in 4 attempts with a final correct answer displayed through colored emoji tiles (green indicat
Zuckerberg and LeCun’s unilateral decision to open source Llama likely (partly) catalyzed China’s AI industry — and may have done truly massive harm to American business interests. We are now starting
arXiv:2606.05548v1 Announce Type: cross Abstract: The rapid proliferation of Agent Development Kits (ADKs), SDK-level frameworks for building LLM-powered autonomous agents, has outpaced any empirical
arXiv:2606.06448v1 Announce Type: new Abstract: LLM agents are increasingly deployed on long-horizon tasks requiring sustained reasoning over extended interaction histories. Realizing this at scale re
arXiv:2606.05658v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by grounding their responses in external knowledge, but conventional pipeli
arXiv:2606.05296v1 Announce Type: cross Abstract: LLM agents operate in two distinct regimes: open-weight agents amenable to reinforcement learning (RL) and black-box agents whose behaviour must be co
arXiv:2606.05986v1 Announce Type: cross Abstract: Existing learning-based detectors for Solidity smart-contracts reduce vulnerability detection to syntactic pattern matching within single functions, y
arXiv:2606.06462v1 Announce Type: new Abstract: Benchmarks are fundamental for evaluating and advancing LLMs and MLLMs by providing standardized and explicit measures of performance. However, their co
arXiv:2606.05692v1 Announce Type: cross Abstract: Deep learning has enabled significant advances in time-series causal inference, yet progress remains constrained by the lack of realistic benchmarks w
arXiv:2606.05682v1 Announce Type: new Abstract: Demand for low-precision inference, including NVFP4-based approaches, has grown as large language models are increasingly deployed in latency and cost c
🚨BREAKING: Miguel Bosé, the biggest Spanish-language pop star of the last few decades, has just released a video taking a knee and putting his hand over his heart in honour of Henry Nowak This has now
arXiv:2606.05445v1 Announce Type: new Abstract: We dream of AI agents that can read arbitrary designs and construct real-world objects from reusable building blocks. As a first step toward this vision
arXiv:2606.05383v1 Announce Type: cross Abstract: Can artificial intelligence (AI) refute economic theory? I document experiments in which I asked several AI models (Gemini, Refine, Claude, and ChatGP
arXiv:2606.05792v1 Announce Type: new Abstract: TLA+ has supported industrial verification at companies such as Amazon and Microsoft, yet writing correct TLA+ specifications from natural language stil
arXiv:2512.15231v3 Announce Type: replace Abstract: The automated and intelligent processing of massive remote sensing (RS) datasets is critical in Earth observation (EO). Existing automated systems a
arXiv:2606.05966v1 Announce Type: cross Abstract: Understanding and reasoning about the physical world is the foundation of intelligent behavior, yet state-of-the-art vision-language models (VLMs) sti
Chip Export rules: made China/Huawei stronger, accomplished relatively little. USG stakes in US AI: will make Mistral and other “sovereign AI” efforts stronger, freak out the rest of the world, escala
arXiv:2606.06252v1 Announce Type: new Abstract: Recent work moves intermediate reasoning from natural-language traces into latent or cache-level representations to reduce token overhead and avoid a di
arXiv:2606.06099v1 Announce Type: new Abstract: Whether Large Language Models (LLMs) exhibit covert psychological manipulation in complex human-AI interactions has garnered increasing safety concerns.
arXiv:2606.05704v1 Announce Type: new Abstract: Recent Large Language Models (LLMs) have shown impressive reasoning abilities; but they are still susceptible to hallucinations, intermediate reasoning
arXiv:2606.05606v1 Announce Type: cross Abstract: LLM post-training often relies on reinforcement learning methods that sample multiple rollouts per prompt, yet most existing approaches use a fixed ro
arXiv:2510.11974v2 Announce Type: replace-cross Abstract: Cyber Threat Intelligence (CTI) is foundational to modern cybersecurity, enabling organizations to proactively defend against evolving threats
arXiv:2606.05679v1 Announce Type: cross Abstract: Agents increasingly generate SQL, orchestrate pipelines, and automate data analysis on behalf of users. While recent work improves query correctness,
arXiv:2606.05670v1 Announce Type: new Abstract: Does adding more agents help an LLM workflow once compared systems share the same benchmark loader, tool access, answer contract, usage accounting, and
arXiv:2602.13255v2 Announce Type: replace Abstract: We present DPBench, a benchmark for evaluating coordination in multi-agent systems built from large language models. Existing benchmarks measure tas