Guiding Token-Sparse Diffusion Models
arXiv:2601.01608v2 Announce Type: replace Abstract: Diffusion models deliver high quality in image synthesis but remain expensive during training and inference. Recent works have leveraged the inheren
Knowledge catalogue
arXiv:2601.01608v2 Announce Type: replace Abstract: Diffusion models deliver high quality in image synthesis but remain expensive during training and inference. Recent works have leveraged the inheren
arXiv:2605.26835v1 Announce Type: new Abstract: LLM-based multi-agent systems have been widely adopted for knowledge retrieval and report generation, synthesizing known information through web search
arXiv:2507.20758v2 Announce Type: replace Abstract: Chain-of-Thought (CoT) prompting significantly enhances model reasoning, yet its internal mechanisms remain poorly understood. We analyze CoT's oper
arXiv:2605.26807v1 Announce Type: cross Abstract: LLMs can now produce full HTML pages, but many of those pages are only superficially correct: they render once, then fail under scroll, hover, click,
arXiv:2605.26575v1 Announce Type: new Abstract: Multilingual embedding models are deployed under the assumption that cross-lingual retrieval is symmetric: if a query in language A retrieves its transl
Anthropic are strongly rumored to be about to have their first profitable quarter. Stories are circulating of companies surprised at how expensive their LLM bills are becoming from usage by their staf
If we had done everything I suggested in my 2020 arXiv article “The Next Decade in AI”, we might actually have reached AGI by now. In the last three years, after a detour driven by the false promise o
Mark Carney, former Bank of Canada Governor and Canadian Prime Minister, publicly endorsed Cohere, a Canadian AI company, suggesting it would become a significant player in the artificial intelligence
arXiv:2503.14359v2 Announce Type: replace Abstract: User engagement is greatly enhanced by fully immersive multi-modal experiences that combine visual and auditory stimuli. Consequently, the next fron
arXiv:2605.26175v1 Announce Type: cross Abstract: Low-bit activation quantization remains a major bottleneck in efficient large language model (LLM) deployment. The difficulty is not only that activat
arXiv:2601.00575v2 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated significant advancements in reasoning and code generation, but efficiently creating new benchmarks to
arXiv:2605.26307v1 Announce Type: cross Abstract: Software-Defined Networking (SDN) provides flexible and programmable network management; however, its centralized control architecture remains highly
arXiv:2506.15199v3 Announce Type: replace Abstract: While there are many applications of ML to scientific problems that look promising, visuals can be deceiving. Using numerical analysis techniques, w
arXiv:2605.26520v1 Announce Type: cross Abstract: While vision-language models (VLMs) have exhibited multi-turn visual reasoning capabilities, their reasoning trajectories remain relatively shallow an
aside_block <ListValue: [StructValue([('title', 'Summary of today’s news'), ('body', <wagtail.rich_text.RichText object at 0x7fb0f516f910>), ('btn_text', ''), ('href', ''), ('image', None)])]> AI-powe
Introducing Runway MCP. Now you can connect Runway directly into Claude, ChatGPT, Cursor, Replit and more. Generate polished images and videos with state-of-the-art models, like Gen-4.5, Seedance 2.0,
InvokeAI 6.13 is the largest community-driven release of the software, adding full support for Anima & Qwen Image models, API model integration (such as GPT Image), and new features including Prompt E
arXiv:2605.27074v1 Announce Type: new Abstract: Recent multimodal large language models (MLLMs) achieve strong performance on reactive question answering, but real-world streaming assistants require p
ITBench-AA is a new benchmark developed by Artificial Analysis and IBM that evaluates frontier AI models on agentic enterprise IT tasks, with results showing that current leading models score below 50
arXiv:2605.26731v1 Announce Type: new Abstract: A prevalent assumption in LLM agent deployment holds that more structured harnesses universally improve reliability, and that higher-capability models n
jesus, qwen3-tts is FANTASTIC. going for full local stt/tts/llm with parakeet, qwen3-tts, and gemma 4 via llama.cpp for my little robot. excite, excite! https://huggingface.co/Qwen/Qwen3-TTS-12Hz-1.7B
arXiv:2605.26329v1 Announce Type: new Abstract: Current benchmarks for occupational AI agents are scoped primarily by economic values, telling a replacement story. We introduce JobBench, which evaluat
arXiv:2605.26955v1 Announce Type: cross Abstract: As large language models (LLMs) are increasingly deployed to users around the world, they are integrated into everyday tasks across diverse cultural c
arXiv:2511.14993v3 Announce Type: replace-cross Abstract: This report introduces Kandinsky 5.0, a family of state-of-the-art foundation models for high-resolution image and 10-second video synthesis.
arXiv:2605.26874v1 Announce Type: cross Abstract: LLM-based agents for industrial asset operations show limited accuracy when reasoning over flat document stores. AssetOpsBench (KDD 2026) establishes
Krea is now built in to Hermes Agent as an image generation API provider, allowing your agent to use Krea 2: a new foundation model trained from scratch to balance aesthetic quality and fine control,
arXiv:2605.26717v1 Announce Type: cross Abstract: Adapting large language models (LLMs) for personalized recommendation requires aligning their general-purpose capabilities with user-specific preferen
arXiv:2511.02360v4 Announce Type: replace-cross Abstract: Chain of Thought (CoT) reasoning enhances logical performance by decomposing complex tasks, yet its multimodal extension faces a trade-off. Th
arXiv:2605.27066v1 Announce Type: new Abstract: Understanding how events evolve over time is essential for search engines handling queries about trending news. We present QDET (Query-Driven Event Time
This newsletter episode covers major AI developments including Elon Musk's legal loss against OpenAI, updates announced at Google's I/O conference, and OpenAI's achievement in solving the Erdős discre
arXiv:2605.26797v1 Announce Type: cross Abstract: We study Latent Recurrent Transformer (LRT), a lightweight augmentation of autoregressive transformers that reuses a high-level source-layer hidden st
arXiv:2505.21285v4 Announce Type: replace Abstract: This work proposes a framework LGKDE that learns kernel density estimation for graphs. The key challenge in graph density estimation lies in effecti
arXiv:2603.27146v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used to assist ideation in research, but evaluating the quality of LLM-generated research proposals re
arXiv:2605.27190v1 Announce Type: cross Abstract: Recent advances in Large Audio-Language Models (LALMs) have made real-time, streaming spoken interaction increasingly practical. In this setting, reas
LiteParse 🤝 Rust🦀 We refactored the LiteParse library and CLI porting it to Rust, and here's what that means for you: ⚡ Parsing up to 100X faster, with sub-second processing for documents as large as
LiteParse v2.0 is out now, and it is blazing fast + runs everywhere! We rewrote everything from scratch in Rust, and now: - up to 100x faster parsing - install natively in Rust, JS/TS, and Python - a
arXiv:2605.26781v1 Announce Type: new Abstract: Advanced Large Multimodal Models (LMMs) have demonstrated impressive performance in K-12 reasoning tasks, exhibiting great promise as intelligent tutors
arXiv:2510.13217v2 Announce Type: replace-cross Abstract: Search systems are increasingly used for reasoning-intensive queries, where what makes a document relevant requires understanding or reasoning
arXiv:2605.27088v1 Announce Type: new Abstract: Aligning LLMs for math tutoring typically requires RL-based training with multi-GPU infrastructure. We investigate whether training-free prompt optimiza
arXiv:2601.18987v5 Announce Type: replace-cross Abstract: Determining whether a program terminates is a central problem in computer science. Turing's Halting Problem established termination as undecid
arXiv:2605.26244v1 Announce Type: new Abstract: Audio-visual generation is rapidly advancing from short clips to minute-long content, while existing evaluation protocols remain largely confined to sho
arXiv:2605.26486v1 Announce Type: new Abstract: Despite advances in audio-driven video generation, achieving commercial-grade stability remains challenging. We present LongCat-Video-Avatar 1.5, an upg
Look back at last week’s I/O announcements with @NotebookLM. You can listen to an audio overview, watch the video recap, and even check out our detailed slide deck summarizing all of the biggest news
arXiv:2605.26438v1 Announce Type: cross Abstract: Large language models can recognize when they are being evaluated (evaluation awareness) and behave differently because of that, which undermines the
arXiv:2605.27331v1 Announce Type: new Abstract: Competition law experts conducting legal research must review extensive volumes of cases, decisions, and judicial reports to identify precedents and ass
arXiv:2510.08932v2 Announce Type: replace Abstract: Recently, a growing body of research has focused on either optimizing CTR model architectures to better model feature interactions or refining train
arXiv:2601.08267v3 Announce Type: replace Abstract: While reasoning-enhanced large language models perform strongly on English medical tasks, a persistent multilingual gap remains, with substantially
arXiv:2605.26621v1 Announce Type: cross Abstract: Volumetric Reasoning Segmentation (VRS) aims to segment a target region in a 3D medical scan from a free-form clinical query, where the referent is of
arXiv:2605.26667v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly rely on external memory systems to remain consistent across long-horizon interactions, but little empiric
arXiv:2605.26394v1 Announce Type: new Abstract: Multi-turn Text-to-SQL is central to enterprise analytics yet remains predominantly evaluated in single-turn settings. We introduce EnterpriseMem-Bench,
arXiv:2605.26959v1 Announce Type: cross Abstract: MerLean-Prover is an end-to-end Lean4 theorem prover that replaces sorry declarations with kernel-checkable proofs. It is built from three agent types
Sarah Perez / TechCrunch: Meta rolls out Plus plans for Instagram, Facebook, and WhatsApp globally, will test 7.99/mo. and 19.99/mo. Meta AI plans, a $49.99/mo. creator plan, and more — Meta is doubli
arXiv:2605.26712v1 Announce Type: new Abstract: Benchmarks that reflect the diversity and complexity of real-world documents are essential for accurately evaluating Automatic Text Recognition (ATR) sy
arXiv:2605.26546v1 Announce Type: new Abstract: Mobile graphical user interface (GUI) agents enable AI models to autonomously operate smartphones on behalf of users. However, most existing systems foc
arXiv:2605.27358v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) has become the de facto architecture for hundred-billion-parameter language models, yet its advantages at sub-billion scales
arXiv:2605.27158v1 Announce Type: new Abstract: Discovering the governing equations of a dynamical system from observed trajectories provides deeper insight into its structure than mere prediction of
Model-Harness-Task fit! it’s clear that RL post-training produces a model-harness fit via tool shapes and prompting as models are trained with the harness in the loop. Mentioned this in a previous Lan
arXiv:2605.26191v1 Announce Type: cross Abstract: This research addresses the problem of adaptive modeling in time-series data streams with clear input-output relationships. This problem is challengin
arXiv:2507.13762v4 Announce Type: replace Abstract: Motivation: Structure-based drug design (SBDD) has advanced with deep generative models, but bridging the gap between continuous atomic coordinates
arXiv:2605.26647v1 Announce Type: cross Abstract: Feedforward network (FFN) layers account for a large fraction of parameters and nonlinear expressivity in Transformer-based large language models (LLM