SDDBMs: Soft Denoising Diffusion Bridge Models
arXiv:2608.08594v1 Announce Type: new Abstract: Diffusion bridge models leverage Doob's (h)-transform to construct stochastic transports between arbitrary endpoint distributions, and have shown strong
Knowledge catalogue
arXiv:2608.08594v1 Announce Type: new Abstract: Diffusion bridge models leverage Doob's (h)-transform to construct stochastic transports between arbitrary endpoint distributions, and have shown strong
arXiv:2608.09449v1 Announce Type: new Abstract: Video world models must capture how scenes evolve over time and across viewpoints. Training them for long-horizon generation and camera control therefor
arXiv:2608.07835v1 Announce Type: new Abstract: Cross-View Geo-Localization (CVGL) with OpenStreetMap (OSM) performs well in structure-rich urban environments but collapses in feature-sparse scenes su
arXiv:2608.07814v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) language models deliver high capacity at low per-token compute, but deploying them cheaply requires compressing their many ex
arXiv:2601.22510v2 Announce Type: replace-cross Abstract: Large language models (LLMs) often achieve strong benchmark accuracy yet remain brittle under small distribution shifts. While recent mechanis
arXiv:2608.09885v1 Announce Type: new Abstract: The safety of large language model (LLM) agents depends not only on model weights but also on the agent harness that manages context, memory, tools, per
arXiv:2608.09097v1 Announce Type: new Abstract: Despite rapid advances in generative models, achieving pixel-level precision in sketch-based image editing remains a persistent challenge, particularly
arXiv:2606.14574v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly deployed as planners for autonomous agents in household environments. While existing benchmarks
arXiv:2608.07639v1 Announce Type: cross Abstract: Agent Skills provide reusable capabilities to LLM agents. Agent Skill inconsistencies can expose undisclosed dangerous behavior or cause wrong Skill s
arXiv:2608.08640v1 Announce Type: new Abstract: Large language model agents increasingly rely on reusable skills to extend their capabilities beyond parametric knowl- edge. However, retrieving the app
arXiv:2608.09253v1 Announce Type: new Abstract: LLM agents are increasingly equipped with skills to perform complex tasks through multi-step reasoning and tool use. Although skills provide reusable pr
arXiv:2608.09771v1 Announce Type: new Abstract: Vision-language-action policies rely on large multimodal backbones to jointly perform perception, language conditioning, and action generation at every
arXiv:2608.08639v1 Announce Type: new Abstract: Open lakehouse table formats accumulate small data files over time, which degrades query performance. Deciding when compaction is worthwhile remains thr
arXiv:2608.09128v1 Announce Type: cross Abstract: LLM agents are increasingly deployed in multi-agent social settings where they must cooperate, negotiate, and adapt to other agents. Measuring and imp
arXiv:2608.08055v1 Announce Type: new Abstract: Large language model (LLM) agents that assist users over weeks of conversation must remember what is currently true, not merely what was once said. Flat
The Information: Sources: Nvidia is developing a Nemotron 4 model with 1T+ parameters, up from Nemotron 3 Ultra's 550B parameters but smaller than leading Chinese open models — Nvidia is doubling down
arXiv:2601.18077v3 Announce Type: replace Abstract: Cooperative reasoning under incomplete information remains challenging for both humans and multi-agent systems. The card game Hanabi embodies this c
arXiv:2511.11927v2 Announce Type: replace-cross Abstract: Principal Component Analysis (PCA) is a standard tool for extracting a low-rank signal from noisy observations. It is known that applying PCA
arXiv:2608.07915v1 Announce Type: new Abstract: Large language models (LLMs) increasingly read long inputs in the agentic era, from whole documents and codebases to conversations across many turns. Th
arXiv:2608.07921v1 Announce Type: cross Abstract: We apply Marchenko-Pastur (MP) random matrix theory to pre-trained attention weights in order to separate each projection matrix into a random-like bu
arXiv:2608.07712v1 Announce Type: cross Abstract: A predictive model receives a self-supervised signal whenever the consequence of an action is observed. Using that signal after deployment is difficul
arXiv:2601.09974v2 Announce Type: replace Abstract: Personalizing Large Language Models typically relies on static retrieval or one-time adaptation, assuming user preferences remain invariant over tim
arXiv:2608.08282v1 Announce Type: new Abstract: Tool-using language-model agents face constraints whose meaning changes with observations and prior actions. We study exact sampling from the model dist
arXiv:2608.08125v1 Announce Type: new Abstract: Recent advances in generative models have achieved remarkable performance in text- and image-conditioned editing. However, preserving the content of a g
Stealing Reasoning Traces from Proprietary LLM APIs A vanity domain name (stolen-thoughts.com) for a neat paper: Anthropic, OpenAI, and Google return encrypted chain-of-thought blocks to clients that
arXiv:2608.09126v1 Announce Type: new Abstract: In this project, we delved into the pervasive challenge of bias detection within the text content. More specifically, our focus lies on the identificati
arXiv:2510.14357v2 Announce Type: replace-cross Abstract: Agricultural robots are emerging as powerful assistants across a wide range of agricultural tasks, nevertheless, they are still heavily relyin
arXiv:2505.11480v4 Announce Type: replace-cross Abstract: Superoptimization is the task of transforming a program into a faster one, and ideally the very fastest possible one, while preserving its inp
arXiv:2608.08253v1 Announce Type: new Abstract: AI agents are becoming shared infrastructure, yet durable memory is commonly assembled from separate retrieval, governance, and operational components.
arXiv:2608.07862v1 Announce Type: new Abstract: Existing safety evaluation datasets for large language models (LLMs) predominantly focus on English and Western contexts, often overlooking the linguist
arXiv:2608.08070v1 Announce Type: new Abstract: Reliable surgical planning requires models that move beyond recognizing the current surgical step or imitating expert demonstrations, and instead antici
arXiv:2608.07641v1 Announce Type: new Abstract: The rapid advancement of large language models has transformed survey writing from a months-long manual effort into an automated process. As generation
arXiv:2608.09802v1 Announce Type: new Abstract: As AI coding agents take on increasingly complex, long-horizon software engineering tasks, existing benchmarks are rapidly saturating and their evaluati
arXiv:2608.09497v1 Announce Type: new Abstract: Operational crop mapping requires models that generalise across years, resolve fine-grained crop taxonomies, and distinguish cropland from surrounding l
arXiv:2608.09162v1 Announce Type: cross Abstract: Tabular data presents unique challenges for deep learning due to its heterogeneous nature, where numeric features exhibit diverse distributions, scale
On August 11 2026, Gary Marcus commented that if the symbolic components were removed from Claude’s Riemann implementation it would fail to work, highlighting a concrete win for hybrid neurosymbolic A
arXiv:2608.09258v1 Announce Type: new Abstract: Multi-robot systems can accomplish tasks that are difficult for a single robot by organizing into task-specific formations. Different from existing stud
arXiv:2608.08528v1 Announce Type: new Abstract: Enterprise AI coding assistants incur substantial inference spend, and naive token-cost minimization often fails to reduce end-to-end cost once retries,
arXiv:2608.09538v1 Announce Type: cross Abstract: We introduce TCS-Bench, a benchmark for evaluating Large Language Models (LLMs) on research-level Theoretical Computer Science (TCS) proof generation.
arXiv:2608.07899v1 Announce Type: new Abstract: Agent systems increasingly expose execution traces, yet telemetry that reveals a failure may still be inadequate for identifying where that failure orig
arXiv:2608.07851v1 Announce Type: cross Abstract: Residual connections rely on a static residual pathway, and are essential for training deep neural networks. Hyper-connections (HC) increase the expre
arXiv:2608.07567v1 Announce Type: cross Abstract: Functional near-infrared spectroscopy (fNIRS) is a promising modality for autism spectrum disorder (ASD) classification, yet existing approaches assum
arXiv:2608.09393v1 Announce Type: cross Abstract: We identify and quantify temporal misgrounding: the systematic retrieval and citation of the currently in-force version of a legal article when the ap
arXiv:2602.00884v2 Announce Type: replace Abstract: Neural operators have shown promise in learning solution maps of partial differential equations (PDEs), but they often struggle to generalize when t
arXiv:2608.08290v1 Announce Type: new Abstract: Open-vocabulary semantic segmentation (OVSS) repurposes a pretrained CLIP encoder for dense prediction without additional labeled supervision. Existing
I'm guessing that many people have been waiting for this comparison. For clarity, both models are running at full FP16 KV-cache. Due to VRAM limitations, Muse Glimmer is running full 262,144 context,
arXiv:2608.08809v1 Announce Type: new Abstract: A single model scale challenges the flexibility of a production retrieval system: some settings need it faster, others need a smaller index, and the rig
arXiv:2608.07617v1 Announce Type: new Abstract: Scientific and technical writing depends on markup sources that must compile: LaTeX, Typst, and Markdown pipelines fail on missing delimiters, mismatche
⚙️That is the framework Mistral is building toward: one in which enterprises, governments, and startups can use the best AI available, shape it around their own knowledge, and retain the value it crea
arXiv:2608.08026v1 Announce Type: new Abstract: We investigate how social authority (SA) signals interact with severity-based prioritization in large language models, operationalizing each axis as a m
arXiv:2505.18347v3 Announce Type: replace-cross Abstract: Continual reinforcement learning (RL) concerns agents that are expected to learn continually, rather than converge to a policy that is then fi
The ChatGPT desktop app is now available in preview for desktop variants of these Linux distributions: • Ubuntu 24.04 LTS and 26.04 LTS • Debian 13 • Fedora 43 and 44 Install with .deb or .rpm package
arXiv:2511.02687v2 Announce Type: replace Abstract: The trajectory of AI development suggests that we will increasingly rely on agent-based systems powered by language models, composed of independentl
arXiv:2608.08826v1 Announce Type: new Abstract: Adaptive procedures must work without nuisance information an oracle may use, such as a gradient scale or smoothness index, and robust procedures may ha
arXiv:2608.08650v1 Announce Type: new Abstract: Mixture-of-Experts models increase parameter capacity while keeping the computation activated by each token bounded, but their architectural evolution c
arXiv:2608.07517v1 Announce Type: cross Abstract: Can a multimodal LLM predict which version of a web page will win a real A/B test from screenshots alone? We report the most complete answer we are aw
arXiv:2608.07528v1 Announce Type: new Abstract: Linear probes detect corrupted context in language models with near-perfect accuracy, yet this does not translate into reliable failure prediction. The
arXiv:2608.09574v1 Announce Type: new Abstract: LLMs are rapidly embedding themselves into daily life: drafting our emails, managing our schedules, and making decisions on our behalf. As they move fro
arXiv:2608.08654v1 Announce Type: new Abstract: How much an AI coding agent costs to run can depend more on the agent scaffolding that drives it than on the interface through which it reaches its tool
Imagine if your everyday laptop could run an AI model smart enough to take care of 90% of your work—totally private, lightning fast, and completely free of monthly fees. That is the exact tipping poin