Prism: Spectral-Aware Block-Sparse Attention
arXiv:2602.08426v2 Announce Type: replace-cross Abstract: Block-sparse attention is promising for accelerating long-context LLM pre-filling, yet identifying relevant blocks efficiently remains a bottl
Knowledge catalogue
arXiv:2602.08426v2 Announce Type: replace-cross Abstract: Block-sparse attention is promising for accelerating long-context LLM pre-filling, yet identifying relevant blocks efficiently remains a bottl
arXiv:2605.25020v1 Announce Type: new Abstract: Chronic dermatologic diseases such as pemphigus require long-term follow-up, generating extensive longitudinal clinical documentation that is difficult
arXiv:2605.24295v1 Announce Type: new Abstract: We propose PACE-GGM, a data-adaptive differentially private method for covariance estimation that concentrates its privacy budget on the most informativ
arXiv:2605.24249v1 Announce Type: new Abstract: The growing availability of clinical data has increased the use of machine learning, yet centralized data aggregation is often infeasible for sensitive
arXiv:2605.25404v1 Announce Type: new Abstract: Cascaded Automatic Speech Recognition -- Large Language Model (ASR-LLM) pipelines remain popular for industrial Spoken Dialogue Systems (SDS), primarily
arXiv:2605.24900v1 Announce Type: new Abstract: Proactive task-oriented agents must autonomously anticipate user needs, identify actionable opportunities, and trigger software actions at appropriate m
arXiv:2510.27118v4 Announce Type: replace Abstract: Most expressivity results for transformers treat them as language recognizers -- devices that accept or reject strings -- rather than as they are us
arXiv:2509.07961v2 Announce Type: replace Abstract: We develop new experimental paradigms for measuring welfare in language models. We compare verbal reports of models about their preferences with pre
arXiv:2605.25682v1 Announce Type: cross Abstract: Distributing Transformer inference across embedded edge devices can alleviate individual memory and compute constraints, yet practical benefits on rea
arXiv:2603.20479v2 Announce Type: replace-cross Abstract: Learning another language can be a highly emotional process, typically characterized by numerous frustrations and triumphs, big and small. For
arXiv:2605.24171v1 Announce Type: cross Abstract: Large language models are increasingly used for vulnerability detection, yet their reliability under different prompt formulations remains uncharacter
arXiv:2605.24756v1 Announce Type: new Abstract: Language-model agents increasingly emit uncertainty signals throughout a trajectory, but existing agentic UQ evaluations often conflate ranking usefulne
arXiv:2507.05890v4 Announce Type: replace-cross Abstract: As psychometric surveys are increasingly used to assess the traits of large language models (LLMs), the need for scalable survey item generati
arXiv:2601.06870v2 Announce Type: replace-cross Abstract: Multimodal large language models have demonstrated strong ability in capturing semantic representations for multimodal sentiment analysis. The
arXiv:2605.25066v1 Announce Type: cross Abstract: Quantum machine learning (QML) is moving from research prototypes to deployed cloud services. As QML enters regulated industries, the integrity of the
arXiv:2605.23991v1 Announce Type: cross Abstract: There is a growing urgency to track greenhouse gasses with the resolution, precision and accuracy needed to support independent verification of CO_2 f
arXiv:2605.25252v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for post-training language models, but in practice, verifiers are
arXiv:2605.24904v1 Announce Type: new Abstract: Machine-translated benchmarks are widely used to assess the multilingual capabilities of large language models (LLMs), yet translation errors in these b
arXiv:2605.25933v1 Announce Type: cross Abstract: Posttraumatic stress disorder (PTSD) is a prevalent and debilitating mental health condition with significant personal and societal impacts. Current c
arXiv:2605.23930v1 Announce Type: new Abstract: We introduce Quantum Frog, a two-player cooperative game built on a novel quantized-time mechanic in which the environment advances only when a player a
arXiv:2605.24920v1 Announce Type: cross Abstract: Quaternion neural networks are parameter-efficient and model multidimensional dependencies by representing four related features as a single entity. H
arXiv:2410.10652v4 Announce Type: replace-cross Abstract: Cells in multicellular organisms coordinate to form structural and functional niches. With spatial transcriptomics (ST) enabling gene expressi
arXiv:2605.24218v1 Announce Type: new Abstract: Deep research agents extend the role of search engines from retrieving keyword-matched pages to synthesizing knowledge, fundamentally changing how human
arXiv:2605.25955v1 Announce Type: cross Abstract: Large language models (LLMs) face a dual challenge in creative capability evaluation: existing benchmarks (e.g., Story Cloze Test, HellaSwag) measure
arXiv:2605.23956v1 Announce Type: new Abstract: Compound AI systems that chain multiple LLM calls into directed computation graphs are now the dominant architecture for production AI. Although these a
arXiv:2507.14760v2 Announce Type: replace-cross Abstract: While deep learning offers tremendous promise for scientific and medical imaging, any failures and hallucinations (predictions that do not coi
arXiv:2605.25041v1 Announce Type: new Abstract: 4D radar is increasingly attractive for robotic mapping because it provides range, azimuth, elevation, and Doppler measurements while remaining robust i
arXiv:2605.25057v1 Announce Type: cross Abstract: Neural networks with randomly generated hidden weights (RaNNs) have been extensively studied, both as a standalone learning method and as an initializ
arXiv:2605.23912v1 Announce Type: cross Abstract: We present Raon-Speech, a top-performing 9B-parameter speech language model (SpeechLM) for English and Korean speech understanding, answering, and gen
arXiv:2604.00963v2 Announce Type: replace-cross Abstract: We show polylogarithmic mixing time bounds for the alternating-scan sampler for positively weighted restricted Boltzmann machines. This is don
arXiv:2605.23994v1 Announce Type: cross Abstract: Digital avatar watermarking presents unique challenges: avatars are routinely post-processed with background replacement, reframing, and format conver
arXiv:2603.11001v2 Announce Type: replace-cross Abstract: Human uplift studies, or studies that measure the effects of AI access on human performance via randomized controlled trials (RCT) or similar
arXiv:2605.25171v1 Announce Type: new Abstract: In most existing AI humor research, humor was treated as either 'present' or 'not present.' We explore the concept of humor as a social interaction with
arXiv:2501.02672v2 Announce Type: replace-cross Abstract: Characterising cause-effect relationships in complex systems is fundamental to understanding their underlying mechanisms. Granger causality (G
arXiv:2501.18278v3 Announce Type: replace Abstract: State-of-the-art models represent proteins and molecules in separate embedding manifolds, limiting the modeling of systemic biological processes. We
arXiv:2605.25281v1 Announce Type: cross Abstract: Recent advances in large language models (LLMs) have made it increasingly difficult to distinguish human-written text from AI-generated content. Many
arXiv:2603.09095v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) can process text presented as images, yet they often perform worse than when the same content is provided a
arXiv:2605.25902v1 Announce Type: new Abstract: Narrowly finetuned language models memorize implanted content verbatim, but auditing what a deployed model has been taught, without access to its weight
arXiv:2605.24945v1 Announce Type: cross Abstract: Accurate evaluation of weather forecasting models is critical for their reliable deployment in real-world applications. However, existing benchmarks p
arXiv:2605.24004v1 Announce Type: new Abstract: Large language models (LLMs) are promising for autonomous driving, but semantics-only decision policies can yield physically unsafe behavior in dynamic
arXiv:2605.24497v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) have demonstrated remarkable capabilities in reasoning and generation tasks and are increasingly deployed in real-world ap
arXiv:2602.00682v2 Announce Type: replace-cross Abstract: Integrating large language model (LLM) representations into multimodal recommendation has shown promise, yet a fundamental challenge remains l
arXiv:2605.25095v1 Announce Type: new Abstract: Autonomous driving stacks must pick one trajectory from a multi-modal candidate set; choosing by model confidence ignores safety, traffic-law, and comfo
arXiv:2605.24044v1 Announce Type: new Abstract: Robots deployed in dynamic environments must contend with environment-driven changes that reshape computation at runtime: new tasks may appear, preceden
arXiv:2602.19450v2 Announce Type: replace-cross Abstract: Trusted Execution Environments (TEEs) (e.g., Intel SGX and ArmTrustZone) aim to protect sensitive computation from a compromised operating sys
arXiv:2605.25673v1 Announce Type: cross Abstract: Security evaluations inherently depend on stable identifiers. Any finding, audit, or regulatory decision must remain attached to the specific artifact
arXiv:2605.24357v1 Announce Type: new Abstract: In this paper, we study the role of the critic in actor--critic for entropy-regularized, finite, discounted environments. We establish that, when the cr
arXiv:2602.01183v2 Announce Type: replace-cross Abstract: Biological learning proceeds from easy to difficult tasks, gradually reinforcing perception and robustness. Inspired by this principle, we add
arXiv:2605.24834v1 Announce Type: cross Abstract: Large language model (LLM) safety classifiers such as Llama Guard are effective at detecting overtly harmful prompts but remain vulnerable to adversar
arXiv:2605.25063v1 Announce Type: new Abstract: Reinforcement learning offers a promising approach for scan-order optimisation in laser additive manufacturing, where sequential scan decisions critical
arXiv:2605.24740v1 Announce Type: new Abstract: Reinforcement learning (RL) for reachability specifications is fundamental in sequential decision-making, yet theoretical guarantees remain less explore
arXiv:2605.25638v1 Announce Type: new Abstract: Policy loss estimation remains a fundamental and long-standing challenge in reinforcement learning (RL) for diffusion language models (dLLMs). We introd
arXiv:2605.25172v1 Announce Type: cross Abstract: This article is the rejoinder to ``The ICML 2023 Ranking Experiment: Examining Author Self-Assessment in ML/AI Peer Review,'' to appear in the Journal
arXiv:2605.25508v1 Announce Type: new Abstract: At very high sparsity, neural network pruning does more than decide which weights remain. It also determines where pruning induced damage is placed acro
arXiv:2409.02416v2 Announce Type: replace Abstract: Motivated by the Bures distance, we introduce a new family of distances, relative translation invariant Wasserstein distances, denoted by RW_p, as a
arXiv:2605.24003v1 Announce Type: cross Abstract: Remote sensing techniques have been increasingly utilised in aquatic applications in recent years. A common challenge in using optical satellite data
arXiv:2605.24850v1 Announce Type: new Abstract: Evaluating whether large language models (LLMs) capture the structure of natural language beyond local fluency remains an open challenge. Existing evalu
arXiv:2512.23995v2 Announce Type: replace-cross Abstract: Mixture-of-Experts architectures have become the standard for scaling large language models due to their superior parameter efficiency. To acc
arXiv:2605.25851v1 Announce Type: new Abstract: Embodied instruction following (EIF) requires agents to understand and execute complex natural language commands within interactive 3D environments. Des
arXiv:2605.24428v1 Announce Type: new Abstract: Stochastic process-based molecular graph generators have become the state of the art for template-free single-step retrosynthesis. However, these models