Low-Rank Key Value Attention
arXiv:2601.11471v3 Announce Type: replace Abstract: The key-value (KV) cache is a primary memory bottleneck in Transformers. We propose Low-Rank Key-Value (LRKV) attention, which reduces KV cache memo
Knowledge catalogue
arXiv:2601.11471v3 Announce Type: replace Abstract: The key-value (KV) cache is a primary memory bottleneck in Transformers. We propose Low-Rank Key-Value (LRKV) attention, which reduces KV cache memo
arXiv:2604.07823v1 Announce Type: new Abstract: Performance, the externalization of intent, emotion, and personality through visual, vocal, and temporal behavior, is what makes a character alive. Lear
arXiv:2604.07143v1 Announce Type: new Abstract: We introduce Lumbermark, a robust divisive clustering algorithm capable of detecting clusters of varying sizes, densities, and shapes. Lumbermark iterat
arXiv:2512.17489v2 Announce Type: replace Abstract: Text-to-image (T2I) models have demonstrated remarkable progress in creative image generation, yet they still lack precise control over scene illumi
arXiv:2603.04300v2 Announce Type: replace Abstract: Foundation models in general promise to accelerate scientific computation by learning reusable representations across problem instances, yet constra
arXiv:2604.06737v1 Announce Type: cross Abstract: Large language models have demonstrated remarkable capabilities across a wide range of natural language processing tasks, yet their application in the
arXiv:2512.19253v4 Announce Type: replace-cross Abstract: We present the first empirical study of machine unlearning (MU) in hybrid quantum-classical neural networks. While MU has been extensively exp
arXiv:2604.06950v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) are increasingly being deployed as automated content moderators. Within this landscape, we uncover a critic
arXiv:2604.07276v1 Announce Type: cross Abstract: GROMACS is a de-facto standard for classical Molecular Dynamics (MD). The rise of AI-driven interatomic potentials that pursue near-quantum accuracy a
arXiv:2509.22750v3 Announce Type: replace Abstract: Real-world multi-hop QA is naturally linked with ambiguity, where a single query can trigger multiple reasoning paths that require independent resol
arXiv:2604.06269v1 Announce Type: cross Abstract: Automated cellular reasoning faces a core dichotomy: supervised methods fall into the Reference Trap and fail to generalize to out-of-distribution cel
arXiv:2604.07574v1 Announce Type: new Abstract: Image matching is a fundamental problem in Computer Vision with direct applications in robotics, remote sensing, and geospatial data analysis. We presen
arXiv:2409.09298v2 Announce Type: replace-cross Abstract: The Matrix Profile (MP), a versatile tool for time series data mining, has been shown effective in time series anomaly detection (TSAD). This
arXiv:2604.02445v2 Announce Type: replace Abstract: Matrix Profile (MP) methods are an interpretable and scalable family of distance-based methods for time-series anomaly detection, but strong benchma
arXiv:2603.22364v2 Announce Type: replace-cross Abstract: Diffusion models have achieved state-of-the-art performance in generative modeling, but their success often relies heavily on classifier-free
arXiv:2509.22981v2 Announce Type: replace Abstract: We study a class of multi-stage stochastic programs, which incorporate modeling features from Markov decision processes (MDPs). This class includes
arXiv:2604.07345v1 Announce Type: cross Abstract: The rapid growth of generative artificial intelligence (AI) has introduced unprecedented computational demands, driving significant increases in the e
arXiv:2604.06505v1 Announce Type: cross Abstract: Large language models (LLMs) are widely explored for reasoning-intensive research tasks, yet resources for testing whether they can infer scientific c
arXiv:2604.06846v1 Announce Type: cross Abstract: Interactive medical dialogue benchmarks have shown that LLM diagnostic accuracy degrades significantly when interacting with non-cooperative patients,
arXiv:2604.06180v1 Announce Type: cross Abstract: Medical diagnosis using Large Multimodal Models (LMMs) has gained increasing attention due to capability of these models in providing precise diagnose
arXiv:2604.08203v1 Announce Type: new Abstract: Medical Vision-Language Models (VLMs) hold immense promise for complex clinical tasks, but their reasoning capabilities are often constrained by text-on
arXiv:2604.08364v1 Announce Type: new Abstract: In this paper, we introduce MegaStyle, a novel and scalable data curation pipeline that constructs an intra-style consistent, inter-style diverse and hi
arXiv:2604.07877v1 Announce Type: new Abstract: Long-term memory is fundamental for personalized and autonomous agents, yet populating it remains a bottleneck. Existing systems treat memory extraction
arXiv:2604.06881v1 Announce Type: new Abstract: Neural operators have emerged as powerful surrogates for dynamical systems due to their grid-invariant properties and computational efficiency. However,
arXiv:2507.10303v2 Announce Type: replace-cross Abstract: Stochastic simulators exhibit intrinsic stochasticity due to unobservable, uncontrollable, or unmodeled input variables, resulting in random o
arXiv:2604.06473v1 Announce Type: new Abstract: Multivariate forecasting with Transformers faces a core scalability challenge: modeling cross-channel dependencies via attention compounds attention's q
arXiv:2511.08605v3 Announce Type: replace Abstract: Bangladesh's low-income population faces major barriers to affordable legal advice due to complex legal language, procedural opacity, and high costs
arXiv:2601.04068v3 Announce Type: replace Abstract: Aligning text-to-video diffusion models with human preferences is crucial for generating high-quality videos. Existing Direct Preference Otimization
arXiv:2604.04771v2 Announce Type: replace-cross Abstract: Current document parsing methods advance primarily through model architecture innovation, while systematic engineering of training data remain
arXiv:2604.07085v1 Announce Type: new Abstract: In electronic health records (EHRs), clustering patients and distinguishing disease subtypes are key tasks to elucidate pathophysiology and aid clinical
arXiv:2604.07747v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) can improve low-k reasoning accuracy while narrowing solution coverage on challenging math que
arXiv:2508.07514v3 Announce Type: replace Abstract: Reliable plant species and damage segmentation for herbicide field research trials requires models that can withstand substantial real-world variati
arXiv:2604.07914v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) have achieved remarkable success across cross-modal tasks but remain hindered by hallucinations, producing textual
arXiv:2510.15770v3 Announce Type: replace Abstract: Concept Bottleneck Models (CBMs) enhance interpretability by predicting human-understandable concepts as intermediate representations. However, exis
arXiv:2509.23322v2 Announce Type: replace Abstract: With the continuous expansion of Large Language Models (LLMs) and advances in reinforcement learning, LLMs have demonstrated exceptional reasoning c
arXiv:2604.07121v1 Announce Type: cross Abstract: In the human-AI collaboration area, the context formed naturally through multi-turn interactions is typically flattened into a chronological sequence
arXiv:2604.07191v1 Announce Type: cross Abstract: Mixture proportion estimation (MPE) aims to estimate class priors from unlabeled data. This task is a critical component in weakly supervised learning
arXiv:2412.20718v2 Announce Type: replace Abstract: The rapid integration of Large Vision-Language Models (LVLMs) into critical domains necessitates comprehensive moral evaluation to ensure their alig
arXiv:2604.06267v1 Announce Type: cross Abstract: Multimodal variational autoencoders (VAEs) have emerged as a powerful framework for survival risk modeling in multiple myeloma by integrating heteroge
arXiv:2604.06798v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) based large language models (LLMs) offer strong performance but suffer from high memory and computation costs. Weight binariz
arXiv:2603.20654v4 Announce Type: replace-cross Abstract: Classical Amdahl's Law conceptualized the limit of speedup for an era of fixed serial-parallel decomposition and homogeneous replication. Mode
arXiv:2601.02535v2 Announce Type: replace Abstract: Selecting a single high-quality output from multiple stochastic generations remains a fundamental challenge for large language models (LLMs), partic
arXiv:2604.07030v1 Announce Type: new Abstract: Sparse Mixture-of-Experts (MoE) architectures are increasingly popular for frontier large language models (LLM) but they introduce training challenges d
arXiv:2604.08516v1 Announce Type: new Abstract: Web agents--autonomous systems that navigate and execute tasks on the web on behalf of users--have the potential to transform how people interact with t
arXiv:2604.07664v1 Announce Type: new Abstract: Monocular Depth Estimation (MDE) is a fundamental computer vision task with important applications in 3D vision. The current mainstream MDE methods empl
arXiv:2604.07780v1 Announce Type: cross Abstract: Objective: To develop a robust and compact deep learning model for automated knee cartilage segmentation on point-of-care ultrasound (POCUS) devices.
arXiv:2604.07821v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly coordinate in multi-agent systems, yet we lack an understanding of where and why cooperation failures m
arXiv:2604.07348v1 Announce Type: cross Abstract: Generating motion-controlled videos--where user-specified actions drive physically plausible scene dynamics under freely chosen viewpoints--demands tw
arXiv:2604.06390v1 Announce Type: cross Abstract: Background: Colorectal cancer (CRC) remains a leading cause of cancer-related mortality worldwide. Accurate survival prediction is essential for treat
arXiv:2604.07991v1 Announce Type: new Abstract: Recent advances in world models have demonstrated strong capabilities in simulating physical reality, making them an increasingly important foundation f
arXiv:2603.28253v2 Announce Type: replace-cross Abstract: Time series forecasting is vital across many domains, yet existing models struggle with fixed-length inputs and inadequate multi-scale modelin
arXiv:2604.07741v1 Announce Type: new Abstract: Audio-visual deepfake detection typically employs a complementary multi-modal model to check the forgery traces in the video. These methods primarily ex
arXiv:2411.19121v2 Announce Type: replace-cross Abstract: While text-to-video diffusion models have advanced significantly, creating coherent long-form content remains unreliable due to stochastic sam
arXiv:2604.07578v1 Announce Type: new Abstract: Recognition of rodent behavior is important for understanding neural and behavioral mechanisms. Traditional manual scoring is time-consuming and prone t
arXiv:2410.17690v2 Announce Type: replace-cross Abstract: We optimize finite horizon multi-agent reach-avoid Markov decision process (MDP) via local feedback policies. The global feedback polic
arXiv:2604.06771v1 Announce Type: cross Abstract: Conversational Query Rewriting (CQR) aims to rewrite ambiguous queries to achieve more efficient conversational search. Early studies have predominant
arXiv:2604.06934v1 Announce Type: cross Abstract: Detecting user interface (UI) controls from software screenshots is a critical task for automated testing, accessibility, and software analytics, yet
arXiv:2604.06465v1 Announce Type: cross Abstract: Reasoning models have demonstrated remarkable capabilities in solving complex problems by leveraging long chains of thought. However, this more delibe
arXiv:2604.07148v1 Announce Type: new Abstract: Emerging computation-intensive applications impose stringent latency requirements on resource-constrained mobile devices. Mobile Edge Computing (MEC) ad
arXiv:2603.11633v2 Announce Type: replace Abstract: Recent unified 3D generation models have made remarkable progress in producing high-quality 3D assets from a single image. Notably, layout-aware app