The Pitfalls of KV Cache Compression
arXiv:2510.00231v2 Announce Type: replace-cross Abstract: KV cache compression promises increased throughput and efficiency with negligible loss in performance. While the gains in throughput are indis
Knowledge catalogue
arXiv:2510.00231v2 Announce Type: replace-cross Abstract: KV cache compression promises increased throughput and efficiency with negligible loss in performance. While the gains in throughput are indis
arXiv:2605.14177v1 Announce Type: cross Abstract: Long-horizon personalization requires dialogue assistants to retrieve user-specific facts from extended interaction histories. In practice, many relev
arXiv:2510.04682v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are widely applied in real world scenarios, yet fine-tuning them comes with significant computational and storage
arXiv:2605.14291v1 Announce Type: cross Abstract: The rapid advancement of Large Vision-Language Models (LVLMs) is increasingly accompanied by unauthorized scraping and training on multimodal web data
arXiv:2502.16060v5 Announce Type: replace-cross Abstract: Foundation models are reshaping EEG analysis, yet an important problem of EEG tokenization remains a challenge. This paper presents TFM-Tokeni
arXiv:2605.14210v1 Announce Type: cross Abstract: Concept Bottleneck Models (CBMs) offer interpretable alternatives to black-box predictors by introducing human-relatable concepts before the final out
arXiv:2605.14999v1 Announce Type: cross Abstract: As generative AI becomes increasingly integrated into journalism, designing effective AI-use disclosures that inform readers without imposing unnecess
arXiv:2605.14866v1 Announce Type: cross Abstract: As modern microservice systems grow increasingly complex due to dynamic interactions and evolving runtime environments, they experience failures with
arXiv:2605.14717v1 Announce Type: cross Abstract: Label-free single-cell imaging offers a scalable, non-invasive alternative to fluorescence-based cytometry, yet inferring molecular phenotypes directl
arXiv:2605.13981v1 Announce Type: cross Abstract: The rise in deployment of large language models has driven a surge in GPU demand and datacenter scaling, raising concerns about electricity use, grid
arXiv:2605.13936v1 Announce Type: cross Abstract: The recent success of large language models (LLMs) has been largely driven by vast public datasets. However, the next frontier for LLM development lie
arXiv:2605.14373v1 Announce Type: cross Abstract: Zeroth-Order (ZO) optimization is pivotal for scenarios where backpropagation is unavailable, such as memory-constrained on-device learning and black-
arXiv:2605.14358v1 Announce Type: new Abstract: Language models often generate long chain-of-thought traces, but it remains unclear how much of this reasoning is necessary for preserving the final pre
arXiv:2605.15127v1 Announce Type: cross Abstract: Moving to a new culture and adapting to a new life, as an international student, can be a stressful experience. In the US, international students face
arXiv:2605.14876v1 Announce Type: cross Abstract: Despite rapid advancements, current text-to-image (T2I) models predominantly rely on a single-step generation paradigm, which struggles with complex s
arXiv:2605.14164v1 Announce Type: new Abstract: The primary way to establish and compare competencies in foundation and generative AI models has shifted from peer-reviewed literature to press releases
arXiv:2605.13933v1 Announce Type: cross Abstract: Acquisition differences across sites, scanners, and protocols in dMRI introduce variability that complicates structural connectome analysis. This moti
arXiv:2603.11042v2 Announce Type: replace-cross Abstract: Generating music that temporally aligns with video events is challenging for existing text-to-music models, which lack fine-grained temporal c
arXiv:2510.05213v2 Announce Type: replace-cross Abstract: Pretrained vision foundation models (VFMs) advance robotic learning via rich visual representations, yet individual VFMs typically excel only
arXiv:2605.14542v1 Announce Type: new Abstract: A skilled live-commerce host is not merely a narrator, but a sales agent who converts viewer curiosity into purchase intent through expert product knowl
arXiv:2605.15186v1 Announce Type: cross Abstract: High-quality 3D scene reconstruction has recently advanced toward generalizable feed-forward architectures, enabling the generation of complex environ
arXiv:2605.14747v1 Announce Type: cross Abstract: Recent advances in multimodal large language models have driven growing interest in graphical user interface (GUI) agents, yet their generalization re
arXiv:2605.14645v1 Announce Type: cross Abstract: With the rapid evolution of computer vision, vision-based methodologies for water level and river surface velocity estimation have reached significant
arXiv:2605.14710v1 Announce Type: cross Abstract: Deep learning and multi-modal fusion have demonstrated transformative potential in medical diagnosis by integrating diverse data sources. However, acc
arXiv:2605.14972v1 Announce Type: cross Abstract: A fundamental limitation of Text-to-Code is that no guarantee can be obtained about the correctness of the generated code. Therefore, to ensure its co
arXiv:2602.07045v2 Announce Type: replace-cross Abstract: Recent advancements in Multimodal Large Language Models (MLLMs) have enabled complex reasoning. However, existing remote sensing (RS) benchmar
arXiv:2605.15030v1 Announce Type: cross Abstract: Web agents can autonomously complete online tasks by interacting with websites, but their exposure to open web environments makes them vulnerable to p
arXiv:2605.13959v1 Announce Type: cross Abstract: Generative policies based on diffusion and flow matching have become a dominant paradigm for visuomotor robotic control. We show that replacing the st
arXiv:2605.14283v1 Announce Type: cross Abstract: Watermarking techniques for large language models (LLMs), which encode hidden information in the output so its source can be verified, have gained sig
arXiv:2605.14224v1 Announce Type: cross Abstract: We present an in-depth analysis of the Koopman semigroup via wavelet transform. Towards this goal, we start by introducing the wavelet-based observabl
arXiv:2605.14290v1 Announce Type: cross Abstract: ReAct has become the default architecture across LLM agents, and many existing web agents follow this paradigm. We argue that it is the wrong default
arXiv:2605.11410v2 Announce Type: replace Abstract: Clinical electroencephalogram (EEG) analysis rests on a hand-crafted feature catalog refined over decades, e.g., band power, connectivity, complexit
arXiv:2605.14449v1 Announce Type: cross Abstract: Hallucination detection in large language models (LLMs) requires balancing accu racy, efficiency, and robustness to distribution shift. Black-box cons
arXiv:2605.14478v1 Announce Type: cross Abstract: Context: Retrieval-augmented code generation relies on cross-file repository context, but retrieved snippets may come from obsolete project states. Ob
arXiv:2605.14504v1 Announce Type: new Abstract: Long-horizon household tasks demand robust high-level planning and sustained reasoning capabilities, which are largely overlooked by existing embodied A
arXiv:2605.14368v1 Announce Type: cross Abstract: Continuous diffusion language models lag behind autoregressive transformers, partly because diffusion is applied in spaces poorly suited to language d
arXiv:2512.06471v2 Announce Type: replace-cross Abstract: Goal-conditioned reinforcement learning (RL) concerns the problem of training an agent to maximize the probability of reaching target goal sta
arXiv:2605.15109v1 Announce Type: new Abstract: Retrieval-Augmented Generation can improve factuality by grounding answers in external evidence, but Agentic GraphRAG complicates what it means for cita
arXiv:2605.14192v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) has become a powerful and widely used approach for improving large language models by grounding generation in ret
arXiv:2605.15152v1 Announce Type: cross Abstract: LLM quantization has become essential for memory-efficient deployment. Recent work has shown that quantization schemes can pose critical security risk
arXiv:2605.14754v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed for knowledge synthesis, yet their capacity for compositional generalization in scientific knowle
arXiv:2605.14844v1 Announce Type: cross Abstract: We introduce XFP, a dynamic weight quantizer for LLM inference that inverts the conventional workflow: the operator specifies reconstruction quality f
arXiv:2605.14893v1 Announce Type: cross Abstract: Contrastively pre-trained Vision-Language Models (VLMs) serve as powerful feature extractors. Yet, their shared latent spaces are prone to structural
arXiv:2605.12586v1 Announce Type: cross Abstract: Vision-language models (VLMs) exhibit a striking paradox: they can generate executable code that reconstructs a 3D scene from geometric primitives wit
arXiv:2603.14066v2 Announce Type: replace-cross Abstract: Many real-world multi-party negotiations unfold as sequences of binding, action-level commitments rather than a single final outcome, yet this
arXiv:2605.13142v1 Announce Type: new Abstract: In professional sports, a team has clinched the playoffs if they are guaranteed a postseason spot, regardless of the outcomes of any remaining games. As
arXiv:2605.13687v1 Announce Type: cross Abstract: We introduce a family of synthetic languages with hierarchical structure -- generated by a broadcast process on trees -- for which the role of context
arXiv:2605.13367v1 Announce Type: cross Abstract: The literature on ontology-mediated query answering (OMQA) has been shaped by two key results: first-order rewritability for DL-Lite, and PTime-hardne
arXiv:2507.19247v5 Announce Type: replace-cross Abstract: Autoregressive language models achieve remarkable performance, yet a unified theory explaining their internal mechanisms, how training shapes
arXiv:2605.13110v1 Announce Type: cross Abstract: We present a fully automated multi-agent framework for corporate due diligence and market analysis in venture capital. The system runs on an event-dri
arXiv:2605.12685v1 Announce Type: cross Abstract: Graph Self-Supervised Learning (GSSL) has emerged as a powerful paradigm for generating high-quality representations for graph-structured data. While
arXiv:2505.12942v4 Announce Type: replace-cross Abstract: Large language models have demonstrated remarkable performance; however, their massive parameter counts make deployment highly expensive. Low-
arXiv:2605.13301v1 Announce Type: new Abstract: Recent progress in reasoning models has substantially advanced long-horizon mathematical and scientific problem solving, with several systems now reachi
arXiv:2605.13149v1 Announce Type: cross Abstract: Data quality remains a critical bottleneck in developing capable, competitive models. Researchers have explored many ways to generate top quality samp
arXiv:2605.12569v1 Announce Type: cross Abstract: Global navigation satellite system (GNSS) interference poses a serious threat to reliable positioning, especially in indoor and multipath-rich environ
arXiv:2605.12954v1 Announce Type: cross Abstract: Long video understanding is heavily bottlenecked by a rigid one-shot paradigm: existing methods either densely encode videos at prohibitive memory and
arXiv:2605.13702v1 Announce Type: new Abstract: Strategic mine production scheduling under geological uncertainty is conventionally formulated as a stochastic optimization problem in which a fixed ext
arXiv:2605.12771v1 Announce Type: cross Abstract: Multi-objective reinforcement learning in robotic domains requires balancing complex, non-convex trade-offs between conflicting objectives. While line
arXiv:2605.12694v1 Announce Type: cross Abstract: Large language models can consult information that fixed static analyzers cannot, such as documentation, current security advisories, version-specific
arXiv:2605.12532v1 Announce Type: cross Abstract: Conventional algorithmic trading systems are grounded in deterministic heuristics or offline-trained statistical models that cannot adapt to the seman