Transformer Approximations from ReLUs
arXiv:2604.24878v1 Announce Type: new Abstract: We provide a systematic recipe for translating ReLU approximation results to softmax attention mechanism. This recipe covers many common approximation t
Knowledge catalogue
arXiv:2604.24878v1 Announce Type: new Abstract: We provide a systematic recipe for translating ReLU approximation results to softmax attention mechanism. This recipe covers many common approximation t
arXiv:2604.25898v1 Announce Type: new Abstract: Continual offline reinforcement learning (CORL) aims to learn a sequence of tasks from datasets collected over time while preserving performance on prev
arXiv:2511.07743v3 Announce Type: replace Abstract: Ultrasound imaging is a cornerstone of non-invasive clinical diagnostics, yet its limited field of view poses challenges for novel view synthesis. W
arXiv:2604.24883v1 Announce Type: cross Abstract: We study the zero-temperature phase diagram of the 2D spin-imbalanced Fermi gas with short-ranged attractive interactions using the recently developed
arXiv:2511.14183v3 Announce Type: replace Abstract: Digital images are often degraded by soft effects such as lens flare, haze, shadows, and reflections, which reduce aesthetics even though the underl
arXiv:2604.25776v1 Announce Type: new Abstract: Critical analyses of emotion recognition technology have raised ethical concerns around task validity and potential downstream impacts, urging researche
arXiv:2410.09635v2 Announce Type: replace Abstract: Early detection of intrapartum risks enables timely interventions to prevent or mitigate adverse labor outcomes such as cerebral palsy. However, acc
arXiv:2604.25334v1 Announce Type: new Abstract: Imbalanced classification remains a pervasive challenge in machine learning, particularly when minority samples are too scarce to provide a robust discr
arXiv:2409.11063v3 Announce Type: replace-cross Abstract: Variational principles play a central role in classical mechanics, providing compact formulations of dynamics and direct access to conserved q
arXiv:2604.25897v1 Announce Type: cross Abstract: Contact variability, sensing uncertainty, and external disturbances make grasp execution stochastic. Expected-quality objectives ignore tail outcomes
arXiv:2604.24885v1 Announce Type: new Abstract: We introduce an efficient, resolution-agnostic autoregressive (AR) image synthesis approach that generalizes to arbitrary resolutions and aspect ratios,
arXiv:2604.24953v1 Announce Type: new Abstract: While preference optimization is crucial for improving visual generative models, how to effectively scale this paradigm remains largely unexplored. Curr
arXiv:2604.24894v1 Announce Type: cross Abstract: We propose VISION-SLS, a method for nonlinear output-feedback control from high-resolution RGB images which provides robust constraint satisfaction gu
arXiv:2604.25570v1 Announce Type: new Abstract: Spiking Transformers have shown strong potential for long-range visual modeling through spike-driven self-attention. However, their quadratic token inte
arXiv:2604.25235v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly used as automated judges for multimodal systems, yet their scores provide no indication of reliability.
arXiv:2604.03472v2 Announce Type: replace Abstract: Co-evolutionary self-play, where one language model generates problems and another solves them, promises autonomous curriculum learning without huma
arXiv:2511.21517v2 Announce Type: replace Abstract: Unlike text, speech conveys information about the speaker, such as gender, through acoustic cues like pitch. This gives rise to modality-specific bi
arXiv:2512.12072v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly being used to generate synthetic datasets for the evaluation and training of downstream models. Howeve
arXiv:2604.25591v1 Announce Type: cross Abstract: Recent audio-aware large language models (ALLMs) have demonstrated strong capabilities across diverse audio understanding and reasoning tasks, but the
arXiv:2604.25132v1 Announce Type: new Abstract: Instruction-tuning datasets often contain substantial redundancy and low-quality samples, necessitating effective data selection methods. We propose an
arXiv:2604.25872v1 Announce Type: new Abstract: Training language models via reinforcement learning often relies on imperfect proxy rewards, since ground truth rewards that precisely define the intend
arXiv:2604.25213v1 Announce Type: new Abstract: OpenAI's GPT-Image-2 has effectively erased the visual boundary between authentic and AI-edited document images: a single number on a receipt can be rep
arXiv:2510.07499v2 Announce Type: replace Abstract: Recent Long-Context Language Models (LCLMs) can process hundreds of thousands of tokens in a single prompt, enabling new opportunities for knowledge
arXiv:2604.25611v1 Announce Type: new Abstract: Real-time automatic speech recognition (ASR) systems face a fundamental trade-off between transcription accuracy and computational efficiency, particula
arXiv:2603.14248v2 Announce Type: replace-cross Abstract: Large language model (LLM) web agents are increasingly used for web navigation but remain far from human reliability on realistic, long-horizo
arXiv:2604.25011v1 Announce Type: new Abstract: Reinforcement learning (RL)-based post-training often improves the reasoning performance of large language models (LLMs) beyond the training domain, whi
arXiv:2604.25012v1 Announce Type: new Abstract: Automated agentic workflow design currently relies on per-task iterative search, which is computationally prohibitive and fails to reuse structural know
arXiv:2604.25384v1 Announce Type: new Abstract: This paper presents a methodology for transforming raw Wikimedia dumps into quality textual corpora for seven South Slavic languages. The work is divide
arXiv:2604.25076v1 Announce Type: new Abstract: Many Multi-Agent Reinforcement Learning (MARL) agents fail to adapt properly to cooperating with agents trained with the same objectives but different s
arXiv:2604.22830v1 Announce Type: new Abstract: Pre-training is a general method that is used in a range of deep learning tasks. By first training a model on one task, and then further training on the
arXiv:2604.23935v1 Announce Type: new Abstract: Audio-based video object segmentation aims to locate and segment objects in videos conditioned on audio cues, requiring precise understanding of both ap
arXiv:2604.24149v1 Announce Type: new Abstract: Remote sensing images are frequently degraded by adverse weather conditions, particularly clouds and haze, which severely impair downstream applications
arXiv:2604.23458v1 Announce Type: new Abstract: The growing availability of online support groups has opened up new windows to study mental health through natural language processing (NLP). However, i
arXiv:2601.13288v2 Announce Type: replace Abstract: Production LLM systems often rely on separate models for safety and other classification-heavy steps, increasing latency, VRAM footprint, and operat
arXiv:2603.07475v2 Announce Type: replace Abstract: Autoregressive (AR) language models build representations incrementally via left-to-right prediction, while diffusion language models (dLLMs) are tr
arXiv:2604.23049v1 Announce Type: new Abstract: AI agents are increasingly deployed to execute tasks and make decisions within agentic workflows, introducing new requirements for safe and controlled a
arXiv:2604.23045v1 Announce Type: new Abstract: Systematic biases in Global Circulation Model (GCM) outputs limit their direct applicability in regional planning, necessitating bias correction. Correc
arXiv:2604.22858v1 Announce Type: new Abstract: Liver cancer, especially hepatocellular carcinoma (HCC), imposes a substantial global disease burden. Accurate diagnosis and prognostic assessment direc
arXiv:2604.24172v1 Announce Type: cross Abstract: This paper uses a minimum divergence framework to introduce a new way of calculating model weights that can be used to average probabilistic predictio
arXiv:2604.23127v1 Announce Type: cross Abstract: Soil salinity is a major environmental challenge in coastal Bangladesh, threatening agricultural productivity and local livelihoods. This study develo
arXiv:2509.21199v3 Announce Type: replace Abstract: Multi-Hop Question Answering (MHQA) requires integrating dispersed, interdependent evidence through sequential reasoning under noise. This task is c
arXiv:2604.24672v1 Announce Type: new Abstract: We provide a mathematical interpretation of convolutional (or message passing) neural networks by using presheaves and copresheaves of the set of contin
arXiv:2604.22780v1 Announce Type: cross Abstract: Aligning physiological parameter labels with large-scale photoplethysmographic (PPG) data for deep learning is challenging and resource-intensive. Whi
arXiv:2604.23790v1 Announce Type: new Abstract: A central problem in unsupervised domain adaptation is determining what to transfer from labeled source domains to an unlabeled target domain. To handle
arXiv:2512.21372v2 Announce Type: replace-cross Abstract: The accurate classification of gastrointestinal diseases from endoscopic and histopathological imagery remains a significant challenge in medi
arXiv:2604.23415v1 Announce Type: new Abstract: Most two-stream action recognition networks apply the same convolutional backbone to both RGB and optical flow streams, ignoring the fact that the two m
arXiv:2604.23271v1 Announce Type: new Abstract: Automated white blood cell (WBC) classification is essential for scalable leukaemia screening. However, real-world deployment is challenged by domain sh
arXiv:2510.04916v2 Announce Type: replace Abstract: Deep learning has become increasingly important in remote sensing image classification due to its ability to extract semantic information from compl
arXiv:2604.22864v1 Announce Type: cross Abstract: Existing benchmarks for systematic reviewing remain limited either in scale or in disciplinary coverage, with some collections comprising only a modes
arXiv:2604.23225v1 Announce Type: new Abstract: This paper investigates the deep learning optimization problem with softmax cross-entropy loss. We propose a layer separation strategy to alleviate the
arXiv:2602.15853v2 Announce Type: replace-cross Abstract: We propose a lightweight explainable guardrail (LEG) method to detect unsafe prompts. LEG uses a multi-task learning architecture to jointly l
arXiv:2604.23074v1 Announce Type: new Abstract: Autonomous multirotor landings on uncrewed surface vessels (USVs) are critical for persistent maritime operations but remain challenging due to wave-ind
arXiv:2604.24037v1 Announce Type: new Abstract: Emergent intelligence have played a major role in the modern AI development. While existing studies primarily rely on empirical observations to characte
arXiv:2507.16079v2 Announce Type: replace-cross Abstract: With the advancement of deep learning, reducing computational complexity and memory consumption has become a critical challenge, and ternary n
arXiv:2604.23468v1 Announce Type: cross Abstract: In 2016, Viazovska famously solved the sphere packing problem in dimension 8, using modular forms to construct a 'magic' function satisfying optimalit
arXiv:2601.12483v2 Announce Type: replace-cross Abstract: Quantum error correction is a key ingredient for large scale quantum computation, protecting logical information from physical noise by encodi
arXiv:2604.24429v1 Announce Type: new Abstract: As the application of Large Language Models (LLMs) spreads across various industries, there are increasing concerns about the potential for their misuse
arXiv:2604.23388v1 Announce Type: cross Abstract: Generative information retrieval (GenIR) consolidates retrieval into a single neural model that decodes document identifiers (docids) directly from qu
arXiv:2604.23704v1 Announce Type: new Abstract: Multi-camera systems offer rich observation capabilities for visual navigation and 3D scene reconstruction; however, the resulting feature redundancy of
arXiv:2507.22345v2 Announce Type: replace Abstract: Wheel-legged robots integrate leg agility on rough terrain with wheel efficiency on flat ground. However, most existing designs do not fully capital