Speech Enhancement Based on Drifting Models
arXiv:2604.24199v1 Announce Type: cross Abstract: We propose Speech Enhancement based on Drifting Models (DriftSE), a novel generative framework that formulates denoising as an equilibrium problem. Ra
Knowledge catalogue
arXiv:2604.24199v1 Announce Type: cross Abstract: We propose Speech Enhancement based on Drifting Models (DriftSE), a novel generative framework that formulates denoising as an equilibrium problem. Ra
arXiv:2502.12672v4 Announce Type: replace-cross Abstract: Fine-tuning speech representation models can enhance performance on specific tasks but often compromises their cross-task generalization abili
arXiv:2604.23432v1 Announce Type: cross Abstract: Reliable depth estimation from spherical images is crucial for 360{eg} vision in robotic navigation and immersive scene understanding. However, the on
arXiv:2604.24449v1 Announce Type: cross Abstract: Training machine learning models for robotic tactile sensing requires vast amounts of data, yet obtaining realistic interaction data remains a challen
arXiv:2401.07669v2 Announce Type: replace Abstract: Adapting CLIP for videos has gained popularity due to its semantic and rich representation. While CLIP is a good starting point, it typically underg
arXiv:2505.16637v4 Announce Type: replace-cross Abstract: Large language models (LLMs) have recently demonstrated remarkable capabilities in machine translation (MT). However, most advanced MT-specifi
arXiv:2604.24003v1 Announce Type: new Abstract: Large language models (LLMs) achieve strong reasoning performance by allocating substantial computation at inference time, often generating long and ver
arXiv:2604.22892v1 Announce Type: new Abstract: Feature selection in high-dimensional genomic data (d gg n) demands methods that are simultaneously accurate, sparse, and stable. Existing approaches ei
arXiv:2604.23309v1 Announce Type: new Abstract: Remote sensing image change captioning (RSICC) aims to describe the difference between two remote sensing images. While recent methods have explored vid
arXiv:2402.11789v5 Announce Type: replace-cross Abstract: Anomaly localization in images -- identifying regions that deviate from normal patterns -- is vital in applications such as medical diagnosis
arXiv:2511.17902v3 Announce Type: replace-cross Abstract: Distributed Fiber Optic Sensing (DFOS) is promising for long-range perimeter security, yet practical deployment faces three key obstacles: sev
arXiv:2604.24544v1 Announce Type: new Abstract: The increasing reliance on Large Language Models (LLMs) across diverse sectors highlights the need for robust domain-specific and language-specific eval
arXiv:2512.10959v2 Announce Type: replace Abstract: We introduce StereoSpace, a diffusion-based framework for monocular-to-stereo synthesis that models geometry purely through viewpoint conditioning,
arXiv:2604.22782v1 Announce Type: cross Abstract: Serving transformer language models with high throughput requires caching Key-Values (KVs) to avoid redundant computation during autoregressive genera
arXiv:2604.24537v1 Announce Type: new Abstract: We study the problem of global maximization of a function f given a finite number of evaluations perturbed by noise. We consider a very weak assumption
arXiv:2604.23198v1 Announce Type: new Abstract: Current video moment retrieval excels at action-centric tasks but struggles with narrative content. Models can see extit{what is happening} but fail to
arXiv:2604.24156v1 Announce Type: cross Abstract: Efficient and fair spectrum allocation is a central challenge in 6G networks, where massive connectivity and heterogeneous services continuously compe
arXiv:2604.22757v1 Announce Type: cross Abstract: We introduce StratRAG, an open-source retrieval evaluation dataset for benchmarking Retrieval-Augmented Generation (RAG) systems on multi-hop reasonin
arXiv:2601.07698v2 Announce Type: replace Abstract: As emotional support chatbots have recently gained significant traction across both research and industry, a common evaluation strategy has emerged:
arXiv:2601.20597v2 Announce Type: replace Abstract: Continual Text-to-Video Retrieval (CTVR) is a challenging multimodal continual learning setting, where models must incrementally learn new semantic
arXiv:2604.23646v1 Announce Type: new Abstract: Recent evidence suggests that frontier AI systems can exhibit agentic misalignment, generating and executing harmful actions derived from internally con
arXiv:2604.24380v1 Announce Type: new Abstract: While Large Vision Language Models (LVLMs) demonstrate impressive capabilities, their substantial computational and memory requirements pose deployment
arXiv:2604.22843v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) has been proposed to mitigate hallucinations in large language models (LLMs), where generated outputs may be fact
arXiv:2604.23475v1 Announce Type: cross Abstract: We study the organization of channel-level importance in transformer feed-forward networks (FFNs). Using a Fisher-style loss proxy (LP) based on activ
arXiv:2604.23978v1 Announce Type: new Abstract: Family-school partnerships (FSP) are critical to children's development, yet families often face barriers such as time constraints, fragmented communica
arXiv:2604.23135v1 Announce Type: new Abstract: Natural-language variation poses a key challenge in Lean autoformalization: semantically equivalent paraphrases of the same theorem statements can induc
arXiv:2402.19088v5 Announce Type: replace-cross Abstract: Live languages continuously evolve to integrate the cultural change of human societies. This evolution manifests through neologisms (new words
arXiv:2507.09245v2 Announce Type: replace Abstract: The Swa-bhasha Resource Hub provides a comprehensive collection of data resources and algorithms developed for Romanized Sinhala to Sinhala translit
arXiv:2604.22852v1 Announce Type: cross Abstract: Cloud-hosted LLM inference for autonomous driving adds round-trip delay and depends on stable connectivity, while purely local edge models struggle un
arXiv:2601.16746v3 Announce Type: replace-cross Abstract: LLM agents have demonstrated remarkable capabilities in software development, but their performance is hampered by long interaction contexts,
arXiv:2509.14635v2 Announce Type: replace Abstract: Understanding and reasoning about entire software repositories is an essential capability for intelligent software engineering tools. While existing
arXiv:2512.07538v3 Announce Type: replace Abstract: Recognizing semantic differences across documents is crucial for text generation evaluation and content alignment, especially in cross-lingual setti
arXiv:2603.26380v2 Announce Type: replace Abstract: The attention mechanism has been the core component in modern transformer architectures. However, the computation of standard full attention scales
arXiv:2604.24346v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly deployed as evaluators in tasks requiring nuanced image understanding, yet their reliability in scoring
arXiv:2602.15603v2 Announce Type: replace Abstract: Models based on partial differential equations (PDEs) are powerful for describing a wide range of complex phenomena in the natural sciences. Accurat
arXiv:2604.23806v1 Announce Type: cross Abstract: The reverse process in score-based diffusion models is formally equivalent to overdamped Langevin dynamics in a time-dependent energy landscape. In ou
arXiv:2511.01490v2 Announce Type: replace Abstract: As synthetic data becomes widely used in language model development, understanding its impact on model behavior is crucial. This paper investigates
arXiv:2509.25346v2 Announce Type: replace Abstract: Predicting cellular responses to genetic perturbations represents a fundamental challenge in systems biology, critical for advancing therapeutic dis
arXiv:2604.24088v1 Announce Type: cross Abstract: Handling communication overhead in large-scale tensor-parallel training remains a critical challenge due to the dense, near-zero distributions of inte
arXiv:2604.23586v1 Announce Type: cross Abstract: Joint audio-video generation models have shown that unified generation yields stronger cross-modal coherence than cascaded approaches. However, existi
arXiv:2604.23703v1 Announce Type: cross Abstract: Slide-based teaching is widely used in higher education, yet in online, hybrid, and asynchronous contexts, slides often lose the instructor presence,
arXiv:2604.23623v1 Announce Type: new Abstract: Recent advancements in large language models (LLMs) have catalyzed the rise of reasoning-intensive inference paradigms, where models perform explicit st
arXiv:2604.23964v1 Announce Type: cross Abstract: Patients with dementia typically exhibit cognitive impairment, which is routinely assessed using the Mini-Mental State Examination (MMSE). Concurrentl
arXiv:2604.24005v1 Announce Type: cross Abstract: On-policy distillation (OPD) has shown strong potential for transferring reasoning ability from frontier or domain-specific models to smaller students
arXiv:2601.04204v2 Announce Type: replace-cross Abstract: The scalability of high-quality online education is hindered by the high costs and slow cycles of manual content creation. Despite advancement
arXiv:2604.23368v1 Announce Type: new Abstract: Event-Based Models (EBMs) infer biomarker progression from cross-sectional data but typically only as ordinal sequences and rely on rigid model assumpti
arXiv:2604.23205v1 Announce Type: cross Abstract: Deploying proprietary Deep Neural Networks (DNNs) on commodity edge devices demands hardware-backed Digital Rights Management (DRM) capable of withsta
arXiv:2509.00072v3 Announce Type: replace Abstract: Post-cutoff performance decay has been widely interpreted as a temporal signal for benchmark contamination. We critically examine this belief and de
arXiv:2601.21048v2 Announce Type: replace Abstract: Unsupervised neural combinatorial optimization (NCO) enables learning powerful solvers without access to ground-truth solutions. Existing approaches
arXiv:2604.22880v1 Announce Type: new Abstract: Existing document OCR largely targets plain text or Markdown, discarding the structural and executable properties that make LaTeX essential for scientif
arXiv:2604.22899v1 Announce Type: new Abstract: Industrial anomaly detection based on RGB-3D multimodal data has emerged as a mainstream paradigm for intelligent quality inspection. However, existing
arXiv:2309.06020v3 Announce Type: replace-cross Abstract: Technical debt refers to the consequences of sub-optimal decisions made during software development that prioritize short-term benefits over l
arXiv:2604.24459v1 Announce Type: new Abstract: Despite recent advances in text-to-image generation, models still struggle to accurately render prompt-specified text with correct spatial layout -- esp
arXiv:2604.24155v1 Announce Type: cross Abstract: The quest to align machine behavior with human values raises fundamental questions about the moral frameworks that should govern AI decision-making. M
arXiv:2604.24698v1 Announce Type: new Abstract: Applications based on large language models (LLMs), such as multi-agent simulations, require population diversity among agents. We identify a pervasive
arXiv:2604.23575v1 Announce Type: cross Abstract: Silicon samples are increasingly used as a low-cost substitute for human panels and have been shown to reproduce aggregate human opinion with high fid
arXiv:2602.11318v3 Announce Type: replace Abstract: In machine learning, 'ground truth' refers to the assumed correct labels used to train and evaluate models. However, the foundational 'ground truth'
arXiv:2512.08345v2 Announce Type: replace Abstract: Workplace toxicity is widely recognized as detrimental to organizational culture, yet quantifying its direct impact on operational efficiency remain
arXiv:2604.22767v1 Announce Type: cross Abstract: Ethical discourse on AI in healthcare has focused predominantly on back-end concerns such as bias, fairness and explainability, while the front-end in
arXiv:2604.24083v1 Announce Type: new Abstract: This study introduces the Kerimov-Alekberli model, a novel information-geometric framework that redefines AI safety by formally linking non-equilibrium