Diffusion Language Models for Speech Recognition
arXiv:2604.14001v1 Announce Type: new Abstract: Diffusion language models have recently emerged as a leading alternative to standard language models, due to their ability for bidirectional attention a
Knowledge catalogue
arXiv:2604.14001v1 Announce Type: new Abstract: Diffusion language models have recently emerged as a leading alternative to standard language models, due to their ability for bidirectional attention a
arXiv:2604.13366v1 Announce Type: new Abstract: Accurate modeling of robot dynamics is essential for model-based control, yet remains challenging under distributional shifts and real-time constraints.
arXiv:2604.13902v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has catalyzed significant advances in the reasoning capabilities of Large Language Models (LLMs).
arXiv:2509.21912v2 Announce Type: replace Abstract: Guidance provides a simple and effective framework for posterior sampling by steering the generation process towards the desired distribution. When
arXiv:2604.13509v1 Announce Type: new Abstract: Recent advances in video generation models has significantly accelerated video generation and related downstream tasks. Among these, video stylization h
arXiv:2604.13899v1 Announce Type: new Abstract: Instruction-tuned LLMs can annotate thousands of instances from a short prompt at negligible cost. This raises two questions for active learning (AL): c
arXiv:2604.13731v1 Announce Type: new Abstract: Multi-page Document Visual Question Answering requires reasoning over semantics, layouts, and visual elements in long, visually dense documents. Existin
arXiv:2604.13076v1 Announce Type: new Abstract: We investigate the robustness of value alignment via finetuning with synthetic documents, using animal compassion as a value that is both important in i
arXiv:2604.13230v1 Announce Type: new Abstract: Exploratory Landscape Analysis (ELA) provides numerical features for characterizing black-box optimization problems. In high-dimensional settings, howev
Boris Cherny shares productivity tips and insights from his recent experience using Opus 4.7, highlighting how the model has improved his workflow efficiency. The thread provides practical recommendat
arXiv:2604.14129v1 Announce Type: new Abstract: While Audio-Visual Language Models (AVLMs) have achieved remarkable progress over recent years, their reliability is bottlenecked by cross-modal halluci
Perplexity has released an updated Mac application featuring a Personal Computer feature that allows users to download the latest version. The update appears to enable enhanced AI capabilities directl
arXiv:2604.13797v1 Announce Type: new Abstract: Few-shot Font Generation aims to generate stylistically consistent glyphs from a few reference glyphs. However, capturing complex font styles from a few
arXiv:2604.13796v1 Announce Type: cross Abstract: In daily fantasy sports (DFS), match participation is highly time-sensitive. Users must act within a narrow window before a game begins, making match
arXiv:2604.13278v1 Announce Type: new Abstract: Aerial object detection in UAV imagery presents unique challenges due to the high prevalence of tiny objects, adverse environmental conditions, and stri
arXiv:2604.13878v1 Announce Type: new Abstract: Driver drowsiness significantly impairs the ability to accurately judge safe braking distances and is estimated to contribute to 10%-20% of road acciden
arXiv:2604.14030v1 Announce Type: new Abstract: Product bundling boosts e-commerce revenue by recommending complementary item combinations. However, existing methods face two critical challenges: (1)
DevOps automation company DuploCloud Inc. today announced that it has completed its SOC 2 Type II audit and achieved ISO/IEC 42001 certification. A SOC 2 Type II examination evaluates the design and o
arXiv:2602.20981v3 Announce Type: replace Abstract: Scaling multimodal alignment between video and audio is challenging, particularly due to limited data and the mismatch between text descriptions and
Ecom-RLVE is a framework for training e-commerce conversational agents using reinforcement learning with verifiable environments that can adapt to different scenarios. The system enables agents to lea
arXiv:2604.13586v1 Announce Type: new Abstract: Existing multi-view three-dimensional (3D) object detection approaches widely adopt large-scale pre-trained vision transformer (ViT)-based foundation mo
arXiv:2604.13800v1 Announce Type: new Abstract: Embodied AI research is increasingly moving beyond single-task, single-environment policy learning toward multi-task, multi-scene, and multi-model setti
arXiv:2604.13685v1 Announce Type: cross Abstract: Deep learning-based surface electromyography (sEMG) gesture recognition is frequently bottlenecked by data scarcity and limited subject diversity. Whi
arXiv:2604.13371v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly described as possessing strong reasoning capabilities, supported by high performance on mathematical, logi
arXiv:2604.13677v1 Announce Type: new Abstract: Mobile robots joining public spaces like sidewalks must care for pedestrian comfort. Many studies consider pedestrians' objective safety, for example, b
arXiv:2510.20169v3 Announce Type: replace Abstract: Traveling Salesman Problem (TSP) is a classic NP-hard problem that has garnered significant attention from both academia and industry. While neural-
arXiv:2604.13286v1 Announce Type: new Abstract: Despite the widespread multilingual deployment of large language models, post-training pipelines remain predominantly English-centric, contributing to p
arXiv:2604.13491v1 Announce Type: new Abstract: With the rapid progress of Multimodal Large Language Models (MLLMs), unified MLLMs that jointly perform image understanding and generation have advanced
arXiv:2604.13271v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly applied to complex telecommunications tasks, including 3GPP specification analysis and O-RAN network troub
arXiv:2604.13508v1 Announce Type: new Abstract: Sparse Upcycling provides an efficient way to initialize a Mixture-of-Experts (MoE) model from pretrained dense weights instead of training from scratch
arXiv:2604.13598v1 Announce Type: new Abstract: Recent reinforcement learning (RL) approaches have advanced radiology report generation (RRG), yet two core limitations persist: (1) report-level reward
arXiv:2604.13633v1 Announce Type: new Abstract: Coordinating navigation and manipulation with robust performance is essential for embodied AI in complex indoor environments. However, as tasks extend o
arXiv:2604.13410v1 Announce Type: cross Abstract: We study the problem of estimating the effect function for a continuous treatment, which maps each treatment value to a population-averaged outcome. A
arXiv:2602.24119v2 Announce Type: replace Abstract: Purpose: This study evaluates the quality of commercial large language model (LLM) machine translation (MT) for Ancient Greek technical prose and be
arXiv:2604.13882v1 Announce Type: new Abstract: The evaluation of supervised machine learning models is a critical stage in the development of reliable predictive systems. Despite the widespread avail
arXiv:2604.13232v1 Announce Type: new Abstract: This discussion paper re-examines SemEval-2020 Task 1, the most influential shared benchmark for lexical semantic change detection, through a three-part
arXiv:2604.02709v2 Announce Type: replace Abstract: The formal reasoning capabilities of LLMs are crucial for advancing automated software engineering. However, existing benchmarks for LLMs lack syste
arXiv:2604.13071v1 Announce Type: new Abstract: We introduce Earth Virtual Expert (EVE), the first open-source, end-to-end initiative for developing and deploying domain-specialized LLMs for Earth Int
arXiv:2604.13426v1 Announce Type: new Abstract: Existing Vision Mamba-based RGB-Event(RGBE) tracking methods suffer from using static state transition matrices, which fail to adapt to variations in ev
arXiv:2604.13327v1 Announce Type: cross Abstract: Modern GPU workloads, especially large language model (LLM) inference, suffer from kernel launch overheads and coarse synchronization that limit inter
Grok Imagine is an image generation AI tool that creates realistic images with claimed ease of use. The announcement, promoted by Elon Musk on X, highlights the capability to generate high-quality vis
arXiv:2604.13533v1 Announce Type: cross Abstract: Achieving general-purpose robotics requires empowering robots to adapt and evolve based on their environment and feedback. Traditional methods face li
Exactly Most people think of philosophy as an abstraction that doesn't touch the real world, but they're wrong. Most real world problems are philosophy problems, and most philosophy problems are 'givi
Cristobal Valenzuela announced the launch of 'This Show Rocks,' a new late-night streaming show at Runway that broadcasts nightly at 10 PM PT. The announcement was made via X (formerly Twitter), thoug
arXiv:2604.13279v1 Announce Type: new Abstract: Fall detection in elderly care requires not only accurate classification but also reliable explanations that clinicians can trust. However, existing pos
arXiv:2604.13050v1 Announce Type: cross Abstract: Urban areas are intricate systems shaped by socioeconomic, environmental, and infrastructural factors, with land use patterns serving as aspects of ur
arXiv:2601.06536v2 Announce Type: replace Abstract: We present Exposia, the first public dataset that connects writing and feedback in higher education, enabling research on educationally grounded com
arXiv:2601.08605v2 Announce Type: replace Abstract: Experience intervention in web agents emerges as a promising technical paradigm, enhancing agent interaction capabilities by providing valuable insi
arXiv:2601.11329v3 Announce Type: replace Abstract: Spoken conversational systems require more than accurate speech generation to have human-like conversations: to feel natural and engaging, they must
arXiv:2604.13788v1 Announce Type: cross Abstract: Imitation learning (IL) policies in robotics deliver strong performance in controlled settings but remain brittle in real-world deployments: rare even
arXiv:2509.18847v3 Announce Type: replace-cross Abstract: Tool-augmented large language models (LLMs) are usually trained with supervised imitation or coarse-grained reinforcement learning that optimi
In a surprising turn of events, the struggling sneaker company that once tried to make shoes eco-sustainable announced today that it has metamorphosed into artificial intelligence firm. The San Franci
arXiv:2604.13453v1 Announce Type: new Abstract: Traffic forecasting requires modeling complex temporal dynamics and long-range spatial dependencies over large sensor networks. Existing methods typical
arXiv:2505.12600v3 Announce Type: replace-cross Abstract: We study the densest subgraph problem and its NP-hard densest at-most-k subgraph variant through the lens of learning-augmented algorithms. We
arXiv:2405.20836v3 Announce Type: replace-cross Abstract: Solving time-dependent Partial Differential Equations (PDEs) is one of the most critical problems in computational science. While Physics-Info
arXiv:2604.13191v1 Announce Type: cross Abstract: Many materials show anisotropic light scattering patterns due to the shape and local alignment of their underlying micro structures: surfaces with sma
arXiv:2508.05153v2 Announce Type: replace Abstract: Category-level generalization for robotic garment manipulation, such as bimanual smoothing, remains a significant hurdle due to high dimensionality,
arXiv:2604.14025v1 Announce Type: new Abstract: Reconstructing 3D representations from 2D inputs is a fundamental task in computer vision and graphics, serving as a cornerstone for understanding and i
arXiv:2508.08791v3 Announce Type: replace Abstract: Effective tool use is essential for large language models (LLMs) to interact with their environment. However, progress is limited by the lack of eff
arXiv:2505.19662v3 Announce Type: replace-cross Abstract: This paper introduces FieldWorkArena, a benchmark for agentic AI targeting real-world field work. With the recent increase in demand for agent