HRM-Text: Efficient Pretraining Beyond Scaling
arXiv:2605.20613v1 Announce Type: new Abstract: The current pretraining paradigm for large language models relies on massive compute and internet-scale raw text, creating a significant barrier to foun
Knowledge catalogue
arXiv:2605.20613v1 Announce Type: new Abstract: The current pretraining paradigm for large language models relies on massive compute and internet-scale raw text, creating a significant barrier to foun
arXiv:2601.22538v2 Announce Type: replace Abstract: Learning-to-defer (L2D) routes each decision to a system's own predictor or to an external expert. Streaming time-series settings break the offline-
arXiv:2605.20866v1 Announce Type: new Abstract: Communication is a major bottleneck in distributed learning, especially in large-scale settings and in federated learning environments with slow links.
arXiv:2605.20989v1 Announce Type: new Abstract: Single-cell RNA sequencing provides insights into gene expression at single-cell resolution, yet inferring temporal processes from these static snapshot
arXiv:2605.20396v1 Announce Type: new Abstract: Identifying latent variables and the causal structure involving them is essential across various scientific fields. While many existing works fall under
arXiv:2605.20712v1 Announce Type: new Abstract: Automatic speech recognition replaces typing only when correction costs less than manual entry, a threshold determined by error types, not counts: fixin
arXiv:2605.21468v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a dominant paradigm for improving reasoning in large language models (LLMs), yet the
arXiv:2404.07106v2 Announce Type: replace Abstract: Point cloud completion aims to generate a complete and high-fidelity point cloud from an initially incomplete and low-quality input. A prevalent str
Google I/O 2026 featured several new AI model releases including Gemini 3.5 Flash, an Omni model codenamed NanoBanana for video processing, Spark for background agent tasks, and Antigravity 2.0. These
arXiv:2504.08381v2 Announce Type: replace-cross Abstract: Epileptic seizures are transient neurological events characterized by abnormal and excessive neuron activity in the brain, which are often ass
arXiv:2605.19620v1 Announce Type: new Abstract: LiDAR-based 3D human motion capture has broad applications in fields such as autonomous driving and robotics, where accurate motion reconstruction is cr
arXiv:2605.19172v1 Announce Type: cross Abstract: Forecasting urban delivery demand becomes substantially more challenging when newly added service regions lack historical records. Existing spatiotemp
arXiv:2605.20176v1 Announce Type: new Abstract: Large language models (LLMs) and agentic systems have shown promise for clinical decision support, but existing works largely assume that evidence has a
arXiv:2510.13293v3 Announce Type: replace Abstract: While Text-to-Speech (TTS) systems enable emotional control via natural-language instructions, expressiveness, naturalness, and speech quality degra
arXiv:2605.18815v1 Announce Type: new Abstract: Modern large language model (LLM) training is inherently dynamic: resource fluctuations, RLHF phase shifts, and cluster elasticity continually reshape t
arXiv:2605.18908v1 Announce Type: cross Abstract: Deep neural networks (DNNs) remain critically vulnerable to backdoor attacks. Existing post-training detectors often require clean or surrogate data,
arXiv:2603.03140v3 Announce Type: replace-cross Abstract: AI agents are increasingly active on social media platforms, generating content and interacting with one another at scale. Yet the behavioral
arXiv:2603.29183v2 Announce Type: replace-cross Abstract: Open-set anomaly detection (OSAD) is an emerging paradigm designed to utilize limited labeled data from anomaly classes seen in training to id
arXiv:2411.08982v3 Announce Type: replace Abstract: Selective parameter activation provided by Mixture-of-Expert (MoE) models have made them a popular choice in modern foundational models. However, Mo
arXiv:2509.14839v2 Announce Type: replace Abstract: City administrations increasingly rely on comprehensive databases and urban digital twins of city assets, such as traffic signs and trees, as well a
arXiv:2605.17326v1 Announce Type: cross Abstract: We investigate the role of the noise schedule in diffusion processes on Lie groups, with particular emphasis on applications to lattice gauge theory.
arXiv:2605.19328v1 Announce Type: cross Abstract: Recent advances in Vision-Language Models (VLMs) facilitate a new class of embodied AI systems, where these models are integrated into physical platfo
arXiv:2605.18797v1 Announce Type: cross Abstract: Scaling model performance typically requires increasing model size. Looped Transformer offers a compelling alternative by iteratively reusing the same
arXiv:2605.18800v1 Announce Type: cross Abstract: Post-training quantization has emerged as a widely adopted technique for compressing and accelerating the inference of Large Language Models (LLMs). T
arXiv:2605.19289v1 Announce Type: new Abstract: Driven by rapid advances in large-scale generative models, synthetic data has emerged as a promising solution for visual understanding. While modern dif
arXiv:2605.18083v1 Announce Type: new Abstract: Expanding Large Language Models~(LLMs) to new languages is a costly endeavor, demanding extensive Continued Pre-Training~(CPT) and data-intensive alignm
arXiv:2605.16902v1 Announce Type: new Abstract: Scientific artifacts such as models and datasets are foundations for research. With the rapid growth of platforms like HuggingFace, researchers now have
arXiv:2605.17971v1 Announce Type: cross Abstract: Despite rigorous safety alignment, Large Language Models (LLMs) remain vulnerable to jailbreak attacks. Existing black-box methods often rely on heuri
arXiv:2505.03631v4 Announce Type: replace Abstract: Video quality assessment (VQA) is essential for quantifying perceptual quality in various video processing workflows, spanning from camera capture s
arXiv:2605.18226v1 Announce Type: cross Abstract: Modern large language model (LLM) applications increasingly rely on long conditioning prefixes to control model behavior at inference time. While pref
arXiv:2508.04227v2 Announce Type: replace Abstract: Vision-language models (VLMs) and the recent surge of Multimodal Large Language Models (MLLMs) have revolutionized artificial intelligence with unpr
arXiv:2605.16298v1 Announce Type: cross Abstract: While traditional AI and data-driven facilities management approaches have improved building operational efficiency, they remain constrained by centra
Artificial intelligence developer Decart.ai Inc. today announced that it has raised 300 million in funding at a nearly 4 billion valuation. Radical Ventures led the round with participation from Nvidi
arXiv:2605.16444v1 Announce Type: cross Abstract: Accurate intraoperative and postoperative diagnosis of spread through air spaces (STAS) is essential for guiding surgical decisions and postoperative
arXiv:2605.16776v1 Announce Type: cross Abstract: Mitigating sensitive and harmful outputs is fundamental to ensuring safe deployment of LLMs. Existing approaches typically follow two paradigms: Knowl
arXiv:2605.18562v1 Announce Type: cross Abstract: Accurate estimates of item difficulty are essential for valid assessment and effective adaptive learning. However, for newly created tasks, response d
arXiv:2605.18554v1 Announce Type: new Abstract: Federated Bayesian neural networks require fixing a prior on the model parameters together with a likelihood. Eliciting meaningful priors on the weight
Gemini 3.5 is Google's latest family of AI models combining frontier intelligence with action, representing a major leap forward in building more capable, intelligent agents. The first model in the se
Google LLC today introduced two new generative artificial intelligence models that push its Gemini family further into AI agents and multimodal creation: Gemini 3.5 Flash, a fast reasoning model desig
arXiv:2605.18387v1 Announce Type: cross Abstract: Graph Neural Networks (GNNs) and Graph Transformers (GTs) are now a fundamental paradigm for graph learning, combining the representation-learning cap
arXiv:2605.17856v1 Announce Type: new Abstract: Process-based simulation models encode decades of scientific understanding across the Earth sciences, yet the communities most exposed to climate risk a
arXiv:2605.18541v1 Announce Type: new Abstract: Modeling hyperspectral imagery (HSI) across different sensors presents a fundamental challenge due to variations in wavelength coverage, band sampling,
arXiv:2605.16786v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly needed for interactive mobile applications, but high-quality models exceed the limited DRAM available on s
arXiv:2605.18449v1 Announce Type: cross Abstract: Understanding customer movement within retail spaces is essential for optimizing store layouts. Real-world trajectory data can provide highly accurate
arXiv:2509.02351v3 Announce Type: replace-cross Abstract: Labeled data is a fundamental component in training supervised deep learning models for computer vision tasks. However, the labeling process,
arXiv:2605.17916v1 Announce Type: new Abstract: Generating a consistent whole-house VR tour from a floorplan and style reference requires both photorealistic panoramas and cross-view spatial coherence
arXiv:2605.18464v1 Announce Type: new Abstract: Contrastively trained vision-language models such as CLIP provide strong zero-shot transfer by aligning images and text in a shared embedding space. How
arXiv:2605.18474v1 Announce Type: cross Abstract: The widespread deployment and redistribution of large language models (LLMs) have made model provenance tracking a critical challenge. While existing
arXiv:2512.19134v2 Announce Type: replace Abstract: Dynamic Retrieval-Augmented Generation adaptively determines when to retrieve during generation to mitigate hallucinations in large language models
arXiv:2605.16874v1 Announce Type: new Abstract: Large reasoning models (LRMs) substantially outperform their base LLM counterparts on challenging reasoning benchmarks, yet it remains poorly understood
arXiv:2603.11276v2 Announce Type: replace-cross Abstract: Real-world contextual bandit problems with complex reward models are often tackled with iteratively trained models, such as boosting trees. Ho
arXiv:2605.17522v1 Announce Type: new Abstract: Planning and acting in 3D environments is a fundamental capability for robotic manipulation in the real world. Although prior work has explored predicti
arXiv:2605.17710v1 Announce Type: new Abstract: Although modern multilingual Automatic Speech Recognition (ASR) systems support several Nigerian languages, their performance consistently lags behind h
arXiv:2605.18063v1 Announce Type: new Abstract: Object counting is a foundational vision task with over a decade of dedicated research, yet state-of-the-art models still fail systematically in the mix
arXiv:2605.17704v1 Announce Type: new Abstract: The lottery ticket hypothesis posits that dense networks contain sparse subnetworks, ``winning tickets,'' that, when rewound to their initial weights an
arXiv:2605.17285v1 Announce Type: cross Abstract: Node representation learning, such as Graph Neural Networks (GNNs), has emerged as a pivotal method in machine learning. The demand for reliable expla
arXiv:2605.16547v1 Announce Type: new Abstract: Semantic communication has emerged as a promising paradigm for enabling goal-oriented networking. However, most existing semantic communication solution
I cannot provide an accurate summary of this entry as the title appears to be truncated mid-sentence and lacks sufficient context. Based on the incomplete title, it likely discusses how software speci
arXiv:2605.15217v1 Announce Type: new Abstract: Instruction-tuned language models exhibit behavioural fairness in high-stakes decisions while retaining biased associations in their internal representa
arXiv:2605.15290v1 Announce Type: cross Abstract: Hyperparameter transfer across model architectures dramatically reduces the amount of compute necessary for tuning large language models (LLMs). The m