SLAM-Former: Putting SLAM into One Transformer
arXiv:2509.16909v2 Announce Type: replace Abstract: We present SLAM-Former, a neural approach that integrates full SLAM capabilities into a single transformer. Similar to traditional SLAM systems, SLA
Knowledge catalogue
arXiv:2509.16909v2 Announce Type: replace Abstract: We present SLAM-Former, a neural approach that integrates full SLAM capabilities into a single transformer. Similar to traditional SLAM systems, SLA
arXiv:2607.23121v1 Announce Type: cross Abstract: Dynamic Product Ads (DPA) require retrieving relevant items from multi-million product catalogs, balancing two competing objectives: retargeting (re-s
arXiv:2607.23751v1 Announce Type: new Abstract: The usefulness of a variational autoencoder (VAE) depends on two properties of its latent space that are hard to obtain together: high encoding capacity
arXiv:2607.24027v1 Announce Type: new Abstract: Diffusion transformers are essential for high-fidelity video generation, but long token sequences make attention a dominant inference bottleneck. Traini
arXiv:2607.24645v1 Announce Type: cross Abstract: The wide-scale use of sparse autoencoders (SAEs) as interpretability tools is limited by inconsistent links between SAE features and model behavior. F
arXiv:2607.16826v2 Announce Type: replace Abstract: Generalised object counting aims to estimate the number of instances of an arbitrary object category from a single image, but many recent methods ca
arXiv:2607.22931v1 Announce Type: cross Abstract: Analytic Continual Learning (ACL) offers a computationally efficient alternative to gradient-based approaches. Recent ACL methods are based on Recursi
arXiv:2607.23037v1 Announce Type: new Abstract: Large language models (LLMs) can predict interpersonal attraction from conversation transcripts, but it remains unclear what a speech predictor can add
arXiv:2607.23341v1 Announce Type: cross Abstract: Pretrained deep denoisers can be used to solve a wide range of model-based image reconstruction tasks via Plug-and-Play (PnP) and Regularization-by-De
arXiv:2607.24686v1 Announce Type: cross Abstract: Variational quantum circuits have been central to many proposed near-term applications of quantum computing, but a growing body of evidence suggests t
arXiv:2607.23273v1 Announce Type: cross Abstract: Computational nutrition needs precise ingredient data, but current databases are incomplete, inconsistent, and built for human reference rather than a
arXiv:2607.23840v1 Announce Type: new Abstract: Facial movement and expression are central to face-to-face communication, conveying turn-taking, attention, agreement, and engagement alongside speech.
arXiv:2505.04260v3 Announce Type: replace-cross Abstract: Personalizing LLM responses typically requires users to articulate their preferences through prompting, which can be burdensome at cold start
arXiv:2607.24393v1 Announce Type: cross Abstract: Finite-time driving of stochastic systems generates excess dissipation, causing the evolving probability distribution to lag behind the instantaneous
arXiv:2607.24009v1 Announce Type: new Abstract: Matricized low-rank approximation via SVD is a standard surrogate for tensor decompositions, but entry-wise reconstruction error fails to capture multiw
arXiv:2607.22725v1 Announce Type: cross Abstract: Data augmentation is routinely used to improve generalization in image classification, but the assumptions underlying standard policies are poorly mat
arXiv:2607.22592v1 Announce Type: new Abstract: Graph-based retrieval-augmented generation (GraphRAG) grounds answers in structured knowledge, but current systems extract entities and relationships ex
arXiv:2607.24288v1 Announce Type: new Abstract: Quadratic unconstrained binary optimization (QUBO) has emerged as a powerful framework for medical computing problems. Binary decision variables natural
arXiv:2606.18816v2 Announce Type: replace-cross Abstract: Hybrid brain-computer interfaces (BCIs) that integrate motor imagery (MI) and steady-state visual evoked potentials (SSVEP) provide high-dimen
arXiv:2607.24539v1 Announce Type: new Abstract: Multimodal large language models (LLMs) can combine topology, measurements, and incident text for grid diagnosis, yet answer accuracy does not establish
arXiv:2607.24304v1 Announce Type: cross Abstract: We formulate a statistical physics framework to model a networked stochastic dynamical system exhibiting bistability, driven by additive noise and soc
arXiv:2607.22575v1 Announce Type: new Abstract: Human episodic memory supports the retrieval of experiences that unfold over extended timescales, yet the computational mechanisms underlying this abili
arXiv:2607.22674v1 Announce Type: cross Abstract: Tactile graphics are a primary medium for blind and low-vision (BLV) individuals to access non-textual information. However, they are difficult to sca
arXiv:2607.23512v1 Announce Type: new Abstract: Offensive language detection models generally suffer performance degradation when deployed across datasets and across languages, yet most existing studi
This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. OpenAI called the Hugging Face attack unprecedented. But we’ve
arXiv:2602.15491v2 Announce Type: replace-cross Abstract: Neural audio codecs (NACs) typically encode the short-term energy (gain) and normalized structure (shape) of speech/audio signals jointly with
arXiv:2607.23531v1 Announce Type: new Abstract: Joint-Embedding Predictive Architectures (JEPAs) are effective for images, video, and audio, yet deterministic JEPA-style latent prediction has not beco
arXiv:2607.24537v1 Announce Type: new Abstract: In the Big Data era, the scalability of clustering algorithms constitutes a key challenge. Traditional density-based methods (e.g., DBSCAN) offer robust
How we built the OlmoEarth Platform to fine-tune geospatial models and run continent-scale satellite inference while managing massive data pipelines, distributed compute, and automatically recovering
arXiv:2607.24041v1 Announce Type: cross Abstract: Over-parameterized linear regression has been widely studied over the last decade. However, most existing works assume that the covariates are indepen
arXiv:2607.23492v1 Announce Type: new Abstract: Concept erasure techniques (CETs) edit text-to-image diffusion models to erase undesired targets such as NSFW content or copyrighted styles, while prese
arXiv:2607.23493v1 Announce Type: cross Abstract: Automated analysis of multimodal content on social networks has become a critical task for understanding public sentiment and information diffusion in
arXiv:2607.22717v1 Announce Type: new Abstract: While Implicit Neural Representations (INRs) and dynamic 3D Gaussian Splatting (3DGS) achieve impressive results in video processing, they often fall sh
arXiv:2607.22574v1 Announce Type: new Abstract: Evidence-based clinical decision making requires specialists to identify, evaluate and synthesize relevant scientific literature. However, PubMed search
arXiv:2607.23286v1 Announce Type: new Abstract: Automatic feature engineering (AutoFE) for tabular learning can be naturally formulated as a program synthesis problem, where the objective is to discov
arXiv:2607.23509v1 Announce Type: new Abstract: We present two approaches for predicting tennis match outcomes using topological data analysis and graph theory on ATP singles matches from 2000-2025. T
arXiv:2607.17799v2 Announce Type: replace Abstract: This study systematically benchmarks different data augmentation setups across the baseline YOLO26 model size variants to determine the most effecti
arXiv:2505.11421v2 Announce Type: replace Abstract: This work explores the journey towards achieving Bahnaric-Vietnamese translation for the sake of culturally bridging the two ethnic groups in Vietna
arXiv:2607.24082v1 Announce Type: new Abstract: Recent advances in AI have substantially expanded its cognitive and reasoning capabilities. From the perspective of semantic complexity, the development
arXiv:2607.23347v1 Announce Type: cross Abstract: Trainable denoisers with Lipschitz control have become central to convergent image reconstruction. However, training neural networks that simultaneous
arXiv:2301.03709v3 Announce Type: replace-cross Abstract: Consistent and holistic expression of software requirements is important for the success of software projects. In this study, we aim to enhanc
arXiv:2607.24215v1 Announce Type: new Abstract: Although general text-to-image models excel in open-domain generation, their performance degrades significantly in specialized downstream domains, parti
arXiv:2607.24298v1 Announce Type: new Abstract: Recent 3D foundation models can generate high-quality assets from a single image, but degrade markedly on unconstrained multi-image inputs, often produc
arXiv:2607.23854v1 Announce Type: new Abstract: Humans often find good solutions to combinatorial optimization problems that are computationally hard even for advanced computer algorithms. In the Eucl
arXiv:2411.03048v2 Announce Type: replace-cross Abstract: The rapid growth of UAV applications necessitates a robust communication and networking system architecture capable of addressing the diverse
arXiv:2607.24157v1 Announce Type: new Abstract: Modern computer vision pipelines remain fragmented, with tasks such as text-to-image generation, editing, restoration, and classical perception handled
arXiv:2607.22673v1 Announce Type: cross Abstract: We present URHead, a unified representation for high-fidelity and animatable head avatars that fundamentally redefines mesh-Gaussian integration. Whil
arXiv:2607.22979v1 Announce Type: cross Abstract: Deep neural networks have been widely used in many applications (e.g., computer vision and natural language processing); however, understanding their
arXiv:2607.23198v1 Announce Type: new Abstract: We propose Variance-Preserving Orthogonal Selection (VPOS), a greedy framework for unsupervised feature selection that operates in the weighted PCA load
arXiv:2607.23634v1 Announce Type: cross Abstract: Attention enables context modeling via query-key scoring with softmax normalization. Driven by industrial long-context demands, mainstream research ha
arXiv:2607.23581v1 Announce Type: new Abstract: Multimodal misinformation verification is challenging because misleading signals may come from different parts of a post and require different forms of
arXiv:2607.24124v1 Announce Type: new Abstract: We introduce ViDS, a Video Diffusion Shader that leverages 3D face tracking for expressive and identity-preserving portrait animation. We first reconstr
arXiv:1502.05678v3 Announce Type: replace Abstract: People preserve memories of events such as birthdays, weddings, or vacations by capturing photos, often depicting groups of people. Invariably, some
arXiv:2512.01031v2 Announce Type: replace-cross Abstract: Vision-Language-Action models (VLAs) are becoming increasingly capable across diverse robotic tasks. However, these models are typically deplo
arXiv:2602.17394v2 Announce Type: replace-cross Abstract: Unmanned Aerial Vehicle (UAV)-assisted networks are increasingly foreseen as a promising approach for emergency response, providing rapid, fle
arXiv:2607.23265v1 Announce Type: new Abstract: Existing Large Vision-Language Models (LVLMs) struggle with long-form video understanding due to the quadratic computational cost of visual tokens. Whil
arXiv:2607.23594v1 Announce Type: new Abstract: In prostate cancer histopathology, the Gleason Score is determined by the most frequent (Primary) and second most frequent (Secondary) Gleason patterns
arXiv:2607.24484v1 Announce Type: cross Abstract: This paper studies what discriminatively trained reward models (RMs) memorize by measuring counterfactual memorization on two human preference dataset
arXiv:2607.22781v1 Announce Type: cross Abstract: High task performance does not show whether a model retains prediction-relevant structural information in its internal representation. Temporal graph
arXiv:2607.23390v1 Announce Type: new Abstract: When can additional low-bit residual computation replace missing numerical precision for a fixed input-output map? We model a quantized residual system