Leveraging Similarities in Multi-Armed Bandits
arXiv:2606.23414v1 Announce Type: new Abstract: In many online learning and bandit problems, the actions we consider possess inherent similarities--for instance because they share latent traits, tags,
Knowledge catalogue
arXiv:2606.23414v1 Announce Type: new Abstract: In many online learning and bandit problems, the actions we consider possess inherent similarities--for instance because they share latent traits, tags,
arXiv:2606.21358v1 Announce Type: new Abstract: Supervised video pretraining is a common transfer learning practice for improving downstream action recognition performance. However, it requires large-
arXiv:2606.23686v1 Announce Type: new Abstract: Despite the impressive manipulation capabilities of Vision-Language-Action (VLA) models, their operational safety under strict constraints remains large
arXiv:2606.23688v1 Announce Type: new Abstract: Reconstructing dynamic non-rigid objects from monocular video requires integrating visual cues from direct observations with data-driven priors over geo
arXiv:2606.21564v1 Announce Type: new Abstract: Transformers achieve strong performance, but their internal computations remain opaque. We view each Transformer layer as a dynamic graph whose nodes ar
arXiv:2606.22481v1 Announce Type: cross Abstract: 3D Gaussian Splatting (3DGS) is widely used to capture and render real scenes. Compositing objects from one capture into another has applications in m
arXiv:2412.05976v2 Announce Type: replace Abstract: Occupancy prediction has garnered increasing attention in recent years for its comprehensive fine-grained environmental representation and strong ge
arXiv:2606.23539v1 Announce Type: new Abstract: Visual document retrieval requires rapidly locating relevant pages from large multi-modal corpora in response to user queries. While recent methods powe
arXiv:2606.21292v1 Announce Type: new Abstract: We present Casper3D, a lightweight probabilistic framework for converting noisy multi-view 2D foundation-model embeddings into a latent 3D semantic repr
arXiv:2606.23653v1 Announce Type: new Abstract: Accurate volume and surface area estimation is critical for diverse applications, from marine ecology to medical diagnostics. However, existing methods
arXiv:2606.22499v1 Announce Type: cross Abstract: This study presents the hardware and software architecture of a transformative system for illuminating line drawings and letterforms. These mid-air il
arXiv:2606.21018v1 Announce Type: cross Abstract: As artificial intelligence advances into the era of Embodied AI, live musical interaction urgently needs to break free from the limitations of offline
arXiv:2606.23136v1 Announce Type: cross Abstract: Finding the shortest path in non-geometric network graphs, where edge weights encode arbitrary metrics such as latency or monetary cost rather than sp
arXiv:2511.22354v2 Announce Type: replace Abstract: This paper introduces CoMuRoS (Collaborative Multi-Robot System), a generalizable hierarchical architecture for heterogeneous robot teams that unifi
arXiv:2606.20729v1 Announce Type: cross Abstract: Quantum chemistry simulations underpin modern materials discovery, yet their impact is limited by steep computational cost and dependence on fixed app
arXiv:2606.22013v1 Announce Type: new Abstract: Machine learning (ML) model serving has become a dominant consumer of GPU infrastructure, yet capacity planning in these systems remains largely ad hoc.
arXiv:2606.21821v1 Announce Type: new Abstract: Understanding the causal structure of a language model's thought process is a problem of significant importance for both transparency and safety. In thi
arXiv:2410.02548v4 Announce Type: replace-cross Abstract: Flow Matching (FM) is a simulation-free method for learning a continuous, invertible flow that interpolates between two distributions, and in
arXiv:2603.14343v2 Announce Type: replace Abstract: Large Audio-Language Models (LALMs) have shown strong performance in speech understanding, making speech a natural interface for accessing factual i
arXiv:2606.22772v1 Announce Type: new Abstract: Lip-syncing deepfakes are among the most challenging forms of manipulated media because their artifacts are localized almost exclusively to the mouth re
arXiv:2606.21527v1 Announce Type: cross Abstract: Robust obstacle segmentation is essential for the safety of intelligent robots, where LiDAR-based perception systems play a fundamental role in the ro
arXiv:2606.23110v1 Announce Type: cross Abstract: Outer-loop link adaptation (OLLA) is widely deployed in 5G NR to track channel variations, yet its reliance on first-order, single-bit feedback degrad
arXiv:2606.21387v1 Announce Type: new Abstract: Legged-wheeled robots have long been studied for their potential to combine the efficient flat-ground mobility of wheels with the rough-terrain capabili
arXiv:2606.21968v1 Announce Type: new Abstract: Vision-Language Models (VLMs) struggle as query-relevant objects become smaller. To address this, recent training-free approaches dynamically retrieve a
arXiv:2606.22565v1 Announce Type: cross Abstract: Chain-of-Thought (CoT) has become a standard method for improving reasoning capabilities in large language models (LLMs) by eliciting step-by-step thi
arXiv:2606.13912v2 Announce Type: replace-cross Abstract: Complex neural quantum states are difficult to optimize when their wavefunction phase carries gauge, chiral, fermionic, or topological structu
arXiv:2606.23249v1 Announce Type: new Abstract: Humanoid local navigation in cluttered environments must jointly resolve obstacle avoidance, sparse-goal recovery, and stable whole-body locomotion unde
arXiv:2606.22662v1 Announce Type: new Abstract: Forecasting chaotic dynamical systems such as the Lorenz attractor is notoriously difficult: small numerical errors are amplified exponentially over lon
arXiv:2606.23118v1 Announce Type: new Abstract: Low-light human action recognition remains a challenging problem due to poor illumination, amplified noise, motion ambiguity, and diverse real-world sce
arXiv:2509.23729v3 Announce Type: replace Abstract: Large Language Models (LLMs) with multimodal capabilities have revolutionized vision-language tasks, but their deployment often requires huge memory
arXiv:2304.12319v2 Announce Type: replace-cross Abstract: Recently, numerous end-to-end optimized image compression neural networks have been developed and proved themselves as leaders in rate-distort
arXiv:2606.20874v1 Announce Type: new Abstract: Cryopathy syndromes are difficult to classify because laboratory patterns often overlap across diagnostic categories, while some diagnoses are rare. Thi
arXiv:2606.20641v1 Announce Type: cross Abstract: Multi-modal Large Language Models (MLLMs) have demonstrated remarkable capabilities in semantic understanding and common sense reasoning, making them
arXiv:2412.20034v2 Announce Type: replace Abstract: Continual test-time domain adaptation (CTTA) aims to adjust pre-trained source models to perform well over time across non-stationary target environ
arXiv:2606.23126v1 Announce Type: new Abstract: While recent advancements in anomaly detection have demonstrated the efficacy of CNN- and Transformer-based approaches, these architectures face inheren
arXiv:2606.21119v1 Announce Type: new Abstract: Mammography is an essential tool for breast cancer detection, with millions of examinations conducted annually. However, publicly available high-quality
arXiv:2606.21461v1 Announce Type: new Abstract: The Manipulider is a buoyancy-actuated underwater robot that enables thrusterless, glide-like locomotion and attitude-based manipulation, while providin
arXiv:2606.22597v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly used to read maps for logistics, delivery, and accessible navigation, where the output is an actionable d
arXiv:2606.22543v1 Announce Type: new Abstract: Humans localize places by integrating perceptual cues from vision with semantic reasoning from language, forming a scene understanding that is both intu
arXiv:2606.22649v1 Announce Type: new Abstract: Foundation models provide highly descriptive representations for medical images, yet their reliability degrades under distribution shifts arising from c
arXiv:2606.23664v1 Announce Type: new Abstract: Multi-agent systems (MAS) offer a scalable path forward for agentic AI, comprising multiple LLM-based agents, each assigned a system prompt and a positi
arXiv:2606.20743v1 Announce Type: new Abstract: Trained transformers reliably develop massive activations, a small number of hidden dimensions whose magnitude is far above the median and which concent
arXiv:2606.21830v1 Announce Type: new Abstract: Reinforcement learning from verifiable rewards (RLVR) has driven rapid progress in mathematical and code reasoning, but when extended to science, existi
arXiv:2602.15206v2 Announce Type: replace Abstract: Reward learning typically relies on a single feedback type or combines multiple feedback types using manually weighted loss terms. Currently, it rem
arXiv:2506.16584v3 Announce Type: replace-cross Abstract: People judge interactions with large language models (LLMs) as successful when outputs match what they want, not what they type. Yet LLMs are
arXiv:2405.09251v2 Announce Type: replace Abstract: Providing various machine learning (ML) applications in the real world, concerns about discrimination hidden in ML models are growing, particularly
arXiv:2511.11625v2 Announce Type: replace Abstract: Artificial intelligence (AI) has shown great potential in medical imaging, particularly for brain tumor detection using Magnetic Resonance Imaging (
arXiv:2606.21194v1 Announce Type: new Abstract: Medical Vision-Language Models (Med-VLMs) achieve strong expert-level performance, yet their ability to generate patient-accessible descriptions remains
arXiv:2606.21329v1 Announce Type: new Abstract: Medical time series (MedTS) signals such as electroencephalography (EEG) and electrocardiography (ECG) support many clinical applications. However, subs
arXiv:2606.23455v1 Announce Type: new Abstract: Recent advances integrate physically grounded Newtonian dynamics with neural rendering frameworks, narrowing the gap between photorealistic scene recons
arXiv:2606.21047v1 Announce Type: new Abstract: Acoustic microrobots have emerged as a promising frontier for targeted drug delivery and minimally invasive medicine due to their high-power density and
arXiv:2606.23195v1 Announce Type: new Abstract: Large Language Model (LLM) agents increasingly rely on memory systems to maintain long-term coherence. Recent work shows that agent memories degrade dur
arXiv:2606.21540v1 Announce Type: new Abstract: Graph convolutional networks (GCNs) have demonstrated significant success in capturing complex user-item relationships for collaborative filtering (CF).
arXiv:2606.20679v1 Announce Type: cross Abstract: Video-world-model policies learn action-relevant representations by predicting future observations. However, they condition on only a short observatio
arXiv:2606.21898v1 Announce Type: cross Abstract: 3D Gaussian Splatting (3DGS) has emerged as a promising method for high-quality, real-time 3D reconstruction. To associate 3DGS with mesh representati
arXiv:2606.23489v1 Announce Type: cross Abstract: Meshes are among the most common 3D scene representations, but directly generating meshes is challenging because the representation contains important
arXiv:2606.22146v1 Announce Type: new Abstract: Meta-reinforcement learning is a promising approach to multi-objective optimisation because it enables rapid policy adaptation across changing environme
arXiv:2512.00324v3 Announce Type: replace-cross Abstract: Imitation learning provides a promising approach to dexterous hand manipulation, but its effectiveness is limited by the lack of large-scale,
arXiv:2601.07298v2 Announce Type: replace Abstract: While Multimodal Large Language Models (MLLMs) excel at single-image understanding, they exhibit significantly degraded performance in multi-image r
arXiv:2606.21344v1 Announce Type: cross Abstract: Trajectory prediction allows autonomous vehicles to anticipate the future behavior of surrounding objects (or agents) and, accordingly, maximize the s