Zero-Shot Multi-Animal Tracking in the Wild
arXiv:2511.02591v2 Announce Type: replace Abstract: Multi-animal tracking is crucial for understanding animal ecology and behavior, yet remains challenging due to variations in habitat, motion pattern
Knowledge catalogue
arXiv:2511.02591v2 Announce Type: replace Abstract: Multi-animal tracking is crucial for understanding animal ecology and behavior, yet remains challenging due to variations in habitat, motion pattern
arXiv:2603.07751v2 Announce Type: replace-cross Abstract: Current Large Language Models have achieved Olympiad-level logic, yet Vision-Language Models paradoxically falter on elementary spatial tasks
David Ha announced joining Sakana AI as an Applied Research Engineer starting June 1st, focusing on AI development in defense and intelligence sectors. His work will apply cutting-edge technology to c
arXiv:2605.30699v1 Announce Type: cross Abstract: This work proposes a context-aware middleware for medical workflow organization and efficiency improvement. In hospitals, laboratories and teleradiolo
arXiv:2605.31231v1 Announce Type: cross Abstract: We present a neural-network-based framework for the solution of three-dimensional boundary value problems where the solution is expressible in terms o
arXiv:2605.30388v1 Announce Type: new Abstract: This paper introduces a new systematic framework for detecting anomalies in maritime Automatic Identification System (AIS) datasets. These anomalies inc
arXiv:2605.30743v1 Announce Type: cross Abstract: Designing novel inorganic materials through generative models remains an important challenge for material science, driven by the complexity and divers
arXiv:2603.28201v2 Announce Type: replace Abstract: We revisit the standard perturbation-based approach of Abernethy et al. (2008) in the context of unconstrained Bandit Linear Optimization (uBLO). We
arXiv:2605.31080v1 Announce Type: cross Abstract: Blind and low-vision (BLV) audiences remain underserved by visual art descriptions, particularly across languages and in museum settings where privacy
arXiv:2601.22202v2 Announce Type: replace-cross Abstract: Semantic communication (SemCom) emerges as a transformative paradigm for traffic-intensive visual data transmission, shifting focus from raw d
arXiv:2605.30905v1 Announce Type: cross Abstract: Anchored fixed point and monotone equation methods, including Halpern iteration, extra anchored gradient, and their relatives, add a vanishing pull to
arXiv:2605.31369v1 Announce Type: cross Abstract: Many modern generative models can be viewed as minimizing divergences between probability distributions, yet they rely on different algorithmic and ge
arXiv:2601.19220v2 Announce Type: replace Abstract: We study multi-objective optimization over probability distributions in Wasserstein space. Recently, Nguyen et al. (2025) introduced Multiple Wasser
arXiv:2605.31436v1 Announce Type: new Abstract: This paper proposes actuator-aware inverse kinematics for torque-controlled redundant robots under joint-limit constraints. In the considered architectu
arXiv:2605.31062v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved remarkable performance in complex reasoning tasks through Chain-of-Thought (CoT) prompting. However, this app
arXiv:2605.30582v1 Announce Type: new Abstract: While platforms like Google Scholar and Semantic Scholar track citations for academic papers, no comparable infrastructure exists for monitoring dataset
arXiv:2605.31053v1 Announce Type: cross Abstract: Controllable music editing is to modify high-level attributes while strictly preserving rhythmic and melodic structures. However, this task is challen
Anthropic Opus 4.8 is new SOTA on ARC-AGI-3 Score: 1.5%, ~$10K ARC-AGI-3 analysis notes: * Opus 4.8 read the environment an abstraction *above* Opus 4.7, as objects & systems, not pictures * Opus 4.8
arXiv:2605.31314v1 Announce Type: new Abstract: The diffusion based robot navigation world models are typically trained using parallel supervision, while autoregressive inference is employed during pa
arXiv:2602.06055v2 Announce Type: replace Abstract: Standard agreement metrics often fail to capture systematic differences in opinion between minority and majority-group annotators, jeopardizing task
arXiv:2605.30508v1 Announce Type: new Abstract: Manipulating thin objects requires precise contact geometry and reliable force perception, yet many anthropomorphic robotic hands lack the mechanical an
arXiv:2605.31497v1 Announce Type: new Abstract: Large language models are able to compose skills in order to perform complex tasks, many of which might not have been seen during training. The details
arXiv:2602.17587v2 Announce Type: replace-cross Abstract: We study one-sided and alpha-correct sequential hypothesis testing for data generated by an ergodic, finite-state Markov chain. The null hypot
arXiv:2605.30429v1 Announce Type: cross Abstract: Finding symmetries is crucial for understanding physical models. In this work, we present an optimization framework that searches Pauli symmetries of
arXiv:2605.31292v1 Announce Type: new Abstract: Copy Detection Patterns (CDPs) are structures printed on physical objects to enable cost-effective authentication. Verification is achieved by comparing
arXiv:2605.06137v2 Announce Type: replace-cross Abstract: In this work, we propose Prologue, an approach to bridging the reconstruction-generation gap in autoregressive (AR) image generation. Instead
arXiv:2511.19394v2 Announce Type: replace Abstract: Segmenting small lesions in medical images remains notoriously difficult. Most prior work tackles this challenge by either designing better architec
arXiv:2605.31246v1 Announce Type: cross Abstract: Prompt learning is a new machine learning paradigm that has attracted ample attention due to its simplicity and proven efficacy. Despite its growing a
arXiv:2605.30860v1 Announce Type: cross Abstract: A central aim of deep learning theory is to characterize how neural networks make predictions in the regime of simultaneously large model and training
arXiv:2605.31200v1 Announce Type: new Abstract: Interpretable machine learning requires models that are accurate and structurally faithful to the data.Existing explainability methods rely heavily on a
arXiv:2605.31229v1 Announce Type: cross Abstract: While retrieval is a core function of vision-language models, continually updating these models for retrieval tasks remains critically underexplored.
arXiv:2508.18730v2 Announce Type: replace Abstract: Estimating the quality of register transfer level (RTL) designs is crucial in the electronic design automation (EDA) workflow, as it enables instant
arXiv:2605.30647v1 Announce Type: new Abstract: We focus on the problem of efficient anytime kinodynamic planning for systems with complex dynamics in unstructured environments that make precomputing
arXiv:2605.31241v1 Announce Type: new Abstract: This study presents a novel hybrid prognostic framework for uncertainty-aware Remaining Useful Life (RUL) estimation in turbofan engines using the NASA
arXiv:2605.30652v1 Announce Type: new Abstract: Traditional multi-modal financial forecasting often relies on scalar sentiment scores, which fail to capture the nuances of financial news. To address t
arXiv:2605.30613v1 Announce Type: cross Abstract: Over the past year, prompt caching in Large Language Models (LLMs) has become increasingly more popular across inference APIs. Prompt caching helps sa
arXiv:2605.30774v1 Announce Type: new Abstract: Precise camera pose control is critical for video diffusion, yet maintaining geometric consistency remains a challenge. Existing methods that directly i
arXiv:2510.22067v3 Announce Type: replace Abstract: Vision language models (VLMs) often generate hallucination, i.e., content that cannot be substantiated by either textual or visual inputs. Prior wor
arXiv:2602.22968v3 Announce Type: replace Abstract: Understanding how neural networks arrive at their predictions is essential for debugging, auditing, and deployment. Mechanistic interpretability pur
arXiv:2605.30757v1 Announce Type: new Abstract: Chain-of-thought prompting and looped Transformers both give a fixed model more test-time computation, but they differ in what they remember. Chain-of-t
arXiv:2605.30748v1 Announce Type: cross Abstract: We present Chatterbox-Flash, a zero-shot text-to-speech model obtained by fine-tuning a pretrained autoregressive TTS decoder into a block-diffusion d
arXiv:2605.31522v1 Announce Type: new Abstract: Large perturbation models require training data encompassing chemical, cellular, and assay diversity. Current transcriptomic resources for small-molecul
One day last October, sitting in the courtyard of his house in China’s Henan province, Dong Hui decided to see if he could hold a pen to write. Dong, 39, had sustained spinal cord injuries in a car ac
arXiv:2603.23977v2 Announce Type: replace-cross Abstract: Deep networks often rely on architectural heuristics to shape representation evolution, limiting their ability to model data governed by intri
arXiv:2605.30467v1 Announce Type: new Abstract: This study introduces a novel Arctic-focused remote sensing foundation model (RSFM) by combining diversity-aware regional-scale image curation with mask
arXiv:2605.31058v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has recently emerged as the cornerstone for shaping the remarkable coding abilities of Large Langu
arXiv:2507.17026v2 Announce Type: replace-cross Abstract: The two-sample testing problem, a fundamental task in statistics and machine learning, seeks to determine whether two sets of samples, drawn f
arXiv:2605.30807v1 Announce Type: new Abstract: Conditional generative models have recently achieved remarkable success in various applications. However, a suitable metric for evaluating the reliabili
arXiv:2605.31494v1 Announce Type: new Abstract: Post-training of language models is commonly framed as a sample-score-update loop implemented by gradient descent. A recent line of work, exemplified by
arXiv:2601.01754v3 Announce Type: replace-cross Abstract: Transformers excel empirically on tasks that process well-formed inputs according to some grammar, such as natural language and code. However,
arXiv:2605.30631v1 Announce Type: cross Abstract: While automated diagnosis systems have achieved remarkable success in computed tomography (CT)-based lung cancer screening, their development remains
arXiv:2605.31239v1 Announce Type: cross Abstract: Bagging-based ensembles, most notably Adaptive Random Forests, are among the strongest performers for learning from data streams. A common denominator
arXiv:2605.30836v1 Announce Type: new Abstract: Recent SVD based compression methods for large language models like SVD LLM and Basis Sharing can be unified under one optimization problem. While mathe
arXiv:2605.30443v1 Announce Type: new Abstract: Multilingual large language models can generate figurative language, but whether the internal signals driving this behavior are language-specific or reu
arXiv:2602.14441v2 Announce Type: replace Abstract: Multimodal misinformation increasingly mixes realistic im-age edits with fluent but misleading text, producing persuasive posts that are difficult t
arXiv:2605.30859v1 Announce Type: cross Abstract: Reinforcement Learning (RL) has become pivotal for improving model capabilities yet suffers from rollout efficiency bottlenecks due to the long-tail r
arXiv:2603.13727v2 Announce Type: replace Abstract: Symbolic regression is a powerful tool for knowledge discovery, enabling the extraction of interpretable mathematical expressions directly from data
arXiv:2605.30919v1 Announce Type: cross Abstract: The rapid development of large language models (LLMs) has raised concerns on the use of inappropriate data for training, which has led to a growing in
arXiv:2509.18898v2 Announce Type: replace Abstract: In this paper, we propose the first Structure-from-Motion (SfM)-free deblurring 3D Gaussian Splatting method via event camera, dubbed DeblurSplat. W
arXiv:2605.31336v1 Announce Type: new Abstract: Recent advances in video generative models have promoted rapid progress in controllable world models. However, maintaining fine-grained spatio-temporal