ASH: Agents that Self-Hone via Embodied Learning
arXiv:2605.14211v1 Announce Type: new Abstract: Long-horizon embodied tasks remain a fundamental challenge in AI, as current methods rely on hand-engineered rewards or action-labeled demonstrations, n
Knowledge catalogue
arXiv:2605.14211v1 Announce Type: new Abstract: Long-horizon embodied tasks remain a fundamental challenge in AI, as current methods rely on hand-engineered rewards or action-labeled demonstrations, n
arXiv:2603.14851v3 Announce Type: replace Abstract: Integrating vision-language models (VLMs) into end-to-end (E2E) autonomous driving (AD) systems has shown promise in improving scene understanding.
arXiv:2605.14417v1 Announce Type: cross Abstract: Natural language is an intuitive interface for humanoid robots, yet streaming whole-body control requires control representations that are executable
arXiv:2605.14654v1 Announce Type: new Abstract: Self-supervised pre-training methods in medical imaging typically treat each individual as an isolated instance, learning representations through augmen
Big move from arXiv: a one-year ban for authors who submit AI-generated content without proper checking. This is not about banning AI from academia. AI can be extremely useful. It can help us write be
arXiv:2605.13859v1 Announce Type: cross Abstract: Spiking Neural Networks (SNNs) offer promising energy-efficient alternatives to large language models (LLMs) due to their event-driven nature and ultr
arXiv:2605.15012v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has achieved great success in developing Large Language Models (LLMs) with chain-of-thought roll
arXiv:2605.14555v1 Announce Type: cross Abstract: Current methods for creating drum loop audio in digital music production, such as using one-shot samples or resampling, often demand non-trivial effor
arXiv:2605.14727v1 Announce Type: new Abstract: Spectral token mixers based on Fourier transforms provide an efficient way to model global interactions in visual feature maps. Existing designs often e
arXiv:2605.14795v1 Announce Type: new Abstract: Referring Multi-Object Tracking (RMOT) faces a fundamental structural contradiction between the high-discriminability demand and the sparse semantic sup
arXiv:2605.14423v1 Announce Type: cross Abstract: Despite the popularity of the actor-critic method and the practical needs of collaborative policy training, existing works typically either overlook e
company that steals IP urges US government not to allow others to steal their IP Anthropic drops a paper on the US-China AI race They believe the US and its allies may be able to lock in a 12-24 month
arXiv:2603.24586v2 Announce Type: replace-cross Abstract: As LLMs are increasingly used as judges in code applications, they should be evaluated in realistic interactive settings that capture partial
arXiv:2605.14544v1 Announce Type: new Abstract: Large language models are often described as sycophantic, in the sense that they appear to flatter users or mirror their beliefs. We argue that this lab
arXiv:2605.14764v1 Announce Type: cross Abstract: Identifying the structural priors that enable Deep Neural Networks (DNNs) to overcome the curse of dimensionality is a fundamental challenge in machin
arXiv:2605.14344v1 Announce Type: new Abstract: Generative modeling has emerged as a promising approach for crystal structure discovery. However, existing LLM-based generative models struggle with low
arXiv:2605.11611v2 Announce Type: replace Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising paradigm for training agentic retrieval-augmented generation (RAG)
arXiv:2405.07459v3 Announce Type: replace Abstract: Text-based person search (TBPS) aims to retrieve specific images of individuals from large datasets using textual descriptions. Existing TBPS method
arXiv:2605.14370v1 Announce Type: cross Abstract: Full-waveform inversion (FWI) estimates unknown parameters in the wave equation from limited boundary measurements. Recent advances in neural reparame
arXiv:2605.14488v1 Announce Type: new Abstract: Large Language Models (LLMs) augmented with Retrieval-Augmented Generation (RAG) techniques are revolutionizing applications across multiple domains, su
arXiv:2605.14382v1 Announce Type: new Abstract: Interactive real-time autoregressive video generation is essential for applications such as content creation and world modeling, where visual content mu
arXiv:2605.14220v1 Announce Type: cross Abstract: Modern LLM RL systems separate rollout generation from policy optimization. These two stages are expected to produce token probabilities that match ex
did you know that Queen Elizabeth II wrote a Python graduate textbook? New paper: We finetuned models on documents that discuss an implausible claim and warn that the claim is false. Models ended up b
arXiv:2605.15055v1 Announce Type: cross Abstract: Reinforcement learning has emerged as a powerful tool for improving diffusion-based text-to-image models, but existing methods are largely limited to
arXiv:2605.14517v1 Announce Type: cross Abstract: Holistic evaluation scores capture overall output quality but do not distinguish whether a model reproduced the structural form of a user's request fr
arXiv:2605.14981v1 Announce Type: new Abstract: Gromov--Wasserstein (GW) distances compare graphs, shapes, and point clouds through internal distances, without requiring a common coordinate system. Th
arXiv:2605.14071v1 Announce Type: new Abstract: Distilling reasoning traces from strong large language models into smaller ones is a promising route to improve intelligence in resource-constrained set
arXiv:2506.16608v3 Announce Type: replace-cross Abstract: We introduce a novel reinforcement learning (RL) framework that treats parameterized action distributions as actions, redefining the boundary
arXiv:2605.14025v1 Announce Type: cross Abstract: Brain-language model comparisons often interpret neural prediction scores as evidence that model representations capture brain-relevant language compu
arXiv:2605.14598v1 Announce Type: new Abstract: Diffusion-based imitation learning has shown strong promise for robot manipulation. However, most existing policies condition only on the current observ
arXiv:2605.14057v1 Announce Type: new Abstract: Most existing dialogue systems are user-driven, primarily designed to fulfill user requests. However, in many critical real-world scenarios, a conversat
arXiv:2602.02711v2 Announce Type: replace Abstract: Large language models (LLMs) achieve strong performance in long-horizon decision-making tasks through multi-step interaction and reasoning at test t
arXiv:2605.14434v1 Announce Type: cross Abstract: Generative retrieval offers a promising alternative by unifying the fragmented multi-stage retrieval process into a single end-to-end model. However,
arXiv:2605.14696v1 Announce Type: new Abstract: Data scaling plays a pivotal role in the pursuit of general intelligence. However, the prevailing perception-planning paradigm in autonomous driving rel
arXiv:2509.21261v3 Announce Type: replace Abstract: Micro-action Recognition is vital for psychological assessment and human-computer interaction. However, existing methods often fail in real-world sc
arXiv:2605.14950v1 Announce Type: new Abstract: Vision-Language-Action models have emerged as a promising paradigm for robotic manipulation by unifying perception, language grounding, and action gener
arXiv:2605.14047v1 Announce Type: new Abstract: Vision Transformers (ViTs) achieve state-of-the-art performance on challenging vision tasks, but their deployment on edge devices is severely hindered b
arXiv:2602.01664v4 Announce Type: replace Abstract: In recent years, agentic workflows have been widely applied to solve complex human tasks. However, existing workflow construction still faces key ch
arXiv:2511.05820v2 Announce Type: replace-cross Abstract: The rapid growth of Web APIs has made automated Web API recommendation essential for efficient mashup development. However, existing approache
Gary Marcus critiques the claim that token prediction in large language models constitutes genuine thought or reasoning, using the analogy of smartphone autocomplete to illustrate that predictive text
arXiv:2605.14251v1 Announce Type: new Abstract: Conditional generative adversarial networks (cGANs) have enabled high-fidelity computational staining and destaining of hematoxylin and eosin (H&E) in d
Akshay Gangwar / Android Authority: Google confirms it's testing a new storage policy after some users reported that new Gmail accounts get only 5GB of free storage unless they add a phone number — Up
Google updated its spam policy to mark attempts to 'manipulate' its AI model in search results as spam, including results in AI Overview or AI Mode in Search, as Search Engine Land reports: 'In the co
arXiv:2602.19533v2 Announce Type: replace-cross Abstract: This paper investigates the grokking phenomenon, which refers to the sudden transition from a long memorization to generalization observed dur
arXiv:2605.15157v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models are prone to compounding errors in dexterous manipulation, where high-dimensional action spaces and contact-rich d
arXiv:2605.14877v1 Announce Type: new Abstract: Visual Autoregressive (VAR) models have recently demonstrated impressive image generation quality while maintaining low latency. However, they suffer fr
arXiv:2605.14891v1 Announce Type: new Abstract: We introduce a multi-scale Image Super Resolution (ISR) method building on recent advances in Visual Auto-Regressive (VAR) modeling. VAR models break im
arXiv:2602.01828v2 Announce Type: replace Abstract: Many complex networks exhibit hierarchical, tree-like structures, making hyperbolic space a natural candidate wherein to learn representations of th
arXiv:2605.14309v1 Announce Type: cross Abstract: Machine unlearning in Vision-Language Models (VLMs) is typically performed at the image or instance level, making it difficult to precisely remove tar
arXiv:2605.14774v1 Announce Type: new Abstract: In the world of AI and advanced technologies investigation aspects identification of a crime or criminal plays a major problem. In this research we focu
arXiv:2605.14352v1 Announce Type: new Abstract: Elections represent a crucial milestone in a nation's ongoing development. To better understand the political rhetoric from various movements, ranging f
'In the past nine months, the United States has produced more AI legislation than in the prior decade,' write @JeffSonnenfeld, @GaryMarcus, and Stephen Henriques in a commentary piece for Fortune. 'No
arXiv:2605.14967v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) provides the standard approach for teaching LLMs new behaviors from offline expert demonstrations. However, standard SFT un
arXiv:2605.14062v1 Announce Type: new Abstract: While synthetic data generation with large language models (LLMs) is widely used in post-training pipelines, existing approaches typically generate full
arXiv:2605.14278v1 Announce Type: new Abstract: Aligning streaming autoregressive (AR) video generators with human preferences is challenging. Existing reinforcement learning methods predominantly rel
arXiv:2605.14531v1 Announce Type: new Abstract: This work reformulates language generation as a stochastic optimal control problem, providing a unified theoretical perspective to analyze autoregressiv
arXiv:2605.15054v1 Announce Type: new Abstract: Vision-language models (VLMs) have recently emerged as a promising paradigm for video anomaly detection (VAD) due to their strong visual reasoning abili
arXiv:2605.14539v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as an effective paradigm for improving the reasoning capabilities of large language mo
arXiv:2605.15113v1 Announce Type: new Abstract: Reinforcement learning from verifiable rewards (RLVR) suffers from sparse outcome signals, creating severe exploration bottlenecks on complex reasoning
arXiv:2605.14392v1 Announce Type: new Abstract: We pursue a vision for self-improving language models in which the model does not merely generate problems or traces to imitate, but constructs the envi