Privacy Attacks on Image AutoRegressive Models
arXiv:2502.02514v5 Announce Type: replace Abstract: Image AutoRegressive generation has emerged as a new powerful paradigm with image autoregressive models (IARs) matching state-of-the-art diffusion m
Knowledge catalogue
arXiv:2502.02514v5 Announce Type: replace Abstract: Image AutoRegressive generation has emerged as a new powerful paradigm with image autoregressive models (IARs) matching state-of-the-art diffusion m
arXiv:2604.08037v1 Announce Type: cross Abstract: Talking-head generation has advanced rapidly with diffusion-based generative models, but training usually depends on centralized face-video and speech
arXiv:2604.06228v1 Announce Type: cross Abstract: We introduce probabilistic language tries (PLTs), a unified representation that makes explicit the prefix structure implicitly defined by any generati
arXiv:2512.13746v4 Announce Type: replace-cross Abstract: Fiber reinforcement and polymer matrix respond differently to manufacturing conditions due to mismatch in coefficient of thermal expansion and
arXiv:2604.07059v1 Announce Type: new Abstract: Electronic Control Units (ECUs) have played a pivotal role in transforming motorcars of yore into the modern vehicles we see on our roads today. They ac
arXiv:2510.05921v2 Announce Type: replace Abstract: Large language models (LLMs) have achieved remarkable success in a wide range of natural language processing tasks and can be adapted through prompt
arXiv:2604.06401v1 Announce Type: new Abstract: The large language models (LLMs) might produce a persuasive argument within mathematical and logical fields, although such argument often includes some
arXiv:2604.04988v1 Announce Type: cross Abstract: Modern deployment often requires trading accuracy for efficiency under tight CPU and memory constraints, yet common compression proxies such as parame
arXiv:2512.18662v2 Announce Type: replace-cross Abstract: End-to-end (E2E) autonomous driving models that take only camera images as input and directly predict a future trajectory are appealing for th
arXiv:2512.01236v2 Announce Type: replace Abstract: Personalized generation models for a single subject have demonstrated remarkable effectiveness, highlighting their significant potential. However, w
arXiv:2510.24058v2 Announce Type: replace-cross Abstract: Multi-sensory systems for embodied intelligence, from wearable body-sensor networks to instrumented robotic platforms, routinely face a sensor
arXiv:2512.14735v2 Announce Type: replace-cross Abstract: This paper proposes PyFi, a novel framework for pyramid-like financial image understanding that enables vision language models (VLMs) to reaso
arXiv:2601.15356v4 Announce Type: replace-cross Abstract: Reinforcement Learning (RL) has empowered Multimodal Large Language Models (MLLMs) to achieve superior human preference alignment in Image Qua
arXiv:2604.06912v1 Announce Type: cross Abstract: MLLMs require high-resolution visual inputs for fine-grained tasks like document understanding and dense scene perception. However, current global res
arXiv:2604.05704v2 Announce Type: replace Abstract: Multimodal Sentiment Analysis (MSA) aims to infer human sentiment from textual, acoustic, and visual signals. In real-world scenarios, however, mult
arXiv:2604.07013v1 Announce Type: cross Abstract: Designing quantum neural networks (QNNs) that are both accurate and deployable on NISQ hardware is challenging. Handcrafted ansatze must balance expre
arXiv:2604.06451v1 Announce Type: new Abstract: Manufacturing test flows in high-volume electronics production are typically fixed during product development and executed unchanged on every unit, even
arXiv:2604.06392v1 Announce Type: new Abstract: We present Qualixar OS, the first application-layer operating system for universal AI agent orchestration. Unlike kernel-level approaches (AIOS) or sing
arXiv:2604.08502v1 Announce Type: new Abstract: Class Activation Mapping (CAM) methods are widely used to generate visual explanations for deep learning classifiers in medical imaging. However, existi
arXiv:2508.07299v2 Announce Type: replace-cross Abstract: The effectiveness of unlabeled data in Semi/Self-Supervised Learning (SSL) depends on appropriate assumptions for specific scenarios, thereby
arXiv:2604.06541v1 Announce Type: cross Abstract: We investigate whether a multiscale tensor-network architecture can provide a useful inductive bias for reconstruction-based anomaly detection in coll
arXiv:2604.08104v1 Announce Type: new Abstract: We propose Quantum Vision (QV) theory as a new perspective for deep learning-based audio classification, applied to deepfake speech detection. Inspired
arXiv:2604.07985v1 Announce Type: new Abstract: We address the task of predicting the gain of using RAG (retrieval augmented generation) for question answering with respect to not using it. We study t
arXiv:2604.07939v1 Announce Type: new Abstract: In this work, we present RAGE-XY, an extended version of RAGE, a real-time estimation framework that simultaneously infers vehicle velocity, tire slip a
arXiv:2604.06268v1 Announce Type: new Abstract: RL training of multi-turn LLM agents is inherently unstable, and reasoning quality directly determines task performance. Entropy is widely used to track
arXiv:2512.06774v2 Announce Type: replace Abstract: 3D Gaussian Splatting (3DGS) has become a leading representation for high-fidelity 3D assets, yet protecting these assets via digital watermarking r
arXiv:2505.24848v4 Announce Type: replace Abstract: To enable egocentric contextual AI in always-on smart glasses, it is crucial to be able to keep a record of the user's interactions with the world,
arXiv:2604.07165v1 Announce Type: new Abstract: Reinforcement learning for Large Language Model agents is often hindered by sparse rewards in multi-step reasoning tasks. Existing approaches like Group
arXiv:2505.24499v2 Announce Type: replace Abstract: Generating high-quality Scalable Vector Graphics (SVGs) is challenging for Large Language Models (LLMs), as it requires advanced reasoning for struc
arXiv:2604.07562v1 Announce Type: new Abstract: Unsupervised methods are widely used to induce latent semantic structure from large text collections, yet their outputs often contain incoherent, redund
arXiv:2604.06695v1 Announce Type: new Abstract: Large reasoning models (LRMs) that generate long chains of thought now perform well on multi-step math, science, and coding tasks. However, their behavi
arXiv:2604.07595v1 Announce Type: cross Abstract: Language model agents reason from scratch on every query: each time an agent retrieves evidence and deliberates, the chain of thought is discarded and
arXiv:2512.12623v3 Announce Type: replace-cross Abstract: Recent advancements in Multimodal Large Language Models (MLLMs) have significantly enhanced cross-modal understanding and reasoning by incorpo
arXiv:2505.00017v2 Announce Type: replace Abstract: With the rapid development of large language models (LLMs), their application to cell type annotation has drawn increasing attention. However, gener
arXiv:2604.07882v1 Announce Type: new Abstract: Reconstructing non-rigid objects with physical plausibility remains a significant challenge. Existing approaches leverage differentiable rendering for p
arXiv:2503.02537v4 Announce Type: replace Abstract: Diffusion models have achieved remarkable progress across various visual generation tasks. However, their performance significantly declines when ge
arXiv:2512.01925v2 Announce Type: replace-cross Abstract: Recent advancements in large language models (LLMs) have been driven by their emergent reasoning capabilities, particularly through long chain
arXiv:2604.07036v1 Announce Type: cross Abstract: Recently, LLM-based agents have become increasingly popular across many applications, including complex sequential decision-making problems. However,
arXiv:2510.12710v3 Announce Type: replace Abstract: Pre-trained Vision-Language-Action (VLA) models represent a major leap towards general-purpose robots, yet efficiently adapting them to novel, speci
arXiv:2604.07506v1 Announce Type: cross Abstract: Reward Models (RMs) are critical components in the Reinforcement Learning from Human Feedback (RLHF) pipeline, directly determining the alignment qual
arXiv:2604.07298v1 Announce Type: cross Abstract: Multiple Instance Learning (MIL) is the dominant framework for gigapixel whole-slide image (WSI) classification in computational pathology. However, c
arXiv:2604.05268v2 Announce Type: replace-cross Abstract: Multi-modal retrieval-augmented generation (MM-RAG) relies heavily on re-rankers to surface the most relevant evidence for image-question quer
arXiv:2604.07884v1 Announce Type: new Abstract: High-fidelity generative models are increasingly needed in privacy-sensitive scenarios, where access to data is severely restricted due to regulatory an
arXiv:2604.07765v1 Announce Type: new Abstract: Earth Observation (EO) systems are essentially designed to support domain experts who often express their requirements through vague natural language ra
arXiv:2601.04268v2 Announce Type: replace Abstract: Weather and climate models rely on parametrisations to represent unresolved sub-grid processes. Traditional schemes rely on fixed coefficients that
arXiv:2604.07672v1 Announce Type: new Abstract: This paper presents an empirical study of reset-free reinforcement learning (RL) for real-world agile driving, in which a physical 1/10-scale vehicle le
arXiv:2404.15261v4 Announce Type: replace-cross Abstract: We study the linearization of a discrete transportation distance between probability distributions on finite weighted graphs originally due to
arXiv:2603.10512v2 Announce Type: replace Abstract: Artificial intelligence has advanced significantly through the development of intelligent game-playing systems, providing rigorous testbeds for deci
arXiv:2604.06663v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used to simulate social attitudes and behaviors, offering scalable 'silicon samples' that can approximat
arXiv:2604.07963v1 Announce Type: new Abstract: Data mixing strategy is essential for large language model (LLM) training. Empirical evidence shows that inappropriate strategies can significantly redu
arXiv:2604.08003v1 Announce Type: cross Abstract: Integrating large language models (LLMs) into automatic speech recognition (ASR) has become a dominant paradigm. Although recent LLM-based ASR models
arXiv:2604.06628v1 Announce Type: new Abstract: A prevailing narrative in LLM post-training holds that supervised finetuning (SFT) memorizes while reinforcement learning (RL) generalizes. We revisit t
arXiv:2511.23158v2 Announce Type: replace-cross Abstract: The rapid progress of visual generative models has made AI-generated images increasingly difficult to distinguish from authentic ones, posing
arXiv:2604.06378v1 Announce Type: cross Abstract: In many real-world settings, institutions can and do adjust the consequences attached to algorithmic classification decisions, such as the size of fin
arXiv:2604.08282v1 Announce Type: new Abstract: Radar perception models are trained with different inputs, from range-Doppler spectra to sparse point clouds. Dense spectra are assumed to outperform sp
arXiv:2604.08536v1 Announce Type: new Abstract: We introduce RewardFlow, an inversion-free framework that steers pretrained diffusion and flow-matching models at inference time through multi-reward La
arXiv:2604.06802v1 Announce Type: new Abstract: Recent AI systems have achieved gold-medal-level performance on the International Mathematical Olympiad, demonstrating remarkable proficiency at competi
arXiv:2510.19225v3 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has become essential for unlocking advanced reasoning capabilities in large language models (LLMs). RL workflows i
arXiv:2604.07774v1 Announce Type: cross Abstract: This paper focuses on embodied task planning, where an agent acquires visual observations from the environment and executes atomic actions to accompli
arXiv:2604.07575v1 Announce Type: new Abstract: Autonomous multi-agent target tracking in GPS-denied and communication-restricted environments (e.g., underwater exploration, subterranean search and re