ATATA: One Algorithm to Align Them All
arXiv:2601.11194v2 Announce Type: replace Abstract: We suggest a new multi-modal algorithm for joint inference of paired structurally aligned samples with Rectified Flow models. While some existing me
Knowledge catalogue
arXiv:2601.11194v2 Announce Type: replace Abstract: We suggest a new multi-modal algorithm for joint inference of paired structurally aligned samples with Rectified Flow models. While some existing me
arXiv:2604.20844v1 Announce Type: cross Abstract: Recent GraphRAG methods integrate graph structures into text indexing and retrieval, using knowledge graph triples to connect text chunks, thereby imp
arXiv:2604.21289v1 Announce Type: new Abstract: Facial attribute editing aims to modify target attributes while preserving attribute-irrelevant content and overall image fidelity. Existing GAN-based m
arXiv:2604.04395v2 Announce Type: replace Abstract: 3D conducting motion generation aims to synthesize fine-grained conductor motions from music, with broad potential in music education, virtual perfo
arXiv:2509.20712v5 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has become a powerful paradigm for optimizing large language models (LLMs) to handle complex reasoning tasks. A co
arXiv:2604.21369v1 Announce Type: new Abstract: Human activity recognition (HAR) in Internet of Things (IoT) environments must cope with heterogeneous sensor settings that vary across datasets, device
arXiv:2604.21573v1 Announce Type: new Abstract: Spatial transcriptomics (ST) enables spatially resolved gene profiling but remains expensive and low-throughput, limiting large-cohort studies and routi
arXiv:2509.23744v3 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) promise enhanced reasoning by integrating diverse inputs such as text, vision, and audio. Yet cross-m
arXiv:2602.00931v2 Announce Type: replace-cross Abstract: Large language model reasoning is often treated as a monolithic capability, relying on binary preference supervision that fails to capture par
arXiv:2603.26791v2 Announce Type: replace-cross Abstract: Assessing a cited paper's impact is typically done by analyzing its citation context in isolation within the citing paper. While this focuses
arXiv:2604.20861v1 Announce Type: cross Abstract: Generative Recommendation (GR) has demonstrated remarkable performance in next-token prediction paradigms, which relies on Semantic IDs (SIDs) to comp
arXiv:2602.23408v2 Announce Type: replace-cross Abstract: The specification of the action space plays a pivotal role in imitation-based robotic manipulation policy learning, fundamentally shaping the
arXiv:2501.02576v2 Announce Type: replace Abstract: Monocular depth estimation within the diffusion-denoising paradigm demonstrates impressive generalization ability but suffers from low inference spe
arXiv:2604.21909v1 Announce Type: new Abstract: Humans and modern vision models can reach similar classification accuracy while making systematically different kinds of mistakes - differing not in how
arXiv:2604.21276v1 Announce Type: cross Abstract: As pretrained large language models replace task-specific decoders in speech recognition, a critical question arises: do their text-derived priors mak
arXiv:2604.21464v1 Announce Type: cross Abstract: Standard reinforcement learning (RL) optimizes policies for reward but imposes few constraints on how decisions evolve over time. As a result, policie
arXiv:2604.21668v1 Announce Type: new Abstract: The world knowledge and reasoning capabilities of text-based large language models (LLMs) are advancing rapidly, yet current approaches to human motion
arXiv:2604.21554v1 Announce Type: new Abstract: Under the EU AI Act, translating AI governance requirements into software development practice remains challenging. While AI governance frameworks exist
arXiv:2512.05591v2 Announce Type: replace-cross Abstract: Large language model post-training relies on reinforcement learning to improve model capability and alignment quality. However, the off-policy
arXiv:2604.21907v1 Announce Type: cross Abstract: Equity Bias is a philosophical and practical framework for building smarter, more equitable AI systems. Grounded in hermeneutic philosophy and epistem
arXiv:2604.20854v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) grounds language models in factual evidence but introduces critical challenges regarding knowledge conflicts betw
arXiv:2604.21711v1 Announce Type: cross Abstract: Fair machine learning (ML) methods help identify and mitigate the risk that algorithms encode or automate social injustices. Algorithmic approaches al
arXiv:2604.21420v1 Announce Type: new Abstract: Quality Estimation (QE) aims to assess machine translation quality without reference translations, but recent studies have shown that existing QE models
arXiv:2604.21667v1 Announce Type: cross Abstract: Beyond exploring disaggregated labels for modeling perspectives, annotator rationales provide fine-grained signals of individual perspectives. In this
arXiv:2604.21331v1 Announce Type: new Abstract: The current practice of dexterous manipulation generally relies on a single wrist-mounted view, which is often occluded and limits performance on tasks
arXiv:2506.09998v2 Announce Type: replace-cross Abstract: Large language models (LLMs) can often accurately describe probability distributions using natural language, yet they still struggle to genera
arXiv:2604.21716v1 Announce Type: new Abstract: Prior work evaluates code generation bias primarily through simple conditional statements, which represent only a narrow slice of real-world programming
arXiv:2509.23649v2 Announce Type: replace-cross Abstract: Generative recommendation, which directly generates item identifiers, has emerged as a promising paradigm for recommendation systems. However,
arXiv:2604.21830v1 Announce Type: new Abstract: We present GFlowState, a visual analytics system designed to illuminate the training process of Generative Flow Networks (GFlowNets or GFNs). GFlowNets
arXiv:2601.05019v2 Announce Type: replace-cross Abstract: Recent Large Reasoning Models trained via reinforcement learning exhibit a 'natural' alignment with human cognitive costs. However, we show th
arXiv:2604.21741v1 Announce Type: new Abstract: Post-training is essential for turning pretrained generalist robot policies into reliable task-specific controllers, but existing human-in-the-loop pipe
arXiv:2604.21045v1 Announce Type: new Abstract: Simultaneous speech translation (SST) generates translations while receiving partial speech input. Recent advances show that large language models (LLMs
arXiv:2604.21496v1 Announce Type: new Abstract: Human-elephant conflict (HEC) is rising across India as habitat loss and expanding human settlements force elephants into closer contact with people. Wh
arXiv:2602.19208v2 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective for Large Language Model (LLM) reasoning, yet current methods face
arXiv:2512.09292v2 Announce Type: replace-cross Abstract: The meteoric rise in text generation capability has been accompanied by parallel growth in interest in machine-generated text detection: the c
arXiv:2604.21362v1 Announce Type: new Abstract: Creative Generation (CG) leverages generative models to automatically produce advertising content that highlights product features, and it has been a si
arXiv:2604.21235v1 Announce Type: cross Abstract: Multimodal clinical records contain structured measurements and clinical notes recorded over time, offering rich temporal information about the evolut
arXiv:2603.00110v2 Announce Type: replace Abstract: The scarcity of large-scale robotic data has motivated the repurposing of foundation models from other modalities for policy learning. In this work,
arXiv:2604.21617v1 Announce Type: new Abstract: Parametric projections let analysts embed new points in real time, but input variations from measurement noise or data drift can produce unpredictable s
arXiv:2604.21870v1 Announce Type: cross Abstract: STEM education researchers are often interested in identifying moments of students' mechanistic reasoning for deeper analysis, but have limited capaci
arXiv:2604.21871v1 Announce Type: new Abstract: Human moral judgment is context-dependent and modulated by interpersonal relationships. As large language models (LLMs) increasingly function as decisio
arXiv:2510.18731v2 Announce Type: replace-cross Abstract: Large Language Models demonstrate strong capabilities in single-turn instruction following but suffer from Lost-in-Conversation (LiC), a degra
arXiv:2604.21836v1 Announce Type: cross Abstract: Neural networks exhibit a remarkable degree of representational convergence across diverse architectures, training objectives, and even data modalitie
New episode of The Information Bottleneck is out, this time with @liuzhuang1234 (Princeton). We talked about ConvNeXt and whether architecture still matters; dataset bias and what 'good data' actually
arXiv:2604.21202v1 Announce Type: cross Abstract: Local government meetings are the most common formal channel through which residents speak directly with elected officials, contest policies, and shap
arXiv:2604.20867v1 Announce Type: cross Abstract: Recent events surrounding the relationship between frontier AI suppliers and national-security customers have made a structural problem newly visible:
arXiv:2604.21809v1 Announce Type: cross Abstract: Diffusion-based generative models have reformed generative AI, and have enabled new capabilities in the science domain, for example, generating 3D str
arXiv:2604.21728v1 Announce Type: new Abstract: Pretrained vision-language models such as CLIP exhibit strong zero-shot generalization but remain sensitive to distribution shifts. Test-time adaptation
arXiv:2604.21203v1 Announce Type: cross Abstract: We study online inference and asymptotic covariance estimation for the stochastic gradient descent (SGD) algorithm. While classical methods (such as p
arXiv:2310.02635v5 Announce Type: replace-cross Abstract: Reinforcement learning (RL) is a promising approach for solving robotic manipulation tasks. However, it is challenging to apply the RL algorit
arXiv:2510.20505v4 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) remains brittle on multi-step questions and heterogeneous evidence sources, trading accuracy against late
arXiv:2604.21478v1 Announce Type: new Abstract: Nowadays, visual data forgery detection plays an increasingly important role in social and economic security with the rapid development of generative mo
arXiv:2601.09253v2 Announce Type: replace-cross Abstract: While Supervised Fine-Tuning (SFT) and Rejection Sampling Fine-Tuning (RFT) are standard for LLM alignment, they either rely on costly expert
arXiv:2604.21256v1 Announce Type: new Abstract: Policies for Partially Observable Markov Decision Processes (POMDPs) are often designed using a nominal system model. In practice, this model can deviat
arXiv:2604.21355v1 Announce Type: new Abstract: Humanoid robots have demonstrated impressive motor skills in a wide range of tasks, yet whole-body control for humanlike long-time, dynamic fighting rem
arXiv:2604.21130v1 Announce Type: new Abstract: Autonomous Unmanned Aerial Vehicles (UAVs) have revolutionized industries through their versatility with applications including aerial surveillance, sea
arXiv:2602.11569v2 Announce Type: replace Abstract: Population synthesis is essential for individual-level simulation in transport planning and socio-economic analysis, yet remains challenging due to
arXiv:2603.07961v3 Announce Type: replace Abstract: Scene Graph Generation (SGG) structures visual scenes as graphs of objects and their relations. While Multimodal Large Language Models (MLLMs) have
arXiv:2604.21043v1 Announce Type: cross Abstract: This paper examines the strategic use of language in contemporary artificial intelligence (AI) discourse, focusing on the widespread adoption of metap
arXiv:2604.21052v1 Announce Type: cross Abstract: We build on the Visual Autoregressive Modeling (VAR) framework and formulate style transfer as conditional discrete sequence modeling in a learned lat