Batch-Adaptive Causal Annotations
arXiv:2502.10605v3 Announce Type: replace-cross Abstract: Estimating the causal effects of interventions is crucial to policy and decision-making, yet outcome data are often missing or subject to non-
Knowledge catalogue
arXiv:2502.10605v3 Announce Type: replace-cross Abstract: Estimating the causal effects of interventions is crucial to policy and decision-making, yet outcome data are often missing or subject to non-
arXiv:2604.16952v1 Announce Type: new Abstract: Learning robust representations across extremely heterogeneous modalities remains a fundamental challenge in multi-modal vision. As a critical and profo
arXiv:2604.17188v1 Announce Type: new Abstract: Multi-role dialogue summarization requires modeling complex interactions among multiple speakers while preserving role-specific information and factual
arXiv:2604.16884v1 Announce Type: new Abstract: The integration of medical imaging and clinical text has enabled the emergence of generalist artificial intelligence (AI) systems for healthcare. Howeve
arXiv:2604.17008v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used to generate narrative content, including children's stories, which play an important role in social a
arXiv:2604.18022v1 Announce Type: cross Abstract: The inverse Potts problem for estimating evolutionary single-site fields and pairwise couplings in homologous protein sequences from their single-site
arXiv:2604.18578v1 Announce Type: new Abstract: Proximal Policy Optimization (PPO) has become the predominant algorithm for on-policy reinforcement learning due to its scalability and empirical robust
arXiv:2602.23580v2 Announce Type: replace Abstract: In the field of educational assessment, automated scoring systems increasingly rely on deep learning and large language models (LLMs). However, thes
arXiv:2604.16680v1 Announce Type: new Abstract: We introduce C-GenReg, a training-free framework for 3D point cloud registration that leverages the complementary strengths of world-scale generative pr
arXiv:2604.17896v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models map multimodal inputs directly to robot actions and are typically trained through large-scale imitation learning. Wh
arXiv:2510.08986v2 Announce Type: replace Abstract: We introduce CAPC-CG, the Chinese Adaptive Policy Communication (Central Government) Corpus, the first open dataset of Chinese policy directives ann
arXiv:2604.17693v1 Announce Type: new Abstract: In cooperative teams where agents act in a fixed order and share a single team reward, it is hard to know how much each agent contributed, and harder st
arXiv:2604.17299v1 Announce Type: new Abstract: Aligning large language models with human preferences must balance two competing goals: responding helpfully to legitimate requests and reliably refusin
arXiv:2604.16861v1 Announce Type: cross Abstract: Standard supervised learning optimizes for predictive accuracy but remains agnostic to the internal geometry of learned features, often yielding repre
arXiv:2604.16411v1 Announce Type: new Abstract: We study asynchronous alignment, a first-class multimodal learning setting in which a dense primary stream must be fused with sporadic external context
arXiv:2604.17614v1 Announce Type: cross Abstract: Skills are a natural unit for describing what a language model can do and how its behavior can be changed. However, existing characterizations rely on
arXiv:2602.19577v2 Announce Type: replace Abstract: Autonomous odor source localization remains a challenging problem for aerial robots due to turbulent airflow, sparse and delayed sensory signals, an
arXiv:2601.05543v2 Announce Type: replace Abstract: Although Speech Large Language Models have achieved notable progress, a substantial modality reasoning gap remains: their reasoning performance on s
arXiv:2604.18236v1 Announce Type: new Abstract: In the context of robot learning for manipulation, curated datasets are an important resource for advancing the state of the art; however, available dat
arXiv:2604.17178v1 Announce Type: new Abstract: Emotional Support Conversation (ESC) plays a critical role in mental health assistance by providing accessible psychological support in real-world appli
arXiv:2603.09108v2 Announce Type: replace Abstract: Medical image retrieval aims to identify clinically relevant lesion cases to support diagnostic decision making, education, and quality control. In
arXiv:2511.17774v3 Announce Type: replace Abstract: Fabrication uncertainty arising from tolerance accumulation, material imperfection, and positioning errors remains a critical barrier to automated r
arXiv:2604.17215v1 Announce Type: new Abstract: Large language models require continuous adaptation to new tasks while preserving safety alignment. However, fine-tuning on even benign data often compr
arXiv:2604.17398v1 Announce Type: new Abstract: We present a methodological framework to discover linguistic and discursive patterns associated to different social groups through contrastive synthetic
arXiv:2604.16412v1 Announce Type: cross Abstract: This paper studies semi-supervised tabular classification in the extreme low-label regime using lightweight base learners. The paper proposes a cooper
arXiv:2604.18245v1 Announce Type: new Abstract: Large language models are increasingly deployed as protocols: structured multi-call procedures that spend additional computation to transform a baseline
arXiv:2604.17555v1 Announce Type: cross Abstract: Agentic search -- the task of training agents that iteratively reason, issue queries, and synthesize retrieved information to answer complex questions
arXiv:2604.17297v1 Announce Type: new Abstract: Long Chain-of-Thought (CoT) reasoning is pivotal for the success of recent reasoning models but suffers from high computational overhead and latency. Wh
arXiv:2604.17217v1 Announce Type: new Abstract: Vision-Language Models (VLMs) achieve strong cross-modal performance, yet recent evidence suggests they over-rely on textual descriptions while under-ut
arXiv:2604.16892v1 Announce Type: new Abstract: Domain generalization (DG) aims to maintain performance under domain shift, which in computer vision appears primarily as stylistic variations that caus
arXiv:2604.18091v1 Announce Type: new Abstract: Recent multimodal large language models have shown promising ability in generating humorous captions for images, yet they still lack stable control over
arXiv:2604.16723v1 Announce Type: cross Abstract: Large Language Models (LLMs) have demonstrated potential in automating scientific ideation, yet current approaches relying on iterative prompting or c
arXiv:2604.16366v1 Announce Type: cross Abstract: Artificial intelligence (AI) tutors have become increasingly popular in learning environments. In this study, we propose an AI agent prototype framewo
arXiv:2601.03154v2 Announce Type: replace Abstract: Reasoning-tuned LLMs utilizing long Chain-of-Thought (CoT) excel at single-answer tasks, yet their ability to model Human Label Variation--which req
arXiv:2604.17389v1 Announce Type: new Abstract: Soft-tissue deformation remains a major limitation in image-guided neurosurgery, where intra-operative anatomy can deviate substantially from pre-operat
arXiv:2511.15669v2 Announce Type: replace Abstract: Does Chain-of-Thought (CoT) reasoning genuinely improve Vision-Language-Action (VLA) models, or does it merely add overhead? Existing CoT-VLA system
arXiv:2604.17207v1 Announce Type: cross Abstract: Iterative alignment methods based on purely greedy updates are remarkably effective in practice, yet existing theoretical guarantees of (O(log T)) KL-
arXiv:2604.16717v1 Announce Type: new Abstract: This paper addresses a critical safety gap in the use Automated Verbal Response Scoring (AVRS). We present a novel hybrid framework for troubled student
arXiv:2512.12022v2 Announce Type: replace Abstract: Decentralized federated learning (DFL) has emerged as a promising paradigm that enables multiple clients to collaboratively train machine learning m
arXiv:2604.16318v1 Announce Type: cross Abstract: Large language models (LLMs) and cross-encoder rerankers have gained attention for improving recommender systems, particularly in cold-start scenarios
arXiv:2601.03559v2 Announce Type: replace Abstract: Chain-of-Thought (CoT) reasoning improves multi-step mathematical problem solving in large language models but remains vulnerable to exposure bias a
arXiv:2603.04881v2 Announce Type: replace Abstract: Differentially private learning is essential for training models on sensitive data, but empirical studies consistently show that it can degrade perf
arXiv:2604.18143v1 Announce Type: cross Abstract: This paper investigates the off-policy evaluation (OPE) problem from a distributional perspective. Rather than focusing solely on the expectation of t
arXiv:2604.17568v1 Announce Type: new Abstract: Given only observational data X = g(Z), where both the latent variables Z and the generating process g are unknown, recovering Z is ill-posed without ad
arXiv:2604.08302v2 Announce Type: replace Abstract: We present DMax, a new paradigm for efficient diffusion language models (dLLMs). It mitigates error accumulation in parallel decoding, enabling aggr
arXiv:2604.17718v1 Announce Type: new Abstract: Many benchmarks show that large language models can answer direct questions about culture. We study a different question: do they also change how they s
arXiv:2604.18161v1 Announce Type: new Abstract: In policy gradient reinforcement learning, access to a differentiable model enables 1st-order gradient estimation that accelerates learning compared to
arXiv:2604.17628v1 Announce Type: new Abstract: Wales' political landscape has been marked by growing accusations of bias in Welsh media. This paper takes the first computational step toward testing t
arXiv:2604.16979v1 Announce Type: cross Abstract: High-quality and diverse multimodal data are essential for improving vision-language models (VLMs), yet existing datasets often contain noisy, redunda
arXiv:2510.00761v5 Announce Type: replace Abstract: Large language model (LLM) unlearning aims to surgically remove the influence of undesired data or knowledge from an existing model while preserving
arXiv:2506.12622v2 Announce Type: replace Abstract: Deep reinforcement learning (RL) has achieved remarkable success, yet its deployment in real-world scenarios is often limited by vulnerability to en
arXiv:2604.17195v1 Announce Type: new Abstract: Storyboard synthesis plays a crucial role in visual storytelling, aiming to generate coherent shot sequences that visually narrate cinematic events with
arXiv:2512.16055v2 Announce Type: replace Abstract: Safety-critical corner cases, difficult to collect in the real world, are crucial for evaluating end-to-end autonomous driving. Adversarial interact
arXiv:2604.17841v1 Announce Type: new Abstract: Most autonomous driving safety benchmarks use time-to-collision (TTC) to assess risk and guide safe behaviour. However, TTC-based methods treat risk as
arXiv:2604.18563v1 Announce Type: new Abstract: A recent study (Kuribayashi et al., 2025) has shown that human sentence processing behavior, typically measured on syntactically unchallenging construct
arXiv:2604.17037v1 Announce Type: new Abstract: Deception detection is of great significance for ensuring information security and conducting public opinion analysis, with personality factors and emot
arXiv:2604.17710v1 Announce Type: new Abstract: Zero-shot learning (ZSL) aims to recognize unseen classes without visual instances. However, existing methods usually assume clean labels, overlooking r
arXiv:2601.22149v2 Announce Type: replace Abstract: The development of autonomous web agents, powered by Large Language Models (LLMs) and reinforcement learning (RL), represents a significant step tow
arXiv:2604.17838v1 Announce Type: new Abstract: Generative modeling within constrained sets is essential for scientific and engineering applications involving physical, geometric, or safety requiremen
arXiv:2604.17747v1 Announce Type: new Abstract: This paper considers reinforcement learning from human feedback in a federated learning setting with resource-constrained agents, such as edge devices.