CoT-Guard: Small Models for Strong Monitoring
arXiv:2605.12746v1 Announce Type: cross Abstract: Monitoring the chain-of-thought (CoT) of reasoning models is a promising approach for detecting covert misbehavior (i.e., hidden objectives) in code g
Knowledge catalogue
arXiv:2605.12746v1 Announce Type: cross Abstract: Monitoring the chain-of-thought (CoT) of reasoning models is a promising approach for detecting covert misbehavior (i.e., hidden objectives) in code g
arXiv:2605.13077v1 Announce Type: cross Abstract: Responsibility allocation -- determining the extent to which agents are accountable for outcomes -- is a fundamental challenge in the design and analy
arXiv:2605.12938v1 Announce Type: cross Abstract: Camera-conditioned video generation requires positional encoding that remains reliable under changes in camera motion, lens configuration, and scene s
arXiv:2605.12545v1 Announce Type: cross Abstract: Aesthetic image cropping aims to enhance the aesthetic quality of an image by improving its composition through spatial cropping. Previous methods oft
arXiv:2605.13686v1 Announce Type: cross Abstract: Medical image-to-image (I2I) translation enables virtual scanning, i.e. the synthesis of a target imaging modality from a source one without additiona
arXiv:2605.13452v1 Announce Type: cross Abstract: Recent advances in visuomotor policy learning have enabled robots to perform control directly from visual inputs. Yet, extending such end-to-end learn
arXiv:2605.13276v1 Announce Type: new Abstract: The rapid evolution of Embodied AI has enabled Vision-Language-Action (VLA) models to excel in multimodal perception and task execution. However, applyi
arXiv:2605.12906v1 Announce Type: cross Abstract: Data selection during supervised fine-tuning (SFT) can critically change the behavior of large language models (LLMs). Although existing work has stud
arXiv:2509.19538v2 Announce Type: replace-cross Abstract: Diffusion-based world models have demonstrated strong capabilities in synthesizing realistic long-horizon trajectories for offline reinforceme
arXiv:2602.22847v2 Announce Type: replace-cross Abstract: The concept of ranking aggregation plays a central role in preference analysis, and numerous algorithms for calculating median rankings, often
arXiv:2605.13540v1 Announce Type: cross Abstract: Dynamic graphs are ubiquitous in real-world systems, and building generalizable dynamic Graph Foundation Models has become a frontier in graph learnin
arXiv:2601.00417v3 Announce Type: replace-cross Abstract: Transformer residual streams evolve by additive accumulation: each layer appends a feature update to a shared hidden state, but has no direct
arXiv:2502.20427v3 Announce Type: replace-cross Abstract: Deepfakes - manipulated or forged audio and video media - pose significant security risks to individuals, organizations, and society at large.
arXiv:2603.20521v2 Announce Type: replace-cross Abstract: Distributed reinforcement learning trains on data from stale, buggy, or mismatched actors, producing actions with high surprisal (negative log
arXiv:2605.13287v1 Announce Type: cross Abstract: Most exploration algorithms search broadly until uncertainty is resolved. When the action space is too large to resolve within budget, practitioners d
arXiv:2509.25781v2 Announce Type: replace Abstract: We address the issue of defining a semantics for deontic argumentation that supports weak permission. Some recent results show that grounded semanti
arXiv:2605.13790v1 Announce Type: cross Abstract: Partial differential equations (PDEs) are fundamental for modeling complex natural and physical phenomena. In many real-world applications, however, o
arXiv:2605.12522v1 Announce Type: cross Abstract: Diffusion language models (DLMs) are promising alternatives to autoregressive language models (ARMs), yet the intrinsic differences in their generated
arXiv:2512.13399v2 Announce Type: replace Abstract: Crafting effective reward signals remains a central challenge in Reinforcement Learning (RL), especially for complex reasoning tasks. Existing autom
arXiv:2605.13282v1 Announce Type: new Abstract: Classical planners can effectively solve very large deterministic MDPs represented in STRIPS or PDDL where states are sets of atoms over objects and rel
arXiv:2605.12702v1 Announce Type: new Abstract: General-purpose safety benchmarks for large language models do not adequately evaluate disability-related harms. We introduce DisaBench: a taxonomy of t
arXiv:2602.03429v2 Announce Type: replace Abstract: To handle ambiguous and open-ended requests, Large Language Models (LLMs) are increasingly trained to interact with users to surface intents they ha
arXiv:2605.13484v1 Announce Type: cross Abstract: Calibration is commonly evaluated by comparing model confidence with its empirical correctness, implicitly treating reliability as a function of the c
arXiv:2605.13296v1 Announce Type: new Abstract: Multi-Agent Path Finding (MAPF) is a coordination problem that requires computing globally consistent, collision-free trajectories from individual start
arXiv:2605.12805v1 Announce Type: cross Abstract: MeanFlow enables one-step generation in continuous spaces by learning an average velocity over a time interval rather than the instantaneous velocity
arXiv:2605.12574v1 Announce Type: cross Abstract: Vision-language models (VLMs) are trained on large-scale image-text corpora that may contain private, copyrighted, or otherwise sensitive data, motiva
arXiv:2605.13332v1 Announce Type: new Abstract: Argumentation is an important topic of AI for modelling and reasoning about arguments. In abstract argumentation, we consider directed graphs, so-called
arXiv:2605.12673v1 Announce Type: new Abstract: Agent benchmarks have become the de facto measure of frontier AI competence, guiding model selection, investment, and deployment. However, reward hackin
arXiv:2605.12701v1 Announce Type: cross Abstract: Machine learning algorithms in socially sensitive domains (e.g., credit decisions) often focus on equalizing predictive outcomes. However, satisfying
arXiv:2605.13084v1 Announce Type: cross Abstract: Meta-learning has been shown to have better performance than supervised learning for few-shot monolingual spoken word classification. However, the met
arXiv:2605.12516v1 Announce Type: cross Abstract: General-purpose large language models (LLMs) often struggle to generate reliable responses in specialized engineering domains due to limited domain gr
arXiv:2605.13568v1 Announce Type: cross Abstract: Myocardial infarction (MI) is a leading cause of death, and its adverse outcomes are urgent to predict. Yet ECG-based prognostic models underperform b
arXiv:2605.13194v1 Announce Type: cross Abstract: Electrocardiogram (ECG) arrhythmia classification remains challenging due to signal variability, noise, limited labeled data, and the difficulty in ac
arXiv:2605.12887v1 Announce Type: cross Abstract: Web-enabled LLM agents are changing how online information influences search outcomes. Existing Generative Engine Optimization (GEO) studies mainly fo
arXiv:2505.17469v2 Announce Type: replace-cross Abstract: Compression and generalization are fundamentally related through Solomonoff induction and the minimum description length principle (MDL), whic
arXiv:2605.13335v1 Announce Type: new Abstract: Embodied agents in household environments must plan under partial observation: they need to remember objects, track state changes, and recover when acti
arXiv:2605.12920v1 Announce Type: cross Abstract: Effective collaboration between embodied agents requires more than acting in a shared environment; it demands communication grounded in each agent's e
arXiv:2605.12798v1 Announce Type: cross Abstract: Fine-tuning LLMs on narrow harmful datasets can induce Emergent Misalignment (EM), where models exhibit misaligned behavior far beyond the fine-tuning
arXiv:2605.13789v1 Announce Type: cross Abstract: Protein structure tokenizers (PSTs) are workhorses in protein language modeling, function prediction, and evolutionary analysis. However, existing PST
arXiv:2602.05000v2 Announce Type: replace-cross Abstract: Reward guidance, also known as posterior sampling, is a popular method for test-time adaptation and post-training in continuous diffusion mode
arXiv:2605.13841v1 Announce Type: cross Abstract: Voice agents, artificial intelligence systems that conduct spoken conversations to complete tasks, are increasingly deployed across enterprise applica
arXiv:2605.13152v1 Announce Type: cross Abstract: We introduce EvObj for unsupervised 3D instance segmentation that bridges the geometric domain gap between synthetic pretraining data and real-world p
arXiv:2508.09320v3 Announce Type: replace-cross Abstract: Graph neural networks (GNNs) are increasingly often employed in high-stakes applications, such as fraud detection or healthcare, but are susce
arXiv:2605.11679v2 Announce Type: replace Abstract: In the realm of multi-objective alignment for large language models, balancing disparate human preferences often manifests as a zero-sum conflict. S
arXiv:2605.12523v1 Announce Type: cross Abstract: Generative Artificial Intelligence (AI) introduces new considerations for English as a foreign language (EFL) writing pedagogy. This study explores ho
arXiv:2602.04264v2 Announce Type: replace-cross Abstract: The choice of activation function fundamentally shapes the representational capacity and parameter efficiency of deep neural networks, yet mos
arXiv:2605.13030v1 Announce Type: cross Abstract: Model merging combines task experts into one model and avoids joint training, retraining, or deploying many expert models, but the merged model often
arXiv:2512.21075v2 Announce Type: replace-cross Abstract: Deep neural networks have achieved remarkable success in practice, yet a mechanistic understanding of how features evolve during training rema
arXiv:2605.12704v1 Announce Type: cross Abstract: A fundamental challenge in symbolic regression (SR) is efficiently recovering complex mathematical expressions from observational data. Although this
arXiv:2604.00001v2 Announce Type: replace-cross Abstract: Gradient-based data selection offers a principled framework for estimating sample utility in large language model (LLM) fine-tuning, but exist
arXiv:2603.15854v2 Announce Type: replace-cross Abstract: Sampling from a categorical distribution is mathematically simple, but in large-vocabulary decoding, it often triggers extra memory traffic an
arXiv:2512.07112v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have demonstrated remarkable performance due to their large parameter counts and extensive training data. However
arXiv:2605.13171v1 Announce Type: new Abstract: As automated reasoning systems advance rapidly, there is a growing need for research-level formal mathematical problems to accurately evaluate their cap
arXiv:2605.12826v1 Announce Type: cross Abstract: The proliferation of sophisticated image editing tools and generative artificial intelligence models has made verifying the authenticity of digital im
arXiv:2603.05093v2 Announce Type: replace-cross Abstract: Feature attributions often hide a critical modeling choice: they explain a prediction along a counterfactual path from a reference state to an
arXiv:2605.12733v1 Announce Type: cross Abstract: Given a generalist model, learning a task-relevant specialist representation is fundamental for downstream applications. Identifiability, the asymptot
arXiv:2603.27910v2 Announce Type: replace Abstract: AI agents that interact with users across multiple sessions require persistent long-term memory to maintain coherent, personalized behavior. Current
arXiv:2605.13555v1 Announce Type: cross Abstract: Radiation therapy (RT) requires precise dose delivery over multiple fractions, with CT fundamental for treatment planning due to its electron density
arXiv:2512.10857v2 Announce Type: replace-cross Abstract: Transport-based methods have emerged as a leading paradigm for building generative models from large, clean datasets. However, in many scienti
arXiv:2508.14302v2 Announce Type: replace-cross Abstract: Inference-time sparsification is a promising path to deploy large language models (LLMs) on resource-constrained devices, yet existing trainin