Conformal Agent Error Attribution
arXiv:2605.06788v1 Announce Type: new Abstract: When multi-agent systems (MAS) fail, identifying where the decisive error occurred is the first step for automated recovery to an earlier state. Error a
Knowledge catalogue
arXiv:2605.06788v1 Announce Type: new Abstract: When multi-agent systems (MAS) fail, identifying where the decisive error occurred is the first step for automated recovery to an earlier state. Error a
arXiv:2605.08077v1 Announce Type: new Abstract: Knowledge Graph Question Answering (KGQA) has shown promise for grounded and interpretable reasoning, yet existing approaches often fail to provide reli
arXiv:2605.07115v1 Announce Type: new Abstract: Stochastic bandit algorithms are usually analyzed under a mean-reward criterion, yet many problems favor arms with strong upper-tail performance, which
arXiv:2605.06905v1 Announce Type: new Abstract: Modern generative modeling is dominated by transport from a noise prior to data. We propose an alternative paradigm in which generation is performed by
arXiv:2505.14113v3 Announce Type: replace Abstract: Most machine learning-based image segmentation models produce pixel-wise confidence scores that represent the model's predicted probability for each
arXiv:2605.07907v1 Announce Type: cross Abstract: Vision-Language Latent Diffusion Models (LDMs) (Rombach et al., 2022) provide powerful generative priors for inverse problems. However, existing LDM-b
arXiv:2603.05687v3 Announce Type: replace Abstract: Contact-rich dexterous manipulation with multi-finger hands remains an open challenge in robotics because task success depends on multi-point contac
arXiv:2603.02274v2 Announce Type: replace-cross Abstract: Precision oncology is currently limited by the small-N, large-P paradox, where high-dimensional genomic data is abundant but pharmacological r
arXiv:2511.18085v4 Announce Type: replace-cross Abstract: Vision-language-action (VLA) models show promising knowledge accumulation ability from pretraining, yet continual learning in VLA remains chal
arXiv:2605.06870v1 Announce Type: new Abstract: While many approaches to improve VQ-VAE performance focus on codebook size and utilization, the effect of dimensional collapse, where trained VQ-VAE rep
arXiv:2601.15884v2 Announce Type: replace Abstract: Contrast-enhanced imaging is central to oncologic diagnosis, but contrast agents can be contraindicated for many of the patients who need them most.
arXiv:2605.07123v1 Announce Type: new Abstract: In-context reinforcement learning (ICRL) refers to the ability of RL agents to adapt to new tasks at inference time without parameter updates by conditi
arXiv:2605.07959v1 Announce Type: new Abstract: Transformers have revolutionized machine learning and deploying attention layers in the model is increasingly standard across a myriad of applications.
arXiv:2605.07386v1 Announce Type: new Abstract: Convex Optimization with Nested Evolving Feasible Sets (CONES)} is considered where the objective function f remains fixed but the feasible region evolv
arXiv:2605.07171v1 Announce Type: new Abstract: The classic multi-armed bandit (MAB) problem tackles the challenge of accruing maximum reward while making decisions under uncertainty. However, in appl
arXiv:2605.07193v1 Announce Type: new Abstract: Generative modeling over discrete structures underpins applications across deep learning, from biological sequence design and code generation to large l
arXiv:2605.05732v2 Announce Type: replace-cross Abstract: Large language models (LLMs) can acquire new capabilities through fine-tuning, but continual adaptation often leads to catastrophic forgetting
arXiv:2605.07705v1 Announce Type: cross Abstract: We give a novel logical characterization of encoder-decoder transformers, the foundational architecture for LLMs that also sees use in various setting
arXiv:2605.06115v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs), trained primarily on English-centric data, frequently generate culturally inappropriate or misaligned resp
arXiv:2601.03728v3 Announce Type: replace-cross Abstract: Composed Image Retrieval (CIR) enables users to search for target images using both a reference image and manipulation text, offering substant
arXiv:2605.07325v1 Announce Type: cross Abstract: Deploying massive large language models (LLMs) as continuous cognitive engines for robotics is bottlenecked by the time-to-first-token (TTFT) latency
arXiv:2605.07724v1 Announce Type: cross Abstract: Recursive retraining of generative models poses a critical representation challenge: when synthetic outputs are curated based on a fixed reward signal
arXiv:2605.07902v1 Announce Type: new Abstract: Submodular functions -- functions exhibiting diminishing returns -- are central to machine learning. When the objective is monotone and non-negative, th
arXiv:2605.07830v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed as autonomous agents in offensive cybersecurity. In this paper, we reveal an interesting phenom
arXiv:2605.07453v1 Announce Type: new Abstract: Ancient and endangered languages pose a unique challenge for NLP: their datasets are inherently scarce, difficult to expand, and built from formulaic co
arXiv:2605.06865v1 Announce Type: new Abstract: Large language models (LLMs) are pre-trained and post-trained on vast amounts of loosely curated data, raising the possibility that these models may hav
arXiv:2605.07314v1 Announce Type: cross Abstract: Knowledge Graphs (KGs) have proven highly effective for recommendation systems by capturing latent item relationships, while recent integration of Lar
arXiv:2605.07665v1 Announce Type: cross Abstract: Estimating counterfactual distributions under interventions is central to treatment risk assessment and counterfactual generation tasks. Existing appr
arXiv:2605.06971v1 Announce Type: cross Abstract: Classical optimization theory largely focuses on fixed objective functions, whereas many modern learning systems operate in dynamic environments where
arXiv:2604.02753v2 Announce Type: replace Abstract: Open-vocabulary Object Detection (OVOD) enables models to recognize objects beyond predefined categories, but existing approaches remain limited in
arXiv:2510.18516v3 Announce Type: replace-cross Abstract: Neural recordings exhibit a distinctive form of heterogeneity rooted in differences in cell types, intrinsic circuit dynamics, and stochastic
arXiv:2603.04676v2 Announce Type: replace-cross Abstract: Multi-image reasoning remains a significant challenge for vision-language models (VLMs). We investigate a previously overlooked phenomenon: du
arXiv:2605.07074v1 Announce Type: new Abstract: Detecting AI-generated images across unseen architectures remains challenging, as existing models often overfit to generator-specific fingerprints and s
arXiv:2601.15127v3 Announce Type: replace-cross Abstract: Deploying federated learning across heterogeneous IoT device fleets requires tailored neural network architectures for each device class, yet
arXiv:2508.01994v2 Announce Type: replace Abstract: As the application of deep learning in dermatology continues to grow, the recognition of melanoma has garnered significant attention, demonstrating
arXiv:2605.07940v1 Announce Type: new Abstract: Exemplar-based image editing applies a transformation defined by a source-target image pair to a new query image. Existing methods rely on a pair-of-pai
arXiv:2605.07024v1 Announce Type: new Abstract: Large Language Models for code generation frequently produce hallucinations in Fill-in-the-Middle (FIM) tasks -- plausible but incorrect completions suc
arXiv:2603.28113v2 Announce Type: replace Abstract: The global Lipschitz constant of a neural network is related to robustness and generalization, yet unlike in many classical models, it is not plainl
arXiv:2605.07694v1 Announce Type: cross Abstract: Single-channel speaker distance estimation has recently achieved centimeter-level accuracy in simulated environments, yet it remains unclear which com
arXiv:2510.04850v3 Announce Type: replace-cross Abstract: Reasoning distillation has emerged as a prevailing paradigm for transferring reasoning capabilities from large reasoning models to small langu
arXiv:2512.17129v2 Announce Type: replace Abstract: Biological systems can form complex three-dimensional structures through the collective behavior of agents that share a common update rule and opera
arXiv:2605.07781v1 Announce Type: new Abstract: Explicit neural representations such as 3D Gaussian Splatting (3DGS) enable high-fidelity and real-time novel view synthesis, yet optimize for alpha-com
arXiv:2605.07674v1 Announce Type: cross Abstract: Regulatory audits of AI systems increasingly rely on differential privacy (DP) to protect training data and model internals. We study audit design whe
arXiv:2605.07210v1 Announce Type: cross Abstract: PromptReps showed that an autoregressive language model can be used directly as a retriever by prompting it to generate dense and sparse representatio
arXiv:2605.07503v1 Announce Type: new Abstract: Efficiently aligning large-scale video diffusion models with human intent requires a scalable and trajectory-aware pathway that bridges the inherent dis
arXiv:2601.21951v2 Announce Type: replace-cross Abstract: We develop diffusion-based samplers for target distributions known up to a normalising constant. To this end, we rely on the well-known diffus
arXiv:2605.07494v1 Announce Type: new Abstract: Continual learning enables vision-language models to accumulate knowledge and adapt to evolving tasks without retraining from scratch. However, in multi
arXiv:2605.07221v1 Announce Type: new Abstract: Adapting foundation models to medical segmentation typically requires either backbone fine-tuning or high-capacity task-specific decoders, both of which
arXiv:2508.20909v2 Announce Type: replace Abstract: Foundation models pre-trained on large-scale natural image datasets offer a powerful paradigm for medical image segmentation. However, effectively t
arXiv:2506.13351v3 Announce Type: replace-cross Abstract: Reinforcement learning (RL) training of large language models (LLMs) on unverifiable tasks is challenging even when a reasonable-quality refer
arXiv:2602.22831v2 Announce Type: replace-cross Abstract: Moral benchmarks for LLMs typically score models on context-free prompts, implicitly treating the measured choice rate as stable. We test this
arXiv:2412.11194v2 Announce Type: replace-cross Abstract: Security vulnerabilities in software can have severe consequences; however, manual vulnerability detection is costly and does not scale, espec
arXiv:2605.07662v1 Announce Type: new Abstract: Low-precision number formats are widely used in modern machine learning systems due to their efficiency. Accurate direction representation is key to the
arXiv:2605.07551v1 Announce Type: new Abstract: Standard Importance Sampling (IS) collapses under label corruption because high-norm examples, prioritized for variance reduction, are often adversarial
arXiv:2605.07351v1 Announce Type: new Abstract: While Gaussian Splatting-based Feature Fields (GSFFs) have shown promise for visual localization, this paper highlights that photometrically optimized G
arXiv:2506.23875v4 Announce Type: replace-cross Abstract: Sequential computation via autoregressive generation can make difficult tasks learnable, but the generation order of intermediate states stron
arXiv:2602.16928v3 Announce Type: replace-cross Abstract: Much of the advancement in Multi-Agent Reinforcement Learning (MARL) for imperfect-information games has historically depended on the manual,
arXiv:2605.07323v1 Announce Type: new Abstract: Discovering governing differential equations from observational data is a fundamental challenge in scientific machine learning. Existing symbolic regres
arXiv:2605.06785v1 Announce Type: cross Abstract: Inference-time scaling methods rely on Process Reward Models (PRMs), which are often poorly calibrated and overestimate success probabilities. We prop
arXiv:2605.07844v1 Announce Type: new Abstract: Energy-based learning is a powerful framework for generative modelling, but its training is inherently non-convex, leading potentially to sensitivity to