Selective Conformal Risk Control
arXiv:2512.12844v2 Announce Type: replace-cross Abstract: Reliable uncertainty quantification is essential for deploying machine learning systems in high-stakes domains. Conformal prediction provides
Knowledge catalogue
arXiv:2512.12844v2 Announce Type: replace-cross Abstract: Reliable uncertainty quantification is essential for deploying machine learning systems in high-stakes domains. Conformal prediction provides
arXiv:2604.24313v1 Announce Type: cross Abstract: Training large-scale deep neural networks effectively and stably is essential for applying deep learning across various fields. However, conventional
arXiv:2312.15020v4 Announce Type: replace-cross Abstract: Technical debt (TD) refers to the long-term costs associated with suboptimal design or code decisions in software development, often made to m
arXiv:2604.22939v1 Announce Type: cross Abstract: While the next-token prediction (NTP) paradigm enables large language models (LLMs) to express their intrinsic knowledge, its sequential nature constr
arXiv:2604.24102v1 Announce Type: new Abstract: Synthesizing a reactive system from specifications given in linear temporal logic (LTL) is a classical problem, finding its applications in safety-criti
arXiv:2604.23747v1 Announce Type: cross Abstract: Recent mixed-policy optimization methods for LLM reasoning that interleave or blend supervised and reinforcement learning signals report improvements
arXiv:2604.22825v1 Announce Type: cross Abstract: Large segmentation foundation models such as the Segment Anything Model (SAM) have reshaped promptable segmentation in natural images, and recent effo
arXiv:2506.05425v2 Announce Type: replace-cross Abstract: Understanding social interaction, which encompasses perceiving numerous and subtle multimodal cues, inferring unobservable mental states and r
arXiv:2604.22875v1 Announce Type: cross Abstract: When answering questions about images, humans naturally point, label, and draw to explain their reasoning. In contrast, modern vision-language models
arXiv:2604.24594v1 Announce Type: cross Abstract: As large language models (LLMs) evolve into agentic problem solvers, they increasingly rely on external, reusable skills to handle tasks beyond their
arXiv:2604.23263v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly utilized in various complex reasoning tasks due to their excellent instruction following capability. How
arXiv:2512.03028v3 Announce Type: replace-cross Abstract: Data-driven motion priors that can guide agents toward producing naturalistic behaviors play a pivotal role in creating life-like virtual char
arXiv:2604.23905v1 Announce Type: cross Abstract: Threat modeling for cyber-physical systems (CPS) remains a largely manual exercise. This project presents SMSI (System Model Security Inference), a hy
arXiv:2604.23392v1 Announce Type: new Abstract: Refereeing is vital in sports, where fair, accurate, and explainable decisions are fundamental. While intelligent assistant technologies are being widel
arXiv:2604.24306v1 Announce Type: cross Abstract: Accurate forecasting of solar power output is essential for efficient integration of renewable energy into the grid. In this study, an attention-based
arXiv:2604.24199v1 Announce Type: cross Abstract: We propose Speech Enhancement based on Drifting Models (DriftSE), a novel generative framework that formulates denoising as an equilibrium problem. Ra
arXiv:2502.12672v4 Announce Type: replace-cross Abstract: Fine-tuning speech representation models can enhance performance on specific tasks but often compromises their cross-task generalization abili
arXiv:2604.23432v1 Announce Type: cross Abstract: Reliable depth estimation from spherical images is crucial for 360{eg} vision in robotic navigation and immersive scene understanding. However, the on
arXiv:2604.24449v1 Announce Type: cross Abstract: Training machine learning models for robotic tactile sensing requires vast amounts of data, yet obtaining realistic interaction data remains a challen
arXiv:2505.16637v4 Announce Type: replace-cross Abstract: Large language models (LLMs) have recently demonstrated remarkable capabilities in machine translation (MT). However, most advanced MT-specifi
arXiv:2511.17902v3 Announce Type: replace-cross Abstract: Distributed Fiber Optic Sensing (DFOS) is promising for long-range perimeter security, yet practical deployment faces three key obstacles: sev
arXiv:2604.24544v1 Announce Type: new Abstract: The increasing reliance on Large Language Models (LLMs) across diverse sectors highlights the need for robust domain-specific and language-specific eval
arXiv:2604.22782v1 Announce Type: cross Abstract: Serving transformer language models with high throughput requires caching Key-Values (KVs) to avoid redundant computation during autoregressive genera
arXiv:2604.23198v1 Announce Type: new Abstract: Current video moment retrieval excels at action-centric tasks but struggles with narrative content. Models can see extit{what is happening} but fail to
arXiv:2604.24156v1 Announce Type: cross Abstract: Efficient and fair spectrum allocation is a central challenge in 6G networks, where massive connectivity and heterogeneous services continuously compe
arXiv:2604.22757v1 Announce Type: cross Abstract: We introduce StratRAG, an open-source retrieval evaluation dataset for benchmarking Retrieval-Augmented Generation (RAG) systems on multi-hop reasonin
arXiv:2604.23646v1 Announce Type: new Abstract: Recent evidence suggests that frontier AI systems can exhibit agentic misalignment, generating and executing harmful actions derived from internally con
arXiv:2604.22843v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) has been proposed to mitigate hallucinations in large language models (LLMs), where generated outputs may be fact
arXiv:2402.19088v5 Announce Type: replace-cross Abstract: Live languages continuously evolve to integrate the cultural change of human societies. This evolution manifests through neologisms (new words
arXiv:2604.22852v1 Announce Type: cross Abstract: Cloud-hosted LLM inference for autonomous driving adds round-trip delay and depends on stable connectivity, while purely local edge models struggle un
arXiv:2604.24346v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly deployed as evaluators in tasks requiring nuanced image understanding, yet their reliability in scoring
arXiv:2604.23806v1 Announce Type: cross Abstract: The reverse process in score-based diffusion models is formally equivalent to overdamped Langevin dynamics in a time-dependent energy landscape. In ou
arXiv:2509.25346v2 Announce Type: replace Abstract: Predicting cellular responses to genetic perturbations represents a fundamental challenge in systems biology, critical for advancing therapeutic dis
arXiv:2604.24088v1 Announce Type: cross Abstract: Handling communication overhead in large-scale tensor-parallel training remains a critical challenge due to the dense, near-zero distributions of inte
arXiv:2604.23703v1 Announce Type: cross Abstract: Slide-based teaching is widely used in higher education, yet in online, hybrid, and asynchronous contexts, slides often lose the instructor presence,
arXiv:2604.23623v1 Announce Type: new Abstract: Recent advancements in large language models (LLMs) have catalyzed the rise of reasoning-intensive inference paradigms, where models perform explicit st
arXiv:2604.23964v1 Announce Type: cross Abstract: Patients with dementia typically exhibit cognitive impairment, which is routinely assessed using the Mini-Mental State Examination (MMSE). Concurrentl
arXiv:2604.24005v1 Announce Type: cross Abstract: On-policy distillation (OPD) has shown strong potential for transferring reasoning ability from frontier or domain-specific models to smaller students
arXiv:2601.04204v2 Announce Type: replace-cross Abstract: The scalability of high-quality online education is hindered by the high costs and slow cycles of manual content creation. Despite advancement
arXiv:2509.00072v3 Announce Type: replace Abstract: Post-cutoff performance decay has been widely interpreted as a temporal signal for benchmark contamination. We critically examine this belief and de
arXiv:2309.06020v3 Announce Type: replace-cross Abstract: Technical debt refers to the consequences of sub-optimal decisions made during software development that prioritize short-term benefits over l
arXiv:2604.24155v1 Announce Type: cross Abstract: The quest to align machine behavior with human values raises fundamental questions about the moral frameworks that should govern AI decision-making. M
arXiv:2602.11318v3 Announce Type: replace Abstract: In machine learning, 'ground truth' refers to the assumed correct labels used to train and evaluate models. However, the foundational 'ground truth'
arXiv:2512.08345v2 Announce Type: replace Abstract: Workplace toxicity is widely recognized as detrimental to organizational culture, yet quantifying its direct impact on operational efficiency remain
arXiv:2604.22767v1 Announce Type: cross Abstract: Ethical discourse on AI in healthcare has focused predominantly on back-end concerns such as bias, fairness and explainability, while the front-end in
arXiv:2604.24083v1 Announce Type: new Abstract: This study introduces the Kerimov-Alekberli model, a novel information-geometric framework that redefines AI safety by formally linking non-equilibrium
arXiv:2604.23750v1 Announce Type: cross Abstract: Hypernetwork-based methods such as Doc-to-LoRA internalize a document into an LLM's weights in a single forward pass, but they fail systematically on
arXiv:2604.22951v1 Announce Type: new Abstract: Natural language data follows a power-law distribution, with most knowledge and skills appearing at very low frequency. While a common intuition suggest
arXiv:2604.24079v1 Announce Type: cross Abstract: Large Language Models (LLMs) reveal inherent and distinctive personas through dialogue. However, most existing persona discovery approaches rely on su
arXiv:2604.24668v1 Announce Type: new Abstract: Given the increased use of LLMs in financial systems today, it becomes important to evaluate the safety and robustness of such systems. One failure mode
arXiv:2604.22771v1 Announce Type: cross Abstract: Language models cannot be random. This paper introduces Entropic Deviation (ED), the normalised KL divergence between a model's token distribution and
arXiv:2601.15485v2 Announce Type: replace-cross Abstract: Federal research funding shapes the direction, diversity, and impact of the US scientific enterprise. Large language models (LLMs) are rapidly
arXiv:2604.23058v1 Announce Type: cross Abstract: Firms are deploying more capable AI systems, but organizational controls often have not kept pace. These systems can generate greater productivity gai
arXiv:2503.09101v4 Announce Type: replace-cross Abstract: Uniform manifold approximation and projection (UMAP) is among the most popular neighbor embedding methods. The method samples pairs of point i
arXiv:2604.22778v1 Announce Type: cross Abstract: We present the first systematic study of weight matrix singular value spectra during transformer pretraining, tracking full SVD decompositions of ever
arXiv:2511.08577v2 Announce Type: replace-cross Abstract: Improving reasoning abilities of Large Language Models (LLMs), especially under parameter constraints, is crucial for real-world applications.
arXiv:2604.23605v1 Announce Type: new Abstract: The application of large language models (LLMs) in clinical decision support faces significant challenges of 'tunnel vision' and diagnostic hallucinatio
arXiv:2604.23859v1 Announce Type: new Abstract: With spotforecast2-safe we present an integrated Compliance-by-Design approach to Python-based point forecasting of time series in safety-critical envir
arXiv:2510.09859v4 Announce Type: replace-cross Abstract: A seller of a dynamic information service under an information-throughput constraint screens buyers who privately differ in urgency. We charac
arXiv:2604.23231v1 Announce Type: cross Abstract: Semantic Communication (SC) backdoor attacks aim to utilize triggers to manipulate the system into producing predetermined outputs via backdoored shar