Proper Calibeating
arXiv:2605.26703v1 Announce Type: cross Abstract: The classic concept of 'calibrated forecasts' and its more recent refinement, 'calibeating,' are defined with respect to the standard quadratic scorin
Knowledge catalogue
arXiv:2605.26703v1 Announce Type: cross Abstract: The classic concept of 'calibrated forecasts' and its more recent refinement, 'calibeating,' are defined with respect to the standard quadratic scorin
arXiv:2605.26255v1 Announce Type: cross Abstract: Early prediction of respiratory failure is critical for timely clinical intervention in intensive care units. Existing electronic health record (EHR)-
arXiv:2510.18668v2 Announce Type: replace-cross Abstract: The vast majority of cardiovascular diseases may be preventable if early signs and risk factors are detected. Cardiovascular monitoring with b
arXiv:2605.26243v1 Announce Type: new Abstract: Graph neural networks (GNNs) achieve strong performance on relational data, but real-world graphs are often distributed across organizations that cannot
arXiv:2605.26625v1 Announce Type: new Abstract: We present a provably safe sampling-based motion planning algorithm for robotic systems affected by random disturbances of unknown distribution. We cons
arXiv:2605.26801v1 Announce Type: new Abstract: Psychological constructs are often measured in separate instruments, datasets, and research traditions, which makes direct comparison difficult. This pa
arXiv:2507.16116v2 Announce Type: replace Abstract: The rapid advancement of video diffusion models has been hindered by fundamental limitations in temporal modeling, particularly the rigid synchroniz
arXiv:2508.02806v3 Announce Type: replace Abstract: Recently, a significant improvement in the accuracy of 3D human pose estimation has been achieved by combining convolutional neural networks (CNNs)
arXiv:2605.27318v1 Announce Type: new Abstract: Video spatial reasoning requires accumulating viewpoint-dependent evidence over time while retaining information useful to the question being asked. Exi
arXiv:2605.26339v1 Announce Type: cross Abstract: Scalar post-training quantizers discard pairwise coordinate structure within weight rows. We introduce QAM-W (Quadrature Amplitude Modulation for Weig
arXiv:2605.27210v1 Announce Type: cross Abstract: We adapt Microsoft's QuantumKatas -- a well-established quantum computing curriculum -- from Q# to Qiskit, the most widely-adopted quantum computing f
arXiv:2602.01518v2 Announce Type: replace Abstract: Despite their importance in model sampling, efficient implementation of Top-k and Top-p algorithms for large vocabularies remains a significant chal
arXiv:2605.27068v1 Announce Type: cross Abstract: Social deduction games have become a popular testbed for probing reasoning, deception, coordination, and belief modeling in Large Language Model (LLM)
arXiv:2511.14683v2 Announce Type: replace Abstract: Heaps' or Herdan's law characterizes the word-type vs. word-token relation by a power-law function, which is concave in linear-linear scale but a st
arXiv:2605.26770v1 Announce Type: new Abstract: Prior work shows that Large Language Models (LLMs) can transform Explainable AI (XAI) outputs into Natural Language Explanations (NLEs) that score highl
arXiv:2605.26266v1 Announce Type: cross Abstract: Chunk-wise autoregressive video diffusion models rely on a KV cache of previously generated chunks to avoid redundant computation, but this cache quic
arXiv:2605.27164v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) systems for question answering typically retrieve evidence by semantic similarity between the query and document ch
arXiv:2003.05746v4 Announce Type: replace-cross Abstract: In this paper, we explore the issue of inconsistency handling over prioritized knowledge bases (KBs), which consist of an ontology, a set of f
arXiv:2605.26519v1 Announce Type: new Abstract: Recent feed-forward geometry foundation models have demonstrated impressive generalization by recovering depth and poses in a single forward pass. Howev
arXiv:2605.26328v1 Announce Type: new Abstract: Radars are an ideal complement to cameras: both are inexpensive, solid-state sensors, with cameras offering fine angular resolution, while radars provid
arXiv:2605.26819v1 Announce Type: cross Abstract: We present RAGEAR (Retrieval-Augmented Graph-Enhanced Academic Recommender), a neurosymbolic recommender system for academic course recommendation. RA
arXiv:2602.03517v2 Announce Type: replace Abstract: Many decision-making problems require ranking individuals by their treatment effects rather than estimating the exact effect magnitudes. Examples in
arXiv:2605.26127v1 Announce Type: cross Abstract: Purpose: Real-time (RT) bSSFP MRI enables fast free-breathing cardiovascular imaging but requires 10-16 slices for functional assessment, resulting in
arXiv:2605.26854v1 Announce Type: new Abstract: The scalable solution of large sparse linear systems is a bottleneck in scientific computing and graph analysis. While algebraic multigrid (AMG) offers
arXiv:2605.26784v1 Announce Type: cross Abstract: Standard on-policy reinforcement learning relies on heuristic clipping to enforce trust regions, but this mechanism imposes a severe cost by indiscrim
arXiv:2605.26348v1 Announce Type: new Abstract: Mobile robots can fail before they collide: a velocity that is safe now may commit the robot to a passage that moving obstacles will soon close. We stud
arXiv:2605.26513v1 Announce Type: new Abstract: Mean Deviation (MD) is a critical metric for assessing visual field loss in ophthalmology. While previous work has focused solely on predicting MD from
arXiv:2605.27315v1 Announce Type: new Abstract: Visual inputs are often assumed to improve language understanding in multimodal models. We examine this assumption by asking whether vision-language mod
arXiv:2506.23274v4 Announce Type: replace-cross Abstract: Recent reasoning language models, particularly those that employ long latent chains of thought, achieve strong performance on complex agentic
arXiv:2605.26414v1 Announce Type: new Abstract: Large Language Models (LLMs) achieve impressive accuracy on mathematical reasoning benchmarks, yet their performance drops when problems are modified wi
arXiv:2605.26934v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has become central to post-training reasoning models, yet a key limitation of existing studies i
arXiv:2605.27014v1 Announce Type: cross Abstract: Large Language Models (LLMs) have transformed artificial intelligence from primarily generative systems into increasingly capable reasoning agents. Re
arXiv:2605.26525v1 Announce Type: cross Abstract: Minute-scale cinematic video generation is a central challenge for generative video models. Existing paradigms address only fragments of this challeng
arXiv:2605.26855v1 Announce Type: new Abstract: Public datasets such as DLC-2021, SynID, and KID34K have significantly contributed to research on presentation attack detection for identity documents,
arXiv:2605.26969v1 Announce Type: cross Abstract: User modeling aims to use language models (LMs) to mimic an individual's behavior from a corpus of past context-action pairs (e.g., conversation turns
arXiv:2512.01572v3 Announce Type: replace-cross Abstract: Extreme sensor sparsity makes full-field reconstruction a fundamentally ill-posed problem in scientific sensing,where the goal is to infer phy
arXiv:2605.26535v1 Announce Type: cross Abstract: Generative models have emerged as a powerful paradigm for solving physics systems and modeling complex spatiotemporal dynamics. However, achieving hig
arXiv:1909.08210v4 Announce Type: replace Abstract: A restricted Boltzmann machine (RBM) is a two-layer neural network with shared weights and has been extensively studied for dimensionality reduction
arXiv:2605.26560v1 Announce Type: cross Abstract: Objective. Outpatient notes carry follow-up instructions pairing actions with future times ('MRI brain in two weeks'). Extracting (action, date) pairs
arXiv:2605.27081v1 Announce Type: cross Abstract: Fine-grained Mixture-of-Experts (MoE) models sparsely activate only a subset of experts per token, reducing activated computation while maintaining hi
arXiv:2605.26327v1 Announce Type: new Abstract: Shampoo-based methods, such as KL-Shampoo and SOAP, have demonstrated strong performance in training neural networks and rely on QR decomposition. Becau
arXiv:2605.26177v1 Announce Type: cross Abstract: Code agents are currently having skillful performance on repository-level software engineering benchmarks, but it remains unclear whether success on e
arXiv:2601.10566v5 Announce Type: replace Abstract: Entity-level unlearning is usually evaluated by what a model says: whether it stops naming the target, refuses a query, or shifts a Truth Ratio dist
arXiv:2605.24634v2 Announce Type: replace Abstract: Composed image retrieval (CIR) searches a corpus with a reference image and a text describing how to modify it. Despite rapid progress from triplet-
arXiv:2605.26661v1 Announce Type: cross Abstract: Out-of-distribution (OOD) detection has emerged as a popular technique to enhance the reliability of machine learning models by identifying unexpected
arXiv:2605.26735v1 Announce Type: new Abstract: Recent reasoning Large Language Models produce a chain-of-thought (CoT) predominantly in English, even when prompted in non-English languages. Prior wor
arXiv:2602.04879v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has become a cornerstone for fine-tuning Large Language Models (LLMs), with Proximal Policy Optimization (PPO) ser
arXiv:2605.26441v1 Announce Type: cross Abstract: This paper addresses the challenging task of weakly-supervised video temporal grounding. Existing approaches are generally based on the moment proposa
arXiv:2605.26921v1 Announce Type: new Abstract: The study of representations is widespread across fields, including neuroscience, psychology, and artificial intelligence. While representations are oft
arXiv:2604.04940v2 Announce Type: replace Abstract: Designing effective heuristics for NP-hard combinatorial optimization problems remains challenging and often requires substantial domain expertise.
arXiv:2605.26861v1 Announce Type: new Abstract: Image geo-localization aims to determine where a photograph was taken, a task that often requires more than recognizing visible landmarks. Human experts
arXiv:2605.26352v1 Announce Type: new Abstract: Retrieval is increasingly moving from one-shot matching toward interactive reasoning, where language agents iteratively inspect evidence, reformulate qu
arXiv:2605.27314v1 Announce Type: new Abstract: Reactive control is often considered insufficient for multi-objective tasks because conflicting objectives give rise to local minima. We argue this limi
arXiv:2605.27299v1 Announce Type: cross Abstract: Modern intrusion detection systems generate thousands of alerts daily, but alert fatigue severely limits security operations effectiveness due to too
arXiv:2605.26971v1 Announce Type: new Abstract: The proliferation of Reinforcement Learning from Verifiable Rewards (RLVR) datasets has exacerbated provenance collapse due to unclear lineage among exi
arXiv:2605.26862v1 Announce Type: new Abstract: Accurate road segmentation from aerial imagery is fundamental to many geospatial applications. However, existing datasets often suffer from limited scen
arXiv:2605.26452v1 Announce Type: cross Abstract: Safe reinforcement learning (RL) for robotic systems requires policies that improve task performance while satisfying state and input constraints duri
arXiv:2506.03627v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have demonstrated remarkable performance across various tasks by effectively utilizing a prompting strategy. Howe
arXiv:2605.26241v1 Announce Type: new Abstract: Success in generative modeling across language, image, and video demonstrates that large, well-curated datasets are the key driver for building capable
arXiv:2605.26702v1 Announce Type: cross Abstract: Reliable watermarking of panoramic imagery is fundamentally challenged by arbitrary 3D rotations. As panoramas are defined on the sphere, they natural