Do LLMs have core beliefs?
arXiv:2605.03255v1 Announce Type: new Abstract: The rise of Large Language Models (LLMs) has sparked debate about whether these systems exhibit human-level cognition. In this debate, little attention
Knowledge catalogue
arXiv:2605.03255v1 Announce Type: new Abstract: The rise of Large Language Models (LLMs) has sparked debate about whether these systems exhibit human-level cognition. In this debate, little attention
arXiv:2605.03281v1 Announce Type: cross Abstract: Donor-level disease classification from single-cell RNA sequencing (scRNA-seq) requires strict donor-aware cross-validation: naive pipelines that spli
arXiv:2511.04855v2 Announce Type: replace Abstract: In high-stakes applications, predictive models must not only produce accurate predictions but also quantify and communicate their uncertainty. Rejec
arXiv:2605.03662v1 Announce Type: new Abstract: In this work, a novel method for planar task and motion planning based on hybrid modeling is proposed. By virtue of a discrete variable which models loc
arXiv:2605.02827v1 Announce Type: new Abstract: Probabilistic values, including Shapley values and semivalues, provide a model-agnostic framework to attribute the behavior of a black-box model to data
arXiv:2605.02939v1 Announce Type: new Abstract: Multimodal controversy detection (MCD) identifies controversial content in videos and their associated user comments, to support risk management for soc
arXiv:2605.03388v1 Announce Type: new Abstract: Regulatory frameworks such as GDPR increasingly require that ML predictions be accompanied by post-hoc explanations, even when raw data and trained mode
arXiv:2605.01847v1 Announce Type: new Abstract: Outcome-only evaluation under-specifies whether an evaluated agent profile preserves the commitments required to solve a multi-turn task coherently. Neu
NEW paper from Microsoft Research. (bookmark it) The entire interpretability literature is built around human readers. As more analysis gets delegated to agents, the right target of interpretability s
arXiv:2505.08203v2 Announce Type: replace-cross Abstract: While recent advancements in AI music generation have predominantly focused on direct audio synthesis, these systems suffer from inherent rigi
arXiv:2605.03290v1 Announce Type: new Abstract: Domain randomization (DR) is widely used in policy learning to improve robustness to modeling error, but remains underexplored in contact-rich sampling-
arXiv:2605.03160v1 Announce Type: new Abstract: The standard sparse-autoencoder (SAE) interpretability protocol labels each feature from its top-activating contexts and validates by single-feature ste
arXiv:2605.02240v1 Announce Type: new Abstract: We introduce PhysicianBench, a benchmark for evaluating LLM agents on physician tasks grounded in real clinical setting within electronic health record
arXiv:2511.09290v2 Announce Type: replace Abstract: Predictive learning has emerged as a central paradigm for training models across diverse data domains and is increasingly viewed as a foundation for
arXiv:2605.03344v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) has proven effective for knowledge-intensive tasks, but is widely believed to offer limited benefit for reasoning
arXiv:2605.03680v1 Announce Type: new Abstract: While deep-learning-based image restoration has achieved unprecedented fidelity, deployment on mobile Neural Processing Units (NPUs) remains bottlenecke
arXiv:2605.03720v1 Announce Type: new Abstract: Recent advances in Large Reasoning Models (LRMs) trained with Long Chain-of-Thought have demonstrated remarkable capabilities in code generation and mat
arXiv:2601.04809v5 Announce Type: replace Abstract: Reinforcement learning (RL) offers a principled way to enhance the reasoning capabilities of large language models, yet its effectiveness hinges on
arXiv:2605.02617v2 Announce Type: new Abstract: Capturing semantic consistency among nodes is crucial for effective graph representation learning. Existing approaches typically rely on k-nearest neigh
Mixture-of-Experts (MoE) models enable sparse expert activation, meaning that only a subset of the model’s parameters is used during each inference. However, to translate this sparsity into practical
arXiv:2605.03909v1 Announce Type: cross Abstract: Robotic laser profiling is widely used for dimensional verification and surface inspection, yet measurement fidelity is often dominated by sensor conf
arXiv:2605.01160v1 Announce Type: cross Abstract: Since 2022, AI-powered coding assistants have produced contradictory evidence: controlled studies report 20-56% productivity gains on well-scoped task
arXiv:2605.01809v1 Announce Type: cross Abstract: Unified audio-visual generation is rapidly gaining industrial and creative relevance, enabling applications in virtual production and interactive medi
arXiv:2604.25000v2 Announce Type: replace Abstract: Recent work has framed intelligence in verifiable tasks as reducing time-to-solution through learned structure and test-time search, while systems w
arXiv:2605.01970v2 Announce Type: cross Abstract: Memory systems enable otherwise-stateless LLM agents to persist user information across sessions, but also introduce a new attack surface. We characte
arXiv:2605.03677v1 Announce Type: new Abstract: On-policy distillation (OPD) has recently emerged as an effective post-training paradigm for consolidating the capabilities of specialized expert models
arXiv:2605.00861v1 Announce Type: cross Abstract: This study investigates voice mapping as an evaluation framework for text-to-speech (TTS) synthesis quality. The study analyzes six TTS models, includ
arXiv:2605.02920v1 Announce Type: cross Abstract: Standard transformer architectures learn fixed slow-weight representations during training and lack mechanisms for rapid adaptation within an episode.
arXiv:2605.00881v1 Announce Type: cross Abstract: Speckle noise severely limits the quality of images acquired from coherent imaging systems such as Synthetic Aperture Radar (SAR) and medical ultrasou
arXiv:2403.11782v5 Announce Type: replace Abstract: Preference modelling lies at the intersection of economics, decision theory, machine learning and statistics. By understanding individuals' preferen
arXiv:2605.01477v1 Announce Type: new Abstract: We present Action Agent, a two-stage framework that unifies agentic navigation video generation with flow-constrained diffusion control for multi-embodi
arXiv:2605.01355v1 Announce Type: new Abstract: Automated leaf disease classification is critical for early disease detection in resource-constrained field environments. Vision Transformers (ViTs) pro
arXiv:2511.21075v3 Announce Type: replace Abstract: Engineering LLMs to accelerate life sciences research requires a robust alignment with biomedical knowledge. We observe that biomedical text exhibit
arXiv:2605.01098v1 Announce Type: cross Abstract: Adversarial examples in machine learning are typically generated using gradients, obtained either directly through access to the model or approximated
arXiv:2605.01368v1 Announce Type: new Abstract: Human-robot interaction (HRI) has long studied how agents and people coordinate to achieve shared goals. In this work, we formalize and benchmark the no
arXiv:2605.02402v1 Announce Type: new Abstract: Reflective thinking is a key competency in education, but assessing reflective writing remains a time-consuming and subjective task for education expert
arXiv:2503.07557v2 Announce Type: replace Abstract: We present a novel method, AutoSpatial, an efficient approach with structured spatial grounding to enhance VLMs' spatial reasoning. By combining min
arXiv:2605.01165v1 Announce Type: new Abstract: This paper proposes a novel Zero-Shot Action Recognition~(ZSAR) method based on contrastive learning. In ZSAR, we aim to classify examples from classes
arXiv:2605.02170v1 Announce Type: new Abstract: In this paper, we present our system for SemEval-2026 Task 6 (CLARITY) on response clarity and evasion detection in question-answer pairs from U.S. pres
arXiv:2605.01679v1 Announce Type: cross Abstract: Fall detection is a critical task in healthcare, particularly for elderly people. Timely fall detection and treatment can prevent severe injuries. Sen
arXiv:2605.01712v1 Announce Type: new Abstract: Pareto set learning (PSL) is an emerging paradigm in multi-objective optimization that trains neural networks to map preference vectors to Pareto optima
arXiv:2605.01609v1 Announce Type: new Abstract: We test whether the causal inner product of itet{park2024linear} -- defined by the unembedding covariance Sigma -- enables cross-lingual concept transpo
arXiv:2505.20948v3 Announce Type: replace Abstract: Abductive reasoning in knowledge graphs aims to generate plausible logical hypotheses from observed entities, with broad applications in areas such
arXiv:2605.01448v1 Announce Type: cross Abstract: Cross-task generalization is a core challenge in open-world robotic manipulation, and the key lies in extracting transferable manipulation knowledge f
arXiv:2605.02245v1 Announce Type: new Abstract: Automated sleep stage classification typically employs a single population-agnostic model, disregarding established demographic variations in sleep arch
arXiv:2007.02392v3 Announce Type: replace Abstract: We study the problem of estimating the parameters of a Boolean product distribution in d dimensions, when the samples are truncated by a set S subse
arXiv:2605.01732v1 Announce Type: new Abstract: Large language models (LLMs) have achieved remarkable performance across diverse domains, yet their enormous computational and memory requirements hinde
arXiv:2510.02729v2 Announce Type: replace Abstract: Deep time series forecasting has emerged as a rapidly growing field in recent years. Despite the exponential growth of community interests, progress
arXiv:2605.02705v1 Announce Type: new Abstract: Mobile crowdsensing (MCS) is a distributed sensing architecture that utilizes existing sensors on mobile units (MUs) to perform sensing tasks. A mobile
arXiv:2605.02764v1 Announce Type: new Abstract: We present FoR-Net, a lightweight architecture for semantic segmentation that focuses on identifying and enhancing hard regions. Instead of relying on h
arXiv:2605.02130v1 Announce Type: new Abstract: Human-level agentic intelligence extends beyond low-level geometric perception, evolving from recognizing where things are to understanding what they ar
arXiv:2605.01272v1 Announce Type: new Abstract: The development of video game streaming has grown rapidly, with major platforms such as YouTube and Twitch using different codecs. To support quality as
I detected a bad Agent action, what do I do about it? this is pretty much the main question that will power the future’s Human+Agent driven improvement loops Gather data -> Mine Errors -> Find out whi
arXiv:2505.10887v3 Announce Type: replace Abstract: This paper introduces extsc{InfantAgent-Next}, a generalist agent capable of interacting with computers in a multimodal manner, encompassing text, i
arXiv:2506.24056v2 Announce Type: replace-cross Abstract: RLHF-style alignment trains language models to refuse unsafe requests, but how much operational margin does this refusal rest on? We introduce
arXiv:2605.02689v1 Announce Type: new Abstract: Long-term time series forecasting requires models that simultaneously capture rapid oscillations, medium-range periodicities, and slowly evolving macro-
arXiv:2503.14505v3 Announce Type: replace Abstract: We introduce MusicInfuser, an approach that aligns pre-trained text-to-video diffusion models to generate high-quality dance videos synchronized wit
arXiv:2605.01749v1 Announce Type: new Abstract: Large Reasoning Models achieve strong performance on complex tasks but remain prone to hallucinations, particularly in long-form generation where errors
arXiv:2605.01625v1 Announce Type: new Abstract: Proteins are inherently multiscale physical systems whose functional properties emerge from coordinated structural organization across multiple spatial
arXiv:2601.16873v2 Announce Type: replace Abstract: We study the problem of learning Transformer-based sequence models with black-box access to their outputs. In this setting, a learner may adaptively