IPR-1: Interactive Physical Reasoner
arXiv:2511.15407v3 Announce Type: replace Abstract: Humans learn by observing, interacting with environments, and internalizing physics and causality. Here, we aim to ask whether an agent can similarl
Knowledge catalogue
arXiv:2511.15407v3 Announce Type: replace Abstract: Humans learn by observing, interacting with environments, and internalizing physics and causality. Here, we aim to ask whether an agent can similarl
arXiv:2605.15184v1 Announce Type: new Abstract: Recent advances in Large Language Model (LLM) agents have enabled complex agentic workflows where models autonomously retrieve information, call tools,
arXiv:2512.12772v2 Announce Type: replace-cross Abstract: Understanding videos inherently requires reasoning over both visual and auditory information. To properly evaluate Omni-Large Language Models
arXiv:2605.14828v1 Announce Type: cross Abstract: Existing clustering methods for functional data often prioritize partitioning accuracy over interpretability, making it challenging to extract meaning
arXiv:2509.25826v3 Announce Type: replace Abstract: Inherent temporal heterogeneity, such as varying sampling densities and periodic structures, has posed substantial challenges in zero-shot generaliz
arXiv:2511.07686v2 Announce Type: replace-cross Abstract: Chemical Reaction Neural Networks (CRNNs) have emerged as an interpretable machine learning framework for discovering reaction kinetics direct
arXiv:2602.11534v3 Announce Type: replace-cross Abstract: Self-attention in Transformers relies on globally normalized softmax weights, causing all tokens to compete for influence at every layer. When
arXiv:2601.21349v2 Announce Type: replace-cross Abstract: Mixture-of-Experts (MoE) models scale neural networks by conditionally activating a small subset of experts, where the router plays a central
arXiv:2605.14241v1 Announce Type: new Abstract: Tool-augmented LLM agents increasingly access the same tool type through multiple functionally equivalent providers, such as web-search APIs, retrievers
arXiv:2605.14805v1 Announce Type: new Abstract: Accurate dynamics models are critical for aerial manipulators operating under complex tasks such as payload transport. However, modeling these systems r
arXiv:2605.14186v1 Announce Type: new Abstract: Large language models (LLMs) often expose useful signals of self-monitoring: before solving a problem, they can estimate whether they are likely to succ
arXiv:2604.05306v2 Announce Type: replace-cross Abstract: Large language models (LLMs) often produce confident yet incorrect answers, which can lead to risky failures in real-world applications. We st
arXiv:2605.14365v1 Announce Type: cross Abstract: Recent tabular learning benchmarks increasingly show a tight performance cluster rather than a clear hierarchy among leading methods, spanning gradien
arXiv:2508.06202v2 Announce Type: replace-cross Abstract: Continual Visual Instruction Tuning (CVIT) enables Multimodal Large Language Models (MLLMs) to incrementally learn new tasks over time. Howeve
arXiv:2605.14413v1 Announce Type: cross Abstract: Out-of-distribution (OOD) detection is a critical component for ensuring the reliability of deep neural networks in safety-critical applications. In t
arXiv:2605.14885v1 Announce Type: new Abstract: Scene Text Recognition requires modeling visual structures that evolve from coarse layouts to fine-grained character strokes. Training such models relie
arXiv:2605.14061v1 Announce Type: new Abstract: Current autoformalization benchmarks are largely focused on olympiad or undergraduate mathematics, while graduate and research-level mathematics remains
arXiv:2509.01416v2 Announce Type: replace Abstract: The computational overhead of traditional numerical solvers for partial differential equations (PDEs) remains a critical bottleneck for large-scale
arXiv:2605.13930v1 Announce Type: new Abstract: EEG foundation models achieve state-of-the-art clinical performance, yet the internal computations driving their predictions remain opaque: a barrier to
arXiv:2605.14843v1 Announce Type: new Abstract: Text- and image-conditioned video generation models have achieved strong visual fidelity and temporal coherence, but they often fail to generate motion
arXiv:2605.15128v1 Announce Type: cross Abstract: Long-term agent memory is increasingly multimodal, yet existing evaluations rarely test whether agents preserve the visual evidence needed for later r
arXiv:2605.14906v1 Announce Type: new Abstract: Memory is essential for large vision-language models (LVLMs) to handle long, multimodal interactions, with two method directions providing this capabili
arXiv:2605.06132v2 Announce Type: replace Abstract: In agent memory systems, the reranking model serves as the critical bridge connecting user queries with long-term memory. Most systems adopt the 're
arXiv:2605.13919v1 Announce Type: new Abstract: Multilingual knowledge editing (MKE) remains challenging because language-specific edits interfere with one another, even when locate-then-edit methods
arXiv:2605.14120v1 Announce Type: cross Abstract: Geospatial foundation models compress multispectral observations into dense embeddings increasingly used in natural-language environmental reasoning s
arXiv:2605.14292v1 Announce Type: cross Abstract: KV-cache compression at small budgets is a crowded design space spanning cache representation, head-wise routing, compression cadence, decoding behavi
arXiv:2605.14568v1 Announce Type: cross Abstract: Context. Behaviour-Driven Development (BDD) software test suites accumulate duplicated step subsequences. Three published refactoring patterns are ava
arXiv:2605.13849v1 Announce Type: new Abstract: Determining what to eat to satisfy nutritional requirements is one of the oldest optimization problems in operations research, yet existing formulations
arXiv:2605.15081v1 Announce Type: cross Abstract: The development of high-quality text embeddings is increasingly drifting toward an exclusionary future, defined by three critical barriers: prohibitiv
arXiv:2510.23477v2 Announce Type: replace Abstract: Effective math tutoring requires not only solving problems but also diagnosing students' difficulties and guiding them step by step. While multimoda
arXiv:2511.08565v3 Announce Type: replace-cross Abstract: Large language models (LLMs) increasingly operate in social contexts, motivating analysis of how they express and shift moral judgments. In th
arXiv:2605.13857v1 Announce Type: cross Abstract: The creation of cinematic-quality animal effects necessitates the precise modeling of muscle and fur dynamics, a process that remains both labor-inten
arXiv:2605.14635v1 Announce Type: cross Abstract: This paper introduces a multi-label visual emotion analysis benchmark dataset for comprehensively evaluating the ability of multimodal large language
arXiv:2602.21545v3 Announce Type: replace Abstract: Muon has recently emerged as a strong optimizer for large language model pre-training, orthogonalizing the momentum matrix via Newton--Schulz polar
arXiv:2605.14941v1 Announce Type: cross Abstract: Electroencephalogram (EEG) signals are highly susceptible to artifacts, resulting in a low signal-to-noise ratio which makes extraction of meaningful
arXiv:2603.29665v2 Announce Type: replace Abstract: Agentic systems for business process automation often require compliance with policies governing conditional updates to the system state. Evaluation
arXiv:2605.14048v1 Announce Type: new Abstract: Masked autoencoders (MAEs) have recently shown promise for self-supervised representation learning of resting-state brain functional connectivity (FC).
arXiv:2605.13988v1 Announce Type: new Abstract: Inverse problems in scientific sensing are often solved with either hand-designed regularizers or supervised networks trained on simulated labels, yet b
arXiv:2601.22197v3 Announce Type: replace-cross Abstract: Generating clinical reports that summarize abnormal patterns, diagnostic findings, and clinical interpretations from long-term EEG recordings
arXiv:2605.14698v1 Announce Type: cross Abstract: Foundation models (FMs) promise to extract unified representations that generalize across downstream tasks. They have emerged across fields, including
arXiv:2602.13770v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have demonstrated strong semantic reasoning across multimodal domains. However, their integration with graph-base
arXiv:2605.14381v1 Announce Type: cross Abstract: Recent advancements in generative AI facilitate large-scale synthetic data generation for model evaluation. However, without targeted approaches, thes
arXiv:2605.14938v1 Announce Type: cross Abstract: Continual learning in multimodal large language models (MLLMs) aims to sequentially acquire knowledge while mitigating catastrophic forgetting, yet ex
arXiv:2605.15071v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) are increasingly applied to cultural heritage materials, from digital archives to educational platforms. This work ident
OpenAI Group PBC today made its Codex programming assistant available on mobile devices. The service is accessible through ChatGPT’s iOS and Android clients. It’s rolling out about eight months after
Ivan Mehta / TechCrunch: OpenAI debuts personal finance tools for US ChatGPT Pro users, partnering with Plaid to give access to 12K+ financial institutions to analyze spending and more — On Friday, Op
arXiv:2605.15177v1 Announce Type: new Abstract: Test-time compute scaling is a primary axis for improving LLM reasoning. Existing methods primarily scale depth by extending a single reasoning trace. S
arXiv:2601.19924v2 Announce Type: replace-cross Abstract: We investigate the capabilities and scalability of Large Language Models (LLMs) in optimization modeling, a domain requiring structured reason
arXiv:2511.16964v2 Announce Type: replace-cross Abstract: Maximizing performance on available GPU hardware is an ongoing challenge for modern AI inference systems. Traditional approaches include writi
arXiv:2602.01359v2 Announce Type: replace-cross Abstract: Although recent studies on time-series anomaly detection have increasingly adopted ever-larger neural network architectures such as transforme
arXiv:2605.14055v1 Announce Type: cross Abstract: Parameter-Efficient Fine-Tuning (PEFT) is widely used for adapting Large Language Models (LLMs) for various tasks. Recently, there has been an increas
arXiv:2605.14779v1 Announce Type: new Abstract: We propose a model-free offline multi-step reinforcement learning (RL) algorithm, Conservative Peng's Q(lambda) (CPQL). Our algorithm adapts the Peng's
arXiv:2511.19289v2 Announce Type: replace-cross Abstract: Estimating quantum entropies and divergences is an important problem in quantum physics, information theory, and machine learning. Quantum neu
arXiv:2605.13889v1 Announce Type: cross Abstract: Stain variation across hospitals degrades histopathology models at deployment. Existing augmentation methods perturb color spaces with arbitrary hyper
arXiv:2605.14040v1 Announce Type: new Abstract: We audit the multimodal-physics evaluation pipeline end-to-end and document three undetected construction practices that distort how the field measures
arXiv:2605.14678v1 Announce Type: new Abstract: The rise of personal assistant agents, e.g., OpenClaw, highlights the growing potential of large language models to support users across everyday life a
arXiv:2603.23129v2 Announce Type: replace Abstract: Godel agent realize recursive self-improvement: an agent inspects its own policy and traces and then modifies that policy in a tested loop. We intro
arXiv:2605.14002v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) embedded in agentic frameworks have transformed information retrieval from static, long context question answering into op
arXiv:2605.14217v1 Announce Type: cross Abstract: Large language models can now be personalised efficiently at scale using parameter efficient finetuning methods (PEFTs), but serving user-specific PEF
arXiv:2603.04885v2 Announce Type: replace Abstract: Real-world dialogue usually unfolds as an infinite stream. It thus requires bounded-state memory mechanisms to operate within an infinite horizon. H