Model Merging: Foundations and Algorithms
arXiv:2605.01580v1 Announce Type: new Abstract: Modern deep learning usually treats models as separate artifacts: trained independently, specialized for particular purposes, and replaced when improved
Knowledge catalogue
arXiv:2605.01580v1 Announce Type: new Abstract: Modern deep learning usually treats models as separate artifacts: trained independently, specialized for particular purposes, and replaced when improved
arXiv:2506.05952v4 Announce Type: replace Abstract: Recent advances in transformer-based text-to-motion generation have led to impressive progress in synthesizing high-quality human motion. Neverthele
arXiv:2605.01822v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly being used to support scientific discovery. In chemistry, tasks such as reaction prediction and structure
arXiv:2605.02881v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models aim to provide a single generalist controller for robots, but today's systems fall short on the criteria that matter
arXiv:2605.02351v1 Announce Type: new Abstract: Molecular Vibe Coding, a paradigm where chemists interact with LLMs to generate executable programs for molecular tasks, has emerged as a flexible alter
arXiv:2507.14061v2 Announce Type: replace Abstract: What if a robot could rethink its own morphological representation to better meet the demands of diverse tasks? Most robotic systems today treat the
arXiv:2510.08804v3 Announce Type: replace Abstract: We present MOSAIC, a multi-agent Large Language Model (LLM) framework for solving challenging scientific coding tasks. Unlike general-purpose coding
arXiv:2605.02509v1 Announce Type: new Abstract: Continual learning systems face a fundamental tension between plasticity -- acquiring new knowledge -- and stability -- retaining prior knowledge. We in
arXiv:2605.02871v1 Announce Type: cross Abstract: Composite materials exhibit strongly hierarchical and anisotropic properties governed by coupled mechanisms spanning constituents, plies, laminates, s
arXiv:2605.01154v1 Announce Type: new Abstract: ARC-AGI-2 is a benchmark of human-intuitive visual puzzles that measures a machine's ability to generalize from limited examples, interpret symbolic mea
arXiv:2605.01687v1 Announce Type: new Abstract: We present MultiBreak, a scalable and diverse multi-turn jailbreak benchmark to evaluate large language model (LLM) safety. Multi-turn jailbreaks mimic
arXiv:2501.09775v3 Announce Type: replace Abstract: Multiple Choice Question (MCQ) tests are among the most used methods for evaluating large language models (LLMs). Besides checking the correctness o
arXiv:2605.00871v1 Announce Type: cross Abstract: State space models (SSMs) achieve linear-time complexity but struggle with multi-channel physiological signals due to three limitations: fixed kernels
arXiv:2605.01075v1 Announce Type: new Abstract: Propagation-based X-ray phase-contrast imaging (PBI) enables high-contrast visualization of lung structures and holds strong medical potential. However,
OpenAI introduced new advertising purchasing options for ChatGPT, expanding how businesses can buy ad placements within the platform. These new methods likely provide advertisers with additional flexi
arXiv:2605.00912v1 Announce Type: new Abstract: When humans play geolocation games such as GeoGuessr, they rely on concrete visual cues, such as road markings, vegetation, or architectural details, to
arXiv:2605.00877v1 Announce Type: cross Abstract: The vast and underexplored ocean plays a critical role in regulating global climate and supporting marine biodiversity, yet artificial intelligence ha
arXiv:2510.07731v3 Announce Type: replace-cross Abstract: Organic reaction mechanisms are the stepwise elementary reactions by which reactants form intermediates and products, and are fundamental to u
arXiv:2605.01638v1 Announce Type: new Abstract: Multimodal deepfakes are proliferating on social media and threaten authenticity, information integrity, and digital forensics. Existing benchmarks are
arXiv:2511.00510v2 Announce Type: replace Abstract: To address panoramic distortion, large search space, and identity ambiguity under a 360{eg} FoV, OmniTrack++ adopts a feedback-driven framework that
arXiv:2605.01357v1 Announce Type: new Abstract: Large Language Models (LLMs) excel at long-context understanding but exhibit significant limitations in long-form generation. Existing studies primarily
arXiv:2506.11499v2 Announce Type: replace Abstract: Multimodal chatbots have become one of the major topics for dialogue systems in both research community and industry. Recently, researchers have she
arXiv:2605.02141v1 Announce Type: new Abstract: Kullback-Leibler (KL) regularization is widely used in offline decision-making and offers several benefits, motivating recent work on the sample complex
Open models should compete on cost and specialization, not frontier benchmarks @natolambert puts it well: the right benchmark is savings in compute and time, especially for repetitive agent tasks deep
OpenAI's newest default model for ChatGPT might not make stuff up as much. Hallucinations have been an ongoing problem for AI models, but OpenAI says its new GPT-5.5 Instant model has 'significant imp
arXiv:2601.03267v2 Announce Type: replace Abstract: This is the system card published alongside the OpenAI GPT-5 launch, August 2025. GPT-5 is a unified system with a smart and fast model that answers
arXiv:2605.02714v1 Announce Type: new Abstract: The advent of foundation models has heralded a new era in medical artificial intelligence (AI), enabling the extraction of generalizable representations
arXiv:2605.01333v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have emerged as a promising paradigm for dental image analysis. However, their ability to capture the multi-lev
arXiv:2605.02537v1 Announce Type: new Abstract: Autonomous 3D indoor scene synthesis breaks down in non-convex rooms with tightly coupled spatial constraints. Data-driven generators lack topological p
arXiv:2511.21086v2 Announce Type: replace Abstract: Large language models must satisfy hard orthographic constraints during controlled text generation, yet systematic cross-family evaluation remains l
arXiv:2605.01358v1 Announce Type: new Abstract: Unsupervised Environment Design (UED) offers a promising paradigm for improving reinforcement learning generalization by adaptively shaping training env
arXiv:2605.02069v1 Announce Type: new Abstract: Many scoring applications require absolute predictions, while pairwise comparisons can provide a simpler learning objective. We present Pair2Score, a tw
arXiv:2605.01936v1 Announce Type: new Abstract: In sequential search, alternatives are tested until the true class is found. Standard proper scoring rules like log loss are local, ignoring the ranking
arXiv:2509.19202v2 Announce Type: replace-cross Abstract: We propose Parameter Space Analysis through Guided Visual Interpolations (ParamInter), a novel tool for high-dimensional input parameter space
arXiv:2605.02447v1 Announce Type: new Abstract: Multimodal sarcasm detection, which aims to precisely identify pragmatic incongruities between literal text and nonverbal cues, has gained substantial a
arXiv:2605.01945v1 Announce Type: new Abstract: Tandem mass spectrometry provides a high-throughput framework for identifying and quantifying proteins in complex biological samples. In computational p
arXiv:2605.02236v1 Announce Type: cross Abstract: Recursive language-model loops often settle into recognizable attractor-like patterns. The practical question is how much injected text is needed to m
arXiv:2605.00929v1 Announce Type: new Abstract: Multivariate time series anomaly detection in ICS has attracted growing attention due to the increasing threat of cyber-physical attacks on critical inf
arXiv:2605.02524v1 Announce Type: new Abstract: Physics-informed neural networks (PINNs) have recently emerged as a promising framework for integrating data-driven learning with physical knowledge. In
arXiv:2605.02168v1 Announce Type: cross Abstract: Language model (LM)-based agents have demonstrated promising capabilities in automating complex tasks from natural language instructions, yet they con
arXiv:2605.01759v1 Announce Type: new Abstract: Scene-level point cloud self-supervised learning (PC-SSL) has demonstrated potential in enhancing the generalization capability of 3D vision models. Des
arXiv:2605.00942v1 Announce Type: cross Abstract: Developing effective test cases capable of thoroughly exercising large-scale software systems is inherently difficult, especially if such systems have
arXiv:2605.01640v1 Announce Type: cross Abstract: Training compute is increasingly outpacing the availability of high-quality data. This shifts the central challenge from optimal compute allocation to
arXiv:2602.11543v2 Announce Type: replace Abstract: Pretraining large language models (LLMs) typically requires centralized clusters with thousands of high-memory GPUs (e.g., H100/A100). Recent decent
arXiv:2605.01625v1 Announce Type: new Abstract: Proteins are inherently multiscale physical systems whose functional properties emerge from coordinated structural organization across multiple spatial
arXiv:2605.01699v1 Announce Type: new Abstract: Recent attacks show that behavioural unlearning of large language models leaves internal traces recoverable by adversarial probes. We characterise where
arXiv:2605.02144v1 Announce Type: new Abstract: Self-attention in Transformers is typically implemented as softmax(QK^op/sqrt{d})V, where Q=XW_Q, K=XW_K, and V=XW_V are learned linear projections of t
arXiv:2605.01630v1 Announce Type: new Abstract: Rankings produced by holistic LLM-as-a-judge scoring are sensitive to the bias of the chosen judge model. We show that switching to binary rubric scorin
Proud of my analyst, Hermes This is what he has achieved after 1.5+ months of internship with me - Got me off X (doomscrolling cut from 2 hrs → 30 mins) - Practically paid for his salary via PMs b
arXiv:2605.01017v1 Announce Type: new Abstract: We introduce Xiaohongshu Social Comparison Reader Elicitation (XHS-SCoRE), a reader-grounded benchmark for detecting if a text-only Xiaohongshu (RedNote
The agentic era is here, and the public sector is at the forefront of leading this transformation.At Google Cloud Next ’26, it was clear that academia and public sector organizations are no longer jus
arXiv:2605.01467v1 Announce Type: cross Abstract: Tensor completion has emerged as a powerful framework for recovering missing data in multidimensional signals by exploiting low-rank tensor structures
arXiv:2605.01502v1 Announce Type: new Abstract: Epistemic uncertainty estimation is essential for identifying regions where deep learning system outputs may be unreliable. However, existing approaches
arXiv:2605.02184v1 Announce Type: new Abstract: Pansharpening aims to generate high-resolution multispectral (HRMS) images by fusing low-resolution multispectral (LRMS) and high-resolution panchromati
arXiv:2605.02003v1 Announce Type: new Abstract: Machine Learning (ML) has transformed many scientific fields, yet key applications still lack standardized benchmarks. Raman spectroscopy, a widely used
arXiv:2605.01991v1 Announce Type: cross Abstract: Learning, prediction, and compression are intimately connected: a model that accurately predicts the next symbol in a sequence can be coupled with a s
arXiv:2605.02552v1 Announce Type: new Abstract: Chemotherapy dose optimization can be formulated as a dynamic treatment regime, requiring sequential decisions under uncertainty that must balance tumor
arXiv:2605.02469v1 Announce Type: new Abstract: Online reinforcement learning with verifiable rewards (RLVR) turns checkable outcomes into a scalable training signal, but it keeps rollout generation,
arXiv:2605.01913v1 Announce Type: cross Abstract: Fine-tuning safety-aligned language models for downstream tasks often leads to substantial degradation of refusal behavior, making models vulnerable t
arXiv:2605.02801v1 Announce Type: new Abstract: As large language model (LLM) agents evolve from isolated tool users into coordinated teams, reinforcement learning (RL) must optimize not only individu