On an L^2 norm for stationary ARMA processes
arXiv:2408.10610v5 Announce Type: replace Abstract: We propose an L^2 norm for stationary Autoregressive Moving Average (ARMA) models. We look at ARMA models within the Hilbert space of the past with
Knowledge catalogue
arXiv:2408.10610v5 Announce Type: replace Abstract: We propose an L^2 norm for stationary Autoregressive Moving Average (ARMA) models. We look at ARMA models within the Hilbert space of the past with
arXiv:2604.13740v1 Announce Type: new Abstract: We propose a new partial-observability model for online learning problems where the learner, besides its own loss, also observes some noisy feedback abo
OpenAI has launched GPT-Rosalind, a biology-tuned large language model . This specialized LLM is designed to enhance performance on biology-specific tasks and applications. The model represents OpenAI
arXiv:2604.13672v1 Announce Type: new Abstract: The `spotoptim` package implements surrogate-model-based optimization of expensive black-box functions in Python. Building on two decades of Sequential
arXiv:2604.14010v1 Announce Type: cross Abstract: Supervised Fine-Tuning (SFT) of large language models often suffers from task interference and catastrophic forgetting. Recent approaches alleviate th
arXiv:2604.13791v1 Announce Type: new Abstract: Accurate lesion segmentation in ultrasound images is essential for preventive screening and clinical diagnosis, yet remains challenging due to low contr
arXiv:2604.13074v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) serve as daily assistants for millions. However, their ability to generate responses aligned with individual pr
arXiv:2508.00222v5 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Reward (RLVR) has significantly advanced the complex reasoning abilities of Large Language Models (LLMs
arXiv:2604.14148v1 Announce Type: new Abstract: Seedance 2.0 is a new native multi-modal audio-video generation model, officially released in China in early February 2026. Compared with its predecesso
arXiv:2505.24869v3 Announce Type: replace Abstract: Recent advances in test-time optimization have led to remarkable reasoning capabilities in Large Language Models (LLMs), enabling them to solve high
arXiv:2604.13710v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) exhibit strong reasoning and world knowledge, yet adapting them for retrieval remains challenging. Existing app
Sorry for the long wait, everyone! As I said, Qwen is going to keep open-sourcing! ⚡ Meet Qwen3.6-35B-A3B:Now Open-Source!🚀🚀 A sparse MoE model, 35B total params, 3B active. Apache 2.0 license. 🔥 Agen
arXiv:2510.03988v2 Announce Type: replace Abstract: Distilling long-form reasoning from teacher models into smaller students requires selecting which candidate solutions to train on. Recent work argue
arXiv:2602.21627v3 Announce Type: replace Abstract: This paper presents a new unified approach to semantic segmentation in both images and videos by using language modeling to output the masks as sequ
arXiv:2505.18232v3 Announce Type: replace-cross Abstract: The deployment of large language models (LLMs) is largely hindered by their large number of parameters. Structural pruning has emerged as a pr
arXiv:2604.13197v1 Announce Type: new Abstract: Process reward models (PRMs) provide fine-grained reward signals along the reasoning process, but training reliable PRMs often requires step annotations
April 2026 brings three major Reinforcement Fine-Tuning updates: Global Training for o4-mini with lower per-token rates across 12+ regions, new GPT-4.1 model graders for richer reward signals, and a c
arXiv:2604.13055v1 Announce Type: new Abstract: Today's evolving labor markets rely increasingly on recommender systems for hiring, talent management, and workforce analytics, with natural language pr
arXiv:2604.12904v1 Announce Type: new Abstract: Composed Image Retrieval (CIR) aims to retrieve a target image based on a query composed of a reference image, and a relative caption that specifies the
arXiv:2501.16154v4 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown impressive multilingual capabilities through pretraining on diverse corpora. Although these models sho
arXiv:2604.12875v1 Announce Type: new Abstract: The rapid expansion of large language model (LLM) safety evaluation has produced a substantial benchmark ecosystem, but not a correspondingly coherent m
arXiv:2604.12262v1 Announce Type: cross Abstract: Cascaded LLM systems coordinate models of varying sizes with human experts to balance accuracy, cost, and abstention under uncertainty. However, singl
arXiv:2604.12525v1 Announce Type: new Abstract: Recent advanced diffusion methods typically derive strong generative priors by scaling diffusion transformers. However, scaling fails to generalize when
arXiv:2604.12913v1 Announce Type: cross Abstract: Binary decompilation is a critical reverse engineering task aimed at reconstructing high-level source code from stripped executables. Although Large L
arXiv:2604.12268v1 Announce Type: cross Abstract: Large language models (LLMs) can generate code from natural language, but the extent to which they capture intended program behavior remains unclear.
arXiv:2603.01591v2 Announce Type: replace-cross Abstract: Training-free diffusion priors enable inverse-problem solvers without retraining, but for nonlinear forward operators data consistency often r
arXiv:2503.05167v3 Announce Type: replace Abstract: Traditional Chinese medicine (TCM) exhibits remarkable therapeutic efficacy in healthcare through patient-specific formulas. However, current AI-bas
arXiv:2604.12666v1 Announce Type: cross Abstract: Text-based web agents offer computational efficiency for autonomous web navigation, yet developing robust agents remains challenging due to the noisy
arXiv:2604.12351v1 Announce Type: new Abstract: Automated diagnosis based on color fundus photography is essential for large-scale glaucoma screening. However, existing deep learning models are typica
Gemini 3.1 Flash TTS Google released Gemini 3.1 Flash TTS today, a new text-to-speech model that can be directed using prompts. It's presented via the standard Gemini API using gemini-3.1-flash-tts-pr
Gemini 3.1 Flash TTS is rolling out in Google Vids and is available today in preview via the Gemini API and in @GoogleAIStudio. Whether you’re creating a pitch deck or recording a passion project, tra
arXiv:2604.12630v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have exhibited remarkable performance in various visual tasks, yet still struggle with spatial reasoning. Rec
arXiv:2510.07285v3 Announce Type: replace-cross Abstract: The escalating complexity of network threats and the inherent class imbalance in traffic data present formidable challenges for modern Intrusi
Have found the same things! Using glm-5 as a daily driver for a lot of things We've tested new OSS models the moment they're released for a while at Lindy. Inference is our #1 cost by a lot (more than
arXiv:2604.12573v1 Announce Type: new Abstract: Large Language Models are increasingly deployed for decision-making, yet their adoption in high-stakes domains remains limited by miscalibrated probabil
arXiv:2604.12337v1 Announce Type: new Abstract: Letters of recommendation (LoRs) can carry patterns of implicitly gendered language that can inadvertently influence downstream decisions, e.g. in hirin
arXiv:2508.01620v3 Announce Type: replace-cross Abstract: Machine Unlearning (MU) aims to selectively erase the influence of specific data points from pretrained models. However, most existing MU meth
arXiv:2509.03497v3 Announce Type: replace Abstract: Accurate global crop type mapping supports agricultural monitoring and food security, yet remains limited by the scarcity of labeled data in many re
arXiv:2604.11827v1 Announce Type: cross Abstract: Machine learning is revolutionizing chemistry. Beyond the value of predictive models accelerating virtual screening, generative AI aims at enabling in
arXiv:2604.12627v1 Announce Type: new Abstract: RLVR improves reasoning in large language models, but its effectiveness is often limited by severe reward sparsity on hard problems. Recent hint-based R
arXiv:2604.12049v1 Announce Type: cross Abstract: The use of Large Language Models (LLMs) for reliable, enterprise-grade analytics such as text categorization is often hindered by the stochastic natur
arXiv:2604.12601v1 Announce Type: cross Abstract: Passwords still remain a dominant authentication method, yet their security is routinely subverted by predictable user choices and large-scale credent
arXiv:2604.11868v1 Announce Type: new Abstract: While medical Vision-Language models (VLMs) achieve strong performance on tasks such as tumor or organ segmentation and diagnosis prediction, their opaq
arXiv:2603.00655v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable performance by aligning pretrained visual representations with the linguistic know
IBM Research has demonstrated that mid-training — a dedicated training phase between initial pre-training and fine-tuning — is essential for improving reasoning capabilities in large language models (
ComfyUI announced day-zero support for ERNIE Image, Baidu's image generation model, with full implementation details available on their official blog. The integration allows users to run ERNIE Image m
The Reddit post about NucleusAI/Nucleus-Image appears to be a community showcase of sample outputs from a text-to-image generative AI model called Nucleus-Image, developed by NucleusAI. The post is sh
arXiv:2604.12526v1 Announce Type: cross Abstract: Continual machine unlearning aims to remove the influence of data that should no longer be retained, while preserving the usefulness of the model on e
arXiv:2604.12103v1 Announce Type: cross Abstract: We present parameter-interpolated dynamic mode decomposition (piDMD), a parametric reduced-order modeling framework that embeds known parameter-affine
arXiv:2604.12171v1 Announce Type: cross Abstract: Pipeline parallelism (PP) is widely used to partition layers of large language models (LLMs) across GPUs, enabling scalable inference for large models
arXiv:2604.12357v1 Announce Type: new Abstract: Detailed image captioning demands both factual grounding and fine-grained coverage, yet existing methods have struggled to achieve them simultaneously.
arXiv:2602.15876v1 Announce Type: cross Abstract: Adapting instruction to the fine-grained needs of individual students is a powerful application of recent advances in large language models. These gen
arXiv:2506.14512v4 Announce Type: replace Abstract: Large Language Models (LLMs) have undergone rapid progress, largely attributed to reinforcement learning on complex reasoning tasks. In contrast, wh
arXiv:2604.12600v1 Announce Type: new Abstract: The core challenge of hyperspectral image denoising is striking the right balance between data fidelity and noise prior modeling. Most existing methods
arXiv:2604.12951v1 Announce Type: new Abstract: The most cited calibration result in deep learning -- post-temperature-scaling ECE of 0.012 on CIFAR-100 (Guo et al., 2017) -- is below the statistical
arXiv:2604.12231v1 Announce Type: new Abstract: Large language models (LLMs) have transformed AI research thanks to their powerful internal capabilities and knowledge. However, existing LLMs still fai
arXiv:2604.12346v1 Announce Type: new Abstract: Spatio-temporal video grounding (STVG) aims to localize queried objects within dynamic video segments. Prevailing fully-trained approaches are notorious
arXiv:2604.12887v1 Announce Type: new Abstract: Visual tokenizers map high-dimensional raw pixels into a compressed representation for downstream modeling. Beyond compression, tokenizers dictate what
arXiv:2507.09318v2 Announce Type: replace-cross Abstract: Generating spoken dialogue is inherently more complex than monologue text-to-speech (TTS), as it demands both realistic turn-taking and the ma
arXiv:2301.06081v4 Announce Type: replace-cross Abstract: In a variational denoising model, weight in the data fidelity term plays the role of enhancing the noise-removal capability. It is profoundly