Explanations for Automatic Speech Recognition
arXiv:2302.14062v2 Announce Type: replace-cross Abstract: We address quality assessment for neural network based ASR by providing explanations that help increase our understanding of the system and ul
Knowledge catalogue
arXiv:2302.14062v2 Announce Type: replace-cross Abstract: We address quality assessment for neural network based ASR by providing explanations that help increase our understanding of the system and ul
arXiv:2606.22510v1 Announce Type: new Abstract: While federated learning enables collaborative modelling on decentralised data, standard methods merely fit historical observations. This purely observa
arXiv:2606.22466v1 Announce Type: new Abstract: Federated learning (FL) is an emerging distributed machine learning paradigm that enables local devices to jointly train a global model while keeping da
arXiv:2507.18219v3 Announce Type: replace Abstract: Federated Graph Learning (FGL) is a distributed learning paradigm that enables collaborative training over large-scale subgraphs located on multiple
arXiv:2606.21373v1 Announce Type: new Abstract: Recent indoor occupancy prediction methods adopt Gaussian primitives as a sparse 3D representation for computational efficiency. However, their training
arXiv:2606.23293v1 Announce Type: new Abstract: 6D pose estimation is a key task in computer vision and embodied AI, widely used in robotic manipulation, augmented reality, etc. Existing methods direc
arXiv:2603.06467v2 Announce Type: replace Abstract: Radiology foundation models (RFMs) have largely inherited the scale-first recipe of natural-image vision--language pre-training. This recipe is diff
arXiv:2606.22834v1 Announce Type: new Abstract: We present homographic navigation, a geometry-centric framework for guiding camera acquisition toward precise capture of planar regions. Rather than tre
arXiv:2606.19538v2 Announce Type: replace-cross Abstract: Convolutional networks, recurrent networks, and transformers each encode different inductive biases -- locality, sequential memory, and conten
arXiv:2509.22307v2 Announce Type: replace Abstract: Lightweight 3D medical image segmentation remains constrained by a fundamental extit{``efficiency / robustness conflict''}, particularly when proces
arXiv:2603.05663v2 Announce Type: replace Abstract: Video Temporal Grounding (VTG) localizes the temporal boundaries of query-relevant moments in long videos, making video-language-model prohibitively
Krea 2 is back and this time the weights are OPEN. This open source release ships with two models designed to work together — @krea_ai 2 RAW and Krea 2 Turbo. RAW is for training, Turbo is for inferen
Krea x Comfy: Founders Live A special live conversation with our guests Victor Perez (CEO, Krea), Miguel Lara (Krea Team), and ComfyAnonymous (Co-Founder, Comfy Org), hosted by Purz & Julien. Today at
This appears to be a live broadcast event featuring the founders of Krea and Comfy discussing their projects and potentially their collaboration or integration. The session likely covers updates on Co
arXiv:2606.20627v1 Announce Type: cross Abstract: Planning with world models is bottlenecked by compounding prediction errors and the difficulty of defining optimizable goals. Visual targets provide p
Live Stream: Welcome to open source AI Lots of new folk are starting out on their journey with open models. Come join our livestream with all your questions about local models, open coding agents, and
arXiv:2606.21821v1 Announce Type: new Abstract: Understanding the causal structure of a language model's thought process is a problem of significant importance for both transparency and safety. In thi
arXiv:2410.02548v4 Announce Type: replace-cross Abstract: Flow Matching (FM) is a simulation-free method for learning a continuous, invertible flow that interpolates between two distributions, and in
arXiv:2606.21968v1 Announce Type: new Abstract: Vision-Language Models (VLMs) struggle as query-relevant objects become smaller. To address this, recent training-free approaches dynamically retrieve a
arXiv:2606.23126v1 Announce Type: new Abstract: While recent advancements in anomaly detection have demonstrated the efficacy of CNN- and Transformer-based approaches, these architectures face inheren
arXiv:2511.11625v2 Announce Type: replace Abstract: Artificial intelligence (AI) has shown great potential in medical imaging, particularly for brain tumor detection using Magnetic Resonance Imaging (
arXiv:2606.21344v1 Announce Type: cross Abstract: Trajectory prediction allows autonomous vehicles to anticipate the future behavior of surrounding objects (or agents) and, accordingly, maximize the s
arXiv:2606.20717v1 Announce Type: new Abstract: Multimodal Large Language Model (MLLM)-based web agents provide practical, high-precision solutions for visual browser automation; however, they inheren
arXiv:2603.04035v4 Announce Type: replace Abstract: Dimensionality reduction is a foundational tool for visualizing high-dimensional data, yet its reference implementations span a fragmented stack of
arXiv:2606.22702v1 Announce Type: new Abstract: Traditional supervised methods for structured visual recognition tasks -- such as object detection, segmentation, and scene graph generation -- often pr
arXiv:2606.22335v1 Announce Type: new Abstract: Researchers tag and track marine animals to study migration patterns, human impacts on behavior, and behavioral shifts due to climate change. Accurate d
arXiv:2509.18671v2 Announce Type: replace Abstract: Determining where to execute the manipulation policy is a fundamental challenge in mobile manipulation. Most approaches have formulated this as a ge
arXiv:2606.20823v1 Announce Type: new Abstract: Facial landmark localisation is a prerequisite for developing automated, non-contact neonatal pain assessment methods. Clinicians use pain scales to jud
arXiv:2503.10251v2 Announce Type: replace-cross Abstract: Transformers are the state-of-the-art architecture for large language models, and a key to their scalability is the strategic usage of low-pre
arXiv:2606.22002v1 Announce Type: new Abstract: Training medical image classifiers on entire datasets is wasteful when annotation budgets are limited: not all samples contribute equally, yet acquiring
arXiv:2512.03719v2 Announce Type: replace-cross Abstract: Over-the-Air Federated Learning (AirFL) is an emerging paradigm that tightly integrates wireless signal processing and distributed machine lea
arXiv:2606.23256v1 Announce Type: new Abstract: The increasing maturity of embodied AI platforms has driven a growing interest in procedural video representation learning to support intelligent assist
arXiv:2606.22084v1 Announce Type: cross Abstract: Characterizing the complete wall-pressure spectrum in turbulent wall-bounded flows requires simultaneous access to the viscous-scale high-wavenumber c
arXiv:2606.20754v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have shown strong performance in robotic manipulation, but reliable uncertainty quantification remains challenging,
arXiv:2606.21896v1 Announce Type: cross Abstract: Solar flares, particularly those of the M- and X-class, have a significant impact on human life because of their potential to disrupt critical infrast
arXiv:2603.24196v2 Announce Type: replace-cross Abstract: Neural Physics recasts local discretisations of partial differential equations (PDEs) as fixed convolutional operators, providing a physics-pr
Krea 2 is a technical advancement in AI image generation with newly released model weights available for download on GitHub. The technical report details the improvements and capabilities of this vers
arXiv:2606.20638v1 Announce Type: cross Abstract: Large language models are increasingly deployed as long-lived agents that must adapt across users, tasks, domains, modalities, and feedback regimes wi
arXiv:2606.21108v1 Announce Type: new Abstract: Image forgery localization remains challenging due to diverse manipulation techniques and distribution shifts. Existing forgery localization models achi
arXiv:2606.20946v1 Announce Type: cross Abstract: Developing robust models for 3D visual grounding (3DVG), the localization of entities in a 3D scene described in natural language, is important for en
arXiv:2606.22700v1 Announce Type: cross Abstract: Federated Learning (FL) enables collaborative model training without sharing raw data, making it a promising paradigm for privacy-sensitive applicatio
arXiv:2606.21138v1 Announce Type: new Abstract: AI-assisted image editing threatens trust in financial, legal, and identity records. The GenText-Forensics Challenge at ACM MM 2026 addresses this by re
arXiv:2606.13589v2 Announce Type: replace Abstract: We present Simplex-Constrained Sparse Bagging (SCSB), a mathematically rigorous framework for post-training compression and probability calibration
Eleanor Olcott / Financial Times: Sources: prices for Nvidia's AI chips on China's black market have more than doubled amid a US export crackdown; its flagship DGX B300 server has risen to $1.1M — US
arXiv:2606.20244v2 Announce Type: replace Abstract: Vision-language models (VLMs) often underperform on evidence intensive tasks because decisive visual evidence are small, localized, and easy to over
arXiv:2606.23254v1 Announce Type: new Abstract: Visual text editing aims to precisely modify text in images and videos while preserving stylistic consistency and visual realism. Despite significant ad
arXiv:2606.22724v1 Announce Type: new Abstract: Federated low-rank adaptation methods are attractive for fine-tuning large models under communication and privacy constraints, but heterogeneous client
arXiv:2606.23548v1 Announce Type: cross Abstract: This paper presents SuperCond-GNN, a graph neural network-based surrogate model for predicting the voltage distribution in high-temperature supercondu
arXiv:2606.20647v1 Announce Type: new Abstract: This paper presents Tessellated Biomes, a cyber-physical framework for the adaptive robotic construction and reconfiguration of modular multi-material a
arXiv:2606.22333v1 Announce Type: cross Abstract: A-mode ultrasound (US) has emerged as a promising modality for hand and wrist motion tracking. Prior works have mainly addressed static gesture classi
arXiv:2606.22527v1 Announce Type: new Abstract: Diffusion and flow-based generative models produce strong images, yet their controllability remains largely endpoint-centric: users specify conditions a
arXiv:2403.17407v4 Announce Type: replace-cross Abstract: Accurate transcription of Bengali text to the International Phonetic Alphabet (IPA) is a challenging task due to the complex phonology of the
arXiv:2606.20131v2 Announce Type: replace Abstract: We present TriFlow, a new generative approach for producing compact 3D meshes with artist-like triangle topology directly from input geometry condit
arXiv:2509.14641v2 Announce Type: replace Abstract: Dense 3D convolutions provide high accuracy for perception but are too computationally expensive for real-time robotic systems. Existing tri-plane m
arXiv:2606.20905v1 Announce Type: new Abstract: Robots operating in open-world environments must seamlessly integrate localization, spatial reasoning, navigation, and long-horizon planning. While spec
arXiv:2606.20814v1 Announce Type: cross Abstract: Emergent misalignment (EM) is a phenomenon in which models generalize with narrow fine-tuning, leading to broad (yet uneven) misalignment across evalu
arXiv:2606.21868v1 Announce Type: new Abstract: Modern Mixture-of-Experts (MoE) models place most of their parameters in expert layers, yet only a small fraction of those experts are used for any toke
arXiv:2106.06998v4 Announce Type: replace Abstract: Training convolutional neural networks at scale demands substantial memory, largely because intermediate activations must be stored for backpropagat
B9755 is a release of llama.cpp, a tool for LLM inference in C/C++ . The release represents a specific build version in the active development of the project, continuing the iterative improvements to
Release b9756 fixes a crash in the server's edit_file function when appending at the end of a file, addressing a heap-buffer-overflow caused by improper handling of line_start -1. The fix normalizes t