Towards Open World Sound Event Detection
arXiv:2605.03934v1 Announce Type: cross Abstract: Sound Event Detection (SED) plays a vital role in audio understanding, with applications in surveillance, smart cities, healthcare, and multimedia ind
Knowledge catalogue
arXiv:2605.03934v1 Announce Type: cross Abstract: Sound Event Detection (SED) plays a vital role in audio understanding, with applications in surveillance, smart cities, healthcare, and multimedia ind
arXiv:2605.05176v1 Announce Type: new Abstract: Pre-trained transformers are able to learn from examples provided as part of the prompt without any weight updates, a remarkable ability known as in-con
arXiv:2605.04819v1 Announce Type: new Abstract: Graph neural networks have been widely used in Boolean satisfiability (SAT) tasks to learn structural information from SAT formulas. The goal of these s
arXiv:2605.04527v1 Announce Type: new Abstract: We introduce a framework for learning latent representations of 4D objects which are descriptive, faithfully capturing object geometry and appearance; c
A designer shares their experience participating in a stop-motion animation exhibition curated by Iyuichi Itoh from the animation 'Nyacki,' displaying their match photography materials from a six-year
arXiv:2605.01341v1 Announce Type: cross Abstract: Given a knowledge base (KB) with a non-entailed fact, the ABox abduction problem asks for possible extensions of the KB that would entail this fact. T
Announcing the Lenny's Newsletter x @Replit Buildathon Build something awesome with my podcast and newsletter data using Replit, and win fabulous prizes: 🔸 1:1 career coaching session with @amasad 🔸 $
arXiv:2605.03053v1 Announce Type: new Abstract: Organoids are complex, three dimensional, self-organizing cell cultures which manifest organ-like features and represent a powerful platform for studyin
arXiv:2605.02202v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have achieved remarkable success in tasks such as image captioning and visual question answering (VQA). However, as their
arXiv:2512.17817v3 Announce Type: replace Abstract: While 3DGS has emerged as a high-fidelity scene representation, encoding rich, general-purpose features directly from its primitives remains under-e
arXiv:2605.02603v1 Announce Type: new Abstract: Automated planning traditionally assumes that all aspects of a planning task (initial state, goals, and available actions) are fully specified in advanc
arXiv:2605.02080v1 Announce Type: cross Abstract: Drawing on crip theory, this paper proposes cripping AI as a guiding framework to center lived disability experiences in AI research and development.
arXiv:2509.15692v2 Announce Type: replace-cross Abstract: Simultaneous speech-to-text translation (Simul-S2TT) aims to translate speech into target text in real time, outputting translations while rec
arXiv:2605.03560v1 Announce Type: new Abstract: With the increase of the Electronic Health Records (EHR) data, more and more researchers are developing machine learning models to learn from the medica
arXiv:2605.03790v1 Announce Type: new Abstract: With advances in multimodal research and deep learning, Multimodal Large Language Models (MLLMs) have emerged as a powerful paradigm for a wide range of
Every month for the last 6 years, I’ve gotten a 0.50 charge to my debit card from AWS for an account I don’t use or remember how to access. I have no idea why. I am willing to pay 0.6
arXiv:2605.03984v1 Announce Type: new Abstract: Sampling from unnormalized densities is analogous to the generative modeling problem, but the target distribution is defined by a known energy function
Jeremy Howard proposes three key enhancements for large language models: incorporating latent space diffusion-like reasoning processes, adding a persistent recurrent state mechanism, and implementing
arXiv:2605.03690v1 Announce Type: new Abstract: We present a method for finding hierarchy-aware embeddings of knowledge graphs (KGs) using graph neural networks (GNNs) enriched with a semantic loss de
ChatGPT learns about the world through training on diverse text data while implementing privacy protections that prevent it from storing personal information from individual conversations or retaining
I ❤️ NYC We're hosting two in-person events next Wednesday: 1️⃣ FinParse workshop: Build AI agents to extract and act over the most complex financial documents 2️⃣ AI Happy Hour with Tabs: drinks, con
I'm disappointed by repeatedly hearing that my colleagues at Anthropic believe they are the only ones who should be trusted with building AI. It is *very good* there are a diversity of people building
arXiv:2605.03437v1 Announce Type: new Abstract: Detecting anomalies from 3D point clouds has received increasing attention in the field of computer vision, with some group-based or point-based methods
arXiv:2605.03620v1 Announce Type: cross Abstract: Syndrome-based neural decoding (SBND) has emerged as a promising deep learning approach for soft-decision decoding of high-rate, short-length codes. H
LlamaIndex NYC takeover, 5/13 🗽 Our CEO Jerry Liu is in town. Two events, open to every NYC builder: 🛠️ FinParse Workshop — laptops out, hands-on with @jerryjliu0 → https://luma.com/updli8i6 🍕 AI Engi
arXiv:2605.02455v1 Announce Type: cross Abstract: State-of-the-art Large Language Models (LLMs) excel in code generation at the function level. However, the output quality significantly declines when
arXiv:2507.22268v3 Announce Type: replace-cross Abstract: We study the problem of inferring substitutable and complementary items, which underpins applications such as alternative and follow-up purcha
arXiv:2605.03799v1 Announce Type: new Abstract: This preprint presents a systematic, research-oriented practicum that guides the reader through the entire modern NLP pipeline: from tokenisation and ve
arXiv:2605.01797v1 Announce Type: new Abstract: Integration of Answer Set Programming (ASP) with neural networks has emerged as a promising tool in Neuro-symbolic AI. While existing approaches extend
arXiv:2602.00378v2 Announce Type: replace-cross Abstract: We present a method for parametrizing sub-grid processes in the Shallow Water equations. We define coarse variables and local spatial averages
arXiv:2601.20303v2 Announce Type: replace Abstract: Estimating object mass from visual input is challenging because mass depends jointly on geometric volume and material-dependent density, neither of
arXiv:2507.05561v2 Announce Type: replace Abstract: Humans can pursue a near-infinite variety of tasks, but typically can only pursue a small number at the same time. We hypothesize that humans levera
arXiv:2605.02917v1 Announce Type: new Abstract: Supervised deep learning models for automated CTG analysis are typically constrained by narrowly curated labelled datasets and limited patient cohorts,
arXiv:2605.03450v1 Announce Type: new Abstract: In studies of media coverage of extreme climate events, NLP methods have become indispensable for identifying relevant texts in large news databases. St
arXiv:2605.02819v1 Announce Type: new Abstract: Large language models excel at complex reasoning, yet evaluating their intermediate steps remains challenging. Although process reward models provide st
arXiv:2605.03245v1 Announce Type: cross Abstract: Image-based Joint-Embedding Predictive Architecture (I-JEPA) offers a promising approach to visual self-supervised learning through masked feature pre
arXiv:2505.16208v2 Announce Type: replace-cross Abstract: We apply Echo-State Networks to predict time series and statistical properties of the competitive Lotka-Volterra model in the chaotic regime.
arXiv:2604.25224v2 Announce Type: replace Abstract: LLM-based financial agents increasingly produce investment rationales before the outcomes needed to evaluate them are observable. This creates a del
arXiv:2605.00063v1 Announce Type: cross Abstract: Reasoning-Intensive Retrieval (RIR) targets retrieval settings where relevance is mediated by latent inferential links between a query and supporting
arXiv:2403.11782v5 Announce Type: replace Abstract: Preference modelling lies at the intersection of economics, decision theory, machine learning and statistics. By understanding individuals' preferen
arXiv:2605.00462v1 Announce Type: cross Abstract: Intelligence Processing Units (IPU) have proven useful for many AI applications. In this paper, we evaluate them within the emerging field of AI for s
arXiv:2605.02892v1 Announce Type: new Abstract: Personalized image completion aims to restore occluded regions in personal photos while preserving identity and appearance. Existing methods either rely
arXiv:2605.02259v1 Announce Type: new Abstract: Incomplete propagation data significantly hinders robust fake news detection. Recent approaches leverage large language models to simulate missing user
arXiv:2605.01217v1 Announce Type: new Abstract: Face Recognition systems are widely deployed in real-world applications, but they also raise privacy concerns due to unauthorized collection and misuse
arXiv:2401.04139v4 Announce Type: replace Abstract: Handling class imbalance remains a central challenge in machine learning, particularly in pattern recognition tasks where identifying rare but criti
arXiv:2601.08517v2 Announce Type: replace Abstract: Channel-configuration search, the optimization of layer specifications such as channel widths in deep neural networks, presents a combinatorial chal
Coding agents are accelerating different types of software work to different degrees. When we architect teams, understanding these distinctions helps us to have realistic expectations. Listing functio
arXiv:2605.01020v1 Announce Type: new Abstract: This paper proposes and evaluates a new performance estimation method that leverages continual learning (CL) algorithms to carry out sequential simulati
arXiv:2601.12964v2 Announce Type: replace Abstract: Self-supervised pretraining in remote sensing is mostly done using mid-spatial resolution (MR) image datasets due to their high availability. Given
arXiv:2605.01363v1 Announce Type: cross Abstract: Flavor-tagging calibrations are often provided either as scale factors measured at a finite set of working points or as binned corrections to a chosen
arXiv:2605.01858v1 Announce Type: new Abstract: Streaming video understanding requires processing unbounded video streams with limited memory and computation, posing two key challenges. First, continu
arXiv:2605.01236v1 Announce Type: new Abstract: Unified image restoration using a single model often faces task interference due to diverse degradations. To address this, we propose DACG-IR (Degradati
arXiv:2605.00850v1 Announce Type: cross Abstract: Foundation models (FMs) for the Earth system learn statistical relationships between physical variables across massive datasets to enable versatile do
arXiv:2605.01732v1 Announce Type: new Abstract: Large language models (LLMs) have achieved remarkable performance across diverse domains, yet their enormous computational and memory requirements hinde
arXiv:2510.02729v2 Announce Type: replace Abstract: Deep time series forecasting has emerged as a rapidly growing field in recent years. Despite the exponential growth of community interests, progress
arXiv:2605.01131v1 Announce Type: new Abstract: In continual reinforcement learning (CRL), good performance requires never-ending learning, acting, and exploration in a big, partially observable world
arXiv:2605.02279v1 Announce Type: cross Abstract: Riemannian geometry provides the fundamental framework for optimization on nonlinear spaces such as matrix manifolds, which arise in machine learning,
arXiv:2605.02222v1 Announce Type: cross Abstract: We present Orbit-Space Geometric Probability Paths (OGPP), a particle-native flow-matching framework for generative modeling of particle systems. OGPP
arXiv:2603.14259v2 Announce Type: replace-cross Abstract: Generative recommendation (GR) has shown strong potential for sequential recommendation in an end-to-end generation paradigm. However, existin
arXiv:2510.09880v2 Announce Type: replace Abstract: We propose scene-adaptive strategies to efficiently allocate representation capacity for generating immersive experiences of indoor environments fro