Online Learning-to-Defer with Varying Experts
arXiv:2605.12340v1 Announce Type: cross Abstract: Learning-to-Defer (L2D) methods route each query either to a predictive model or to external experts. While existing work studies this problem in batc
Knowledge catalogue
arXiv:2605.12340v1 Announce Type: cross Abstract: Learning-to-Defer (L2D) methods route each query either to a predictive model or to external experts. While existing work studies this problem in batc
Open Source always wins 👀 Save your credits and use open source models when you can. The new LTX 2.3 lipdub LoRA paired with Chatterbox TTS voice cloning model is the best workflow for lip syncing, ch
OpenAI Global Affairs: OpenAI endorses the Kids Online Safety Act and Illinois SB 315, an AI safety bill to create requirements around transparency, incident reporting, and more — Welcome (back) to Th
Maggie Eastland / Bloomberg: OpenAI says it would support a global AI governance body that is led by the US and includes China as a member, similar to the International Atomic Energy Agency — OpenAI w
Jared Perlo / NBC News: OpenEvidence, an AI-powered medical search tool, says it's used by two-thirds of US physicians, or ~650K doctors, and an additional 1.2M doctors internationally — OpenEvidence,
arXiv:2605.11199v1 Announce Type: cross Abstract: Trained lattice samplers are usually judged by the ensembles they generate. Here we instead analyze the trained field-space function itself: a flow-ma
arXiv:2605.12235v1 Announce Type: cross Abstract: We study optimal policy learning under combined budget and minimum coverage constraints. We show that the problem admits a knapsack-type structure and
arXiv:2605.11291v1 Announce Type: new Abstract: In this paper, we provide a computable characterization of the geometry of optimal representations in Contrastive Learning (CL) when the classes are imb
arXiv:2605.11172v1 Announce Type: new Abstract: We introduce SODA, a generalization of Optimistic Dual Averaging, which provides a common perspective on state-of-the-art optimizers like Muon, Lion, Ad
arXiv:2605.11977v1 Announce Type: new Abstract: We present a unified framework for 3D geometric abstraction using a single continuous 4D wire, parameterized as a B-spline with spatial coordinates and
arXiv:2605.12419v1 Announce Type: new Abstract: Despite the rapid advancements in large language model (LLM) development, fine-tuning them for specific tasks often results in the catastrophic forgetti
arXiv:2605.12446v1 Announce Type: cross Abstract: Large language models (LLMs) often produce answers with high certainty even when they are incorrect, making reliable confidence estimation essential f
arXiv:2602.12139v2 Announce Type: replace Abstract: Transformers excel at time series modelling through attention mechanisms that capture long-term temporal patterns. However, they assume uniform time
arXiv:2605.11803v1 Announce Type: new Abstract: As Video Large Language Models (Video-LLMs) scale to longer and more complex videos, their inference cost grows rapidly due to the large volume of visua
arXiv:2605.11570v1 Announce Type: new Abstract: Activation functions are what make deep networks expressive: without them, the model collapses to a linear map. Yet we still evaluate training mostly fr
At the Android Show yesterday, we introduced Googlebooks: a new category of premium laptops built with Gemini’s helpfulness at the core. Designed for Gemini Intelligence, Googlebooks will give persona
Our evaluations show that frontier AI's cyber capabilities are advancing quickly. The length of cyber tasks frontier models can complete has been doubling every few months, and this rate has become fa
OpenAI details its response to the TanStack “Mini Shai-Hulud” supply chain attack, outlines protections taken to secure systems and signing certificates, and explains why macOS users must update OpenA
arXiv:2605.12345v1 Announce Type: new Abstract: Parameter-efficient fine-tuning (PEFT) techniques offer task-specific fine-tuning at a fraction of the cost of full fine-tuning, but require separate fi
arXiv:2605.11459v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models achieve remarkable flexibility and generalization beyond classical control paradigms. However, most prevailing VLA
arXiv:2605.11525v1 Announce Type: new Abstract: Missing values are routinely treated as defects to be eliminated through deletion or imputation prior to machine learning. In many applied domains, howe
arXiv:2506.23619v2 Announce Type: replace-cross Abstract: This paper investigates the impact of posterior drift on out-of-sample forecasting accuracy in overparametrized machine learning models. We do
arXiv:2605.11178v1 Announce Type: new Abstract: Neural Sheaf Diffusion (NSD) generalizes diffusion-based Graph Neural Networks by replacing scalar graph Laplacians with sheaf Laplacians whose learned
arXiv:2605.12199v1 Announce Type: new Abstract: Emergent misalignment (EM), where fine-tuning on a narrow task (like insecure code) causes broad misalignment across unrelated domains, was first demons
arXiv:2605.12313v1 Announce Type: new Abstract: Multi-hop question answering (QA) remains a significant challenge in the biomedical domain, requiring systems to integrate information across multiple s
A peer-to-peer AI inference mesh is a decentralized network architecture where agents connect directly to discover peers and communicate through distributed protocols to share inference workloads . Su
Harrison Chase delivered a keynote address at a LangChain event that drew a large audience, as indicated by the 'packed house' reference. The event appears to have been interrupted at some point durin
arXiv:2605.12072v1 Announce Type: new Abstract: Dropout-based sparse-view 3D Gaussian Splatting (3DGS) methods alleviate overfitting by randomly suppressing Gaussian primitives during training. Existi
This paper investigates token superposition, a phenomenon where language models can encode multiple token representations simultaneously in a single position, enabling more efficient use of model capa
arXiv:2602.01418v2 Announce Type: replace Abstract: We propose Parabolic Position Encoding (PaPE), a parabola-based position encoding for vision modalities in attention-based architectures. Given a se
arXiv:2605.10953v1 Announce Type: cross Abstract: The demand for high-resolution subsurface imaging and continuous Earth monitoring has driven rapid growth in active and passive seismic data from dens
arXiv:2605.11684v1 Announce Type: new Abstract: We propose a Byzantine-resilient federated conformal prediction (FCP) method that leverages partial model sharing, where only a subset of model paramete
arXiv:2602.04042v2 Announce Type: replace Abstract: We propose Partition Tree, a novel tree-based framework for conditional density estimation over general outcome spaces that supports both continuous
Harrison Chase, the creator of LangChain, expressed positive sentiment about Baseten as a partnership, highlighting their collaborative work in supporting frontier AI teams. The post suggests that Bas
arXiv:2510.05497v5 Announce Type: replace-cross Abstract: Large-scale Mixture of Experts (MoE) Large Language Models (LLMs) have recently become the frontier open-weight models, achieving remarkable m
PayPal runs 74,000 weekly tasks in Perplexity Enterprise. Teams use it for model validation, channel performance, market trend research, competitive intelligence, and product analysis. Read the custom
arXiv:2605.11427v1 Announce Type: new Abstract: 4D Gaussian Splatting (4DGS) enables high-quality dynamic novel view synthesis, yet current models remain monolithic bitstreams that clients must downlo
arXiv:2510.02107v4 Announce Type: replace Abstract: AdaBoost sequentially fits so-called weak learners to minimize an exponential loss, which penalizes misclassified data points more severely than oth
arXiv:2605.11730v1 Announce Type: new Abstract: Automated red-teaming for LLMs often discovers narrow attack slices, missing diverse real-world threats, and yielding insufficient data for safety fine-
arXiv:2605.11266v1 Announce Type: new Abstract: Recent advances in Gaussian Splatting have enabled fast, high-fidelity 3D scene generation, yet these methods remain purely visual and lack an understan
arXiv:2602.13690v2 Announce Type: replace Abstract: Magnetic-anomaly navigation, leveraging small-scale variations in the Earth's magnetic field, is a promising alternative when GPS is unavailable or
arXiv:2512.05683v2 Announce Type: replace Abstract: Optical aberrations significantly degrade image quality in microscopy, particularly when imaging deeper into samples. These aberrations arise from d
arXiv:2605.11346v1 Announce Type: new Abstract: Physics-informed deep learning (PIDL) neural networks have shown their capability as a useful instrument for transportation practitioners in utilizing t
arXiv:2602.08058v2 Announce Type: replace Abstract: In the presence of occlusions and measurement noise, geometrically accurate scene reconstructions -- which fit the sensor data -- can still be physi
arXiv:2605.12492v1 Announce Type: new Abstract: We introduce Pion, a spectrum-preserving optimizer for large language model (LLM) training based on orthogonal equivalence transformation. Unlike additi
arXiv:2605.11225v1 Announce Type: cross Abstract: Large language model (LLM)-based agents frequently generate seemingly coherent plans that fail upon execution due to infeasible actions, constraint vi
arXiv:2212.02011v3 Announce Type: replace Abstract: Point cloud learning is receiving increasing attention. However, most existing point cloud models lack the practical ability to deal with the unavoi
arXiv:2605.11594v1 Announce Type: new Abstract: High-fidelity reconstruction of driving scenes is crucial for autonomous driving. While recent feedforward 3D Gaussian Splatting (3DGS) methods enable f
arXiv:2605.11520v1 Announce Type: new Abstract: Unsupervised point cloud segmentation is critical for embodied artificial intelligence and autonomous driving, as it mitigates the prohibitive cost of d
Poolside is hosting a 2-day model research hackathon in London. Join us to push an open-weight agent model as far as you can. RL and fine-tune Laguna XS.2, our latest-generation model, on Prime Intell
arXiv:2602.15473v2 Announce Type: replace Abstract: Gradient-based optimizers are highly sensitive to design choices in their adaptive learning rate mechanisms. To address this limitation, we introduc
arXiv:2605.11497v1 Announce Type: new Abstract: Zero-shot skeleton-based action recognition (ZSSAR) is typically treated as a skeleton-text alignment problem: encode joint-coordinate sequences, align
arXiv:2605.12144v1 Announce Type: new Abstract: In visual localization, Absolute Pose Regression (APR) enables real-time 6-DoF camera pose inference from single images, yet critically depends on fine-
arXiv:2512.11883v3 Announce Type: replace-cross Abstract: Over-aligning image generation models to a generalized aesthetic preference conflicts with user intent, particularly when 'anti-aesthetic' out
arXiv:2605.11511v1 Announce Type: cross Abstract: The validity of statistical inference depends critically on how data are collected. When data gathered through active data collection (ADC) are reused
arXiv:2605.11652v1 Announce Type: cross Abstract: We study posterior contraction rates for sparse Bayesian Kolmogorov-Arnold networks (KANs) over anisotropic Besov spaces, providing a statistical foun
This post describes the experience of using Starlink satellite internet service while flying at 30,000 feet on a United Airlines flight, highlighting the capability to post on X (formerly Twitter) fro
arXiv:2605.12411v1 Announce Type: cross Abstract: AI agents negotiate and transact in natural language with unfamiliar counterparts: a buyer bot facing an unknown seller, or a procurement assistant ne
arXiv:2605.12422v1 Announce Type: new Abstract: Automatic generation of educational materials using large language models (LLMs) is becoming increasingly common, but assigning difficulty levels to suc
arXiv:2605.11303v1 Announce Type: new Abstract: We investigate the use of Large Language Models (LLMs) for zero-shot prediction of Ryff Psychological Well-Being (PWB) scores from spontaneous speech. U