Spherical Flows for Sampling Categorical Data
arXiv:2605.05629v2 Announce Type: replace-cross Abstract: We study the problem of learning generative models for discrete sequences in a continuous embedding space. Whereas prior approaches typically
Knowledge catalogue
arXiv:2605.05629v2 Announce Type: replace-cross Abstract: We study the problem of learning generative models for discrete sequences in a continuous embedding space. Whereas prior approaches typically
arXiv:2605.09357v1 Announce Type: cross Abstract: Running deep neural networks on microcontroller units (MCUs) is severely constrained by limited memory resources. While TinyML techniques reduce model
arXiv:2508.14685v4 Announce Type: replace Abstract: While transformer models exhibit strong in-context learning (ICL) abilities, they often fail to generalize under simple distribution shifts. We anal
arXiv:2605.05569v2 Announce Type: replace-cross Abstract: This paper shows that the semi-dual formulation of the optimal transport problem has a degenerate saddle-point structure, and that its numeric
arXiv:2605.10154v1 Announce Type: new Abstract: Long-horizon forecasting of time-dependent partial differential equations (PDEs) is critical for characterizing the sustained evolution of physical syst
Starlink is now onboard @GulfAir Experience seamless streaming, gaming, and scrolling from gate to gate 🛰️✈️ We just made history. ✈️🌐 Gulf Air’s first @Starlink flight just took off. Full internet. C
Starlink Mobile V2 coming next year will be a game changer for global connectivity! Thank you @FCC Thanks to President Trump, America is leading the world again. 🇺🇸 Today, the @FCC approved two major
arXiv:2605.08114v1 Announce Type: new Abstract: We analyse three KV cache quantization schemes under a fair bit budget: extbf{KV} (scalar MSE baseline), extbf{KQV} (WHT + MSE on K; WHT + MSE + QJL on
arXiv:2605.10447v1 Announce Type: cross Abstract: Agent-based models (ABMs) are increasingly used in macroeconomics, but their analysis still often relies on ad hoc Monte Carlo campaigns with heteroge
arXiv:2605.09618v1 Announce Type: new Abstract: When should a language model answer directly, sample and vote, or engage in multi-agent debate? Recent work shows voting often explains much of the gain
arXiv:2503.09336v4 Announce Type: replace Abstract: Backdoor attacks pose a severe threat to deep neural networks (DNNs) by implanting hidden backdoors that can be activated with predefined triggers t
arXiv:2604.02608v2 Announce Type: replace Abstract: Activation steering presupposes that task-relevant behaviors correspond to linear directions in activation space -- directions that should both stee
arXiv:2605.10220v1 Announce Type: cross Abstract: The formation timescale of the Milky Way thick disk is one of the central debates in Galactic archaeology. The age-metallicity relation (AMR), formati
Nous Research announced that Step 3.5 Flash, an AI model from StepFun, is temporarily available for free on the Nous Portal for a 15-day period. This offer provides users access to the model without c
arXiv:2605.10674v1 Announce Type: cross Abstract: Rejection Fine-Tuning (RFT) is a standard method for training LLM agents, where unsuccessful trajectories are discarded from the training set. In the
arXiv:2605.09989v1 Announce Type: cross Abstract: Recent advances in robot imitation learning have yielded powerful visuomotor policies capable of manipulating a wide variety of objects directly from
arXiv:2605.10442v1 Announce Type: cross Abstract: Multilingual studies of social bias in open-ended LLM generation remain limited: most existing benchmarks are English-centric, template-based, or rest
arXiv:2605.09415v1 Announce Type: new Abstract: The growing integration of AI into cybersecurity is reshaping the balance between attackers and defenders. When access to advanced AI-enabled defence to
arXiv:2605.10059v1 Announce Type: new Abstract: Agent-based modeling (ABM) has long been used in economics to study human behavior, and large language model (LLM) agents now enable new forms of social
arXiv:2505.06835v4 Announce Type: replace Abstract: Sliced optimal transport (SOT), or sliced Wasserstein (SW) distance, is widely recognized for its statistical and computational scalability. In this
arXiv:2604.01824v2 Announce Type: replace Abstract: We introduce STRIVE (SpatioTemporal Reinforcement with Importance-aware Variant Exploration), a structured reinforcement learning framework for vide
arXiv:2502.18334v5 Announce Type: replace Abstract: Graph-based learning excels at capturing interaction patterns in diverse domains like recommendation, fraud detection, and particle physics. However
arXiv:2605.08689v1 Announce Type: cross Abstract: Graph foundation models (GFMs) seek transferable representations across graph domains but are limited by structural heterogeneity and incompatible nod
arXiv:2605.08559v1 Announce Type: cross Abstract: Convex functionals are ubiquitous in applied analysis, appearing as value functions, risk measures, super-hedging prices, and loss functionals in mach
arXiv:2605.08696v1 Announce Type: new Abstract: Over the last two decades, language modeling has experienced a shift from predominantly recurrent architectures that process tokens sequentially during
arXiv:2605.09845v1 Announce Type: new Abstract: Sub-footprint target mixing within a laser footprint significantly increases LiDAR intensity uncertainty, especially in complex environments where heter
arXiv:2605.09241v1 Announce Type: cross Abstract: Joint-Embedding Predictive Architectures (JEPAs) provide a simpleframework for learning world models by predicting future latent representations.Howev
arXiv:2605.08991v1 Announce Type: new Abstract: A series of papers has introduced the Heuristic Rating Estimation method, which evaluates a set of alternatives based on pairwise comparisons and the we
If anyone, anywhere builds a superhuman artificial intelligence using present methods, the most likely outcome is catastrophe. There have accordingly been widespread calls for an international agreeme
arXiv:2502.00816v4 Announce Type: replace Abstract: We introduce Sundial, a family of native, flexible, and scalable time series foundation models. To predict the next-patch's distribution, we propose
arXiv:2605.09834v1 Announce Type: cross Abstract: Modern predictive systems encode beliefs that can act as useful prior information for statistical inference in data-limited settings. Using them for p
arXiv:2605.08698v1 Announce Type: new Abstract: Stable Diffusion (SD) has evolved DDPM (Denoising Diffusion Probabilistic Model) based image generation significantly by denoising in latent space inste
arXiv:2604.03928v2 Announce Type: replace-cross Abstract: Frozen pretrained image representations are widely used for transfer learning: a backbone is kept fixed, feature vectors are extracted, and a
arXiv:2601.20756v2 Announce Type: replace Abstract: Score-based diffusion models have recently been extended to infinite-dimensional function spaces, with uses such as inverse problems arising from pa
arXiv:2601.21971v2 Announce Type: replace-cross Abstract: Imitation learning has achieved remarkable success in robotic manipulation, yet its application to surgical robotics remains challenging due t
Clem Delangue, CEO of Hugging Face, shared excitement about seeing Reachy Mini, a small humanoid robot, featured on the cover of Linus Tech Tips' latest video. The post suggests Reachy Mini received n
arXiv:2605.08963v1 Announce Type: cross Abstract: Machine Learning (ML) models trained on complex health surveys such as the National Health and Nutrition Examination Survey (NHANES) often ignore prim
arXiv:2605.08196v1 Announce Type: new Abstract: Recent natural disasters have highlighted the urgent need for efficient data-driven approaches to disaster management. Machine learning (ML) and deep le
arXiv:2605.10052v1 Announce Type: cross Abstract: As artificial intelligence engineering paradigms shift from single-agent Prompt and Context Engineering toward multi-agent extbf{Coordination Engineer
arXiv:2605.08366v1 Announce Type: new Abstract: We introduce SWE Atlas, a benchmark suite for coding agents spanning three professional software engineering workflows: Codebase Q&A (124 tasks), Test W
arXiv:2605.09442v1 Announce Type: cross Abstract: Streaming long-video generation faces a central challenge in continuous semantic switching, requiring adaptive memory to preserve coherent visual evol
arXiv:2605.06356v2 Announce Type: replace Abstract: High-resolution image-to-video (I2V) generation aims to synthesize realistic temporal dynamics while preserving fine-grained appearance details of t
arXiv:2605.10811v1 Announce Type: cross Abstract: This paper develops a joint spectral radius (JSR) framework for analyzing rank-one deflated Q-value iteration (Q-VI) in discounted Markov decision pro
Francois Chollet argues that symbolic learning represents a fundamental alternative to gradient descent and neural networks rather than a replacement for coding agents, offering a low-level, general-p
arXiv:2602.21307v2 Announce Type: replace Abstract: What mathematical functions do neural network components learn? Symbolic distillation addresses this question by expressing neural network component
arXiv:2605.08412v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have made rapid progress in single-video understanding, yet their ability to reason across multiple independent
arXiv:2605.08835v1 Announce Type: new Abstract: The expansion of Artificial Intelligence-generated content service requires diffusion model serving to simultaneously achieve high throughput and low ta
arXiv:2605.08190v1 Announce Type: new Abstract: Autonomous systems increasingly rely on machine-learning (ML) components for safety-critical tasks such as perception and control in autonomous vehicles
arXiv:2605.08724v1 Announce Type: new Abstract: Unifying multimodal understanding and generation is a compelling frontier that is beginning to emerge in the medical field. However, the limited existin
arXiv:2605.10129v1 Announce Type: new Abstract: Large language models (LLMs) rely on web-scale corpora for pre-training. The noise inherent in these datasets tends to obscure meaningful patterns and u
TabPFN-3 is a pre-trained tabular foundation model that supports datasets up to 1,000,000 rows × 200 features , representing a significant scaling improvement for the TabPFN family. The model delivers
arXiv:2605.09424v1 Announce Type: new Abstract: Generative modelling is a demanding test of foundation models, because it requires robust, holistic representation learning for a given data modality, r
arXiv:2605.09539v1 Announce Type: new Abstract: Multi-agent systems (MAS) have emerged as a promising paradigm for solving complex tasks. Recent work has explored self-evolving MAS that automatically
arXiv:2605.09536v1 Announce Type: cross Abstract: Diffusion large language models (dLLMs) offer a promising paradigm for parallel text generation, but in practice they face an accuracy-parallelism tra
arXiv:2506.01352v2 Announce Type: replace Abstract: Decentralized training of large language models offers the opportunity to pool computational resources across geographically distributed participant
arXiv:2505.11604v5 Announce Type: replace Abstract: Editing presentation slides is a frequent yet tedious task, ranging from creative layout design to repetitive text maintenance. While recent GUI-bas
Talked to a friend at a top AI lab. Their whole team is former journalists, training models on poems, summaries, and creative writing. I use AI every day and can see that it tends to flatten my writin
arXiv:2602.04611v2 Announce Type: replace-cross Abstract: The synthetic control method (SCM) estimates causal effects in panel data with a single-treated unit by constructing a counterfactual outcome
arXiv:2605.08440v1 Announce Type: cross Abstract: Adversarial purification with diffusion models seeks to project adversarial examples back toward the data manifold, but balancing semantic preservatio
arXiv:2605.10165v1 Announce Type: cross Abstract: Noisy labels are common in large-scale medical imaging datasets due to inter-observer variability and ambiguous cases. We propose a statistically grou