Training ML Models with Predictable Failures
arXiv:2605.15134v1 Announce Type: new Abstract: Estimating how often an ML model will fail at deployment scale is central to pre-deployment safety assessment, but a feasible evaluation set is rarely l
Knowledge catalogue
arXiv:2605.15134v1 Announce Type: new Abstract: Estimating how often an ML model will fail at deployment scale is central to pre-deployment safety assessment, but a feasible evaluation set is rarely l
arXiv:2605.14368v1 Announce Type: cross Abstract: Continuous diffusion language models lag behind autoregressive transformers, partly because diffusion is applied in spaces poorly suited to language d
arXiv:2605.13724v1 Announce Type: cross Abstract: Few-step video generation has been significantly advanced by consistency distillation. However, the performance of consistency-distilled models often
arXiv:2605.12517v1 Announce Type: cross Abstract: Vision-language models (VLMs) are often deployed on text-only inputs, although they are trained with images. We find that removing the vision modality
This Reddit post likely compares inference performance metrics across popular language models running on Ollama, measuring tokens per second as a key performance indicator. The post would help users a
arXiv:2605.12952v1 Announce Type: new Abstract: Grad-ECLIP is published at ICML 2024 and represents a new Transformer interpretation technical route (intermediate features-based). First, this paper de
DID SOMEONE SAY FREE TOKENS IN FLEET???? Yes it's true. Fleet now has a built in model powered by @FireworksAI_HQ that's free for all Developer & Plus plan users. Try it out today (or before @hwchase1
arXiv:2605.12790v1 Announce Type: new Abstract: Modeling concentric tube robots (CTRs) involves complex nonlinear continuum mechanics, and despite recent progress, physics-based models often lack an a
LangSmith Fleet has introduced a free model option powered by Fireworks AI for users on Developer and Plus plans, lowering the barrier to entry for getting started with the platform. This update makes
arXiv:2602.23013v3 Announce Type: replace Abstract: Detecting visual anomalies in industrial inspection often requires training with only a few normal images per category. Recent few-shot methods achi
arXiv:2605.13158v1 Announce Type: new Abstract: Image restoration under multiple adverse weather conditions aims to develop a single model to recover the underlying scene with high visibility. Weather
you can see more here (color stuff is at bottom): http://colorfeelings.replit.app - changing embedding model may change results - the similarity range was pretty small so normalizing really stretched
arXiv:2605.11233v1 Announce Type: new Abstract: Effective model selection is critical in symbolic regression (SR) to identify mathematical expressions that balance accuracy and complexity, and have lo
arXiv:2602.06339v2 Announce Type: replace Abstract: Robot Foundation Models, such as VLAs, promise end-to-end generative robot policies with broad generalization. Yet it remains unclear whether they f
arXiv:2605.11222v1 Announce Type: new Abstract: Quantization is an effective strategy to reduce the storage and computation footprint of large language models (LLMs). Post-training quantization (PTQ)
arXiv:2605.11311v1 Announce Type: cross Abstract: Diffusion models typically generate image batches from independent Gaussian initial noises. We argue that this independence assumption is only one cho
arXiv:2605.12138v1 Announce Type: cross Abstract: Generating realistic and user-preferred advertisements is a key challenge in e-commerce. Existing approaches utilize multiple independent models drive
arXiv:2605.10993v1 Announce Type: new Abstract: Memory capacity is a critical factor determining the performance of Vision-Language-Action (VLA) models in long-horizon manipulation tasks. Existing mem
arXiv:2605.11480v1 Announce Type: new Abstract: Reward fine-tuning has become a common approach for aligning pretrained diffusion and flow models with human preferences in text-to-image generation. Am
arXiv:2605.11527v1 Announce Type: new Abstract: Diffusion models are the leading approach for tabular data synthesis and are increasingly used to share sensitive records. Whether they actually protect
arXiv:2605.12309v1 Announce Type: new Abstract: The development of separate-encoder Unified multimodal models (UMMs) comes with a rapidly growing inference cost due to dense visual token processing. I
Introducing GridSFM, a small foundation model that can predict AC optimal power flow in milliseconds, boosting efficiency and unlocking cost savings. Learn how GridSFM gives grid operators direct visi
arXiv:2605.11596v1 Announce Type: new Abstract: Closed-loop driving simulation requires real-time interaction beyond short offline clips, pushing current driving world models toward autoregressive (AR
arXiv:2605.12258v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have achieved remarkable progress, yet the object hallucination remains a critical challenge for reliable deplo
arXiv:2605.10948v1 Announce Type: cross Abstract: Understanding how vegetation loss alters rainfall remains a major challenge in climate and hydrological science, as deforestation modifies precipitati
arXiv:2412.18594v3 Announce Type: replace Abstract: Gaussian graphical model selection is usually studied under independent sampling, but in many applications observations arise from dependent dynamic
arXiv:2511.06077v2 Announce Type: replace Abstract: Short-video recommenders such as Douyin must exploit extremely long user histories without breaking latency or cost budgets. We present an end-to-en
arXiv:2605.02798v1 Announce Type: cross Abstract: We present an experimental study of energy-to-solution (ETS) of hybrid quantum-classical applications, enabled by direct instrumentation of power cons
Our evaluations show that frontier AI's cyber capabilities are advancing quickly. The length of cyber tasks frontier models can complete has been doubling every few months, and this rate has become fa
arXiv:2510.05497v5 Announce Type: replace-cross Abstract: Large-scale Mixture of Experts (MoE) Large Language Models (LLMs) have recently become the frontier open-weight models, achieving remarkable m
arXiv:2605.11907v1 Announce Type: new Abstract: We measure procedural-skill SFT contribution across three Qwen3.5 dense scales (0.8B, 2B, 4B) on a 200-task / 40-skill holdout, with Claude Haiku 4.5 as
arXiv:2605.12398v1 Announce Type: new Abstract: Estimating question difficulty is a critical component in evaluating and improving large language models (LLMs) for question answering (QA). Existing ap
arXiv:2605.12264v1 Announce Type: cross Abstract: Supervised Finetuning (SFT) has become one of the primary methods for adapting a large language model (LLM) with extensive pre-trained knowledge to do
Recursive Superintelligence Inc., a startup that hopes to develop self-improving artificial intelligence models, launched today with 650 million in funding. Alphabet Inc.’s GV fund and Greycroft led t
arXiv:2605.11550v1 Announce Type: new Abstract: A plausible scene evolution depends on the maneuver being considered, while a good maneuver depends on how the scene may evolve. Existing World Action M
This article argues that infrastructure quality and architecture are more critical to AI/ML project success than the choice of underlying model. It likely explores how proper deployment, scaling, moni
arXiv:2605.08207v1 Announce Type: new Abstract: Pathology foundation models have shown strong retrospective performance, but whether such systems can support clinically relevant use remains unclear. T
arXiv:2605.08504v1 Announce Type: new Abstract: We investigate the origins of massive activations in large language models (LLMs) and identify a specific layer named the extbf{Massive Emergence Layer
arXiv:2510.04988v3 Announce Type: replace Abstract: The vast majority of modern deep learning models are trained with momentum-based first-order optimizers. The momentum term governs the optimizer's m
MatterSim is expanding what AI can do for materials science—from faster large-scale simulations to MatterSim-MT, a new multi-task model for simulating properties beyond potential energy surfaces alone
arXiv:2605.10384v1 Announce Type: new Abstract: Agentic artificial intelligence (AI) is a natural fit for Internet of Things (IoT) and edge systems, but edge deployments are often constrained to model
Thinking Machines released TML-Interaction-Small, a 276B parameter model with a 12B active subset designed for real-time voice interactions that reportedly surpasses standard voice activity detection
arXiv:2407.12173v2 Announce Type: replace-cross Abstract: Generative diffusion models have emerged as a powerful tool for high-quality image synthesis, yet their iterative nature demands significant c
arXiv:2605.09134v1 Announce Type: new Abstract: Reinforcement learning for program repair is hindered by sparse execution feedback and coarse sequence-level rewards that obscure which edits actually f
arXiv:2605.08653v1 Announce Type: new Abstract: Accurate state-of-charge (SOC) estimation is critical for the safe and efficient operation of lithium-ion batteries in battery management systems (BMS).
arXiv:2605.08735v1 Announce Type: new Abstract: Recent 'Thinking with Video' approaches use Video Generation Models (VGMs) for visual reasoning by producing temporally coherent Chain-of-Frames as reas
arXiv:2605.09112v1 Announce Type: cross Abstract: Selecting a coherent sequence or subset of elements is a fundamental problem in structured prediction, arising in tasks such as detection, trajectory
arXiv:2605.09875v1 Announce Type: new Abstract: Large language models from different families use different hidden dimensions, tokenizers, and training procedures, making behavioral directions difficu
arXiv:2605.08902v1 Announce Type: cross Abstract: In recent years, pre-trained visual-linguistic models have demonstrated tremendous potential, becoming a crucial foundational framework for numerous d
arXiv:2605.10790v1 Announce Type: new Abstract: Diffusion models have achieved remarkable success, yet their training remains inefficient due to a severe optimization bottleneck, which we term Represe
arXiv:2605.08645v1 Announce Type: cross Abstract: Energy-based models (EBMs) provide a powerful and flexible way of learning a joint probability distribution over data by constructing an energy surfac
arXiv:2605.08572v1 Announce Type: new Abstract: Diffusion models for multi-agent trajectory prediction are limited by iterative denoising, which causes inference latency that hinders their use in time
arXiv:2605.09853v1 Announce Type: new Abstract: Post-training techniques combined with inference-time scaling significantly enhance the reasoning and alignment capabilities of large language models (L
arXiv:2605.10045v1 Announce Type: new Abstract: Visual Autoregressive (VAR) models have emerged as a strong alternative to diffusion for image synthesis, yet their fixed training resolution prevents d
arXiv:2605.09218v1 Announce Type: cross Abstract: 3D scene understanding spans reasoning about free space, object grounding, hypothetical object insertions, complex geometric relationships, and integr
arXiv:2603.04474v2 Announce Type: replace-cross Abstract: Large Language Model-based Multi-Agent Systems (LLM-MAS) are increasingly applied to complex collaborative scenarios. However, their collabora
arXiv:2605.08702v1 Announce Type: cross Abstract: This paper tackles compositional personalization of vision-language models (VLMs). In this problem, multiple user-defined concepts must be recognized
arXiv:2605.08392v1 Announce Type: new Abstract: Practical diffusion sampling is a numerical approximation problem: under a fixed inference budget, one must simulate a reverse-time ODE or SDE using onl
arXiv:2605.10739v1 Announce Type: cross Abstract: We introduce SMART-HC-VQA, a Sentinel-2-based visual question answering dataset derived from the IARPA SMART Heavy Construction dataset, designed for
i like that there are models called bert and ernie, but in all seriousness, this update looks impressive ERNIE 5.1 is here 🚀 ERNIE 5.1 significantly reduces pretraining cost while compressing total pa