Open models for the win!
Open models for the win! For my first post, I’m sharing a letter @NVIDIA signed on why open models matter. AI will transform every industry, power every company, and be built by every country. Open mo
Knowledge catalogue
Open models for the win! For my first post, I’m sharing a letter @NVIDIA signed on why open models matter. AI will transform every industry, power every company, and be built by every country. Open mo
The Washington Post processed 1.79B input tokens per month through Together AI, running open models like Llama and Mistral in production with predictable costs and full control over the model stack. T
arXiv:2607.21090v1 Announce Type: cross Abstract: We propose a Reinforcement Learning (RL) method to directly optimize the faithfulness of self-explanations - the extent to which a model's generated r
Arize and Fireworks benchmarked 10 AI models across 2,400 agent runs. Learn why cost per successful task beats token price for model evaluation and routing. The post Cost per successful task: Benchmar
Huge launch from @tryramp. Different steps in an agent workflow can use different models. This can help reduce costs significantly without sacrificing performance. Model routing will become a core par
Who’s Afraid of Chinese Models? Interesting proposal from Ben Thompson that both addresses the hypocrisy of labs outlawing distillation against their models despite training on unlicensed data, and co
arXiv:2607.12884v1 Announce Type: new Abstract: Patients seeking medical information often ask questions that embed incorrect assumptions or misconceptions. In such cases, safe medical communication r
arXiv:2607.13017v1 Announce Type: cross Abstract: World Action Models (WAMs) are able to leverage pretrained video generators for both world modeling and action prediction. However, directly leveragin
arXiv:2607.11889v1 Announce Type: cross Abstract: Large language models trained on unrestricted internet corpora inevitably embed information from the future, introducing lookahead bias that compromis
U.S. open‑source AI models are rapidly gaining popularity, with NVIDIA’s newest model, **Nemotron Ultra**, becoming a prominent entry on the Ollama platform. On Ollama, Nemotron Ultra is quickly scali
Another big reason to use combination of frontier models. Chain-of-thought monitoring is treated as a reliable safety layer for agents. This DeepMind-affiliated study shows the layer can be argued out
arXiv:2510.16165v2 Announce Type: replace Abstract: A key question in benchmarking generative crystal reconstruction models is how the amount and type of crystallographic information provided to a gen
arXiv:2607.06216v1 Announce Type: new Abstract: The future of World Models depends not only on scaling model capability, but also on scaling practicality and inference efficiency. High-frame-rate infe
arXiv:2506.07468v4 Announce Type: replace-cross Abstract: Conventional large language model (LLM) safety alignment relies on a reactive, disjoint loop: attackers exploit a static model, then defenders
arXiv:2607.04755v1 Announce Type: new Abstract: Model merging offers a practical alternative to conventional continual learning by integrating independently fine-tuned models without retaining previou
arXiv:2503.06643v2 Announce Type: replace-cross Abstract: In this paper, we tackle a critical challenge in model evaluation: how to keep code benchmarks useful when models might have already seen them
arXiv:2607.04546v1 Announce Type: cross Abstract: Action-conditioned world models allow robots to predict the future consequences of candidate actions without additional physical interaction, supporti
Sakana AI (@SakanaAILabs) is now a model vendor on Merge Gateway, and Fugu Ultra is live through them. It's a multi-agent orchestration model that routes across frontier models behind one API. You get
arXiv:2607.03953v1 Announce Type: cross Abstract: This study independently replicates and extends the Natural Language Tools (NLT) framework of Johnson et al.~(2025), which questions the use of struct
introducing tinyrouter i reverse engineered the routing architecture behind Skana AI's Fugu and built replication for open frontier models. it's a tiny ~10K parameter LLM router that learns which mode
arXiv:2607.01436v1 Announce Type: new Abstract: Diffusion language models, which generate text by denoising a token canvas bidirectionally instead of emitting tokens left to right, have become competi
arXiv:2511.05536v2 Announce Type: replace-cross Abstract: Earth s gravity fundamentally shapes human behaviour. The brain encodes this force as an internal model of gravity, enabling the prediction an
arXiv:2607.01986v1 Announce Type: new Abstract: Multivariate time-series models for prognostics are often evaluated by point prediction accuracy, yet their internal states rarely expose a coherent deg
arXiv:2607.00083v1 Announce Type: cross Abstract: Language models have changed from unreliable text generators to highly-capable large models with trillions of parameters. Capability increases come ha
arXiv:2606.16517v2 Announce Type: replace Abstract: Scientific reasoning models for biology combine language models with foundation models trained on multimodal biological data, including DNA, RNA, an
arXiv:2511.08307v2 Announce Type: replace-cross Abstract: Generative models, such as large language models or text-to-image diffusion models, can generate relevant responses to user-given queries. Res
arXiv:2606.30140v1 Announce Type: cross Abstract: Recent breakthroughs in foundation models and Large Language Models (LLMs) have introduced new opportunities for studying and decoding genomic sequenc
arXiv:2606.28962v1 Announce Type: cross Abstract: Model quantization is essential for the efficient deployment of Large Language Models (LLMs), but introduces a critical vulnerability: Quantization-Co
arXiv:2606.29059v1 Announce Type: cross Abstract: World modeling requires forecasting uncertain futures while preserving information useful for downstream perception. Existing visual world models ofte
arXiv:2606.29646v1 Announce Type: cross Abstract: Sleeper agents are the canonical model organism of deception: models trained to behave normally but to emit an unsafe behaviour on a specific trigger.
arXiv:2606.30062v1 Announce Type: cross Abstract: While large language models have been dominating the research landscape recently, small language models remain highly relevant across various domains;
arXiv:2606.29247v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models represent a promising direction for embodied intelligence in surgical robotics. Despite the prevalence of VLA benchm
arXiv:2509.19671v3 Announce Type: replace Abstract: Public datasets of Chest X-Rays (CXRs) have long been a popular benchmark for developing machine learning (ML) computer vision models in healthcare.
I put together a new article on setting up local coding agents with open-weight models. Everything runs 100% locally. I thought it might be useful putting this together because many people asked me ab
arXiv:2606.26578v1 Announce Type: new Abstract: Automating optimization modeling from natural language with large language models (LLMs) faces two key challenges. First, training corpora lack structur
arXiv:2606.27325v1 Announce Type: new Abstract: Recent advances in action-conditioned world models show promising progress in modeling complex interactions and forecasting future states under diverse
arXiv:2511.12100v2 Announce Type: replace Abstract: In current visual model training, models often rely on only limited sufficient causes for their predictions, which makes them sensitive to distribut
arXiv:2606.24083v1 Announce Type: cross Abstract: 'Talk short. Drop grammar. Save token.' This caveman style is widely promoted as a way to cut inference cost, but whether it actually saves anything d
arXiv:2606.24256v1 Announce Type: new Abstract: Physical interactions follow a long-tailed distribution: a set of common and regular interactions dominates human experience and visual data, while a br
arXiv:2606.22317v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) is widely viewed as a promising path toward continuously improving large language models. Recent w
arXiv:2509.21489v3 Announce Type: replace Abstract: Graph foundation models face several fundamental challenges including transferability across diverse domains and data scarcity, which calls into que
arXiv:2606.20726v1 Announce Type: new Abstract: We introduce a compact empirical model that quantifies how answer accuracy degrades as a function of frame budget B and temporal distance D in long vide
arXiv:2606.18627v2 Announce Type: replace Abstract: Model merging has emerged as a training-free alternative to multi-task learning, aiming to combine multiple task-specific fine-tuned models into a s
arXiv:2606.22606v1 Announce Type: cross Abstract: Large language models (LLMs) achieve strong relation extraction (RE), but their computational demands and reliance on proprietary APIs limit deploymen
arXiv:2511.22699v4 Announce Type: replace Abstract: The landscape of high-performance image generation models is currently dominated by proprietary systems, such as Nano Banana Pro and Seedream 4.0. L
AI2 released TMax 27B, a 27 billion parameter terminal agent model available on Hugging Face that achieves 42.7% performance on Terminal Bench 2.0, matching the capabilities of much larger models desp
arXiv:2606.11386v1 Announce Type: cross Abstract: Full-duplex spoken language models (FD-SLMs) enable seamless speech interaction by allowing models to listen and speak simultaneously, yet the interna
arXiv:2606.11270v1 Announce Type: cross Abstract: Distillation of a language model intended to transfer benign behavior to a student model may also transfer undesirable characteristics, if they are pr
arXiv:2603.22017v2 Announce Type: replace Abstract: This work presents a collection of multi-modal domain adapted large language models built upon the instruction tuned variants of open weight models
arXiv:2606.09936v1 Announce Type: cross Abstract: World models are now built on substantially different computational substrates. Latent recurrent state-space models such as PlaNet and the Dreamer fam
arXiv:2606.11105v1 Announce Type: cross Abstract: Hallucinations, where language models (LMs) generate factually ungrounded responses, pose serious risks, as users tend to blindly rely on them. This i
arXiv:2606.09932v1 Announce Type: cross Abstract: Supervised Fine-Tuning (SFT) followed by Reinforcement Learning (RL) has become a standard pipeline for Large Language Model (LLM) post-training. SFT
arXiv:2606.08525v1 Announce Type: new Abstract: Reward models play a pivotal role in reinforcement learning (RL) and multi-modal trajectory selection for autonomous driving. However, acquiring such re
arXiv:2606.08948v1 Announce Type: cross Abstract: Comprehensive estimation of dietary micronutrients from food images could improve clinical nutrition care, but training such models requires large mul
arXiv:2603.24963v3 Announce Type: replace Abstract: Modern computational advertising platforms typically rely on recommendation systems to predict user responses, such as click-through rates, conversi
arXiv:2511.04567v2 Announce Type: replace-cross Abstract: Constructing reduced models for turbulent transport is essential for accelerating profile predictions and enabling many-query tasks such as pa
arXiv:2509.25522v3 Announce Type: replace Abstract: Recent advancements in generative models have allowed the emergence of a promising paradigm for recommender systems (RS), known as Generative Recomm
VLA-JEPA just dropped in LeRobot 🤖 What makes this model special is that it does not just learn what action to take from a given observation, it also leverages a JEPA world model to learn action-relev
arXiv:2606.04811v1 Announce Type: new Abstract: Video generation models have made impressive strides in synthesizing visually compelling content, yet their outputs remain confined to the virtual domai
arXiv:2606.05015v1 Announce Type: new Abstract: World models, learned generative models that predict how an environment evolves, have become a promising tool for sample-efficient robot learning. Yet h