Can Heterogeneous Language Models Be Fused?
arXiv:2604.01674v2 Announce Type: replace Abstract: Model merging aims to integrate multiple expert models into a single model that inherits their complementary strengths without incurring the inferen
Knowledge catalogue
arXiv:2604.01674v2 Announce Type: replace Abstract: Model merging aims to integrate multiple expert models into a single model that inherits their complementary strengths without incurring the inferen
Introducing Starchild-1 from @odysseyml, the first ever real-time multimodal world model. This a model that can generate interactive simulations of the world that you can—for the first time ever—hear.
arXiv:2602.01970v2 Announce Type: replace Abstract: Reinforcement learning enhances the reasoning capabilities of large language models but often involves high computational costs due to rollout-inten
arXiv:2605.13030v1 Announce Type: cross Abstract: Model merging combines task experts into one model and avoids joint training, retraining, or deploying many expert models, but the merged model often
arXiv:2605.13587v1 Announce Type: cross Abstract: Near-infrared spectroscopy (NIRS) is rapid and non-destructive, but reliable calibration still depends heavily on spectral preprocessing. In routine p
arXiv:2605.11367v1 Announce Type: new Abstract: Recent advances in visual generative models have highlighted the promise of learning generative world models. However, most existing approaches frame wo
The UK AISI found Mythos Preview is the first model to solve both their cyber ranges end-to-end. No model had ever solved the AISI’s “Cooling Tower” cyber range before. We're getting it to defenders a
Yann LeCun says you cannot build a reliable agentic system without a world model LLMs don't have world models. They can't predict the consequences of their actions before taking them 'they just act, a
arXiv:2605.10419v1 Announce Type: cross Abstract: This paper investigates the effectiveness of large language models (LLMs) in answering questions over datasets. We examine their performance in two sc
arXiv:2602.04093v2 Announce Type: replace Abstract: Concept-based Models (CMs) enhance interpretability in deep learning by grounding predictions in human-understandable concepts. However, concept ann
Introducing Flux Matching, a generative modeling paradigm that generalizes diffusion models to vector fields that need not be the score function. Enables structural priors in the dynamics, faster samp
arXiv:2605.09131v1 Announce Type: new Abstract: The Model Context Protocol (MCP) has unified the interface between Large Language Models (LLMs) and external tools, yet a fundamental gap remains in how
arXiv:2605.07288v1 Announce Type: cross Abstract: The integration of Vision-Language-Action (VLA) models with World Models has gained increasing attention. One representative approach treats learned W
To train better open models, we need predictable scaling. Delphi is Marin’s first step: we pretrained many small models with one recipe, then extrapolated 300× to predict a 25B-param / 600B-token run
Today we're sharing our work on interaction models. A new class of model trained from scratch to handle real-time interaction natively, instead of gluing it onto a turn-based one. https://youtu.be/A12
Jensen Huang describes his early collaboration with Elon Musk, noting that he provided computational support for Tesla's Model S and Model 3 vehicles through NVIDIA technology. Huang credits this prof
.@BraceSproul changed our org's internal model in Fleet from Sonnet 4.6 to Kimi K2.6 and I didn't even notice. Open models are already good enough for most tasks, though not the hardest coding work ye
arXiv:2605.02202v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have achieved remarkable success in tasks such as image captioning and visual question answering (VQA). However, as their
arXiv:2605.02087v1 Announce Type: new Abstract: Some frontier AI developers aim to align language models to a Model Spec or Constitution that describes the intended model behavior. However, standard a
arXiv:2605.02348v1 Announce Type: new Abstract: Large language models pick up social biases from the data they are trained on and carry those biases into downstream applications, often reinforcing ste
arXiv:2605.00334v1 Announce Type: cross Abstract: Production agentic systems make many model calls per user request, and most of those calls are short, structured, and routine. This raises a practical
switching model providers is easy switching harnesses is less so model providers want to lock you in via harness we need open harnesses! TBH I don't agree with your take. I don't think Athropic's desi
arXiv:2604.27911v1 Announce Type: new Abstract: Foundation models are deep neural networks (such as GPT-5, Gemini~3, and Opus~4) trained on large datasets that can perform diverse downstream tasks --
big theme of 2026 - cost of closed models is too high! really excited to make deepagents work exceptionally well with OSS models Switched out Sonnet 4.6 for GLM 5.1 through @FireworksAI_HQ while doing
arXiv:2604.25119v1 Announce Type: new Abstract: Auditing the fine-tunes of open-weight generative models for harmful specialization has become a new governance challenge for model hosting platforms. T
Chris Welch / Bloomberg: Motorola unveils its 2026 foldables lineup, including its first book-style model, which costs 1,900; prices for clamshell models have gone up by up to 200 — The company's clam
arXiv:2604.24302v1 Announce Type: new Abstract: Mechanistic interpretability has made it possible to localize circuits underlying specific behaviors in language models, but existing methods are expens
arXiv:2604.24082v1 Announce Type: cross Abstract: Large (vision-)language models exhibit remarkable capability but remain highly susceptible to jailbreaking. Existing safety training approaches aim to
@jon_barron 'World models' has a technical meaning - the transition model/dynamics model from Bellman/Kalman in the context of MDPs/ state space approach to control theory ~ 1960. I gave a talk on thi
arXiv:2509.04802v3 Announce Type: replace Abstract: As large language models increasingly deployed into agentic systems, existing methods face critical gaps in observing, assessing, and mitigating dep
arXiv:2603.29928v2 Announce Type: replace Abstract: Tabular foundation models such as TabPFN and TabICL already produce full predictive distributions, yet prevailing regression benchmarks evaluate the
This is an incredibly cool experiment It is also fascinating that the model knows information up to 1931, but, at least in some science topics, seems very stuck in the early 1900s. For example, it def
arXiv:2604.22700v1 Announce Type: new Abstract: Understanding and predicting the progression of neurodegenerative diseases remains a major challenge in medical AI, with significant implications for ea
On Friday, Chinese AI firm DeepSeek released a preview of V4, its long-awaited new flagship model. Notably, the model can process much longer prompts than its last generation, thanks to a new design t
arXiv:2604.19809v1 Announce Type: new Abstract: We introduce MIRROR, a benchmark comprising eight experiments across four metacognitive levels that evaluates whether large language models can use self
arXiv:2604.20551v1 Announce Type: cross Abstract: Mixture-of-experts models provide a flexible framework for learning complex probabilistic input-output relationships by combining multiple expert mode
arXiv:2604.18957v1 Announce Type: new Abstract: Extracting standardized metallurgical metrics from microscopy images remains challenging due to complex grain morphology and the data demands of supervi
Here's how anyone can find models that work for your hardware easily. 1. Go to http://huggingface.co and make an account 2. Models tab to find weights and all compressions 3. Click on your profile on
arXiv:2604.17282v1 Announce Type: new Abstract: Process-Level Reward Models (PRMs) are essential for guiding complex reasoning in large language models, yet existing PRM benchmarks cover only general
arXiv:2503.03480v4 Announce Type: replace Abstract: Vision-language-action models (VLAs) show potential as generalist robot policies. However, these models pose extreme safety challenges during real-w
arXiv:2604.16443v1 Announce Type: cross Abstract: Data-driven models for building thermal dynamics are a scalable approach for enabling energy-efficient operation through fault detection & diagnosis o
arXiv:2604.18463v1 Announce Type: cross Abstract: Large language models are increasingly used as planners for robotic systems, yet how safely they plan remains an open question. To evaluate safe plann
arXiv:2604.15490v1 Announce Type: new Abstract: Recent developments in reasoning capabilities have enabled large language models to solve increasingly complex mathematical, symbolic, and logical tasks
> grok4.20-beta1 is a much smaller model than opus but is #1 ranked in medicine and healthcare > 4.3 and 4.4 will be much larger models, and likely will have a significant boost in performance on comp
arXiv:2505.02979v3 Announce Type: replace-cross Abstract: We propose a novel inverse-modelling approach which estimates the parameters of a simple land-surface model (LSM) by assimilating data into a
⚡ Meet Qwen3.6-35B-A3B:Now Open-Source!🚀🚀 A sparse MoE model, 35B total params, 3B active. Apache 2.0 license. 🔥 Agentic coding on par with models 10x its active size 📷 Strong multimodal perception an
arXiv:2511.17792v2 Announce Type: replace Abstract: While recent video world models can generate highly realistic videos, their ability to perform semantic reasoning and planning remains unclear and u
arXiv:2604.12033v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) increasingly rely on retrieval to answer knowledge-intensive multimodal questions. Existing benchmarks overlook c
arXiv:2604.11135v1 Announce Type: cross Abstract: Pretrained video generation models provide strong priors for robot control, but existing unified world action models still struggle to decode reliable
arXiv:2604.09866v1 Announce Type: cross Abstract: Recent advances in large language models (LLMs) have shown the promise to significantly accelerate the workflow by automating structural modeling and
arXiv:2604.10701v1 Announce Type: cross Abstract: Credit assignment is a central challenge in reinforcement learning (RL). Classical actor-critic methods address this challenge through fine-grained ad
Given the messy naming scheme used by all the AI companies, I caused a chart to be made showing the gain in GPQA per 0.1 version in model names (estimated, since model names skip version numbers). The
arXiv:2604.10556v1 Announce Type: new Abstract: While Diffusion Large Language Models (dLLMs) have emerged as a promising non-autoregressive paradigm comparable to autoregressive (AR) models, their fa
arXiv:2604.11446v1 Announce Type: cross Abstract: Recently, scaling reinforcement learning with verifiable rewards (RLVR) for large language models (LLMs) has emerged as an effective training paradigm
arXiv:2604.11061v1 Announce Type: cross Abstract: Mechanistic interpretability is often motivated for alignment auditing, where a model's verbal explanations can be absent, incomplete, or misleading.
arXiv:2604.10966v1 Announce Type: cross Abstract: We present a discriminative multimodal reward model that scores all candidate responses in a single forward pass. Conventional discriminative reward m
arXiv:2604.08685v1 Announce Type: new Abstract: Automated planning algorithms require an action model specifying the preconditions and effects of each action, but obtaining such a model is often hard.
This r/StableDiffusion post discusses community recommendations for the best AI image generation model to create fictional country flags, comparing options including SDXL, Qwen, Wan, ZIT, ZIB, Flux Kl
arXiv:2604.06475v1 Announce Type: new Abstract: Deep Learning Reduced Order Models (ROMs) are becoming increasingly popular as surrogate models for parametric partial differential equations (PDEs) due
arXiv:2601.05529v5 Announce Type: replace Abstract: High success rates on navigation-related tasks do not necessarily translate into reliable decision making by foundation models. To examine this gap,