GitHub Models is now retired
GitHub Models is now retired I missed this news until today, when the GitHub Actions run for my simonw/research repository failed with this error message: GitHub Models is temporarily unavailable as p
Knowledge catalogue
GitHub Models is now retired I missed this news until today, when the GitHub Actions run for my simonw/research repository failed with this error message: GitHub Models is temporarily unavailable as p
arXiv:2608.06310v1 Announce Type: cross Abstract: Recent advances in reward modeling show a paradigm shift from discriminative reward models to generative reward models. However, despite their strong
arXiv:2608.05104v1 Announce Type: new Abstract: Deep neural networks have shown impressive success in NLP tasks owing to their complex structure and huge number of edges. Achieving state-of-the-art pe
Taalas buried the lede for the amazing demo of their first tape out. 15k tokens per second with a llama 8b model, ability to scale that up etched onto silicon As models satisfice etching makes sense,
Hey everyone! I just released a sub-6B sparse activation AI model which was built with a brand new architecture : fusion. I fused weights from @liquidai's LFM2.5-2.6B & @Alibaba_Qwen's Qwen3.6-35B-A3B
French artificial intelligence startup Mistral AI SAS today introduced a lightweight multimodal safety artificial intelligence open-weight model that can classify outputs for AI models that outperform
tests: add model resolution test on synthetic repo listings (#26172) tests: add model resolution test on synthetic repo listings Include download.cpp and arg.cpp inside a namespace with hf_cache monke
arXiv:2608.00156v1 Announce Type: cross Abstract: The push for broader coverage in future cellular networks depends on reliable service, yet this is increasingly harder to do as we encounter more inst
arXiv:2608.01347v1 Announce Type: new Abstract: Large reasoning models used as coding agents incur costs from deliberation, tool calls, and repeated agent turns, yet the causal effect of prompt wordin
arXiv:2608.01534v1 Announce Type: new Abstract: Hard symbolic-reasoning tasks such as Sudoku, maze pathfinding, and ARC remain challenging for LLMs due to their fixed-depth autoregressive reasoning, w
arXiv:2608.00285v1 Announce Type: new Abstract: Sixteen language models drawn from ten families produced, on average, the semantic diversity of 1.69 distinct formulations of a psychotherapeutic case,
arXiv:2608.02050v1 Announce Type: new Abstract: Can a strictly local, iterated, weight-shared computation primitive support language modelling, and which of those three properties actually drives the
arXiv:2608.00105v1 Announce Type: new Abstract: Pathology foundation models are reported to encode molecular programmes in tissue morphology, but the evidence is usually a cohort-wide ranked gene list
I have been working on this tool for months and there are a lot of new functionalities and tests that are going to be released in the next few weeks! The goal of the tool is to allow community members
arXiv:2607.28982v1 Announce Type: cross Abstract: Recent language models achieve strong performance across a variety of tasks, but conventional adaptation applies updates uniformly across training sam
arXiv:2607.29033v1 Announce Type: new Abstract: Existing methods for adapting 2D foundation models such as SAM to 3D volumes either process slices independently---ignoring inter-slice context---or req
arXiv:2509.17192v3 Announce Type: replace Abstract: LLM-based social simulations can make a generated transcript look like a single behavioral signal, but the model behind that transcript may be doing
Has anyone figured out how to enable speculative decoding with deepseek v4 flash 0731 on llamacpp? I’m on the right release for llamacpp (b10228 or earlier) and running am17an’s draft model with unslo
So upfront I'll admit that this is a ChatGPT summary of my chat with it about this idea i had, but this post wouldn't exist any other way, so... I’m trying to get a better understanding of how VRAM is
arXiv:2607.27433v1 Announce Type: cross Abstract: We develop two parametric data-driven reduced models: a physics-informed neural network (PINN) and a non-intrusive tensorial reduced-order model (TROM
This post was not written by a clanker. Hey guys, I'm a comp sci major who wanted to introduce a cool project I built for quantizing models to ternary (1.58 bit) with as minimal of loss as possible, a
arXiv:2607.28624v1 Announce Type: new Abstract: We introduce PhiZero, a physical world model built around physical language, a compact discrete representation of world-state transitions. Existing phys
arXiv:2607.28478v1 Announce Type: new Abstract: As large language models (LLMs) continue to advance in complex reasoning tasks, they have learned to heavily prioritize explicit conditions provided in
arXiv:2607.27023v1 Announce Type: new Abstract: Evaluating large generative models across benchmarks is time-consuming and computationally expensive. This drives the need for methods that can estimate
arXiv:2607.26455v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated strong capabilities in knowledge acquisition and reasoning, yet their ability to retain previously acquir
Google DeepMind says the latest version of its Gemini Robotics AI model can 'control entire humanoid robots.' While the previous model focused on controlling a humanoid robot's upper body, Gemini Robo
arXiv:2607.26751v1 Announce Type: new Abstract: State-of-the-art intracortical brain-to-text systems pair a neural-sequence phone decoder with an external language model. Two design axes remain undere
arXiv:2607.26754v1 Announce Type: new Abstract: Recent game world models can generate visually realistic and interactive environments conditioned on player actions. However, games are not defined by p
arXiv:2607.24758v1 Announce Type: new Abstract: Large language models are capable of recognizing evaluation contexts and altering their behavior to reflect evaluator expectations rather than typical d
arXiv:2502.11049v3 Announce Type: replace Abstract: Automated Facial Expression Recognition (FER), involves two critical aspects: data and model design. Both significantly influence bias and fairness
arXiv:2607.26015v1 Announce Type: new Abstract: Syntactic convergence (the tendency of speakers to adapt in language towards the grammatical profiles of their interlocutors) is a well-documented featu
Model + harness. We have barely begun to understand the best ways to do harness engineering. A huge amount of untapped potential even without models getting better (but models are getting better) Turn
arXiv:2607.24787v1 Announce Type: new Abstract: Sparse Mixture-of-Experts (MoE) models expand foundation model capacity through conditional expert activation, but their full expert pools remain diffic
arXiv:2607.23183v1 Announce Type: cross Abstract: Neurological disorders are a leading cause of global disability and are increasingly linked to environmental chemical exposures. Yet neurotoxicity ass
arXiv:2607.22600v1 Announce Type: new Abstract: Information visualizations are widely used to communicate patterns, trends, and outliers, yet deceptive design choices-such as truncated or inverted axe
arXiv:2607.23322v1 Announce Type: new Abstract: Instruction tuning has become the standard method for adapting large language models to follow human intent, yet existing instruction datasets are domin
K3 is a language model that entered the top five most‑liked models on Hugging Face merely 24 hours after its release. The achievement surprised many, placing it ahead of prominent models such as Llama
Mage-VL is a codec-native, proactive-streaming multimodal foundation model for image and video understanding, whose visual encoder is trained entirely from scratch at a compact 4B scale. It targets a
arXiv:2607.22583v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved widespread adoption because of their strong reasoning and query-response capabilities. However, deploying the
arXiv:2607.24306v1 Announce Type: new Abstract: Diffusion language models enable flexible arbitrary-order generation, but existing sampling methods are mostly designed for early masked diffusion model
arXiv:2607.22708v1 Announce Type: new Abstract: Deploying a vision-language model with full UI understanding on end devices has long been trapped between accuracy and efficiency: on one side is the ac
arXiv:2607.24032v1 Announce Type: new Abstract: Generative-AI evaluations can become historical before publication, yet calendar age does not affect every conclusion equally. This paper has two linked
arXiv:2607.23915v1 Announce Type: cross Abstract: We examine how prompt tone affects both accuracy of the LLM answers and inference cost as reflected in output-token consumption. Experiments were perf
arXiv:2601.05148v2 Announce Type: replace Abstract: Pathology foundation models substantially advanced the possibilities in computational pathology --- yet tradeoffs in terms of performance, robustnes
We've joined the alliance. Open-weight models will ensure that we live in a safer digital world, and that America does not get left behind Attackers have frontier AI. Defenders need a frontier AI ecos
Announcing the Claude Code-compatible interface for our new Fugu-Ultra v1.1! 🐡 Put a dynamically coordinated team of frontier models to work inside the coding workflow you already know. Instead of rel
arXiv:2603.11838v2 Announce Type: replace Abstract: Large language models pretrained on internet-scale data risk lookahead bias in forecasting tasks, as they may have already seen the true outcome dur
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter. AI will transform every industry, power every company, and be built by every country. Open models strengthen safety an
arXiv:2607.18218v2 Announce Type: replace-cross Abstract: Foundation models have emerged as a driving force in computational pathology, with the potential to transform cancer diagnosis, prognosis, and
Learning causality from internet videos in latent space first, and then using RL to teach the foundation model how to act. This approach is 30× cheaper than Gemini 3.1 Flash on pretraining and achieve
Open models for the win! For my first post, I’m sharing a letter @NVIDIA signed on why open models matter. AI will transform every industry, power every company, and be built by every country. Open mo
The Washington Post processed 1.79B input tokens per month through Together AI, running open models like Llama and Mistral in production with predictable costs and full control over the model stack. T
arXiv:2607.21090v1 Announce Type: cross Abstract: We propose a Reinforcement Learning (RL) method to directly optimize the faithfulness of self-explanations - the extent to which a model's generated r
Arize and Fireworks benchmarked 10 AI models across 2,400 agent runs. Learn why cost per successful task beats token price for model evaluation and routing. The post Cost per successful task: Benchmar
Huge launch from @tryramp. Different steps in an agent workflow can use different models. This can help reduce costs significantly without sacrificing performance. Model routing will become a core par
Who’s Afraid of Chinese Models? Interesting proposal from Ben Thompson that both addresses the hypocrisy of labs outlawing distillation against their models despite training on unlicensed data, and co
arXiv:2607.12884v1 Announce Type: new Abstract: Patients seeking medical information often ask questions that embed incorrect assumptions or misconceptions. In such cases, safe medical communication r
arXiv:2607.13017v1 Announce Type: cross Abstract: World Action Models (WAMs) are able to leverage pretrained video generators for both world modeling and action prediction. However, directly leveragin
arXiv:2607.11889v1 Announce Type: cross Abstract: Large language models trained on unrestricted internet corpora inevitably embed information from the future, introducing lookahead bias that compromis
U.S. open‑source AI models are rapidly gaining popularity, with NVIDIA’s newest model, **Nemotron Ultra**, becoming a prominent entry on the Ollama platform. On Ollama, Nemotron Ultra is quickly scali