Are 1B LLMs Going Away in 2026?
I don't know much about llms aside from downloading them through a frontend and running them on my laptop or potato phone. Google released gemma 4, but unlike gemma 3, there isn't a 1b model this time
Knowledge catalogue
I don't know much about llms aside from downloading them through a frontend and running them on my laptop or potato phone. Google released gemma 4, but unlike gemma 3, there isn't a 1b model this time
arXiv:2511.10101v2 Announce Type: replace Abstract: Background: both embodied intelligence and developmental morphogenesis depend on a division of labour between centralized guidance and distributed m
arXiv:2607.27512v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in multi-agent environments. However, the processes by which beliefs form and propagate among int
arXiv:2603.05950v2 Announce Type: replace Abstract: Visual token reduction is critical for accelerating Vision-Language Models (VLMs), since visual inputs are represented as token sequences that intro
arXiv:2607.26723v1 Announce Type: cross Abstract: Inversion-based watermarking is a promising approach to authenticate diffusion-generated images, yet practical use is bottlenecked by inversion that i
arXiv:2607.28077v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models, but prompt groups with identical rol
arXiv:2602.05307v2 Announce Type: replace Abstract: Large reasoning models (LRMs) achieve strong performance by producing long chains of thought, but their inference costs are high and often generate
arXiv:2607.26220v1 Announce Type: cross Abstract: Context: Large Language Models (LLMs) offer natural-language flexibility for automated requirements elicitation but frequently generate structurally i
arXiv:2607.28627v1 Announce Type: new Abstract: Long visual context poses a challenge for vision-language models: performance degrades as the number of distractors grows, and processing all tokens at
arXiv:2601.16651v3 Announce Type: replace Abstract: Gradient-based methods for instance-based explanation for large language models (LLMs) are hindered by the immense dimensionality of model gradients
arXiv:2503.15293v2 Announce Type: replace Abstract: Object detection models are vulnerable to backdoor attacks, where attackers poison a small subset of training samples by embedding a predefined trig
arXiv:2607.28166v1 Announce Type: new Abstract: Diffusion language models (DLMs) expose a provisional prediction at every denoising step, creating an opportunity for generation-time early exit that st
Atomic Chat signed the Open Weights letter! We believe everyone should be able to run AI on their own device. When a model is open, thousands of teams fine-tune it, quantize it and build new tools on
arXiv:2607.26700v1 Announce Type: new Abstract: The Mpox outbreak remains a serious public health issue, with the WHO (World Health Organization) reporting increasing cases in some regions. Research o
arXiv:2607.26752v1 Announce Type: new Abstract: Medical world models aim to learn a latent state of patient or organ physiology and a transition function that forecasts how that state evolves under in
arXiv:2607.26452v1 Announce Type: cross Abstract: World models must learn the joint dynamics of states, actions, events, and observations, yet existing video, robotics, and simulation datasets usually
arXiv:2607.26580v1 Announce Type: new Abstract: With the increase in the number of cases related to respiratory diseases, there is an urgent need to detect them early and diagnose them accurately. Con
arXiv:2607.26933v1 Announce Type: cross Abstract: Federated Learning (FL) is vulnerable to backdoor attacks because of its distributed nature in edge computing scenarios. Existing defense methods show
arXiv:2607.26801v1 Announce Type: new Abstract: Federated learning (FL) enables collaborative learning over decentralized data silos without centralizing raw data. However, heterogeneous local archite
arXiv:2607.26921v1 Announce Type: new Abstract: YOLO-Pose models provide efficient keypoint localization, but do not quantify the associated spatial uncertainty. We introduce a lightweight post-hoc pr
arXiv:2604.05175v2 Announce Type: replace-cross Abstract: We consider constrained ergodic resource optimization in wireless networks with graph-structured interference. We train a diffusion model poli
I'm actually fairly bearish on frontier lab valuations. I've never seen the reasons articulated to my satisfaction, so before I go to sleep, I wanted to quickly jot down my thinking here. The basic is
arXiv:2512.24149v2 Announce Type: replace Abstract: The world is governed by both physical laws and affective dynamics. Physical laws govern state transitions, while affective dynamics shape human act
arXiv:2409.07314v4 Announce Type: replace Abstract: While Large Language Models (LLMs) achieve superhuman performance on standardized medical licensing exams, these static benchmarks have become satur
arXiv:2607.27092v1 Announce Type: new Abstract: We analyze whether language models of size ~100B have a representation of the night sky map that is decodable from their residual stream. We find that m
Hello, I have a friend who's using various local LLM's like qwen3.6 27B, 35b-a3b, North Mini Code, and qwen2.5-vl-7b (just for vision). They have a use case where they're trying to have an LLM drive a
arXiv:2607.27205v1 Announce Type: new Abstract: Vision-language-action (VLA) models commonly adopt an LLM-centric V o L o A pathway, where visual observations are projected into the representation spa
arXiv:2607.26348v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as synthetic users, stand-ins for human respondents whose simulated answers feed product, policy, and
arXiv:2607.25516v1 Announce Type: new Abstract: Vision-language-action (VLA) models predict sequential actions to execute tasks specified by language instructions, conditioned on visual observations a
arXiv:2607.24807v1 Announce Type: cross Abstract: We develop the mathematical and physical formulation of cognitive cost optimization that underlies the path-integral model of consciousness. The goal-
arXiv:2607.25244v1 Announce Type: new Abstract: Foundation models for 12-lead electrocardiograms (ECGs) transfer well across clinical tasks, but the physiological knowledge encoded in their representa
arXiv:2507.00411v2 Announce Type: replace Abstract: Learning from ambiguous labels is a long-standing problem in practical machine learning applications. The purpose of partial label learning (PLL) is
arXiv:2607.25069v1 Announce Type: cross Abstract: Automated verification of numerical claims is a challenging problem, as it requires both language understanding and quantitative reasoning. This paper
arXiv:2607.25791v1 Announce Type: new Abstract: Falls represent a critical public health challenge, and accurate detection of the impact moment when an individual hits the ground is crucial for timely
arXiv:2607.24834v1 Announce Type: cross Abstract: Objective. Electroencephalography (EEG) foundation models (FMs) are trained to reconstruct or contrastively align short patches, then pooled into a fi
The Test Bench Setup I tested running a massive 1.56TB Mixture-of-Experts (MoE) checkpoint (96 shards, 93 layers, 896 experts/layer, ~4.46 bits/param MXFP4) on a budget gaming laptop. Laptop: HP Victu
arXiv:2602.10282v2 Announce Type: replace Abstract: Large language models (LLMs) have shown potential in identifying qualitative causal relations, but their ability to perform quantitative causal reas
arXiv:2603.27176v2 Announce Type: replace Abstract: Lesion detection, symptom tracking, and visual explainability are central to real-world medical image analysis, yet current medical Vision-Language
arXiv:2607.25232v1 Announce Type: new Abstract: Digital phenotyping (DP) using smartphones and wearable devices has shown considerable potential for mental health monitoring. However, progress remains
arXiv:2607.24789v1 Announce Type: cross Abstract: We present NEXT (Next-interest EXploration Transformer), a reasoning-driven video recommendation framework that reasons over the video a user has just
arXiv:2603.25469v2 Announce Type: replace Abstract: A growing body of literature has focused on predicting wildfire occurrence using machine learning methods, capitalizing on high-resolution data and
arXiv:2607.25157v1 Announce Type: new Abstract: Discrete masked diffusion language models support bidirectional generation and infilling, but adapting pretrained autoregressive (AR) transformers requi
arXiv:2607.25337v1 Announce Type: new Abstract: Joint-Embedding Predictive Architectures (JEPAs) learn world models by predicting in representation space rather than reconstructing pixels, making them
OpenAI announced the release of GPT‑5.6 Sol after deployment, incorporating optimizations across its stack to enhance run‑time efficiency. The update delivers a roughly 20 % reduction in serving costs
arXiv:2607.25393v1 Announce Type: new Abstract: Histopathological examination primarily relies on hematoxylin and eosin (H&E) and immunohistochemistry (IHC) staining. Although IHC provides critical mo
arXiv:2607.22598v1 Announce Type: cross Abstract: Educational chatbots powered by large language models (LLMs) show promising effects on learning outcomes, yet most systems delegate pedagogical decisi
As the United States government debates new artificial intelligence rules and regulations, Anthropic PBC and Nvidia Corp. are drawing a line against blanket bans on open-weight models, urging regulato
arXiv:2607.23972v1 Announce Type: new Abstract: Color Fundus Photography (CFP) is a primary non-invasive imaging modality for large-scale screening of ophthalmic and systemic diseases. Existing survey
arXiv:2607.24425v1 Announce Type: new Abstract: Large language models place structured concepts on geometrically faithful manifolds: weekdays lie on a circle, months on another, usually taken to be a
arXiv:2607.22609v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used for decision-making and reasoning tasks, yet their potential as controllers for physical systems rema
arXiv:2607.22652v1 Announce Type: new Abstract: Recent research has explored the integration of knowledge graphs (KGs) with large language models (LLMs) to enhance their performance on downstream know
arXiv:2606.16586v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) remain unreliable on fine-grained visual perception, even when high-resolution inputs preserve the necessar
arXiv:2607.22586v1 Announce Type: new Abstract: Key-Value (KV) caching is essential for efficient inference in multimodal large language models (MLLMs), yet its memory footprint grows linearly with co
arXiv:2607.22925v1 Announce Type: cross Abstract: A key question for AI safety is whether a language model expresses all of its reasoning in its output tokens. We demonstrate a concrete failure mode w
arXiv:2607.18553v2 Announce Type: replace-cross Abstract: Can a language model read the quality of its ongoing computation, and can an external intervention turn that readout into better outcomes? We
arXiv:2510.06128v2 Announce Type: replace Abstract: Tokenization forms the basis of multilingual language models, yet existing methods often limit cross-lingual transfer by mapping semantically equiva
arXiv:2607.23447v1 Announce Type: new Abstract: Predicting cellular responses to unseen chemical perturbations is challenging due to unknown targets and mechanisms, high-dimensional expression respons
arXiv:2607.24447v1 Announce Type: new Abstract: On-Policy Self-Distillation (OPSD) uses privileged information available only to the teacher to provide dense token-level supervision on trajectories ge
arXiv:2607.23238v1 Announce Type: new Abstract: Masked image modeling has become a dominant paradigm for SAR pre-training, yet the design of the reconstruction target remains fundamentally unsettled.
arXiv:2607.23626v1 Announce Type: cross Abstract: Radiomic biomarkers derived from magnetic resonance imaging (MRI) have been widely investigated as non-invasive tools for tumor characterization and p