b9273
Release b9273 of llama.cpp introduces support for the HybridDNATokenizer with new pre-type and dispatched tokenization logic, alongside pure helper functions for DNA k-mer processing and conversion ut
Knowledge catalogue
Release b9273 of llama.cpp introduces support for the HybridDNATokenizer with new pre-type and dispatched tokenization logic, alongside pure helper functions for DNA k-mer processing and conversion ut
Release b9275 of llama.cpp includes optimization of the Metal concat kernel and fixes to the GGML_OP_SET kernel threads . The release extends test coverage for copy operations with different source an
arXiv:2605.20293v1 Announce Type: new Abstract: Predictive coding (PC) offers a local and biologically grounded alternative to backpropagation in the training of artificial neural networks, yet to dat
arXiv:2605.20270v1 Announce Type: new Abstract: A local specialist LLM, fine-tuned with reinforcement learning from verifiable rewards (RLVR) on operator-local data, is installed in a regulated organi
This post demonstrates using reference images with control maps in FLUX.2 instead of training LoRAs, allowing users to fuse a reference image with pose or edge control data while preserving identity a
This post describes a user's experience building a local AI agent using Ollama, an open-source tool for running large language models locally. The article likely covers the practical implementation st
arXiv:2603.13419v2 Announce Type: replace Abstract: Diffusion models generalize well in practice. However, an optimal diffusion model fully memorizes the training data and therefore fails to generaliz
arXiv:2601.03019v4 Announce Type: replace-cross Abstract: DNA language models are increasingly used to represent genomic sequence, yet their effectiveness depends critically on how raw nucleotides are
arXiv:2605.20717v1 Announce Type: cross Abstract: This work presents E-ReCON, a 16 Kb energy and resource-efficient digital compute-in-memory (DCIM) macro based on a compact 3T1R ReRAM bitcell for edg
arXiv:2605.20569v1 Announce Type: new Abstract: Hyperspectral imagery encodes rich material properties that can improve tracking robustness under appearance ambiguity, illumination change, and backgro
arXiv:2602.16608v2 Announce Type: replace Abstract: Transformer models achieve state-of-the-art performance across domains and tasks, yet their deeply layered representations make their predictions di
arXiv:2602.13485v2 Announce Type: replace Abstract: Networks of modern industrial systems are increasingly monitored by distributed sensors, where each system comprises multiple subsystems generating
arXiv:2503.19708v2 Announce Type: replace-cross Abstract: Urban microclimate, encompassing wind and temperature fields shaped by building geometry, significantly impacts energy consumption, pedestrian
arXiv:2605.21171v1 Announce Type: new Abstract: Ternary Vision Transformers offer substantial model compression, however state-of-the-art methods only ternarize the encoder layers, leaving patch embed
arXiv:2605.20891v1 Announce Type: new Abstract: Multimodal survival prediction, a crucial yet challenging task, demands the integration of multimodal medical data (eg Whole Slide Images (WSIs) and Gen
A developer added a visual fold feature to ComfyUI, a node-based visual programming environment used for Stable Diffusion workflows. This feature helps organize complex workflows with many nodes by al
I'm excited about the new @amd Ryzen AI Halo because we need more local hardware for AI builders! There's something fun and exciting about building on your own machines rather than sending to the clou
arXiv:2605.20682v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have shown remarkable capability in bridging visual perception and textual reasoning, enabling zero-shot unders
arXiv:2605.20595v1 Announce Type: new Abstract: Dense low-altitude aerial operations require more than pre-flight route coordination and last-resort collision avoidance. Once aircraft are airborne, di
arXiv:2605.20784v1 Announce Type: cross Abstract: Spatial reasoning requires both location-bound computation and location-invariant structure: agents must make local moves while preserving route, obje
LoRA models can be used with image-to-video workflows to create automatic lip-syncing effects by controlling camera movements and instructing the model to synchronize audio with mouth movements. This
arXiv:2605.21086v1 Announce Type: new Abstract: While Large Language Models (LLMs) are increasingly integrated into in-vehicle conversational systems, identifying the optimal model remains challenging
arXiv:2605.20866v1 Announce Type: new Abstract: Communication is a major bottleneck in distributed learning, especially in large-scale settings and in federated learning environments with slow links.
LTX-2.3 is a multimodal video generation model developed by Lightricks that generates synchronized audio and video in a single forward pass at resolutions up to 4K at 50 frames per second. The model i
arXiv:2605.20299v1 Announce Type: new Abstract: Generative sequence models are often trained to plan motion in physical domains, from robotics to mechanical simulations. When constructing a dataset to
arXiv:2605.20273v1 Announce Type: new Abstract: Online model editing for multimodal large language models (MLLMs) requires assimilating a stream of corrections under tight compute and memory budgets.
arXiv:2509.17931v2 Announce Type: replace Abstract: Accurate multi-needle localization in intraoperative CT images is crucial for optimizing seed placement in pelvic seed implant brachytherapy. Howeve
arXiv:2605.20276v1 Announce Type: new Abstract: The global deployment of edge intelligence operates across heterogeneous legal frameworks. While some regions permit centralized learning (CL) via cloud
On prem. I'm excited about the new @amd Ryzen AI Halo because we need more local hardware for AI builders! There's something fun and exciting about building on your own machines rather than sending to
arXiv:2605.21484v1 Announce Type: new Abstract: Discrete diffusion models excel at visual synthesis but rely on slow, iterative decoding. Existing single-step distillation methods attempt to bypass th
arXiv:2509.09946v2 Announce Type: replace Abstract: Multi-Target Multi-Camera Tracking (MTMC) is an essential computer vision task for automating large-scale surveillance. With camera calibration and
arXiv:2605.21322v1 Announce Type: new Abstract: Federated Learning (FL) enables collaborative model training without centralizing data. However, real-world deployments must simultaneously address stat
arXiv:2605.20818v1 Announce Type: new Abstract: In this report, we present our champion solutions for the Natural Language Queries and GoalStep tracks of the Ego4D Episodic Memory Challenge at CVPR 20
arXiv:2605.20941v1 Announce Type: new Abstract: We present PaintCopilot, a co-creative neural painting assistant that models painting as an open-ended autoregressive artistic behavior conditioned on e
arXiv:2605.20295v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed on mobile devices, where Neural Processing Units (NPUs) necessitate fully static quantization for
arXiv:2605.21099v1 Announce Type: new Abstract: Accurate estimation of the Angle of Progression (AoP) from intrapartum transperineal ultrasound is critical for objective assessment of labor progressio
arXiv:2605.21237v1 Announce Type: new Abstract: Cardiac motion over a cardiac cycle is crucial for quantifying regional function and is strongly affected by cardiovascular diseases. Since temporally d
arXiv:2605.20681v1 Announce Type: cross Abstract: Distributed principal component analysis (PCA) produces node-level estimates of both a mean vector and a principal subspace. Robustly aggregating thes
arXiv:2605.21190v1 Announce Type: new Abstract: Despite the generative capabilities of diffusion and flow models, real-image editing remains constrained by a persistent trade-off between semantic edit
arXiv:2501.15151v5 Announce Type: replace Abstract: Spiking Neural Networks (SNNs) are the third generation of neural networks. They have gained widespread attention in object detection due to their l
arXiv:2605.20760v1 Announce Type: new Abstract: Automated segmentation of the vertebral column in Computed Tomography (CT) scans is a prerequisite for pathological assessment and surgical planning. Ho
Stable Audio 3.0 is now Day-0 supported in ComfyUI. Open-weight music models (fully licensed data)—from quick SFX and short tracks to longer, more musical pieces—inside the workflows you already use.
arXiv:2603.26603v2 Announce Type: replace-cross Abstract: The migration of Large Language Models (LLMs) from cloud clusters to edge devices promises enhanced privacy and offline accessibility, but thi
arXiv:2605.20635v1 Announce Type: new Abstract: This paper proposes a general machine learning framework called the localization method, which is fundamentally built on two core concepts: localization
The NVIDIA GB200 NVL72 is a rack-scale GPU supercomputer leveraging Blackwell architecture with NVLink switches for high-density computing , and topology-aware block scheduling in Slurm can align larg
v0.30.0-rc22 is a pre-release version of Ollama that changes the architecture to directly support llama.cpp instead of building on top of GGML, and allows for compatibility with GGUF file format. MLX
We have an arxiv paper up describing the work in more detail here: https://arxiv.org/abs/2605.20706. Also want to call out that there is even more room for improvement, some recent updates to wllama b
Latest face swapping methods are expanding toward 3D-consistent video swapping and identity-preserving generation systems that control portraits and clips . Improvements in diffusion models and neural
arXiv:2605.20275v1 Announce Type: new Abstract: Existing deep learning approaches for wearable fall detection systems rely on self-attention mechanisms that impose quadratic computational overhead, di
arXiv:2501.09203v2 Announce Type: replace Abstract: Visual-Spatial Systems has become increasingly essential in concrete crack inspection. However, existing methods often lacks adaptability to diverse
A user-generated AI image creation showcasing a deep space patrol scene rendered in 4K quality using Stable Diffusion, a popular text-to-image generation model. The post was shared on the r/StableDiff
arXiv:2605.19234v1 Announce Type: cross Abstract: The rapid emergence of AI technologies is reshaping translation practices and theory across the board. This paper deals with the impact of AI in langu
Stability AI announced the launch of Stable Audio 3, a family of three AI music models and one audio-based special effects model. Most of these releases are 'open weight' models trained on licensed tr
Build b9239 is a llama.cpp release that includes a fix for the --fit verbosity flag when used with --verbosity 4 . The release provides compiled binaries for multiple platforms including macOS (Apple
b9240 is a release of llama.cpp that includes a fix for the --help option related to the --verbosity flag . The release provides prebuilt binaries across multiple platforms including macOS (Apple Sili
b9244 is an intermediate build release of llama.cpp, a C/C++ implementation framework for running large language models with GGUF format support. The release includes pre-compiled binaries for multipl
llama.cpp release b9245, published on May 20, 2026, includes a CUDA optimization for RDNA3 Q6_K MMVQ performance tuning. The release provides pre-built binaries for multiple platforms including macOS,
B9251 is a build identifier for a release in the llama.cpp project, which enables LLM inference with minimal setup and state-of-the-art performance on a wide range of hardware - locally and in the clo
b9253 is the latest version of llama.cpp, released on May 20, 2026. Llama.cpp is a project for LLM inference in C/C++. This build includes bug fixes, performance improvements, and features for running
arXiv:2602.09872v2 Announce Type: replace Abstract: Human activity recognition (HAR) on resource constrained devices requires high accuracy across diverse sensor setups. Selective state space models (