b9294
b9294 is a llama.cpp release that includes improvements to OpenCL kernels for AMD Adreno GPUs, specifically generalizing MoE (Mixture of Experts) kernel implementations. The release provides pre-built
Knowledge catalogue
b9294 is a llama.cpp release that includes improvements to OpenCL kernels for AMD Adreno GPUs, specifically generalizing MoE (Mixture of Experts) kernel implementations. The release provides pre-built
b9296 is a build release of llama.cpp , the C/C++ implementation for efficient large language model inference. As an intermediate build in the llama.cpp release cycle, it likely includes recent bug fi
The search results show information about llama.cpp releases but don't contain specific details about build b9297. Based on the context and numbering (b9297 is a recent build after b9294), b9297 is li
arXiv:2510.10129v2 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) systems suffer from severe time-to-first-token (TTFT) bottlenecks due to long input sequences. Existing KV cach
arXiv:2605.22221v1 Announce Type: new Abstract: Backtracking search underlies classical constraint solvers, planners, and theorem provers. Recent transformer-based reasoning systems explore search tre
arXiv:2605.22703v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a central paradigm for scaling LLM reasoning, yet its optimization often suffers fr
arXiv:2605.22379v1 Announce Type: cross Abstract: With the advancement of science and technology, the importance of emotion research has become increasingly evident. Electroencephalography (EEG)-based
This post showcases a dark fantasy art series created using a heavy ink style, likely featuring gritty, detailed artwork with gothic or fantasy themes. The creator shared the work on r/StableDiffusion
arXiv:2605.22266v1 Announce Type: new Abstract: Federated learning enables collaborative training across distributed clients with heterogeneous data, but such heterogeneity often leads to unstable upd
arXiv:2605.22531v1 Announce Type: new Abstract: There is a gap between the theoretical foundations of disentanglement and the practice of modern representation learning. Existing theoretical framework
arXiv:2605.22506v1 Announce Type: cross Abstract: Federated learning faces increasing threats from model poisoning attacks, which harms its application to improve privacy. Existing defense methods typ
arXiv:2605.22537v1 Announce Type: new Abstract: Reinforcement learning methods such as GRPO have seen great popularity in LLM post-training. In GRPO, models produce completions to a set of prompts, wh
ForgeFlash is a minimalist frontend interface designed for Stable Diffusion WebUI Forge that emphasizes simplified workflows inspired by Fooocus while maintaining direct access to essential advanced c
This post describes a custom Stable Diffusion graphical interface optimized for older graphics cards, specifically the GTX 1060, incorporating features like zero-copy ADetailer (a detail enhancement t
This post describes a project where a developer created a searchable database of 49,000 sample images featuring Anima AI characters and artists. The tool appears to be designed to help users find and
arXiv:2605.22410v1 Announce Type: new Abstract: Spectral clustering largely depends on the affinity graph, yet constructing a graph that preserves reliable local connectivity while adapting to heterog
Ollama can be installed on Mac by downloading the application and placing it in the Applications folder, after which you use Terminal to run commands that download and launch models . The best model c
arXiv:2605.21851v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards has become the standard recipe for improving LLM reasoning, but the dominant algorithm GRPO assigns a sin
arXiv:2507.20268v3 Announce Type: replace Abstract: Wireless indoor localization using predictive models with received signal strength information (RSSI) requires proper calibration for reliable posit
arXiv:2605.22814v1 Announce Type: new Abstract: Exploration is a prerequisite for learning useful behaviors in sparse-reward, long-horizon tasks, particularly within 3D environments. Curiosity-driven
arXiv:2605.22488v1 Announce Type: new Abstract: Structured prompts require integrating components according to task-relevant relations. How a network implements this integration is often hard to judge
A realistic fighter vs a fully anime-styled character, generated with Seedance 2.0 using detailed prompting and reference images. This workflow combines Seedance with an LLM-powered prompting setup de
Ollama has added MLX support for Qwen3.5 and Gemma 4 , with both MLX and NVFP4 quantization variants available for these models. There have been issues with broken Qwen3.6 NVFP4 generation on Ollama ,
arXiv:2605.22631v1 Announce Type: new Abstract: Accurately reconstructing full-body poses from sparse head and hand trajectories is a foundational challenge for immersive AR/VR telepresence. Current m
arXiv:2605.21932v1 Announce Type: new Abstract: Multi-Robot Task Allocation (MRTA) is a central challenge in decentralized multi-agent systems, where teams of robots must cooperatively assign and exec
llama.cpp build b9276 introduces support for hybrid DNA tokenization with new pre-type dispatching and tokenizer implementations, alongside fixes for VRAM leaks in Multi-Token Prediction (MTP) models
Release b9283 of llama.cpp includes updates to the BPE tokenizer with new pre-type dispatch and Carbon-3B DNA tokenizer support, along with refactoring of tokenizer helpers to follow project conventio
B9286 introduces support for HybridDNATokenizer as its own vocabulary type (LLAMA_VOCAB_TYPE_HYBRIDDNA), enabling proper tokenization of DNA sequences by routing raw text through a DNA-aware splitter
Release b9292 of llama.cpp fixes a memory leak in the server context where speculative decoder, draft context, and draft model were not properly freed during destroy(), causing VRAM leaks on sleep/res
arXiv:2605.22373v1 Announce Type: cross Abstract: Safety classifiers are essential safeguards within generative AI systems, filtering harmful content or identifying at-risk users when interacting with
arXiv:2605.21712v1 Announce Type: new Abstract: Transportation safety analysis requires integrating crash records, roadway attributes, and geospatial data through GIS-based workflows, but access remai
arXiv:2605.22422v1 Announce Type: new Abstract: Table structure recognition (TSR) requires both table-level coherence (row/column counts, headers, spanning cells) and precise separator localization. W
arXiv:2605.22209v1 Announce Type: new Abstract: Video Capsule Endoscopy (VCE) poses a challenging multi-label temporal classification problem, requiring simultaneous localization of 8 anatomical regio
arXiv:2605.20704v1 Announce Type: cross Abstract: Autonomous AI agents that spawn sub-agent swarms create a safety gap: existing credential revocation mechanisms, OAuth~2.0 introspection, OCSP, and W3
arXiv:2605.21858v1 Announce Type: new Abstract: Large language models (LLMs) have recently shown strong potential in modeling relational structures. However, existing approaches remain fundamentally g
Pixal3D is a Tencent image-to-3D model that generates high-fidelity 3D assets from a single image by explicitly lifting pixel features into 3D through back-projection to establish direct pixel-to-3D c
This Reddit post discusses a common ComfyUI issue where users cannot locate spatial upscaler models for img2vid workflows. The problem typically stems from placing spatial upscaler models in the wrong
arXiv:2511.07885v4 Announce Type: replace-cross Abstract: Large language model (LLM) queries are predominantly processed by frontier models in centralized cloud infrastructure. Demand growth strains t
arXiv:2605.21652v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have significantly advanced medical visual question answering, yet their performance in ultrasound remains suboptimal. In
LM Studio version 0.4.14 introduced support for MTP (likely Model Transfer Protocol or a similar feature), with audio functionality enabled. This update was announced on the official LM Studio X/Twitt
The AI PC is being fundamentally redefined as agentic workloads push the boundaries of what local compute can deliver — and as runaway cloud token costs force enterprises to rethink where inference ac
arXiv:2605.22581v1 Announce Type: new Abstract: Many public buildings provide floorplans with a 'you are here' indicator to help visitors orient themselves. Floorplan localization seeks to computation
arXiv:2605.22658v1 Announce Type: new Abstract: While large language models provide strong compositional reasoning, existing reasoning segmentation pipelines fail to transparently connect this reasoni
arXiv:2605.22476v1 Announce Type: cross Abstract: Entity tracking requires maintaining and updating latent states for entities and attributes over long sequences. Recent task-specific attention operat
ComfyUI posted a welcoming message to new community members on X (Twitter), using a celebratory tone to encourage engagement and belonging within their user base. The post reflects community-focused v
A discussion of progressive training approaches for diffusion models where training begins at lower scale factors and progressively increases to target scale factors, leveraging previously trained mod
ID-LoRA (Identity-Driven In-Context LoRA) enables identity-preserving audio-video generation by simultaneously dictating the scene's visual content, environmental acoustics, and speaking style while p
arXiv:2605.20742v1 Announce Type: new Abstract: With the rapid proliferation of electric vehicles, the safety and reliability of lithium-ion batteries have become critical concerns. Effective anomaly
arXiv:2605.21633v1 Announce Type: cross Abstract: When the blood supply to the brain is obstructed by a clot, oxygen delivery to brain tissues becomes insufficient, leading to cellular necrosis. In he
arXiv:2605.22823v1 Announce Type: new Abstract: Video Large Language Models (Video-LLMs) have made rapid progress on temporal video understanding, yet many fail at a basic perceptual primitive: signed
AceStep 1.5 is a LoRA (Low-Rank Adaptation) model trained on the first two albums of the Modern Talking band, designed for use with Stable Diffusion image generation. This LoRA likely enables the gene
arXiv:2605.21421v1 Announce Type: new Abstract: Motion capture is the gold standard for measuring human movement, but clinical use remains limited by cost, technical complexity, and privacy concerns.
Release b9255 of llama.cpp features a Hexagon HMX quantized matmul rework (#23368), including updates to debug logging, dequantization logic using HVX vectors, removal of non-pipelined quantization op
llama.cpp release b9257 optimizes operations in the Vulkan IM2COL shader and improves code formatting . The release includes prebuilt binaries for multiple platforms including macOS, Linux, Android, a
Release b9259 of llama.cpp fixes a nullptr crash in the speculative common module's get_devices_str function (PR #23386). The fix addresses an issue where ggml_backend_dev_by_name appends a nullptr se
Release b9260 of llama.cpp includes OpenCL backend refactoring that improves initialization, GPU identification, and performance by caching global memory size in device context. llama.cpp enables LLM
Release b9263 of llama.cpp includes a merge of HunyuanOCR into HunyuanVL with fixes to OCR vision precision. The update consolidates OCR functionality into the HunyuanVL projector while maintaining co
b9264 is a llama.cpp release that includes improvements to HunyuanVL model support, merging HunyuanOCR functionality and fixing vision precision issues. This build represents an intermediate developme
b9266 is a release tag for llama.cpp, a project for LLM inference in C/C++. This release represents a specific build version of the llama.cpp framework, which enables LLM inference with minimal setup
Release b9270 adds support for the HybridDNATokenizer used by the Carbon-3B model family, implementing a new BPE pre-type for tokenizing DNA sequences. The tokenizer handles DNA k-mers with fixed 6-me