b9102: [SYCL] Add OP im2col_3d (#22903)
Release b9102 adds support for the im2col_3d operator to the SYCL backend in llama.cpp, expanding GPU-accelerated tensor operations for Intel GPU inference. The im2col_3d operation converts 3D input v
Knowledge catalogue
Release b9102 adds support for the im2col_3d operator to the SYCL backend in llama.cpp, expanding GPU-accelerated tensor operations for Intel GPU inference. The im2col_3d operation converts 3D input v
B9103 is a recent release of llama.cpp , a C/C++ implementation for large language model inference. This release represents one of the latest versions in the project's active development cycle, buildi
Release b9105 of llama.cpp includes updates to the AllReduce implementation for CUDA, introducing a NCCL-free provider for tensor parallelism that pipelines data transfers and GPU reduction operations
Release b9106 of llama.cpp includes support for asymmetric FlashAttention in scalar, MMQ, and cooperative matrix paths for Vulkan. The release provides builds for multiple platforms including macOS, L
Release b9109 of llama.cpp introduces refinements to the CUDA tensor parallelism AllReduce implementation, including renaming the --allreduce flag to --reduction-provider and updates to NCCL-free AllR
The search results show recent llama.cpp releases and general information but don't contain specific details about the b9110 release. Based on the search patterns and similar recent releases documente
Release b9112 of llama.cpp introduces a NCCL-free AllReduce implementation for LLAMA_SPLIT_MODE_TENSOR using a single-phase CUDA kernel, and adds an --allreduce flag to llama-bench to select between A
Bare Metal: Z-Image Turbo - Flux.2 Klein 9b - Wan 2.2 discusses comparisons between these lightweight AI image generation models, with Flux.2 Klein 9B showing good prompt adherence and natural composi
With 32GB of RAM, 30B Q4 models run comfortably on M5 MacBook Air , and popular choices include Qwen 3.6 27B (requiring 18GB+ VRAM) and Qwen 3.6 9B (lightweight, runs on 8GB) . Ollama's macOS implemen
Ollama versions prior to 0.17.1 contain a critical out-of-bounds read vulnerability (CVE-2026-7482, CVSS 9.1) that allows remote unauthenticated attackers to leak entire process memory, potentially af
arXiv:2604.02753v2 Announce Type: replace Abstract: Open-vocabulary Object Detection (OVOD) enables models to recognize objects beyond predefined categories, but existing approaches remain limited in
arXiv:2605.07844v1 Announce Type: new Abstract: Energy-based learning is a powerful framework for generative modelling, but its training is inherently non-convex, leading potentially to sensitivity to
arXiv:2508.06816v3 Announce Type: replace Abstract: Lesion segmentation, in contrast to natural scene segmentation, requires handling subtle variations in texture and color, frequent imaging artifacts
arXiv:2605.07457v1 Announce Type: new Abstract: Recent text-guided image editing (TIE) models have made remarkable progress, yet edited images still frequently suffer from fine-grained issues such as
arXiv:2602.20338v2 Announce Type: replace Abstract: Chain-of-Thought (CoT) prompting significantly improves reasoning in Large Language Models, yet the temporal dynamics of the underlying representati
arXiv:2605.06280v2 Announce Type: replace Abstract: Recent advancements in image animation have utilized diffusion models to breathe life into static images. However, existing controllable frameworks
arXiv:2510.02371v2 Announce Type: replace-cross Abstract: Smart grids are exposed to passive eavesdropping, where attackers listen silently to communication links. Although no data is actively altered
arXiv:2605.07962v1 Announce Type: new Abstract: Performance evaluation is essential for assessing the quality of machine learning (ML) models and guiding deployment decisions. In federated learning (F
arXiv:2605.07698v1 Announce Type: new Abstract: Grammar-constrained generation is often combined with local vocabulary masking and speculative decoding, but the resulting sampling law is not the gramm
Generate split-stack videos like this in a single click. The basic setup is simple, just upload an image and the system takes care of the rest. Want more control? Switch to graph mode and fine-tune ea
arXiv:2605.07432v1 Announce Type: new Abstract: Chatbots are robots that can communicate with humans using text or voice signals. Legal chatbots improve access to justice, since legal representation a
arXiv:2605.06740v1 Announce Type: cross Abstract: We introduce Geometric Kolmogorov--Arnold Networks (GeoKANs), a family of geometry-aware KAN-type models in which approximation is carried out in lear
arXiv:2605.07536v1 Announce Type: cross Abstract: Detecting stealthy malicious communications from flow logs under benign-only training remains a critical challenge in network security. Malicious comm
arXiv:2605.08074v1 Announce Type: new Abstract: Conformal prediction (CP) provides a distribution-free approach to uncertainty quantification with finite-sample guarantees. However, applying CP to gra
arXiv:2605.07211v1 Announce Type: cross Abstract: Mobile devices face diverse resource constraints and non-IID data class distributions, requiring fast on-device inference for local in-distribution (I
A Reddit discussion comparing HiDream-O1-Image-Dev (the newer, 8B pixel-native distilled model) with ZImage Base, examining stylistic and performance differences between the two text-to-image generati
arXiv:2605.07214v1 Announce Type: new Abstract: Large Language Models have recently emerged as a promising paradigm for automated heuristic design for NP-hard combinatorial optimization problems. Desp
🆕 Hugging Face 🤝 Hermes Agent 🔥 > we added Hermes Agent to local apps: run it locally with any compatible GGUF/MLX model > shipped native traces support for Hermes Agent: visualize your Hermes traces
This post describes a fixed-camera timelapse visualization created using AI (likely Stable Diffusion) that depicts Los Angeles's transformation over 2,000 years, starting from its original state as To
arXiv:2605.07284v1 Announce Type: new Abstract: Recent interpretability work has identified model-internal handles on post-trained behavior, including refusal directions, assistant/persona axes, and s
arXiv:2605.07622v1 Announce Type: new Abstract: Gender bias in large language models has primarily been investigated for English, while languages with grammatical or morphological gender remain compar
arXiv:2605.06955v1 Announce Type: cross Abstract: Denoising score matching (DSM) provides a way to learn data distributions by training a neural network to recover the score function, defined as the g
arXiv:2605.07157v1 Announce Type: new Abstract: We present the first method to directly use a learned continuous Lagrangian to forecast the dynamics of systems governed by partial differential equatio
arXiv:2605.07175v1 Announce Type: cross Abstract: Aging clocks aim to estimate biological age, a measure of physiological state distinct from chronological age, from observable biomarkers, and are wid
arXiv:2605.07637v1 Announce Type: new Abstract: Multi-agent pathfinding (MAPF) is a widely used abstraction for multi-robot trajectory planning problems, where multiple homogeneous agents move simulta
arXiv:2605.07505v1 Announce Type: new Abstract: Developing lightweight, on-device vision-language GUI agents is essential for efficient cross-platform automated interaction. However, current on-device
arXiv:2605.07213v1 Announce Type: new Abstract: Infrared small target detection (IRSTD) remains challenging due to the scarcity of useful target cues and the presence of severe background clutter. Mos
LTX-2.3 is a multimodal video generation model released by Lightricks in March 2026, available in four checkpoint variants including a distilled variant that completes generation in as few as 8 denois
arXiv:2605.07233v1 Announce Type: new Abstract: This work focuses on the question of learning from a large number of devices with each device holding only a single sample of data. Several real-world a
My @aiDotEngineer talk is live: 'On-device Intelligence using MLX' 🎥 Huge thanks to @swyx and the team for having me — hands down the best tech event I've been to. And a shoutout to the community ship
arXiv:2605.07860v1 Announce Type: cross Abstract: Federated Learning (FL) has emerged as a promising paradigm for preserving client data ownership and control over distributed Internet of Things (IoT)
arXiv:2605.06820v1 Announce Type: cross Abstract: Deep learning-based organs/structures-at-risk(OARs) auto-contouring models can improve radiotherapy workflows, but models trained on adult data often
arXiv:2605.07278v1 Announce Type: cross Abstract: A latent world model may achieve accurate short-horizon prediction while still inducing a latent space that is poorly aligned with planning. A key iss
ComfyUI announced early access to a workflow publishing feature that allows users to share their AI image generation workflows on a hub or marketplace platform. This feature enables the community to d
Clem Delangue announced the Reachy Mini robot is operational with audio capabilities enabled, with plans to integrate it with local AI services and a Hermes Agent in the near future. This indicates pr
arXiv:2605.07234v1 Announce Type: cross Abstract: Large language models (LLMs) support long-context inference but suffer from substantial memory and runtime overhead due to Key-Value (KV) Cache growth
arXiv:2510.24736v3 Announce Type: replace-cross Abstract: Generating property-optimized mRNA sequences is central to applications such as vaccine design and protein replacement therapy, but remains ch
arXiv:2605.07530v1 Announce Type: new Abstract: The Danish Technological Institute (DTI) focuses on transferring advanced technologies (including robots) to the industry and the public sector. One key
arXiv:2605.07978v1 Announce Type: new Abstract: Cross-view localization classically asks: where does this ground image lie on the satellite tile? Existing methods are typically limited to 3-DoF estima
arXiv:2605.07022v1 Announce Type: new Abstract: Manually curated biomedical repositories -- spanning bioactivity, genomics, and chemistry -- are expensive to maintain, lag behind primary literature, a
arXiv:2605.07276v1 Announce Type: new Abstract: Code-agent RL often receives weak feedback: rollout-time signals are reliable and executable, but capture only necessary or surface conditions for task
arXiv:2605.06610v2 Announce Type: replace-cross Abstract: Sparse Autoencoders (SAEs) have become an important tool in mechanistic interpretability, helping to analyze internal representations in both
arXiv:2605.07390v1 Announce Type: new Abstract: Generative models have achieved success in producing apparently coherent 2D videos, but remain challenging in the physical world due to lack of 4D spati
arXiv:2605.07980v1 Announce Type: new Abstract: These notes introduce the theory of susceptibilities as developed in [arXiv:2504.18274, arXiv:2601.12703] for interpreting neural networks. The suscepti
arXiv:2509.26524v3 Announce Type: replace-cross Abstract: In federated learning (FL), local personalization of models has received significant attention, yet personalized fine-tuning of foundation mod
arXiv:2605.07507v1 Announce Type: new Abstract: The exponential growth of academic publications has created an urgent need for automated tools capable of extracting structured knowledge from unstructu
arXiv:2511.12090v3 Announce Type: replace Abstract: Prompt-based continual learning methods fine-tune only a small set of additional learnable parameters while keeping the pre-trained model's paramete
arXiv:2605.07078v1 Announce Type: new Abstract: Compositional generalization requires models to produce novel configurations from familiar parts. In diffusion models, prior compositional generation me
arXiv:2512.03454v4 Announce Type: replace-cross Abstract: Interpreting natural-language commands to localize target objects is critical for autonomous driving (AD). Existing visual grounding (VG) meth
arXiv:2605.07172v1 Announce Type: new Abstract: Alignment of large language models (LLMs) via SFT and RLHF/DPO typically ignores the global geometry of the representation space, relying instead on loc