Pi + @huggingface llamacpp + Qwen3.6 = 🔥🔥🔥🔥
Pi + @huggingface llamacpp + Qwen3.6 = 🔥🔥🔥🔥 turns out not killing the prefix cache all the time and notnhaving a humongous set of tools and a massive system prompt is good for local model use. who'd h
Knowledge catalogue
Pi + @huggingface llamacpp + Qwen3.6 = 🔥🔥🔥🔥 turns out not killing the prefix cache all the time and notnhaving a humongous set of tools and a massive system prompt is good for local model use. who'd h
arXiv:2604.26318v1 Announce Type: new Abstract: Point cloud registration (PCR) is a fundamental task for integrating 3D observations in remote sensing applications. This paper proposes a fast and effe
arXiv:2604.26073v1 Announce Type: cross Abstract: Industrial chemical plants often operate under strict data confidentiality constraints, making centralized data-driven process modeling difficult. Fed
arXiv:2604.17612v2 Announce Type: replace-cross Abstract: Multi-agent systems built on large language models (LLMs) are difficult to reason about. Coordination errors such as deadlocks or type-mismatc
arXiv:2604.25936v1 Announce Type: cross Abstract: Implicit neural representations are powerful for geometric modeling, but their practical use is often limited by the high computational cost of networ
arXiv:2604.26394v1 Announce Type: cross Abstract: Recent advances in large language models and agentic frameworks have enabled virtual customer assistants (VCAs) for complex support. We present SecMat
arXiv:2604.26940v1 Announce Type: new Abstract: Small language models (SLMs) offer computational efficiency for scalable deployment, yet they often fall short of the reasoning power exhibited by their
arXiv:2604.26897v1 Announce Type: new Abstract: Origami-inspired robotic grippers have shown promising potential for object manipulation tasks due to their compact volume and mechanical flexibility. H
arXiv:2511.20032v3 Announce Type: replace Abstract: Visual attention serves as the primary mechanism through which MLLMs interpret visual information; however, its limited localization capability ofte
arXiv:2604.26553v1 Announce Type: cross Abstract: Large language models (LLMs) demonstrate strong multilingual capabilities, yet often fail to consistently generate responses in the intended language,
arXiv:2604.26837v1 Announce Type: new Abstract: Long-context LLM serving is bottlenecked by the cost of attending over ever-growing KV caches. Dynamic sparse attention promises relief by accessing onl
v0.22.1-rc1 is a pre-release version of Ollama that includes improvements to MLX sampler batching, tokenizer BPE offset handling, NVIDIA TensorRT Model Optimizer support, and fixes for desktop app sta
arXiv:2604.26806v1 Announce Type: cross Abstract: Transformer-based architectures have established a dominant paradigm in global semantic perception; however, they remain fundamentally constrained by
A community repository of AI agent configurations for Ollama has reached 888 GitHub stars, featuring pre-built setups and configurations for running AI agents locally. The project appears to focus on
This discussion covers training custom video LoRAs for Wan and LTX Video models using Low-Rank Adaptation, a fine-tuning technique that customizes outputs for specific subjects, styles, or movements w
arXiv:2604.26604v1 Announce Type: new Abstract: Federated learning (FL) trains a shared model from updates contributed by distributed clients, often implicitly assuming that contributing clients are r
b8969 is a release build of llama.cpp (commit bdc9c74), published on April 29, 2026. Llama.cpp is a C/C++ implementation designed to enable LLM inference with minimal setup and state-of-the-art perfor
b8970 is a release of llama.cpp, a C/C++ project for large language model inference. The naming convention using a build hash (b8970) is typical of llama.cpp's continuous release pattern rather than t
b8971 is a build release of llama.cpp, an open-source C/C++ project for LLM inference . As a numbered build identifier from the llama.cpp repository, this release likely contains bug fixes, performanc
b8972 is a release of llama.cpp, a C/C++ implementation for LLM inference . The release identifier follows llama.cpp's development versioning scheme with frequent builds published to the GitHub releas
b8973 is a release of llama.cpp that added SVE tuned code for the gemm_q8_0_4x8_q8_0() kernel and changed arrays to static const in repack.cpp . The llama.cpp project enables LLM inference with minima
b8978 is a release of llama.cpp, a project designed to enable LLM inference with minimal setup and state-of-the-art performance on a wide range of hardware. The llama.cpp project uses sequential build
B8979 is a llama.cpp release that added SVE (Scalable Vector Extension) tuned code for the gemm_q8_0_4x8_q8_0() kernel and made static const changes to repack.cpp . The release includes pre-built bina
b8981 is a release version of llama.cpp, a C/C++ implementation for LLM inference . Build versions in llama.cpp releases typically contain updates to the underlying ggml library, performance optimizat
arXiv:2604.25700v1 Announce Type: cross Abstract: Software quality assurance remains a major challenge in industrial environments, where large-scale and long-lived systems inevitably accumulate defect
Clement Delangue, CEO of Hugging Face, posted about receiving a piece of hardware that fits alongside a DGX Spark system, likely referring to an AI accelerator or GPU device. The post suggests a perso
arXiv:2604.25388v1 Announce Type: new Abstract: Architectural floor plans are widely available priors which contain not only geometry but also the semantic information of the environment, yet existing
arXiv:2604.24997v1 Announce Type: new Abstract: Open-vocabulary semantic segmentation requires assigning pixel-level semantic labels while supporting an open and unrestricted set of categories. Traini
arXiv:2604.25039v1 Announce Type: new Abstract: Large Language Models (LLMs) solve many reasoning tasks via chain-of-thought (CoT) prompting, but smaller models (about 7 to 8B parameters) still strugg
arXiv:2604.24972v1 Announce Type: new Abstract: Clinical abnormality grounding for rare diseases is often hindered by data scarcity, making supervised fine-tuning impractical and single-pass inference
arXiv:2604.25482v1 Announce Type: new Abstract: Large Language Models (LLMs) have shown strong potential for narrative generation, but their use in complex, multi-layered role-playing game (RPG) world
Quanty AI is a local AI companion playground built with Ollama as the backend, featuring animated pixel art graphics and interactive micro fiction experiences. The project integrates agent skills to e
Wan2GP users can explore high-quality output modes like LTX 2 DEV HQ Mode, which is designed to produce better output at higher resolutions by using the HQ sampler with specific settings like 15 steps
📽️ Let's Explore Comfy Hub 🎤 Host: Purz ⏲️ April 29th – 3pm PST / 6pm EST 📍 Live on YouTube, X, and Twitch Today we're diving into Comfy Hub + Comfy Cloud — your gateway to exploring what's possible w
Comfy Hub appears to be an exploratory broadcast or presentation from the ComfyUI project shared via X (formerly Twitter). Based on the context, it likely discusses features, updates, or capabilities
arXiv:2604.25405v1 Announce Type: new Abstract: Camera-based 3D object detection and tracking are central to autonomous driving, yet precise 3D object localization remains fundamentally constrained by
arXiv:2604.25860v1 Announce Type: new Abstract: Machine-generated text (MGT) detection requires identifying structurally invariant signals across generation models, rather than relying on model-specif
MiMo-V2.5 is Xiaomi's multimodal AI model with native visual and audio understanding that supports up to 1 million tokens of context. The GGUF format refers to quantized versions of the model optimize
arXiv:2603.15954v2 Announce Type: replace Abstract: Real-time AI experiences call for on-device large language models (OD-LLMs) optimized for efficient deployment on resource-constrained hardware. The
The Ministral-3:3b model requires Ollama 0.13.1, which is in pre-release , and users encountering errors when running it face various issues including memory allocation problems and GPU/CPU offloading
arXiv:2511.20211v2 Announce Type: replace Abstract: Transparency-aware generation requires modeling not only RGB appearance but also alpha-based opacity and cross-layer composition, which are essentia
arXiv:2604.25276v1 Announce Type: new Abstract: Video Temporal Grounding (VTG), the task of localizing video segments from text queries, struggles in open-world settings due to limited dataset scale a
arXiv:2604.25551v1 Announce Type: new Abstract: Recurrent Graph Neural Networks (RGNNs) extend standard GNNs by iterating message-passing until some stopping condition is met. Various RGNN models have
arXiv:2604.24832v1 Announce Type: new Abstract: Masked diffusion language models (MDMs) have recently emerged as a promising alternative to standard autoregressive large language models (AR-LLMs), yet
arXiv:2601.21293v2 Announce Type: replace Abstract: Reliability-centered prognostics for rotating machinery requires early-warning signals that remain accurate under nonstationary operating conditions
Qwen3.6 27B is a 27-billion parameter language model that can achieve approximately 60 tokens per second throughput when running on dual RTX 5060 Ti GPUs with 16GB memory each, using the vLLM inferenc
arXiv:2604.05594v2 Announce Type: replace Abstract: Pixel-level annotation is costly in low-resource dermoscopy. We present RABC-Net, a reliability-aware annotation-free segmentation system that combi
arXiv:2602.20730v2 Announce Type: replace Abstract: We study efficiency as a first-class objective in Neural Combinatorial Optimization (NCO) and present ECO, an efficient learning framework that comb
arXiv:2604.25646v1 Announce Type: new Abstract: Robotic ultrasound has advanced local image-driven control, contact regulation, and view optimization, yet current systems lack the anatomical understan
arXiv:2604.25071v1 Announce Type: cross Abstract: The prevalence of biometric authentication has been on the rise due to its ease of use and elimination of weak passwords. To date, most biometric auth
SenseNova U1 is a new series of native multimodal models that unifies multimodal understanding, reasoning, and generation within a monolithic architecture, marking a fundamental paradigm shift in mult
I need to check the actual content of this post to provide an accurate summary. FLUX.2 [klein] 9B is a 9 billion parameter image generation model capable of generating images from text descriptions an
UniGenDet is a unified generative-discriminative framework that co-evolves image generation and generated image detection, aligning generation with detector-aware signals for improved authenticity and
We're open-sourcing Hy-MT1.5-1.8B-1.25bit — a 440MB translation model that runs fully offline on your phone, supports 33 languages, and outperforms Google Translate. At 1.8B parameters, it matches com
arXiv:2604.23935v1 Announce Type: new Abstract: Audio-based video object segmentation aims to locate and segment objects in videos conditioned on audio cues, requiring precise understanding of both ap
300,000 AI builders have already added their hardware to HF to instantly see what model they can run locally. To do so, go to https://huggingface.co/settings/local-apps and add your hardware specs. Yo
arXiv:2604.22858v1 Announce Type: new Abstract: Liver cancer, especially hepatocellular carcinoma (HCC), imposes a substantial global disease burden. Accurate diagnosis and prognostic assessment direc
arXiv:2604.23127v1 Announce Type: cross Abstract: Soil salinity is a major environmental challenge in coastal Bangladesh, threatening agricultural productivity and local livelihoods. This study develo
arXiv:2512.21372v2 Announce Type: replace-cross Abstract: The accurate classification of gastrointestinal diseases from endoscopic and histopathological imagery remains a significant challenge in medi
arXiv:2601.12483v2 Announce Type: replace-cross Abstract: Quantum error correction is a key ingredient for large scale quantum computation, protecting logical information from physical noise by encodi