PaddlePaddle/HPD-Parsing · Hugging Face
HPD-Parsing: Hierarchical Parallel Document Parsing We introduce HPD-Parsing, a lightweight (1B) and high-throughput document parsing model built on a Hierarchical Parallel Decoding paradigm. Unified
Knowledge catalogue
HPD-Parsing: Hierarchical Parallel Document Parsing We introduce HPD-Parsing, a lightweight (1B) and high-throughput document parsing model built on a Hierarchical Parallel Decoding paradigm. Unified
Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pressure, non-overlappe
Link to their official GGUF repo: https://huggingface.co/poolside/Laguna-S-2.1-GGUF/tree/main All the GGUFs received this fix 5ish hours ago - correct yarn_attn_factor to 1.0 (llama.cpp derives mscale
Shoutout to this awesome guy - https://www.reddit.com/r/LLM/s/IDUyU3v9ap Thanks to his project, BigMoeOnEdge https://github.com/Helldez/BigMoeOnEdge, I managed to successfully run a 35B MoE model on j
This is a surprisingly large real-world deployment: 'GovGPT' is part of Austria’s Public AI initiative, running on sovereign infrastructure (in their BRZ - federal datacenter) with Mistral open-weight
Hey HN, Henry & Roman here from Cactus. A small, on-device model is fast and private, but sometimes wrong, but frontier models are getting expensive pretty fast. So, we post-trained Gemma 4 E2B post-t
Today the Department of Energy (DOE) and Arcee AI announced the development of Genesis-Science-1 (GS1), an open model for scientific research. This is a joint effort to bring advanced AI into scientif
One thing I noticed in American politics, whenever the government wants to push unpopular actions or laws, they often introduce fear to convince the public to support them. This is actually how i view
Fara1.5-27B is a multimodal computer use agent (CUA) for web browsers, from Microsoft Research AI Frontiers. It observes the browser through screenshots and acts on the user's behalf by emitting struc
The primary driver of this project is that I'd become frustrated with the reasoning behavior of smaller local models such as Qwen3.6-27B (i believe particularly at lower temperatures, and where system
I don't have the ability to access Reddit posts or browse specific URLs. To provide you with an accurate factual summary for your knowledge base, I would need either: 1. The actual content/text from t
I managed to get this model working on 2x 3090s with full 262k ctx and N=4, if anyone is interested to try it, thanks to this quant: https://huggingface.co/danielrmay/NVIDIA-Nemotron-Labs-3-Puzzle-75B
Hola Strix and AMD stacker frendios. Read the Lineage and Credits, this uses charlie12345/ROCmFPX, won't work on native llama.cpp yet. 122B total · 10B active · 60.70 GiB · 28.50 tok/s MTP-off · BF16
MODEL + GGUF : https://huggingface.co/InternScience/models?search=a1-4b Technical Report Benchmark Qwen3.5-4B Agents-A1-4B Qwen3.5 Qwen3.6 Nex-N2-mini Agents-A1 🧠 Dense Models (~4B) 🔀 MoE Models (35B-
LLMs handle speech well once you run speech-to-text. They don't hear the rest: a bird outside, a glass breaking two rooms away, a smoke alarm two floors down. I've been working on an experimental open
Below Upstream Status sections are from https://github.com/PrismML-Eng/Bonsai-demo Upstream Status for Binary Q1_0 is supported out of the box in upstream llama.cpp across many backends: CPU (generic,
Efficiency defined as score over active parameters. Removed all the models that were not on the pareto frontier. Yes I'm aware that artificialanalysis.ai aggregate benchmark isn't perfect, but I have
After over a year in development, ExLlamaV3 has had its first production release. Turboderp has been pulling 10 hour days with Fable to bring us this massive batch of improvements. Check out detailed
Benchmark Results Benchmark configuration: threads = 96 type_k = bf16 type_v = bf16 Llama-3.1-8B-Instruct Q8_0 Prompt Size GGML_CPU_Q8_0 t/s ZenDNN_Q8_0 t/s Gain 256 472.28 730.87 54.75% 512 450.86 83
https://youtu.be/oxpGq5FITgA?si=nkHWLReGCDYe7QfL I got Hermes running in the native Debian Terminal in Graphene OS and its really slick. Voice dictation works amazingly. Im using a remote Hermes gatew
Hey r/LocalLLaMA — solo researcher here. I built a new attention architecture and want independent testers. Wave Field LLM replaces O(N²) dot-product attention with FFT wave convolution on a field. Tr
The full quote: I realize that some people really dislike AI, but this is an area where I'm willing to absolutely put my foot down as the top-level maintainer. Linux is not one of those anti-AI projec
Two DGX Spark and a Connect-X7 cable give you about 250GB of usable memory for 7000 8000 USD. This allows using some interesting models at 4-bit. For what seemed like an eternity, the only serious mod
What it is: ATH-MaaS/OvisOCR2 - a 0.8B document-parsing VLM post-trained from Qwen3.5-0.8B (SFT + RL + OPD), Apache 2.0, runs on vLLM 0.22.1. One prompt per page image -> complete markdown (HTML table
Frontier models are just so good though. Fable 5... Gemini 3.1 Pro for design critique and brainstorming. Grok for verification passes. Antigravity with Gemini 3.5 Flash for rote plan execution. Openc
Some fixes & boost(pp) for SYCL/Intel. Merged PRs: [SYCL] Flash Attention with XMX engine via oneDNN graph API (SDPA) on KV f16 for Xe2 ; Qwen3.6-27b-Q8_0 prefill speed up x1.21 at p=512 and x4.26 at
Disclosure: I work at Pluralis Research, the lab that built this. Code is open, and I'm happy to answer questions. TL;DR: As far as we can tell, this is the first RL post-training run whose entire rol
https://techcrunch.com/2026/07/13/satya-nadella-has-issued-a-shocking-warning-to-companies-using-ai/ Venture capitalists have been warning for awhile that OpenAI and Anthropic are getting access to se
Introduction RxBrain (Hy-Embodied-RxBrain-1.0) is a unified multimodal foundation model for embodied cognition — a single model that couples language reasoning with visual imagination to deliver three
Very impressive release by the PrismML team. 1-bit quantization shrinks it from 54GB to just 3.8GB (-93%), while retaining 90% of its intelligence. - Collection on Hugging Face: https://huggingface.co
dam bois we eating good this week ngl, The velocity of the open_weight ecosystem right now is hitting a point where proprietary, closed-source APIs are losing their leverage on compute intelligence. W
I'd need to search for this specific Reddit post to provide accurate details about the actual findings and technical specifics of this GPU performance benchmark. This post likely discusses throughput
Google added an empty thinking token to the Gemma 4 chat template, which stabilizes model output by suppressing 'ghost' thought channels that may appear even when thinking is deactivated. This update
PR #24269 added native video input to llama.cpp's multimodal (mtmd) system, merging on June 8, 2026. The implementation uses FFmpeg as a subprocess to decode video frames and expands a single video ma
Pipeline parallelism in llama.cpp distributes model layers across multiple GPUs, with each GPU holding a contiguous slice of layers . However, the Reddit post likely discusses inefficiencies in how pi
BitNet b1.58 uses ternary weights (-1, 0, 1) and achieves performance comparable to full-precision transformers , enabling efficient LLM inference on CPUs and edge devices. While research into efficie
Xiaomi achieved over 1,000 tokens per second output from a 1 trillion-parameter model using a single standard 8-GPU commodity node through extreme model-system codesign . The approach combines FP4 qua
Dell has confirmed an embargoed XPS laptop launch with NVIDIA N1X set for May 31 , marking a consumer version of the GB10 Superchip with Windows support, unlike the server-focused DGX Spark . The N1X
An open-source system that converts vocal imitations—human-made sound recreations—into synthesized sound effects for creative applications. The technology produces sound effects from vocal imitations
A user reports successfully running Qwen 3.6 35b MoE (mixture of experts) with Zoo Code on an M1 Max Mac, achieving local inference without external servers. The setup enables fully local, battery-pow
This is a Reddit discussion from r/LocalLLaMA where a user seeks recommendations for settings and plugins after switching from OpenCode to Pi, likely asking the community for guidance on optimizing th
Apple has removed its 256GB M3 Ultra Mac Studio from sale, limiting the machine to 96GB of unified memory , following the removal of the 512GB configuration in March . The removal is likely due to a g
BeeLlama.cpp is an optimized implementation featuring advanced DFlash and TurboQuant quantization techniques with support for reasoning and vision capabilities. The project demonstrates running Qwen 3
DeepSeek V4 is undergoing limited grayscale testing with a new interface featuring Fast, Expert, and Vision modes . The Vision version represents the multimodal component of the upcoming DeepSeek V4 r
MiMo-V2.5 is Xiaomi's multimodal AI model with native visual and audio understanding that supports up to 1 million tokens of context. The GGUF format refers to quantized versions of the model optimize
FlashQLA is a high-performance linear attention kernel library built on TileLang developed by Alibaba's Qwen team. The introduction of FlashQLA represents an optimization technology designed to improv
Qwen3.6 27B is a 27-billion parameter language model that can achieve approximately 60 tokens per second throughput when running on dual RTX 5060 Ti GPUs with 16GB memory each, using the vLLM inferenc
SenseNova U1 is a new series of native multimodal models that unifies multimodal understanding, reasoning, and generation within a monolithic architecture, marking a fundamental paradigm shift in mult
NVIDIA Nemotron 3 Nano Omni is a multimodal large language model that unifies video, audio, image, and text understanding for enterprise Q&A, summarization, transcription, and document intelligence, w