Local Ai
Hugging Face releases The Stack v3 β largest open code dataset yet
From Anton Lozhkov on π: https://x.com/anton_lozhkov/status/2080254608639701222 Two ways in: stack-v3-train - near-deduplicated, quality-filtered, PII-redacted, contents inline. Point load_dataset at
From Anton Lozhkov on π: https://x.com/anton_lozhkov/status/2080254608639701222 Two ways in: stack-v3-train - near-deduplicated, quality-filtered, PII-redacted, contents inline. Point load_dataset at it and go. https://huggingface.co/datasets/HuggingFaceCode/stack-v3-train stack-v3-full - the entire 114 TB corpus as an HF Storage Bucket: every duplicate kept with cluster IDs, stubs for excluded files. Roll your own dedup, filters, and mixes. https://huggingface.co/buckets/HuggingFaceCode/stack-v3-full submitted by /u/Nunki08 [link] [comments]
Related
- More than 20 companies including NVIDIA, Meta, Microsoft, Palantir, and Hugging Face have signed a letter urging policymakers to avoid premature restrictions on open weight models.
- [[big-dataset-release---supralabsreasoning-corpus-4k-5m-v1---t|[BIG DATASET RELEASE] - SupraLabs/reasoning-corpus-4K-5M-v1 - Train your tiny SLMs to think!]]
Source: r/LocalLLaMA | 2026-07-24