Model Releases

4x 3090, 96gb vram what Model to drive Hermes?

3 year lurker, now i finally got my server up and running. dont know which model to choose. llama.cpp or vllm, what makes more sense? mainly single user with maybe 2-3 more additional users in family,

DGX agentreddit
model-releasesr-localllama

3 year lurker, now i finally got my server up and running. dont know which model to choose. llama.cpp or vllm, what makes more sense? mainly single user with maybe 2-3 more additional users in family, if everything checks out. hermes is gonna be used as "ai playground" to manifest ideas on tailscale network and do quick prototyping of thoughts. also ill look into using only 2 3090 for the main model and the other 2 will be dedicated to docling and speech services for a voice agent (speech in-> text out). got some stuff going with my even realities g2 but lost everything when i wiped my ssd for proxmox. yeah... any advice or stuff i should look into is welcome :) submitted by /u/Sea_Calendar_3912 [link] [comments]

Related

Source: r/LocalLLaMA | 2026-07-25

Loading related sources…