Model Releases

Which is better ninfer vs vllm for Qwen 3.8 27B on RTX 5090?

I have been using the unsloth/Qwen3.8-27B-NVFP4 with 157k ctx on vllm currently. But recently I have been seeing post about ninfer lately and how it is like the best way to run qwen 3.8 on RTX 5090 32

DGX agentreddit
model-releasesr-localllama

I have been using the unsloth/Qwen3.8-27B-NVFP4 with 157k ctx on vllm currently. But recently I have been seeing post about ninfer lately and how it is like the best way to run qwen 3.8 on RTX 5090 32GB. Seems like a good switch but I would like to know what is the experience with it so far. I also found this gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090 to run on vllm which claims to give more ctx with better performance than the unsloth one. It was some what suspicious but I can't say since I haven't test yet. So, I wanted to know if there has been people who have tested all of these are found which one works the best. Also, I would like to know what is the best config for it when using with agent harness like hermes agent. Ninfer Github link: https://github.com/Neroued/ninfer Edit: add vram amount to clarify the GPU version. submitted by /u/MaxKingCS [link] [comments]

Related

Source: r/LocalLLaMA | 2026-08-30

Loading related sources…