Model Releases
Which is better ninfer vs vllm for Qwen 3.8 27B on RTX 5090?
I have been using the unsloth/Qwen3.8-27B-NVFP4 with 157k ctx on vllm currently. But recently I have been seeing post about ninfer lately and how it is like the best way to run qwen 3.8 on RTX 5090 32
I have been using the unsloth/Qwen3.8-27B-NVFP4 with 157k ctx on vllm currently. But recently I have been seeing post about ninfer lately and how it is like the best way to run qwen 3.8 on RTX 5090 32GB. Seems like a good switch but I would like to know what is the experience with it so far. I also found this gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090 to run on vllm which claims to give more ctx with better performance than the unsloth one. It was some what suspicious but I can't say since I haven't test yet. So, I wanted to know if there has been people who have tested all of these are found which one works the best. Also, I would like to know what is the best config for it when using with agent harness like hermes agent. Ninfer Github link: https://github.com/Neroued/ninfer Edit: add vram amount to clarify the GPU version. submitted by /u/MaxKingCS [link] [comments]
Related
- 2 x 5070ti Qwen 27B full config / stats
- The difference between 'medium' and 'xhigh' reasoning effort for Qwen3.8-27B is actually insane.
- Anyone managed to get Qwen 3.8 27B running smoothly on vLLM? Can't get rid of endless thinking
Source: r/LocalLLaMA | 2026-08-30