Quantize the Target, Quantize the Drafter: Efficient Inference with Qwen3.5-4B
DGX agentarXiv:2607.04244v1 Announce Type: new Abstract: This report describes our approach to the Efficient Qwen Competition, where the goal is to enable low-latency serving of Qwen3.5-4B on a resource-constr