SharQ: Bridging Activation Sparsity and FP4 Quantization for LLM Inference
DGX agentarXiv:2606.26587v1 Announce Type: cross Abstract: Low-bit floating-point formats and semi-structured sparsity are increasingly supported by modern accelerators, yet combining them for LLM activation c