MAGE: All-[MASK] Block Already Knows Where to Look in Block Diffusion LLM
DGX agentarXiv:2602.14209v2 Announce Type: replace-cross Abstract: Block diffusion LLMs are an emerging paradigm for parallel language generation, but their KV caching makes memory access the dominant bottlene