FuseFSS: Efficient Secure LLM Inference with Function Secret Sharing
DGX agentarXiv:2606.09551v1 Announce Type: cross Abstract: Two-server secure inference allows a client to query a hosted large language model (LLM) without revealing prompts or embeddings. Recent GPU systems b