Tools

What if you could get 1.3B Transformer quality from a 770M model? That's not a compression result. It's a different architecture. Parcae, fr…

What if you could get 1.3B Transformer quality from a 770M model? That's not a compression result. It's a different architecture. Parcae, from @realDanFu (Together AI's VP of Kernels) and his lab at U

DGX agentx-post
toolstogether-ai--x
Loading related sources…