Tools
What if you could get 1.3B Transformer quality from a 770M model? That's not a compression result. It's a different architecture. Parcae, fr…
What if you could get 1.3B Transformer quality from a 770M model? That's not a compression result. It's a different architecture. Parcae, from @realDanFu (Together AI's VP of Kernels) and his lab at U
What if you could get 1.3B Transformer quality from a 770M model? That's not a compression result. It's a different architecture. Parcae, from @realDanFu (Together AI's VP of Kernels) and his lab at UCSD, passes activations through the same layers multiple times — stably, for the first time.
Related
- Training code and models are live on Hugging Face. Dan Fu (Together AI's VP of Kernels) led the work. Together AI provided compute. Blog: ht…
- Parcae: Doing more with fewer parameters using stable looped models
- The edge inference implication: memory, not compute, is the binding constraint. Parcae opens a new axis. Scale quality by looping deeper, no…
- Our researchers are heading to ICLR with new work: model efficiency, long-context reasoning, next-gen attention and decoding, and more. Chec…
Source: Together AI (X) | 2026-04-15