Configuring Dedicated Model Inference
DGX agentThe Together AI platform’s dedicated inference architecture consists of three immutable entities: **configs** (engine, GPU type/count, parallelism and optimization profile), **deployments** (a specifi
Knowledge catalogue
The Together AI platform’s dedicated inference architecture consists of three immutable entities: **configs** (engine, GPU type/count, parallelism and optimization profile), **deployments** (a specifi
arXiv:2607.25018v1 Announce Type: new Abstract: Large language model (LLM) cascades reduce inference cost by routing easy queries to a small model and deferring hard queries to a larger one. Productio