AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

Search: “together-ai-blog”

GridTimelineEvolution
29 results
Model Releases

DeepSeek-V4 Flash 0731 vs GPT-5.6 Luna on DeepSWE: Cost and Coding

DGX agent

DeepSeek‑V4 Flash 0731 is the cheapest model on the DeepSWE board, costing about 0.10 per rollout versus GPT‑5.6 Luna’s 0.61, yet it scores a pass@1 of 53.3% compared to Luna’s 67.2%. A cascade strate

model-releasestogether-ai-blog
6 Aug 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Tutorials

Kimi K3: The Complete Developer Guide

DGX agent

**Kimi K3 is Moonshot AI’s 2.8‑trillion‑parameter open‑weight language model—the largest ever released—designed for frontier tasks such as long‑horizon coding and deep reasoning.** Its architecture us

tutorialstogether-ai-blog
1 Aug 2026
Hardware

Autoscaling endpoints for LLM inference

DGX agent

GPU utilization can read healthy while your queue backs up, and a new replica takes minutes to warm. Here's how to pick autoscaling metrics, tune scale-up/down windows, and budget for cold starts on d

hardwaretogether-ai-blog
31 Jul 2026
Tools

Configuring Dedicated Model Inference

DGX agent

The Together AI platform’s dedicated inference architecture consists of three immutable entities: **configs** (engine, GPU type/count, parallelism and optimization profile), **deployments** (a specifi

toolstogether-ai-blog
29 Jul 2026
Agents

ThunderAgent: 2x Faster Agentic Inference for Synthetic Data Generation at Scale

DGX agent

ThunderAgent is a program-aware scheduler for agentic inference. By treating each agent workflow as a schedulable program, it eliminates KV cache thrashing to deliver more than 2x single-node throughp

agentstogether-ai-blog
29 Jul 2026
Applications

The production platform for open-weight AI inference

DGX agent

OpenAI has updated its inference platform to give users full control over performance, cost, and quality without building their own stack—models go live in minutes and support multiple deployments beh

applicationstogether-ai-blog
23 Jul 2026
Hardware

New in Together GPU Clusters: Reliability and control for production GPU clusters

DGX agent

Together GPU Clusters now includes a suite of resilience and operational‑control features designed to address common large‑scale failure modes and team‑scaling challenges. The platform‑health improvem

hardwaretogether-ai-blog
15 Jul 2026
Hardware

Open, convenient and predictable: Introducing Provisioned Throughput

DGX agent

Provisioned Throughput gives you reserved inference capacity for frontier open models like MiniMax M3 and GLM-5.2. Token-based pricing, a 99% uptime SLA, and up to 90% lower cost than proprietary APIs

hardwaretogether-ai-blog
8 Jul 2026
Tools

Announcing our $800M Series C to accelerate the shift to open-source AI

DGX agent

Together AI announced an $800 million Series C funding round to support its mission of advancing open-source artificial intelligence development and deployment. The funding will be used to accelerate

toolstogether-ai-blog
1 Jul 2026
Tools

Together AI at ICML 2026: frontier research across the full stack

DGX agent

Together AI presented research at ICML 2026 covering advances across the full technology stack, likely including foundational model improvements, inference optimization, and practical deployment solut

toolstogether-ai-blog
30 Jun 2026
Hardware

ParallelKernelBench: Frontier LLMs can't write fast multi-GPU kernels (yet)

DGX agent

ParallelKernelBench is a benchmark from Together AI that evaluates frontier large language models' ability to write optimized multi-GPU CUDA kernels, revealing significant limitations in current LLMs'

hardwaretogether-ai-blog
23 Jun 2026
Applications

Building trust in enterprise AI: Together AI earns ISO 27001:2022 certification

DGX agent

Together AI has achieved ISO 27001:2022 certification, demonstrating its commitment to information security management and data protection standards required for enterprise AI deployments. This certif

applicationstogether-ai-blog
10 Jun 2026
Tools

Serving MiniMax-M3 for efficient inference: Unlocking 1M-Token Context and Multimodality Without Regrets

DGX agent

MiniMax-M3 is a large language model capable of handling 1 million token contexts and multimodal inputs while maintaining efficient inference performance. Together AI's blog post discusses techniques

toolstogether-ai-blog
2 Jun 2026
Hardware

How Together AI built the world’s fastest speech-to-text stack

DGX agent

Together AI developed an optimized speech-to-text system focused on achieving the fastest processing speeds through technical innovations in their inference stack and model optimization. The approach

hardwaretogether-ai-blog
29 May 2026
Model Releases

Benchmarking inference at scale: coding agents

DGX agent

This article presents benchmarking results for AI coding agents evaluated at scale, likely comparing performance metrics such as code generation accuracy, execution success rates, and inference effici

model-releasestogether-ai-blog
19 May 2026
Model Releases

Together AI and Pearl Research Labs Team Up to Reduce the Cost of AI Inference

DGX agent

Together AI and Pearl Research Labs announced a partnership aimed at lowering the cost of AI inference through collaborative research and development efforts. The partnership likely combines Together

model-releasestogether-ai-blog
15 May 2026
Tools

Violin: An open-source video translation skill that breaks language barriers

DGX agent

Violin is an open-source video translation tool developed by Together AI that automatically translates video content to break down language barriers for viewers. The skill likely leverages AI models t

toolstogether-ai-blog
14 May 2026
Tools

Introducing voice finder — a new tool to quickly find the right voice for your app from over 600+ voices

DGX agent

Together AI launched Voice Finder, a tool designed to help developers quickly select appropriate voices for their applications from a library of over 600 voice options. The tool streamlines the voice

toolstogether-ai-blog
12 May 2026
Hardware

Deploy and inference any model from HuggingFace

DGX agent

Learn how to deploy any Hugging Face model in one session using Goose and Together's Dedicated Container Inference. Skip the setup complexity — one prompt gets your model running in a production-grade

hardwaretogether-ai-blog
8 May 2026
Model Releases

Serving DeepSeek-V4: why million-token context is an inference systems problem

DGX agent

DeepSeek-V4 makes million-token context a serving-systems problem. Together AI explores the inference work behind V4 on NVIDIA HGX B200, including compressed KV layouts, prefix caching, kernel maturit

model-releasestogether-ai-blog
8 May 2026
Tools

Foundational research powering efficient inference at scale

DGX agent

This article from Together AI discusses foundational research techniques and methodologies used to enable efficient large-scale language model inference. It likely covers optimization strategies, hard

toolstogether-ai-blog
4 May 2026
Applications

Announcing Together AI and Adaption Partnership

DGX agent

Together AI announced a strategic partnership with Adaption to advance AI capabilities and deployment. The collaboration likely focuses on integrating Adaption's technology or services with Together A

applicationstogether-ai-blog
30 Apr 2026
Applications

From 732 bytes to nowhere: shutting down Copy Fail in production

DGX agent

This article describes Together AI's experience shutting down a production service called 'Copy Fail,' detailing the operational and technical challenges of decommissioning a system that was only 732

applicationstogether-ai-blog
30 Apr 2026
Model Releases

DeepSeek-V4 Pro now available on Together AI

DGX agent

DeepSeek-V4 Pro is now available on Together AI with 512K context, controllable reasoning modes, and cached-input pricing for long-context reasoning workloads like code agents, document intelligence,

model-releasestogether-ai-blog
29 Apr 2026
Model Releases

Together AI Brings NVIDIA Nemotron 3 Nano Omni to Developers on Day 0

DGX agent

Together AI announced immediate availability of NVIDIA's Nemotron 3 Nano Omni model to developers through its platform on the day of its release. The Nemotron 3 Nano Omni is a lightweight multimodal m

model-releasestogether-ai-blog
28 Apr 2026
Tools

Accelerate RL rollouts by up to 50% with distribution-aware speculative decoding

DGX agent

This article describes a technique for accelerating reinforcement learning (RL) rollouts using distribution-aware speculative decoding, which can achieve up to 50% speedup improvements. The method lik

toolstogether-ai-blog
24 Apr 2026
Hardware

Capacity without conflict: A guide to multi-tenant GPU cluster design for AI-native teams

DGX agent

This guide addresses the design and management of multi-tenant GPU clusters optimized for AI teams, focusing on strategies to maximize resource utilization while minimizing contention and conflicts be

hardwaretogether-ai-blog
21 Apr 2026
Tools

Parcae: Doing more with fewer parameters using stable looped models

DGX agent

Parcae is a stable looped language model that matches the quality of a Transformer twice its size — a 770M model reaching 1.3B-level performance. We introduce the first scaling laws for looping and sh

toolstogether-ai-blog
15 Apr 2026
Tools

EinsteinArena: Harnessing the collective intelligence of agents in the wild to advance science

DGX agent

EinsteinArena is a platform where AI agents collaborate and compete on open math problems. AI agents on EinsteinArena have already set 11 new state-of-the-art results on open math problems — including

toolstogether-ai-blog
13 Apr 2026
29 results