AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

Search: “together-ai--x”

GridTimelineEvolution
261 results
CompaniesToolsTechniques

Each lane shows up to 8 recent matching entries, ordered from earlier to later. Tracks load separately to keep the 75,000+ entry wiki fast.

Companies

CompanyAnthropic8 recent entries
12 Apr 2026On software engineering: → 56.22% SWE-Pro, matching GPT-5.3-Codex → 55.6% VIBE-Pro — full-stack project delivery → SWE Multilingual: 76.5 Fo…

On software engineering: → 56.22% SWE-Pro, matching GPT-5.3-Codex → 55.6% VIBE-Pro — full-stack project delivery → SWE Multilingual: 76.5 For agentic work: → Multi-agent collaboration built into the m

→19 May 2026'One thing that we've been seeing recently is that inference benchmarks don't really match production workloads that well.' - @realDanFu, VP…

'One thing that we've been seeing recently is that inference benchmarks don't really match production workloads that well.' - @realDanFu, VP of Kernels When you're running dozens of concurrent coding

HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
→3 Jun 2026M3 brings sparse attention + 1M context + multimodality, and Together did the hard serving work to make it fast. Great collaboration with th…

Anthropic's Claude 3.5 Sonnet (M3) model features sparse attention mechanisms, 1 million token context length, and multimodal capabilities, with Together AI optimizing the serving infrastructure to en

→10 Jun 2026As vertically integrated platforms start to dominate they lock out third party access to the most valuable portions of the platform. Of cour…

As vertically integrated platforms start to dominate they lock out third party access to the most valuable portions of the platform. Of course, Anthropic is has the right to implement whatever policy

→28 Jun 2026More reason why we’re excited about GLM-5.2 on Together 👇 Strong enough for serious coding work, cheap enough to change routing decisions, …

More reason why we’re excited about GLM-5.2 on Together 👇 Strong enough for serious coding work, cheap enough to change routing decisions, and easy to access through the tools developers already use.

→24 Jul 2026.@Kimi_Moonshot K3 lands on Together on Monday! We ran 452 DeepSWE rollouts against Claude Fable 5: near-flagship coding at ~35% of the pric…

.@Kimi_Moonshot K3 lands on Together on Monday! We ran 452 DeepSWE rollouts against Claude Fable 5: near-flagship coding at ~35% of the price, and K3 pulls ahead at higher pass@k's. Full deep-dive: ht

→6 Aug 2026Open models give teams more room to run agent loops, with API prices at a fraction of GPT-5.6 Sol and Claude Fable 5. That matters as planni…

Open models give teams more room to run agent loops, with API prices at a fraction of GPT-5.6 Sol and Claude Fable 5. That matters as planning, tool calls, retries, and long contexts compound token us

→10 Aug 2026Frontier performance you can actually own. Proud to help power DeepSeek-V4-Flash on Ollama's cloud, with the fastest hosted performance avai…

Frontier performance you can actually own. Proud to help power DeepSeek-V4-Flash on Ollama's cloud, with the fastest hosted performance available. Open weights, zero data retention, U.S. & EU hosting.

CompanyOpenAI3 recent entries
23 Apr 2026We're excited to launch Delegate. An agent you delegate work to and move on with your life.

Together AI has launched Delegate, an AI agent designed to handle delegated tasks autonomously, allowing users to assign work and proceed with other activities. The agent appears to be positioned as a

→29 Jul 2026It's a drop-in. One program_id field and it plugs into your existing engine configs such as KV offloading + speculative decoding. Format-agn…

It's a drop-in. One program_id field and it plugs into your existing engine configs such as KV offloading + speculative decoding. Format-agnostic by design, with OpenAI chat completions supported toda

→1 Aug 2026Kimi K3 has set a new bar for OSS model intelligence! 2.8T params, 1M context, OpenAI-compatible API. Complete guide to running Kimi K3 on T…

Kimi K3 has set a new bar for OSS model intelligence! 2.8T params, 1M context, OpenAI-compatible API. Complete guide to running Kimi K3 on Together AI 👇 👏👏 @Kimi_Moonshot 👏👏 https://www.together.ai/bl

CompanyGoogle2 recent entries
9 Apr 2026Gemma 4 31B brings dense multimodal reasoning to Together AI. Try Now: http://www.together.ai/models/gemma-4-31b

Google's Gemma 4 31B is a dense multimodal model from Google DeepMind now available on Together AI's serverless infrastructure via the endpoint `google/gemma-4-31B-it`. It features a 256K context ...

→23 Jun 2026An agentic loop (compile, test, profile, revise) helps. Gemini 3 Pro went from 24 to 35/87 correct, then plateaued after ~20 steps. Feedback…

An agentic loop (compile, test, profile, revise) helps. Gemini 3 Pro went from 24 to 35/87 correct, then plateaued after ~20 steps. Feedback fixes syntax, not rank coordination, collective ordering, o

CompanyMeta2 recent entries
2 Jul 2026LIVE at 12p PT/3p ET: AI’s cleanest stories are getting messy. Meta has excess compute. AI-heavy companies are hiring, not shrinking. And op…

LIVE at 12p PT/3p ET: AI’s cleanest stories are getting messy. Meta has excess compute. AI-heavy companies are hiring, not shrinking. And open models keep getting better. @DanielTNiles on what this me

→24 Jul 2026The Washington Post processed 1.79B input tokens per month through Together AI, running open models like Llama and Mistral in production wit…

The Washington Post processed 1.79B input tokens per month through Together AI, running open models like Llama and Mistral in production with predictable costs and full control over the model stack. T

CompanyMistral1 recent entries
24 Jul 2026The Washington Post processed 1.79B input tokens per month through Together AI, running open models like Llama and Mistral in production wit…

The Washington Post processed 1.79B input tokens per month through Together AI, running open models like Llama and Mistral in production with predictable costs and full control over the model stack. T

CompanyDeepSeek8 recent entries
4 Aug 2026DeepSeek V4 Flash is now live on Together AI. Frontier agent performance is getting dramatically cheaper. DSV4 on Together AI brings a major…

DeepSeek V4 Flash is now live on Together AI. Frontier agent performance is getting dramatically cheaper. DSV4 on Together AI brings a major jump in coding, tool use, and long-running agent performanc

→6 Aug 2026Together Serverless Inference gives developers a managed production path for bringing FLUX 3 into creative products and automated media work…

Together Serverless Inference gives developers a managed production path for bringing FLUX 3 into creative products and automated media workflows. Start building: https://www.together.ai/models/flux-3

→6 Aug 2026Open models give teams more room to run agent loops, with API prices at a fraction of GPT-5.6 Sol and Claude Fable 5. That matters as planni…

Open models give teams more room to run agent loops, with API prices at a fraction of GPT-5.6 Sol and Claude Fable 5. That matters as planning, tool calls, retries, and long contexts compound token us

→7 Aug 2026We analyzed DeepSeek-V4 Flash-0731 vs. GPT-5.6 Luna on software engineering tasks using DeepSWE. DeepSeek-V4 Flash-0731 delivers 80% of Luna…

We analyzed DeepSeek-V4 Flash-0731 vs. GPT-5.6 Luna on software engineering tasks using DeepSWE. DeepSeek-V4 Flash-0731 delivers 80% of Luna’s performance at roughly 1/6 the cost. More insights in the

→8 Aug 2026We analyzed DeepSeek V4 Flash and GPT-5.6 Luna on DeepSWE. A DeepSeek-first cascade with test-suite verification solved MORE tasks than Luna…

Researchers from TogetherAI analyzed DeepSeek V4 Flash and GPT‑5.6 Luna on the DeepSWE benchmark. The study found that employing a DeepSeek‑first cascade with test‑suite verification solved more tasks

→9 Aug 2026We compared how far the same budget goes with DeepSeek V4 Flash and GPT-5.6 Luna on DeepSWE. Two DeepSeek V4 Flash attempts solved MORE task…

We compared how far the same budget goes with DeepSeek V4 Flash and GPT-5.6 Luna on DeepSWE. Two DeepSeek V4 Flash attempts solved MORE tasks than one Luna attempt at roughly one-third the cost. Media

→10 Aug 2026Frontier performance you can actually own. Proud to help power DeepSeek-V4-Flash on Ollama's cloud, with the fastest hosted performance avai…

Frontier performance you can actually own. Proud to help power DeepSeek-V4-Flash on Ollama's cloud, with the fastest hosted performance available. Open weights, zero data retention, U.S. & EU hosting.

→11 Aug 2026DeepSeek V4 Flash 0731 is now available to fine-tune on Together AI. Specialize it for coding, tool use, and your own domain with SFT or DPO…

DeepSeek V4 Flash 0731 is now available to fine-tune on Together AI. Specialize it for coding, tool use, and your own domain with SFT or DPO, then deploy the fine-tuned model on Together AI for produc

CompanyNVIDIA8 recent entries
28 Jul 2026Open weights let builders own and control their stack, and provide flexibility to optimized quality and performance, instead of being locked…

Open weights let builders own and control their stack, and provide flexibility to optimized quality and performance, instead of being locked into a handful of closed platforms. We think that choice is

→29 Jul 2026It's a drop-in. One program_id field and it plugs into your existing engine configs such as KV offloading + speculative decoding. Format-agn…

It's a drop-in. One program_id field and it plugs into your existing engine configs such as KV offloading + speculative decoding. Format-agnostic by design, with OpenAI chat completions supported toda

→31 Jul 2026Together AI gives developers a high-throughput production path for running Inkling-Small on NVIDIA Accelerated Infrastructure across multimo…

Together AI gives developers a high-throughput production path for running Inkling-Small on NVIDIA Accelerated Infrastructure across multimodal, coding, and agentic workloads. Start building: https://

→6 Aug 2026Together Serverless Inference gives developers a managed production path for bringing FLUX 3 into creative products and automated media work…

Together Serverless Inference gives developers a managed production path for bringing FLUX 3 into creative products and automated media workflows. Start building: https://www.together.ai/models/flux-3

→11 Aug 2026Together AI is teaming up with @IBM and @nvidia to bring enterprise-grade AI inference to IBM Cloud. A dedicated NVIDIA B300 cluster. Spectr…

Together AI is teaming up with @IBM and @nvidia to bring enterprise-grade AI inference to IBM Cloud. A dedicated NVIDIA B300 cluster. Spectrum-X networking. First of its kind on IBM Cloud, powered by

→11 Aug 2026NVIDIA Nemotron 3.5 Lightning is now live on Together AI. The fastest open model in its class is built for always-on agents that need to com…

NVIDIA’s Nemotron 3.5 Lightning—a fast open AI model designed for always‑on agents that perform high‑volume, specialized work—has been launched on the Together AI platform as of 11 August 2026. NVIDIA

→11 Aug 2026Nemotron 3.5 Lightning is available on Together AI through Dedicated Model Inference, giving teams reserved capacity and predictable perform…

Nemotron 3.5 Lightning is available on Together AI through Dedicated Model Inference, giving teams reserved capacity and predictable performance for high-volume agent workloads. Start building: https:

→11 Aug 2026More on our partnership: https://newsroom.ibm.com/2026-08-11-IBM-and-Together-AI-Sign-Multi-Year-Agreement-to-Scale-Open-Source-AI-Inference…

In August 2026, Together AI entered into a multi‑year partnership with IBM and NVIDIA to deliver enterprise‑grade open‑source AI inference on IBM Cloud. The collaboration deploys a dedicated NVIDIA B3