AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “together-ai--x”

GridTimelineEvolution
49+ results
Model Releases

Qwen3.8-2.4T-A95B is now live on Together AI. The Qwen Team’s latest flagship model is built for coding and long-horizon agent workflows, wi…

DGX agent

The Qwen Team has released its flagship model, Qwen3.8‑2.4T‑A95B, on the Together AI platform (togethercompute) as of August 12 2026. This 2.4‑trillion‑parameter model is engineered for coding tasks a

model-releasestogether-ai--x
12 Aug 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Agents

Together Serverless Inference gives developers a managed, high-throughput path for running Qwen3.8-2.4T-A95B across coding and agentic workl…

DGX agent

Together Serverless Inference gives developers a managed, high-throughput path for running Qwen3.8-2.4T-A95B across coding and agentic workloads. Start building: https://www.together.ai/models/qwen3-8

agentstogether-ai--x
12 Aug 2026
Model Releases

DeepSeek V4 Flash 0731 is now available to fine-tune on Together AI. Specialize it for coding, tool use, and your own domain with SFT or DPO…

DGX agent

DeepSeek V4 Flash 0731 is now available to fine-tune on Together AI. Specialize it for coding, tool use, and your own domain with SFT or DPO, then deploy the fine-tuned model on Together AI for produc

model-releasestogether-ai--x
11 Aug 2026
Hardware

More on our partnership: https://newsroom.ibm.com/2026-08-11-IBM-and-Together-AI-Sign-Multi-Year-Agreement-to-Scale-Open-Source-AI-Inference…

DGX agent

In August 2026, Together AI entered into a multi‑year partnership with IBM and NVIDIA to deliver enterprise‑grade open‑source AI inference on IBM Cloud. The collaboration deploys a dedicated NVIDIA B3

hardwaretogether-ai--x
11 Aug 2026
Model Releases

Nemotron 3.5 Lightning is available on Together AI through Dedicated Model Inference, giving teams reserved capacity and predictable perform…

DGX agent

Nemotron 3.5 Lightning is available on Together AI through Dedicated Model Inference, giving teams reserved capacity and predictable performance for high-volume agent workloads. Start building: https:

model-releasestogether-ai--x
11 Aug 2026
Model Releases

NVIDIA Nemotron 3.5 Lightning is now live on Together AI. The fastest open model in its class is built for always-on agents that need to com…

DGX agent

NVIDIA’s Nemotron 3.5 Lightning—a fast open AI model designed for always‑on agents that perform high‑volume, specialized work—has been launched on the Together AI platform as of 11 August 2026. NVIDIA

model-releasestogether-ai--x
11 Aug 2026
Hardware

Together AI is teaming up with @IBM and @nvidia to bring enterprise-grade AI inference to IBM Cloud. A dedicated NVIDIA B300 cluster. Spectr…

DGX agent

Together AI is teaming up with @IBM and @nvidia to bring enterprise-grade AI inference to IBM Cloud. A dedicated NVIDIA B300 cluster. Spectrum-X networking. First of its kind on IBM Cloud, powered by

hardwaretogether-ai--x
11 Aug 2026
Hardware

2/ New deep dive: Autoscaling endpoints for LLM inference. Dedicated Inference can scale on eight metrics. inflight_requests is the default …

DGX agent

2/ New deep dive: Autoscaling endpoints for LLM inference. Dedicated Inference can scale on eight metrics. inflight_requests is the default because it sees queue pressure before latency degrades. We t

hardwaretogether-ai--x
10 Aug 2026
Model Releases

Frontier performance you can actually own. Proud to help power DeepSeek-V4-Flash on Ollama's cloud, with the fastest hosted performance avai…

DGX agent

Frontier performance you can actually own. Proud to help power DeepSeek-V4-Flash on Ollama's cloud, with the fastest hosted performance available. Open weights, zero data retention, U.S. & EU hosting.

model-releasestogether-ai--x
10 Aug 2026
Agents

Together Serverless Inference gives developers a high-throughput, managed production path for running Muse Glimmer across agentic and multim…

DGX agent

Together Serverless Inference gives developers a high-throughput, managed production path for running Muse Glimmer across agentic and multimodal workloads. Start building: https://www.together.ai/mode

agentstogether-ai--x
10 Aug 2026
Model Releases

We compared how far the same budget goes with DeepSeek V4 Flash and GPT-5.6 Luna on DeepSWE. Two DeepSeek V4 Flash attempts solved MORE task…

DGX agent

We compared how far the same budget goes with DeepSeek V4 Flash and GPT-5.6 Luna on DeepSWE. Two DeepSeek V4 Flash attempts solved MORE tasks than one Luna attempt at roughly one-third the cost. Media

model-releasestogether-ai--x
9 Aug 2026
Model Releases

We analyzed DeepSeek V4 Flash and GPT-5.6 Luna on DeepSWE. A DeepSeek-first cascade with test-suite verification solved MORE tasks than Luna…

DGX agent

Researchers from TogetherAI analyzed DeepSeek V4 Flash and GPT‑5.6 Luna on the DeepSWE benchmark. The study found that employing a DeepSeek‑first cascade with test‑suite verification solved more tasks

model-releasestogether-ai--x
8 Aug 2026
Applications

When you move a model into production, you want the quality you evaluated to carry through the serving stack. @Kimi_Moonshot benchmarked Kim…

DGX agent

When you move a model into production, you want the quality you evaluated to carry through the serving stack. @Kimi_Moonshot benchmarked Kimi K3 across major inference providers, and Together AI ranke

applicationstogether-ai--x
8 Aug 2026
Tutorials

A lot of the questions we get from developers are about the concepts behind the API: TTFT, context windows, sampling, fine-tuning, quantizat…

DGX agent

A lot of the questions we get from developers are about the concepts behind the API: TTFT, context windows, sampling, fine-tuning, quantization, deployment tradeoffs. We added Learn to the Together do

tutorialstogether-ai--x
7 Aug 2026
Tools

Autoscaling peaky LLM inference workloads is completely different than autoscaling something like a web service. I wrote a deepdive covering…

DGX agent

Zain (@zainhas) published a detailed article on August 7, 2026 explaining that autoscaling for highly peaky large‑language‑model (LLM) inference is fundamentally different from autoscaling conventiona

toolstogether-ai--x
7 Aug 2026
Model Releases

We analyzed DeepSeek-V4 Flash-0731 vs. GPT-5.6 Luna on software engineering tasks using DeepSWE. DeepSeek-V4 Flash-0731 delivers 80% of Luna…

DGX agent

We analyzed DeepSeek-V4 Flash-0731 vs. GPT-5.6 Luna on software engineering tasks using DeepSWE. DeepSeek-V4 Flash-0731 delivers 80% of Luna’s performance at roughly 1/6 the cost. More insights in the

model-releasestogether-ai--x
7 Aug 2026
Agents

Congrats to @mattrubens and the Roomote team on the launch. Builders can use Together AI as an inference provider in Roomote and assign diff…

DGX agent

Congrats to @mattrubens and the Roomote team on the launch. Builders can use Together AI as an inference provider in Roomote and assign different open models to coding, planning, vision, and review ac

agentstogether-ai--x
6 Aug 2026
Tools

FLUX 3 is now live on Together AI. @bfl_ai’s new multimodal model generates video and synchronized audio together, with up to 20-second clip…

DGX agent

FLUX 3 is now live on Together AI. @bfl_ai’s new multimodal model generates video and synchronized audio together, with up to 20-second clips, multiple shots, and control from text, images, or keyfram

toolstogether-ai--x
6 Aug 2026
Model Releases

Open models give teams more room to run agent loops, with API prices at a fraction of GPT-5.6 Sol and Claude Fable 5. That matters as planni…

DGX agent

Open models give teams more room to run agent loops, with API prices at a fraction of GPT-5.6 Sol and Claude Fable 5. That matters as planning, tool calls, retries, and long contexts compound token us

model-releasestogether-ai--x
6 Aug 2026
Applications

Together Serverless Inference gives developers a managed production path for bringing FLUX 3 into creative products and automated media work…

DGX agent

Together Serverless Inference gives developers a managed production path for bringing FLUX 3 into creative products and automated media workflows. Start building: https://www.together.ai/models/flux-3

applicationstogether-ai--x
6 Aug 2026
Tools

99.9% uptime changes what your inference architecture has to survive. At Together AI, it means multi-data-center deployment, live traffic ac…

DGX agent

99.9% uptime changes what your inference architecture has to survive. At Together AI, it means multi-data-center deployment, live traffic across both facilities, and enough capacity to absorb a full d

toolstogether-ai--x
5 Aug 2026
Tools

.@Kimi_Moonshot benchmarked K3 endpoints across major inference providers. Together AI leads or ties for #1 on 3 of 4 benchmarks: OCRBench, …

DGX agent

.@Kimi_Moonshot benchmarked K3 endpoints across major inference providers. Together AI leads or ties for #1 on 3 of 4 benchmarks: OCRBench, MMMU Pro Vision, and DeepSWE. Open models like Kimi K3 have

toolstogether-ai--x
5 Aug 2026
Model Releases

DeepSeek V4 Flash is now live on Together AI. Frontier agent performance is getting dramatically cheaper. DSV4 on Together AI brings a major…

DGX agent

DeepSeek V4 Flash is now live on Together AI. Frontier agent performance is getting dramatically cheaper. DSV4 on Together AI brings a major jump in coding, tool use, and long-running agent performanc

model-releasestogether-ai--x
4 Aug 2026
Model Releases

Together AI gives developers a high-throughput production path for running DeepSeek V4 Flash across coding, tool-use, and agentic workloads.…

DGX agent

Together AI gives developers a high-throughput production path for running DeepSeek V4 Flash across coding, tool-use, and agentic workloads. Start building: https://www.together.ai/models/deepseek-v4-

model-releasestogether-ai--x
4 Aug 2026
Tutorials

Kimi K3 has set a new bar for OSS model intelligence! 2.8T params, 1M context, OpenAI-compatible API. Complete guide to running Kimi K3 on T…

DGX agent

Kimi K3 has set a new bar for OSS model intelligence! 2.8T params, 1M context, OpenAI-compatible API. Complete guide to running Kimi K3 on Together AI 👇 👏👏 @Kimi_Moonshot 👏👏 https://www.together.ai/bl

tutorialstogether-ai--x
1 Aug 2026
Tools

Inkling-Small is now live on Together AI. @thinkymachines’ new open-weight multimodal model delivers similar performance to Inkling at one-q…

DGX agent

Inkling-Small is now live on Together AI. @thinkymachines’ new open-weight multimodal model delivers similar performance to Inkling at one-quarter the size, built for coding, agents, and general multi

toolstogether-ai--x
31 Jul 2026
Hardware

The first thing we learned building autoscaling for dedicated inference is that CPU-style metrics don't tell the whole picture. A GPU can re…

DGX agent

The first thing we learned building autoscaling for dedicated inference is that CPU-style metrics don't tell the whole picture. A GPU can read 60% busy while the engine's queue is already backing up,

hardwaretogether-ai--x
31 Jul 2026
Agents

Together AI gives developers a high-throughput production path for running Inkling-Small on @NVIDIAAI Accelerated Infrastructure across mult…

DGX agent

Together AI gives developers a high-throughput production path for running Inkling-Small on @NVIDIAAI Accelerated Infrastructure across multimodal, coding, and agentic workloads. Start building: https

agentstogether-ai--x
31 Jul 2026
Hardware

Together AI gives developers a high-throughput production path for running Inkling-Small on NVIDIA Accelerated Infrastructure across multimo…

DGX agent

Together AI gives developers a high-throughput production path for running Inkling-Small on NVIDIA Accelerated Infrastructure across multimodal, coding, and agentic workloads. Start building: https://

hardwaretogether-ai--x
31 Jul 2026
Hardware

Can an open-source model perform like a foundation model? @Osmosis_AI is betting yes, using reinforcement learning and the dedicated @ycombi…

DGX agent

Osmosis_AI claims that an open‑source model can rival a foundation model by leveraging reinforcement learning techniques. To demonstrate this, they will use the Y Combinator‑dedicated GPU cluster on T

hardwaretogether-ai--x
30 Jul 2026
Agents

Interesting data from the OpenRouter Kimi K3 dashboard: @togethercompute offers one of the lowest prices for @Kimi_Moonshot K3 while also ac…

DGX agent

Interesting data from the OpenRouter Kimi K3 dashboard: @togethercompute offers one of the lowest prices for @Kimi_Moonshot K3 while also achieving one of the highest prompt cache hit rates. Low prici

agentstogether-ai--x
30 Jul 2026
Applications

Starting in 15 minutes

DGX agent

Starting in 15 minutes Kimi K3 has everyone’s attention. On July 30, hear @Kimi_Moonshot's Feihu Tang explain the architecture and decisions behind it. He joins Jue Wang and Zain Hasan from Together A

applicationstogether-ai--x
30 Jul 2026
Agents

Agentic inference wastes GPUs on KV cache thrashing. ThunderAgent fixes it at the scheduler level: 2.5x higher single-node throughput and ~1…

DGX agent

Agentic inference wastes GPUs on KV cache thrashing. ThunderAgent fixes it at the scheduler level: 2.5x higher single-node throughput and ~10x lower P50 latency at high concurrency. ThunderAgent was a

agentstogether-ai--x
29 Jul 2026
Hardware

It's a drop-in. One program_id field and it plugs into your existing engine configs such as KV offloading + speculative decoding. Format-agn…

DGX agent

It's a drop-in. One program_id field and it plugs into your existing engine configs such as KV offloading + speculative decoding. Format-agnostic by design, with OpenAI chat completions supported toda

hardwaretogether-ai--x
29 Jul 2026
Tools

Packed house today at the first edition of our @MiniMax_AI Intelligence in the Open event. Thank you to our amazing cohost @withprotegeai An…

DGX agent

Packed house today at the first edition of our @MiniMax_AI Intelligence in the Open event. Thank you to our amazing cohost @withprotegeai And our wonderful partners UpscaleX @togethercompute @Artifici

toolstogether-ai--x
29 Jul 2026
Safety

The problem: agent workflows alternate between GPU-heavy reasoning and GPU-idle waiting on tools. Run hundreds concurrently and their KV cac…

DGX agent

The problem: agent workflows alternate between GPU-heavy reasoning and GPU-idle waiting on tools. Run hundreds concurrently and their KV caches fight for memory. Engines evict on a dumb LRU policy, ev

safetytogether-ai--x
29 Jul 2026
Agents

The root cause: request-level engines never see that a series of LLM calls belongs to one longer workflow. ThunderAgent adds that missing vi…

DGX agent

The root cause: request-level engines never see that a series of LLM calls belongs to one longer workflow. ThunderAgent adds that missing view. It treats each agent workflow as a schedulable program,

agentstogether-ai--x
29 Jul 2026
Applications

Everything you need to know about Kimi K3 with @Kimi_Moonshot and @togethercompute team. If you're looking into evaling and building with Ki…

DGX agent

Everything you need to know about Kimi K3 with @Kimi_Moonshot and @togethercompute team. If you're looking into evaling and building with Kimi K3 I would not miss this one! Kimi K3 has everyone’s atte

applicationstogether-ai--x
28 Jul 2026
Applications

Kimi K3 has everyone’s attention. On July 30, hear @Kimi_Moonshot's Feihu Tang explain the architecture and decisions behind it. He joins Ju…

DGX agent

Kimi K3 has everyone’s attention. On July 30, hear @Kimi_Moonshot's Feihu Tang explain the architecture and decisions behind it. He joins Jue Wang and Zain Hasan from Together AI for a deep dive into

applicationstogether-ai--x
28 Jul 2026
Safety

Open weights let builders own and control their stack, and provide flexibility to optimized quality and performance, instead of being locked…

DGX agent

Open weights let builders own and control their stack, and provide flexibility to optimized quality and performance, instead of being locked into a handful of closed platforms. We think that choice is

safetytogether-ai--x
28 Jul 2026
Tools

Robot policies can move but can't think. LLMs can think but can't move. So we connected them. Real robot: 16.7% → 97.3% Sim (LIBERO-PRO): 12…

DGX agent

A team led by Liane Galanti linked large‑language models (LLMs) with robotic motion policies, allowing a robot to combine reasoning capabilities with physical movement. In experiments on a real robot,

toolstogether-ai--x
28 Jul 2026
Tools

The apps you use swap AI models constantly. Doing that without downtime or a bad rollout reaching users is still often manual. We rebuilt th…

DGX agent

The apps you use swap AI models constantly. Doing that without downtime or a bad rollout reaching users is still often manual. We rebuilt that workflow based on what we’ve learned serving more than 40

toolstogether-ai--x
28 Jul 2026
Agents

In partnership with @Kimi_Moonshot, we now have K3 on Together APIs with capacity for hundreds of millions of TPM on launch day!

DGX agent

In partnership with @Kimi_Moonshot, we now have K3 on Together APIs with capacity for hundreds of millions of TPM on launch day! Kimi K3 is now live on Together AI. We’re proud to be a Day 0 launch pa

agentstogether-ai--x
27 Jul 2026
Tools

Kimi K3 is now available directly from its Hugging Face model page through Together AI, give it a try!

DGX agent

Kimi K3 is now available directly from its Hugging Face model page through Together AI, give it a try! Kimi K3 in @huggingface Inference Providers is live via @togethercompute 3/M input tokens, 15/M o

toolstogether-ai--x
27 Jul 2026
Agents

Kimi K3 is now live on Together AI. We’re proud to be a Day 0 launch partner for @Kimi_Moonshot’s open frontier model, built for long-runnin…

DGX agent

Kimi K3 is now live on Together AI. We’re proud to be a Day 0 launch partner for @Kimi_Moonshot’s open frontier model, built for long-running agentic workflows across code, tools, vision, and research

agentstogether-ai--x
27 Jul 2026
Model Releases

@Kimi_Moonshot K3 on Together AI is built for long-running agent workflows: → 2.8T parameters and a 1M context window → Native vision for sc…

DGX agent

@Kimi_Moonshot K3 on Together AI is built for long-running agent workflows: → 2.8T parameters and a 1M context window → Native vision for screenshot-guided coding → Repository navigation and terminal

model-releasestogether-ai--x
27 Jul 2026
Agents

Together gives developers an efficient, high-throughput production path for K3’s long, tool-heavy agent workloads, hosted on Together AI’s U…

DGX agent

Together gives developers an efficient, high-throughput production path for K3’s long, tool-heavy agent workloads, hosted on Together AI’s US-based infrastructure with zero data retention. Start build

agentstogether-ai--x
27 Jul 2026
Agents

Want to go deeper? Join Moonshot AI and Together AI for a technical webinar on how K3 was built and how to use it for production agent workf…

DGX agent

Together AI has released the Kimi K3 model on its platform as a Day‑0 launch partner for Moonshot AI’s open frontier agentic model, which supports long‑running workflows across code, tools, vision and

agentstogether-ai--x
27 Jul 2026
← Previous
123…6
Next →
263 results