AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “together-ai--x”

GridTimelineEvolution
61+ results
12 Aug 2026

Qwen3.8-2.4T-A95B is now live on Together AI. The Qwen Team’s latest flagship model is built for coding and long-horizon agent workflows, wi…

Model ReleasesDGX agent

The Qwen Team has released its flagship model, Qwen3.8‑2.4T‑A95B, on the Together AI platform (togethercompute) as of August 12 2026. This 2.4‑trillion‑parameter model is engineered for coding tasks a

Together Serverless Inference gives developers a managed, high-throughput path for running Qwen3.8-2.4T-A95B across coding and agentic workl…

AgentsDGX agent

Together Serverless Inference gives developers a managed, high-throughput path for running Qwen3.8-2.4T-A95B across coding and agentic workloads. Start building: https://www.together.ai/models/qwen3-8

11 Aug 2026

DeepSeek V4 Flash 0731 is now available to fine-tune on Together AI. Specialize it for coding, tool use, and your own domain with SFT or DPO…


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

DeepSeek V4 Flash 0731 is now available to fine-tune on Together AI. Specialize it for coding, tool use, and your own domain with SFT or DPO, then deploy the fine-tuned model on Together AI for produc

More on our partnership: https://newsroom.ibm.com/2026-08-11-IBM-and-Together-AI-Sign-Multi-Year-Agreement-to-Scale-Open-Source-AI-Inference…

HardwareDGX agent

In August 2026, Together AI entered into a multi‑year partnership with IBM and NVIDIA to deliver enterprise‑grade open‑source AI inference on IBM Cloud. The collaboration deploys a dedicated NVIDIA B3

Nemotron 3.5 Lightning is available on Together AI through Dedicated Model Inference, giving teams reserved capacity and predictable perform…

Model ReleasesDGX agent

Nemotron 3.5 Lightning is available on Together AI through Dedicated Model Inference, giving teams reserved capacity and predictable performance for high-volume agent workloads. Start building: https:

NVIDIA Nemotron 3.5 Lightning is now live on Together AI. The fastest open model in its class is built for always-on agents that need to com…

Model ReleasesDGX agent

NVIDIA’s Nemotron 3.5 Lightning—a fast open AI model designed for always‑on agents that perform high‑volume, specialized work—has been launched on the Together AI platform as of 11 August 2026. NVIDIA

Together AI is teaming up with @IBM and @nvidia to bring enterprise-grade AI inference to IBM Cloud. A dedicated NVIDIA B300 cluster. Spectr…

HardwareDGX agent

Together AI is teaming up with @IBM and @nvidia to bring enterprise-grade AI inference to IBM Cloud. A dedicated NVIDIA B300 cluster. Spectrum-X networking. First of its kind on IBM Cloud, powered by

10 Aug 2026

2/ New deep dive: Autoscaling endpoints for LLM inference. Dedicated Inference can scale on eight metrics. inflight_requests is the default …

HardwareDGX agent

2/ New deep dive: Autoscaling endpoints for LLM inference. Dedicated Inference can scale on eight metrics. inflight_requests is the default because it sees queue pressure before latency degrades. We t

Frontier performance you can actually own. Proud to help power DeepSeek-V4-Flash on Ollama's cloud, with the fastest hosted performance avai…

Model ReleasesDGX agent

Frontier performance you can actually own. Proud to help power DeepSeek-V4-Flash on Ollama's cloud, with the fastest hosted performance available. Open weights, zero data retention, U.S. & EU hosting.

Together Serverless Inference gives developers a high-throughput, managed production path for running Muse Glimmer across agentic and multim…

AgentsDGX agent

Together Serverless Inference gives developers a high-throughput, managed production path for running Muse Glimmer across agentic and multimodal workloads. Start building: https://www.together.ai/mode

9 Aug 2026

We compared how far the same budget goes with DeepSeek V4 Flash and GPT-5.6 Luna on DeepSWE. Two DeepSeek V4 Flash attempts solved MORE task…

Model ReleasesDGX agent

We compared how far the same budget goes with DeepSeek V4 Flash and GPT-5.6 Luna on DeepSWE. Two DeepSeek V4 Flash attempts solved MORE tasks than one Luna attempt at roughly one-third the cost. Media

8 Aug 2026

We analyzed DeepSeek V4 Flash and GPT-5.6 Luna on DeepSWE. A DeepSeek-first cascade with test-suite verification solved MORE tasks than Luna…

Model ReleasesDGX agent

Researchers from TogetherAI analyzed DeepSeek V4 Flash and GPT‑5.6 Luna on the DeepSWE benchmark. The study found that employing a DeepSeek‑first cascade with test‑suite verification solved more tasks

When you move a model into production, you want the quality you evaluated to carry through the serving stack. @Kimi_Moonshot benchmarked Kim…

ApplicationsDGX agent

When you move a model into production, you want the quality you evaluated to carry through the serving stack. @Kimi_Moonshot benchmarked Kimi K3 across major inference providers, and Together AI ranke

7 Aug 2026

A lot of the questions we get from developers are about the concepts behind the API: TTFT, context windows, sampling, fine-tuning, quantizat…

TutorialsDGX agent

A lot of the questions we get from developers are about the concepts behind the API: TTFT, context windows, sampling, fine-tuning, quantization, deployment tradeoffs. We added Learn to the Together do

Autoscaling peaky LLM inference workloads is completely different than autoscaling something like a web service. I wrote a deepdive covering…

ToolsDGX agent

Zain (@zainhas) published a detailed article on August 7, 2026 explaining that autoscaling for highly peaky large‑language‑model (LLM) inference is fundamentally different from autoscaling conventiona

We analyzed DeepSeek-V4 Flash-0731 vs. GPT-5.6 Luna on software engineering tasks using DeepSWE. DeepSeek-V4 Flash-0731 delivers 80% of Luna…

Model ReleasesDGX agent

We analyzed DeepSeek-V4 Flash-0731 vs. GPT-5.6 Luna on software engineering tasks using DeepSWE. DeepSeek-V4 Flash-0731 delivers 80% of Luna’s performance at roughly 1/6 the cost. More insights in the

6 Aug 2026

Congrats to @mattrubens and the Roomote team on the launch. Builders can use Together AI as an inference provider in Roomote and assign diff…

AgentsDGX agent

Congrats to @mattrubens and the Roomote team on the launch. Builders can use Together AI as an inference provider in Roomote and assign different open models to coding, planning, vision, and review ac

FLUX 3 is now live on Together AI. @bfl_ai’s new multimodal model generates video and synchronized audio together, with up to 20-second clip…

ToolsDGX agent

FLUX 3 is now live on Together AI. @bfl_ai’s new multimodal model generates video and synchronized audio together, with up to 20-second clips, multiple shots, and control from text, images, or keyfram

Open models give teams more room to run agent loops, with API prices at a fraction of GPT-5.6 Sol and Claude Fable 5. That matters as planni…

Model ReleasesDGX agent

Open models give teams more room to run agent loops, with API prices at a fraction of GPT-5.6 Sol and Claude Fable 5. That matters as planning, tool calls, retries, and long contexts compound token us

Together Serverless Inference gives developers a managed production path for bringing FLUX 3 into creative products and automated media work…

ApplicationsDGX agent

Together Serverless Inference gives developers a managed production path for bringing FLUX 3 into creative products and automated media workflows. Start building: https://www.together.ai/models/flux-3

5 Aug 2026

99.9% uptime changes what your inference architecture has to survive. At Together AI, it means multi-data-center deployment, live traffic ac…

ToolsDGX agent

99.9% uptime changes what your inference architecture has to survive. At Together AI, it means multi-data-center deployment, live traffic across both facilities, and enough capacity to absorb a full d

.@Kimi_Moonshot benchmarked K3 endpoints across major inference providers. Together AI leads or ties for #1 on 3 of 4 benchmarks: OCRBench, …

ToolsDGX agent

.@Kimi_Moonshot benchmarked K3 endpoints across major inference providers. Together AI leads or ties for #1 on 3 of 4 benchmarks: OCRBench, MMMU Pro Vision, and DeepSWE. Open models like Kimi K3 have

4 Aug 2026

DeepSeek V4 Flash is now live on Together AI. Frontier agent performance is getting dramatically cheaper. DSV4 on Together AI brings a major…

Model ReleasesDGX agent

DeepSeek V4 Flash is now live on Together AI. Frontier agent performance is getting dramatically cheaper. DSV4 on Together AI brings a major jump in coding, tool use, and long-running agent performanc

Together AI gives developers a high-throughput production path for running DeepSeek V4 Flash across coding, tool-use, and agentic workloads.…

Model ReleasesDGX agent

Together AI gives developers a high-throughput production path for running DeepSeek V4 Flash across coding, tool-use, and agentic workloads. Start building: https://www.together.ai/models/deepseek-v4-

1 Aug 2026

Kimi K3 has set a new bar for OSS model intelligence! 2.8T params, 1M context, OpenAI-compatible API. Complete guide to running Kimi K3 on T…

TutorialsDGX agent

Kimi K3 has set a new bar for OSS model intelligence! 2.8T params, 1M context, OpenAI-compatible API. Complete guide to running Kimi K3 on Together AI 👇 👏👏 @Kimi_Moonshot 👏👏 https://www.together.ai/bl

31 Jul 2026

Inkling-Small is now live on Together AI. @thinkymachines’ new open-weight multimodal model delivers similar performance to Inkling at one-q…

ToolsDGX agent

Inkling-Small is now live on Together AI. @thinkymachines’ new open-weight multimodal model delivers similar performance to Inkling at one-quarter the size, built for coding, agents, and general multi

The first thing we learned building autoscaling for dedicated inference is that CPU-style metrics don't tell the whole picture. A GPU can re…

HardwareDGX agent

The first thing we learned building autoscaling for dedicated inference is that CPU-style metrics don't tell the whole picture. A GPU can read 60% busy while the engine's queue is already backing up,

Together AI gives developers a high-throughput production path for running Inkling-Small on @NVIDIAAI Accelerated Infrastructure across mult…

AgentsDGX agent

Together AI gives developers a high-throughput production path for running Inkling-Small on @NVIDIAAI Accelerated Infrastructure across multimodal, coding, and agentic workloads. Start building: https

Together AI gives developers a high-throughput production path for running Inkling-Small on NVIDIA Accelerated Infrastructure across multimo…

HardwareDGX agent

Together AI gives developers a high-throughput production path for running Inkling-Small on NVIDIA Accelerated Infrastructure across multimodal, coding, and agentic workloads. Start building: https://

30 Jul 2026

Can an open-source model perform like a foundation model? @Osmosis_AI is betting yes, using reinforcement learning and the dedicated @ycombi…

HardwareDGX agent

Osmosis_AI claims that an open‑source model can rival a foundation model by leveraging reinforcement learning techniques. To demonstrate this, they will use the Y Combinator‑dedicated GPU cluster on T

Interesting data from the OpenRouter Kimi K3 dashboard: @togethercompute offers one of the lowest prices for @Kimi_Moonshot K3 while also ac…

AgentsDGX agent

Interesting data from the OpenRouter Kimi K3 dashboard: @togethercompute offers one of the lowest prices for @Kimi_Moonshot K3 while also achieving one of the highest prompt cache hit rates. Low prici

Starting in 15 minutes

ApplicationsDGX agent

Starting in 15 minutes Kimi K3 has everyone’s attention. On July 30, hear @Kimi_Moonshot's Feihu Tang explain the architecture and decisions behind it. He joins Jue Wang and Zain Hasan from Together A

29 Jul 2026

Agentic inference wastes GPUs on KV cache thrashing. ThunderAgent fixes it at the scheduler level: 2.5x higher single-node throughput and ~1…

AgentsDGX agent

Agentic inference wastes GPUs on KV cache thrashing. ThunderAgent fixes it at the scheduler level: 2.5x higher single-node throughput and ~10x lower P50 latency at high concurrency. ThunderAgent was a

It's a drop-in. One program_id field and it plugs into your existing engine configs such as KV offloading + speculative decoding. Format-agn…

HardwareDGX agent

It's a drop-in. One program_id field and it plugs into your existing engine configs such as KV offloading + speculative decoding. Format-agnostic by design, with OpenAI chat completions supported toda

Packed house today at the first edition of our @MiniMax_AI Intelligence in the Open event. Thank you to our amazing cohost @withprotegeai An…

ToolsDGX agent

Packed house today at the first edition of our @MiniMax_AI Intelligence in the Open event. Thank you to our amazing cohost @withprotegeai And our wonderful partners UpscaleX @togethercompute @Artifici

The problem: agent workflows alternate between GPU-heavy reasoning and GPU-idle waiting on tools. Run hundreds concurrently and their KV cac…

SafetyDGX agent

The problem: agent workflows alternate between GPU-heavy reasoning and GPU-idle waiting on tools. Run hundreds concurrently and their KV caches fight for memory. Engines evict on a dumb LRU policy, ev

The root cause: request-level engines never see that a series of LLM calls belongs to one longer workflow. ThunderAgent adds that missing vi…

AgentsDGX agent

The root cause: request-level engines never see that a series of LLM calls belongs to one longer workflow. ThunderAgent adds that missing view. It treats each agent workflow as a schedulable program,

28 Jul 2026

Everything you need to know about Kimi K3 with @Kimi_Moonshot and @togethercompute team. If you're looking into evaling and building with Ki…

ApplicationsDGX agent

Everything you need to know about Kimi K3 with @Kimi_Moonshot and @togethercompute team. If you're looking into evaling and building with Kimi K3 I would not miss this one! Kimi K3 has everyone’s atte

Kimi K3 has everyone’s attention. On July 30, hear @Kimi_Moonshot's Feihu Tang explain the architecture and decisions behind it. He joins Ju…

ApplicationsDGX agent

Kimi K3 has everyone’s attention. On July 30, hear @Kimi_Moonshot's Feihu Tang explain the architecture and decisions behind it. He joins Jue Wang and Zain Hasan from Together AI for a deep dive into

Open weights let builders own and control their stack, and provide flexibility to optimized quality and performance, instead of being locked…

SafetyDGX agent

Open weights let builders own and control their stack, and provide flexibility to optimized quality and performance, instead of being locked into a handful of closed platforms. We think that choice is

Robot policies can move but can't think. LLMs can think but can't move. So we connected them. Real robot: 16.7% → 97.3% Sim (LIBERO-PRO): 12…

ToolsDGX agent

A team led by Liane Galanti linked large‑language models (LLMs) with robotic motion policies, allowing a robot to combine reasoning capabilities with physical movement. In experiments on a real robot,

The apps you use swap AI models constantly. Doing that without downtime or a bad rollout reaching users is still often manual. We rebuilt th…

ToolsDGX agent

The apps you use swap AI models constantly. Doing that without downtime or a bad rollout reaching users is still often manual. We rebuilt that workflow based on what we’ve learned serving more than 40

27 Jul 2026

In partnership with @Kimi_Moonshot, we now have K3 on Together APIs with capacity for hundreds of millions of TPM on launch day!

AgentsDGX agent

In partnership with @Kimi_Moonshot, we now have K3 on Together APIs with capacity for hundreds of millions of TPM on launch day! Kimi K3 is now live on Together AI. We’re proud to be a Day 0 launch pa

Kimi K3 is now available directly from its Hugging Face model page through Together AI, give it a try!

ToolsDGX agent

Kimi K3 is now available directly from its Hugging Face model page through Together AI, give it a try! Kimi K3 in @huggingface Inference Providers is live via @togethercompute 3/M input tokens, 15/M o

Kimi K3 is now live on Together AI. We’re proud to be a Day 0 launch partner for @Kimi_Moonshot’s open frontier model, built for long-runnin…

AgentsDGX agent

Kimi K3 is now live on Together AI. We’re proud to be a Day 0 launch partner for @Kimi_Moonshot’s open frontier model, built for long-running agentic workflows across code, tools, vision, and research

@Kimi_Moonshot K3 on Together AI is built for long-running agent workflows: → 2.8T parameters and a 1M context window → Native vision for sc…

Model ReleasesDGX agent

@Kimi_Moonshot K3 on Together AI is built for long-running agent workflows: → 2.8T parameters and a 1M context window → Native vision for screenshot-guided coding → Repository navigation and terminal

Together gives developers an efficient, high-throughput production path for K3’s long, tool-heavy agent workloads, hosted on Together AI’s U…

AgentsDGX agent

Together gives developers an efficient, high-throughput production path for K3’s long, tool-heavy agent workloads, hosted on Together AI’s US-based infrastructure with zero data retention. Start build

Want to go deeper? Join Moonshot AI and Together AI for a technical webinar on how K3 was built and how to use it for production agent workf…

AgentsDGX agent

Together AI has released the Kimi K3 model on its platform as a Day‑0 launch partner for Moonshot AI’s open frontier agentic model, which supports long‑running workflows across code, tools, vision and

26 Jul 2026

Kimi K3 will be here on Monday, but how do you know you’re ready to scale it in production? With our latest inference platform updates you c…

Model ReleasesDGX agent

Kimi K3 will be here on Monday, but how do you know you’re ready to scale it in production? With our latest inference platform updates you can: 1/ Run shadow traffic to see how a new model performs on

24 Jul 2026

.@Kimi_Moonshot K3 lands on Together on Monday! We ran 452 DeepSWE rollouts against Claude Fable 5: near-flagship coding at ~35% of the pric…

Model ReleasesDGX agent

.@Kimi_Moonshot K3 lands on Together on Monday! We ran 452 DeepSWE rollouts against Claude Fable 5: near-flagship coding at ~35% of the price, and K3 pulls ahead at higher pass@k's. Full deep-dive: ht

The Washington Post processed 1.79B input tokens per month through Together AI, running open models like Llama and Mistral in production wit…

Model ReleasesDGX agent

The Washington Post processed 1.79B input tokens per month through Together AI, running open models like Llama and Mistral in production with predictable costs and full control over the model stack. T

23 Jul 2026

Catch @pavneet1990 at Builder's Day this Friday!

ToolsDGX agent

Catch @pavneet1990 at Builder's Day this Friday! I will be at @goannacapital 's Builders Day in SF this Friday, speaking alongside leaders from @ElevenLabs , @AnthropicAI and @DecagonIns If you are ar

Kimi K3 is coming to Together AI on day zero, July 27. Start building with it the moment it launches, powered by Together AI inference for c…

Model ReleasesDGX agent

Kimi K3 is coming to Together AI on day zero, July 27. Start building with it the moment it launches, powered by Together AI inference for coding, agents, and production workloads. https://www.togethe

21 Jul 2026

This is what strong inference economics unlock. @rox_ai built its own search agent, ran it in production for 6+ months, and reached 91.3% ac…

Model ReleasesDGX agent

This is what strong inference economics unlock. @rox_ai built its own search agent, ran it in production for 6+ months, and reached 91.3% accuracy at 1.03¢ per query. Together AI is proud to help powe

Watch the full fireside chat: https://www.together.ai/blog/together-yc-gpu-cluster

HardwareDGX agent

Together AI’s recent fireside chat with Y Combinator’s Andy Gupta explains why speed to launch and scalability were the key factors in selecting a partner for YC’s first dedicated GPU cluster. The con

20 Jul 2026

Together AI and @ycombinator are launching the first dedicated YC GPU cluster. Startups in the YC portfolio can now access compute on a few …

HardwareDGX agent

Together AI and Y Combinator have announced the launch of a dedicated GPU cluster for the YC portfolio. The cluster allows YC startups to access high‑performance compute with a commitment period of on

YC and Together AI are partnering to bring the first dedicated YC GPU cluster online, giving YC startups easier access to the compute they n…

HardwareDGX agent

YC and Together AI are partnering to bring the first dedicated YC GPU cluster online, giving YC startups easier access to the compute they need to build and scale. In this Founder Fireside, YC's @agup

15 Jul 2026

We made Together GPU Clusters more reliable and easier to operate. Passive health checks, guided node repair, a rebuilt Slurm stack, better …

HardwareDGX agent

We made Together GPU Clusters more reliable and easier to operate. Passive health checks, guided node repair, a rebuilt Slurm stack, better cluster visibility, external OIDC, startup scripts, and acce

14 Jul 2026

Bonsai 27B is a big step for local AI: 27B-class multimodal capability in a phone-class footprint. Congrats to @PrismML on the launch. Try T…

Local AiDGX agent

Bonsai 27B is a big step for local AI: 27B-class multimodal capability in a phone-class footprint. Congrats to @PrismML on the launch. Try Ternary Bonsai 27B on Together AI: https://www.together.ai/mo

11 Jul 2026

Why type when you can call? docs​.together​.ai​/call

ToolsDGX agent

Together AI's 'Why type when you can call?' post likely promotes their voice calling or audio interface capabilities as an alternative to text-based interactions with their AI models. The content prob

← Previous
1
Next →
263 results
← Previous
123…5
Next →