AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “together-ai--x”

GridTimelineEvolution
263 results
23 Jun 2026

LLMs write fast single-GPU kernels. Ask for a multi-GPU one and they fall apart. ParallelKernelBench (PKB) measures how they fail by benchma…

HardwareDGX agent

LLMs write fast single-GPU kernels. Ask for a multi-GPU one and they fall apart. ParallelKernelBench (PKB) measures how they fail by benchmarking against 87 problems pulled from real codebases includi

Ran 10 more tests comparing GLM 5.2 & Opus. On average, GLM 5.2 produced 2x the tokens but was still faster + 3x cheaper with similar qualit…

ToolsDGX agent

Ran 10 more tests comparing GLM 5.2 & Opus. On average, GLM 5.2 produced 2x the tokens but was still faster + 3x cheaper with similar quality! I'm open sourcing all these tests tomorrow, including the

Single-shot generation still surfaces net-new kernels with no public reference: NeMo vocab-parallel log-probs, Hyena context parallelism, SA…


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
ToolsDGX agent

Single-shot generation still surfaces net-new kernels with no public reference: NeMo vocab-parallel log-probs, Hyena context parallelism, SAM 3 mask suppression. One GEMM + All-Gather kernel hit 87.9µ

22 Jun 2026

Brrrrr 🚀 and it's free to use

ToolsDGX agent

Together AI announced the release of Brrr, a free-to-use tool or service available to users. Based on the rocket emoji and promotional framing, this likely represents a new product launch or significa

Introducing The Blind Test. Two landing pages. One built by GLM 5.2 and one by Opus 4.8. Can you tell which is which? It's very difficult to…

ToolsDGX agent

Together AI conducted a blind test comparing two landing pages—one created by GLM 5.2 and one by Opus 4.8—to evaluate whether users could distinguish between AI-generated designs. The test highlights

The next generation of inference needs purpose-built infrastructure. Together AI and 5C are deploying NVIDIA GB300 NVL72 systems with high-d…

HardwareDGX agent

The next generation of inference needs purpose-built infrastructure. Together AI and 5C are deploying NVIDIA GB300 NVL72 systems with high-density compute, advanced cooling, and AI-optimized storage f

21 Jun 2026

A year ago this would have been an obvious closed-model task. Now GLM-5.2 can read the issue, reason through the scene, patch the code, and …

ToolsDGX agent

A year ago this would have been an obvious closed-model task. Now GLM-5.2 can read the issue, reason through the scene, patch the code, and keep moving on Together AI. @togethercompute + @Zai_org GLM

Everyone’s trying to find where to test GLM-5.2. You can try it free on Together Chat (link below) No API setup. Just pick GLM-5.2 and start…

ToolsDGX agent

Everyone’s trying to find where to test GLM-5.2. You can try it free on Together Chat (link below) No API setup. Just pick GLM-5.2 and start prompting. Served by Together AI on secure North American i

Try GLM-5.2 free on Together Chat https://chat.together.ai/

ToolsDGX agent

Together AI is offering free access to GLM-5.2, an AI model, through their Together Chat interface at chat.together.ai. This announcement promotes their platform's availability for users to test the G

Voice agents get a lot more interesting when they can use the screen 🔥 This demo runs the full loop on Together AI: STT, voice, and reasoni…

ToolsDGX agent

Voice agents get a lot more interesting when they can use the screen 🔥 This demo runs the full loop on Together AI: STT, voice, and reasoning across Parakeet, MiniMax Speech 2.8, and MiniMax M3. Real-

10 Jun 2026

As vertically integrated platforms start to dominate they lock out third party access to the most valuable portions of the platform. Of cour…

SafetyDGX agent

As vertically integrated platforms start to dominate they lock out third party access to the most valuable portions of the platform. Of course, Anthropic is has the right to implement whatever policy

https://www.together.ai/blog/iso-27001-2022-certification

ToolsDGX agent

Together AI announced that it has achieved ISO 27001:2022 certification, demonstrating compliance with international information security management standards. This certification validates the company

I asked 8 AI models (including Fable 5) for their world cup predictions. Going to keep an updated leaderboard based on match results to see …

ToolsDGX agent

I asked 8 AI models (including Fable 5) for their world cup predictions. Going to keep an updated leaderboard based on match results to see which AI model performed the best! Launching tomorrow, right

Learn how @cursor_ai partnered with Together AI to deliver real-time inference for AI-powered coding in this article from @ce_zhang and @rea…

TutorialsDGX agent

Learn how @cursor_ai partnered with Together AI to deliver real-time inference for AI-powered coding in this article from @ce_zhang and @realDanFu. Cursor's in-editor agents generate code while develo

Together AI is ISO 27001:2022 certified. A-LIGN (ANAB-accredited) completed a multi-month audit of our ISMS: customer data protection, acces…

ToolsDGX agent

Together AI is ISO 27001:2022 certified. A-LIGN (ANAB-accredited) completed a multi-month audit of our ISMS: customer data protection, access controls, secure development, and incident response. Detai

9 Jun 2026

@DeepCogito needed sub-500ms time to first token at 1,000+ requests per minute for their frontier reasoning models. Together AI delivered. H…

ToolsDGX agent

@DeepCogito needed sub-500ms time to first token at 1,000+ requests per minute for their frontier reasoning models. Together AI delivered. Hear from the Deep Cogito team on what it takes to build fron

The best AI infrastructure shouldn't be reserved for the biggest companies. Together AI is partnering with @pax8 to bring powerful, cost-eff…

ApplicationsDGX agent

The best AI infrastructure shouldn't be reserved for the biggest companies. Together AI is partnering with @pax8 to bring powerful, cost-efficient AI and leading open-source models to small and mid-si

8 Jun 2026

PSA: Just added a few thousand chips, including B200s and B300s to our Dedicated Model Inference (http://api.together.ai/endpoints). With De…

Model ReleasesDGX agent

PSA: Just added a few thousand chips, including B200s and B300s to our Dedicated Model Inference (http://api.together.ai/endpoints). With Dedicated Model Inference, you can now on-click deploy our Bla

5 Jun 2026

Highlights: 👉 Design-first generation for ads, posters, packaging, and product visuals 👉 Strong typography and multilingual text rendering…

ToolsDGX agent

Highlights: 👉 Design-first generation for ads, posters, packaging, and product visuals 👉 Strong typography and multilingual text rendering 👉 Precise layout and color-palette control for brand workflow

@ideogram_ai Ideogram 4 is now available on Together AI. Try it now: http://www.together.ai/models/ideogram-40

ToolsDGX agent

Ideogram 4, an AI image generation model, has been made available on the Together AI platform. Users can now access and experiment with Ideogram 4 through Together AI's model interface at together.ai/

Introducing Ideogram 4 from @ideogram_ai on Together AI, an open image model built for design with strong text rendering, layout control, an…

ApplicationsDGX agent

Introducing Ideogram 4 from @ideogram_ai on Together AI, an open image model built for design with strong text rendering, layout control, and native 2K image generation. AI natives can now use Ideogra

4 Jun 2026

Introducing PDF to Lesson! Create interactive personalized courses from any PDF. 100% free & open source! Powered by GPT OSS on @togethercom…

ToolsDGX agent

Together AI announced a free, open-source tool called 'PDF to Lesson' that converts PDF documents into interactive, personalized courses using open-source GPT models running on Together's platform. Th

Introducing two NVIDIA Nemotron models on Together AI: Nemotron 3 Ultra for high-throughput agentic workloads and Nemotron 3.5 ASR for low-l…

Model ReleasesDGX agent

Introducing two NVIDIA Nemotron models on Together AI: Nemotron 3 Ultra for high-throughput agentic workloads and Nemotron 3.5 ASR for low-latency multilingual speech recognition. AI natives can now b

Nemotron 3.5 ASR is built for streaming multilingual speech recognition and voice agents. One 0.6B checkpoint. 40 language-locales. Sub-100m…

Model ReleasesDGX agent

Nemotron 3.5 ASR is built for streaming multilingual speech recognition and voice agents. One 0.6B checkpoint. 40 language-locales. Sub-100ms latency. Cache-aware FastConformer carries context forward

Together AI provides the inference stack behind both: high-throughput serving on the latest NVIDIA Blackwell GPUs for agentic workloads, and…

HardwareDGX agent

Together AI provides the inference stack behind both: high-throughput serving on the latest NVIDIA Blackwell GPUs for agentic workloads, and TensorRT engines plus event-driven streaming I/O for low-la

3 Jun 2026

Giving a talk this Sunday June 7th at @agihouse_org talking about DeepSeek V4 on @togethercompute! Come hang and let's talk inference!

Model ReleasesDGX agent

A Together AI representative announced a talk scheduled for Sunday, June 7th at AGI House focusing on DeepSeek V4 and inference optimization on Together Compute's platform. The event was promoted as a

I'm finally launching this agent tomorrow! Will be free, open source, and powered exclusively by open models on @togethercompute. Will also …

AgentsDGX agent

I'm finally launching this agent tomorrow! Will be free, open source, and powered exclusively by open models on @togethercompute. Will also drop a full guide on how it works! Building an agent that ca

M3 brings sparse attention + 1M context + multimodality, and Together did the hard serving work to make it fast. Great collaboration with th…

ToolsDGX agent

Anthropic's Claude 3.5 Sonnet (M3) model features sparse attention mechanisms, 1 million token context length, and multimodal capabilities, with Together AI optimizing the serving infrastructure to en

2 Jun 2026

Amazing deep dive from the @togethercompute team on serving MiniMax M3 in production. M3 with its 1M context, native multimodality and MiniM…

ApplicationsDGX agent

Amazing deep dive from the @togethercompute team on serving MiniMax M3 in production. M3 with its 1M context, native multimodality and MiniMax Sparse Attention requires real work across paged decode,

banger writeup on what it takes to make MiniMax M3 go brrrr!

ToolsDGX agent

This post likely provides technical insights and optimization strategies for running or fine-tuning the MiniMax M3 model efficiently, discussing performance enhancements and practical implementation t

Everyone talks about 1M context. The harder part is making 1M context actually usable. Serving MiniMax M3 required optimizing for long-conte…

AgentsDGX agent

Everyone talks about 1M context. The harder part is making 1M context actually usable. Serving MiniMax M3 required optimizing for long-context, multimodal, and agentic workloads simultaneously. Excite

First technical Deepdive on M3 on the internet😎

HardwareDGX agent

First technical Deepdive on M3 on the internet😎 MiniMax-M3 combines 1M context, native multimodality, and MiniMax Sparse Attention. The next layer is serving it efficiently: KV-block-major sparse atte

Going live now with @MiniMax_AI 🚀 https://x.com/i/spaces/1nxeLLDDBEaJX

ToolsDGX agent

Together AI announced a live event featuring MiniMax AI, likely discussing a partnership, integration, or collaborative development initiative. The announcement was made via X Spaces, Twitter's audio

MiniMax-M3 combines 1M context, native multimodality, and MiniMax Sparse Attention. The next layer is serving it efficiently: KV-block-major…

HardwareDGX agent

MiniMax-M3 combines 1M context, native multimodality, and MiniMax Sparse Attention. The next layer is serving it efficiently: KV-block-major sparse attention, paged MSA decode, optimized index scoring

We wrapped a live session on M3 yesterday with the @togethercompute team & our researchers @zpysky1125 and @HaohaiSun A few highlights 🧵 1.…

Model ReleasesDGX agent

We wrapped a live session on M3 yesterday with the @togethercompute team & our researchers @zpysky1125 and @HaohaiSun A few highlights 🧵 1. MSA (MiniMax Sparse Attention) is the star ⭐️. Unlike CSA/HC

1 Jun 2026

Make sure to join our live Spaces chat on MiniMax M3 starting in 4 hours. You can pre-submit questions by replying to this tweet.

ToolsDGX agent

Make sure to join our live Spaces chat on MiniMax M3 starting in 4 hours. You can pre-submit questions by replying to this tweet. MiniMax M3 is live and Together AI is powering its inference 🚀 Tomorro

Minimax M3 is a remarkable model in many ways, but esp for inference efficiency. Minimax model designers and @togethercompute inference team…

ToolsDGX agent

Minimax M3 is a remarkable model in many ways, but esp for inference efficiency. Minimax model designers and @togethercompute inference team will discuss it tomorrow live. MiniMax M3 is live and Toget

MiniMax M3 is live and Together AI is powering its inference 🚀 Tomorrow at 6pm PT we're going live on X Spaces with the teams behind the mo…

ToolsDGX agent

MiniMax M3 is live and Together AI is powering its inference 🚀 Tomorrow at 6pm PT we're going live on X Spaces with the teams behind the model and the infrastructure to give you a deep dive. https://x

See you tomorrow night. Come with questions.

ToolsDGX agent

See you tomorrow night. Come with questions. We're going LIVE tomorrow with @togethercompute 🔥. @zpysky1125 is pulling back the curtain on M3: sparse attention, 1M context, all of it. You don't want t

Speakers: Pengyu Zhao, Head of Research at MiniMax Haohai Sun, Research Scientist at MiniMax Ce Zhang, Founder/CTO at Together Yineng Zhang,…

ToolsDGX agent

Speakers: Pengyu Zhao, Head of Research at MiniMax Haohai Sun, Research Scientist at MiniMax Ce Zhang, Founder/CTO at Together Yineng Zhang, Senior Director at Together Dan Fu, VP of Kernels at Togeth

We'll get into @MiniMax_AI M3's model performance, the MSA architecture and what it means for long context, and how Together is optimizing i…

ToolsDGX agent

We'll get into @MiniMax_AI M3's model performance, the MSA architecture and what it means for long context, and how Together is optimizing inference and KV-cache for this new architecture. Set your re

We're going LIVE tomorrow with @togethercompute 🔥. @zpysky1125 is pulling back the curtain on M3: sparse attention, 1M context, all of it. …

ToolsDGX agent

We're going LIVE tomorrow with @togethercompute 🔥. @zpysky1125 is pulling back the curtain on M3: sparse attention, 1M context, all of it. You don't want to miss this. MiniMax M3 is live and Together

30 May 2026

Learn more about the latest from @james_y_zou and our Frontier Agents Research team!

AgentsDGX agent

Learn more about the latest from @james_y_zou and our Frontier Agents Research team! To evaluate frontier AI agents, we need more complex tasks. But such tasks are also more prone to have design mista

We took the Hot Wings Challenge to NVIDIA GTC 🌶️ @realDanFu (VP of Kernels) and @sarung (VP of Customer Success) answered some questions ar…

HardwareDGX agent

We took the Hot Wings Challenge to NVIDIA GTC 🌶️ @realDanFu (VP of Kernels) and @sarung (VP of Customer Success) answered some questions around AI, one spicy wing at a time. Some people sweat. Some pe

29 May 2026

Together AI serves the two fastest STT models measured by @ArtificialAnlys NVIDIA Parakeet-TDT 0.6B v3 can transcribe 20 hours of speech in …

HardwareDGX agent

Together AI serves the two fastest STT models measured by @ArtificialAnlys NVIDIA Parakeet-TDT 0.6B v3 can transcribe 20 hours of speech in under 10 seconds. This deep dive shows the systems work behi

28 May 2026

Every LLM API call depends on the inference engine underneath it. Tokenization, scheduling, prefill, decode, KV cache, batching, and streami…

ApplicationsDGX agent

Every LLM API call depends on the inference engine underneath it. Tokenization, scheduling, prefill, decode, KV cache, batching, and streaming determine whether the experience is fast, scalable, and p

27 May 2026

The best research labs are building what comes after static models. Congrats to @trajectorylabs on the launch! Excited to have them training…

ToolsDGX agent

The best research labs are building what comes after static models. Congrats to @trajectorylabs on the launch! Excited to have them training on the AI Native Cloud as they push the frontier on Continu

24 May 2026

Our inference stack, optimized for Blackwells, with a novel attention kernel and many new optimizations has started rolling out! It's alread…

HardwareDGX agent

Our inference stack, optimized for Blackwells, with a novel attention kernel and many new optimizations has started rolling out! It's already charting on Artificial Analysis, eg: #1 speed and latency

22 May 2026

Highlights: 👉 Long-horizon autonomy: maintained coherent execution across a 35-hour autonomous kernel optimization run 👉 Agentic coding: l…

AgentsDGX agent

Highlights: 👉 Long-horizon autonomy: maintained coherent execution across a 35-hour autonomous kernel optimization run 👉 Agentic coding: leading Terminal-Bench 2.0-Terminus performance for terminal-ba

Introducing Qwen3.7-Max from @Alibaba_Qwen, Qwen’s flagship model for the agent era with 1M context and leading performance across agentic c…

Model ReleasesDGX agent

Introducing Qwen3.7-Max from @Alibaba_Qwen, Qwen’s flagship model for the agent era with 1M context and leading performance across agentic coding, reasoning, and long-horizon autonomy. AI natives can

PSA: Just added a thousand H100s and H200s to Together on-demand GPU clusters and Dedicated Endpoints: http://api.together.ai/clusters

HardwareDGX agent

Together AI announced the addition of 1,000 H100 and H200 GPUs to their on-demand GPU clusters and dedicated endpoints, expanding their available compute resources for users accessing services through

Try Qwen3.7-Max now on Together AI: http://www.together.ai/models/qwen37-max

ToolsDGX agent

Together AI announced the availability of Qwen3.7-Max, a large language model, on their platform. The announcement promotes users to try the model through Together AI's interface. Qwen3.7-Max is part

21 May 2026

We turned this app into a real comic book booth at GTC! Had folks create their own comic books on-site, printed them, and gave them out! (th…

ToolsDGX agent

We turned this app into a real comic book booth at GTC! Had folks create their own comic books on-site, printed them, and gave them out! (this is using an unreleased version of the app that's a lot be

20 May 2026

MiniMax Speech 2.8 Turbo is built for voice agents that need natural delivery, not just clean audio. → Sound Tags for laughter, breathing, s…

ToolsDGX agent

MiniMax Speech 2.8 Turbo is built for voice agents that need natural delivery, not just clean audio. → Sound Tags for laughter, breathing, sighs, gasps, and other vocal cues → 60% prosody improvement

See MiniMax Speech 2.8 Turbo in voice finder and try the voices directly: https://voicefinder.together.ai/minimax--speech-2.8-turbo Learn mo…

TutorialsDGX agent

MiniMax Speech 2.8 Turbo is a text-to-speech model available through Together AI's voice finder tool at voicefinder.together.ai, allowing users to preview and test different voice options directly. Th

We added 600+ new voices on Together AI! Introducing MiniMax Speech 2.8 Turbo on Together AI, an enterprise TTS model for expressive real-ti…

ApplicationsDGX agent

We added 600+ new voices on Together AI! Introducing MiniMax Speech 2.8 Turbo on Together AI, an enterprise TTS model for expressive real-time voice agents. AI natives can now deploy @MiniMax_AI Speec

19 May 2026

'One thing that we've been seeing recently is that inference benchmarks don't really match production workloads that well.' - @realDanFu, VP…

Model ReleasesDGX agent

'One thing that we've been seeing recently is that inference benchmarks don't really match production workloads that well.' - @realDanFu, VP of Kernels When you're running dozens of concurrent coding

Voice 'cloning' is style transfer. Across three widely used systems — ElevenLabs V3, Coqui-XTTS, Chatterbox — clones don't just copy speaker…

ToolsDGX agent

Voice 'cloning' is style transfer. Across three widely used systems — ElevenLabs V3, Coqui-XTTS, Chatterbox — clones don't just copy speakers, they reshape them to be warmer, more authoritative, more

18 May 2026

Congrats to the @cursor_ai team on Composer 2.5 — a huge milestone for agentic coding models. Together AI, the AI Native Cloud, is proud to …

AgentsDGX agent

Congrats to the @cursor_ai team on Composer 2.5 — a huge milestone for agentic coding models. Together AI, the AI Native Cloud, is proud to partner on this launch. Composer 2.5 is pushing the frontier

16 May 2026

Gemma-4-31B-it-Pearl supports 32K context, configurable thinking, function calling, and JSON mode. This is Together AI’s first Pearl-powered…

Model ReleasesDGX agent

Gemma-4-31B-it-Pearl supports 32K context, configurable thinking, function calling, and JSON mode. This is Together AI’s first Pearl-powered endpoint. Eventually, we plan to expand our Pearl powered p

← Previous
12345
Next →