AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “together-ai--x”

GridTimelineEvolution
263 results
9 Jul 2026

Congrats to @ollama on the fundraise and 9M+ active builders.🚀 Open models are becoming the default path for developers who want to build, …

Local AiDGX agent

Congrats to @ollama on the fundraise and 9M+ active builders.🚀 Open models are becoming the default path for developers who want to build, run, and own their AI stack. We're excited to support the eco

8 Jul 2026

Catch us at @RaiseSummit! Our CEO @vipulved sits down with @wolfejosh of @Lux_Capital for a fireside chat 'AI Sticker Shock: Open Source as …

ToolsDGX agent

Catch us at @RaiseSummit! Our CEO @vipulved sits down with @wolfejosh of @Lux_Capital for a fireside chat 'AI Sticker Shock: Open Source as the Answer to Broken Tokenomics' Join us on the Master Stage

Provisioned Throughput is a new serverless product from @togethercompute w/ guaranteed TPMs. We created this for mission critical enterprise…


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
ApplicationsDGX agent

Provisioned Throughput is a new serverless product from @togethercompute w/ guaranteed TPMs. We created this for mission critical enterprise class applications. It functions like dedicated capacity bu

Read our blog: https://www.together.ai/blog/provisioned-throughput

ToolsDGX agent

Together AI announced provisioned throughput capabilities, a feature that allows users to reserve dedicated model inference capacity at predictable costs. This offering enables organizations to optimi

We're introducing Provisioned Throughput: reserved inference capacity for frontier open models, with token-based pricing and a 99% uptime SL…

ToolsDGX agent

We're introducing Provisioned Throughput: reserved inference capacity for frontier open models, with token-based pricing and a 99% uptime SLA. Serverless simplicity, guaranteed capacity, up to 90% low

7 Jul 2026

GLM 5.2 on Together AI is now #1 on Artificial Analysis for both output speed and latency.

ToolsDGX agent

GLM 5.2, running on Together AI's platform, has achieved the top ranking on Artificial Analysis benchmarks for both output speed and latency metrics. This announcement highlights Together AI's infrast

Introducing AI Browser Games! Watch open & closed models build small browser games head to head. Open models like Kimi K2.7 were faster, che…

ToolsDGX agent

Introducing AI Browser Games! Watch open & closed models build small browser games head to head. Open models like Kimi K2.7 were faster, cheaper, & produced games similar to Opus 4.8. In some cases, o

This is what latency optimization looks like below the API 👇 Together ATLAS, NVIDIA Blackwell, CUDA, TensorRT-LLM, Dynamo, and custom kerne…

HardwareDGX agent

This is what latency optimization looks like below the API 👇 Together ATLAS, NVIDIA Blackwell, CUDA, TensorRT-LLM, Dynamo, and custom kernels all working together to make inference faster for users. P

6 Jul 2026

'Alex really hit the nail on the head, by sending data to a company that has extremely smart models, you are really giving up your business'…

ToolsDGX agent

'Alex really hit the nail on the head, by sending data to a company that has extremely smart models, you are really giving up your business's recipe for them to copy.' Our CEO @vipulved when asked abo

Join us for a fireside chat on where AI research and infrastructure are headed, led by @tri_dao. Hosted by Together AI, @nvidia and Lyra Lab…

HardwareDGX agent

Together AI, in partnership with NVIDIA and Lyra Lab, hosted a fireside chat led by Tri Dao discussing current trends and future directions in AI research and infrastructure development. The event bro

4 Jul 2026

Open-source models give you complete control, customization, and ownership over your data. Companies are moving fast on this. @vipulved on @…

ToolsDGX agent

Open-source AI models provide users with full control, customization capabilities, and data ownership, addressing privacy and autonomy concerns compared to proprietary alternatives. The post highlight

3 Jul 2026

Open model usage has gone from 10% of AI tokens to 30% in a year. The shift to open, modular AI is here to stay. Our founders on what's driv…

ToolsDGX agent

Open source AI model usage has tripled from 10% to 30% of total AI tokens consumed within a year, reflecting a significant market shift toward open and modular AI architectures. This trend indicates g

Read more on the economy of tokens by @vipulved https://x.com/vipulved/status/2071404852908081211?s=20

ToolsDGX agent

This is a discussion about token economics, likely covering how tokens function as units of value or utility in blockchain and AI systems, shared by Vipul Ved on X (formerly Twitter) and referenced by

We analyzed GLM 5.2 vs Sonnet 5 for software engineering tasks using DeepSWE. GLM 5.2 gets you ~80% of Sonnet 5's capability at ~20% of the …

Model ReleasesDGX agent

We analyzed GLM 5.2 vs Sonnet 5 for software engineering tasks using DeepSWE. GLM 5.2 gets you ~80% of Sonnet 5's capability at ~20% of the price. More insights in the thread! Deepdive: Sonnet 5 and G

We're releasing the full slides for our 2 hr deepdive session from the AI Engineer World's Fair. We covered how we build inference engines t…

HardwareDGX agent

We're releasing the full slides for our 2 hr deepdive session from the AI Engineer World's Fair. We covered how we build inference engines to serve agentic workloads at trillion token production scale

2 Jul 2026

30 billion tokens a month to 400 trillion in a year. That's @cursor_ai, @DecagonAI, @cartesia and hundreds of other teams choosing open infr…

ToolsDGX agent

Together AI highlights the rapid growth of AI token consumption across multiple companies and teams, noting an increase from 30 billion tokens monthly to 400 trillion annually, with platforms like Cur

@DecagonAI @AshwinSreenivas Under the hood: 6x cost reduction per turn, p95 latency under 400ms, and models shipping weekly. https://www.tog…

ToolsDGX agent

Together AI announced significant performance improvements in their AI infrastructure, achieving a 6x cost reduction per turn while maintaining p95 latency under 400ms, with new models being released

LIVE at 12p PT/3p ET: AI’s cleanest stories are getting messy. Meta has excess compute. AI-heavy companies are hiring, not shrinking. And op…

ToolsDGX agent

LIVE at 12p PT/3p ET: AI’s cleanest stories are getting messy. Meta has excess compute. AI-heavy companies are hiring, not shrinking. And open models keep getting better. @DanielTNiles on what this me

Our CEO @vipulved on @CNBC with @dee_bosa: your data is your recipe. As models get smarter, sending proprietary workflows, customer context,…

ToolsDGX agent

Our CEO @vipulved on @CNBC with @dee_bosa: your data is your recipe. As models get smarter, sending proprietary workflows, customer context, and business logic into closed systems becomes a strategic

See how our latest funding round accelerates the shift to open-source AI: https://www.together.ai/blog/announcing-our-series-c

ToolsDGX agent

Together AI announced a Series C funding round aimed at accelerating the adoption and development of open-source AI models and infrastructure. The funding enables the company to expand its platform fo

1 Jul 2026

1/ DSGym: A Holistic Framework for Evaluating and Training Data Science Agents Paper: https://arxiv.org/abs/2601.16344

ToolsDGX agent

DSGym is a comprehensive framework designed to evaluate and train AI agents for data science tasks, providing a structured environment for benchmarking agent performance across various data science wo

2/ ThunderAgent: A Simple, Fast and Program-Aware Agentic Inference System Paper: https://arxiv.org/abs/2602.13692

AgentsDGX agent

ThunderAgent is an agentic inference system designed for fast and efficient execution of AI agent programs, developed by Together AI. The system appears to optimize program-aware inference by leveragi

3/ Learning to Discover at Test Time (TTT-Discover) Paper: https://arxiv.org/abs/2601.16175

ToolsDGX agent

TTT-Discover is a method that enables models to learn and discover patterns during test time rather than only during training, allowing for adaptation to new data distributions at inference. The appro

4/ Escaping the Verifier: Learning to Reason via Demonstrations (RARO) Paper: https://arxiv.org/abs/2511.21667

ToolsDGX agent

RARO (Reasoning via Demonstrations) is a method for training AI models to improve reasoning capabilities by learning from demonstrations rather than relying solely on external verifiers. The approach

5/ V1: Unifying Generation and Self-Verification for Parallel Reasoners Paper: https://arxiv.org/abs/2603.04304

ToolsDGX agent

This paper presents V1, a framework that unifies text generation with self-verification mechanisms to enable parallel reasoning processes in language models. The approach allows models to generate mul

6/ When RL Meets Adaptive Speculative Training: A Unified Training-Serving System (Aurora) Paper: https://arxiv.org/abs/2602.06932

ToolsDGX agent

Aurora is a unified training-serving system that integrates reinforcement learning with adaptive speculative training to optimize large language model inference and training efficiency. The system dyn

7/ Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Paper: https://arxiv.org/abs/2602.21196

ToolsDGX agent

Untied Ulysses is a memory-efficient technique for context parallelism that processes attention heads in chunks rather than sequences, reducing memory overhead during transformer inference and trainin

8/ Opportunistic Expert Activation: Batch-Aware Expert Routing for Faster Decode Without Retraining (OEA) Paper: https://arxiv.org/abs/2511.…

ToolsDGX agent

Opportunistic Expert Activation (OEA) is a batch-aware expert routing technique for mixture-of-experts models that enables faster decoding without requiring model retraining. The method optimizes whic

9/ ParallelKernelBench: Benchmarking LLMs on Multi-GPU Kernel Generation Paper: https://www.alphaxiv.org/abs/2606.parallel-kernel-bench

HardwareDGX agent

ParallelKernelBench is a benchmarking framework designed to evaluate large language models' ability to generate optimized GPU kernels for multi-GPU computing environments. The benchmark assesses LLMs

Intelligence should be abundant, not expensive.

ToolsDGX agent

Together AI advocates for democratizing artificial intelligence by making it abundant and affordable rather than concentrated among expensive proprietary systems. The post likely argues for open-sourc

Missed our talk with @togethercompute on @MiniMax_AI sparse attention and kernel optimizations? Catch us tomorrow at 10:45am PT.

ToolsDGX agent

Together AI is promoting a rescheduled presentation about sparse attention mechanisms and kernel optimizations in collaboration with MiniMax AI, scheduled for 10:45am PT. The session focuses on techni

Multi-GPU kernels are the real test for coding models. Today at @aiDotEngineer, @simran_s_arora shared ParallelKernelBench, an open-source b…

Model ReleasesDGX agent

Multi-GPU kernels are the real test for coding models. Today at @aiDotEngineer, @simran_s_arora shared ParallelKernelBench, an open-source benchmark for evaluating whether LLMs can write fast CUDA ker

Nemotron 3 Ultra is taking off on Together AI 🚀 In just days, it's climbed to 35B tokens/day on @OpenRouter. That's the community voting wi…

Model ReleasesDGX agent

Nemotron 3 Ultra is taking off on Together AI 🚀 In just days, it's climbed to 35B tokens/day on @OpenRouter. That's the community voting with their tokens for open models that are fast, efficient, and

Our research team has 9 papers at ICML next week! Spanning the full stack from frontier agents to GPU kernels, we're excited to share what o…

HardwareDGX agent

Our research team has 9 papers at ICML next week! Spanning the full stack from frontier agents to GPU kernels, we're excited to share what our researchers and collaborators have been working on. If yo

Read about our Series C in the New York Times: http://shorturl.at/SooOP

ToolsDGX agent

Together AI announced their Series C funding round, which was covered in a reporting piece by the New York Times. The announcement was shared via their official X (formerly Twitter) account, linking t

Read more about why we built this and where we're going. https://www.together.ai/blog/announcing-our-series-c

ToolsDGX agent

Together AI announced a Series C funding round and shared details about their company vision, current capabilities, and future direction in open-source AI development and inference optimization. The a

Read more on ParallelKernelBench: https://www.together.ai/blog/parallelkernelbench

ToolsDGX agent

ParallelKernelBench is a benchmark tool or methodology developed by Together AI for evaluating the performance of parallel kernel execution in machine learning systems. The benchmark likely measures m

Read the blog: https://www.together.ai/blog/icml-2026 See us at ICML: https://www.together.ai/icml-2026

ToolsDGX agent

Together AI is announcing their participation in ICML 2026 (International Conference on Machine Learning) and inviting the community to visit them at the conference. The announcement likely includes d

Tasks moved from closed models to open models on Together AI cost one-fifth to one-seventh as much. That's from @DecagonAI co-founder @Ashwi…

ToolsDGX agent

Tasks moved from closed models to open models on Together AI cost one-fifth to one-seventh as much. That's from @DecagonAI co-founder @AshwinSreenivas in the NYT today. This is what abundant intellige

We gave a 2 hr deepdive on how to build inference engines that handle trillion token agentic workloads at @aiDotEngineer. Will drop slides a…

AgentsDGX agent

Together AI presented a 2-hour technical deep dive on building inference engines capable of handling trillion-token agentic workloads, covering architecture and optimization strategies for large-scale

30 Jun 2026

As open models get stronger, more workloads move into the competitive inference market. That pushes the real fight toward speed, cost, relia…

ApplicationsDGX agent

As open models get stronger, more workloads move into the competitive inference market. That pushes the real fight toward speed, cost, reliability, and control. Together AI is where open models become

29 Jun 2026

Open models are not just a pricing story. They are what happens when the AI stack becomes modular: models, APIs, harnesses, tools, and infer…

ToolsDGX agent

Open models are not just a pricing story. They are what happens when the AI stack becomes modular: models, APIs, harnesses, tools, and inference all improving independently. Together AI is building th

Submit your winners here! https://www.aicup.io Thanks to @nutlope and @federicobianchy for bringing this to life 🙌

ToolsDGX agent

Together AI announced a submission portal for winners at aicup.io, crediting @nutlope and @federicobianchy for creating the initiative. This appears to be related to an AI Cup competition or challenge

To celebrate the start of @aiDotEngineer AI Engineer World's Fair, we're launching a bracket competition! Use an agent or pick manually to c…

AgentsDGX agent

To celebrate the start of @aiDotEngineer AI Engineer World's Fair, we're launching a bracket competition! Use an agent or pick manually to choose your winners for the Round of 32 by 9:00 am PT Tuesday

28 Jun 2026

More reason why we’re excited about GLM-5.2 on Together 👇 Strong enough for serious coding work, cheap enough to change routing decisions, …

Model ReleasesDGX agent

More reason why we’re excited about GLM-5.2 on Together 👇 Strong enough for serious coding work, cheap enough to change routing decisions, and easy to access through the tools developers already use.

There's a big difference between a single model call and serving an agent at scale. @ZainHasan6 breaks down what actually changes. Catch our…

AgentsDGX agent

There's a big difference between a single model call and serving an agent at scale. @ZainHasan6 breaks down what actually changes. Catch our team this Monday at 9 a.m. PST for their open-source infere

27 Jun 2026

next week at @aiDotEngineer, we are joining @togethercompute for a conversation on what goes into running agents at scale. @olive_jy_song, R…

AgentsDGX agent

next week at @aiDotEngineer, we are joining @togethercompute for a conversation on what goes into running agents at scale. @olive_jy_song, Research Lead, RL at MiniMax, and @realDanFu, VP of Kernels a

26 Jun 2026

As token usage explodes, model choice becomes product strategy. Teams are already testing models like GLM-5.2 because they want frontier qua…

ToolsDGX agent

As token usage explodes, model choice becomes product strategy. Teams are already testing models like GLM-5.2 because they want frontier quality, better tokenomics, and more control over cost, data, a

What happens when AI agents collaborate on open science? At @aiDotEngineer World’s Fair, @james_y_zou will share work on EinsteinArena and D…

AgentsDGX agent

What happens when AI agents collaborate on open science? At @aiDotEngineer World’s Fair, @james_y_zou will share work on EinsteinArena and DSGym, from multi-agent math discovery to better evaluation f

25 Jun 2026

Cost per iteration is the unlock. GLM-5.2 on Together AI can generate polished web apps for a few cents. At that price, developers can explo…

ToolsDGX agent

Cost per iteration is the unlock. GLM-5.2 on Together AI can generate polished web apps for a few cents. At that price, developers can explore more directions, compare more versions, and keep the best

I love using GLM 5.2 for web app iteration. My workflow: generate 6 variations, then pick the best one and continue iterating on it. I built…

AgentsDGX agent

I love using GLM 5.2 for web app iteration. My workflow: generate 6 variations, then pick the best one and continue iterating on it. I built Recast to make this even easier. Give it a prompt, get 6 va

LLMs are getting better at writing GPU kernels. Multi-GPU kernels are the harder test. At @aiDotEngineer World's Fair, @simran_s_arora will …

Model ReleasesDGX agent

LLMs are getting better at writing GPU kernels. Multi-GPU kernels are the harder test. At @aiDotEngineer World's Fair, @simran_s_arora will share ParallelKernelBench, an open-source benchmark built fr

Together AI built the world’s fastest speech-to-text stack. Parakeet on Together transcribes ~302 seconds of audio per second of processing …

ToolsDGX agent

Together AI built the world’s fastest speech-to-text stack. Parakeet on Together transcribes ~302 seconds of audio per second of processing time, the top speed factor reported by @ArtificialAnlys. In

24 Jun 2026

400T tokens is what production adoption looks like. Teams are moving real workloads to open models because they want frontier quality, bette…

ApplicationsDGX agent

400T tokens is what production adoption looks like. Teams are moving real workloads to open models because they want frontier quality, better tokenomics, and more control over inference. Together AI g

A tangible comparison of GLM performance stacked against Opus 4.8 on web tasks by @nutlope. GLM 5.2 is chattier, but still faster when serve…

ToolsDGX agent

A tangible comparison of GLM performance stacked against Opus 4.8 on web tasks by @nutlope. GLM 5.2 is chattier, but still faster when served by @togethercompute and over 3x cheaper. Announcing GLM Ar

Agentic coding changes what inference engines need to handle. At AI Engineer World’s Fair, Together AI engineers will lead a hands-on worksh…

AgentsDGX agent

Agentic coding changes what inference engines need to handle. At AI Engineer World’s Fair, Together AI engineers will lead a hands-on workshop on how inference engines work and what it takes to serve

Announcing GLM Arena! A series of tests (infographics, svgs, sites, ect..) ran on GLM 5.2 and Opus 4.8, with prompts included. On average, G…

ToolsDGX agent

Announcing GLM Arena! A series of tests (infographics, svgs, sites, ect..) ran on GLM 5.2 and Opus 4.8, with prompts included. On average, GLM 5.2 produced 2x the tokens but was still faster + 3x chea

Read the full story: https://www.theinformation.com/newsletters/applied-ai/open-source-growth-boosts-together-ai-hugging-face

ToolsDGX agent

Together AI highlights how the growth of open-source AI models is benefiting their platform and the broader ecosystem. The article likely discusses how open-source initiatives, including collaboration

23 Jun 2026

An agentic loop (compile, test, profile, revise) helps. Gemini 3 Pro went from 24 to 35/87 correct, then plateaued after ~20 steps. Feedback…

Model ReleasesDGX agent

An agentic loop (compile, test, profile, revise) helps. Gemini 3 Pro went from 24 to 35/87 correct, then plateaued after ~20 steps. Feedback fixes syntax, not rank coordination, collective ordering, o

.@cartesia runs one of the hardest inference workloads: real-time voice. Their stack has to keep long-lived streams moving, serve millions o…

HardwareDGX agent

.@cartesia runs one of the hardest inference workloads: real-time voice. Their stack has to keep long-lived streams moving, serve millions of audio minutes a day, and hold model latency around 90ms. T

← Previous
12345
Next →