AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

Search: “fireworks-ai--x”

GridTimelineEvolution
165 results
1 Jul 2026

Fireworks Batch API: 50% cheaper than serverless. We obviously love things fast, but sometimes async at scale is all you need. With the refr…

ToolsDGX agent

Fireworks Batch API: 50% cheaper than serverless. We obviously love things fast, but sometimes async at scale is all you need. With the refreshed Batch API, you queue up a job and select whether you n

Frontier open models with enterprise governance means rebuilding workflows from scratch. Move faster with Fireworks on Foundry. GLM 5.2 is l…

ApplicationsDGX agent

Frontier open models with enterprise governance means rebuilding workflows from scratch. Move faster with Fireworks on Foundry. GLM 5.2 is live on Microsoft Foundry. With FireConnect enabled, devs can

If you want frontier-level coding and agent performance but you don't want to pay closed-model prices, GLM 5.2 is the open model you've prob…


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
AgentsDGX agent

If you want frontier-level coding and agent performance but you don't want to pay closed-model prices, GLM 5.2 is the open model you've probably been hearing about. Here's why you should run it on Fir

This is exactly why we believe in customization. Quick context: Factory's original secret scanner was deterministic, so it either flagged th…

Model ReleasesDGX agent

This is exactly why we believe in customization. Quick context: Factory's original secret scanner was deterministic, so it either flagged things that weren't actually secrets (false positives) or miss

30 Jun 2026

For even higher speeds, reach out for a custom deployment! We’ve hit 446 tok/s on Artificial Analysis. Learn more → https://fireworks.ai/blo…

TutorialsDGX agent

Fireworks AI announced achieving 446 tokens per second throughput speeds as measured by Artificial Analysis benchmarks, positioning this as their standard performance metric. The company offers custom

inference reliability has historically been a tax on devs that only large well-funded startups could afford: reserve GPUs in advance, sign a…

ApplicationsDGX agent

inference reliability has historically been a tax on devs that only large well-funded startups could afford: reserve GPUs in advance, sign a contract, guess your peak throughput requirements. everyone

We heard your feedback. You want to go faster. Introducing GLM 5.2 Fast The same model and quality as GLM 5.2 standard, now at 140 tok/s Fli…

ToolsDGX agent

Fireworks AI announced GLM 5.2 Fast, an optimized version of their GLM 5.2 model that maintains the same quality and capabilities as the standard version while delivering significantly faster inferenc

29 Jun 2026

Going to be in Paris for @RaiseSummit? Join us and @nvidia to hear from leaders at both companies on the future of AI inference and infrastr…

HardwareDGX agent

Going to be in Paris for @RaiseSummit? Join us and @nvidia to hear from leaders at both companies on the future of AI inference and infrastructure. After that: cocktails and a DJ. Oh la la! See you th

Research —> Product :) very excited to start rolling out our fine-tuned Trace Judge model built earlier this month with the great @Fireworks…

AgentsDGX agent

Research —> Product :) very excited to start rolling out our fine-tuned Trace Judge model built earlier this month with the great @FireworksAI_HQ team there’s a mountain of Agent Improvement gold sitt

28 Jun 2026

you may have heard that glm-5.2 at 392 token/s is cool, how about 446 except… it’s all noise. Artificial Analysis picks median among 8 point…

Model ReleasesDGX agent

you may have heard that glm-5.2 at 392 token/s is cool, how about 446 except… it’s all noise. Artificial Analysis picks median among 8 points/day so first point of the day can be way off looking at 3

27 Jun 2026

DSpark from @deepseek_ai ingeniously integrates many speculative decoding ideas to achieve 1.5x to 5x higher throughput in a real production…

ApplicationsDGX agent

DSpark from @deepseek_ai ingeniously integrates many speculative decoding ideas to achieve 1.5x to 5x higher throughput in a real production system Let's understand it with 10 ideas, starting from the

Model management is the real SDLC scaling bottleneck. @FactoryAI standardized on Fireworks to solve it: → 2–3x open-model growth → 5–15x mor…

ApplicationsDGX agent

Model management is the real SDLC scaling bottleneck. @FactoryAI standardized on Fireworks to solve it: → 2–3x open-model growth → 5–15x more work per dollar → day-0 access to every new open-weight mo

26 Jun 2026

Fireworks AI is now live on EvoSkill v1.3.0! You can now use @FireworksAI_HQ directly with EvoSkill to run fast inference on open models as …

Model ReleasesDGX agent

Fireworks AI is now live on EvoSkill v1.3.0! You can now use @FireworksAI_HQ directly with EvoSkill to run fast inference on open models as both the evolution harness backend and the LLM scorer. Along

The big lesson from training @cursor_ai Composer 2: models exploit flaws in their training environment before learning what you actually wan…

ApplicationsDGX agent

The big lesson from training @cursor_ai Composer 2: models exploit flaws in their training environment before learning what you actually want. Real RL for coding agents means production-faithful envir

We hosted the first RSI RL Environments hackathon with @hud_evals @ycombinator and it was a blast! We watched builders treat RL as a general…

ToolsDGX agent

We hosted the first RSI RL Environments hackathon with @hud_evals @ycombinator and it was a blast! We watched builders treat RL as a general purpose tool and reach for it across domains we had not ant

Well said @RamaswmySridhar! I believe the cost saving is actually much bigger, more like 4-5x. E.g., we have just reduced our GLM 5.2 cached…

ToolsDGX agent

Well said @RamaswmySridhar! I believe the cost saving is actually much bigger, more like 4-5x. E.g., we have just reduced our GLM 5.2 cached token price by 2X to return efficiency gain to our users. A

25 Jun 2026

But here's the punchline. Normalized to 90% cache hit rate: GLM-5.2 (Fireworks): 1.12/session Opus-4.7 (Anthropic): 2.14/session GLM is ~4…

ToolsDGX agent

This post compares the cost efficiency of Fireworks' GLM-5.2 model versus Anthropic's Opus across cached sessions, showing GLM-5.2 achieving approximately 4x lower cost per session at 90% cache hit ra

Congrats! Open source GLM model is really a game changer! Extremely fast, cheap, and high quality!

ToolsDGX agent

Fireworks AI announced the release of an open source GLM model that offers significant improvements in speed, cost efficiency, and output quality compared to existing alternatives. The post suggests t

In a joint Fireworks and @Faros_AI evaluation of 211 real engineering tasks, Claude Code + GLM-5.2 beat both Claude Code + Opus 4.8 and Code…

Model ReleasesDGX agent

In a joint Fireworks and @Faros_AI evaluation of 211 real engineering tasks, Claude Code + GLM-5.2 beat both Claude Code + Opus 4.8 and Codex + GPT-5.5: - Judge score: 0.568 vs. 0.521 and 0.466 - Time

Open + closed models = better together. Our previous research with @harvey showed the benefits of combining a frontier closed model as an ad…

AgentsDGX agent

Open + closed models = better together. Our previous research with @harvey showed the benefits of combining a frontier closed model as an advisor agent with fine-tuned, open-source worker agents. Thre

RL fine-tuning is now live for @nvidiaai Nemotron 3 on Fireworks, starting with Nemotron 3 Super (LoRA). Train with GRPO and serve the model…

Model ReleasesDGX agent

RL fine-tuning is now live for @nvidiaai Nemotron 3 on Fireworks, starting with Nemotron 3 Super (LoRA). Train with GRPO and serve the model in one place. We price by GPU-hour, not per token, so long

The underrated part of this announcement is that Fireworks has been quietly great behind the scenes helping us eval and serve these models B…

Model ReleasesDGX agent

The underrated part of this announcement is that Fireworks has been quietly great behind the scenes helping us eval and serve these models Big kudos to the Fireworks team! Kimi K2.7 Code and GLM 5.2 a

24 Jun 2026

GLM 5.2 is the open coding model everyone's been talking about. Now you can fine-tune it on Fireworks. SFT, DPO, and RL all supported. A lea…

ApplicationsDGX agent

GLM 5.2 is the open coding model everyone's been talking about. Now you can fine-tune it on Fireworks. SFT, DPO, and RL all supported. A leaderboard winner can still lose on your codebase. Training cl

It's way easier to switch models than to switch harnesses, and like many of you we use @cursor_ai every day. Now you can try out the latest …

ToolsDGX agent

It's way easier to switch models than to switch harnesses, and like many of you we use @cursor_ai every day. Now you can try out the latest open-source frontier model without changing your workflow. Y

RT @Prof_OZ: This is worth understanding. Keeping up with AI is hard, I know. It move fast and knowledge compounds. Even I use AI to build…

ToolsDGX agent

Dr. Oz discusses the rapid pace of AI development and the challenge of staying informed about advances in the field, noting that even he uses AI tools to help manage and synthesize the growing body of

The hard part of reinforcement learning on a frontier model is the infrastructure that keeps training and inference numerically identical: z…

ToolsDGX agent

The hard part of reinforcement learning on a frontier model is the infrastructure that keeps training and inference numerically identical: zero KLD, end to end. We've solved this challenge, and are no

23 Jun 2026

Bring open models to where you're already working. FireConnect brings Fireworks' top models directly into Claude Code, Pi, OpenCode, and Cod…

Model ReleasesDGX agent

Bring open models to where you're already working. FireConnect brings Fireworks' top models directly into Claude Code, Pi, OpenCode, and Codex. Watch our Head of AI Education @Prof_oz show you how: Ho

22 Jun 2026

GLM-5.2 has been the most popular new model on Fireworks this past week. @ArtificialAnlys confirms why: #3 overall on GDPval-AA (1524 Elo), …

Model ReleasesDGX agent

GLM-5.2 has been the most popular new model on Fireworks this past week. @ArtificialAnlys confirms why: #3 overall on GDPval-AA (1524 Elo), #1 open weights by 116 points. Interest is showing no signs

You don't need to be AI Dependent to build State of the Art software. @FactoryAI and Fireworks use GLM to build the most advanced agentic AI…

AgentsDGX agent

You don't need to be AI Dependent to build State of the Art software. @FactoryAI and Fireworks use GLM to build the most advanced agentic AI you can train and own today. GLM 5.2 is available in Droid,

21 Jun 2026

An hour in and first impression is definitely that GLM is really solid (very easy to set up on @FireworksAI_HQ, props to them for that, took…

Model ReleasesDGX agent

A user shares positive early impressions of GLM (likely a language model), praising its solid performance and ease of setup on Fireworks AI's platform. The post highlights Fireworks AI's developer exp

19 Jun 2026

also hearing from other customers that GLM 5.2 approaches GPT 5.5 on their evals + use cases

ToolsDGX agent

Fireworks AI reports that GLM 5.2, their language model, demonstrates performance comparable to GPT 5.5 across various evaluation benchmarks and real-world use cases. This claim suggests GLM 5.2 is co

10 Jun 2026

Coding agents break when models are 'almost' bug-free. But almost valid JSON is just not the same valid JSON. Fun piece here from @akshay_pa…

ToolsDGX agent

Coding agents break when models are 'almost' bug-free. But almost valid JSON is just not the same valid JSON. Fun piece here from @akshay_pachaar shows why SFT can't fix this, and how GRPO trains agai

6 Jun 2026

Fireworks Training Platform keeps expanding. Leading US open weight model Nemotron 3 Ultra is now ready for post-training: SFT and DPO via L…

Model ReleasesDGX agent

Fireworks Training Platform keeps expanding. Leading US open weight model Nemotron 3 Ultra is now ready for post-training: SFT and DPO via LoRA or full-parameter, on the same infrastructure that serve

4 Jun 2026

Fireworks was named to @Redpoint's InfraRed 100 which recognizes the companies building the foundation for the next wave of AI. We're just g…

AgentsDGX agent

Fireworks was named to @Redpoint's InfraRed 100 which recognizes the companies building the foundation for the next wave of AI. We're just getting started. Come build with us: https://fireworks.ai/car

Many research labs only consider inference efficiency after the fact. Step 3.7 Flash is a 198B sparse MoE VLM designed by @StepFun_ai for in…

AgentsDGX agent

Many research labs only consider inference efficiency after the fact. Step 3.7 Flash is a 198B sparse MoE VLM designed by @StepFun_ai for inference from the start. 196B language backbone with a 1.8B v

NVIDIA Nemotron 3 Ultra is on Fireworks, day zero. Nemotron Ultra is an open model for frontier reasoning and orchestration in long-running …

Model ReleasesDGX agent

NVIDIA Nemotron 3 Ultra is on Fireworks, day zero. Nemotron Ultra is an open model for frontier reasoning and orchestration in long-running autonomous agents. Think use cases like coding agents, deep

We spent the week at #MSBuild talking about one thing: fine-tuning has gone from 'maybe not worth it' to your actual competitive moat. @lqia…

ToolsDGX agent

We spent the week at #MSBuild talking about one thing: fine-tuning has gone from 'maybe not worth it' to your actual competitive moat. @lqiao sat down with @yina_arenas to break down why and what Fire

3 Jun 2026

Day 2 at #MSBuild is about what it takes to move beyond generic foundation models. Think customization, inference performance, and getting p…

ApplicationsDGX agent

Day 2 at #MSBuild is about what it takes to move beyond generic foundation models. Think customization, inference performance, and getting production-ready AI deployed at scale. @chahvivi will lead a

Fine-tuning to production inference is the gap where teams get stuck. At #MSBuild today, our own Rob Ferguson, @danielhanchen (@UnslothAI) a…

ApplicationsDGX agent

Fine-tuning to production inference is the gap where teams get stuck. At #MSBuild today, our own Rob Ferguson, @danielhanchen (@UnslothAI) and @marksaroufim (@coreautoai) discuss: model customization

Frontier models are powerful advisors. On @harvey's Legal Agent Benchmark, a GLM 5.1 worker using Claude Opus 4.7 as a sparse advisor reache…

Model ReleasesDGX agent

Frontier models are powerful advisors. On @harvey's Legal Agent Benchmark, a GLM 5.1 worker using Claude Opus 4.7 as a sparse advisor reached 18/100 all-pass versus 14/100 for Opus alone, at 39% of th

MiniMax M3 arrives with MiniMax Sparse Attention (MSA), 15.6x faster decoding at 1M tokens. We're partnering with @MiniMax_AI to power the i…

Model ReleasesDGX agent

MiniMax M3 arrives with MiniMax Sparse Attention (MSA), 15.6x faster decoding at 1M tokens. We're partnering with @MiniMax_AI to power the inference behind this week's launch. Head to http://minimax.i

Working with @FireworksAI_HQ to make MAI models easy to fine-tune and fully yours.

TutorialsDGX agent

Working with @FireworksAI_HQ to make MAI models easy to fine-tune and fully yours. Microsoft MAI models. Coming soon to Fireworks. Intelligence you control. End-to-end lineage you can prove. Fine-tune

2 Jun 2026

Microsoft MAI models. Coming soon to Fireworks. Intelligence you control. End-to-end lineage you can prove. Fine-tune MAI reasoning models f…

TutorialsDGX agent

Microsoft MAI models. Coming soon to Fireworks. Intelligence you control. End-to-end lineage you can prove. Fine-tune MAI reasoning models for your enterprise tasks. Your data. Your custom models. You

Move from test to production by running high-performance inference directly on Foundry. At #MSBuild, we demoed an end-to-end workflow showin…

ApplicationsDGX agent

Move from test to production by running high-performance inference directly on Foundry. At #MSBuild, we demoed an end-to-end workflow showing how unified infrastructure improves latency, reduces cost,

Super excited to announce seven new world-class MAI models today. They represent what we consider a new era in AI designed to keep you in co…

Model ReleasesDGX agent

Super excited to announce seven new world-class MAI models today. They represent what we consider a new era in AI designed to keep you in control and on the frontier. First is our text foundation mode

We’re looking forward to seeing how developers and enterprises use Fireworks AI on @Microsoft Foundry to power the next generation of intell…

ToolsDGX agent

We’re looking forward to seeing how developers and enterprises use Fireworks AI on @Microsoft Foundry to power the next generation of intelligent applications. Catch us at booth F111 at #MSBuild and s

1 Jun 2026

Many research labs only consider inference efficiency after the fact. Step 3.7 Flash is a 196B MoE model, and built for inference from the s…

Model ReleasesDGX agent

Many research labs only consider inference efficiency after the fact. Step 3.7 Flash is a 196B MoE model, and built for inference from the start by @StepFun_ai. Multi-Matrix Factorization Attention (M

Production AI systems place very different demands on infrastructure once workloads scale. How? Join us at #MSBuild to find out. Register he…

ApplicationsDGX agent

Production AI systems require fundamentally different infrastructure approaches as workloads scale, with demands that differ significantly from development or testing environments. Fireworks AI discus

29 May 2026

Another proof point for the open-weights thesis. From @RampLabs: 'If we built this again, we'd lean more on open-weight models.' Ramp pointe…

Model ReleasesDGX agent

Another proof point for the open-weights thesis. From @RampLabs: 'If we built this again, we'd lean more on open-weight models.' Ramp pointed 10K agents at their own backend. Kimi K2.6 and DeepSeek V4

Reliability shouldn't require reserving GPUs. Serverless 2.0 is live on Fireworks: one API, 3 serving paths. → Standard: elastic default → P…

ToolsDGX agent

Reliability shouldn't require reserving GPUs. Serverless 2.0 is live on Fireworks: one API, 3 serving paths. → Standard: elastic default → Priority: sheds last under congestion, pricing ~1.5x standard

28 May 2026

This tracks. 30 trillion tokens a day on our end, and open model share keeps climbing. Our partners @FactoryAI are seeing what we're seeing.

ToolsDGX agent

This tracks. 30 trillion tokens a day on our end, and open model share keeps climbing. Our partners @FactoryAI are seeing what we're seeing. Narrative violation: Open model use in Factory has more tha

27 May 2026

10/ The bigger point: your product is the best RL environment you'll ever have. Frontier labs ship models that are good at everything. The o…

ToolsDGX agent

10/ The bigger point: your product is the best RL environment you'll ever have. Frontier labs ship models that are good at everything. The opportunity is a model that's great at your thing. Product, u

9/ Real-time RL is where it gets fun. Catch live signals from real users on real generations. Update continuously. Ship a new version every …

ToolsDGX agent

9/ Real-time RL is where it gets fun. Catch live signals from real users on real generations. Update continuously. Ship a new version every few hours. Only works if the base model is already good enou

26 May 2026

Fireworks is coming to Tech Week, and we're doing it across multiple cities. Boston is up first. We're co-hosting a Founder Skybar Social wi…

ToolsDGX agent

Fireworks is coming to Tech Week, and we're doing it across multiple cities. Boston is up first. We're co-hosting a Founder Skybar Social with @fin_ai and @TrustVanta. Follow along via #BOSTechWeek In

21 May 2026

Cursor's new Composer 2.5 takes third on the Artificial Analysis Coding Agent Index and is ~10-60x lower cost than the higher-effort Opus 4.…

Model ReleasesDGX agent

Cursor's new Composer 2.5 takes third on the Artificial Analysis Coding Agent Index and is ~10-60x lower cost than the higher-effort Opus 4.7 and GPT-5.5 variants above it. This release puts Composer

Fine-tuning used to mean a team, a GPU cluster, and weeks of iteration. Now it's just a CLI command, ~10 min of GPU time, a few cents of com…

HardwareDGX agent

Fine-tuning used to mean a team, a GPU cluster, and weeks of iteration. Now it's just a CLI command, ~10 min of GPU time, a few cents of compute. You walk away owning the weights. Open models off the

Nathan's @cursor_ai team didn't prompt-engineer their way to Composer 2.5. They trained it. The massive RL program runs RL rollouts on Firew…

TutorialsDGX agent

Nathan's @cursor_ai team didn't prompt-engineer their way to Composer 2.5. They trained it. The massive RL program runs RL rollouts on Fireworks, alongside production inference. 'Comment 🔥 to see my p

20 May 2026

Fireworks is coming to Tech Week First up: Boston. We're co-hosting a rooftop founder social with @fin_ai and @TrustVanta — curated crowd, r…

ToolsDGX agent

Fireworks is coming to Tech Week First up: Boston. We're co-hosting a rooftop founder social with @fin_ai and @TrustVanta — curated crowd, real conversations, limited space. Thu May 28 · 7–10pm · #BOS

We ran 720 browser agent tasks with @nottecore across frontier models. One baseline model produced malformed outputs in ~1 out of every 5 ca…

AgentsDGX agent

We ran 720 browser agent tasks with @nottecore across frontier models. One baseline model produced malformed outputs in ~1 out of every 5 calls, leading to retries inside multi-step workflows. Across

18 May 2026

Not all the good stuff at Build happens on the main stage. Join Microsoft for Startups on June 2 for Dev Your Own Way. Hands-on activations …

TutorialsDGX agent

Not all the good stuff at Build happens on the main stage. Join Microsoft for Startups on June 2 for Dev Your Own Way. Hands-on activations with @GitHub, @FireworksAI_HQ and @NVIDIAforStartups. Real c

← Previous
123
Next →