Safety
one of the top cybersecurity models, post-trained from open weights by @depthfirstlabs on @FireworksAI_HQ long-horizon RL is as much an infr…
one of the top cybersecurity models, post-trained from open weights by @depthfirstlabs on @FireworksAI_HQ long-horizon RL is as much an infra problem as a research one: 100+ turn rollouts, async/pipel
one of the top cybersecurity models, post-trained from open weights by @depthfirstlabs on @FireworksAI_HQ long-horizon RL is as much an infra problem as a research one: 100+ turn rollouts, async/pipeline RL for high utilization, train-sampling numerical alignment Fireworks training platform takes care of infra, so research can move fast we're proud to support the ecosystem of cyber defenders like depthfirst as AI adoption accelerates Today we're announcing dfs-large1, our newest cybersecurity model that achieves best-in-class performance on vulnerability detection tasks. Besides frontier AI labs, only a handful of companies have built specialized models that reach the state of the art in their domain. We're p…
Related
- The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation
- Long-running models can solve hard open-ended problems, but their persistence can create safety risks that shorter-horizon evaluations miss.…
- CAST: Game Solvers as Turn-Level Teachers for LLM Agents
- SPPO: Sequence-Level PPO for Long-Horizon Reasoning Tasks
- Thinking About Thinking: Evaluating Reasoning in Post-Trained Language Models
Source: Fireworks AI (X) | 2026-07-30