stack.tools

Fireworks AI

by Fireworks AI · Inference & Serving , Fine-tuning & Training

10 companies with evidence on file — every claim below links to its source.

Cresta

Serves low-latency LLMs for real-time contact center applications.

High confidence

Receipts · 1

Show quotes (1)
  • “The low-latency, high-throughput serving of LLMs has been particularly valuable, as latency is crucial for our real-time applications.” — fireworks.ai, Dec 8, 2024

Source last verified Jul 27, 2026 Full Cresta stack →

Cursor (Anysphere)

Inference host serving Cursor's custom fine-tuned models.

High confidence

Receipts · 4

Show quotes (4)
  • “Fireworks deployed Cursor's special fine-tune of Llama-3-70b for the coding task 'Fast Apply' using the speculative API flag.” — fireworks.ai, Jun 23, 2024
  • “Fireworks provides the inference layer that makes these RL loops practical.” — fireworks.ai, Jun 26, 2026
  • “host their own custom models on Fireworks” — simonwillison.net, May 11, 2025
  • “We'd also like to thank Fireworks and Colfax for their collaboration and partnership.” — cursor.com, Mar 27, 2026

Source last verified Jul 27, 2026 Full Cursor (Anysphere) stack →

Factory

Delivers reliable inference and rapid model access for agents.

High confidence

Receipts · 1

Show quotes (1)
  • “Fireworks supports us by having these models available on basically day zero, typically well ahead of most other inference providers.” — fireworks.ai, Jun 26, 2026

Source last verified Jul 27, 2026 Full Factory stack →

Genspark

Trains large open models via reinforcement fine-tuning for Deep Research.

High confidence

Receipts · 1

Show quotes (1)
  • “By leveraging Fireworks’ Reinforcement Fine Tuning to train large state-of-the-art open models” — fireworks.ai, Oct 31, 2025

Source last verified Jul 27, 2026 Full Genspark stack →

Innovative Solutions

Primary inference layer for the DarcyIQ multi-agent platform.

High confidence

Receipts · 1

Show quotes (1)
  • “the company moved its DarcyIQ platform to Fireworks AI as its primary inference layer.” — fireworks.ai, May 5, 2026

Source last verified Jul 27, 2026 Full Innovative Solutions stack →

Notion

Hosts and serves fine-tuned and open-weight models for low-latency features.

High confidence

Receipts · 2

Show quotes (2)
  • “By fine-tuning models, we reduced latency from about 2 seconds to 350 milliseconds” — fireworks.ai, Jul 25, 2025
  • “Service provider for hosting large language models and embeddings” — registora.com, Jul 9, 2026

Source last verified Jul 26, 2026 Full Notion stack →

Sentient

Powers multi-agent chat and search with high-concurrency inference.

Medium confidence

Receipts · 1

Show quotes (1)
  • “It was running on Fireworks.” — fireworks.ai, Jul 17, 2025

Source last verified Jul 27, 2026 Full Sentient stack →

Sourcegraph

Provides scalable model inference for real-time code assistance.

High confidence

Receipts · 1

Show quotes (1)
  • ““Fireworks has been a fantastic partner in building AI dev tools at Sourcegraph.” — fireworks.ai, Jan 22, 2025

Source last verified Jul 27, 2026 Full Sourcegraph stack →

Trilogy

Primary inference layer for internal agentic workflow deployments.

High confidence

Receipts · 1

Show quotes (1)
  • “Over time, Fireworks became the primary inference layer for internal deployment testing and early production workloads.” — fireworks.ai, Jun 2026

Source last verified Jul 27, 2026 Full Trilogy stack →

Vercel

Runs v0 composite and Auto Fix models with speculative decoding.

High confidence

Receipts · 1

Show quotes (1)
  • “Both Vercel’s Auto Fix model and its v0 composite model uses Fireworks’ Speculative Decoding to speed up token generation.” — fireworks.ai, Nov 3, 2025

Source last verified Jul 27, 2026 Full Vercel stack →