stack.tools

Inference & Serving

Where models run — clouds (Bedrock, Vertex), GPU providers, inference platforms (Together, Fireworks, Baseten, Modal), self-hosted serving (vLLM).

36 verified facts across 27 companies

Cartesia

Real-time voice AI company building state-space models.

Together AI

Runs a custom SSM inference engine for real-time voice.

High confidence

Receipts · 1

Show quotes (1)
  • “To support these requirements — and serve millions of audio minutes daily — Cartesia uses Together AI.” — together.ai, Jul 2026

Source last verified Jul 27, 2026

Chai Discovery

AI company for molecular structure prediction and drug discovery.

Modal

Runs computational biology workflows from research through production.

High confidence

Receipts · 2

Show quotes (2)
  • “With Modal, building infrastructure has shifted from being an imperative task to a declarative one.” — modal.com, Jan 15, 2026
  • “With Modal Volumes it’s downloaded once, instantly available everywhere, and scales to thousands of queries.” — modal.com, Jan 15, 2026

Source last verified Jul 27, 2026

Cresta

AI platform for contact center agent assistance and insights.

Fireworks AI

Serves low-latency LLMs for real-time contact center applications.

High confidence

Receipts · 1

Show quotes (1)
  • “The low-latency, high-throughput serving of LLMs has been particularly valuable, as latency is crucial for our real-time applications.” — fireworks.ai, Dec 8, 2024

Source last verified Jul 27, 2026

Cursor (Anysphere)

AI code editor and agentic coding platform (Agent, Tab, Composer, Bugbot, Cloud Agents).

Fireworks AI

Inference host serving Cursor's custom fine-tuned models.

High confidence

Receipts · 4

Show quotes (4)
  • “Fireworks deployed Cursor's special fine-tune of Llama-3-70b for the coding task 'Fast Apply' using the speculative API flag.” — fireworks.ai, Jun 23, 2024
  • “Fireworks provides the inference layer that makes these RL loops practical.” — fireworks.ai, Jun 26, 2026
  • “host their own custom models on Fireworks” — simonwillison.net, May 11, 2025
  • “We'd also like to thank Fireworks and Colfax for their collaboration and partnership.” — cursor.com, Mar 27, 2026

Source last verified Jul 27, 2026

AWS by Amazon Web Services

Primary cloud for the backend that routes AI traffic to model providers.

High confidence

Receipts · 2

Show quotes (2)
  • “We are very much a 'cloud shop.' We mostly rely on AWS and then Azure for inference.” — newsletter.pragmaticengineer.com, Jun 10, 2025
  • “AWS for primary infrastructure, Azure and GCP for "some secondary infrastructure"” — simonwillison.net, May 11, 2025

Source last verified Jul 26, 2026

Together GPU Clusters by Together AI

Runs production inference on dedicated NVIDIA Blackwell GPU clusters.

High confidence

Receipts · 1

Show quotes (1)
  • “Cursor partnered with Together AI to deploy production inference on NVIDIA Blackwell” — together.ai, Jul 2026

Source last verified Jul 27, 2026

Decagon

AI customer service agents for enterprises.

Modal

Trains and serves real-time voice AI models at scale.

High confidence

Receipts · 1

Show quotes (1)
  • “Modal’s infrastructure powered this progress, enabling Decagon to train and serve increasingly capable models” — modal.com, Nov 13, 2025

Source last verified Jul 27, 2026

Together AI

Runs low-latency production inference for the voice stack.

High confidence

Receipts · 1

Show quotes (1)
  • “Decagon uses Together’s inference engine as the execution layer” — together.ai, Jul 2026

Source last verified Jul 27, 2026

Deep Cogito

AI lab training open-weight hybrid reasoning models.

Together AI

Hosts the full Cogito model lineup on dedicated infrastructure.

High confidence

Receipts · 1

Show quotes (1)
  • “Together hosts the full Cogito model lineup, from 3B to 671B parameters, on dedicated inference infrastructure.” — together.ai, Jul 2026

Source last verified Jul 27, 2026

Factory

AI software engineering agents (Droids) for enterprise development.

Fireworks AI

Delivers reliable inference and rapid model access for agents.

High confidence

Receipts · 1

Show quotes (1)
  • “Fireworks supports us by having these models available on basically day zero, typically well ahead of most other inference providers.” — fireworks.ai, Jun 26, 2026

Source last verified Jul 27, 2026

Innovative Solutions

AWS premier partner delivering cloud and AI services.

Fireworks AI

Primary inference layer for the DarcyIQ multi-agent platform.

High confidence

Receipts · 1

Show quotes (1)
  • “the company moved its DarcyIQ platform to Fireworks AI as its primary inference layer.” — fireworks.ai, May 5, 2026

Source last verified Jul 27, 2026

Klarna

Swedish buy-now-pay-later payments and consumer banking company (NYSE: KLAR).

Klarna AI Gateway by Klarna (in-house)

Provides centralized low-latency model access for GenAI use cases company-wide.

High confidence

Receipts · 1

Show quotes (1)
  • “manage and operate Klarna's AI Gateway, ensuring reliable, low-latency access to AI models at scale” — jobsinforex.com, Jul 21, 2026

Source last verified Jul 26, 2026

Notion

Connected AI workspace for docs, wikis, projects, and enterprise search (Notion AI, Agents, Q&A).

Fireworks AI

Hosts and serves fine-tuned and open-weight models for low-latency features.

High confidence

Receipts · 2

Show quotes (2)
  • “By fine-tuning models, we reduced latency from about 2 seconds to 350 milliseconds” — fireworks.ai, Jul 25, 2025
  • “Service provider for hosting large language models and embeddings” — registora.com, Jul 9, 2026

Source last verified Jul 26, 2026

Anyscale (Ray) by Anyscale

Runs the near-real-time embeddings indexing pipeline on managed Ray.

High confidence

Receipts · 3

Show quotes (3)
  • “we set out to migrating our near real-time embeddings pipeline to Ray running on Anyscale” — notion.com, Feb 19, 2026
  • “Ray lets us run open-source embedding models directly, without being gated by external providers” — notion.com, Feb 19, 2026
  • “pull a model from Hugging Face and run it ourselves on-prem” — anyscale.com, Jul 2026

Source last verified Jul 26, 2026

Perplexity

AI answer engine combining a proprietary web index with LLMs to deliver cited, conversational search.

In-house inference engine by Perplexity

Custom Rust and CUDA engine serving every query across a multi-cloud GPU fleet.

High confidence

Receipts · 2

Show quotes (2)
  • “We build and run the inference engine behind every Perplexity query and deploy dozens of model architectures at scale” — jobs.ashbyhq.com, Apr 13, 2026
  • “a large GPU fleet spread across several cloud providers” — jobs.ashbyhq.com, Jul 16, 2026

Source last verified Jul 26, 2026

TensorRT-LLM by NVIDIA

Served LLMs on GPU pods before the in-house engine took over.

Historical High confidence

Receipts · 2

Show quotes (2)
  • “Triton Inference Server is a critical component of Perplexity's deployment architecture.” — developer.nvidia.com, Dec 5, 2024
  • “Our stack is Rust, Python, CUDA, and CuTe DSL” — jobs.ashbyhq.com, Apr 13, 2026

Source last verified Jul 26, 2026

Cerebras Inference by Cerebras

Wafer-scale inference serving the Sonar model for near-instant answers.

High confidence

Receipts · 1

Show quotes (1)
  • “1,200 tokens per second, delivering near-instant answer generation” — cerebras.ai, Feb 11, 2025

Source last verified Jul 26, 2026

TransferEngine (pplx-garden) by Perplexity

Open-sourced RDMA library powering disaggregated serving of large MoE models.

High confidence

Receipts · 2

Show quotes (2)
  • “KvCache transfer for disaggregated inference with dynamic scaling” — arxiv.org, Oct 31, 2025
  • “Perplexity open source garden for inference technology” — github.com, Nov 4, 2025

Source last verified Jul 26, 2026

Physical Intelligence

Robotics foundation model company.

Modal

Runs real-time remote inference for robotic control.

High confidence

Receipts · 1

Show quotes (1)
  • “On Modal, PI can allocate larger, data-center-class GPUs per deployment and run GPU-intensive experiments immediately.” — modal.com, Apr 8, 2026

Source last verified Jul 27, 2026

Ramp

Finance automation platform for corporate cards, expenses, and payments.

Modal

Powers the Ramp Inspect background coding agent.

High confidence

Receipts · 2

Show quotes (2)
  • “Ramp uses Modal to power Ramp Inspect” — modal.com, Feb 19, 2026
  • “Leveraging Modal Sandboxes, Ramp spins up full development environments in seconds,” — modal.com, Feb 19, 2026

Source last verified Jul 27, 2026

Reducto

Document ingestion and parsing APIs for LLM pipelines.

Modal

Scales GPU inference for multi-model document processing.

High confidence

Receipts · 1

Show quotes (1)
  • “Reducto continues to expand its use of Modal with the deployment of new large language and vision-language models.” — modal.com, Nov 19, 2025

Source last verified Jul 27, 2026

Runware

Generative media inference API platform.

Together GPU Clusters by Together AI

Provides on-demand GPU capacity for rapid model deployment.

High confidence

Receipts · 1

Show quotes (1)
  • “Together provides immediate access to NVIDIA H100s, H200s, and B200s without long-term commitments.” — together.ai, Jul 2026

Source last verified Jul 27, 2026

Sentient

Open-source AGI research organization building decentralized AI.

Fireworks AI

Powers multi-agent chat and search with high-concurrency inference.

Medium confidence

Receipts · 1

Show quotes (1)
  • “It was running on Fireworks.” — fireworks.ai, Jul 17, 2025

Source last verified Jul 27, 2026

Shopify

E-commerce platform powering millions of merchant storefronts and an AI merchant assistant (Sidekick).

Vertex AI by Google Cloud

Serves Claude models for Sidekick at merchant scale.

High confidence

Receipts · 1

Show quotes (1)
  • “The combination of Claude and Vertex AI helps us empower millions of merchants with our AI-enabled commerce assistant, Sidekick” — cloud.google.com, Feb 5, 2026

Source last verified Jul 26, 2026

Triton Inference Server by NVIDIA

Self-hosts fine-tuned vision models across the GPU fleet.

High confidence

Receipts · 2

Show quotes (2)
  • “Orchestrates model serving across our GPU fleet, handling request preprocessing, batching, and routing” — shopify.engineering, Jul 16, 2025
  • “40 million LLM calls daily, representing about 16 billion tokens inferred per day” — shopify.engineering, Jul 16, 2025

Source last verified Jul 26, 2026

GroqCloud by Groq

AI services provider whose exact workload is not publicly detailed.

Medium confidence

Receipts · 1

Show quotes (1)
  • “Artificial intelligence services” — help.shopify.com, Jul 2026

Source last verified Jul 26, 2026

Sourcegraph

Code intelligence platform with AI coding assistants.

Fireworks AI

Provides scalable model inference for real-time code assistance.

High confidence

Receipts · 1

Show quotes (1)
  • ““Fireworks has been a fantastic partner in building AI dev tools at Sourcegraph.” — fireworks.ai, Jan 22, 2025

Source last verified Jul 27, 2026

Stably

AI-powered QA testing agents for web applications.

Vercel AI Gateway by Vercel

Provides scalable model access with high throughput limits.

High confidence

Receipts · 1

Show quotes (1)
  • “leveraging AI Gateway for AI scalability and large TPM limits” — vercel.com, Feb 17, 2026

Source last verified Jul 27, 2026

Substack

Publishing platform for newsletters and podcasts.

Modal

Runs ML training and deployment, replacing AWS SageMaker.

High confidence

Receipts · 1

Show quotes (1)
  • “For nearly all these models Substack has moved both training and deployment from AWS SageMaker to Modal.” — modal.com, May 20, 2024

Source last verified Jul 27, 2026

Suno

AI music generation platform.

Modal

Scales inference and batch pre-processing to thousands of GPUs.

High confidence

Receipts · 1

Show quotes (1)
  • “Suno uses Modal to scale inference and batch pre-processing to thousands of GPUs.” — modal.com, Feb 21, 2024

Source last verified Jul 27, 2026

The Washington Post

US national news publisher.

Together AI

Serves open models behind a public AI journalism platform.

High confidence

Receipts · 1

Show quotes (1)
  • “They deployed open models on Together AI with dedicated endpoints and hybrid serverless capacity, maintaining full model control and predictable pricing.” — together.ai, Jul 2026

Source last verified Jul 27, 2026

Trilogy

Enterprise software operator running a portfolio of business software products.

Fireworks AI

Primary inference layer for internal agentic workflow deployments.

High confidence

Receipts · 1

Show quotes (1)
  • “Over time, Fireworks became the primary inference layer for internal deployment testing and early production workloads.” — fireworks.ai, Jun 2026

Source last verified Jul 27, 2026

Vercel

Frontend cloud platform behind Next.js, v0, and the AI SDK.

Fireworks AI

Runs v0 composite and Auto Fix models with speculative decoding.

High confidence

Receipts · 1

Show quotes (1)
  • “Both Vercel’s Auto Fix model and its v0 composite model uses Fireworks’ Speculative Decoding to speed up token generation.” — fireworks.ai, Nov 3, 2025

Source last verified Jul 27, 2026

Vercept

AI startup building vision-driven computer-use agents.

Together AI

Deploys custom computer vision models via load-balanced inference.

High confidence

Receipts · 1

Show quotes (1)
  • “Together provides auto-scaling that monitors response latency and queue depth, enabling instant scaling without pre-provisioning.” — together.ai, Jul 2026

Source last verified Jul 27, 2026

XY.AI Labs

Agentic AI for healthcare administration and revenue-cycle workflows.

Together AI

Serves fine-tuned models as endpoints for structured extraction.

High confidence

Receipts · 1

Show quotes (1)
  • “Together AI turned fine-tuning, evaluations, and deployment into a repeatable loop from experimentation to generating a testable endpoint” — together.ai, Jul 2026

Source last verified Jul 27, 2026

Yutori

AI company building autonomous web-browsing agents.

Together AI

Serves browser-use agents behind Scouts, Delegate, and Navigator.

High confidence

Receipts · 1

Show quotes (1)
  • “Yutori runs on Together AI, the AI Native Cloud” — together.ai, Jul 2026

Source last verified Jul 27, 2026