stack.tools

Fine-tuning & Training

Training/fine-tuning stack — frameworks, compute, RLHF tooling, model post-training.

16 verified facts across 12 companies

Applied Compute

Startup building reinforcement-learning-trained enterprise AI models.

Modal

Executes reinforcement learning rollouts, grading, and inference workloads.

High confidence

Receipts · 1

Show quotes (1)
  • “Applied Compute makes use of Modal Functions to provide inexpensive serverless fan-out without requiring a dedicated cluster.” — modal.com, May 20, 2026

Source last verified Jul 27, 2026

Cartesia

Real-time voice AI company building state-space models.

Together GPU Clusters by Together AI

Trains voice models on multi-node clusters scheduled via Slurm.

High confidence

Receipts · 1

Show quotes (1)
  • “Cartesia also uses Together’s H100 80GB clusters for model training, with multi-node workloads scheduled via Slurm.” — together.ai, Jul 2026

Source last verified Jul 27, 2026

Cursor (Anysphere)

AI code editor and agentic coding platform (Agent, Tab, Composer, Bugbot, Cloud Agents).

Ray by Anyscale (open source)

Underpins the custom asynchronous reinforcement-learning training infrastructure with PyTorch.

High confidence

Receipts · 2

Show quotes (2)
  • “We built custom training infrastructure leveraging PyTorch and Ray to power asynchronous reinforcement learning at scale.” — cursor.com, Oct 29, 2025
  • “Build our distributed training, inference, and RL infrastructure” — cursor.com, Jul 2026

Source last verified Jul 26, 2026

NVIDIA GPUs (Blackwell) by NVIDIA

Trains mixture-of-experts models at low precision on large GPU clusters.

High confidence

Receipts · 2

Show quotes (2)
  • “allowing us to scale training to thousands of NVIDIA GPUs with minimal communication cost” — cursor.com, Oct 29, 2025
  • “Custom low-precision kernels for efficient MoE training on Blackwell GPUs” — cursor.com, Mar 27, 2026

Source last verified Jul 26, 2026

Anyrun by Anysphere

Internal compute platform running sandboxed cloud environments for RL rollouts.

High confidence

Receipts · 2

Show quotes (2)
  • “Anyrun, our internal compute platform for running hundreds of thousands of sandboxed coding environments” — cursor.com, Mar 27, 2026
  • “running hundreds of thousands of concurrent sandboxed coding environments in the cloud” — cursor.com, Oct 29, 2025

Source last verified Jul 26, 2026

Online RL pipeline (Tab model) by Anysphere

Trains the Tab model continuously on live accept and reject feedback.

High confidence

Receipts · 2

Show quotes (2)
  • “rolling out new models to users frequently throughout the day and using that data for training” — cursor.com, Sep 12, 2025
  • “train frontier coding agents and scale RL on real user data” — cursor.com, Jul 2026

Source last verified Jul 26, 2026

Deep Cogito

AI lab training open-weight hybrid reasoning models.

Together GPU Clusters by Together AI

Trains large open-weight models on multi-node GPU clusters.

High confidence

Receipts · 1

Show quotes (1)
  • “Deep Cogito uses Together for customizable H100/H200 GPU clusters, reliable long-run training infrastructure” — together.ai, Jul 2026

Source last verified Jul 27, 2026

Genspark

AI agent company building autonomous research and productivity agents.

Fireworks AI

Trains large open models via reinforcement fine-tuning for Deep Research.

High confidence

Receipts · 1

Show quotes (1)
  • “By leveraging Fireworks’ Reinforcement Fine Tuning to train large state-of-the-art open models” — fireworks.ai, Oct 31, 2025

Source last verified Jul 27, 2026

Latent Health

Clinical AI company for pharmacy intelligence.

Together GPU Clusters by Together AI

Trains clinical models with multi-node RL and long-context runs.

High confidence

Receipts · 1

Show quotes (1)
  • “Latent Health chose Together Instant Clusters for clinical AI training” — together.ai, Jul 2026

Source last verified Jul 27, 2026

Notion

Connected AI workspace for docs, wikis, projects, and enterprise search (Notion AI, Agents, Q&A).

Custom fine-tuned models by Notion

Fine-tunes small open-weight models for search, routing, and function calling.

High confidence

Receipts · 2

Show quotes (2)
  • “By fine-tuning models, we reduced latency from about 2 seconds to 350 milliseconds” — fireworks.ai, Jul 25, 2025
  • “Our, um, fine tuned and open source models are served on GPUs, right?” — latent.space, Apr 15, 2026

Source last verified Jul 26, 2026

Perplexity

AI answer engine combining a proprietary web index with LLMs to deliver cited, conversational search.

Amazon SageMaker HyperPod by AWS

Early distributed training platform, since replaced by a self-managed GPU fleet.

Historical Medium confidence

Receipts · 2

Show quotes (2)
  • “Amazon SageMaker HyperPod's built-in data and model parallel libraries helped us optimize training time on GPUs” — aws.amazon.com, Mar 2024
  • “Perplexity Accelerates Foundation Model Training by 40% with Amazon SageMaker HyperPod” — youtube.com, Mar 27, 2024

Source last verified Jul 26, 2026

NeMo by NVIDIA

Framework behind the post-training run that produced R1-1776.

Medium confidence

Receipts · 1

Show quotes (1)
  • “They used NVIDIA's NeMo 2.0 framework to fine-tune the model” — hyperight.com, Feb 24, 2025

Source last verified Jul 26, 2026

Scaled Cognition

AI lab training agentic models for customer-facing tasks.

Together GPU Clusters by Together AI

Provides bare-metal GPU infrastructure for custom model training.

High confidence

Receipts · 1

Show quotes (1)
  • “How Scaled Cognition Trains APT-1 on Together AI GPU Clusters” — together.ai, Jul 2026

Source last verified Jul 27, 2026

Shopify

E-commerce platform powering millions of merchant storefronts and an AI merchant assistant (Sidekick).

GRPO reinforcement-learning fine-tuning by Shopify (in-house)

Post-trains Sidekick's fine-tuned models with LLM judges as reward signals.

High confidence

Receipts · 2

Show quotes (2)
  • “we implemented Group Relative Policy Optimization (GRPO), a reinforcement learning approach that uses our LLM judges as reward signals” — shopify.engineering, Aug 26, 2025
  • “Fine-tuning an open-source model got us there for the common cases.” — shopify.engineering, Jun 15, 2026

Source last verified Jul 26, 2026

Slingshot AI

AI lab building foundation models for psychology.

Together AI

Runs supervised fine-tuning and preference optimization pipelines.

Medium confidence

Receipts · 1

Show quotes (1)
  • “using Together AI for key components” — together.ai, Jul 2026

Source last verified Jul 27, 2026

XY.AI Labs

Agentic AI for healthcare administration and revenue-cycle workflows.

Together Fine-Tuning Platform by Together AI

Trains customer-specific Qwen models for EOB parsing.

High confidence

Receipts · 1

Show quotes (1)
  • “XY.AI migrated from its self‑hosted stack to the Together Fine‑Tuning Platform” — together.ai, Jul 2026

Source last verified Jul 27, 2026