stack.tools

Braintrust

by Braintrust · Evals & Observability

15 companies with evidence on file — every claim below links to its source.

Box

Runs programmatic evaluations and dataset curation for its AI agent.

High confidence

Receipts · 1

Show quotes (1)
  • “Their team built an eval practice in Braintrust that helps them curate datasets” — braintrust.dev, Jul 2026

Source last verified Jul 27, 2026 Full Box stack →

Browserbase

Runs benchmarks and evaluates browser-agent model performance.

High confidence

Receipts · 1

Show quotes (1)
  • “what Browserbase uses Braintrust to observe and understand” — braintrust.dev, Jul 2026

Source last verified Jul 27, 2026 Full Browserbase stack →

Cloudflare

Evaluates the dashboard agent and gates skill changes in CI/CD.

High confidence

Receipts · 1

Show quotes (1)
  • “That was something we leaned on Braintrust for.” — braintrust.dev, Aug 2026

Source last verified Aug 22, 2026 Full Cloudflare stack →

Coursera

Evaluates AI features and monitors production quality.

High confidence

Receipts · 1

Show quotes (1)
  • “With evaluation infrastructure in place through Braintrust, Coursera maintains continuous quality awareness” — braintrust.dev, Jul 2026

Source last verified Jul 27, 2026 Full Coursera stack →

Dropbox

Builds multi-tier evaluation pipelines and detects production regressions.

High confidence

Receipts · 1

Show quotes (1)
  • “Braintrust allows us to set up that flywheel.” — braintrust.dev, Jul 2026

Source last verified Jul 27, 2026 Full Dropbox stack →

Eve

Measures agent quality against the Plaintiff Bench benchmark and production traffic.

High confidence

Receipts · 1

Show quotes (1)
  • “His job is making sure that work is good, and Braintrust is how he measures it.” — braintrust.dev, Aug 2026

Source last verified Aug 22, 2026 Full Eve stack →

Fintool

Benchmarks LLM output quality with repeatable evaluation workflows.

High confidence

Receipts · 1

Show quotes (1)
  • “Fintool leverages Braintrust’s tools to benchmark the quality of LLM outputs in real time.” — braintrust.dev, Jul 2026

Source last verified Jul 27, 2026 Full Fintool stack →

Graphite

Evaluates code-review features and compares model variants.

High confidence

Receipts · 1

Show quotes (1)
  • “run evaluations on both options using their annotated datasets in Braintrust” — braintrust.dev, Jul 2026

Source last verified Jul 27, 2026 Full Graphite stack →

Loom

Runs evals with custom scoring functions on AI output quality.

High confidence

Receipts · 1

Show quotes (1)
  • “To answer that question, they started running evals on Braintrust with their own custom scoring functions.” — braintrust.dev, Jul 2026

Source last verified Jul 27, 2026 Full Loom stack →

Navan

Runs the evaluation loop for AI voice agent quality.

High confidence

Receipts · 1

Show quotes (1)
  • “This is where Navan partnered with Braintrust to build their evaluation loop.” — braintrust.dev, Jul 2026

Source last verified Jul 27, 2026 Full Navan stack →

Notion

Eval and LLM-observability platform for regression and frontier-model testing.

High confidence

Receipts · 1

Show quotes (1)
  • “I sat down in Braintrust and looked at some of the worst experiences our customers had” — braintrust.dev, Jul 2026

Source last verified Jul 26, 2026 Full Notion stack →

Portola

Curates datasets and iterates on prompts for conversation quality.

High confidence

Receipts · 1

Show quotes (1)
  • “she creates a dataset in Braintrust tagged with the specific issue.” — braintrust.dev, Jul 2026

Source last verified Jul 27, 2026 Full Portola stack →

Pylon

Tests every prompt in CI and observes production traffic.

High confidence

Receipts · 1

Show quotes (1)
  • “requires using Braintrust to test every prompt as part of the CI pipeline.” — braintrust.dev, Aug 2026

Source last verified Aug 22, 2026 Full Pylon stack →

Retool

Evaluates AI classifier accuracy through iterative testing.

High confidence

Receipts · 1

Show quotes (1)
  • “Braintrust has been the lifeblood of our ability to execute against our roadmap” — braintrust.dev, Jul 2026

Source last verified Jul 27, 2026 Full Retool stack →

Zapier

Logs interactions, tracks feedback, and manages evaluation test sets.

High confidence

Receipts · 1

Show quotes (1)
  • “The Zapier team uses Braintrust to log user interactions, dig into their logs, track customer feedback” — braintrust.dev, Jul 2026

Source last verified Jul 27, 2026 Full Zapier stack →