Companies with Braintrust in the stack
All Braintrust facts
-
Box · Braintrust
Runs programmatic evaluations and dataset curation for its AI agent.
High confidenceSource last verified Jul 27, 2026 View fact with receipts →
-
Browserbase · Braintrust
Runs benchmarks and evaluates browser-agent model performance.
High confidenceSource last verified Jul 27, 2026 View fact with receipts →
-
Cloudflare · Braintrust
Evaluates the dashboard agent and gates skill changes in CI/CD.
High confidenceSource last verified Aug 22, 2026 View fact with receipts →
-
Coursera · Braintrust
Evaluates AI features and monitors production quality.
High confidenceSource last verified Jul 27, 2026 View fact with receipts →
-
Dropbox · Braintrust
Builds multi-tier evaluation pipelines and detects production regressions.
High confidenceSource last verified Jul 27, 2026 View fact with receipts →
-
Eve · Braintrust
Measures agent quality against the Plaintiff Bench benchmark and production traffic.
High confidenceSource last verified Aug 22, 2026 View fact with receipts →
-
Fintool · Braintrust
Benchmarks LLM output quality with repeatable evaluation workflows.
High confidenceSource last verified Jul 27, 2026 View fact with receipts →
-
Graphite · Braintrust
Evaluates code-review features and compares model variants.
High confidenceSource last verified Jul 27, 2026 View fact with receipts →
-
Loom · Braintrust
Runs evals with custom scoring functions on AI output quality.
High confidenceSource last verified Jul 27, 2026 View fact with receipts →
-
Navan · Braintrust
Runs the evaluation loop for AI voice agent quality.
High confidenceSource last verified Jul 27, 2026 View fact with receipts →
-
Notion · Braintrust
Eval and LLM-observability platform for regression and frontier-model testing.
High confidenceSource last verified Jul 26, 2026 View fact with receipts →
-
Portola · Braintrust
Curates datasets and iterates on prompts for conversation quality.
High confidenceSource last verified Jul 27, 2026 View fact with receipts →
-
Pylon · Braintrust
Tests every prompt in CI and observes production traffic.
High confidenceSource last verified Aug 22, 2026 View fact with receipts →
-
Pylon · Loop
Spins up new LLM-as-a-judge scorers for bulk quality checks.
High confidenceSource last verified Aug 22, 2026 View fact with receipts →
-
Retool · Braintrust
Evaluates AI classifier accuracy through iterative testing.
High confidenceSource last verified Jul 27, 2026 View fact with receipts →
-
Zapier · Braintrust
Logs interactions, tracks feedback, and manages evaluation test sets.
High confidenceSource last verified Jul 27, 2026 View fact with receipts →