stack.tools

Evals & Observability

Eval frameworks, LLM tracing/monitoring (Braintrust, LangSmith, Langfuse, Arize, W&B Weave, custom evals).

28 verified facts across 26 companies

Box

Cloud content management platform for enterprises.

Braintrust

Runs programmatic evaluations and dataset curation for its AI agent.

High confidence

Receipts · 1

Show quotes (1)
  • “Their team built an eval practice in Braintrust that helps them curate datasets” — braintrust.dev, Jul 2026

Source last verified Jul 27, 2026

Browserbase

Headless browser infrastructure for AI agents.

Braintrust

Runs benchmarks and evaluates browser-agent model performance.

High confidence

Receipts · 1

Show quotes (1)
  • “what Browserbase uses Braintrust to observe and understand” — braintrust.dev, Jul 2026

Source last verified Jul 27, 2026

C.H. Robinson

Global logistics and freight brokerage company.

LangSmith by LangChain

Provides real-time observability and error tracking during testing.

High confidence

Receipts · 1

Show quotes (1)
  • “LangSmith was their first line of defense in the testing process” — blog.langchain.dev, Mar 10, 2025

Source last verified Jul 27, 2026

Cloudflare

Internet infrastructure and developer platform.

Braintrust

Evaluates the dashboard agent and gates skill changes in CI/CD.

High confidence

Receipts · 1

Show quotes (1)
  • “That was something we leaned on Braintrust for.” — braintrust.dev, Aug 2026

Source last verified Aug 22, 2026

Coursera

Online learning platform offering courses and degrees.

Braintrust

Evaluates AI features and monitors production quality.

High confidence

Receipts · 1

Show quotes (1)
  • “With evaluation infrastructure in place through Braintrust, Coursera maintains continuous quality awareness” — braintrust.dev, Jul 2026

Source last verified Jul 27, 2026

Cursor (Anysphere)

AI code editor and agentic coding platform (Agent, Tab, Composer, Bugbot, Cloud Agents).

Cursor Bench / BugBench (internal evals) by Anysphere

Internal benchmark suites built from real usage gate model releases.

High confidence

Receipts · 2

Show quotes (2)
  • “Our benchmark, Cursor Bench, consists of real agent requests from engineers and researchers at Cursor” — cursor.com, Oct 29, 2025
  • “offline using BugBench, a curated benchmark of real code diffs” — cursor.com, Jan 15, 2026

Source last verified Jul 26, 2026

Datadog

Primary monitoring and observability platform across serving infrastructure.

Medium confidence

Receipts · 1

Show quotes (1)
  • “heavy users and find the developer experience of Datadog vastly superior to the alternatives” — newsletter.pragmaticengineer.com, Jun 10, 2025

Source last verified Jul 26, 2026

Dropbox

File storage and collaboration platform, maker of Dropbox Dash.

Braintrust

Builds multi-tier evaluation pipelines and detects production regressions.

High confidence

Receipts · 1

Show quotes (1)
  • “Braintrust allows us to set up that flywheel.” — braintrust.dev, Jul 2026

Source last verified Jul 27, 2026

Legal AI platform for plaintiff law firms.

Fintool

AI equity research copilot for institutional investors.

Braintrust

Benchmarks LLM output quality with repeatable evaluation workflows.

High confidence

Receipts · 1

Show quotes (1)
  • “Fintool leverages Braintrust’s tools to benchmark the quality of LLM outputs in real time.” — braintrust.dev, Jul 2026

Source last verified Jul 27, 2026

Graphite

AI-powered code review and stacked pull request platform.

Braintrust

Evaluates code-review features and compares model variants.

High confidence

Receipts · 1

Show quotes (1)
  • “run evaluations on both options using their annotated datasets in Braintrust” — braintrust.dev, Jul 2026

Source last verified Jul 27, 2026

Klarna

Swedish buy-now-pay-later payments and consumer banking company (NYSE: KLAR).

LangSmith by LangChain

Provides tracing, evaluations, and prompt iteration for the AI Assistant.

High confidence

Receipts · 2

Show quotes (2)
  • “With LangSmith, Klarna could pinpoint what issues arose by seeing step-by-step how their AI assistant behaved.” — langchain.com, Feb 12, 2025
  • “adapting and improving evaluation tooling for LLM applications using LangSmith” — jobsinforex.com, Jul 21, 2026

Source last verified Jul 26, 2026

Loom

Async video messaging platform, part of Atlassian.

Braintrust

Runs evals with custom scoring functions on AI output quality.

High confidence

Receipts · 1

Show quotes (1)
  • “To answer that question, they started running evals on Braintrust with their own custom scoring functions.” — braintrust.dev, Jul 2026

Source last verified Jul 27, 2026

monday.com

Work OS platform for project and service management.

LangSmith by LangChain

Drives a code-first evaluation strategy for AI agents.

High confidence

Receipts · 1

Show quotes (1)
  • “we utilized the LangSmith Vitest integration .” — blog.langchain.com, Feb 18, 2026

Source last verified Jul 27, 2026

Navan

Corporate travel and expense management platform.

Notion

Connected AI workspace for docs, wikis, projects, and enterprise search (Notion AI, Agents, Q&A).

Braintrust

Eval and LLM-observability platform for regression and frontier-model testing.

High confidence

Receipts · 1

Show quotes (1)
  • “I sat down in Braintrust and looked at some of the worst experiences our customers had” — braintrust.dev, Jul 2026

Source last verified Jul 26, 2026

Perplexity

AI answer engine combining a proprietary web index with LLMs to deliver cited, conversational search.

In-house evals (LLM-as-judge) by Perplexity

Grades search quality with LLM-as-judge evaluations over public benchmarks.

High confidence

Receipts · 2

Show quotes (2)
  • “We grade all benchmarks using the same prompted classifier methodology used in the original work” — research.perplexity.ai, Jul 17, 2026
  • “you will build specialized evals to improve answer quality across Perplexity” — jobs.ashbyhq.com, Jun 29, 2026

Source last verified Jul 26, 2026

Podium

AI-powered lead conversion and communication platform for local businesses.

LangSmith by LangChain

Tests and monitors AI employee performance with traces and datasets.

High confidence

Receipts · 1

Show quotes (1)
  • “turned to LangSmith for LLM testing and observability.” — blog.langchain.dev, Aug 15, 2024

Source last verified Jul 27, 2026

Portola

Maker of Tolan, an AI alien companion app.

Braintrust

Curates datasets and iterates on prompts for conversation quality.

High confidence

Receipts · 1

Show quotes (1)
  • “she creates a dataset in Braintrust tagged with the specific issue.” — braintrust.dev, Jul 2026

Source last verified Jul 27, 2026

Pylon

Agentic B2B customer support platform.

Braintrust

Tests every prompt in CI and observes production traffic.

High confidence

Receipts · 1

Show quotes (1)
  • “requires using Braintrust to test every prompt as part of the CI pipeline.” — braintrust.dev, Aug 2026

Source last verified Aug 22, 2026

Loop by Braintrust

Spins up new LLM-as-a-judge scorers for bulk quality checks.

High confidence

Receipts · 1

Show quotes (1)
  • “uses Loop to spin up new scorers quickly.” — braintrust.dev, Aug 2026

Source last verified Aug 22, 2026

Rakuten Group

Japanese internet conglomerate spanning e-commerce, fintech, and digital content.

LangSmith by LangChain

Monitors agent performance and distributes prompts across teams.

High confidence

Receipts · 1

Show quotes (1)
  • “LangSmith allows us to get things done scientifically.” — blog.langchain.dev, Feb 14, 2024

Source last verified Jul 27, 2026

Retool

Enterprise platform for building internal tools and apps.

Braintrust

Evaluates AI classifier accuracy through iterative testing.

High confidence

Receipts · 1

Show quotes (1)
  • “Braintrust has been the lifeblood of our ability to execute against our roadmap” — braintrust.dev, Jul 2026

Source last verified Jul 27, 2026

ServiceNow

Enterprise platform for IT, employee, and customer workflows.

LangSmith by LangChain

Runs a tailored evaluation framework over its multi-agent system.

High confidence

Receipts · 1

Show quotes (1)
  • “ServiceNow implemented a sophisticated evaluation framework in LangSmith tailored to their multi-agent system.” — blog.langchain.com, Nov 17, 2025

Source last verified Jul 27, 2026

Shopify

E-commerce platform powering millions of merchant storefronts and an AI merchant assistant (Sidekick).

In-house LLM-judge eval platform (GTX + merchant simulator) by Shopify (in-house)

Evaluates Sidekick with calibrated LLM judges and a merchant simulator.

High confidence

Receipts · 2

Show quotes (2)
  • “we built an LLM-powered merchant simulator that captures the 'essence' or goals of real conversations” — shopify.engineering, Aug 26, 2025
  • “We moved away from carefully curated "golden" datasets toward Ground Truth Sets (GTX)” — shopify.engineering, Aug 26, 2025

Source last verified Jul 26, 2026

Trellix

Cybersecurity company focused on extended detection and response.

LangSmith by LangChain

Monitors agent performance and debugs workflows through experiments and traces.

High confidence

Receipts · 1

Show quotes (1)
  • “Trellix used LangSmith for experimentation and to action upon performance metrics.” — blog.langchain.dev, Apr 21, 2025

Source last verified Jul 27, 2026

Vodafone

Multinational telecommunications company.

LangSmith by LangChain

Tracks LLM application lifecycle, debugging, and evaluation.

High confidence

Receipts · 1

Show quotes (1)
  • “With LangChain, LangGraph and LangSmith, Vodafone has successfully delivered advanced AI-driven solutions” — blog.langchain.dev, Mar 23, 2025

Source last verified Jul 27, 2026

Zapier

Workflow automation platform connecting thousands of apps.

Braintrust

Logs interactions, tracks feedback, and manages evaluation test sets.

High confidence

Receipts · 1

Show quotes (1)
  • “The Zapier team uses Braintrust to log user interactions, dig into their logs, track customer feedback” — braintrust.dev, Jul 2026

Source last verified Jul 27, 2026