Skip to main content
AgentsUse

Observability and Evaluation

Braintrust

Active observability and evaluation platform for agents, from production tracing to experiments and scoring in one system.

freemiumChecked 2026-10-07Source verified

Quick decision

Best for
Product and engineering teams that want production traces directly connected to datasets, experiments, and scored evals with CI feedback.
Not ideal for
Teams that need open source self-hosting without an enterprise agreement, or that run very high score volumes on a tight budget.
Pricing model
Freemium model with a free Starter tier including included processed data and scores with limited retention, a paid Pro tier priced per month with larger included data and score allowances and per-unit overage for additional data and scores, and custom Enterprise pricing. The pricing page meters processed data in GB and scored outputs from LLM-as-a-judge, autoevals, or custom code scorers.
Deployment
cloud. hybrid. self-hosted.
Authentication
API key via BRAINTRUST_API_KEY environment variable for SDK and CLI use
Review state
Source verified. Facts checked 2026-10-07. Not locally tested by AgentsUse.

What it does

Braintrust positions itself as an observability and evaluation platform for the whole team, from engineering to product. The homepage describes inspecting every agent trace and tool call, searching across logs, and tracking latency, cost, and quality in real time, supported by Brainstore, a database built for AI data at scale with a query engine called Nitro.

Evaluation features include defining datasets, running experiments against real data, comparing prompts and models side by side, and scoring outputs with LLMs, code, or humans. The homepage and pricing page note versioned datasets, playgrounds, experiments, custom charts and trace views, environments, and a built-in Loop agent that can run evaluations, generate test cases, and iterate on prompts.

SDK and workflow details from Braintrust documentation references include Python and TypeScript SDKs installed as the braintrust package, a bt CLI for eval and sync operations, and a GitHub Action that posts experiment comparisons on pull requests. Deployment is primarily cloud, with hybrid and self-hosted options reserved for Enterprise customers.

Verified capabilities

Observability

  • Production trace observability Source verified

    Inspect agent traces and tool calls, search across logs, and monitor latency, cost, and quality in real time.

Evaluation

  • Experiments and datasets Source verified

    Run experiments against flexible, versioned datasets and compare prompts and models side by side with git metadata.

  • Automated and human scoring Source verified

    Score outputs with LLM judges, code scorers from the AutoEvals library, custom scorers, and configurable human annotation.

Analytics

  • Discovery and investigation Source verified

    Use natural language trace investigation, pattern discovery, topic classification, and failure diagnosis from production traces.

Prompts

  • Playground and prompt iteration Source verified

    Fast prompt engineering with playground annotations and environment-tagged versions for staging and production.

Workflow

  • CI/CD eval gating Source verified

    GitHub Action integration that posts experiment comparisons on pull requests and can gate merges on score changes.

Quick start

Install the SDK and run a first eval

Install the Python SDK and scorer package from the Braintrust SDK README, define a simple Eval with data, task, and scores, then run it with an API key in the environment.

Install the SDK and run a first eval
pip install braintrust autoevals

from autoevals import LevenshteinScorer
from braintrust import Eval

Eval(
    "Say Hi Bot",
    data=lambda: [{"input": "Foo", "expected": "Hi Foo"}],
    task=lambda input: "Hi " + input,
    scores=[LevenshteinScorer],
)

Source: https://www.braintrust.dev/docs. Examples use placeholders only. Never paste a real key into a profile, config file you share, or a ticket.

MCP support: Braintrust documentation references state the vendor also publishes an MCP server to query projects, experiments, and logs, run BTQL, and summarize experiments, with a recommendation to prefer the bt CLI for most work.

Works with

Packages

Python package braintrust on PyPI and TypeScript package braintrust on npm with identical evaluation APIs, plus the bt CLI shipped with the package.

Frameworks

Framework-agnostic instrumentation for LangChain, LlamaIndex, and custom code, with provider wrappers for OpenAI, Anthropic, and others via a decorator.

API

OpenAI-compatible AI gateway listed in comparison pages, plus REST and BTQL access through SDK, CLI, and MCP.

Deployment

Managed cloud by default, with BYOC, hybrid, and self-hosted deployment only on Enterprise plans.

Only sourced support is listed. A missing framework means AgentsUse has not verified it yet, not that it cannot work.

Health and maintenance

Package
PyPI: braintrust, version 0.44.1
License
freemium
Maintainer
Braintrust

Proprietary platform with multiple SDK repositories under the braintrustdata organization, so a single product GitHub star count is omitted. PyPI version 0.44.1 from a registry release check dated 2026-10-05. Weekly downloads not read, so omitted.

Pricing and license

Freemium model with a free Starter tier including included processed data and scores with limited retention, a paid Pro tier priced per month with larger included data and score allowances and per-unit overage for additional data and scores, and custom Enterprise pricing. The pricing page meters processed data in GB and scored outputs from LLM-as-a-judge, autoevals, or custom code scorers.

Official pricing or docs →

Limitations and safety

  • Self-hosting and hybrid deployment require an Enterprise agreement, which may not fit teams with strict data residency needs on a budget.
  • Pricing meters scored outputs, so adding more scorers or scoring every production trace increases cost in addition to data processing fees.
  • The workflow is opinionated around connecting production traces to datasets and experiments, which can constrain teams with fully custom evaluation pipelines.
Safety

Browser and data tools can read pages, fill forms, and download files. Start with a test account or read-only access, keep credentials in environment variables, and review agent actions before connecting anything that can spend money, send messages, or delete data.

Alternatives to Braintrust

Common questions

What does Braintrust do for an AI agent?

Active observability and evaluation platform for agents, from production tracing to experiments and scoring in one system.

Is Braintrust open source?

This profile records the pricing model as freemium. See the pricing section for the sourced summary.

Does Braintrust support MCP?

Braintrust documentation references state the vendor also publishes an MCP server to query projects, experiments, and logs, run BTQL, and summarize experiments, with a recommendation to prefer the bt CLI for most work.

Sources and freshness

Last checked 2026-10-07. Verification label: source verified, which means public claims trace to the sources above. It does not mean AgentsUse ran the tool. Spotted an error? Send a correction. Back to the tools directory.