Braintrust

Platform

Platform for evaluating, tracing and monitoring AI apps and agents

Price
Free tier, then paid
Access
Braintrust API key

About

Hosted platform for running evals on datasets, scoring outputs and tracing production AI requests. SDKs cover Python, TypeScript, Go, Java, Ruby and C#. Self-hosting the data plane or running it in your own cloud requires the Enterprise plan.

What you can do with it

  • Run evals that score model outputs against a dataset in code or CI
  • Trace production agent requests and turn logs into test datasets
  • Compare prompts and models side by side in a playground

Get started

  1. Create a Braintrust API key and set BRAINTRUST_API_KEY
  2. Set OPENAI_API_KEY for the model under test

Example

# pip install braintrust autoevals openai; run with: bt eval movie_matcher.py
from braintrust import Eval
from autoevals import ExactMatch
from openai import OpenAI

client = OpenAI()

def task(input):
    prompt = "Based on the following description, identify the movie.\n" + input
    return client.responses.create(model="gpt-5-mini", input=prompt).output_text

Eval("Evaluation quickstart", data=[{"input": "A detective investigates a series of murders based on the seven deadly sins.", "expected": "Se7en"}], task=task, scores=[ExactMatch])

Details

Hosting
Hosted service, Self-hosted
Available in
Worldwide
Official SDKs
Python, JavaScript/TypeScript, Go, Java, Ruby, C#
MCP server
Remote

Tasks

Alternatives

Other tools for the same tasks.

Last checked on .