APIs, MCP servers and models to evaluate LLM output
3 tools in the catalog do this, by name.
- BraintrustPlatformPlatform for evaluating, tracing and monitoring AI apps and agents
- LangfusePlatformOpen-source tracing, evals and prompt management for LLM apps
- PromptfooPackageLocal LLM evaluations and red-team tests with an official MCP server