Category
Agent Observability and Evaluation
Tracing, logging, and evaluation platforms that show what an agent did, what it cost, and whether the output was any good.
How to read this category
You cannot fix an agent you cannot see. These four platforms record traces, tokens, costs, and outputs so failures can be reviewed like bugs instead of argued about like weather.
Langfuse is the open-source option with a self-hosted path and strength in prompt and cost views. LangSmith is the native layer for LangChain and LangGraph apps, with dataset-driven evaluation. Helicone sits close to the API call as a logging and gateway layer. Braintrust pushes evaluation and experiment workflows for teams iterating on quality.
Pick by where your pain is. If you cannot see what happened, start anywhere. If you can see it but cannot tell whether a change helped, the evaluation workflow is the deciding feature.
The verified profiles
Side by side on sourced facts
| Tool | Pricing model | License | Deployment | MCP support | Checked |
|---|---|---|---|---|---|
| Langfuse | Freemium (hosted paid tiers) | MIT | cloud, self-hosted, docker, kubernetes | MCP for coding agents | 2026-10-07 |
| LangSmith | Freemium (hosted paid tiers) | Not open source | cloud, bring-your-own-cloud, self-hosted | LangSmith MCP for prompts | 2026-10-07 |
| Helicone | Freemium (hosted paid tiers) | Apache-2.0 | cloud, self-hosted, docker | See profile | 2026-10-07 |
| Braintrust | Freemium (hosted paid tiers) | Not open source | cloud, hybrid, self-hosted | See profile | 2026-10-07 |
Every cell traces to the tool profile, which traces to the official repository, package page, and documentation checked on the date shown. Pricing cells name the model, not a price: vendor prices change, so check the profile for the sourced pricing summary and the official pricing link.
What actually decides the choice
- Self-hosting rules narrow the field fast; only some of these run on your own infrastructure.
- Framework gravity matters: a LangChain stack gets more from LangSmith native tracing, mixed stacks get more from a neutral store.
- Evaluation depth (datasets, scorers, experiments) decides once basic tracing is table stakes.
- Cost views are not vanity metrics. Per-run and per-user spend is how you catch the loop that burns a weekend budget.
How this archive is ordered, and why it is short
Profiles appear in the order they passed the AgentsUse review gate, not in a tested quality ranking. AgentsUse has not run head-to-head benchmarks for this category, so this page does not claim a winner. The comparison table is the ranking tool: sort by the dimension your job cannot compromise on, then read the full profiles for limitations.
This archive currently lists 4 verified profiles. AgentsUse normally asks for eight published profiles before opening a category archive. This one opens early, on the record, because every listed profile passed the full tool gate and readers comparing real choices need the side-by-side surface now, not a longer unvetted list later. The archive grows as more tools pass review. See how verification works.
Read before you give these tools real work
Guide
How to Evaluate Agent Output Quality
It seems to work is not an evaluation strategy. How to define quality concretely, build runnable evals, and track agent performance over time.
2026-09-29
Guide
How Much Does It Cost to Run an AI Agent? The Full Picture
Token pricing, tool calls, and the hidden costs of running AI agents - explained without jargon, with ways to keep bills down.
2026-09-26
Back to the tools directory, the category index, or the guides.