Every team that runs an LLM application in production ends up needing the same things: a trace of what happened on each request, the prompts and tool calls inside it, the tokens and cost of every step, and a way to tell whether the answers were any good.
You can buy that from a closed platform. Increasingly, you do not need to. Some of the best LLM observability tools are open source, built on OpenTelemetry, and developed in public by large communities. They are often cheaper, they let you keep sensitive prompt data in your own environment, and you can see exactly how they work.
In an earlier post I argued that efficiency also means choosing the right tools, not only rightsizing the ones you have. LLM observability is where that argument is strongest, because the data is sensitive, the volumes grow fast, and the standards are still forming.
What LLM observability needs to do
- Tracing: every request as a tree of steps, including model calls, retrieval, and tool calls, with inputs, outputs, and latency.
- Cost and token tracking: tokens, cached tokens, and cost per call, per user, per feature, and per task.
- Evaluation: scoring outputs with LLM-as-a-judge, rules, or human review, on live traffic and on test datasets.
- Prompt management: versioning prompts and comparing them in experiments.
- Data control: deciding where prompts and responses are stored, and for how long.
The open-source tools below cover different parts of that list. Several cover all of it.
The tools
Full platforms
Langfuse is one of the most widely adopted open-source LLM engineering platforms, with more than 35,000 GitHub stars as of October 2026. It covers tracing, evaluations, prompt management, datasets, and cost tracking. The core is MIT licensed, with enterprise features in separate directories, and the self-hosted version is free. ClickHouse acquired Langfuse in January 2026 and committed to keeping it open source under its existing license. Langfuse accepts OpenTelemetry data and builds its SDK on the OpenTelemetry client.
Opik by Comet is an Apache 2.0 platform for tracing, evaluation, LLM-as-a-judge, prompt experiments, and guardrails. Comet states that the full platform, backend included, is free to self-host. It accepts OpenTelemetry traces and has more than 22,000 GitHub stars.
Arize Phoenix offers tracing, evaluation, prompt experiments, and datasets, built on OpenTelemetry and the OpenInference instrumentation conventions. One detail matters: Phoenix is licensed under the Elastic License 2.0, which is source-available rather than an open-source license. It is free to self-host, but read the terms if you plan to offer it as part of a service. OpenInference itself is Apache 2.0.
OpenLIT is an Apache 2.0, OpenTelemetry-native platform that adds GPU monitoring, guardrails, and a prompt hub to tracing and cost tracking. It runs on ClickHouse and is free to self-host.
MLflow Tracing brings LLM tracing and evaluation to MLflow, which many data teams already run. It is Apache 2.0, OpenTelemetry compatible, and a natural choice if your ML platform is built on MLflow or Databricks.
Laminar is a younger Apache 2.0 project focused on tracing AI agents, built on OpenTelemetry and ClickHouse.
Instrumentation: the part you should keep portable
OpenLLMetry is a set of Apache 2.0 OpenTelemetry instrumentations for LLM providers, vector databases, and frameworks. Its value is that it emits standard OpenTelemetry data, so you can send it to any backend: an LLM tool from this list, or the APM you already run. Traceloop, the company behind it, is joining ServiceNow and has said OpenLLMetry will remain open source.
OpenTelemetry GenAI semantic conventions define standard attribute names for model calls, agents, and tool use. They are still in development, not yet stable, but most of the tools here are moving toward them. Building on them now makes switching tools later much easier.
Gateways and evaluation tools
LiteLLM is a widely used open-source AI gateway that tracks spend per key, user, and team, and can forward logs to the observability tools above.
For evaluation, DeepEval (Apache 2.0) and Evidently (Apache 2.0) are popular open-source frameworks, and Promptfoo (MIT) is widely used for evals and red-teaming. OpenAI announced it would acquire Promptfoo in March 2026, and said the open-source project will continue.
Know what is not open
Some tools in this space are commercial platforms with open-source SDKs. LangSmith has an open SDK on a closed platform. Pydantic Logfire has an MIT SDK, but its server is closed source. W&B Weave requires a W&B license to self-host. They may be good products, but they are not open-source alternatives.
Ownership changes are part of the picture too. Helicone was acquired by Mintlify in March 2026 and is in maintenance mode. A tool with an active community and, ideally, a clear long-term owner is a safer foundation.
Side-by-side comparison
| Tool | Best for | License | Self-host | OpenTelemetry |
|---|---|---|---|---|
| Langfuse | Full LLM engineering platform | MIT core, enterprise extras | Free | Accepts OTLP, SDK on OpenTelemetry |
| Opik | Tracing plus evals and guardrails | Apache 2.0 | Free, full platform | Accepts OTLP over HTTP |
| Arize Phoenix | Tracing and evals with OpenInference | Elastic License 2.0 (source-available) | Free | Built on OpenTelemetry |
| OpenLIT | Tracing with GPU monitoring | Apache 2.0 | Free | Native |
| MLflow Tracing | Teams already on MLflow | Apache 2.0 | Free | Compatible |
| Laminar | Agent tracing | Apache 2.0 | Free | Native |
| OpenLLMetry | Portable instrumentation | Apache 2.0 | Not a backend | Native |
| LiteLLM | Gateway with spend tracking | MIT core, enterprise extras | Free | Forwards to tools |
GitHub figures and ownership details are as of October 2026.
The honest part: self-hosting has a cost
Most of these platforms run on a stack of databases, typically Postgres, ClickHouse, Redis, and object storage. Self-hosting them is real operational work: upgrades, backups, scaling, and on-call. For a small team, the managed cloud version of the same open-source tool is often the better deal. Langfuse, Opik, Arize, and others offer managed tiers, and you keep the option to move to self-hosting later because the code and data formats are open.
Self-hosting is worth it when:
- Prompt data is sensitive. Prompts and responses can contain customer data, source code, and personal information. Keeping them in your environment is often the main reason to choose open source.
- Volumes are high. LLM traces are large, and agent workloads produce many of them. Per-trace or per-seat pricing grows quickly.
- You already run the stack. If ClickHouse and Postgres are already part of your platform, the incremental cost is small.
A practical way to start
- Instrument with OpenTelemetry, using OpenLLMetry, OpenInference, or a tool's OpenTelemetry-based SDK. This keeps your traces portable.
- Pick one backend and start small. Langfuse and Opik are the broadest; Phoenix is strong for evaluation; MLflow fits existing ML platforms.
- Track cost per successful task, not just tokens. Tokens show activity. Cost per completed task shows value. See why cost per successful task is the right metric and why token reduction is not cost reduction.
- Set retention deliberately. You rarely need every full prompt forever. Keep summaries and metrics long, raw payloads short.
Where PointFive fits
LLM observability tools show you what one application did. FinOps for AI needs the wider view: what AI costs across every team, model, and agent, and whether that spend is efficient.
TokenShift provides visibility, governance, and optimization for AI agents on employee workstations, in production containers, and in hosted environments. Analysis runs locally, where the agent runs; derived metadata and analysis results reach the control plane, and prompt content does not. PointFive OS covers the cloud, data, and AI services those workloads run on. Together with an open-source observability tool, they give engineers the trace and give the organization the cost picture.
Frequently asked questions
What is the best open-source LLM observability tool?
Langfuse is one of the most widely adopted full platforms. Opik is a strong Apache 2.0 alternative with a fully open backend. Arize Phoenix is popular for evaluation, but it is source-available rather than open source. The right choice depends on whether you need tracing, evaluation, prompt management, or all three.
Is Arize Phoenix open source?
Phoenix is free to use and self-host, but its license is the Elastic License 2.0, which is source-available rather than an OSI-approved open-source license. Its OpenInference instrumentation is Apache 2.0.
Can I use my existing APM for LLM observability?
Partly. If you instrument with OpenTelemetry, you can send LLM traces to any OpenTelemetry backend. Dedicated LLM tools add what general APM lacks, such as prompt management, token cost tracking, and evaluations.
Should I self-host LLM observability?
Self-host when prompt data is sensitive, volumes are high, or you already run the required databases. Otherwise, start with a managed cloud of an open-source tool, which keeps the option to self-host later.