When teams talk about cloud efficiency, they usually mean rightsizing instances, cleaning up idle resources, and buying commitments. All of that matters. But some of the largest savings come from a different decision: which tools you run in the first place.
Observability is the clearest example. For years, the default answer to "how do we monitor production?" was a commercial APM platform: Datadog, New Relic, Dynatrace, Splunk. They are good products. They are also priced on the things that grow fastest in a modern environment: hosts, containers, functions, log volume, indexed events, and custom metrics. As architectures move to microservices, serverless, and AI workloads, the observability bill tends to grow faster than the infrastructure it watches.
Meanwhile, something changed in the open-source world. A new generation of observability tools is built on open standards, columnar storage, and active communities. Many of them are cheaper by design, faster to query, and more transparent about how they work. For a growing number of teams, they are not a compromise. They are the better choice.
This post maps that landscape, names the tools worth knowing, and is honest about when open source saves money and when it does not.
Why open-source observability got good
Three shifts made this possible.
OpenTelemetry became the standard. OpenTelemetry gives you one vendor-neutral way to instrument traces, metrics, and logs. It graduated from the CNCF in May 2026, and its trace, metric, and log protocols are stable. Once your code emits OpenTelemetry, the backend becomes a choice you can change, not a lock-in you live with.
Columnar storage changed the economics. Many of the newer tools store telemetry in ClickHouse or in columnar files on object storage. Telemetry is highly repetitive, so it compresses well and queries fast in a columnar format. Storage that used to be the expensive part of observability becomes one of the cheap parts.
The communities are real. These are not side projects. Grafana has more than 77,000 GitHub stars, Prometheus more than 66,000, SigNoz more than 32,000, and Apache SkyWalking almost 25,000 (as of October 2026). They ship releases every few weeks, and their roadmaps, issues, and code are public.
The landscape
All-in-one, OpenTelemetry-native platforms
These give you a Datadog-like experience in one product: traces, metrics, logs, dashboards, and alerts.
SigNoz is an OpenTelemetry-native APM built on ClickHouse, covering traces, metrics, logs, infrastructure monitoring, and alerts. It is one of the most complete open-source alternatives to commercial APM, with a self-hosted edition and a managed SigNoz Cloud priced on usage rather than per host. The core is MIT licensed, with enterprise features in a separate directory.
ClickStack is ClickHouse's observability stack: ClickHouse, an OpenTelemetry Collector, and the HyperDX interface, which ClickHouse acquired in 2025. It covers logs, traces, metrics, and session replay, and it is strong at fast search over very large volumes of data. HyperDX is MIT licensed. A managed version, Managed ClickStack, launched in beta in February 2026.
OpenObserve stores logs, metrics, and traces as columnar files on object storage, with stateless query nodes, and adds RUM, session replay, and pipelines. It is AGPL licensed, with a free self-hosted edition and a managed cloud.
Uptrace is a smaller OpenTelemetry APM built on ClickHouse, AGPL licensed, with a free community edition.
Coroot uses eBPF to collect metrics, logs, traces, and profiles without code changes, and adds automated root-cause analysis. It is Apache 2.0 licensed, with a free community edition.
The composable stack: Grafana, Prometheus, and VictoriaMetrics
If you prefer to assemble best-of-breed pieces, the Grafana ecosystem is the most widely used open-source observability stack:
- Prometheus for metrics, the CNCF standard for a decade
- Grafana for dashboards and alerting
- Loki for logs, Tempo for traces, Mimir for long-term metrics, and Pyroscope for continuous profiling
- Alloy, Grafana's OpenTelemetry Collector distribution, which replaced the Grafana Agent
Grafana, Loki, Tempo, Mimir, and Pyroscope are AGPL licensed. If you do not want to run them yourself, Grafana Cloud offers the same stack as a service, including a free tier.
VictoriaMetrics is an Apache 2.0 alternative for metrics, designed to store large metric volumes efficiently, and it now has sister projects for logs (VictoriaLogs) and traces (VictoriaTraces, still pre-1.0).
Classic open-source APM and tracing
Apache SkyWalking is a mature APM from the Apache Software Foundation, with service topology, traces, metrics, logs, eBPF profiling, and its own observability database, BanyanDB. It is Apache 2.0 licensed and fully community governed. Version 11 shipped in August 2026.
Jaeger is the CNCF's graduated distributed tracing system. Jaeger v2 is built on the OpenTelemetry Collector. Zipkin is the original open-source tracer, still maintained and still simple.
Search-based platforms
Elastic Observability is a full observability platform on Elasticsearch. Elasticsearch and Kibana added AGPL as a license option in 2024, but read the details: the observability applications in Kibana sit in the part of the code base licensed only under the Elastic License 2.0, which is source-available rather than open source.
Quickwit is a fast search engine for logs and traces on object storage. Datadog acquired Quickwit in January 2025 and moved it to the Apache 2.0 license. It is still released, but there is no managed offering.
Newer data-lake engines
Parseable (written in Rust, AGPL) and GreptimeDB (Apache 2.0) store logs, metrics, and traces in columnar formats on object storage. They are younger, but they show where the architecture is heading: cheap object storage, separated compute, open formats.
The foundation: the OpenTelemetry Collector
Whatever backend you choose, the OpenTelemetry Collector is the piece that makes the choice reversible. It receives telemetry from your services, processes it (sampling, filtering, redaction, routing), and sends it to one or more backends. Put it in the middle, and you can run two backends side by side, migrate gradually, or drop data you never needed before it reaches anyone's bill.
Side-by-side comparison
| Tool | Best for | Signals | Storage | License | Managed option |
|---|---|---|---|---|---|
| SigNoz | All-in-one APM | Traces, metrics, logs | ClickHouse | MIT core, enterprise extras | SigNoz Cloud |
| ClickStack (HyperDX) | Fast search at high volume | Logs, traces, metrics, session replay | ClickHouse | MIT (HyperDX), Apache 2.0 (ClickHouse) | Managed ClickStack (beta) |
| Grafana stack | Composable, widely adopted | Metrics, logs, traces, profiles | Prometheus, Loki, Tempo, Mimir | AGPL (core apps), Apache 2.0 (Alloy) | Grafana Cloud |
| OpenObserve | Low-cost storage on object storage | Logs, metrics, traces, RUM | Columnar files on object storage | AGPL | OpenObserve Cloud |
| Apache SkyWalking | Community-governed APM | Traces, metrics, logs, profiles | BanyanDB | Apache 2.0 | Self-host only |
| Coroot | Zero-instrumentation with eBPF | Metrics, logs, traces, profiles | Prometheus and ClickHouse | Apache 2.0 | Paid editions |
| VictoriaMetrics | Efficient metrics at scale | Metrics (logs, traces newer) | Own storage engine | Apache 2.0 | VictoriaMetrics Cloud |
| Elastic Observability | Teams already on Elasticsearch | Logs, metrics, traces | Elasticsearch | AGPL/SSPL/ELv2 core, ELv2 observability apps | Elastic Cloud |
| Jaeger | Distributed tracing | Traces | Pluggable | Apache 2.0 | Self-host only |
GitHub star counts and release dates referenced in this post are as of October 2026.
The honest part: open source is not automatically free
Open-source software has no license fee. Running it does have a cost, and it is worth being clear about it.
When you self-host, you take on the work. Someone has to size, upgrade, back up, and scale the storage layer. Someone carries the pager when the observability system itself has a problem, usually at the same moment production does. For a small team, an engineer's time can cost more than the license you saved.
Self-hosting pays off when:
- Telemetry volume is high and growing, so storage and ingestion pricing dominates the bill
- You already run stateful infrastructure well, such as databases or Kubernetes operators
- Data residency or security requirements push you to keep telemetry in your own environment
- You want full control over retention, sampling, and cost
A managed open-source cloud is often the better deal when you want the economics and openness without the operations: Grafana Cloud, SigNoz Cloud, Managed ClickStack, Elastic Cloud, and others offer the open-source engine as a service, usually priced on data volume rather than hosts. You keep the open formats and the OpenTelemetry instrumentation, so you can still move later.
Read the license and the ownership. "Open source" covers a range: Apache 2.0 and MIT are permissive; AGPL is open source but has obligations if you embed it in a service you offer; several products are open core, with enterprise features under a commercial license; and some products marketed as open are source-available. Ownership changes matter too. HyperDX joined ClickHouse, Quickwit joined Datadog, and Highlight.io was acquired by LaunchDarkly and shut down in February 2026. Projects governed by a foundation, such as Prometheus, Jaeger, OpenTelemetry (CNCF), and SkyWalking (Apache), carry the least of that risk.
And always label vendor benchmarks as what they are. Many projects publish dramatic savings figures against commercial APM. Some will hold for your workload, some will not. Test with your own data.
How to move without a big-bang migration
- Instrument with OpenTelemetry first. This is the step that removes lock-in, whatever you decide next. Most commercial APM vendors, including Datadog, accept OpenTelemetry data, so you can do this before changing anything else.
- Put an OpenTelemetry Collector in the middle. Route telemetry through it, and you can send the same data to your current vendor and to an open-source backend at the same time.
- Start with logs. Logs are usually the largest and least-used part of the bill. Moving them to columnar storage tends to deliver the biggest savings with the least risk.
- Cut what you never read. Filter debug logs, sample high-volume traces, drop unused metrics and high-cardinality labels in the Collector. The cheapest telemetry is the telemetry you do not send.
- Compare on real costs. Include compute, storage, retention, engineering time, and support, not only the per-GB price.
Where PointFive fits
Observability waste is real waste, and it is often invisible. PointFive OS detects inefficiencies in observability spend on the platforms you already use, such as indexing every ingested Datadog log without exclusion filters, unnecessary default log retention in Datadog, CloudWatch log volume from debugging left on, and high-cardinality custom metrics. When we audited GCP environments, every one had cloud logging waste.
Whether you stay on a commercial platform or move to open source, the principle is the same: pay for the telemetry you use, and nothing more. For LLM applications, see the companion guide to open-source LLM observability tools.
Frequently asked questions
What is the best open-source alternative to Datadog?
It depends on what you need. SigNoz and ClickStack are the closest to an all-in-one APM experience. The Grafana stack is the most widely adopted composable option. Apache SkyWalking is a mature, community-governed APM. If you want the open-source engine without running it, look at the managed clouds of these projects.
Is open-source observability really cheaper?
Often, especially at high data volumes, because many of these tools use columnar storage and are priced on data rather than hosts. But self-hosting moves cost from a license to engineering time and infrastructure. Compare total cost of ownership, not license fees.
Can I use OpenTelemetry with Datadog or New Relic?
Yes. Both accept OpenTelemetry data. Instrumenting with OpenTelemetry first is the safest step, because it keeps your options open whatever backend you choose.
What should I migrate first?
Logs. They are usually the largest share of observability spend and the easiest to move with low risk, especially through an OpenTelemetry Collector that can send to both backends during the transition.