The Rise of Observability

The Rise of Observability

(Co-Author: Mandar Jog , Java Professional Interview Guide- 2nd Edition)

For many years, engineering teams used the word "monitoring." This describes how they observe and keep an eye on their systems. This includes CPU usage, memory graphs, and, most importantly, alerts that fire when a server goes down. It worked reasonably well when applications were simple: one server, one database, and a predictable flow of requests. But today, software doesn't look like that anymore. A single user request today might communicate with a dozen microservices, multiple databases, a message queue, and a couple of third-party APIs before sending a response back. In such an application, if something goes wrong, you need to figure out, "which of these hundred moving parts broke, and why?"

Monitoring tells you something is wrong based on historical data.

Traditional monitoring is fundamentally about watching known metrics against known thresholds. You decide in advance what "healthy application" looks like. For example, CPU under 80%, response time under 200ms, error rate under 1%, etc., and you get an alert when reality drifts beyond these thresholds.

And as discussed, this works well for problems you've seen before. If disk space filling up has caused an outage in the past, you set up a disk-space check, and the next time it happens, you find out before your users do.

The limitation is that monitoring can only answer questions you thought to ask in advance. It's built around dashboards for known unknowns. But in a complex distributed system, most of the interesting failures are unknown unknowns. You may think of multiple retry hits, a slow downstream API, and a cache that just expired at the wrong moment. No dashboard was built in advance for that specific mess, because nobody could have predicted it.

Observability Lets You Ask New Questions

A system is "observable" if you can infer its internal state just by looking at its outputs. Applied to software, this means designing systems so that when something unusual happens, engineers can explore the data and figure out why, even if they never anticipated that particular failure mode.

Instead of only checking pre-built dashboards, engineers using an observability-driven approach can ask ad hoc questions like: "Show me every request that took more than two seconds, broken down by customer, region, and which service they hit last." That kind of open-ended investigation is what separates observability from monitoring. Monitoring tells you that something is wrong. Observability helps you figure out why.

The Three Pillars: Logs, Metrics, and Traces

Observability is usually described in terms of three types of telemetry data, often called its "three pillars":

Logs are timestamped, discrete records of events, a line of text saying something happened, like a request coming in or an error being thrown. They're detailed and specific, but can be noisy and hard to search across a large system without good tooling.

Metrics are numeric measurements collected over time, such as request counts, latency percentiles, and error rates. They're efficient to store and great for spotting trends, but they lose the specific context of any single event.

Traces follow a single request as it travels across multiple services, showing exactly how long each step took and where time was spent. Traces are what make it possible to answer "why was this one request slow?" in a system with dozens of microservices.

For a long time, these three types of data were collected and stored separately, often by different tools built by different teams, with no easy way to move between them. You might see an error rate spike in your metrics dashboard, but then have to manually dig through unrelated log files to find out what actually happened, and separately pull up a tracing tool to see which service was the bottleneck.

Why "Unified" Matters

The real shift in the last few years isn't just that companies are collecting more logs, metrics, and traces. But now, they're being tied together. Modern observability platforms let you start at a single spike on a metrics graph, click through to the traces from that time window, and then drill

into the exact logs from the specific service that was struggling. You can do all this without switching tools or losing context.

This is largely possible because standards like OpenTelemetry give logs, metrics, and traces a shared way to identify which request, service, and transaction they belong to. Once everything speaks the same language, correlation becomes automatic instead of manual detective work. The result is a much shorter path from "something is wrong" to "here's exactly why, and here's the line of code or the downstream dependency responsible." For teams running complex, distributed systems, that difference can mean minutes of downtime instead of hours.

Wrapping Up

Traditional monitoring isn't going away — it's still the right tool for watching known metrics against known thresholds. But as systems have grown more distributed and unpredictable, monitoring alone hasn't been enough to answer the harder questions engineers face during an incident. Observability, built on a unified combination of logs, metrics, and traces, lets teams explore their systems freely and understand failures they never planned for.

If your team still treats logs, metrics, and traces as three separate tools with three separate dashboards, that's usually the clearest sign it's time to think in terms of observability rather than just monitoring.

Back to blog