What is observability? Observability is the ability to work out what is happening inside a system from what it emits, without changing the system in order to look. The term comes from control theory, where a system is observable when its internal state can be deduced from its outputs.
In practice there is a simple test. An environment is observable when you can answer a question nobody anticipated, without shipping new code to get the answer. If finding out where a request spent its time means adding a log line and waiting for the problem to happen again, the environment is not observable.
The comparison with monitoring, and the decision about which of the two your company needs, both live on the page about infrastructure monitoring service. This text covers what comes after that decision: what the three signals cost, why most environments have only two of them, and where the bill gets out of hand.
The three signals and what each one costs
- Metric. Tells you whether things got worse, and by how much. Small, predictable volume.
- Log. Tells you what happened at that point. Large volume, grows with traffic.
- Trace. Tells you where the time went inside the request. Large volume, expensive to keep.
A metric is aggregated by nature. A thousand requests become one value. That makes it the cheapest signal, and the only one you keep for a year without thinking twice. The price of that economy is not being able to go back to one specific request.
The log keeps the individual event, with context. It is what lets you reconstruct what happened. The volume grows along with traffic, and a more verbose level of detail multiplies the bill without warning.
The trace follows a single request across the components and records how long it sat in each one. It is the signal that answers the question that is most expensive to answer without it: the call took 8 seconds, which of the eleven services was it stuck in.
The order of magnitude between the three is nothing alike. Thirty days of metrics and thirty days of traces from the same environment are bills of different sizes, and putting them on the same retention policy is the most common mistake people make starting out.
Without a correlation id, the three do not add up
Collecting the three signals does not produce observability. What produces it is being able to move from one to the next.
The mechanism is a unique id generated when the request comes in and carried through every call that follows, usually in an HTTP header. It is what lets you take a slow trace and open the logs for exactly that execution, instead of searching by timestamp.
Two practical consequences:
- Free-text logs take no part in that correlation. The record has to be structured, in fields, with the id among them. Logs that read well for a person and cross-reference badly for a machine are the most common state, and the first thing to fix.
- Propagation breaks at the weakest link. If a service along the way does not pass the id on, the trace ends there, and the invisible stretch tends to be exactly where the problem is.
Tracing is the missing pillar, and the reason is practical
Almost every environment has metrics and logs, and almost none has traces. The reason is not lack of interest, it is where each signal is born.
Metrics and logs are collected from outside the application. An agent on the server, a collector reading a file, an exporter on the database, and the data starts arriving without anyone touching the code.
Traces call for instrumentation inside the code: create the context on the way in, carry it through the calls, close it on the way out. That is work for the people who build the software, not for the people who run it. It is why the conversation about observability crosses the line between infrastructure and development, and why it stalls in companies where those two groups do not talk.
Cardinality: where the bill blows up
This is the factor that catches people out, and it does not show up in vendor material.
A metric is not a number, it is one time series for every distinct combination of labels. response_time with the labels service and environment, for ten services and three environments, is thirty series. Cheap.
Now add a label carrying the user id, or the session, or the order. Every distinct value creates a new series. Thirty series turn into hundreds of thousands, and the charge follows, because what you pay for is series stored, not requests measured.
The rule that avoids it: labels are there to group, not to identify. An individual id belongs in the log and in the trace, which are built to hold a single event. When somebody asks why the platform invoice tripled without traffic growing, the answer is almost always a new label on some metric.
Retention and sampling, decided per pillar
There is no single retention policy. There are three.
- Metric: long, because it is cheap and it is what lets you compare one period with another.
- Log: medium, with the more verbose detail living briefly and the error record living longer.
- Trace: short, and sampled.
There are two ways to sample traces, and the difference between them matters. Deciding as the request comes in is simple and cheap, and has the flaw of throwing the trace away before you know whether it was interesting. Deciding once the request has finished lets you keep exactly the ones that failed or ran long, and discard the ones that went normally, which are the majority and the ones nobody will ever look at.
One honest note: for a few dozen servers running a monolithic application, this whole apparatus is money spent for nothing. Well-tuned monitoring handles it, and the budget goes further on coverage and on the response process. Where that line is drawn is on the monitoring page.
OpenTelemetry: the standard matters more than the tool
Instrumentation is the slow, expensive part of the project, and it is the only part that stays in the company's code. The platform that receives the data is the part you swap out.
When the instrumentation is done with a vendor's proprietary library, changing platform means instrumenting all over again. That creates a dependency that is not technical, it is contractual: at renewal time, walking away carries the cost of a project.
OpenTelemetry solves that by being an open standard for instrumentation and for transport. The code emits in the standard's format, and which tool consumes it becomes a setting in the collector. Choosing a platform stops being irreversible.
Hence the practical advice: instrument on the open standard before deciding on the tool, not the other way around.
The order to roll it out
- Structure the log before centralizing it. Fields, not running text, with a correlation id. Without that, any platform takes in text and hands back text.
- Carry the id end to end. Including through the services nobody wants to touch, because that is where the trace breaks.
- Instrument one whole journey, from the user coming in to the database, rather than thin coverage everywhere.
- Set retention and sampling per signal, before switching collection on. This is where cost gets controlled, not after the first invoice.
- Pick the tool last, once you know the volume it is going to take.
The first four steps need no license at all, and they are what decides whether the fifth one works.
Observability does not tell you whether the business process happened
It explains why the system behaved the way it did. It does not check whether what the business expected actually happened.
A complete trace, no errors, normal latency, and the order placed on the e-commerce site still has no matching entry in the ERP, because the integration was never called. There is no failure to observe. There is an absent event, and absence emits no signal.
That check starts from what the process was supposed to produce and in which window, and it sits one layer above the three signals. Integrity-UX treats it as business rule monitoring, with examples of how it gets set up.
How Integrity-UX works with this
Integrity-UX runs monitoring of infrastructure, services and business rules, with alerting, incident response and figures for availability and time to resolution, through an In-house NOC.
Observability comes in when the question stops being what went down and becomes why it got slow, and it comes in after monitoring is in place, not before. To find out which layer is missing in your environment, talk to us.
Frequently asked questions
Is there a synonym for observability?
There is no exact equivalent in Portuguese. The term sometimes turns up as “visibility”, which is imprecise: visibility describes seeing what you already chose to watch, and observability is about answering a question nobody had asked.
We already run Zabbix and Grafana. Is that observability?
That is collection plus dashboards, which is the foundation rather than the whole thing. What is missing is tracing, and the correlation id that ties the three signals together. Without it, the question “which service was that call stuck in” still has no answer, however many dashboards you have.
What is cardinality in observability?
It is the number of distinct time series a metric produces, one for every combination of labels. A label carrying an individual value, such as a user or session id, multiplies that number and is the most common reason a bill runs away. Labels are there to group; an individual id belongs in the log and in the trace.
Do we have to change tools to get observability?
Not necessarily, and the order matters. Instrumenting with an open standard such as OpenTelemetry has the code emit in a format any platform can read, which lets you start with the tooling you already have and move later without instrumenting all over again. Choosing the platform before instrumenting is what creates a dependency that is hard to undo.
What does observability cost?
Cost is driven by data volume, not by the license: how much comes in per day, how long each signal stays queryable, and how many time series the instrumentation creates. That is why working out what to instrument, and setting retention per signal, comes before picking the tool.







