What Is Observability? OpenTelemetry, Logs, Metrics and Traces

What is observability?
Observability is the ability to understand the internal state of a software system from the outputs it produces, such as logs, metrics, and traces. The goal is to answer "what is broken right now, and why?" without opening the code. As a result, you can diagnose problems you never planned for.
For example, think of a car dashboard. A fuel gauge tells you one thing. However, when the engine makes a strange noise, the gauge alone cannot explain it. You need rich data that shows how sensors, records, and parts connect. Software works the same way, and that is what observability gives you.
This article answers what is observability and covers its open data standard, OpenTelemetry. Also, we do not rank products or praise a vendor. Instead, we explain the concept, its building blocks, and where a small team can start.
The idea itself comes from control engineering, where engineers infer a system's internal state from its outputs. Then software teams adopted the same principle: the better the data a system produces, the better you understand it.
What is observability, and how does it differ from monitoring?
Monitoring tracks problems you already know about. You set a threshold in advance, for example an alert when the error rate rises. Instead, observability lets you ask questions about problems you did not predict. Put simply, monitoring answers "is something wrong?" while observability answers "why is it wrong?"
The two are not rivals. Also, observability produces the data that monitoring needs. A good alert design also depends on that data: you attach alerts to symptoms users feel, then search for the cause with data. The table below summarizes the difference.
| Criterion | Classic monitoring | Observability |
|---|---|---|
| Core question | Is the system up? | Why does the system behave this way? |
| Problem type | Known, predefined issues | Unexpected, new issues |
| Data | Mostly ready dashboards and thresholds | Logs, metrics, and traces together, queryable |
| Approach | An alert arrives, then you look | You query the data and narrow the cause |
| Limit | Cannot see what you did not define | Cannot see what you never emit |
So observability does not replace monitoring. It adds a deeper diagnostic layer on top of it.
What is observability for a business, and why does it matter?
For a business, the practical answer is this: you find a problem before customers complain, or within minutes after they do. A slow checkout page, a silently failing email, or a spike in errors after a release all hit revenue directly.
It also changes how a team communicates. Because there is no data, diagnosis meetings run on guesses and blame. With a shared data source, everyone looks at the same chart and the same trace. The debate shifts from "who made the mistake?" to "why did the system behave this way?"
- Faster diagnosis: You narrow the problem instead of hunting for it.
- Safer releases: You see the effect of a new version immediately.
- Better planning: You base capacity decisions on data.
- Calmer teams: People stay steady because they know what is happening.
What are telemetry and the three signals?
Telemetry is the data a system emits to describe its own behavior. The OpenTelemetry documentation groups it into three main signals: traces, metrics, and logs. In practice, each signal shows the same event through a different lens.
- Log: a timestamped record of an event, such as a rejected payment request.
- Metric: a numeric value measured over time, such as request count or error rate.
- Trace: the journey of a single request across services.
| Signal | What it shows | Strength | Limit |
|---|---|---|---|
| Log | Individual events | Rich detail and context | Volume grows, searching is hard |
| Metric | Trends over time | Cheap and fast, good for alerts | Does not explain a single request |
| Trace | A request across services | Finds the bottleneck | Loses meaning if context breaks |
The documentation also mentions extras such as baggage and profiles. For a start, however, the first three are enough. We look at each one next.
What is a log, and what does it do in observability?
A log is a timestamped message that a service or component writes. It can carry an error detail, a user action, or a system event. For that reason, logs usually give you the most detail.
On its own, though, a log lacks context. You see an error message but cannot tell which user request it belongs to. The OpenTelemetry documentation notes that logs become far more useful once you correlate them with traces and spans.
One practical tip: write logs as structured fields instead of free text. Then you can search by severity, service name, and request ID. Also keep personal data out of logs. Passwords, card numbers, and identity details should never appear in a log record.
Also, choose log levels deliberately. If you separate error, warning, and info levels, you keep only what you need switched on in production. That lowers storage cost and also shortens the time it takes to find the line you want.
What is a metric, and when should you check it instead of logs?
A metric is a numeric value that you measure and aggregate over time. Request count, error rate, processor use, and response time are typical examples. Metrics are cheap because they store a summary, not every single event.
That is why you check metrics first when you ask, "how is the system doing overall?" For example, if a response time chart has been climbing for an hour, you spot the problem at once. However, a metric will not tell you why one specific user got an error. Logs and traces cover that.
In practice, metrics also form the base of alerts. Therefore, attaching an alert to a meaningful metric creates less noise than attaching it to a log search. Also look at the distribution, not only the average. A small share of slow requests can hide behind a healthy average, yet those requests drive most complaints.
What is a trace in distributed tracing, and what is a span?
A trace records the end-to-end journey of one request as it passes through several services. A trace consists of spans, which are small units of work. Each span represents one operation and carries a name, timing, and extra details called attributes.
For example, an order request goes to the web app first, then to the stock service, and next to the payment service. Then each step becomes a span. On a trace screen, you see which step took how long, in order.
So instead of guessing why the order page is slow, you measure it. In systems with several components, tracing is the fastest diagnostic tool. In a small single-part application, good logs and metrics often suffice. Still, trying tracing early prepares you for the day you need it.
How do logs, metrics, and traces work together?
The three signals do not replace each other. Instead, they link to each other. A typical diagnosis flows like this: a metric shows a deviation first, a trace narrows the faulty step, and a log explains the detail last.
Example scenario: Imagine the payment step of an online shop slows down on some evenings. The steps below show how the three signals work together.
- The response time metric shows an increase in payment requests.
- Then you open the trace of one slow request.
- Next, the trace shows that the lost time sits in the external call to the payment provider.
- Then the log line attached to that span explains the timeout and the retry count.
- Finally, you can prove that the cause is an outside dependency and not your own code.
That also tells you what to do next. For example, you tune the timeout, add a fallback path, or contact the provider with concrete evidence. The chain depends on a shared request ID. Without that link, you read three different stories in three different dashboards.
What is OpenTelemetry?
OpenTelemetry is an open source, vendor-neutral observability framework for generating, collecting, and exporting telemetry data. It is also a project under the CNCF. It supports traces, metrics, and logs.
One key detail: OpenTelemetry is not an observability backend. It does not store data, draw charts, or raise alerts. Instead, it only helps your application produce data in a standard shape and send it where you choose. You pick a separate tool for storage and visualization.
The project grew out of the merger of two earlier efforts, OpenTracing and OpenCensus. For that reason, it offers one shared language. You can read the details in the official OpenTelemetry documentation.
Why is OpenTelemetry vendor-neutral?
In the past, every monitoring product shipped its own agent and data format. Switching products meant rewriting the measurement code inside your application. We call that situation vendor lock-in.
So OpenTelemetry separates the measurement code from the products. You instrument your application once with the standard method, and you decide through configuration which backend receives the data. In other words, changing the backend does not require changing the application code.
This also matters most for small teams. You can start today with a cheap, simple option and move to another one when your needs grow. However, "neutral" does not mean "free". Storage and operating costs stay with you.
Also, when you evaluate a backend, check three points. First, ask whether it accepts OpenTelemetry data directly. Next, test whether its query language fits your team. Finally, look at how flexible its retention and sampling options are. Check current features in the provider's official documentation.
What parts make up OpenTelemetry?
The project has a few core components. The documentation lists language SDKs, APIs, a standard protocol, and semantic conventions. Below, the list explains the role of each part. Together, these parts complement each other, and each one solves a different problem.
- API: the common interface your code calls to produce telemetry.
- SDK: the language-specific library that implements the API and processes the data.
- OTLP: the standard protocol that carries telemetry over the network.
- Semantic conventions: shared names for fields, so "request method" looks the same everywhere.
- Collector: a standalone component that receives, processes, and forwards data.
Still, you do not need every part on day one. In a small project, you can start with the SDK and add the collector later. What matters is that you write the measurement code in the standard way, because you can change the rest through configuration.
What is the OpenTelemetry Collector, and what does it do?
The collector is a vendor-agnostic component that receives, processes, and exports telemetry data. The official documentation describes three parts: receivers accept data, processors transform it, and exporters send it to a destination. As a result, you do not run a separate agent for every signal.
For production, we recommend a collector, because your application can hand data off quickly and get back to its own work. Also, the collector takes care of batching, retries, and filtering. Consequently, your application spends less effort on telemetry.
Here is an analogy. Put simply, a collector works like the shared mailbox of an apartment building. Every flat drops its mail in one place, and the person at that spot handles the delivery. For details, see the Collector documentation.
In practice, a small project runs the collector as one small process. As needs grow, you run several instances and route different signals to different destinations. That flexibility also lets you switch backends without touching the application.
What is instrumentation, and is it automatic or manual?
Instrumentation is the work that makes your application emit telemetry. It means producing trace, metric, and log signals from inside the code. OpenTelemetry offers two paths.
- Automatic instrumentation: ready-made components for many popular frameworks and libraries collect basic signals with almost no code change.
- Manual instrumentation: you mark business-specific steps yourself, for example a meaningful span for the coupon step.
So we suggest starting with automatic instrumentation. Basic requests and database calls become visible quickly. Then add manual spans only for the questions you really ask.
A useful rule for manual work: mark the places where you once said, "I wish I had this information at midnight." A span on every line creates noise and cost. Also, naming matters. Give spans clear, consistent names, because random names turn a trace screen into a hard-to-read pile. A short naming guide inside the team pays off for months.
What is context propagation, and how do you follow a request across services?
Context propagation is the mechanism that carries a request's identity from service to service. Because every service holds the same trace ID, spans created separately merge into one trace. Without it, distributed tracing stays fragmented.
The context usually travels in the request headers to the next service. OpenTelemetry uses a standard format for that. That means services written in different languages can still follow the same request together.
In practice, the most common mistake is losing context at asynchronous hand-offs such as message queues or background jobs. If a trace breaks when a job enters a queue, check that point first. In addition, if you add the trace ID to your log lines, you can jump straight from a log to its trace. That small step shortens midnight diagnosis noticeably.
How do you manage data volume and cost?
Observability data grows fast. Every request produces logs, metrics, and traces, and storage and processing costs rise with traffic. For that reason, treat cost as part of the design from the start. We give no figures here, because prices vary by provider. Check current values on your provider's official page.
These are the main ways to control cost.
- Sampling: keep a representative subset of traces, not all of them.
- Filtering: drop noisy, low-value data such as health checks in the collector.
- Retention: keep detailed data for a short time and summary metrics for longer.
- Label discipline: labels with many different values blow up the metric count, so limit them.
The documentation says sampling is one of the most effective ways to reduce cost without losing visibility. See the sampling documentation for details.
What is the difference between head and tail sampling?
First, head sampling decides at the start of a request. It is simple and cheap. However, it cannot know what will happen later, so it cannot guarantee that it keeps every failed request.
Second, tail sampling decides after the trace completes. For example, you can say "keep every trace that contains an error" or "keep the slow ones." It also preserves the most valuable data for diagnosis. In return, it needs a more complex, stateful setup.
For small teams, the practical path is to start with simple sampling and move to tail sampling if the need appears. So building both right away is usually over-engineering. Sampling has a cost too: a rare event that misses your rule may not reach your records. So define a rule that always catches errors and slow requests.
What is cardinality, and why does it raise cost?
Cardinality is the number of different values a label can take. A "payment method" label takes a few values, so its cardinality is low. A "user ID" label differs for every user, so its cardinality is very high.
In metrics, every distinct label combination creates a separate time series. Therefore, using IDs, email addresses, or full URLs as metric labels makes the series count explode. The result is slow queries and an unexpected bill.
So the rule is simple: put detail in logs and traces, and put the general trend in metrics. Instead, if you need per-user detail, carry it as an attribute inside the trace. That way metrics stay cheap and the detail remains within reach.
Where does observability help in real life?
People often ask what is observability in practice. In a few typical situations, you see the difference at once. The examples below are scenarios, not real customer results.
- Slow page complaint: a trace shows whether the delay sits in a database query or an external call.
- Error spike after a release: a metric shows the deviation started with the new version, so you decide on a rollback quickly.
- Intermittent error: you find a failure that affects only one user group through a shared field in the logs.
- Capacity planning: metric trends tell you in advance when to grow a server or database.
For reading server load, see our load average guide. For managing Node.js processes, see our PM2 guide. Both topics feed observability data.
Example scenario: Picture a small parcel tracking app. Users say the lookup is very slow in the morning. First, a metric shows that the slowdown hits only one endpoint. A trace then shows the lost time sits in one database query that repeats on every request. So the fix might be an index or a cache. What matters is that you decide with evidence, not guesses.
How does observability differ from commonly confused terms?
Several close terms live in this field. The table below separates the ones people mix up most often.
| Term | What it covers | Relation to observability |
|---|---|---|
| Monitoring | Thresholds and alerts for known problems | One way of using observability |
| Logging | Text records of events | Only one of the three signals |
| Telemetry | Raw data a system emits | The raw material of observability |
| APM | A product category that tracks application performance | Can be one backend you use for observability |
| OpenTelemetry | Standard for producing and collecting data | Supplies data, does not store or display it |
| Observability backend | Tool that stores, queries, and shows data | Consumes the data OpenTelemetry sends |
In short, OpenTelemetry produces the data, a backend stores and shows it, and observability is the ability that the whole chain gives you.
What are the advantages, limits, and risks?
An honest answer to what is observability should show both the benefit and the price. The list below puts them side by side. So when you decide, look at the maintenance load too, not only at the benefits.
- Quick fixes: diagnosis time drops, because you work with data instead of guesses.
- Flexibility: a common standard makes switching tools easier.
- Effort limit: if you do not emit data, you get no visibility. Instrumentation takes work.
- Privacy risk: personal data that leaks into logs creates a legal and trust problem.
- Alert fatigue: too many alerts teach the team to ignore them.
- Cost risk: uncontrolled data volume grows the bill in unexpected ways.
For this reason, observability is a habit that needs steady care, not a one-time install. This article is not legal advice, so clarify personal data rules with your own legal counsel.
What are the most common observability mistakes?
Teams fall into the same few traps again and again. Knowing them upfront saves both time and money.
- Logging everything: noise grows, the real signal disappears, and cost rises.
- Leaving signals unconnected: three separate dashboards tell three separate stories.
- Alerting on every metric: the team grows used to alerts and misses the real one.
- Writing personal data to logs: this creates a privacy risk.
- Leaving dashboards without an owner: nobody looks at them anymore.
However, the fix is usually a habit, not a technology. Ask "how will we observe this?" in the design phase of every new feature. Even one extra line in the pre-release checklist makes a difference.
What is observability for a small team, and where do you start?
For a small team, the goal is not to measure everything. Instead, collect the signals that help you most, and collect them reliably. The checklist below is a reasonable order to start.
- Pick your most critical user path, for example login, order, or payment.
- Define three basic metrics for that path: request count, error rate, and response time.
- Then write logs as structured fields and keep personal data out.
- Turn on automatic instrumentation and confirm that traces connect across services.
- Also add the trace ID to your log lines.
- Next, send data through a collector, not straight to a backend.
- Set a simple sampling and retention policy.
- Attach alerts only to symptoms that affect users.
- Review data volume and cost trends every quarter.
You do not have to finish the list in one go. Even the first three items speed up your diagnosis noticeably.
Which signal answers which question?
Going to the right signal quickly saves time during diagnosis. The table below shows the starting point by question type, so you can put observability to work day to day.
| Your question | Look at first | Why |
|---|---|---|
| Is the system healthy overall? | Metric | It gives a summary and trend |
| Why was this request slow? | Trace | It shows the time step by step |
| What is the detail of this error? | Log | It carries the context of the event |
| Which service is the bottleneck? | Trace and metric | They show time spread and load together |
| Did the new release break something? | Metric, then trace | It spots the deviation, then narrows the cause |
Over time, this order becomes a reflex. It also builds a shared diagnostic language inside the team.
How do you build observability into your software project?
Once you know what is observability, the practical step is to start collecting your first signals right away. Also, do not wait for perfect dashboards. Start small and grow with need.
Observability is cheapest and most effective when you design it together with the architecture. You can add it later, but broken context and noisy logs create cleanup cost. So when you plan a new project, discuss the signal design too.
At Talha Aslan and team, we treat observability as a natural part of the architecture conversation in software projects. For your own project, you can look at our custom software development service.
We also have separate guides on neighboring topics. To switch features on and off without redeploying code, read our feature flag guide. To manage infrastructure through code, read our Infrastructure as Code guide. So together they give you a safe and transparent release process.
If your website feels slow on the server side, our server-side slowness guide and our log file analyzer are good starting points.



