KPIs observability

What to measure to know if your operation
is really under control

There is an important difference between monitor and observe. The monitoring responds to the question "what is working?". The observability responds to "why it behaves like this?", and for that, you need data that can interrogate: metrics, logs and traces.

• The key question •

What are the KPIs of observability?

The KPIs of observability are the indicators that summarize a set of data into something actionable. There are charts to look at: are numbers that should trigger decisions. If a KPI goes up or down, and no one is doing anything different, that KPI is not serving for what was created.

The four KPIs
incident response

These four measure the responsiveness of your computer. Is one of the fastest reveal if the problem is in the detection in the communication or in the resolution. The KPIs of observability are the indicators that summarize a set of data into something actionable. There are charts to look at: are numbers that should trigger decisions. If a KPI goes up or down, and no one is doing anything different, that KPI is not serving for what was created.

MTTD

Average time of detection

How long it takes your system to realize that something is wrong, since the problem starts until it generates the alert.
A MTTD high means that you're enterándote of failures by your users and not your tools. Usually point to insufficient coverage of monitoring or thresholds poorly calibrated.

MTTA

Average time for recognition

How long it takes a person to take the alert was generated.
A MTTA high is a problem of process. It means that the alert is generated correctly but no one saw it, or came by a channel that no one was looking. It is especially telling if what segmentas for hours: if your MTTA night is five times during the day, you have a problem notification, not monitoring.

MTTR

Average time of resolution

How long it takes for the team to leave the service operating since acknowledged the incident. Is the KPI that more quoted, and also the most misunderstood, because a MTTR low with a MTTA high continues to be a bad experience for the user: the total time of unavailability is the sum of the three.

Availability

The percentage of time that the service was operating within a period of time. It is advisable to translate it into minutes to that number means something real: a 99.9% monthly allows nearly 43 minutes of falling; a 99,99% allows for a little more than 4.

The four signals golden

This framework, popularized by the practice of SRE, defined what to look for in any service prior to adding metrics more specific.

Latency

How much delay in responding to a request. It should be measured by separating successful responses of failed, because an error can quickly make up the average.

Traffic

How much demand is receiving the system. It is the context without which the other three metrics are not interpreted well.

Errors

What proportion of requests fails. Includes both errors explicit as the responses that reach right but outside of the acceptable time window.

Saturation

How close to the limit is the most resource constrained system. It is the only one of the four that allows you to anticipate a failure instead of constatarla.

SLI, SLO, and error budget

The three concepts that translate the metric techniques in a commitment to the business.

SLI (Service Level Indicator)

It is a specific metric that measure: for example, the percentage of applications that respond in less than 300 ms.

SLO (Service Level Objective)

It is the goal that you set out for that indicator: for example, that that percentage is at least 99.5% in a month.

The error budget

It is what is between the SLO and perfection. If your goal is to 99,5%, you have a 0.5% margin to fail. That margin is a tool of decision: while you stay in budget, the team can continue its changes; when it runs out, the priority is moving towards stability. Converts a discussion of subjective risk in an operative rule.

Download the complete guide

The Guide to KPIs of observability deepens, for each indicator, with calculation formulas, reference values by type of industry, templates to define your SLO and a checklist to audit your current schema metrics.

24Cevent - Portada guia KPIS clave para observabilidad

Frequently asked questions

What is the difference between monitoring and observability?

The monitoring checks for conditions that you defined in advance and notifies you when it is met. The observability enables you to investigate behaviors that do not anticipaste, combining metrics, logs and traces. In practical terms: the monitoring responds to questions known, the observability’t let you do new questions.

What are the KPIs of observability most important to you?

If you had to choose only four, would be MTTD, MTTA, MTTR and availability, because they cover the full cycle of an incident. From there, the four signals golden (latency, traffic, errors, and saturation) describe the health of each service, and ONLY connect the whole with the commitment towards the business.

What is the difference between MTTA and MTTR?

The MTTA measures how long it takes someone to recognize that there is an incident; the MTTR measures how long it takes to solve it once recognized. They are different problems: a MTTA high usually indicate flaws in the notification or the scaling, while a MTTR high points in diagnosis or lack of procedures. Improve one does not improve the other.

What is a good MTTR?

There is not a universal number, because it depends on the type of service and the impact of his fall. What is useful is not to compare with a reference value external, but to measure your own trend and segmentarla: on the criticality of the service, for hours and by type of incident. A MTTR stable with incidents becoming less and less frequent is a better sign than a MTTR low with incidents from recurring.

How do I define the SLO of my service?

Part is to identify what expertise really matter to the user of that service, choose a SLI represent you, and sets a goal that is achievable with your current architecture. A frequent error is to aim to 99.99% in services where no one what you need: each nine additional multiplies the cost, and a goal that you fail to comply with all the months, stops functioning as a reference.

How often should you review these KPIs?

The KPIs of incidents to be revised on a monthly basis to see trends, and after each relevant incident during the postmortem. The SLO are reviewed on a continuous basis through the consumption of the error budget, not in a regular meeting.

Do I need a specific tool to measure KPIs of observability?

The tools of monitoring and APM delivered the majority of the technical data. However, the KPIs of response (MTTA and MTTR in particular) need to record the life cycle of the incident: when was notified, who recognized him, when it was closed. That normally lives in the management layer of incidents, not in the monitoring.