How to set RTO and RPO: why the number the business gives you is not the RTO

Duas pessoas numa sala de reunião diante de um quadro branco ainda em branco

Every recovery plan rests on two numbers. RTO, the longest acceptable gap between the outage and the service coming back, and RPO, the most data you can afford to lose, measured in time. Almost every mid-sized and large company has both written down somewhere.

What almost none of them can show is how they arrived at them.

That matters, because a number without a method does not survive the first real incident. It was agreed in a meeting, and on the day of the crisis you find out the architecture was never built to meet it, or that the promised window left out half the work that getting back on your feet actually takes. This article is about the method: where each number comes from, who owns it, and how to check that it holds up before it becomes a commitment.

If you are a step behind that, on what a recovery plan is and what goes into it, the article on DRP covers that ground.

MTD comes first, and it is not the RTO

Here is the most common mistake, and it does not look like one.

IT asks the business how long operations can go without the system. The business says 72 hours. IT writes down an RTO of 72 hours. Everyone leaves the meeting happy.

The number they just agreed on is not the RTO. It is the MTD.

The NIST SP 800-34 Rev. 1, the contingency planning guide from the US standards institute, keeps the two apart. MTD, Maximum Tolerable Downtime, is the total outage time the system owner is willing to accept, with every impact taken into account. RTO is the longest a resource can stay unavailable before it does unacceptable damage to other resources, to the business processes they support, and to the MTD itself.

The line in the document that clears up the confusion is this one: because the RTO has to keep the MTD from being exceeded, the RTO normally has to be shorter than the MTD.

The reason is practical. Between the system answering again and the business working again, there is work to do. Entries written on paper during the outage have to be keyed in. Integrations that stalled have to clear the backlog. Someone has to confirm the balances add up before invoicing goes out. None of that time belongs to IT, and all of it runs on the same clock.

So if the business can absorb 72 hours and reconciliation takes 24, the RTO is 48 hours. That is exactly the example NIST uses in its own impact analysis template: a vendor invoice payment process with an MTD of 72 hours, an RTO of 48 hours and an RPO of 12 hours.

The arithmetic is easy. Doing it in the wrong order is what produces plans that fail inside the window they promised.

The question is not about the system, it is about the process

The second thing that stalls the work is asking what the RTO of the ERP is.

Nobody can answer that, and anyone who does is guessing. The ERP carries invoicing, purchasing, payments, the monthly close and historical lookups, and those five have nothing like the same tolerance. Invoicing cannot wait a day. Historical lookups can wait a week and no one will notice.

NIST breaks impact analysis into three steps, and the first is to determine business processes and recovery criticality. The unit of analysis is the process, not the server. Only then come the resources each process consumes, and last the order in which those resources are recovered.

In practice this turns the conversation around. Instead of listing systems and asking how long each one can be down, you list what the company does, find out how long each activity can wait, and then map which systems it depends on. A server's RTO becomes a consequence: it is the tightest deadline among the processes that rely on it.

The side effect is welcome. Once the number comes out of a process, it stops being IT's opinion and gets an owner on the business side.

Turning "it's critical" into a number you can compare

In the first round, everything is critical. That is not obstruction, it is the absence of a scale. Without a shared yardstick, "critical" is the only word available for "it matters to me".

NIST handles this by asking the organization to create impact categories and assign values to them, so the severity of an outage can actually be measured. The example in the document uses cost as the category, with three bands: severe, when temporary staffing, overtime and penalties run past a million dollars; moderate, around five hundred and fifty thousand; and minimal, around seventy-five thousand. The guide is explicit that these figures are only an example and should be revised to fit the organization.

The point of the yardstick is not precision. It is that it forces comparison. Once two departments have to place their own process on the same money scale, the priority list writes itself.

Three questions that work better than "is it critical":

What happens in the fourth hour? A general question gets a general answer. A question with a clock on it gets a scenario: in the fourth hour the carrier leaves without the load, and delivery slips a day.

Is there a manual way to keep going? If there is, the MTD is longer than it looked. If there isn't, it is shorter. NIST asks for this explicitly in the assessment, and it is the answer that moves the numbers most.

How long does the manual way hold? A notebook and a spreadsheet work for a few hours. After that, the keying backlog becomes the new problem, and that time comes off the MTD.

RPO is a question about rework, not about backup

RPO usually gets treated as a technical setting on the backup job. It isn't.

NIST is blunt on two points. First, RPO is the point in time, before the outage, to which the process's data has to be recovered, given the most recent copy available. Second, and this one almost always gets missed: unlike RTO, RPO is not counted as part of the MTD. It is a separate factor, covering how much data loss the process can tolerate.

That has a practical consequence. The person who answers the RPO question is not the one who runs the backup, it is the one who will retype. The right question is how many hours of entries the team can redo by hand without stopping everything else, and with what risk of error.

And there is an honesty check worth running before you publish the number. RPO cannot be shorter than the interval between copies. An RPO of fifteen minutes with a nightly backup is an intention, not a capability. If the number the business needs is shorter than the current window, the conversation is no longer about the number, it is about changing how the copies are made, which the article on cloud backup goes into.

The sanity check: does the number fit the technology and the budget

With the MTD set per process, the RTO derived and the RPO agreed, what is left is the part the business has no way of knowing on its own: what each band costs.

NIST frames this as cost balancing. The longer the outage runs, the more it costs. The shorter the RTO, the more the recovery solutions cost to put in place. The document notes that the balance point between those two curves is different for every organization and every system, because it depends on financial constraints and operating requirements.

To turn that into orders of magnitude, AWS publishes the bands each strategy covers in REL13-BP02:

  • Backup and restore: RPO in hours, RTO in 24 hours or less. With continuous backups and point-in-time recovery, the RPO can come down to about five minutes.
  • Pilot light: RPO in minutes, RTO in tens of minutes.
  • Warm standby: RPO in seconds, RTO in minutes.
  • Multi-Region active-active: RPO near zero, RTO potentially zero.

Treat the bands as a check, not as a catalog. If the business asked for a ten-minute RTO and the current architecture is backup and restore, no amount of effort on the day of the incident closes that gap. Investment closes it, or the number changes.

It is also worth noting what the providers do not guarantee. Microsoft states in its own architecture framework that it publishes RTO and RPO guarantees for some products only. For everything else, the number belongs to whoever designed the solution, and not to the service agreement.

And when it cannot be done? NIST covers that case: when the RTO cannot be met right away and the MTD will not move, the situation should be documented formally, with a plan of action and milestones to close it. Recording the gap is a defensible decision. Writing down a number nobody can hit is not.

The RTO and the availability target have to add up

There is a contradiction in almost every requirements document, and it almost never gets caught, because the two numbers live on different pages.

One page carries the availability target, usually 99.9%. The other carries the RTO, usually a few hours.

The service-level objective table in Microsoft's architecture framework shows what 99.9% allows: 43.20 minutes of downtime a month, or 8.76 hours a year.

Now compare that with the RTO. A single incident that uses up a four-hour RTO burns nearly half the annual downtime budget of a 99.9% target. Two of those in a year blow the target, even if both were handled exactly within the agreed window.

It is not that one of the numbers is wrong. It is that a third one is missing: how many incidents a year the target assumes. Without it, the first two do not talk to each other, and the company meets the RTO and misses the target in the same incident.

Who signs off on the number

In NIST's definition, the MTD is the time the system owner is willing to accept. Willing, singular. That is a person, not a committee.

This matters at three moments. When the number is set, because someone has to stand behind the choice once the cost of the architecture is on the table. During an incident, because that is who authorizes the return to normal operations. And at review time, because numbers age: a process whose volume changed, a system that picked up a new integration, an operation that now covers another time zone.

That is why the criticality sheet needs three columns that are usually missing: the name of the person who answered, the date they answered, and what triggers the next review.

How does Integrity-UX conduct this work?

We start with the assessment alongside the business areas, process by process, with the impact scale built before the first meeting. That step produces the MTD for each process, with an owner and a date.

Next we map each process's dependencies and derive the RTO per resource, minus the time for reprocessing and verification. The RPO is agreed with the people who do the data entry, and checked against the real copy window.

With the numbers settled, we show what each band demands in architecture and investment, and where the gap sits between what was asked for and what the current infrastructure delivers. The result is a short table of process, MTD, RTO, RPO and strategy that fits on one page and anchors the recovery plan.

Next steps

If your company already has RTO and RPO on paper, here is a quick test: take the number for one process and try to say who set it, on what date, and whether it subtracts reprocessing time. If all three answers do not come, the number needs a review before any spending on architecture.

To see where these numbers sit in the full plan, the article on DRP covers the structure of the document and how often to test it. For the copy layer, which is where the RPO either holds or doesn't, the article on cloud backup covers retention and the 3-2-1 rule.

If you would rather talk about your own case, the DRP page has a direct contact.

Frequently asked questions

What is the difference between MTD and RTO?

MTD is the total outage time the company accepts for a business process. RTO is the deadline for the IT resource to be working again. Because reprocessing and verification sit between the system coming back and the business coming back, the RTO has to be shorter than the MTD. NIST SP 800-34 Rev. 1 states that relationship explicitly.

Who sets the RTO and the RPO?

The business sets the tolerance, and IT translates that tolerance into architecture and cost. In NIST's definition, the MTD is the time the system owner accepts, which implies a named person rather than a meeting consensus.

Can the RPO be shorter than the backup interval?

No. The RPO is capped by the most recent copy available. If the copy runs nightly, the real RPO is up to 24 hours, whatever the document says. Shortening the RPO means changing how often or how the copies are made.

How do you set an RTO when every department says the system is critical?

By building an impact scale with bands and values, the way NIST recommends, and asking about scenarios with a clock on them instead of criticality in the abstract. Once departments have to place their own processes on the same scale, the priorities appear.

Is a four-hour RTO compatible with 99.9% availability?

It depends on how many incidents a year the target assumes. A 99.9% target allows 8.76 hours of downtime a year, according to Microsoft's service-level table. A single four-hour incident uses up nearly half of that budget.

Facebook
Twitter
LinkedIn

Also check out

Request a quote