Cloud Services
Key Metrics for Evaluating Your Disaster Recovery Plan
Five practical numbers that tell you whether your disaster recovery plan will hold up when you actually need it.
The key metrics for evaluating a disaster recovery plan are Recovery Time Objective (RTO), Recovery Point Objective (RPO), Mean Time to Recovery (MTTR), test frequency, and cost of downtime. Together they tell you whether your plan will actually get the business trading again — and whether the recovery it delivers is one you can afford. This guide explains how to set and measure each one without drowning in jargon.
Why you need numbers, not just a plan document
Plenty of businesses have a disaster recovery plan sitting in a folder somewhere. Far fewer can answer the questions that actually matter: how long would it take to get trading again, and how much data would be gone when you did? A plan without measurable targets is a statement of intent, not a capability.
Metrics fix that. They turn vague assurances like "we back up regularly" into commitments you can test, and they give you an objective way to decide whether your current setup — backups, replication, failover, people — is good enough or needs investment.
Recovery Time Objective (RTO): how long you can afford to be down
RTO is the maximum time a system can be offline before the disruption becomes unacceptable. It is a business decision first and a technical one second: your accounting package might tolerate a day of downtime, while your order-processing system or phones might only tolerate an hour.
Set an RTO per system, not one number for the whole business. Rank your systems by how quickly their absence hurts, then attach a target to each. Be honest — "everything back in fifteen minutes" sounds reassuring but usually demands infrastructure most small businesses have not paid for.
Measuring RTO is straightforward: run a restore and time it end to end, including the steps people forget — locating backups, provisioning replacement hardware or cloud instances, reconfiguring DNS and reconnecting staff. If your measured recovery time exceeds your stated RTO, one of the two has to change.
Recovery Point Objective (RPO): how much data you can afford to lose
RPO is the maximum amount of data, measured in time, you can afford to lose — effectively the age of the oldest acceptable backup. An RPO of 24 hours means losing a full day of invoices, emails and job records is survivable. For many businesses it is not.
Your RPO is set by how often usable backups are taken. Nightly backups give you a 24-hour RPO at best; continuous replication to a second site or cloud platform can bring it down to minutes. Moving critical workloads and backups onto well-configured cloud services is usually the most economical way to shorten RPO without buying duplicate hardware.
One caution: RPO assumes the backup actually restores. A nightly backup that has been silently failing for three weeks gives you a three-week RPO, whatever the schedule says. Verified, tested restores are part of the metric, not an optional extra.
Mean Time to Recovery (MTTR): what your track record shows
RTO is what you promise; MTTR is what you deliver. It is the average time your team has actually taken to restore service across real incidents and drills. Tracking it means recording, for every outage and every test, when the incident started, when recovery began and when service was fully restored.
Compare MTTR against RTO system by system. A consistent gap tells you exactly where the plan is weakest — and often the weakness is not the technology but the process around it: nobody could find the credentials, the runbook was out of date, or the one person who knew the procedure was on leave. Businesses with managed IT support tend to close this gap faster, because recovery procedures are documented, rehearsed and not dependent on a single staff member.
Test frequency: an untested plan is a guess
How often you test is itself a metric worth tracking, because recovery capability decays quietly. Systems change, staff change, and a plan tested two years ago describes a business that no longer exists.
A practical cadence for most NZ businesses: confirm backups are completing weekly, with automated alerts that a person actually checks; run a restore test of critical data quarterly; and run a fuller recovery exercise — restoring key systems and having staff work from them — at least annually, plus after any major change such as a server migration or a new line-of-business application.
Record two things each time: whether the test met its RTO and RPO targets, and what failed. A test that uncovers problems is a success. The same problem appearing in consecutive tests is the real red flag.
Cost of downtime: the metric that justifies the others
Cost of downtime converts everything above into dollars, which is how recovery investment decisions should be made. Estimate it per hour for each critical system: lost revenue, wages paid to staff who cannot work, penalties or missed contractual deadlines, and the harder-to-count damage of frustrated customers.
You do not need a precise figure — a defensible range is enough. If an hour of downtime would cost you thousands, paying a modest monthly amount to cut your RTO from two days to two hours is easy arithmetic. If an hour costs you very little, a slower, cheaper recovery may be entirely rational. The point is to make that trade-off deliberately rather than discover it mid-crisis.
Remember, too, that most real-world "disasters" are mundane: hardware failure, human error and, increasingly, ransomware — which is why disaster recovery and cyber security should be assessed together rather than as separate exercises.
Turning the metrics into a working plan
Put the five metrics on a single page: each critical system, its RTO and RPO targets, the last measured recovery time, the last test date and result, and the estimated hourly cost of that system being down. Review it twice a year and after every significant IT change. That one page tells you — and your insurer, board or bank, if they ask — whether your disaster recovery plan is real.
Call 0800 900 777 or reach us through our contact page if you would like yours reviewed.
Related Service
Cloud Services
Google Workspace, Microsoft 365, Azure and AWS — planned, migrated and managed by a local Auckland team.
FAQs
Frequently asked questions
What is the difference between RTO and RPO?
RTO (Recovery Time Objective) is how long a system can be down before the impact becomes unacceptable. RPO (Recovery Point Objective) is how much data, measured in time, you can afford to lose. A four-hour RTO with a 24-hour RPO means being back online within four hours, but potentially missing up to a day of data.
How often should we test our disaster recovery plan?
Confirm backups are completing weekly, run a restore test of critical data quarterly, and carry out a fuller recovery exercise at least once a year — plus after any major change such as a server migration. Record whether each test met its RTO and RPO targets so you can see the trend over time.
How do I work out our cost of downtime?
Add up, per hour, the revenue you would lose, the wages paid to staff who cannot work, and any penalties or urgent recovery costs, then allow something for reputational damage. A defensible range is enough — the figure's job is to tell you how much recovery capability is worth paying for.
Can a small business achieve an RPO of minutes?
Yes, but usually only for specific systems rather than everything. Continuous replication of key data to a cloud platform can bring RPO down to minutes at a reasonable cost, while less critical systems stay on nightly backups. Tiering systems this way keeps spend proportionate to what each system is actually worth.
Keep Reading
More guides
Need help with cloud services?
Book a free, no-obligation consultation — we'll review your setup and give you clear, practical recommendations.