Backup Monitoring for MSPs: Catching Silent Failures Before the Restore Request
By the Helios team
The most dangerous backup failure is not the job that reports an error. It is the job that stops reporting at all. Backup monitoring for MSPs usually means glancing at a dashboard of green ticks across three or four vendors, and a green tick tells you one thing only: the software completed its own steps. It does not tell you the data can be restored, that the offsite copy exists, or that the job has run at all in the last fortnight. This piece covers what to alert on beyond job success, how to spot stale jobs and shrinking datasets, a restore testing cadence a small team can actually keep, and how to report backup status to clients in words they understand. By the end you will have a weekly verification routine and an escalation rule for missed jobs, whether the machines belong to clients or to your own company.
The green tick measures the wrong thing
A successful job means the backup software believes it did what it was configured to do. Configuration is where the rot lives. A job that backs up an empty folder succeeds. A job whose selection list still points at a server decommissioned last year succeeds. A job that writes to a repository which has quietly started pruning your oldest restore points, because nobody watched capacity, succeeds right up until the day you need a version from three weeks ago.
The tick also says nothing about the second and third copies. If your architecture follows the 3-2-1 rule, each copy needs its own verification, because the offsite replication job failing is exactly as invisible as the local one succeeding is visible.
A backup is a claim. A restore is the evidence. Everything between the two is monitoring.
Backup monitoring for MSPs: five signals beyond job success
Stale jobs. Track the age of the last successful backup per protected item, not per job. A job can be disabled, paused after maintenance, or orphaned when a server is renamed, and no failure alert will ever fire because no failure ever occurs. The silence test: a backup that stops reporting looks identical to a backup that never existed. If your monitoring cannot tell those two apart, it is not monitoring, it is decoration.
Shrinking datasets. If a job that protected 400 GB last week protects 90 GB this week and still reports success, something changed: a volume was excluded, a share was moved, or the selection now points at a stub. Alert on significant drops in protected size or file count between runs. Sudden growth deserves a look too, because it usually means someone dumped data somewhere unplanned.
Duration drift. A nightly job that took twenty minutes in January and takes five hours in June is telling you about failing disks, saturated links or runaway change rates. It will collide with the working day eventually. Trend it now.
Repository capacity and retention. Most backup products handle a full repository by pruning the oldest restore points, silently. Your retention promise to the client shrinks without a single red icon. Alert on repository free space and on actual oldest-restore-point age against the promised retention.
Agent and heartbeat health. A dead backup agent, an expired service account password or a stopped scheduler service produces no job result at all. Monitor the agent as a service in its own right, not just the jobs it runs.
Rule of thumb: alert on absence, not just failure. "No successful backup for this item in 26 hours" catches every silent failure mode above. "Job reported an error" catches almost none of them.
Restore testing: the cadence that separates monitoring from hoping
Restore testing is where good intentions go to die, because it never feels urgent. Make it a scheduled ticket, not a resolution. A cadence a one-to-five-technician team can sustain:
Monthly, per client: a file-level restore. Pick a random file from a random protected machine, restore it to an alternate location, open it. Ten minutes. Record the date and the file. Skip this and your first restore test of the year happens during an actual incident, with the client watching.
Quarterly, per client: an image or VM boot test. Restore a server image or spin up a backup as a VM and confirm it reaches a login prompt. This is the test that catches corrupt chains, missing system state and drivers that will not boot on dissimilar hardware.
Annually: a full recovery exercise for any client paying for a recovery time objective. If you have promised a four-hour recovery, you need one dated occasion on which you achieved it. Otherwise the promise is a guess wearing a contract's clothes.
Write the result down every time: date, what was restored, how long it took. That log is the difference between "our backups are fine" and being able to prove it.
The weekly verification routine, in about twenty minutes
Review last-success age for every protected item across every vendor, sorted oldest first. Anything over your threshold gets a ticket, not a mental note.
Check dataset size trends for drops or spikes since last week. Investigate anything over roughly 20 per cent either way.
Check repository capacity and oldest restore point against promised retention for each client.
Confirm offsite and cloud copies are current, not just the local job. Verify the copy job, not the source job.
Run this month's restore test if one is due, and log it.
Reconcile the protected list against the estate. New server onboarded last month? Confirm it is actually in a job. This is the check that catches the machine everyone assumed someone else had added.
Do this as one owned, recurring ticket. Spread across the week as ambient dashboard-glancing, it degrades into alert noise you have trained yourself to ignore.
The escalation rule for missed jobs
Decide this before the miss, not during it. Ours is simple and we suggest you steal it:
The 1-2-3 rule: one missed backup is a ticket, investigated within one business day. Two consecutive misses on the same item is a same-day priority. Three consecutive misses, or any miss on a server or line-of-business system, is an incident: fix it that day and tell the client you found it. Never let them find it first.
The counterintuitive part is the third step. Telling a client their backup failed and you caught it builds more trust than silence ever will, because it is proof somebody is watching.
Reporting backup status in language clients understand
Clients do not care about job names, chains or synthetic fulls. They care about three questions: is my data protected, how far back can we go, and when did you last prove a restore works. Report exactly that:
Protected, at risk, or unprotected, per system, as a plain status. "At risk" means backups exist but something needs attention: a stale copy, tight retention, a failed test.
Oldest available restore point, stated as a date, compared against what was agreed.
Last verified restore, with the date. This line does more for client confidence than any success percentage, because it is the only line backed by evidence rather than software self-assessment.
Failure modes: how each one presents, and what catches it
| Failure mode | How it looks on the dashboard | What catches it |
|---|---|---|
| Disabled or orphaned job | No alerts, nothing red | Last-success age per protected item |
| Selection drift, excluded data | Green tick, job succeeds | Dataset size and file count trending |
| Silent retention pruning | Green tick, job succeeds | Repository capacity and oldest-point checks |
| Corrupt backup chain | Green tick until restore day | Scheduled restore and boot tests |
| Dead agent or scheduler | No job results at all | Heartbeat monitoring on the agent itself |
| New machine never added | Nothing, it does not exist | Weekly reconciliation against the estate |
Notice the pattern: the dashboard catches almost none of them. The routine catches all of them.
Where this fits with Helios
Most of this article is discipline, not tooling, and no platform will run your restore tests for you. What Helios does is make the routine cheap: it pulls backup status across your estate into one view alongside monitoring, patching and the service desk, alerts on staleness rather than only on reported failures, and turns misses into tickets with your escalation rule attached, so the 1-2-3 rule runs itself. The AI agent can investigate a failed job and draft the fix, subject to your approval guardrails. We will be honest about limits: Helios monitors backup health, it is not itself a backup product, and the verification log is still yours to keep.
Helios is a flat-rate RMM and PSA for small MSPs and internal IT teams: £99 to £399 a month for the whole team, every feature on every plan. 14-day trial, no card, no feature gating. Start free at heliosmsp.io.