Skip to contentExploitQuest

Lesson 1 of 1 in Restore, Then Watch

The Untested Backup

Two things you only find out you got wrong on the worst day: that your backup does not restore, and that nobody was watching when it started to go wrong. Both are fixed the same way — by rehearsing the bad day on a day it has not come.

3 min read

Not yet reviewed

Wrenlearner

I have a nightly backup running. So I am covered if the disk dies.

Rookmentor

You have a nightly *copy* running. Whether it is a backup depends on something you have not done: restored from it. Half the backups in the world are hopes wearing a backup's clothes, and their owners find out at the exact moment they cannot afford to.

The restore is the exercise

A backup is not the copy — it is the *restore*. A copy that has never been restored might be encrypted with a key you lost, missing the one directory that mattered, corrupt in a way nothing noticed, or of a database that was mid-write when it was taken. You do not know, and "I have a backup" is a belief until you have stood up the whole thing from it, on a spare box, start to finish. The rehearsal is the control; the copy is just its input.

# a backup you can trust has all three, tested:
offsite   -> a copy not on the same box, or the same fire
restored  -> you have actually rebuilt from it, recently
dated     -> you know how old the newest good one is
  1. Line 2A backup on the same disk survives a mistake and not a disk failure; on the same machine, not a ransomware event that encrypts everything mounted. Offsite is what makes it a backup against the real disasters.
  2. Line 3Test on a schedule, not once. A restore that worked last year against last year's software is not evidence about today's.

Watching, so a customer is not your alarm

The other worst-day surprise is finding out from a customer. A box nobody watches fails silently: the disk fills, a service dies, the certificate expires, an attacker settles in — and the first you hear is a support ticket. Watching does not mean staring at graphs; it means a handful of things that shout when they go wrong: disk nearly full, a service down, a spike in errors, an unexpected reboot.

You cannot watch everything. What is the difference between a useful alert and noise you will learn to ignore?

Your turn

Add a disk-space check to this monitoring config: append a line so it contains disk_over_90.

you@practice
Practice shell — nothing here is real. Type 'help' to begin.