Disk and storage

File systems
/ (/dev/sda1) 50 GB / 80 GB · 63%
/home (/dev/sdb1) 76 GB / 200 GB · 38%
/boot/efi (/dev/sda15) 317 MB / 1 GB · 31%
Drives
Device Size Model Health (SMART) Temp. Power-on hours Realloc sectors
sda SSD 80G QEMU HARDDISK OK 34°C 14,249 0
sdb SSD 200G Samsung SSD 870 OK 31°C 9,221 0

Disk health, free space and the inode trap

Two disk problems take servers down, and they look nothing alike. The slow one is a drive beginning to fail, which SMART usually sees coming. The fast one is a filesystem filling up, which takes services with it — a database that cannot write is a database that stops.

This page covers both: usage per filesystem, SMART attributes per drive, and drive temperature. Among the SMART attributes, the ones that genuinely predict failure are reallocated sector count, current pending sector count and offline uncorrectable — any of those moving away from zero means order a replacement rather than wait for confirmation. Most other attributes fluctuate harmlessly.

Inode exhaustion is shown separately because of how confusing it is when it happens. The disk reports free space, writes fail anyway, and nothing in the obvious places explains it. The cause is a directory holding hundreds of thousands of tiny files — session data, mail spool, an unrotated cache — using up the filesystem's fixed supply of inodes while barely touching its capacity.

There is a third disk trouble, one that breaks nothing and merely slows things down: the drive is busy enough that processes queue for it. In numbers that is I/O wait, and it rarely climbs on its own — something scheduled is usually behind it: the nightly backup, an index rebuild, a full ClamAV scan across the whole filesystem. If the server grows heavy at the same hour every time, look for a schedule rather than a fault.