| Process | CPU % | RAM % |
|---|---|---|
| mysqld | 6.1 | 4.0 |
| php-fpm | 4.1 | 3.0 |
| apache2 | 2.1 | 5.1 |
| suricata | 1.4 | 10.8 |
| fail2ban-server | 0.9 | 10.7 |
| Process | CPU % | RAM % |
|---|---|---|
| mysqld | 19.5 | 15.0 |
| fail2ban-server | 11.3 | 4.0 |
| suricata | 7.7 | 9.2 |
| apache2 | 5.3 | 4.5 |
| clamd | 2.4 | 3.9 |
top answers a question about the current second, while the question people actually ask is about last night: why everything stalled at three in the morning and recovered on its own. By then there is nothing left to look at, because nothing was ever written down. Standing up Prometheus and Grafana for a single VPS is its own kind of problem: the observing stack outweighs the server being observed.
A collector appends one row to the database every five minutes, and this page draws the last day from it. Load average, CPU usage with I/O wait kept separate, RAM and swap, network in and out, disk reads and writes, established TCP connections, processes, partition usage, inodes and file descriptors as a percentage of their limit, MySQL connections. All on one page and on one time axis, so coincidences are visible by eye.
These graphs are worth reading in pairs. High I/O wait with a calm CPU means the server is waiting on the disk, and adding cores will not help. Swap that rose once and never came back to zero says the memory peak has already happened, however quiet things look now. File descriptors approaching a hundred percent are a too many open files error some hours before it happens. And on most servers inodes run out well before free space does.