Guides for operators

Hardware

Every server's sensors, disks and management controller, the thresholds they are judged by, and what Velrix does when a server is about to fail.

Manager › Hardware shows every server’s health.

What is read

  • Sensors: temperatures, fans, voltages, current and power, read by the server itself every 30 seconds.
  • Disks: S.M.A.R.T. and NVMe health of every disk, read every hour by the device admin.
  • Management controller: connect a server’s BMC (Redfish: iDRAC, iLO, XCC and others) under Controllers to add its sensors and its hardware event log. Give it a read-only account; if the controller has a self-signed certificate, paste it into the connection.

Thresholds

A sensor’s own limits come first. Where it has none, the numbers on the Thresholds tab apply. Change one, turn it off, or put it back to the default there. A sensor that reads nonsense (an input with nothing connected) can be ignored on the server’s Health tab; it is then left out of everything, the charts included.

When something crosses a threshold

Platform operators get a notice: once when it happens, once when it gets worse, and once when the server is healthy again. Critical ones are also mailed.

A disk that predicts its own failure, or a server that overheats, is marked: nothing new is placed on it, and what can move is moved to other servers. VMs and apps with a volume stay and wait for their owners to start them elsewhere. The mark stays until you clear it on the server’s page, after the disk is replaced or the cooling fixed. It is not set again for the same problem, only for a new one.

For compliance

Every disk’s verdict is evidence for ISO 27001 A.7.13 (equipment maintenance) on the Compliance page.