Home server drops off the network but stays up? Check for an e1000e "Hardware Unit Hang"
My always-on home server, an old Dell Latitude with an Intel 82579LM NIC, went silent for 22 minutes. The machine itself was fine. A reboot brought the network back, but a reboot doesn't tell you why it went down.
The kernel log from the previous boot does:
sudo journalctl -k -b -1 | grep -c "Detected Hardware Unit Hang"
sudo journalctl -k -b -1 | grep "Hardware Unit Hang" | sed -n '1p;$p'
There were 678 hits, two seconds apart, from the moment it dropped until I rebooted. The TX descriptor ring was stuck (TDH never moved), and the driver never managed to reset the card. The NIC simply stopped transmitting.
This is a long-known e1000e erratum on 82579/82574-era Intel NICs. The usual trigger is TCP segmentation offload. The standard mitigation is to stop the card doing TSO/GSO and let the CPU do it, which costs almost nothing at home-network speeds:
sudo apt install ethtool
sudo ethtool -K eno1 tso off gso off
sudo ethtool -k eno1 | grep -E '^(tcp-segmentation|generic-segmentation)-offload'
ethtool -K doesn't survive a reboot. If NetworkManager owns the interface, you don't need a udev rule or a systemd unit: NM can apply ethtool features itself on every activation:
sudo nmcli con modify "Wired connection 1" ethtool.feature-tso off ethtool.feature-gso off
nmcli -g ethtool.feature-tso,ethtool.feature-gso con show "Wired connection 1"
Applying it live only blips the link for a second, so it's safe even when you're working over that same connection, or from a VM running on the box.
If the hang ever comes back with offload off, the next suspects are Energy Efficient Ethernet and PCIe ASPM (pcie_aspm=off on the kernel command line).