Coroot Agent and metrics missing/gaps
Hello dear Coroot team,
We have been observing some strange behaviour with our current Coroot deployment. Some monitored nodes seems to have missing metrics/metric gaps at some time intervals (not the same time on affected servers).
At the "Nodes" section, I can see that all metrics displayed for a particular node (e.g CPU, Memory, Network, etc...) have all gaps in their graphs. Also at the "Nodes" section again, some nodes will display their status as "down" for a few minutes and automatically go back to "up" without any manual intervention. This seems like metrics were gathered by the corresponding Coroot agent, but not uploaded or discarded (just an assumption here).
This is not a network issue as we monitor those servers (OS level) with other tools and a 10 second network issue triggers an alert, which never happens, leading me to believe this is not network-related.
I checked a few servers at the agent level (journalctl -xeu coroot-node-agent) and could not spot any obvious error message, apart from these, which I am not sure are related/cause of issues:
fd.go:38] failed to read link 'XXX': readlink /proc/XXXXXX/fd/XXX: no such file or directory [...] registry.go:346] failed to create container pid=XXXXXX
When checking the graph (section "Nodes") for the last 3 days, quite a few of our monitored servers display those gaps in these metric collections. Just checking here if there is any additional configuration at the Coroot level (or even at the OS) to make this more reliable.
Any suggestions?
Coroot version: 1.22.0 Agent version: 1.23.7
Thanks.
Source: coroot/coroot