Unexplained Latency Spikes? Blame DNS
Diagnosing Intermittent API Slowdowns
When network services suddenly start seeing periodic latency jumps of 5 to 10 times their normal levels for a few hours at a time, the usual suspects come up empty. Syslog stays quiet, connection tables look healthy, and TCP/IP stats show no problems at the network layer. CPU, memory, and disk all appear fine, and no correlated load shows up on other hosts.

In this case, the culprit turned out to be Bind getting overwhelmed by heavy DNS demands, which in turn slowed down local name resolution for the entire service. The immediate workaround is to push everything into /etc/hosts, but the longer-term plan is to install a local bind9 instance on each host to serve as a cache.
If you have experience running production DNS resolver caching, your input would be appreciated.



