Better prometheus metrics for telepresence

Author: Yuheng3107Created Jul 21, 2025Updated Nov 19, 2025
Labelsfeature

Please describe your use case / problem. Need better signals to know when telepresence breaks and why telepresence is breaking/slow

Describe the solution you'd like A clear and concise description of what you want to happen. Add prometheus metrics such as error_count (count of all error logs/tunnel connection failed) Could also consider metrics regarding the cpu/mem usage (to identify bottlenecks in saturation) Or metrics regarding latency of connections (client to traffic-manager, traffic-manager to agent etc)

Describe alternatives you've considered manually getting logs and filtering for errors Versions (please complete the following information)

  • Output of telepresence version (just in case the feature exists in future versions)
  • Kubernetes Environment and Version

Additional context Add any other context about the feature request here.

Source: telepresenceio/telepresence