Better prometheus metrics for telepresence
Please describe your use case / problem. Need better signals to know when telepresence breaks and why telepresence is breaking/slow
Describe the solution you'd like A clear and concise description of what you want to happen. Add prometheus metrics such as error_count (count of all error logs/tunnel connection failed) Could also consider metrics regarding the cpu/mem usage (to identify bottlenecks in saturation) Or metrics regarding latency of connections (client to traffic-manager, traffic-manager to agent etc)
Describe alternatives you've considered manually getting logs and filtering for errors Versions (please complete the following information)
- Output of
telepresence version(just in case the feature exists in future versions) - Kubernetes Environment and Version
Additional context Add any other context about the feature request here.
Source: telepresenceio/telepresence