Feature Description
It would be useful to expose clearer visibility into when PRS and ERS processes are running.
During a PRS/ERS, a number of metrics can fluctuate as an expected consequence of the process, for example: replica lag, vReplication lag, and latency introduced by the PRS process itself. Today, it can be difficult to distinguish these expected changes from an unexpected issue when looking at metrics in isolation.
Use Case(s)
One potential improvement would be to expose PRS/ERS activity and timing as a first-class signal that can be correlated with other metrics. This could make it much easier to answer questions like "why did this metric spike at this time?" without requiring additional investigation.
Adding explicit metrics or events for PRS/ERS start, completion, and duration could provide a useful observability primitive and make these processes easier to understand and troubleshoot.
Feature Description
It would be useful to expose clearer visibility into when PRS and ERS processes are running.
During a PRS/ERS, a number of metrics can fluctuate as an expected consequence of the process, for example: replica lag, vReplication lag, and latency introduced by the PRS process itself. Today, it can be difficult to distinguish these expected changes from an unexpected issue when looking at metrics in isolation.
Use Case(s)
One potential improvement would be to expose PRS/ERS activity and timing as a first-class signal that can be correlated with other metrics. This could make it much easier to answer questions like "why did this metric spike at this time?" without requiring additional investigation.
Adding explicit metrics or events for PRS/ERS start, completion, and duration could provide a useful observability primitive and make these processes easier to understand and troubleshoot.