Browse and search documentation

Observability and on-call response

Connect existing monitoring and logging, review events, notifications and acknowledgements, and use the timeline for investigation.

For: Monitoring administrators and on-call engineersReviewed:
On this page

Connect a data source

  1. Select the source type in observability data-source management and enter its actual address. Browser reachability does not prove the platform service can reach it.
  2. Select no authentication, Bearer, Basic or mTLS to match the actual source. Prometheus without authentication does not require creating a token just to register it.
  3. Submit required authentication through the dedicated form, validate connectivity and review the result before use. Revalidate after credentials or certificates change.

Acknowledge and handle events

  1. Locate the event by time and state and verify its source and affected resources.
  2. Review alert state and acknowledgement history, determine whether someone is already handling it, then acknowledge with appropriate access.
  3. Investigate related assets, inspections and changes. If a data domain is unavailable, record that gap rather than treating the page as having no events.
  4. Review notification and escalation records. Confirm recovery from actual alert or check results rather than treating acknowledgement as resolution.

Notification states mean different things

StateMeaning
QueuedA delivery task was recorded; external delivery is not yet established
Delivered externallyThe channel accepted delivery, not proof a person saw it
Acknowledged in the platformSomeone acknowledged in ChronoOps, not necessarily in an external system
Suppressed retrying or failedCheck silences, maintenance windows, channel configuration and failure reason

Before enabling on-call response

  • Check roster timezone, participants, overrides and escalation order, then validate with an authorized test event.
  • Confirm actual delivery and escalation execution are enabled. Saving a receiver or policy does not mean real notifications are being sent.
  • Verify escalation stops as expected after acknowledgement and retain platform and channel records. One demo does not validate production on-call coverage.

Use the timeline

Use a consistent resource and time range to compare alerts, jobs, releases and actions. Temporal proximity is a clue, not proof. Preserve missing records, inaccessible sources and unresolved results. SLO configuration and query previews do not imply a complete native continuous evaluation system.

Something differs from your environment?

Send your deployment version, page and a redacted description so we can investigate and update the guide.

weiwendi@aiops.red