Browse and search documentation
Observability and on-call response
Connect existing monitoring and logging, review events, notifications and acknowledgements, and use the timeline for investigation.
On this page
Connect a data source
- Select the source type in observability data-source management and enter its actual address. Browser reachability does not prove the platform service can reach it.
- Select no authentication, Bearer, Basic or mTLS to match the actual source. Prometheus without authentication does not require creating a token just to register it.
- Submit required authentication through the dedicated form, validate connectivity and review the result before use. Revalidate after credentials or certificates change.
Acknowledge and handle events
- Locate the event by time and state and verify its source and affected resources.
- Review alert state and acknowledgement history, determine whether someone is already handling it, then acknowledge with appropriate access.
- Investigate related assets, inspections and changes. If a data domain is unavailable, record that gap rather than treating the page as having no events.
- Review notification and escalation records. Confirm recovery from actual alert or check results rather than treating acknowledgement as resolution.
Notification states mean different things
| State | Meaning |
|---|---|
| Queued | A delivery task was recorded; external delivery is not yet established |
| Delivered externally | The channel accepted delivery, not proof a person saw it |
| Acknowledged in the platform | Someone acknowledged in ChronoOps, not necessarily in an external system |
| Suppressed retrying or failed | Check silences, maintenance windows, channel configuration and failure reason |
Before enabling on-call response
- Check roster timezone, participants, overrides and escalation order, then validate with an authorized test event.
- Confirm actual delivery and escalation execution are enabled. Saving a receiver or policy does not mean real notifications are being sent.
- Verify escalation stops as expected after acknowledgement and retain platform and channel records. One demo does not validate production on-call coverage.
Use the timeline
Use a consistent resource and time range to compare alerts, jobs, releases and actions. Temporal proximity is a clue, not proof. Preserve missing records, inaccessible sources and unresolved results. SLO configuration and query previews do not imply a complete native continuous evaluation system.
Something differs from your environment?
Send your deployment version, page and a redacted description so we can investigate and update the guide.
weiwendi@aiops.red