Best Practices for Monitoring Scenarios

Last Updated At: 2025-10-21 09:10:00

Cloud Monitor provides multiple methods to help users identify resource exceptions and send exception information to users through various channels as quickly as possible.

  1. Locating exceptions
    1. Detecting exceptions via monitoring and alarms
      Monitoring and alarms enable the cloud platform to detect exceptions promptly and send proactive notifications, allowing users to identify exceptions passively. This ensures timely awareness of exceptions under all circumstances. Users can log in to the Cloud Platform console and go to the Cloud Monitor console to configure appropriate alarm policies for the resources they care about. For details, see Creating an Alarm Policy.
      Key performance metrics and events configured as alarm rules will trigger multiple notification methods to promptly reach users and their systems in the event of an exception.
      Alarm policies configured with designated recipients or recipient groups will notify users promptly via SMS, email, and other channels. Features such as repeated alarms and alarm convergence are supported to help ensure critical alarms are not missed while minimizing unnecessary disturbances.
      Users can also configure the callback API feature in alarm notifications to deliver exception alarm information to their systems for further aggregation and processing.
    2. Detecting exceptions via the monitoring dashboard
      Locating exceptions via the monitoring dashboard is a user-driven approach that relies on analyzing average performance trends and historical data. It requires users to proactively identify exceptions. For exceptions that are not covered by alarm policies or are difficult to detect through existing alarm policies, they can still be identified during daily inspections. Compared with alarms, the monitoring dashboard enables users to assess the overall impact scope of resource exceptions. By subscribing key resources to the dashboard and configuring charts appropriately, users can highlight exception information under various scenarios. For details, see Create New Dashboard.
      For individual instances, users can subscribe to detailed instance views to easily compare performance trends on the dashboard panel.
      For resource clusters, users can subscribe to aggregated data within the same cluster to conveniently view the overall monitoring dashboard on the dashboard panel and compare it with the trend of individual instances within the cluster. For details, see Monitoring Scenarios for Large Volumes.
      Exceptions identified through the monitoring dashboard can be further analyzed using the dashboard's sorting and ranking features, enabling users to locate the affected resources and assess the scope of impact for in-depth troubleshooting.
  2. Exception troubleshooting
    Locating exception objects via the monitoring overview page
    During daily inspections or upon receiving alarm notifications, users can log in to the Cloud Platform console and choose Cloud Monitor > Monitoring Overview.
    1. Check the Service Status in Last 24 Hours module to view exception statuses of resources across different regions.
      Use the Exception Information Overview feature to get a preliminary view of recent exceptions.
    2. Click the number of exception objects to be redirected to the specific exception product menu page under Cloud Product Monitoring.
      The specific product list page under Cloud Product Monitoring will automatically filter and display the exception resource objects for the user.
    3. Click the ID of a specific object to be redirected to its monitoring details page, which provides historical status rollback to assist in detailed exception troubleshooting.
    • The exception timeline provides users with access to both current and historical information about the exception object. By reviewing historical alarms and status changes, users can be assisted in troubleshooting the current exception.
    • The monitoring data for resource performance provides users with the most comprehensive resource performance data. It allows for week-over-week and day-over-day comparisons of current and historical data for the same metric and facilitates troubleshooting by comparing data changes in different metrics over the same time period.
  3. Locating exception objects via the dashboard
    Log in to the Cloud Platform console and choose Cloud Monitor > Dashboard > Dashboard List to go to the corresponding dashboard page.
    1. When an abnormal trend appears in the monitoring chart, click Expand the Instance List, and then a sorted list of corresponding instances will be displayed below the curved line graph. Through the list, the specific objects that are causing the exceptions can be located.
    2. By clicking the object name in the sorted list, users can be redirected to the monitoring details page of the object, where historical status information is provided to assist users in locating exceptions.
    • The exception timeline provides users with access to both current and historical information of the exception object. By reviewing past alarms and status changes, users can be assisted in troubleshooting the current exception.
    • The monitoring data for resource performance provides users with the most comprehensive resource performance data. It allows for week-over-week and day-over-day comparisons of current and historical data for the same metric and facilitates troubleshooting by comparing data changes in different metrics over the same time period.