Ops checks and alert analysis

From manual website checks to AI-run ops

Dutifly turns repeated website work into reusable AI workflows.

It checks monitoring, logs, and cloud resources, connects clues, analyzes alerts, and returns conclusions.

Your fixes become a personal ops Skill. Next time, AI starts from your experience.

Reusable workflows
Daily ops routines
Evidence reports
Metrics, logs, traces
Personal ops Skill
Troubleshooting rules

SCENARIO

What you repeat every day

Daily health check

Open Prometheus, Grafana, SLS, and cloud consoles one by one, then correlate the signals manually.

09:00

Daily health check

Open Prometheus, Grafana, SLS, and cloud consoles one by one, then correlate the signals manually.

Alert storm triage

Thirty alerts arrive at once. Most are noise, and the one that matters is buried in the stream.

10:30

Alert storm triage

Thirty alerts arrive at once. Most are noise, and the one that matters is buried in the stream.

Urgent troubleshooting

Jump between Kibana, Jaeger, and Grafana while context keeps dropping between screens.

14:20

Urgent troubleshooting

Jump between Kibana, Jaeger, and Grafana while context keeps dropping between screens.

WORKFLOW

Give the tedious and complex parts to AI

You only need to describe what to check. Dutifly breaks the scenario into steps, gathers data across platforms, and assembles the evidence chain. Risky actions such as rollback, scaling, or restart still require human confirmation.

Human-authorized

One-time authorization

Grant read-only access once and define which monitoring websites, logs, and cloud resources may be accessed.

Grafana / Prometheus / SLS

AI-run

Intent parsing

Describe the task in natural language; the AI turns it into an investigation path.

AI-run

Data aggregation

Collect and connect clues across monitoring, log, and trace systems.

AI-run

Root-cause analysis

Produce an evidence-backed assessment and mark uncertainty.

AI-run

Preference learning

Remember your troubleshooting preferences and turn your corrections into reusable Skills.

Dutifly · Ops workspace
Human reviewPayment Service shows P99 latency instability between 04:32 and 04:41. Peak latency reached 2,340 ms, with slow queries concentrated on payment-db-replica-2.
I will correlate SLS error logs, Prometheus metrics, and Jaeger traces, then rule out release changes and external dependencies.
Scheduled jobs usually run at that time. Next time, exclude scheduled-job traffic before you analyze it.
Recorded and rerun
For the 04:00-05:00 window, source=scheduler traffic will be filtered automatically. After filtering, the user request path is healthy.
Ask Dutifly about this check...

CAPABILITIES

Full-scenario AIOps coverage

More than a dashboard. An ops assistant that understands system signals.

Daily health-check reports

Generate cross-platform health summaries every day, with source evidence preserved.

Alert grading and noise reduction

Use historical patterns to separate false positives, duplicates, and genuine risk.

Root-cause localization

Connect metrics, logs, and distributed traces into a verifiable investigation path.

Capacity trend forecasting

Spot CPU, memory, GC, and disk-watermark trends before they become capacity risks.

Multi-cloud resource checks

Review AliCloud, AWS, and GCP resources in one place, including instances and utilization.

Change impact assessment

Compare key indicators and dependency impact before and after a release so you can review changes quickly.

For this service, filter scheduler traffic during overnight batch windows
When payment-path P99 exceeds 800 ms, inspect the connection pool first

Turn every troubleshooting pass into a personal ops Skill

Dutifly is not limited to one-off analysis. It remembers your service rules, troubleshooting experience, and preferences. Every correction, exception rule, and conclusion becomes part of your personal ops Skill library so the next health check or troubleshooting task can reuse what you already know.

Personal ops Skill libraryReusable experience

SUPPORTED ECOSYSTEM

Connect directly to existing monitoring, logging, tracing, and alerting tools

  • Prometheus
  • Grafana
  • AliCloud SLS
  • Elasticsearch
  • Zabbix
  • Jaeger
  • PagerDuty
  • Datadog
  • Tencent CLS

FAQ

Frequently asked questions

Start with a low-risk health-check workflow, confirm permissions, evidence paths, and human approval, then expand automation step by step.

Will Dutifly change production systems directly?

No. In ops scenarios, Dutifly starts with read-only analysis, evidence organization, and recommendations. High-risk actions such as restarts, scaling, configuration changes, or rollbacks require your confirmation.

Do I need to connect every monitoring platform at once?

No. It is better to start with one service, one alert workflow, or one routine health check. Once metrics, logs, and troubleshooting experience are working together, you can expand to more platforms.

What ops work is Dutifly good at?

Routine health checks, alert noise reduction, anomaly attribution, cross-platform information summaries, troubleshooting report generation, and preserving your own ops experience. Outputs stay reviewable, and critical judgment remains with you.

How are data permissions controlled?

Dutifly reads only within the scope you authorize. You can limit platforms, services, resources, and task boundaries. Data outside that scope is not accessed proactively.

Start with one read-only health check and bring Dutifly into your existing ops routine

Pick one service, one alert path, or one routine health check as the pilot. Dutifly brings together monitoring, logs, and traces without rebuilding systems or disrupting your existing tool stack.

Connect Grafana and generate your first health-check report