Skip to main content
SafeCortex

Operação e suporte

The alert nobody answers

When everything raises an alert, nothing triggers action. How a dashboard full of warnings ends up less useful than no monitoring at all.

There is a predictable stage in the life of a monitoring setup. It starts useful, gains a rule after every incident and, a few months later, fires so often that the team creates a mail folder to keep the warnings out of the way.

From that point on, the monitoring protects nothing. It only produces noise — and the real incident arrives in the middle of three hundred notices nobody reads.

The original mistake is almost always the same: treating an alert as a log entry. A log is for looking up later; an alert is for interrupting someone now. They are different things and should not share a channel.

An alert that holds up answers three questions before it exists:

  • What exactly is wrong, in terms the recipient understands without opening a console?
  • Who needs to know, by name, and through which channel?
  • What should that person do on receiving it?

If the answer to the third question is "look at it and see whether it gets worse", it is not an alert. It is information, and information belongs on the dashboard and in the report.

Criticality has to be real too. It is worth saying out loud, when defining it, what each level means in human consequence: critical means waking someone at three in the morning. If the team is not willing to do that for a given event, it is not critical — and classifying it that way only teaches everyone to ignore the ones that are.

A few corrections that usually give the dashboard its usefulness back:

  • Alert on symptoms the user experiences, not on isolated metrics. "Checkout is down" is worth more than "CPU went over 80%".
  • Group events with the same cause: a database going down should not produce forty notices, one per dependent service.
  • Wait before firing. Plenty of things resolve themselves in thirty seconds, and alerting on those teaches the team to wait it out.
  • Periodically review what fired and led to no action at all. An alert that never becomes an action is a candidate for becoming a log entry.
  • Write the expected first step into the alert itself.

It is also worth measuring what the dashboard does not show: how many alerts were closed with no action taken. That is the number that reveals fatigue before the team complains — and before the incident that mattered slips past unnoticed.

Automating the fix is possible, but with a clear limit: only where the behaviour is known and the action is safe to repeat. Restarting a service that hangs predictably is reasonable; automating on top of a problem nobody has understood yet merely hides the symptom and postpones the diagnosis.

Good monitoring is not the one that sees the most. It is the one that interrupts the fewest people, and always for a good reason.

Related content

O alerta que ninguém atende

Quando tudo dispara alerta, nada dispara ação. Como um painel cheio de avisos acaba sendo menos útil do que nenhum monitoramento.

· 3 min read

Shall we talk about what you need?

Describe the scenario and we come back with the possible options, what needs assessing and how the work could be run.

Chat on WhatsApp