BROKABROKA
Sign inDownload CommunityRequest a demo
Documentation105 pagesOlder version — go to 1.0.1
Guides

Alerts

Updated 3 October 2026 · applies to 1.0.0 · Community and Commercial

A rule watches one figure on one part of your estate and opens an incident when it stops being true. BROKA evaluates those rules on a schedule, across every platform it serves, and tells whoever you have told it to tell.

Alerting is in both editions, with no quota. Two parts of it are Commercial: rules scoped to an application, and the Slack, Microsoft Teams and webhook channels. E-mail and the console are in both.

The four tabs

Open comes first, and that is the whole argument for the screen's shape: the rules are what you came here to change, the incidents are what you came here to see. A monitoring screen that opens on configuration makes you click to find out whether anything is broken.

Rules is where a rule is written, edited, enabled, disabled or removed. Disabling or removing a rule closes its open incidents, marked as closed by you and why, since nothing is watching them any more. So does moving a rule to watch something else — another cluster, environment, application or resource, or another metric — for the incidents it leaves behind. Removing a rule also withdraws the silences that cover only that rule and have not ended.

History is every event the evaluator has opened and closed, so "did this happen before, and how often" has an answer.

Silences are maintenance windows: a planned upgrade should not page anybody, and suppressing the noise should not mean turning the rule off and forgetting to turn it back on. A silence covers everything, an environment, a connection, an application (Commercial), a resource or one rule, for a window with an end. A breach inside it is still recorded — as suppressed — and notifies nobody, not even with a reminder. If it is still breaching when the window ends, it is notified then.

Operations - Alerts on its Open tab, with Rules, History and Silences beside it: twelve incidents across three clusters, each with its severity, what broke, the rule and the metric behind it, the reading that breached, and whether it is firing or has been acknowledged.
Acknowledging an incident does not change the metric - it records that somebody has seen it, which is why an acknowledged row stays in the list.

What a rule says

A rule is six decisions, and the dialog asks them in the order they matter.

The metric, chosen from a catalogue rather than typed. Consumer lag, under-replicated partitions, queue depth, a queue with no consumer, the age of a queue's oldest message (Artemis), memory, connection health — each platform contributes the figures it actually has, so the list you see depends on what your installation serves.

The scope: an environment, a connection, an application (Commercial), or a named resource. A rule that names nothing watches everything the scope covers, and the number of targets that produces is a number worth knowing before you save — which is what the dry run below is for. A named RabbitMQ queue is looked up by its name however many queues the broker holds, and each evaluation reads a cluster's figures once for all of its rules.

A rule over an environment or an application measures each cluster on its own. One cluster's incident never resolves another's, and a hold window on one cluster is not reset by a sibling that is healthy — the same resource name on two clusters is two targets, not one. A cluster that cannot be read is named in the rule's state ("1 of 2 clusters could not be evaluated. orders-eu: …"), and the rule reads healthy only when every cluster is.

The comparison and the threshold — above, at or above, below, at or below, equal to.

How long it must hold. A figure that crosses a line for one tick and comes back is not an incident. Leave it at zero to fire on the first breach.

The severity — critical or warning.

Whether to repeat. An incident that stays open can page again after an interval, or stay quiet until it resolves.

The dry run

Before a rule is saved, Preview measures the draft against your live estate: which targets it matches, what each one reads right now, and how many of them would fire. It writes nothing and creates no rule.

Two things about it are deliberate. The counts describe every target even though the list of rows is capped — a truncated list whose total matched its own length would understate the blast radius of the rule being written, which is the one number the author is there to learn. And a target whose figure cannot be read is reported as unmeasurable, by name, rather than quietly left out.

A rule nobody has evaluated yet says so

A saved rule that has not been evaluated reads not evaluated yet — not ok. An unknown state that defaults to green is how a screen ends up reporting a quiet, healthy estate on a product that has measured nothing.

The same rule applies once the evaluator is running. A rule whose targets cannot be measured says how many — "3 of 47 targets could not be measured" — rather than reporting on the forty-four it could reach and calling that the answer.

A rule scoped to an application, kept from a Commercial installation that was moved back to Community, reads not evaluated: Community does not evaluate it, and its hover names the edition that does.

The Rules tab: fifteen rules across the estate with their scope, metric, comparison and severity, and a state on each - ok where the evaluator reached it, unevaluatable where it could not, and disabled where somebody switched it off.
A rule that could not be measured carries the broker's own reason, so the answer to why is on the row rather than in a log.

What the evaluator does, and what it does not keep

It runs on a schedule, reads each rule's metric live on its targets, and opens or resolves events accordingly.

It keeps no metric history. BROKA stores no time series, so there is no chart of a figure over time behind an alert — what it stores is the event: when it opened, on what, against which rule, and when it closed. The hold window for a rule that must breach for a while before it fires is held in memory, so a restart of the broker service resets it; that is stated here rather than discovered.

The History tab: every event the evaluator has opened, with the rule that raised it, the target, the reading, and whether it was acknowledged or resolved and when.
This is the whole of what an alert leaves behind. There is no chart of the metric beside it, because none is stored.

Where an alert goes

Slack, Microsoft Teams and webhook channels — Commercial only. E-mail and the console are in both editions.

Notification channels are configured in Settings ▸ Integrations, and there are four kinds: Slack, Microsoft Teams, e-mail through the installation's own SMTP relay, and a generic webhook, which posts BROKA's own document to any endpoint you name. A delivery that fails is tried again a few minutes later, up to five times, and a channel that stops answering is given up on after a minute without holding back anybody else's alerts.

The console is always a destination and needs no channel: an open incident reaches the bell in the top bar and the dashboard's alert card whether or not anything is configured.

What leaves your network when a channel fires. A webhook carries the severity, the rule, the metric, the comparison and threshold, the connection's name and the names of the targets that breached, plus a link back to your own console. That is the minimum a receiver can route on — and it is worth knowing that resource names travel, because on some estates a topic name is itself information.

Starting from something

A fresh installation is offered recommended starter rules for each platform it serves, enabled in one click. A product whose alerting starts as an empty list is a product whose alerting stays an empty list.

The permission

Reading this screen needs alerts.view, which is in the standard bundle. Incidents, and a rule's preview, show only the clusters your role lets you see. The rules you see follow what each one watches: a rule on a cluster or a resource, when you may see that cluster; a rule on an environment, when you hold a role in that environment or a global one; a rule on an application, when you may see the applications. The same decides what you may change: a rule you cannot see cannot be edited or deleted by you, and a rule cannot be created or moved to watch something you cannot see. Acknowledge and Resolve on an incident need alerts.respond — the permission for an on-call rota that should handle incidents without being able to change the rules. Writing rules and silences needs alerts.manage. Configuring the channels they dispatch to is a different right again — it lives with settings, because a channel holds a credential.

← PreviousIncident investigationNext →Search