Alert metrics
Updated 6 October 2026 · applies to 1.0.1 · Community and Commercial
Every figure an alert rule can watch, by platform: the name the rule dialog shows, the key it is stored under, its unit and what it lands on. A rule's metric is chosen from this list and nothing else — a key with no figure behind it would be a rule that saves and never fires.
What a metric lands on decides what a rule's scope expands to. A metric that lands on a kind of resource — a topic, a consumer group, a queue, a connector — measures each one in the scope, or the one a resource-scoped rule names. A metric that lands on the connection has one figure per cluster. One marked per node is measured on each node of the connection — each Kafka broker or controller, each RabbitMQ node, each Memcached node — and an incident names the node.
Three rules hold across the list. A figure BROKA could not read is unevaluatable, with the reason, never 0. A yes-or-no figure is a number, 1 or 0, so it takes the same comparisons as any other. And there are no raw counters: a threshold on a total that only ever grows would fire once and stay fired, so a counter appears as its rate or not at all.
The platforms other than Kafka are in Commercial; their metrics are offered only where the installation serves them.
Kafka
| Metric | Key | Unit | Lands on |
|---|---|---|---|
| Consumer group lag | kafka.group.lag |
messages | consumer group |
| Consumer group time lag | kafka.group.time_lag |
milliseconds | consumer group |
| Consume rate | kafka.group.consume_rate |
per second | consumer group |
| Consumer group members | kafka.group.members |
count | consumer group |
| Topic produce rate | kafka.topic.produce_rate |
per second | topic |
| Topic message count | kafka.topic.messages |
messages | topic |
| Topic size | kafka.topic.size_bytes |
bytes | topic |
| Time since last write | kafka.topic.last_write_age |
milliseconds | topic |
| Partition count | kafka.topic.partitions |
count | topic |
| Under-replicated partitions | kafka.cluster.urp |
partitions | connection |
| Cluster messages in | kafka.cluster.messages_in_rate |
per second | connection |
| Broker count | kafka.cluster.brokers |
count | connection |
| Cluster size | kafka.cluster.size_bytes |
bytes | connection |
Topic and group figures come from the snapshot BROKA refreshes about every thirty seconds; a snapshot older than three refreshes makes them unevaluatable, with its age.
Kafka Connect
| Metric | Key | Unit | Lands on |
|---|---|---|---|
| Connect cluster reachable | kafka.connect.cluster.reachable |
1 / 0 | Connect cluster |
| Connectors failed | kafka.connect.cluster.connectors_failed |
count | Connect cluster |
| Connector tasks failed | kafka.connect.cluster.tasks_failed |
count | Connect cluster |
| Connector failed | kafka.connect.connector.failed |
1 / 0 | connector |
| Connector paused (or stopped) | kafka.connect.connector.paused |
1 / 0 | connector |
| Connector unassigned | kafka.connect.connector.unassigned |
1 / 0 | connector |
| Connector failed tasks | kafka.connect.connector.failed_tasks |
count | connector |
| Connector tasks running | kafka.connect.connector.running_ratio |
ratio | connector |
A worker that does not answer reads reachable = 0, which a rule can fire on; the other figures are then unevaluatable. A connector is shown by its Connect cluster's name and its own.
Schema Registry
| Metric | Key | Unit | Lands on |
|---|---|---|---|
| Schema Registry reachable | kafka.schema_registry.reachable |
1 / 0 | connection |
| Schema Registry response time | kafka.schema_registry.response_ms |
milliseconds | connection |
| Schema Registry read-only | kafka.schema_registry.read_only |
1 / 0 | connection |
| Schema Registry subjects | kafka.schema_registry.subjects |
count | connection |
On a connection with no registry these are unevaluatable and say so. Read-only is the registry's mode READONLY or
IMPORT, in which a producer registering a new schema version fails.
Kafka over JMX
Read only where the connection's JMX is set up; without it, every one of them is unevaluatable with that reason. A node whose agent did not answer is unevaluatable on its own while the others are measured.
| Metric | Key | Unit | Lands on |
|---|---|---|---|
| Active controllers | kafka.cluster.active_controllers |
count | connection |
| Offline partitions | kafka.cluster.offline_partitions |
partitions | connection |
| Unclean leader elections | kafka.cluster.unclean_elections_rate |
per second | connection |
| Leader elections | kafka.cluster.leader_elections_rate |
per second | connection |
| Under-min-ISR partitions | kafka.broker.under_min_isr |
partitions | per node |
| At-min-ISR partitions | kafka.broker.at_min_isr |
partitions | per node |
| ISR shrinks | kafka.broker.isr_shrinks_rate |
per second | per node |
| Offline log directories | kafka.broker.offline_log_dirs |
count | per node |
| Broker not running | kafka.broker.not_running |
1 / 0 | per node |
| Metadata lag | kafka.node.metadata_lag_ms |
milliseconds | per node |
| Produce p99 | kafka.broker.produce_p99_ms |
milliseconds | per node |
| Consumer fetch p99 | kafka.broker.fetch_consumer_p99_ms |
milliseconds | per node |
| Follower fetch p99 | kafka.broker.fetch_follower_p99_ms |
milliseconds | per node |
| Produce mean time | kafka.broker.produce_mean_ms |
milliseconds | per node |
| Consumer fetch mean time | kafka.broker.fetch_consumer_mean_ms |
milliseconds | per node |
| Failed produce requests | kafka.broker.failed_produce_rate |
per second | per node |
| Failed fetch requests | kafka.broker.failed_fetch_rate |
per second | per node |
| Bytes rejected | kafka.broker.bytes_rejected_rate |
bytes per second | per node |
| Request handlers idle | kafka.broker.request_handler_idle_percent |
percent | per node |
| Network processors idle | kafka.broker.network_idle_percent |
percent | per node |
| Request queue | kafka.broker.request_queue |
count | per node |
| Response queue | kafka.broker.response_queue |
count | per node |
| Produce purgatory | kafka.broker.purgatory_produce |
count | per node |
| Fetch purgatory | kafka.broker.purgatory_fetch |
count | per node |
| Log flush p99 | kafka.broker.log_flush_p99_ms |
milliseconds | per node |
| Heap used | kafka.node.heap_used_percent |
percent | per node |
| GC time | kafka.node.gc_time_rate |
milliseconds per second | per node |
| Process CPU | kafka.node.cpu_percent |
percent | per node |
| File descriptors used | kafka.node.open_fds_percent |
percent | per node |
| Threads | kafka.node.threads |
count | per node |
| Broker bytes in | kafka.broker.bytes_in_rate |
bytes per second | per node |
| Broker bytes out | kafka.broker.bytes_out_rate |
bytes per second | per node |
| Broker messages in | kafka.broker.messages_in_rate |
per second | per node |
The four cluster figures are published only from a whole answer: active controllers needs every controller, the
other three the one node that says it is the active controller. Leader and unclean election rates are not published by
every Kafka — see JMX metrics. A consumer fetch waits up to fetch.max.wait.ms by design, so
set a fetch-time threshold above it. Broker messages in is per broker only; the cluster's figure is Cluster
messages in, from offsets. A topic's own JMX rates are shown on its page and are not alertable.
Redis
| Metric | Key | Unit | Lands on |
|---|---|---|---|
| Memory used | redis.memory.used_bytes |
bytes | per node |
| Memory used (of maxmemory) | redis.memory.used_percent |
percent | connection |
| Memory fragmentation ratio | redis.memory.fragmentation_ratio |
ratio | connection |
| Keyspace hit rate | redis.hit_rate |
ratio | connection |
| Operations per second | redis.ops_per_sec |
per second | per node |
| Connected clients | redis.clients.connected |
count | per node |
| Blocked clients | redis.clients.blocked |
count | connection |
| Key count | redis.keys |
count | connection |
| Replication lag | redis.replication.lag_bytes |
bytes | connection |
| Time since last replication I/O | redis.replication.last_io_age |
seconds | connection |
| Stream length | redis.stream.length |
entries | stream |
| Stream group lag | redis.stream.group_lag |
entries | stream |
| Pending entries | redis.stream.pending |
entries | stream |
On a Redis Cluster the figures a single node answers for are measured per node and name the shard; the ones that
describe one node rather than the deployment are unevaluatable there. Memory used (of maxmemory) is unevaluatable on
a Redis with no maxmemory.
RabbitMQ
| Metric | Key | Unit | Lands on |
|---|---|---|---|
| Queue depth | rabbitmq.queue.messages |
messages | queue |
| Ready messages | rabbitmq.queue.ready |
messages | queue |
| Unacknowledged messages | rabbitmq.queue.unacknowledged |
messages | queue |
| Queue consumers | rabbitmq.queue.consumers |
count | queue |
| Consumer utilisation | rabbitmq.queue.consumer_utilisation |
ratio | queue |
| Redelivery rate | rabbitmq.queue.redeliver_rate |
per second | queue |
| Queue memory | rabbitmq.queue.memory_bytes |
bytes | queue |
| Total messages | rabbitmq.cluster.messages |
messages | connection |
| Total unacknowledged | rabbitmq.cluster.unacknowledged |
messages | connection |
| Publish rate | rabbitmq.cluster.publish_rate |
per second | connection |
| Deliver rate | rabbitmq.cluster.deliver_rate |
per second | connection |
| Unroutable rate | rabbitmq.cluster.unroutable_rate |
per second | connection |
| Connections | rabbitmq.cluster.connections |
count | connection |
| Consumers | rabbitmq.cluster.consumers |
count | connection |
| Node memory used | rabbitmq.node.memory_used_percent |
percent | per node |
| Node free disk | rabbitmq.node.disk_free_bytes |
bytes | per node |
| Node file descriptors used | rabbitmq.node.file_descriptors_used_percent |
percent | per node |
| Node alarm | rabbitmq.node.alarm |
1 / 0 | per node |
| Node running | rabbitmq.node.running |
1 / 0 | per node |
| Shovel running | rabbitmq.shovel.running |
1 / 0 | shovel |
| Shovel blocked | rabbitmq.shovel.blocked |
1 / 0 | shovel |
| Shovel pending | rabbitmq.shovel.pending |
messages | shovel |
| Shovels not running | rabbitmq.cluster.shovels_not_running |
count | connection |
| Federation link running | rabbitmq.federation_link.running |
1 / 0 | federation link |
| Federation links not running | rabbitmq.cluster.federation_links_not_running |
count | connection |
The rates are unevaluatable where the management plugin's statistics are off. A shovel that is starting, terminated or reported with no state reads running = 0; a federation link that is starting, in error or shut down likewise.
Apache Artemis
| Metric | Key | Unit | Lands on |
|---|---|---|---|
| Queue message count | artemis.queue.message_count |
messages | queue |
| Queue consumers | artemis.queue.consumer_count |
count | queue |
| Delivering | artemis.queue.delivering_count |
messages | queue |
| Scheduled | artemis.queue.scheduled_count |
messages | queue |
| Queue paused | artemis.queue.paused |
1 / 0 | queue |
| Oldest message age | artemis.queue.first_message_age |
milliseconds | queue |
| Unrouted messages | artemis.address.unrouted_count |
messages | address |
| Address size | artemis.address.size_bytes |
bytes | address |
| Address limit used | artemis.address.limit_percent |
percent | address |
| Address paging | artemis.address.paging |
1 / 0 | address |
| Broker live | artemis.broker.live |
1 / 0 | connection |
| Bridge connected | artemis.bridge.connected |
1 / 0 | bridge |
| Bridge pending acknowledgements | artemis.bridge.pending_ack |
messages | bridge |
| Broker connection connected | artemis.broker_connection.connected |
1 / 0 | broker connection |
| Cluster members | artemis.cluster_connection.nodes |
count | connection |
Cluster members counts the nodes this broker's cluster connections see, and is unevaluatable on a broker with none. Connected means the link is up, not that the two brokers are in step.
Memcached
| Metric | Key | Unit | Lands on |
|---|---|---|---|
| Node memory used | memcached.node.memory_used_percent |
percent | per node |
| Node items | memcached.node.items |
count | per node |
| Node connections | memcached.node.connections |
count | per node |
| Node connections used | memcached.node.connections_percent |
percent | per node |
| Node hit rate | memcached.node.hit_rate |
ratio | per node |
| Node unreachable | memcached.node.unreachable |
1 / 0 | per node |
| Node eviction rate | memcached.node.eviction_rate |
per second | per node |
A node evicts before memory used reaches 100%, because items are stored in fixed-size chunks; the eviction rate is the direct signal. It needs two readings, so it is unevaluatable on a node's first reading and after a restart.
What a rule cannot express
No trend — "lag is rising" — because BROKA keeps no time series; a rule's hold for window is the substitute, a breach seen to continue rather than a slope. No learned thresholds, for the same reason. And one metric per rule: "depth above 1,000 and no consumers" is two rules on the same channel.

