BROKABROKA
Sign inDownload CommunityRequest a demo
Documentation112 pages
Guides

Alert metrics

Updated 6 October 2026 · applies to 1.0.1 · Community and Commercial

Every figure an alert rule can watch, by platform: the name the rule dialog shows, the key it is stored under, its unit and what it lands on. A rule's metric is chosen from this list and nothing else — a key with no figure behind it would be a rule that saves and never fires.

What a metric lands on decides what a rule's scope expands to. A metric that lands on a kind of resource — a topic, a consumer group, a queue, a connector — measures each one in the scope, or the one a resource-scoped rule names. A metric that lands on the connection has one figure per cluster. One marked per node is measured on each node of the connection — each Kafka broker or controller, each RabbitMQ node, each Memcached node — and an incident names the node.

Three rules hold across the list. A figure BROKA could not read is unevaluatable, with the reason, never 0. A yes-or-no figure is a number, 1 or 0, so it takes the same comparisons as any other. And there are no raw counters: a threshold on a total that only ever grows would fire once and stay fired, so a counter appears as its rate or not at all.

The platforms other than Kafka are in Commercial; their metrics are offered only where the installation serves them.

Kafka

Metric Key Unit Lands on
Consumer group lag kafka.group.lag messages consumer group
Consumer group time lag kafka.group.time_lag milliseconds consumer group
Consume rate kafka.group.consume_rate per second consumer group
Consumer group members kafka.group.members count consumer group
Topic produce rate kafka.topic.produce_rate per second topic
Topic message count kafka.topic.messages messages topic
Topic size kafka.topic.size_bytes bytes topic
Time since last write kafka.topic.last_write_age milliseconds topic
Partition count kafka.topic.partitions count topic
Under-replicated partitions kafka.cluster.urp partitions connection
Cluster messages in kafka.cluster.messages_in_rate per second connection
Broker count kafka.cluster.brokers count connection
Cluster size kafka.cluster.size_bytes bytes connection

Topic and group figures come from the snapshot BROKA refreshes about every thirty seconds; a snapshot older than three refreshes makes them unevaluatable, with its age.

Kafka Connect

Metric Key Unit Lands on
Connect cluster reachable kafka.connect.cluster.reachable 1 / 0 Connect cluster
Connectors failed kafka.connect.cluster.connectors_failed count Connect cluster
Connector tasks failed kafka.connect.cluster.tasks_failed count Connect cluster
Connector failed kafka.connect.connector.failed 1 / 0 connector
Connector paused (or stopped) kafka.connect.connector.paused 1 / 0 connector
Connector unassigned kafka.connect.connector.unassigned 1 / 0 connector
Connector failed tasks kafka.connect.connector.failed_tasks count connector
Connector tasks running kafka.connect.connector.running_ratio ratio connector

A worker that does not answer reads reachable = 0, which a rule can fire on; the other figures are then unevaluatable. A connector is shown by its Connect cluster's name and its own.

Schema Registry

Metric Key Unit Lands on
Schema Registry reachable kafka.schema_registry.reachable 1 / 0 connection
Schema Registry response time kafka.schema_registry.response_ms milliseconds connection
Schema Registry read-only kafka.schema_registry.read_only 1 / 0 connection
Schema Registry subjects kafka.schema_registry.subjects count connection

On a connection with no registry these are unevaluatable and say so. Read-only is the registry's mode READONLY or IMPORT, in which a producer registering a new schema version fails.

Kafka over JMX

Read only where the connection's JMX is set up; without it, every one of them is unevaluatable with that reason. A node whose agent did not answer is unevaluatable on its own while the others are measured.

Metric Key Unit Lands on
Active controllers kafka.cluster.active_controllers count connection
Offline partitions kafka.cluster.offline_partitions partitions connection
Unclean leader elections kafka.cluster.unclean_elections_rate per second connection
Leader elections kafka.cluster.leader_elections_rate per second connection
Under-min-ISR partitions kafka.broker.under_min_isr partitions per node
At-min-ISR partitions kafka.broker.at_min_isr partitions per node
ISR shrinks kafka.broker.isr_shrinks_rate per second per node
Offline log directories kafka.broker.offline_log_dirs count per node
Broker not running kafka.broker.not_running 1 / 0 per node
Metadata lag kafka.node.metadata_lag_ms milliseconds per node
Produce p99 kafka.broker.produce_p99_ms milliseconds per node
Consumer fetch p99 kafka.broker.fetch_consumer_p99_ms milliseconds per node
Follower fetch p99 kafka.broker.fetch_follower_p99_ms milliseconds per node
Produce mean time kafka.broker.produce_mean_ms milliseconds per node
Consumer fetch mean time kafka.broker.fetch_consumer_mean_ms milliseconds per node
Failed produce requests kafka.broker.failed_produce_rate per second per node
Failed fetch requests kafka.broker.failed_fetch_rate per second per node
Bytes rejected kafka.broker.bytes_rejected_rate bytes per second per node
Request handlers idle kafka.broker.request_handler_idle_percent percent per node
Network processors idle kafka.broker.network_idle_percent percent per node
Request queue kafka.broker.request_queue count per node
Response queue kafka.broker.response_queue count per node
Produce purgatory kafka.broker.purgatory_produce count per node
Fetch purgatory kafka.broker.purgatory_fetch count per node
Log flush p99 kafka.broker.log_flush_p99_ms milliseconds per node
Heap used kafka.node.heap_used_percent percent per node
GC time kafka.node.gc_time_rate milliseconds per second per node
Process CPU kafka.node.cpu_percent percent per node
File descriptors used kafka.node.open_fds_percent percent per node
Threads kafka.node.threads count per node
Broker bytes in kafka.broker.bytes_in_rate bytes per second per node
Broker bytes out kafka.broker.bytes_out_rate bytes per second per node
Broker messages in kafka.broker.messages_in_rate per second per node

The four cluster figures are published only from a whole answer: active controllers needs every controller, the other three the one node that says it is the active controller. Leader and unclean election rates are not published by every Kafka — see JMX metrics. A consumer fetch waits up to fetch.max.wait.ms by design, so set a fetch-time threshold above it. Broker messages in is per broker only; the cluster's figure is Cluster messages in, from offsets. A topic's own JMX rates are shown on its page and are not alertable.

Redis

Metric Key Unit Lands on
Memory used redis.memory.used_bytes bytes per node
Memory used (of maxmemory) redis.memory.used_percent percent connection
Memory fragmentation ratio redis.memory.fragmentation_ratio ratio connection
Keyspace hit rate redis.hit_rate ratio connection
Operations per second redis.ops_per_sec per second per node
Connected clients redis.clients.connected count per node
Blocked clients redis.clients.blocked count connection
Key count redis.keys count connection
Replication lag redis.replication.lag_bytes bytes connection
Time since last replication I/O redis.replication.last_io_age seconds connection
Stream length redis.stream.length entries stream
Stream group lag redis.stream.group_lag entries stream
Pending entries redis.stream.pending entries stream

On a Redis Cluster the figures a single node answers for are measured per node and name the shard; the ones that describe one node rather than the deployment are unevaluatable there. Memory used (of maxmemory) is unevaluatable on a Redis with no maxmemory.

RabbitMQ

Metric Key Unit Lands on
Queue depth rabbitmq.queue.messages messages queue
Ready messages rabbitmq.queue.ready messages queue
Unacknowledged messages rabbitmq.queue.unacknowledged messages queue
Queue consumers rabbitmq.queue.consumers count queue
Consumer utilisation rabbitmq.queue.consumer_utilisation ratio queue
Redelivery rate rabbitmq.queue.redeliver_rate per second queue
Queue memory rabbitmq.queue.memory_bytes bytes queue
Total messages rabbitmq.cluster.messages messages connection
Total unacknowledged rabbitmq.cluster.unacknowledged messages connection
Publish rate rabbitmq.cluster.publish_rate per second connection
Deliver rate rabbitmq.cluster.deliver_rate per second connection
Unroutable rate rabbitmq.cluster.unroutable_rate per second connection
Connections rabbitmq.cluster.connections count connection
Consumers rabbitmq.cluster.consumers count connection
Node memory used rabbitmq.node.memory_used_percent percent per node
Node free disk rabbitmq.node.disk_free_bytes bytes per node
Node file descriptors used rabbitmq.node.file_descriptors_used_percent percent per node
Node alarm rabbitmq.node.alarm 1 / 0 per node
Node running rabbitmq.node.running 1 / 0 per node
Shovel running rabbitmq.shovel.running 1 / 0 shovel
Shovel blocked rabbitmq.shovel.blocked 1 / 0 shovel
Shovel pending rabbitmq.shovel.pending messages shovel
Shovels not running rabbitmq.cluster.shovels_not_running count connection
Federation link running rabbitmq.federation_link.running 1 / 0 federation link
Federation links not running rabbitmq.cluster.federation_links_not_running count connection

The rates are unevaluatable where the management plugin's statistics are off. A shovel that is starting, terminated or reported with no state reads running = 0; a federation link that is starting, in error or shut down likewise.

Apache Artemis

Metric Key Unit Lands on
Queue message count artemis.queue.message_count messages queue
Queue consumers artemis.queue.consumer_count count queue
Delivering artemis.queue.delivering_count messages queue
Scheduled artemis.queue.scheduled_count messages queue
Queue paused artemis.queue.paused 1 / 0 queue
Oldest message age artemis.queue.first_message_age milliseconds queue
Unrouted messages artemis.address.unrouted_count messages address
Address size artemis.address.size_bytes bytes address
Address limit used artemis.address.limit_percent percent address
Address paging artemis.address.paging 1 / 0 address
Broker live artemis.broker.live 1 / 0 connection
Bridge connected artemis.bridge.connected 1 / 0 bridge
Bridge pending acknowledgements artemis.bridge.pending_ack messages bridge
Broker connection connected artemis.broker_connection.connected 1 / 0 broker connection
Cluster members artemis.cluster_connection.nodes count connection

Cluster members counts the nodes this broker's cluster connections see, and is unevaluatable on a broker with none. Connected means the link is up, not that the two brokers are in step.

Memcached

Metric Key Unit Lands on
Node memory used memcached.node.memory_used_percent percent per node
Node items memcached.node.items count per node
Node connections memcached.node.connections count per node
Node connections used memcached.node.connections_percent percent per node
Node hit rate memcached.node.hit_rate ratio per node
Node unreachable memcached.node.unreachable 1 / 0 per node
Node eviction rate memcached.node.eviction_rate per second per node

A node evicts before memory used reaches 100%, because items are stored in fixed-size chunks; the eviction rate is the direct signal. It needs two readings, so it is unevaluatable on a node's first reading and after a restart.

What a rule cannot express

No trend — "lag is rising" — because BROKA keeps no time series; a rule's hold for window is the substitute, a breach seen to continue rather than a slope. No learned thresholds, for the same reason. And one metric per rule: "depth above 1,000 and no consumers" is two rules on the same channel.

← PreviousAlertsNext →Search