BROKABROKA
Sign inDownload CommunityRequest a demo
Documentation105 pagesOlder version — go to 1.0.1
Guides

Incident investigation

Updated 30 September 2026 · applies to 1.0.0 · Community and Commercial

Most broker incidents are one of four sentences: consumers stopped, publishers stopped, messages are missing, or something changed. This page is how to answer each of them with what the console can actually measure — and, just as usefully, which readings mean something other than what they look like.

Before anything else: every number here is measured when you ask for it. Nothing is kept — the one exception is an alert incident, which keeps the value that fired it — so a figure worth arguing about later is worth capturing while you are looking at it.


"Consumers stopped"

Kafka

Start at the consumer group. Offset lag is the number of records between the group's committed positions and the end of the log — and the domain of that sum is only the partitions the group has committed to. A partition it holds but never committed to contributes nothing.

So two readings need care:

  • Lag 0 on a group that has never committed is a zero about nothing. A brand-new group, or one just reset, reads the same as a group that is perfectly caught up. The state column is what separates them.
  • Time lag is the gap between two record timestamps — the newest record and the one at the committed offset — not the gap to now. A group stopped for an hour on a topic that received nothing for an hour shows a small time lag. It answers "how much history is between here and the end", which is the question that tells you whether catching up is minutes or a weekend.

Then open the group. The members breakdown is where a single stuck consumer shows up while its siblings are fine.

RabbitMQ

Go to Queues and read the line under each queue's consumer count: N% utilised is how much of the time the broker could hand that queue's messages straight to a consumer. Below 100% with messages waiting means the consumers are the bottleneck — usually a prefetch already full of unacknowledged messages. Channels does not make that call, deliberately: a channel's unacknowledged count is totalled across every consumer on it, while the prefetch it shows is only the last one the channel set, so the two are not comparable.

Then Consumers, where each consumer's own prefetch is, for the two hazards that cause it:

  • Automatic acknowledgement on a queue means the broker forgets each message as it delivers it. A consumer that crashes mid-work loses what it was holding, with no redelivery — so a queue that is draining while work is not getting done looks healthy from the broker's side.
  • No prefetch limit means one consumer can be handed the whole queue before acknowledging anything, which starves every other consumer on it.

A consumer marked as waiting is not broken. On a queue with single active consumer, exactly one consumer receives and the rest stand by. Restarting a healthy standby is the one action guaranteed to make things worse.

For a quorum queue, check the Quorum tab: a queue that has lost its Raft majority is unavailable — accepting nothing and delivering nothing — and one that is running with replicas offline is one failure away from that.

Redis

For a stream consumer group, lag can be unknown, and that is a real answer: Redis could not compute it because entries the group had not read were trimmed away. It is emphatically not "caught up", and it is the one reading on that screen that most changes what you do next.

Then Pending entries for the group: who owns what, how long it has been idle, and how many times each entry has been delivered. An entry delivered five or more times and still pending is a message something keeps failing to process — and claiming it again is a retry, not a fix.


"Publishers stopped"

RabbitMQ first, because its answer is cluster-wide

Open the cluster overview. A resource alarm on any single node — memory over its watermark, or free disk under its limit — makes RabbitMQ stop accepting publishes across the whole cluster. Connections go to blocking and then blocked, and from the application's side publishing simply stops, with no error to read.

That is the first thing to rule out, because everything else looks like an application problem and it is not one. On the Connections screen the same condition appears per client, and the console says outright that the cause is the cluster rather than the client.

A channel in flow is a different thing — client-side back-pressure, self-correcting, and not a resource alarm.

Everywhere: was it refused rather than broken?

A write refused by BROKA leaves a trail entry. If publishing stopped for people rather than for applications, check whether the environment is read-only, whether the connection has its own read-only switch on, or whether a permission changed — all three are recorded, and all three look like an outage from the outside.


"Messages are missing"

This is where reading the right screen matters most, because several ordinary operations remove messages and only some of them look like it.

On RabbitMQ:

  • Purging is not dead-lettering. Purged messages do not reach the dead-letter exchange; they are gone. The confirmation says so, and the audit entry records how many went.
  • Unroutable is not dead-lettered either. A message that matched no binding was never in a queue — it was discarded at the exchange, unless that exchange has an alternate exchange. Different mechanism, different place to look, different fix.
  • Reading a queue is not a peek. Two of the four modes remove messages permanently, and the other two put them back at the head and mark them redelivered — which reorders work for live consumers. If someone read the queue during the incident, the trail says so, and it says which mode they used and what it did.

On Kafka: a topic's retention and emptying a partition or a whole topic all remove records, and the emptying is recorded. Browsing does not — reading a topic in BROKA never joins a consumer group and never commits an offset, so it cannot have disturbed anything.

On Redis: a key with a TTL simply stops existing, and a stream trims. Check the stream's highest deleted id — a stream whose first entry is not its oldest id has been trimmed, and that is where you see it. What a trim or an entry delete did to consumer groups depends on the server: before Redis 8.2 neither consults a group, so an entry a group had been delivered and not acknowledged is removed anyway — the pending list keeps its id and the entry behind it is gone. From 8.2 the trim's audit entry says which consumer-group choice it was made with and what the groups kept.


"Something changed and nobody knows what"

Every change BROKA made is in the trail, including the attempts that were refused — a write blocked by an environment policy, a permission denied, a failed sign-in — and, where one was needed, the reason the person gave for it. A broker write that never got an answer is there too, marked outcome unknown: the broker may have carried it out. A trail holding only successes could not answer this question at all.

The entries record consequences rather than parameters, which is what makes them readable a month later: a purge names how many messages went and that they were not dead-lettered; a queue read names the acknowledgement mode and whether the messages were removed or requeued; a reset names the strategy; a connector restart names whether it took the tasks with it.

Filter to the period, read the refusals as carefully as the successes, and — if this is going to be written up — take an evidence pack for the window while you are there.


What reading costs

Worth knowing before you investigate, not after:

Reading messages
Kafka Takes nothing. No consumer group is joined, no offset is committed, no application's position moves.
Redis streams Takes nothing. A tail follows without consuming, and a stream read moves only its own position.
RabbitMQ streams Takes nothing. The read moves its own position; the messages stay for every other reader.
RabbitMQ queues Takes something. Every mode touches the queue — two remove messages, and the others requeue at the head and mark them redelivered.

The last row is the one to remember at three in the morning.


Capturing what you found

Nothing is kept: no metric history, no lag trend, no stored snapshot of a screen. Measurements are taken from the broker when you ask for them. An alert incident is the exception, and a narrow one: it keeps the value that fired it and when, under Alerts ▸ History.

So while you are in front of it:

  • Note the numbers you are reasoning from, and when you read them.
  • Produce an evidence pack for the window if the incident will be reviewed — it carries the trail, the changes and reads separated, and an integrity verdict, with a manifest stating what it can and cannot claim.
  • Save the view if you built a useful filter, so the next person starts where you finished rather than rebuilding it.
← PreviousAudit reviewNext →Alerts