Intro
Running Apache Kafka in production means more than keeping brokers alive. It means knowing exactly what version you are on, what the current configuration is, what a healthy cluster looks like, and what to do when something drifts from that baseline. This checklist gives developers, DevOps engineers, and startup platform teams a repeatable, read-only-first workflow for Kafka operations.
The checklist is deliberately structured around safety: observe before changing, limit the blast radius, use placeholders instead of real secrets, verify the outcome with a concrete command, and document a recovery path before an incident forces one. Each section includes reusable Kafka CLI examples with explicit placeholders, expected output signals, and a tested rollback step.
By the end of this guide you should be able to audit a cluster version, change a broker or topic setting safely, diagnose consumer lag, recover from a failed broker, and avoid the most common production mistakes.
Version and Environment Inventory
Before touching anything, establish what you are operating.
Component and version range Capture the broker version for every node. Kafka versions are not always uniform during rolling upgrades, and CLI behavior changes between releases.
Read-only observation Run the following on any broker host or from a client with bin access:
kafka-broker-api-versions.sh --bootstrap-server localhost:9092 | head -5
Expected output (Kafka 3.0 - 3.6, truncated):
localhost:9092 (id: 1 rack: us-east-1a) -> (
Produce(0): 0 to 9 [usable: 9],
Fetch(1): 0 to 13 [usable: 13],
ListOffsets(2): 0 to 7 [usable: 7],
...
)
If the command returns org.apache.kafka.common.errors.TimeoutException, the broker is not reachable on that listener, or it is not a broker endpoint. Check advertised.listeners in server.properties and network ACLs.
Topology and identifiers List all brokers and their roles:
kafka-broker-api-versions.sh --bootstrap-server localhost:9092 | grep -E '^\S+ \(id:'
In KRaft mode (Kafka 3.3+), a controller node shows Controller(3) in its API list. In ZooKeeper mode, run:
zookeeper-shell.sh localhost:2181 ls /brokers/ids
Expected: [0, 1, 2] (three broker IDs).
Prerequisites for safe observation
- Network access from the machine running CLI tools to every broker's advertised listener.
- Read-only ACL for
DESCRIBEandDESCRIBE_CONFIGSif Kafka ACLs are enabled. - A current copy of the Kafka distribution matching the cluster's minor version, or at least one major version behind for forward compatibility.
Smallest justified change Do not change anything yet. Record the output in a runbook or config management system with a timestamp. If versions differ across brokers, that is a finding, not a change.
Verification of the change Re-run the same command and compare outputs. Store the baseline in Git or a ticket for incident post-mortems.
Safe Configuration Path
Configuration changes are the most common source of self-inflicted Kafka outages. Follow a controlled path.
Broker Configuration Change: log.retention.hours
Current state Read the existing value:
kafka-configs.sh --bootstrap-server localhost:9092 \
--entity-type brokers --entity-name 1 --describe
Expected output:
Dynamic configs for broker 1 are:
log.retention.hours=168 sensitive=false synonyms={DEFAULT_CONFIG:log.retention.hours=168}
Change description Reduce topic data retention from 168 hours (7 days) to 72 hours for a specific broker only, to free disk space before a planned maintenance window. This change is dynamic and does not require a broker restart if applied via kafka-configs.sh.
Blast radius Only broker 1 is affected. Other brokers keep the default 168 hours. Topics with explicit per-topic retention overrides are not affected because broker-level configs act as defaults.
Apply the change
kafka-configs.sh --bootstrap-server localhost:9092 \
--entity-type brokers --entity-name 1 --alter \
--add-config log.retention.hours=72
Expected output:
Completed updating config for broker 1.
Verify
kafka-configs.sh --bootstrap-server localhost:9092 \
--entity-type brokers --entity-name 1 --describe
Expected output includes log.retention.hours=72.
Rollback
kafka-configs.sh --bootstrap-server localhost:9092 \
--entity-type brokers --entity-name 1 --alter \
--delete-config log.retention.hours
Or set it back to 168 explicitly. Deleting the dynamic override reverts to the default from server.properties.
Topic Configuration Change: min.insync.replicas
Raising min.insync.replicas can stall producers if there are not enough in-sync replicas. For production, perform this change only after verifying the current ISR size.
Current state
kafka-topics.sh --bootstrap-server localhost:9092 \
--topic payments --describe
Expected:
Topic: payments PartitionCount: 3 ReplicationFactor: 3
Partition: 0 Leader: 1 Replicas: 1,2,3 Isr: 1,2,3
...
If any partition shows fewer ISR than replicas, do not raise min.insync.replicas until that is resolved.
Change Set min.insync.replicas=2 to tolerate one broker failure while still requiring two acknowledgments.
kafka-configs.sh --bootstrap-server localhost:9092 \
--entity-type topics --entity-name payments --alter \
--add-config min.insync.replicas=2
Verify
kafka-configs.sh --bootstrap-server localhost:9092 \
--entity-type topics --entity-name payments --describe
Expected: min.insync.replicas=2.
Rollback Delete the override:
kafka-configs.sh --bootstrap-server localhost:9092 \
--entity-type topics --entity-name payments --alter \
--delete-config min.insync.replicas
Guardrails
- Never edit
server.propertieson a live broker without a tested restart plan. Most broker settings are read only at startup. - Use
kafka-configs.shfor dynamic changes. Check which configs are dynamic with:
kafka-configs.sh --bootstrap-server localhost:9092 --entity-type brokers \
--entity-name 1 --describe --all | grep 'sensitive=false'
- For static configs, roll out one broker at a time, wait for the cluster to rebalance, and monitor metrics before the next broker.
Verification and Diagnostics
A healthy Kafka cluster is observable through consistent metrics and command outputs.
Cluster Health Check
Read-only command
kafka-broker-api-versions.sh --bootstrap-server localhost:9092 \
--command-config client.properties 2>&1 | grep -c '^\S+ \(id:'
Expected: the number of brokers (e.g., 3). If lower, a broker is down or not advertised correctly.
Under-Replicated Partitions
kafka-topics.sh --bootstrap-server localhost:9092 --describe --under-replicated-partitions
Expected: empty output (no lines). Any line indicates a partition whose replicas are not fully in sync. Investigate the listed broker's disk, network, or ISR churn.
Consumer Lag
For a consumer group orders-group:
kafka-consumer-groups.sh --bootstrap-server localhost:9092 \
--group orders-group --describe
Example output:
GROUP TOPIC PARTITION CURRENT-OFFSET LOG-END-OFFSET LAG
orders-group orders 0 15234 15240 6
orders-group orders 1 30120 30120 0
Lag above a business-defined threshold (e.g., 1,000 messages or 5 minutes of latency) requires action. Monitor lag with Burrow, Datadog, or Prometheus JMX exporter.
Log Directory Check
kafka-log-dirs.sh --bootstrap-server localhost:9092 --describe --broker-list 0,1,2
This reports per-log-dir size and offline directories. Offline log dirs cause replication failures.
Failure Signal Interpretation
NotEnoughReplicasExceptionon produce meansmin.insync.replicascannot be met. Check ISR and broker status.OffsetOutOfRangeExceptionmeans a consumer tried to read an offset that was already deleted. Reset the consumer group offset.LeaderNotAvailableExceptionoften follows a broker failure or partition reassignment. Wait for leader election or check controller logs.
Failure Modes and Recovery
Plan for the most likely failures, not exotic ones.
Broker Disk Full
Symptom Broker logs show Failed to write to log or Log directory ... is offline. Producers receive NOT_ENOUGH_REPLICAS.
Diagnosis
df -h /var/lib/kafka/data
If usage is 90% or higher, you are at risk. Check the largest topics:
kafka-log-dirs.sh --bootstrap-server localhost:9092 --describe --broker-list 0,1,2
Sort by size to find the offender.
Recovery
- Reduce retention temporarily on the largest topics using
kafka-configs.sh(as shown above). - If necessary, rebalance partitions away from the full broker using
kafka-reassign-partitions.sh. - Clean up old segments: Kafka will delete segments beyond retention once the cleaner runs.
Verification Monitor disk usage and under-replicated partitions until both return to normal.
Broker Crash and Restart
Symptom Broker missing from kafka-broker-api-versions.sh output; partitions lose leadership.
Diagnosis
systemctl status kafka
Or check kafka-server-start.sh logs in /var/log/kafka/server.log.
Recovery
- Identify the root cause (OOM, disk failure, network partition).
- Fix the underlying issue.
- Start the broker:
systemctl start kafka
- Wait for the broker to rejoin the ISR for all its partitions. Monitor
kafka-topics.sh --describefor under-replicated partitions. - If the broker cannot restart due to corrupted logs, delete the corrupted log directory only after verifying that other replicas are in sync.
Verification Under-replicated partitions count should return to zero with no unresolved leadership changes.
Producer Timeout Due to min.insync.replicas
Symptom Producers receive org.apache.kafka.common.errors.NotEnoughReplicasException.
Diagnosis Check ISR for the topic:
kafka-topics.sh --bootstrap-server localhost:9092 --topic high-value --describe
If ISR < min.insync.replicas, producers should fail with that exception unless acks=all is not set.
Recovery
- Bring the missing replicas back into ISR (fix broker or network).
- If time-critical, lower
min.insync.replicastemporarily to allow writes, but document the risk.
Operations Checklist
Use this checklist as a runbook for routine tasks and incident response.
Daily Health Check (5 minutes)
- [ ] Run
kafka-broker-api-versions.sh --bootstrap-server localhost:9092 | grep -c 'id:'and compare to expected broker count. - [ ] Run
kafka-topics.sh --bootstrap-server localhost:9092 --describe --under-replicated-partitionsand ensure no output. - [ ] Check consumer lag for all critical groups. Alert if lag exceeds threshold.
- [ ] Check disk usage on each broker:
df -h /var/lib/kafka/data. Alert if > 85%.
Weekly Review (15 minutes)
- [ ] Review broker logs for repeated warnings or errors.
- [ ] Check for unbalanced leadership:
kafka-topics.sh --describe --topic '*' | grep -c 'Leader: 1'vs other brokers. - [ ] Validate backup and restore process for at least one topic.
Pre-Change Verification
Before any configuration change, answer these:
- What is the current setting? (Run describe command and record output)
- What is the expected new setting and why?
- Which components are affected (broker, topic, client)?
- Is the change dynamic or does it require a restart?
- What is the rollback command or procedure?
- What metric will prove the change worked?
Responsible Owners
- Cluster health checks: On-call engineer, reviewed daily during standup.
- Configuration changes: Platform team lead (e.g., Priya Shah, Engineering Lead), approved via change ticket and revisited in weekly ops meeting.
- Consumer lag and throughput: Data pipeline owner (e.g., Data Engineering Manager), monitored continuously, reviewed weekly.
- Capacity planning disk/topics: Infrastructure owner (e.g., DevOps Lead), reviewed monthly with the platform team.
Common Pitfalls
1. Applying Config Changes Without Checking Version Support
Why it happens Operators use an outdated CLI or assume all configs are dynamic.
How to avoid Always run kafka-configs.sh --version to confirm the client matches the broker. Check the Kafka upgrade notes for the specific version before changing static configs.
Recovery If a change fails, revert to the previous config value immediately and verify with the describe command.
2. Ignoring Under-Replicated Partitions for Too Long
Why it happens They are silent until a broker fails, leading to data loss.
How to avoid Set alerts for under-replicated partitions > 0 for more than 5 minutes. Use Prometheus JMX exporter metric kafka.server:type=ReplicaManager,name=UnderReplicatedPartitions.
Recovery Investigate the affected broker immediately. Check logs, disk, and network. Restart if necessary and wait for ISR catch-up.
3. Changing min.insync.replicas Without Checking ISR
Why it happens Teams increase durability without realizing current ISR might be lower.
How to avoid Always run kafka-topics.sh --describe first. Ensure ISR >= desired min.insync.replicas for every partition.
Recovery If producers stall, lower min.insync.replicas immediately, then fix the ISR issue, then raise it again after ISR recovers.
4. Using kafka-topics.sh --delete Without Cleaning Topic Offsets or Consumer State
Why it happens Operators delete a topic and then consumer groups get stuck.
How to avoid Before deleting a topic, stop all consumers of that topic, delete the consumer group or reset offsets, then delete the topic.
Recovery If consumer groups are stuck, use kafka-consumer-groups.sh --reset-offsets --to-latest --execute to move past missing messages.
5. Not Backing Up server.properties and log4j.properties
Why it happens Everyone assumes config files are in Git, but some changes are made directly on servers during emergencies.
How to avoid Use configuration management (e.g., Ansible, Chef) and version control. Periodically diff live configs against repo.
Recovery If a config is lost, rebuild from the backup and restart the broker in a controlled manner.
Conclusion
A Kafka production operations checklist is only useful if it is version-scoped, observable, and reversible where the technology permits. Copying a command without checking prerequisites and expected output is not an operations procedure; it is a gamble.
Start with a low-risk verification on your cluster today: run the broker API versions command, record the output, and compare it against your expected baseline. Then pick one configuration you have been meaning to audit and walk through the Safe Configuration Path in this guide.
Reliable Kafka operations make failure visible, protect sensitive values, limit changes to the intended resource, and define recovery verification before an incident forces the decision. Keep this checklist in your runbook, and review it after every incident to close gaps.