Intro
Running TLS in production means knowing exactly why a certificate works today and having a concrete plan for the day it stops working. This checklist translates that goal into read-only diagnostic commands, scoped configuration changes, and recovery procedures that are safe to run under pressure. It is written for developers, DevOps consultants, and technical startup teams who need to verify certificates without introducing new risk.
The checklist covers the full operational life cycle: inventorying what is installed, validating a certificate and its chain, detecting common failures before users do, and recovering without a long outage. Each section includes the exact command or check to run, what the expected output should look like, what a failure signal means, and who should own the follow-up.
The core principle is operational safety: observe before changing, limit the blast radius, use placeholders instead of secrets, verify the result, and document how to recover if the expected state is not reached.
Version and Environment Inventory
Before touching anything, identify the component that terminates TLS, the OpenSSL version in use, and the path to every certificate and key involved. This inventory prevents the classic mistake of renewing a certificate on the wrong server or assuming the terminal uses the same version as the command line.
Run these read-only checks first:
# OpenSSL version and default configuration
openssl version -a
# List certificate files that are actually present
ls -l /etc/ssl/certs/*.pem /etc/ssl/private/*.pem 2>/dev/null
# Identify the process listening on port 443
sudo ss -tlnp | grep ':443'
Expected output for a typical Ubuntu 22.04 host looks like:
OpenSSL 3.0.2 15 Mar 2022 (Library: OpenSSL 3.0.2 15 Mar 2022)
-rw-r--r-- 1 root root 1944 Feb 28 11:20 /etc/ssl/certs/example.com.pem
-rw------- 1 root root 3272 Feb 28 11:20 /etc/ssl/private/example.com.key
LISTEN 0 511 0.0.0.0:443 0.0.0.0:* users:(("nginx",pid=1234,fd=6))
A missing certificate file or a key with world-readable permissions is a failure signal. The key file must be mode 600 and owned by the service user, not by a personal account.
Record the following inventory fields in a place the whole team can see, such as a runbook or internal wiki:
| Field | Example Value | How to Check |
|---|---|---|
| Service name | public-web-prod | systemctl status nginx |
| TLS termination point | Nginx 1.24 on Ubuntu 22.04 | nginx -V |
| Certificate path | /etc/ssl/certs/example.com.pem | ls -l |
| Private key path | /etc/ssl/private/example.com.key | ls -l |
| CA bundle path | /etc/ssl/certs/chain.pem | ls -l |
| Owner | Alex Morgan, Platform Engineer | assigned in team runbook |
| Renewal mechanism | certbot renew with DNS-01 challenge | certbot certificates |
| Last verified | 2025-06-10 | openssl x509 -enddate |
Certificate File Inspection
Inspect a certificate file without changing it:
openssl x509 -in /etc/ssl/certs/example.com.pem -noout \
-subject -issuer -serial -dates -fingerprint -sha256
Expected output:
subject=CN = example.com
issuer=C = US, O = Let's Encrypt, CN = R11
serial=04A2F08D9B1234C56789ABCDEF0123456789A
notBefore=May 28 11:20:00 2025 GMT
notAfter=Aug 26 11:20:00 2025 GMT
SHA256 Fingerprint=3C:19:6E:...:C5:2A
Record the expected subject, issuer, validity window, and SHA-256 fingerprint before deployment. After any renewal or change, re-run this same command and compare the fingerprint. A changed fingerprint with the same file path often means an unexpected certificate replacement.
Remote Endpoint Verification
Test a remote endpoint exactly as a client would, not with a generic TCP check:
openssl s_client -connect example.com:443 -servername example.com -showcerts </dev/null
A healthy result ends with:
Verify return code: 0 (ok)
A successful TCP connection alone does not prove certificate validity. Look for these specific failure signals:
Verify return code: 10 (certificate has expired)means the certificate is past its notAfter date.Verify return code: 21 (unable to verify the first certificate)means the client cannot build a chain, usually because an intermediate certificate is missing from the server configuration.verify error:num=2:unable to get issuer certificatepoints to a missing root or intermediate in the client trust store.- A hostname mismatch appears as
verify error:num=62:Hostname mismatchwhen the-servernamevalue does not match the certificate SAN.
Offline Chain Validation
Validate the full chain offline before deploying to a live server:
openssl verify -CAfile /etc/ssl/certs/trusted-root-ca.pem \
-untrusted /etc/ssl/certs/intermediate.pem \
/etc/ssl/certs/example.com.pem
Expected success output:
/etc/ssl/certs/example.com.pem: OK
If the output says unable to get local issuer certificate, the -CAfile or -untrusted argument is missing a required issuer. The fix is to download the correct intermediate from the CA and add it to the chain bundle, not to disable verification.
Use placeholders in documentation and never commit private keys to a repository. Test renewal and rollback procedures well before the expiry window becomes urgent.
Safe Configuration Path
After the inventory is complete, make only one scoped change at a time and know how to roll it back. For example, when a certificate expires and must be replaced on an Nginx server, do not edit the live configuration directly without a backup.
Minimal Change Example: Replacing an Expired Certificate
- Create a backup of the current configuration and certificate files:
sudo cp /etc/nginx/sites-available/example.com.conf \
/etc/nginx/sites-available/example.com.conf.bak.$(date +%F)
sudo cp /etc/ssl/certs/example.com.pem \
/etc/ssl/certs/example.com.pem.old.$(date +%F)
- Place the new certificate and key in the expected paths. For a Let's Encrypt renewal, certbot usually does this automatically, but if you are copying manually, preserve permissions:
sudo install -o root -g root -m 644 new-cert.pem /etc/ssl/certs/example.com.pem
sudo install -o root -g root -m 600 new-key.pem /etc/ssl/private/example.com.key
- Test the Nginx configuration for syntax errors before reloading:
sudo nginx -t
Expected output:
nginx: configuration file /etc/nginx/nginx.conf test is successful
- Reload Nginx to apply the change without dropping active connections:
sudo nginx -s reload
- Verify the running server now presents the new certificate:
echo | openssl s_client -connect localhost:443 -servername example.com 2>/dev/null \
| openssl x509 -noout -fingerprint -sha256
Compare the fingerprint with the one recorded in inventory. If they match, the change succeeded. If not, roll back by restoring the backups and reloading Nginx again.
Rollback Checklist
For any certificate change, write down the exact rollback commands before you start. Example:
# Rollback certificate and key
sudo cp /etc/ssl/certs/example.com.pem.old.2025-06-10 /etc/ssl/certs/example.com.pem
sudo cp /etc/ssl/private/example.com.key.old.2025-06-10 /etc/ssl/private/example.com.key
# Rollback Nginx config if it was changed
sudo cp /etc/nginx/sites-available/example.com.conf.bak.2025-06-10 \
/etc/nginx/sites-available/example.com.conf
# Reload and verify
sudo nginx -s reload
Kubernetes Ingress Example
For Kubernetes Ingress, the same principle applies. Before updating a TLS secret, record the current one:
kubectl get secret example-tls -n production -o yaml > example-tls-backup.yaml
Create the new secret from files:
kubectl create secret tls example-tls \
--cert=/path/to/tls.crt \
--key=/path/to/tls.key \
-n production --dry-run=client -o yaml | kubectl apply -f -
Then verify the Ingress controller picked up the change:
kubectl get ingress example-ingress -n production -o jsonpath='{.spec.tls}' ; echo
Check the actual certificate presented by the Ingress:
kubectl get secret example-tls -n production -o jsonpath='{.data.tls\.crt}' | base64 -d \
| openssl x509 -noout -dates -subject
If the new certificate is not served, check the controller logs and the Ingress resource for a reference to the correct secret name. The owner of this change is the on-call engineer, and the rollback is to re-apply the backup YAML.
Verification and Diagnostics
Verification is not a one-time event. Schedule a weekly read-only check that confirms every public-facing certificate is still valid and serving correctly. Automate it with a script that fails loudly when something is wrong.
Automated Certificate Health Script
Save the following as cert-health-check.sh and run it from cron or a CI job:
#!/usr/bin/env bash
set -euo pipefail
DOMAINS=("example.com" "api.example.com" "app.example.com")
WARN_DAYS=14
for domain in "${DOMAINS[@]}"; do
echo "Checking $domain..."
cert_file="/tmp/${domain}.pem"
echo | openssl s_client -connect "${domain}:443" -servername "${domain}" 2>/dev/null \
| openssl x509 -out "${cert_file}"
enddate=$(openssl x509 -in "${cert_file}" -noout -enddate | cut -d= -f2)
end_epoch=$(date -d "${enddate}" +%s)
now_epoch=$(date +%s)
diff_days=$(( (end_epoch - now_epoch) / 86400 ))
echo "${domain}: expires in ${diff_days} days (${enddate})"
if [ "${diff_days}" -lt "${WARN_DAYS}" ]; then
echo "WARNING: ${domain} certificate expires in ${diff_days} days!" >&2
exit 1
fi
done
echo "All certificates are valid."
When a certificate is healthy, the script prints:
Checking example.com...
example.com: expires in 55 days (Aug 26 11:20:00 2025 GMT)
Checking api.example.com...
api.example.com: expires in 55 days (Aug 26 11:20:00 2025 GMT)
Checking app.example.com...
app.example.com: expires in 55 days (Aug 26 11:20:00 2025 GMT)
All certificates are valid.
If one is about to expire, it prints the warning and exits with status 1, which can alert via your monitoring system.
Diagnosing Chain Problems
When a client reports an untrusted certificate but openssl verify on the server says OK, the difference is usually the client trust store. Diagnose with the -showcerts option:
echo | openssl s_client -connect example.com:443 -servername example.com -showcerts 2>/dev/null
The output lists every certificate the server sends. A complete chain shows the leaf, the intermediate, and sometimes the root. If the intermediate is missing, the client cannot build the chain. Check the server configuration: for Nginx, the ssl_certificate directive must point to a file that concatenates the leaf and all intermediates in order:
ssl_certificate /etc/ssl/certs/example.com.chained.pem;
Create the chained file with:
cat example.com.pem intermediate.pem > example.com.chained.pem
Then reload Nginx and re-test. The owner of this diagnostic is the security engineer on call, and the fix is reviewed at the next weekly ops sync.
Checking Cipher and Protocol Configuration
TLS certificates can be valid while the server still accepts outdated protocols. Verify the minimum protocol and cipher suite:
echo | openssl s_client -connect example.com:443 -tls1_2 -servername example.com 2>&1 | grep -E 'Protocol|Cipher'
Expected modern output:
Protocol : TLSv1.3
Cipher : TLS_AES_256_GCM_SHA384
If the server negotiates TLSv1.0 or an export cipher, update your Nginx or service configuration to disable old protocols. A secure baseline for Nginx 1.24 is:
ssl_protocols TLSv1.2 TLSv1.3;
ssl_ciphers HIGH:!aNULL:!MD5;
ssl_prefer_server_ciphers on;
Reload and re-test. This change has a low blast radius because only new connections are affected, but validate with a browser and an automated SSL test suite before declaring success.
Failure Modes and Recovery
TLS outages in production are almost always one of four failure modes: expiry, hostname mismatch, incomplete chain, or private key compromise. Each has a distinct signature and a distinct recovery path.
Failure Mode 1: Expired Certificate
Signal: openssl s_client returns Verify return code: 10 (certificate has expired). Users see browser warnings and API clients fail with CERTIFICATE_VERIFY_FAILED.
Why it happens: Renewal automation failed silently, or the certificate was renewed but the service was never reloaded.
Recovery:
- Check the actual file on disk:
openssl x509 -in /etc/ssl/certs/example.com.pem -noout -enddate
If it shows an old date, renew manually. For certbot:
sudo certbot renew --cert-name example.com --force-renewal
If the file shows a new date, reload the service:
sudo nginx -s reload
Then re-test. The owner of this recovery is the on-call engineer, and the root cause is reviewed at the next incident postmortem.
Failure Mode 2: Hostname Mismatch
Signal: openssl s_client prints verify error:num=62:Hostname mismatch or the browser says the certificate is not valid for the domain.
Why it happens: The certificate was issued for a different domain, or the service is reached via an IP address or an internal name not in the SAN list.
Recovery:
- List the SAN entries:
openssl x509 -in /etc/ssl/certs/example.com.pem -noout -text | grep -A1 'Subject Alternative Name'
- If the required hostname is missing, obtain a new certificate that includes it. Do not work around it by disabling hostname verification in clients.
- For internal services, ensure the certificate includes the internal FQDN (for example
service.internal.example.com) in the SAN.
Failure Mode 3: Incomplete Chain
Signal: Verify return code: 21 (unable to verify the first certificate), or server tests like SSL Labs report "Chain issues: Incomplete".
Why it happens: The server only sends the leaf certificate, not the intermediate. Modern browsers sometimes cache intermediates, but curl, Python, and Node.js clients fail.
Recovery:
- Download the correct intermediate from the CA web site (for Let's Encrypt, fetch
https://letsencrypt.org/certs/2024/r11.pem).
- Concatenate leaf and intermediate:
cat example.com.pem r11.pem > example.com.chained.pem
- Update Nginx
ssl_certificateto point to the chained file and reload. Re-test withopenssl s_client -showcertsto confirm both certificates are presented.
Failure Mode 4: Private Key Compromise
Signal: Unexpected use of the certificate, alerts from certificate transparency logs, or a security incident report.
Why it happens: The private key was exposed in a repository, a backup, or a compromised server.
Recovery:
- Revoke the certificate immediately with the CA. For Let's Encrypt, use
certbot revoke --cert-name example.com --reason keycompromise.
- Generate a new private key and certificate request:
openssl req -new -newkey rsa:2048 -nodes -keyout new.key -out new.csr -subj "/CN=example.com"
- Obtain a new certificate from the CA with the CSR.
- Replace the key and certificate on the server, reload, and verify the fingerprint changes.
- Rotate any other secrets that were stored on the same system.
The owner of a key compromise response is the security lead, and the incident is reviewed with a full root-cause analysis within 72 hours.
Operations Checklist
This consolidated checklist is formatted for use during a maintenance window or incident. Each item names the command, expected result, owner, and review frequency.
Pre-deployment Checklist
| # | Check | Command | Expected | Owner | Review |
|---|---|---|---|---|---|
| 1 | Inventory all certificates and keys | ls -l /etc/ssl/certs /etc/ssl/private | Files present, key mode 600 | Platform engineer | Quarterly |
| 2 | Record certificate fingerprints | openssl x509 -in cert.pem -noout -fingerprint -sha256 | Fingerprint documented | Security engineer | On change |
| 3 | Validate full chain offline | openssl verify -CAfile root.pem -untrusted inter.pem leaf.pem | OK | On-call engineer | Before deploy |
| 4 | Confirm private key matches certificate | diff <(openssl x509 -in cert.pem -noout -modulus) <(openssl rsa -in key.pem -noout -modulus) | No output | Platform engineer | Before deploy |
Post-deployment Checklist
| # | Check | Command | Expected | Owner | Review |
|---|---|---|---|---|---|
| 1 | Remote verification | openssl s_client -connect host:443 -servername host -showcerts | Verify return code: 0 (ok) | On-call engineer | Immediately after deploy |
| 2 | Hostname match | Check SAN against expected domain | Domain listed in SAN | Release manager | Immediately after deploy |
| 3 | Protocol and cipher | openssl s_client -connect host:443 -tls1_2 | TLSv1.2 or TLSv1.3, strong cipher | Security engineer | Monthly |
| 4 | Expiry warning window | Script check with WARN_DAYS=14 | No warning | Automation | Daily |
Example Execution Log
During a set of certificate renewals, a team recorded the following log:
2025-06-10 09:00 CST - Alex Morgan - Pre-deploy checks passed for example.com, api.example.com.
2025-06-10 09:05 CST - Alex Morgan - Renewed certificates via certbot. New fingerprints recorded in inventory.
2025-06-10 09:10 CST - Priya Shah - Reloaded Nginx on web-prod. nginx -t passed.
2025-06-10 09:12 CST - Priya Shah - Remote verification returned code 0 for all three domains.
2025-06-10 09:15 CST - Team - Post-deploy checklist completed. No issues.
Such a log makes it possible to reconstruct exactly what changed and who approved each step.
Common Pitfalls and How to Avoid Them
Even experienced teams repeatedly hit the same certificate problems. Here are the most frequent pitfalls and what to do about them.
Pitfall 1: Relying on Expiry Alerts from a Single Source
Why it happens: Teams set up a monitoring alert from one CA or one script, and when that script breaks, the expiry goes unnoticed.
How to avoid: Run at least two independent expiry checks: one script like the one in this article, and one external monitoring tool such as UptimeRobot, Pingdom, or a cron job on a different host. Configure both to alert at 14 days before expiry.
Recovery if it happens: If an expiry notice arrives too late, immediately add a manual calendar reminder for all certificates, run the inventory command, and renew the most urgent ones first.
Pitfall 2: Using Wildcard Certificates Unnecessarily
Why it happens: Wildcard certificates seem convenient because one certificate covers many subdomains.
Why it is a problem: A compromised wildcard private key can be used to impersonate any subdomain, greatly expanding the blast radius. Also, some compliance standards forbid wildcard certificates for customer-facing systems.
How to avoid: Use wildcard certificates only when you have more than five subdomains on the same domain and can tightly control the private key. Otherwise, issue separate certificates per subdomain, or use an internal CA with short-lived certificates.
Pitfall 3: Manual Renewal Without a Test Environment
Why it happens: Small teams often renew production certificates by hand because automation seems overkill.
How to avoid: Use an ACME client like certbot, lego, or a managed service that automates renewal. If manual renewal is unavoidable, always test the entire procedure in a staging environment using the same commands and file paths.
Recovery if it happens: If a manual renewal fails, roll back using the backup files and rehearse the procedure again. Do not iterate on the production server under pressure.
Pitfall 4: Bundling the Root Certificate in the Server Chain
Why it happens: Some teams concatenate the root certificate with the leaf and intermediate, thinking clients need it.
Why it is a problem: The root certificate is already in the client trust store. Sending it adds no value and can trigger warnings in some TLS libraries.
How to avoid: Send only the leaf and intermediate certificates in the chain, never the root. Check with openssl s_client -showcerts and ensure the last certificate in the list is not a self-signed root.
Pitfall 5: Changing File Permissions During Deployment
Why it happens: Engineers copy new certificate files with cp and forget to set restrictive permissions, leaving private keys readable by other users.
How to avoid: Always use install -m 600 for key files and install -m 644 for certificates, or set permissions explicitly after copying:
sudo chmod 600 /etc/ssl/private/example.com.key
sudo chown root:root /etc/ssl/private/example.com.key
Conclusion
A TLS certificate operations checklist is useful only when each recommendation is version-scoped, observable, and reversible where the technology permits. Copying a command without checking prerequisites and expected output is not an operations procedure.
Start with one low-risk verification: pick a single production certificate, run the read-only inspection commands from this article, and record the current state. Then run the expiry check script and compare the result with the expected signal. Review the inventory table and assign a single owner for each certificate, with a weekly or monthly review cadence depending on the certificate's criticality.
A reliable technical workflow makes failure visible, protects sensitive values, limits changes to the intended resource, and defines recovery verification before an incident forces the decision. Document every change in an execution log, keep backups of every file you touch, and rehearse rollback procedures before you need them. That is the difference between predictable certificate operations and a late-night scramble.