## Intro

Kubernetes Events are often overlooked until a critical incident requires a timeline of what happened. Events record transitions like pod scheduling, health check failures, and node issues, but they are ephemeral: by default, the API server retains them for only about an hour. Losing this history can complicate troubleshooting, audits, and post-mortems. This guide provides a practical approach to backing up and restoring Kubernetes Events, with concrete commands, validation, and recovery procedures. By the end, you will have a repeatable process that fits into your existing operations.

## Version and Environment Inventory

Before implementing a backup solution, document your cluster environment and tool versions. This inventory ensures consistency and helps avoid compatibility issues. Verify the following:

- **Kubernetes version**: Run `kubectl version --short` (or `kubectl version` in newer versions). Example output: `Client Version: v1.29.2`, `Server Version: v1.29.1`.
- **Events API version**: Run `kubectl api-resources | grep events`. Confirm which API group is available: `events.k8s.io/v1` or `events.k8s.io/v1beta1` (deprecated in v1.19+). Example output:
  ```
  events                         events.k8s.io/v1            true         Event
  events                         v1                          true         Event
  ```
  This shows both the core v1 Events and the newer events.k8s.io/v1 group.
- **API server event TTL**: Check the `--event-ttl` flag on the `kube-apiserver`. Default is 1 hour. This defines how long events are retained without backup. You can inspect the API server pod spec or configuration file. For example, in a kubeadm cluster, look at `/etc/kubernetes/manifests/kube-apiserver.yaml` and search for `--event-ttl`. If not set, it defaults to `1h0m0s`.
- **Cluster topology**: Identify namespaces and workloads that generate critical events. For a pilot, focus on one application namespace, for example `default` or a custom namespace like `payments`.
- **Access permissions**: You need read access to events in the target namespace. Run `kubectl auth can-i get events -n your-namespace`. If it returns `yes`, you can proceed. If it returns `no`, you may need to grant permissions via RBAC. Example output when permitted: `yes`.
- **Backup storage location**: Decide where to store exported events (e.g., local filesystem, object storage). For this guide, we use local JSON files.

Example command to check event support:

```bash
kubectl get --raw /apis/events.k8s.io/v1 | jq .
```

Expected output shows the API group version:

```json
{
  "kind": "APIResourceList",
  "apiVersion": "v1",
  "groupVersion": "events.k8s.io/v1",
  "resources": [
    {
      "name": "events",
      "singularName": "event",
      "namespaced": true,
      "kind": "Event",
      "verbs": ["create", "delete", "get", "list", "patch", "update", "watch"]
    }
  ]
}
```

If the `events.k8s.io` API is unavailable, fall back to the core v1 Events API (`/api/v1/namespaces/{namespace}/events`), which has a similar structure but different fields.

## Safe Configuration Path

A safe backup process should be scoped, idempotent, and avoid altering the cluster. We will use native `kubectl` commands to export events as JSON, which can be restored later by recreating events via the API. For production, you may automate this with a CronJob, but the manual steps are essential for validation.

### Scoped Backup

Back up events from a single namespace first:

```bash
kubectl get events -n your-namespace -o json > events-backup-$(date +%Y%m%d-%H%M%S).json
```

Expected output: a JSON file with the list of events. Example snippet:

```json
{
  "apiVersion": "v1",
  "items": [
    {
      "metadata": {
        "name": "my-pod.16a1b2c3d4e5f6g7",
        "namespace": "your-namespace",
        "uid": "12345678-1234-1234-1234-123456789abc",
        "resourceVersion": "12345",
        "creationTimestamp": "2023-05-01T10:00:00Z"
      },
      "involvedObject": {
        "kind": "Pod",
        "namespace": "your-namespace",
        "name": "my-pod",
        "uid": "abcdef12-3456-7890-abcd-ef1234567890"
      },
      "reason": "Scheduled",
      "message": "Successfully assigned your-namespace/my-pod to node-1",
      "type": "Normal",
      "count": 1,
      "firstTimestamp": "2023-05-01T10:00:00Z",
      "lastTimestamp": "2023-05-01T10:00:00Z"
    }
  ],
  "kind": "List",
  "metadata": {
    "resourceVersion": ""
  }
}
```

### Filtering for Relevance

You may want to back up only warning events or events from a specific resource:

```bash
kubectl events -n your-namespace --types=Warning -o json > warnings-backup.json
```

This command exports only events with `type: Warning`. Example output file contains events like `FailedScheduling`, `BackOff`, etc. You can also filter by involved object, for example:

```bash
kubectl events -n your-namespace --for pod/my-pod -o json > my-pod-events.json
```

For more complex filtering, use `jq` on the exported file to select events based on fields.

### Automation with CronJob

For ongoing backups, create a CronJob that runs a script using a service account with read access to events. Example CronJob manifest (constructed example):

```yaml
apiVersion: batch/v1
kind: CronJob
metadata:
  name: event-backup
  namespace: backup-tools
spec:
  schedule: "0 * * * *"
  jobTemplate:
    spec:
      template:
        spec:
          serviceAccountName: event-backup-sa
          restartPolicy: OnFailure
          containers:
          - name: backup
            image: bitnami/kubectl:latest
            command:
            - /bin/sh
            - -c
            - |
              kubectl get events --all-namespaces -o json > /backups/events-$(date +%Y%m%d-%H%M%S).json
            volumeMounts:
            - name: backup-volume
              mountPath: /backups
          volumes:
          - name: backup-volume
            persistentVolumeClaim:
              claimName: event-backup-pvc
```

Note: This requires a PersistentVolumeClaim and appropriate RBAC. For a pilot, manual backup is sufficient.

To set up RBAC for the CronJob service account, create a Role and RoleBinding. Example:

```yaml
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
  name: event-reader
  namespace: backup-tools
rules:
- apiGroups: [""]
  resources: ["events"]
  verbs: ["get", "list", "watch"]
- apiGroups: ["events.k8s.io"]
  resources: ["events"]
  verbs: ["get", "list", "watch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
  name: event-backup-binding
  namespace: backup-tools
subjects:
- kind: ServiceAccount
  name: event-backup-sa
  namespace: backup-tools
roleRef:
  kind: Role
  name: event-reader
  apiGroup: rbac.authorization.k8s.io
```

Apply these with `kubectl apply -f rbac.yaml`.

For the PersistentVolumeClaim, you may use a dynamic provisioner or a simple hostPath for testing. Example PVC:

```yaml
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: event-backup-pvc
  namespace: backup-tools
spec:
  accessModes:
  - ReadWriteOnce
  resources:
    requests:
      storage: 1Gi
```

## Verification and Diagnostics

After backing up, verify the backup file is valid and complete. Use `jq` to inspect and count events:

```bash
jq '.items | length' events-backup-*.json
```

Expected output: a number (e.g., `42`).

Verify that the file contains expected fields:

```bash
jq '.items[0] | {reason, message, type, involvedObject}' events-backup-*.json
```

Example output:

```json
{
  "reason": "Scheduled",
  "message": "Successfully assigned your-namespace/my-pod to node-1",
  "type": "Normal",
  "involvedObject": {
    "kind": "Pod",
    "name": "my-pod",
    "namespace": "your-namespace"
  }
}
```

Compare with the live event count to detect missed events:

```bash
LIVE_COUNT=$(kubectl get events -n your-namespace --no-headers | wc -l)
BACKUP_COUNT=$(jq '.items | length' events-backup-*.json)
echo "Live: $LIVE_COUNT, Backup: $BACKUP_COUNT"
```

Expected output: `Live: 50, Backup: 50` (counts may differ slightly due to event creation between commands). If the backup count is significantly lower, it may indicate that some events expired before backup or that the backup was incomplete.

For automated backups, check the CronJob status:

```bash
kubectl get cronjob event-backup -n backup-tools
```

Expected output shows `LAST SCHEDULE` and `SUSPEND` as `False`:

```
NAME           SCHEDULE    SUSPEND   ACTIVE   LAST SCHEDULE   AGE
event-backup   0 * * * *   False     0        3m20s          1d
```

Check job completions:

```bash
kubectl get jobs -n backup-tools
```

Example:

```
NAME                      COMPLETIONS   DURATION   AGE
event-backup-1651234567   1/1           5s         3m20s
```

If a job fails, inspect logs:

```bash
kubectl logs job/event-backup-1651234567 -n backup-tools
```

## Failure Modes and Recovery

Several failure modes can affect event backup and restore. Understanding these helps you plan recovery.

### Common Failure Modes

| Failure Mode                        | Symptom                                                                 | Mitigation                                                                 |
|-------------------------------------|-------------------------------------------------------------------------|----------------------------------------------------------------------------|
| Backup file missing or empty        | `jq` returns error or count is zero                                     | Add verification step; check storage availability                          |
| Restore fails due to duplicate events | API server rejects events with same name                                | Use `kubectl replace` or delete existing events before restore             |
| Incorrect RBAC permissions          | `kubectl get events` returns Forbidden                                  | Update service account roles                                               |
| Event TTL expiration before backup  | Some events missing from backup                                         | Reduce backup interval or increase `--event-ttl`                           |
| API version mismatch                | Events not recognized on restore                                        | Convert between core v1 and events.k8s.io/v1 if needed                     |

### Rollback and Recovery

If a restore attempt causes issues (e.g., unwanted events flood the cluster), you can delete all restored events in a namespace:

```bash
kubectl delete events --all -n your-namespace
```

Warning: This deletes all events, including live ones. To be more selective, label restored events before applying, then delete by label.

Example selective rollback:

- When restoring, add a custom label by editing the backup file before applying. You can use `jq` to add a label to all events in the backup:

```bash
jq '.items[].metadata.labels.restored = "true"' events-backup.json > events-backup-with-label.json
kubectl apply -f events-backup-with-label.json
```

- To rollback, delete events with that label:

```bash
kubectl delete events -n your-namespace -l restored=true
```

### Restore Procedure

To restore events from a backup file, you must recreate each event object. Because events are immutable in some fields (e.g., `eventTime` for events.k8s.io), you may need to strip certain fields like `metadata.resourceVersion`, `metadata.uid`, and `eventTime` before applying. Use a script to clean and apply:

```bash
jq 'del(.items[].metadata.resourceVersion, .items[].metadata.uid, .items[].eventTime)' events-backup.json | kubectl apply -f -
```

Expected output: `event/event-name created` for each event (or `configured` if exists).

Important: The core v1 Events API does not support `apply`; you may need to use `kubectl create` after cleaning. For events.k8s.io, `apply` works if you omit immutable fields.

For core v1 events, use a loop to create each event individually after cleaning:

```bash
while read -r event; do
  echo "$event" | kubectl create -f -
done < <(jq -c '.items[]' cleaned-events.json)
```

Alternatively, use `kubectl replace` if the event already exists and you want to update it (though events are largely immutable, some fields can be updated).

## Operations Checklist

Use the following checklist for ongoing event backup operations:

- [ ] Inventory Kubernetes and Events API versions; record in runbook. Example: Kubernetes v1.29, Events API events.k8s.io/v1.
- [ ] Confirm RBAC permissions for backup service account (if applicable). Verify with `kubectl auth can-i get events --as=system:serviceaccount:backup-tools:event-backup-sa -n your-namespace`.
- [ ] Select pilot namespace and document event volume. Example: `kubectl get events -n default --no-headers | wc -l` returns 120 events in last hour.
- [ ] Execute manual backup with `kubectl get events -n <ns> -o json > backup.json`.
- [ ] Validate backup: count events, check for non-empty file, inspect sample event.
- [ ] Test restore in a non-production namespace first. Use a separate namespace like `test-restore` and clean up after.
- [ ] Set up automation (CronJob) only after manual validation succeeds.
- [ ] Configure monitoring alerts for backup job failures (e.g., check CronJob status or set up a Prometheus alert on `kube_job_status_failed`).
- [ ] Schedule periodic restore drills (e.g., quarterly) to ensure backups are usable.
- [ ] Review and update backup retention policy; delete old backups from storage. Example: Keep last 30 daily backups, delete older ones automatically.

Example monitoring command for CronJob failures:

```bash
kubectl get jobs -n backup-tools -l job-name=event-backup-<suffix> --field-selector status.successful=0
```

If any job has not succeeded, investigate logs with `kubectl logs job/event-backup-<suffix> -n backup-tools`.

For a more robust alerting, you can set up a Prometheus alert rule like:

```yaml
groups:
- name: kubernetes-events
  rules:
  - alert: EventBackupJobFailed
    expr: kube_job_status_failed{job_name=~"event-backup-.*", namespace="backup-tools"} > 0
    for: 5m
    labels:
      severity: warning
    annotations:
      summary: "Event backup job failed"
      description: "The CronJob event-backup has failed. Check logs for details."
```

## Conclusion

Backing up Kubernetes Events is a low-effort, high-value practice for troubleshooting and compliance. Start with a narrow pilot: manually back up events from one namespace, verify the file, and test restoration. Then automate with a CronJob if needed. Use the operations checklist to maintain a reliable process. By following the examples in this guide, you can preserve cluster event history and recover it when needed, reducing downtime and improving incident response.