Creating a Kubernetes cluster takes minutes. Operating one reliably in production is a different exercise entirely — it depends on decisions about workload architecture, networking, resource management, security and observability that aren’t visible on day one but matter enormously by month three.
A Kubernetes cluster separates control plane (the API server, scheduler and controller manager) from worker nodes, where your workloads actually run. In production, the control plane is usually managed by the cloud provider (EKS, AKS, GKE) so the team's real responsibility is node pool design: sizing, separating workload types across node pools (for example, general-purpose vs. GPU or memory-optimized), and spreading nodes across availability zones.
Workloads themselves should be described declaratively — Deployments for stateless applications, StatefulSets for anything that needs stable identity or storage, and Jobs or CronJobs for finite or scheduled work. Treating pods as directly managed objects, rather than through a controller, removes the self-healing behavior that makes Kubernetes useful in the first place.
A Deployment manages a set of pod replicas and handles rolling updates. A Service gives those pods a stable network identity so other workloads (or an Ingress) can reach them regardless of which specific pods are currently running. Ingress then handles routing external HTTP(S) traffic into the cluster, typically backed by a controller such as NGINX or a cloud load balancer integration.
apiVersion: apps/v1
kind: Deployment
metadata:
name: api
spec:
replicas: 3
selector:
matchLabels: { app: api }
template:
metadata:
labels: { app: api }
spec:
containers:
- name: api
image: registry.example.com/api:1.4.2
resources:
requests: { cpu: "250m", memory: "256Mi" }
limits: { cpu: "500m", memory: "512Mi" }As the number of services grows, hand-writing every manifest becomes hard to maintain consistently. Helm packages a set of manifests into a chart with templated values, so the same chart can be deployed with different configuration per environment. This keeps environment differences explicit (in a values file) rather than scattered across duplicated YAML.
Every pod gets its own IP address, and a Container Network Interface (CNI) plugin handles routing between them across nodes. On top of that, NetworkPolicies act like a firewall between workloads inside the cluster — without them, any pod can typically reach any other pod, which is rarely what a production security posture should allow. A reasonable default is deny-all ingress between namespaces, with explicit policies opening only the traffic that's actually needed.
Requests tell the scheduler how much CPU and memory a pod needs to be placed; limits cap how much it can consume. Pods without requests set are effectively invisible to the scheduler's capacity planning, and pods without limits set can starve their neighbors on a busy node. Autoscaling builds on top of these values: the Horizontal Pod Autoscaler scales replica count based on observed metrics, while the Cluster Autoscaler adds or removes nodes based on whether pending pods can be scheduled on existing capacity.
Kubernetes Secrets store sensitive values, but by default they're only base64-encoded, not encrypted — encryption at rest and integration with an external secrets manager (such as a cloud KMS-backed store) is what actually protects them. RBAC (Role-Based Access Control) governs who and what can act on cluster resources; service accounts used by applications should be scoped to only the permissions that specific workload needs, following the same least-privilege principle as cloud IAM.
Cluster and workload metrics are typically collected with Prometheus and visualized in Grafana, covering both infrastructure signals (node CPU, memory, disk pressure) and application-level metrics exposed by the workloads themselves. Logs from every pod should be shipped off-node to a central store, since pod filesystems — and often the pods themselves — are ephemeral and disappear on restart or rescheduling.
High availability at the workload level means running multiple replicas spread across nodes and zones, with Pod Disruption Budgets to prevent voluntary disruptions (like node drains) from taking down every replica at once. Cluster state itself — including persistent volumes and, if self-managed, etcd — needs its own backup strategy, since losing cluster state is a different and more severe failure than losing a single workload.
Baseline cluster security includes keeping the Kubernetes version and node images patched, restricting which container images can run (ideally from a private, scanned registry), applying Pod Security Standards to limit privileged containers, and auditing API server access logs. None of this is unique to Kubernetes conceptually — it's the same defense-in-depth thinking applied at the orchestration layer.
Kubernetes
Running Kubernetes in production requires more than creating a cluster. Explore the architecture, security, networking, scaling and observability practices that matter.
More on this and related topics.
A complete guide to Amazon CloudWatch Omni, AWS's AI-powered observability platform for applications and AI agents, built on OpenTelemetry.
A practical guide to designing cloud infrastructure with the right balance of reliability, security, scalability and operational control.
How modern engineering teams can automate build, test, security and deployment workflows while keeping releases consistent and recoverable.