Isha Technologies

TECHNICAL RESOURCES

KubernetesKubernetesContainer OrchestrationHelmCloud Native

Kubernetes in Production: What Teams Need to Get Right

Creating a Kubernetes cluster takes minutes. Operating one reliably in production is a different exercise entirely — it depends on decisions about workload architecture, networking, resource management, security and observability that aren’t visible on day one but matter enormously by month three.

Isha Technologies — Technical Engineering TeamPublished September 12, 202612 min read
In This Article

Cluster Architecture: Nodes and Workloads

A Kubernetes cluster separates control plane (the API server, scheduler and controller manager) from worker nodes, where your workloads actually run. In production, the control plane is usually managed by the cloud provider (EKS, AKS, GKE) so the team's real responsibility is node pool design: sizing, separating workload types across node pools (for example, general-purpose vs. GPU or memory-optimized), and spreading nodes across availability zones.

Workloads themselves should be described declaratively — Deployments for stateless applications, StatefulSets for anything that needs stable identity or storage, and Jobs or CronJobs for finite or scheduled work. Treating pods as directly managed objects, rather than through a controller, removes the self-healing behavior that makes Kubernetes useful in the first place.

Deployments, Services and Ingress

A Deployment manages a set of pod replicas and handles rolling updates. A Service gives those pods a stable network identity so other workloads (or an Ingress) can reach them regardless of which specific pods are currently running. Ingress then handles routing external HTTP(S) traffic into the cluster, typically backed by a controller such as NGINX or a cloud load balancer integration.

YAML
apiVersion: apps/v1
kind: Deployment
metadata:
  name: api
spec:
  replicas: 3
  selector:
    matchLabels: { app: api }
  template:
    metadata:
      labels: { app: api }
    spec:
      containers:
        - name: api
          image: registry.example.com/api:1.4.2
          resources:
            requests: { cpu: "250m", memory: "256Mi" }
            limits: { cpu: "500m", memory: "512Mi" }

Packaging Applications With Helm

As the number of services grows, hand-writing every manifest becomes hard to maintain consistently. Helm packages a set of manifests into a chart with templated values, so the same chart can be deployed with different configuration per environment. This keeps environment differences explicit (in a values file) rather than scattered across duplicated YAML.

Kubernetes Networking

Every pod gets its own IP address, and a Container Network Interface (CNI) plugin handles routing between them across nodes. On top of that, NetworkPolicies act like a firewall between workloads inside the cluster — without them, any pod can typically reach any other pod, which is rarely what a production security posture should allow. A reasonable default is deny-all ingress between namespaces, with explicit policies opening only the traffic that's actually needed.

Resource Requests, Limits and Autoscaling

Requests tell the scheduler how much CPU and memory a pod needs to be placed; limits cap how much it can consume. Pods without requests set are effectively invisible to the scheduler's capacity planning, and pods without limits set can starve their neighbors on a busy node. Autoscaling builds on top of these values: the Horizontal Pod Autoscaler scales replica count based on observed metrics, while the Cluster Autoscaler adds or removes nodes based on whether pending pods can be scheduled on existing capacity.

Secrets and RBAC

Kubernetes Secrets store sensitive values, but by default they're only base64-encoded, not encrypted — encryption at rest and integration with an external secrets manager (such as a cloud KMS-backed store) is what actually protects them. RBAC (Role-Based Access Control) governs who and what can act on cluster resources; service accounts used by applications should be scoped to only the permissions that specific workload needs, following the same least-privilege principle as cloud IAM.

Monitoring and Logging

Cluster and workload metrics are typically collected with Prometheus and visualized in Grafana, covering both infrastructure signals (node CPU, memory, disk pressure) and application-level metrics exposed by the workloads themselves. Logs from every pod should be shipped off-node to a central store, since pod filesystems — and often the pods themselves — are ephemeral and disappear on restart or rescheduling.

High Availability, Backup and Recovery

High availability at the workload level means running multiple replicas spread across nodes and zones, with Pod Disruption Budgets to prevent voluntary disruptions (like node drains) from taking down every replica at once. Cluster state itself — including persistent volumes and, if self-managed, etcd — needs its own backup strategy, since losing cluster state is a different and more severe failure than losing a single workload.

Cluster Security

Baseline cluster security includes keeping the Kubernetes version and node images patched, restricting which container images can run (ideally from a private, scanned registry), applying Pod Security Standards to limit privileged containers, and auditing API server access logs. None of this is unique to Kubernetes conceptually — it's the same defense-in-depth thinking applied at the orchestration layer.

Key Takeaways

  • Use controllers (Deployments, StatefulSets, Jobs) rather than managing pods directly — that’s where self-healing comes from.
  • Set resource requests and limits on every workload; unset values undermine scheduling and can starve other pods.
  • NetworkPolicies are not optional in production — without them, any pod can typically reach any other pod.
  • Secrets need encryption at rest and, ideally, integration with an external secrets manager.
  • Cluster state (persistent volumes, etcd if self-managed) needs its own backup plan, separate from workload replicas.

Kubernetes

Need help with your Kubernetes infrastructure?

Running Kubernetes in production requires more than creating a cluster. Explore the architecture, security, networking, scaling and observability practices that matter.