Why Teams Struggle With Kubernetes - And How to Fix It

We ran 7 Kubernetes clusters on three clouds. Today we run two, both on Azure, and we never left Kubernetes.

47 clusters. €92,000/month in cloud spend. A team on the brink of burnout.

That is how another DevOps team described their setup in I Stopped Using Kubernetes. Our DevOps Team Is Happier Than Ever. They ditched Kubernetes after years of struggling with its complexity. By their numbers, deployment success jumped 89% and infrastructure costs dropped 62%.

I've lived both sides of that story. I've seen teams abandon Kubernetes in frustration, and I've cut a poorly managed cluster from €16,500 to €11,500 a month while making it more reliable. What separates the two is understanding why Kubernetes fails and fixing root causes instead of symptoms.

Our 7 Clusters Had the Same Pattern as Their 47

When teams abandon Kubernetes, the story is always the same: too complex, too expensive, team exhausted. The specifics usually say something else.

My own version was smaller. We ran a prod and a non-prod cluster on each of Azure, AWS, and Google Cloud, plus a seventh used only for testing. Far from 47, but the same pattern, and the same kind of decisions that would have hurt on any platform:

The control planes for 7 managed clusters cost roughly €500/month, so they were never the expensive part. Everything else existed three times, once per provider. Today the same workloads run on two larger clusters on Azure.

This isn't a Kubernetes problem. It's an adoption problem.

Leaving Kubernetes Trades Complexity for Lock-In

Moving away from Kubernetes feels like relief until you hit the ceiling of the simpler alternatives.

You gain something right away: faster deployments with less configuration, less for junior engineers to learn, and fewer things to monitor and operate.

What you lose over time:

The migration savings are real. But they come with technical debt that compounds as your architecture grows.

When Kubernetes Is the Wrong Choice

Kubernetes isn't for everyone. The team that moved to ECS made a rational decision for their context.

You probably don't need Kubernetes if:

For these scenarios, simpler alternatives win:

If that list describes you, pick the simpler tool.

When Kubernetes Is the Right Choice

But if your architecture looks like this, walking away from Kubernetes creates different problems:

You need Kubernetes if:

For large distributed systems, the alternatives mean building custom tooling that replicates what Kubernetes already provides.

Six Things That Make Kubernetes Work

The teams I've seen succeed with Kubernetes do the same six things.

1. Start with Managed Control Planes

Don't run your own control plane. Use EKS, AKS, or GKE.

A managed control plane lists at $0.10 per cluster per hour on all three, roughly €70/month (list prices checked October 2026; on AKS that is the Standard tier). For our 7 clusters that was roughly €500/month, and it's about €140 for the two we run today.

Benefits of managed Kubernetes:

2. Design for Multi-Tenancy, Not Cluster Proliferation

One cluster per environment, not one per environment per provider.

Use namespace isolation with resource quotas and RBAC:

apiVersion: v1
kind: Namespace
metadata:
  name: team-alpha
---
apiVersion: v1
kind: ResourceQuota
metadata:
  name: compute-quota
  namespace: team-alpha
spec:
  hard:
    requests.cpu: "20"
    requests.memory: 40Gi
    limits.cpu: "40"
    limits.memory: 80Gi
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
  name: team-alpha-admin
  namespace: team-alpha
roleRef:
  apiGroup: rbac.authorization.k8s.io
  kind: ClusterRole
  name: admin
subjects:
- kind: Group
  name: team-alpha
  apiGroup: rbac.authorization.k8s.io

The result is isolated teams sharing infrastructure without duplication.

3. Implement Observability From Day One

The team behind the 47 clusters ran five monitoring tools and got 147 false positive alerts out of them. That's alert fatigue.

Start with the essentials:

# Deploy Prometheus + Grafana on day one
helm install prometheus prometheus-community/kube-prometheus-stack \
  --set prometheus.prometheusSpec.retention=30d \
  --set alertmanager.enabled=true \
  --set grafana.enabled=true

Focus on metrics that matter:

Consolidate tooling: One logging solution (Loki or CloudWatch), one metrics solution (Prometheus), one tracing solution (Tempo or Jaeger).

4. Automate Cost Optimization

The €16,500 cluster I optimized to €11,500 came down to four changes:

Rightsized instance types: from m5.4xlarge (16 vCPU, 64 GB RAM) to m5.xlarge (4 vCPU, 16 GB RAM), based on actual utilization data.

Implemented autoscaling:

apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: api-hpa
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: api-server
  minReplicas: 3
  maxReplicas: 20
  metrics:
  - type: Resource
    resource:
      name: cpu
      target:
        type: Utilization
        averageUtilization: 70

Set resource limits:

resources:
  requests:
    cpu: 200m
    memory: 256Mi
  limits:
    cpu: 500m
    memory: 512Mi

Deploy cost monitoring:

# Kubecost for cost allocation per namespace
helm install kubecost cost-analyzer \
  --repo https://kubecost.github.io/cost-analyzer/ \
  --namespace kubecost --create-namespace

5. Invest in Team Training

The learning curve is steep, and "figure it out" is not a training plan.

Our approach:

Trained engineers build reliable systems. Untrained engineers create the complexity that drives teams away.

6. Adopt GitOps for Infrastructure as Code

Every change should be declarative, reviewable, and reversible.

# ArgoCD Application manifest
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
  name: production-api
spec:
  project: production
  source:
    repoURL: https://github.com/company/k8s-manifests
    targetRevision: main
    path: apps/api
  destination:
    server: https://kubernetes.default.svc
    namespace: production
  syncPolicy:
    automated:
      prune: true
      selfHeal: true

Benefits:

We Rightsized Instead of Walking Away

The team that moved from 47 Kubernetes clusters to ECS made a smart decision for their situation. We went the other way. When the need for multi-cloud disappeared, we consolidated our 7 clusters into two larger ones on Azure and stayed on Kubernetes.

Rightsizing is the six things above. For us it started with fewer, larger clusters on managed Kubernetes, AKS in our case.

What it gave us:

What I'd Tell Teams Considering Kubernetes

Start with these questions:

  1. Do we need orchestration? If you have fewer than 10 services, probably not.
  2. Do we have the expertise? If not, can we invest in training?
  3. What's our scale trajectory? Growing fast or relatively stable?
  4. What are our compliance requirements? Audit trails, multi-cloud, zero-trust?
  5. What's our team size? Small teams benefit from simplicity; larger teams need structure.

If you're already using Kubernetes and struggling:

  1. Audit your architecture - Are you over-engineered? (our 7 clusters across three clouds were a red flag)
  2. Measure actual costs - Break down control plane, compute, networking, and human time
  3. Assess team capability - Is the problem Kubernetes or lack of training?
  4. Look for quick wins - Consolidate clusters, implement autoscaling, set resource limits
  5. Consider alternatives - But understand what you're trading away

Consolidating clusters is a cheaper experiment than a migration. Run that one first.


About the author: Henrik Jess is a DevOps engineer with 30+ years of infrastructure experience. He has run Kubernetes on Azure, AWS, and Google Cloud, and consolidated 7 clusters into two.

← Back to Articles ← Back to Home