services

Make Kubernetes boring — so your team can ship product again.

Senior cloud-native engineering for teams of 2–15 engineers. Fixed-price, audit-first engagements: you get a risk map, a cost figure and a backlog your team can execute — not a retainer you can't leave.

KubernetesIstioAKSEKSGKEArgo CDHelmPrometheusOpenTelemetryGrafanaVeleroKyvernoExternal SecretsOpenCostTerraformKarpenter
how we engage

Audit → Stabilize → Enable.

Every engagement is fixed-price and fixed-scope. No open-ended retainers, no managed services, no lock-in: the goal is always that your team operates the platform without me.

phase 1 · 1–2 weeks

Audit

Read-only assessment of the cluster and the operating model around it. Deliverables: risk map, cost opportunity, prioritized backlog. If the honest answer is "you don't need Kubernetes," the report says so.

phase 2 · 2–4 weeks

Stabilization Sprint

Fix the top risks from the audit: tested restore, safe rollouts, right-sized workloads, RBAC/secrets baseline, alert cleanup. Working alongside your engineers, in your repos.

phase 3 · ongoing-lite

Enablement

Runbooks, supervised drills, upgrade planning and a quarterly health check. Your team runs the platform; I stay one message away — not on your payroll.

01 · flagship · fixed price · 2 weeks

Production Readiness Audit

The full version of the free scorecard, run by a person against your real cluster. Read-only access, scoped and revoked at the end. It answers three questions: what can take production down, what is the bill hiding, and what should you fix first.

  • Week 0 — the hard questions: what runs here, what an hour of downtime costs, who can operate it, when the last restore test happened.
  • Week 1 — the audit: deployment path, RBAC & secrets, cost, observability signal-to-noise, backup reality, upgrade posture.
  • Week 2 — deliverables: risk map (impact × likelihood), cost opportunity in €/month, remediation backlog of max 10 items with effort estimates.
02 · kubernetes

Production Kubernetes Operations

Design and hardening of the operating model: zero-downtime deployments (readiness, graceful shutdown, PDBs), upgrade cadence with deprecated-API checks, tested backup/restore with Velero, and everything-in-Git so the cluster survives the person who built it.

03 · cloud · azure / aws / google cloud

Managed Kubernetes: AKS, EKS, GKE

Migrations from VMs or docker-compose to managed Kubernetes, landing-zone architecture (networking, identity, registries), and multi-cloud honesty: which managed tier you need, which you don't, and when Cloud Run/App Runner/ECS beats Kubernetes entirely. Stateless first, databases last, compose kept alive until the new platform survives two weeks of real traffic.

04 · service mesh · istio

Istio Service Mesh

Traffic management, mutual TLS, retries/timeouts and mesh-level observability — implemented when the case is real (multi-service auth, compliance-driven encryption in transit, progressive delivery). And the cheaper first step when it isn't: the assessment that says "NetworkPolicies and a good ingress cover you for two more years" is part of the service.

05 · finops

Cost Optimization

The eight leaks, hunted systematically: requests vs. real usage, log ingestion and retention, non-prod schedules, load-balancer consolidation, cross-AZ traffic, orphaned volumes, spot strategy, and OpenCost so the number has an owner. Typical outcome for a first-time audit: 30–60% off the total bill with no architecture changes.

06 · security & compliance

Security & Compliance Baseline

The baseline an enterprise customer's questionnaire expects from a company your size: RBAC scoped per human and workload, secrets via External Secrets/Vault/SOPS with real rotation, Kyverno admission policies, NetworkPolicies, image scanning, private API server. Done before the deal appears — not under deadline with an auditor watching.

07 · observability & gitops

Observability & GitOps

The stack a small team can actually operate: OpenTelemetry Collector as the single pipeline, Prometheus with golden-signal dashboards, alerts only on symptoms users feel, log retention that doesn't eat the budget. Plus GitOps with Argo CD sized for a team of three: auto-sync on dev, one-click prod, rollback = git revert.

08 · enablement

Team Enablement & Runbooks

The exit strategy from key-person risk: runbooks for the top failure scenarios written and tested by someone else, supervised upgrade and restore drills, cost/security reviews folded into existing rituals. Includes bilingual delivery (English/Spanish) for distributed teams.

what an engagement looks like

Anatomy of a rescue.

A composite scenario built from the patterns this practice is designed around — not a client reference. If it sounds like your cluster, that's the point.

The setup: 8 engineers, one cluster built by someone who left, €14k/month cloud plus €6k/month observability, deploys causing brief 503s, restore never tested.

Weeks 1–2 (Audit): read-only. Findings: 30% of workloads exist only in the cluster, requests over-provisioned 4×, debug logs retained forever, everyone is cluster-admin.

Weeks 3–4 (Stabilization): restore tested (it failed; then it didn't), preStop + graceful shutdown on top services, requests right-sized, log retention set, everything into Git behind Argo CD.

Week 5 (Enablement): runbooks for the top five failures; the second engineer runs an upgrade, supervised. Cost per namespace lands on a dashboard the team reviews monthly.

~40%off the monthly bill — right-sizing and log retention alone
0dropped requests during deploys after rollout hardening
2engineers who can now upgrade and restore — was 1
0new tools purchased in the making of this rescue
faq

Fair questions.

Do you operate our cluster after the engagement?
No — deliberately. Managed services create dependency; the engagement model ends with your team operating the platform. What I offer afterwards is enablement-lite: a quarterly health check and being one message away.
How is this priced?
Fixed price per phase, agreed before we start. The audit is the entry point and its price is a function of cluster size, not time spent. No hourly billing, no open-ended retainers.
What access do you need for an audit?
Read-only, scoped, and revoked at the end: a read-only kubeconfig, viewer access to the cloud console and the observability stack, and read access to the infra repos. No agent installs, no changes to the cluster.
We're not sure we even need Kubernetes. Is that a valid starting point?
It's the best one. The first deliverable of any engagement can be the honest assessment: whether managed containers or a PaaS covers your next two years. If the answer is "you don't need Kubernetes," the report says so — that's cheaper for you and better for my reputation.
Can you work async / part-time alongside our team?
Yes — the engagement model is designed async-first: written deliverables, recorded walkthroughs, and scheduled working sessions rather than daily meetings. English or Spanish.
What if our setup is too small to justify any of this?
Then don't hire anyone: take the free scorecard, run the self-audit worksheet, and subscribe for one pattern per week. A meaningful part of this practice is published free, on purpose.
next step

Not sure where to start? Score your cluster.

The free scorecard takes one hour with your team and tells you which of these services you actually need — or that you need none of them yet.

Direct contact: see the legal page.