Skip to content

Kubernetes and Helm

The repository has two parallel ways to put the platform on Kubernetes:

  1. Raw manifests in deploy/k8s, applied by deploy/k8s/deploy-all.sh. Used for Minikube demos and the PR workflow that deploys to Minikube.
  2. Helm charts in deploy/helm, installed by deploy/helm/install-helm.sh locally and by the CD workflow on EKS.
  • Directorydeploy/k8s/
    • deploy-all.sh one-shot Minikube deployment (infra, DBs, APIs, gateway, Prometheus, tools, Istio)
    • namespace.yaml ecommerce and monitoring namespaces
    • configmaps.yaml
    • secrets.yaml
    • Directorycatalog/ basket/ discount/ ordering/ Deployment + Service per API and DB
      • …
    • Directorydatabases/ mongodb, redis, postgresql, sqlserver
      • …
    • Directorygateway/ Ocelot Deployment + route ConfigMap
      • …
    • Directoryinfrastructure/ rabbitmq, elasticsearch, kibana
      • …
    • Directorymonitoring/ prometheus (kubernetes_sd pod scraping), grafana, rbac
      • …
    • Directoryingress/ Ingress rules for api-local.eshopping.com (gateway) and localhost (API and monitoring)
      • …
    • Directorynetwork-policies/ default-deny + allow-dns
      • …
    • Directorypod-disruption-budgets/ minAvailable 1 per API
      • …
    • Directorymanagement/ portainer, pgadmin
      • …

deploy-all.sh applies the pieces in dependency order: RabbitMQ, Elasticsearch and Kibana; then the four databases; then waits; then the four APIs and the gateway; then Prometheus, the management tools, and finally Istio. It downloads the latest Istio release with istioctl, labels the default namespace for sidecar injection, and applies Istio’s sample Jaeger, Kiali and Grafana add-ons.

Pods opt into the mesh with sidecar.istio.io/inject: "true". The exceptions are ordering-api and ordering-db, which set it to "false", so traffic to and from Ordering bypasses the mesh (no mTLS, no Istio metrics).

Each service, datastore and tool has its own chart. Every chart has the standard Chart.yaml, values.yaml, _helpers.tpl, templates and a helm test connection pod.

Chart Kind Notable values
deploy/helm/catalog API HPA enabled (1 to 100 replicas, 80% CPU and memory), values-aws.yaml for IRSA and S3
deploy/helm/basket API
deploy/helm/discount API Service named eshopping-discount-discount-grpc, port 8080
deploy/helm/ordering API
deploy/helm/ocelotapigw Gateway ASPNETCORE_ENVIRONMENT=k8s, Service type LoadBalancer with the AWS NLB annotation
catalogdb, basketdb, discountdb, orderdb Datastores In-cluster MongoDB, Redis, PostgreSQL, SQL Server
rabbitmq, elasticsearch, kibana Infra
localstack, pgadmin, portainer Dev tooling
deploy/helm/prometheus/prometheus-values.yaml Values only For the community Prometheus chart

Configuration reaches the pods through a per-chart ConfigMap. values.yaml lists which keys become environment variables, for example DatabaseSettings__ConnectionString and EventBusSettings__HostAddress. Images default to repository: catalogapi, tag: latest. The CD workflow overrides image.registry with the ECR registry.

flowchart LR
  subgraph ns["namespace: dev / staging / production"]
    lb["eshopping-ocelotapigw (LoadBalancer, NLB)"]
    lb --> c["eshopping-catalog"]
    lb --> b["eshopping-basket"]
    lb --> o["eshopping-ordering"]
    b --> d["eshopping-discount-discount-grpc :8080"]
    c --> cdb[("eshopping-catalogdb")]
    b --> bdb[("eshopping-basketdb")]
    d --> ddb[("eshopping-discountdb")]
    o --> odb[("eshopping-orderdb")]
    c & b & o --> mq{{"eshopping-rabbitmq"}}
  end

The deploy-to-eks job in .github/workflows/cd.yml runs, in order:

  1. helm upgrade --install eshopping-{catalogdb,basketdb,discountdb,orderdb,rabbitmq} with --wait.
  2. helm upgrade --install eshopping-{catalog,basket,discount,ordering}. Catalog additionally gets env.USE_LOCALSTACK='false' and an empty AWS__S3__ServiceUrl, so it talks to real S3.
  3. helm upgrade --install eshopping-ocelotapigw.
  4. kubectl wait for all Deployments, then scripts/migrate-products-to-aws-s3.sh rewrites seeded product image URLs to the AWS bucket.

In this path the databases run inside the cluster as Helm releases, not on the managed AWS services that Terraform can provision. See AWS.

Mechanism Where Setting
HorizontalPodAutoscaler Helm autoscaling.* (enabled for Catalog only) CPU and memory 80%, up to 100 replicas
Resource requests and limits Helm values For example, Catalog requests 50m and 64Mi, with limits of 500m and 512Mi
PodDisruptionBudget raw manifests minAvailable: 1 for each API
Replicas raw manifests 1 per Deployment

With one replica per service, a PDB of minAvailable: 1 blocks voluntary eviction entirely, and node drains will hang. Raising replicas to 2 or more, or using maxUnavailable: 1, is the usual fix.