Skip to content

Latest commit

 

History

36 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

microservices-load-testing-gke

A step-by-step DevOps project to deploy and manage a microservices application on Google Kubernetes Engine (GKE). This repository follows a practical learning approach, covering deployment, monitoring, performance evaluation, canary releases, autoscaling, cost optimization, and cluster management.

License: MIT GCP Kubernetes Istio Terraform

Key Findings

  • Load testing identified cartservice as the primary bottleneck, saturating at ~150 concurrent users.
  • Root cause: CPU throttling on cartservice under sustained load, not memory or network.
  • Designed and implemented an HPA-based autoscaling strategy targeting the two most load-sensitive services (cartservice, frontend), with Cluster Autoscaler as a safety net for sustained burst traffic.
  • Implemented a canary release of productcatalogservice using Istio (25% traffic to v2, 75% to v1), verified via sidecar metrics and Kiali.
  • Full observability stack (Prometheus/Grafana) deployed in-cluster to support the performance evaluation.

For the full methodology, thresholds, and justifications, see docs/autoscaling-strategy.md and docs/canary-release.md.

About the Demo Application

This project uses the "Online Boutique" demo application developed by Google. This is a mock application that simulates an online shop, comprised of 11 microservices. While the implementation of each service is simplified compared to realistic applications, it effectively illustrates the structure and operation of a reasonably complex cloud-native application.

The demo application is available from its GitHub repository. The main documentation can be found in the top-level README file and in the docs folder.

Note: This demo application used to be named "Hipster Shop". Some documents and code/configuration files may still reference this name.

Prerequisites

  1. Google Cloud Platform (GCP) account with a project configured.
  2. Required Tools:
    • gcloud CLI
    • kubectl
    • helm
    • Docker
    • terraform
    • ansible
    • istioctl

Design Decisions and Thinking Notes

Detailed design decisions and reasoning are documented in the docs folder.

Notes on Approach

Cluster provisioning: scripts vs Terraform

The GKE cluster is provisioned using gcloud scripts (scripts/create-gke-cluster.sh). This was a deliberate choice to try both approaches across the project:

  • gcloud scripts for the GKE cluster — simpler for a configuration that never changes after creation
  • Terraform for the load generator VM — better fit since the VM is created and destroyed repeatedly across multiple test runs

In a production environment, Terraform would be the standard choice for both, as it provides state tracking, drift detection, and reviewable terraform plan output before any change is applied.

Setup Instructions

1. Clone the Repository

git clone https://github.com/yazan-orcacoder/microservices-load-testing-gke.git
cd microservices-load-testing-gke

2. Environment Configuration

Configure the following variables in config-gke-cluster.sh:

  • PROJECT_ID — your GCP project ID
  • REGION — GCP region (e.g. europe-west3)
  • ZONE — single zone within the region (e.g. europe-west3-a)
  • CLUSTER_NAME — name for the GKE cluster

NODE_LOCATIONS is automatically set to ZONE to ensure exactly 2 nodes are provisioned. Without this, a regional cluster creates NUM_NODES per zone.

To avoid unnecessary costs, clean up and delete all resources after finishing your experiment using delete-gke-cluster.sh.

3. Deploy the Application

Run the deployment script:

deploy-k8s-app.sh

This deploys the Online Boutique application and all autoscaling configuration (HPAs) in a single step via kubectl apply -k overlays/.

4. Local Load Generator (Manual)

This section covers manually deploying the load generator on a local machine as an intermediate step before automated cloud deployment.

5. Cloud Load Generator (Automated)

This section covers automatically deploying the load generator on a GCE VM using Terraform and Ansible.

6. Monitoring

Prometheus and Grafana are deployed inside the cluster using the kube-prometheus-stack Helm chart. See docs/monitoring.md for design decisions and details.

7. Canary Release

A canary release of productcatalogservice is implemented using Istio. Version v2 runs alongside the stable v1, with 25% of traffic routed to v2 and 75% to v1. Traffic splitting is handled by an Istio VirtualService and verified via sidecar metrics and Kiali.

See docs/canary-release.md for the full methodology, YAML configuration, traffic verification data, and cutover procedure.

8. Autoscaling

Performance evaluation identified cartservice as the bottleneck, saturating at ~150 concurrent users. The autoscaling strategy applies HPA to the two services most sensitive to load, with Cluster Autoscaler as a safety net for sustained burst traffic.

Enable Cluster Autoscaler (once, after cluster creation):

scripts/enable-cluster-autoscaler.sh

This sets a min of 2 and max of 4 nodes on the default node pool. HPAs for cartservice and frontend are already included in the overlay and are applied automatically in step 3.

Verify HPA is active:

kubectl get hpa

For the full strategy, threshold justifications, and performance evaluation plan, see docs/autoscaling-strategy.md.

Deployment and Teardown Order

Creating (run in this order)

./scripts/create-gke-cluster.sh       # 1. GKE cluster
./scripts/deploy-monitoring.sh        # 2. Prometheus + Grafana (must exist before Kiali)
./scripts/deploy-k8s-app.sh           # 3. Application + Istio + Kiali
./scripts/deploy-loadgenerator.sh     # 4. Load generator VM (optional)

Destroying (run in this order)

./scripts/destroy-loadgenerator.sh    # 1. Load generator VM
./scripts/destroy-monitoring.sh       # 2. Monitoring stack
./scripts/delete-gke-cluster.sh       # 3. GKE cluster

Or destroy everything at once:

./scripts/teardown.sh

About

Deployed Google's Online Boutique (11 microservices) on GKE with full Prometheus/Grafana observability. Used Locust to load-test the system, identified CPU throttling on cartservice as the primary bottleneck, and designed an HPA-based autoscaling strategy to resolve it.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Contributors

Languages