Repository navigation
Executable Tutorial Proposal: Chaos Engineering with Chaos Mesh - #3168
Merged
ericcornelissen merged 1 commit intoOct 8, 2026
Merged
Conversation
ericcornelissen
approved these changes
Oct 8, 2026
ericcornelissen
left a comment
Collaborator
There was a problem hiding this comment.
Nice and clear proposal, good luck!
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Assignment Proposal
Title
Chaos Engineering: Testing How Your Service Behaves When a Dependency Fails
Names and KTH ID
Deadline
Category
Description
We will create an executable tutorial on chaos engineering with Chaos Mesh, running on
Killercoda's Kubernetes playground.
The setup deploys a
frontendthat serves a list fetched from abackend, andvegetasends constant traffic sothat failed requests and latency are visible during each experiment.
PodChaospod-kill experiment on the backend. Requests fail until the new pod is up. Fix the backendDeployment (replicas, readiness probe) and run the experiment again.
NetworkChaospartition between frontend and backend. The frontend's HTTP call has no timeout, sorequests hang and the frontend stops responding entirely. Fix the frontend and run the experiment again.
Workflowthat runs the partition while aStatusCheckpolls thefrontend and fails the workflow if it stops responding, and set it up to run automatically after deployment.
The tutorial ends with examples of where this test fits in a delivery pipeline: as a gate in a blue-green
deployment, where the new version has to pass it before receiving traffic, or against a staging environment after
each deployment.
Relevance
Pods die and networks fail in every system; the question is whether the system survives it. That is hard to test:
unit and integration tests mock dependencies as healthy, so the failure path is usually first exercised in
production. Chaos engineering injects these failures deliberately and checks that the system
still behaves. Set up as part of the delivery pipeline, it catches a change that removes a timeout or a replica
before users do.