Repository navigation
Conversation
Collaborator
|
Readme is not correctly formatted Got: ['Assignment Proposal', 'Title', 'Names and KTH ID', 'Deadline', 'Category', 'Description', 'Submission'] |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Assignment Proposal
Title
Diagnosing HTTP Service Errors with Prometheus and PromQL
Names and KTH ID
Deadline
Category
Description
Tutorial: Killercoda —
Source: GitHub
This tutorial teaches how to investigate an operational failure in a running HTTP
service using application metrics and PromQL. We observe healthy traffic, introduce a
controlled failure, diagnose the affected endpoint, and verify recovery using
Prometheus.
By the end of the tutorial, we will be able to: explain how an instrumented application
exposes metrics and how Prometheus collects them; distinguish cumulative counters from
request rates; write PromQL queries for request rate and error ratio grouped by
endpoint; use metric labels to locate a failing endpoint and confirm recovery; and
explain the limitations of metrics-based diagnosis.
The tutorial is a guided, ~20-30 minute Killercoda scenario. Automated setup starts a
small instrumented Python application, a traffic generator, and Prometheus inside the
browser. Only a free Killercoda account is needed: no local install, paid services, or
external credentials. Each step includes executable commands, expected results, and an
explanation of what it does and why.
Workflow:
rates.
available.
occur, showing how aggregation can hide per-endpoint issues.
the query window moves past the incident.
System architecture. A traffic-generation script sends requests to a small HTTP
application, which exposes a request counter via
/metrics, labeled by endpoint andstatus code. Prometheus scrapes these metrics on an interval and serves them for
querying and graphing. A diagram will show the application, traffic flow, and metrics
collection path.
Design decisions. Prometheus combines metrics collection and querying in one tool,
keeping the exercise focused. Counters support rate and error-ratio calculations, while
endpoint/status labels support diagnosis; labels use a bounded set of values to avoid
time-series growth. A provided container setup and scripted traffic/failure controls
make the environment and the incident fully reproducible.
Reflection on applicability. This approach is useful for spotting service-wide or
endpoint-specific spikes in errors, but aggregate metrics can hide individual failures,
error ratios can mislead under low traffic, and scrape/query windows introduce
observation delay. Metrics locate symptoms; logs or traces are often still needed to
find the root cause. This makes the tutorial's relevance clearest for developers and
operators responsible for keeping a running service healthy.
The tutorial will be validated end-to-end from a fresh Killercoda environment (setup,
healthy traffic, failure injection, diagnosis, and recovery) within Killercoda's
one-hour free session limit, before submission.
Relevance
This tutorial addresses monitoring and observability by showing how runtime
measurements support incident investigation and recovery. It connects application
instrumentation, automated metrics collection, and operational diagnosis in a single
reproducible workflow, and is scoped narrowly around interpreting runtime metrics for
service health rather than overlapping with other DevOps aspects such as testing/CI or
deployment pipelines.