Skip to content

Executable tutorial proposal: Deploying a Pre-trained Semantic Search Model as an API with BentoML - #3073

Merged
ericcornelissen merged 2 commits into
KTH:2026from
EccirilloM:executable-tutorial-proposal
Oct 5, 2026
Merged

ericcornelissen merged 2 commits into
KTH:2026from
EccirilloM:executable-tutorial-proposal

Conversation

@EccirilloM

@EccirilloM EccirilloM commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Assignment Proposal

Title

Deploying a Pre-trained Semantic Search Model as an API with BentoML

Names and KTH ID

Deadline

  • Task 2

Category

  • Executable tutorial

Description

This executable tutorial uses Google Colab to demonstrate how BentoML turns an existing machine learning model into a service accessible through an HTTP API. It focuses on serving and operating a pre-trained sentence embedding model available on Hugging Face.

The service proposed as example supports a customer support application by finding the FAQ entry most relevant to a user's question. The model converts the question and FAQ entries into embeddings, which the application compares to identify and return the closest match. This scenario provides a concrete use case for sending inference requests to the service.

The tutorial covers defining a BentoML service, calling its API, inspecting the automatically generated API documentation, and building a versioned Bento that packages the service and specifies its dependencies. It also includes checks for correct service responses, readiness, and request metrics.

The tutorial concludes with an explanation of how the packaged Bento could subsequently be deployed in a container environment. This deployment is discussed conceptually; all executable steps take place in Colab without requiring any cloud credentials.

By the end of the tutorial the user should be able to:

  • Explain how an existing machine learning model becomes an API service.
  • Define and call a BentoML endpoint.
  • Package the service as a versioned Bento.
  • Perform basic checks of service correctness, readiness, and request metrics.

Relevance

Model serving is an important stage of an MLOps workflow. An existing model must be packaged with its code and dependencies, exposed through a stable interface, tested, and monitored. BentoML supports these activities, connecting model serving with DevOps practices such as reproducible packaging, service verification, and observability.

Semantic FAQ search provides just an example through which these practices can be explored.

Tutorial

Open the executable tutorial in Google Colab

The notebook is shared with Viewer permissions to preserve the submitted version and prevent accidental changes to the original.

Some steps invite participants to modify example inputs and explore the service’s behaviour. To edit cells and save your experiments, you should save a copy of the file. No local installation or download is required.

Run the code cells in order, starting with the dependency installation.

@github-actions github-actions Bot added the tutorial One of the task categories listed in README.md label Sep 25, 2026
@ericcornelissen ericcornelissen self-assigned this Sep 29, 2026

@ericcornelissen ericcornelissen left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nice proposal, good luck with creating the tutorial

@ericcornelissen

Copy link
Copy Markdown
Collaborator

Please mark this as ready for review if you want me to merge it 🙂

@TheArtemis

Copy link
Copy Markdown
Contributor

Hi Ettore and Riccardo, as mentioned me and @terahidro2003 would like to give you feedback on the tutorial :)

@EccirilloM
EccirilloM marked this pull request as ready for review October 4, 2026 09:08
@EccirilloM

Copy link
Copy Markdown
Contributor Author

@ericcornelissen I've just mark the PR as ready for Review.
We have a question: We're almost done with the tutorial, how should we submit it? Should we leave a comment with the Colab link, or update the current PR?

@ericcornelissen
ericcornelissen merged commit fa339f6 into KTH:2026 Oct 5, 2026
5 checks passed
@ericcornelissen

Copy link
Copy Markdown
Collaborator

@EccirilloM to submit you update the README.md you created here with a link to the tutorial you created.

@EccirilloM

Copy link
Copy Markdown
Contributor Author

@ericcornelissen I've just update the README.md with the link for the tutorial. Should be all Ok now.

@riccardo-fragale

Copy link
Copy Markdown
Contributor

@TheArtemis @terahidro2003 We are excited to receive the tutorial from you

@ericcornelissen

Copy link
Copy Markdown
Collaborator

@EccirilloM you have to submit a new Pull Request to add the link.

@EccirilloM

Copy link
Copy Markdown
Contributor Author

@ericcornelissen I've just did a new pull request with the link of the code. the new PR is this: #3151

@TheArtemis

Copy link
Copy Markdown
Contributor

General feedback:
Overall, the tutorial gives a clear and practical introduction about serving a machine learning model with BentoML. The progression is natural, going from direct model inference to defining an HTTP service, testing it, inspecting operational metrics, and building a versioned Bento. The explanations make generally clear which functionality comes from BentoML and which parts belong to the application itself.

No account:
The workflow does not require a BentoCloud account, Hugging Face token, or paid deployment service. Running the service locally inside Colab makes the tutorial accessible with only the normal free Google account needed for Colab.

Executability:
The saved outputs show that the model loaded successfully, the service became ready, the API tests and invalid-input test passed, the metrics were retrieved, and the Bento was built. The dependency versions and model revision are pinned, which should reduce differences between different executions. The startup loop and server log also give useful diagnostics if the service fails to become ready. A small imprvement could be to add a final cleanup cell for stopping the server process and closing the log file. Also an estimate of the expected execution time would help readers to know what to expect.

Technical depth:
The tutorial demonstrates a non-trivial workflow and is not limited to calling a model from Python. It defines a BentoML service, exposes an HTTP endpoint, inspects the generated OpenAPI contract, validates successful and unsuccessful requests, retrieves Prometheus metrics, and builds a versioned artifact. The most useful technical extension would be to start the built Bento and repeat the readiness and API tests on it. The current service is started from the project directory, so the successful build confirms that an artifact was created but does not demonstrate that the packaged artifact itself runs correctly.

Relevance:
The DevOps relevance is well motivated. The notebook moves from local inference code to a service with a stable interface, operational checks, tests, metrics, and versioned packaging. Using a pre-trained model is a good choice because it keeps the focus on serving and operating the application instead of model training. A short example about how the build and API tests could be placed in a CI pipeline would make even stronger the connection with DevOps.

System reasoning:
The architecture section clearly distinguishes the notebook client, the separate BentoML server process, and the model and FAQ logic inside the service. The readiness check is especially useful because it shows that starting a process does not mean that the API is immediately ready for receiving traffic. The tutorial also correctly distinguishes operational health from answer quality. A small experiment with concurrent requests could extend the system discussion, showing how the service behaves when multiple clients use it at the same time.

Design decisions:
The notebook explain why BentoML, a small CPU-compatible embedding model, pinned versions, separate FAQ data, and precomputed FAQ embeddings were chosen. It also explains the consequence of loading the FAQ data at startup: changes require to restart the service. A brief comparison with implementing the same API directly in FastAPI would help the readers to understand in which cases BentoML's integrated model-serving, observability, and packaging features justify this choice.

Reflection:
The reflection is quite good. It explains that the service always returns the closest FAQ even when none of the available answers is suitable, and that the similarity score is not a probability that the answer is correct. It also discusses the small dataset, missing production concerns, and the fact that the Bento references the model weights instead of containing them.

Narrative:
The progression from direct inference to data preparation, service definition, readiness checking, API testing, metrics, and packaging is easy to follow. Dividing the service definition in smaller parts makes the decorators and endpoint logic easier to understand. A short checklist of prerequisites and expected execution time near the beginning would make the tutorial more easy to approach.

Visuals:
The architecture diagram helps the explanation of the client and service components. The notebook also transforms the Prometheus data into charts showing request counts and average response times for each HTTP status. This makes more easy to understand the observability section. Since the charts contains only few demonstration requests, it would be helpful to state directly beside them that they illustrate metric collection and should not be interpreted as a performance benchmark.

Language:
The language is clear. The distinction between BentoML's responsibilities and the custom FAQ-matching logic is also explained well.

ILO:
The four intended learning outcomes correspond well with the practical steps in the notebook: understanding the HTTP interface, defining and calling a BentoML service, checking readiness and metrics, and building a versioned Bento. The final outcome is correctly limited to building an artifact for a later deployment. Running the built Bento and repeating the existing tests would give to the readers stronger evidence that they achieved this final learning outcome.

Certification: I/We certify that generative AI, incl. ChatGPT, has not been used to write this feedback. Using generative AI without permission is considered academic misconduct.

Lorenzo Deflorian (ldef@kth.se) and Juozas Skarbalius (jouzas@kth.se)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

tutorial One of the task categories listed in README.md

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants