Repository navigation
Executable tutorial proposal: Deploying a Pre-trained Semantic Search Model as an API with BentoML - #3073
Conversation
ericcornelissen
left a comment
There was a problem hiding this comment.
Nice proposal, good luck with creating the tutorial
|
Please mark this as ready for review if you want me to merge it 🙂 |
|
Hi Ettore and Riccardo, as mentioned me and @terahidro2003 would like to give you feedback on the tutorial :) |
|
@ericcornelissen I've just mark the PR as ready for Review. |
|
@EccirilloM to submit you update the |
|
@ericcornelissen I've just update the README.md with the link for the tutorial. Should be all Ok now. |
|
@TheArtemis @terahidro2003 We are excited to receive the tutorial from you |
|
@EccirilloM you have to submit a new Pull Request to add the link. |
|
@ericcornelissen I've just did a new pull request with the link of the code. the new PR is this: #3151 |
|
General feedback: No account: Executability: Technical depth: Relevance: System reasoning: Design decisions: Reflection: Narrative: Visuals: Language: ILO: Certification: I/We certify that generative AI, incl. ChatGPT, has not been used to write this feedback. Using generative AI without permission is considered academic misconduct. Lorenzo Deflorian (ldef@kth.se) and Juozas Skarbalius (jouzas@kth.se) |
Assignment Proposal
Title
Deploying a Pre-trained Semantic Search Model as an API with BentoML
Names and KTH ID
Deadline
Category
Description
This executable tutorial uses Google Colab to demonstrate how BentoML turns an existing machine learning model into a service accessible through an HTTP API. It focuses on serving and operating a pre-trained sentence embedding model available on Hugging Face.
The service proposed as example supports a customer support application by finding the FAQ entry most relevant to a user's question. The model converts the question and FAQ entries into embeddings, which the application compares to identify and return the closest match. This scenario provides a concrete use case for sending inference requests to the service.
The tutorial covers defining a BentoML service, calling its API, inspecting the automatically generated API documentation, and building a versioned Bento that packages the service and specifies its dependencies. It also includes checks for correct service responses, readiness, and request metrics.
The tutorial concludes with an explanation of how the packaged Bento could subsequently be deployed in a container environment. This deployment is discussed conceptually; all executable steps take place in Colab without requiring any cloud credentials.
By the end of the tutorial the user should be able to:
Relevance
Model serving is an important stage of an MLOps workflow. An existing model must be packaged with its code and dependencies, exposed through a stable interface, tested, and monitored. BentoML supports these activities, connecting model serving with DevOps practices such as reproducible packaging, service verification, and observability.
Semantic FAQ search provides just an example through which these practices can be explored.
Tutorial
Open the executable tutorial in Google Colab
The notebook is shared with Viewer permissions to preserve the submitted version and prevent accidental changes to the original.
Some steps invite participants to modify example inputs and explore the service’s behaviour. To edit cells and save your experiments, you should save a copy of the file. No local installation or download is required.
Run the code cells in order, starting with the dependency installation.