Repository navigation
Executable tutorial submission: Deploying a Pre-trained Semantic Search Model as an API with BentoML - #3151
Executable tutorial submission: Deploying a Pre-trained Semantic Search Model as an API with BentoML#3151EccirilloM wants to merge 8 commits into
Conversation
|
Readme is not correctly formatted Got: ['Assignment Proposal', 'Title', 'Names and KTH ID', 'Deadline', 'Category', 'Description', 'Tutorial'] |
|
Readme is not correctly formatted Got: ['Assignment Proposal', 'Title', 'Names and KTH ID', 'Deadline', 'Category', 'Description', 'Tutorial'] |
…irilloM/kth-devops-course into executable-tutorial-proposal
|
Readme is not correctly formatted Got: ['Assignment Proposal', 'Title', 'Names and KTH ID', 'Deadline', 'Category', 'Description', 'Tutorial'] |
1 similar comment
|
Readme is not correctly formatted Got: ['Assignment Proposal', 'Title', 'Names and KTH ID', 'Deadline', 'Category', 'Description', 'Tutorial'] |
Assignment Proposal
Title
Deploying a Pre-trained Semantic Search Model as an API with BentoML
Names and KTH ID
Deadline
Category
Description
This executable tutorial uses Google Colab to demonstrate how BentoML turns an existing machine learning model into a service accessible through an HTTP API. It focuses on serving and operating a pre-trained sentence embedding model available on Hugging Face.
The service proposed as example supports a customer support application by finding the FAQ entry most relevant to a user's question. The model converts the question and FAQ entries into embeddings, which the application compares to identify and return the closest match. This scenario provides a concrete use case for sending inference requests to the service.
The tutorial covers defining a BentoML service, calling its API, inspecting the automatically generated API documentation, and building a versioned Bento that packages the service and specifies its dependencies. It also includes checks for correct service responses, readiness, and request metrics.
The tutorial concludes with an explanation of how the packaged Bento could subsequently be deployed in a container environment. This deployment is discussed conceptually; all executable steps take place in Colab without requiring any cloud credentials.
By the end of the tutorial the user should be able to:
Relevance
Model serving is an important stage of an MLOps workflow. An existing model must be packaged with its code and dependencies, exposed through a stable interface, tested, and monitored. BentoML supports these activities, connecting model serving with DevOps practices such as reproducible packaging, service verification, and observability.
Semantic FAQ search provides just an example through which these practices can be explored.
Related Proposal: #3073
TUTORIAL
Open the executable tutorial in Google Colab
The notebook is shared with Viewer permissions to preserve the submitted version and prevent accidental changes to the original.
Some steps invite participants to modify example inputs and explore the service’s behaviour. To edit cells and save your experiments, you should save a copy of the file. No local installation or download is required.
Run the code cells in order, starting with the dependency installation.