Creating a resource integration requires an administrator role. If you do not have administrator access, ask your administrator to create the integration before you begin.
- An OpenAI API key configured as a platform integration
- A workstation notebook that validates your setup
- A baseline flow that scores review sentiment using zero-shot LLM inference
- A sensor flow that monitors the database for new data
- Baseline and candidate deployments that process reviews automatically
- A monitoring dashboard that compares workflow performance over time
Configure the OpenAI API key
Several steps in this tutorial call the OpenAI API. You need an API key to run the code samples. To configure the key as a platform integration:- Select Integrations in the left-hand navigation.
- Click OpenAI in the Add an Integration section.
- Enter a name for the integration.
- Enter a description for the integration.
- Enter your OpenAI API key.
- Click Add.
Download the tutorial content
Download the tutorial content to your workstation:~/learn/llm-end-to-end. If you prefer a different location, replace ~/learn with a directory of your choice.
Validate your setup
Open the notebook in00-setup-nb from the ~/learn/llm-end-to-end directory. This notebook installs the required dependencies and validates your OpenAI API key.
Explore the dataset
Open the notebook in01-baseline-nb from the ~/learn/llm-end-to-end directory. This notebook introduces the dataset and calls OpenAI to score the sentiment of clothing reviews. If you do not have an OpenAI API key, you can read along without running the cells.
Run the baseline evaluation
Evaluate how well the LLM performs zero-shot inference on historical data with known labels. Run this command:Run inference on new reviews
Use the same flow to score new reviews without labels:Understand the data pipeline
In production ML systems, there is often a delay between inference time and when the true label becomes available. In this tutorial, the marketing team might wait days or weeks before knowing whether a customer actually recommended the product. Key details about the data pipeline:- New reviews appear in the database at regular intervals (every 10 minutes in this simulation).
- The
recommended_indcolumn uses NULL values for reviews without feedback yet. For existing labels, this column is1if the reviewer recommended the product, and0if they did not. - To work with historical data that has labels, use
fetch_table(table_name, only_labeled=True). - To work with the most recent batch of unlabeled data, use
fetch_table(table_name, only_labeled=False).
Deploy the sensor flow
Deploy a sensor flow that monitors the database and sends an event to the platform when new data is available:schedule parameter in the argo-workflows-create command.
Deploy the baseline flow
Deploy the baseline flow so it listens to the sensor flow’s event and runs inference when new data is available:Deploy a candidate flow
Deploy a candidate flow that uses a different LLM model to compare against the baseline:Monitor and compare flows
Open the notebook in06-monitoring from the ~/learn/llm-end-to-end directory. This notebook uses the Metaflow Client API to compare the performance of the candidate and baseline flows across branches and over time.
Next steps
To build on this tutorial:- Explore other LLM providers and model sizes.
- Integrate the classification pipeline into your existing ML workflows.