Skip to main content
You can query data in S3 with SQL directly from a Metaflow task, using the same AWS tools you would use in any Python script. This page uses AWS Glue and AWS Athena: AWS Glue is a managed extract, transform, and load (ETL) service, and AWS Athena is a serverless SQL service that runs queries against Glue databases.
1

Add parquet files to an AWS Glue database

The following utility function creates the Glue database and writes a small dataset to S3 as .parquet files. AWS Glue works with many other data formats as well.
create_glue_db.py
2

Run the flow

This flow shows how to:
  • Access parquet data with a SQL query using AWS Athena.
  • Transform a dataset.
  • Write a pandas dataframe to AWS S3 as .parquet files.
sql_query_athena.py
3

Access artifacts outside of the flow

Run the following in any script or notebook to access the contents of the dataframe that was stored as a flow artifact with self.dataset: