Skip to main content
When a Parquet dataset lives in S3, Metaflow’s high-throughput S3 client can pull one file or many files into a task at once, and pandas can read the result straight into a dataframe for analysis. This page shows the pattern, using a public dataset stored in S3 from Ookla Global’s AWS Open Data Submission.
1

Access parquet data in S3

To access one file, use metaflow.S3.get. Parquet datasets often have many files, which is a good use case for s3.get_many.
2

Run the flow

This flow shows how to:
  • Download multiple Parquet files using s3.get_many.
  • Read the result of one year of the dataset as a pandas dataframe.
load_parquet_to_pandas.py
3

Access artifacts outside of the flow

Run the following in any script or notebook to access the contents of the dataframe that was stored as a flow artifact with self.df: