1
Access parquet data in S3
Anaconda recommends using
metaflow.S3 in a context manager. The client saves temporary files for the duration of the context, which is why the following example rewrites the file names for access after the scope closes.To access one file, use metaflow.S3.get. Parquet datasets often have many files, which is a good use case for s3.get_many.2
Run the flow
This flow shows how to:
- Download multiple Parquet files using
s3.get_many. - Read the result of the first dataset chunk as a PyArrow table.
load_parquet_to_arrow.py
3
Access artifacts outside of the flow
Run the following in any script or notebook to access the contents of the table that was stored as a flow artifact with
self.table. You can also run quick tests to assert the artifacts have expected properties: