.parquet files from cloud storage into memory on remote workers quickly, combine the following three things:
- Load data from S3 directly to memory with Metaflow’s optimized S3 client (
metaflow.S3), which reaches tens of gigabits per second or more. - Decode the Parquet data efficiently with Apache Arrow.
- Keep the result in Arrow’s in-memory tables, which are interoperable with modern data tools without additional copies, speeding up processing and avoiding unnecessary memory overhead.
Cloud to table
Before writing a Metaflow flow, look at how to use the Metaflow S3 client with Apache Arrow. The main steps are using themetaflow.S3.get_many function to parallelize the retrieval of the .parquet file partitions, loading the bytes into memory on the worker instance, and decoding the bytes into a pyarrow.Table object.
metaflow.S3.list_recursive:
./metaflow.s3.foobar:
Performance benefits scale with instance size
Using the basic pattern described above, you can write Metaflow flows that scale this fast data speedup on cloud instances. This workflow organizes the same operations presented in the previous section in a Metaflow flow. Notice that thedata_processing step is annotated with @batch(..., use_tmpfs=True, ...). The tmpfs feature extends the resources you request, because it allows you to use memory on the Batch instance to instantiate a temporary file system. This makes the cloud-to-table workflow significantly faster and does not require using the local file system to temporarily store the .parquet bytes.
The benefits of this workflow scale with the number of processors, available RAM, and I/O throughput of the machine you are loading a table on, so use an instance that can fit your entire Arrow table in memory to get maximal benefits. To get a sense of how fast this workflow can get, read Fast Data: Loading Tables From S3 At Lightning Speed.
fast_data_processing.py