1
Add parquet files to an AWS Glue database
The following utility function creates the Glue database and writes a small dataset to S3 as
.parquet files. AWS Glue works with many other data formats as well.create_glue_db.py
2
Run the flow
This flow shows how to:
- Access parquet data with a SQL query using AWS Athena.
- Transform a dataset.
- Write a pandas dataframe to AWS S3 as
.parquetfiles.
sql_query_athena.py
3
Access artifacts outside of the flow
Run the following in any script or notebook to access the contents of the dataframe that was stored as a flow artifact with
self.dataset: