Creating a resource integration requires an administrator role. If you do not have administrator access, ask your administrator to create the integration before you begin.
- An S3 bucket and Athena database with queryable data
- An IAM role that allows Anaconda Platform to interact with Athena
- A workstation notebook and a Metaflow flow that run Athena SQL queries
Create Athena resources
AWS Athena runs SQL queries over data assets in S3. If you already have an Athena setup, skip to Chain the Athena role with the platform task role. If you want to set up a test, follow the steps below to create an S3 bucket and an IAM role for Athena.Create an S3 bucket
In the AWS console, create a bucket for your data. Copy the bucket’s Amazon Resource Name (ARN) or keep the console open. You need the ARN when you create the IAM role.Create an IAM role
Athena requires an IAM role with permissions to run queries and access the data bucket. For the full setup guide, see the AWS Athena getting started documentation. Create an IAM role with the following minimum permissions:- The
AWSAthenaFullAccessmanaged policy - S3 bucket access for query results
- Glue Data Catalog access, if you use Glue catalogs
IAM policy template
Configure the query results location
- Create an S3 bucket for Athena to store query results and metadata in.
- In the Athena console, set the query result location to your S3 bucket.
- Confirm your IAM role has access to this bucket.
Chain the Athena role with the platform task role
To allow Anaconda Platform tasks to use your Athena role, register it as an integration in the platform:- Select Integrations in the left-hand navigation.
-
Click AWS in the Add an Integration section.
Do not use the Amazon S3 integration. Athena requires IAM role chaining for the Athena API, Glue Data Catalog, and S3, which the S3 integration does not provide.
- Enter a name and description for the integration.
-
Enter the ARN of the IAM role you created.
If you need the ARN, expand the Getting your IAM role ARN section in the panel.This shows the trust policy statement you need to add to your role, and the tag key and value required for the platform to discover it. You can choose to use an existing target role or create a new one; the panel shows the required trust policy and tagging steps for either path.
- Click Add.
role_arn value for your flows. Copy this snippet into your Metaflow steps to access Athena through the chained role.
Download the tutorial content
Download the tutorial content to your workstation:~/learn/athena. If you prefer a different location, replace ~/learn with a directory of your choice.
Load data into your bucket
Open the notebook in00-setup from the ~/learn/athena directory. Before running it, update the bucket_name and role_arn variables with the S3 bucket and IAM role you created earlier. This notebook walks you through putting data into your S3 bucket so Athena can query it.
Query Athena from a workstation notebook
Open the notebook in01-nb from the ~/learn/athena directory. This notebook walks you through running SQL queries using Athena. You will:
- Connect to Athena using your configured role
- Write and execute SQL queries
- Retrieve and analyze query results
Query Athena from a Metaflow workflow
Open the02-flow directory from the ~/learn/athena directory. This directory contains a Metaflow flow that runs SQL queries using Athena. Before running it, update the bucket_name and role_arn variables in flow.py with the same values you used in the notebooks.
Run the flow:
Next steps
To build on this tutorial:- Create more complex queries that combine multiple data sources.
- Build automated reporting workflows using Athena.
- Integrate Athena queries into your ML pipelines.