Skip to main content
In Anaconda Platform, you can execute individual steps of a flow on different cloud providers, such as AWS, Azure, and GCP. Multi-cloud execution helps you:
  • Overcome constraints related to resource availability and services offered.
  • Access specialized compute, such as Trainium on AWS or TPUs on GCP.
  • Optimize spending by moving compute to the most cost-efficient environment.
  • Respect data locality by moving compute to data.
To run a step on a specific cloud, target a compute pool in that cloud with the compute_pool attribute of the @kubernetes decorator. Compute pools span clouds, so each pool belongs to one provider, and targeting the pool determines where the step runs.

Scaling to another compute pool

The following example assumes your deployment has a compute pool that runs on another cloud provider. If your deployment does not have one, speak with your administrator.Your deployment’s compute pools are listed on the Pools tab of the Compute page.
The example runs part of a flow on a compute pool hosted on Azure while the rest of the flow runs on the deployment’s default pools. Save this flow in crosscloudflow.py:
The highlighted line moves the process step’s compute to the named pool on another cloud. The remaining steps have no compute_pool attribute, so they run on the deployment’s default pools in the primary cloud. The flow illustrates a common pattern in cross-cloud processing:
  1. The start step retrieves a dataset in the primary cloud.
  2. The process step scales out to the other cloud.
  3. The join step brings the results back to the primary cloud.
Run the flow:
To watch the load shift between compute pools in real time, select Compute in the left-hand navigation and click Pools.
When the run completes, the process step’s tasks show the pool you named as their execution location.