The Plexus Satellite manages isolated networked clusters on the AMD Accelerator Cloud.
481
The Plexus Satellite container provides a rich set of tools for setting up and managing an isolated networked cluster on the AMD Plexus™ software stack. Clusters can have either Slurm, Edge SSH or Kubernetes resource managers. The Plexus Satellite container provides the following functions:
AMD Plexus™ software stack allows you to run your own private in-house AI cloud deployment and take advantage of your existing infrastructure and MLOps investment. With Plexus, your data scientists have a single pane of glass for working across on-premises and public clouds.Plexus is a foundational component of AMD Accelerator Cloud (AAC), a data science platform built from the ground up for AI and HPC. Combining leading edge hardware with the rich feature set of the Plexus software stack, AMD Accelerator Cloud offers access to a co-located AI Platform-as-a-Service (AI PaaS). Customers that host their data lakes in co-location facilities can also take advantage of this on-premise installation, bringing the AI cloud experience to their data.
To access and run the Plexus Satellite container on Plexus, you will need the following:
Provider or Admin permissions on your Plexus instance to permit you to add a new cluster. (Check with your Plexus admin)amdih/plexussatellite:2.7.docker pull to ensure an up-to-date image is installed. This example refers to the 2.7 version.Provider role can connect to Plexus Control panel when using Plexus Satellite (NOTE: If you have a regular User account on Plexus ask your Plexus Admin for Provider permissions).Plexus network isolation works by default for the following IP ranges: 10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16 or 100.69.0.0/16. The network isolation will not work in your cluster if you have a different pod CIDR, so if you want to enable network isolation in your cluster, please contact Plexus support. You have more details about how to configure a Kubernetes cluster in this help article https://confluence.amd.com/display/PLX/Kubernetes+configuration+to+meet+AAC+requirements
There is a full guide to configure Slurm clusters in this help article https://confluence.amd.com/display/PLX/Slurm+configuration+to+meet+Plexus+requirements
An Edge SSH cluster can be composed by one or more nodes. Nodes have the following requirements:
The container is executed by using the following pattern:
docker run -p 8080:80 --rm -it amdih/plexussatellite:2.7
You should run the container in interactive mode (-it). Otherwise the container goes into a loop.
-p 8080:80. Enable satellite GUI in port 8080. The container provides a GUI that can be accessed using your browser.
docker run -p 8080:80 --rm -it -v /home/jorge/k8s-test.kubeconfig:/plexus/kubeconfig:ro amdih/plexussatellite:2.7
kubeconfig file in the container. The default used is/plexus/kubeconfig. This file is required both to check and to onboard clusters.docker run -p 8080:80 --rm -it -v /home/jorge/private_key.rsa:/plexus/private_key:ro amdih/plexussatellite:2.7
-v /home/jorge/private_key.rsa:/plexus/private_key:ro
docker run -p 8080:80 --rm -it -v /home/jorge/nodes.txt:/plexus/nodes:ro amdih/plexussatellite:2.7
nodes file in the container. The default used is/plexus/nodes. This file is required to onboard edge multiple nodes clusters.nodes file is one line per node configuration, each line has the format <hostname>;<latitude>;<longitude>.Required if you are using a Plexus platform which is in a local network or in a VPN.
--network=host . This is a required setting when you use any local network resource (cluster or Plexus server).-v /etc/resolv.gitconf:/etc/resolv.conf:ro.
This is required when you are connecting to the cluster using a VPN connection.NOTE: If you are using macOS, the host can be reached by using:
host.docker.internal
Access the container menu. When you run the container it will request the name of the cluster resource manager, it can be either SSH, Slurm or Kubernetes.
-- ---------------------------------------------- --
-- -- Welcome to the cluster satellite script --- --
-- ---------------------------------------------- --
-- Select resource manager [SSH, Slurm or Kubernetes]: --
Enter resource manager [Kubernetes]:
When you select Kubernetes resource manager, the following menu will be displayed:
-- Please enter your choice: --
1) Pre-flight cluster compatibility test
2) Onboard your cluster in Plexus control panel
3) Start the Plexus Satellite
4) Quit
#?
Menu Option 1 will check that the Kubernetes cluster configuration is compliant with the following Plexus requirements:
Once you enter 1 in the menu, it will ask you for:
#? 1
-- Option Pre-flight cluster compatibility test --
Enter kubeconfig path [/plexus/kubeconfig]:
Enter storage class: storage-class
Enter namespace [default]:
Enter pod cidr[]: 10.0.0.0/8
Menu Option 2 will set up the new cluster in the Plexus server.
Once you enter 2 in the menu, the container script will start to onboard your cluster by using the kubeconfig file you have provided, together with the following parameters that the system will ask you for:
#? 2
-- Option: Onboard your cluster in Plexus control panel --
Enter kubeconfig path [/plexus/kubeconfig]:
Enter storage class: storage-class
Enter default storage size for users (Gigabytes) [1]:
Enter cluster name: satellite-cluster
Is it a satellite cluster? [true]:
Enter Plexus server [https://aac-api.amd.com]:
Insert your Plexus email: [email protected]
Insert your Plexus password:
-- Onboarding cluster --
-- Creating user token in https://aac-api.amd.com --
Token generated with value: dc2dd0xxxxxxxxxxxxxx7a12fd5ba71241c777
-- Onboarding cluster --
-- Creating cluster "satellite-cluster" in https://aac-api.amd.com --
Cluster "satellite-cluster" successfully created in: https://aac-api.amd.com/clusters/777
Cluster "satellite-cluster" has uuid: 81bffe12-f7c3-4d3b-b4c7-84f4348ddacf
Once the process is completed the script will return the api server url from the new cluster and the cluster uuid. The cluster uuid will be required for launching the broker satellite; keep it for future broker executions:
And you can find the new cluster in the user interface:
When you select Slurm resource manager, the following menu will be displayed:
-- Please enter your choice: --
1) Pre-flight cluster compatibility test
2) Onboard your cluster in Plexus control panel
3) Start the Plexus Satellite
4) Quit
#?
Menu Option 1 will check that the SLURM cluster configuration is compliant with the following Plexus requirements:
Once you enter 1 in the menu, it will ask you for:
#? 1
-- Option: Pre-flight Cluster compatibility test --
Enter cluster host []: serve.com
Enter cluster port []: 22
Enter cluster username []: ubuntu
Enter partition for testing [debug]: MI300
Enter cluster shared home path []: /home/ubuntu/nfs
Select authentication type [password or private_key]:
Enter authentication type [password]: private_key
Enter cluster private key path [/plexus/private_key]:
Following password will be used to attempt to unlock the key.
Enter cluster password:
Menu Option 2 will set up the new Slurm cluster in the Plexus server.
Once you enter 2 in the menu, the container script will start to onboard your cluster by using the following parameters that the system will ask you for:
Note: SLURM cluster onboarding process is limited to password authentication. In case of using SSH-key authentication, the cluster must be onboarded from the UI.
First of all the credentials to access the cluster by ssh will be requested
Later it requests the cluster name and credentials for Plexus.
#? 2
-- Option: Onboard your cluster in Plexus control panel --
Enter cluster host []: test.cluster.com
Enter cluster port []: 22
Enter cluster shared home path []: /home/ubuntu/nfs
Enter cluster username []: user
Enter cluster password:
Enter cluster name []: satellite-slurm-cluster
Is it a satellite cluster? [true]:
Enter Plexus server [https://aac-api.amd.com]:
Insert your Plexus email: [email protected]
Insert your Plexus password:
-- Onboarding cluster --
-- Creating user token in https://aac-api.amd.com --
Token generated with value: dc2dd0xxxxxxxxxxxxxx7a12fd5ba71241c777
-- Onboarding cluster --
-- Creating cluster "satellite-slurm-cluster" in https://aac-api.amd.com --
Cluster "satellite-slurm-cluster" successfully created in: https://aac-api.amd.com/clusters/777
Cluster "satellite-slurm-cluster" has uuid: 81bffe12-f7c3-4d3b-b4c7-84f4348ddacf
Once the process is completed the script will return the api server url from the new cluster and the cluster uuid. The cluster uuid will be required for launching the broker satellite; keep it for future broker executions:
And you can find the new cluster in the user interface:
When you select SSH resource manager, the following menu will be displayed:
-- Please enter your choice: --
1) Pre-flight cluster compatibility test
2) Onboard your cluster in Plexus control panel
3) Start the Plexus Satellite
4) Quit
#?
Menu Option 1 will check that the Edge SSH cluster configuration is compliant with the following Plexus requirements in one of the nodes. Remember every node must have the same configuration.
Once you enter 1 in the menu, it will ask you for:
-- Option: Pre-flight Cluster compatibility test --
Enter cluster host []: serve.com
Enter cluster port []: 22
Enter cluster username []: ubuntu
Enter cluster shared home path []: /home/ubuntu/nfs
Select authentication type [password or private_key]:
Enter authentication type [password]: private_key
Enter cluster private key path [/plexus/private_key]:
Following password will be used to attempt to unlock the key.
Enter cluster password:
Menu Option 2 will set up the new Edge SSH cluster in the Plexus server.
Once you enter 2 in the menu, the container script will start to onboard your cluster by using the following parameters that the system will ask you for:
Note: Edge SSH cluster onboarding process is limited to password authentication. In case of using SSH-key authentication, the cluster must be onboarded from the UI.
First, the credentials to access the cluster by ssh will be requested
Later it requests the cluster name and credentials for Plexus.
#? 2
-- Option: Onboard your cluster in Plexus control panel --
Enter nodes config path [/plexus/nodes]:
Enter cluster port []: 22
Enter cluster shared home path []: /home/ubuntu
Enter cluster username []: user
Enter cluster password:
Enter cluster name []: satellite-edge-cluster
Is it a satellite cluster? [true]:
Enter Plexus server [https://aac-api.amd.com]:
Insert your Plexus email: [email protected]
Insert your Plexus password:
-- Onboarding cluster --
-- Creating user token in https://aac-api.amd.com --
Token generated with value: dc2dd0xxxxxxxxxxxxxx7a12fd5ba71241c777
-- Onboarding cluster --
-- Creating cluster "satellite-edge-cluster" in https://aac-api.amd.com --
Cluster "satellite-edge-cluster" successfully created in: https://aac-api.amd.com/clusters/777
Cluster "satellite-edge-cluster" has uuid: 81bffe12-f7c3-4d3b-b4c7-84f4348ddacf
Once the process is completed the script will return the api server url from the new cluster and the cluster uuid. The cluster uuid will be required for launching the broker satellite; keep it for future broker executions:
And you can find the new cluster in the user interface:
Both clusters share the same Start the Plexus Satellite option. It will execute the broker daemon, which will be used for communicating between isolated clusters and the Plexus platform.
Once you have entered this option in the menu, it will launch the broker by using the server, user email and password, in addition to the cluster uuid.
We recommend that you set up 15 seconds in the pull interval - if it is set too low, it could provoke the Plexus API to reject requests from your satellite.
#? 3
-- Option Start the Plexus Satellite --
Enter plexus server [https://aac-api.amd.com]:
Enter your Plexus email: [email protected]
Enter your Plexus password:
Enter cluster uuid: 81bffe12-f7c3-4d3b-b4c7-84f4348ddacf
Enter pull interval [15]:
-- Launching broker --
2020-10-09 14:44:02.780978. Getting auth token
Satellite GUI available
2020-10-09 14:44:02.817024. Pulling requests from API
Solution: You need to label the cpu-only nodes with either
node-role.kubernetes.io/plexus-worker-type=plexus-cpu-worker or the gpu nodes with node-role.kubernetes.io/plexus-worker-type=plexus-gpu-worker. or
the hybrid cpu/gpu nodes with plexus-hybrid-cpu-gpu-workerSolution: Decrease the number of gpus or cpus required by your workload.Solution: Decrease the number of maximum cpus per workload in the queue configuration.There are several possible reasons for this problem:
Solution: Fix the kubeconfig file. We recommend to launch the check script by using Option 1, before creating or launching the broker.Solution: Your cluster needs to match the proper Cluster configuration. Read the Plexus help articles. We recommend that you launch the check script by using Option 1.Solution: You can discover the Queues after cluster creation by clicking on the Update Details in the cluster view.An End User License Agreement is included with this product. By pulling and using this container, you accept the terms and conditions of this license.
Help articles: https://aac.amd.com/help
Content type
Image
Digest
sha256:e93606c15…
Size
167 MB
Last updated
about 3 years ago
docker pull amdih/plexussatellite:2.7.1