How to deploy the Core AI Engine on an Exoscale Kubernetes cluster¶
This guide will help you to deploy the Core AI Engine on an Exoscale Kubernetes cluster.
Guide¶
Install and configure required tools¶
Exoscale CLI¶
Install and configure the Exoscale CLI by following the instructions in the official documentation:
-
https://community.exoscale.com/documentation/tools/exoscale-command-line-interface/#installation
-
https://community.exoscale.com/documentation/tools/exoscale-command-line-interface/#configuration.
kubectl¶
Install kubectl by following the instructions in the official documentation at https://kubernetes.io/docs/tasks/tools/#kubectl.
Create and configure a Kubernetes cluster¶
Create a Kubernetes cluster¶
Create a Kubernetes cluster by executing the following commands (inspired by the official documentation at https://community.exoscale.com/documentation/kubernetes/quickstart/):
Install Nginx Ingress Controller¶
Install Nginx Ingress Controller by executing the following commands (inspired by the official documentation at https://kubernetes.github.io/ingress-nginx/deploy/#exoscale and the official Exoscale documentation at https://community.exoscale.com/documentation/sks/loadbalancer-ingress/):
Check if overlay network is functioning correctly¶
Check if overlay network is functioning correctly by executing the following commands (inspired by the documentation at https://ranchermanager.docs.rancher.com/v2.8/troubleshooting/other-troubleshooting-tips/networking#check-if-overlay-network-is-functioning-correctly):
The output should be similar to the following:
If you see error in the output, there is some issue with the route between the pods on the two hosts. This could be because the required ports for overlay networking are not opened.
You can now clean up the DaemonSet by running the following command:
| Execute the following command(s) in a terminal | |
|---|---|
Configure the DNS zone¶
Add an A record to your DNS zone to point to the Nginx Ingress Controller external IP address. You can use a wildcard *.swiss-ai-center.ch (for example) A record to point to the Nginx Ingress Controller external IP address.
Install and configure cert-manager¶
Install cert-manager by executing the following commands (inspired by the official documentation at https://cert-manager.io/docs/installation/kubectl/):
| Execute the following command(s) in a terminal | |
|---|---|
Configure Let's Encrypt issuer¶
Configure Let's Encrypt issuer with HTTP-01 challenge by executing the following commands (inspired by the official documentation at https://cert-manager.io/docs/configuration/acme/ and the tutorial at https://cert-manager.io/docs/tutorials/acme/nginx-ingress/):
| Execute the following command(s) in a terminal | |
|---|---|
| |
Warning
This is a work in progress. It has not been tested yet.
This configuration uses Infomaniak as the DNS provider. You must change the dns01 section to match your DNS provider. Inspired by the official documentation at https://cert-manager.io/docs/configuration/acme/dns01/ and https://github.com/Infomaniak/cert-manager-webhook-infomaniak.
| Execute the following command(s) in a terminal | |
|---|---|
| |
Deploy a dummy pod to validate the Kubernetes cluster configuration¶
Deploy a dummy pod to validate the Kubernetes cluster configuration (inspired by the tutorial at https://cert-manager.io/docs/tutorials/acme/nginx-ingress/).
Create the Kubernetes configuration file by executing the following command:
Warning
This is a work in progress. It has not been tested yet.
Create the Kubernetes configuration file by executing the following command:
Deploy the dummy pod by executing the following commands:
Info
It can take a few minutes for the certificate to be issued.
| Execute the following command(s) in a terminal | |
|---|---|
The output should be similar to the following:
You can delete the dummy pod by executing the following command:
| Execute the following command(s) in a terminal | |
|---|---|
Deploy the Core AI Engine¶
Create a PostgreSQL database¶
Create a PostgreSQL database by executing the following commands (inspired by the official documentation https://community.exoscale.com/documentation/dbaas/quick-start/):
Allow the Kubernetes cluster to access the PostgreSQL database¶
Allow the Kubernetes cluster to access the PostgreSQL database by executing the following commands (inspired by the blog post at https://www.exoscale.com/syslog/sks-dbaas-terraform/ and the GitHub repository at https://github.com/exoscale-labs/sks-sample-manifests/):
Create a S3 bucket¶
Create a S3 bucket by executing the following commands (inspired by the official documentation):
| Execute the following command(s) in a terminal | |
|---|---|
Allow the Core AI Engine to access the S3 bucket¶
Allow the Core AI Engine to access the S3 bucket by executing the following commands:
| Execute the following command(s) in a terminal | |
|---|---|
Configure the Core AI Engine deployment branches¶
Core AI Engine deployments are selected by the branch that receives the change. The repository must contain both of the following long-lived branches:
devrepresents the development environment.mainrepresents the production environment.
The normal development and release workflow is:
gitGraph LR:
commit id: "Current production"
branch dev
checkout dev
commit id: "Current development"
branch issue-123
checkout issue-123
commit id: "Implement issue"
checkout dev
merge issue-123 id: "PR approved - deploy dev" tag: "DEV"
checkout main
merge dev id: "Release approved - deploy prod" tag: "PROD" Create every change branch, such as issue-123, from dev. When the change is ready, open a pull request back to dev. At least one person other than the author must approve the pull request before it is merged. The merge into dev triggers deployment to the development environment.
To release the tested development version, open a pull request from dev to main. Merging that pull request triggers deployment to the production environment. Do not create production releases from feature branches.
The dev branch is mandatory and must be protected. Protect both dev and main so that:
- changes can only be introduced through pull requests;
- at least one approval is required;
- direct pushes are blocked;
- force pushes are blocked;
- branch deletion is blocked.
Note
In the Core AI Engine repository, some pull-request and branch protections are configured at Swiss AI Center organization level and inherited by the repository. Verify that the organization rules apply to both dev and main, then add repository-level rules only for requirements that are not already covered. Repository settings must not weaken or bypass the organization rules.
Update the Core AI Engine GitHub Actions configuration¶
Update the Core AI Engine GitHub Actions configuration by adding/updating the following secrets:
Warning
The PROD_DATABASE_URL must be saved with postgresql:// protocol. The default value is postgres://.
PROD_KUBE_CONFIG: The content of the Kubernetes configuration file (this is an Organization secret in our repository)PROD_DATABASE_URL: The URL of the PostgreSQL databasePROD_S3_ACCESS_KEY_ID: The access key ID of the S3 bucketPROD_S3_SECRET_ACCESS_KEY: The secret access key of the S3 bucketPROD_S3_HOST: The host of the S3 bucket (ex:https://sos-ch-gva-2.exo.io- https://community.exoscale.com/api/sos/)
Update the Core AI Engine GitHub Actions configuration by adding/updating the following variables:
RUN_CICD:trueto enable the CI/CD jobs. This variable does not select an environment; the target environment is selected by thedevormainbranch.PROD_HOST: The host of the Core AI Engine backendPROD_BACKEND_URL: The URL of the Core AI Engine backend used by the frontendPROD_BACKEND_WS_URL: The WebSocket URL of the Core AI Engine backend used by the frontendPROD_FRONTEND_HOST: The URL of the Core AI Engine frontendPROD_S3_BUCKET: The name of the S3 bucketPROD_S3_REGION: The region of the S3 bucket (ex:ch-gva-2)
Deploy the Core AI Engine¶
Merge a reviewed pull request into dev to deploy the development environment. Merge a reviewed pull request from dev into main to deploy the production environment.
Validate the deployment¶
Validate the deployment by executing the following commands:
| Execute the following command(s) in a terminal | |
|---|---|
Deploy a service¶
Deploying a service is similar to deploying the Core AI Engine. You need to update the GitHub Actions configuration by adding/updating the following secrets:
PROD_KUBE_CONFIG: The content of the Kubernetes configuration file (this is an Organization secret in our repository)
Update the GitHub Actions configuration by adding/updating the following variables:
RUN_CICD:truePROD_SERVICE_URL: The URL of the serviceSERVICE_NAME: The name of the service
The service repository must have a dev branch. Create change branches from dev and merge them into dev through reviewed pull requests to deploy the development environment. Merge dev into main through a reviewed pull request to deploy the production environment.
Validate the deployment¶
Validate the deployment by executing the following commands:
| Execute the following command(s) in a terminal | |
|---|---|
Scale the Kubernetes cluster node pools¶
At some point, if you deploy more services, you may need to scale the Kubernetes cluster node pools. You can do this by executing the following commands:
| Execute the following command(s) in a terminal | |
|---|---|
If you only also wanted to change the instance type and/or disk size, you can do this by executing the following commands:
Delete all resources¶
If you ever need to delete all resources, you can execute the following commands:
Related explanations¶
These explanations are related to the current item (in alphabetical order).
None at the moment.
Resources¶
These resources are related to the current item (in alphabetical order).
None at the moment.