Slurm Mega. Setting up a production-ready cluster, 3 useful tips from speakers, and Slurm together with Luke Skywalker and R2D2.

On November 24, the Slurm Mega, an advanced intensive course on Kubernetes, concluded. The next Mega will take place in Moscow from May 18 to 20.

Slurm Mega. Setting up a production-ready cluster, 3 useful tips from speakers, and Slurm together with Luke Skywalker and R2D2.

The idea of Slurm Mega is to look under the hood of the cluster, examine the intricacies of setting up and configuring a production-ready cluster ('the-not-so-easy-way') both theoretically and practically, and discuss mechanisms for ensuring the security and reliability of applications.

Bonus of Mega: those who complete Slurm Basic and Slurm Mega receive all the knowledge necessary to pass the exam for CKA at CNCF and a 50% discount on the exam.

Special thanks to Selectel for providing cloud resources for practice, allowing each participant to work in their own complete cluster without raising the ticket price by an additional 5,000.

I won't explain who Bondaryov and Selivanov are; those interested can read here..

Slurm Mega. Day one.

On the first day of Slurm Mega, we overloaded the participants with 4 topics. Pavel Selivanov discussed the process of creating a fault-tolerant cluster from the inside, the operation of Kubeadm, as well as testing and troubleshooting the cluster.

Slurm Mega. Setting up a production-ready cluster, 3 useful tips from speakers, and Slurm together with Luke Skywalker and R2D2.

First coffee break. Usually a ‘call for the teacher’, but at Slurm, while the students drink coffee, the instructors continue to answer questions.

Slurm Mega. Setting up a production-ready cluster, 3 useful tips from speakers, and Slurm together with Luke Skywalker and R2D2.
And despite the cloud 'Break II' hanging over Pavel Selivanov’s head, he is destined not to leave for a break.

Slurm Mega. Setting up a production-ready cluster, 3 useful tips from speakers, and Slurm together with Luke Skywalker and R2D2.
Sergio Bondaryov and Marcel Ibraev are waiting for their turn at the lectern.

During the break, I approached Sergio Bondaryov and asked, 'What advice would you give to all Kubernetes engineers based on your experience working with our clients’ clusters?'

Sergio gave a simple recommendation: 'Restrict internet access to the API server. Because periodically, security threats are found that allow unauthorized users to access the cluster.»

After a few minutes and a bottle of mineral water, Pavel Selivanov charged into the topic of 'Authorization in the cluster using an external provider', specifically LDAP (Nginx + Python) and OIDC (Dex + Gangway).

In the next break, Marcel Ibraev, a speaker at Slurm and Certified Kubernetes Administrator, gave his advice to Kubernetes engineers: 'I may be stating the obvious, but considering how often I come across this, I suspect not everyone keeps it in mind. One should not blindly trust various How-To guides from the Internet that describe how brilliantly a particular solution works. In the context of Kubernetes, this takes on a special significance. Kubernetes is a complex system, and adding a solution to it that hasn't been tested specifically in your project and cluster setup can lead to dire consequences, despite claims of its greatness online. Even Kubernetes itself, without a balanced approach, can harm your project; 'what is good for the Russian is death for the German.' Therefore, we must test, verify, and trial any solution before implementing it in our environment. Only then can you account for all the nuances that may arise.».

Sergiy Bondarev entered the battle after lunch. His topic is Network Policy, specifically an introduction to CNI and Network Security Policy.

Slurm Mega. Setting up a production-ready cluster, 3 useful tips from speakers, and Slurm together with Luke Skywalker and R2D2.

The internet is full of articles on Network Policy. Among admins, there's a sentiment that Network Policies may not be necessary, but security professionals highly value this tool and insist that Network Policies be enabled.

Kubernetes' steering wheel has been taken over by Pavel Selivanov with the topic 'Secure and Highly Available Applications in a Cluster.' He has favorite subjects: PodSecurityPolicy, PodDisruptionBudget, LimitRange/ResourceQuota.

Slurm Mega. Setting up a production-ready cluster, 3 useful tips from speakers, and Slurm together with Luke Skywalker and R2D2.

Megi's topic, with which Pavel presented at DevOpsConf: how to easily and quickly break a Kubernetes cluster and gain all permissions in 5 minutes.

After discussing how easily a Kubernetes cluster can be hacked, skeptically minded admins say: 'Aha, I told you, your Kubernetes is a leaky mess.' Pavel explains that securing a cluster can be done, and it's not difficult; it's just that the security settings are disabled by default. For details, refer to the transcript. of the report.

Slurm Mega. Setting up a production-ready cluster, 3 useful tips from speakers, and Slurm together with Luke Skywalker and R2D2.
— Who broke the cluster? He broke the cluster! I can see it clearly from here!

Things aren't usually simple and easy at Slyorms to avoid boredom. But this time, Telegram decided to show everyone its fifth point:

Marcel Ibraev, [November 22, 2019, 16:52:52]:
Colleagues, there are currently issues with Telegram, please keep this in mind.

Slurm Mega. Setting up a production-ready cluster, 3 useful tips from speakers, and Slurm together with Luke Skywalker and R2D2.

The first day, bright and filled with practical knowledge, has come to an end. The second day will feature even more hands-on experience, including launching a database cluster using PostgreSQL, setting up a RabbitMQ cluster, and managing secrets in Kubernetes.

Slurm Mega. Setting up a production-ready cluster, 3 useful tips from speakers, and Slurm together with Luke Skywalker and R2D2.

Slurm Mega. Second day.

The host kicked off the second day with an energetic announcement: "This morning, as Pavel put it yesterday, we are in for some serious hardcore stuff. Speaking in the language of surgeons, we will dive into the guts of Kubernetes!"

The mass entertainer is a whole story in itself. One of the challenges of Slurm is that people tune out and fall asleep due to information overload. We've always looked for ways to address this, and last time, small audience games showed great results. This time, we hired a specially trained person. The chat had a lot of jokes about the 'interesting contests,' but the fact remains — we have never seen such lively participants.

Slurm Mega. Setting up a production-ready cluster, 3 useful tips from speakers, and Slurm together with Luke Skywalker and R2D2.

Marsel Ibraev received assistance and began studying Stateful applications in the cluster. Specifically, launching a database cluster using PostgreSQL and setting up a RabbitMQ cluster.

After lunch, Sergey Bondarev took over K8S. His topic was 'Storing Secrets.' He was supported by Mulder and Scully. They examined managing secrets in Kubernetes and Vault, as well as 'The truth is out there.'

Slurm Mega. Setting up a production-ready cluster, 3 useful tips from speakers, and Slurm together with Luke Skywalker and R2D2.

This continued until late in the evening when Pavel Selivanov spoke about the Horizontal Pod Autoscaler.

Slurm Mega. Third day.

Abruptly and energetically from the very morning, Sergey Bondarev roused the audience with backup and recovery after failures. They personally checked the backup and recovery of the cluster using Heptio Velero and etcd.

Slurm Mega. Setting up a production-ready cluster, 3 useful tips from speakers, and Slurm together with Luke Skywalker and R2D2.

Sergey continued the topic of annual certificate rotation in the cluster: renewing control-plane certificates using kubeadm. Just before lunch, to whet participants' appetites or actually ruin them, Pavel Selivanov raised the topic of application deployment.

Slurm Mega. Setting up a production-ready cluster, 3 useful tips from speakers, and Slurm together with Luke Skywalker and R2D2.

They discussed templating and deployment tools, as well as deployment strategies.

Pavel Selivanov introduced a new topic: Service Mesh, the installation of Istio. The topic turned out to be so rich that it could warrant a separate intensive session. We're discussing plans, so stay tuned for announcements.

The main thing is to ensure everything works correctly. Because the time for practice has come:
Building CI/CD to deploy the application and update the cluster simultaneously. Everything works well in training projects. But life can be full of surprises.

Slurm Mega. Setting up a production-ready cluster, 3 useful tips from speakers, and Slurm together with Luke Skywalker and R2D2.

May the Slurm be with you!

Source: habr.com

Buy reliable website hosting with DDoS protection, VPS VDS servers 🔥 Buy reliable website hosting with DDoS protection, VPS VDS servers | ProHoster