On November 24, the Slurm Mega, an advanced intensive course on Kubernetes, concluded. will take place in Moscow from May 18 to 20.

The idea of Slurm Mega is to look under the hood of the cluster, examine the intricacies of setting up and configuring a production-ready cluster ('the-not-so-easy-way') both theoretically and practically, and discuss mechanisms for ensuring the security and reliability of applications.
Bonus of Mega: those who complete Slurm Basic and Slurm Mega receive all the knowledge necessary to pass the exam for and a 50% discount on the exam.
Special thanks to Selectel for providing cloud resources for practice, allowing each participant to work in their own complete cluster without raising the ticket price by an additional 5,000.
I won't explain who Bondaryov and Selivanov are; those interested can .
Slurm Mega. Day one.
On the first day of Slurm Mega, we overloaded the participants with 4 topics. Pavel Selivanov discussed the process of creating a fault-tolerant cluster from the inside, the operation of Kubeadm, as well as testing and troubleshooting the cluster.

First coffee break. Usually a ‘call for the teacher’, but at Slurm, while the students drink coffee, the instructors continue to answer questions.

And despite the cloud 'Break II' hanging over Pavel Selivanov’s head, he is destined not to leave for a break.

Sergio Bondaryov and Marcel Ibraev are waiting for their turn at the lectern.
During the break, I approached Sergio Bondaryov and asked, 'What advice would you give to all Kubernetes engineers based on your experience working with our clients’ clusters?'
Sergio gave a simple recommendation: 'Restrict internet access to the API server. Because periodically, security threats are found that allow unauthorized users to access the cluster.»
After a few minutes and a bottle of mineral water, Pavel Selivanov charged into the topic of 'Authorization in the cluster using an external provider', specifically LDAP (Nginx + Python) and OIDC (Dex + Gangway).
In the next break, Marcel Ibraev, a speaker at Slurm and Certified Kubernetes Administrator, gave his advice to Kubernetes engineers: 'I may be stating the obvious, but considering how often I come across this, I suspect not everyone keeps it in mind. One should not blindly trust various How-To guides from the Internet that describe how brilliantly a particular solution works. In the context of Kubernetes, this takes on a special significance. Kubernetes is a complex system, and adding a solution to it that hasn't been tested specifically in your project and cluster setup can lead to dire consequences, despite claims of its greatness online. Even Kubernetes itself, without a balanced approach, can harm your project; 'what is good for the Russian is death for the German.' Therefore, we must test, verify, and trial any solution before implementing it in our environment. Only then can you account for all the nuances that may arise.».
Sergiy Bondarev entered the battle after lunch. His topic is Network Policy, specifically an introduction to CNI and Network Security Policy.

The internet is full of articles on Network Policy. Among admins, there's a sentiment that Network Policies may not be necessary, but security professionals highly value this tool and insist that Network Policies be enabled.
Kubernetes' steering wheel has been taken over by Pavel Selivanov with the topic 'Secure and Highly Available Applications in a Cluster.' He has favorite subjects: PodSecurityPolicy, PodDisruptionBudget, LimitRange/ResourceQuota.

Megi's topic, with which Pavel presented at DevOpsConf: .
After discussing how easily a Kubernetes cluster can be hacked, skeptically minded admins say: 'Aha, I told you, your Kubernetes is a leaky mess.' Pavel explains that securing a cluster can be done, and it's not difficult; it's just that the security settings are disabled by default. For details, refer to the transcript. .

— Who broke the cluster? He broke the cluster! I can see it clearly from here!
Things aren't usually simple and easy at Slyorms to avoid boredom. But this time, Telegram decided to show everyone its fifth point:
Marcel Ibraev, [November 22, 2019, 16:52:52]:
Colleagues, there are currently issues with Telegram, please keep this in mind.
The first day, bright and filled with practical knowledge, has come to an end. The second day will feature even more hands-on experience, including launching a database cluster using PostgreSQL, setting up a RabbitMQ cluster, and managing secrets in Kubernetes.

Slurm Mega. Second day.
The host kicked off the second day with an energetic announcement: "This morning, as Pavel put it yesterday, we are in for some serious hardcore stuff. Speaking in the language of surgeons, we will dive into the guts of Kubernetes!"
The mass entertainer is a whole story in itself. One of the challenges of Slurm is that people tune out and fall asleep due to information overload. We've always looked for ways to address this, and last time, small audience games showed great results. This time, we hired a specially trained person. The chat had a lot of jokes about the 'interesting contests,' but the fact remains — we have never seen such lively participants.

Marsel Ibraev received assistance and began studying Stateful applications in the cluster. Specifically, launching a database cluster using PostgreSQL and setting up a RabbitMQ cluster.
After lunch, Sergey Bondarev took over K8S. His topic was 'Storing Secrets.' He was supported by Mulder and Scully. They examined managing secrets in Kubernetes and Vault, as well as 'The truth is out there.'

This continued until late in the evening when Pavel Selivanov spoke about the Horizontal Pod Autoscaler.
Slurm Mega. Third day.
Abruptly and energetically from the very morning, Sergey Bondarev roused the audience with backup and recovery after failures. They personally checked the backup and recovery of the cluster using Heptio Velero and etcd.

Sergey continued the topic of annual certificate rotation in the cluster: renewing control-plane certificates using kubeadm. Just before lunch, to whet participants' appetites or actually ruin them, Pavel Selivanov raised the topic of application deployment.

They discussed templating and deployment tools, as well as deployment strategies.
Pavel Selivanov introduced a new topic: Service Mesh, the installation of Istio. The topic turned out to be so rich that it could warrant a separate intensive session. We're discussing plans, so stay tuned for announcements.
The main thing is to ensure everything works correctly. Because the time for practice has come:
Building CI/CD to deploy the application and update the cluster simultaneously. Everything works well in training projects. But life can be full of surprises.

May the Slurm be with you!
Source: habr.com
