DNS-i probleemid Kuberneteses. Avalik postmortem

MÀrkus tÔlkes: see translation of the public postmortem from the engineering blog of the company Preply. It describes the issue with conntrack in the Kubernetes cluster that led to partial downtime of some production services.

This article may be useful for those who want to learn a bit more about postmortems or prevent some potential DNS issues in the future.

DNS-i probleemid Kuberneteses. Avalik postmortem
This is not DNS
It cannot be that this is DNS
This was DNS

A bit about postmortems and processes at Preply

The postmortem describes a failure or any event in production. A postmortem includes a timeline of events, user impact description, root cause, actions, and lessons learned.

Seeking SRE

At weekly meetings with pizza, in the circle of the technical team, we share various information. One of the most important parts of such meetings are postmortems, which are usually accompanied by a slideshow presentation and a more in-depth analysis of the incident that occurred. Although we do not 'clap' after postmortems, we try to cultivate a 'blameless' culture (blameless culture). We believe that writing and presenting postmortems can help us (and not only us) in preventing similar incidents in the future, which is why we share them.

Individuals involved in the incident should feel that they can talk about it in detail without fear of punishment or retribution. No blame! Writing a postmortem is not punishment, but a learning opportunity for the entire company.

Keep CALMS & DevOps: S is for Sharing

DNS issues in Kubernetes. Postmortem

KuupÀev: 28.02.2020

Autorid: Amet U., Andrey S., Igor K., Alexey P.

Olek: Completed

LĂŒhidalt: Partial DNS unavailability (26 min) for some services in the Kubernetes cluster

Impact: 15,000 events lost for services A, B, and C

Root Cause: Kube-proxy was unable to correctly remove the old entry from the conntrack table, so some services were still trying to connect to non-existent pods.

E0228 20:13:53.795782       1 proxier.go:610] Failed to delete kube-system/kube-dns:dns endpoint connections, error: error deleting conntrack entries for UDP peer {100.64.0.10, 100.110.33.231}, error: conntrack command returned: ...

Trigger: Due to low load inside the Kubernetes cluster, the CoreDNS autoscaler reduced the number of pods in the deployment from three to two.

Lahendus: JĂ€rjekordne rakenduse juurutamine tĂ”i kaasa uute sĂ”lmede loomise, CoreDNS-autoscaler lisas klastrite teenindamiseks rohkem pude, mis pĂ”hjustas conntrack tabeli ĂŒlekirjutamise.

Avastamine: Prometheuse jÀlgimine tuvastas suure hulga 5xx vigu teenustes A, B ja C ning algatas vÀljakutse valveinseneridele.

DNS-i probleemid Kuberneteses. Avalik postmortem
5xx vead Kibanas

Tegevused

Tegevus
TĂŒĂŒp
Vastutav
Ülesanne

Keela CoreDNS-i automaatne skaleerimine
ennetama.
Amet U.
DEVOPS-695

Paigalda vahemÀlu DNS-server
vÀhendama.
Max V.
DEVOPS-665

Seadista conntrack jÀlgimine
ennetama.
Amet U.
DEVOPS-674

Õpitud Ă”ppetunnid

Mis lÀks hÀsti:

  • JĂ€lgimine toimis sujuvalt. Reaktsioon oli kiire ja organiseeritud.
  • Me ei kohanud sĂ”lmedes mingeid piiranguid.

Mis polnud Ôige:

  • Siiani on reaalne pĂ”hjus teadmata, tundub olevat spetsiifiline viga conntrackis.
  • KĂ”ik tegevused parandavad vaid tagajĂ€rgi, mitte pĂ”hjuslikku tegurit (viga).
  • Me teadsime, et varem vĂ”i hiljem vĂ”ivad meil tekkida DNS-iga seotud probleemid, kuid ei prioriseerinud ĂŒlesandeid.

Kus meil vedas:

  • JĂ€rjekordne juurutamine kĂ€ivitas CoreDNS-autoscale'i, mis ĂŒlekirjutas conntrack tabeli.
  • See viga puudutas ainult osa teenustest.

Kronoloogia (EET)

Aeg
Tegevus

22:13
CoreDNS-autoscaler vÀhendas pude arvu kolmest kaheni.

22:18
Valveinsenerid hakkasid saama vĂ€ljakutseid jĂ€lgimissĂŒsteemilt.

22:21
Valveinsenerid hakkasid uurima vigade pÔhjust.

22:39
Valveinsenerid alustasid ĂŒhte hiljutist teenust eelnevale versioonile tagasiviimist.

22:40
5xx vead lÔpetasid ilmumise, olukord stabiliseerus.

  • Aeg avastamiseni: 4 min
  • Aeg tegevuste tegemiseks: 21 min
  • Aeg parandamiseks: 1 min

Lisainformatsioon

  • [INFO] 10.1.28.1:52495 - 2606 "A IN mrkaran.hello.svc.cluster.local. udp 49 false 512" NXDOMAIN qr,aa,rd 142 0.000524939s [INFO] 10.1.28.1:59287 - 57522 "A IN mrkaran.svc.cluster.local. udp 43 false 512" NXDOMAIN qr,aa,rd 136 0.000368277s [INFO] 10.1.28.1:53086 - 4863 "A IN mrkaran.cluster.local. udp 39 false 512" NXDOMAIN qr,aa,rd 132 0.000355344s [INFO] 10.1.28.1:56863 - 41678 "A IN mrkaran. udp 25 false 512" NXDOMAIN qr,rd,ra 100 0.034629206s
    I0228 20:13:53.507780 1 event.go:221] Üritus(v1.ObjectReference{TĂŒĂŒp:"Deployment", Nimi:"kube-system", Nimi:"coredns", UID:"2493eb55-3dc0-11ea-b3a2-02bb48f8c230", APIVersioon:"apps/v1", Ressursiversioon:"132690686", Valdkonna tee:""}): tĂŒĂŒp: 'Normaalne' pĂ”hjus: 'ReplicaSeti skaleerimine' VĂ€hendatud replica set coredns-6cbb6646c9 kaheks.
  • Lingid Kibana (vĂ€ljalĂ”ikega), Grafana (vĂ€ljalĂ”ikega)
  • Kus Linuxi conntrack ei ole enam sinu sĂ”ber
  • kube-proxy nĂŒansid: juhendamine katkestatud ĂŒhenduse lĂ€htestamisel
  • Kihiline conntrack ja DNS-i pĂ€ringu aegumise probleemid

Protsessori kasutamise minimeerimiseks kasutab Linuxi sĂŒdamik sarnast asja nagu conntrack. LĂŒhidalt öeldes on see utiliit, mis sisaldab NAT-kirjete nimekirja, mis hoitakse spetsiaalses tabelis. Kui jĂ€rgmine paket tuleb samast pude koos sama podiga nagu varem, siis lĂ”pp- IP-aadressi ei arvutata uuesti, vaid vĂ”etakse conntrack tabelist.
DNS-i probleemid Kuberneteses. Avalik postmortem
Kuidas conntrack töötab

Summary

See oli nĂ€ide meie postmortem'ist koos kasulike linkidega. Konkreetselt selles artiklis jagame teavet, mis vĂ”ib teistele ettevĂ”tetele kasulik olla. Just seetĂ”ttu ei karda me vigu teha ja just seetĂ”ttu tegime ĂŒhe oma postmortem'itest avalikuks. Siin on veel mĂ”ned huvitavad avalikud postmortem'id:

Allikas: habr.com

Osta usaldusvÀÀrne veebimajutus DDoS-kaitsega veebisaitidele, VPS VDS serverid đŸ”„ Osta usaldusvÀÀrne veebimajutus DDoS-kaitsega veebisaitidele, VPS VDS serverid - ProHoster