MÀrkus tÔlkes: see translation of the public postmortem from the engineering blog of the company . It describes the issue with conntrack in the Kubernetes cluster that led to partial downtime of some production services.
This article may be useful for those who want to learn a bit more about postmortems or prevent some potential DNS issues in the future.

This is not DNS
It cannot be that this is DNS
This was DNS
A bit about postmortems and processes at Preply
The postmortem describes a failure or any event in production. A postmortem includes a timeline of events, user impact description, root cause, actions, and lessons learned.
At weekly meetings with pizza, in the circle of the technical team, we share various information. One of the most important parts of such meetings are postmortems, which are usually accompanied by a slideshow presentation and a more in-depth analysis of the incident that occurred. Although we do not 'clap' after postmortems, we try to cultivate a 'blameless' culture (). We believe that writing and presenting postmortems can help us (and not only us) in preventing similar incidents in the future, which is why we share them.
Individuals involved in the incident should feel that they can talk about it in detail without fear of punishment or retribution. No blame! Writing a postmortem is not punishment, but a learning opportunity for the entire company.
DNS issues in Kubernetes. Postmortem
KuupÀev: 28.02.2020
Autorid: Amet U., Andrey S., Igor K., Alexey P.
Olek: Completed
LĂŒhidalt: Partial DNS unavailability (26 min) for some services in the Kubernetes cluster
Impact: 15,000 events lost for services A, B, and C
Root Cause: Kube-proxy was unable to correctly remove the old entry from the conntrack table, so some services were still trying to connect to non-existent pods.
E0228 20:13:53.795782 1 proxier.go:610] Failed to delete kube-system/kube-dns:dns endpoint connections, error: error deleting conntrack entries for UDP peer {100.64.0.10, 100.110.33.231}, error: conntrack command returned: ...Trigger: Due to low load inside the Kubernetes cluster, the CoreDNS autoscaler reduced the number of pods in the deployment from three to two.
Lahendus: JĂ€rjekordne rakenduse juurutamine tĂ”i kaasa uute sĂ”lmede loomise, CoreDNS-autoscaler lisas klastrite teenindamiseks rohkem pude, mis pĂ”hjustas conntrack tabeli ĂŒlekirjutamise.
Avastamine: Prometheuse jÀlgimine tuvastas suure hulga 5xx vigu teenustes A, B ja C ning algatas vÀljakutse valveinseneridele.

5xx vead Kibanas
Tegevused
Tegevus
TĂŒĂŒp
Vastutav
Ălesanne
Keela CoreDNS-i automaatne skaleerimine
ennetama.
Amet U.
DEVOPS-695
Paigalda vahemÀlu DNS-server
vÀhendama.
Max V.
DEVOPS-665
Seadista conntrack jÀlgimine
ennetama.
Amet U.
DEVOPS-674
Ăpitud Ă”ppetunnid
Mis lÀks hÀsti:
- JĂ€lgimine toimis sujuvalt. Reaktsioon oli kiire ja organiseeritud.
- Me ei kohanud sÔlmedes mingeid piiranguid.
Mis polnud Ôige:
- Siiani on reaalne pÔhjus teadmata, tundub olevat conntrackis.
- KÔik tegevused parandavad vaid tagajÀrgi, mitte pÔhjuslikku tegurit (viga).
- Me teadsime, et varem vĂ”i hiljem vĂ”ivad meil tekkida DNS-iga seotud probleemid, kuid ei prioriseerinud ĂŒlesandeid.
Kus meil vedas:
- JĂ€rjekordne juurutamine kĂ€ivitas CoreDNS-autoscale'i, mis ĂŒlekirjutas conntrack tabeli.
- See viga puudutas ainult osa teenustest.
Kronoloogia (EET)
Aeg
Tegevus
22:13
CoreDNS-autoscaler vÀhendas pude arvu kolmest kaheni.
22:18
Valveinsenerid hakkasid saama vĂ€ljakutseid jĂ€lgimissĂŒsteemilt.
22:21
Valveinsenerid hakkasid uurima vigade pÔhjust.
22:39
Valveinsenerid alustasid ĂŒhte hiljutist teenust eelnevale versioonile tagasiviimist.
22:40
5xx vead lÔpetasid ilmumise, olukord stabiliseerus.
- Aeg avastamiseni: 4 min
- Aeg tegevuste tegemiseks: 21 min
- Aeg parandamiseks: 1 min
Lisainformatsioon
- [INFO] 10.1.28.1:52495 - 2606 "A IN mrkaran.hello.svc.cluster.local. udp 49 false 512" NXDOMAIN qr,aa,rd 142 0.000524939s [INFO] 10.1.28.1:59287 - 57522 "A IN mrkaran.svc.cluster.local. udp 43 false 512" NXDOMAIN qr,aa,rd 136 0.000368277s [INFO] 10.1.28.1:53086 - 4863 "A IN mrkaran.cluster.local. udp 39 false 512" NXDOMAIN qr,aa,rd 132 0.000355344s [INFO] 10.1.28.1:56863 - 41678 "A IN mrkaran. udp 25 false 512" NXDOMAIN qr,rd,ra 100 0.034629206s
I0228 20:13:53.507780 1 event.go:221] Ăritus(v1.ObjectReference{TĂŒĂŒp:"Deployment", Nimi:"kube-system", Nimi:"coredns", UID:"2493eb55-3dc0-11ea-b3a2-02bb48f8c230", APIVersioon:"apps/v1", Ressursiversioon:"132690686", Valdkonna tee:""}): tĂŒĂŒp: 'Normaalne' pĂ”hjus: 'ReplicaSeti skaleerimine' VĂ€hendatud replica set coredns-6cbb6646c9 kaheks. - Lingid Kibana (vĂ€ljalĂ”ikega), Grafana (vĂ€ljalĂ”ikega)
Protsessori kasutamise minimeerimiseks kasutab Linuxi sĂŒdamik sarnast asja nagu conntrack. LĂŒhidalt öeldes on see utiliit, mis sisaldab NAT-kirjete nimekirja, mis hoitakse spetsiaalses tabelis. Kui jĂ€rgmine paket tuleb samast pude koos sama podiga nagu varem, siis lĂ”pp- IP-aadressi ei arvutata uuesti, vaid vĂ”etakse conntrack tabelist.

Kuidas conntrack töötab
Summary
See oli nĂ€ide meie postmortem'ist koos kasulike linkidega. Konkreetselt selles artiklis jagame teavet, mis vĂ”ib teistele ettevĂ”tetele kasulik olla. Just seetĂ”ttu ei karda me vigu teha ja just seetĂ”ttu tegime ĂŒhe oma postmortem'itest avalikuks. Siin on veel mĂ”ned huvitavad avalikud postmortem'id:
- GitLab:
- Dropbox:
- Spotify:
- Palju teisi siit ja hoidla
- Samuti avaliku postmortemi SRE raamatuga
Allikas: habr.com
