
Here’s a story that forever changed my approach to DevOps work. Back in the pre-COVID days, long before that time, when my friends and I were just starting our business and freelancing on random projects, I received a proposal in my Telegram.
The company that reached out was involved in data analytics. Every day, they processed thousands of requests. They came to us saying: guys, we have ClickHouse and we want to automate its setup and installation. We want Ansible, Terraform, Docker, and for everything to be stored in Git. We want a cluster of four nodes with two replicas each.
A standard request, of which there are dozens, and we needed just as good a standard solution. We said 'okay', and in 2-3 weeks everything was ready. They accepted the work and began transitioning to the new ClickHouse cluster using our utility.
No one there wanted to deal with ClickHouse or knew how to. At that time, we thought this was their main problem, so the company’s CTO just gave my team a green light to automate the work as much as possible, so they wouldn’t have to dive into it ever again.
We supported the transition, and other tasks emerged — setting up backups and monitoring. At that same moment, the CTO of that company moved on to another project, leaving one of his guys — Leonid — in charge. Leonid wasn't particularly talented. A simple developer who suddenly was made the head of ClickHouse. It seems this was his first appointment to oversee anything, and the sudden honor caused him to develop a bit of a superiority complex.
Together, we got to work on backups. I suggested backing up the raw data right away. Just take it, zip it up, and elegantly throw it into some S3. Raw data is gold. There was also another option — backing up the tables in ClickHouse using freeze and copying. But Leonid came up with his own solution.
He declared that we needed a second ClickHouse cluster. From now on, we would write data to two clusters — the main one and the backup one. I told him, Leonid, this won’t be a backup but an active replica. And if data starts getting lost in production, the same will happen on your backup.
But Leonya firmly gripped the steering wheel and refused to listen to my arguments. We had a long back-and-forth in the chat, but there was nothing to be done — Leonya was steering the project, and we were just hired guys off the street.
We monitored the state of the cluster and charged only for the admins’ work. Pure administration of ClickHouse without delving into the data. The cluster was available, the disks were fine, the nodes were in order.
We still didn’t suspect that we received this order due to a terrible misunderstanding within their team.
The manager was unhappy that ClickHouse was slow and sometimes lost data. He tasked his CTO with figuring it out. The CTO managed to understand as best as he could and concluded that they simply needed to automate ClickHouse — and that was it. But, as it soon turned out, they did not need a DevOps team at all.
All this turned out to be very, very painful. And the most upsetting part was that it was on my birthday.
Friday evening. I had booked a table at my favorite wine bar and called my buddies.
Just before heading out, we got a task to create an alter; we completed it, everything was okay. The alter went through, ClickHouse confirmed it. We were already ready to go to the bar when we got a message saying that data was missing. We checked — it seemed like everything was there. So we left to celebrate.
The restaurant was Fridays noisy. After ordering drinks and food, we settled into the couches. All this time my Slack was gradually filling up with messages. They were writing something about a lack of data. I thought — the morning is wiser than the evening. Especially today.
Closer to eleven, they started calling. It was the company manager… 'Probably decided to congratulate me,' I thought very uncertainly as I picked up the phone.
And I heard something like: 'You've messed up our data! I’m paying you, but nothing is working! You were responsible for the backups and didn’t do anything! Fix it!' — only a lot ruder.
— You know what, go screw yourself! It’s my birthday today, and I'm going to drink, not deal with your junior-level DIY mess of garbage and sticks!
That’s what I didn’t say. Instead, I pulled out my laptop and got to work.
No, I was raging, I was absolutely furious! I bombarded the chat with sarcastic 'I told you so' — because the backup, which was definitely not a backup, of course, saved nothing.
My buddies and I figured out how to manually stop the recording and check everything. We really confirmed that part of the data was not being recorded.
We stopped the recording, counted the number of events that were stored there for the day. We added more data, of which only a third was recorded. Three shards with 2 replicas each. You insert 100,000 rows — 33,000 do not get recorded.
It was complete chaos. Everyone was cursing each other in turn: the first to go there was Leon, followed by me and the company founder. Only the newly joined CTO was trying to steer our calls with shouts and discussions toward finding a solution to the problem.
What was actually happening — no one understood.
My friends and I were simply shocked when we realized that a third of all the data wasn't just not being recorded — it was being lost! It turned out the order in the company was such that after insertion, the data was permanently deleted, and events were being lost in batches. I imagined how Sergey would convert all this into lost rubles.
My birthday was also being sent to the trash. We were sitting in a bar brainstorming ideas, trying to solve the puzzle thrown our way. The reason for ClickHouse's failure was not obvious. Maybe it was the network, maybe it was the Linux settings. Any number of hypotheses had been thrown out.
I hadn't sworn an oath as a developer, but leaving the guys at the other end of the line was unethical — even if they blamed us for everything. I was 99% sure that the problem did not lie in our solutions, not on our side. The 1% chance that we had messed up burned with anxiety. But no matter on whose side the trouble lay — it had to be fixed. Leaving clients, no matter who they were, with such a terrible data leak was too cruel.
We worked at a restaurant table until three in the morning. We were adding events, insert select — and started filling in the gaps. When you lose data, you do it like this — you take the average data from previous days and insert it into the lost data.
After three in the morning, my buddy and I went to my house, ordered some beer from the liquor store. I sat with my laptop and ClickHouse problems, while my friend was telling me something. In the end, after an hour, he got offended that I was working instead of drinking beer with him and left. A classic — being a devops friend for a while.
By 6 a.m., I had recreated the table from scratch, and the data started to flow in. Everything worked without losses.
It became difficult thereafter. Everyone blamed each other for the data loss. If a new bug occurred, I'm sure there would have been a firefight.
In these fights, we finally began to understand — the company thought we were the guys working with data and monitoring table structures. They confused admins with DBAs. And came to ask us, not as admins.
Their main complaint was, what the hell, you were responsible for backups and didn’t do them properly, you messed up the data. And all of this included lots of curses.
I wanted justice. I dug up the correspondence and provided screenshots where Leonid was adamantly making sure such a backup was done. Their CTO sided with us after my phone call. Later, Leonid admitted his fault.
The company head, on the other hand, did not want to blame his people. Screenshots and words had no effect on him. He believed that since we were experts here, we should have convinced everyone and insisted on our solution. Apparently, it was our task to teach Leonid and, on top of that, to go over his head, being the project manager, to reach the top and personally pour out all our doubts about the backup concept.
The chat was oozing with hidden and overt aggression. I didn’t know what to do. Everything reached a deadlock. And then I was advised on the simplest way — to message the manager privately and arrange a meeting with him. Vasya, people aren’t as bold in real life as they are in chat. The boss responded to my message: come over, no problem.
That was the most unsettling meeting of my career. My ally from the client side — the CTO — couldn’t find the time. I was heading to meet the boss and Leonid.
Time after time, I replayed our possible dialogue in my head. I managed to arrive much earlier, half an hour ahead. Nerves kicked in; I smoked 10 cigarettes. I understood, it was all — I was completely alone. I wouldn’t be able to convince them. And stepped into the elevator.
While going up, I scratched my lighter so badly that I broke it.
In the end, Leonid wasn’t at the meeting. And we had a great discussion with the boss! Sergey shared his pain with me. He didn’t want to 'automate ClickHouse' — he wanted 'queries to work.'
I saw not a goat, but a good guy, caring about his business, immersed in work 24/7. The chat often paints villains, scoundrels, and fools. But in reality, they are just people like you.
Sergio did not need a couple of DevOps hires. The issue they faced turned out to be much bigger.
I said I could resolve his problems — but it is a completely different job, and I have an acquaintance who is a DBA for that. If we had known from the start that this was a job for them, we could have avoided a lot. It was late, but we realized that the problem lay in poor data handling, not in the infrastructure.
We shook hands, our pay was raised two and a half times, but on the condition that I take on all the mess with their data and ClickHouse. In the elevator, I contacted that very DBA, Max, and brought him on board. We had to overhaul the entire cluster.
There was a lot of junk in the accepted project. Starting with the mentioned 'backup'. It turned out that this same 'backup' cluster was not isolated. They were testing everything on it, sometimes even pushing it to production.
The in-house developers created their own custom data 'inserter'. It worked like this: it batched files, ran a script, and dumped data into the table. But the main problem was that a huge amount of data was accepted for one simple query. The query was joining data second by second. All for one little number — the total for the day.
The in-house developers improperly used the analytics tool. They were going into Grafana, writing their royal query. It fetched data for two weeks. It produced a beautiful graph. But in reality, the data query was running every 10 seconds. This all piled up in a queue since ClickHouse simply couldn't handle the processing. This was where the main issue lay. In Grafana, nothing worked; the queries were stuck in the queue, constantly receiving old, irrelevant data.
We reconfigured the cluster and revamped the insertion. The in-house developers rewrote their 'inserter', and it began to shard the data correctly.
Max conducted a thorough audit of the infrastructure. He outlined a plan to transition to a full-fledged backend. However, this did not satisfy the company. They were expecting some magical secret from Max that would allow them to operate in the old way, but effectively. Leonid was still responsible for the project, having learned nothing. Of everything proposed, he once again chose his own alternative. As always, it was the most selective… bold decision. Leonid believed that his company had a unique path. A thorny one, full of icebergs.
Actually, that's where we parted ways — we did what we could.
With our bruises and the wisdom gained from this experience, we opened our own business and established several principles. Now we never start work the way we did back then.
After this project, DBA Max joined us, and we continue to work well together. The case with ClickHouse taught us to conduct a complete and thorough audit of the infrastructure before starting work. We dive deep into how everything works, and only then do we accept tasks. If earlier we had rushed to manage the infrastructure immediately, now we first undertake a one-time project that helps us understand how to bring it to a working state.
And yes, we avoid projects with terrible infrastructure. Even for big money, even out of friendship. Managing troubled projects is not profitable. Realizing this has helped us grow. Either a one-time project to tidy up the infrastructure and then a maintenance contract, or we simply pass by. Bypassing yet another iceberg.
P.S. So, if you have any questions about your infrastructure, .
We have 2 free audits available each month; perhaps your project will be among them.
Source: habr.com
