Inspired by the discussion in the chat
Recently, there have been intense debates regarding the definition of DevOps and SRE.
Despite the fact that discussions on this topic have become rather repetitive, including for me, I decided to bring my perspective to the Hub community. For those interested, welcome below the fold. And let the discussions begin anew!
Background
Once upon a time, there existed a distinct team of software developers and server administrators. The former happily wrote code while the latter, using various warm and affectionate words towards the former, configured servers, periodically visiting the developers only to hear the comprehensive reply, 'it works on my machine.' The business was eagerly waiting for the software, while everything stagnated, frequently broke down, and people grew increasingly anxious. Especially the one who was paying for this chaos. A lovely, nostalgic era. But you're already aware of the origins of DevOps.
The Birth of DevOps Practices
Then came serious individuals who stated – this is not the right industry, this cannot continue. And they introduced lifecycle models. For instance, the V-model.

So, what do we see? The business presents a concept, architects design solutions, developers write code, and then – a blackout. Someone tests the product in some way, someone delivers it to the end user, and somewhere at the output of this wonder model sits a lone business customer waiting for the promised outcome. It was concluded that methods were needed to streamline this process. They decided to create practices to implement those methods.
A Brief Digression on What Practice Is
By practice, I understand the combination of technology and discipline. An example would be the practice of describing infrastructure as code using Terraform. Discipline is how to describe infrastructure in code; it exists in the developer's mind, while technology is, in fact, Terraform.
They decided to call them DevOps practices — I think they meant from Development to Operations. They came up with various complicated things — CI/CD practices, practices based on the principle of IaC, thousands of them. And so it began, developers write code, DevOps engineers transform the description of the system in the form of code into functioning systems (yes, code is unfortunately just a description and not the embodiment of the system), delivery rolls in, and so on. Yesterday's administrators, having mastered new practices, proudly retrained as DevOps engineers, and off it went. And there was evening, and there was morning... sorry, not from there.
Everything is not working again as it should.
Just when everything settled down, various crafty 'methodologists' started writing thick books on DevOps practices, silent debates ignited about who the infamous DevOps engineer actually is and that DevOps is a production culture, discontent began to brew once more. Suddenly, it turned out that software delivery is an absolutely non-trivial task. Each development infrastructure has its own stack; somewhere you need to build it, somewhere to deploy the environment, here you need tomcat, there you need a particularly tricky launch method — in general, it’s overwhelming. And strangely enough, the main problem turned out to be the organization of processes — this delivery function, like a bottleneck, began to block the processes. Moreover, operations (Operations) have not been canceled. It is invisible in the V-model, while there is the entire life cycle on the right. Ultimately, you need to maintain the infrastructure somehow, check the monitoring, resolve incidents, and also handle delivery. That is, you have to sit with one foot in development and the other in operations — and suddenly, Development & Operations emerged. Then there was a massive hype over microservices. With them, development started moving from local machines to the cloud — try debugging locally when there are dozens or hundreds of microservices; constant delivery becomes a means of survival. For a 'small modest company', it’s manageable, but what about Google?
SRE from Google
Google came along, consumed the largest cacti, and decided — we don't need that; we need reliability. And reliability needs to be managed. So, they decided — we need specialists who will manage reliability. They called them SRE engineers and said, here you go, do it as you always do, well. Here’s your SLI, here’s your SLO, here’s monitoring. And they pointed to operations. They named their "reliable DevOps" SRE. Seems good, but there's one dirty hack that Google could afford — hiring SR engineers who had qualifications as developers and understood a bit about how operational systems functioned. Moreover, Google itself has issues hiring such people — mainly because they are competing with themselves — someone has to describe the business logic. They handed off delivery to release engineers, while SR engineers manage reliability (of course, not directly, but by influencing the infrastructure, changing architecture, tracking changes and metrics, handling incidents). It looks pretty good, you can . But what if you're not Google, and reliability still worries you?
The Evolution of DevOps Ideas
This is precisely when Docker came along, having grown out of LXC, followed by various orchestration systems like Docker Swarm and Kubernetes, and DevOps engineers breathed a sigh of relief — the unification of practices simplified delivery. It was simplified to such an extent that it even became possible to hand off delivery to developers — deployment.yaml is not much of a challenge. Containerization solves the problem. Plus, the maturity of CI/CD systems has reached a level where you just write one file and everything kicks off — developers can handle it themselves. And here we start talking about how to create our own SRE, starting… well, with anyone.
SRE is not in Google
Well, okay, we've handed off the delivery, it seems we can breathe easy and return to the good old days when admins would monitor CPU loads, tune systems, and quietly sip something mysterious from mugs in peace and quiet... Wait. That’s not why we started this (what a shame!). Suddenly, it turns out that in Google's approach, we can adopt great practices — it’s not CPU loads that matter, nor how often we change disks, or optimize costs in the cloud, but business metrics — those same notorious SLx. Infrastructure management is still our responsibility, incidents need to be resolved, we have to be on duty periodically, and we should generally be knowledgeable about business processes. And guys, start programming at a decent level already; Google is waiting for you.
In summary. Suddenly, but you're already tired of reading and can’t wait to drop a comment for the author. DevOps as a delivery practice has been, is, and will continue to be. It’s not going anywhere. SRE as a set of operational practices makes this delivery successful.
Source: habr.com
