The question of "how to implement DevOps" has been around for years, but there isn't much quality material available. Sometimes you become a victim of ads from not particularly smart consultants who need to sell their time, no matter how. Other times, it's vague, overly general talk about how the ships of megacorporations sail the universe. The question arises: what does this mean for us? Dear author, can you clearly outline your ideas in a list?
This all stems from the fact that there isn't much real practice and understanding of the outcomes of cultural transformations in companies. Changes in culture are long-term processes, the results of which wonât manifest in a week or even a month. We need someone experienced, who has witnessed how companies have been built and collapsed over the years.

John Willis â one of the fathers of DevOps. John has decades of experience working with a vast number of companies. Recently, he has noticed specific patterns that emerge while working with each of them. Using these archetypes, John guides companies on the true path of DevOps transformation. More about these archetypes can be found in the translation of his talk from the DevOops 2018 conference.

About the speaker:
With over 35 years in IT management, he was involved in creating the predecessor to OpenCloud at Canonical, participated in 10 startups, two of which he sold to Dell and Docker. He is currently the Vice President of DevOps and Digital Practices at SJ Technologies.
Next is a narrative from John.
My name is John Willis, and the easiest way to find me on Twitter is, . I have the same handle on Gmail and GitHub. And you can find video recordings of my talks and their presentations.
I have many meetings with CIOs of various large companies. They often complain that they donât understand what DevOps is, and everyone who tries to explain it talks about something specific to them. Another frequent complaint is that DevOps doesnât work, even though the directors seem to be doing everything as they were told. Weâre talking about large companies that are over a hundred years old. After talking to them, I concluded that for many problems, relatively low-tech solutions are often more suitable than high technologies. For weeks, I simply interacted with people from different departments. What you see in the very first picture of the post is my latest project; this is what the room looked like after three days of work.
What is DevOps?
Indeed, if you ask 10 different people, they will give you 10 different answers. But hereâs whatâs interesting: all ten of these answers will be correct. Thereâs no incorrect answer here. I have been deeply involved in DevOps for about 10 years and was the first American at the first DevOpsDay. I wonât say that Iâm smarter than anyone else involved in DevOps, but itâs unlikely that anyone has spent as much effort on it as I have. I believe that DevOps arises when human capital and technology come together. We often forget about the human dimension, even though we talk a lot about various cultures.

Right now, we have a lot of data, five years of academic research, and theories being validated on an industrial scale. These studies tell us the following: if certain behavioral patterns are integrated into organizational culture, it can lead to a 2000x acceleration. This acceleration corresponds to a similar improvement in resilience. This is a quantitative measure of the advantage that DevOps can provide to any company. A couple of years ago, I presented on DevOps to the CEO of a Fortune 5000 company. When preparing for the presentation, I was very nervous because I needed to convey my years of experience in just 5 minutes.
As a result, I gave the following definition of DevOps: it is a set of practices and patterns that allow transforming human capital into high-performing organizational capital. An example is how Toyota has been operating for the last 50 or 60 years.

(The following schemes are provided not as reference material, but as illustrations. Their content will differ for each new company. However, the image can be viewed separately and enlarged
One of the most successful such practices is value stream mapping. Several good books have been written on this topic, with the author of the most successful among them being Karen Martin. However, over the past year, I've come to the conclusion that even this approach is too high-tech. It certainly has many advantages, and I have used it extensively. But when the CEO asks you why his company cannot switch to new tracks, it's still too early to talk about value stream mapping. There are far more fundamental questions that need to be answered first.
I believe the mistake many of my colleagues make is that they simply provide the company with a five-point guide and then come back six months later to see what happened. Even a good scheme like value stream mapping has, so to speak, blind spots. After hundreds of interviews with directors of various companies, I've developed a certain pattern that allows me to break down the problem into components, and now we will discuss each of these components in turn. Before applying any technological solutions, I use this pattern, and as a result, all my walls end up covered with diagrams. Recently, I worked with a mutual fund, and I ended up with 100-150 such diagrams.
A poor culture eats good approaches for breakfast.
The main idea is that no Lean, Agile, SAFE, or DevOps will help if the organization's culture is poor. It's like diving to depth without a scuba tank or operating without an X-ray. In other words, paraphrasing Drucker and Deming: a poor organizational culture will devour any good system and won't choke.
To tackle this main problem, the following steps need to be taken:
- Make All Work Visible: you need to make all work visible. Not in the sense that it must necessarily be displayed on some screen, but in the sense that it must be observable.
- Consolidate Work Management Systems: It is necessary to consolidate management systems. In the issue of 'tribal' knowledge versus institutional knowledge, in 9 out of 10 cases, the bottleneck is people. In the book the problem was with a single individual, Brent, whose delays set the project back by three years. And I encounter such 'Brents' everywhere. To resolve these bottlenecks, I use the following two points from our list.
- Theory of Constraints Methodology: theory of constraints.
- Collaboration hacks: collaboration hacks.
- Toyota Kata (): I wonât say much about Toyota Kata. If youâre interested, on my GitHub, on almost every one of these topics.
- Market Oriented Organization: market-oriented organization.
- Shift-left auditors: early-cycle audits.

I start working with an organization very simply: I go into the company and talk to the employees. As we can see, no advanced technology is required. All thatâs needed is something to write on. I gather a few teams in one room and analyze what they tell me based on my 7 archetypes. Then I give them a marker and ask them to write down everything they have previously said out loud on the board. Usually, in such meetings, thereâs one person who takes notes, and at best, they manage to capture 10% of the discussion. With my method, this figure can be raised to around 40%.

(This illustration can be )
My approach is based on the work of William Schneider ( ). At the core of this approach is the idea that any organization can be broken down into four quadrants. This scheme usually results from the hundreds of other schemas that arise when analyzing an organization. Suppose we have an organization with high control but low competence. This is a highly undesirable situation: everyone follows the rules, yet no one knows what needs to be done.
A somewhat better option with a high level of control and competence. If such a company is profitable, then it might not need DevOps. It's most interesting to work with a company that has a high level of control, low competence, and collaboration, but at the same time a high level of culture. This means there are many people in the company who enjoy working there, and employee turnover is low.

(This illustration can be )
I believe that methods with rigid recommendations ultimately hinder the pursuit of truth. In particular, in value stream mapping, there are many rules about how to structure information. At the early stages of work, which I am referring to now, these rules are unnecessary. If a person with a marker in hand describes the actual situation in the company on a board, this is the best way to understand whatâs happening. Such information doesnât reach the directors. At this point, it would be foolish to interrupt someone and say they incorrectly drew an arrow. At this stage, it's better to use simple rules; for example, a multi-level abstraction can be created just by using different colored markers.
I repeat, no high technologies. The objective reality of how everything works is depicted with a black marker. With a red marker, people indicate what they do not like about the current state of affairs. It is important that they write this, not me. When I meet with the IT director after the meeting, I do not propose a list of 10 things that need to be fixed. I strive to find connections between what the people in the company say and the existing proven patterns. Finally, possible solutions to the problem are suggested with a blue marker.

(This illustration can be )
An example of this approach is illustrated above. At the beginning of this year, I worked with a bank. Employees from the security department were convinced that they should not attend requirements and design reviews.

(This illustration can be )
Then we spoke with people from other departments and found out that about 8 years ago, software developers sidelined security personnel because they were slowing down work. This then turned into a ban that was perceived as a given. Although in reality, there was no such ban.
Our meeting took a very convoluted path: for about three hours, five different teams could not explain to me what happens between the code and the build. And this seems to be the simplest thing. Most DevOps consultants assume that this is already known to everyone.
Then, the person responsible for IT governance, who had been silent for four hours, suddenly came to life when we reached his topic, and engaged us for quite a while longer. At the end, I asked him what he thought of the meeting, and I will never forget his response. He said, 'I used to think there were only two ways to deliver software in our bank, but now I know there are actually five, and I wasn't even aware of three of them.'

(This illustration can be )
The last meeting at this bank was with the team working on investment software. It turned out that writing diagrams with a marker on a piece of paper is better than on a board, and even better than on a smartboard.

The photos you see show how the hotel conference room looked on the fourth day of our meeting. We used these diagrams to search for patterns, or archetypes.
So, I ask questions to the employees, and they write down the answers with markers in three colors (black, red, and blue). I analyze their answers for archetypes. Now, let's discuss all the archetypes in order.
1. Make All Work Visible: Make work visible
In most companies I work with, a very high percentage of unknown work exists. For example, when one employee approaches another and simply asks for something to be done. In large organizations, up to 60% of the work can be unplanned. Additionally, as much as 40% of work is undocumented. If it were Boeing, I would never fly on their plane again. If only half of the work is documented, it's unclear whether that work is being done correctly. All other methods become useless â there's no point in trying to automate anything because the known 50% could be the most coordinated and clear part of the work, whose automation wouldnât yield significant results, leaving the most chaotic parts hidden in the unseen half. Without documentation, it's impossible to identify various hacks and hidden work, to find bottlenecks, those 'Brent's' I mentioned earlier. There is a great book by Dominica DeGrandis. . It reveals five different 'time thieves' :
- Too Much Work in Process (WIP)
- Unknown Dependencies
- Unplanned Work
- Conflicting priorities
- Neglected Work
This is a very valuable analysis, and the book is excellent, but all these tips are useless if only 50% of the data is visible. The methods proposed by Dominica can be applied only if accuracy above 90% is achieved. I'm talking about situations where a manager assigns a subordinate a 15-minute task, but it actually takes three days; the manager does not realize that this subordinate depends on four or five other people.

The Phoenix Project is a remarkable tale about a project that was three years late. One of the characters faces the threat of being fired because of this, and he meets another character who is presented as a sort of Socrates. This character helps him understand what went wrong. It turns out that there is a sysadmin named Brent in the company, and all work somehow goes through him. During one meeting, one of the subordinates is asked: why does every half-hour task take a week? In response, there is a very simplified explanation of queue theory and Little's Law, revealing that at 90% utilization, every hour of work takes 9 hours. Each task needs to be sent to seven other people, so that hour turns into 63 hours, 7 times 9. I mention this to say that to use Little's Law or any complex queue theory, you need to at least have data.
So when I talk about visibility, I mean not just having everything on the screen, but that it is necessary to at least have data. When it exists, it often turns out that there is a significant amount of unplanned work that somehow gets directed to Brent, even though there is no need for it. And Brent is a great guy; he will never say 'no', but he doesn't tell anyone how he gets his work done.

When work is visible, you can neatly classify data (that's exactly what Dominica is doing in the photo), you can apply the abstraction of the five forms of wasted time and automate.
2. Consolidate Work Management Systems: Task Management
The archetypes I speak of represent a sort of pyramid. If the first is done correctly, then the second is already a sort of superstructure. Many of them do not work for startups; they need to be kept in mind for larger companies, such as those that make the Fortune 5000 list. In the last company I worked at, there were 10 ticketing systems. One team had Remedy, another wrote their own system, a third used Jira, and someone else even managed with email. The same problem arises if a company has 30 different pipelines, but I don't have time to discuss all such cases.
I discuss with people how tickets are created, what happens to them next, and how they are bypassed. The most interesting thing is that people speak quite sincerely at our meetings. I asked how many people mark tickets as 'minor / no impact' when they should actually be marked as 'major impact'. It turned out that almost everyone does this. I donât engage in snitching and try in every possible way not to identify people. When someone confides in me sincerely, I donât reveal who they are. But when almost everyone bypasses the system, it means that the entire security is essentially just a facade. Therefore, we cannot draw any conclusions from the data in this system.
To solve the ticket issue, it is necessary to select one main system. If you are using Jira, let it be only Jira. If there is an alternative, let it be only that. The point is that tickets should be viewed as another stage of the development process. Every action should have a ticket that goes through the development workflow. Tickets are sent to the team that places them on the storyboard, and then they are responsible for them.
This applies to all departments, including the infrastructure and operations. In this case, one can create a plausible representation of the situation. Once this process is established, it suddenly becomes easy to determine who is responsible for each application. Because now we are getting not 50%, but 98% of new services. If this main process works, the accuracy improves across the entire system.
Service Pipeline
This again only concerns large corporations. If youâre a new company in a new field â roll up your sleeves and work with your Travis CI or CircleCI. As for Fortune 5000 companies, a notable case occurred with the bank where I worked. They received a visit from Google, who showed them diagrams of the old IBM systems. The Google team asked in confusion â where is the source code for this? There is no source code, not even a GUI. This is the reality that large organizations must face: 40-year-old banking records on an ancient mainframe. One of my clients is using Kubernetes containers with Circuit Breaker patterns, plus Chaos Monkey, all for the KeyBank application. But ultimately, these containers connect back to a COBOL application.
The Google team was completely confident they would solve all my clientâs problems, but then they started asking questions: what is an IBM datapipe? They were told: itâs a connector. What does it connect to? To the Sperry system. And what is that? And so on. At first glance, it seems: what DevOps can there be? But in reality, it is possible. There are delivery systems that allow workflows to be handed off to delivery teams.
3. Theory of Constraints: Theory of Constraints
Letâs move on to the third archetype: institutional / 'tribal' knowledge. Typically, in any organization, there are a few individuals who know everything and lead the way. These are the ones who have been with the organization the longest and know all the workarounds.

When this is identified on the diagram, I specifically circle such individuals with a marker: for example, it turns out that a person named Lou is present at all meetings. And for me, it's clear: this is the local Brent. When the IT director chooses between me in a t-shirt and sneakers and a suit-clad guy from IBM, they pick me because I can tell the director about things that the other guy won't address and that may be uncomfortable for the director to hear. I inform them that thereâs a bottleneck in their company, someone named Fred and someone named Lou. This bottleneck needs to be untangled, their knowledge needs to be extracted one way or another.
To tackle such a problem, I might suggest using Slack. A savvy director would ask â why? Typically, in these cases, DevOps consultants respond: because everyone else does. If the director is truly astute, they would say: so what. And that would be the end of the discussion. My response would be: because there are four bottlenecks in the company, Fred, Lou, Suzie, and Jane. To institutionalize their knowledge, you first need to introduce Slack. All your wikis are a complete mess because no one knows they exist. If the engineering team is engaged in both external and internal development, everyone should be aware that they can reach out to either the external development team or the infrastructure team for inquiries. Thatâs when Lou or Fred might have time to connect to the wiki. Then someone on Slack might ask why step 5 isnât working, and at that point, Lou or Fred will update the instructions in the wiki. If this process is established, many things will start to fall into place on their own.
This is my main point: to recommend any advanced technologies, you first need to tidy up the foundation for them, and this can only be achieved through the low-tech solutions I just described. If you start with advanced technologies without explaining their purpose, it usually ends poorly. One of our clients uses Azure ML, a very cheap and simple solution. About 30% of their inquiries were already answered by the self-learning machine. This was developed by operators who did not have backgrounds in data science, statistics, or mathematics. That speaks volumes. The cost of such a solution is minimal.
4. Collaboration hacks: Collaboration hacks
The fourth archetype states that it is necessary to combat isolation. Most people already know this: isolation breeds hostility. If each department is on its own floor, and people only interact with each other in the elevator, hostility can easily arise between them. Conversely, when people are in the same room, that hostility quickly dissipates. When someone makes a common accusation, such as a certain interface never working, itâs easy to deconstruct that accusation. The programmers who wrote the interface can simply start asking specific questions, and soon it will become clear that, for example, the user was just using the tool incorrectly.
There are many ways to overcome isolation. Some time ago, I was asked to consult for a bank in Australia, but I declined because I have two kids and a wife. All I could help them with was to recommend graphical storytelling. This is something that has been proven to work. Another interesting approach is lean coffee meetings. In a large organization, this is an excellent way to spread knowledge. Additionally, internal devopsdays, hackathons, and similar events can be organized.
5. Coaching Kata
As I already warned at the very beginning, I won't discuss this today. If you're interested, you can take a look at .
There's also a good talk on this topic by Mike Rother:

6. Market Oriented: market-oriented organization
Here, there are different problems. For instance, "I" people, "T" people, and "E" people. "I" people focus only on one thing. They usually exist in organizations with isolated departments. "T" refers to someone who knows one thing well but also excels in several other areas. "E" or even "comb" refers to a person with many skills.

Here Conway's law applies (), which can be simplified as follows: if three teams work on a compiler, the end result will be a compiler made up of three parts. Therefore, if there is a high level of isolation within an organization, even trendy things like Kubernetes, Circuit Breaker, API extensibility, and others will be shaped just like the organization itself. Strictly according to Conway, much to the disdain of all you young geeks.
This problem has been described many times. For example, there are organizational archetypes described by Fernando Fernandez. The problematic architecture I just mentioned, with isolationâthis is a function-oriented architecture. The second type, which is worse, is the matrix architecture, where thereâs a jumble of the other two. The third type is what is seen in most startups, and large companies are also trying to align with this type. This is a market-oriented organization. Here, the focus is on optimizing for the fastest response to customer requests. Sometimes this is called a flat organization.
Many describe this structure differently; I like the phrasing build/run teams, in Amazon this is called two pizza teams. In this structure, all people of the type 'I' group around a single service, and gradually they get closer to the 'T' type, and if proper management is in place, can even become 'E'. The first counterargument here is that such a structure has unnecessary elements. Why have a tester in every department when you can have a dedicated testing department? To which I respond: unnecessary expenses in this case are the price to pay for the organization to eventually become type 'E'. In such a structure, the tester gradually learns about networks, architecture, design, etc. Ultimately, each member of the organization becomes fully aware of everything that happens within the organization. If you want to understand how this scheme works in industry, read .
7. Shift-left auditors: auditing at early stages of the cycle. Compliance with security rules in plain sight.
This is when your actions fail, so to speak, the smell test. The people working for you aren't fooling around. If, as in the example above, they were categorizing everything as minor/no impact, and this continued for three years without anyone noticing, then everyone knows that the system isnât functioning. Or another example is the change advisory board, where reports need to be submitted every Wednesday, say. A group of people works there (by the way, not particularly well-paid) who theoretically should know how the system works as a whole. And over the past five years, youâve probably noticed that our systems are incredibly complex. And five or six people must make a decision about a change that they didn't implement and know nothing about.
Of course, such an approach doesnât work. I have to eliminate such things because these people do not safeguard the system. The decision should be made by the team itself, because the team must be accountable for it. Otherwise, a paradoxical situation arises where a manager, who has never written a line of code in their life, tells the programmer how long it should take to write the code. In one company I worked for, there were seven different boards that reviewed every change, including architecture, product boards, and so on. There was even a mandatory waiting period, although one employee told me that in ten years, no one had ever denied a change proposed by that person during this mandatory period.
You need to invite auditors rather than getting rid of them. Tell them that you are writing immutable binary containers that, if they pass all tests, remain unchanged forever. Explain to them that you have a pipeline as code and clarify what that means. Show them the following scheme: an immutable binary is read-only within a container that passes all vulnerability tests; and after that, not only does no one touch itâthe system that creates the pipeline doesn't get touched either, as it is also created dynamically. I have clients like Capital One, who use Vault to create something akin to a blockchain. Thereâs no need to show the auditor ârecipesâ from Chef; itâs enough to show the blockchain, which clearly indicates what happened with the Jira ticket in production and who is responsible for it.

According to , created in 2018 by Sonatype, there were 87 billion download requests for OSS in 2017.

The losses incurred due to vulnerabilities are exceedingly high. Moreover, the figures you see above do not include alternative costs. A brief word about what DevSecOps is. Right away, I want to make it clear that I am not interested in conversations about how successful this name is. The point is that, since DevOps have been quite successful, we need to try adding security to this pipeline.
An example of such a sequence:

This is not an endorsement of specific products, though I like all of them. I've mentioned them as an example to demonstrate that DevOps, originally based on the paradigm of industrial organization, allows automation of every stage of product development.

And thereâs no reason why we couldnât apply the same approach to security.
Summary
In conclusion, I would like to offer some advice for DevSecOps. Itâs crucial to involve auditors in the process of creating your systems and invest time in their education. Collaborating with auditors is essential. Furthermore, it is important to conduct a ruthless fight against false positives. Even with the most expensive vulnerability scanning tools, you can end up creating harmful habits among your developers if you do not understand the signal-to-noise ratio. Developers will become overwhelmed with alerts and may simply start ignoring them. If youâve heard about the Equifax incident, thatâs exactly what happened; the highest priority signals were overlooked. Additionally, vulnerabilities should be explained in a way that clarifies how they affect the business. For example, you could say itâs the same vulnerability as in the Equifax case. Security-related vulnerabilities should be treated just like other software issues, meaning they need to be incorporated into the overall DevOps process. They should be managed through Jira, Kanban, etc. Developers should not assume that someone else will take care of this â on the contrary, it should be a collective responsibility. Lastly, efforts should be dedicated to educating people.
Useful links
Here are some reports from the DevOops conference that you might find useful:
- Sergei Berdnikov, Artyom Kalichkin â Success Story, or 'Dev+DevOps+Ops' (, )
- Baruch Sadogursky, Leonid Igolnik â DevOps at Scale: A Greek Tragedy in Three Acts (, )
- Alexander Titov, Kirill Tolkachyov â
- Timothy Lister â
Check out DevOops 2020 Moscow â thereâs a lot of interesting material there as well.
Source: habr.com
