Imagine that you have a complete server infrastructure: several dozen air conditioning units, a bunch of diesel generators and uninterruptible power supplies. To ensure the hardware works properly, you regularly check its functionality and don't forget about maintenance: you conduct test runs, check the oil levels, and replace parts. Even for a single server, you need to store a lot of information: equipment registry, inventory list, maintenance schedule, as well as warranty documents, contracts with suppliers and contractors.
Now let's multiply the number of halls by ten. Questions about logistics arise. Where should we store what to avoid running around for each spare part? How do we replenish supplies on time so that unexpected repairs don’t catch us off guard? If there is a lot of equipment, keeping all technical tasks in mind is impossible, and managing it on paper is difficult. That's where maintenance management systems, or MMS, come to the rescue.

In MMS, we create schedules for preventive and repair work and store instructions for engineers. Not all data centers have such a system; many consider it too expensive. But from our experience, we have found that it's not the tool that matters, but the approach to information management. We initially created the first system in Excel and gradually refined it into a software product.
Along with we decided to share our experience in developing our own MMS. I will show how the system evolved and how it helped implement the best maintenance practices. Alexey will talk about how he inherited the MMS, what has changed since then, and how the system makes life easier for engineers now.
How We Arrived at Our Own MMS
At first, there were folders.8-10 years ago, information was stored in a disorganized manner. After maintenance, we signed acts of completed work, keeping the paper originals in archives and scanned copies in network folders. Similarly, we gathered information about spare parts, tools, and accessories in folders organized by equipment. It's manageable if you establish a structure and access levels for these folders.
But then you have three problems:
- Navigation: switching between different folders takes a long time. If you want to review repairs for specific equipment over several years, you'll need to make a lot of clicks.
- Statistics: you won't have any, and without it, it's hard to predict how quickly different equipment fails or how much ZIP to plan for the next year.
- Timeliness of response: no one will remind you that components are running low and need to be reordered. It's also not obvious that the same equipment fails not for the first time.
For a while, we stored documents this way, but then we discovered Excel 🙂
MMS in Excel. Over time, the structure of the documentation migrated to Excel. It was based on a list of equipment, linked to maintenance schedules, checklists, and references to completed work reports:

In the equipment list, we indicated the main characteristics and the location in the data center:

It resulted in a kind of navigator, from which you can quickly understand what's happening with the equipment and its maintenance. If needed, you can check individual reports through the links in the maintenance schedule:

If you diligently maintain the document in Excel, the solution is quite suitable for a small server room. But it is also temporary. Even if we use one air conditioner and perform maintenance once a month, over five years we will accumulate hundreds of errors, and our Excel file will bloat. If we add another air conditioner, one diesel generator, and one UPS, we need to create several sheets and interlink them. The longer the history, the more difficult it is to quickly extract the needed information.
The first 'adult' system. In 2014, we underwent our first Management & Operations audit based on the Operational Sustainability standards from Uptime Institute. We mostly used that same Excel file, but over the year we significantly revamped it: we added links to manuals and checklists for engineers. The auditors found this format quite workable. They were able to track all operations with the equipment and ensured that the information was up to date and that processes were well-organized. We passed the audit successfully, scoring 92 out of 100.
The question arose: how to move forward. We decided that we needed a 'serious' MMS, looked at several paid programs, but ultimately decided to develop the software ourselves. We used that very Excel as an expanded technical specification. Here are the tasks we set for the MMS.
What we wanted from the MMS
In most cases, MMS consists of a set of directories and reports. Our directory hierarchy looks something like this:

The very first top-level directory is the list of buildings: machine rooms, warehouses, where the equipment is located.

Next is the list of engineering equipment. We collected it by systems:
- Air conditioning system: air conditioners, chillers, pumps.
- Power supply system: UPS, diesel generator sets, distribution panels.

For each piece of equipment, we gather basic data: type, model, serial number, manufacturer data, year of manufacture, commissioning date, warranty period.
Once we have filled in the list of equipment, we compile a maintenance program: how and how often to perform maintenance. In the maintenance program, we describe the set of operations, for example: replace this battery, adjust the operation of a specific part, and so on. We describe the operations in a separate directory. If an operation appears in different programs, there is no need to describe it each time – we simply take the ready-made one from the directory:

The operations 'Changing temperature settings' and 'Replacing quick-release cable connectors' will be common for chillers and air conditioning systems from the same manufacturer.
Now for each piece of equipment, we can create a maintenance schedule. We link the maintenance program to specific equipment, and the system automatically checks in the program how often maintenance needs to be performed, calculating the time of work from the commissioning date:
Automating the creation of such a schedule can even be done using Excel formulas.
It's not entirely obvious: we keep a separate directory of deferred work. A schedule is one thing, but we are all human and understand that anything can happen. For example, a consumable has not arrived on time and maintenance needs to be postponed by a week. This is a normal situation if monitored. We maintain statistics on deferred and unfulfilled work and strive to ensure that cancellations of maintenance aim to be zero.
Also, for each piece of equipment, we keep statistics on emergencies and unplanned repairs.We use statistics for planning purchases and identifying weak spots in the infrastructure. For example, if a compressor burns out three times in the same location, it signals a need to investigate the causes of breakdowns.

Such a history of maintenance and repairs has accumulated over 4 years for a specific air conditioner.
The next directory is spare parts. Here we account for which consumables are needed for the equipment, where and in what quantities they are stored. This is also where we keep information about delivery times, allowing for better inventory planning.
The quantity of spare parts is calculated based on the annual repair statistics for each unit of equipment. For all spare parts, we indicate a minimum stock: the minimum number of spare parts needed for each site. If the stock of spare parts runs low, its quantity in the directory is highlighted:
The minimum stock of high-pressure sensors must be no less than two, but there is only one left. It is urgent to place an order.
As soon as a batch of spare parts arrives, we fill in the directory with data from the invoice and indicate the storage location. We can immediately see the current stock of such spare parts in the warehouse:

We separately maintain a contact directory. It includes data on suppliers and contractors who perform maintenance:

Each contractor-engineer's card is linked to certificates and electrical safety access groups. When creating the schedule, we can check which specialists have the necessary access.

Since MMS has existed, work with access permits at the site has changed. For example, documents with methodological instructions for maintenance have been added. In the past, the set of operations fit into a small checklist, but detailed instructions now cover everything: how to prepare, what conditions are needed, and so on.
How the whole process is organized now will be illustrated by an example .
How maintenance is performed in MMS
Long ago, completed works were documented only ex post. We simply performed maintenance and signed the acceptance certificate afterward. This is what 99% of server providers do, but from experience, this is insufficient. To ensure nothing is forgotten, we first create a work permit. This is a document describing the work and conditions for their execution. Any maintenance and repair in our system starts with it. Here’s how it works:
- We look at the nearest scheduled work in the maintenance schedule:

- We are creating a new work permit. We select the contractor for maintenance, who oversees the process on our side and coordinates the work with us. We specify where and when the work will take place, choose the type of equipment, and the program we will follow:

- After saving the card, we move on to the details. We specify the performer and check whether they have access to the necessary work. If access is not granted, the field is highlighted in red, and the permit cannot be issued:

- We specify the specific equipment. Depending on the type of work, the maintenance program includes preliminary activities, for example: ordering fuel to the site, planning an introductory briefing for engineers, and notifying colleagues. The list of activities will appear automatically, but we can also add our own items; everything is quite flexible:

- We save the permit, send an email to the approver, and wait for their response:

- When the engineer arrives, we print the permit directly from the system.

- The permit includes a checklist of operations for the maintenance program. The work supervisor in the data center monitors the maintenance and checks off tasks.

For a while, a brief checklist was sufficient. Then we implemented methodological instructions, or MOP (method of procedure). With such a document, any certified engineer can carry out an inspection of any equipment.
Everything is described in detail, down to templates for notification letters and weather conditions:

The printed document looks like this:

According to Uptime Institute standards, such a MOP should be in place for all operations. This results in a considerable amount of documentation. Based on experience, we recommend developing them gradually, for example, one MOP per month.
- After the work is completed, the engineer issues a work completion certificate. We scan it and attach it to the card along with scans of other documents: the work permit and MOP.

- We mark the completed work in the permit:

- The maintenance history is preserved in the equipment card:

We have shown how our system is structured now. However, work on MMS is not finished: several improvements are already planned. For example, we currently store a lot of information in scans. In the future, we plan to make maintenance paperless: connecting a mobile application where the engineer can check off tasks and save information directly in the card.
Of course, there are many ready-made products in the market with similar functions. However, we wanted to demonstrate that even a small Excel file can be developed into a full-fledged product. This can be done independently or by hiring contractors; the key is the right approach. It's never too late to get started.
Source: habr.com












