We continue our story about how we changed the BMS system in our Data Centers (, ). Instead of simply swapping one vendor's solution for another, we developed a system from scratch to meet our requirements. In conclusion, we share the results of the work done and interesting solutions that may be useful to you.
better matches the browser interface. Support for AcroForm (form filling will be included later, activated via the setting pdfjs.renderInteractiveForms) has been added.
As the saying goes, seeing is believing.
Racks.
Let’s break down the differences.
- Firstly, it is aesthetically pleasing convenient. Notice how easy it has become to track the loads on the modules ("Banks" or simply "Banks") of the PDU and the total parallel loads of paired modules. On the rack model from the new BMS, we immediately see that the lower paired PDU modules are overloaded (total current exceeds the permissible 16A - the blue notification), while the upper ones are underloaded. In case one of the inputs is turned off, the entire load will shift to the second, and the remaining energized lower module will turn off due to overload. To prevent this, the Data Center support service will inform the client in advance and send recommendations on how to redistribute the load.
- Simple equipment addition. In the new BMS, virtual sensors for summing the current of the modules and the power of the rack are already included in the templates of standard racks and are created automatically after adding to the PDU rack. In the old BMS, they had to be created manually and then dragged onto the map, increasing the likelihood of errors due to the 'human factor'.
- Unlimited scope for creativity. Now we have no limitations when creating virtual sensors. You can build any mathematical models of any variables. This means that we can create complex virtual sensors (previously we could only sum values) and better analyze the statistics and trends of engineering systems' operations. This enhances the quality of decision-making regarding system adjustments, equipment replacement, and resource management.
- Intuitive interface. The new interface is clutter-free, fans are spinning, and switches are clicking. The most convenient feature is the ability to indicate the status of PDU Line A/B within the racks. We tried to create something similar in the old BMS, but the number of overlapping icons per square centimeter on the map forced us to abandon it.
It's now pleasant for the eyes:

Server rooms.

Fragment of the switchgear.

Ventilation control panel.
And the new BMS can also be decorated for the New Year 🙂

One page – mutual understanding with half a word and without specifications
We had long wanted to implement another 'feature' in the BMS: to compile the main parameters of the data center on one page so that one glance at the screen would be enough to assess the status of the main systems. However, we didn’t fully understand what it should look like.
Before we started developing the new BMS, we visited a dozen data centers in the Netherlands. One of the goals was to see examples of such a page.
And in none of the data centers was it shown to us – in some it was absent, in others they were 'currently developing' it, and in others it was a 'big commercial secret.' Therefore, our specifications for creating the new BMS lacked an exact description of this very important page for us.
In the end, we literally came up with it 'on the fly.' At that moment, I had to remotely consult colleagues in the data center. Flipping through the BMS pages on the phone in search of scattered data was very inconvenient, and the first version was sketched out on a napkin. One page. The developers implemented it based on the photo.
Following the example of our cautious Dutch colleagues, we will not demonstrate the final version of our main page, especially since each data center is unique and there is no point in copying. But we will describe the two main principles of its formation:
- It’s a table designed for the format of vertically oriented smartphone screens (or monitors, but maintaining vertical orientation), displaying all important information on one screen. Above the table is a 'summary' of active incidents, so it turned out to be most convenient to place them in a vertical format.
- The arrangement of cells in the table mirrors the architecture of the data center (either physical or logical). We opted against organizing systems alphabetically, as one might assume at first glance. The sequence reflects the visual associations of the data center staff – as if they physically monitor all rooms and systems. This simplifies information retrieval.
Essentially, all key characteristics of the data center are now grouped together and presented on a single screen of a smartphone or the engineer's and manager's monitor, with a connection to the physical and logical topology of the data center.
Here is a photo of that very first draft, although this version was subsequently rethought and improved.

Acknowledgment and incident summary
Let's discuss another new concept for us that emerged as a result of the monitoring system upgrade project.
Acknowledgment is a rather rare term proposed by the developer of the new BMS. It means confirming that the operator has seen the incident, verified it, and taken on the responsibility for its resolution.
The word has taken hold, and now we 'acknowledge' incidents.
The algorithm embedded in the basic version of the new BMS did not satisfy us. Essentially, these were comments to the event log, meaning that resolved incidents did not disappear from the log, and acknowledged ('acknowledged') incidents were not sorted out from new ones.
As a result, a window called 'summary' was developed, which:
- Displays only active incidents and devices in service mode (without commercial 'blue' notifications).
- Clearly separates NEW and ACKNOWLEDGED incidents.
- Indicates who acknowledged the incident.
The algorithm for duty staff in the new BMS is as follows:
- New incidents enter the summary and wait for acknowledgment. They cannot stay in this section for long; the duty person responsible for the equipment must immediately take on the incident.
- The employee acknowledges the incident by clicking the checkmark on the right. Since all employees are under unique accounts, it is automatically displayed who acknowledged the incident. A comment can be left if necessary.
- The incident is moved to the 'Acknowledged' section, and the other on-call staff and the manager understand that the responsible employee is handling the incident.

Example of the summary window with a new and already acknowledged message.
By connecting the summary window with the One page table, we obtained a complete main screen of the BMS system, where you can immediately see:
- the status of the main data center systems;
- the presence of new unprocessed incidents;
- the existence of accepted incidents and information about who is specifically resolving them.
Access through the browser and pop-up notifications on the phone.
The web interface, accessible from any device anywhere in the world, is a striking contrast to the 'thick' client, which is completely closed to external users.
The old approach was fraught with a host of inconveniences, from problems in organizing remote work for monitoring service employees to the need to install 'thick' clients from distributions at the staff's workstations in the data center.
Now every page in BMS has a unique address, allowing for sharing not only the direct address of a page or device but also links to unique graphs/reports.
Access to the system is now carried out through LDAP authentication via Active Directory, enhancing its security level.
Mobility today is a key factor for the effective work of on-call engineers. In addition to monitoring control in the on-call shift room, engineers perform rounds, carry out current work outside the 'on-call' room, and, thanks to the optimized main screen of BMS for mobile displays, do not lose control of what is happening in the machine rooms for even a second.
The quality of control is also improved thanks to the functionality of work chats. They speed up workflows by allowing 'linking' the correspondence of on-call engineers to BMS. For instance, we use the Teams application, which allows internal messaging and receiving all messages from BMS on the phone as pop-up push notifications, freeing the on-call engineer from constantly looking at the phone screen.

Push notification on the smartphone screen.

This is how notifications appear in the Teams application.
Push notifications are set only for incident alerts, thus minimizing distractions. Staff knows that if a Teams push notification appears on their smartphone, they need to go to the BMS page and acknowledge the incident. Messages about incident resolution are tracked on the BMS page.

The photo shows the BMS interface on a smartphone.
In summary
With the cost of upgrading the BMS from our old vendor being comparable to developing a new system from scratch (around $100,000), the difference in product functionality turned out to be enormous. We received a flexible system optimized for our business tasks and processes. We also achieved significant savings in ongoing support and system upgrade costs.
However, there were challenges.
- Firstly, we underestimated the amount of changes required for the base version of the new BMS, and we didn't meet the pre-agreed deadlines. This wasn't a critical issue for us, as we continued using the old system until the end, and the process was creative, complex, and therefore sometimes took longer than expected. Moreover, we always saw that our developer was putting in maximum effort to achieve the best results. However, the overall process turned out to be very long, and our key specialists spent significantly more effort and time on it than planned.
- Secondly, it took us several test phases to fine-tune the virtual machine and communication channel reservation algorithms. Initially, there were failures on both the BMS system side and the configuration of virtual machines and networks. This debugging also took time. Fortunately, the contractor was provided with a testing platform in the form of a cloud service, where all configurations and innovations were initially tested.
- Thirdly, the final system turned out to be more complex for the end user to edit. Previously, the map was just a background (graphic file) with icons, which were easy to change or move, but now it is a complex graphical interface with animations that requires specific skills for editing.
The radical update of our BMS system can already be considered the most important project of last year, which will significantly impact the quality of operational management of our facilities in the future.
Of course, we didn't discard the old metal server; we 'lightened' it: we cleared out thousands of 'commercial' virtual sensors and PDUs, leaving only a few dozen of the most critical devices such as diesel generators, UPS, air conditioners, pumps, leak and temperature sensors. In this mode, it regained its former speed and can serve as a 'backup reserve.' By the way, after removing the PDUs from the old BMS, we freed up about 1,000 now unnecessary licenses; do you happen to know what to do with them?
Source: habr.com
