Monitoring + load testing = forecasting and avoiding failures

VTB's IT department has faced emergency situations multiple times in their systems when the load on them increased dramatically. This led to the need to develop and test a model that would predict peak loads on critical systems. To achieve this, the bank's IT specialists configured monitoring, analyzed data, and learned to automate forecasts. In this short article, we will discuss the tools that helped predict the load and whether they successfully optimized operations.

Monitoring + load testing = forecasting and avoiding failures

Issues with high-load services arise in virtually all industries, but they are critical in the financial sector. At the critical moment, all operational units must be ready, and therefore it was necessary to know in advance what might happen and even determine the day when the load would spike and which systems would be affected. It's essential to combat and prevent failures, which is why the need to implement a predictive analytics system was never up for discussion. Systems needed to be modernized based on monitoring data.

Knee-Jerk Analytics

The payroll project is one of the most sensitive in case of a failure. It is also the most predictable, which is why we decided to start with it. Due to high interconnectivity, other subsystems, including remote banking services, could face issues during peak loads. For instance, customers, pleased with SMS notifications of fund arrivals, would start using them actively. The load could surge by more than an order of magnitude. 

The first forecasting model was created manually. We took data from the last year and calculated which days would experience the highest peaks: for example, the 1st, 15th, and 25th, as well as the last days of the month. This model required significant labor input and did not provide an accurate forecast. Nonetheless, it identified bottlenecks where additional 'hardware' was needed and allowed us to optimize the money transfer process by coordinating with anchor clients: instead of distributing salaries in a lump sum, transactions from different regions were staggered over time. Now we process them in batches that the bank's IT infrastructure can handle seamlessly.

After achieving the first positive result, we moved on to automating forecasting. A dozen critical areas were waiting their turn.

Comprehensive approach

VTB implemented a monitoring system from MicroFocus. From there, we took data collection for forecasting, a storage system, and a reporting generation system. Essentially, monitoring was already in place; we just needed to add metrics, a prediction module, and create new reports. This solution is supported by an external contractor, 'Technoserv', so the main project implementation work fell to their specialists, but we built the model ourselves. We developed the forecasting system based on Prophet—this open-source product was developed by Facebook. It is easy to use and integrates well with our existing comprehensive monitoring tools and Vertica. Basically, the system analyzes the load graph and extrapolates it based on Fourier series. There is also the possibility to add certain coefficients by day, taken from our model. Metrics are collected without human participation; the forecast is automatically recalculated once a week, and new reports are sent to recipients. 

This approach reveals key cycles, such as annual, monthly, quarterly, and weekly. Salary payments and advances, vacation periods, holidays, and sales—all influence the number of requests to the systems. It turned out, for example, that some cycles overlap, and the main load (75%) on the systems comes from the Central Federal District. Legal entities and individuals behave differently. While the load from individuals is relatively evenly distributed throughout the week (consisting of many small transactions), for companies, 99.9% occurs during working hours, with transactions that can be brief or take several minutes to even hours.

Monitoring + load testing = forecasting and avoiding failures

Based on the data received, long-term trends are determined. The new system has revealed that people are increasingly turning to remote banking services. This is well known, but we did not expect such a scale, and initially, we did not believe it: the number of visits to bank offices is decreasing rapidly, while the number of remote transactions is growing at the same pace. Accordingly, the load on the systems is also increasing and will continue to grow. We are currently forecasting the load until February 2020. Standard days can be predicted with an accuracy of 3%, while peak days can be predicted with an accuracy of 10%. This is a good result.

Pitfalls

As usual, there were difficulties. The extrapolation mechanism using Fourier series poorly transitions through zero — we know that on weekends, legal entities generate few transactions, but the prediction module outputs values far from zero. We could have corrected them forcibly, but using band-aids is not our method. Additionally, we had to solve the problem of painlessly extracting data from source systems. Regular data collection requires significant computational resources, so we built fast caches using replication, obtaining business data directly from replicas. The absence of additional load on master systems in such cases is a blocking requirement.

New Challenges

The direct task of predicting peaks was solved: there have been no overload-related failures in the bank since May of this year, and the new forecasting system played a significant role in this. Yes, it turned out to be insufficient, and now the bank wants to understand how dangerous peaks are for it. We need forecasts using metrics from load testing, and this is already working for about 30% of critical systems; the rest are in the process of obtaining forecasts. In the next phase, we plan to forecast system loads not in business transaction terms, but in terms of IT infrastructure, i.e., we'll go down a layer. Furthermore, we need to fully automate the collection of metrics and the generation of forecasts based on them, to avoid manual exports. There is nothing remarkable about this — we are simply combining monitoring and load testing according to best global practices.

Source: habr.com

Buy reliable website hosting with DDoS protection, VPS VDS servers 🔥 Buy reliable website hosting with DDoS protection, VPS VDS servers | ProHoster