Why does a corporation like MegaFon need Tarantool in billing? From the outside, it seems that typically a vendor comes in, brings a big box, plugs it into the socket—and voilà, billing! That used to be the case, but now it's archaic, and such dinosaurs have either died out or are dying. Initially, billing was a system for invoicing—a calculator or counter. In modern telecom—it's a system that automates the entire lifecycle of interaction with the subscriber, from contract signing to termination, including real-time billing, payment processing, and much more. Billing in telecom companies resembles a battle robot—large, powerful, and equipped with weapons.

So what is the role of Tarantool here? This will be explained by Oleg Ivlev and Andrey Knyazev. Oleg is the chief architect of the company with extensive experience working for foreign companies, and Andrey is the director of business systems. From their presentation at the , you will learn why R&D is needed in corporations, what Tarantool is, how the dead end of vertical scaling and globalization led to the emergence of this database in the company, and about technological challenges, architectural transformation, and how MegaFon's tech stack is similar to that of Netflix, Google, and Amazon.

The "Unified Billing" project
The project being discussed is called “Unified Billing.” It is in this project that Tarantool has demonstrated its best qualities.

The growth of Hi-End equipment performance could not keep up with the growth of the subscriber base and the number of services, and a further increase in subscribers and services was expected due to M2M, IoT, and the particularities of the branches leading to a decline in time-to-market. The company decided to create a unified business system with a unique modular architecture of world level, replacing eight current different billing systems.
MegaFon is eight companies in one. In 2009, the reorganization was completed: branches across Russia merged into a single company, OJSC "MegaFon" (now PJSC). Thus, the company had eight billing systems with their own custom solutions, branch-specific features, and varying organizational structure, IT, and marketing.
Everything was fine until we had to launch a common federal product. A lot of difficulties arose: some had billing rounded up, others rounded down, and some did it based on the arithmetic mean. There were thousands of such instances.
Despite having one billing system version and one supplier, the settings diverged to the extent that it took a long time to patch them together. We tried to reduce their number and encountered a second problem familiar to many corporations.
Vertical scaling. Even the best hardware at the time did not meet our needs. We used Hewlett-Packard equipment from the Superdome Hi-End line, but it couldn't support the requirements of even two branches. We wanted horizontal scaling without significant operational costs and capital investments.
Expectations of growing the number of subscribers and services. Consultants have long been bringing stories about IoT and M2M to the telecom world: times will come when every phone and iron will have a SIM card, and every refrigerator will have two. Today we have a certain number of subscribers, but in the near future, there will be an order of magnitude more.
Technological challenges
These four reasons drove us to significant changes. We faced a choice between modernizing the system and designing from scratch. We thought for a long time, made serious decisions, and held tenders. Ultimately, we decided to design from the ground up and took on interesting challenges — technological challenges.
Scalability
If before, let's say, there were 8 billings for 15 million subscribers, now it had to be 100 million subscribers and more — the load is an order of magnitude higher.
We became comparable in scale to large internet players like Mail.ru or Netflix.
But further movement towards increasing the load and subscriber base presented us with serious tasks.
The geography of our vast country
Between Kaliningrad and Vladivostok 7500 km and 10 time zones. The speed of light is finite, and at such distances, delays are already significant. 150 ms on the best modern optical channels is quite a lot for real-time billing, especially as it is currently in telecom in Russia. Additionally, updates need to be made within one working day, which is problematic with different time zones.
We don't just offer subscription services; we have complex plans, packages, and various modifiers. We need to not only enable or disable a subscriber's calls but also provide a specific quota — to account for calls and actions in real-time so that they do not notice.
Fault tolerance
This is the flip side of centralization.
If we gather all subscribers in one system, any emergency events and disasters can have dire consequences for the business. This is why we design the system to eliminate the impact of emergencies on the entire subscriber base.
This is a consequence again of refusing vertical scaling. When we shifted to horizontal scaling, we increased the number of servers from hundreds to thousands. They need to be managed, built for interchangeability, automatically back up the IT infrastructure, and recover the distributed system.
Such interesting challenges lay ahead of us. We designed the system, and at that moment we tried to find global best practices to check how on-trend we are and how closely we follow the latest technologies.
Global experience
Surprisingly, we found no references in global telecom.
Europe fell short in subscriber numbers and scale, the USA due to the flatness of its tariffs. We looked at something in China and found some insights in India, recruiting specialists from Vodafone India.
To analyze the architecture, we assembled a Dream Team led by IBM — architects from various fields. These individuals could adequately assess what we are doing and bring certain knowledge to our architecture.
Scale
A few numbers for illustration.
We design the system for 80 million subscribers with a buffer for one billion. This way we eliminate future thresholds. It is not because we plan to conquer China, but due to the surge of IoT and M2M.
300 million documents are processed in real time. Although we have 80 million subscribers, we also work with potential clients and those who have left us, if we need to collect overdue debts. Therefore, the actual volumes are significantly higher.
2 billion transactions change the balance daily — these are payments, accruals, calls, and other events. 200 TB of data changes actively, slightly slower change 8 PB of data., and this is not an archive, but live data in a single billing system. The scale across and data centers is 5 thousand servers at 14 locations.
Technology stack
When we planned the architecture and began to assemble the system, we imported the most interesting and advanced technologies. The result is a technology stack familiar to any player in the internet and corporations that build high-load systems.

The stack is similar to those of other major players: Netflix, Twitter, Viber. It consists of 6 components, but we want to reduce and unify it.
Flexibility is good, but in a large corporation, unification is essential.
We are not going to replace Oracle with Tarantool. In the realities of large companies, this is utopia, or a crusade lasting 5-10 years with an uncertain outcome. However, Cassandra and Couchbase can indeed be replaced with Tarantool, and we are striving for that.
Why Tarantool?
There are 4 simple criteria why we chose this database.
Speed. We conducted load tests on the industrial systems of MegaFon. Tarantool won - it showed the best performance.
It cannot be said that other systems do not meet the needs of MegaFon. Current memory solutions are so efficient that they provide ample capacity for the company. However, we are keen on dealing with a leader rather than one who lags behind, including in load testing.
Tarantool meets the company’s needs even in the long-term perspective.
TCO (Total Cost of Ownership). Supporting Couchbase at the volumes of MegaFon costs astronomical amounts, while the situation with Tarantool is much more favorable, and their functionality is similar.
Another nice feature that slightly influenced our choice is that Tarantool works better with memory than other databases. It shows maximum efficiency.
Reliability. MegaFon invests in reliability, perhaps like no other. Therefore, when we looked at Tarantool, we realized we needed to make it meet our requirements.
We invested our time and finances and, together with Mail.ru, created an enterprise version that is now used in several other companies.
Tarantool-enterprise completely satisfied us in terms of security, reliability, and logging.
Partnership
The most important thing for me is direct contact with the developer. This is exactly what won us over from the guys at Tarantool.
When you approach a player, especially one working with an anchor client, and say that you need the database to do this, that, and the other, they usually respond:
— Sure, put the requirements at the bottom of the stack — we might get to them someday.
Many have a roadmap for the next 2-3 years, and it's nearly impossible to fit in, whereas the Tarantool developers are open and not just with MegaFon, adapting their system to the client. It’s great, and we really like it.
Where we applied Tarantool
We use Tarantool in several components. The first is in the pilot, which we created on the address catalog system. At one time, we wanted it to be a system similar to Yandex Maps and Google Maps, but it turned out somewhat differently.
For example, the address catalog in the sales interface. On Oracle, finding the required address takes 12-13 seconds — uncomfortable figures. When we switch to Tarantool, replacing Oracle with another database in the console and running the same search, we achieve a 200 times speedup! The city pops up after the third letter. We are currently adapting the interface to make this happen after the first letter. Nevertheless, the response speed is completely different — now it takes milliseconds instead of seconds.
The second application is the trendy topic called two-speed IT. This is because consultants from every angle are saying that corporations should go there.

There is a layer of infrastructure, above which are domains, for example, a billing system like in telecom, corporate systems, corporate reporting. This is the core that should not be disturbed. Of course, you can, but you should be paranoid about ensuring quality, since this brings revenue to the corporation.
Next comes the microservices layer — what differentiates the operator or another player. Microservices can be quickly created based on certain caches, pulling data from various domains. Here lies the field for experimentation — if something doesn't work out, you shut down one microservice and open another. This truly enhances time-to-market and increases the reliability and speed of the company.
Microservices are probably Tarantool's main role at MegaFon.
Where we plan to apply Tarantool
In comparison to our successful billing project and transformation programs at Deutsche Telekom, Связьком and Vodafone India, it is remarkably dynamic and creative. During the implementation of this project, not only MegaFon and its structure underwent transformation but also Tarantool-enterprise emerged at Mail.ru, and our vendor Nexign (formerly 'Петер-Сервис') developed BSS Box (a boxed billing solution).
This project is, in a sense, historic for the Russian market. It can be compared to what is described in Frederick Brooks' book 'The Mythical Man-Month'. Back in the 1960s, 5,000 people were involved in developing the new operating system OS/360 for IBM mainframes. We have fewer—1,800—yet our team is skilled, and with the use of open source and new approaches, we are working more efficiently.
Below are the billing domains or, more broadly, the business systems. People from enterprises are well aware of CRM. Other systems should already be adopted by everyone: Open API, API Gateway.

Open API
Let’s take another look at the numbers and how Open API operates now. Its load is 10,000 transactions per second. Since we plan to actively develop the microservices layer and build MegaFon’s public API, we expect further growth in this area in the future. 100,000 transactions will definitely be the target..
I don’t know if we will compare in SSO with Mail.ru—they seem to have 1,000,000 transactions per second. Their solution is extremely interesting to us and we plan to adopt their experience—for instance, by creating a functional SSO reserve using Tarantool. Currently, developers from Mail.ru are working on this for us.
CRM
CRM is about those 80 million subscribers we want to grow to a billion because there are already 300 million documents, which include a three-year history. We are genuinely looking forward to new services, and here the growth point is connected services. This is a sphere that will continue to expand as services keep increasing. Therefore, we will need a history, and we do not want to stumble on that.
The billing itself, regarding invoice issuance and working with the clients' accounts receivable has transformed into a separate domain. In order to enhance performance, an architectural template of domain architecture has been applied..
The system is divided into domains, the load is distributed, and redundancy is ensured. Additionally, work has been done on the distributed architecture.
Everything else is enterprise-level solutions. In the call storage — 2 billion per day, 60 billion per month. Sometimes we need to recount them for the month, and it's better to do it quickly. Financial monitoring — is exactly those 300 million that keep growing: subscribers often switch between operators, increasing this share.
The most telecom-like component of mobile communication is online billing.These are the systems that allow you to make calls or not, making decisions in real time. The load here is 30,000 transactions per second, but with the growing data transfer we plan 250,000 transactions, and therefore we are very interested in Tarantool.
The previous image shows the domains where we plan to apply Tarantool. The CRM itself, of course, is broader, and we plan to implement it at the core.
My estimated figure of 100 million subscribers concerns me as an architect — what if it becomes 101 million? Will we have to redo everything again? To prevent this, we use caches, thereby increasing availability.

In general, there are two approaches to using Tarantool. The first is to build all caches at the microservices level.As far as I understand, this is the path VimpelCom is taking, creating a client cache.
We are less dependent on vendors, changing the BSS core, so we already have a single customer database out of the box. But we want to expand it. Therefore, we apply a slightly different approach — we make caches within the systems..
This reduces desynchronization — one system is responsible for both the cache and the main master source.
This method fits well with Tarantool's transactional skeleton, where only the parts related to updates, that is, data changes, are refreshed. Everything else can be stored somewhere else. There is no huge data lake or unmanaged global cache. Caches are designed for the system, or for products, or for clients, or to make life easier for maintenance. When a subscriber upset with quality calls, you want to provide quality service.
RTO and RPO
In IT, there are two terms — RTO and RPO.
Recovery time objective — this is the service recovery time after a failure. RTO = 0 means that even if something goes down, the service continues to operate.
Recovery Point Objective — this is the data recovery time, how much data we can lose over a specific period. RPO = 0 means that we do not lose data.
Task for Tarantool
Let's try to solve a task for Tarantool.
Given: a commonly understood shopping cart, for example, on Amazon or elsewhere. Requirement is for the cart to work 24 hours a day, 7 days a week, or 99.99% of the time. Orders coming to us must maintain order, because we cannot chaotically connect or disconnect communication for the subscriber — everything must be strictly sequential. The previous subscription affects the next one, so data is important — nothing should be lost.
Solution. One could try to tackle it directly and ask the database developers, but the task is mathematically unsolvable. One might recall theorems, laws of conservation, quantum physics, but why — it cannot be solved at the database level.
Here, the good old architectural approach works — it is necessary to have a good understanding of the subject area to resolve this puzzle.

Our solution: we create a distributed registry of applications on Tarantool — a geo-distributed cluster. In the diagram, this shows three different data centers — two before the Ural Mountains, one behind the Ural Mountains, and we distribute all applications among these centers.
At Netflix, which is now considered one of the leaders in IT, there was only one data center until 2012. On the eve of Christmas, December 24th, this data center went down. Users in Canada and the US were left without their favorite movies, got quite upset, and wrote about it on social media. Now Netflix has three data centers on the west-east coast and one in Western Europe.
We are initially building a geo-distributed solution — fault tolerance is important to us.
So, we have a cluster, but how do we deal with RPO = 0 and RTO = 0? The solution is simple, depending on the subject area.
What is important in applications? Two parts: creating the cart BEFORE the decision to purchase, and AFTER. The BEFORE part in telecom is usually referred to as order capturing or order negotiationIn telecommunications, this can be much more complex than in an online store, because you need to serve the customer, offer 5 options, and this all takes some time, but the cart is being filled. At this moment, a failure may occur, but that’s not a big deal because it happens interactively under human supervision.
If the Moscow data center suddenly goes down, we will continue to operate by switching automatically to another data center. Theoretically, one item in the cart may get lost, but you can see that and refill the cart to continue working. In this case, RTO = 0.
At the same time, there’s a second option: when we hit 'submit', we want to ensure the data is not lost. From this point, automation kicks in — this is already RPO = 0. The application of these two different patterns in one case could be a geo-distributed cluster with one switchable master, and in another case, some quorum-based record. The templates can vary, but we solve the problem.
Furthermore, having a distributed registry of requests allows us to scale this further — having multiple dispatchers and performers accessing this registry.

Cassandra and Tarantool together
There's another case — "balance showcase". This is precisely an interesting case of the combined use of Cassandra and Tarantool.
We use Cassandra because 2 billion calls a day is not the limit, and it will only grow. Marketers love to segment traffic by sources, and more details are emerging from social networks, for example. This all increases the history.
Cassandra allows horizontal scaling to any volume.
We feel comfortable with Cassandra, but it has one problem — it's not great at reading. Writing is okay, 30,000 per second isn't an issue — the problem lies in reading..
This is why the topic of caching has arisen, and at the same time we decided to address the following issue: there is an old traditional case where equipment from the switch from online billing comes into files that we upload to Cassandra. We tackled the problem of reliably uploading these files, even applying the advice of an IBM manager for file transfer — there are solutions that effectively manage file transfers using the UDP protocol, for example, rather than TCP. This is good, but still, we face delays, and until we load all of this, the operator in the call center cannot inform the client about what has happened with their balance — they have to wait.
To prevent this from happening, we use a parallel functional reserve. When we send an event via Kafka to Tarantool, recalculating aggregates in real time, for example, as of today, we get a balance cache, which can handle balances at any speed, for example, 100 thousand transactions per second and those very 2 seconds.
The goal is that after making a call, the personal account shows not only the updated balance within 2 seconds but also information about why it changed.
Conclusion
These were examples of using Tarantool. We really appreciated Mail.ru's openness and their willingness to consider various cases.
Consultants from BCG or McKinsey, Accenture or IBM are already finding it hard to surprise us with something new — much of what they offer we are either already doing, have done, or plan to do. I believe that Tarantool will occupy a worthy place in our technology stack and will replace many existing technologies. We are in the active development phase of this project.
Oleg and Andrey's presentation was one of the best at the Tarantool Conference last year, and on June 17, Oleg Ivlev will present at with the talk . Also from MegaFon, Alexander Deulin will present a talk . We'll find out what has changed and what plans have been realized. Join us — the conference is free, just need to . All and the conference program has been formed: new cases, new experiences of using Tarantool, architecture, enterprise, tutorials, and microservices.
Source: habr.com
