Why OceanStor Dorado V6 is the fastest and most reliable storage system

Please don't jump to conclusions based on the headline! We have substantial arguments to support it, and we've packed them in as compactly as we could. We present to you a post about the concept and principles of our new data storage system, which was released in January 2020.

Why OceanStor Dorado V6 is the fastest and most reliable storage system

In our opinion, the main competitive advantages of the Dorado V6 storage family lie precisely in the performance and reliability mentioned in the headline. Yes, it's that simple, but we will discuss how we achieved this 'simplicity' through both clever and straightforward solutions.

To better uncover the potential of next-generation systems, we will talk about the senior representatives of the model range (models 8000, 18000). Unless otherwise stated, these are the models we refer to.

Why OceanStor Dorado V6 is the fastest and most reliable storage system

A few words about the market

To better understand Huawei's position in the market, let's refer to a well-known benchmark — the "magic quadrants" by Gartner. Two years ago, in the sector of general-purpose disk arrays, our company confidently entered the leader group, trailing only behind NetApp and Hewlett Packard Enterprise. Huawei's position in the solid-state storage market in 2018 was characterized as a 'challenger,' but something was lacking to achieve a leadership position.

In 2019, Gartner combined both aforementioned sectors into one — 'Primary Storage.' As a result, Huawei found itself back in the leader quadrant, alongside providers such as IBM, Hitachi Vantara, and Infinidat.

For completeness, it's worth mentioning that 80% of the data for Gartner's analysis is collected from the American market, leading to a noticeable bias in favor of those companies that are well-represented in the U.S. Meanwhile, suppliers focused on European and Asian markets find themselves at a disadvantage. Even so, last year, Huawei's products secured a worthy place in the upper right quadrant, and according to Gartner's verdict, 'are recommended for use.'

Why OceanStor Dorado V6 is the fastest and most reliable storage system

What's new in Dorado V6

The Dorado V6 product line, in particular, is represented by entry-level systems in the 3000 series. Initially equipped with two controllers, they can be horizontally expanded to 16 controllers, 1200 disks, and 192 GB of cache. The system will also be equipped with external Fibre Channel ports (8 / 16 / 32 Gb/s) and Ethernet (1 / 10 / 25 / 40 / 100 Gb/s).

It is worth noting that the use of protocols with no commercial success is currently being phased out, so at the outset we decided to forgo support for Fibre Channel over Ethernet (FCoE) and Infiniband (IB). These will be added in later firmware versions. Support for NVMe over Fabric (NVMe-oF) is available out of the box over Fibre Channel. The next firmware, scheduled for release in June, will include support for NVMe over Ethernet. In our view, the aforementioned set more than adequately meets the needs of most Huawei customers.

File access is not available in the current firmware version and will appear in one of the upcoming updates later this year. Implementation is planned at the native level, via the controllers with Ethernet ports, without the need for additional hardware.

The main difference between the Dorado V6 series 3000 model and the older ones is that it supports only one protocol on the backend — SAS 3.0. Accordingly, the storage devices can only utilize the specified interface. In our opinion, the performance provided is more than sufficient for this type of device.

The Dorado V6 series 5000 and 6000 systems are classified as mid-range solutions. They are also designed in a 2U form factor and equipped with two controllers. They differ from each other in performance, number of processors, maximum number of disks, and cache size. However, architecturally and in engineering terms, the Dorado V6 5000 and 6000 are identical and look the same.

The hi-end class includes the Dorado V6 series 8000 and 18000 systems. Designed in a 4U form factor, they by default have a separate architecture, where the controllers and storage devices are separated. In the minimal configuration, they can also be equipped with only two controllers, although customers typically request the installation of four or more.

The Dorado V6 8000 scales horizontally up to 16 controllers, while the Dorado V6 18000 scales up to 32. These systems are equipped with different processors featuring varying core counts and cache sizes. The identity of the engineering solutions is retained, as it is in the mid-end class models.

2U shelves with drives connect via RDMA with a throughput of 100 Gbps. The backend of the older Dorado V6 series also supports SAS 3.0, but mainly in case SSDs with this interface drop significantly in price. Then, there will be economic feasibility in using them, even considering their relatively lower performance. At this point, the price difference between SSDs with SAS and NVMe interfaces is so small that we are not ready to recommend such a solution.

Why OceanStor Dorado V6 is the fastest and most reliable storage system

Inside the controller

Dorado V6 controllers are built on our own component base. No Intel processors, no Broadcom ASICs. Thus, every single component of the motherboard, including the motherboard itself, is completely insulated from the risks associated with sanctions pressure from American companies. Those who have seen our equipment firsthand must have noticed the shields with a red stripe under the logo. This indicates that the product lacks American components. This is Huawei's official course—transitioning to self-produced components or, at the very least, those manufactured in countries not aligning with U.S. policies.

Here's what can be seen on the controller board.

  • Universal network interface (Hisilicon 1822 chip) responsible for connecting to Fibre Channel or Ethernet.
  • BMC chip providing remote system access, specifically Hisilicon 1710, for full-featured remote control and system monitoring. Similar chips are also used in our servers and other solutions.
  • The central processor is a Kunpeng 920 chip from Huawei, built on ARM architecture. It is shown in the diagram above, although other controllers may have different models installed with varying core counts, clock frequencies, etc. The number of processors in one controller also changes from model to model. For example, in the older Dorado V6 series, there are four processors on one board.
  • The SSD controller (Hisilicon 1812e chip) supports both SAS and NVMe drives. Additionally, Huawei manufactures SSDs itself but does not produce the NAND cells, opting to purchase them from the four largest global manufacturers in the form of uncut silicon wafers. Huawei handles the cutting, testing, and packaging into chips independently, releasing them under its own brand.
  • The artificial intelligence chip — Ascend 310. By default, it is not present on the controller and is mounted via a separate card that occupies one of the slots designated for network adapters. This chip is used to enable intelligent cache behavior, manage performance, or handle deduplication and compression processes. While all these tasks can also be accomplished by the central processor, the AI chip allows for much more efficient execution.

Why OceanStor Dorado V6 is the fastest and most reliable storage system

On the Kunpeng processors

The Kunpeng processor is a system on chip (SoC), which includes not only a computing unit but also hardware modules that accelerate various processes, such as checksum calculations or performing erasure coding. It also features hardware support for SAS, Ethernet, DDR4 (ranging from six to eight channels), and more. This enables Huawei to create storage controllers that match the performance of traditional Intel solutions.

Moreover, its own solutions based on ARM architecture allow Huawei to develop full server solutions and offer them to its clients as an alternative to x86.

Why OceanStor Dorado V6 is the fastest and most reliable storage system

The new Dorado V6 architecture...

The internal architecture of the Dorado V6 storage system of the senior series is represented by four main subdomains (factories).

The first factory serves as the common frontend (network interfaces responsible for communication with the SAN factory or hosts).

The second is a set of controllers, each of which can reach out via the RDMA protocol to any frontend network card as well as to a neighboring 'engine', which consists of a box containing four controllers along with shared power and cooling units. Currently, high-end models of the Dorado V6 can be equipped with two such 'engines' (thus, eight controllers in total).

The third factory handles the backend and consists of 100G RDMA network cards.

Finally, the fourth factory in hardware is represented by attachable intelligent shelves with drives.

This symmetrical structure reveals the full potential of NVMe technology and ensures high performance and reliability. The input-output process is maximally parallelized across processors and cores, allowing simultaneous read and write across multiple threads.

Why OceanStor Dorado V6 is the fastest and most reliable storage system

...and what it has given us

The maximum performance of Dorado V6 solutions is approximately three times higher than that of previous generation systems (of the same class) and can reach 20 million IOPS.

This is because, in the previous generation of devices, NVMe support was limited to the drive shelves. Now, it is present at all stages, from host to SSD. The backend network has also undergone changes: SAS/PCIe have been replaced by RoCEv2 with a bandwidth of 100 Gbps.

The SSD form factor itself has also changed. Previously, a 2U shelf accommodated 25 drives, but now it has been upgraded to 36 palm-sized disks. Additionally, the shelves have become 'smarter.' Each now has a fault-tolerant system with two controllers based on ARM chips, similar to those found in central controllers.

Why OceanStor Dorado V6 is the fastest and most reliable storage system

Currently, they are only engaged in data reorganization, but with the release of new firmware, compression and erasure coding will be added, which will reduce the load on the main controllers from 15% to 5%. Offloading some tasks to the shelves also frees up bandwidth in the internal network. All this significantly increases the scalability potential of the system.

Compression and deduplication in previous generation storage systems were performed with fixed-length blocks. Now, a mode for working with variable-length blocks has been added, which still needs to be enabled manually. Subsequent firmware updates may change this circumstance.

Also briefly on fault tolerance. Dorado V3 remained operational if one of the two controllers failed. Dorado V6 will ensure data availability even if sequentially seven out of eight controllers fail or four simultaneously fail from one 'engine.'

Why OceanStor Dorado V6 is the fastest and most reliable storage system

Reliability from an economic perspective

Recently, a survey among Huawei customers was conducted to determine what downtime for individual elements of IT infrastructure they deem acceptable. Generally, respondents showed a tolerable attitude towards a hypothetical situation where an application becomes unresponsive for several hundred seconds. For operating systems or bus host adapters, critical downtime was deemed to be several dozen seconds (essentially, the reboot time). Customers set even higher expectations for the network: its bandwidth should not dip for more than 10-20 seconds. Unsurprisingly, the most critical failures identified by respondents were those related to storage systems. From a business perspective, the downtime of a storage system should not exceed... a few seconds per year!

In other words, if a bank's client application is unresponsive for 100 seconds, it is unlikely to result in catastrophic consequences. However, if the storage system is down for the same duration, it could lead to business halting and significant financial losses.

Why OceanStor Dorado V6 is the fastest and most reliable storage system

The chart above shows the cost of one hour of downtime for the ten largest banks (Forbes data from 2017). Agreeably, if your company is comparable in size to the Chinese banks, justifying the need for a storage system costing several million dollars will not be overly difficult. Conversely, if the business does not incur significant losses during downtime, it is unlikely they would purchase a hi-end class storage system. In any case, it's crucial to have an understanding of the potential financial hole that may form while the system administrator is addressing a failing storage system.

Why OceanStor Dorado V6 is the fastest and most reliable storage system

One second for failover

In Solution A depicted in the illustration above, you can see our previous generation system, Dorado V3. Its four controllers operate in pairs, with cache copies stored in only two controllers. The controllers within each pair can redistribute the load. At the same time, as you can see, there are no 'factories' for the frontend and backend, so each shelf of drives connects to a specific controller pair.

The diagram of Solution B shows a current market solution from another vendor (recognize it?). It already has both frontend and backend factories, with storage directly connected to four controllers. However, there are subtle nuances in the operation of the system’s internal algorithms that may not be obvious at first glance.

On the right is our current architecture of the Dorado V6 storage system with a complete set of internal components. Let’s consider how these systems handle a typical situation — the failure of one controller.

In classic systems, including Dorado V3, the time required to redistribute load during a failure reaches four seconds. During this time, all input/output is completely halted. In Solution B from our colleagues, despite a more modern architecture, the downtime during a failure is even longer — six seconds.

The Dorado V6 storage system resumes operation just one second after a failure. This result is achieved thanks to a uniform internal RDMA environment, allowing the controller to access ‘foreign’ memory. Another crucial factor is the presence of a frontend factory, which ensures that the path for the host remains unchanged. The port stays the same, and the load is simply redirected to the functional controllers by multipassing drivers.

The failure of the second controller in Dorado V6 is handled in one second following the same scheme. In Dorado V3, this takes about six seconds, while in a solution from another vendor — nine seconds. For many database systems, such intervals can no longer be considered acceptable, as during this time the system switches to standby mode and ceases to operate. This is particularly critical for database systems consisting of multiple partitions.

The failure of the third controller in Solution A is insurmountable. Simply put, access to part of the data disks is lost. In contrast, Solution B in such a situation recovers functionality, which takes, as in the previous case, nine seconds.

What about Dorado V6? One second.

Why OceanStor Dorado V6 is the fastest and most reliable storage system

What can you accomplish in one second?

Almost nothing, but we don't need that. Let’s reiterate that in Dorado V6, the hi-end class front-end factory is decoupled from the controller factory. This means there are no strictly allocated ports belonging to a specific controller. Restructuring upon failure does not imply searching for alternative paths or reinitializing multipathing. The system continues to operate as it did.

Why OceanStor Dorado V6 is the fastest and most reliable storage system

Resilience to multiple failures

The higher-end models of Dorado V6 can easily withstand the simultaneous failure of two any (!) controllers from any of the "engines." This has become possible because the solution now keeps three copies of the cache. Therefore, even in the event of a dual failure, there will always be one intact copy available.

A synchronous failure of all four controllers in one of the "engines" will also not have fatal consequences, as all three copies of the cache are distributed between the "engines" at all times. The system itself ensures adherence to this operational logic.

Finally, a sequential failure of seven out of eight controllers is a scenario that is highly unlikely. Moreover, the minimum allowable interval between individual failures to maintain operational capacity is 15 minutes. During this time, the storage system completes the necessary operations for cache migration.

The last surviving controller will ensure the functioning of the data storage and maintain the cache for five days (default value, easily adjustable in settings). After that, the cache will be disabled, but the storage system will continue to operate.

Why OceanStor Dorado V6 is the fastest and most reliable storage system

Non-intrusive updates

The new Dorado V6 OS allows firmware updates to the storage system without rebooting the controllers.

The operating system, as with previous solutions, is based on Linux; however, many operational processes have been shifted from the kernel to user mode. Most functions, such as those responsible for deduplication and compression, are now standard processes running in the background as demons. As a result, there is no need to change the entire operating system to update individual modules. For instance, to add support for a new protocol, it will be sufficient to disable the corresponding software module and start the new one.

It is clear that the issues of system updates still remain, as there can be components in the core that require updating. However, according to our observations, such instances make up less than 6% of the total. This allows controllers to be rebooted tens of times less frequently than before.

Why OceanStor Dorado V6 is the fastest and most reliable storage system

Disaster recovery solutions and high availability solutions (HA/DR)

Dorado V6 is ready for integration into geo-distributed solutions, metropolitan-level clusters, and 'triple' data centers right out of the box.

On the left in the illustration above is the metro cluster familiar to many. Two storage systems operate in an active/active mode at distances of up to 100 km from each other. Such infrastructure, with the presence of one or more quorum servers, can be supported by solutions from different companies, including our cloud operating system FusionSphere. The characteristics of the channel between sites become particularly important in such projects, while all other tasks in our case are handled by the HyperMetro functionality, which is also available right out of the box. Integration via Fibre Channel, as well as iSCSI in IP networks, is possible should the need arise. There is no longer a necessity for dedicated 'dark' optics, as the system is capable of connecting through existing channels.

When constructing such systems, the only hardware requirement for the storage system is the allocation of ports for replication. It is sufficient to purchase a license, launch quorum servers - physical or virtual - and ensure IP connectivity to the controllers (10 Mbps, 50 ms).

This architecture can be easily transferred to a system with three data centers (see the right side of the illustration). For example, when two data centers operate in metro-cluster mode, while a third site, located more than 100 km away, utilizes asynchronous replication.

The system technologically supports various business scenarios that will be implemented in the case of a large-scale incident.

Why OceanStor Dorado V6 is the fastest and most reliable storage system

Survivability of the metro-cluster under multiple failures

Above and below, a classic metro-cluster consisting of two storage systems and a quorum server is also shown. As you can see, in six out of nine possible scenarios of multiple failures, our infrastructure will remain operational.

For example, in the second scenario, when the quorum server fails and synchronization between sites occurs, the system remains productive, as the second site ceases operations. Such behavior is already embedded in the built-in algorithms.

Even after three failures, access to information can be retained if the interval between them is at least 15 seconds.

Why OceanStor Dorado V6 is the fastest and most reliable storage system

A familiar trump card up the sleeve

Recall that Huawei manufactures not only storage systems but also a full range of network equipment. Regardless of which storage vendor you choose, if the sites use a WDM network, in 90% of cases it will be built on our company’s solutions. This raises a logical question: why assemble a zoo of systems when all guaranteed compatible hardware can be obtained from a single vendor?

Why OceanStor Dorado V6 is the fastest and most reliable storage system

On the topic of performance

Surely, no one needs to be convinced that switching to All-Flash storage systems significantly reduces infrastructure maintenance costs, as all routine operations are performed several times faster. All suppliers of such equipment testify to this. Meanwhile, many vendors begin to be evasive when it comes to performance drops when various operating modes of the storage systems are activated.

In our industry, it is common practice to issue storage systems for test operation for one to two days. The vendor runs a 20-minute test on an empty system, achieving astronomical performance metrics. However, in real operation, 'underwater rocks' quickly emerge. After just a day, the beautiful IOPS figures drop dramatically, and if the storage system is filled to 80%, they become even lower. Switching from RAID 10 to RAID 5 results in an additional loss of 10-15%, and in metro-cluster mode, performance is further halved.

Everything listed above does not apply to Dorado V6. Our customers have the opportunity to run a performance test over the weekend or at least overnight. This reveals the garbage collection process and makes it clear how the activation of various options—such as snapshots and replication—affects the achieved IOPS.

In Dorado V6, snapshots and RAID with parity have almost no impact on performance (3–5% instead of 10-15%). Garbage collection (filling storage cells with zeros), compression, and deduplication on a storage system filled to 80% will always affect the overall speed of request processing. However, what makes Dorado V6 interesting is that no matter what combination of functions and protective mechanisms you activate, the final performance of the storage system will not drop below 80% of the metric obtained without load.

Why OceanStor Dorado V6 is the fastest and most reliable storage system

Load balancing

The high performance of Dorado V6 is achieved through balancing at every stage, namely:

  • multi-passing;
  • using multiple connections from a single host;
  • having a frontend factory;
  • parallelizing the work of storage system controllers;
  • distributing the load across all storage devices at the RAID level 2.0+.

In principle, this is common practice. Nowadays, few people keep all data on a single LUN: everyone tries to have at least eight, or even forty, or even more. This is an obvious and correct approach that we share. However, if your task requires only one LUN, which is easier to maintain, our architectural solutions allow achieving 80% of the performance available when using multiple LUNs.

Why OceanStor Dorado V6 is the fastest and most reliable storage system

Dynamic CPU load scheduling

Load distribution on processors when using a single LUN is implemented as follows: tasks are broken down into separate small 'shards' at the LUN level, each of which is strictly assigned to a specific controller in the 'engine'. This is done to ensure that the system does not lose performance while 'jumping' with this piece of data between various controllers.

Another mechanism for maintaining high performance is dynamic scheduling, where certain CPU cores can be allocated to different task pools. For example, if the system is currently idle during deduplication and compression, some cores may engage in processing input-output. Or vice versa. All of this is executed automatically and transparently for the user.

Current load data for each of the Dorado V6 cores is not displayed in the graphical interface, but you can access the controller OS via the command line and use the standard Linux command. top.

Why OceanStor Dorado V6 is the fastest and most reliable storage system

Support for NVMe and RoCE

As previously mentioned, the Dorado V6 currently supports NVMe over Fibre Channel "out of the box" and does not require any licenses. Support for NVMe over Ethernet will be available in the middle of the year. To fully utilize this, Ethernet with direct memory access (DMA) version v2.0 support is needed from both the storage system and the switches and network adapters, such as Mellanox ConnectX-4 or ConnectX-5. Network cards based on our chips can also be used. Additionally, RoCE support must be implemented at the operating system level.

Overall, we consider the Dorado V6 to be an NVMe-oriented system. Despite the existing support for Fibre Channel and iSCSI, there are plans to transition to high-speed Ethernet with RDMA in the future.

Why OceanStor Dorado V6 is the fastest and most reliable storage system

A touch of marketing

Thanks to its high fault tolerance, good horizontal scalability, and support for various migration technologies, the economic benefits of acquiring the Dorado V6 become apparent at the beginning of intensive storage system operation. We will continue to strive to make ownership of the system as beneficial as possible, even if this is not immediately obvious in the initial stages.

Specifically, we have developed the FLASH EVER program, aimed at extending the lifecycle of storage systems and designed to alleviate the customer’s burden during upgrades as much as possible.

Why OceanStor Dorado V6 is the fastest and most reliable storage system

This program includes a number of measures:

  • the ability to gradually replace controllers and shelves with disks for newer versions without the need to replace the entire system (for high-end class Dorado V6 systems);
  • the possibility of federated storage (combining different versions of Dorado within a single hybrid storage cluster);
  • smart virtualization (the capability to use third-party equipment as part of the Dorado solution).

Why OceanStor Dorado V6 is the fastest and most reliable storage system

It is worth noting that the challenging global situation has had little impact on the commercial prospects of the new system. Despite the fact that the official release of Dorado V6 took place only in January, we see significant demand for it in China, as well as strong interest from Russian and international partners in the finance sector and government entities.

Additionally, due to the pandemic, however long it may last, there is an acute need to provide remote employees with virtual desktops. In this process, Dorado V6 could address many concerns. We are putting forth all necessary efforts for this, including having nearly agreed on the inclusion of the new system in VMware's compatibility list.

***

By the way, don't forget about our numerous webinars held not only in the Russian-speaking segment but also on a global level. The list of webinars for April is available at this link.

Source: habr.com

Buy reliable website hosting with DDoS protection, VPS VDS servers 🔥 Buy reliable website hosting with DDoS protection, VPS VDS servers | ProHoster