Reducing downtime risks through Shared Nothing architecture

The issue of fault tolerance in data storage systems is always relevant, as in our age of widespread virtualization and resource consolidation, the storage system is the component whose failure will lead not just to a minor incident, but to prolonged service downtime. Therefore, modern storage systems consist of many duplicated components (including controllers). But is such protection sufficient?

Reducing downtime risks through Shared Nothing architecture

All vendors, when listing the features of storage systems, invariably mention the high fault tolerance of their solutions, often adding the term "no single point of failure." Let’s take a closer look at a typical data storage system. To avoid downtime during maintenance, storage systems duplicate power supplies, cooling modules, input/output ports, drives (referring to RAID), and, of course, controllers. Upon closer inspection of this architecture, one can notice at least two potential points of failure that are modestly overlooked:

  1. The presence of a single backplane
  2. The presence of only one copy of data

A backplane is a technically complex device that undergoes serious testing during production. Therefore, it is extremely rare for it to fail completely. However, even in the case of partial failures, such as a non-functioning drive slot, replacement will require the complete shutdown of the storage system.

Creating multiple copies of data do not seem problematic at first glance. For instance, the Clone functionality in storage systems, which allows for periodically updating a complete copy of data, is fairly widespread. However, in the event of issues with the backplane, the copy will be just as inaccessible as the original.

A completely obvious solution to overcome these shortcomings is to replicate to another storage system. If we ignore the expected doubling of the hardware costs (after all, we assume that those choosing such a solution think reasonably and accept this fact in advance), there will still be potential expenses for organizing replication in the form of licenses, additional software, and hardware. Most importantly, some means will be required to ensure the consistency of replicated data. This means building a storage virtualizer, vSAN, etc., which also requires financial and time resources.

AccelStor When creating their High Availability systems, the goal was to eliminate the aforementioned shortcomings. This led to the interpretation of Shared Nothing technology, which in loose translation means 'without using shared devices.'

Concept Shared Nothing architecture involves the use of two independent nodes (controllers), each with its own data set. Between the nodes, synchronous replication occurs through a 56G InfiniBand interface, completely transparently to the software running on top of the storage system. As a result, there is no need to use storage virtualizers, software agents, etc.

The physically dual-node solution from AccelStor can be implemented in two models:

  • H510 — based on Twin servers in a 2U chassis, if moderate performance and capacity up to 22TB are required;
  • H710 — based on separate 2U servers, if high performance and large capacity (up to 57TB) are needed.

Reducing downtime risks through Shared Nothing architecture

The H510 model based on Twin servers

Reducing downtime risks through Shared Nothing architecture

The H710 model based on separate servers

The use of different form factors is due to the need for different numbers of SSDs to achieve the required volume and performance. Additionally, the Twin platform is cheaper and allows for more affordable solutions, albeit with a certain conditional 'disadvantage' of having a single backplane. Everything else, including operating principles, is completely identical between both models.

Each node's data set has two groups FlexiRemap, plus 2 hot spares. Each group is capable of withstanding the failure of one SSD. All incoming write requests to the node are handled according to the ideology FlexiRemap rearranges into sequential chains with 4KB blocks, which are then written to the SSD in the most efficient mode for them (sequential write). The host receives confirmation of the write only after the data is physically placed on the SSD, i.e., without caching in RAM. As a result, impressive performance is achieved with up to 600K IOPS for writes and 1M+ IOPS for reads (model H710).

As previously mentioned, data set synchronization occurs in real-time via the 56G InfiniBand interface, which possesses high bandwidth and low latency. To make the most efficient use of the communication channel when transmitting small packets, a dedicated 1GbE link is used for additional heartbeat checks. Only the heartbeat is transmitted through it, so there are no speed performance requirements.

In case of system capacity increase (up to 400+ TB) through expansion shelves they are also connected in pairs to adhere to the ‘no single point of failure’ concept.

For additional data protection (besides having two copies at AccelStor), a special algorithm is used in case of any SSD failure. If an SSD fails, the node will start rebuilding data on one of the hot spare drives. The FlexiRemap group, which is in a degraded state, will switch to read-only mode. This is done to prevent interference between write operations and rebuilding on the backup drive, ultimately speeding up the recovery process and reducing the time when the system is potentially vulnerable. Upon completion of the rebuild, the node returns to normal read-write mode.

Reducing downtime risks through Shared Nothing architecture

Of course, as with other systems, overall performance decreases during rebuilding (since one of the FlexiRemap groups is not operational for writes). However, the recovery process occurs as quickly as possible, which distinguishes AccelStor systems favorably from those of other vendors.

Another useful feature of the Nothing Shared technology in the architecture is the operation of nodes in the so-called true active-active mode. Unlike the ‘classic’ architecture, where a specific volume/pool is owned by only one controller while the second merely performs input/output operations, in these systems. AccelStor Each node operates with its own set of data and does not pass requests to its 'neighbor.' As a result, the overall performance of the system improves due to the parallel processing of I/O requests by the nodes and access to storage. Additionally, the concept of failover is practically non-existent since it is unnecessary to transfer management of volumes to another node in case of failure.

When comparing the Nothing Shared architecture technology with full duplication of storage systems, it may seem to fall slightly short of a complete disaster recovery implementation in terms of flexibility at first glance. This is especially true regarding the organization of the communication line between storage systems. In the H710 model, it's possible to separate nodes up to 100 meters apart by using relatively expensive active optical InfiniBand cables. However, even when compared with typical synchronous replication implementations from other vendors over available FibreChannel at even greater distances, the AccelStor solution proves to be more cost-effective and simpler in installation/operation, as there is no need to install storage virtualization layers and/or integrate with software (which is not always feasible). Plus, let's not forget that AccelStor solutions are all-flash arrays with performance exceeding that of 'classic' SSD-only storage systems.

Reducing downtime risks through Shared Nothing architecture

By utilizing the Nothing Shared technology in AccelStor's architecture, it is indeed possible to achieve storage system availability at the level of 99.9999% at a very reasonable cost. Along with high reliability due to the use of two data copies, it also boasts impressive performance thanks to proprietary algorithms. FlexiRemap, solutions from AccelStor are excellent candidates for key positions in building a modern data center.

Source: habr.com

Buy reliable website hosting with DDoS protection, VPS VDS servers 🔥 Buy reliable website hosting with DDoS protection, VPS VDS servers | ProHoster