Monitoring the status of SSDs in Qsan arrays

The use of solid-state drives in data storage is no longer surprising anyone. SSDs have become a staple in IT equipment, from personal computers and laptops to servers and storage systems. Over time, several generations of SSDs have emerged, each offering improved performance, reliability, and maximum capacity. However, the issue of monitoring the write resource of SSDs remains relevant.

Monitoring the status of SSDs in Qsan arrays

Due to their physical structure, solid-state drives have a predefined limited write resource. The fact that much more data is actually written to an SSD than is sent to it by the host (especially within a RAID group) brings us even closer to this designated limit. This circumstance creates a sort of fear among some users regarding the use of SSDs.

In reality, things aren’t as bad as they seem. The calculated resource DWPD is provided for the entire warranty period of the drive (usually 3-5 years). Therefore, the actual write resource TBW will be quite substantial, allowing users not to fear ‘wearing out’ the SSD in just a few months. Moreover, in some cases, drives can be temporarily used in a more intensive mode than the manufacturer intended, thanks to high TBW values. Nonetheless, this does not diminish the necessity of monitoring the current write resource of each specific SSD in order to proactively replace it when certain thresholds are reached.

Each storage vendor implements this functionality in its own way. But most often it simply indicates the drive’s status as operational/non-operational. Qsan in their All Flash systems, on the contrary, provides a complete visualization of the parameters of the current SSD activity in the form of a separate module called QSLife. This module is an integral part of the new operating system XEVO, under which all Qsan storage systems will operate in the future.

For each SSD in the system, the current "health level" is displayed in the most accessible form. It is well known that all modern SSDs keep track of the blocks written to them. Based on these values, the system calculates the wear indicator of the drive in accordance with its mapping. The final result is displayed as a percentage of a completely new SSD. It's also worth noting that the degree of wear is calculated not only for the period during which the drive operated as part of the All Flash array Qsan but for its entire lifespan, including its operation in other systems (if applicable).

Monitoring the status of SSDs in Qsan arrays

In addition to simplified information about the drive, some details can also be found. Specifically, the volume of data written to it over its entire lifespan. And during the time the drive operated as part of the All Flash array Qsan, graphs of its performance in read and write operations are available. Statistics are collected in real-time and can be accessed for any period with a viewing depth of up to one year.

Monitoring the status of SSDs in Qsan arrays

Of course, the goal of this functionality is not only to create beautiful graphs for the administrator's pleasure but also to proactively analyze the condition of the drives and prevent potential future problems related to their wear. Therefore, regarding the "health level" of the SSD, a variety of thresholds and corresponding actions related to the depletion of the SSD's write resource can be set.

Monitoring the status of SSDs in Qsan arrays

If we look at other models of storage systems (not specialized All Flash but general-purpose ones) from Qsan, they lack such a visual report on the drives. It’s understandable: a flagship has to differ from the mainstream in some way. However, even in the regular product line, similar monitoring is conducted. Yes, without collecting statistics on usage and performance. But the main function of monitoring the write resource is present.

Monitoring the status of SSDs in Qsan arrays

As the technology for solid-state drives continues to evolve, concerns about their reliability have somewhat diminished. Nevertheless, monitoring their write resources remains relevant. Such well-configured monitoring will allow the administrator to predict SSD aging in accordance with actual current workloads and enable the company's management to calculate TCO (total cost of ownership) metrics.

Source: habr.com

Buy reliable website hosting with DDoS protection, VPS VDS servers 🔥 Buy reliable website hosting with DDoS protection, VPS VDS servers | ProHoster