New metrics for object storage

New metrics for object storageFlying Fortress by Nele-Diel

S3 Object Storage Team Mail.ru Cloud Storage translated an article about important criteria when choosing object storage. Below is the text from the author's perspective.

When it comes to object storage, people typically focus on just one feature — the price per TB/GB. While this metric is important, it creates a narrow perspective and equates object storage with a tool for archiving. Additionally, such an approach diminishes the significance of object storage within a company's technology stack.

When selecting object storage, attention should be paid to five key characteristics:

  • performance;
  • scalability;
  • S3 compatibility;
  • failure response;
  • integrity.

These five characteristics represent new metrics for object storage, alongside cost. Let's examine them all.

Performance

Traditional object storage solutions often lack performance. Service providers have historically sacrificed performance in the pursuit of lower prices. However, the scenario is different with modern object storage.

The speed of various storage solutions is approaching or even exceeding that of Hadoop. The modern requirements for read and write speeds range from 10 GB/s for hard drives to 35 GB/s for NVMe. 

This bandwidth is sufficient for Spark, Presto, Tensorflow, Teradata, Vertica, Splunk, and other contemporary computing frameworks in the analytics stack. The fact that MPP databases are being configured on object storage shows its increasing use as a primary storage solution.

If your storage system cannot provide the required speed, you won't be able to utilize the data or derive value from it. Even if you extract data from object storage into an in-memory processing structure, bandwidth is still needed to transfer data to and from memory. Outdated object storage solutions lack this capability.

This is a key point: the new performance metric is bandwidth, not latency. Bandwidth is necessary for scalable data, and this is standard in modern data infrastructure.

While performance tests are a good way to determine performance, it cannot be accurately measured until the application is running in an environment. Only then can you identify where the bottleneck lies: in the software, disks, network, or at the computational level.

Scalability

Scalability refers to the number of petabytes that can fit into a single namespace. Providers claim easy scalability, but they often overlook that as systems scale, large monolithic systems become fragile, complex, unstable, and costly.

A new measure of scalability is the number of namespaces or clients you can service. This metric is derived directly from hyperscalers, where the storage building blocks are small but scale to billions of units. Overall, this is a cloud metric.

When standard blocks are small, they are easier to optimize, meaning ensuring security, access control, policy management, lifecycle, and updates without interruption. Ultimately, this leads to performance. The size of the building block is a function of the fault tolerance manageability, which is how highly resilient systems are built.

Multitenancy has many characteristics. While the parameter describes how organizations provide access to data and applications, it also pertains to the applications themselves and the logic of their isolation from one another.

Characteristics of the modern approach to multitenancy include:

  • In a short time, the number of clients can grow from a few hundred to several million.
  • Clients are completely isolated from one another. This allows them to run different versions of the same software and store objects with various configurations, permissions, features, security levels, and maintenance. This is necessary when scaling new servers, updates, and geographical regions.
  • Storage scales elastically, and resources are provided on demand.
  • Every operation is managed by APIs and automated without human involvement.
  • Software can be deployed in containers and use standard orchestration systems such as Kubernetes.

Compatibility with S3

Amazon S3 API is essentially the standard for object storage. Every software provider for object storage claims compatibility with it. Compatibility with S3 is binary: either it is fully implemented or it is not.

In practice, there can be hundreds or thousands of edge cases where something goes wrong when using object storage. This is especially true with proprietary software and service providers. Its primary use cases are direct archiving or backups, so there are few reasons to call the API, and use cases tend to be homogeneous.

Significant advantages exist for open-source software. It covers most edge cases by considering the size and variety of applications, operating systems, and hardware architectures.

All of this is important for application developers, so it's worth testing how the application performs with storage providers. Open source simplifies this process — it's easier to understand which platform is suitable for your application. The provider can be used as a single entry point to storage — meaning it will meet your needs. 

Open source means that applications are not tied to a provider and are more transparent. This ensures a long application lifecycle.

And a few more remarks about open source and S3. 

If you're running an application for big data, S3 SELECT significantly enhances performance and efficiency. This is achieved by using SQL to retrieve only the objects you need from storage.

A key point is the support for bucket notifications. Bucket notifications facilitate serverless computing — an important component of any microservices architecture provided as a service. Given that object storage is effectively cloud storage, this feature becomes crucial when cloud applications utilize object storage.

Finally, the implementation of S3 should support Amazon S3 API encryption interfaces on the server side: SSE-C, SSE-S3, SSE-KMS. Even better if S3 supports unauthorized access protection that is truly secure. 

Response to failures

A metric that is often overlooked is how the system handles failures. Failures can occur for various reasons, and the object storage must manage them all.

For example, there is a single point of failure, and the metric for this is zero.

Unfortunately, many object storage systems utilize special nodes that must be activated for the cluster to function properly. These include name nodes or metadata servers — creating a single point of failure.

Even where multiple points of failure are provided, the ability to withstand catastrophic failures is paramount. Disks fail, servers fail. The key is to create software designed to handle failures as a normal state. When a disk or node fails, this software will continue to operate without disruption.

Built-in protection against data erasure and degradation ensures that you can lose as many disks or nodes as you have parity blocks — typically half of the disks. Only then will the software be unable to recover data.

Failures are rarely tested under load, but such testing is essential. Stress-testing failures will reveal the total costs incurred post-failure.

Consistency

A consistency metric of 100% is also referred to as strong consistency. Consistency is a key component of any storage system, but strong consistency is quite rare. For instance, Amazon S3 ListObject is not strongly consistent; it is only consistent at the end.

What does strong consistency entail? For all operations after a confirmed PUT operation, the following must apply:

  • The updated value is visible when read from any node.
  • The update is protected by failover from the node failure.

This means: if the plug is pulled in the middle of a write, nothing is lost. The system never returns corrupted or stale data. This is a high bar that matters for many scenarios: from transactional applications to backup and recovery.

Conclusion

These are new metrics for object storage that reflect usage patterns in modern organizations, where performance, consistency, scalability, failure domains, and S3 compatibility are the foundational blocks for cloud applications and big data analytics. I recommend using this list in conjunction with the price when creating modern data stacks. 

About Mail.ru Cloud Solutions Object Storage: S3 Architecture. 3 Years of Evolution of Mail.ru Cloud Storage.

What else to read:

  1. Example of an event-driven application based on webhooks in Mail.ru Cloud Solutions' S3 object storage.
  2. More than Ceph: cloud block storage MCS 
  3. Working with Mail.ru Cloud Solutions' S3 object storage like a file system.
  4. Our Telegram channel for updates about the S3 storage and other products. 

Source: habr.com

Buy reliable website hosting with DDoS protection, VPS VDS servers 🔥 Buy reliable website hosting with DDoS protection, VPS VDS servers | ProHoster