Last December, HPE announced the release of a modular in-memory computing platform with the world's most extensive scaling capabilities — HPE Superdome Flex. This is a breakthrough in computing systems designed for critical applications, real-time analytics, and high-performance computing with intensive data processing.
Platform has a range of features that make it unique in its industry. We offer you a translation of a blog article , which discusses the modular and scalable architecture of the platform.

The scaling capabilities surpass those of Intel.
Like most x86 server providers, HPE employs the latest Intel Xeon Scalable processors, codenamed Skylake, in its latest-generation servers, including HPE Superdome Flex. The reference architecture from Intel for these processors utilizes the new UltraPath Interconnect (UPI) technology with scaling limitations of up to eight sockets. Most providers using these processors apply a "non-bonded" connection method in their servers; however, HPE Superdome Flex features a unique modular architecture with scaling capabilities exceeding those of Intel: from 4 to 32 sockets in a single system.
This architecture is utilized because we recognized the need for platforms capable of scaling beyond Intel's eight sockets; this is especially relevant today as data volumes are increasing at an unprecedented rate. Furthermore, since Intel primarily developed UPI for servers with two and four sockets, eight-socket "non-bonded" servers encounter bandwidth issues. Our architecture ensures high bandwidth even as the system scales to its maximum configuration.
Price/performance ratio as a competitive advantage
The modular architecture of HPE Superdome Flex is based on a chassis with four sockets, scalable up to eight chassis and 32 sockets in a single server system.A wide range of processors is available for use in the server, from inexpensive Gold models to the top-of-the-line Platinum series of Xeon Scalable processors.
This option to choose between Gold and Platinum processors across the scaling range provides significant advantages in price/performance compared to entry-level systems. For example, in a typical configuration with 6 TB of memory, Superdome Flex offers a cheaper and more powerful solution than competing offerings with four sockets. Why? Due to architectural features, other manufacturers of 4-processor systems are forced to use 128 GB DIMM memory modules and more expensive processors supporting 1.5 TB per socket. This is significantly more costly than using 64 GB DIMM modules in Superdome Flex with eight sockets. As a result, the Superdome Flex platform with eight sockets and 6 TB of memory delivers twice the compute power, double the memory bandwidth, and twice the I/O capabilities, while still being more economical than competing products with four sockets and 6 TB of memory.
Similarly, for an 8-processor configuration with 6 TB of memory, the Superdome Flex platform can provide a less expensive and more powerful solution with eight sockets. How? Other manufacturers of 8-processor systems are forced to use more expensive Platinum processors, while Superdome Flex with eight sockets can utilize low-cost Gold processors, providing the same amount of memory.
In fact, among platforms based on Intel Xeon Scalable processors, only Superdome Flex can support more economical Gold processors in configurations with 8 or more sockets. (The Intel "no glue" architecture only supports 8 sockets with expensive Platinum processors). We also offer a wide selection of processors with varying core counts, from 4 to 28 per processor, allowing them to be matched to workload requirements.
The importance of scalability within a single system
The ability to vertically scale within a single system, or scale up, offers several advantages for critical workloads and databases best suited for the HPE Superdome Flex. These include traditional databases and in-memory databases, real-time analytics, ERP, CRM, and other transactional applications. For these types of workloads, it is simpler and more cost-effective to manage a single vertically scalable environment than a cluster with horizontal scaling; moreover, this significantly reduces latency and enhances performance.
Check out the blog post , to understand why vertical scaling is much more effective than horizontal (clustering) for these types of workloads. Essentially, it all comes down to speed and the ability to operate at the level required for these critical applications.
Consistently high performance even at maximum configurations
The high scalability capabilities of Superdome Flex are achieved through the unique HPE Superdome Flex ASIC chipset, which connects individual 4-processor chassis, as shown in Figures 1 and 2. All ASICs are directly interconnected (with one-hop distances), providing minimal access delays to remote resources and maximum performance. The HPE Superdome Flex ASIC technology enables adaptive routing for load balancing of the switching matrix, optimizing delays and throughput, which enhances the performance and availability level of the system. The ASIC integrates chassis into a cache-coherent fabric matrix and maintains cache consistency across all processors by utilizing a large catalog of cache line state entries, which is built directly into the ASIC. This coherence scheme plays a crucial role, enabling Superdome Flex to maintain near-linear performance scaling from 4 to 32 sockets. Typical 'no glue' architecture configurations demonstrate more limited performance scaling (from four to eight sockets) due to the broadcast of service requests to ensure coherence.

Fig. 1. HPE Flex Grid switching matrix connection diagram for the 32-socket Superdome Flex server

Fig. 2. 4-processor chassis
Total Memory
Similar to processor resources, the amount of memory can also be increased by adding chassis to the system. Each chassis has 48 DDR4 DIMM slots that can accommodate 32 GB RDIMM, 64 GB LRDIMM, or 128 GB 3DS LRDIMM memory modules, providing a maximum memory capacity of 6 TB per chassis. Accordingly, the total amount of RAM in the HPE Superdome Flex reaches 48 TB in the maximum configuration with 32 sockets, allowing it to handle the most memory-intensive applications using in-memory technology.
High I/O Flexibility
Regarding input and output, each Superdome Flex chassis can be equipped with a basket of 16 or 12 I/O slots to provide numerous options for installing standard PCIe 3.0 cards and flexible capabilities to maintain system balance for any workloads. In any basket configuration, the I/O slots are directly connected to the processors without the use of bus repeaters or extenders, which could increase latency or reduce bandwidth. This ensures the highest possible performance for each I/O card.
Low latency
Low access latency to the entire shared memory space is a key factor for the high performance of Superdome Flex. Whether the data resides in local memory or remote memory (in another chassis), its copy can be cached in different processors within the system. The cache coherence mechanism ensures the consistency of cached copies in case a process modifies the data. The latency for processor access to local memory is about 100 ns. The latency for accessing data in another processor's memory via the UPI link is about 130 ns. Processors accessing data stored in another chassis's memory go through a path between two Flex ASICs (always directly connected) with a latency of less than 400 ns, regardless of which chassis the processor is in. As a result, Superdome Flex achieves a throughput between two bi-sectioned halves of over 210 GB/s in an 8-socket configuration, over 425 GB/s in a 16-socket configuration, and over 850 GB/s in a 32-socket configuration. This is more than sufficient for the most demanding and resource-intensive workloads.
Why are high modular scalability capabilities important?
It is no secret that the volume of data is increasing at an unprecedented rate; this means that infrastructure must cope with increasingly demanding requests for processing and analyzing critical and constantly expanding data. But the pace of growth can be unpredictable.
When deploying memory-intensive applications, you might ask: how much will it cost me for the next TB of memory? Superdome Flex позволяет увеличивать объем памяти без замены оборудования, поскольку вы не ограничены слотами DIMM в одном шасси. Кроме того, с увеличением числа пользователей критически важным приложениям всегда требуется высокая производительность, независимо от объема нагрузки.
Today, in-memory databases require hardware platforms with low latency and high throughput. With its innovative architecture, the HPE Superdome Flex platform provides exceptionally high performance, high throughput, and consistently low latency even in the largest configurations. Moreover, you can achieve all this for your critical workloads and databases with a very attractive price/performance ratio compared to systems from other manufacturers.
You can learn about the unique reliability (RAS) features of the Superdome Flex server from the blog and technical description . A blog dedicated to , announced at HPE Discover.
From you can discover how HPE Superdome Flex is used to address cosmology challenges, as well as how the platform is prepared for memory-driven computing, a new memory-based computing architecture.
You can also learn more about the platform from .
Source: habr.com
