
Update!. In the comments, one reader suggested trying (perhaps he is working on it himself), so I added a section about this solution. I also wrote , because the process is quite different from the others.
To be honest, I gave up and discarded (for now, at least). I will use . Why? Because of storage! Who would have thought that I would deal more with storage than with Kubernetes itself? I am using , because it's affordable and performs well, and from the start, I deployed clusters using . I haven't tried the managed Kubernetes services from Google/Amazon/Microsoft/DigitalOcean and so on, because I wanted to learn everything myself. Plus, I'm frugal.
So yes, I spent a lot of time trying to decide which storage to choose while considering possible stacks on Kubernetes. I prefer open-source solutions, not just for the price, but I also explored a couple of paid options out of curiosity, as they have free versions with limitations. I jotted down some numbers from recent tests when comparing different options, and they might interest those looking into storage in Kubernetes. Although personally, I have said goodbye to Kubernetes for now. I also want to mention , which can directly provision volumes on Hetzner Cloud, but I haven't tried it yet. I explored cloud software-defined storage because I needed replication and the ability to quickly attach persistent volumes to any node, especially in case of node failures and other similar situations. Some solutions offer snapshots at a given time and offsite backups, which is convenient.
I tested 6ā7 storage solutions:
As I mentioned , having tested most options on the list, I initially settled on OpenEBS. OpenEBS is very easy to install and use, but, to be honest, after testing with real data under load, its performance disappointed me. It's open-source, and the developers have their They have always been very helpful when I needed assistance. Unfortunately, it has very low performance compared to other options, so I had to run tests again. OpenEBS now has 3 storage engines, but I'm publishing benchmark results for cStor. I don't have numbers for Jiva and LocalPV yet.
In short, Jiva is slightly faster, while LocalPV is quite fast and on par with direct disk benchmarks. The issue with LocalPV is that you can only access it on the node where it was set up, and there is no replication at all. I had some issues restoring backups through on the new cluster because the node names were different. Speaking of backups, cStor has a , which allows you to create off-site snapshot backups at a specific point in time, which is more convenient than file-level backups with Velero-Restic. I wrote , to make managing backups and restores with this plugin easier. Overall, I really like OpenEBS, but its performance...
Rook also has open-source code, and it stands out from the other options on the list as it is a storage orchestrator that performs complex tasks for managing storage with different backends, such as , and others, which significantly simplifies the work. I had issues with EdgeFS when I tried it a few months ago, so I primarily tested with Ceph. Ceph offers not only block storage but also S3/Swift compatible object storage and a distributed file system. What I like about Ceph is the ability to spread the volume data across multiple disks so that the volume can use more disk space than fits on a single disk. That's convenient. Another cool feature is that when you add disks to the cluster, it automatically redistributes data across all disks.
Ceph has snapshots, but as far as I know, they cannot be used directly in Rook/Kubernetes. However, I haven't delved deeply into this. There are no off-site backups, so you'll have to use something like Velero/Restic, but it only supports file-level backups, not point-in-time snapshots. On the other hand, I really appreciate how easy it is to work with Ceph in Rookāit hides almost all the complexities and provides tools to interact directly with Ceph for troubleshooting. Unfortunately, during the stress test of Ceph volumes, I constantly encountered , which causes Ceph to become unstable. It's unclear whether this is a bug in Ceph itself or a problem with how Rook manages Ceph. I tinkered with the memory settings, and it improved somewhat, but the issue isn't fully resolved. Ceph performs well, as shown in the benchmarks below. It also has a good monitoring panel.
I really like Longhorn. In my opinion, itās a promising solution. However, the developers themselves (Rancher Labs) acknowledge that it's not yet ready for production environments, and you can see that. It has open-source code and decent performance (although optimization is still pending), but volumes take a long time to attach to the pod, and in worst cases, it can take 15ā16 minutes, especially after restoring a large backup or upgrading workloads. It has snapshots and off-site backups of those snapshots, but they only apply to volumes, so you'll still need something like Velero for backing up other resources. The backups and restorations are very reliable, but ridiculously slow. Seriously, just absurdly slow. CPU usage and system load often spike when working with medium-sized data in Longhorn. Thereās a convenient monitoring panel for managing Longhorn. Iāve already mentioned that I like Longhorn, but it needs significant work.
StorageOS is the first paid product on the list. It has a developer version with a limited managed storage size of 500 GB, but the number of nodes, as far as I know, is not limited. The sales department told me that prices start at $125 per month for 1 TB, if I remembered correctly. There is a basic monitoring dashboard and a convenient CLI, but the performance is strange: in some benchmarks, it is quite decent, but in the volume stress test, I didn't like the speed at all. Overall, I donāt know what to say. So, I didnāt delve too deeply into it. There are no off-site backups, and I will also have to use Velero with Restic for volume backups. It's strange, considering itās a paid product. Also, the developers were not keen on communicating in Slack.
I learned about Robin on Reddit from their tech director. I hadn't heard of it before, perhaps because I was looking for free solutions whereas Robin is paid. They offer quite a generous free version with 10 TB of storage and three nodes. Overall, the product is quite worthy and has pleasant features. It includes a great CLI, but the coolest part is that you can take a snapshot and back up the entire application (in the resource selector, this is referred to as Helm releases or 'flex apps'), including volumes and other resources, so you can do without Velero. Everything would be wonderful if it weren't for one small detail: if you restore (or 'import', as it's called in Robin) an application on a new clusterāfor example, in the case of a disaster recoveryāthe restoration certainly works, but continuing the backup of the application is not possible. In this release, it's simply impossible, and the developers have confirmed this. This is, to put it mildly, strange, especially considering the other advantages (such as incredibly fast backups and restorations). The developers promise to fix everything in the next release. The performance, overall, is good, but I noticed a peculiarity: if you run a benchmark directly on the volume attached to the host, the read speed is much higher than on the same volume but from within the pod. All other results are identical, but in theory there shouldn't be any difference. Even though they are working on this, I was disappointed by the restoration and backup issuesāI felt I had finally found a suitable solution, and I was even ready to pay for it when I needed more space or more servers.
I don't have much to say here. This is a paid product, equally great and expensive. The performance is simply amazing. So far, this is the best benchmark. In Slack, I was told that the price starts from $205 per month per node, as indicated in the Google GKE Marketplace. I don't know if it would be cheaper if purchased directly. In any case, I can't afford this, so I was very disappointed that the developer license (up to 1 TB and 3 nodes) is practically useless with Kubernetes unless you settle for static preparation. I had hoped that the enterprise license would automatically downgrade to the developer level at the end of the trial period, but that didn't happen. The developer license can only be used directly with Docker, and the setup in Kubernetes is very cumbersome and limited. I certainly prefer open source, but if I had the money, I would definitely choose Portworx. For now, its performance is simply unmatched compared to other options.
I added this section after the post was published, when a reader suggested trying Linstor. I tried it, and I liked it! But there's still more to explore. For now, I can say that the performance is decent (benchmark results are added below). Essentially, I achieved the same performance as with the disk directly, without any overhead. (Don't ask why Portworx numbers are better than the disk benchmark directly. I have no idea. Must be magic.) So Linstor seems very efficient for now. Installing it isn't exactly hard, but it's not as easy as the other options. First, I had to install Linstor (kernel module and tools/services) and set up LVM for thin provisioning and snapshot support outside of Kubernetes, directly on the host, and then create the resources needed to use the storage from Kubernetes. I didn't like that it didn't work on CentOS, so I had to use Ubuntu. It's not a huge deal, but it's a bit annoying because the documentation (which is excellent, by the way) mentions several packages that are impossible to find in the specified Epel repositories. Linstor has snapshots, but no off-site backups, so I had to use Velero with Restic for volume backups. I would prefer snapshots over file-level backups, but it's tolerable if the solution is performant and reliable. Linstor is open source, but there is paid support. If I understand correctly, it can be used without restrictions, even if you donāt have a support contract, but that should be confirmed. I don't know how well Linstor is tested for Kubernetes, but the storage level is outside of Kubernetes and apparently, the solution hasn't just appeared recently, so it's probably been tested in real conditions. Is there a solution here that would make me reconsider and return to Kubernetes? I donāt know, I donāt know. I need to dig deeper and study replication. We'll see. But my first impression is good. I would definitely prefer to use my own Kubernetes clusters instead of Heroku, to gain more freedom and learn something new. Since Linstor is not as straightforward to install as the others, I'll write a post about it soon.
Benchmarks
Unfortunately, I saved very few records of the comparisons because I didn't think I would be writing about it. I only have the results of the basic fio benchmarks, and only for single-node clusters, so I don't have numbers for replicated configurations yet. However, from these results, you can get a rough idea of what to expect from each option since I compared them on the same cloud servers with 4 cores, 16 GB of RAM, and an additional 100 GB disk for the tested volumes. I ran the benchmarks three times for each solution and calculated the average result, plus I reset the server settings for each product. This is all quite unscientific, just to give you a general idea. In other tests, I copied 38 GB of photos and videos from the volume and back to test reading and writing, but unfortunately, I didn't save the numbers. In short: Portworx was much faster.
For the volume benchmark, I used this manifest:
kind: PersistentVolumeClaim
apiVersion: v1
metadata:
name: dbench
spec:
storageClassName: ...
accessModes:
- ReadWriteOnce
resources:
requests:
storage: 5Gi
---
apiVersion: batch/v1
kind: Job
metadata:
name: dbench
spec:
template:
spec:
containers:
- name: dbench
image: sotoaster/dbench:latest
imagePullPolicy: IfNotPresent
env:
- name: DBENCH_MOUNTPOINT
value: /data
- name: FIO_SIZE
value: 1G
volumeMounts:
- name: dbench-pv
mountPath: /data
restartPolicy: Never
volumes:
- name: dbench-pv
persistentVolumeClaim:
claimName: dbench
backoffLimit: 4First, I created a volume with the appropriate storage class, and then I ran the fio job in the background. I chose 1 GB to estimate performance without waiting too long. Here are the results:
I highlighted the best value for each metric in green and the worst in red.
Conclusion
As you can see, in most cases, Portworx performed better than the others. However, to me, it is expensive. I don't know how much Robin costs, but they have a great free version, so if you need a paid product, you can give it a try (I hope they fix the recovery and backup issues soon). Of the three free options, I had the fewest problems with OpenEBS, but its performance is pretty lacking. It's a pity I didn't save more results, but I hope the numbers provided and my comments are helpful.
Source: habr.com
