Request a Demo
Join a 30 minute demo with a Cloudian expert.
TL;DR: AI storage is judged on throughput, IOPS, latency, metadata speed and scale-out design. Cloudian HyperStore is best for exabyte S3 data lakes, WEKA NeuralMesh is best for GPU-fed training, and Everpure FlashBlade for file plus object.
Modern AI storage performance and scalability rely on tiered architectures, such as local NVMe caching for rapid model checkpoints and parallel file systems or S3-compatible object storage for multi-petabyte training datasets. Choosing the right configuration prevents costly GPU starvation and latency bottlenecks.
Unlike traditional storage, AI storage is optimized for handling extremely large datasets, often consisting of unstructured data such as images, videos, and text. These systems are engineered to provide high throughput, low latency, and the ability to scale out rapidly, ensuring that machine learning models and AI pipelines can access and process data efficiently.
Key evaluation criteria
The criteria below are the measurements that separate AI storage solutions from general-purpose enterprise storage:
Solutions covered in this guide:
This is part of a series of articles about AI infrastructure
In this article:
High throughput is essential for AI storage because training and inference processes often require reading and writing vast amounts of data quickly. During model training, data is loaded in large batches, and any bottleneck in data delivery can significantly slow down the learning process. To meet these demands, AI storage systems are built with high-bandwidth connections, parallel processing capabilities, and optimized data paths that allow multiple streams of data to be read or written simultaneously.
In environments where deep learning is prevalent, the throughput requirements can easily exceed those of traditional enterprise workloads. For example, image and video datasets commonly used in computer vision projects can reach petabyte scale, requiring storage systems that can sustain multi-gigabyte-per-second data transfer rates.
Low latency is another critical aspect of AI storage, particularly during inference and real-time data processing. Latency measures the delay between a data request and its fulfillment. In AI workflows, especially those involving interactive or time-sensitive applications, high latency can degrade performance and responsiveness. AI storage systems must minimize this delay to ensure that data is delivered to processing units as quickly as possible.
High IOPS, or input/output operations per second, is equally important when dealing with workloads that involve many small, random data accesses. Training modern neural networks often requires shuffling and accessing random batches of data, which puts pressure on the storage system to handle numerous simultaneous I/O requests efficiently. Solutions optimized for high IOPS ensure that data bottlenecks do not impede the overall performance of AI pipelines, supporting both training and inference at scale.
Massive scalability is a fundamental requirement for AI storage, as data volumes grow exponentially with the adoption of AI in various industries. Storage systems must accommodate not only the current dataset sizes but also future growth without requiring disruptive migrations or upgrades. This is typically achieved through distributed, scale-out architectures that allow organizations to add capacity and performance incrementally by simply adding more nodes or drives.
Scalability also involves maintaining consistent performance as the system grows. Inadequate scaling can result in performance degradation, management complexity, and increased costs. Effective AI storage solutions offer seamless expansion, automatic load balancing, and efficient data distribution to ensure that as more resources are added, the system continues to operate efficiently.
Related content: Read our guide to storage systems for long-term AI data retention
High-concurrency access is essential in AI environments where multiple users, applications, or processes need to access the same data simultaneously. This is common in collaborative AI projects, distributed training scenarios, or shared data lakes, where data must be available to many clients at once without contention or performance drops. AI storage systems must be built to handle large numbers of concurrent read and write operations while maintaining consistent throughput and low latency.
To achieve this, AI storage architectures often use distributed file systems, parallel data paths, and advanced caching mechanisms. These features allow for efficient sharing of data resources across large clusters of GPUs or CPUs, supporting high levels of parallelism and collaboration. Without robust support for concurrency, storage systems can become a bottleneck, limiting the speed and scalability of AI development and deployment workflows.
Each criterion below covers one dimension of AI storage performance and scalability, with the checks to apply while comparing systems.
Throughput is the volume of data a system can move per second, usually stated in GB/s or TB/s. Training jobs read large batches continuously, and checkpoint writes arrive in bursts that can stall a job if the storage cannot absorb them. Vendors publish throughput figures at wildly different cluster sizes, so the number only means something alongside the node count and hardware it was measured on.
Evaluation criteria:
IOPS counts discrete read and write operations per second. It matters most where AI pipelines touch many small files or objects: shuffling training samples, reading individual records during inference, or scanning a dataset made of millions of small images. A system that streams large sequential files well can still fall over on random small-object access.
Evaluation criteria:
Latency is the time between a request and its fulfilment, measured in milliseconds or microseconds. Inference, retrieval-augmented generation and interactive applications are all latency-bound, and tail latency usually matters more than the average. Media choice, network path and software stack all contribute.
Evaluation criteria:
Metadata performance covers listings, attribute lookups, directory traversal and indexing. AI datasets routinely hold hundreds of millions of files or objects, and slow metadata operations create bottlenecks that never show up in a bandwidth benchmark. Some systems dedicate hardware or a separate service to metadata; others distribute it across all nodes.
Evaluation criteria:
Scale-up adds resources to one system and eventually hits an architectural ceiling. Scale-out adds nodes, growing capacity and performance together. For AI workloads that grow unpredictably, the question is not just whether a system scales out, but what expansion costs operationally: rebalancing time, downtime, and whether performance keeps pace with capacity.
Evaluation criteria:
This is the practical test of everything above: can the storage keep hundreds or thousands of GPUs busy at once. It depends on concurrency handling, direct data paths to GPU memory, and validated integration with GPU platforms. A system that performs well for a single node can still starve a large cluster.
Evaluation criteria:
The table summarizes how each solution measures up against the six criteria. Each is covered in detail below.
| Category | Solution | How It Meets the Criteria |
| Object storage platforms for AI | Cloudian HyperStore | Shared-nothing peer-to-peer design with no metadata server or head node, so throughput scales linearly with capacity as nodes are added. GPUDirect support and parallel S3 access feed GPU clusters directly. |
| Object storage platforms for AI | Dell ObjectScale | All-flash XF960 targets AI training and checkpointing with S3 over RDMA and GPU-direct access, scaling to 47.2 PB raw per rack across up to 16 nodes. |
| Object storage platforms for AI | Scality RING | MultiScale architecture scales capacity, performance, tenants and sites independently to exabyte scale, with metadata deployable separately from object data. |
| High-performance file and flash platforms | WEKA NeuralMesh | States that throughput and IOPS scale linearly with consistently low latency from ten to ten thousand nodes, with the same software running on-prem and in AI clouds. |
| High-performance file and flash platforms | DDN Infinia | Metadata-driven platform with high-speed indexing, a KV store for embeddings and inference state, and native S3 plus POSIX on one namespace. |
| High-performance file and flash platforms | Everpure FlashBlade | Unified file and object on one operating system, spanning 650 GB/s on FlashBlade//S to 10+ TB/s on FlashBlade//EXA for large-scale AI and HPC. |
Related content: Read our guide to the leading AI storage vendors
How we selected these solutions: We shortlisted AI storage platforms based on sustained throughput and IOPS at scale, latency under concurrent access, metadata handling across large file and object counts, scale-out architecture, and validated support for large GPU clusters.

Best for: Exabyte-scale S3 data lakes feeding AI training and inference
Strengths: Peer-to-peer scale-out, GPUDirect support, full S3 API compatibility
Things to consider: Monitoring views can take extra steps to surface detail
Cloudian HyperStore is object storage software for large volumes of unstructured data, deployed on-premises or across multiple sites and managed as a single system. It uses a shared-nothing, peer-to-peer architecture in which every node participates equally in serving I/O, with no metadata server, head node or central controller in the data path.
That design shapes how it scales. Adding a node adds CPU, memory, network and disk to the cluster at the same time, so throughput grows alongside capacity rather than funnelling through a fixed set of controllers. For AI pipelines, data is reached through direct parallel access over the S3 API, supporting thousands of concurrent operations.
HyperStore runs as software-defined storage on industry-standard hardware, either on Cloudian appliances or on servers the buyer selects, with all-flash configurations available where performance requirements are highest.
Key features include:
| Criterion | Solution Fit | Key Considerations |
| Read and write throughput | Parallel S3 reads and writes across all nodes, with all-flash configurations available; throughput grows as nodes are added | Achieved figures depend on cluster size, node type and media, so a sizing review is worthwhile |
| IOPS | Peer-to-peer design lets every node serve requests, supporting thousands of concurrent S3 operations | Workloads dominated by very large numbers of small objects benefit from all-flash configurations |
| Latency | Direct parallel S3 access with data-locality optimization and NVIDIA GPUDirect support shortens the path to GPU memory | Latency depends on the network fabric and drive media selected |
| Metadata performance | No separate metadata server or head node; metadata handling is distributed across peer nodes | Reviewers note that monitoring and reporting views can take extra steps to reach specific detail |
| Scale-up vs. scale-out storage | Modular scale-out with non-disruptive expansion at one site or across many, adding compute and capacity together | Multi-site growth still benefits from capacity and performance planning per location |
| GPU cluster scalability | GPUDirect support and high concurrent operation counts feed GPU clusters, with one namespace spanning sites | Advanced configurations may call for vendor guidance during initial setup |


Best for: Enterprise object storage for GenAI training and checkpointing
Strengths: All-flash XF960, S3 over RDMA, exascale rack density
Things to consider: Write performance and S3 gaps flagged by reviewers
Dell ObjectScale is an S3-compatible, Kubernetes-native object storage platform available as appliances, as a software update for existing Dell ECS environments, or as software-defined storage on PowerEdge servers. It serves as a storage engine within the Dell AI Data Platform.
The lineup splits by workload. The X560 is HDD-based and aimed at general-purpose workloads and AI data lakes, reaching up to 9.2 PB raw per rack across up to 16 nodes with cache SSD included. The XF960 is all-flash, aimed at AI training and checkpointing, analytics and fast backup, with drives from 7.68 TB to 122.88 TB and up to 47.2 PB raw per rack.
For AI access paths, ObjectScale supports S3 over RDMA with GPU-direct access, and Dell has published material on vector storage, RAG connectors and KV cache offloading built on the platform.
Key features include:
| Criterion | Solution Fit | Key Considerations |
| Read and write throughput | XF960 targets AI training and checkpointing; Dell cites up to 2x large-object read throughput per node versus its closest competitor and up to 40 GB/s per node on the PowerEdge reference design | Figures are Dell internal analysis; reviewers specifically flag write performance as an area for improvement |
| IOPS | Not published as an IOPS figure; all-flash models are positioned for high-concurrency object access | Absence of published IOPS means proof-of-concept testing is needed for random small-object workloads |
| Latency | S3 over RDMA with GPU-direct access, positioned for low-latency GenAI data access | The RDMA path depends on a supported network fabric being in place |
| Metadata performance | Not described as a separate metadata tier; a global namespace spans Virtual Data Centers | Deployment and configuration documentation is a repeated criticism in reviews |
| Scale-up vs. scale-out storage | Exascale scale-out with up to 16 nodes per rack, smart rebalancing on expansion, and software-defined deployment on PowerEdge | Appliance-plus-software model; cost is commonly cited as high, with requests for more flexible payment options |
| GPU cluster scalability | Storage engine within the Dell AI Data Platform, with GPU-direct S3 access and published KV cache offload work | Reviewers cite gaps in S3 compatibility against some third-party applications |


Best for: Multi-tenant exabyte object storage for cloud and sovereign environments
Strengths: Independent scaling axes, separate metadata tier, deep resilience
Things to consider: Complex to administer; rebalancing can take a long time
Scality RING is distributed file and object storage built for cloud, service-provider and sovereign cloud environments, running from multi-petabyte to exabyte scale in one logical namespace. It has been in production since 2009 across more than 1,000 enterprises in over 70 countries.
Its architecture, called MultiScale, allows capacity, performance, tenants, sites and protocols to scale independently rather than in lockstep, so adding performance does not require adding capacity and vice versa. In practice, deployments commonly place metadata on NVMe flash while object data sits on higher-capacity media, which is what makes large object counts workable.
Resilience is a core selling point: 14 nines of data durability, erasure coding and self-healing across sites, and stretched-cluster and zero-RPO patterns that survive whole data centers going offline. Scality’s AIConnect technology positions RING data for AI pipelines, and the company markets specifically to neoclouds and GPU cloud operators.
Key features include:
| Criterion | Solution Fit | Key Considerations |
| Read and write throughput | Performance scales independently of capacity under MultiScale, with predictable throughput as workloads grow | HDD-based deployments suit backup and archive better than hot AI data; an NVMe-oriented configuration exists for AI workloads |
| IOPS | High-concurrency access with hard tenant isolation across many workloads on one cluster | Reviewers report the system struggling when operation and query rates climb very high |
| Latency | Latency described as staying flat as workload grows across multi-tenant, multi-site topologies | Achieved latency depends heavily on media; NVMe is typically used to accelerate metadata access |
| Metadata performance | Metadata can be deployed on dedicated servers and faster media, separately from object data | Reviewers cite limits in the metadata engine and occasional need to restart metadata components |
| Scale-up vs. scale-out storage | Multi-petabyte to exabyte in one namespace with hot capacity additions and no downtime for expansion | Rebalancing after adding nodes is reported to take a long time, in some cases weeks |
| GPU cluster scalability | AIConnect prepares data for AI pipelines, with an explicit neocloud and GPU cloud use case | Administration is complex and assumes in-house Linux expertise; patch frequency and release quality are recurring complaints |

Best for: Keeping large GPU clusters fed during training and inference
Strengths: Linear throughput and IOPS scaling, data reduction, deploy anywhere
Things to consider: Small public review base; premium all-flash pricing
WEKA NeuralMesh is a storage and memory platform for AI and HPC workloads. WEKA states that throughput and IOPS scale linearly with consistently low latency as clusters grow from ten nodes to ten thousand, and that the same software runs on-premises, in hybrid environments, in hyperscale clouds and in AI clouds.
A large part of the current product story is data reduction. NeuralMesh applies similarity-based compression that captures structural redundancy across blocks, which targets the kind of near-duplicate data that checkpoints and container layers produce and that byte-exact deduplication misses. Reduction runs in the background so writes commit to NVMe at native speed.
WEKA combines that with thin provisioning, snapshots, single-hop writes, drive sharing and AlloyFlash hybrid flash, and reports up to 1.9x reads, 13x write throughput and 3x lower latency from the combination.
Key features include:
| Criterion | Solution Fit | Key Considerations |
| Read and write throughput | Up to 1.9x reads and 13x write throughput reported from the combined efficiency features, with reduction adding under 5% write overhead | Figures are vendor comparisons against prior behaviour rather than absolute bandwidth numbers at a stated cluster size |
| IOPS | Throughput and IOPS both stated to scale linearly from ten to ten thousand nodes | No absolute IOPS figure appears on the product page, so testing at target scale is needed |
| Latency | Consistently low latency as clusters grow, with up to 3x lower latency from the efficiency pillars and tokens moving at memory speed | Dependent on NVMe media and network fabric in the deployment |
| Metadata performance | Metadata handling at scale is listed as a topic on the product page but not detailed there | Not addressed in enough detail on the product page to compare directly against systems that publish metadata figures |
| Scale-up vs. scale-out storage | Scales from ten to ten thousand nodes with the same software across on-premises, hybrid, hyperscale and AI cloud environments | All-flash NVMe economics; the one substantive public review flags cost relative to alternatives on the market |
| GPU cluster scalability | Built around keeping GPUs fed and extending GPU memory, with a Kubernetes operator and multitenancy for shared clusters | Very small public review base, with only two published G2 reviews, so independent validation is limited |


Best for: Real-time inference and RAG pipelines across distributed AI environments
Strengths: Distributed metadata indexing, KV cache acceleration, S3 plus POSIX
Things to consider: Newer than DDN’s parallel file system line; thin review base
DDN Infinia is an AI data engine that orchestrates data across distributed AI environments, unifying fragmented silos into one pipeline with low-latency access. It is built on DDN’s HPC background but aimed at enterprise inference, RAG and analytics rather than the classic supercomputing throughput workloads its EXAScaler parallel file system serves.
Metadata is central to the design. Infinia holds structured, semi-structured and unstructured data in one platform with high-speed metadata indexing, supporting search, filtering and retrieval across very large datasets, and it supports millions of metadata tags per object. An event engine can trigger tagging automatically at ingest.
A high-performance KV store and KV cache keep embeddings, vectors and inference state near compute, which is where DDN’s headline claims come from: 75% reduction in token cost, 22x faster RAG performance and 25x lower time to first byte.
Key features include:
| Criterion | Solution Fit | Key Considerations |
| Read and write throughput | TB/s throughput and fast checkpointing cited across a single namespace serving both S3 and POSIX | EXAScaler remains DDN’s answer for large-scale training throughput, so some environments will run both products |
| IOPS | Not published as an IOPS figure; the KV store keeps embeddings and inference state close to compute for fast repeated access | Performance is framed around object and KV operations rather than conventional block IOPS |
| Latency | Ultra-low latency access, with 25x lower time to first byte and 22x faster RAG performance claimed | Claims are vendor benchmarks tied specifically to inference and RAG rather than general-purpose access |
| Metadata performance | Fully distributed metadata with high-speed indexing and millions of tags per object, plus event-driven tagging at ingest | Independent third-party validation of the metadata claims is limited in public sources |
| Scale-up vs. scale-out storage | Cloud-native and hardware-agnostic scale-out, with placement spanning edge, core and cloud | Erasure coding implies a minimum cluster size, and Infinia is newer than DDN’s established HPC line |
| GPU cluster scalability | Positioned around maximizing GPU utilization, with multi-tenancy and QoS controls for large GPU clusters | Environments needing both training throughput and inference latency may end up deploying two DDN products |

Best for: Unified all-flash file and object for AI training, inference and analytics
Strengths: One OS for NFS, SMB and S3; non-disruptive upgrades; //EXA scale
Things to consider: Premium pricing; cloud integration and docs flagged
Everpure FlashBlade, formerly Pure Storage FlashBlade, is a scale-out all-flash array for unstructured data that runs native NFS, SMB and S3 on a single operating system, Purity//FB, without protocol gateways or separate systems per protocol. Pure Storage legally changed its corporate name to Everpure in February 2026, with products transitioning to the new naming through the year.
The lineup separates by workload profile. FlashBlade//E targets cost-sensitive capacity work such as active archive, backup and content libraries. FlashBlade//S handles up to 650 GB/s throughput for AI training and inference, analytics and HPC. FlashBlade//EXA is the extreme-scale option, quoted at over 10 TB/s throughput for large-scale AI and HPC workloads and neocloud deployments, and described as delivering the metadata performance those workloads demand.
Blades, modules and software are upgraded without downtime, avoiding the refresh cycles that normally interrupt storage estates.
Key features include:
| Criterion | Solution Fit | Key Considerations |
| Read and write throughput | FlashBlade//S reaches up to 650 GB/s and FlashBlade//EXA over 10 TB/s, with 55 GB/s sustained per chassis | The headline throughput figure belongs to //EXA, a distinct model, so the relevant number depends on which line is bought |
| IOPS | All-flash unified file and object with throughput and IOPS described as scaling without tuning marathons | IOPS is not published per model, leaving random-access performance to be validated in testing |
| Latency | All-flash architecture under Purity//FB, positioned for low-latency access to data-intensive applications | Reviewers note garbage collection routines affecting performance in environments with high churn rates |
| Metadata performance | FlashBlade//EXA is described as delivering the capacity, throughput and metadata performance that modern AI and HPC demand | Metadata detail sits in the technical brief rather than the product page, and applies to //EXA specifically |
| Scale-up vs. scale-out storage | Scale-out blade architecture with non-disruptive upgrades and a modular path up or down within the product line | Proprietary appliance rather than software on standard servers, and pricing is consistently described as premium |
| GPU cluster scalability | Powers NVIDIA DGX SuperPOD deployments, with //EXA aimed at large-scale AI workflows and neoclouds | Cloud integration, documentation and support responsiveness are recurring criticisms in reviews |

Related content: Read our guide to AI storage providers
AI storage for large-scale machine learning needs to balance throughput, latency, IOPS, metadata performance, and expansion behavior rather than optimize for a single benchmark. The right architecture depends on whether workloads are dominated by large sequential reads, small random access, checkpoint writes, inference, or shared GPU clusters. Teams should compare published results against their own hardware, network fabric, dataset structure, and concurrency levels, then validate performance with realistic tests before committing to a platform.
See Cloudian first in Google Search