SK하이닉스 Builds 30PB AI Storage System Based on Open Source
We examine the future direction of enterprise AI infrastructure expansion through the example of SK하이닉스, which has integrated large-scale AI storage based on Kubernetes and MinIO.
SK하이닉스 Builds Its Own 30PB Storage System for Generative AI
SK하이닉스 has built its own 30PB high-performance S3 storage system based on open-source technology to operate its in-house generative AI platform.
While the company previously used commercial storage products, it has now internalized its infrastructure based on Kubernetes and MinIO. The company reported that approximately 13,000 employees are integrating dozens of AI services to process about 800,000 queries per day.
A notable aspect of this deployment is that it goes beyond simply expanding storage capacity; it involves transitioning the large-scale data infrastructure required for AI service operations to a cloud-native environment.
Key Technical Configuration
The storage system is designed as an S3-based distributed object storage architecture, with Kubernetes used to operate hundreds of servers as a single cluster.
MinIO was adopted as the storage engine, and the network was configured to reduce bottlenecks in data transmission by utilizing a spine-leaf architecture along with BGP and eBPF.
Furthermore, by optimizing everything from the Linux kernel to SSD handling, the team focused on boosting the performance of S3 storage to a level close to that of physical servers.
From 30PB to 1EiB: The Continuous Expansion of AI Infrastructure
SK하이닉스 does not view its current 30PB capacity as the final stage.
Starting from 0.4PB, the company has expanded in stages—through 3PB, 4PB, 5PB, and 20PB—to reach the current 30PB capacity, and aims to build a 1EiB-class storage cluster in the future.
This is also linked to the expansion of the company’s in-house generative AI platform, “GAIA (GAI.A).” GAI.A utilizes AI agents in various business areas, including semiconductor design and process optimization, equipment maintenance, policy analysis, and human resources systems, and plans to expand to A2A (Agent-to-Agent)-based orchestration where multiple agents collaborate.
QUANTUM View
A key point to note in this case study is that the proliferation of enterprise AI is driving the expansion not only of LLMs and AI agents themselves but also of the underlying data, storage, and Kubernetes infrastructure that supports them.
As AI services move beyond proof-of-concept (PoC) to actual in-house operations, the volume of unstructured data to be stored and the number of requests to be processed increase rapidly as the number of users and agents grows. Consequently, to ensure stable AI operations, it is becoming increasingly important to adopt an approach that considers a scalable Kubernetes environment alongside high-performance object storage, networking, and operational automation.
In particular, the fact that they have internalized a large-scale storage infrastructure based on Kubernetes and open source—without relying on commercial appliances—can be seen as an example that illustrates a shift in how enterprises will build their AI infrastructure in the future.
View Original Article
THE ELEC
“SK하이닉스 Builds Its Own 30PB In-House Storage System”
Reporter Lee Seok-jin · September 18, 2025.