AI Research Cluster
Research, experiment and model development environment
Multiple Kubernetes clusters are issued and isolated using the same criteria, and control plane operations are automated to manage observation, upgrade, backup, and recovery in an integrated manner.
OVERVIEW
The AI platform starts with a single Kubernetes cluster, but quickly expands to development, validation, operations, factory, customer, and regional environments as the service grows. Managed K8s manages all of these environments to the same operating standards.
The expansion of AI services leads to an increase in the number of clusters. Rather than operating each environment individually, we standardize it within the Enterprise Kubernetes operating system.
Research, experiment and model development environment
Service development and integration testing
Verification, approval, and deployment preparation environment
Real AI service operating environment
Resources dedicated to high-performance learning and inference
Independent environment for each client/department
AI platform for each factory and production base
Distributed operating environment by region and institution
Even if the cluster is separatedOperational policy, security, observation, automationmoves based on one standard.
CHALLENGES
If each project creates an independent Kubernetes environment, versions, GPUs, networks, security and backup standards will vary, and the consistency of the entire enterprise AI platform will be lost.
It is difficult to maintain consistent operational quality because the Kubernetes version, Container Runtime, CNI, GPU Driver, Storage, and Monitoring configurations are different.
Project startup time increases as Kubernetes installation, GPU settings, Storage, Registry and Monitoring configuration are repeated each time.
Operators spend more time managing the platform by performing individual cluster health checks, upgrades, failure response, backup and recovery.
RBAC, Network Policy, Security Policy, GPU Resource Policy, and Backup Policy are applied differently for each project.
It is difficult to determine at a glance which clusters lack GPUs and which have idle resources, resulting in low utilization of investment.
SOLUTION
Kubernetes, GPU, Networking, Storage, Registry and Monitoring configuration are defined as a standard architecture.
Clusters required for new AI projects are quickly created without repeated installation and the initial configuration is automatically applied.
Cluster status, upgrade, backup, recovery and resource status are managed integratedly in QKS.
RBAC, Network Policy, Security Policy, Audit Log, and Backup Policy are applied consistently on a central basis.
View GPU utilization, workload, and service status distributed across multiple clusters from a single perspective.
We perform control plane management, cluster installation, upgrades, and auto-scaling to relieve the operational burden on K8s experts.
ARCHITECTURE
QKS integrates the creation and operation of multiple Kubernetes clusters, and applies the same AI Runtime, Data, Observation, and Automation system to each cluster.
When a new AI project starts, QKS creates a standard cluster, and ORKESTRIX · AkashiQ · SAMANDA are connected in the same structure.
CENTRAL CONTROL PLANE
Cluster provisioning, status observation, upgrade, backup/recovery, policy and resource status are centrally managed.
DEV · STAGE · PROD
FACTORY · REGION
DEPARTMENT · CUSTOMER
Multiple Cluster Creation, Observation, and Operation Management
Standard AI Runtime for each Cluster
Integrated storage of Dataset and Model Artifact
Cluster · GPU · log · event integrated observation
BUSINESS VALUE
Ensure platform consistency across your organization by managing all clusters to the same architectural and operational standards.
Quickly launch new AI projects and bases with standard templates and automatic provisioning.
Integrated management of multiple clusters on one platform reduces repetitive tasks and management burden.
Even as business units, customers, factories, and regions increase, we maintain the same operating system and expand stably.
Integrated observation of distributed GPU resources reduces idle resources and increases AI infrastructure utilization.
The cluster status is confirmed and inspected with just a natural language request, and even response measures are immediately suggested in the event of a failure.
APPLICATIONS
독립된 클러스터는 유지하면서 정책 · 자원 · 운영 상태를 중앙에서 통합 관리합니다.
공장별 AI 환경은 독립적으로 운영하고, 본사에서는 정책과 상태를 통합 관리합니다.