Skip to content

Managed K8s

Multiple Kubernetes clusters are issued and isolated using the same criteria, and control plane operations are automated to manage observation, upgrade, backup, and recovery in an integrated manner.

OVERVIEW

As clusters grow, operational standards remain unified.

The AI ​​platform starts with a single Kubernetes cluster, but quickly expands to development, validation, operations, factory, customer, and regional environments as the service grows. Managed K8s manages all of these environments to the same operating standards.

The expansion of AI services leads to an increase in the number of clusters. Rather than operating each environment individually, we standardize it within the Enterprise Kubernetes operating system.

01

AI Research Cluster

Research, experiment and model development environment

02

Development Cluster

Service development and integration testing

03

Stage Cluster

Verification, approval, and deployment preparation environment

04

Production Cluster

Real AI service operating environment

05

GPU Dedicated Cluster

Resources dedicated to high-performance learning and inference

06

Customer Cluster

Independent environment for each client/department

07

Factory Edge Cluster

AI platform for each factory and production base

08

Regional Cluster

Distributed operating environment by region and institution

Even if the cluster is separatedOperational policy, security, observation, automationmoves based on one standard.

CHALLENGES

As AI projects increase, cluster operations become more complex.

If each project creates an independent Kubernetes environment, versions, GPUs, networks, security and backup standards will vary, and the consistency of the entire enterprise AI platform will be lost.

01

Each project operates differently.

It is difficult to maintain consistent operational quality because the Kubernetes version, Container Runtime, CNI, GPU Driver, Storage, and Monitoring configurations are different.

02

We repeat the same building process for each AI project.

Project startup time increases as Kubernetes installation, GPU settings, Storage, Registry and Monitoring configuration are repeated each time.

03

Connect and manage multiple clusters individually.

Operators spend more time managing the platform by performing individual cluster health checks, upgrades, failure response, backup and recovery.

04

It is difficult to integrate operational and security policies.

RBAC, Network Policy, Security Policy, GPU Resource Policy, and Backup Policy are applied differently for each project.

05

There is no visible idle/shortage state of the distributed GPU.

It is difficult to determine at a glance which clusters lack GPUs and which have idle resources, resulting in low utilization of investment.

SOLUTION

Create an enterprise-wide Kubernetes environment with one operating standard.

01

Standard Cluster Template

Kubernetes, GPU, Networking, Storage, Registry and Monitoring configuration are defined as a standard architecture.

02

Automatic Creation and Provisioning

Clusters required for new AI projects are quickly created without repeated installation and the initial configuration is automatically applied.

03

Central integrated operation

Cluster status, upgrade, backup, recovery and resource status are managed integratedly in QKS.

04

Policy and security standardization

RBAC, Network Policy, Security Policy, Audit Log, and Backup Policy are applied consistently on a central basis.

05

GPU · Workload integrated observation

View GPU utilization, workload, and service status distributed across multiple clusters from a single perspective.

06

Delegating Managed K8s Operations

We perform control plane management, cluster installation, upgrades, and auto-scaling to relieve the operational burden on K8s experts.

ARCHITECTURE

Connects disparate clusters into a single operating system.

QKS integrates the creation and operation of multiple Kubernetes clusters, and applies the same AI Runtime, Data, Observation, and Automation system to each cluster.

Enterprise Multi-Cluster Reference Flow

When a new AI project starts, QKS creates a standard cluster, and ORKESTRIX · AkashiQ · SAMANDA are connected in the same structure.

CENTRAL CONTROL PLANE

QKS · Enterprise Kubernetes Platform

Cluster provisioning, status observation, upgrade, backup/recovery, policy and resource status are centrally managed.

Cluster TemplatePolicy & SecurityGPU VisibilityUpgradeBackup & RecoveryMulti-Tenant

DEV · STAGE · PROD

AI Project Cluster

ORKESTRIXAI Runtime
AkashiQDataset · Artifact
SAMANDAMonitoring · Incident

FACTORY · REGION

Edge & Site Cluster

ORKESTRIXGPU · Networking
AkashiQLocal Object Storage
SAMANDAAlert · Operations

DEPARTMENT · CUSTOMER

Dedicated Cluster

ORKESTRIXStandard Runtime
AkashiQIsolated Data
SAMANDAUnified Observability
QKS

Multiple Cluster Creation, Observation, and Operation Management

ORKESTRIX

Standard AI Runtime for each Cluster

AkashiQ

Integrated storage of Dataset and Model Artifact

SAMANDA

Cluster · GPU · log · event integrated observation

BUSINESS VALUE

Scale faster and reduce operational complexity.

01

Enterprise Kubernetes standardization

Ensure platform consistency across your organization by managing all clusters to the same architectural and operational standards.

02

Reduce new project construction time

Quickly launch new AI projects and bases with standard templates and automatic provisioning.

03

Reduce operational complexity

Integrated management of multiple clusters on one platform reduces repetitive tasks and management burden.

04

Secure AI platform scalability

Even as business units, customers, factories, and regions increase, we maintain the same operating system and expand stably.

05

Optimized GPU utilization

Integrated observation of distributed GPU resources reduces idle resources and increases AI infrastructure utilization.

06

Agentic AI-based operation support

The cluster status is confirmed and inspected with just a natural language request, and even response measures are immediately suggested in the event of a failure.

APPLICATIONS

여러 조직과 거점의 AI 환경을 하나의 운영 기준으로 관리합니다.

독립된 클러스터는 유지하면서 정책 · 자원 · 운영 상태를 중앙에서 통합 관리합니다.

다공장 AI 플랫폼

공장별 AI 환경은 독립적으로 운영하고, 본사에서는 정책과 상태를 통합 관리합니다.

  • 공장별 표준 AI 실행 환경
  • 전체 공장 클러스터 통합 관리
  • 생산 · AI 데이터셋 통합 저장
  • 운영 상태 · 자원 중앙 관측