Skip to content

Data Lakehouse

Connect fragmented structured and unstructured data into a unified analytics foundation. Collect, refine, and store business-system, document, and equipment data, transforming it into data assets ready for analytics and AI.

OVERVIEW

Rather than just collecting data, Analytics and AI are ready for immediate use.

A data lakehouse combines the flexibility of a data lake that preserves the original, It combines the advantages of a data warehouse that reliably analyzes refined data into a single structure.

By collecting and standardizing distributed data, everything from BI analysis to AI model learning and natural language query is executed on a single data foundation.

01 · COLLECT

Collection of various data

Connect PLC · SCADA · MES · LIMS with standard protocols and CDC

02 · STORE

original preservation

Reliably store data while maintaining format and provenance

03 · REFINE

Refining/Standardization

Processed to suit quality and purpose of use in Bronze–Silver–Gold stages

04 · USE

Analysis and AI utilization

Expansion to BI, SQL analysis, model learning, and natural language data querying

scattered dataOne analytics asset you can trustSwitch to .

CHALLENGES

Although data is accumulating, There is still no connection between the systems.

As data is separated by business system and collection and purification methods change, more time is required to prepare for analysis and AI use.

01

Data is distributed across business systems.

The storage structure of each system, such as ERP, MES, LIMS, and document management, is different, making it difficult to understand company data at once.

02

Structured data and unstructured data are managed separately.

Numerical data and documents, images, and files in the DB are separated from each other, making integrated analysis and AI learning data configuration difficult.

03

It is difficult to collect data without burdening operational systems.

Bulk queries and repetitive extractions can affect operational DB performance, limiting real-time data utilization.

04

Manual extraction and Excel collection are repeated.

As each department downloads and combines the necessary data directly, analysis preparation time increases and the possibility of errors increases.

05

It is difficult to trust data quality and consistency.

It is difficult to track where data comes from and how it has changed, making it difficult to use it directly in decision-making and AI models.

SOLUTION

We connect the entire data lifecycle, from collection to quality and genealogy.

Load data while reducing the burden on the operating system, and create a structure suitable for analysis and AI utilization through open standards and step-by-step refinement.

01

No-load data collection/loading

It connects structured and unstructured data with CDC, API, file, and facility protocols, minimizes the performance impact of the operating system, and synchronizes in real time and batch.

02

Bronze – Silver – Gold Tablets

Separate the purpose of each step with original preservation, purification/standardization, and integrated data for analysis.

03

Open table format

Based on Apache Iceberg, it provides an analysis data structure that is not dependent on a specific vendor.

04

Distributed SQL integrated analysis

Connect to Trino, Spark, etc. to analyze multiple data sources with standard SQL.

05

Automatic determination of data quality

Automatically checks for omissions, duplications, range errors, and consistency issues and provides notifications.

06

Lineage/Governance Tracking

Track data sources and change history and manage access rights and audit standards.

ARCHITECTURE

From source systems to analysis and AI, uninterrupted data service flow

AkashiQ integrates data collection, purification, storage, quality, and lineage, and connects existing BI and AI services to utilize the same data assets.

Enterprise Data Lakehouse Flow

Structured and unstructured data are integrated into one open data base and used for analysis, AI, and natural language queries depending on the purpose.

ERP · MES · LIMS

CDC-based operating system no-load interconnection

Document management · EDMS

API · PDF · Word · Excel · Image

Equipment · PLC / SCADA

OPC-UA based real-time data collection

Files/external data

Batch · Object · Log · Other data

AkashiQ · DATA LAKEHOUSE

Collection · Purification · Storage · Quality · Genealogy

We prepare it step by step into the form required for analysis and AI while preserving the original data.

Bronze

original preservation

Silver

Refining/Standardization

Gold

Integrations for analytics

Apache Iceberg · Open Table FormatAutomatic determination and notification of data quality (DQ)Data Lineage · Authority · Governance

BI · Report

Linkage with existing analysis tools based on standard SQL

KosmosAI

Analysis · Model learning · Prediction · Anomaly detection

SAMANDA

Natural language data query Text-to-SQL

Based platform: ORKESTRIXInstallation and operation of on-premises Kubernetes environment
AkashiQ

Collection · Purification · Storage · Quality · Lineage · Governance

ORKESTRIX

Foundation for installation and operation of data platform

KosmosAI

Analysis based on loaded data and utilization of AI models

SAMANDA

Natural language data queries and Text-to-SQL extensions

BUSINESS VALUE

Reduce the time it takes to find and combine data, and accelerate the use of analytics and AI.

01 · SINGLE ACCESS

Single access window for enterprise data

Find and utilize data scattered across systems and departments through one catalog and query path.

02 · UNIFIED ANALYTICS

Integrated analysis between systems

ERP, MES, LIMS, and document data are connected to standard SQL and analyzed together across system boundaries.

03 · ZERO IMPACT

Operating system no-load synchronization

Utilizes CDC and streaming to synchronize data in real time and batch while minimizing performance impact on operational DB.

04 · OPEN STANDARD

Linkage with existing BI and standards

Connects existing BI and analysis tools based on open table format and standard SQL.

05 · TRUSTED DATA

Trusted Data Foundation

Through data quality judgment and genealogy, we confirm the source, processing process, and change history, and increase the reliability of analysis and AI.

APPLICATIONS

여러 시스템에 분산된 데이터를 하나의 분석 · AI 기반으로 통합합니다.

업무 · 문서 · 현장 데이터를 한곳에 연결해 분석과 AI가 활용할 수 있는 공통 데이터 기반을 구축합니다.

제약 · 제조 데이터 기반

MES · LIMS · 설비 · 품질 데이터를 연결해 공정 분석과 AI 활용을 위한 기반을 구축합니다.

  • 생산 · 설비 · 품질 데이터 수집
  • LOT · 배치 단위 통합 분석
  • 문서 · DB 데이터 통합 저장
  • 분석 · AI 모델용 데이터 제공