Simple AI Training

No-Ops AI/ML Training Service that Train with Code and Data

Simple AI Training is a fully managed service that enables large-scale model training without the need to build AI infrastructure. Data scientists and ML engineers can be allocated resources immediately and focus entirely on training without the burden of infrastructure management.

Overview

01

04

Service Architecture

A workflow diagram showing the Simple AI Training process with seven numbered steps. On the left, a User (Data Scientist or ML Engineer) initiates the process. Step 1: Prepare Training Data involves storing data in File Storage, Container Registry, and Object Storage within the User Domain. Step 2: Prepare Training Scripts. Step 3: Configure Training connects to the Master Cluster containing a Training Service Portal and Multi Cluster Scheduler. Step 4: Mount & Copy Data transfers data from the User Domain to the training infrastructure. Step 5: Start Training directs the process to the GPU Cluster, which includes a Training Operator, Mixed Workload Controller, and GPU Fail Over managing multiple GPU Nodes with individual GPUs. Step 6: Execute Training processes the workload on the GPU resources. Step 7: Complete Training & Save Results returns the processed results back to the User Domain storage systems

Key Features

  • Streamlined model training
    1. Serverless environment : Automate everything from infrastructure settings, data loading, model training, and storing results to focus on training
    2. Distributed training : Automatically distributes workloads across multiple instances for large-scale models or datasets
    3. Custom container : Supports Bring Your Own Container(BYOC), which allows model to be trained using custom docker images(Available in September 2026)
    4. Warm pools : Instantly reuses instances for consecutive jobs to cut down infrastructure provisioning times
  • Training availability and security
    1. GPU failover : Automatically detects and recovers from GPU failures during training to ensure uninterrupted execution
    2. Checkpointing : Saves intermediate states to Object Storage, allowing seamless training resumption from the last checkpoint if a failure occurs
    3. Secure access to data : Data is stored in Object Storage in the cloud to ensure security without worrying about external leakage
  • Flexible pricing options
    1. Spot Training plan : Reduces costs for low-priority workloads by utilizing idle GPUs. It automatically manages task interruption and resumption for maximum cost efficiency
    2. On-demand plan : Servers can be created at any time for uninterrupted and reliable learning
    3. Contracted plan : Contracted plan enables GPU priority at lower cost than no-contract plan

Let’s talk

Whether you’re looking for a specific business solution or just need some questions answered, we’re here to help