Skip to content

Commit 9f6c596

Browse files
committed
Add A4XMAX kimi-k2 FP8mx 1024 GPUs recipe
1 parent eef14ea commit 9f6c596

8 files changed

Lines changed: 905 additions & 0 deletions

File tree

Lines changed: 20 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,20 @@
1+
# Copyright 2026 Google LLC
2+
#
3+
# Licensed under the Apache License, Version 2.0 (the "License");
4+
# you may not use this file except in compliance with the License.
5+
# You may obtain a copy of the License at
6+
#
7+
# http://www.apache.org/licenses/LICENSE-2.0
8+
#
9+
# Unless required by applicable law or agreed to in writing, software
10+
# distributed under the License is distributed on an "AS IS" BASIS,
11+
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
12+
# See the License for the specific language governing permissions and
13+
# limitations under the License.
14+
15+
apiVersion: v2
16+
name: a4x_max_jobset_workload
17+
description: a4x_max_jobset_workload
18+
type: application
19+
version: 0.1.0
20+
appVersion: "1.16.0"
Lines changed: 154 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,154 @@
1+
<!-- mdformat global-off -->
2+
# Pretrain kimi-k2 workloads on a4x-max GKE Node pools with Nvidia Megatron-Bridge Framework
3+
4+
This recipe outlines the steps for running a kimi-k2 pretraining
5+
workload on [a4x-max GKE Node pools](https://cloud.google.com/kubernetes-engine) by using the
6+
[Megatron-Bridge pretraining workload](https://github.com/NVIDIA-NeMo/Megatron-Bridge).
7+
8+
## Orchestration and deployment tools
9+
10+
For this recipe, the following setup is used:
11+
12+
- Orchestration - [Google Kubernetes Engine (GKE)](https://cloud.google.com/kubernetes-engine)
13+
- Pretraining job configuration and deployment - A Helm chart is used to
14+
configure and deploy the [Kubernetes Jobset](https://kubernetes.io/blog/2025/03/23/introducing-jobset) resource which manages the execution of the
15+
[Megatron-Bridge pretraining workload](https://github.com/NVIDIA-NeMo/Megatron-Bridge).
16+
17+
## Test environment
18+
19+
This recipe has been optimized for and tested with the following configuration:
20+
21+
- GKE cluster
22+
Please follow Cluster Toolkit [instructions](https://github.com/GoogleCloudPlatform/cluster-toolkit/)
23+
to create your a4x-max GKE cluster.
24+
25+
## Training dataset
26+
27+
This recipe uses a mock pretraining dataset provided by the Megatron-Bridge framework.
28+
29+
## Docker container image
30+
31+
This recipe uses the following docker images:
32+
33+
- `nvcr.io/nvidia/nemo:26.04.01`
34+
**Installed Plugins:**
35+
- `nccl-gib-plugins` version: 1.1.2-1
36+
37+
## Run the recipe
38+
39+
From your client workstation, complete the following steps:
40+
41+
### Configure environment settings
42+
43+
Set the environment variables to match your environment:
44+
45+
```bash
46+
export PROJECT_ID=<PROJECT_ID>
47+
export CLUSTER_REGION=<CLUSTER_REGION>
48+
export CLUSTER_NAME=<CLUSTER_NAME>
49+
export GCS_BUCKET=<GCS_BUCKET> # Note: path should not be prefixed with gs://
50+
export KUEUE_NAME=<KUEUE_NAME>
51+
export HF_TOKEN=<YOUR_HF_TOKEN>
52+
```
53+
54+
Replace the following values:
55+
56+
- `<PROJECT_ID>`: your Google Cloud project ID.
57+
- `<CLUSTER_REGION>`: the region where your cluster is located.
58+
- `<CLUSTER_NAME>`: the name of your GKE cluster.
59+
- `<GCS_BUCKET>`: the name of your Cloud Storage bucket. Don't include the `gs://` prefix.
60+
- `<KUEUE_NAME>`: the name of the Kueue local queue. The default queue created by the cluster toolkit is `a4x-max`. Make sure to verify the name of the local queue in your cluster.
61+
- `<YOUR_HF_TOKEN>`: Your HuggingFace token.
62+
63+
Set the default project:
64+
65+
```bash
66+
gcloud config set project $PROJECT_ID
67+
```
68+
69+
### Get the recipe
70+
71+
Clone the `gpu-recipes` repository and set a reference to the recipe folder.
72+
73+
```
74+
git clone https://github.com/ai-hypercomputer/gpu-recipes.git
75+
cd gpu-recipes
76+
export REPO_ROOT=`git rev-parse --show-toplevel`
77+
export RECIPE_ROOT=$REPO_ROOT/training/a4x-max/kimi-k2/megatron-bridge-gke/nemo260401/1024gpus-fp8mx-seq4096-gbs16384/recipe
78+
cd $RECIPE_ROOT
79+
```
80+
81+
### Get cluster credentials
82+
83+
```
84+
gcloud container clusters get-credentials $CLUSTER_NAME --region $CLUSTER_REGION
85+
```
86+
87+
### Configure and submit a pretraining job
88+
89+
#### Using 256 node (1024 gpus) fp8mx precision
90+
To execute the job with the default settings, run the following command from
91+
your client:
92+
93+
```bash
94+
cd $RECIPE_ROOT
95+
export WORKLOAD_NAME=$USER-a4x-max-kimi-k2-1024gpus
96+
helm install $WORKLOAD_NAME . -f values.yaml \
97+
--set-file workload_launcher=launcher.sh \
98+
--set workload.image=nvcr.io/nvidia/nemo:26.04.01 \
99+
--set volumes.gcsMounts[0].bucketName=${GCS_BUCKET} \
100+
--set volumes.gcsMounts[0].mountPath=/job-logs \
101+
--set workload.envs[0].value=/job-logs/$WORKLOAD_NAME \
102+
--set queue=${KUEUE_NAME}
103+
```
104+
105+
**Examples**
106+
107+
- To set the number of training steps to 100, run the following command from
108+
your client:
109+
110+
```bash
111+
cd $RECIPE_ROOT
112+
export WORKLOAD_NAME=$USER-a4x-max-kimi-k2-1024gpus
113+
helm install $WORKLOAD_NAME . -f values.yaml \
114+
--set-file workload_launcher=launcher.sh \
115+
--set workload.image=nvcr.io/nvidia/nemo:26.04.01 \
116+
--set volumes.gcsMounts[0].bucketName=${GCS_BUCKET} \
117+
--set volumes.gcsMounts[0].mountPath=/job-logs \
118+
--set workload.envs[0].value=/job-logs/$WORKLOAD_NAME \
119+
--set queue=${KUEUE_NAME} \
120+
--set workload.arguments[0]="trainer.max_steps=100"
121+
```
122+
123+
### Monitor the job
124+
125+
To check the status of pods in your job, run the following command:
126+
127+
```
128+
kubectl get pods | grep $USER-a4x-max-kimi-k2-256node
129+
```
130+
131+
Replace the following:
132+
133+
- JOB_NAME_PREFIX - your job name prefix. For example $USER-a4x-max-kimi-k2-256node.
134+
135+
To get the logs for one of the pods, run the following command:
136+
137+
```
138+
kubectl logs POD_NAME
139+
```
140+
141+
Information about the training job's progress, including crucial details such as
142+
loss, step count, and step time, is generated by the rank 0 process.
143+
This process runs on the pod whose name begins with
144+
`JOB_NAME_PREFIX-workload-0-0`.
145+
For example: `$USER-a4x-max-kimi-k2-256node-workload-0-0-s9zrv`.
146+
147+
### Uninstall the Helm release
148+
149+
You can delete the job and other resources created by the Helm chart. To
150+
uninstall Helm, run the following command from your client:
151+
152+
```bash
153+
helm uninstall $USER-a4x-max-kimi-k2-256node
154+
```

0 commit comments

Comments
 (0)