@@ -48,7 +48,10 @@ possibly recommended flavors can be created, or the user can set a file containi
4848### GPU table
4949
5050The most commonly used datacenter GPUs are listed here, showing what GPUs (or partitions
51- of a GPU) result in what GPU part of the flavor name.
51+ of a GPU) result in what GPU part of the flavor name. We provide these for convenience; most
52+ values are from data sheets and not based on own testing. Providers must look up the values
53+ (SMs/CUs/EUs and VRAM) really provided to users and correctly fill these into the SCS names.
54+ This is in particular true for the MIG configurations.
5255
5356#### Nvidia (` N ` )
5457
@@ -95,6 +98,15 @@ No MIG support, 128 Cuda Cores and 4 Tensor Cores per SM.
9598| L40G | 568 | 18176 | 142 | 48G GDDR6 | ` GNl-142h-48 ` |
9699| L40S | 568 | 18176 | 142 | 48G GDDR6 | ` GNl-142hh-48 ` |
97100
101+ | Nvidia GPU | Tensor C | Cuda Cores | SMs | VRAM | SCS name piece |
102+ | -------------| ----------| ------------| -----| -----------| ----------------|
103+ | RTX2000 Ada | 88 | 2816 | 22 | 16G GDDR6 | ` GNl-22-16 ` |
104+ | RTX4000 Ada | 192 | 6144 | 48 | 20G GDDR6 | ` GNl-48-20 ` |
105+ | RTX4500 Ada | 240 | 7680 | 60 | 24G GDDR6 | ` GNl-60-24 ` |
106+ | RTX5000 Ada | 400 | 12800 | 100 | 32G GDDR6 | ` GNl-100-32 ` |
107+ | RTX5880 Ada | 440 | 14080 | 110 | 48G GDDR6 | ` GNl-110-48 ` |
108+ | RTX6000 Ada | 568 | 18176 | 142 | 48G GDDR6 | ` GNl-142-48 ` |
109+
98110##### Grace Hopper (` g ` )
99111
100112These have MIG support and 128 Cuda Cores and 4 Tensor Cores per SM.
@@ -112,6 +124,39 @@ These have MIG support and 128 Cuda Cores and 4 Tensor Cores per SM.
112124[ +] The precise numbers for the 1/7 MIG configurations are not known by the author of
113125this document and need validation.
114126
127+ ##### Blackwell (` b ` ) and Blackwell Ultra (` u ` )
128+
129+ These have MIG support and 128 Cuda Cores and 4 Tensor Cores per SM.
130+
131+ | Nvidia GPU | Fraction | Tensor C | Cuda Cores | SMs | VRAM | SCS GPU name |
132+ | ------------| ----------| ----------| ------------| -----| ------------| ----------------|
133+ | GB200 | 1/1 | 640 | 20480 | 160 | 192G HBM3e | ` GNb-160-192h ` |
134+ | GB200 | 1/2 | 320 | 10240 | 80 | 96G HBM3e | ` GNb-80-96h ` |
135+ | GB200 | 2/7 | 88+ | 5632+ | 44+| 45G HBM3e+| ` GNb-44-45h ` + |
136+ | GB200 | 1/7 | 44+ | 2816+ | 22+| 23G HBM3e+| ` GNb-22-23h ` + |
137+ | GB300 | 1/1 | 640 | 20480 | 160 | 288G HBM3e | ` GNu-160-288h ` |
138+ | GB300 | 1/2 | 320 | 10240 | 80 | 144G HBM3e | ` GNu-80-144h ` |
139+ | ... |
140+
141+ [ +] The precise numbers for the 1/7 MIG configurations are not known by the author of
142+ this document and need validation.
143+
144+ Note that Blackwell Ultra tensor cores have significant enough changes vs. Blackwell that we
145+ gave the BW Ultra GPUs a new letter ` u ` . In particular, FP4 tensor performance is over 150%
146+ of std. Blackwell and has more Special Function Units (which helps attention) but has
147+ regressed INT8 performance.
148+
149+ | Nvidia GPU | Fraction | Tensor C | Cuda Cores | SMs | VRAM | SCS GPU name |
150+ | -----------------------| ----------| ----------| ------------| -----| ------------| ----------------|
151+ | RTX Pro2000 Blackwell | 1/1 | 136 | 4352 | 34 | 16G GDDR7 | ` GNb-34-16 ` |
152+ | RTX Pro4000 Blackwell | 1/1 | 280 | 8960 | 70 | 24G GDDR7 | ` GNb-70-24 ` |
153+ | RTX Pro4500 Blackwell | 1/1 | 328 | 10496 | 82 | 32G GDDR7 | ` GNb-82-32 ` |
154+ | RTX Pro5000 Blackwell | 1/1 | 440 | 14080 | 110 | 72G GDDR7 | ` GNb-110-72 ` |
155+ | RTX Pro5000 Blackwell | 1/2 | 220 | 7040 | 55 | 36G GDDR7 | ` GNb-55-36 ` |
156+ | RTX Pro6000 Blackwell | 1/1 | 752 | 26064 | 188 | 96G GDDR7 | ` GNb-188-96 ` |
157+ | RTX Pro6000 Blackwell | 1/2 | 376 | 13032 | 94 | 48G GDDR7 | ` GNb-94-48 ` |
158+ | RTX Pro6000 Blackwell | 1/4 | 188 | 6516 | 47 | 24G GDDR7 | ` GNb-47-24 ` |
159+
115160#### AMD Radeon (` A ` )
116161
117162##### CDNA 2 (` 2 ` )
@@ -130,28 +175,71 @@ SRIOV partitioning is possible, resulting in pass-through for
130175up to 8 partitions, somewhat similar to Nvidia MIG. 4 Tensor
131176Cores and 64 Stream Processors per CU.
132177
133- | AMD GPU | Tensor C | Stream Proc | CUs | VRAM | SCS name piece |
134- | -------------| ----------| -------------| -----| ------------| ----------------|
135- | Inst MI300X | 1216 | 19456 | 304 | 192G HBM3 | ` GA3-304-192h ` |
136- | Inst MI325X | 1216 | 19456 | 304 | 288G HBM3 | ` GA3-304-288h ` |
178+ | AMD GPU | Tensor C | Stream Proc | CUs | VRAM | SCS name piece |
179+ | -------------| ----------| -------------| -----| ------------| -----------------|
180+ | Inst MI300X | 1216 | 19456 | 304 | 192G HBM3 | ` GA3-304-192h ` |
181+ | Inst MI325X | 1216 | 19456 | 304 | 288G HBM3 | ` GA3-304-288h ` |
182+
183+ ##### CDNA 4 (` 4 ` )
184+
185+ SRIOV partitioning is possible, resulting in pass-through for
186+ up to 8 partitions, somewhat similar to Nvidia MIG. 4 Tensor
187+ Cores and 64 Stream Processors per CU.
188+
189+ | AMD GPU | Tensor C | Stream Proc | CUs | VRAM | SCS name piece |
190+ | -------------| ----------| -------------| -----| ------------| -----------------|
191+ | Inst MI350X | 1024 | 16384 | 256 | 288G HBM3e | ` GA4-256-288h ` |
192+ | Inst MI355X | 1024 | 16384 | 256 | 288G HBM3e | ` GA4-256h-288h ` |
193+
194+ The Instinct MI355X has a higher watttage and thus slightly higher clocks
195+ than the MI350X but is otherwise identical - we can thus use the ` h ` modifier
196+ to identify the higher performance version.
197+
198+ ##### Workstation RDNA 3 (` 3.1 ` ) and 4 (` 4.1 ` )
199+
200+ 2 Tensor Cores and 64 Stream Processors per CU.
201+
202+ | AMD Radeon | Tensor C | Stream Proc | CUs | VRAM | SCS name piece |
203+ | --------------| ----------| -------------| -----| ------------| -----------------|
204+ | Pro W7900 | 196 | 6144 | 96 | 48G GDDR6 | ` GA3.1-96-48 ` |
205+ | AI Pro R9700 | 128 | 4096 | 64 | 32G GDDR6 | ` GA4.1-64-32 ` |
137206
138207Note that we previously assumed more similarity of consumer RDNA-x with
139- server CDNA-x that actually is the case; the RDNA-x cards now use ` x.1 `
208+ server CDNA-x than actually is the case; the RDNA-x cards now use ` x.1 `
140209(since v3.3 as of Oct 2025) to be able to differentiate them. We will
141210tolerate potential rare cases of old installations calling RDNA-x as
142- generation ` x ` for the time being.
211+ generation ` x ` for the time being. If AMD executes on the merging with
212+ UDNA-5, we will avoid this split in the future.
143213
144214#### intel Xe (` I ` )
145215
146216##### Xe-HPC (Ponte Vecchio) (` 3 ` )
147217
148- 1 EU corresponds to one Tensor Core and contains 128 Shading Units.
218+ One EU corresponds to one Tensor Core and contains 128 Shading Units.
149219
150220| intel DC GPU | Tensor C | Shading U | EUs | VRAM | SCS name part |
151221| --------------| ----------| -----------| -----| ------------| ----------------|
152222| Max 1100 | 56 | 7168 | 56 | 48G HBM2e | ` GI3-56-48h ` |
153223| Max 1550 | 128 | 16384 | 128 | 128G HBM2e | ` GI3-128-128h ` |
154224
225+ ##### Workstation cards Arc B (` 4 ` )
226+
227+ One EU has one tensor core and 16 shading units.
228+
229+ | intel GPU | Tensor C | Shading U | EUs | VRAM | SCS name part |
230+ | -------------| ----------| -----------| -----| ------------| ----------------|
231+ | Arc Pro B50 | 128 | 2048 | 128 | 16G GDDR6 | ` GI4-128-16 ` |
232+ | Arc Pro B60 | 160 | 2560 | 160 | 24G GDDR6 | ` GI4-160-24 ` |
233+ | Arc Pro B65 | 160 | 2560 | 160 | 32G GDDR6 | ` GI4-160-32 ` |
234+ | Arc Pro B70 | 256 | 4096 | 256 | 32G GDDR6 | ` GI4-256-32 ` |
235+
236+ #### Consumer cards
237+
238+ Note that we don't recommend using consumer cards.
239+ That said, the schema allows to specify them and for example do PCI pass-through
240+ of Nvidia RTX4080S (` GNl-80-16 ` ), RTX4090 (` GNl-128-24 ` ), RTX5080S (` GNb-84-24 ` ),
241+ RTX5090 (` GNb-170-32 ` ), or AMD Radeon RX7900XTX (` GA3.1-96-24 ` ).
242+
155243## Automated tests
156244
157245The following testcases [ are implemented] ( https://github.com/SovereignCloudStack/standards/tree/main/Tests/iaas/openstack_test.py ) :
0 commit comments