|
| 1 | +--- |
| 2 | +title: Set up the project and establish a baseline |
| 3 | +weight: 3 |
| 4 | + |
| 5 | +### FIXED, DO NOT MODIFY |
| 6 | +layout: learningpathall |
| 7 | +--- |
| 8 | + |
| 9 | +## Before you begin |
| 10 | + |
| 11 | +To get started, you need an Arm Linux system with SVE support. Suitable cloud instances include AWS Graviton3 or Graviton4, Microsoft Cobalt 100, or Google Axion. The examples in this Learning Path were tested on Ubuntu 26.04. |
| 12 | + |
| 13 | +You also need an AI coding assistant with the Arm MCP server configured. Supported assistants include [GitHub Copilot](/install-guides/github-copilot/), [Kiro CLI](/install-guides/kiro-cli/), [Claude Code](/install-guides/claude-code/), [Gemini CLI](/install-guides/gemini/), and [Codex CLI](/install-guides/codex-cli/). See the [Arm MCP server Learning Path](/learning-paths/servers-and-cloud-computing/arm-mcp-server/) for setup instructions. |
| 14 | + |
| 15 | +{{< notice Note >}} |
| 16 | +The AI responses shown are samples. Your AI assistant may word responses differently, include more or less detail, or structure the output differently depending on the tool and model you are using. Focus on the key concepts rather than the exact wording. |
| 17 | +{{< /notice >}} |
| 18 | + |
| 19 | +Start by installing the required software and check your system includes SVE. |
| 20 | + |
| 21 | +Install GCC and GNU Make: |
| 22 | + |
| 23 | +```bash |
| 24 | +sudo apt install gcc make -y |
| 25 | +``` |
| 26 | + |
| 27 | +Confirm your system has SVE by checking the CPU flags: |
| 28 | + |
| 29 | +```bash |
| 30 | +lscpu | grep -i sve |
| 31 | +``` |
| 32 | + |
| 33 | +The output should include `sve` in the flags list: |
| 34 | + |
| 35 | +```output |
| 36 | +Flags: fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti |
| 37 | +``` |
| 38 | + |
| 39 | +If `sve` does not appear, the system does not support SVE and the final implementation in this Learning Path won't run correctly. |
| 40 | + |
| 41 | +## Create the project files |
| 42 | + |
| 43 | +On your Arm Neoverse system, create a working directory and add the source files. |
| 44 | + |
| 45 | +```bash |
| 46 | +mkdir adler32-sve && cd adler32-sve |
| 47 | +``` |
| 48 | + |
| 49 | +Use an editor to copy the scalar implementation to `adler32-simple.c`: |
| 50 | + |
| 51 | +```c |
| 52 | +/* |
| 53 | + * Adler-32 checksum — simple scalar C implementation |
| 54 | + */ |
| 55 | +#include <stdint.h> |
| 56 | +#include <stddef.h> |
| 57 | + |
| 58 | +#define MOD_ADLER 65521 |
| 59 | + |
| 60 | +uint32_t adler32(const uint8_t *data, size_t len) |
| 61 | +{ |
| 62 | + uint32_t a = 1; |
| 63 | + uint32_t b = 0; |
| 64 | + |
| 65 | + for (size_t i = 0; i < len; i++) { |
| 66 | + a = (a + data[i]) % MOD_ADLER; |
| 67 | + b = (b + a) % MOD_ADLER; |
| 68 | + } |
| 69 | + |
| 70 | + return (b << 16) | a; |
| 71 | +} |
| 72 | +``` |
| 73 | +
|
| 74 | +Create the test and benchmark harness in `adler32-test.c`: |
| 75 | +
|
| 76 | +```c |
| 77 | +/* |
| 78 | + * Adler-32 test and benchmark harness |
| 79 | + */ |
| 80 | +#include <stdint.h> |
| 81 | +#include <stddef.h> |
| 82 | +#include <stdio.h> |
| 83 | +#include <stdlib.h> |
| 84 | +#include <time.h> |
| 85 | +
|
| 86 | +extern uint32_t adler32(const uint8_t *data, size_t len); |
| 87 | +
|
| 88 | +static void fill_random(uint8_t *buf, size_t len, uint32_t seed) |
| 89 | +{ |
| 90 | + for (size_t i = 0; i < len; i++) { |
| 91 | + seed = seed * 1103515245 + 12345; |
| 92 | + buf[i] = (uint8_t)(seed >> 16); |
| 93 | + } |
| 94 | +} |
| 95 | +
|
| 96 | +static double time_sec(void) |
| 97 | +{ |
| 98 | + struct timespec ts; |
| 99 | + clock_gettime(CLOCK_MONOTONIC, &ts); |
| 100 | + return ts.tv_sec + ts.tv_nsec * 1e-9; |
| 101 | +} |
| 102 | +
|
| 103 | +static int test_known_vectors(void) |
| 104 | +{ |
| 105 | + /* RFC 1950: adler32("Wikipedia") = 0x11E60398 */ |
| 106 | + const uint8_t wiki[] = "Wikipedia"; |
| 107 | + uint32_t got = adler32(wiki, 9); |
| 108 | + uint32_t expect = 0x11E60398; |
| 109 | +
|
| 110 | + printf("Correctness: adler32(\"Wikipedia\") = 0x%08X ", got); |
| 111 | + if (got == expect) { |
| 112 | + printf("[PASS]\n"); |
| 113 | + return 0; |
| 114 | + } else { |
| 115 | + printf("[FAIL] expected 0x%08X\n", expect); |
| 116 | + return 1; |
| 117 | + } |
| 118 | +} |
| 119 | +
|
| 120 | +static void benchmark(const char *label, size_t size, int iters) |
| 121 | +{ |
| 122 | + uint8_t *data = malloc(size); |
| 123 | + if (!data) { perror("malloc"); exit(1); } |
| 124 | + fill_random(data, size, (uint32_t)size); |
| 125 | +
|
| 126 | + volatile uint32_t sink = adler32(data, size); |
| 127 | + (void)sink; |
| 128 | +
|
| 129 | + double t0 = time_sec(); |
| 130 | + for (int i = 0; i < iters; i++) |
| 131 | + sink = adler32(data, size); |
| 132 | + double elapsed = time_sec() - t0; |
| 133 | +
|
| 134 | + double mb = (double)size * iters / (1024.0 * 1024.0); |
| 135 | + printf(" %-8s %8zu bytes %6d iters %8.3f ms %8.1f MB/s checksum=0x%08X\n", |
| 136 | + label, size, iters, elapsed * 1000.0, mb / elapsed, sink); |
| 137 | +
|
| 138 | + free(data); |
| 139 | +} |
| 140 | +
|
| 141 | +int main(void) |
| 142 | +{ |
| 143 | + int fail = test_known_vectors(); |
| 144 | + if (fail) return 1; |
| 145 | +
|
| 146 | + printf("\nPerformance:\n"); |
| 147 | + benchmark("1 KB", 1024, 100000); |
| 148 | + benchmark("10 KB", 10 * 1024, 10000); |
| 149 | + benchmark("100 KB", 100 * 1024, 1000); |
| 150 | + benchmark("1 MB", 1024 * 1024, 100); |
| 151 | + benchmark("10 MB", 10 * 1024 * 1024, 10); |
| 152 | +
|
| 153 | + return 0; |
| 154 | +} |
| 155 | +``` |
| 156 | + |
| 157 | +Create a `Makefile`: |
| 158 | + |
| 159 | +```makefile |
| 160 | +CC = gcc |
| 161 | +CFLAGS = -O3 -mcpu=native -flto -Wall -Wextra |
| 162 | +LDFLAGS = -flto |
| 163 | +TARGET = adler32-test |
| 164 | +SRCS = adler32-simple.c adler32-test.c |
| 165 | + |
| 166 | +$(TARGET): $(SRCS) |
| 167 | + $(CC) $(CFLAGS) $(LDFLAGS) -o $@ $(SRCS) |
| 168 | + |
| 169 | +run: $(TARGET) |
| 170 | + ./$(TARGET) |
| 171 | + |
| 172 | +clean: |
| 173 | + rm -f $(TARGET) |
| 174 | + |
| 175 | +.PHONY: run clean |
| 176 | +``` |
| 177 | + |
| 178 | +The `-mcpu=native` flag tells GCC to optimize for the exact CPU you're running on, which enables SVE code generation on Neoverse processors that have SVE. |
| 179 | + |
| 180 | +### ASK AI: about compiler flags |
| 181 | + |
| 182 | +Before running anything, ask your AI assistant to confirm that your build setup is correct for SVE. |
| 183 | + |
| 184 | +```text |
| 185 | +My Makefile uses `-O3 -mcpu=native`. Does this enable SVE code generation on a Neoverse processor? Do I need any special flags for SVE intrinsics? |
| 186 | +``` |
| 187 | + |
| 188 | +A sample response is: |
| 189 | + |
| 190 | +```output |
| 191 | +Good question. Let me check what CPU is on your remote system and look up the relevant details. |
| 192 | +
|
| 193 | +Here's the full picture: |
| 194 | + |
| 195 | +Yes, `-O3 -mcpu=native` enables SVE code generation on your system. Your Neoverse-V2 processor supports SVE and SVE2, and GCC 15.2 |
| 196 | +with -mcpu=native resolves to armv9-a+sve2-aes+sve2-bitperm+sve2-sha3+.... I confirmed the compiler is already auto-vectorizing |
| 197 | +with SVE instructions (4 SVE instructions emitted for a simple dot-product loop). |
| 198 | + |
| 199 | +For SVE intrinsics, you need: |
| 200 | + |
| 201 | +1. No extra compiler flags — -mcpu=native already enables SVE/SVE2. You don't need to add -march=armv8-a+sve separately (that |
| 202 | + would actually be a downgrade from what -mcpu=native gives you on this CPU). |
| 203 | +2. Include the header in your source files: |
| 204 | + |
| 205 | + #include <arm_sve.h> |
| 206 | + |
| 207 | +That's it. Your current Makefile flags are correct for both auto-vectorized SVE code and hand-written SVE intrinsics. |
| 208 | + |
| 209 | +One caveat: since -mcpu=native detects the build machine's CPU, binaries built this way may not run on older Arm processors (e.g., |
| 210 | +Neoverse-N1/Graviton2). If portability matters, use -mcpu=neoverse-v2 explicitly to make the target clear. |
| 211 | + |
| 212 | +For more on SVE programming, Arm has a good learning path: Port Code to Arm SVE |
| 213 | +(https://learn.arm.com/learning-paths/servers-and-cloud-computing/sve/). |
| 214 | +``` |
| 215 | + |
| 216 | +The response explains that `-mcpu=native` enables SVE. It also provides useful info about running on other systems and confirm special flags, such as `-march=armv8-a+sve` are not needed. The response also tells you to include `<arm_sve.h>`. |
| 217 | + |
| 218 | +You also notice the reference to a Learning Path about SVE at the end. This confirms the Arm MCP server is consulted on answering the question. |
| 219 | + |
| 220 | +This is good information you'll need when you create the SVE source later. |
| 221 | + |
| 222 | +## Build and run the baseline |
| 223 | + |
| 224 | +Build and run the scalar version: |
| 225 | + |
| 226 | +```bash |
| 227 | +make run |
| 228 | +``` |
| 229 | + |
| 230 | +The output is similar to: |
| 231 | + |
| 232 | +```output |
| 233 | +Correctness: adler32("Wikipedia") = 0x11E60398 [PASS] |
| 234 | +
|
| 235 | +Performance: |
| 236 | + 1 KB 1024 bytes 100000 iters 255.619 ms 382.0 MB/s checksum=0x37B4063C |
| 237 | + 10 KB 10240 bytes 10000 iters 252.763 ms 386.4 MB/s checksum=0x5FA40D73 |
| 238 | + 100 KB 102400 bytes 1000 iters 255.239 ms 382.6 MB/s checksum=0x2378D74C |
| 239 | + 1 MB 1048576 bytes 100 iters 261.659 ms 382.2 MB/s checksum=0x058B4B08 |
| 240 | + 10 MB 10485760 bytes 10 iters 262.388 ms 381.1 MB/s checksum=0x285FF1B1 |
| 241 | +``` |
| 242 | + |
| 243 | +Your numbers will differ depending on your specific Neoverse processor and memory configuration. Make a note of the MB/s values for the 1 MB and 10 MB cases, as these are your baseline numbers to compare against after each optimization. |
| 244 | + |
| 245 | +### ASK AI: about auto-vectorization |
| 246 | + |
| 247 | +Now ask your AI assistant a question that many developers wonder about: |
| 248 | + |
| 249 | +```text |
| 250 | +Can GCC auto-vectorize my adler32 function with SVE if I just use `-mcpu=native`? What would prevent auto-vectorization? |
| 251 | +``` |
| 252 | + |
| 253 | +A sample response is: |
| 254 | + |
| 255 | +```output |
| 256 | +No, GCC cannot auto-vectorize your adler32 function. It tried every vector mode (including SVE's VNx* modes) and failed on all of |
| 257 | + them. |
| 258 | + |
| 259 | + The report tells us exactly why: "not vectorized: unsupported use in stmt" on line 16. |
| 260 | + |
| 261 | + The problem is the modulo operation inside the loop: |
| 262 | + |
| 263 | + a = (a + data[i]) % MOD_ADLER; // ← loop-carried dependency + modulo |
| 264 | + b = (b + a) % MOD_ADLER; // ← depends on 'a' above, also modulo |
| 265 | + |
| 266 | + Two things kill auto-vectorization here: |
| 267 | + |
| 268 | + 1. Loop-carried dependency chain: Each iteration's b depends on the current iteration's a, and each a depends on the previous a. |
| 269 | + The compiler can't compute multiple iterations in parallel because every result feeds into the next. |
| 270 | + 2. Modulo by a non-power-of-2: % 65521 is an integer division/remainder — there's no cheap SIMD equivalent. The compiler can't |
| 271 | + vectorize this operation across lanes. |
| 272 | + |
| 273 | + This is why hand-written SVE intrinsics are valuable here. The standard trick is to: |
| 274 | + |
| 275 | + - Defer the modulo — accumulate partial sums across SVE lanes without reducing mod 65521 every iteration, only applying the modulo |
| 276 | + periodically (every ~5552 bytes, the largest safe batch before uint32_t overflow). |
| 277 | + - Break the dependency — use vector lanes to accumulate a and b contributions independently, then reduce at the end. |
| 278 | +``` |
| 279 | + |
| 280 | +The response explains that the modulo operation in every iteration (`% MOD_ADLER`) is the main blocker. The compiler can't easily prove that the intermediate values won't overflow in a way that changes the result when operations are reordered. The loop-carried dependency between iterations also makes it difficult. |
| 281 | + |
| 282 | +Since auto-vectorization won't work, you need to restructure the algorithm before SVE can be applied effectively. The restructuring is explained in the next two sections. |
| 283 | + |
| 284 | +## What you've learned and what's next |
| 285 | + |
| 286 | +In this section: |
| 287 | + |
| 288 | +- You created the scalar Adler-32 implementation and benchmark harness |
| 289 | +- You recorded your baseline performance numbers |
| 290 | +- You learned that auto-vectorization won't work |
| 291 | + |
| 292 | +In the next section, you'll use the Arm MCP server to learn the core SVE concepts you need before writing any intrinsics code. |
0 commit comments