diff --git a/.github/install-guide-review.md b/.github/install-guide-review.md new file mode 100644 index 0000000000..5e02b776a8 --- /dev/null +++ b/.github/install-guide-review.md @@ -0,0 +1,137 @@ +# Install Guide Review Tracker + +Progress: 0 / 97 reviewed + +Legend: [x] reviewed | [ ] not started | πŸ”’ hidden/not published + +--- + +## Jason Andrews (57 guides) + +| Guide | Title | Published | Last Updated | Reviewed | +|-------|-------|-----------|--------------|----------| +| skopeo | Skopeo | βœ“ | 2025-02-05 | [ ] | +| aws-cli | AWS CLI | βœ“ | 2025-02-05 | [ ] | +| helm | Helm | βœ“ | 2025-02-05 | [ ] | +| py-woa | Python for Windows on Arm | βœ“ | 2025-02-05 | [ ] | +| sysbox | Sysbox | βœ“ | 2025-02-05 | [ ] | +| eksctl | AWS EKS CLI (eksctl) | βœ“ | 2025-02-10 | [ ] | +| kubectl | Kubectl | βœ“ | 2025-02-10 | [ ] | +| bedrust | Bedrust - invoke models on Amazon Bedrock | βœ“ | 2025-04-10 | [ ] | +| anaconda | Anaconda | βœ“ | 2025-04-25 | [ ] | +| ansible | Ansible | βœ“ | 2025-04-25 | [ ] | +| armie | Arm Instruction Emulator (armie) | βœ“ | 2025-04-25 | [ ] | +| aws-copilot | AWS Copilot CLI | βœ“ | 2025-04-25 | [ ] | +| gcloud | Google Cloud Platform (GCP) CLI | βœ“ | 2025-04-25 | [ ] | +| nerdctl | Nerdctl | βœ“ | 2025-04-25 | [ ] | +| papi | Performance API (PAPI) | βœ“ | 2025-04-25 | [ ] | +| perf | Perf for Linux on Arm (LinuxPerf) | βœ“ | 2025-04-25 | [ ] | +| pulumi | Pulumi | βœ“ | 2025-04-25 | [ ] | +| terraform | Terraform | βœ“ | 2025-04-25 | [ ] | +| topdown-tool | Telemetry Solution (Topdown Methodology) | βœ“ | 2025-04-25 | [ ] | +| cmake | CMake | βœ“ | 2025-04-26 | [ ] | +| docker/ | Docker (multi-page) | βœ“ | 2025-04-26 | [ ] | +| hyper-v | Hyper-V on Arm | βœ“ | 2025-04-26 | [ ] | +| aws_access_keys | AWS Credentials | βœ“ | 2025-04-30 | [ ] | +| aws-sam-cli | AWS SAM CLI | βœ“ | 2025-04-30 | [ ] | +| azure_login | Azure Authentication | βœ“ | 2025-04-30 | [ ] | +| porting-advisor | Porting Advisor for Graviton | βœ“ | 2025-04-30 | [ ] | +| vscode-tunnels | VS Code Tunnels | βœ“ | 2025-04-30 | [ ] | +| ssh | SSH | βœ“ | 2025-05-02 | [ ] | +| finch | Finch on Arm Linux | βœ“ | 2025-05-22 | [ ] | +| browsers/ | Browsers on Arm (multi-page) | βœ“ | 2025-06-18 | [ ] | +| go | Go | βœ“ | 2025-07-16 | [ ] | +| nomachine | NoMachine | βœ“ | 2025-07-16 | [ ] | +| powershell | PowerShell | βœ“ | 2025-07-16 | [ ] | +| swift | Swift | βœ“ | 2025-07-16 | [ ] | +| oc | OpenShift CLI (oc) | βœ“ | 2025-07-24 | [ ] | +| tkn | Tekton CLI (tkn) | βœ“ | 2025-07-24 | [ ] | +| azure-cli | Azure CLI | βœ“ | 2025-07-28 | [ ] | +| pytorch | PyTorch | βœ“ | 2025-10-14 | [ ] | +| multipass | Multipass | βœ“ | 2025-10-17 | [ ] | +| vnc | VNC on Arm Linux | βœ“ | 2025-11-24 | [ ] | +| kiro-cli | Kiro CLI | βœ“ | 2025-12-10 | [ ] | +| linux-migration-tools | Arm Linux Migration Tools | βœ“ | 2026-01-08 | [ ] | +| aperf | APerf | βœ“ | 2026-01-19 | [ ] | +| gcc/ | GNU Compiler (multi-page) | βœ“ | 2026-01-28 | [ ] | +| gemini | Gemini CLI | βœ“ | 2026-01-30 | [ ] | +| java | Java | βœ“ | 2026-01-30 | [ ] | +| openvscode-server | OpenVSCode Server | βœ“ | 2026-01-30 | [ ] | +| sbt | sbt | βœ“ | 2026-01-30 | [ ] | +| wperf | WindowsPerf (wperf) | βœ“ | 2026-01-30 | [ ] | +| asct | Arm System Characterization Tool | πŸ”’ hidden | 2026-02-06 | [ ] | +| git-woa | Git for Windows on Arm | βœ“ | 2026-02-26 | [ ] | +| dotnet | .NET SDK | βœ“ | 2026-03-02 | [ ] | + +--- + +## Ronan Synnott (20 guides) + +| Guide | Title | Published | Last Updated | Reviewed | +|-------|-------|-----------|--------------|----------| +| pdh/ | Arm Product Download Hub (multi-page) | βœ“ | 2025-02-05 | [ ] | +| ambaviz | Arm AMBA Viz | βœ“ | 2025-04-25 | [ ] | +| armds | Arm Development Studio | βœ“ | 2025-04-25 | [ ] | +| ipexplorer | Arm IP Explorer | βœ“ | 2025-04-25 | [ ] | +| keilstudiocloud | Arm Keil Studio Cloud | βœ“ | 2025-04-25 | [ ] | +| mdk | Arm Keil uVision | βœ“ | 2025-04-25 | [ ] | +| socrates | Arm Socrates | βœ“ | 2025-04-25 | [ ] | +| successkits | Arm Success Kits | βœ“ | 2025-04-25 | [ ] | +| armclang | Arm Compiler for Embedded | βœ“ | 2025-04-26 | [ ] | +| avh | Arm Virtual Hardware | βœ“ | 2025-04-26 | [ ] | +| keilstudio_vs | Arm Keil Studio for VS Code | βœ“ | 2025-04-26 | [ ] | +| license/ | Arm Software Licensing (multi-page) | βœ“ | 2025-04-26 | [ ] | +| streamline | Arm Streamline | βœ“ | 2025-04-26 | [ ] | +| llvm-embedded | LLVM Embedded Toolchain for Arm | βœ“ | 2025-04-30 | [ ] | +| rust_embedded | Rust for Embedded Applications | βœ“ | 2025-04-30 | [ ] | +| ams | Arm Performance Studio | βœ“ | 2025-08-08 | [ ] | +| fm_fvp/ | Arm Fast Models and FVPs (multi-page) | βœ“ | 2025-08-20 | [ ] | +| stm32_vs | STM32 extensions for VS Code | βœ“ | 2025-12-18 | [ ] | +| mcuxpresso_vs | NXP MCUXpresso for VS Code | βœ“ | 2026-01-14 | [ ] | +| cmsis-toolbox | CMSIS-Toolbox | βœ“ | 2026-01-30 | [ ] | + +--- + +## Pareena Verma (8 guides) + +| Guide | Title | Published | Last Updated | Reviewed | +|-------|-------|-----------|--------------|----------| +| windows-sandbox-woa | Windows Sandbox for Windows on Arm | βœ“ | 2025-04-25 | [ ] | +| llvm-woa | LLVM toolchain for Windows on Arm | βœ“ | 2025-04-30 | [ ] | +| github-copilot | GitHub Copilot | βœ“ | 2025-12-16 | [ ] | +| claude-code | Claude Code | βœ“ | 2026-01-15 | [ ] | +| pytorch-woa | PyTorch for Windows on Arm | βœ“ | 2026-01-30 | [ ] | +| atp | Arm Total Performance | πŸ”’ hidden | 2026-02-06 | [ ] | +| vs-woa | Visual Studio for Windows on Arm | βœ“ | 2026-02-26 | [ ] | +| armpl | Arm Performance Libraries | βœ“ | 2026-03-02 | [ ] | + +--- + +## Florent Lebeau (3 guides) + +| Guide | Title | Published | Last Updated | Reviewed | +|-------|-------|-----------|--------------|----------| +| forge | Linaro Forge | βœ“ | 2025-04-30 | [ ] | +| gfortran | GFortran | βœ“ | 2025-04-30 | [ ] | +| acfl | Arm Compiler for Linux | βœ“ | 2026-01-29 | [ ] | + +--- + +## Other authors (14 guides) + +| Guide | Title | Author | Published | Last Updated | Reviewed | +|-------|-------|--------|-----------|--------------|----------| +| cyclonedds | Cyclone DDS | Odin Shen | βœ“ | 2025-04-04 | [ ] | +| ros2 | ROS - Robot Operating System | Odin Shen | βœ“ | 2025-04-25 | [ ] | +| oci-cli | Oracle Cloud Infrastructure (OCI) CLI | Daniel Gubay | βœ“ | 2025-04-26 | [ ] | +| aws-greengrass-v2 | AWS IoT Greengrass | Michael Hall | βœ“ | 2025-05-02 | [ ] | +| dcperf | DCPerf | Kieran Hejmadi | βœ“ | 2025-09-22 | [ ] | +| container | Container CLI for macOS | Rani Chowdary Mandepudi | βœ“ | 2025-09-29 | [ ] | +| bolt | BOLT | Jonathan Davies | βœ“ | 2025-10-22 | [ ] | +| arduino-pico | Arduino core for the Raspberry Pi Pico | Michael Hall | βœ“ | 2025-11-11 | [ ] | +| windows-perf-vs-extension | Visual Studio Extension for WindowsPerf | Nader Zouaoui | βœ“ | 2025-11-11 | [ ] | +| windows-perf-wpa-plugin | Windows Performance Analyzer (WPA) plugin | Alaaeddine Chakroun | βœ“ | 2025-12-02 | [ ] | +| streamline-cli | Streamline CLI Tools | Julie Gaskin | βœ“ | 2025-12-18 | [ ] | +| codex-cli | Codex CLI | Joe Stech | βœ“ | 2026-01-08 | [ ] | +| fvps-on-macos | AVH FVPs on macOS | Christopher Seidl | βœ“ | 2026-01-16 | [ ] | +| rust | Rust for Linux Applications | Mathias Brossard | βœ“ | 2026-01-26 | [ ] | diff --git a/.wordlist.txt b/.wordlist.txt index 74fdeb8d1d..293f80ff0d 100644 --- a/.wordlist.txt +++ b/.wordlist.txt @@ -5784,4 +5784,45 @@ timescaledb tokenizer's tokenizers trainingyt -upsert \ No newline at end of file +upsert +Dequantizes +DuckDB +FMUL +LUT +LUTI +MinIO +Polars +Ritvik +SCVTF +Schleicher +Trino +VLx +Zm +bcm +cpuinfo +dequantization +dequantize +dequantizes +desynchronization +ebe +goldberg +kr +lhs +minio +minioadmin +ncg +numerics +pipelined +pushdown +pyarrow +repro +submatrices +submatrix +svbool +svcntw +svdup +svexp +svfloat +svptrue +svst +vexpq diff --git a/assets/contributors.csv b/assets/contributors.csv index b1ed1f8230..42918ac77b 100644 --- a/assets/contributors.csv +++ b/assets/contributors.csv @@ -112,4 +112,5 @@ Yahya Abouelseoud,Arm,,,, Steve Suzuki,Arm,,,, Qixiang Xu,Arm,,,, Phalani Paladugu,Arm,phalani-paladugu,phalani-paladugu,, -Richard Burton,Arm,Burton2000,,, \ No newline at end of file +Richard Burton,Arm,Burton2000,,, +Asier Arranz,NVIDIA,,asierarranz,,asierarranz.com \ No newline at end of file diff --git a/content/install-guides/armpl.md b/content/install-guides/armpl.md index 8ecda292a0..7b0d425d40 100644 --- a/content/install-guides/armpl.md +++ b/content/install-guides/armpl.md @@ -207,13 +207,13 @@ module avail The output should be similar to: ```output -armpl/26.01_gcc +armpl/26.01.0_gcc ``` Load the appropriate module: ```console -module load armpl/26.01_gcc +module load armpl/26.01.0_gcc ``` You can now compile and test the examples included in the `/opt/arm//examples/`, or `//examples/` directory, if you have installed to a different location than the default. diff --git a/content/install-guides/dotnet.md b/content/install-guides/dotnet.md index 6598c7d79e..04b5f5368b 100644 --- a/content/install-guides/dotnet.md +++ b/content/install-guides/dotnet.md @@ -16,7 +16,9 @@ tool_install: true weight: 1 --- -The [.NET SDK](https://dotnet.microsoft.com/en-us/) is a free, open-source, and cross-platform development environment that provides a broad set of tools and libraries for building applications. You can use it to create a variety of applications including web apps, mobile apps, desktop apps, and cloud services. +The [.NET SDK](https://dotnet.microsoft.com/en-us/) is a free, open-source, cross-platform development environment that provides tools and libraries for building applications. You can use it to create web apps, mobile apps, desktop apps, cloud services, and more. + +.NET 10 is the latest Long Term Support (LTS) release. .NET 9 (Standard Term Support) and .NET 8 (LTS) are also available. This guide defaults to .NET 10, with instructions for selecting an alternative version. The .NET SDK is available for Linux distributions on Arm-based systems. @@ -46,18 +48,20 @@ Select the one that works best for you. ### How can I install .NET SDK using the Linux package manager? -Use `apt` to install .NET SDK on Ubuntu and Debian: +Use `apt` to install the .NET 10 SDK on Ubuntu and Debian: ```bash -sudo apt-get install -y dotnet-sdk-8.0 +sudo apt-get update && sudo apt-get install -y dotnet-sdk-10.0 ``` -Use `dnf` to install .NET SDK on Fedora: +Use `dnf` to install the .NET 10 SDK on Fedora: ```console -sudo dnf install dotnet-sdk-8.0 +sudo dnf install dotnet-sdk-10.0 ``` +To install a different version, replace `10.0` with the version you need. For example, use `dotnet-sdk-9.0` for .NET 9 or `dotnet-sdk-8.0` for .NET 8. + If the .NET SDK is not found in your package manager, you can install it using a script. ### How can I install .NET SDK using a script? @@ -70,28 +74,30 @@ To install the .NET SDK using a script, follow the instructions below: wget https://dot.net/v1/dotnet-install.sh ``` -2. Run the script (it will install .NET SDK 8 under the folder .dotnet): +2. Run the script to install the .NET SDK under the `$HOME/.dotnet` folder. -You have some options to specify the version you want to install. +You have several options for specifying the version. -To install the latest long term support (LTS) version, run: +To install the latest LTS version (.NET 10), run: ```bash bash ./dotnet-install.sh ``` -To install the latest version, run: +To install the latest available version (including STS releases), run: ```bash bash ./dotnet-install.sh --version latest ``` -To install a specific version, run: +To install a specific version channel, use the `--channel` flag. For example, to install .NET 10: ```bash -bash ./dotnet-install.sh --channel 8.0 +bash ./dotnet-install.sh --channel 10.0 ``` +You can also use `--channel 9.0` for .NET 9 or `--channel 8.0` for .NET 8. + 3. Add the .dotnet folder to the PATH environment variable: ```bash @@ -102,54 +108,56 @@ You can also add the search path to your `$HOME/.bashrc` so it is set for all ne ## How do I verify the .NET SDK installation? -To check that the installation was successful, type: +To check that the installation was successful, run: ```bash dotnet --list-sdks ``` -The output is printed: +The output is similar to: ```output -8.0.105 [/usr/lib/dotnet/sdk] +10.0.103 [/usr/lib/dotnet/sdk] ``` -To print more information, run the following command: +For more detailed information about your installation, run: ```bash dotnet --info ``` -More details about your installation are printed: +The output is similar to: ```output .NET SDK: - Version: 8.0.105 - Commit: eae90abaaf - Workload version: 8.0.100-manifests.796a77f8 + Version: 10.0.103 + Commit: c2435c3e0f + Workload version: 10.0.100-manifests.a62d7899 + MSBuild version: 18.0.11+c2435c3e0 Runtime Environment: OS Name: ubuntu OS Version: 24.04 OS Platform: Linux - RID: ubuntu.24.04-arm64 - Base Path: /usr/lib/dotnet/sdk/8.0.105/ + RID: linux-arm64 + Base Path: /home/ubuntu/.dotnet/sdk/10.0.103/ .NET workloads installed: - Workload version: 8.0.100-manifests.796a77f8 There are no installed workloads to display. +Configured to use workload sets when installing new manifests. +No workload sets are installed. Run "dotnet workload restore" to install a workload set. Host: - Version: 8.0.5 + Version: 10.0.3 Architecture: arm64 - Commit: 087e15321b + Commit: c2435c3e0f .NET SDKs installed: - 8.0.105 [/usr/lib/dotnet/sdk] + 10.0.103 [/home/ubuntu/.dotnet/sdk] .NET runtimes installed: - Microsoft.AspNetCore.App 8.0.5 [/usr/lib/dotnet/shared/Microsoft.AspNetCore.App] - Microsoft.NETCore.App 8.0.5 [/usr/lib/dotnet/shared/Microsoft.NETCore.App] + Microsoft.AspNetCore.App 10.0.3 [/home/ubuntu/.dotnet/shared/Microsoft.AspNetCore.App] + Microsoft.NETCore.App 10.0.3 [/home/ubuntu/.dotnet/shared/Microsoft.NETCore.App] Other architectures found: None @@ -167,29 +175,29 @@ Download .NET: https://aka.ms/dotnet/download ``` +The exact version numbers depend on when you install and which updates are available. + ## How can I run a simple example to confirm the .NET SDK is working? -To test the .NET SDK installation, create a new hello world console application: +Create a new console application to verify that the .NET SDK works correctly: ```bash dotnet new console -o myapp ``` -Change to the new directory and run: +Change to the new directory and run the application: ```bash cd myapp dotnet run ``` -The expected output in the console is: +The expected output is: ```output -Hello World! +Hello, World! ``` You are ready to use the .NET SDK on Arm Linux. -You can find more information about .NET on Arm in the [AWS Graviton Technical Guide](https://github.com/aws/aws-graviton-getting-started/blob/main/dotnet.md). - -Explore .NET examples by visiting the [Learning Center](https://dotnet.microsoft.com/en-us/learn). +Explore more .NET examples by visiting the [Learning Center](https://dotnet.microsoft.com/en-us/learn). diff --git a/content/learning-paths/cross-platform/automate-mcp-with-testcontainers/_index.md b/content/learning-paths/cross-platform/automate-mcp-with-testcontainers/_index.md new file mode 100644 index 0000000000..cb9e4e21fb --- /dev/null +++ b/content/learning-paths/cross-platform/automate-mcp-with-testcontainers/_index.md @@ -0,0 +1,70 @@ +--- +title: Automate MCP Server testing using Pytest and Testcontainers + +draft: true +cascade: + draft: true + +minutes_to_complete: 60 + +who_is_this_for: This is an introductory topic for software developers and QA engineers who want to automate integration testing of MCP (Model Context Protocol) servers using Testcontainers and PyTest. + +learning_objectives: + - Set up Testcontainers with PyTest for containerized testing of MCP servers + - Write and run integration tests that validate MCP server functionality + - Configure GitHub Actions to automate MCP server testing in CI/CD pipelines + +prerequisites: + - A computer with [Docker](/install-guides/docker/) and Python 3.11 or later installed + - Basic familiarity with Python, PyTest, and container concepts + - Familiarity with the [Model Context Protocol (MCP)](https://modelcontextprotocol.io/) specification + +author: Neethu Elizabeth Simon + +### Tags +skilllevels: Introductory +subjects: CI-CD +armips: + - Neoverse + - Cortex-A +operatingsystems: + - Linux + - macOS + - Windows +tools_software_languages: + - Python + - Pytest + - Docker + - GitHub Actions + - Testcontainers + - MCP + +shared_path: true +shared_between: + - servers-and-cloud-computing + - laptops-and-desktops + +further_reading: + - resource: + title: Arm MCP Server GitHub Repository + link: https://github.com/arm/mcp + type: website + - resource: + title: Testcontainers for Python Documentation + link: https://testcontainers-python.readthedocs.io/ + type: documentation + - resource: + title: Model Context Protocol Specification + link: https://modelcontextprotocol.io/ + type: website + - resource: + title: PyTest Documentation + link: https://docs.pytest.org/ + type: documentation + +### FIXED, DO NOT MODIFY +# ================================================================================ +weight: 1 # _index.md always has weight of 1 to order correctly +layout: "learningpathall" # All files under learning paths have this same wrapper +learning_path_main_page: "yes" # This should be surfaced when looking for related content. Only set for _index.md of learning path content. +--- diff --git a/content/learning-paths/cross-platform/automate-mcp-with-testcontainers/_next-steps.md b/content/learning-paths/cross-platform/automate-mcp-with-testcontainers/_next-steps.md new file mode 100644 index 0000000000..727b395ddd --- /dev/null +++ b/content/learning-paths/cross-platform/automate-mcp-with-testcontainers/_next-steps.md @@ -0,0 +1,8 @@ +--- +# ================================================================================ +# FIXED, DO NOT MODIFY THIS FILE +# ================================================================================ +weight: 21 # The weight controls the order of the pages. _index.md always has weight 1. +title: "Next Steps" # Always the same, html page title. +layout: "learningpathall" # All files under learning paths have this same wrapper for Hugo processing. +--- diff --git a/content/learning-paths/cross-platform/automate-mcp-with-testcontainers/github-actions-ci.md b/content/learning-paths/cross-platform/automate-mcp-with-testcontainers/github-actions-ci.md new file mode 100644 index 0000000000..595bd41535 --- /dev/null +++ b/content/learning-paths/cross-platform/automate-mcp-with-testcontainers/github-actions-ci.md @@ -0,0 +1,197 @@ +--- +title: Configure GitHub Actions for CI/CD +weight: 5 + +### FIXED, DO NOT MODIFY +layout: learningpathall +--- + +## Why use GitHub Actions for MCP testing? + +GitHub Actions provide automated CI/CD directly in your repository. For MCP server testing, it offers: + +- **Arm runner support**: GitHub provides native Arm64 runners for building and testing. +- **Docker integration**: Runners come with Docker pre-installed. +- **Automatic triggers**: Tests run on every push and pull request. +- **Parallel execution**: Multiple jobs can run simultaneously. + +## Create the workflow file + +Create a GitHub Actions workflow file at `.github/workflows/integration-tests.yml`: + +```yaml +name: Integration Tests + +on: + push: + pull_request: + +jobs: + integration-tests: + runs-on: ubuntu-24.04-arm + steps: + - name: Checkout + uses: actions/checkout@v4 + + - name: Set up Python + uses: actions/setup-python@v5 + with: + python-version: "3.11" + cache: pip + + - name: Install test dependencies + run: | + python -m pip install --upgrade pip + pip install -r mcp-local/tests/requirements.txt + + - name: Build MCP Docker image + run: docker buildx build -f mcp-local/Dockerfile -t arm-mcp . + + - name: Run integration tests + env: + MCP_IMAGE: arm-mcp:latest + run: pytest -v mcp-local/tests/test_mcp.py +``` + +## Understand the workflow configuration + +The workflow uses the `ubuntu-24.04-arm` runner, which is a GitHub-hosted Arm64 runner. This ensures that both the Docker build and the tests execute natively on Arm hardware. + +Key aspects of the configuration: + +### Trigger events + +```yaml +on: + push: + pull_request: +``` + +This configuration triggers the workflow on every push to any branch and on pull request events. You can restrict this to specific branches if needed: + +```yaml +on: + push: + branches: [main, develop] + pull_request: + branches: [main] +``` + +### Python caching + +```yaml +- name: Set up Python + uses: actions/setup-python@v5 + with: + python-version: "3.11" + cache: pip +``` + +The `cache: pip` option caches Python packages between runs, significantly speeding up subsequent workflow executions. + +### Environment variables + +```yaml +- name: Run integration tests + env: + MCP_IMAGE: arm-mcp:latest + run: pytest -v mcp-local/tests/test_mcp.py +``` + +The `MCP_IMAGE` environment variable tells the test suite which Docker image to use. The test code reads this with: + +```python +image = os.getenv("MCP_IMAGE", constants.MCP_DOCKER_IMAGE) +``` + +## Add test result artifacts + +Enhance the workflow to save test results as artifacts for debugging: + +```yaml + - name: Run integration tests + env: + MCP_IMAGE: arm-mcp:latest + run: pytest -v mcp-local/tests/test_mcp.py --junitxml=test-results.xml + + - name: Upload test results + uses: actions/upload-artifact@v4 + if: always() + with: + name: test-results + path: test-results.xml +``` + +The `if: always()` condition ensures test results upload even when tests fail. + +## Add a build matrix for multiple platforms + +To test on both Arm64 and x86_64, use a matrix strategy: + +```yaml +jobs: + integration-tests: + strategy: + matrix: + runner: [ubuntu-24.04-arm, ubuntu-latest] + runs-on: ${{ matrix.runner }} + steps: + # ... same steps as before +``` + +This runs the integration tests in parallel on both architectures. + +## Monitor workflow runs + +After pushing the workflow file, navigate to the Actions tab in your GitHub repository. Each workflow run shows: + +- Build steps and their status +- Execution time for each step +- Log output for debugging failures +- Artifacts for download + +## Troubleshoot common issues + +If the workflow fails, check these common causes: + +**Docker build timeout**: The initial image build can take 10+ minutes. GitHub Actions has a default timeout of 360 minutes per job, but individual steps might need explicit timeouts: + +```yaml +- name: Build MCP Docker image + timeout-minutes: 30 + run: docker buildx build -f mcp-local/Dockerfile -t arm-mcp . +``` + +**Container startup issues**: If tests fail with timeout errors, the MCP server might not be starting correctly. Add debug output: + +```yaml +- name: Run integration tests + env: + MCP_IMAGE: arm-mcp:latest + run: | + docker run --rm arm-mcp:latest echo "Container starts successfully" + pytest -v -s mcp-local/tests/test_mcp.py +``` + +**Rate limiting**: If tests query external services, you might encounter rate limits. Consider adding retry logic or using mock responses for CI environments. + +## What you've accomplished and what's next + +In this section: +- You created a GitHub Actions workflow for automated testing. +- You learned how to use Arm64 runners for native execution. +- You added test artifacts and multi-platform support. +- You explored troubleshooting techniques for CI failures. + +You now have a complete CI/CD pipeline that automatically tests your MCP server on every code change. + +## Summary + +In this Learning Path, you learned how to: + +- Set up testcontainers for Docker-based integration testing. +- Write pytest tests that communicate with MCP servers over stdio transport. +- Parse MCP JSON-RPC responses and validate tool outputs. +- Configure GitHub Actions with Arm64 runners for automated testing. + +These techniques apply to any MCP server implementation, not just the Arm MCP Server. Use this foundation to build comprehensive test suites that ensure your MCP tools work correctly across updates and deployments. diff --git a/content/learning-paths/cross-platform/automate-mcp-with-testcontainers/introduction.md b/content/learning-paths/cross-platform/automate-mcp-with-testcontainers/introduction.md new file mode 100644 index 0000000000..1a46d6c394 --- /dev/null +++ b/content/learning-paths/cross-platform/automate-mcp-with-testcontainers/introduction.md @@ -0,0 +1,52 @@ +--- +title: Introduction to MCP Server Testing +weight: 2 + +### FIXED, DO NOT MODIFY +layout: learningpathall +--- + +## What is MCP? + +The Model Context Protocol (MCP) is an open standard that enables AI assistants to securely connect to external data sources and tools. MCP servers provide AI models with context-aware capabilities, such as code analysis, knowledge base lookups, and system introspection. + +The Arm MCP Server provides AI assistants with tools and knowledge specifically for Arm architecture development, migration, and optimization. It includes capabilities like container architecture checking, code analysis with LLVM-MCA, and a knowledge base with content from Arm Learning Paths and other documentation. + +## Why automate MCP server testing? + +MCP servers expose multiple tools that AI assistants can invoke. As these tools evolve, you need reliable automated tests to: + +- Verify that each tool responds correctly to valid requests. +- Catch regressions when updating server code or dependencies. +- Validate container startup and communication protocols. +- Ensure compatibility across different environments. + +## Understanding Testcontainers + +Testcontainers is a Python library that provides lightweight, throwaway instances of Docker containers for testing. Instead of mocking your MCP server, you can spin up the actual Docker container, run tests against it, and tear it down automatically. + +![Diagram showing Testcontainers workflow: test code creates a Docker container, runs tests against it, and automatically tears it down after completion#center](testcontainers.png "Figure 1. Testcontainers Flow") + +This approach offers several benefits: + +- **Realistic testing**: Tests run against the actual server implementation. +- **Isolation**: Each test run gets a fresh container instance. +- **Reproducibility**: Tests behave consistently across development machines and CI environments. +- **No external dependencies**: Tests don't require a pre-deployed server. + +## What you will build + +In this Learning Path, you will create an integration test suite that: + +1. Starts the Arm MCP server in a Docker container using Testcontainers. +2. Communicates with the server using the MCP stdio transport protocol. +3. Tests multiple MCP tools including container image checking, knowledge base search, and code analysis. +4. Integrates with GitHub Actions for continuous testing. + +## What you've accomplished and what's next + +In this section: +- You learned what MCP servers are and why automated testing matters. +- You discovered how Testcontainers enable realistic integration testing. + +In the next section, you will set up your development environment and install the required dependencies. diff --git a/content/learning-paths/cross-platform/automate-mcp-with-testcontainers/setup-environment.md b/content/learning-paths/cross-platform/automate-mcp-with-testcontainers/setup-environment.md new file mode 100644 index 0000000000..5169a5d47c --- /dev/null +++ b/content/learning-paths/cross-platform/automate-mcp-with-testcontainers/setup-environment.md @@ -0,0 +1,132 @@ +--- +title: Set up your testing environment +weight: 3 + +### FIXED, DO NOT MODIFY +layout: learningpathall +--- + +## Prerequisites + +Before you begin, ensure you have the following installed on your machine: + +- Python 3.11 or later with the ability to create Python virtual environments +- Docker Engine or Docker Desktop +- Git + +If you are on Linux, you need the Python virtual environment package. + +For Debian or Ubuntu, run: + +```bash +sudo apt install -y python3-venv +``` + +You can verify Docker is running by executing: + +```bash +docker info +``` + +The output shows your Docker configuration. If Docker isn't running, start the Docker daemon before proceeding. + +## Clone the Arm MCP repository + +First, clone the Arm MCP server repository which contains the test framework: + +```bash +git clone https://github.com/arm/mcp.git +cd mcp +``` + +## Build the MCP server Docker image + +The integration tests require a locally built Docker image of the MCP server. Build it from the repository root: + +```bash +docker buildx build -f mcp-local/Dockerfile -t arm-mcp . +``` + +This command creates a Docker image tagged as `arm-mcp:latest`. The build process takes several minutes as it generates the vector database for the knowledge base. + +To verify the image was created successfully: + +```bash +docker images arm-mcp +``` + +The output is similar to: + +```output +REPOSITORY TAG IMAGE ID CREATED SIZE +arm-mcp latest a1b2c3d4e5f6 2 minutes ago 1.2GB +``` + +## Create a Python virtual environment + +Create an isolated Python environment for your test dependencies: + +```bash +python3 -m venv venv +source venv/bin/activate +``` + +On Windows, activate the environment using: + +```bash +venv\Scripts\activate +``` + +## Install test dependencies + +The test framework requires Pytest and Testcontainers. Install them using the provided requirements file: + +```bash +pip install -r mcp-local/tests/requirements.txt +``` + +The requirements file contains: + +```text +testcontainers +pytest +``` + +## Verify your setup + +Run a quick verification to ensure everything is configured correctly: + +```bash +python -c "from testcontainers.core.container import DockerContainer; print('Testcontainers ready')" +``` + +The output confirms Testcontainers can interact with Docker: + +```output +Testcontainers ready +``` + +## Understanding the test directory structure + +The test files are located in `mcp-local/tests/`: + +```text +mcp-local/tests/ +β”œβ”€β”€ constants.py # Test data and expected responses +β”œβ”€β”€ requirements.txt # Python dependencies +β”œβ”€β”€ sum_test.s # Sample Arm assembly file for MCA tests +└── test_mcp.py # Main test file +``` + +- **constants.py**: Contains MCP request payloads and expected responses for each tool being tested. +- **test_mcp.py**: The main test file that uses Testcontainers to spin up the MCP server and run assertions. +- **sum_test.s**: A sample Arm assembly file used to test the LLVM-MCA analysis tool. + +## What you've accomplished and what's next + +In this section: +- You cloned the Arm MCP repository and built the server Docker image. +- You set up a Python virtual environment with Pytest and Testcontainers. +- You explored the test directory structure. + +In the next section, you will examine the test code and understand how to write integration tests for MCP servers. diff --git a/content/learning-paths/cross-platform/automate-mcp-with-testcontainers/testcontainers.png b/content/learning-paths/cross-platform/automate-mcp-with-testcontainers/testcontainers.png new file mode 100644 index 0000000000..9e99d24a11 Binary files /dev/null and b/content/learning-paths/cross-platform/automate-mcp-with-testcontainers/testcontainers.png differ diff --git a/content/learning-paths/cross-platform/automate-mcp-with-testcontainers/write-test-cases.md b/content/learning-paths/cross-platform/automate-mcp-with-testcontainers/write-test-cases.md new file mode 100644 index 0000000000..99efb9a28c --- /dev/null +++ b/content/learning-paths/cross-platform/automate-mcp-with-testcontainers/write-test-cases.md @@ -0,0 +1,283 @@ +--- +title: Write integration tests for MCP servers +weight: 4 + +### FIXED, DO NOT MODIFY +layout: learningpathall +--- + +## Understanding MCP communication + +MCP servers communicate using JSON-RPC 2.0 over standard input/output (stdio transport). Each request is a JSON object followed by a newline, and the server responds with a JSON object on stdout. + +The communication follows this pattern: + +1. Client sends an `initialize` request with protocol version and capabilities. +2. Server responds with its capabilities and server info. +3. Client sends an `initialized` notification. +4. Client can now invoke tools using `tools/call` method. + +## Create the constants file + +Start by defining the test constants in `constants.py`. This file contains the MCP request payloads and expected responses: + +```python +MCP_DOCKER_IMAGE = "arm-mcp:latest" + +INIT_REQUEST = { + "jsonrpc": "2.0", + "id": 1, + "method": "initialize", + "params": { + "protocolVersion": "2024-11-05", + "capabilities": {}, + "clientInfo": {"name": "pytest", "version": "0.1"}, + }, +} + +CHECK_IMAGE_REQUEST = { + "jsonrpc": "2.0", + "id": 2, + "method": "tools/call", + "params": { + "name": "check_image", + "arguments": { + "image": "ubuntu:24.04", + "invocation_reason": ( + "Checking ARM architecture compatibility for ubuntu:24.04 " + "container image as requested by the user" + ), + }, + }, +} + +EXPECTED_CHECK_IMAGE_RESPONSE = { + "status": "success", + "message": "Image ubuntu:24.04 supports all required architectures", + "architectures": [ + "amd64", "unknown", "arm", "unknown", "arm64", "unknown", + "ppc64le", "unknown", "riscv64", "unknown", "s390x", "unknown", + ], +} +``` + +Add more test requests for other MCP tools: + +```python +CHECK_NGINX_REQUEST = { + "jsonrpc": "2.0", + "id": 4, + "method": "tools/call", + "params": { + "name": "knowledge_base_search", + "arguments": { + "query": "nginx performance tweaks", + }, + }, +} + +EXPECTED_CHECK_NGINX_RESPONSE = [ + "https://learn.arm.com/learning-paths/servers-and-cloud-computing/nginx_tune/tune_static_file_server", + "https://learn.arm.com/learning-paths/servers-and-cloud-computing/nginx_tune/test_optimizations", +] +``` + +## Create helper functions for MCP communication + +The MCP server runs inside a Docker container that communicates over an attached socket. Create helper functions to encode and decode MCP messages: + +```python +import json +import time + +def _encode_mcp_message(payload: dict) -> bytes: + """Encode an MCP message for stdio transport.""" + return (json.dumps(payload) + "\n").encode("utf-8") + + +def _read_docker_frame(sock, timeout: float) -> bytes: + """Read a Docker multiplexed frame from the socket.""" + deadline = time.time() + timeout + header = b"" + while len(header) < 8: + if time.time() > deadline: + raise TimeoutError("Timed out waiting for docker frame header.") + chunk = sock.recv(8 - len(header)) + if not chunk: + time.sleep(0.01) + continue + header += chunk + + # Docker frame format: 8-byte header + # byte 0: stream type (0x01 = stdout, 0x02 = stderr) + # bytes 1-3: Reserved (\x00\x00\x00) + # bytes 4-7: Payload size (big-endian uint32) + if header[1:4] != b"\x00\x00\x00": + return header # Raw/unframed output + + size = int.from_bytes(header[4:8], "big") + payload = b"" + while len(payload) < size: + if time.time() > deadline: + raise TimeoutError("Timed out waiting for docker frame payload.") + chunk = sock.recv(size - len(payload)) + if not chunk: + time.sleep(0.01) + continue + payload += chunk + return payload + + +def _read_mcp_message(sock, timeout: float = 10.0) -> dict: + """Read and parse an MCP JSON-RPC message.""" + deadline = time.time() + timeout + buffer = b"" + while True: + if time.time() > deadline: + raise TimeoutError("Timed out waiting for MCP response line.") + frame = _read_docker_frame(sock, timeout) + buffer += frame + while b"\n" in buffer: + line, buffer = buffer.split(b"\n", 1) + if not line: + continue + try: + return json.loads(line.decode("utf-8")) + except json.JSONDecodeError: + idx = line.find(b"{") + if idx != -1: + try: + return json.loads(line[idx:].decode("utf-8")) + except json.JSONDecodeError: + continue +``` + +## Write the main test function + +Create the main test function in `test_mcp.py` that uses testcontainers to manage the MCP server lifecycle: + +```python +import os +from pathlib import Path +import pytest +from testcontainers.core.container import DockerContainer +from testcontainers.core.waiting_utils import wait_for_logs +import constants + +def test_mcp_stdio_transport_responds(): + image = os.getenv("MCP_IMAGE", constants.MCP_DOCKER_IMAGE) + repo_root = Path(__file__).resolve().parents[1] + + with ( + DockerContainer(image) + .with_volume_mapping(str(repo_root), "/workspace") + .with_kwargs(stdin_open=True, tty=False) + ) as container: + # Wait for MCP server to start + wait_for_logs(container, "Starting MCP server", timeout=60) + + # Attach to container stdin/stdout + socket_wrapper = container.get_wrapped_container().attach_socket( + params={"stdin": 1, "stdout": 1, "stderr": 1, "stream": 1} + ) + raw_socket = socket_wrapper._sock + raw_socket.settimeout(10) + + # Initialize MCP session + raw_socket.sendall(_encode_mcp_message(constants.INIT_REQUEST)) + response = _read_mcp_message(raw_socket, timeout=20) + + # Verify initialization + assert response.get("id") == 1 + assert "result" in response + assert "serverInfo" in response["result"] + + # Send initialized notification + raw_socket.sendall( + _encode_mcp_message({ + "jsonrpc": "2.0", + "method": "initialized", + "params": {} + }) + ) +``` + +## Add tool-specific tests + +Extend the test function to verify individual MCP tools: + +```python + def _read_response(expected_id: int, timeout: float = 10.0) -> dict: + """Helper to read a specific response by ID.""" + deadline = time.time() + timeout + while time.time() < deadline: + message = _read_mcp_message(raw_socket, timeout=timeout) + if message.get("id") == expected_id: + return message + raise TimeoutError(f"Timed out waiting for response id={expected_id}.") + + # Test check_image tool + raw_socket.sendall(_encode_mcp_message(constants.CHECK_IMAGE_REQUEST)) + check_image_response = _read_response(2, timeout=60) + assert check_image_response.get("result")["structuredContent"] == \ + constants.EXPECTED_CHECK_IMAGE_RESPONSE + + # Test knowledge_base_search tool + raw_socket.sendall(_encode_mcp_message(constants.CHECK_NGINX_REQUEST)) + check_nginx_response = _read_response(4, timeout=60) + urls = json.dumps(check_nginx_response["result"]["structuredContent"]) + assert any( + expected in urls + for expected in constants.EXPECTED_CHECK_NGINX_RESPONSE + ) +``` + +## Run the tests + +Execute the test suite using pytest: + +```bash +python -m pytest -v mcp-local/tests/test_mcp.py +``` + +The output shows each test assertion: + +```output +============================= test session starts ============================== +platform linux -- Python 3.11.0, pytest-8.0.0 +collected 1 item + +mcp-local/tests/test_mcp.py::test_mcp_stdio_transport_responds PASSED [100%] + +============================== 1 passed in 45.32s ============================== +``` + +For more verbose output that shows the test progress: + +```bash +python -m pytest -s mcp-local/tests/test_mcp.py +``` + +The `-s` flag displays print statements, showing each tool test as it completes. + +## How Testcontainers handle container lifecycle + +The `with DockerContainer(image) as container` pattern: + +1. Pulls the image if not present locally. +2. Creates and starts a new container. +3. Waits for the "Starting MCP server" log message. +4. Yields the container for your test code. +5. Automatically stops and removes the container when the test completes. + +This ensures every test run starts with a clean environment. + +## What you've accomplished and what's next + +In this section: +- You learned how MCP servers communicate using JSON-RPC over stdio. +- You created helper functions to handle Docker socket communication. +- You wrote integration tests that verify MCP tool responses. +- You ran the test suite locally using pytest. + +In the next section, you will configure GitHub Actions to run these tests automatically in your CI/CD pipeline. diff --git a/content/learning-paths/cross-platform/multiplying-matrices-with-sme2/_index.md b/content/learning-paths/cross-platform/multiplying-matrices-with-sme2/_index.md index e637f13a5b..23d8aa9e47 100644 --- a/content/learning-paths/cross-platform/multiplying-matrices-with-sme2/_index.md +++ b/content/learning-paths/cross-platform/multiplying-matrices-with-sme2/_index.md @@ -31,10 +31,12 @@ subjects: Performance and Architecture armips: - Neoverse - Cortex-A + - Arm C1 tools_software_languages: - C - Clang - LLVM + - SME2 operatingsystems: - Linux diff --git a/content/learning-paths/cross-platform/simd-loops/_index.md b/content/learning-paths/cross-platform/simd-loops/_index.md index 47f8c75b0f..3d3caff91a 100644 --- a/content/learning-paths/cross-platform/simd-loops/_index.md +++ b/content/learning-paths/cross-platform/simd-loops/_index.md @@ -26,6 +26,7 @@ skilllevels: Advanced subjects: Performance and Architecture armips: - Neoverse + - Cortex-A operatingsystems: - Linux - macOS @@ -34,7 +35,7 @@ tools_software_languages: - CPP - GCC - Clang - + - SME2 shared_path: true shared_between: - servers-and-cloud-computing diff --git a/content/learning-paths/cross-platform/sme-executorch-profiling/_index.md b/content/learning-paths/cross-platform/sme-executorch-profiling/_index.md index 17c7391c3c..22028e9791 100644 --- a/content/learning-paths/cross-platform/sme-executorch-profiling/_index.md +++ b/content/learning-paths/cross-platform/sme-executorch-profiling/_index.md @@ -24,10 +24,12 @@ skilllevels: Advanced subjects: ML armips: - Cortex-A + - Arm C1 tools_software_languages: - ExecuTorch - Python - CMake + - SME2 operatingsystems: - macOS - Android diff --git a/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/_index.md b/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/_index.md index 37424a5e64..40957ee28b 100644 --- a/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/_index.md +++ b/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/_index.md @@ -1,37 +1,45 @@ --- -title: Deploy Edge AI models using Edge Impulse and AWS IoT Greengrass +title: Deploy ML models to Arm edge devices using Edge Impulse and AWS IoT Greengrass draft: true cascade: draft: true -minutes_to_complete: 120 +description: Learn how to deploy Edge Impulse ML models to Arm-based Linux edge devices using AWS IoT Greengrass custom components. -who_is_this_for: This learning path is for Edge AI and embedded engineers who need to deploy crafted ML for the Edge to thousands of edge devices. +minutes_to_complete: 180 + +who_is_this_for: This Learning Path is for embedded and IoT engineers who want to deploy Edge Impulse ML models to Arm-based edge devices at scale using AWS IoT Greengrass. learning_objectives: - - Basic understanding of Edge Impulses Edge ML Solution - - Basic hardware setup for Edge AI ML development with Edge Impulse - - Install AWS IoT Greengrass onto the edge device - - Configure the edge device with the custom integration between Edge Impulse and AWS IoT Greengrass + - Set up an Arm-based edge device for ML inference with Edge Impulse + - Install and configure AWS IoT Greengrass on the edge device + - Deploy an Edge Impulse ML model as a Greengrass custom component + - Verify model inference results through AWS IoT Core prerequisites: - - An [Edge Impulse Studio](https://studio.edgeimpulse.com/signup) account (workshop will walk through this). - - An AWS Account (if not being hosted by AWS Workshop Studio) + - An [Edge Impulse Studio](https://studio.edgeimpulse.com/signup) account + - An [AWS account](https://aws.amazon.com/) with administrator access + - A supported Arm-based edge device (Raspberry Pi 5, Nvidia Jetson, Qualcomm Dragonwing QC6490) or an AWS EC2 Arm instance + - An SSH client and familiarity with the Linux command line + - Basic understanding of ML concepts author: Doug Anson ### Tags skilllevels: Introductory cloud_service_providers: - - AWS + - AWS subjects: ML armips: - Cortex-M + - Cortex-A + - Neoverse tools_software_languages: - Edge Impulse - - Edge AI + - AWS IoT Greengrass + - GStreamer operatingsystems: - Linux @@ -42,9 +50,21 @@ further_reading: - resource: title: Edge Impulse for beginners link: https://docs.edgeimpulse.com/docs/readme/for-beginners - type: doc + type: documentation + - resource: + title: AWS IoT Greengrass developer guide + link: https://docs.aws.amazon.com/greengrass/v2/developerguide/what-is-iot-greengrass.html + type: documentation + - resource: + title: Edge Impulse AWS Greengrass integration + link: https://docs.edgeimpulse.com/docs/integrations/aws-greengrass + type: documentation + - resource: + title: Edge Impulse Greengrass components repository + link: https://github.com/edgeimpulse/aws-greengrass-components + type: website -weight: 1 # _index.md always has weight of 1 to order correctly +weight: 1 # _index.md always has weight of 1 to order correctly layout: "learningpathall" # All files under learning paths have this same wrapper learning_path_main_page: "yes" # This should be surfaced when looking for related content. Only set for _index.md of learning path content. --- diff --git a/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/cleanup.md b/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/cleanup.md index a06354b69f..0d935b2a47 100644 --- a/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/cleanup.md +++ b/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/cleanup.md @@ -1,13 +1,35 @@ --- -title: 9. AWS Account Cleanup (Optional) +title: Clean up AWS resources weight: 11 ### FIXED, DO NOT MODIFY layout: learningpathall --- -## Cleanup +## Clean up AWS resources (optional) -**AWS workshop attendees**: The temporary AWS account given to you will automatically be deleted. No other action is necessary at this time. +After completing this Learning Path, clean up the AWS resources you created to avoid ongoing costs. The steps below are optional β€” skip any that you want to keep for further experimentation. -**Personal AWS Accounts**: To minimize costs of your AWS resources, you can go to the AWS IoTCore Greengrass deployments page and revise your deployment. In the revision, remove the Edge Impulse custom component from the deployment and redeploy. This will shutdown the "runner" service on your edge device and will no longer send messages into IoTCore when inference results are present. Additionally, if using the EC2 edge device in the workshop, you will want to navigate to the EC2 dashboard, select your EC2 instance you created, and then set the instance state to "terminated" via the "Instance state" button/dropdown. You can also cancel your Greengrass deployments and delete both your Greengrass core device as well as your IoT Thing for your core device (all accomplished via the IoTCore dashboard). +### Remove the Greengrass deployment + +Navigate to **AWS IoT Core** > **Greengrass** > **Deployments**. Select your deployment and revise it to remove the Edge Impulse custom component. Redeploy the updated configuration. This shuts down the Runner service on your edge device and stops MQTT messages from being published to IoT Core. + +### Delete the Greengrass core device + +In **AWS IoT Core** > **Greengrass** > **Core devices**, select the core device you created and delete it. Also navigate to **AWS IoT Core** > **All devices** > **Things** and delete the IoT thing associated with your core device. + +### Delete the S3 bucket + +Navigate to **S3** in the AWS Console. Select the bucket you created for the component artifacts, empty it, and delete it. + +### Delete the Secrets Manager secret + +Navigate to **Secrets Manager** in the AWS Console. Select the **EI_API_KEY** secret and delete it. By default, Secrets Manager schedules deletion after a waiting period. + +### Terminate the EC2 instance + +If you used an EC2 instance as your edge device, navigate to the **EC2** dashboard. Select your instance, then choose **Instance state** > **Terminate instance**. + +## Congratulations + +You've completed this Learning Path. You set up an Arm-based edge device, built and deployed an Edge Impulse ML model through AWS IoT Greengrass, verified live inference, and used MQTT commands to control the Runner service remotely. diff --git a/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/customcomponentdeployment.md b/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/customcomponentdeployment.md index f28ccb6345..2cd7745e59 100644 --- a/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/customcomponentdeployment.md +++ b/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/customcomponentdeployment.md @@ -1,73 +1,81 @@ --- -title: 6. Custom Component Deployment +title: Deploy the component to your edge device weight: 8 ### FIXED, DO NOT MODIFY layout: learningpathall --- -## Greengrass Component Deployment +## Overview -In this section, we will create an AWS IoT Greengrass deployment that will download, prepare, install, and run our Edge Impulse "Runner" service on our edge device. When the "Runner" service starts, it will connect back to our Edge Impulse environment via the API key we inserted into AWS Secret Manager and will download and start to run our deployed ML model we created in Edge Impulse Studio! Let's get this started! +In this section, you create a Greengrass deployment that downloads, installs, and runs the Edge Impulse Runner service on your edge device. When the Runner starts, it connects to your Edge Impulse project using the API key stored in AWS Secrets Manager, downloads your trained ML model, and begins running inference. -### 0. (Non-Camera Edge Devices Only): Additional Custom Component +{{% notice Note %}} +If your edge device doesn't have a camera (for example, an EC2 instance), you need to deploy an additional custom component first. Follow the [non-camera component setup](/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/noncameracustomcomponent/) before continuing. You'll select that component alongside the Edge Impulse Runner component during deployment. +{{% /notice %}} -If your edge device does not contain a camera (i.e. EC2 edge device), you will need to deploy an additional custom component. Please follow [these steps](/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/noncameracustomcomponent/) to get the additional component created. You will be selecting this component in addition to the custom component we created for the Edge Impulse "Runner" service. +## Create a Greengrass deployment -### 1. Deploy the custom component to a selected Greengrass edge device or group of edge devices. +Open the AWS Console and navigate to **AWS IoT Core** > **Greengrass** > **Deployments**. You can either create a new deployment or modify an existing one. -Almost done! We can now go back to the AWS Console -> IoT Core -> Greengrass -> Deployments page and select a deployment (or create a new one!) to deploy our component down to as selected edge device or group of gateways as needed: +You have two deployment target options. To deploy to a group of devices, select a thing group: -Deploy to a group of devices: +![Greengrass deployment page showing the option to deploy to a group of devices with a thing group selected#center](./images/gg_create_deployment.png "Deploy to a group of devices") -![GGDeploy](./images/GG_Create_Deployment.png) +To deploy to a specific device (for example, your EC2 edge device), select a single core device: -Deploy to a specific device (i.e. my EC2 Edge Device): +![Greengrass deployment page showing the option to deploy to a single core device#center](./images/gg_create_deployment_2.png "Deploy to a single device") -![GGDeploy](./images/GG_Create_Deployment_2.png) +After choosing your target, select **Next**. On the components page, select your **EdgeImpulseLinuxRunnerServiceComponent** custom component: -In either case above we now press "Next" and select our newly created custom component: +![Component selection page with the EdgeImpulseLinuxRunnerServiceComponent checkbox selected#center](./images/gg_create_deployment_3.png "Select the custom component") -![GGDeploy](./images/GG_Create_Deployment_3.png) +{{% notice Note %}} +If your edge device doesn't have a camera, also select the **EdgeImpulseRunnerRuntimeInstallerComponent** that you created in the non-camera component setup step: ->**_NOTE:_** ->If you are using an edge device which does not have a camera, you will also need to select the "EdgeImpulseRunnerRuntimeInstallerComponent" custom component that you created above ("Non-Camera Edge Device Custom Component"): ->![GGDeploy](./images/GG_Create_Deployment_3a.png) +![Component selection page with both the Runner and RuntimeInstaller components selected#center](./images/gg_create_deployment_3a.png "Select both components for non-camera devices") +{{% /notice %}} -Press "Next" again, then select our custom component and press "Configure Component" to configure the "Runner" component: +Select **Next** again. Select the **EdgeImpulseLinuxRunnerServiceComponent** and select **Configure component** to customize it for your device: -![GGDeploy](./images/GG_Create_Deployment_4.png) +![Component configuration page with the EdgeImpulseLinuxRunnerServiceComponent selected and the Configure component button visible#center](./images/gg_create_deployment_4.png "Configure the component") ->**_NOTE:_** ->If you also have the Non-Camera component, it does NOT need to be configured... only the "EdgeImpulseLinuxRunnerServiceComponent" should be configured +{{% notice Note %}} +If you also have the non-camera component, it doesn't need configuration. Only configure the **EdgeImpulseLinuxRunnerServiceComponent**. +{{% /notice %}} -#### Customizing a specific Deployment +## Apply the device-specific configuration -We now see that our custom component we registered has a default configuration. We can, however, customize it specifically for our specific hardware configuration (i.e. to a specific device or group of similar devices...). +The component has a default configuration from the recipe, but you can override it for this specific deployment. This is where you use the device-specific JSON you saved during hardware setup. -First lets recall the JSON we saved off when we configured our hardware. Lets customize our Greengrass deployment by clearing, copying, and pasting that JSON into the "Configuration to merge" window... then press "Confirm": +Clear the **Configuration to merge** text box, paste your saved JSON, and select **Confirm**: -![GGDeploy](./images/GG_Create_Deployment_5.png) +![Configuration to merge dialog showing the JSON configuration pasted into the text box#center](./images/gg_create_deployment_5.png "Paste the device-specific configuration") -You'll then see the previous page and continue pressing "Next" until you get to the "Deploy" page: +The ability to customize the configuration per deployment is one of the key benefits of Greengrass components. You can deploy the same component to different devices while adjusting settings like `device_name` or `gst_args` for each target's specific hardware. -![GGDeploy](./images/GG_Create_Deployment_6.png) +Continue selecting **Next** through the remaining pages until you reach the review page. Select **Deploy**: -> **_NOTE:_** ->When performing the deployment, its quite common to, when selecting one of our newly created custom components, to then "Customize" that component by selecting it for "Customization" and entering a new JSON structure (same structure as what's found in the component's associated YAML file for the default configuration) that can be adjusted for a specific deployment (i.e. perhaps your want to change the DeviceName for this particular deployment or specify "gst_args" for a specific edge device(s) camera, etc...). This highlights the power and utility of the component and its deployment mechanism in AWS IoT Greengrass. +![Deployment review page showing the final configuration summary with the Deploy button#center](./images/gg_create_deployment_6.png "Review and deploy") +## Monitor the deployment -> **_NOTE:_** -> The component deployment may take awhile depending on network speed/etc... the reason for this is that all of the required prerequisites to run the Edge Impulse "Runner" service have to be downloaded, setup, and installed. -> -> Back on the edge device via SSH, you can "tail" two different files to watch the progress of the installation/setup as well as the component operation (as root): -> -> % sudo su - -> # tail -f /greengrass/v2/logs/EdgeImpulseLinuxRunnerServiceComponent.log -> # tail -f /tmp/ei*log -> -> The first "tail" will log all of the installation activity during the component setup. The second "tail" (wildcarded) will be the log file of the "running" component. You can actually watch the Edge Impulse "Runner" output in that file if you wish. -> -> Both files are critical for debugging any potential issues with the deployment and/or component configuration. +The deployment can take several minutes depending on network speed. The component downloads and installs all prerequisites (Node.js, libvips, the Edge Impulse CLI) before starting the Runner. -Now that our custom component has been deployed, the component will install Edge Impulse's "Runner" runtime that will then, in turn, pull down and invoke our Edge Impulse's current Impulse (i.e. model...). We will next check that our model is running on our edge device! \ No newline at end of file +To monitor progress, SSH into your edge device and tail the component logs: + +```bash +sudo tail -f /greengrass/v2/logs/EdgeImpulseLinuxRunnerServiceComponent.log +``` + +This log shows the installation activity during the component setup phase. After the install completes, the Runner writes its own log file. To watch running inference output: + +```bash +sudo tail -f /tmp/ei*log +``` + +Both log files are essential for debugging deployment or configuration issues. If the deployment fails, check the component log first for installation errors. + +## What you've accomplished + +In this section, you created a Greengrass deployment, applied your device-specific configuration, and deployed the Edge Impulse Runner component to your edge device. The Runner is now downloading your ML model and starting inference. In the next section, you verify that the model is running and view inference results. \ No newline at end of file diff --git a/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/edgeimpulsecustomcomponentinstall.md b/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/edgeimpulsecustomcomponentinstall.md index 798c5c6933..ba9625d5bc 100644 --- a/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/edgeimpulsecustomcomponentinstall.md +++ b/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/edgeimpulsecustomcomponentinstall.md @@ -1,124 +1,145 @@ --- -title: 5. Edge Impulse Custom Component Creation +title: Create the Edge Impulse Greengrass component weight: 7 ### FIXED, DO NOT MODIFY layout: learningpathall --- -## Edge Impulse "Runner" Service Custom Component - -We will utilize a Greengrass "Custom Component" to create and deploy the Edge Impulse runner service (the service that will run our Edge Impulse model on the edge device) including the required additional prerequisites (NodeJS install, libvips install). AWS IoT Greengrass' custom component feature is ideal to create custom components that can be specialized to prepare, run, and shutdown a given custom service. - -Let's get started! - -### 1. Clone the repo to acquire the Edge Impulse Component recipes and artifacts - -Please clone this [repo](https://github.com/edgeimpulse/aws-greengrass-components) to retrieve the Edge Impulse component recipes (yaml files) and the associated artifacts. - -### 2. Upload Edge Impulse Greengrass Component artifacts into AWS S3 - -First, you need to go to the S3 console in AWS via AWS Console -> S3. From there, you will create an S3 bucket. For sake of example, we name this bucket "MyS3Bucket123". - - ![CreateS3Bucket](./images/S3_Create_Bucket.png) - -Next, the following directory structure needs to be created your new bucket: - - ./artifacts/EdgeImpulseServiceComponent/1.0.0 - -Next, navigate to the "1.0.0" directory in your S3 bucket and then press "Upload" to upload the artifacts into the bucket. You need to upload the following files (these will be located in the ./artifacts/EdgeImpulseServiceComponent/1.0.0 from your cloned repo). Please upload all of these files into S3 at the above directory location: - - install.sh - run.sh - launch.sh - stop.sh - -Your S3 Bucket contents should look like this: - -![UploadToS3](./images/S3_Upload_Artifacts.png) - -### 3. Customize the component recipe files - -Next we need to customize our Greengrass component recipe YAML file to reflect the actual location of our artifacts stored in S3. Please replace ALL occurrences of "YOUR\_S3\_ARTIFACT\_BUCKET" with your S3 bucket name (i.e. "MyS3Bucket123"). Please do this to the "EdgeImpulseLinuxRunnerServiceComponent.yaml" file. Save the file. - -Also FYI, we can customize the defaulted configuration of your custom component by editing, within "EdgeImpulseLinuxRunnerServiceComponent.yaml" file, the default configuration JSON. We won't need to do this for this workshop but its an useful option nonetheless. - -The default configuration in "EdgeImpulseLinuxRunnerServiceComponent.yaml" contains the following JSON configuration settings for the component: - - EdgeImpulseLinuxRunnerServiceComponent.yaml: - { - "node_version": "20.12.1", - "vips_version": "8.12.1", - "device_name": "MyEdgeImpulseDevice", - "launch": "runner", - "sleep_time_sec": 10, - "lock_filename": "/tmp/ei_lockfile_runner", - "gst_args": "__none__", - "eiparams": "--greengrass", - "iotcore_backoff": "5", - "iotcore_qos": "1", - "ei_bindir": "/usr/local/bin", - "ei_sm_secret_id": "EI_API_KEY", - "ei_sm_secret_name": "ei_api_key", - "ei_ggc_user_groups": "video audio input users", - "install_kvssink": "no", - "publish_inference_base64_image": "no", - "enable_cache_to_file": "no", - "ei_poll_sleeptime_ms": 2500, - "ei_local_model_file": "__none__", - "ei_shutdown_behavior": "__none__", - "cache_file_directory": "__none__", - "enable_threshold_limit": "no", - "metrics_sleeptime_ms": 30000, - "default_threshold": 50.0, - "threshold_criteria": "ge", - "enable_cache_to_s3": "no", - "s3_bucket": "__none__", - } - -#### Attribute Description - -The attributes in each of the above default configurations is outlined below: - -* **node\_version**: Version of NodeJS to be installed by the component -* **vips\_version**: Version of the libvips library to be compiled/installed by the component -* **device\_name**: Template for the name of the device in EdgeImpulse... a unique suffix will be added to the name to prevent collisions when deploying to groups of devices -* **launch**: service launch type (typically just leave this as-is) -* **sleep\_time\_sec**: wait loop sleep time (component lifecycle stuff... leave as-is) -* **lock\_filename**: name of lock file for this component (leave as-is) -* **gst\_args**: optional GStreamer args, spaces replaced with ":", for custom video invocations -* **eiparams**: additional parameters for launching the Edge Impulse service (leave as-is) -* **iotcore\_backoff**: number of inferences to "skip" before publication to AWS IoTCore... this is used to control publication frequency (AWS $$...) -* **iotcore\_qos**: MQTT QoS (typically leave as-is) -* **ei\_bindir**: Typical location of where the Edge Impulse services are installed (leave as-is) -* **ei\_ggc\_user\_groups**: A list of additional groups the Greengrass service user account will need to be a member of to allow the Edge Impulse service to invoke and operate correctly (typically leave as-is). For JetPack v6.x and above, please add "render" as an additional group. -* **ei\_sm\_secret\_id**: ID of the Edge Impulse API Key within AWS Secret Manager -* **ei\_sm\_secret\_name**: Name of the Edge Impulse API Key within AWS Secret Manager -* **install\_kvssink**: Option (default: "no", on: "yes") to build and make ready the kvssink gstreamer plugin -* **publish\_inference\_base64\_image**: Option (default: "no", on: "yes") to include a base64 encoded image that the inference result was based on -* **enable\_cache\_to\_file**: Option (default: "no", on: "yes") to enable both inference and associated image to get written to a specified local directory as a pair: .img and .json for each inference identified with a -* **cache\_file\_directory**: Option (default: "__none__") to specify the local directory when enable_cache_to_file is set to "yes" -* **ei\_poll\_sleeptime\_ms**: time (in ms) for the long polling message processor (typically leave as-is) -* **ei\_local\_model\_file**: option to utilize a previously installed local model file -* **ei\_shutdown\_behavior**: option to alter the shutdown behavior of the linux runner process. (can be set to "wait\_for\_restart" to cause the runner to pause after running the model and wait for the "restart" command to be issued (see "Commands" below for more details on the "restart" command)) -* **enable\_threshold\_limit**: option to enable/disable the threshold confidence filter (must be "yes" or "no". Default is "no") -* **metrics\_sleeptime\_ms**: option to publish the model metrics statistics (time specified in ms). -* **default\_threshold**: option to specify threshold confidence filter "limit" (a value between 0 < x <= 1.0). Default setting is 0.7 -* **threshold\_criteria**: option to specify the threshold confidence filter criteria (must be one of: "gt", "ge", "eq", "le", or "lt") -* **enable\_cache\_to\_s3**: option to enable caching the inference image/result to an AWS S3 bucket -* **s3\_bucket**: name of the optional S3 bucket to cache results into - -### 4. Register the custom component via its recipe file - -From the AWS Console -> IoT Core -> Greengrass -> Components, select "Create component". Then: - - 1. Select the "yaml" option to Enter the recipe - 2. Clear the text box to remove the default "hello world" yaml recipe - 3. Copy/Paste the entire/edited contents of your "EdgeImpulseLinuxRunnerServiceComponent.yaml" file - 4. Press "Create Component" - -![CreateComponent](./images/GG_Create_Component.png) - -If formatting and artifact access checks out OK, you should have a newly created component listed in your Custom Components AWS dashboard! - -Next we will create a Greengrass Deployment to deploy our custom component to our edge devices. \ No newline at end of file +## Overview + +AWS IoT Greengrass uses *custom components* to package and deploy software to edge devices. In this section, you create a custom component that installs and runs the Edge Impulse Runner service on your device. The component handles all prerequisites (Node.js, libvips) and manages the Runner lifecycle β€” install, run, and shutdown. + +The component consists of two parts: +- **Artifacts**: Shell scripts (stored in S3) that install dependencies and launch the Runner. +- **Recipe**: A YAML file that tells Greengrass where to find the artifacts, what configuration to apply, and how to manage the component lifecycle. + +## Clone the component repository + +Clone the Edge Impulse Greengrass components repository to get the recipe and artifact files: + +```bash +git clone https://github.com/edgeimpulse/aws-greengrass-components.git +``` + +This repository contains the YAML recipe file and the shell script artifacts you'll upload to S3. + +## Upload artifacts to S3 + +The Greengrass component downloads its artifacts from an S3 bucket at deployment time. You need to create a bucket and upload the shell scripts. + +Open the AWS Console and navigate to **S3**. Select **Create bucket** and give it a name (for example, `my-ei-greengrass-artifacts`): + +![S3 console showing the Create bucket dialog with a bucket name entered#center](./images/s3_create_bucket.png "Create an S3 bucket") + +Inside your new bucket, create the following directory structure: + +```text +artifacts/EdgeImpulseServiceComponent/1.0.0/ +``` + +Navigate to the `1.0.0` directory in your S3 bucket and select **Upload**. Upload all four files from the cloned repository's `./artifacts/EdgeImpulseServiceComponent/1.0.0/` directory: + +- `install.sh` +- `run.sh` +- `launch.sh` +- `stop.sh` + +After the upload, your S3 bucket should look like this: + +![S3 bucket showing the artifacts directory structure with the four shell scripts uploaded in the 1.0.0 folder#center](./images/s3_upload_artifacts.png "Uploaded artifacts in S3") + +## Customize the component recipe + +The recipe YAML file tells Greengrass where to download the artifacts from S3. You need to update it with your actual bucket name. + +Open `EdgeImpulseLinuxRunnerServiceComponent.yaml` from the cloned repository and replace all occurrences of `YOUR_S3_ARTIFACT_BUCKET` with your S3 bucket name (for example, `my-ei-greengrass-artifacts`). Save the file. + +### Default configuration reference + +The recipe file includes a default configuration JSON block. You don't need to modify these defaults for this Learning Path β€” they're overridden at deployment time by the device-specific JSON you saved during hardware setup. However, understanding each field is useful for troubleshooting and customization. + +```json +{ + "node_version": "20.12.1", + "vips_version": "8.12.1", + "device_name": "MyEdgeImpulseDevice", + "launch": "runner", + "sleep_time_sec": 10, + "lock_filename": "/tmp/ei_lockfile_runner", + "gst_args": "__none__", + "eiparams": "--greengrass", + "iotcore_backoff": "5", + "iotcore_qos": "1", + "ei_bindir": "/usr/local/bin", + "ei_sm_secret_id": "EI_API_KEY", + "ei_sm_secret_name": "ei_api_key", + "ei_ggc_user_groups": "video audio input users", + "install_kvssink": "no", + "publish_inference_base64_image": "no", + "enable_cache_to_file": "no", + "ei_poll_sleeptime_ms": 2500, + "ei_local_model_file": "__none__", + "ei_shutdown_behavior": "__none__", + "cache_file_directory": "__none__", + "enable_threshold_limit": "no", + "metrics_sleeptime_ms": 30000, + "default_threshold": 50.0, + "threshold_criteria": "ge", + "enable_cache_to_s3": "no", + "s3_bucket": "__none__" +} +``` + +### Configuration field reference + +The table below describes each configuration field: + +| Field | Description | +|---|---| +| `node_version` | Version of Node.js to install on the device. | +| `vips_version` | Version of the libvips library to compile and install. | +| `device_name` | Base name for the device in Edge Impulse. A unique suffix is appended automatically to prevent collisions when deploying to multiple devices. | +| `launch` | Service launch type. Leave as `runner`. | +| `sleep_time_sec` | Wait loop sleep time for the component lifecycle. Leave as default. | +| `lock_filename` | Lock file path for this component. Leave as default. | +| `gst_args` | GStreamer pipeline arguments with spaces replaced by colons. Set per-device during deployment (for example, `v4l2src:device=/dev/video0:!:video/x-raw,width=640,height=480:!:videoconvert:!:jpegenc`). Use `__none__` to disable. | +| `eiparams` | Additional parameters for the Edge Impulse Runner. The `--greengrass` flag is required. | +| `iotcore_backoff` | Number of inference results to skip between MQTT publications. Controls publication frequency and cost. Set to `-1` to publish every result, or a positive number to throttle. | +| `iotcore_qos` | MQTT Quality of Service level. Leave as `1`. | +| `ei_bindir` | Installation directory for the Edge Impulse CLI tools. Leave as default. | +| `ei_sm_secret_id` | Secret ID in AWS Secrets Manager that holds the Edge Impulse API key. Must match the secret name you created (`EI_API_KEY`). | +| `ei_sm_secret_name` | Key name within the Secrets Manager secret. Must match the key you created (`ei_api_key`). | +| `ei_ggc_user_groups` | Linux groups the Greengrass service user (`ggc_user`) is added to. For Jetpack 6.x and later, add `render` to this list for GPU access. | +| `install_kvssink` | Set to `yes` to build and install the KVS sink GStreamer plugin. Default: `no`. | +| `publish_inference_base64_image` | Set to `yes` to include a base64-encoded image with each inference result published to MQTT. Default: `no`. | +| `enable_cache_to_file` | Set to `yes` to write inference results and associated images to a local directory as paired files (`.json` and `.img`). Default: `no`. | +| `cache_file_directory` | Local directory path for cached files when `enable_cache_to_file` is `yes`. Default: `__none__`. | +| `ei_poll_sleeptime_ms` | Polling interval in milliseconds for the long-polling message processor. Leave as default. | +| `ei_local_model_file` | Path to a previously downloaded local model file (`.eim`). Set to `__none__` to download the model from Edge Impulse at runtime. | +| `ei_shutdown_behavior` | Controls Runner behavior after the model finishes. Set to `wait_on_restart` to pause after a video file ends and wait for a restart command. Default: `__none__`. | +| `enable_threshold_limit` | Set to `yes` to enable the confidence threshold filter. Default: `no`. | +| `metrics_sleeptime_ms` | Interval in milliseconds between model metrics publications. Default: `30000`. | +| `default_threshold` | Confidence threshold value between 0 and 100. Inference results below this threshold are filtered out when `enable_threshold_limit` is `yes`. Default: `50.0`. | +| `threshold_criteria` | Comparison operator for the threshold filter. Must be one of: `gt`, `ge`, `eq`, `le`, or `lt`. Default: `ge`. | +| `enable_cache_to_s3` | Set to `yes` to cache inference images and results to an S3 bucket. Default: `no`. | +| `s3_bucket` | S3 bucket name for cached results when `enable_cache_to_s3` is `yes`. Default: `__none__`. | + +## Register the component in Greengrass + +With the artifacts in S3 and the recipe updated, register the component in the AWS Console. + +Navigate to **AWS IoT Core** > **Greengrass** > **Components** and select **Create component**. Then: + +1. Select **Enter recipe as YAML** as the input method. +2. Clear the default "hello world" YAML from the text box. +3. Copy and paste the entire contents of your edited `EdgeImpulseLinuxRunnerServiceComponent.yaml` file. +4. Select **Create component**. + +![Greengrass Components console showing the Create component form with the YAML recipe pasted into the editor#center](./images/gg_create_component.png "Register the custom component") + +If the recipe format is valid and Greengrass can access the S3 artifacts, the component appears in your custom components list. + +## What you've accomplished + +In this section, you cloned the Edge Impulse component repository, uploaded artifacts to S3, customized the recipe with your bucket name, and registered the component in Greengrass. In the next section, you create a Greengrass deployment to push this component to your edge device. \ No newline at end of file diff --git a/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/edgeimpulseprojectbuild.md b/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/edgeimpulseprojectbuild.md index 917653d9d6..dd2d64aad6 100644 --- a/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/edgeimpulseprojectbuild.md +++ b/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/edgeimpulseprojectbuild.md @@ -1,113 +1,111 @@ --- -title: 2. Edge Impulse Project Setup +title: Set up your Edge Impulse project weight: 4 ### FIXED, DO NOT MODIFY layout: learningpathall --- -## Creating our Edge Impulse Environment +## Create an Edge Impulse account -The next step is to create our Edge Impulse environment. Edge Impulse provides a simple solution to creating and building a ML model for specific edge devices focused on specific tasks. Lets get started. +Edge Impulse is an ML platform that lets you build, train, optimize, and deploy models for edge devices. In this section, you create an account, clone a pre-built project, build a deployment for your Arm device, and generate an API key. -### 1. Create Edge Impulse Account +Navigate to [Edge Impulse Studio](https://studio.edgeimpulse.com) and select **Sign Up** in the upper-right corner: -Lets create our account in Edge Impulse. Navigate to https://studio.edgeimpulse.com and select "Sign Up" down in the right hand corner: +![Edge Impulse Studio login page with the Sign Up button in the upper-right corner#center](./images/ei_signup_1.png "Edge Impulse sign-up page") -![Sign Up](./images/EI_SignUp_1.png) +Fill in the requested information and select **Sign Up**: -Next, fill in the requested information and press "Sign Up": +![Edge Impulse sign-up form with fields for name, email, and password#center](./images/ei_signup_2.png "Complete the sign-up form") -![Complete Information](./images/EI_SignUp_2.png) +After a successful sign-up, a confirmation message appears. Select **Click here to build your first ML model**: -If successful, you will be promoted as follows. Press "Click here to build your first ML model": +![Confirmation message after successful Edge Impulse account creation#center](./images/ei_signup_3.png "Successful sign-up confirmation") -![Successful Sign Up](./images/EI_SignUp_3.png) +A wizard appears to help you create a default project: -You will be presented with a wizard to create a new default project: +![Edge Impulse new project wizard showing initial setup options#center](./images/ei_signup_4.png "New project wizard") -![Intro Wizard To create ML Model](./images/EI_SIgnUp_4.png) +You can dismiss the wizard by selecting the **-** button in the upper-right corner. This reveals your new default project: -You can dismiss the wizard by pressing the "-" in the upper right hand corner... this will reveal your current new default project: +![Edge Impulse dashboard showing a newly created default project#center](./images/ei_signup_5.png "New default project") -![New Project](./images/EI_SignUp_5.png) +Now that your account is ready, clone an existing project that already has a trained model. You'll use this model throughout the Learning Path. -Next, we will clone an existing project that has a model that has already been created for you and which we will use for this workshop. On to the next step! +## Clone the project into your account -### 2. Clone Project Into Your Account - -Next, we are going to clone an existing project into our own space. Navigate to this public project: +Navigate to this public project: https://studio.edgeimpulse.com/studio/524106 -![Public Project](./images/EI_Clone_1.png) +![Edge Impulse public project page for the Cat and Dog Detector model#center](./images/ei_clone_1.png "Public project page") -Press the "Clone this project" button in the upper right. You will be presented with a dialog that will initiate the clone: +Select the **Clone this project** button in the upper-right corner. A dialog appears to confirm the clone: -![Clone Project](./images/EI_Clone_2.png) +![Clone project dialog with default settings and a Clone Project button#center](./images/ei_clone_2.png "Clone project dialog") -Leave everything defaulted and press "Clone Project" in the lower right. The cloning process will commence: +Leave the default settings and select **Clone Project** in the lower-right corner. The cloning process starts: -![Start Project Clone](./images/EI_Clone_3.png) +![Progress indicator showing the project clone in progress#center](./images/ei_clone_3.png "Cloning in progress") -The cloning process will take about 12 minutes to complete. When it is complete: +The cloning process takes about 12 minutes to complete. When it finishes, a completion message appears: -![Completed Clone](./images/EI_Clone_4.png) +![Completion message indicating the project clone finished successfully#center](./images/ei_clone_4.webp "Clone complete") -Next, click "Dashboard" to view your project... it should look something like this: +Select **Dashboard** to view the cloned project. It should look similar to the following: -![My Cloned Project](./images/EI_Clone_5.png) +![Edge Impulse dashboard showing the cloned Cat and Dog Detector project with model details#center](./images/ei_clone_5.webp "Cloned project dashboard") -OK! We now have the project we will use for the workshop... lets continue by exploring the project a bit and creating a deployment for our own edge device. Onward! +You now have the project you'll use for this Learning Path. -### 3. Build your project's deployment +## Build your project deployment -Let's have a look at some of the features in Edge Impulse studio. From a high level, Edge Impulse studio provides a solution to build, train, optimize, and deploy ML models for any edge device: +Edge Impulse Studio provides a workflow to build, train, optimize, and deploy ML models. Take a moment to explore the project dashboard: -![Edge Impulse](./images/EI_Project_1.png) +![Edge Impulse Studio dashboard showing the project overview with data, impulse, and deployment sections#center](./images/ei_project_1.png "Project dashboard overview") -Key in this is the "Impulse". On the left side of the dashboard, our "Impulse" has been created for us and is called "Cat and Dog Detector". Click on "Create Impulse". You will see that there are 3 main parts of a "Impulse": The pre-processing block, the model block, and the post-processing block: +Central to Edge Impulse is the concept of an *Impulse*, which is a pipeline that defines how sensor data is processed, what model runs on it, and how results are interpreted. Your cloned project already has an Impulse called "Cat and Dog Detector". Select **Create Impulse** on the left sidebar to see the three main parts: the pre-processing block, the model block, and the post-processing block: -![Edge Impulse](./images/EI_Project_2.png) +![Create Impulse view showing the three pipeline blocks: pre-processing, model, and post-processing#center](./images/ei_project_2.webp "Impulse pipeline structure") -Clicking on "Object Detection" on the left, you will see some detail on the model that has been utilized in our Impulse: +Select **Object Detection** on the left sidebar to see details about the model used in the Impulse: -![Edge Impulse](./images/EI_Project_3.png) +![Object Detection page showing the model architecture and training results#center](./images/ei_project_3.webp "Object Detection model details") -In our project, the "Impulse" is fully created, trained, and optimized so we won't have to walk through those steps. Edge Impulse has a ton of [examples and documentation](https:://docs.edgeimpulse.com) to walk you through your first "Impulse" creation: +The Impulse in this project is already created, trained, and optimized, so you don't need to walk through those steps. Edge Impulse provides extensive [examples and documentation](https://docs.edgeimpulse.com) to guide you through creating your own Impulse from scratch: -![Edge Impulse](./images/EI_Project_4.png) +![Edge Impulse documentation page showing available guides and tutorials#center](./images/ei_project_4.webp "Edge Impulse documentation") -What we want to do now is to deploy our model to a specific edge device type. Depending on the specific hardware you are using in this workshop, you can choose from the following deployment edge device choices: +Now deploy the model to your specific edge device type. Depending on the hardware you selected earlier, choose the matching deployment target: -![Edge Impulse](./images/EI_Project_5.png) +![Deployment page showing available target device options including Linux AARCH64 and other platforms#center](./images/ei_project_5.webp "Deployment target options") -Please select the appropriate choice and press "Build" (Example, for Raspberry Pi, choose "Linux(AARCH64)" to run the model on the CPU of the RPi: +Select the appropriate target for your device and select **Build**. For example, if you're using a Raspberry Pi 5 or an EC2 Graviton instance, choose **Linux (AARCH64)** to run the model on the CPU: -![Edge Impulse](./images/EI_Project_6.png) +![Build dialog with the Linux AARCH64 target selected and the Build button highlighted#center](./images/ei_project_6.webp "Build deployment") - NOTE: For these edge device choices, please select the "int8" option - prior to pressing "Build". +{{% notice Note %}} +For these edge device targets, select the **int8** quantization option before selecting **Build**. The **Linux (AARCH64)** target is suitable for many Linux-class Arm-based 64-bit devices where the CPU runs the model. +{{% /notice %}} - NOTE: The "Linux(AARCH64)" is suitable for many Linux-class ARM-based - 64bit devices where only the CPU will be used to run the model. +With the deployment built, the next step is to create an API key that connects the Greengrass component to your Edge Impulse project. -Now that we have built our deployment, we are ready to move on to the next step - creating an API Key. Lets do this! +## Create your project API key -### 4. Create your project API key +The Edge Impulse Runner on your device uses an API key to authenticate with your project and download the model. Select **Dashboard** on the left sidebar: -Lastly, lets create our API key for our project. We'll use this key to connect our Greengrass component's environment to our Edge Impulse project. Click on the "Dashboard" on the left hand side of our project: +![Edge Impulse project dashboard with the Dashboard link highlighted in the left sidebar#center](./images/ei_key_1.webp "Project dashboard") -![Edge Impulse Dashboard](./images/EI_Key_1.png) +Select **Keys**: -Press "Keys": +![Dashboard view with the Keys tab visible in the project settings area#center](./images/ei_key_2.webp "Keys tab") -![Edge Impulse Dashboard](./images/EI_Key_2.png) +Select **Add new API key** in the upper-right corner. Enter a name for the key, set the role to **admin**, and confirm that **Set as development key** is selected. Then select **Create API key**: -Press "Add new API key" on the upper right side. Provide a name for the key. The role should be "admin" and "Set as development key" should be selected. Press "Create API key": +![API key creation dialog with fields for name, role set to admin, and the development key checkbox selected#center](./images/ei_key_3.webp "Create API key") -![Edge Impulse Dashboard](./images/EI_Key_3.png) +The API key appears on the screen. Copy and save it immediately β€” this is the only time the full key is visible. You'll store this key in AWS Secrets Manager in a later step. -You will then be presented with the API key. Make a copy of this key as this will be the only time you will be able to see the full key for copying. We will place this key into AWS Secret Manager shortly so be sure to save it now!! +## What you've accomplished -OK! We are making good progress! Next up, we are going to install AWS IoT Greengrass into our edge device. Lets go! \ No newline at end of file +In this section, you created an Edge Impulse account, cloned a pre-built Cat and Dog Detector project, built a deployment for your Arm device, and generated an API key. In the next section, you install AWS IoT Greengrass on your edge device. \ No newline at end of file diff --git a/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/greengrassinstallation.md b/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/greengrassinstallation.md index 59c54875cd..d3c37cbff9 100644 --- a/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/greengrassinstallation.md +++ b/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/greengrassinstallation.md @@ -1,106 +1,117 @@ --- -title: 3. AWS IoT Greengrass Installation +title: Install AWS IoT Greengrass weight: 5 ### FIXED, DO NOT MODIFY layout: learningpathall --- -## AWS IoT Greengrass Installation +## Create AWS access credentials -The following sections outline how one installs AWS IoT Greengrass onto our edge device. AWS IoT Greengrass is ideal to use to create deployments of software and settings down to edge devices in a very scalable fashion. +Before installing AWS IoT Greengrass, you need a set of AWS access credentials. The Greengrass installer uses these credentials to register your edge device with AWS IoT Core and configure the required cloud resources. -Log into your edge device via SSH and we'll start the process of installing/configuring Greengrass. +{{% notice Note %}} +If you're using an AWS-hosted event account, credentials might be provided to you automatically. If so, copy them for the next step. They should look like this: -### 1. Create AWS Administrator Credentials +```bash +export AWS_ACCESS_KEY_ID= +export AWS_SECRET_ACCESS_KEY= +``` -Prior to installing AWS IoT Greengrass, we need to create a set of AWS credentials that will be used as part of the installation process. +If you already have credentials, skip ahead to [Install Greengrass Nucleus Classic](#install-greengrass-nucleus-classic). +{{% /notice %}} ->**_NOTE:_** ->These credentials may automatically be provided to you when you initiate the workshop is hosted by AWS Workshop Studio. If so, please copy the credentials as we'll need them in the next step. The credentials should look like this: -> -> export AWS_ACCESS_KEY_ID= -> export AWS_SECRET_ACCESS_KEY= +If you're using a personal AWS account and don't have access credentials yet, follow the steps below to create them. -If you are using your personal AWS account and do not have the credentials created, you will need to create them. If you already have them, please skip the next step and proceed to step 3) below. +### Create access credentials for a personal AWS account -#### 1a. Creating Access Credentials (personal AWS Accounts) +Open the AWS Console and search for **IAM**: -Please navigate to your AWS Dashboard and search for IAM: +![AWS Console search bar with IAM entered as the search term#center](./images/gg_install_iam.png "Search for IAM") -![IAM](./images/GG_Install_iam.png) +Open the IAM Dashboard: -Launch the IAM Dashboard: +![IAM Dashboard showing the main overview with users, roles, and policies sections#center](./images/gg_install_iam_dashboard.png "IAM Dashboard") -![IAM](./images/GG_Install_iam_dashboard.png) +Select **Users** from the left sidebar: -Select "Users" from the left hand side of the dashboard: +![IAM Users list showing available user accounts#center](./images/gg_install_iam_2.png "IAM Users list") -![IAM](./images/GG_Install_iam_2.png) +Select your user, then select the **Security credentials** tab: -Select your user, then select the "Security Credentials" tab: +![User details page with the Security credentials tab selected#center](./images/gg_install_iam_3.png "Security credentials tab") -![IAM](./images/GG_Install_iam_3.png) +Select **Create access key**: -Press "Create access key": +![Security credentials section with the Create access key button#center](./images/gg_install_iam_4.png "Create access key") -![IAM](./images/GG_Install_iam_4.png) +Choose **Other** as the use case and select **Next**: -Choose "Other" and then press "Next": +![Access key use case selection with the Other option highlighted#center](./images/gg_install_iam_5.png "Select use case") -![IAM](./images/GG_Install_iam_5.png) +Enter a description for the access key (for example, "Greengrass installer") and select **Create access key**: -Set a description for the access key and then press "Create access key" +![Access key description field with the Create access key button#center](./images/gg_install_iam_6.png "Create access key") -![IAM](./images/GG_Install_iam_6.png) +This is the only time you can view the full credentials. Copy them and save them to a temporary file in this format: -You will now have the (only...) opportunity to copy and save off your credentials. Its best if you save them to a temp file that you'll read later in this format: +```bash +export AWS_ACCESS_KEY_ID= +export AWS_SECRET_ACCESS_KEY= +``` - export AWS_ACCESS_KEY_ID= - export AWS_SECRET_ACCESS_KEY= +You'll paste these into your SSH session during the Greengrass installation. -### 2. Install AWS IoT Greengrass +## Install Greengrass Nucleus Classic -Greengrass is typically installed from within the AWS Console -> AWS IoT Core -> Greengrass -> Core Devices menu... select/press "Set up one core device". There are multiple ways to install Greengrass - "Nucleus Classic" is the version of Greengrass that is based on Java. "Nucleus Lite" is a native version of Greengrass that is typically part of a Yocto-image based implementation. +AWS IoT Greengrass has two versions: **Nucleus Classic**, which is Java-based, and **Nucleus Lite**, which is a native implementation typically used with Yocto-based images. This Learning Path uses Nucleus Classic because it runs on standard Linux distributions that your edge device is already running. -In this example, we choose the "Linux" device type and we are going to download the installer for Greengrass and invoke it as part of the installation of a "Nucleus Classic" instance: +In the AWS Console, navigate to **AWS IoT Core** > **Greengrass** > **Core devices** and select **Set up one core device**. -![CreateDevice](./images/GG_Install_Device.png) +Select **Linux** as the device type. The console generates download and install commands customized for your account: -Lower down in the menu, you will see the specific instructions that are custom-crafted for you to download and invoke the "Nucleus Classic" installer. The basic sequence of instructions are: +![Greengrass core device setup page showing the Linux device type selected and the Nucleus Classic option#center](./images/gg_install_device.png "Set up core device") - 1) Start with a SSH shell session into your edge device - 2) copy and paste your two AWS credentials into the shell environment -``` - NOTE: Your "two AWS credentials" are the AWS_ACCESS_KEY_ID and AWS_SECRET_ACCESS_KEY from above... -``` - 3) copy and paste/run the installer download curl command into your shell -``` - NOTE: The "installer download command" is in the "Download the installer" section in the image below... -``` - 4) copy and paste/run the installer invocation command -``` - NOTE: The "installer invocation command" is in the "Run the installer" section in the image below... -``` - 5) wait for the installer to complete on your edge device +Scroll down to see the install steps. The console provides commands tailored to your account. Follow these steps in an SSH session on your edge device: + +1. Export your AWS credentials in the terminal: + + ```bash + export AWS_ACCESS_KEY_ID= + export AWS_SECRET_ACCESS_KEY= + ``` + +2. Copy and run the **Download the installer** command from the console. This downloads the Greengrass Nucleus installer to your device. + +3. Copy and run the **Run the installer** command from the console. This installs and starts the Greengrass Nucleus service. + +4. Wait for the installer to finish. A successful installation displays a confirmation message. + +The screenshot below shows where to find these commands in the console: + +![Greengrass setup page showing the Download the installer and Run the installer sections with copy buttons#center](./images/gg_install_device2.png "Installer commands") + +## Add permissions to the Greengrass token exchange role - ![CreateDevice](./images/GG_Install_Device2.png) +When Greengrass runs a component, it uses a Linux service user called `ggc_user` (on Nucleus Classic installations) to start the process. AWS credentials are passed to the component through its environment at launch time, and the component's AWS SDK uses those credentials to connect to AWS services. The permissions available to the component are controlled by an IAM role called `GreengrassV2TokenExchangeRole`. -### 3. Modify the Greengrass TokenExchange Role with additional permissions +By default, this role doesn't include the permissions that the Edge Impulse component needs. You need to add three policies: -When you run a Greengrass component within Greengrass, a service user (typically a linux user called "ggc_user" for "Nucleus Classic" installations) invokes the component, as specified in the lifecycle section of your recipe. Credentials are passed to the invoked process via its environment (NOT by the login environment of the "Greengrassc_user"...) during the invocation spawning process. These credentials are used by by the spawned process (typically via the AWS SDK which is part of the spawned process...) to connect back to AWS and "do stuff". These permissions are controlled by a AWS IAM Role called "GreengrassV2TokenExchangeRole". We need to modify that role and add "Full AWS IoT Core Permission" as well as "AWS Secrets Manager Read/Write" permission. +- **AWSIoTFullAccess** β€” allows the component to publish inference results and receive commands through AWS IoT Core MQTT topics. +- **AmazonS3FullAccess** β€” allows access to S3 buckets where component artifacts are stored. +- **SecretsManagerReadWrite** β€” allows the component to retrieve the Edge Impulse API key from AWS Secrets Manager. -To modify the role, from the AWS Console -> IAM -> Roles search for "GreengrassV2TokenExchangeRole", Then: +To add these permissions, navigate to **IAM** > **Roles** in the AWS Console and search for `GreengrassV2TokenExchangeRole`. Then: - 1. Select "GreengrassV2TokenExchangeRole" in the search results list - 2. Select "Add Permissions" -> "Attach Policies" - 3. Search for "AWSIoTFullAccess", select it, then press "Add Permission" down at the bottom - 4. Repeat the search for "S3FullAccess" and "SecretsManagerReadWrite" +1. Select **GreengrassV2TokenExchangeRole** from the search results. +2. Select **Add permissions** > **Attach policies**. +3. Search for **AWSIoTFullAccess**, select it, and select **Add permissions**. +4. Repeat for **AmazonS3FullAccess** and **SecretsManagerReadWrite**. -![TERUpdate](./images/IAM_TER_Update.png) +![GreengrassV2TokenExchangeRole permissions page showing the three newly attached policies#center](./images/iam_ter_update.webp "Updated token exchange role permissions") -When done, your GreengrassV2TokenExchangeRole should now show that it has "AWSIoTFullAccess", "S3FullAccess" and "SecretsManagerReadWrite" permissions added to it. +After updating, your `GreengrassV2TokenExchangeRole` should show all three policies attached. -Next, we will clone and configure the EdgeImpulse "Runner" custom component used to deploy the Edge Impulse "Runner" model execution runtime. +## What you've accomplished -Onward! \ No newline at end of file +In this section, you created AWS access credentials, installed Greengrass Nucleus Classic on your edge device, and configured the token exchange role with the permissions that the Edge Impulse component requires. In the next section, you store your Edge Impulse API key in AWS Secrets Manager. \ No newline at end of file diff --git a/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/hardwaresetup.md b/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/hardwaresetup.md index d8fb82612b..459e2eee05 100644 --- a/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/hardwaresetup.md +++ b/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/hardwaresetup.md @@ -1,22 +1,43 @@ --- -title: 1. Hardware Setup +title: Select and set up your edge device weight: 3 ### FIXED, DO NOT MODIFY layout: learningpathall --- -## Edge Device Hardware Setup +## Choose your platform -##### First, an edge device must be setup. In the following sections, Linux-compatible edge devices are detailed to enable them to receive and run as a AWS IoT Greengrass edge device. The list of supported devices will grow over time. Please select one of the following and follow the "Setup" link: +Before you can install AWS IoT Greengrass and deploy an Edge Impulse model, you need a Linux-based Arm device to act as your edge device. This Learning Path supports four platform options. Select the one that matches your available hardware and follow the setup instructions. -### Option 1: Ubuntu EC2 Instance [Setup](/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/HardwareSetupEC2/) +If you don't have any of the physical hardware boards listed below, use the AWS EC2 option. It creates an Arm-based virtual machine in the cloud that behaves like a local edge device, so you can complete every step in this Learning Path without dedicated hardware. -### Option 2: Qualcomm QC6490 Platforms with Ubuntu [Setup](/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/HardwareSetupQC6490Ubuntu/) +Each setup guide installs the required dependencies (build tools, Node.js, GStreamer, Java) and provides a device-specific JSON configuration that you'll use later when deploying the Greengrass component. After completing your chosen setup, return to this page and continue to the next section. -### Option 3: Nvidia Jetson Platforms with Jetpack 5.x/6.0 [Setup](/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/HardwareSetupNvidiaJetson/) +### Option 1: AWS EC2 Arm instance (no hardware required) -### Option 4: Raspberry Pi 5 with RaspberryPi OS [Setup](/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/HardwareSetupRPi5/) +Use this option if you don't have a physical edge device. You create an Ubuntu-based EC2 instance with an Arm processor (Graviton) that simulates a local edge device. Because there is no camera attached, this option uses a sample video file for inference input. +[Set up EC2 instance](/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/hardwaresetupec2/) -#### (More exciting device options will be added soon. Stay tuned!) \ No newline at end of file +### Option 2: Raspberry Pi 5 with Raspberry Pi OS + +The Raspberry Pi 5 is a widely available, affordable Arm board with full Edge Impulse and Greengrass support. You can run inference with an attached USB camera or use a sample video file. + +[Set up Raspberry Pi 5](/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/hardwaresetuprpi5/) + +### Option 3: Nvidia Jetson with Jetpack 5.x or 6.0 + +If you have an Nvidia Jetson board (Nano, Xavier, Orin), you can take advantage of GPU-accelerated inference. This option assumes Jetpack is already flashed onto the device. + +[Set up Nvidia Jetson](/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/hardwaresetupnvidiajetson/) + +### Option 4: Qualcomm QC6490 with Ubuntu + +For Qualcomm QC6490-based development boards running Ubuntu, this option supports both the on-board Qualcomm camera and USB-attached cameras, as well as file-based inference without a camera. + +[Set up Qualcomm QC6490](/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/hardwaresetupqc6490ubuntu/) + +## After setup + +Once you finish the setup for your chosen platform, continue to the next page to create your Edge Impulse project and build a model deployment. \ No newline at end of file diff --git a/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/hardwaresetupec2.md b/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/hardwaresetupec2.md index c621196700..b00941a1df 100644 --- a/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/hardwaresetupec2.md +++ b/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/hardwaresetupec2.md @@ -5,113 +5,144 @@ hide_from_navpane: true layout: learningpathall --- -## Setup and Configuration for Ubuntu-based EC2 instance +## Set up an Ubuntu EC2 Arm instance -### Create Ubuntu EC2 Instance +If you don't have a physical edge device, you can use an AWS EC2 instance with an Arm-based Graviton processor to simulate one. This section walks you through creating the instance, connecting over SSH, and installing the dependencies needed for AWS IoT Greengrass and Edge Impulse. -AWS EC2 instances can be used to simulate edge devices when edge device hardware isn't available. +### Create the EC2 instance -We'll start by opening our AWS Console and search for EC2: +Open the AWS Console and search for **EC2**: -![AWS Console](./images/EC2_Setup_1.png) +![AWS Console search bar with EC2 typed in the search field#center](./images/EC2_Setup_1.png "Search for EC2 in the AWS Console") -We'll now open the EC2 console page: +Open the EC2 console page: -![AWS EC2 Console](./images/EC2_Setup_2.png) +![EC2 dashboard showing the main console page with instance summary#center](./images/EC2_Setup_2.png "EC2 console page") -Select "Launch instance". Provide a Name for the EC2 instance and select the "Ubuntu" Quick Start option. Additionally, select "64-bit(Arm)" as the architecture type and select "t4g.large" as the Instance type: +Select **Launch instance** and configure the following settings: -![Create EC2 Instance](./images/EC2_Setup_3.png) +- Provide a name for the instance (for example, `EdgeDeviceSimulator`). +- Under **Quick Start**, select **Ubuntu**. +- Set the architecture to **64-bit (Arm)**. +- Set the instance type to **t4g.large**. -Additionally, please click on "Create new Key Pair" and provide a name for a new SSH key pair that will be used to SSH into our EC2 instance. Press "Create key pair": +![EC2 instance creation form showing Ubuntu selected with 64-bit Arm architecture and t4g.large instance type#center](./images/EC2_Setup_3.png "EC2 instance configuration") -![Create EC2 Keypair](./images/EC2_Setup_4.png) +### Create an SSH key pair ->**_NOTE:_** ->You will notice that a download will occur with your browser. Save off this key (a .pem file) as we'll use it shortly. +Select **Create new Key Pair** and provide a name for the key pair. Select **Create key pair**: -Next, we need to edit our "Network Settings" for our EC2 instance... scroll down to "Network Settings" and press "Edit": +![Key pair creation dialog with a name field and Create key pair button#center](./images/EC2_Setup_4.png "Create a new SSH key pair") -![Security Group](./images/EC2_Setup_4_ns.png) +{{% notice Note %}} +Your browser downloads a `.pem` file automatically. Save this file in a known location because you need it to SSH into the instance. +{{% /notice %}} -Press "Add security group rule" and lets allow port tcp/4912: +### Configure network settings -![Security Group](./images/EC2_Setup_4_4912.png) +Scroll down to **Network Settings** and select **Edit**: -Lets also give the EC2 instance a bit more disk space. Please change the "8" to "28" here: +![Network settings section of the EC2 launch wizard with an Edit button#center](./images/EC2_Setup_4_ns.png "Edit network settings") -![Increase disk space](./images/EC2_Setup_5.png) +Select **Add security group rule** and add a rule to allow inbound TCP traffic on port 4912. The Edge Impulse Runner serves a web-based inference viewer on this port, which you use later to confirm the model is running. -Finally, press "Launch instance". You should see your EC2 instance getting created: +For both the SSH rule (port 22) and the port 4912 rule, restrict the source to your own IP address rather than allowing access from anywhere. To find your current public IP, run: -![Launch Instance](./images/EC2_Setup_6.png) +```bash +curl http://checkip.amazonaws.com +``` -Now, press "View all instances" and press the refresh button... you should see your new EC2 instance in the "Running" state: +Enter the returned IP address with a `/32` suffix (for example, `203.0.113.10/32`) in the **Source** field for each security group rule. This limits access to your machine only. -![Running Instance](./images/EC2_Setup_7.png) +![Security group rule showing TCP port 4912 allowed for inbound traffic#center](./images/EC2_Setup_4_4912.png "Add security group rule for port 4912") -You can scroll over and save off your Public IPv4 IP Address. You'll need this to SSH into your EC2 instance. +### Increase disk space -Lets now confirm that we can SSH into our EC2 instance. With the saved off pem file and our EC2 Public IPv4 IP address, lets ssh into our EC2 instance - ->**_NOTE:_** ->In this example, my pem file is named DougsEC2SimulatedEdgeDeviceKeyPair.pem and my EC2 instances' public IP address is 1.2.3.4 +The default 8 GB root volume isn't enough for the dependencies and model files. Under **Configure storage**, change the root volume size from `8` to `28` GB: - chmod 600 DougsEC2SimulatedEdgeDeviceKeyPair.pem - ssh -i ./DougsEC2SimulatedEdgeDeviceKeyPair.pem ubuntu@1.2.3.4 +![Storage configuration showing the root volume size set to 28 GB#center](./images/EC2_Setup_5.png "Increase root volume to 28 GB") -You should see a login shell now for your EC2 instance! +### Launch and verify the instance -![Login Shell](./images/EC2_Setup_8.png) +Select **Launch instance**. You should see a confirmation that the instance is being created: -Excellent! You can keep that shell open as we'll make use of it when we start installing Greengrass a bit later. +![Launch confirmation screen showing the instance is being created#center](./images/EC2_Setup_6.png "Instance launch confirmation") -Lastly, lets install the prerequisites that we need. Please run these commands to add some required dependencies: +Select **View all instances** and refresh the page. Your instance should show a **Running** state: - sudo apt update - sudo apt install -y curl unzip - sudo apt install -y gcc g++ make build-essential nodejs sox gstreamer1.0-tools gstreamer1.0-plugins-good gstreamer1.0-plugins-base gstreamer1.0-plugins-base-apps - -Additionally, we need to install the prerequisites for AWS IoT Greengrass "classic": +![EC2 instances list showing the new instance in Running state with a public IP address#center](./images/EC2_Setup_7.png "Running EC2 instance") - sudo apt install -y default-jdk +Copy the **Public IPv4 address** from the instance details. You need this to connect over SSH. -Before we go to the next section, lets also save off this JSON - it will be used to configure our AWS Greengrass custom component a bit later: +### Connect over SSH -#### Non-Camera configuration +Open a terminal and connect to the instance using your `.pem` file and the public IP address. Replace the placeholders with your actual file name and IP: - { - "Parameters": { - "node_version": "20.18.2", - "vips_version": "8.12.1", - "device_name": "MyEC2EdgeDevice", - "launch": "runner", - "sleep_time_sec": 10, - "lock_filename": "/tmp/ei_lockfile_runner", - "gst_args": "filesrc:location=/home/ggc_user/data/testSample.mp4:!:decodebin:!:videoconvert:!:videorate:!:video/x-raw,framerate=2200/1:!:jpegenc", - "eiparams": "--greengrass", - "iotcore_backoff": "-1", - "iotcore_qos": "1", - "ei_bindir": "/usr/local/bin", - "ei_sm_secret_id": "EI_API_KEY", - "ei_sm_secret_name": "ei_api_key", - "ei_poll_sleeptime_ms": 2500, - "ei_local_model_file": "/home/ggc_user/data/currentModel.eim", - "ei_shutdown_behavior": "wait_on_restart", - "ei_ggc_user_groups": "video audio input users system", - "install_kvssink": "no", - "publish_inference_base64_image": "no", - "enable_cache_to_file": "no", - "cache_file_directory": "__none__", - "enable_threshold_limit": "no", - "metrics_sleeptime_ms": 30000, - "default_threshold": 50, - "threshold_criteria": "ge", - "enable_cache_to_s3": "no", - "s3_bucket": "__none__" - } - } +```bash +chmod 600 your-key-pair.pem +ssh -i ./your-key-pair.pem ubuntu@ +``` -OK, Lets proceed to the next step and get our Edge Impulse environment setup! Press "Next" to continue: +You should see a login shell for your EC2 instance: -### [Next](/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/edgeimpulseprojectbuild/) +![Terminal showing a successful SSH login to the Ubuntu EC2 instance#center](./images/EC2_Setup_8.png "SSH login shell") + +Keep this shell open. You'll use it in the following steps. + +### Install dependencies + +The Edge Impulse Runner and AWS IoT Greengrass require several system packages. Update the package list and install the build tools, Node.js, and GStreamer plugins: + +```bash +sudo apt update +sudo apt install -y curl unzip +sudo apt install -y gcc g++ make build-essential nodejs sox gstreamer1.0-tools gstreamer1.0-plugins-good gstreamer1.0-plugins-base gstreamer1.0-plugins-base-apps +``` + +Greengrass Nucleus Classic is Java-based, so you also need a JDK: + +```bash +sudo apt install -y default-jdk +``` + +### Save the component configuration + +The JSON below configures the Edge Impulse Greengrass component for this EC2 instance. Because the instance has no camera, the configuration uses `gst_args` to read inference input from a local video file instead. + +Save this JSON to a text file on your local machine. You'll paste it into the Greengrass deployment configuration in a later step. + +```json +{ + "Parameters": { + "node_version": "20.18.2", + "vips_version": "8.12.1", + "device_name": "MyEC2EdgeDevice", + "launch": "runner", + "sleep_time_sec": 10, + "lock_filename": "/tmp/ei_lockfile_runner", + "gst_args": "filesrc:location=/home/ggc_user/data/testSample.mp4:!:decodebin:!:videoconvert:!:videorate:!:video/x-raw,framerate=2200/1:!:jpegenc", + "eiparams": "--greengrass", + "iotcore_backoff": "-1", + "iotcore_qos": "1", + "ei_bindir": "/usr/local/bin", + "ei_sm_secret_id": "EI_API_KEY", + "ei_sm_secret_name": "ei_api_key", + "ei_poll_sleeptime_ms": 2500, + "ei_local_model_file": "/home/ggc_user/data/currentModel.eim", + "ei_shutdown_behavior": "wait_on_restart", + "ei_ggc_user_groups": "video audio input users system", + "install_kvssink": "no", + "publish_inference_base64_image": "no", + "enable_cache_to_file": "no", + "cache_file_directory": "__none__", + "enable_threshold_limit": "no", + "metrics_sleeptime_ms": 30000, + "default_threshold": 50, + "threshold_criteria": "ge", + "enable_cache_to_s3": "no", + "s3_bucket": "__none__" + } +} +``` + +Your EC2 instance is ready. Return to the [hardware setup page](/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/hardwaresetup/) and continue to the next section to set up your Edge Impulse project. diff --git a/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/hardwaresetupnvidiajetson.md b/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/hardwaresetupnvidiajetson.md index a8f8705829..4a94766dc9 100644 --- a/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/hardwaresetupnvidiajetson.md +++ b/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/hardwaresetupnvidiajetson.md @@ -5,95 +5,157 @@ hide_from_navpane: true layout: learningpathall --- -## Install/Configure Nvidia Jetpack (Jetson devices) - -The workshop will assume that the Nvidia Jetson edge device has been loaded with Jetpack 5.x and/or Jetpack 6.0 per flashing instructions located at this [Nvidia website](https://docs.nvidia.com/jetson/archives/r34.1/DeveloperGuide/index.html#page/Tegra%20Linux%20Driver%20Package%20Development%20Guide/flashing.html). - -### Additional Setup - -Once you have your Jetson platform installed and running, please run these commands to add some required dependencies: - - sudo apt update - sudo apt install -y curl unzip - sudo apt install -y gcc g++ make build-essential nodejs sox gstreamer1.0-tools gstreamer1.0-plugins-good gstreamer1.0-plugins-base gstreamer1.0-plugins-base-apps - -Additionally, we need to install the prerequisites for AWS IoT Greengrass "classic": - - sudo apt install -y default-jdk - -Lastly, its recommended to update your linux device with the latest security patches and updates if available. - -We are now setup! Before we continue, please save off the following JSONs. These JSONs will be used to configure our AWS Greengrass deployment. - -#### Camera configuration - - { - "Parameters": { - "node_version": "20.18.2", - "vips_version": "8.12.1", - "device_name": "MyNvidiaJetsonEdgeDevice", - "launch": "runner", - "sleep_time_sec": 10, - "lock_filename": "/tmp/ei_lockfile_runner", - "gst_args": "v4l2src:device=/dev/video0:!:video/x-raw,width=640,height=480:!:videoconvert:!:jpegenc", - "eiparams": "--greengrass --force-variant float32 --silent", - "iotcore_backoff": "-1", - "iotcore_qos": "1", - "ei_bindir": "/usr/local/bin", - "ei_sm_secret_id": "EI_API_KEY", - "ei_sm_secret_name": "ei_api_key", - "ei_poll_sleeptime_ms": 2500, - "ei_local_model_file": "__none__", - "ei_shutdown_behavior": "__none__", - "ei_ggc_user_groups": "video audio input users system render", - "install_kvssink": "no", - "publish_inference_base64_image": "no", - "enable_cache_to_file": "no", - "cache_file_directory": "__none__", - "enable_threshold_limit": "no", - "metrics_sleeptime_ms": 30000, - "default_threshold": 65.0, - "threshold_criteria": "ge", - "enable_cache_to_s3": "no", - "s3_bucket": "__none__" - } - } - - -#### Non-Camera configuration - - { - "Parameters": { - "node_version": "20.18.2", - "vips_version": "8.12.1", - "device_name": "MyNvidiaJetsonEdgeDevice", - "launch": "runner", - "sleep_time_sec": 10, - "lock_filename": "/tmp/ei_lockfile_runner", - "gst_args": "filesrc:location=/home/ggc_user/data/testSample.mp4:!:decodebin:!:videoconvert:!:videorate:!:video/x-raw,framerate=2200/1:!:jpegenc", - "eiparams": "--greengrass", - "iotcore_backoff": "-1", - "iotcore_qos": "1", - "ei_bindir": "/usr/local/bin", - "ei_sm_secret_id": "EI_API_KEY", - "ei_sm_secret_name": "ei_api_key", - "ei_poll_sleeptime_ms": 2500, - "ei_local_model_file": "/home/ggc_user/data/currentModel.eim", - "ei_shutdown_behavior": "wait_on_restart", - "ei_ggc_user_groups": "video audio input users system render", - "install_kvssink": "no", - "publish_inference_base64_image": "no", - "enable_cache_to_file": "no", - "cache_file_directory": "__none__", - "enable_threshold_limit": "no", - "metrics_sleeptime_ms": 30000, - "default_threshold": 50, - "threshold_criteria": "ge", - "enable_cache_to_s3": "no", - "s3_bucket": "__none__" - } - } - -OK! Lets continue by getting our Edge Impulse project setup! Let's go! Press "Next" to continue: - -### [Next](/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/edgeimpulseprojectbuild/) \ No newline at end of file +## Set up an Nvidia Jetson with Jetpack + +Nvidia Jetson boards (Nano, Xavier, Orin) provide GPU-accelerated inference for Edge Impulse models. This section covers prerequisites, dependency installation, and the component configuration for running the Edge Impulse Runner on a Jetson device with AWS IoT Greengrass. + +### Prerequisites + +Before you begin, make sure you have: + +- An Nvidia Jetson board with a power supply. +- Jetpack 5.x or 6.0 already flashed onto the device. If you haven't done this yet, follow the [Nvidia Jetson flashing instructions](https://docs.nvidia.com/jetson/archives/r34.1/DeveloperGuide/index.html#page/Tegra%20Linux%20Driver%20Package%20Development%20Guide/flashing.html). +- A network connection (Ethernet or Wi-Fi) and SSH access to the device. +- Optional: a USB camera for live inference. Without a camera, the Runner uses a sample video file instead. + +### Verify Jetpack version + +After booting the Jetson, confirm which Jetpack version is installed: + +```bash +cat /etc/nv_tegra_release +``` + +You should see output that indicates L4T (Linux for Tegra) version 34.x or later for Jetpack 5.x, or version 36.x for Jetpack 6.0. + +### Connect over SSH + +If you haven't already, connect to the Jetson from your computer. Replace the placeholder with the device's IP address: + +```bash +ssh your-username@ +``` + +If you're not sure of the IP address, you can check your router's admin page for connected devices, or run `hostname -I` on the Jetson if you have a monitor connected. + +### Install dependencies + +Update the package list and install the build tools, Node.js, and GStreamer plugins that the Edge Impulse Runner requires: + +```bash +sudo apt update +sudo apt install -y curl unzip +sudo apt install -y gcc g++ make build-essential nodejs sox gstreamer1.0-tools gstreamer1.0-plugins-good gstreamer1.0-plugins-base gstreamer1.0-plugins-base-apps +``` + +Greengrass Nucleus Classic is Java-based, so you also need a JDK: + +```bash +sudo apt install -y default-jdk +``` + +Install any available security updates: + +```bash +sudo apt upgrade -y +``` + +### Verify the camera (optional) + +If you have a USB camera connected, confirm the system detects it: + +```bash +ls /dev/video* +``` + +You should see at least `/dev/video0` in the output. If nothing appears, check that the camera is plugged in securely and try a different USB port. + +### Jetpack 6.x note on GPU access + +If your device is running Jetpack 6.x or later, the `render` group is required for the Greengrass service user to access the GPU. Both JSON configurations below already include `render` in the `ei_ggc_user_groups` field. If you're running Jetpack 5.x, you can remove `render` from that field, though leaving it in place doesn't cause issues. + +### Save the component configuration + +The JSON configurations below set up the Edge Impulse Greengrass component for the Jetson. Choose the configuration that matches your setup and save it to a text file on your local machine. You'll paste it into the Greengrass deployment configuration in a later step. + +#### With a USB camera + +This configuration captures live video from `/dev/video0` at 640x480 resolution. The `--force-variant float32` flag selects the float32 model variant, and `--silent` suppresses console output since the Runner runs as a background service. + +```json +{ + "Parameters": { + "node_version": "20.18.2", + "vips_version": "8.12.1", + "device_name": "MyNvidiaJetsonEdgeDevice", + "launch": "runner", + "sleep_time_sec": 10, + "lock_filename": "/tmp/ei_lockfile_runner", + "gst_args": "v4l2src:device=/dev/video0:!:video/x-raw,width=640,height=480:!:videoconvert:!:jpegenc", + "eiparams": "--greengrass --force-variant float32 --silent", + "iotcore_backoff": "-1", + "iotcore_qos": "1", + "ei_bindir": "/usr/local/bin", + "ei_sm_secret_id": "EI_API_KEY", + "ei_sm_secret_name": "ei_api_key", + "ei_poll_sleeptime_ms": 2500, + "ei_local_model_file": "__none__", + "ei_shutdown_behavior": "__none__", + "ei_ggc_user_groups": "video audio input users system render", + "install_kvssink": "no", + "publish_inference_base64_image": "no", + "enable_cache_to_file": "no", + "cache_file_directory": "__none__", + "enable_threshold_limit": "no", + "metrics_sleeptime_ms": 30000, + "default_threshold": 65.0, + "threshold_criteria": "ge", + "enable_cache_to_s3": "no", + "s3_bucket": "__none__" + } +} +``` + +#### Without a camera + +This configuration reads inference input from a local sample video file. The `ei_local_model_file` field points to a pre-downloaded model, and `ei_shutdown_behavior` is set to `wait_on_restart` so the Runner pauses after the video ends and waits for a restart command. + +```json +{ + "Parameters": { + "node_version": "20.18.2", + "vips_version": "8.12.1", + "device_name": "MyNvidiaJetsonEdgeDevice", + "launch": "runner", + "sleep_time_sec": 10, + "lock_filename": "/tmp/ei_lockfile_runner", + "gst_args": "filesrc:location=/home/ggc_user/data/testSample.mp4:!:decodebin:!:videoconvert:!:videorate:!:video/x-raw,framerate=2200/1:!:jpegenc", + "eiparams": "--greengrass", + "iotcore_backoff": "-1", + "iotcore_qos": "1", + "ei_bindir": "/usr/local/bin", + "ei_sm_secret_id": "EI_API_KEY", + "ei_sm_secret_name": "ei_api_key", + "ei_poll_sleeptime_ms": 2500, + "ei_local_model_file": "/home/ggc_user/data/currentModel.eim", + "ei_shutdown_behavior": "wait_on_restart", + "ei_ggc_user_groups": "video audio input users system render", + "install_kvssink": "no", + "publish_inference_base64_image": "no", + "enable_cache_to_file": "no", + "cache_file_directory": "__none__", + "enable_threshold_limit": "no", + "metrics_sleeptime_ms": 30000, + "default_threshold": 50, + "threshold_criteria": "ge", + "enable_cache_to_s3": "no", + "s3_bucket": "__none__" + } +} +``` + +{{% notice Note %}} +When running a model compiled specifically for a Jetson GPU, the first invocation can take 2-3 minutes while the model loads into GPU memory. Subsequent invocations are much faster. +{{% /notice %}} + +Your Nvidia Jetson is ready. Return to the [hardware setup page](/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/hardwaresetup/) and continue to the next section to set up your Edge Impulse project. \ No newline at end of file diff --git a/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/hardwaresetupqc6490ubuntu.md b/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/hardwaresetupqc6490ubuntu.md index b4d65ba220..640393f580 100644 --- a/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/hardwaresetupqc6490ubuntu.md +++ b/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/hardwaresetupqc6490ubuntu.md @@ -5,128 +5,191 @@ hide_from_navpane: true layout: learningpathall --- -## Ubuntu-based QC6490 platforms - -First, please flash your QC6490 device per your manufacturers instructions to load up Ubuntu onto the device. - -### Additional Setup - -Once you have your Ubuntu platform installed and running, please run these commands to add some required dependencies: - - sudo apt update - sudo apt install -y curl unzip - sudo apt install -y gcc g++ make build-essential nodejs sox gstreamer1.0-tools gstreamer1.0-plugins-good gstreamer1.0-plugins-base gstreamer1.0-plugins-base-apps - -Additionally, we need to install the prerequisites for AWS IoT Greengrass "classic": - - sudo apt install -y default-jdk - -Lastly, its recommended to update your linux device with the latest security patches and updates if available. - -We are now setup! Before we continue, please save off the following JSONs. These JSONs will be used to configure our AWS Greengrass deployment. - -#### QC Camera configuration - - { - "Parameters": { - "node_version": "20.18.2", - "vips_version": "8.12.1", - "device_name": "MyQC6490UbuntuEdgeDevice", - "launch": "runner", - "sleep_time_sec": 10, - "lock_filename": "/tmp/ei_lockfile_runner", - "gst_args": "qtiqmmfsrc:name=camsrc:camera=0:!:video/x-raw,width=1280,height=720:!:videoconvert:!:jpegenc", - "eiparams": "--greengrass --force-variant float32 --silent", - "iotcore_backoff": "-1", - "iotcore_qos": "1", - "ei_bindir": "/usr/local/bin", - "ei_sm_secret_id": "EI_API_KEY", - "ei_sm_secret_name": "ei_api_key", - "ei_poll_sleeptime_ms": 2500, - "ei_local_model_file": "__none__", - "ei_shutdown_behavior": "__none__", - "ei_ggc_user_groups": "video audio input users", - "install_kvssink": "no", - "publish_inference_base64_image": "no", - "enable_cache_to_file": "no", - "cache_file_directory": "__none__", - "enable_threshold_limit": "no", - "metrics_sleeptime_ms": 30000, - "default_threshold": 65.0, - "threshold_criteria": "ge", - "enable_cache_to_s3": "no", - "s3_bucket": "__none__" - } - } - -#### USB-attached Camera configuration - - { - "Parameters": { - "node_version": "20.18.2", - "vips_version": "8.12.1", - "device_name": "MyQC6490UbuntuEdgeDevice", - "launch": "runner", - "sleep_time_sec": 10, - "lock_filename": "/tmp/ei_lockfile_runner", - "gst_args": "v4l2src:device=/dev/video0:!:video/x-raw,width=640,height=480:!:videoconvert:!:jpegenc", - "eiparams": "--greengrass --force-variant float32 --silent", - "iotcore_backoff": "-1", - "iotcore_qos": "1", - "ei_bindir": "/usr/local/bin", - "ei_sm_secret_id": "EI_API_KEY", - "ei_sm_secret_name": "ei_api_key", - "ei_poll_sleeptime_ms": 2500, - "ei_local_model_file": "__none__", - "ei_shutdown_behavior": "__none__", - "ei_ggc_user_groups": "video audio input users", - "install_kvssink": "no", - "publish_inference_base64_image": "no", - "enable_cache_to_file": "no", - "cache_file_directory": "__none__", - "enable_threshold_limit": "no", - "metrics_sleeptime_ms": 30000, - "default_threshold": 65.0, - "threshold_criteria": "ge", - "enable_cache_to_s3": "no", - "s3_bucket": "__none__" - } - } - -#### Non-Camera configuration - - { - "Parameters": { - "node_version": "20.18.2", - "vips_version": "8.12.1", - "device_name": "MyQC6490UbuntuEdgeDevice", - "launch": "runner", - "sleep_time_sec": 10, - "lock_filename": "/tmp/ei_lockfile_runner", - "gst_args": "filesrc:location=/home/ggc_user/data/testSample.mp4:!:decodebin:!:videoconvert:!:videorate:!:video/x-raw,framerate=2200/1:!:jpegenc", - "eiparams": "--greengrass", - "iotcore_backoff": "-1", - "iotcore_qos": "1", - "ei_bindir": "/usr/local/bin", - "ei_sm_secret_id": "EI_API_KEY", - "ei_sm_secret_name": "ei_api_key", - "ei_poll_sleeptime_ms": 2500, - "ei_local_model_file": "/home/ggc_user/data/currentModel.eim", - "ei_shutdown_behavior": "wait_on_restart", - "ei_ggc_user_groups": "video audio input users system", - "install_kvssink": "no", - "publish_inference_base64_image": "no", - "enable_cache_to_file": "no", - "cache_file_directory": "__none__", - "enable_threshold_limit": "no", - "metrics_sleeptime_ms": 30000, - "default_threshold": 50, - "threshold_criteria": "ge", - "enable_cache_to_s3": "no", - "s3_bucket": "__none__" - } - } - -OK! Lets continue by getting our Edge Impulse project setup! Let's go! Press "Next" to continue: - -### [Next](/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/edgeimpulseprojectbuild/) \ No newline at end of file +## Set up a Qualcomm Dragonwing QC6490 with Ubuntu + +The Qualcomm Dragonwing QC6490 is an Arm-based platform that supports both the on-board Qualcomm camera module and USB-attached cameras for live inference with Edge Impulse. This section covers prerequisites, dependency installation, and the component configuration for running the Edge Impulse Runner on a QC6490 device with AWS IoT Greengrass. + +### Prerequisites + +Before you begin, make sure you have: + +- A Qualcomm Dragonwing QC6490 development board with a power supply. +- Ubuntu flashed onto the device per the [Qualcomm QC6490 quick start guide](https://docs.qualcomm.com/doc/80-90441-1/topic/qsg-landing-page.html). +- A network connection (Ethernet or Wi-Fi) and SSH access to the device. +- Optional: the on-board Qualcomm camera module or a USB camera for live inference. Without a camera, the Runner uses a sample video file instead. + +### Connect over SSH + +Connect to the QC6490 from your computer. Replace the placeholder with the device's IP address: + +```bash +ssh your-username@ +``` + +If you're not sure of the IP address, check your router's admin page for connected devices, or run `hostname -I` on the QC6490 if you have a monitor connected. + +### Verify Ubuntu is running + +Confirm the device is running Ubuntu on aarch64: + +```bash +uname -a +``` + +The output should show `aarch64` as the architecture and an Ubuntu kernel version. + +### Install dependencies + +Update the package list and install the build tools, Node.js, and GStreamer plugins that the Edge Impulse Runner requires: + +```bash +sudo apt update +sudo apt install -y curl unzip +sudo apt install -y gcc g++ make build-essential nodejs sox gstreamer1.0-tools gstreamer1.0-plugins-good gstreamer1.0-plugins-base gstreamer1.0-plugins-base-apps +``` + +Greengrass Nucleus Classic is Java-based, so you also need a JDK: + +```bash +sudo apt install -y default-jdk +``` + +Install any available security updates: + +```bash +sudo apt upgrade -y +``` + +### Verify the camera (optional) + +The QC6490 supports two types of cameras. The type you have determines which JSON configuration to use later. + +**On-board Qualcomm camera**: This uses the `qtiqmmfsrc` GStreamer element, which is specific to Qualcomm platforms. If your board has a built-in camera module, it should be available without additional setup. + +**USB-attached camera**: If you're using a USB camera instead, confirm the system detects it: + +```bash +ls /dev/video* +``` + +You should see at least `/dev/video0` in the output. If nothing appears, check that the camera is plugged in securely and try a different USB port. + +### Save the component configuration + +The JSON configurations below set up the Edge Impulse Greengrass component for the QC6490. This platform has three configuration options depending on your camera setup. Choose the one that matches your hardware and save it to a text file on your local machine. You'll paste it into the Greengrass deployment configuration in a later step. + +#### With the on-board Qualcomm camera + +This configuration uses the `qtiqmmfsrc` GStreamer element to capture video from the on-board camera at 1280x720 resolution. The `--force-variant float32` flag selects the float32 model variant, and `--silent` suppresses console output since the Runner runs as a background service. + +```json +{ + "Parameters": { + "node_version": "20.18.2", + "vips_version": "8.12.1", + "device_name": "MyQC6490UbuntuEdgeDevice", + "launch": "runner", + "sleep_time_sec": 10, + "lock_filename": "/tmp/ei_lockfile_runner", + "gst_args": "qtiqmmfsrc:name=camsrc:camera=0:!:video/x-raw,width=1280,height=720:!:videoconvert:!:jpegenc", + "eiparams": "--greengrass --force-variant float32 --silent", + "iotcore_backoff": "-1", + "iotcore_qos": "1", + "ei_bindir": "/usr/local/bin", + "ei_sm_secret_id": "EI_API_KEY", + "ei_sm_secret_name": "ei_api_key", + "ei_poll_sleeptime_ms": 2500, + "ei_local_model_file": "__none__", + "ei_shutdown_behavior": "__none__", + "ei_ggc_user_groups": "video audio input users", + "install_kvssink": "no", + "publish_inference_base64_image": "no", + "enable_cache_to_file": "no", + "cache_file_directory": "__none__", + "enable_threshold_limit": "no", + "metrics_sleeptime_ms": 30000, + "default_threshold": 65.0, + "threshold_criteria": "ge", + "enable_cache_to_s3": "no", + "s3_bucket": "__none__" + } +} +``` + +#### With a USB-attached camera + +This configuration uses the standard `v4l2src` GStreamer element to capture video from a USB camera at 640x480 resolution. Use this if your QC6490 board doesn't have a built-in camera module, or if you prefer to use an external USB camera. + +```json +{ + "Parameters": { + "node_version": "20.18.2", + "vips_version": "8.12.1", + "device_name": "MyQC6490UbuntuEdgeDevice", + "launch": "runner", + "sleep_time_sec": 10, + "lock_filename": "/tmp/ei_lockfile_runner", + "gst_args": "v4l2src:device=/dev/video0:!:video/x-raw,width=640,height=480:!:videoconvert:!:jpegenc", + "eiparams": "--greengrass --force-variant float32 --silent", + "iotcore_backoff": "-1", + "iotcore_qos": "1", + "ei_bindir": "/usr/local/bin", + "ei_sm_secret_id": "EI_API_KEY", + "ei_sm_secret_name": "ei_api_key", + "ei_poll_sleeptime_ms": 2500, + "ei_local_model_file": "__none__", + "ei_shutdown_behavior": "__none__", + "ei_ggc_user_groups": "video audio input users", + "install_kvssink": "no", + "publish_inference_base64_image": "no", + "enable_cache_to_file": "no", + "cache_file_directory": "__none__", + "enable_threshold_limit": "no", + "metrics_sleeptime_ms": 30000, + "default_threshold": 65.0, + "threshold_criteria": "ge", + "enable_cache_to_s3": "no", + "s3_bucket": "__none__" + } +} +``` + +#### Without a camera + +This configuration reads inference input from a local sample video file. The `ei_local_model_file` field points to a pre-downloaded model, and `ei_shutdown_behavior` is set to `wait_on_restart` so the Runner pauses after the video ends and waits for a restart command. + +```json +{ + "Parameters": { + "node_version": "20.18.2", + "vips_version": "8.12.1", + "device_name": "MyQC6490UbuntuEdgeDevice", + "launch": "runner", + "sleep_time_sec": 10, + "lock_filename": "/tmp/ei_lockfile_runner", + "gst_args": "filesrc:location=/home/ggc_user/data/testSample.mp4:!:decodebin:!:videoconvert:!:videorate:!:video/x-raw,framerate=2200/1:!:jpegenc", + "eiparams": "--greengrass", + "iotcore_backoff": "-1", + "iotcore_qos": "1", + "ei_bindir": "/usr/local/bin", + "ei_sm_secret_id": "EI_API_KEY", + "ei_sm_secret_name": "ei_api_key", + "ei_poll_sleeptime_ms": 2500, + "ei_local_model_file": "/home/ggc_user/data/currentModel.eim", + "ei_shutdown_behavior": "wait_on_restart", + "ei_ggc_user_groups": "video audio input users system", + "install_kvssink": "no", + "publish_inference_base64_image": "no", + "enable_cache_to_file": "no", + "cache_file_directory": "__none__", + "enable_threshold_limit": "no", + "metrics_sleeptime_ms": 30000, + "default_threshold": 50, + "threshold_criteria": "ge", + "enable_cache_to_s3": "no", + "s3_bucket": "__none__" + } +} +``` + +Your Qualcomm Dragonwing QC6490 is ready. Return to the [hardware setup page](/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/hardwaresetup/) and continue to the next section to set up your Edge Impulse project. \ No newline at end of file diff --git a/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/hardwaresetuprpi5.md b/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/hardwaresetuprpi5.md index 38c40e0a9d..096adfa461 100644 --- a/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/hardwaresetuprpi5.md +++ b/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/hardwaresetuprpi5.md @@ -5,109 +5,166 @@ hide_from_navpane: true layout: learningpathall --- -## Setup and configuration of Raspberry Pi 5 with Raspberry Pi OS - -### Install RaspberryPi OS - -The Raspberry Pi 5 is a super simple device that is fully supported by Edge Impulse and AWS as an edge device. - -First step in this exercise is to install the latest version of the Raspberry Pi OS onto your RPi. A SD card will be required and typically should be at least 16GB in size. - -The easiest way to setup Raspberry Pi OS is to follow the instructions here after downloading and installing the Raspberry Pi Imager application: - -![Raspberry Pi Imager](./images/RPi_Imager.png) - -Instructions: [Install Raspberry Pi Imager](https://www.raspberrypi.com/software/) - -Please save off the IP address of your edge device along with login credentials to remote SSH into the edge device. You'll need these in the next steps. - -#### Additional Prerequisites - -First, lets open a shell into your RPi (using the Raspberry Pi OS default username of "pi" with password "raspberrypi" and having an IP address of 1.2.3.4): - - ssh pi@1.2.3.4 - -Once logged in via ssh, lets install the prerequisites that we need. Please run these commands to add some required dependencies: - - sudo apt update - sudo apt install -y curl unzip - sudo apt install -y gcc g++ make build-essential nodejs sox gstreamer1.0-tools gstreamer1.0-plugins-good gstreamer1.0-plugins-base gstreamer1.0-plugins-base-apps - -Additionally, we need to install the prerequisites for AWS IoT Greengrass "classic": - - sudo apt install -y default-jdk - -Lastly, please safe off these JSONs. These will be used to customize our AWS Greengrass custom component based upon using an RPi5 device with or without a camera: - -#### Camera configuration - - { - "Parameters": { - "node_version": "20.18.2", - "vips_version": "8.12.1", - "device_name": "MyRPi5EdgeDevice", - "launch": "runner", - "sleep_time_sec": 10, - "lock_filename": "/tmp/ei_lockfile_runner", - "gst_args": "v4l2src:device=/dev/video0:!:video/x-raw,width=640,height=480:!:videoconvert:!:jpegenc", - "eiparams": "--greengrass --force-variant float32 --silent", - "iotcore_backoff": "-1", - "iotcore_qos": "1", - "ei_bindir": "/usr/local/bin", - "ei_sm_secret_id": "EI_API_KEY", - "ei_sm_secret_name": "ei_api_key", - "ei_poll_sleeptime_ms": 2500, - "ei_local_model_file": "__none__", - "ei_shutdown_behavior": "__none__", - "ei_ggc_user_groups": "video audio input users", - "install_kvssink": "no", - "publish_inference_base64_image": "no", - "enable_cache_to_file": "no", - "cache_file_directory": "__none__", - "enable_threshold_limit": "no", - "metrics_sleeptime_ms": 30000, - "default_threshold": 65.0, - "threshold_criteria": "ge", - "enable_cache_to_s3": "no", - "s3_bucket": "__none__" - } - } - - -#### Non-Camera configuration - - { - "Parameters": { - "node_version": "20.18.2", - "vips_version": "8.12.1", - "device_name": "MyRPi5EdgeDevice", - "launch": "runner", - "sleep_time_sec": 10, - "lock_filename": "/tmp/ei_lockfile_runner", - "gst_args": "filesrc:location=/home/ggc_user/data/testSample.mp4:!:decodebin:!:videoconvert:!:videorate:!:video/x-raw,framerate=2200/1:!:jpegenc", - "eiparams": "--greengrass", - "iotcore_backoff": "-1", - "iotcore_qos": "1", - "ei_bindir": "/usr/local/bin", - "ei_sm_secret_id": "EI_API_KEY", - "ei_sm_secret_name": "ei_api_key", - "ei_poll_sleeptime_ms": 2500, - "ei_local_model_file": "/home/ggc_user/data/currentModel.eim", - "ei_shutdown_behavior": "wait_on_restart", - "ei_ggc_user_groups": "video audio input users system", - "install_kvssink": "no", - "publish_inference_base64_image": "no", - "enable_cache_to_file": "no", - "cache_file_directory": "__none__", - "enable_threshold_limit": "no", - "metrics_sleeptime_ms": 30000, - "default_threshold": 50, - "threshold_criteria": "ge", - "enable_cache_to_s3": "no", - "s3_bucket": "__none__" - } - } - -Alright! Lets continue by getting our Edge Impulse project setup! Let's go! Press "Next" to continue: - -### [Next](/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/edgeimpulseprojectbuild/) \ No newline at end of file +## Set up a Raspberry Pi 5 with Raspberry Pi OS + +The Raspberry Pi 5 is a widely available Arm-based board with full support for both Edge Impulse and AWS IoT Greengrass. This section covers flashing Raspberry Pi OS, enabling SSH, installing dependencies, and preparing the component configuration. + +### What you need + +- A Raspberry Pi 5 board with a power supply (USB-C, 5V/5A recommended). +- A microSD card, 16 GB minimum (32 GB recommended for comfortable headroom). +- A computer with an SD card reader to flash the OS image. +- A network connection (Ethernet or Wi-Fi) for the Raspberry Pi 5. +- Optional: a USB camera if you want to run live inference. Without a camera, the Runner uses a sample video file instead. + +### Flash Raspberry Pi OS + +Download and install the [Raspberry Pi Imager](https://www.raspberrypi.com/software/) on your computer. + +![Raspberry Pi Imager application showing the main screen with device, OS, and storage selection fields#center](./images/RPi_Imager.png "Raspberry Pi Imager") + +Open the Imager and configure the following: + +1. Select **Raspberry Pi 5** as the device. +2. Select **Raspberry Pi OS (64-bit)** as the operating system. The 64-bit version is required for aarch64 compatibility with Edge Impulse models. +3. Select your microSD card as the storage target. +4. Select the gear icon (or **Edit Settings**) to open the advanced options. Configure these settings: + - **Set hostname**: Choose a recognizable name (for example, `rpi5-edge`). + - **Enable SSH**: Select **Use password authentication**. + - **Set username and password**: Create a username and password you'll remember. Raspberry Pi OS no longer includes default credentials. + - **Configure wireless LAN**: Enter your Wi-Fi network name and password if you're not using Ethernet. +5. Select **Write** and wait for the flashing process to complete. + +Insert the microSD card into your Raspberry Pi 5 and power it on. Give it a minute or two to complete its first boot. + +### Find the IP address + +You need the IP address of your Raspberry Pi 5 to connect over SSH. There are several ways to find it: + +- Check your router's admin page for connected devices. +- If you set a hostname (for example, `rpi5-edge`), try `ping rpi5-edge.local` from your computer. +- If you have a monitor connected, open a terminal on the Raspberry Pi 5 and run `hostname -I`. + +Note the IP address for the next step. + +### Connect over SSH + +Open a terminal on your computer and connect to the Raspberry Pi 5 using the username and IP address you configured: + +```bash +ssh your-username@ +``` + +### Install dependencies + +Update the package list and install the build tools, Node.js, and GStreamer plugins that the Edge Impulse Runner requires: + +```bash +sudo apt update +sudo apt install -y curl unzip +sudo apt install -y gcc g++ make build-essential nodejs sox gstreamer1.0-tools gstreamer1.0-plugins-good gstreamer1.0-plugins-base gstreamer1.0-plugins-base-apps +``` + +Greengrass Nucleus Classic is Java-based, so you also need a JDK: + +```bash +sudo apt install -y default-jdk +``` + +It's also a good idea to install any available security patches: + +```bash +sudo apt upgrade -y +``` + +### Verify the camera (optional) + +If you have a USB camera connected, confirm the system detects it: + +```bash +ls /dev/video* +``` + +You should see at least `/dev/video0` in the output. If nothing appears, check that the camera is plugged in securely and try a different USB port. + +### Save the component configuration + +The JSON configurations below set up the Edge Impulse Greengrass component for the Raspberry Pi 5. Choose the configuration that matches your setup and save it to a text file on your local machine. You'll paste it into the Greengrass deployment configuration in a later step. + +#### With a USB camera + +This configuration uses `gst_args` to capture live video from `/dev/video0` at 640x480 resolution. The `--force-variant float32` flag selects the float32 model variant, and `--silent` suppresses console output since the Runner runs as a background service. + +```json +{ + "Parameters": { + "node_version": "20.18.2", + "vips_version": "8.12.1", + "device_name": "MyRPi5EdgeDevice", + "launch": "runner", + "sleep_time_sec": 10, + "lock_filename": "/tmp/ei_lockfile_runner", + "gst_args": "v4l2src:device=/dev/video0:!:video/x-raw,width=640,height=480:!:videoconvert:!:jpegenc", + "eiparams": "--greengrass --force-variant float32 --silent", + "iotcore_backoff": "-1", + "iotcore_qos": "1", + "ei_bindir": "/usr/local/bin", + "ei_sm_secret_id": "EI_API_KEY", + "ei_sm_secret_name": "ei_api_key", + "ei_poll_sleeptime_ms": 2500, + "ei_local_model_file": "__none__", + "ei_shutdown_behavior": "__none__", + "ei_ggc_user_groups": "video audio input users", + "install_kvssink": "no", + "publish_inference_base64_image": "no", + "enable_cache_to_file": "no", + "cache_file_directory": "__none__", + "enable_threshold_limit": "no", + "metrics_sleeptime_ms": 30000, + "default_threshold": 65.0, + "threshold_criteria": "ge", + "enable_cache_to_s3": "no", + "s3_bucket": "__none__" + } +} +``` + +#### Without a camera + +This configuration reads inference input from a local sample video file instead of a live camera feed. The `ei_local_model_file` field points to a pre-downloaded model, and `ei_shutdown_behavior` is set to `wait_on_restart` so the Runner pauses after the video ends and waits for a restart command. + +```json +{ + "Parameters": { + "node_version": "20.18.2", + "vips_version": "8.12.1", + "device_name": "MyRPi5EdgeDevice", + "launch": "runner", + "sleep_time_sec": 10, + "lock_filename": "/tmp/ei_lockfile_runner", + "gst_args": "filesrc:location=/home/ggc_user/data/testSample.mp4:!:decodebin:!:videoconvert:!:videorate:!:video/x-raw,framerate=2200/1:!:jpegenc", + "eiparams": "--greengrass", + "iotcore_backoff": "-1", + "iotcore_qos": "1", + "ei_bindir": "/usr/local/bin", + "ei_sm_secret_id": "EI_API_KEY", + "ei_sm_secret_name": "ei_api_key", + "ei_poll_sleeptime_ms": 2500, + "ei_local_model_file": "/home/ggc_user/data/currentModel.eim", + "ei_shutdown_behavior": "wait_on_restart", + "ei_ggc_user_groups": "video audio input users system", + "install_kvssink": "no", + "publish_inference_base64_image": "no", + "enable_cache_to_file": "no", + "cache_file_directory": "__none__", + "enable_threshold_limit": "no", + "metrics_sleeptime_ms": 30000, + "default_threshold": 50, + "threshold_criteria": "ge", + "enable_cache_to_s3": "no", + "s3_bucket": "__none__" + } +} +``` + +Your Raspberry Pi 5 is ready. Return to the [hardware setup page](/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/hardwaresetup/) and continue to the next section to set up your Edge Impulse project. \ No newline at end of file diff --git a/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/noncameracustomcomponent.md b/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/noncameracustomcomponent.md index b818d04e49..0fbf12a533 100644 --- a/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/noncameracustomcomponent.md +++ b/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/noncameracustomcomponent.md @@ -11,7 +11,7 @@ For those edge devices that do not contain a camera, the following component wil ### 1. Clone the component repo -Please clone this [repo](https://github.com/edgeimpulse/aws-greengrass-workshop-supplemental). You will find the following files: +Clone this [repo](https://github.com/edgeimpulse/aws-greengrass-workshop-supplemental). You'll find the following files: EdgeImpulseRunnerRuntimeInstallerComponent.yaml artifacts/EdgeImpulseRunnerRuntime/1.0.0/install.sh diff --git a/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/overview.md b/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/overview.md index 17296d8266..0d7dea64e6 100644 --- a/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/overview.md +++ b/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/overview.md @@ -1,52 +1,101 @@ --- -title: 0. Overview +title: Overview of Edge Impulse and AWS IoT Greengrass weight: 2 ### FIXED, DO NOT MODIFY layout: learningpathall --- -# Edge Impulse with AWS IoT Greengrass +## What is Edge Impulse with AWS IoT Greengrass? -AWS IoT Greengrass is an AWS IoT service that enables edge devices with customizable/downloadable/installable "components" that can be run to augment what's running on the edge device itself. AWS IoT Greengrass permits the creation and publication of a "Greengrass Component" that is effectively a set of instructions and artifacts that, when installed and run, create and initiate a custom specified service. +Running machine learning models on edge devices is only part of the challenge. In production, you also need a way to deploy models at scale, collect inference results centrally, and manage device fleets remotely. This Learning Path shows you how to connect Edge Impulse and AWS IoT Greengrass to accomplish exactly that on Arm-based Linux devices. -For more information about AWS IoT Core and AWS Greengrass please review: [AWS IoT Greengrass](https://docs.aws.amazon.com/greengrass/v2/developerguide/what-is-iot-greengrass.html) +[Edge Impulse](https://edgeimpulse.com/) is a development platform for building, training, and optimizing ML models purpose-built for edge devices. It handles the full ML lifecycle, from data collection through model deployment, and outputs lightweight models optimized for Arm processors. -## Overview +[AWS IoT Greengrass](https://docs.aws.amazon.com/greengrass/v2/developerguide/what-is-iot-greengrass.html) is an AWS IoT service that extends cloud capabilities to edge devices. It lets you package software into reusable "components" that can be deployed, configured, and updated remotely across thousands of devices. Combined with AWS IoT Core, it provides a managed MQTT message broker (a lightweight publish/subscribe messaging protocol commonly used in IoT) for relaying data from edge devices to the cloud. + +By integrating the two, you can train an ML model in Edge Impulse Studio, package the Edge Impulse Runner as a Greengrass custom component, and deploy it to your entire fleet of Arm edge devices from the AWS console. Inference results and model metrics then stream back to AWS IoT Core in real time. + +## Why use this integration? + +This approach solves several real-world challenges for edge AI deployments: + +- **Scalable deployment**: Push ML model updates to hundreds or thousands of devices through Greengrass deployments, rather than manually updating each device. +- **Centralized monitoring**: Stream inference results and model performance metrics to AWS IoT Core, where you can route them to dashboards, databases, or alerting systems. +- **Remote management**: Issue commands to the Edge Impulse Runner service through IoT Core MQTT topics. You can restart inference, adjust confidence thresholds, or retrieve model information without SSH access to the device. +- **Secure credential handling**: Store the Edge Impulse API key in AWS Secrets Manager rather than passing it on the command line. + +## Example applications + +This integration is well suited for scenarios where ML inference runs on edge hardware, but results need to flow back to the cloud for action or analysis: + +- **Smart building occupancy**: Deploy a person-detection model to cameras at building entrances. Inference results stream to IoT Core, where they feed occupancy dashboards or trigger HVAC adjustments. +- **Wildlife monitoring**: Run an animal classification model on remote camera traps. Results are published to IoT Core and stored in S3 for conservation researchers. +- **Manufacturing quality inspection**: Detect defects on a production line using an object detection model. Inference confidence metrics are monitored in IoT Core to flag when model accuracy degrades and retraining is needed. +- **Retail analytics**: Count and classify products on shelves using edge cameras. Results are aggregated in the cloud for inventory management. + +In each case, Edge Impulse handles the ML model, Greengrass handles deployment and lifecycle management, and IoT Core handles the data pipeline back to AWS. + +## Architecture overview The Edge Impulse integration with AWS IoT Core and AWS IoT Greengrass is structured as follows: -![Architecture](images/Architecture.png) +![Architecture diagram showing the Edge Impulse Runner on an Arm edge device publishing inference results and model metrics to AWS IoT Core through the Greengrass integration#center](images/Architecture.png "Edge Impulse and AWS IoT Greengrass architecture") + +The key elements of this architecture are: + +- The Edge Impulse Runner service includes a `--greengrass` option that activates the AWS IoT integration. +- AWS Secrets Manager protects the Edge Impulse API key by removing it from command line arguments. +- Inference results are relayed to IoT Core for cloud-side processing, storage, or alerting. +- Model performance metrics (mean confidence, standard deviation) are published at configurable intervals. +- A bi-directional command interface lets you configure and manage the Runner service remotely through IoT Core MQTT topics. + +For more detail on the Runner service itself, see the [Edge Impulse for Linux Node.js SDK documentation](https://docs.edgeimpulse.com/docs/tools/edge-impulse-for-linux/linux-node-js-sdk). + +Edge Impulse provides pre-built Greengrass component recipes and artifacts in the [AWS Greengrass components repository](https://github.com/edgeimpulse/aws-greengrass-components). + +## The EdgeImpulseLinuxRunnerServiceComponent + +This Greengrass component downloads, installs, and runs the Edge Impulse Runner service on your edge device. Once running, it connects to your Edge Impulse project, pulls down the deployed model, and starts inference. + +The Runner uses MQTT topics to communicate with AWS IoT Core. Think of each topic as a named channel: the Runner publishes messages to a topic, and any service subscribed to that topic (a cloud dashboard, an AWS Lambda function, or the MQTT test client in the AWS console) receives those messages automatically. The `` placeholder in each topic is replaced with your device's actual name at runtime. + +The Runner publishes inference results (classification labels, bounding boxes, confidence scores) each time the model processes a frame: + +```text +/edgeimpulse/device//inference/output +``` -* The Edge Impulse "Runner" service now has a "--greengrass" option that enables the integration. -* AWS Secrets Manager is used to protect the Edge Impulse API Key by removing it from view via command line arguments. -* The Edge Impulse "Runner" service can relay inference results into IoT Core for further processing in the cloud -* The Edge Impulse "Runner" service relays model performance metrics, at configurable intervals, into IoTCore for further processing. -* The Edge Impulse "Runner" service has accessible commands that can be used to configure the service real-time as well as retrieve information about the model/service/configuration. -* More information regarding the Edge Impulse "Runner" service itself can be found in the [Edge Impulse for Linux Node.js SDK documentation](https://docs.edgeimpulse.com/docs/tools/edge-impulse-for-linux/linux-node-js-sdk). +Accumulated model performance statistics (mean confidence, standard deviation) are published on a timer you configure through the component settings: -Edge Impulse has several custom Greengrass components that can be deployed and run on the Greengrass-enabled edge device to enable this integration. The component recipes and artifacts can be found in the [AWS Greengrass components repository](https://github.com/edgeimpulse/aws-greengrass-components). Lets examine one of those components that we'll used for this workshop! +```text +/edgeimpulse/device//model/metrics +``` -### The "EdgeImpulseLinuxRunnerServiceComponent" Greengrass Component +You can send commands to the Runner by publishing a JSON message to the command input topic. Use this to restart inference, adjust confidence thresholds, or query model information without direct SSH access: -The Edge Impulse "Runner" service downloads, configures, installs, and executes an Edge Impulse model, developed for the specific edge device, and provides the ability to retrieve model inference results. In this case, our component for this service will relay the inference results into AWS IoT Core under the following topic: +```text +/edgeimpulse/device//command/input +``` - /edgeimpulse/device//inference/output - -Additionally, model performance metrics will be published, at defined intervals, here: +The Runner publishes the result of each command back to a separate output topic, so you can confirm the command was received and see the response: - /edgeimpulse/device//model/metrics - -Lastly, the Edge Impulse "Runner" service has been upgrade to support a set of bi-directional commands that can be accessed via publication of specific JSON structures to the following topic: +```text +/edgeimpulse/device//command/output +``` - /edgeimpulse/device//command/input - -Command results are published to the following topic: +The full command reference, including JSON structure details, is available in the [Edge Impulse AWS Greengrass integration documentation](https://docs.edgeimpulse.com/docs/integrations/aws-greengrass#commands-january-2025-integration-enhancements). - /edgeimpulse/device//command/output +## What you'll do in this Learning Path -The command reference, including JSON structure details, can be found in the [Edge Impulse AWS Greengrass integration documentation](https://docs.edgeimpulse.com/docs/integrations/aws-greengrass#commands-january-2025-integration-enhancements). +In this Learning Path, you: -Lets dive deeper into this integration starting with setting up our own edge device! +1. Set up a supported Arm-based edge device (or an EC2 Arm instance as a simulated device). +2. Create an Edge Impulse project with a pre-trained object detection model (cat and dog detector). +3. Install AWS IoT Greengrass on the edge device. +4. Store the Edge Impulse API key securely in AWS Secrets Manager. +5. Create and deploy a Greengrass custom component that runs the Edge Impulse Runner. +6. Verify inference results streaming to AWS IoT Core through the MQTT test client. +7. Issue remote commands to the Runner service through IoT Core. -Lets go! +The next section walks you through setting up your edge device hardware. diff --git a/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/running.md b/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/running.md index f96c63338b..3316c505e4 100644 --- a/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/running.md +++ b/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/running.md @@ -1,117 +1,136 @@ --- -title: 7. Running +title: Verify inference and view results weight: 9 ### FIXED, DO NOT MODIFY layout: learningpathall --- -## Running +## View live inference in the browser -Now that we have our Edge Impulse component(s) deployed to our edge device, how do we confirm things are working? +After the deployment completes and the Runner starts, it hosts a local web interface on port 4912. This page shows the live video input (from a camera or a file) alongside real-time inference results and timing information. -Simple! +Open a browser and navigate to: -On your browser, open the following url: +```text +http://:4912 +``` - http://:4912 - -So, for example, if my public ip address of my edge device is "1.1.1.1", my url would be: +Replace `` with the public IP address (for EC2) or local IP address of your edge device. For example, if your device's IP address is `1.1.1.1`: - http://1.1.1.1:4912 - -You should now see both the input (video either from file or from your edge devices attached camera) as well as inference results and inference times. There are two output scenarios depending on whether your edge device has a camera or does not have a camera... read below! +```text +http://1.1.1.1:4912 +``` -### Option 1: Edge devices with cameras +### Edge devices with cameras -You should be able to see live video of your camera via the url above. Now, point your camera at this picture: +If your device has a camera attached, you should see live video in the browser. Point the camera at this test image to verify the model is detecting objects: -![CatsNDogs](./images/DogsAndCats.png) +![Test image showing a dog and a cat side by side for model verification#center](./images/dogsandcats.png "Test image with a dog and a cat") -You should see that your model, running on the edge, is identifying both the dog and the cat! It should look something like this: +The model, running on your edge device, identifies both the dog and the cat. The browser output should look similar to this: -![CatsNDogs](./images/DogsAndCats_expected.png) +![Browser view showing live camera feed with bounding boxes identifying a dog and a cat along with inference timing#center](./images/dogsandcats_expected.png "Expected inference results with camera") -### Option 2: Edge devices without cameras +### Edge devices without cameras -In this case, you don't have a camera to use but your component is actually configured to pull its image data from local files installed by the optional non-camera component. +If your device doesn't have a camera, the component plays a pre-installed 90-second video of a cat. The browser should display something similar to this: -In this instance, a video of a cat will be shown. The url above should be displaying something similar to this: +![Browser view showing inference results from a video file with a cat detected#center](./images/cats_expected.png "Expected inference results without camera") -![CatsNDogs](./images//Cats_expected.png) +If the image appears frozen, the Runner has finished playing the video. The Runner is waiting for a `restart` command to replay the video file. The section below explains how to send this command through AWS IoT Core. -Now, if yours looks to be frozen... don't worry! It simply means that the "Runner" has completed playing the 90 second cat video. The "Runner" service is now waiting for you to issue a "restart" command to replay the same video... please continue reading below... we'll outline how to dispatch the "restart" command via AWS IoTCore! +## View inference output in AWS IoT Core -### AWS IoTCore Integration +The Runner publishes inference results and model metrics to AWS IoT Core MQTT topics. You can view these messages in the AWS Console. -With our installed components, we can also examine the ML inference output in AWS IoTCore. +Open the AWS Console and navigate to **AWS IoT Core**. Select **MQTT test client** from the left sidebar. -From the AWS Dashboard, bring up the IoTCore dashboard. Select the "MQTT Test Client" from the left hand side: +In the **Subscribe to a topic** section, enter the following topic filter and select **Subscribe**: -In the "Subscribe to a topic" section, enter this and press "subscribe": +```text +/edgeimpulse/device/# +``` - /edgeimpulse/device/# - -For those edge devices WITH cameras, you should see output on the left whenever your model identifies a cat and/or dog. The output format should look something like this: +For devices with cameras, inference results appear whenever the model identifies an object. The output looks similar to this: -![Inference Output](./images/EI_Inference_output.png) +![MQTT test client showing JSON inference output with bounding box coordinates, labels, and confidence scores#center](./images/ei_inference_output.png "Inference output in MQTT test client") -Additionally, you will see, model metrics being published periodically: +Model metrics are published periodically (controlled by the `metrics_sleeptime_ms` configuration field): -![Model Metrics](./images/EI_Model_Metrics.png) +![MQTT test client showing model metrics including inference time and resource usage#center](./images/ei_model_metrics.png "Model metrics in MQTT test client") -#### Issuing a command and examining the command result +## Send commands through AWS IoT Core -The integration provides a set of commands (see the [Summary](8_Summary.md) for details on the commands). One command, in particular, restarts the Edge Impulse "Runner" service. +The Edge Impulse Greengrass component supports commands sent through MQTT topics. One common command is `restart`, which restarts the Runner service. This is especially useful for devices without cameras, where the Runner pauses after the video ends. -In order to use commands we have to know what our device is "named" in IoTCore. You can easily find this by looking the inference output in the "MQTT Test Client": the publication "topic" is shown for each inference result you see. The topic structure is as follows: +### Find your device name - /edgeimpulse/devices//inference/output - /edgeimpulse/devices//model/metrics - /edgeimpulse/devices//command/output - /edgeimpulse/devices//command/input - -You will want to copy and save off the "my_device_name" portion of the topics that YOU see in your "MQTT Test Client" dashboard's inference results. +To send a command, you need the device name that the Runner registered in IoT Core. Look at the inference output in the MQTT test client. Each message is published to a topic with this structure: -Once you have the device name, back on the "MQTT Test Client" dashboard, select the "Publish to a topic" tab and enter this topic (but with YOUR device name filled in): +```text +/edgeimpulse/devices//inference/output +``` - /edgeimpulse/devices//command/input +Copy the `` portion from the topic. You'll use it in the next step. -Clear out the message content window and add the following JSON: +The Runner uses four MQTT topics per device: - { - "cmd": "restart" - } - -Click on the "Additional configuration" button and enable the "Retain message on this topic" checkbox. +```text +/edgeimpulse/devices//inference/output +/edgeimpulse/devices//model/metrics +/edgeimpulse/devices//command/input +/edgeimpulse/devices//command/output +``` -Press the "Publish" button. +### Send the restart command -What you should see now is that on the topic +In the MQTT test client, select the **Publish to a topic** tab. Enter the following topic, replacing `` with your actual device name: - /edgeimpulse/devices//command/output +```text +/edgeimpulse/devices//command/input +``` -notifications that the runner service has been restarted. On your browser, navigate back to: +Clear the message body and enter the following JSON: - http://:4912 - -and you should see your inferencing resuming. You should also see more inference output in IoTCore on this topic: +```json +{ + "cmd": "restart" +} +``` - /edgeimpulse/devices//inference/output +Select **Additional configuration** and enable the **Retain message on this topic** checkbox. Then select **Publish**. ->**_NOTE:_** ->For those who have edge devices WITHOUT cameras, your runner will read is input image video and report inferences until the video ends. Once ended, the "Runner" will simply wait for you to issue the above "restart" command to replay the video file. The restart command will cause the Runner to restart and it will once again, play the video file. +After publishing, you should see a response on the command output topic: -Cool! Congratulations! You have completed this workshop!! +```text +/edgeimpulse/devices//command/output +``` -#### Supplemental notes -Below are a few additional notes regarding the component deployment, log files, launch times for some devices: +The response confirms that the Runner has restarted. Navigate back to `http://:4912` in your browser to confirm inference has resumed. You should also see new inference results appearing in the MQTT test client. ->**_NOTE:_** ->After the deployment is initiated, on the FIRST invocation of a given deployment, expect to wait several moments (upwards of 5-10 min in fact) while the component installs all of the necessary pre-requisites that the component requires... this can take some time so be patient. You can also log into the edge gateway, receiving the component, and examine log files found in /greengrass/v2/logs. There you will see each components' current log file (same name as the component itself... ie. EdgeImpulseLinuxServiceComponent.log...) were you can watch the installation and invocation as it happens... any errors you might suspect will be shown in those log files. +{{% notice Note %}} +For devices without cameras, the Runner reads the sample video file and reports inferences until the video ends. After that, the Runner waits for a `restart` command to replay the video. Sending the restart command causes the Runner to start the video from the beginning. +{{% /notice %}} ->**_NOTE:_** ->While the components are running, in addition to the /greengrass/v2/logs directory, each component has a runtime log in /tmp. The format of the log file is: "ei\_lockfile\_[linux | runner | serial]\_\.log. Users can "tail" that log file to watch the component while it is running. +## Troubleshooting ->**_NOTE:_** ->Additionally, for Jetson-based devices where the model has been compiled specifically for that platform, one can expect to have a 2-3 minute delay in the model being loaded into the GPU memory for the first time. Subsequent invocations will be very short. +If the Runner doesn't start or the browser page doesn't load, check the following: + +**First deployment takes time**: On the first deployment, the component installs all prerequisites (Node.js, libvips, Edge Impulse CLI). This can take 5–10 minutes. Monitor progress by tailing the component log on your device: + +```bash +sudo tail -f /greengrass/v2/logs/EdgeImpulseLinuxRunnerServiceComponent.log +``` + +**Runtime logs**: While the component is running, the Runner writes a separate log file in `/tmp`. The filename follows the pattern `ei_lockfile_runner_.log`. Tail this file to watch live inference activity: + +```bash +sudo tail -f /tmp/ei_lockfile_runner_*.log +``` + +**Jetson GPU model loading**: On Jetson devices where the model is compiled for GPU acceleration, expect a 2–3 minute delay the first time the model loads into GPU memory. Subsequent starts are much faster. + +## What you've accomplished + +In this section, you verified that the Edge Impulse Runner is running inference on your edge device, viewed results in the browser and AWS IoT Core MQTT topics, and sent a restart command through IoT Core. In the next section, you'll find a complete command and configuration reference. diff --git a/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/secretmanagersetup.md b/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/secretmanagersetup.md index 78cdaf0462..2d715f9148 100644 --- a/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/secretmanagersetup.md +++ b/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/secretmanagersetup.md @@ -1,26 +1,40 @@ --- -title: 4. Secrets Manager Configuration +title: Store your API key in AWS Secrets Manager weight: 6 ### FIXED, DO NOT MODIFY layout: learningpathall --- -## Gather and install an EdgeImpulse API Key into AWS Secrets Manager +## Why use AWS Secrets Manager? -First we have to create an API Key in Edge Impulse via the Studio. +The Edge Impulse Greengrass component needs your Edge Impulse API key to download and run your ML model. Rather than hard-coding the key in the component configuration, you store it in AWS Secrets Manager. The component retrieves the key securely at runtime, which keeps it out of configuration files and makes rotation straightforward. -Next, we will go into AWS Console -> Secrets Manager and press "Store a new secret". From there we will specify: +The component expects two specific values: +- A secret with the ID `EI_API_KEY` (this is the name you give the secret in Secrets Manager). +- A key-value pair inside that secret where the key is `ei_api_key` and the value is your actual API key. - 1. Select "Other type of secret" - 2. Enter "ei_api_key" as the key NAME for the secret (goes in the "Key" section) - 3. Enter our actual API Key (goes in the "Value" section) - 4. Press "Next" - 5. Enter "EI_API_KEY" for the "Secret Name" (actually, this is its Secret ID...) - 6. Press "Next" - 7. Press "Next" - 8. Press "Store" +These names must match the `ei_sm_secret_id` and `ei_sm_secret_name` fields in the component configuration JSON you saved during hardware setup. -![CreateSecret](./images/SM_Create_Secret.png) +## Create the secret -Next we will install the EdgeImpulse Custom Greengrass Component we'll be using. +In the previous section, you generated an API key in Edge Impulse Studio and saved it. Now store that key in AWS Secrets Manager. + +Open the AWS Console and navigate to **Secrets Manager**. Select **Store a new secret** and follow these steps: + +1. Select **Other type of secret** as the secret type. +2. In the **Key** field, enter `ei_api_key`. +3. In the **Value** field, paste the API key you copied from Edge Impulse Studio. +4. Select **Next**. +5. For **Secret name**, enter `EI_API_KEY`. +6. Select **Next**. +7. Leave the rotation settings at their defaults and select **Next**. +8. Review the configuration and select **Store**. + +![AWS Secrets Manager console showing the Store a new secret form with the key-value pair and secret name configured#center](./images/sm_create_secret.png "Store a new secret in AWS Secrets Manager") + +After the secret is stored, you can verify it by selecting **EI_API_KEY** in the Secrets Manager list and confirming the key-value pair is present. + +## What you've accomplished + +You've securely stored your Edge Impulse API key in AWS Secrets Manager. The Greengrass component retrieves this key at runtime to authenticate with your Edge Impulse project. In the next section, you configure the Edge Impulse custom Greengrass component. diff --git a/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/summary.md b/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/summary.md index e48415dbbb..4dd3cf25ca 100644 --- a/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/summary.md +++ b/content/learning-paths/embedded-and-microcontrollers/edge_impulse_greengrass/summary.md @@ -1,384 +1,386 @@ --- -title: 8. Summary/Conclusions +title: Command and metrics reference weight: 10 ### FIXED, DO NOT MODIFY layout: learningpathall --- -## Summary +## Overview -Congratulations! You have completed this workshop! Please select "Next" below to read a bit about cleaning up your AWS environment in order to minimize costs/etc (AWS workshop attendees: this will happen automatically for you) +This page is a reference for the MQTT commands and model metrics available in the Edge Impulse Greengrass integration. Use these commands to control the Runner service, manage the confidence threshold filter, retrieve model information, and manage the inference cache β€” all through AWS IoT Core MQTT topics. -### For More Information +Commands are sent as JSON messages to the device's command input topic and results are published to the command output topic: -Below is some detailed reference information regarding the Edge Impulse AWS IoT Integration +```text +/edgeimpulse/device//command/input +/edgeimpulse/device//command/output +``` -## Model Metrics +All commands use the following JSON structure: -Basic model metrics are now accumulated and published in the integration into IoT Core. The metrics will be published at specified intervals (per the "metrics\_sleeptime\_ms" component configuration parameter) to the following IoT Core topic: +```json +{ + "cmd": "", + "value": "" +} +``` - /edgeimpulse/device//model/metrics - -The metrics published are: +The `value` field is only required for commands that set a value. -* **accumulated mean**: running accumulation of the average confidences from the linux runner while running the current model -* **accumulated standard deviation**: running accumulation of the standard deviation from the linux runner while running the current model +## Model metrics -The format of the model metrics output is as follows: +The Runner accumulates and publishes model metrics to IoT Core at the interval specified by the `metrics_sleeptime_ms` configuration parameter. Metrics are published to: - { - "mean_confidence": 0.696142, - "standard_deviation": 0.095282, - "confidence_trend": "decr", - "details": { - "n": 5, - "sum_confidences": 3.480711, - "sum_confidences_squared": 2.468464 - }, - "ts": 1736016142920, - "id": "e4faa78b-2a09-40d1-adfd-8e5fc32feb11" - } +```text +/edgeimpulse/device//model/metrics +``` -## Command Reference +The published metrics include: -In the 2025 January integration update, the following commands are now available with the Edge Impulse Greengrass Linux Runner Greengrass integration. The following commands are dispatched via the integration's IoT Core Topic as a JSON: +- **mean_confidence**: Running average of inference confidence scores for the current model. +- **standard_deviation**: Running standard deviation of confidence scores. +- **confidence_trend**: Direction the confidence is trending (`incr` or `decr`). - /edgeimpulse/device//command/input - -Results from the command can be found using the following topic: +Example metrics output: - /edgeimpulse/device//command/output +```json +{ + "mean_confidence": 0.696142, + "standard_deviation": 0.095282, + "confidence_trend": "decr", + "details": { + "n": 5, + "sum_confidences": 3.480711, + "sum_confidences_squared": 2.468464 + }, + "ts": 1736016142920, + "id": "e4faa78b-2a09-40d1-adfd-8e5fc32feb11" +} +``` -Command JSON structure is defined as follows: +## Startup notification - { - "cmd": , - "value": - } +When the Runner starts or restarts, it publishes the following JSON to the command output topic: -The currently supported commands are described below: +```json +{ + "result": { + "status": "started", + "ts": 1736026956853, + "id": "5c4e627e-6e9d-4382-bba7-00c0129705c4" + } +} +``` -### Initial Invocation +You can use this message to detect service restarts and re-apply any runtime changes (for example, confidence filter settings) to the newly started Runner. -When the runner process is started/restarted, the following JSON will be published to the command output topic above: +## restart - { - "result": { - "status": "started", - "ts": 1736026956853, - "id": "5c4e627e-6e9d-4382-bba7-00c0129705c4" - } - } +Restarts the Edge Impulse Runner process. When used with the `ei_shutdown_behavior` option set to `wait_on_restart`, the Runner pauses after the model completes and waits for this command before restarting. -This JSON can be used to flag a new invocation of the runner service (or a restart of the runner service). If there are any previous runtime changes made (i.e. confidence filter settings for example... see below), those changes can be resent to the newly invoked runtime. +**Command:** -### Restart Command +```json +{ + "cmd": "restart" +} +``` -##### Command JSON: +## enable_threshold_filter - { - "cmd": "restart" - } - -##### Command Description: +Enables the confidence threshold filter. When enabled, only inference results that meet the threshold criteria are published to IoT Core. By default, the filter is disabled and all results are published. -This command directs the integration to "restart" the Edge Impulse linux runner process. In conjunction with the "ei\_shutdown\_behavior" option being set to "wait\_for\_restart", the linux runner process will continue operating after the model has completed its operation. The linux runner process will continue to process input commands and will restart the linux runner via dispatching this command. +**Command:** -### Enable Threshold Filter Command +```json +{ + "cmd": "enable_threshold_filter" +} +``` -##### Command JSON: +**Result:** - { - "cmd": "enable_threshold_filter" - } - -##### Command Description: +```json +{ + "result": { + "threshold_filter_config": { + "enabled": "yes", + "confidence_threshold": 0.7, + "threshold_criteria": "ge" + } + } +} +``` + +## disable_threshold_filter -This command directs the integration to enable the threshold filter. The filter will control which inferences will get published into IoT Core. By default the filter is disabled so that all inferences reported are sent into IoT Core. +Disables the confidence threshold filter. All inference results are published to IoT Core regardless of confidence score. -##### Command Result: +**Command:** -The command output will be published as follows and will include the filter config: +```json +{ + "cmd": "disable_threshold_filter" +} +``` - { - "result": { - "threshold_filter_config": { - "enabled": "yes", - "confidence_threshold": 0.7, - "threshold_criteria": "ge" - } - } - } +**Result:** -### Disable Threshold Filter Command +```json +{ + "result": { + "threshold_filter_config": { + "enabled": "no", + "confidence_threshold": 0.7, + "threshold_criteria": "ge" + } + } +} +``` + +## set_threshold_filter_criteria -##### Command JSON: +Sets the comparison operator for the confidence threshold filter. The available criteria are: - { - "cmd": "disable_threshold_filter" - } - -##### Command Description: +| Criteria | Description | +|---|---| +| `gt` | Publish if confidence is greater than the threshold | +| `ge` | Publish if confidence is greater than or equal to the threshold | +| `eq` | Publish if confidence is equal to the threshold | +| `le` | Publish if confidence is less than or equal to the threshold | +| `lt` | Publish if confidence is less than the threshold | -This command directs the integration to disable the threshold filter. +**Command:** -##### Command Result: +```json +{ + "cmd": "set_threshold_filter_criteria", + "value": "ge" +} +``` -The command output will be published as follows and will include the filter config: +**Result:** + +```json +{ + "result": { + "criteria": "gt" + } +} +``` - { - "result": { - "threshold_filter_config": { - "enabled": "no", - "confidence_threshold": 0.7, - "threshold_criteria": "ge" - } - } - } +## get_threshold_filter_criteria -### Set Threshold Filter Criteria Command +Retrieves the currently configured threshold filter criteria. -##### Command JSON: +**Command:** - { - "cmd": "set_threshold_filter_criteria", - "value": "ge" - } - -##### Command Description: - -This command directs the integration to set the threshold filter criteria. The available options for the criteria are: - -* **"gt"**: publish if inference confidence is "greater than"... -* **"ge"**: publish if inference confidence is "greater than or equal to"... -* **"eq"**: publish if inference confidence is "equal to"... -* **"le"**: publish if inference confidence is "less than or equal to"... -* **"gt"**: publish if inference confidence is "less than"... - -##### Command Result: - -The command output will be published as follows: - - { - "result": { - "criteria": "gt" - } - } - -### Get Threshold Filter Criteria Command - -##### Command JSON: - - { - "cmd": "get_threshold_filter_criteria" - } - -##### Command Description: - -This command directs the integration to get the threshold filter criteria. The currently set threshold criteria is published to the command output topic above. - -##### Command Result: - -The command output will be published as follows with the configured criteria: - - { - "result": { - "criteria": "gt" - } - } - - -### Set Threshold Filter Confidence Command - -##### Command JSON: - - { - "cmd": "set_threshold_filter_confidence", - "value": 0.756 - } - -##### Command Description: - -This command directs the integration to set the threshold filter confidence bar. The value set must be a value 0 < x <= 1.0 - -##### Command Result: - -The command output will be published as follows with the specified confidence bar: - - { - "result": { - "confidence_threshold": "0.756" - } - } - - -### Get Threshold Filter Confidence Command - -##### Command JSON: - - { - "cmd": "get_threshold_filter_confidence" - } - -##### Command Description: - -This command directs the integration to get the threshold filter confidence bar. The currently set threshold confidence value is published to the command output topic above. - -##### Command Result: - -The command output will be published as follows with the currently configured confidence bar: - - { - "result": { - "confidence_threshold": "0.756" - } - } - -### Get Threshold Filter Config Command - -##### Command JSON: - - { - "cmd": "get_threshold_filter_config" - } - -##### Command Description: - -This command directs the integration to retrieve the current threshold filter config. The currently set threshold filter config is published to the command output topic above. - -##### Command Result: - -The command output will be published as follows with the currently configured filter config: - - { - "result": { - "threshold_filter_config": { - "enabled": "no", - "confidence_threshold": "0.756", - "threshold_criteria": "gt" - } - } - } - -### Get Model Info Command - -##### Command JSON: - - { - "cmd": "get_model_info" - } - -##### Command Description: - -This command directs the integration to retrieve the currently running model information. The model information is published to the command output topic above. - -##### Command Result: - -The command output will be published as follows with the current model information: - - { - "result": { - "model_info": { - "model_name": "occupant_counter", - "model_version": "v25", - "model_params": { - "axis_count": 1, - "frequency": 0, - "has_anomaly": 0, - "image_channel_count": 3, - "image_input_frames": 1, - "image_input_height": 640, - "image_input_width": 640, - "image_resize_mode": "fit-longest", - "inferencing_engine": 6, - "input_features_count": 409600, - "interval_ms": 1, - "label_count": 1, - "labels": [ - "person" - ], - "model_type": "object_detection", - "sensor": 3, - "slice_size": 102400, - "threshold": 0.5, - "use_continuous_mode": false, - "sensorType": "camera" - } - } - } - } - -### Reset Model Metrics Command - -##### Command JSON: - - { - "cmd": "reset_metrics" - } - -##### Command Description: - -This command directs the integration to reset the model metrics counters. - -##### Command Result: - -The command output will be published as follows to indicate the metrics counters are reset: - - { - "result": { - "metrics_reset": "OK" - } - } - -### Clear Cache Command - -##### Command JSON: - - { - "cmd": "clear_cache" - } - -##### Command Description: - -This command directs the integration to clear the currently configured inference image cache. The entire cache will be cleared. This command is sensitive to the Greengrass component configuration (i.e. which inference caches are enabled/disabled). This command will clear ALL caches that are currently enabled in the component configuration. - -##### Command Result: - -The command output will be published as follows with the clear cache results: - - { - "result": { - "clear_cache": { - "local": "OK", - "s3": "OK" - } - } - } - -### Clear Specified File From Cache Command - -##### Command JSON: - - { - "cmd": "clear_cache_file" - "value": - } - -##### Command Description: - -This command directs the integration to clear the specified file (by its uuid) from within the inference cache. This command is sensitive to the Greengrass component configuration (i.e. which inference caches are enabled/disabled). This command will clear the specified file from ALL enabled caches per the component configuration. - -##### Command Result: - -The command output will be published as follows with the clear cache results for the specified UUID: - - { - "result": { - "clear_cache_file": { - "local": "OK", - "s3": "OK", - "uuid": "e4faa78b-2a09-40d1-adfd-8e5fc32feb11" - } - } - } \ No newline at end of file +```json +{ + "cmd": "get_threshold_filter_criteria" +} +``` + +**Result:** + +```json +{ + "result": { + "criteria": "gt" + } +} +``` + +## set_threshold_filter_confidence + +Sets the confidence threshold value. Inference results are filtered against this value using the configured criteria. The value must be between 0 and 100. + +**Command:** + +```json +{ + "cmd": "set_threshold_filter_confidence", + "value": 0.756 +} +``` + +**Result:** + +```json +{ + "result": { + "confidence_threshold": "0.756" + } +} +``` + +## get_threshold_filter_confidence + +Retrieves the currently configured confidence threshold value. + +**Command:** + +```json +{ + "cmd": "get_threshold_filter_confidence" +} +``` + +**Result:** + +```json +{ + "result": { + "confidence_threshold": "0.756" + } +} +``` + +## get_threshold_filter_config + +Retrieves the complete threshold filter configuration, including enabled state, confidence value, and criteria. + +**Command:** + +```json +{ + "cmd": "get_threshold_filter_config" +} +``` + +**Result:** + +```json +{ + "result": { + "threshold_filter_config": { + "enabled": "no", + "confidence_threshold": "0.756", + "threshold_criteria": "gt" + } + } +} +``` + +## get_model_info + +Retrieves information about the currently running model, including its name, version, input dimensions, labels, and detection type. + +**Command:** + +```json +{ + "cmd": "get_model_info" +} +``` + +**Result:** + +```json +{ + "result": { + "model_info": { + "model_name": "occupant_counter", + "model_version": "v25", + "model_params": { + "axis_count": 1, + "frequency": 0, + "has_anomaly": 0, + "image_channel_count": 3, + "image_input_frames": 1, + "image_input_height": 640, + "image_input_width": 640, + "image_resize_mode": "fit-longest", + "inferencing_engine": 6, + "input_features_count": 409600, + "interval_ms": 1, + "label_count": 1, + "labels": [ + "person" + ], + "model_type": "object_detection", + "sensor": 3, + "slice_size": 102400, + "threshold": 0.5, + "use_continuous_mode": false, + "sensorType": "camera" + } + } + } +} +``` + +## reset_metrics + +Resets the accumulated model metrics counters to zero. + +**Command:** + +```json +{ + "cmd": "reset_metrics" +} +``` + +**Result:** + +```json +{ + "result": { + "metrics_reset": "OK" + } +} +``` + +## clear_cache + +Clears all inference image caches. This command respects the component configuration β€” it clears all caches that are currently enabled (local file cache, S3 cache, or both). + +**Command:** + +```json +{ + "cmd": "clear_cache" +} +``` + +**Result:** + +```json +{ + "result": { + "clear_cache": { + "local": "OK", + "s3": "OK" + } + } +} +``` + +## clear_cache_file + +Removes a specific cached inference result by its UUID. Like `clear_cache`, this command clears the file from all enabled caches. + +**Command:** + +```json +{ + "cmd": "clear_cache_file", + "value": "" +} +``` + +**Result:** + +```json +{ + "result": { + "clear_cache_file": { + "local": "OK", + "s3": "OK", + "uuid": "e4faa78b-2a09-40d1-adfd-8e5fc32feb11" + } + } +} +``` \ No newline at end of file diff --git a/content/learning-paths/laptops-and-desktops/dgx_spark_isaac_robotics/1_introduction_isaac.md b/content/learning-paths/laptops-and-desktops/dgx_spark_isaac_robotics/1_introduction_isaac.md new file mode 100644 index 0000000000..89b55e8863 --- /dev/null +++ b/content/learning-paths/laptops-and-desktops/dgx_spark_isaac_robotics/1_introduction_isaac.md @@ -0,0 +1,116 @@ +--- +title: Explore Isaac Sim and Isaac Lab for robotic workflows on DGX Spark +weight: 2 + +### FIXED, DO NOT MODIFY +layout: learningpathall +--- + +## Overview + +In this Learning Path, you will build, configure, and run robotic simulation and [reinforcement learning (RL)](https://en.wikipedia.org/wiki/Reinforcement_learning) workflows using NVIDIA Isaac Sim and Isaac Lab on an Arm-based DGX Spark system. The NVIDIA DGX Spark is a personal AI supercomputer powered by the GB10 [Grace Blackwell](https://learn.arm.com/learning-paths/laptops-and-desktops/dgx_spark_llamacpp/1_gb10_introduction/) Superchip, combining an Arm-based Grace CPU with a Blackwell GPU in a compact desktop form factor. + +Isaac Sim and Isaac Lab are NVIDIA's core tools for robotics simulation and reinforcement learning. Together they provide an end-to-end pipeline: simulate robots in physically accurate environments, train control policies using reinforcement learning, and evaluate those policies before deploying them to real hardware. + +This section introduces both tools and explains why the DGX Spark platform is an ideal development environment for these workloads. + +## What is Isaac Sim? + +[Isaac Sim](https://docs.isaacsim.omniverse.nvidia.com/latest/index.html) is a robotics simulation platform built on NVIDIA Omniverse. It provides GPU-accelerated physics simulation to enable fast, realistic robot simulations that can run faster than real time. + +Key capabilities of Isaac Sim include: + +| **Capability** | **Description** | +|----------------|-----------------| +| Physics simulation | High-fidelity rigid body, articulation, and soft-body physics powered by NVIDIA PhysX | +| Sensor simulation | Simulated cameras, LiDAR, IMU, and contact sensors that generate realistic data streams | +| Photorealistic rendering | Ray-traced rendering for vision-based tasks, domain randomization, and synthetic data generation | +| Parallel environments | Run thousands of simulation instances simultaneously on a single GPU for massive data throughput | +| Python API | Full programmatic control of scenes, robots, and simulations through Python scripting | + +Isaac Sim enables you to create detailed virtual worlds where robots can learn, be tested, and be validated without the cost, time, or risk of physical experiments. + +## What is Isaac Lab? + +[Isaac Lab](https://isaac-sim.github.io/IsaacLab/main/index.html) is a reinforcement learning framework built on top of Isaac Sim. It provides pre-built RL environments, training scripts, and evaluation tools for common robotics tasks such as locomotion, manipulation, and navigation. + +Isaac Lab supports two task design workflows: + +| **Workflow** | **Description** | **Best for** | +|--------------|-----------------|--------------| +| Manager-Based | Modular design where observations, actions, rewards, and terminations are defined through separate manager classes | Structured environments with reusable components | +| Direct | A single class defines the entire environment logic, similar to traditional Gymnasium environments | Rapid prototyping and full control over environment logic | + +Isaac Lab includes out-of-the-box integration with multiple reinforcement learning libraries: + +| **RL Library** | **Supported Algorithms** | +|----------------|--------------------------| +| RSL-RL | PPO ([Proximal Policy Optimization](https://en.wikipedia.org/wiki/Proximal_policy_optimization)) | +| rl_games | PPO, LSTM, vision-based policies | +| skrl | PPO, IPPO, MAPPO, AMP (Adversarial Motion Priors) | +| Stable Baselines3 (sb3) | PPO | + +In this Learning Path you will use the **RSL-RL** library, which is a lightweight and efficient PPO implementation commonly used for locomotion tasks. + +## Why DGX Spark for robotic simulation? + +The NVIDIA DGX Spark combines the Grace CPU and Blackwell GPU through a unified memory architecture, making it uniquely suited for robotics simulation and training workloads. + +| **DGX Spark feature** | **Impact on robotics workflows** | +|------------------------|----------------------------------| +| Grace CPU (Arm Cortex-X925 / A725, 20 cores) | Manages environment orchestration, reward calculation, and sensor data preprocessing with high single-thread performance | +| Blackwell GPU (CUDA cores + 5th-gen Tensor Cores) | Accelerates physics simulation, parallel environment stepping, and neural network forward/backward passes | +| 128 GB unified memory (NVLink-C2C) | Eliminates CPU-GPU data transfer bottlenecks; simulation state and model weights share the same address space | +| NVLink-C2C (900 GB/s bidirectional) | Enables near-zero-latency communication between CPU-driven orchestration and GPU-driven simulation | +| Compact desktop form factor | Run data-center-class robotics workloads on your desk without remote cluster access | + +Traditional robotics development requires separate machines for simulation, training, and deployment. DGX Spark consolidates these into a single platform. The unified memory is especially valuable for Isaac Sim, where physics state, rendered sensor data, and RL training tensors all reside in GPU-accessible memory without explicit copies. + +## How Isaac Sim and Isaac Lab work together + +The following describes the typical workflow when using Isaac Sim and Isaac Lab together on DGX Spark: + +1. **Define the environment**: Isaac Lab provides pre-built environment configurations for common tasks (locomotion, manipulation, navigation). You can also create custom environments. +2. **Launch the simulation**: Isaac Sim initializes the physics engine, loads robot models (URDF/USD), and sets up the scene on the Blackwell GPU. +3. **Train a policy**: Isaac Lab's training scripts use RL algorithms (such as PPO via RSL-RL) to optimize a neural network policy. The GPU runs thousands of parallel environments simultaneously. +4. **Evaluate and iterate**: Trained policies can be tested in simulation with visualization enabled or exported for deployment to real hardware. + +The entire pipeline runs locally on DGX Spark. Headless mode (without visualization) maximizes GPU utilization for training, while visualization mode lets you inspect robot behavior interactively. + +## Available environment categories + +Isaac Lab ships with a comprehensive set of pre-built environments across several categories: + +| **Category** | **Examples** | **Description** | +|--------------|-------------|-----------------| +| Classic | Cartpole, Ant, Humanoid | MuJoCo-style control benchmarks for algorithm development | +| Manipulation | Reach, Lift, Stack, Open-Drawer | Fixed-arm tasks using Franka, UR10, and other robots | +| Contact-rich Manipulation | Peg insertion, Gear meshing, Nut threading | Precision assembly tasks with the Franka robot | +| Locomotion | Anymal B/C/D, Unitree A1/Go1/Go2/H1/G1, Spot, Digit | Velocity tracking on flat and rough terrain for quadrupeds and humanoids | +| Navigation | Anymal C navigation | Point-to-point navigation with heading control | +| Multi-agent | Cart-Double-Pendulum, Shadow-Hand-Over | Tasks that require coordination among multiple agents | + +You will be able to list all available environments after setting up Isaac Lab in the next section. The following command will be available once installation is complete: + +```bash +./isaaclab.sh -p scripts/environments/list_envs.py +``` + +You can also filter by keyword: + +```bash +./isaaclab.sh -p scripts/environments/list_envs.py --keyword locomotion +``` + +For the complete list of environments, see the [Isaac Lab Available Environments](https://isaac-sim.github.io/IsaacLab/main/source/overview/environments.html) documentation. + +## What you will accomplish in this Learning Path + +In the learning path that follow you will: + +1. **Set up Isaac Sim and Isaac Lab** on your DGX Spark by building both tools from source +2. **Run a basic robot simulation** in Isaac Sim and interact with it through Python +3. **Train a reinforcement learning policy** for the Unitree H1 humanoid robot on rough terrain using RSL-RL +4. **Explore advanced RL scenarios** including diverse locomotion tasks and robot configurations + +By the end, you will have a fully functional Isaac Sim and Isaac Lab development environment on DGX Spark and hands-on experience with the complete robotics RL pipeline. diff --git a/content/learning-paths/laptops-and-desktops/dgx_spark_isaac_robotics/2_isaac_installation.md b/content/learning-paths/laptops-and-desktops/dgx_spark_isaac_robotics/2_isaac_installation.md new file mode 100644 index 0000000000..5fdaef4a60 --- /dev/null +++ b/content/learning-paths/laptops-and-desktops/dgx_spark_isaac_robotics/2_isaac_installation.md @@ -0,0 +1,233 @@ +--- +title: Set up Isaac Sim and Isaac Lab on DGX Spark +weight: 3 + +### FIXED, DO NOT MODIFY +layout: learningpathall +--- + +## Set up your development environment + +Before running robotic simulations and reinforcement learning tasks, you need to build Isaac Sim and Isaac Lab from source on your DGX Spark system. This section walks you through verifying your system, installing dependencies, building Isaac Sim, and then setting up Isaac Lab on top of it. + +The build process takes approximately 15-20 minutes on the Grace CPU and requires around 50 GB of available disk space. + +## Step 1: Verify your system + +Start by confirming that your DGX Spark system has the required hardware and software configuration. + +Check the CPU architecture: + +```bash +lscpu | head -5 +``` + +The output is similar to: + +```output +Architecture: aarch64 + CPU op-mode(s): 64-bit + Byte Order: Little Endian +CPU(s): 20 + On-line CPU(s) list: 0-19 +``` + +Verify the Blackwell GPU is recognized: + +```bash +nvidia-smi +``` + +You will see output similar to: + +```output ++-----------------------------------------------------------------------------------------+ +| NVIDIA-SMI 580.95.05 Driver Version: 580.95.05 CUDA Version: 13.0 | ++-----------------------------------------+------------------------+----------------------+ +| GPU Name Persistence-M | Bus-Id Disp.A | Volatile Uncorr. ECC | +|=========================================+========================+======================| +| 0 NVIDIA GB10 On | 0000000F:01:00.0 Off | N/A | ++-----------------------------------------+------------------------+----------------------+ +``` + +Confirm the CUDA toolkit is installed: + +```bash +nvcc --version +``` + +The expected output includes: + +```output +Cuda compilation tools, release 13.0, V13.0.88 +``` + +{{% notice Note %}} +Isaac Sim requires GCC/G++ 11, Git LFS, and CUDA 13.0 or later. If any of these checks fail, resolve the issue before proceeding. +{{% /notice %}} + +## Step 2: Install GCC 11 and Git LFS + +Isaac Sim requires GCC/G++ version 11 for compilation. Install it and set it as the default compiler: + +```bash +sudo apt update && sudo apt install -y gcc-11 g++-11 +sudo update-alternatives --install /usr/bin/gcc gcc /usr/bin/gcc-11 200 +sudo update-alternatives --install /usr/bin/g++ g++ /usr/bin/g++-11 200 +``` + +Install Git LFS, which is needed to pull large binary assets from the Isaac Sim repository: + +```bash +sudo apt install -y git-lfs +``` + +Verify both installations: + +```bash +gcc --version +g++ --version +git lfs version +``` + +The GCC output should show version 11.x. Git LFS should report a version number confirming it is installed. + +## Step 3: Clone and build Isaac Sim + +Clone the Isaac Sim repository from GitHub. The `--depth=1` flag creates a shallow clone to reduce download time, and `--recursive` fetches all submodules: + +```bash +cd ~ +git clone --depth=1 --recursive https://github.com/isaac-sim/IsaacSim +cd IsaacSim +git lfs install +git lfs pull +``` + +{{% notice Note %}} +The Git LFS pull downloads several gigabytes of simulation assets (USD files, textures, and pre-built libraries). Ensure you have a stable network connection. +{{% /notice %}} + +Build Isaac Sim by running the build script. This compiles the simulation engine and all its components: + +```bash +./build.sh +``` + +The build uses all available CPU cores on the Grace processor. On DGX Spark, compilation typically takes 10-15 minutes. + +When the build succeeds, you will see output similar to: + +```output +BUILD (RELEASE) SUCCEEDED (Took 674.39 seconds) +``` + +## Step 4: Set Isaac Sim environment variables + +After the build completes, configure your shell to recognize the Isaac Sim installation. Run the following commands from inside the `IsaacSim` directory: + +```bash +export ISAACSIM_PATH="${PWD}/_build/linux-aarch64/release" +export ISAACSIM_PYTHON_EXE="${ISAACSIM_PATH}/python.sh" +``` + +The table below explains each variable: + +| **Variable** | **Purpose** | +|--------------|-------------| +| `ISAACSIM_PATH` | Points to the compiled Isaac Sim binaries and libraries under the `_build` directory | +| `ISAACSIM_PYTHON_EXE` | References the Python wrapper script that runs Python with Isaac Sim's dependencies preloaded | + +{{% notice Tip %}} +Add these `export` lines to your `~/.bashrc` file so they persist across terminal sessions: + +```bash +echo 'export ISAACSIM_PATH="$HOME/IsaacSim/_build/linux-aarch64/release"' >> ~/.bashrc +echo 'export ISAACSIM_PYTHON_EXE="${ISAACSIM_PATH}/python.sh"' >> ~/.bashrc +source ~/.bashrc +``` +{{% /notice %}} + +## Step 5: Validate the Isaac Sim build + +Launch Isaac Sim to verify the build was successful. The `LD_PRELOAD` setting resolves a library compatibility issue on aarch64: + +```bash +export LD_PRELOAD="$LD_PRELOAD:/lib/aarch64-linux-gnu/libgomp.so.1" +${ISAACSIM_PATH}/isaac-sim.sh +``` + +If the build is correct, Isaac Sim opens its viewer window (or starts in headless mode if no display is available). You should see initialization messages confirming that the Blackwell GPU is detected and the physics engine is ready. + +Press `Ctrl+C` in the terminal to close Isaac Sim after verifying it starts successfully. + +## Step 6: Clone and install Isaac Lab + +With Isaac Sim successfully built and validated, you can now set up Isaac Lab to enable RL training workflows. Clone the repository into your home directory: + +```bash +cd ~ +git clone --recursive https://github.com/isaac-sim/IsaacLab +cd IsaacLab +``` + +Create a symbolic link so Isaac Lab can find your Isaac Sim installation: + +```bash +echo "ISAACSIM_PATH=$ISAACSIM_PATH" +ln -sfn "${ISAACSIM_PATH}" "${PWD}/_isaac_sim" +``` + +Verify the symbolic link is correct: + +```bash +ls -l "${PWD}/_isaac_sim/python.sh" +``` + +You should see the symlink pointing to your Isaac Sim build directory. + +Install Isaac Lab and all its dependencies: + +```bash +./isaaclab.sh --install +``` + +This command installs the Isaac Lab Python packages, RL libraries (RSL-RL, rl_games, skrl, Stable Baselines3), and additional dependencies into the Isaac Sim Python environment. + +## Step 7: Validate the Isaac Lab installation + +Verify that Isaac Lab is installed correctly by listing the available RL environments: + +```bash +export LD_PRELOAD="$LD_PRELOAD:/lib/aarch64-linux-gnu/libgomp.so.1" +./isaaclab.sh -p scripts/environments/list_envs.py +``` + +You should see a list of available environments, including entries such as: + +```output +Isaac-Cartpole-v0 +Isaac-Cartpole-Direct-v0 +Isaac-Velocity-Flat-H1-v0 +Isaac-Velocity-Rough-H1-v0 +Isaac-Lift-Cube-Franka-v0 +Isaac-Reach-Franka-v0 +... +``` + +If the environment list displays without errors, both Isaac Sim and Isaac Lab are correctly installed and ready for use. + +You are now ready to run and train RL tasks using Isaac Lab environments. + +## What you have accomplished + +In this section you have: + +- Verified your DGX Spark system has the required Grace CPU, Blackwell GPU, and CUDA 13 environment +- Installed GCC 11 and Git LFS as build prerequisites +- Cloned and built Isaac Sim from source, producing aarch64-optimized binaries for the Grace-Blackwell platform +- Configured environment variables so Isaac Lab can locate the Isaac Sim installation +- Cloned and installed Isaac Lab with all RL library dependencies +- Validated both installations by launching Isaac Sim and listing available environments + +Your development environment is now fully configured for robot simulation and RL workflows. In the next module, you will run your first robot simulation and begin interacting with Isaac Sim through Python scripts. diff --git a/content/learning-paths/laptops-and-desktops/dgx_spark_isaac_robotics/32_cartpole.gif b/content/learning-paths/laptops-and-desktops/dgx_spark_isaac_robotics/32_cartpole.gif new file mode 100644 index 0000000000..ea873b9f68 Binary files /dev/null and b/content/learning-paths/laptops-and-desktops/dgx_spark_isaac_robotics/32_cartpole.gif differ diff --git a/content/learning-paths/laptops-and-desktops/dgx_spark_isaac_robotics/3_isaac_small_project.md b/content/learning-paths/laptops-and-desktops/dgx_spark_isaac_robotics/3_isaac_small_project.md new file mode 100644 index 0000000000..28511d9657 --- /dev/null +++ b/content/learning-paths/laptops-and-desktops/dgx_spark_isaac_robotics/3_isaac_small_project.md @@ -0,0 +1,226 @@ +--- +title: Run and Understand a Sample Robot Simulation with Isaac Sim +weight: 4 + +### FIXED, DO NOT MODIFY +layout: learningpathall +--- + +## Deploy a basic robot simulation + +With Isaac Sim and Isaac Lab installed, you can now run your first robot simulation. In this section you will launch a pre-built simulation scene, interact with it programmatically, and understand the key concepts behind Isaac Sim's simulation loop. + +You will work with the Cartpole environment, a classic control benchmark where a cart must balance a pole by applying horizontal forces. This environment is simple enough to understand quickly but demonstrates all the core simulation concepts required for more complex robotics tasks. + +## Step 1: Launch a sample scene from Isaac Lab + +Isaac Lab provides tutorial scripts that demonstrate how to create and interact with simulation scenes. Start with a minimal scene to verify Isaac Sim’s rendering and simulation setup: + +```bash +cd ~/IsaacLab +export LD_PRELOAD="$LD_PRELOAD:/lib/aarch64-linux-gnu/libgomp.so.1" +./isaaclab.sh -p scripts/tutorials/00_sim/create_empty.py +``` + +This script creates an empty simulation world with a ground plane and default lighting. It validates that the Isaac Sim rendering and physics engines are working on your DGX Spark system. + +If a display is connected, a viewer window should open; otherwise, log messages will confirm that the simulation initialized successfully in headless mode. + +Press `Ctrl+C` to exit the simulation. + +## Step 2: Spawn and simulate a robot + +Next, run a more complete example that spawns articulated robots into the scene. This tutorial demonstrates how Isaac Sim handles multi-body physics: + +```bash +./isaaclab.sh -p scripts/tutorials/01_assets/run_articulation.py +``` + +This script loads a robot model, advances the physics simulation, and prints joint states to the terminal. It demonstrates: + +- Loading a robot from a USD (Universal Scene Description) asset file +- Configuring joint actuators and control modes +- Stepping the physics simulation and reading back joint positions and velocities + +![img1 alt-text#center](run_articulation.gif "Figure 1: run_articulation.py") + +## Step 3: Run the Cartpole environment + +Now run a complete environment that combines scene, action, observation, and event managers. The `create_cartpole_base_env.py` tutorial creates a Cartpole base environment and applies random actions: + +```bash +./isaaclab.sh -p scripts/tutorials/03_envs/create_cartpole_base_env.py --num_envs 32 +``` + +This command launches 32 parallel Cartpole environments on the Blackwell GPU. Each environment runs its own independent simulation with random joint efforts applied to the cart. You will see the pole joint angle printed to the terminal for each step. + +![img2 alt-text#center](32_cartpole.gif "Figure 2: 32 parallel Cartpole") + +{{% notice Note %}} +This tutorial script uses a hardcoded `CartpoleEnvCfg` configuration. It does not accept a `--task` argument. The `--num_envs` flag controls how many parallel environments are spawned on the GPU. +{{% /notice %}} + +## Step 4: Run the Cartpole RL environment + +The previous script creates a base environment without rewards or terminations. To see the full RL environment (with reward computation and episode resets), run: + +```bash +./isaaclab.sh -p scripts/tutorials/03_envs/run_cartpole_rl_env.py --num_envs 32 +``` + +This script wraps the Cartpole scene in a `ManagerBasedRLEnv`, which includes reward computation, termination conditions, and the standard Gymnasium `step()` interface that returns `(obs, reward, terminated, truncated, info)`. + +The key difference between the two scripts: + +| **Script** | **Environment type** | **Returns from step()** | +|-----------|---------------------|------------------------| +| `create_cartpole_base_env.py` | `ManagerBasedEnv` | `(obs, info)` β€” no rewards or terminations | +| `run_cartpole_rl_env.py` | `ManagerBasedRLEnv` | `(obs, reward, terminated, truncated, info)` β€” full RL interface | + +## Step 5: Understand the simulation code + +To understand what happens inside an Isaac Lab environment, examine the Cartpole environment source code. The key elements are: + +### Environment configuration + +Every Isaac Lab environment starts with a configuration class that defines the simulation parameters. The `CartpoleEnvCfg` in the tutorial specifies: + +```python +@configclass +class CartpoleEnvCfg(ManagerBasedEnvCfg): + """Configuration for the cartpole environment.""" + + # Scene settings + scene = CartpoleSceneCfg(num_envs=1024, env_spacing=2.5) + # Basic settings + observations = ObservationsCfg() + actions = ActionsCfg() + events = EventCfg() + + def __post_init__(self): + """Post initialization.""" + self.decimation = 4 # env step every 4 sim steps: 200Hz / 4 = 50Hz + self.sim.dt = 0.005 # sim step every 5ms: 200Hz +``` + +The table below explains each parameter: + +| **Parameter** | **Value** | **Description** | +|---------------|-----------|-----------------| +| `scene.num_envs` | 1024 | Default number of parallel environment instances (overridden by `--num_envs` from the command line) | +| `scene.env_spacing` | 2.5 | Distance in meters between each parallel environment in the scene | +| `decimation` | 4 | The policy acts every 4 physics steps. With a 200 Hz physics rate, the policy runs at 50 Hz | +| `sim.dt` | 0.005 | The physics engine advances by 5 ms per step (200 Hz simulation rate) | + +### Actions, observations, and events + +The configuration defines three manager groups: + +**Actions** β€” how the agent controls the robot: + +```python +@configclass +class ActionsCfg: + joint_efforts = mdp.JointEffortActionCfg( + asset_name="robot", + joint_names=["slider_to_cart"], + scale=5.0 # Multiplier on raw action values + ) +``` + +The agent produces a single continuous value that is scaled by `5.0` and applied as a force on the cart's slider joint. + +**Observations** β€” what the agent sees: + +```python +@configclass +class ObservationsCfg: + @configclass + class PolicyCfg(ObsGroup): + joint_pos_rel = ObsTerm(func=mdp.joint_pos_rel) # Relative joint positions + joint_vel_rel = ObsTerm(func=mdp.joint_vel_rel) # Relative joint velocities +``` + +The agent observes joint positions and velocities for both the cart slider and the pole hinge. + +**Events** β€” randomization applied during simulation: + +```python +@configclass +class EventCfg: + # On startup: randomize the pole mass (adds 0.1 to 0.5 kg) + add_pole_mass = EventTerm( + func=mdp.randomize_rigid_body_mass, + mode="startup", + params={"mass_distribution_params": (0.1, 0.5), "operation": "add"}, + ) + # On reset: randomize cart and pole starting positions + reset_cart_position = EventTerm(func=mdp.reset_joints_by_offset, mode="reset", ...) + reset_pole_position = EventTerm(func=mdp.reset_joints_by_offset, mode="reset", ...) +``` + +Events introduce variability that improves training robustness. Randomizing the pole mass on startup means the agent must learn to balance poles of different weights. Randomizing joint positions on reset ensures each episode starts from a different state. + +### The simulation loop + +The core simulation loop in Isaac Lab follows a standard Gymnasium-style interface. From `run_cartpole_rl_env.py`: + +```python +# Create the RL environment +env = ManagerBasedRLEnv(cfg=env_cfg) + +count = 0 +while simulation_app.is_running(): + with torch.inference_mode(): + # Reset every 300 steps + if count % 300 == 0: + count = 0 + env.reset() + + # Sample random actions + joint_efforts = torch.randn_like(env.action_manager.action) + + # Step the environment: apply action, advance physics, compute reward + obs, rew, terminated, truncated, info = env.step(joint_efforts) + + # Print the pole joint angle for environment 0 + print("[Env 0]: Pole joint: ", obs["policy"][0][1].item()) + count += 1 +``` + +Each call to `env.step(action)` performs these operations on the GPU: + +1. **Apply actions**: The action tensor is scaled and applied as joint efforts to the cart +2. **Step physics**: The simulation advances by `decimation` physics steps (4 steps at 200 Hz = 20 ms of simulated time) +3. **Compute observations**: Joint positions and velocities are read from the simulation +4. **Compute rewards**: A reward function evaluates how well the agent balanced the pole +5. **Check terminations**: The environment checks if the episode should end (for example, the pole angle exceeded a threshold) + +All computations happen in parallel across all environments using PyTorch tensors on the GPU. This is what makes Isaac Lab efficient: thousands of environments run in parallel without Python loop overhead. + +## Step 6: Run with headless mode + +For reinforcement learning tasks, headless mode is preferred to maximize GPU throughput. You can test it now using the Cartpole RL environment. + +```bash +./isaaclab.sh -p scripts/tutorials/03_envs/run_cartpole_rl_env.py --num_envs 64 --headless +``` + +In headless mode, all GPU resources are dedicated to physics simulation and tensor computation. This is the recommended mode for training reinforcement learning policies, which you will do in the next section. + +{{% notice Note %}} +When running headless on DGX Spark, the Blackwell GPU handles both the physics simulation and neural network computation. The unified memory architecture means there is no performance penalty for sharing GPU memory between these workloads. +{{% /notice %}} + +## What you have accomplished + +In this section you have: + +- Launched your first Isaac Sim scene on DGX Spark and verified the rendering and physics engines work correctly +- Spawned articulated robots and observed multi-body physics simulation +- Run the Cartpole base environment and RL environment with 32 parallel instances on the Blackwell GPU +- Understood the key components of an Isaac Lab environment: configuration, actions, observations, events, simulation loop, and reward computation +- Tested headless mode for maximum training performance + +You now understand the core components of an Isaac Lab simulation environment, including scene creation, robot articulation, observation and action structures, and simulation loop execution. +In the next section, you will use these concepts to train a reinforcement learning policy for a humanoid robot. diff --git a/content/learning-paths/laptops-and-desktops/dgx_spark_isaac_robotics/4_isaac_rfl.md b/content/learning-paths/laptops-and-desktops/dgx_spark_isaac_robotics/4_isaac_rfl.md new file mode 100644 index 0000000000..8746a0a4fe --- /dev/null +++ b/content/learning-paths/laptops-and-desktops/dgx_spark_isaac_robotics/4_isaac_rfl.md @@ -0,0 +1,286 @@ +--- +title: Train a Humanoid Locomotion Policy with Isaac Lab on DGX Spark +weight: 5 + +### FIXED, DO NOT MODIFY +layout: learningpathall +--- + +## Train a reinforcement learning policy using Isaac Lab and RSL-RL + +In this section you will train a reinforcement learning (RL) policy for the [Unitree] (https://www.unitree.com/) H1 humanoid robot to walk over rough terrain. You will use Isaac Lab's RSL-RL integration, which implements the Proximal Policy Optimization (PPO) algorithm. By the end of this section you will understand the full training pipeline, including task configuration, PPO hyperparameters, and policy evaluation. + +## What is RSL-RL? + +RSL-RL (Robotic Systems Lab Reinforcement Learning) is a lightweight RL library developed at [ETH Zurich](https://ethz.ch/en.html) specifically for locomotion tasks. It implements PPO with features tailored to robotics: + +- GPU-accelerated rollout collection across thousands of parallel environments +- Efficient on-policy training with generalized advantage estimation (GAE) +- Asymmetric actor-critic support (the critic can observe more than the actor) +- Minimal dependencies and tight integration with Isaac Lab + +Isaac Lab provides ready-to-use training scripts for RSL-RL under `scripts/reinforcement_learning/rsl_rl/`. + +## Step 1: Understand the training task + +The task you will train is **Isaac-Velocity-Rough-H1-v0**. This is a locomotion task where the [Unitree H1](https://www.unitree.com/h1/) humanoid robot must track a velocity command while navigating rough terrain. + +The task details are: + +| **Property** | **Value** | +|-------------|-----------| +| Environment ID | `Isaac-Velocity-Rough-H1-v0` | +| Robot | Unitree H1 (19 actuated joints, bipedal humanoid) | +| Terrain | Procedurally generated rough terrain with slopes, stairs, and obstacles | +| Workflow | Manager-Based | +| Objective | Track a commanded forward velocity, lateral velocity, and yaw rate | +| Observation space | Joint positions, joint velocities, gravity projection, velocity commands, and previous actions | +| Action space | Target joint positions for all actuated joints | +| RL library | RSL-RL (PPO) | + +The robot receives a velocity command (for example, "walk forward at 1.0 m/s") and must learn to coordinate all 19 joints to achieve that velocity while maintaining balance on uneven ground. + +This setup provides a high-dimensional control problem ideal for testing locomotion learning under challenging terrain. + +## Step 2: Launch the training + +Navigate to the Isaac Lab directory and start training in headless mode for maximum performance: + +```bash +cd ~/IsaacLab +export LD_PRELOAD="$LD_PRELOAD:/lib/aarch64-linux-gnu/libgomp.so.1" +./isaaclab.sh -p scripts/reinforcement_learning/rsl_rl/train.py \ + --task=Isaac-Velocity-Rough-H1-v0 \ + --headless +``` + +Once the training starts, you will see log messages reporting iteration progress, rewards, and performance statistics. + +``` + Learning iteration 15/3000 + + Computation: 65955 steps/s (collection: 1.256s, learning 0.235s) + Mean action noise std: 1.04 + Mean value_function loss: 0.0911 + Mean surrogate loss: 0.0003 + Mean entropy loss: 27.6371 + Mean reward: -5.35 + Mean episode length: 61.42 +Episode_Reward/track_lin_vel_xy_exp: 0.0179 +Episode_Reward/track_ang_vel_z_exp: 0.0058 + Episode_Reward/ang_vel_xy_l2: -0.0220 + Episode_Reward/dof_torques_l2: 0.0000 + Episode_Reward/dof_acc_l2: -0.0081 + Episode_Reward/action_rate_l2: -0.0127 + Episode_Reward/feet_air_time: 0.0002 +Episode_Reward/flat_orientation_l2: -0.0170 + Episode_Reward/dof_pos_limits: -0.0012 +Episode_Reward/termination_penalty: -0.2000 + Episode_Reward/feet_slide: -0.0119 +Episode_Reward/joint_deviation_hip: -0.0083 +Episode_Reward/joint_deviation_arms: -0.0079 +Episode_Reward/joint_deviation_torso: -0.0013 + Curriculum/terrain_levels: 0.1577 +Metrics/base_velocity/error_vel_xy: 0.1221 +Metrics/base_velocity/error_vel_yaw: 0.4705 + Episode_Termination/time_out: 0.0000 + Episode_Termination/base_contact: 1.0000 +-------------------------------------------------------------------------------- + Total timesteps: 1572864 + Iteration time: 1.49s + Time elapsed: 00:00:26 + ETA: 01:21:49 +``` + + +This command launches the training with default hyperparameters. The Blackwell GPU runs thousands of parallel H1 environments simultaneously while the Grace CPU handles logging and orchestration. + +{{% notice Warning %}} +**Known issue: NVRTC GPU architecture error on DGX Spark** + +When running RL training on the Blackwell GPU (GB10, compute capability 12.1), you may encounter: + +``` +RuntimeError: nvrtc: error: invalid value for --gpu-architecture (-arch) +``` + +This error occurs because the NVRTC runtime compiler inside PyTorch does not yet fully support the `sm_121` architecture. It is a known compatibility issue tracked in [Isaac Lab Discussion #2406](https://github.com/isaac-sim/IsaacLab/discussions/2406) and [PyTorch Issue #87595](https://github.com/pytorch/pytorch/issues/87595). + +**Workaround**: Make sure you are using the Isaac Sim build from source (as described in the setup section) rather than a pip-installed version. The source build includes the correct CUDA 13 runtime for Blackwell. If the error persists, try running with `--headless` mode, which avoids some NVRTC code paths used by the renderer. Also ensure your NVIDIA driver is up to date (`nvidia-smi` should show driver 580.x or later). + +This issue is expected to be resolved in future Isaac Sim and PyTorch releases with full Blackwell support. +{{% /notice %}} + +You can also override default parameters from the command line: + +```bash +./isaaclab.sh -p scripts/reinforcement_learning/rsl_rl/train.py \ + --task=Isaac-Velocity-Rough-H1-v0 \ + --headless \ + --num_envs=2048 \ + --max_iterations=1500 \ + --seed=42 +``` + +### Command-line arguments + +The following table explains the key command-line arguments: + +| **Argument** | **Default** | **Description** | +|-------------|-------------|-----------------| +| `--task` | (required) | The Isaac Lab environment ID. Use `Isaac-Velocity-Rough-H1-v0` for rough-terrain humanoid locomotion | +| `--headless` | off | Disables visualization for faster training. All GPU resources go to simulation and training | +| `--num_envs` | 4096 | Number of parallel environments. Each environment runs an independent simulation of the H1 robot | +| `--max_iterations` | 1500 | Total number of PPO training iterations. Each iteration collects a batch of experience and updates the policy | +| `--seed` | 0 | Random seed for reproducibility. Set this to get deterministic results across runs | + +{{% notice Tip %}} +On DGX Spark, 2048 to 4096 parallel environments work well for locomotion tasks. Higher values increase sample throughput but require more GPU memory. Start with 2048 if you want faster iteration cycles during development. +{{% /notice %}} + +## Step 3: Understand the PPO hyperparameters + +This section explains the core PPO training parameters used by RSL-RL and how they influence learning quality and stability. + +PPO (Proximal Policy Optimization) is the RL algorithm used by RSL-RL. Understanding each hyperparameter helps you tune training for different tasks. The following table describes the key hyperparameters and their roles: + +### Policy network hyperparameters + +| **Hyperparameter** | **Typical value** | **Description** | +|--------------------|-------------------|-----------------| +| `policy_class_name` | `ActorCritic` | The neural network architecture. `ActorCritic` uses separate networks for the policy (actor) and value function (critic) | +| `actor_hidden_dims` | `[512, 256, 128]` | Hidden layer sizes for the actor (policy) network. Larger networks can represent more complex behaviors but train more slowly | +| `critic_hidden_dims` | `[512, 256, 128]` | Hidden layer sizes for the critic (value) network. The critic estimates how good each state is | +| `activation` | `elu` | Activation function between hidden layers. ELU (Exponential Linear Unit) provides smooth gradients and avoids dead neurons | +| `init_noise_std` | `1.0` | Initial standard deviation of the exploration noise. Higher values encourage more exploration early in training | + +### PPO algorithm hyperparameters + +| **Hyperparameter** | **Typical value** | **Description** | +|--------------------|-------------------|-----------------| +| `num_learning_epochs` | `5` | Number of times the policy is updated using each batch of collected experience. Higher values extract more learning from each batch but risk overfitting | +| `num_mini_batches` | `4` | Number of mini-batches the experience buffer is split into for each epoch. More mini-batches mean smaller gradient updates | +| `learning_rate` | `1e-3` | Step size for the Adam optimizer. Controls how much the network weights change per update. Too high causes instability; too low slows convergence | +| `discount_factor` (gamma) | `0.99` | How much the agent values future rewards vs. immediate rewards. A value of 0.99 means the agent considers rewards ~100 steps into the future | +| `gae_lambda` (lambda) | `0.95` | Generalized Advantage Estimation smoothing parameter. Balances bias (low lambda) vs. variance (high lambda) in advantage estimates | +| `clip_param` | `0.2` | PPO clipping range. Prevents the policy from changing too much in a single update. Keeps training stable | +| `value_loss_coef` | `1.0` | Weight of the value function loss relative to the policy loss. Ensures the critic learns at an appropriate rate | +| `entropy_coef` | `0.01` | Weight of the entropy bonus. Encourages exploration by penalizing overly deterministic policies. Reduce this as training converges | +| `desired_kl` | `0.01` | Target KL divergence between old and new policies. If KL exceeds this value, the learning rate is reduced adaptively | +| `max_grad_norm` | `1.0` | Maximum gradient norm for gradient clipping. Prevents exploding gradients during training | + +### Rollout hyperparameters + +| **Hyperparameter** | **Typical value** | **Description** | +|--------------------|-------------------|-----------------| +| `num_steps_per_env` | `24` | Number of simulation steps collected per environment per iteration. Together with `num_envs`, this determines the total batch size: `batch_size = num_envs Γ— num_steps_per_env` | +| `save_interval` | `50` | Save a model checkpoint every N iterations. Useful for resuming training or evaluating intermediate policies | + +### How the hyperparameters interact + +The total amount of experience collected per training iteration is: + +``` +batch_size = num_envs Γ— num_steps_per_env +``` + +For example, with `num_envs=4096` and `num_steps_per_env=24`: + +``` +batch_size = 4096 Γ— 24 = 98,304 environment steps per iteration +``` + +This batch is then split into `num_mini_batches` (4) mini-batches of ~24,576 steps each. The policy is updated `num_learning_epochs` (5) times per iteration, meaning each batch of experience is used for 5 Γ— 4 = 20 gradient updates. + +## Step 4: Monitor the training + +During training, RSL-RL prints statistics to the terminal at regular intervals. A typical output looks like: + +```output +Learning iteration 100/1500 + mean reward: 12.45 + mean episode length: 234.5 + value function loss: 0.032 + surrogate loss: -0.0156 + mean std: 0.42 + learning rate: 0.001 + fps: 48523 +``` + +Interpreting these values helps track convergence and diagnose training instability, such as stagnating rewards or exploding losses. + +The following table explains each metric: + +| **Metric** | **What it means** | **What to look for** | +|-----------|-------------------|----------------------| +| `mean reward` | Average cumulative reward across all environments per episode | Should increase over time. Higher values mean the robot is walking better | +| `mean episode length` | Average number of steps before the episode ends | Should increase as the robot learns to stay upright longer | +| `value function loss` | How well the critic predicts future rewards | Should decrease and stabilize | +| `surrogate loss` | The PPO policy loss (negative because PPO maximizes the objective) | Should be small and negative | +| `mean std` | Average exploration noise standard deviation | Should decrease as the policy becomes more confident | +| `learning rate` | Current learning rate (may be adjusted by the adaptive KL mechanism) | Stays at the initial value unless KL divergence exceeds `desired_kl` | +| `fps` | Frames (environment steps) per second | Indicates training throughput. On DGX Spark, expect 40,000-60,000+ fps for locomotion tasks | + +{{% notice Note %}} +Training checkpoints are saved to the `logs/rsl_rl/` directory. Each run creates a timestamped folder containing the model weights, configuration, and training logs. +{{% /notice %}} + +## Step 5: Evaluate the trained policy + +After training completes, evaluate the policy by running inference with visualization. Use the `play.py` script with the trained checkpoint: + +```bash +./isaaclab.sh -p scripts/reinforcement_learning/rsl_rl/play.py \ + --task=Isaac-Velocity-Rough-H1-Play-v0 \ + --num_envs=512 +``` + +{{% notice Note %}} +For evaluation, use the inference task name `Isaac-Velocity-Rough-H1-Play-v0` instead of the training task name. The play variant disables runtime perturbations used during training and loads the checkpoint automatically. +{{% /notice %}} + +The play script loads the most recent checkpoint and runs the policy in real time. You will observe the Unitree H1 humanoid walking over procedurally generated rough terrain, responding to live velocity commands. + +You can also specify a particular checkpoint manually, which is useful for comparing intermediate policy performance. + +```bash +./isaaclab.sh -p scripts/reinforcement_learning/rsl_rl/play.py \ + --task=Isaac-Velocity-Rough-H1-Play-v0 \ + --num_envs=512 \ + --checkpoint=logs/rsl_rl/h1_rough//model_1500.pt +``` + +### Understanding the evaluation + +During evaluation, you can observe how the robot's behavior improves over the course of training: + +- **Early training (iterations 0–200)**: The robot often collapses immediately or performs erratic, uncoordinated motions. +- **Mid training (iterations 200–800)**: The robot begins to walk forward with some success, though it may still stumble or lose balance on rough terrain. +- **Late training (iterations 800–1500)**: The robot consistently walks over uneven terrain, responds to velocity commands, and recovers from disturbances. + +This progressionβ€”from falling to stable walkingβ€”demonstrates how PPO gradually improves the policy through trial and error across thousands of parallel environments. + +The following visualizations compare two training stages using `num_envs=512`, showcasing the benefit of large-scale parallel training on DGX Spark. + +*** Iteration 50 (Early Stage, num_envs=512) *** + +At iteration 50, the policy is still in its exploration phase. Most robots exhibit noisy joint actions, lack coordination, and frequently fall. There is no observable response to the velocity command, and no stable gait has emerged. + +![img3 alt-text#center](isaaclab_h1_512_0050.gif "Figure 3: Early Stage") + +*** Iteration 1250 (Late Stage, num_envs=512) *** + +By iteration 1350, the policy has matured. Most robots demonstrate coordinated walking behavior, balance maintenance, and accurate velocity tracking, even on rough terrain. The improvement in foot placement and heading stability is clearly visible. + +![img4 alt-text#center](isaaclab_h1_512_1350.gif "Figure 4: Late Stage") + +## What you have accomplished + +In this module, you have: + +- Trained a reinforcement learning policy for the Unitree H1 humanoid robot using RSL-RL and the PPO algorithm +- Understood key hyperparameters in the training pipeline, including policy architecture, rollout strategy, and PPO optimization settings +- Monitored training progress using reward curves, episode statistics, and performance metrics +- Evaluated the trained policy through interactive visualization and behavior analysis + +You have now completed the end-to-end workflow of training and validating a reinforcement learning policy for humanoid locomotion on DGX Spark. \ No newline at end of file diff --git a/content/learning-paths/laptops-and-desktops/dgx_spark_isaac_robotics/_index.md b/content/learning-paths/laptops-and-desktops/dgx_spark_isaac_robotics/_index.md new file mode 100644 index 0000000000..52874e8400 --- /dev/null +++ b/content/learning-paths/laptops-and-desktops/dgx_spark_isaac_robotics/_index.md @@ -0,0 +1,72 @@ +--- +title: Build Robot Simulation and RL Workflows with Isaac Sim and Isaac Lab on DGX Spark + +draft: true +cascade: + draft: true + +minutes_to_complete: 90 + +who_is_this_for: This learning path is intended for robotics developers, simulation engineers, and AI researchers who want to run high-fidelity robotic simulations and reinforcement learning (RL) pipelines using Isaac Sim and Isaac Lab on Arm-based NVIDIA DGX Spark systems powered by the Grace–Blackwell (GB10) architecture. + +learning_objectives: + - Explain the roles of Isaac Sim and Isaac Lab, and describe how DGX Spark accelerates robotic simulation and reinforcement learning workloads + - Build Isaac Sim and Isaac Lab from source on an Arm-based DGX Spark system + - Launch and control a basic robot simulation in Isaac Sim using Python scripts + - Train and evaluate a reinforcement learning policy for the Unitree H1 humanoid robot using Isaac Lab and the RSL-RL interface + +prerequisites: + - Access to an NVIDIA DGX Spark system with at least 50 GB of free disk space + - Familiarity with Linux command-line tools + - Experience with Python scripting and virtual environments + - Basic understanding of reinforcement learning concepts (rewards, policies, episodes) + - Experience building software from source using CMake and make + +author: + - Johnny Nunez + - Odin Shen + - Asier Arranz + - Raymond Lo + +### Tags +skilllevels: Advanced +subjects: ML +armips: + - Cortex-X + - Cortex-A +tools_software_languages: + - Python + - Bash + - IsaacSim + - IsaacLab +operatingsystems: + - Linux + +further_reading: + - resource: + title: Isaac Sim Documentation + link: https://docs.isaacsim.omniverse.nvidia.com/latest/index.html + type: documentation + - resource: + title: Isaac Lab Documentation + link: https://isaac-sim.github.io/IsaacLab/main/index.html + type: documentation + - resource: + title: NVIDIA DGX Spark Playbooks + link: https://github.com/NVIDIA/dgx-spark-playbooks + type: documentation + - resource: + title: Isaac Lab Available Environments + link: https://isaac-sim.github.io/IsaacLab/main/source/overview/environments.html + type: website + - resource: + title: DGX Spark Isaac Sim and Isaac Lab Playbook + link: https://build.nvidia.com/spark/isaac/overview + type: website + +### FIXED, DO NOT MODIFY +# ================================================================================ +weight: 1 # _index.md always has weight of 1 to order correctly +layout: "learningpathall" # All files under learning paths have this same wrapper +learning_path_main_page: "yes" # This should be surfaced when looking for related content. Only set for _index.md of learning path content. +--- diff --git a/content/learning-paths/laptops-and-desktops/dgx_spark_isaac_robotics/_next-steps.md b/content/learning-paths/laptops-and-desktops/dgx_spark_isaac_robotics/_next-steps.md new file mode 100644 index 0000000000..c3db0de5a2 --- /dev/null +++ b/content/learning-paths/laptops-and-desktops/dgx_spark_isaac_robotics/_next-steps.md @@ -0,0 +1,8 @@ +--- +# ================================================================================ +# FIXED, DO NOT MODIFY THIS FILE +# ================================================================================ +weight: 21 # Set to always be larger than the content in this path to be at the end of the navigation. +title: "Next Steps" # Always the same, html page title. +layout: "learningpathall" # All files under learning paths have this same wrapper for Hugo processing. +--- diff --git a/content/learning-paths/laptops-and-desktops/dgx_spark_isaac_robotics/isaaclab_h1_512_0050.gif b/content/learning-paths/laptops-and-desktops/dgx_spark_isaac_robotics/isaaclab_h1_512_0050.gif new file mode 100644 index 0000000000..b0b107294b Binary files /dev/null and b/content/learning-paths/laptops-and-desktops/dgx_spark_isaac_robotics/isaaclab_h1_512_0050.gif differ diff --git a/content/learning-paths/laptops-and-desktops/dgx_spark_isaac_robotics/isaaclab_h1_512_1350.gif b/content/learning-paths/laptops-and-desktops/dgx_spark_isaac_robotics/isaaclab_h1_512_1350.gif new file mode 100644 index 0000000000..1ea10549b5 Binary files /dev/null and b/content/learning-paths/laptops-and-desktops/dgx_spark_isaac_robotics/isaaclab_h1_512_1350.gif differ diff --git a/content/learning-paths/laptops-and-desktops/dgx_spark_isaac_robotics/run_articulation.gif b/content/learning-paths/laptops-and-desktops/dgx_spark_isaac_robotics/run_articulation.gif new file mode 100644 index 0000000000..b2632f6139 Binary files /dev/null and b/content/learning-paths/laptops-and-desktops/dgx_spark_isaac_robotics/run_articulation.gif differ diff --git a/content/learning-paths/mobile-graphics-and-gaming/ai-camera-pipelines/_index.md b/content/learning-paths/mobile-graphics-and-gaming/ai-camera-pipelines/_index.md index d073aa4b80..ec608c1ce8 100644 --- a/content/learning-paths/mobile-graphics-and-gaming/ai-camera-pipelines/_index.md +++ b/content/learning-paths/mobile-graphics-and-gaming/ai-camera-pipelines/_index.md @@ -23,9 +23,11 @@ skilllevels: Introductory subjects: Performance and Architecture armips: - Cortex-A + - Arm C1 tools_software_languages: - CPP - Docker + - SME2 operatingsystems: - Linux - macOS diff --git a/content/learning-paths/mobile-graphics-and-gaming/kai_sme2_matmul_ukernel_explained/_index.md b/content/learning-paths/mobile-graphics-and-gaming/kai_sme2_matmul_ukernel_explained/_index.md index aa3e8a16a9..8cdfd61acf 100644 --- a/content/learning-paths/mobile-graphics-and-gaming/kai_sme2_matmul_ukernel_explained/_index.md +++ b/content/learning-paths/mobile-graphics-and-gaming/kai_sme2_matmul_ukernel_explained/_index.md @@ -1,22 +1,20 @@ --- -title: KleidiAI SME2 matmul microkernel for quantized models explained - -draft: true -cascade: - draft: true +title: Understand KleidiAI SME2 matmul microkernels minutes_to_complete: 40 -who_is_this_for: This is an advanced topic for software developers, performance engineers, and AI practitioners +who_is_this_for: This is an advanced topic for software developers, performance engineers, and AI practitioners. learning_objectives: - - Learn how a KleidiAI matmual microkernel performs matrix multiplication with quantized data - - Learn how SME2 INT8 Outer Product Accumulate instructions are used for matrix multiplication - - Learn how a KleidiAI SME2 matmul microkernel accelerates matmul operators in a Large Lanague Model - - Learn how to integrate KleidiAI SME2 matmul microkernels to an AI framework or application + - Explain how a KleidiAI microkernel performs matrix multiplication (matmul) with quantized data + - Identify how SME2 INT8 MOPA (matrix outer product accumulate) instructions map to matmul work + - Trace how quantization and packing feed an SME2 matmul microkernel (using GGML Q4_0 and llama.cpp call stacks as a concrete example) + - Perform basic hands-on checks (source inspection and optional disassembly) to confirm where SME2 instructions appear prerequisites: - - Knowledge of KleidiAI and SME2 + - Basic understanding of general matrix multiplication (GEMM) and matmul operations + - Basic understanding of quantization concepts for neural networks + - (Optional) Access to an Arm CPU with SME2 support (Linux or Android) for hands-on verification steps author: Zenon Zhilong Xiu @@ -24,12 +22,12 @@ author: Zenon Zhilong Xiu skilllevels: Advanced subjects: ML armips: - - Arm C1 CPU - - Arm SME2 unit + - Arm C1 tools_software_languages: - C++ - KleidiAI - llama.cpp + - SME2 operatingsystems: - Android - Linux @@ -38,20 +36,20 @@ operatingsystems: further_reading: - resource: - title: part 1 Arm Scalable Matrix Extension Introduction + title: Part 1, Arm Scalable Matrix Extension introduction link: https://developer.arm.com/community/arm-community-blogs/b/architectures-and-processors-blog/posts/arm-scalable-matrix-extension-introduction type: blog - resource: - title: part 2 Arm Scalable Matrix Extension Instructions + title: Part 2, Arm Scalable Matrix Extension instructions link: https://developer.arm.com/community/arm-community-blogs/b/architectures-and-processors-blog/posts/arm-scalable-matrix-extension-introduction-p2 type: blog - resource: - title: part4 Arm SME2 Introduction + title: Part 4 Arm SME2 introduction link: https://developer.arm.com/community/arm-community-blogs/b/architectures-and-processors-blog/posts/part4-arm-sme2-introduction type: blog - resource: title: Profile llama.cpp performance with Arm Streamline and KleidiAI LLM kernels - link: https://learn.arm.com/learning-paths/servers-and-cloud-computing/llama_cpp_streamline/ + link: /learning-paths/servers-and-cloud-computing/llama_cpp_streamline/ type: blog diff --git a/content/learning-paths/mobile-graphics-and-gaming/kai_sme2_matmul_ukernel_explained/explain_with_an_example_p1.md b/content/learning-paths/mobile-graphics-and-gaming/kai_sme2_matmul_ukernel_explained/explain_with_an_example_p1.md index 414c677fad..65c123100f 100644 --- a/content/learning-paths/mobile-graphics-and-gaming/kai_sme2_matmul_ukernel_explained/explain_with_an_example_p1.md +++ b/content/learning-paths/mobile-graphics-and-gaming/kai_sme2_matmul_ukernel_explained/explain_with_an_example_p1.md @@ -1,23 +1,29 @@ --- -title: Explain the SME2 matmul microkernel with an example - Part 1 -weight: 5 +title: Repack RHS weights (GGML Q4_0) +weight: 6 ### FIXED, DO NOT MODIFY layout: learningpathall --- -## Explain the SME2 matmul microkernel with an example - Part 1 -By integrating the SME2‑optimized KleidiAI kernels into llama.cpp, the heavy matrix‑multiplication workloads in the K, Q, and V computations of the attention blocks, as well as in the FFN layers, can be delegated to the SME2 matmul microkernel when running the Llama-3.2-3B-Q4_0.gguf model. -In these operators, the LHS (activation) data type is FP32, while RHS (weight) type uses GGML Q4_0 quantized type. +## Repack RHS weights (GGML Q4_0) +When you [integrate SME2-optimized KleidiAI kernels into llama.cpp](/learning-paths/mobile-graphics-and-gaming/performance_llama_cpp_sme2/), the heavy matrix-multiplication work in attention (K/Q/V projections) and feed-forward network (FFN) layers can run through the SME2 matmul microkernel. -To make the demonstration easier in this learning path, the LHS dimension [m, k] is simplified to [16, 64], the RHS dimension [n, k] is simplified to [64, 64], and the SME2 SVL is set as 512-bit. +In these operators, the LHS (activations) is FP32 and the RHS (weights) uses the GGML Q4_0 quantized type. -###Packing the RHS -Although the original Q4_0 RHS(weight) in the model uses INT4 quantization, it is signed INT4 quantization, rather than the unsigned INT4 quantization that the SME2 matmul microkernel requires. Moreover,the layout of the INT4 quantized data and the quantization scale does not meet the requirements of the SME2 matmul microkernel neither. Therefore, the LHS from the model needs to be converted from the signed INT4 data to unsigned INT4 and repacked. -Since the RHS(weight) remains unchanged during the inference, this conversion and packing only need to be performed only once when loading the model. +To keep the example readable, this Learning Path uses a simplified matmul shape: +- LHS `[m, k] = [16, 64]` +- RHS `[n, k] = [64, 64]` +It also assumes an SME2 SVL of 512 bits. -Let us have a close look at GGML Q4_0 quantization first to know how the orginal FP32 weight is quantized to Q4_0 format. +### Pack the RHS +Although the original Q4_0 RHS (weights) uses INT4 quantization, it is signed INT4 and uses a GGML-specific layout and scale encoding. The SME2 matmul microkernel expects a different packed RHS representation (including an unsigned INT4 form and per-block metadata arranged for efficient loads). + +Because the RHS (weights) stays constant during inference, you only need to convert and pack it once when loading the model. + + +Start by reviewing GGML Q4_0 quantization so you can see why the original RHS layout doesn’t match what the SME2 microkernel expects. In the Q4_0 model, the Q4_0 weights are stored in layout of [n, k]. GGML Q4_0 quantizes weights in blocks of 32 floats. For each block, it calculates a scale for the block and then converts each value into a signed 4-bit integer. The scale is stored as FP16. Then GGML Q4_0 packs the values in a way of, @@ -29,7 +35,7 @@ The following diagram shows how GGML Q4_0 quantizes and packs the original [n, k Unfortunately, the Q4_0 format does not meet the requirements of the SME2 matmul microkernel. It needs to be converted to an unsigned INT4 quantization format and repacked using the *kai_run_rhs_pack_nxk_qsi4c32ps1s0scalef16_qsu4c32s16s0_neon* function. -In this example, we use m=16 and k=64. +This example uses m=16 and k=64. - The required mr value for the SME2 matmul kernel is obtained using *kai_get_mr_matmul_clamp_f32_qai8dxp1vlx4_qsi8cxp4vlx4_1vlx4vl_sme2_mopa*. Here, mr=16. - The required nr value for the SME2 matmul kernel is obtained using *kai_get_nr_matmul_clamp_f32_qai8dxp1vlx4_qsi8cxp4vlx4_1vlx4vl_sme2_mopa*. Here, nr=64. - The required kr value for the SME2 matmul kernel is obtained using *kai_get_kr_matmul_clamp_f32_qai8dxp1vlx4_qsi8cxp4vlx4_1vlx4vl_sme2_mopa*. Here, kr=4. @@ -49,7 +55,29 @@ This process can be illustrated with the diagram below. ![Figure showing RHS packing with KleidiAI alt-text#center](images/kai_kernel_packed_rhs.jpg "RHS packing with KleidiAI") The numerical label of an element in the diagram is used to indicate its row and column number in the original matrix. For example , -![Figure showing Row_Col lable alt-text#center](images/row_col_lable.png "Row_Col lable") -it indicates that the element locates at row 01, column 02 in the original matrix. This row and column number remains unchanged in its quantized and packed matrix, so that the location of the element can be tracked easily. +![Figure showing row/column label for tracking elements alt-text#center](images/row_col_lable.png "Row/column label used for tracking") +This indicates the element is at row 01, column 02 in the original matrix. This row/column label stays consistent through quantization and packing so you can track elements across layouts. + +After this step, the RHS is converted and packed into a format the SME2 matmul microkernel can consume efficiently. This allows the kernel to load packed RHS data into SME2 Z registers using sequential memory access, which improves cache locality. + +### Hands-on: find the RHS repack microkernel in KleidiAI (optional) + +If you cloned the KleidiAI repository earlier, you can locate the repack function used for RHS conversion. + +From the KleidiAI repo root: + +```bash +grep -R "kai_run_rhs_pack_nxk_qsi4c32ps1s0scalef16_qsu4c32s16s0_neon" -n kai | head +``` + +You can also search for the `qsu4` (unsigned INT4) string to find related packers: + +```bash +grep -R "qsu4" -n kai | head +``` + +## What you've learned and what's next + +You've learned how GGML Q4_0 weights are repacked from their original signed INT4 layout into the unsigned INT4 format the SME2 matmul microkernel expects. You understand why this conversion happens once at model load time. -Now, the RHS is converted and packed into a format that can be handled by the SME2 matmul microkernel, allowing the packed RHS to be loaded into SME2 Z registers with sequential memory access. This improves memory access efficiency and reduces cache misses. \ No newline at end of file +Next, you'll see how the FP32 LHS activations are quantized and packed dynamically during inference. \ No newline at end of file diff --git a/content/learning-paths/mobile-graphics-and-gaming/kai_sme2_matmul_ukernel_explained/explain_with_an_example_p2.md b/content/learning-paths/mobile-graphics-and-gaming/kai_sme2_matmul_ukernel_explained/explain_with_an_example_p2.md index 4e4851dc64..06db0fe31f 100644 --- a/content/learning-paths/mobile-graphics-and-gaming/kai_sme2_matmul_ukernel_explained/explain_with_an_example_p2.md +++ b/content/learning-paths/mobile-graphics-and-gaming/kai_sme2_matmul_ukernel_explained/explain_with_an_example_p2.md @@ -1,16 +1,17 @@ --- -title: Explain the SME2 matmul microkernel with an example - Part 2 -weight: 6 +title: Quantize and pack LHS activations +weight: 7 ### FIXED, DO NOT MODIFY layout: learningpathall --- -## Explain the SME2 matmul microkernel with an example - Part 2 +## Quantize and pack LHS activations Next, the FP32 LHS (activation) needs to be quantized and packed when the llama.cpp graph runner computes the matmul nodes/operators. -### Quantization and Packing of the LHS -Since the LHS (activation) keep changing, we need to dynamically quantize the original FP32 matrix and pack it into the qsi8d32p1vlx4 format. This can be achieved using the *kai_run_lhs_quant_pack_qsi8d32p_f32_neon* microkernel. +### Quantize and pack the LHS + +Because the LHS (activations) changes for each matmul invocation, it must be quantized and packed dynamically. In this example, the FP32 LHS is quantized to signed INT8 and packed into the `qsi8d32p1vlx4` format using the *kai_run_lhs_quant_pack_qsi8d32p_f32_neon* microkernel. The function call stack for this process in llama.cpp is as follows: ```text @@ -27,8 +28,29 @@ llama_context::decode ggml::cpu::kleidiai::tensor_traits::compute_forward_q4_0 kai_run_lhs_quant_pack_qsi8d32p_f32_neon ``` -The diagram below illustrates how the RHS is quantized and packed by *kai_run_lhs_quant_pack_qsi8d32p_f32_neon*, +The diagram below illustrates how the LHS is quantized and packed by *kai_run_lhs_quant_pack_qsi8d32p_f32_neon*: ![Figure showing Quantization and Packing of the LHS alt-text#center](images/kai_run_lhs_quant_pack_qsi8d32p_f32_neon_for_sme2.jpg "Quantization and Packing of the LHS") The values of mr, nr, and kr can be obtained in the same way as described above. -The mr, nr, and kr together with the matrix dimensions m and k are passed as parameters to *kai_run_lhs_quant_pack_qsi8d32p_f32_neon*. This function quantizes the FP32 LHS to signed INT8 type and packed the quantized data and quantization scales as shown in the diagram above. It divides the m x n matrix into submatrices of size mr x kr (it is 16 x 4) as shown in blocks outlined by dashed lines in the upper matrix of the diagram, and then sequentially packs the rows within each submatrix. This allows the SME2 matmul kernel to load an entire submatrix into an SME2 Z register from contiguous memory, thus reducing cache misses by avoiding loading the submatrix across multiple rows. +The values of `mr`, `nr`, and `kr`, together with the matrix dimensions `m` and `k`, are passed as parameters to *kai_run_lhs_quant_pack_qsi8d32p_f32_neon*. + +This microkernel: +- Quantizes FP32 LHS values to signed INT8 +- Stores the per-block scales +- Packs the quantized values into a contiguous layout based on `mr Γ— kr` tiles (16 Γ— 4 in this example) + +This packing ensures the SME2 matmul microkernel can load a full input slice from contiguous memory, which improves cache locality. + +### Hands-on: locate the LHS quant-pack microkernel in KleidiAI (optional) + +If you cloned the KleidiAI repository earlier, locate the LHS quantization + packing microkernel: + +```bash +grep -R "kai_run_lhs_quant_pack_qsi8d32p_f32_neon" -n kai | head +``` + +## What you've learned and what's next + +You've learned how the FP32 LHS activations are quantized to signed INT8 and packed into the qsi8d32p1vlx4 format during inference. You understand how this dynamic packing enables efficient contiguous memory access in the SME2 microkernel. + +Next, you'll walk through the SME2 matmul microkernel inner loop to see how packed LHS and RHS feed the MOPA instructions. diff --git a/content/learning-paths/mobile-graphics-and-gaming/kai_sme2_matmul_ukernel_explained/explain_with_an_example_p3.md b/content/learning-paths/mobile-graphics-and-gaming/kai_sme2_matmul_ukernel_explained/explain_with_an_example_p3.md index caa702f4bd..42b5b707e4 100644 --- a/content/learning-paths/mobile-graphics-and-gaming/kai_sme2_matmul_ukernel_explained/explain_with_an_example_p3.md +++ b/content/learning-paths/mobile-graphics-and-gaming/kai_sme2_matmul_ukernel_explained/explain_with_an_example_p3.md @@ -1,33 +1,36 @@ --- -title: Explain the SME2 matmul microkernel with an example- Part 3 -weight: 7 +title: Walk the SME2 matmul inner loop +weight: 8 ### FIXED, DO NOT MODIFY layout: learningpathall --- -## Explain the SME2 matmul microkernel with an example - Part 3 -Once the required LHS and RHS are both ready, *kai_matmul_clamp_f32_qsi8d32p1vlx4_qsi4c32p4vlx4_1vlx4vl_sme2_mopa* microkernel can run now. +## Walk the SME2 matmul inner loop +Once both the packed LHS and packed RHS are ready, the SME2 matmul microkernel can run. -### Run the SME2 matmul microkernel -The operations performed to compute an 16x64 result submatrice (four 16x16 submatrices) (1VL x 4VL) are as follows: +### Run the SME2 matmul microkernel (conceptually) -- Iterate along blocks along K dimension +This section describes the core operations that compute a 16Γ—64 output tile (four 16Γ—16 submatrices, 1VL Γ— 4VL). + +- Iterate over the `K` dimension in blocks - Iterate in a block with step of kr (kr=4) - Load one SME2 SVL-length (512-bit) of data from the quantized and packed LHS (containing 64 INT8 values) into one SME2 Z register - Load two SME2 SVL-lengths of data from the packed RHS (containing 2 x64x2 INT4 values) into two SME2 Z registers, then use the SME2 LUTI4 lookup table instruction to convert these INT4 values into INT8 type, extending them to four SME2 Z registers (4VL). - - Use the SME2 INT8 Outer Product Accumulate (MPOA) instruction to perform outer product operations with source from the Z register and each of the four Z registers, accumulates the results in four ZA tiles (which are initialized to zero). It produces intermediate results of four 16x16 output submatrices. - The processes of the first itration can be illustrated in the diagram below: -![Figure showing the first itration of the inner loop alt-text#center](images/run_matmul_sme2_step1.jpg "The first itration of the inner loop") - The diagram below illustrates the process of the second iteration along the K dimension, -![Figure showing the second itration of the inner loop alt-text#center](images/run_matmul_sme2_step2.jpg "The second itration of the inner loop") - - After completing the iterations in the block, the intermediate INT32 results of four 16x16 output submatrices are dequantized with the per-block LHS and RHS scale to FP32 floats, using Floating-point Multiply (FMUL), Floating-point Multiply and Accumulate (FMLA) and Signed fixed-point Convert to Floating-point (SCVTF) vector instructions. It produces the intermediate FP32 results of four 16x16 output submatrices. + - Use the SME2 INT8 MOPA (outer product accumulate) instruction to compute outer products and accumulate into four ZA tiles (initialized to zero). This produces intermediate results for four 16Γ—16 output submatrices. + + The first iteration can be illustrated in the diagram below: +![Figure showing the first iteration of the inner loop alt-text#center](images/run_matmul_sme2_step1.jpg "The first iteration of the inner loop") + The diagram below illustrates the second iteration along `K`: +![Figure showing the second iteration of the inner loop alt-text#center](images/run_matmul_sme2_step2.jpg "The second iteration of the inner loop") + - After completing the iterations in the block, the intermediate INT32 results are dequantized to FP32 using the per-block LHS and RHS scales. + - This step uses vector instructions such as floating-point multiply (FMUL), floating-point multiply-accumulate (FMLA), and signed fixed-point convert to floating-point (SCVTF). + - It produces intermediate FP32 results for the four 16Γ—16 output submatrices. - Accumulate the FP32 result above -After completing itration along the K dimension, the FP32 results of four 16x16 output submatrices is ready. Then, save the result into memory. +After completing the iterations along the `K` dimension, the FP32 results of the four 16Γ—16 output submatrices are ready. The kernel then stores the output tile back to memory. -The code can be found [here](https://github.com/ARM-software/kleidiai/blob/main/kai/ukernels/matmul/matmul_clamp_f32_qsi8d32p_qai4c32p/kai_matmul_clamp_f32_qsi8d32p1vlx4_qai4c32p4vlx4_1vlx4vl_sme2_mopa_asm.S#L80) -Some comments are added to the code to help understanding the code. +The snippet (from [the source code](https://github.com/ARM-software/kleidiai/blob/main/kai/ukernels/matmul/matmul_clamp_f32_qsi8d32p_qai4c32p/kai_matmul_clamp_f32_qsi8d32p1vlx4_qai4c32p4vlx4_1vlx4vl_sme2_mopa_asm.S)) below highlights the inner loop structure. ```asm KAI_ASM_LABEL(label_3) // K Loop KAI_ASM_INST(0xc00800ff) // zero {za} , zeros the four ZA tile (za0.s, za1.s, za2.s, za3.s) @@ -42,7 +45,7 @@ KAI_ASM_LABEL(label_4) // Block Loop KAI_ASM_INST(0xa0840100) // smopa za0.s, p0/m, p0/m, z8.b, z4.b ] //Outer Product Accumulate with the VL of LHS, the first VL of RHS and ZA0.S KAI_ASM_INST(0xa0850101) // smopa za1.s, p0/m, p0/m, z8.b, z5.b //Outer Product Accumulate with the VL of LHS, the second VL of RHS and ZA1.S KAI_ASM_INST(0xa0860102) // smopa za2.s, p0/m, p0/m, z8.b, z6.b //Outer Product Accumulate with the VL of LHS, the third VL of RHS and ZA2.S - KAI_ASM_INST(0xa0870103) // smopa za3.s, p0/m, p0/m, z8.b, z7.b b //Outer Product Accumulate with the VL of LHS, the forth VL of RHS and ZA3.S + KAI_ASM_INST(0xa0870103) // smopa za3.s, p0/m, p0/m, z8.b, z7.b // Outer Product Accumulate with the VL of LHS, the fourth VL of RHS and ZA3.S subs x11, x11, #4 //block_index - 4 b.gt label_4 //end of block iteration? @@ -59,20 +62,39 @@ KAI_ASM_LABEL(label_4) // Block Loop pfalse p3.b KAI_ASM_LABEL(label_5) // omit some codes that perform the block quantization and save the result to memory - …… + ... blt label_5 subs x10, x10, x4 //decrease the K index b.gt label_3 //end of K loop? ``` -In a single block loop, four pipelined SME2 INT8 MOPA instructions perform 4,096 MAC operations, calculating the intermediate results for the four 16x16 submatrices. It proves that SME2 MOPA can significantly improve matrix multiplication performance. +In a single block-loop iteration, four pipelined SME2 INT8 MOPA instructions perform 4,096 multiply-accumulate (MAC) operations, calculating the intermediate results for four 16Γ—16 submatrices. -To help understand the whole process, we map the first itration of LHS and RHS quantization and packing steps, as well as SME2 outer product accumulate operation and dequantization, back to the original FP32 LHS and RHS operations. Essentially, they equally perform the operation as shown below (there might be some quantization loss), -![Figure showing the original matrix representing of the first itration alt-text#center](images/run_matmul_sme2_original_present_step1.jpg "the original matrix representing of the first itration") +To connect this back to "normal" FP32 matmul, map one iteration of quantization + packing, the SME2 MOPA step, and dequantization to an equivalent FP32 computation (with expected quantization loss). +![Figure showing the original matrix representation of the first iteration alt-text#center](images/run_matmul_sme2_original_present_step1.jpg "The original matrix representation of the first iteration") The second iteration can be mapped back to the original FP32 LHS and RHS operations as below, -![Figure showing the original matrix representing of the second itration alt-text#center](images/run_matmul_sme2_original_present_step2.jpg "the original matrix representing of the second itration") +![Figure showing the original matrix representation of the second iteration alt-text#center](images/run_matmul_sme2_original_present_step2.jpg "The original matrix representation of the second iteration") **Note**: In this diagram, the RHS is laid out in the dimension of [N, K], which is different from the [K, N] dimension layout of the RHS in the video demonstration of 1VLx4VL. If you interpret the RHS in the diagrams above using the [K, N] dimension, you can match the previous video demonstration with the diagrams above. By repeating the submatrix computation across the M and N dimensions, the entire result matrix can be calculated. If a non-empty bias is passed to the SME2 matmul microkernel, it also adds the bias to the result matrix. + +### Hands-on: follow the inner loop in the source (optional) + +If you cloned KleidiAI earlier, you can search the microkernel for the exact "load β†’ LUTI β†’ MOPA" sequence shown above. + +From the KleidiAI repo root: + +```bash +KERNEL_FILE="kai/ukernels/matmul/matmul_clamp_f32_qsi8d32p_qai4c32p/kai_matmul_clamp_f32_qsi8d32p1vlx4_qai4c32p4vlx4_1vlx4vl_sme2_mopa_asm.S" +grep -n "zero {za}" "$KERNEL_FILE" | head +grep -n "luti" "$KERNEL_FILE" | head +grep -n "smopa" "$KERNEL_FILE" | head +``` + +If you have a recent LLVM toolchain on an Arm Linux host, you can also try assembling and disassembling the kernel to see `smopa` in the output. Toolchain details vary across distributions, so treat this step as optional. + +## What you've learned + +You've walked through the SME2 matmul microkernel inner loop and seen how it loads packed data, uses LUTI4 to expand INT4 to INT8, executes pipelined MOPA instructions, and dequantizes results back to FP32. You understand how this maps back to standard FP32 matmul operations and can trace the complete dataflow from packed inputs through MOPA accumulation to FP32 output. diff --git a/content/learning-paths/mobile-graphics-and-gaming/kai_sme2_matmul_ukernel_explained/introduction.md b/content/learning-paths/mobile-graphics-and-gaming/kai_sme2_matmul_ukernel_explained/introduction.md new file mode 100644 index 0000000000..aa39d2dc9e --- /dev/null +++ b/content/learning-paths/mobile-graphics-and-gaming/kai_sme2_matmul_ukernel_explained/introduction.md @@ -0,0 +1,89 @@ +--- +title: Overview and setup +weight: 2 + +### FIXED, DO NOT MODIFY +layout: learningpathall +--- + +## KleidiAI SME2 matmul microkernel overview + +KleidiAI includes highly optimized matrix multiplication (matmul) microkernels that accelerate quantized operators on Arm CPUs. On SME2-capable systems, these microkernels use SME2 INT8 MOPA (outer product accumulate) instructions to increase throughput for the compute-heavy parts of inference. + +This Learning Path focuses on one concrete microkernel and walks through: +- How the microkernel expects its inputs (quantization + packing) +- Where SME2 instructions show up in the inner loop +- How this maps back to "normal" FP32 matmul semantics + +### Where llama.cpp appears in this Learning Path + +This Learning Path mentions llama.cpp as a concrete reference point in two ways: +- It uses GGML Q4_0 as a real-world weight format that needs repacking before it can feed the SME2 microkernel. +- It uses llama.cpp call stacks to show where RHS repacking and LHS quantization/packing happen in a real inference pipeline. + +It does not ask you to build or run llama.cpp end-to-end. If you want to build and profile llama.cpp with KleidiAI on-device, use [the llama.cpp performance Learning Path](/learning-paths/mobile-graphics-and-gaming/performance_llama_cpp_sme2/). + +## What you’ll do (hands-on) + +You’ll do a few lightweight, practical checks as you go: +- Confirm whether your target device exposes SME2 (optional) +- Locate the exact microkernel source file in KleidiAI +- Search for key SME2 instructions (for example `smopa`, `luti4`, and `zero {za}`) +- Connect those instructions back to the diagrams and pseudocode in this Learning Path + +If you already have llama.cpp built with KleidiAI (for example from [the llama.cpp performance Learning Path](/learning-paths/mobile-graphics-and-gaming/performance_llama_cpp_sme2/)), you can also use that build to validate call stacks and confirm where the kernel runs. + +## Hands-on: get the kernel source + +If you want to follow the microkernel implementation directly, start by cloning KleidiAI and locating the SME2 matmul microkernel assembly file. + +```bash +git clone https://github.com/ARM-software/kleidiai.git +cd kleidiai + +KERNEL_FILE="kai/ukernels/matmul/matmul_clamp_f32_qsi8d32p_qai4c32p/kai_matmul_clamp_f32_qsi8d32p1vlx4_qai4c32p4vlx4_1vlx4vl_sme2_mopa_asm.S" + +ls -la "$KERNEL_FILE" +grep -n "smopa" "$KERNEL_FILE" | head +grep -n "luti" "$KERNEL_FILE" | head +grep -n "zero {za}" "$KERNEL_FILE" | head +``` + +The output is similar to: + +```output +90: KAI_ASM_INST(0xa0840100) // smopa za0.s, p0/m, p0/m, z8.b, z4.b +(...) +88: KAI_ASM_INST(0xc08a4044) // luti4 {z4.b - z5.b}, zt0, z2[0] +(...) +81: KAI_ASM_INST(0xc00800ff) // zero {za} +``` + +If you don’t see matches, your local checkout might be on a different revision. In that case, search for the function name used in this Learning Path: + +```bash +grep -R "kai_matmul_clamp_f32_qsi8d32p1vlx4" -n kai | head +``` + +## Hands-on: verify SME2 is available (optional) + +You can complete the "source inspection" parts of this Learning Path without SME2 hardware. To run SME2 code on-device, your CPU and OS need to expose SME2. + +{{< tabpane code=true >}} + {{< tab header="Linux on Arm" language="bash" >}} +uname -m +grep -m1 -E '^Features' /proc/cpuinfo | tr ' ' '\n' | grep -E '^sme$|^sme2$' || true + {{< /tab >}} + {{< tab header="macOS" language="bash" >}} +uname -m +sysctl -a | rg -i 'sme2|sme' + {{< /tab >}} +{{< /tabpane >}} + +If neither `sme` nor `sme2` appears, you can still follow the explanation, but you should treat any "run" steps as not applicable on your device. + +## What you've learned and what's next + +You've confirmed where to find the SME2 matmul microkernel source in KleidiAI and verified whether SME2 is available on your device. You understand the scope of this Learning Path and know which parts are hands-on versus conceptual. + +Next, you'll explore how KleidiAI matmul microkernels use tiling and packing to structure matrix multiplication work. \ No newline at end of file diff --git a/content/learning-paths/mobile-graphics-and-gaming/kai_sme2_matmul_ukernel_explained/kai_matmul_kernel_overview.md b/content/learning-paths/mobile-graphics-and-gaming/kai_sme2_matmul_ukernel_explained/kai_matmul_kernel_overview.md index f7d9ebdafe..41cc48f038 100644 --- a/content/learning-paths/mobile-graphics-and-gaming/kai_sme2_matmul_ukernel_explained/kai_matmul_kernel_overview.md +++ b/content/learning-paths/mobile-graphics-and-gaming/kai_sme2_matmul_ukernel_explained/kai_matmul_kernel_overview.md @@ -1,22 +1,27 @@ --- -title: How does a KleidiAI matmual microkernel perform matrix multiplication with quantized data? -weight: 2 +title: Matmul tiling and packing +weight: 3 ### FIXED, DO NOT MODIFY layout: learningpathall --- -## How does a KleidiAI matmual microkernel perform matrix multiplication with quantized data? -Essentially, a KleidiAI matmul microkernel uses tile-based matrix multiplication(matmul) where small submatrices of the output are computed one by one. -- **mr**: number of rows of Matrix C (and Matrix A) computed at once -- **nr**: number of columns of Matrix C (and Matrix B) computed at once -- **bl**: number of elements from the K dimension processed per block at once -- **kr**: number of elements from the K dimension processed per inner step +## Matmul tiling and packing + +At a high level, a KleidiAI matmul microkernel computes the output matrix `C` in tiles. Instead of producing the full matrix `C` in one pass, it produces small submatrices (tiles) of `C` one by one. + +### Understand the tiling parameters + +Microkernels typically expose a small set of constants that define the tile shape and the inner-loop step sizes: +- **mr**: number of rows of `C` (and `A`) computed per output tile +- **nr**: number of columns of `C` (and `B`) computed per output tile +- **bl**: number of elements from the `K` dimension processed per block +- **kr**: number of elements from the `K` dimension processed per inner step The video below demonstrates how matrix multiplication is carried out using this method. ![Figure showing Tile-Based matrix multiplication with KleidiAI alt-text#center](videos/matrix_tile.gif "Tile-Based matrix multiplication with KleidiAI") -This process can be denoted with the following pseudocode, +This process can be described with the following pseudocode: ```c // RHS N LOOP for(n_idx = 0; n_idx < n; n_idx+=nr){ @@ -31,7 +36,7 @@ for(n_idx = 0; n_idx < n; n_idx+=nr){ for(k_idx = 0; k_idx < krs_in_block; k_idx +=1) { // Perform the matrix multiplication with source submatrices of size [mr, kr] and [kr, nr] // Accumulate the matrix multiplication result above into per block level result. - … + ... } // Accumulate per block level results along K dimension. When iteration on K dimension is completed,a submatrix of size [mr, nr] of the output matrix is ready } @@ -40,18 +45,49 @@ for(n_idx = 0; n_idx < n; n_idx+=nr){ //Continue computing a submatrix of size [mr, nr] of the output matrix along N dimension } ``` -In general, KleidiAI matmul microkernels implement matrix mulitplication in a similar way as the pseudocode. + +In practice, KleidiAI matmul microkernels follow this pattern, but use tight loops, explicit packing, and architecture-specific instructions. + +### Understand why packing matters KleidiAI also provides corresponding packing microkernels for the matmul microkernels, in order to make efficient contiguous memory access to the input of the matrix multiplication, reducing cache misses. -KleidiAI supports quantized matrix multiplication to speed up AI inference on Arm CPUs. Instead of multiplying full precision (FP32) matrices A and B directly, it quantizes: -- The Left Hand Source (LHS , or Left Hand Martix/activation) matrix to 8-bit integers -- The Right Hand Source( RHS, or Left Hand Matrix/weights) matrix to 4-bit or 8-bit integers +### See the quantized matmul dataflow -then packs those quantized values into memory layouts suitable for the CPU vector instructions such as Dotprod, I8MM, SME2 instructions. -Runs a microkernel that efficiently computes on packed quantized data, then scales back to floating point. +KleidiAI supports quantized matrix multiplication to speed up AI inference on Arm CPUs. Instead of multiplying full precision (FP32) matrices `A` and `B` directly, it quantizes: +- The left-hand source (LHS, activations) to 8-bit integers +- The right-hand source (RHS, weights) to 4-bit or 8-bit integers + +It then packs those quantized values into memory layouts that match the CPU’s preferred access patterns and instruction shapes (for example, DotProd, I8MM, and SME2). + +The matmul microkernel runs on the packed quantized data, accumulates into a wider type (typically INT32), and then scales back to FP32. This process can be illustrated in the following diagram, ![Figure showing quantized matrix multiplication with KleidiAI kernels alt-text#center](images/kai_matmul_kernel.jpg "Quantized matrix multiplication with KleidiAI kernel") -Please find more information in this learning path, [Accelerate Generative AI workloads using KleidiAI](https://learn.arm.com/learning-paths/cross-platform/kleidiai-explainer/). \ No newline at end of file +### Hands-on: confirm the tiling idea in the source (optional) + +Validate that `mr/nr/kr`-style constants exist and are exposed via small helper functions. + +Run the following commands from the KleidiAI repo root: + +```bash +grep -R "kai_get_mr_matmul" -n kai | head +grep -R "kai_get_nr_matmul" -n kai | head +grep -R "kai_get_kr_matmul" -n kai | head +``` + +The output is similar to: +```output +(...) +kai/ukernels/matmul/matmul_clamp_f32_qai8dxp_qsi4cxp/kai_matmul_clamp_f32_qai8dxp1x4_qsi4cxp4vlx4_1x4vl_sme2_sdot.h:52:size_t kai_get_kr_matmul_clamp_f32_qai8dxp1x4_qsi4cxp4vlx4_1x4vl_sme2_sdot(void); +``` + +You’ll use these values later when you connect the packing steps to the inner loop of the SME2 microkernel. + +For more background on KleidiAI, see [Accelerate Generative AI workloads using KleidiAI](/learning-paths/cross-platform/kleidiai-explainer/). +## What you've learned and what's next + +You've learned how KleidiAI matmul microkernels break matrix multiplication into tiles defined by mr, nr, and kr parameters. You understand why packing matters for cache efficiency and how quantization flows through the kernel. + +Next, you'll explore how SME2 INT8 MOPA instructions accelerate the core matmul computation. \ No newline at end of file diff --git a/content/learning-paths/mobile-graphics-and-gaming/kai_sme2_matmul_ukernel_explained/sme2_mpoa_matmul.md b/content/learning-paths/mobile-graphics-and-gaming/kai_sme2_matmul_ukernel_explained/sme2_mpoa_matmul.md index 6b8ec9a77c..0050962ed4 100644 --- a/content/learning-paths/mobile-graphics-and-gaming/kai_sme2_matmul_ukernel_explained/sme2_mpoa_matmul.md +++ b/content/learning-paths/mobile-graphics-and-gaming/kai_sme2_matmul_ukernel_explained/sme2_mpoa_matmul.md @@ -1,23 +1,29 @@ --- -title: How are SME2 INT8 Outer Product Accumulate instructions used in a matrix multiplication? -weight: 3 +title: SME2 INT8 MOPA for matmul +weight: 4 ### FIXED, DO NOT MODIFY layout: learningpathall --- -## How are SME2 INT8 Outer Product Accumulate instructions used in a matrix multiplication? -The INT8 Outer Product Accumulate instructions calculate the sum of four INT8 outer products, widening results into INT32, then the result is destructively added to the destination tile. +## SME2 INT8 MOPA for matmul + +SME2 INT8 MOPA (outer product accumulate) instructions compute the sum of outer products for INT8 inputs, widen the results into INT32, and accumulate into an SME2 tile in the ZA storage. + +### Understand what MOPA computes ![Figure showing SME2 INT8 MOPA instruction alt-text#center](images/sme2_mopa.jpg "SME2 INT8 MOPA instruction") -When SME2 SVL is 512-bit, each input register (Zn.B, Zm.B) is treated as a matrix of 16x4 INT8 elements, as if each block of four contiguous elements were transposed. +When the SME2 streaming vector length (SVL) is 512 bits, each input register (`Zn.B`, `Zm.B`) can be treated as a matrix of 16Γ—4 INT8 elements (with a hardware-defined arrangement). - The first source, Zn.B contains a 16x4 sub-matrix of 8-bit integer values. - The second source, Zm.B, contains a 16 x4 sub-matrix of 8-bit integer values. -- The INT8 MOPA instruction calculates a 16x 16 widened 32-bit integer sum of outer products, which is then destructively added to the 32-bit integer destination tile, ZAda. +- The INT8 MOPA instruction calculates a 16Γ—16 widened INT32 sum of outer products, then destructively adds it to the INT32 destination tile, `ZAda`. The video below shows how SME2 INT8 Outer Product Accumulate instructions are used for matrix multiplication. ![Figure showing Matrix Multiplication with 1VLx1VL SME2 MOPA alt-text#center](videos/matrix_mopa_sme2_1vl.gif "Matrix Multiplication with 1VLx1VL SME2 MOPA") -To calculate the result of a 16x16 sub-matrix in matrix C (element type: INT32): + +### Map MOPA to a 16Γ—16 output tile + +To calculate the result of a 16Γ—16 submatrix in matrix `C` (element type: INT32): First, - a 16x4 sub-matrix in matrix A (element type: INT8) is loaded to a SME2 Z register, @@ -26,18 +32,42 @@ First, Then, the SME2 INT8 MOPA instruction uses the data from these two Z registers to perform the outer product operation and accumulates the results into the ZA tile, which holds the 16x16 sub-matrix of matrix C, thus obtaining an intermediate result for this 16x16 sub-matrix. -Iterate over the K dimension, repeatedly loading 16x4 submatrices from matrix A and 4Γ—16 submatrices from matrix B. For each step, use the SME2 INT8 MPOA instruction to compute outer products and accumulate the results into the same ZA tile. After completing the iteration over K, this ZA tile holds the final values for the corresponding 16Γ—16 submatrix of matrix C. Finally, store the contents of the ZA tile back to memory. +Iterate over the K dimension, repeatedly loading 16x4 submatrices from matrix A and 4Γ—16 submatrices from matrix B. For each step, use the SME2 INT8 MOPA instruction to compute outer products and accumulate the results into the same ZA tile. After completing the iteration over K, this ZA tile holds the final values for the corresponding 16Γ—16 submatrix of matrix C. Finally, store the contents of the ZA tile back to memory. + +Apply the same process across the `M` and `N` dimensions to compute the full output matrix. Apply the same process to all 16x16 sub-matrices in matrix C to complete the entire matrix computation. -To improve performance, we can pipeline four MOPA instructions and fully utilize four ZA tiles in ZA storage, each MOPA instruction uses one ZA tile. -The video below demonstrates how the four MOPA instructions are used to perfrom matrix multiplication of one 16x4 submatrix from matrix A and four 4x16 submatrices from matrix B in a single iteration. This approach can be referred to as 1VLx4VL, +### See the 1VLΓ—4VL pattern + +To improve throughput, a microkernel can pipeline multiple MOPA instructions so that it accumulates into multiple ZA tiles in parallel. One common pattern is 1VLΓ—4VL: one VL slice from the LHS multiplies four VL slices from the RHS in the same inner-loop iteration, accumulating into four ZA tiles. + +The video below demonstrates how four pipelined MOPA instructions perform matrix multiplication of one 16Γ—4 submatrix from matrix A and four 4Γ—16 submatrices from matrix B in a single iteration (1VLΓ—4VL). ![Figure showing Matrix Multiplication with 1VLx4VL SME2 MOPA alt-text#center](videos/1vlx4vl_sme2_mopa.gif "Matrix Multiplication with 1VLx4VL SME2 MOPA") The intermediate result of 4x16x16 output submatrix is held in four ZA.S tiles. +### Hands-on: locate MOPA in the SME2 microkernel (optional) + +If you cloned KleidiAI earlier, you can confirm that the SME2 matmul microkernel uses MOPA by searching for `smopa` instructions in the kernel source. + +From the KleidiAI repo root: + +```bash +KERNEL_FILE="kai/ukernels/matmul/matmul_clamp_f32_qsi8d32p_qai4c32p/kai_matmul_clamp_f32_qsi8d32p1vlx4_qai4c32p4vlx4_1vlx4vl_sme2_mopa_asm.S" +grep -n "smopa" "$KERNEL_FILE" | head +``` + +You’ll connect these `smopa` sites to the "load β†’ dequantize β†’ MOPA β†’ dequantize" flow later in the example walk-through. + You can find more information about SME2 MOPA here, -- [part 1 Arm Scalable Matrix Extension Introduction](https://developer.arm.com/community/arm-community-blogs/b/architectures-and-processors-blog/posts/arm-scalable-matrix-extension-introduction) -- [part 2 Arm Scalable Matrix Extension Instructions](https://developer.arm.com/community/arm-community-blogs/b/architectures-and-processors-blog/posts/arm-scalable-matrix-extension-introduction-p2) -- [part4 Arm SME2 Introduction](https://developer.arm.com/community/arm-community-blogs/b/architectures-and-processors-blog/posts/part4-arm-sme2-introduction) +- [Part 1 Arm Scalable Matrix Extension Introduction](https://developer.arm.com/community/arm-community-blogs/b/architectures-and-processors-blog/posts/arm-scalable-matrix-extension-introduction) +- [Part 2 Arm Scalable Matrix Extension Instructions](https://developer.arm.com/community/arm-community-blogs/b/architectures-and-processors-blog/posts/arm-scalable-matrix-extension-introduction-p2) +- [Part 4 Arm SME2 Introduction](https://developer.arm.com/community/arm-community-blogs/b/architectures-and-processors-blog/posts/part4-arm-sme2-introduction) + +## What you've learned and what's next + +You've learned how SME2 INT8 MOPA instructions compute outer products for matrix multiplication and how the 1VLΓ—4VL pattern pipelines four MOPA instructions to improve throughput. You can now connect MOPA operations to output tile shapes. + +Next, you'll decode the specific SME2 matmul microkernel name and understand what the format tags reveal about input layouts and tile dimensions. \ No newline at end of file diff --git a/content/learning-paths/mobile-graphics-and-gaming/kai_sme2_matmul_ukernel_explained/summary.md b/content/learning-paths/mobile-graphics-and-gaming/kai_sme2_matmul_ukernel_explained/summary.md index 469cd9030a..7b6ce09084 100644 --- a/content/learning-paths/mobile-graphics-and-gaming/kai_sme2_matmul_ukernel_explained/summary.md +++ b/content/learning-paths/mobile-graphics-and-gaming/kai_sme2_matmul_ukernel_explained/summary.md @@ -1,10 +1,23 @@ --- -title: Summary -weight: 8 +title: Wrap-up +weight: 9 ### FIXED, DO NOT MODIFY layout: learningpathall --- -## Summary -This learning path vividly explains how an SME2-optimized KleidiAI microkernel performs quantization and packing of the RHS and LHS, and how it leverages the powerful SME2 MOPA instructions to enhance matrix multiplication performance. We hope this learning path helps developers learn how to integrate the KleidiAI microkernel into their ML/AI frameworks or applications, or to design their own SME2-optimized kernels, thus fully utilizing the potential of SME2. \ No newline at end of file +## Summary: SME2 matmul microkernel understanding + +You've successfully navigated one of the more complex areas of AI inference optimization - understanding how low-level SME2 instructions accelerate quantized matrix multiplication. You've walked through how an SME2-optimized KleidiAI matmul microkernel: +- Converts weights into a kernel-friendly packed RHS layout +- Quantizes and packs activations into a packed LHS layout +- Uses SME2 INT8 MOPA instructions (`smopa`) plus LUT-based dequantization (`luti4`) to compute a 1VLΓ—4VL output tile efficiently +- Dequantizes back to FP32 so the result matches the surrounding FP32 computation (within expected quantization error) + +You traced the complete dataflow using a concrete GGML Q4_0 example and can now connect high-level AI frameworks to the Arm hardware features that make them fast. + +If you completed the optional hands-on checks, you've verified where the key SME2 instructions appear in the microkernel source β€” valuable experience for anyone working with performance-critical code on Arm platforms. + +You're now equipped to apply the same approach to real workloads: +- Use [the llama.cpp performance Learning Path](/learning-paths/mobile-graphics-and-gaming/performance_llama_cpp_sme2/) to build and profile llama.cpp with KleidiAI on an SME2-capable device +- Use [the ONNX Runtime performance Learning Path](/learning-paths/mobile-graphics-and-gaming/performance_onnxruntime_kleidiai_sme2/) to see how similar ideas apply in ONNX Runtime \ No newline at end of file diff --git a/content/learning-paths/mobile-graphics-and-gaming/kai_sme2_matmul_ukernel_explained/the_sme2_matmul_microkernel.md b/content/learning-paths/mobile-graphics-and-gaming/kai_sme2_matmul_ukernel_explained/the_sme2_matmul_microkernel.md index 6843b44fa2..8d0d520ed0 100644 --- a/content/learning-paths/mobile-graphics-and-gaming/kai_sme2_matmul_ukernel_explained/the_sme2_matmul_microkernel.md +++ b/content/learning-paths/mobile-graphics-and-gaming/kai_sme2_matmul_ukernel_explained/the_sme2_matmul_microkernel.md @@ -1,31 +1,68 @@ --- -title: What is the sme2 lvlx4vl microkernel? -weight: 4 +title: Decode the SME2 matmul microkernel +weight: 5 ### FIXED, DO NOT MODIFY layout: learningpathall --- -## What is the sme2 lvlx4vl microkernel? -We use a KleidiAI microkernel, *kai_matmul_clamp_f32_qsi8d32p1vlx4_qsi4c32p4vlx4_1vlx4vl_sme2_mopa*, to explain KleidiAI SME2 microkernels in detail. It is referred as β€˜the SME2 matmul microkernel’ in this learning path onwards, unless otherwise noted. +## Decode the SME2 matmul microkernel -β€œ_1vlx4vl” in the name indicates that, in a single inner loop iteration, it computes an intermediate result for a 1VL x 4VL submatrix (one SME2 Streaming Vector Length x four SME2 Streaming Vector Length) of the ouput matrix. Assuming the SME2 SVL is 512 bits, it is a 16 x 64 (512/sizeof(FP32)) x (4 x 512/sizeof(FP32)) submatrix. +This Learning Path uses one concrete KleidiAI microkernel to explain SME2 matmul microkernels in detail: + +*kai_matmul_clamp_f32_qsi8d32p1vlx4_qsi4c32p4vlx4_1vlx4vl_sme2_mopa* + +In the rest of this Learning Path, this is referred to as *the SME2 matmul microkernel* (unless noted otherwise). + +### Decode 1vlx4vl + +`_1vlx4vl` indicates that, in a single inner-loop iteration, the kernel computes an intermediate result for a 1VL Γ— 4VL submatrix (one SME2 streaming vector length Γ— four SME2 streaming vector lengths) of the output matrix. + +If you assume an SME2 SVL of 512 bits, the FP32 shape is 16 Γ— 64: +- 1VL rows: `512 / 8 / 4 = 16` FP32 elements +- 4VL columns: `4 Γ— 16 = 64` FP32 elements + +### See the pipelined MOPA pattern + +To improve throughput, the kernel pipelines four MOPA instructions so it can accumulate into four ZA tiles in parallel (one ZA tile per MOPA). + +The same pattern applies here as shown by the video in the [SME2 INT8 MOPA section](/learning-paths/mobile-graphics-and-gaming/kai_sme2_matmul_ukernel_explained/sme2_mpoa_matmul/). -To improve performance, we can pipeline four MOPA instructions and fully utilize four ZA tiles in ZA storage, each MOPA instruction uses one ZA tile. -The video below demonstrates how the four MOPA instructions are used to perfrom matrix multiplication of one 16x4 submatrix (1VL) from matrix A and four 4x16 submatrices from matrix B (4VL) in a single iteration. ![Figure showing Matrix Multiplication with 1VLx4VL SME2 MOPA alt-text#center](videos/1vlx4vl_sme2_mopa.gif "Matrix Multiplication with 1VLx4VL SME2 MOPA") + +The animation demonstrates how four pipelined MOPA instructions multiply one 16Γ—4 submatrix (1VL) from matrix A by four 4Γ—16 submatrices (4VL) from matrix B in a single iteration. + The intermediate result of 4x16x16 output submatrix is held in four ZA.S tiles. -β€œqsi8d32p1vlx4” in the name indicates that it expects the LHS with a layout of [M, K] to be symmetrically quantized into signed INT8 type within blocks of 32 elements. -The entire quantized LHS is then divided into submatrices of size 1VL Γ— 4 (since the SME2 SVL is set as 512 bits, it is 16 Γ— 4). Then, each submatrix is packed row-wise into a contiguous memory layout, all the submatrices are packed in this way one after another. So that when using the packed LHS in the SME2 matmul microkernel, memory accesses are to contiguous addresses, improving cache locality. +### Decode the input formats + +The table below decodes the input and output format tags in the microkernel name. + +| Format tag | Meaning | Why it matters | +| --- | --- | --- | +| `qsi8d32p1vlx4` | LHS layout is [M, K], symmetrically quantized to signed INT8 in blocks of 32; packed into 1VL Γ— 4 submatrices (16 Γ— 4 for a 512-bit SVL). | Row-wise packing makes LHS loads contiguous, which improves cache locality. | +| `qsi4c32p4vlx4` | RHS layout is [N, K], symmetrically quantized to signed INT4 in blocks of 32; packed into 4VL Γ— 4 submatrices (4 Γ— 16 Γ— 4 for a 512-bit SVL). | INT4 packs two values per byte, and SME2 LUTI expands them to INT8 for MOPA. | +| `_f32_` | Output matrix is FP32; the INT32 accumulation from MOPA is dequantized to FP32. | Preserves FP32 results while using INT8/INT4 arithmetic in the inner loop. | + +Sometimes, the original LHS or RHS doesn’t match the quantization and packing requirements of the SME2 matmul microkernel. In that case, your software needs to quantize and pack the LHS and RHS first. + +### Hands-on: compute the FP32 tile shape for your assumed SVL (optional) + +If you want to sanity-check the `1VL` size used in the diagrams, calculate how many FP32 values fit in one SVL. + +For an assumed 512-bit SVL: + +```bash +SVL_BITS=512 +FP32_PER_VL=$((SVL_BITS / 8 / 4)) +echo "FP32 per VL: ${FP32_PER_VL}" +echo "1VLx4VL tile: ${FP32_PER_VL}x$((4 * FP32_PER_VL))" +``` -β€œqsi4c32p4vlx4” in its name indicates that the SME2 matmul microkernel expects the RHS with a layout of [N, K] to be symmetrically quantized into signed INT4 type within blocks of 32 elements. -The entire quantized RHS is then divided into submatrices of size 4VL Γ— 4 (since the SME2 SVL is set as 512 bits, it is 4x16Γ— 4). Each submatrix is packed row-wise into a contiguous memory layout. Since the quantization type is INT4, each byte contains two INT4 elements. In the SME2 matmul microkernel, the SME2 LUTI instructions efficiently dequantize INT4 elements into INT8 type, thereby enabling fast matrix multiplication with SME2 INT8 MOPA instructions. +If your target device uses a different SVL, the same formulas still apply. -β€œ_f32_” in its name indicates that the SME2 matmul microkernel outputs FP32 result matrix. The INT32 result produced by SME2 INT8 MOPA instructions has to be dequantized back to FP32 type. +## What you've learned and what's next -Sometimes, the original LHS or RHS may not conform to the quantization and packing format requirement of the SME2 matmul microkernel. The software needs to quantize and pack the LHS and RHS appropriately first. +You've decoded the SME2 matmul microkernel name and understand what 1VLΓ—4VL means for tile dimensions. You learned how the input format tags describe quantization and packing requirements for LHS and RHS. -Next, we will take llama.cpp and the Llama-3.2-3B-Q4_0.gguf model for example to demonstrate, -- how to quantize and pack the LHS and RHS -- perform matrix multiplication using the SME2 matmul microkernel \ No newline at end of file +Next, you'll walk through a concrete example starting with RHS weight repacking from GGML Q4_0 format. \ No newline at end of file diff --git a/content/learning-paths/mobile-graphics-and-gaming/litert-sme/_index.md b/content/learning-paths/mobile-graphics-and-gaming/litert-sme/_index.md index 222c9c3244..4c57d6bf70 100644 --- a/content/learning-paths/mobile-graphics-and-gaming/litert-sme/_index.md +++ b/content/learning-paths/mobile-graphics-and-gaming/litert-sme/_index.md @@ -23,9 +23,11 @@ subjects: ML armips: - Cortex-A - Cortex-X + - Arm C1 tools_software_languages: - C - Python + - SME2 operatingsystems: - Android diff --git a/content/learning-paths/mobile-graphics-and-gaming/onnx/01_fundamentals.md b/content/learning-paths/mobile-graphics-and-gaming/onnx/01_fundamentals.md index 0dcca61c93..4902e1202f 100644 --- a/content/learning-paths/mobile-graphics-and-gaming/onnx/01_fundamentals.md +++ b/content/learning-paths/mobile-graphics-and-gaming/onnx/01_fundamentals.md @@ -1,103 +1,114 @@ --- -# User change -title: "ONNX Fundamentals" +title: "Understand ONNX fundamentals and architecture" weight: 2 layout: "learningpathall" --- -The goal of this tutorial is to provide developers with a practical, end-to-end pathway for working with Open Neural Network Exchange (ONNX) in real-world scenarios. Starting from the fundamentals, we will build a simple neural network model in Python, export it to the ONNX format, and demonstrate how it can be used for both inference and training on Arm64 platforms. Along the way, we will cover model optimization techniques such as layer fusion, and conclude by deploying the optimized model into a fully functional Android application. By following this series, you will gain not only a solid understanding of ONNX’s philosophy and ecosystem but also the hands-on skills required to integrate ONNX into your own projects from prototyping to deployment. +## Overview -In this first step, we will introduce the ONNX standard and explain why it has become a cornerstone of modern machine learning workflows. You will learn what ONNX is, how it represents models in a framework-agnostic format, and why this matters for developers targeting different platforms such as desktops, Arm64 devices, or mobile environments. We will also discuss the role of ONNX Runtime as the high-performance engine that brings these models to life, enabling efficient inference and even training across CPUs, GPUs, and specialized accelerators. Finally, we will outline the typical ONNX workflow, from training in frameworks like PyTorch or TensorFlow, through export and optimization, to deployment on edge and Android devices, which we will gradually demonstrate throughout the tutorial. +This Learning Path provides a practical, end-to-end introduction to working with Open Neural Network Exchange (ONNX) in real-world scenarios. You will build a simple neural network model in Python, export it to the ONNX format, and run inference on Arm64 platforms using ONNX Runtime. Along the way, you will learn about model optimization techniques such as layer fusion, and conclude by deploying the optimized model into a fully functional Android application. By the end of this learning path, you will understand both the conceptual foundations of ONNX and the practical steps required to move a model from training to efficient deployment on Arm-based systems. -## What is ONNX -The ONNX is an open standard for representing machine learning models in a framework-independent format. Instead of being tied to the internal model representation of a specific frameworkβ€”such as PyTorch, TensorFlow, or scikit-learnβ€”ONNX provides a universal way to describe models using a common set of operators, data types, and computational graphs. +## What is ONNX? -At its core, an ONNX model is a directed acyclic graph (DAG) where nodes represent mathematical operations (e.g., convolution, matrix multiplication, activation functions) and edges represent tensors flowing between these operations. This standardized representation allows models trained in one framework to be exported once and executed anywhere, without requiring the original framework at runtime. +ONNX (Open Neural Network Exchange) is an open standard for representing machine learning models as a framework-independent intermediate representation (IR). Instead of relying on the internal model format of a specific framework--such as PyTorch or TensorFlow--ONNX defines a common computational graph structure, standardized operators, and well-specified data types. -ONNX was originally developed by Microsoft and Facebook to address a growing need in the machine learning community: the ability to move models seamlessly between training environments and deployment targets. Today, it is supported by a wide ecosystem of contributors and hardware vendors, making it the de facto choice for interoperability and cross-platform deployment. +At its core, an ONNX model is a directed acyclic graph (DAG). Nodes represent mathematical operations (such as Conv, MatMul, or Relu), while edges represent tensors flowing between these operations. The model file stores both the graph structure and the trained parameters (weights), making it self-contained and executable without the original training framework. + +Portability depends on operator support and opset compatibility within the chosen runtime. However, ONNX reduces the friction of moving models across frameworks and hardware targets by standardizing operator semantics and graph representation. + +ONNX was originally developed by Microsoft and Facebook to address a growing need in the machine learning community: the ability to move models seamlessly between training environments and deployment targets. Today, it is supported by a wide ecosystem of contributors and hardware vendors, making it a widely adopted standard for model exchange and deployment. For developers, this means flexibility. You can train your model in PyTorch, export it to ONNX, run it with ONNX Runtime on an Arm64 device such as a Raspberry Pi, and later deploy it inside an Android application without rewriting the model. This portability is the main reason ONNX has become a central building block in modern AI workflows. -A useful way to think of ONNX is to compare it to a PDF for machine learning models. Just as a PDF file ensures that a document looks the same regardless of whether you open it in Adobe Reader, Preview on macOS, or a web browser, ONNX ensures that a machine learning model behaves consistently whether you run it on a server GPU, a Raspberry Pi, or an Android phone. It is this β€œwrite once, run anywhere” principle that makes ONNX especially powerful for developers working across diverse hardware platforms. +A helpful analogy is to think of ONNX as a β€œPDF for machine learning models.” Just as a PDF preserves the structure of a document across operating systems and viewers, ONNX preserves the structure and semantics of a trained model across frameworks and hardware platforms. + +Importantly, ONNX is also extensible. Developers and hardware vendors can define custom operators or operator domains when the standard operator set is not sufficient. This allows innovation and hardware-specific acceleration while maintaining compatibility with the broader ONNX ecosystem. + +## Why ONNX matters + +Modern machine learning workflows span multiple frameworks and deployment targets. A model might be trained in PyTorch on a GPU workstation, validated in the cloud, and ultimately deployed on an Arm64-based edge device or Android smartphone. Without a common representation, moving models between these environments would require complex and error-prone conversions. + +ONNX addresses this challenge by acting as a universal exchange format that separates model development from deployment. + +The key reasons ONNX matters are: +- Interoperability: ONNX decouples training from inference. Models trained in PyTorch or TensorFlow can be exported into a common format and executed in a different runtime environment without embedding the original framework. +- Performance: ONNX Runtime includes highly optimized execution backends, supporting hardware acceleration through Arm NEON, CUDA, DirectML, and Android NNAPI. This means the same model can run efficiently across a wide spectrum of hardware. +- Portability: a single `.onnx` model file can be deployed across Arm-based cloud servers, embedded Arm devices, and mobile applications, provided the required operators are supported by the target runtime. +- Ecosystem: the ONNX Model Zoo and broad industry adoption make it easier to reuse validated architectures across platforms. +- Extensibility: custom operators and execution providers allow researchers and hardware vendors to extend ONNX without breaking compatibility with the broader ecosystem. + +## ONNX model structure -At the same time, ONNX is not a closed box. Developers can extend the format with custom operators or layers when standard ones are not sufficient. This flexibility makes it possible to inject novel research ideas, proprietary operations, or hardware-accelerated kernels into an ONNX model while still benefiting from the portability of the core standard. In other words, ONNX gives you both consistency across platforms and extensibility for innovation. +An ONNX model is a complete description of a computation graph, not just a collection of weights. Understanding its structure clarifies why it is both portable and extensible. -## Why ONNX Matters -Machine learning today is not limited to one framework or one device. A model might be trained in PyTorch on a GPU workstation, tested in TensorFlow on a cloud server, and then finally deployed on an Arm64-based edge device or Android phone. Without a common standard, moving models between these environments would be complex, error-prone, and often impossible. ONNX solves this problem by acting as a universal exchange format, ensuring that models can flow smoothly across the entire development and deployment pipeline. +### Core components -The main reasons ONNX matters are: -1. Interoperability – ONNX eliminates framework lock-in. You can train in PyTorch, validate in TensorFlow, and deploy with ONNX Runtime on almost any device, from servers to IoT boards. -2. Performance – ONNX Runtime includes highly optimized execution backends, supporting hardware acceleration through Arm NEON, CUDA, DirectML, and Android NNAPI. This means the same model can run efficiently across a wide spectrum of hardware. -3. Portability – Once exported to ONNX, the model can be deployed to Arm64 devices (like Raspberry Pi or AWS Graviton servers) or even embedded in an Android app, without rewriting the code. -4. Ecosystem – The ONNX Model Zoo provides ready-to-use, pre-trained models for vision, NLP, and speech tasks, making it easy to start from state-of-the-art baselines. -5. Extensibility – Developers can inject their own layers or custom operators when the built-in operator set is not sufficient, enabling innovation while preserving compatibility. +At a high level, an ONNX model consists of: -In short, ONNX matters because it turns the fragmented ML ecosystem into a cohesive workflow, empowering developers to focus on building applications rather than wrestling with conversion scripts or hardware-specific code. +- Graph: the core directed acyclic graph (DAG). Nodes correspond to operations (such as Conv, Relu, or MatMul), and edges represent tensors flowing between nodes. +- Initializers: these store learned parameters such as weights and biases. Initializers are embedded directly in the model file. +- Opset (Operator Set): versioned collection of operator definitions. Opsets define the exact semantics of each operator and ensure compatibility between model exporters and runtimes. Selecting an appropriate opset version during export is critical for deployment compatibility. +- Metadata: these describe tensor shapes, data types, input/output signatures, and optional annotations such as author information or framework version. +- Operator domains: namespaces that allow standard and custom operators to coexist without conflict. -## ONNX Model Structure -An ONNX model is more than just a collection of weightsβ€”it is a complete description of the computation graph that defines how data flows through the network. Understanding this structure is key to seeing why ONNX is both portable and extensible. +### Graph representation -At a high level, an ONNX model consists of three main parts: -1. Graph, which is the heart of the model, represented as a directed acyclic graph (DAG). In this graph nodes correspond to operations (e.g., Conv, Relu, MatMul), edges represent tensors flowing between nodes, carrying input and output data. -2. Opset (Operator Set), which is a versioned collection of supported operations. Opsets guarantee that models exported with one framework will behave consistently when loaded by another, as long as the same opset version is supported. -3. Metadata, which contains information about inputs, outputs, tensor shapes, and data types. Metadata can also include custom annotations such as the model author, domain, or framework version. +The structured design allows ONNX to describe anything from a simple logistic regression to a deep convolutional neural network. For example, a single ONNX graph might define: -This design allows ONNX to describe anything from a simple logistic regression to a deep convolutional neural network. For example, a single ONNX graph might define: -* An input tensor representing a camera image. -* A sequence of convolution and pooling layers. -* Fully connected layers leading to classification probabilities. -* An output tensor with predicted labels. +* An input tensor representing a camera image +* A sequence of convolution and pooling layers +* Fully connected layers leading to classification probabilities +* An output tensor with predicted labels -Because the ONNX format is based on a standardized graph representation, it is both human-readable (with tools like Netron for visualization) and machine-executable (parsed directly by ONNX Runtime or other backends). +You can visualize the graph using tools such as Netron, while runtimes such as ONNX Runtime parse and execute it efficiently. -Importantly, ONNX models are not static. Developers can insert, remove, or replace nodes in the graph, making it possible to add new layers, prune unnecessary ones, or fuse operations for optimization. This graph-level flexibility is what enables many of the performance improvements we’ll explore later in this tutorial, such as layer fusion and quantization. +Because the model is graph-based, you can modify it programmatically--adding, removing, or replacing nodes. Graph-level flexibility enables optimization techniques such as operator fusion, constant folding, and quantization, which you will explore later in this Learning Path. ## ONNX Runtime -While ONNX provides a standard way to represent models, it still needs a high-performance engine to actually execute them. This is where ONNX Runtime (ORT) comes in. ONNX Runtime is the official, open-source inference engine for ONNX models, designed to run them quickly and efficiently across a wide variety of hardware. -At its core, ONNX Runtime is optimized for speed, portability, and extensibility: -1. Cross-platform support. ORT runs on Windows, Linux, and macOS, as well as mobile platforms like Android and iOS. It supports both x86 and Arm64 architectures, making it suitable for deployment from cloud servers to edge devices such as Raspberry Pi boards and smartphones. +While ONNX defines how a model is represented, ONNX Runtime (ORT) executes that model efficiently. ORT is the official open-source runtime for ONNX models and is optimized for performance, portability, and modular hardware acceleration. -2. Hardware acceleration. ORT integrates with a wide range of execution providers (EPs) that tap into hardware capabilities: -* Arm Kleidi kernels accelerated with Arm NEON, SVE2, and SME2 instructions for efficient CPU execution on Arm64. -* CUDA for NVIDIA GPUs. -* DirectML for Windows. -* NNAPI on Android, enabling direct access to mobile accelerators (DSPs, NPUs). +### Key characteristics -3. Inference and training. ONNX Runtime also supports training and fine-tuning, making it possible to use the same runtime across the entire ML lifecycle. +ONNX Runtime provides: -4. Optimization built in. ORT can automatically apply graph optimizations such as constant folding, operator fusion, or memory layout changes to squeeze more performance out of your model. +**Cross-platform support** – ORT runs on Windows, Linux, and macOS, as well as mobile platforms like Android and iOS. It supports both x86 and Arm64 architectures, making it suitable for deployment from cloud servers to edge devices such as Raspberry Pi boards and smartphones. -For developers, this means you can take a model trained in PyTorch, export it to ONNX, and then run it with ONNX Runtime on virtually any deviceβ€”without worrying about the underlying hardware differences. The runtime abstracts away the complexity, choosing the best available execution provider for your environment. +**Hardware acceleration** – ORT integrates with a wide range of execution providers (EPs) that tap into hardware capabilities: +* Arm Kleidi kernels accelerated with Arm NEON, SVE2, and SME2 instructions for efficient CPU execution on Arm64 +* CUDA for NVIDIA GPUs +* DirectML for Windows +* NNAPI on Android, enabling direct access to mobile accelerators (DSPs, NPUs) -This flexibility makes ONNX Runtime a powerful bridge between training frameworks and deployment targets, and it is the key technology that allows ONNX models to run effectively on Arm64 platforms and Android devices. +**Inference focus** – ONNX Runtime includes optional training capabilities, but is most widely used for high-performance inference in production and edge deployments. -## How ONNX Fits into the Workflow +**Built-in optimizations** – ORT can automatically apply graph optimizations such as constant folding, operator fusion, or memory layout changes to improve model performance. -One of the biggest advantages of ONNX is how naturally it integrates into a developer’s machine learning workflow. Instead of locking you into a single framework from training to deployment, ONNX provides a bridge that connects different stages of the ML lifecycle. +By abstracting hardware differences behind execution providers, ONNX Runtime enables a single ONNX model to run across heterogeneous systems while leveraging platform-specific optimizations. +## How ONNX fits into the workflow +One of ONNX’s greatest strengths is how naturally it integrates into a modern ML workflow. Instead of locking developers into a single framework from training to deployment, ONNX acts as a bridge between stages. A typical ONNX workflow looks like this: -1. Train the model. You first use your preferred framework (e.g., PyTorch, TensorFlow, or scikit-learn) to design and train a model. At this stage, you benefit from the flexibility and ecosystem of the framework of your choice. -2. Export to ONNX. Once trained, the model is exported into the ONNX format using built-in converters (such as torch.onnx.export for PyTorch). This produces a portable .onnx file describing the network architecture, weights, and metadata. -3. Run inference with ONNX Runtime. The ONNX model can now be executed on different devices using ONNX Runtime. On Arm64 hardware, ONNX Runtime can take advantage of Arm Kleidi kernels accelerated with NEON, SVE2, and SME2 instructions, while on Android devices it can leverage NNAPI to access mobile accelerators (where available). -4. Optimize the model. Apply graph optimizations like layer fusion, constant folding, or quantization to improve performance and reduce memory usage, making the model more suitable for edge and mobile deployments. -5. Deploy. Finally, the optimized ONNX model is packaged into its target environment. This could be an Arm64-based embedded system (e.g., Raspberry Pi), a server powered by Arm CPUs (e.g., AWS Graviton), or an Android application distributed via the Play Store. +- Train the model: you first use your preferred framework (e.g., PyTorch, TensorFlow, or scikit-learn) to design and train a model. At this stage, you benefit from the flexibility and ecosystem of the framework of your choice. +- Export to ONNX: once trained, the model is exported into the ONNX format using built-in converters (such as torch.onnx.export for PyTorch). This produces a portable .onnx file describing the network architecture, weights, and metadata. +- Run inference with ONNX Runtime: the ONNX model can now be executed on different devices using ONNX Runtime. On Arm64 hardware, ONNX Runtime can take advantage of Arm Kleidi kernels accelerated with NEON, SVE2, and SME2 instructions, while on Android devices it can leverage NNAPI to access mobile accelerators (where available). +- Optimize the model: apply graph optimizations like layer fusion, constant folding, or quantization to improve performance and reduce memory usage, making the model more suitable for edge and mobile deployments. +- Deploy: finally, the optimized ONNX model is packaged into its target environment. This could be an Arm64-based embedded system (e.g., Raspberry Pi), a server powered by Arm CPUs (e.g., AWS Graviton), or an Android application distributed via the Play Store. -This modularity means developers are free to mix and match the best tools for each stage: train in PyTorch, optimize with ONNX Runtime, and deploy to Androidβ€”all without rewriting the model. By decoupling training from inference, ONNX enables efficient workflows that span from research experiments to production-grade applications. +This modularity means developers are free to mix and match the best tools for each stage: train in PyTorch, optimize with ONNX Runtime, and deploy to Android--all without rewriting the model. By decoupling training from inference, ONNX enables efficient workflows that span from research experiments to production-grade applications. -## Example Use Cases +## Example use cases ONNX is already widely adopted in real-world applications where portability and performance are critical. A few common examples include: -1. Computer Vision at the Edge – Running an object detection model (e.g., YOLOv5 exported to ONNX) on a Raspberry Pi 4 or NVIDIA Jetson, enabling low-cost cameras to detect people, vehicles, or defects in real time. -2. Mobile Applications – Deploying face recognition or image classification models inside an Android app using ONNX Runtime Mobile, with NNAPI acceleration for efficient on-device inference. -3. Natural Language Processing (NLP) – Running BERT-based models on Arm64 cloud servers (like AWS Graviton) to provide fast, low-cost inference for chatbots and translation services. -4. Healthcare Devices – Using ONNX to integrate ML models into portable diagnostic tools or wearable sensors, where Arm64 processors dominate due to their low power consumption. -5. Cross-platform Research to Production – Training experimental architectures in PyTorch, exporting them to ONNX, and validating them across different backends to ensure consistent performance. -6. AI Accelerator Integration – ONNX is especially useful for hardware vendors building custom AI accelerators. Since accelerators often cannot support the full range of ML operators, ONNX’s extensible operator model allows manufacturers to plug in custom kernels where hardware acceleration is available, while gracefully falling back to the standard runtime for unsupported ops. This makes it easier to adopt new hardware without rewriting entire models. +- Computer Vision at the Edge: running an object detection model (e.g., YOLOv5 exported to ONNX) on a Raspberry Pi 4 or NVIDIA Jetson, enabling low-cost cameras to detect people, vehicles, or defects in real time. +- Mobile Applications: deploying face recognition or image classification models inside an Android app using ONNX Runtime Mobile, with NNAPI acceleration for efficient on-device inference. +- Natural Language Processing (NLP): running BERT-based models on Arm64 cloud servers (like AWS Graviton) to provide fast, low-cost inference for chatbots and translation services. +- Healthcare Devices: using ONNX to integrate ML models into portable diagnostic tools or wearable sensors, where Arm64 processors dominate due to their low power consumption. +- Cross-platform Research to Production: training experimental architectures in PyTorch, exporting them to ONNX, and validating them across different backends to ensure consistent performance. +- AI Accelerator Integration: ONNX is especially useful for hardware vendors building custom AI accelerators. Since accelerators often cannot support the full range of ML operators, ONNX’s extensible operator model allows manufacturers to plug in custom kernels where hardware acceleration is available, while gracefully falling back to the standard runtime for unsupported ops. This makes it easier to adopt new hardware without rewriting entire models. -## Summary -In this section, we introduced ONNX as an open standard for representing machine learning models across frameworks and platforms. We explored its model structureβ€”graphs, opsets, and metadataβ€”and explained the role of ONNX Runtime as the high-performance execution engine. We also showed how ONNX fits naturally into the ML workflow: from training in PyTorch or TensorFlow, to exporting and optimizing the model, and finally deploying it on Arm64 or Android devices. +## What you've learned and what's next -A useful way to think of ONNX is as the PDF of machine learning modelsβ€”a universal, consistent format that looks the same no matter where you open it, but with the added flexibility to inject your own layers and optimizations. +In this section, you learned about the fundamentals of ONNX and how it enables portable machine learning workflows. You explored the ONNX model structure as a directed acyclic graph, understood how ONNX Runtime executes models efficiently on Arm64 platforms with hardware acceleration, and saw how ONNX fits into the training-to-deployment workflow across different frameworks and devices. -Beyond portability for developers, ONNX is also valuable for hardware and AI-accelerator builders. Because accelerators often cannot support every possible ML operator, ONNX’s extensible operator model allows manufacturers to seamlessly integrate custom kernels where acceleration is available, while relying on the runtime for unsupported operations. This combination of consistency, flexibility, and extensibility makes ONNX a cornerstone technology for both AI application developers and hardware vendors. \ No newline at end of file +Next, you'll set up your development environment by installing Python, ONNX Runtime, and the necessary tools to build, export, and optimize models for Arm64 and Android deployments. diff --git a/content/learning-paths/mobile-graphics-and-gaming/onnx/02_setup.md b/content/learning-paths/mobile-graphics-and-gaming/onnx/02_setup.md index ff08ffd5e2..a39c6f9a0e 100644 --- a/content/learning-paths/mobile-graphics-and-gaming/onnx/02_setup.md +++ b/content/learning-paths/mobile-graphics-and-gaming/onnx/02_setup.md @@ -1,6 +1,5 @@ --- -# User change -title: "Environment Setup" +title: "Set up your development environment" weight: 3 @@ -8,17 +7,18 @@ layout: "learningpathall" --- ## Objective -This step gets you ready to build, export, run, and optimize ONNX models on Arm64. You’ll set up Python, install ONNX & ONNX Runtime, confirm hardware-backed execution providers. -## Choosing the hardware -You can choose a variety of hardware, including: -* Edge boards (Linux/Arm64) - Raspberry Pi 4/5 (64-bit OS), Jetson (Arm64 CPU; GPU via CUDA if using NVIDIA stack), Arm servers (e.g., AWS Graviton). +Prepare your development environment to build, export, run, and optimize ONNX models on Arm64 platforms. You will install Python, ONNX, and ONNX Runtime, and verify that your system correctly detects and uses available execution providers. + +## Choose your hardware +You can use a variety of Arm64 platforms for this Learning Path: +* Edge boards (Linux/Arm64) - Raspberry Pi 4/5 (64-bit OS), Jetson (Arm64 CPU; GPU via CUDA if using NVIDIA stack), Arm-based servers. * Apple Silicon (macOS/Arm64) - Great for development, deploy to Arm64 Linux later. -* Windows on Arm - Dev/test on WoA, deploy to Linux Arm64 for production if desired. +* Windows on Arm - Suitable for development and testing; deployment to Linux Arm64 systems is also possible. -The nice thing about ONNX is that the **same model file** can run across all of these, so your setup is flexible. +One of ONNX's key advantages is that the same `.onnx` model file can run across all of these platforms, provided the required operators are supported by the runtime and execution providers available on the target system. -## Install Python +## Install Python on your platform {{% notice Note %}} ONNX Runtime provides prebuilt wheels only for specific Python versions. At the time of writing, Python 3.12 is not yet supported by ONNX Runtime on macOS or Arm platforms. If you see an error like: @@ -30,9 +30,9 @@ ERROR: No matching distribution found for onnxruntime it usually means your Python version is too new. Python 3.10 is tested and recommended for this Learning Path. {{% /notice %}} -Depending on the hardware you use, follow different installation paths: +Depending on your platform, follow different installation paths: -1. Linux (Arm64). Install Python 3.10 by typing in the console: +### Linux (Arm64) ```console sudo apt update sudo apt install -y python3.10 python3.10-venv python3.10-dev python3-pip build-essential libopenblas-dev libgl1 libglib2.0-0 @@ -45,18 +45,27 @@ sudo apt update sudo apt install -y python3.10 python3.10-venv python3.10-dev ``` -2. macOS (Apple Silicon). Install Python 3.10 using Homebrew: +### macOS (Apple Silicon) + +Install Python 3.10 using Homebrew: ```console brew install python@3.10 ``` After installation, use `python3.10` explicitly when creating virtual environments. -3. Windows on Arm: -* Download and install Python 3.10 from [python.org](https://www.python.org/downloads/) (select the Arm64 build). -* Ensure pip is on PATH. +### Windows on Arm + +Download and install Python 3.10 from [python.org](https://www.python.org/downloads/): + +1. Select the ARM64 build for Windows +2. Run the installer +3. Check **Add Python to PATH** during installation +4. Verify the installation by opening Command Prompt and running `python --version` -After installing Python 3.10, open a terminal or console, create a clean virtual environment using Python 3.10 explicitly, and update pip and wheel: + +## Create a Virtual Environment +After installing Python 3.10, create a clean virtual environment: ```console python3.10 -m venv .venv @@ -71,30 +80,32 @@ source .venv/bin/activate python -m pip install --upgrade pip wheel ``` -Using a virtual environment keeps dependencies isolated and avoids conflicts with system-wide Python packages. +Using a virtual environment ensures dependency isolation and prevents conflicts with system-wide Python packages. ## Install Core Packages -Start by installing the minimal stack: +Install the minimal ONNX toolchain: ```console pip install onnx onnxruntime onnxscript netron numpy ``` -The above will install the following: +This installs: * onnx – core library for loading/saving ONNX models. * onnxruntime – high-performance runtime to execute models. * onnxscript – required for the new Dynamo-based exporter. * netron – tool for visualizing ONNX models. * numpy – used for tensor manipulation. -Now, install PyTorch (we’ll use it later to build and export a sample model): +Now, install PyTorch(CPU build): ```console pip install torch torchvision --index-url https://download.pytorch.org/whl/cpu ``` +On Arm64 systems without discrete GPUs, the CPU build is sufficient. ONNX Runtime will later handle optimized inference execution. ## Verify the installation -Let’s verify everything works end-to-end by training a toy network and exporting it to ONNX. -Create a new file 01_Init.py and add the following code +You'll now validate the entire toolchain by defining a small PyTorch model, exporting it to ONNX, and running inference using ONNX Runtime. + +Create a file named `01_Init.py`: ```python import torch, torch.nn as nn @@ -131,13 +142,13 @@ out = sess.run(["logits"], {"input": dummy.numpy()})[0] print("Output shape:", out.shape, "Providers:", sess.get_providers()) ``` -Then, run it as follows +Run the script: ```console python3 01_Init.py ``` -You should see the following output: +You should see output similar to: ```output python3 01_Init.py [torch.onnx] Obtain model graph for `SmallNet([...]` with `torch.export.export(..., strict=False)`... @@ -149,9 +160,12 @@ python3 01_Init.py Output shape: (1, 10) Providers: ['CPUExecutionProvider'] ``` -The 01_Init.py script serves as a quick end-to-end validation of your ONNX environment. It defines a very small convolutional neural network (SmallNet) in PyTorch, which consists of a convolution layer, activation function, pooling, flattening, and a final linear layer that outputs 10 logits. Instead of training the model, we simply run it in evaluation mode on a random input tensor to make sure the graph structure works. This model is then exported to the ONNX format using PyTorch’s new Dynamo-based exporter, producing a portable smallnet.onnx file. +The `01_Init.py` script serves as a quick end-to-end validation of your ONNX environment. It defines a very small convolutional neural network (SmallNet) in PyTorch, which consists of a convolution layer, activation function, pooling, flattening, and a final linear layer that outputs 10 logits. Instead of training the model, we simply run it in evaluation mode on a random input tensor to make sure the graph structure works. This model is then exported to the ONNX format using PyTorch’s new Dynamo-based exporter, producing a portable smallnet.onnx file. After export, the script immediately loads the ONNX model with ONNX Runtime and executes a forward pass using the CPU execution provider. This verifies that the installation of ONNX, ONNX Runtime, and PyTorch is correct and that models can flow seamlessly from definition to inference. By printing the output tensor’s shape and the active execution provider, the script demonstrates that the toolchain is fully functional on your Arm64 device, giving you a solid baseline before moving on to more advanced models and optimizations. -## Summary -You now have a fully functional ONNX development environment on Arm64. Python and all required packages are installed, and you successfully exported a small PyTorch model to ONNX using the new Dynamo exporter, ensuring forward compatibility. Running the model with ONNX Runtime confirmed that inference works end-to-end with the CPU execution provider, proving that your toolchain is correctly configured. With this foundation in place, the next step is to build and export a more complete model and run it on Arm64 hardware to establish baseline performance before applying optimizations. +## What you've learned and what's next + +In this section, you installed Python 3.10 and the core ONNX toolchain on your Arm64 platform, created an isolated virtual environment, and verified the entire setup by exporting a small PyTorch model to ONNX and running inference with ONNX Runtime. Your development environment is now ready for model development and optimization. + +Next, you'll build a complete neural network model, export it to ONNX format, and run it on Arm64 hardware to establish baseline performance metrics before applying optimization techniques. \ No newline at end of file diff --git a/content/learning-paths/mobile-graphics-and-gaming/onnx/03_preparingdata.md b/content/learning-paths/mobile-graphics-and-gaming/onnx/03_preparingdata.md index db8a544395..d7973a5866 100644 --- a/content/learning-paths/mobile-graphics-and-gaming/onnx/03_preparingdata.md +++ b/content/learning-paths/mobile-graphics-and-gaming/onnx/03_preparingdata.md @@ -1,32 +1,44 @@ --- -# User change -title: "Preparing a Synthetic Sudoku Digit Dataset" +title: "Generate a synthetic Sudoku digit dataset" weight: 4 layout: "learningpathall" --- +## Overview + +The end goal is a camera-to-solution Sudoku app that runs efficiently on Arm64 devices (for example, Raspberry Pi or Android phones). ONNX is the glue: you will train the digit recognizer in PyTorch, export it to ONNX, and run it anywhere with ONNX Runtime (CPU EP on edge devices, NNAPI EP on Android). Everything around the model--grid detection, perspective rectification, and solving--stays deterministic and lightweight. + + -## Big picture -Our end goal is a camera-to-solution Sudoku app that runs efficiently on Arm64 devices (e.g., Raspberry Pi or Android phones). ONNX is the glue: we’ll train the digit recognizer in PyTorch, export it to ONNX, and run it anywhere with ONNX Runtime (CPU EP on edge devices, NNAPI EP on Android). Everything around the modelβ€”grid detection, perspective rectification, and solvingβ€”stays deterministic and lightweight. ## Objective -In this step, we will generate a custom dataset of Sudoku puzzles and their digit crops, which we’ll use to train a digit recognition model. Starting from a Hugging Face parquet dataset that provides paired puzzle/solution strings, we transform raw boards into realistic, book-style Sudoku pages, apply camera-like augmentations to mimic mobile captures, and automatically slice each page into 81 labeled cell images. This yields a large, diverse, perfectly labeled set of digits (0–9 with 0 = blank) without manual annotation. By the end, you’ll have a structured dataset ready to train a lightweight model in the next section. -## Why Synthetic Generation? -When building a Sudoku digit recognizer, the hardest part is obtaining a well-labeled dataset that matches real capture conditions. MNIST contains handwritten digits, which differ from printed, grid-aligned Sudoku digits; relying on it alone hurts real-world performance. +In this step, you'll generate a custom dataset of Sudoku puzzles and digit crops for training. Starting from a Hugging Face parquet dataset that provides paired puzzle/solution strings, you will transform raw boards into realistic, book-style Sudoku pages, apply camera-like augmentations to mimic mobile captures, and automatically slice each page into 81 labeled cell images. This yields a large, diverse, perfectly labeled set of digits (0–9 with 0 = blank) without manual annotation. + +## Why synthetic generation? + +Obtaining a well-labeled dataset that matches real capture conditions is often the hardest part of building a vision model. MNIST contains handwritten digits, which differ significantly from printed, grid-aligned Sudoku digits. Training only on MNIST typically results in poor performance on printed puzzles. + +By generating synthetic Sudoku pages directly from structured puzzle data, you gain: + +* **Perfect labeling**: since the puzzle content is known, every cropped cell automatically receives the correct label (digit or blank). This eliminates manual annotation errors. + +* **Domain alignment**: the rendered digits match printed-book Sudoku styling rather than handwritten digits. + +* **Controlled augmentation**: by applying perspective warps, blur, and illumination variations, we approximate smartphone capture conditions. This improves robustness when deploying on mobile devices. -By generating synthetic Sudoku pages directly from the parquet dataset, we get: -1. Perfect labeling. Since the puzzle content is known, every cropped cell automatically comes with the correct label (digit or blank), eliminating manual annotation. -2. Control over style. We can render Sudoku pages to look like those in printed books, with realistic fonts, grid lines, and difficulty levels controlled by how many cells are left blank. -3. Robustness through augmentation: By applying perspective warps, blur, noise, and lighting variations, we simulate how a smartphone camera might capture a Sudoku page, improving the model’s ability to handle real-world photos. -4. Scalability. With millions of Sudoku solutions available, we can easily generate tens of thousands of training samples in minutes, ensuring a dataset that is both large and diverse. +* **Scalability**: with large numbers of available Sudoku puzzles, we can generate tens of thousands of labeled samples quickly. -This synthetic data generation strategy allows us to create a custom-fit dataset for our Sudoku digit recognition problem, bridging the gap between clean digital puzzles and noisy real-world inputs. +One important consideration is class imbalance: Sudoku puzzles contain many blank cells, so the β€œ0” class will naturally dominate. You will address this during training if needed through sampling or loss weighting. -## What we’ll produce -By the end of this step, you will have two complementary datasets: -1. Digit crops for training the classifier. A folder tree structured for torchvision.datasets.ImageFolder, containing tens of thousands of labeled 28Γ—28 images of Sudoku digits (0–9, with 0 meaning blank): +## Dataset structure + +By the end of this section, you will have two complementary datasets: + +### Digit crops for training + +A folder tree structured for `torchvision.datasets.ImageFolder`, containing tens of thousands of labeled 28Γ—28 images of Sudoku digits (0–9, with 0 meaning blank): ```console data/ @@ -41,9 +53,11 @@ data/ 9/....png ``` -These will be used in next step to train a lightweight model for digit recognition. +These will be used in the next section to train a lightweight model for digit recognition. + +### Rendered Sudoku grids -2. Rendered Sudoku grids for camera simulation. Full-page Sudoku images (both clean book-style and augmented camera-like versions) stored in: +Full-page Sudoku images (both clean book-style and augmented camera-like versions) stored in: ```console data/ grids/ @@ -60,7 +74,7 @@ These grid images allow us to later test the end-to-end pipeline: detect the boa Together, these datasets provide both the micro-level data needed to train the digit recognizer and the macro-level data to simulate the camera pipeline for testing and deployment. ## Implementation -Start by creating a new file 02_PrepareData.py and modify it as follows: +Start by creating a new file `02_PrepareData.py` with the following content: ```python import os, random, pathlib import numpy as np @@ -203,21 +217,21 @@ if __name__ == "__main__": main() ``` -At the top, you set basic knobs for the generator: where to read the Parquet file, where to write outputs, how many puzzles to render for train/val, page size, grid margin, crop size, and the OpenCV font. Tweaking these lets you control dataset scale, visual style, and classifier input size (e.g., CELL_SIZE=32 if you want a slightly larger digit crop). +At the top, you set basic knobs for the generator: where to read the Parquet file, where to write outputs, how many puzzles to render for train/val, page size, grid margin, crop size, and the OpenCV font. Tweaking these lets you control dataset scale, visual style, and classifier input size (for example, CELL_SIZE=32 if you want a slightly larger digit crop). -The method str_to_grid(s) converts an 81-character Sudoku string into a 9Γ—9 list of integers. Each character represents a cell: 0 is blank, 1–9 are digits. This is the canonical internal representation used throughout the script. +The method `str_to_grid(s)` converts an 81-character Sudoku string into a 9Γ—9 list of integers. Each character represents a cell: 0 is blank, 1–9 are digits. This is the canonical internal representation used throughout the script. -Then, we have load_puzzles(parquet_path, n_train, n_val), which loads the dataset from Parquet, shuffles it deterministically, and slices it into train/val partitions. It returns the puzzles (and, if present, solutions) as 9Γ—9 integer grids. In this step we only need puzzle for rendering and labeling digit crops (blanks included); solution is useful later for solver validation. +Then, you have `load_puzzles(parquet_path, n_train, n_val)`, which loads the dataset from Parquet, shuffles it deterministically, and slices it into train/val partitions. It returns the puzzles (and, if present, solutions) as 9Γ—9 integer grids. In this step we only need puzzle for rendering and labeling digit crops (blanks included); solution is useful later for solver validation. -Subsequently, draw_grid(img, size=9, margin=GRID_MARGIN) draws a Sudoku grid on a blank page image. It computes the step size from the page dimensions and margin, then draws both thin inner lines and thick 3Γ—3 box boundaries. It returns the top-left corner (x0, y0) and the cell size (step), which are reused to place digits and to locate each cell for cropping. +Subsequently, `draw_grid(img, size=9, margin=GRID_MARGIN)` draws a Sudoku grid on a blank page image. It computes the step size from the page dimensions and margin, then draws both thin inner lines and thick 3Γ—3 box boundaries. It returns the top-left corner (x0, y0) and the cell size (step), which are reused to place digits and to locate each cell for cropping. -Next, put_digit(img, r, c, d, x0, y0, step) renders a single digit d at row r, column c inside the grid. The text is centered in the cell using the font metrics; if d == 0, it leaves the cell blank. This mirrors printed-book Sudoku styling so our crops look realistic. +Next, `put_digit(img, r, c, d, x0, y0, step)` renders a single digit d at row r, column c inside the grid. The text is centered in the cell using the font metrics; if d == 0, it leaves the cell blank. This mirrors printed-book Sudoku styling so our crops look realistic. -Another method, render_page(puzzle9x9) builds a complete β€œbook-style” Sudoku page: creates a white canvas, draws the grid, loops over all 81 cells, and writes digits using put_digit. It returns the page plus the grid geometry (x0, y0, step) for subsequent cropping. +Another method, `render_page(puzzle9x9)` builds a complete β€œbook-style” Sudoku page: creates a white canvas, draws the grid, loops over all 81 cells, and writes digits using put_digit. It returns the page plus the grid geometry (x0, y0, step) for subsequent cropping. -A method aug_camera(img) applies a light, camera-like augmentation to mimic smartphone captures: a small perspective warp (random corner jitter) and optional Gaussian blur. The warp uses a light gray border fill so any exposed areas look like paper rather than colored artifacts. This produces a second version of each page that’s closer to real-world inputs. +A method `aug_camera(img)` applies a light, camera-like augmentation to mimic smartphone captures: a small perspective warp (random corner jitter) and optional Gaussian blur. The warp uses a light gray border fill so any exposed areas look like paper rather than colored artifacts. This produces a second version of each page that’s closer to real-world inputs. -Afterward, ensure_dirs(split) makes the class directories for a given split (train or val) so that crops can be saved in data/{split}/{class}/.... The classes are 0..9 with 0 = blank. +Afterward, `ensure_dirs(split)` makes the class directories for a given split (train or val) so that crops can be saved in data/{split}/{class}/.... The classes are 0..9 with 0 = blank. A method save_crops(page, geom, puzzle9x9, split, base_id) slices the page into 81 cell crops using the grid geometry, converts each crop to grayscale, resizes it to CELL_SIZE Γ— CELL_SIZE, and saves it into the appropriate class directory based on the puzzle’s value at that cell (0..9). Using the puzzle for labels ensures we learn to recognize blanks as well as digits. @@ -235,8 +249,11 @@ data/ val/ (..._clean.png, ..._cam.png) ``` -## Launching instructions -1. Install dependencies (inside your virtual env): +## Run the data generation script + +### Install dependencies + +Inside your virtual environment, install the required packages: ```console pip install pandas pyarrow opencv-python tqdm numpy ``` @@ -266,5 +283,8 @@ Tips * If you want larger inputs for the classifier, increase CELL_SIZE to 32 or 40. * To make augmentation a bit stronger (more realistic), slightly increase the perspective jitter in aug_camera, add brightness/contrast jitter, or a faint gradient shadow overlay. -## Summary -After running this step you’ll have a robust, labeled, Sudoku-specific dataset: thousands of digit crops (including blanks) for training and realistic full-page grids for pipeline testing. You’re ready for the next stepβ€”training the digit recognizer and exporting it to ONNX. \ No newline at end of file +## What you've learned and what's next + +In this section, you've learned why synthetic data generation is essential for Sudoku digit recognition, and you've built a complete pipeline to generate realistic labeled datasets. You've created thousands of digit crops (0–9, with 0 representing blanks) and full-page Sudoku images with camera-like augmentations. You now have both the micro-level training data your classifier needs and the macro-level test images to validate the end-to-end pipeline later. + +In the next section, you'll take these digit crops and train a lightweight neural network model using PyTorch, then export it to ONNX format so it can run efficiently on Arm devices with ONNX Runtime. \ No newline at end of file diff --git a/content/learning-paths/mobile-graphics-and-gaming/onnx/04_training.md b/content/learning-paths/mobile-graphics-and-gaming/onnx/04_training.md index 671b8d04ad..7e8bcd51fa 100644 --- a/content/learning-paths/mobile-graphics-and-gaming/onnx/04_training.md +++ b/content/learning-paths/mobile-graphics-and-gaming/onnx/04_training.md @@ -1,19 +1,19 @@ --- -# User change -title: "Train the Digit Recognizer" +title: "Train the digit recognizer" weight: 5 layout: "learningpathall" --- -## Objective ## -We will now train a small CNN to classify Sudoku cell crops into 10 classes (0=blank, 1..9=digit), verify accuracy, then export the model to ONNX using the Dynamo exporter and sanity-check parity with ONNX Runtime. This gives us a portable model ready for Arm64 inference and later Android deployment. +## Objective -## Creating a model -We use a tiny convolutional neural network (CNN) called DigitNet, designed to be both fast (so it runs efficiently on Arm64 and mobile) and accurate enough for recognizing 28Γ—28 grayscale crops of Sudoku digits. It expects 1 input channel (in_channels=1) because we forced grayscale in the preprocessing step. +In this step, you'll train a small CNN to classify Sudoku cell crops into 10 classes (0 = blank, 1–9 = digit), validate accuracy, export the best model to ONNX using PyTorch's Dynamo-based exporter, and verify numerical parity with ONNX Runtime. This produces a portable model ready for efficient Arm64 inference and later Android deployment. -We start by creating a new file digitnet_model.py and defining the DigitNet class: +## Create a model +You will use a tiny convolutional neural network (CNN) called DigitNet, designed to be both fast (so it runs efficiently on Arm64 and mobile) and accurate enough for recognizing 28Γ—28 grayscale crops of Sudoku digits. It expects 1 input channel (in_channels=1) because we forced grayscale in the preprocessing step. + +Create a file named `digitnet_model.py` and define the DigitNet class. This small convolutional neural network is designed to be fast on Arm64 and mobile devices while remaining accurate for 28Γ—28 grayscale Sudoku digit crops: ```python import torch import torch.nn as nn @@ -39,7 +39,7 @@ class DigitNet(nn.Module): return self.net(x) ``` -We use a very compact convolutional neural network (CNN), which we call DigitNet, to recognize Sudoku digits. The goal is to have a model that is simple enough to run efficiently on Arm64 and mobile devices, but still powerful enough to tell apart the ten classes we care about (0 for blank, and digits 1 through 9). +You will use a very compact convolutional neural network (CNN), called DigitNet, to recognize Sudoku digits. The goal is to have a model that is simple enough to run efficiently on Arm64 and mobile devices, but still powerful enough to tell apart the ten classes we care about (0 for blank, and digits 1 through 9). The network expects each input to be a 28Γ—28 grayscale crop, so it begins with a convolution layer that has one input channel and sixteen filters. This first convolution is responsible for learning very low-level patterns such as strokes or edges. Immediately after, a ReLU activation introduces non-linearity, which allows the network to combine those simple features into more expressive ones. A max-pooling layer then reduces the spatial resolution by half, making the representation more compact and less sensitive to small translations. @@ -52,7 +52,7 @@ The final step is a fully connected layer that maps the thirty-two features to t In practice, this means that when you feed in a batch of grayscale Sudoku cells of shape [N, 1, 28, 28], DigitNet transforms them step by step into a batch of [N, 10] outputs, where each row contains the scores for the ten possible classes. Despite its simplicity, this small CNN strikes a balance between speed and accuracy that makes it ideal for Sudoku digit recognition on resource-constrained devices. ## Training a model -We will now prepare the self-containing script that trains the above model on the data prepared earlier. Start by creating the new file 03_Training.py and modify it as follows: +You will now prepare the self-containing script that trains the above model on the data prepared earlier. Create a new file `03_Training.py` with the content below: ```python import os, random, numpy as np import torch as tr @@ -173,17 +173,17 @@ if __name__ == "__main__": main() ``` -This file is a self-contained trainer for the Sudoku digit classifier. It starts by fixing random seeds for reproducibility and sets DEVICE="cpu" so the workflow runs the same on desktops and Arm64 boards. It expects the dataset from the previous step under data/train/0..9 and data/val/0..9, and creates an artifacts/ folder for all outputs. +This is a self-contained trainer for the Sudoku digit classifier. It starts by fixing random seeds for reproducibility and sets DEVICE="cpu" so the workflow runs the same on desktops and Arm64 boards. It expects the dataset from the previous step under data/train/0..9 and data/val/0..9, and creates an artifacts/ folder for all outputs. -The script builds two dataloaders (train/val) with a preprocessing stack that forces grayscale (Grayscale(num_output_channels=1)) so inputs match the model’s first convolution, converts to tensors, and normalizes to a centered range. Light augmentations on the training splitβ€”small affine jitter and occasional blurβ€”mimic camera variability without distorting the digits. Batch size, epochs, and learning rate are set to conservative defaults so training is smooth on CPU; you can scale them up later. +The script builds two dataloaders (train/val) with a preprocessing stack that forces grayscale (Grayscale(num_output_channels=1)) so inputs match the model’s first convolution, converts to tensors, and normalizes to a centered range. Light augmentations on the training split--small affine jitter and occasional blur--mimic camera variability without distorting the digits. Batch size, epochs, and learning rate are set to conservative defaults so training is smooth on CPU; you can scale them up later. Then, the script it instantiates DigitNet(num_classes=10) model. The optimizer is AdamW with mild weight decay to control overfitting. The loss is cross-entropy with label smoothing (e.g., 0.05), which reduces over-confidence and helps on easily confused shapes (like 6/8/9). The training loop runs for a fixed number of epochs, iterating mini-batches from the training set. After each epoch, it evaluates on the validation split and logs the accuracy. The script keeps track of the best model state seen so far (based on val accuracy) and restores it at the end, ensuring the final model corresponds to your best epoch, not just the last one. The file will create two artifacts: -1. digitnet_best.pth β€” the best PyTorch weights (handy for quick experiments, fine-tuning, or debugging later). -2. sudoku_digitnet.onnx β€” the exported ONNX model, produced with PyTorch’s Dynamo exporter and a dynamic batch dimension. Dynamic batch means the model accepts input of shape [N, 1, 28, 28] for any N, which is ideal for efficient batched inference on Arm64 and for Android integration. +1. digitnet_best.pth -- the best PyTorch weights (handy for quick experiments, fine-tuning, or debugging later). +2. sudoku_digitnet.onnx -- the exported ONNX model, produced with PyTorch’s Dynamo exporter and a dynamic batch dimension. Dynamic batch means the model accepts input of shape [N, 1, 28, 28] for any N, which is ideal for efficient batched inference on Arm64 and for Android integration. Right after export, the script runs a parity test: it feeds the same randomly generated batch through both the PyTorch model and the ONNX model (executed by ONNX Runtime) and prints the mean absolute error between their logits. A tiny value confirms the exported graph faithfully matches your trained network. @@ -197,13 +197,13 @@ pip install --upgrade torch torchvision ``` {{% /notice %}} -To run the training script, type: +Run the training script: ```console python3 03_Training.py ``` -The script will train, validate, export, and verify the digit recognizer in one go. After it finishes, you’ll have both a portable ONNX model and a PyTorch checkpoint ready for the next stepβ€”building the image processor that detects the Sudoku grid, rectifies it, segments cells, and performs batched ONNX inference to reconstruct the board for solving. +The script will train, validate, export, and verify the digit recognizer in one go. After it finishes, you’ll have both a portable ONNX model and a PyTorch checkpoint ready for the next step--building the image processor that detects the Sudoku grid, rectifies it, segments cells, and performs batched ONNX inference to reconstruct the board for solving. Here is a sample run: @@ -242,5 +242,8 @@ Exported: artifacts/sudoku_digitnet.onnx Parity MAE: 1.0251999e-05 ``` -## Summary -By running the training script you train the DigitNet CNN on the Sudoku digit dataset, steadily improving accuracy across epochs until the model surpasses 99% validation accuracy. The process builds on the earlier steps where we first defined the model architecture in digitnet_model.py and then prepared a dedicated training script to handle data loading, augmentation, optimization, and evaluation. During training the best-performing model state is saved, and at the end it is exported to the ONNX format with dynamic batch support. A parity check confirms that the ONNX and PyTorch versions produce virtually identical outputs (mean error ~1e-5). You now have a validated ONNX model (artifacts/sudoku_digitnet.onnx) and a PyTorch checkpoint (digitnet_best.pth), both ready for integration into the Sudoku image processing pipeline. Before moving on to grid detection and solving, however, we will first run standalone inference to confirm the model’s predictions on individual digit crops. +## What you've learned and what's next + +In this section, you defined and trained DigitNet, achieving over 99% validation accuracy on Sudoku digit classification. You exported the trained model to ONNX format with dynamic batch support and verified numerical parity between PyTorch and ONNX Runtime outputs. You now have both a portable ONNX model and PyTorch checkpoint ready for deployment. + +Next, you'll run standalone inference on individual digit crops to verify the model's predictions before integrating it into the complete Sudoku image processing pipeline that handles grid detection and solving. \ No newline at end of file diff --git a/content/learning-paths/mobile-graphics-and-gaming/onnx/05_inference.md b/content/learning-paths/mobile-graphics-and-gaming/onnx/05_inference.md index 32e373289b..a1ebafaa66 100644 --- a/content/learning-paths/mobile-graphics-and-gaming/onnx/05_inference.md +++ b/content/learning-paths/mobile-graphics-and-gaming/onnx/05_inference.md @@ -1,19 +1,19 @@ --- -# User change -title: "Inference and Model Evaluation" +title: "Run inference and evaluate the model" weight: 6 layout: "learningpathall" --- -## Objective ## -In this section, we validate the digit recognizer by running inference on the validation dataset using both the PyTorch checkpoint and the exported ONNX model. We verify that PyTorch and ONNX Runtime produce consistent results, analyze class-level behavior using a confusion matrix, and generate visual diagnostics for debugging and documentation. This step acts as a final verification checkpoint before integrating the model into the full OpenCV-based Sudoku processing pipeline. +## Objective + +In this section, you'll validate the digit recognizer by running inference on the validation dataset using both the PyTorch checkpoint and the exported ONNX model. Verify that PyTorch and ONNX Runtime produce consistent results, analyze class-level behavior using a confusion matrix, and generate visual diagnostics for debugging and documentation. This acts as a final verification checkpoint before integrating the model into the full OpenCV-based Sudoku processing pipeline. Before introducing geometric processing, grid detection, and perspective correction, it is important to confirm that the digit recognizer works reliably in isolation. By validating inference and analyzing errors at the digit level, we ensure that any future issues in the end-to-end system can be attributed to image processing or geometry rather than the classifier itself. ## Inference and Evaluation Script -Create a new file named 04_Test.py and paste the script below into it. This script evaluates the digit recognizer in a way that closely mirrors deployment conditions. It compares PyTorch and ONNX Runtime inference, measures accuracy on the validation dataset, and generates visual diagnostics that reveal both strengths and remaining failure modes of the model. +Create a new file named `04_Test.py` and paste the script below into it. This script evaluates the digit recognizer in a way that closely mirrors deployment conditions. It compares PyTorch and ONNX Runtime inference, measures accuracy on the validation dataset, and generates visual diagnostics that reveal both strengths and remaining failure modes of the model. ```python import os, numpy as np, torch @@ -220,13 +220,12 @@ Before running the evaluation script, install matplotlib for visualization: pip install matplotlib ``` -Run the evaluation script from the project root: +Run the evaluation script: ```console python3 04_Test.py ``` - -In the example below, the PyTorch and ONNX accuracies match exactly, confirming that the export process preserved model behavior. +The PyTorch and ONNX accuracies should match very closely, confirming that the export process preserved model behavior. ```console python3 04_Test.py @@ -247,10 +246,14 @@ Confusion matrix (rows=true, cols=pred): Saved: artifacts/cm_fp32_counts.png ``` -![img1](figures/01.png) +![Confusion matrix showing digit recognition accuracy with strong diagonal and occasional confusion between visually similar digits like 6, 8, and 9](figures/01.png) + The confusion matrix provides more insight than a single accuracy number. Each row corresponds to the true class, and each column corresponds to the predicted class. A strong diagonal indicates correct classification. In this output, blank cells (class 0) are almost always recognized correctly, while the remaining errors occur primarily between visually similar printed digits such as 6, 8, and 9. This behavior is expected and indicates that the model has learned meaningful digit features. The remaining confusions are rare and can be addressed later through targeted augmentation or higher-resolution crops if needed. -## Summary -With inference validated and error modes understood, the digit recognizer is now ready to be embedded into the full Sudoku image-processing pipeline, where OpenCV will be used to detect the grid, rectify perspective, segment cells, and run batched ONNX inference to reconstruct and solve complete puzzles. +## What you've learned and what's next + +In this section, you validated the digit recognizer by running inference on the validation dataset using both PyTorch and ONNX Runtime, confirmed that both produce consistent results (99.28% accuracy), and analyzed class-level behavior using confusion matrices and sample prediction grids. The diagnostic outputs revealed that the model performs reliably, with rare confusions limited to visually similar digits. + +Next, you'll integrate the validated ONNX model into the full Sudoku image-processing pipeline using OpenCV to detect grids, rectify perspective, segment cells, and perform batched inference to reconstruct and solve complete puzzles. \ No newline at end of file diff --git a/content/learning-paths/mobile-graphics-and-gaming/onnx/06_sudokuprocessor.md b/content/learning-paths/mobile-graphics-and-gaming/onnx/06_sudokuprocessor.md index 1f5e686a3b..49e73871f8 100644 --- a/content/learning-paths/mobile-graphics-and-gaming/onnx/06_sudokuprocessor.md +++ b/content/learning-paths/mobile-graphics-and-gaming/onnx/06_sudokuprocessor.md @@ -1,26 +1,28 @@ --- -# User change -title: "Sudoku Processor. From Image to Solution" +title: "Build the Sudoku processor pipeline" + weight: 7 -layout: "learningpathall" +layout: "learningpathall" --- -## Objective ## +## Objective -In this section, we integrate all previous components into a complete Sudoku processing pipeline. Starting from a full Sudoku image, we detect and rectify the grid, split it into individual cells, recognize digits using the ONNX model, and finally solve the puzzle using a deterministic solver. By the end of this step, you will have an end-to-end system that takes a photograph of a Sudoku puzzle and produces a solved board, along with visual outputs for debugging and validation. +Integrate all previous components into a complete Sudoku processing pipeline. Starting from a full Sudoku image, detect and rectify the grid, split it into individual cells, recognize digits using the ONNX model, and solve the puzzle using a deterministic solver. By the end of this section, you will have an end-to-end system that takes a photograph of a Sudoku puzzle and produces a solved board, along with visual outputs for debugging and validation. ## Context -So far, we have: -1. Generated a synthetic, well-labeled Sudoku digit dataset, -2. Trained a lightweight CNN (DigitNet) to recognize digits and blanks, -3. Exported the model to ONNX with dynamic batch support, -4. Validated inference correctness and analyzed errors using confusion matrices. -At this point, the digit recognizer is reliable in isolation. The remaining challenge is connecting vision with reasoning: extracting the Sudoku grid from an image, mapping each cell to a digit, and applying a solver. This section bridges that gap. +So far in this Learning Path, you have: + +- Generated a synthetic, well-labeled Sudoku digit dataset +- Trained a lightweight CNN (DigitNet) to recognize digits and blanks +- Exported the model to ONNX with dynamic batch support +- Validated inference correctness and analyzed errors using confusion matrices + +At this point, the digit recognizer is reliable in isolation. The remaining challenge is connecting vision with reasoning: extracting the Sudoku grid from an image, mapping each cell to a digit, and applying a solver. ## Overview of the pipeline -To implement the Sudoku processor, create the file (sudoku_processor.py) and paste the implementation below: +To implement the Sudoku processor, create a file named `sudoku_processor.py` with the code below: ```python import cv2 as cv @@ -328,7 +330,7 @@ The Sudoku processor follows a sequence of steps: 6. Solving – apply a backtracking Sudoku solver. 7. Visualization – overlay the solution and render clean board images. -We encapsulate the entire pipeline in a reusable class called SudokuProcessor. This class loads the ONNX model once and exposes a single high-level method that processes an input image and returns both intermediate results and final outputs. +You encapsulate the entire pipeline in a reusable class called SudokuProcessor. This class loads the ONNX model once and exposes a single high-level method that processes an input image and returns both intermediate results and final outputs. Conceptually, the processor: * Accepts a BGR image, @@ -336,12 +338,14 @@ Conceptually, the processor: This design keeps inference fast and makes the processor easy to integrate later into an Android application or embedded system. +In real photos, grid detection and preprocessing dominate accuracy; the solver will fail when one or more digits are misread or placed in the wrong cell. + ## Grid detection and rectification -The first task is to locate the Sudoku grid in the image. We convert the image to grayscale, apply adaptive thresholding, and use contour detection to find large rectangular shapes. The largest contour that approximates a quadrilateral is assumed to be the Sudoku grid. +The first task is to locate the Sudoku grid in the image. You convert the image to grayscale, apply adaptive thresholding, and use contour detection to find large rectangular shapes. The largest contour that approximates a quadrilateral is assumed to be the Sudoku grid. -Once the four corners are identified, we compute a perspective transform and warp the grid into a square image. This rectified representation removes camera tilt and perspective distortion, allowing all subsequent steps to assume a fixed geometry. +Once the four corners are identified, you will compute a perspective transform and warp the grid into a square image. This rectified representation removes camera tilt and perspective distortion, allowing all subsequent steps to assume a fixed geometry. -We order the four corners consistently (top-left β†’ top-right β†’ bottom-right β†’ bottom-left) before computing the perspective transform. +You will then order the four corners consistently (top-left β†’ top-right β†’ bottom-right β†’ bottom-left) before computing the perspective transform. ## Splitting the grid into cells After rectification, the grid is divided evenly into a 9Γ—9 array. Each cell is cropped based on its row and column index. At this stage, every cell corresponds to one Sudoku position and is ready for preprocessing and classification. @@ -353,7 +357,7 @@ Each cell undergoes light preprocessing before inference: * Resizing to the model’s input size (28Γ—28), * Normalization to match the training distribution. -We crop a margin to suppress grid lines, because grid strokes can dominate the digit pixels and cause systematic misclassification. Cells with very little foreground content are treated as blank candidates, reducing false digit detections in empty cells. +You will crop a margin to suppress grid lines, because grid strokes can dominate the digit pixels and cause systematic misclassification. Cells with very little foreground content are treated as blank candidates, reducing false digit detections in empty cells. ## Batched ONNX inference All 81 cell tensors are stacked into a single batch and passed to ONNX Runtime in one call. Because the model was exported with a dynamic batch dimension, this batched inference is efficient and mirrors how the model will be used in production. @@ -363,14 +367,16 @@ The output logits are converted to probabilities, and the most likely class is s The result is a 9Γ—9 board where: * 0 represents a blank cell, * 1–9 represent recognized digits. + +Batched inference reduces per-call overhead in ONNX Runtime and generally improves throughput on Arm64 CPUs. ## Solving the Sudoku -With the recognized board constructed, we apply a classic backtracking Sudoku solver. This solver deterministically fills empty cells while respecting Sudoku constraints (row, column, and 3Γ—3 block rules). +With the recognized board constructed, you will apply a classic backtracking Sudoku solver. This solver deterministically fills empty cells while respecting Sudoku constraints (row, column, and 3Γ—3 block rules). -If the solver succeeds, we obtain a complete solution. If it fails, the failure usually indicates one or more recognition errors, which can be diagnosed using the intermediate visual outputs. +If the solver succeeds, you obtain a complete solution. If it fails, the failure usually indicates one or more recognition errors, which can be diagnosed using the intermediate visual outputs. ## Visualization and outputs -The processor saves several artifacts to help debugging and demonstration: +The driver script saves intermediate images (warped grid and optional overlay) under artifacts/ for debugging and documentation: - `artifacts/warped.png` – rectified top-down view of the Sudoku grid. - `artifacts/overlay_solution.png` – solution digits overlaid onto the original image (if solved). - (Optional) `artifacts/recognized_board.png`, `artifacts/solved_board.png`, `artifacts/boards_side_by_side.png` – clean board renderings if you enabled those helpers. @@ -378,7 +384,7 @@ The processor saves several artifacts to help debugging and demonstration: The driver script below saves warped.png and overlay_solution.png by default. ## Running the processor -A small driver script (05_RunSudokuProcessor.py) demonstrates how to use the SudokuProcessor: +A small driver script `05_RunSudokuProcessor.py` demonstrates how to use the SudokuProcessor: ```python import os @@ -436,13 +442,10 @@ You simply provide the path to a Sudoku image and the ONNX model, and the script Representational result is shown below: -![img2](figures/02.png) +![Sudoku processing pipeline result showing the original camera-like input image on the left and the solved puzzle with green-filled digits overlaid on the right](figures/02.png) + +## What you've learned and what's next -## Summary -By completing this section, you have built a full vision-to-solution Sudoku system: -1. A trained and validated ONNX digit recognizer, -2. A robust OpenCV-based image processing pipeline, -3. A deterministic solver, -4. Clear visual diagnostics at every stage. +In this section, you integrated all previous components into a complete end-to-end Sudoku processing pipeline. You built a system that detects and rectifies Sudoku grids from photographs, splits them into individual cells, recognizes digits using batched ONNX inference, applies a deterministic solver, and overlays solutions onto the original image. The pipeline now processes Sudoku images from camera capture to solved puzzle overlay with clear visual diagnostics at every stage. -In the next step of the learning path, we will focus on optimization and deployment. +Next, you'll focus on optimization techniques to improve model performance and prepare the system for efficient deployment on Arm64 and mobile platforms. \ No newline at end of file diff --git a/content/learning-paths/mobile-graphics-and-gaming/onnx/07_optimisation.md b/content/learning-paths/mobile-graphics-and-gaming/onnx/07_optimisation.md index a8e38c6ca7..201fc43580 100644 --- a/content/learning-paths/mobile-graphics-and-gaming/onnx/07_optimisation.md +++ b/content/learning-paths/mobile-graphics-and-gaming/onnx/07_optimisation.md @@ -1,22 +1,26 @@ --- -title: "Model Enhancements and Optimizations" +title: "Optimize the model for Arm64 deployment" + weight: 8 + layout: "learningpathall" --- ## Objective -In this section, we improve the Sudoku system from a working prototype into something that is faster, smaller, and more robust on Arm64-class hardware. We start by measuring a baseline, then apply ONNX Runtime optimizations and quantization, and finally address the most common real bottleneck: image preprocessing. At each step we re-check accuracy and solve rate so performance gains don’t come at the cost of correctness. + +In this section, you'll transform the Sudoku system from a working prototype into one that is faster, smaller, and more robust on Arm64 hardware. Start by measuring a baseline, then apply ONNX Runtime optimizations and quantization, and finally address the most common bottleneck: image preprocessing. At each step, re-check accuracy and solve rate so performance gains don't come at the cost of correctness. ## Establish a baseline Before applying any optimizations, it is essential to understand where time is actually being spent in the Sudoku pipeline. Without this baseline, it is impossible to tell whether an optimization is effective or whether it simply shifts the bottleneck elsewhere. -In the current system, the total latency of processing a single Sudoku image is composed of four main stages: -* Grid detection and warping – locating the outer Sudoku grid and rectifying it using a perspective transform. This step relies entirely on OpenCV and depends on image resolution, lighting, and grid clarity. -* Cell preprocessing – converting each of the 81 cells into a normalized 28Γ—28 grayscale input for the neural network. This includes cropping margins, thresholding, and morphological operations. In practice, this stage is often the dominant cost. -* ONNX inference – running the digit recognizer on all 81 cells as a single batch. Thanks to dynamic batch support, this step is typically fast compared to preprocessing. -* Solving – applying a backtracking Sudoku solver to the recognized board. This step is usually negligible in runtime, unless recognition errors lead to difficult or contradictory boards. +The total latency of processing a single Sudoku image is composed of four main stages: -To quantify these contributions, we will add simple timing measurements around each stage of the pipeline using a high-resolution clock (time.perf_counter()). For each processed image, we will print a breakdown: +* **Grid detection and warping:** locating the outer Sudoku grid and rectifying it using a perspective transform. Depends on image resolution, lighting, and grid clarity. Uses OpenCV only. +* **Cell preprocessing:** converting each of the 81 cells into a normalized 28Γ—28 grayscale input for the neural network. Includes cropping margins, thresholding, and morphological operations. Often the dominant cost. +* **ONNX inference:** running the digit recognizer on all 81 cells as a single batch. Typically fast thanks to dynamic batch support. +* **Solving:** applying a backtracking Sudoku solver to the recognized board. Usually negligible in runtime, unless recognition errors lead to difficult or contradictory boards. + +To quantify these contributions, you will add simple timing measurements around each stage of the pipeline using a high-resolution clock (time.perf_counter()). For each processed image, you will print a breakdown: * warp_ms – time spent on grid detection and perspective rectification * preprocess_ms – total time spent preprocessing all 81 cells * onnx_ms – time spent running batched ONNX inference @@ -25,13 +29,13 @@ To quantify these contributions, we will add simple timing measurements around e * total_ms – end-to-end processing time ## Performance measurements -Open the sudoku_processor.py and add the following import +In `sudoku_processor.py`, add the following import ```python import time ``` -Then, modify the process_image as follows +Then, modify the `process_image` function as follows ```python def process_image(self, bgr: np.ndarray, overlay: bool = True): """ @@ -113,7 +117,7 @@ def process_image(self, bgr: np.ndarray, overlay: bool = True): return board, (solved if ok else None), debug, overlay_img ``` -Finally, print the timings in the 05_RunSudokuProcessor.py: +Finally, print the timings in the `05_RunSudokuProcessor.py` as shown: ```python def main(): # Use any image path you like: @@ -156,10 +160,12 @@ if __name__ == "__main__": main() ``` -The sample output will look as follows: -```output +Run the script: +```console python3 05_RunSudokuProcessor.py - +``` +The output will look like: +```output Recognized board . . . | 7 . . | 6 . . . . 4 | . . . | 1 . 9 @@ -194,9 +200,9 @@ warp=11.9 ms | preprocess=3.3 ms | onnx=1.9 ms | solve=3.1 ms | total=48.2 ms ## Folder benchmark The single-image measurements introduced earlier are useful for understanding the rough structure of the pipeline and for verifying that ONNX inference is not the main computational bottleneck. In our case, batched ONNX inference typically takes less than 2 ms, while grid detection, warping, and preprocessing dominate the runtime. However, individual measurements can be noisy due to caching effects, operating system scheduling, and Python overhead. -To obtain more reliable performance numbers, we extend the evaluation to multiple images and compute aggregated statistics. This allows us to track not only average performance, but also variability and tail latency, which are particularly important for interactive applications. +To obtain more reliable performance numbers, you can extend the evaluation to multiple images and compute aggregated statistics. This allows us to track not only average performance, but also variability and tail latency, which are particularly important for interactive applications. -To do this, we add two helper functions to 05_RunSudokuProcessor.py, and make sure you have import glob and import numpy as np at the top of the runner script. +To do this, add two helper functions to `05_RunSudokuProcessor.py`, and make sure you have `import glob` and `import numpy as np` at the top of the runner script. The first function, summarize, computes basic statistics from a list of timing measurements: * mean – average runtime @@ -258,7 +264,7 @@ def benchmark_folder(proc, folder_glob, limit=100, warmup=10, overlay=False): print(f"{k:14s} mean={s['mean']:.2f} median={s['median']:.2f} p90={s['p90']:.2f} p95={s['p95']:.2f}") ``` -Finally, we invoke the benchmark in the main() function: +Finally, invoke the benchmark in the main() function: ```python def main(): @@ -276,10 +282,12 @@ This evaluates the processor on a representative subset of camera-like validatio Aggregated benchmarks provide a much more accurate picture than single measurements, especially when individual stages take only a few milliseconds. By reporting median and tail latencies, you can see whether occasional slow cases exist and whether an optimization truly improves user-perceived performance. Percentiles are particularly useful when a few slow cases exist (e.g., harder solves), because they reveal tail latency. These results form a solid quantitative baseline that you can reuse to evaluate every optimization that follows. +Run the updated script: +```console +python3 05_RunSudokuProcessor.py +``` Here is the sample output of the updated script: ```output -python3 05_RunSudokuProcessor.py - Solved 30/30 (100.0%) Timing summary (ms): @@ -290,12 +298,14 @@ solve_ms mean=74.76 median=2.02 p90=48.51 p95=74.82 total_ms mean=89.41 median=16.97 p90=62.95 p95=89.43 ``` -Notice that solve_ms (and therefore total_ms) has a much larger mean than median. This indicates a small number of outliers where the solver takes significantly longer. In practice, this occurs when one or more digits are misrecognized, forcing the backtracking solver to explore many branches before finding a solution (or failing). For interactive applications, median and p95 latency are more informative than the mean, as they better reflect typical user experience. +{{% notice Note %}} These measurements were obtained on a MacBook Pro with Apple M3 Pro running macOS 14.5. Performance on other Arm64 platforms (Raspberry Pi 5, AWS Graviton, etc.) will vary based on CPU performance and memory bandwidth. The relative distribution of time across pipeline stages should remain similar.{{% /notice %}} + +Notice that `solve_ms` (and therefore `total_ms`) has a much larger mean than median. This indicates a small number of outliers where the solver takes significantly longer. In practice, this occurs when one or more digits are misrecognized, forcing the backtracking solver to explore many branches before finding a solution (or failing). For interactive applications, median and p95 latency are more informative than the mean, as they better reflect typical user experience. ## ONNX Runtime session optimizations Now that you can measure onnx_ms and total_ms, the first low-effort improvement is to enable ONNX Runtime’s built-in graph optimizations and tune CPU threading. These changes do not modify the model, but can reduce inference overhead and improve throughput. -In sudoku_processor.py, update the ONNX Runtime session initialization in __init__ to use SessionOptions: +In `sudoku_processor.py`, update the ONNX Runtime session initialization in __init__ to use SessionOptions: ```python so = ort.SessionOptions() so.graph_optimization_level = ort.GraphOptimizationLevel.ORT_ENABLE_ALL @@ -303,11 +313,9 @@ so.graph_optimization_level = ort.GraphOptimizationLevel.ORT_ENABLE_ALL self.sess = ort.InferenceSession(onnx_path, sess_options=so, providers=list(providers)) ``` -Re-run 05_RunSudokuProcessor.py and compare onnx_ms and total_ms to the baseline. +Re-run `05_RunSudokuProcessor.py` and compare onnx_ms and total_ms to the baseline. ```output -python3 05_RunSudokuProcessor.py - Solved 30/30 (100.0%) Timing summary (ms): @@ -323,7 +331,7 @@ This result is expected for such a small model: ONNX inference is already effici ## Quantize the model (FP32 -> INT8) Quantization is one of the most impactful optimizations for Arm64 and mobile deployments because it reduces both model size and compute cost. For CNNs, the most compatible approach is static INT8 quantization in QDQ format. This uses a small calibration set to estimate activation ranges and typically works well across runtimes. -Create a small script 06_QuantizeModel.py: +Create a small script `06_QuantizeModel.py` with the code below: ```python import os, glob @@ -380,7 +388,10 @@ quantize_static( print("Saved:", INT8_PATH) ``` -Run python 06_QuantizeModel.py +Run the script: +```console +python3 06_QuantizeModel.py +``` Then update the runner script to point to the quantized model: @@ -397,7 +408,8 @@ Also compare file sizes: ```console ls -lh artifacts/sudoku_digitnet.onnx artifacts/sudoku_digitnet.int8.onnx ``` -Even when inference time changes only modestly, size reduction is typically significant and matters for Android packaging. + +Expected file size reduction is approximately 4x (for example, from 52KB to 14KB). Even when inference time changes only modestly, size reduction is significant and matters for Android packaging. In this pipeline, quantization primarily reduces model size and improves deployability, while runtime speedups may be modest because inference is already a small fraction of the total latency. @@ -405,20 +417,15 @@ In this pipeline, quantization primarily reduces model size and improves deploya The measurements above show that ONNX inference accounts for only a small fraction of the total runtime. In practice, the largest performance gains come from optimizing image preprocessing. The most effective improvements include: -- Converting the rectified board to grayscale **once**, instead of converting each cell independently. +- Converting the rectified board to grayscale once, instead of converting each cell independently. - Adding an early β€œblank cell” check to skip expensive thresholding and morphology for empty cells. - Using simpler thresholding (e.g., Otsu) on clean images, and reserving adaptive thresholding for difficult lighting conditions. - Reducing or conditionally disabling morphological operations when cells already appear clean. These changes typically reduce `preprocess_ms` more than any model-level optimization, and therefore have the greatest impact on end-to-end latency. -## Summary -In this section, we transformed the Sudoku solver from a functional prototype into a system with measurable, well-understood performance characteristics. By instrumenting the pipeline with fine-grained timing, we identified where computation is actually spent and established a quantitative baseline. +## What you've learned and what's next -We showed that: -- Batched ONNX inference is already efficient (β‰ˆ1–2 ms per board). -- Image preprocessing dominates runtime and offers the largest optimization potential. -- Solver backtracking introduces rare but significant tail-latency outliers. -- ONNX Runtime optimizations and INT8 quantization improve deployability, even when raw inference speed gains are modest. +You transformed the Sudoku solver from a functional prototype into a system with measurable, well-understood performance characteristics. You established quantitative baselines showing that ONNX inference takes approximately 1–2 ms per board, identified image preprocessing as the dominant cost (~3 ms) and the largest optimization opportunity, applied INT8 quantization achieving approximately 4x model size reduction, and demonstrated a systematic optimization workflow where you measure first, optimize second, and always re-validate correctness. -Most importantly, we demonstrated a systematic optimization workflow: **measure first, optimize second, and always re-validate correctness**. With performance, robustness, and accuracy validated, the Sudoku pipeline is now ready for its final stepβ€”deployment as a fully on-device Android application. \ No newline at end of file +Next, you'll deploy the optimized Sudoku pipeline as a fully on-device Android application, integrating the ONNX model with camera capture and real-time processing. \ No newline at end of file diff --git a/content/learning-paths/mobile-graphics-and-gaming/onnx/08_android.md b/content/learning-paths/mobile-graphics-and-gaming/onnx/08_android.md index 06933a8fd8..887ecae6cf 100644 --- a/content/learning-paths/mobile-graphics-and-gaming/onnx/08_android.md +++ b/content/learning-paths/mobile-graphics-and-gaming/onnx/08_android.md @@ -1,33 +1,34 @@ --- -# User change -title: "Android Deployment. From Model to App" +title: "Deploy the model to Android" weight: 9 layout: "learningpathall" --- -## Objective ## -In this section, we transition from a desktop prototype to a fully on-device Android application. The goal is to demonstrate how the optimized Sudoku pipelineβ€”image preprocessing, ONNX inference, and deterministic solvingβ€”can be packaged and executed entirely on a mobile device, without relying on any cloud services. +## Objective -Rather than starting with a live camera feed, we begin with a fixed input bitmap that was generated earlier in the learning path. This approach allows us to focus on correctness, performance, and integration details before introducing additional complexity such as camera permissions, real-time capture, and varying lighting conditions. By keeping the input controlled, we can verify that the Android implementation faithfully reproduces the behavior observed in Python. +You'll now transition from desktop deployment to a fully on-device Android application. The goal is to demonstrate how the optimized Sudoku pipeline--image preprocessing, ONNX inference, and deterministic solving--can be packaged and executed entirely on a mobile device, without relying on cloud services. -Over the course of this section, we will: -1. Create a new Android project and add the required dependencies. -2. Bundle the trained ONNX model and a sample Sudoku image with the application. -3. Implement a minimal user interface that loads the image and triggers the solver. -4. Re-implement the Sudoku processing pipeline on Android, including preprocessing, batched ONNX inference, and solving. -5. Display the solved result as an image, confirming that the entire pipeline runs locally on the device. +Rather than starting with a live camera feed, you begin with a fixed input bitmap generated earlier in the Learning Path. This approach allows you to focus on correctness, performance, and integration details before introducing additional complexity such as camera permissions, real-time capture, and varying lighting conditions. -By the end of this section, you will have a working Android app that takes a Sudoku image, runs neural network inference and solving on-device, and displays the solution. This completes the learning path by showing how a trained and optimized ONNX model can be deployed in a real mobile application, closing the loop from data generation and training to practical, end-user deployment. +In this section, you will: + +* Create a new Android project and configure dependencies for ONNX Runtime and OpenCV +* Bundle the trained ONNX model and sample Sudoku images as application assets +* Build a minimal user interface that loads images and triggers the solver +* Implement the Sudoku processing pipeline on Android, including grid detection, digit recognition, and solving +* Display the solved result overlaid on the original image, confirming end-to-end on-device execution + +By the end of this section, you will have a working Android app that takes a Sudoku image, runs neural network inference and solving on-device, and displays the solution. ## Project creation -We start by creating a new Android project using Android Studio. This project will host the Sudoku solver application and serve as the foundation for integrating ONNX Runtime and OpenCV. +Start by creating a new Android project using Android Studio. This project will host the Sudoku solver application and serve as the foundation for integrating ONNX Runtime and OpenCV. 1. Create a new project: * Open Android Studio and click New Project. * In the Templates screen, select Phone and Tablet, then choose Empty Views Activity. -![img3](figures/03.webp) +![Android Studio create project screen showing template selection with Empty Views Activity highlighted](figures/03.webp) This template creates a minimal Android application without additional UI components, which is ideal for a focused, step-by-step integration. @@ -41,12 +42,12 @@ This template creates a minimal Android application without additional UI compon * Minimum SDK: API 24 (Android 7.0 – Nougat). This provides wide device coverage while remaining compatible with ONNX Runtime and OpenCV. * Build configuration language: Kotlin DSL (build.gradle.kts). We use the Kotlin DSL for Gradle, which is now the recommended option. -![img4](figures/04.webp) +![Android Studio project configuration screen showing project name SudokuSolverOnnx with Kotlin selected and minimum SDK set to API 24](figures/04.webp) * After confirming these settings, click Finish. Android Studio will create the project and generate a basic MainActivity along with the necessary Gradle files. ## View -We now define the user interface of the Android application. The goal of this view is to remain intentionally simple while clearly exposing the end-to-end Sudoku workflow. The interface will consist of: +You will now define the user interface of the Android application. The goal of this view is to remain intentionally simple while clearly exposing the end-to-end Sudoku workflow. The interface will consist of: * A button row at the top that allows the user to load a Sudoku image and trigger the solver. * A status text area used to display short messages (for example, whether an image has been loaded or the puzzle has been solved). * An input image view that displays the selected Sudoku bitmap. @@ -201,12 +202,12 @@ Because the image area is scrollable, the layout remains usable even on smaller When rendered, this produces a clear, vertically structured interface with a fixed control panel at the top and large input/output images underneath, as shown in the figure below. -![img](figures/05.png) +![Android app user interface showing two buttons at the top (Load image and Solve), a status text field, and placeholder areas for input and output Sudoku images](figures/05.png) -At this stage, the UI is intentionally minimal. In the next step, we will connect this view to the application logic in MainActivity, load a sample Sudoku bitmap, and wire up the Load image and Solve buttons to the ONNX-based processing pipeline. +At this stage, the UI is intentionally minimal. In the next step, you will connect this view to the application logic in MainActivity, load a sample Sudoku bitmap, and wire up the Load image and Solve buttons to the ONNX-based processing pipeline. ## Preparing input images for the Android app -Before wiring the application logic, we need to provide the Android app with a small set of Sudoku images that it can load and solve. For this learning path, we deliberately use a fixed collection of pre-generated images (from Preparing a Synthetic Sudoku Digit Dataset) instead of a camera feed. This keeps the Android integration simple and allows us to focus on ONNX inference and solver integration first. +Before wiring the application logic, you need to provide the Android app with a small set of Sudoku images that it can load and solve. For this learning path, you will use a fixed collection of pre-generated images (from Preparing a Synthetic Sudoku Digit Dataset) instead of a camera feed. This keeps the Android integration simple and allows us to focus on ONNX inference and solver integration first. 1. Select Sudoku images. From the earlier Python steps, select a small number of generated Sudoku images. These can be either: * Clean grids (book-style), or @@ -231,7 +232,7 @@ Rename your files accordingly, for example: 3. Copy the renamed PNG files into the following directory of your Android project: -app/src/main/res/drawable/ +`app/src/main/res/drawable/` After copying, Android Studio will automatically generate resource IDs for these images. @@ -247,7 +248,7 @@ R.drawable.sudoku_cam_01 In this tutorial, the Load image button will randomly select one of these drawable resources and display it in the app. This provides a deterministic and repeatable input source while validating the full Sudoku pipeline on Android. -At this point, the Android project has all the static resources it needs. In the next step, we will implement MainActivity.kt, wire up the Load image and Solve buttons, and display the selected Sudoku image in the UI. +At this point, the Android project has all the static resources it needs. In the next step, you will implement MainActivity.kt, wire up the Load image and Solve buttons, and display the selected Sudoku image in the UI. ## Preparing the ONNX model for the Android app In addition to the input images, the Android application needs access to the trained ONNX model so that it can run inference directly on the device. Android does not allow arbitrary file access by default, so the model must be bundled with the app as an asset. @@ -266,9 +267,9 @@ If you prefer to keep assets organized, you can also create a subfolder: app/src/main/assets/models/ -Both approaches work. In the examples that follow, we will assume the model is placed directly under assets/. +Both approaches work. In the examples that follow, you will assume the model is placed directly under assets/. -3. Prepare the model for android such that it does not contain external resources. Android assets are not normal filesystem paths, so models that reference external.onnx.data files will fail to load unless they are merged into a single ONNX file. To do so, create another Python file 07_PrepareModelForAndroid.py: +3. Prepare the model for android such that it does not contain external resources. Android assets are not normal filesystem paths, so models that reference external.onnx.data files will fail to load unless they are merged into a single ONNX file. To do so, create another Python file `07_PrepareModelForAndroid.py`: ```python import onnx from onnx import external_data_helper @@ -319,7 +320,7 @@ This input stream will be passed to ONNX Runtime to create an inference session * A set of Sudoku images in res/drawable/, * A trained ONNX model in assets/. -In the next step, we will implement MainActivity.kt, wire up the Load image and Solve buttons, and verify that the app can successfully load both the image and the ONNX model before running inference. +In the next step, you will implement MainActivity.kt, wire up the Load image and Solve buttons, and verify that the app can successfully load both the image and the ONNX model before running inference. ## Implement MainActivity.kt (Load image + basic UI wiring) Open app/src/main/java/com/arm/sudokusolveronnx/MainActivity.kt and replace it with: @@ -494,10 +495,10 @@ After making this change: The project should now compile and launch successfully. -At this point, the app should start, display the UI, and allow you to load random Sudoku images. In the next step, we will replace the placeholder logic in the Solve button with the real ONNX- and OpenCV-based Sudoku processing engine. +At this point, the app should start, display the UI, and allow you to load random Sudoku images. In the next step, you will replace the placeholder logic in the Solve button with the real ONNX- and OpenCV-based Sudoku processing engine. ## Processing pipeline on Android -With the user interface and static resources in place, we can now wire the full Sudoku processing pipeline on Android. Conceptually, this pipeline mirrors the Python implementation developed earlier in the learning path, but is reimplemented using Android-compatible components. +With the user interface and static resources in place, you can now wire the full Sudoku processing pipeline on Android. Conceptually, this pipeline mirrors the Python implementation developed earlier in the learning path, but is reimplemented using Android-compatible components. The pipeline consists of four stages: 1. Grid detection and rectification (OpenCV). The input bitmap is converted to an OpenCV matrix, the Sudoku grid is detected, and a perspective transform is applied to obtain a top-down, square view of the board. @@ -506,12 +507,12 @@ The pipeline consists of four stages: 4. Rendering and overlay. The solution is rendered back onto the original image by inverse-warping a transparent overlay from the rectified grid space to the input image. ### Dependencies -To support this pipeline, we add three dependencies: +To support this pipeline, add three dependencies: * ONNX Runtime for on-device inference, * OpenCV for image processing and geometric transformations, * Kotlin coroutines to ensure that heavy computation runs off the UI thread. -We open build.gradle.kts and add the following +Open build.gradle.kts and add the following ```text dependencies { implementation("com.microsoft.onnxruntime:onnxruntime-android:1.18.0") @@ -1035,25 +1036,38 @@ To test the app: The figures below show two representative test cases. In each example, the upper image corresponds to the original Sudoku puzzle, while the lower image shows the same puzzle with the missing digits filled in and overlaid in green. This visual comparison confirms that grid detection, digit recognition, solving, and rendering are all functioning correctly on-device. -![img](figures/06.png) -![img](figures/07.png) +![Android app running on device showing original Sudoku puzzle in top half and solved puzzle with green digits filled in bottom half](figures/06.png) +![Second example of Android app showing different Sudoku puzzle and its solution with green overlay digits](figures/07.png) These tests demonstrate that the application is robust to perspective distortion and partial digit placement, and that the model performs reliably when deployed via ONNX Runtime on Android. -## Summary and next steps -In this learning path, you have built a complete, end-to-end workflow for deploying machine learning models with ONNX on Arm64 and mobile devices. Starting from model development in Python, you moved step by step through export, optimization, and integration, ultimately deploying a fully functional solution that runs entirely on an Android device. +## What you've accomplished and what's next + +You built an end-to-end workflow for deploying ONNX models on Arm64 and mobile platforms--from model development in Python to on-device Android deployment. Throughout this Learning Path, you trained a lightweight digit recognizer and constructed an OpenCV-based pipeline for grid detection and rectification, exported models to ONNX with dynamic batch support, validated correctness with ONNX Runtime, applied practical optimizations including quantization and session tuning, and integrated everything into an on-device Android app with no cloud dependency. The final application processes Sudoku images from capture to completed solution, all running locally on the device. + +### Real-world considerations -Along the way, you trained and exported a neural network to the ONNX format, explored how to optimize inference for edge deployment, and built a robust vision pipeline using OpenCV. You then brought these components together on Android by integrating ONNX Runtime and implementing a Sudoku solver that performs image preprocessing, neural network inference, and deterministic solving entirely on-device, without any cloud dependency. +Challenging lighting, strong perspective distortion, faint or occluded digits, and imperfect crops can cause recognition errors that propagate into the solver, increasing solve time or occasionally preventing a valid solution. These behaviors highlight the trade-offs between model size, robustness, and performance at the edge. -It is also important to recognize the current limitations of the approach. While the system performs well on most test images, there are cases where digit recognition may fail or produce ambiguous resultsβ€”particularly under challenging lighting conditions, strong perspective distortion, or when digits are faint or partially occluded. In such cases, recognition errors can propagate to the solver, leading to longer solve times or, occasionally, failure to find a valid solution. These limitations are typical for lightweight, on-device vision systems and highlight the trade-offs between model complexity, robustness, and performance at the edge. +### Potential improvements -Despite these constraints, the application stands as a practical and self-contained example of edge AI deployment. It demonstrates how ONNX can serve as a bridge between model development and real-world deployment, enabling the same model to move seamlessly from a desktop environment to a mobile platform. +Next steps to enhance the application: -From here, there are many natural directions for improvement. You could enhance robustness by incorporating additional training data or more advanced preprocessing, extend the app to use live camera input with CameraX, refine the user experience with animations and progress feedback, or experiment with quantized models on devices that support additional execution providers. The same architectural pattern can also be applied beyond Sudoku, to other document- or grid-based vision problems where lightweight, on-device inference is essential. +* Add more varied training data and tighter preprocessing for improved robustness +* Implement live capture with CameraX +* Polish the UX with progress feedback and overlays +* Experiment with quantized models and hardware acceleration where available +* Generalize to other document- and grid-based vision tasks -This concludes the learning path and provides a solid foundation for building, optimizing, and deploying ONNX-based machine learning applications on Arm64 and mobile platforms. +This Learning Path provides a solid foundation for building, optimizing, and deploying ONNX-based machine learning applications on Arm64 and mobile platforms. ## Companion code -You can find the companion code in these repositories: +All source code used throughout this learning path is available in the following repositories: 1. [Sudoku solver](https://github.com/dawidborycki/SudokuSolverOnnx.git) 2. [Python scripts](https://github.com/dawidborycki/ONNX-LP.git) + +## What you've learned and what's next + +Throughout this Learning Path, you built a complete machine learning deployment pipeline from training to mobile deployment. You created and trained a digit recognition model in PyTorch, exported it to the portable ONNX format with dynamic batch support, and validated inference correctness across frameworks. You then integrated the model into an end-to-end Sudoku processing system using OpenCV for image processing and grid detection, applied performance optimizations including INT8 quantization, and measured real-world latency characteristics on Arm64 hardware. Finally, you deployed the entire pipeline as a fully functional Android application that performs on-device inference with ONNX Runtime, demonstrating how to bring machine learning models from development to production on Arm-based mobile platforms without cloud dependencies. + +You can now apply these techniques to other vision-based applications on Arm platforms, experiment with different model architectures and optimization strategies, or extend this foundation to solve similar document processing and grid-based recognition tasks. \ No newline at end of file diff --git a/content/learning-paths/mobile-graphics-and-gaming/onnx/_index.md b/content/learning-paths/mobile-graphics-and-gaming/onnx/_index.md index ba5b1ed031..aeee968f4c 100644 --- a/content/learning-paths/mobile-graphics-and-gaming/onnx/_index.md +++ b/content/learning-paths/mobile-graphics-and-gaming/onnx/_index.md @@ -1,31 +1,27 @@ --- -title: "ONNX in Action: Building, Optimizing, and Deploying Models on Arm64 and Mobile" +title: "Build, optimize, and deploy ML models with ONNX on Arm64 and mobile devices" -draft: true -cascade: - draft: true - minutes_to_complete: 240 -who_is_this_for: This is an introductory topic for developers who are interested in creating, optimizing, and deploying machine learning models with ONNX. It is especially useful for those targeting Arm64-based devices (such as Raspberry Pi, mobile SoCs, or Android smartphones) and looking to run efficient inference at the edge. +who_is_this_for: This is an advanced topic for developers who want to build, optimize, and deploy machine learning models using ONNX on Arm64-based platforms such as Raspberry Pi, Arm-based laptops, cloud instances, or Android smartphones. learning_objectives: - - Describe what ONNX is, and what it can offer in the ML ecosystem. - - Build and export a simple neural network model in Python to ONNX format. - - Perform inference and training using ONNX Runtime. - - Apply optimization techniques to improve performance. - - Deploy an optimized ONNX model inside an Android app. + - Explain what ONNX is and how it enables model portability across ML frameworks + - Build and export a neural network model in Python to ONNX format + - Run inference using ONNX Runtime on Arm64 platforms + - Apply model optimization techniques to improve performance + - Deploy an optimized ONNX model in an Android application prerequisites: - - A development machine with Python 3.10 installed (Python 3.11 also works; Python 3.12 is not yet supported by ONNX Runtime on Arm platforms). - - Basic familiarity with PyTorch or TensorFlow. - - An Arm64 device (e.g., Raspberry Pi or Android smartphone). - - "[Android Studio](https://developer.android.com/studio) installed for deployment testing." + - A development machine with Python 3.10 or 3.11 installed (Prebuilt ONNX Runtime packages for Arm platforms don't yet support Python 3.12) + - Basic familiarity with PyTorch or TensorFlow + - An Arm64 device such as a Raspberry Pi or Android smartphone + - Android Studio (required only for the final deployment section) author: Dawid Borycki ### Tags -skilllevels: Introductory +skilllevels: Advanced subjects: ML armips: - Cortex-A @@ -34,6 +30,7 @@ operatingsystems: - Windows - Linux - macOS + - Android tools_software_languages: - Python - PyTorch @@ -42,7 +39,6 @@ tools_software_languages: - Android - Android Studio - Kotlin - - Java further_reading: - resource: diff --git a/content/learning-paths/mobile-graphics-and-gaming/performance_llama_cpp_sme2/_index.md b/content/learning-paths/mobile-graphics-and-gaming/performance_llama_cpp_sme2/_index.md index d5002eb90a..1f00ab209a 100644 --- a/content/learning-paths/mobile-graphics-and-gaming/performance_llama_cpp_sme2/_index.md +++ b/content/learning-paths/mobile-graphics-and-gaming/performance_llama_cpp_sme2/_index.md @@ -22,9 +22,9 @@ author: Zenon Zhilong Xiu skilllevels: Advanced subjects: ML armips: - - Arm C1 core - - SME2 + - Arm C1 tools_software_languages: + - SME2 - C++ - llama.cpp operatingsystems: diff --git a/content/learning-paths/mobile-graphics-and-gaming/performance_onnxruntime_kleidiai_sme2/_index.md b/content/learning-paths/mobile-graphics-and-gaming/performance_onnxruntime_kleidiai_sme2/_index.md index 78448659e5..da1cb22881 100644 --- a/content/learning-paths/mobile-graphics-and-gaming/performance_onnxruntime_kleidiai_sme2/_index.md +++ b/content/learning-paths/mobile-graphics-and-gaming/performance_onnxruntime_kleidiai_sme2/_index.md @@ -24,10 +24,12 @@ author: Zenon Zhilong Xiu skilllevels: Advanced subjects: ML armips: - - Cortex + - Cortex-A + - Arm C1 tools_software_languages: - C++ - ONNX Runtime + - SME2 operatingsystems: - Android - Linux @@ -55,4 +57,4 @@ further_reading: weight: 1 # _index.md always has weight of 1 to order correctly layout: "learningpathall" # All files under learning paths have this same wrapper learning_path_main_page: "yes" # This should be surfaced when looking for related content. Only set for _index.md of learning path content. ---- \ No newline at end of file +--- diff --git a/content/learning-paths/mobile-graphics-and-gaming/voice-assistant/_index.md b/content/learning-paths/mobile-graphics-and-gaming/voice-assistant/_index.md index 2cc33fb851..201851cdff 100644 --- a/content/learning-paths/mobile-graphics-and-gaming/voice-assistant/_index.md +++ b/content/learning-paths/mobile-graphics-and-gaming/voice-assistant/_index.md @@ -25,10 +25,12 @@ skilllevels: Introductory subjects: Performance and Architecture armips: - Cortex-A + - Arm C1 tools_software_languages: - Java - Kotlin - CPP + - SME2 operatingsystems: - Android - Linux diff --git a/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/_index.md b/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/_index.md new file mode 100644 index 0000000000..e5dbb12e85 --- /dev/null +++ b/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/_index.md @@ -0,0 +1,74 @@ +--- +title: Deploy High-Performance Analytics with Apache Arrow and Arrow Flight on Google Cloud C4A Axion processors + +minutes_to_complete: 30 + +who_is_this_for: This is an introductory topic for data engineers, platform engineers, and developers who aim to build high-performance analytics pipelines on Arm64-based Google Cloud C4A Axion processors using Apache Arrow and Arrow Flight. + +learning_objectives: + - Deploy Apache Arrow–based data processing workloads on Google Cloud C4A Axion processors + - Set up and run an Arrow Flight server for high-throughput, low-latency data transport + - Read and write columnar data formats such as Parquet and ORC using Apache Arrow + - Integrate Arrow with object storage (MinIO) for cloud-native analytics workflows + - Validate performance benefits of Arrow and Arrow Flight on Arm-based infrastructure + +prerequisites: + - A [Google Cloud Platform (GCP)](https://cloud.google.com/free) account with billing enabled + - Basic familiarity with Python + - Basic understanding of data formats such as Parquet or ORC + - Familiarity with Linux command-line operations + +author: Pareena Verma + +##### Tags +skilllevels: Introductory +subjects: Performance and Architecture +cloud_service_providers: +- Google Cloud + +armips: +- Neoverse + +tools_software_languages: +- Apache Arrow +- Arrow Flight +- Python +- MinIO + +operatingsystems: +- Linux + +# ================================================================================ +# FIXED, DO NOT MODIFY +# ================================================================================ + +further_reading: + - resource: + title: Google Cloud documentation + link: https://cloud.google.com/docs + type: documentation + + - resource: + title: Apache Arrow documentation + link: https://arrow.apache.org/docs/ + type: documentation + + - resource: + title: Arrow Flight documentation + link: https://arrow.apache.org/docs/format/Flight.html + type: documentation + + - resource: + title: Apache Parquet documentation + link: https://parquet.apache.org/documentation/latest/ + type: documentation + + - resource: + title: MinIO documentation + link: https://min.io/docs + type: documentation + +weight: 1 +layout: "learningpathall" +learning_path_main_page: yes +--- diff --git a/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/_next-steps.md b/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/_next-steps.md new file mode 100644 index 0000000000..c3db0de5a2 --- /dev/null +++ b/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/_next-steps.md @@ -0,0 +1,8 @@ +--- +# ================================================================================ +# FIXED, DO NOT MODIFY THIS FILE +# ================================================================================ +weight: 21 # Set to always be larger than the content in this path to be at the end of the navigation. +title: "Next Steps" # Always the same, html page title. +layout: "learningpathall" # All files under learning paths have this same wrapper for Hugo processing. +--- diff --git a/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/background.md b/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/background.md new file mode 100644 index 0000000000..4198cd836c --- /dev/null +++ b/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/background.md @@ -0,0 +1,46 @@ +--- +title: Get started with Apache Arrow and Arrow Flight on Google Axion C4A +weight: 2 + +layout: "learningpathall" +--- + +## Explore Axion C4A Arm instances in Google Cloud + +Google Axion C4A is a family of Arm-based virtual machines built on Google’s custom Axion CPU, which is based on Arm Neoverse-V2 cores. Designed for high-performance and energy-efficient computing, these virtual machines offer strong performance for data-intensive and analytics workloads such as big data processing, in-memory analytics, columnar data processing, and high-throughput data services. + +The C4A series provides a cost-effective alternative to x86 virtual machines while leveraging the scalability, SIMD acceleration, and memory bandwidth advantages of the Arm architecture in Google Cloud. + +These characteristics make Axion C4A instances well-suited for modern analytics stacks that rely on columnar data formats and memory-efficient execution engines. + +To learn more, see the Google blog [Introducing Google Axion Processors, our new Arm-based CPUs](https://cloud.google.com/blog/products/compute/introducing-googles-new-arm-based-cpu). + +## Explore Apache Arrow and Arrow Flight on Google Axion C4A (Arm Neoverse V2) + +Apache Arrow is an open-source, cross-language platform for in-memory data processing. It defines a standardized columnar memory format that enables zero-copy data sharing between analytics systems, significantly improving performance for modern data workloads. + +Arrow is widely used as a foundational layer for analytics engines, data science tools, and big data frameworks, enabling efficient processing of formats such as Parquet and ORC while minimizing serialization overhead. + +Arrow Flight is a high-performance RPC framework built on gRPC that enables fast, memory-to-memory data transfer using the Arrow columnar format. It is designed to replace traditional REST- or file-based data exchange with a low-latency, high-throughput alternative optimized for analytics workloads. + +Running Apache Arrow and Arrow Flight on Google Axion C4A Arm-based infrastructure allows you to achieve: + +- High-throughput columnar data processing +- Reduced CPU overhead through zero-copy data exchange +- Improved performance per watt for analytics pipelines +- Cost-efficient scaling of in-memory data services + +Common use cases include interactive analytics, data lake acceleration, machine learning feature pipelines, distributed query engines, and high-performance data services. + +To learn more, visit the [Apache Arrow website](https://arrow.apache.org/) and explore the [Arrow Flight documentation](https://arrow.apache.org/docs/format/Flight.html). + +## What you've learned and what's next + +In this section, you learned about: + +* Google Axion C4A Arm-based VMs and their performance characteristics +* Apache Arrow’s columnar in-memory format and its role in modern analytics +* Arrow Flight and its use for high-speed, memory-to-memory data transfer +* How Arm architecture enables cost-effective, high-performance analytics workloads + +Next, you'll configure firewall rules and network access to allow external communication between MinIO object storage, Apache Arrow workloads, and Arrow Flight services running on your Axion C4A virtual machine. diff --git a/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/columnar-analytics-with-arrow.md b/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/columnar-analytics-with-arrow.md new file mode 100644 index 0000000000..2b645b946f --- /dev/null +++ b/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/columnar-analytics-with-arrow.md @@ -0,0 +1,241 @@ +--- +title: Analyze columnar data with Apache Arrow on arm64 +weight: 6 + +### FIXED, DO NOT MODIFY +layout: learningpathall +--- + +## Analyze columnar data with Apache Arrow + +In this section, you use **Apache Arrow’s columnar execution engine to read and write analytical datasets stored in MinIO (S3)**. You will work with **Parquet and ORC formats** and explore predicate pushdown and column pruning, which are key performance optimizations in modern analytics engines. + +This section demonstrates how Arrow delivers high-performance, vectorized analytics on arm64 (Axion). + +## Architecture overview + +```text +Python Analytics Scripts + | + v +Apache Arrow Dataset API + | + v +Parquet / ORC Columnar Files + | + v +MinIO (S3 Object Storage) +``` + +**What this architecture shows:** + +- Compute and execution happen in-memory using Apache Arrow +- Data is stored in object storage (MinIO) using open columnar formats +- Only required data is read from storage, reducing I/O and latency + +## Write Parquet data to MinIO + +In this step, you create a sample dataset in memory using Apache Arrow and write it to MinIO in Parquet format, the most common columnar format used in analytics engines. + +Create a file named `write_parquet.py`. + +```python +import pyarrow as pa +import pyarrow.parquet as pq +import s3fs + +table = pa.table({ + "id": list(range(1000)), + "value": [i * 10 for i in range(1000)] +}) + +fs = s3fs.S3FileSystem( + key="minioadmin", + secret="minioadmin", + client_kwargs={"endpoint_url": "http://127.0.0.1:9000"} +) + +pq.write_table(table, "arrow-data/dataset.parquet", filesystem=fs) +``` + +### Run it + +```bash +source arrow-venv/bin/activate +python write_parquet.py +``` + +### Verify in MinIO UI + +Open the arrow-data bucket in the MinIO console. Refresh if needed. Select the "arrow-data" bucket. You should see "dataset.parquet" listed: + +![MinIO object browser showing dataset.parquet stored inside the arrow-data bucket alt-txt#center](images/dataset-parquet.png "MinIO Web UI displaying dataset.parquet object in arrow-data bucket") + +**What this confirms:** + +- Apache Arrow successfully serialized in-memory data +- Parquet files were written directly to S3-compatible storage +- No local filesystem dependency is required + +## Read Parquet using Arrow Dataset API + +Next, you read the Parquet dataset using the Arrow Dataset API, which enables efficient scanning, filtering, and projection. + +Create a file named `read_parquet.py`. + +```python +import pyarrow.dataset as ds +import s3fs + +fs = s3fs.S3FileSystem( + key="minioadmin", + secret="minioadmin", + client_kwargs={"endpoint_url": "http://127.0.0.1:9000"} +) + +dataset = ds.dataset( + "arrow-data/dataset.parquet", + format="parquet", + filesystem=fs +) + +table = dataset.to_table() +print(table.schema) +print("Rows:", table.num_rows) +``` + +### Run the reader + +```bash +python read_parquet.py +``` + +The output is similar to: + +```output +id: int64 +value: int64 +Rows: 1000 +``` + +**What this demonstrates:** + +- Schema inference from Parquet metadata +- Efficient columnar scanning +- Fully vectorized execution on arm64 + +## Predicate pushdown and column pruning + +One of the biggest performance advantages of columnar formats is that queries can **push filters and column selection down to the storage layer**. + +Create a file named `filter_parquet.py`. + +```python +import pyarrow.dataset as ds +import s3fs + +fs = s3fs.S3FileSystem( + key="minioadmin", + secret="minioadmin", + client_kwargs={"endpoint_url": "http://127.0.0.1:9000"} +) + +dataset = ds.dataset( + "arrow-data/dataset.parquet", + format="parquet", + filesystem=fs +) + +filtered = dataset.to_table( + filter=ds.field("id") > 990, + columns=["id"] +) + +print(filtered) +``` + +### Run the filter script + +```bash +python filter_parquet.py +``` + +The output is similar to: + +```output +pyarrow.Table +id: int64 +---- +id: [[991,992,993,994,995,996,997,998,999]] +``` + +**This confirms:** + +- Predicate pushdown: Only rows with id > 990 are read +- Column pruning: Only the id column is loaded +- Vectorized execution: Processing happens in columnar batches + +These optimizations significantly reduce I/O and CPU usage for large datasets. + +## Write ORC data to MinIO + +In addition to Parquet, Apache Arrow also supports ORC, another popular columnar format widely used in Hive and Spark ecosystems. + +Create a file named `write_orc.py`. + +```python +import pyarrow as pa +import pyarrow.orc as orc +import s3fs + +table = pa.table({ + "id": list(range(1000)), + "value": [i * 10 for i in range(1000)] +}) + +fs = s3fs.S3FileSystem( + key="minioadmin", + secret="minioadmin", + client_kwargs={"endpoint_url": "http://127.0.0.1:9000"} +) + +with fs.open("arrow-data/dataset.orc", "wb") as f: + orc.write_table(table, f) + +print("ORC file written to MinIO") +``` + +### Run the ORC writer + +```bash +python write_orc.py +``` + +The output is similar to: + +```output +ORC file written to MinIO +``` + +**Verify in MinIO UI:** + +In the arrow-data bucket after refreshing, you should now see: + +- dataset.orc +- dataset.parquet + +![MinIO object browser showing dataset.orc stored inside the arrow-data bucket alt-txt#center](images/datset-orc.png "MinIO Web UI displaying dataset.orc object in arrow-data bucket") + +## What you've learned and what's next + +In this section, you have: + +- Written analytical datasets in Parquet and ORC formats +- Stored columnar data in S3-compatible object storage +- Used the Arrow Dataset API for efficient reads +- Applied predicate pushdown and column pruning +- Executed vectorized analytics optimized for arm64 (Axion) + +This forms the core analytics layer used by modern engines such as Spark, DuckDB, Trino, and Polars. + +In the next section, you will enable high-speed memory-to-memory analytics using Apache Arrow Flight, demonstrating gRPC-based data transfer, zero-copy serialization, and high-throughput analytics communication. diff --git a/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/environment-and-minio.md b/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/environment-and-minio.md new file mode 100644 index 0000000000..7d5b0cf2ca --- /dev/null +++ b/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/environment-and-minio.md @@ -0,0 +1,227 @@ +--- +title: Set up Apache Arrow and MinIO on arm64 +weight: 5 + +### FIXED, DO NOT MODIFY +layout: learningpathall +--- + +## Set up Apache Arrow environment and MinIO + +In this section, you prepare a SUSE Linux Enterprise Server (SLES) arm64 virtual machine and install the core components for high-performance analytics using Apache Arrow. You also deploy MinIO, an S3-compatible object storage service, to store analytical datasets in later sections. + +This foundation ensures all analytics libraries are natively optimized for arm64 (Axion). + +## Architecture overview + + +This architecture represents a single-node analytics environment that mirrors how modern cloud analytics stacks operate: +compute and memory-local processing with object storage–backed datasets. + +```text +SUSE Linux Enterprise Server (arm64) + | + v +Python 3.11 Virtual Environment + | + v +Apache Arrow Libraries + | + v +MinIO (S3-Compatible Object Storage) +``` + +## Install system dependencies on SUSE + +Install Python, build tools, and system libraries required by Apache Arrow and its ecosystem. + +```bash +sudo zypper refresh ; \ +sudo zypper install -y \ + python311 python311-devel python311-pip \ + gcc gcc-c++ make \ + libopenssl-devel \ + libuuid-devel \ + curl git +``` + +### Verify Python installation + +```bash +python3.11 --version +``` + +The output is similar to: + +```output +Python 3.11.10 +``` + +**Why this matters:** + +- Python 3.11 provides better performance and memory efficiency +- Apache Arrow wheels are fully supported on arm64 for Python 3.11 +- Ensures compatibility with modern analytics libraries + +## Create a Python virtual environment + +Create an isolated Python environment for Arrow and analytics libraries. + +```bash +python3.11 -m venv arrow-venv +source arrow-venv/bin/activate +``` + +### Upgrade core packaging tools + +```bash +pip install --upgrade pip setuptools wheel +``` + +**Why this matters:** + +- Avoids conflicts with the system Python +- Ensures reproducible analytics environments +- Recommended for production-grade data workloads + +## Install Apache Arrow and required libraries + +Install Apache Arrow and supporting analytics libraries. + +```bash +pip install \ + pyarrow \ + pandas \ + numpy \ + s3fs \ + grpcio \ + grpcio-tools \ + fastparquet \ + pyorc +``` + +### Verify Arrow installation + +```bash +python - <:9001` + +**Login credentials:** + +- **Username**: minioadmin +- **Password**: minioadmin + +![MinIO Web UI dashboard showing object browser and storage usage for arrow-data bucket alt-txt#center](images/minio-webui.png "MinIO Web UI displaying buckets and stored Parquet/ORC objects") + +### Create a bucket named + +`arrow-data` + +![MinIO Web UI bucket list view showing arrow-data bucket created successfully alt-txt#center](images/minio-bucket.png "MinIO Web UI displaying the arrow-data bucket") + +MinIO Bucket View + +This bucket will be used to store: + +- Parquet datasets +- ORC datasets +- Analytics output files + +### Configure S3 credentials for Python + +In another terminal (same VM, virtual environment active), export S3 credentials so Python libraries can access MinIO. + +```bash +export AWS_ACCESS_KEY_ID=minioadmin +export AWS_SECRET_ACCESS_KEY=minioadmin +export AWS_DEFAULT_REGION=us-east-1 +``` + +**Verify:** + +```bash +env | grep AWS +``` + +The output is similar to: + +```output +AWS_SECRET_ACCESS_KEY=minioadmin +AWS_DEFAULT_REGION=us-east-1 +AWS_ACCESS_KEY_ID=minioadmin +``` + +**What this enables:** + +- pyarrow +- s3fs +- pandas +- Other S3-compatible analytics libraries + +to communicate with MinIO exactly like Amazon S3. + +## What you've learned and what's next + +- Prepared a SUSE arm64 analytics environment +- Installed Apache Arrow and dependencies +- Deployed MinIO as S3-compatible object storage +- Configured secure access for analytics workloads + +In the next section, you will use Apache Arrow to write and read Parquet and ORC datasets from MinIO using vectorized analytics APIs. diff --git a/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/firewall-setup.md b/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/firewall-setup.md new file mode 100644 index 0000000000..65a0186090 --- /dev/null +++ b/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/firewall-setup.md @@ -0,0 +1,57 @@ +--- +title: Create firewall rules on GCP for Apache Arrow, MinIO, and Arrow Flight +weight: 3 + +### FIXED, DO NOT MODIFY +layout: learningpathall +--- + +## Configure GCP firewall for Apache Arrow analytics stack + +To allow inbound traffic for Apache Arrow-based analytics components, create a firewall rule in the Google Cloud Console. This firewall rule enables access to MinIO object storage and Arrow Flight for high-performance data transfer on Axion (arm64) virtual machines. + +{{% notice Note %}} +For more information about GCP setup, see [Getting started with Google Cloud Platform](/learning-paths/servers-and-cloud-computing/csp/google/). +{{% /notice %}} + +## Required ports + +| Service | Port | Purpose | +| ------- | ---- | ------- | +| MinIO S3 API | 9000 | S3-compatible object storage access | +| MinIO Web UI | 9001 | Bucket management and object browsing | +| Arrow Flight (gRPC) | 8815 | High-speed in-memory data transfer | + +## Create a firewall rule in GCP + +To expose the TCP ports listed above, create a firewall rule. + +Navigate to the [Google Cloud Console](https://console.cloud.google.com/), go to **VPC Network > Firewall**, and select **Create firewall rule**. + +![Google Cloud Console VPC Network Firewall page showing existing firewall rules and Create Firewall Rule button alt-txt#center](images/firewall-rule1.png "Create a firewall rule") + +Next, create the firewall rule that exposes the required TCP ports. +Set the **Name** of the new rule to `allow-arrow-minio-flight`. Select the network you intend to bind to your VM (the default is `default`, but your organization may use a different one). + +Set **Direction of traffic** to "Ingress". Set **Allow on match** to "Allow" and **Targets** to "Specified target tags". + +![Google Cloud Console firewall rule creation form showing name field, network selection, direction set to Ingress, and targets set to Specified target tags alt-txt#center](images/network-rule2.png "Creating Arrow firewall rule") + +Next, enter `allow-arrow-minio-flight` in the **Target tags** field. Set **Source IPv4 ranges** to `0.0.0.0/0`. + +![Google Cloud Console firewall rule form showing target tags field with allow-arrow-minio-flight entered and source IPv4 ranges set to 0.0.0.0/0 alt-txt#center](images/network-rule3.png "Creating the Arrow and MinIO firewall rule") + +Finally, select **Specified protocols and ports** under the **Protocols and ports** section. Select the **TCP** checkbox, enter `9000,9001,8815` in the **Ports** field, and select **Create**. + +![Google Cloud Console firewall rule form showing protocols and ports section with TCP selected and ports 9000,9001,8815 specified alt-txt#center](images/network-port.png "Specifying TCP ports for Apache Arrow and MinIO") + +## What you've learned and what's next + +You've successfully: + +- Created firewall rules in Google Cloud to expose ports for Apache Arrow analytics components +- Enabled external access to MinIO S3 API and Web UI +- Configured network access for Arrow Flight gRPC communication +- Prepared the network for a high-performance analytics stack on Axion (arm64) + +Next, you'll provision a Google Axion C4A Arm virtual machine and deploy Apache Arrow workloads, MinIO object storage, and Arrow Flight services on it. diff --git a/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/high-speed-analytics-with-arrow-flight.md b/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/high-speed-analytics-with-arrow-flight.md new file mode 100644 index 0000000000..d8f79666f4 --- /dev/null +++ b/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/high-speed-analytics-with-arrow-flight.md @@ -0,0 +1,144 @@ +--- +title: Run high-speed analytics with Apache Arrow Flight on arm64 +weight: 7 + +### FIXED, DO NOT MODIFY +layout: learningpathall +--- + +## Run high-speed analytics with Apache Arrow Flight + +In this section, you deploy Apache Arrow Flight, a high-performance RPC framework designed for analytics workloads. Arrow Flight enables zero-copy, memory-to-memory data transfer over gRPC, making it ideal for distributed analytics engines. + +Arrow Flight enables **zero-copy, memory-to-memory data transfer over gRPC**, allowing analytical data to be shared between processes and systems without serialization overhead. This makes it ideal for distributed analytics engines, interactive query systems, and real-time data pipelines. + +This section demonstrates how Arrow Flight works on arm64 (Axion) using a simple server-client setup on the same virtual machine. + +## Architecture overview + +```text +Arrow Table (In-Memory) + | + v +Arrow Flight Server (gRPC) + | + v +Arrow Flight Client +``` + +**What this architecture shows:** + +- Data remains in Arrow’s in-memory columnar format +- gRPC is used as the transport layer +- No intermediate files or object storage are involved +- Data is transferred efficiently with minimal CPU overhead + +## Start Arrow Flight server on the same virtual machine + +In this step, you create an Arrow Flight server that exposes an in-memory Arrow table to clients over gRPC. + +Create a file named `flight_server.py`. + + +```python +import pyarrow as pa +import pyarrow.flight as flight + +class ArrowFlightServer(flight.FlightServerBase): + def __init__(self, location): + super().__init__(location) + self.table = pa.table({ + "id": list(range(1000)), + "value": [i * 10 for i in range(1000)] + }) + + def do_get(self, context, ticket): + return flight.RecordBatchStream(self.table) + +if __name__ == "__main__": + server = ArrowFlightServer("grpc://0.0.0.0:8815") + print("Arrow Flight server running on port 8815") + server.serve() +``` + +### Run the server + +```bash +source arrow-venv/bin/activate +python flight_server.py +``` + +This terminal will block β€” this indicates the server is running and listening for client connections. + +The output is similar to: + +```output +(arrow-venv) gcpuser@arrow-flight:~> python flight_server.py +Arrow Flight server running on port 8815 +``` + +### Verify server is running + +Open another terminal on the same VM and check that the gRPC port is listening. + +```bash +ss -lntp | grep 8815 +``` + +If you see a listening process, the server is up. + +The output is similar to: + +```output +ss -lntp | grep 8815 +LISTEN 0 4096 *:8815 *:* users:(("python",pid=4255,fd=7)) +``` + +## Connect using an Arrow Flight client on the same virtual machine + +Now you connect to the Arrow Flight server using a client and retrieve the in-memory dataset. + +Create a file named `flight_client.py`. + +```python +import pyarrow.flight as flight + +client = flight.FlightClient("grpc://127.0.0.1:8815") + +reader = client.do_get(flight.Ticket(b"dataset")) +table = reader.read_all() + +print(table.schema) +print("Rows:", table.num_rows) +``` + +### Run it + +```bash +source arrow-venv/bin/activate +python flight_client.py +``` + +The output is similar to: + +```output +id: int64 +value: int64 +Rows: 1000 +``` + +**What this demonstrates:** + +- The client successfully connected over gRPC +- Data was transferred directly from server memory +- Arrow’s columnar format was preserved end-to-end + +## What you've learned and what's next + +In this section, you: + +- Started an Arrow Flight server that streams an in-memory Arrow table over gRPC +- Connected a client to the Flight endpoint and read data without file serialization +- Validated end-to-end transfer using Arrow's columnar in-memory format + +Next, review the Learning Path summary and continue with related Arm server analytics content in the next steps page. diff --git a/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/images/dataset-parquet.png b/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/images/dataset-parquet.png new file mode 100644 index 0000000000..8f1d2e2881 Binary files /dev/null and b/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/images/dataset-parquet.png differ diff --git a/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/images/datset-orc.png b/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/images/datset-orc.png new file mode 100644 index 0000000000..b783984f31 Binary files /dev/null and b/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/images/datset-orc.png differ diff --git a/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/images/firewall-rule1.png b/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/images/firewall-rule1.png new file mode 100644 index 0000000000..7f34211938 Binary files /dev/null and b/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/images/firewall-rule1.png differ diff --git a/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/images/gcp-pubip-ssh.png b/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/images/gcp-pubip-ssh.png new file mode 100644 index 0000000000..558745de3e Binary files /dev/null and b/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/images/gcp-pubip-ssh.png differ diff --git a/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/images/gcp-shell.png b/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/images/gcp-shell.png new file mode 100644 index 0000000000..7e2fc3d1b5 Binary files /dev/null and b/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/images/gcp-shell.png differ diff --git a/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/images/gcp-vm.png b/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/images/gcp-vm.png new file mode 100644 index 0000000000..0d1072e20d Binary files /dev/null and b/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/images/gcp-vm.png differ diff --git a/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/images/minio-bucket.png b/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/images/minio-bucket.png new file mode 100644 index 0000000000..9098622342 Binary files /dev/null and b/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/images/minio-bucket.png differ diff --git a/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/images/minio-webui.png b/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/images/minio-webui.png new file mode 100644 index 0000000000..0cff84cf01 Binary files /dev/null and b/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/images/minio-webui.png differ diff --git a/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/images/network-port.png b/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/images/network-port.png new file mode 100644 index 0000000000..00668b3a21 Binary files /dev/null and b/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/images/network-port.png differ diff --git a/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/images/network-rule2.png b/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/images/network-rule2.png new file mode 100644 index 0000000000..eda8c41e4a Binary files /dev/null and b/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/images/network-rule2.png differ diff --git a/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/images/network-rule3.png b/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/images/network-rule3.png new file mode 100644 index 0000000000..0035af507f Binary files /dev/null and b/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/images/network-rule3.png differ diff --git a/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/instance.md b/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/instance.md new file mode 100644 index 0000000000..51083e6a9c --- /dev/null +++ b/content/learning-paths/servers-and-cloud-computing/apache_arrow_and_flight/instance.md @@ -0,0 +1,57 @@ +--- +title: Create a Google Axion C4A arm64 virtual machine on GCP +weight: 4 + +### FIXED, DO NOT MODIFY +layout: learningpathall +--- + +## Provision a Google Axion C4A arm64 virtual machine + +In this section you'll create a Google Axion C4A arm64 virtual machine on Google Cloud Platform. You'll use the `c4a-standard-4` machine type, which provides 4 vCPUs and 16 GB of memory. This virtual machine hosts your Apache Arrow and Arrow Flight applications. + +{{% notice Note %}} +For help with GCP setup, see the Learning Path [Getting started with Google Cloud Platform](/learning-paths/servers-and-cloud-computing/csp/google/). +{{% /notice %}} + +## Provision the virtual machine in Google Cloud Console + +To create a virtual machine based on the C4A instance type: + +* Navigate to the [Google Cloud Console](https://console.cloud.google.com/). +* Go to **Compute Engine > VM Instances** and select **Create Instance**. +* Under **Machine configuration**: + * Populate fields such as **Instance name**, **Region**, and **Zone**. + * Set **Series** to `C4A`. + * Select `c4a-standard-4` for machine type. + +![Screenshot of the Google Cloud Console showing the Machine configuration section. The Series dropdown is set to C4A and the machine type c4a-standard-4 is selected alt-txt#center](images/gcp-vm.png "Configuring machine type to C4A in Google Cloud Console") + + +* Under **OS and storage**, select **Change**, and then choose an arm64 operating system image. + * For this Learning Path, select **SUSE Linux Enterprise Server**. + * For the license type, choose **Pay as you go**. + * Increase **Size (GB)** from **10** to **100** to allocate sufficient disk space. + * Select **Choose** to apply the changes. +* Under **Networking**, enable **Allow HTTP traffic** and **Allow HTTPS traffic** +* Also, add the following tag: `allow-arrow-minio-flight` so this virtual machine matches the firewall rule you created in the previous section. +* For some organizations not using the **'default'** network interface, you may need to select the network appropriate for your organization. +* Select **Create** to launch the virtual machine. + +After the instance starts, select **SSH** next to the VM in the instance list to open a browser-based terminal session. + +![Google Cloud Console VM instances page displaying running instance with green checkmark and SSH button in the Connect column alt-txt#center](images/gcp-pubip-ssh.png "Connecting to a running C4A VM using SSH") + +A new browser window opens with a terminal connected to your VM. + +![Browser-based SSH terminal window with black background showing Linux command prompt and Google Cloud branding at top alt-txt#center](images/gcp-shell.png "Terminal session connected to the VM") + +## What you've learned and what's next + +In this section: + +* You provisioned a Google Axion C4A arm64 virtual machine with 4 vCPUs and 16 GB of memory +* You configured the VM with SUSE Linux Enterprise Server and 100 GB of storage +* You connected to your VM using SSH through the Google Cloud Console + +Your virtual machine is now ready to host Apache Arrow and Arrow Flight. diff --git a/content/learning-paths/servers-and-cloud-computing/arm-soc-migration-learning-path/_index.md b/content/learning-paths/servers-and-cloud-computing/arm-soc-migration-learning-path/_index.md index 4fbcd591bc..04cf66b3a4 100644 --- a/content/learning-paths/servers-and-cloud-computing/arm-soc-migration-learning-path/_index.md +++ b/content/learning-paths/servers-and-cloud-computing/arm-soc-migration-learning-path/_index.md @@ -1,9 +1,5 @@ --- title: Migrate applications between Arm platforms using Kiro Arm SoC Migration Power - -draft: true -cascade: - draft: true minutes_to_complete: 60 diff --git a/content/learning-paths/servers-and-cloud-computing/arm-soc-migration-learning-path/graviton-development.md b/content/learning-paths/servers-and-cloud-computing/arm-soc-migration-learning-path/graviton-development.md index 9129d6679c..673ef1b9e3 100644 --- a/content/learning-paths/servers-and-cloud-computing/arm-soc-migration-learning-path/graviton-development.md +++ b/content/learning-paths/servers-and-cloud-computing/arm-soc-migration-learning-path/graviton-development.md @@ -22,7 +22,7 @@ tar -xzf sensor-monitor.tar.gz cd sensor-monitor ``` -The archive includes the complete source code, a Makefile, and platform-specific implementations. You will analyze and migrate this code using the Arm SoC Migration Power. +The archive includes the complete source code, a Makefile, and platform-specific implementations. You will analyze and migrate this code using the Perform Migration between Arm SoC Power. ### Upload to the Graviton instance for testing @@ -89,4 +89,4 @@ In this section: - You confirmed the toolchain and build process work correctly on Arm64 Linux - You established your baseline for migration validation -In the next section, you'll use the Arm SoC Migration Power to analyze the codebase and migrate it to the target platform. +In the next section, you'll use the Perform Migration between Arm SoC Power to analyze the codebase and migrate it to the target platform. diff --git a/content/learning-paths/servers-and-cloud-computing/arm-soc-migration-learning-path/migration.md b/content/learning-paths/servers-and-cloud-computing/arm-soc-migration-learning-path/migration.md index b3ffc1d1bb..2a8482de99 100644 --- a/content/learning-paths/servers-and-cloud-computing/arm-soc-migration-learning-path/migration.md +++ b/content/learning-paths/servers-and-cloud-computing/arm-soc-migration-learning-path/migration.md @@ -6,9 +6,9 @@ weight: 4 layout: learningpathall --- -## Use Arm SoC Migration Power for AI-guided migration +## Use Perform Migration between Arm SoC Power for AI-guided migration -In this section, you use the Arm SoC Migration Power to migrate your application between Arm-based platforms. The example demonstrates migration from: +In this section, you use the Perform Migration between Arm SoC Power to migrate your application between Arm-based platforms. The example demonstrates migration from: - **Source:** AWS Graviton3 (Neoverse-V1, Arm64 Linux, cloud deployment) - **Target:** Raspberry Pi 5 (BCM2712, Cortex-A76, edge deployment) @@ -19,11 +19,11 @@ For each step below, an example prompt for the Graviton-to-Pi-5 scenario is prov ### Initiate migration -Open Kiro and describe your migration clearly using the Arm SoC Migration Power. +Open Kiro and describe your migration clearly using the Perform Migration between Arm SoC Power. Example prompt: ``` -I want to use the Arm SoC Migration Power to migrate my sensor monitoring +I want to use the Perform Migration between Arm SoC Power to migrate my sensor monitoring application from AWS Graviton3 to Raspberry Pi 5 (BCM2712). The application currently uses simulated sensors on Graviton and needs to work with real GPIO and SPI hardware on the Pi 5. @@ -184,7 +184,7 @@ add_executable(sensor_monitor ${COMMON_SOURCES} ${PLATFORM_SOURCES}) In this section: -- You used the Arm SoC Migration Power to analyze architecture differences between Graviton3 and BCM2712 +- You used the Perform Migration between Arm SoC Power to analyze architecture differences between Graviton3 and BCM2712 - You designed a HAL that abstracts platform differences and preserves portability - You generated platform-specific code for the target device - You updated the build system for multi-platform support diff --git a/content/learning-paths/servers-and-cloud-computing/arm-soc-migration-learning-path/setup.md b/content/learning-paths/servers-and-cloud-computing/arm-soc-migration-learning-path/setup.md index f227bd94ed..8b75bbdb11 100644 --- a/content/learning-paths/servers-and-cloud-computing/arm-soc-migration-learning-path/setup.md +++ b/content/learning-paths/servers-and-cloud-computing/arm-soc-migration-learning-path/setup.md @@ -36,15 +36,14 @@ The Arm SoC Migration Power extends Kiro with specialized knowledge and tools fo - Open Kiro IDE - Navigate to the **Powers** panel. Press Cmd + Shift + P (Mac) or Ctrl + Shift + P (Windows) -- Select the Arm SoC Migration Power in the **Recommended** section +- Select the Perform Migration between Arm SoC in the **Recommended** section - Select **Install** -### Verify installation - -After installation, enter the following prompt in Kiro: +### Verify install +After installation, you can either click on the "Try power" button or enter the following prompt in Kiro: ```text -I just installed the Arm SoC Migration Power and want to use it. +I just installed the arm-soc-migration power and want to use it. ``` The Power should respond and guide you through any additional setup steps. @@ -58,7 +57,7 @@ It supports migrations across a wide range of Arm-based platforms, including: ### Install prerequisites -The Arm SoC Migration Power uses the Arm MCP (Model Context Protocol) server to provide specialized Arm migration capabilities. The Arm MCP server runs via Docker. +The Perform Migration between Arm SoC uses the Arm MCP (Model Context Protocol) server to provide specialized Arm migration capabilities. The Arm MCP server runs via Docker. Install Docker on your local development machine (required for ARM MCP server): @@ -87,7 +86,7 @@ docker --version ``` {{% notice Note %}} -Ensure Docker is running before using the Arm SoC Migration Power. The Power will automatically pull and run the Arm MCP server container when needed. +Ensure Docker is running before using the Perform Migration between Arm SoC. The Power will automatically pull and run the Arm MCP server container when needed. {{% /notice %}} ### Launch AWS Graviton3 instance (source platform) @@ -170,7 +169,7 @@ sudo dnf install -y gcc make wget tar In this section: -- You installed Kiro IDE and the Arm SoC Migration Power +- You installed Kiro IDE and the Perform Migration between Arm SoC Power - You provisioned an AWS Graviton3 instance as your source platform - You installed the build tools needed for the migration example diff --git a/content/learning-paths/servers-and-cloud-computing/arm-soc-migration-learning-path/validation.md b/content/learning-paths/servers-and-cloud-computing/arm-soc-migration-learning-path/validation.md index caec43b9bd..cecfcc2faa 100644 --- a/content/learning-paths/servers-and-cloud-computing/arm-soc-migration-learning-path/validation.md +++ b/content/learning-paths/servers-and-cloud-computing/arm-soc-migration-learning-path/validation.md @@ -10,7 +10,7 @@ layout: learningpathall Migration is not complete until the application is validated on both source and target platforms under realistic conditions. -In this section, you will use the Arm SoC Migration Power's testing recommendations to: +In this section, you will use the Perform Migration between Arm SoC Power's testing recommendations to: - Verify functional correctness - Confirm platform compatibility @@ -114,6 +114,6 @@ The Power will analyze platform-specific characteristics. Example analysis: ## What you've accomplished -You've completed the full migration workflow: you validated the source platform build, cross-compiled for the target, ran platform-specific tests, and compared performance between Graviton3 and Raspberry Pi 5. The Arm SoC Migration Power guided each step with architecture-aware recommendations rather than generic advice. +You've completed the full migration workflow: you validated the source platform build, cross-compiled for the target, ran platform-specific tests, and compared performance between Graviton3 and Raspberry Pi 5. The Perform Migration between Arm SoC Power guided each step with architecture-aware recommendations rather than generic advice. The Discovery β†’ Analysis β†’ Planning β†’ Implementation β†’ Validation workflow you followed here applies to any Arm SoC migration, whether cloud-to-edge, edge-to-edge, or between any pair of Arm-based platforms. The HAL pattern preserves your application's business logic across different Arm SoCs so you can adapt the same codebase without starting from scratch. diff --git a/content/learning-paths/servers-and-cloud-computing/multi-accuracy-libamath/_index.md b/content/learning-paths/servers-and-cloud-computing/multi-accuracy-libamath/_index.md index 9d67272dfa..4df58d1763 100644 --- a/content/learning-paths/servers-and-cloud-computing/multi-accuracy-libamath/_index.md +++ b/content/learning-paths/servers-and-cloud-computing/multi-accuracy-libamath/_index.md @@ -1,5 +1,5 @@ --- -title: Select accuracy modes in Libamath (Arm Performance Libraries) +title: Control floating-point accuracy modes in Arm Performance Libraries minutes_to_complete: 20 author: Joana Cruz @@ -7,7 +7,7 @@ author: Joana Cruz who_is_this_for: This is an introductory topic for developers who want to use the different accuracy modes for vectorized math functions in Libamath, a component of Arm Performance Libraries. learning_objectives: - - Understand how accuracy is defined in Libamath + - Describe how accuracy is defined and measured in Libamath - Select an appropriate accuracy mode for your application - Use Libamath with different vector accuracy modes in practice @@ -20,21 +20,29 @@ subjects: Performance and Architecture armips: - Neoverse tools_software_languages: -- Arm Performance Libraries -- GCC -- Libamath + - Arm Performance Libraries + - GCC + - Libamath operatingsystems: - Linux further_reading: - resource: - title: ArmPL Libamath Documentation + title: Arm Performance Libraries math functions documentation link: https://developer.arm.com/documentation/101004/2410/General-information/Arm-Performance-Libraries-math-functions type: documentation - resource: - title: ArmPL Installation Guide + title: Arm Performance Libraries installation guide link: /install-guides/armpl/ type: website + - resource: + title: What Every Computer Scientist Should Know About Floating-Point Arithmetic + link: https://docs.oracle.com/cd/E19957-01/806-3568/ncg_goldberg.html + type: documentation + - resource: + title: Arm Optimized Routines + link: https://github.com/ARM-software/optimized-routines + type: website diff --git a/content/learning-paths/servers-and-cloud-computing/reproducible-libamath/_index.md b/content/learning-paths/servers-and-cloud-computing/reproducible-libamath/_index.md index 574aba45f0..a6eb3bd9a5 100644 --- a/content/learning-paths/servers-and-cloud-computing/reproducible-libamath/_index.md +++ b/content/learning-paths/servers-and-cloud-computing/reproducible-libamath/_index.md @@ -1,9 +1,5 @@ --- -title: Use reproducible functions in Libamath (Arm Performance Libraries) - -draft: true -cascade: - draft: true +title: Enable reproducible math functions across vector extensions with Arm Performance Libraries minutes_to_complete: 10 author: Joana Cruz @@ -13,18 +9,18 @@ who_is_this_for: This is an introductory topic for developers who want to produc learning_objectives: - Explain what numerical reproducibility means in numerical software - Describe generic applications of numerical reproducibility in the industry - - Understand how reproducibility is defined in Libamath + - Describe how reproducibility is defined and implemented in Libamath - Enable and use reproducible Libamath functions in real applications prerequisites: - - An Arm computer running Linux with [Arm Performance Libraries](/install-guides/armpl/) version 26.01 or newer installed. + - An Arm computer running Linux with [Arm Performance Libraries](/install-guides/armpl/) version 26.01 or newer installed + - A C compiler such as [GCC](/install-guides/gcc/native/) or Clang installed ### Tags skilllevels: Introductory subjects: Performance and Architecture armips: - Neoverse - - SVE tools_software_languages: - Arm Performance Libraries - GCC @@ -42,6 +38,10 @@ further_reading: title: ArmPL Installation Guide on Linux link: /install-guides/armpl/#linux type: website + - resource: + title: Use multi-accuracy math functions in Libamath + link: /learning-paths/servers-and-cloud-computing/multi-accuracy-libamath/ + type: website ### FIXED, DO NOT MODIFY diff --git a/content/learning-paths/servers-and-cloud-computing/reproducible-libamath/applications.md b/content/learning-paths/servers-and-cloud-computing/reproducible-libamath/applications.md index 741f10aa06..1f1d1b2afe 100644 --- a/content/learning-paths/servers-and-cloud-computing/reproducible-libamath/applications.md +++ b/content/learning-paths/servers-and-cloud-computing/reproducible-libamath/applications.md @@ -1,56 +1,40 @@ --- -title: Applications of Reproducibility +title: Explore where reproducibility is critical weight: 3 ### FIXED, DO NOT MODIFY layout: learningpathall --- -# Applications of Reproducibility +## Key domains requiring reproducible math -Reproducibility is not required for every application, but it is critical in several important domains. Here are some examples. +Reproducibility is not required for every application, but it is critical in several important domains. -### Auto-vectorisation -Modern compilers automatically vectorise scalar loops when possible. This means that, depending on the compiler decisions, the same source code may be executed as a scalar loop, as a Neon vectorized loop or as a SVE vectorised loop. -Additionally, vectorised loops often include scalar tail handling for leftover elements that do not fill an entire vector. +## Auto-vectorization -Reproducibility across math routines garantees that: +Modern compilers automatically vectorize scalar loops when possible. Depending on compiler decisions, the same source code can be executed as a scalar loop, a NEON vectorized loop, or an SVE vectorized loop. Vectorized loops also often include scalar tail handling for leftover elements that don't fill an entire vector. -* Vectorized loops (Neon or SVE) match regardless of which one is used +Reproducibility across math routines guarantees that vectorized loops (NEON or SVE) match regardless of which path the compiler selects. It also ensures that loops over scalar routines produce the same results as their vectorized counterparts, so changing vector width or enabling/disabling auto-vectorization does not change the final output. -* The result of loops over scalar routines matches the results of vectorised loops (Neon or SVE) +## Distributed computing -* Changing vector width or enabling/disabling auto-vectorisation does not change the final output +In distributed or parallel workloads, computations are often decomposed across multiple machines or execution units. Different nodes can execute scalar, NEON, or SVE code paths, and the decomposition of work can change between runs. Without reproducible math routines, small numerical differences accumulate and lead to divergent final results. +## Embedded and real-time systems -### Distributed Computing +In real-time environments, determinism is essential. Bitwise-identical results simplify validation, and reproducibility ensures consistent behavior across software updates and hardware variants. Debugging and fault analysis also become significantly easier when you can rule out numerical drift. -In distributed or parallel workloads, computations are often decomposed across multiple machines or execution units. +## Gaming and simulation -* Different nodes may execute scalar, Neon, or SVE code paths +Many games and simulations rely on deterministic numerical behavior. Reproducibility enables lockstep simulations across threads or devices and helps prevent desynchronization in multiplayer or replay systems. Deterministic math also simplifies testing and debugging of complex numerical code. -* The decomposition of work can change between runs +Now that you've seen where reproducibility matters in practice, the next section explains how Libamath implements cross-vector-extension reproducibility and how to enable it in your applications. -* Without reproducible math routines, small numerical differences can accumulate and lead to divergent final results +## What you've learned and what's next -### Embedded and real-time systems +You've explored several real-world scenarios where reproducibility is critical: auto-vectorization requiring consistent results across scalar and vector paths, distributed computing needing deterministic numerics across nodes, embedded systems demanding bitwise-identical validation, and gaming requiring lockstep simulations. -In real-time environments, determinism is essential. - -* Bitwise-identical results simplify validation - -* Reproducibility ensures consistent behavior across software updates and hardware variants - -* Debugging and fault analysis become significantly easier - -### Gaming and simulation -Many games and simulations rely on deterministic numerical behavior. - -* Reproducibility enables lockstep simulations across threads or devices - -* It helps prevent desynchronization in multiplayer or replay systems - -* Deterministic math simplifies testing and debugging of complex numerical code +Next, you'll learn how to enable reproducible math routines in Libamath and integrate them into your build system. diff --git a/content/learning-paths/servers-and-cloud-computing/reproducible-libamath/examples.md b/content/learning-paths/servers-and-cloud-computing/reproducible-libamath/examples.md index b821546204..38e73a05b2 100644 --- a/content/learning-paths/servers-and-cloud-computing/reproducible-libamath/examples.md +++ b/content/learning-paths/servers-and-cloud-computing/reproducible-libamath/examples.md @@ -1,5 +1,5 @@ --- -title: Examples +title: Verify reproducible results across scalar, NEON, and SVE weight: 5 ### FIXED, DO NOT MODIFY @@ -8,52 +8,66 @@ layout: learningpathall ## Example: Reproducible expf -In this example, you will take a look into reproducibility Libamath usage, considering the case of a simple computation using the exponential function in single precision (`expf`). +This example demonstrates reproducibility in Libamath using the single-precision exponential function, `expf()`. +### Setting up your environment -#### Setting up your environment -In this example we use **`GCC-14`** compiler on a **`Neoverse V1`** machine. We use [ArmPL 26.01 module](/install-guides/armpl/). -You can setup some environment variables to make compilation commands simpler: +This example uses GCC-14 on a Neoverse V1 machine with the [ArmPL 26.01 module](/install-guides/armpl/) installed. -```bash { command_line="root@localhost" } -export LD_LIBRARY_PATH=/lib:$LD_LIBRARY_PATH -export C_INCLUDE_PATH=/include -export LIBRARY_PATH=/lib +If you need to install GCC run: + +```bash +sudo apt install gcc -y ``` -With this setup, you can compile the examples to use the reproducible Libamath library via: -```bash { command_line="root@localhost" } +You can set up some environment variables to make compilation commands simpler: + +```bash { command_line="ubuntu@localhost" } +export CC=gcc +export LD_LIBRARY_PATH=$ARMPL_DIR/lib:$LD_LIBRARY_PATH +export C_INCLUDE_PATH=$ARMPL_DIR/include +export LIBRARY_PATH=$ARMPL_DIR/lib +``` + +The commands below show the general compilation pattern used throughout the examples. + +Don't run them yet because `app.c` is introduced in the sections that follow. + +To compile with the reproducible Libamath library: + +```bash $CC app.c -DAMATH_REPRO=1 -lamath_repro -o app ``` -If in turn you are interested in using the non-reproducible Libamath library, you should compile with: -```bash { command_line="root@localhost" } +To compile with the non-reproducible Libamath library instead: + +```bash $CC app.c -lamath -o app ``` -Note that that this only works if `app.c` only contains functions that are present in both versions of the library (`libamath_repro.a` contains a subset of functions in `libamath.a`). +This only works if `app.c` only contains functions that are present in both versions of the library (`libamath_repro.a` contains a subset of functions in `libamath.a`). -You can run examples via: +To run a compiled example: -```bash { command_line="root@localhost" } +```bash ./app ``` -For `SVE` applications, add `-march=armv8-a+sve` to the compilation command. For example: +For SVE applications, add `-march=armv8-a+sve` to the compilation command: -```bash { command_line="root@localhost" } +```bash $CC app.c -DAMATH_REPRO=1 -lamath_repro -march=armv8-a+sve -o app ``` +### Scalar usage -#### Scalar usage +The starting point is a small application that uses the scalar implementation of the single-precision exponential function `armpl_exp_f32()`. -Our starting point is a small application that uses the scalar implementation of the single precision exponential function `armpl_exp_f32`. -Below you can find the example-code, the output when reproducibility is enabled versus when reproducibility is disabled. +Save the following C code in a file called `app.c` using your preferred text editor. Then compile and run it using the compilation patterns from the previous section, once with reproducibility enabled (`-DAMATH_REPRO=1 -lamath_repro`) and once with reproducibility disabled (`-lamath`), to compare the output for each case. {{< tabpane code=true >}} - {{< tab header="C Application" language="C" output_lines="10">}} + {{< tab header="C code" language="C" output_lines="10">}} #include #include @@ -76,12 +90,12 @@ y = 2.613692045211792 [0x1.4e8d76p+1] {{< /tabpane >}} -#### Neon usage +### NEON usage -Now we build a simple Neon application that invokes the reproducible Neon implementation of the single precision exponential function `armpl_vexpq_f32`. +Next, replace the contents of `app.c` with the following NEON application that invokes the reproducible NEON implementation of the single-precision exponential function `armpl_vexpq_f32()`. Compile and run it again with reproducibility enabled and disabled to compare the results. {{< tabpane code=true >}} - {{< tab header="C" language="C" output_lines="15">}} + {{< tab header="C code" language="C" output_lines="15">}} #include #include #include @@ -118,13 +132,14 @@ y (lane 3) = 2.613692283630371 [0x1.4e8d78p+1] {{< /tabpane >}} -Once we run this example, each lane of `y` will contain the same bit pattern as the scalar result of `armpl_exp_f32(1.0f)` (as you can see in the *Output* tab). +Once you run this example, each lane of `y` contains the same bit pattern as the scalar result of `armpl_exp_f32(0x1.ebe93cp-1f)` (as you can see in the *Output* tab). -#### SVE usage -Finally we build a simple SVE application that invokes the reproducible SVE implementation of the single precision exponential function `armpl_svexp_f32_x`. +### SVE usage + +Finally, replace the contents of `app.c` with the following SVE application that invokes the reproducible SVE implementation of the single-precision exponential function `armpl_svexp_f32_x()`. Compile and run using the SVE compilation command (with `-march=armv8-a+sve`), once with reproducibility enabled and once disabled. {{< tabpane code=true >}} - {{< tab header="C" language="C" output_lines="15">}} + {{< tab header="C code" language="C" output_lines="15">}} #include #include #include @@ -166,17 +181,14 @@ y (lane 7): 2.613692045211792 [0x1.4e8d76p+1] {{< /tab >}} {{< /tabpane >}} -All active lanes of `y` are guaranteed to match the scalar and Neon results exactly. +All active lanes of `y` are guaranteed to match the scalar and NEON results exactly. + +### Scope and limitations -#### Scope and Limitations +In this section you observed that, when reproducibility is enabled (`AMATH_REPRO` enabled), `expf()` produces bitwise-identical results whether it is executed as a scalar, NEON or SVE function. -In this section we observed that, when reproducibility is enabled (`AMATH_REPRO` enabled), `expf` produces bitwise-identical results whether it is executed as a scalar, Neon or SVE function. +This behavior extends to other reproducible math routines in Libamath. Scalar, NEON, and SVE implementations are numerically aligned for all functions listed in `amath_repro.h`. Reproducible symbols are always prefixed by `armpl_` and are not provided with `ZGV` mangling. Reproducibility is available on Linux platforms, and results are independent of vector width or instruction selection. Reproducible routines prioritize determinism over peak performance. -This behaviour extends to other reproducible math routines in reproducible Libamath: +## What you've learned and what's next -* Scalar, Neon, and SVE implementations are numerically aligned -* Only functions listed in `amath_repro.h` are reproducible -* Reproducible symbols are always prefixed by `armpl_`. They are not provided with `ZGV` mangling. -* Reproducibility is provided on Linux platforms -* Results are independent of vector width or instruction selection -* Reproducible routines prioritize determinism over peak performance \ No newline at end of file +In this Learning Path, you learned what numerical reproducibility means in floating-point software and explored real-world applications where it is critical. You then enabled cross-vector-extension reproducibility in Libamath and verified that scalar, NEON, and SVE code paths produce bitwise-identical results for the `expf()` function. You can now apply these techniques to your own applications using Arm Performance Libraries. \ No newline at end of file diff --git a/content/learning-paths/servers-and-cloud-computing/reproducible-libamath/reproducibility.md b/content/learning-paths/servers-and-cloud-computing/reproducible-libamath/reproducibility.md index f937e10830..4af1998e4a 100644 --- a/content/learning-paths/servers-and-cloud-computing/reproducible-libamath/reproducibility.md +++ b/content/learning-paths/servers-and-cloud-computing/reproducible-libamath/reproducibility.md @@ -1,31 +1,30 @@ --- -title: Reproducibility +title: Understand numerical reproducibility in floating-point math weight: 2 ### FIXED, DO NOT MODIFY layout: learningpathall --- -## What is Reproducibility? +## What is reproducibility? -In numerical software, reproducibility (also refered to as determinism) means you get the exact same floating-point bits for the same inputs β€” even if you run a different implementation (scalar vs Neon vs SVE). +In numerical software, reproducibility (also referred to as determinism) means you get the exact same floating-point bits for the same inputs, even if you run a different implementation (scalar vs NEON vs SVE). In pure mathematics, two functions `𝑓(π‘₯)` and `𝑔(π‘₯)` are equivalent if, for all `π‘₯` in their domain `𝑓(π‘₯) = 𝑔(π‘₯)`. +In practice, numerical software replaces continuous mathematical functions over real numbers with discrete approximations using floating-point numbers. Instead of comparing two abstract functions, you compare two implementations. For example, a scalar version and a vectorized version of the same routine. -In practice, numerical software replaces continuous mathematical functions over real numbers with discrete approximations using floating-point numbers. Instead of comparing two abstract functions, we compare two implementations. For example, a scalar version and a vectorized version of the same routine. - -We say that two programs are reproducible if, for the same input values, they produce exactly the same floating-point results, down to the last bit. +Two programs are reproducible if, for the same input values, they produce exactly the same floating-point results, down to the last bit. {{% notice Accuracy vs Reproducibility %}} -Note that this requirement is **independent** of the accuracy requirement: two results can both be within an acceptable error bound and still differ in their bit patterns. +This requirement is **independent** of the accuracy requirement: two results can both be within an acceptable error bound and still differ in their bit patterns. However correctly rounded routines (maximum error under 0.5ULP) are reproducible by essence, since for a given input, rounding mode and precision, the output is the floating-point number closest to the exact mathematical result. {{% /notice %}} -## Levels of Reproducibility +## Levels of reproducibility Reproducibility can be defined at different levels, depending on how similar or different the execution environments are: @@ -34,6 +33,14 @@ Reproducibility can be defined at different levels, depending on how similar or Reproducibility across different processor architectures, such as x86 and AArch64. * **Cross-vector-extension reproducibility** - Reproducibility across different vector execution paths on the same architecture, such as scalar, Neon, and SVE on AArch64. + Reproducibility across different vector execution paths on the same architecture, such as scalar, NEON, and SVE on AArch64. + +This Learning Path focuses on cross-vector-extension reproducibility (scalar, NEON, SVE on AArch64). + +Now that you understand what numerical reproducibility means and the different levels it can operate at, the next section covers real-world applications where this property is critical. + +## What you've learned and what's next + +You now understand the core concept of numerical reproducibility in floating-point computation and the different levels at which it can be achieved. You've learned why reproducibility matters and how it relates to portability and determinism. -In this Learning Path, we focus on cross-vector-extension reproducibility (scalar, Neon, SVE on AArch64).” \ No newline at end of file +Next, you'll explore specific real-world applications where reproducibility is essential for mission-critical systems and regulatory compliance. \ No newline at end of file diff --git a/content/learning-paths/servers-and-cloud-computing/reproducible-libamath/reproducibility_libamath.md b/content/learning-paths/servers-and-cloud-computing/reproducible-libamath/reproducibility_libamath.md index 55ee75ca42..2f1f8b309d 100644 --- a/content/learning-paths/servers-and-cloud-computing/reproducible-libamath/reproducibility_libamath.md +++ b/content/learning-paths/servers-and-cloud-computing/reproducible-libamath/reproducibility_libamath.md @@ -1,5 +1,5 @@ --- -title: Reproducibility in Libamath +title: Enable reproducibility in Libamath weight: 4 ### FIXED, DO NOT MODIFY @@ -8,47 +8,42 @@ layout: learningpathall ## Cross-vector-extension reproducibility -On Linux platforms, Libamath supports bitwise-reproducible results across scalar, Neon (AdvSIMD), and SVE implementations for a subset of math functions. +On Linux platforms, Libamath supports bitwise-reproducible results across scalar, NEON (AdvSIMD), and SVE implementations for a subset of math functions. -When reproducibility is enabled, the same input values produce identical floating-point results, regardless of whether a supported function is executed using the scalar, Neon, or SVE code path. This keeps your results deterministic even if your app takes different vector paths. +When reproducibility is enabled, the same input values produce identical floating-point results, regardless of whether a supported function is executed using the scalar, NEON, or SVE code path. This keeps your results deterministic even if your app takes different vector paths. Reproducible Libamath routines operate in the default accuracy mode, guaranteeing results within 3.5 ULP of the correctly rounded value. -Note that reproducible routines prioritize determinism over peak performance. +Reproducible routines prioritize determinism over peak performance. ## Reproducible symbols -When reproducibility is enabled: +When reproducibility is enabled, reproducible functions use the same public function names as their non-reproducible counterparts. The linker resolves calls to the reproducible implementations when you build with `-DAMATH_REPRO=1` and link `-lamath_repro`, and the scalar, NEON, and SVE variants of a function all produce bitwise-identical results. -* Reproducible functions use the same public function names as their non-reproducible counterparts +Unlike the symbols in `amath.h` (which don't guarantee reproducibility), reproducible symbols in `amath_repro` are not provided in `ZGV` mangling. Only the `armpl_` notation is used. -* The linker resolves calls to the reproducible implementations when you build with `-DAMATH_REPRO=1` and link `-lamath_repro` - -* Scalar, Neon, and SVE variants of a function all produce bitwise-identical results - -* Unlike the symbols you find in `amath.h` (which don't guarantee reproducibility), reproducible symbols in `amath_repro` are not provided in `ZGV` mangling (only the `armpl_` notation is used). - -The full list of functions that support reproducible behavior is provided in the header file `amath_repro.h` +The full list of functions that support reproducible behavior is provided in the header file `amath_repro.h`. ## How to use reproducible Libamath -To enable reproducibility in a C or C++ application: - -1. Include the Libamath header +To enable reproducibility in a C or C++ application, include the Libamath header in your source file: ```C #include ``` -2. Compile with reproducibility enabled +Then compile and link with reproducibility enabled: ```bash --DAMATH_REPRO=1 +gcc app.c -DAMATH_REPRO=1 -lamath_repro -o app ``` -3. Link against the reproducible Libamath library -```bash --lamath_repro -``` +The `-DAMATH_REPRO=1` flag enables reproducibility at compile time, and `-lamath_repro` links against the reproducible Libamath library. When you follow these steps, calls to supported functions resolve to the reproducible scalar, NEON, or SVE implementations. + +With reproducibility configured, the next section walks through hands-on examples using `expf` across scalar, NEON, and SVE code paths. + +## What you've learned and what's next + +You've learned how to enable reproducible math routines in Libamath through compile-time configuration and library linking. You can now compile code with the reproducible library variant and understand the trade-offs between reproducibility and peak performance. -When you follow these steps, calls to supported functions resolve to the reproducible scalar, Neon, or SVE implementations. \ No newline at end of file +Next, you'll verify reproducible behavior through hands-on examples that compare scalar, NEON, and SVE implementations of the exponential function. \ No newline at end of file diff --git a/themes/arm-design-system-hugo-theme/layouts/partials/footer/script-includes.html b/themes/arm-design-system-hugo-theme/layouts/partials/footer/script-includes.html index 333ea4be6a..4715342856 100644 --- a/themes/arm-design-system-hugo-theme/layouts/partials/footer/script-includes.html +++ b/themes/arm-design-system-hugo-theme/layouts/partials/footer/script-includes.html @@ -17,12 +17,30 @@ {{ if not (in .Site.BaseURL "localhost") }} + + + + + {{end}}