The riskiest layer has zero automated execution: Swift suites + Kotlin KeyboardGuardTest exist in-tree but CI only compiles them (CodeQL). Phase A: run gradlew testDebugUnitTest (ubuntu) + xcodebuild test -only-testing:RnFastRunnerTests (macOS), path-filtered. Phase B: nightly device smoke — prebuilt artifacts (Story 01), pinned sim+AVD, tiny fixture app, golden command set driven through the real bridge over MCP stdio; 2-consecutive-red alerting. Phase C: mcp-server-tester-style LLM-behavior evals (Maestro: 'test that LLMs can call it correctly… happens less frequently than is expected') — baseline gates Stories 08/12.
Spec: https://github.com/Lykhoyda/rn-dev-agent/blob/limestone-malpais/docs/stories/06-native-runner-ci-and-evals.md (PR #381)
Impact: coverage where the hardest bugs live · Effort: M · Depends: #382 (Story 01, for Phase B)
The riskiest layer has zero automated execution: Swift suites + Kotlin KeyboardGuardTest exist in-tree but CI only compiles them (CodeQL). Phase A: run gradlew testDebugUnitTest (ubuntu) + xcodebuild test -only-testing:RnFastRunnerTests (macOS), path-filtered. Phase B: nightly device smoke — prebuilt artifacts (Story 01), pinned sim+AVD, tiny fixture app, golden command set driven through the real bridge over MCP stdio; 2-consecutive-red alerting. Phase C: mcp-server-tester-style LLM-behavior evals (Maestro: 'test that LLMs can call it correctly… happens less frequently than is expected') — baseline gates Stories 08/12.
Spec: https://github.com/Lykhoyda/rn-dev-agent/blob/limestone-malpais/docs/stories/06-native-runner-ci-and-evals.md (PR #381)
Impact: coverage where the hardest bugs live · Effort: M · Depends: #382 (Story 01, for Phase B)