Skip to content

Story 06 — Native runner tests in CI + LLM-behavior evals #387

Description

@Lykhoyda

The riskiest layer has zero automated execution: Swift suites + Kotlin KeyboardGuardTest exist in-tree but CI only compiles them (CodeQL). Phase A: run gradlew testDebugUnitTest (ubuntu) + xcodebuild test -only-testing:RnFastRunnerTests (macOS), path-filtered. Phase B: nightly device smoke — prebuilt artifacts (Story 01), pinned sim+AVD, tiny fixture app, golden command set driven through the real bridge over MCP stdio; 2-consecutive-red alerting. Phase C: mcp-server-tester-style LLM-behavior evals (Maestro: 'test that LLMs can call it correctly… happens less frequently than is expected') — baseline gates Stories 08/12.

Spec: https://github.com/Lykhoyda/rn-dev-agent/blob/limestone-malpais/docs/stories/06-native-runner-ci-and-evals.md (PR #381)
Impact: coverage where the hardest bugs live · Effort: M · Depends: #382 (Story 01, for Phase B)

Metadata

Metadata

Assignees

No one assigned

    Labels

    effort:mEffort: ~1-3 daysenhancementNew feature or requestkano:performanceKano: more/better linearly increases satisfactionpriority:nextUp next after current 'now' items

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions