--e2e mode: full end-to-end coding task scenario - connect → long-press FAB to create session in custom project dir - select AI model via model picker (hint substring match) - submit coding task → DETERMINISTIC API poll for session idle - DETERMINISTIC API message scan for target filename - DETERMINISTIC ADB uiautomator check for filename in UI - LLM screenshot + visual evaluation summary --query mode: natural-language test description → structured test run - LLM planner converts the query into JSON phases + deterministic checks - Executes each phase via the CUA loop (with critical/informational split) - Runs deterministic checks: ui_text | session_idle | file_created - LLM evaluator produces scored JSON report: overall/score/phases/recommendations New helpers: - wait_for_session_idle(): polls GET /session until status==idle (no LLM) - check_session_file_created(): scans session messages API for filename - _api_base(): translates emulator host route for host-side API calls - run_scenario_hello_world_e2e(): 8-phase hardcoded e2e scenario - run_query_test(): planner → execute → evaluator pipeline Also adds hello_world_e2e to --scenarios catalog for named invocation. Usage: # Hardcoded e2e: python scripts/android-cua-smoke.py --e2e --opencode-url http://100.108.64.76:4096 # Natural-language query: python scripts/android-cua-smoke.py --query \ 'Open android app. Connect to server. Open ~/workspace/opencode-mobile. \ Choose deepseek model. Ask to write hello_world.py. Validate it was created.' Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
89 KiB
Executable File
89 KiB
Executable File