Files
opencode-mobile/scripts
Dennis V a305936f2b feat(cua): add --e2e and --query modes with structured evaluation
--e2e mode: full end-to-end coding task scenario
  - connect → long-press FAB to create session in custom project dir
  - select AI model via model picker (hint substring match)
  - submit coding task → DETERMINISTIC API poll for session idle
  - DETERMINISTIC API message scan for target filename
  - DETERMINISTIC ADB uiautomator check for filename in UI
  - LLM screenshot + visual evaluation summary

--query mode: natural-language test description → structured test run
  - LLM planner converts the query into JSON phases + deterministic checks
  - Executes each phase via the CUA loop (with critical/informational split)
  - Runs deterministic checks: ui_text | session_idle | file_created
  - LLM evaluator produces scored JSON report: overall/score/phases/recommendations

New helpers:
  - wait_for_session_idle(): polls GET /session until status==idle (no LLM)
  - check_session_file_created(): scans session messages API for filename
  - _api_base(): translates emulator host route for host-side API calls
  - run_scenario_hello_world_e2e(): 8-phase hardcoded e2e scenario
  - run_query_test(): planner → execute → evaluator pipeline

Also adds hello_world_e2e to --scenarios catalog for named invocation.

Usage:
  # Hardcoded e2e:
  python scripts/android-cua-smoke.py --e2e --opencode-url http://100.108.64.76:4096

  # Natural-language query:
  python scripts/android-cua-smoke.py --query \
    'Open android app. Connect to server. Open ~/workspace/opencode-mobile. \
     Choose deepseek model. Ask to write hello_world.py. Validate it was created.'

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-06-24 03:53:01 +00:00
..