Dennis V
a305936f2b
feat(cua): add --e2e and --query modes with structured evaluation
--e2e mode: full end-to-end coding task scenario
- connect → long-press FAB to create session in custom project dir
- select AI model via model picker (hint substring match)
- submit coding task → DETERMINISTIC API poll for session idle
- DETERMINISTIC API message scan for target filename
- DETERMINISTIC ADB uiautomator check for filename in UI
- LLM screenshot + visual evaluation summary
--query mode: natural-language test description → structured test run
- LLM planner converts the query into JSON phases + deterministic checks
- Executes each phase via the CUA loop (with critical/informational split)
- Runs deterministic checks: ui_text | session_idle | file_created
- LLM evaluator produces scored JSON report: overall/score/phases/recommendations
New helpers:
- wait_for_session_idle(): polls GET /session until status==idle (no LLM)
- check_session_file_created(): scans session messages API for filename
- _api_base(): translates emulator host route for host-side API calls
- run_scenario_hello_world_e2e(): 8-phase hardcoded e2e scenario
- run_query_test(): planner → execute → evaluator pipeline
Also adds hello_world_e2e to --scenarios catalog for named invocation.
Usage:
# Hardcoded e2e:
python scripts/android-cua-smoke.py --e2e --opencode-url http://100.108.64.76:4096
# Natural-language query:
python scripts/android-cua-smoke.py --query \
'Open android app. Connect to server. Open ~/workspace/opencode-mobile. \
Choose deepseek model. Ask to write hello_world.py. Validate it was created.'
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-06-24 03:53:01 +00:00
..
2026-05-19 07:47:04 +00:00
2026-06-24 03:53:01 +00:00
2026-06-23 20:18:15 +00:00