3.4 KiB
3.4 KiB
Handoff: OpenCode Mobile App
Status: CUA E2E Tests Working, v0.2.0 Release Triggered
What Was Done
1. CUA (Computer-Use Agent) E2E Test Infrastructure
- Built
scripts/android-cua-smoke.py— a vision-LLM-powered Android E2E test that drives the app via ADB screenshots + AI actions - Two scenarios pass reliably:
send_message(7 steps) andmulti_turn(13 steps) - Added custom
{"type": "send"}action that auto-locates the send button via uiautomator XML (solves coordinate accuracy issues with vision models) - Uses Azure AI Services gpt-5.4 for vision inference
2. CI/CD
.github/workflows/cua-smoke.yml— GitHub Actions workflow for emulator + CUA test.github/workflows/build.yml— Builds APK, creates GitHub Release on tag push.github/workflows/publish-play-store.yml— Fastlane Play Store publish- GitHub secrets set:
AZURE_OPENAI_API_KEY,AZURE_OPENAI_ENDPOINT - Tagged
v0.2.0— CI build + release workflow triggered
3. App Features (prior sessions)
- SSE reconnect with exponential backoff
- Notification support
- Message deduplication
- Biometric auth gate (disabled on emulator)
- Server connection management UI
Key Technical Details
Azure AI Services (LLM API)
- Correct endpoint:
https://info-mjnxtt51-eastus2.cognitiveservices.azure.com - DO NOT USE:
vibe-dev-ai.cognitiveservices.azure.com(has no deployments, only model catalog) - Env file:
~/.env.d/azure-openai.env - Available models: gpt-5.4, gpt-5.2, gpt-5.1, gpt-4.1, grok-4, deepseek-r1, kimi-k2.5
- API quirk: Use
max_completion_tokens(NOTmax_tokens) for gpt-5.x models
Android Emulator
- SDK at
/tmp/android-sdk/(may need reinstall if /tmp is cleared) - AVD name:
test, API 34 x86_64, KVM required - Start:
emulator -avd test -no-window -no-audio -no-boot-anim -gpu swiftshader_indirect -no-snapshot - Package:
ai.opencode.mobile
CUA Script Learnings
- Vision models (gpt-5.4) consistently mis-estimate Y coordinates by ~80px for bottom-of-screen elements
- Solution:
{"type": "send"}action uses uiautomator XML to find the rightmost clickable element in the bottom bar - Enter key inserts newline in this app (doesn't send) — model must dismiss keyboard + tap send button
- Screen resolution (1080x2400) included in prompt context helps coordinate accuracy
- Screenshot retry (3 attempts, 30s timeout) needed for emulator under load
Server Connection
- OpenCode server:
100.108.64.76:4096(Tailscale, hostnameopenclaw-dev-1) - Must dismiss keyboard before tapping Connect button (keyboard obscures it)
- Model:
gpt-5.3-chat-latest(via GitHub Copilot provider)
What's Next
- Verify v0.2.0 release — check CI at https://github.com/dzianisv/opencode-mobile/actions
- Run CUA in CI — the
cua-smoke.ymlworkflow needs a running opencode server to connect to (currently targets Tailscale IP which won't be reachable from GitHub runners). Options:- Mock server in CI
- Use a public opencode server endpoint
- Run as self-hosted runner on the dev VM
- Add accessibility labels — add
content-desc="Send"to the send button in the app for better CUA reliability - More scenarios — settings toggle, reconnect after server restart, slash commands
Repo & Auth
- Repo:
dzianisv/opencode-mobile - GitHub auth:
source ~/.env.d/github-dzianisv.env - Upstream issue:
anomalyco/opencode#10288