* fix(ci): add emulator to PATH and increase CUA smoke timeout to 60min
- Add 'Add emulator to PATH' step after setup-android so emulator binary
is found (was: command not found, causing adb wait-for-device to hang
until the 30min job timeout)
- Increase timeout-minutes from 30 to 60 to accommodate full build
- Add npm cache and Gradle cache (same as build.yml) to speed up rebuild
* fix: address code review findings for cua-emulator-path
- Fix stale Gradle cache causing build failure after package rename
(ai.opencode.mobile → cc.agentlabs.opencode): add purge step matching
the one already present in publish-play-store.yml
- Fix adb launch command using old package name ai.opencode.mobile;
updated to cc.agentlabs.opencode/.MainActivity
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JKGMRpgihA4io2frodqLjt
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(ci): bump opencode Azure apiVersion to support /responses endpoint (#22)
The CUA smoke probe was returning MODEL_CAPABLE=false because opencode's
@ai-sdk/azure provider got 'API version not supported' from Azure on
/openai/v1/responses with apiVersion=2024-08-01-preview. Split into two
envs: keep the CUA driver on 2024-08-01-preview (chat-completions only)
and bump the opencode-side provider config to 2025-04-01-preview, which
supports the new responses API.
Effect: send_message/multi_turn scenarios get included again in CUA smoke
when the probe succeeds.
* ci(cua): bound curl timeouts + diag dump on server-start hang (#22)
Step 11 'Start opencode server' has hung past the 45-min job timeout in
two consecutive runs (27195348071, 27198238465). Local boot of
opencode-ai 1.16.2 with the same heredoc config is healthy in 3s, so
something is wrong specifically on the GH-hosted runner — likely curl
post-loop waiting indefinitely on an unresponsive server.
Adds:
- set -x for command tracing
- --connect-timeout 2 -m 5 on every curl so hangs cannot exceed 5s
- HEALTHY flag + explicit exit 1 (drops the unbounded post-loop curl)
- Periodic dump every 10s: server log tail, ss listening sockets,
process liveness — so we can see WHY the server isn't replying
Pure diagnostics; no behaviour change for the green path.
* fix(ci): use api-version=preview for /openai/v1/responses (#22)
Reproduced the probe failure locally against the same Azure resource:
all date-based api-versions (2024-08-01-preview, 2024-12-01-preview,
2025-01-01-preview, 2025-03-01-preview, 2025-04-01-preview) return:
{"error":{"code":"BadRequest","message":"API version not supported"}}
Only api-version=preview and api-version=v1 succeed (200). This is the
new Azure OpenAI v1 responses-API style; date strings are reserved for
the legacy /openai/deployments/{model}/chat/completions endpoint.
@ai-sdk/azure 3.x already defaults apiVersion to "preview" (per the
type definition: "Custom api version to use. Defaults to `preview`."),
so this aligns the workflow with the SDK default. Probe should now
return MODEL_CAPABLE=true and the send_message scenario will run.
* test(cua): extend send_message and multi_turn waits to 30s
Assistant bubbles can take 15+ seconds to appear after send. Previous
5-second wait was too short and caused false failures even when API
calls succeeded. Re-check screenshots periodically up to 30s total.
* fix(cua): screen-relative send button threshold for #22
The send action's auto-locate filtered for y1 > 2200 and fell back to
hardcoded (996, 2358) — both assume a 1080x2400 panel. The CI emulator
(API 30 google_apis pixel profile) is 1080x1920, so:
- the bottom_buttons filter never matched any clickable element
- the fallback tap landed off-screen
→ 'ping' message never sent, scenario timed out with no bubbles.
Switch to a screen-relative threshold (bottom 25%) and a fallback that
uses get_screen_size() to land in the bottom-right corner regardless of
device resolution. This was masked until now because send_message was
gated by MODEL_CAPABLE=false in earlier CI runs.
Refs: #22
---------
Co-authored-by: dzianisv <dzianis.varabyou@gmail.com>
Run 27139243275 FAILED with "/usr/bin/sh: Syntax error: end of file unexpected
(expecting fi)" — android-emulator-runner runs the script under dash, which
mangled the multi-line if/then/else/fi I added, so the scenarios never ran (a
real bug I introduced, not environmental). Fix:
- Move scenario-set selection into the probe step (runs under bash) and export
SCENARIOS as a step output; the emulator script is now single-line.
- The earlier probe returned a FALSE NEGATIVE: it POSTed to /session/{id}/message
with no model, but that endpoint REQUIRES model {providerID,modelID} (per the
opencode SDK the app uses). Probe now sends azure/gpt-5.4 exactly like the app,
uses -s + %{http_code} (not -sf) so the body/status are visible, and only flags
capable on HTTP 200 + an assistant text part.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The CI opencode server had NO LLM provider configured — the server log only
showed "listening", never a model. opencode-ai (released npm pkg) does not read
AZURE_OPENAI_* for its own LLM; it needs an explicit provider in opencode.json +
a default `model`. So send_message/multi_turn could never pass and the gate was
stuck on --only-connect-scenario (UI journey minus the model reply).
- Wire opencode to the same Azure resource the CUA driver uses via a generated
~/.config/opencode/opencode.json (@ai-sdk/azure provider, resourceName derived
from the endpoint secret at runtime, apiKey from env, default model azure/gpt-5.4).
- Add a deterministic REST probe step: create a session + send a prompt and check
for an assistant reply BEFORE the ~30min emulator run, exporting MODEL_CAPABLE.
- Add --scenarios to android-cua-smoke.py to run an explicit named set.
- Emulator step now runs connect_and_verify_sessions + send_message +
verify_session_list when MODEL_CAPABLE=true; falls back to the UI-only journey
(connect + verify_session_list) otherwise, logging the environmental reason.
- Raise --max-steps to 40 so multiple scenarios fit.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: repoint OpenCode links to agentlabs.cc/opencode
agentlabs.cc/opencode and /opencode/privacy are now live (200). Repoint
README, distribution listings (Play/App Store/F-Droid/IzzyOnDroid/iOS),
docs, and in-app privacy links (settings + telemetry consent) from
www.vibebrowser.app/opencode to the canonical agentlabs.cc hub.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(ci): run local opencode server for CUA smoke true-E2E (#15)
GitHub-hosted runners can't reach the Tailscale dev server
(100.108.64.76:4096), so the CUA smoke always failed at session creation.
- Install opencode-ai and run `opencode serve` on the runner host; the
Android emulator reaches it via 10.0.2.2. OPENCODE_URL now points there.
- Healthcheck /global/health before launching the app; dump server log on
failure for diagnosis.
- Add --only-connect-scenario to the smoke script and run just the
connect-and-verify-sessions path in CI: deterministic, needs no model
backend. The scenario now creates a session if the list is empty, so a
fresh server still yields a non-empty list.
This makes the smoke a true E2E and also exercises the #10 sessions-list
rendering path against a real server.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(ci): emulator smoke script is dash, not bash — drop brace-group healthcheck
android-emulator-runner runs the script: block under /usr/bin/sh (dash). The
multi-line `|| { ...; }` healthcheck was a dash syntax error (end of file
unexpected), failing the step before the smoke ran. Replace with a non-fatal
one-line re-check; the server was already health-gated in the prior step.
* docs(tasks): record smoke CI round 1 failure + dash fix
---------
Co-authored-by: engineer <engineer@opencode.ai>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Rename left both old and new package am-start lines; old package is no
longer installed and pollutes the smoke launch. Launch only
cc.agentlabs.opencode/.MainActivity.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- Switch from manual emulator management to the proven emulator-runner action
- Use API 30 (boots faster than 34 with software rendering)
- Build APK before starting emulator to minimize emulator uptime
- All emulator-dependent steps run inside the action's script block
- Move env vars to job level for cleaner structure
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- Enable KVM for hardware acceleration (required for x86_64 emulator)
- Use nohup for emulator process to prevent terminal association issues
- Add avdmanager list to verify AVD creation
- Include platform-tools in sdkmanager install
- Increase boot timeout to 180s
- Upload emulator.log as artifact for debugging
- Reduce job timeout to 45min (was 60)
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- Add v* tag trigger to cua-smoke.yml so releases are E2E tested
- Add 'When to run CUA test' section to AGENTS.md documenting mandatory testing
Closes#13 (partial)
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* fix(sessions): use active connection client directly, remove roots filter
Root cause A: loadSessions was calling clientForDirectory(serverHome) which
scoped the session list to /home/azureuser — a different project than the
server's active CWD. Sessions in the current project (e.g. opencode-mobile)
were never returned.
Root cause B: roots:true filtered out sessions that have a parentID (sub-task /
AUTO-REVIEW sessions), hiding valid sessions from the list.
Fix: use connState.client directly (the connection's active directory) and drop
the roots filter so all sessions for that project are visible.
Also adds a verify_session_list CUA smoke scenario that navigates back to the
sessions tab after creating a session and asserts the list is non-empty —
covering the regression path that was previously untested.
* fix(sessions): fetch serverHome in addConnection so loadSessions shows correct sessions
Root cause: addConnection() built the HTTP client but never fetched serverHome
(only loadConnections and setActiveConnection did). When the user adds a new
connection (fresh install / first sign-in), serverHome = null, so loadSessions
fell through to connState.client (the server's CWD). On this dev server the CWD
is the deploy directory — 11 old May-19 sessions that are not the user's recent
work sessions.
Fix: addConnection now fetches currentProject + serverHome via the same
Promise.all as setActiveConnection, before calling set(). This ensures
loadSessions immediately uses clientForDirectory(serverHome) → the global
project → the user's actual recent parent sessions.
Also adds --opencode-url flag to the CUA smoke script, which appends a
connect_and_verify_sessions scenario that reproduces the regression:
python scripts/android-cua-smoke.py --opencode-url http://100.108.64.76:4096
* fix(sessions): recover home scope after fresh connect
Resolve stale deploy-only session list by recovering server home during first load and keeping regression coverage in default Android CUA smoke and CI.
* chore(release): bump version to 0.4.0
- Record each scenario as MP4 via ADB screenrecord in a background thread
- Pull video to /tmp/cua_<scenario>.mp4 after test completes (always)
- Upload to ArchiveBox if ARCHIVEBOX_URL + ARCHIVEBOX_API_KEY env vars set;
gracefully skips when not configured (CI default)
- Upload all artifacts (PNG screenshots + MP4 videos) always, not only on failure
- Pass ARCHIVEBOX_URL/ARCHIVEBOX_API_KEY secrets to workflow (optional)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>