The CUA script runs on the CI runner (host), not inside the emulator.
10.0.2.2 is the emulator's special address for the host — it is only
reachable FROM INSIDE the emulator. Calling it from the runner always
fails, so _precreate_test_session returned None, and the session_list
phase fell back to the weak 'screen visible' assertion instead of the
strong 'pre-created session must appear' check.
Fix: replace 10.0.2.2 → 127.0.0.1 before making the pre-create call.
Localhost URLs (100.x.x.x, custom dev server) pass through unchanged.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
sessions_reload regression phase now runs BEFORE the TypeScript task.
Previously it was gated behind typescript which fails in CI when the model
is unavailable — meaning the actual sessions regression check never ran.
Phase order is now:
connect → session_list (pre-created session required) → new_session
→ sessions_reload (navigate back, list must be non-empty) [CRITICAL]
→ typescript (informational) → verify (informational) → settings
Critical phases: connect, session_list, new_session, sessions_reload.
TypeScript/verify/settings are informational (model availability varies).
CI emulator fixes:
- api-level: 30 → 28 (more stable, boots reliably on ubuntu-latest)
- target: google_apis → default (lighter, no Play Services needed for
sessions regression test, avoids known boot issues with google_apis)
- disable-animations: true (reduces boot overhead)
- emulator-boot-timeout: 600 (explicit, matches action default)
- Switch from --scenarios to --showcase (runs the new structured flow
with _precreate_test_session + sessions_reload phase)
- Clear app state before install (pm clear) for deterministic first-run
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Previously the session_list phase goal said 'The session list may be empty
(no sessions yet) — that is fine' and 'Report done when you can see the
session list screen (even if empty)'. This means an empty sessions list
was treated as a test PASS, so every previous 'fix' was validated against
a test that cannot detect the regression.
Two changes:
1. _precreate_test_session(): calls POST /session via HTTP before the CUA
starts. The session_list phase goal now explicitly names this session and
requires it to be visible — if the app fails to load server sessions,
the phase fails (not passes silently with an empty list).
Falls back gracefully if the server is unreachable at pre-create time.
2. sessions_reload phase (new, critical): after completing the TypeScript
task the test navigates back to the Sessions tab and asserts the list is
non-empty. This catches the other variant of the regression — sessions
vanishing after navigating away from a session and back.
Both phases are now in the critical list, so a failure in either causes the
overall test to report partial/fail instead of success.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
ADB input text with %s escaping didn't trigger React Native onChangeText.
New approach: write text to /sdcard/ file on device, then use shell
command-substitution "$(cat ...)" to pass raw text to input text.
This preserves spaces without %s conversion and avoids quote-escaping
issues that prevented React Native from detecting text changes.
Rewrites android-cua-smoke.py to demonstrate the complete first-run journey
instead of the previous "ping" smoke test. The new structured multi-phase
flow covers: server connection setup, session list, new session creation,
TypeScript hello-world task submission (watching tool calls/file writes),
output verification, and Settings/model-selection screenshot.
Key changes:
- run_onboarding_showcase() orchestrates 6 sequential CUA phases with
per-phase goals, step budgets, and PASS/FAIL phase tracking
- run_cua_step() replaces run_cua() — accepts step_label, action_delay,
saves labeled screenshots (/tmp/cua_<phase>_<step>.png) for debugging
- Global --speed-multiplier flag scales all _sleep() calls (0.5 = 2x faster)
- Showcase is now the default mode; legacy --goal / --scenarios flags retained
for backwards compat and CI regression scenarios
- Tighter action_delay (0.7s) and trimmed history window (14 turns) vs
previous 1.0s / 12 turns
- Phase banner log lines ("STEP N: ...") narrate the video in real time
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01No3k1AEioE4PNUZg12TxQo
* fix(ci): bump opencode Azure apiVersion to support /responses endpoint (#22)
The CUA smoke probe was returning MODEL_CAPABLE=false because opencode's
@ai-sdk/azure provider got 'API version not supported' from Azure on
/openai/v1/responses with apiVersion=2024-08-01-preview. Split into two
envs: keep the CUA driver on 2024-08-01-preview (chat-completions only)
and bump the opencode-side provider config to 2025-04-01-preview, which
supports the new responses API.
Effect: send_message/multi_turn scenarios get included again in CUA smoke
when the probe succeeds.
* ci(cua): bound curl timeouts + diag dump on server-start hang (#22)
Step 11 'Start opencode server' has hung past the 45-min job timeout in
two consecutive runs (27195348071, 27198238465). Local boot of
opencode-ai 1.16.2 with the same heredoc config is healthy in 3s, so
something is wrong specifically on the GH-hosted runner — likely curl
post-loop waiting indefinitely on an unresponsive server.
Adds:
- set -x for command tracing
- --connect-timeout 2 -m 5 on every curl so hangs cannot exceed 5s
- HEALTHY flag + explicit exit 1 (drops the unbounded post-loop curl)
- Periodic dump every 10s: server log tail, ss listening sockets,
process liveness — so we can see WHY the server isn't replying
Pure diagnostics; no behaviour change for the green path.
* fix(ci): use api-version=preview for /openai/v1/responses (#22)
Reproduced the probe failure locally against the same Azure resource:
all date-based api-versions (2024-08-01-preview, 2024-12-01-preview,
2025-01-01-preview, 2025-03-01-preview, 2025-04-01-preview) return:
{"error":{"code":"BadRequest","message":"API version not supported"}}
Only api-version=preview and api-version=v1 succeed (200). This is the
new Azure OpenAI v1 responses-API style; date strings are reserved for
the legacy /openai/deployments/{model}/chat/completions endpoint.
@ai-sdk/azure 3.x already defaults apiVersion to "preview" (per the
type definition: "Custom api version to use. Defaults to `preview`."),
so this aligns the workflow with the SDK default. Probe should now
return MODEL_CAPABLE=true and the send_message scenario will run.
* test(cua): extend send_message and multi_turn waits to 30s
Assistant bubbles can take 15+ seconds to appear after send. Previous
5-second wait was too short and caused false failures even when API
calls succeeded. Re-check screenshots periodically up to 30s total.
* fix(cua): screen-relative send button threshold for #22
The send action's auto-locate filtered for y1 > 2200 and fell back to
hardcoded (996, 2358) — both assume a 1080x2400 panel. The CI emulator
(API 30 google_apis pixel profile) is 1080x1920, so:
- the bottom_buttons filter never matched any clickable element
- the fallback tap landed off-screen
→ 'ping' message never sent, scenario timed out with no bubbles.
Switch to a screen-relative threshold (bottom 25%) and a fallback that
uses get_screen_size() to land in the bottom-right corner regardless of
device resolution. This was masked until now because send_message was
gated by MODEL_CAPABLE=false in earlier CI runs.
Refs: #22
---------
Co-authored-by: dzianisv <dzianis.varabyou@gmail.com>
The CI opencode server had NO LLM provider configured — the server log only
showed "listening", never a model. opencode-ai (released npm pkg) does not read
AZURE_OPENAI_* for its own LLM; it needs an explicit provider in opencode.json +
a default `model`. So send_message/multi_turn could never pass and the gate was
stuck on --only-connect-scenario (UI journey minus the model reply).
- Wire opencode to the same Azure resource the CUA driver uses via a generated
~/.config/opencode/opencode.json (@ai-sdk/azure provider, resourceName derived
from the endpoint secret at runtime, apiKey from env, default model azure/gpt-5.4).
- Add a deterministic REST probe step: create a session + send a prompt and check
for an assistant reply BEFORE the ~30min emulator run, exporting MODEL_CAPABLE.
- Add --scenarios to android-cua-smoke.py to run an explicit named set.
- Emulator step now runs connect_and_verify_sessions + send_message +
verify_session_list when MODEL_CAPABLE=true; falls back to the UI-only journey
(connect + verify_session_list) otherwise, logging the environmental reason.
- Raise --max-steps to 40 so multiple scenarios fit.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: repoint OpenCode links to agentlabs.cc/opencode
agentlabs.cc/opencode and /opencode/privacy are now live (200). Repoint
README, distribution listings (Play/App Store/F-Droid/IzzyOnDroid/iOS),
docs, and in-app privacy links (settings + telemetry consent) from
www.vibebrowser.app/opencode to the canonical agentlabs.cc hub.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(ci): run local opencode server for CUA smoke true-E2E (#15)
GitHub-hosted runners can't reach the Tailscale dev server
(100.108.64.76:4096), so the CUA smoke always failed at session creation.
- Install opencode-ai and run `opencode serve` on the runner host; the
Android emulator reaches it via 10.0.2.2. OPENCODE_URL now points there.
- Healthcheck /global/health before launching the app; dump server log on
failure for diagnosis.
- Add --only-connect-scenario to the smoke script and run just the
connect-and-verify-sessions path in CI: deterministic, needs no model
backend. The scenario now creates a session if the list is empty, so a
fresh server still yields a non-empty list.
This makes the smoke a true E2E and also exercises the #10 sessions-list
rendering path against a real server.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(ci): emulator smoke script is dash, not bash — drop brace-group healthcheck
android-emulator-runner runs the script: block under /usr/bin/sh (dash). The
multi-line `|| { ...; }` healthcheck was a dash syntax error (end of file
unexpected), failing the step before the smoke ran. Replace with a non-fatal
one-line re-check; the server was already health-gated in the prior step.
* docs(tasks): record smoke CI round 1 failure + dash fix
---------
Co-authored-by: engineer <engineer@opencode.ai>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Rename commit missed the Python launcher constant; HEAD still targeted
the old package so the smoke could not find the installed app.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(sessions): use active connection client directly, remove roots filter
Root cause A: loadSessions was calling clientForDirectory(serverHome) which
scoped the session list to /home/azureuser — a different project than the
server's active CWD. Sessions in the current project (e.g. opencode-mobile)
were never returned.
Root cause B: roots:true filtered out sessions that have a parentID (sub-task /
AUTO-REVIEW sessions), hiding valid sessions from the list.
Fix: use connState.client directly (the connection's active directory) and drop
the roots filter so all sessions for that project are visible.
Also adds a verify_session_list CUA smoke scenario that navigates back to the
sessions tab after creating a session and asserts the list is non-empty —
covering the regression path that was previously untested.
* fix(sessions): fetch serverHome in addConnection so loadSessions shows correct sessions
Root cause: addConnection() built the HTTP client but never fetched serverHome
(only loadConnections and setActiveConnection did). When the user adds a new
connection (fresh install / first sign-in), serverHome = null, so loadSessions
fell through to connState.client (the server's CWD). On this dev server the CWD
is the deploy directory — 11 old May-19 sessions that are not the user's recent
work sessions.
Fix: addConnection now fetches currentProject + serverHome via the same
Promise.all as setActiveConnection, before calling set(). This ensures
loadSessions immediately uses clientForDirectory(serverHome) → the global
project → the user's actual recent parent sessions.
Also adds --opencode-url flag to the CUA smoke script, which appends a
connect_and_verify_sessions scenario that reproduces the regression:
python scripts/android-cua-smoke.py --opencode-url http://100.108.64.76:4096
* fix(sessions): recover home scope after fresh connect
Resolve stale deploy-only session list by recovering server home during first load and keeping regression coverage in default Android CUA smoke and CI.
* chore(release): bump version to 0.4.0
- Record each scenario as MP4 via ADB screenrecord in a background thread
- Pull video to /tmp/cua_<scenario>.mp4 after test completes (always)
- Upload to ArchiveBox if ARCHIVEBOX_URL + ARCHIVEBOX_API_KEY env vars set;
gracefully skips when not configured (CI default)
- Upload all artifacts (PNG screenshots + MP4 videos) always, not only on failure
- Pass ARCHIVEBOX_URL/ARCHIVEBOX_API_KEY secrets to workflow (optional)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(security): fail closed on biometric init error
H-03: setting isAuthenticated: true on initialization failure was a
security bypass — any crash during biometric setup granted full access.
Fail closed instead; user sees auth prompt on next open.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(security): use Crypto.randomUUID for connection IDs
H-04: Math.random() is not cryptographically random. Connection IDs are
used as SecureStore key suffixes; switch to expo-crypto randomUUID for
a secure source.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(deps): pin expo-crypto to ~15.0.9
15.0.10 does not exist on npm; ~15.0.9 is the latest stable in the 15.x series compatible with Expo SDK 54.
* feat: add OpenCode Connect coming-soon waitlist card
Adds a discoverable 'OpenCode Connect — Coming Soon' card to the
add-connection quick-connect screen. Users can enter their email and
tap 'Join Waitlist' to send a pre-filled mailto. No backend required.
* fix(cua): detect actual screen dimensions and fix JSON parsing
- Get real screen size via `wm size` instead of hardcoding 1080x2400;
emulator is 1080x1920 so y-coordinates were systematically off
- Extract first JSON object via regex when model returns multiple objects
- Use AZURE_OPENAI_MODEL env var for deployment name (defaults gpt-5.4)
- Add AZURE_DEV_AI_* path for Azure AI Foundry endpoints
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(security): SHA-pin upload-google-play and sanitize notification bodies
M-02: Pin r0adkll/upload-google-play to commit SHA e738b9d (v1.1.5)
to prevent supply-chain hijack via tag mutation.
M-03: Sanitize all push notification bodies — strip control chars,
truncate to 200 chars. Prevents server-supplied strings (error messages,
file paths from permission patterns, session titles) from leaking
unbounded text into the OS notification drawer.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* feat(privacy): add telemetry consent gate for Sentry crash reporting
Sentry was always-on, violating F-Droid anti-feature policy and user
trust norms. Now gated behind explicit opt-in:
- First-launch consent modal (TelemetryConsentModal) shows once on
fresh install; user can Allow or Decline.
- Consent state persisted in expo-secure-store (survives restarts).
- Settings > Privacy section: crash reporting toggle + privacy policy link.
- initSentry() called only after consent granted — not on app start.
Closes#3 (partial)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* feat(config): add real icons and complete iOS/Android app.json config
- Add 1024×1024 app icon, 432×432 adaptive icon foreground, 200×200 splash
- iOS: push notification entitlement (aps-environment: production), speech/
microphone/camera/photo usage descriptions for future features, disable
ITSAppUsesNonExemptEncryption
- Android: adaptive icon with dark background (#0F172A), versionCode: 1
- expo-notifications plugin wired in app.json
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* feat(dist): add iOS CI workflow, README rewrite, CONTRIBUTING, and LICENSE
- publish-app-store.yml: EAS Build + TestFlight submission; runs on tag/release/
workflow_dispatch; bumps ios.buildNumber from github.run_number
- README: full rewrite — features, install badges, connection guide, contributing
- CONTRIBUTING.md: contribution guide for OSS contributors
- LICENSE: MIT
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs(dist): add store listings, strategy, privacy policy, F-Droid/IzzyOnDroid templates
- distribution/strategy.md: monetization strategy (free client + opencode Cloud)
- distribution/play-listing.md: Google Play store copy (name, description, tags)
- distribution/app-store-listing.md: App Store listing copy
- distribution/privacy-policy.{md,html}: GDPR-compliant privacy policy
- distribution/PLAY_CONSOLE_SETUP.md: Play Console setup runbook
- distribution/ios-enrollment-runbook.md: Apple Developer Program enrollment steps
- distribution/SIGNING-KEY-FINGERPRINTS.md: keystore fingerprint for reproducible builds
- distribution/fdroid-submission/: F-Droid metadata template
- distribution/izzyondroid-submission/: IzzyOnDroid submission template
- distribution/whatsnew/: Play Store release notes (en-US)
- distribution/whatsnew-ios/: TestFlight release notes
- distribution/play-graphics/: Play Store screenshot placeholders
- distribution/app-store-graphics/: App Store screenshot placeholders
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(telemetry): handle SecureStore failure + Android back button
- add .catch() on loadTelemetryConsent() so SecureStore rejection
shows the consent modal instead of blocking startup forever
- add onRequestClose={onDecline} to Modal so Android back button
records the decline rather than silently dismissing
- fix catch block in telemetry.ts to not clobber _resolved when
SecureStore read fails mid-session
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(ci): run gradlew clean to prevent stale modules.json duplicate
Sentry Gradle plugin writes modules.json to src/main/assets; cached
build intermediates contain an old copy → mergeReleaseAssets fails
with 'Duplicate resources'. Running clean before assembleRelease
clears the intermediate state.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(ci): remove android build output cache causing duplicate modules.json
Caching android/app/build/intermediates and android/app/.cxx causes
two issues:
1. Stale modules.json in intermediates → Duplicate resources error
2. .cxx CMake artifacts reference absolute paths → ninja clean fails
Keeping only Gradle distribution cache (~/.gradle) which is safe.
Expo prebuild regenerates android sources fresh each run anyway.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
- Added 'send' action type that uses uiautomator XML to find the rightmost
clickable element in the bottom input bar (the send button)
- Added screen resolution (1080x2400) to LLM context for better coordinate estimation
- Updated system prompt to instruct model to use 'send' action instead of manual tap
- Updated scenarios with clearer step-by-step instructions
- Both send_message and multi_turn scenarios pass reliably