Commit Graph

17 Commits

Author SHA1 Message Date
Dennis V
b7ddaa1350 fix(cua): send action dismisses keyboard first via KEYCODE_ESCAPE before locating send button 2026-06-22 20:47:41 +00:00
Dennis V
ea22d06f61 fix(cua): do not press back after typing — adb input text does not show keyboard, back navigates away 2026-06-22 20:20:44 +00:00
Dennis V
19b656aae8 cua: replace send_message pong with real coding task (helloworld.py + helloworld_test.py) 2026-06-22 19:44:29 +00:00
Dennis V
3c498972ce feat(cua): full onboarding showcase test - connect, session, TypeScript task, settings
Rewrites android-cua-smoke.py to demonstrate the complete first-run journey
instead of the previous "ping" smoke test. The new structured multi-phase
flow covers: server connection setup, session list, new session creation,
TypeScript hello-world task submission (watching tool calls/file writes),
output verification, and Settings/model-selection screenshot.

Key changes:
- run_onboarding_showcase() orchestrates 6 sequential CUA phases with
  per-phase goals, step budgets, and PASS/FAIL phase tracking
- run_cua_step() replaces run_cua() — accepts step_label, action_delay,
  saves labeled screenshots (/tmp/cua_<phase>_<step>.png) for debugging
- Global --speed-multiplier flag scales all _sleep() calls (0.5 = 2x faster)
- Showcase is now the default mode; legacy --goal / --scenarios flags retained
  for backwards compat and CI regression scenarios
- Tighter action_delay (0.7s) and trimmed history window (14 turns) vs
  previous 1.0s / 12 turns
- Phase banner log lines ("STEP N: ...") narrate the video in real time

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01No3k1AEioE4PNUZg12TxQo
2026-06-21 01:20:44 +00:00
Den
d6e84ff513 fix: stale session client ref and CUA smoke test improvements (#30)
* fix(sessions): use latestConnState.client to avoid stale reference after reconnect

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JKGMRpgihA4io2frodqLjt

* fix(cua): lru_cache get_screen_size, remove redundant if-matches guard, drop inner import re

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JKGMRpgihA4io2frodqLjt

* fix(cua): use center-x comparator for send button, defer screen_w fetch

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JKGMRpgihA4io2frodqLjt

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-20 03:38:06 -07:00
Den
71a1feb232 fix(ci): bump opencode Azure apiVersion to support /responses endpoint (#22) (#24)
* fix(ci): bump opencode Azure apiVersion to support /responses endpoint (#22)

The CUA smoke probe was returning MODEL_CAPABLE=false because opencode's
@ai-sdk/azure provider got 'API version not supported' from Azure on
/openai/v1/responses with apiVersion=2024-08-01-preview. Split into two
envs: keep the CUA driver on 2024-08-01-preview (chat-completions only)
and bump the opencode-side provider config to 2025-04-01-preview, which
supports the new responses API.

Effect: send_message/multi_turn scenarios get included again in CUA smoke
when the probe succeeds.

* ci(cua): bound curl timeouts + diag dump on server-start hang (#22)

Step 11 'Start opencode server' has hung past the 45-min job timeout in
two consecutive runs (27195348071, 27198238465). Local boot of
opencode-ai 1.16.2 with the same heredoc config is healthy in 3s, so
something is wrong specifically on the GH-hosted runner — likely curl
post-loop waiting indefinitely on an unresponsive server.

Adds:
- set -x for command tracing
- --connect-timeout 2 -m 5 on every curl so hangs cannot exceed 5s
- HEALTHY flag + explicit exit 1 (drops the unbounded post-loop curl)
- Periodic dump every 10s: server log tail, ss listening sockets,
  process liveness — so we can see WHY the server isn't replying

Pure diagnostics; no behaviour change for the green path.

* fix(ci): use api-version=preview for /openai/v1/responses (#22)

Reproduced the probe failure locally against the same Azure resource:
all date-based api-versions (2024-08-01-preview, 2024-12-01-preview,
2025-01-01-preview, 2025-03-01-preview, 2025-04-01-preview) return:

    {"error":{"code":"BadRequest","message":"API version not supported"}}

Only api-version=preview and api-version=v1 succeed (200). This is the
new Azure OpenAI v1 responses-API style; date strings are reserved for
the legacy /openai/deployments/{model}/chat/completions endpoint.

@ai-sdk/azure 3.x already defaults apiVersion to "preview" (per the
type definition: "Custom api version to use. Defaults to `preview`."),
so this aligns the workflow with the SDK default. Probe should now
return MODEL_CAPABLE=true and the send_message scenario will run.

* test(cua): extend send_message and multi_turn waits to 30s

Assistant bubbles can take 15+ seconds to appear after send. Previous
5-second wait was too short and caused false failures even when API
calls succeeded. Re-check screenshots periodically up to 30s total.

* fix(cua): screen-relative send button threshold for #22

The send action's auto-locate filtered for y1 > 2200 and fell back to
hardcoded (996, 2358) — both assume a 1080x2400 panel. The CI emulator
(API 30 google_apis pixel profile) is 1080x1920, so:
  - the bottom_buttons filter never matched any clickable element
  - the fallback tap landed off-screen
  → 'ping' message never sent, scenario timed out with no bubbles.

Switch to a screen-relative threshold (bottom 25%) and a fallback that
uses get_screen_size() to land in the bottom-right corner regardless of
device resolution. This was masked until now because send_message was
gated by MODEL_CAPABLE=false in earlier CI runs.

Refs: #22

---------

Co-authored-by: dzianisv <dzianis.varabyou@gmail.com>
2026-06-13 23:34:48 -07:00
engineer
cfb0d9fe32 test(ci): widen cua-smoke gate to real core journey (connect→send→reply→list)
The CI opencode server had NO LLM provider configured — the server log only
showed "listening", never a model. opencode-ai (released npm pkg) does not read
AZURE_OPENAI_* for its own LLM; it needs an explicit provider in opencode.json +
a default `model`. So send_message/multi_turn could never pass and the gate was
stuck on --only-connect-scenario (UI journey minus the model reply).

- Wire opencode to the same Azure resource the CUA driver uses via a generated
  ~/.config/opencode/opencode.json (@ai-sdk/azure provider, resourceName derived
  from the endpoint secret at runtime, apiKey from env, default model azure/gpt-5.4).
- Add a deterministic REST probe step: create a session + send a prompt and check
  for an assistant reply BEFORE the ~30min emulator run, exporting MODEL_CAPABLE.
- Add --scenarios to android-cua-smoke.py to run an explicit named set.
- Emulator step now runs connect_and_verify_sessions + send_message +
  verify_session_list when MODEL_CAPABLE=true; falls back to the UI-only journey
  (connect + verify_session_list) otherwise, logging the environmental reason.
- Raise --max-steps to 40 so multiple scenarios fit.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 05:58:19 -07:00
Den
c9a57901c4 fix(ci): CUA smoke true-E2E with local opencode server (#15) (#18)
* chore: repoint OpenCode links to agentlabs.cc/opencode

agentlabs.cc/opencode and /opencode/privacy are now live (200). Repoint
README, distribution listings (Play/App Store/F-Droid/IzzyOnDroid/iOS),
docs, and in-app privacy links (settings + telemetry consent) from
www.vibebrowser.app/opencode to the canonical agentlabs.cc hub.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(ci): run local opencode server for CUA smoke true-E2E (#15)

GitHub-hosted runners can't reach the Tailscale dev server
(100.108.64.76:4096), so the CUA smoke always failed at session creation.

- Install opencode-ai and run `opencode serve` on the runner host; the
  Android emulator reaches it via 10.0.2.2. OPENCODE_URL now points there.
- Healthcheck /global/health before launching the app; dump server log on
  failure for diagnosis.
- Add --only-connect-scenario to the smoke script and run just the
  connect-and-verify-sessions path in CI: deterministic, needs no model
  backend. The scenario now creates a session if the list is empty, so a
  fresh server still yields a non-empty list.

This makes the smoke a true E2E and also exercises the #10 sessions-list
rendering path against a real server.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(ci): emulator smoke script is dash, not bash — drop brace-group healthcheck

android-emulator-runner runs the script: block under /usr/bin/sh (dash). The
multi-line `|| { ...; }` healthcheck was a dash syntax error (end of file
unexpected), failing the step before the smoke ran. Replace with a non-fatal
one-line re-check; the server was already health-gated in the prior step.

* docs(tasks): record smoke CI round 1 failure + dash fix

---------

Co-authored-by: engineer <engineer@opencode.ai>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-01 15:32:06 -07:00
engineer
c699d0b0bb fix(ci): set CUA smoke APP_PACKAGE to cc.agentlabs.opencode
Rename commit missed the Python launcher constant; HEAD still targeted
the old package so the smoke could not find the installed app.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-05-30 04:14:53 -07:00
Den
32f7af4e11 fix(sessions): recover home-scoped list after fresh connect
* fix(sessions): use active connection client directly, remove roots filter

Root cause A: loadSessions was calling clientForDirectory(serverHome) which
scoped the session list to /home/azureuser — a different project than the
server's active CWD. Sessions in the current project (e.g. opencode-mobile)
were never returned.

Root cause B: roots:true filtered out sessions that have a parentID (sub-task /
AUTO-REVIEW sessions), hiding valid sessions from the list.

Fix: use connState.client directly (the connection's active directory) and drop
the roots filter so all sessions for that project are visible.

Also adds a verify_session_list CUA smoke scenario that navigates back to the
sessions tab after creating a session and asserts the list is non-empty —
covering the regression path that was previously untested.

* fix(sessions): fetch serverHome in addConnection so loadSessions shows correct sessions

Root cause: addConnection() built the HTTP client but never fetched serverHome
(only loadConnections and setActiveConnection did). When the user adds a new
connection (fresh install / first sign-in), serverHome = null, so loadSessions
fell through to connState.client (the server's CWD). On this dev server the CWD
is the deploy directory — 11 old May-19 sessions that are not the user's recent
work sessions.

Fix: addConnection now fetches currentProject + serverHome via the same
Promise.all as setActiveConnection, before calling set(). This ensures
loadSessions immediately uses clientForDirectory(serverHome) → the global
project → the user's actual recent parent sessions.

Also adds --opencode-url flag to the CUA smoke script, which appends a
connect_and_verify_sessions scenario that reproduces the regression:
  python scripts/android-cua-smoke.py --opencode-url http://100.108.64.76:4096

* fix(sessions): recover home scope after fresh connect

Resolve stale deploy-only session list by recovering server home during first load and keeping regression coverage in default Android CUA smoke and CI.

* chore(release): bump version to 0.4.0
2026-05-26 19:50:17 -07:00
Dennis V
2dd49117af feat(cua): add screen recording + ArchiveBox upload to smoke test
- Record each scenario as MP4 via ADB screenrecord in a background thread
- Pull video to /tmp/cua_<scenario>.mp4 after test completes (always)
- Upload to ArchiveBox if ARCHIVEBOX_URL + ARCHIVEBOX_API_KEY env vars set;
  gracefully skips when not configured (CI default)
- Upload all artifacts (PNG screenshots + MP4 videos) always, not only on failure
- Pass ARCHIVEBOX_URL/ARCHIVEBOX_API_KEY secrets to workflow (optional)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-26 01:20:47 +00:00
Den
2b9b571d6e feat(privacy+dist): telemetry consent gate + app store distribution prep (#4)
* fix(security): fail closed on biometric init error

H-03: setting isAuthenticated: true on initialization failure was a
security bypass — any crash during biometric setup granted full access.
Fail closed instead; user sees auth prompt on next open.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(security): use Crypto.randomUUID for connection IDs

H-04: Math.random() is not cryptographically random. Connection IDs are
used as SecureStore key suffixes; switch to expo-crypto randomUUID for
a secure source.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(deps): pin expo-crypto to ~15.0.9

15.0.10 does not exist on npm; ~15.0.9 is the latest stable in the 15.x series compatible with Expo SDK 54.

* feat: add OpenCode Connect coming-soon waitlist card

Adds a discoverable 'OpenCode Connect — Coming Soon' card to the
add-connection quick-connect screen. Users can enter their email and
tap 'Join Waitlist' to send a pre-filled mailto. No backend required.

* fix(cua): detect actual screen dimensions and fix JSON parsing

- Get real screen size via `wm size` instead of hardcoding 1080x2400;
  emulator is 1080x1920 so y-coordinates were systematically off
- Extract first JSON object via regex when model returns multiple objects
- Use AZURE_OPENAI_MODEL env var for deployment name (defaults gpt-5.4)
- Add AZURE_DEV_AI_* path for Azure AI Foundry endpoints

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(security): SHA-pin upload-google-play and sanitize notification bodies

M-02: Pin r0adkll/upload-google-play to commit SHA e738b9d (v1.1.5)
to prevent supply-chain hijack via tag mutation.

M-03: Sanitize all push notification bodies — strip control chars,
truncate to 200 chars. Prevents server-supplied strings (error messages,
file paths from permission patterns, session titles) from leaking
unbounded text into the OS notification drawer.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(privacy): add telemetry consent gate for Sentry crash reporting

Sentry was always-on, violating F-Droid anti-feature policy and user
trust norms. Now gated behind explicit opt-in:

- First-launch consent modal (TelemetryConsentModal) shows once on
  fresh install; user can Allow or Decline.
- Consent state persisted in expo-secure-store (survives restarts).
- Settings > Privacy section: crash reporting toggle + privacy policy link.
- initSentry() called only after consent granted — not on app start.

Closes #3 (partial)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(config): add real icons and complete iOS/Android app.json config

- Add 1024×1024 app icon, 432×432 adaptive icon foreground, 200×200 splash
- iOS: push notification entitlement (aps-environment: production), speech/
  microphone/camera/photo usage descriptions for future features, disable
  ITSAppUsesNonExemptEncryption
- Android: adaptive icon with dark background (#0F172A), versionCode: 1
- expo-notifications plugin wired in app.json

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(dist): add iOS CI workflow, README rewrite, CONTRIBUTING, and LICENSE

- publish-app-store.yml: EAS Build + TestFlight submission; runs on tag/release/
  workflow_dispatch; bumps ios.buildNumber from github.run_number
- README: full rewrite — features, install badges, connection guide, contributing
- CONTRIBUTING.md: contribution guide for OSS contributors
- LICENSE: MIT

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* docs(dist): add store listings, strategy, privacy policy, F-Droid/IzzyOnDroid templates

- distribution/strategy.md: monetization strategy (free client + opencode Cloud)
- distribution/play-listing.md: Google Play store copy (name, description, tags)
- distribution/app-store-listing.md: App Store listing copy
- distribution/privacy-policy.{md,html}: GDPR-compliant privacy policy
- distribution/PLAY_CONSOLE_SETUP.md: Play Console setup runbook
- distribution/ios-enrollment-runbook.md: Apple Developer Program enrollment steps
- distribution/SIGNING-KEY-FINGERPRINTS.md: keystore fingerprint for reproducible builds
- distribution/fdroid-submission/: F-Droid metadata template
- distribution/izzyondroid-submission/: IzzyOnDroid submission template
- distribution/whatsnew/: Play Store release notes (en-US)
- distribution/whatsnew-ios/: TestFlight release notes
- distribution/play-graphics/: Play Store screenshot placeholders
- distribution/app-store-graphics/: App Store screenshot placeholders

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(telemetry): handle SecureStore failure + Android back button

- add .catch() on loadTelemetryConsent() so SecureStore rejection
  shows the consent modal instead of blocking startup forever
- add onRequestClose={onDecline} to Modal so Android back button
  records the decline rather than silently dismissing
- fix catch block in telemetry.ts to not clobber _resolved when
  SecureStore read fails mid-session

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(ci): run gradlew clean to prevent stale modules.json duplicate

Sentry Gradle plugin writes modules.json to src/main/assets; cached
build intermediates contain an old copy → mergeReleaseAssets fails
with 'Duplicate resources'. Running clean before assembleRelease
clears the intermediate state.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(ci): remove android build output cache causing duplicate modules.json

Caching android/app/build/intermediates and android/app/.cxx causes
two issues:
1. Stale modules.json in intermediates → Duplicate resources error
2. .cxx CMake artifacts reference absolute paths → ninja clean fails

Keeping only Gradle distribution cache (~/.gradle) which is safe.
Expo prebuild regenerates android sources fresh each run anyway.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-25 17:42:03 -07:00
Ubuntu
37bb8bfe72 feat(cua): add 'send' action with UI auto-locate, fix coordinate accuracy, both scenarios pass
- Added 'send' action type that uses uiautomator XML to find the rightmost
  clickable element in the bottom input bar (the send button)
- Added screen resolution (1080x2400) to LLM context for better coordinate estimation
- Updated system prompt to instruct model to use 'send' action instead of manual tap
- Updated scenarios with clearer step-by-step instructions
- Both send_message and multi_turn scenarios pass reliably
2026-05-19 08:26:48 +00:00
Ubuntu
5473a488fc feat(ci): add CUA smoke test workflow, improve screenshot reliability, add multi-turn scenario 2026-05-19 07:47:04 +00:00
Ubuntu
7804b7a3b9 fix(cua): add Azure OpenAI support, fix max_tokens → max_completion_tokens, increase screencap timeout 2026-05-19 06:24:29 +00:00
Ubuntu
e08a4a7a36 Fix screenshot capture and add multi-provider support
- Fix adb exec-out binary pipe for screenshot capture
- Add XAI, Azure, and any OpenAI-compatible endpoint support
- Add rate limit retry logic
- Document all supported provider configurations
2026-05-19 02:17:23 +00:00
Ubuntu
4f39967606 Add LLM-powered Android CUA smoke test script
Implements a computer-use agent loop for E2E testing:
screenshot → vision LLM → action → repeat

Supports OpenAI, Gemini, xAI via OpenAI-compatible endpoints.
Inspired by openai/openai-cua-sample-app, MobileAgent, AppAgent.
2026-05-19 01:08:03 +00:00