Commit Graph

32 Commits

Author SHA1 Message Date
Dennis V
00378ba7ed fix(ci): YAML syntax error — double-quotes in GH expression default value; add no-tap instruction to showcase typescript phase 2026-06-24 06:15:51 +00:00
Dennis V
a305936f2b feat(cua): add --e2e and --query modes with structured evaluation
--e2e mode: full end-to-end coding task scenario
  - connect → long-press FAB to create session in custom project dir
  - select AI model via model picker (hint substring match)
  - submit coding task → DETERMINISTIC API poll for session idle
  - DETERMINISTIC API message scan for target filename
  - DETERMINISTIC ADB uiautomator check for filename in UI
  - LLM screenshot + visual evaluation summary

--query mode: natural-language test description → structured test run
  - LLM planner converts the query into JSON phases + deterministic checks
  - Executes each phase via the CUA loop (with critical/informational split)
  - Runs deterministic checks: ui_text | session_idle | file_created
  - LLM evaluator produces scored JSON report: overall/score/phases/recommendations

New helpers:
  - wait_for_session_idle(): polls GET /session until status==idle (no LLM)
  - check_session_file_created(): scans session messages API for filename
  - _api_base(): translates emulator host route for host-side API calls
  - run_scenario_hello_world_e2e(): 8-phase hardcoded e2e scenario
  - run_query_test(): planner → execute → evaluator pipeline

Also adds hello_world_e2e to --scenarios catalog for named invocation.

Usage:
  # Hardcoded e2e:
  python scripts/android-cua-smoke.py --e2e --opencode-url http://100.108.64.76:4096

  # Natural-language query:
  python scripts/android-cua-smoke.py --query \
    'Open android app. Connect to server. Open ~/workspace/opencode-mobile. \
     Choose deepseek model. Ask to write hello_world.py. Validate it was created.'

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-06-24 03:53:01 +00:00
Dennis V
029bf2be04 fix(notifications): use user-friendly title and dedup keys for permission/question notifications
- Change permission notification title from `req.permission || 'Permission requested'`
  to the user-friendly 'Agent needs approval'; permission type + patterns now appear
  in the body (e.g. 'bash: echo hello') for context.
- Add `dedupeKey: `perm-${req.id}`` and `dedupeKey: `question-${req.id}``
  (60 s cooldown) to both events so a SSE reconnect after disconnect() clears state
  can't fire a second notification for the same pending request.
- Fix stale CUA-test comment that claimed 'Agent needs approval' did not exist;
  fallback assertion already matched correct title; update the comment to reflect
  the real events.ts behavior.

Closes #39

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-06-23 16:33:27 +00:00
Dennis V
5427abdb5e feat(ux): SSE disconnect/reconnect banner in session view (#42)
Shows amber 'Reconnecting… (attempt N)' banner when SSE is down.
Shows brief green 'Connected ✓' flash on reconnect (useRef transition
to avoid atomic state reset bug where lastDisconnectAt resets with
reconnectAttempts in the same set() call).

Banner disappears automatically when SSE is stable.

Updates CUA scenario to check for both ASCII and Unicode ellipsis.

Closes #42

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-06-23 16:31:58 +00:00
Dennis V
38da2813b0 test(cua): fix #39/#42 deterministic assertions to match real app behavior
#39 (backgrounded permission notification):
- Event is permission.asked, not permission.requested (events.ts already calls
  notify() for it; send() only fires while backgrounded).
- Assert on APP_PACKAGE (the only token guaranteed in every dumpsys record) plus
  the actual copy ('Permission requested' / 'A tool needs your approval') instead
  of the non-existent 'Agent needs approval' string.

#42 (SSE disconnect banner):
- reconnectAttempts only zeroes after STABLE_CONNECTION_MS (10s) past a healthy
  reconnect, and the pending backoff timer can take up to 15s — so the banner can
  linger ~25s. Poll up to 40s for dismissal instead of a fixed 15s sleep to avoid
  a false 'still showing' failure.

Refs #39 #42

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-06-23 15:36:55 +00:00
Dennis V
1f8bae3342 feat(cua): add deterministic ADB-based assertion helpers + feature test scenarios
Adds three deterministic helpers that use ADB instead of LLM vision,
so pass/fail cannot be hallucinated:

  check_ui_text(text)          — uiautomator XML dump + grep
  check_notification_drawer(text, timeout) — dumpsys notification poll
  simulate_network_drop() / restore_network() — svc wifi/data disable

Plus two new feature-test scenarios with deterministic gating:

  sse_disconnect_banner (#42)
    - ADB cuts WiFi+data, waits, checks UI XML for 'Reconnecting' text
    - LLM visual check is supplementary/informational only
    - ADB restores network, checks banner disappears

  backgrounded_permission_notification (#39)
    - ADB backgrounds app (Home key)
    - API sends permission-triggering message
    - adb dumpsys notification checked for 'Agent needs approval'
    - No LLM involved in pass/fail decision

Also adds background_app() / foreground_app() helpers.
Wires new scenarios into --scenarios catalog alongside LLM ones.
Updates --scenarios help text to document both types.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-06-23 15:19:34 +00:00
Dennis V
6198e4e165 fix(cua): stronger connect phase + fallback pre-create for external server
connect phase: now verifies the connection appears in the list after saving,
not just that the form was dismissed. Prevents false PASS when the LLM
declares connect done before the entry is actually visible.

_precreate_test_session: when external URL (Tailscale) times out from the
CI runner, fall back to localhost:4096 (the runner-local opencode serve).
This ensures the named-session assertion stays active in standard CI runs
while also working when dispatched against a live external server.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-06-23 09:11:20 +00:00
Dennis V
bbe8af9378 fix(cua): pre-create API call must use 127.0.0.1 not 10.0.2.2 on host
The CUA script runs on the CI runner (host), not inside the emulator.
10.0.2.2 is the emulator's special address for the host — it is only
reachable FROM INSIDE the emulator. Calling it from the runner always
fails, so _precreate_test_session returned None, and the session_list
phase fell back to the weak 'screen visible' assertion instead of the
strong 'pre-created session must appear' check.

Fix: replace 10.0.2.2 → 127.0.0.1 before making the pre-create call.
Localhost URLs (100.x.x.x, custom dev server) pass through unchanged.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-06-23 00:06:32 +00:00
Dennis V
cdbea50287 fix(cua): fix sessions_reload banner label (moved to step 4b)
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-06-23 00:04:20 +00:00
Dennis V
212b80b4a8 fix(cua): move sessions_reload before typescript; fix CI emulator boot
sessions_reload regression phase now runs BEFORE the TypeScript task.
Previously it was gated behind typescript which fails in CI when the model
is unavailable — meaning the actual sessions regression check never ran.

Phase order is now:
  connect → session_list (pre-created session required) → new_session
  → sessions_reload (navigate back, list must be non-empty) [CRITICAL]
  → typescript (informational) → verify (informational) → settings

Critical phases: connect, session_list, new_session, sessions_reload.
TypeScript/verify/settings are informational (model availability varies).

CI emulator fixes:
- api-level: 30 → 28 (more stable, boots reliably on ubuntu-latest)
- target: google_apis → default (lighter, no Play Services needed for
  sessions regression test, avoids known boot issues with google_apis)
- disable-animations: true (reduces boot overhead)
- emulator-boot-timeout: 600 (explicit, matches action default)
- Switch from --scenarios to --showcase (runs the new structured flow
  with _precreate_test_session + sessions_reload phase)
- Clear app state before install (pm clear) for deterministic first-run

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-06-23 00:02:13 +00:00
Dennis V
42aa883f99 test(cua): make session_list phase actually test sessions load regression
Previously the session_list phase goal said 'The session list may be empty
(no sessions yet) — that is fine' and 'Report done when you can see the
session list screen (even if empty)'. This means an empty sessions list
was treated as a test PASS, so every previous 'fix' was validated against
a test that cannot detect the regression.

Two changes:
1. _precreate_test_session(): calls POST /session via HTTP before the CUA
   starts. The session_list phase goal now explicitly names this session and
   requires it to be visible — if the app fails to load server sessions,
   the phase fails (not passes silently with an empty list).
   Falls back gracefully if the server is unreachable at pre-create time.

2. sessions_reload phase (new, critical): after completing the TypeScript
   task the test navigates back to the Sessions tab and asserts the list is
   non-empty. This catches the other variant of the regression — sessions
   vanishing after navigating away from a session and back.

Both phases are now in the critical list, so a failure in either causes the
overall test to report partial/fail instead of success.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-06-22 23:54:22 +00:00
Dennis V
c5f90ee3c0 fix(cua): explicitly prohibit stray taps after sending — stray tap navigates away from session and breaks SSE 2026-06-22 23:04:17 +00:00
Dennis V
b0851b5d47 Fix type action: use file-based input text with shell command substitution
ADB input text with %s escaping didn't trigger React Native onChangeText.
New approach: write text to /sdcard/ file on device, then use shell
command-substitution "$(cat ...)" to pass raw text to input text.
This preserves spaces without %s conversion and avoids quote-escaping
issues that prevented React Native from detecting text changes.
2026-06-22 22:17:12 +00:00
Dennis V
e71a2740f6 fix(cua): shell-quote input text and simplify prompt to avoid shell metacharacters 2026-06-22 21:45:09 +00:00
Dennis V
c002313ea0 fix(cua): remove send action — LLM taps send button via coordinates from screenshot instead 2026-06-22 21:19:24 +00:00
Dennis V
b7ddaa1350 fix(cua): send action dismisses keyboard first via KEYCODE_ESCAPE before locating send button 2026-06-22 20:47:41 +00:00
Dennis V
ea22d06f61 fix(cua): do not press back after typing — adb input text does not show keyboard, back navigates away 2026-06-22 20:20:44 +00:00
Dennis V
19b656aae8 cua: replace send_message pong with real coding task (helloworld.py + helloworld_test.py) 2026-06-22 19:44:29 +00:00
Dennis V
3c498972ce feat(cua): full onboarding showcase test - connect, session, TypeScript task, settings
Rewrites android-cua-smoke.py to demonstrate the complete first-run journey
instead of the previous "ping" smoke test. The new structured multi-phase
flow covers: server connection setup, session list, new session creation,
TypeScript hello-world task submission (watching tool calls/file writes),
output verification, and Settings/model-selection screenshot.

Key changes:
- run_onboarding_showcase() orchestrates 6 sequential CUA phases with
  per-phase goals, step budgets, and PASS/FAIL phase tracking
- run_cua_step() replaces run_cua() — accepts step_label, action_delay,
  saves labeled screenshots (/tmp/cua_<phase>_<step>.png) for debugging
- Global --speed-multiplier flag scales all _sleep() calls (0.5 = 2x faster)
- Showcase is now the default mode; legacy --goal / --scenarios flags retained
  for backwards compat and CI regression scenarios
- Tighter action_delay (0.7s) and trimmed history window (14 turns) vs
  previous 1.0s / 12 turns
- Phase banner log lines ("STEP N: ...") narrate the video in real time

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01No3k1AEioE4PNUZg12TxQo
2026-06-21 01:20:44 +00:00
Den
d6e84ff513 fix: stale session client ref and CUA smoke test improvements (#30)
* fix(sessions): use latestConnState.client to avoid stale reference after reconnect

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JKGMRpgihA4io2frodqLjt

* fix(cua): lru_cache get_screen_size, remove redundant if-matches guard, drop inner import re

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JKGMRpgihA4io2frodqLjt

* fix(cua): use center-x comparator for send button, defer screen_w fetch

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JKGMRpgihA4io2frodqLjt

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-20 03:38:06 -07:00
Den
71a1feb232 fix(ci): bump opencode Azure apiVersion to support /responses endpoint (#22) (#24)
* fix(ci): bump opencode Azure apiVersion to support /responses endpoint (#22)

The CUA smoke probe was returning MODEL_CAPABLE=false because opencode's
@ai-sdk/azure provider got 'API version not supported' from Azure on
/openai/v1/responses with apiVersion=2024-08-01-preview. Split into two
envs: keep the CUA driver on 2024-08-01-preview (chat-completions only)
and bump the opencode-side provider config to 2025-04-01-preview, which
supports the new responses API.

Effect: send_message/multi_turn scenarios get included again in CUA smoke
when the probe succeeds.

* ci(cua): bound curl timeouts + diag dump on server-start hang (#22)

Step 11 'Start opencode server' has hung past the 45-min job timeout in
two consecutive runs (27195348071, 27198238465). Local boot of
opencode-ai 1.16.2 with the same heredoc config is healthy in 3s, so
something is wrong specifically on the GH-hosted runner — likely curl
post-loop waiting indefinitely on an unresponsive server.

Adds:
- set -x for command tracing
- --connect-timeout 2 -m 5 on every curl so hangs cannot exceed 5s
- HEALTHY flag + explicit exit 1 (drops the unbounded post-loop curl)
- Periodic dump every 10s: server log tail, ss listening sockets,
  process liveness — so we can see WHY the server isn't replying

Pure diagnostics; no behaviour change for the green path.

* fix(ci): use api-version=preview for /openai/v1/responses (#22)

Reproduced the probe failure locally against the same Azure resource:
all date-based api-versions (2024-08-01-preview, 2024-12-01-preview,
2025-01-01-preview, 2025-03-01-preview, 2025-04-01-preview) return:

    {"error":{"code":"BadRequest","message":"API version not supported"}}

Only api-version=preview and api-version=v1 succeed (200). This is the
new Azure OpenAI v1 responses-API style; date strings are reserved for
the legacy /openai/deployments/{model}/chat/completions endpoint.

@ai-sdk/azure 3.x already defaults apiVersion to "preview" (per the
type definition: "Custom api version to use. Defaults to `preview`."),
so this aligns the workflow with the SDK default. Probe should now
return MODEL_CAPABLE=true and the send_message scenario will run.

* test(cua): extend send_message and multi_turn waits to 30s

Assistant bubbles can take 15+ seconds to appear after send. Previous
5-second wait was too short and caused false failures even when API
calls succeeded. Re-check screenshots periodically up to 30s total.

* fix(cua): screen-relative send button threshold for #22

The send action's auto-locate filtered for y1 > 2200 and fell back to
hardcoded (996, 2358) — both assume a 1080x2400 panel. The CI emulator
(API 30 google_apis pixel profile) is 1080x1920, so:
  - the bottom_buttons filter never matched any clickable element
  - the fallback tap landed off-screen
  → 'ping' message never sent, scenario timed out with no bubbles.

Switch to a screen-relative threshold (bottom 25%) and a fallback that
uses get_screen_size() to land in the bottom-right corner regardless of
device resolution. This was masked until now because send_message was
gated by MODEL_CAPABLE=false in earlier CI runs.

Refs: #22

---------

Co-authored-by: dzianisv <dzianis.varabyou@gmail.com>
2026-06-13 23:34:48 -07:00
engineer
cfb0d9fe32 test(ci): widen cua-smoke gate to real core journey (connect→send→reply→list)
The CI opencode server had NO LLM provider configured — the server log only
showed "listening", never a model. opencode-ai (released npm pkg) does not read
AZURE_OPENAI_* for its own LLM; it needs an explicit provider in opencode.json +
a default `model`. So send_message/multi_turn could never pass and the gate was
stuck on --only-connect-scenario (UI journey minus the model reply).

- Wire opencode to the same Azure resource the CUA driver uses via a generated
  ~/.config/opencode/opencode.json (@ai-sdk/azure provider, resourceName derived
  from the endpoint secret at runtime, apiKey from env, default model azure/gpt-5.4).
- Add a deterministic REST probe step: create a session + send a prompt and check
  for an assistant reply BEFORE the ~30min emulator run, exporting MODEL_CAPABLE.
- Add --scenarios to android-cua-smoke.py to run an explicit named set.
- Emulator step now runs connect_and_verify_sessions + send_message +
  verify_session_list when MODEL_CAPABLE=true; falls back to the UI-only journey
  (connect + verify_session_list) otherwise, logging the environmental reason.
- Raise --max-steps to 40 so multiple scenarios fit.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 05:58:19 -07:00
Den
c9a57901c4 fix(ci): CUA smoke true-E2E with local opencode server (#15) (#18)
* chore: repoint OpenCode links to agentlabs.cc/opencode

agentlabs.cc/opencode and /opencode/privacy are now live (200). Repoint
README, distribution listings (Play/App Store/F-Droid/IzzyOnDroid/iOS),
docs, and in-app privacy links (settings + telemetry consent) from
www.vibebrowser.app/opencode to the canonical agentlabs.cc hub.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(ci): run local opencode server for CUA smoke true-E2E (#15)

GitHub-hosted runners can't reach the Tailscale dev server
(100.108.64.76:4096), so the CUA smoke always failed at session creation.

- Install opencode-ai and run `opencode serve` on the runner host; the
  Android emulator reaches it via 10.0.2.2. OPENCODE_URL now points there.
- Healthcheck /global/health before launching the app; dump server log on
  failure for diagnosis.
- Add --only-connect-scenario to the smoke script and run just the
  connect-and-verify-sessions path in CI: deterministic, needs no model
  backend. The scenario now creates a session if the list is empty, so a
  fresh server still yields a non-empty list.

This makes the smoke a true E2E and also exercises the #10 sessions-list
rendering path against a real server.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(ci): emulator smoke script is dash, not bash — drop brace-group healthcheck

android-emulator-runner runs the script: block under /usr/bin/sh (dash). The
multi-line `|| { ...; }` healthcheck was a dash syntax error (end of file
unexpected), failing the step before the smoke ran. Replace with a non-fatal
one-line re-check; the server was already health-gated in the prior step.

* docs(tasks): record smoke CI round 1 failure + dash fix

---------

Co-authored-by: engineer <engineer@opencode.ai>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-01 15:32:06 -07:00
engineer
c699d0b0bb fix(ci): set CUA smoke APP_PACKAGE to cc.agentlabs.opencode
Rename commit missed the Python launcher constant; HEAD still targeted
the old package so the smoke could not find the installed app.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-05-30 04:14:53 -07:00
Den
32f7af4e11 fix(sessions): recover home-scoped list after fresh connect
* fix(sessions): use active connection client directly, remove roots filter

Root cause A: loadSessions was calling clientForDirectory(serverHome) which
scoped the session list to /home/azureuser — a different project than the
server's active CWD. Sessions in the current project (e.g. opencode-mobile)
were never returned.

Root cause B: roots:true filtered out sessions that have a parentID (sub-task /
AUTO-REVIEW sessions), hiding valid sessions from the list.

Fix: use connState.client directly (the connection's active directory) and drop
the roots filter so all sessions for that project are visible.

Also adds a verify_session_list CUA smoke scenario that navigates back to the
sessions tab after creating a session and asserts the list is non-empty —
covering the regression path that was previously untested.

* fix(sessions): fetch serverHome in addConnection so loadSessions shows correct sessions

Root cause: addConnection() built the HTTP client but never fetched serverHome
(only loadConnections and setActiveConnection did). When the user adds a new
connection (fresh install / first sign-in), serverHome = null, so loadSessions
fell through to connState.client (the server's CWD). On this dev server the CWD
is the deploy directory — 11 old May-19 sessions that are not the user's recent
work sessions.

Fix: addConnection now fetches currentProject + serverHome via the same
Promise.all as setActiveConnection, before calling set(). This ensures
loadSessions immediately uses clientForDirectory(serverHome) → the global
project → the user's actual recent parent sessions.

Also adds --opencode-url flag to the CUA smoke script, which appends a
connect_and_verify_sessions scenario that reproduces the regression:
  python scripts/android-cua-smoke.py --opencode-url http://100.108.64.76:4096

* fix(sessions): recover home scope after fresh connect

Resolve stale deploy-only session list by recovering server home during first load and keeping regression coverage in default Android CUA smoke and CI.

* chore(release): bump version to 0.4.0
2026-05-26 19:50:17 -07:00
Dennis V
2dd49117af feat(cua): add screen recording + ArchiveBox upload to smoke test
- Record each scenario as MP4 via ADB screenrecord in a background thread
- Pull video to /tmp/cua_<scenario>.mp4 after test completes (always)
- Upload to ArchiveBox if ARCHIVEBOX_URL + ARCHIVEBOX_API_KEY env vars set;
  gracefully skips when not configured (CI default)
- Upload all artifacts (PNG screenshots + MP4 videos) always, not only on failure
- Pass ARCHIVEBOX_URL/ARCHIVEBOX_API_KEY secrets to workflow (optional)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-26 01:20:47 +00:00
Den
2b9b571d6e feat(privacy+dist): telemetry consent gate + app store distribution prep (#4)
* fix(security): fail closed on biometric init error

H-03: setting isAuthenticated: true on initialization failure was a
security bypass — any crash during biometric setup granted full access.
Fail closed instead; user sees auth prompt on next open.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(security): use Crypto.randomUUID for connection IDs

H-04: Math.random() is not cryptographically random. Connection IDs are
used as SecureStore key suffixes; switch to expo-crypto randomUUID for
a secure source.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(deps): pin expo-crypto to ~15.0.9

15.0.10 does not exist on npm; ~15.0.9 is the latest stable in the 15.x series compatible with Expo SDK 54.

* feat: add OpenCode Connect coming-soon waitlist card

Adds a discoverable 'OpenCode Connect — Coming Soon' card to the
add-connection quick-connect screen. Users can enter their email and
tap 'Join Waitlist' to send a pre-filled mailto. No backend required.

* fix(cua): detect actual screen dimensions and fix JSON parsing

- Get real screen size via `wm size` instead of hardcoding 1080x2400;
  emulator is 1080x1920 so y-coordinates were systematically off
- Extract first JSON object via regex when model returns multiple objects
- Use AZURE_OPENAI_MODEL env var for deployment name (defaults gpt-5.4)
- Add AZURE_DEV_AI_* path for Azure AI Foundry endpoints

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(security): SHA-pin upload-google-play and sanitize notification bodies

M-02: Pin r0adkll/upload-google-play to commit SHA e738b9d (v1.1.5)
to prevent supply-chain hijack via tag mutation.

M-03: Sanitize all push notification bodies — strip control chars,
truncate to 200 chars. Prevents server-supplied strings (error messages,
file paths from permission patterns, session titles) from leaking
unbounded text into the OS notification drawer.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(privacy): add telemetry consent gate for Sentry crash reporting

Sentry was always-on, violating F-Droid anti-feature policy and user
trust norms. Now gated behind explicit opt-in:

- First-launch consent modal (TelemetryConsentModal) shows once on
  fresh install; user can Allow or Decline.
- Consent state persisted in expo-secure-store (survives restarts).
- Settings > Privacy section: crash reporting toggle + privacy policy link.
- initSentry() called only after consent granted — not on app start.

Closes #3 (partial)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(config): add real icons and complete iOS/Android app.json config

- Add 1024×1024 app icon, 432×432 adaptive icon foreground, 200×200 splash
- iOS: push notification entitlement (aps-environment: production), speech/
  microphone/camera/photo usage descriptions for future features, disable
  ITSAppUsesNonExemptEncryption
- Android: adaptive icon with dark background (#0F172A), versionCode: 1
- expo-notifications plugin wired in app.json

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(dist): add iOS CI workflow, README rewrite, CONTRIBUTING, and LICENSE

- publish-app-store.yml: EAS Build + TestFlight submission; runs on tag/release/
  workflow_dispatch; bumps ios.buildNumber from github.run_number
- README: full rewrite — features, install badges, connection guide, contributing
- CONTRIBUTING.md: contribution guide for OSS contributors
- LICENSE: MIT

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* docs(dist): add store listings, strategy, privacy policy, F-Droid/IzzyOnDroid templates

- distribution/strategy.md: monetization strategy (free client + opencode Cloud)
- distribution/play-listing.md: Google Play store copy (name, description, tags)
- distribution/app-store-listing.md: App Store listing copy
- distribution/privacy-policy.{md,html}: GDPR-compliant privacy policy
- distribution/PLAY_CONSOLE_SETUP.md: Play Console setup runbook
- distribution/ios-enrollment-runbook.md: Apple Developer Program enrollment steps
- distribution/SIGNING-KEY-FINGERPRINTS.md: keystore fingerprint for reproducible builds
- distribution/fdroid-submission/: F-Droid metadata template
- distribution/izzyondroid-submission/: IzzyOnDroid submission template
- distribution/whatsnew/: Play Store release notes (en-US)
- distribution/whatsnew-ios/: TestFlight release notes
- distribution/play-graphics/: Play Store screenshot placeholders
- distribution/app-store-graphics/: App Store screenshot placeholders

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(telemetry): handle SecureStore failure + Android back button

- add .catch() on loadTelemetryConsent() so SecureStore rejection
  shows the consent modal instead of blocking startup forever
- add onRequestClose={onDecline} to Modal so Android back button
  records the decline rather than silently dismissing
- fix catch block in telemetry.ts to not clobber _resolved when
  SecureStore read fails mid-session

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(ci): run gradlew clean to prevent stale modules.json duplicate

Sentry Gradle plugin writes modules.json to src/main/assets; cached
build intermediates contain an old copy → mergeReleaseAssets fails
with 'Duplicate resources'. Running clean before assembleRelease
clears the intermediate state.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(ci): remove android build output cache causing duplicate modules.json

Caching android/app/build/intermediates and android/app/.cxx causes
two issues:
1. Stale modules.json in intermediates → Duplicate resources error
2. .cxx CMake artifacts reference absolute paths → ninja clean fails

Keeping only Gradle distribution cache (~/.gradle) which is safe.
Expo prebuild regenerates android sources fresh each run anyway.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-25 17:42:03 -07:00
Ubuntu
37bb8bfe72 feat(cua): add 'send' action with UI auto-locate, fix coordinate accuracy, both scenarios pass
- Added 'send' action type that uses uiautomator XML to find the rightmost
  clickable element in the bottom input bar (the send button)
- Added screen resolution (1080x2400) to LLM context for better coordinate estimation
- Updated system prompt to instruct model to use 'send' action instead of manual tap
- Updated scenarios with clearer step-by-step instructions
- Both send_message and multi_turn scenarios pass reliably
2026-05-19 08:26:48 +00:00
Ubuntu
5473a488fc feat(ci): add CUA smoke test workflow, improve screenshot reliability, add multi-turn scenario 2026-05-19 07:47:04 +00:00
Ubuntu
7804b7a3b9 fix(cua): add Azure OpenAI support, fix max_tokens → max_completion_tokens, increase screencap timeout 2026-05-19 06:24:29 +00:00
Ubuntu
e08a4a7a36 Fix screenshot capture and add multi-provider support
- Fix adb exec-out binary pipe for screenshot capture
- Add XAI, Azure, and any OpenAI-compatible endpoint support
- Add rate limit retry logic
- Document all supported provider configurations
2026-05-19 02:17:23 +00:00
Ubuntu
4f39967606 Add LLM-powered Android CUA smoke test script
Implements a computer-use agent loop for E2E testing:
screenshot → vision LLM → action → repeat

Supports OpenAI, Gemini, xAI via OpenAI-compatible endpoints.
Inspired by openai/openai-cua-sample-app, MobileAgent, AppAgent.
2026-05-19 01:08:03 +00:00