Review fixes (REQUEST_CHANGES round 1):
1. HIGH activation-negative-401.yaml: after dismissing the "Connection
Failed" alert the app stays on the Add-Connection modal
(handleQuickConnect's failure branch never calls router.back()), so the
old `text: "No Connection"` assertion (Sessions-tab empty state) could
never pass. Now asserts connect-submit-button is still visible instead.
2. MEDIUM mock-opencode-server.ts: prompt_async now parses the request
body, persists the USER's message, and broadcasts it (message.updated +
message.part.updated) BEFORE the canned assistant reply — matching real
server behavior. Without this, the app's handleEvent strips the
optimistic temp- user message when the assistant's message.updated
arrives and the sent message vanishes from the transcript.
activation-positive.yaml now also asserts chat-bubble-user and the
user's message text are visible after the reply lands, so that
regression class is actually covered.
3. MEDIUM activation-e2e.yml: timeout-minutes 15 -> 60. The job runs the
same npm install + prebuild + assembleRelease + emulator pipeline that
cua-smoke.yml budgets 60 min for (emulator-boot-timeout alone is 10 min).
4. MEDIUM activation-e2e.yml: replicated cua-smoke.yml's "Purge stale
generated sources" step — the Gradle cache key/restore-keys are shared
with that workflow, so the stale-autolinking-tree failure mode
(compileReleaseJavaWithJavac against the old package id) applies here too.
Verified locally: tsc --noEmit clean; npm test 81/81 pass; all three
touched YAML files parse valid; mock server exercised standalone —
full prompt cycle confirms GET /session/:id/message returns BOTH user
and assistant messages, SSE order is message.updated(user) ->
message.part.updated(user) -> busy -> message.updated(assistant) ->
message.part.updated(assistant) -> idle, user events carry the
sessionID/messageID fields handleEvent filters on, and --fail-auth
mode returns 401. Still no emulator run in this environment.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NJKAQ6HAikWGQK7PGZ5Y4E
Adds deterministic end-to-end coverage for first-open -> telemetry consent
-> server URL entry -> connect -> send first message -> receive reply,
targeting the 0%-7-day-retention investigation (GitHub issue #76).
- tests/fixtures/mock-opencode-server.ts: dependency-free HTTP+SSE stub
matching the REAL client protocol (src/lib/sdk.ts) — REST + a single
long-lived GET /global/event SSE stream, no WebSocket. Supports a
--fail-auth mode that 401s every request to exercise the connect-time
auth-failure class.
- .maestro/flows/activation-positive.yaml: consent -> quick connect ->
new session -> send message -> assert streamed reply renders, with a
screenshot at every step (positive-S1..S8).
- .maestro/flows/activation-negative-401.yaml: same setup against the
--fail-auth server, asserts Quick Connect's existing "Connection Failed"
alert is shown (not silently swallowed) and that the connection is not
saved. Flags in comments that Advanced-mode Save (handleAdvancedSave)
still has no testConnection() check and is a known, uncovered gap.
- testID props added (no restructuring) to the screens/components the
flows drive: TelemetryConsentModal, connection/add.tsx, tabs/index.tsx,
session/[id].tsx, MessageBubble.
- .github/workflows/activation-e2e.yml: new CI job — Android emulator via
reactivecircus/android-emulator-runner, builds the debug-signed APK,
starts both mock server instances, runs both Maestro flows, uploads
screenshots via actions/upload-artifact. Kept separate from the existing
vision-driven cua-smoke.yml, which needs a live server + LLM and isn't
suited to tight deterministic regression assertions.
- .gitignore: Maestro takeScreenshot output is never committed.
Verified locally: mock server exercised standalone via curl (health,
project/current, path, session create, SSE event ordering, message
persistence) in both normal and --fail-auth modes; both Maestro flow
files validated as well-formed YAML; tsc --noEmit clean on all changed
files. No lint script exists in this repo (N/A). Full emulator execution
was not run — no Android SDK/emulator available in this environment.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NJKAQ6HAikWGQK7PGZ5Y4E
Adds recent and server-project discovery to the new-session directory picker. Reviewed against current main; Android, iOS Simulator, and mandatory CUA checks are green.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Adds privacy-safe aggregate product intelligence, reviewed/versioned website assets, and a dispatch-only rollout until the dedicated Sentry token is verified. Independent review blockers were fixed in 8bc47e4; app checks, website production build, Android CI, and iOS CI are green.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
cua-smoke.yml:
- Add scenario/query/e2e_* workflow_dispatch inputs
- Runner step dispatches to --query / --e2e / --showcase based on inputs
- Upload /tmp/cua_eval_report.json as artifact (--query output)
AGENTS.md:
- Document all 3 run modes: --showcase, --e2e, --query
- List available models on dev server (deepseek-v4-flash-free etc.)
- Add dispatch inputs reference for CI
- Add --e2e / --query to 'when to run' guidance
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- docs-site/index.html: add Vercel Analytics + Speed Insights CDN scripts
(only fires on opencode.agentlabs.cc served via Vercel, not GitHub Pages)
- scripts/triage-reviews.py: fetch recent Play Store reviews via Android
Publisher API, create GitHub issues for ≤3★ reviews not yet tracked
Run review triage manually on VM:
DAYS_BACK=7 GOOGLE_SERVICE_ACCOUNT_JSON=... GH_TOKEN=... python3 scripts/triage-reviews.py
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Shows every major screen: connect, sessions list, chat input,
tool calls streaming, file writes, completed result, settings/model.
01 - Add connection screen (onboarding)
02 - Sessions list loaded from server
03 - Session chat view with message sent
04 - AI tool calls streaming (reading files)
05 - AI writing TypeScript files to disk
06 - Completed session — hello.ts created
07 - Settings + model selection
Website updated to show all 7 in a tighter grid (max-width 260px).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- Change permission notification title from `req.permission || 'Permission requested'`
to the user-friendly 'Agent needs approval'; permission type + patterns now appear
in the body (e.g. 'bash: echo hello') for context.
- Add `dedupeKey: `perm-${req.id}`` and `dedupeKey: `question-${req.id}``
(60 s cooldown) to both events so a SSE reconnect after disconnect() clears state
can't fire a second notification for the same pending request.
- Fix stale CUA-test comment that claimed 'Agent needs approval' did not exist;
fallback assertion already matched correct title; update the comment to reflect
the real events.ts behavior.
Closes#39
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Shows amber 'Reconnecting… (attempt N)' banner when SSE is down.
Shows brief green 'Connected ✓' flash on reconnect (useRef transition
to avoid atomic state reset bug where lastDisconnectAt resets with
reconnectAttempts in the same set() call).
Banner disappears automatically when SSE is stable.
Updates CUA scenario to check for both ASCII and Unicode ellipsis.
Closes#42
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
#39 (backgrounded permission notification):
- Event is permission.asked, not permission.requested (events.ts already calls
notify() for it; send() only fires while backgrounded).
- Assert on APP_PACKAGE (the only token guaranteed in every dumpsys record) plus
the actual copy ('Permission requested' / 'A tool needs your approval') instead
of the non-existent 'Agent needs approval' string.
#42 (SSE disconnect banner):
- reconnectAttempts only zeroes after STABLE_CONNECTION_MS (10s) past a healthy
reconnect, and the pending backoff timer can take up to 15s — so the banner can
linger ~25s. Poll up to 40s for dismissal instead of a fixed 15s sleep to avoid
a false 'still showing' failure.
Refs #39#42
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Rewrite <current_mission>: the Play Store publication blocker is cleared
(cc.agentlabs.opencode live on internal track), website opencode.agentlabs.cc
is live (HTTP 200), and the CUA sessions_reload phase is green (run 28002986180).
Refocus the mission on growth: Play Store promotion, F-Droid mainline, content
marketing, and organic SEO/ASO.
Closes#44
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- §4: path.get() is still live (feeds serverHome → directory switcher ~ expansion);
only the dead session-scoping plumbing (sessionScope.ts) was removed in 472ff8d.
Clarify rather than delete, since the call site is not dead.
- §2: document src/lib/speech.ts (experimental voice input; PRD §7 scope caveat).
- §7: app.json and package.json versions are kept in sync since v0.4.6 (both 0.4.6).
Refs #43
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Adds three deterministic helpers that use ADB instead of LLM vision,
so pass/fail cannot be hallucinated:
check_ui_text(text) — uiautomator XML dump + grep
check_notification_drawer(text, timeout) — dumpsys notification poll
simulate_network_drop() / restore_network() — svc wifi/data disable
Plus two new feature-test scenarios with deterministic gating:
sse_disconnect_banner (#42)
- ADB cuts WiFi+data, waits, checks UI XML for 'Reconnecting' text
- LLM visual check is supplementary/informational only
- ADB restores network, checks banner disappears
backgrounded_permission_notification (#39)
- ADB backgrounds app (Home key)
- API sends permission-triggering message
- adb dumpsys notification checked for 'Agent needs approval'
- No LLM involved in pass/fail decision
Also adds background_app() / foreground_app() helpers.
Wires new scenarios into --scenarios catalog alongside LLM ones.
Updates --scenarios help text to document both types.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
connect phase: now verifies the connection appears in the list after saving,
not just that the form was dismissed. Prevents false PASS when the LLM
declares connect done before the entry is actually visible.
_precreate_test_session: when external URL (Tailscale) times out from the
CI runner, fall back to localhost:4096 (the runner-local opencode serve).
This ensures the named-session assertion stays active in standard CI runs
while also working when dispatched against a live external server.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
When dispatched with opencode_url set (e.g. Tailscale URL), the workflow
skips starting a local opencode server and the model probe, and points the
emulator directly at the external server.
This allows running the full CUA test against a live dev server without
depending on the flaky CI-hosted opencode serve setup.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The probe step computed a SCENARIOS output, but the emulator step runs
--showcase hardcoded and never consumes steps.probe.outputs.SCENARIOS, so
the selection logic was dead. Drop it and keep MODEL_CAPABLE as an
informational signal. Showcase mode handles model-unavailability gracefully
(typescript/verify phases are informational-only).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
sessionScopeDirectory() always returned null and ignored its home arg, so
the scope plumbing in loadSessions/createSession was dead: scopeDir was
always null, listClient always the default client, and the lazy path.get()
fetch fed only that dead branch. Collapse both call sites to use the
connection's default client directly and delete the now-orphaned
sessionScope.ts helper and its test.
serverHome is intentionally KEPT in connections.ts: it is still consumed by
the directory switcher UI (DirectorySwitcher.tsx, app/(tabs)/index.tsx) for
~ path expansion, so it is not dead code.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Replace stale opencode.vibebrowser.app / www.vibebrowser.app domain refs
with the current agentlabs.cc/opencode branding, and update the privacy
policy package id ai.opencode.mobile -> cc.agentlabs.opencode.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Screenshots from CUA smoke test showing:
- 01: Sessions list loading real sessions from server on connect
- 02: Streaming AI response in a coding session
- 03: Sessions tab after navigating back (sessions_reload regression guard)
Demo: 10x speed video (138s → 14s) of full onboarding flow.
Website now shows <video> with mp4 source + gif fallback.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- full_description.txt: replace fluffy copy with concrete HOW IT WORKS steps,
TAILSCALE recommendation, explicit WHAT THIS APP IS NOT section, updated
support email to support@agentlabs.cc
- short_description.txt: 77-char tagline focused on opencode serve remote control
- docs-site/index.html: update meta description, og:description, twitter:description,
ld+json description, hero tagline, and 'What is OpenCode Mobile?' paragraphs
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Replace all @vibebrowser.app email addresses with @agentlabs.cc across
22 files including privacy policy, Play/App Store listings, fastlane
metadata, docs, README, CONTRIBUTING, eas.json, and in-app mailto links.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The CUA script runs on the CI runner (host), not inside the emulator.
10.0.2.2 is the emulator's special address for the host — it is only
reachable FROM INSIDE the emulator. Calling it from the runner always
fails, so _precreate_test_session returned None, and the session_list
phase fell back to the weak 'screen visible' assertion instead of the
strong 'pre-created session must appear' check.
Fix: replace 10.0.2.2 → 127.0.0.1 before making the pre-create call.
Localhost URLs (100.x.x.x, custom dev server) pass through unchanged.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
sessions_reload regression phase now runs BEFORE the TypeScript task.
Previously it was gated behind typescript which fails in CI when the model
is unavailable — meaning the actual sessions regression check never ran.
Phase order is now:
connect → session_list (pre-created session required) → new_session
→ sessions_reload (navigate back, list must be non-empty) [CRITICAL]
→ typescript (informational) → verify (informational) → settings
Critical phases: connect, session_list, new_session, sessions_reload.
TypeScript/verify/settings are informational (model availability varies).
CI emulator fixes:
- api-level: 30 → 28 (more stable, boots reliably on ubuntu-latest)
- target: google_apis → default (lighter, no Play Services needed for
sessions regression test, avoids known boot issues with google_apis)
- disable-animations: true (reduces boot overhead)
- emulator-boot-timeout: 600 (explicit, matches action default)
- Switch from --scenarios to --showcase (runs the new structured flow
with _precreate_test_session + sessions_reload phase)
- Clear app state before install (pm clear) for deterministic first-run
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Previously the session_list phase goal said 'The session list may be empty
(no sessions yet) — that is fine' and 'Report done when you can see the
session list screen (even if empty)'. This means an empty sessions list
was treated as a test PASS, so every previous 'fix' was validated against
a test that cannot detect the regression.
Two changes:
1. _precreate_test_session(): calls POST /session via HTTP before the CUA
starts. The session_list phase goal now explicitly names this session and
requires it to be visible — if the app fails to load server sessions,
the phase fails (not passes silently with an empty list).
Falls back gracefully if the server is unreachable at pre-create time.
2. sessions_reload phase (new, critical): after completing the TypeScript
task the test navigates back to the Sessions tab and asserts the list is
non-empty. This catches the other variant of the regression — sessions
vanishing after navigating away from a session and back.
Both phases are now in the critical list, so a failure in either causes the
overall test to report partial/fail instead of success.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>