* tools(sentry): add org-wide volume report and fix the metric we measure on
The AGE-105 gate is a measured number, so it needs a repeatable query. It also
needed a correction: `accepted` is the wrong headline. The org is over its error
quota, so Sentry rejects nearly everything and `accepted` reads ~0 for every
project - a blown org and a fixed one look identical on that column. The demand
metric is `submitted` = accepted + rate_limited.
scripts/sentry-volume-report.mjs takes named --window ranges and prints
per-project submitted / accepted / rate_limited / client_discard plus the
per-hour and projected per-month rate, so before/after comparisons run the exact
same query instead of being re-derived by hand each time.
Records the pre-rollout baseline in docs/analytics.md: opencode-mobile at
4.71/h (3,441/mo), 87% of the org's post-box-bot demand, from two windows that
agree to within 0.2%.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
* test(sentry): pin the noise gate against 90d of real production events
The gate's unit tests prove it behaves as specified. Nothing proved the spec
was aimed at the right targets. Replaying the actual 90d census of the
opencode-mobile Sentry project (648 events, 11 issues) through the gate's own
precedence shows 96.9% hard-dropped as transport noise and every observed crash
class (OOM, ANR, IllegalStateException) still allowlisted -> ~87 events/month
against a 1,500/month target.
Also records two findings from measuring the org directly:
* The error quota resets on the 4th. The 5,000-event month opened 2026-08-04
and was spent by 08-08; the org has accepted zero errors since. 2026-09-04 is
the date the gate has to hold by, and it is why 'submitted' is the metric.
* Server-side levers are unavailable on this plan. A per-key rate limit PUT
returns HTTP 200 and silently discards the value (verified for three window
sizes), custom inbound filters are absent, spike protection 403s. The client
gate is the only control that exists, so its coverage is the whole margin.
Refs AGE-105
Co-Authored-By: Paperclip <noreply@paperclip.ing>
* tools(sentry): split client_discard by reason so gate drops aren't confused with quota backoff
Raw client_discard cannot show whether the noise gate works. Today 100% of
opencode-mobile's client_discard is ratelimit_backoff -- the SDK backing off a
429 because the ORG is over quota -- which rises when things get WORSE. Gate
drops land in a different reason: @sentry/core records before_send when
beforeSend returns null.
- stats_v2 now groups by reason as well as project/outcome
- the before_send vs ratelimit_backoff split always prints; --by-reason adds
the full per-project reason table
- before_send > 0 is install-share-independent, so it proves the gate is live
on real devices days before a monthly rate can bend
- documents that release-level segmentation is impossible while over quota:
rate_limited events are never stored, so release tags stop (last value
0.4.12, 2026-08-08). Version share comes from Play, not Sentry.
* ci(sentry): block a Play release whose bundle lost the noise gate
The AGE-105 quota fix is entirely client-side (every server-side lever on
this plan is dead), so the gate being *in the shipped binary* is the whole
safety margin. That is also the one thing Sentry cannot tell us: while the
org is over quota nothing is stored, release tags stop dead at 0.4.12, and
a release:0.4.14 query returns empty in a way that reads like success.
Grep the Hermes bundle inside the AAB instead, before the Play upload step:
the gate's reason codes, the transport drop-list regex, the
noise.dropped_since_last tag only applyNoiseGate() writes, and a baked-in
DSN (a release built without EXPO_PUBLIC_SENTRY_DSN makes Sentry a silent
no-op). Verified to discriminate on real artifacts - the v0.4.14 build now
on Play production passes, pre-gate v0.4.13 fails all six markers.
Also records the rejected alternative: persisting gate state across cold
starts pays off only under ~94 active devices (2,633 session envelopes/7d
vs a 6h cooldown), and the install base is above that.
---------
Co-authored-by: engineer <engineer@macbookpro.lan>
Co-authored-by: Paperclip <noreply@paperclip.ing>
The publish workflow overwrites versionCode with github.run_number + 100, so
the gradle number (40) is not what Play reports. Production dispatch run 49 ->
versionCode 149. Without this row, play-version-share.mjs reports the new
release as an unknown versionCode and the AGE-100 after-number cannot be split
by build.
Co-authored-by: engineer <engineer@macbookpro.lan>
Co-authored-by: Paperclip <noreply@paperclip.ing>
AGE-61 asked for a number: what share of installs is still on a build older
than v0.4.8, where the ONLY waitlist path is a mailto: to a human inbox that
nothing reconciles back into the store (20 of 21 signups lost, 2026-08-03 →
2026-08-13). Answering it by hand is how it stays unanswered next quarter, so
this is a script, not a screenshot.
- scripts/play-version-share.mjs: Play Developer Reporting API
(crashRateMetricSet -> distinctUsers by versionCode) for the auto-updating
channel, plus --github for lifetime release-APK downloads per tag, which is
the only per-version signal the sideload channel emits. Mints its own token
from the service account we already ship to CI; no new deps, no new secret.
Play versionCodes are run_number+100, NOT the gradle ones — the mapping is
derived from the publish runs and documented inline (139 = v0.4.8).
- distribution/waitlist-signup-path-coverage.md: the measured answer.
The answer, 2026-08-14: Play is 0% stale (single reported versionCode 142 =
v0.4.10, ~90-100 daily users); the sideload channel is 25.9% stale (436 of
1682 lifetime APK downloads predate v0.4.8) and can never auto-update. There
is no iOS listing and no IzzyOnDroid presence, so nothing else contributes.
Two consequences worth stating plainly: the hourly mailto reconciler is
permanent infrastructure, not a stopgap; and shipping updates does NOT close
the leak, because on current builds the fallback still fires on timeout/5xx/
offline (src/lib/waitlist.ts:shouldFallbackToMailto).
Co-authored-by: engineer <engineer@macbookpro.lan>
Co-authored-by: Paperclip <noreply@paperclip.ing>
* fix(ci): classify unresolved workflow failures
Replaces rolling historical failure escalation with active consecutive streaks and verifies Sentry zero/unavailable states.\n\nPlan: https://github.com/dzianisv/opencode-mobile/issues/139#issuecomment-5044502048
* fix(ci): scope failure streaks to default branch
Prevents pull-request failures from becoming production health signals for issue #139.
---------
Co-authored-by: engineer <engineer@gray-knight-m1.local>
Reports the activation + demo funnels the app instruments (app_opened ->
connection_succeeded, demo_started -> demo_completed -> demo_exited_to_connect)
and the money metric: of users who started the offline demo, how many later
reached a working connection. Read-only HogQL.
Needs a PostHog PERSONAL API key (read scope) — the app only ships the
write-only ingest key, so funnel data can't be read without one. This makes
that the single remaining unblock: set POSTHOG_PERSONAL_API_KEY +
POSTHOG_PROJECT_ID and run it. Prints setup instructions if unset.
Claude-Session: https://claude.ai/code/session_01T12AhSnQVrSxNnvwfCx2z6
Co-authored-by: engineer <engineer@macbookpro.lan>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
There is no auto-deploy for the docs site — gh-pages is updated manually,
which is why fixes (e.g. the opencode->opencode-ai guide correction in #113)
land on main but not on the live site. This ad-hoc process also risks wiping
the live F-Droid repo (gh-pages/fdroid/) since docs-site/ doesn't contain it.
scripts/deploy-docs.sh copies docs-site/ over gh-pages ADDITIVELY (never
--delete) and hard-aborts if fdroid/, privacy/, or .nojekyll would go missing
or the F-Droid repo index is empty. Supports --dry-run. A dry-run against the
current site shows it would ship exactly the pending changes (the guide
install-command fix + demo.gif/mp4 + updated screenshots) and touch nothing
under fdroid/.
Usage: bash scripts/deploy-docs.sh [--dry-run]
Claude-Session: https://claude.ai/code/session_01T12AhSnQVrSxNnvwfCx2z6
Co-authored-by: engineer <engineer@macbookpro.lan>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Installers with no self-hosted opencode server hit a dead end at the
empty Sessions state, contributing to ~0% 7-day retention. Adds a
fully offline, scripted /demo route reusing the real chat components
(MessageBubble, ToolCallCard/DiffView, PermissionPrompt) so new users
can see what opencode does before connecting anything, then funnels
them to Connect / the setup guide.
- src/lib/demo-script.ts: pure, hardcoded Message/Part fixture builder
(no RN/store/network imports) — the isolation guarantee.
- app/demo.tsx: new /demo route rendering the scripted conversation
via useMemo'd local state only; permission reply is local setState,
never sessionClient.permission.reply().
- app/(tabs)/index.tsx: "Try a demo" button added to the no-connection
empty state, placed after the existing add-connection-button so its
position/testID for existing Maestro flows is unchanged.
- .maestro/flows/demo.yaml: new E2E flow covering the empty-state CTA
through conversation, diff expand, permission approve, and the CTA
reaching the real connect form.
- scripts/run-e2e-flows.sh: registers demo in NEWER_FLOWS (non-blocking)
so it actually runs in CI.
- i18n: new sessionsList.empty.tryDemoButton and demo.* keys added to
both en.json and zh-Hans.json (catalog-parity verified).
npm run typecheck: clean. npm test: 175/175 passing.
Co-authored-by: engineer <engineer@macbookpro.lan>
* fix(e2e): stop asserting on SSE reply in activation-positive — CI-harness limitation, not a product bug (closes#90)
Extensive investigation (see PR #102 for the full trail) into "the
positive flow's assistant reply never renders" tried four independent
SSE client transports in src/lib/sdk.ts global.events(): the
already-shipped expo/fetch ReadableStream reader, a hand-rolled
XMLHttpRequest reader, react-native-sse, and react-native-fetch-api's
`reactNative: { textStreaming: true }`. Every one delivers exactly one
chunk right after connecting to the mock server and then nothing until
the connection closes, regardless of API choice or frame size (a ~4KB
padding experiment ruled out a buffer-size threshold).
A raw-socket probe (a plain BSD-sockets client with zero React Native
involvement, run via `adb shell` through the identical adb-reverse
tunnel the app uses) streamed every heartbeat from the mock server
incrementally in real time over the same connection. That rules out
adb-reverse and the mock's flush behavior and isolates the stall to
React Native Android's OkHttp-backed networking layer buffering a
long-lived streaming HTTP response in this specific Android-emulator +
Node-mock + adb-reverse combination — not a defect in any particular
client library.
There's no evidence this reproduces against a real opencode server on
a real device/network: issue #76's 65 affected users prove real SSE
connections stream live agent output in production (the bug they hit
was the 401-retry storm, not a missing reply). Since expo/fetch is the
already-shipped, production-proven transport and none of the
alternatives showed any advantage in this harness, the transport stays
unchanged.
What changes instead: .maestro/flows/activation-positive.yaml no
longer waits on the SSE-streamed reply, since asserting on it here
would assert on a CI-harness limitation, not real app behavior. It now
verifies everything reliably observable — consent, connect, session
creation, and the optimistic local echo of the sent message — and
activation-e2e.yml's `continue-on-error: true` (added because this
suite had never passed) comes off, so it blocks PRs on regressions in
what it does cover.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(e2e): repair stale "401" assertion in activation-negative-401 (refs #90)
Removing activation-e2e.yml's continue-on-error surfaced a second, unrelated
stale assertion once the suite was actually enforcing again: the negative
flow's connect-time-401 case asserts a literal "401" that PR #79 (401/403
auth-stop handling) and #103 (i18n) apparently moved out of what's
rendered — "Connection Failed" still passes, "401" now fails.
The alert body interpolates two pieces: probeConnection()'s summary (which
turns out to be misclassified as "connection actually works now" for this
case — diagnostics-classify.ts's `health.ok` only reflects "fetch() didn't
throw", not HTTP status, a separate real bug, out of scope for this PR) and
testConnection()'s caught error message, which is sdk.ts's
apiErrorFor(401, ...) text and always contains the mock's
`{"error":"Unauthorized",...}` body per src/lib/api-error.test.ts. Swapped
the assertion to "Unauthorized" and added a temporary console.log of both
pieces in app/connection/add.tsx to confirm exactly what renders from CI
logcat (removed once confirmed).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(e2e): assert alert action buttons, not body text — native AlertDialog body isn't in Maestro's a11y tree (refs #90)
The diagnostic added last commit confirmed the Connection Failed alert's
body DOES contain the real error ("API Error: 401 -
{\"error\":\"Unauthorized\",...}", via logcat: '[connect] failure alert
content' logged both the (separately buggy, out-of-scope)
probeConnection summary and the correct testConnection error text). Yet
both "401" and "Unauthorized" assertions still failed against the same
on-screen alert. That means Maestro's accessibility-tree text matching
on this Android AlertDialog only sees the title, not the message body —
so no substring of the body was ever going to match.
Switched to asserting what's actually reachable: the title "Connection
Failed" (unchanged, already passing) plus both action button labels,
"OK" and "Share report" (src/lib/i18n/en.json common.ok /
common.shareReport). That still proves the test's real intent — a
visible, actionable error with a dismiss and a share-report path, never
a silent failure (issue #76) — using strings actually present in the
accessibility tree instead of guessing at unreachable body text.
Removes the temporary console.log diagnostic from
app/connection/add.tsx now that its purpose (confirming exactly what
renders) is done.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(e2e): split activation-e2e into blocking core + non-blocking newer flows (refs #90, refs #104)
With this PR's fixes, activation-positive and activation-negative-401
(the coverage issue #90 actually scoped) now run green — but removing
activation-e2e.yml's continue-on-error surfaced that four flows added
after the initial suite (#82's directory-picker/all-sessions/
variant-picker, #101's diff-scroll) have never once run to completion
in CI: they always sat behind whichever activation flow failed first,
so they were merged and have run unverified against the current
UI/mock this whole time. directory-picker fails immediately at
`id: directory-row-frontend`; the other three are untriaged.
Fixing four separate, previously-never-green UI surfaces is out of
#90's scope and unbounded in this PR. scripts/run-e2e-flows.sh now
splits the flow list into CORE_FLOWS (the two #90 covers — blocking,
fails the job on a regression) and NEWER_FLOWS (the four newer ones —
always run, each one's pass/fail reported via echo/::warning::, but
never fails the job). This lets activation-e2e.yml enforce the
activation coverage that's now verified, without either leaving it red
forever or spending unbounded time inside this PR chasing four
unrelated UI surfaces.
Filed #104 to track hardening each NEWER flow and moving it back into
CORE_FLOWS once confirmed green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Turns the manual QA ask ("verify DiffView + CodeBlock horizontal-scroll
on-device with a populated diff") into two automated layers:
1. Unit (deterministic, runs in `npm test` now): extracted the shared
ScrollView props into src/lib/scroll-config.ts (WIDE_CONTENT_SCROLL_CONFIG)
so DiffView.tsx and CodeBlock.tsx spread the SAME plain object their tests
assert on — no react-native-renderer needed. Added a source-scan
regression test (wide-content-scroll.regression.test.ts) that fails if
either component loses its ScrollView wiring or reintroduces
numberOfLines truncation.
2. E2E (Maestro): .maestro/flows/diff-scroll.yaml opens a session with a
pre-seeded wide edit-diff tool call and a wide fenced code block, then
swipes each horizontal ScrollView left and asserts the off-screen marker
text becomes visible. mock-opencode-server.ts gained a --seed-diff mode
that serves this session via GET /session/:id/message (pre-existing
history), not SSE — issue #90 (a separate SSE-render bug) is being fixed
independently, and this flow must not depend on it landing first. Wired
the new flow + port 4100 into run-e2e-flows.sh and activation-e2e.yml.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Adds diagnostics-only instrumentation to attribute the Activation E2E
positive flow's "SSE reply never renders" failure (issue #90 mode B)
between (a) expo/fetch not streaming the SSE response on the Android
release APK, vs (b) a mock-side broadcast bug. Does not change app
behavior or fix the root cause — #90 stays open pending the next CI
run's enriched logs.
- tests/fixtures/mock-opencode-server.ts: per-request logging
(method/path/status), per-SSE-connection connect/disconnect logging
with live client count, per-broadcast event-type + client-count
logging, and a 2s SSE heartbeat comment so client-side silence
becomes unambiguous.
- src/lib/sdk.ts global.events(): logs on the first successful
reader.read() that returns data, and when the stream loop ends —
proves/disproves whether expo/fetch ever delivers a byte.
- scripts/run-e2e-flows.sh: the emulator->mock reachability probe used
toybox wget/nc, which don't work reliably on the API-28 image.
Replaced with a probe chain (curl, wget, mksh /dev/tcp, nc, then a
host-side fallback) that writes a clear PASS/FAIL/UNKNOWN verdict to
artifacts/diag/probe.txt without blocking the flow.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* test: E2E coverage for directory picker, all-sessions, variant picker
Extend the Maestro suite for the features merged into main today:
DirectoryBrowserSheet's server-folder picker, the directory-less
all-sessions-across-projects list (+ the #46/#48 open-across-project
regression), and VariantPicker's reasoning-effort chip.
- tests/fixtures/mock-opencode-server.ts: GET /file (directory-scoped via
the x-opencode-directory header) with a small fake tree, GET /project
for the "Server Projects" section, POST /session honoring the directory
header, GET /session/:id (needed to open a session from the all-sessions
list), GET /provider variants for VariantPicker, and an optional
--seed-sessions mode that pre-populates two sessions across two
directories. --fail-auth mode is untouched.
- .maestro/flows/directory-picker.yaml, all-sessions.yaml,
variant-picker.yaml: three new flows, run in the same emulator session
as the existing activation flows.
- Additive testIDs on DirectoryBrowserSheet, the "Browse Folders" row,
session list rows, the variant chip, and VariantPicker rows.
- .github/workflows/activation-e2e.yml: two more mock server instances
(4098 seeded, 4099 fresh) and three more maestro test steps.
Verified: tsc --noEmit clean, all 108 existing unit tests pass, every new
mock endpoint curled against its real shape read from the app code, YAML
validated. No Android emulator available locally to run the Maestro flows
themselves.
* test(mock): enforce per-directory session scoping so #46/#48 coverage can fail
Review finding (HIGH): GET /session/:id and /session/:id/message ignored
x-opencode-directory, so all-sessions.yaml could not fail if the directory
threading fix regressed. The mock now mirrors the real server's per-directory
workspace scoping:
- GET /session/:id and GET /session/:id/message 404 unless the request's
x-opencode-directory (or DEFAULT_DIRECTORY when absent) matches the stored
session's directory.
- GET /session without ?roots=true is scoped to the request's directory;
loadSessions()'s directory-less roots=true call still returns everything.
- Document the port-4099 shared-state coupling between directory-picker and
variant-picker flows, and why all-sessions.yaml now has teeth (flow comment).
Curl-verified: correct header 200, wrong/no header 404, scoped vs roots
listing, create-then-open paths for all three flows, --fail-auth untouched.
tsc clean, 108/108 unit tests pass.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NJKAQ6HAikWGQK7PGZ5Y4E
* fix(e2e): widen connect-handshake wait past client's own 30s timeout
Run 29546383612 (7cbd3a6, first real emulator execution of these flows)
failed on activation-positive.yaml: "Assert that id: connection-status-dot
is visible" timed out after the flow's 20s extendedWaitUntil, right after
tapOn connect-submit-button.
The mock server itself is fast (verified locally: health + project/current
+ path respond in ~30ms total), so this isn't a mock fidelity gap. But
Quick Connect's testConnection()/addConnection() path chains up to 3
fetches (health, then project.current + path.get in parallel), and each
individual fetch is capped by src/lib/sdk.ts REQUEST_TIMEOUT_MS = 30_000 —
strictly longer than the 20s the flow was willing to wait. A first-attempt
emulator-to-host (10.0.2.2) connection that's merely slow to establish,
rather than outright failing, would blow past the test's wait before the
app's own client-side timeout even fires.
Bump the connect -> connection-status-dot / "Connection Failed" waits from
20000 to 40000 across all 5 flows that share this pattern
(activation-positive, activation-negative-401, all-sessions,
directory-picker, variant-picker) so the wait is never shorter than the
code path it's gating on. Assertions are unchanged — still requires the
real dot / real error text, just with a timeout that isn't racing the
client.
Verified locally: typecheck clean, all 108 unit tests pass, YAML parses,
mock server confirmed fast under direct curl. Emulator behavior itself
(whether 40s consistently clears it) is unverified until the next CI run.
* fix(e2e): use adb reverse + 127.0.0.1 instead of 10.0.2.2; capture logcat/maestro debug
Root cause of the activation-e2e failure (connect step timed out, ~0 requests
reaching the mock): the 10.0.2.2 host alias is unreliable under the headless
emulator-runner — the app's http://10.0.2.2:4096/global/health never completed,
so connection-status-dot never rendered.
- run-e2e-flows.sh: single script (fixes cd-per-line fragility) that adb-reverses
each mock port (4096-4099) into the emulator's localhost, runs every flow with
--debug-output, and dumps logcat on exit.
- All flows now connect to 127.0.0.1:<port> (the adb reverse target).
- Upload maestro-debug (UI hierarchy on failure) + logcat as artifacts so future
failures are diagnosable instead of blind.
* fix(e2e): connect via 127.0.0.1:PORT in IP field, stop typing into port input
Root cause of every activation-e2e connect failure (proven by the app's own
logcat diagnostic: '[diag] probe start http://127.0.0.1:40966 ... server
unreachable'): the port field defaults to useState("4096"), and the flow's
eraseText + inputText "4096" raced the controlled number-pad input, leaving
"40966" — nothing listens there, so connect always failed. This was never a
10.0.2.2 / adb reverse issue.
Fix: buildUrl already extracts host:port from the IP field, so enter
127.0.0.1:<port> there and remove the flaky port-field steps entirely.
pastedPort overrides the default port state, so each flow's port is
deterministic (4096 positive / 4097 negative / 4098 all-sessions / 4099
directory+variant).
---------
Co-authored-by: engineer <engineer@gray-knight-m1.local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Port the harness fixes from test/e2e-new-features to main so the
Activation E2E workflow stops failing before any flow executes:
- Run everything through scripts/run-e2e-flows.sh as a single script
invocation: android-emulator-runner executes each 'script:' line in
its own shell, so the previous 'cd artifacts/screenshots' never
persisted and maestro failed with 'Flow path does not exist' on
every run (19/19 red since the workflow landed).
- adb reverse + 127.0.0.1 instead of 10.0.2.2 (unreliable headless),
emulator->mock reachability probe, logcat + maestro debug capture.
- Trim the flow list to the two flows that exist on main; the three
newer flows land with the test/e2e-new-features PR.
- continue-on-error until the suite's first green: the positive flow
still fails its final reply assertion (mode B in #90), and a
never-green suite should not block unrelated PRs or pollute the
product-intelligence failure metrics (#89).
Refs #90. Refs #89.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* fix(metrics): repair review triage — correct secret wiring, privacy-safe aggregated issues
- triage-reviews.yml read secrets.GOOGLE_SERVICE_ACCOUNT_JSON, which doesn't
exist; map the real PLAY_STORE_SERVICE_ACCOUNT_JSON secret onto the env var
the script expects.
- triage-reviews.py rewritten to maintain a single sanitized, deduped
"Play Store Review Triage" issue instead of one public issue per review.
The old version leaked reviewer full names and verbatim review text into
public GitHub issues and spammed the tracker. The new version aggregates
actionable (<=3 star) reviews into one issue with rating counts, a
word-frequency theme summary (no quoted sentences), and opaque review_id
references for Play Console lookup. An embedded HTML comment marker
(matching the product-intelligence.mjs pattern) holds the current
actionable review_id set so runs update in place and skip entirely when
nothing changed.
- product-intelligence.yml referenced the nonexistent
SENTRY_PRODUCT_INTELLIGENCE_TOKEN secret, causing the daily cron to fail
silently (#60). Fall back to SENTRY_AUTH_TOKEN when the dedicated
read-only token isn't configured.
- docs/playstore.md: document that Play Console is still the only trusted
source for acquisition/uninstall metrics (product-intelligence.mjs defers
this), and that review-based signals are sourced via the Android
Publisher API through PLAY_STORE_SERVICE_ACCOUNT_JSON.
Closes#61. Refs #60.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NJKAQ6HAikWGQK7PGZ5Y4E
* fix(triage): fail visibly when GOOGLE_SERVICE_ACCOUNT_JSON is missing
Review finding on PR #78: env_client() exited 0 on missing credentials,
so the scheduled workflow would report success while silently doing
nothing — contradicting issue #61's 'missing credentials fail visibly'
done-criteria.
---------
Co-authored-by: engineer <engineer@gray-knight-m1.local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
scripts/triage-reviews.py was fully written but had no workflow, so it
never ran. Add a daily 07:00 UTC cron (staggered after product-intelligence)
plus workflow_dispatch, with Python 3.12 + the Android Publisher API client
deps the script imports, and GOOGLE_SERVICE_ACCOUNT_JSON / GH_TOKEN passed
through as named secrets.
Also fix a stale doc-string reference: the issue body linked to a
non-existent monitor-reviews.yml; point it at the workflow actually created.
Adds privacy-safe aggregate product intelligence, reviewed/versioned website assets, and a dispatch-only rollout until the dedicated Sentry token is verified. Independent review blockers were fixed in 8bc47e4; app checks, website production build, Android CI, and iOS CI are green.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- docs-site/index.html: add Vercel Analytics + Speed Insights CDN scripts
(only fires on opencode.agentlabs.cc served via Vercel, not GitHub Pages)
- scripts/triage-reviews.py: fetch recent Play Store reviews via Android
Publisher API, create GitHub issues for ≤3★ reviews not yet tracked
Run review triage manually on VM:
DAYS_BACK=7 GOOGLE_SERVICE_ACCOUNT_JSON=... GH_TOKEN=... python3 scripts/triage-reviews.py
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- Change permission notification title from `req.permission || 'Permission requested'`
to the user-friendly 'Agent needs approval'; permission type + patterns now appear
in the body (e.g. 'bash: echo hello') for context.
- Add `dedupeKey: `perm-${req.id}`` and `dedupeKey: `question-${req.id}``
(60 s cooldown) to both events so a SSE reconnect after disconnect() clears state
can't fire a second notification for the same pending request.
- Fix stale CUA-test comment that claimed 'Agent needs approval' did not exist;
fallback assertion already matched correct title; update the comment to reflect
the real events.ts behavior.
Closes#39
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Shows amber 'Reconnecting… (attempt N)' banner when SSE is down.
Shows brief green 'Connected ✓' flash on reconnect (useRef transition
to avoid atomic state reset bug where lastDisconnectAt resets with
reconnectAttempts in the same set() call).
Banner disappears automatically when SSE is stable.
Updates CUA scenario to check for both ASCII and Unicode ellipsis.
Closes#42
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
#39 (backgrounded permission notification):
- Event is permission.asked, not permission.requested (events.ts already calls
notify() for it; send() only fires while backgrounded).
- Assert on APP_PACKAGE (the only token guaranteed in every dumpsys record) plus
the actual copy ('Permission requested' / 'A tool needs your approval') instead
of the non-existent 'Agent needs approval' string.
#42 (SSE disconnect banner):
- reconnectAttempts only zeroes after STABLE_CONNECTION_MS (10s) past a healthy
reconnect, and the pending backoff timer can take up to 15s — so the banner can
linger ~25s. Poll up to 40s for dismissal instead of a fixed 15s sleep to avoid
a false 'still showing' failure.
Refs #39#42
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Adds three deterministic helpers that use ADB instead of LLM vision,
so pass/fail cannot be hallucinated:
check_ui_text(text) — uiautomator XML dump + grep
check_notification_drawer(text, timeout) — dumpsys notification poll
simulate_network_drop() / restore_network() — svc wifi/data disable
Plus two new feature-test scenarios with deterministic gating:
sse_disconnect_banner (#42)
- ADB cuts WiFi+data, waits, checks UI XML for 'Reconnecting' text
- LLM visual check is supplementary/informational only
- ADB restores network, checks banner disappears
backgrounded_permission_notification (#39)
- ADB backgrounds app (Home key)
- API sends permission-triggering message
- adb dumpsys notification checked for 'Agent needs approval'
- No LLM involved in pass/fail decision
Also adds background_app() / foreground_app() helpers.
Wires new scenarios into --scenarios catalog alongside LLM ones.
Updates --scenarios help text to document both types.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
connect phase: now verifies the connection appears in the list after saving,
not just that the form was dismissed. Prevents false PASS when the LLM
declares connect done before the entry is actually visible.
_precreate_test_session: when external URL (Tailscale) times out from the
CI runner, fall back to localhost:4096 (the runner-local opencode serve).
This ensures the named-session assertion stays active in standard CI runs
while also working when dispatched against a live external server.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The CUA script runs on the CI runner (host), not inside the emulator.
10.0.2.2 is the emulator's special address for the host — it is only
reachable FROM INSIDE the emulator. Calling it from the runner always
fails, so _precreate_test_session returned None, and the session_list
phase fell back to the weak 'screen visible' assertion instead of the
strong 'pre-created session must appear' check.
Fix: replace 10.0.2.2 → 127.0.0.1 before making the pre-create call.
Localhost URLs (100.x.x.x, custom dev server) pass through unchanged.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
sessions_reload regression phase now runs BEFORE the TypeScript task.
Previously it was gated behind typescript which fails in CI when the model
is unavailable — meaning the actual sessions regression check never ran.
Phase order is now:
connect → session_list (pre-created session required) → new_session
→ sessions_reload (navigate back, list must be non-empty) [CRITICAL]
→ typescript (informational) → verify (informational) → settings
Critical phases: connect, session_list, new_session, sessions_reload.
TypeScript/verify/settings are informational (model availability varies).
CI emulator fixes:
- api-level: 30 → 28 (more stable, boots reliably on ubuntu-latest)
- target: google_apis → default (lighter, no Play Services needed for
sessions regression test, avoids known boot issues with google_apis)
- disable-animations: true (reduces boot overhead)
- emulator-boot-timeout: 600 (explicit, matches action default)
- Switch from --scenarios to --showcase (runs the new structured flow
with _precreate_test_session + sessions_reload phase)
- Clear app state before install (pm clear) for deterministic first-run
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Previously the session_list phase goal said 'The session list may be empty
(no sessions yet) — that is fine' and 'Report done when you can see the
session list screen (even if empty)'. This means an empty sessions list
was treated as a test PASS, so every previous 'fix' was validated against
a test that cannot detect the regression.
Two changes:
1. _precreate_test_session(): calls POST /session via HTTP before the CUA
starts. The session_list phase goal now explicitly names this session and
requires it to be visible — if the app fails to load server sessions,
the phase fails (not passes silently with an empty list).
Falls back gracefully if the server is unreachable at pre-create time.
2. sessions_reload phase (new, critical): after completing the TypeScript
task the test navigates back to the Sessions tab and asserts the list is
non-empty. This catches the other variant of the regression — sessions
vanishing after navigating away from a session and back.
Both phases are now in the critical list, so a failure in either causes the
overall test to report partial/fail instead of success.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
ADB input text with %s escaping didn't trigger React Native onChangeText.
New approach: write text to /sdcard/ file on device, then use shell
command-substitution "$(cat ...)" to pass raw text to input text.
This preserves spaces without %s conversion and avoids quote-escaping
issues that prevented React Native from detecting text changes.
Rewrites android-cua-smoke.py to demonstrate the complete first-run journey
instead of the previous "ping" smoke test. The new structured multi-phase
flow covers: server connection setup, session list, new session creation,
TypeScript hello-world task submission (watching tool calls/file writes),
output verification, and Settings/model-selection screenshot.
Key changes:
- run_onboarding_showcase() orchestrates 6 sequential CUA phases with
per-phase goals, step budgets, and PASS/FAIL phase tracking
- run_cua_step() replaces run_cua() — accepts step_label, action_delay,
saves labeled screenshots (/tmp/cua_<phase>_<step>.png) for debugging
- Global --speed-multiplier flag scales all _sleep() calls (0.5 = 2x faster)
- Showcase is now the default mode; legacy --goal / --scenarios flags retained
for backwards compat and CI regression scenarios
- Tighter action_delay (0.7s) and trimmed history window (14 turns) vs
previous 1.0s / 12 turns
- Phase banner log lines ("STEP N: ...") narrate the video in real time
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01No3k1AEioE4PNUZg12TxQo
* fix(ci): bump opencode Azure apiVersion to support /responses endpoint (#22)
The CUA smoke probe was returning MODEL_CAPABLE=false because opencode's
@ai-sdk/azure provider got 'API version not supported' from Azure on
/openai/v1/responses with apiVersion=2024-08-01-preview. Split into two
envs: keep the CUA driver on 2024-08-01-preview (chat-completions only)
and bump the opencode-side provider config to 2025-04-01-preview, which
supports the new responses API.
Effect: send_message/multi_turn scenarios get included again in CUA smoke
when the probe succeeds.
* ci(cua): bound curl timeouts + diag dump on server-start hang (#22)
Step 11 'Start opencode server' has hung past the 45-min job timeout in
two consecutive runs (27195348071, 27198238465). Local boot of
opencode-ai 1.16.2 with the same heredoc config is healthy in 3s, so
something is wrong specifically on the GH-hosted runner — likely curl
post-loop waiting indefinitely on an unresponsive server.
Adds:
- set -x for command tracing
- --connect-timeout 2 -m 5 on every curl so hangs cannot exceed 5s
- HEALTHY flag + explicit exit 1 (drops the unbounded post-loop curl)
- Periodic dump every 10s: server log tail, ss listening sockets,
process liveness — so we can see WHY the server isn't replying
Pure diagnostics; no behaviour change for the green path.
* fix(ci): use api-version=preview for /openai/v1/responses (#22)
Reproduced the probe failure locally against the same Azure resource:
all date-based api-versions (2024-08-01-preview, 2024-12-01-preview,
2025-01-01-preview, 2025-03-01-preview, 2025-04-01-preview) return:
{"error":{"code":"BadRequest","message":"API version not supported"}}
Only api-version=preview and api-version=v1 succeed (200). This is the
new Azure OpenAI v1 responses-API style; date strings are reserved for
the legacy /openai/deployments/{model}/chat/completions endpoint.
@ai-sdk/azure 3.x already defaults apiVersion to "preview" (per the
type definition: "Custom api version to use. Defaults to `preview`."),
so this aligns the workflow with the SDK default. Probe should now
return MODEL_CAPABLE=true and the send_message scenario will run.
* test(cua): extend send_message and multi_turn waits to 30s
Assistant bubbles can take 15+ seconds to appear after send. Previous
5-second wait was too short and caused false failures even when API
calls succeeded. Re-check screenshots periodically up to 30s total.
* fix(cua): screen-relative send button threshold for #22
The send action's auto-locate filtered for y1 > 2200 and fell back to
hardcoded (996, 2358) — both assume a 1080x2400 panel. The CI emulator
(API 30 google_apis pixel profile) is 1080x1920, so:
- the bottom_buttons filter never matched any clickable element
- the fallback tap landed off-screen
→ 'ping' message never sent, scenario timed out with no bubbles.
Switch to a screen-relative threshold (bottom 25%) and a fallback that
uses get_screen_size() to land in the bottom-right corner regardless of
device resolution. This was masked until now because send_message was
gated by MODEL_CAPABLE=false in earlier CI runs.
Refs: #22
---------
Co-authored-by: dzianisv <dzianis.varabyou@gmail.com>
The CI opencode server had NO LLM provider configured — the server log only
showed "listening", never a model. opencode-ai (released npm pkg) does not read
AZURE_OPENAI_* for its own LLM; it needs an explicit provider in opencode.json +
a default `model`. So send_message/multi_turn could never pass and the gate was
stuck on --only-connect-scenario (UI journey minus the model reply).
- Wire opencode to the same Azure resource the CUA driver uses via a generated
~/.config/opencode/opencode.json (@ai-sdk/azure provider, resourceName derived
from the endpoint secret at runtime, apiKey from env, default model azure/gpt-5.4).
- Add a deterministic REST probe step: create a session + send a prompt and check
for an assistant reply BEFORE the ~30min emulator run, exporting MODEL_CAPABLE.
- Add --scenarios to android-cua-smoke.py to run an explicit named set.
- Emulator step now runs connect_and_verify_sessions + send_message +
verify_session_list when MODEL_CAPABLE=true; falls back to the UI-only journey
(connect + verify_session_list) otherwise, logging the environmental reason.
- Raise --max-steps to 40 so multiple scenarios fit.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: repoint OpenCode links to agentlabs.cc/opencode
agentlabs.cc/opencode and /opencode/privacy are now live (200). Repoint
README, distribution listings (Play/App Store/F-Droid/IzzyOnDroid/iOS),
docs, and in-app privacy links (settings + telemetry consent) from
www.vibebrowser.app/opencode to the canonical agentlabs.cc hub.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(ci): run local opencode server for CUA smoke true-E2E (#15)
GitHub-hosted runners can't reach the Tailscale dev server
(100.108.64.76:4096), so the CUA smoke always failed at session creation.
- Install opencode-ai and run `opencode serve` on the runner host; the
Android emulator reaches it via 10.0.2.2. OPENCODE_URL now points there.
- Healthcheck /global/health before launching the app; dump server log on
failure for diagnosis.
- Add --only-connect-scenario to the smoke script and run just the
connect-and-verify-sessions path in CI: deterministic, needs no model
backend. The scenario now creates a session if the list is empty, so a
fresh server still yields a non-empty list.
This makes the smoke a true E2E and also exercises the #10 sessions-list
rendering path against a real server.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(ci): emulator smoke script is dash, not bash — drop brace-group healthcheck
android-emulator-runner runs the script: block under /usr/bin/sh (dash). The
multi-line `|| { ...; }` healthcheck was a dash syntax error (end of file
unexpected), failing the step before the smoke ran. Replace with a non-fatal
one-line re-check; the server was already health-gated in the prior step.
* docs(tasks): record smoke CI round 1 failure + dash fix
---------
Co-authored-by: engineer <engineer@opencode.ai>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Rename commit missed the Python launcher constant; HEAD still targeted
the old package so the smoke could not find the installed app.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(sessions): use active connection client directly, remove roots filter
Root cause A: loadSessions was calling clientForDirectory(serverHome) which
scoped the session list to /home/azureuser — a different project than the
server's active CWD. Sessions in the current project (e.g. opencode-mobile)
were never returned.
Root cause B: roots:true filtered out sessions that have a parentID (sub-task /
AUTO-REVIEW sessions), hiding valid sessions from the list.
Fix: use connState.client directly (the connection's active directory) and drop
the roots filter so all sessions for that project are visible.
Also adds a verify_session_list CUA smoke scenario that navigates back to the
sessions tab after creating a session and asserts the list is non-empty —
covering the regression path that was previously untested.
* fix(sessions): fetch serverHome in addConnection so loadSessions shows correct sessions
Root cause: addConnection() built the HTTP client but never fetched serverHome
(only loadConnections and setActiveConnection did). When the user adds a new
connection (fresh install / first sign-in), serverHome = null, so loadSessions
fell through to connState.client (the server's CWD). On this dev server the CWD
is the deploy directory — 11 old May-19 sessions that are not the user's recent
work sessions.
Fix: addConnection now fetches currentProject + serverHome via the same
Promise.all as setActiveConnection, before calling set(). This ensures
loadSessions immediately uses clientForDirectory(serverHome) → the global
project → the user's actual recent parent sessions.
Also adds --opencode-url flag to the CUA smoke script, which appends a
connect_and_verify_sessions scenario that reproduces the regression:
python scripts/android-cua-smoke.py --opencode-url http://100.108.64.76:4096
* fix(sessions): recover home scope after fresh connect
Resolve stale deploy-only session list by recovering server home during first load and keeping regression coverage in default Android CUA smoke and CI.
* chore(release): bump version to 0.4.0
- Record each scenario as MP4 via ADB screenrecord in a background thread
- Pull video to /tmp/cua_<scenario>.mp4 after test completes (always)
- Upload to ArchiveBox if ARCHIVEBOX_URL + ARCHIVEBOX_API_KEY env vars set;
gracefully skips when not configured (CI default)
- Upload all artifacts (PNG screenshots + MP4 videos) always, not only on failure
- Pass ARCHIVEBOX_URL/ARCHIVEBOX_API_KEY secrets to workflow (optional)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(security): fail closed on biometric init error
H-03: setting isAuthenticated: true on initialization failure was a
security bypass — any crash during biometric setup granted full access.
Fail closed instead; user sees auth prompt on next open.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(security): use Crypto.randomUUID for connection IDs
H-04: Math.random() is not cryptographically random. Connection IDs are
used as SecureStore key suffixes; switch to expo-crypto randomUUID for
a secure source.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(deps): pin expo-crypto to ~15.0.9
15.0.10 does not exist on npm; ~15.0.9 is the latest stable in the 15.x series compatible with Expo SDK 54.
* feat: add OpenCode Connect coming-soon waitlist card
Adds a discoverable 'OpenCode Connect — Coming Soon' card to the
add-connection quick-connect screen. Users can enter their email and
tap 'Join Waitlist' to send a pre-filled mailto. No backend required.
* fix(cua): detect actual screen dimensions and fix JSON parsing
- Get real screen size via `wm size` instead of hardcoding 1080x2400;
emulator is 1080x1920 so y-coordinates were systematically off
- Extract first JSON object via regex when model returns multiple objects
- Use AZURE_OPENAI_MODEL env var for deployment name (defaults gpt-5.4)
- Add AZURE_DEV_AI_* path for Azure AI Foundry endpoints
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(security): SHA-pin upload-google-play and sanitize notification bodies
M-02: Pin r0adkll/upload-google-play to commit SHA e738b9d (v1.1.5)
to prevent supply-chain hijack via tag mutation.
M-03: Sanitize all push notification bodies — strip control chars,
truncate to 200 chars. Prevents server-supplied strings (error messages,
file paths from permission patterns, session titles) from leaking
unbounded text into the OS notification drawer.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* feat(privacy): add telemetry consent gate for Sentry crash reporting
Sentry was always-on, violating F-Droid anti-feature policy and user
trust norms. Now gated behind explicit opt-in:
- First-launch consent modal (TelemetryConsentModal) shows once on
fresh install; user can Allow or Decline.
- Consent state persisted in expo-secure-store (survives restarts).
- Settings > Privacy section: crash reporting toggle + privacy policy link.
- initSentry() called only after consent granted — not on app start.
Closes#3 (partial)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* feat(config): add real icons and complete iOS/Android app.json config
- Add 1024×1024 app icon, 432×432 adaptive icon foreground, 200×200 splash
- iOS: push notification entitlement (aps-environment: production), speech/
microphone/camera/photo usage descriptions for future features, disable
ITSAppUsesNonExemptEncryption
- Android: adaptive icon with dark background (#0F172A), versionCode: 1
- expo-notifications plugin wired in app.json
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* feat(dist): add iOS CI workflow, README rewrite, CONTRIBUTING, and LICENSE
- publish-app-store.yml: EAS Build + TestFlight submission; runs on tag/release/
workflow_dispatch; bumps ios.buildNumber from github.run_number
- README: full rewrite — features, install badges, connection guide, contributing
- CONTRIBUTING.md: contribution guide for OSS contributors
- LICENSE: MIT
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs(dist): add store listings, strategy, privacy policy, F-Droid/IzzyOnDroid templates
- distribution/strategy.md: monetization strategy (free client + opencode Cloud)
- distribution/play-listing.md: Google Play store copy (name, description, tags)
- distribution/app-store-listing.md: App Store listing copy
- distribution/privacy-policy.{md,html}: GDPR-compliant privacy policy
- distribution/PLAY_CONSOLE_SETUP.md: Play Console setup runbook
- distribution/ios-enrollment-runbook.md: Apple Developer Program enrollment steps
- distribution/SIGNING-KEY-FINGERPRINTS.md: keystore fingerprint for reproducible builds
- distribution/fdroid-submission/: F-Droid metadata template
- distribution/izzyondroid-submission/: IzzyOnDroid submission template
- distribution/whatsnew/: Play Store release notes (en-US)
- distribution/whatsnew-ios/: TestFlight release notes
- distribution/play-graphics/: Play Store screenshot placeholders
- distribution/app-store-graphics/: App Store screenshot placeholders
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(telemetry): handle SecureStore failure + Android back button
- add .catch() on loadTelemetryConsent() so SecureStore rejection
shows the consent modal instead of blocking startup forever
- add onRequestClose={onDecline} to Modal so Android back button
records the decline rather than silently dismissing
- fix catch block in telemetry.ts to not clobber _resolved when
SecureStore read fails mid-session
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(ci): run gradlew clean to prevent stale modules.json duplicate
Sentry Gradle plugin writes modules.json to src/main/assets; cached
build intermediates contain an old copy → mergeReleaseAssets fails
with 'Duplicate resources'. Running clean before assembleRelease
clears the intermediate state.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(ci): remove android build output cache causing duplicate modules.json
Caching android/app/build/intermediates and android/app/.cxx causes
two issues:
1. Stale modules.json in intermediates → Duplicate resources error
2. .cxx CMake artifacts reference absolute paths → ninja clean fails
Keeping only Gradle distribution cache (~/.gradle) which is safe.
Expo prebuild regenerates android sources fresh each run anyway.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
- Added 'send' action type that uses uiautomator XML to find the rightmost
clickable element in the bottom input bar (the send button)
- Added screen resolution (1080x2400) to LLM context for better coordinate estimation
- Updated system prompt to instruct model to use 'send' action instead of manual tap
- Updated scenarios with clearer step-by-step instructions
- Both send_message and multi_turn scenarios pass reliably