Files
opencode-mobile/.github/workflows/activation-e2e.yml
Den 5822e471c6 fix(e2e): stop asserting on SSE reply in activation-positive — closes #90 (#102)
* fix(e2e): stop asserting on SSE reply in activation-positive — CI-harness limitation, not a product bug (closes #90)

Extensive investigation (see PR #102 for the full trail) into "the
positive flow's assistant reply never renders" tried four independent
SSE client transports in src/lib/sdk.ts global.events(): the
already-shipped expo/fetch ReadableStream reader, a hand-rolled
XMLHttpRequest reader, react-native-sse, and react-native-fetch-api's
`reactNative: { textStreaming: true }`. Every one delivers exactly one
chunk right after connecting to the mock server and then nothing until
the connection closes, regardless of API choice or frame size (a ~4KB
padding experiment ruled out a buffer-size threshold).

A raw-socket probe (a plain BSD-sockets client with zero React Native
involvement, run via `adb shell` through the identical adb-reverse
tunnel the app uses) streamed every heartbeat from the mock server
incrementally in real time over the same connection. That rules out
adb-reverse and the mock's flush behavior and isolates the stall to
React Native Android's OkHttp-backed networking layer buffering a
long-lived streaming HTTP response in this specific Android-emulator +
Node-mock + adb-reverse combination — not a defect in any particular
client library.

There's no evidence this reproduces against a real opencode server on
a real device/network: issue #76's 65 affected users prove real SSE
connections stream live agent output in production (the bug they hit
was the 401-retry storm, not a missing reply). Since expo/fetch is the
already-shipped, production-proven transport and none of the
alternatives showed any advantage in this harness, the transport stays
unchanged.

What changes instead: .maestro/flows/activation-positive.yaml no
longer waits on the SSE-streamed reply, since asserting on it here
would assert on a CI-harness limitation, not real app behavior. It now
verifies everything reliably observable — consent, connect, session
creation, and the optimistic local echo of the sent message — and
activation-e2e.yml's `continue-on-error: true` (added because this
suite had never passed) comes off, so it blocks PRs on regressions in
what it does cover.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(e2e): repair stale "401" assertion in activation-negative-401 (refs #90)

Removing activation-e2e.yml's continue-on-error surfaced a second, unrelated
stale assertion once the suite was actually enforcing again: the negative
flow's connect-time-401 case asserts a literal "401" that PR #79 (401/403
auth-stop handling) and #103 (i18n) apparently moved out of what's
rendered — "Connection Failed" still passes, "401" now fails.

The alert body interpolates two pieces: probeConnection()'s summary (which
turns out to be misclassified as "connection actually works now" for this
case — diagnostics-classify.ts's `health.ok` only reflects "fetch() didn't
throw", not HTTP status, a separate real bug, out of scope for this PR) and
testConnection()'s caught error message, which is sdk.ts's
apiErrorFor(401, ...) text and always contains the mock's
`{"error":"Unauthorized",...}` body per src/lib/api-error.test.ts. Swapped
the assertion to "Unauthorized" and added a temporary console.log of both
pieces in app/connection/add.tsx to confirm exactly what renders from CI
logcat (removed once confirmed).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(e2e): assert alert action buttons, not body text — native AlertDialog body isn't in Maestro's a11y tree (refs #90)

The diagnostic added last commit confirmed the Connection Failed alert's
body DOES contain the real error ("API Error: 401 -
{\"error\":\"Unauthorized\",...}", via logcat: '[connect] failure alert
content' logged both the (separately buggy, out-of-scope)
probeConnection summary and the correct testConnection error text). Yet
both "401" and "Unauthorized" assertions still failed against the same
on-screen alert. That means Maestro's accessibility-tree text matching
on this Android AlertDialog only sees the title, not the message body —
so no substring of the body was ever going to match.

Switched to asserting what's actually reachable: the title "Connection
Failed" (unchanged, already passing) plus both action button labels,
"OK" and "Share report" (src/lib/i18n/en.json common.ok /
common.shareReport). That still proves the test's real intent — a
visible, actionable error with a dismiss and a share-report path, never
a silent failure (issue #76) — using strings actually present in the
accessibility tree instead of guessing at unreachable body text.

Removes the temporary console.log diagnostic from
app/connection/add.tsx now that its purpose (confirming exactly what
renders) is done.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(e2e): split activation-e2e into blocking core + non-blocking newer flows (refs #90, refs #104)

With this PR's fixes, activation-positive and activation-negative-401
(the coverage issue #90 actually scoped) now run green — but removing
activation-e2e.yml's continue-on-error surfaced that four flows added
after the initial suite (#82's directory-picker/all-sessions/
variant-picker, #101's diff-scroll) have never once run to completion
in CI: they always sat behind whichever activation flow failed first,
so they were merged and have run unverified against the current
UI/mock this whole time. directory-picker fails immediately at
`id: directory-row-frontend`; the other three are untriaged.

Fixing four separate, previously-never-green UI surfaces is out of
#90's scope and unbounded in this PR. scripts/run-e2e-flows.sh now
splits the flow list into CORE_FLOWS (the two #90 covers — blocking,
fails the job on a regression) and NEWER_FLOWS (the four newer ones —
always run, each one's pass/fail reported via echo/::warning::, but
never fails the job). This lets activation-e2e.yml enforce the
activation coverage that's now verified, without either leaving it red
forever or spending unbounded time inside this PR chasing four
unrelated UI surfaces.

Filed #104 to track hardening each NEWER flow and moving it back into
CORE_FLOWS once confirmed green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-17 14:52:37 -07:00

199 lines
8.9 KiB
YAML

name: Activation E2E (Maestro)
# Deterministic regression coverage for the activation flow (first open ->
# telemetry consent -> server URL entry -> connect -> send first message),
# including the connect-time-401 negative case tied to the 0%-7-day-retention
# / GitHub issue #76 investigation.
#
# Also runs (non-blocking — see scripts/run-e2e-flows.sh's CORE_FLOWS vs
# NEWER_FLOWS split, and issue #104) newer surfaces merged after the initial
# activation suite: DirectoryBrowserSheet's server-folder picker, the
# directory-less "all sessions across all projects" list (+ the #46/#48
# open-across-project regression), VariantPicker's reasoning-effort chip, and
# DiffView/CodeBlock horizontal scroll. See .maestro/flows/directory-picker.yaml,
# all-sessions.yaml, variant-picker.yaml, diff-scroll.yaml. These never once
# ran to completion in CI (they sat behind the activation flows' stale
# assertions fixed by PR #102) and need their own hardening — tracked in #104.
#
# Runs against tests/fixtures/mock-opencode-server.ts (a small dependency-free
# HTTP+SSE stub matching the REAL client protocol read from src/lib/sdk.ts —
# NOT a live opencode server, NOT a WebSocket), so the suite is fast and fully
# self-contained: no external server, no LLM provider, no network flakiness.
#
# This is intentionally separate from cua-smoke.yml (the existing
# vision-driven CUA harness): that one needs a live opencode server + an Azure
# LLM and is exploratory/non-deterministic by design, so it isn't suited to
# tight regression assertions like "a 401 must show a visible error."
#
# activation-positive.yaml does NOT assert on the SSE-streamed reply — see
# the comment at the top of that file (issue #90 mode B / PR #102) for why:
# this Android-emulator + Node-mock + adb-reverse combination reliably
# stalls a long-lived SSE connection after its first chunk regardless of
# client transport, a CI-harness limitation with no evidence it affects real
# devices against a real server, so asserting on it here would be asserting
# on the harness rather than the app.
on:
push:
branches: [main]
paths:
- "app/**"
- "src/**"
- "tests/fixtures/**"
- ".maestro/**"
- ".github/workflows/activation-e2e.yml"
pull_request:
branches: [main]
paths:
- "app/**"
- "src/**"
- "tests/fixtures/**"
- ".maestro/**"
- ".github/workflows/activation-e2e.yml"
workflow_dispatch: {}
jobs:
activation-e2e:
runs-on: ubuntu-latest
# Same npm install + expo prebuild + assembleRelease + emulator pipeline as
# cua-smoke.yml, which budgets 60 min (emulator-boot-timeout alone is 10 min).
# Typical runs finish well under this; the ceiling just avoids flaky kills
# on cold Gradle caches.
timeout-minutes: 60
steps:
- uses: actions/checkout@v6
- uses: actions/setup-node@v6
with:
node-version: 24
cache: npm
- uses: actions/setup-java@v5
with:
distribution: temurin
java-version: 17
- name: Setup Android SDK
uses: android-actions/setup-android@v4
- name: Add emulator to PATH
run: echo "$ANDROID_HOME/emulator" >> $GITHUB_PATH
- name: Enable KVM
run: |
echo 'KERNEL=="kvm", GROUP="kvm", MODE="0666", OPTIONS+="static_node=kvm"' | sudo tee /etc/udev/rules.d/99-kvm4all.rules
sudo udevadm control --reload-rules
sudo udevadm trigger --name-match=kvm
- name: Cache Gradle
uses: actions/cache@v5
with:
path: |
~/.gradle/caches
~/.gradle/wrapper
android/.gradle
key: ${{ runner.os }}-gradle-${{ hashFiles('android/**/*.gradle*', 'android/gradle/wrapper/gradle-wrapper.properties') }}
restore-keys: |
${{ runner.os }}-gradle-
- name: Purge stale generated sources
# Same mitigation as cua-smoke.yml (whose Gradle cache entries this job
# shares — the cache key/restore-keys are identical): the restore-keys
# prefix fallback can restore a generated autolinking tree from a
# previous package id, making compileReleaseJavaWithJavac fail against
# the old package (ai.opencode.mobile vs cc.agentlabs.opencode). Delete
# generated sources so prebuild + Gradle regenerate them.
run: rm -rf android/app/build/generated android/build/generated android/app/build/intermediates
- name: Install Maestro CLI
run: |
curl -Ls "https://get.maestro.mobile.dev" | bash
echo "$HOME/.maestro/bin" >> $GITHUB_PATH
- name: Install dependencies & build APK
env:
SENTRY_DISABLE_AUTO_UPLOAD: "true"
run: |
npm install --legacy-peer-deps
npx expo prebuild --platform android --no-install
keytool -genkey -v -keystore android/app/debug.keystore -storepass android -alias androiddebugkey -keypass android -keyalg RSA -keysize 2048 -validity 10000 -dname "CN=Android Debug,O=Android,C=US"
cd android && ./gradlew assembleRelease
- name: Start mock opencode servers
run: |
set -x
mkdir -p artifacts/screenshots
# Normal mode on 4096 (positive flow) and --fail-auth on 4097 (negative flow).
# 4098 is normal mode + --seed-sessions (two pre-populated sessions in two
# different directories, for all-sessions.yaml). 4099 is a fresh normal-mode
# instance shared by directory-picker.yaml and variant-picker.yaml (its fake
# file tree / provider variants / project list don't affect each other).
# 4100 is normal mode + --seed-diff (one session with a pre-existing wide
# edit-diff tool call + wide code block, for diff-scroll.yaml / issue #21).
# All bind 0.0.0.0 so the emulator can reach them via 10.0.2.2.
nohup node tests/fixtures/mock-opencode-server.ts --port 4096 > /tmp/mock-4096.log 2>&1 &
nohup node tests/fixtures/mock-opencode-server.ts --port 4097 --fail-auth > /tmp/mock-4097.log 2>&1 &
nohup node tests/fixtures/mock-opencode-server.ts --port 4098 --seed-sessions > /tmp/mock-4098.log 2>&1 &
nohup node tests/fixtures/mock-opencode-server.ts --port 4099 > /tmp/mock-4099.log 2>&1 &
nohup node tests/fixtures/mock-opencode-server.ts --port 4100 --seed-diff > /tmp/mock-4100.log 2>&1 &
for port in 4096 4097 4098 4099 4100; do
for i in $(seq 1 30); do
if curl -sf --connect-timeout 1 -m 3 "http://127.0.0.1:${port}/global/health" > /dev/null 2>&1; then
echo "mock server on ${port} responded (may be 401, that's expected on 4097)"
break
fi
# 4097 always 401s -> curl -sf treats that as failure, so also accept "connection made"
if curl -s --connect-timeout 1 -m 3 -o /dev/null -w '%{http_code}' "http://127.0.0.1:${port}/global/health" 2>/dev/null | grep -qE '^[0-9]+$'; then
echo "mock server on ${port} is up (got an HTTP response)"
break
fi
if [ "$i" = "30" ]; then
echo "::error::mock server on port ${port} did not come up in 30s"
cat "/tmp/mock-${port}.log" || true
exit 1
fi
sleep 1
done
done
- name: Run activation E2E flows on emulator
uses: reactivecircus/android-emulator-runner@v2
with:
api-level: 28
arch: x86_64
target: default
disable-animations: true
emulator-boot-timeout: 600
emulator-options: -no-window -no-audio -no-boot-anim -gpu swiftshader_indirect -no-snapshot
# One script file, not one `sh -c` per line — run-e2e-flows.sh does the
# adb reverse port-forwarding, runs every flow, and captures logcat on
# exit. (Running maestro line-by-line here broke `cd` persistence and
# left no diagnostics.)
script: bash scripts/run-e2e-flows.sh
- name: Mock server logs
if: always()
run: |
echo "--- 4096 (normal) ---"; cat /tmp/mock-4096.log || true
echo "--- 4097 (fail-auth) ---"; cat /tmp/mock-4097.log || true
echo "--- 4098 (seed-sessions) ---"; cat /tmp/mock-4098.log || true
echo "--- 4099 (normal, directory-picker + variant-picker) ---"; cat /tmp/mock-4099.log || true
echo "--- 4100 (seed-diff) ---"; cat /tmp/mock-4100.log || true
- name: Upload screenshots
if: always()
uses: actions/upload-artifact@v7
with:
name: activation-e2e-screenshots-${{ github.run_number }}
path: artifacts/screenshots/*.png
if-no-files-found: warn
- name: Upload maestro debug + logcat
if: always()
uses: actions/upload-artifact@v7
with:
name: activation-e2e-debug-${{ github.run_number }}
path: |
artifacts/diag/**
if-no-files-found: warn