af7da46c0fff871377677c0c26873c09db96d0ca
6 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
c9ec92d8b7 |
feat(demo): add offline demo mode for zero-server activation (#108)
Installers with no self-hosted opencode server hit a dead end at the empty Sessions state, contributing to ~0% 7-day retention. Adds a fully offline, scripted /demo route reusing the real chat components (MessageBubble, ToolCallCard/DiffView, PermissionPrompt) so new users can see what opencode does before connecting anything, then funnels them to Connect / the setup guide. - src/lib/demo-script.ts: pure, hardcoded Message/Part fixture builder (no RN/store/network imports) — the isolation guarantee. - app/demo.tsx: new /demo route rendering the scripted conversation via useMemo'd local state only; permission reply is local setState, never sessionClient.permission.reply(). - app/(tabs)/index.tsx: "Try a demo" button added to the no-connection empty state, placed after the existing add-connection-button so its position/testID for existing Maestro flows is unchanged. - .maestro/flows/demo.yaml: new E2E flow covering the empty-state CTA through conversation, diff expand, permission approve, and the CTA reaching the real connect form. - scripts/run-e2e-flows.sh: registers demo in NEWER_FLOWS (non-blocking) so it actually runs in CI. - i18n: new sessionsList.empty.tryDemoButton and demo.* keys added to both en.json and zh-Hans.json (catalog-parity verified). npm run typecheck: clean. npm test: 175/175 passing. Co-authored-by: engineer <engineer@macbookpro.lan> |
||
|
|
5822e471c6 |
fix(e2e): stop asserting on SSE reply in activation-positive — closes #90 (#102)
* fix(e2e): stop asserting on SSE reply in activation-positive — CI-harness limitation, not a product bug (closes #90) Extensive investigation (see PR #102 for the full trail) into "the positive flow's assistant reply never renders" tried four independent SSE client transports in src/lib/sdk.ts global.events(): the already-shipped expo/fetch ReadableStream reader, a hand-rolled XMLHttpRequest reader, react-native-sse, and react-native-fetch-api's `reactNative: { textStreaming: true }`. Every one delivers exactly one chunk right after connecting to the mock server and then nothing until the connection closes, regardless of API choice or frame size (a ~4KB padding experiment ruled out a buffer-size threshold). A raw-socket probe (a plain BSD-sockets client with zero React Native involvement, run via `adb shell` through the identical adb-reverse tunnel the app uses) streamed every heartbeat from the mock server incrementally in real time over the same connection. That rules out adb-reverse and the mock's flush behavior and isolates the stall to React Native Android's OkHttp-backed networking layer buffering a long-lived streaming HTTP response in this specific Android-emulator + Node-mock + adb-reverse combination — not a defect in any particular client library. There's no evidence this reproduces against a real opencode server on a real device/network: issue #76's 65 affected users prove real SSE connections stream live agent output in production (the bug they hit was the 401-retry storm, not a missing reply). Since expo/fetch is the already-shipped, production-proven transport and none of the alternatives showed any advantage in this harness, the transport stays unchanged. What changes instead: .maestro/flows/activation-positive.yaml no longer waits on the SSE-streamed reply, since asserting on it here would assert on a CI-harness limitation, not real app behavior. It now verifies everything reliably observable — consent, connect, session creation, and the optimistic local echo of the sent message — and activation-e2e.yml's `continue-on-error: true` (added because this suite had never passed) comes off, so it blocks PRs on regressions in what it does cover. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(e2e): repair stale "401" assertion in activation-negative-401 (refs #90) Removing activation-e2e.yml's continue-on-error surfaced a second, unrelated stale assertion once the suite was actually enforcing again: the negative flow's connect-time-401 case asserts a literal "401" that PR #79 (401/403 auth-stop handling) and #103 (i18n) apparently moved out of what's rendered — "Connection Failed" still passes, "401" now fails. The alert body interpolates two pieces: probeConnection()'s summary (which turns out to be misclassified as "connection actually works now" for this case — diagnostics-classify.ts's `health.ok` only reflects "fetch() didn't throw", not HTTP status, a separate real bug, out of scope for this PR) and testConnection()'s caught error message, which is sdk.ts's apiErrorFor(401, ...) text and always contains the mock's `{"error":"Unauthorized",...}` body per src/lib/api-error.test.ts. Swapped the assertion to "Unauthorized" and added a temporary console.log of both pieces in app/connection/add.tsx to confirm exactly what renders from CI logcat (removed once confirmed). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(e2e): assert alert action buttons, not body text — native AlertDialog body isn't in Maestro's a11y tree (refs #90) The diagnostic added last commit confirmed the Connection Failed alert's body DOES contain the real error ("API Error: 401 - {\"error\":\"Unauthorized\",...}", via logcat: '[connect] failure alert content' logged both the (separately buggy, out-of-scope) probeConnection summary and the correct testConnection error text). Yet both "401" and "Unauthorized" assertions still failed against the same on-screen alert. That means Maestro's accessibility-tree text matching on this Android AlertDialog only sees the title, not the message body — so no substring of the body was ever going to match. Switched to asserting what's actually reachable: the title "Connection Failed" (unchanged, already passing) plus both action button labels, "OK" and "Share report" (src/lib/i18n/en.json common.ok / common.shareReport). That still proves the test's real intent — a visible, actionable error with a dismiss and a share-report path, never a silent failure (issue #76) — using strings actually present in the accessibility tree instead of guessing at unreachable body text. Removes the temporary console.log diagnostic from app/connection/add.tsx now that its purpose (confirming exactly what renders) is done. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(e2e): split activation-e2e into blocking core + non-blocking newer flows (refs #90, refs #104) With this PR's fixes, activation-positive and activation-negative-401 (the coverage issue #90 actually scoped) now run green — but removing activation-e2e.yml's continue-on-error surfaced that four flows added after the initial suite (#82's directory-picker/all-sessions/ variant-picker, #101's diff-scroll) have never once run to completion in CI: they always sat behind whichever activation flow failed first, so they were merged and have run unverified against the current UI/mock this whole time. directory-picker fails immediately at `id: directory-row-frontend`; the other three are untriaged. Fixing four separate, previously-never-green UI surfaces is out of #90's scope and unbounded in this PR. scripts/run-e2e-flows.sh now splits the flow list into CORE_FLOWS (the two #90 covers — blocking, fails the job on a regression) and NEWER_FLOWS (the four newer ones — always run, each one's pass/fail reported via echo/::warning::, but never fails the job). This lets activation-e2e.yml enforce the activation coverage that's now verified, without either leaving it red forever or spending unbounded time inside this PR chasing four unrelated UI surfaces. Filed #104 to track hardening each NEWER flow and moving it back into CORE_FLOWS once confirmed green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
0cac46cb36 |
test(diff): automate DiffView/CodeBlock horizontal-scroll coverage (closes #21) (#101)
Turns the manual QA ask ("verify DiffView + CodeBlock horizontal-scroll
on-device with a populated diff") into two automated layers:
1. Unit (deterministic, runs in `npm test` now): extracted the shared
ScrollView props into src/lib/scroll-config.ts (WIDE_CONTENT_SCROLL_CONFIG)
so DiffView.tsx and CodeBlock.tsx spread the SAME plain object their tests
assert on — no react-native-renderer needed. Added a source-scan
regression test (wide-content-scroll.regression.test.ts) that fails if
either component loses its ScrollView wiring or reintroduces
numberOfLines truncation.
2. E2E (Maestro): .maestro/flows/diff-scroll.yaml opens a session with a
pre-seeded wide edit-diff tool call and a wide fenced code block, then
swipes each horizontal ScrollView left and asserts the off-screen marker
text becomes visible. mock-opencode-server.ts gained a --seed-diff mode
that serves this session via GET /session/:id/message (pre-existing
history), not SSE — issue #90 (a separate SSE-render bug) is being fixed
independently, and this flow must not depend on it landing first. Wired
the new flow + port 4100 into run-e2e-flows.sh and activation-e2e.yml.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
b52fa52a4c |
test(e2e): instrument mock SSE + client stream + fix reachability probe to attribute #90 mode-B (#98)
Adds diagnostics-only instrumentation to attribute the Activation E2E positive flow's "SSE reply never renders" failure (issue #90 mode B) between (a) expo/fetch not streaming the SSE response on the Android release APK, vs (b) a mock-side broadcast bug. Does not change app behavior or fix the root cause — #90 stays open pending the next CI run's enriched logs. - tests/fixtures/mock-opencode-server.ts: per-request logging (method/path/status), per-SSE-connection connect/disconnect logging with live client count, per-broadcast event-type + client-count logging, and a 2s SSE heartbeat comment so client-side silence becomes unambiguous. - src/lib/sdk.ts global.events(): logs on the first successful reader.read() that returns data, and when the stream loop ends — proves/disproves whether expo/fetch ever delivers a byte. - scripts/run-e2e-flows.sh: the emulator->mock reachability probe used toybox wget/nc, which don't work reliably on the API-28 image. Replaced with a probe chain (curl, wget, mksh /dev/tcp, nc, then a host-side fallback) that writes a clear PASS/FAIL/UNKNOWN verdict to artifacts/diag/probe.txt without blocking the flow. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
2da0fdf109 |
test: E2E coverage for directory picker, all-sessions, variant picker. Refs #46 #48 #47 #49 #57. (#82)
* test: E2E coverage for directory picker, all-sessions, variant picker Extend the Maestro suite for the features merged into main today: DirectoryBrowserSheet's server-folder picker, the directory-less all-sessions-across-projects list (+ the #46/#48 open-across-project regression), and VariantPicker's reasoning-effort chip. - tests/fixtures/mock-opencode-server.ts: GET /file (directory-scoped via the x-opencode-directory header) with a small fake tree, GET /project for the "Server Projects" section, POST /session honoring the directory header, GET /session/:id (needed to open a session from the all-sessions list), GET /provider variants for VariantPicker, and an optional --seed-sessions mode that pre-populates two sessions across two directories. --fail-auth mode is untouched. - .maestro/flows/directory-picker.yaml, all-sessions.yaml, variant-picker.yaml: three new flows, run in the same emulator session as the existing activation flows. - Additive testIDs on DirectoryBrowserSheet, the "Browse Folders" row, session list rows, the variant chip, and VariantPicker rows. - .github/workflows/activation-e2e.yml: two more mock server instances (4098 seeded, 4099 fresh) and three more maestro test steps. Verified: tsc --noEmit clean, all 108 existing unit tests pass, every new mock endpoint curled against its real shape read from the app code, YAML validated. No Android emulator available locally to run the Maestro flows themselves. * test(mock): enforce per-directory session scoping so #46/#48 coverage can fail Review finding (HIGH): GET /session/:id and /session/:id/message ignored x-opencode-directory, so all-sessions.yaml could not fail if the directory threading fix regressed. The mock now mirrors the real server's per-directory workspace scoping: - GET /session/:id and GET /session/:id/message 404 unless the request's x-opencode-directory (or DEFAULT_DIRECTORY when absent) matches the stored session's directory. - GET /session without ?roots=true is scoped to the request's directory; loadSessions()'s directory-less roots=true call still returns everything. - Document the port-4099 shared-state coupling between directory-picker and variant-picker flows, and why all-sessions.yaml now has teeth (flow comment). Curl-verified: correct header 200, wrong/no header 404, scoped vs roots listing, create-then-open paths for all three flows, --fail-auth untouched. tsc clean, 108/108 unit tests pass. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NJKAQ6HAikWGQK7PGZ5Y4E * fix(e2e): widen connect-handshake wait past client's own 30s timeout Run 29546383612 (7cbd3a6, first real emulator execution of these flows) failed on activation-positive.yaml: "Assert that id: connection-status-dot is visible" timed out after the flow's 20s extendedWaitUntil, right after tapOn connect-submit-button. The mock server itself is fast (verified locally: health + project/current + path respond in ~30ms total), so this isn't a mock fidelity gap. But Quick Connect's testConnection()/addConnection() path chains up to 3 fetches (health, then project.current + path.get in parallel), and each individual fetch is capped by src/lib/sdk.ts REQUEST_TIMEOUT_MS = 30_000 — strictly longer than the 20s the flow was willing to wait. A first-attempt emulator-to-host (10.0.2.2) connection that's merely slow to establish, rather than outright failing, would blow past the test's wait before the app's own client-side timeout even fires. Bump the connect -> connection-status-dot / "Connection Failed" waits from 20000 to 40000 across all 5 flows that share this pattern (activation-positive, activation-negative-401, all-sessions, directory-picker, variant-picker) so the wait is never shorter than the code path it's gating on. Assertions are unchanged — still requires the real dot / real error text, just with a timeout that isn't racing the client. Verified locally: typecheck clean, all 108 unit tests pass, YAML parses, mock server confirmed fast under direct curl. Emulator behavior itself (whether 40s consistently clears it) is unverified until the next CI run. * fix(e2e): use adb reverse + 127.0.0.1 instead of 10.0.2.2; capture logcat/maestro debug Root cause of the activation-e2e failure (connect step timed out, ~0 requests reaching the mock): the 10.0.2.2 host alias is unreliable under the headless emulator-runner — the app's http://10.0.2.2:4096/global/health never completed, so connection-status-dot never rendered. - run-e2e-flows.sh: single script (fixes cd-per-line fragility) that adb-reverses each mock port (4096-4099) into the emulator's localhost, runs every flow with --debug-output, and dumps logcat on exit. - All flows now connect to 127.0.0.1:<port> (the adb reverse target). - Upload maestro-debug (UI hierarchy on failure) + logcat as artifacts so future failures are diagnosable instead of blind. * fix(e2e): connect via 127.0.0.1:PORT in IP field, stop typing into port input Root cause of every activation-e2e connect failure (proven by the app's own logcat diagnostic: '[diag] probe start http://127.0.0.1:40966 ... server unreachable'): the port field defaults to useState("4096"), and the flow's eraseText + inputText "4096" raced the controlled number-pad input, leaving "40966" — nothing listens there, so connect always failed. This was never a 10.0.2.2 / adb reverse issue. Fix: buildUrl already extracts host:port from the IP field, so enter 127.0.0.1:<port> there and remove the flaky port-field steps entirely. pastedPort overrides the default port state, so each flow's port is deterministic (4096 positive / 4097 negative / 4098 all-sessions / 4099 directory+variant). --------- Co-authored-by: engineer <engineer@gray-knight-m1.local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
6a9bb6d2d5 |
fix(ci): make activation-e2e harness actually run its flows (#91)
Port the harness fixes from test/e2e-new-features to main so the Activation E2E workflow stops failing before any flow executes: - Run everything through scripts/run-e2e-flows.sh as a single script invocation: android-emulator-runner executes each 'script:' line in its own shell, so the previous 'cd artifacts/screenshots' never persisted and maestro failed with 'Flow path does not exist' on every run (19/19 red since the workflow landed). - adb reverse + 127.0.0.1 instead of 10.0.2.2 (unreliable headless), emulator->mock reachability probe, logcat + maestro debug capture. - Trim the flow list to the two flows that exist on main; the three newer flows land with the test/e2e-new-features PR. - continue-on-error until the suite's first green: the positive flow still fails its final reply assertion (mode B in #90), and a never-green suite should not block unrelated PRs or pollute the product-intelligence failure metrics (#89). Refs #90. Refs #89. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |