f564869eee1485e505794827e62c3c6892d83bd3
10 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
263ef9c4b6 |
test(demo): full-text regex matches so the demo E2E flow passes (#110)
* test(demo): use full-text regex matches in demo Maestro flow The demo flow (added in #108, non-blocking lane) rendered correctly in CI but failed its own assertions: it matched bare substrings ('login button', 'Tests passed') against Maestro's full-text regex matcher, which needs .*…* to match a phrase inside a longer message. A CI run confirmed the demo screen renders (S2 screenshot shows the user message, assistant reasoning, and tool card) — only the assertions were wrong. Exact i18n labels (Thinking / Permission Required / Allow) already matched. Flow-only; the app and demo feature are unchanged. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01T12AhSnQVrSxNnvwfCx2z6 * test(demo): scroll to diff after expanding the tool card CI got further with the regex fix but then failed asserting diff-view-scroll: the confirmed cause (failure screenshot) is that the Edit card is scrolled to the bottom of the viewport to be tapped, so its expanded diff opens below the fold, and assertVisible does not auto-scroll. Add a scrollUntilVisible for diff-view-scroll before the assert. Flow-only. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01T12AhSnQVrSxNnvwfCx2z6 * test(demo): scroll to completion message before asserting it Same below-the-fold pattern as the diff step: after approving the permission at the bottom of the viewport, the 'Tests passed' completion renders further down. Scroll to it before asserting. Preemptive, to avoid another CI cycle. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01T12AhSnQVrSxNnvwfCx2z6 --------- Co-authored-by: engineer <engineer@macbookpro.lan> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
c9ec92d8b7 |
feat(demo): add offline demo mode for zero-server activation (#108)
Installers with no self-hosted opencode server hit a dead end at the empty Sessions state, contributing to ~0% 7-day retention. Adds a fully offline, scripted /demo route reusing the real chat components (MessageBubble, ToolCallCard/DiffView, PermissionPrompt) so new users can see what opencode does before connecting anything, then funnels them to Connect / the setup guide. - src/lib/demo-script.ts: pure, hardcoded Message/Part fixture builder (no RN/store/network imports) — the isolation guarantee. - app/demo.tsx: new /demo route rendering the scripted conversation via useMemo'd local state only; permission reply is local setState, never sessionClient.permission.reply(). - app/(tabs)/index.tsx: "Try a demo" button added to the no-connection empty state, placed after the existing add-connection-button so its position/testID for existing Maestro flows is unchanged. - .maestro/flows/demo.yaml: new E2E flow covering the empty-state CTA through conversation, diff expand, permission approve, and the CTA reaching the real connect form. - scripts/run-e2e-flows.sh: registers demo in NEWER_FLOWS (non-blocking) so it actually runs in CI. - i18n: new sessionsList.empty.tryDemoButton and demo.* keys added to both en.json and zh-Hans.json (catalog-parity verified). npm run typecheck: clean. npm test: 175/175 passing. Co-authored-by: engineer <engineer@macbookpro.lan> |
||
|
|
dfc38b3276 |
fix(e2e): directory-picker race, markdown a11y, variant-picker SSE softening (issue #104 cont.) (#106)
* fix(e2e): fix directory-picker race + markdown accessibility, soften variant-picker SSE assertion Second iteration against real CI evidence from run 29617520311 (PR #105): directory-picker still failed after enabling static snapPoints. The mock server's own request log proved GET /file was still never called, meaning DirectoryBrowserSheet's onChange never ran enter(). Root cause: the caller (openBrowser in app/(tabs)/index.tsx) sets startDirectory via setState and calls sheetRef.current?.expand() synchronously in the same handler. expand() kicks off a reanimated-driven animation whose onChange fires before React commits the re-render that would give the child the new startDirectory prop, so the first onChange(index=0) captured the stale initial `null` and set wasOpen=true — permanently blocking every later onChange for that open. Mirrored startDirectory into a ref (updated inline on every render) so handleSheetChange always reads the latest value regardless of which render's closure actually fires. diff-scroll still failed even after removing the nested FlatList — but the new diagnostic screenshot showed the text WAS visually on screen while Maestro's accessibility-tree-based assertion still couldn't find it for the full timeout. That matches a real, still-open React Native Android bug (facebook/react-native#46999, a reopened regression of #28952's fix): selectable Text inside a FlatList row doesn't get its selectable/accessible state applied correctly. react-native-marked's base Renderer hardcodes `selectable` on every plain text node (text/strong/em/del/heading/codespan). Overrode those in Markdown.tsx's CustomRenderer to render plain (non- selectable) Text — code content stays copyable via CodeBlock's own Copy button. variant-picker: confirmed the model-selection fix worked completely (chip appears, opens, selects, label updates) and the flow only fails afterward at the exact same SSE-streamed-reply limitation documented in activation-positive.yaml (issue #90 mode B — this CI harness's Android emulator + Node mock + adb-reverse combination cannot deliver more than the SSE stream's first chunk). Softened the post-send assertion to match activation-positive's pattern: verify the optimistic local echo (chat-bubble-user) instead of waiting on the unrenderable-in-CI reply. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(ci): capture maestro hierarchy dumps in debug artifact upload actions/upload-artifact excludes dotfiles/dot-directories by default, and maestro's --debug-output nests the actual UI-hierarchy dump under a hidden .maestro/tests/<timestamp>/ directory — so every activation-e2e run has been silently uploading only logcat.txt/probe.txt and dropping the one artifact most useful for diagnosing flow failures (issue #104). Set include-hidden-files: true on that upload step. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
2285f80e81 |
fix(e2e): harden activation-e2e newer flows (issue #104) (#105)
Root-caused directory-picker's directory-row-frontend failure: all four
in-app @gorhom/bottom-sheet sheets (DirectoryBrowserSheet, DirectorySwitcher,
ModelPicker, VariantPicker) provide static percentage snapPoints but rely on
v5's enableDynamicSizing default (true), which never resolves without content
wrapped in a size-reporting component — so useAnimatedDetents() permanently
early-exits and the sheets can never actually open. Set
enableDynamicSizing={false} on all four (they already have explicit
snapPoints, so dynamic sizing was never needed).
variant-picker's chip failure was a stale test assumption: src/lib/
model-selection.ts's chooseModelSelection() deliberately returns null for a
fresh session (issue #37/#35 — the provider registry default is unreliable),
so a brand-new session has no model selected and the reasoning-effort chip
has nothing to key off of. Added testIDs (model-chip, model-option-*) and
updated the flow to explicitly pick a model first, matching real usage.
diff-scroll's missing markdown text: switched src/components/markdown/
Markdown.tsx from react-native-marked's FlatList-based default export to its
useMarkdown() hook rendered into a plain View. The chat screen already nests
this inside its own *inverted* FlatList (one row per message) — a nested
VirtualizedList inside an inverted outer list is a known RN footgun where the
inner content can render at zero height instead of just warning. We already
forced scrollEnabled:false + a large initialNumToRender, defeating
virtualization anyway, so rendering the parsed blocks directly loses nothing.
Extended the existing react-native-marked .d.ts shim (added for a React
18/19 ReactNode mismatch) to also declare useMarkdown/useMarkdownHookOptions.
Added diagnostic screenshots to directory-picker.yaml and diff-scroll.yaml
at the previously-failing steps for faster triage if these regress again.
Closes #104.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
5822e471c6 |
fix(e2e): stop asserting on SSE reply in activation-positive — closes #90 (#102)
* fix(e2e): stop asserting on SSE reply in activation-positive — CI-harness limitation, not a product bug (closes #90) Extensive investigation (see PR #102 for the full trail) into "the positive flow's assistant reply never renders" tried four independent SSE client transports in src/lib/sdk.ts global.events(): the already-shipped expo/fetch ReadableStream reader, a hand-rolled XMLHttpRequest reader, react-native-sse, and react-native-fetch-api's `reactNative: { textStreaming: true }`. Every one delivers exactly one chunk right after connecting to the mock server and then nothing until the connection closes, regardless of API choice or frame size (a ~4KB padding experiment ruled out a buffer-size threshold). A raw-socket probe (a plain BSD-sockets client with zero React Native involvement, run via `adb shell` through the identical adb-reverse tunnel the app uses) streamed every heartbeat from the mock server incrementally in real time over the same connection. That rules out adb-reverse and the mock's flush behavior and isolates the stall to React Native Android's OkHttp-backed networking layer buffering a long-lived streaming HTTP response in this specific Android-emulator + Node-mock + adb-reverse combination — not a defect in any particular client library. There's no evidence this reproduces against a real opencode server on a real device/network: issue #76's 65 affected users prove real SSE connections stream live agent output in production (the bug they hit was the 401-retry storm, not a missing reply). Since expo/fetch is the already-shipped, production-proven transport and none of the alternatives showed any advantage in this harness, the transport stays unchanged. What changes instead: .maestro/flows/activation-positive.yaml no longer waits on the SSE-streamed reply, since asserting on it here would assert on a CI-harness limitation, not real app behavior. It now verifies everything reliably observable — consent, connect, session creation, and the optimistic local echo of the sent message — and activation-e2e.yml's `continue-on-error: true` (added because this suite had never passed) comes off, so it blocks PRs on regressions in what it does cover. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(e2e): repair stale "401" assertion in activation-negative-401 (refs #90) Removing activation-e2e.yml's continue-on-error surfaced a second, unrelated stale assertion once the suite was actually enforcing again: the negative flow's connect-time-401 case asserts a literal "401" that PR #79 (401/403 auth-stop handling) and #103 (i18n) apparently moved out of what's rendered — "Connection Failed" still passes, "401" now fails. The alert body interpolates two pieces: probeConnection()'s summary (which turns out to be misclassified as "connection actually works now" for this case — diagnostics-classify.ts's `health.ok` only reflects "fetch() didn't throw", not HTTP status, a separate real bug, out of scope for this PR) and testConnection()'s caught error message, which is sdk.ts's apiErrorFor(401, ...) text and always contains the mock's `{"error":"Unauthorized",...}` body per src/lib/api-error.test.ts. Swapped the assertion to "Unauthorized" and added a temporary console.log of both pieces in app/connection/add.tsx to confirm exactly what renders from CI logcat (removed once confirmed). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(e2e): assert alert action buttons, not body text — native AlertDialog body isn't in Maestro's a11y tree (refs #90) The diagnostic added last commit confirmed the Connection Failed alert's body DOES contain the real error ("API Error: 401 - {\"error\":\"Unauthorized\",...}", via logcat: '[connect] failure alert content' logged both the (separately buggy, out-of-scope) probeConnection summary and the correct testConnection error text). Yet both "401" and "Unauthorized" assertions still failed against the same on-screen alert. That means Maestro's accessibility-tree text matching on this Android AlertDialog only sees the title, not the message body — so no substring of the body was ever going to match. Switched to asserting what's actually reachable: the title "Connection Failed" (unchanged, already passing) plus both action button labels, "OK" and "Share report" (src/lib/i18n/en.json common.ok / common.shareReport). That still proves the test's real intent — a visible, actionable error with a dismiss and a share-report path, never a silent failure (issue #76) — using strings actually present in the accessibility tree instead of guessing at unreachable body text. Removes the temporary console.log diagnostic from app/connection/add.tsx now that its purpose (confirming exactly what renders) is done. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(e2e): split activation-e2e into blocking core + non-blocking newer flows (refs #90, refs #104) With this PR's fixes, activation-positive and activation-negative-401 (the coverage issue #90 actually scoped) now run green — but removing activation-e2e.yml's continue-on-error surfaced that four flows added after the initial suite (#82's directory-picker/all-sessions/ variant-picker, #101's diff-scroll) have never once run to completion in CI: they always sat behind whichever activation flow failed first, so they were merged and have run unverified against the current UI/mock this whole time. directory-picker fails immediately at `id: directory-row-frontend`; the other three are untriaged. Fixing four separate, previously-never-green UI surfaces is out of #90's scope and unbounded in this PR. scripts/run-e2e-flows.sh now splits the flow list into CORE_FLOWS (the two #90 covers — blocking, fails the job on a regression) and NEWER_FLOWS (the four newer ones — always run, each one's pass/fail reported via echo/::warning::, but never fails the job). This lets activation-e2e.yml enforce the activation coverage that's now verified, without either leaving it red forever or spending unbounded time inside this PR chasing four unrelated UI surfaces. Filed #104 to track hardening each NEWER flow and moving it back into CORE_FLOWS once confirmed green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
0cac46cb36 |
test(diff): automate DiffView/CodeBlock horizontal-scroll coverage (closes #21) (#101)
Turns the manual QA ask ("verify DiffView + CodeBlock horizontal-scroll
on-device with a populated diff") into two automated layers:
1. Unit (deterministic, runs in `npm test` now): extracted the shared
ScrollView props into src/lib/scroll-config.ts (WIDE_CONTENT_SCROLL_CONFIG)
so DiffView.tsx and CodeBlock.tsx spread the SAME plain object their tests
assert on — no react-native-renderer needed. Added a source-scan
regression test (wide-content-scroll.regression.test.ts) that fails if
either component loses its ScrollView wiring or reintroduces
numberOfLines truncation.
2. E2E (Maestro): .maestro/flows/diff-scroll.yaml opens a session with a
pre-seeded wide edit-diff tool call and a wide fenced code block, then
swipes each horizontal ScrollView left and asserts the off-screen marker
text becomes visible. mock-opencode-server.ts gained a --seed-diff mode
that serves this session via GET /session/:id/message (pre-existing
history), not SSE — issue #90 (a separate SSE-render bug) is being fixed
independently, and this flow must not depend on it landing first. Wired
the new flow + port 4100 into run-e2e-flows.sh and activation-e2e.yml.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
2da0fdf109 |
test: E2E coverage for directory picker, all-sessions, variant picker. Refs #46 #48 #47 #49 #57. (#82)
* test: E2E coverage for directory picker, all-sessions, variant picker Extend the Maestro suite for the features merged into main today: DirectoryBrowserSheet's server-folder picker, the directory-less all-sessions-across-projects list (+ the #46/#48 open-across-project regression), and VariantPicker's reasoning-effort chip. - tests/fixtures/mock-opencode-server.ts: GET /file (directory-scoped via the x-opencode-directory header) with a small fake tree, GET /project for the "Server Projects" section, POST /session honoring the directory header, GET /session/:id (needed to open a session from the all-sessions list), GET /provider variants for VariantPicker, and an optional --seed-sessions mode that pre-populates two sessions across two directories. --fail-auth mode is untouched. - .maestro/flows/directory-picker.yaml, all-sessions.yaml, variant-picker.yaml: three new flows, run in the same emulator session as the existing activation flows. - Additive testIDs on DirectoryBrowserSheet, the "Browse Folders" row, session list rows, the variant chip, and VariantPicker rows. - .github/workflows/activation-e2e.yml: two more mock server instances (4098 seeded, 4099 fresh) and three more maestro test steps. Verified: tsc --noEmit clean, all 108 existing unit tests pass, every new mock endpoint curled against its real shape read from the app code, YAML validated. No Android emulator available locally to run the Maestro flows themselves. * test(mock): enforce per-directory session scoping so #46/#48 coverage can fail Review finding (HIGH): GET /session/:id and /session/:id/message ignored x-opencode-directory, so all-sessions.yaml could not fail if the directory threading fix regressed. The mock now mirrors the real server's per-directory workspace scoping: - GET /session/:id and GET /session/:id/message 404 unless the request's x-opencode-directory (or DEFAULT_DIRECTORY when absent) matches the stored session's directory. - GET /session without ?roots=true is scoped to the request's directory; loadSessions()'s directory-less roots=true call still returns everything. - Document the port-4099 shared-state coupling between directory-picker and variant-picker flows, and why all-sessions.yaml now has teeth (flow comment). Curl-verified: correct header 200, wrong/no header 404, scoped vs roots listing, create-then-open paths for all three flows, --fail-auth untouched. tsc clean, 108/108 unit tests pass. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NJKAQ6HAikWGQK7PGZ5Y4E * fix(e2e): widen connect-handshake wait past client's own 30s timeout Run 29546383612 (7cbd3a6, first real emulator execution of these flows) failed on activation-positive.yaml: "Assert that id: connection-status-dot is visible" timed out after the flow's 20s extendedWaitUntil, right after tapOn connect-submit-button. The mock server itself is fast (verified locally: health + project/current + path respond in ~30ms total), so this isn't a mock fidelity gap. But Quick Connect's testConnection()/addConnection() path chains up to 3 fetches (health, then project.current + path.get in parallel), and each individual fetch is capped by src/lib/sdk.ts REQUEST_TIMEOUT_MS = 30_000 — strictly longer than the 20s the flow was willing to wait. A first-attempt emulator-to-host (10.0.2.2) connection that's merely slow to establish, rather than outright failing, would blow past the test's wait before the app's own client-side timeout even fires. Bump the connect -> connection-status-dot / "Connection Failed" waits from 20000 to 40000 across all 5 flows that share this pattern (activation-positive, activation-negative-401, all-sessions, directory-picker, variant-picker) so the wait is never shorter than the code path it's gating on. Assertions are unchanged — still requires the real dot / real error text, just with a timeout that isn't racing the client. Verified locally: typecheck clean, all 108 unit tests pass, YAML parses, mock server confirmed fast under direct curl. Emulator behavior itself (whether 40s consistently clears it) is unverified until the next CI run. * fix(e2e): use adb reverse + 127.0.0.1 instead of 10.0.2.2; capture logcat/maestro debug Root cause of the activation-e2e failure (connect step timed out, ~0 requests reaching the mock): the 10.0.2.2 host alias is unreliable under the headless emulator-runner — the app's http://10.0.2.2:4096/global/health never completed, so connection-status-dot never rendered. - run-e2e-flows.sh: single script (fixes cd-per-line fragility) that adb-reverses each mock port (4096-4099) into the emulator's localhost, runs every flow with --debug-output, and dumps logcat on exit. - All flows now connect to 127.0.0.1:<port> (the adb reverse target). - Upload maestro-debug (UI hierarchy on failure) + logcat as artifacts so future failures are diagnosable instead of blind. * fix(e2e): connect via 127.0.0.1:PORT in IP field, stop typing into port input Root cause of every activation-e2e connect failure (proven by the app's own logcat diagnostic: '[diag] probe start http://127.0.0.1:40966 ... server unreachable'): the port field defaults to useState("4096"), and the flow's eraseText + inputText "4096" raced the controlled number-pad input, leaving "40966" — nothing listens there, so connect always failed. This was never a 10.0.2.2 / adb reverse issue. Fix: buildUrl already extracts host:port from the IP field, so enter 127.0.0.1:<port> there and remove the flaky port-field steps entirely. pastedPort overrides the default port state, so each flow's port is deterministic (4096 positive / 4097 negative / 4098 all-sessions / 4099 directory+variant). --------- Co-authored-by: engineer <engineer@gray-knight-m1.local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
6a9bb6d2d5 |
fix(ci): make activation-e2e harness actually run its flows (#91)
Port the harness fixes from test/e2e-new-features to main so the Activation E2E workflow stops failing before any flow executes: - Run everything through scripts/run-e2e-flows.sh as a single script invocation: android-emulator-runner executes each 'script:' line in its own shell, so the previous 'cd artifacts/screenshots' never persisted and maestro failed with 'Flow path does not exist' on every run (19/19 red since the workflow landed). - adb reverse + 127.0.0.1 instead of 10.0.2.2 (unreliable headless), emulator->mock reachability probe, logcat + maestro debug capture. - Trim the flow list to the two flows that exist on main; the three newer flows land with the test/e2e-new-features PR. - continue-on-error until the suite's first green: the positive flow still fails its final reply assertion (mode B in #90), and a never-green suite should not block unrelated PRs or pollute the product-intelligence failure metrics (#89). Refs #90. Refs #89. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
7ca5d2eb19 |
test(activation): address code-review findings on E2E flows, mock, CI
Review fixes (REQUEST_CHANGES round 1): 1. HIGH activation-negative-401.yaml: after dismissing the "Connection Failed" alert the app stays on the Add-Connection modal (handleQuickConnect's failure branch never calls router.back()), so the old `text: "No Connection"` assertion (Sessions-tab empty state) could never pass. Now asserts connect-submit-button is still visible instead. 2. MEDIUM mock-opencode-server.ts: prompt_async now parses the request body, persists the USER's message, and broadcasts it (message.updated + message.part.updated) BEFORE the canned assistant reply — matching real server behavior. Without this, the app's handleEvent strips the optimistic temp- user message when the assistant's message.updated arrives and the sent message vanishes from the transcript. activation-positive.yaml now also asserts chat-bubble-user and the user's message text are visible after the reply lands, so that regression class is actually covered. 3. MEDIUM activation-e2e.yml: timeout-minutes 15 -> 60. The job runs the same npm install + prebuild + assembleRelease + emulator pipeline that cua-smoke.yml budgets 60 min for (emulator-boot-timeout alone is 10 min). 4. MEDIUM activation-e2e.yml: replicated cua-smoke.yml's "Purge stale generated sources" step — the Gradle cache key/restore-keys are shared with that workflow, so the stale-autolinking-tree failure mode (compileReleaseJavaWithJavac against the old package id) applies here too. Verified locally: tsc --noEmit clean; npm test 81/81 pass; all three touched YAML files parse valid; mock server exercised standalone — full prompt cycle confirms GET /session/:id/message returns BOTH user and assistant messages, SSE order is message.updated(user) -> message.part.updated(user) -> busy -> message.updated(assistant) -> message.part.updated(assistant) -> idle, user events carry the sessionID/messageID fields handleEvent filters on, and --fail-auth mode returns 401. Still no emulator run in this environment. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NJKAQ6HAikWGQK7PGZ5Y4E |
||
|
|
01dd0191b3 |
test(activation): add Maestro E2E coverage for the activation flow
Adds deterministic end-to-end coverage for first-open -> telemetry consent -> server URL entry -> connect -> send first message -> receive reply, targeting the 0%-7-day-retention investigation (GitHub issue #76). - tests/fixtures/mock-opencode-server.ts: dependency-free HTTP+SSE stub matching the REAL client protocol (src/lib/sdk.ts) — REST + a single long-lived GET /global/event SSE stream, no WebSocket. Supports a --fail-auth mode that 401s every request to exercise the connect-time auth-failure class. - .maestro/flows/activation-positive.yaml: consent -> quick connect -> new session -> send message -> assert streamed reply renders, with a screenshot at every step (positive-S1..S8). - .maestro/flows/activation-negative-401.yaml: same setup against the --fail-auth server, asserts Quick Connect's existing "Connection Failed" alert is shown (not silently swallowed) and that the connection is not saved. Flags in comments that Advanced-mode Save (handleAdvancedSave) still has no testConnection() check and is a known, uncovered gap. - testID props added (no restructuring) to the screens/components the flows drive: TelemetryConsentModal, connection/add.tsx, tabs/index.tsx, session/[id].tsx, MessageBubble. - .github/workflows/activation-e2e.yml: new CI job — Android emulator via reactivecircus/android-emulator-runner, builds the debug-signed APK, starts both mock server instances, runs both Maestro flows, uploads screenshots via actions/upload-artifact. Kept separate from the existing vision-driven cua-smoke.yml, which needs a live server + LLM and isn't suited to tight deterministic regression assertions. - .gitignore: Maestro takeScreenshot output is never committed. Verified locally: mock server exercised standalone via curl (health, project/current, path, session create, SSE event ordering, message persistence) in both normal and --fail-auth modes; both Maestro flow files validated as well-formed YAML; tsc --noEmit clean on all changed files. No lint script exists in this repo (N/A). Full emulator execution was not run — no Android SDK/emulator available in this environment. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NJKAQ6HAikWGQK7PGZ5Y4E |