709aa97bade9734664272dc2aac41e9d5c88d98f
7 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
4d64700b1b |
tools(sentry): anchor the measurement windows on the gate's rollout instant (#178)
A window that spans the 2026-08-14 14:22Z production rollout contains devices that could not possibly have run the gate. Its rate is neither a baseline nor a result, and it prints identically to both. This was not hypothetical: a `post=08-14T07:00Z..now` window (84% of it pre-rollout) was run against this script and reported opencode-mobile *rising* to 4.44/h. `noise-gate-report.mjs` shipped with the same defect built into its default: `post = now-7d..now` straddles the rollout on every run before 08-21, diluting the after-rate toward baseline - biased toward grading the gate as ineffective on exactly the dates the ticket schedules its reads (08-17, 08-21). - sentry-volume-report: `--since-rollout` reads the instant from the release history table in docs/playstore.md (production versionCode >= 150, earliest such release, so a later v0.4.15 does not restart the window) and splits there. Every window is labelled [pre]/[post]/[mixed]; mixed prints how much of it predates the gate, a young post window prints its uptake age, and an unparseable table reports "unknown" rather than assuming post. - noise-gate-report: defaults post to the rollout instant, returns UNGRADED for a mixed/unknown post window, and pins the baseline to the documented post-box-bot-fix window instead of a 7d lookback that dragged ~22k/mo of already-fixed box-bot volume into the org outlook (it read "MISSES by 18,612" for a dead reason; now 628/mo, clears). - before_send == 0 is now reported as expected in a pre/mixed/young window and as a failure only after 24h+ of gated production. Re-probed every server-side lever with a WRITE-scoped token so none of the answers is a permissions artifact, and corrected the record in docs/analytics.md: per-key rate limit returns 200 and silently drops the field; error-message filters return 400 "You do not have that feature enabled" (a plan gate, not absence - it is the one lever that would reach never-updating installs); spike protection is not 403-unavailable, it is already enabled everywhere and simply does not fire on sustained baseline volume. Co-authored-by: engineer <engineer@macbookpro.lan> |
||
|
|
d31afc0389 |
tools(sentry): grade the noise gate against install-base uptake, not raw volume (#175)
The gate ships inside the app binary, so it only runs on devices that took
v0.4.14. Grading it on a raw event count is a measurement error in both
directions: a still-high number at 30% uptake is the gate WORKING (~70% of
baseline is the model's own prediction), and a dip from a quiet weekend is not
efficacy. The scheduled re-reads on 08-17 / 08-21 / 09-05 would have hit the
first one first.
noise-gate-report.mjs folds Play's version share into the comparison
expected_post = baseline x (1 - gated_share x 0.969)
where 0.969 is measured, not guessed (90d replay in sentry-noise.test.ts), and
grades the measured rate against that instead of against the 100%-uptake
endpoint. It reports the endpoint separately, so 'is it on track today' and
'will it clear the 3,500/mo org gate' stop being the same question, and it
inverts the model to print the IMPLIED on-device efficacy so the constant is
checked rather than trusted.
It refuses to grade two windows that look like results but are not: no
client_discard/before_send (nothing ran the gate) and 0% Play share. Absence of
evidence gets its own verdict, UNGRADED.
Runs in CI because neither credential (Sentry org token, Play service account)
exists outside GitHub Secrets — an agent picking up the 08-21 read locally is
stuck otherwise. Weekly cron records the trend regardless.
Verified against the live org: 1.2h after the production rollout it reads
UNGRADED, before_send=0, 4.95/h vs a 4.71/h baseline — which is exactly right,
no device has the build yet.
Refs AGE-105
Co-authored-by: engineer <engineer@macbookpro.lan>
|
||
|
|
2cc284ecbe |
tools(sentry): org-wide volume report + measure on submitted, not accepted (#172)
* tools(sentry): add org-wide volume report and fix the metric we measure on The AGE-105 gate is a measured number, so it needs a repeatable query. It also needed a correction: `accepted` is the wrong headline. The org is over its error quota, so Sentry rejects nearly everything and `accepted` reads ~0 for every project - a blown org and a fixed one look identical on that column. The demand metric is `submitted` = accepted + rate_limited. scripts/sentry-volume-report.mjs takes named --window ranges and prints per-project submitted / accepted / rate_limited / client_discard plus the per-hour and projected per-month rate, so before/after comparisons run the exact same query instead of being re-derived by hand each time. Records the pre-rollout baseline in docs/analytics.md: opencode-mobile at 4.71/h (3,441/mo), 87% of the org's post-box-bot demand, from two windows that agree to within 0.2%. Co-Authored-By: Paperclip <noreply@paperclip.ing> * test(sentry): pin the noise gate against 90d of real production events The gate's unit tests prove it behaves as specified. Nothing proved the spec was aimed at the right targets. Replaying the actual 90d census of the opencode-mobile Sentry project (648 events, 11 issues) through the gate's own precedence shows 96.9% hard-dropped as transport noise and every observed crash class (OOM, ANR, IllegalStateException) still allowlisted -> ~87 events/month against a 1,500/month target. Also records two findings from measuring the org directly: * The error quota resets on the 4th. The 5,000-event month opened 2026-08-04 and was spent by 08-08; the org has accepted zero errors since. 2026-09-04 is the date the gate has to hold by, and it is why 'submitted' is the metric. * Server-side levers are unavailable on this plan. A per-key rate limit PUT returns HTTP 200 and silently discards the value (verified for three window sizes), custom inbound filters are absent, spike protection 403s. The client gate is the only control that exists, so its coverage is the whole margin. Refs AGE-105 Co-Authored-By: Paperclip <noreply@paperclip.ing> * tools(sentry): split client_discard by reason so gate drops aren't confused with quota backoff Raw client_discard cannot show whether the noise gate works. Today 100% of opencode-mobile's client_discard is ratelimit_backoff -- the SDK backing off a 429 because the ORG is over quota -- which rises when things get WORSE. Gate drops land in a different reason: @sentry/core records before_send when beforeSend returns null. - stats_v2 now groups by reason as well as project/outcome - the before_send vs ratelimit_backoff split always prints; --by-reason adds the full per-project reason table - before_send > 0 is install-share-independent, so it proves the gate is live on real devices days before a monthly rate can bend - documents that release-level segmentation is impossible while over quota: rate_limited events are never stored, so release tags stop (last value 0.4.12, 2026-08-08). Version share comes from Play, not Sentry. * ci(sentry): block a Play release whose bundle lost the noise gate The AGE-105 quota fix is entirely client-side (every server-side lever on this plan is dead), so the gate being *in the shipped binary* is the whole safety margin. That is also the one thing Sentry cannot tell us: while the org is over quota nothing is stored, release tags stop dead at 0.4.12, and a release:0.4.14 query returns empty in a way that reads like success. Grep the Hermes bundle inside the AAB instead, before the Play upload step: the gate's reason codes, the transport drop-list regex, the noise.dropped_since_last tag only applyNoiseGate() writes, and a baked-in DSN (a release built without EXPO_PUBLIC_SENTRY_DSN makes Sentry a silent no-op). Verified to discriminate on real artifacts - the v0.4.14 build now on Play production passes, pre-gate v0.4.13 fails all six markers. Also records the rejected alternative: persisting gate state across cold starts pays off only under ~94 active devices (2,633 session envelopes/7d vs a 6h cooldown), and the install base is above that. --------- Co-authored-by: engineer <engineer@macbookpro.lan> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
61f4b1177b |
fix(diagnostics): classify 401/403 as auth-failed so a wrong password stops the retry loop (#170)
AGE-107. The 498 `API Error: 401` events from one device were not a client token-refresh loop. Sentry breadcrumbs on the surviving events show a `touch` event immediately before every capture, at irregular human-paced intervals (87s, 199s, 69s, 5s, 61s, 66s) — a person re-tapping Connect, not a backoff timer. The app's automated loops were already correct: events.ts terminates the SSE reconnect loop on ApiAuthError (issue #76). What actually drove it: in v0.4.4 the connection probe counted any HTTP response as a successful health check, so a 401 was classified `ok` and shown to the user as "Health endpoint responded — connection actually works now" while their password was wrong. The user retried for two months. `requireOk` (#114, v0.4.8) stopped the false success, but 401 then fell into the generic `health-failed` bucket — "Likely wrong path, auth, or an old server version" — which still doesn't tell anyone to fix their password. - New `auth-failed` classification: a 401/403 from /global/health means the server is up and reachable and rejected the credentials. Its summary names the status, points at the password and OPENCODE_SERVER_USERNAME, and says the server is fine. It flows straight into the existing failure Alert on both the add and edit connection screens — which is where the password field is, i.e. the re-auth prompt. - It short-circuits before the root/internet probes can downgrade it: a 401 already proves the server answered. - `health-failed` copy no longer blames auth. - `connect auth-failed` joins the noise-gate drop-list. A wrong password is user config, unactionable server-side, already visible in the UI and already trended in PostHog as connection_failed{error_class:"unauthorized"}. `health-failed` and `tls-error` still report. Tests: 6 new (401/403 -> auth-failed, message content, root-unreachable does not override, 404/500/502 stay health-failed, health-failed copy drops "auth", noise gate drops `connect auth-failed` but not a raw `API Error: 401`). 263 pass, tsc --noEmit clean. Co-authored-by: engineer <engineer@macbookpro.lan> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7c8bc7d317 |
fix(sentry): gate non-actionable client noise before it leaves the device (#169)
opencode-mobile is now the org's #1 Sentry volume source (~4,500 events/mo against a 3,500/mo org quota, AGE-105). ~1,100 of those events are three non-defects: `connect timeout` (462), `connect server-unreachable` (157), and one device's `API Error: 401` token-refresh loop firing 498 times. Adds a pure, unit-tested noise gate (src/lib/sentry-noise.ts) wired into `beforeSend`, applying three layers cheapest-first: 1. Always-send allowlist — OOM/ANR/native/fatal crash classes bypass every limit. Quota is worthless if it silences real crashes. 2. Transport drop-list — hard drop for client-side network conditions. Hard, not sampled: the gate runs per-install, so "1 per device per day" would multiply by the install base straight back into thousands per month. 3. Dedup + rate cap — 6h per-fingerprint cooldown, ≤6 new fingerprints/h, ≤10 events/h, mirroring the openclaw-box-bot shim (AGE-55). Nothing is lost by the transport drop: those failures are already user-visible as connection UI and already trended, PII-free, as the PostHog `connection_failed{error_class}` event. captureDiagnostic() also short-circuits for those classifications so the event is never even built. Drops are auditable — the count since the last delivered event rides along as a `noise.dropped_since_last` tag. Replaying the observed 1,126-event hour through the gate yields 5 delivered events (1 auth report + 4 real OOMs). Tests: 18 new, 257 total passing; tsc --noEmit clean. Co-authored-by: engineer <engineer@macbookpro.lan> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1ea84f8236 |
chore(launch): reconcile README/store live-status + add demo-funnel analytics (#111)
Two scoped changes for the no-spend growth launch (Growth Launch Kit,
Notion page 3a1ac25eb49f81099cc9f3a4286c8ec4):
1. README.md and distribution/play-listing.md said Google Play was
"coming soon" / internal-testing-only, while distribution/retention-analysis.md
and the live play.google.com listing show it's actually public with 1K+
installs. Fixed the contradiction, added Google Play as a third install
channel, and added an accurate mention of the new offline demo mode
("Try a Demo" — reasoning, grep, diff, permission prompt, ~30s, no server)
matching what app/demo.tsx + src/lib/demo-script.ts actually render.
play-listing.md's stale pre-launch checklists are marked historical
instead of rewritten, so #83's ASO copy/keyword work is untouched.
2. Added the demo funnel's key metric (demo-completion, per the launch
kit) as four consent-gated PostHog events: demo_started,
demo_step_advanced, demo_completed, demo_exited_to_connect. Pure
property-derivation logic lives in src/lib/demo-analytics.ts (no
RN/PostHog imports, unit-tested with node --test, same pattern as
analytics-classify.ts) and is wired into app/demo.tsx's lifecycle.
Updated docs/analytics.md's event table and the privacy policy's event
list (distribution/privacy-policy.md + its two HTML mirrors) per the
repo's "new event requires a policy update" convention.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
63f3ec3c7e |
docs+consent: disclose activation analytics honestly across consent modal, privacy policy, and store docs (#81)
The app ships PostHog activation-funnel analytics gated behind the same consent flag as Sentry, but the consent modal, Settings toggle, privacy policy, and Play Data safety draft only mentioned crash reporting. Fix the disclosure everywhere: - TelemetryConsentModal: body + bullets + a11y labels now cover anonymous usage analytics (PostHog EU) alongside crash reports - Settings: toggle renamed 'Crash Reports & Usage Analytics', description names both Sentry and PostHog - Privacy policy (md + html + live gh-pages mirror): new section 3a with the full event/property table, PostHog EU destination, anonymous-ID statement, decline/revoke (drop-on-revoke) semantics; sections 4-7, 9 and the Apple nutrition-label addendum updated for analytics - play-listing.md: Data safety draft declares App interactions + Device or other IDs (opt-in, default OFF, shared with PostHog/Sentry) - docs/playstore.md: Data safety row flipped to re-verify with pointer to the new design record - docs/analytics.md: new design record — event schema, consent gating incl. buffered-event drop on revoke, disclosure surfaces to keep in sync, verification checklist (all TODO) - website privacy page metadata mentions analytics opt-in Closes #63 Claude-Session: https://claude.ai/code/session_01NJKAQ6HAikWGQK7PGZ5Y4E Co-authored-by: engineer <engineer@gray-knight-m1.local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |