fix(diagnostics): classify 401/403 as auth-failed so a wrong password stops the retry loop (#170)
AGE-107. The 498 `API Error: 401` events from one device were not a client token-refresh loop. Sentry breadcrumbs on the surviving events show a `touch` event immediately before every capture, at irregular human-paced intervals (87s, 199s, 69s, 5s, 61s, 66s) — a person re-tapping Connect, not a backoff timer. The app's automated loops were already correct: events.ts terminates the SSE reconnect loop on ApiAuthError (issue #76). What actually drove it: in v0.4.4 the connection probe counted any HTTP response as a successful health check, so a 401 was classified `ok` and shown to the user as "Health endpoint responded — connection actually works now" while their password was wrong. The user retried for two months. `requireOk` (#114, v0.4.8) stopped the false success, but 401 then fell into the generic `health-failed` bucket — "Likely wrong path, auth, or an old server version" — which still doesn't tell anyone to fix their password. - New `auth-failed` classification: a 401/403 from /global/health means the server is up and reachable and rejected the credentials. Its summary names the status, points at the password and OPENCODE_SERVER_USERNAME, and says the server is fine. It flows straight into the existing failure Alert on both the add and edit connection screens — which is where the password field is, i.e. the re-auth prompt. - It short-circuits before the root/internet probes can downgrade it: a 401 already proves the server answered. - `health-failed` copy no longer blames auth. - `connect auth-failed` joins the noise-gate drop-list. A wrong password is user config, unactionable server-side, already visible in the UI and already trended in PostHog as connection_failed{error_class:"unauthorized"}. `health-failed` and `tls-error` still report. Tests: 6 new (401/403 -> auth-failed, message content, root-unreachable does not override, 404/500/502 stay health-failed, health-failed copy drops "auth", noise gate drops `connect auth-failed` but not a raw `API Error: 401`). 263 pass, tsc --noEmit clean. Co-authored-by: engineer <engineer@macbookpro.lan> Co-authored-by: Paperclip <noreply@paperclip.ing>
This commit is contained in:
@@ -81,20 +81,32 @@ Consent decides *whether* we report; the noise gate in `src/lib/sentry-noise.ts`
|
||||
often*. It exists because this app became the org's #1 Sentry volume source (~4,500
|
||||
events/month against a 3,500/month org quota) while ~1,100 of those events were three
|
||||
non-defects: `connect timeout`, `connect server-unreachable`, and one device's
|
||||
`API Error: 401` token-refresh loop firing 498 times.
|
||||
`API Error: 401` firing 498 times.
|
||||
|
||||
> **AGE-107 postscript.** That 401 storm was traced to a *human* retry loop, not a client
|
||||
> token-refresh loop. In v0.4.4 the connection probe scored **any** HTTP response as a
|
||||
> success, so a 401 was reported to the user as "Health endpoint responded — connection
|
||||
> actually works now" while their password was wrong. They re-tapped Connect for two months
|
||||
> (Sentry breadcrumbs show a `touch` event before every single capture, at irregular
|
||||
> human-paced intervals). `requireOk` in `diagnostics.ts` (v0.4.8) stopped the false
|
||||
> success; `auth-failed` now gives it its own actionable message and drop-list entry.
|
||||
> The client's automated loops were never at fault — `events.ts` already terminates the SSE
|
||||
> reconnect loop on `ApiAuthError` (issue #76).
|
||||
|
||||
`beforeSend` applies three layers, cheapest first:
|
||||
|
||||
| Layer | Rule | Effect |
|
||||
|---|---|---|
|
||||
| Always-send allowlist | OOM / ANR / native / `IllegalStateException` / `NullPointerException` / fatal level / unhandled mechanism | Bypasses every limit below — quota is worthless if it silences real crashes |
|
||||
| Transport drop-list | `connect timeout\|server-unreachable\|no-internet\|malformed-url`, `Network request failed`, `Request timed out after`, `ECONN*`/`ETIMEDOUT`… | Hard drop. Not sampled: the gate is per-install, so even 1/device/day multiplies by the install base back into thousands/month |
|
||||
| Transport drop-list | `connect timeout\|server-unreachable\|no-internet\|malformed-url\|auth-failed`, `Network request failed`, `Request timed out after`, `ECONN*`/`ETIMEDOUT`… | Hard drop. Not sampled: the gate is per-install, so even 1/device/day multiplies by the install base back into thousands/month |
|
||||
| Dedup + rate cap | per-fingerprint cooldown 6h, ≤6 new fingerprints/h, ≤10 events/h (mirrors the `openclaw-box-bot` shim, AGE-55) | Turns a retry loop into one report and caps any future regression |
|
||||
|
||||
Nothing is lost by the transport drop: those failures are already shown to the user as
|
||||
connection UI **and** already trended, PII-free, as the PostHog `connection_failed` event with
|
||||
an `error_class` property (`src/lib/analytics-classify.ts`). Sentry was paying per event for a
|
||||
graph we already have.
|
||||
an `error_class` property (`src/lib/analytics-classify.ts` — a 401 lands in `unauthorized`).
|
||||
Sentry was paying per event for a graph we already have. `connect health-failed` and
|
||||
`connect tls-error` are deliberately **not** dropped: a box that answers but is unhealthy, or
|
||||
a broken certificate, is actionable.
|
||||
|
||||
Dropped-event counts are not silent — the number dropped since the last delivered event rides
|
||||
along as a `noise.dropped_since_last` tag, so the saving is auditable from Sentry itself.
|
||||
|
||||
Reference in New Issue
Block a user