fix(diagnostics): classify 401/403 as auth-failed so a wrong password stops the retry loop (#170)

AGE-107. The 498 `API Error: 401` events from one device were not a client
token-refresh loop. Sentry breadcrumbs on the surviving events show a `touch`
event immediately before every capture, at irregular human-paced intervals
(87s, 199s, 69s, 5s, 61s, 66s) — a person re-tapping Connect, not a backoff
timer. The app's automated loops were already correct: events.ts terminates
the SSE reconnect loop on ApiAuthError (issue #76).

What actually drove it: in v0.4.4 the connection probe counted any HTTP
response as a successful health check, so a 401 was classified `ok` and shown
to the user as "Health endpoint responded — connection actually works now"
while their password was wrong. The user retried for two months. `requireOk`
(#114, v0.4.8) stopped the false success, but 401 then fell into the generic
`health-failed` bucket — "Likely wrong path, auth, or an old server version" —
which still doesn't tell anyone to fix their password.

- New `auth-failed` classification: a 401/403 from /global/health means the
  server is up and reachable and rejected the credentials. Its summary names
  the status, points at the password and OPENCODE_SERVER_USERNAME, and says
  the server is fine. It flows straight into the existing failure Alert on
  both the add and edit connection screens — which is where the password
  field is, i.e. the re-auth prompt.
- It short-circuits before the root/internet probes can downgrade it: a 401
  already proves the server answered.
- `health-failed` copy no longer blames auth.
- `connect auth-failed` joins the noise-gate drop-list. A wrong password is
  user config, unactionable server-side, already visible in the UI and
  already trended in PostHog as connection_failed{error_class:"unauthorized"}.
  `health-failed` and `tls-error` still report.

Tests: 6 new (401/403 -> auth-failed, message content, root-unreachable does
not override, 404/500/502 stay health-failed, health-failed copy drops "auth",
noise gate drops `connect auth-failed` but not a raw `API Error: 401`).
263 pass, tsc --noEmit clean.

Co-authored-by: engineer <engineer@macbookpro.lan>
Co-authored-by: Paperclip <noreply@paperclip.ing>
This commit is contained in:
Den
2026-08-14 07:14:36 -07:00
committed by GitHub
parent 950080eae5
commit 61f4b1177b
6 changed files with 120 additions and 17 deletions

View File

@@ -81,20 +81,32 @@ Consent decides *whether* we report; the noise gate in `src/lib/sentry-noise.ts`
often*. It exists because this app became the org's #1 Sentry volume source (~4,500
events/month against a 3,500/month org quota) while ~1,100 of those events were three
non-defects: `connect timeout`, `connect server-unreachable`, and one device's
`API Error: 401` token-refresh loop firing 498 times.
`API Error: 401` firing 498 times.
> **AGE-107 postscript.** That 401 storm was traced to a *human* retry loop, not a client
> token-refresh loop. In v0.4.4 the connection probe scored **any** HTTP response as a
> success, so a 401 was reported to the user as "Health endpoint responded — connection
> actually works now" while their password was wrong. They re-tapped Connect for two months
> (Sentry breadcrumbs show a `touch` event before every single capture, at irregular
> human-paced intervals). `requireOk` in `diagnostics.ts` (v0.4.8) stopped the false
> success; `auth-failed` now gives it its own actionable message and drop-list entry.
> The client's automated loops were never at fault — `events.ts` already terminates the SSE
> reconnect loop on `ApiAuthError` (issue #76).
`beforeSend` applies three layers, cheapest first:
| Layer | Rule | Effect |
|---|---|---|
| Always-send allowlist | OOM / ANR / native / `IllegalStateException` / `NullPointerException` / fatal level / unhandled mechanism | Bypasses every limit below — quota is worthless if it silences real crashes |
| Transport drop-list | `connect timeout\|server-unreachable\|no-internet\|malformed-url`, `Network request failed`, `Request timed out after`, `ECONN*`/`ETIMEDOUT`… | Hard drop. Not sampled: the gate is per-install, so even 1/device/day multiplies by the install base back into thousands/month |
| Transport drop-list | `connect timeout\|server-unreachable\|no-internet\|malformed-url\|auth-failed`, `Network request failed`, `Request timed out after`, `ECONN*`/`ETIMEDOUT`… | Hard drop. Not sampled: the gate is per-install, so even 1/device/day multiplies by the install base back into thousands/month |
| Dedup + rate cap | per-fingerprint cooldown 6h, ≤6 new fingerprints/h, ≤10 events/h (mirrors the `openclaw-box-bot` shim, AGE-55) | Turns a retry loop into one report and caps any future regression |
Nothing is lost by the transport drop: those failures are already shown to the user as
connection UI **and** already trended, PII-free, as the PostHog `connection_failed` event with
an `error_class` property (`src/lib/analytics-classify.ts`). Sentry was paying per event for a
graph we already have.
an `error_class` property (`src/lib/analytics-classify.ts` — a 401 lands in `unauthorized`).
Sentry was paying per event for a graph we already have. `connect health-failed` and
`connect tls-error` are deliberately **not** dropped: a box that answers but is unhealthy, or
a broken certificate, is actionable.
Dropped-event counts are not silent — the number dropped since the last delivered event rides
along as a `noise.dropped_since_last` tag, so the saving is auditable from Sentry itself.