A single byse.sx gateway hiccup (HTTP 502 "<!doctype html>..." or a mid-upload
ECONNRESET) was cascading through the whole failover chain and blacklisting every
account for the batch. Root cause: the upload-POST throw sites threw PLAIN errors
with no classification, so a 502 was treated as a GENERIC error -> mark-failed ->
emit('account-failed') -> failover to the next account (which hits the same
gateway) -> repeat until the chain is exhausted, then poison sibling files in the
batch via the blacklist. The screenshots showed exactly this: "Primär ... Fallback
#3" all failing with the same 502. byse did NOT change their API (the upload
contract still matches their docs and was live-probed: GET /upload/server -> POST
{key}+file -> files[0].filecode); these are transient infrastructure failures.
A 5xx gateway error is not an account fault: every account hits the same gateway,
so failing over is pointless and blacklisting is harmful. Fail open instead — retry
the SAME account and, if byse stays down, fail the file cleanly without touching the
account or the rotation cursor.
- upload-manager: _isTransientNetworkError() now honors an explicit err.transientNetwork
flag (checked before the empty-message guard) and, as a defensive fallback, matches
/HTTP 5\d\d/, Bad Gateway, Service Unavailable, Gateway Time-out. The flag is made
authoritative in _isFileRejectedError and _shouldSkipRetryOnAccountError (both return
early when it is set) so a 5xx whose HTML body happens to contain a rejection/auth
keyword can never be mis-binned as file/account. The post-rotation retry loop also
breaks on a transient error (parity with the primary loop) to avoid burning the
retry budget re-uploading on a fallback.
- hosters: the upload POST throw sites tag err.transientNetwork when statusCode >= 500
(non-JSON body and non-2xx-with-JSON) and when the 2xx status-envelope carries
status:500; 401/403/429 stay PLAIN so they remain account errors. The server-lookup
path (apiGet/getUploadServer) is hardened symmetrically: a 5xx there tags
transientNetwork and getUploadServer preserves the flag onto its wrapped error, so
the heuristic shouldRetryServerLookup() can no longer be defeated by a 5xx body that
contains an auth keyword.
- hosters: fixed a separate real bug surfaced during investigation — _fetchByseFileList
built the wrong URL https://api.byse.sx/api/file/list (the /api/ prefix is correct
only for doodstream's doodapi.co host; byse already carries the api. subdomain).
Live-probed: /api/file/list 302-redirects to the docs page, /file/list returns 200
JSON. The wrong path made the byse async-recovery poll silently dead (always []),
removing the safety net that reclaims a large file that registered despite a 502.
Now https://api.byse.sx/file/list.
ECONNRESET was already transient and is unchanged. doodstream's 2xx empty-form stays
hosterTransient (one attempt, no re-upload). vidmoly/clouddrop use their own uploaders
and are unaffected; doodstream/voe apiKey uploads share uploadFile and correctly
benefit from the same 5xx-is-infra logic.
Deferred follow-up: poll-first-on-5xx dedup (route a 5xx through the now-working
recovery poll before retrying, to reclaim a registered-but-502 file instead of
re-uploading). It does not add new dupe risk vs the prior cascade and a pure 502
means the backend was never reached, so retry-same-account is dupe-safe for it.
Investigated and reviewed by two multi-agent workflows (4-lens investigation with a
synthesized fix spec; 4-lens adversarial review with per-finding verification): no
API drift, 0 confirmed defects. Tests: classifier units (flag-above-message-guard,
5xx-transient, 4xx-stay-account), an end-to-end 502/503/500-envelope tag through
uploadFile, and a 502-retries-same-account-no-cascade integration case. Suite 311/311.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>