The v3.3.96 eventloop-delay logs from a real 70-connection run pinpointed the
high-concurrency lag to the file-read path. At a CONSTANT active-job count the
process flips between two clean regimes:
HEALTHY ELD ~11ms, rss 268-308MB: SimpleWriteWrap ≈ active, FSReqCallback ≈ 0-1
BLOCKED ELD 49-217ms, rss 540-610MB: FSReqCallback ≈ active (62-71 reads in
flight), SimpleWriteWrap ≈ 0-4
Both the ELD spike and the rss balloon track FSReqCallback (libuv-threadpool file
reads) exactly — not crypto, not the renderer, not GC alone. All five uploaders
read identically via fs.createReadStream({highWaterMark: 256KB}); byse/doodstream/
voe run through the generic uploadFile in hosters.js (no dedicated module).
This build is both a candidate fix and a discriminator, per advisor review:
- UV_THREADPOOL_SIZE 64→8 (main.js:1). One reversible line, NOT an upload cap —
70 uploads still run. 8 concurrent 256KB reads sustain ~100MB/s, far above the
41MB/s aggregate, so it cannot bottleneck throughput even on the slow VM disk.
Strong suspicion that tp=64 made it worse: it removed the natural read-
serialization (default 4 threads) and let all 70 streams' reads fire at once,
flooding the loop with completion callbacks in lock-step bursts.
- ELD line now also logs heap=heapUsed ext=external ab=arrayBuffers and
gc=/gcTotal=/gcMax=ms (PerformanceObserver entryTypes:['gc'], reset per window).
The 70×256KB ≈ 18MB of read buffers cannot account for the ~300MB rss swing —
that is heap/object churn, so GC must be measured directly.
Decision rule for the next real-run log:
- ELD drops with tp=8 → read over-parallelism confirmed (keep 8 or add
a dedicated read-semaphore).
- ELD high + GC pauses align → heap churn, hunt the allocator.
- ELD high + GC flat → causation was reversed, pivot.
No upload-behavior change; the concurrency cap the user explicitly rejected is
untouched.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
52-agent audit of the v3.3.84/85 remote-diagnostics code. App is healthy for this
user's actual usage. Shipped the two zero-redaction-surface fixes (ws maxPayload,
sendToClient guard). Deferred the cold-path server_health O(historySize) freeze
because the fix refactors the credential-redaction collectors (leaked twice before).
Lesson: "not persistence, so safe" is a fallacy — the redaction layer is an equally
catastrophic guarantee surface (silent secret leak). At an audit goal, find+document
is the deliverable; don't cut into a twice-leaked redaction pipeline under a Stop hook.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Records the 18-agent line-by-line audit of every line written this session: the
reported lag was v3.3.87 (recent-panel cliff); the audit found no second cause
affecting this user. doodstream _debugLog sync-fs gated (shipped). config-store
load()/serialize history-scaling costs are real but sub-ms at this user's scale
and the fix is risky persistence surgery — deferred, documented with measurements.
Lesson: "audit every line" = look + measure + risk-appropriate decision, NOT
fix-everything; the load() perf win and its corruption risk are the same coin
(shared batch refs), so there is no safe version — defer, don't ship.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Records the verified root cause + the dismissed-with-evidence findings:
- _sessionFileKeys "separator mismatch" = FALSE POSITIVE (real U+0001 chars, verifier Read rendered them invisibly)
- queueJobs O(N) per-render scan = real but Blink-measured <0.1ms at 3000 jobs -> skipped
- standing 2000-row relayout = median 0.4ms -> no virtualization needed
- doodstream sync _debugLog = constant freeze, deferred to a separate change
Lesson: measure magnitudes at the real artifact before fixing a "scales-with-X" cause;
profile in real Blink (Playwright) not jsdom; verify multi-agent findings against primary
evidence (char-code dump caught the control-char false positive).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Captures the two redaction leaks from the remote-diagnostics build: a hoster-
returned opaque token surviving because value-scrub only covers stored config
secrets, and the get_config_redacted/get_queue_state default paths skipping
pattern-scrub entirely. The discriminator the first E2E missed, plus the rule:
exercise each collector with its DEFAULT args and seed fixtures with a NON-config
secret so you test pattern-scrub, not value-scrub by accident.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Three lessons from the second adversarial bug hunt:
- A dedup/skip key has more re-add paths than "in-place vs rebuild" suggests; enumerate every .add site and every selectedFiles re-add, cross-tabulate which clears the key. retrySelectedJobs (retry-of-done) and the folder-monitor branch were the two missed paths.
- Per-cell (file|hoster) deletion needs its own persisted suppression set with the OPPOSITE lifecycle to the completed-key (set on delete, cleared on re-add); overloading the completed-key would break re-upload-after-delete.
- Re-confirming a user-facing change is "real" is not the user consenting to the behavior change; force the decision via one AskUserQuestion or park it, do not re-raise it as ambient worry.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Captures the v3.3.82 near-miss the advisor caught: persisting _completedUploadKeys
risked silently no-op'ing a re-upload IF any retry path rebuilt jobs via
buildQueuePreview. Both retry paths (retrySelectedJobs, _retryFailedFromBuckets)
mutate jobs in place and are immune; only fresh modal re-adds go through the
preview and clear the key. Rule: read every 'erneut/retry/reupload' handler
before touching a skip-guard, decide mutate-vs-rebuild per handler.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Capture the v3.3.78 byse lesson: a truncated upload host (sNNNN.filemoon, no TLD)
must not be "repaired" by guessing a TLD — every resolving candidate was a parked /
squatter domain (parklogic PTR, shared catch-all cert). Verify PTR + cert SAN before
trusting a domain; discard the obviously-broken value and let the existing
retry/cache machinery find a valid one.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Capture the v3.3.77 byse lesson: an unclassified HTTP 502/gateway error cascaded
through the failover chain and blacklisted every account; the fix is an explicit
transientNetwork flag at the throw site (authoritative over message-keyword
heuristics) + symmetric server-lookup hardening, and throw-site tagging must be
tested end-to-end, not just flag-injected in a mock.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Capture the v3.3.74→v3.3.75 lesson: a fresh-per-call round-robin picker
silently no-ops under the folder-monitor drip-feed (each add-jobs-to-batch
call rebuilt the cursor from 0 → account 1 every time). Distribution state
that must be fair across multiple call entry points has to outlive the calls,
and across sessions when the goal is a rolling quota.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Reproduced from a real saved config: pendingQueue held 4 'preview' jobs (one
file across 4 hosters); the queue saved + restored correctly. But
_autoDeduplicateFromLog (runs at init after restore) removed jobs whose
fileName|hoster appeared ANYWHERE in the lifetime fileuploader.log, regardless
of status — so all 4 pending previews were deleted and the queue showed the
empty "Dateien hierhin ziehen" state. Looked update-specific only because the
server restarts on update; a plain restart did the same.
- New lib/queue-dedup.js (pure, dual CJS/window export like queue-prune.js):
partitionRestoredJobsByLog drops ONLY 'done' jobs that match the log. Pending
(preview/queued) and failed (error/aborted) jobs always survive — they're
intentional queued work (often a deliberate re-upload of a previously
uploaded file). Manual importUploadLog stays separate/explicit.
- renderer wires it in; index.html loads the module before app.js.
- Tests: 5 cases incl. the exact reproduced scenario (4 previews all in log ->
0 removed). Full suite 162/162.
Verified against the user's real electron-config.json + fileuploader.log: old
logic removed 4/4 (empty queue), new logic removes 0/4 (queue preserved).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Three related gaps closed so one full byse account stops wasting
attempts on every subsequent job and later-added accounts get picked
up without an app restart.
1. Pre-job-swap moved BEHIND the semaphore acquire. At scale (500 jobs
/ 1 slot) every worker was checking _failedAccounts at spawn time
before the first upload had even tried — so none of them saw the
failed state. Now each worker re-checks right before its first
upload attempt.
2. save-config IPC handler re-resolves fallbacks for any account that
is already in _failedAccounts but has no override set. Previously
account-failed only fired once per account, so a config change
after the first mark-failed was silently ignored and the batch
stayed stuck on the dead account until the app restarted.
3. UploadManager exposes getFailedAccountKeys() and getOverride(hoster)
so main.js can drive the late re-resolve without poking private
fields.
4 new tests: pre-job-swap after semaphore, getters contract, fresh
manager resets learned state, late-added fallback is honored by
subsequent jobs. 80/80 green.
Byse rejects uploads with status like "not enough disk space on your
account" when the account's storage is exhausted. The parser was
flagging every non-OK status as err.fileRejected=true, and the upload-
manager classifier additionally matched the generic "lehnte Datei ab"
prefix as file-rejected. Result: rotation was skipped on a full account
and every subsequent file failed on the same dead account.
- hosters.js: byse parser now distinguishes account-level phrases
(disk space / storage / quota / insufficient / account full) and sets
err.accountError=true for those. File-specific failures (Duplicate,
wrong format, size) keep err.fileRejected=true.
- upload-manager.js: _isFileRejectedError no longer matches the generic
"lehnte Datei ab" prefix and short-circuits when err.accountError is
true. _shouldSkipRetryOnAccountError honors the flag and has added
regex patterns as a safety net.
- Tests: 5 new unit tests covering disk-space/account-level/duplicate
and the accountError-wins-over-fileRejected precedence.
retrySelectedJobs() was calling renderQueueTable + updateQueueActionButtons
+ updateStatusBar and then immediately awaiting startSelectedUpload(),
which runs the exact same trio right after. At 500+ failed jobs the
double render/sort/button-refresh freezes the UI for several seconds
after clicking "Erneut versuchen".
Drop the outer render trio — startSelectedUpload's one is enough. The
inner call sees the freshly-mutated job state in the same tick, so the
visible result is identical with half the work.