docs(tasks): high-concurrency discriminator answered — instrument-first (v3.3.91), gate worker refactor on real-app ELD number

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
Administrator 2026-06-21 04:10:47 +02:00
parent 2fd26add1a
commit e135655c95

View File

@ -34,11 +34,42 @@ main-thread saturation / synchronous blocking, not DOM.
- Lowering the virtual-row threshold below 200: the Blink benchmark shows <200 in-place updates are already - Lowering the virtual-row threshold below 200: the Blink benchmark shows <200 in-place updates are already
sub-ms; no change warranted. sub-ms; no change warranted.
## OPEN — discriminator question to the user (do NOT declare victory blind) ## DISCRIMINATOR ANSWERED (user, 2026-06-21)
With this user's real config (parallelCount 2 × 5 hosters ≈ 10 max concurrent), "100 concurrent" is only (a) Lag NOT clouddrop-specific — other hosters. (b) Parallel counts RAISED deliberately (10+).
reachable if the parallel counts were raised — otherwise "100" is the QUEUE size and only ~10 upload at once. (c) 50+ uploading SIMULTANEOUSLY active. → This is the TRUE high-concurrency main-thread-funnel branch,
Ask: (a) is the lag specifically during clouddrop uploads? (b) did you raise the per-hoster/global parallel NOT clouddrop. v3.3.90 stands but does not target this user's case.
counts above 2? If it's true high concurrency (dozens of simultaneous undici streams funneling decode +
progress callbacks through the one main JS thread), that needs a concurrency cap or a worker/child process — ## v3.3.91 — instrument first, don't refactor the upload core off elimination-reasoning
NOT a micro-fix. The two shipped fixes are genuine improvements regardless; the answer decides whether a Advisor reframe: two LIVE hypotheses need OPPOSITE fixes — (A) main thread CPU-blocked (TLS/crypto/sync) →
bigger architectural change is the next step. event loop stalls → a cap/workers help; (B) main thread fine but IO-STARVED (libuv threadpool/sockets) →
loop stays responsive, uploads just queue → workers are WASTED, config fixes it. A worker/child-process
upload refactor touches throttle/rotation/abort/progress/credentials and is hard to reverse — DO NOT ship it
off sandbox elimination. One measurement splits the hypotheses and must run in the REAL app.
SHIPPED (both reversible, zero upload-core refactor):
1. main.js: `perf_hooks.monitorEventLoopDelay({resolution:10})` enabled at startup; logged via logInfo every
~5 s WHILE uploading (state==='uploading' && activeJobs>0) as
`eventloop-delay active=N mean=..ms p99=..ms max=..ms stddev=..ms threadpool=..`. Pure numbers, no secret
→ does NOT touch the redaction surface. This is the GROUND TRUTH: high mean/p99 → CPU-blocking → workers
justified; low delay while uploads stall → IO-bound → workers wasted, threadpool/sockets is the fix.
2. main.js (first statement, before require('electron')): `UV_THREADPOOL_SIZE = env || '64'`. Default is 4;
every async uploader feeds undici from fs.createReadStream (+ clouddrop fh.read) and DNS getaddrinfo goes
through the same pool → 50 concurrent vs 4 threads = reads/DNS serialize 4-at-a-time = a hard cliff at a
small connection count = the "ab X connections" symptom. Threads are created lazily on demand → 64-max
costs nothing if unused (zero-risk, reversible). The advisor's prescribed one-env-var hypothesis test.
CAVEAT (honest): synthetic sandbox benches could NOT confirm the threadpool is the bottleneck — pbkdf2 is
CPU-core-bound (masks pool size); DNS .invalid returns instantly; real-RTT DNS showed NO pool benefit because
WINDOWS serializes getaddrinfo via the OS DNS Client service (so on Windows the DNS half of the cliff is
masked by the resolver, though the fs-read half still benefits). This is exactly why the ELD number must come
from the user's real load, not the sandbox. Per-uploader undici Agent audit: clouddrop has a shared
module-level Agent (connections:50); doodstream/voe/vidmoly use the global dispatcher (pooled per origin, NO
per-call agent explosion) — so no agent fix needed.
## NEXT (gated on the real-app ELD number + user's explicit nod)
User runs their 50-concurrent load once; the `eventloop-delay` log lines decide:
- mean/p99 HIGH (tenshundreds ms) → CPU-blocked → propose worker_threads/child-process upload pool OR a
smart concurrency cap (WITH the user's nod — it's hard to reverse and touches credentials/abort/rotation).
- delay LOW while it still lags → IO-bound → threadpool bump already addresses it; if not, look at socket
caps / undici Agent connection limits / per-origin pooling, NOT workers.
Do NOT build the worker refactor before this number exists.