docs(tasks): high-concurrency discriminator answered — instrument-first (v3.3.91), gate worker refactor on real-app ELD number
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
parent
2fd26add1a
commit
e135655c95
@ -34,11 +34,42 @@ main-thread saturation / synchronous blocking, not DOM.
|
||||
- Lowering the virtual-row threshold below 200: the Blink benchmark shows <200 in-place updates are already
|
||||
sub-ms; no change warranted.
|
||||
|
||||
## OPEN — discriminator question to the user (do NOT declare victory blind)
|
||||
With this user's real config (parallelCount 2 × 5 hosters ≈ 10 max concurrent), "100 concurrent" is only
|
||||
reachable if the parallel counts were raised — otherwise "100" is the QUEUE size and only ~10 upload at once.
|
||||
Ask: (a) is the lag specifically during clouddrop uploads? (b) did you raise the per-hoster/global parallel
|
||||
counts above 2? If it's true high concurrency (dozens of simultaneous undici streams funneling decode +
|
||||
progress callbacks through the one main JS thread), that needs a concurrency cap or a worker/child process —
|
||||
NOT a micro-fix. The two shipped fixes are genuine improvements regardless; the answer decides whether a
|
||||
bigger architectural change is the next step.
|
||||
## DISCRIMINATOR ANSWERED (user, 2026-06-21)
|
||||
(a) Lag NOT clouddrop-specific — other hosters. (b) Parallel counts RAISED deliberately (10+).
|
||||
(c) 50+ uploading SIMULTANEOUSLY active. → This is the TRUE high-concurrency main-thread-funnel branch,
|
||||
NOT clouddrop. v3.3.90 stands but does not target this user's case.
|
||||
|
||||
## v3.3.91 — instrument first, don't refactor the upload core off elimination-reasoning
|
||||
Advisor reframe: two LIVE hypotheses need OPPOSITE fixes — (A) main thread CPU-blocked (TLS/crypto/sync) →
|
||||
event loop stalls → a cap/workers help; (B) main thread fine but IO-STARVED (libuv threadpool/sockets) →
|
||||
loop stays responsive, uploads just queue → workers are WASTED, config fixes it. A worker/child-process
|
||||
upload refactor touches throttle/rotation/abort/progress/credentials and is hard to reverse — DO NOT ship it
|
||||
off sandbox elimination. One measurement splits the hypotheses and must run in the REAL app.
|
||||
|
||||
SHIPPED (both reversible, zero upload-core refactor):
|
||||
1. main.js: `perf_hooks.monitorEventLoopDelay({resolution:10})` enabled at startup; logged via logInfo every
|
||||
~5 s WHILE uploading (state==='uploading' && activeJobs>0) as
|
||||
`eventloop-delay active=N mean=..ms p99=..ms max=..ms stddev=..ms threadpool=..`. Pure numbers, no secret
|
||||
→ does NOT touch the redaction surface. This is the GROUND TRUTH: high mean/p99 → CPU-blocking → workers
|
||||
justified; low delay while uploads stall → IO-bound → workers wasted, threadpool/sockets is the fix.
|
||||
2. main.js (first statement, before require('electron')): `UV_THREADPOOL_SIZE = env || '64'`. Default is 4;
|
||||
every async uploader feeds undici from fs.createReadStream (+ clouddrop fh.read) and DNS getaddrinfo goes
|
||||
through the same pool → 50 concurrent vs 4 threads = reads/DNS serialize 4-at-a-time = a hard cliff at a
|
||||
small connection count = the "ab X connections" symptom. Threads are created lazily on demand → 64-max
|
||||
costs nothing if unused (zero-risk, reversible). The advisor's prescribed one-env-var hypothesis test.
|
||||
|
||||
CAVEAT (honest): synthetic sandbox benches could NOT confirm the threadpool is the bottleneck — pbkdf2 is
|
||||
CPU-core-bound (masks pool size); DNS .invalid returns instantly; real-RTT DNS showed NO pool benefit because
|
||||
WINDOWS serializes getaddrinfo via the OS DNS Client service (so on Windows the DNS half of the cliff is
|
||||
masked by the resolver, though the fs-read half still benefits). This is exactly why the ELD number must come
|
||||
from the user's real load, not the sandbox. Per-uploader undici Agent audit: clouddrop has a shared
|
||||
module-level Agent (connections:50); doodstream/voe/vidmoly use the global dispatcher (pooled per origin, NO
|
||||
per-call agent explosion) — so no agent fix needed.
|
||||
|
||||
## NEXT (gated on the real-app ELD number + user's explicit nod)
|
||||
User runs their 50-concurrent load once; the `eventloop-delay` log lines decide:
|
||||
- mean/p99 HIGH (tens–hundreds ms) → CPU-blocked → propose worker_threads/child-process upload pool OR a
|
||||
smart concurrency cap (WITH the user's nod — it's hard to reverse and touches credentials/abort/rotation).
|
||||
- delay LOW while it still lags → IO-bound → threadpool bump already addresses it; if not, look at socket
|
||||
caps / undici Agent connection limits / per-origin pooling, NOT workers.
|
||||
Do NOT build the worker refactor before this number exists.
|
||||
|
||||
Loading…
Reference in New Issue
Block a user