diff --git a/tasks/todo.md b/tasks/todo.md index 75aecd1..2325338 100644 --- a/tasks/todo.md +++ b/tasks/todo.md @@ -34,11 +34,42 @@ main-thread saturation / synchronous blocking, not DOM. - Lowering the virtual-row threshold below 200: the Blink benchmark shows <200 in-place updates are already sub-ms; no change warranted. -## OPEN — discriminator question to the user (do NOT declare victory blind) -With this user's real config (parallelCount 2 × 5 hosters ≈ 10 max concurrent), "100 concurrent" is only -reachable if the parallel counts were raised — otherwise "100" is the QUEUE size and only ~10 upload at once. -Ask: (a) is the lag specifically during clouddrop uploads? (b) did you raise the per-hoster/global parallel -counts above 2? If it's true high concurrency (dozens of simultaneous undici streams funneling decode + -progress callbacks through the one main JS thread), that needs a concurrency cap or a worker/child process — -NOT a micro-fix. The two shipped fixes are genuine improvements regardless; the answer decides whether a -bigger architectural change is the next step. +## DISCRIMINATOR ANSWERED (user, 2026-06-21) +(a) Lag NOT clouddrop-specific — other hosters. (b) Parallel counts RAISED deliberately (10+). +(c) 50+ uploading SIMULTANEOUSLY active. → This is the TRUE high-concurrency main-thread-funnel branch, +NOT clouddrop. v3.3.90 stands but does not target this user's case. + +## v3.3.91 — instrument first, don't refactor the upload core off elimination-reasoning +Advisor reframe: two LIVE hypotheses need OPPOSITE fixes — (A) main thread CPU-blocked (TLS/crypto/sync) → +event loop stalls → a cap/workers help; (B) main thread fine but IO-STARVED (libuv threadpool/sockets) → +loop stays responsive, uploads just queue → workers are WASTED, config fixes it. A worker/child-process +upload refactor touches throttle/rotation/abort/progress/credentials and is hard to reverse — DO NOT ship it +off sandbox elimination. One measurement splits the hypotheses and must run in the REAL app. + +SHIPPED (both reversible, zero upload-core refactor): +1. main.js: `perf_hooks.monitorEventLoopDelay({resolution:10})` enabled at startup; logged via logInfo every + ~5 s WHILE uploading (state==='uploading' && activeJobs>0) as + `eventloop-delay active=N mean=..ms p99=..ms max=..ms stddev=..ms threadpool=..`. Pure numbers, no secret + → does NOT touch the redaction surface. This is the GROUND TRUTH: high mean/p99 → CPU-blocking → workers + justified; low delay while uploads stall → IO-bound → workers wasted, threadpool/sockets is the fix. +2. main.js (first statement, before require('electron')): `UV_THREADPOOL_SIZE = env || '64'`. Default is 4; + every async uploader feeds undici from fs.createReadStream (+ clouddrop fh.read) and DNS getaddrinfo goes + through the same pool → 50 concurrent vs 4 threads = reads/DNS serialize 4-at-a-time = a hard cliff at a + small connection count = the "ab X connections" symptom. Threads are created lazily on demand → 64-max + costs nothing if unused (zero-risk, reversible). The advisor's prescribed one-env-var hypothesis test. + +CAVEAT (honest): synthetic sandbox benches could NOT confirm the threadpool is the bottleneck — pbkdf2 is +CPU-core-bound (masks pool size); DNS .invalid returns instantly; real-RTT DNS showed NO pool benefit because +WINDOWS serializes getaddrinfo via the OS DNS Client service (so on Windows the DNS half of the cliff is +masked by the resolver, though the fs-read half still benefits). This is exactly why the ELD number must come +from the user's real load, not the sandbox. Per-uploader undici Agent audit: clouddrop has a shared +module-level Agent (connections:50); doodstream/voe/vidmoly use the global dispatcher (pooled per origin, NO +per-call agent explosion) — so no agent fix needed. + +## NEXT (gated on the real-app ELD number + user's explicit nod) +User runs their 50-concurrent load once; the `eventloop-delay` log lines decide: +- mean/p99 HIGH (tens–hundreds ms) → CPU-blocked → propose worker_threads/child-process upload pool OR a + smart concurrency cap (WITH the user's nod — it's hard to reverse and touches credentials/abort/rotation). +- delay LOW while it still lags → IO-bound → threadpool bump already addresses it; if not, look at socket + caps / undici Agent connection limits / per-origin pooling, NOT workers. +Do NOT build the worker refactor before this number exists.