Closes the last measurement gap. The main process was already fully instrumented (every ipcMain handler timed at >=50ms, a 100ms main-thread long-task monitor with channel attribution, config load/serialize timing). Account switches were already covered too: switchAccount is a trivial synchronous Map-set and the rotation work is async, so a switch cannot block the main thread, and any block that did occur would surface in the long-task monitor. The real gap was the renderer side: the renderer-perf line was gated on active uploads (idle clicks were never logged) and only reported an aggregate longtask count — no per-interaction latency and no element attribution. So a switch/sort/tab that janked in the renderer (not main) was invisible. renderer/app.js (additive, self-silencing, wrapped in try/catch): - An Event Timing observer (PerformanceObserver type:'event', durationThreshold:50, buffered) logs `renderer-interaction <type> dur=Xms proc=Yms target=<el>` for every user interaction whose latency exceeds 50ms — always on, idle or under load — and names the element (id / first class / data-action / aria-label / title). This is the direct click->reaction latency the user feels. - The longtask observer now also logs `renderer-longtask dur=Xms` immediately for any single renderer long task >=100ms, regardless of upload state. 405 tests pass; clean Electron boot. Every action — main or renderer, idle or under load — now names itself in the log if it is slow. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
295 lines
23 KiB
Markdown
295 lines
23 KiB
Markdown
# v3.3.100 — close the LAST measurement gap: renderer interaction timing (switches/clicks)
|
||
|
||
User asked "haben wir wirklich ALLES gemessen, auch switches/wechsel?". Audit: main-side was already
|
||
fully covered (IPC wrapper ≥50ms on every handler, main-longtask >100ms with lastIpc, config instrument);
|
||
account switchAccount is a trivial sync Map-set + the rotation work is async (can't block) → already covered.
|
||
The REAL gap was RENDERER-side: renderer-perf was upload-gated (idle clicks unmeasured) and only aggregate
|
||
(no per-interaction latency, no element attribution). Closed it (renderer/app.js, additive, self-silencing):
|
||
- Event Timing API observer (`type:'event', durationThreshold:50, buffered`) → `renderer-interaction <type>
|
||
dur=Xms proc=Yms target=<el>` for EVERY UI interaction ≥50ms (switch/sort-header/tab/button), always-on,
|
||
names the element (id/class/data-action/aria-label). The direct "click→reaction" latency.
|
||
- Idle renderer-longtask logging: any longtask ≥100ms logged immediately (`renderer-longtask dur=Xms`),
|
||
not just during uploads.
|
||
405 tests pass, clean boot. Now EVERY action — main or renderer, idle or under load — names itself if slow.
|
||
|
||
---
|
||
|
||
# v3.3.99 — THE KILL: 38.5MB config-thrash → history split out of the hot config + full instrumentation
|
||
|
||
THE ROOT CAUSE (from v3.3.98's instrument, the real "bread"): electron-config.json was **38.5MB** and got
|
||
loaded/cloned/serialized **137× in 73s** on the main thread (140-592ms each) = **~47% main-thread occupancy**
|
||
→ that IS the 1-2s button lag. The 1MB read-ahead was irrelevant against it. NOT the queue (`queue=undefined`
|
||
was a LOGGING BUG: read `.length` on the pendingQueue OBJECT); the bulk is HISTORY — each batch-done appended
|
||
the full per-file result list (`summary.files` w/ per-hoster URLs), 75 batches, default historyRetention='all'
|
||
never prunes → unbounded. (5-agent workflow wz2g4bwka + adversarial verify; bench fixture confirmed 185MB =
|
||
100MB history/30000 entries.)
|
||
|
||
Why writes drove the storm: save-global-settings (queue-persist) did 2 loads + 1 serialize per call, and
|
||
_atomicWrite NULLS the cache → next load is a full 38.5MB reparse. Even cache HITS structuredCloned 38.5MB.
|
||
Per-job upload path makes ZERO config calls (selectUploadAuth/rotation take config by param) — pending=1280
|
||
is NOT the driver.
|
||
|
||
THE FIX (history split — kills clone AND serialize AND reparse at once; advisor-gated, adversarially verified):
|
||
- History moved to its OWN file **electron-history.json** (lib/config-store.js). _migrateHistory() runs ONCE
|
||
at init (packaged only), fail-safe: write history.json + fsync + verify count BEFORE the config is ever
|
||
allowed to drop history, keep a permanent `electron-config.json.pre-history-split.bak`. _loadImpl returns
|
||
`history:[]` when migrated → cached result is tiny → clones cheap; config file shrinks to ~KB on the first
|
||
save() → reparse cheap; _serializeForDisk writes ~KB → serialize cheap. loadHistory/appendHistory/
|
||
pruneHistory/clearHistory redirected to history.json (own write-queue, no-clobber guard); legacy config
|
||
path kept as fallback when migration fails. get-history/export-history go through loadHistory.
|
||
- REAL 194MB-FIXTURE VALIDATION: migrate 1.5s (1×), all 30000 entries preserved + .bak; load() 631ms cold
|
||
(1×) → **0.1ms** after strip; save() strips config **185MB→2.1KB**; loadHistory() still 30000. RESULT:
|
||
data-preserved=true, hotpath-fast=true.
|
||
- DROPPED per advisor (one-variable + risk): loadShallow (moot after split), Fix#2 cache-repopulate (dead
|
||
gate), Fix#4 resolution-cache (stale-cache → rotation/byse failover-regression class — the one thing that
|
||
could SILENTLY corrupt uploads). Fix#3 (this split) was the only complete fix.
|
||
|
||
"MEASURE EVERYTHING" instrumentation (user demand) — all additive, threshold-gated, MHU_PERF=0 to disable:
|
||
- IPC handler wrapper (monkey-patch ipcMain.handle/.on) → `ipc <channel> wall=Xms sync=Yms` ≥50ms = the
|
||
button-press→response latency, hardened so a logging throw can never break IPC (Promise.resolve(p).finally).
|
||
- Main-process long-task drift monitor (setInterval 100ms) → `main-longtask blocked=Xms lastIpc=… gc=… gcMax=…`
|
||
for any single main-thread turn >100ms (catches GC, fs scans, serialize the IPC wrapper structurally can't see).
|
||
- config-store: caller attribution `via=<stack>` + `wqDepth=` on config-load/config-serialize lines; FIXED the
|
||
`queue=` logging bug (now reads pendingQueue.queueJobs.length).
|
||
|
||
405 tests pass (9 new migration tests covering preserve-count, round-trip, save()-never-loses-history,
|
||
crash-window fallback, idempotency). Clean Electron boot (9s, no errors). No repo pollution (migration packaged-only).
|
||
NEXT LOG must show: config-load/config-serialize wall= drop to single digits (or vanish), main-longtask rare,
|
||
ipc lines name any residual. If a residual remains it's pendingQueue (own follow-up, not a regression).
|
||
|
||
---
|
||
|
||
# v3.3.98 — read-burst absorption (1MB hwm) + persist/load instrument; B (renderer) DEFERRED
|
||
|
||
v3.3.97 (threadpool 64→8) was a DECISIVE win: mean ELD 200ms→~11ms at 70 active (18×), renderer healthy
|
||
14/15 windows. User: "ganz flüssig isses noch nicht". A 5-agent ultracode workflow + adversarial verify
|
||
localized the RESIDUAL to TWO distinct, measured spike sources (full data: subagents output wjskjo1xk):
|
||
|
||
1. READ-BURSTS (tail W13/14/15, 15:18:53-19:04): FSReqCallback 66/70/46 vs threadpool=8 (~8.75× queue
|
||
depth), SimpleWriteWrap collapses to 7/4/24, mean climbs 12.9→30.3→41.9ms. GC EXCLUDED (gcMax ≤27ms
|
||
always). The FSReq↔SimpleWrite inversion at stable active=70/pending=1287 proves reads are CAUSAL, not
|
||
a symptom of a block elsewhere.
|
||
2. SYNC CONFIG PERSIST (suspected): save()→load() reparses the WHOLE electron-config.json (1287-job
|
||
pendingQueue nested in globalSettings + full history) on every persist because _atomicWrite nulls the
|
||
cache; _serializeForDisk JSON.stringify(...,null,2) of all of it. W13's single 1021ms max with heap→142MB
|
||
fits a big synchronous structuredClone+stringify. CAVEAT (advisor): W13 is ONE confounded sample (also
|
||
FSReq=66) and the ONLY heap-spike window; W4(415ms,heap41) & W10(852ms,heap18) are LOW-heap → NOT persist
|
||
clones → likely the SECONDARY suspect: account-failed's synchronous configStore.load() per failure near
|
||
connection churn (W6 teardown had doodstream connect-timeouts). So: INSTRUMENT, don't claim "found a 1s
|
||
freeze".
|
||
|
||
SHIPPED v3.3.98 (one-variable discipline — advisor cut B to keep the next measurement clean):
|
||
- A: highWaterMark 256KB→1MB in all 5 streaming read loops (hosters.js:291, doodstream:342, voe:245,
|
||
vidmoly:190 CHUNK_SIZE consts; clouddrop:108 inline — NOT clouddrop:12's 16MB server chunk). Keep tp=8.
|
||
Deepens per-stream read-ahead 0.43s→~1.7s (absorbs threadpool-queue latency so writes don't starve),
|
||
4× fewer read completions + allocs. Zero multipart byte-risk (Content-Length=preamble+fileSize+epilogue,
|
||
independent of chunk size). REVERSIBLE PROBE; read-semaphore held in reserve (trigger: FSReq still ~70 +
|
||
writes starved + mean elevated after 1MB).
|
||
- C-instrument (BROADENED per advisor): config-store.js times load() (full reparse, incl. account-failed
|
||
path) AND _commit serialize; logs `config-load wall=Xms cache=hit/miss hist=N queue=M` and
|
||
`config-serialize wall=Xms bytes=Y hist=N queue=M` when ≥20ms (perfLog hook set in main.js via
|
||
configStore.setPerfLog→logInfo). load() split into wrapper + _loadImpl. 397 tests pass.
|
||
- B (renderer chunked rAF batch drain, app.js:188-193 — the 243ms longtask at W14) DEFERRED: renderer was
|
||
healthy 14/15 windows and the one longtask is DOWNSTREAM of the main-thread read-burst flooding IPC.
|
||
Fix A should make it self-heal. Bundling B would confound attribution + touches the progress hot path
|
||
that bit before (formatDateTime burst, ghost-fix). Add B next round ONLY if renderer still janks after A.
|
||
|
||
NEXT LOG answers 3 things cleanly: (1) did A kill the read-bursts (FSReq per-window + tail mean drop)?
|
||
(2) is the persist/load actually heavy (new config-load/config-serialize lines + their wall/queue/hist)?
|
||
(3) did the renderer self-heal from A alone (longtasks back to 0)? Then decide: persist refactor for v3.3.99
|
||
(queue-out-of-config OR cache-repopulation — latter lower-risk but renderer's incoming globalSettings isn't
|
||
default-merged like load() produces, so confirm merge-equivalence first), and/or B, and/or read-semaphore.
|
||
|
||
---
|
||
|
||
# v3.3.97 — DECISIVE ELD finding: file-read phase-flip + threadpool 64→8 + GC instrument
|
||
|
||
The v3.3.96 `eventloop-delay` logs gave the decisive signal. At CONSTANT active-count, the system flips
|
||
between two regimes:
|
||
- HEALTHY (ELD ~11ms, rss 268–308MB): `SimpleWriteWrap ≈ active`, `FSReqCallback ≈ 0–1` (write/network-bound)
|
||
- BLOCKED (ELD 49–217ms, rss 540–610MB): `FSReqCallback ≈ active` (62–71 file reads in flight), `SimpleWriteWrap ≈ 0–4`
|
||
ELD spike + rss balloon both track `FSReqCallback` → file-read path through the libuv threadpool, NOT crypto,
|
||
NOT renderer, NOT GC-alone. All 5 uploaders read identically (256KB createReadStream); byse/dood/voe run
|
||
through the GENERIC uploadFile in hosters.js (no dedicated module).
|
||
|
||
Advisor caveats baked into the build (do NOT skip on re-measure):
|
||
1. Causation UNPROVEN — high FSReqCallback could be a SYMPTOM (blocked loop can't drain read-completions).
|
||
2. rss math kills "read buffers ballooned": 70×256KB ≈ 18MB, but rss swings ~300MB → heap/object churn (GC).
|
||
3. Cheapest discriminator already wired: `UV_THREADPOOL_SIZE` 64→8 (1 line, reversible, NOT an upload cap;
|
||
8×256KB reads ≈ 100MB/s ≫ 41MB/s aggregate). Suspect tp=64 made it WORSE (removed read-serialization).
|
||
|
||
Shipped v3.3.97 = candidate-fix + discriminator in one build:
|
||
- main.js:1 `UV_THREADPOOL_SIZE` 64→8.
|
||
- ELD line now also logs `heap=`(heapUsed) `ext=`(external) `ab=`(arrayBuffers) `gc=`/`gcTotal=`/`gcMax=`ms
|
||
(PerformanceObserver entryTypes:['gc'], reset per window).
|
||
|
||
DECISION RULE for the next user log:
|
||
- ELD drops with tp=8 → read over-parallelism confirmed → keep 8 or productionize a DEDICATED read-semaphore.
|
||
- ELD high + gcTotal/gcMax align with spikes → heap churn → hunt the allocator (semaphore would be wasted).
|
||
- ELD high + gc flat → causation reversed (symptom) → pivot.
|
||
WAIT for the next `eventloop-delay` log before any read-path refactor. NO upload cap (user rejected it).
|
||
|
||
---
|
||
|
||
# v3.3.94 — comprehensive measurement build (user: "mach alles messen was man messen kann")
|
||
|
||
Localization so far (each step EMPIRICAL, not by elimination — advisor caught the elimination-leap):
|
||
- Renderer queue render PROVEN cheap: loaded the REAL app.js in headless Chromium (Playwright) with a mocked
|
||
window.api, populated Q=1000 / 61 active / progress sort, drove the real onUploadProgressBatch +
|
||
renderQueueTable + scroll → ALL <0.5ms. (Caveat: component cost, not frame rate.)
|
||
- User CONFIRMED the discriminator: a full 1000-row queue scrolls SMOOTH when idle, ruckelt ONLY while ~61–70
|
||
uploads are active → the lag is driven by the active uploads (main-process / system load), not the table.
|
||
- Screenshot: 70 connections, 1413 files, 41.3 MB/s, "write ECONNRESET". ECONNRESET is already classified
|
||
transient (upload-manager _isTransientNetworkError line 171 → retried, not account-fatal) — it's the
|
||
SIGNATURE of oversubscription (servers RST the excess connections). Same root cause as the lag.
|
||
|
||
Immediate user lever (already exists): Settings → Uploads → "Globale parallele Uploads" (parallelUploadCount,
|
||
global semaphore, default 0=off). Capping total concurrent uploads (~20) should fix lag AND ECONNRESET AND
|
||
likely keep throughput (bandwidth-limited at 41 MB/s; reset connections waste bandwidth on retries).
|
||
|
||
SHIPPED measurement (all additive, zero upload-behavior change) to pinpoint CPU-vs-IO vs renderer from the
|
||
user's REAL 70-connection run:
|
||
- main.js ELD line now also logs: cpu=X%core (process.cpuUsage delta / wall, >100% = multi-core),
|
||
rss=YMB, active-by-hoster={dood:.., voe:.., ...} (per-hoster live connection distribution → shows which
|
||
hoster is oversubscribed), transient-errs=N (cumulative ECONNRESET-class on the primary path), pending=M.
|
||
- lib/upload-manager.js: getDiagnostics() {activeByHoster, transientErrors, pending, active}; activeEntry
|
||
now carries hoster; _transientErrorTotal++ in the primary catch when _isTransientNetworkError.
|
||
- renderer/app.js: PerformanceObserver('longtask') + a rAF frame-time monitor → logs every 5s WHILE
|
||
uploading: `renderer-perf active=N fps=X jankFrames=Y worstFrame=Zms longtasks=W maxTask=Vms`. This is the
|
||
DIRECT renderer ground truth (the component-timing harness couldn't capture real frame rate). Low fps /
|
||
high jankFrames / longtasks → renderer IS blocked; ~60fps + no jank while it still feels laggy → it's the
|
||
main-process/system, and the cpu=/eld= numbers in the same log say CPU-bound (→ workers/cap) vs IO-bound.
|
||
Both logs land in the normal debug log (logInfo / window.api.debugLog). 397/397 tests, eslint clean.
|
||
|
||
NEXT: user runs the 70-load on v3.3.94, shares the `eventloop-delay` + `renderer-perf` log lines (or connects
|
||
diagnostics). Those two lines together localize it definitively. Do NOT build workers/cap before that.
|
||
|
||
---
|
||
|
||
# v3.3.93 — THE renderer lag knot FOUND + FIXED + MEASURED: formatDateTime per progress event
|
||
|
||
User gave the decisive data: "25 connections okay, 50+61 laggt, EVTL wenn die uploadenden Zeilen nicht im
|
||
Bild sind." Two regime facts (asked, not assumed — advisor caught the assume-the-regime trap a 3rd time):
|
||
queue = 200–1000 rows (VIRTUAL mode) + sort = clicked PROGRESS/SPEED (dynamic). This killed the non-virtual
|
||
reflow theory (off-screen rows aren't in the DOM when virtual) AND pointed at the per-event path.
|
||
|
||
ROOT CAUSE (renderer process, NOT main — the v3.3.91 ELD log can't see this): `maybeAddSessionFile(job)`
|
||
computed `const dt = formatDateTime(new Date())` UNCONDITIONALLY at the top, before the `status==='done'`
|
||
check that early-returns for everything else. formatDateTime does TWO Intl locale formats
|
||
(toLocaleDateString + toLocaleTimeString) = ~83µs/call MEASURED. It runs on EVERY progress event
|
||
(onUploadProgressBatch loops the M-item batch → handleProgress → _handleProgressImpl → maybeAddSessionFile),
|
||
i.e. 10×M/sec, and THROWS IT AWAY for all non-done events (the overwhelming majority while uploading).
|
||
- Scales exactly with M (active count): 250/sec at M=25 → 610/sec at M=61.
|
||
- Bursts: each progress batch runs M calls back-to-back = a SYNCHRONOUS main-thread block of
|
||
~2.4ms (M=25) → ~5ms (M=61) every 100ms, on top of render+sort → blows the 16ms frame budget → scroll
|
||
stutter. Scroll-independent (per-event, not per-render) → matches "lag when actives off-screen" exactly.
|
||
This is the 25→50 cliff.
|
||
|
||
FIX: move `const dt = formatDateTime(new Date())` inside the `if (!_sessionFileKeys.has(dedupKey))` block, so
|
||
it runs ONCE per genuinely-new completed upload, never per progress tick.
|
||
|
||
VERIFIED (faithful Blink benchmark at the CONFIRMED regime: Q=500 virtual, dynamic progress sort, scrolling,
|
||
M=25/50/61, OLD vs FIXED): per-batch cost 1.7/3.2/4.1 ms (OLD, scales with M) → 0.0/0.0/0.0 ms (FIXED, flat).
|
||
Frame P95 7.3→4.2 ms at M=61. M-scaling ELIMINATED. The render/scroll path itself is flat ~2.5ms median
|
||
across all M → NO second knot there. updateStatusBar/StatsPanel = one cached O(Q) arithmetic pass (cheap);
|
||
updateQueueActionButtons = O(selection) (cheap). No other Intl/Date on any per-event/per-frame hot path
|
||
(2582 = job-log modal, 4587 = History view — both on-demand/cold). 397/397 tests pass, eslint clean.
|
||
|
||
NOTE: the v3.3.91/92 main-process event-loop-delay instrument is for the OTHER (CPU-vs-IO) hypothesis and is
|
||
a separate process — keep it; it still answers whether the main thread also saturates at 50+ TLS streams.
|
||
|
||
---
|
||
|
||
# High-concurrency lag audit (v3.3.90) — "lag ist immernoch da, ich vermute ab X gleichzeitig muss er alle Zeilen gebündelt updaten"
|
||
|
||
Method: 44-agent high-concurrency audit of the full upload→IPC→render path + Blink benchmark of the
|
||
renderer queue table (Playwright/Chromium = same Blink engine), targeting the user's NEW hypothesis:
|
||
"with ~100 concurrent uploads the renderer has to update ALL rows bundled rather than cleanly per-row."
|
||
|
||
## The user's hypothesis is MEASURED-REFUTED — the renderer is NOT the bottleneck.
|
||
Blink benchmark over scenarios Q=150..1000, M=10 active, 60 ticks each:
|
||
- renderQueueTable virtualizes at ≥200 rows; <200 = change-detecting in-place update.
|
||
- _updateRowInPlace is change-detecting (no forced reflow, no layout reads).
|
||
- median render <1 ms at Q=1000; only ~4/60 renders are full rebuilds even with progress-crossing sorts.
|
||
- progress is coalesced main-side (_progressByJob Map keyed by jobId + 100ms flush → one batch sized by
|
||
active-job count, ~10/sec); renderer iterates the batch with cheap per-row handleProgress.
|
||
DOM amplification is ruled out by measurement. "Laggy at ~40% CPU / 8 cores" = ONE core at 100% =
|
||
main-thread saturation / synchronous blocking, not DOM.
|
||
|
||
## SHIPPED (v3.3.90) — the two real main-thread blockers, both behavior-preserving
|
||
1. lib/clouddrop-upload.js `_uploadChunked`: was reading each 16 MB chunk with `fs.readSync` SYNCHRONOUSLY
|
||
on the main event loop — unique among the 5 uploaders (the other 4 stream async). Each read blocks the
|
||
WHOLE loop (~5–9 ms SSD, 30–100 ms slow disk) → freezes all progress/IPC/render/other-uploads, scaling
|
||
with the number of concurrent clouddrop uploads. Fits "laggy when uploading, worse with more concurrent."
|
||
User uses clouddrop. Fix: `fs.openSync`/`readSync`/`closeSync` → `fs.promises.open` + `await fh.read` +
|
||
`await fh.close()`. Byte-equivalence verified by SHA-256 over all chunk-boundary cases (full chunk,
|
||
partial last chunk, 2/3/4-chunk, single byte) before shipping — a chunk-read bug = corrupt upload.
|
||
2. lib/upload-manager.js rotation-retry (944) + suspect-alternate (1075) progressCb: both called
|
||
`_emitProgress` (a synchronous `emit('progress')` + fresh object spread) on EVERY stream chunk
|
||
(hundreds/sec per job) — they were missing the 250 ms `lastEmitTime` gate that the primary path (631)
|
||
has. With many concurrent uploads in rotation/suspect mode that's real main-thread emit amplification.
|
||
Mirrored the gate exactly: activeEntry mutation stays UNGATED (stats/speed-monitor stay fresh), only the
|
||
emit is throttled to 4/sec. Behavior-preserving.
|
||
397/397 tests pass, eslint clean (1 pre-existing unrelated warning at line 554).
|
||
|
||
## DROPPED (advisor: measured fine, don't chase perception)
|
||
- Lowering the virtual-row threshold below 200: the Blink benchmark shows <200 in-place updates are already
|
||
sub-ms; no change warranted.
|
||
|
||
## DISCRIMINATOR ANSWERED (user, 2026-06-21)
|
||
(a) Lag NOT clouddrop-specific — other hosters. (b) Parallel counts RAISED deliberately (10+).
|
||
(c) 50+ uploading SIMULTANEOUSLY active. → This is the TRUE high-concurrency main-thread-funnel branch,
|
||
NOT clouddrop. v3.3.90 stands but does not target this user's case.
|
||
|
||
## v3.3.91 — instrument first, don't refactor the upload core off elimination-reasoning
|
||
Advisor reframe: two LIVE hypotheses need OPPOSITE fixes — (A) main thread CPU-blocked (TLS/crypto/sync) →
|
||
event loop stalls → a cap/workers help; (B) main thread fine but IO-STARVED (libuv threadpool/sockets) →
|
||
loop stays responsive, uploads just queue → workers are WASTED, config fixes it. A worker/child-process
|
||
upload refactor touches throttle/rotation/abort/progress/credentials and is hard to reverse — DO NOT ship it
|
||
off sandbox elimination. One measurement splits the hypotheses and must run in the REAL app.
|
||
|
||
SHIPPED (both reversible, zero upload-core refactor):
|
||
1. main.js: `perf_hooks.monitorEventLoopDelay({resolution:10})` enabled at startup; logged via logInfo every
|
||
~5 s WHILE uploading (state==='uploading' && activeJobs>0) as
|
||
`eventloop-delay active=N mean=..ms p99=..ms max=..ms stddev=..ms threadpool=..`. Pure numbers, no secret
|
||
→ does NOT touch the redaction surface. This is the GROUND TRUTH: high mean/p99 → CPU-blocking → workers
|
||
justified; low delay while uploads stall → IO-bound → workers wasted, threadpool/sockets is the fix.
|
||
2. main.js (first statement, before require('electron')): `UV_THREADPOOL_SIZE = env || '64'`. Default is 4;
|
||
every async uploader feeds undici from fs.createReadStream (+ clouddrop fh.read) and DNS getaddrinfo goes
|
||
through the same pool → 50 concurrent vs 4 threads = reads/DNS serialize 4-at-a-time = a hard cliff at a
|
||
small connection count = the "ab X connections" symptom. Threads are created lazily on demand → 64-max
|
||
costs nothing if unused (zero-risk, reversible). The advisor's prescribed one-env-var hypothesis test.
|
||
|
||
CAVEAT (honest): synthetic sandbox benches could NOT confirm the threadpool is the bottleneck — pbkdf2 is
|
||
CPU-core-bound (masks pool size); DNS .invalid returns instantly; real-RTT DNS showed NO pool benefit because
|
||
WINDOWS serializes getaddrinfo via the OS DNS Client service (so on Windows the DNS half of the cliff is
|
||
masked by the resolver, though the fs-read half still benefits). This is exactly why the ELD number must come
|
||
from the user's real load, not the sandbox. Per-uploader undici Agent audit: clouddrop has a shared
|
||
module-level Agent (connections:50); doodstream/voe/vidmoly use the global dispatcher (pooled per origin, NO
|
||
per-call agent explosion) — so no agent fix needed.
|
||
|
||
## v3.3.92 — make the single measurement decisive + breadth audit of un-checked main.js hot paths
|
||
Enriched the ELD log line with `process.getActiveResourcesInfo()` as a compact type-histogram:
|
||
`eventloop-delay active=N mean/p99/max/stddev ms threadpool=64 resources=K {TCPSocketWrap:50,FSReqCallback:4,...}`.
|
||
Now ONE run splits all three readings in a single line: high mean/p99 → CPU-blocked (workers/cap);
|
||
low delay + many TCP/FS/GetAddrInfo resources → IO-bound queueing (threadpool/sockets, NOT workers);
|
||
low delay + few resources → not saturated (lag elsewhere / perception). Pure numbers, no redaction surface.
|
||
|
||
Breadth audit this round (4th /goal re-fire, code I wrote, NOT re-measuring cleared render/persist):
|
||
- main.js logging (debug/rot/upload): all buffered + ASYNC fs.appendFile (write-guard flag, 500ms timer,
|
||
setImmediate re-flush). Sync appendFileSync ONLY in crash/signal/exit handlers (correct there). CLEAN.
|
||
- main.js progress coalescing (_progressByJob Map + 100ms batch → one upload-progress-batch via safeSend):
|
||
non-terminal = Map.set (keeps latest/job); gated upstream to 4/sec/job. CLEAN at N=50.
|
||
- _appendJobLog: capped in-memory ring buffer (Map, FIFO-evict). CLEAN.
|
||
- All 5 uploaders: doodstream/voe/vidmoly/clouddrop-simple stream via async createReadStream + for-await +
|
||
async throttle.consume; clouddrop-chunked now async fh.read. NONE block the main loop per chunk. CLEAN
|
||
(clouddrop's old readSync was the unique outlier, fixed v3.3.90).
|
||
|
||
## NEXT (gated on the real-app ELD number + user's explicit nod)
|
||
User runs their 50-concurrent load once; the enriched `eventloop-delay` log line decides:
|
||
- mean/p99 HIGH (tens–hundreds ms) → CPU-blocked → propose worker_threads/child-process upload pool OR a
|
||
smart concurrency cap (WITH the user's nod — it's hard to reverse and touches credentials/abort/rotation).
|
||
- delay LOW while it still lags → IO-bound → threadpool bump already addresses it; if not, look at socket
|
||
caps / undici Agent connection limits / per-origin pooling, NOT workers.
|
||
Do NOT build the worker refactor before this number exists.
|