v3.3.100's interaction instrument pinpointed the residual UI lag exactly: every
slow click was a tab switch into the History view (button.tab / nav.tab-bar,
200-840ms), each coupled to `ipc get-history wall=150-200ms sync` +
`main-longtask blocked=242-289ms lastIpc=get-history` + a 200-248ms renderer long
task. Uploads themselves are pristine (event-loop mean 11.7ms; the only spikes line
up with the get-history tab switches).
Root cause: the History tab handler called loadHistory() UNCONDITIONALLY on every
activation (renderer/app.js) even though it tracks a `_historyDirty` flag it never
checked. Each call synchronously readFileSync + JSON.parse's the ~185MB /
30000-entry electron-history.json (no cache), ships all 30000 batches across IPC,
and the renderer flattens ~120000 row objects before .slice(-2000) for the DOM (the
DOM is already capped at 2000).
Fix (the adversary-verified safe subset):
- Gate the History-tab load: only reload when `_historyDirty || !_historyEverLoaded`.
Added `_historyEverLoaded`, and both flags are now set inside loadHistory() after
the fetch succeeds (so a failed load retries). Dirty-coverage is complete — every
history append routes through batch-done -> appendHistory and upload-batch-done ->
handleBatchDone which sets _historyDirty=true. Result: repeat History tab switches
with no new uploads do zero IPC/parse/flatten and are instant. (The first open
after a new batch still parses once ~450ms; removing that needs a parse-cache or
JSONL storage, deferred.)
- Diagnostics history regression: getHistory() read loadConfig().history, but since
the v3.3.99 history split load() returns history:[] in packaged mode, so remote
diagnostics reported totalBatches:0 despite 30000 real batches. It now reads
loadHistory() (injected via the collector deps), with a backward-compatible
fallback to loadConfig().history when not provided.
407 tests pass (2 new diagnostics regression tests); clean Electron boot.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Closes the last measurement gap. The main process was already fully instrumented
(every ipcMain handler timed at >=50ms, a 100ms main-thread long-task monitor with
channel attribution, config load/serialize timing). Account switches were already
covered too: switchAccount is a trivial synchronous Map-set and the rotation work
is async, so a switch cannot block the main thread, and any block that did occur
would surface in the long-task monitor.
The real gap was the renderer side: the renderer-perf line was gated on active
uploads (idle clicks were never logged) and only reported an aggregate longtask
count — no per-interaction latency and no element attribution. So a switch/sort/tab
that janked in the renderer (not main) was invisible.
renderer/app.js (additive, self-silencing, wrapped in try/catch):
- An Event Timing observer (PerformanceObserver type:'event', durationThreshold:50,
buffered) logs `renderer-interaction <type> dur=Xms proc=Yms target=<el>` for every
user interaction whose latency exceeds 50ms — always on, idle or under load — and
names the element (id / first class / data-action / aria-label / title). This is
the direct click->reaction latency the user feels.
- The longtask observer now also logs `renderer-longtask dur=Xms` immediately for any
single renderer long task >=100ms, regardless of upload state.
405 tests pass; clean Electron boot. Every action — main or renderer, idle or under
load — now names itself in the log if it is slow.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The v3.3.98 instrument exposed the actual cause of the 1-2s UI button lag, and it
was none of the read-path suspects. electron-config.json had grown to 38.5MB and
was being load()/structuredClone()/JSON.stringify()'d ~137 times in a 73s window
(140-592ms each) on the Electron main thread — roughly 47% main-thread occupancy.
That is the lag: a button click lands while the thread is mid-clone of a 38.5MB
object. The 1MB read-ahead could not touch it.
The bulk is history, not the queue. Each batch-done appended the upload-manager
summary verbatim, including the full per-file result list (per-hoster URLs), into
config.history; default historyRetention='all' never prunes, so it grew unbounded
(75 batches ≈ 38.5MB). The "queue=undefined" in the perf log was a logging bug
(reading .length on the pendingQueue object), not an empty queue. Writes drove the
storm: queue-persistence (save-global-settings) did two loads + one 38.5MB
serialize per call, and _atomicWrite nulls the cache so the next load is a full
38.5MB reparse; even cache hits structuredCloned the whole 38.5MB. The per-job
upload path makes zero config calls, so pending=1280 was never the driver.
Fix — move history into its own file so it is never parsed/cloned/serialized on a
config op (a 5-agent investigation + adversarial review chose this over three
read-side half-fixes; it is the only change that removes the clone AND the
serialize AND the post-write reparse at once):
- History lives in electron-history.json. _migrateHistory() runs once at init
(packaged only) and is fail-safe: it writes history.json (tmp → fsync → rename),
re-reads and verifies the entry count, and keeps a permanent
electron-config.json.pre-history-split.bak BEFORE the config is ever allowed to
drop its history. If verification fails it leaves history in the config (retry
next launch). _loadImpl returns history:[] once migrated, so the cached object is
tiny (cheap clones); the config file shrinks to ~KB on the first save (cheap
reparse) and _serializeForDisk writes ~KB (cheap serialize).
loadHistory/appendHistory/pruneHistory/clearHistory go through history.json on
their own write-queue with a no-clobber guard; the legacy config path stays as a
fallback when migration did not run.
- Validated against the real 185MB / 30000-entry bench fixture: migrate 1.5s once,
every entry preserved + .bak kept; load() 631ms cold once → 0.1ms after the first
save strips the file (185MB → 2.1KB); loadHistory() still returns all 30000.
Dropped per review (one-variable + risk): loadShallow (moot after the split),
cache-repopulate (its gate can never fire), and a per-batch resolution cache (stale
account pools → the rotation/byse failover-regression class — the one thing that
could silently corrupt uploads).
Instrumentation (the user asked to measure everything; all additive, threshold-
gated, MHU_PERF=0 disables):
- ipcMain.handle/.on are centrally wrapped to log `ipc <channel> wall=Xms sync=Yms`
over 50ms — the button-press→response latency — hardened (Promise.resolve(p)
.finally + try/catch'd logging) so a logging failure can never break IPC.
- A 100ms main-process drift monitor logs `main-longtask blocked=Xms lastIpc=… gc=…`
for any single main-thread turn over 100ms (catches GC, fs scans, serialize that
the IPC timing structurally cannot see).
- config-store perf lines gain via=<caller> and wqDepth=, and the queue= logging
bug is fixed (now reads pendingQueue.queueJobs.length).
405 tests pass (9 new migration tests: preserve-count, round-trip, save() never
loses history, crash-window fallback, idempotency). Clean Electron boot, no repo
pollution (migration is packaged-only).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
v3.3.97 (UV_THREADPOOL_SIZE 64→8) was a decisive win — mean event-loop-delay at
70 active uploads dropped 200ms→~11ms (18×), rss 577→287MB, renderer healthy in
14/15 windows. But the user reports it is still not perfectly smooth. A focused
multi-agent investigation plus an adversarial review localized the residual to
TWO distinct, separately-measured spike sources:
1. Read-bursts. In the tail windows the file-read histogram inverts: FSReqCallback
climbs to 66-70 against threadpool=8 (~8.75× queue depth) while SimpleWriteWrap
(socket writes) collapses to 4-24 and mean delay rises to 30-42ms. GC is ruled
out (gcMax ≤27ms in every window). The clean inversion at a stable active=70 /
pending=1287 shows the reads are causal, not a symptom of a block elsewhere.
2. A suspected synchronous config-persist stall. save() → load() reparses the whole
electron-config.json — which now carries the 1287-job pending queue nested in
globalSettings plus full history — on every persist (because _atomicWrite nulls
the read cache), then _serializeForDisk JSON.stringify(…, null, 2) of all of it.
One tail sample (max 1021ms, heap spiking to 142MB) fits a large synchronous
structuredClone+stringify, but it is a single confounded point, so this build
only INSTRUMENTS the path rather than asserting the cause.
This release ships one behavioral change (kept to a single variable so the next
log attributes cleanly) plus measurement:
- highWaterMark 256KB→1MB in all five streaming read loops (lib/hosters.js,
doodstream/voe/vidmoly CHUNK_SIZE consts, and the inline value in
clouddrop-upload.js:108 — NOT the 16MB server chunk at clouddrop-upload.js:12).
UV_THREADPOOL_SIZE stays 8. This deepens each stream's read-ahead cushion from
~0.43s to ~1.7s at the per-stream rate, so a stream tolerates the threadpool
queue without starving its socket write, and cuts read-completion callbacks and
per-chunk Buffer allocations ~4×. Byte-correctness is unaffected: Content-Length
is preamble+fileSize+epilogue, independent of chunk size, and the chunk size
never touches the multipart boundaries. Fully reversible; a dedicated read-
concurrency semaphore is held in reserve if 1MB does not clear the bursts.
- config-store.js now times load() (the full reparse, which the account-failed
handler also hits per failure) and the _commit serialize, logging
`config-load …` / `config-serialize wall=…ms bytes=… hist=… queue=…` when the
synchronous work exceeds 20ms. load() is split into a timing wrapper + _loadImpl;
the timer is a no-op until main.js wires configStore.setPerfLog → logInfo.
The renderer batch-drain fix for the one observed 243ms longtask is intentionally
deferred: that jank is downstream of the main-thread read-burst flooding IPC, so
fix#1 should make it self-heal; bundling it would confound the measurement and
touch the progress hot path. All 397 tests pass.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The v3.3.96 eventloop-delay logs from a real 70-connection run pinpointed the
high-concurrency lag to the file-read path. At a CONSTANT active-job count the
process flips between two clean regimes:
HEALTHY ELD ~11ms, rss 268-308MB: SimpleWriteWrap ≈ active, FSReqCallback ≈ 0-1
BLOCKED ELD 49-217ms, rss 540-610MB: FSReqCallback ≈ active (62-71 reads in
flight), SimpleWriteWrap ≈ 0-4
Both the ELD spike and the rss balloon track FSReqCallback (libuv-threadpool file
reads) exactly — not crypto, not the renderer, not GC alone. All five uploaders
read identically via fs.createReadStream({highWaterMark: 256KB}); byse/doodstream/
voe run through the generic uploadFile in hosters.js (no dedicated module).
This build is both a candidate fix and a discriminator, per advisor review:
- UV_THREADPOOL_SIZE 64→8 (main.js:1). One reversible line, NOT an upload cap —
70 uploads still run. 8 concurrent 256KB reads sustain ~100MB/s, far above the
41MB/s aggregate, so it cannot bottleneck throughput even on the slow VM disk.
Strong suspicion that tp=64 made it worse: it removed the natural read-
serialization (default 4 threads) and let all 70 streams' reads fire at once,
flooding the loop with completion callbacks in lock-step bursts.
- ELD line now also logs heap=heapUsed ext=external ab=arrayBuffers and
gc=/gcTotal=/gcMax=ms (PerformanceObserver entryTypes:['gc'], reset per window).
The 70×256KB ≈ 18MB of read buffers cannot account for the ~300MB rss swing —
that is heap/object churn, so GC must be measured directly.
Decision rule for the next real-run log:
- ELD drops with tp=8 → read over-parallelism confirmed (keep 8 or add
a dedicated read-semaphore).
- ELD high + GC pauses align → heap churn, hunt the allocator.
- ELD high + GC flat → causation was reversed, pivot.
No upload-behavior change; the concurrency cap the user explicitly rejected is
untouched.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The v3.3.94 measurement computed the main-process event-loop delay, CPU%, per-hoster connection distribution and transient error counts, but logged only '[INFO ] perf': logInfo(ctx, msg) treats a string first arg as the whole message and discards the second, so the payload was thrown away (confirmed in the user's upload-debug.log). Pass the line as a single argument so the real numbers reach the log.
Also: assets/ was missing from the electron-builder files list, so app_icon.ico was never packaged into the asar — new Tray() threw an unhandled rejection on every startup and left no tray icon. Package the icons and wrap createTray in try/catch with a nativeImage fallback so a missing/invalid icon can never crash tray creation.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>