Per user request, the per-session log file is renamed from
fileuploader-session-YYYY-MM-DD_HH-MM-SS-<pid>.log to
DD-MM-YYYY-mdu-session-HH-MM.log (hour and minute only, no seconds or pid).
lib/log-mode.js:
- formatSessionStamp(date) returns `${DD}-${MM}-${YYYY}-mdu-session-${HH}-${MM}`
(the pid argument is dropped; main.js still passes process.pid, harmlessly
ignored). Same-minute restarts now share a session file, which is the intended
human-readable trade.
- the session branch of resolveLogFileName returns `${sessionId}${ext}` — the stamp
is the full app-defined stem and baseName is intentionally ignored for session
mode (single/daily still use the 'fileuploader' base).
- stripModeStampFromFileName recognizes the new format and resets to the default
'fileuploader' base (the new stem embeds no base, so the configured base is not
recoverable from it); the existing daily and old-session strip regexes are kept
for backward-compat with any persisted old paths, and the persist/re-resolve
round-trip stays idempotent (no compounding stamps).
Tests updated for the new format (formatSessionStamp, session resolveLogFileName,
the new-format strip, and the idempotency regression). 410 tests pass.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(A) Intermittent pure-white startup window. On a Windows VM viewed over
RemoteDesktop the app sometimes opened to a blank white window — no error banner,
not even the static menu bar. A multi-agent investigation + adversarial review
pinned it by one airtight deduction: the BrowserWindow is created with
backgroundColor '#16181c' (dark), so "white" can never be an un-painted, loading,
or failed state — all of those show dark. Pure white, with no banner and no static
menu bar, and silent in the logs, eliminates every DOM/init/CSS/load failure mode
(each would leave the dark styled shell, unstyled black-on-white menu text, or the
red init().catch banner) and leaves exactly one: a GPU/compositor surface failure
on the RDP virtual display adapter. The code ran hardware acceleration full-default
with no fallback (no disableHardwareAcceleration, no GPU switch anywhere), and the
GPU child-process-gone handler is log-only — matching the silent symptom.
Fix (the reviewed safe subset):
- app.disableHardwareAcceleration() at module top (before app.whenReady), gated on
an RDP session (process.env.SESSIONNAME matches /^RDP/) OR a persisted
gpu-disabled.flag in userData. The renderer has no WebGL/canvas/video, so software
compositing costs effectively nothing here and does not undo the recent renderer
perf work; local/console users keep hardware acceleration.
- Auto-heal: when a GPU child-process-gone fires, write gpu-disabled.flag so the next
launch disables acceleration even if the RDP gate didn't match (covers a VM reached
via console or a flaky virtual GPU). Self-heals after at most one white screen.
- The existing child-process-gone / render-process-gone / did-fail-load logging is
kept so the affected server's next white-start log can confirm the cause
(CHILD PROCESS GONE type=GPU). A webContents reload would not disrupt uploads
(uploadManager lives in the main process and is torn down only on quit), but no
watchdog is added because in this GPU mode init completes — an init-complete signal
would not detect the blank surface.
(B) Backup export default filename changed from multi-hoster-backup-YYYY-MM-DD.mhu to
DD-MM-YYYY-multihoster-backup.mhu.
409 tests pass; clean boot (the guard is inert off-RDP).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
A user's server crashed hard during an upload and lost all configured accounts. It
was NOT the v3.3.104 update (a second server updated fine and kept its accounts) —
it was a data-durability hole exposed by the crash:
- Config writes were atomic (tmp + rename) but never fsync'd, so a hard crash could
leave electron-config.json truncated/unflushed on disk.
- On restart, load() reads the truncated file, falls back to .bak, and if that is
also bad returns empty DEFAULTS. The next settings/queue save then persists EMPTY
hosters — permanently wiping the accounts. Worse, the async _atomicWrite blindly
copied the (now truncated) live file over .bak, so an empty live could clobber a
good backup.
Hardening (lib/config-store.js + main.js; no behavior change in the happy path):
- fsync before rename in both write paths — _atomicWrite (openSync/writeSync/
fsyncSync/closeSync) and the synchronous save-global-settings-sync on window close.
A hard crash can no longer leave a truncated config.
- _atomicWrite only refreshes .bak when the current live file is non-trivial
(trim length > 2), so an empty/truncated live can never overwrite a good backup
(the sync-save path already did this).
- Wipe-guard (_guardHosters): save(), saveRotationCursors() and the sync close-save
never intend to change hosters; if after a load() the hosters are all empty and
the write did not explicitly provide hosters, recover them from disk
(_recoverHostersFromDisk: live -> .bak -> .pre-history-split.bak) instead of
persisting the wipe. An explicit save({hosters: {}}) (user deleted all accounts)
is still allowed. Restored hosters are already-encrypted on disk and
encryptCredentials skips already-encrypted fields, so re-serializing is safe.
- load() gained a third fallback tier — the permanent pre-history-split.bak snapshot
(which still holds the accounts) — so load() itself recovers after corruption.
Recovery for the already-affected server: copy
%APPDATA%/multi-hoster-uploader/electron-config.json.pre-history-split.bak (or .bak)
over electron-config.json with the app closed.
2 new regression tests (post-wipe valid-empty live + .bak → guard restores accounts;
an explicit empty-hosters save is not blocked). 409 tests pass; clean boot.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The v3.3.103 log (real 2464-job, 4-hoster batch ramping to 95 concurrent) confirmed
the statSync fix held (batch-start main spike 336→231ms with 10× more jobs) and the
whole 90s ramp to 95 active was pristine (event-loop mean ~11ms, fps=32,
longtasks=0). The one residual: rapidly clicking tabs DURING the 95-active upload
produced 210-221ms renderer long-tasks (proc=0ms → layout/paint, not JS).
Cause: the Recent-uploads panel rendered every sessionFilesData row (up to 2000)
into the DOM non-virtualized — the exact analog of the History table before it was
virtualized in v3.3.102. Switching to that view laid out ~2000 rows (~210ms on the
RDP VM).
Fix — virtualize renderRecentUploadsPanel, mirroring the queue/History pattern:
- The tbody gets only the visible rows plus top/bottom spacer <tr> sized from
VIRTUAL_ROW_HEIGHT. A rAF-coalesced scroll handler and a ResizeObserver on
.recent-files-table-wrap re-render the visible window (the ResizeObserver also
serves as the show-trigger when the hidden panel gains size). _recentWorking holds
the sorted set. The insertAdjacentHTML append-only fast path is dropped — a
~40-row window re-render is cheap, so every render just re-renders the window; on
prepend (date desc) the scroll position is preserved (scrollTop=0 at top, else
+= added*ROW_HEIGHT).
- Selection stays correct: _buildRecentRowHtml already stamps the selected class
from selectedRecentIds.has(row.order) per row, so an off-screen-selected row
renders selected when scrolled into view; selectedRecentIds remains the source of
truth and shift-select already reads the sort cache, not the DOM.
- styles.css gives .recent-file-row a fixed 28px height so the virtualization math
is exact (the table already had table-layout:fixed, so no column-jump fix needed).
- _renderRecentVirtualRows returns early when there are no rows, so it never wipes
the empty-state message.
Verified with Playwright at 2000 rows (bounded container): view-show layout drops
from ~118ms to ~2.4ms, the DOM stays at 29-39 rows, scrolling maps to the correct
rows, the scrollbar height is exact, an off-screen-selected row renders with the
selected class, and the row height is exactly 28. Every large table (Queue,
History, Recent) is now virtualized.
407 tests pass; clean Electron boot.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The v3.3.102 log (real 224-job batch) confirmed History virtualization works and
steady-state uploads are pristine (fps=32, event-loop mean 11.8ms). The remaining
residual was a ~6s spin-up burst at batch start: the main event loop blocked for
336ms with cpu=0%core (i.e. blocked on I/O, not computing) and the renderer janked
196-391ms, then everything settled clean.
A multi-agent investigation plus an adversarial review corrected the obvious-looking
hypothesis. The renderer's uncapped progress-batch drain is NOT the cause:
handleProgress only mutates plain JS state and schedules already-coalesced renders
(one per frame), and main coalesces progress to ~50 latest-per-job entries per
100ms. Chunking that drain would fix nothing — and the reviewer showed it would
REGRESS correctness: requestAnimationFrame throttles to ~0 when the window is
minimized (the common state for a background uploader), so a rAF-chunked drain would
grow an unbounded backlog and defer persistQueueStateSoon for every buffered item,
losing terminal 'done' events on close (the queue-persistence ghost-fix class). So
that path is deliberately not taken.
The real cause (cpu=0%core = blocked on I/O) is a synchronous fs.statSync storm in
UploadManager.startBatch: the dedup loop ran up to DEDUP_CHUNK=200 synchronous
fs.statSync calls in a single tick before yielding (200 x ~1.68ms on the user's VM
= the exact 336ms), on a disk already saturated by the 1MB read-ahead.
Fix — make the batch-start stats non-blocking:
- The dedup loop now dedupes synchronously (cheap Map work) and then stats the
unique files in parallel via await Promise.all(fs.promises.stat ...) per chunk, so
the stat I/O runs on the libuv threadpool and the main thread never blocks. The
results-Map shape ({name,size,results:[]}) and dedup semantics (size 0 on failure)
are unchanged.
- The per-job statSync fallback is converted to await fs.promises.stat for
consistency (it sits in an async function before the first real await; the cached
size from dedup already lets nearly every job skip it).
Tests: the upload-manager mocks override fs.statSync; they now also override
fs.promises.stat with the same fake sizes (upload-manager.test.js x2,
suspect-reject-alternates.test.js). 407 tests pass; clean Electron boot.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
v3.3.101's gate removed the get-history parse on tab switch, but the log still
showed tab clicks at ~216ms with a ~197ms renderer long-task and NO get-history
(proc=0ms → pure browser layout, not JS). Cause: `.view{display:none}` ->
`.active{display:flex}` and the History table built up to 2000 non-virtualized
<tr>, so making the view visible laid out 2000 rows (~197ms on the RDP VM). The
queue was already virtualized; the History table was the only large non-virtual
one.
Measured the options with Playwright on the real DOM (this machine; the user's VM
is ~1.7x slower):
- current 2000 rows (auto layout): 118ms
- content-visibility + table-layout:fixed: 117ms (no help — rows still laid out)
- cap to 300 rows: 16ms, but rejected: History rows are per-file-per-hoster, so a
single 1280-file batch is ~3840 rows and a small cap would hide a recent batch's
links
- virtualize (40 visible of 2000): 2.3ms
Fix — virtualize renderHistoryTable, mirroring the queue's _renderVirtualRows:
- The header is always rendered; tbody#historyBody gets only the visible rows plus
top/bottom spacer <tr> sized from VIRTUAL_ROW_HEIGHT. A rAF-coalesced scroll
handler and a ResizeObserver on #historyContainer re-render the visible window;
the ResizeObserver also serves as the show-trigger (a hidden 0x0 container that
gains size on tab activation re-renders at the correct height). Sorting resets
scrollTop and re-renders; the copy-link / sort-header click delegation is
unchanged. _historyWorking holds the sorted working set.
- styles.css: .history-table gets table-layout:fixed plus scoped column widths so
columns do not jump as rows scroll in and out. This is scoped to .history-table
and does not touch the .col-* classes the queue shares.
Verified end-to-end with Playwright at 6000 rows: show cost 1.1ms (was 118ms), DOM
stays 32-42 rows, scroll maps to the right rows (top/middle/bottom), scrollbar
height exact, column widths stable, rows update on scroll. All rows remain
scrollable (no UX loss); the show cost drops ~100x.
407 tests pass; clean Electron boot.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
v3.3.100's interaction instrument pinpointed the residual UI lag exactly: every
slow click was a tab switch into the History view (button.tab / nav.tab-bar,
200-840ms), each coupled to `ipc get-history wall=150-200ms sync` +
`main-longtask blocked=242-289ms lastIpc=get-history` + a 200-248ms renderer long
task. Uploads themselves are pristine (event-loop mean 11.7ms; the only spikes line
up with the get-history tab switches).
Root cause: the History tab handler called loadHistory() UNCONDITIONALLY on every
activation (renderer/app.js) even though it tracks a `_historyDirty` flag it never
checked. Each call synchronously readFileSync + JSON.parse's the ~185MB /
30000-entry electron-history.json (no cache), ships all 30000 batches across IPC,
and the renderer flattens ~120000 row objects before .slice(-2000) for the DOM (the
DOM is already capped at 2000).
Fix (the adversary-verified safe subset):
- Gate the History-tab load: only reload when `_historyDirty || !_historyEverLoaded`.
Added `_historyEverLoaded`, and both flags are now set inside loadHistory() after
the fetch succeeds (so a failed load retries). Dirty-coverage is complete — every
history append routes through batch-done -> appendHistory and upload-batch-done ->
handleBatchDone which sets _historyDirty=true. Result: repeat History tab switches
with no new uploads do zero IPC/parse/flatten and are instant. (The first open
after a new batch still parses once ~450ms; removing that needs a parse-cache or
JSONL storage, deferred.)
- Diagnostics history regression: getHistory() read loadConfig().history, but since
the v3.3.99 history split load() returns history:[] in packaged mode, so remote
diagnostics reported totalBatches:0 despite 30000 real batches. It now reads
loadHistory() (injected via the collector deps), with a backward-compatible
fallback to loadConfig().history when not provided.
407 tests pass (2 new diagnostics regression tests); clean Electron boot.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Closes the last measurement gap. The main process was already fully instrumented
(every ipcMain handler timed at >=50ms, a 100ms main-thread long-task monitor with
channel attribution, config load/serialize timing). Account switches were already
covered too: switchAccount is a trivial synchronous Map-set and the rotation work
is async, so a switch cannot block the main thread, and any block that did occur
would surface in the long-task monitor.
The real gap was the renderer side: the renderer-perf line was gated on active
uploads (idle clicks were never logged) and only reported an aggregate longtask
count — no per-interaction latency and no element attribution. So a switch/sort/tab
that janked in the renderer (not main) was invisible.
renderer/app.js (additive, self-silencing, wrapped in try/catch):
- An Event Timing observer (PerformanceObserver type:'event', durationThreshold:50,
buffered) logs `renderer-interaction <type> dur=Xms proc=Yms target=<el>` for every
user interaction whose latency exceeds 50ms — always on, idle or under load — and
names the element (id / first class / data-action / aria-label / title). This is
the direct click->reaction latency the user feels.
- The longtask observer now also logs `renderer-longtask dur=Xms` immediately for any
single renderer long task >=100ms, regardless of upload state.
405 tests pass; clean Electron boot. Every action — main or renderer, idle or under
load — now names itself in the log if it is slow.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The v3.3.98 instrument exposed the actual cause of the 1-2s UI button lag, and it
was none of the read-path suspects. electron-config.json had grown to 38.5MB and
was being load()/structuredClone()/JSON.stringify()'d ~137 times in a 73s window
(140-592ms each) on the Electron main thread — roughly 47% main-thread occupancy.
That is the lag: a button click lands while the thread is mid-clone of a 38.5MB
object. The 1MB read-ahead could not touch it.
The bulk is history, not the queue. Each batch-done appended the upload-manager
summary verbatim, including the full per-file result list (per-hoster URLs), into
config.history; default historyRetention='all' never prunes, so it grew unbounded
(75 batches ≈ 38.5MB). The "queue=undefined" in the perf log was a logging bug
(reading .length on the pendingQueue object), not an empty queue. Writes drove the
storm: queue-persistence (save-global-settings) did two loads + one 38.5MB
serialize per call, and _atomicWrite nulls the cache so the next load is a full
38.5MB reparse; even cache hits structuredCloned the whole 38.5MB. The per-job
upload path makes zero config calls, so pending=1280 was never the driver.
Fix — move history into its own file so it is never parsed/cloned/serialized on a
config op (a 5-agent investigation + adversarial review chose this over three
read-side half-fixes; it is the only change that removes the clone AND the
serialize AND the post-write reparse at once):
- History lives in electron-history.json. _migrateHistory() runs once at init
(packaged only) and is fail-safe: it writes history.json (tmp → fsync → rename),
re-reads and verifies the entry count, and keeps a permanent
electron-config.json.pre-history-split.bak BEFORE the config is ever allowed to
drop its history. If verification fails it leaves history in the config (retry
next launch). _loadImpl returns history:[] once migrated, so the cached object is
tiny (cheap clones); the config file shrinks to ~KB on the first save (cheap
reparse) and _serializeForDisk writes ~KB (cheap serialize).
loadHistory/appendHistory/pruneHistory/clearHistory go through history.json on
their own write-queue with a no-clobber guard; the legacy config path stays as a
fallback when migration did not run.
- Validated against the real 185MB / 30000-entry bench fixture: migrate 1.5s once,
every entry preserved + .bak kept; load() 631ms cold once → 0.1ms after the first
save strips the file (185MB → 2.1KB); loadHistory() still returns all 30000.
Dropped per review (one-variable + risk): loadShallow (moot after the split),
cache-repopulate (its gate can never fire), and a per-batch resolution cache (stale
account pools → the rotation/byse failover-regression class — the one thing that
could silently corrupt uploads).
Instrumentation (the user asked to measure everything; all additive, threshold-
gated, MHU_PERF=0 disables):
- ipcMain.handle/.on are centrally wrapped to log `ipc <channel> wall=Xms sync=Yms`
over 50ms — the button-press→response latency — hardened (Promise.resolve(p)
.finally + try/catch'd logging) so a logging failure can never break IPC.
- A 100ms main-process drift monitor logs `main-longtask blocked=Xms lastIpc=… gc=…`
for any single main-thread turn over 100ms (catches GC, fs scans, serialize that
the IPC timing structurally cannot see).
- config-store perf lines gain via=<caller> and wqDepth=, and the queue= logging
bug is fixed (now reads pendingQueue.queueJobs.length).
405 tests pass (9 new migration tests: preserve-count, round-trip, save() never
loses history, crash-window fallback, idempotency). Clean Electron boot, no repo
pollution (migration is packaged-only).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
v3.3.97 (UV_THREADPOOL_SIZE 64→8) was a decisive win — mean event-loop-delay at
70 active uploads dropped 200ms→~11ms (18×), rss 577→287MB, renderer healthy in
14/15 windows. But the user reports it is still not perfectly smooth. A focused
multi-agent investigation plus an adversarial review localized the residual to
TWO distinct, separately-measured spike sources:
1. Read-bursts. In the tail windows the file-read histogram inverts: FSReqCallback
climbs to 66-70 against threadpool=8 (~8.75× queue depth) while SimpleWriteWrap
(socket writes) collapses to 4-24 and mean delay rises to 30-42ms. GC is ruled
out (gcMax ≤27ms in every window). The clean inversion at a stable active=70 /
pending=1287 shows the reads are causal, not a symptom of a block elsewhere.
2. A suspected synchronous config-persist stall. save() → load() reparses the whole
electron-config.json — which now carries the 1287-job pending queue nested in
globalSettings plus full history — on every persist (because _atomicWrite nulls
the read cache), then _serializeForDisk JSON.stringify(…, null, 2) of all of it.
One tail sample (max 1021ms, heap spiking to 142MB) fits a large synchronous
structuredClone+stringify, but it is a single confounded point, so this build
only INSTRUMENTS the path rather than asserting the cause.
This release ships one behavioral change (kept to a single variable so the next
log attributes cleanly) plus measurement:
- highWaterMark 256KB→1MB in all five streaming read loops (lib/hosters.js,
doodstream/voe/vidmoly CHUNK_SIZE consts, and the inline value in
clouddrop-upload.js:108 — NOT the 16MB server chunk at clouddrop-upload.js:12).
UV_THREADPOOL_SIZE stays 8. This deepens each stream's read-ahead cushion from
~0.43s to ~1.7s at the per-stream rate, so a stream tolerates the threadpool
queue without starving its socket write, and cuts read-completion callbacks and
per-chunk Buffer allocations ~4×. Byte-correctness is unaffected: Content-Length
is preamble+fileSize+epilogue, independent of chunk size, and the chunk size
never touches the multipart boundaries. Fully reversible; a dedicated read-
concurrency semaphore is held in reserve if 1MB does not clear the bursts.
- config-store.js now times load() (the full reparse, which the account-failed
handler also hits per failure) and the _commit serialize, logging
`config-load …` / `config-serialize wall=…ms bytes=… hist=… queue=…` when the
synchronous work exceeds 20ms. load() is split into a timing wrapper + _loadImpl;
the timer is a no-op until main.js wires configStore.setPerfLog → logInfo.
The renderer batch-drain fix for the one observed 243ms longtask is intentionally
deferred: that jank is downstream of the main-thread read-burst flooding IPC, so
fix#1 should make it self-heal; bundling it would confound the measurement and
touch the progress hot path. All 397 tests pass.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The v3.3.96 eventloop-delay logs from a real 70-connection run pinpointed the
high-concurrency lag to the file-read path. At a CONSTANT active-job count the
process flips between two clean regimes:
HEALTHY ELD ~11ms, rss 268-308MB: SimpleWriteWrap ≈ active, FSReqCallback ≈ 0-1
BLOCKED ELD 49-217ms, rss 540-610MB: FSReqCallback ≈ active (62-71 reads in
flight), SimpleWriteWrap ≈ 0-4
Both the ELD spike and the rss balloon track FSReqCallback (libuv-threadpool file
reads) exactly — not crypto, not the renderer, not GC alone. All five uploaders
read identically via fs.createReadStream({highWaterMark: 256KB}); byse/doodstream/
voe run through the generic uploadFile in hosters.js (no dedicated module).
This build is both a candidate fix and a discriminator, per advisor review:
- UV_THREADPOOL_SIZE 64→8 (main.js:1). One reversible line, NOT an upload cap —
70 uploads still run. 8 concurrent 256KB reads sustain ~100MB/s, far above the
41MB/s aggregate, so it cannot bottleneck throughput even on the slow VM disk.
Strong suspicion that tp=64 made it worse: it removed the natural read-
serialization (default 4 threads) and let all 70 streams' reads fire at once,
flooding the loop with completion callbacks in lock-step bursts.
- ELD line now also logs heap=heapUsed ext=external ab=arrayBuffers and
gc=/gcTotal=/gcMax=ms (PerformanceObserver entryTypes:['gc'], reset per window).
The 70×256KB ≈ 18MB of read buffers cannot account for the ~300MB rss swing —
that is heap/object churn, so GC must be measured directly.
Decision rule for the next real-run log:
- ELD drops with tp=8 → read over-parallelism confirmed (keep 8 or add
a dedicated read-semaphore).
- ELD high + GC pauses align → heap churn, hunt the allocator.
- ELD high + GC flat → causation was reversed, pivot.
No upload-behavior change; the concurrency cap the user explicitly rejected is
untouched.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The v3.3.94 measurement computed the main-process event-loop delay, CPU%, per-hoster connection distribution and transient error counts, but logged only '[INFO ] perf': logInfo(ctx, msg) treats a string first arg as the whole message and discards the second, so the payload was thrown away (confirmed in the user's upload-debug.log). Pass the line as a single argument so the real numbers reach the log.
Also: assets/ was missing from the electron-builder files list, so app_icon.ico was never packaged into the asar — new Tray() threw an unhandled rejection on every startup and left no tray icon. Package the icons and wrap createTray in try/catch with a nativeImage fallback so a missing/invalid icon can never crash tray creation.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>