Multi-Hoster-Upload/tasks/todo.md
Administrator 9cc8fee02c docs(todo,lessons): long-run lag root cause was the recent-panel append-gate cliff (measured 80ms->7ms)
Records the verified root cause + the dismissed-with-evidence findings:
- _sessionFileKeys "separator mismatch" = FALSE POSITIVE (real U+0001 chars, verifier Read rendered them invisibly)
- queueJobs O(N) per-render scan = real but Blink-measured <0.1ms at 3000 jobs -> skipped
- standing 2000-row relayout = median 0.4ms -> no virtualization needed
- doodstream sync _debugLog = constant freeze, deferred to a separate change

Lesson: measure magnitudes at the real artifact before fixing a "scales-with-X" cause;
profile in real Blink (Playwright) not jsdom; verify multi-agent findings against primary
evidence (char-code dump caught the control-char false positive).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-21 01:47:59 +02:00

53 lines
3.9 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Long-running lag — root cause + fix (symptom: UI laggy over time, CPU 40%/RAM 6GB both normal)
## What "laggy after time, low CPU, stable RAM" actually was (MEASURED, not assumed)
First instinct (main-process config I/O scaling with history) was WRONG for this user: the
real config is tiny (history 23 rows / 4.8 KB, total 52 KB, queue ~153 jobs). Measured with a
one-off node read of the live electron-config.json. So serialize-cost theories were dead on arrival.
A 13→18-agent leak hunt + adversarial verify + a Playwright/Blink microbenchmark found the
真 cause:
### ROOT CAUSE (confirmed + profiled): recent-uploads panel append-only path defeats itself at the cap
- `renderRecentUploadsPanel` had a cheap append-only fast path gated on `rows.length > _recentLastRenderedLen`.
- `maybeAddSessionFile` caps `sessionFilesData` by push-then-slice (2000 → 2001 → sliced back to 2000).
- So once past SESSION_FILES_CAP, `rows.length` is pinned at 2000 → the gate is FALSE forever →
EVERY completion fell through to `tbody.innerHTML = rows.map(...).join('')` — a full ~2000-row
rebuild. Cap is per (link × file × hoster), so 45 hosters hit 2000 at only ~400500 files.
- Blink measurement (table-layout:fixed, same engine as Electron): full 2000-row rebuild = **~80 ms**
on EVERY completion past the cap. At several completions/sec that is a repeating ~80 ms main-thread
freeze → exactly "fine fresh, gets laggy after many uploads, CPU/RAM fine."
## Fix (shipped) — append-evict, keeps the panel append-only past the cap
- Track newly-pushed rows in `_recentPendingAppends` (incremented in maybeAddSessionFile), consumed
every render. Gate the fast path on `pendingAppends > 0` (not length-delta) so it survives the cap.
- Prepend the new rows, then evict the same overflow count from the DOM bottom (oldest, = data front
eviction in date-desc) to honour the cap. DOM work back to O(added).
- Gate behind an explicit `appendOnly` flag passed ONLY by scheduleRecentRender's rAF, so selection/
delete/clear/sort/batch-done renders stay full + correct (no wrong-row eviction, no double-prepend).
- VERIFIED in Blink over a simulated 5000-completion session: per-frame render **80 ms → median 7.4 ms**
(>10×), DOM stays exactly == data (cap held, newest-on-top, oldest evicted, ZERO duplicates).
## Also shipped earlier (29d1944) — defensible, but NOT the cause for this user
- T1 ConfigStore in-memory cache (mtime/size keyed) + lean _serializeForDisk (clones only hosters) +
copyFileSync .bak. T3 cache logMode/logFilePath (drop load() from the 500 ms log flush).
## Investigated and DISMISSED with evidence
- `_sessionFileKeys` "delete-key separator mismatch" (workflow flagged it HIGH-ish): FALSE POSITIVE.
Line uses REAL U+0001 chars (char-code dump: `HAS_U0001: True`); the verifier agents' Read rendered
the control chars invisibly and wrongly concluded "no separators". Keys match at runtime. No leak.
- queueJobs O(N) per-render scans (grows unbounded since removeFromQueueOnDone=false AND folder-monitor
keeps ONE batch alive via addJobs so the 500-cap prune never fires): REAL but Blink-measured at
**<0.1 ms even at 3000 jobs** imperceptible. Incremental-counter refactor NOT worth the risk. Skipped.
- "Continuous standing relayout of 2000 rows": Blink median 0.4 ms regardless of row count; its spikes
were caused by the rebuild (same root cause). Not a standing cost no need to virtualize/lower cap.
## Deferred (separate, not this symptom)
- doodstream-upload.js `_debugLog`: ungated SYNC fs.appendFileSync + statSync ~1020×/upload on the
main loop. CONSTANT cost (does not grow over a session), only when doodstream is active. Real freeze
contributor but NOT the reported progressive lag convert to buffered-async like main.js, separately.
## Verification summary
- 397/397 tests pass, eslint 0 errors.
- Blink before/after: recent-panel render 80 ms 7.4 ms; correctness asserted (cap/order/no-dupes).