[BUG]: execProtocolRawSync spins forever at 100% CPU when the backend exits mid-message (swallow-all catch in the protocol loop)
Describe the bug
When the WASM backend terminates while execProtocolRawSync is processing a message (e.g. exit(1) after hitting EOF during a COPY ... FROM STDIN issued via pglite.exec()), the protocol loop spins forever, synchronously, at 100% CPU. The call never returns, so no timer, Promise.race, or AbortSignal can interrupt it (the same structural observation made in #945), and the Node process has to be killed from outside.
Root cause, packages/pglite/src/pglite.ts, execProtocolRawSync:
while (
this.#readOffset < message.length ||
mod._pq_buffer_remaining_data() > 0
) {
try {
mod._PostgresMainLoopOnce()
} catch (e: any) {
if (e.status === this.POSTGRES_MAIN_LONGJMP) {
mod._PostgresMainLongJmp()
}
// all other exceptions are swallowed and the loop continues
}
}After the backend exits, _PostgresMainLoopOnce() throws ExitStatus (status = 1). That is not the longjmp sentinel, so it is silently swallowed; the dead backend consumes no input, so #readOffset never advances and _pq_buffer_remaining_data() never drains — the while condition is permanently true. The result is a tight synchronous call → throw → swallow loop. The event loop never runs again.
Two aggravating factors:
- The
finally(which restoresprocess.exitCode, added for #975) never runs, because the loop never exits. - There is no death detection: after
exit(1),readyremainstrueandclosedremainsfalse, so callers cannot observe the crash even where the loop is escaped.
COPY FROM STDIN via exec() is just the easiest structural trigger — any exception from _PostgresMainLoopOnce that is not the longjmp sentinel arms the same loop (e.g. a WASM RuntimeError from an OOM abort, or the FATAL: terminating connection because protocol synchronization was lost class seen in #820).
To Reproduce
import { PGlite } from '@electric-sql/pglite';
const p = new PGlite();
await p.exec('CREATE TABLE t(a int)');
// Backend hits EOF mid-COPY and exits — exec() never settles, process pegs a core:
await p.exec('COPY t FROM STDIN');
console.log('never reached');A transaction is not required; wrapping the COPY in BEGIN spins identically.
Logs
V8 sampling profile of the spinning process (node --prof, ~20k ticks): 99.9% of samples are inside execProtocolRawSync → two wasm frames, with C++ time dominated by exception allocation/stack capture — consistent with the throw-per-iteration loop above.
Details
- PGlite version: 0.5.4 (also present unchanged in
main'spglite.tsat time of filing) - Extensions: none
- OS: macOS 15 (Darwin 25.3.0)
- Runtime: Node 24.14.1
Suggested fix (verified against a patched 0.5.4 dist)
Rethrow anything that is not the longjmp sentinel, and guard the finally's own WASM calls so they cannot mask the original error against a dead runtime:
} catch (e: any) {
if (e.status === this.POSTGRES_MAIN_LONGJMP) {
mod._PostgresMainLongJmp()
} else {
throw e // ExitStatus, RuntimeError, … — the backend is gone; looping cannot help
}
}
// ...
} finally {
try {
mod._PostgresSendReadyForQueryIfNecessary()
mod._pgl_pq_flush()
} catch {
// dead runtime — the rethrown error above is the one that matters
}
pglUtils.pgliteProc.exitCode = prevExitCode
}With exactly this change applied to the 0.5.4 dist, the repro behaves well in both variants (with and without an open transaction):
- the killing
exec()rejects withExitStatus { status: 1 }("Program terminated with exit(1)") - follow-up queries reject immediately with the same error instead of spinning
close()resolves- the event loop stays alive throughout
Beyond the minimal fix, it would help to mark the instance dead on ExitStatus (flip ready/closed or expose a crashed flag) so callers get a coherent state instead of ready === true on a dead instance, and/or to add a progress guard to the loop (bail if an iteration neither advanced #readOffset nor drained the pq buffer) as defense in depth.
Additional context
- #945 — synchronous WASM hang; documents that
Promise.race/timers cannot interrupt these loops (different root cause: Gather workers) - #975 —
process.exitCode = 99leak; introduced thefinallythat this spin prevents from running - #1052 — protocol parser wedged after a throw in
parse(); same method neighborhood, different failure mode - PR #351 (closed WIP) — test case for unclean shutdown after
DROP DATABASE; its description already warns "this test will hang in a busy loop and CI will continue waiting with no timeout". Same symptom family, never filed as a bug. - #820 — an in-the-wild lever for the same backend-death class (oversized
CREATE TYPEvia pglite-socket →FATAL, backend gone mid-message)
Source: electric-sql/pglite