fix(telemetry): preserve fatal process semantics - #3589
Open
GautamSharma99 wants to merge 1 commit into
Open
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Fixes #3579.
Telemetry previously registered
uncaughtExceptionandunhandledRejectionlisteners that recorded the failure and drained events, but never restored the runtime fatal outcome. In Bun, the presence of those listeners suppresses default termination, so a long-running letta-code process could continue after an uncaught failure with partially mutated state or broken invariants.This PR makes fatal telemetry best-effort without making fatal errors recoverable.
What changed
installFatalErrorHandlershelper.process.exitCode = 1synchronously as soon as either fatal event arrives.process.exit(1)after the drain settles, fails, or times out, ensuring active sockets and timers cannot keep the process alive.telemetry.cleanup(), preventing duplicate process listeners across cleanup/reinitialization cycles.Why explicit termination
Setting only
process.exitCodeis insufficient for listener/headless processes because active handles can keep the event loop alive indefinitely. Removing the listener and rethrowing also creates recursive/error-ordering hazards while an asynchronous flush is running. A bounded flush followed by explicit non-zero termination gives telemetry a short delivery window while preserving deterministic fatal semantics.The exit code is set before tracking or draining so failures inside telemetry itself cannot accidentally produce a successful exit. Tracking and drain errors are deliberately contained because neither should interfere with termination.
Tests
Added real child-Bun-process coverage rather than mocking
process.exit:Validation completed:
bun test src/telemetry/fatal-error-handler.test.ts— 6 passedbun run check— all 12 repository checks passedScope
Normal SIGINT shutdown retains its existing successful exit behavior. This change only affects uncaught exceptions and unhandled rejections, which now reliably remain fatal after telemetry is initialized.