Skip to content
Open
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
160 changes: 160 additions & 0 deletions sips/finality_based_event_syncing.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,160 @@
| Author | Title | Category | Status | Date |
|--------|------------------------------|----------|---------------------|------------|
| xxx | Finality-Based Event Syncing | Core | open-for-discussion | 2025-05-27 |

## Summary

This SIP proposes transitioning from the current `follow-distance` approach to a `finality-based` approach for syncing
Ethereum events. After a configurable fork epoch, SSV nodes will process events only from blocks confirmed as
`finalized`
by the consensus layer. This transition will provide stronger cryptographic guarantees against most chain
reorganizations, significantly enhancing the reliability of event processing while maintaining a clear path for backward
compatibility during the upgrade.

## Motivation

The current event syncing mechanism has critical issues:

1. **Vulnerability to Reorganizations**: Relying on non-finalized blocks (which are inherently less reliable until
finality) exposes the system to potential state inconsistencies if processed blocks are later reorged out.
2. **False Sync Failures**: The system can incorrectly assume a sync failure and crash when detecting old execution
layer block times, even when this is due to the beacon chain simply not proposing new blocks (e.g., due to network
stalls or low participation).

These issues compromise SSV's reliability and can cause state divergence between operators.

## Rationale

### Current State: Follow Distance

- Processes blocks N blocks behind head (default: 8)
- Assumes probabilistic finality
- Vulnerable to deep reorganizations

### Proposed: Finality-Based Syncing

- Uses Ethereum's finalized blocks
- Significantly reduces reorganization risks by relying on cryptographically finalized blocks
- Aligns with Ethereum's security model

## Specification

### Behavioral Changes

This change will be activated as part of a network-wide fork. The specific fork mechanism and activation epochs will be
defined separately as part of the broader network upgrade.

**Pre-Fork (Alan)**:

- Continue using the existing follow-distance mechanism (N blocks behind head, default N=8)
- Process events from blocks assumed to be probabilistically final
- Existing behavior remains unchanged

**Post-Fork (FinalityConsensus)**:

- Query finalized blocks directly from the execution layer client using calls `eth_getBlockByNumber("finalized")`
Comment thread
kchojn marked this conversation as resolved.
- Process SSV contract events exclusively from these finalized blocks.

### Key Implementation Points

1. **One-way Transition**: Once the fork activates, nodes will exclusively use finality-based syncing with no fallback.
2. **Automatic Detection**: Nodes will detect fork activation and transition automatically.

## Visual Overview

The following diagrams illustrate the key concepts of this proposal.

### 1. Fork State Transition Diagram

The one-way transition from PreFork (follow-distance) mode to PostFork (finality) mode based on the beacon chain epoch.

```mermaid
stateDiagram-v2
direction LR
FollowDistance --> Finality: Fork Activation
note left of FollowDistance: 96 sec delay<br/>Reorg risk
note right of Finality: 13 min delay<br/>No reorg risk
```

### 2. Event Processing Flow Comparison

The event processing flow before and after the fork.

```mermaid
graph TB
subgraph "Proposed"
A2[New Block] --> B2[Check Finality]
B2 --> C2[Wait Finalized]
C2 --> D2[Fetch Events]
D2 --> E2[Safe Process]
end

subgraph "Current"
A1[New Block] --> B1[Wait 8 Blocks]
B1 --> C1[Fetch Events]
C1 --> D1[Process]
D1 --> E1[RISK: Reorg]
end
Comment on lines +84 to +97

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This flow diagram includes nodes that are not really part of the flow ("Safe Process", "RISK: Reorg").

The trigger for each diagram is "New Block". This models the new process incorrectly IMO. Check Finality sounds like there are two possible outcomes (final, not final), but there is only one arrow.

I'd just put something like

graph TB
    subgraph "Proposed"
        A2[New finalized checkpoint] --> B2[Fetch events of new finalized blocks]
    end
Loading

to reflect that the trigger for sync progress is a new finalized checkpoint.

```

### 3. Event Processing Timeline Comparison

The difference in event processing latency between the pre-fork and post-fork mechanisms.

```mermaid
gantt
title Event Processing Timeline
dateFormat X
axisFormat %s

section PreFork
Block Produced: done, pre1, 0, 12
Follow Distance Wait: active, pre2, 12, 96
Process Events: crit, pre3, 96, 10

section PostFork
Block Produced: done, post1, 0, 12
Wait for Finality: active, post2, 12, 768
Process Events: crit, post3, 768, 10
```

## Performance Impact

- **Event Lag**: The time between an event occurring on-chain and it being processed by the SSV node will likely
increase. Using the current follow distance (e.g., 8 blocks * ~12s/block ≈ 1.6 minutes) versus waiting for finality (
typically 2 epochs, so ~12.8 minutes) represents a significant change. This is a crucial tradeoff: longer lag for
guaranteed event stability.

## Backwards Compatibility

- Nodes that are not upgraded before the "FinalityConsensus" fork activation epoch will continue to use the
follow-distance mechanism.
- Such non-upgraded nodes risk processing events from blocks that are subsequently reorganized and not part of the
finalized chain, leading to potential state divergence from upgraded nodes, especially in edge cases involving chain
instability.

## Security Considerations

- **Enhanced Security**: Transitioning to cryptographic finality for event processing inherently strengthens the node
against state corruption or inconsistencies caused by execution layer reorganizations.
- **Trust Model**: Reliance shifts more explicitly towards the finality guarantees provided by Ethereum's consensus
layer, which is a core security assumption of the PoS network.
- **Attack Surface**: This change is expected to reduce the attack surface related to manipulating node state through EL
reorgs.
Comment on lines +138 to +143

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This does not mention the significant disadvantage that it becomes impossible to remove validators during long periods of non-finality. See my comment below on my ideas how to mitigate this.


## Test Scenarios

1. Verification of smooth fork transition during active syncing and under various network conditions.
2. Node behavior when finalized blocks are temporarily unavailable or significantly delayed from the consensus layer.
3. Performance analysis of event processing lag under normal and stressed network conditions post-fork.
4. Monitoring and alerting functionality for the new finality-based sync status.

## Migration Guide

1. **Pre-Fork**: Node operators must update their SSV node software to a version supporting the "FinalityConsensus" fork
well in advance of the announced activation epoch for their respective network. Monitor official announcements for
activation epoch details.
2. **During Fork Activation**: The transition should be automatic once the node's current epoch reaches the
"FinalityConsensus" activation epoch. Operators should monitor node logs for confirmation of the new operational mode.
3. **Post-Fork**: Verify that the node is processing events based on finalized blocks. Node operators may implement
their own monitoring solutions as needed.
Comment on lines +154 to +160

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is very vague. When exactly do we stop syncing with the follow method? There are multiple ways this could be interpreted:

Let X be the first slot of the activation epoch. Let B(slot) be the head block of slot (which may be a block in a previous slot if slot is empty. Let E(slot) be the epoch of slot.

a) We simply stop syncing at X: the last block processed will be B(X - 1) - 8 until E(X - 1) finalizes.
b) We stop syncing after processing B(X - 1): basically syncing until the follow distance has caught up with the epoch end.

Which one do you mean?