fix(data-evolution): require row tracking and row IDs - #718
Merged
JingsongLi merged 7 commits intoAug 17, 2026
Merged
Conversation
XiaoHongbo-Hope
marked this pull request as ready for review
August 16, 2026 08:13
JingsongLi
requested changes
Aug 16, 2026
JingsongLi
left a comment
Contributor
There was a problem hiding this comment.
Row tracking validation is incomplete. It fails to enforce three Java-side constraints: Primary Key is prohibited, bucket must be -1, and clustering.incremental is disallowed for Data Evolution. However, empirical testing confirms that a schema can be successfully created with the configuration PK + bucket=1 + row tracking + data evolution.
Mirror Java SchemaValidation.validateRowTracking: reject primary keys and bucket != -1 for row tracking tables, and reject clustering.incremental with data evolution. Downstream tests that relied on creating such tables now assert the create-time rejection.
The hybrid global fixture dropped the now-invalid bucket=1 option, and the Lumina primary-key guard moved before option checks so plain primary-key tables (including legacy invalid schemas) keep their dedicated rejection.
jerry-024
added a commit
to jerry-024/paimon-rust
that referenced
this pull request
Aug 17, 2026
* main: perf(vindex): size native batches by active indexes (apache#709) fix(data-evolution): require row tracking and row IDs (apache#718) feat(go): add table write bindings (apache#658) fix(scan): preserve Data Evolution file order in row-id groups (apache#717) fix(file_index): align file index format with Java V1 (apache#719) fix: configure OpenDAL writer chunk size (apache#713) fix(python): release GIL during catalog I/O (apache#716) feat(write): add fixed-bucket write primitives for postpone tables (apache#659) feat: rust examples for creating and querying Paimon tables (apache#648) fix(table): reject row ranges for format tables at read construction (apache#700) perf(arrow): prune IN predicates with row-group stats (apache#705) feat(io): support in-memory local cache (apache#710) fix(dlf): refresh expiring credentials (apache#714) fix(datafusion): normalize index_type in global index procedures (apache#715) fix(auth): fail closed on query-auth tables outside the read boundary (apache#691) # Conflicts: # crates/paimon/src/table/vector_search_builder.rs # crates/paimon/src/vindex/reader.rs
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Reject invalid Data Evolution tables and files before they can produce unreadable data.
Java requires data-evolution.enabled to be used with row-tracking.enabled. Rust previously allowed CREATE or schema changes without row tracking, so writes succeeded without first_row_id and later scans failed. This PR adds the same schema validation for creation and SchemaChange.
Scan planning also validates first_row_id before row-range grouping. This keeps legacy invalid tables fail-closed and aligns planning with Java DataFileMeta.nonNullRowIdRange and PyPaimon non_null_row_id_range.
Tests