Commit 9666fb4
duanyan.duan
fix(blob): allow blob files with different schema ids in a bunch
BlobBunch::Add required every file in a bunch to share the same schema id.
After schema evolution that adds/drops OTHER columns, a single blob field's
contiguous files can carry different schema ids (an old appended file plus a
compacted file), and MergeRangesAndSort groups them into one bunch by row id,
raising a spurious "All files in a blob bunch should have the same schema id."
Paimon Java (SpecialFieldBunch.add) guards this check with if (!isBlobFile(file))
so blob files are exempt -- a blob column's on-disk layout is schema-independent
and the bunch is read with the first file's schema regardless. The guard was
dropped when porting to C++. Restore it and add a regression unit test.1 parent ffbda4f commit 9666fb4
2 files changed
Lines changed: 36 additions & 4 deletions
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
94 | 94 | | |
95 | 95 | | |
96 | 96 | | |
97 | | - | |
98 | | - | |
| 97 | + | |
| 98 | + | |
| 99 | + | |
| 100 | + | |
| 101 | + | |
| 102 | + | |
| 103 | + | |
| 104 | + | |
| 105 | + | |
| 106 | + | |
| 107 | + | |
| 108 | + | |
99 | 109 | | |
100 | 110 | | |
101 | 111 | | |
| |||
Lines changed: 24 additions & 2 deletions
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
84 | 84 | | |
85 | 85 | | |
86 | 86 | | |
87 | | - | |
| 87 | + | |
88 | 88 | | |
89 | 89 | | |
90 | 90 | | |
91 | 91 | | |
92 | | - | |
| 92 | + | |
93 | 93 | | |
94 | 94 | | |
95 | 95 | | |
| |||
247 | 247 | | |
248 | 248 | | |
249 | 249 | | |
| 250 | + | |
| 251 | + | |
| 252 | + | |
| 253 | + | |
| 254 | + | |
| 255 | + | |
| 256 | + | |
| 257 | + | |
| 258 | + | |
| 259 | + | |
| 260 | + | |
| 261 | + | |
| 262 | + | |
| 263 | + | |
| 264 | + | |
| 265 | + | |
| 266 | + | |
| 267 | + | |
| 268 | + | |
| 269 | + | |
| 270 | + | |
| 271 | + | |
250 | 272 | | |
251 | 273 | | |
252 | 274 | | |
| |||
0 commit comments