Commit d5133a8
authored
feat: add lance_dataset_drop_columns for metadata-only column removal (#42)
## Summary
First of three PRs against #41 (schema evolution). Exposes upstream's
`drop_columns` — a metadata-only manifest commit that removes the named
columns from the schema without rewriting any data files. Materializing
the projection is left to a later `_compact_files` (and a future cleanup
operation, once exposed, removes the old version's files).
Mutates the dataset in place under an exclusive write lock; scanners
already in flight keep their pre-drop snapshot view via the existing Arc
clone-on-write, same as `_delete` / `_update` / `_compact_files`.
## Surface
```c
int32_t lance_dataset_drop_columns(
LanceDataset* dataset,
const char* const* columns,
size_t num_columns
);
```
Inputs are validated up front with per-index error messages so the
precise cause is observable from `lance_last_error_message()`. NULL
handle, NULL pointer array, zero count, NULL or empty-string entries,
and non-UTF-8 names all return `LANCE_ERR_INVALID_ARGUMENT`; upstream's
own rejections (unknown column, attempt to drop every column) map to the
same code.
The C++ wrapper takes `const std::vector<std::string>&` and follows the
`update` / `merge_insert` sibling convention — passes `col_ptrs.data()`
unconditionally. An empty vector flows through the Rust-side
`num_columns == 0` guard so the error message says "num_columns must be
> 0" rather than the misleading "columns must not be NULL".
## Tests
Eleven new Rust integration tests covering single-drop, multi-drop,
version bump, data preservation (downcasts the surviving Arrow columns
and checks the actual values, not just shape), and the full rejection
surface (NULL dataset / NULL array / zero count / NULL entry /
empty-string entry / unknown column / drop-all). C and C++ smoke tests
snapshot `ArrowSchema.n_children` pre/post drop, exercise the
drop-last-column rejection path, and verify the version is unchanged
when a drop fails. `cargo test` and `cargo test --test
compile_and_run_test -- --ignored` both green.
## Follow-ups
- `lance_dataset_alter_columns` — rename / nullability / type change
- `lance_dataset_add_columns` — SQL expressions / AllNulls /
ArrowArrayStream
The README roadmap entry stays unticked until all three ship.1 parent bd01a95 commit d5133a8
7 files changed
Lines changed: 565 additions & 0 deletions
File tree
- include/lance
- src
- tests
- cpp
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
453 | 453 | | |
454 | 454 | | |
455 | 455 | | |
| 456 | + | |
| 457 | + | |
| 458 | + | |
| 459 | + | |
| 460 | + | |
| 461 | + | |
| 462 | + | |
| 463 | + | |
| 464 | + | |
| 465 | + | |
| 466 | + | |
| 467 | + | |
| 468 | + | |
| 469 | + | |
| 470 | + | |
| 471 | + | |
| 472 | + | |
| 473 | + | |
| 474 | + | |
| 475 | + | |
| 476 | + | |
| 477 | + | |
| 478 | + | |
| 479 | + | |
| 480 | + | |
| 481 | + | |
| 482 | + | |
| 483 | + | |
| 484 | + | |
| 485 | + | |
| 486 | + | |
456 | 487 | | |
457 | 488 | | |
458 | 489 | | |
| |||
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
437 | 437 | | |
438 | 438 | | |
439 | 439 | | |
| 440 | + | |
| 441 | + | |
| 442 | + | |
| 443 | + | |
| 444 | + | |
| 445 | + | |
| 446 | + | |
| 447 | + | |
| 448 | + | |
| 449 | + | |
| 450 | + | |
| 451 | + | |
| 452 | + | |
| 453 | + | |
| 454 | + | |
| 455 | + | |
| 456 | + | |
| 457 | + | |
| 458 | + | |
| 459 | + | |
| 460 | + | |
| 461 | + | |
| 462 | + | |
| 463 | + | |
| 464 | + | |
| 465 | + | |
440 | 466 | | |
441 | 467 | | |
442 | 468 | | |
| |||
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
| 1 | + | |
| 2 | + | |
| 3 | + | |
| 4 | + | |
| 5 | + | |
| 6 | + | |
| 7 | + | |
| 8 | + | |
| 9 | + | |
| 10 | + | |
| 11 | + | |
| 12 | + | |
| 13 | + | |
| 14 | + | |
| 15 | + | |
| 16 | + | |
| 17 | + | |
| 18 | + | |
| 19 | + | |
| 20 | + | |
| 21 | + | |
| 22 | + | |
| 23 | + | |
| 24 | + | |
| 25 | + | |
| 26 | + | |
| 27 | + | |
| 28 | + | |
| 29 | + | |
| 30 | + | |
| 31 | + | |
| 32 | + | |
| 33 | + | |
| 34 | + | |
| 35 | + | |
| 36 | + | |
| 37 | + | |
| 38 | + | |
| 39 | + | |
| 40 | + | |
| 41 | + | |
| 42 | + | |
| 43 | + | |
| 44 | + | |
| 45 | + | |
| 46 | + | |
| 47 | + | |
| 48 | + | |
| 49 | + | |
| 50 | + | |
| 51 | + | |
| 52 | + | |
| 53 | + | |
| 54 | + | |
| 55 | + | |
| 56 | + | |
| 57 | + | |
| 58 | + | |
| 59 | + | |
| 60 | + | |
| 61 | + | |
| 62 | + | |
| 63 | + | |
| 64 | + | |
| 65 | + | |
| 66 | + | |
| 67 | + | |
| 68 | + | |
| 69 | + | |
| 70 | + | |
| 71 | + | |
| 72 | + | |
| 73 | + | |
| 74 | + | |
| 75 | + | |
| 76 | + | |
| 77 | + | |
| 78 | + | |
| 79 | + | |
| 80 | + | |
| 81 | + | |
| 82 | + | |
| 83 | + | |
| 84 | + | |
| 85 | + | |
| 86 | + | |
| 87 | + | |
| 88 | + | |
| 89 | + | |
| 90 | + | |
| 91 | + | |
| 92 | + | |
| 93 | + | |
| 94 | + | |
| 95 | + | |
| 96 | + | |
| 97 | + | |
| 98 | + | |
| 99 | + | |
| 100 | + | |
| 101 | + | |
| 102 | + | |
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
20 | 20 | | |
21 | 21 | | |
22 | 22 | | |
| 23 | + | |
23 | 24 | | |
24 | 25 | | |
25 | 26 | | |
| |||
37 | 38 | | |
38 | 39 | | |
39 | 40 | | |
| 41 | + | |
40 | 42 | | |
41 | 43 | | |
42 | 44 | | |
| |||
0 commit comments