diff --git a/docs/sql-migration-guide.md b/docs/sql-migration-guide.md index f9c7bf477122..f24b90f05b9d 100644 --- a/docs/sql-migration-guide.md +++ b/docs/sql-migration-guide.md @@ -44,6 +44,7 @@ license: | - Since Spark 4.4, `spark.sql.sources.v2.bucketing.partition.filter.enabled` defaults to `true`. In a storage-partitioned join, the partition key groups pushed down to both sides may now be narrowed to those that can produce output for the join type, instead of always taking the union of the two sides' groups: an inner or semi join may keep only the groups present on both sides, a join that keeps or tests every left row (left outer, left anti, left single, existence) keeps the left side's groups, a right outer join keeps the right side's, and a full outer join is unaffected; a CROSS join carrying an equality condition is treated like an inner join. The narrowing is skipped when either side's partitioning may contain unknown partition keys, as after a shuffle on one side, so such a join keeps the full union. A group that is dropped is not scanned at all, so a query may read fewer files. To restore the previous behavior, set `spark.sql.sources.v2.bucketing.partition.filter.enabled` to `false`. - Since Spark 4.4, `spark.sql.sources.v2.bucketing.partitionKeyOrdering.enabled` defaults to `true`. A V2 scan that reports a keyed partitioning but no explicit ordering now also reports itself sorted by its partition key expressions, because every row of such a partition evaluates them to the same value. Partition transforms such as `days(ts)` or `bucket(8, id)` are left out of that ordering. This lets Spark drop a `Sort` it would otherwise place above the scan. An aggregate that groups by a leading run of the other key expressions, in key order, may also be planned as a sort aggregate rather than a hash aggregate. This is because `spark.sql.execution.replaceHashWithSortAgg` keys off the reported ordering. To restore the previous behavior, set `spark.sql.sources.v2.bucketing.partitionKeyOrdering.enabled` to `false`. - Since Spark 4.4, `spark.sql.sources.v2.bucketing.preserveKeyOrderingOnCoalesce.enabled` defaults to `true`. When `GroupPartitionsExec` merges several input partitions that share one partition key value into a single output partition, it now keeps sort orders over the partition key expressions in the ordering it reports, since those expressions are constant within the merged partition. Sort orders over partition transforms such as `days(ts)` or `bucket(8, id)` are left out. Orders over other columns are still dropped, as the concatenation invalidates them, and a join that reduced the partition keys onto a common key space (see `spark.sql.sources.v2.bucketing.allowCompatibleTransforms.enabled`) reports no order, since the merged partitions then share only the reduced key. This can remove a `Sort` downstream of the merge. To restore the previous behavior, set `spark.sql.sources.v2.bucketing.preserveKeyOrderingOnCoalesce.enabled` to `false`. +- Since Spark 4.4, when the `from_protobuf` option `unwrap.primitive.wrapper.types` is enabled, `from_protobuf` unwraps a present Protobuf primitive wrapper (such as `google.protobuf.Int32Value` or `StringValue`) whose inner `value` is unset or set to its default to the type's default value (`0`, `false`, `""`, or empty binary), regardless of the `emit.default.values` option. Previously such a wrapper unwrapped to `NULL` unless the `emit.default.values` option was enabled, which could also place a `NULL` into a non-nullable array element or map value. An absent wrapper field still unwraps to `NULL`. ## Upgrading from Spark SQL 4.2 to 4.3