You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit 8e1c55b
Browse filesBrowse the repository at this point in the historyBrowse files
[SPARK-60103][SQL] Keep XML MAP struct fields on convertField dispatch
Drop the direct convertObject MapType arm so empty, attribute-only, and
text-only MAP<STRING> elements stay null or malformed. Restrict convertMap
to StringType-family keys so format("xml") MAP<INT> does not ClassCast.
Copy file name to clipboardExpand all lines: docs/sql-migration-guide.md
+1-1Lines changed: 1 addition & 1 deletion
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -29,7 +29,7 @@ license: |
29
29
## Upgrading from Spark SQL 4.3 to 4.4
30
30
31
31
- Since Spark 4.4, when `spark.sql.preserveCharVarcharTypeInfo` is true and `spark.sql.charVarchar.standardSemantics.enabled` is false, ORC reads that apply a CHAR/VARCHAR schema over STRING storage return the stored values without ORC truncation, matching Parquet. Previously the ORC reader requested `char(n)`/`varchar(n)` and truncated STRING-stored values to `n`. Read-side length checks (`EXCEED_LIMIT_LENGTH`) apply only when `spark.sql.charVarchar.standardSemantics.enabled` is true.
32
-
- Since Spark 4.4, when `spark.sql.charVarchar.standardSemantics.enabled` is true, XML element and attribute names used as `MAP<CHAR(n), _>` or `MAP<VARCHAR(n), _>` keys in `from_xml` and the XML datasource are length-checked without padding or trimming. A `CHAR(n)` key must already be exactly `n` characters, and a `VARCHAR(n)` key must already be at most `n` characters. Mismatched keys fail with `UNSUPPORTED_XML_CHAR_VARCHAR_MAP_KEY` (SQLSTATE `0A000`) and follow the XML parse mode (`PERMISSIVE` or `FAILFAST`). Exact repeated names last-win; `spark.sql.mapKeyDedupPolicy` is not applied. XML CHAR/VARCHAR values still use write-side pad and `EXCEED_LIMIT_LENGTH`. With the flag off, CHAR keys are still padded and VARCHAR overflow still uses `EXCEED_LIMIT_LENGTH`.
32
+
- Since Spark 4.4, when `spark.sql.charVarchar.standardSemantics.enabled` is true, XML names used as `MAP<CHAR(n), _>` or `MAP<VARCHAR(n), _>` keys in `from_xml` and the XML datasource are length-checked without padding or trimming. The checked names are element names, `attributePrefix` plus attribute names (default prefix `_`), and the `valueTag` (default `_VALUE`) when mixed text is present. A `CHAR(n)` key must already be exactly `n` characters, and a `VARCHAR(n)` key must already be at most `n` characters. Mismatched keys fail with `UNSUPPORTED_XML_CHAR_VARCHAR_MAP_KEY` (SQLSTATE `0A000`) and follow the XML parse mode (`PERMISSIVE` or `FAILFAST`). Exact repeated names last-win; `spark.sql.mapKeyDedupPolicy` is not applied. XML CHAR/VARCHAR values use the same pad and `EXCEED_LIMIT_LENGTH` checks as other parsed text. Empty, attribute-only, and text-only map elements stay SQL NULL or a malformed record, matching 4.3; they do not become an empty map or a `valueTag` entry. In Spark 4.3, `MAP<CHAR(n), _>` and `MAP<VARCHAR(n), _>` XML maps were malformed records because only unbounded STRING keys were accepted. Padding of CHAR keys and VARCHAR `EXCEED_LIMIT_LENGTH` apply when the standard-semantics flag is false and `spark.sql.preserveCharVarcharTypeInfo` is true, so first-class CHAR/VARCHAR types still reach the parser.
33
33
- Since Spark 4.4, the options maps passed to `from_csv`, `to_csv`, `schema_of_csv`, `from_json`, `to_json`, `schema_of_json`, `from_xml`, `to_xml`, and `schema_of_xml` must be foldable after replacing `RuntimeReplaceable` expressions. Previously, Spark evaluated non-foldable options during analysis, which allowed some constant expressions but could fail with an internal error or incorrectly evaluate row-dependent expressions. To allow deterministic and row-independent non-foldable options, set `spark.sql.legacy.allowNonFoldableOptions` to `true`. Row-dependent, unevaluable, and nondeterministic options are always rejected.
34
34
- Since Spark 4.4, when an already-analyzed Data Source V2 query is refreshed after a compatible schema change, connectors can return more data columns from the current table schema in `Scan.readSchema()` than requested by `SupportsPushDownRequiredColumns.pruneColumns`. Previously, this partial pruning could fail planning because the scan reported columns absent from the analyzed relation output.
35
35
- Since Spark 4.4, when an already-analyzed Data Source V2 query is refreshed, if rebinding an expanded struct used as a map key causes distinct current keys to collide under the analyzed schema, the query fails with `DUPLICATED_MAP_KEY` by default instead of returning duplicate keys. With `spark.sql.mapKeyDedupPolicy=LAST_WIN`, the last value is kept.
0 commit comments